igneum/docs/analysis/asic-resistance-history.md

80 KiB

ASIC resistance, 2011 to 2026: the history, the papers, the lessons, and the audit of Igneum against them

5 October 2026 (night), branch asic-history. Asked by the project lead at 20:05 UTC: "do a full on deep dive into the full history of 'asic resistance' and see if we can add or upgrade anything." Baseline for the audit: the Counter ASIC 2.0 final class decided tonight (docs/plans/counter-asic-2-status.md on ca2-coord, entries 20:16 to 22:25 UTC; docs/analysis/chip-model-v3.md on ca2-mixer 1ab8b21). Every figure about another chain cites a repo file, a paper or a dated article, or is labelled approximate. Hash-per-joule gains are computed from the cited hashrate and watt figures of the chip and of the best consumer GPU of the same year, and are approximate by construction (GPU figures vary by tuning). Research gathered by four sub-agents between 20:10 and 20:45 UTC; the fetch failures they reported are listed in section 6.

0. One page for the project lead

What the history says Igneum is doing right.

# What The evidence
1 Binding the hash to random reads over a dataset larger than any on-chip cache, with the dataset growing on a schedule, and measuring the latency-bound share per card Every compute-bound hash fell to a chip at 20x to 1,200x per joule within 16 to 37 months (rows Scrypt, X11, Blake, kHeavyHash, Blake3). The memory-bound hashes capped the chip at 1.1x to 4.8x (Ethash rows) or saw no chip at all (KawPow, Verthash, Autolykos, FishHash). RandomX's own design chose a 2 GiB dataset because a 2 GiB SRAM die was "questionable" at 7 nm in 2019 (RandomX doc/design.md)
2 A random program per epoch from a VDF seed, a weak-program filter, era draws from chain state, instruction families unlocked by height, and no scheduled human fork Monero forked four times in 20 months and lost 85% of its hashrate at each fork to chips that returned within months (CryptoNight rows). Vertcoin forked three times and was 51%-attacked after two of them. Ravencoin's X16Rv2 fork was followed by FPGA bitstreams within weeks. Grin's six-monthly tweaks worked only because the lane was scheduled to die. The only random-program hashes with no chip after five years are the ones that never needed a fork (KawPow, FiroPoW, ProgPowZ)
3 Pricing the on-die-cache recompute chip and spending the mixer budget against it (x8: 0.92x with the 3x factor), with the model public This is the "light-evaluation attack" Least Authority flagged on ProgPoW in 2019 and ProgPoW never fixed; Bob Rao's hardware audit put an on-die-DAG ProgPoW chip at "<< 0.1x" the energy per hash of a GPU. Kik's 2020 exploit was the same attack through a 64-bit seed. Igneum has a number against it; ProgPoW had a suggestion

What the history says Igneum is missing or under-weighting.

# What The evidence
1 The partial-store chip with a custom memory system (HBM or many narrow DRAM channels) is not in the chip model. The model prices only the f = 0 endpoint (all SRAM, recompute everything) The only chip class that ever beat a memory-bound GPU hash did it this way: Ethash chips reached 2.1x (Linzhi, 2020), 2.9x (E9, 2022) and 4.8x (Jasminer X4, 2021) per joule through custom memory controllers and on-package memory, with no on-die dataset at all. The time-memory curve between f = 0 and f = 1 is open item O-1.6 and has never been drawn (proto-metal/MEMHARD.md section 3 item 2). Cuckoo Cycle's "tmto-hard" claim fell to a 50x memory cut for 2x time within two months of publication (Andersen, 2014)
2 The item-derivation mixer has a fixed shape. That fixed shape is exactly what hands the recompute chip its 3x fixed-function factor (bare 0.31x becomes 0.92x) RandomX made the item derivation itself a random program (SuperscalarHash: about 450 instructions generated per seed, scheduled for a superscalar core, 170-cycle latency to match DRAM), so a chip cannot hard-wire it. Igneum draws the mixer's constants per day and keeps the shape; a chip hard-wires the shape. CryptoNight-R's random math per block raised chip latency only 2.5x; the lever is in the memory path, which for the recompute chip is the mixer
3 No clock and no detector. Chips have appeared at market caps from $18M (Radiant) and $23M (Grin) upward, and Vorick's 2018 rule was "any coin with over $20M of block reward in a year has a secret ASIC on it". Monero's secret chips held 85% of the hashrate before anyone saw them. Igneum's bounty is unfunded (D11) and the public benchmark is January 2027 Rows Monero, Radiant, Grin, Handshake, Kadena; section 2.5 (the market-cap table); MoneroCrusher's nonce analysis (February 2019) found the chips by their share pattern, which is the detector Igneum can run from day one on the observer

The one change I would make first. Add the partial-store HBM chip to the chip model and draw the time-memory curve before genesis (ranked addition 1, section 4.3). It is an analysis, it costs the GPU nothing, and it is the one chip class that history shows beating memory-bound GPU work. If that row comes out under 2x the x8 decision stands as measured; if it does not, the next lever is known before the vectors are frozen.

Everything else in this document is the evidence behind that page.

1. The history, one row per attempt

Columns: what the hash relied on; when it went live; the first chip that beat it (vendor, model, date, rated hashrate and watts); the gain in hash per joule against the best consumer GPU of the time (approximate, derived); how long it held (months from live to first chip); what the chain did. "None" in the chip column means no shipped chip was found in any source as of October 2026.

1.1 The rows

# Hash (chain) Live Relied on First chip (vendor, model, date, rate, watts) Gain per joule vs GPU (approximate) Held (months) Response Sources
1 Scrypt (Tenebrix, Litecoin, Dogecoin) Sep and Oct 2011 128 KB scratchpad, meant to fit a CPU cache and not a 2011 GPU; latency at SRAM scale Gridseed GC3355, early 2014, 360 kH/s at 7 to 8 W; Innosilicon A2 Terminator, Apr 2014, 28 nm; KnC Titan, 2014, 300 MH/s at 850 W Gridseed 19x, A2 68x, Titan 140x (vs Radeon 7970 at 700 kH/s, 285 W); Antminer L7 (2021) 1,100x 27 Embraced. Dogecoin merge-mined with Litecoin from Sep 2014 [S1] [S2] [S3] [S4]
2 X11 (Dash) and the chain family X13, X15, X17, Quark, Nist5, C11 Jan 2014 A chain of 11 SHA-3 candidates; compute only. Duffield said it was meant to replay Bitcoin's CPU to GPU to ASIC path, not to prevent it iBeLink DM384M, Mar 2016, 384 MH/s at 715 W; Baikal Giant A900 (2016); Antminer D3, Sep 2017, 19.3 GH/s at 1,200 W DM384M 33x, D3 1,000x (vs R9 280X at 4 MH/s, 250 W) 26 Embraced. Baikal's multi-algo units covered the whole family by 2017 [S5] [S6] [S7]
3 Ethash (Ethereum) Jul 2015 DAG of 1 GB growing per epoch, 128-byte random reads; memory bandwidth Antminer E3, announced Apr 2018, 180 MH/s at 800 W, 4 GB DDR3; Innosilicon A10 Pro (2020) 500 MH/s at 950 W; Linzhi Phoenix (Dec 2020) 2,733 MH/s at about 3,000 W; Jasminer X4 (Oct 2021) 2.5 GH/s at 1,200 W; Antminer E9 (Jun 2022) 2.4 GH/s at 1,920 W E3 1.1x to 1.6x (a tuned 1080 Ti beat it per watt); A10 Pro 1.2x vs RTX 3080; Phoenix 2.1x; X4 4.8x; E9 2.9x 32 to the first chip; about 65 to a chip over 2x ProgPoW (EIP-1057) debated 2018 to 2020 and shelved; PoS at the Merge, 15 Sep 2022. ASIC share of hashrate stayed small (one 2018 estimate: 3%) [S8] [S9] [S10] [S11] [S12] [S13]
4 Etchash (Ethereum Classic) Nov 2020 (Thanos, ECIP-1099) Ethash with the DAG cut to 2.5 GB to keep 3 to 4 GB cards mining; not an anti-chip change Post-Merge Ethash chips moved over: Antminer E9 Pro (Feb 2023) 3.68 GH/s at 2,200 W; Jasminer X16-P (2023) 5.8 GH/s at 1,900 W E9 Pro about 4x, X16-P about 7x vs RTX 3090 (approximate) n/a Embraced [S14] [S15]
5 Equihash 200,9 (Zcash, Horizen, Pirate) Oct 2016 Generalised birthday problem (Wagner); 144 MB in practice; memory size, with a claimed 1,000x compute penalty for halving memory Antminer Z9 mini, announced 3 May 2018, 10 kSol/s at 300 W; Innosilicon A9 ZMaster (Jun 2018) 50 kSol/s at 620 W; Z11 (2019) 135 kSol/s at 1,418 W; Z15 (2020) 420 kSol/s at 1,510 W Z9 mini 12x, A9 29x, Z11 34x, Z15 100x (vs GTX 1080 Ti at 700 Sol/s, 250 W) 18 Zcash: no fork (Zcon0 vote 45 to 19 against prioritising resistance, Jun 2018; ECC chose Sapling over resistance); PoS plan announced Nov 2021. Horizen: stayed after a 51% attack (Jun 2018). Pirate: stayed [S16] [S17] [S18] [S19] [S20]
6 Equihash parameter forks: Zhash 144,5 (Bitcoin Gold), ZelHash 125,4 (Flux), BeamHash I to III 150,5 (Beam), 210,9 (Aion), 192,7 (Zero) Jul 2018 (BTG), Jun 2019 (Flux), Jan 2019 (Beam) Larger memory per solver than 200,9 (Beam and Flux also changed the datapath against FPGA bitstreams) None found for 144,5, 125,4 or 150,5. Vorick (May 2018) wrote that an Equihash chip able to follow any parameter fork had been designed n/a BTG 7 years, Flux 7 years, Beam 7 years, with small prizes (section 2.5) BTG: forked after the May 2018 51% attack, attacked again Jan 2020. Beam: planned "one or two hard forks" then BeamHash III (Jun 2020) as the last. Flux: stayed on ZelHash [S21] [S22] [S23] [S24]
7 Scrypt-N, Lyra2RE, Lyra2REv2 (Vertcoin) Jan 2014, Dec 2014, Aug 2015 Memory-hard sponge (Lyra2) inside a hash chain; Scrypt-N grew N over time Dayun Zig Z1, Sep 2018, 6.8 GH/s at 1,200 W (FPGA bitstreams of about 216 MH/s per board preceded it in 2018) Z1 20x (vs GTX 1080 Ti at 59 MH/s, 207 W) 37 (Lyra2REv2) Lyra2REv3, Feb 2019, "to rid the network of the current generation of ASICs and FPGAs"; 51% attacks via rented hash in Oct to Dec 2018 (22 reorgs) and Dec 2019 [S25] [S26] [S27] [S28]
8 Verthash (Vertcoin) Jan 2021 A 1.2 GB file generated from the chain's own block headers; random reads; bandwidth, like Ethash None found n/a 69 and counting, small prize No fork since [S29] [S30]
9 X16R (Ravencoin) Jan 2018 16 hashes in an order set by the previous block hash; compute, with order randomised OW Miner OW1, Sep 2019, about 182 MH/s at 1,400 W; SKC Turing R1 claimed. Widely thought FPGA-based OW1 1.3x (vs GTX 1080 Ti at 18 to 31 MH/s, 190 to 284 W) 20 X16Rv2, 1 Oct 2019 (hashrate fell 70%); FPGA bitstreams for X16Rv2 within weeks (BittWare CVP-13 at 240 MH/s) [S31] [S32] [S33]
10 KawPow (Ravencoin; Neoxa, Clore, Meowcoin, Neurai) 6 May 2020 ProgPoW 0.9.4 variant: random math per block, DAG, 16 KB cache reads; targets the GPU datapath None found as of 2026 (no vendor lists one; one retailer listing naming an "Antminer X9" for KawPow is an error) n/a 77 and counting Roadmap: "No additional future algorithm forks are envisaged" [S34] [S35] [S36]
11 MTP (Zcoin, now Firo) Dec 2018 Argon2d memory array with a Merkle tree (Biryukov and Khovratovich, "Egalitarian computing"); 4 GB; memory size None n/a 34, then replaced Dinur and Nadler broke the 2 GB instance to under 1 MB at a 170x compute penalty before launch (2017); MTP 1.2 patched it. FiroPoW (ProgPoW variant) Oct 2021 for block size and GPU fairness, not for a chip [S37] [S38] [S39] [S40]
12 FiroPoW (Firo) 26 Oct 2021 ProgPoW 0.9.4 with a per-block program None found n/a 59 and counting Nov 2025 fork cut the maximum DAG to 6.76 GB to keep 8 GB cards [S40] [S41]
13 Blake-256 14r (Decred) Feb 2016 Compute; chosen to be ASIC-friendly ("easy and fast implementation of hardware is the main design goal") Innosilicon D9, about Apr 2018, 2.4 TH/s at 1,000 W; Obelisk DCR1 (Jun 2018); Antminer DR5 (Dec 2018) 35 TH/s at 1,610 W D9 130x, DR5 1,200x (vs GTX 1080 Ti at 4.6 GH/s, 250 W) 26 Embraced by design [S42] [S43] [S44]
14 Blake2b (Sia) Jun 2015 Compute Antminer A3, Jan 2018, 815 GH/s at 1,186 W; Obelisk SC1 (Jul 2018) 550 GH/s at 500 W; Innosilicon S11 (2018) 3.83 TH/s at 1,380 W A3 58x, S11 230x (vs GTX 1080 Ti at 2.96 GH/s, 250 W) 31 Fork at block 179,000 (Oct 2018) to brick Bitmain and Innosilicon units and keep Obelisk's; Innosilicon then held about 37% of hashrate [S45] [S46] [S47]
15 Blake2s (Kadena) Nov 2019 Compute; chosen to be "GPU mineable and not immediately ASIC mineable (but for which an ASIC can be made)" Goldshell KD2 and KD5, Mar 2021, 18 TH/s at 2,250 W; Antminer KA3 (Sep 2022) 166 TH/s at 3,154 W KD5 195x, KA3 1,280x (vs RTX 3080 at 8.9 GH/s, 217 W) 16 Embraced by design [S48] [S49] [S50]
16 CryptoNight (Bytecoin, Monero) Jul 2012, Apr 2014 2 MB scratchpad sized to a per-core L3, AES rounds, random reads; latency at SRAM scale Secret chips from about late 2017 (85% of the hashrate vanished at the April 2018 fork); Antminer X3, announced Mar 2018, 220 kH/s at 550 W; Baikal Giant-N X3 40x to 50x (vs Vega 64 at about 2 kH/s, 200 to 250 W, approximate) 43 to the secret chips, 47 to the announced one Forks: CryptoNight v7 (6 Apr 2018), v8 (18 Oct 2018), CryptoNight-R (9 Mar 2019, random math per block seeded by height, chip latency up 2.5x), RandomX (30 Nov 2019). MoneroCrusher's nonce analysis (Feb 2019) found chips at over 85% of the hashrate again, four months after v8 [S51] [S52] [S53] [S54] [S55]
17 RandomX (Monero; Wownero, ArQmA, Zephyr, Tari) 30 Nov 2019 A VM running 8 chained random programs per hash on a superscalar CPU with floating point; 2 GiB dataset derived from a 256 MiB cache by a random superscalar program (SuperscalarHash); 2 MiB scratchpad in L1, L2, L3 tiers Antminer X5, Sep 2023, 212 kH/s at 1,350 W (RISC-V cores); Antminer X9, Jul 2026 delivery, 1 MH/s at 2,472 W; Pinecone INIBOX R1X, Mar 2026, 1.2 MH/s at 2,055 W X5 at parity with a Ryzen 9 7950X (157 against about 200 H/J, approximate); X9 and R1X 2x to 3x over the best CPU (approximate). GPUs are 25x worse per joule than CPUs on it 46 to parity hardware, about 75 to a 2x to 3x chip No fork as of Oct 2026. Four audits in 2019 (Trail of Bits, X41, Kudelski, QuarksLab) found nothing critical [S56] [S57] [S58] [S59] [S60]
18 Cuckoo Cycle (Grin Cuckaroo lane, Aeternity, Cortex) Jan 2019 (Grin) Find a 42-cycle in a random graph; lean solver one bit per edge; memory latency, or bandwidth in the mean solver None on the Cuckaroo lane (Cuckaroo29 tweaked every 6 months: Cuckarood Jul 2019, Cuckaroom Jan 2020, Cuckarooz Jul 2020) n/a 24, retired on schedule The lane was built to die: 90% of reward at launch falling to 0% in Jan 2021 (HF4) [S61] [S62] [S63]
19 Cuckatoo31+ (Grin's chip lane) Jan 2019 Same, with plain bits in place of ternary counters to simplify chips; "Proof of SRAM" per Tromp Obelisk GRN1 announced Jan 2019 and cancelled Jul 2019; Innosilicon G32 announced 2019 and never shipped; iPollo G1, Dec 2020, 36 GPS Cuckatoo32 at 2,800 W G1 about 4x (vs RTX 3090 at about 1 GPS, 300 W, approximate) 23, by design Surrender by schedule. Tromp's $10,000 linear TMTO bounty was claimed in Apr 2025 (N/k bits at about k + 1,000 hashes per edge) [S61] [S64] [S65] [S66]
20 ProgPoW (Ethereum proposal; Bitcoin Interest, Sero, Zano as ProgPowZ, Quai) EIP May 2018; Bitcoin Interest 2018; Quai Jan 2025 Random math per period on a 32-register file, 16 KB cache reads, 256-byte DAG loads, keccak-f800; "saturate the GPU" None n/a 8 years across its adopters Ethereum: tentatively approved Jan 2019 and Feb 2020, petition 27 Feb 2020, left "approved" and unscheduled on 6 Mar 2020, dead. Audits: Least Authority (Sep 2019) and Bob Rao (Sep 2019). Kik's 64-bit-seed exploit (Mar 2020) patched in 0.9.4 [S67] [S68] [S69] [S70] [S71] [S72]
21 Autolykos v1 and v2 (Ergo) Jul 2019; v2 Feb 2021 v1: memory-hard with a per-miner secret key, so puzzles could not be outsourced to pools. v2: the secret removed (contract pools bypassed it); a 2 GB table that grows 5% per 51,200 blocks from block 614,400 None n/a 87 and counting, small prize v2 by EIP-0009 at block 417,792 [S73] [S74]
22 Octopus (Conflux) Oct 2020 Ethash-style DAG; the "dense matrix step" could not be verified in conflux-rust tonight (unverified) None n/a 72 and counting CIP-102 (Aug 2022) proposed switching to Ethash to attract post-Merge miners; dormant [S75] [S76]
23 kHeavyHash (Kaspa; Bugna kept it) Nov 2021 cSHAKE256, a 64x64 4-bit matrix multiply from the pre-PoW hash, cSHAKE256; compute, designed for optical and specialised hardware IceRiver KS0, Jul 2023, 100 GH/s at 65 W; KS1, KS2 (Sep 2023); Antminer KS3, Aug 2023, 8.3 TH/s at 3,188 W; KS5 Pro (Mar 2024) 21 TH/s at 3,150 W KS0 250x, KS5 Pro 1,100x (vs RTX 3090 at 910 MH/s, 150 W) 17 to 20 Embraced (Sompolinsky, May 2023: "an overall positive"). Hashrate went from under 100 PH/s to over 700 PH/s in months; the GPU share was negligible by late 2023 (approximate). Forks that left: Karlsen (FishHashPlus, Sep 2024), Pyrin (PyrinHash v2, Sep 2024), Spectre (CPU AstroBWTv3), Nexellia, Waglayla, Cryptix, Hoosat [S77] [S78] [S79] [S80] [S81]
24 NexaPow (Nexa) 2023 SHA-256 plus a secp256k1 Schnorr signature per attempt; framed as "useful ASICs" DragonBall A21, Jan 2025, 3.4 GH/s at 1,800 W 3x to 4x (vs RTX 3090 at 123 to 137 MH/s, 230 W, approximate) 24 None [S82] [S83]
25 Blake3 (Alephium; Iron Fish until 2024) Nov 2021 Double Blake3; compute; chosen as ASIC-friendly Goldshell AL-BOX, 2023, 360 GH/s at 180 W; IceRiver AL0; Antminer AL1 (2024) 15.6 TH/s at 3,510 W; AL3 AL-BOX 156x, AL1 350x (vs RTX 3090 at 2.3 GH/s, 180 W) 22 to 24 Alephium embraced. Iron Fish forked to FishHash (Apr 2024, FIP-3: Ethash-derived, fixed 4.6 GB dataset, 512 iterations, 128-byte mix); Karlsen adopted FishHashPlus (Sep 2024) [S84] [S85] [S86] [S87]
26 FishHash (Iron Fish, Karlsen) Apr 2024 Ethash-derived, 4.6 GB fixed dataset; bandwidth None found n/a 30 and counting, small prize None [S87]
27 Eaglesong (Nervos) Nov 2019 Compute (a new sponge) Toddminer C1 (Feb 2020); Antminer K5, Mar 2020, 1.13 TH/s at 1,580 W; Goldshell CK5 (Mar 2021) 12 TH/s at 2,400 W K5 70x, CK5 500x (vs RTX 3090 at about 2.1 GH/s, approximate) 4 Embraced [S88] [S89]
28 Blake2b + SHA3 (Handshake) Feb 2020 Compute Goldshell HS1, Jun 2020; HS3 (Jul 2020) 2 TH/s at 2,000 W; HS5 Over 100x (approximate) 5 Embraced [S90] [S91]
29 SHA512/256d (Radiant) 2022 Compute DragonBall A11; IceRiver RX0, Sep 2024, 260 GH/s at 100 W About 550x (vs RTX 3090 at 1.3 to 1.5 GH/s, approximate) About 24 None [S92] [S93]
30 ProgPowZ (Zano), DynexSolve (Dynex), Janushash (Warthog), XelisHash v1 and v2 (Xelis), VerusHash 2.2 (Verus) 2019 to 2024 ProgPoW variant; GPU "neuromorphic" useful work; a product of VerusHash and SHA256t to balance CPU and GPU; CPU and GPU balanced; AES-based CPU hash None found for any of them n/a Small prizes throughout Xelis forked to v2 (Jul 2024) for FPGA resistance [S94] [S95] [S96] [S97] [S98]
31 Ethash on EthereumPoW (ETHW) after the Merge Sep 2022 As Ethash The Ethash chips above About 4x (E9 Pro, X16-P vs RTX 3090, approximate) n/a Embraced [S15]

1.2 What the rows say when sorted

Class of hash Rows Months to first chip First-chip gain per joule Best gain reached
Compute only (chains of hashes, Blake family, SHA-3 family, matrix multiply) 2, 13, 14, 15, 23, 25, 27, 28, 29 4 to 31 (median about 24) 33x to 250x 500x to 1,280x
Memory at SRAM scale (128 KB scrypt, 2 MB CryptoNight) 1, 16 27, 43 19x, 40x to 50x 1,100x (Scrypt, 2021)
Memory size without a bandwidth bound (Equihash, Lyra2REv2) 5, 7 18, 37 12x, 20x 100x
Memory bandwidth at DRAM scale (Ethash, Verthash, FishHash, Etchash) 3, 4, 8, 26 32 (Ethash); none for the others 1.1x to 1.6x 2.9x to 4.8x
Random program on a commodity datapath (RandomX, ProgPoW family, X16R's order randomisation) 9, 10, 12, 17, 20, 30 X16R 20 (FPGA-class, 1.3x); RandomX 46 to parity; none for ProgPoW's adopters in 8 years 1.3x (X16R), 1x (RandomX 2023) 2x to 3x (RandomX 2026, approximate)
Graph search (Cuckoo) 18, 19 23 on the chip lane; never on the tweaked lane 4x 4x

Two caveats on the random-program rows. The prizes were small: Ravencoin, Firo and Zano never reached the market caps at which the 2018 chips appeared (section 2.5), so "no chip" is partly an economic fact. And RandomX's chips arrived once Monero's reward justified them: parity hardware at 46 months, a 2x to 3x chip at about 75 months (approximate), on a hash whose whole purpose was to make the CPU the chip.

2. The academic side

2.1 Memory-hard functions

Paper Result What it means for Igneum
Abadi, Burrows, Manasse, Wobber, "Moderately hard, memory-bound functions", NDSS 2003 and ACM TOIT 2005 [P1]; Dwork, Goldberg, Naor, "On memory-bound functions for fighting spam", CRYPTO 2003 [P2] The origin of the idea: CPU speed varies 100x across machines, memory latency does not, so a cost function bound by cache misses is fairer than one bound by cycles Igneum's latency-bound rule is this argument from 2003 applied to GPUs and DRAM: the DRAM row cycle is the same physics for a chip and a card (section 2.6)
Percival, "Stronger key derivation via sequential memory-hard functions", BSDCan 2009 [P3] Defines sequential memory-hardness; ROMix is sequential memory-hard in the random-oracle model; cost measured in area-time (dollar-seconds) The area-time measure is the one the chip model uses (equal silicon); scrypt's 2011 deployment at 128 KB ignored the paper's own scale
Alwen and Serbinenko, "High parallel complexity graphs and memory-hard functions", STOC 2015 [P4] Cumulative memory complexity (CMC) in the parallel random-oracle model; earlier sequential measures fail against parallel, amortising adversaries A chip is a parallel, amortising adversary; any Igneum claim about the dataset must be made in a parallel model
Alwen and Blocki, "Efficiently computing data-independent memory-hard functions", CRYPTO 2016 [P5]; "Towards practical attacks on Argon2i and Balloon hashing", EuroS&P 2017 [P6] Any data-independent MHF can be computed in less than n^2 cumulative memory; Argon2i at O(n^1.75 log n), Catena and Balloon at O(n^1.67); the attacks are practical at real parameters Igneum's addresses are data-dependent (register state), which is the right side of this result; the price is cache-timing leakage, which does not matter for a PoW
Alwen, Chen, Pietrzak, Reyzin, Tessaro, "Scrypt is maximally memory-hard", EUROCRYPT 2017 [P7] scrypt's CMC is Omega(n^2 w) in the parallel ROM, optimal, against parallel amortising adversaries Data-dependent chains of reads are the construction with the proof; Igneum's item derivation (8 dependent cache reads) is a short chain of this kind, with no proof
Biryukov, Dinu, Khovratovich, "Argon2", EuroS&P 2016 [P8]; Boneh, Corrigan-Gibbs, Schechter, "Balloon hashing", ASIACRYPT 2016 [P9] Argon2d: a one-pass adversary can cut memory at most 3x at equal area-time; Argon2i needs over 10 passes to resist the Alwen-Blocki attack. Balloon: provable in the sequential model only; the paper says parallel ASIC attacks are outside its model A "memory-hard" label without a stated adversary model has been wrong three times in this list (Argon2i, Balloon, Catena)
Biryukov and Khovratovich, "Tradeoff cryptanalysis of memory-hard functions", ASIACRYPT 2015 [P10]; Forler, Lucks, Wenzel, "Catena", 2013 [P11]; Simplicio et al., "Lyra2", IEEE TC 2016 [P12] The ranking trade-off attack on Lyra2, yescrypt and Argon2; Catena's proofs flawed, 25x area-time cut; designers changed their algorithms Lyra2REv2 (row 7) carried this construction into a PoW and still fell to a chip at 20x; the cryptanalysis found the shortcut before the chip did

2.2 Bandwidth-hard functions

Paper Result What it means for Igneum
Ren and Devadas, "Bandwidth hard functions for ASIC resistance", TCC 2017 [P13] Memory-hardness (CMC) bounds a chip's area advantage and says nothing about energy; energy spent on off-chip memory traffic is comparable for a chip and a CPU, so bandwidth-hardness is the lever; scrypt, Catena-BRG and Balloon are bandwidth-hard with suitable parameters; the stacked double butterfly is capacity-hard and not bandwidth-hard The chip model's "equal silicon" row is an area argument. The energy argument is the one the Ethash chips answered: they moved the same bytes at lower energy per byte with custom memory controllers (rows 3 and 4). Igneum's hash moves 128 x 64 B = 8 KB of DRAM lines per hash on AMD and 128 x 32 B on NVIDIA; a chip with 4-byte access granularity moves 512 B for the same work. That is the bandwidth-per-watt gain the plan's last section warns about, stated in Ren-Devadas's units
Blocki, Ren, Zhou, "Bandwidth-hard functions: reductions and lower bounds", CCS 2018 [P14] Bandwidth cost in the parallel ROM equals the red-blue pebbling cost of the graph; high CMC implies high bandwidth cost; Argon2i and DRSample are maximally bandwidth-hard; a tight lower bound on scrypt's energy The right formal target for a future proof about the item derivation, if one is ever attempted; none exists today
Alwen, Blocki, Harsha, "Practical graphs for optimal side-channel resistant MHFs", CCS 2017 [P15] DRSample: a practical graph with maximal depth-robustness Not applicable: Igneum does not need side-channel resistance

2.3 Asymmetric, egalitarian and graph proofs of work

Paper Result What it means for Igneum
Biryukov and Khovratovich, "Equihash", NDSS 2016 [P16] Wagner's generalised birthday with algorithm binding; claimed 1,000x compute for halving memory The claim did not survive contact with a chip design: 144 MB in practice fitted the Z9's memory system (row 5)
Biryukov and Khovratovich, "Egalitarian computing", USENIX Security 2016 [P17]; Dinur and Nadler, "Time-memory tradeoff attacks on the MTP proof-of-work scheme", CRYPTO 2017 [P18] MTP: Argon2d plus a Merkle tree. Dinur-Nadler: malicious proofs with under 1 MB in place of 2 GB at a 170x compute penalty, by injecting blocks that steer Argon2d's data-dependent addressing The attacker who controls the memory's contents controls the addresses. In Igneum the day key comes from a VDF of chain state and the cache fill is a chained block function, so no miner chooses the contents. The analogy still holds for the unreviewed mixer: a structural weakness in M_r is the shortcut this paper found in MTP
Tromp, "Cuckoo Cycle", BITCOIN 2015 [P19]; Andersen, "A public review of Cuckoo Cycle", 31 Mar 2014, and "Exploiting time-memory tradeoffs in Cuckoo Cycle", 1 Aug 2014 [P20]; the linear TMTO bounty, claimed Apr 2025 [S66] Edge trimming cut memory about 50x for about 2x time, two months after publication; Tromp adopted it. The 2025 bounty result: an N/k-bit chip must hash each edge about k + 1,000 times A time-memory claim is a curve, and the curve was wrong by 50x until someone drew it. Igneum's curve between "store everything" and "recompute everything" has not been drawn (O-1.6)
Georghiades, Flolid, Vishwanath, "HashCore", 2019 [P21] "Inverted benchmarking": random widgets modelled on SPEC CPU workloads so the CPU is already the chip The same idea as RandomX and ProgPoW stated generally: the hash is a benchmark of the target hardware

2.4 Program-based proofs of work and their audits

Document What it says What it means for Igneum
RandomX doc/design.md and doc/specs.md (tevador) [S56] [S57] A VM so that the work is "data and code"; 8 chained programs per hash so a miner cannot filter (filtering 25% of programs at a 50% speedup yields 0.44x honest speed); SuperscalarHash of about 450 instructions with 155 multiplies, scheduled for a superscalar core at a 170-cycle latency to match DRAM, so a light-mode chip with the 256 MiB cache on die pays 760 cycles and 1,240 multiplies per item, "energy comparable to loading 64 bytes from DRAM"; a 2,080 MiB dataset because a 2 GiB SRAM die was "questionable" at 7 nm in 2019; cache-to-dataset ratio capped at 8 to keep the area-time product constant; double-precision floating point to force the whole CPU; "DRAM cannot do more than about 25 million random accesses per second per bank group" Igneum rebuilt the idea for a GPU. The parts that carried over: the dataset above on-chip cache, the dependent item derivation, the cache-to-dataset ratio (4 at genesis, 8 at year 4 under option C). The parts that did not: per-hash programs (a GPU cannot JIT per hash and stay a GPU), floating point (vendor rounding), a random item-derivation program (Igneum's mixer has a fixed shape with drawn constants). Section 4.3 ranks the last of these
Trail of Bits audit of RandomX, 2 Jul 2019 [S60]; Kudelski, X41, QuarksLab (2019) Two low findings and 47 brittle parameters; the design affirmed Four paid external reviews before launch, for a hash whose whole value was the resistance claim. Igneum has had none (ledger M7)
EIP-1057 ProgPoW [S67]; Least Authority audit, 9 Sep 2019 [S69]; Bob Rao hardware audit, Sep 2019 [S70] Claimed chip gain 1.1x to 1.2x. Least Authority: no issues, five suggestions, one of them the light-evaluation attack (on-the-fly DAG generation with the 16 MB cache in on-die SRAM) "may become possible within a few years" once about 100 MB of fast on-die SRAM is feasible. Rao: energy per hash is the only meaningful metric; shipping Ethash chips show about 1.6x hashrate per watt; conventional compute chips gain little on ProgPoW; integrating the DAG on die cuts data-movement energy by over 10x, so "ProgPOW ASICs with << 0.1X E/H over GPUs can be built"; an advanced-node chip is "$20M+" and "1+ year"; a 16-die split holding a 2.78 GB DAG was about $172 per board in 2019 against about $240 for a GPU board, and a monolithic die "viable around 2025" The light-evaluation attack is Igneum's M16 recompute chip. ProgPoW left it as a suggestion; Igneum priced it and spent the mixer against it (x8). Rao's "$172 per board for a 16-die split" is the HBM-class partial-store chip in a different form, and it is the row the Igneum model lacks
Kik, "ProgPoW exploit", 4 Mar 2020 [S71] A 64-bit seed lets a chip skip memory access with a cooperating node; patched in 0.9.4 Igneum's seed is 256 bits and the program is the epoch's; the nearest analogue is header grinding for cache locality, unmeasured (section 4.3, check 4)

2.5 The economics of a chip

What a chip costs to make (design plus masks, by node; all figures from the cited articles, which disagree with each other by 2x and say so):

Node Mask set Full design, IBS as quoted by Semiengineering (2018 and 2021) Full design, other estimates Sources
65 nm MPW shuttle n/a n/a Europractice 2025: about €51,000 minimum (9 mm^2 at €5,720 per mm^2) [E1]
28 nm "beyond $1M" (SemiAnalysis 2022); $1M to $3M (Silicon Analysts 2026) $51.3M (2018); $40M (2021) $5M to $30M total NRE for a small chip (Silicon Analysts) [E2] [E3] [E4]
16/12 nm n/a $106M (2018 revision of a 2014 $310M figure) Europractice MPW 16 nm: about €125,000 minimum [E1] [E4]
7 nm "beyond $10M" (SemiAnalysis); $5M to $10M (Silicon Analysts) $297.8M (2018); $217M "mainstream" (2021); Semiengineering's own 2023 discount: about $160M Startups shipped 7 nm chips for "$50M to $75M" all-in (SemiAnalysis); a 10 nm-class mining chip "$20M+" (Rao 2019) [E2] [E3] [E4] [S70]
5 nm $10M to $20M $542.2M (2018); $416M (2021); about $280M discounted (2023) Marvell 2023: $449M (secondary source, approximate) [E2] [E3] [E5]
3 nm "$40M range" $500M to $1.5B (2018); $590M (2021) Marvell 2023: $581M (secondary, approximate) [E2] [E3] [E5]

Miner-makers' own numbers: Bitmain's 2018 filing shows R&D of $73M in 2017 and $86M in the first half of 2018 and three failed chips at a reported combined cost of about $500M [E6]; Canaan's 2019 prospectus shows R&D of $26.5M in 2018 and "seven tape-outs" at a 100% success rate [E7]; Vorick wrote that Bitmain brought the Sia A3 to market for "less than $10 million" and took over $20M of orders within eight minutes [S47]; Obelisk's DCR1 was a 28 nm part [S43]; Taylor's 2013 survey gives $150,000 for a 130 nm and $500,000 for a 65 nm Bitcoin chip NRE in 2012 [E8].

Where the 256 MiB cache lands a chip. The Counter ASIC 2.0 analysis priced the 256 MiB SRAM mirror at 128 mm^2 and $46 per good die at N5 on the shipped-product density (sram-mirror.md revision 2, from AMD V-Cache 64 MB on 41 mm^2 at N7 [E9], TSMC N5 HD macro 31.8 Mib/mm^2 [E10]). The history adds the node question: a cheap chip is a 28 nm chip ($1M to $3M of masks, a $5M to $30M project), and a 28 nm bit cell is about 6x an N7 cell (approximate, from memory: TSMC 28 nm HD about 0.127 um^2 against N7's 0.027 [E10]), so 256 MiB at 28 nm is about 1,000 mm^2 of SRAM on the V-Cache density: more than a reticle. The cache forces the recompute chip onto a 7 nm or better node, which moves its project from the $5M class to the $50M class (SemiAnalysis's 7 nm startup figure). That is a stronger statement than the $46 per die, and it is the reason the cache size matters more than its per-die cost. Option C (the cache doubles with the dataset) keeps it true as nodes shrink: at the 6% per year density trend the status file cites, a 512 MiB mirror in year 4 costs more mm^2 than 256 MiB today.

When chips appeared (CoinMarketCap historical snapshots pulled by the research agent; daily issuance is arithmetic from each chain's schedule; all approximate):

Chain First public chip Market cap then Daily issuance then (USD)
Litecoin Gridseed, Dec 2013 $817M $1.0M
Dash PinIdea DR-100, Aug 2017 (iBeLink 2016 widely cited, date unverified) $2.2B $0.6M
Siacoin Obelisk SC1 announced Jun 2017; Antminer A3 Jan 2018 $430M; $1.5B $0.43M; $1.1M
Decred Obelisk DCR1 Jun 2017; Innosilicon D9 Apr 2018 $216M; $353M $0.18M
Monero Antminer X3, Mar 2018 (secret chips from early 2017 per Vorick, unverified) $3.3B $0.75M
Ethereum Antminer E3, Apr 2018 $37.4B $7.6M
Zcash Z9 mini, May 2018 $1.1B $2.1M
Bitcoin Gold the same chips, May 2018 $1.3B $0.14M
Grin GRN1 announced Jan 2019 (cancelled); G32 Apr 2019 (never shipped); iPollo G1 Dec 2020 $23M (Apr 2019); $23M (Dec 2020) $0.24M; $33K
Nervos Toddminer C1, Feb 2020 $75M $65K to $85K
Handshake Goldshell HS1, Jun 2020 $30M $31K
Kadena Goldshell KD5, Mar 2021 $42M $21K
Kaspa IceRiver KS0, Jul 2023 $480M $0.42M
Alephium Goldshell AL-BOX, May 2024 $179M $0.1M
Radiant IceRiver RX0, Sep 2024 $18M $22K

Sources: [E11] (the research agent's CoinMarketCap pulls, dates in the table) and the chip rows above. Reading: a compute-bound hash gets a chip at $20K to $30K of daily issuance (Radiant, Kadena, Handshake); the 2018 cluster sat at $0.15M to $2M a day. Vorick's rule from May 2018: any coin with over $20M of block reward in a year (about $55K a day) should assume a secret chip [S47]. For Igneum the clock is the day its issuance in dollars crosses about $50K; a memory-bound hash buys time against that clock (Ethash: 32 months at the largest prize in the table), and the random program buys more (section 1.2), but nothing in the table says it buys forever.

2.6 Latency as the resource

Source What it says What it means for Igneum
Li, Reddy, Jacob, "A performance and power comparison of modern high-speed DRAM architectures", MEMSYS 2018 [L1] Row timings from datasheets: DDR4 tRCD 14, tRAS 33, tRP 14 ns; GDDR5 tRCD 14, tRAS 28, tRP 12; HBM and HBM2 tRCD 14, tRAS 34, tRP 14. Row cycle tRC about 40 ns (GDDR5) to 48 ns (DDR4, HBM2). "The memory-latency problem does still remain" The row cycle is the floor under every dependent random read whatever the controller; HBM does not shorten it. What HBM and a custom controller change is the number of rows that can be opened per second per watt (channels, banks, pseudo-channels), which is the throughput of random reads in flight, which is what the 9070 XT probe measured as the card's ceiling (2.4 G reads/s against the 5090's 17.5 G)
Chang, CMU thesis, Dec 2017 [L2] Over two decades DRAM capacity improved 128x, bandwidth 20x, latency 1.3x The latency-bound rule has a long half-life; the bandwidth-per-watt lever (the Ethash chips) does not stand still
NVIDIA profiling guide and the Ampere tuning deck [L3] L1 and L2 lines are 128 bytes in four 32-byte sectors; DRAM-to-L2 transactions default to 64 bytes since Volta, configurable 32, 64 or 128 on A100 The 5090's 32-byte sector per 4-byte read is where a custom controller gains bandwidth efficiency (8x fewer bytes), the Ren-Devadas energy lever; the 9070 XT's 64-byte line is 16x. Neither changes the row cycle
RandomX doc/design.md [S56] About 25 million random accesses per second per DRAM bank group; "all Dataset accesses read one CPU cache line (64 bytes) and are fully prefetched"; one program iteration tuned to "typical DRAM access latency (50-100 ns)" The same arithmetic Igneum uses (128 dependent reads per hash against the card's random-read ceiling), from the design that held longest
Condrey, "PoSME", arXiv Apr 2026 (single author, not peer reviewed) [L4] Latency-bound pointer chasing with hash compute under 3.5% of the step cost; GPUs 14x to 19x slower than a consumer CPU The only dedicated latency-bound PoW paper found; its GPU-vs-CPU gap is the cost RandomX pays, and the cost Igneum avoids by keeping thousands of loads in flight per card

The paper that does not exist: nothing found treats cache timing as a feature; Catena and Argon2i treat it as a leak. No peer-reviewed survey of ASIC resistance as such surfaced; the nearest are Cho's 2018 multi-hash evaluation [P22] (X11-style resistance "is not strong enough"), Feng and Luo's 2020 three-processor study [P23] (GPUs dominate CryptoNight, Ethash and Cuckoo on CPU, GPU and Xeon Phi) and Yaish and Zohar's 2023 pricing of mining hardware as a bundle of options [P24].

3. The lessons

Each lesson is stated once, with the rows it comes from.

  1. Compute-bound work loses by 30x to 1,000x within two years, whatever its shape. Chains of eleven hashes (row 2), sixteen hashes in a random order (row 9), a 64x64 matrix multiply (row 23), a new sponge (row 27), a signature per attempt (row 24): every one got a chip, the random-order chain at 1.3x by an FPGA within 20 months and the rest at 33x to 1,100x. Multiplying the number of fixed functions multiplies the chip's die, not its difficulty. For Igneum: nothing in the program's ALU work is a defence and the design already says so (ledger M1); the defence is the memory path.

  2. Memory at SRAM scale is compute-bound with extra steps. Scrypt's 128 KB (row 1) and CryptoNight's 2 MB (row 16) were sized to a 2011 and a 2014 CPU cache; a chip put the same memory on die and won 19x and 40x. Igneum's answer is the 256 MiB cache growing with the dataset (option C) and the 1 GiB to 2 GiB dataset; section 2.5 shows the cache size also sets the chip's node and therefore its project cost. The hot table (layer 5) was a step back toward SRAM scale, and the measurement agreed (the honest card paid 7% to 16%, the chip paid $0.23 per MB).

  3. Bandwidth-bound work gets a memory chip at 2x to 5x. Ethash held 32 months and then got chips whose whole design was the memory system: DDR3 (E3, no gain), GDDR6 (A10 Pro, 1.2x), custom controllers (Linzhi, 2.1x), on-package memory (Jasminer X4, 4.8x) (rows 3, 4). Rao's audit explains why in energy terms: the chip moves the same bytes at lower energy per byte, and a split-die design holding the DAG was already cheaper than a GPU board in 2019. Ren and Devadas give the bound: a chip's energy advantage on a bandwidth-hard function is the ratio of its memory energy per bit to the GPU's. Igneum's rule "avoid leaning on bandwidth" is right; its model has no row for this chip (section 4.3, addition 1).

  4. Latency-bound and random-program work held longest, and the prize was usually small. CryptoNight's latency bound at SRAM scale held 43 months, then fell to secret chips (row 16). RandomX's at DRAM scale held 46 months to parity hardware and about 75 to a 2x to 3x chip (row 17, approximate), on the largest prize any resistant hash has carried. ProgPoW's adopters have had no chip in 8 years on small prizes (rows 10, 12, 20). Verthash, Autolykos, Octopus and FishHash have none on small prizes (rows 8, 21, 22, 26). The honest reading: the random program on a DRAM-latency bound is the strongest construction the history has, and nobody has tested it at Ethereum's prize.

  5. Periodic human forks fail as a defence. Monero: four forks in 20 months; chips were back at 85% of the hashrate within four months of the v8 fork (row 16), and Vorick wrote that a chip able to survive forks at under a 5x hit had been designed. Vertcoin: three forks, two followed by rented-hash 51% attacks within weeks, because each fork reset the hashrate to a rentable size (row 7). Ravencoin: FPGA bitstreams for the new order within weeks (row 9). Sia: the fork bricked competitors' chips and left one vendor at 37% (row 14). Grin: the tweaks worked because the lane was scheduled to die (row 18). The chip's design cycle is 5 months for Bitmain (Vorick) and 13 for a startup; a fork every 6 months is a race the chip wins on the second lap, and each fork is a governance event. Igneum's draws are automatic and scheduled at genesis; that is the right side of this lesson, and section 4.3 asks whether the epoch can also be shorter than a bitstream (addition 5).

  6. RandomX got the target right and the derivation right, and it costs GPUs 25x. Right: the work is "code and data" so a fixed circuit cannot serve it; chained programs defeat filtering (0.44x); the dataset is above any SRAM die; the item derivation is a random superscalar program tuned to DRAM latency so the light-mode chip pays as much energy per item as a DRAM read (section 2.4). The cost: a GPU runs the VM at 25x worse per joule than a CPU (row 17), which is the cost Igneum refuses, and the reason the program is per hour and compiled. What Igneum did not take: the random item derivation (addition 2) and four external audits before launch (addition 3).

  7. ProgPoW got the datapath right and lost on governance and one unpriced attack. Right: target the commodity hardware's whole datapath (random math, register file, cache reads, DAG loads) so a chip has to be a GPU; the hardware audit agreed for compute-only chips (1.1x to 1.2x, Rao). Unpriced: the DAG on die (Least Authority suggestion 2, Rao's "<< 0.1x"), the same attack Igneum calls M16. Not adopted: two tentative approvals, a petition, bugs found late (Kik), authorship disputes and a PoS roadmap; the change needed a contentious fork on a live chain (row 20). Igneum's lesson is the one it already follows: every layer goes in before the public testnet as a genesis rule or a reserve, so no adoption vote is ever needed.

  8. What a "GPU-friendly" chain lost when its GPU miner fell behind: the miners, then the chain's shape. Kaspa's hashrate rose 7x in months and its GPU share went to nothing; seven forks left to re-resist (row 23). Alephium, Nervos, Handshake, Kadena and Radiant went the same way without the forks (rows 25, 27, 28, 15, 29). Iron Fish forked away from its own Blake3 within a year of the first box (row 25). The chains kept their security budget and lost the fleet that had launched them; the fleet's hardware went to the next GPU chain. For Igneum the metric is the share of hashrate on consumer cards by model, which is what the January 2027 benchmark should report and what the observer can estimate earlier (addition 4).

  9. An unreviewed memory-hard construction has a shortcut until someone looks. MTP fell from 2 GB to under 1 MB before launch (row 11); Catena's proofs were flawed; Argon2i's parameters were attackable at the IRTF's "paranoid" setting; Cuckoo's memory claim was off by 50x within two months (section 2.3). Igneum's M_r and chained cache have had no cryptanalysis (MEMHARD.md section 3, ledger M7); the acceptance rule is a statistical filter, not a proof. The x8 decision multiplies the mixer's weight in the chip model, which multiplies the cost of a structural weakness in it (addition 3).

  10. The secret chip is found by its share, and it is on the chain before the announcement. Monero's chips held 85% before anyone saw them; a nonce-pattern analysis found them (row 16). Zcash's Z9 was "5x to 10x below" what Obelisk's own study said the hash allowed, which is Vorick's evidence that better secret chips existed (section 1.1 row 5). A detector costs an observer query; a bounty costs escrow (addition 4).

4. The audit of Igneum against the history

4.1 The first-generation layers (live in class v2 and carried into v3)

Layer Answers which failure Does not answer Evidence
Random program per epoch from a VDF seed, 12 integer families, nonce-dependent select Fixed-function chips (lesson 1); program filtering and seed grinding (RandomX's 0.44x, spec 04's 130-to-1) A "GPU without graphics": a programmable sequencer over 12 ops and 8 registers (ledger M1); an FPGA overlay or bitstream compiled within the hour (rows 7, 9: FPGAs were the first adversary of Lyra2REv2 and X16R); the 7.5x AMD gap is a one-vendor fleet Rows 9, 10, 16, 17, 20
Weak-program acceptance, exact 16 loads, fresh-source rule Per-program hash-rate spread (1.10x residual) that a chip could pick Nothing it claims to; a chip's advantage cannot come from the program (status 20:16) Census [I1]
1 GiB to 2 GiB dataset of 4-byte random reads, latency-bound, cache 256 MiB SRAM-scale memory (lesson 2); the bandwidth lever at the honest card (lesson 3: 128 x 4 B keeps the 5090 at 9% of its stream bandwidth) The partial-store chip with a custom memory system (lesson 3); the time-memory curve (O-1.6) Rows 1, 3, 16; [P13]
8 dependent cache reads per item, fixed-shape mixer with drawn constants The on-die recompute chip (Least Authority's light-evaluation attack), priced at 2.45x bare under v2 The fixed shape gives the chip its 3x factor (lesson 6); no cryptanalysis (lesson 9) [S69] [S70]; M16
Era draws from chain state (op weights, fold rotations), reserve families by height, dataset growth, no human release Fork fatigue and fork-reset attacks (lesson 5); the chip that "survives forks at under 5x" (Vorick) is the chip that the draws are meant to outlast The draws touch the program, not the item derivation, so they cost the recompute chip nothing (chip-model-v3.md section 2) Rows 7, 16, 18
Warp-unit CPU verification without the dataset (2.1 ms per warp under x8) Keeps the verifier light, the Equihash and Cuckoo goal Caps every lever: the mixer budget stops at the 10 ms gate [P16] [P19]

4.2 The Counter ASIC 2.0 layers as decided tonight

Layer Decision (status file) Answers Does not answer History's verdict
1 Load width 4, 16, 64 B Keep 4 B (w16 closes nothing) Keeps the 5090 latency-bound (9% of stream) The AMD 7.5x gap (2.4 G reads/s at every width) Right by lesson 3; the vendor gap is a 3.0 question and a soft form of lesson 8
2 Per-program width mix Out (spread over 5% on every card) n/a n/a Right: a per-program spread is what a chip picks (lesson 1's X16R: randomised order gave 1.3x, the shape still fixed)
3 Per-warp scratch with RMW Out (does not move the recompute chip; 2.4x at every share; costs GPUs 12% to 48%) n/a n/a Right: SRAM-tier work favours the chip (lesson 2; Rao: SRAM is the chip's weapon)
4 + 8 Era layout (stride, interleave) and per-site windows In (era inside the class) A hard-wired layout tuned to one era A programmable address decoder (era-layout.md section 8 says so); costs the recompute chip nothing Small by itself; its value is in lesson 5 (automatic change without a fork)
5 Hot table sized to GPU cache Measured, not adopted (honest card pays g = 0.84 to 0.93; chip pays SRAM) n/a n/a Right by lesson 2
6 Cache growth Option C: doubles with the dataset (256 MiB, 512 MiB year 4, 1 GiB year 12) Keeps the mirror on a leading node (section 2.5) n/a Right; RandomX's cache-to-dataset ratio of 8 is reached at year 4
7 INT8 matrix family Reserve R1 = mm8, W_new 4, unlock era 4 or 90% signal A family that a 12-op chip lacks Matrix hardware is the most abundant custom silicon on earth; Least Authority's suggestion 5 was "watch ML hardware"; Apple's emulation costs 1.6x to 4.7x per op Keep in reserve, order it last (addition 6)
9 Epoch length as an era parameter (10 min to 2 h) Reserve only, design on ca2-epoch The bitstream-per-epoch FPGA (rows 7, 9) The FPGA overlay (a soft GPU) and the HBM FPGA Rank it up (addition 5)
Mixer x8 (M16's lever) In: 0.31x bare, 0.92x with the 3x factor, verifier 2.1 ms per warp, daily build 23 to 77 ms The on-die recompute chip (lesson 6, Least Authority's attack) Its own fixed shape (the 3x factor stays) and its lack of review (lesson 9) The right lever; additions 2 and 3 are what the history says to do to it next

4.3 Ranked additions and upgrades

Ranked by how much the history says each would change the outcome, with the cost to GPUs and the risk. "Genesis" means a rule fixed before the public testnet; "reserve" means a named family or parameter in the genesis reserve, unlockable by height or 90% signal; "nowhere" means do not add.

Rank Addition What it does Evidence Cost to GPUs Risk Where
1 Price the partial-store chip and draw the time-memory curve. A chip that stores a fraction f of the dataset in HBM or on many narrow DRAM channels, recomputes the rest from a 256 MiB on-die cache under x8, and reads with 4-byte granularity. Rows for f = 0.25, 0.5, 1 at HBM3 and at GDDR7 random-read rates, priced in energy per hash (Rao's metric) and in reads in flight per watt The only chip class that beat a memory-bound GPU hash: Ethash's 2.1x to 4.8x came from the memory system with no on-die dataset (rows 3, 4); Rao priced a 16-die DAG holder under a GPU board in 2019; Cuckoo's curve was wrong by 50x until drawn [P20]; O-1.6 is open and MEMHARD.md section 3 item 2 says the curve was never drawn None (analysis) The row may come out over 2x, which would qualify the public claim before anyone else does Genesis (before the vectors freeze)
2 A random item-derivation program per day in place of the fixed-shape mixer: a SuperscalarHash-style generator, integer only, drawn from the day key, with its own acceptance test, compiled once a day by miners and verifiers RandomX's reason for SuperscalarHash: a fixed derivation is hard-wired by a chip; a random one makes the light-mode chip a CPU (section 2.4). In Igneum's model the fixed shape is the 3x factor that turns 0.31x into 0.92x; removing the factor is worth more than x8 to x16 would be (x16: 0.46x with the factor by M16's table) None per hash (the daily build is 23 to 77 ms at x8 and would roughly double); the verifier needs a per-day compiled derivation (a JIT, or a round schedule drawn from a fixed set of reviewed rounds), measured against the 10 ms gate Cryptanalysis of random ARX programs; weak draws; a JIT in the verifier is new attack surface; the vendors must agree bit-exactly on a program they compile Reserve (named family, unlock by height or signal) now; genesis if the verifier cost is measured under the gate before the freeze
3 External cryptanalysis of M_r, the chained cache and the acceptance rule before genesis, with the x8 shape as the target Lesson 9 (MTP, Catena, Argon2i, Cuckoo); RandomX bought four audits for $141,000 before launch [S60]; the x8 decision multiplies the mixer's weight in the chip model, so a shortcut inside the mixer is now worth 8x more to a chip None Finding something late moves the vectors; not finding it in time moves nothing Genesis gate (ledger M7, raised in priority)
4 The clock and the detector. (a) A share-pattern detector on the observer: per-program hash-rate spread, nonce-group patterns and per-card-model rate bands, with an alert when a population behaves like one fixed design (MoneroCrusher's method); (b) a stated trigger: the bounty escrowed and the benchmark live before daily issuance crosses about $50K (Vorick's rule), not on a calendar date Lesson 10 (85% secret share); section 2.5's table (chips at $20K to $30K a day on compute-bound hashes); D11 (the bounty is unfunded) None A detector with false positives; a trigger the project lead has to fund Not a layer; genesis-independent; do it before the public testnet
5 Rank layer 9 (the epoch length) up, and measure the FPGA lane: the compile-ahead cost per card at a 10-minute epoch (the ca2-epoch work), plus an estimate of a soft-overlay FPGA miner with HBM (reads in flight per watt against the 5090's 17.5 G/s) FPGAs were the first adversary of Lyra2REv2 and X16R and came back within weeks of X16Rv2 (rows 7, 9); Xelis forked for FPGA resistance (row 30); a per-hour program is a bitstream target in a way a per-hash program is not At 10-minute epochs: 6x the compile work per card (measured on ca2-epoch); the VDF lead shrinks A short epoch moves the difficulty window (spec 1.12) and the seed path Reserve (as decided), with the measurement before the public testnet
6 Order the reserve by chip-unfriendliness: families that force a full 32-bit datapath per lane first (byte permute, bit-field extract, variable shifts, popcount, select, the second shuffle form), mm8 last Least Authority's "watch ML hardware"; int8 matrix blocks are licensable IP at every node; Apple pays 1.6x to 4.7x per emulated dot4 (status 20:38) None at launch None Reserve ordering, genesis
7 A vendor-share metric and a 3.0 target for the AMD gap: the share of hashrate by vendor published with the benchmark, and the line-width question kept open as the plan says Lesson 8: a one-vendor fleet is a softer version of chip capture; Equihash's NVIDIA tilt and Ethash's balance were part of each chain's miner politics (rows 3, 5) n/a A width that closes the gap makes the 5090 bandwidth-bound (status 20:27) Counter ASIC 3.0

Checks the history suggests that are not layers:

Check Why Source
1. Header grinding for cache locality: can a miner search the pre-PoW header hash H for 32-lane groups whose 128 loads cluster into fewer DRAM rows or cache lines, at a search cost below the gain? Kik's ProgPoW exploit and Dinur-Nadler's MTP attack were both "the attacker steers the addresses" [S71] [P18]
2. The chip detector's baseline: the per-program spread per card model, from the first week of the public testnet Needed before addition 4(a) can alert [S54]
3. The 28 nm SRAM density figure in section 2.5 (approximate, from memory) and the node-cost consequence, cited properly It is the argument that the cache size sets the chip's project cost [E10]

Evaluated and placed nowhere, with the reason:

Candidate Verdict Reason
Program entropy per hash (RandomX) instead of per hour Nowhere Per-hash programs need an interpreter or JIT on the GPU, which is the 25x GPU penalty RandomX pays (row 17) and the reason Igneum compiles per epoch. The filtering attack per-hash chaining prevents is already closed by the VDF seed and the acceptance rule. Per-hour's residual exposure is the FPGA lane, which addition 5 addresses with a shorter epoch, not with per-hash programs
Superscalar dependency-chain design for the program itself Nowhere, beyond what exists The program's ALU work is not the defence (lesson 1); the dependency chain that matters is the 8 dependent cache reads per item and the 128 dependent loads per hash, both in place. The superscalar idea belongs in the item derivation (addition 2)
Verthash's table from the blockchain; a dataset derived from chain history Nowhere Against the recompute chip and the partial-store chip it changes nothing: both build the table from the same public inputs the GPU does. The day key already comes from a VDF of chain state, which gives the unpredictability without a history dependency; a history dependency costs the verifier the history (Verthash needs the headers) and ties the hash to pruning (spec 10)
Grin's dual PoW with a shifting split Nowhere It is a scheduled surrender (row 19). Igneum's automatic schedules (dataset growth, reserve unlocks, cache doubling) are the shifting split applied to one hash; a second lane would hand a chip a lane
Autolykos v1's non-outsourceability Nowhere It stops pools, not chips, and Ergo removed it after 19 months because contract pools bypassed it (row 21); Igneum needs pools (spec 09)
A per-hash VRF against nonce grinding Nowhere A signature per attempt is what NexaPow did and it got a 3x to 4x chip (row 24): EC arithmetic is fixed-function work. The grinding Igneum must guard is the header-locality search (check 1), which a VRF does not touch
Ternary or variable-precision integer ops Nowhere, beyond the reserve Every family must be bit-exact on three vendors; dot4 is native on NVIDIA and AMD and emulated on Apple at 1.6x to 4.7x (status 20:38), so each precision added is paid by the weakest vendor. The reserve already holds the integer-exact candidates; adding more does not change lesson 1
Cache-timing-bound reads (ProgPoW's 16 KB cache, RandomX's L1 tier) Nowhere Measured out tonight at the L2 tier (layer 5) and the per-warp tier (layer 3): the honest card pays and the chip buys SRAM at $0.23 per MB. Rao's audit says the same about ProgPoW's cache reads
Divergent data-dependent branches Nowhere (already excluded) Branches cost a GPU divergence and a chip nothing; RandomX's single predictable branch targets speculative CPUs, which Igneum does not have
Floating point Nowhere (already excluded) Vendor rounding splits the chain (spec 1.14); RandomX could afford it because its target is one ISA family with IEEE semantics

5. Decisions this raises for the project lead

# Decision Recommendation
1 Add the partial-store chip rows to chip-model-v3.md and draw the time-memory curve before the public testnet Yes, before the vectors freeze (addition 1)
2 Name a random item-derivation program as a reserve family, and fund the verifier measurement that would move it to genesis Reserve now; genesis if the verifier lands under the gate (addition 2)
3 Commission the external cryptanalysis of M_r and the chained cache before genesis, with the x8 shape as the target Yes (addition 3; ledger M7)
4 Escrow the bounty and set its trigger to daily issuance, not to a date; build the share-pattern detector on the observer Yes to the detector now; the escrow is the project lead's (D11)
5 Rank the epoch-length reserve above the mm8 reserve, and measure the FPGA lane Yes (additions 5 and 6)

6. Sources and limits of this research

Research was gathered by four sub-agents between 20:10 and 20:45 UTC on 5 October 2026 and checked against the citations below. Fetch failures they reported: eprint.iacr.org PDFs sit behind a challenge page (abstract pages worked), so the Ren-Devadas energy figures, the Alwen-Blocki EuroS&P tables and the Lyra2 exponent come from abstracts; medium.com and bitcointalk.org returned 403 (Vorick's post was read through archive.sia.tech and secondary coverage; the IfDefElse posts through the Veil interview); Bitmain's prospectus PDF was blocked; the Dash iBeLink date and Octopus's "matrix step" are unverified; the 28 nm SRAM bit cell is from memory. Every hash-per-joule gain is derived from the cited rate and watt figures and is approximate.

Igneum sources: [I1] docs/analysis/weak-program-census-2026-10-03.md; docs/plans/counter-asic-2.md (be4b295); docs/plans/counter-asic-2-status.md and docs/plans/counter-asic-2-rollout.md (ca2-coord); docs/analysis/chip-model-v3.md (ca2-mixer 1ab8b21); docs/analysis/m16-recompute-attacker-2026-10-05.md; docs/analysis/sram-mirror.md revision 2 (ca2-analysis); docs/plans/era-layout.md (ca2-era); docs/plans/hot-table.md (ca2-cache); docs/plans/read-width.md (readwidth); docs/spec/01-lottery-hash.md, 04-seeds-and-vdf.md; docs/bench-log.md ("the 9070 XT on the eGPU", opencl-rdna4); proto-metal/MEMHARD.md; docs/fud-ledger.md M1, M3, M7, M16, C2, D11.

History rows:

Papers:

Economics and silicon:

Latency: