igneum/docs/analysis/sram-mirror.md

18 KiB

Layer 6: the SRAM mirror of the cache against published SRAM density, year 0 to 10

5 October 2026 (night), Counter ASIC 2.0 (docs/plans/counter-asic-2.md, layer 6), branch ca2-analysis. Every figure below is either cited (paper, vendor document, URL, date) or labelled approximate. Nothing here is a measurement of a chip. Numbers in this file were computed with the arithmetic shown; the script is in section 9.

1. The question

The lottery hash derives every dataset item from a 256 MiB cache (spec 01 sections 1.5 and 1.8). A chip that holds the cache in on-die SRAM can recompute items instead of reading the dataset (ledger M16, the recompute attacker). Layer 6 asks whether the cache size, as the specification schedules it, keeps that SRAM mirror unaffordable for ten years of the genesis schedule, and if not what growth rule would.

Two things also sit in a chip's SRAM budget if it mirrors the full read-only working set: the layer 5 hot table (32, 64 or 96 MB, a class parameter on readwidth b970dda, coordinator's note of 5 October) beside the 256 MiB cache. The per-warp scratch of layer 3 (32 or 128 KB per warp, written, not read-only) is not mirrorable and is left out of the mirror; it is counted in the 6 GB working-set budget in section 7.

2. What the specification schedules for the cache

Quantity Rule Where
Dataset 2 GiB at genesis plus 0.5 GiB per year (N_d grows about 23 KiB per day) spec 01 section 1.13.3, Designed
Cache 256 MiB, "prototype value, to be fixed at gate 1"; the rule that fixes it: "the cache must exceed the largest on-chip cache of any card that mines, and 96 MiB of L2 on the 5090 is the figure to beat" spec 01 sections 1.5 and 1.16
Cache growth None. No section of docs/spec/ grows the cache (grep of docs/spec for cache growth, schedule, doubling: only the dataset rule of 1.13.3 and the README's "growth" word, which refers to it) this analysis, 5 October 2026

So the plan's layer 6 row ("already in the design; confirm the schedule") is half right: dataset growth is in the design, cache growth is not. The cache is flat at 256 MiB for every year of the schedule as the spec stands. M16's closing line names the rule the cache should get ("exceeds what one die can hold, and grows") as a gate 1 decision that has not been taken.

3. SRAM bit cell per node, cited

Node (vendor) HD 6T bit cell, um^2 Raw density, Mbit/mm^2 (1/cell) Year of volume (approximate) Source
N7 (TSMC) 0.027 37.0 2018 WikiChip, "TSMC Details 5 nm" (ISSCC/IEDM disclosures), https://fuse.wikichip.org/news/3398/tsmc-details-5-nm/
N5 (TSMC) 0.021 47.6 2020 same (two N5 cells: HD 0.021, HP 0.025)
N3B (TSMC) 0.0199 50.3 2022 to 2023 WikiChip, "IEDM 2022: Did We Just Witness The Death Of SRAM?", https://fuse.wikichip.org/news/7343/iedm-2022-did-we-just-witness-the-death-of-sram/ (TSMC's IEDM 2022 N3 paper)
N3E (TSMC) 0.021 47.6 2023 same; Tom's Hardware, "TSMC's 3nm Node: No SRAM Scaling", https://www.tomshardware.com/news/no-sram-scaling-implies-on-more-expensive-cpus-and-gpus
N2 (TSMC) 0.0175 57.1 2025 to 2026 TSMC at IEDM 2024, reported by Tom's Hardware, https://www.tomshardware.com/tech-industry/tsmc-shares-deep-dive-details-about-its-cutting-edge-2nm-process-node-at-iedm-2024-35-percent-less-power-or-15-percent-more-performance ; ISSCC 2025 paper "A 38.1Mb/mm2 SRAM in a 2nm-CMOS-Nanosheet Technology", https://research.tsmc.com/page/memory/4.html
Intel 18A 0.021 47.6 2025 to 2026 ISSCC 2025 paper 29.2, "A 0.021 um^2 High-Density SRAM in Intel 18A RibbonFET Technology with PowerVia", https://www.researchgate.net/publication/389644177 ; IEEE Spectrum 26 Feb 2025, https://spectrum.ieee.org/sram-intel-tsmc
Samsung SF3 / SF2 not disclosed as a bit cell area in anything found tonight (Samsung's ISSCC papers give assist circuits and macro figures, not the HD cell) search of ISSCC 2021 to 2025 coverage, 5 October 2026; left out of the tables

The stall. N3B's cell is 5% smaller than N5's and N3E's is the same size as N5's (0.021 um^2 both): zero SRAM scaling from N5 to N3E (WikiChip IEDM 2022 article above; Tom's Hardware above; SemiAnalysis "TSMC's 3nm Conundrum", https://newsletter.semianalysis.com/p/tsmcs-3nm-conundrum-does-it-even). N2's nanosheet cell recovers 17% (0.021 to 0.0175 um^2). So across 2020 to 2026 the HD bit cell shrank once, by 17%.

Array efficiency (bit cell to macro). The usable density of a macro is below 1/cell because of word-line and bit-line drivers, sense amplifiers, decoders and redundancy. The factor used here is 0.70, WikiChip's convention (their 31.8 Mib/mm^2 for the 0.021 um^2 N3E cell is 1/0.021 x 0.70 in Mib). The two ISSCC 2025 macros bracket it: TSMC N2 38.1 Mb/mm^2 at a 0.0175 um^2 cell is 67%; Intel 18A 38.1 Mb/mm^2 array density and 34.3 Mb/mm^2 for the volume macro at a 0.021 um^2 cell are 80% and 72% (ISSCC 2025 29.2, above). Both lie within 10% of 0.70.

GPU on-die SRAM for scale: the RTX 5090 carries 96 MB of L2 (98,304 KB) on a 750 mm^2 TSMC 4N die with 92.2 billion transistors; the full GB202 has 128 MB; the RTX 4090 had 72 MB and the RTX 3090 6 MB (NVIDIA, "RTX Blackwell GPU Architecture" whitepaper v1.1, appendix table "L2 Cache Size", https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf). At the N5-class cell and 0.70 that L2 is about 24 mm^2 of the 750 (3%), approximate. The RX 9070 XT carries 64 MB of Infinity Cache plus 8 MB of L2 (vendor figures, approximate, bench-log "the 9070 XT on the eGPU").

Reticle: the EUV field is 26 x 33 mm = 858 mm^2, about 830 mm^2 usable after scribe lanes (SemiAnalysis, "Die Size And Reticle Conundrum", https://newsletter.semianalysis.com/p/die-size-and-reticle-conundrum-cost ; WikiChip "Mask", https://en.wikichip.org/wiki/mask). The 5090's 750 mm^2 is 90% of it.

Wafer prices (approximate; TSMC publishes none, every figure is supply-chain reporting): N7 about $9,500, N5 and N3 about $20,000 (Silicon Analysts, "Wafer Pricing by Node", September 2026, https://siliconanalysts.com/data/wafer-pricing); N2 about $30,000 (Tom's Hardware, https://www.tomshardware.com/tech-industry/semiconductors/tsmc-could-charge-up-to-usd45-000-for-1-6nm-wafers-rumors-allege-a-50-percent-increase-in-pricing-over-prior-gen-wafers).

4. Die area to mirror the cache, per node

Area = bits / (raw density x 0.70). The columns are the 256 MiB cache alone, the cache plus each hot-table size of layer 5 (32, 64, 96 MB taken as MiB), and the larger caches of the options in section 6.

Node Macro Mbit/mm^2 at 0.70 256 MiB 256 + 32 256 + 64 256 + 96 512 MiB 1 GiB 2 GiB 4 GiB
N7 25.9 83 mm^2 93 104 114 166 331 663 1,325 (2 dies)
N5 33.3 64 72 81 89 129 258 515 1,031 (2 dies)
N3B 35.2 61 69 76 84 122 244 488 977 (2 dies)
N3E, Intel 18A 33.3 64 72 81 89 129 258 515 1,031 (2 dies)
N2 40.0 54 60 67 74 107 215 429 859 (2 dies)

One reticle (830 mm^2) holds 2.5 GiB of SRAM at N7, 3.2 GiB at N5, N3E and 18A, 3.9 GiB at N2 (same arithmetic).

Against the figures the ledger carries: M16's "100 to 300 mm^2" (low end from a 0.02 um^2 cell with overhead, high end from wafer-scale parts at about 1 MB/mm^2) and the plan's "about 45 mm^2 at a leading node" both bracket the cited 54 to 64 mm^2; the wafer-scale high end is a different efficiency (Cerebras-class arrays sit beside logic) and is not the right number for a pure SRAM die. The right figure for the ledger is 54 to 83 mm^2 depending on node, cited above.

5. Cost per good die

Dies per 300 mm wafer by the usual approximation pi x 150^2 / A minus the edge term pi x 300 / sqrt(2A); yield by Poisson exp(-A x D0) with D0 = 0.1 defects per cm^2 (an assumption, approximate; SRAM arrays carry redundancy so real yield is higher, which lowers these costs). Cost per good die = wafer price / (dies x yield). Packaging, test, the logic beside the SRAM and the design (masks at N5 and below run into the tens of millions of dollars, approximate) are not in these numbers; they are per-die silicon only.

Node, wafer price 256 MiB 256 + 96 MiB 1 GiB 4 GiB (2 dies)
N7, $9,500 83 mm^2, 780 dies, yield 0.92, $13 $19 331 mm^2, 177 dies, 0.72, $75 $456
N5, $20,000 64 mm^2, 1,014 dies, 0.94, $21 $30 258 mm^2, 233 dies, 0.77, $111 $621
N3E, $20,000 $21 $30 $111 $621
N2, $30,000 54 mm^2, 1,226 dies, 0.95, $26 $37 215 mm^2, 284 dies, 0.81, $131 $696

Reading. The silicon for a 256 MiB mirror is tens of dollars per die on any node from N7 up. With the hot table it is still under $40. It was never the SRAM that priced the recompute attacker out; the plan's premise for layer 6 ("the SRAM mirror stays unaffordable") does not hold for the cache as a mirror and did not hold at genesis either.

6. What the mirror buys the attacker, year by year

From M16 (docs/analysis/m16-recompute-attacker-2026-10-05.md): with the cache on die the attacker recomputes 128 items per hash at about 1,170 integer operations and 8 dependent 64-byte cache reads each, about 150,000 operations and 1,024 dependent SRAM reads per hash. At a 5090-class integer budget (about 50 T op/s, approximate) that is 0.33 Ghash/s against the honest 141 Mhash/s projected for version 2 programs: 2.4x at equal silicon before any fixed-function factor, 3x to 6x with one (approximate). The SRAM is 54 to 83 mm^2 of that chip (7 to 11% of a 750 mm^2 die), so the mirror is cheap and the recompute route is bound by integer throughput, not by SRAM.

The layer 5 hot table changes nothing in that arithmetic: the hot table is read-only and derived from the day key like the cache, so a chip mirrors it in the same SRAM (another 32 to 96 MB, 7 to 24 mm^2 at N5) and reads it at SRAM latency, which is exactly what a GPU's L2 does with it. Layer 5 taxes the DRAM-only chip (the one without SRAM); it does not tax the SRAM chip.

Dataset growth does not touch the recompute attacker: the attacker never holds the dataset. It taxes the partial-store attacker (O-1.6, the time-memory curve, not drawn) and the honest card.

Year by year under the schedule as it stands (flat 256 MiB), the mirror's area at the best node available that year. Node years are approximate; the density trend from 2018 to 2025 is 37.0 to 57.1 Mbit/mm^2 raw, 1.54x in 7 years, about 6% per year, and it came in one step (N2); the extrapolation past 2026 assumes that average holds (approximate, and optimistic for the attacker: A16 and A14 have no disclosed SRAM cell yet).

Year Calendar (approximate) Dataset, GiB Cache (spec) Best node, raw Mbit/mm^2 Mirror of the cache, mm^2 With a 96 MiB hot table, mm^2 Mirror as a share of a 750 mm^2 die
0 2027 2.0 256 MiB N2, 57.1 (cited) 54 74 7%
1 2028 2.5 256 MiB N2 or A16, 57 to 61 50 to 54 69 to 74 7%
2 2029 3.0 256 MiB about 64 (trend) 48 66 6%
3 2030 3.5 256 MiB about 68 45 62 6%
4 2031 4.0 256 MiB about 72 43 59 6%
5 2032 4.5 256 MiB about 76 40 55 5%
6 2033 5.0 256 MiB about 81 38 52 5%
7 2034 5.5 256 MiB about 86 36 49 5%
8 2035 6.0 256 MiB about 91 34 46 5%
9 2036 6.5 256 MiB about 97 32 44 4%
10 2037 7.0 256 MiB about 102 30 41 4%

Reading. A flat cache's mirror shrinks from 7% to 4% of a large die over the decade, and a 5090-class consumer GPU already carries 96 MB of L2 on one die with the full GB202 at 128 MB; at the 2020 to 2025 pace of GPU L2 growth (6 MB, 72 MB, 96 MB on the three NVIDIA flagships in the whitepaper table) a consumer GPU could hold 256 MiB on die within the decade. The spec's own rule for the cache ("must exceed the largest on-chip cache of any card that mines") would then be broken by a flat cache. That is the real reason to grow it: not to price a chip out (section 5 shows the SRAM cannot do that) but to keep the cache out of every GPU's own cache, so the honest hash stays DRAM-latency-bound and the recompute route stays a route only a custom chip can take.

7. Answer to the layer 6 question, and the options

Does the flat 256 MiB cache keep the SRAM mirror unaffordable through year 10? No. It is affordable at year 0 (tens of dollars of silicon per die, section 5) and gets cheaper. What keeps the recompute attacker near 1x is M16's integer arithmetic and the mixer-cost lever (4x the mixer cost puts the equal-silicon gain at 0.36x, bounded by the CPU verify gate), not the cache size. The cache size does one other job, keeping the cache larger than any GPU's L2, and that job needs growth.

Options for the cache rule, with the honest costs each implies. Verifier fill time is 0.2 s per 256 MiB on one core (spec 1.12: "a 0.2 s CPU cache fill", from the measured 175 to 190 ms of section 1.8.3), scaled linearly; the verifier holds the whole cache (section 1.11), so its memory is the cache size plus the program and the interpreter. GPU fill: 0.67 ms per 256 MiB on the 5090 (section 1.8.3), linear. The GPU dataset build (13.4 ms per 1 GiB on the 5090, section 1.8.3) depends on the dataset size, not the cache size; a larger cache spreads the build's 8 dependent reads per item over more memory, which on a GPU means more of them miss L2 and the build slows by some factor between 1x and the L2-to-DRAM latency ratio, which is a measurement to take (approximate; owed). Mirror area is at N2 (cited density), the node of the first years; at the trend's year-10 density divide by about 1.8.

Option Rule Cache at year 0 / 4 / 10 Mirror at N2, year 0 / 4 / 10 (mm^2) Dies at year 10 (830 mm^2 reticle) Verifier fill, one core, year 0 / 10 Verifier memory, year 10 GPU cache fill (5090), year 10 Keeps the cache above a 96 MB L2 at year 10 Keeps it above a 256 MB L2
A, as specified flat 256 MiB 256 / 256 / 256 MiB 54 / 54 / 54 1 0.2 / 0.2 s 256 MiB 0.7 ms yes, 2.7x no
B cache = dataset / 8 (today's ratio) 256 / 512 / 896 MiB 54 / 107 / 188 1 0.2 / 0.7 s 896 MiB 2.3 ms yes, 9.3x yes, 3.5x
C cache doubles when the dataset doubles (the dataset's own clock: year 4, then year 12) 256 / 512 / 512 MiB 54 / 107 / 107 1 0.2 / 0.4 s 512 MiB 1.3 ms yes, 5.3x yes, 2x
D cache = dataset / 4 512 / 1,024 / 1,792 MiB 107 / 215 / 376 1 0.4 / 1.4 s 1.75 GiB 4.7 ms yes yes, 7x
E, one reticle cache sized so the mirror exceeds one reticle at the node of the day: 4 GiB at N2 (section 4), growing with density 4 GiB / about 4.5 / about 7 GiB 859 / 860 / 860 (by construction) 2 3.2 / 5.6 s 7 GiB 11 / 19 ms yes yes

Where the working set enters (coordinator's budget: 1 GiB table + hot table + scratch for every resident warp + buffers under 6 GB on an 8 GB card): the cache is not in the miner's working set at hash time (the dataset is built from it once a day and the cache can be dropped or kept), so options A to D do not move that budget; the dataset's own growth does (2 GiB at genesis, 4 GiB at year 4, 7 GiB at year 10, which is past an 8 GB card at about year 8 on its own). Option E's 4 GiB cache would have to be built on the card and dropped, which is fine for a 16 GB card and tight on an 8 GB one at build time (4 GiB cache + 2 GiB dataset + hot table). The per-warp scratch at 170 SMs x 64 warps (approximate, readwidth) is 340 MB at 32 KB and 1.36 GB at 128 KB per warp; with the 1 GiB table, a 96 MB hot table and buffers that is 1.5 to 2.5 GB at the prototype dataset size, 2.5 to 3.5 GB at the 2 GiB genesis size, inside 6 GB either way.

Recommendation. Option C (the cache doubles when the dataset doubles) is the one that keeps the spec's own rule true with the smallest verifier cost: it ties the cache to a clock the spec already has, keeps AND MASK (a power of two every step, which is the 1.13.3 option (b) argument again), costs the verifier 0.4 s and 512 MiB at year 4 and nothing more until year 12, and keeps the cache 2x above a 256 MB GPU L2 if one appears. It does not price a chip out; nothing about cache size does (section 5). The lever that does is the mixer cost multiplier of M16, which is the gate 1 decision to take beside this one. Option B is the same idea in a smooth form and costs the verifier 0.7 s at year 10. Option E is the only one that makes the mirror a multi-die part and it costs every verifier 3.2 s and 4 GiB at genesis, which fails the spirit of the 10 ms verify gate (the fill is once a day, but a light node joining pays it on every day it syncs across).

Decision for Josh, at gate 1: A, B, C, D or E above, together with M16's mixer multiplier. Nothing here changes a vector today: the cache size is a prototype value of spec 1.16 and the growth rule would be a new sentence in 1.13.3.

8. What is cited, what is approximate, what is owed

Item Status
Bit cells for N7, N5, N3B, N3E, N2, Intel 18A cited (section 3)
Samsung SF2 or SF3 bit cell not found; left out
Array efficiency 0.70 WikiChip's convention, bracketed by two ISSCC 2025 macros (67 to 80%)
Wafer prices approximate, supply-chain reporting, cited
D0 = 0.1 per cm^2, Poisson yield assumption, stated
Node years and the 6% per year density trend past 2026 approximate, extrapolated from cited 2018 to 2025 points
GPU L2 sizes cited (NVIDIA whitepaper); AMD Infinity Cache approximate
Recompute attacker arithmetic M16, which is itself arithmetic on measured rates, not a chip measurement
Dataset-build slowdown at a larger cache on a GPU owed, a measurement (5090 at a 512 MiB and 1 GiB cache)
The on-die emulation of M16 (inline kernel with a 64 MiB cache inside the 5090's L2) still a PC job (M16)

9. The arithmetic

MiB = 2^20; bits = cache_MiB * MiB * 8
raw_Mbit_per_mm2 = 1 / cell_um2            (1e6 cells per mm2 per um2 of cell)
area_mm2 = bits / (raw * 0.70 * 1e6)
dies_per_wafer = pi * 150^2 / area - pi * 300 / sqrt(2 * area)
yield = exp(-area_mm2 * 0.001)             (D0 = 0.1 per cm2)
cost_per_good_die = wafer_price / (dies * yield)
reticle_GiB = 830 * raw * 0.70 * 1e6 / 8 / 2^30

Run on 5 October 2026 with Python 3 on the M5 Max; the printed tables are the ones above, rounded.