igneum/docs/analysis/sram-mirror.md

24 KiB

Layer 6: the SRAM mirror of the cache against published SRAM density, year 0 to 10

5 October 2026 (night), Counter ASIC 2.0 (docs/plans/counter-asic-2.md, layer 6), branch ca2-analysis. Every figure below is either cited (paper, vendor document, URL, date) or labelled approximate. Nothing here is a measurement of a chip. Numbers in this file were computed with the arithmetic shown; the script is in section 10.

Revision 2 (same night): the first draft priced the mirror from bit-cell area times a 0.70 array factor. The coordinator's chip-economics research (sources below) showed that shipped cache-only dies land at about half that density once assist circuits, redundancy, TSVs, power and test are in. Every table now carries two columns: the shipped-product density as the headline and the bit-cell figure as the lower bound. The conclusion did not move; the cost per die rose 2 to 3x.

1. The question

The lottery hash derives every dataset item from a 256 MiB cache (spec 01 sections 1.5 and 1.8). A chip that holds the cache in on-die SRAM can recompute items instead of reading the dataset (ledger M16, the recompute attacker). Layer 6 asks whether the cache size, as the specification schedules it, keeps that SRAM mirror unaffordable for ten years of the genesis schedule, and if not what growth rule would.

Two things also sit in a chip's SRAM budget if it mirrors the full read-only working set: the layer 5 hot table (32, 64 or 96 MB, a class parameter on readwidth b970dda, coordinator's note of 5 October) beside the 256 MiB cache. The per-warp scratch of layer 3 (32 or 128 KB per warp, written, not read-only) is not mirrorable and is left out of the mirror; it is counted in the 6 GB working-set budget in section 7.

2. What the specification schedules for the cache

Quantity Rule Where
Dataset 2 GiB at genesis plus 0.5 GiB per year (N_d grows about 23 KiB per day) spec 01 section 1.13.3, Designed
Cache 256 MiB, "prototype value, to be fixed at gate 1"; the rule that fixes it: "the cache must exceed the largest on-chip cache of any card that mines, and 96 MiB of L2 on the 5090 is the figure to beat" spec 01 sections 1.5 and 1.16
Cache growth None. No section of docs/spec/ grows the cache (grep of docs/spec for cache growth, schedule, doubling: only the dataset rule of 1.13.3 and the README's "growth" word, which refers to it) this analysis, 5 October 2026

So the plan's layer 6 row ("already in the design; confirm the schedule") is half right: dataset growth is in the design, cache growth is not. The cache is flat at 256 MiB for every year of the schedule as the spec stands. M16's closing line names the rule the cache should get ("exceeds what one die can hold, and grows") as a gate 1 decision that has not been taken.

3. SRAM density, cited: bit cells per node and shipped cache dies

3.1 Bit cells

Node (vendor) HD 6T bit cell, um^2 Raw density, Mbit/mm^2 (1/cell) Year of volume (approximate) Source
N7 (TSMC) 0.027 37.0 2018 WikiChip, "TSMC Details 5 nm" (ISSCC/IEDM disclosures), https://fuse.wikichip.org/news/3398/tsmc-details-5-nm/
N5 (TSMC) 0.021 47.6 2020 same (two N5 cells: HD 0.021, HP 0.025)
N3B (TSMC) 0.0199 50.3 2022 to 2023 WikiChip, "IEDM 2022: Did We Just Witness The Death Of SRAM?", https://fuse.wikichip.org/news/7343/iedm-2022-did-we-just-witness-the-death-of-sram/ (TSMC's IEDM 2022 N3 paper)
N3E (TSMC) 0.021 47.6 2023 same; Tom's Hardware, "TSMC's 3nm Node: No SRAM Scaling", https://www.tomshardware.com/news/no-sram-scaling-implies-on-more-expensive-cpus-and-gpus
N2 (TSMC) 0.0175 57.1 2025 to 2026 TSMC at IEDM 2024, reported by Tom's Hardware, https://www.tomshardware.com/tech-industry/tsmc-shares-deep-dive-details-about-its-cutting-edge-2nm-process-node-at-iedm-2024-35-percent-less-power-or-15-percent-more-performance ; ISSCC 2025 paper "A 38.1Mb/mm2 SRAM in a 2nm-CMOS-Nanosheet Technology", https://research.tsmc.com/page/memory/4.html
Intel 18A 0.021 47.6 2025 to 2026 ISSCC 2025 paper 29.2, "A 0.021 um^2 High-Density SRAM in Intel 18A RibbonFET Technology with PowerVia", https://www.researchgate.net/publication/389644177 ; IEEE Spectrum 26 Feb 2025, https://spectrum.ieee.org/sram-intel-tsmc
Samsung SF3 / SF2 not disclosed as a bit cell area in anything found tonight (Samsung's ISSCC papers give assist circuits and macro figures, not the HD cell) search of ISSCC 2021 to 2025 coverage, 5 October 2026; left out of the tables

The stall. N3B's cell is 5% smaller than N5's and N3E's is the same size as N5's (0.021 um^2 both): zero SRAM scaling from N5 to N3E (WikiChip IEDM 2022 article above; Tom's Hardware above; SemiAnalysis "TSMC's 3nm Conundrum", https://newsletter.semianalysis.com/p/tsmcs-3nm-conundrum-does-it-even). N2's nanosheet cell recovers 17% (0.021 to 0.0175 um^2). So across 2020 to 2026 the HD bit cell shrank once, by 17%.

Macro density from the bit cell. WikiChip's and SemiAnalysis's convention is bit-cell density times about 0.70 for the assist and periphery overhead (SemiAnalysis, December 2022: TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30% assist overhead; WikiChip's 31.8 Mib/mm^2 for the 0.021 um^2 cell is the same arithmetic). The two ISSCC 2025 macros bracket it: TSMC N2 38.1 Mb/mm^2 at a 0.0175 um^2 cell is 67%; Intel 18A 38.1 Mb/mm^2 array density and 34.3 Mb/mm^2 for the volume macro at a 0.021 um^2 cell are 80% and 72%. That is a macro on a test chip. It is the LOWER BOUND on die area, not the die.

3.2 Shipped cache dies (what a whole die of SRAM really holds)

Product SRAM Die Node MB per mm^2 Source
AMD 3D V-Cache (Zen 3 SRAM chiplet) 64 MB 41 mm^2 TSMC 7 nm 1.56 AMD at Hot Chips 33, reported by Tom's Hardware, August 2021, https://www.tomshardware.com/news/amd-unveils-more-ryzen-3d-packaging-and-v-cache-details-at-hot-chips ("the 3D V-Cache SRAM measures 41 mm^2", "64 MB of 7 nm SRAM"); the densest cache-only die that has shipped
Graphcore GC200 (with compute) 900 MB 823 mm^2 7 nm 1.09 coordinator's chip-economics research, 5 October 2026 (vendor figures)
Groq TSP 220 MB 725 mm^2 14 nm 0.30 same

The V-Cache die is a pure SRAM die with its TSVs, redundancy, test and power: 1.56 MB/mm^2 at N7 against the bit-cell figure 37.0 Mbit/mm^2 = 4.6 MB/mm^2 and the 0.70-macro figure 3.2 MB/mm^2. The shipped die is 0.48 of the macro figure. The headline column below scales the V-Cache density to other nodes by the bit-cell ratio (0.027 / cell), an approximation that assumes the periphery and TSV overheads scale with the cell, which they do not fully (so the headline column is itself slightly optimistic for the attacker at N5 and below).

3.3 GPU on-die SRAM, the reticle, wafer prices

GPU on-die SRAM for scale: the RTX 5090 carries 96 MB of L2 (98,304 KB) on a 750 mm^2 TSMC 4N die with 92.2 billion transistors; the full GB202 has 128 MB; the RTX 4090 had 72 MB and the RTX 3090 6 MB (NVIDIA, "RTX Blackwell GPU Architecture" whitepaper v1.1, appendix table "L2 Cache Size", https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf). At the V-Cache density scaled to N5 (2.0 MB/mm^2) that L2 is about 48 mm^2 of the 750 (6%), approximate. The RX 9070 XT carries 64 MB of Infinity Cache plus 8 MB of L2 (vendor figures, approximate, bench-log "the 9070 XT on the eGPU").

Reticle: the EUV field is 26 x 33 mm = 858 mm^2, about 830 mm^2 usable after scribe lanes (SemiAnalysis, "Die Size And Reticle Conundrum", https://newsletter.semianalysis.com/p/die-size-and-reticle-conundrum-cost ; WikiChip "Mask", https://en.wikichip.org/wiki/mask). The 5090's 750 mm^2 is 90% of it.

Wafer prices (approximate; TSMC publishes none, every figure is supply-chain reporting): N7 about $9,500, N5 and N3 about $20,000 (Silicon Analysts, "Wafer Pricing by Node", September 2026, https://siliconanalysts.com/data/wafer-pricing); N2 about $30,000 (Tom's Hardware, https://www.tomshardware.com/tech-industry/semiconductors/tsmc-could-charge-up-to-usd45-000-for-1-6nm-wafers-rumors-allege-a-50-percent-increase-in-pricing-over-prior-gen-wafers).

4. Die area to mirror the cache, per node, two columns

Headline = V-Cache density (41 mm^2 per 64 MiB at N7) scaled by the bit-cell ratio. Lower bound = bits / (raw density x 0.70). Columns: the 256 MiB cache alone, the cache plus the 96 MB hot table of layer 5 (as MiB), and the larger caches of the options in section 7. Area in mm^2; a figure over 830 is split into the dies shown.

Node 256 MiB, headline 256 MiB, lower bound 256 + 96, headline 256 + 96, lower bound 512 MiB, headline / lower 1 GiB, headline / lower 4 GiB, headline / lower
N7 164 83 226 114 328 / 166 656 / 331 2,624 (4 dies) / 1,325 (2 dies)
N5 128 64 175 89 255 / 129 510 / 258 2,041 (3 dies) / 1,031 (2 dies)
N3B 121 61 166 84 242 / 122 483 / 244 1,934 (3 dies) / 977 (2 dies)
N3E, Intel 18A 128 64 175 89 255 / 129 510 / 258 2,041 (3 dies) / 1,031 (2 dies)
N2 106 54 146 74 213 / 107 425 / 215 1,701 (3 dies) / 859 (2 dies)

One reticle (830 mm^2) holds, at the headline density, 1.3 GiB of SRAM at N7, 1.6 GiB at N5, N3E and 18A, 1.9 GiB at N2 (lower-bound column: 2.5, 3.2, 3.9 GiB).

Against the figures the ledger carries: M16's "100 to 300 mm^2" (low end from a 0.02 um^2 cell with overhead, high end from wafer-scale parts at about 1 MB per mm^2) brackets the headline 106 to 164 mm^2 well; the plan's "about 45 mm^2 at a leading node" is below even the lower bound and should be read as the bit-cell area with no overhead. The right figures for the ledger are 106 to 164 mm^2 (shipped density) with 54 to 83 mm^2 as the floor.

5. Cost per good die, two columns

Dies per 300 mm wafer by the usual approximation pi x 150^2 / A minus the edge term pi x 300 / sqrt(2A); yield by Poisson exp(-A x D0) with D0 = 0.1 defects per cm^2 (an assumption, approximate; SRAM arrays carry redundancy so real yield is higher, which lowers these costs). Cost per good die = wafer price / (dies x yield). Packaging, test, the logic beside the SRAM and the design (masks at N5 and below run into the tens of millions of dollars, approximate) are not in these numbers; they are per-die silicon only. Headline / lower bound in each cell.

Node, wafer price 256 MiB 256 + 96 MiB 1 GiB 4 GiB
N7, $9,500 164 mm^2, 379 dies, yield 0.85: $30 / $13 $44 / $19 $224 / $75 $896 (4 dies) / $456 (2 dies)
N5, $20,000 128 mm^2, 495 dies, 0.88: $46 / $21 $68 / $30 $306 / $111 $1,512 (3 dies) / $621 (2 dies)
N3B, $20,000 121 mm^2, 524 dies, 0.89: $43 / $20 $63 / $28 $280 / $103 $1,371 (3 dies) / $569 (2 dies)
N3E, 18A, $20,000 $46 / $21 $68 / $30 $306 / $111 $1,512 / $621
N2, $30,000 106 mm^2, 600 dies, 0.90: $56 / $26 $81 / $37 $343 / $131 $1,641 (3 dies) / $696 (2 dies)

Reading. The silicon for a 256 MiB mirror is $30 to $56 per die at shipped density (2 to 3x the first draft's figure), under $90 with the hot table. A funded chip programme pays that without noticing: it was never the SRAM that priced the recompute attacker out, and the plan's premise for layer 6 ("the SRAM mirror stays unaffordable") does not hold for the cache as a mirror and did not hold at genesis either. A 1 GiB cache is a 425 to 656 mm^2 die ($224 to $343), affordable too; 4 GiB is a 3 to 4 die part at about $900 to $1,600 of silicon, which is a different product but not an impossible one (the attacker's problem at that size is the 1,024 dependent cross-die reads per hash, section 6).

6. What the mirror buys the attacker, year by year

From M16 (docs/analysis/m16-recompute-attacker-2026-10-05.md): with the cache on die the attacker recomputes 128 items per hash at about 1,170 integer operations and 8 dependent 64-byte cache reads each, about 150,000 operations and 1,024 dependent SRAM reads per hash. At a 5090-class integer budget (about 50 T op/s, approximate) that is 0.33 Ghash/s against the honest 141 Mhash/s projected for version 2 programs: 2.4x at equal silicon before any fixed-function factor, 3x to 6x with one (approximate). The SRAM is 106 to 164 mm^2 of that chip at the headline density (14 to 22% of a 750 mm^2 die; the m16 model's 13 to 40% band holds), so the mirror is cheap and the recompute route is bound by integer throughput, not by SRAM.

The layer 5 hot table changes nothing in that arithmetic: the hot table is read-only and derived from the day key like the cache, so a chip mirrors it in the same SRAM (another 32 to 96 MB, 24 to 48 mm^2 at N5 headline) and reads it at SRAM latency, which is exactly what a GPU's L2 does with it. Layer 5 taxes the DRAM-only chip (the one without SRAM); it does not tax the SRAM chip.

Dataset growth does not touch the recompute attacker: the attacker never holds the dataset. It taxes the partial-store attacker (O-1.6, the time-memory curve, not drawn) and the honest card.

Year by year under the schedule as it stands (flat 256 MiB), the mirror's area at the best node available that year, headline density. Node years are approximate; the density trend from 2018 to 2025 is 37.0 to 57.1 Mbit/mm^2 raw, 1.54x in 7 years, about 6% per year, and it came in one step (N2); the extrapolation past 2026 assumes that average holds (approximate, and optimistic for the attacker: A16 and A14 have no disclosed SRAM cell yet).

Year Calendar (approximate) Dataset, GiB Cache (spec) Best node Mirror of the cache, headline (lower bound), mm^2 With a 96 MiB hot table, headline, mm^2 Mirror as a share of a 750 mm^2 die
0 2027 2.0 256 MiB N2 (cited) 106 (54) 146 14%
1 2028 2.5 256 MiB N2 or A16 103 (52) 142 14%
2 2029 3.0 256 MiB trend 95 (48) 130 13%
3 2030 3.5 256 MiB trend 89 (45) 123 12%
4 2031 4.0 256 MiB trend 84 (43) 116 11%
5 2032 4.5 256 MiB trend 79 (40) 109 11%
6 2033 5.0 256 MiB trend 75 (38) 103 10%
7 2034 5.5 256 MiB trend 71 (36) 97 9%
8 2035 6.0 256 MiB trend 67 (34) 92 9%
9 2036 6.5 256 MiB trend 63 (32) 87 8%
10 2037 7.0 256 MiB trend 59 (30) 82 8%

Reading. A flat cache's mirror shrinks from 14% to 8% of a large die over the decade, and a 5090-class consumer GPU already carries 96 MB of L2 on one die with the full GB202 at 128 MB; at the 2020 to 2025 pace of GPU L2 growth (6 MB, 72 MB, 96 MB on the three NVIDIA flagships in the whitepaper table) a consumer GPU could hold 256 MiB on die within the decade. The spec's own rule for the cache ("must exceed the largest on-chip cache of any card that mines") would then be broken by a flat cache. That is the real reason to grow it: not to price a chip out (section 5 shows the SRAM cannot do that) but to keep the cache out of every GPU's own cache, so the honest hash stays DRAM-latency-bound and the recompute route stays a route only a custom chip can take.

7. Answer to the layer 6 question, and the options

Does the flat 256 MiB cache keep the SRAM mirror unaffordable through year 10? No. It is affordable at year 0 ($30 to $56 of silicon per die at shipped density, section 5) and gets cheaper. What keeps the recompute attacker near 1x is M16's integer arithmetic and the mixer-cost lever (4x the mixer cost puts the equal-silicon gain at 0.36x, bounded by the CPU verify gate), not the cache size. The cache size does one other job, keeping the cache larger than any GPU's L2, and that job needs growth.

Options for the cache rule, with the honest costs each implies. Verifier fill time is 0.2 s per 256 MiB on one core (spec 1.12: "a 0.2 s CPU cache fill", from the measured 175 to 190 ms of section 1.8.3), scaled linearly; the verifier holds the whole cache (section 1.11), so its memory is the cache size plus the program and the interpreter. GPU fill: 0.67 ms per 256 MiB on the 5090 (section 1.8.3), linear. The GPU dataset build (13.4 ms per 1 GiB on the 5090, section 1.8.3) depends on the dataset size, not the cache size; a larger cache spreads the build's 8 dependent reads per item over more memory, which on a GPU means more of them miss L2 and the build slows by some factor between 1x and the L2-to-DRAM latency ratio, which is a measurement to take (approximate; owed). Mirror area is at N2 headline density (lower bound in brackets), the node of the first years; at the trend's year-10 density divide by about 1.8.

Option Rule Cache at year 0 / 4 / 10 Mirror at N2, headline (lower bound), year 0 / 4 / 10, mm^2 Dies at year 10 (830 mm^2 reticle), headline Verifier fill, one core, year 0 / 10 Verifier memory, year 10 GPU cache fill (5090), year 10 Keeps the cache above a 96 MB L2 at year 10 Keeps it above a 256 MB L2
A, as specified flat 256 MiB 256 / 256 / 256 MiB 106 (54) / 106 / 106 1 0.2 / 0.2 s 256 MiB 0.7 ms yes, 2.7x no
B cache = dataset / 8 (today's ratio) 256 / 512 / 896 MiB 106 (54) / 213 (107) / 372 (188) 1 0.2 / 0.7 s 896 MiB 2.3 ms yes, 9.3x yes, 3.5x
C cache doubles when the dataset doubles (the dataset's own clock: year 4, then year 12) 256 / 512 / 512 MiB 106 (54) / 213 (107) / 213 (107) 1 0.2 / 0.4 s 512 MiB 1.3 ms yes, 5.3x yes, 2x
D cache = dataset / 4 512 / 1,024 / 1,792 MiB 213 (107) / 425 (215) / 744 (376) 1 0.4 / 1.4 s 1.75 GiB 4.7 ms yes yes, 7x
E, one reticle cache sized so the mirror exceeds one reticle at the node of the day: 2 GiB at N2 headline density (section 4; 4 GiB on the lower bound), growing with density 2 GiB / about 2.3 / about 3.5 GiB 850 / 850 / 850 (by construction) 2 1.6 / 2.8 s 3.5 GiB 5.4 / 9.4 ms yes yes

Where the working set enters (coordinator's budget: 1 GiB table + hot table + scratch for every resident warp + buffers under 6 GB on an 8 GB card): the cache is not in the miner's working set at hash time (the dataset is built from it once a day and the cache can be dropped or kept), so options A to D do not move that budget; the dataset's own growth does (2 GiB at genesis, 4 GiB at year 4, 7 GiB at year 10, which is past an 8 GB card at about year 8 on its own). Option E's 2 GiB cache would have to be built on the card and dropped, which is fine for a 16 GB card and tight on an 8 GB one at build time (2 GiB cache + 2 GiB dataset + hot table). The per-warp scratch at 170 SMs x 64 warps (approximate, readwidth) is 340 MB at 32 KB and 1.36 GB at 128 KB per warp; with the 1 GiB table, a 96 MB hot table and buffers that is 1.5 to 2.5 GB at the prototype dataset size, 2.5 to 3.5 GB at the 2 GiB genesis size, inside 6 GB either way.

Recommendation. Option C (the cache doubles when the dataset doubles) is the one that keeps the spec's own rule true with the smallest verifier cost: it ties the cache to a clock the spec already has, keeps AND MASK (a power of two every step, which is the 1.13.3 option (b) argument again), costs the verifier 0.4 s and 512 MiB at year 4 and nothing more until year 12, and keeps the cache 2x above a 256 MB GPU L2 if one appears. It does not price a chip out; nothing about cache size does (section 5). The lever that does is the mixer cost multiplier of M16, which is the gate 1 decision to take beside this one. Option B is the same idea in a smooth form and costs the verifier 0.7 s at year 10. Option E is the only one that makes the mirror a multi-die part and it costs every verifier 1.6 s and 2 GiB at genesis (at the headline density; the lower-bound density would ask for 4 GiB and 3.2 s), which fails the spirit of the 10 ms verify gate (the fill is once a day, but a light node joining pays it on every day it syncs across).

Decision for Josh, at gate 1: A, B, C, D or E above, together with M16's mixer multiplier. Nothing here changes a vector today: the cache size is a prototype value of spec 1.16 and the growth rule would be a new sentence in 1.13.3.

8. Why the latency bound is the property to lean on (citations behind the plan's rule)

The plan's "what stays true" paragraph says DRAM latency is the same physics for everyone and bandwidth per watt is what a custom memory chip buys. The sources behind that:

Claim Figure Source
Random-access DRAM latency is the same across memory types Row cycle time 40 to 48 ns across DDR4, GDDR5 and HBM2 Li, Reddy and Jacob, "A Performance and Power Comparison of Contemporary DRAM Architectures", MEMSYS 2018 (coordinator's chip-economics research, 5 October 2026)
Latency does not scale, bandwidth does DRAM latency improved about 1.3x in two decades while bandwidth improved about 20x K. Chang, "Understanding and Improving the Latency of DRAM-Based Memory Systems", PhD thesis, CMU, 2017 (same research)
No mining chip has bought latency with exotic memory No shipped mining chip has used HBM or stacked memory; the Ethash chips used DDR3, GDDR6 and undisclosed types same research; the Ethash chip gain of about 3x in the plan came from bandwidth per watt, not latency
The honest hash is latency-bound on every card measured The hash runs within a few percent of 1/128 of each card's dependent random-read ceiling (5090, 9070 XT, M5 Max) docs/bench-log.md, "the 9070 XT on the eGPU", 5 October 2026 (measured)

Reading for layer 6: an SRAM mirror beats DRAM latency by about 10x per read (a 64 MiB buffer inside the 9070 XT's Infinity Cache chased at 9.2 G loads/s against 2.5 in GDDR6, the same bench-log entry; the 5090's L2 at 5.8x the hash rate of its 1 GiB dataset, M16), which is why the recompute attacker is bound by the 1,024 dependent SRAM reads and the 150,000 integer operations per hash and not by the SRAM's size or price. The cache size decides whether the mirror is one die or several (section 4); it does not decide whether the mirror exists.

9. What is cited, what is approximate, what is owed

Item Status
Bit cells for N7, N5, N3B, N3E, N2, Intel 18A cited (section 3.1)
Shipped cache-die density (AMD V-Cache 64 MB on 41 mm^2 at 7 nm; Graphcore GC200; Groq TSP) cited (section 3.2; V-Cache checked against Tom's Hardware's Hot Chips 33 report, 5 October 2026; the Graphcore and Groq rows are from the coordinator's research and were not re-checked tonight)
Scaling the V-Cache density to other nodes by the bit-cell ratio approximate, stated
Samsung SF2 or SF3 bit cell not found; left out
Array efficiency 0.70 WikiChip's and SemiAnalysis's convention, bracketed by two ISSCC 2025 macros (67 to 80%); a macro figure, used only as the lower bound
Wafer prices approximate, supply-chain reporting, cited
D0 = 0.1 per cm^2, Poisson yield assumption, stated
Node years and the 6% per year density trend past 2026 approximate, extrapolated from cited 2018 to 2025 points
GPU L2 sizes cited (NVIDIA whitepaper); AMD Infinity Cache approximate
Latency citations (MEMSYS 2018, Chang 2017, mining-chip memory types) from the coordinator's research, not re-read tonight
Recompute attacker arithmetic M16, which is itself arithmetic on measured rates, not a chip measurement
Dataset-build slowdown at a larger cache on a GPU owed, a measurement (5090 at a 512 MiB and 1 GiB cache)
The on-die emulation of M16 (inline kernel with a 64 MiB cache inside the 5090's L2) still a PC job (M16)

10. The arithmetic

MiB = 2^20; bits = cache_MiB * MiB * 8
headline_mm2 = cache_MiB * (41 / 64) * (cell_um2 / 0.027)      (V-Cache: 41 mm2 per 64 MiB at N7, scaled by cell)
raw_Mbit_per_mm2 = 1 / cell_um2                                 (1e6 cells per mm2 per um2 of cell)
lower_bound_mm2 = bits / (raw * 0.70 * 1e6)
dies_per_wafer = pi * 150^2 / area - pi * 300 / sqrt(2 * area)
yield = exp(-area_mm2 * 0.001)                                  (D0 = 0.1 per cm2)
cost_per_good_die = wafer_price / (dies * yield); over 830 mm2: k = ceil(area / 830) dies of area / k, cost x k
reticle_GiB = 830 / (mm2 per MiB) / 1024

Run on 5 October 2026 with Python 3 on the M5 Max; the printed tables are the ones above, rounded.