diff --git a/docs/analysis/sram-mirror.md b/docs/analysis/sram-mirror.md index 970239776..e1798fa88 100644 --- a/docs/analysis/sram-mirror.md +++ b/docs/analysis/sram-mirror.md @@ -2,7 +2,13 @@ 5 October 2026 (night), Counter ASIC 2.0 (`docs/plans/counter-asic-2.md`, layer 6), branch `ca2-analysis`. Every figure below is either cited (paper, vendor document, URL, date) or labelled approximate. Nothing here is a measurement of a -chip. Numbers in this file were computed with the arithmetic shown; the script is in section 9. +chip. Numbers in this file were computed with the arithmetic shown; the script is in section 10. + +Revision 2 (same night): the first draft priced the mirror from bit-cell area times a 0.70 array factor. The +coordinator's chip-economics research (sources below) showed that shipped cache-only dies land at about half that +density once assist circuits, redundancy, TSVs, power and test are in. Every table now carries two columns: the +shipped-product density as the headline and the bit-cell figure as the lower bound. The conclusion did not move; the +cost per die rose 2 to 3x. ## 1. The question @@ -29,7 +35,9 @@ design, cache growth is not. The cache is flat at 256 MiB for every year of the closing line names the rule the cache should get ("exceeds what one die can hold, and grows") as a gate 1 decision that has not been taken. -## 3. SRAM bit cell per node, cited +## 3. SRAM density, cited: bit cells per node and shipped cache dies + +### 3.1 Bit cells | Node (vendor) | HD 6T bit cell, um^2 | Raw density, Mbit/mm^2 (1/cell) | Year of volume (approximate) | Source | |---|---|---|---|---| @@ -46,17 +54,35 @@ scaling from N5 to N3E (WikiChip IEDM 2022 article above; Tom's Hardware above; https://newsletter.semianalysis.com/p/tsmcs-3nm-conundrum-does-it-even). N2's nanosheet cell recovers 17% (0.021 to 0.0175 um^2). So across 2020 to 2026 the HD bit cell shrank once, by 17%. -Array efficiency (bit cell to macro). The usable density of a macro is below 1/cell because of word-line and -bit-line drivers, sense amplifiers, decoders and redundancy. The factor used here is 0.70, WikiChip's convention -(their 31.8 Mib/mm^2 for the 0.021 um^2 N3E cell is 1/0.021 x 0.70 in Mib). The two ISSCC 2025 macros bracket it: -TSMC N2 38.1 Mb/mm^2 at a 0.0175 um^2 cell is 67%; Intel 18A 38.1 Mb/mm^2 array density and 34.3 Mb/mm^2 for the -volume macro at a 0.021 um^2 cell are 80% and 72% (ISSCC 2025 29.2, above). Both lie within 10% of 0.70. +Macro density from the bit cell. WikiChip's and SemiAnalysis's convention is bit-cell density times about 0.70 for +the assist and periphery overhead (SemiAnalysis, December 2022: TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30% +assist overhead; WikiChip's 31.8 Mib/mm^2 for the 0.021 um^2 cell is the same arithmetic). The two ISSCC 2025 macros +bracket it: TSMC N2 38.1 Mb/mm^2 at a 0.0175 um^2 cell is 67%; Intel 18A 38.1 Mb/mm^2 array density and 34.3 Mb/mm^2 +for the volume macro at a 0.021 um^2 cell are 80% and 72%. That is a macro on a test chip. It is the LOWER BOUND on +die area, not the die. + +### 3.2 Shipped cache dies (what a whole die of SRAM really holds) + +| Product | SRAM | Die | Node | MB per mm^2 | Source | +|---|---|---|---|---|---| +| AMD 3D V-Cache (Zen 3 SRAM chiplet) | 64 MB | 41 mm^2 | TSMC 7 nm | 1.56 | AMD at Hot Chips 33, reported by Tom's Hardware, August 2021, https://www.tomshardware.com/news/amd-unveils-more-ryzen-3d-packaging-and-v-cache-details-at-hot-chips ("the 3D V-Cache SRAM measures 41 mm^2", "64 MB of 7 nm SRAM"); the densest cache-only die that has shipped | +| Graphcore GC200 (with compute) | 900 MB | 823 mm^2 | 7 nm | 1.09 | coordinator's chip-economics research, 5 October 2026 (vendor figures) | +| Groq TSP | 220 MB | 725 mm^2 | 14 nm | 0.30 | same | + +The V-Cache die is a pure SRAM die with its TSVs, redundancy, test and power: 1.56 MB/mm^2 at N7 against the bit-cell +figure 37.0 Mbit/mm^2 = 4.6 MB/mm^2 and the 0.70-macro figure 3.2 MB/mm^2. The shipped die is 0.48 of the macro +figure. The headline column below scales the V-Cache density to other nodes by the bit-cell ratio (0.027 / cell), an +approximation that assumes the periphery and TSV overheads scale with the cell, which they do not fully (so the +headline column is itself slightly optimistic for the attacker at N5 and below). + +### 3.3 GPU on-die SRAM, the reticle, wafer prices GPU on-die SRAM for scale: the RTX 5090 carries 96 MB of L2 (98,304 KB) on a 750 mm^2 TSMC 4N die with 92.2 billion transistors; the full GB202 has 128 MB; the RTX 4090 had 72 MB and the RTX 3090 6 MB (NVIDIA, "RTX Blackwell GPU Architecture" whitepaper v1.1, appendix table "L2 Cache Size", https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf). -At the N5-class cell and 0.70 that L2 is about 24 mm^2 of the 750 (3%), approximate. The RX 9070 XT carries 64 MB -of Infinity Cache plus 8 MB of L2 (vendor figures, approximate, bench-log "the 9070 XT on the eGPU"). +At the V-Cache density scaled to N5 (2.0 MB/mm^2) that L2 is about 48 mm^2 of the 750 (6%), approximate. The +RX 9070 XT carries 64 MB of Infinity Cache plus 8 MB of L2 (vendor figures, approximate, bench-log "the 9070 XT on the +eGPU"). Reticle: the EUV field is 26 x 33 mm = 858 mm^2, about 830 mm^2 usable after scribe lanes (SemiAnalysis, "Die Size And Reticle Conundrum", https://newsletter.semianalysis.com/p/die-size-and-reticle-conundrum-cost ; WikiChip "Mask", @@ -66,45 +92,51 @@ Wafer prices (approximate; TSMC publishes none, every figure is supply-chain rep about $20,000 (Silicon Analysts, "Wafer Pricing by Node", September 2026, https://siliconanalysts.com/data/wafer-pricing); N2 about $30,000 (Tom's Hardware, https://www.tomshardware.com/tech-industry/semiconductors/tsmc-could-charge-up-to-usd45-000-for-1-6nm-wafers-rumors-allege-a-50-percent-increase-in-pricing-over-prior-gen-wafers). -## 4. Die area to mirror the cache, per node +## 4. Die area to mirror the cache, per node, two columns -Area = bits / (raw density x 0.70). The columns are the 256 MiB cache alone, the cache plus each hot-table size of -layer 5 (32, 64, 96 MB taken as MiB), and the larger caches of the options in section 6. +Headline = V-Cache density (41 mm^2 per 64 MiB at N7) scaled by the bit-cell ratio. Lower bound = bits / (raw +density x 0.70). Columns: the 256 MiB cache alone, the cache plus the 96 MB hot table of layer 5 (as MiB), and the +larger caches of the options in section 7. Area in mm^2; a figure over 830 is split into the dies shown. -| Node | Macro Mbit/mm^2 at 0.70 | 256 MiB | 256 + 32 | 256 + 64 | 256 + 96 | 512 MiB | 1 GiB | 2 GiB | 4 GiB | -|---|---|---|---|---|---|---|---|---|---| -| N7 | 25.9 | 83 mm^2 | 93 | 104 | 114 | 166 | 331 | 663 | 1,325 (2 dies) | -| N5 | 33.3 | 64 | 72 | 81 | 89 | 129 | 258 | 515 | 1,031 (2 dies) | -| N3B | 35.2 | 61 | 69 | 76 | 84 | 122 | 244 | 488 | 977 (2 dies) | -| N3E, Intel 18A | 33.3 | 64 | 72 | 81 | 89 | 129 | 258 | 515 | 1,031 (2 dies) | -| N2 | 40.0 | 54 | 60 | 67 | 74 | 107 | 215 | 429 | 859 (2 dies) | +| Node | 256 MiB, headline | 256 MiB, lower bound | 256 + 96, headline | 256 + 96, lower bound | 512 MiB, headline / lower | 1 GiB, headline / lower | 4 GiB, headline / lower | +|---|---|---|---|---|---|---|---| +| N7 | 164 | 83 | 226 | 114 | 328 / 166 | 656 / 331 | 2,624 (4 dies) / 1,325 (2 dies) | +| N5 | 128 | 64 | 175 | 89 | 255 / 129 | 510 / 258 | 2,041 (3 dies) / 1,031 (2 dies) | +| N3B | 121 | 61 | 166 | 84 | 242 / 122 | 483 / 244 | 1,934 (3 dies) / 977 (2 dies) | +| N3E, Intel 18A | 128 | 64 | 175 | 89 | 255 / 129 | 510 / 258 | 2,041 (3 dies) / 1,031 (2 dies) | +| N2 | 106 | 54 | 146 | 74 | 213 / 107 | 425 / 215 | 1,701 (3 dies) / 859 (2 dies) | -One reticle (830 mm^2) holds 2.5 GiB of SRAM at N7, 3.2 GiB at N5, N3E and 18A, 3.9 GiB at N2 (same arithmetic). +One reticle (830 mm^2) holds, at the headline density, 1.3 GiB of SRAM at N7, 1.6 GiB at N5, N3E and 18A, 1.9 GiB at +N2 (lower-bound column: 2.5, 3.2, 3.9 GiB). Against the figures the ledger carries: M16's "100 to 300 mm^2" (low end from a 0.02 um^2 cell with overhead, high -end from wafer-scale parts at about 1 MB/mm^2) and the plan's "about 45 mm^2 at a leading node" both bracket the -cited 54 to 64 mm^2; the wafer-scale high end is a different efficiency (Cerebras-class arrays sit beside logic) and -is not the right number for a pure SRAM die. The right figure for the ledger is 54 to 83 mm^2 depending on node, -cited above. +end from wafer-scale parts at about 1 MB per mm^2) brackets the headline 106 to 164 mm^2 well; the plan's "about +45 mm^2 at a leading node" is below even the lower bound and should be read as the bit-cell area with no overhead. +The right figures for the ledger are 106 to 164 mm^2 (shipped density) with 54 to 83 mm^2 as the floor. -## 5. Cost per good die +## 5. Cost per good die, two columns Dies per 300 mm wafer by the usual approximation pi x 150^2 / A minus the edge term pi x 300 / sqrt(2A); yield by Poisson exp(-A x D0) with D0 = 0.1 defects per cm^2 (an assumption, approximate; SRAM arrays carry redundancy so real yield is higher, which lowers these costs). Cost per good die = wafer price / (dies x yield). Packaging, test, the logic beside the SRAM and the design (masks at N5 and below run into the tens of millions of dollars, -approximate) are not in these numbers; they are per-die silicon only. +approximate) are not in these numbers; they are per-die silicon only. Headline / lower bound in each cell. -| Node, wafer price | 256 MiB | 256 + 96 MiB | 1 GiB | 4 GiB (2 dies) | +| Node, wafer price | 256 MiB | 256 + 96 MiB | 1 GiB | 4 GiB | |---|---|---|---|---| -| N7, $9,500 | 83 mm^2, 780 dies, yield 0.92, $13 | $19 | 331 mm^2, 177 dies, 0.72, $75 | $456 | -| N5, $20,000 | 64 mm^2, 1,014 dies, 0.94, $21 | $30 | 258 mm^2, 233 dies, 0.77, $111 | $621 | -| N3E, $20,000 | $21 | $30 | $111 | $621 | -| N2, $30,000 | 54 mm^2, 1,226 dies, 0.95, $26 | $37 | 215 mm^2, 284 dies, 0.81, $131 | $696 | +| N7, $9,500 | 164 mm^2, 379 dies, yield 0.85: $30 / $13 | $44 / $19 | $224 / $75 | $896 (4 dies) / $456 (2 dies) | +| N5, $20,000 | 128 mm^2, 495 dies, 0.88: $46 / $21 | $68 / $30 | $306 / $111 | $1,512 (3 dies) / $621 (2 dies) | +| N3B, $20,000 | 121 mm^2, 524 dies, 0.89: $43 / $20 | $63 / $28 | $280 / $103 | $1,371 (3 dies) / $569 (2 dies) | +| N3E, 18A, $20,000 | $46 / $21 | $68 / $30 | $306 / $111 | $1,512 / $621 | +| N2, $30,000 | 106 mm^2, 600 dies, 0.90: $56 / $26 | $81 / $37 | $343 / $131 | $1,641 (3 dies) / $696 (2 dies) | -Reading. The silicon for a 256 MiB mirror is tens of dollars per die on any node from N7 up. With the hot table it -is still under $40. It was never the SRAM that priced the recompute attacker out; the plan's premise for layer 6 -("the SRAM mirror stays unaffordable") does not hold for the cache as a mirror and did not hold at genesis either. +Reading. The silicon for a 256 MiB mirror is $30 to $56 per die at shipped density (2 to 3x the first draft's +figure), under $90 with the hot table. A funded chip programme pays that without noticing: it was never the SRAM +that priced the recompute attacker out, and the plan's premise for layer 6 ("the SRAM mirror stays unaffordable") +does not hold for the cache as a mirror and did not hold at genesis either. A 1 GiB cache is a 425 to 656 mm^2 die +($224 to $343), affordable too; 4 GiB is a 3 to 4 die part at about $900 to $1,600 of silicon, which is a different +product but not an impossible one (the attacker's problem at that size is the 1,024 dependent cross-die reads per +hash, section 6). ## 6. What the mirror buys the attacker, year by year @@ -112,37 +144,38 @@ From M16 (`docs/analysis/m16-recompute-attacker-2026-10-05.md`): with the cache items per hash at about 1,170 integer operations and 8 dependent 64-byte cache reads each, about 150,000 operations and 1,024 dependent SRAM reads per hash. At a 5090-class integer budget (about 50 T op/s, approximate) that is 0.33 Ghash/s against the honest 141 Mhash/s projected for version 2 programs: 2.4x at equal silicon before any -fixed-function factor, 3x to 6x with one (approximate). The SRAM is 54 to 83 mm^2 of that chip (7 to 11% of a -750 mm^2 die), so the mirror is cheap and the recompute route is bound by integer throughput, not by SRAM. +fixed-function factor, 3x to 6x with one (approximate). The SRAM is 106 to 164 mm^2 of that chip at the headline +density (14 to 22% of a 750 mm^2 die; the m16 model's 13 to 40% band holds), so the mirror is cheap and the recompute +route is bound by integer throughput, not by SRAM. The layer 5 hot table changes nothing in that arithmetic: the hot table is read-only and derived from the day key -like the cache, so a chip mirrors it in the same SRAM (another 32 to 96 MB, 7 to 24 mm^2 at N5) and reads it at -SRAM latency, which is exactly what a GPU's L2 does with it. Layer 5 taxes the DRAM-only chip (the one without SRAM); -it does not tax the SRAM chip. +like the cache, so a chip mirrors it in the same SRAM (another 32 to 96 MB, 24 to 48 mm^2 at N5 headline) and reads +it at SRAM latency, which is exactly what a GPU's L2 does with it. Layer 5 taxes the DRAM-only chip (the one without +SRAM); it does not tax the SRAM chip. Dataset growth does not touch the recompute attacker: the attacker never holds the dataset. It taxes the partial-store attacker (O-1.6, the time-memory curve, not drawn) and the honest card. Year by year under the schedule as it stands (flat 256 MiB), the mirror's area at the best node available that -year. Node years are approximate; the density trend from 2018 to 2025 is 37.0 to 57.1 Mbit/mm^2 raw, 1.54x in -7 years, about 6% per year, and it came in one step (N2); the extrapolation past 2026 assumes that average holds -(approximate, and optimistic for the attacker: A16 and A14 have no disclosed SRAM cell yet). +year, headline density. Node years are approximate; the density trend from 2018 to 2025 is 37.0 to 57.1 Mbit/mm^2 +raw, 1.54x in 7 years, about 6% per year, and it came in one step (N2); the extrapolation past 2026 assumes that +average holds (approximate, and optimistic for the attacker: A16 and A14 have no disclosed SRAM cell yet). -| Year | Calendar (approximate) | Dataset, GiB | Cache (spec) | Best node, raw Mbit/mm^2 | Mirror of the cache, mm^2 | With a 96 MiB hot table, mm^2 | Mirror as a share of a 750 mm^2 die | +| Year | Calendar (approximate) | Dataset, GiB | Cache (spec) | Best node | Mirror of the cache, headline (lower bound), mm^2 | With a 96 MiB hot table, headline, mm^2 | Mirror as a share of a 750 mm^2 die | |---|---|---|---|---|---|---|---| -| 0 | 2027 | 2.0 | 256 MiB | N2, 57.1 (cited) | 54 | 74 | 7% | -| 1 | 2028 | 2.5 | 256 MiB | N2 or A16, 57 to 61 | 50 to 54 | 69 to 74 | 7% | -| 2 | 2029 | 3.0 | 256 MiB | about 64 (trend) | 48 | 66 | 6% | -| 3 | 2030 | 3.5 | 256 MiB | about 68 | 45 | 62 | 6% | -| 4 | 2031 | 4.0 | 256 MiB | about 72 | 43 | 59 | 6% | -| 5 | 2032 | 4.5 | 256 MiB | about 76 | 40 | 55 | 5% | -| 6 | 2033 | 5.0 | 256 MiB | about 81 | 38 | 52 | 5% | -| 7 | 2034 | 5.5 | 256 MiB | about 86 | 36 | 49 | 5% | -| 8 | 2035 | 6.0 | 256 MiB | about 91 | 34 | 46 | 5% | -| 9 | 2036 | 6.5 | 256 MiB | about 97 | 32 | 44 | 4% | -| 10 | 2037 | 7.0 | 256 MiB | about 102 | 30 | 41 | 4% | +| 0 | 2027 | 2.0 | 256 MiB | N2 (cited) | 106 (54) | 146 | 14% | +| 1 | 2028 | 2.5 | 256 MiB | N2 or A16 | 103 (52) | 142 | 14% | +| 2 | 2029 | 3.0 | 256 MiB | trend | 95 (48) | 130 | 13% | +| 3 | 2030 | 3.5 | 256 MiB | trend | 89 (45) | 123 | 12% | +| 4 | 2031 | 4.0 | 256 MiB | trend | 84 (43) | 116 | 11% | +| 5 | 2032 | 4.5 | 256 MiB | trend | 79 (40) | 109 | 11% | +| 6 | 2033 | 5.0 | 256 MiB | trend | 75 (38) | 103 | 10% | +| 7 | 2034 | 5.5 | 256 MiB | trend | 71 (36) | 97 | 9% | +| 8 | 2035 | 6.0 | 256 MiB | trend | 67 (34) | 92 | 9% | +| 9 | 2036 | 6.5 | 256 MiB | trend | 63 (32) | 87 | 8% | +| 10 | 2037 | 7.0 | 256 MiB | trend | 59 (30) | 82 | 8% | -Reading. A flat cache's mirror shrinks from 7% to 4% of a large die over the decade, and a 5090-class consumer GPU +Reading. A flat cache's mirror shrinks from 14% to 8% of a large die over the decade, and a 5090-class consumer GPU already carries 96 MB of L2 on one die with the full GB202 at 128 MB; at the 2020 to 2025 pace of GPU L2 growth (6 MB, 72 MB, 96 MB on the three NVIDIA flagships in the whitepaper table) a consumer GPU could hold 256 MiB on die within the decade. The spec's own rule for the cache ("must exceed the largest on-chip cache of any card that @@ -152,8 +185,8 @@ DRAM-latency-bound and the recompute route stays a route only a custom chip can ## 7. Answer to the layer 6 question, and the options -Does the flat 256 MiB cache keep the SRAM mirror unaffordable through year 10? No. It is affordable at year 0 -(tens of dollars of silicon per die, section 5) and gets cheaper. What keeps the recompute attacker near 1x is +Does the flat 256 MiB cache keep the SRAM mirror unaffordable through year 10? No. It is affordable at year 0 ($30 to +$56 of silicon per die at shipped density, section 5) and gets cheaper. What keeps the recompute attacker near 1x is M16's integer arithmetic and the mixer-cost lever (4x the mixer cost puts the equal-silicon gain at 0.36x, bounded by the CPU verify gate), not the cache size. The cache size does one other job, keeping the cache larger than any GPU's L2, and that job needs growth. @@ -165,22 +198,23 @@ GPU fill: 0.67 ms per 256 MiB on the 5090 (section 1.8.3), linear. The GPU datas 5090, section 1.8.3) depends on the dataset size, not the cache size; a larger cache spreads the build's 8 dependent reads per item over more memory, which on a GPU means more of them miss L2 and the build slows by some factor between 1x and the L2-to-DRAM latency ratio, which is a measurement to take (approximate; owed). Mirror area is at N2 -(cited density), the node of the first years; at the trend's year-10 density divide by about 1.8. +headline density (lower bound in brackets), the node of the first years; at the trend's year-10 density divide by +about 1.8. -| Option | Rule | Cache at year 0 / 4 / 10 | Mirror at N2, year 0 / 4 / 10 (mm^2) | Dies at year 10 (830 mm^2 reticle) | Verifier fill, one core, year 0 / 10 | Verifier memory, year 10 | GPU cache fill (5090), year 10 | Keeps the cache above a 96 MB L2 at year 10 | Keeps it above a 256 MB L2 | +| Option | Rule | Cache at year 0 / 4 / 10 | Mirror at N2, headline (lower bound), year 0 / 4 / 10, mm^2 | Dies at year 10 (830 mm^2 reticle), headline | Verifier fill, one core, year 0 / 10 | Verifier memory, year 10 | GPU cache fill (5090), year 10 | Keeps the cache above a 96 MB L2 at year 10 | Keeps it above a 256 MB L2 | |---|---|---|---|---|---|---|---|---|---| -| A, as specified | flat 256 MiB | 256 / 256 / 256 MiB | 54 / 54 / 54 | 1 | 0.2 / 0.2 s | 256 MiB | 0.7 ms | yes, 2.7x | no | -| B | cache = dataset / 8 (today's ratio) | 256 / 512 / 896 MiB | 54 / 107 / 188 | 1 | 0.2 / 0.7 s | 896 MiB | 2.3 ms | yes, 9.3x | yes, 3.5x | -| C | cache doubles when the dataset doubles (the dataset's own clock: year 4, then year 12) | 256 / 512 / 512 MiB | 54 / 107 / 107 | 1 | 0.2 / 0.4 s | 512 MiB | 1.3 ms | yes, 5.3x | yes, 2x | -| D | cache = dataset / 4 | 512 / 1,024 / 1,792 MiB | 107 / 215 / 376 | 1 | 0.4 / 1.4 s | 1.75 GiB | 4.7 ms | yes | yes, 7x | -| E, one reticle | cache sized so the mirror exceeds one reticle at the node of the day: 4 GiB at N2 (section 4), growing with density | 4 GiB / about 4.5 / about 7 GiB | 859 / 860 / 860 (by construction) | 2 | 3.2 / 5.6 s | 7 GiB | 11 / 19 ms | yes | yes | +| A, as specified | flat 256 MiB | 256 / 256 / 256 MiB | 106 (54) / 106 / 106 | 1 | 0.2 / 0.2 s | 256 MiB | 0.7 ms | yes, 2.7x | no | +| B | cache = dataset / 8 (today's ratio) | 256 / 512 / 896 MiB | 106 (54) / 213 (107) / 372 (188) | 1 | 0.2 / 0.7 s | 896 MiB | 2.3 ms | yes, 9.3x | yes, 3.5x | +| C | cache doubles when the dataset doubles (the dataset's own clock: year 4, then year 12) | 256 / 512 / 512 MiB | 106 (54) / 213 (107) / 213 (107) | 1 | 0.2 / 0.4 s | 512 MiB | 1.3 ms | yes, 5.3x | yes, 2x | +| D | cache = dataset / 4 | 512 / 1,024 / 1,792 MiB | 213 (107) / 425 (215) / 744 (376) | 1 | 0.4 / 1.4 s | 1.75 GiB | 4.7 ms | yes | yes, 7x | +| E, one reticle | cache sized so the mirror exceeds one reticle at the node of the day: 2 GiB at N2 headline density (section 4; 4 GiB on the lower bound), growing with density | 2 GiB / about 2.3 / about 3.5 GiB | 850 / 850 / 850 (by construction) | 2 | 1.6 / 2.8 s | 3.5 GiB | 5.4 / 9.4 ms | yes | yes | Where the working set enters (coordinator's budget: 1 GiB table + hot table + scratch for every resident warp + buffers under 6 GB on an 8 GB card): the cache is not in the miner's working set at hash time (the dataset is built from it once a day and the cache can be dropped or kept), so options A to D do not move that budget; the dataset's own growth does (2 GiB at genesis, 4 GiB at year 4, 7 GiB at year 10, which is past an 8 GB card at about year 8 on its -own). Option E's 4 GiB cache would have to be built on the card and dropped, which is fine for a 16 GB card and tight -on an 8 GB one at build time (4 GiB cache + 2 GiB dataset + hot table). The per-warp scratch at 170 SMs x 64 warps +own). Option E's 2 GiB cache would have to be built on the card and dropped, which is fine for a 16 GB card and tight +on an 8 GB one at build time (2 GiB cache + 2 GiB dataset + hot table). The per-warp scratch at 170 SMs x 64 warps (approximate, readwidth) is 340 MB at 32 KB and 1.36 GB at 128 KB per warp; with the 1 GiB table, a 96 MB hot table and buffers that is 1.5 to 2.5 GB at the prototype dataset size, 2.5 to 3.5 GB at the 2 GiB genesis size, inside 6 GB either way. @@ -191,37 +225,59 @@ every step, which is the 1.13.3 option (b) argument again), costs the verifier 0 more until year 12, and keeps the cache 2x above a 256 MB GPU L2 if one appears. It does not price a chip out; nothing about cache size does (section 5). The lever that does is the mixer cost multiplier of M16, which is the gate 1 decision to take beside this one. Option B is the same idea in a smooth form and costs the verifier 0.7 s at year 10. -Option E is the only one that makes the mirror a multi-die part and it costs every verifier 3.2 s and 4 GiB at -genesis, which fails the spirit of the 10 ms verify gate (the fill is once a day, but a light node joining pays it on -every day it syncs across). +Option E is the only one that makes the mirror a multi-die part and it costs every verifier 1.6 s and 2 GiB at +genesis (at the headline density; the lower-bound density would ask for 4 GiB and 3.2 s), which fails the spirit of +the 10 ms verify gate (the fill is once a day, but a light node joining pays it on every day it syncs across). Decision for Josh, at gate 1: A, B, C, D or E above, together with M16's mixer multiplier. Nothing here changes a vector today: the cache size is a prototype value of spec 1.16 and the growth rule would be a new sentence in 1.13.3. -## 8. What is cited, what is approximate, what is owed +## 8. Why the latency bound is the property to lean on (citations behind the plan's rule) + +The plan's "what stays true" paragraph says DRAM latency is the same physics for everyone and bandwidth per watt is +what a custom memory chip buys. The sources behind that: + +| Claim | Figure | Source | +|---|---|---| +| Random-access DRAM latency is the same across memory types | Row cycle time 40 to 48 ns across DDR4, GDDR5 and HBM2 | Li, Reddy and Jacob, "A Performance and Power Comparison of Contemporary DRAM Architectures", MEMSYS 2018 (coordinator's chip-economics research, 5 October 2026) | +| Latency does not scale, bandwidth does | DRAM latency improved about 1.3x in two decades while bandwidth improved about 20x | K. Chang, "Understanding and Improving the Latency of DRAM-Based Memory Systems", PhD thesis, CMU, 2017 (same research) | +| No mining chip has bought latency with exotic memory | No shipped mining chip has used HBM or stacked memory; the Ethash chips used DDR3, GDDR6 and undisclosed types | same research; the Ethash chip gain of about 3x in the plan came from bandwidth per watt, not latency | +| The honest hash is latency-bound on every card measured | The hash runs within a few percent of 1/128 of each card's dependent random-read ceiling (5090, 9070 XT, M5 Max) | `docs/bench-log.md`, "the 9070 XT on the eGPU", 5 October 2026 (measured) | + +Reading for layer 6: an SRAM mirror beats DRAM latency by about 10x per read (a 64 MiB buffer inside the 9070 XT's +Infinity Cache chased at 9.2 G loads/s against 2.5 in GDDR6, the same bench-log entry; the 5090's L2 at 5.8x the +hash rate of its 1 GiB dataset, M16), which is why the recompute attacker is bound by the 1,024 dependent SRAM reads +and the 150,000 integer operations per hash and not by the SRAM's size or price. The cache size decides whether the +mirror is one die or several (section 4); it does not decide whether the mirror exists. + +## 9. What is cited, what is approximate, what is owed | Item | Status | |---|---| -| Bit cells for N7, N5, N3B, N3E, N2, Intel 18A | cited (section 3) | +| Bit cells for N7, N5, N3B, N3E, N2, Intel 18A | cited (section 3.1) | +| Shipped cache-die density (AMD V-Cache 64 MB on 41 mm^2 at 7 nm; Graphcore GC200; Groq TSP) | cited (section 3.2; V-Cache checked against Tom's Hardware's Hot Chips 33 report, 5 October 2026; the Graphcore and Groq rows are from the coordinator's research and were not re-checked tonight) | +| Scaling the V-Cache density to other nodes by the bit-cell ratio | approximate, stated | | Samsung SF2 or SF3 bit cell | not found; left out | -| Array efficiency 0.70 | WikiChip's convention, bracketed by two ISSCC 2025 macros (67 to 80%) | +| Array efficiency 0.70 | WikiChip's and SemiAnalysis's convention, bracketed by two ISSCC 2025 macros (67 to 80%); a macro figure, used only as the lower bound | | Wafer prices | approximate, supply-chain reporting, cited | | D0 = 0.1 per cm^2, Poisson yield | assumption, stated | | Node years and the 6% per year density trend past 2026 | approximate, extrapolated from cited 2018 to 2025 points | | GPU L2 sizes | cited (NVIDIA whitepaper); AMD Infinity Cache approximate | +| Latency citations (MEMSYS 2018, Chang 2017, mining-chip memory types) | from the coordinator's research, not re-read tonight | | Recompute attacker arithmetic | M16, which is itself arithmetic on measured rates, not a chip measurement | | Dataset-build slowdown at a larger cache on a GPU | owed, a measurement (5090 at a 512 MiB and 1 GiB cache) | | The on-die emulation of M16 (inline kernel with a 64 MiB cache inside the 5090's L2) | still a PC job (M16) | -## 9. The arithmetic +## 10. The arithmetic ``` MiB = 2^20; bits = cache_MiB * MiB * 8 -raw_Mbit_per_mm2 = 1 / cell_um2 (1e6 cells per mm2 per um2 of cell) -area_mm2 = bits / (raw * 0.70 * 1e6) +headline_mm2 = cache_MiB * (41 / 64) * (cell_um2 / 0.027) (V-Cache: 41 mm2 per 64 MiB at N7, scaled by cell) +raw_Mbit_per_mm2 = 1 / cell_um2 (1e6 cells per mm2 per um2 of cell) +lower_bound_mm2 = bits / (raw * 0.70 * 1e6) dies_per_wafer = pi * 150^2 / area - pi * 300 / sqrt(2 * area) -yield = exp(-area_mm2 * 0.001) (D0 = 0.1 per cm2) -cost_per_good_die = wafer_price / (dies * yield) -reticle_GiB = 830 * raw * 0.70 * 1e6 / 8 / 2^30 +yield = exp(-area_mm2 * 0.001) (D0 = 0.1 per cm2) +cost_per_good_die = wafer_price / (dies * yield); over 830 mm2: k = ceil(area / 830) dies of area / k, cost x k +reticle_GiB = 830 / (mm2 per MiB) / 1024 ``` Run on 5 October 2026 with Python 3 on the M5 Max; the printed tables are the ones above, rounded.