Counter ASIC 2.0 layer 6, revision 2: shipped cache-die density (AMD V-Cache 64 MB on 41 mm^2 at 7 nm, Hot Chips 33) as the headline column beside the bit-cell lower bound; mirror 106 to 164 mm^2 and $30 to $56 per die at 256 MiB; year 0 to 10, hot-table and option tables redone; latency citations section (MEMSYS 2018, Chang 2017, mining-chip memory types)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
eeef9cd056
commit
16cfea717a
1 changed files with 134 additions and 78 deletions
|
|
@ -2,7 +2,13 @@
|
|||
|
||||
5 October 2026 (night), Counter ASIC 2.0 (`docs/plans/counter-asic-2.md`, layer 6), branch `ca2-analysis`. Every figure
|
||||
below is either cited (paper, vendor document, URL, date) or labelled approximate. Nothing here is a measurement of a
|
||||
chip. Numbers in this file were computed with the arithmetic shown; the script is in section 9.
|
||||
chip. Numbers in this file were computed with the arithmetic shown; the script is in section 10.
|
||||
|
||||
Revision 2 (same night): the first draft priced the mirror from bit-cell area times a 0.70 array factor. The
|
||||
coordinator's chip-economics research (sources below) showed that shipped cache-only dies land at about half that
|
||||
density once assist circuits, redundancy, TSVs, power and test are in. Every table now carries two columns: the
|
||||
shipped-product density as the headline and the bit-cell figure as the lower bound. The conclusion did not move; the
|
||||
cost per die rose 2 to 3x.
|
||||
|
||||
## 1. The question
|
||||
|
||||
|
|
@ -29,7 +35,9 @@ design, cache growth is not. The cache is flat at 256 MiB for every year of the
|
|||
closing line names the rule the cache should get ("exceeds what one die can hold, and grows") as a gate 1 decision
|
||||
that has not been taken.
|
||||
|
||||
## 3. SRAM bit cell per node, cited
|
||||
## 3. SRAM density, cited: bit cells per node and shipped cache dies
|
||||
|
||||
### 3.1 Bit cells
|
||||
|
||||
| Node (vendor) | HD 6T bit cell, um^2 | Raw density, Mbit/mm^2 (1/cell) | Year of volume (approximate) | Source |
|
||||
|---|---|---|---|---|
|
||||
|
|
@ -46,17 +54,35 @@ scaling from N5 to N3E (WikiChip IEDM 2022 article above; Tom's Hardware above;
|
|||
https://newsletter.semianalysis.com/p/tsmcs-3nm-conundrum-does-it-even). N2's nanosheet cell recovers 17% (0.021 to
|
||||
0.0175 um^2). So across 2020 to 2026 the HD bit cell shrank once, by 17%.
|
||||
|
||||
Array efficiency (bit cell to macro). The usable density of a macro is below 1/cell because of word-line and
|
||||
bit-line drivers, sense amplifiers, decoders and redundancy. The factor used here is 0.70, WikiChip's convention
|
||||
(their 31.8 Mib/mm^2 for the 0.021 um^2 N3E cell is 1/0.021 x 0.70 in Mib). The two ISSCC 2025 macros bracket it:
|
||||
TSMC N2 38.1 Mb/mm^2 at a 0.0175 um^2 cell is 67%; Intel 18A 38.1 Mb/mm^2 array density and 34.3 Mb/mm^2 for the
|
||||
volume macro at a 0.021 um^2 cell are 80% and 72% (ISSCC 2025 29.2, above). Both lie within 10% of 0.70.
|
||||
Macro density from the bit cell. WikiChip's and SemiAnalysis's convention is bit-cell density times about 0.70 for
|
||||
the assist and periphery overhead (SemiAnalysis, December 2022: TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30%
|
||||
assist overhead; WikiChip's 31.8 Mib/mm^2 for the 0.021 um^2 cell is the same arithmetic). The two ISSCC 2025 macros
|
||||
bracket it: TSMC N2 38.1 Mb/mm^2 at a 0.0175 um^2 cell is 67%; Intel 18A 38.1 Mb/mm^2 array density and 34.3 Mb/mm^2
|
||||
for the volume macro at a 0.021 um^2 cell are 80% and 72%. That is a macro on a test chip. It is the LOWER BOUND on
|
||||
die area, not the die.
|
||||
|
||||
### 3.2 Shipped cache dies (what a whole die of SRAM really holds)
|
||||
|
||||
| Product | SRAM | Die | Node | MB per mm^2 | Source |
|
||||
|---|---|---|---|---|---|
|
||||
| AMD 3D V-Cache (Zen 3 SRAM chiplet) | 64 MB | 41 mm^2 | TSMC 7 nm | 1.56 | AMD at Hot Chips 33, reported by Tom's Hardware, August 2021, https://www.tomshardware.com/news/amd-unveils-more-ryzen-3d-packaging-and-v-cache-details-at-hot-chips ("the 3D V-Cache SRAM measures 41 mm^2", "64 MB of 7 nm SRAM"); the densest cache-only die that has shipped |
|
||||
| Graphcore GC200 (with compute) | 900 MB | 823 mm^2 | 7 nm | 1.09 | coordinator's chip-economics research, 5 October 2026 (vendor figures) |
|
||||
| Groq TSP | 220 MB | 725 mm^2 | 14 nm | 0.30 | same |
|
||||
|
||||
The V-Cache die is a pure SRAM die with its TSVs, redundancy, test and power: 1.56 MB/mm^2 at N7 against the bit-cell
|
||||
figure 37.0 Mbit/mm^2 = 4.6 MB/mm^2 and the 0.70-macro figure 3.2 MB/mm^2. The shipped die is 0.48 of the macro
|
||||
figure. The headline column below scales the V-Cache density to other nodes by the bit-cell ratio (0.027 / cell), an
|
||||
approximation that assumes the periphery and TSV overheads scale with the cell, which they do not fully (so the
|
||||
headline column is itself slightly optimistic for the attacker at N5 and below).
|
||||
|
||||
### 3.3 GPU on-die SRAM, the reticle, wafer prices
|
||||
|
||||
GPU on-die SRAM for scale: the RTX 5090 carries 96 MB of L2 (98,304 KB) on a 750 mm^2 TSMC 4N die with 92.2 billion
|
||||
transistors; the full GB202 has 128 MB; the RTX 4090 had 72 MB and the RTX 3090 6 MB (NVIDIA, "RTX Blackwell GPU
|
||||
Architecture" whitepaper v1.1, appendix table "L2 Cache Size", https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf).
|
||||
At the N5-class cell and 0.70 that L2 is about 24 mm^2 of the 750 (3%), approximate. The RX 9070 XT carries 64 MB
|
||||
of Infinity Cache plus 8 MB of L2 (vendor figures, approximate, bench-log "the 9070 XT on the eGPU").
|
||||
At the V-Cache density scaled to N5 (2.0 MB/mm^2) that L2 is about 48 mm^2 of the 750 (6%), approximate. The
|
||||
RX 9070 XT carries 64 MB of Infinity Cache plus 8 MB of L2 (vendor figures, approximate, bench-log "the 9070 XT on the
|
||||
eGPU").
|
||||
|
||||
Reticle: the EUV field is 26 x 33 mm = 858 mm^2, about 830 mm^2 usable after scribe lanes (SemiAnalysis, "Die Size
|
||||
And Reticle Conundrum", https://newsletter.semianalysis.com/p/die-size-and-reticle-conundrum-cost ; WikiChip "Mask",
|
||||
|
|
@ -66,45 +92,51 @@ Wafer prices (approximate; TSMC publishes none, every figure is supply-chain rep
|
|||
about $20,000 (Silicon Analysts, "Wafer Pricing by Node", September 2026, https://siliconanalysts.com/data/wafer-pricing);
|
||||
N2 about $30,000 (Tom's Hardware, https://www.tomshardware.com/tech-industry/semiconductors/tsmc-could-charge-up-to-usd45-000-for-1-6nm-wafers-rumors-allege-a-50-percent-increase-in-pricing-over-prior-gen-wafers).
|
||||
|
||||
## 4. Die area to mirror the cache, per node
|
||||
## 4. Die area to mirror the cache, per node, two columns
|
||||
|
||||
Area = bits / (raw density x 0.70). The columns are the 256 MiB cache alone, the cache plus each hot-table size of
|
||||
layer 5 (32, 64, 96 MB taken as MiB), and the larger caches of the options in section 6.
|
||||
Headline = V-Cache density (41 mm^2 per 64 MiB at N7) scaled by the bit-cell ratio. Lower bound = bits / (raw
|
||||
density x 0.70). Columns: the 256 MiB cache alone, the cache plus the 96 MB hot table of layer 5 (as MiB), and the
|
||||
larger caches of the options in section 7. Area in mm^2; a figure over 830 is split into the dies shown.
|
||||
|
||||
| Node | Macro Mbit/mm^2 at 0.70 | 256 MiB | 256 + 32 | 256 + 64 | 256 + 96 | 512 MiB | 1 GiB | 2 GiB | 4 GiB |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| N7 | 25.9 | 83 mm^2 | 93 | 104 | 114 | 166 | 331 | 663 | 1,325 (2 dies) |
|
||||
| N5 | 33.3 | 64 | 72 | 81 | 89 | 129 | 258 | 515 | 1,031 (2 dies) |
|
||||
| N3B | 35.2 | 61 | 69 | 76 | 84 | 122 | 244 | 488 | 977 (2 dies) |
|
||||
| N3E, Intel 18A | 33.3 | 64 | 72 | 81 | 89 | 129 | 258 | 515 | 1,031 (2 dies) |
|
||||
| N2 | 40.0 | 54 | 60 | 67 | 74 | 107 | 215 | 429 | 859 (2 dies) |
|
||||
| Node | 256 MiB, headline | 256 MiB, lower bound | 256 + 96, headline | 256 + 96, lower bound | 512 MiB, headline / lower | 1 GiB, headline / lower | 4 GiB, headline / lower |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| N7 | 164 | 83 | 226 | 114 | 328 / 166 | 656 / 331 | 2,624 (4 dies) / 1,325 (2 dies) |
|
||||
| N5 | 128 | 64 | 175 | 89 | 255 / 129 | 510 / 258 | 2,041 (3 dies) / 1,031 (2 dies) |
|
||||
| N3B | 121 | 61 | 166 | 84 | 242 / 122 | 483 / 244 | 1,934 (3 dies) / 977 (2 dies) |
|
||||
| N3E, Intel 18A | 128 | 64 | 175 | 89 | 255 / 129 | 510 / 258 | 2,041 (3 dies) / 1,031 (2 dies) |
|
||||
| N2 | 106 | 54 | 146 | 74 | 213 / 107 | 425 / 215 | 1,701 (3 dies) / 859 (2 dies) |
|
||||
|
||||
One reticle (830 mm^2) holds 2.5 GiB of SRAM at N7, 3.2 GiB at N5, N3E and 18A, 3.9 GiB at N2 (same arithmetic).
|
||||
One reticle (830 mm^2) holds, at the headline density, 1.3 GiB of SRAM at N7, 1.6 GiB at N5, N3E and 18A, 1.9 GiB at
|
||||
N2 (lower-bound column: 2.5, 3.2, 3.9 GiB).
|
||||
|
||||
Against the figures the ledger carries: M16's "100 to 300 mm^2" (low end from a 0.02 um^2 cell with overhead, high
|
||||
end from wafer-scale parts at about 1 MB/mm^2) and the plan's "about 45 mm^2 at a leading node" both bracket the
|
||||
cited 54 to 64 mm^2; the wafer-scale high end is a different efficiency (Cerebras-class arrays sit beside logic) and
|
||||
is not the right number for a pure SRAM die. The right figure for the ledger is 54 to 83 mm^2 depending on node,
|
||||
cited above.
|
||||
end from wafer-scale parts at about 1 MB per mm^2) brackets the headline 106 to 164 mm^2 well; the plan's "about
|
||||
45 mm^2 at a leading node" is below even the lower bound and should be read as the bit-cell area with no overhead.
|
||||
The right figures for the ledger are 106 to 164 mm^2 (shipped density) with 54 to 83 mm^2 as the floor.
|
||||
|
||||
## 5. Cost per good die
|
||||
## 5. Cost per good die, two columns
|
||||
|
||||
Dies per 300 mm wafer by the usual approximation pi x 150^2 / A minus the edge term pi x 300 / sqrt(2A); yield by
|
||||
Poisson exp(-A x D0) with D0 = 0.1 defects per cm^2 (an assumption, approximate; SRAM arrays carry redundancy so
|
||||
real yield is higher, which lowers these costs). Cost per good die = wafer price / (dies x yield). Packaging, test,
|
||||
the logic beside the SRAM and the design (masks at N5 and below run into the tens of millions of dollars,
|
||||
approximate) are not in these numbers; they are per-die silicon only.
|
||||
approximate) are not in these numbers; they are per-die silicon only. Headline / lower bound in each cell.
|
||||
|
||||
| Node, wafer price | 256 MiB | 256 + 96 MiB | 1 GiB | 4 GiB (2 dies) |
|
||||
| Node, wafer price | 256 MiB | 256 + 96 MiB | 1 GiB | 4 GiB |
|
||||
|---|---|---|---|---|
|
||||
| N7, $9,500 | 83 mm^2, 780 dies, yield 0.92, $13 | $19 | 331 mm^2, 177 dies, 0.72, $75 | $456 |
|
||||
| N5, $20,000 | 64 mm^2, 1,014 dies, 0.94, $21 | $30 | 258 mm^2, 233 dies, 0.77, $111 | $621 |
|
||||
| N3E, $20,000 | $21 | $30 | $111 | $621 |
|
||||
| N2, $30,000 | 54 mm^2, 1,226 dies, 0.95, $26 | $37 | 215 mm^2, 284 dies, 0.81, $131 | $696 |
|
||||
| N7, $9,500 | 164 mm^2, 379 dies, yield 0.85: $30 / $13 | $44 / $19 | $224 / $75 | $896 (4 dies) / $456 (2 dies) |
|
||||
| N5, $20,000 | 128 mm^2, 495 dies, 0.88: $46 / $21 | $68 / $30 | $306 / $111 | $1,512 (3 dies) / $621 (2 dies) |
|
||||
| N3B, $20,000 | 121 mm^2, 524 dies, 0.89: $43 / $20 | $63 / $28 | $280 / $103 | $1,371 (3 dies) / $569 (2 dies) |
|
||||
| N3E, 18A, $20,000 | $46 / $21 | $68 / $30 | $306 / $111 | $1,512 / $621 |
|
||||
| N2, $30,000 | 106 mm^2, 600 dies, 0.90: $56 / $26 | $81 / $37 | $343 / $131 | $1,641 (3 dies) / $696 (2 dies) |
|
||||
|
||||
Reading. The silicon for a 256 MiB mirror is tens of dollars per die on any node from N7 up. With the hot table it
|
||||
is still under $40. It was never the SRAM that priced the recompute attacker out; the plan's premise for layer 6
|
||||
("the SRAM mirror stays unaffordable") does not hold for the cache as a mirror and did not hold at genesis either.
|
||||
Reading. The silicon for a 256 MiB mirror is $30 to $56 per die at shipped density (2 to 3x the first draft's
|
||||
figure), under $90 with the hot table. A funded chip programme pays that without noticing: it was never the SRAM
|
||||
that priced the recompute attacker out, and the plan's premise for layer 6 ("the SRAM mirror stays unaffordable")
|
||||
does not hold for the cache as a mirror and did not hold at genesis either. A 1 GiB cache is a 425 to 656 mm^2 die
|
||||
($224 to $343), affordable too; 4 GiB is a 3 to 4 die part at about $900 to $1,600 of silicon, which is a different
|
||||
product but not an impossible one (the attacker's problem at that size is the 1,024 dependent cross-die reads per
|
||||
hash, section 6).
|
||||
|
||||
## 6. What the mirror buys the attacker, year by year
|
||||
|
||||
|
|
@ -112,37 +144,38 @@ From M16 (`docs/analysis/m16-recompute-attacker-2026-10-05.md`): with the cache
|
|||
items per hash at about 1,170 integer operations and 8 dependent 64-byte cache reads each, about 150,000 operations
|
||||
and 1,024 dependent SRAM reads per hash. At a 5090-class integer budget (about 50 T op/s, approximate) that is
|
||||
0.33 Ghash/s against the honest 141 Mhash/s projected for version 2 programs: 2.4x at equal silicon before any
|
||||
fixed-function factor, 3x to 6x with one (approximate). The SRAM is 54 to 83 mm^2 of that chip (7 to 11% of a
|
||||
750 mm^2 die), so the mirror is cheap and the recompute route is bound by integer throughput, not by SRAM.
|
||||
fixed-function factor, 3x to 6x with one (approximate). The SRAM is 106 to 164 mm^2 of that chip at the headline
|
||||
density (14 to 22% of a 750 mm^2 die; the m16 model's 13 to 40% band holds), so the mirror is cheap and the recompute
|
||||
route is bound by integer throughput, not by SRAM.
|
||||
|
||||
The layer 5 hot table changes nothing in that arithmetic: the hot table is read-only and derived from the day key
|
||||
like the cache, so a chip mirrors it in the same SRAM (another 32 to 96 MB, 7 to 24 mm^2 at N5) and reads it at
|
||||
SRAM latency, which is exactly what a GPU's L2 does with it. Layer 5 taxes the DRAM-only chip (the one without SRAM);
|
||||
it does not tax the SRAM chip.
|
||||
like the cache, so a chip mirrors it in the same SRAM (another 32 to 96 MB, 24 to 48 mm^2 at N5 headline) and reads
|
||||
it at SRAM latency, which is exactly what a GPU's L2 does with it. Layer 5 taxes the DRAM-only chip (the one without
|
||||
SRAM); it does not tax the SRAM chip.
|
||||
|
||||
Dataset growth does not touch the recompute attacker: the attacker never holds the dataset. It taxes the
|
||||
partial-store attacker (O-1.6, the time-memory curve, not drawn) and the honest card.
|
||||
|
||||
Year by year under the schedule as it stands (flat 256 MiB), the mirror's area at the best node available that
|
||||
year. Node years are approximate; the density trend from 2018 to 2025 is 37.0 to 57.1 Mbit/mm^2 raw, 1.54x in
|
||||
7 years, about 6% per year, and it came in one step (N2); the extrapolation past 2026 assumes that average holds
|
||||
(approximate, and optimistic for the attacker: A16 and A14 have no disclosed SRAM cell yet).
|
||||
year, headline density. Node years are approximate; the density trend from 2018 to 2025 is 37.0 to 57.1 Mbit/mm^2
|
||||
raw, 1.54x in 7 years, about 6% per year, and it came in one step (N2); the extrapolation past 2026 assumes that
|
||||
average holds (approximate, and optimistic for the attacker: A16 and A14 have no disclosed SRAM cell yet).
|
||||
|
||||
| Year | Calendar (approximate) | Dataset, GiB | Cache (spec) | Best node, raw Mbit/mm^2 | Mirror of the cache, mm^2 | With a 96 MiB hot table, mm^2 | Mirror as a share of a 750 mm^2 die |
|
||||
| Year | Calendar (approximate) | Dataset, GiB | Cache (spec) | Best node | Mirror of the cache, headline (lower bound), mm^2 | With a 96 MiB hot table, headline, mm^2 | Mirror as a share of a 750 mm^2 die |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 0 | 2027 | 2.0 | 256 MiB | N2, 57.1 (cited) | 54 | 74 | 7% |
|
||||
| 1 | 2028 | 2.5 | 256 MiB | N2 or A16, 57 to 61 | 50 to 54 | 69 to 74 | 7% |
|
||||
| 2 | 2029 | 3.0 | 256 MiB | about 64 (trend) | 48 | 66 | 6% |
|
||||
| 3 | 2030 | 3.5 | 256 MiB | about 68 | 45 | 62 | 6% |
|
||||
| 4 | 2031 | 4.0 | 256 MiB | about 72 | 43 | 59 | 6% |
|
||||
| 5 | 2032 | 4.5 | 256 MiB | about 76 | 40 | 55 | 5% |
|
||||
| 6 | 2033 | 5.0 | 256 MiB | about 81 | 38 | 52 | 5% |
|
||||
| 7 | 2034 | 5.5 | 256 MiB | about 86 | 36 | 49 | 5% |
|
||||
| 8 | 2035 | 6.0 | 256 MiB | about 91 | 34 | 46 | 5% |
|
||||
| 9 | 2036 | 6.5 | 256 MiB | about 97 | 32 | 44 | 4% |
|
||||
| 10 | 2037 | 7.0 | 256 MiB | about 102 | 30 | 41 | 4% |
|
||||
| 0 | 2027 | 2.0 | 256 MiB | N2 (cited) | 106 (54) | 146 | 14% |
|
||||
| 1 | 2028 | 2.5 | 256 MiB | N2 or A16 | 103 (52) | 142 | 14% |
|
||||
| 2 | 2029 | 3.0 | 256 MiB | trend | 95 (48) | 130 | 13% |
|
||||
| 3 | 2030 | 3.5 | 256 MiB | trend | 89 (45) | 123 | 12% |
|
||||
| 4 | 2031 | 4.0 | 256 MiB | trend | 84 (43) | 116 | 11% |
|
||||
| 5 | 2032 | 4.5 | 256 MiB | trend | 79 (40) | 109 | 11% |
|
||||
| 6 | 2033 | 5.0 | 256 MiB | trend | 75 (38) | 103 | 10% |
|
||||
| 7 | 2034 | 5.5 | 256 MiB | trend | 71 (36) | 97 | 9% |
|
||||
| 8 | 2035 | 6.0 | 256 MiB | trend | 67 (34) | 92 | 9% |
|
||||
| 9 | 2036 | 6.5 | 256 MiB | trend | 63 (32) | 87 | 8% |
|
||||
| 10 | 2037 | 7.0 | 256 MiB | trend | 59 (30) | 82 | 8% |
|
||||
|
||||
Reading. A flat cache's mirror shrinks from 7% to 4% of a large die over the decade, and a 5090-class consumer GPU
|
||||
Reading. A flat cache's mirror shrinks from 14% to 8% of a large die over the decade, and a 5090-class consumer GPU
|
||||
already carries 96 MB of L2 on one die with the full GB202 at 128 MB; at the 2020 to 2025 pace of GPU L2 growth
|
||||
(6 MB, 72 MB, 96 MB on the three NVIDIA flagships in the whitepaper table) a consumer GPU could hold 256 MiB on die
|
||||
within the decade. The spec's own rule for the cache ("must exceed the largest on-chip cache of any card that
|
||||
|
|
@ -152,8 +185,8 @@ DRAM-latency-bound and the recompute route stays a route only a custom chip can
|
|||
|
||||
## 7. Answer to the layer 6 question, and the options
|
||||
|
||||
Does the flat 256 MiB cache keep the SRAM mirror unaffordable through year 10? No. It is affordable at year 0
|
||||
(tens of dollars of silicon per die, section 5) and gets cheaper. What keeps the recompute attacker near 1x is
|
||||
Does the flat 256 MiB cache keep the SRAM mirror unaffordable through year 10? No. It is affordable at year 0 ($30 to
|
||||
$56 of silicon per die at shipped density, section 5) and gets cheaper. What keeps the recompute attacker near 1x is
|
||||
M16's integer arithmetic and the mixer-cost lever (4x the mixer cost puts the equal-silicon gain at 0.36x, bounded
|
||||
by the CPU verify gate), not the cache size. The cache size does one other job, keeping the cache larger than any
|
||||
GPU's L2, and that job needs growth.
|
||||
|
|
@ -165,22 +198,23 @@ GPU fill: 0.67 ms per 256 MiB on the 5090 (section 1.8.3), linear. The GPU datas
|
|||
5090, section 1.8.3) depends on the dataset size, not the cache size; a larger cache spreads the build's 8 dependent
|
||||
reads per item over more memory, which on a GPU means more of them miss L2 and the build slows by some factor
|
||||
between 1x and the L2-to-DRAM latency ratio, which is a measurement to take (approximate; owed). Mirror area is at N2
|
||||
(cited density), the node of the first years; at the trend's year-10 density divide by about 1.8.
|
||||
headline density (lower bound in brackets), the node of the first years; at the trend's year-10 density divide by
|
||||
about 1.8.
|
||||
|
||||
| Option | Rule | Cache at year 0 / 4 / 10 | Mirror at N2, year 0 / 4 / 10 (mm^2) | Dies at year 10 (830 mm^2 reticle) | Verifier fill, one core, year 0 / 10 | Verifier memory, year 10 | GPU cache fill (5090), year 10 | Keeps the cache above a 96 MB L2 at year 10 | Keeps it above a 256 MB L2 |
|
||||
| Option | Rule | Cache at year 0 / 4 / 10 | Mirror at N2, headline (lower bound), year 0 / 4 / 10, mm^2 | Dies at year 10 (830 mm^2 reticle), headline | Verifier fill, one core, year 0 / 10 | Verifier memory, year 10 | GPU cache fill (5090), year 10 | Keeps the cache above a 96 MB L2 at year 10 | Keeps it above a 256 MB L2 |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| A, as specified | flat 256 MiB | 256 / 256 / 256 MiB | 54 / 54 / 54 | 1 | 0.2 / 0.2 s | 256 MiB | 0.7 ms | yes, 2.7x | no |
|
||||
| B | cache = dataset / 8 (today's ratio) | 256 / 512 / 896 MiB | 54 / 107 / 188 | 1 | 0.2 / 0.7 s | 896 MiB | 2.3 ms | yes, 9.3x | yes, 3.5x |
|
||||
| C | cache doubles when the dataset doubles (the dataset's own clock: year 4, then year 12) | 256 / 512 / 512 MiB | 54 / 107 / 107 | 1 | 0.2 / 0.4 s | 512 MiB | 1.3 ms | yes, 5.3x | yes, 2x |
|
||||
| D | cache = dataset / 4 | 512 / 1,024 / 1,792 MiB | 107 / 215 / 376 | 1 | 0.4 / 1.4 s | 1.75 GiB | 4.7 ms | yes | yes, 7x |
|
||||
| E, one reticle | cache sized so the mirror exceeds one reticle at the node of the day: 4 GiB at N2 (section 4), growing with density | 4 GiB / about 4.5 / about 7 GiB | 859 / 860 / 860 (by construction) | 2 | 3.2 / 5.6 s | 7 GiB | 11 / 19 ms | yes | yes |
|
||||
| A, as specified | flat 256 MiB | 256 / 256 / 256 MiB | 106 (54) / 106 / 106 | 1 | 0.2 / 0.2 s | 256 MiB | 0.7 ms | yes, 2.7x | no |
|
||||
| B | cache = dataset / 8 (today's ratio) | 256 / 512 / 896 MiB | 106 (54) / 213 (107) / 372 (188) | 1 | 0.2 / 0.7 s | 896 MiB | 2.3 ms | yes, 9.3x | yes, 3.5x |
|
||||
| C | cache doubles when the dataset doubles (the dataset's own clock: year 4, then year 12) | 256 / 512 / 512 MiB | 106 (54) / 213 (107) / 213 (107) | 1 | 0.2 / 0.4 s | 512 MiB | 1.3 ms | yes, 5.3x | yes, 2x |
|
||||
| D | cache = dataset / 4 | 512 / 1,024 / 1,792 MiB | 213 (107) / 425 (215) / 744 (376) | 1 | 0.4 / 1.4 s | 1.75 GiB | 4.7 ms | yes | yes, 7x |
|
||||
| E, one reticle | cache sized so the mirror exceeds one reticle at the node of the day: 2 GiB at N2 headline density (section 4; 4 GiB on the lower bound), growing with density | 2 GiB / about 2.3 / about 3.5 GiB | 850 / 850 / 850 (by construction) | 2 | 1.6 / 2.8 s | 3.5 GiB | 5.4 / 9.4 ms | yes | yes |
|
||||
|
||||
Where the working set enters (coordinator's budget: 1 GiB table + hot table + scratch for every resident warp +
|
||||
buffers under 6 GB on an 8 GB card): the cache is not in the miner's working set at hash time (the dataset is built
|
||||
from it once a day and the cache can be dropped or kept), so options A to D do not move that budget; the dataset's own
|
||||
growth does (2 GiB at genesis, 4 GiB at year 4, 7 GiB at year 10, which is past an 8 GB card at about year 8 on its
|
||||
own). Option E's 4 GiB cache would have to be built on the card and dropped, which is fine for a 16 GB card and tight
|
||||
on an 8 GB one at build time (4 GiB cache + 2 GiB dataset + hot table). The per-warp scratch at 170 SMs x 64 warps
|
||||
own). Option E's 2 GiB cache would have to be built on the card and dropped, which is fine for a 16 GB card and tight
|
||||
on an 8 GB one at build time (2 GiB cache + 2 GiB dataset + hot table). The per-warp scratch at 170 SMs x 64 warps
|
||||
(approximate, readwidth) is 340 MB at 32 KB and 1.36 GB at 128 KB per warp; with the 1 GiB table, a 96 MB hot table
|
||||
and buffers that is 1.5 to 2.5 GB at the prototype dataset size, 2.5 to 3.5 GB at the 2 GiB genesis size, inside
|
||||
6 GB either way.
|
||||
|
|
@ -191,37 +225,59 @@ every step, which is the 1.13.3 option (b) argument again), costs the verifier 0
|
|||
more until year 12, and keeps the cache 2x above a 256 MB GPU L2 if one appears. It does not price a chip out; nothing
|
||||
about cache size does (section 5). The lever that does is the mixer cost multiplier of M16, which is the gate 1
|
||||
decision to take beside this one. Option B is the same idea in a smooth form and costs the verifier 0.7 s at year 10.
|
||||
Option E is the only one that makes the mirror a multi-die part and it costs every verifier 3.2 s and 4 GiB at
|
||||
genesis, which fails the spirit of the 10 ms verify gate (the fill is once a day, but a light node joining pays it on
|
||||
every day it syncs across).
|
||||
Option E is the only one that makes the mirror a multi-die part and it costs every verifier 1.6 s and 2 GiB at
|
||||
genesis (at the headline density; the lower-bound density would ask for 4 GiB and 3.2 s), which fails the spirit of
|
||||
the 10 ms verify gate (the fill is once a day, but a light node joining pays it on every day it syncs across).
|
||||
|
||||
Decision for the project lead, at gate 1: A, B, C, D or E above, together with M16's mixer multiplier. Nothing here changes a
|
||||
vector today: the cache size is a prototype value of spec 1.16 and the growth rule would be a new sentence in 1.13.3.
|
||||
|
||||
## 8. What is cited, what is approximate, what is owed
|
||||
## 8. Why the latency bound is the property to lean on (citations behind the plan's rule)
|
||||
|
||||
The plan's "what stays true" paragraph says DRAM latency is the same physics for everyone and bandwidth per watt is
|
||||
what a custom memory chip buys. The sources behind that:
|
||||
|
||||
| Claim | Figure | Source |
|
||||
|---|---|---|
|
||||
| Random-access DRAM latency is the same across memory types | Row cycle time 40 to 48 ns across DDR4, GDDR5 and HBM2 | Li, Reddy and Jacob, "A Performance and Power Comparison of Contemporary DRAM Architectures", MEMSYS 2018 (coordinator's chip-economics research, 5 October 2026) |
|
||||
| Latency does not scale, bandwidth does | DRAM latency improved about 1.3x in two decades while bandwidth improved about 20x | K. Chang, "Understanding and Improving the Latency of DRAM-Based Memory Systems", PhD thesis, CMU, 2017 (same research) |
|
||||
| No mining chip has bought latency with exotic memory | No shipped mining chip has used HBM or stacked memory; the Ethash chips used DDR3, GDDR6 and undisclosed types | same research; the Ethash chip gain of about 3x in the plan came from bandwidth per watt, not latency |
|
||||
| The honest hash is latency-bound on every card measured | The hash runs within a few percent of 1/128 of each card's dependent random-read ceiling (5090, 9070 XT, M5 Max) | `docs/bench-log.md`, "the 9070 XT on the eGPU", 5 October 2026 (measured) |
|
||||
|
||||
Reading for layer 6: an SRAM mirror beats DRAM latency by about 10x per read (a 64 MiB buffer inside the 9070 XT's
|
||||
Infinity Cache chased at 9.2 G loads/s against 2.5 in GDDR6, the same bench-log entry; the 5090's L2 at 5.8x the
|
||||
hash rate of its 1 GiB dataset, M16), which is why the recompute attacker is bound by the 1,024 dependent SRAM reads
|
||||
and the 150,000 integer operations per hash and not by the SRAM's size or price. The cache size decides whether the
|
||||
mirror is one die or several (section 4); it does not decide whether the mirror exists.
|
||||
|
||||
## 9. What is cited, what is approximate, what is owed
|
||||
|
||||
| Item | Status |
|
||||
|---|---|
|
||||
| Bit cells for N7, N5, N3B, N3E, N2, Intel 18A | cited (section 3) |
|
||||
| Bit cells for N7, N5, N3B, N3E, N2, Intel 18A | cited (section 3.1) |
|
||||
| Shipped cache-die density (AMD V-Cache 64 MB on 41 mm^2 at 7 nm; Graphcore GC200; Groq TSP) | cited (section 3.2; V-Cache checked against Tom's Hardware's Hot Chips 33 report, 5 October 2026; the Graphcore and Groq rows are from the coordinator's research and were not re-checked tonight) |
|
||||
| Scaling the V-Cache density to other nodes by the bit-cell ratio | approximate, stated |
|
||||
| Samsung SF2 or SF3 bit cell | not found; left out |
|
||||
| Array efficiency 0.70 | WikiChip's convention, bracketed by two ISSCC 2025 macros (67 to 80%) |
|
||||
| Array efficiency 0.70 | WikiChip's and SemiAnalysis's convention, bracketed by two ISSCC 2025 macros (67 to 80%); a macro figure, used only as the lower bound |
|
||||
| Wafer prices | approximate, supply-chain reporting, cited |
|
||||
| D0 = 0.1 per cm^2, Poisson yield | assumption, stated |
|
||||
| Node years and the 6% per year density trend past 2026 | approximate, extrapolated from cited 2018 to 2025 points |
|
||||
| GPU L2 sizes | cited (NVIDIA whitepaper); AMD Infinity Cache approximate |
|
||||
| Latency citations (MEMSYS 2018, Chang 2017, mining-chip memory types) | from the coordinator's research, not re-read tonight |
|
||||
| Recompute attacker arithmetic | M16, which is itself arithmetic on measured rates, not a chip measurement |
|
||||
| Dataset-build slowdown at a larger cache on a GPU | owed, a measurement (5090 at a 512 MiB and 1 GiB cache) |
|
||||
| The on-die emulation of M16 (inline kernel with a 64 MiB cache inside the 5090's L2) | still a PC job (M16) |
|
||||
|
||||
## 9. The arithmetic
|
||||
## 10. The arithmetic
|
||||
|
||||
```
|
||||
MiB = 2^20; bits = cache_MiB * MiB * 8
|
||||
raw_Mbit_per_mm2 = 1 / cell_um2 (1e6 cells per mm2 per um2 of cell)
|
||||
area_mm2 = bits / (raw * 0.70 * 1e6)
|
||||
headline_mm2 = cache_MiB * (41 / 64) * (cell_um2 / 0.027) (V-Cache: 41 mm2 per 64 MiB at N7, scaled by cell)
|
||||
raw_Mbit_per_mm2 = 1 / cell_um2 (1e6 cells per mm2 per um2 of cell)
|
||||
lower_bound_mm2 = bits / (raw * 0.70 * 1e6)
|
||||
dies_per_wafer = pi * 150^2 / area - pi * 300 / sqrt(2 * area)
|
||||
yield = exp(-area_mm2 * 0.001) (D0 = 0.1 per cm2)
|
||||
cost_per_good_die = wafer_price / (dies * yield)
|
||||
reticle_GiB = 830 * raw * 0.70 * 1e6 / 8 / 2^30
|
||||
yield = exp(-area_mm2 * 0.001) (D0 = 0.1 per cm2)
|
||||
cost_per_good_die = wafer_price / (dies * yield); over 830 mm2: k = ceil(area / 830) dies of area / k, cost x k
|
||||
reticle_GiB = 830 / (mm2 per MiB) / 1024
|
||||
```
|
||||
Run on 5 October 2026 with Python 3 on the M5 Max; the printed tables are the ones above, rounded.
|
||||
|
|
|
|||
Loading…
Reference in a new issue