100 lines
9.4 KiB
Markdown
100 lines
9.4 KiB
Markdown
# The on-die-cache recompute chip against the RTX 5090, class v2 and class v3, everything combined
|
|
|
|
5 October 2026 (night), Counter ASIC 2.0, worker ca2-mixer. The model is M16's
|
|
(`docs/analysis/m16-recompute-attacker-2026-10-05.md`): the strongest chip the plan has priced holds the whole
|
|
cache in SRAM and derives every dataset item instead of reading it, so its cost per hash is item derivations,
|
|
and its rate at a 50 T op/s integer budget (an RTX 5090's, approximate) is `50 T / (ops per hash)`. Nothing here
|
|
is a measurement of a chip; every GPU figure says where it was measured. "Approximate" marks a figure from memory.
|
|
|
|
## 1. Inputs
|
|
|
|
| Input | Value | Source |
|
|
|---|---|---|
|
|
| Items per hash | 128 (one item per load, 128 loads per hash, median 128.00 distinct) | spec 01 sections 1.4.2 and 1.8.5; the 20,000-program census |
|
|
| Integer operations per mixer application | about 130 | spec 01 section 1.8.4 |
|
|
| Mixer applications per item | 9 under v2; 36 under v3 (`m = 4`, `docs/plans/mixer-x4.md`) | `memhard::Shape::mixers_per_item` |
|
|
| Integer operations per item | 1,170 (v2); 4,680 (v3) | 9 x 130; 36 x 130 |
|
|
| Integer operations per hash | 149,760 (v2, "150,000"); 599,040 (v3, "600,000") | 128 x the above |
|
|
| Chip integer budget | 50 T op/s (approximate: 21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock) | M16 section 3 |
|
|
| Fixed-function factor | 3x (approximate, from memory: 2x to 5x is the usual credit for a pipeline with no scheduling or divergence) | M16 section 3 |
|
|
| RTX 5090, version 2 programs, measured | 136.1 MH/s (readwidth, tonight, `docs/plans/read-width.md`, pack w4 on PC 2); 139.7 MH/s (M11, 4 October, `docs/bench-log.md`) | this analysis uses tonight's 136.1 as the denominator and quotes both |
|
|
| RTX 5090 at w16 (16-byte loads), measured | 139.8 MH/s | readwidth table, tonight (the width stays 4 B: w16 closes nothing) |
|
|
| Cache mirror, 256 MiB, N5 headline density | 128 mm^2, $46 per good die (64 mm^2, $21 at the bit-cell lower bound) | `docs/analysis/sram-mirror.md` revision 2, sections 4 and 5 (`ca2-analysis` e6085c6) |
|
|
| Cache mirror plus a 96 MB hot table, N5 headline | 175 mm^2, $68 | same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate) |
|
|
| 512 MiB and 1 GiB mirrors, N5 headline | 255 mm^2 and 510 mm^2; $111 to $306 | same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node) |
|
|
| GPU-class die | 750 mm^2 (the equal-silicon comparison) | M16 section 3 |
|
|
| CPU verifier, one M5 Max core (loaded, load average 5.6; ratios are the measurement) | v2 1.31 to 1.36 ms per unit, x4 1.92 to 1.96 (1.45x), x8 2.79 (2.1x); worst cold 1.58 / 2.04 / 2.94 ms | `docs/plans/mixer-x4.md` section 6.4, 5 October 2026 21:40 UTC |
|
|
|
|
## 2. The rows
|
|
|
|
Chip rate = 50 T op/s / ops per hash. "Bare" = chip rate / 136.1 MH/s. "With the factor" = bare x 3. "Equal
|
|
silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does not get, the M16 convention
|
|
("minus the area the SRAM takes"). SRAM in mm^2 and dollars at the N5 headline density.
|
|
|
|
| Row | Mixer | Ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | With the 3x factor | Equal silicon, SRAM deducted, with the factor |
|
|
|---|---|---|---|---|---|---|---|---|
|
|
| v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x |
|
|
| v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x |
|
|
| x4 (the candidate measured beside v3; not v3) | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x |
|
|
| MEASURED, NOT ADOPTED (layer 5 decided out of v3 on the PC rows, coordinator 21:40 UTC): v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.66x at the Mac's g = 0.93 (126.6 MH/s); 0.70x at the 5090's g = 0.87 (118.4); the 9070 XT's g 0.84 | 1.98x (Mac g), 2.11x (5090 g) | 1.60x, 1.71x |
|
|
| MEASURED, NOT ADOPTED: v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.71x at the Mac's g = 0.87 (118.4 MH/s); 0.73x at the 5090's g = 0.84 (114.3); the 9070 XT's g 0.80 | 2.12x (Mac g), 2.19x (5090 g) | 1.67x, 1.73x |
|
|
| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), no hot table | x4 | 599,040 | 83.5 | 512 MiB | 255 / $111 | 0.61x | 1.84x | 1.21x |
|
|
| v3 at year 12 (cache 1 GiB, dataset 8 GiB) | x4 | 599,040 | 83.5 | 1 GiB | 510 / $306 | 0.61x | 1.84x | 0.59x |
|
|
| **v3: mixer x8** (decided 22:05 UTC under the delegated rule: verify 2.1 ms per unit on one Mac core against the 10 ms gate, the daily 1 GiB build 23 to 77 ms on the 5090 and the 9070 XT) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x |
|
|
| x8 at year 4 | x8 | 1,198,080 | 41.7 | 512 MiB | 255 / $111 | 0.31x | 0.92x | 0.61x |
|
|
|
|
The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op
|
|
weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The
|
|
width rule (4-byte loads kept) changes nothing either: w16 would have moved the honest denominator by 2.7% and the
|
|
chip's cost not at all.
|
|
|
|
Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6
|
|
hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53.
|
|
Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of `sram-mirror.md`
|
|
scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. The honest denominator in the added
|
|
form is the v2 rate times `g`, the card's measured ratio with the hot loads added: on the M5 Max tonight
|
|
`g = 0.93 / 0.87 / 0.83` at 32 / 64 / 96 MiB (the cache agent, relayed by the coordinator at 21:23 UTC;
|
|
`docs/plans/hot-table.md` carries the runs); the 5090's and the 9070 XT's `g` are the PC rows, owed, and until they
|
|
land the row carries the Mac's `g` against the 5090's rate, which is a mixed figure and is marked so. Year 4 and 12 rows: the mirror of
|
|
`sram-mirror.md` section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of
|
|
those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area,
|
|
not a forecast).
|
|
|
|
## 3. The margin, plainly
|
|
|
|
The combined headline row is the mixer row alone (layer 5 is out: the added form costs the 5090 13 to 16 percent
|
|
and the 9070 XT 16 to 20 percent against the 0.97 bar, coordinator 21:40 UTC; the width stays 4 bytes; the era
|
|
draws and the cache growth cost this chip nothing at year 0), and class v3 is x8 (decided 22:05 UTC). The headline:
|
|
**the on-die-cache recompute chip at 50 T op/s reaches 41.7 MH/s against the 5090's 136.1, 0.31x bare, 0.92x with
|
|
the 3x fixed-function factor, 0.76x with the mirror's area deducted: under 1x with the factor, 0.92x, a margin of 8
|
|
percent on the factor (a 3.3x factor reads 1.0x) and of 9 percent on the budget (55 T op/s reads 1.0x).** The x4
|
|
candidate, measured beside it, read 1.84x and 1.53x. The hot-table rows above are kept as measured, not adopted:
|
|
against THIS chip an added hot table is a cost to the honest card and none to the chip, so it would have moved the
|
|
row the wrong way by the card's own `g`. The margin, plainly:
|
|
|
|
- the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x;
|
|
- the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart);
|
|
- the 50 T op/s budget is approximate; a chip at 55 T op/s reads 2.0x;
|
|
- the hot table in the added form lowers the honest denominator by whatever the hot loads cost the GPU (owed from
|
|
the PC rows), which raises the chip's gain by the same share, 1.84x or more if the hot loads are free, higher if
|
|
not; the hot table's only cost to this chip is 16 to 32 mm^2 of die.
|
|
|
|
What keeps it under 1x is the mixer, and nothing else in Counter ASIC 2.0 moves this chip (the scratch at any share
|
|
gave 2.4x, `docs/analysis/scratch-soundness.md` section 3.4; the hot table taxes the DRAM-only chip, not this one;
|
|
the cache growth taxes it only in die area, which is cheap at year 0 and real at year 12). The next levers, in
|
|
order:
|
|
|
|
1. Mixer x16 (the next step of the same lever): 0.16x bare and 0.46x with the factor against 136.1; the verifier
|
|
by the measured increments (+0.63 ms at x4, +1.46 at x8 on the M5 Max core: about +3.1 ms at x16, 3.7 ms per
|
|
unit, 9 ms on a 2.5x slower laptop core, approximate) is at the edge of the 10 ms gate, so a 2019-class laptop
|
|
core measurement (O-1.14) decides it, not this model.
|
|
2. The hot table: adopted or not on the PC rows (`docs/plans/hot-table.md`); in the added form it costs the GPU
|
|
7 to 17 percent on the Mac and the chip die area only, so against this chip it is a lever in the wrong
|
|
direction and against a DRAM-only chip the first lever; if it is adopted, the mixer must carry the extra `1/g`
|
|
(x8 at g = 0.87 reads 1.06x at the equal budget, 0.84x with the SRAM deducted).
|
|
|
|
## 4. What this does not settle
|
|
|
|
The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point
|
|
under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had
|
|
no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM.
|