igneum/docs/analysis/chip-model-v3.md

7.4 KiB

The on-die-cache recompute chip against the RTX 5090, class v2 and class v3, everything combined

5 October 2026 (night), Counter ASIC 2.0, worker ca2-mixer. The model is M16's (docs/analysis/m16-recompute-attacker-2026-10-05.md): the strongest chip the plan has priced holds the whole cache in SRAM and derives every dataset item instead of reading it, so its cost per hash is item derivations, and its rate at a 50 T op/s integer budget (an RTX 5090's, approximate) is 50 T / (ops per hash). Nothing here is a measurement of a chip; every GPU figure says where it was measured. "Approximate" marks a figure from memory.

1. Inputs

Input Value Source
Items per hash 128 (one item per load, 128 loads per hash, median 128.00 distinct) spec 01 sections 1.4.2 and 1.8.5; the 20,000-program census
Integer operations per mixer application about 130 spec 01 section 1.8.4
Mixer applications per item 9 under v2; 36 under v3 (m = 4, docs/plans/mixer-x4.md) memhard::Shape::mixers_per_item
Integer operations per item 1,170 (v2); 4,680 (v3) 9 x 130; 36 x 130
Integer operations per hash 149,760 (v2, "150,000"); 599,040 (v3, "600,000") 128 x the above
Chip integer budget 50 T op/s (approximate: 21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock) M16 section 3
Fixed-function factor 3x (approximate, from memory: 2x to 5x is the usual credit for a pipeline with no scheduling or divergence) M16 section 3
RTX 5090, version 2 programs, measured 136.1 MH/s (readwidth, tonight, docs/plans/read-width.md, pack w4 on PC 2); 139.7 MH/s (M11, 4 October, docs/bench-log.md) this analysis uses tonight's 136.1 as the denominator and quotes both
RTX 5090 at w16 (16-byte loads), measured 139.8 MH/s readwidth table, tonight (the width stays 4 B: w16 closes nothing)
Cache mirror, 256 MiB, N5 headline density 128 mm^2, $46 per good die (64 mm^2, $21 at the bit-cell lower bound) docs/analysis/sram-mirror.md revision 2, sections 4 and 5 (ca2-analysis e6085c6)
Cache mirror plus a 96 MB hot table, N5 headline 175 mm^2, $68 same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate)
512 MiB and 1 GiB mirrors, N5 headline 255 mm^2 and 510 mm^2; $111 to $306 same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node)
GPU-class die 750 mm^2 (the equal-silicon comparison) M16 section 3

2. The rows

Chip rate = 50 T op/s / ops per hash. "Bare" = chip rate / 136.1 MH/s. "With the factor" = bare x 3. "Equal silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does not get, the M16 convention ("minus the area the SRAM takes"). SRAM in mm^2 and dollars at the N5 headline density.

Row Mixer Ops per hash Chip rate at 50 T op/s SRAM the chip holds mm^2 / $ (N5 headline) Bare gain against 136.1 MH/s With the 3x factor Equal silicon, SRAM deducted, with the factor
v2 as shipped (the M16 and scratch-soundness row) x1 149,760 334 MH/s 256 MiB 128 / $46 2.45x (2.39x against 139.7) 7.4x 6.1x
v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) x1 149,760 334 256 MiB 128 / $46 2.39x against 139.8 7.2x 5.9x
v3: mixer x4 x4 599,040 83.5 MH/s 256 MiB 128 / $46 0.61x 1.84x 1.53x
v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads) x4 599,040 (a hot load is one SRAM read, no item) 83.5 288 MiB 144 / $53 0.61x or below (owed: the 5090's added-form rate; the hot loads cost it something, the chip nothing) 1.84x or below 1.49x
v3 plus a 64 MiB hot table, added form x4 599,040 83.5 320 MiB 160 / $61 0.61x or below (owed) 1.84x or below 1.45x
v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), 64 MiB hot table x4 599,040 83.5 576 MiB 287 / $130 0.61x 1.84x 1.14x
v3 at year 12 (cache 1 GiB, dataset 8 GiB), 64 MiB hot table x4 599,040 83.5 1,088 MiB 542 / $330 0.61x 1.84x 0.51x
v3 with the mixer at x8 instead (the next lever, not adopted) x8 1,198,080 41.7 256 MiB 128 / $46 0.31x 0.92x 0.76x

The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The width rule (4-byte loads kept) changes nothing either: w16 would have moved the honest denominator by 2.7% and the chip's cost not at all.

Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6 hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53. Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of sram-mirror.md scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. Year 4 and 12 rows: the mirror of sram-mirror.md section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area, not a forecast).

3. The margin, plainly

The combined headline row reads 1.84x with the 3x factor at an equal integer budget, 1.5x with the SRAM area deducted. The claim is "under 2x", and the margin is thin:

  • the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x;
  • the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart);
  • the 50 T op/s budget is approximate; a chip at 55 T op/s reads 2.0x;
  • the hot table in the added form lowers the honest denominator by whatever the hot loads cost the GPU (owed from the PC rows), which raises the chip's gain by the same share, 1.84x or more if the hot loads are free, higher if not; the hot table's only cost to this chip is 16 to 32 mm^2 of die.

What keeps it under 2x is the mixer, and nothing else in Counter ASIC 2.0 moves this chip (the scratch at any share gave 2.4x, docs/analysis/scratch-soundness.md section 3.4; the hot table taxes the DRAM-only chip, not this one; the cache growth taxes it only in die area, which is cheap at year 0 and real at year 12). The next levers, in order:

  1. Mixer x8: 0.18x bare and 0.55x with the factor in the M16 table (0.31x and 0.92x against 136.1 here; the M16 table's denominator is 229 MH/s); the CPU verifier at 3.3 to 9.6 ms per warp scaled from the version 1 range, at the edge of the 10 ms gate; the measured v3 row of docs/plans/mixer-x4.md section 6 is what to scale from now, and whether a 2019-class laptop core (unmeasured, O-1.14) passes 10 ms is what decides it.
  2. The hot table: adopted or not on the PC rows (docs/plans/hot-table.md); in the added form it costs the GPU nothing it was not already paying in cache misses and the chip die area only, so it is the second lever for the chip only through area and the first against a DRAM-only chip.

4. What this does not settle

The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM.