8.3 KiB
M16: the recompute attacker with the 256 MiB cache on a die, a cost model
5 October 2026 (evening), FUD ledger sweep round 6. Ledger M16, review R3.5, the chip designer's attack 3.
The claim: "put 256 MiB of SRAM on a die and the dataset is never needed: 128 items per hash at about 1,170 integer operations and 8 near-free reads each; integer operations per dollar is where silicon beats a GPU."
This file prices that device from the rules in the specification and the rates measured so far. Nothing here is a measurement of a chip. Every figure says where it comes from; "approximate" marks a figure from memory.
1. What the honest miner pays per hash
| Quantity | Value | Source |
|---|---|---|
| Loads per hash | 128 (16 load slots x 8 iterations), every accepted program | igneum-pow/src/generator.rs (LOAD_SLOTS, ITERATIONS), spec 01 section 1.4.2 |
| Distinct addresses per hash | 120.05 to 128, median 128.00, over 20,000 accepted programs | 20,000-program census, docs/bench-log.md "generator version 2"; docs/analysis/weak-program-census-2026-10-03.md |
| Bytes per load | 4 (one dataset word) | spec 01 section 1.8.5 |
| Dataset | 1 GiB (2^28 words), built once a day from the 256 MiB cache | spec 01 section 1.8.5; CLAUDE.md |
| Honest rate, RTX 5090, 1 GiB dataset | 228.95 Mhash/s, 23.8 G random loads/s, 95.2 GB/s useful | docs/bench-log.md, "3 October 2026, RTX 5090, memory-hard dataset" |
| Honest rate, RTX 5090, 64 MiB dataset (fits the 96 MiB L2) | 1,352 Mhash/s, 5.8x the 1 GiB rate | docs/bench-log.md, RTX 5090 dataset sweep (cited in ledger M1) |
| Dataset build, RTX 5090 | 13.4 ms for 1 GiB (1,253 M items/s); cache fill 0.67 ms | the same entry |
| Projected honest rate for version 2 programs, RTX 5090 | 141 Mhash/s (approximate: 18.0 G distinct loads/s at 128 distinct loads per hash; not yet run) | docs/bench-log.md, weak-program census, "Hash-rate spread" |
The honest hash is 128 dependent random 4-byte reads over a buffer larger than any on-chip cache. The ALU work of the program (64 instructions x 8 iterations per lane) is not what bounds it: the 5.8x step between the 64 MiB and 1 GiB datasets on the same card is the memory system, not the arithmetic.
2. What the recompute attacker pays per hash
The attacker holds the 256 MiB cache (on a die, the premise) and derives each dataset word on demand instead of reading it.
| Quantity | Value | Source |
|---|---|---|
| Items per hash | 128 (one dataset item of 16 words per load; two loads in one item share it, so at most 128 and in the census's accepted population about 128) | spec 01 section 1.8.5, dataset[w] = item(w >> 4)[w AND 15] |
| Mixer applications per item | 9 (ITEM_ROUNDS = 8 dependent cache reads, nine mixer applications) |
spec 01 sections 1.8.4 and 1.8.5 |
| Operations per mixer application | about 130 integer operations (16 xor-add-multiply steps and 8 ChaCha quarter rounds on a 16-word state) | spec 01 section 1.8.4 |
| Operations per item | about 1,170 | 9 x 130 |
| Cache reads per item | 8 dependent 64-byte lines (each address depends on every earlier read) | spec 01 section 1.8.5 |
| Operations per hash | about 150,000 (128 x 1,170) | arithmetic |
| Cache bytes per hash | 65,536 (128 x 8 x 64) in 1,024 dependent reads | arithmetic |
Measured on Apple silicon (the only inline kernel run so far): the inline kernel does 9.48 Mhash/s on the M5 Max at
both a 256 MiB and a 1 GiB dataset, against 94.8 honest at 256 MiB and 45.2 at 1 GiB (proto-metal/MEMHARD.md
section 2.2: 10x and 4.8x slower). The same inline kernel under heavy load on 3 October ran 17x slower than honest at
a 256 MiB dataset (docs/bench-log.md, "R3.26 / M15", M16 note). The Mac's 256 MiB cache sits in DRAM, so its
inline kernel is bound by the 1,024 dependent cache-line reads per hash and does not price a die; it only shows
the kernel exists and is bit-exact.
3. The die, priced in the GPU's own units
To match ONE RTX 5090 at its honest 1 GiB rate the attacker's chip must deliver, per second:
| Need | Value | Arithmetic |
|---|---|---|
| Integer operations | 34 T op/s | 228.95 M hash/s x 150,000 op/hash |
| Cache-line reads | 234 G reads/s, 15 TB/s of SRAM bandwidth | 228.95 M x 1,024 reads x 64 B |
| SRAM | 256 MiB, with the next day's cache under construction beside it | spec 01 (the cache is per day) |
What the GPU itself has (approximate, from memory, for scale): the RTX 5090's integer throughput is about 50 T op/s (21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock), the figure the ledger entry already carries; on-die SRAM at 256 MiB costs about 100 to 300 mm^2 on a current node (the low end from a 0.02 um^2 bit cell with array overhead, the high end from wafer-scale parts at about 1 MB per mm^2), against a 750 mm^2 class GPU die; 15 TB/s of on-die SRAM bandwidth is within what wafer-scale parts quote and is not the bound.
So, at equal silicon and equal integer throughput, the recompute attacker reaches 50 T / 150,000 = about 0.33 Ghash/s:
| Against | Honest 5090 rate | Attacker gain at equal integer budget |
|---|---|---|
| Closed-form and version 1 programs as measured | 229 Mhash/s | 1.5x |
| Version 2 programs as projected (distinct-load bound) | 141 Mhash/s | 2.4x (the figure in the ledger entry) |
Before any chip-versus-GPU efficiency factor. A fixed-function pipeline with no instruction scheduling and no warp divergence is usually credited with 2x to 5x over a GPU on integer work (approximate, from memory; the honest target in the ledger is "under 2x"). Taking 3x: 4.5x to 7x over a 5090 at equal die area, minus the area the SRAM takes (13% to 40% of the die), so about 3x to 6x. That is the exposure as the parameters stand, and it is arithmetic, not a measurement.
4. The lever, and why it costs the honest miner nothing
The attacker's cost is linear in operations per item. The honest miner pays the mixer once per day in the
dataset build (2^26 items x 1,170 op = 78 G op, 13.4 ms on the 5090 as measured) and never per hash. The
verifier pays it per load it checks (spec 01 section 1.11: the CPU verifier derives the distinct items of a warp
from the cache, 0.41 to 1.2 ms per warp with the cache on one M5 Max core, docs/bench-log.md 3 October).
| Mixer cost multiplier m | Operations per hash | Attacker rate at 50 T op/s | Gain against 229 Mhash/s (equal silicon, no efficiency factor) | Gain with a 3x fixed-function factor | Honest daily dataset build, 5090 | CPU verify per warp (scaled from 0.41 to 1.2 ms) |
|---|---|---|---|---|---|---|
| 1 (today, 1,170 op/item) | 150,000 | 0.33 Ghash/s | 1.5x | 4.4x | 13.4 ms | 0.4 to 1.2 ms |
| 2 | 300,000 | 0.17 Ghash/s | 0.73x | 2.2x | 27 ms | 0.8 to 2.4 ms |
| 4 | 600,000 | 0.083 Ghash/s | 0.36x | 1.1x | 54 ms | 1.6 to 4.8 ms |
| 8 | 1,200,000 | 0.042 Ghash/s | 0.18x | 0.55x | 107 ms | 3.3 to 9.6 ms |
| 16 | 2,400,000 | 0.021 Ghash/s | 0.09x | 0.27x | 214 ms | 6.6 to 19 ms |
Reading. The mixer cost is the one parameter that moves the recompute attacker and leaves the honest hash rate untouched; it is bounded by the 10 ms CPU verification gate (ledger M9, spec 1.16), which at today's verify time allows about 8x before the slow core of the gate (a 2019-class laptop core, unmeasured, O-1.14) is at risk. The cache size is the other lever and is linear in SRAM area, which is the cheaper side for the attacker: 256 MiB to 1 GiB moves the die from about 100 to 300 mm^2 to 400 to 1,200 mm^2, which is a multi-die part. Both are prototype values of spec 1.16, fixed at gate 1.
5. What this does not settle
- The inline kernel on NVIDIA at a 64 MiB cache inside the 5090's 96 MiB L2 (the on-die SRAM emulation the ledger entry names) has not run; it is a PC job. It would put a measured point under the "50 T op/s" row: the rate a real integer engine reaches on the actual mixer with the cache in SRAM-class memory.
- The time-memory curve (O-1.6: store a fraction f of the dataset, recompute the rest) is not drawn; the model above is the f = 0 endpoint. A partial-store attacker with HBM instead of SRAM is a different device and may be the cheaper one.
- The mixer has had no cryptanalysis (spec 1.8.4,
MEMHARD.mdsection 3); a shortcut inside the mixer would cut the 1,170 directly. - Monero's seven years do not price this device (ledger C13).
Decision at gate 1 (owner: the project lead): the cache size rule "exceeds what one die can hold, and grows", and the mixer cost multiplier, against the CPU verify gate.