The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh). The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge. Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
45 KiB
The on-die-cache recompute chip against the RTX 5090, class v2 and class v3, everything combined
5 October 2026 (night), Counter ASIC 2.0, worker ca2-mixer. The model is M16's
(docs/analysis/m16-recompute-attacker-2026-10-05.md): the strongest chip the plan has priced holds the whole
cache in SRAM and derives every dataset item instead of reading it, so its cost per hash is item derivations,
and its rate at a 50 T op/s integer budget (an RTX 5090's, approximate) is 50 T / (ops per hash). Nothing here
is a measurement of a chip; every GPU figure says where it was measured. "Approximate" marks a figure from memory.
1. Inputs
| Input | Value | Source |
|---|---|---|
| Items per hash | 128 (one item per load, 128 loads per hash, median 128.00 distinct) | spec 01 sections 1.4.2 and 1.8.5; the 20,000-program census |
| Integer operations per mixer application | about 130 | spec 01 section 1.8.4 |
| Mixer applications per item | 9 under v2; 36 under v3 (m = 4, docs/plans/mixer-x4.md) |
memhard::Shape::mixers_per_item |
| Integer operations per item | 1,170 (v2); 4,680 (v3) | 9 x 130; 36 x 130 |
| Integer operations per hash | 149,760 (v2, "150,000"); 599,040 (v3, "600,000") | 128 x the above |
| Chip integer budget | 50 T op/s (approximate: 21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock) | M16 section 3 |
| Fixed-function factor | 3x (approximate, from memory: 2x to 5x is the usual credit for a pipeline with no scheduling or divergence) | M16 section 3 |
| RTX 5090, version 2 programs, measured | 136.1 MH/s (readwidth, tonight, docs/plans/read-width.md, pack w4 on PC 2); 139.7 MH/s (M11, 4 October, docs/bench-log.md) |
this analysis uses tonight's 136.1 as the denominator and quotes both |
| RTX 5090 at w16 (16-byte loads), measured | 139.8 MH/s | readwidth table, tonight (the width stays 4 B: w16 closes nothing) |
| Cache mirror, 256 MiB, N5 headline density | 128 mm^2, $46 per good die (64 mm^2, $21 at the bit-cell lower bound) | docs/analysis/sram-mirror.md revision 2, sections 4 and 5 (ca2-analysis e6085c6) |
| Cache mirror plus a 96 MB hot table, N5 headline | 175 mm^2, $68 | same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate) |
| 512 MiB and 1 GiB mirrors, N5 headline | 255 mm^2 and 510 mm^2; $111 to $306 | same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node) |
| GPU-class die | 750 mm^2 (the equal-silicon comparison) | M16 section 3 |
| CPU verifier, one M5 Max core (loaded, load average 5.6; ratios are the measurement) | v2 1.31 to 1.36 ms per unit, x4 1.92 to 1.96 (1.45x), x8 2.79 (2.1x); worst cold 1.58 / 2.04 / 2.94 ms | docs/plans/mixer-x4.md section 6.4, 5 October 2026 21:40 UTC |
2. The rows
Chip rate = 50 T op/s / ops per hash. "Bare" = chip rate / 136.1 MH/s. "With the factor" = bare x 3. "Equal silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does not get, the M16 convention ("minus the area the SRAM takes"). SRAM in mm^2 and dollars at the N5 headline density.
| Row | Mixer | Ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | With the 3x factor | Equal silicon, SRAM deducted, with the factor |
|---|---|---|---|---|---|---|---|---|
| v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x |
| v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x |
| x4 (the candidate measured beside v3; not v3) | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x |
| MEASURED, NOT ADOPTED (layer 5 decided out of v3 on the PC rows, coordinator 21:40 UTC): v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.66x at the Mac's g = 0.93 (126.6 MH/s); 0.70x at the 5090's g = 0.87 (118.4); the 9070 XT's g 0.84 | 1.98x (Mac g), 2.11x (5090 g) | 1.60x, 1.71x |
| MEASURED, NOT ADOPTED: v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.71x at the Mac's g = 0.87 (118.4 MH/s); 0.73x at the 5090's g = 0.84 (114.3); the 9070 XT's g 0.80 | 2.12x (Mac g), 2.19x (5090 g) | 1.67x, 1.73x |
| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), no hot table | x4 | 599,040 | 83.5 | 512 MiB | 255 / $111 | 0.61x | 1.84x | 1.21x |
| v3 at year 12 (cache 1 GiB, dataset 8 GiB) | x4 | 599,040 | 83.5 | 1 GiB | 510 / $306 | 0.61x | 1.84x | 0.59x |
| v3: mixer x8 (decided 22:05 UTC under the delegated rule: verify 2.1 ms per unit on one Mac core against the 10 ms gate, the daily 1 GiB build 23 to 77 ms on the 5090 and the 9070 XT) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x |
| x8 at year 4 | x8 | 1,198,080 | 41.7 | 512 MiB | 255 / $111 | 0.31x | 0.92x | 0.61x |
The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The width rule (4-byte loads kept) changes nothing either: w16 would have moved the honest denominator by 2.7% and the chip's cost not at all.
Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6
hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53.
Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of sram-mirror.md
scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. The honest denominator in the added
form is the v2 rate times g, the card's measured ratio with the hot loads added: on the M5 Max tonight
g = 0.93 / 0.87 / 0.83 at 32 / 64 / 96 MiB (the cache agent, relayed by the coordinator at 21:23 UTC;
docs/plans/hot-table.md carries the runs); the 5090's and the 9070 XT's g are the PC rows, owed, and until they
land the row carries the Mac's g against the 5090's rate, which is a mixed figure and is marked so. Year 4 and 12 rows: the mirror of
sram-mirror.md section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of
those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area,
not a forecast).
3. The margin, plainly
The combined headline row is the mixer row alone (layer 5 is out: the added form costs the 5090 13 to 16 percent
and the 9070 XT 16 to 20 percent against the 0.97 bar, coordinator 21:40 UTC; the width stays 4 bytes; the era
draws and the cache growth cost this chip nothing at year 0), and class v3 is x8 (decided 22:05 UTC). The headline:
the on-die-cache recompute chip at 50 T op/s reaches 41.7 MH/s against the 5090's 136.1, 0.31x bare, 0.92x with
the 3x fixed-function factor, 0.76x with the mirror's area deducted: under 1x with the factor, 0.92x, a margin of 8
percent on the factor (a 3.3x factor reads 1.0x) and of 9 percent on the budget (55 T op/s reads 1.0x). The x4
candidate, measured beside it, read 1.84x and 1.53x. The hot-table rows above are kept as measured, not adopted:
against THIS chip an added hot table is a cost to the honest card and none to the chip, so it would have moved the
row the wrong way by the card's own g. The margin, plainly:
- the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x;
- the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart);
- the 50 T op/s budget is approximate; a chip at 55 T op/s reads 2.0x;
- the hot table in the added form lowers the honest denominator by whatever the hot loads cost the GPU (owed from the PC rows), which raises the chip's gain by the same share, 1.84x or more if the hot loads are free, higher if not; the hot table's only cost to this chip is 16 to 32 mm^2 of die.
What keeps it under 1x is the mixer, and nothing else in Counter ASIC 2.0 moves this chip (the scratch at any share
gave 2.4x, docs/analysis/scratch-soundness.md section 3.4; the hot table taxes the DRAM-only chip, not this one;
the cache growth taxes it only in die area, which is cheap at year 0 and real at year 12). The next levers, in
order:
- Mixer x16 (the next step of the same lever): 0.16x bare and 0.46x with the factor against 136.1; the verifier by the measured increments (+0.63 ms at x4, +1.46 at x8 on the M5 Max core: about +3.1 ms at x16, 3.7 ms per unit, 9 ms on a 2.5x slower laptop core, approximate) is at the edge of the 10 ms gate, so a 2019-class laptop core measurement (O-1.14) decides it, not this model.
- The hot table: adopted or not on the PC rows (
docs/plans/hot-table.md); in the added form it costs the GPU 7 to 17 percent on the Mac and the chip die area only, so against this chip it is a lever in the wrong direction and against a DRAM-only chip the first lever; if it is adopted, the mixer must carry the extra1/g(x8 at g = 0.87 reads 1.06x at the equal budget, 0.84x with the SRAM deducted).
4. What this does not settle
The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM. The time-memory curve and the partial-store chip (f = 0.25, 0.5, 0.75, 1 on GDDR7 and HBM3) are now drawn in section 5 (Counter ASIC 3.0 item 1, 6 October 2026): the f = 1 chip is over 2x per joule on both memory systems, and the mixer does not touch it.
5. The partial-store chip and the time-memory curve (Counter ASIC 3.0 item 1, 6 October 2026)
Counter ASIC 3.0 item 1 (docs/plans/counter-asic-3.md, row 1; the history's addition 1,
docs/analysis/asic-resistance-history.md section 4.3). The chip priced here holds a fraction f of the dataset in
off-die DRAM (GDDR7 or HBM3) and derives the other 1 - f of its items from the 256 MiB on-die cache under class v3
(x8), reading DRAM at the hash's 4-byte granularity through its own controller. Sections 2 and 3 priced only f = 0.
Nothing below is a measurement of a chip. Every GPU figure says where it was measured; every chip figure is arithmetic
on cited memory and logic figures, and "approximate" marks a figure from memory or an estimate. The history's rows 3
and 4 (asic-resistance-history.md section 1.1) are the precedent: Ethash chips reached 2.1x (Linzhi Phoenix, 2020),
2.9x (Antminer E9, 2022) and 4.8x per joule (Jasminer X4, 2021) with custom memory controllers and on-package memory
and no on-die dataset, which is the f = 1 end of this curve.
5.1 Inputs
| Input | Value | Source (URL read 6 October 2026 unless a file is named) |
|---|---|---|
| RTX 5090 class v3, the denominator | 136.1 MH/s at 326 W (power approximate: 328.6 W peak in the 5 October prover-cost run, 323 W after the M11 race); 2.40 microjoules per hash; 17.5 G dependent 4-byte reads per second (CUDA wall), 18.2 (OpenCL event), 415 ns at 256 lanes; a 32-byte sector per read | docs/bench-log.md "Counter ASIC 2.0, the numbers" and M11; docs/benchmarks/repro.md 2.2 (branch repro-bench) |
| RTX 5090 memory system | 32 GB GDDR7 on a 512-bit bus at 28 Gbps, 1,792 GB/s; 16 devices of 2 GB (16 Gb); 575 W TGP; $1,999 at launch; 96 MB L2; 750 mm^2 | https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/ (512-bit, 32 GB GDDR7, 21,760 cores, 575 W); https://en.wikipedia.org/wiki/GeForce_RTX_50_series (28 Gbps, 1,792 GB/s, $1,999, 96 MB, 750 mm^2); the 16-device count is the brief's and the 32 GB / 2 GB arithmetic |
| RTX 5090 integer rate | 45.2 T op/s (OpenCL event, integer chain) | repro.md 2.2 |
| Program work per hash | 64 instructions x 8 iterations = 512, of which 128 loads | docs/spec/01-lottery-hash.md 1.4 (the table at lines 70 to 71) |
| GDDR7 organisation | four 10-bit channels per device (8 data bits each); 16 banks per channel; PAM3; 1.1 to 1.2 V; 28 to 32 Gbps per pin now, 48 on the roadmap | https://www.rambus.com/blogs/all-you-need-to-know-about-gddr7/ (channels, PAM3, voltage, rates); https://www.smart-dv.com/memory/gddr7.html (4 channels, 16 banks per channel); so 64 channels and 1,024 banks on the 5090's 16 devices |
| GDDR7 access granularity | 32 bytes per channel access (8 data bits x a burst of 32 beats; the 5090's measured 32-byte sector agrees) | arithmetic on the channel width; the sector is repro.md 2.2; the JEDEC burst length itself is behind the paywall, so the 32 B is marked approximate |
| GDDR7 energy, streaming | 4.5 pJ per bit average device power (GDDR6X 6, GDDR6 6.5) | Micron, quoted in the search results for the GDDR7 product brief (the brief's own PDF refused the fetch); Rambus: "over 10 percent less power per bit than GDDR6X"; approximate |
| GDDR7 price | about $20 per 2 GB device (16 Gb, 28 Gbps), September 2026; 3 GB devices $60 to $70 | https://www.trendforce.com/news/2026/09/24/news-micron-reportedly-ends-2gb-gddr7-narrowing-supply-options-for-nvidias-rtx-50-series/ (quoting Tom's Hardware and VideoCardz) |
| HBM3 organisation | 1,024-bit interface, 16 64-bit channels, 32 32-bit pseudo-channels; burst of 8 beats, 32-byte packet; up to 64 banks per channel; 6.4 Gbps per pin, 819 GB/s per stack; core 1.1 V, I/O 0.4 V; 64 GB per stack maximum | https://www.synopsys.com/glossary/what-is-high-bandwitdth-memory-3.html and https://www.synopsys.com/articles/hbm3-ip-dwtb.html; https://www.tomshardware.com/news/hbm3-spec-reaches-819-gbps-of-bandwidth-and-64gb-of-capacity |
| HBM3E | the same organisation, 9.2 to 9.8 Gbps per pin, 1.15 to 1.2 TB/s per stack; 24 GB (8-high) and 36 GB (12-high) | https://en.wikipedia.org/wiki/High_Bandwidth_Memory; https://blogs.sw.siemens.com/semiconductor-packaging/2026/04/24/hbm3e-hbm4-ic-design-guide/ |
| HBM energy per bit, streaming | Samsung's roadmap: HBM2 6.25, HBM3 4.12, HBM3E 4.05 pJ per bit; O'Connor et al. (NVIDIA, MICRO 2017): HBM2 3.97 pJ per bit | https://eureka.patsnap.com/insight/the-hbm-wars-sk-hynixs-dominance-samsungs-roadmap-and-the-looming-threat-of-cyclicality (20 August 2025); https://www.cs.utexas.edu/~skeckler/pubs/MICRO_2017_Fine_Grained_DRAM.pdf |
| HBM random-access energy, the breakdown | HBM2: row activation 909 pJ per 1 KB row; data movement 1.51 + 1.17 pJ per bit; I/O 0.80 pJ per bit at 50 percent activity; 32-byte atom; 16 banks per channel; tRC 45, tRCD 16, tRP 16, tRAS 29, tRRD 2, tFAW 12 ns, 8 activates per tFAW per channel | O'Connor et al., Tables 2 and 3 of the PDF above (text extracted with pdftotext) |
| DRAM row cycle across types | DDR4 tRCD 14, tRAS 33, tRP 14; GDDR5 14, 28, 12; HBM and HBM2 14, 34, 14 ns; 16 banks per rank; GDDR5 page 2 KB, HBM2 page 2 KB | Li, Reddy, Jacob, MEMSYS 2018, Table 2: https://terpconnect.umd.edu/~blj/papers/memsys2018-dramsim.pdf (text extracted with pdftotext); the history's [L1] |
| HBM random-access ceiling, the literature | "Folded Banks" (AMD, ISCA 2025): fine-grained random access on HBM is bound by activate parallelism (tRC, tRRD, tFAW), and a redesign with 8x the activate parallelism gives 6.7x the irregular bandwidth, which says the stock stack sits far under its streaming figure on random reads | https://dl.acm.org/doi/10.1145/3695053.3731111 (abstract; the PDF refused the fetch) |
| HBM3 price | about $200 per 24 GB stack factory gate (HBM3E $300 per 36 GB), October 2026; contract pricing about twice that | https://siliconanalysts.com/data/hbm-pricing (no external source cited there; approximate) |
| Interposer and packaging | CoWoS-S $600 to $900 per H100-class package (8 stacks, about 800 mm^2 of silicon interposer); CoWoS-L 20 to 47 percent more; September 2026 | https://siliconanalysts.com/tools/packaging (public sources only, approximate); a one-stack package is taken at $200, approximate |
| 32-bit integer op energy, N5-class | int32 add 0.06 pJ, int32 multiply 0.52 pJ at 5 nm (7 nm: 0.10 and 0.80; 45 nm, Horowitz 2014: 0.1 and 3.1) | https://mlsysbook.ai/vol1/backmatter/appendix_assumptions.html Table 13, citing Horowitz 2014 and Dally 2021; datapath only, so a 2x pipeline and clock overhead is applied below, approximate |
| On-die SRAM read, 64 bytes from a 256 MiB array | 0.5 nJ, range 0.2 to 1.0 (approximate: Horowitz's 45 nm 1 MB cache at 100 pJ per 64-bit read scaled to N5, plus about 0.6 pJ per bit of global wire across a 128 mm^2 array) | from memory; the sensitivity is shown in every row |
| Cache mirror and recompute die | 256 MiB = 128 mm^2, $46 at N5 headline density; the whole recompute die (50 T op/s of integer logic beside the mirror) taken as a 750 mm^2-class N5 die at about $600 of silicon (70 dies per $20,000 wafer at the sram-mirror.md yield model) |
docs/analysis/sram-mirror.md sections 4 and 5; the $600 is arithmetic on its wafer price and D0, approximate |
| Chip project cost | a 7 nm-class project $50M to $75M all-in; a 28 nm project $5M to $30M; masks $1M to $3M at 28 nm, $10M to $20M at 5 nm | asic-resistance-history.md section 2.5 and its [E2] [E3] [E4] |
| Ethash chips, the precedent | Phoenix 2,733 MH/s at about 3,000 W, 2.1x; X4 2,500 MH/s at 1,200 W, 4.8x; E9 2,400 MH/s at 1,920 W, 2.9x per joule | asic-resistance-history.md rows 3 and 4 and their [S11] [S12] |
5.2 The op count, counted from the code
igneum-pow/src/memhard.rs, mixer (lines 290 to 303) and derive_items_mask (lines 503 to 536), under class v3's
mixer_mult = 8:
| Where | Operations | Count |
|---|---|---|
Per mixer application, the 16-word prologue (s[i] ^ (rc[i] + rk)) * mul[i] |
16 xor, 16 add, 16 mul | 48 (32 when rc[i] + rk, a per-round constant, is hoisted; a chip hoists it) |
| Per mixer application, 8 quarter rounds of 4 add, 4 xor, 4 rotate | 32 add, 32 xor, 32 rotate | 96 |
| Per mixer application, total | 144 unhoisted, 128 hoisted | |
Mixer applications per item, (ITEM_ROUNDS + 1) x m |
9 x 8 | 72 |
Per item, the cache-line fold s[i] ^= line[i] over 8 rounds, and the 16-word init |
128 xor, 8 mul, 8 add | 144 |
| Per item, total | 72 x 128 + 144 (hoisted) to 72 x 144 + 144 | 9,360 to 10,512 |
| Per hash, 128 items | 1,198,080 to 1,345,536, plus the program's 512 |
The spec's "about 130" per application (1.8.4) is the hoisted count plus the fold spread over the applications: 9,360
per item exactly, so the figures of sections 1 and 2 stand. The energy per application at N5 datapath figures is
16 x 0.52 + 128 x 0.06 = 8.3 + 7.7 = 16.0 pJ (a rotate by a per-day constant is a wire mux on a chip, counted at the
add's 0.06); with the 2x pipeline overhead, 32 pJ. Per item: 72 x 32 pJ = 2.30 nJ of logic plus 8 cache reads at
0.5 nJ = 4.0 nJ plus the fold, 6.3 nJ (3.9 at 0.2 nJ per read, 10.3 at 1.0). Per hash at f = 0: 128 x 6.3 = 0.81
microjoules, before static power. The on-die cache reads cost this chip more energy than the mixer does, which is the
first thing the curve says: the mixer's 9,360 ops are 2.3 nJ of a 6.3 nJ item.
The rows below use the hoisted count, 9,360 per item (72 x 128 plus the fold), because a chip pays the cheapest
form and the section 1 and 2 rows already price that figure. The two counts differ by 12 percent, so the f = 0 row
at the unhoisted 10,512 per item (1,346,048 ops per hash) is: 50 T / 1,346,048 = 37.1 MH/s, 0.27x bare, 0.82x with
the 3x factor (against 0.31x and 0.92x); its energy per item 2.44 nJ of logic instead of 2.30, 6.5 nJ in all, 50.7 W at 37.1 MH/s, 1.37 microjoules,
1.75x per joule instead of 1.86x (the 20 W of static power spread over fewer hashes). The f = 1 rows do not move: they contain no mixer. The multiplier is class v3's shipped
mixer_mult = 8 (72 applications per item), not the 4 of mixer-x4.md section 2.
5.3 The two memory systems as random-read engines
A dependent 4-byte read opens a row (tRCD), reads one 32-byte atom (tCL and the burst) and must close it before the same bank opens another (tRC). The rate of random reads a memory system can sustain is the smaller of two ceilings: banks divided by tRC, and activates per tFAW window per channel times the channels. Neither ceiling moves with the pin speed, so HBM3E is HBM3 here, and 48 Gbps GDDR7 is 28 Gbps GDDR7.
| GDDR7, 16 devices, 512-bit (the 5090's board) | HBM3, one stack | HBM3, eight stacks (an H100-class package) | |
|---|---|---|---|
| Channels, banks | 64 channels, 1,024 banks | 16 channels (32 pseudo-channels), up to 1,024 banks | 128 channels, 8,192 banks |
| Bank-bound ceiling, banks / 45 ns (tRC, the HBM2 and GDDR5-class figure, approximate for both) | 22.8 G reads/s | 22.8 | 182 |
Activate-bound ceiling (GDDR7: 4 per 12 ns per channel, approximate; HBM: 8 per 12 ns per channel, O'Connor Table 2). UNMEASURED (6 October 2026, the Horizon lane analysis docs/analysis/horizon/algorithm.md section 5.1): the JEDEC HBM2 table gives tFAW 28 ns, 4 activates per channel per window (ICCAD 2021 Table I), which is 2.3 G reads/s for a 16-channel stack, and the one measured random-read rate of an HBM2 part (Shuhai, Alveo U280, FCCM 2020 Fig 7) is 2.4 G, equal to that tFAW ceiling; the 12 ns figure holds only if a bank-interleaved mapping lifts tFAW, which the die enforces per channel. The HBM columns below carry the 10.7 G row as the model's ceiling, unmeasured; an AWS F2 hour (Virtex UltraScale+ VU47P, the same HBM2 subsystem) is the measurement |
21.3 G reads/s | 10.7 (unmeasured; 2.3 at JEDEC tFAW) | 85.3 (unmeasured) |
| The ceiling carried below | 21.3 G reads/s (the 5090 measures 17.5, 82 percent of it: the card is already near its memory's activate limit) | 10.7 | 85.3 |
| Energy per random 32-byte read (approximate) | 2.0 nJ: 909 pJ activation (one atom per row opened, the HBM2 1 KB row taken for GDDR7's row) plus 4.5 pJ per bit x 256 bits of movement and I/O = 1,150 pJ | 1.2 nJ: the HBM2 sum (909 + (1.51 + 1.17 + 0.80) x 256 = 1,800 pJ) scaled by Samsung's 4.12 / 6.25 | 1.2 nJ |
| Static power (refresh, standby, PLLs; approximate, from memory) | 20 W (about 1.25 W per device) | 4 W | 32 W |
| Controller and PHY die beside it (approximate) | 15 W, $50 | 10 W, $50 | 40 W, $200 |
| Memory dollars | $320 (16 x $20) | $200 plus a $200 one-stack interposer | $1,600 plus $750 CoWoS-S |
| Lanes in flight needed at the ceiling, at a 55 ns controller latency (tRCD 16 + tCL 16 + burst 1.25 + about 20 of controller, approximate) | 1,172 lanes, 73 KB of lane state at 64 B | 588 lanes, 37 KB | 4,692 lanes, 293 KB |
| Reads per second per watt at the ceiling (memory, static and controller) | 0.27 G | 0.40 G | 0.49 G |
| The 5090 for comparison | 17.5 G reads/s at 326 W = 0.054 G per W; 7,262 reads in flight (17.5 G x 415 ns), 22 per watt |
The FPGA line, public (6 October 2026, the Horizon lane analysis section 5.1): an HBM2 FPGA soft overlay (Alveo U280 or U55C class) carries only the measured row, 2.4 G reads/s per card (Shuhai, FCCM 2020 Fig 7, equal to the JEDEC tFAW ceiling of 2.3 G at 28 ns), which is 0.30x to 0.39x of the RTX 5090 per watt (U55C at 115 to 150 W; 0.20x on the U280). The 11.4 G bank-bound row and the 12.2 G ceiling quoted elsewhere rest on a 12 ns tFAW the JEDEC HBM2 table does not give and are unmeasured until an AWS F2 hour (f2.6xlarge, VU47P, 16 GB HBM2, USD 1.98 an hour on demand) runs the chase kernel at 1 GiB across all 32 pseudo-channels; the lane's pass line is 15 to 25 M reads/s/W, its alarm line 27 (0.5x of the 5090), and over 54 (1.0x) the FPGA lane becomes a Counter ASIC 4.0 item. Consequence per tier: none today (no FPGA mines); if the measured row holds, a soft-overlay FPGA at USD 4,000 to 5,000 a card (approximate) mines at an RX 9070 XT's rate per watt for 7x the price, so no home or rig tier is displaced by it. Ledger M33.
The GDDR7 system's own power at the 5090's 17.5 G reads/s is 17.5 x 2.0 nJ = 35 W plus 20 W static, 55 W: about 17
percent of the card's 326 W (approximate). The other 83 percent is the GPU: 21,760 ALUs spinning at 92.9 percent
utilisation on 512 program ops per hash, their register files, schedulers, L1 and L2, and the clock trees, against a
45 T op/s integer budget of which the hash uses 512 x 136.1 M = 0.07 T op/s, 0.15 percent. That is the whole case for
the f = 1 chip: it is the 55 W without the 271.
5.4 The curve
Per row: reads per hash = 128 f; items recomputed = 128 (1 - f); ops per hash = 128 (1 - f) x 9,360 + 512.
Memory-bound rate = the ceiling / (128 f). Compute-bound rate = 50 T op/s / ops per hash (the section 1 budget).
The rate is the smaller; "binding" names it. Power = rate x (128 f x E_read + 128 (1 - f) x 6.3 nJ) + static (memory,
controller, and 20 W for the recompute die's clocks and leakage when f < 1). Energy per hash = power / rate. "Gain,
rate" = rate / 136.1 MH/s (per chip, the section 2 metric); "with the 3x factor" multiplies the compute-bound rate by
3 on the recompute share only, the memory ceiling unchanged. "Gain, joule" = 2.40 microjoules / energy per hash, the
Ethash chips' metric, with the on-die read energy at 0.5 nJ and, in brackets, at 0.2 and 1.0. Dollars = memory +
interposer + controller + $100 of board, plus the $600 recompute die when f < 1; no project cost (section 5.6).
Arithmetic: scratchpad curve.py, reproduced by hand for the first and last rows below the table.
| Memory | f | Reads per hash | Items recomputed | Ops per hash | Memory-bound MH/s | Compute-bound MH/s | Binding | Power W | Energy per hash, microjoules | Gain, rate, bare | With the 3x factor | Gain, joule (0.2 / 1.0 nJ reads) | Silicon and memory dollars | $ per MH/s |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| none (sections 1 to 3) | 0 | 0 | 128 | 1,198,592 | n/a | 41.7 | compute | 53.7 | 1.29 | 0.31x | 0.92x (125.1) | 1.86x (2.44 / 1.33) | $700 | $16.8 |
| GDDR7 | 0.25 | 32 | 96 | 899,072 | 665.6 | 55.6 | compute | 92.3 | 1.66 | 0.41x | 1.23x (166.8) | 1.44x (1.68 / 1.17) | $1,070 | $19.2 |
| GDDR7 | 0.5 | 64 | 64 | 599,552 | 332.8 | 83.4 | compute | 99.4 | 1.19 | 0.61x | 1.84x (250.2) | 2.01x (2.31 / 1.65) | $1,070 | $12.8 |
| GDDR7 | 0.75 | 96 | 32 | 300,032 | 221.9 | 166.6 | compute | 120.7 | 0.72 | 1.22x | 1.63x (221.9, memory) | 3.31x (3.70 / 2.81) | $1,070 | $6.4 |
| GDDR7 | 1 | 128 | 0 | 512 | 166.4 | 97,656 | memory | 77.6 | 0.47 | 1.22x | 1.22x | 5.14x | $470 | $2.8 |
| HBM3, one stack | 0.25 | 32 | 96 | 899,072 | 334.4 | 55.6 | compute | 69.9 | 1.26 | 0.41x | 1.23x (166.8) | 1.91x (2.33 / 1.46) | $1,150 | $20.7 |
| HBM3, one stack | 0.5 | 64 | 64 | 599,552 | 167.2 | 83.4 | compute | 74.1 | 0.89 | 0.61x | 1.23x (167.2, memory) | 2.69x (3.26 / 2.09) | $1,150 | $13.8 |
| HBM3, one stack | 0.75 | 96 | 32 | 300,032 | 111.5 | 166.6 | memory | 69.4 | 0.62 | 0.82x | 0.82x | 3.85x (4.39 / 3.19) | $1,150 | $10.3 |
| HBM3, one stack | 1 | 128 | 0 | 512 | 83.6 | 97,656 | memory | 26.8 | 0.32 | 0.61x | 0.61x | 7.46x | $550 | $6.6 |
| HBM3, eight stacks | 0.25 | 32 | 96 | 899,072 | 2,665.6 | 55.6 | compute | 127.9 | 2.30 | 0.41x | 1.23x (166.8) | 1.04x (1.16 / 0.89) | $3,250 | $58.4 |
| HBM3, eight stacks | 0.5 | 64 | 64 | 599,552 | 1,332.8 | 83.4 | compute | 132.1 | 1.58 | 0.61x | 1.84x (250.2) | 1.51x (1.67 / 1.30) | $3,250 | $39.0 |
| HBM3, eight stacks | 0.75 | 96 | 32 | 300,032 | 888.5 | 166.6 | compute | 144.9 | 0.87 | 1.22x | 3.67x (499.9) | 2.75x (3.02 / 2.40) | $3,250 | $19.5 |
| HBM3, eight stacks | 1 | 128 | 0 | 512 | 666.4 | 97,656 | memory | 174.4 | 0.26 | 4.90x | 4.90x | 9.15x | $2,650 | $4.0 |
| HBM3E, any row | the same: the activate ceiling does not move with the pin rate | 1 percent lower (4.05 against 4.12 pJ per bit) | $300 per 36 GB stack |
Arithmetic, f = 0: 50 x 10^12 / 1,198,592 = 41.7 x 10^6; power 41.7 M x 128 x 6.3 nJ = 33.7 W, plus 20 W static =
53.7 W; 53.7 / 41.7 M = 1.29 microjoules; 2.40 / 1.29 = 1.86. Arithmetic, GDDR7 f = 1: 21.3 G / 128 = 166.4 MH/s;
power 166.4 M x 128 x 2.0 nJ = 42.6 W, plus 20 + 15 static = 77.6 W; 77.6 / 166.4 M = 0.466 microjoules; 2.40 / 0.466 =
5.14. HBM3 one stack f = 1: 10.7 G / 128 = 83.6 MH/s; 83.6 M x 128 x 1.2 nJ = 12.8 W, plus 14 = 26.8 W; 0.321
microjoules; 7.46x. Lanes: 21.3 G x 55 ns = 1,172.
What the curve says:
- It is monotone. Every memory system's cheapest point is
f = 1, and the partial rows are worse than both ends on dollars per MH/s. A recomputed item costs 6.3 nJ and 9,360 ops; a stored one costs 1.2 to 2.0 nJ and no ops. At $8 to $10 per GB nothing makes a chip maker recompute a 1 to 2 GiB dataset; the partial-store chip is not the threat and will not be built. The curve matters again only if the dataset outgrows cheap memory, and the schedule of 1.13.3 (2 GiB plus 0.5 GiB a year; 4 GiB at year 4) stays under one HBM3 stack's 24 GB for the chain's life. - The
f = 0chip of sections 1 to 3 reads 0.31x per chip at the op budget and 0.92x with the factor, and those rows stand. Per joule the same chip reads 1.3x to 2.4x (1.86x at 0.5 nJ per on-die read), because its 50 T op/s of fixed-function logic draws about 54 W where the 5090 draws 326 to do the same work. The two metrics disagree because they measure different things: the rate row says how many chips match one card, the joule row says what each hash costs to run. The public claim "under 2x" has so far been the rate row. - The
f = 1rows are over 2x per joule on both memory systems, at every on-die read energy (the on-die cache is not in them), and over 2x per dollar on GDDR7: 5.1x and $2.8 per MH/s against the 5090's $14.7 at launch price. The history's band for exactly this chip class is 2.1x to 4.8x per joule (Ethash, rows 3 and 4); the model reads 5.1x (GDDR7) to 9.2x (eight HBM3 stacks) with an ideal controller and the static allowances above. Read the history's band as the floor a first chip reaches and the model as the ceiling.
5.5 The f = 1 chip: a GPU's memory system without the GPU
A dependent read chain cannot be pipelined within a hash: read r + 1's address is read r's data through the
program's registers, so one hash advances one read per memory latency. The rate per chip is therefore lanes in
flight divided by latency, and the only question is how many lanes each memory system lets a controller keep in
flight per watt. The 5090 keeps 7,262 (17.5 G reads/s x 415 ns) at 326 W, 22 per watt. The f = 1 chip's lanes are
64 bytes of registers each (the hash's eight 32-bit registers and the program counter and nonce), so 1,172 lanes
are 73 KB of SRAM, and its latency is the controller's, about 55 ns, not the GPU's 415 ns of queueing: the chip
holds 6x fewer reads in flight and still reaches the memory's activate ceiling. What it needs to beat the 5090 is
more than 17.5 G reads per second per 326 W, 0.054 G per watt; the GDDR7 system alone gives 0.27 and one HBM3 stack
0.40 (section 5.3). An HBM3 part at the 5090's 326 W: twelve stacks, about 128 G reads/s, 1,000 MH/s, 7.3x per chip
(the eight-stack row at 174 W is 4.9x). HBM3's bank count per stack (1,024, in 32 pseudo-channels) equals the whole
GDDR7 system's 1,024 banks over 64 channels, so its random-read rate per stack is half the 5090 board's (10.7
against 21.3 G) on the activate count; what HBM wins is energy per read (1.2 against 2.0 nJ: 0.8 pJ per bit of
interposer I/O against a PCB) and the dollars per read per second favour GDDR7 ($320 for 21.3 G against $400 for
10.7 G). Both beat the card per joule by more than 2x because the card's memory system is 17 percent of its power.
5.6 Verdict
Over 2x. The worst case for us is f = 1 on HBM3 (7.5x per joule for one stack, 9.2x for eight; GDDR7 5.1x), and
the cheapest chip for an attacker is f = 1 on GDDR7 ($2.8 per MH/s of silicon and memory, a controller die that
needs no advanced node because the mixer is not on it: a 28 nm-class project at $5M to $30M, not the $50M 7 nm
project the on-die cache forces on the f = 0 chip). The on-die-cache recompute chip stays at 0.92x with the factor
and 1.3x to 2.4x per joule, between "under 1x" and "1 to 2x", and it is not the chip anyone builds. The verdict
changes the public claim: "under 2x" held for the chip that recomputes the dataset; it does not hold for the chip
that stores it.
5.7 What it means for item 2, and what does protect
Item 2 (a random item-derivation program per day in place of the fixed mixer) removes the 3x fixed-function factor
from the recompute share. The f = 1 chip has no recompute share: it reads every item from DRAM and never derives
one, so item 2 moves none of the rows that decide the verdict, and neither does a mixer at x16 or x64. The mixer
earns its place against the f = 0 chip, which at $8 per GB of DRAM nobody builds. Item 2's urgency is therefore
low on this result; it stays a reserve family. What does move the f = 1 rows, with the arithmetic:
| Lever | What it does to the f = 1 chip | What it costs the honest cards | Reading |
|---|---|---|---|
| Dataset size | Nothing until the dataset exceeds what one stack or one board holds: 24 GB (HBM3, one stack) or 32 GB (the 5090's own board). The schedule reaches 4 GiB at year 4 | Everything: a 4 GiB step already retires 4 GB cards (card-lifetime-2026-10-05.md) |
Not a lever against this chip |
| Read granularity | The chip pays the same 32-byte atom the 5090 pays; the 9070 XT pays 64. Wider honest reads (w16, measured, layer 1) give the chip nothing and the 5090 nothing; w64 made the 5090 bandwidth-bound (71.9 MH/s) | w64 costs the 5090 47 percent | Not a lever; the decision to stay at 4 B stands |
| Latency | A longer chain (more reads per hash) scales the chip's rate and the card's rate together; lane state is 64 B, so lanes are free to the chip | Nothing per se | Not a lever: the rate per chip is lanes / latency on both sides and the chip has more lanes per watt |
| The denominator: the 5090's watts at the hash | The gain is 2.40 microjoules over the chip's 0.47; the card's 326 W is 92.9 percent utilisation spinning on loads. At a 250 W cap holding 136.1 MH/s the gain reads 3.9x (GDDR7) and 5.7x (HBM3); at 200 W, 3.2x and 4.6x | None if the rate holds under the cap; the measurement is one PC 2 job (nvidia-smi -pl 200, 250, 326, two minutes each, STATUS lines as the rate) |
The first measurement to run; it moves every row and costs nothing. Owed (PC 1 is not released; PC 2's budget is the coordinator's) |
| Program work in the latency shadow | The hash hides 512 ops per hash behind 128 reads; the 5090 could hide 330,000 (45.2 T / 136.1 M) before compute binds, the M5 Max about 290,000 and the 9070 XT about 650,000 (their ALU budgets approximate, from memory). Work in the shadow is free in hash rate and costs the card watts it now wastes: at N ops per hash the card rises from 326 toward 575 W (linear, approximate) and the chip must add a core that runs the per-epoch random program, at k times the GPU's 5.5 pJ per op (the 5090's marginal ALU energy, (575 - 326) / 45.2 T). At N = 100,000: the card 401 W, 2.95 microjoules; the chip 1.02 at k = 1, 0.83 at k = 1.5; gain 2.9x and 3.5x. At N = 200,000: 477 W, 3.50; chip 1.57 and 1.20; gain 2.2x and 2.9x. At N = 330,000: 575 W, 4.22; chip 2.28 and 1.68; gain 1.85x and 2.5x | Hash rate none while every card stays latency-bound (under about 290,000 on the M5 Max); watts up to TGP; the verifier N x 32 ops per warp: 3.2 M at N = 100,000, under 1 ms at the 18 G op/s the x8 verifier shows (38 M ops in 2.08 ms), inside the 10 ms gate; INSTR_COUNT and ITERATIONS are prototype values to be fixed at gate 1 (spec 1.4) |
The only lever that moves the f = 1 row toward 2x, and only if the chip's core is no better than a GPU's on a random program (k near 1: RandomX's argument, and the founder's goal in the brief's words, "build a better GPU than NVIDIA"). It reaches 1.85x at the 5090's full ALU budget and k = 1, not under; combined with a 250 W cap it reads about 1.4x (approximate). It is item 2's idea applied to the program, not to the item derivation |
| The clock (item 4) | The f = 1 chip is a commodity-memory controller project: by the Ethash precedent, 32 months to a first chip at the largest prize, and a chip over 2x at 65 months | None | The issuance trigger and the share-pattern detector matter more than any item-derivation change |
So: item 2 can wait; the power-cap measurement runs first; the program-length lever is the Counter ASIC 3.0 design item this analysis adds, with its own six gates (the hash rate per card at N = 50,000, 100,000, 200,000; the watts; the verifier on a 2019-class core; bit-exactness on three vendors); and the public claim is re-worded until the measurements land.
5.8 Consequences per user tier
| Tier | What the verdict means | What is being done |
|---|---|---|
| Home miner, one 8 GB card (any vendor, any OS) | Today, nothing changes: no chip exists, the devnet pays nothing, and the f = 1 chip is a project of $5M to $30M and about 32 months by the precedent. When one lands it runs at 0.3 to 0.5 microjoules per hash; an 8 GB card runs 10 to 20 (approximate: the 9070 XT's 18.6 MH/s at its 304 W board power, from memory, is 16 microjoules) and is the first tier out | The power-cap and program-length measurements (above) decide how far the honest card's joules can fall; the share-pattern detector (item 4) is what tells this miner a chip has arrived |
| One 12 GB card | As the 8 GB tier; the dataset size gives it no protection, since the chip's memory is 24 to 32 GB whatever the card holds | The same |
| One 16 GB card (9070 XT class) | AMD RDNA 4 at 2.4 G reads/s and 304 W (approximate) is 7x worse per joule than the 5090 and 30x worse than the f = 1 chip; a chip ends AMD home mining first | The vendor-share metric (item 7) will show it; nothing in the hash fixes AMD's dependent-read rate |
| One 24 or 32 GB card (5090 class, Apple M5 Max) | The 5090 is the honest best at 2.40 microjoules; the M5 Max at 27.9 MH/s and a GPU power of about 60 to 80 W (approximate, unmeasured) is 2 to 3 microjoules, the same class per joule. Against the f = 1 chip both are 5x to 9x behind per joule in the model, 2x to 5x by the precedent | The two measurements; the 5090 at a 200 W cap (if its rate holds) is 1.47 microjoules and halves the gap |
| A rig | A rig's cost is electricity; per joule it is its cards. Once chips hold the hashrate, a rig at the same tariff earns 1/2 to 1/9 of a chip per watt and leaves. The precedent: ASIC share of Ethash stayed small (about 3 percent) for years because the chips were not cheap enough per dollar at scale, and the dollars-per-MH/s row says this chip is ($2.8 against $14.7) | The issuance trigger (item 4): the bounty and the benchmark live before daily issuance crosses about $50K |
| A pool user | A chip fleet is a few operators at 0.3 microjoules; MoneroCrusher found Monero's at 85 percent of the hashrate by the share pattern | The detector on the observer, item 4, before the public testnet |
| The public claim "under 2x" | It held for the recompute chip at the op budget. Per joule and against the stored-dataset chip the model reads over 2x on both memory systems and the precedent reads 2.1x to 4.8x. The claim as worded is not safe to publish | Re-word to the measured fact (the 5090 runs at 1/128 of its dependent-read ceiling and a chip must out-read it per watt) until the power-cap and program-length rows land; nothing goes to the devnet or the site from this analysis |
5.9 Unverified and owed
- Every energy figure is a published streaming or breakdown figure applied to random 32-byte reads: HBM2's 909 pJ activation and 1 KB row stand in for HBM3 and for GDDR7 (GDDR6 rows are 2 KB per 16-bit channel, Li-Reddy-Jacob Table 2; GDDR7's 8-bit channel row is unknown to me); the static allowances (20 W, 4 W, 32 W), the controller powers, the 55 ns controller latency and the 2x pipeline overhead are from memory.
- The 5090's 326 W is a peak from a run with the prover on (328.6) and a post-run reading (323); the power at the hash alone, and under a cap, is unmeasured. PC 1 is not released (the 9070 XT rows are OWED) and no PC 2 job was published for this item (analysis only; the measurement is named, not run).
- The 50 T op/s budget, the 3x factor and the $600 die are section 1's approximations; the on-die read energy (0.2 to
1.0 nJ) moves the
f < 1rows by up to 1.8x and thef = 1rows not at all. - Prices are factory-gate figures from a secondary source (contract pricing about 2x), September to October 2026; the 5090's $1,999 is the launch price and its 2026 street price is higher (reports of $4,000 and more, approximate), which makes the chip's dollar advantage larger, not smaller.
- The Ethash chips' internals are not read (no teardown); their gains are the history's rows.
- The GDDR7 burst length and bank-group timing, and HBM3's tFAW per pseudo-channel, are behind the JEDEC paywall; the activate ceilings use the HBM2 figures and the 5090's measured 17.5 G (82 percent of the GDDR7 ceiling) as the check that they are the right order.
- The ALU budgets of the M5 Max and the 9070 XT, their power at the hash, and the verifier's cost at N = 100,000 program ops are estimates; the program-length lever is a design item with its own measurements, not a result.
6. The per-day derivation (item 2)
6 October 2026, Counter ASIC 3.0 item 2, worker derive (docs/plans/counter-asic-3-derivation.md; everything
PROPOSED, a prototype behind load class dr736). The fixed-shape mixer of section 2's rows is replaced by nine
straight-line programs of 736 instructions per item drawn from the day key stream (twelve two-register forms,
the chain rule, an acceptance test with the x8 mixer's counts as floors). The chip's cost per hash is still item
derivations; what changes is the fixed-function factor, because the chip must now execute an arbitrary program
of the day from a 12-form set over 16 registers (a sequencer: instruction store, register file, operand muxes, a
32-bit ALU with a multiplier and a rotator) instead of a wired pipeline of 72 mixer stages with the day's
constants in the wires. The counts are from the code (memhard::mixer: 144 ops per application as written, 128
with the round constants hoisted, 16 multiplies; the x8 item is 10,368 / 9,216 / 1,152), not the 130 of section 1;
the day program's floor is those counts, so the bare row cannot fall below x8's.
| Row | Derivation | Chip ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | Allowance 1.2x (ProgPoW's claimed range, history 2.4 [S67] [S70]) | Allowance 1.5x (cautious upper bound, approximate) | The old 3x (the fixed shape's; does not apply) | Equal silicon, SRAM deducted, at 1.2x / 1.5x |
|---|---|---|---|---|---|---|---|---|---|---|
| x8 as shipped (section 2's v3 row, re-counted from the code with constants hoisted) | fixed mixer, 72 x 128 | 1,179,648 | 42.4 MH/s | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
| dr736, the genesis day's draw (9,992 chip ops, 1,461 multiplies per item; the floor is x8's 9,216) | the day program, 9 x 736 instructions | 1,278,976 | 39.1 MH/s | 256 MiB | 128 / $46 | 0.29x | 0.34x | 0.43x | 0.86x | 0.29x / 0.36x |
| dr736 at the floor (a day whose draw sits exactly on the acceptance floor) | the day program | 1,179,648 | 42.4 | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
| dr368, the fallback (the x4-equivalent count: 5,004 chip ops per item on the genesis day) | the day program, 9 x 368 | 640,512 | 78.1 | 256 MiB | 128 / $46 | 0.57x | 0.69x | 0.86x | 1.72x | 0.57x / 0.71x |
| dr736 at year 4 (cache 512 MiB) | the day program | 1,278,976 | 39.1 | 512 MiB | 255 / $111 | 0.29x | 0.34x | 0.43x | 0.86x | 0.23x / 0.28x |
Arithmetic, row dr736: 9,992 x 128 = 1,278,976; 50 x 10^12 / 1,278,976 = 39.1 x 10^6; 39.1 / 136.1 = 0.287; x 1.2 = 0.345; x 1.5 = 0.431; x 3 = 0.862; equal silicon (750 - 128) / 750 = 0.829, x 0.345 = 0.286, x 0.431 = 0.357. The allowance argument, plainly: the 3x of section 1 was the credit for "a pipeline with no scheduling or divergence", which a fixed dataflow earns because the chip wires the 72 applications and bakes the constants in; with a program that changes daily the chip keeps no divergence (the GPU has none here either: the item function is straight-line), the constants folded into an instruction store, and no warp scheduler or operand collector, and it loses the wiring. That residual is what ProgPoW's audits priced at 1.1x to 1.2x for a conventional compute chip (Rao: "conventional compute chips gain little on ProgPoW", history section 2.4); 1.5x is a cautious upper bound of mine (approximate) for a chip that also drops the GPU's float and graphics area. The chain rule (every instruction reads the register the previous one wrote) adds a cost the row does not credit: with no intra-item parallelism a single engine completes one dependent instruction per cycle at best and must interleave items to keep its multiplier busy, which is a register file per item in flight (RandomX's light-mode argument, history 2.4). The measured costs that buy this: the verifier 4.88 ms per unit on one M5 Max core against x8's 2.06 (the derivation document's section 5.1), the Mac's daily build 29 ms against 22, the hash rate unchanged; the 5090's build and compile are the PC 2 job, the 9070 XT's OWED.
What this does not settle: the rows are the same 50 T op/s budget and the same denominator as section 2 (their
margins apply); no chip has been priced for its instruction store or its register files per item in flight; the
random ARX programs have had no cryptanalysis (the item 3 brief should name them beside M_r); the 2019-class
core measurement (O-1.14) decides whether 736 or 368 is the length, and the derivation document's section 0
carries that verdict.