From 6039ae8d9e94326081b420987b0b2e63fa3996a7 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 07:32:00 +0000 Subject: [PATCH] Counter ASIC 3.0 item 1: the partial-store chip and the time-memory curve in chip-model-v3.md section 5 f = 0, 0.25, 0.5, 0.75, 1 on GDDR7 (the 5090's board), one HBM3 stack and eight HBM3 stacks, scored in energy per hash, reads in flight per watt, rate per chip and dollars per MH/s against the RTX 5090 at 136.1 MH/s and 326 W. The curve is monotone toward f = 1; the f = 1 chip reads 5.1x (GDDR7) to 9.2x (HBM3) per joule in the model and 2.1x to 4.8x by the Ethash precedent: over 2x. The mixer and item 2 do not touch it; the levers named are the 5090's watts under a power cap (a PC 2 job, owed) and program work in the latency shadow. Inputs cited with URLs read 6 October 2026; the mixer counted from memhard.rs at 128 hoisted / 144 unhoisted ops per application. Co-Authored-By: Claude Fable 5.1 --- docs/analysis/chip-model-v3.md | 234 +++++++++++++++++++++++++++++++++ 1 file changed, 234 insertions(+) diff --git a/docs/analysis/chip-model-v3.md b/docs/analysis/chip-model-v3.md index 1d47b2f72..43a864728 100644 --- a/docs/analysis/chip-model-v3.md +++ b/docs/analysis/chip-model-v3.md @@ -98,3 +98,237 @@ order: The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM. +The time-memory curve and the partial-store chip (f = 0.25, 0.5, 0.75, 1 on GDDR7 and HBM3) are now drawn in section 5 +(Counter ASIC 3.0 item 1, 6 October 2026): the f = 1 chip is over 2x per joule on both memory systems, and the mixer does +not touch it. + +## 5. The partial-store chip and the time-memory curve (Counter ASIC 3.0 item 1, 6 October 2026) + +Counter ASIC 3.0 item 1 (`docs/plans/counter-asic-3.md`, row 1; the history's addition 1, +`docs/analysis/asic-resistance-history.md` section 4.3). The chip priced here holds a fraction `f` of the dataset in +off-die DRAM (GDDR7 or HBM3) and derives the other `1 - f` of its items from the 256 MiB on-die cache under class v3 +(x8), reading DRAM at the hash's 4-byte granularity through its own controller. Sections 2 and 3 priced only `f = 0`. +Nothing below is a measurement of a chip. Every GPU figure says where it was measured; every chip figure is arithmetic +on cited memory and logic figures, and "approximate" marks a figure from memory or an estimate. The history's rows 3 +and 4 (`asic-resistance-history.md` section 1.1) are the precedent: Ethash chips reached 2.1x (Linzhi Phoenix, 2020), +2.9x (Antminer E9, 2022) and 4.8x per joule (Jasminer X4, 2021) with custom memory controllers and on-package memory +and no on-die dataset, which is the `f = 1` end of this curve. + +### 5.1 Inputs + +| Input | Value | Source (URL read 6 October 2026 unless a file is named) | +|---|---|---| +| RTX 5090 class v3, the denominator | 136.1 MH/s at 326 W (power approximate: 328.6 W peak in the 5 October prover-cost run, 323 W after the M11 race); 2.40 microjoules per hash; 17.5 G dependent 4-byte reads per second (CUDA wall), 18.2 (OpenCL event), 415 ns at 256 lanes; a 32-byte sector per read | `docs/bench-log.md` "Counter ASIC 2.0, the numbers" and M11; `docs/benchmarks/repro.md` 2.2 (branch repro-bench) | +| RTX 5090 memory system | 32 GB GDDR7 on a 512-bit bus at 28 Gbps, 1,792 GB/s; 16 devices of 2 GB (16 Gb); 575 W TGP; $1,999 at launch; 96 MB L2; 750 mm^2 | https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/ (512-bit, 32 GB GDDR7, 21,760 cores, 575 W); https://en.wikipedia.org/wiki/GeForce_RTX_50_series (28 Gbps, 1,792 GB/s, $1,999, 96 MB, 750 mm^2); the 16-device count is the brief's and the 32 GB / 2 GB arithmetic | +| RTX 5090 integer rate | 45.2 T op/s (OpenCL event, integer chain) | `repro.md` 2.2 | +| Program work per hash | 64 instructions x 8 iterations = 512, of which 128 loads | `docs/spec/01-lottery-hash.md` 1.4 (the table at lines 70 to 71) | +| GDDR7 organisation | four 10-bit channels per device (8 data bits each); 16 banks per channel; PAM3; 1.1 to 1.2 V; 28 to 32 Gbps per pin now, 48 on the roadmap | https://www.rambus.com/blogs/all-you-need-to-know-about-gddr7/ (channels, PAM3, voltage, rates); https://www.smart-dv.com/memory/gddr7.html (4 channels, 16 banks per channel); so 64 channels and 1,024 banks on the 5090's 16 devices | +| GDDR7 access granularity | 32 bytes per channel access (8 data bits x a burst of 32 beats; the 5090's measured 32-byte sector agrees) | arithmetic on the channel width; the sector is `repro.md` 2.2; the JEDEC burst length itself is behind the paywall, so the 32 B is marked approximate | +| GDDR7 energy, streaming | 4.5 pJ per bit average device power (GDDR6X 6, GDDR6 6.5) | Micron, quoted in the search results for the GDDR7 product brief (the brief's own PDF refused the fetch); Rambus: "over 10 percent less power per bit than GDDR6X"; approximate | +| GDDR7 price | about $20 per 2 GB device (16 Gb, 28 Gbps), September 2026; 3 GB devices $60 to $70 | https://www.trendforce.com/news/2026/09/24/news-micron-reportedly-ends-2gb-gddr7-narrowing-supply-options-for-nvidias-rtx-50-series/ (quoting Tom's Hardware and VideoCardz) | +| HBM3 organisation | 1,024-bit interface, 16 64-bit channels, 32 32-bit pseudo-channels; burst of 8 beats, 32-byte packet; up to 64 banks per channel; 6.4 Gbps per pin, 819 GB/s per stack; core 1.1 V, I/O 0.4 V; 64 GB per stack maximum | https://www.synopsys.com/glossary/what-is-high-bandwitdth-memory-3.html and https://www.synopsys.com/articles/hbm3-ip-dwtb.html; https://www.tomshardware.com/news/hbm3-spec-reaches-819-gbps-of-bandwidth-and-64gb-of-capacity | +| HBM3E | the same organisation, 9.2 to 9.8 Gbps per pin, 1.15 to 1.2 TB/s per stack; 24 GB (8-high) and 36 GB (12-high) | https://en.wikipedia.org/wiki/High_Bandwidth_Memory; https://blogs.sw.siemens.com/semiconductor-packaging/2026/04/24/hbm3e-hbm4-ic-design-guide/ | +| HBM energy per bit, streaming | Samsung's roadmap: HBM2 6.25, HBM3 4.12, HBM3E 4.05 pJ per bit; O'Connor et al. (NVIDIA, MICRO 2017): HBM2 3.97 pJ per bit | https://eureka.patsnap.com/insight/the-hbm-wars-sk-hynixs-dominance-samsungs-roadmap-and-the-looming-threat-of-cyclicality (20 August 2025); https://www.cs.utexas.edu/~skeckler/pubs/MICRO_2017_Fine_Grained_DRAM.pdf | +| HBM random-access energy, the breakdown | HBM2: row activation 909 pJ per 1 KB row; data movement 1.51 + 1.17 pJ per bit; I/O 0.80 pJ per bit at 50 percent activity; 32-byte atom; 16 banks per channel; tRC 45, tRCD 16, tRP 16, tRAS 29, tRRD 2, tFAW 12 ns, 8 activates per tFAW per channel | O'Connor et al., Tables 2 and 3 of the PDF above (text extracted with pdftotext) | +| DRAM row cycle across types | DDR4 tRCD 14, tRAS 33, tRP 14; GDDR5 14, 28, 12; HBM and HBM2 14, 34, 14 ns; 16 banks per rank; GDDR5 page 2 KB, HBM2 page 2 KB | Li, Reddy, Jacob, MEMSYS 2018, Table 2: https://terpconnect.umd.edu/~blj/papers/memsys2018-dramsim.pdf (text extracted with pdftotext); the history's [L1] | +| HBM random-access ceiling, the literature | "Folded Banks" (AMD, ISCA 2025): fine-grained random access on HBM is bound by activate parallelism (tRC, tRRD, tFAW), and a redesign with 8x the activate parallelism gives 6.7x the irregular bandwidth, which says the stock stack sits far under its streaming figure on random reads | https://dl.acm.org/doi/10.1145/3695053.3731111 (abstract; the PDF refused the fetch) | +| HBM3 price | about $200 per 24 GB stack factory gate (HBM3E $300 per 36 GB), October 2026; contract pricing about twice that | https://siliconanalysts.com/data/hbm-pricing (no external source cited there; approximate) | +| Interposer and packaging | CoWoS-S $600 to $900 per H100-class package (8 stacks, about 800 mm^2 of silicon interposer); CoWoS-L 20 to 47 percent more; September 2026 | https://siliconanalysts.com/tools/packaging (public sources only, approximate); a one-stack package is taken at $200, approximate | +| 32-bit integer op energy, N5-class | int32 add 0.06 pJ, int32 multiply 0.52 pJ at 5 nm (7 nm: 0.10 and 0.80; 45 nm, Horowitz 2014: 0.1 and 3.1) | https://mlsysbook.ai/vol1/backmatter/appendix_assumptions.html Table 13, citing Horowitz 2014 and Dally 2021; datapath only, so a 2x pipeline and clock overhead is applied below, approximate | +| On-die SRAM read, 64 bytes from a 256 MiB array | 0.5 nJ, range 0.2 to 1.0 (approximate: Horowitz's 45 nm 1 MB cache at 100 pJ per 64-bit read scaled to N5, plus about 0.6 pJ per bit of global wire across a 128 mm^2 array) | from memory; the sensitivity is shown in every row | +| Cache mirror and recompute die | 256 MiB = 128 mm^2, $46 at N5 headline density; the whole recompute die (50 T op/s of integer logic beside the mirror) taken as a 750 mm^2-class N5 die at about $600 of silicon (70 dies per $20,000 wafer at the `sram-mirror.md` yield model) | `docs/analysis/sram-mirror.md` sections 4 and 5; the $600 is arithmetic on its wafer price and D0, approximate | +| Chip project cost | a 7 nm-class project $50M to $75M all-in; a 28 nm project $5M to $30M; masks $1M to $3M at 28 nm, $10M to $20M at 5 nm | `asic-resistance-history.md` section 2.5 and its [E2] [E3] [E4] | +| Ethash chips, the precedent | Phoenix 2,733 MH/s at about 3,000 W, 2.1x; X4 2,500 MH/s at 1,200 W, 4.8x; E9 2,400 MH/s at 1,920 W, 2.9x per joule | `asic-resistance-history.md` rows 3 and 4 and their [S11] [S12] | + +### 5.2 The op count, counted from the code + +`igneum-pow/src/memhard.rs`, `mixer` (lines 290 to 303) and `derive_items_mask` (lines 503 to 536), under class v3's +`mixer_mult = 8`: + +| Where | Operations | Count | +|---|---|---| +| Per mixer application, the 16-word prologue `(s[i] ^ (rc[i] + rk)) * mul[i]` | 16 xor, 16 add, 16 mul | 48 (32 when `rc[i] + rk`, a per-round constant, is hoisted; a chip hoists it) | +| Per mixer application, 8 quarter rounds of 4 add, 4 xor, 4 rotate | 32 add, 32 xor, 32 rotate | 96 | +| Per mixer application, total | | 144 unhoisted, 128 hoisted | +| Mixer applications per item, `(ITEM_ROUNDS + 1) x m` | 9 x 8 | 72 | +| Per item, the cache-line fold `s[i] ^= line[i]` over 8 rounds, and the 16-word init | 128 xor, 8 mul, 8 add | 144 | +| Per item, total | 72 x 128 + 144 (hoisted) to 72 x 144 + 144 | 9,360 to 10,512 | +| Per hash, 128 items | | 1,198,080 to 1,345,536, plus the program's 512 | + +The spec's "about 130" per application (1.8.4) is the hoisted count plus the fold spread over the applications: 9,360 +per item exactly, so the figures of sections 1 and 2 stand. The energy per application at N5 datapath figures is +16 x 0.52 + 128 x 0.06 = 8.3 + 7.7 = 16.0 pJ (a rotate by a per-day constant is a wire mux on a chip, counted at the +add's 0.06); with the 2x pipeline overhead, 32 pJ. Per item: 72 x 32 pJ = 2.30 nJ of logic plus 8 cache reads at +0.5 nJ = 4.0 nJ plus the fold, 6.3 nJ (3.9 at 0.2 nJ per read, 10.3 at 1.0). Per hash at `f = 0`: 128 x 6.3 = 0.81 +microjoules, before static power. The on-die cache reads cost this chip more energy than the mixer does, which is the +first thing the curve says: the mixer's 9,360 ops are 2.3 nJ of a 6.3 nJ item. + +The rows below use the hoisted count, 9,360 per item (72 x 128 plus the fold), because a chip pays the cheapest +form and the section 1 and 2 rows already price that figure. The two counts differ by 12 percent, so the `f = 0` row +at the unhoisted 10,512 per item (1,346,048 ops per hash) is: 50 T / 1,346,048 = 37.1 MH/s, 0.27x bare, 0.82x with +the 3x factor (against 0.31x and 0.92x); its energy per item 2.44 nJ of logic instead of 2.30, 6.5 nJ in all, 50.7 W at 37.1 MH/s, 1.37 microjoules, +1.75x per joule instead of 1.86x (the 20 W of static power spread over fewer hashes). The `f = 1` rows do not move: they contain no mixer. The multiplier is class v3's shipped +`mixer_mult = 8` (72 applications per item), not the 4 of `mixer-x4.md` section 2. + +### 5.3 The two memory systems as random-read engines + +A dependent 4-byte read opens a row (tRCD), reads one 32-byte atom (tCL and the burst) and must close it before the +same bank opens another (tRC). The rate of random reads a memory system can sustain is the smaller of two ceilings: +banks divided by tRC, and activates per tFAW window per channel times the channels. Neither ceiling moves with the pin +speed, so HBM3E is HBM3 here, and 48 Gbps GDDR7 is 28 Gbps GDDR7. + +| | GDDR7, 16 devices, 512-bit (the 5090's board) | HBM3, one stack | HBM3, eight stacks (an H100-class package) | +|---|---|---|---| +| Channels, banks | 64 channels, 1,024 banks | 16 channels (32 pseudo-channels), up to 1,024 banks | 128 channels, 8,192 banks | +| Bank-bound ceiling, banks / 45 ns (tRC, the HBM2 and GDDR5-class figure, approximate for both) | 22.8 G reads/s | 22.8 | 182 | +| Activate-bound ceiling (GDDR7: 4 per 12 ns per channel, approximate; HBM: 8 per 12 ns per channel, O'Connor Table 2) | 21.3 G reads/s | 10.7 | 85.3 | +| The ceiling carried below | 21.3 G reads/s (the 5090 measures 17.5, 82 percent of it: the card is already near its memory's activate limit) | 10.7 | 85.3 | +| Energy per random 32-byte read (approximate) | 2.0 nJ: 909 pJ activation (one atom per row opened, the HBM2 1 KB row taken for GDDR7's row) plus 4.5 pJ per bit x 256 bits of movement and I/O = 1,150 pJ | 1.2 nJ: the HBM2 sum (909 + (1.51 + 1.17 + 0.80) x 256 = 1,800 pJ) scaled by Samsung's 4.12 / 6.25 | 1.2 nJ | +| Static power (refresh, standby, PLLs; approximate, from memory) | 20 W (about 1.25 W per device) | 4 W | 32 W | +| Controller and PHY die beside it (approximate) | 15 W, $50 | 10 W, $50 | 40 W, $200 | +| Memory dollars | $320 (16 x $20) | $200 plus a $200 one-stack interposer | $1,600 plus $750 CoWoS-S | +| Lanes in flight needed at the ceiling, at a 55 ns controller latency (tRCD 16 + tCL 16 + burst 1.25 + about 20 of controller, approximate) | 1,172 lanes, 73 KB of lane state at 64 B | 588 lanes, 37 KB | 4,692 lanes, 293 KB | +| Reads per second per watt at the ceiling (memory, static and controller) | 0.27 G | 0.40 G | 0.49 G | +| The 5090 for comparison | 17.5 G reads/s at 326 W = 0.054 G per W; 7,262 reads in flight (17.5 G x 415 ns), 22 per watt | | | + +The GDDR7 system's own power at the 5090's 17.5 G reads/s is 17.5 x 2.0 nJ = 35 W plus 20 W static, 55 W: about 17 +percent of the card's 326 W (approximate). The other 83 percent is the GPU: 21,760 ALUs spinning at 92.9 percent +utilisation on 512 program ops per hash, their register files, schedulers, L1 and L2, and the clock trees, against a +45 T op/s integer budget of which the hash uses 512 x 136.1 M = 0.07 T op/s, 0.15 percent. That is the whole case for +the `f = 1` chip: it is the 55 W without the 271. + +### 5.4 The curve + +Per row: reads per hash = 128 f; items recomputed = 128 (1 - f); ops per hash = 128 (1 - f) x 9,360 + 512. +Memory-bound rate = the ceiling / (128 f). Compute-bound rate = 50 T op/s / ops per hash (the section 1 budget). +The rate is the smaller; "binding" names it. Power = rate x (128 f x E_read + 128 (1 - f) x 6.3 nJ) + static (memory, +controller, and 20 W for the recompute die's clocks and leakage when `f < 1`). Energy per hash = power / rate. "Gain, +rate" = rate / 136.1 MH/s (per chip, the section 2 metric); "with the 3x factor" multiplies the compute-bound rate by +3 on the recompute share only, the memory ceiling unchanged. "Gain, joule" = 2.40 microjoules / energy per hash, the +Ethash chips' metric, with the on-die read energy at 0.5 nJ and, in brackets, at 0.2 and 1.0. Dollars = memory + +interposer + controller + $100 of board, plus the $600 recompute die when `f < 1`; no project cost (section 5.6). +Arithmetic: `scratchpad curve.py`, reproduced by hand for the first and last rows below the table. + +| Memory | f | Reads per hash | Items recomputed | Ops per hash | Memory-bound MH/s | Compute-bound MH/s | Binding | Power W | Energy per hash, microjoules | Gain, rate, bare | With the 3x factor | Gain, joule (0.2 / 1.0 nJ reads) | Silicon and memory dollars | $ per MH/s | +|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| +| none (sections 1 to 3) | 0 | 0 | 128 | 1,198,592 | n/a | 41.7 | compute | 53.7 | 1.29 | 0.31x | 0.92x (125.1) | 1.86x (2.44 / 1.33) | $700 | $16.8 | +| GDDR7 | 0.25 | 32 | 96 | 899,072 | 665.6 | 55.6 | compute | 92.3 | 1.66 | 0.41x | 1.23x (166.8) | 1.44x (1.68 / 1.17) | $1,070 | $19.2 | +| GDDR7 | 0.5 | 64 | 64 | 599,552 | 332.8 | 83.4 | compute | 99.4 | 1.19 | 0.61x | 1.84x (250.2) | 2.01x (2.31 / 1.65) | $1,070 | $12.8 | +| GDDR7 | 0.75 | 96 | 32 | 300,032 | 221.9 | 166.6 | compute | 120.7 | 0.72 | 1.22x | 1.63x (221.9, memory) | 3.31x (3.70 / 2.81) | $1,070 | $6.4 | +| GDDR7 | 1 | 128 | 0 | 512 | 166.4 | 97,656 | memory | 77.6 | 0.47 | 1.22x | 1.22x | 5.14x | $470 | $2.8 | +| HBM3, one stack | 0.25 | 32 | 96 | 899,072 | 334.4 | 55.6 | compute | 69.9 | 1.26 | 0.41x | 1.23x (166.8) | 1.91x (2.33 / 1.46) | $1,150 | $20.7 | +| HBM3, one stack | 0.5 | 64 | 64 | 599,552 | 167.2 | 83.4 | compute | 74.1 | 0.89 | 0.61x | 1.23x (167.2, memory) | 2.69x (3.26 / 2.09) | $1,150 | $13.8 | +| HBM3, one stack | 0.75 | 96 | 32 | 300,032 | 111.5 | 166.6 | memory | 69.4 | 0.62 | 0.82x | 0.82x | 3.85x (4.39 / 3.19) | $1,150 | $10.3 | +| HBM3, one stack | 1 | 128 | 0 | 512 | 83.6 | 97,656 | memory | 26.8 | 0.32 | 0.61x | 0.61x | 7.46x | $550 | $6.6 | +| HBM3, eight stacks | 0.25 | 32 | 96 | 899,072 | 2,665.6 | 55.6 | compute | 127.9 | 2.30 | 0.41x | 1.23x (166.8) | 1.04x (1.16 / 0.89) | $3,250 | $58.4 | +| HBM3, eight stacks | 0.5 | 64 | 64 | 599,552 | 1,332.8 | 83.4 | compute | 132.1 | 1.58 | 0.61x | 1.84x (250.2) | 1.51x (1.67 / 1.30) | $3,250 | $39.0 | +| HBM3, eight stacks | 0.75 | 96 | 32 | 300,032 | 888.5 | 166.6 | compute | 144.9 | 0.87 | 1.22x | 3.67x (499.9) | 2.75x (3.02 / 2.40) | $3,250 | $19.5 | +| HBM3, eight stacks | 1 | 128 | 0 | 512 | 666.4 | 97,656 | memory | 174.4 | 0.26 | 4.90x | 4.90x | 9.15x | $2,650 | $4.0 | +| HBM3E, any row | | | | | the same: the activate ceiling does not move with the pin rate | | | | 1 percent lower (4.05 against 4.12 pJ per bit) | | | | $300 per 36 GB stack | | + +Arithmetic, `f = 0`: 50 x 10^12 / 1,198,592 = 41.7 x 10^6; power 41.7 M x 128 x 6.3 nJ = 33.7 W, plus 20 W static = +53.7 W; 53.7 / 41.7 M = 1.29 microjoules; 2.40 / 1.29 = 1.86. Arithmetic, GDDR7 `f = 1`: 21.3 G / 128 = 166.4 MH/s; +power 166.4 M x 128 x 2.0 nJ = 42.6 W, plus 20 + 15 static = 77.6 W; 77.6 / 166.4 M = 0.466 microjoules; 2.40 / 0.466 = +5.14. HBM3 one stack `f = 1`: 10.7 G / 128 = 83.6 MH/s; 83.6 M x 128 x 1.2 nJ = 12.8 W, plus 14 = 26.8 W; 0.321 +microjoules; 7.46x. Lanes: 21.3 G x 55 ns = 1,172. + +What the curve says: + +1. It is monotone. Every memory system's cheapest point is `f = 1`, and the partial rows are worse than both ends on + dollars per MH/s. A recomputed item costs 6.3 nJ and 9,360 ops; a stored one costs 1.2 to 2.0 nJ and no ops. At + $8 to $10 per GB nothing makes a chip maker recompute a 1 to 2 GiB dataset; the partial-store chip is not the + threat and will not be built. The curve matters again only if the dataset outgrows cheap memory, and the + schedule of 1.13.3 (2 GiB plus 0.5 GiB a year; 4 GiB at year 4) stays under one HBM3 stack's 24 GB for the + chain's life. +2. The `f = 0` chip of sections 1 to 3 reads 0.31x per chip at the op budget and 0.92x with the factor, and those + rows stand. Per joule the same chip reads 1.3x to 2.4x (1.86x at 0.5 nJ per on-die read), because its 50 T op/s + of fixed-function logic draws about 54 W where the 5090 draws 326 to do the same work. The two metrics disagree + because they measure different things: the rate row says how many chips match one card, the joule row says what + each hash costs to run. The public claim "under 2x" has so far been the rate row. +3. The `f = 1` rows are over 2x per joule on both memory systems, at every on-die read energy (the on-die cache is + not in them), and over 2x per dollar on GDDR7: 5.1x and $2.8 per MH/s against the 5090's $14.7 at launch price. + The history's band for exactly this chip class is 2.1x to 4.8x per joule (Ethash, rows 3 and 4); the model reads + 5.1x (GDDR7) to 9.2x (eight HBM3 stacks) with an ideal controller and the static allowances above. Read the + history's band as the floor a first chip reaches and the model as the ceiling. + +### 5.5 The f = 1 chip: a GPU's memory system without the GPU + +A dependent read chain cannot be pipelined within a hash: read `r + 1`'s address is read `r`'s data through the +program's registers, so one hash advances one read per memory latency. The rate per chip is therefore lanes in +flight divided by latency, and the only question is how many lanes each memory system lets a controller keep in +flight per watt. The 5090 keeps 7,262 (17.5 G reads/s x 415 ns) at 326 W, 22 per watt. The `f = 1` chip's lanes are +64 bytes of registers each (the hash's eight 32-bit registers and the program counter and nonce), so 1,172 lanes +are 73 KB of SRAM, and its latency is the controller's, about 55 ns, not the GPU's 415 ns of queueing: the chip +holds 6x fewer reads in flight and still reaches the memory's activate ceiling. What it needs to beat the 5090 is +more than 17.5 G reads per second per 326 W, 0.054 G per watt; the GDDR7 system alone gives 0.27 and one HBM3 stack +0.40 (section 5.3). An HBM3 part at the 5090's 326 W: twelve stacks, about 128 G reads/s, 1,000 MH/s, 7.3x per chip +(the eight-stack row at 174 W is 4.9x). HBM3's bank count per stack (1,024, in 32 pseudo-channels) equals the whole +GDDR7 system's 1,024 banks over 64 channels, so its random-read rate per stack is half the 5090 board's (10.7 +against 21.3 G) on the activate count; what HBM wins is energy per read (1.2 against 2.0 nJ: 0.8 pJ per bit of +interposer I/O against a PCB) and the dollars per read per second favour GDDR7 ($320 for 21.3 G against $400 for +10.7 G). Both beat the card per joule by more than 2x because the card's memory system is 17 percent of its power. + +### 5.6 Verdict + +**Over 2x.** The worst case for us is `f = 1` on HBM3 (7.5x per joule for one stack, 9.2x for eight; GDDR7 5.1x), and +the cheapest chip for an attacker is `f = 1` on GDDR7 ($2.8 per MH/s of silicon and memory, a controller die that +needs no advanced node because the mixer is not on it: a 28 nm-class project at $5M to $30M, not the $50M 7 nm +project the on-die cache forces on the `f = 0` chip). The on-die-cache recompute chip stays at 0.92x with the factor +and 1.3x to 2.4x per joule, between "under 1x" and "1 to 2x", and it is not the chip anyone builds. The verdict +changes the public claim: "under 2x" held for the chip that recomputes the dataset; it does not hold for the chip +that stores it. + +### 5.7 What it means for item 2, and what does protect + +Item 2 (a random item-derivation program per day in place of the fixed mixer) removes the 3x fixed-function factor +from the recompute share. The `f = 1` chip has no recompute share: it reads every item from DRAM and never derives +one, so item 2 moves none of the rows that decide the verdict, and neither does a mixer at x16 or x64. The mixer +earns its place against the `f = 0` chip, which at $8 per GB of DRAM nobody builds. Item 2's urgency is therefore +low on this result; it stays a reserve family. What does move the `f = 1` rows, with the arithmetic: + +| Lever | What it does to the f = 1 chip | What it costs the honest cards | Reading | +|---|---|---|---| +| Dataset size | Nothing until the dataset exceeds what one stack or one board holds: 24 GB (HBM3, one stack) or 32 GB (the 5090's own board). The schedule reaches 4 GiB at year 4 | Everything: a 4 GiB step already retires 4 GB cards (`card-lifetime-2026-10-05.md`) | Not a lever against this chip | +| Read granularity | The chip pays the same 32-byte atom the 5090 pays; the 9070 XT pays 64. Wider honest reads (w16, measured, layer 1) give the chip nothing and the 5090 nothing; w64 made the 5090 bandwidth-bound (71.9 MH/s) | w64 costs the 5090 47 percent | Not a lever; the decision to stay at 4 B stands | +| Latency | A longer chain (more reads per hash) scales the chip's rate and the card's rate together; lane state is 64 B, so lanes are free to the chip | Nothing per se | Not a lever: the rate per chip is lanes / latency on both sides and the chip has more lanes per watt | +| The denominator: the 5090's watts at the hash | The gain is 2.40 microjoules over the chip's 0.47; the card's 326 W is 92.9 percent utilisation spinning on loads. At a 250 W cap holding 136.1 MH/s the gain reads 3.9x (GDDR7) and 5.7x (HBM3); at 200 W, 3.2x and 4.6x | None if the rate holds under the cap; the measurement is one PC 2 job (`nvidia-smi -pl 200, 250, 326`, two minutes each, STATUS lines as the rate) | The first measurement to run; it moves every row and costs nothing. Owed (PC 1 is not released; PC 2's budget is the coordinator's) | +| Program work in the latency shadow | The hash hides 512 ops per hash behind 128 reads; the 5090 could hide 330,000 (45.2 T / 136.1 M) before compute binds, the M5 Max about 290,000 and the 9070 XT about 650,000 (their ALU budgets approximate, from memory). Work in the shadow is free in hash rate and costs the card watts it now wastes: at N ops per hash the card rises from 326 toward 575 W (linear, approximate) and the chip must add a core that runs the per-epoch random program, at k times the GPU's 5.5 pJ per op (the 5090's marginal ALU energy, (575 - 326) / 45.2 T). At N = 100,000: the card 401 W, 2.95 microjoules; the chip 1.02 at k = 1, 0.83 at k = 1.5; gain 2.9x and 3.5x. At N = 200,000: 477 W, 3.50; chip 1.57 and 1.20; gain 2.2x and 2.9x. At N = 330,000: 575 W, 4.22; chip 2.28 and 1.68; gain 1.85x and 2.5x | Hash rate none while every card stays latency-bound (under about 290,000 on the M5 Max); watts up to TGP; the verifier N x 32 ops per warp: 3.2 M at N = 100,000, under 1 ms at the 18 G op/s the x8 verifier shows (38 M ops in 2.08 ms), inside the 10 ms gate; `INSTR_COUNT` and `ITERATIONS` are prototype values to be fixed at gate 1 (spec 1.4) | The only lever that moves the f = 1 row toward 2x, and only if the chip's core is no better than a GPU's on a random program (k near 1: RandomX's argument, and the project lead's goal in the brief's words, "build a better GPU than NVIDIA"). It reaches 1.85x at the 5090's full ALU budget and k = 1, not under; combined with a 250 W cap it reads about 1.4x (approximate). It is item 2's idea applied to the program, not to the item derivation | +| The clock (item 4) | The f = 1 chip is a commodity-memory controller project: by the Ethash precedent, 32 months to a first chip at the largest prize, and a chip over 2x at 65 months | None | The issuance trigger and the share-pattern detector matter more than any item-derivation change | + +So: item 2 can wait; the power-cap measurement runs first; the program-length lever is the Counter ASIC 3.0 design +item this analysis adds, with its own six gates (the hash rate per card at N = 50,000, 100,000, 200,000; the watts; +the verifier on a 2019-class core; bit-exactness on three vendors); and the public claim is re-worded until the +measurements land. + +### 5.8 Consequences per user tier + +| Tier | What the verdict means | What is being done | +|---|---|---| +| Home miner, one 8 GB card (any vendor, any OS) | Today, nothing changes: no chip exists, the devnet pays nothing, and the f = 1 chip is a project of $5M to $30M and about 32 months by the precedent. When one lands it runs at 0.3 to 0.5 microjoules per hash; an 8 GB card runs 10 to 20 (approximate: the 9070 XT's 18.6 MH/s at its 304 W board power, from memory, is 16 microjoules) and is the first tier out | The power-cap and program-length measurements (above) decide how far the honest card's joules can fall; the share-pattern detector (item 4) is what tells this miner a chip has arrived | +| One 12 GB card | As the 8 GB tier; the dataset size gives it no protection, since the chip's memory is 24 to 32 GB whatever the card holds | The same | +| One 16 GB card (9070 XT class) | AMD RDNA 4 at 2.4 G reads/s and 304 W (approximate) is 7x worse per joule than the 5090 and 30x worse than the f = 1 chip; a chip ends AMD home mining first | The vendor-share metric (item 7) will show it; nothing in the hash fixes AMD's dependent-read rate | +| One 24 or 32 GB card (5090 class, Apple M5 Max) | The 5090 is the honest best at 2.40 microjoules; the M5 Max at 27.9 MH/s and a GPU power of about 60 to 80 W (approximate, unmeasured) is 2 to 3 microjoules, the same class per joule. Against the f = 1 chip both are 5x to 9x behind per joule in the model, 2x to 5x by the precedent | The two measurements; the 5090 at a 200 W cap (if its rate holds) is 1.47 microjoules and halves the gap | +| A rig | A rig's cost is electricity; per joule it is its cards. Once chips hold the hashrate, a rig at the same tariff earns 1/2 to 1/9 of a chip per watt and leaves. The precedent: ASIC share of Ethash stayed small (about 3 percent) for years because the chips were not cheap enough per dollar at scale, and the dollars-per-MH/s row says this chip is ($2.8 against $14.7) | The issuance trigger (item 4): the bounty and the benchmark live before daily issuance crosses about $50K | +| A pool user | A chip fleet is a few operators at 0.3 microjoules; MoneroCrusher found Monero's at 85 percent of the hashrate by the share pattern | The detector on the observer, item 4, before the public testnet | +| The public claim "under 2x" | It held for the recompute chip at the op budget. Per joule and against the stored-dataset chip the model reads over 2x on both memory systems and the precedent reads 2.1x to 4.8x. The claim as worded is not safe to publish | Re-word to the measured fact (the 5090 runs at 1/128 of its dependent-read ceiling and a chip must out-read it per watt) until the power-cap and program-length rows land; nothing goes to the devnet or the site from this analysis | + +### 5.9 Unverified and owed + +- Every energy figure is a published streaming or breakdown figure applied to random 32-byte reads: HBM2's 909 pJ + activation and 1 KB row stand in for HBM3 and for GDDR7 (GDDR6 rows are 2 KB per 16-bit channel, Li-Reddy-Jacob + Table 2; GDDR7's 8-bit channel row is unknown to me); the static allowances (20 W, 4 W, 32 W), the controller + powers, the 55 ns controller latency and the 2x pipeline overhead are from memory. +- The 5090's 326 W is a peak from a run with the prover on (328.6) and a post-run reading (323); the power at the hash + alone, and under a cap, is unmeasured. PC 1 is not released (the 9070 XT rows are OWED) and no PC 2 job was + published for this item (analysis only; the measurement is named, not run). +- The 50 T op/s budget, the 3x factor and the $600 die are section 1's approximations; the on-die read energy (0.2 to + 1.0 nJ) moves the `f < 1` rows by up to 1.8x and the `f = 1` rows not at all. +- Prices are factory-gate figures from a secondary source (contract pricing about 2x), September to October 2026; + the 5090's $1,999 is the launch price and its 2026 street price is higher (reports of $4,000 and more, approximate), + which makes the chip's dollar advantage larger, not smaller. +- The Ethash chips' internals are not read (no teardown); their gains are the history's rows. +- The GDDR7 burst length and bank-group timing, and HBM3's tFAW per pseudo-channel, are behind the JEDEC paywall; + the activate ceilings use the HBM2 figures and the 5090's measured 17.5 G (82 percent of the GDDR7 ceiling) as the + check that they are the right order. +- The ALU budgets of the M5 Max and the 9070 XT, their power at the hash, and the verifier's cost at N = 100,000 + program ops are estimates; the program-length lever is a design item with its own measurements, not a result.