From 59441ccbd7dd43f20a33101004b7a6f5f4a781dd Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Thu, 8 Oct 2026 11:17:50 +0100 Subject: [PATCH] Class v6 research lane B: the hardware future, five years out (first cut and the full report in one file) docs/analysis/class-v6/hardware-future.md: one table over PIM and PNM (HBM-PIM, AiM, LPDDR5X-PIM, UPMEM), HBM3E, HBM4 and the custom HBM4E base die, fine-grained DRAM on logic, LPDDR6, 3D DRAM, RLDRAM 3 and Folded Banks, FPGA with HBM2e, wafer-scale, CXL, optical I/O, chiplets and UCIe, and SRAM at N2: energy per dependent read, latency, cost per GB, availability to a non-hyperscaler, the k_read and shadow k bands, the edge per joule at zero shadow against the 5090 and the M5 Max, and the one layer of the four that blunts each with a number. The three findings: a 2 GiB SRAM full store on one N2 reticle is the strongest five-year chip (13x to 17x per joule, about USD 0.3 per MH/s) and no layer reaches it (layer 2 moves its capex, not its joules; the shadow at k = 0.5 holds 4.8x); per-bank PIM is blind (1.6 percent of reads in-bank at 2 GiB) and the custom HBM4E base die is the f = 1 chip with its controller in the stack (6.5x to 14x), untouched by layers 1 to 4; the denominator is the wrong card (the M5 Max at 0.78 microjoules halves to thirds every chip edge). Section 5 carries the k bands for the research lane's chip rows; section 7 the four decisions with defaults and deadlines; section 9 the sources with URLs and dates. tools/ci/export-exclude.txt: docs/analysis/class-v6 is research, not public export. Co-Authored-By: Claude Fable 5.1 --- docs/analysis/class-v6/hardware-future.md | 443 ++++++++++++++++++++++ tools/ci/export-exclude.txt | 3 + 2 files changed, 446 insertions(+) create mode 100644 docs/analysis/class-v6/hardware-future.md diff --git a/docs/analysis/class-v6/hardware-future.md b/docs/analysis/class-v6/hardware-future.md new file mode 100644 index 000000000..2ac76e04d --- /dev/null +++ b/docs/analysis/class-v6/hardware-future.md @@ -0,0 +1,443 @@ +# Class v6 research lane B: the hardware future, five years out + +Lane B of the class v6 rotating-family research (the founder's word of 8 October 2026, 11:1x UK: "see if anything can be +optimised, added or invented"). First cut landed 8 October 2026, 11:5x UK; the full report fills the same file. Every +figure carries a label: **measured** (a number read off an instrument in this repository, with the file), **claimed** +(a vendor's or a paper's number, with the URL and the date read), **modelled** (arithmetic on claimed figures by the +method of `docs/analysis/chip-model-v3.md` section 5), **approximate** (from memory or an estimate; the sensitivity is +given). Nothing here is a measurement of a chip. Reading public research is in-house; nothing was paid for or asked of +anyone outside. + +The question, as the coordinator put it: for each memory or packaging line, what does it do to a chip's cost per +dependent random read over a dataset of 1 to 8 GiB that grows with chain state (energy per read, latency, capacity cost +per GB, availability to a non-hyperscaler), what k band does it give the chip five years out, and which ONE of the four +v6 layers (1: per-era parameter draws; 2: the state-sized dataset with a floor; 3: scheduled family epochs; 4: the +(c''') acceptance floor and the F8 uniformity test per era) blunts it, with a number. Where a line beats every layer, +this file says so with the number. + +## 0. One page + +The reads are the hash. Under class v3 the RTX 5090 spends 2.40 microjoules per hash on 128 dependent reads, 18.8 nJ +per read all-in, and the memory system itself spends 2.0 nJ of that (chip-model-v3 5.3, modelled); the card's own +marginal per dependent DRAM read is 10.9 nJ unlocked and 8.7 nJ at the 1,300 MHz lock (measured 8 October 2026, +counter-asic-4-research 15.1a). Every chip in this file is a machine that pays the memory's nanojoule and not the +card's ten, plus whatever the shadow (the program work drawn into the memory wait) forces it to pay at `k` times the +GPU's cost per op. That identity does not change with any technology below; what changes is the memory's nanojoule, +the rate a chip can read at, what a GB costs, and who can buy it. + +**The three findings that change v6's design** + +1. **The strongest five-year chip is not a DRAM chip. It is a 2 GiB SRAM full store on one reticle of merchant N2, and + none of the four layers reaches it.** TSMC N2 reads 38 Mb/mm^2 of SRAM (claimed, IEEE Spectrum, 12 December 2024, + volume in 2025), so 2 GiB is about 452 mm^2 of macro, one die under the 858 mm^2 reticle, roughly USD 400 to 600 of + silicon at a USD 30,000 wafer (approximate). It has no activate ceiling, so its rate is power-bound: about 2,100 + MH/s at 300 W, 0.14 microjoules per hash, **13x to 17x the 5090 per joule at zero shadow and about USD 0.3 per + MH/s** against the 5090's 14.7 and the GDDR7 chip's 2.8 (modelled; the wire energy, 0.5 to 2.0 nJ per read, is the + sensitivity). Layer 2's floor moves its capex, not its joules: at 4 GiB it is two dies, at 8 GiB four, USD 1,000 to + 2,500 of silicon, and its edge per joule falls only from 17x to about 13x because an inter-die hop costs 0.27 nJ + (UCIe, claimed 0.5 pJ/bit). A floor that would blunt it on joules (16 to 32 GiB, eight to sixteen reticles) retires + every honest card under 32 GB first. The only lever that reaches it is the shadow at `k`: with the shadow core on + the same N2 die the ALU band is 0.3 to 0.8 (the record's), and the class v4 premium holds **4.8x at k = 0.5, 2.7x + at k = 1**; at the 5090's whole ALU budget (about 330,000 ops per hash, 575 W) 3.7x and 2.0x. The design change: + layer 1's program-length draw is sized against this chip, not the GDDR7 board, with its lower bound at the shadow + that holds the record's 2.1x today and its upper bound at the honest cards' full latency shadow, re-based every + family epoch (layer 3) on the cards then mining; and the public "2x" line is not reachable in this model against + this chip at any `k` under 1. What slows it is money and time (an N2 project, USD 100 M to 500 M and 18 to 24 + months, approximate), which is the clock of Counter ASIC 3.0 item 4, not a hash property. + +2. **Per-bank processing-in-memory is structurally blind to this hash; the real near-memory threat is the custom + HBM4E base die, and layers 1 to 4 do nothing to it.** HBM-PIM, AiM, LPDDR5X-PIM and UPMEM put a compute unit beside + each bank (or a DPU per 64 MB); a dependent read's next address is uniform over the dataset, so it lands in the same + bank with probability bank bytes over dataset bytes: **1.6 percent at 2 GiB on a 32 MB bank, 0.4 percent at 8 + GiB**, and UPMEM has "no direct communication channel among DPUs" (claimed, the PrIM paper), so the other 98 + percent of reads go to the host. The unit that can follow the chain across banks is a controller on the stack's + base die, which is exactly what TSMC's custom C-HBM4E is: "the custom base die will integrate memory controllers and + PHY" on N3P at 0.75 V, "2x the power efficiency", Micron production 2027, SK hynix HBM4E in 2026 (claimed, + TrendForce, 1 December 2025). That is the `f = 1` chip of the record with its controller moved into the stack: about + 0.9 to 1.0 nJ per read, and HBM4's 32 channels (JEDEC JESD270-4, April 2025) double the activate-bound ceiling per + stack, so **6.5x to 14x per joule at zero shadow** (the ceiling unmeasured, as the record's HBM rows are). The four + layers act on the program, the item map and the capacity; this chip runs any program and holds 36 to 64 GB. What + brakes it until about 2028 is availability (HBM allocated to AI, 20 to 26 week leads, Samsung asking USD 4 to 5 per + Gbit for HBM4 against 1.5 for HBM3E, October 2026) and the shadow at `k`: 4.4x at k = 0.5, 2.6x at k = 1. + +3. **The denominator is the wrong card.** Every chip edge in the record is quoted against the 5090 at 2.40 microjoules. + The Apple M5 Max measured 0.78 microjoules per hash at the GPU-plus-DRAM meter (latency-shadow-2026-10-06, 6 + October 2026), 3.1x the 5090 per joule, and LPDDR6 SoCs (JEDEC JESD209-6, 2025; 14.4 Gbps, 32-byte atoms) are that + tier's next step. Against the M5 Max the GDDR7 chip reads 1.7x, one HBM3 stack 2.4x, the HBM4 base die 3.6x to + 4.4x, the N2 SRAM die about 5x, and with the shadow at k = 0.5 the SRAM die reads about 3x. The honest joule, not + the 5090's, is the chain's resistance, and the same chips are 2x to 3x less frightening against it. The design + change: v6's acceptance floor (layer 4) and the shadow sizing (layer 1) are scored per card tier with the + unified-memory SoC tier as the reference joule, the dataset is kept inside 16 GB unified memory (8 GiB at the top of + the schedule does that), and the public text states the edge over the best honest joule, which is the number a + miner can act on. + +The honest line, in one sentence: five years out a 2 to 8 GiB dataset fits in one to four reticles of merchant SRAM at +USD 500 to 2,500 of silicon, every DRAM line converges on the same 0.5 to 1.0 nJ per read with its controller in the +stack, and no layer of the four touches either; the shadow at the measured `k` band holds 3x to 5x, the honest tier's +own efficiency halves that again, and the clock (project cost against daily issuance) is the wall that is left. + +## 1. The method and the denominators + +| Quantity | Value | Label | Source | +|---|---|---|---| +| The 5090 at class v3 | 136.1 MH/s at 326 W, 2.40 microjoules per hash, 17.5 G dependent reads per second, 415 ns at 256 lanes, a 32-byte sector per 4-byte read | measured | `docs/bench-log.md`; chip-model-v3 5.1 | +| The 5090 per dependent DRAM read, the whole card's marginal | 10.9 nJ unlocked, 8.7 nJ at the 1,300 MHz lock (dram_chase_1g, 18.17 G reads per second) | measured, 8 October 2026 | counter-asic-4-research 15.1a | +| The 5090 L2 hit, per dependent read | 2.4 nJ unlocked, 1.4 nJ at the lock | measured | the same | +| The 5090 per counted ALU op (the class v4 shadow's mix) | 11.3 pJ unlocked, 6.2 pJ at the lock (int_arx); the shadow's own 10.4 and 6.5 | measured | the same; counter-asic-4-research section 1 | +| The 5090 at the 1,300 MHz lock, class v3 | 134.6 MH/s at 223 W, 1.67 microjoules | measured, 7 October 2026 | counter-asic-3-status 7c, the clock grid | +| The class v4 premium `F` on the 5090 | 1.10 microjoules unlocked (145 W), 0.65 at the lock (82 to 88 W) | measured | the same | +| Apple M5 Max, class v3, GPU plus DRAM channels | 27.08 MH/s at 21.0 W, 0.78 microjoules; class v4 at 100,000 ops 1.40 (+16 W); 6.9 pJ per counted op | measured, 6 October 2026 (the SoC's other rails and the wall are not in the 21 W) | `docs/analysis/latency-shadow-2026-10-06.md` section 3 | +| RX 9070 XT, class v3 | about 18.6 MH/s at about 304 W, 10.6 microjoules; 2.4 G reads per second | rate measured, watts approximate | `docs/analysis/horizon/algorithm.md` 5.1 | +| The memory system's own cost per random 32-byte read | GDDR7 2.0 nJ, HBM3 1.2 nJ | modelled on O'Connor et al. 2017 Tables 2 and 3 and Samsung's pJ per bit roadmap | chip-model-v3 5.3 | +| The record's `f = 1` chips at zero shadow | GDDR7 board 166 MH/s at 77.6 W, 0.466 microjoules, 5.1x; one HBM3 stack 83.6 MH/s at 26.8 W, 0.321, 7.5x; eight stacks 666 MH/s at 174 W, 0.262, 9.2x | modelled | chip-model-v3 5.4 to 5.6 | +| The chip's `k` on the shadow's work | ALU 0.3 to 0.8; int8 tile 0.03 to 0.3 (the worse lever); L2 hit 0.1 to 0.3; shuffle 0.4 to 0.7 | the GPU side measured, the chip side claimed | counter-asic-4-research 15.1a | + +The chip's energy per hash is `E = 128 x E_read + P_static / rate + k x F`, the rate the smaller of the memory's +activate-bound ceiling over 128 and the power budget over the per-hash energy, and the edge is the card's microjoules +over the chip's. Two `k` columns appear below: `k_read`, the memory's energy per dependent read over the 5090's measured +10.9 nJ (what the technology itself buys, independent of any program), and the shadow `k` band, the chip's cost per +forced op over the GPU's, which no memory technology changes (it is a logic-node question: a chip core at N2 against a +GPU at N3 or N4 sits at the band's low end, 0.3; at the same node, 0.5 to 0.8; approximate). + +Reads are counted at the part's atom: 32 bytes on GDDR7, HBM3, HBM4 and LPDDR6 (each gives a 32-byte minimum access; +JEDEC, claimed), 64 bytes on an SRAM macro or a line-based part. The class v6 read-width draw (4 to 64 bytes per read +under layer 1) does not move any row: a chip's controller fetches the atom whatever the hash asks for, as the 5090 +fetches its 32-byte sector for a 4-byte load, and the record's w64 reading (the 5090 bandwidth-bound at 71.9 MH/s, +chip-model-v3 5.7) says a wider honest read costs the card and gives the chip nothing. The dataset's size enters only +through capacity cost and die count; its growth (layer 2) enters through the number of dies or stacks a chip must carry +at each family epoch. + +## 2. The table + +Energy is per dependent random read at the part's atom. Latency is the read's own (controller to data), not the GPU's +queueing. Cost per GB is factory-gate where a source gives it, with contract pricing about 2x (Silicon Analysts, October +2026). "Edge" is per joule against the 5090 at class v3 and zero shadow; the bracket is against the M5 Max's 0.78. +"Year" is when a non-hyperscaler could put the part on a board. Every chip-side figure is modelled unless labelled. + +| Technology | Year | Energy per dependent read | Latency | Cost per GB | Availability to a non-hyperscaler | `k_read` (over 10.9 nJ) | Edge at zero shadow, per joule | The layer that blunts it, and by how much | +|---|---|---|---|---|---|---|---|---| +| GDDR7 on a PCB, 16 devices, a 28 nm controller (the record's `f = 1` chip) | now | 2.0 nJ (modelled); 4.5 pJ per bit streaming (claimed, Micron) | about 55 ns controller, tRC about 45 ns | USD 10 (2 GB parts, ending) to 20 to 23 (3 GB parts at USD 60 to 70), September 2026 (claimed, TrendForce) | anyone, through distribution; the 2026 DRAM price cycle (LTAs USD 7.8 to 21 per GB, May 2026, claimed) roughly triples the chip's memory bill against 2025 | 0.18 | 5.1x (1.7x); 166 MH/s at 78 W | none of the four; the shadow at k = 1 holds 2.1x (the record); the dataset's size does not matter because the chip over-provisions capacity for channels (16 devices whatever the dataset) | +| GDDR7 at 36 to 48 Gbps, 4 and 6 GB devices | 2027 to 2028 (claimed, Micron roadmap via Guru3D and OC3D) | the same 2.0 nJ: the activate ceiling and the row energy do not move with the pin rate | the same | per-GB falls, per-channel rises: a random-read chip buys channels, not bytes | anyone | 0.18 | 5.1x | none; denser devices HURT the chip (fewer channels per GB), so this line goes the chain's way | +| HBM3E, one stack, a 28 nm controller on a one-stack interposer | now | 1.2 nJ (modelled); 4.05 pJ per bit streaming (claimed, Samsung) | about 50 ns | USD 8.3 factory gate (USD 300 per 36 GB), about 2x on contract (claimed, Silicon Analysts, October 2026); plus about USD 200 of interposer | Tier-1 volume pricing, 20 to 26 week leads, "supply constrained through 2026" (claimed); a one-stack buyer pays broker prices | 0.11 | 7.5x (2.4x); 84 MH/s at 27 W (ceiling unmeasured, 10.7 G reads per second; 2.3 G at the JEDEC tFAW, which would read 1.8x) | none; the shadow at k = 1 holds 2.4x (the record) | +| HBM4 (JESD270-4, April 2025): 2,048-bit, 32 channels x 2 pseudo-channels, 8 Gbps, 2 TB/s, 36 to 64 GB, VDDQ 0.7 to 0.9 V, a TSMC N12 base die at 0.8 V "1.5x efficiency" (claimed) | 2027 to 2028 for a non-hyperscaler | 1.0 to 1.1 nJ (modelled: the I/O term at 0.8 V against 1.1 V, the 909 pJ row activation unchanged) | about 50 ns | USD 15.3 (USD 550 per 36 GB, October 2026 estimate); Samsung asking USD 4 to 5 per Gbit for 2027 (USD 32 to 40 per GB) (claimed) | allocated to AI accelerators through 2027; "mid-to-high $4 per gigabit" in this month's negotiations (claimed, 2 October 2026) | 0.10 | 4.5x to 11x (1.5x to 3.6x): the 32 channels double the activate-bound ceiling per stack if tFAW is per channel (unmeasured) | none of the four; the shadow at k = 0.5 holds 4.4x, at k = 1 2.6x | +| Custom HBM4E base die (C-HBM4E): the controller and PHY in the stack on N3P at 0.75 V, "2x the power efficiency" (claimed, TrendForce, 1 December 2025); Micron 2027, SK hynix 2026 | 2028 or later for a non-hyperscaler | 0.9 to 1.0 nJ (modelled: no interposer crossing for the controller's traffic) | about 45 ns | HBM4E class, USD 20 to 40, plus a custom base die (USD 50 to 100 per stack, approximate) | by design a per-customer product of the three DRAM makers; the named customers are NVIDIA and Google; a miner-maker of Bitmain's size could commission one (approximate) | 0.09 | 6.5x to 14x (2.1x to 4.4x) | **beats all four layers**; the shadow at k = 0.5 holds 4.4x, at k = 1 2.6x; the shadow core sits on the same N3P base die | +| Fine-grained, hybrid-bonded DRAM on logic (FGDRAM-class 256-byte rows, "tFAW effectively eliminated", 51 percent lower energy per access, 4x bandwidth, claimed, O'Connor et al., MICRO 2017; hybrid-bonded stacks and 3D DRAM on the 2028 to 2030 roadmaps, approximate) | 2029 to 2031 | 0.5 to 0.7 nJ (modelled: 230 pJ activation for a 256-byte row plus about 1.2 pJ per bit of movement over bonded pads) | about 40 ns | HBM class, USD 15 to 40 (approximate) | custom, the three DRAM makers, top customers first | 0.05 | 12x to 20x (4x to 6x); about 125 MH/s per stack at 19 W, 1,000 MH/s at 150 W for eight | **beats all four layers**; the shadow at k = 0.5 holds 5.1x, at k = 1 2.8x | +| Per-bank PIM: Samsung HBM-PIM (a PCU per bank, 2021), SK hynix GDDR6-AiM (16 Gbps, 1.25 V, 2022) and AiMX (32 GB card), Samsung LPDDR5X-PIM (Hot Chips, 25 August 2026, 614 GB/s internal) and LPDDR6-PIM (JEDEC work "substantial progress") | now to 2027 | the host memory's own: 1.8 nJ (HBM2), 2.0 (GDDR6 class); the PIM unit never sees the read it would need | the same | the host part's plus a premium | HBM-PIM and AiM were samples and prototypes; LPDDR5X-PIM mass production "as early as 2027" (claimed) | n/a: no path for a cross-bank dependent read | none beyond the base-die row: the unit serves 1.6 percent of reads at 2 GiB (a 32 MB bank), 0.4 percent at 8 GiB | layer 2, trivially: any dataset over a bank; the number is bank bytes over dataset bytes | +| UPMEM DPU-in-DRAM: 128 DPUs per 8 GB DIMM, 64 MB MRAM per DPU, 350 MHz, MRAM by DMA at most 628 MB/s per DPU, "no direct communication channel among DPUs" (claimed, the PrIM paper, 2021); 1.2 W per 4 Gb chip, 10x DRAM's price at sampling, 1.5x projected (claimed, The Next Platform, 2020) | now | about 0.3 microseconds per 64-byte DMA at 350 MHz (approximate from the paper's alpha-plus-beta model); 2 GiB spans 32 DPUs and every cross-DPU hop is a host round trip of microseconds | microseconds | 1.5x to 10x DDR4 | anyone, in server DIMMs | n/a | none: dead at any dataset over 64 MB | layer 2's floor alone (2 GiB = 32 DPUs with no path between them) | +| LPDDR6 (JESD209-6, 2025): 24-bit channels as two 12-bit sub-channels, 32-byte minimum access, 10.7 to 14.4 Gbps, 4 to 64 Gb dies (claimed, JEDEC via PCWorld and HotHardware) | 2026 to 2028 | 1.5 to 2.0 nJ (approximate: short rows, low-voltage I/O over a PCB, 16 to 32 banks per channel) | about 60 ns | USD 8 to 21 in the 2026 cycle (LTAs, claimed); about USD 3 to 4 in 2024 (approximate) | anyone; it is also the honest SoC tier's memory | 0.15 to 0.18 | about 4x to 5x for a 32-channel controller chip (1.3x to 1.6x against the M5 Max, which already IS this part with a GPU) | none needed: the honest tier converges on it; see finding 3 | +| 3D DRAM (Samsung VS-DRAM 16 layers, VCT 4F2; Neo 3D X-DRAM 230 layers; SK hynix and Micron stacked cells), "from 2030" (claimed, heise, 29 May 2024) | 2030 or later | no vendor claims a latency or activate change; the row energy class is the planar part's (approximate) | the same | lower per GB after 2030 | the three makers | 0.1 to 0.2 | the DRAM rows above | none needed inside five years; it lowers the chip's and the card's capacity cost alike | +| RLDRAM 3 (Micron, 2011): tRC under 10 ns, 16 banks, no activate command, "SRAM-like random access" (claimed, Micron); 576 Mb and 1.125 Gb devices at 2,133 Mb/s (approximate) | now, in small volume | 2 to 4 nJ (approximate: a full small-row access per read; no energy figure found, the datasheet's power calculator exists) | under 10 ns | USD 400 to 800 (USD 50 to 100 per 1.125 Gb device, approximate, a networking part) | anyone, catalogue | 0.2 to 0.4 | about 2 G random reads per second PER DEVICE (16 banks over 8 ns), 32 G for 16 devices: 250 MH/s per board, about 5x per joule (approximate) | layer 2: USD 3,000 to 6,000 of memory at 8 GiB; the line to watch is AMD's "Folded Banks" (ISCA 2025): 8x the activate parallelism in HBM gives 6.7x the irregular bandwidth (claimed, the abstract) | +| FPGA with HBM2e: AMD Alveo V80 (Versal XCV80, 32 GB, 820 GB/s, 190 W, USD 9,495 MSRP, May 2024, claimed), Versal HBM VH1782, Altera Agilex 7 M (two HBM2e stacks, 820 GB/s), Alveo U55C USD 4,747 | now | HBM2 class, 1.8 nJ, but the fabric's controller reaches only the JEDEC tFAW ceiling: 2.3 to 2.4 G reads per second per stack (Shuhai, FCCM 2020; the record) | about 100 ns | USD 300 per GB of card | anyone, 10 to 24 week leads | 0.17 | 0.6x (0.2x): 4.8 G reads per second for two stacks, about 37 MH/s at 150 W, at 5x the 5090's price | none needed; no HBM3E FPGA was found announced (unverified) | +| Wafer-scale SRAM: Cerebras WSE-3, 44 GB SRAM, 21 PB/s, 900,000 cores, 46,225 mm^2 on N5 (claimed, Cerebras and The Next Platform, March 2024); CS-3 about 23 kW and "maybe $2.5 million" (claimed, approximate) | now, by the system | about 7 nJ (approximate): a 512-bit reply crossing about 140 mm of mesh on average at about 0.1 pJ per bit per mm | about 0.5 microseconds across the wafer | USD 50,000 per GB | by the system only | 0.6 | 0.1x to 0.4x: a uniformly random dependent chain is bisection-bound on a 2D mesh (about 110 G reads per second per wafer, up to 500 G with the dataset replicated twenty times), 5 to 22 M reads per second per watt against the 5090's 54 M; 1/1,000 per dollar | none needed; layer 2 removes the replicas as the dataset grows | +| CXL memory pools (CXL 2.0 and 3.x; about 70 ns of controller on top of local DDR5, 100 to 160 ns on Xeon 6 against 75 local, claimed, Introl, February 2026; USD 4 to 7 per GB before the 2026 cycle) | now | DDR5's 2 to 3 nJ plus the link | 170 to 250 ns through a host | USD 4 to 21 | anyone | 0.2 to 0.3 | 0.4x: 500 M 64-byte transactions per second per x8 device at about 25 W (approximate) | none needed; a capacity tool, not a random-read engine | +| Optical I/O: Ayar Labs TeraPHY (UCIe-compliant, about 5 pJ per bit, USD 500 M raised 3 March 2026, claimed); Celestial AI Photonic Fabric (6.2 pJ per bit, about 120 ns round trip, claimed) | 2027 or later | 6x to 8x the interposer's 0.8 pJ per bit per bit moved | plus 120 ns | n/a | n/a | worse than copper for this traffic | no edge: a pooling fabric; the dependent chain wants the memory closer, not farther | none needed | +| Chiplets and die-to-die: UCIe 0.25 to 0.5 pJ per bit (claimed, UCIe via SNIA); BoW 0.5 to 0.7; CoWoS-S USD 600 to 900 per H100-class package, a one-stack interposer about USD 200 (the record) | now | adds 0.27 nJ per 544-bit read that crosses a die boundary | plus 5 to 10 ns per hop | the package: USD 100 to 200 organic with UCIe-S, USD 200 to 900 for 2.5D | UCIe IP from several vendors; CoWoS capacity booked by AI through 2027 (approximate) | n/a | moves the SRAM chip's multi-die rows (below) and lets a 28 nm controller sit beside memory on an organic package | none needed | +| SRAM scaling at 2 nm: TSMC N2 38 Mb/mm^2 HD, +11 percent over N3E (claimed, IEEE Spectrum, 12 December 2024); wafers about USD 30,000, booked to 2028 (claimed, 2026) | now (N2 in volume from 2025) | 1.0 nJ for a 64-byte read from a 452 mm^2 array (approximate; 0.5 to 2.0: the global wire at about 1.3 pJ per bit across a 24 mm die is the term) | 10 to 20 ns | USD 200 to 300 per GB of silicon (one 600 mm^2 die, about USD 400 to 600, holds 2 GiB) | N2 is a merchant node: any customer with a project (Apple, AMD, NVIDIA, MediaTek, Qualcomm; the Bitcoin chip makers are on N3 and N4 class already, approximate) | 0.09 | **13x to 17x (4x to 5x)**: no activate ceiling, power-bound at about 2,100 MH/s per die at 300 W, USD 0.3 per MH/s of silicon; two dies at 4 GiB and four at 8 GiB cost USD 1,000 to 2,500 and read about 13x | **beats layers 1, 3 and 4; layer 2 cuts its capex, not its joules**: one reticle at 2 GiB, two at 4 GiB, four at 8 GiB; the shadow at k = 0.5 holds 4.8x, at k = 1 2.7x; on the M5 Max's joule about 3x at k = 0.5 | + +Density against the schedule: SRAM gains 6 to 11 percent per node every two to three years (N3E +6, N2 +11, claimed), +the dataset doubles every four years (spec 1.13.3: 2 GiB at genesis, 4 GiB at year 4). The SRAM chip loses that race +slowly: its die count doubles each doubling and its per-read energy rises 0.1 to 0.3 nJ per extra hop. It does not lose +it inside five years. + +## 3. What each line does to the four layers, read the other way + +| Layer | What it does to the lines above | Number | +|---|---|---| +| 1. Per-era draws (mixer rounds, op-mix weights, read width, program length, shadow placement) | The read width and the mixer draws move no chip row: every chip in the table stores the dataset and reads the atom. The program length IS the shadow, and it is the only draw that reaches the memory-system chips | at `F` = 1.10 microjoules (class v4) the strongest DRAM chip reads 2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5; at the full 5090 shadow (`F` about 2.0) 2.0x and 3.7x | +| 2. The state-sized dataset with a floor | Kills per-bank PIM and UPMEM outright; raises the SRAM chip's die count and capex; does nothing to any DRAM chip, which carries 24 to 64 GB per stack and over-provisions capacity for channels anyway | PIM serves 1.6 percent of reads at 2 GiB, 0.4 at 8; the SRAM chip's silicon USD 500 at 2 GiB, 1,000 at 4, 2,500 at 8, its joules 17x to 13x | +| 3. Scheduled family epochs every 180 days | None of the chips in the table is fixed-function; the controller chip runs any program and the SRAM chip's shadow core is a sequencer | 0 | +| 4. The (c''') floor and the F8 uniformity test per era | Keeps the hot-set cache at the record's 1.067x ceiling; the L2 hit at 2.4 nJ on the 5090 against 0.2 to 0.5 on a chip means the uniformity test is what stops a small cache from being the chip's edge | 1.067x at the ceiling (the record) | +| A fifth: the denominator | The honest tier's own operating point is the resistance: the 5090 at the lock reads 1.67 microjoules, the M5 Max 0.78; every chip edge halves to thirds against them | the SRAM die 5x and the base die 3.6x to 4.4x against the M5 Max at zero shadow; about 3x with the shadow at k = 0.5 | + +## 4. The lines in detail + +### 4.1 Processing-in-memory and processing-near-memory + +Samsung's HBM-PIM (Aquabolt-XL, February 2021) places a programmable computing unit inside each bank of an HBM2 stack +and reports 2.5x system performance and over 60 percent energy saved on a Xilinx Alveo host (claimed, Samsung, 24 +August 2021). SK hynix's GDDR6-AiM (February 2022) adds compute to a 16 Gbps GDDR6 die at 1.25 V, claims up to 16x on +some AI operations and 80 percent less power, and the AiMX card (2023) carries 32 GB of it (claimed, SK hynix). Samsung's +LPDDR5X-PIM (Hot Chips, 25 August 2026) reaches 614 GB/s inside the package against 76.8 GB/s over the external +interface at LPDDR5X-9600, 3x the tokens per second on Llama 3.1 8B, with mass production "as early as 2027" and +LPDDR6-PIM in JEDEC work (claimed, TrendForce and Sammyfans, 25 to 26 August 2026). UPMEM's DPU-in-DRAM ships: 128 +DPUs per 8 GB DIMM, 64 MB per DPU, 350 MHz, MRAM reached only by DMA with a fixed cost plus a per-byte cost and at most +628 MB/s per DPU for 2,048-byte transfers, and "there is no direct communication channel among DPUs" (claimed, the PrIM +characterisation paper, arXiv 2105.03814, read 8 October 2026). The academic line (IMPICA, ICCD 2016: a pointer-chasing +engine on a 3D stack's logic layer, 1.2x to 1.9x and 10 to 41 percent less energy on linked lists, hash tables and +B-trees, claimed) puts the chaser on the base die, not in the bank, for the same reason this hash defeats the bank +units: the next address is anywhere. + +What it does to the chip's cost per dependent read: nothing good for the attacker at the bank level. The PIM unit's +arithmetic is the wrong kind (FP16 SIMD, not 32-bit integer ARX) and the wrong place: with uniform addresses the chain +leaves the bank after one read with probability 1 minus bank bytes over dataset bytes (98.4 percent at 2 GiB on a 32 MB +bank), and in UPMEM's case leaves the DPU with no path but the host. The near-memory version (a chaser with its own +controller on the base die) is the record's `f = 1` chip, and that is where PNM becomes real: see 4.2. + +Blunted by: layer 2, by the dataset's size alone. k: none (no path). + +### 4.2 HBM3E, HBM4 and the custom base die + +JEDEC's JESD270-4 (16 April 2025; read via eeNews Europe, 18 April 2025, and the search summaries, the Business Wire +and All About Circuits pages refusing the fetch) doubles the channel count to 32 with two pseudo-channels each on a +2,048-bit interface at up to 8 Gbps, 2 TB/s per stack, 4 to 16-high stacks of 24 or 32 Gbit dies up to 64 GB, VDDQ 0.7 +to 0.9 V and VDDC 1.0 to 1.05 V (claimed). TSMC builds the standard HBM4 base die on N12 at 0.8 V for "roughly 1.5x" +efficiency and the custom C-HBM4E base die on N3P at 0.75 V for "2x the power efficiency of today's DRAM +manufacturing", and "the custom base die will integrate memory controllers and PHY components typically housed +separately"; Micron targets 2027 production, SK hynix a first tailored HBM4E in the second half of 2026 with 12 nm for +mainstream and 3 nm for NVIDIA and Google premium designs, Samsung 4 nm now and 2 nm for custom HBM (claimed, +TrendForce, 1 December 2025 and 23 January 2026). Prices: HBM2e USD 120 per 16 GB, HBM3 200 per 24, HBM3E 300 per 36, +HBM4 about 550 per 36 (estimate), factory gate, with contract about 2x and 20 to 26 week leads (claimed, Silicon +Analysts, October 2026); Samsung is asking "mid-to-high $4 per gigabit" for 2027 HBM4 against about 1.5 for HBM3E +(claimed, Sammyfans, 2 October 2026). + +What it does to the chip's cost per read: the row activation (909 pJ per 1 KB row, HBM2, O'Connor Table 3) does not +move; the movement and I/O terms fall with the voltage and the base-die node; a 32-channel stack doubles the activate +parallelism the record's HBM rows are bound by (10.7 G reads per second per HBM3 stack, unmeasured; 2.3 G at the JEDEC +tFAW). The arithmetic on the record's method: HBM4 1.0 to 1.1 nJ per read and 21.3 G reads per second per stack (166 +MH/s at about 36 W, 0.22 microjoules, 11x; at the JEDEC-tFAW ceiling 36 MH/s at 19 W, 0.53, 4.5x); the custom base die +0.9 to 1.0 nJ with its controller inside (166 MH/s at 29 W, 0.18, 14x; 6.5x at the low ceiling). The latency stays in +the 45 to 50 ns class; the capacity (36 to 64 GB) is 4x to 8x any floor layer 2 could set without retiring the honest +cards. + +Who can buy it: through 2027 the stacks are allocated to AI accelerators at Tier-1 volume terms; a custom base die is a +per-customer engagement with the DRAM maker. A chip maker of Bitmain's or Canaan's size (the history's rows: tape-outs +on 7 nm-class nodes, R&D in the tens of millions of dollars a year) could commission one from 2028 (approximate); a +USD 5 M startup cannot. The brake is money and queue, not physics, and it expires. + +Blunted by: nothing among the four. The shadow at `k` = 0.5 holds 4.4x, at `k` = 1 2.6x; the shadow core is logic on +the N3P base die and sits at the band's low end against a GPU on an older node. + +### 4.3 LPDDR6 + +JESD209-6 (2025) gives 10,667 to 14,400 MT/s on a 24-bit channel split into two 12-bit sub-channels, a 32-byte minimum +access with 32 and 64-byte bursts, 4 to 64 Gb dies, lower voltages than LPDDR5, a dynamic efficiency mode and on-die +ECC (claimed, JEDEC via PCWorld, HotHardware and MicrocontrollerTips, 2025). The bank count per channel and the tFAW +are not in the public summaries read (unverified); LPDDR5's 16 banks per channel and a tFAW near 20 ns are the +assumption (approximate). A 32-channel, 64-sub-channel controller chip then reads about 12.8 G dependent reads per +second (approximate), 100 MH/s, at 1.5 to 2.0 nJ per read: about 4x to 5x the 5090 per joule, the GDDR7 board's class +at a lower price per channel. The dies are the cheapest random-access memory a non-hyperscaler can buy outside the +2026 price cycle (USD 3 to 4 per GB in 2024, approximate; USD 8 to 21 in 2026 LTAs, claimed). + +The point of this line is not the chip. The M5 Max already is an LPDDR5X part with a GPU, at 0.78 microjoules per hash +measured at the GPU-plus-DRAM meter, 3.1x the 5090 per joule; the Windows-class LPDDR5X SoCs (NVIDIA's and Qualcomm's +desktop parts, approximate) and the LPDDR6 generation after them are the honest tier's floor. A controller chip on the +same memory beats that tier by 1.3x to 1.6x at zero shadow. The resistance of the chain is set by this tier, and v6 +should say so (finding 3). + +Blunted by: none needed. + +### 4.4 3D DRAM + +Samsung's VS-DRAM (VLSI 2023), 16 stacked layers demonstrated against Micron's 8, VCT 4F2 cells as the stepping stone +with prototypes in 2025 and commercial 3D DRAM "by approximately 2030"; Neo Semiconductor's 3D X-DRAM at 230 layers +and 128 Gbit per die as a concept (claimed, heise 29 May 2024, Yole, ComputerBase). No source read claims a latency or +activate-rate change; the gain is capacity per area (about 3x). For a chip that reads 1 to 8 GiB at random, capacity +per die is not the constraint (channels are), so 3D DRAM lowers the honest card's and the chip's cost per GB alike and +changes no row. Outside the five-year edge. + +### 4.5 CXL memory pools + +CXL 2.0 expanders on Xeon 6 measure 100 to 160 ns against 75 ns local DDR5; a CXL 3.1 controller adds about 70 ns; +pooled DDR5 was USD 4 to 7 per GB before the 2026 cycle (claimed, Introl, 1 February 2026, and the search summaries). +A dependent chain through a host CPU and a PCIe-class link at 64 bytes per transaction is bound by the link's +transaction rate (about 500 M per second per x8 device at about 25 W, approximate): 0.4x the 5090 per watt. A capacity +tool. No edge, no layer needed. + +### 4.6 Wafer-scale + +The WSE-3 holds 44 GB of SRAM at 21 PB/s across 900,000 cores on 46,225 mm^2 of N5; the CS-3 draws about 23 kW and +costs "maybe $2.5 million" (claimed, Cerebras and The Next Platform, 14 March 2024; the power from the search summaries). +A 2 GiB dataset spread over the wafer is read by a dependent chain whose next address is uniformly random across 215 +mm of mesh: the traffic is all-to-all, the mesh is bisection-bound, and the average reply crosses about 140 mm of wire. +At about 0.1 pJ per bit per mm (approximate) a 512-bit reply costs about 7 nJ before routers, 3x the GDDR7 board's +2.0; at about 950 links across the bisection at 32 bits per cycle and about 1 GHz (approximate), the wafer completes +about 110 G reads per second, 500 G with the dataset replicated twenty times in regions, 5 to 22 M reads per second per +watt against the 5090's 54 M. Per joule 0.1x to 0.4x, per dollar one thousandth. The wafer is a streaming machine; this +hash is not streaming. No layer needed; layer 2 removes the replicas as the dataset grows. + +### 4.7 Chiplets, interposers and optical I/O + +UCIe gives 0.25 to 0.5 pJ per bit by package type (claimed, UCIe consortium via SNIA SDC 2022 and 2025 pages); BoW 0.5 +to 0.7; CoWoS-S is USD 600 to 900 per H100-class package and a one-stack interposer about USD 200 (the record). A +read that crosses a die boundary pays about 0.27 nJ (544 bits at 0.5 pJ) and 5 to 10 ns, which is what makes the +multi-die SRAM chip of 4.9 cost 1.1 to 1.3 nJ per read instead of 1.0. Optical I/O (Ayar Labs' TeraPHY at about 5 pJ +per bit, UCIe-compliant, a USD 500 M round on 3 March 2026; Celestial AI's Photonic Fabric at 6.2 pJ per bit and about +120 ns round trip; claimed) is 6x to 8x the interposer's energy per bit and adds latency: it pools memory across +packages, which this traffic never wants. No row moves. + +### 4.8 FPGA with HBM + +The Alveo V80 (Versal XCV80, 32 GB HBM2e as two 16 GB stacks, 820 GB/s, 190 W, USD 9,495 MSRP, May 2024) and the U55C +(USD 4,747) are what a non-hyperscaler can buy today with HBM on it; Altera's Agilex 7 M-series carries the same two +HBM2e stacks at 820 GB/s (claimed, AMD, Wccftech, The Next Platform). The record's reading stands: an HBM2 stack under a +soft controller reaches the JEDEC tFAW ceiling, 2.3 to 2.4 G random reads per second (Shuhai, FCCM 2020, Figure 7), +so two stacks give about 4.8 G, about 37 MH/s at about 150 W (approximate), 0.6x the 5090 per joule at 5x its price. No +HBM3E FPGA was found announced in the pages read (unverified). No layer needed. + +### 4.9 SRAM at 2 nm: the full store on one reticle + +TSMC's N2 reads 38 Mb/mm^2 of high-density SRAM, 11 percent over N3E (claimed, IEEE Spectrum, 12 December 2024; +volume from 2025); wafers are about USD 30,000 and "booked to 2028" (claimed, tech-insider, 2026). The arithmetic: 2 +GiB is 17,180 Mbit, 452 mm^2 of macro; with periphery, a controller, the lanes' registers and a shadow core, a 550 to +650 mm^2 die under the 858 mm^2 reticle; about 95 gross dies per wafer, 55 to 70 percent good with row and column +repair (approximate), USD 400 to 600 of silicon. No DRAM, no interposer, an organic package. Under class v5 the +dataset's items are leaves of the chain state refreshed per window; the chip rewrites 2 GiB per window at on-die +bandwidth, the same 32 ms every GPU pays (the record, chip-model-v3 5.10), and holds a node or shares one across a farm +as the record prices. + +Energy per read: the macro's 64-byte read, about 0.1 nJ (approximate), plus the global wire. The record took 0.6 pJ per +bit of wire across a 128 mm^2 array; scaling with the side of the die gives about 1.3 pJ per bit across 600 mm^2, 0.68 +nJ for 512 bits, so about 1.0 nJ per read with the controller, range 0.5 to 2.0 (the sensitivity of every SRAM figure +here). Latency 10 to 20 ns. No activate ceiling, no tFAW, no refresh: the rate is power-bound. At 300 W with 30 W of +static and controller power: 270 W over 128 nJ per hash is about 2,100 MH/s per die, 0.14 microjoules per hash, 17x +the 5090 at zero shadow (13x at 1.3 nJ, 8x at 2.0, 30x at 0.5); USD 0.25 per MH/s of silicon, about 0.4 with the +board. Two dies at 4 GiB: half the reads cross one UCIe hop, 1.14 nJ, 15x, USD 1,000. Four dies at 8 GiB: 1.3 nJ, 13x, +USD 2,000 to 2,500, a 4-die organic package. The honest cards hold 8 GiB fine (16 GB and up); a floor that pushed the +dataset past what a package can hold (16 to 32 GiB, eight to sixteen reticles at USD 5,000 to 10,000 and 1.5 to 2.0 +nJ per read, still 7x to 9x) would retire every honest card under 32 GB first. Layer 2 therefore sets this chip's +capex and die count, not its joules, and the floor's number is a card-lifetime decision, not a chip decision. + +The project: the history's IBS figures (5 nm USD 416 M to 542 M, 3 nm 590 M) price an SoC; an SRAM array with a +controller and a sequencer core is simpler and the startup figure ("$50M to $75M" for 7 nm, SemiAnalysis) is the +better guide, so USD 100 M to 500 M at N2 (approximate) and 18 to 24 months to a first chip (approximate). That is the +clock of Counter ASIC 3.0 item 4 and the daily-issuance threshold of the history (chips at USD 20 K to 50 K of daily +issuance for compute-bound hashes; 32 months for Ethash at the largest prize): the SRAM chip arrives when the prize +pays for an N2 project, and nothing in the hash moves that date. + +Blunted by: layer 2 on capex only (USD 500 to 2,500 across the floor's range); the shadow at `k` on joules: with the +shadow core on the same N2 die against a GPU on N3 or N4 the ALU band sits at 0.3 to 0.5; at `F` = 1.10 the chip reads +4.8x at k = 0.5 and 2.7x at k = 1; at the 5090's whole latency shadow (about 330,000 ops per hash, `F` about 2.0, +575 W) 3.7x and 2.0x; against the M5 Max's joule about 3x at k = 0.5. + +### 4.10 RLDRAM and the activate-free line + +RLDRAM 3 (Micron, 2011; ISSI second-sourced) has a tRC under 10 ns, 16 banks per device, no separate activate +command and "SRAM-like random access" (claimed, Micron's product page and the 2011 announcements), in 576 Mb and 1.125 +Gb devices at up to 2,133 Mb/s (approximate). One device completes about 2 G random reads per second (16 banks over 8 +ns), six times a GDDR7 device's share of the 5090 board's 21.3 G; sixteen devices, 2 GiB, about 32 G reads per second, +250 MH/s per board. The energy per read is not published in anything read (a power calculator exists); a full +small-row access per read at an old node reads 2 to 4 nJ (approximate), so about 5x per joule, at USD 400 to 800 per +GB (approximate, a low-volume networking part). It is the proof that an activate-free DRAM exists, and AMD's "Folded +Banks" (ISCA 2025, with AMD Research: 8x the activate parallelism in a 3D-stacked HBM gives 6.7x the irregular +bandwidth, claimed from the abstract; the PDF refused the fetch) is the same idea on the HBM roadmap. If a DRAM maker +ships it in HBM, the activate ceilings in the record's HBM rows rise 6x to 8x and the HBM chip's rate per stack with +them; its energy per read falls by the row-size term (O'Connor's FGDRAM: 51 percent). That is the DRAM-on-logic row. + +Blunted by: layer 2 on RLDRAM's capacity cost (USD 3,000 to 6,000 at 8 GiB); nothing on the HBM version. + +## 5. The k bands for the research lane's chip rows + +For `docs/design/class-v6-rotating-family.md`. The shadow `k` is the chip core's energy per forced op over the GPU's +at the same operating point (the record's ALU band 0.3 to 0.8 on the 5090's measured 6.2 to 11.3 pJ per op); the +memory `k_read` is the technology's energy per dependent read over the 5090's measured 10.9 nJ. "Edge" is per joule at +zero shadow; "with the shadow" is at the class v4 premium `F` = 1.10 microjoules on the 5090 (`E_chip` = `E_mem` + `k F`). + +| Chip row | Year | `E_read` nJ | Rate per chip, MH/s | `E_hash` at zero shadow, microjoules | Edge over the 5090 (2.40) | Edge over the M5 Max (0.78) | `k_read` | Shadow `k` band | With the shadow, k = 0.5 / 1 | Silicon and memory, USD per MH/s | +|---|---|---|---|---|---|---|---|---|---|---| +| GDDR7 board, 28 nm controller (the record) | now | 2.0 | 166 | 0.466 | 5.1x | 1.7x | 0.18 | 0.5 to 0.8 (28 nm core: the high end) | 3.3x / 2.1x | 2.8 (2025 memory), about 7 in the 2026 cycle | +| HBM3E, one stack | now | 1.2 | 84 (ceiling unmeasured) | 0.321 | 7.5x | 2.4x | 0.11 | 0.3 to 0.8 | 3.8x / 2.4x | 6.6 | +| HBM3E, eight stacks | now | 1.2 | 666 | 0.262 | 9.2x | 3.0x | 0.11 | 0.3 to 0.8 | 4.1x / 2.5x | 4.0 | +| HBM4, one stack, N12 base die | 2027 to 2028 | 1.0 to 1.1 | 166 (36 at the JEDEC tFAW) | 0.22 (0.53) | 11x (4.5x) | 3.6x (1.5x) | 0.10 | 0.3 to 0.8 | 4.4x / 2.6x | about 5 | +| Custom HBM4E base die, N3P, controller in the stack | 2028 or later | 0.9 to 1.0 | 166 (36) | 0.18 (0.37) | 14x (6.5x) | 4.4x (2.1x) | 0.09 | 0.3 to 0.5 (an N3P core) | 4.4x / 2.6x | about 5 | +| DRAM on logic, FGDRAM-class rows, hybrid bonded | 2029 to 2031 | 0.5 to 0.7 | 125 per stack, 1,000 for eight | 0.15 | 16x (12x to 20x) | 5.2x | 0.05 | 0.3 to 0.5 | 5.1x / 2.8x | about 4 (approximate) | +| SRAM full store, one N2 reticle, 2 GiB | 2027 to 2028 (an N2 project) | 1.0 (0.5 to 2.0) | about 2,100 at 300 W | 0.14 | 17x (8x to 30x) | 5.6x | 0.09 | 0.3 to 0.5 (an N2 core) | 4.8x / 2.7x | 0.25 to 0.4 | +| SRAM full store, two N2 dies, 4 GiB | the same | 1.14 | about 1,850 | 0.16 | 15x | 4.9x | 0.10 | 0.3 to 0.5 | 4.6x / 2.7x | 0.55 | +| SRAM full store, four N2 dies, 8 GiB | the same | 1.3 | about 1,600 | 0.185 | 13x | 4.2x | 0.12 | 0.3 to 0.5 | 4.5x / 2.6x | 1.3 to 1.6 | +| LPDDR6 controller chip, 32 channels | 2027 | 1.5 to 2.0 | about 100 | 0.25 to 0.30 | 4x to 5x | 1.3x to 1.6x | 0.15 to 0.18 | 0.3 to 0.8 | 3.1x / 2.0x | about 3 (approximate) | +| Per-bank PIM, UPMEM, FPGA with HBM2e, wafer-scale, CXL, optical | | | | | under 1x or no path | | | | | | + +Arithmetic, the SRAM row: 270 W over (128 x 1.0 nJ) = 2.11 G hashes per second; 300 W over 2.11 G = 0.142 +microjoules; 2.40 over 0.142 = 16.9x; with the shadow at k = 0.5: (0.142 + 0.55) over (2.26 + 1.10) = 0.692 over 3.36 = +4.86x; at k = 1: 1.242 over 3.36 = 2.71x. The base-die row: 166 M x 128 x 0.95 nJ = 20.2 W plus 4 static plus 5 +controller = 29.2 W; 29.2 over 166 M = 0.176 microjoules; 2.40 over 0.176 = 13.6x. The M5 Max column divides 0.78 by +the same `E_hash`. Every chip-side figure is modelled; the GPU-side figures are the record's measurements. + +## 6. Consequences per tier + +| Tier | What this file means | What is being done | +|---|---|---| +| Home miner, one 8 GB card | Nothing changes today: no chip exists, and the first one in this file (an N2 SRAM die or an HBM4 base-die chip) is a USD 100 M-class project with a 2028-class date. When one lands it runs at 0.14 to 0.22 microjoules per hash against this card's 10 to 20; this tier is the first out, as the record says. The dataset's floor decides this tier's life more than any chip does: 4 GiB at year 4 (the spec's schedule) retires it then | the shadow sizing and the floor are the founder's numbers to set (section 7); the share-pattern detector (Counter ASIC 3.0 item 4) is what tells this miner a chip has arrived | +| One 16 GB card (9070 XT class) | Holds 8 GiB with room; AMD's 2.4 G reads per second at 304 W is 7x behind the 5090 per joule and 50x to 100x behind the chips here | the vendor-share metric; nothing in the hash moves AMD's dependent-read rate | +| One 24 or 32 GB card, the 5090 at the lock | 1.67 microjoules at the 1,300 MHz lock; the chips here are 8x to 12x ahead per joule at zero shadow and 2.5x to 4x with the class v4 shadow at k = 0.5 | the Ember knob carries the lock rows; the shadow's upper bound at the full ALU budget is the lever this file sizes | +| The unified-memory SoC (M5 Max, LPDDR5X and LPDDR6 desktops) | 0.78 microjoules at the GPU-plus-DRAM meter: the honest tier the chips beat least (1.7x to 5.6x at zero shadow, about 3x with the shadow). Capex-poor (27 MH/s per USD 4,000 machine) but joule-rich | finding 3: v6 scores the floor and the shadow per tier with this tier as the reference joule; the dataset stays inside 16 GB unified memory | +| A rig | A rig's cost is electricity; against an N2 SRAM chip at USD 0.3 per MH/s and 0.14 microjoules it earns 1/10 to 1/17 of a chip per watt and leaves when chips hold the hashrate | the issuance trigger: the bounty and the benchmark live before daily issuance crosses about USD 50 K (the record) | +| A pool user | A chip fleet is a few operators; the share-pattern detector is the warning | the detector on the observer, before the public testnet (the record) | +| The public claim | "Under 2x" is not reachable against any chip in this file at a `k` under 1. The honest sentence is: the strongest chip five years out beats a 5090 by 2.6x to 2.8x per joule at k = 1 and about 4.5x at k = 0.5 with the shadow on, and the best honest SoC by about 3x; and it costs an N2 project | the research lane's synthesis carries the number; nothing from this file goes to the site or the devnet | + +## 7. Decisions this raises for the founder + +Each carries a default and a deadline; silence means the default. + +1. **Size the shadow against the SRAM chip, not the GDDR7 board.** Layer 1's program-length draw gets a lower bound at + the length that holds the record's 2.1x on the GDDR7 chip today and an upper bound at the honest cards' full latency + shadow (the 5090's about 330,000 ops unlocked, about 150,000 at the lock; the M5 Max about 290,000; the 9070 XT + about 650,000; approximate from the record), re-based at every family epoch on the cards then mining. Default: the + research lane writes the draw with these bounds into the synthesis; the hash lane measures the rate and the watts at + the upper bound on the 5090 and the M5 Max before the first v6 era is cut. Deadline: 20:00 UK today (the synthesis). +2. **The floor (layer 2).** 2 GiB at genesis is one N2 reticle; 4 GiB two; 8 GiB four. The floor moves the chip's + capex (USD 500 to 2,500), not its joules (17x to 13x), and 8 GiB is the last size inside 16 GB unified memory and + the 16 GB card tier. Default: the spec's schedule stands (2 GiB plus 0.5 GiB a year, doubling at year 4), chosen on + card lifetime; the floor is not a chip lever and this file does not ask to raise it. Deadline: none; a note in the + synthesis. +3. **The reference joule (a fifth layer, or a rule).** Resistance is stated against the best honest joule (the + unified-memory SoC tier, then the 5090 at the lock), not the unlocked 5090. Default: the research lane adopts it in + the synthesis's chip rows (section 5's M5 Max column) and the public text, when there is one, carries the edge over + the best honest joule. Deadline: 20:00 UK today. +4. **The clock.** The chips in this file are USD 100 M-class projects with 2028-class dates; the issuance trigger, the + benchmark and the share-pattern detector of Counter ASIC 3.0 item 4 are what decide when they are built. Default: + unchanged from the record. Deadline: none. + +## 8. Unverified and owed + +- Every chip-side energy figure is modelled on the record's method (O'Connor's HBM2 breakdown, Samsung's pJ per bit + roadmap, the 909 pJ row activation standing in for HBM3, HBM4 and GDDR7); the SRAM wire figure (0.6 pJ per bit at 128 + mm^2, scaled with the die's side) is from memory and moves the SRAM rows by 2x either way; the UCIe hop (0.5 pJ per + bit) is the consortium's claim; the FGDRAM row size (256 bytes) and its 51 percent are the paper's simulation. +- HBM4's banks per pseudo-channel and tFAW per channel are behind the JEDEC paywall (the Business Wire and All About + Circuits pages refused the fetch); the "2x the activate ceiling" rests on 32 channels with tFAW per channel, as the + record's HBM rows rest on the same assumption at 16; both unmeasured. The AWS F2 hour the record names (chip-model-v3 + 5.3) is still the one measurement that would settle the HBM2 figure, and an HBM3 or HBM4 part is not rentable at the + controller level by anyone outside a hyperscaler today. +- LPDDR6's bank count and tFAW are not in the summaries read; the row is LPDDR5's structure (approximate). +- RLDRAM 3's energy per read and its 2026 price are not published in anything read; the row is approximate. +- The Cerebras mesh figures (link width, clock, bisection) are approximate; the fabric bandwidth was not on the pages + read (The Next Platform gives only its change against WSE-2). +- Dates: the DRAM-on-logic and hybrid-bonded HBM timelines ("2028 to 2031") are approximate; the sources read give + 3D DRAM "from 2030" and C-HBM4E production in 2027; the hybrid-bonding pages were not fetched (the TrendForce page + refused). The N2 project cost and time are approximate. +- The web-search budget of this session ran out at about 11:3x UK after 26 searches; the remaining facts were read by + direct fetch of the pages named in section 9. Pages that refused (403): JEDEC's Business Wire release, All About + Circuits, ACM's Folded Banks page, OC3D's GDDR7 capacity note, All About Circuits' HBM-PIM note. Their figures are + carried from the search summaries and marked claimed. +- Nothing was run on the Mac; nothing was built or benchmarked anywhere. The one measurement this file would want next + is on the record's queue already: the 5090 and the M5 Max at the shadow's upper bound (decision 1). + +## 9. Sources (URL and the date read; all read 8 October 2026 unless a file is named) + +Repository: `docs/analysis/chip-model-v3.md` (sections 5.1 to 5.11, 6); `docs/analysis/counter-asic-4-research.md` +on branch counter-asic-4 at fb61ed4b (sections 1, 15.1a, 20.4); `docs/analysis/asic-resistance-history.md` (2.5, 2.6); +`docs/plans/counter-asic-3-status.md` (7c, the clock grid); `docs/spec/01-lottery-hash.md` (1.8, 1.13.3, 1.14); +`docs/analysis/latency-shadow-2026-10-06.md`; `docs/analysis/sram-mirror.md`. + +- JEDEC HBM4 (JESD270-4, 16 April 2025): https://www.eenewseurope.com/en/hbm4-standard-doubles-channel-count-for-ai-boost (18 April 2025); https://hothardware.com/news/jedec-finalizes-hbm4-spec; https://www.businesswire.com/news/home/20250416843598/en (refused) +- TSMC HBM4 and C-HBM4E base dies: https://www.trendforce.com/news/2025/12/01/news-tsmc-unveils-custom-c-hbm4e-details-n3p-logic-dies-reportedly-target-2x-efficiency-gain/ (1 December 2025); https://www.trendforce.com/news/2026/01/23/news-samsungs-custom-hbm4e-design-reportedly-aimed-for-mid-2026-parallels-sk-hynix-and-micron/ (23 January 2026) +- HBM prices: https://siliconanalysts.com/data/hbm-pricing (October 2026); https://www.sammyfans.com/2026/10/02/samsung-seeks-more-than-3x-hbm3e-pricing-for-hbm4/ (2 October 2026); https://www.trendforce.com/news/?p=62555 +- GDDR7 prices and roadmap: https://www.trendforce.com/news/2026/09/24/news-micron-reportedly-ends-2gb-gddr7-narrowing-supply-options-for-nvidias-rtx-50-series/ (24 September 2026); https://www.guru3d.com/story/micron-fiveyear-roadmap-shows-24gb-36gbps-gddr7-in-2026; https://overclock3d.net/?p=314548 (refused; the 4 and 6 GB parts in 2027 to 2028 from the search summary) +- DRAM prices 2026: https://wccftech.com/mobile-dram-prices-expected-to-increase-by-100-quarter-over-quarter-as-long-term-agreements-now-getting-signed-at-prices-as-high-as-21-gb/ (4 May 2026); https://tech-insider.org/ddr5-ram-prices-2026/ +- Samsung HBM-PIM: https://news.samsung.com/global/samsung-brings-in-memory-processing-power-to-wider-range-of-applications (24 August 2021); https://www.allaboutcircuits.com/news/beyond-high-bandwidth-memory-samsung-breaks-processing-in-memory-into-AI-applications/ (refused) +- SK hynix AiM and AiMX: https://news.skhynix.com/developed-processing-in-memory (2022); https://www.hc2024.hotchips.org/assets/program/conference/day1/11_HC2024.SKhynix.GuhyunKim.rev920240822.pdf +- Samsung LPDDR5X-PIM and LPDDR6-PIM: https://www.sammyfans.com/2026/08/25/samsung-unveils-lpddr5x-pim-dram/ (25 August 2026); https://www.trendforce.com/news/2026/08/26/news-samsungs-4nm-gaia-could-mark-first-pim-commercialization-in-ai-pcs-mass-production-as-early-as-2027/ (26 August 2026); https://en.fnnews.com/news/202609060934024282 +- UPMEM: https://arxiv.org/pdf/2105.03814 (the PrIM characterisation, text extracted with pdftotext: Table 1, section 3.2, the inter-DPU statement); https://www.nextplatform.com/2020/02/04/putting-in-memory-processing-through-the-paces/ (4 February 2020); https://old.hotchips.org/hc31/HC31_1.4_UPMEM.FabriceDevaux.v2_1.pdf +- IMPICA: https://ghose.cs.illinois.edu/papers/16iccd_impica.pdf (ICCD 2016) +- Fine-grained DRAM: https://www.cs.utexas.edu/~skeckler/pubs/MICRO_2017_Fine_Grained_DRAM.pdf (text extracted with pdftotext: 3.92 pJ per bit per HBM2 access, 1.21 of it activation, the 256-byte row, tFAW "effectively eliminated", 51 percent) +- Folded Banks (ISCA 2025): https://dl.acm.org/doi/10.1145/3695053.3731111 (refused today; the abstract as the record read it on 6 October 2026); https://wantongli.ucr.edu/news/announcement22-ISCA%202025 +- LPDDR6: https://www.pcworld.com/article/2845759/lpddr6-memory-standard-announced-as-ddr5-dram-takes-over.html; https://hothardware.com/news/jedec-lpddr6-standard-released; https://www.microcontrollertips.com/what-is-jesd209-6-and-why-is-it-important-for-edge-ai/ +- CXL: https://introl.com/blog/cxl-memory-expansion-pooling-disaggregated-memory-ai-data-center-2025 (1 February 2026); https://www.snia.org/sites/default/files/2025-09/SNIA-SDC25-Peethambaran-CXL-as-scalable-cost-effective-Memory.pdf +- Cerebras: https://www.cerebras.ai/chip; https://www.nextplatform.com/2024/03/14/cerebras-goes-hyperscale-with-third-gen-waferscale-supercomputers/ (14 March 2024); https://sacra.com/c/cerebras-systems ("a couple million per system") +- FPGA with HBM: https://wccftech.com/amd-announces-mass-production-of-the-alveo-v80-compute-accelerator-9495-price-tag/ (17 May 2024); https://www.amd.com/en/products/accelerators/alveo/u55c/a-u55c-p00g-pq-g.html; https://www.nextplatform.com/2022/03/08/a-cornucopia-of-memory-and-bandwidth-in-the-agilex-m-fpga; https://arxiv.org/abs/2005.04324 (Shuhai) +- 3D DRAM: https://heise.de/en/news/Huge-RAM-3D-DRAM-with-multiple-layers-planned-from-2030-9738064.html (29 May 2024); https://www.yolegroup.com/industry-news/samsung-reveals-16-layer-3d-dram-plans-with-vct-dram-as-a-stepping-stone/ +- RLDRAM 3: https://www.micron.com/products/memory/dram-components/rldram-memory; https://newelectronics.co.uk/content/news/micron-unveils-third-generation-rldram-technology (2011); https://arxiv.org/pdf/1810.07059 +- SRAM at N2: https://spectrum.ieee.org/tsmc-n2-2670436570 (12 December 2024); https://marklapedus.substack.com/p/intel-tsmc-tout-sram-breakthroughs; wafer prices https://tech-insider.org/tsmc-2nm-wafer-price-2026/ +- UCIe and BoW: https://www.snia.org/sites/default/files/2025-05/SNIA-SDC22-Sharma-Universal-Chiplet-Interconnect-Express.pdf; https://www.3dtested.com/news/new-ucie-chiplet-standard-supported-by-intel-amd-and-arm +- Optical: https://www.theregister.com/2026/03/03/ayar_labs_500m/ (3 March 2026); https://www.allpcb.com/allelectrohub/photonic-interconnects-aim-to-solve-ai-memory-bottlenecks (the Celestial AI figures) +- H100 pointer-chase latency (353 ns, claimed): https://arxiv.org/pdf/2608.15764 diff --git a/tools/ci/export-exclude.txt b/tools/ci/export-exclude.txt index 76af5a1a3..eeeb9b8f6 100644 --- a/tools/ci/export-exclude.txt +++ b/tools/ci/export-exclude.txt @@ -19,3 +19,6 @@ docs/analysis/mission # copied logs are research and operations documents, not public export; the public text is the served sentence main landed. docs/analysis/cryptanalysis docs/plans/cryptanalysis +# 8 October 2026: the class v6 research lanes (the hardware future, the rotating-family research): internal research written +# for the founder on the chip model, quoting the operations record; the public text is whatever main lands from the synthesis. +docs/analysis/class-v6