diff --git a/docs/analysis/class-v6/floor/invention.md b/docs/analysis/class-v6/floor/invention.md new file mode 100644 index 000000000..481147ee3 --- /dev/null +++ b/docs/analysis/class-v6/floor/invention.md @@ -0,0 +1,298 @@ +# Class v6 floor lane 5, invention: what none of the four floor lanes covers, checked against the identity + +8 October 2026, 12:48 to 19:xx UK, branch `class-v6-floor-invention` from `counter-asic-4` (base 91117093, the sound per-load form's 5090 rows). The open-ended lane of the floor research: four other lanes work the known levers (SM-sparse gating, the shadow's `k` from RTL synthesis, the SRAM full-store chip and the dataset floor, the honest denominator per tier); this file finds what they do not cover. Every candidate is checked against the identity first and killed or kept with a number. Every number carries a label: **measured** (a card on a named job, the file named), **modelled** (arithmetic on the chip model's cited inputs, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling). Nothing here changes a consensus object, a pack or a served number. Research and reasoning ran on this Mac; no build, no benchmark and no miner ran anywhere for this file. + +## 0. One page + +**The identity, as the record states it** (`docs/analysis/counter-asic-4-research.md` section 2): against the chip that stores the dataset, `edge = (E_card + F) / (E_mem + k F)`; at zero premium the edge is `E_card / E_mem` (3.6x on a 5090 at its 1,300 MHz knee against the GDDR7 board, measured card, modelled chip), and a shadow op with `k` under 1 never fully closes it. The one thing with `k = 1` by construction is the DRAM itself: the same 16 devices on both sides. The record forced one DRAM operation per read (the activate and one 32-byte sector) and called every wider read "nothing to force: the chip already pays the sector". + +**The one new lever this lane finds: the second sector.** The dataset item is 64 bytes (16 words, spec 1.5 and the item table). The hash reads one word of it, which moves one 32-byte sector on the 5090 and on the chip's GDDR7. Reading the whole item (the read-width branch's class `w16`, built and measured 5 October, bit-exact on four runtimes, fingerprint `e7c890445b47af60`) moves two sectors: on the chip model's own inputs that is the data-movement term paid twice (4.5 pJ per bit x 256 bits = 1.15 nJ, Micron, claimed) on one activate (909 pJ), so the chip's energy per read rises from 2.0 to 3.2 nJ (+56 percent, modelled) and its energy per hash from 0.466 to 0.62 microjoules (+33 percent; the static and controller watts are unchanged and its activate-bound rate is unchanged at 166 MH/s, 1.36 TB/s of the board's 1.79). The honest 5090 pays the same DRAM joules (+20 W at 17.9 G reads per second) plus one more sector through its L2 and crossbar (about +6 to +10 W, approximate), on a rate that MEASURED 2.7 percent faster at w16 (139.8 against 136.1 MH/s, 5 October); the RX 9070 XT (a 64-byte line per read, measured) and the M5 Max (64 B at the 4 B rate, measured) pay nothing. Card energy per hash at the knee: 1.67 to about 1.84 microjoules (+10 percent, modelled; the w16 job carried no power sampling, so this is the one row owed). The edge at zero shadow: 3.6x to 3.0x on GDDR7, 5.2x to 4.6x on one HBM3 stack; with the class v4 shadow at `k = 0.3`: 3.5x to 3.05x; at `k = 0.5`: 2.9x to 2.6x; at `k = 1`: 2.1x to 2.0x. The verifier pays nothing (0.610 against 0.604 ms per 32 hashes, measured). It is the only lever on the table whose forced work is `k = 1`, it stacks with the operating point and with the shadow, and it is against the SRAM full store alone that it does nothing (that chip's macro read is already 64 bytes). **KEEP, rank 1, into class v6 now** as the read-width floor (W = 16 words, the whole item), replacing the layer-1 band {1, 4} words whose warrant was that wider reads are harmless and worthless. + +**Everything else on the brief is killed, with the number.** The tensor core: the record's `k` 0.03 to 0.3 is not defensible at its floor (the 0.04 pJ figure is INT4 at 0.46 V on a test chip with no memory system); shipping merchant silicon at nominal voltage reads 0.30 to 0.56 pJ per INT8 MAC (Meta MTIA v2 0.51, Qualcomm AI 100 Ultra 0.34, AMD MI355X 0.56, NVIDIA B200 0.44, Apple M4 ANE 0.30 measured; all claimed unless marked), and a Hopper tensor core under a pure MMA loop measured 0.34, so the honest band against the tile the class runs (1.5 pJ per MAC at the knee, measured) is `k` 0.2 to 0.4 and against the dense wide tile (0.83 at the knee) 0.4 to 0.7: equal to or under the ALU shadow's 0.3 to 0.8, never above it, and the Apple cost (-35 percent of rate at 1,024 tiles per hash, measured) kills it as content regardless. The RT core: traversal results are implementation-defined on every API (Vulkan "no ordering guarantee", DXR "no defined order", NVIDIA "could change depending on which driver, which GPU"), the acceleration structure is opaque, and a dedicated unit reads 4 to 33 nJ per ray against 290 to 750 nJ measured board-level on an RTX 2080 (`k` 0.01 to 0.2): dead on bit-exactness and on `k`. Memory-level parallelism per dollar: no asymmetry per channel, but a correction to the chip model: no 28 nm GDDR7 PHY exists (the shipped ones are N3 and FinFET; the oldest GDDR6 PHYs are 12 nm), so the record's cheapest chip (a USD 5 M project, a USD 17 M break-even cap) is not buildable; the floor is a 12 nm GDDR6 part at about USD 20 M to 30 M (cap USD 70 M to 100 M) or an N5-class GDDR7 part at USD 50 M to 75 M (cap 170 M to 250 M), modelled. Proof of latency: the chain checks values, never time; the block is 1 s and a DRAM round trip 0.1 to 0.4 microseconds; the farm's node is on its own LAN; a sequential per-block prefix favours the lower-latency side, which is the chip (about 4x). The time dimension on the address map: a bit permutation is firmware (a few hundred gates), no prefetch exists for a dependent chain, and the literature has no such scheme. The literature since 2023 holds nothing outside the identity; what it adds is a higher Ethash precedent (iPollo V2H, 0.14 J per MH, about 13x a 5090, vendor claim), the measured PAM3 I/O energies (SK hynix, ISSCC 2024) and the PHY-node fact above. Of this lane's own ideas, four more die on the identity (video decode, row-straddling, scratch in DRAM, independent chains per hash) and one is held for v7 (two items per read, W = 32: the chip goes bandwidth-bound at 110 MH/s and 1.02 microjoules, the edge at zero shadow 2.4x and 2.5x at `k = 0.3` with the shadow, at a cost of about 20 percent of the 5090's rate, modelled). + +### The KEEP list, ranked + +| Rank | Item | Chip edge at the 5090's knee, GDDR7 (zero shadow / with class v4 at k 0.3 / 0.5 / 1) | Cost to the honest tiers | Verifier | Where | Label | +|---|---|---|---|---|---|---| +| 1 | **The second sector: W = 16 words, the whole 64-byte item folded (class `w16`)** | 3.0x / 3.05x / 2.6x / 2.0x (from 3.6 / 3.5 / 2.9 / 2.1) | 5090 about +10 percent energy per hash (+26 to 30 W at the knee), rate +2.7 percent measured; 9070 XT and M5 Max 0 (measured rate; the line is 64 B already); 5070 Ti and 4070 as the 5090 (32 B sectors, approximate) | +1 percent (0.610 against 0.604 ms, measured) | class v6 now: layer 1's read-width band becomes the floor W = 16 | card measured (rate) and modelled (watts); chip modelled | +| 2 | **The chip model's project floor corrected: no 28 nm GDDR7 PHY** | unchanged | none | none | chip-model-v3 5.4 and the record's 16.2: the cheapest row's project USD 20 M to 75 M, cap USD 70 M to 250 M | claimed (vendor IP pages), modelled (the mission lane's cap) | +| 3 | **Two items per read, W = 32 (128 B, one L2 line)** | 2.4x / 2.5x / 2.3x / 1.9x; the chip bandwidth-bound at 110 MH/s | 5090 about -20 percent rate and +23 percent watts (modelled from the 1.79 TB/s peak); M5 Max 0 to -14 percent (approximate); 9070 XT two lines per read (unmeasured) | about +1 percent (the w64 row 0.630 ms) | v7, after the 5090 and M5 Max rate and watts rows on a `w32` export | modelled | +| 4 | **The tensor `k` band corrected** (0.2 to 0.4 on the class's tile, 0.4 to 0.7 on the dense tile; not 0.03 to 0.3) | no change to any served row (the tile stays a worse-or-equal lever) | none | none | a wording correction to the record's 15.1a and 20.4 and chip-model 5.11 | claimed and measured | +| 5 | **The Ethash precedent's top row** (iPollo V2H, 3.4 GH/s at 475 W, about 13x a 5090 per joule) | the precedent band 2.1x to 13x, not 2.1x to 6.8x | none | none | `asic-resistance-history.md` and chip-model 5.1 | claimed | + +### The five-sentence reading for the 20:00 close + +The identity leaves one piece of work a chip pays at the card's own price, the DRAM's data movement, and the record forced only one sector of it per read; reading the whole 64-byte item (W = 16, the measured `w16` class) forces the second sector at `k = 1`, costs the 5090 about 10 percent of energy per hash and no rate, costs AMD and Apple nothing, costs the verifier nothing, and takes 10 to 15 percent off every chip column (3.6x to 3.0x at zero shadow, 2.9x to 2.6x with the shadow at `k = 0.5`), so it goes into class v6 now as the read-width floor. The tensor core, the RT core, proof of latency and the drawn address map are dead on the identity or on bit-exactness, with the tensor `k` band corrected upward to 0.2 to 0.7 (still never above the ALU shadow's) and the RT core's traversal implementation-defined on every vendor. The chip model's cheapest chip does not exist as priced: a GDDR7 PHY needs a FinFET node (N3 shipped; 12 nm is the oldest GDDR6 PHY), so the floor project is USD 20 M to 75 M and the floor break-even cap USD 70 M to 250 M instead of 17 M, which goes into the model now. Two items per read (W = 32) walks both sides into the bandwidth regime where the DRAM's joules are most of both bills (the edge 2.4x at zero shadow, 2.5x at `k = 0.3`) at a cost of about a fifth of the 5090's rate, so it is a v7 candidate behind two measured rows, not a v6 change. Nothing in the literature since 2023 offers a lever outside the identity, and the one measurement this file owes is the 5090's watts at w16 at the knee and unlocked: one PC job on a pack that already exists. + +## 1. The frame + +| Term | Value at the 5090's 1,300 MHz knee | Label | Source | +|---|---|---|---| +| `E_card`, class v3 | 1.67 microjoules (127.3 MH/s at 213.0 W; 134.6 at 223.3 on the efficiency pass) | measured | the record 20.3; `docs/plans/counter-asic-3-status.md` | +| `E_mem`, the `f = 1` GDDR7 chip | 0.466 microjoules: 166 MH/s at 77.6 W (42.6 W of reads at 2.0 nJ, 20 W static, 15 W controller) | modelled | chip-model-v3 5.3 and 5.4 | +| `E_mem`, one HBM3 stack | 0.321 microjoules: 83.6 MH/s at 26.8 W (12.8 W of reads at 1.2 nJ, 14 W static and controller) | modelled | the same | +| The read's energy, split | GDDR7: 909 pJ activation plus 4.5 pJ per bit x 256 bits = 1,150 pJ of movement and I/O (44 / 56 percent); HBM3: 909 plus 3.48 x 256 = 890 scaled by 4.12 / 6.25 | modelled on O'Connor et al. 2017 Tables 2 and 3 and Micron's and Samsung's pJ per bit (claimed) | chip-model-v3 5.3 | +| `F`, the class v4 premium | 0.652 microjoules (82.8 W over 126.9 M x 102,100 ops: 6.4 pJ per counted op) | measured | the record 20.3 | +| `k`, the ALU shadow | 0.3 to 0.8 (the 5090's 6.2 to 11.3 pJ per op measured against a 5 nm SIMD array's 2 to 5, approximate) | GPU measured, chip approximate | the record 15.1a | +| The dataset item | 16 words, 64 bytes; `dataset[w] = item(w >> 4)[w AND 15]`; the lottery hash reads one word per load | design | spec 1.5 and the item table; `docs/plans/read-width.md` | +| The 5090's dependent-read rate | 17.5 to 18.2 G reads per second at 4 B and 16 B, the best over lanes in flight; 9.1 to 15.7 at 64 B in the probe; the w16 hash at 17.9 G (139.8 MH/s x 128) | measured, 5 October | `docs/bench-log.md`, the read-width entry | +| The card's whole-card marginal per dependent DRAM read | 10.9 nJ unlocked, 8.7 nJ at the lock | measured, 8 October | the record 15.1a | + +The four doors a candidate can open (the invention lane's section 1, kept): raise `k`; raise the chip's capex or project; shorten the chip's useful life; lower the honest card's cost at the same `F`. This file adds the fifth door the identity itself names: **force work whose `k` is 1 by construction, which is only the DRAM's own.** + +## 2. The candidates on the brief + +### 2.1 Work the GPU already does at near-ASIC efficiency + +**The tensor core, re-read.** The question: is "a merchant N2 project beats Blackwell's tensor core per joule by 3x to 30x" defensible, or is the tensor shadow at full rate closer to `k` 0.7 to 1? + +| Chip | Node | Year | Dense INT8 TOPS (FP8 where marked) | W | pJ per MAC (2 / (TOPS per W)) | Label | URL (read 8 October 2026) | +|---|---|---|---|---|---|---|---| +| RTX 5090 | 4N | 2025 | 838 | 575 | 1.37 at spec; **1.36 measured** on the dense s8 m16n8k32 tile at 80 percent of peak unlocked, 0.83 at the 1,300 lock; the dependent u8 m8n8k16 tile the class runs 2.9 unlocked, 1.5 at the lock (the packs job), 4.1 / 2.2 (the microbench) | claimed spec; measured rows the record's 15.1a and 20.3 | https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/ | +| RTX 4090 | 4N | 2022 | 661 | 450 | 1.36 | claimed | NVIDIA Ada whitepaper | +| H100 SXM | 4N | 2022 | 1,979 | 700 | 0.71 | claimed | https://www.nvidia.com/en-us/data-center/h100/ | +| H800, mma and wgmma INT8 dense under a pure MMA loop | 4N | 2025 | | sub-TDP loop | **0.34 / 0.54** | measured | https://arxiv.org/html/2501.12084v2 Tables 10 and 11 | +| B200 (FP8 / INT8) | N4P | 2024 | 4,500 | 1,000 | 0.44 | claimed | https://lenovopress.lenovo.com/LP2226 | +| AMD MI300X | N5 | 2023 | 2,615 | 750 | 0.57 | claimed | AMD data sheet | +| AMD MI355X | N3P | 2025 | 5,033 | 1,400 | 0.56 | claimed | https://glennklockwood.com/garden/processors/mi355x | +| RX 9070 XT (RDNA 4 WMMA) | N4P | 2025 | 389 | 304 | 1.56 | claimed | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070xt.html | +| Qualcomm Cloud AI 100 Ultra | 7 nm | 2023 | 870 | 150 | 0.34 | claimed | Qualcomm product brief | +| Meta MTIA v2 | 5 nm | 2024 | 354 | 90 | 0.51 | claimed | https://www.eenewseurope.com/en/meta-launches-second-generation-custom-ai-chip | +| Intel Gaudi 3 (FP8) | 5 nm | 2024 | 1,835 | 900 | 0.98 | claimed | https://www.anandtech.com/show/21342 | +| Google TPU Ironwood (FP8) | n/d | 2025 | 4,614 | about 1,100 inferred | about 0.48 | inferred from the Hot Chips 2025 slides | https://hc2025.hotchips.org/ | +| AWS Trainium2 (FP8) | 5 nm | 2024 | 1,299 | about 500 (third party) | 0.77 | inferred | https://arxiv.org/pdf/2608.28048 | +| Tenstorrent Blackhole p150 (FP8) | N6 | 2025 | 774 | 300 | 0.78 | claimed | https://docs.tenstorrent.com/aibs/blackhole/specifications.html | +| Apple M4 ANE (FP16 rate, INT8 runs at it) | N3E | 2024 | 19 | 2.8 | **0.30** | measured | https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615 | +| Hailo-8 | 16 nm | 2019 | 26 | 2.5 | 0.19 at spec, 0.71 measured on ResNet-50 | claimed / measured | https://www.cnx-software.com/2020/10/07/ | +| Untether speedAI 240 (FP8, at-memory) | 7 nm | 2022 | 2,000 | 66 | 0.067 | claimed, never shipped at volume | https://www.design-reuse.com/news/52539/ | +| Axelera Metis (digital in-memory) | 12 nm | 2023 | 214 | about 14 | 0.13 | claimed | https://www.hackster.io/news/ | +| NVIDIA VSQ test chip (INT4) | 5 nm | 2023 | | | 0.021 at 0.46 V on a benchmarking layer; **0.052 at the nominal 0.67 V on full networks**; no INT8 figure | measured, a test chip without HBM or a NoC | https://research.nvidia.com/publication/2023-01_956-topsw-deep-learning-inference-accelerator-vector-scaled-4-bit-quantization | + +The reading. Every shipping systolic chip with a real memory system and fabric lands at 0.3 to 0.6 pJ per INT8 MAC at nominal voltage; the figures under 0.15 are in-memory-compute vendor claims and an INT4 test chip at low voltage. Dally's own energy pie for that test chip (Hot Chips 2023, slide 24) has the datapath at 47 percent and the buffers, collector and movement at 53 percent, so a whole chip is about 2x its datapath before any NoC or HBM. The record's `k` floor of 0.03 took the 0.46 V INT4 figure against the 5090's 1.5 to 4 pJ; at nominal and INT8 the chip side is 0.3 to 0.6. The honest band, labelled: + +| Against | GPU pJ per MAC (measured) | Chip pJ per MAC at nominal (claimed, the cluster above) | `k` | Reading | +|---|---|---|---|---| +| The tile the class runs (dependent u8 m8n8k16, the `mm1430` pack) at the knee | 1.5 | 0.3 to 0.6 | **0.2 to 0.4** | the forcing lever's `k`; equal to or under the ALU shadow's 0.3 to 0.8 | +| The same unlocked | 2.9 | 0.3 to 0.6 | 0.1 to 0.2 | | +| The dense wide tile (s8 m16n8k32) at the knee | 0.83 | 0.3 to 0.6 | **0.4 to 0.7** | a tensor shadow built of wide tiles would sit at the ALU band's centre; the H800's 0.34 under a pure loop says a GPU tensor core at its best is inside the merchant cluster | +| The dense wide tile unlocked | 1.36 | 0.3 to 0.6 | 0.25 to 0.45 | | + +So "3x to 30x" is not defensible: at nominal voltage merchant silicon beats the 5090's dense tile by 2.3x to 4.5x and its dependent tile by 2.5x to 10x, and the 30x needs an INT4 layer at 0.46 V. And "`k` 0.7 to 1 at full rate" is not reached either: the top of the dense band is 0.7. The identity check: the tile's `k` is at best the ALU shadow's and the joules it forces are the same by construction (the record 20.3: 71.5 W against 82.8 at the knee for the same tile count); the chip's downside bet (its `k` floor) is 0.2 against the ALU's 0.3, so the tile remains the worse-or-equal lever, not the 4x to 30x worse lever the record wrote. The cost to the honest tiers is what kills it as content: the M5 Max at -35 percent of rate at 1,024 tiles per hash and -78 percent at 4,096 (measured, the record 20.2a), with no integer matrix path in Metal. Bit-exact on NVIDIA and the CPU (measured); the AMD WMMA layout unverified. **KILL as v6 content; KEEP the band correction (rank 4).** FP8, FP16 and BF16 tiles: the accumulation order inside a tile is unspecified by PTX, so not bit-exact across vendors or generations: KILL. + +**The texture unit's filtered gather.** Measured 0.19 nJ per filtered fetch unlocked and 0.10 at the lock (the microbench `tex_linear_f32_256k`); a 9-bit fixed-point interpolator on NVIDIA (CUDA Programming Guide, claimed), vendor-specific fraction widths elsewhere: not bit-exact across the three vendors; a chip's interpolator is one 9-bit multiply-add, `k` far under 1. KILL. The point-sampled fetch is the L2 chase under another name (the same checksum, measured): the L2 lever, dead at `k` 0.1 to 0.3. + +**The rasteriser and ROPs.** Unreachable from CUDA, OpenCL and Metal compute; fill conventions, sample positions and depth precision differ by vendor by design. KILL. + +**The hardware video codec.** A conformant decoder of a standard (H.264, HEVC, AV1) is normatively bit-exact, which is the one fixed-function block on a card whose output a CPU verifier could reproduce, and a chip would license the same decoder block at the same node (`k` about 0.5 to 1, the same circuit, approximate). It dies on throughput and on the verifier: a card carries a fixed count of decode engines (two to four on current cards, approximate) against any count on a chip, so the honest card's hash rate would be bound by its decoder, not its memory, at a ratio that differs per vendor (the 5090 against the M5 Max's media engine); and a software decode on the verifier runs at milliseconds per frame on one core (approximate), against a 10 ms gate that holds the whole warp. KILL. (This lane's own candidate, recorded here because it is the only fixed-function block that passes the bit-exactness test.) + +### 2.2 The RT core + +| Unit (node) | Throughput | Power or area | Energy per ray | Label | URL (read 8 October 2026) | +|---|---|---|---|---|---| +| RTX 5090, 170 fourth-generation RT cores | 317.5 RT TFLOPS; "double the throughput for ray-triangle intersection" over Ada; no per-ray or per-clock rate published, no RT-core area or power share | 575 W | not published | claimed | https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf | +| RTX 2080 running DXR (Mach-RT, Table 3) | 288 to 745 Mrays per second over six scenes | 215 W TDP | **290 to 750 nJ per ray, board-level** | measured rays, TDP assumed | https://hwrt.cs.utah.edu/papers/mach-rt_TVCG.pdf | +| Turing TU102 | 10 Giga rays per second claimed | about 260 W | about 26 nJ per ray at the claimed peak | claimed | Hot Chips 31 slides | +| AMD RDNA 4 | 8 ray-box and 2 ray-triangle tests per clock per CU, BVH8, OBBs | no RT power figure | none | claimed | https://hc2025.hotchips.org/assets/program/conference/day1/8_amd_pomianowski_final.pdf | +| SiliconArts RayCore, 6 RTUs, 28 nm | 239 Mrays per second | 1 W, 18 mm^2 | about 4.2 nJ per ray | claimed, FPGA-derived | https://www.embedded.com/?p=4433770 | +| Imagination PowerVR GR6500, 28 nm | 300 Mrays per second | 5 to 10 W for the whole GPU | 17 to 33 nJ per ray | claimed | https://www.anandtech.com/show/7870 | +| RayFlex RTL, 15 nm PDK, 1 GHz | one ray-box or ray-triangle test per cycle | 60 to 85 mW per datapath | 60 to 85 pJ per intersection test | synthesised | https://arxiv.org/pdf/2409.06000 | +| TRaX study, 65 nm | | memory 60 to 95 percent of ray energy, DRAM alone up to 80, compute 1 to 2 | 0.5 to 5 microjoules per ray | simulated | https://hwrt.cs.utah.edu/papers/rt_performance_CGI18.pdf | + +Per node visit (Ylitie et al., NVIDIA, HPG 2017, measured): about 15 ray-node and 7 to 9 ray-triangle tests per ray on a compressed 8-wide BVH, 1.5 to 3.2 KB fetched per ray, "memory traffic is the main limiting factor with incoherent rays"; Chou et al. (MICRO 2023) call it "a latency-bound" pointer chase. The identity check: a traversal step over a dataset-sized tree is one dependent DRAM read (the identity's `E_mem`, which both sides pay) plus a box test the RT core does in hardware; the box test is the only forcing, and the chip's cost to match it is 60 to 85 pJ at 15 nm (synthesised) against a GPU share nobody has measured below 290 nJ per ray board-level, so `k` reads 0.01 to 0.2 on the public figures: the chip builds a cheaper RT unit, not a dearer one. + +Bit-exactness, which is the first question and the one that decides it: the Vulkan specification gives "no ordering guarantee between operations performed on different intersection candidates", makes watertightness a "should", and performs the hit test "in an implementation specific manner" (https://docs.vulkan.org/spec/latest/chapters/raytraversal.html); DXR: "There is no defined order of execution of any hit shaders for the intersections along a ray path" (https://microsoft.github.io/DirectX-Specs/d3d/Raytracing.html); NVIDIA's own answer (forum, 14 October 2024): which edge or vertex wins is "implementation defined and not guaranteed" and "could theoretically change depending on which API you use, which driver, which GPU"; Blender's Cycles records that matching OptiX or MetalRT bit for bit was impossible. The acceleration structure is opaque in DXR, Vulkan and Metal, and Vulkan's serialised form carries a driver UUID that the next driver may refuse. A CPU verifier therefore cannot reproduce a hardware traversal, and a design that defines its own traversal in software and uses the RT core as an untrusted accelerator forces nothing (the honest card then does the box tests in software or re-checks them). **KILL**, on bit-exactness first and on `k` second. Price of the RT unit a chip would build: 18 mm^2 at 28 nm for six units (claimed), about USD 5 of N5 (approximate). + +### 2.3 Memory-level parallelism as the resource: channels per dollar, the controller and PHY + +The claim to check: a card has more channels per dollar than a chip built from the same devices plus a controller, because the chip's controller and PHY are a cost the card amortises over gaming. + +| Item | Finding | Number | Label | URL (read 8 October 2026) | +|---|---|---|---|---| +| Channels per device | GDDR7: four 8-bit channels and 64 banks per 2 GB device; the 5090's 16 devices are 64 channels and 1,024 banks; a chip built from the same devices has the same count; the next parts (3 GB, 4 GB, 6 GB) carry the same four channels, so channels per GB FALL for the chip and the card alike (Micron ends the 2 GB part) | 64 channels per 32 GB today; 64 per 48 GB on 3 GB parts | claimed (TrendForce, Rambus) | chip-model-v3 5.1; https://www.trendforce.com/news/2026/09/24/ | +| Does a chip buy more channels per dollar | No: the channel count is bought with capacity on both sides (16 devices whatever the dataset), and the chip's only path to more activates per second is more devices, which the card cannot add and the chip pays for at the same price per device | 0 asymmetry per channel | modelled | | +| The PHY's node | No 28 nm GDDR7 PHY exists: Cadence's GDDR7 PHY is silicon-proven on TSMC N3 at 32 Gbps (September 2024); Innosilicon's is "advanced FinFET"; the oldest GDDR6 PHYs are 12 nm (Cadence 12FFC and Six Semiconductor, 16 Gbps); nothing at 28 nm at any GDDR6 or GDDR7 rate | the chip's cheapest controller is a 12 nm GDDR6 part or an N5-class GDDR7 part | claimed (vendor pages) | https://www.businesswire.com/news/home/20240925780123/en ; https://innosilicon.com/html/ip-solution/gddr7.html ; https://www.businesswire.com/news/home/20221115005674/en | +| The PHY's energy | GDDR7 PAM3 I/O measured on silicon: TX 1.103 pJ per bit, RX 0.764 pJ per bit (SK hynix, ISSCC 2024 paper 13.1); at the hash's 17.9 G reads x 256 bits = 4.6 Tb/s the controller-side RX is 3.5 W, inside the model's 15 W controller allowance with its clocking and termination (PAM3 termination is a large share: arXiv 2410.12990, simulated) | `E_mem` moves by under 0.06 microjoules | measured I/O; the allowance approximate | https://www.researchgate.net/publication/378947349 ; https://arxiv.org/abs/2410.12990 | +| The GB202 die | 761.56 mm^2 with 16 x 32-bit controllers and PHYs along the edges; no public mm^2 for the PHY blocks | the card's PHY is a real area it amortises; the chip's is the same area at the same node | measured die, no annotation | https://kurnal-insights.com/en/dieshot/nvidia-gb2025090/ | + +The identity check: the chip and the card pay the same per-bit I/O and the same controller physics, so nothing moves `E_mem` beyond the allowance already in the model. What moves is the chip's PROJECT: the record's cheapest row (chip-model-v3 5.6 "a 28 nm-class project at USD 5 M to 30 M", the record's 16.2 "28 nm controller, USD 5 M, break-even cap USD 17 M") prices a PHY that is not made. The corrected floor, on the mission lane's model (cap = C_proj / 0.30 in years 1 to 2): + +| Cheapest buildable chip | Memory | Project | Break-even cap, years 1 to 2 | Label | +|---|---|---|---|---| +| A 12 nm GDDR6 controller (16 Gbps PHY IP exists): 32 banks per device, so 32 devices of GDDR6 for the same 1,024 banks and 21.3 G activates per second; about USD 10 per 2 GB device (approximate); a 1,024-bit board | USD 320 | USD 20 M to 30 M (a 12 nm project with licensed PHY IP; the history's 2.5 figures scaled, approximate) | **USD 70 M to 100 M** | modelled, approximate | +| An N5-class GDDR7 controller (the 5090's own 16 devices) | USD 320 | USD 50 M to 75 M (the history's 7 nm-class startup figure, claimed) | **USD 170 M to 250 M** | modelled | +| The same plus the N5 shadow core (class v4 as shipped; the record's 16.2) | | USD 50 M to 75 M (the core joins the controller's die, which is N5 anyway) | USD 170 M to 250 M, not 100 M | modelled | + +**KEEP as a correction to the chip model (rank 2), not a lever.** Consequence: the record's break-even row for the chip that matters (the N5 core on a 28 nm controller, USD 30 M, cap 100 M) understates the floor by about 2x, because the controller cannot sit on 28 nm; the shadow per load's "doubling" (USD 60 M, cap 200 M) therefore adds less than it seemed, since the single-N5-die form is already the only buildable form. The 85 W node per farm and the leaves' bandwidth are unchanged. + +### 2.4 Proof of latency + +A per-block random challenge whose answer window is shorter than one DRAM round trip plus network latency, so a stored copy in a remote farm cannot answer in time but a local card can. + +| Check | Finding | Number | Label | +|---|---|---|---| +| The identity | Not an energy lever: it changes who may answer, not what a hash costs. If it bound anything it would bind the chip's node placement | 0 on `E_mem`, `F` and `k` | arithmetic | +| What the chain can enforce | Values, never time (the record's section 6, new-pow 3.2): a block header is the challenge, blocks arrive over the WAN at tens to hundreds of milliseconds, the interval is 1 s and a timestamp is the miner's own; no node can check that an answer arrived within microseconds of a challenge it saw itself at a different time | the shortest enforceable window is a block interval, 1 s, against a DRAM round trip of 0.4 microseconds on the 5090 (415 ns queued, measured) and about 0.1 on a chip (55 ns of controller plus tRCD and tCL, modelled): 2.5 million round trips per block | measured card latency; the chip's modelled | +| The threat model | The stored-dataset chip's node is on its own LAN (class-v5 2a.2: one node serves a farm, the leaves at 16.5 KB/s); there is no "remote chip farm" to lock out; a farm is as local as a card | 0 | design | +| The one form with teeth: a sequential per-block prefix (every miner must run a chain of dependent reads from the block hash before any nonce counts) | Favours the side with the lower unloaded latency, which is the chip: a 100 ms prefix is 250,000 reads at 400 ns on the 5090 and 100 ms of its time; the chip at about 100 ns per read finishes in 25 ms and hashes for the rest | the chip's share of the block rises about 8 percent at a 100 ms prefix; the card loses 10 percent of its time | modelled | +| Fresh join | Unchanged: a joiner needs the previous block, which it needs anyway | 0 | design | +| The verifier | Nothing to verify: time is not in the proof | 0 | design | +| The literature | No proof of work with a sub-round-trip window was found 2023 to 2026; the nearest are watchtower-ping location proofs (BFT-PoLoc, arXiv 2403.13230: about 2 percent Byzantine challengers tolerated, 12 percent misclassification, measured on the Internet; Witness Chain) and PoSME (already in the record); Proximity Signatures (eprint 2026/694) refused the fetch and is unread | null | claimed | + +**KILL.** The chain cannot measure time at the scale a DRAM round trip lives on, the farm is local, and the sequential form helps the chip. + +### 2.5 The time dimension: the address map drawn often enough to break a chip's prefetch and banking plan + +| Check | Finding | Number | Label | +|---|---|---|---| +| The identity | A mapping change is firmware for the stored-dataset chip; no `F`, no `k` | 0 | arithmetic | +| Prefetch | A dependent chain has nothing to prefetch on either side: the next address is the previous read's data | 0 | design | +| Banking plan | The chip's controller maps address bits to channel, bank and row through a table; a drawn permutation of the index bits costs a few hundred gates (the invention lane's 2.7, the spec's 1.13.1 draws of `M`, `R` and `pos` already); the card pays whatever the map does to its own bank parallelism, measured as a 0.8 to 3.2 percent spread across six eras | 0 to the chip; 0.8 to 3.2 percent to the cards | measured (Counter ASIC 2.0 layers 4 and 8) | +| A fixed-function chip | Dead at the first draw, which is layer 1 already; a programmable chip pays USD 25 to 40 of N5 core, already in the record's 16.1 | nothing new | modelled | +| The cadence | Drawing per epoch instead of per era changes nothing above: a table reload is microseconds | 0 | design | +| The literature | No scheme since 2023 rotates an address map to defeat banking or prefetch; Ethash's 30,000-block reseed and RandomX's dataset rebuild are the only moving layouts, both firmware to every chip that shipped | null | claimed | + +**KILL.** Covered by layer 1 and the invention lane's row; nothing a programmable chip pays beyond the core it already carries. + +### 2.6 The literature since 2023, what the record's section 8 missed + +| Item | What it found | Number | Label | URL (read 8 October 2026) | +|---|---|---|---|---| +| Choe, Moreshet, Bahar, Herlihy, "Attacking memory-hard scrypt with near-data processing", MEMSYS 2019 | the one paper that runs a memory-hard function on a PIM model: one in-order core per 128 MB vault | 1.5x over the host, no energy figure | simulated | https://par.nsf.gov/servlets/purl/10156963 | +| Blocki and Smearsoll, TCC 2025; Blocki et al., CRYPTO 2026 (arXiv 2508.06795) | pebbling bounds on MTP-style and data-dependent memory-hard functions | complexity only, no joules | theory | https://arxiv.org/abs/2508.06795 | +| SemiAnalysis, "The Memory Wall", 3 September 2024 | off-chip movement about 20x the cell read; an HBM access about 95 percent interface, 5 percent cell, for streaming | 2 pJ per bit interface against 0.18 pJ per bit activation amortised over a streamed row; for THIS hash's random 32-byte read the activation is 909 pJ against 1,150 of movement (44 / 56 percent), so a shorter link (lane B's base-die and bonded rows) cuts the larger half, which the model already carries | claimed; the split modelled | https://newsletter.semianalysis.com/p/the-memory-wall | +| SK hynix GDDR7 at 35.4 Gb/s per pin, ISSCC 2024 paper 13.1 | PAM3 TX 1.103 pJ per bit, RX 0.764 pJ per bit | the I/O share of the 4.5 pJ per bit device figure | measured silicon | https://www.researchgate.net/publication/378947349 | +| ETH SAFARI real-chip DRAM studies 2024 to 2026 | all on DDR4; no GDDR6, GDDR7 or HBM3 activate-energy measurement exists in public | the record's 909 pJ stands as the only figure | measured, DDR4 | https://arxiv.org/pdf/2405.06081 | +| iPollo V2H, V2, V2X (Ethash, November 2024) | 3.4 GH/s at 475 W; 10 GH/s at 1,500 W; 1.2 GH/s at 165 W | 0.14 J per MH: about 13x an RTX 5090 at 1.8 J per MH on Ethash; the precedent's top row, above the record's X16-P at 6.8x; all of it the memory system | vendor spec, claimed | https://www.asicminervalue.com/en/miners/ipollo/v2h ; https://www.whattomine.com/coins/382-egaz-etchash/gpus/92-nvidia-geforce-rtx-5090 | +| Jasminer X16-Q (Ethash, ETC) | 1,950 MH/s at 620 to 630 W | 0.32 J per MH, about 5.7x a 5090 | vendor spec | https://pool.kryptex.com/device/asic/Jasminer/x16-q | +| Innosilicon A11 Pro (2021); iPollo V1 (June 2022) | 1.5 GH/s at 2,350 W; 3.6 GH/s at 3,100 W | 1.3x and 2.3x a 4090 | vendor spec | https://www.asicminervalue.com/miners/innosilicon/a11-pro-eth-1500mh | +| Pinecone R1X (RandomX, January 2026) | 1.2 MH/s at 2,055 W, listed against the X9's 1 MH/s at 2,472 W | 1.71 J per kH; unverified, no unit seen | vendor spec, unverified | https://github.com/monero-project/monero/issues/10270 | +| Qubic, 2026 | 23 to 34 percent of Monero's hash in withholding windows, never a sustained majority; pivot to Dogecoin ASICs 1 April 2026, Monero mining phased out | 27,000 XMR blocks, about USD 3.5 M | measured events | https://arxiv.org/abs/2512.01437v2 ; https://decrypt.co/363018 | +| Karlsen, Iron Fish, Pyrin after the FishHash forks | no ASIC since April 2024; a ColEngine P2 FPGA did 32 GH/s Karlsen at 1,484 W before the fork | FPGA only | vendor | https://pool.kryptex.com/articles/ironfish-hardfork-en | +| Conflux, Flux, Nexa, Ergo, Ravencoin, Beam, Tari, Zephyr | no dedicated chip 2024 to 2026 (Tari and Zephyr mined by the RandomX X5 class) | null | vendor | | +| GDDR7 PHY nodes and GDDR6's oldest | Cadence N3 (2024); Innosilicon FinFET; GDDR6 at 12 nm (Cadence 12FFC, Six Semiconductor); nothing at 28 nm | section 2.3's correction | claimed | as 2.3 | +| Proof of latency; era-rotating address maps | no scheme found, 2023 to 2026 | null | | | + +Reading: nothing published since 2023 offers a lever outside the identity; every energy figure is a streaming pJ per bit, and no one has measured a per-random-read joule on GDDR6, GDDR7 or HBM3 at the controller, so the record's 2.0 and 1.2 nJ stand as the only model and are owed a measurement nobody can rent (the AWS F2 hour of chip-model 5.3 is still the one HBM2 instrument). The new rows change two precedents: the Ethash band's top is about 13x (claimed) not 6.8x, and the Monero chip story has a shipped-but-unverified successor to the withdrawn X9. + +## 3. This lane's own ideas + +### 3.1 The second sector: W = 16 words, the whole item (KEEP, rank 1) + +The construction is the read-width branch's `w16` class (`docs/plans/read-width.md` section 1.5 and the fold of its section 1): a load reads the 64-byte-aligned item and folds its 16 words into the destination register with a rotate-multiply between words (`verify::fold_words`, mirrored in Metal, CUDA and OpenCL), so no function of the line alone replaces the dataset and the dependent chain is unchanged. Built 5 October; packs in `proto-cuda/packs-readwidth/`; three vector units and the cache and dataset checks passed on Metal, Apple OpenCL, the RTX 5090 (NVRTC) and the RX 9070 XT; the 2^24 fingerprint `e7c890445b47af60` equal on all four; the acceptance rule's rejection 0 to 14 of 60 as version 2; the CPU verifier 0.610 ms per 32 hashes against 0.604 (one item per lane, derived whole either way). What was never priced is the chip side. + +The chip side, on the model's own inputs (chip-model-v3 5.3: a random 32-byte read is 909 pJ of activation plus 256 bits at 4.5 pJ per bit of movement and I/O; HBM3 the same shape at O'Connor's 3.48 pJ per bit scaled by 4.12 / 6.25): + +| Memory | Read energy at W = 1 (one 32 B sector) | At W = 16 (two sectors in the open row: one activate, two column accesses) | `E_mem` per hash, W = 1 to W = 16 | Rate | Label | +|---|---|---|---|---|---| +| GDDR7, 16 devices | 2.0 nJ | 3.2 nJ (+56 percent; 2.9 at O'Connor's 3.5 pJ per bit, +45) | 0.466 to **0.62** (+33 percent; 0.58 at the lower movement figure, +25); the 35 W of static and controller unchanged | activate-bound at 166 MH/s, unchanged; 1.36 TB/s of the board's 1.79 | modelled | +| HBM3, one stack | 1.2 nJ | 1.8 nJ (+50 percent) | 0.321 to **0.40** (+24 percent) | 83.6 MH/s, unchanged; 0.68 TB/s of 0.82 | modelled | +| HBM3, eight stacks | 1.2 | 1.8 | 0.262 to 0.33 (+26 percent) | unchanged | modelled | +| The SRAM full store (lane B) | 1.0 nJ for a 64-byte line | 1.0 (the macro reads the line) | 0.14, unchanged | unchanged | modelled | +| The custom HBM4E base die (lane B) | 0.9 to 1.0 | about 1.5 (the activation 909 pJ at 0.75 V scaled, the movement doubled) | 0.18 to about 0.27 (+50 percent) | unchanged | modelled, approximate | + +The card side: + +| Card | Rate at W = 16 | DRAM joules | The card's own fabric | Energy per hash | Label | +|---|---|---|---|---|---| +| RTX 5090 at the 1,300 knee | +2.7 percent (139.8 against 136.1 MH/s unlocked, 5 October; the knee row unmeasured) | the same second sector: 1.15 nJ x 17.9 G reads per second = +20 W | one more 32-byte sector through L2 and the crossbar to the SM: about 0.3 to 0.5 nJ per sector (approximate; the measured L2 hit of 1.4 nJ per dependent read at the lock includes the wait), +6 to +10 W | 213 W to about 240 to 243 W at 130.7 MH/s: **1.67 to about 1.84 microjoules (+10 percent)**; the w16 job carried no power sampling (read-width.md 4.1), so this is modelled and is the one row owed | rate measured; watts modelled | +| RTX 5090 unlocked | the same +2.7 | +20 W | +6 to +10 | 311 to about 340 W at 139.8 MH/s: 2.26 to 2.43 microjoules (+8 percent) | modelled | +| RX 9070 XT | 17.90 against 18.15 MH/s (-1.4 percent, measured) | 0: every 4-byte read already costs a 64-byte line (measured, the probe) | 0 | +1.4 percent | measured | +| Apple M5 Max | 28.26 against 27.74 (+1.9 percent, measured, under load) | 0: 64 B at the 4 B rate (3.51 G per second at both, Apple OpenCL probe) | 0 | about -2 percent | measured | +| RTX 5070 Ti, 4070, 4060 Ti (32-byte sectors) | as the 5090 (approximate) | the same second sector | the same | about +10 percent (approximate) | approximate | + +The edge, every column (the 5090 at the knee; GDDR7 chip): + +| Row | W = 1 (the record) | W = 16 | Change | +|---|---|---|---| +| Zero shadow, GDDR7 | 3.59x | **3.0x** (1.84 / 0.62) | -16 percent | +| Zero shadow, one HBM3 stack | 5.2x | 4.6x | -12 | +| Class v4 shadow at `k = 0.3` | 3.52x | **3.05x** ((1.84 + 0.652) / (0.62 + 0.196)) | -13 | +| at `k = 0.5` | 2.94x | **2.63x** | -11 | +| at `k = 1` | 2.08x | 1.96x | -6 | +| The 5090 unlocked, zero shadow | 4.85x | 3.9x | -20 | +| Against the M5 Max's joule (0.78, unchanged) | 1.7x | 1.26x | -25 | +| The SRAM full store, zero shadow (lane B's 17x against 2.40) | 17x | 17x against the stock point; 13x against the W = 16 5090 at 2.43 | 0 on the chip; the denominator moves | + +The identity check, in the identity's own terms: `F` here is the second sector's joules on the card (about 1.5 nJ per read, 0.19 microjoules per hash at the knee) and `k` is the chip's cost for the same sector over the card's: 1.15 to 1.2 nJ over 1.45 to 1.65, **`k` about 0.75 to 0.8**, the highest `k` of any forcing work on the table, because 1.15 of it is the DRAM device's own movement at `k = 1` and only the card's fabric share is `k` under 1. The slope at `F = 0` from the record's section 2 (`(E_mem - k E_card) / E_mem^2`) at `k = 0.78`: -1.9x per microjoule, so 0.19 microjoules buys about 0.36x, and the full arithmetic above gives 0.6x because the chip's `E_mem` rises by the same joules it would have charged to `k F`. It stacks with the ALU shadow (the DRAM and the ALUs are different hardware; the fold adds about 5,760 ALU ops per hash, 6 percent of the shadow's count, inside the wait) and with the operating point. + +What a chip can do about it: nothing cheaper than paying. Items are pseudo-random (incompressible); the fold uses all 16 words with a state-dependent map (read-width.md section 1), so a pre-folded dataset does not exist; a narrower access atom does not exist on any DRAM (32 B is the minimum on GDDR7, HBM3, HBM4 and LPDDR6, JEDEC, claimed); spreading the words over devices multiplies the sectors; the HBM parts pay the movement at their own lower pJ per bit, which is the `E_mem` lever the model already carries. The one chip untouched is the SRAM full store (64-byte macro lines), which is lane 3's chip and not this lane's. + +Cost to the honest tiers and the rule of 5 October: the 5090 class pays about 10 percent of energy per hash for no rate (a rig at the knee about +27 W per card); the 5070 Ti and the 12 GB and 8 GB NVIDIA tiers the same share (approximate); AMD and Apple nothing (their lines are 64 B or wider already); a pool user nothing; a node about 1 percent on the verifier; the public claim moves from "3.6x at the knee" to "3.0x" at zero shadow and from "2.1x at `k = 1`" to "2.0x". The known-failed case for the acceptance rule: the w16 class's census (0 to 14 of 60 rejected) and its F8 uniformity row are owed at 2^24 on 64 seeds under the sub-version 3 rule, since the 5 October census ran the version 2 rule. + +What it changes in class v6: layer 1's read-width band is {1, 4} words with the warrant "w16 moved the chip not at all" (the design's section 2 and lane B's method note). That warrant is the error this lane corrects; the band becomes the floor W = 16 (one item), and the lane A ranking's row 3 ("keep the read width out of the era draw: a draw over {4 B, 16 B} is harmless and worthless") keeps its first clause and loses its second. One measurement owed before the cut: the 5090's watts at w16 at the knee and unlocked (the pack exists; one PC job of four rows) and the M5 Max's power channels at w16. + +### 3.2 Two items per read: W = 32, 128 bytes, one NVIDIA L2 line (KEEP for v7, rank 3) + +| Side | At W = 32 | Label | +|---|---|---| +| GDDR7 chip | one activate plus four column accesses: 0.909 + 4 x 1.15 = 5.5 nJ per read; 21.3 G activates x 128 B = 2.7 TB/s is over the board's 1.79, so the chip becomes bandwidth-bound at 14 G reads per second, 110 MH/s; `E_mem` = 128 x 5.5 nJ + 35 W / 110 M = 0.70 + 0.32 = **1.02 microjoules** (+120 percent); capex per MH/s +50 percent (USD 2.8 to 4.3) | modelled | +| HBM3 one stack | 0.909 x 0.66 + 4 x 0.59 = 2.96 nJ; 10.7 G x 128 B = 1.37 TB/s over the stack's 0.82, bound at 6.4 G reads, 50 MH/s; `E_mem` = 0.38 + 14 W / 50 M = 0.66 microjoules (+105 percent) | modelled | +| RTX 5090 at the knee | 17.9 G x 128 B = 2.3 TB/s over 1.79, so about 14 G reads per second, 110 MH/s (-20 percent; the w64 row measured the bind at 256 B: 71.9 MH/s, 0.58 share); DRAM +3 sectors x 1.15 nJ x 14 G = +48 W, fabric +15 W: about 276 W at 110 MH/s, **2.5 microjoules** (+50 percent) | modelled from the measured w64 bind and the rated peak; the w32 pack does not exist yet | +| Apple M5 Max | 3.5 G x 128 B = 448 GB/s against a measured 522 GB/s stream: 0 to -14 percent of rate (approximate; the Apple fetch granularity at 128 B unmeasured) | approximate | +| RX 9070 XT | two 64-byte lines per read: 2.4 G x 128 B = 307 GB/s of 636; unmeasured whether the second line costs it a second fetch | unmeasured | +| The edge, 5090 knee, GDDR7 | zero shadow 2.5 / 1.02 = **2.4x**; with the class v4 shadow at `k = 0.3`: (2.5 + 0.652) / (1.02 + 0.196) = **2.6x**; at `k = 0.5`: 2.3x; at `k = 1`: 1.9x | modelled | + +Reading: wider reads walk both sides into the regime where the DRAM's own joules are most of both bills, and there the edge tends toward the ratio of two bandwidth-bound machines (about 2.5x, modelled) at the price of the honest card's rate (the Ethash shape the design avoided on 5 October for that reason). W = 16 takes most of the per-joule gain at no rate; W = 32 takes the rest at a fifth of the 5090's rate and a 50 percent rise in the chip's capex per MH/s. It is a v7 candidate behind two measured rows (the 5090 and the M5 Max at a `w32` export, rate and watts), not a v6 change. + +### 3.3 Independent chains per hash (memory-level parallelism as the card's lever) (KILL, one probe would reopen it) + +The idea: four independent dependent chains of 32 reads per hash, folded at the end, so each resident hash keeps four reads in flight and an occupancy-bound card climbs toward its memory's activate ceiling while the chip (already at the ceiling) gains nothing; an `E_card` lever. The number: the 5090 reads 17.5 to 18.2 G per second at 4 B as "the best over lanes in flight" (the 5 October probe), which means more lanes did not help, so the bind is the memory system, not the in-flight count; the bound is the ceiling's 21.3 G, +17 to +22 percent, and the expectation is 0 to 5 percent (the microbench's `l2_indep4` against `l2_chase` read +2 percent on the L2-bound pattern). The 9070 XT caps at 2.63 G at 4,096 lanes (measured) with the same plateau shape. A class change for an expected few percent, which also cuts the sequential depth per hash to 32. KILL; the probe that would reopen it is one `dram_indep4_1g` row in the microbench (60 s on PC 1). + +### 3.4 Row-straddling items (two activates per read) (KILL, dominated) + +A 64-byte item laid across a row boundary costs the chip two activates: 2 x 0.909 + 2 x 1.15 = 4.1 nJ per read (+105 percent) and halves its activate-bound rate (83 MH/s), so `E_mem` = 128 x 4.1 nJ + 35 W / 83 M = 0.52 + 0.42 = 0.94 microjoules; the card's activate-bound rate halves too (about 83 MH/s if it sits at 82 percent of the same ceiling), at about +22 W: 2.95 microjoules. Edge 3.1x at a 39 percent rate loss, against the same-row second sector's 3.0x at no rate loss. Dominated by 3.1. KILL (the invention lane's 2.13 killed its 256-byte form on the w64 row; this is the 64-byte form, and it loses to the aligned item). + +### 3.5 Scratch state in DRAM (KILL) + +A per-lane scratch line in DRAM, read-modify-written per step, would be `k = 1` work if it reached the DRAM on both sides; it does not: 7,262 hashes in flight x 64 B is 465 KB, inside the 5090's 96 MB L2 (an L2 hit at 1.4 nJ, `k` 0.1 to 0.3 against the chip's lane SRAM), and a scratch large enough to miss the L2 (13 KB per lane) is 15 MB of SRAM on the chip (about USD 5 of N5). The 5 October scratch rows measured the card's side: -12 to -48 percent of rate on the 5090. KILL (`docs/analysis/scratch-soundness.md` and the read-width entry say the same). + +### 3.6 A note on the sunk card (not a candidate) + +Both sides are capex-dominated (the record's 16.1: 7x to 10x their electricity per MH/s-hour), and the chip's economic edge is capex per MH/s (USD 2.8 against 14.7). For a home miner whose card was bought for something else the capex is sunk, so that tier's cost per MH/s-hour is its electricity alone, USD 0.000084 at the knee and USD 0.05 per kWh, against the chip's all-in USD 0.00018 (capex over two years plus electricity): the sunk card is 2.2x CHEAPER per MH/s-hour than the chip until the chip's capex is amortised. The honest-denominator lane owns this reading; it is recorded here because it is the one place the identity's per-joule framing understates the honest tier. + +## 4. The KILL list, with the number + +| Candidate | The identity check | The number | Label | +|---|---|---|---| +| Tensor tiles as the shadow (int8) | `k` 0.2 to 0.4 on the class's tile, 0.4 to 0.7 on the dense tile: never above the ALU shadow's band | the M5 Max -35 percent at 1,024 tiles, -78 at 4,096 (measured); the same premium as the ALU shadow by construction | measured card, claimed chip | +| FP8, FP16, BF16 tiles; the texture interpolator; the rasteriser | not bit-exact across vendors | | spec and vendor documents | +| The hardware video decoder | bit-exact by standard, but the honest card is decoder-bound (a fixed engine count) and the verifier decodes at milliseconds per frame | engines per card two to four (approximate) against any count on a chip | approximate | +| The RT core | traversal implementation-defined on Vulkan, DXR, OptiX and Metal; the BVH opaque; `k` 0.01 to 0.2 on the public per-ray figures | 4 to 33 nJ per ray on a dedicated unit (claimed) against 290 to 750 measured board-level on an RTX 2080 | claimed and measured | +| Memory-level parallelism per dollar | no asymmetry per channel; the correction is the PHY's node (section 2.3, kept as rank 2) | `E_mem` under +0.06 microjoules | claimed, modelled | +| Proof of latency | the chain checks values; the block is 1 s against a 0.4 microsecond round trip; the farm's node is local; a sequential prefix favours the chip about 4x | 2.5 million round trips per block | measured card latency, modelled chip | +| The drawn address map | firmware; no prefetch exists for a dependent chain | 0 to the chip, 0.8 to 3.2 percent to the cards | measured spread | +| Independent chains per hash | the 5090's read rate plateaus over lanes; the bind is the memory system | expected 0 to 5 percent, bound 22 | measured probe | +| Row-straddling items | dominated by the aligned item | 3.1x at -39 percent of rate against 3.0x at 0 | modelled | +| Scratch in DRAM | the card's scratch sits in L2; the chip's in SRAM | 465 KB against 96 MB | measured rows | + +## 5. Consequences per tier (the standing rule of 5 October 2026) + +| Tier | What this file means | What is being done | +|---|---|---| +| Home miner, one 8 GB NVIDIA card (32-byte sectors) | W = 16 costs it about 10 percent of energy per hash (approximate, the 5090's share scaled) for no rate; against the GDDR7 chip its row improves by the same 10 to 15 percent as the 5090's | the w16 watts row on the 5090 first; the 4060 Ti class scaled from it | +| One 12 GB card (4070, 5070) | the same share; the 4070 at its tune point about +8 W (approximate) | the same | +| One 16 GB card (RX 9070 XT) | nothing: its 64-byte line already moves the whole item (measured 5 October); its row against the chip improves by the chip's +33 percent alone, 23x to about 17x | nothing to measure; the vendor-share metric | +| One 24 or 32 GB card (5090 class) | +26 to 30 W at the knee (modelled), +2.7 percent of rate (measured); 3.6x to 3.0x at zero shadow, 2.9x to 2.6x with the shadow at `k = 0.5` | one PC job: the w16 pack at the 1,300 lock and unlocked, nvidia-smi at 1 Hz, four rows | +| Apple (M5 Max and the unified tiers) | nothing: 64 B at the 4 B rate (measured); the honest best per joule improves against every DRAM chip by the chip's cost alone, 1.7x to about 1.3x on GDDR7 | the M5 Max power channels at w16 (one Mac measurement under the measure lock, deferred until the Mac rule allows) | +| A rig | +12 percent of watts on NVIDIA cards for no rate; the chip's bill per MH/s-hour up a third on energy | design 1 (the knee lock) as the default, then W = 16 | +| A pool user | nothing changes in shares; the class change is announced by the 95 percent signal | | +| A node operator (the verifier) | +1 percent (0.610 against 0.604 ms per 32 hashes, measured); W = 32 about +4 percent | nothing | +| A chip | GDDR7 +33 percent of energy per hash, HBM3 +24, the base die +50, the SRAM store 0; the cheapest project is a 12 nm GDDR6 part at USD 20 M to 30 M, not a 28 nm one at 5 M | the chip model's rows 5.4 and 5.6 and the record's 16.2 corrected (owed to the Counter lane) | +| The public claim | "3.6x at the knee, 2.1x at `k = 1`" becomes "3.0x and 2.0x" on W = 16 (modelled card side until the job); the tensor `k` sentence loses "4x to 30x"; the Ethash precedent gains the 13x row | the Counter lane's texts, after the 20:00 word | + +## 6. Unverified and owed + +- The 5090's watts at w16 are modelled (+26 to 30 W at the knee); the 5 October job carried no power sampling. One PC job on the existing pack decides the card side of rank 1. The M5 Max's power channels at w16 are unmeasured (the rate is). +- The chip side of rank 1 is the model's own data-movement term applied twice (4.5 pJ per bit, Micron, claimed, streaming; O'Connor's 3.48 for HBM2); no one has measured a second-sector column read's energy on GDDR7 (section 2.6: the public record has no random-read joule on any current part). +- The fabric share of the card's second sector (0.3 to 0.5 nJ) is approximate. +- The W = 32 rows rest on the rated 1.79 TB/s and the measured 256-byte bind; no w32 pack exists. +- The tensor chip figures are vendor specifications at TDP (claimed), the M4 ANE and the H800 loop measured; the RT figures are claims and one measured RTX 2080 row; the PHY node facts are vendor IP pages read today. +- The project-cost corrections use the history's figures (a 7 nm-class startup project USD 50 M to 75 M, claimed; the 12 nm figure scaled, approximate) and the mission lane's cap model. +- The w16 class's census under the sub-version 3 rule and its F8 row at 2^24 on 64 seeds are owed before any cut (the 5 October census ran version 2's rule). +- Nothing ran on the Mac, on a box or on a PC for this file. + +## 7. Sources + +Internal: `docs/analysis/counter-asic-4-research.md` (sections 2, 4, 5, 8, 15.1a, 16, 20), `docs/analysis/chip-model-v3.md` (5.1, 5.3, 5.4, 5.6, 5.10 to 5.12), `docs/design/class-v6-rotating-family.md` (sections 2, 7, 7a, 7c), `docs/analysis/class-v6/hardware-future.md` at 7618e729 (sections 0, 1, 2, 4.9, 5), `docs/analysis/class-v6/history.md` (sections 3 and 5), `docs/analysis/class-v6/invention.md` on `class-v6-invention` (sections 1 and 2), `docs/plans/read-width.md` (sections 1, 1.5, 4.1), `docs/bench-log.md` (the 5 October read-width entry), `docs/analysis/scratch-soundness.md`, spec 01 (1.5, the item table). + +External, all read 8 October 2026 by this lane's three research sub-agents (tensor, RT, literature) and labelled where carried: the NVIDIA RTX 5090 page and the Blackwell architecture whitepaper; the Ada whitepaper; arXiv 2501.12084v2 (Hopper tensor-core microbenchmarks); Lenovo LP2226 (B200); the AMD MI300X data sheet and MI355X notes; the RX 9070 XT page; Qualcomm's AI 100 Ultra brief; eeNews Europe (MTIA v2); AnandTech 21342 (Gaudi 3); the Hot Chips 2025 Ironwood slides; arXiv 2608.28048 (Trainium2); Tenstorrent's Blackhole specification; the M4 ANE measurement (maderix.substack.com); cnx-software (Hailo-8); design-reuse 52539 (Untether); hackster.io (Axelera); NVIDIA Research's VSQ accelerator (JSSC 2023); Dally, Hot Chips 2023 keynote, slides 12 and 24; the Vulkan ray traversal chapter; the DXR specification; the NVIDIA developer forum thread 309730 (14 October 2024); the Blender Cycles source note; the Mach-RT TVCG preprint; Hot Chips 31 (Turing); the AMD Hot Chips 2025 RDNA 4 slides; embedded.com (RayCore); AnandTech 7870 (GR6500); arXiv 2409.06000 (RayFlex); the TRaX CGI 2018 study; Ylitie et al., HPG 2017; Chou et al., MICRO 2023; par.nsf.gov 10156963 (Choe et al., MEMSYS 2019); arXiv 2508.06795; SemiAnalysis "The Memory Wall"; the SK hynix ISSCC 2024 GDDR7 paper (ResearchGate 378947349); arXiv 2410.12990; arXiv 2405.06081; asicminervalue (iPollo V2H, V1, A11 Pro); kryptex (Jasminer X16-Q, FishHash); whattomine (the 5090 on Etchash); monero-project issue 10270; arXiv 2512.01437v2 and decrypt.co 363018 (Qubic); Business Wire 20240925780123 (Cadence GDDR7 on N3) and 20221115005674 (GDDR6 at 12 nm); innosilicon.com GDDR7; kurnal-insights (GB202 die); TrendForce 24 September 2026 (the 2 GB GDDR7 part); arXiv 2403.13230 (BFT-PoLoc); eprint 2026/694 (unread, refused).