Merge class-v6-floor-invention-docs ca381a45 into master (gate: green on ca381a45, recorded by tools/ci/pre-push.sh; landed on the box mirror under the exception declared by main: main's ruling, 7 Oct 2026 19:5x UK: the GitHub account is suspended, lanes land on the box mirror's master, the box gate stamp is the verdict; GitHub gets the fast-forward when it answers)

This commit is contained in:
igneum-labs 2026-10-08 13:02:33 +00:00
commit e266d8a376

View file

@ -0,0 +1,281 @@
# Class v6 floor lane 5, invention: what none of the four floor lanes covers, checked against the identity
8 October 2026, 12:48 to 19:xx UK, branch `class-v6-floor-invention` from `counter-asic-4` (base 171618c5 after the fast-forward; first cut at 13:0x UK; corrected 13:1x UK for the read-width naming: `w16` is 16 bytes, `w64` is the 64-byte item; the rented-card rows of 13:2x to 13:4x UK carried in section 3.1, which reverse rank 1 at stock; the 5090 row at 13:4x UK). The open-ended lane of the floor research: four other lanes work the known levers (SM-sparse gating, the shadow's `k` from RTL synthesis, the SRAM full-store chip and the dataset floor, the honest denominator per tier); this file finds what they do not cover. Every candidate is checked against the identity first and killed or kept with a number. Every number carries a label: **measured** (a card on a named job, the file named), **modelled** (arithmetic on the chip model's cited inputs, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling). Nothing here changes a consensus object, a pack or a served number. Research and reasoning ran on this Mac; no build, no benchmark and no miner ran anywhere for this file.
## 0. One page
**The identity, as the record states it** (`docs/analysis/counter-asic-4-research.md` section 2): against the chip that stores the dataset, `edge = (E_card + F) / (E_mem + k F)`; at zero premium the edge is `E_card / E_mem` (3.6x on a 5090 at its 1,300 MHz knee against the GDDR7 board, measured card, modelled chip), and a shadow op with `k` under 1 never fully closes it. The one thing with `k = 1` by construction is the memory itself: the same devices on both sides. The record forced one DRAM operation per read (the activate and one 32-byte sector) and called every wider read "nothing to force: the chip already pays the sector". That is true up to 32 bytes; at 64 bytes the chip pays a second sector, and this lane's question was whether the honest card pays it at the DRAM's price or at its own.
**The candidate this lane found, measured today, and where it stands.** The dataset item is 64 bytes (16 words; spec 1.5). Reading the whole item (W = 16 words, the read-width branch's class `w64`, bit-exact on four runtimes, fingerprint `836e56e7d496e980`, the verifier +4 percent) forces a second 32-byte sector per read in the same open row: on the chip model's own inputs (909 pJ of activation plus 1,150 pJ of movement per sector, Micron's 4.5 pJ per bit, claimed) the GDDR7 chip's energy per hash rises from 0.466 to 0.62 microjoules (+33 percent, its rate unchanged), one HBM3 stack's from 0.321 to 0.40 (+24), and on floor lane 3's wire figure the N2 SRAM die's from 0.036 to 0.126 (66x to 19x). The design's "never 16 words" rested on one row, the 5090's 71.9 MH/s (-47 percent) in the exported kernel at full occupancy, while the same card's own 64-byte probe read -13 percent; so the lane rented cards and measured the card side at stock (section 3.1): **the rate question is settled per architecture, and the energy question kills the lever at stock.** A one-request form of the same kernel text (PTX `ld.global.L2::64B` on the first load of each item; bit-exact, the same fingerprint) holds the 4-byte rate on the RTX 4090 at full occupancy (63.65 against 64.73 MH/s, 98 percent, PASS) and on the RTX 5090 at a sparse shape (145.7 against 142.6 MH/s, sp170-w4, PASS); on the H100 the best form is -21 percent and the hint does nothing (FAIL); on the RTX 3090 the best form is -13 percent (FAIL). But the card pays the second sector at its own fabric's price, not the DRAM's 1.15 nJ: 8.6 nJ on the 4090 (+71 W at 8.3 G reads per second for no rate) and 4.7 nJ on the 5090 (+88 W at 18.6 G), so the energy per hash rises 34 percent on the 4090, 25 on the 5090, 33 on the H100's best form and 25 on the 3090, against the modelled chip's 33 (GDDR7) and 50 (HBM3). The sector's `k` is about 0.13 (4090) to 0.24 (5090) on GDDR7, under the ALU shadow's 0.3 to 0.8; the edge at zero shadow moves 0 on Ada, 7 percent on Blackwell (4.7x to 4.4x at stock) and 0.9x in the chip's favour on Hopper, and 0 with the shadow on (3.3x at `k` 0.5 either way). **KILL at stock on measured rows.** The one row that could revive it is the honest card's knee (a smaller fixed share and a lower fabric voltage; a rented pod refuses `-lgc`), owed as a single PC 1 row and not blocking anything: W = 16 stays out of class v6.
**Everything else on the brief is killed, with the number.** The tensor core: the record's `k` 0.03 to 0.3 is not defensible at its floor (the 0.04 pJ figure is INT4 at 0.46 V on a test chip with no memory system); shipping merchant silicon at nominal voltage reads 0.30 to 0.56 pJ per INT8 MAC (Meta MTIA v2 0.51, Qualcomm AI 100 Ultra 0.34, AMD MI355X 0.56, NVIDIA B200 0.44, Apple M4 ANE 0.30 measured; all claimed unless marked), and a Hopper tensor core under a pure MMA loop measured 0.34, so the honest band against the tile the class runs (1.5 pJ per MAC at the knee, measured) is `k` 0.2 to 0.4 and against the dense wide tile (0.83 at the knee) 0.4 to 0.7: equal to or under the ALU shadow's 0.3 to 0.8, never above it, and the Apple cost (-35 percent of rate at 1,024 tiles per hash, measured) kills it as content regardless. The RT core: traversal results are implementation-defined on every API (Vulkan "no ordering guarantee", DXR "no defined order", NVIDIA "could change depending on which driver, which GPU"), the acceleration structure is opaque, and a dedicated unit reads 4 to 33 nJ per ray against 290 to 750 nJ measured board-level on an RTX 2080 (`k` 0.01 to 0.2): dead on bit-exactness and on `k`. Memory-level parallelism per dollar: no asymmetry per channel, but a correction to the chip model: no 28 nm GDDR7 PHY exists (the shipped ones are N3 and FinFET; the oldest GDDR6 PHYs are 12 nm), so the record's cheapest chip (a USD 5 M project, a USD 17 M break-even cap) is not buildable; the floor is a 12 nm GDDR6 part at about USD 20 M to 30 M (cap USD 70 M to 100 M) or an N5-class GDDR7 part at USD 50 M to 75 M (cap 170 M to 250 M), modelled. Proof of latency: the chain checks values, never time; the block is 1 s and a DRAM round trip 0.1 to 0.4 microseconds; the farm's node is on its own LAN; a sequential per-block prefix favours the lower-latency side, which is the chip (about 4x). The time dimension on the address map: a bit permutation is firmware (a few hundred gates), no prefetch exists for a dependent chain, and the literature has no such scheme. The literature since 2023 holds nothing outside the identity; what it adds is a higher Ethash precedent (iPollo V2H, 0.14 J per MH, about 13x a 5090, vendor claim), the measured PAM3 I/O energies (SK hynix, ISSCC 2024) and the PHY-node fact above. Of this lane's own ideas, four more die on the identity (video decode, row-straddling, scratch in DRAM, independent chains per hash), and two items per read (W = 32) dies with W = 16.
### The KEEP list, ranked
| Rank | Item | What it moves | Cost to the honest tiers | Where | Label |
|---|---|---|---|---|---|
| 1 | **The chip model's project floor corrected: no 28 nm GDDR7 PHY** | the cheapest buildable chip's project USD 20 M to 75 M (cap USD 70 M to 250 M), not USD 5 M (cap 17 M); the N5-core row's cap 170 M to 250 M, not 100 M | none | chip-model-v3 5.4 and 5.6 and the record's 16.2 | claimed (vendor IP pages), modelled (the mission lane's cap) |
| 2 | **The tensor `k` band corrected** (0.2 to 0.4 on the class's tile, 0.4 to 0.7 on the dense tile; not 0.03 to 0.3) | no served row (the tile stays a worse-or-equal lever) | none | the record's 15.1a and 20.4 and chip-model 5.11 | claimed and measured |
| 3 | **The card's cost of a second sector, measured: about 8.6 nJ on a 4090 at stock** | closes the read-width question on the identity's own terms (`k` about 0.13 on the sector); the hinted one-request load form is kept as a miner-side kernel fact for Ada (98 percent of the 4-byte rate at 64 bytes) | none (nothing ships) | the read-width plan's 4.1 and lane 3's width table; the knee row owed on PC 1 | measured |
| 4 | **The Ethash precedent's top row** (iPollo V2H, 3.4 GH/s at 475 W, about 13x a 5090 per joule) | the precedent band 2.1x to 13x, not 2.1x to 6.8x | none | `asic-resistance-history.md` and chip-model 5.1 | claimed |
No item on this list is a class v6 change; nothing here reads `k` at or above 1 on a measured GPU figure.
### The five-sentence reading for the 15:45 close
The identity leaves one piece of work a chip pays at the card's own price, the memory's data movement, and this lane's candidate was the second sector of the 64-byte item (W = 16 words); measured today on rented cards at stock, a one-request form holds the 4-byte rate on Ada (98 percent on a 4090) and fails on Hopper and Ampere (-21 and -13 percent at best), but the card pays that sector at about 8.6 nJ through its own fabric against the DRAM's 1.15, so its energy per hash rises 25 to 34 percent, the same as the modelled chip's 33, and the lever reads `k` about 0.13: dead at stock, W = 16 stays out of class v6, with the knee row owed on PC 1 as a formality. The tensor core, the RT core, proof of latency and the drawn address map are dead on the identity or on bit-exactness, with the tensor `k` band corrected upward to 0.2 to 0.7 (still never above the ALU shadow's) and the RT core's traversal implementation-defined on every vendor. The chip model's cheapest chip does not exist as priced: a GDDR7 PHY needs a FinFET node (N3 shipped; 12 nm is the oldest GDDR6 PHY), so the floor project is USD 20 M to 75 M and the floor break-even cap USD 70 M to 250 M instead of 17 M, which goes into the model now. Nothing in the literature since 2023 offers a lever outside the identity, and this lane's own four other ideas die on it with numbers. The floor section's default for this lane therefore stands as written: no new lever; what the lane adds is two corrections to the model, one measured card-side figure, and the measured closing of the read-width question.
## 1. The frame
| Term | Value at the 5090's 1,300 MHz knee | Label | Source |
|---|---|---|---|
| `E_card`, class v3 | 1.67 microjoules (127.3 MH/s at 213.0 W; 134.6 at 223.3 on the efficiency pass) | measured | the record 20.3; `docs/plans/counter-asic-3-status.md` |
| `E_mem`, the `f = 1` GDDR7 chip | 0.466 microjoules: 166 MH/s at 77.6 W (42.6 W of reads at 2.0 nJ, 20 W static, 15 W controller) | modelled | chip-model-v3 5.3 and 5.4 |
| `E_mem`, one HBM3 stack | 0.321 microjoules: 83.6 MH/s at 26.8 W (12.8 W of reads at 1.2 nJ, 14 W static and controller) | modelled | the same |
| `E_mem`, the N2 SRAM full store at the hash's width | 0.036 microjoules at W = 1 (0.25 nJ per read: 1.3 pJ per bit of wire over 80 bits plus 0.1 nJ of macro and 0.05 of controller), 66x the 5090's stock point; 0.126 at W = 16 (0.88 nJ, 560 bits), 19x | modelled (floor lane 3, section 10.3 of the class v6 design) | `docs/analysis/class-v6/floor/sram-and-floor.md` |
| The read's energy, split | GDDR7: 909 pJ activation plus 4.5 pJ per bit x 256 bits = 1,150 pJ of movement and I/O (44 / 56 percent); HBM3: 909 plus 3.48 x 256 = 890 scaled by 4.12 / 6.25 | modelled on O'Connor et al. 2017 Tables 2 and 3 and Micron's and Samsung's pJ per bit (claimed) | chip-model-v3 5.3 |
| `F`, the class v4 premium | 0.652 microjoules (82.8 W over 126.9 M x 102,100 ops: 6.4 pJ per counted op) | measured | the record 20.3 |
| `k`, the ALU shadow | 0.3 to 0.8 (the 5090's 6.2 to 11.3 pJ per op measured against a 5 nm SIMD array's 2 to 5, approximate) | GPU measured, chip approximate | the record 15.1a |
| The dataset item | 16 words, 64 bytes; `dataset[w] = item(w >> 4)[w AND 15]`; the lottery hash reads one word per load | design | spec 1.5 and the item table; `docs/plans/read-width.md` |
| The read-width classes, by name | `w4` = 4 bytes (W = 1 word, the lottery hash); `w16` = 16 bytes (W = 4, one sector); `w64` = 64 bytes (W = 16, the whole item, two sectors on the 5090); `w64x4` = 64 bytes at 32 loads | design | `docs/plans/read-width.md` 1.5; the bench-log entry of 5 October |
| The 5090's dependent-read probe, best over lanes in flight | 17.5 to 18.2 G reads per second at 4 B; 18.0 to 19.9 at 16 B; 9.1 to 15.7 at 64 B (9.1 at 4 M lanes) | measured, 5 October | `docs/bench-log.md`, the read-width entry |
| The hash rates at the widths, unlocked, no power sampling | 5090: w4 136.1, w16 139.8, w64 71.9 MH/s; 9070 XT: 18.15, 17.90, 17.59; M5 Max: 27.74, 28.26, 28.27 | measured, 5 October | the same |
| The card's whole-card marginal per dependent DRAM read | 10.9 nJ unlocked, 8.7 nJ at the lock | measured, 8 October | the record 15.1a |
The four doors a candidate can open (the invention lane's section 1, kept): raise `k`; raise the chip's capex or project; shorten the chip's useful life; lower the honest card's cost at the same `F`. This file adds the fifth door the identity itself names: **force work whose `k` is 1 by construction, which is only the memory's own.**
## 2. The candidates on the brief
### 2.1 Work the GPU already does at near-ASIC efficiency
**The tensor core, re-read.** The question: is "a merchant N2 project beats Blackwell's tensor core per joule by 3x to 30x" defensible, or is the tensor shadow at full rate closer to `k` 0.7 to 1?
| Chip | Node | Year | Dense INT8 TOPS (FP8 where marked) | W | pJ per MAC (2 / (TOPS per W)) | Label | URL (read 8 October 2026) |
|---|---|---|---|---|---|---|---|
| RTX 5090 | 4N | 2025 | 838 | 575 | 1.37 at spec; **1.36 measured** on the dense s8 m16n8k32 tile at 80 percent of peak unlocked, 0.83 at the 1,300 lock; the dependent u8 m8n8k16 tile the class runs 2.9 unlocked, 1.5 at the lock (the packs job), 4.1 / 2.2 (the microbench) | claimed spec; measured rows the record's 15.1a and 20.3 | https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/ |
| RTX 4090 | 4N | 2022 | 661 | 450 | 1.36 | claimed | NVIDIA Ada whitepaper |
| H100 SXM | 4N | 2022 | 1,979 | 700 | 0.71 | claimed | https://www.nvidia.com/en-us/data-center/h100/ |
| H800, mma and wgmma INT8 dense under a pure MMA loop | 4N | 2025 | | sub-TDP loop | **0.34 / 0.54** | measured | https://arxiv.org/html/2501.12084v2 Tables 10 and 11 |
| B200 (FP8 / INT8) | N4P | 2024 | 4,500 | 1,000 | 0.44 | claimed | https://lenovopress.lenovo.com/LP2226 |
| AMD MI300X | N5 | 2023 | 2,615 | 750 | 0.57 | claimed | AMD data sheet |
| AMD MI355X | N3P | 2025 | 5,033 | 1,400 | 0.56 | claimed | https://glennklockwood.com/garden/processors/mi355x |
| RX 9070 XT (RDNA 4 WMMA) | N4P | 2025 | 389 | 304 | 1.56 | claimed | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070xt.html |
| Qualcomm Cloud AI 100 Ultra | 7 nm | 2023 | 870 | 150 | 0.34 | claimed | Qualcomm product brief |
| Meta MTIA v2 | 5 nm | 2024 | 354 | 90 | 0.51 | claimed | https://www.eenewseurope.com/en/meta-launches-second-generation-custom-ai-chip |
| Intel Gaudi 3 (FP8) | 5 nm | 2024 | 1,835 | 900 | 0.98 | claimed | https://www.anandtech.com/show/21342 |
| Google TPU Ironwood (FP8) | n/d | 2025 | 4,614 | about 1,100 inferred | about 0.48 | inferred from the Hot Chips 2025 slides | https://hc2025.hotchips.org/ |
| AWS Trainium2 (FP8) | 5 nm | 2024 | 1,299 | about 500 (third party) | 0.77 | inferred | https://arxiv.org/pdf/2608.28048 |
| Tenstorrent Blackhole p150 (FP8) | N6 | 2025 | 774 | 300 | 0.78 | claimed | https://docs.tenstorrent.com/aibs/blackhole/specifications.html |
| Apple M4 ANE (FP16 rate, INT8 runs at it) | N3E | 2024 | 19 | 2.8 | **0.30** | measured | https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615 |
| Hailo-8 | 16 nm | 2019 | 26 | 2.5 | 0.19 at spec, 0.71 measured on ResNet-50 | claimed / measured | https://www.cnx-software.com/2020/10/07/ |
| Untether speedAI 240 (FP8, at-memory) | 7 nm | 2022 | 2,000 | 66 | 0.067 | claimed, never shipped at volume | https://www.design-reuse.com/news/52539/ |
| Axelera Metis (digital in-memory) | 12 nm | 2023 | 214 | about 14 | 0.13 | claimed | https://www.hackster.io/news/ |
| NVIDIA VSQ test chip (INT4) | 5 nm | 2023 | | | 0.021 at 0.46 V on a benchmarking layer; **0.052 at the nominal 0.67 V on full networks**; no INT8 figure | measured, a test chip without HBM or a NoC | https://research.nvidia.com/publication/2023-01_956-topsw-deep-learning-inference-accelerator-vector-scaled-4-bit-quantization |
The reading. Every shipping systolic chip with a real memory system and fabric lands at 0.3 to 0.6 pJ per INT8 MAC at nominal voltage; the figures under 0.15 are in-memory-compute vendor claims and an INT4 test chip at low voltage. Dally's own energy pie for that test chip (Hot Chips 2023, slide 24) has the datapath at 47 percent and the buffers, collector and movement at 53 percent, so a whole chip is about 2x its datapath before any NoC or HBM. The record's `k` floor of 0.03 took the 0.46 V INT4 figure against the 5090's 1.5 to 4 pJ; at nominal and INT8 the chip side is 0.3 to 0.6. The honest band, labelled:
| Against | GPU pJ per MAC (measured) | Chip pJ per MAC at nominal (claimed, the cluster above) | `k` | Reading |
|---|---|---|---|---|
| The tile the class runs (dependent u8 m8n8k16, the `mm1430` pack) at the knee | 1.5 | 0.3 to 0.6 | **0.2 to 0.4** | the forcing lever's `k`; equal to or under the ALU shadow's 0.3 to 0.8 |
| The same unlocked | 2.9 | 0.3 to 0.6 | 0.1 to 0.2 | |
| The dense wide tile (s8 m16n8k32) at the knee | 0.83 | 0.3 to 0.6 | **0.4 to 0.7** | a tensor shadow built of wide tiles would sit at the ALU band's centre; the H800's 0.34 under a pure loop says a GPU tensor core at its best is inside the merchant cluster |
| The dense wide tile unlocked | 1.36 | 0.3 to 0.6 | 0.25 to 0.45 | |
So "3x to 30x" is not defensible: at nominal voltage merchant silicon beats the 5090's dense tile by 2.3x to 4.5x and its dependent tile by 2.5x to 10x, and the 30x needs an INT4 layer at 0.46 V. And "`k` 0.7 to 1 at full rate" is not reached either: the top of the dense band is 0.7. The identity check: the tile's `k` is at best the ALU shadow's and the joules it forces are the same by construction (the record 20.3: 71.5 W against 82.8 at the knee for the same tile count); the chip's downside bet (its `k` floor) is 0.2 against the ALU's 0.3, so the tile remains the worse-or-equal lever, not the 4x to 30x worse lever the record wrote. The cost to the honest tiers is what kills it as content: the M5 Max at -35 percent of rate at 1,024 tiles per hash and -78 percent at 4,096 (measured, the record 20.2a), with no integer matrix path in Metal. Bit-exact on NVIDIA and the CPU (measured); the AMD WMMA layout unverified. **KILL as v6 content; KEEP the band correction (rank 3).** FP8, FP16 and BF16 tiles: the accumulation order inside a tile is unspecified by PTX, so not bit-exact across vendors or generations: KILL.
**The texture unit's filtered gather.** Measured 0.19 nJ per filtered fetch unlocked and 0.10 at the lock (the microbench `tex_linear_f32_256k`); a 9-bit fixed-point interpolator on NVIDIA (CUDA Programming Guide, claimed), vendor-specific fraction widths elsewhere: not bit-exact across the three vendors; a chip's interpolator is one 9-bit multiply-add, `k` far under 1. KILL. The point-sampled fetch is the L2 chase under another name (the same checksum, measured): the L2 lever, dead at `k` 0.1 to 0.3.
**The rasteriser and ROPs.** Unreachable from CUDA, OpenCL and Metal compute; fill conventions, sample positions and depth precision differ by vendor by design. KILL.
**The hardware video codec.** A conformant decoder of a standard (H.264, HEVC, AV1) is normatively bit-exact, which is the one fixed-function block on a card whose output a CPU verifier could reproduce, and a chip would license the same decoder block at the same node (`k` about 0.5 to 1, the same circuit, approximate). It dies on throughput and on the verifier: a card carries a fixed count of decode engines (two to four on current cards, approximate) against any count on a chip, so the honest card's hash rate would be bound by its decoder, not its memory, at a ratio that differs per vendor (the 5090 against the M5 Max's media engine); and a software decode on the verifier runs at milliseconds per frame on one core (approximate), against a 10 ms gate that holds the whole warp. KILL. (This lane's own candidate, recorded here because it is the only fixed-function block that passes the bit-exactness test.)
### 2.2 The RT core
| Unit (node) | Throughput | Power or area | Energy per ray | Label | URL (read 8 October 2026) |
|---|---|---|---|---|---|
| RTX 5090, 170 fourth-generation RT cores | 317.5 RT TFLOPS; "double the throughput for ray-triangle intersection" over Ada; no per-ray or per-clock rate published, no RT-core area or power share | 575 W | not published | claimed | https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf |
| RTX 2080 running DXR (Mach-RT, Table 3) | 288 to 745 Mrays per second over six scenes | 215 W TDP | **290 to 750 nJ per ray, board-level** | measured rays, TDP assumed | https://hwrt.cs.utah.edu/papers/mach-rt_TVCG.pdf |
| Turing TU102 | 10 Giga rays per second claimed | about 260 W | about 26 nJ per ray at the claimed peak | claimed | Hot Chips 31 slides |
| AMD RDNA 4 | 8 ray-box and 2 ray-triangle tests per clock per CU, BVH8, OBBs | no RT power figure | none | claimed | https://hc2025.hotchips.org/assets/program/conference/day1/8_amd_pomianowski_final.pdf |
| SiliconArts RayCore, 6 RTUs, 28 nm | 239 Mrays per second | 1 W, 18 mm^2 | about 4.2 nJ per ray | claimed, FPGA-derived | https://www.embedded.com/?p=4433770 |
| Imagination PowerVR GR6500, 28 nm | 300 Mrays per second | 5 to 10 W for the whole GPU | 17 to 33 nJ per ray | claimed | https://www.anandtech.com/show/7870 |
| RayFlex RTL, 15 nm PDK, 1 GHz | one ray-box or ray-triangle test per cycle | 60 to 85 mW per datapath | 60 to 85 pJ per intersection test | synthesised | https://arxiv.org/pdf/2409.06000 |
| TRaX study, 65 nm | | memory 60 to 95 percent of ray energy, DRAM alone up to 80, compute 1 to 2 | 0.5 to 5 microjoules per ray | simulated | https://hwrt.cs.utah.edu/papers/rt_performance_CGI18.pdf |
Per node visit (Ylitie et al., NVIDIA, HPG 2017, measured): about 15 ray-node and 7 to 9 ray-triangle tests per ray on a compressed 8-wide BVH, 1.5 to 3.2 KB fetched per ray, "memory traffic is the main limiting factor with incoherent rays"; Chou et al. (MICRO 2023) call it "a latency-bound" pointer chase. The identity check: a traversal step over a dataset-sized tree is one dependent DRAM read (the identity's `E_mem`, which both sides pay) plus a box test the RT core does in hardware; the box test is the only forcing, and the chip's cost to match it is 60 to 85 pJ at 15 nm (synthesised) against a GPU share nobody has measured below 290 nJ per ray board-level, so `k` reads 0.01 to 0.2 on the public figures: the chip builds a cheaper RT unit, not a dearer one.
Bit-exactness, which is the first question and the one that decides it: the Vulkan specification gives "no ordering guarantee between operations performed on different intersection candidates", makes watertightness a "should", and performs the hit test "in an implementation specific manner" (https://docs.vulkan.org/spec/latest/chapters/raytraversal.html); DXR: "There is no defined order of execution of any hit shaders for the intersections along a ray path" (https://microsoft.github.io/DirectX-Specs/d3d/Raytracing.html); NVIDIA's own answer (forum, 14 October 2024): which edge or vertex wins is "implementation defined and not guaranteed" and "could theoretically change depending on which API you use, which driver, which GPU"; Blender's Cycles records that matching OptiX or MetalRT bit for bit was impossible. The acceleration structure is opaque in DXR, Vulkan and Metal, and Vulkan's serialised form carries a driver UUID that the next driver may refuse. A CPU verifier therefore cannot reproduce a hardware traversal, and a design that defines its own traversal in software and uses the RT core as an untrusted accelerator forces nothing (the honest card then does the box tests in software or re-checks them). **KILL**, on bit-exactness first and on `k` second. Price of the RT unit a chip would build: 18 mm^2 at 28 nm for six units (claimed), about USD 5 of N5 (approximate).
### 2.3 Memory-level parallelism as the resource: channels per dollar, the controller and PHY
The claim to check: a card has more channels per dollar than a chip built from the same devices plus a controller, because the chip's controller and PHY are a cost the card amortises over gaming.
| Item | Finding | Number | Label | URL (read 8 October 2026) |
|---|---|---|---|---|
| Channels per device | GDDR7: four 8-bit channels and 64 banks per 2 GB device; the 5090's 16 devices are 64 channels and 1,024 banks; a chip built from the same devices has the same count; the next parts (3 GB, 4 GB, 6 GB) carry the same four channels, so channels per GB FALL for the chip and the card alike (Micron ends the 2 GB part) | 64 channels per 32 GB today; 64 per 48 GB on 3 GB parts | claimed (TrendForce, Rambus) | chip-model-v3 5.1; https://www.trendforce.com/news/2026/09/24/ |
| Does a chip buy more channels per dollar | No: the channel count is bought with capacity on both sides (16 devices whatever the dataset), and the chip's only path to more activates per second is more devices, which the card cannot add and the chip pays for at the same price per device | 0 asymmetry per channel | modelled | |
| The PHY's node | No 28 nm GDDR7 PHY exists: Cadence's GDDR7 PHY is silicon-proven on TSMC N3 at 32 Gbps (September 2024); Innosilicon's is "advanced FinFET"; the oldest GDDR6 PHYs are 12 nm (Cadence 12FFC and Six Semiconductor, 16 Gbps); nothing at 28 nm at any GDDR6 or GDDR7 rate | the chip's cheapest controller is a 12 nm GDDR6 part or an N5-class GDDR7 part | claimed (vendor pages) | https://www.businesswire.com/news/home/20240925780123/en ; https://innosilicon.com/html/ip-solution/gddr7.html ; https://www.businesswire.com/news/home/20221115005674/en |
| The PHY's energy | GDDR7 PAM3 I/O measured on silicon: TX 1.103 pJ per bit, RX 0.764 pJ per bit (SK hynix, ISSCC 2024 paper 13.1); at the hash's 17.9 G reads x 256 bits = 4.6 Tb/s the controller-side RX is 3.5 W, inside the model's 15 W controller allowance with its clocking and termination (PAM3 termination is a large share: arXiv 2410.12990, simulated) | `E_mem` moves by under 0.06 microjoules | measured I/O; the allowance approximate | https://www.researchgate.net/publication/378947349 ; https://arxiv.org/abs/2410.12990 |
| The GB202 die | 761.56 mm^2 with 16 x 32-bit controllers and PHYs along the edges; no public mm^2 for the PHY blocks | the card's PHY is a real area it amortises; the chip's is the same area at the same node | measured die, no annotation | https://kurnal-insights.com/en/dieshot/nvidia-gb2025090/ |
The identity check: the chip and the card pay the same per-bit I/O and the same controller physics, so nothing moves `E_mem` beyond the allowance already in the model. What moves is the chip's PROJECT: the record's cheapest row (chip-model-v3 5.6 "a 28 nm-class project at USD 5 M to 30 M", the record's 16.2 "28 nm controller, USD 5 M, break-even cap USD 17 M") prices a PHY that is not made. The corrected floor, on the mission lane's model (cap = C_proj / 0.30 in years 1 to 2):
| Cheapest buildable chip | Memory | Project | Break-even cap, years 1 to 2 | Label |
|---|---|---|---|---|
| A 12 nm GDDR6 controller (16 Gbps PHY IP exists): 32 banks per device, so 32 devices of GDDR6 for the same 1,024 banks and 21.3 G activates per second; about USD 10 per 2 GB device (approximate); a 1,024-bit board | USD 320 | USD 20 M to 30 M (a 12 nm project with licensed PHY IP; the history's 2.5 figures scaled, approximate) | **USD 70 M to 100 M** | modelled, approximate |
| An N5-class GDDR7 controller (the 5090's own 16 devices) | USD 320 | USD 50 M to 75 M (the history's 7 nm-class startup figure, claimed) | **USD 170 M to 250 M** | modelled |
| The same plus the N5 shadow core (class v4 as shipped; the record's 16.2) | | USD 50 M to 75 M (the core joins the controller's die, which is N5 anyway) | USD 170 M to 250 M, not 100 M | modelled |
**KEEP as a correction to the chip model (rank 2), not a lever.** Consequence: the record's break-even row for the chip that matters (the N5 core on a 28 nm controller, USD 30 M, cap 100 M) understates the floor by about 2x, because the controller cannot sit on 28 nm; the shadow per load's "doubling" (USD 60 M, cap 200 M) therefore adds less than it seemed, since the single-N5-die form is already the only buildable form. The 85 W node per farm and the leaves' bandwidth are unchanged.
### 2.4 Proof of latency
A per-block random challenge whose answer window is shorter than one DRAM round trip plus network latency, so a stored copy in a remote farm cannot answer in time but a local card can.
| Check | Finding | Number | Label |
|---|---|---|---|
| The identity | Not an energy lever: it changes who may answer, not what a hash costs. If it bound anything it would bind the chip's node placement | 0 on `E_mem`, `F` and `k` | arithmetic |
| What the chain can enforce | Values, never time (the record's section 6, new-pow 3.2): a block header is the challenge, blocks arrive over the WAN at tens to hundreds of milliseconds, the interval is 1 s and a timestamp is the miner's own; no node can check that an answer arrived within microseconds of a challenge it saw itself at a different time | the shortest enforceable window is a block interval, 1 s, against a DRAM round trip of 0.4 microseconds on the 5090 (415 ns queued, measured) and about 0.1 on a chip (55 ns of controller plus tRCD and tCL, modelled): 2.5 million round trips per block | measured card latency; the chip's modelled |
| The threat model | The stored-dataset chip's node is on its own LAN (class-v5 2a.2: one node serves a farm, the leaves at 16.5 KB/s); there is no "remote chip farm" to lock out; a farm is as local as a card | 0 | design |
| The one form with teeth: a sequential per-block prefix (every miner must run a chain of dependent reads from the block hash before any nonce counts) | Favours the side with the lower unloaded latency, which is the chip: a 100 ms prefix is 250,000 reads at 400 ns on the 5090 and 100 ms of its time; the chip at about 100 ns per read finishes in 25 ms and hashes for the rest | the chip's share of the block rises about 8 percent at a 100 ms prefix; the card loses 10 percent of its time | modelled |
| Fresh join | Unchanged: a joiner needs the previous block, which it needs anyway | 0 | design |
| The verifier | Nothing to verify: time is not in the proof | 0 | design |
| The literature | No proof of work with a sub-round-trip window was found 2023 to 2026; the nearest are watchtower-ping location proofs (BFT-PoLoc, arXiv 2403.13230: about 2 percent Byzantine challengers tolerated, 12 percent misclassification, measured on the Internet; Witness Chain) and PoSME (already in the record); Proximity Signatures (eprint 2026/694) refused the fetch and is unread | null | claimed |
**KILL.** The chain cannot measure time at the scale a DRAM round trip lives on, the farm is local, and the sequential form helps the chip.
### 2.5 The time dimension: the address map drawn often enough to break a chip's prefetch and banking plan
| Check | Finding | Number | Label |
|---|---|---|---|
| The identity | A mapping change is firmware for the stored-dataset chip; no `F`, no `k` | 0 | arithmetic |
| Prefetch | A dependent chain has nothing to prefetch on either side: the next address is the previous read's data | 0 | design |
| Banking plan | The chip's controller maps address bits to channel, bank and row through a table; a drawn permutation of the index bits costs a few hundred gates (the invention lane's 2.7, the spec's 1.13.1 draws of `M`, `R` and `pos` already); the card pays whatever the map does to its own bank parallelism, measured as a 0.8 to 3.2 percent spread across six eras | 0 to the chip; 0.8 to 3.2 percent to the cards | measured (Counter ASIC 2.0 layers 4 and 8) |
| A fixed-function chip | Dead at the first draw, which is layer 1 already; a programmable chip pays USD 25 to 40 of N5 core, already in the record's 16.1 | nothing new | modelled |
| The cadence | Drawing per epoch instead of per era changes nothing above: a table reload is microseconds | 0 | design |
| The literature | No scheme since 2023 rotates an address map to defeat banking or prefetch; Ethash's 30,000-block reseed and RandomX's dataset rebuild are the only moving layouts, both firmware to every chip that shipped | null | claimed |
**KILL.** Covered by layer 1 and the invention lane's row; nothing a programmable chip pays beyond the core it already carries.
### 2.6 The literature since 2023, what the record's section 8 missed
| Item | What it found | Number | Label | URL (read 8 October 2026) |
|---|---|---|---|---|
| Choe, Moreshet, Bahar, Herlihy, "Attacking memory-hard scrypt with near-data processing", MEMSYS 2019 | the one paper that runs a memory-hard function on a PIM model: one in-order core per 128 MB vault | 1.5x over the host, no energy figure | simulated | https://par.nsf.gov/servlets/purl/10156963 |
| Blocki and Smearsoll, TCC 2025; Blocki et al., CRYPTO 2026 (arXiv 2508.06795) | pebbling bounds on MTP-style and data-dependent memory-hard functions | complexity only, no joules | theory | https://arxiv.org/abs/2508.06795 |
| SemiAnalysis, "The Memory Wall", 3 September 2024 | off-chip movement about 20x the cell read; an HBM access about 95 percent interface, 5 percent cell, for streaming | 2 pJ per bit interface against 0.18 pJ per bit activation amortised over a streamed row; for THIS hash's random 32-byte read the activation is 909 pJ against 1,150 of movement (44 / 56 percent), so a shorter link (lane B's base-die and bonded rows) cuts the larger half, which the model already carries | claimed; the split modelled | https://newsletter.semianalysis.com/p/the-memory-wall |
| SK hynix GDDR7 at 35.4 Gb/s per pin, ISSCC 2024 paper 13.1 | PAM3 TX 1.103 pJ per bit, RX 0.764 pJ per bit | the I/O share of the 4.5 pJ per bit device figure | measured silicon | https://www.researchgate.net/publication/378947349 |
| ETH SAFARI real-chip DRAM studies 2024 to 2026 | all on DDR4; no GDDR6, GDDR7 or HBM3 activate-energy measurement exists in public | the record's 909 pJ stands as the only figure | measured, DDR4 | https://arxiv.org/pdf/2405.06081 |
| iPollo V2H, V2, V2X (Ethash, November 2024) | 3.4 GH/s at 475 W; 10 GH/s at 1,500 W; 1.2 GH/s at 165 W | 0.14 J per MH: about 13x an RTX 5090 at 1.8 J per MH on Ethash; the precedent's top row, above the record's X16-P at 6.8x; all of it the memory system | vendor spec, claimed | https://www.asicminervalue.com/en/miners/ipollo/v2h ; https://www.whattomine.com/coins/382-egaz-etchash/gpus/92-nvidia-geforce-rtx-5090 |
| Jasminer X16-Q (Ethash, ETC) | 1,950 MH/s at 620 to 630 W | 0.32 J per MH, about 5.7x a 5090 | vendor spec | https://pool.kryptex.com/device/asic/Jasminer/x16-q |
| Innosilicon A11 Pro (2021); iPollo V1 (June 2022) | 1.5 GH/s at 2,350 W; 3.6 GH/s at 3,100 W | 1.3x and 2.3x a 4090 | vendor spec | https://www.asicminervalue.com/miners/innosilicon/a11-pro-eth-1500mh |
| Pinecone R1X (RandomX, January 2026) | 1.2 MH/s at 2,055 W, listed against the X9's 1 MH/s at 2,472 W | 1.71 J per kH; unverified, no unit seen | vendor spec, unverified | https://github.com/monero-project/monero/issues/10270 |
| Qubic, 2026 | 23 to 34 percent of Monero's hash in withholding windows, never a sustained majority; pivot to Dogecoin ASICs 1 April 2026, Monero mining phased out | 27,000 XMR blocks, about USD 3.5 M | measured events | https://arxiv.org/abs/2512.01437v2 ; https://decrypt.co/363018 |
| Karlsen, Iron Fish, Pyrin after the FishHash forks | no ASIC since April 2024; a ColEngine P2 FPGA did 32 GH/s Karlsen at 1,484 W before the fork | FPGA only | vendor | https://pool.kryptex.com/articles/ironfish-hardfork-en |
| Conflux, Flux, Nexa, Ergo, Ravencoin, Beam, Tari, Zephyr | no dedicated chip 2024 to 2026 (Tari and Zephyr mined by the RandomX X5 class) | null | vendor | |
| GDDR7 PHY nodes and GDDR6's oldest | Cadence N3 (2024); Innosilicon FinFET; GDDR6 at 12 nm (Cadence 12FFC, Six Semiconductor); nothing at 28 nm | section 2.3's correction | claimed | as 2.3 |
| Proof of latency; era-rotating address maps | no scheme found, 2023 to 2026 | null | | |
Reading: nothing published since 2023 offers a lever outside the identity; every energy figure is a streaming pJ per bit, and no one has measured a per-random-read joule on GDDR6, GDDR7 or HBM3 at the controller, so the record's 2.0 and 1.2 nJ stand as the only model and are owed a measurement nobody can rent (the AWS F2 hour of chip-model 5.3 is still the one HBM2 instrument). The new rows change two precedents: the Ethash band's top is about 13x (claimed) not 6.8x, and the Monero chip story has a shipped-but-unverified successor to the withdrawn X9.
## 3. This lane's own ideas
### 3.1 The second sector: W = 16 words, the whole 64-byte item (KILL at stock on measured rows; the knee row owed)
**The construction** is the read-width branch's `w64` class (`docs/plans/read-width.md` section 1 and 1.5, the fold of its section 1): a load reads the 64-byte-aligned item and folds its 16 words into the destination register with a rotate-multiply between words (`verify::fold_words`, mirrored in Metal, CUDA and OpenCL: four `uint4` loads from the aligned line, then fifteen rotl-mul-xor steps), so no function of the line alone replaces the dataset and the dependent chain is unchanged. Built 5 October; the pack in `proto-cuda/packs-readwidth/w64/`; three vector units and the cache and dataset checks passed on Metal, Apple OpenCL, the RTX 5090 (NVRTC) and the RX 9070 XT; the 2^24 fingerprint `836e56e7d496e980` equal on all four; the acceptance rule's rejection 0 to 14 of 60 as version 2; the CPU verifier 0.630 ms per 32 hashes against 0.604 (+4 percent). Two things were never priced before today: the chip side, and the card side at any occupancy but the exported one, with watts.
**The chip side, on the model's own inputs** (chip-model-v3 5.3: a random 32-byte read is 909 pJ of activation plus 256 bits at 4.5 pJ per bit of movement and I/O; HBM3 the same shape at O'Connor's 3.48 pJ per bit scaled by 4.12 / 6.25; the SRAM die on floor lane 3's wire figure):
| Memory | Read energy at W = 1 (one 32 B sector) | At W = 16 (two sectors in the open row: one activate, two column accesses) | `E_mem` per hash, W = 1 to W = 16 | Rate | Label |
|---|---|---|---|---|---|
| GDDR7, 16 devices | 2.0 nJ | 3.2 nJ (+56 percent; 2.9 at O'Connor's 3.5 pJ per bit, +45) | 0.466 to **0.62** (+33 percent; 0.58 at the lower movement figure); the 35 W of static and controller unchanged | activate-bound at 166 MH/s, unchanged; 1.36 TB/s of the board's 1.79 | modelled |
| HBM3, one stack | 1.2 nJ | 1.8 nJ (+50 percent) | 0.321 to **0.40** (+24 percent) | 83.6 MH/s, unchanged | modelled |
| HBM3, eight stacks | 1.2 | 1.8 | 0.262 to 0.33 (+26 percent) | unchanged | modelled |
| The custom HBM4E base die (lane B) | 0.9 to 1.0 | about 1.5 | 0.18 to about 0.27 (+50 percent) | unchanged | modelled, approximate |
| The N2 SRAM full store (floor lane 3's rows) | 0.25 nJ (80 bits of wire) | 0.88 nJ (560 bits) | 0.036 to 0.126 (+250 percent): 66x to 19x against the 5090's stock point | power-bound: 8.3 to 2.4 GH/s per die | modelled (lane 3) |
**The card side, measured today** (8 October 2026, 13:2x to 13:4x UK; four cards on five pods; RunPod one-shot pods under the fleet's `oneshot.py`, the image `nvidia/cuda:12.8.1-devel-ubuntu24.04`; the worker built on each pod from this branch's `proto-cuda/nvrtc/worker.cpp` with `g++ -O2 -std=c++17 ... -ldl` through its Linux `dlopen` path, sha256 prefix 30eea48b; stock clocks, no power cap change; the packs `w4` (the lottery hash), `w64` (the item) and `w64-l2` (the `w64` kernel text with the first `uint4` load of each item replaced by inline PTX `ld.global.L2::64B.v4.u32`, so the L2 fetches the 64-byte line as one request; nothing else changed); every row's self-test PASS (96 of 96 vector lanes) and the 2^24 fingerprint `836e56e7d496e980` on every `w64` and `w64-l2` row, `25f96e7dce90bd4e` on every `w4` row, so the hinted text is bit-exact; rows of 2^24 nonces x 60 batches with `nvidia-smi` at 1 Hz over the row, the watts the mean of the upper half of the row's samples; a first pass of 5-batch rows for the shape sweep, a second of 60-batch rows for the watts; the sparse shapes are the worker's `sp<N>-w<W>` persistent grids, N = the card's SM count, W warps per block, so lanes in flight = N x W x 32):
| Card (architecture, memory) | `w4`: MH/s at W | `w64` exported, full occupancy | `w64` best sparse shape | `w64-l2` best form | On the 95 percent line | Energy per hash, `w4` to the best `w64` form | Label |
|---|---|---|---|---|---|---|---|
| RTX 4090 (Ada, GDDR6X, 128 SMs; driver 595.71) | 64.73 at 225 W (3.48 microjoules); idle 29 W | 33.75 (-48 percent) at 242 W | sp128-w1 (4,096 lanes) 53.98 (-17) at 258 W | **full occupancy 63.65 (-1.7 percent) at 296 W**; sp128-w1 58.5 (-10) at 264 | **PASS on rate** | 3.48 to 4.66: **+34 percent** | measured |
| H100 SXM (Hopper, HBM3, 132 SMs; driver 580.126) | 253.4 at 442 W (1.74); idle 70 W | 171.4 (-32) at 469 W | sp132-w8 200.3 (-21) at 463 W | 171.4 at full occupancy, 200.3 at sp132-w8: **the hint does nothing on Hopper** | FAIL (79 percent) | 1.74 to 2.31 (best form): +33 percent; the full-occupancy form 2.74, +57 | measured |
| RTX 3090 (Ampere, GDDR6X, 82 SMs; driver 580.159) | 60.00 at 321 W (5.35) | 23.95 (-60) at 330 W | sp82-w2 42.19 (-30) at 334 W | sp82-w2 52.34 (-13) at 349 W; sp82-w4 51.0 (5-batch row) | FAIL (87 percent) | 5.35 to 6.67: +25 percent | measured |
| RTX 5090 (Blackwell, GDDR7, 170 SMs; a secure-cloud pod, driver 580.126, bus 02:00.0; the community host 216.249.100.66 answered `cuInit` 999 on three pods, all destroyed) | 142.5 to 142.7 at 311 to 319 W (2.21 microjoules; the record's 2.26 at 311 W); idle not settled on the pod | 75.3 (-47 percent) at 344 W (the 5 October row, 71.9, reproduced) | sp170-w1 (5,440 lanes) 101.5 (-29) at 343 W; sp170-w8 88.2 | full occupancy 110.4 (-22.5) at 392 W; **sp170-w4 (21,760 lanes) 145.7 (+2.3 percent) at 401 W**; sp170-w8 141.9 (-0.4) at 408 | **PASS on rate at the sparse shape** | 2.21 to 2.75: **+25 percent** | measured |
| RX 9070 XT, Apple M5 Max | 5 October: -3 percent and 0 at 64 bytes (a 64-byte line per read already) | | | not re-run | free | 0 on the memory side | measured 5 October |
The rows are the record: the 4090 and H100 pods were destroyed before their raw logs were copied home (the copy step failed quietly on a shell expansion), so their figures above are the extracted rows read off the pods in the session; the 3090's raw logs are at `~/igneum-fleet/fl5-results/3090-*` (result.log, result2.log, power.csv, power2.csv). Spend: about USD 6 of the USD 150 allowed.
**What the rows say, on the identity.** (1) The rate question: the exported form's -47 to -60 percent is a request-shape and occupancy effect, as the 5090's probe said; one PTX qualifier on the first load makes the 64-byte item one L2 request and holds the 4-byte rate on Ada at full occupancy (98 percent), while on Hopper the hint changes nothing (the H100's loss is elsewhere: its HBM3 pseudo-channel moves a 64-byte access as two bursts and its best shape is 79 percent) and on Ampere the best shape is 87 percent. (2) The energy question, which decides it: the card pays the second sector at the cost of a whole dependent read, not at the DRAM's movement term. On the 4090 the hinted form draws 71 W more at the same 64.7 MH/s: 71 W over 8.3 G reads per second is **8.6 nJ per second sector**; on the 5090 at sp170-w4, 88 W more at 145.7 MH/s: **4.7 nJ**; against the 1.15 nJ the DRAM device charges for it (the 4090's GDDR6X at 6 pJ per bit would be 1.5), so the other 7 nJ is the card's own L2, crossbar, L1 and the SMs' longer wait, the same anatomy as the record's 8.7 nJ whole-card marginal per dependent read (15.1a). In the identity's terms `F` is 128 x 8.6 nJ = 1.1 microjoules per hash on the 4090 (0.6 on the 5090) and the chip's cost for the same work is 128 x 1.15 = 0.15: **`k` about 0.13 on the 4090, 0.24 on the 5090**, under the ALU shadow's 0.3 to 0.8 and no better than the L2 hot table's. The energy per hash rises 34 percent on the 4090 and 33 percent on the H100's best form against the modelled chip's 33 (GDDR7) and 50 (HBM3): the edge at zero shadow moves 0 on Ada against the GDDR7 chip and 0.9x in the chip's favour against HBM3 on Hopper; against the SRAM die the chip's 66x to 19x becomes about 25x against the card's new energy, not 19x. (3) The one row that could change the sign is the knee: at the 1,300 MHz lock the 5090's fixed share is smaller and its fabric runs at a lower voltage, so the sector's 8.6 nJ would fall with the measured 10.9 to 8.7 nJ per read (15.1a) and the card's rise would be a larger fraction of a smaller base; the lever lives only if the card's energy per hash rises under about 20 percent there, and a rented pod cannot read it (`nvidia-smi -lgc` refused: "does not have permission to change clocks"). That row is one PC 1 run (the `w4` and `w64-l2` packs at `--block-warps 1`, 60 batches, at the lock) and is owed to the hash lane as a formality, not a gate.
**Verdict: KILL at stock on measured rows; W = 16 words stays out of class v6; the design's "never 16" stands, for the measured reason (the card's fabric price of a sector) rather than the plan's (the pins).** What is kept: the hinted one-request load as a kernel fact for Ada (a 64-byte dependent read at the 4-byte rate), which a future width decision can use; the measured 8.6 nJ per sector as the card-side figure the chip model's width rows lacked; and the per-architecture rate table above.
Cost to the honest tiers and the rule of 5 October: nothing ships, so nothing changes for any tier; had it shipped, the Ada and Blackwell tiers would have paid about a third more energy per hash for no rate, Hopper a third more for a fifth of its rate, Ampere a quarter more for an eighth of its rate, AMD and Apple nothing, the verifier 4 percent, and every DRAM chip the same third.
### 3.2 Two items per read: W = 32, 128 bytes, one NVIDIA L2 line (KILL with 3.1)
The chip at W = 32: one activate plus four column accesses, 5.5 nJ per read; 21.3 G activates x 128 B = 2.7 TB/s is over the board's 1.79, so the GDDR7 chip becomes bandwidth-bound at 14 G reads per second, 110 MH/s, `E_mem` 1.02 microjoules (+120 percent); one HBM3 stack 0.66 (+105); the SRAM die about 0.21 (11x). The card: the pins bind the 5090 at 2.3 TB/s before any fabric cost (-20 percent of rate at best), and the measured sector price of 3.1 (8.6 nJ per sector at stock) puts three more sectors at about 3.3 microjoules per hash on the card against the chip's 0.44: `k` about 0.13 again, on a card that has also lost a fifth of its rate. Dead with 3.1; no `w32` pack is needed.
### 3.3 Independent chains per hash (memory-level parallelism as the card's lever) (KILL, one probe would reopen it)
The idea: four independent dependent chains of 32 reads per hash, folded at the end, so each resident hash keeps four reads in flight and an occupancy-bound card climbs toward its memory's activate ceiling while the chip (already at the ceiling) gains nothing; an `E_card` lever. The number: the 5090 reads 17.5 to 18.2 G per second at 4 B as "the best over lanes in flight" (the 5 October probe), which means more lanes did not help, so the bind is the memory system, not the in-flight count; the bound is the ceiling's 21.3 G, +17 to +22 percent, and the expectation is 0 to 5 percent (the microbench's `l2_indep4` against `l2_chase` read +2 percent on the L2-bound pattern). Today's rows agree from the other side: on the 4090 the `w4` hash at 64.7 MH/s is 8.3 G reads per second, the probe's own ceiling (8.3 to 8.9), at every occupancy from one warp per block to the sparse grids (64.70 to 64.74). KILL; the probe that would reopen it is one `dram_indep4_1g` row in the microbench.
### 3.4 Row-straddling items (two activates per read) (KILL, dominated)
A 64-byte item laid across a row boundary costs the chip two activates: 2 x 0.909 + 2 x 1.15 = 4.1 nJ per read (+105 percent) and halves its activate-bound rate (83 MH/s), so `E_mem` = 0.94 microjoules; the card's activate-bound rate halves too, at the fabric price of 3.1 on top. Dominated by the aligned item, which is itself dead. KILL.
### 3.5 Scratch state in DRAM (KILL)
A per-lane scratch line in DRAM, read-modify-written per step, would be `k = 1` work if it reached the DRAM on both sides; it does not: 7,262 hashes in flight x 64 B is 465 KB, inside the 5090's 96 MB L2 (an L2 hit at 1.4 nJ, `k` 0.1 to 0.3 against the chip's lane SRAM), and a scratch large enough to miss the L2 (13 KB per lane) is 15 MB of SRAM on the chip (about USD 5 of N5). The 5 October scratch rows measured the card's side: -12 to -48 percent of rate on the 5090. KILL (`docs/analysis/scratch-soundness.md` and the read-width entry say the same).
### 3.6 A note on the sunk card (not a candidate)
Both sides are capex-dominated (the record's 16.1: 7x to 10x their electricity per MH/s-hour), and the chip's economic edge is capex per MH/s (USD 2.8 against 14.7). For a home miner whose card was bought for something else the capex is sunk, so that tier's cost per MH/s-hour is its electricity alone, USD 0.000084 at the knee and USD 0.05 per kWh, against the chip's all-in USD 0.00018 (capex over two years plus electricity): the sunk card is 2.2x CHEAPER per MH/s-hour than the chip until the chip's capex is amortised. The honest-denominator lane owns this reading; it is recorded here because it is the one place the identity's per-joule framing understates the honest tier.
## 4. The KILL list, with the number
| Candidate | The identity check | The number | Label |
|---|---|---|---|
| The second sector, W = 16 words (the whole item) | the card pays the sector at its own fabric's price: `k` about 0.13; energy per hash +34 percent (4090), +33 (H100 best form), +25 (3090) against the chip's +33 | 8.6 nJ per sector on a 4090 at stock; the rate holds on Ada with the hint (98 percent), fails on Hopper (79) and Ampere (87) | measured card, modelled chip |
| Two items per read, W = 32 | the same `k` on three more sectors, on a card the pins bind at -20 percent | chip 1.02 microjoules; card about 3.3 of sector work | modelled on the measured sector price |
| Tensor tiles as the shadow (int8) | `k` 0.2 to 0.4 on the class's tile, 0.4 to 0.7 on the dense tile: never above the ALU shadow's band | the M5 Max -35 percent at 1,024 tiles, -78 at 4,096 (measured); the same premium as the ALU shadow by construction | measured card, claimed chip |
| FP8, FP16, BF16 tiles; the texture interpolator; the rasteriser | not bit-exact across vendors | | spec and vendor documents |
| The hardware video decoder | bit-exact by standard, but the honest card is decoder-bound (a fixed engine count) and the verifier decodes at milliseconds per frame | engines per card two to four (approximate) against any count on a chip | approximate |
| The RT core | traversal implementation-defined on Vulkan, DXR, OptiX and Metal; the BVH opaque; `k` 0.01 to 0.2 on the public per-ray figures | 4 to 33 nJ per ray on a dedicated unit (claimed) against 290 to 750 measured board-level on an RTX 2080 | claimed and measured |
| Memory-level parallelism per dollar | no asymmetry per channel; the correction is the PHY's node (section 2.3, kept as rank 1) | `E_mem` under +0.06 microjoules | claimed, modelled |
| Proof of latency | the chain checks values; the block is 1 s against a 0.4 microsecond round trip; the farm's node is local; a sequential prefix favours the chip about 4x | 2.5 million round trips per block | measured card latency, modelled chip |
| The drawn address map | firmware; no prefetch exists for a dependent chain | 0 to the chip, 0.8 to 3.2 percent to the cards | measured spread |
| Independent chains per hash | the card's read rate sits at its probe's ceiling at every occupancy | the 4090 64.70 to 64.74 MH/s across six shapes; expected 0 to 5 percent on the 5090, bound 22 | measured |
| Row-straddling items | dominated by the aligned item, which is dead | | modelled |
| Scratch in DRAM | the card's scratch sits in L2; the chip's in SRAM | 465 KB against 96 MB | measured rows |
## 5. Consequences per tier (the standing rule of 5 October 2026)
| Tier | What this file means | What is being done |
|---|---|---|
| Home miner, one 8 GB NVIDIA card (32-byte sectors) | nothing ships; had W = 16 shipped, this tier would have paid about a third more energy per hash for no rate (Ada) and an eighth of its rate too (Ampere) | nothing; the knee row on PC 1 is a formality |
| One 12 GB card (4070, 5070) | the same | the same |
| One 16 GB card (RX 9070 XT) | nothing (its 64-byte line is free at W = 16, measured 5 October); its row against the chip is unchanged at 22.7x on class v3 | the vendor-share metric |
| One 24 or 32 GB card (5090 class) | nothing ships; had W = 16 shipped, the 5090 would have held its rate at a sparse shape and paid +88 W (+25 percent per hash) at stock; the record's 3.6x at the knee and 2.1x at `k = 1` stand | the knee row on PC 1 (the hash lane, one row, a formality) |
| Apple (M5 Max and the unified tiers) | nothing (free at 64 bytes, measured 5 October); the honest best per joule is unchanged | nothing |
| A rig | nothing ships; the two model corrections raise the chip's project floor (USD 20 M to 75 M), which is the rig's real wall, by 4x to 15x over the record's cheapest row | the chip model's rows corrected (the Counter lane) |
| A pool user | nothing changes | |
| A node operator (the verifier) | nothing changes (the +4 percent of `w64` never ships) | |
| A chip | the cheapest project is a 12 nm GDDR6 part at USD 20 M to 30 M or an N5 GDDR7 part at 50 M to 75 M, not a 28 nm one at 5 M; its `k` on tile work 0.2 to 0.7, not 0.03 to 0.3; its cost of the second sector (+33 percent) was matched by the card's | chip-model-v3 5.4, 5.6 and 5.11 and the record's 16.2 corrected (owed to the Counter lane) |
| The public claim | unchanged by this lane on the hash; the tensor `k` sentence loses "4x to 30x"; the Ethash precedent gains the 13x row; the cheapest-chip project line gains a floor | the Counter lane's texts |
## 6. Unverified and owed
- The knee row (the 5090 at the 1,300 MHz lock, `w4` against `w64-l2`, 60 batches, watts): a rented pod refuses the lock; one PC 1 row with the hash lane; the lever lives only under a 20 percent energy rise there, which the stock rows make unlikely.
- The 4090's and H100's raw logs were not copied before their pods were destroyed; the rows in 3.1 are the figures read off the pods during the session. The 5090's and the 3090's raw logs are at `~/igneum-fleet/fl5-results/5090-*` and `3090-*` (result.log, result2.log, result3.log, power*.csv).
- The watts are `nvidia-smi` board power at 1 Hz, the upper half of each row's samples averaged; idle on the rented pods read 29 W (4090) and 70 W (H100); the 3090's idle row was taken straight after a run and is not a settled idle.
- The chip side of 3.1 is the model's own data-movement term applied twice (4.5 pJ per bit, Micron, claimed, streaming; O'Connor's 3.48 for HBM2) and lane 3's wire figure for the SRAM die; no one has measured a second-sector column read's energy on GDDR7 (section 2.6: the public record has no random-read joule on any current part).
- The tensor chip figures are vendor specifications at TDP (claimed), the M4 ANE and the H800 loop measured; the RT figures are claims and one measured RTX 2080 row; the PHY node facts are vendor IP pages read today.
- The project-cost corrections use the history's figures (a 7 nm-class startup project USD 50 M to 75 M, claimed; the 12 nm figure scaled, approximate) and the mission lane's cap model.
- Nothing ran on the Mac, on a Hetzner box or on a PC for this file; the measurements ran on four RunPod one-shot pods, all destroyed on their done (the 5090 pod when its rows are in).
## 7. Sources
Internal: `docs/analysis/counter-asic-4-research.md` (sections 2, 4, 5, 8, 15.1a, 16, 20), `docs/analysis/chip-model-v3.md` (5.1, 5.3, 5.4, 5.6, 5.10 to 5.12), `docs/design/class-v6-rotating-family.md` (sections 2, 7, 7a, 7c, 10 and 10.3), `docs/analysis/class-v6/hardware-future.md` at 7618e729 (sections 0, 1, 2, 4.9, 5), `docs/analysis/class-v6/history.md` (sections 3 and 5), `docs/analysis/class-v6/invention.md` on `class-v6-invention` (sections 1 and 2), `docs/analysis/class-v6/floor/sram-and-floor.md` on `class-v6-floor-sram` as carried in section 10.3, `docs/plans/read-width.md` (sections 1, 1.5, 4.1), `docs/bench-log.md` (the 5 October read-width entry: the probes and the hash rates), `proto-cuda/packs-readwidth/w64/kernel_bound.cu` (the four `uint4` loads per item), `proto-cuda/nvrtc/worker.cpp` (`variantSource`, the `sp<N>-w<W>` shapes, the Linux `dlopen` path), the fleet's `oneshot.py` and the rented rows' logs under `~/igneum-fleet/fl5-results/`, `docs/analysis/scratch-soundness.md`, spec 01 (1.5, the item table).
External, all read 8 October 2026 by this lane's three research sub-agents (tensor, RT, literature) and labelled where carried: the NVIDIA RTX 5090 page and the Blackwell architecture whitepaper; the Ada whitepaper; arXiv 2501.12084v2 (Hopper tensor-core microbenchmarks); Lenovo LP2226 (B200); the AMD MI300X data sheet and MI355X notes; the RX 9070 XT page; Qualcomm's AI 100 Ultra brief; eeNews Europe (MTIA v2); AnandTech 21342 (Gaudi 3); the Hot Chips 2025 Ironwood slides; arXiv 2608.28048 (Trainium2); Tenstorrent's Blackhole specification; the M4 ANE measurement (maderix.substack.com); cnx-software (Hailo-8); design-reuse 52539 (Untether); hackster.io (Axelera); NVIDIA Research's VSQ accelerator (JSSC 2023); Dally, Hot Chips 2023 keynote, slides 12 and 24; the Vulkan ray traversal chapter; the DXR specification; the NVIDIA developer forum thread 309730 (14 October 2024); the Blender Cycles source note; the Mach-RT TVCG preprint; Hot Chips 31 (Turing); the AMD Hot Chips 2025 RDNA 4 slides; embedded.com (RayCore); AnandTech 7870 (GR6500); arXiv 2409.06000 (RayFlex); the TRaX CGI 2018 study; Ylitie et al., HPG 2017; Chou et al., MICRO 2023; par.nsf.gov 10156963 (Choe et al., MEMSYS 2019); arXiv 2508.06795; SemiAnalysis "The Memory Wall"; the SK hynix ISSCC 2024 GDDR7 paper (ResearchGate 378947349); arXiv 2410.12990; arXiv 2405.06081; asicminervalue (iPollo V2H, V1, A11 Pro); kryptex (Jasminer X16-Q, FishHash); whattomine (the 5090 on Etchash); monero-project issue 10270; arXiv 2512.01437v2 and decrypt.co 363018 (Qubic); Business Wire 20240925780123 (Cadence GDDR7 on N3) and 20221115005674 (GDDR6 at 12 nm); innosilicon.com GDDR7; kurnal-insights (GB202 die); TrendForce 24 September 2026 (the 2 GB GDDR7 part); arXiv 2403.13230 (BFT-PoLoc); eprint 2026/694 (unread, refused); the PTX ISA (`ld` with the `.L2::64B` prefetch-size qualifier, sm_75 and later).