Class v6 floor lane 5 (invention): the second sector measured on rented 4090, H100 and 3090 at stock and KILLED: a hinted one-request load (ld.global.L2::64B) holds the 4-byte rate on Ada (63.65 against 64.73 MH/s) and fails on Hopper (-21 percent) and Ampere (-13), but the card pays the sector at about 8.6 nJ through its own fabric (296 against 225 W on the 4090), +34 percent energy per hash against the modelled chip's +33, k about 0.13; W = 16 stays out of class v6, W = 32 dies with it; the KEEP list is now two model corrections, one measured card figure and the Ethash precedent row

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 12:33:45 +00:00
parent 43fcf0fed6
commit 68315bf83b

View file

@ -1,28 +1,29 @@
# Class v6 floor lane 5, invention: what none of the four floor lanes covers, checked against the identity
8 October 2026, 12:48 to 19:xx UK, branch `class-v6-floor-invention` from `counter-asic-4` (base 171618c5 after the fast-forward; first cut at 13:0x UK, corrected 13:3x UK for the read-width naming: `w16` is 16 bytes, `w64` is the 64-byte item). The open-ended lane of the floor research: four other lanes work the known levers (SM-sparse gating, the shadow's `k` from RTL synthesis, the SRAM full-store chip and the dataset floor, the honest denominator per tier); this file finds what they do not cover. Every candidate is checked against the identity first and killed or kept with a number. Every number carries a label: **measured** (a card on a named job, the file named), **modelled** (arithmetic on the chip model's cited inputs, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling). Nothing here changes a consensus object, a pack or a served number. Research and reasoning ran on this Mac; no build, no benchmark and no miner ran anywhere for this file.
8 October 2026, 12:48 to 19:xx UK, branch `class-v6-floor-invention` from `counter-asic-4` (base 171618c5 after the fast-forward; first cut at 13:0x UK; corrected 13:1x UK for the read-width naming: `w16` is 16 bytes, `w64` is the 64-byte item; the rented-card rows of 13:2x to 13:4x UK carried in section 3.1, which reverse rank 1 at stock). The open-ended lane of the floor research: four other lanes work the known levers (SM-sparse gating, the shadow's `k` from RTL synthesis, the SRAM full-store chip and the dataset floor, the honest denominator per tier); this file finds what they do not cover. Every candidate is checked against the identity first and killed or kept with a number. Every number carries a label: **measured** (a card on a named job, the file named), **modelled** (arithmetic on the chip model's cited inputs, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling). Nothing here changes a consensus object, a pack or a served number. Research and reasoning ran on this Mac; no build, no benchmark and no miner ran anywhere for this file.
## 0. One page
**The identity, as the record states it** (`docs/analysis/counter-asic-4-research.md` section 2): against the chip that stores the dataset, `edge = (E_card + F) / (E_mem + k F)`; at zero premium the edge is `E_card / E_mem` (3.6x on a 5090 at its 1,300 MHz knee against the GDDR7 board, measured card, modelled chip), and a shadow op with `k` under 1 never fully closes it. The one thing with `k = 1` by construction is the memory itself: the same devices on both sides. The record forced one DRAM operation per read (the activate and one 32-byte sector) and called every wider read "nothing to force: the chip already pays the sector". That is true up to 32 bytes and false at 64.
**The identity, as the record states it** (`docs/analysis/counter-asic-4-research.md` section 2): against the chip that stores the dataset, `edge = (E_card + F) / (E_mem + k F)`; at zero premium the edge is `E_card / E_mem` (3.6x on a 5090 at its 1,300 MHz knee against the GDDR7 board, measured card, modelled chip), and a shadow op with `k` under 1 never fully closes it. The one thing with `k = 1` by construction is the memory itself: the same devices on both sides. The record forced one DRAM operation per read (the activate and one 32-byte sector) and called every wider read "nothing to force: the chip already pays the sector". That is true up to 32 bytes; at 64 bytes the chip pays a second sector, and this lane's question was whether the honest card pays it at the DRAM's price or at its own.
**The one lever this lane finds: the second sector, which is the whole item.** The dataset item is 64 bytes (16 words; spec 1.5 and the item table). The hash reads one word of it, which moves one 32-byte sector on the 5090 and on the chip's GDDR7; the read-width band of class v6 (W in {1, 4} words, 4 to 16 bytes) stays inside that sector, so a DRAM chip pays nothing more at any width in the band, as the record said. Reading the whole item (W = 16 words, the read-width branch's class `w64`, built and measured 5 October, bit-exact on four runtimes, fingerprint `836e56e7d496e980`, the verifier 0.630 against 0.604 ms) moves two sectors in one open row: on the chip model's own inputs that is the data-movement term paid twice (4.5 pJ per bit x 256 bits = 1.15 nJ, Micron, claimed) on one activate (909 pJ), so the GDDR7 chip's energy per read rises from 2.0 to 3.2 nJ (+56 percent, modelled), its energy per hash from 0.466 to 0.62 microjoules (+33 percent; static, controller and the activate-bound rate unchanged), one HBM3 stack's from 0.321 to 0.40 (+24 percent), and, on floor lane 3's wire figure, the N2 SRAM full store's from 0.036 to 0.126 microjoules (66x to 19x against the 5090's stock point, that lane's own rows): **the second sector is the one lever that moves every chip in the model at `k` near 1, and the SRAM die most of all.** What decides its worth is the honest card's rate at 64 bytes, which is measured in two places that disagree: the hash at full occupancy ran at 71.9 MH/s on the 5090 (-47 percent, the row the design calls dead) while the same card's 64-byte dependent-read probe reached 15.7 G reads per second at its best lane count (-10 to -14 percent against its 4-byte probe), both 5 October; the 589 GB/s the hash moved is a third of the stream, so the bind at full occupancy is the request path (two serialised sector misses, a second activate at the controller, or queueing), not the pins. The RX 9070 XT (-3 percent) and the M5 Max (0) pay nothing at 64 bytes (measured). Three outcomes, priced at the 5090's knee against the GDDR7 chip (modelled): the exported form at full occupancy, 5.3x (dead: the card halves, the chip does not); the probe's best lane count, 3.4x (not worth a class change); a one-request form that holds the 4-byte rate within 5 percent (PTX `ld.global.L2::64B` on the first word, or the lane count tuned by the race), 3.0x at zero shadow and 2.7x with the class v4 shadow at `k = 0.5`, with every chip column down 10 to 15 percent and the SRAM die's down 3.5x. **KEEP, rank 1, as the one measurement this lane asks for before the 20:00 close's v6 word**: the `w64` pack on PC 1's 5090, an occupancy sweep and the `.L2::64B` variant, watts at 1 Hz, unlocked and at the 1,300 lock; pass line 95 percent of the 4-byte rate (then W = 16 words becomes layer 1's floor in v6), kill line under 95 percent (then the band stays {1, 4} and this row closes).
**The candidate this lane found, measured today, and where it stands.** The dataset item is 64 bytes (16 words; spec 1.5). Reading the whole item (W = 16 words, the read-width branch's class `w64`, bit-exact on four runtimes, fingerprint `836e56e7d496e980`, the verifier +4 percent) forces a second 32-byte sector per read in the same open row: on the chip model's own inputs (909 pJ of activation plus 1,150 pJ of movement per sector, Micron's 4.5 pJ per bit, claimed) the GDDR7 chip's energy per hash rises from 0.466 to 0.62 microjoules (+33 percent, its rate unchanged), one HBM3 stack's from 0.321 to 0.40 (+24), and on floor lane 3's wire figure the N2 SRAM die's from 0.036 to 0.126 (66x to 19x). The design's "never 16 words" rested on one row, the 5090's 71.9 MH/s (-47 percent) in the exported kernel at full occupancy, while the same card's own 64-byte probe read -13 percent; so the lane rented cards and measured the card side at stock (section 3.1): **the rate question is settled per architecture, and the energy question kills the lever at stock.** A one-request form of the same kernel text (PTX `ld.global.L2::64B` on the first load of each item; bit-exact, the same fingerprint) holds the 4-byte rate on the RTX 4090 at full occupancy (63.65 against 64.73 MH/s, 98 percent, PASS); on the H100 the best form is -21 percent and the hint does nothing (FAIL); on the RTX 3090 the best form is -13 percent (FAIL). But the card pays the second sector at about 8.6 nJ (the 4090: +71 W at 8.3 G reads per second for no rate), the cost of a whole dependent read through its L2 and crossbar, not the DRAM's 1.15 nJ: the 4090's energy per hash rises 34 percent, the H100's 33 percent in its best form, the 3090's 25 percent, against the modelled chip's 33 (GDDR7) and 50 (HBM3). The sector's `k` is about 0.13 on GDDR7 (1.15 over 8.6), under the ALU shadow's 0.3 to 0.8, and the edge at zero shadow moves 0 on Ada and 0.9x in the chip's favour on Hopper. **KILL at stock on measured rows.** The one row that could revive it is the honest card's knee (a smaller fixed share and a lower fabric voltage; a rented pod refuses `-lgc`), owed as a single PC 1 row and not blocking anything: W = 16 stays out of class v6.
**Everything else on the brief is killed, with the number.** The tensor core: the record's `k` 0.03 to 0.3 is not defensible at its floor (the 0.04 pJ figure is INT4 at 0.46 V on a test chip with no memory system); shipping merchant silicon at nominal voltage reads 0.30 to 0.56 pJ per INT8 MAC (Meta MTIA v2 0.51, Qualcomm AI 100 Ultra 0.34, AMD MI355X 0.56, NVIDIA B200 0.44, Apple M4 ANE 0.30 measured; all claimed unless marked), and a Hopper tensor core under a pure MMA loop measured 0.34, so the honest band against the tile the class runs (1.5 pJ per MAC at the knee, measured) is `k` 0.2 to 0.4 and against the dense wide tile (0.83 at the knee) 0.4 to 0.7: equal to or under the ALU shadow's 0.3 to 0.8, never above it, and the Apple cost (-35 percent of rate at 1,024 tiles per hash, measured) kills it as content regardless. The RT core: traversal results are implementation-defined on every API (Vulkan "no ordering guarantee", DXR "no defined order", NVIDIA "could change depending on which driver, which GPU"), the acceleration structure is opaque, and a dedicated unit reads 4 to 33 nJ per ray against 290 to 750 nJ measured board-level on an RTX 2080 (`k` 0.01 to 0.2): dead on bit-exactness and on `k`. Memory-level parallelism per dollar: no asymmetry per channel, but a correction to the chip model: no 28 nm GDDR7 PHY exists (the shipped ones are N3 and FinFET; the oldest GDDR6 PHYs are 12 nm), so the record's cheapest chip (a USD 5 M project, a USD 17 M break-even cap) is not buildable; the floor is a 12 nm GDDR6 part at about USD 20 M to 30 M (cap USD 70 M to 100 M) or an N5-class GDDR7 part at USD 50 M to 75 M (cap 170 M to 250 M), modelled. Proof of latency: the chain checks values, never time; the block is 1 s and a DRAM round trip 0.1 to 0.4 microseconds; the farm's node is on its own LAN; a sequential per-block prefix favours the lower-latency side, which is the chip (about 4x). The time dimension on the address map: a bit permutation is firmware (a few hundred gates), no prefetch exists for a dependent chain, and the literature has no such scheme. The literature since 2023 holds nothing outside the identity; what it adds is a higher Ethash precedent (iPollo V2H, 0.14 J per MH, about 13x a 5090, vendor claim), the measured PAM3 I/O energies (SK hynix, ISSCC 2024) and the PHY-node fact above. Of this lane's own ideas, four more die on the identity (video decode, row-straddling, scratch in DRAM, independent chains per hash) and one is held behind rank 1 for v7 (two items per read, W = 32).
**Everything else on the brief is killed, with the number.** The tensor core: the record's `k` 0.03 to 0.3 is not defensible at its floor (the 0.04 pJ figure is INT4 at 0.46 V on a test chip with no memory system); shipping merchant silicon at nominal voltage reads 0.30 to 0.56 pJ per INT8 MAC (Meta MTIA v2 0.51, Qualcomm AI 100 Ultra 0.34, AMD MI355X 0.56, NVIDIA B200 0.44, Apple M4 ANE 0.30 measured; all claimed unless marked), and a Hopper tensor core under a pure MMA loop measured 0.34, so the honest band against the tile the class runs (1.5 pJ per MAC at the knee, measured) is `k` 0.2 to 0.4 and against the dense wide tile (0.83 at the knee) 0.4 to 0.7: equal to or under the ALU shadow's 0.3 to 0.8, never above it, and the Apple cost (-35 percent of rate at 1,024 tiles per hash, measured) kills it as content regardless. The RT core: traversal results are implementation-defined on every API (Vulkan "no ordering guarantee", DXR "no defined order", NVIDIA "could change depending on which driver, which GPU"), the acceleration structure is opaque, and a dedicated unit reads 4 to 33 nJ per ray against 290 to 750 nJ measured board-level on an RTX 2080 (`k` 0.01 to 0.2): dead on bit-exactness and on `k`. Memory-level parallelism per dollar: no asymmetry per channel, but a correction to the chip model: no 28 nm GDDR7 PHY exists (the shipped ones are N3 and FinFET; the oldest GDDR6 PHYs are 12 nm), so the record's cheapest chip (a USD 5 M project, a USD 17 M break-even cap) is not buildable; the floor is a 12 nm GDDR6 part at about USD 20 M to 30 M (cap USD 70 M to 100 M) or an N5-class GDDR7 part at USD 50 M to 75 M (cap 170 M to 250 M), modelled. Proof of latency: the chain checks values, never time; the block is 1 s and a DRAM round trip 0.1 to 0.4 microseconds; the farm's node is on its own LAN; a sequential per-block prefix favours the lower-latency side, which is the chip (about 4x). The time dimension on the address map: a bit permutation is firmware (a few hundred gates), no prefetch exists for a dependent chain, and the literature has no such scheme. The literature since 2023 holds nothing outside the identity; what it adds is a higher Ethash precedent (iPollo V2H, 0.14 J per MH, about 13x a 5090, vendor claim), the measured PAM3 I/O energies (SK hynix, ISSCC 2024) and the PHY-node fact above. Of this lane's own ideas, four more die on the identity (video decode, row-straddling, scratch in DRAM, independent chains per hash), and two items per read (W = 32) dies with W = 16.
### The KEEP list, ranked
| Rank | Item | Chip edge at the 5090's knee, GDDR7 (zero shadow / with class v4 at k 0.3 / 0.5 / 1) | Cost to the honest tiers | Verifier | Where | Label |
|---|---|---|---|---|---|---|
| 1 | **The second sector: W = 16 words, the whole 64-byte item (class `w64`)**, conditional on one PC 1 job | if the 5090 holds 95 percent of its 4-byte rate: 3.0 to 3.2x / 3.1x / 2.7x / 2.0x (from 3.6 / 3.5 / 2.9 / 2.1); at the probe's -13 percent: 3.4x at zero shadow; at the exported form's -47 percent: 5.3x (dead) | 5090 about +10 percent energy at the pass line (+28 W at the knee, modelled) plus whatever rate it loses; 9070 XT -3 percent and M5 Max 0 (measured); the SRAM die 66x to 19x and the GDDR7 chip +33 percent whatever the card does | +4 percent (0.630 against 0.604 ms, measured) | the job first; then class v6's layer-1 width row (the floor W = 16) if it passes | card measured at two occupancies (unlocked, no watts); chip modelled |
| 2 | **The chip model's project floor corrected: no 28 nm GDDR7 PHY** | unchanged | none | none | chip-model-v3 5.4 and 5.6 and the record's 16.2: the cheapest row's project USD 20 M to 75 M, cap USD 70 M to 250 M | claimed (vendor IP pages), modelled (the mission lane's cap) |
| 3 | **The tensor `k` band corrected** (0.2 to 0.4 on the class's tile, 0.4 to 0.7 on the dense tile; not 0.03 to 0.3) | no change to any served row (the tile stays a worse-or-equal lever) | none | none | a wording correction to the record's 15.1a and 20.4 and chip-model 5.11 | claimed and measured |
| 4 | **Two items per read, W = 32 (128 B)**, only if rank 1 passes | the chip bandwidth-bound at 110 MH/s and 1.02 microjoules; 2.4x at zero shadow if the card held its rate, which at 128 B it will not | the 5090 at least -20 percent (the pins bind at 2.3 TB/s), likely far more by the 64-byte row | about +4 percent | v7 at the earliest, after a `w32` export and its rows | modelled |
| 5 | **The Ethash precedent's top row** (iPollo V2H, 3.4 GH/s at 475 W, about 13x a 5090 per joule) | the precedent band 2.1x to 13x, not 2.1x to 6.8x | none | none | `asic-resistance-history.md` and chip-model 5.1 | claimed |
| Rank | Item | What it moves | Cost to the honest tiers | Where | Label |
|---|---|---|---|---|---|
| 1 | **The chip model's project floor corrected: no 28 nm GDDR7 PHY** | the cheapest buildable chip's project USD 20 M to 75 M (cap USD 70 M to 250 M), not USD 5 M (cap 17 M); the N5-core row's cap 170 M to 250 M, not 100 M | none | chip-model-v3 5.4 and 5.6 and the record's 16.2 | claimed (vendor IP pages), modelled (the mission lane's cap) |
| 2 | **The tensor `k` band corrected** (0.2 to 0.4 on the class's tile, 0.4 to 0.7 on the dense tile; not 0.03 to 0.3) | no served row (the tile stays a worse-or-equal lever) | none | the record's 15.1a and 20.4 and chip-model 5.11 | claimed and measured |
| 3 | **The card's cost of a second sector, measured: about 8.6 nJ on a 4090 at stock** | closes the read-width question on the identity's own terms (`k` about 0.13 on the sector); the hinted one-request load form is kept as a miner-side kernel fact for Ada (98 percent of the 4-byte rate at 64 bytes) | none (nothing ships) | the read-width plan's 4.1 and lane 3's width table; the knee row owed on PC 1 | measured |
| 4 | **The Ethash precedent's top row** (iPollo V2H, 3.4 GH/s at 475 W, about 13x a 5090 per joule) | the precedent band 2.1x to 13x, not 2.1x to 6.8x | none | `asic-resistance-history.md` and chip-model 5.1 | claimed |
### The five-sentence reading for the 20:00 close
No item on this list is a class v6 change; nothing here reads `k` at or above 1 on a measured GPU figure.
The identity leaves one piece of work a chip pays at the card's own price, the memory's data movement, and the record forced one sector of it per read; reading the whole 64-byte item (W = 16 words, the measured `w64` class) forces the second sector at `k` near 1, costs AMD and Apple nothing (measured), costs the verifier 4 percent, raises the GDDR7 chip's energy per hash by a third and the N2 SRAM die's by 3.5x (modelled), and is dead or alive on one number nobody has yet read: the 5090's rate at 64 bytes in a one-request form at its best occupancy, which the 5 October probe put at -13 percent and the exported hash at -47 percent. That job (the `w64` pack, an occupancy sweep, the `.L2::64B` variant, watts at both clocks) is the one measurement this lane asks for, with its pass line at 95 percent of the 4-byte rate; on a pass W = 16 words becomes layer 1's width floor in class v6 (3.0x at zero shadow against the GDDR7 chip, 2.7x with the shadow at `k = 0.5`, the SRAM die 19x instead of 66x), and on a fail the band stays {1, 4} and this row closes. The tensor core, the RT core, proof of latency and the drawn address map are dead on the identity or on bit-exactness, with the tensor `k` band corrected upward to 0.2 to 0.7 (still never above the ALU shadow's) and the RT core's traversal implementation-defined on every vendor. The chip model's cheapest chip does not exist as priced: a GDDR7 PHY needs a FinFET node (N3 shipped; 12 nm is the oldest GDDR6 PHY), so the floor project is USD 20 M to 75 M and the floor break-even cap USD 70 M to 250 M instead of 17 M, which goes into the model now. Nothing in the literature since 2023 offers a lever outside the identity, and two items per read (W = 32) waits for v7 behind rank 1's result.
### The five-sentence reading for the 15:45 close
The identity leaves one piece of work a chip pays at the card's own price, the memory's data movement, and this lane's candidate was the second sector of the 64-byte item (W = 16 words); measured today on rented cards at stock, a one-request form holds the 4-byte rate on Ada (98 percent on a 4090) and fails on Hopper and Ampere (-21 and -13 percent at best), but the card pays that sector at about 8.6 nJ through its own fabric against the DRAM's 1.15, so its energy per hash rises 25 to 34 percent, the same as the modelled chip's 33, and the lever reads `k` about 0.13: dead at stock, W = 16 stays out of class v6, with the knee row owed on PC 1 as a formality. The tensor core, the RT core, proof of latency and the drawn address map are dead on the identity or on bit-exactness, with the tensor `k` band corrected upward to 0.2 to 0.7 (still never above the ALU shadow's) and the RT core's traversal implementation-defined on every vendor. The chip model's cheapest chip does not exist as priced: a GDDR7 PHY needs a FinFET node (N3 shipped; 12 nm is the oldest GDDR6 PHY), so the floor project is USD 20 M to 75 M and the floor break-even cap USD 70 M to 250 M instead of 17 M, which goes into the model now. Nothing in the literature since 2023 offers a lever outside the identity, and this lane's own four other ideas die on it with numbers. The floor section's default for this lane therefore stands as written: no new lever; what the lane adds is two corrections to the model, one measured card-side figure, and the measured closing of the read-width question.
## 1. The frame
@ -179,72 +180,49 @@ Reading: nothing published since 2023 offers a lever outside the identity; every
## 3. This lane's own ideas
### 3.1 The second sector: W = 16 words, the whole 64-byte item (KEEP, rank 1, conditional on one measurement)
### 3.1 The second sector: W = 16 words, the whole 64-byte item (KILL at stock on measured rows; the knee row owed)
**The construction** is the read-width branch's `w64` class (`docs/plans/read-width.md` section 1 and 1.5, the fold of its section 1): a load reads the 64-byte-aligned item and folds its 16 words into the destination register with a rotate-multiply between words (`verify::fold_words`, mirrored in Metal, CUDA and OpenCL: four `uint4` loads from the aligned line, then fifteen rotl-mul-xor steps), so no function of the line alone replaces the dataset and the dependent chain is unchanged. Built 5 October; the pack in `proto-cuda/packs-readwidth/w64/`; three vector units and the cache and dataset checks passed on Metal, Apple OpenCL, the RTX 5090 (NVRTC) and the RX 9070 XT; the 2^24 fingerprint `836e56e7d496e980` equal on all four; the acceptance rule's rejection 0 to 14 of 60 as version 2; the CPU verifier 0.630 ms per 32 hashes against 0.604 (+4 percent; a lane's words lie in one item, derived whole either way). Two things were never priced: the chip side, and the card side at any occupancy but the exported one.
**The construction** is the read-width branch's `w64` class (`docs/plans/read-width.md` section 1 and 1.5, the fold of its section 1): a load reads the 64-byte-aligned item and folds its 16 words into the destination register with a rotate-multiply between words (`verify::fold_words`, mirrored in Metal, CUDA and OpenCL: four `uint4` loads from the aligned line, then fifteen rotl-mul-xor steps), so no function of the line alone replaces the dataset and the dependent chain is unchanged. Built 5 October; the pack in `proto-cuda/packs-readwidth/w64/`; three vector units and the cache and dataset checks passed on Metal, Apple OpenCL, the RTX 5090 (NVRTC) and the RX 9070 XT; the 2^24 fingerprint `836e56e7d496e980` equal on all four; the acceptance rule's rejection 0 to 14 of 60 as version 2; the CPU verifier 0.630 ms per 32 hashes against 0.604 (+4 percent). Two things were never priced before today: the chip side, and the card side at any occupancy but the exported one, with watts.
**The chip side, on the model's own inputs** (chip-model-v3 5.3: a random 32-byte read is 909 pJ of activation plus 256 bits at 4.5 pJ per bit of movement and I/O; HBM3 the same shape at O'Connor's 3.48 pJ per bit scaled by 4.12 / 6.25; the SRAM die on floor lane 3's wire figure):
| Memory | Read energy at W = 1 (one 32 B sector) | At W = 16 (two sectors in the open row: one activate, two column accesses) | `E_mem` per hash, W = 1 to W = 16 | Rate | Label |
|---|---|---|---|---|---|
| GDDR7, 16 devices | 2.0 nJ | 3.2 nJ (+56 percent; 2.9 at O'Connor's 3.5 pJ per bit, +45) | 0.466 to **0.62** (+33 percent; 0.58 at the lower movement figure); the 35 W of static and controller unchanged | activate-bound at 166 MH/s, unchanged; 1.36 TB/s of the board's 1.79 | modelled |
| HBM3, one stack | 1.2 nJ | 1.8 nJ (+50 percent) | 0.321 to **0.40** (+24 percent) | 83.6 MH/s, unchanged; 0.68 TB/s of 0.82 | modelled |
| HBM3, one stack | 1.2 nJ | 1.8 nJ (+50 percent) | 0.321 to **0.40** (+24 percent) | 83.6 MH/s, unchanged | modelled |
| HBM3, eight stacks | 1.2 | 1.8 | 0.262 to 0.33 (+26 percent) | unchanged | modelled |
| The custom HBM4E base die (lane B) | 0.9 to 1.0 | about 1.5 | 0.18 to about 0.27 (+50 percent) | unchanged | modelled, approximate |
| **The N2 SRAM full store (floor lane 3's rows)** | 0.25 nJ (80 bits of wire) | 0.88 nJ (560 bits) | **0.036 to 0.126 (+250 percent): 66x to 19x against the 5090's stock point** | power-bound: 8.3 to 2.4 GH/s per die | modelled (lane 3) |
| The N2 SRAM full store (floor lane 3's rows) | 0.25 nJ (80 bits of wire) | 0.88 nJ (560 bits) | 0.036 to 0.126 (+250 percent): 66x to 19x against the 5090's stock point | power-bound: 8.3 to 2.4 GH/s per die | modelled (lane 3) |
A chip's controller merges the two sectors into one row cycle because it was built for this hash; a controller that did not (a second activate per sector) would pay 4.1 nJ and halve its rate, which no maker ships. The 5090's controller is the open question below.
**The card side, measured today** (8 October 2026, 13:2x to 13:4x UK; RunPod one-shot pods under the fleet's `oneshot.py`, the image `nvidia/cuda:12.8.1-devel-ubuntu24.04`; the worker built on each pod from this branch's `proto-cuda/nvrtc/worker.cpp` with `g++ -O2 -std=c++17 ... -ldl` through its Linux `dlopen` path, sha256 prefix 30eea48b; stock clocks, no power cap change; the packs `w4` (the lottery hash), `w64` (the item) and `w64-l2` (the `w64` kernel text with the first `uint4` load of each item replaced by inline PTX `ld.global.L2::64B.v4.u32`, so the L2 fetches the 64-byte line as one request; nothing else changed); every row's self-test PASS (96 of 96 vector lanes) and the 2^24 fingerprint `836e56e7d496e980` on every `w64` and `w64-l2` row, `25f96e7dce90bd4e` on every `w4` row, so the hinted text is bit-exact; rows of 2^24 nonces x 60 batches with `nvidia-smi` at 1 Hz over the row, the watts the mean of the upper half of the row's samples; a first pass of 5-batch rows for the shape sweep, a second of 60-batch rows for the watts; the sparse shapes are the worker's `sp<N>-w<W>` persistent grids, N = the card's SM count, W warps per block, so lanes in flight = N x W x 32):
**The card side, measured at two occupancies and unmeasured at the one that matters:**
| Card (architecture, memory) | `w4`: MH/s at W | `w64` exported, full occupancy | `w64` best sparse shape | `w64-l2` best form | On the 95 percent line | Energy per hash, `w4` to the best `w64` form | Label |
|---|---|---|---|---|---|---|---|
| RTX 4090 (Ada, GDDR6X, 128 SMs; driver 595.71) | 64.73 at 225 W (3.48 microjoules); idle 29 W | 33.75 (-48 percent) at 242 W | sp128-w1 (4,096 lanes) 53.98 (-17) at 258 W | **full occupancy 63.65 (-1.7 percent) at 296 W**; sp128-w1 58.5 (-10) at 264 | **PASS on rate** | 3.48 to 4.66: **+34 percent** | measured |
| H100 SXM (Hopper, HBM3, 132 SMs; driver 580.126) | 253.4 at 442 W (1.74); idle 70 W | 171.4 (-32) at 469 W | sp132-w8 200.3 (-21) at 463 W | 171.4 at full occupancy, 200.3 at sp132-w8: **the hint does nothing on Hopper** | FAIL (79 percent) | 1.74 to 2.31 (best form): +33 percent; the full-occupancy form 2.74, +57 | measured |
| RTX 3090 (Ampere, GDDR6X, 82 SMs; driver 580.159) | 60.00 at 321 W (5.35) | 23.95 (-60) at 330 W | sp82-w2 42.19 (-30) at 334 W | sp82-w2 52.34 (-13) at 349 W; sp82-w4 51.0 (5-batch row) | FAIL (87 percent) | 5.35 to 6.67: +25 percent | measured |
| RTX 5090 (Blackwell, GDDR7, 170 SMs) | the 5 October rows: 136.1 MH/s; `w64` 71.9 (-47) at full occupancy; the probe -13 at its best lane count | | | pending: RunPod's one community 5090 host (216.249.100.66, bus 85:00.0, driver 590.48) answered `cuInit` 999 on three pods in a row (destroyed); a secure-cloud pod is the fourth try, its row amended here when it lands | | | measured 5 October (rate only) |
| RX 9070 XT, Apple M5 Max | 5 October: -3 percent and 0 at 64 bytes (a 64-byte line per read already) | | | not re-run | free | 0 on the memory side | measured 5 October |
| Card | At W = 16 (`w64`), unlocked, 5 October | What it says | Label |
|---|---|---|---|
| RTX 5090, the hash at full occupancy (the exported kernel, one warp per block, the race's default) | 71.9 MH/s against 136.1 (**-47 percent**); 9.2 G reads per second, 589 GB/s of sectors | the row the read-width plan and class v6's layer 1 call "bandwidth-bound" and dead; but 589 GB/s is a third of the 1.79 TB/s stream, so the bind is the request path: two serialised sector misses per item, or a second activate at the controller when the second sector's request finds the row already closed (close-page on a random stream), or queueing at 4,080 resident warps | measured rate; the cause unmeasured |
| RTX 5090, the 64-byte dependent-read probe, best over lanes in flight | **15.7 G reads per second** (9.1 at 4 M lanes) against 17.5 to 18.2 at 4 B: **-10 to -14 percent** | the same memory system, at its best lane count, does 64-byte dependent reads at 1.0 TB/s and within 14 percent of its 4-byte rate; the hash's -47 percent is therefore an occupancy and request-shape effect, not the pins | measured |
| RTX 5090, a one-request form (PTX `ld.global.L2::64B` prefetch qualifier on the first word, sm_75 and later, so the L2 fetches the whole line as one transaction and the next three loads hit; or `cp.async.bulk` on sm_90 and later) | unmeasured | if the request reaches the controller as one 64-byte access the card pays one activate and two column reads as the chip does; the 4-byte rate would then hold (17.9 G x 64 B = 1.15 TB/s, 64 percent of the stream) | unmeasured; the worker's `variantSource` mechanism rewrites kernel text by anchor, as the hot-table `ldcs` and SM-sparse variants did |
| RX 9070 XT | 17.59 against 18.15 MH/s (**-3 percent**); every 4-byte read already fetches a 64-byte line (the probe: 2.47 to 2.87 G lines per second at every width) | free | measured |
| Apple M5 Max | 28.27 against 27.74 (**0**); 64 B at the 4 B rate (3.51 G per second at both, Apple OpenCL probe) | free | measured |
| RTX 5070 Ti, 4070, 4060 Ti (32-byte sectors, the same controller family) | as the 5090, whichever row it turns out to be | | approximate |
The rows are the record: the 4090 and H100 pods were destroyed before their raw logs were copied home (the copy step failed quietly on a shell expansion), so their figures above are the extracted rows read off the pods in the session; the 3090's raw logs are at `~/igneum-fleet/fl5-results/3090-*` (result.log, result2.log, power.csv, power2.csv). Spend: about USD 6 of the USD 150 allowed.
**The three outcomes, priced** (the 5090 at the 1,300 knee, class v3 213.0 W at 127.3 MH/s; the card's DRAM joules for the second sector +1.15 nJ per read, its fabric +0.3 to 0.5 nJ per sector, approximate; the GDDR7 chip at 0.62):
**What the rows say, on the identity.** (1) The rate question: the exported form's -47 to -60 percent is a request-shape and occupancy effect, as the 5090's probe said; one PTX qualifier on the first load makes the 64-byte item one L2 request and holds the 4-byte rate on Ada at full occupancy (98 percent), while on Hopper the hint changes nothing (the H100's loss is elsewhere: its HBM3 pseudo-channel moves a 64-byte access as two bursts and its best shape is 79 percent) and on Ampere the best shape is 87 percent. (2) The energy question, which decides it: the card pays the second sector at the cost of a whole dependent read, not at the DRAM's movement term. On the 4090 the hinted form draws 71 W more at the same 64.7 MH/s: 71 W over 8.3 G reads per second is **8.6 nJ per second sector**, against the 1.15 nJ the DRAM device charges for it (the 4090's GDDR6X at 6 pJ per bit would be 1.5), so the other 7 nJ is the card's own L2, crossbar, L1 and the SMs' longer wait, the same anatomy as the record's 8.7 nJ whole-card marginal per dependent read (15.1a). In the identity's terms `F` is 128 x 8.6 nJ = 1.1 microjoules per hash on the card and the chip's cost for the same work is 128 x 1.15 = 0.15: **`k` about 0.13**, under the ALU shadow's 0.3 to 0.8 and no better than the L2 hot table's. The energy per hash rises 34 percent on the 4090 and 33 percent on the H100's best form against the modelled chip's 33 (GDDR7) and 50 (HBM3): the edge at zero shadow moves 0 on Ada against the GDDR7 chip and 0.9x in the chip's favour against HBM3 on Hopper; against the SRAM die the chip's 66x to 19x becomes about 25x against the card's new energy, not 19x. (3) The one row that could change the sign is the knee: at the 1,300 MHz lock the 5090's fixed share is smaller and its fabric runs at a lower voltage, so the sector's 8.6 nJ would fall with the measured 10.9 to 8.7 nJ per read (15.1a) and the card's rise would be a larger fraction of a smaller base; the lever lives only if the card's energy per hash rises under about 20 percent there, and a rented pod cannot read it (`nvidia-smi -lgc` refused: "does not have permission to change clocks"). That row is one PC 1 run (the `w4` and `w64-l2` packs at `--block-warps 1`, 60 batches, at the lock) and is owed to the hash lane as a formality, not a gate.
| Outcome | The 5090's rate at W = 16 | Card watts | `E_card` | Edge at zero shadow | With the class v4 shadow at k 0.3 / 0.5 / 1 | Reading |
|---|---|---|---|---|---|---|
| A. The exported form at full occupancy (-47 percent, measured unlocked) | 67.5 MH/s | about 225 W (the per-second terms unchanged; +10 to 18 W of DRAM depending on whether the second sector re-activates) | 3.3 | **5.3x** | 4.5x / 4.1x / 3.1x | dead: the card halves, the chip does not |
| B. The probe's best lane count (-13 percent, measured as a probe) | 110.8 | about 234 W | 2.11 | **3.4x** | 3.4x / 3.0x / 2.2x | not worth a class change: 0.2x at zero shadow |
| C. A one-request form within 5 percent of the 4-byte rate (unmeasured) | 121 to 127 | about 241 W | 1.90 to 1.99 | **3.05x to 3.2x** | 3.1x / 2.7x / 2.0x | the record's best new lever: every column down 10 to 15 percent, and it stacks with the operating point and the shadow |
| Pass line | 95 percent of the 4-byte rate at both clocks with the `.L2::64B` or tuned-occupancy variant | | | 3.2x or better | | W = 16 words into class v6 as layer 1's width floor |
| Kill line | under 95 percent on every variant | | | | | the band stays {1, 4}; the row closes |
**Verdict: KILL at stock on measured rows; W = 16 words stays out of class v6; the design's "never 16" stands, for the measured reason (the card's fabric price of a sector) rather than the plan's (the pins).** What is kept: the hinted one-request load as a kernel fact for Ada (a 64-byte dependent read at the 4-byte rate), which a future width decision can use; the measured 8.6 nJ per sector as the card-side figure the chip model's width rows lacked; and the per-architecture rate table above.
Against the honest denominators the lever is unconditional, because those cards pay nothing: the M5 Max against the GDDR7 chip 1.67x to 1.26x, against one HBM3 stack 2.4x to 1.95x, against the SRAM die (lane 3's 66x on the 5090 is about 21x on the Mac's 0.78) about 6.2x; the 9070 XT against the GDDR7 chip 22.7x to about 17.6x. The identity check, in the identity's terms: `F` is the second sector's joules on the card (about 1.5 nJ per read, 0.19 microjoules per hash at the knee) and `k` the chip's cost for the same sector over the card's: 1.15 nJ over 1.45 to 1.65, **`k` about 0.75 to 0.8 on GDDR7**, the highest `k` of any forcing work on the table, because 1.15 of it is the DRAM device's own movement at `k = 1` and only the card's fabric share is under 1; on the SRAM die `k` is about 0.4 (0.63 nJ over 1.5) but on a base of 0.036 microjoules, which is why that chip moves 3.5x. It stacks with the ALU shadow (different hardware; the fold adds about 5,760 ALU ops per hash, 6 percent of the shadow's count, inside the wait) and with the operating point. What a chip can do about it: nothing cheaper than paying. Items are pseudo-random (incompressible); the fold is dst-keyed and uses all 16 words, so a pre-folded dataset does not exist (lane 3 reads the same); a narrower access atom does not exist on any DRAM (32 B is the minimum on GDDR7, HBM3, HBM4 and LPDDR6, JEDEC, claimed); spreading the words over devices multiplies the sectors; the HBM parts pay the movement at their own lower pJ per bit, which is the `E_mem` lever the model already carries.
Cost to the honest tiers and the rule of 5 October: nothing ships, so nothing changes for any tier; had it shipped, the Ada and Blackwell tiers would have paid about a third more energy per hash for no rate, Hopper a third more for a fifth of its rate, Ampere a quarter more for an eighth of its rate, AMD and Apple nothing, the verifier 4 percent, and every DRAM chip the same third.
**Why the design's "never 16 words" and lane 3's "dead" are one measurement away from reversing.** Both rest on the -47 percent row, which was the exported kernel at the race's default occupancy with four separate 16-byte loads per item and no request-size hint, and the plan read 589 GB/s as a bandwidth bind. The same card's probe did 64-byte dependent reads at 1.0 TB/s. The job: the `w64` pack on PC 1's 5090 (the read-width exe or the installed worker with the pack), (i) the race's occupancy variants (warps per block 1, 2, 4, 8 and the persistent shapes) to find the probe's lane count inside the hash, (ii) a kernel-text variant with `ld.global.L2::64B.v4.u32` on the first of the four loads (one anchor rewrite in `variantSource`; the compile either takes the qualifier or drops the variant with "compile:" on the race line, which is the known-failed case), (iii) nvidia-smi at 1 Hz, unlocked and at the 1,300 lock, the `w4` pack beside it as the control, 250 batches per row; six to ten rows, under an hour. Rate and MH/W per row; the pass line above. The M5 Max's power channels at `w64` are the second row owed (its rate is measured free; its joules move by at most the DRAM's own movement term). A Windows build is not needed if the installed worker takes the pack; otherwise the cross-compile is on igneum-build-1, never this Mac.
### 3.2 Two items per read: W = 32, 128 bytes, one NVIDIA L2 line (KILL with 3.1)
Cost to the honest tiers and the rule of 5 October, on outcome C: the 5090 class pays about 10 percent of energy per hash for no rate (a rig at the knee about +28 W per card); the 5070 Ti and the 12 GB and 8 GB NVIDIA tiers the same share (approximate); AMD and Apple nothing (measured); a pool user nothing; a node about 4 percent on the verifier (measured); the public claim moves from "3.6x at the knee" to "3.0x to 3.2x" at zero shadow and from "2.1x at `k = 1`" to "2.0x", and the SRAM die's line from "66x at the hash's width" to "19x". On outcome A nothing moves and the row closes. The known-failed case for the acceptance rule: the `w64` class's census (0 to 14 of 60 rejected under version 2) and its F8 uniformity row are owed at 2^24 on 64 seeds under the sub-version 3 rule before any cut.
What it changes in class v6 on a pass: layer 1's read-width row ("pinned at 4 words, 8 after the owed rows, 16 never") becomes "the floor W = 16 words"; lane A's ranking row 3 ("keep the read width out of the era draw: harmless and worthless") keeps its first clause and loses its second; lane 3's width table keeps its chip column and replaces its card column's last row.
### 3.2 Two items per read: W = 32, 128 bytes, one NVIDIA L2 line (KEEP for v7 behind rank 1, rank 4)
| Side | At W = 32 | Label |
|---|---|---|
| GDDR7 chip | one activate plus four column accesses: 0.909 + 4 x 1.15 = 5.5 nJ per read; 21.3 G activates x 128 B = 2.7 TB/s is over the board's 1.79, so the chip becomes bandwidth-bound at 14 G reads per second, 110 MH/s; `E_mem` = 128 x 5.5 nJ + 35 W / 110 M = 0.70 + 0.32 = **1.02 microjoules** (+120 percent); capex per MH/s +50 percent (USD 2.8 to 4.3) | modelled |
| HBM3 one stack | 0.909 x 0.66 + 4 x 0.59 = 2.96 nJ; 10.7 G x 128 B = 1.37 TB/s over the stack's 0.82, bound at 6.4 G reads, 50 MH/s; `E_mem` = 0.38 + 14 W / 50 M = 0.66 microjoules (+105 percent) | modelled |
| The SRAM die | about 1.5 nJ (1,070 bits of wire), 0.21 microjoules per hash: 11x against the 5090's stock point | modelled on lane 3's figure |
| RTX 5090 at the knee | even in a one-request form the pins bind: 17.9 G x 128 B = 2.3 TB/s over 1.79, so at best 14 G reads per second, 110 MH/s (-20 percent); in the exported form far worse by the 64-byte row; DRAM +3 sectors x 1.15 nJ x 14 G = +48 W, fabric +15 W: about 276 W at 110 MH/s, 2.5 microjoules (+50 percent) at best | modelled from the rated peak; no `w32` pack exists |
| Apple M5 Max | 3.5 G x 128 B = 448 GB/s against a measured 522 GB/s stream: 0 to -14 percent of rate (approximate; the Apple fetch granularity at 128 B unmeasured) | approximate |
| RX 9070 XT | two 64-byte lines per read: 2.4 G x 128 B = 307 GB/s of 636; unmeasured whether the second line costs it a second fetch | unmeasured |
| The edge, 5090 knee, GDDR7, at the card's best case | zero shadow 2.5 / 1.02 = 2.4x; with the class v4 shadow at `k = 0.3`: 2.6x; at `k = 0.5`: 2.3x; at `k = 1`: 1.9x | modelled |
Reading: wider reads walk both sides into the regime where the DRAM's own joules are most of both bills, and there the edge tends toward the ratio of two bandwidth-bound machines (about 2.5x, modelled) at the price of the honest card's rate (the Ethash shape the design avoided on 5 October for that reason). W = 16 takes most of the per-joule gain, possibly at no rate; W = 32 takes the rest at a fifth of the 5090's rate at best and a 50 percent rise in the chip's capex per MH/s. A v7 candidate only if rank 1 passes, behind a `w32` export and its rows; not a v6 change.
The chip at W = 32: one activate plus four column accesses, 5.5 nJ per read; 21.3 G activates x 128 B = 2.7 TB/s is over the board's 1.79, so the GDDR7 chip becomes bandwidth-bound at 14 G reads per second, 110 MH/s, `E_mem` 1.02 microjoules (+120 percent); one HBM3 stack 0.66 (+105); the SRAM die about 0.21 (11x). The card: the pins bind the 5090 at 2.3 TB/s before any fabric cost (-20 percent of rate at best), and the measured sector price of 3.1 (8.6 nJ per sector at stock) puts three more sectors at about 3.3 microjoules per hash on the card against the chip's 0.44: `k` about 0.13 again, on a card that has also lost a fifth of its rate. Dead with 3.1; no `w32` pack is needed.
### 3.3 Independent chains per hash (memory-level parallelism as the card's lever) (KILL, one probe would reopen it)
The idea: four independent dependent chains of 32 reads per hash, folded at the end, so each resident hash keeps four reads in flight and an occupancy-bound card climbs toward its memory's activate ceiling while the chip (already at the ceiling) gains nothing; an `E_card` lever. The number: the 5090 reads 17.5 to 18.2 G per second at 4 B as "the best over lanes in flight" (the 5 October probe), which means more lanes did not help, so the bind is the memory system, not the in-flight count; the bound is the ceiling's 21.3 G, +17 to +22 percent, and the expectation is 0 to 5 percent (the microbench's `l2_indep4` against `l2_chase` read +2 percent on the L2-bound pattern). The 9070 XT caps at 2.63 G at 4,096 lanes (measured) with the same plateau shape. A class change for an expected few percent, which also cuts the sequential depth per hash to 32. KILL; the probe that would reopen it is one `dram_indep4_1g` row in the microbench (60 s on PC 1). (The 64-byte case of 3.1 is different: there the probe's own best row is 14 percent under the 4-byte row and the hash sits 47 percent under it, so the occupancy gap is measured, not conjectured.)
The idea: four independent dependent chains of 32 reads per hash, folded at the end, so each resident hash keeps four reads in flight and an occupancy-bound card climbs toward its memory's activate ceiling while the chip (already at the ceiling) gains nothing; an `E_card` lever. The number: the 5090 reads 17.5 to 18.2 G per second at 4 B as "the best over lanes in flight" (the 5 October probe), which means more lanes did not help, so the bind is the memory system, not the in-flight count; the bound is the ceiling's 21.3 G, +17 to +22 percent, and the expectation is 0 to 5 percent (the microbench's `l2_indep4` against `l2_chase` read +2 percent on the L2-bound pattern). Today's rows agree from the other side: on the 4090 the `w4` hash at 64.7 MH/s is 8.3 G reads per second, the probe's own ceiling (8.3 to 8.9), at every occupancy from one warp per block to the sparse grids (64.70 to 64.74). KILL; the probe that would reopen it is one `dram_indep4_1g` row in the microbench.
### 3.4 Row-straddling items (two activates per read) (KILL, dominated)
A 64-byte item laid across a row boundary costs the chip two activates: 2 x 0.909 + 2 x 1.15 = 4.1 nJ per read (+105 percent) and halves its activate-bound rate (83 MH/s), so `E_mem` = 128 x 4.1 nJ + 35 W / 83 M = 0.52 + 0.42 = 0.94 microjoules; the card's activate-bound rate halves too (about 83 MH/s if it sits at 82 percent of the same ceiling), at about +22 W: 2.95 microjoules. Edge 3.1x at a 39 percent rate loss, against the aligned item's 3.0x to 3.2x at no rate loss on outcome C. Dominated by 3.1 whenever 3.1 passes, and worse than it on every outcome. KILL (the invention lane's 2.13 killed the 256-byte form on the w64x4 row; this is the 64-byte form).
A 64-byte item laid across a row boundary costs the chip two activates: 2 x 0.909 + 2 x 1.15 = 4.1 nJ per read (+105 percent) and halves its activate-bound rate (83 MH/s), so `E_mem` = 0.94 microjoules; the card's activate-bound rate halves too, at the fabric price of 3.1 on top. Dominated by the aligned item, which is itself dead. KILL.
### 3.5 Scratch state in DRAM (KILL)
@ -258,48 +236,47 @@ Both sides are capex-dominated (the record's 16.1: 7x to 10x their electricity p
| Candidate | The identity check | The number | Label |
|---|---|---|---|
| The second sector, W = 16 words (the whole item) | the card pays the sector at its own fabric's price: `k` about 0.13; energy per hash +34 percent (4090), +33 (H100 best form), +25 (3090) against the chip's +33 | 8.6 nJ per sector on a 4090 at stock; the rate holds on Ada with the hint (98 percent), fails on Hopper (79) and Ampere (87) | measured card, modelled chip |
| Two items per read, W = 32 | the same `k` on three more sectors, on a card the pins bind at -20 percent | chip 1.02 microjoules; card about 3.3 of sector work | modelled on the measured sector price |
| Tensor tiles as the shadow (int8) | `k` 0.2 to 0.4 on the class's tile, 0.4 to 0.7 on the dense tile: never above the ALU shadow's band | the M5 Max -35 percent at 1,024 tiles, -78 at 4,096 (measured); the same premium as the ALU shadow by construction | measured card, claimed chip |
| FP8, FP16, BF16 tiles; the texture interpolator; the rasteriser | not bit-exact across vendors | | spec and vendor documents |
| The hardware video decoder | bit-exact by standard, but the honest card is decoder-bound (a fixed engine count) and the verifier decodes at milliseconds per frame | engines per card two to four (approximate) against any count on a chip | approximate |
| The RT core | traversal implementation-defined on Vulkan, DXR, OptiX and Metal; the BVH opaque; `k` 0.01 to 0.2 on the public per-ray figures | 4 to 33 nJ per ray on a dedicated unit (claimed) against 290 to 750 measured board-level on an RTX 2080 | claimed and measured |
| Memory-level parallelism per dollar | no asymmetry per channel; the correction is the PHY's node (section 2.3, kept as rank 2) | `E_mem` under +0.06 microjoules | claimed, modelled |
| Memory-level parallelism per dollar | no asymmetry per channel; the correction is the PHY's node (section 2.3, kept as rank 1) | `E_mem` under +0.06 microjoules | claimed, modelled |
| Proof of latency | the chain checks values; the block is 1 s against a 0.4 microsecond round trip; the farm's node is local; a sequential prefix favours the chip about 4x | 2.5 million round trips per block | measured card latency, modelled chip |
| The drawn address map | firmware; no prefetch exists for a dependent chain | 0 to the chip, 0.8 to 3.2 percent to the cards | measured spread |
| Independent chains per hash | the 5090's read rate plateaus over lanes; the bind is the memory system | expected 0 to 5 percent, bound 22 | measured probe |
| Row-straddling items | dominated by the aligned item | 3.1x at -39 percent of rate against 3.0x to 3.2x at 0 | modelled |
| Independent chains per hash | the card's read rate sits at its probe's ceiling at every occupancy | the 4090 64.70 to 64.74 MH/s across six shapes; expected 0 to 5 percent on the 5090, bound 22 | measured |
| Row-straddling items | dominated by the aligned item, which is dead | | modelled |
| Scratch in DRAM | the card's scratch sits in L2; the chip's in SRAM | 465 KB against 96 MB | measured rows |
| The whole item in the exported form (outcome A of 3.1) | the card halves, the chip does not | 5.3x | measured rate, modelled chip |
## 5. Consequences per tier (the standing rule of 5 October 2026)
| Tier | What this file means | What is being done |
|---|---|---|
| Home miner, one 8 GB NVIDIA card (32-byte sectors) | on a pass of rank 1, about 10 percent of energy per hash (approximate, the 5090's share scaled) for no rate, and its row against every chip improves by 10 to 15 percent; on a fail, nothing | the 5090 job first; the 4060 Ti class scaled from it |
| One 12 GB card (4070, 5070) | the same share; the 4070 at its tune point about +8 W (approximate) | the same |
| One 16 GB card (RX 9070 XT) | nothing in watts or rate at any outcome (-3 percent of rate measured at 64 B); against the GDDR7 chip 22.7x to about 17.6x and against the SRAM die 3.5x better, on a pass | nothing to measure; the vendor-share metric |
| One 24 or 32 GB card (5090 class) | the whole question: -47 percent of rate as exported (measured), -13 percent at the probe's best lane count (measured), unmeasured in a one-request form; on a pass +28 W at the knee (modelled) and 3.6x to 3.0x at zero shadow | one PC 1 job: the `w64` pack, occupancy variants, the `.L2::64B` variant, both clocks, watts at 1 Hz, the `w4` control; under an hour |
| Apple (M5 Max and the unified tiers) | nothing at any outcome (0 at 64 B, measured); the honest best per joule improves against every chip by the chip's cost alone on a pass: 1.7x to 1.26x on GDDR7, about 21x to 6.2x against the SRAM die (modelled) | the M5 Max power channels at `w64` (one Mac measurement under the measure lock, when the Mac rule allows) |
| A rig | on a pass, +12 percent of watts on NVIDIA cards for no rate; the chip's bill per MH/s-hour up a third on energy | design 1 (the knee lock) as the default, then W = 16 |
| A pool user | nothing changes in shares; the class change is announced by the 95 percent signal | |
| A node operator (the verifier) | +4 percent (0.630 against 0.604 ms per 32 hashes, measured); W = 32 about the same again | nothing |
| A chip | GDDR7 +33 percent of energy per hash, HBM3 +24, the base die +50, the SRAM die +250 (all modelled) on a pass; the cheapest project is a 12 nm GDDR6 part at USD 20 M to 30 M, not a 28 nm one at 5 M, on any outcome | the chip model's rows 5.4 and 5.6 and the record's 16.2 corrected (owed to the Counter lane) |
| The public claim | on a pass "3.6x at the knee, 2.1x at `k = 1`" becomes "3.0x to 3.2x and 2.0x" and the SRAM die's "66x" becomes "19x"; the tensor `k` sentence loses "4x to 30x" on any outcome; the Ethash precedent gains the 13x row | the Counter lane's texts, after the 20:00 word |
| Home miner, one 8 GB NVIDIA card (32-byte sectors) | nothing ships; had W = 16 shipped, this tier would have paid about a third more energy per hash for no rate (Ada) and an eighth of its rate too (Ampere) | nothing; the knee row on PC 1 is a formality |
| One 12 GB card (4070, 5070) | the same | the same |
| One 16 GB card (RX 9070 XT) | nothing (its 64-byte line is free at W = 16, measured 5 October); its row against the chip is unchanged at 22.7x on class v3 | the vendor-share metric |
| One 24 or 32 GB card (5090 class) | nothing ships; the 5090's own stock row is pending a working pod and amends this file; the record's 3.6x at the knee and 2.1x at `k = 1` stand | the secure 5090 pod; the knee row on PC 1 (the hash lane, one row) |
| Apple (M5 Max and the unified tiers) | nothing (free at 64 bytes, measured 5 October); the honest best per joule is unchanged | nothing |
| A rig | nothing ships; the two model corrections raise the chip's project floor (USD 20 M to 75 M), which is the rig's real wall, by 4x to 15x over the record's cheapest row | the chip model's rows corrected (the Counter lane) |
| A pool user | nothing changes | |
| A node operator (the verifier) | nothing changes (the +4 percent of `w64` never ships) | |
| A chip | the cheapest project is a 12 nm GDDR6 part at USD 20 M to 30 M or an N5 GDDR7 part at 50 M to 75 M, not a 28 nm one at 5 M; its `k` on tile work 0.2 to 0.7, not 0.03 to 0.3; its cost of the second sector (+33 percent) was matched by the card's | chip-model-v3 5.4, 5.6 and 5.11 and the record's 16.2 corrected (owed to the Counter lane) |
| The public claim | unchanged by this lane on the hash; the tensor `k` sentence loses "4x to 30x"; the Ethash precedent gains the 13x row; the cheapest-chip project line gains a floor | the Counter lane's texts |
## 6. Unverified and owed
- The 5090's rate at 64 bytes in a one-request form and at its best occupancy inside the hash: unmeasured; the two measured rows (-47 percent as exported, -13 percent as a probe) bracket it. Its watts at 64 bytes at either clock: unmeasured (the 5 October job carried no power sampling). This is the one job this lane asks for; the pass and kill lines are stated in 3.1.
- Whether the 5090's controller re-activates for the second sector (close-page on a random stream) or the loss is queueing: the same job reads it from the rate per variant.
- The M5 Max's power channels at `w64` are unmeasured (the rate is).
- The chip side of rank 1 is the model's own data-movement term applied twice (4.5 pJ per bit, Micron, claimed, streaming; O'Connor's 3.48 for HBM2) and lane 3's wire figure for the SRAM die; no one has measured a second-sector column read's energy on GDDR7 (section 2.6: the public record has no random-read joule on any current part).
- The fabric share of the card's second sector (0.3 to 0.5 nJ) is approximate.
- The W = 32 rows rest on the rated 1.79 TB/s; no `w32` pack exists.
- The 5090's own stock rows (`w4`, `w64`, `w64-l2` at full occupancy and the sparse shapes, with watts): pending the secure-cloud pod after three community pods on the same host failed `cuInit` with 999; amended into 3.1 when they land.
- The knee row (the 5090 at the 1,300 MHz lock, `w4` against `w64-l2`, 60 batches, watts): a rented pod refuses the lock; one PC 1 row with the hash lane; the lever lives only under a 20 percent energy rise there, which the stock rows make unlikely.
- The 4090's and H100's raw logs were not copied before their pods were destroyed; the rows in 3.1 are the figures read off the pods during the session. The 3090's raw logs are at `~/igneum-fleet/fl5-results/3090-*`.
- The watts are `nvidia-smi` board power at 1 Hz, the upper half of each row's samples averaged; idle on the rented pods read 29 W (4090) and 70 W (H100); the 3090's idle row was taken straight after a run and is not a settled idle.
- The chip side of 3.1 is the model's own data-movement term applied twice (4.5 pJ per bit, Micron, claimed, streaming; O'Connor's 3.48 for HBM2) and lane 3's wire figure for the SRAM die; no one has measured a second-sector column read's energy on GDDR7 (section 2.6: the public record has no random-read joule on any current part).
- The tensor chip figures are vendor specifications at TDP (claimed), the M4 ANE and the H800 loop measured; the RT figures are claims and one measured RTX 2080 row; the PHY node facts are vendor IP pages read today.
- The project-cost corrections use the history's figures (a 7 nm-class startup project USD 50 M to 75 M, claimed; the 12 nm figure scaled, approximate) and the mission lane's cap model.
- The `w64` class's census under the sub-version 3 rule and its F8 row at 2^24 on 64 seeds are owed before any cut (the 5 October census ran version 2's rule).
- Nothing ran on the Mac, on a box or on a PC for this file.
- Nothing ran on the Mac, on a Hetzner box or on a PC for this file; the measurements ran on four RunPod one-shot pods, all destroyed on their done (the 5090 pod when its rows are in).
## 7. Sources
Internal: `docs/analysis/counter-asic-4-research.md` (sections 2, 4, 5, 8, 15.1a, 16, 20), `docs/analysis/chip-model-v3.md` (5.1, 5.3, 5.4, 5.6, 5.10 to 5.12), `docs/design/class-v6-rotating-family.md` (sections 2, 7, 7a, 7c, 10 and 10.3), `docs/analysis/class-v6/hardware-future.md` at 7618e729 (sections 0, 1, 2, 4.9, 5), `docs/analysis/class-v6/history.md` (sections 3 and 5), `docs/analysis/class-v6/invention.md` on `class-v6-invention` (sections 1 and 2), `docs/analysis/class-v6/floor/sram-and-floor.md` on `class-v6-floor-sram` as carried in section 10.3, `docs/plans/read-width.md` (sections 1, 1.5, 4.1), `docs/bench-log.md` (the 5 October read-width entry: the probes and the hash rates), `proto-cuda/packs-readwidth/w64/kernel_bound.cu` (the four `uint4` loads per item), `proto-cuda/nvrtc/worker.cpp` (`variantSource`), `docs/analysis/scratch-soundness.md`, spec 01 (1.5, the item table).
Internal: `docs/analysis/counter-asic-4-research.md` (sections 2, 4, 5, 8, 15.1a, 16, 20), `docs/analysis/chip-model-v3.md` (5.1, 5.3, 5.4, 5.6, 5.10 to 5.12), `docs/design/class-v6-rotating-family.md` (sections 2, 7, 7a, 7c, 10 and 10.3), `docs/analysis/class-v6/hardware-future.md` at 7618e729 (sections 0, 1, 2, 4.9, 5), `docs/analysis/class-v6/history.md` (sections 3 and 5), `docs/analysis/class-v6/invention.md` on `class-v6-invention` (sections 1 and 2), `docs/analysis/class-v6/floor/sram-and-floor.md` on `class-v6-floor-sram` as carried in section 10.3, `docs/plans/read-width.md` (sections 1, 1.5, 4.1), `docs/bench-log.md` (the 5 October read-width entry: the probes and the hash rates), `proto-cuda/packs-readwidth/w64/kernel_bound.cu` (the four `uint4` loads per item), `proto-cuda/nvrtc/worker.cpp` (`variantSource`, the `sp<N>-w<W>` shapes, the Linux `dlopen` path), the fleet's `oneshot.py` and the rented rows' logs under `~/igneum-fleet/fl5-results/`, `docs/analysis/scratch-soundness.md`, spec 01 (1.5, the item table).
External, all read 8 October 2026 by this lane's three research sub-agents (tensor, RT, literature) and labelled where carried: the NVIDIA RTX 5090 page and the Blackwell architecture whitepaper; the Ada whitepaper; arXiv 2501.12084v2 (Hopper tensor-core microbenchmarks); Lenovo LP2226 (B200); the AMD MI300X data sheet and MI355X notes; the RX 9070 XT page; Qualcomm's AI 100 Ultra brief; eeNews Europe (MTIA v2); AnandTech 21342 (Gaudi 3); the Hot Chips 2025 Ironwood slides; arXiv 2608.28048 (Trainium2); Tenstorrent's Blackhole specification; the M4 ANE measurement (maderix.substack.com); cnx-software (Hailo-8); design-reuse 52539 (Untether); hackster.io (Axelera); NVIDIA Research's VSQ accelerator (JSSC 2023); Dally, Hot Chips 2023 keynote, slides 12 and 24; the Vulkan ray traversal chapter; the DXR specification; the NVIDIA developer forum thread 309730 (14 October 2024); the Blender Cycles source note; the Mach-RT TVCG preprint; Hot Chips 31 (Turing); the AMD Hot Chips 2025 RDNA 4 slides; embedded.com (RayCore); AnandTech 7870 (GR6500); arXiv 2409.06000 (RayFlex); the TRaX CGI 2018 study; Ylitie et al., HPG 2017; Chou et al., MICRO 2023; par.nsf.gov 10156963 (Choe et al., MEMSYS 2019); arXiv 2508.06795; SemiAnalysis "The Memory Wall"; the SK hynix ISSCC 2024 GDDR7 paper (ResearchGate 378947349); arXiv 2410.12990; arXiv 2405.06081; asicminervalue (iPollo V2H, V1, A11 Pro); kryptex (Jasminer X16-Q, FishHash); whattomine (the 5090 on Etchash); monero-project issue 10270; arXiv 2512.01437v2 and decrypt.co 363018 (Qubic); Business Wire 20240925780123 (Cadence GDDR7 on N3) and 20221115005674 (GDDR6 at 12 nm); innosilicon.com GDDR7; kurnal-insights (GB202 die); TrendForce 24 September 2026 (the 2 GB GDDR7 part); arXiv 2403.13230 (BFT-PoLoc); eprint 2026/694 (unread, refused); the PTX ISA (`ld` with the `.L2::64B` prefetch-size qualifier, sm_75 and later).