Class v6 floor lane 5 (invention), corrected for the read-width naming: the whole 64-byte item is class w64 (W = 16 words), not w16; its card side is measured at two occupancies that disagree (the hash -47 percent at full occupancy, the probe -13 percent at its best lane count) and unmeasured in a one-request form, so rank 1 is conditional on one PC 1 job with a 95 percent pass line; the chip side unchanged (+33 percent GDDR7, +24 HBM3) and, on floor lane 3's wire figure, the SRAM die 66x to 19x; the three outcomes priced (5.3x dead, 3.4x, 3.0x to 3.2x)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
cf602a0314
commit
ef79c5e71f
1 changed files with 69 additions and 62 deletions
|
|
@ -1,28 +1,28 @@
|
|||
# Class v6 floor lane 5, invention: what none of the four floor lanes covers, checked against the identity
|
||||
|
||||
8 October 2026, 12:48 to 19:xx UK, branch `class-v6-floor-invention` from `counter-asic-4` (base 91117093, the sound per-load form's 5090 rows). The open-ended lane of the floor research: four other lanes work the known levers (SM-sparse gating, the shadow's `k` from RTL synthesis, the SRAM full-store chip and the dataset floor, the honest denominator per tier); this file finds what they do not cover. Every candidate is checked against the identity first and killed or kept with a number. Every number carries a label: **measured** (a card on a named job, the file named), **modelled** (arithmetic on the chip model's cited inputs, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling). Nothing here changes a consensus object, a pack or a served number. Research and reasoning ran on this Mac; no build, no benchmark and no miner ran anywhere for this file.
|
||||
8 October 2026, 12:48 to 19:xx UK, branch `class-v6-floor-invention` from `counter-asic-4` (base 171618c5 after the fast-forward; first cut at 13:0x UK, corrected 13:3x UK for the read-width naming: `w16` is 16 bytes, `w64` is the 64-byte item). The open-ended lane of the floor research: four other lanes work the known levers (SM-sparse gating, the shadow's `k` from RTL synthesis, the SRAM full-store chip and the dataset floor, the honest denominator per tier); this file finds what they do not cover. Every candidate is checked against the identity first and killed or kept with a number. Every number carries a label: **measured** (a card on a named job, the file named), **modelled** (arithmetic on the chip model's cited inputs, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling). Nothing here changes a consensus object, a pack or a served number. Research and reasoning ran on this Mac; no build, no benchmark and no miner ran anywhere for this file.
|
||||
|
||||
## 0. One page
|
||||
|
||||
**The identity, as the record states it** (`docs/analysis/counter-asic-4-research.md` section 2): against the chip that stores the dataset, `edge = (E_card + F) / (E_mem + k F)`; at zero premium the edge is `E_card / E_mem` (3.6x on a 5090 at its 1,300 MHz knee against the GDDR7 board, measured card, modelled chip), and a shadow op with `k` under 1 never fully closes it. The one thing with `k = 1` by construction is the DRAM itself: the same 16 devices on both sides. The record forced one DRAM operation per read (the activate and one 32-byte sector) and called every wider read "nothing to force: the chip already pays the sector".
|
||||
**The identity, as the record states it** (`docs/analysis/counter-asic-4-research.md` section 2): against the chip that stores the dataset, `edge = (E_card + F) / (E_mem + k F)`; at zero premium the edge is `E_card / E_mem` (3.6x on a 5090 at its 1,300 MHz knee against the GDDR7 board, measured card, modelled chip), and a shadow op with `k` under 1 never fully closes it. The one thing with `k = 1` by construction is the memory itself: the same devices on both sides. The record forced one DRAM operation per read (the activate and one 32-byte sector) and called every wider read "nothing to force: the chip already pays the sector". That is true up to 32 bytes and false at 64.
|
||||
|
||||
**The one new lever this lane finds: the second sector.** The dataset item is 64 bytes (16 words, spec 1.5 and the item table). The hash reads one word of it, which moves one 32-byte sector on the 5090 and on the chip's GDDR7. Reading the whole item (the read-width branch's class `w16`, built and measured 5 October, bit-exact on four runtimes, fingerprint `e7c890445b47af60`) moves two sectors: on the chip model's own inputs that is the data-movement term paid twice (4.5 pJ per bit x 256 bits = 1.15 nJ, Micron, claimed) on one activate (909 pJ), so the chip's energy per read rises from 2.0 to 3.2 nJ (+56 percent, modelled) and its energy per hash from 0.466 to 0.62 microjoules (+33 percent; the static and controller watts are unchanged and its activate-bound rate is unchanged at 166 MH/s, 1.36 TB/s of the board's 1.79). The honest 5090 pays the same DRAM joules (+20 W at 17.9 G reads per second) plus one more sector through its L2 and crossbar (about +6 to +10 W, approximate), on a rate that MEASURED 2.7 percent faster at w16 (139.8 against 136.1 MH/s, 5 October); the RX 9070 XT (a 64-byte line per read, measured) and the M5 Max (64 B at the 4 B rate, measured) pay nothing. Card energy per hash at the knee: 1.67 to about 1.84 microjoules (+10 percent, modelled; the w16 job carried no power sampling, so this is the one row owed). The edge at zero shadow: 3.6x to 3.0x on GDDR7, 5.2x to 4.6x on one HBM3 stack; with the class v4 shadow at `k = 0.3`: 3.5x to 3.05x; at `k = 0.5`: 2.9x to 2.6x; at `k = 1`: 2.1x to 2.0x. The verifier pays nothing (0.610 against 0.604 ms per 32 hashes, measured). It is the only lever on the table whose forced work is `k = 1`, it stacks with the operating point and with the shadow, and it is against the SRAM full store alone that it does nothing (that chip's macro read is already 64 bytes). **KEEP, rank 1, into class v6 now** as the read-width floor (W = 16 words, the whole item), replacing the layer-1 band {1, 4} words whose warrant was that wider reads are harmless and worthless.
|
||||
**The one lever this lane finds: the second sector, which is the whole item.** The dataset item is 64 bytes (16 words; spec 1.5 and the item table). The hash reads one word of it, which moves one 32-byte sector on the 5090 and on the chip's GDDR7; the read-width band of class v6 (W in {1, 4} words, 4 to 16 bytes) stays inside that sector, so a DRAM chip pays nothing more at any width in the band, as the record said. Reading the whole item (W = 16 words, the read-width branch's class `w64`, built and measured 5 October, bit-exact on four runtimes, fingerprint `836e56e7d496e980`, the verifier 0.630 against 0.604 ms) moves two sectors in one open row: on the chip model's own inputs that is the data-movement term paid twice (4.5 pJ per bit x 256 bits = 1.15 nJ, Micron, claimed) on one activate (909 pJ), so the GDDR7 chip's energy per read rises from 2.0 to 3.2 nJ (+56 percent, modelled), its energy per hash from 0.466 to 0.62 microjoules (+33 percent; static, controller and the activate-bound rate unchanged), one HBM3 stack's from 0.321 to 0.40 (+24 percent), and, on floor lane 3's wire figure, the N2 SRAM full store's from 0.036 to 0.126 microjoules (66x to 19x against the 5090's stock point, that lane's own rows): **the second sector is the one lever that moves every chip in the model at `k` near 1, and the SRAM die most of all.** What decides its worth is the honest card's rate at 64 bytes, which is measured in two places that disagree: the hash at full occupancy ran at 71.9 MH/s on the 5090 (-47 percent, the row the design calls dead) while the same card's 64-byte dependent-read probe reached 15.7 G reads per second at its best lane count (-10 to -14 percent against its 4-byte probe), both 5 October; the 589 GB/s the hash moved is a third of the stream, so the bind at full occupancy is the request path (two serialised sector misses, a second activate at the controller, or queueing), not the pins. The RX 9070 XT (-3 percent) and the M5 Max (0) pay nothing at 64 bytes (measured). Three outcomes, priced at the 5090's knee against the GDDR7 chip (modelled): the exported form at full occupancy, 5.3x (dead: the card halves, the chip does not); the probe's best lane count, 3.4x (not worth a class change); a one-request form that holds the 4-byte rate within 5 percent (PTX `ld.global.L2::64B` on the first word, or the lane count tuned by the race), 3.0x at zero shadow and 2.7x with the class v4 shadow at `k = 0.5`, with every chip column down 10 to 15 percent and the SRAM die's down 3.5x. **KEEP, rank 1, as the one measurement this lane asks for before the 20:00 close's v6 word**: the `w64` pack on PC 1's 5090, an occupancy sweep and the `.L2::64B` variant, watts at 1 Hz, unlocked and at the 1,300 lock; pass line 95 percent of the 4-byte rate (then W = 16 words becomes layer 1's floor in v6), kill line under 95 percent (then the band stays {1, 4} and this row closes).
|
||||
|
||||
**Everything else on the brief is killed, with the number.** The tensor core: the record's `k` 0.03 to 0.3 is not defensible at its floor (the 0.04 pJ figure is INT4 at 0.46 V on a test chip with no memory system); shipping merchant silicon at nominal voltage reads 0.30 to 0.56 pJ per INT8 MAC (Meta MTIA v2 0.51, Qualcomm AI 100 Ultra 0.34, AMD MI355X 0.56, NVIDIA B200 0.44, Apple M4 ANE 0.30 measured; all claimed unless marked), and a Hopper tensor core under a pure MMA loop measured 0.34, so the honest band against the tile the class runs (1.5 pJ per MAC at the knee, measured) is `k` 0.2 to 0.4 and against the dense wide tile (0.83 at the knee) 0.4 to 0.7: equal to or under the ALU shadow's 0.3 to 0.8, never above it, and the Apple cost (-35 percent of rate at 1,024 tiles per hash, measured) kills it as content regardless. The RT core: traversal results are implementation-defined on every API (Vulkan "no ordering guarantee", DXR "no defined order", NVIDIA "could change depending on which driver, which GPU"), the acceleration structure is opaque, and a dedicated unit reads 4 to 33 nJ per ray against 290 to 750 nJ measured board-level on an RTX 2080 (`k` 0.01 to 0.2): dead on bit-exactness and on `k`. Memory-level parallelism per dollar: no asymmetry per channel, but a correction to the chip model: no 28 nm GDDR7 PHY exists (the shipped ones are N3 and FinFET; the oldest GDDR6 PHYs are 12 nm), so the record's cheapest chip (a USD 5 M project, a USD 17 M break-even cap) is not buildable; the floor is a 12 nm GDDR6 part at about USD 20 M to 30 M (cap USD 70 M to 100 M) or an N5-class GDDR7 part at USD 50 M to 75 M (cap 170 M to 250 M), modelled. Proof of latency: the chain checks values, never time; the block is 1 s and a DRAM round trip 0.1 to 0.4 microseconds; the farm's node is on its own LAN; a sequential per-block prefix favours the lower-latency side, which is the chip (about 4x). The time dimension on the address map: a bit permutation is firmware (a few hundred gates), no prefetch exists for a dependent chain, and the literature has no such scheme. The literature since 2023 holds nothing outside the identity; what it adds is a higher Ethash precedent (iPollo V2H, 0.14 J per MH, about 13x a 5090, vendor claim), the measured PAM3 I/O energies (SK hynix, ISSCC 2024) and the PHY-node fact above. Of this lane's own ideas, four more die on the identity (video decode, row-straddling, scratch in DRAM, independent chains per hash) and one is held for v7 (two items per read, W = 32: the chip goes bandwidth-bound at 110 MH/s and 1.02 microjoules, the edge at zero shadow 2.4x and 2.5x at `k = 0.3` with the shadow, at a cost of about 20 percent of the 5090's rate, modelled).
|
||||
**Everything else on the brief is killed, with the number.** The tensor core: the record's `k` 0.03 to 0.3 is not defensible at its floor (the 0.04 pJ figure is INT4 at 0.46 V on a test chip with no memory system); shipping merchant silicon at nominal voltage reads 0.30 to 0.56 pJ per INT8 MAC (Meta MTIA v2 0.51, Qualcomm AI 100 Ultra 0.34, AMD MI355X 0.56, NVIDIA B200 0.44, Apple M4 ANE 0.30 measured; all claimed unless marked), and a Hopper tensor core under a pure MMA loop measured 0.34, so the honest band against the tile the class runs (1.5 pJ per MAC at the knee, measured) is `k` 0.2 to 0.4 and against the dense wide tile (0.83 at the knee) 0.4 to 0.7: equal to or under the ALU shadow's 0.3 to 0.8, never above it, and the Apple cost (-35 percent of rate at 1,024 tiles per hash, measured) kills it as content regardless. The RT core: traversal results are implementation-defined on every API (Vulkan "no ordering guarantee", DXR "no defined order", NVIDIA "could change depending on which driver, which GPU"), the acceleration structure is opaque, and a dedicated unit reads 4 to 33 nJ per ray against 290 to 750 nJ measured board-level on an RTX 2080 (`k` 0.01 to 0.2): dead on bit-exactness and on `k`. Memory-level parallelism per dollar: no asymmetry per channel, but a correction to the chip model: no 28 nm GDDR7 PHY exists (the shipped ones are N3 and FinFET; the oldest GDDR6 PHYs are 12 nm), so the record's cheapest chip (a USD 5 M project, a USD 17 M break-even cap) is not buildable; the floor is a 12 nm GDDR6 part at about USD 20 M to 30 M (cap USD 70 M to 100 M) or an N5-class GDDR7 part at USD 50 M to 75 M (cap 170 M to 250 M), modelled. Proof of latency: the chain checks values, never time; the block is 1 s and a DRAM round trip 0.1 to 0.4 microseconds; the farm's node is on its own LAN; a sequential per-block prefix favours the lower-latency side, which is the chip (about 4x). The time dimension on the address map: a bit permutation is firmware (a few hundred gates), no prefetch exists for a dependent chain, and the literature has no such scheme. The literature since 2023 holds nothing outside the identity; what it adds is a higher Ethash precedent (iPollo V2H, 0.14 J per MH, about 13x a 5090, vendor claim), the measured PAM3 I/O energies (SK hynix, ISSCC 2024) and the PHY-node fact above. Of this lane's own ideas, four more die on the identity (video decode, row-straddling, scratch in DRAM, independent chains per hash) and one is held behind rank 1 for v7 (two items per read, W = 32).
|
||||
|
||||
### The KEEP list, ranked
|
||||
|
||||
| Rank | Item | Chip edge at the 5090's knee, GDDR7 (zero shadow / with class v4 at k 0.3 / 0.5 / 1) | Cost to the honest tiers | Verifier | Where | Label |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 1 | **The second sector: W = 16 words, the whole 64-byte item folded (class `w16`)** | 3.0x / 3.05x / 2.6x / 2.0x (from 3.6 / 3.5 / 2.9 / 2.1) | 5090 about +10 percent energy per hash (+26 to 30 W at the knee), rate +2.7 percent measured; 9070 XT and M5 Max 0 (measured rate; the line is 64 B already); 5070 Ti and 4070 as the 5090 (32 B sectors, approximate) | +1 percent (0.610 against 0.604 ms, measured) | class v6 now: layer 1's read-width band becomes the floor W = 16 | card measured (rate) and modelled (watts); chip modelled |
|
||||
| 2 | **The chip model's project floor corrected: no 28 nm GDDR7 PHY** | unchanged | none | none | chip-model-v3 5.4 and the record's 16.2: the cheapest row's project USD 20 M to 75 M, cap USD 70 M to 250 M | claimed (vendor IP pages), modelled (the mission lane's cap) |
|
||||
| 3 | **Two items per read, W = 32 (128 B, one L2 line)** | 2.4x / 2.5x / 2.3x / 1.9x; the chip bandwidth-bound at 110 MH/s | 5090 about -20 percent rate and +23 percent watts (modelled from the 1.79 TB/s peak); M5 Max 0 to -14 percent (approximate); 9070 XT two lines per read (unmeasured) | about +1 percent (the w64 row 0.630 ms) | v7, after the 5090 and M5 Max rate and watts rows on a `w32` export | modelled |
|
||||
| 4 | **The tensor `k` band corrected** (0.2 to 0.4 on the class's tile, 0.4 to 0.7 on the dense tile; not 0.03 to 0.3) | no change to any served row (the tile stays a worse-or-equal lever) | none | none | a wording correction to the record's 15.1a and 20.4 and chip-model 5.11 | claimed and measured |
|
||||
| 1 | **The second sector: W = 16 words, the whole 64-byte item (class `w64`)**, conditional on one PC 1 job | if the 5090 holds 95 percent of its 4-byte rate: 3.0 to 3.2x / 3.1x / 2.7x / 2.0x (from 3.6 / 3.5 / 2.9 / 2.1); at the probe's -13 percent: 3.4x at zero shadow; at the exported form's -47 percent: 5.3x (dead) | 5090 about +10 percent energy at the pass line (+28 W at the knee, modelled) plus whatever rate it loses; 9070 XT -3 percent and M5 Max 0 (measured); the SRAM die 66x to 19x and the GDDR7 chip +33 percent whatever the card does | +4 percent (0.630 against 0.604 ms, measured) | the job first; then class v6's layer-1 width row (the floor W = 16) if it passes | card measured at two occupancies (unlocked, no watts); chip modelled |
|
||||
| 2 | **The chip model's project floor corrected: no 28 nm GDDR7 PHY** | unchanged | none | none | chip-model-v3 5.4 and 5.6 and the record's 16.2: the cheapest row's project USD 20 M to 75 M, cap USD 70 M to 250 M | claimed (vendor IP pages), modelled (the mission lane's cap) |
|
||||
| 3 | **The tensor `k` band corrected** (0.2 to 0.4 on the class's tile, 0.4 to 0.7 on the dense tile; not 0.03 to 0.3) | no change to any served row (the tile stays a worse-or-equal lever) | none | none | a wording correction to the record's 15.1a and 20.4 and chip-model 5.11 | claimed and measured |
|
||||
| 4 | **Two items per read, W = 32 (128 B)**, only if rank 1 passes | the chip bandwidth-bound at 110 MH/s and 1.02 microjoules; 2.4x at zero shadow if the card held its rate, which at 128 B it will not | the 5090 at least -20 percent (the pins bind at 2.3 TB/s), likely far more by the 64-byte row | about +4 percent | v7 at the earliest, after a `w32` export and its rows | modelled |
|
||||
| 5 | **The Ethash precedent's top row** (iPollo V2H, 3.4 GH/s at 475 W, about 13x a 5090 per joule) | the precedent band 2.1x to 13x, not 2.1x to 6.8x | none | none | `asic-resistance-history.md` and chip-model 5.1 | claimed |
|
||||
|
||||
### The five-sentence reading for the 20:00 close
|
||||
|
||||
The identity leaves one piece of work a chip pays at the card's own price, the DRAM's data movement, and the record forced only one sector of it per read; reading the whole 64-byte item (W = 16, the measured `w16` class) forces the second sector at `k = 1`, costs the 5090 about 10 percent of energy per hash and no rate, costs AMD and Apple nothing, costs the verifier nothing, and takes 10 to 15 percent off every chip column (3.6x to 3.0x at zero shadow, 2.9x to 2.6x with the shadow at `k = 0.5`), so it goes into class v6 now as the read-width floor. The tensor core, the RT core, proof of latency and the drawn address map are dead on the identity or on bit-exactness, with the tensor `k` band corrected upward to 0.2 to 0.7 (still never above the ALU shadow's) and the RT core's traversal implementation-defined on every vendor. The chip model's cheapest chip does not exist as priced: a GDDR7 PHY needs a FinFET node (N3 shipped; 12 nm is the oldest GDDR6 PHY), so the floor project is USD 20 M to 75 M and the floor break-even cap USD 70 M to 250 M instead of 17 M, which goes into the model now. Two items per read (W = 32) walks both sides into the bandwidth regime where the DRAM's joules are most of both bills (the edge 2.4x at zero shadow, 2.5x at `k = 0.3`) at a cost of about a fifth of the 5090's rate, so it is a v7 candidate behind two measured rows, not a v6 change. Nothing in the literature since 2023 offers a lever outside the identity, and the one measurement this file owes is the 5090's watts at w16 at the knee and unlocked: one PC job on a pack that already exists.
|
||||
The identity leaves one piece of work a chip pays at the card's own price, the memory's data movement, and the record forced one sector of it per read; reading the whole 64-byte item (W = 16 words, the measured `w64` class) forces the second sector at `k` near 1, costs AMD and Apple nothing (measured), costs the verifier 4 percent, raises the GDDR7 chip's energy per hash by a third and the N2 SRAM die's by 3.5x (modelled), and is dead or alive on one number nobody has yet read: the 5090's rate at 64 bytes in a one-request form at its best occupancy, which the 5 October probe put at -13 percent and the exported hash at -47 percent. That job (the `w64` pack, an occupancy sweep, the `.L2::64B` variant, watts at both clocks) is the one measurement this lane asks for, with its pass line at 95 percent of the 4-byte rate; on a pass W = 16 words becomes layer 1's width floor in class v6 (3.0x at zero shadow against the GDDR7 chip, 2.7x with the shadow at `k = 0.5`, the SRAM die 19x instead of 66x), and on a fail the band stays {1, 4} and this row closes. The tensor core, the RT core, proof of latency and the drawn address map are dead on the identity or on bit-exactness, with the tensor `k` band corrected upward to 0.2 to 0.7 (still never above the ALU shadow's) and the RT core's traversal implementation-defined on every vendor. The chip model's cheapest chip does not exist as priced: a GDDR7 PHY needs a FinFET node (N3 shipped; 12 nm is the oldest GDDR6 PHY), so the floor project is USD 20 M to 75 M and the floor break-even cap USD 70 M to 250 M instead of 17 M, which goes into the model now. Nothing in the literature since 2023 offers a lever outside the identity, and two items per read (W = 32) waits for v7 behind rank 1's result.
|
||||
|
||||
## 1. The frame
|
||||
|
||||
|
|
@ -31,14 +31,17 @@ The identity leaves one piece of work a chip pays at the card's own price, the D
|
|||
| `E_card`, class v3 | 1.67 microjoules (127.3 MH/s at 213.0 W; 134.6 at 223.3 on the efficiency pass) | measured | the record 20.3; `docs/plans/counter-asic-3-status.md` |
|
||||
| `E_mem`, the `f = 1` GDDR7 chip | 0.466 microjoules: 166 MH/s at 77.6 W (42.6 W of reads at 2.0 nJ, 20 W static, 15 W controller) | modelled | chip-model-v3 5.3 and 5.4 |
|
||||
| `E_mem`, one HBM3 stack | 0.321 microjoules: 83.6 MH/s at 26.8 W (12.8 W of reads at 1.2 nJ, 14 W static and controller) | modelled | the same |
|
||||
| `E_mem`, the N2 SRAM full store at the hash's width | 0.036 microjoules at W = 1 (0.25 nJ per read: 1.3 pJ per bit of wire over 80 bits plus 0.1 nJ of macro and 0.05 of controller), 66x the 5090's stock point; 0.126 at W = 16 (0.88 nJ, 560 bits), 19x | modelled (floor lane 3, section 10.3 of the class v6 design) | `docs/analysis/class-v6/floor/sram-and-floor.md` |
|
||||
| The read's energy, split | GDDR7: 909 pJ activation plus 4.5 pJ per bit x 256 bits = 1,150 pJ of movement and I/O (44 / 56 percent); HBM3: 909 plus 3.48 x 256 = 890 scaled by 4.12 / 6.25 | modelled on O'Connor et al. 2017 Tables 2 and 3 and Micron's and Samsung's pJ per bit (claimed) | chip-model-v3 5.3 |
|
||||
| `F`, the class v4 premium | 0.652 microjoules (82.8 W over 126.9 M x 102,100 ops: 6.4 pJ per counted op) | measured | the record 20.3 |
|
||||
| `k`, the ALU shadow | 0.3 to 0.8 (the 5090's 6.2 to 11.3 pJ per op measured against a 5 nm SIMD array's 2 to 5, approximate) | GPU measured, chip approximate | the record 15.1a |
|
||||
| The dataset item | 16 words, 64 bytes; `dataset[w] = item(w >> 4)[w AND 15]`; the lottery hash reads one word per load | design | spec 1.5 and the item table; `docs/plans/read-width.md` |
|
||||
| The 5090's dependent-read rate | 17.5 to 18.2 G reads per second at 4 B and 16 B, the best over lanes in flight; 9.1 to 15.7 at 64 B in the probe; the w16 hash at 17.9 G (139.8 MH/s x 128) | measured, 5 October | `docs/bench-log.md`, the read-width entry |
|
||||
| The read-width classes, by name | `w4` = 4 bytes (W = 1 word, the lottery hash); `w16` = 16 bytes (W = 4, one sector); `w64` = 64 bytes (W = 16, the whole item, two sectors on the 5090); `w64x4` = 64 bytes at 32 loads | design | `docs/plans/read-width.md` 1.5; the bench-log entry of 5 October |
|
||||
| The 5090's dependent-read probe, best over lanes in flight | 17.5 to 18.2 G reads per second at 4 B; 18.0 to 19.9 at 16 B; 9.1 to 15.7 at 64 B (9.1 at 4 M lanes) | measured, 5 October | `docs/bench-log.md`, the read-width entry |
|
||||
| The hash rates at the widths, unlocked, no power sampling | 5090: w4 136.1, w16 139.8, w64 71.9 MH/s; 9070 XT: 18.15, 17.90, 17.59; M5 Max: 27.74, 28.26, 28.27 | measured, 5 October | the same |
|
||||
| The card's whole-card marginal per dependent DRAM read | 10.9 nJ unlocked, 8.7 nJ at the lock | measured, 8 October | the record 15.1a |
|
||||
|
||||
The four doors a candidate can open (the invention lane's section 1, kept): raise `k`; raise the chip's capex or project; shorten the chip's useful life; lower the honest card's cost at the same `F`. This file adds the fifth door the identity itself names: **force work whose `k` is 1 by construction, which is only the DRAM's own.**
|
||||
The four doors a candidate can open (the invention lane's section 1, kept): raise `k`; raise the chip's capex or project; shorten the chip's useful life; lower the honest card's cost at the same `F`. This file adds the fifth door the identity itself names: **force work whose `k` is 1 by construction, which is only the memory's own.**
|
||||
|
||||
## 2. The candidates on the brief
|
||||
|
||||
|
|
@ -77,7 +80,7 @@ The reading. Every shipping systolic chip with a real memory system and fabric l
|
|||
| The dense wide tile (s8 m16n8k32) at the knee | 0.83 | 0.3 to 0.6 | **0.4 to 0.7** | a tensor shadow built of wide tiles would sit at the ALU band's centre; the H800's 0.34 under a pure loop says a GPU tensor core at its best is inside the merchant cluster |
|
||||
| The dense wide tile unlocked | 1.36 | 0.3 to 0.6 | 0.25 to 0.45 | |
|
||||
|
||||
So "3x to 30x" is not defensible: at nominal voltage merchant silicon beats the 5090's dense tile by 2.3x to 4.5x and its dependent tile by 2.5x to 10x, and the 30x needs an INT4 layer at 0.46 V. And "`k` 0.7 to 1 at full rate" is not reached either: the top of the dense band is 0.7. The identity check: the tile's `k` is at best the ALU shadow's and the joules it forces are the same by construction (the record 20.3: 71.5 W against 82.8 at the knee for the same tile count); the chip's downside bet (its `k` floor) is 0.2 against the ALU's 0.3, so the tile remains the worse-or-equal lever, not the 4x to 30x worse lever the record wrote. The cost to the honest tiers is what kills it as content: the M5 Max at -35 percent of rate at 1,024 tiles per hash and -78 percent at 4,096 (measured, the record 20.2a), with no integer matrix path in Metal. Bit-exact on NVIDIA and the CPU (measured); the AMD WMMA layout unverified. **KILL as v6 content; KEEP the band correction (rank 4).** FP8, FP16 and BF16 tiles: the accumulation order inside a tile is unspecified by PTX, so not bit-exact across vendors or generations: KILL.
|
||||
So "3x to 30x" is not defensible: at nominal voltage merchant silicon beats the 5090's dense tile by 2.3x to 4.5x and its dependent tile by 2.5x to 10x, and the 30x needs an INT4 layer at 0.46 V. And "`k` 0.7 to 1 at full rate" is not reached either: the top of the dense band is 0.7. The identity check: the tile's `k` is at best the ALU shadow's and the joules it forces are the same by construction (the record 20.3: 71.5 W against 82.8 at the knee for the same tile count); the chip's downside bet (its `k` floor) is 0.2 against the ALU's 0.3, so the tile remains the worse-or-equal lever, not the 4x to 30x worse lever the record wrote. The cost to the honest tiers is what kills it as content: the M5 Max at -35 percent of rate at 1,024 tiles per hash and -78 percent at 4,096 (measured, the record 20.2a), with no integer matrix path in Metal. Bit-exact on NVIDIA and the CPU (measured); the AMD WMMA layout unverified. **KILL as v6 content; KEEP the band correction (rank 3).** FP8, FP16 and BF16 tiles: the accumulation order inside a tile is unspecified by PTX, so not bit-exact across vendors or generations: KILL.
|
||||
|
||||
**The texture unit's filtered gather.** Measured 0.19 nJ per filtered fetch unlocked and 0.10 at the lock (the microbench `tex_linear_f32_256k`); a 9-bit fixed-point interpolator on NVIDIA (CUDA Programming Guide, claimed), vendor-specific fraction widths elsewhere: not bit-exact across the three vendors; a chip's interpolator is one 9-bit multiply-add, `k` far under 1. KILL. The point-sampled fetch is the L2 chase under another name (the same checksum, measured): the L2 lever, dead at `k` 0.1 to 0.3.
|
||||
|
||||
|
|
@ -176,71 +179,72 @@ Reading: nothing published since 2023 offers a lever outside the identity; every
|
|||
|
||||
## 3. This lane's own ideas
|
||||
|
||||
### 3.1 The second sector: W = 16 words, the whole item (KEEP, rank 1)
|
||||
### 3.1 The second sector: W = 16 words, the whole 64-byte item (KEEP, rank 1, conditional on one measurement)
|
||||
|
||||
The construction is the read-width branch's `w16` class (`docs/plans/read-width.md` section 1.5 and the fold of its section 1): a load reads the 64-byte-aligned item and folds its 16 words into the destination register with a rotate-multiply between words (`verify::fold_words`, mirrored in Metal, CUDA and OpenCL), so no function of the line alone replaces the dataset and the dependent chain is unchanged. Built 5 October; packs in `proto-cuda/packs-readwidth/`; three vector units and the cache and dataset checks passed on Metal, Apple OpenCL, the RTX 5090 (NVRTC) and the RX 9070 XT; the 2^24 fingerprint `e7c890445b47af60` equal on all four; the acceptance rule's rejection 0 to 14 of 60 as version 2; the CPU verifier 0.610 ms per 32 hashes against 0.604 (one item per lane, derived whole either way). What was never priced is the chip side.
|
||||
**The construction** is the read-width branch's `w64` class (`docs/plans/read-width.md` section 1 and 1.5, the fold of its section 1): a load reads the 64-byte-aligned item and folds its 16 words into the destination register with a rotate-multiply between words (`verify::fold_words`, mirrored in Metal, CUDA and OpenCL: four `uint4` loads from the aligned line, then fifteen rotl-mul-xor steps), so no function of the line alone replaces the dataset and the dependent chain is unchanged. Built 5 October; the pack in `proto-cuda/packs-readwidth/w64/`; three vector units and the cache and dataset checks passed on Metal, Apple OpenCL, the RTX 5090 (NVRTC) and the RX 9070 XT; the 2^24 fingerprint `836e56e7d496e980` equal on all four; the acceptance rule's rejection 0 to 14 of 60 as version 2; the CPU verifier 0.630 ms per 32 hashes against 0.604 (+4 percent; a lane's words lie in one item, derived whole either way). Two things were never priced: the chip side, and the card side at any occupancy but the exported one.
|
||||
|
||||
The chip side, on the model's own inputs (chip-model-v3 5.3: a random 32-byte read is 909 pJ of activation plus 256 bits at 4.5 pJ per bit of movement and I/O; HBM3 the same shape at O'Connor's 3.48 pJ per bit scaled by 4.12 / 6.25):
|
||||
**The chip side, on the model's own inputs** (chip-model-v3 5.3: a random 32-byte read is 909 pJ of activation plus 256 bits at 4.5 pJ per bit of movement and I/O; HBM3 the same shape at O'Connor's 3.48 pJ per bit scaled by 4.12 / 6.25; the SRAM die on floor lane 3's wire figure):
|
||||
|
||||
| Memory | Read energy at W = 1 (one 32 B sector) | At W = 16 (two sectors in the open row: one activate, two column accesses) | `E_mem` per hash, W = 1 to W = 16 | Rate | Label |
|
||||
|---|---|---|---|---|---|
|
||||
| GDDR7, 16 devices | 2.0 nJ | 3.2 nJ (+56 percent; 2.9 at O'Connor's 3.5 pJ per bit, +45) | 0.466 to **0.62** (+33 percent; 0.58 at the lower movement figure, +25); the 35 W of static and controller unchanged | activate-bound at 166 MH/s, unchanged; 1.36 TB/s of the board's 1.79 | modelled |
|
||||
| GDDR7, 16 devices | 2.0 nJ | 3.2 nJ (+56 percent; 2.9 at O'Connor's 3.5 pJ per bit, +45) | 0.466 to **0.62** (+33 percent; 0.58 at the lower movement figure); the 35 W of static and controller unchanged | activate-bound at 166 MH/s, unchanged; 1.36 TB/s of the board's 1.79 | modelled |
|
||||
| HBM3, one stack | 1.2 nJ | 1.8 nJ (+50 percent) | 0.321 to **0.40** (+24 percent) | 83.6 MH/s, unchanged; 0.68 TB/s of 0.82 | modelled |
|
||||
| HBM3, eight stacks | 1.2 | 1.8 | 0.262 to 0.33 (+26 percent) | unchanged | modelled |
|
||||
| The SRAM full store (lane B) | 1.0 nJ for a 64-byte line | 1.0 (the macro reads the line) | 0.14, unchanged | unchanged | modelled |
|
||||
| The custom HBM4E base die (lane B) | 0.9 to 1.0 | about 1.5 (the activation 909 pJ at 0.75 V scaled, the movement doubled) | 0.18 to about 0.27 (+50 percent) | unchanged | modelled, approximate |
|
||||
| The custom HBM4E base die (lane B) | 0.9 to 1.0 | about 1.5 | 0.18 to about 0.27 (+50 percent) | unchanged | modelled, approximate |
|
||||
| **The N2 SRAM full store (floor lane 3's rows)** | 0.25 nJ (80 bits of wire) | 0.88 nJ (560 bits) | **0.036 to 0.126 (+250 percent): 66x to 19x against the 5090's stock point** | power-bound: 8.3 to 2.4 GH/s per die | modelled (lane 3) |
|
||||
|
||||
The card side:
|
||||
A chip's controller merges the two sectors into one row cycle because it was built for this hash; a controller that did not (a second activate per sector) would pay 4.1 nJ and halve its rate, which no maker ships. The 5090's controller is the open question below.
|
||||
|
||||
| Card | Rate at W = 16 | DRAM joules | The card's own fabric | Energy per hash | Label |
|
||||
|---|---|---|---|---|---|
|
||||
| RTX 5090 at the 1,300 knee | +2.7 percent (139.8 against 136.1 MH/s unlocked, 5 October; the knee row unmeasured) | the same second sector: 1.15 nJ x 17.9 G reads per second = +20 W | one more 32-byte sector through L2 and the crossbar to the SM: about 0.3 to 0.5 nJ per sector (approximate; the measured L2 hit of 1.4 nJ per dependent read at the lock includes the wait), +6 to +10 W | 213 W to about 240 to 243 W at 130.7 MH/s: **1.67 to about 1.84 microjoules (+10 percent)**; the w16 job carried no power sampling (read-width.md 4.1), so this is modelled and is the one row owed | rate measured; watts modelled |
|
||||
| RTX 5090 unlocked | the same +2.7 | +20 W | +6 to +10 | 311 to about 340 W at 139.8 MH/s: 2.26 to 2.43 microjoules (+8 percent) | modelled |
|
||||
| RX 9070 XT | 17.90 against 18.15 MH/s (-1.4 percent, measured) | 0: every 4-byte read already costs a 64-byte line (measured, the probe) | 0 | +1.4 percent | measured |
|
||||
| Apple M5 Max | 28.26 against 27.74 (+1.9 percent, measured, under load) | 0: 64 B at the 4 B rate (3.51 G per second at both, Apple OpenCL probe) | 0 | about -2 percent | measured |
|
||||
| RTX 5070 Ti, 4070, 4060 Ti (32-byte sectors) | as the 5090 (approximate) | the same second sector | the same | about +10 percent (approximate) | approximate |
|
||||
**The card side, measured at two occupancies and unmeasured at the one that matters:**
|
||||
|
||||
The edge, every column (the 5090 at the knee; GDDR7 chip):
|
||||
|
||||
| Row | W = 1 (the record) | W = 16 | Change |
|
||||
| Card | At W = 16 (`w64`), unlocked, 5 October | What it says | Label |
|
||||
|---|---|---|---|
|
||||
| Zero shadow, GDDR7 | 3.59x | **3.0x** (1.84 / 0.62) | -16 percent |
|
||||
| Zero shadow, one HBM3 stack | 5.2x | 4.6x | -12 |
|
||||
| Class v4 shadow at `k = 0.3` | 3.52x | **3.05x** ((1.84 + 0.652) / (0.62 + 0.196)) | -13 |
|
||||
| at `k = 0.5` | 2.94x | **2.63x** | -11 |
|
||||
| at `k = 1` | 2.08x | 1.96x | -6 |
|
||||
| The 5090 unlocked, zero shadow | 4.85x | 3.9x | -20 |
|
||||
| Against the M5 Max's joule (0.78, unchanged) | 1.7x | 1.26x | -25 |
|
||||
| The SRAM full store, zero shadow (lane B's 17x against 2.40) | 17x | 17x against the stock point; 13x against the W = 16 5090 at 2.43 | 0 on the chip; the denominator moves |
|
||||
| RTX 5090, the hash at full occupancy (the exported kernel, one warp per block, the race's default) | 71.9 MH/s against 136.1 (**-47 percent**); 9.2 G reads per second, 589 GB/s of sectors | the row the read-width plan and class v6's layer 1 call "bandwidth-bound" and dead; but 589 GB/s is a third of the 1.79 TB/s stream, so the bind is the request path: two serialised sector misses per item, or a second activate at the controller when the second sector's request finds the row already closed (close-page on a random stream), or queueing at 4,080 resident warps | measured rate; the cause unmeasured |
|
||||
| RTX 5090, the 64-byte dependent-read probe, best over lanes in flight | **15.7 G reads per second** (9.1 at 4 M lanes) against 17.5 to 18.2 at 4 B: **-10 to -14 percent** | the same memory system, at its best lane count, does 64-byte dependent reads at 1.0 TB/s and within 14 percent of its 4-byte rate; the hash's -47 percent is therefore an occupancy and request-shape effect, not the pins | measured |
|
||||
| RTX 5090, a one-request form (PTX `ld.global.L2::64B` prefetch qualifier on the first word, sm_75 and later, so the L2 fetches the whole line as one transaction and the next three loads hit; or `cp.async.bulk` on sm_90 and later) | unmeasured | if the request reaches the controller as one 64-byte access the card pays one activate and two column reads as the chip does; the 4-byte rate would then hold (17.9 G x 64 B = 1.15 TB/s, 64 percent of the stream) | unmeasured; the worker's `variantSource` mechanism rewrites kernel text by anchor, as the hot-table `ldcs` and SM-sparse variants did |
|
||||
| RX 9070 XT | 17.59 against 18.15 MH/s (**-3 percent**); every 4-byte read already fetches a 64-byte line (the probe: 2.47 to 2.87 G lines per second at every width) | free | measured |
|
||||
| Apple M5 Max | 28.27 against 27.74 (**0**); 64 B at the 4 B rate (3.51 G per second at both, Apple OpenCL probe) | free | measured |
|
||||
| RTX 5070 Ti, 4070, 4060 Ti (32-byte sectors, the same controller family) | as the 5090, whichever row it turns out to be | | approximate |
|
||||
|
||||
The identity check, in the identity's own terms: `F` here is the second sector's joules on the card (about 1.5 nJ per read, 0.19 microjoules per hash at the knee) and `k` is the chip's cost for the same sector over the card's: 1.15 to 1.2 nJ over 1.45 to 1.65, **`k` about 0.75 to 0.8**, the highest `k` of any forcing work on the table, because 1.15 of it is the DRAM device's own movement at `k = 1` and only the card's fabric share is `k` under 1. The slope at `F = 0` from the record's section 2 (`(E_mem - k E_card) / E_mem^2`) at `k = 0.78`: -1.9x per microjoule, so 0.19 microjoules buys about 0.36x, and the full arithmetic above gives 0.6x because the chip's `E_mem` rises by the same joules it would have charged to `k F`. It stacks with the ALU shadow (the DRAM and the ALUs are different hardware; the fold adds about 5,760 ALU ops per hash, 6 percent of the shadow's count, inside the wait) and with the operating point.
|
||||
**The three outcomes, priced** (the 5090 at the 1,300 knee, class v3 213.0 W at 127.3 MH/s; the card's DRAM joules for the second sector +1.15 nJ per read, its fabric +0.3 to 0.5 nJ per sector, approximate; the GDDR7 chip at 0.62):
|
||||
|
||||
What a chip can do about it: nothing cheaper than paying. Items are pseudo-random (incompressible); the fold uses all 16 words with a state-dependent map (read-width.md section 1), so a pre-folded dataset does not exist; a narrower access atom does not exist on any DRAM (32 B is the minimum on GDDR7, HBM3, HBM4 and LPDDR6, JEDEC, claimed); spreading the words over devices multiplies the sectors; the HBM parts pay the movement at their own lower pJ per bit, which is the `E_mem` lever the model already carries. The one chip untouched is the SRAM full store (64-byte macro lines), which is lane 3's chip and not this lane's.
|
||||
| Outcome | The 5090's rate at W = 16 | Card watts | `E_card` | Edge at zero shadow | With the class v4 shadow at k 0.3 / 0.5 / 1 | Reading |
|
||||
|---|---|---|---|---|---|---|
|
||||
| A. The exported form at full occupancy (-47 percent, measured unlocked) | 67.5 MH/s | about 225 W (the per-second terms unchanged; +10 to 18 W of DRAM depending on whether the second sector re-activates) | 3.3 | **5.3x** | 4.5x / 4.1x / 3.1x | dead: the card halves, the chip does not |
|
||||
| B. The probe's best lane count (-13 percent, measured as a probe) | 110.8 | about 234 W | 2.11 | **3.4x** | 3.4x / 3.0x / 2.2x | not worth a class change: 0.2x at zero shadow |
|
||||
| C. A one-request form within 5 percent of the 4-byte rate (unmeasured) | 121 to 127 | about 241 W | 1.90 to 1.99 | **3.05x to 3.2x** | 3.1x / 2.7x / 2.0x | the record's best new lever: every column down 10 to 15 percent, and it stacks with the operating point and the shadow |
|
||||
| Pass line | 95 percent of the 4-byte rate at both clocks with the `.L2::64B` or tuned-occupancy variant | | | 3.2x or better | | W = 16 words into class v6 as layer 1's width floor |
|
||||
| Kill line | under 95 percent on every variant | | | | | the band stays {1, 4}; the row closes |
|
||||
|
||||
Cost to the honest tiers and the rule of 5 October: the 5090 class pays about 10 percent of energy per hash for no rate (a rig at the knee about +27 W per card); the 5070 Ti and the 12 GB and 8 GB NVIDIA tiers the same share (approximate); AMD and Apple nothing (their lines are 64 B or wider already); a pool user nothing; a node about 1 percent on the verifier; the public claim moves from "3.6x at the knee" to "3.0x" at zero shadow and from "2.1x at `k = 1`" to "2.0x". The known-failed case for the acceptance rule: the w16 class's census (0 to 14 of 60 rejected) and its F8 uniformity row are owed at 2^24 on 64 seeds under the sub-version 3 rule, since the 5 October census ran the version 2 rule.
|
||||
Against the honest denominators the lever is unconditional, because those cards pay nothing: the M5 Max against the GDDR7 chip 1.67x to 1.26x, against one HBM3 stack 2.4x to 1.95x, against the SRAM die (lane 3's 66x on the 5090 is about 21x on the Mac's 0.78) about 6.2x; the 9070 XT against the GDDR7 chip 22.7x to about 17.6x. The identity check, in the identity's terms: `F` is the second sector's joules on the card (about 1.5 nJ per read, 0.19 microjoules per hash at the knee) and `k` the chip's cost for the same sector over the card's: 1.15 nJ over 1.45 to 1.65, **`k` about 0.75 to 0.8 on GDDR7**, the highest `k` of any forcing work on the table, because 1.15 of it is the DRAM device's own movement at `k = 1` and only the card's fabric share is under 1; on the SRAM die `k` is about 0.4 (0.63 nJ over 1.5) but on a base of 0.036 microjoules, which is why that chip moves 3.5x. It stacks with the ALU shadow (different hardware; the fold adds about 5,760 ALU ops per hash, 6 percent of the shadow's count, inside the wait) and with the operating point. What a chip can do about it: nothing cheaper than paying. Items are pseudo-random (incompressible); the fold is dst-keyed and uses all 16 words, so a pre-folded dataset does not exist (lane 3 reads the same); a narrower access atom does not exist on any DRAM (32 B is the minimum on GDDR7, HBM3, HBM4 and LPDDR6, JEDEC, claimed); spreading the words over devices multiplies the sectors; the HBM parts pay the movement at their own lower pJ per bit, which is the `E_mem` lever the model already carries.
|
||||
|
||||
What it changes in class v6: layer 1's read-width band is {1, 4} words with the warrant "w16 moved the chip not at all" (the design's section 2 and lane B's method note). That warrant is the error this lane corrects; the band becomes the floor W = 16 (one item), and the lane A ranking's row 3 ("keep the read width out of the era draw: a draw over {4 B, 16 B} is harmless and worthless") keeps its first clause and loses its second. One measurement owed before the cut: the 5090's watts at w16 at the knee and unlocked (the pack exists; one PC job of four rows) and the M5 Max's power channels at w16.
|
||||
**Why the design's "never 16 words" and lane 3's "dead" are one measurement away from reversing.** Both rest on the -47 percent row, which was the exported kernel at the race's default occupancy with four separate 16-byte loads per item and no request-size hint, and the plan read 589 GB/s as a bandwidth bind. The same card's probe did 64-byte dependent reads at 1.0 TB/s. The job: the `w64` pack on PC 1's 5090 (the read-width exe or the installed worker with the pack), (i) the race's occupancy variants (warps per block 1, 2, 4, 8 and the persistent shapes) to find the probe's lane count inside the hash, (ii) a kernel-text variant with `ld.global.L2::64B.v4.u32` on the first of the four loads (one anchor rewrite in `variantSource`; the compile either takes the qualifier or drops the variant with "compile:" on the race line, which is the known-failed case), (iii) nvidia-smi at 1 Hz, unlocked and at the 1,300 lock, the `w4` pack beside it as the control, 250 batches per row; six to ten rows, under an hour. Rate and MH/W per row; the pass line above. The M5 Max's power channels at `w64` are the second row owed (its rate is measured free; its joules move by at most the DRAM's own movement term). A Windows build is not needed if the installed worker takes the pack; otherwise the cross-compile is on igneum-build-1, never this Mac.
|
||||
|
||||
### 3.2 Two items per read: W = 32, 128 bytes, one NVIDIA L2 line (KEEP for v7, rank 3)
|
||||
Cost to the honest tiers and the rule of 5 October, on outcome C: the 5090 class pays about 10 percent of energy per hash for no rate (a rig at the knee about +28 W per card); the 5070 Ti and the 12 GB and 8 GB NVIDIA tiers the same share (approximate); AMD and Apple nothing (measured); a pool user nothing; a node about 4 percent on the verifier (measured); the public claim moves from "3.6x at the knee" to "3.0x to 3.2x" at zero shadow and from "2.1x at `k = 1`" to "2.0x", and the SRAM die's line from "66x at the hash's width" to "19x". On outcome A nothing moves and the row closes. The known-failed case for the acceptance rule: the `w64` class's census (0 to 14 of 60 rejected under version 2) and its F8 uniformity row are owed at 2^24 on 64 seeds under the sub-version 3 rule before any cut.
|
||||
|
||||
What it changes in class v6 on a pass: layer 1's read-width row ("pinned at 4 words, 8 after the owed rows, 16 never") becomes "the floor W = 16 words"; lane A's ranking row 3 ("keep the read width out of the era draw: harmless and worthless") keeps its first clause and loses its second; lane 3's width table keeps its chip column and replaces its card column's last row.
|
||||
|
||||
### 3.2 Two items per read: W = 32, 128 bytes, one NVIDIA L2 line (KEEP for v7 behind rank 1, rank 4)
|
||||
|
||||
| Side | At W = 32 | Label |
|
||||
|---|---|---|
|
||||
| GDDR7 chip | one activate plus four column accesses: 0.909 + 4 x 1.15 = 5.5 nJ per read; 21.3 G activates x 128 B = 2.7 TB/s is over the board's 1.79, so the chip becomes bandwidth-bound at 14 G reads per second, 110 MH/s; `E_mem` = 128 x 5.5 nJ + 35 W / 110 M = 0.70 + 0.32 = **1.02 microjoules** (+120 percent); capex per MH/s +50 percent (USD 2.8 to 4.3) | modelled |
|
||||
| HBM3 one stack | 0.909 x 0.66 + 4 x 0.59 = 2.96 nJ; 10.7 G x 128 B = 1.37 TB/s over the stack's 0.82, bound at 6.4 G reads, 50 MH/s; `E_mem` = 0.38 + 14 W / 50 M = 0.66 microjoules (+105 percent) | modelled |
|
||||
| RTX 5090 at the knee | 17.9 G x 128 B = 2.3 TB/s over 1.79, so about 14 G reads per second, 110 MH/s (-20 percent; the w64 row measured the bind at 256 B: 71.9 MH/s, 0.58 share); DRAM +3 sectors x 1.15 nJ x 14 G = +48 W, fabric +15 W: about 276 W at 110 MH/s, **2.5 microjoules** (+50 percent) | modelled from the measured w64 bind and the rated peak; the w32 pack does not exist yet |
|
||||
| The SRAM die | about 1.5 nJ (1,070 bits of wire), 0.21 microjoules per hash: 11x against the 5090's stock point | modelled on lane 3's figure |
|
||||
| RTX 5090 at the knee | even in a one-request form the pins bind: 17.9 G x 128 B = 2.3 TB/s over 1.79, so at best 14 G reads per second, 110 MH/s (-20 percent); in the exported form far worse by the 64-byte row; DRAM +3 sectors x 1.15 nJ x 14 G = +48 W, fabric +15 W: about 276 W at 110 MH/s, 2.5 microjoules (+50 percent) at best | modelled from the rated peak; no `w32` pack exists |
|
||||
| Apple M5 Max | 3.5 G x 128 B = 448 GB/s against a measured 522 GB/s stream: 0 to -14 percent of rate (approximate; the Apple fetch granularity at 128 B unmeasured) | approximate |
|
||||
| RX 9070 XT | two 64-byte lines per read: 2.4 G x 128 B = 307 GB/s of 636; unmeasured whether the second line costs it a second fetch | unmeasured |
|
||||
| The edge, 5090 knee, GDDR7 | zero shadow 2.5 / 1.02 = **2.4x**; with the class v4 shadow at `k = 0.3`: (2.5 + 0.652) / (1.02 + 0.196) = **2.6x**; at `k = 0.5`: 2.3x; at `k = 1`: 1.9x | modelled |
|
||||
| The edge, 5090 knee, GDDR7, at the card's best case | zero shadow 2.5 / 1.02 = 2.4x; with the class v4 shadow at `k = 0.3`: 2.6x; at `k = 0.5`: 2.3x; at `k = 1`: 1.9x | modelled |
|
||||
|
||||
Reading: wider reads walk both sides into the regime where the DRAM's own joules are most of both bills, and there the edge tends toward the ratio of two bandwidth-bound machines (about 2.5x, modelled) at the price of the honest card's rate (the Ethash shape the design avoided on 5 October for that reason). W = 16 takes most of the per-joule gain at no rate; W = 32 takes the rest at a fifth of the 5090's rate and a 50 percent rise in the chip's capex per MH/s. It is a v7 candidate behind two measured rows (the 5090 and the M5 Max at a `w32` export, rate and watts), not a v6 change.
|
||||
Reading: wider reads walk both sides into the regime where the DRAM's own joules are most of both bills, and there the edge tends toward the ratio of two bandwidth-bound machines (about 2.5x, modelled) at the price of the honest card's rate (the Ethash shape the design avoided on 5 October for that reason). W = 16 takes most of the per-joule gain, possibly at no rate; W = 32 takes the rest at a fifth of the 5090's rate at best and a 50 percent rise in the chip's capex per MH/s. A v7 candidate only if rank 1 passes, behind a `w32` export and its rows; not a v6 change.
|
||||
|
||||
### 3.3 Independent chains per hash (memory-level parallelism as the card's lever) (KILL, one probe would reopen it)
|
||||
|
||||
The idea: four independent dependent chains of 32 reads per hash, folded at the end, so each resident hash keeps four reads in flight and an occupancy-bound card climbs toward its memory's activate ceiling while the chip (already at the ceiling) gains nothing; an `E_card` lever. The number: the 5090 reads 17.5 to 18.2 G per second at 4 B as "the best over lanes in flight" (the 5 October probe), which means more lanes did not help, so the bind is the memory system, not the in-flight count; the bound is the ceiling's 21.3 G, +17 to +22 percent, and the expectation is 0 to 5 percent (the microbench's `l2_indep4` against `l2_chase` read +2 percent on the L2-bound pattern). The 9070 XT caps at 2.63 G at 4,096 lanes (measured) with the same plateau shape. A class change for an expected few percent, which also cuts the sequential depth per hash to 32. KILL; the probe that would reopen it is one `dram_indep4_1g` row in the microbench (60 s on PC 1).
|
||||
The idea: four independent dependent chains of 32 reads per hash, folded at the end, so each resident hash keeps four reads in flight and an occupancy-bound card climbs toward its memory's activate ceiling while the chip (already at the ceiling) gains nothing; an `E_card` lever. The number: the 5090 reads 17.5 to 18.2 G per second at 4 B as "the best over lanes in flight" (the 5 October probe), which means more lanes did not help, so the bind is the memory system, not the in-flight count; the bound is the ceiling's 21.3 G, +17 to +22 percent, and the expectation is 0 to 5 percent (the microbench's `l2_indep4` against `l2_chase` read +2 percent on the L2-bound pattern). The 9070 XT caps at 2.63 G at 4,096 lanes (measured) with the same plateau shape. A class change for an expected few percent, which also cuts the sequential depth per hash to 32. KILL; the probe that would reopen it is one `dram_indep4_1g` row in the microbench (60 s on PC 1). (The 64-byte case of 3.1 is different: there the probe's own best row is 14 percent under the 4-byte row and the hash sits 47 percent under it, so the occupancy gap is measured, not conjectured.)
|
||||
|
||||
### 3.4 Row-straddling items (two activates per read) (KILL, dominated)
|
||||
|
||||
A 64-byte item laid across a row boundary costs the chip two activates: 2 x 0.909 + 2 x 1.15 = 4.1 nJ per read (+105 percent) and halves its activate-bound rate (83 MH/s), so `E_mem` = 128 x 4.1 nJ + 35 W / 83 M = 0.52 + 0.42 = 0.94 microjoules; the card's activate-bound rate halves too (about 83 MH/s if it sits at 82 percent of the same ceiling), at about +22 W: 2.95 microjoules. Edge 3.1x at a 39 percent rate loss, against the same-row second sector's 3.0x at no rate loss. Dominated by 3.1. KILL (the invention lane's 2.13 killed its 256-byte form on the w64 row; this is the 64-byte form, and it loses to the aligned item).
|
||||
A 64-byte item laid across a row boundary costs the chip two activates: 2 x 0.909 + 2 x 1.15 = 4.1 nJ per read (+105 percent) and halves its activate-bound rate (83 MH/s), so `E_mem` = 128 x 4.1 nJ + 35 W / 83 M = 0.52 + 0.42 = 0.94 microjoules; the card's activate-bound rate halves too (about 83 MH/s if it sits at 82 percent of the same ceiling), at about +22 W: 2.95 microjoules. Edge 3.1x at a 39 percent rate loss, against the aligned item's 3.0x to 3.2x at no rate loss on outcome C. Dominated by 3.1 whenever 3.1 passes, and worse than it on every outcome. KILL (the invention lane's 2.13 killed the 256-byte form on the w64x4 row; this is the 64-byte form).
|
||||
|
||||
### 3.5 Scratch state in DRAM (KILL)
|
||||
|
||||
|
|
@ -262,37 +266,40 @@ Both sides are capex-dominated (the record's 16.1: 7x to 10x their electricity p
|
|||
| Proof of latency | the chain checks values; the block is 1 s against a 0.4 microsecond round trip; the farm's node is local; a sequential prefix favours the chip about 4x | 2.5 million round trips per block | measured card latency, modelled chip |
|
||||
| The drawn address map | firmware; no prefetch exists for a dependent chain | 0 to the chip, 0.8 to 3.2 percent to the cards | measured spread |
|
||||
| Independent chains per hash | the 5090's read rate plateaus over lanes; the bind is the memory system | expected 0 to 5 percent, bound 22 | measured probe |
|
||||
| Row-straddling items | dominated by the aligned item | 3.1x at -39 percent of rate against 3.0x at 0 | modelled |
|
||||
| Row-straddling items | dominated by the aligned item | 3.1x at -39 percent of rate against 3.0x to 3.2x at 0 | modelled |
|
||||
| Scratch in DRAM | the card's scratch sits in L2; the chip's in SRAM | 465 KB against 96 MB | measured rows |
|
||||
| The whole item in the exported form (outcome A of 3.1) | the card halves, the chip does not | 5.3x | measured rate, modelled chip |
|
||||
|
||||
## 5. Consequences per tier (the standing rule of 5 October 2026)
|
||||
|
||||
| Tier | What this file means | What is being done |
|
||||
|---|---|---|
|
||||
| Home miner, one 8 GB NVIDIA card (32-byte sectors) | W = 16 costs it about 10 percent of energy per hash (approximate, the 5090's share scaled) for no rate; against the GDDR7 chip its row improves by the same 10 to 15 percent as the 5090's | the w16 watts row on the 5090 first; the 4060 Ti class scaled from it |
|
||||
| Home miner, one 8 GB NVIDIA card (32-byte sectors) | on a pass of rank 1, about 10 percent of energy per hash (approximate, the 5090's share scaled) for no rate, and its row against every chip improves by 10 to 15 percent; on a fail, nothing | the 5090 job first; the 4060 Ti class scaled from it |
|
||||
| One 12 GB card (4070, 5070) | the same share; the 4070 at its tune point about +8 W (approximate) | the same |
|
||||
| One 16 GB card (RX 9070 XT) | nothing: its 64-byte line already moves the whole item (measured 5 October); its row against the chip improves by the chip's +33 percent alone, 23x to about 17x | nothing to measure; the vendor-share metric |
|
||||
| One 24 or 32 GB card (5090 class) | +26 to 30 W at the knee (modelled), +2.7 percent of rate (measured); 3.6x to 3.0x at zero shadow, 2.9x to 2.6x with the shadow at `k = 0.5` | one PC job: the w16 pack at the 1,300 lock and unlocked, nvidia-smi at 1 Hz, four rows |
|
||||
| Apple (M5 Max and the unified tiers) | nothing: 64 B at the 4 B rate (measured); the honest best per joule improves against every DRAM chip by the chip's cost alone, 1.7x to about 1.3x on GDDR7 | the M5 Max power channels at w16 (one Mac measurement under the measure lock, deferred until the Mac rule allows) |
|
||||
| A rig | +12 percent of watts on NVIDIA cards for no rate; the chip's bill per MH/s-hour up a third on energy | design 1 (the knee lock) as the default, then W = 16 |
|
||||
| One 16 GB card (RX 9070 XT) | nothing in watts or rate at any outcome (-3 percent of rate measured at 64 B); against the GDDR7 chip 22.7x to about 17.6x and against the SRAM die 3.5x better, on a pass | nothing to measure; the vendor-share metric |
|
||||
| One 24 or 32 GB card (5090 class) | the whole question: -47 percent of rate as exported (measured), -13 percent at the probe's best lane count (measured), unmeasured in a one-request form; on a pass +28 W at the knee (modelled) and 3.6x to 3.0x at zero shadow | one PC 1 job: the `w64` pack, occupancy variants, the `.L2::64B` variant, both clocks, watts at 1 Hz, the `w4` control; under an hour |
|
||||
| Apple (M5 Max and the unified tiers) | nothing at any outcome (0 at 64 B, measured); the honest best per joule improves against every chip by the chip's cost alone on a pass: 1.7x to 1.26x on GDDR7, about 21x to 6.2x against the SRAM die (modelled) | the M5 Max power channels at `w64` (one Mac measurement under the measure lock, when the Mac rule allows) |
|
||||
| A rig | on a pass, +12 percent of watts on NVIDIA cards for no rate; the chip's bill per MH/s-hour up a third on energy | design 1 (the knee lock) as the default, then W = 16 |
|
||||
| A pool user | nothing changes in shares; the class change is announced by the 95 percent signal | |
|
||||
| A node operator (the verifier) | +1 percent (0.610 against 0.604 ms per 32 hashes, measured); W = 32 about +4 percent | nothing |
|
||||
| A chip | GDDR7 +33 percent of energy per hash, HBM3 +24, the base die +50, the SRAM store 0; the cheapest project is a 12 nm GDDR6 part at USD 20 M to 30 M, not a 28 nm one at 5 M | the chip model's rows 5.4 and 5.6 and the record's 16.2 corrected (owed to the Counter lane) |
|
||||
| The public claim | "3.6x at the knee, 2.1x at `k = 1`" becomes "3.0x and 2.0x" on W = 16 (modelled card side until the job); the tensor `k` sentence loses "4x to 30x"; the Ethash precedent gains the 13x row | the Counter lane's texts, after the 20:00 word |
|
||||
| A node operator (the verifier) | +4 percent (0.630 against 0.604 ms per 32 hashes, measured); W = 32 about the same again | nothing |
|
||||
| A chip | GDDR7 +33 percent of energy per hash, HBM3 +24, the base die +50, the SRAM die +250 (all modelled) on a pass; the cheapest project is a 12 nm GDDR6 part at USD 20 M to 30 M, not a 28 nm one at 5 M, on any outcome | the chip model's rows 5.4 and 5.6 and the record's 16.2 corrected (owed to the Counter lane) |
|
||||
| The public claim | on a pass "3.6x at the knee, 2.1x at `k = 1`" becomes "3.0x to 3.2x and 2.0x" and the SRAM die's "66x" becomes "19x"; the tensor `k` sentence loses "4x to 30x" on any outcome; the Ethash precedent gains the 13x row | the Counter lane's texts, after the 20:00 word |
|
||||
|
||||
## 6. Unverified and owed
|
||||
|
||||
- The 5090's watts at w16 are modelled (+26 to 30 W at the knee); the 5 October job carried no power sampling. One PC job on the existing pack decides the card side of rank 1. The M5 Max's power channels at w16 are unmeasured (the rate is).
|
||||
- The chip side of rank 1 is the model's own data-movement term applied twice (4.5 pJ per bit, Micron, claimed, streaming; O'Connor's 3.48 for HBM2); no one has measured a second-sector column read's energy on GDDR7 (section 2.6: the public record has no random-read joule on any current part).
|
||||
- The 5090's rate at 64 bytes in a one-request form and at its best occupancy inside the hash: unmeasured; the two measured rows (-47 percent as exported, -13 percent as a probe) bracket it. Its watts at 64 bytes at either clock: unmeasured (the 5 October job carried no power sampling). This is the one job this lane asks for; the pass and kill lines are stated in 3.1.
|
||||
- Whether the 5090's controller re-activates for the second sector (close-page on a random stream) or the loss is queueing: the same job reads it from the rate per variant.
|
||||
- The M5 Max's power channels at `w64` are unmeasured (the rate is).
|
||||
- The chip side of rank 1 is the model's own data-movement term applied twice (4.5 pJ per bit, Micron, claimed, streaming; O'Connor's 3.48 for HBM2) and lane 3's wire figure for the SRAM die; no one has measured a second-sector column read's energy on GDDR7 (section 2.6: the public record has no random-read joule on any current part).
|
||||
- The fabric share of the card's second sector (0.3 to 0.5 nJ) is approximate.
|
||||
- The W = 32 rows rest on the rated 1.79 TB/s and the measured 256-byte bind; no w32 pack exists.
|
||||
- The W = 32 rows rest on the rated 1.79 TB/s; no `w32` pack exists.
|
||||
- The tensor chip figures are vendor specifications at TDP (claimed), the M4 ANE and the H800 loop measured; the RT figures are claims and one measured RTX 2080 row; the PHY node facts are vendor IP pages read today.
|
||||
- The project-cost corrections use the history's figures (a 7 nm-class startup project USD 50 M to 75 M, claimed; the 12 nm figure scaled, approximate) and the mission lane's cap model.
|
||||
- The w16 class's census under the sub-version 3 rule and its F8 row at 2^24 on 64 seeds are owed before any cut (the 5 October census ran version 2's rule).
|
||||
- The `w64` class's census under the sub-version 3 rule and its F8 row at 2^24 on 64 seeds are owed before any cut (the 5 October census ran version 2's rule).
|
||||
- Nothing ran on the Mac, on a box or on a PC for this file.
|
||||
|
||||
## 7. Sources
|
||||
|
||||
Internal: `docs/analysis/counter-asic-4-research.md` (sections 2, 4, 5, 8, 15.1a, 16, 20), `docs/analysis/chip-model-v3.md` (5.1, 5.3, 5.4, 5.6, 5.10 to 5.12), `docs/design/class-v6-rotating-family.md` (sections 2, 7, 7a, 7c), `docs/analysis/class-v6/hardware-future.md` at 7618e729 (sections 0, 1, 2, 4.9, 5), `docs/analysis/class-v6/history.md` (sections 3 and 5), `docs/analysis/class-v6/invention.md` on `class-v6-invention` (sections 1 and 2), `docs/plans/read-width.md` (sections 1, 1.5, 4.1), `docs/bench-log.md` (the 5 October read-width entry), `docs/analysis/scratch-soundness.md`, spec 01 (1.5, the item table).
|
||||
Internal: `docs/analysis/counter-asic-4-research.md` (sections 2, 4, 5, 8, 15.1a, 16, 20), `docs/analysis/chip-model-v3.md` (5.1, 5.3, 5.4, 5.6, 5.10 to 5.12), `docs/design/class-v6-rotating-family.md` (sections 2, 7, 7a, 7c, 10 and 10.3), `docs/analysis/class-v6/hardware-future.md` at 7618e729 (sections 0, 1, 2, 4.9, 5), `docs/analysis/class-v6/history.md` (sections 3 and 5), `docs/analysis/class-v6/invention.md` on `class-v6-invention` (sections 1 and 2), `docs/analysis/class-v6/floor/sram-and-floor.md` on `class-v6-floor-sram` as carried in section 10.3, `docs/plans/read-width.md` (sections 1, 1.5, 4.1), `docs/bench-log.md` (the 5 October read-width entry: the probes and the hash rates), `proto-cuda/packs-readwidth/w64/kernel_bound.cu` (the four `uint4` loads per item), `proto-cuda/nvrtc/worker.cpp` (`variantSource`), `docs/analysis/scratch-soundness.md`, spec 01 (1.5, the item table).
|
||||
|
||||
External, all read 8 October 2026 by this lane's three research sub-agents (tensor, RT, literature) and labelled where carried: the NVIDIA RTX 5090 page and the Blackwell architecture whitepaper; the Ada whitepaper; arXiv 2501.12084v2 (Hopper tensor-core microbenchmarks); Lenovo LP2226 (B200); the AMD MI300X data sheet and MI355X notes; the RX 9070 XT page; Qualcomm's AI 100 Ultra brief; eeNews Europe (MTIA v2); AnandTech 21342 (Gaudi 3); the Hot Chips 2025 Ironwood slides; arXiv 2608.28048 (Trainium2); Tenstorrent's Blackhole specification; the M4 ANE measurement (maderix.substack.com); cnx-software (Hailo-8); design-reuse 52539 (Untether); hackster.io (Axelera); NVIDIA Research's VSQ accelerator (JSSC 2023); Dally, Hot Chips 2023 keynote, slides 12 and 24; the Vulkan ray traversal chapter; the DXR specification; the NVIDIA developer forum thread 309730 (14 October 2024); the Blender Cycles source note; the Mach-RT TVCG preprint; Hot Chips 31 (Turing); the AMD Hot Chips 2025 RDNA 4 slides; embedded.com (RayCore); AnandTech 7870 (GR6500); arXiv 2409.06000 (RayFlex); the TRaX CGI 2018 study; Ylitie et al., HPG 2017; Chou et al., MICRO 2023; par.nsf.gov 10156963 (Choe et al., MEMSYS 2019); arXiv 2508.06795; SemiAnalysis "The Memory Wall"; the SK hynix ISSCC 2024 GDDR7 paper (ResearchGate 378947349); arXiv 2410.12990; arXiv 2405.06081; asicminervalue (iPollo V2H, V1, A11 Pro); kryptex (Jasminer X16-Q, FishHash); whattomine (the 5090 on Etchash); monero-project issue 10270; arXiv 2512.01437v2 and decrypt.co 363018 (Qubic); Business Wire 20240925780123 (Cadence GDDR7 on N3) and 20221115005674 (GDDR6 at 12 nm); innosilicon.com GDDR7; kurnal-insights (GB202 die); TrendForce 24 September 2026 (the 2 GB GDDR7 part); arXiv 2403.13230 (BFT-PoLoc); eprint 2026/694 (unread, refused).
|
||||
External, all read 8 October 2026 by this lane's three research sub-agents (tensor, RT, literature) and labelled where carried: the NVIDIA RTX 5090 page and the Blackwell architecture whitepaper; the Ada whitepaper; arXiv 2501.12084v2 (Hopper tensor-core microbenchmarks); Lenovo LP2226 (B200); the AMD MI300X data sheet and MI355X notes; the RX 9070 XT page; Qualcomm's AI 100 Ultra brief; eeNews Europe (MTIA v2); AnandTech 21342 (Gaudi 3); the Hot Chips 2025 Ironwood slides; arXiv 2608.28048 (Trainium2); Tenstorrent's Blackhole specification; the M4 ANE measurement (maderix.substack.com); cnx-software (Hailo-8); design-reuse 52539 (Untether); hackster.io (Axelera); NVIDIA Research's VSQ accelerator (JSSC 2023); Dally, Hot Chips 2023 keynote, slides 12 and 24; the Vulkan ray traversal chapter; the DXR specification; the NVIDIA developer forum thread 309730 (14 October 2024); the Blender Cycles source note; the Mach-RT TVCG preprint; Hot Chips 31 (Turing); the AMD Hot Chips 2025 RDNA 4 slides; embedded.com (RayCore); AnandTech 7870 (GR6500); arXiv 2409.06000 (RayFlex); the TRaX CGI 2018 study; Ylitie et al., HPG 2017; Chou et al., MICRO 2023; par.nsf.gov 10156963 (Choe et al., MEMSYS 2019); arXiv 2508.06795; SemiAnalysis "The Memory Wall"; the SK hynix ISSCC 2024 GDDR7 paper (ResearchGate 378947349); arXiv 2410.12990; arXiv 2405.06081; asicminervalue (iPollo V2H, V1, A11 Pro); kryptex (Jasminer X16-Q, FishHash); whattomine (the 5090 on Etchash); monero-project issue 10270; arXiv 2512.01437v2 and decrypt.co 363018 (Qubic); Business Wire 20240925780123 (Cadence GDDR7 on N3) and 20221115005674 (GDDR6 at 12 nm); innosilicon.com GDDR7; kurnal-insights (GB202 die); TrendForce 24 September 2026 (the 2 GB GDDR7 part); arXiv 2403.13230 (BFT-PoLoc); eprint 2026/694 (unread, refused); the PTX ISA (`ld` with the `.L2::64B` prefetch-size qualifier, sm_75 and later).
|
||||
|
|
|
|||
Loading…
Reference in a new issue