Class v6: 7a the hardware future beside the four layers (lane B's k bands: HBM4, the custom base die, DRAM on logic, the N2 SRAM store at 17x; the three changes taken into layers 1 and 2 and the denominator), section 0's one line extended
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
837576edc2
commit
d8b2e2ee68
1 changed files with 22 additions and 1 deletions
|
|
@ -15,7 +15,7 @@ The honest frame first, from last night's measured close (`docs/analysis/counter
|
|||
|
||||
Per tier against the stored-dataset chip (GDDR7, 0.466 microjoules per hash, modelled), at stock: the 5090 at 2.26 microjoules (class v3) and 3.36 (class v4) reads 4.9x and 2.2x at `k = 1`; the 5070 Ti at 1.79 and 2.84 microjoules reads 3.8x and 1.9x at `k = 1` (a card whose stock point is nearer its knee); the M5 Max at 0.78 and 1.40 (GPU and DRAM channels) reads 1.7x and 0.9x. The layers do not move these; the operating point and the shadow do.
|
||||
|
||||
What a fully general chip still gets, in one line: the stored-dataset chip with a programmable core at the top of the band keeps 3.6x at zero premium and 2.1x at `k = 1` on a 5090 at its knee (section 2 of the research file), and the four layers cost it about USD 30 to 60 per chip of extra silicon and a project that must be N5-class from the first tape-out (the USD 100 M break-even cap of the mission lane's model).
|
||||
What a fully general chip still gets, in one line: the stored-dataset chip with a programmable core at the top of the band keeps 3.6x at zero premium and 2.1x at `k = 1` on a 5090 at its knee (section 2 of the research file), and the four layers cost it about USD 30 to 60 per chip of extra silicon and a project that must be N5-class from the first tape-out (the USD 100 M break-even cap of the mission lane's model); the chips of 2027 on (lane B, section 7a: a custom HBM4E base die, DRAM on logic, a 2 GiB SRAM store on one N2 reticle at 17x per joule at zero shadow and USD 0.25 to 0.4 per MH/s) are touched by none of the four layers except layer 2's capex, and against the SRAM chip the public 2x needs the honest card's whole latency shadow at a core as good as a GPU lane.
|
||||
|
||||
## 1. Where the four layers sit in the design as it stands
|
||||
|
||||
|
|
@ -146,6 +146,27 @@ The identity of the research file's section 2, with the night's measured rows: `
|
|||
| 4 | none | none | the safety of 1 and 3, not a lever |
|
||||
| All four, against a 5090 at its knee | **3.6x at zero premium, 2.1x at `k = 1` with the class v4 shadow (measured card, modelled chip), unchanged** | about +USD 30 to 60 per chip; the N5 project forced | "useless as soon as it dropped" is true of a fixed-function ASIC and false of the chip anyone builds; what holds the general chip is the price per joule of the honest card's own operating point and the shadow's premium, as last night's close said |
|
||||
|
||||
## 7a. The hardware future beside the four layers (lane B, `docs/analysis/class-v6/hardware-future.md` on master at 34f63b3c, 10:25 UTC; every chip figure modelled on chip-model-v3's method; the first cut, the full report by 09:00 UK tomorrow)
|
||||
|
||||
The chip rows of sections 2 to 7 price the GDDR7 board and one HBM3 stack. Lane B's table carries the memory systems a chip could buy from 2027 on, per joule at zero shadow against the 5090's 2.40 microjoules (the M5 Max's 0.78 in brackets) and with the class v4 shadow at the premium F = 1.10 at `k = 0.5 / k = 1`:
|
||||
|
||||
| Memory system (when) | Energy per random read, modelled | Edge at zero shadow | With the shadow, k 0.5 / 1 | What the four layers do to it |
|
||||
|---|---|---|---|---|
|
||||
| GDDR7 board, 28 nm controller (the record) | 2.0 nJ | 5.1x (1.7x) | 3.3x / 2.1x | layer 1 forces the N5 core (the project to USD 30 M); nothing else |
|
||||
| HBM3E, one stack | 1.2 nJ | 7.5x (2.4x) | 3.8x / 2.4x | the same |
|
||||
| HBM4, one stack, 2027 to 2028 | 1.0 to 1.1 nJ | 4.5x to 11x (the activate ceiling doubles if tFAW is per channel, unmeasured) | 4.4x / 2.6x | the same |
|
||||
| A custom HBM4E base die (controller and PHY in the stack, N3P; a non-hyperscaler from about 2028) | 0.9 to 1.0 nJ | 6.5x to 14x (2.1x to 4.4x) | 4.4x / 2.6x | beats all four layers: untouched by any of them |
|
||||
| DRAM on logic (FGDRAM-class 256-byte rows, hybrid bonded), 2029 to 2031 | 0.5 to 0.7 nJ | 12x to 20x (5.2x) | 5.1x / 2.8x | beats all four layers |
|
||||
| **An SRAM full store on one N2 reticle, 2 GiB (452 mm^2 of macro, USD 400 to 600 of silicon)** | 1.0 nJ (0.5 to 2.0) | power-bound at about 2,100 MH/s per die at 300 W: 17x (8x to 30x; 5.6x against the M5 Max), USD 0.25 to 0.4 per MH/s | 4.8x / 2.7x; at the 5090's whole latency shadow (F about 2.0) 3.7x / 2.0x | beats layers 1, 3 and 4; **layer 2 moves its capex, not its joules** (two dies at 4 GiB 15x and USD 1,000; four at 8 GiB 13x and USD 2,000 to 2,500) |
|
||||
| LPDDR6 controller chip | 1.5 to 2.0 nJ | 4x to 5x (1.3x to 1.6x) | | the honest SoC tier's own memory |
|
||||
| Per-bank PIM, UPMEM, an FPGA with HBM2e, wafer-scale, CXL, optical | | under 1x or no path (PIM is blind: 1.6 percent of reads in-bank at 2 GiB on a 32 MB bank, 0.4 at 8 GiB) | | |
|
||||
|
||||
What this changes in the design, taken into the layers:
|
||||
|
||||
1. **Layer 1's program-length band is sized against the N2 SRAM chip, not the GDDR7 board.** The lower bound is the length that holds the record's 2.1x today (rung 0, 102,100 ops); the upper bound is the honest cards' full latency shadow (the 5090 about 330,000 ops unlocked and 150,000 at the lock, the M5 Max 290,000 by its budget and 130,000 by the 5 percent rule, the 9070 XT 650,000, all measured or budgeted in the ladder's rows), re-based at each family epoch; N still moves by the ladder's signal inside that band, never by an unconditional draw. The public "2x" is not reachable against the SRAM chip at any `k` under 1 (2.0x needs the 5090's whole shadow at `k = 1`), which is the honest line section 0 now carries.
|
||||
2. **Layer 2's floor is a card-lifetime decision, not a chip lever**: the SRAM chip pays capex, not joules, for a larger dataset (USD 400 to 600 per 2 GiB die), and the real processing-near-memory threat is the custom base die, which no layer touches; the brake until about 2028 is HBM allocation and price (claimed: Samsung asking USD 4 to 5 per Gbit for HBM4 against 1.5 for HBM3E, 2 October 2026). The spec's schedule stands; the dataset stays inside 16 GB unified memory (8 GiB at the top of the schedule does).
|
||||
3. **The denominator: resistance is stated against the best honest joule** (the M5 Max at 0.78 microjoules, the 5090 at its knee at 1.67), where every chip edge is 2x to 3x smaller than against the 5090's stock point; section 0's per-tier line already reads so.
|
||||
|
||||
## 8. Unverified and owed
|
||||
|
||||
- Lane D's family harness (the family gate's measured coverage: the 10,000-era random stratum and the three corner strata, the lossy cap, width 4, shape 64) lands in `family-gate.md` by 17:00 UK and is layer 4's coverage table; this document's layer 4 cites it where it is named and does not restate it.
|
||||
|
|
|
|||
Loading…
Reference in a new issue