Class v6 10.0e: the node sentence (2.0x own node, 2.4x a node ahead, 2.8x two ahead; the honest tier moves nodes with every GPU generation, a chip re-tapes-out)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
58c22e9e1e
commit
773c4e0bcb
1 changed files with 2 additions and 2 deletions
|
|
@ -305,7 +305,7 @@ The honest floor in one line, on the synthesised core: against the chip anyone c
|
|||
|
||||
The capex wall's two thresholds (lane 3, 10.3, re-folded 13:18 UK on lane 5's corrected project floor of USD 20 to 75 M for the cheapest DRAM-board chip): no rational chip project of any kind below about USD 23 M a year of miner revenue (IGN 0.03 at launch emission, USD 62 K a day, a cap of about USD 60 M); every DRAM-board project at a third of the network above about USD 340 M a year (IGN 0.44, USD 0.93 M a day, a cap of about USD 0.9 B), where the SRAM project also starts. The project cost moves the threshold 5x, the chip's edge 1.4x.
|
||||
|
||||
The served line in one sentence, on the synthesised core with the 64-register window (10.0e), in the two units: a chip that stores the dataset reaches about 2.2x to 2.4x per joule against the honest NVIDIA tiers at their knee (the 5080 2.2x, the 5090 2.4x; 2.8x on the 5090 without the window) and under 2x only against the Apple tier (the M5 Max 1.5x), about 4x and 2.6x if it is built on an N2 SRAM store for USD 100 M or more, and under 1x per hash over its 180-day class life only above about USD 300 M a year of miner revenue (10.0a), with Monero's measured 1.0x to 1.5x beside it; no rotating layer moves the per-joule figures, and layer 3's class life is what sets the per-hash one; the layers decide which chip can be built and how long its tape-out lives. The worst case beside it, if a chip's core costs no more than its bare units (the floor k of 10.2): 4x and 2.4x against the DRAM board, 14x and 9x against the SRAM die (lane 3's line: 6x to 10x at the full shadow, never under 2x). In the other convention, the card's whole 330,000-op latency shadow priced on the chip at the synthesised core (lane 3's 10.3 table): the SRAM die 3.5x against a 5090 at stock and 2.2x at its knee (4.8x and 3.0x with a 2 nm core), the DRAM board 2.6x at stock. The record's claimed band (k 0.3 to 0.8) is confirmed from the chip side by the core row (0.56 at the lock), so the served 2.1x at k = 1 is the conservative end and 2.8x the measured-core reading.
|
||||
The served line in one sentence, on the synthesised core with the 64-register window (10.0e), in the two units: a chip that stores the dataset reaches about 2.2x to 2.4x per joule against the honest NVIDIA tiers at their knee (the 5080 2.2x, the 5090 2.4x with a chip a node ahead; 2.0x against a chip on the GPU's own node, 2.8x two nodes ahead; 2.8x on the 5090 without the window) and under 2x only against the Apple tier (the M5 Max 1.5x), about 4x and 2.6x if it is built on an N2 SRAM store for USD 100 M or more, and under 1x per hash over its 180-day class life only above about USD 300 M a year of miner revenue (10.0a), with Monero's measured 1.0x to 1.5x beside it; no rotating layer moves the per-joule figures, and layer 3's class life is what sets the per-hash one; the layers decide which chip can be built and how long its tape-out lives. The worst case beside it, if a chip's core costs no more than its bare units (the floor k of 10.2): 4x and 2.4x against the DRAM board, 14x and 9x against the SRAM die (lane 3's line: 6x to 10x at the full shadow, never under 2x). In the other convention, the card's whole 330,000-op latency shadow priced on the chip at the synthesised core (lane 3's 10.3 table): the SRAM die 3.5x against a 5090 at stock and 2.2x at its knee (4.8x and 3.0x with a 2 nm core), the DRAM board 2.6x at stock. The record's claimed band (k 0.3 to 0.8) is confirmed from the chip side by the core row (0.56 at the lock), so the served 2.1x at k = 1 is the conservative end and 2.8x the measured-core reading.
|
||||
|
||||
The three class v6 changes it implies: (1) the op mix stays class v4's with the lossy families capped at their base (lane D: 0 of 3,000 eras exhausted under the band; the k lane: mulhi is the card's worst lever by 6x, so no re-weight helps the card); (2) no SM-sparse default (measured dead on the 5090, 4090 and H100 at 1.3 to 3.9 percent at best; `--sm-sparse auto` ships off by default, on in the Efficiency and Balanced tiers at its measured 1 to 2.5 percent, 10.1 item 6; the per-watt gift of the hot table is the chip's, 20.3c); (3) the dataset schedule 5.5 / 8.5 / 11.5 GiB with the read width pinned at 4 words; and a fourth the sweep adds, (4) the 64-register window per lane as the core shape (k 0.78 at the lock against 0.56; 10.0c, the edge table 10.0e), its GPU side modelled until the generator carries a 64-entry init and fold rule (a half-day line) and a pack is measured (the chip's ticket USD 1,500 / 2,500 / 3,000; a tuned 5090 pays 1 to 5 percent per step; about a quarter of today's cards by count per step).
|
||||
|
||||
|
|
@ -370,7 +370,7 @@ The node column (the k lane, 15:0x UK; the factors claimed from TSMC's per-node
|
|||
| the 64-register file | 9.7 / 6.8 / 4.9 / 3.5 | 1.09 / 0.78 / 0.56 | **2.0x** / 2.4x / 2.8x | synthesised; scaling claimed |
|
||||
| all four together | ABC on build-3, about 16:00 | | | owed |
|
||||
|
||||
In one line: of the 2.8x at k 0.56, the N5-to-N3 node step is 0.4x (2.4x node-for-node, a factor 1.17, claimed); the rest is the memory system's 3.6x at zero shadow less what the shadow takes back on the card's own node, which is the design. **Node-for-node the base core is k 0.78 and the window core 1.09, so with the window a chip on the GPU's own node reaches 2.0x against a 5090 at its knee and 1.8x against the 5080; what a chip project buys back with N3 is 0.4x and with N2 0.8x.** This is the honest form of the served "with a core as good as a GPU lane": on the same node the window core is one.
|
||||
In one line: of the 2.8x at k 0.56, the N5-to-N3 node step is 0.4x (2.4x node-for-node, a factor 1.17, claimed); the rest is the memory system's 3.6x at zero shadow less what the shadow takes back on the card's own node, which is the design. **Node-for-node the base core is k 0.78 and the window core 1.09, so with the window a chip on the GPU's own node reaches 2.0x against a 5090 at its knee and 1.8x against the 5080; what a chip project buys back with N3 is 0.4x and with N2 0.8x.** This is the honest form of the served "with a core as good as a GPU lane": on the same node the window core is one. **The close's sentence (the coordinator's order, 15:1x UK): 2.0x against a chip on the GPU's own node, 2.4x a node ahead, 2.8x two nodes ahead; and the honest tier moves to the next node with every GPU generation while a chip must re-tape-out, so the node step a chip project buys is on loan until the next card ships (the 5090 is N5 class in 2025; its successor on N3 takes the 0.4x back).**
|
||||
|
||||
#### 10.0d The rotation schedule the close adopts (the rotation lane, `docs/design/class-rotation-four-layers.md` on class-v6-rotation at bd43f808, build-3, gate green, 14:4x UK; one line per layer; both of this document's constraints held: the 180-day family epoch not shorter, W = 4 not drawn)
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue