Class v6 floor lane 4: the fold-any-k line (E_mem per chip, the k lane's 1.1 pJ column priced)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay ofbae817f11(6d0498b7e6) for the box mirror master
This commit is contained in:
parent
41914adee1
commit
babbbf2a9e
1 changed files with 8 additions and 0 deletions
|
|
@ -124,6 +124,14 @@ Max; the wall more); the NVIDIA and AMD rows are whole-card.
|
|||
| Apple M4 Pro | Apple | 1.65 | estimated | 2.1x / 1.5x | 2.6x / 1.7x | 3.6x / 2.1x | 256-bit LPDDR5X, a 20-core GPU: about 13.5 MH/s; the shadow's per-hash premium does not shrink with the part, so about 1.65 at the meter |
|
||||
| Apple M3 Max | Apple | 1.75 | estimated | 2.2x / 1.6x | 2.7x / 1.8x | 3.8x / 2.2x | 512-bit LPDDR5-6400, a 40-core GPU on N3B: about 24 MH/s; about 1.75 at the meter |
|
||||
|
||||
To fold any other chip core figure (the research lane's convention, 13:5x): the edge on a row is the card's class v5
|
||||
microjoules over `E_mem + 101,170 x c`, with `E_mem` 0.466 (GDDR7 board), 0.321 (one HBM3 stack), 0.14 (the N2 SRAM
|
||||
die), 0.18 (the custom HBM4E base die) and `c` the core's picojoules per forced op; at the k lane's synthesised
|
||||
sequencer-core floor of 1.1 pJ per op at N3 (its per-unit figure before fetch, decode and the register file, due 15:00)
|
||||
the chip's class v5 energy reads GDDR7 0.577, HBM3 0.432, SRAM 0.251, so the 5080's floor of 2.10 gives 3.6x, 4.9x and
|
||||
8.4x and the M5 Max's 1.43 gives 2.5x, 3.3x and 5.7x; that figure is a floor on the core and the row at it is the
|
||||
pessimistic column.
|
||||
|
||||
What the table says:
|
||||
|
||||
1. The honest floor per tier, measured: 32 GB 2.33 to 2.37 (the 5090 at its knee); 16 GB 2.06 to 2.10 (the 5080 at
|
||||
|
|
|
|||
Loading…
Reference in a new issue