Class v6 floor lane 4: the fold-any-k line (E_mem per chip, the k lane's 1.1 pJ column priced)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 5e77445c7 (2a8572e489) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 12:25:35 +00:00
parent 43c94502ad
commit 75618b8a50

View file

@ -124,6 +124,14 @@ Max; the wall more); the NVIDIA and AMD rows are whole-card.
| Apple M4 Pro | Apple | 1.65 | estimated | 2.1x / 1.5x | 2.6x / 1.7x | 3.6x / 2.1x | 256-bit LPDDR5X, a 20-core GPU: about 13.5 MH/s; the shadow's per-hash premium does not shrink with the part, so about 1.65 at the meter |
| Apple M3 Max | Apple | 1.75 | estimated | 2.2x / 1.6x | 2.7x / 1.8x | 3.8x / 2.2x | 512-bit LPDDR5-6400, a 40-core GPU on N3B: about 24 MH/s; about 1.75 at the meter |
To fold any other chip core figure (the research lane's convention, 13:5x): the edge on a row is the card's class v5
microjoules over `E_mem + 101,170 x c`, with `E_mem` 0.466 (GDDR7 board), 0.321 (one HBM3 stack), 0.14 (the N2 SRAM
die), 0.18 (the custom HBM4E base die) and `c` the core's picojoules per forced op; at the k lane's synthesised
sequencer-core floor of 1.1 pJ per op at N3 (its per-unit figure before fetch, decode and the register file, due 15:00)
the chip's class v5 energy reads GDDR7 0.577, HBM3 0.432, SRAM 0.251, so the 5080's floor of 2.10 gives 3.6x, 4.9x and
8.4x and the M5 Max's 1.43 gives 2.5x, 3.3x and 5.7x; that figure is a floor on the core and the row at it is the
pessimistic column.
What the table says:
1. The honest floor per tier, measured: 32 GB 2.33 to 2.37 (the 5090 at its knee); 16 GB 2.06 to 2.10 (the 5080 at