diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index 8a303e5a5..601904b69 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -429,15 +429,15 @@ The two numbers for the close: no rational chip project of any kind below about | HBM3 one stack | 7.5x | 6.7x | 12.7x | 7.2x | 5.9x | modelled chip, measured card | | SRAM die | 66x | 19x | 36x | 14x | 8.8x | modelled chip, measured card | -That table was the rate question; the energy question answered it (the lane's a670ba28, 13:31 UK, on floor lane 5's measured rows of 13:47 UK): W = 16 is dead on measured energy, not on rate. Netted with the card at 1.34x its energy (the 4090 holds its 4-byte rate at 98 percent but draws 296 W against 225; the H100 -21 percent of rate and +33 percent of energy; the 3090 -15 percent of rate): +That table was the rate question; the energy question answered it (the lane's bbf1b1dc, 13:40 UK, on floor lane 5's measured rows): W = 16 is dead on measured energy, not on rate. Netted with the card at 1.25x its energy on the 5090 (the hinted sp170-w4 form: the rate held at +2 percent for 401 W against 313 to 319; the second sector about 4.7 nJ through the fabric against the DRAM's 1.15, above the 2 nJ line that would reverse the DRAM rows) and 1.34x on the 4090 (in brackets; the H100 -21 percent of rate and +33 percent of energy; the 3090 -15 percent of rate): | Chip | W = 16 zero shadow | Class v4 shadow (the synthesised core) | The full shadow | The same chip at W = 8 (the card free) | Label | |---|---|---|---|---|---| -| GDDR7 | 5.8x (up from 5.1x) | 6.3x | 5.5x | 5.1x / 5.1x / 4.7x | modelled chip, measured card | -| HBM3 | 9.0x (up from 7.5x) | 8.8x | 7.0x | 7.5x / 7.2x / 5.9x | modelled chip, measured card | -| SRAM die | 26x (from 66x) | 17.6x | 10.4x | 31x / 17.9x / 9.7x | modelled chip, measured card | +| GDDR7 | 5.5x on the 5090 (5.8x on the 4090), up from 5.1x | 5.9x | 5.3x | 5.1x / 5.1x / 4.7x | modelled chip, measured card | +| HBM3 | 8.4x (9.0x), up from 7.5x | 8.4x | 6.7x | 7.5x / 7.2x / 5.9x | modelled chip, measured card | +| SRAM die | 24x (26x), from 66x | 16.7x | 9.9x | 31x / 17.9x / 9.7x | modelled chip, measured card | -W = 16 costs the honest card 34 percent of its energy so that the DRAM chips' edges rise and the SRAM die lands exactly where W = 8 puts it for free with the shadow on; the only thing it buys is the die's zero-shadow number, which no served line carries. **W = 8 is the width; "never 16" stands on measured rows at stock**; the 5090 stock row and the knee row are amendments. +W = 16 costs the honest card 25 to 34 percent of its energy so that the DRAM chips' edges rise and the SRAM die lands exactly where W = 8 puts it for free with the shadow on; the only thing it buys is the die's zero-shadow number, which no served line carries. **W = 8 is the width; "never 16" stands on measured rows at stock on the 5090, 4090, H100 and 3090**; the 5090's knee row (it would need the second sector under about 2 nJ) is the one amendment that could move it. (3) The shadowed SRAM rows on the k lane's synthesised core (0.11 microjoules of shadow at N3, absolute k about 0.1, the shuffle open; the lane's 2.2b): under class v4 21x at W = 4 and 18x at W = 8 at stock, 14x and 12x at the knee; at the honest cards' whole latency shadow 10x and 9.7x at stock, 6.5x and 6.1x at the knee; an N2 core about 1.2x more. On the synthesised core the shadow is not the whole hold (the record's claimed band read 2x to 6x, marked beside): the memory is a third to a half of the die's energy, the width is worth 1.3x with the shadow on, and the lane's line for the served sentence is 6x to 10x at the full shadow (2x to 4x on the claimed band, marked), never under 2x. These are the per-unit-floor figures of 10.2; the k lane's sequencer-core row re-folds them.