Class v6: W = 16 closed from the knee side (2.3x the energy per hash at the 5090's lock); nothing on it owed
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
86a9f4f27a
commit
9ff09ecc9f
1 changed files with 2 additions and 2 deletions
|
|
@ -294,7 +294,7 @@ The table, filled with the defaults, each cell replaced as its lane's row lands
|
|||
|
||||
| Honest tier (measured joules per hash, class v4) | GDDR7 board (0.466): core / floor | HBM3 one stack (0.321): core / floor | SRAM die at W = 4 (0.051): core / floor | With SM-sparse | W = 16 variant |
|
||||
|---|---|---|---|---|---|
|
||||
| RTX 5090 stock (3.36; class v3 2.26) | 4.1x core / 5.8x floor | 5.0x / 7.8x | 8.2x / 21x | 1.3 to 3.9 percent at best, MEASURED (10.1, 14:10 UK: four rented 5090s, the ceiling held to 43 of 170 SMs; the 4090 saves 4 W at 16 SMs, the H100 1.4 percent); the residual is the clock domain, which only the lock takes off; the fraction a chip cannot strip is 16 to 19 percent of the 5090's stock joules, 25 percent at the lock | not applied: KILLED on measured rows (10.5, 13:47 UK): the hinted form holds the rate on Ada but the card pays the second sector at 8 to 9 nJ, energy per hash +34 percent (4090), +33 (H100), the chip +33 (GDDR7); the edge does not move; the 5090 measured too (13:40 UK: the hinted sparse form holds the rate at +2.3 percent for +88 W, energy +25 percent, the GDDR7 edge 4.7x to 4.4x at zero shadow and unchanged with the shadow); the PC 1 knee row measured in kit d (15:0x UK): the hinted w64-l2 pack 37.59 MH/s at 164.1 W at the lock, 0.229 MH/W against w4's 0.528; dead at the knee too |
|
||||
| RTX 5090 stock (3.36; class v3 2.26) | 4.1x core / 5.8x floor | 5.0x / 7.8x | 8.2x / 21x | 1.3 to 3.9 percent at best, MEASURED (10.1, 14:10 UK: four rented 5090s, the ceiling held to 43 of 170 SMs; the 4090 saves 4 W at 16 SMs, the H100 1.4 percent); the residual is the clock domain, which only the lock takes off; the fraction a chip cannot strip is 16 to 19 percent of the 5090's stock joules, 25 percent at the lock | not applied: KILLED on measured rows (10.5, 13:47 UK): the hinted form holds the rate on Ada but the card pays the second sector at 8 to 9 nJ, energy per hash +34 percent (4090), +33 (H100), the chip +33 (GDDR7); the edge does not move; the 5090 measured too (13:40 UK: the hinted sparse form holds the rate at +2.3 percent for +88 W, energy +25 percent, the GDDR7 edge 4.7x to 4.4x at zero shadow and unchanged with the shadow); the PC 1 knee row measured in kit d (15:0x UK; lane 5's reading at 3951528d): at the 1,300 lock w4 100.55 MH/s at 190.5 W (1.89 microjoules) against the hinted w64-l2 37.59 at 164.1 W (4.37): the hint recovers nothing in the latency-bound regime and the card pays 2.3x per hash against the 20 percent line; so W = 16 is dead at stock and at the knee on Blackwell, Ada, Hopper and Ampere, and no row on it is owed by anyone |
|
||||
| RTX 5090 at the 1,300 MHz lock, the record's operating point (2.33 = 1.67 + 0.65) | 2.8x / 4.0x | 3.4x / 5.4x | 5.7x / 14x | no change, MEASURED (10.1 item 2: class v4 at the lock saves nothing on the full grid, 2.25 microjoules at 302 W on PC 1 today; the class v3 lock row stays 20.3b's, today's pass partial) | not applied (same default) |
|
||||
| Apple M5 Max (1.40; class v3 0.78; the GPU and DRAM channels) | 1.7x / 2.4x | 2.1x / 3.2x | 3.4x / 8.7x | not applicable (no SM lever on Apple) | not applicable (the M5 Max pays 0 at 64 bytes, measured; the chip +33 percent) |
|
||||
| RTX 5080 at its 1,100 MHz lock, the honest NVIDIA floor (2.06 class v4 measured: 71.20 MH/s at 146.6 W; class v5 2.10; floor lane 4, 10.4) | 2.5x / 3.6x | 3.0x / 4.8x | 5.0x / 13x | no change | not applied |
|
||||
|
|
@ -534,7 +534,7 @@ That table was the rate question; the energy question answered it (the lane's bb
|
|||
| HBM3 | 8.4x (9.0x), up from 7.5x | 8.4x | 6.7x | 7.5x / 7.2x / 5.9x | modelled chip, measured card |
|
||||
| SRAM die | 24x (26x), from 66x | 16.7x | 9.9x | 31x / 17.9x / 9.7x | modelled chip, measured card |
|
||||
|
||||
W = 16 costs the honest card 25 to 34 percent of its energy so that the DRAM chips' edges rise and the SRAM die lands exactly where W = 8 puts it for free with the shadow on; the only thing it buys is the die's zero-shadow number, which no served line carries. **W = 8 is the width; "never 16" stands on measured rows at stock on the 5090, 4090, H100 and 3090**; the 5090's knee row (it would need the second sector under about 2 nJ) is the one amendment that could move it.
|
||||
W = 16 costs the honest card 25 to 34 percent of its energy so that the DRAM chips' edges rise and the SRAM die lands exactly where W = 8 puts it for free with the shadow on; the only thing it buys is the die's zero-shadow number, which no served line carries. **W = 8 is the width; "never 16" stands on measured rows at stock on the 5090, 4090, H100 and 3090**; the 5090's knee row landed (kit d, 15:0x UK: 2.3x the energy per hash at the lock) and closes it; nothing on W = 16 is owed.
|
||||
|
||||
(3) The shadowed SRAM rows on the k lane's synthesised core (0.11 microjoules of shadow at N3, absolute k about 0.1, the shuffle open; the lane's 2.2b): under class v4 21x at W = 4 and 18x at W = 8 at stock, 14x and 12x at the knee; at the honest cards' whole latency shadow 10x and 9.7x at stock, 6.5x and 6.1x at the knee; an N2 core about 1.2x more. On the synthesised core the shadow is not the whole hold (the record's claimed band read 2x to 6x, marked beside): the memory is a third to a half of the die's energy, the width is worth 1.3x with the shadow on, and the lane's line for the served sentence is 6x to 10x at the full shadow (2x to 4x on the claimed band, marked), never under 2x. These are the per-unit-floor figures of 10.2, now the marked worst case; the core-row fold follows.
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue