Class v6 10.1: the 5090 SM ladder on four rented hosts (the ceiling held to 43 of 170 SMs, the occupancy 1.3 to 3.9 percent at best), the H100 power decomposition, the fraction a chip cannot strip per tier
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
9629fd8dbc
commit
f21f256a05
1 changed files with 18 additions and 1 deletions
|
|
@ -294,7 +294,7 @@ The table, filled with the defaults, each cell replaced as its lane's row lands
|
|||
|
||||
| Honest tier (measured joules per hash, class v4) | GDDR7 board (0.466) | HBM3 one stack (0.321) | SRAM die at W = 4 (0.051) | With SM-sparse | W = 16 variant |
|
||||
|---|---|---|---|---|---|
|
||||
| RTX 5090 stock (3.36; class v3 2.26) | 3.3x default / 5.8x floor | 3.9x / 7.8x | 5.6x / 21x | no change, MEASURED on three more cards (10.1, 13:25 UK): the honest floor at stock is the base row within 2 percent on the 4090 (best shape saves 4 W), the H100 (1.4 percent) and the 5090; the residual is the clock domain, which only the lock takes off | not applied (default: outcome A, the card -47 percent, dead; the rented-5090 row owed by 14:45) |
|
||||
| RTX 5090 stock (3.36; class v3 2.26) | 3.3x default / 5.8x floor | 3.9x / 7.8x | 5.6x / 21x | 1.3 to 3.9 percent at best, MEASURED (10.1, 14:10 UK: four rented 5090s, the ceiling held to 43 of 170 SMs; the 4090 saves 4 W at 16 SMs, the H100 1.4 percent); the residual is the clock domain, which only the lock takes off; the fraction a chip cannot strip is 16 to 19 percent of the 5090's stock joules, 25 percent at the lock | not applied (default: outcome A, the card -47 percent, dead; the rented-5090 row owed by 14:45) |
|
||||
| RTX 5090 at the 1,300 MHz lock, the record's operating point (2.33 = 1.67 + 0.65) | 2.9x / 4.0x | 3.6x / 5.4x | 3.9x absolute (6.1x in the record's convention, marked) / 14x | no change (the lock pass on PC 1 running; the knee row by 15:00) | not applied (same default) |
|
||||
| Apple M5 Max (1.40; class v3 0.78; the GPU and DRAM channels) | 1.8x / 2.4x | 2.2x / 3.2x | 3.9x / 8.7x | not applicable (no SM lever on Apple) | not applicable (the M5 Max pays 0 at 64 bytes, measured; the chip +33 percent) |
|
||||
| RTX 5080 at its 1,100 MHz lock, the honest NVIDIA floor (2.06 class v4 measured: 71.20 MH/s at 146.6 W; class v5 2.10; floor lane 4, 10.4) | 2.6x / 3.6x | 3.0x / 4.8x | 3.6x / 13x | no change | not applied |
|
||||
|
|
@ -321,6 +321,23 @@ The three class v6 changes it implies: (1) the op mix stays class v4's with the
|
|||
|
||||
The reading for section 2's term: what an idle SM does not save is the die's clock-domain cost (the clock tree, L2 and the crossbar, the memory controllers, leakage at 2,700 to 2,850 MHz); the 4090 at 8 SMs and 65 MH/s still draws 204 W over idle at 2,715 MHz, while the 5090's lock takes 100 W off at the same read rate. **The honest floor at stock per card is the base row within 2 percent; the lever stays rank 1, the operating point.** The patch that ships the auto-tune anyway (`--sm-sparse auto|off|N` and `--sm-hold` on the worker, the tuning-file keys `sm_sparse` and `sm_hold` per card, the pure pieces tested in `emu/variant-test.cpp` on the H100 pod) is on class-v6-floor-sm at 9e478daf; the design's default is off.
|
||||
|
||||
The 5090 ladder (14:10 UK; four rented 5090s on four hosts, base rows 2.08 to 2.49 microjoules at 142 MH/s by board; the ladder as a percent of each host's own base; bit-exact; placement checked by `%smid`). **At stock the 5090 holds class v3's ceiling down to 43 of 170 SMs (99.5 percent) and 28 (98.3), falls at 21 (96.6) and 16 (91); the occupancy saves 1.3 to 3.9 percent of energy per hash by host and never more; class v4's knee is 43 SMs by rate (compute-bound below it) with no saving.** Host c (idle 7 W), class v3, 32 warps per block, one block per SM:
|
||||
|
||||
| SMs | MH/s (percent of base) | W | Microjoules per hash | GDDR7 / HBM3 / SRAM (lane B's 0.14) |
|
||||
|---|---|---|---|---|
|
||||
| 170 | 142.6 (100) | 296.4 | 2.079 | 4.46x / 6.48x / 14.9x |
|
||||
| 85 | 142.4 (99.9) | 296.0 | 2.078 | |
|
||||
| 43 | 141.8 (99.4) | 291.1 | 2.053 | 4.41x / 6.40x / 14.7x |
|
||||
| 28 | 140.3 (98.3) | 288.6 | 2.058 | |
|
||||
| 21 | 137.8 (96.6) | 288.1 | 2.091 | |
|
||||
| 16 | 130.0 (91.1) | 280.4 | 2.157 | |
|
||||
| 11 | 100.9 (70.7) | 255.5 | 2.533 | 5.4x |
|
||||
|
||||
Host e (idle 12 W): base 142.0 at 352.7 W = 2.485; 43 SMs 141.2 at 337.4 = 2.389 (the 3.9 percent case); 21 SMs 137.2 at 334.9 = 2.440. Full SM count with fewer warps (host c): 8 warps per SM 142.4 at 293.8 W = 2.063; 2 warps 139.3 at 288.7; 1 warp 119.6 at 270.3 = 2.260. Matched warps (43 x 32 against 170 x 8): 2.053 against 2.063, equal: the warps in flight set the rate and nothing in the occupancy sets the watts. Class v4 at stock: host c base 142.6 at 450.0 W (the cap) = 3.157; PC 1 unlocked today: base 137.2 at 452.8 W, 43 SMs 134.5 at 470 W, 36 SMs 120.3, 28 SMs 94.4, 21 SMs 70.7, 16 SMs 54.0 (the shadow compute-bound under 43 SMs; the pass drifted 62 to 75 C). The 1,300 lock ladder on PC 1 (`run-ca4-pc1-floorsm-5090-20261008`) is the amendment if not in by 15:30; the default is 20.3b's lock points.
|
||||
|
||||
The decomposition for section 2's term (the H100 microbench, full residency at the boost clock, measured): idle 80 W; resident SMs issuing nothing 136.6 W (+57 W, the SM clock domain); the ALU probe 437.9; the dependent DRAM chase 450.3 W at 32.2 G reads a second (11.5 nJ per read whole-card); the hash 450.4 W at 32.6 G reads a second. The memory path on the GPU side is 313 of the 370 W over idle (9.6 nJ per read above the HBM's modelled 1.2). The 4090 and 5090 probes land by 15:00; PC 1's memory-clock ladder at the 1,300 lock is queued (`run-ca4-pc1-memclk-5090-20261008`). **The fraction a chip cannot strip** (the memory side's own on chip-model 5.3's rows: DRAM plus static plus a controller, about 0.40 microjoules per hash on GDDR7 at 128 reads, 0.47 on the 4090's GDDR6X at its rate, 0.23 on HBM3): the 5090 at stock 16 to 19 percent, at the 1,300 lock 25 percent, the 4090 13 percent, the H100 13 percent. Everything else is the card's silicon around the read and a chip strips it; the honest floor's real number per tier is that fraction, and the only honest-side lever on the rest is the clock (rank 1).
|
||||
|
||||
|
||||
### 10.2 The shadow's k from RTL (floor lane 2, `docs/analysis/class-v6/floor/shadow-k.md` on class-v6-floor-k, build-4 and build-3, first synthesised rows 13:3x UK; RTL and flow under `tools/chip-model/rtl`; full by 19:30 UK)
|
||||
|
||||
**The first synthesised k is 0.18 at the 5090's lock in the absolute convention at N3, and nothing reads inside the claimed 0.3 to 0.8 band except the unscaled ASAP7 figure at the lock (0.35). These are PER-UNIT FLOORS: no fetch, no decode, no register file beyond an 8-entry window, which is the chip rotation kills (section 1). The coordinator's order (13:5x UK): the headline k for the close is the programmable sequencer-core row (fetch, decode, a 32-register file, the per-era registers, the drawn program), asked of the lane with the M5 Max column; until it lands the default is the record's shadowed rows at k 0.5 with these unit floors beside them as the lower bound, and the GDDR7 board at 4.0x at the lock and 5.8x at stock on the floor k is the WORST CASE the served line must survive, not the reading.** Method: a minimal lane (instruction register, an 8 x 32-bit register window of flops, read muxes, the unit, one write port; no fetch or decode) in Yosys 0.68 plus OpenROAD on ASAP7, routed, SPEF, power from a random-input gate-level VCD with every pin annotated, the TC corner at 0.70 V; so every chip figure is a FLOOR and every k a floor. The node scaling is claimed from TSMC's headline per-node power reductions (N5 x0.70, N3 x0.50, N2 x0.36 of ASAP7; approximate). The GPU side is the research file's 15.1a, measured.
|
||||
|
|
|
|||
Loading…
Reference in a new issue