Class v6 floor lane 1: sm-sparse.md amendment, the memory-clock ladder on PC 1 at the 1,300 lock (the 5090's PHY has two states, 13,801 and 7,001 MHz; the half-rate state loses 43 percent of rate for 37 percent of watts, energy per hash up 12 percent on class v3 and 5 on class v4: the memory clock is not a lever; the memory path read as 11 nJ per read at the margin, equal to the DRAM chase probes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
e7f622f11c
commit
6d666ce9d6
1 changed files with 57 additions and 7 deletions
|
|
@ -40,8 +40,9 @@ M5 Max's 0.78 is measured (6 October). "Edge" is the card's microjoules over the
|
|||
On the H100 the persistent shape itself loses 4 percent of the class v4 rate and the self-tune installs nothing.
|
||||
PC 1's class v3 at stock agrees with the rented ladder at every rung (43 SMs 136.5 of 136.7 MH/s at 10 W under
|
||||
base; 170 SMs x 8 warps 137.1 at 310.8 W, 3.5 percent under base, the best shape on that board).
|
||||
4. **The residual is the core clock domain and the memory path on the GPU side, and the memory clock is the one
|
||||
knob left after the core lock.** The decomposition probes (section 3) put the 5090's draw over idle at the hash
|
||||
4. **The residual is the core clock domain and the memory path on the GPU side, and the memory clock is not a
|
||||
lever either (section 3.1: the 5090's PHY has two states, 13,801 and 7,001 MHz; the half-rate state loses 43
|
||||
percent of the rate for 37 percent of the watts, so the energy per hash rises 12 percent).** The decomposition probes (section 3) put the 5090's draw over idle at the hash
|
||||
into: the resident SMs issuing nothing (the SM clock domain awake), the L2 and crossbar path per read, and the DRAM
|
||||
path per read on top of the memory's own modelled 2.0 nJ; the memory-clock ladder on PC 1 at the 1,300 lock reads
|
||||
the PHY's share directly. The honest floor per tier is the base row at stock within 2 percent, the core lock
|
||||
|
|
@ -536,7 +537,30 @@ The floorsm ladder
|
|||
### PC 1, the memory-clock ladder at the 1,300 MHz lock, job `run-ca4-pc1-memclk-5090-20261008`
|
||||
|
||||
|
||||
The memory-clock ladder: not landed.
|
||||
The memory-clock ladder
|
||||
|
||||
| Pack | State | Shape | MH/s | W | Microjoules | SM MHz | Mem MHz | Temp max | Check |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| v4-devnet-epoch0 | unlocked | base (grid) | 136.552 | 464.3 | 3.400 | 2852.1 | 13801 | 70 | PASS |
|
||||
| mx8-devnet-epoch0 | unlocked | base (grid) | 136.29 | 328.1 | 2.407 | 2850 | 13801 | 65 | PASS |
|
||||
| v4-devnet-epoch0 | 1300 | base (grid) | 134.34 | 309 | 2.300 | 1290 | 13801 | 60 | PASS |
|
||||
| mx8-devnet-epoch0 | 1300 | base (grid) | 134.089 | 222.9 | 1.662 | 1290 | 13801 | 57 | PASS |
|
||||
| v4-devnet-epoch0 | 1300m14001 | base (grid) | 134.513 | 308.1 | 2.290 | 1290 | 13801 | 59 | PASS |
|
||||
| mx8-devnet-epoch0 | 1300m14001 | base (grid) | 134.342 | 222.7 | 1.658 | 1290 | 13801 | 56 | PASS |
|
||||
| v4-devnet-epoch0 | 1300m12001 | base (grid) | 134.528 | 308.6 | 2.294 | 1290 | 13801 | 59 | PASS |
|
||||
| mx8-devnet-epoch0 | 1300m12001 | base (grid) | 134.252 | 223.1 | 1.662 | 1290 | 13801 | 56 | PASS |
|
||||
| v4-devnet-epoch0 | 1300m10001 | base (grid) | 134.487 | 308.9 | 2.297 | 1290 | 13801 | 59 | PASS |
|
||||
| mx8-devnet-epoch0 | 1300m10001 | base (grid) | 134.287 | 223.7 | 1.666 | 1290 | 13801 | 55 | PASS |
|
||||
| v4-devnet-epoch0 | 1300m8001 | base (grid) | 134.541 | 308.7 | 2.294 | 1290 | 13801 | 60 | PASS |
|
||||
| mx8-devnet-epoch0 | 1300m8001 | base (grid) | 134.33 | 223 | 1.660 | 1290 | 13801 | 56 | PASS |
|
||||
| v4-devnet-epoch0 | 1300m6001 | base (grid) | 76.09 | 187 | 2.458 | 1290 | 7001 | 53 | PASS |
|
||||
| mx8-devnet-epoch0 | 1300m6001 | base (grid) | 76.009 | 142.2 | 1.871 | 1290 | 7001 | 51 | PASS |
|
||||
| v4-devnet-epoch0 | 1300m5001 | base (grid) | 76.091 | 184.4 | 2.423 | 1290 | 7001 | 50 | PASS |
|
||||
| mx8-devnet-epoch0 | 1300m5001 | base (grid) | 76.007 | 141.1 | 1.856 | 1290 | 7001 | 48 | PASS |
|
||||
| v4-devnet-epoch0 | 1300m3001 | base (grid) | 76.082 | 183.8 | 2.416 | 1290 | 7001 | 50 | PASS |
|
||||
| mx8-devnet-epoch0 | 1300m3001 | base (grid) | 75.96 | 140.9 | 1.855 | 1290 | 7001 | 47 | PASS |
|
||||
| v4-devnet-epoch0 | unlocked-end | base (grid) | 136.562 | 452.5 | 3.314 | 2866.6 | 13801 | 61 | PASS |
|
||||
| mx8-devnet-epoch0 | unlocked-end | base (grid) | 136.298 | 318.6 | 2.338 | 2865 | 13801 | 60 | PASS |
|
||||
|
||||
|
||||
### The decomposition probes (the worker's --microbench: sleep = full residency and no issue; the L2 chase; the dependent DRAM chase; the ALU probe; 30 s each with the sampler's watts and both clocks)
|
||||
|
|
@ -626,6 +650,35 @@ clock for the PHY. The core lock takes the first three down with the clock (the
|
|||
reads per second on PC 1, 20.3); the memory-clock lock is the only knob on the fourth, and no rented host allows it, so
|
||||
its row is PC 1's (the memory-clock ladder table in section 1, the job `run-ca4-pc1-memclk-5090-20261008`).
|
||||
|
||||
### 3.1 The memory-clock ladder (PC 1, the 5090 alone, job `run-ca4-pc1-memclk-5090-20261008-b`, 15:0x to 15:4x UK, the helper's `lmc` verb at the 1,300 MHz core lock; the amendment of 15:50 UK)
|
||||
|
||||
The driver holds the 5090's memory clock at two points only: every ask at or above 8,001 MHz reads 13,801 on the row
|
||||
and every ask at or below 6,001 reads 7,001 (the GDDR7 PHY's half-rate state). The rows, both bit-exact against the
|
||||
Mac's fingerprints:
|
||||
|
||||
| Pack | Core lock, memory ask | Memory MHz on the row | MH/s | Watts | Microjoules per hash | nJ per dependent read, whole card | Label |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| class v3 (mx8) | 1,300, no memory lock | 13,801 | 134.09 | 222.9 | 1.662 | 13.0 | measured |
|
||||
| class v3 | 1,300, lmc 14001 / 12001 / 10001 / 8001 | 13,801 | 134.3 to 134.5 | 222.7 to 223.7 | 1.658 to 1.666 | 13.0 | measured (no change) |
|
||||
| class v3 | 1,300, lmc 6001 / 5001 / 3001 | 7,001 | 76.0 | 140.9 to 142.2 | 1.855 to 1.871 | 14.5 | measured |
|
||||
| class v4 (sh256x27) | 1,300, no memory lock | 13,801 | 134.34 | 309.0 | 2.300 | | measured |
|
||||
| class v4 | 1,300, lmc 8001 and above | 13,801 | 134.5 | 308.1 to 308.9 | 2.290 to 2.297 | | measured |
|
||||
| class v4 | 1,300, lmc 6001 and below | 7,001 | 76.1 | 183.8 to 187.0 | 2.416 to 2.458 | | measured |
|
||||
| both, unlocked (the pass's own base) | none | 13,801 | 136.3 / 136.6 | 318.6 to 328.1 / 452.5 to 464.3 | 2.34 to 2.41 / 3.31 to 3.40 | | measured (the heat drift across the pass) |
|
||||
|
||||
The reading: the memory clock is not a lever either. At the half-rate PHY state the rate falls 43 percent (134 to 76
|
||||
MH/s: the hash is bound by the memory's activate ceiling, which scales with the memory clock) and the watts fall 37
|
||||
percent (82 W on class v3, 125 W on class v4), so the energy per hash RISES 12 percent on class v3 (1.66 to 1.86) and
|
||||
5 percent on class v4 (2.30 to 2.42). Read as a decomposition, the 82 W between the two memory states on class v3 is
|
||||
the memory path's clock-scaled share at 7.4 G reads per second of rate lost: 11 nJ per read at the margin, the same
|
||||
figure the DRAM chase probes give for the whole memory path on the GPU side (10.9 to 11.8 nJ). So the memory-side term
|
||||
of the 5090's draw is about 80 to 90 W of its 223 W at the lock (the PHY, controllers and L2 at the memory clock plus
|
||||
the devices' own 55 W modelled), and it is bought one for one with the rate: the honest floor at the 1,300 lock stays
|
||||
1.66 microjoules on this board (1.58 on 20.3b's cooler pass), and no clock or occupancy knob on the card takes it lower
|
||||
without taking the rate with it. The class v4 premium at the half-rate state is 43 W for 76 MH/s x 102,100 ops = 5.5
|
||||
pJ per counted op (6.3 at the full memory clock on the same pass): the shadow's marginal is the core domain's, not the
|
||||
memory's, as expected.
|
||||
|
||||
## 4. The honest floor per tier, and the fraction a chip cannot strip
|
||||
|
||||
The chip of the record pays the memory's own energy per read, its static power and a controller and PHY of its own
|
||||
|
|
@ -702,10 +755,7 @@ AMD or Apple owner gets nothing from this lane (the shapes are CUDA's; Metal and
|
|||
reading the sparse shapes lower than 20.3b did on the same board at the same lock (85 SMs 111.0 against 129.8 MH/s);
|
||||
the pass ran across the founder's card swap on PC 1 (14:2x UK), so the class v3 lock row stays 20.3b's (43 SMs
|
||||
126.3 MH/s at 205.4 W against the full grid's 134.0 at 211.4: 1.63 against 1.58 microjoules, nothing saved) and the
|
||||
ladder is re-run on a quiet PC 1 before it replaces it. The memory-clock ladder is queued behind the hash lane's two
|
||||
jobs (run-ca4-pc1-memclk-5090-20261008-b) and lands in section 1 when its upload does.
|
||||
- The decomposition probes on the 4090 and the 5090 host c: in section 1 when their runs end (queued behind the
|
||||
sweeps on the same cards so no row shares the card).
|
||||
ladder is re-run on a quiet PC 1 before it replaces it. The memory-clock ladder landed at 15:4x UK (section 3.1).
|
||||
- The memory side's own share per tier is modelled (chip-model 5.3); the memory-clock ladder is the one measurement of
|
||||
its PHY term, and only on PC 1.
|
||||
- Class v5 on the rented 5090s was not run (the v5-kits worker carries the leaf upload and not the sparse shapes; the
|
||||
|
|
|
|||
Loading…
Reference in a new issue