Class v6 floor lane 1: sm-sparse.md amendment, the memory-clock ladder on PC 1 at the 1,300 lock (the 5090's PHY has two states, 13,801 and 7,001 MHz; the half-rate state loses 43 percent of rate for 37 percent of watts, energy per hash up 12 percent on class v3 and 5 on class v4: the memory clock is not a lever; the memory path read as 11 nJ per read at the margin, equal to the DRAM chase probes)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 14:48:53 +00:00
parent e7f622f11c
commit 6d666ce9d6

View file

@ -40,8 +40,9 @@ M5 Max's 0.78 is measured (6 October). "Edge" is the card's microjoules over the
On the H100 the persistent shape itself loses 4 percent of the class v4 rate and the self-tune installs nothing.
PC 1's class v3 at stock agrees with the rented ladder at every rung (43 SMs 136.5 of 136.7 MH/s at 10 W under
base; 170 SMs x 8 warps 137.1 at 310.8 W, 3.5 percent under base, the best shape on that board).
4. **The residual is the core clock domain and the memory path on the GPU side, and the memory clock is the one
knob left after the core lock.** The decomposition probes (section 3) put the 5090's draw over idle at the hash
4. **The residual is the core clock domain and the memory path on the GPU side, and the memory clock is not a
lever either (section 3.1: the 5090's PHY has two states, 13,801 and 7,001 MHz; the half-rate state loses 43
percent of the rate for 37 percent of the watts, so the energy per hash rises 12 percent).** The decomposition probes (section 3) put the 5090's draw over idle at the hash
into: the resident SMs issuing nothing (the SM clock domain awake), the L2 and crossbar path per read, and the DRAM
path per read on top of the memory's own modelled 2.0 nJ; the memory-clock ladder on PC 1 at the 1,300 lock reads
the PHY's share directly. The honest floor per tier is the base row at stock within 2 percent, the core lock
@ -536,7 +537,30 @@ The floorsm ladder
### PC 1, the memory-clock ladder at the 1,300 MHz lock, job `run-ca4-pc1-memclk-5090-20261008`
The memory-clock ladder: not landed.
The memory-clock ladder
| Pack | State | Shape | MH/s | W | Microjoules | SM MHz | Mem MHz | Temp max | Check |
|---|---|---|---|---|---|---|---|---|---|
| v4-devnet-epoch0 | unlocked | base (grid) | 136.552 | 464.3 | 3.400 | 2852.1 | 13801 | 70 | PASS |
| mx8-devnet-epoch0 | unlocked | base (grid) | 136.29 | 328.1 | 2.407 | 2850 | 13801 | 65 | PASS |
| v4-devnet-epoch0 | 1300 | base (grid) | 134.34 | 309 | 2.300 | 1290 | 13801 | 60 | PASS |
| mx8-devnet-epoch0 | 1300 | base (grid) | 134.089 | 222.9 | 1.662 | 1290 | 13801 | 57 | PASS |
| v4-devnet-epoch0 | 1300m14001 | base (grid) | 134.513 | 308.1 | 2.290 | 1290 | 13801 | 59 | PASS |
| mx8-devnet-epoch0 | 1300m14001 | base (grid) | 134.342 | 222.7 | 1.658 | 1290 | 13801 | 56 | PASS |
| v4-devnet-epoch0 | 1300m12001 | base (grid) | 134.528 | 308.6 | 2.294 | 1290 | 13801 | 59 | PASS |
| mx8-devnet-epoch0 | 1300m12001 | base (grid) | 134.252 | 223.1 | 1.662 | 1290 | 13801 | 56 | PASS |
| v4-devnet-epoch0 | 1300m10001 | base (grid) | 134.487 | 308.9 | 2.297 | 1290 | 13801 | 59 | PASS |
| mx8-devnet-epoch0 | 1300m10001 | base (grid) | 134.287 | 223.7 | 1.666 | 1290 | 13801 | 55 | PASS |
| v4-devnet-epoch0 | 1300m8001 | base (grid) | 134.541 | 308.7 | 2.294 | 1290 | 13801 | 60 | PASS |
| mx8-devnet-epoch0 | 1300m8001 | base (grid) | 134.33 | 223 | 1.660 | 1290 | 13801 | 56 | PASS |
| v4-devnet-epoch0 | 1300m6001 | base (grid) | 76.09 | 187 | 2.458 | 1290 | 7001 | 53 | PASS |
| mx8-devnet-epoch0 | 1300m6001 | base (grid) | 76.009 | 142.2 | 1.871 | 1290 | 7001 | 51 | PASS |
| v4-devnet-epoch0 | 1300m5001 | base (grid) | 76.091 | 184.4 | 2.423 | 1290 | 7001 | 50 | PASS |
| mx8-devnet-epoch0 | 1300m5001 | base (grid) | 76.007 | 141.1 | 1.856 | 1290 | 7001 | 48 | PASS |
| v4-devnet-epoch0 | 1300m3001 | base (grid) | 76.082 | 183.8 | 2.416 | 1290 | 7001 | 50 | PASS |
| mx8-devnet-epoch0 | 1300m3001 | base (grid) | 75.96 | 140.9 | 1.855 | 1290 | 7001 | 47 | PASS |
| v4-devnet-epoch0 | unlocked-end | base (grid) | 136.562 | 452.5 | 3.314 | 2866.6 | 13801 | 61 | PASS |
| mx8-devnet-epoch0 | unlocked-end | base (grid) | 136.298 | 318.6 | 2.338 | 2865 | 13801 | 60 | PASS |
### The decomposition probes (the worker's --microbench: sleep = full residency and no issue; the L2 chase; the dependent DRAM chase; the ALU probe; 30 s each with the sampler's watts and both clocks)
@ -626,6 +650,35 @@ clock for the PHY. The core lock takes the first three down with the clock (the
reads per second on PC 1, 20.3); the memory-clock lock is the only knob on the fourth, and no rented host allows it, so
its row is PC 1's (the memory-clock ladder table in section 1, the job `run-ca4-pc1-memclk-5090-20261008`).
### 3.1 The memory-clock ladder (PC 1, the 5090 alone, job `run-ca4-pc1-memclk-5090-20261008-b`, 15:0x to 15:4x UK, the helper's `lmc` verb at the 1,300 MHz core lock; the amendment of 15:50 UK)
The driver holds the 5090's memory clock at two points only: every ask at or above 8,001 MHz reads 13,801 on the row
and every ask at or below 6,001 reads 7,001 (the GDDR7 PHY's half-rate state). The rows, both bit-exact against the
Mac's fingerprints:
| Pack | Core lock, memory ask | Memory MHz on the row | MH/s | Watts | Microjoules per hash | nJ per dependent read, whole card | Label |
|---|---|---|---|---|---|---|---|
| class v3 (mx8) | 1,300, no memory lock | 13,801 | 134.09 | 222.9 | 1.662 | 13.0 | measured |
| class v3 | 1,300, lmc 14001 / 12001 / 10001 / 8001 | 13,801 | 134.3 to 134.5 | 222.7 to 223.7 | 1.658 to 1.666 | 13.0 | measured (no change) |
| class v3 | 1,300, lmc 6001 / 5001 / 3001 | 7,001 | 76.0 | 140.9 to 142.2 | 1.855 to 1.871 | 14.5 | measured |
| class v4 (sh256x27) | 1,300, no memory lock | 13,801 | 134.34 | 309.0 | 2.300 | | measured |
| class v4 | 1,300, lmc 8001 and above | 13,801 | 134.5 | 308.1 to 308.9 | 2.290 to 2.297 | | measured |
| class v4 | 1,300, lmc 6001 and below | 7,001 | 76.1 | 183.8 to 187.0 | 2.416 to 2.458 | | measured |
| both, unlocked (the pass's own base) | none | 13,801 | 136.3 / 136.6 | 318.6 to 328.1 / 452.5 to 464.3 | 2.34 to 2.41 / 3.31 to 3.40 | | measured (the heat drift across the pass) |
The reading: the memory clock is not a lever either. At the half-rate PHY state the rate falls 43 percent (134 to 76
MH/s: the hash is bound by the memory's activate ceiling, which scales with the memory clock) and the watts fall 37
percent (82 W on class v3, 125 W on class v4), so the energy per hash RISES 12 percent on class v3 (1.66 to 1.86) and
5 percent on class v4 (2.30 to 2.42). Read as a decomposition, the 82 W between the two memory states on class v3 is
the memory path's clock-scaled share at 7.4 G reads per second of rate lost: 11 nJ per read at the margin, the same
figure the DRAM chase probes give for the whole memory path on the GPU side (10.9 to 11.8 nJ). So the memory-side term
of the 5090's draw is about 80 to 90 W of its 223 W at the lock (the PHY, controllers and L2 at the memory clock plus
the devices' own 55 W modelled), and it is bought one for one with the rate: the honest floor at the 1,300 lock stays
1.66 microjoules on this board (1.58 on 20.3b's cooler pass), and no clock or occupancy knob on the card takes it lower
without taking the rate with it. The class v4 premium at the half-rate state is 43 W for 76 MH/s x 102,100 ops = 5.5
pJ per counted op (6.3 at the full memory clock on the same pass): the shadow's marginal is the core domain's, not the
memory's, as expected.
## 4. The honest floor per tier, and the fraction a chip cannot strip
The chip of the record pays the memory's own energy per read, its static power and a controller and PHY of its own
@ -702,10 +755,7 @@ AMD or Apple owner gets nothing from this lane (the shapes are CUDA's; Metal and
reading the sparse shapes lower than 20.3b did on the same board at the same lock (85 SMs 111.0 against 129.8 MH/s);
the pass ran across the founder's card swap on PC 1 (14:2x UK), so the class v3 lock row stays 20.3b's (43 SMs
126.3 MH/s at 205.4 W against the full grid's 134.0 at 211.4: 1.63 against 1.58 microjoules, nothing saved) and the
ladder is re-run on a quiet PC 1 before it replaces it. The memory-clock ladder is queued behind the hash lane's two
jobs (run-ca4-pc1-memclk-5090-20261008-b) and lands in section 1 when its upload does.
- The decomposition probes on the 4090 and the 5090 host c: in section 1 when their runs end (queued behind the
sweeps on the same cards so no row shares the card).
ladder is re-run on a quiet PC 1 before it replaces it. The memory-clock ladder landed at 15:4x UK (section 3.1).
- The memory side's own share per tier is modelled (chip-model 5.3); the memory-clock ladder is the one measurement of
its PHY term, and only on PC 1.
- Class v5 on the rented 5090s was not run (the v5-kits worker carries the leaf upload and not the sparse shapes; the