From 308aa4a7dabebd8e13c3527f4c202582fae8161c Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 14:48:53 +0000 Subject: [PATCH] Class v6 floor lane 1: sm-sparse.md amendment, the memory-clock ladder on PC 1 at the 1,300 lock (the 5090's PHY has two states, 13,801 and 7,001 MHz; the half-rate state loses 43 percent of rate for 37 percent of watts, energy per hash up 12 percent on class v3 and 5 on class v4: the memory clock is not a lever; the memory path read as 11 nJ per read at the margin, equal to the DRAM chase probes) Co-Authored-By: Claude Fable 5.1 Documents-only replay of 6d666ce9d (6d666ce9d60240d19adf390216ef49e387d7e325) for the box mirror master --- docs/analysis/class-v6/floor/sm-sparse.md | 64 ++++++++++++++++++++--- 1 file changed, 57 insertions(+), 7 deletions(-) diff --git a/docs/analysis/class-v6/floor/sm-sparse.md b/docs/analysis/class-v6/floor/sm-sparse.md index 41099203b..2848698e6 100644 --- a/docs/analysis/class-v6/floor/sm-sparse.md +++ b/docs/analysis/class-v6/floor/sm-sparse.md @@ -40,8 +40,9 @@ M5 Max's 0.78 is measured (6 October). "Edge" is the card's microjoules over the On the H100 the persistent shape itself loses 4 percent of the class v4 rate and the self-tune installs nothing. PC 1's class v3 at stock agrees with the rented ladder at every rung (43 SMs 136.5 of 136.7 MH/s at 10 W under base; 170 SMs x 8 warps 137.1 at 310.8 W, 3.5 percent under base, the best shape on that board). -4. **The residual is the core clock domain and the memory path on the GPU side, and the memory clock is the one - knob left after the core lock.** The decomposition probes (section 3) put the 5090's draw over idle at the hash +4. **The residual is the core clock domain and the memory path on the GPU side, and the memory clock is not a + lever either (section 3.1: the 5090's PHY has two states, 13,801 and 7,001 MHz; the half-rate state loses 43 + percent of the rate for 37 percent of the watts, so the energy per hash rises 12 percent).** The decomposition probes (section 3) put the 5090's draw over idle at the hash into: the resident SMs issuing nothing (the SM clock domain awake), the L2 and crossbar path per read, and the DRAM path per read on top of the memory's own modelled 2.0 nJ; the memory-clock ladder on PC 1 at the 1,300 lock reads the PHY's share directly. The honest floor per tier is the base row at stock within 2 percent, the core lock @@ -536,7 +537,30 @@ The floorsm ladder ### PC 1, the memory-clock ladder at the 1,300 MHz lock, job `run-ca4-pc1-memclk-5090-20261008` -The memory-clock ladder: not landed. +The memory-clock ladder + +| Pack | State | Shape | MH/s | W | Microjoules | SM MHz | Mem MHz | Temp max | Check | +|---|---|---|---|---|---|---|---|---|---| +| v4-devnet-epoch0 | unlocked | base (grid) | 136.552 | 464.3 | 3.400 | 2852.1 | 13801 | 70 | PASS | +| mx8-devnet-epoch0 | unlocked | base (grid) | 136.29 | 328.1 | 2.407 | 2850 | 13801 | 65 | PASS | +| v4-devnet-epoch0 | 1300 | base (grid) | 134.34 | 309 | 2.300 | 1290 | 13801 | 60 | PASS | +| mx8-devnet-epoch0 | 1300 | base (grid) | 134.089 | 222.9 | 1.662 | 1290 | 13801 | 57 | PASS | +| v4-devnet-epoch0 | 1300m14001 | base (grid) | 134.513 | 308.1 | 2.290 | 1290 | 13801 | 59 | PASS | +| mx8-devnet-epoch0 | 1300m14001 | base (grid) | 134.342 | 222.7 | 1.658 | 1290 | 13801 | 56 | PASS | +| v4-devnet-epoch0 | 1300m12001 | base (grid) | 134.528 | 308.6 | 2.294 | 1290 | 13801 | 59 | PASS | +| mx8-devnet-epoch0 | 1300m12001 | base (grid) | 134.252 | 223.1 | 1.662 | 1290 | 13801 | 56 | PASS | +| v4-devnet-epoch0 | 1300m10001 | base (grid) | 134.487 | 308.9 | 2.297 | 1290 | 13801 | 59 | PASS | +| mx8-devnet-epoch0 | 1300m10001 | base (grid) | 134.287 | 223.7 | 1.666 | 1290 | 13801 | 55 | PASS | +| v4-devnet-epoch0 | 1300m8001 | base (grid) | 134.541 | 308.7 | 2.294 | 1290 | 13801 | 60 | PASS | +| mx8-devnet-epoch0 | 1300m8001 | base (grid) | 134.33 | 223 | 1.660 | 1290 | 13801 | 56 | PASS | +| v4-devnet-epoch0 | 1300m6001 | base (grid) | 76.09 | 187 | 2.458 | 1290 | 7001 | 53 | PASS | +| mx8-devnet-epoch0 | 1300m6001 | base (grid) | 76.009 | 142.2 | 1.871 | 1290 | 7001 | 51 | PASS | +| v4-devnet-epoch0 | 1300m5001 | base (grid) | 76.091 | 184.4 | 2.423 | 1290 | 7001 | 50 | PASS | +| mx8-devnet-epoch0 | 1300m5001 | base (grid) | 76.007 | 141.1 | 1.856 | 1290 | 7001 | 48 | PASS | +| v4-devnet-epoch0 | 1300m3001 | base (grid) | 76.082 | 183.8 | 2.416 | 1290 | 7001 | 50 | PASS | +| mx8-devnet-epoch0 | 1300m3001 | base (grid) | 75.96 | 140.9 | 1.855 | 1290 | 7001 | 47 | PASS | +| v4-devnet-epoch0 | unlocked-end | base (grid) | 136.562 | 452.5 | 3.314 | 2866.6 | 13801 | 61 | PASS | +| mx8-devnet-epoch0 | unlocked-end | base (grid) | 136.298 | 318.6 | 2.338 | 2865 | 13801 | 60 | PASS | ### The decomposition probes (the worker's --microbench: sleep = full residency and no issue; the L2 chase; the dependent DRAM chase; the ALU probe; 30 s each with the sampler's watts and both clocks) @@ -626,6 +650,35 @@ clock for the PHY. The core lock takes the first three down with the clock (the reads per second on PC 1, 20.3); the memory-clock lock is the only knob on the fourth, and no rented host allows it, so its row is PC 1's (the memory-clock ladder table in section 1, the job `run-ca4-pc1-memclk-5090-20261008`). +### 3.1 The memory-clock ladder (PC 1, the 5090 alone, job `run-ca4-pc1-memclk-5090-20261008-b`, 15:0x to 15:4x UK, the helper's `lmc` verb at the 1,300 MHz core lock; the amendment of 15:50 UK) + +The driver holds the 5090's memory clock at two points only: every ask at or above 8,001 MHz reads 13,801 on the row +and every ask at or below 6,001 reads 7,001 (the GDDR7 PHY's half-rate state). The rows, both bit-exact against the +Mac's fingerprints: + +| Pack | Core lock, memory ask | Memory MHz on the row | MH/s | Watts | Microjoules per hash | nJ per dependent read, whole card | Label | +|---|---|---|---|---|---|---|---| +| class v3 (mx8) | 1,300, no memory lock | 13,801 | 134.09 | 222.9 | 1.662 | 13.0 | measured | +| class v3 | 1,300, lmc 14001 / 12001 / 10001 / 8001 | 13,801 | 134.3 to 134.5 | 222.7 to 223.7 | 1.658 to 1.666 | 13.0 | measured (no change) | +| class v3 | 1,300, lmc 6001 / 5001 / 3001 | 7,001 | 76.0 | 140.9 to 142.2 | 1.855 to 1.871 | 14.5 | measured | +| class v4 (sh256x27) | 1,300, no memory lock | 13,801 | 134.34 | 309.0 | 2.300 | | measured | +| class v4 | 1,300, lmc 8001 and above | 13,801 | 134.5 | 308.1 to 308.9 | 2.290 to 2.297 | | measured | +| class v4 | 1,300, lmc 6001 and below | 7,001 | 76.1 | 183.8 to 187.0 | 2.416 to 2.458 | | measured | +| both, unlocked (the pass's own base) | none | 13,801 | 136.3 / 136.6 | 318.6 to 328.1 / 452.5 to 464.3 | 2.34 to 2.41 / 3.31 to 3.40 | | measured (the heat drift across the pass) | + +The reading: the memory clock is not a lever either. At the half-rate PHY state the rate falls 43 percent (134 to 76 +MH/s: the hash is bound by the memory's activate ceiling, which scales with the memory clock) and the watts fall 37 +percent (82 W on class v3, 125 W on class v4), so the energy per hash RISES 12 percent on class v3 (1.66 to 1.86) and +5 percent on class v4 (2.30 to 2.42). Read as a decomposition, the 82 W between the two memory states on class v3 is +the memory path's clock-scaled share at 7.4 G reads per second of rate lost: 11 nJ per read at the margin, the same +figure the DRAM chase probes give for the whole memory path on the GPU side (10.9 to 11.8 nJ). So the memory-side term +of the 5090's draw is about 80 to 90 W of its 223 W at the lock (the PHY, controllers and L2 at the memory clock plus +the devices' own 55 W modelled), and it is bought one for one with the rate: the honest floor at the 1,300 lock stays +1.66 microjoules on this board (1.58 on 20.3b's cooler pass), and no clock or occupancy knob on the card takes it lower +without taking the rate with it. The class v4 premium at the half-rate state is 43 W for 76 MH/s x 102,100 ops = 5.5 +pJ per counted op (6.3 at the full memory clock on the same pass): the shadow's marginal is the core domain's, not the +memory's, as expected. + ## 4. The honest floor per tier, and the fraction a chip cannot strip The chip of the record pays the memory's own energy per read, its static power and a controller and PHY of its own @@ -702,10 +755,7 @@ AMD or Apple owner gets nothing from this lane (the shapes are CUDA's; Metal and reading the sparse shapes lower than 20.3b did on the same board at the same lock (85 SMs 111.0 against 129.8 MH/s); the pass ran across the founder's card swap on PC 1 (14:2x UK), so the class v3 lock row stays 20.3b's (43 SMs 126.3 MH/s at 205.4 W against the full grid's 134.0 at 211.4: 1.63 against 1.58 microjoules, nothing saved) and the - ladder is re-run on a quiet PC 1 before it replaces it. The memory-clock ladder is queued behind the hash lane's two - jobs (run-ca4-pc1-memclk-5090-20261008-b) and lands in section 1 when its upload does. -- The decomposition probes on the 4090 and the 5090 host c: in section 1 when their runs end (queued behind the - sweeps on the same cards so no row shares the card). + ladder is re-run on a quiet PC 1 before it replaces it. The memory-clock ladder landed at 15:4x UK (section 3.1). - The memory side's own share per tier is modelled (chip-model 5.3); the memory-clock ladder is the one measurement of its PHY term, and only on PC 1. - Class v5 on the rented 5090s was not run (the v5-kits worker carries the leaf upload and not the sparse shapes; the