Class v6 floor lane 1: sm-sparse.md, the decomposition rows (the awake floor 120 to 139 W on every card, the memory path 10 to 13 nJ per dependent read against the devices' modelled 2.0 or 1.2), the 5090's clean self-tune rows (sp37 of 170 at 2.5 percent under base on class v3, nothing on class v4), every pod destroyed
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
b1e940a260
commit
e89fef216c
1 changed files with 75 additions and 7 deletions
|
|
@ -109,8 +109,26 @@ card NVIDIAGeForceRTX5090 driver 610.57.04 power limit 450.00 W idle 7.1 W at 21
|
|||
| sp85-w32 | 85 x 32 | 141.46 (99.2%) | 450.0 | 442.9 | 3.181 | 6.83x / 9.91x / 22.72x / 4.08x | 2833 | 64 | PASS |
|
||||
| sp64-w32 | 64 x 32 | 136.58 (95.8%) | 450.4 | 443.3 | 3.298 | 7.08x / 10.27x / 23.56x / 4.23x | 2850 | 63 | PASS |
|
||||
| sp52-w32 | 52 x 32 | 121.00 (84.9%) | 417.9 | 410.8 | 3.454 | 7.41x / 10.76x / 24.67x / 4.43x | 2873 | 62 | PASS |
|
||||
| sp43-w32 | 43 x 32 | 124.16 (87.1%) | 422.7 | 415.6 | 3.404 | 7.30x / 10.60x / 24.31x / 4.36x | 2872 | 62 | PASS |
|
||||
| sp36-w32 | 36 x 32 | 104.35 (73.2%) | 380.5 | 373.4 | 3.646 | 7.82x / 11.36x / 26.04x / 4.67x | 2877 | 61 | PASS |
|
||||
| sp28-w32 | 28 x 32 | 81.35 (57.1%) | 334.2 | 327.1 | 4.108 | 8.82x / 12.80x / 29.34x / 5.27x | 2880 | 59 | PASS |
|
||||
| sp21-w32 | 21 x 32 | 61.05 (42.8%) | 288.2 | 281.1 | 4.721 | 10.13x / 14.71x / 33.72x / 6.05x | 2878 | 57 | PASS |
|
||||
| sp16-w32 | 16 x 32 | 46.51 (32.6%) | 252.1 | 245.0 | 5.420 | 11.63x / 16.88x / 38.71x / 6.95x | 2872 | 55 | PASS |
|
||||
| sp11-w32 | 11 x 32 | 32.04 (22.5%) | 216.3 | 209.2 | 6.752 | 14.49x / 21.03x / 48.23x / 8.66x | 2880 | 54 | PASS |
|
||||
| sp43-w8 | 43 x 8 | 87.18 (61.2%) | 345.2 | 338.1 | 3.960 | 8.50x / 12.34x / 28.29x / 5.08x | 2872 | 59 | PASS |
|
||||
| sp340-w16 | 340 x 16 | 142.18 (99.7%) | 450.0 | 442.9 | 3.165 | 6.79x / 9.86x / 22.61x / 4.06x | 2800 | 65 | PASS |
|
||||
|
||||
|
||||
The follow-up on host c (the memory-clock try, the first self-tune rows, the microbench) overlapped its own queue on the card (the launcher's wait keyed on a pid file written after it read it), so those rows are discarded; the self-tune was re-run alone on the idle card (idle 5.2 W, 1 MiB used) and these are its rows:
|
||||
|
||||
|
||||
| Hold | Pack | Served shape | MH/s | Watts | Microjoules | Tune line |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 0.99 | mx8-genesis | sp37-w32 | 141.272 | 286.5 | 2.028 | tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.575 ceiling 142.575 hold 0.990 rows 170=142.008 127=142.276 85=142.224 63=141.866 42=141.549 31=140.347 36=141.083 39=141.295 37=141.240 chosen sp37-w32 141.240 (99.1% of the ceiling, 37 of 170 SMs) 21816 ms |
|
||||
| 0.99 | mx8_sh256x27 | sp82-w32 | 141.045 | 448.3 | 3.178 | tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.288 ceiling 142.288 hold 0.990 rows 170=141.986 127=141.787 85=141.302 63=135.712 74=140.204 79=140.561 82=141.046 80=140.855 chosen sp82-w32 141.046 (99.1% of the ceiling, 82 of 170 SMs) 19820 ms |
|
||||
| 0.98 | mx8-genesis | sp27-w32 | 139.717 | 284.1 | 2.033 | tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.277 ceiling 142.277 hold 0.980 rows 170=141.940 127=142.336 85=142.125 63=142.024 42=141.710 31=140.447 21=137.555 26=139.300 28=139.939 27=139.557 chosen sp27-w32 139.557 (98.1% of the ceiling, 27 of 170 SMs) 24133 ms |
|
||||
| 0.98 | mx8_sh256x27 | sp71-w32 | 140.082 | 446.1 | 3.185 | tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.534 ceiling 142.534 hold 0.980 rows 170=142.005 127=141.994 85=141.509 63=136.138 74=140.573 68=139.343 71=140.150 69=139.615 chosen sp71-w32 140.150 (98.3% of the ceiling, 71 of 170 SMs) 19850 ms |
|
||||
|
||||
|
||||
### RTX 5090, rented host d (Vast, Quebec; driver 595.91; power limit 475 W)
|
||||
|
||||
|
|
@ -276,6 +294,21 @@ RESULT memclocks label=4090 lmc=refused
|
|||
| tune0.99 | mx8_sh256x27 | sp61-w32 (61 x 32) | 70.42 | 346.9 | 4.926 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.624 ceiling 70.627 hold 0.990 rows 128=70.627 96=70.574 64=70.573 48=69.092 56=67.021 60=69.890 62=70.459 61=70.423 chosen sp61-w32 70.423 (99.7% of the ceiling, 61 of 128 SMs) 21979 ms |
|
||||
| tune0.98 | mx8-genesis | sp25-w32 (25 x 32) | 69.73 | 250.3 | 3.590 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.636 ceiling 70.636 hold 0.980 rows 128=70.622 96=70.618 64=70.574 48=70.516 32=70.472 24=67.159 28=70.391 26=70.159 25=69.728 chosen sp25-w32 69.728 (98.7% of the ceiling, 25 of 128 SMs) 24483 ms |
|
||||
| tune0.98 | mx8_sh256x27 | sp59-w32 (59 x 32) | 70.23 | 345.9 | 4.925 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.619 ceiling 70.619 hold 0.980 rows 128=70.619 96=70.581 64=70.564 48=69.094 56=67.027 60=69.901 58=69.197 59=70.270 chosen sp59-w32 70.270 (99.5% of the ceiling, 59 of 128 SMs) 22034 ms |
|
||||
| v5 | v5-genesis | ( x ) | 69.91 | 360.1 | 5.151 | 2715 | 10501 | PASS | |
|
||||
| v5 | v4-genesis | ( x ) | 69.89 | 360.0 | 5.151 | 2715 | 10501 | PASS | |
|
||||
| v5 | v4-genesis | base (0 x 1) | 69.91 | 360.6 | 5.158 | 2715 | 10501 | PASS | |
|
||||
|
||||
RESULT after-micro label=4090 watts=134.5 RESULT microbench probe=sleep status=ok unit="none (the SM-resident floor: full occupancy, __nanosleep, no issue)" ops_per_step=0 lanes=196608 regs=11 blocks_per_sm=6 steps=2000 launches=14056 launch_ms=2.1 seconds=30.0 start_utc=2026-10-08T12:32:09Z end_utc=2026-10-08T12:32:39Z G_steps_s=184.227 G_ops_s=0.000 checksum=f4a2875f29883636
|
||||
|
||||
|
||||
RESULT after-micro label=4090 watts=447.1 RESULT microbench probe=int_arx status=ok unit="int32 add, xor or rotate (4 independent chains, 3 ops each per step)" ops_per_step=12 lanes=196608 regs=12 blocks_per_sm=6 steps=16384 launches=20066 launch_ms=1.5 seconds=30.0 start_utc=2026-10-08T12:32:39Z end_utc=2026-10-08T12:33:09Z G_steps_s=2154.515 G_ops_s=25854.179 checksum=e8e10dc19924d864
|
||||
|
||||
|
||||
RESULT after-micro label=4090 watts=268.4 RESULT microbench probe=dram_chase_1g status=ok unit="dependent random 4-byte read in a 1 GiB table (the hash's own pattern, the control)" ops_per_step=1 lanes=196608 regs=16 blocks_per_sm=6 steps=512 launches=2795 launch_ms=10.7 seconds=30.0 start_utc=2026-10-08T12:33:09Z end_utc=2026-10-08T12:33:39Z G_steps_s=9.376 G_ops_s=9.376 checksum=7e9ab598fe7052e5
|
||||
|
||||
|
||||
RESULT after-idle label=4090 watts=65.7 sm_mhz=536.7 mem_mhz=3659.4
|
||||
|
||||
|
||||
|
||||
### H100 80 GB HBM3, rented (RunPod, 132 SMs; driver 580.126; power limit 700 W)
|
||||
|
|
@ -393,6 +426,24 @@ The memory-clock ladder: not landed.
|
|||
|
||||
| Card | Probe | Watts | SM MHz | Mem MHz | G ops or reads per s | Lanes | nJ or pJ per op over the sleep row |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 4090 | idle | 33.9 | 210.0 | 405.0 | | | |
|
||||
| 4090 | sleep | 133.1 | 2725.9 | 10501.0 | 0.000 | 196608 | |
|
||||
| 4090 | int_arx | 448.0 | 2574.1 | 10501.0 | 26076.595 | 196608 | 12.1 pJ |
|
||||
| 4090 | l2_chase_32m | 329.0 | 2715.0 | 10501.0 | 80.816 | 196608 | 2.42 nJ |
|
||||
| 4090 | l2_indep4_32m | 328.2 | 2715.0 | 10501.0 | 82.301 | 196608 | 2370.6 pJ |
|
||||
| 4090 | dram_chase_1g | 263.0 | 2715.0 | 10501.0 | 9.399 | 196608 | 13.82 nJ |
|
||||
| 5090c | idle | 184.2 | 2044.5 | 11141.9 | | | |
|
||||
| 5090c | sleep | 196.5 | 2865.0 | 13801.0 | 0.000 | 261120 | |
|
||||
| 5090c | int_arx | 372.9 | 2853.9 | 13801.0 | 13064.448 | 261120 | 13.5 pJ |
|
||||
| 5090c | l2_chase_32m | 382.2 | 2865.0 | 13801.0 | 53.964 | 261120 | 3.44 nJ |
|
||||
| 5090c | l2_indep4_32m | 401.4 | 2851.3 | 13801.0 | 51.085 | 261120 | 4011.0 pJ |
|
||||
| 5090c | dram_chase_1g | 372.9 | 2865.0 | 13801.0 | 8.019 | 261120 | 22.00 nJ |
|
||||
| 5090e | idle | 8.7 | 180.0 | 405.0 | | | |
|
||||
| 5090e | sleep | 139.0 | 2932.0 | 13801.0 | 0.000 | 261120 | |
|
||||
| 5090e | int_arx | 535.6 | 2854.9 | 13801.0 | 29035.421 | 261120 | 13.7 pJ |
|
||||
| 5090e | l2_chase_32m | 415.1 | 2902.0 | 13801.0 | 104.533 | 261120 | 2.64 nJ |
|
||||
| 5090e | l2_indep4_32m | 434.9 | 2902.0 | 13801.0 | 112.864 | 261120 | 2621.7 pJ |
|
||||
| 5090e | dram_chase_1g | 341.3 | 2904.0 | 13801.0 | 17.617 | 261120 | 11.48 nJ |
|
||||
| h100 | idle | 71.8 | 345.0 | 2619.0 | | | |
|
||||
| h100 | sleep | 130.3 | 1980.0 | 2619.0 | 0.000 | 270336 | |
|
||||
| h100 | int_arx | 440.0 | 1980.0 | 2619.0 | 21499.215 | 270336 | 14.4 pJ |
|
||||
|
|
@ -428,11 +479,23 @@ inside the L2: the L2 and crossbar path, no DRAM), `dram_chase_1g` (the hash's o
|
|||
through the same L2 and crossbar), `int_arx` (the ALU path). The table in section 1 carries the rows; the reading per
|
||||
card, watts over idle:
|
||||
|
||||
| Card | Idle | SM clock domain awake (sleep minus idle) | The L2 path per dependent read (l2 chase minus sleep, over reads per s) | The DRAM path per dependent read (dram chase minus sleep, over reads per s) | The hash's draw over idle at its read rate | The memory's own, modelled (chip-model 5.3) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| RTX 5090, host c | 7 W | see the probe table | | | 289 W at 18.3 G reads per s = 15.8 nJ per read whole-card | 2.0 nJ per read plus 20 W static |
|
||||
| RTX 4090 | 35 W | | | | 219 W at 9.0 G reads per s = 24.2 nJ per read | about 2.0 nJ plus 15 W |
|
||||
| H100 | 71 to 80 W | 57 W | | 11.5 nJ (450.3 W at 32.2 G reads per s) | 379 W at 32.6 G reads per s = 11.6 nJ per read | 1.2 nJ plus about 20 W |
|
||||
| Card | Idle | SM clock domain awake (sleep minus idle) | The L2 path per dependent read (l2 chase minus sleep, over reads per s) | The DRAM path per dependent read (dram chase minus sleep, over reads per s) | The hash's draw over idle at its read rate | Of which the memory path (hash minus sleep) | The memory's own, modelled (chip-model 5.3) | Label |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| RTX 5090, PC 1 (15.1a, 8 October, the microbench at stock) | 74.7 W | 45 W (120.2 at full residency) | 2.4 nJ (378.7 W at 107.6 G reads per s) | 10.9 nJ whole-card (318.4 W at 18.2 G); 8.5 nJ over the L2 path | 311 - 75 = 236 W at 17.6 G reads per s = 13.4 nJ per read (20.3's mx8 row) | 191 W = 10.9 nJ per read | 2.0 nJ per read plus 20 W static = 55 W, 3.1 nJ per read | measured; the memory's own modelled |
|
||||
| RTX 5090, rented host e (this file's probes) | 8.7 W at 180 MHz | 130 W (139.0 W at 2,932 MHz) | 2.6 nJ (415.1 W at 104.5 G reads per s) | 11.5 nJ over the sleep row (341.3 W at 17.6 G); 18.9 nJ whole-card | 352.7 - 8.7 = 344 W at 18.2 G reads per s = 18.9 nJ per read | 214 W = 11.8 nJ per read | the same 55 W | measured; modelled |
|
||||
| RTX 4090, rented | 33.9 W at 210 MHz | 99 W (133.1 W at 2,726 MHz with every warp resident and nothing issued) | 2.4 nJ (329.0 W at 80.8 G reads per s) | 13.8 nJ over the sleep row (263.0 W at 9.4 G); 24.4 nJ whole-card | 253.4 - 33.9 = 219 W at 9.04 G reads per s = 24.2 nJ per read | 120 W = 13.3 nJ per read | 2.0 nJ plus 15 W static = 33 W, 3.7 nJ per read (approximate for GDDR6X) | measured; modelled |
|
||||
| H100 80 GB HBM3, rented | 71.8 to 80 W at 345 MHz | 50 to 58 W (130.3 to 136.6 W at 1,980 MHz) | 2.9 nJ (479.1 W at 121.0 G reads per s) | 10.3 nJ over the sleep row (466.5 W at 32.6 G); 12.1 nJ whole-card | 450.4 - 71.3 = 379 W at 32.6 G reads per s = 11.6 nJ per read | 320 W = 9.8 nJ per read | 1.2 nJ plus about 20 W static = 59 W, 1.8 nJ per read | measured; modelled |
|
||||
|
||||
The reading: at the hash, each card's draw splits into a fixed "awake" floor (the card with every warp resident and
|
||||
nothing issuing: 120 to 139 W on the 5090 whatever its idle, 133 W on the 4090, 130 to 137 W on the H100; the SM clock
|
||||
tree, the GPCs' shared logic and the memory clock domain at their boost clocks; the SM-sparse rows say the SMs' own
|
||||
share of it is 4 to 15 W, because emptying three quarters of them saves that much) and a per-read term on the memory
|
||||
path (10.9 to 11.8 nJ per dependent read on the 5090 on two boards, 13.3 on the 4090, 9.8 on the H100) of which the
|
||||
memory devices themselves are 2.0 or 1.2 nJ modelled; the difference (8 to 11 nJ per read) is the L2 slices, the
|
||||
crossbar, the memory controllers and the PHY, in that order of distance from the SM, and the L2 path alone is 2.4 to
|
||||
2.9 nJ of it (the L2 chase). The ALU path is not in the hash's draw at all (512 ops per hash at 11 to 14 pJ is under
|
||||
1 W). The two 5090 boards agree on the per-read term within 8 percent and differ by 60 W on the idle (7 to 9 W on the
|
||||
rented Linux boards against 75 W on PC 1 with its display), which is where the across-host base spread comes from.
|
||||
|
||||
What the SM-sparse rows add to it: the SMs that go idle at the knee were worth 4 to 15 W on every card (the difference
|
||||
between the sleeping SM clock domain at full residency and the same domain with three quarters of its SMs empty), so
|
||||
|
|
@ -492,8 +555,13 @@ AMD or Apple owner gets nothing from this lane (the shapes are CUDA's; Metal and
|
|||
line goes out with the race line (`sparse tune device ... rows N=mhs ... chosen sp<N>-w32 (x percent of the ceiling)`)
|
||||
and the bench's RESULT line carries `variant=` and `sparse_blocks=`.
|
||||
- Measured: H100 15.2 s, 6 probes, `chosen sp127-w32 253.179 (99.5% of the ceiling)`, the served row 253.2 MH/s at
|
||||
441.6 W; class v4 on the H100 `chosen none` in 6.6 s (the persistent shape at 97.8 percent of the ceiling, under the
|
||||
hold); RTX 4090 24.5 s, `chosen sp26-w32 (99.3%)`, 70.15 MH/s at 250.2 W; the 5090 host c rows are in section 1.
|
||||
441.6 W (1.744 against 1.769 microjoules); class v4 on the H100 `chosen none` in 6.6 s (the persistent shape at 97.8
|
||||
percent of the ceiling, under the hold); RTX 4090 24.5 s, `chosen sp26-w32 (99.3%)`, 70.15 MH/s at 250.2 W (3.567
|
||||
against 3.588); RTX 5090 host c, class v3 at the 0.99 hold `chosen sp37-w32`, 141.27 MH/s at 286.5 W = 2.028
|
||||
against the base's 2.079 (2.5 percent), at 0.98 `sp27-w32` 2.033; class v4 at the 0.99 hold `sp82-w32` 3.178
|
||||
against 3.157 (nothing, inside noise under the board's 450 W cap), at 0.98 `sp71-w32` 3.185. The 0.99 hold is the
|
||||
default: on class v4 it costs 1 percent of rate for no watts, which is why Maximum turns it off and Balanced holds
|
||||
0.995.
|
||||
- Tests: `proto-cuda/nvrtc/emu/variant-test.cpp` case `sparse-tune` (the ladder from 170 SMs, the chooser on the 20.3b
|
||||
rows at holds 0.99, 0.98, 0.5 and 1.0, the mode reader on nine strings, the tuning keys for two cards): PASS beside
|
||||
the three earlier cases, built with the real headers on the H100 pod (`g++ -DIGNEUM_EMU`), 8 October 12:04 UTC. The
|
||||
|
|
|
|||
Loading…
Reference in a new issue