Class v6 floor lane 1: sm-sparse.md, the decomposition rows (the awake floor 120 to 139 W on every card, the memory path 10 to 13 nJ per dependent read against the devices' modelled 2.0 or 1.2), the 5090's clean self-tune rows (sp37 of 170 at 2.5 percent under base on class v3, nothing on class v4), every pod destroyed

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 12:45:48 +00:00
parent b1e940a260
commit e89fef216c

View file

@ -109,8 +109,26 @@ card NVIDIAGeForceRTX5090 driver 610.57.04 power limit 450.00 W idle 7.1 W at 21
| sp85-w32 | 85 x 32 | 141.46 (99.2%) | 450.0 | 442.9 | 3.181 | 6.83x / 9.91x / 22.72x / 4.08x | 2833 | 64 | PASS |
| sp64-w32 | 64 x 32 | 136.58 (95.8%) | 450.4 | 443.3 | 3.298 | 7.08x / 10.27x / 23.56x / 4.23x | 2850 | 63 | PASS |
| sp52-w32 | 52 x 32 | 121.00 (84.9%) | 417.9 | 410.8 | 3.454 | 7.41x / 10.76x / 24.67x / 4.43x | 2873 | 62 | PASS |
| sp43-w32 | 43 x 32 | 124.16 (87.1%) | 422.7 | 415.6 | 3.404 | 7.30x / 10.60x / 24.31x / 4.36x | 2872 | 62 | PASS |
| sp36-w32 | 36 x 32 | 104.35 (73.2%) | 380.5 | 373.4 | 3.646 | 7.82x / 11.36x / 26.04x / 4.67x | 2877 | 61 | PASS |
| sp28-w32 | 28 x 32 | 81.35 (57.1%) | 334.2 | 327.1 | 4.108 | 8.82x / 12.80x / 29.34x / 5.27x | 2880 | 59 | PASS |
| sp21-w32 | 21 x 32 | 61.05 (42.8%) | 288.2 | 281.1 | 4.721 | 10.13x / 14.71x / 33.72x / 6.05x | 2878 | 57 | PASS |
| sp16-w32 | 16 x 32 | 46.51 (32.6%) | 252.1 | 245.0 | 5.420 | 11.63x / 16.88x / 38.71x / 6.95x | 2872 | 55 | PASS |
| sp11-w32 | 11 x 32 | 32.04 (22.5%) | 216.3 | 209.2 | 6.752 | 14.49x / 21.03x / 48.23x / 8.66x | 2880 | 54 | PASS |
| sp43-w8 | 43 x 8 | 87.18 (61.2%) | 345.2 | 338.1 | 3.960 | 8.50x / 12.34x / 28.29x / 5.08x | 2872 | 59 | PASS |
| sp340-w16 | 340 x 16 | 142.18 (99.7%) | 450.0 | 442.9 | 3.165 | 6.79x / 9.86x / 22.61x / 4.06x | 2800 | 65 | PASS |
The follow-up on host c (the memory-clock try, the first self-tune rows, the microbench) overlapped its own queue on the card (the launcher's wait keyed on a pid file written after it read it), so those rows are discarded; the self-tune was re-run alone on the idle card (idle 5.2 W, 1 MiB used) and these are its rows:
| Hold | Pack | Served shape | MH/s | Watts | Microjoules | Tune line |
|---|---|---|---|---|---|---|
| 0.99 | mx8-genesis | sp37-w32 | 141.272 | 286.5 | 2.028 | tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.575 ceiling 142.575 hold 0.990 rows 170=142.008 127=142.276 85=142.224 63=141.866 42=141.549 31=140.347 36=141.083 39=141.295 37=141.240 chosen sp37-w32 141.240 (99.1% of the ceiling, 37 of 170 SMs) 21816 ms |
| 0.99 | mx8_sh256x27 | sp82-w32 | 141.045 | 448.3 | 3.178 | tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.288 ceiling 142.288 hold 0.990 rows 170=141.986 127=141.787 85=141.302 63=135.712 74=140.204 79=140.561 82=141.046 80=140.855 chosen sp82-w32 141.046 (99.1% of the ceiling, 82 of 170 SMs) 19820 ms |
| 0.98 | mx8-genesis | sp27-w32 | 139.717 | 284.1 | 2.033 | tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.277 ceiling 142.277 hold 0.980 rows 170=141.940 127=142.336 85=142.125 63=142.024 42=141.710 31=140.447 21=137.555 26=139.300 28=139.939 27=139.557 chosen sp27-w32 139.557 (98.1% of the ceiling, 27 of 170 SMs) 24133 ms |
| 0.98 | mx8_sh256x27 | sp71-w32 | 140.082 | 446.1 | 3.185 | tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.534 ceiling 142.534 hold 0.980 rows 170=142.005 127=141.994 85=141.509 63=136.138 74=140.573 68=139.343 71=140.150 69=139.615 chosen sp71-w32 140.150 (98.3% of the ceiling, 71 of 170 SMs) 19850 ms |
### RTX 5090, rented host d (Vast, Quebec; driver 595.91; power limit 475 W)
@ -276,6 +294,21 @@ RESULT memclocks label=4090 lmc=refused
| tune0.99 | mx8_sh256x27 | sp61-w32 (61 x 32) | 70.42 | 346.9 | 4.926 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.624 ceiling 70.627 hold 0.990 rows 128=70.627 96=70.574 64=70.573 48=69.092 56=67.021 60=69.890 62=70.459 61=70.423 chosen sp61-w32 70.423 (99.7% of the ceiling, 61 of 128 SMs) 21979 ms |
| tune0.98 | mx8-genesis | sp25-w32 (25 x 32) | 69.73 | 250.3 | 3.590 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.636 ceiling 70.636 hold 0.980 rows 128=70.622 96=70.618 64=70.574 48=70.516 32=70.472 24=67.159 28=70.391 26=70.159 25=69.728 chosen sp25-w32 69.728 (98.7% of the ceiling, 25 of 128 SMs) 24483 ms |
| tune0.98 | mx8_sh256x27 | sp59-w32 (59 x 32) | 70.23 | 345.9 | 4.925 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.619 ceiling 70.619 hold 0.980 rows 128=70.619 96=70.581 64=70.564 48=69.094 56=67.027 60=69.901 58=69.197 59=70.270 chosen sp59-w32 70.270 (99.5% of the ceiling, 59 of 128 SMs) 22034 ms |
| v5 | v5-genesis | ( x ) | 69.91 | 360.1 | 5.151 | 2715 | 10501 | PASS | |
| v5 | v4-genesis | ( x ) | 69.89 | 360.0 | 5.151 | 2715 | 10501 | PASS | |
| v5 | v4-genesis | base (0 x 1) | 69.91 | 360.6 | 5.158 | 2715 | 10501 | PASS | |
RESULT after-micro label=4090 watts=134.5 RESULT microbench probe=sleep status=ok unit="none (the SM-resident floor: full occupancy, __nanosleep, no issue)" ops_per_step=0 lanes=196608 regs=11 blocks_per_sm=6 steps=2000 launches=14056 launch_ms=2.1 seconds=30.0 start_utc=2026-10-08T12:32:09Z end_utc=2026-10-08T12:32:39Z G_steps_s=184.227 G_ops_s=0.000 checksum=f4a2875f29883636
RESULT after-micro label=4090 watts=447.1 RESULT microbench probe=int_arx status=ok unit="int32 add, xor or rotate (4 independent chains, 3 ops each per step)" ops_per_step=12 lanes=196608 regs=12 blocks_per_sm=6 steps=16384 launches=20066 launch_ms=1.5 seconds=30.0 start_utc=2026-10-08T12:32:39Z end_utc=2026-10-08T12:33:09Z G_steps_s=2154.515 G_ops_s=25854.179 checksum=e8e10dc19924d864
RESULT after-micro label=4090 watts=268.4 RESULT microbench probe=dram_chase_1g status=ok unit="dependent random 4-byte read in a 1 GiB table (the hash's own pattern, the control)" ops_per_step=1 lanes=196608 regs=16 blocks_per_sm=6 steps=512 launches=2795 launch_ms=10.7 seconds=30.0 start_utc=2026-10-08T12:33:09Z end_utc=2026-10-08T12:33:39Z G_steps_s=9.376 G_ops_s=9.376 checksum=7e9ab598fe7052e5
RESULT after-idle label=4090 watts=65.7 sm_mhz=536.7 mem_mhz=3659.4
### H100 80 GB HBM3, rented (RunPod, 132 SMs; driver 580.126; power limit 700 W)
@ -393,6 +426,24 @@ The memory-clock ladder: not landed.
| Card | Probe | Watts | SM MHz | Mem MHz | G ops or reads per s | Lanes | nJ or pJ per op over the sleep row |
|---|---|---|---|---|---|---|---|
| 4090 | idle | 33.9 | 210.0 | 405.0 | | | |
| 4090 | sleep | 133.1 | 2725.9 | 10501.0 | 0.000 | 196608 | |
| 4090 | int_arx | 448.0 | 2574.1 | 10501.0 | 26076.595 | 196608 | 12.1 pJ |
| 4090 | l2_chase_32m | 329.0 | 2715.0 | 10501.0 | 80.816 | 196608 | 2.42 nJ |
| 4090 | l2_indep4_32m | 328.2 | 2715.0 | 10501.0 | 82.301 | 196608 | 2370.6 pJ |
| 4090 | dram_chase_1g | 263.0 | 2715.0 | 10501.0 | 9.399 | 196608 | 13.82 nJ |
| 5090c | idle | 184.2 | 2044.5 | 11141.9 | | | |
| 5090c | sleep | 196.5 | 2865.0 | 13801.0 | 0.000 | 261120 | |
| 5090c | int_arx | 372.9 | 2853.9 | 13801.0 | 13064.448 | 261120 | 13.5 pJ |
| 5090c | l2_chase_32m | 382.2 | 2865.0 | 13801.0 | 53.964 | 261120 | 3.44 nJ |
| 5090c | l2_indep4_32m | 401.4 | 2851.3 | 13801.0 | 51.085 | 261120 | 4011.0 pJ |
| 5090c | dram_chase_1g | 372.9 | 2865.0 | 13801.0 | 8.019 | 261120 | 22.00 nJ |
| 5090e | idle | 8.7 | 180.0 | 405.0 | | | |
| 5090e | sleep | 139.0 | 2932.0 | 13801.0 | 0.000 | 261120 | |
| 5090e | int_arx | 535.6 | 2854.9 | 13801.0 | 29035.421 | 261120 | 13.7 pJ |
| 5090e | l2_chase_32m | 415.1 | 2902.0 | 13801.0 | 104.533 | 261120 | 2.64 nJ |
| 5090e | l2_indep4_32m | 434.9 | 2902.0 | 13801.0 | 112.864 | 261120 | 2621.7 pJ |
| 5090e | dram_chase_1g | 341.3 | 2904.0 | 13801.0 | 17.617 | 261120 | 11.48 nJ |
| h100 | idle | 71.8 | 345.0 | 2619.0 | | | |
| h100 | sleep | 130.3 | 1980.0 | 2619.0 | 0.000 | 270336 | |
| h100 | int_arx | 440.0 | 1980.0 | 2619.0 | 21499.215 | 270336 | 14.4 pJ |
@ -428,11 +479,23 @@ inside the L2: the L2 and crossbar path, no DRAM), `dram_chase_1g` (the hash's o
through the same L2 and crossbar), `int_arx` (the ALU path). The table in section 1 carries the rows; the reading per
card, watts over idle:
| Card | Idle | SM clock domain awake (sleep minus idle) | The L2 path per dependent read (l2 chase minus sleep, over reads per s) | The DRAM path per dependent read (dram chase minus sleep, over reads per s) | The hash's draw over idle at its read rate | The memory's own, modelled (chip-model 5.3) |
|---|---|---|---|---|---|---|
| RTX 5090, host c | 7 W | see the probe table | | | 289 W at 18.3 G reads per s = 15.8 nJ per read whole-card | 2.0 nJ per read plus 20 W static |
| RTX 4090 | 35 W | | | | 219 W at 9.0 G reads per s = 24.2 nJ per read | about 2.0 nJ plus 15 W |
| H100 | 71 to 80 W | 57 W | | 11.5 nJ (450.3 W at 32.2 G reads per s) | 379 W at 32.6 G reads per s = 11.6 nJ per read | 1.2 nJ plus about 20 W |
| Card | Idle | SM clock domain awake (sleep minus idle) | The L2 path per dependent read (l2 chase minus sleep, over reads per s) | The DRAM path per dependent read (dram chase minus sleep, over reads per s) | The hash's draw over idle at its read rate | Of which the memory path (hash minus sleep) | The memory's own, modelled (chip-model 5.3) | Label |
|---|---|---|---|---|---|---|---|---|
| RTX 5090, PC 1 (15.1a, 8 October, the microbench at stock) | 74.7 W | 45 W (120.2 at full residency) | 2.4 nJ (378.7 W at 107.6 G reads per s) | 10.9 nJ whole-card (318.4 W at 18.2 G); 8.5 nJ over the L2 path | 311 - 75 = 236 W at 17.6 G reads per s = 13.4 nJ per read (20.3's mx8 row) | 191 W = 10.9 nJ per read | 2.0 nJ per read plus 20 W static = 55 W, 3.1 nJ per read | measured; the memory's own modelled |
| RTX 5090, rented host e (this file's probes) | 8.7 W at 180 MHz | 130 W (139.0 W at 2,932 MHz) | 2.6 nJ (415.1 W at 104.5 G reads per s) | 11.5 nJ over the sleep row (341.3 W at 17.6 G); 18.9 nJ whole-card | 352.7 - 8.7 = 344 W at 18.2 G reads per s = 18.9 nJ per read | 214 W = 11.8 nJ per read | the same 55 W | measured; modelled |
| RTX 4090, rented | 33.9 W at 210 MHz | 99 W (133.1 W at 2,726 MHz with every warp resident and nothing issued) | 2.4 nJ (329.0 W at 80.8 G reads per s) | 13.8 nJ over the sleep row (263.0 W at 9.4 G); 24.4 nJ whole-card | 253.4 - 33.9 = 219 W at 9.04 G reads per s = 24.2 nJ per read | 120 W = 13.3 nJ per read | 2.0 nJ plus 15 W static = 33 W, 3.7 nJ per read (approximate for GDDR6X) | measured; modelled |
| H100 80 GB HBM3, rented | 71.8 to 80 W at 345 MHz | 50 to 58 W (130.3 to 136.6 W at 1,980 MHz) | 2.9 nJ (479.1 W at 121.0 G reads per s) | 10.3 nJ over the sleep row (466.5 W at 32.6 G); 12.1 nJ whole-card | 450.4 - 71.3 = 379 W at 32.6 G reads per s = 11.6 nJ per read | 320 W = 9.8 nJ per read | 1.2 nJ plus about 20 W static = 59 W, 1.8 nJ per read | measured; modelled |
The reading: at the hash, each card's draw splits into a fixed "awake" floor (the card with every warp resident and
nothing issuing: 120 to 139 W on the 5090 whatever its idle, 133 W on the 4090, 130 to 137 W on the H100; the SM clock
tree, the GPCs' shared logic and the memory clock domain at their boost clocks; the SM-sparse rows say the SMs' own
share of it is 4 to 15 W, because emptying three quarters of them saves that much) and a per-read term on the memory
path (10.9 to 11.8 nJ per dependent read on the 5090 on two boards, 13.3 on the 4090, 9.8 on the H100) of which the
memory devices themselves are 2.0 or 1.2 nJ modelled; the difference (8 to 11 nJ per read) is the L2 slices, the
crossbar, the memory controllers and the PHY, in that order of distance from the SM, and the L2 path alone is 2.4 to
2.9 nJ of it (the L2 chase). The ALU path is not in the hash's draw at all (512 ops per hash at 11 to 14 pJ is under
1 W). The two 5090 boards agree on the per-read term within 8 percent and differ by 60 W on the idle (7 to 9 W on the
rented Linux boards against 75 W on PC 1 with its display), which is where the across-host base spread comes from.
What the SM-sparse rows add to it: the SMs that go idle at the knee were worth 4 to 15 W on every card (the difference
between the sleeping SM clock domain at full residency and the same domain with three quarters of its SMs empty), so
@ -492,8 +555,13 @@ AMD or Apple owner gets nothing from this lane (the shapes are CUDA's; Metal and
line goes out with the race line (`sparse tune device ... rows N=mhs ... chosen sp<N>-w32 (x percent of the ceiling)`)
and the bench's RESULT line carries `variant=` and `sparse_blocks=`.
- Measured: H100 15.2 s, 6 probes, `chosen sp127-w32 253.179 (99.5% of the ceiling)`, the served row 253.2 MH/s at
441.6 W; class v4 on the H100 `chosen none` in 6.6 s (the persistent shape at 97.8 percent of the ceiling, under the
hold); RTX 4090 24.5 s, `chosen sp26-w32 (99.3%)`, 70.15 MH/s at 250.2 W; the 5090 host c rows are in section 1.
441.6 W (1.744 against 1.769 microjoules); class v4 on the H100 `chosen none` in 6.6 s (the persistent shape at 97.8
percent of the ceiling, under the hold); RTX 4090 24.5 s, `chosen sp26-w32 (99.3%)`, 70.15 MH/s at 250.2 W (3.567
against 3.588); RTX 5090 host c, class v3 at the 0.99 hold `chosen sp37-w32`, 141.27 MH/s at 286.5 W = 2.028
against the base's 2.079 (2.5 percent), at 0.98 `sp27-w32` 2.033; class v4 at the 0.99 hold `sp82-w32` 3.178
against 3.157 (nothing, inside noise under the board's 450 W cap), at 0.98 `sp71-w32` 3.185. The 0.99 hold is the
default: on class v4 it costs 1 percent of rate for no watts, which is why Maximum turns it off and Balanced holds
0.995.
- Tests: `proto-cuda/nvrtc/emu/variant-test.cpp` case `sparse-tune` (the ladder from 170 SMs, the chooser on the 20.3b
rows at holds 0.99, 0.98, 0.5 and 1.0, the mode reader on nine strings, the tuning keys for two cards): PASS beside
the three earlier cases, built with the real headers on the H100 pod (`g++ -DIGNEUM_EMU`), 8 October 12:04 UTC. The