igneum/docs/analysis/class-v6/floor/sm-sparse.md

70 KiB

Class v6 floor lane 1: SM-sparse, the honest card's watts toward the DRAM's own (8 October 2026)

Branch class-v6-floor-sm off counter-asic-4 (d8ba861e). The term this lane owns is E_card above E_mem in the identity of docs/analysis/counter-asic-4-research.md section 2: at zero shadow the stored-dataset chip's edge is E_card / E_mem, and the only honest-side lever is to take the card's watts toward the DRAM's own at the activate ceiling. Section 20.3b of that file read five shapes on PC 1's RTX 5090 (one block of 32 warps per SM at 170, 85, 43, 21 and 11 SMs) and closed the candidate: the draw follows the work, not the SM count. This file starts from that row, answers the question it left open (what the residual watts between idle-plus-memory and the card's draw ARE, and whether any occupancy shape at full SM count moves them) on four rented RTX 5090s, an RTX 4090 and an H100 plus PC 1 at the 1,300 MHz lock, measures the memory-clock knob where a host allows it, and ships the measurement as a worker option with a self-tune.

Every number here is measured unless marked otherwise: the worker's --bench (250 dispatches of 2^24 nonces, the pack's vector warps through the served kernel first, the 2^24 fingerprint on every row and equal across every shape of a pack) with nvidia-smi at 1 Hz beside it (the mean from 8 s in; power.draw, power.draw.instant and power.draw.average within 2 W on every row). The shapes: base is the shipped grid (one warp per block); w4 and w8 the shipped grid at 4 and 8 warps per block; sp<N>-w<W> the persistent wrapper of N blocks of W warps (the Counter ASIC 4.0 variant, bit-exact), and a %smid probe on every rented card confirmed that N blocks of 1,024 threads land on N distinct SMs, one block each, for every N used. The chip rows are the model's (chip-model-v3.md 5.3 to 5.6 and 5.12; lane B's SRAM die): GDDR7 board 0.466 microjoules per hash, one HBM3 stack 0.321, the N2 SRAM die 0.14; the M5 Max's 0.78 is measured (6 October). "Edge" is the card's microjoules over the chip's. Times are UTC in the logs and UK (BST, UTC+1) in the text. Idle is the card alone after 12 s with nothing running.

0. The answer

  1. The SM count is not the lever. On every card the draw over idle is flat across the SM count and the warps per SM for as long as the rate holds, and falls only when the rate falls. RTX 5090 (four rented hosts, class v3 at stock): 43 of 170 SMs holds 99.4 percent of the ceiling, 36 SMs 99.1, 28 SMs 98.4, 21 SMs 96.6, 16 SMs 91, 11 SMs 71; the best energy per hash on any host is 1.3 to 3.9 percent under its base (host c 2.053 against 2.079 microjoules at 43 SMs; host e 2.389 against 2.485). RTX 4090: 16 of 128 SMs holds 99.0 percent at 4 W under base; the best shape 1.2 percent under. H100: per-SM throughput-bound (the rate falls one for one with the SM count, 1.93 MH/s per SM), so its sparse floor is 127 of 132 SMs at 1.4 percent under base.
  2. Fewer warps per SM at full SM count reads the same as fewer SMs. 170 SMs at 8 warps (1,360 warps) and 43 SMs at 32 warps (1,376 warps) give 2.063 and 2.053 microjoules on the same host: the warps in flight set the rate, nothing in the occupancy sets the watts. One warp per SM (5,440 lanes) still holds 84 percent of the 5090's rate.
  3. Class v4's knee is 43 SMs by rate at stock and the full grid at the 1,300 MHz lock, and nothing is saved at either. Under 43 SMs the shadow is compute-bound at stock (PC 1: 36 SMs 88 percent, 28 SMs 69, 21 SMs 52, 16 SMs 39 of the rate); at 43 SMs the watts equal or exceed the base's. At the lock (PC 1, this file's job) the shadow needs every SM: 134.2 MH/s at 302 W on the full grid, 125 at 85 SMs, 64 at 43 SMs (3.24 against 2.25 microjoules). On the H100 the persistent shape itself loses 4 percent of the class v4 rate and the self-tune installs nothing. PC 1's class v3 at stock agrees with the rented ladder at every rung (43 SMs 136.5 of 136.7 MH/s at 10 W under base; 170 SMs x 8 warps 137.1 at 310.8 W, 3.5 percent under base, the best shape on that board).
  4. The residual is the core clock domain and the memory path on the GPU side, and the memory clock is not a lever either (section 3.1: the 5090's PHY has two states, 13,801 and 7,001 MHz; the half-rate state loses 43 percent of the rate for 37 percent of the watts, so the energy per hash rises 12 percent). The decomposition probes (section 3) put the 5090's draw over idle at the hash into: the resident SMs issuing nothing (the SM clock domain awake), the L2 and crossbar path per read, and the DRAM path per read on top of the memory's own modelled 2.0 nJ; the memory-clock ladder on PC 1 at the 1,300 lock reads the PHY's share directly. The honest floor per tier is the base row at stock within 2 percent, the core lock (rank 1 of the research file) and the memory-clock lock behind it, and the fraction a chip cannot strip is the memory side's own (section 4).
  5. What ships. --sm-sparse auto on the worker (section 5): one NVRTC compile of the persistent shape, a ladder from the SM count down, a bisection, the fewest blocks holding 99 percent of the ceiling installed, bit-exact or nothing; the tuning file's sm_sparse and sm_hold per card. Measured: 15 s on the H100 (chose 127 of 132 SMs, 1.4 percent of watts), 24 s on the 4090 (chose 26 of 128, 1.3 percent), nothing installed on class v4 where the shape loses rate. The design default is off; the tiers map to it (Efficiency on at 0.985, Balanced on at 0.995, Maximum off) and Ember re-runs it on every pack flip because the tune is part of the pair's build.

1. The rows

Each rented card's table: one row per shape per pack, the sampler's watts, microjoules per hash and the edge against the four chip rows. The 5090 hosts differ by board (power limit, idle, cooling), so the ladder is read as a percent of each host's own base and the energy as a difference within the host; the across-host spread of the base row itself (2.08 to 2.49 microjoules at 142 MH/s on class v3) is the rented-against-rented band the denominator lane also reads. Class v4 on hosts c, d and f is power-capped (the board's limit at 450 or 475 W, the clock throttling to 2,770 to 2,810 MHz), so those class v4 rows are a cap reading, not a free-running one; PC 1 (452.8 W, no cap) and host e (533.6 W) are the free-running class v4 rows.

RTX 5090, rented host c (Vast, Alberta; driver 610.57; the board's power limit 450 W of a 600 W default; idle 7 W)

knob pl=refused knob lgc=refused knob lmc=refused card NVIDIAGeForceRTX5090 driver 610.57.04 power limit 450.00 W idle 7.1 W at 218.8 MHz SM, 478.6 MHz memory

mx8-genesis (mx8)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 142.60 296.4 289.3 2.079 4.46x / 6.48x / 14.85x / 2.67x 2880 52 PASS
sp170-w32 170 x 32 142.14 (99.7%) 296.9 289.8 2.089 4.48x / 6.51x / 14.92x / 2.68x 2865 54 PASS
sp170-w16 170 x 16 142.55 (100.0%) 296.0 288.9 2.076 4.45x / 6.47x / 14.83x / 2.66x 2865 57 PASS
sp170-w8 170 x 8 142.41 (99.9%) 293.8 286.7 2.063 4.43x / 6.43x / 14.74x / 2.64x 2865 59 PASS
sp170-w4 170 x 4 141.88 (99.5%) 293.5 286.4 2.069 4.44x / 6.45x / 14.78x / 2.65x 2865 59 PASS
sp170-w2 170 x 2 139.31 (97.7%) 288.7 281.6 2.072 4.45x / 6.45x / 14.80x / 2.66x 2865 59 PASS
sp170-w1 170 x 1 119.59 (83.9%) 270.3 263.2 2.260 4.85x / 7.04x / 16.14x / 2.90x 2865 59 PASS
sp128-w32 128 x 32 142.60 (100.0%) 303.4 296.3 2.128 4.57x / 6.63x / 15.20x / 2.73x 2865 59 PASS
sp96-w32 96 x 32 138.96 (97.4%) 294.5 287.4 2.119 4.55x / 6.60x / 15.14x / 2.72x 2865 59 PASS
sp85-w32 85 x 32 142.42 (99.9%) 296.0 288.9 2.078 4.46x / 6.47x / 14.84x / 2.66x 2865 59 PASS
sp64-w32 64 x 32 142.24 (99.7%) 295.2 288.1 2.075 4.45x / 6.46x / 14.82x / 2.66x 2865 59 PASS
sp52-w32 52 x 32 136.01 (95.4%) 288.8 281.7 2.123 4.56x / 6.61x / 15.16x / 2.72x 2865 59 PASS
sp43-w32 43 x 32 141.77 (99.4%) 291.1 284.0 2.053 4.41x / 6.40x / 14.66x / 2.63x 2865 59 PASS
sp36-w32 36 x 32 141.38 (99.1%) 290.2 283.1 2.053 4.41x / 6.40x / 14.66x / 2.63x 2865 59 PASS
sp28-w32 28 x 32 140.27 (98.4%) 288.6 281.5 2.058 4.42x / 6.41x / 14.70x / 2.64x 2865 59 PASS
sp21-w32 21 x 32 137.81 (96.6%) 288.1 281.0 2.091 4.49x / 6.51x / 14.94x / 2.68x 2865 59 PASS
sp16-w32 16 x 32 129.98 (91.2%) 280.4 273.3 2.157 4.63x / 6.72x / 15.41x / 2.77x 2865 58 PASS
sp11-w32 11 x 32 100.88 (70.7%) 255.5 248.4 2.533 5.44x / 7.89x / 18.09x / 3.25x 2865 57 PASS
sp43-w8 43 x 8 139.54 (97.9%) 286.9 279.8 2.056 4.41x / 6.40x / 14.69x / 2.64x 2865 58 PASS
sp340-w16 340 x 16 142.57 (100.0%) 303.7 296.6 2.130 4.57x / 6.64x / 15.21x / 2.73x 2865 59 PASS

mx8_sh256x27 (mx8+sh256x27)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 142.55 450.0 442.9 3.157 6.77x / 9.83x / 22.55x / 4.05x 2797 66 PASS
sp170-w32 170 x 32 142.12 (99.7%) 450.0 442.9 3.166 6.79x / 9.86x / 22.61x / 4.06x 2810 64 PASS
sp170-w16 170 x 16 142.36 (99.9%) 450.0 442.9 3.161 6.78x / 9.85x / 22.58x / 4.05x 2811 64 PASS
sp170-w8 170 x 8 141.93 (99.6%) 450.0 442.9 3.171 6.80x / 9.88x / 22.65x / 4.07x 2796 64 PASS
sp170-w4 170 x 4 139.72 (98.0%) 450.0 442.9 3.221 6.91x / 10.03x / 23.01x / 4.13x 2802 64 PASS
sp170-w2 170 x 2 99.50 (69.8%) 389.4 382.3 3.913 8.40x / 12.19x / 27.95x / 5.02x 2865 62 PASS
sp170-w1 170 x 1 55.28 (38.8%) 287.7 280.6 5.204 11.17x / 16.21x / 37.17x / 6.67x 2865 57 PASS
sp128-w32 128 x 32 142.41 (99.9%) 450.0 442.9 3.160 6.78x / 9.84x / 22.57x / 4.05x 2823 62 PASS
sp96-w32 96 x 32 139.18 (97.6%) 450.0 442.9 3.233 6.94x / 10.07x / 23.09x / 4.14x 2846 64 PASS
sp85-w32 85 x 32 141.46 (99.2%) 450.0 442.9 3.181 6.83x / 9.91x / 22.72x / 4.08x 2833 64 PASS
sp64-w32 64 x 32 136.58 (95.8%) 450.4 443.3 3.298 7.08x / 10.27x / 23.56x / 4.23x 2850 63 PASS
sp52-w32 52 x 32 121.00 (84.9%) 417.9 410.8 3.454 7.41x / 10.76x / 24.67x / 4.43x 2873 62 PASS
sp43-w32 43 x 32 124.16 (87.1%) 422.7 415.6 3.404 7.30x / 10.60x / 24.31x / 4.36x 2872 62 PASS
sp36-w32 36 x 32 104.35 (73.2%) 380.5 373.4 3.646 7.82x / 11.36x / 26.04x / 4.67x 2877 61 PASS
sp28-w32 28 x 32 81.35 (57.1%) 334.2 327.1 4.108 8.82x / 12.80x / 29.34x / 5.27x 2880 59 PASS
sp21-w32 21 x 32 61.05 (42.8%) 288.2 281.1 4.721 10.13x / 14.71x / 33.72x / 6.05x 2878 57 PASS
sp16-w32 16 x 32 46.51 (32.6%) 252.1 245.0 5.420 11.63x / 16.88x / 38.71x / 6.95x 2872 55 PASS
sp11-w32 11 x 32 32.04 (22.5%) 216.3 209.2 6.752 14.49x / 21.03x / 48.23x / 8.66x 2880 54 PASS
sp43-w8 43 x 8 87.18 (61.2%) 345.2 338.1 3.960 8.50x / 12.34x / 28.29x / 5.08x 2872 59 PASS
sp340-w16 340 x 16 142.18 (99.7%) 450.0 442.9 3.165 6.79x / 9.86x / 22.61x / 4.06x 2800 65 PASS

The follow-up on host c (the memory-clock try, the first self-tune rows, the microbench) overlapped its own queue on the card (the launcher's wait keyed on a pid file written after it read it), so those rows are discarded; the self-tune was re-run alone on the idle card (idle 5.2 W, 1 MiB used) and these are its rows:

Hold Pack Served shape MH/s Watts Microjoules Tune line
0.99 mx8-genesis sp37-w32 141.272 286.5 2.028 tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.575 ceiling 142.575 hold 0.990 rows 170=142.008 127=142.276 85=142.224 63=141.866 42=141.549 31=140.347 36=141.083 39=141.295 37=141.240 chosen sp37-w32 141.240 (99.1% of the ceiling, 37 of 170 SMs) 21816 ms
0.99 mx8_sh256x27 sp82-w32 141.045 448.3 3.178 tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.288 ceiling 142.288 hold 0.990 rows 170=141.986 127=141.787 85=141.302 63=135.712 74=140.204 79=140.561 82=141.046 80=140.855 chosen sp82-w32 141.046 (99.1% of the ceiling, 82 of 170 SMs) 19820 ms
0.98 mx8-genesis sp27-w32 139.717 284.1 2.033 tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.277 ceiling 142.277 hold 0.980 rows 170=141.940 127=142.336 85=142.125 63=142.024 42=141.710 31=140.447 21=137.555 26=139.300 28=139.939 27=139.557 chosen sp27-w32 139.557 (98.1% of the ceiling, 27 of 170 SMs) 24133 ms
0.98 mx8_sh256x27 sp71-w32 140.082 446.1 3.185 tune device NVIDIA_GeForce_RTX_5090 sms 170 installed base 142.534 ceiling 142.534 hold 0.980 rows 170=142.005 127=141.994 85=141.509 63=136.138 74=140.573 68=139.343 71=140.150 69=139.615 chosen sp71-w32 140.150 (98.3% of the ceiling, 71 of 170 SMs) 19850 ms

RTX 5090, rented host d (Vast, Quebec; driver 595.91; power limit 475 W)

knob pl=refused knob lgc=refused knob lmc=refused card NVIDIAGeForceRTX5090 driver 595.91.07 power limit 475.00 W idle 27.4 W at 195.0 MHz SM, 405.0 MHz memory

mx8-genesis (mx8)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 142.44 326.3 298.9 2.291 4.92x / 7.14x / 16.36x / 2.94x 2887 50 PASS
sp170-w32 170 x 32 141.88 (99.6%) 327.5 300.1 2.308 4.95x / 7.19x / 16.49x / 2.96x 2886 53 PASS
sp128-w32 128 x 32 142.28 (99.9%) 327.7 300.3 2.303 4.94x / 7.17x / 16.45x / 2.95x 2872 55 PASS
sp96-w32 96 x 32 138.83 (97.5%) 323.4 296.0 2.330 5.00x / 7.26x / 16.64x / 2.99x 2872 56 PASS
sp85-w32 85 x 32 142.23 (99.9%) 328.1 300.7 2.307 4.95x / 7.19x / 16.48x / 2.96x 2872 57 PASS
sp64-w32 64 x 32 142.03 (99.7%) 329.0 301.6 2.316 4.97x / 7.21x / 16.54x / 2.97x 2872 58 PASS

mx8_sh256x27 (mx8+sh256x27)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 142.36 475.0 447.6 3.337 7.16x / 10.40x / 23.84x / 4.28x 2777 66 PASS
sp170-w32 170 x 32 141.98 (99.7%) 475.0 447.6 3.345 7.18x / 10.42x / 23.89x / 4.29x 2767 68 PASS
sp128-w32 128 x 32 142.22 (99.9%) 475.0 447.6 3.340 7.17x / 10.40x / 23.86x / 4.28x 2764 69 PASS
sp96-w32 96 x 32 139.08 (97.7%) 475.0 447.6 3.415 7.33x / 10.64x / 24.39x / 4.38x 2787 69 PASS
sp85-w32 85 x 32 141.27 (99.2%) 475.0 447.6 3.362 7.21x / 10.47x / 24.01x / 4.31x 2772 70 PASS
sp64-w32 64 x 32 135.94 (95.5%) 475.0 447.6 3.494 7.50x / 10.88x / 24.96x / 4.48x 2803 70 PASS

RTX 5090, rented host e (Vast, United Kingdom; driver 595.91)

knob pl=refused knob lgc=refused knob lmc=refused card NVIDIAGeForceRTX5090 driver 595.91.07 power limit 575.00 W idle 11.7 W at 185.2 MHz SM, 445.5 MHz memory

mx8-genesis (mx8)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 141.96 352.7 341.0 2.485 5.33x / 7.74x / 17.75x / 3.19x 2917 56 PASS
sp52-w32 52 x 32 135.85 (95.7%) 332.0 320.3 2.444 5.24x / 7.61x / 17.46x / 3.13x 2922 56 PASS
sp43-w32 43 x 32 141.21 (99.5%) 337.4 325.7 2.389 5.13x / 7.44x / 17.06x / 3.06x 2917 57 PASS
sp36-w32 36 x 32 140.88 (99.2%) 336.9 325.2 2.391 5.13x / 7.45x / 17.08x / 3.07x 2917 58 PASS
sp28-w32 28 x 32 139.58 (98.3%) 335.6 323.9 2.404 5.16x / 7.49x / 17.17x / 3.08x 2917 58 PASS
sp21-w32 21 x 32 137.23 (96.7%) 334.9 323.2 2.440 5.24x / 7.60x / 17.43x / 3.13x 2917 58 PASS

mx8_sh256x27 (mx8+sh256x27)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 141.90 533.6 521.9 3.760 8.07x / 11.71x / 26.86x / 4.82x 2902 68 PASS
sp52-w32 52 x 32 122.17 (86.1%) 484.4 472.7 3.965 8.51x / 12.35x / 28.32x / 5.08x 2903 68 PASS
sp43-w32 43 x 32 125.11 (88.2%) 491.8 480.1 3.931 8.44x / 12.25x / 28.08x / 5.04x 2902 69 PASS
sp36-w32 36 x 32 105.18 (74.1%) 444.6 432.9 4.227 9.07x / 13.17x / 30.19x / 5.42x 2902 68 PASS
sp28-w32 28 x 32 81.73 (57.6%) 389.2 377.5 4.762 10.22x / 14.83x / 34.01x / 6.11x 2896 66 PASS
sp21-w32 21 x 32 61.61 (43.4%) 340.4 328.7 5.525 11.86x / 17.21x / 39.46x / 7.08x 2902 63 PASS

RTX 5090, rented host f (Vast, Beijing; driver 580.173; power limit 450 W)

knob pl=refused knob lgc=refused knob lmc=refused card NVIDIAGeForceRTX5090 driver 580.173.02 power limit 450.00 W idle 29.8 W at 300.0 MHz SM, 405.0 MHz memory

mx8-genesis (mx8)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 142.49 333.2 303.4 2.338 5.02x / 7.28x / 16.70x / 3.00x 3028 47 PASS
sp16-w32 16 x 32 130.53 (91.6%) 309.8 280.0 2.373 5.09x / 7.39x / 16.95x / 3.04x 3030 48 PASS
sp11-w32 11 x 32 106.04 (74.4%) 288.1 258.3 2.717 5.83x / 8.46x / 19.41x / 3.48x 3030 48 PASS
sp170-w8 170 x 8 142.32 (99.9%) 322.8 293.0 2.268 4.87x / 7.07x / 16.20x / 2.91x 3030 49 PASS
sp170-w4 170 x 4 141.72 (99.5%) 323.3 293.5 2.281 4.89x / 7.11x / 16.29x / 2.92x 3030 50 PASS
sp170-w2 170 x 2 139.28 (97.8%) 320.6 290.8 2.302 4.94x / 7.17x / 16.44x / 2.95x 3030 51 PASS
sp43-w8 43 x 8 139.46 (97.9%) 320.0 290.2 2.295 4.92x / 7.15x / 16.39x / 2.94x 3030 52 PASS

mx8_sh256x27 (mx8+sh256x27)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 142.39 450.0 420.2 3.160 6.78x / 9.84x / 22.57x / 4.05x 2882 58 PASS
sp16-w32 16 x 32 48.94 (34.4%) 298.0 268.2 6.090 13.07x / 18.97x / 43.50x / 7.81x 3022 54 PASS
sp11-w32 11 x 32 33.79 (23.7%) 259.2 229.4 7.671 16.46x / 23.90x / 54.79x / 9.83x 3037 52 PASS
sp170-w8 170 x 8 141.83 (99.6%) 450.0 420.2 3.173 6.81x / 9.88x / 22.66x / 4.07x 2853 59 PASS
sp170-w4 170 x 4 139.72 (98.1%) 450.0 420.2 3.221 6.91x / 10.03x / 23.01x / 4.13x 2854 60 PASS
sp170-w2 170 x 2 102.45 (72.0%) 443.5 413.7 4.329 9.29x / 13.49x / 30.92x / 5.55x 3012 62 PASS
sp43-w8 43 x 8 90.20 (63.3%) 398.7 368.9 4.420 9.48x / 13.77x / 31.57x / 5.67x 3022 60 PASS

RTX 4090, rented (RunPod, 128 SMs, GDDR6X; driver 595.91; power limit 450 W)

knob pl=refused knob lgc=refused knob lmc=refused card NVIDIAGeForceRTX4090 driver 595.91.07 power limit 450.00 W idle 34.7 W at 210.0 MHz SM, 405.0 MHz memory

mx8-genesis (mx8)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 70.63 253.4 218.7 3.588 7.70x / 11.18x / 25.63x / 4.60x 2732 39 PASS
w4 0 x 4 70.62 (100.0%) 254.9 220.2 3.609 7.74x / 11.24x / 25.78x / 4.63x 2730 44 PASS
w8 0 x 8 70.64 (100.0%) 255.6 220.9 3.618 7.76x / 11.27x / 25.84x / 4.64x 2721 47 PASS
sp128-w32 128 x 32 70.64 (100.0%) 255.6 220.9 3.618 7.76x / 11.27x / 25.84x / 4.64x 2715 48 PASS
sp256-w16 256 x 16 70.57 (99.9%) 256.1 221.4 3.629 7.79x / 11.31x / 25.92x / 4.65x 2715 48 PASS
sp512-w8 512 x 8 70.60 (100.0%) 256.0 221.3 3.626 7.78x / 11.30x / 25.90x / 4.65x 2715 49 PASS
sp128-w16 128 x 16 70.64 (100.0%) 254.2 219.5 3.599 7.72x / 11.21x / 25.71x / 4.61x 2715 49 PASS
sp128-w8 128 x 8 70.58 (99.9%) 252.0 217.3 3.570 7.66x / 11.12x / 25.50x / 4.58x 2715 49 PASS
sp128-w4 128 x 4 70.45 (99.7%) 251.3 216.6 3.567 7.65x / 11.11x / 25.48x / 4.57x 2715 48 PASS
sp128-w2 128 x 2 70.17 (99.4%) 250.6 215.9 3.571 7.66x / 11.12x / 25.51x / 4.58x 2715 48 PASS
sp128-w1 128 x 1 66.28 (93.8%) 242.7 208.0 3.662 7.86x / 11.41x / 26.16x / 4.69x 2715 47 PASS
sp96-w32 96 x 32 70.61 (100.0%) 253.7 219.0 3.593 7.71x / 11.19x / 25.66x / 4.61x 2715 48 PASS
sp64-w32 64 x 32 70.59 (100.0%) 253.1 218.4 3.585 7.69x / 11.17x / 25.61x / 4.60x 2715 48 PASS
sp48-w32 48 x 32 70.50 (99.8%) 252.2 217.5 3.578 7.68x / 11.15x / 25.56x / 4.59x 2715 48 PASS
sp32-w32 32 x 32 70.42 (99.7%) 251.4 216.7 3.570 7.66x / 11.12x / 25.50x / 4.58x 2715 48 PASS
sp24-w32 24 x 32 67.15 (95.1%) 247.1 212.4 3.680 7.90x / 11.46x / 26.29x / 4.72x 2715 48 PASS
sp16-w32 16 x 32 69.95 (99.0%) 249.6 214.9 3.568 7.66x / 11.12x / 25.49x / 4.57x 2715 47 PASS
sp12-w32 12 x 32 68.69 (97.3%) 246.5 211.8 3.589 7.70x / 11.18x / 25.64x / 4.60x 2715 47 PASS
sp8-w32 8 x 32 65.07 (92.1%) 239.1 204.4 3.675 7.89x / 11.45x / 26.25x / 4.71x 2715 46 PASS
sp32-w8 32 x 8 70.19 (99.4%) 248.9 214.2 3.546 7.61x / 11.05x / 25.33x / 4.55x 2715 47 PASS

mx8_sh256x27 (mx8+sh256x27)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 70.63 353.5 318.8 5.005 10.74x / 15.59x / 35.75x / 6.42x 2715 56 PASS
w4 0 x 4 70.62 (100.0%) 356.8 322.1 5.052 10.84x / 15.74x / 36.09x / 6.48x 2715 58 PASS
w8 0 x 8 70.63 (100.0%) 357.6 322.9 5.063 10.86x / 15.77x / 36.16x / 6.49x 2715 58 PASS
sp128-w32 128 x 32 70.60 (100.0%) 356.1 321.4 5.044 10.82x / 15.71x / 36.03x / 6.47x 2715 59 PASS
sp256-w16 256 x 16 70.58 (99.9%) 356.9 322.2 5.057 10.85x / 15.75x / 36.12x / 6.48x 2715 59 PASS
sp512-w8 512 x 8 70.57 (99.9%) 358.4 323.7 5.078 10.90x / 15.82x / 36.27x / 6.51x 2715 59 PASS
sp128-w16 128 x 16 70.54 (99.9%) 355.6 320.9 5.041 10.82x / 15.70x / 36.01x / 6.46x 2715 59 PASS
sp128-w8 128 x 8 70.53 (99.9%) 355.0 320.3 5.034 10.80x / 15.68x / 35.96x / 6.45x 2715 59 PASS
sp128-w4 128 x 4 70.30 (99.5%) 354.4 319.7 5.041 10.82x / 15.70x / 36.01x / 6.46x 2715 59 PASS
sp128-w2 128 x 2 63.43 (89.8%) 332.5 297.8 5.242 11.25x / 16.33x / 37.44x / 6.72x 2715 57 PASS
sp128-w1 128 x 1 40.51 (57.4%) 267.5 232.8 6.603 14.17x / 20.57x / 47.16x / 8.47x 2715 54 PASS
sp96-w32 96 x 32 70.57 (99.9%) 350.9 316.2 4.972 10.67x / 15.49x / 35.51x / 6.37x 2715 58 PASS
sp64-w32 64 x 32 70.57 (99.9%) 349.7 315.0 4.955 10.63x / 15.44x / 35.39x / 6.35x 2715 59 PASS
sp48-w32 48 x 32 69.10 (97.8%) 345.1 310.4 4.994 10.72x / 15.56x / 35.67x / 6.40x 2715 59 PASS
sp32-w32 32 x 32 68.22 (96.6%) 341.4 306.7 5.004 10.74x / 15.59x / 35.74x / 6.42x 2715 58 PASS
sp24-w32 24 x 32 66.93 (94.8%) 336.6 301.9 5.029 10.79x / 15.67x / 35.92x / 6.45x 2715 58 PASS
sp16-w32 16 x 32 44.84 (63.5%) 276.4 241.7 6.163 13.23x / 19.20x / 44.02x / 7.90x 2715 54 PASS
sp12-w32 12 x 32 33.66 (47.7%) 242.9 208.2 7.216 15.48x / 22.48x / 51.54x / 9.25x 2715 51 PASS
sp8-w32 8 x 32 22.46 (31.8%) 208.0 173.3 9.261 19.87x / 28.85x / 66.15x / 11.87x 2715 48 PASS
sp32-w8 32 x 8 58.36 (82.6%) 311.3 276.6 5.334 11.45x / 16.62x / 38.10x / 6.84x 2715 54 PASS

follow-up, 4090

Tag Pack Served shape MH/s Watts Microjoules SM MHz Mem MHz Check Tune line

RESULT memclocks label=4090 lmc=refused

| tune0.99 | mx8-genesis | sp26-w32 (26 x 32) | 70.15 | 250.2 | 3.567 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.632 ceiling 70.632 hold 0.990 rows 128=70.609 96=70.616 64=70.590 48=70.524 32=70.476 24=67.148 28=70.377 26=70.143 25=69.726 chosen sp26-w32 70.143 (99.3% of the ceiling, 26 of 128 SMs) 24485 ms | | tune0.99 | mx8_sh256x27 | sp61-w32 (61 x 32) | 70.42 | 346.9 | 4.926 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.624 ceiling 70.627 hold 0.990 rows 128=70.627 96=70.574 64=70.573 48=69.092 56=67.021 60=69.890 62=70.459 61=70.423 chosen sp61-w32 70.423 (99.7% of the ceiling, 61 of 128 SMs) 21979 ms | | tune0.98 | mx8-genesis | sp25-w32 (25 x 32) | 69.73 | 250.3 | 3.590 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.636 ceiling 70.636 hold 0.980 rows 128=70.622 96=70.618 64=70.574 48=70.516 32=70.472 24=67.159 28=70.391 26=70.159 25=69.728 chosen sp25-w32 69.728 (98.7% of the ceiling, 25 of 128 SMs) 24483 ms | | tune0.98 | mx8_sh256x27 | sp59-w32 (59 x 32) | 70.23 | 345.9 | 4.925 | 2715 | 10501 | PASS | RESULT sparse tune device NVIDIA_GeForce_RTX_4090 sms 128 installed base 70.619 ceiling 70.619 hold 0.980 rows 128=70.619 96=70.581 64=70.564 48=69.094 56=67.027 60=69.901 58=69.197 59=70.270 chosen sp59-w32 70.270 (99.5% of the ceiling, 59 of 128 SMs) 22034 ms | | v5 | v5-genesis | ( x ) | 69.91 | 360.1 | 5.151 | 2715 | 10501 | PASS | | | v5 | v4-genesis | ( x ) | 69.89 | 360.0 | 5.151 | 2715 | 10501 | PASS | | | v5 | v4-genesis | base (0 x 1) | 69.91 | 360.6 | 5.158 | 2715 | 10501 | PASS | |

RESULT after-micro label=4090 watts=134.5 RESULT microbench probe=sleep status=ok unit="none (the SM-resident floor: full occupancy, __nanosleep, no issue)" ops_per_step=0 lanes=196608 regs=11 blocks_per_sm=6 steps=2000 launches=14056 launch_ms=2.1 seconds=30.0 start_utc=2026-10-08T12:32:09Z end_utc=2026-10-08T12:32:39Z G_steps_s=184.227 G_ops_s=0.000 checksum=f4a2875f29883636

RESULT after-micro label=4090 watts=447.1 RESULT microbench probe=int_arx status=ok unit="int32 add, xor or rotate (4 independent chains, 3 ops each per step)" ops_per_step=12 lanes=196608 regs=12 blocks_per_sm=6 steps=16384 launches=20066 launch_ms=1.5 seconds=30.0 start_utc=2026-10-08T12:32:39Z end_utc=2026-10-08T12:33:09Z G_steps_s=2154.515 G_ops_s=25854.179 checksum=e8e10dc19924d864

RESULT after-micro label=4090 watts=268.4 RESULT microbench probe=dram_chase_1g status=ok unit="dependent random 4-byte read in a 1 GiB table (the hash's own pattern, the control)" ops_per_step=1 lanes=196608 regs=16 blocks_per_sm=6 steps=512 launches=2795 launch_ms=10.7 seconds=30.0 start_utc=2026-10-08T12:33:09Z end_utc=2026-10-08T12:33:39Z G_steps_s=9.376 G_ops_s=9.376 checksum=7e9ab598fe7052e5

RESULT after-idle label=4090 watts=65.7 sm_mhz=536.7 mem_mhz=3659.4

H100 80 GB HBM3, rented (RunPod, 132 SMs; driver 580.126; power limit 700 W)

knob pl=refused knob lgc=refused knob lmc=allowed card NVIDIAH10080GBHBM3 driver 580.126.09 power limit 700.00 W idle 71.3 W at 345.0 MHz SM, 2619.0 MHz memory

mx8-genesis (mx8)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 254.63 450.4 379.1 1.769 3.80x / 5.51x / 12.64x / 2.27x 1980 44 PASS
w4 0 x 4 254.35 (99.9%) 454.2 382.9 1.786 3.83x / 5.56x / 12.76x / 2.29x 1980 46 PASS
w8 0 x 8 254.52 (100.0%) 457.2 385.9 1.796 3.85x / 5.60x / 12.83x / 2.30x 1980 50 PASS
sp132-w32 132 x 32 253.98 (99.7%) 453.4 382.1 1.785 3.83x / 5.56x / 12.75x / 2.29x 1980 52 PASS
sp264-w16 264 x 16 254.20 (99.8%) 457.4 386.1 1.799 3.86x / 5.60x / 12.85x / 2.31x 1980 54 PASS
sp528-w8 528 x 8 254.20 (99.8%) 461.5 390.2 1.816 3.90x / 5.66x / 12.97x / 2.33x 1980 55 PASS
sp132-w16 132 x 16 253.94 (99.7%) 454.1 382.8 1.788 3.84x / 5.57x / 12.77x / 2.29x 1980 56 PASS
sp132-w8 132 x 8 252.53 (99.2%) 454.7 383.4 1.801 3.86x / 5.61x / 12.86x / 2.31x 1980 56 PASS
sp132-w4 132 x 4 237.53 (93.3%) 430.1 358.8 1.811 3.89x / 5.64x / 12.94x / 2.32x 1980 56 PASS
sp132-w2 132 x 2 178.50 (70.1%) PASS
sp132-w1 132 x 1 105.25 (41.3%) 277.6 206.3 2.637 5.66x / 8.21x / 18.84x / 3.38x 1980 49 PASS
sp99-w32 99 x 32 212.70 (83.5%) 402.8 331.5 1.894 4.06x / 5.90x / 13.53x / 2.43x 1980 50 PASS
sp66-w32 66 x 32 189.20 (74.3%) 372.9 301.6 1.971 4.23x / 6.14x / 14.08x / 2.53x 1980 48 PASS
sp44-w32 44 x 32 150.11 (58.9%) 323.9 252.6 2.158 4.63x / 6.72x / 15.41x / 2.77x 1980 46 PASS
sp33-w32 33 x 32 123.41 (48.5%) 291.0 219.7 2.358 5.06x / 7.35x / 16.84x / 3.02x 1980 44 PASS
sp24-w32 24 x 32 89.90 (35.3%) 250.2 178.9 2.783 5.97x / 8.67x / 19.88x / 3.57x 1980 42 PASS
sp17-w32 17 x 32 63.66 (25.0%) 219.1 147.8 3.442 7.39x / 10.72x / 24.59x / 4.41x 1980 41 PASS
sp12-w32 12 x 32 55.38 (21.7%) 210.8 139.5 3.806 8.17x / 11.86x / 27.19x / 4.88x 1980 40 PASS
sp8-w32 8 x 32 36.93 (14.5%) 189.3 118.0 5.126 11.00x / 15.97x / 36.61x / 6.57x 1980 40 PASS
sp33-w8 33 x 8 123.50 (48.5%) 293.2 221.9 2.374 5.09x / 7.40x / 16.96x / 3.04x 1980 47 PASS

mx8_sh256x27 (mx8+sh256x27)

Shape Blocks x warps MH/s Watts Watts over idle Microjoules per hash Edge: GDDR7 / HBM3 / SRAM / M5 Max SM MHz Temp Check
base 0 x 1 251.90 699.4 628.1 2.776 5.96x / 8.65x / 19.83x / 3.56x 1830 66 PASS
w4 0 x 4 250.56 (99.5%) 698.5 627.2 2.788 5.98x / 8.69x / 19.91x / 3.57x 1748 69 PASS
w8 0 x 8 249.31 (99.0%) 697.6 626.3 2.798 6.00x / 8.72x / 19.99x / 3.59x 1734 70 PASS
sp132-w32 132 x 32 242.15 (96.1%) 696.7 625.4 2.877 6.17x / 8.96x / 20.55x / 3.69x 1744 72 PASS
sp264-w16 264 x 16 237.66 (94.3%) 698.5 627.2 2.939 6.31x / 9.16x / 20.99x / 3.77x 1795 71 PASS
sp528-w8 528 x 8 231.67 (92.0%) 695.7 624.4 3.003 6.44x / 9.36x / 21.45x / 3.85x 1768 69 PASS
sp132-w16 132 x 16 226.18 (89.8%) 699.4 628.1 3.092 6.64x / 9.63x / 22.09x / 3.96x 1857 69 PASS
sp132-w8 132 x 8 173.30 (68.8%) 630.0 558.7 3.635 7.80x / 11.32x / 25.96x / 4.66x 1980 66 PASS
sp132-w4 132 x 4 108.07 (42.9%) 452.5 381.2 4.187 8.98x / 13.04x / 29.91x / 5.37x 1980 57 PASS
sp132-w2 132 x 2 54.90 (21.8%) 301.3 230.0 5.488 11.78x / 17.10x / 39.20x / 7.04x 1980 49 PASS
sp132-w1 132 x 1 29.45 (11.7%) 226.4 155.1 7.688 16.50x / 23.95x / 54.91x / 9.86x 1980 44 PASS
sp99-w32 99 x 32 206.11 (81.8%) 697.2 625.9 3.383 7.26x / 10.54x / 24.16x / 4.34x 1967 63 PASS
sp66-w32 66 x 32 138.62 (55.0%) 520.4 449.1 3.754 8.06x / 11.69x / 26.81x / 4.81x 1980 57 PASS
sp44-w32 44 x 32 92.59 (36.8%) 392.6 321.3 4.240 9.10x / 13.21x / 30.29x / 5.44x 1980 52 PASS
sp33-w32 33 x 32 69.57 (27.6%) 330.5 259.2 4.751 10.20x / 14.80x / 33.94x / 6.09x 1980 49 PASS
sp24-w32 24 x 32 50.64 (20.1%) 279.1 207.8 5.511 11.83x / 17.17x / 39.36x / 7.07x 1980 46 PASS
sp17-w32 17 x 32 35.89 (14.2%) 238.0 166.7 6.631 14.23x / 20.66x / 47.36x / 8.50x 1980 43 PASS
sp12-w32 12 x 32 25.34 (10.1%) 207.5 136.2 8.190 17.58x / 25.51x / 58.50x / 10.50x 1980 41 PASS
sp8-w32 8 x 32 16.90 (6.7%) 182.5 111.2 10.798 23.17x / 33.64x / 77.13x / 13.84x 1980 39 PASS
sp33-w8 33 x 8 45.06 (17.9%) 262.6 191.3 5.827 12.50x / 18.15x / 41.62x / 7.47x 1980 43 PASS

follow-up, h100

Tag Pack Served shape MH/s Watts Microjoules SM MHz Mem MHz Check Tune line

RESULT memclocks label=h100 supported=[2619 1593 ]

| mem2619 | mx8-genesis | base (0 x 1) | 254.84 | 454.6 | 1.784 | 1980 | 2619 | PASS | | | mem2619 | mx8_sh256x27 | base (0 x 1) | 253.00 | 699.4 | 2.764 | 1716 | 2619 | PASS | | | mem1593 | mx8-genesis | base (0 x 1) | 254.57 | 458.8 | 1.802 | 1980 | 2619 | PASS | | | mem1593 | mx8_sh256x27 | base (0 x 1) | 252.77 | 698.4 | 2.763 | 1694 | 2619 | PASS | | | tune0.99 | mx8-genesis | sp127-w32 (127 x 32) | 253.19 | 441.6 | 1.744 | 1980 | 2619 | PASS | RESULT sparse tune device NVIDIA_H100_80GB_HBM3 sms 132 installed base 254.458 ceiling 254.458 hold 0.990 rows 132=254.076 99=212.879 115=230.328 123=244.690 127=253.179 125=248.878 chosen sp127-w32 253.179 (99.5% of the ceiling, 127 of 132 SMs) 15165 ms | | tune0.99 | mx8_sh256x27 | base (0 x 1) | 252.67 | 692.7 | 2.742 | 1752 | 2619 | PASS | RESULT sparse tune device NVIDIA_H100_80GB_HBM3 sms 132 installed base 247.311 ceiling 247.311 hold 0.990 rows 132=242.003 99=205.751 chosen none (no sparse shape held; the installed kernel stays) 6557 ms | | tune0.98 | mx8-genesis | sp127-w32 (127 x 32) | 253.35 | 451.1 | 1.781 | 1980 | 2619 | PASS | RESULT sparse tune device NVIDIA_H100_80GB_HBM3 sms 132 installed base 254.909 ceiling 254.909 hold 0.980 rows 132=254.394 99=212.934 115=230.633 123=244.933 127=252.939 125=249.108 chosen sp127-w32 252.939 (99.2% of the ceiling, 127 of 132 SMs) 15153 ms | | tune0.98 | mx8_sh256x27 | base (0 x 1) | 252.17 | 687.3 | 2.726 | 1782 | 2619 | PASS | RESULT sparse tune device NVIDIA_H100_80GB_HBM3 sms 132 installed base 247.177 ceiling 247.177 hold 0.980 rows 132=241.720 99=204.596 chosen none (no sparse shape held; the installed kernel stays) 6572 ms | | v5 | v5-genesis | ( x ) | 250.24 | 698.8 | 2.793 | 1674 | 2619 | PASS | | | v5 | v4-genesis | ( x ) | 250.59 | 699.7 | 2.792 | 1732 | 2619 | PASS | | | v5 | v4-genesis | base (0 x 1) | 250.58 | 699.6 | 2.792 | 1755 | 2619 | PASS | |

RESULT after-micro label=h100 watts=136.6 RESULT microbench probe=sleep status=ok unit="none (the SM-resident floor: full occupancy, __nanosleep, no issue)" ops_per_step=0 lanes=270336 regs=11 blocks_per_sm=8 steps=2000 launches=13740 launch_ms=2.1 seconds=30.0 start_utc=2026-10-08T12:16:17Z end_utc=2026-10-08T12:16:47Z G_steps_s=247.619 G_ops_s=0.000 checksum=8b671be935472947

RESULT after-micro label=h100 watts=437.9 RESULT microbench probe=int_arx status=ok unit="int32 add, xor or rotate (4 independent chains, 3 ops each per step)" ops_per_step=12 lanes=270336 regs=14 blocks_per_sm=8 steps=16384 launches=11899 launch_ms=2.4 seconds=30.0 start_utc=2026-10-08T12:16:47Z end_utc=2026-10-08T12:17:17Z G_steps_s=1756.760 G_ops_s=21081.115 checksum=fcb635c03d37ed2a

RESULT after-micro label=h100 watts=450.3 RESULT microbench probe=dram_chase_1g status=ok unit="dependent random 4-byte read in a 1 GiB table (the hash's own pattern, the control)" ops_per_step=1 lanes=270336 regs=19 blocks_per_sm=8 steps=512 launches=6982 launch_ms=4.2 seconds=30.0 start_utc=2026-10-08T12:17:17Z end_utc=2026-10-08T12:17:47Z G_steps_s=32.213 G_ops_s=32.213 checksum=7623c6b86570e4c

RESULT after-idle label=h100 watts=80.0 sm_mhz=350.6 mem_mhz=2619.0

PC 1, the RTX 5090 alone, the hash lane's job run-ca4-pc1-floorsm-5090-20261008 (the ca4sparse5 exe, unlocked and at the 1,300 MHz lock through the Power Helper)

The floorsm ladder

Pack State Shape MH/s W Microjoules SM MHz Mem MHz Temp max Check
v4-devnet-epoch0 unlocked base (0) 137.164 452.8 3.301 2865 13801 62 PASS
v4-devnet-epoch0 unlocked sp170-w32 (170) 136.407 462.6 3.391 2842 13801 67 PASS
v4-devnet-epoch0 unlocked sp170-w16 (170) 137.12 480.3 3.503 2842 13801 72 PASS
v4-devnet-epoch0 unlocked sp170-w8 (170) 137.035 482.9 3.524 2827.6 13801 74 PASS
v4-devnet-epoch0 unlocked sp170-w4 (170) 134.941 483.9 3.586 2827 13801 75 PASS
v4-devnet-epoch0 unlocked sp170-w2 (170) 98.779 409.6 4.147 2835 13801 73 PASS
v4-devnet-epoch0 unlocked sp128-w32 (128) 136.422 476.2 3.491 2835 13801 74 PASS
v4-devnet-epoch0 unlocked sp96-w32 (96) 134.373 470.6 3.502 2835 13801 75 PASS
v4-devnet-epoch0 unlocked sp85-w32 (85) 136.278 474.8 3.484 2835 13801 76 PASS
v4-devnet-epoch0 unlocked sp64-w32 (64) 136.056 474.7 3.489 2835 13801 75 PASS
v4-devnet-epoch0 unlocked sp52-w32 (52) 128.568 459.9 3.577 2842 13801 75 PASS
v4-devnet-epoch0 unlocked sp43-w32 (43) 134.477 470 3.495 2842 13801 75 PASS
v4-devnet-epoch0 unlocked sp36-w32 (36) 120.347 440.9 3.664 2842 13801 74 PASS
v4-devnet-epoch0 unlocked sp32-w32 (32) 107.265 413.1 3.851 2850 13801 73 PASS
v4-devnet-epoch0 unlocked sp28-w32 (28) 94.364 384.2 4.071 2857 13801 71 PASS
v4-devnet-epoch0 unlocked sp24-w32 (24) 80.823 353.8 4.377 2850 13801 69 PASS
v4-devnet-epoch0 unlocked sp21-w32 (21) 70.735 327.2 4.626 2850 13801 66 PASS
v4-devnet-epoch0 unlocked sp16-w32 (16) 53.985 290.9 5.389 2858.3 13801 64 PASS
v4-devnet-epoch0 unlocked sp11-w32 (11) 37.34 250.9 6.719 2872 13801 60 PASS
v4-devnet-epoch0 unlocked sp43-w8 (43) 89.595 362.8 4.049 2855.9 13801 65 PASS
mx8-devnet-epoch0 unlocked base (0) 136.721 320.9 2.347 2850 13801 63 PASS
mx8-devnet-epoch0 unlocked sp170-w32 (170) 136.166 319.1 2.343 2850 13801 63 PASS
mx8-devnet-epoch0 unlocked sp170-w16 (170) 136.884 316.8 2.314 2850 13801 63 PASS
mx8-devnet-epoch0 unlocked sp170-w8 (170) 137.149 310.8 2.266 2850 13801 63 PASS
mx8-devnet-epoch0 unlocked sp170-w4 (170) 136.904 310.9 2.271 2865 13801 63 PASS
mx8-devnet-epoch0 unlocked sp170-w2 (170) 135.705 309.9 2.284 2865 13801 62 PASS
mx8-devnet-epoch0 unlocked sp128-w32 (128) 136.627 313.4 2.294 2865 13801 62 PASS
mx8-devnet-epoch0 unlocked sp96-w32 (96) 134.635 309.9 2.302 2865 13801 62 PASS
mx8-devnet-epoch0 unlocked sp85-w32 (85) 136.758 311.6 2.278 2865 13801 62 PASS
mx8-devnet-epoch0 unlocked sp64-w32 (64) 136.733 311.3 2.277 2865 13801 62 PASS
mx8-devnet-epoch0 unlocked sp52-w32 (52) 132.782 307.2 2.314 2865 13801 62 PASS
mx8-devnet-epoch0 unlocked sp43-w32 (43) 136.463 310.6 2.276 2865 13801 61 PASS
mx8-devnet-epoch0 unlocked sp36-w32 (36) 136.046 309.6 2.276 2865 13801 61 PASS
mx8-devnet-epoch0 unlocked sp32-w32 (32) 135.686 309.4 2.280 2865 13801 61 PASS
mx8-devnet-epoch0 unlocked sp28-w32 (28) 135.201 308.6 2.283 2865 13801 61 PASS
mx8-devnet-epoch0 unlocked sp24-w32 (24) 133.819 307 2.294 2865 13801 61 PASS
mx8-devnet-epoch0 unlocked sp21-w32 (21) 132.826 305.6 2.301 2865 13801 61 PASS
mx8-devnet-epoch0 unlocked sp16-w32 (16) 125.545 294.4 2.345 2865 13801 61 PASS
mx8-devnet-epoch0 unlocked sp11-w32 (11) 100.111 275.1 2.748 2865.2 13801 60 PASS
mx8-devnet-epoch0 unlocked sp43-w8 (43) 135.989 306.2 2.252 2865 13801 61 PASS
v4-devnet-epoch0 1300 base (0) 134.234 302 2.250 1290 13801 55 PASS
v4-devnet-epoch0 1300 sp170-w32 (170) 128.882 298.1 2.313 1290 13801 59 PASS
v4-devnet-epoch0 1300 sp170-w16 (170) 130.496 302.4 2.317 1290 13801 61 PASS
v4-devnet-epoch0 1300 sp170-w8 (170) 131.723 303.7 2.306 1290 13801 62 PASS
v4-devnet-epoch0 1300 sp170-w4 (170) 101.305 267.2 2.638 1290 13801 60 PASS
v4-devnet-epoch0 1300 sp170-w2 (170) 54.393 197.9 3.638 1290 13801 57 PASS
v4-devnet-epoch0 1300 sp128-w32 (128) 130.252 299 2.296 1290 13801 59 PASS
v4-devnet-epoch0 1300 sp96-w32 (96) 120.933 286.1 2.366 1290 13801 60 PASS
v4-devnet-epoch0 1300 sp85-w32 (85) 125.221 292.8 2.338 1290 13801 61 PASS
v4-devnet-epoch0 1300 sp64-w32 (64) 95.167 249.9 2.626 1290 13801 59 PASS
v4-devnet-epoch0 1300 sp52-w32 (52) 77.801 228.9 2.942 1290 13801 57 PASS
v4-devnet-epoch0 1300 sp43-w32 (43) 64.343 208.5 3.240 1290 13801 55 PASS
v4-devnet-epoch0 1300 sp36-w32 (36) 53.964 185.3 3.434 1290 13801 53 PASS
v4-devnet-epoch0 1300 sp32-w32 (32) 47.978 174.7 3.641 1290 13801 51 PASS
v4-devnet-epoch0 1300 sp28-w32 (28) 41.985 170.9 4.071 1290 13801 50 PASS
v4-devnet-epoch0 1300 sp24-w32 (24) 35.981 158.5 4.405 1290 13801 49 PASS
v4-devnet-epoch0 1300 sp21-w32 (21) 31.455 155 4.928 1290 13801 48 PASS
v4-devnet-epoch0 1300 sp16-w32 (16) 24.07 137.3 5.704 1290 13801 46 PASS
v4-devnet-epoch0 1300 sp11-w32 (11) 16.551 115.1 6.954 1290 13801 44 PASS
v4-devnet-epoch0 1300 sp43-w8 (43) 45.592 173.4 3.803 1290 13801 46 PASS
mx8-devnet-epoch0 1300 base (0) 134.076 213 1.589 1290 13801 48 PASS
mx8-devnet-epoch0 1300 sp170-w32 (170) 129.08 208.9 1.618 1290 13801 49 PASS
mx8-devnet-epoch0 1300 sp170-w16 (170) 130.441 209.7 1.608 1290 13801 51 PASS
mx8-devnet-epoch0 1300 sp170-w8 (170) 130.788 209.7 1.603 1290 13801 51 PASS
mx8-devnet-epoch0 1300 sp170-w4 (170) 130.789 210.1 1.606 1290 13801 52 PASS
mx8-devnet-epoch0 1300 sp170-w2 (170) 130.278 210.4 1.615 1290 13801 52 PASS
mx8-devnet-epoch0 1300 sp128-w32 (128) 132.634 213.6 1.610 1290 13801 53 PASS
mx8-devnet-epoch0 1300 sp96-w32 (96) 120.795 205.2 1.699 1290 13801 52 PASS
mx8-devnet-epoch0 1300 sp85-w32 (85) 110.96 191.9 1.729 1290 13801 52 PASS
mx8-devnet-epoch0 1300 sp64-w32 (64) 103.449 187.2 1.810 1290 13801 51 PASS
mx8-devnet-epoch0 1300 sp52-w32 (52) 85.585 174.3 2.037 1290 13801 50 PASS
mx8-devnet-epoch0 1300 sp43-w32 error: budget
mx8-devnet-epoch0 1300 sp36-w32 error: budget
mx8-devnet-epoch0 1300 sp32-w32 error: budget
mx8-devnet-epoch0 1300 sp28-w32 error: budget
mx8-devnet-epoch0 1300 sp24-w32 error: budget
mx8-devnet-epoch0 1300 sp21-w32 error: budget
mx8-devnet-epoch0 1300 sp16-w32 error: budget
mx8-devnet-epoch0 1300 sp11-w32 error: budget
mx8-devnet-epoch0 1300 sp43-w8 error: budget
v4-devnet-epoch0 unlocked-end base error: budget
v4-devnet-epoch0 unlocked-end sp170-w32 error: budget
v4-devnet-epoch0 unlocked-end sp170-w16 error: budget
v4-devnet-epoch0 unlocked-end sp170-w8 error: budget
v4-devnet-epoch0 unlocked-end sp170-w4 error: budget
v4-devnet-epoch0 unlocked-end sp170-w2 error: budget
v4-devnet-epoch0 unlocked-end sp128-w32 error: budget
v4-devnet-epoch0 unlocked-end sp96-w32 error: budget
v4-devnet-epoch0 unlocked-end sp85-w32 error: budget
v4-devnet-epoch0 unlocked-end sp64-w32 error: budget
v4-devnet-epoch0 unlocked-end sp52-w32 error: budget
v4-devnet-epoch0 unlocked-end sp43-w32 error: budget
v4-devnet-epoch0 unlocked-end sp36-w32 error: budget
v4-devnet-epoch0 unlocked-end sp32-w32 error: budget
v4-devnet-epoch0 unlocked-end sp28-w32 error: budget
v4-devnet-epoch0 unlocked-end sp24-w32 error: budget
v4-devnet-epoch0 unlocked-end sp21-w32 error: budget
v4-devnet-epoch0 unlocked-end sp16-w32 error: budget
v4-devnet-epoch0 unlocked-end sp11-w32 error: budget
v4-devnet-epoch0 unlocked-end sp43-w8 error: budget
mx8-devnet-epoch0 unlocked-end base error: budget
mx8-devnet-epoch0 unlocked-end sp170-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp170-w16 error: budget
mx8-devnet-epoch0 unlocked-end sp170-w8 error: budget
mx8-devnet-epoch0 unlocked-end sp170-w4 error: budget
mx8-devnet-epoch0 unlocked-end sp170-w2 error: budget
mx8-devnet-epoch0 unlocked-end sp128-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp96-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp85-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp64-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp52-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp43-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp36-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp32-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp28-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp24-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp21-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp16-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp11-w32 error: budget
mx8-devnet-epoch0 unlocked-end sp43-w8 error: budget

PC 1, the memory-clock ladder at the 1,300 MHz lock, job run-ca4-pc1-memclk-5090-20261008

The memory-clock ladder

Pack State Shape MH/s W Microjoules SM MHz Mem MHz Temp max Check
v4-devnet-epoch0 unlocked base (grid) 136.552 464.3 3.400 2852.1 13801 70 PASS
mx8-devnet-epoch0 unlocked base (grid) 136.29 328.1 2.407 2850 13801 65 PASS
v4-devnet-epoch0 1300 base (grid) 134.34 309 2.300 1290 13801 60 PASS
mx8-devnet-epoch0 1300 base (grid) 134.089 222.9 1.662 1290 13801 57 PASS
v4-devnet-epoch0 1300m14001 base (grid) 134.513 308.1 2.290 1290 13801 59 PASS
mx8-devnet-epoch0 1300m14001 base (grid) 134.342 222.7 1.658 1290 13801 56 PASS
v4-devnet-epoch0 1300m12001 base (grid) 134.528 308.6 2.294 1290 13801 59 PASS
mx8-devnet-epoch0 1300m12001 base (grid) 134.252 223.1 1.662 1290 13801 56 PASS
v4-devnet-epoch0 1300m10001 base (grid) 134.487 308.9 2.297 1290 13801 59 PASS
mx8-devnet-epoch0 1300m10001 base (grid) 134.287 223.7 1.666 1290 13801 55 PASS
v4-devnet-epoch0 1300m8001 base (grid) 134.541 308.7 2.294 1290 13801 60 PASS
mx8-devnet-epoch0 1300m8001 base (grid) 134.33 223 1.660 1290 13801 56 PASS
v4-devnet-epoch0 1300m6001 base (grid) 76.09 187 2.458 1290 7001 53 PASS
mx8-devnet-epoch0 1300m6001 base (grid) 76.009 142.2 1.871 1290 7001 51 PASS
v4-devnet-epoch0 1300m5001 base (grid) 76.091 184.4 2.423 1290 7001 50 PASS
mx8-devnet-epoch0 1300m5001 base (grid) 76.007 141.1 1.856 1290 7001 48 PASS
v4-devnet-epoch0 1300m3001 base (grid) 76.082 183.8 2.416 1290 7001 50 PASS
mx8-devnet-epoch0 1300m3001 base (grid) 75.96 140.9 1.855 1290 7001 47 PASS
v4-devnet-epoch0 unlocked-end base (grid) 136.562 452.5 3.314 2866.6 13801 61 PASS
mx8-devnet-epoch0 unlocked-end base (grid) 136.298 318.6 2.338 2865 13801 60 PASS

The decomposition probes (the worker's --microbench: sleep = full residency and no issue; the L2 chase; the dependent DRAM chase; the ALU probe; 30 s each with the sampler's watts and both clocks)

Card Probe Watts SM MHz Mem MHz G ops or reads per s Lanes nJ or pJ per op over the sleep row
4090 idle 33.9 210.0 405.0
4090 sleep 133.1 2725.9 10501.0 0.000 196608
4090 int_arx 448.0 2574.1 10501.0 26076.595 196608 12.1 pJ
4090 l2_chase_32m 329.0 2715.0 10501.0 80.816 196608 2.42 nJ
4090 l2_indep4_32m 328.2 2715.0 10501.0 82.301 196608 2370.6 pJ
4090 dram_chase_1g 263.0 2715.0 10501.0 9.399 196608 13.82 nJ
5090c idle 184.2 2044.5 11141.9
5090c sleep 196.5 2865.0 13801.0 0.000 261120
5090c int_arx 372.9 2853.9 13801.0 13064.448 261120 13.5 pJ
5090c l2_chase_32m 382.2 2865.0 13801.0 53.964 261120 3.44 nJ
5090c l2_indep4_32m 401.4 2851.3 13801.0 51.085 261120 4011.0 pJ
5090c dram_chase_1g 372.9 2865.0 13801.0 8.019 261120 22.00 nJ
5090e idle 8.7 180.0 405.0
5090e sleep 139.0 2932.0 13801.0 0.000 261120
5090e int_arx 535.6 2854.9 13801.0 29035.421 261120 13.7 pJ
5090e l2_chase_32m 415.1 2902.0 13801.0 104.533 261120 2.64 nJ
5090e l2_indep4_32m 434.9 2902.0 13801.0 112.864 261120 2621.7 pJ
5090e dram_chase_1g 341.3 2904.0 13801.0 17.617 261120 11.48 nJ
h100 idle 71.8 345.0 2619.0
h100 sleep 130.3 1980.0 2619.0 0.000 270336
h100 int_arx 440.0 1980.0 2619.0 21499.215 270336 14.4 pJ
h100 l2_chase_32m 479.1 1980.0 2619.0 121.037 270336 2.88 nJ
h100 l2_indep4_32m 538.6 1980.0 2619.0 107.277 270336 3806.0 pJ
h100 dram_chase_1g 466.5 1980.0 2619.0 32.551 270336 10.33 nJ

2. What the rows say, shape by shape

Question 5090 (host c; PC 1 for the lock) 4090 H100 Label
Fewest SMs that hold 99 percent of class v3 43 of 170 (99.4); 28 holds 98.4 16 of 128 (99.0); the tune's 26 at 99.3 127 of 132 (the rate falls with the SM count below it) measured
Watts at that point against base 291.1 against 296.4 (-5 W); host e 337.4 against 352.7 (-15 W) 249.6 against 253.4 (-4 W) 441.6 against 450.4 (-9 W) measured
Energy per hash at that point 2.053 against 2.079 (-1.3 percent); host e -3.9 percent 3.568 against 3.588 (-0.6); the tune's 3.567 1.744 against 1.769 (-1.4) measured
Fewer warps per SM at full SM count 8 warps per SM 2.063 (-0.8 percent); 2 warps 97.7 percent of rate at 2.072; 1 warp 84 percent 8 warps 3.570; 1 warp 94 percent at 3.662 8 warps 99.2 percent at the same draw; 4 warps 93 percent; 1 warp 41 percent measured
Matched warps, packed against spread (43 x 32 against 170 x 8) 2.053 against 2.063: equal 32 x 8 (3.546) against the ladder's 256 warps (3.568): equal 33 x 8 against 33 x 32: 123.5 against 123.4 MH/s, equal measured
Class v4 knee by rate 43 SMs (PC 1: 98 percent; 36 SMs 88; 28 SMs 69) the full grid (all shapes within 0.1 percent; the ladder below 128 not run on class v4) the shipped grid; the persistent shape loses 4 percent measured
Class v4 watts at its knee PC 1 470 W against 452.8 base (heat drift 62 to 75 C across the pass); host c capped at 450 356 against 353.5 capped at 700 measured
At the 1,300 MHz lock (PC 1, this file's job) class v4: the full grid 134.2 MH/s at 302 W = 2.25 microjoules; the persistent shape on every SM 128.9 at 298.1 (the wrapper costs 4 percent); 128 SMs 130.3 at 299; 96 SMs 120.9 at 286; 85 SMs 125.2 at 292.8; 64 SMs 95.2 at 249.9; 43 SMs 64.3 at 208.5 (3.24 microjoules): nothing under the full grid's energy at any point. Class v3: the full grid 134.1 at 213 W = 1.589; 128 SMs 132.6 at 213.6; 96 SMs 120.8 at 205.2; 85 SMs 111.0 at 191.9; 64 SMs 103.4 at 187.2; 52 SMs 85.6 at 174.3; the rows under 52 SMs were cut by the job's 66-minute budget no lock on a rented host no lock on a rented host measured; the class v3 lock ladder is PARTIAL and disagrees with 20.3b (85 SMs 129.8 MH/s at 207.8 W there against 111.0 at 191.9 here, 43 SMs 126.3 there, not reached here): the pass ran across the founder's card swap on PC 1 (14:2x UK) and is re-run before it replaces 20.3b's lock points
Class v3 at stock on PC 1 (this file's job, the same board as 20.3b) base 136.7 at 320.9 W = 2.347; 43 SMs 136.5 at 310.6 = 2.276 (3.0 percent under); 28 SMs 135.2 at 308.6; 21 SMs 132.8 at 305.6; 16 SMs 125.5 at 294.4; 11 SMs 100.1 at 275.1; 170 SMs x 8 warps 137.1 at 310.8 = 2.266 (3.5 percent under, the best shape on PC 1) measured; agrees with the rented ladder within 1 percent at every rung
Class v5 against class v4 not run on the rented 5090s (the v5-kits worker has no sparse shapes); PC 1's v5lock-b of 8 October: v5 at the knee = v4 rate, +2 percent watts not run 250.2 MH/s at 698.8 W against 251.9 at 699.4: equal under the cap measured
The memory-clock knob on a rented host refused (-lmc, -lgc, -pl all refused on Vast) refused (RunPod) -lmc accepted and the clock did not move (2,619 MHz on the row at both settings) measured

Two oddities the ladder shows and this file keeps: on host c the 96-SM and 52-SM shapes read 97.4 and 95.4 percent of the rate where 85 and 43 read 99.9 and 99.4 (and host e's 52 reads 95.7, host d's 96 reads 97.5), so the dip is the shape, not the host; the 5090's 170 SMs sit in 8 GPCs, and a block count that lands unevenly across them (96 and 52 against 85 and 43) leaves some GPCs' L2 ports hotter than others. The persistent wrapper's own cost is visible on the H100 only (sp132-w32 at 96 percent of class v4's base against 99.7 on class v3: a compute-bound kernel pays the loop).

3. The decomposition: what the residual is

The microbench's probes run one kernel each at full residency for 30 s with the sampler beside it: sleep (every warp resident, no instruction issued: the SM clock domain awake and nothing else), l2_chase_32m (a dependent 4-byte chain inside the L2: the L2 and crossbar path, no DRAM), dram_chase_1g (the hash's own pattern over 1 GiB: the DRAM path through the same L2 and crossbar), int_arx (the ALU path). The table in section 1 carries the rows; the reading per card, watts over idle:

Card Idle SM clock domain awake (sleep minus idle) The L2 path per dependent read (l2 chase minus sleep, over reads per s) The DRAM path per dependent read (dram chase minus sleep, over reads per s) The hash's draw over idle at its read rate Of which the memory path (hash minus sleep) The memory's own, modelled (chip-model 5.3) Label
RTX 5090, PC 1 (15.1a, 8 October, the microbench at stock) 74.7 W 45 W (120.2 at full residency) 2.4 nJ (378.7 W at 107.6 G reads per s) 10.9 nJ whole-card (318.4 W at 18.2 G); 8.5 nJ over the L2 path 311 - 75 = 236 W at 17.6 G reads per s = 13.4 nJ per read (20.3's mx8 row) 191 W = 10.9 nJ per read 2.0 nJ per read plus 20 W static = 55 W, 3.1 nJ per read measured; the memory's own modelled
RTX 5090, rented host e (this file's probes) 8.7 W at 180 MHz 130 W (139.0 W at 2,932 MHz) 2.6 nJ (415.1 W at 104.5 G reads per s) 11.5 nJ over the sleep row (341.3 W at 17.6 G); 18.9 nJ whole-card 352.7 - 8.7 = 344 W at 18.2 G reads per s = 18.9 nJ per read 214 W = 11.8 nJ per read the same 55 W measured; modelled
RTX 4090, rented 33.9 W at 210 MHz 99 W (133.1 W at 2,726 MHz with every warp resident and nothing issued) 2.4 nJ (329.0 W at 80.8 G reads per s) 13.8 nJ over the sleep row (263.0 W at 9.4 G); 24.4 nJ whole-card 253.4 - 33.9 = 219 W at 9.04 G reads per s = 24.2 nJ per read 120 W = 13.3 nJ per read 2.0 nJ plus 15 W static = 33 W, 3.7 nJ per read (approximate for GDDR6X) measured; modelled
H100 80 GB HBM3, rented 71.8 to 80 W at 345 MHz 50 to 58 W (130.3 to 136.6 W at 1,980 MHz) 2.9 nJ (479.1 W at 121.0 G reads per s) 10.3 nJ over the sleep row (466.5 W at 32.6 G); 12.1 nJ whole-card 450.4 - 71.3 = 379 W at 32.6 G reads per s = 11.6 nJ per read 320 W = 9.8 nJ per read 1.2 nJ plus about 20 W static = 59 W, 1.8 nJ per read measured; modelled

The reading: at the hash, each card's draw splits into a fixed "awake" floor (the card with every warp resident and nothing issuing: 120 to 139 W on the 5090 whatever its idle, 133 W on the 4090, 130 to 137 W on the H100; the SM clock tree, the GPCs' shared logic and the memory clock domain at their boost clocks; the SM-sparse rows say the SMs' own share of it is 4 to 15 W, because emptying three quarters of them saves that much) and a per-read term on the memory path (10.9 to 11.8 nJ per dependent read on the 5090 on two boards, 13.3 on the 4090, 9.8 on the H100) of which the memory devices themselves are 2.0 or 1.2 nJ modelled; the difference (8 to 11 nJ per read) is the L2 slices, the crossbar, the memory controllers and the PHY, in that order of distance from the SM, and the L2 path alone is 2.4 to 2.9 nJ of it (the L2 chase). The ALU path is not in the hash's draw at all (512 ops per hash at 11 to 14 pJ is under 1 W). The two 5090 boards agree on the per-read term within 8 percent and differ by 60 W on the idle (7 to 9 W on the rented Linux boards against 75 W on PC 1 with its display), which is where the across-host base spread comes from.

What the SM-sparse rows add to it: the SMs that go idle at the knee were worth 4 to 15 W on every card (the difference between the sleeping SM clock domain at full residency and the same domain with three quarters of its SMs empty), so the "SM clock domain awake" term is mostly the clock tree and the GPCs' shared logic, not the SMs' own schedulers and register files. The rest of the draw over idle at the hash is the memory path on the GPU side: the L2 slices, the crossbar, the memory controllers and the GDDR7 or HBM PHY, at the boost clock for the first three and at the memory clock for the PHY. The core lock takes the first three down with the clock (the 5090's 311 to 213 W at the same 17.5 G reads per second on PC 1, 20.3); the memory-clock lock is the only knob on the fourth, and no rented host allows it, so its row is PC 1's (the memory-clock ladder table in section 1, the job run-ca4-pc1-memclk-5090-20261008).

3.1 The memory-clock ladder (PC 1, the 5090 alone, job run-ca4-pc1-memclk-5090-20261008-b, 15:0x to 15:4x UK, the helper's lmc verb at the 1,300 MHz core lock; the amendment of 15:50 UK)

The driver holds the 5090's memory clock at two points only: every ask at or above 8,001 MHz reads 13,801 on the row and every ask at or below 6,001 reads 7,001 (the GDDR7 PHY's half-rate state). The rows, both bit-exact against the Mac's fingerprints:

Pack Core lock, memory ask Memory MHz on the row MH/s Watts Microjoules per hash nJ per dependent read, whole card Label
class v3 (mx8) 1,300, no memory lock 13,801 134.09 222.9 1.662 13.0 measured
class v3 1,300, lmc 14001 / 12001 / 10001 / 8001 13,801 134.3 to 134.5 222.7 to 223.7 1.658 to 1.666 13.0 measured (no change)
class v3 1,300, lmc 6001 / 5001 / 3001 7,001 76.0 140.9 to 142.2 1.855 to 1.871 14.5 measured
class v4 (sh256x27) 1,300, no memory lock 13,801 134.34 309.0 2.300 measured
class v4 1,300, lmc 8001 and above 13,801 134.5 308.1 to 308.9 2.290 to 2.297 measured
class v4 1,300, lmc 6001 and below 7,001 76.1 183.8 to 187.0 2.416 to 2.458 measured
both, unlocked (the pass's own base) none 13,801 136.3 / 136.6 318.6 to 328.1 / 452.5 to 464.3 2.34 to 2.41 / 3.31 to 3.40 measured (the heat drift across the pass)

The reading: the memory clock is not a lever either. At the half-rate PHY state the rate falls 43 percent (134 to 76 MH/s: the hash is bound by the memory's activate ceiling, which scales with the memory clock) and the watts fall 37 percent (82 W on class v3, 125 W on class v4), so the energy per hash RISES 12 percent on class v3 (1.66 to 1.86) and 5 percent on class v4 (2.30 to 2.42). Read as a decomposition, the 82 W between the two memory states on class v3 is the memory path's clock-scaled share at 7.4 G reads per second of rate lost: 11 nJ per read at the margin, the same figure the DRAM chase probes give for the whole memory path on the GPU side (10.9 to 11.8 nJ). So the memory-side term of the 5090's draw is about 80 to 90 W of its 223 W at the lock (the PHY, controllers and L2 at the memory clock plus the devices' own 55 W modelled), and it is bought one for one with the rate: the honest floor at the 1,300 lock stays 1.66 microjoules on this board (1.58 on 20.3b's cooler pass), and no clock or occupancy knob on the card takes it lower without taking the rate with it. The class v4 premium at the half-rate state is 43 W for 76 MH/s x 102,100 ops = 5.5 pJ per counted op (6.3 at the full memory clock on the same pass): the shadow's marginal is the core domain's, not the memory's, as expected.

4. The honest floor per tier, and the fraction a chip cannot strip

The chip of the record pays the memory's own energy per read, its static power and a controller and PHY of its own (chip-model-v3 5.3 to 5.5: GDDR7 2.0 nJ per read plus 20 W static plus a 15 W controller at 166 MH/s = 0.466 microjoules per hash; one HBM3 stack 0.321). The honest card pays the same memory's energy (the devices are the same silicon on both sides) plus everything its die spends around the read. So the fraction of the card's energy per hash a chip cannot strip is the memory side's own share of it, and the floor's real number per tier is that share:

Tier (class v3, stock unless named) Card microjoules per hash The memory side's own (modelled: the DRAM's 2.0 or 1.2 nJ x 128 reads, plus static over the rate) Fraction a chip cannot strip Edge at zero shadow: GDDR7 board / HBM3 stack / SRAM die Label
RTX 5090, rented host c (the best board of four) 2.079 0.256 + 0.140 = 0.40 19 percent 4.5x / 6.5x / 14.9x measured card; modelled memory
RTX 5090, rented host e (the worst of four) 2.485 0.40 16 percent 5.3x / 7.7x / 17.8x the same
RTX 5090, PC 1 at the 1,300 MHz lock (20.3b, class v3 base) 1.58 0.40 25 percent 3.4x / 4.9x / 11.3x measured
RTX 4090, rented (GDDR6X at 9.0 G reads per s) 3.588 0.256 + 0.21 = 0.47 13 percent 7.7x / 11.2x / 25.6x measured card; modelled memory
H100 80 GB, rented (HBM3, 5 stacks) 1.769 0.154 + 0.08 = 0.23 13 percent 3.8x / 5.5x / 12.6x the same
Apple M5 Max (6 October, GPU plus DRAM channels) 0.78 about 0.3 (approximate: LPDDR at 1.5 to 2.0 nJ) about 40 percent 1.7x / 2.4x / 5.6x measured card; approximate memory

The edge columns above are against the model's three chips; the 5090 host c row is the honest 5090 floor at stock this lane measured (4.5x against the GDDR7 board, not the 5.2x of the 7 October PC 2 row: the rented board draws 296 W at 142.6 MH/s where PC 2's drew 330 at 136.6, and PC 1's 311 at 137.5; the three are one card model on three boards and drivers, and the honest denominator is a band, as the denominator lane's file says). The lock row (PC 1, 20.3b) stays the honest 5090 floor: 3.4x on class v3 at zero shadow, 3.6x at the knee on the efficiency pass's 1.67; with class v4's shadow and the record's k band it is section 20.4's 2.1x at k = 1.

Per tier, what it means and what is done: a home miner on one 5090, 4090 or any NVIDIA card gains 1 to 4 percent of energy per hash from the occupancy at stock and never more, so nothing in the occupancy changes what the card earns; the core lock stays the lever (the 5090: 1.67 against 2.26 microjoules, the Ember knob of 0.3.24), and the memory-clock lock behind it is measured on PC 1 in this file; the self-tune ships as the safe default inside Ember's Efficiency tier and costs nothing at Maximum. An H100 owner (a rented or data-centre tier) gets the 1.4 percent from the self-tune and nothing from its memory clock (the knob is accepted and ignored on the rented host). A 4090 owner gets 1.3 percent. An AMD or Apple owner gets nothing from this lane (the shapes are CUDA's; Metal and OpenCL workers run the shipped grid).

5. The patch: --sm-sparse auto and the self-tune

proto-cuda/nvrtc/worker.cpp on this branch:

  • --sm-sparse auto|off|N and --sm-hold <f> (default 0.99); the tuning file's per-card "sm_sparse": "auto" | "off" | "<N>" and "sm_hold": 0.99 (the manifest's tuning object is written as is to tuning.json, so a fleet default needs no app change; the flag wins over the file). A count pins sp<N>-w32 through the existing pinned race (the vector warps gate the install); auto runs tuneSparse at the end of buildPair, on the first pack and on every prepared pack, so a pack flip re-measures.
  • tuneSparse: the persistent shape is one compiled kernel for every block count (N is the launch's grid, not the text), so one NVRTC compile serves the whole search; the installed kernel's rate is the ceiling's first reading, the sparse shape on every SM its second (the vector warps through it first: bit-exact or nothing is installed); the ladder (sparseLadder: the SM count by quarters and eighths down to 2) descends until a shape falls under the hold and a bisection of at most four probes closes the bracket; chooseSparseBlocks installs the fewest blocks that held. One line goes out with the race line (sparse tune device ... rows N=mhs ... chosen sp<N>-w32 (x percent of the ceiling)) and the bench's RESULT line carries variant= and sparse_blocks=.
  • Measured: H100 15.2 s, 6 probes, chosen sp127-w32 253.179 (99.5% of the ceiling), the served row 253.2 MH/s at 441.6 W (1.744 against 1.769 microjoules); class v4 on the H100 chosen none in 6.6 s (the persistent shape at 97.8 percent of the ceiling, under the hold); RTX 4090 24.5 s, chosen sp26-w32 (99.3%), 70.15 MH/s at 250.2 W (3.567 against 3.588); RTX 5090 host c, class v3 at the 0.99 hold chosen sp37-w32, 141.27 MH/s at 286.5 W = 2.028 against the base's 2.079 (2.5 percent), at 0.98 sp27-w32 2.033; class v4 at the 0.99 hold sp82-w32 3.178 against 3.157 (nothing, inside noise under the board's 450 W cap), at 0.98 sp71-w32 3.185. The 0.99 hold is the default: on class v4 it costs 1 percent of rate for no watts, which is why Maximum turns it off and Balanced holds 0.995.
  • Tests: proto-cuda/nvrtc/emu/variant-test.cpp case sparse-tune (the ladder from 170 SMs, the chooser on the 20.3b rows at holds 0.99, 0.98, 0.5 and 1.0, the mode reader on nine strings, the tuning keys for two cards): PASS beside the three earlier cases, built with the real headers on the H100 pod (g++ -DIGNEUM_EMU), 8 October 12:04 UTC. The known-failed shape it guards: a hold of 1.0 or an empty row set must choose nothing (the 7 October fault was a shape installed that was never measured).
  • The tiers (the Ember tiers lane's ember-tiers-25 carries the tier switch; this file names the mapping): Efficiency writes sm_sparse: auto, sm_hold: 0.985; Balanced auto, 0.995; Maximum off. The tune runs inside the pair's build (the prepare thread, ahead of the epoch), so a flip of tier or pack re-measures at the next prepare and the serve loop never stalls; its cost is 7 to 25 s of the prepare lead (600 DAA).
  • What it is not: a lever on the floor. It is a 1 to 4 percent per-watt line on NVIDIA at stock, worth shipping because it is free and measured, and the frame (one compile, a ladder, a hold, bit-exact or nothing) is the shape any later occupancy or clock self-tune takes.

6. Unverified and owed

  • PC 1's 1,300 MHz lock ladder is in section 1 as the job left it: class v4 complete (every point under the full grid's energy, the shadow compute-bound under 128 SMs at 1,290 MHz), class v3 cut at 52 SMs by the job's budget and reading the sparse shapes lower than 20.3b did on the same board at the same lock (85 SMs 111.0 against 129.8 MH/s); the pass ran across the founder's card swap on PC 1 (14:2x UK), so the class v3 lock row stays 20.3b's (43 SMs 126.3 MH/s at 205.4 W against the full grid's 134.0 at 211.4: 1.63 against 1.58 microjoules, nothing saved) and the ladder is re-run on a quiet PC 1 before it replaces it. The memory-clock ladder landed at 15:4x UK (section 3.1).
  • The memory side's own share per tier is modelled (chip-model 5.3); the memory-clock ladder is the one measurement of its PHY term, and only on PC 1.
  • Class v5 on the rented 5090s was not run (the v5-kits worker carries the leaf upload and not the sparse shapes; the two trees merge in 0.3.26); the H100 row and PC 1's v5lock-b stand in.
  • The 4090's class v4 ladder below 128 SMs was not run (the slice pods ran class v3 and the full-grid class v4 rows).
  • Every chip figure is the model's; no chip has been measured.
  • Spend: six pods, USD 25 of the lane's 300 (h100 3.49 per hour, 4090 0.74, four 5090s 0.44 to 0.56; two 5090 hosts destroyed unused: one shared with another tenant's 5.7 GB process at 485 W idle, one with a dead CUDA init; the Korean host vanished mid-run).