Bench log: the class v4 efficiency passes on the RTX 5090 and the RTX 5080 (7 to 8 October 2026): every row, the knee per card, the best MH per watt points, the premium at the lock, the lever and its limits

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 01:58:57 +00:00
parent 584bb6760d
commit 751dd3a268

View file

@ -2698,3 +2698,63 @@ p99 / max 4.610 / 4.948 / 5.606 / 6.194 ms per warp); the worst 200 re-timed at
1,000 on the half-core at 2 reps, the worst 10 at 20: worst half-core cold max 8.708 ms (`attack-f6/87142`), then 8.629
(`attack-f6/88521`), 8.414 (`attack-f6/15781`); the genesis program 8.624. Gate 10 ms: PASS by 1.29 ms. Record
`docs/analysis/attack-pass/f6-verifier.md`; logs `/srv/builds/igneum-wt-attack/target-attack-f6/phase2b.log`, `phase2c.log`.
## 7 to 8 October 2026, the class v4 efficiency passes: the core clock lock on the RTX 5090 and the RTX 5080 (branch ca3-v4-amend, the hash lane)
Main's question of 7 October ("profitable mining is all about efficiency, is there a solution?"): the latency-shadow work hides in the memory wait, so the core clock, and the voltage along the driver's V/F curve, can drop until the work just fits the wait, with no rate loss and the class v4 premium coming back. Measured on PC 1 (machine ae432dc7, driver 617.14, app 0.3.20) with `tools/ca3-v4-amend/pc1-v4-efficiency.ps1`: the installed CUDA worker's `--bench` on the class v4 pack `v4-devnet-epoch0` (id a785001687d8688a) and the class v3 control `mx8-devnet-epoch0` of the same seed (the two differ by the shadow block alone), the card alone through the runner's card switch, 60 s per step (the batch count sized from a 5-batch probe), the clock lock through the installed app's Power Helper task (`nvidia-smi -lgc 0,<MHz>`, a cap; no prompt), nvidia-smi at 1 Hz with the mean from 8 s in (power.draw, power.draw.instant and power.draw.average agree within 0.2 W on every row on 617.14), the 2^24 fingerprint at base 0 checked against the Mac on every step (every step matched), the clocks reset and read back at the end. Jobs: run-ca3-pc1-v4-eff-5090-20261007-b (18:40 to 19:16 UTC, 2,850 to 1,400 MHz), run-ca3-pc1-v4-eff-5090-floor-20261007 (20:21 to 20:41 UTC, 1,400 to 1,100; the steps below were lost to the helper's sequence rule, see the fud ledger), run-ca3-pc1-v4-eff-5080-20261007-d (8 October, 00:45 to 01:54 UTC, unlocked to 1,000 MHz plus 900 on v4; the 75 min budget ended it there). The memory clock stayed at the driver's default throughout (13,801 MHz on the 5090, 14,801 on the 5080).
RTX 5090 (lock MHz: class v4 MH/s / W / MH/W ; class v3 MH/s / W / MH/W; the SM clock read):
| lock | v4 MH/s | v4 W | v4 MH/W | v3 MH/s | v3 W | v3 MH/W | sm MHz |
|---|---|---|---|---|---|---|---|
| unlocked | 136.84 | 475.5 | 0.288 | 136.59 | 330.2 | 0.414 | 2,838 / 2,850 |
| 2850 | 136.89 | 478.9 | 0.286 | 136.61 | 330.3 | 0.414 | 2,833 / 2,842 |
| 2781 | 136.87 | 457.6 | 0.299 | 136.62 | 317.9 | 0.430 | 2,767 |
| 2700 | 136.85 | 443.2 | 0.309 | 136.54 | 312.4 | 0.437 | 2,692 |
| 2550 | 136.77 | 402.9 | 0.339 | 136.43 | 288.5 | 0.473 | 2,542 |
| 2472 | 136.62 | 391.2 | 0.349 | 136.38 | 276.0 | 0.494 | 2,460 |
| 2400 | 136.67 | 382.8 | 0.357 | 136.29 | 269.8 | 0.505 | 2,392 |
| 2250 | 136.43 | 366.5 | 0.372 | 136.07 | 255.0 | 0.534 | 2,242 |
| 2163 | 136.33 | 361.3 | 0.377 | 136.11 | 252.4 | 0.539 | 2,152 |
| 2100 | 136.25 | 359.3 | 0.379 | 136.04 | 251.4 | 0.541 | 2,092 |
| 1950 | 135.99 | 350.1 | 0.388 | 135.78 | 244.6 | 0.555 | 1,942 |
| 1854 | 135.85 | 341.4 | 0.398 | 135.56 | 239.6 | 0.566 | 1,845 |
| 1800 | 135.75 | 337.1 | 0.403 | 135.48 | 237.4 | 0.571 | 1,792 |
| 1650 | 135.37 | 328.1 | 0.413 | 135.10 | 235.1 | 0.575 | 1,642 |
| 1500 | 135.22 | 320.5 | 0.422 | 134.90 | 232.1 | 0.581 | 1,492 |
| 1400 | 134.98 | 316.3 | 0.427 | 134.68 | 228.0 | 0.591 | 1,387 |
| 1300 | 134.76 | 312.5 | 0.431 | 134.62 | 223.3 | 0.603 | 1,290 |
| 1200 | 133.80 | 305.1 | 0.439 | 129.54 | 215.7 | 0.601 | 1,192 |
| 1100 | 122.43 | 287.3 | 0.426 | 118.70 | 209.4 | 0.567 | 1,087 |
Reading (5090): the rate is memory-bound on the whole grid and holds within 1.5 percent of unlocked down to 1,300 MHz on both classes; it falls past 5 percent at 1,200 on class v3 and past 10 percent at 1,100 on class v4, so the knee is 1,300 MHz (42 percent of the 3,090 MHz maximum). Best MH per watt: class v4 at 1,200 MHz (133.80 MH/s at 305.1 W, 0.439 MH/W; 168.6 W recovered for 2.2 percent rate), class v3 at 1,300 MHz (134.62 at 223.3 W, 0.603; 106.6 W for 1.4 percent). The class v4 premium over class v3 is 145.3 W at the unlocked clock and 81.8 W at the best points: 63 W of the premium is the clock and comes back with the lock, 82 W is the shadow's ALU work and stays. The throttle reason reads the power governor (0x400) from unlocked to 2,163 and the clock lock itself (0x4) below. Consequence per tier: a 5090 owner on class v4 locked at 1,200 to 1,300 MHz draws 305 to 313 W instead of 474 for 1.5 to 2.2 percent less rate (MH per watt up 49 to 52 percent).
RTX 5080 (the same shape; 8 October 2026, 00:45 to 01:54 UTC):
| lock | v4 MH/s | v4 W | v4 MH/W | v3 MH/s | v3 W | v3 MH/W | sm MHz |
|---|---|---|---|---|---|---|---|
| unlocked | 71.41 | 253.1 | 0.282 | 71.28 | 169.7 | 0.420 | 2,963 / 2,977 |
| 2850 | 71.41 | 229.5 | 0.311 | 71.29 | 155.3 | 0.459 | 2,842 |
| 2781 | 71.41 | 223.9 | 0.319 | 71.29 | 150.1 | 0.475 | 2,767 |
| 2700 | 71.41 | 209.6 | 0.341 | 71.29 | 149.2 | 0.478 | 2,692 |
| 2550 | 71.41 | 193.0 | 0.370 | 71.29 | 134.4 | 0.531 | 2,542 |
| 2472 | 71.41 | 182.2 | 0.392 | 71.28 | 133.7 | 0.533 | 2,460 |
| 2400 | 71.41 | 176.0 | 0.406 | 71.29 | 128.7 | 0.554 | 2,392 |
| 2250 | 71.41 | 165.6 | 0.431 | 71.28 | 117.3 | 0.608 | 2,242 |
| 2163 | 71.39 | 161.5 | 0.442 | 71.28 | 114.1 | 0.625 | 2,147 / 2,152 |
| 2100 | 71.38 | 157.7 | 0.453 | 71.28 | 112.3 | 0.635 | 2,085 |
| 1950 | 71.37 | 154.0 | 0.463 | 71.26 | 111.6 | 0.639 | 1,942 |
| 1854 | 71.37 | 153.0 | 0.467 | 71.25 | 111.4 | 0.640 | 1,845 |
| 1800 | 71.35 | 150.3 | 0.475 | 71.24 | 113.4 | 0.628 | 1,777 / 1,785 |
| 1650 | 71.33 | 151.4 | 0.471 | 71.22 | 109.7 | 0.649 | 1,635 / 1,642 |
| 1500 | 71.30 | 149.2 | 0.478 | 71.19 | 110.6 | 0.644 | 1,492 |
| 1400 | 71.27 | 149.6 | 0.476 | 71.16 | 107.0 | 0.665 | 1,387 |
| 1300 | 71.24 | 147.9 | 0.482 | 71.14 | 107.7 | 0.661 | 1,275 / 1,282 |
| 1200 | 71.20 | 149.4 | 0.477 | 71.12 | 104.4 | 0.681 | 1,192 |
| 1100 | 71.20 | 146.6 | 0.486 | 71.11 | 105.6 | 0.673 | 1,087 |
| 1000 | 71.19 | 149.0 | 0.478 | 71.11 | 103.7 | 0.686 | 990 |
| 900 | 67.66 | 137.8 | 0.491 | not taken (the budget) | | | 893 |
Reading (5080): the rate holds within 0.3 percent of unlocked down to 1,000 MHz on both classes and falls 5.2 percent at 900 MHz on class v4, so the knee is between 1,000 and 900 MHz, a third of the 2,963 MHz boost and lower than the 5090's (the 5080's 84 SMs have more compute headroom per unit of its memory bandwidth, so the memory wait hides the shadow down to a lower clock). Best MH per watt within the 1 percent rate tolerance: class v4 at 1,100 MHz (71.20 MH/s at 146.6 W, 0.486 MH/W; 106.5 W recovered for 0.29 percent rate), class v3 at 1,000 MHz (71.11 at 103.7 W, 0.686; 66.0 W for 0.25 percent). The class v4 premium is 83.4 W unlocked and 41 W at the best points (146.6 W against 105.6 W at 1,100). The draw floors from 1,500 MHz down (about 147 W on v4, 104 W on v3) with the power governor as the only throttle reason say the clock lever is spent by 1,500 MHz on this card. Consequence per tier: a 5080 owner on class v4 locked near 1,100 MHz pays 147 W instead of 253 for 0.3 percent less rate (MH per watt up 72 percent).
What the lever is and is not: nvidia-smi exposes no voltage offset (that is NVAPI's); the clock lock walks the driver's V/F curve, which is where the watts come from; the Mac has no lever (no clock cap on Apple silicon); AMD has the ADLX tune line through igneum-gpu-telemetry or nothing. The knob goes into Ember Tune for 0.3.24 (src/ember.rs: the clock ladder continues below 45 percent of the maximum in 100 MHz steps to a 20 percent floor, the search stops at the first row more than the tolerance under the cap point's rate or on a faulted row, the best MH per watt within tolerance is the point, the fingerprint checked on every step, the result stored per card as the lock_* fields). The 5080 stock rows against the rented 5080 of 7 October (71.16 MH/s at 145 W on driver 580): the rate agrees to 0.4 percent, the watts do not (253 W here); PC 1's three power fields agree, so the difference sits with the rented card's sampler or its cap, the fleet lane's re-measure owed.