Class v6 floor lane 4: section 2 regenerated from the tiers table (the class v5 stock column measured on 14 card classes, the three core figures), section 6 the sweep's rows and its two faults

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 13:08:55 +00:00
parent 25cf4f33a1
commit e985c1929a

View file

@ -86,43 +86,46 @@ fixed by the class), the occupancy (one warp per block, 24 blocks per SM, fixed
## 2. The table: every card class under class v5, the lowest microjoules per hash with the knobs, and the chip edge
"Floor" is the lowest class v5 energy per hash the card reaches with the knobs available to it today. A measured row
cites its job; an estimated row names its method and its band. The chip columns divide the row's floor by the chip's
class v5 energy of section 1. The Apple rows are at the GPU-plus-DRAM meter (the package adds about 17 W on the M5
"Floor" is the lowest class v5 energy per hash the card reaches with the knobs available to it today; "stock" is the
card unlocked at 100 percent under class v5 (the sweep's measured rows, 13:4x to 14:0x UK, say "stock, lock owed" in
the table: no rented host allowed the lock). A measured row cites its job; an estimated row names its method and its
band. The chip columns divide the row's floor by the chip's class v5 energy at three core figures: 1.1 pJ per forced op
(the k lane's synthesised sequencer-core floor at N3, the pessimistic column), 3.2 pJ (the record's k = 0.5) and 6.4
pJ (k = 1, a core as good as a GPU lane); `E_mem` 0.466, 0.321 and 0.14 microjoules. The Apple rows are at the GPU-plus-DRAM meter (the package adds about 17 W on the M5
Max; the wall more); the NVIDIA and AMD rows are whole-card.
| Card | Tier | Class v5 floor, microjoules per hash | Label | GDDR7 board, k = 0.5 / 1 | HBM3 one stack, k = 0.5 / 1 | N2 SRAM die, k = 0.5 / 1 | The point and the band |
|---|---|---|---|---|---|---|---|
| RTX 5090 (32 GB) | 32 GB | 2.33 to 2.37 | measured | 3.0x / 2.1x | 3.6x / 2.4x | 5.0x / 3.0x | 1,300 MHz lock (Balanced): class v4 134.76 MH/s at 312.5 W, +2.0 percent for v5 = 2.37; 1,200 MHz (Efficiency): 133.80 at 305.1, v5 2.33, rate -2.2 percent; the v5lock job read class v5 125.93 at 299.8 W (2.38) with the one-warp worker. Stock 3.48 |
| RTX 5080 (16 GB) | 16 GB | 2.10 | measured (v4 2.06) | 2.7x / 1.9x | 3.3x / 2.2x | 4.5x / 2.7x | 1,100 MHz lock: class v4 71.20 MH/s at 146.6 W (2.06); the draw floors from 1,500 MHz down; the knee between 1,000 and 900. Stock 3.54 (PC 1; the rented host read 18 percent lower watts) |
| RTX 5070 Ti (16 GB) | 16 GB | 1.70 | estimated | 2.2x / 1.5x | 2.6x / 1.8x | 3.7x / 2.2x | stock class v4 78.8 MH/s at 224.0 W (2.84) measured; the Blackwell shape (the v3 draw x0.63, the premium x0.5 at the lock) gives about 131 W at 78.7: 1.66 v4, 1.70 v5; band 1.6 to 2.0 (the rented stock watts may read low against a PC by 15 to 20 percent) |
| RTX 5070 (12 GB) | 12 GB | 1.75 | estimated | 2.2x / 1.6x | 2.7x / 1.8x | 3.8x / 2.2x | stock class v4 58.9 MH/s at 176.1 W (2.99) measured (the paired class v3 run 113 W); the Blackwell shape gives about 103 W: 1.75 v4; band 1.7 to 2.1 |
| RTX 5060 Ti (16 GB) | 16 GB | 2.33 | estimated | 2.9x / 2.1x | 3.6x / 2.4x | 5.0x / 3.0x | stock class v4 30.9 MH/s at 114.8 W (3.72) measured on PC 2 (the Thunderbolt enclosure); the Blackwell shape gives about 72 W: 2.33; band 2.2 to 2.6; the card-in job measures the ladder |
| RTX 5060 (8 GB) | 8 GB | 2.15 | estimated | 2.7x / 1.9x | 3.3x / 2.2x | 4.6x / 2.7x | stock class v3 31.27 MH/s at 75.4 W (2.41) measured; class v4 stock about 110 W (3.5) and the knee about 67 W: 2.15; band 2.0 to 2.5 |
| RTX 4090 (24 GB) | 24 GB | 3.64 | estimated (stock measured, lock owed) | 4.6x / 3.3x | 5.6x / 3.8x | 7.8x / 4.6x | stock measured by floor lane 1: class v3 70.63 MH/s at 253.4 W (3.59), class v4 70.63 at 353.5 W (5.0), idle 34.7 W; the denominator sweep's RunPod 4090 read class v3 70.09 at 222.6 W (3.18, the host spread); the Ada shape (the 4070's tune: v3 x0.72, the premium x0.7) gives about 257 W at the knee: 3.64 v5; band 3.3 to 4.0; the occupancy knob 1 to 2 percent (lane 1) |
| RTX 4080 (16 GB) | 16 GB | 3.33 | estimated | 4.2x / 3.0x | 5.2x / 3.4x | 7.2x / 4.2x | stock class v3 40.71 MH/s at 128.6 W (3.16) and class v4 40.74 at 185.6 W (4.56) measured; the Ada shape gives about 133 W: 3.26 v4, 3.33 v5; the 4080 Super reads the same (42.6 at 134.9 / 195.7 W) |
| RTX 4070 (12 GB) | 12 GB | 3.51 to 3.58 | measured at the tune point | 4.5x / 3.2x | 5.6x / 3.7x | 7.7x / 4.5x | 1,860 MHz lock + the 50 percent cap: class v4 31.08 MH/s at 109.0 W (3.51), v5 3.58; against 3.58 stock on class v3 and 2.57 tuned; the ladder below 1,860 under class v4 is owed (band 3.2 to 3.6) |
| RTX 4070 Ti (12 GB) | 12 GB | 3.60 | estimated | 4.6x / 3.2x | 5.6x / 3.7x | 7.8x / 4.6x | stock class v3 31.24 MH/s at 95.4 W (3.05) and class v4 31.26 at 155.3 W (4.97) measured; the Ada shape gives about 111 W: 3.6; the 4070 Super at a 110 W host cap held 31.26 MH/s at 108.4 W under class v4 (3.47): a cap row measured |
| RTX 4060 Ti (16 GB) | 16 GB | 3.70 | estimated | 4.7x / 3.3x | 5.7x / 3.8x | 8.0x / 4.7x | stock class v3 20.07 MH/s at 79.2 W (3.95) and class v4 20.10 at 102.3 W (5.09) measured; the Ada shape gives about 73 W: 3.7; the 8 GB card reads the same on class v3 (20.10 at 77.5 W) |
| RTX 4060 (8 GB) | 8 GB | 3.77 | estimated | 4.8x / 3.4x | 5.8x / 3.9x | 8.1x / 4.8x | 19.09 MH/s measured on three rented hosts, the watts unread on every one (no power sensor exposed); the 4060 Ti's watts shape: about 95 W stock on class v4 (5.0), 72 W at the knee (3.77); band 3.7 to 4.2 |
| RTX 3090 (24 GB) | 24 GB | 6.40 | estimated | 8.1x / 5.7x | 9.9x / 6.6x | 13.8x / 8.1x | two rented 3090s in the sweep (section 6); the 0.3.12 row 37.79 MH/s at 228.8 W (6.05) and the standing voters 42 to 50 MH/s; the Ampere lever is the power cap and small (the cap knee near 1,800 MHz against a 1,900 to 2,000 stock clock: 10 to 15 percent); the 3090 Ti read class v3 61.95 at 249.5 W (4.03) |
| RTX 3080 (10 GB) | 10 GB | 5.70 | estimated | 7.2x / 5.1x | 8.8x / 5.9x | 12.3x / 7.2x | the 0.3.12 row 40.82 MH/s at 204.9 W (5.02); the sweep's rows pending (section 6); the 3080 Ti measured class v3 58.9 at 293.5 W (4.98) and, at a 220 W host cap with the SM at 749 MHz, class v4 42.6 at 216 W (5.08) with 28 percent of the rate lost: the Ampere cap's known-failed row |
| RTX 3070 (8 GB) | 8 GB | 4.80 | estimated | 6.1x / 4.3x | 7.4x / 5.0x | 10.4x / 6.1x | stock class v3 37.09 MH/s at 142.3 W (3.84) and class v4 37.11 at 196.4 W (5.29) measured; a 10 percent cap: 4.8; the 3070 Ti 38.98 at 177.1 / 267.6 W (4.54 / 6.84) |
| RTX 3060 (12 GB) | 12 GB | 4.90 | estimated | 6.2x / 4.4x | 7.6x / 5.1x | 10.6x / 6.2x | stock class v3 26.89 MH/s at 111.6 W (4.15) measured; class v4 about 5.4 at stock, 4.9 at a mild cap; the 3060 Ti at a 130 W host cap held 33.06 MH/s at 128.7 W under class v4 (3.89): a cap row measured |
| RX 9070 XT (16 GB) | 16 GB | 8.10 | measured (v4 7.9) | 10.3x / 7.3x | 12.6x / 8.4x | 17.5x / 10.3x | the ADLX grid: 18.9 MH/s at 149.3 W (-500 MHz core, -30 percent power), 24 percent under the 202 W stock point (10.7); the shadow costs this card about 3 W (199 W on class v3, 202 on v4): it is the AMD read path, not the ALU, that sets its joule |
| RX 7900 XTX (24 GB) | 24 GB | 8.00 | estimated | 10.1x / 7.2x | 12.4x / 8.3x | 17.3x / 10.2x | not rentable, not owned: by the 9070 XT's dependent-read rate per channel (150 M per second per 32-bit channel) 24 channels give about 28 MH/s at about 300 W stock (10.7), 7.5 to 8.5 with the knob |
| RX 7800 XT (16 GB) | 16 GB | 9.50 | estimated | 12.0x / 8.5x | 14.7x / 9.8x | 20.5x / 12.1x | 16 channels: about 18 MH/s at about 230 W stock (12.8); the knob 9 to 10 |
| RX 7600 XT (16 GB) | 16 GB | 13.0 | estimated | 16.5x / 11.7x | 20.2x / 13.4x | 28.0x / 16.5x | 8 channels: about 9 MH/s at about 160 W stock (18); the knob 12 to 14; the card-in queue holds one for a PC measurement |
| RX 9060 XT (16 GB) | 16 GB | 10.0 | estimated | 12.7x / 9.0x | 15.5x / 10.3x | 21.6x / 12.7x | 8 channels of GDDR6 on RDNA 4: about 9.5 MH/s at about 130 W stock (13.7); the knob 9.5 to 10.5; the card-in queue holds one |
| H100 SXM (80 GB) | DC | 2.00 | estimated (stock measured, lock owed) | 2.5x / 1.8x | 3.1x / 2.1x | 4.3x / 2.5x | stock class v3 248.9 MH/s at 411.7 W (1.65) and class v4 250.6 at 695.1 W (2.77) measured (floor lane 1 the same: v3 254.6 at 450.4 W, v4 251.9 at 699 W with the 700 W cap binding and the clock throttled to 1,750 MHz; class v5 = class v4 on this card, 250.2 at 698.8 W; its -lmc was accepted and the clock stayed 2,619 MHz); the premium 283 W (11.2 pJ per op); a lock on an owned host (Hopper at 1,980 MHz boost, HBM3-bound): about 500 W (2.0); band 1.9 to 2.4; rented hosts refuse `-lgc` |
| L40S (48 GB) | DC | 3.50 | estimated | 4.4x / 3.1x | 5.4x / 3.6x | 7.5x / 4.4x | stock class v3 56.4 MH/s at 220.7 W (3.91) and class v4 56.4 at 277.6 W (4.92) measured; the Ada shape gives about 198 W (3.5) |
| A100 SXM (80 GB) | DC | 3.10 | estimated | 3.9x / 2.8x | 4.8x / 3.2x | 6.7x / 3.9x | stock class v3 138.3 MH/s at 268.6 W (1.94) and class v4 138.0 at 489.3 W (3.54) measured (the premium 221 W = 15.8 pJ per op, the dearest ALU in the record); the card already runs at 1,410 MHz, a cap recovers about 60 W: 3.1 |
| Intel Arc B580 (12 GB) | 12 GB | 10.4 | estimated | 13.2x / 9.4x | 16.2x / 10.8x | 22.5x / 13.3x | 10.6 to 11 MH/s measured on PC 2 (the enclosure), the watts unread; about 110 W by the board's class (OWED); no clock lever in the app for Intel |
| Intel Arc A750 (8 GB) | 8 GB | 18.8 | estimated | 23.8x / 16.9x | 29.1x / 19.4x | 40.5x / 23.9x | unmeasured; Alchemist at 256-bit GDDR6: about 8 MH/s at about 150 W |
| Apple M5 Max | Apple | 1.40 to 1.43 | measured (v4 1.40) | 1.8x / 1.3x | 2.2x / 1.5x | 3.1x / 1.8x | the GPU and DRAM meter: class v3 27.08 MH/s at 21.0 W (0.78), class v4 26.67 at 37.3 W (1.40); the package about 54 W (2.0); no lever |
| Apple M4 Max | Apple | 1.55 | estimated | 2.0x / 1.4x | 2.4x / 1.6x | 3.3x / 2.0x | the same 512-bit LPDDR5X (8,533 against 9,600 MT/s) and a 40-core GPU on N3E: about 26 MH/s; class v3 about 0.85, class v5 about 1.55 at the meter (the M5 Max's rows scaled); a devnet row from an unnamed Apple laptop read 24.3 MH/s on class v4 |
| Apple M4 Pro | Apple | 1.65 | estimated | 2.1x / 1.5x | 2.6x / 1.7x | 3.6x / 2.1x | 256-bit LPDDR5X, a 20-core GPU: about 13.5 MH/s; the shadow's per-hash premium does not shrink with the part, so about 1.65 at the meter |
| Apple M3 Max | Apple | 1.75 | estimated | 2.2x / 1.6x | 2.7x / 1.8x | 3.8x / 2.2x | 512-bit LPDDR5-6400, a 40-core GPU on N3B: about 24 MH/s; about 1.75 at the meter |
| Card | Tier | Class v5 floor, microjoules per hash | Label | Class v5 stock, microjoules per hash (label) | GDDR7 board at 1.1 / 3.2 / 6.4 pJ per forced op | HBM3 one stack | N2 SRAM die | The point, the band and the source |
|---|---|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 | 32 GB | 2.33 | measured | 3.48 (measured) | 4.0x / 3.0x / 2.1x | 5.4x / 3.6x / 2.4x | 9.3x / 5.0x / 3.0x | 1200 MHz at 100 percent. class v4 best MH per watt (133.80 MH/s at 305.1 W, 2.2 percent of rate given); under class v5 +2.0 percent watts: 311.2 W, 2.33 Stock: the power cap is flat on this hash (0.41 MH/W under every limit), so no power rung is used on any tier |
| NVIDIA GeForce RTX 5080 | 16 GB | 2.06 | measured | 3.48 (measured) | 3.6x / 2.6x / 1.9x | 4.8x / 3.2x / 2.1x | 8.2x / 4.4x / 2.6x | 1100 MHz at 100 percent. class v4 best MH per watt; the draw floors from 1,500 MHz down; the knee is between 1,000 and 900 MHz (900 costs 5.2 percent); class v5 about 2.10 Stock: class v5 measured on a Vast 5080 (13:49 UK): 71.35 MH/s at 248.0 W (3.48, 39 samples, SM 2,769, memory 14,801); class v3 71.16 at 159.4 W on the same host; PC 1 read class v4 71.41 at 253.1 W: the three agree within 3 percent on this card; stock, lock owed on the rented host |
| NVIDIA GeForce RTX 5070 Ti | 16 GB | 1.7 | estimated | 2.84 (measured) | 2.9x / 2.2x / 1.5x | 3.9x / 2.6x / 1.8x | 6.8x / 3.7x / 2.2x | 1100 MHz at 100 percent. the Blackwell shape applied to the measured stock rows (the v3 draw x0.63, the premium x0.5): band 1.6 to 2.0; the rented host refused -lgc, so the knee is the search's to find Stock: rented pod 8 October 10:09Z: class v3 78.69 MH/s at 140.8 W, class v4 78.78 at 224.0 W |
| NVIDIA GeForce RTX 5070 | 12 GB | 1.75 | estimated | 2.99 (measured) | 3.0x / 2.2x / 1.6x | 4.0x / 2.7x / 1.8x | 7.0x / 3.8x / 2.2x | 1100 MHz at 100 percent. the Blackwell shape on the measured stock rows (paired class v3 113 W): band 1.7 to 2.1 Stock: rented v4watts pod 8 October; the 7 October pod read class v3 52.0 MH/s at 102.8 W (the rate differs by host) |
| NVIDIA GeForce RTX 5060 Ti | 16 GB | 2.36 | estimated | 4.02 (measured) | 4.1x / 3.0x / 2.1x | 5.5x / 3.7x / 2.4x | 9.4x / 5.1x / 3.0x | 1100 MHz at 100 percent. the Blackwell shape on the PC 2 stock row: band 2.2 to 2.6; re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: class v5 measured on a Vast 5060 Ti (13:5x UK): 30.84 MH/s at 123.9 W (4.02, 96 samples, SM 2,835); class v3 30.75 at 82.6 W on the same host; PC 2's enclosure row read class v4 30.9 at 114.8 W; stock, lock owed |
| NVIDIA GeForce RTX 5060 | 8 GB | 2.21 | estimated | 3.76 (measured) | 3.8x / 2.8x / 2.0x | 5.1x / 3.4x / 2.3x | 8.8x / 4.8x / 2.8x | 1100 MHz at 100 percent. class v3 measured 31.27 MH/s at 75.4 W; re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: class v5 measured (13:51 UK): 30.80 MH/s at 115.7 W (3.76, 99 samples, SM 2,715, memory 13,801); class v3 30.76 at 78.8 W on the same host (the shadow +37 W = 12 pJ per op); stock, lock owed |
| NVIDIA GeForce RTX 4090 | 24 GB | 3.58 | estimated | 5.0 (measured) | 6.2x / 4.5x / 3.2x | 8.3x / 5.6x / 3.7x | 14.2x / 7.7x / 4.5x | 1860 MHz at 60 percent. the Ada shape (the 4070's measured tune: v3 x0.72, the premium x0.7) on the measured stock rows (class v3 253.4 W, class v4 353.5 W, floor lane 1, 8 October): band 3.3 to 4.0; re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: floor lane 1's rented 4090 at a 450 W limit: class v3 70.63 MH/s at 253.4 W (3.59), class v4 70.63 at 353.5 W (5.0); the denominator sweep's Vast 4090 carried a 250 W host limit: class v5 62.68 at 249.9 W (3.99, the SM at 2,579 MHz) and class v3 62.40 at 200.5 W, so a 250 W cap held 89 percent of the rate at 71 percent of the draw (an Ada cap row measured); stock, lock owed |
| NVIDIA GeForce RTX 4080 | 16 GB | 3.51 | estimated | 4.92 (measured) | 6.1x / 4.4x / 3.2x | 8.1x / 5.4x / 3.6x | 14.0x / 7.6x / 4.5x | 1860 MHz at 60 percent. the Ada shape on the measured stock rows (class v3 128.6 W, class v4 185.6 W): band 3.1 to 3.7; re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: class v5 measured (13:50 UK): 40.80 MH/s at 200.7 W (4.92, 72 samples, SM 2,760); class v3 40.69 at 141.5 W on the same host; the v4watts pod read class v4 40.74 at 185.6 W; stock, lock owed |
| NVIDIA GeForce RTX 4070 | 12 GB | 3.58 | measured | 5.82 (measured) | 6.2x / 4.5x / 3.2x | 8.3x / 5.6x / 3.7x | 14.2x / 7.7x / 4.5x | 1860 MHz at 50 percent. Ember run 6's tune point (1,863 MHz at 50 percent) under class v4 (sh256x27): 31.08 MH/s at 109.0 W, +30 W over class v3 at the same point for no rate; class v5 about 3.58; the ladder below 1,860 under class v4 is owed (band 3.2 to 3.6) Stock: class v5 measured on a Vast 4070 (13:54 UK): 28.43 MH/s at 165.5 W (5.82, 108 samples, SM 2,820, memory 9,801, a 210 W host limit not binding); class v3 28.40 at 116.3 W on the same host (the shadow +49 W = 17 pJ per op at the stock clock); Ember run 6's PC 1 stock on class v3 read 28.72 at 106.0 W; the tune point (1,860 MHz, 50 percent) under class v4 measured 31.08 at 109.0 W (3.51), so the lock takes 34 percent of this card's class v5 draw for no rate |
| NVIDIA GeForce RTX 4070 Ti | 12 GB | 3.6 | estimated | 4.97 (measured) | 6.2x / 4.6x / 3.2x | 8.3x / 5.6x / 3.7x | 14.3x / 7.8x / 4.6x | 1860 MHz at 60 percent. the Ada shape on the measured stock rows (class v3 95.4 W, class v4 155.3 W); the 4070 Super at a 110 W host cap held 31.26 MH/s at 108.4 W under class v4 (3.47): a cap row measured on a sibling Stock: rented pods 8 October |
| NVIDIA GeForce RTX 4060 Ti | 16 GB | 3.7 | estimated | 5.09 (measured) | 6.4x / 4.7x / 3.3x | 8.6x / 5.7x / 3.8x | 14.7x / 8.0x / 4.7x | 1860 MHz at 60 percent. the Ada shape on the measured stock rows (class v3 79.2 W, class v4 102.3 W); the 8 GB card reads the same on class v3 (20.10 MH/s at 77.5 W) Stock: rented pods 7 and 8 October |
| NVIDIA GeForce RTX 4060 | 8 GB | 3.77 | estimated | 4.98 (estimated) | 6.5x / 4.8x / 3.4x | 8.7x / 5.8x / 3.9x | 15.0x / 8.1x / 4.8x | 1860 MHz at 60 percent. 19.09 MH/s measured on three rented hosts, the watts unread on every one (no power sensor): the watts are the 4060 Ti's shape; band 3.7 to 4.2 Stock: watts OWED: every rented 4060 host exposed no power sensor |
| NVIDIA GeForce RTX 3090 | 24 GB | 4.51 | estimated | 5.03 (measured) | 7.8x / 5.7x / 4.1x | 10.4x / 7.0x / 4.7x | 17.9x / 9.7x / 5.7x | 1800 MHz at 80 percent. the Ampere cap lever (10 to 15 percent of the draw; a cap that takes the SM under 1,800 MHz costs rate) on the measured class v5 stock row; the 3090 Ti read class v3 61.95 MH/s at 249.5 W (4.03) Stock: the denominator sweep, RunPod, 8 October 13:25 UK: class v3 60.01 MH/s at 294.6 W (4.91), class v5 61.49 at 309.4 W (5.03, 46 samples, SM 1,673 MHz, memory 9,501); the class v5 premium on this card 5 percent at stock; stock, lock owed (the host refused -lgc and -lmc) |
| NVIDIA GeForce RTX 3080 | 10 GB | 4.2 | estimated | 4.54 (estimated) | 7.3x / 5.3x / 3.8x | 9.7x / 6.5x / 4.3x | 16.7x / 9.1x / 5.3x | 1800 MHz at 80 percent. the Ampere cap lever (10 percent, no lower: a 170 W cap took 34 percent of the class v5 rate on this card, a 180 W cap 11 percent) Stock: the sweep's two 3080 hosts both carried a limit under the class v5 draw: at 170 W class v3 50.71 MH/s at 169.2 W (3.34) held the rate and class v5 fell to 33.37 MH/s at 169.9 W (5.09) with the SM at 689 MHz; at 180 W class v3 49.83 at 177.8 W and class v5 44.31; the uncapped class v5 stock by the 3080 Ti's +30 percent: about 230 W (4.54); the Ampere cap under the shadow costs rate one for one: measured twice |
| NVIDIA GeForce RTX 3070 | 8 GB | 4.8 | estimated | 5.29 (measured) | 8.3x / 6.1x / 4.3x | 11.1x / 7.4x / 5.0x | 19.1x / 10.4x / 6.1x | 1800 MHz at 80 percent. the Ampere cap lever on the measured stock rows (class v3 142.3 W, class v4 196.4 W); the 3070 Ti 38.98 MH/s at 177.1 / 267.6 W Stock: rented pods 8 October |
| NVIDIA GeForce RTX 3060 | 12 GB | 5.77 | estimated | 6.4 (measured) | 10.0x / 7.3x / 5.2x | 13.3x / 8.9x / 6.0x | 23.0x / 12.4x / 7.3x | 1800 MHz at 80 percent. the Ampere cap lever (10 percent) on the measured class v4 stock row; the 3060 Ti at a 130 W host cap held 33.06 MH/s at 128.7 W under class v4 (3.89): a cap row measured on a sibling Stock: class v5 measured on a Vast 3060 (13:53 UK): 26.53 MH/s at 169.8 W (6.40) at the host's 170 W limit (the limit binding: the class v4 pod read 26.89 at 166.1 W, 6.18); class v3 26.53 at 120.4 W on the same host; stock, lock owed |
| AMD Radeon RX 9070 XT | 16 GB | 7.9 | measured | 10.7 (measured) | 13.7x / 10.0x / 7.1x | 18.3x / 12.3x / 8.2x | 31.4x / 17.0x / 10.0x | ADLX -500 MHz, -30 percent. the ADLX grid (24 rows, PC 1): the rate flat at 18.9 MH/s across the grid, the best point -500 MHz core and -30 percent power, 24 percent under stock; class v5 about 8.1 Stock: the app's own power reading at the stock point, 8 October 07:11 UK |
| AMD Radeon RX 7900 XTX | 24 GB | 8.0 | estimated | 10.7 (estimated) | 13.9x / 10.1x / 7.2x | 18.5x / 12.4x / 8.3x | 31.8x / 17.3x / 10.2x | ADLX -500 MHz, -30 percent. not rentable, not owned: the 9070 XT's dependent-read rate per channel (150 M per second per 32-bit channel) on 24 channels; the knob by the 9070 XT grid Stock: modelled |
| AMD Radeon RX 7800 XT | 16 GB | 9.5 | estimated | 12.8 (estimated) | 16.5x / 12.0x / 8.5x | 22.0x / 14.7x / 9.8x | 37.8x / 20.5x / 12.1x | ADLX -500 MHz, -30 percent. 16 channels of GDDR6 at the 9070 XT's per-channel rate; the knob by the 9070 XT grid Stock: modelled |
| AMD Radeon RX 7600 XT | 16 GB | 13.0 | estimated | 17.8 (estimated) | 22.5x / 16.5x / 11.7x | 30.1x / 20.2x / 13.4x | 51.7x / 28.0x / 16.5x | ADLX -500 MHz, -30 percent. 8 channels; the card-in queue holds one for a PC measurement Stock: modelled |
| AMD Radeon RX 9060 XT | 16 GB | 10.0 | estimated | 13.7 (estimated) | 17.3x / 12.7x / 9.0x | 23.1x / 15.5x / 10.3x | 39.8x / 21.6x / 12.7x | ADLX -500 MHz, -30 percent. 8 channels of GDDR6 on RDNA 4; the card-in queue holds one Stock: modelled |
| NVIDIA H100 80GB HBM3 | DC | 2.0 | estimated | 2.58 (measured) | 3.5x / 2.5x / 1.8x | 4.6x / 3.1x / 2.1x | 8.0x / 4.3x / 2.5x | 1400 MHz at 80 percent. the premium 240 to 280 W at stock says half of it is the clock; a lock on an owned host (Hopper at 1,980 MHz boost, HBM3-bound): about 500 W (2.0); band 1.9 to 2.4; rented hosts refuse -lgc Stock: the denominator sweep, RunPod, 13:25 UK: class v3 252.96 MH/s at 382.4 W (1.51), class v5 241.34 at 621.6 W (2.58, 7 samples, the SM throttled to 1,590 MHz under the 700 W cap); the 7 and 8 October pods read class v4 250.6 at 695.1 W (2.77); -lmc answered 'use --lock-memory-clocks-deferred' and the memory clock stayed 2,619; stock, lock owed |
| NVIDIA L40S | DC | 3.78 | estimated | 5.29 (measured) | 6.5x / 4.8x / 3.4x | 8.7x / 5.9x / 3.9x | 15.0x / 8.2x / 4.8x | 1860 MHz at 60 percent. the Ada shape on the measured stock rows (class v3 220.7 W, class v4 277.6 W); re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: class v5 measured (13:49 UK): 56.49 MH/s at 298.7 W (5.29, 46 samples, SM 2,520); class v3 56.35 at 221.5 W on the same host; the v4watts pod read class v4 56.4 at 277.6 W; stock, lock owed |
| NVIDIA A100-SXM4-80GB | DC | 2.89 | estimated | 2.99 (measured) | 5.0x / 3.7x / 2.6x | 6.7x / 4.5x / 3.0x | 11.5x / 6.2x / 3.7x | 1200 MHz at 80 percent. the card already runs at 1,410 MHz and the 400 W host limit binds under class v5 (SM clock gives); a cap under it costs rate on Ampere, so the floor is within 5 percent of the measured stock row Stock: the denominator sweep, RunPod, 13:24 UK: class v3 138.06 MH/s at 294.9 W (2.14), class v5 133.11 at 398.1 W (2.99, 19 samples, at the host's 400 W limit, memory 1,593 MHz); the v4watts pod on a 500 W host read class v4 138.0 at 489.3 W (3.54); -lmc 'not supported' on this card; stock, lock owed |
| Intel Arc B580 | 12 GB | 10.4 | estimated | 10.4 (estimated) | 18.0x / 13.2x / 9.3x | 24.1x / 16.1x / 10.7x | 41.4x / 22.4x / 13.2x | stock, no lever. 10.6 to 11 MH/s measured on PC 2 (the enclosure), the watts unread; about 110 W by the board's class (OWED); no clock lever in the app for Intel |
| Intel Arc A750 | 8 GB | 18.8 | estimated | 18.8 (estimated) | 32.6x / 23.8x / 16.9x | 43.5x / 29.2x / 19.4x | 74.8x / 40.5x / 23.9x | stock, no lever. unmeasured; modelled from the B580 and the Alchemist board |
| Apple M5 Max | Apple | 1.4 | measured | 1.4 (measured) | 2.4x / 1.8x / 1.3x | 3.2x / 2.2x / 1.4x | 5.6x / 3.0x / 1.8x | stock, no lever. the GPU and DRAM channels of the IOReport meter (not the wall): class v3 27.08 MH/s at 21.0 W (0.78), class v4 26.67 at 37.3 W (1.40), class v5 about 1.43; the package about 17 W more; no lever (no clock cap on Apple silicon) |
| Apple M4 Max | Apple | 1.55 | estimated | 1.55 (estimated) | 2.7x / 2.0x / 1.4x | 3.6x / 2.4x / 1.6x | 6.2x / 3.3x / 2.0x | stock, no lever. the same 512-bit LPDDR5X (8,533 against 9,600 MT/s) and a 40-core GPU on N3E: about 26 MH/s; class v3 about 0.85 at the meter, class v5 about 1.55 (the M5 Max's rows scaled); a devnet row from an unnamed Apple laptop read 24.3 MH/s on class v4 |
| Apple M4 Pro | Apple | 1.65 | estimated | 1.65 (estimated) | 2.9x / 2.1x / 1.5x | 3.8x / 2.6x / 1.7x | 6.6x / 3.6x / 2.1x | stock, no lever. 256-bit LPDDR5X and a 20-core GPU: about 13.5 MH/s; the shadow's per-hash premium does not shrink with the part, so class v5 reads about 1.65 at the meter |
| Apple M3 Max | Apple | 1.75 | estimated | 1.75 (estimated) | 3.0x / 2.2x / 1.6x | 4.0x / 2.7x / 1.8x | 7.0x / 3.8x / 2.2x | stock, no lever. 512-bit LPDDR5-6400 and a 40-core GPU on N3B: about 24 MH/s; class v5 about 1.75 at the meter |
To fold any other chip core figure (the research lane's convention, 13:5x): the edge on a row is the card's class v5
microjoules over `E_mem + 101,170 x c`, with `E_mem` 0.466 (GDDR7 board), 0.321 (one HBM3 stack), 0.14 (the N2 SRAM
@ -264,26 +267,51 @@ The re-measure rule after a class flip, as the table carries it (`remeasure_rule
3. The set also re-measures weekly (`PERIOD_S`) and on a driver major change (the prior key carries driver major and
class), and a confirm check that beats its prior by over 1 percent on MH per watt asks for the full plan.
## 6. The denominator sweep (rented, 8 October 13:05 onward, the stock rows the record lacked)
## 6. The denominator sweep (rented, 8 October 13:05 to 14:1x UK, the stock rows the record lacked)
Three drivers of the fleet lane's model sweep (the third, from 13:21, on the coordinator's full-spend order: 16 card classes in parallel on `box-cardbench-v2m.sh`, the class v5 kit's bench run for 200 batches of 2^24 under the sampler so the class v5 watts are a measured mean, and a memory-clock try (`-lmc` to the card's maximum, a second class v5 run when the host accepts it)) (`box-cardbench-v2.sh` for class v3, the v5 kit's fingerprint and the
memprobe ceiling; `box-cardbench-v4watts.sh` for the class v4 shape with the paired class v3 run), one one-shot pod per
card class on RunPod or Vast under the USD 1.00 per hour consumer cap, destroyed at the end of each row; the rows
appended to `~/igneum-fleet/cardbench/rows.jsonl` with `sweep: 2026-10-08-denominator`; spend inside the lane's USD 100.
No rented host allowed `-lgc`, so these are stock rows and the knees above stay estimated. The rows are read into the
table of section 2 as they land; this section lists them.
Three drivers of the fleet lane's model sweep and then a direct ssh runner: `box-cardbench-v2.sh` (class v3, the v5
kit's fingerprint, the memprobe ceiling), `box-cardbench-v4watts.sh` (the class v4 shape with the paired class v3 run)
and, from 13:21 on the coordinator's full-spend order, `box-cardbench-v2m.sh` on 16 card classes in parallel (the
class v5 kit's bench run for 200 batches of 2^24 under the 1 Hz sampler, so the class v5 watts are a measured mean
over busy samples from 8 s in; then `-lmc` to the card's maximum and a second class v5 run when the host accepted it).
One one-shot pod per card class on RunPod or Vast under the consumer cap, destroyed at the end of each row; the rows
in `~/igneum-fleet/cardbench/rows.jsonl` under `sweep: 2026-10-08-denominator` and `-v5`; the lane's spend about USD 5
of its USD 100 (the pods ran 3 to 10 minutes each). No rented host allowed `-lgc`; the two that accepted `-lmc` (the
H100 and A100) left the memory clock where it was, so every row here is stock and the knees of section 2 stay
estimated where no founder machine holds the card.
| Card | Class v3 (MH/s at W, microjoules) | Class v4 shape (MH/s at W, microjoules) | v4 over v3 watts | Host, driver, cost | Reading |
Two faults found on the way, for the fleet lane: (1) the sweep driver's ssh wait read every Vast pod of the 13:21 wave
as "never answered ssh in 10 minutes" while a direct `ssh` to the same host and port answered at once (13 pods, all
live; the rows were then taken by a direct runner, `~/igneum-fleet/denom-direct.sh`, which keys on the provider's
own ssh host and port and nothing in the registry); (2) the registry lock `~/igneum-fleet/boxes.json.lock` was held
from 13:32 by three `dn3-q05.py` processes of another lane, so every `Registry.patch` from this lane blocked (a
re-rent sat 12 minutes on a live pod and then destroyed it); the lane's scripts now write the registry best-effort
under a 15 s alarm. The rented stock rows against the PC rows on the one card measured both ways (the 5080): class v5
71.35 MH/s at 248.0 W rented against class v4 71.41 at 253.1 W on PC 1, within 3 percent, so the 18 percent class v3
spread of the morning was that host and not the method.
| Card | Host | Class v3 (MH/s at W = microjoules) | Class v4 shape or class v5 (MH/s at W = microjoules; samples) | The knobs on the host | Cost |
|---|---|---|---|---|---|
| RTX 4090 24 GB | 70.09 MH/s at 222.6 W (3.18; RunPod, driver 595.91, 13:07); class v5 kit 70.25 MATCH | pending | | USD 0.74 per hour, 0.03 | floor lane 1's rented 4090 read 70.63 at 253.4 W on the same class: the host spread of 12 percent on the watts at the same rate |
| RTX 5060 8 GB | 24.77 MH/s at 80.7 W (3.26; Vast, 13:11); class v5 kit 31.31 MATCH (a rate the v3 run on this host did not reach) | 16.23 MH/s at 87.8 W (5.41; a second Vast host, 13:10, v4 over v3 +20 percent on that host) | +20 percent | USD 0.02 | the rate varies by host 16 to 31 MH/s on this card class: the 7 October pod read 31.27 on class v3; the per-hash energy at the matched rate stands at 2.4 (v3) |
| RTX 3090 24 GB | pending (two Vast hosts rented 13:05) | pending | | USD 0.14 per hour | |
| RTX 3080 10 GB | pending | pending | | | the first Vast offer vanished at the rent (HTTP 410) |
| RTX 5080 16 GB | the RunPod host failed `cuInit` (a bad host, destroyed); Vast next | pending | | | |
| RTX 5060 Ti 16 GB | pending | pending | | | |
| RTX 5060 8 GB | pending | pending | | | |
| RTX 3060 12 GB | pending | pending | | | |
| RTX 4070 12 GB | pending | pending | | | |
| RTX 4090 24 GB | runpod m1tstl1h1b7r0o (595.91.07), limit 450.0 W | 70.094 at 222.6 W = 3.18 | class v5 70.253 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.032 |
| RTX 5060 8 GB | vast 54840005 (580.126.09), limit 145.0 W | | class v4 16.234 at 87.8 W = 5.41 (v4 over v3 20.4 percent on this host) | | USD 0.009 |
| RTX 5060 8 GB | vast 54840004 (580.126.09), limit 145.0 W | 24.774 at 80.7 W = 3.26 | class v5 31.312 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.011 |
| RTX 5080 16 GB | vast 54840017 (570.133.07), limit 300.0 W | | class v4 70.951 at 239.8 W = 3.38 (v4 over v3 41.1 percent on this host) | | USD 0.052 |
| RTX 3080 10 GB | vast 54840023 (580.159.03), limit 130.0 W | | class v4 24.007 at 129.2 W = 5.38 (v4 over v3 3.0 percent on this host) | | USD 0.025 |
| RTX 3060 12 GB | vast 54840846 (580.126.20), limit 170.0 W | | class v4 26.886 at 166.1 W = 6.18 (v4 over v3 44.6 percent on this host) | | USD 0.007 |
| RTX 4070 12 GB | vast 54840856 (580.173.02), limit 200.0 W | 30.829 at 107.1 W = 3.47 | class v5 30.878 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.023 |
| A100 80 GB | runpod miuoe6gkcb1rhz (580.126.16), limit 400.0 W | 138.062 at 294.9 W = 2.14 | class v5 133.548 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.075 |
| H100 80 GB | runpod zosqj1s802dk3x (580.126.09), limit 700.0 W | 252.959 at 382.4 W = 1.51 | class v5 241.342 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.216 |
| RTX 3090 24 GB | runpod sbrhhunqui25d0 (580.65.06), limit 310.0 W | 60.008 at 294.6 W = 4.91 | class v5 61.486 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.014 |
| RTX 3080 10 GB | vast 54841582 (580.159.03), limit 180.0 W | 49.832 at 177.8 W = 3.57 | class v5 44.308 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.024 |
| RTX 5080 16 GB | vast 54845199 (570.133.07), limit 300.0 W | 71.16 at 159.4 W = 2.24 | class v5 71.349 at 248.0 W = 3.476; 39 samples, SM 2769, memory 14801 | -lgc refused; -lmc refused | USD 0.014 |
| L40S 48 GB | vast 54845608 (595.71.05), limit 350.0 W | 56.354 at 221.5 W = 3.93 | class v5 56.493 at 298.7 W = 5.287; 46 samples, SM 2520, memory 9001 | -lgc refused; -lmc refused | USD 0.04 |
| RTX 4080 16 GB | vast 54845190 (595.84), limit 320.0 W | 40.694 at 141.5 W = 3.48 | class v5 40.8 at 200.7 W = 4.919; 72 samples, SM 2760, memory 10801 | -lgc refused; -lmc refused | USD 0.017 |
| RTX 3080 10 GB | vast 54845161 (580.178.04), limit 170.0 W | 50.708 at 169.2 W = 3.34 | class v5 33.37 at 169.9 W = 5.091; 90 samples, SM 689, memory 9251 | -lgc refused; -lmc refused | USD 0.008 |
| RTX 4090 24 GB | vast 54845886 (595.84), limit 250.0 W | 62.401 at 200.5 W = 3.21 | class v5 62.684 at 249.9 W = 3.987; 46 samples, SM 2579, memory 10251 | -lgc refused; -lmc refused | USD 0.02 |
| RTX 5060 8 GB | vast 54845177 (595.84), limit 140.0 W | 30.76 at 78.8 W = 2.56 | class v5 30.804 at 115.7 W = 3.756; 99 samples, SM 2715, memory 13801 | -lgc refused; -lmc refused | USD 0.011 |
| RTX 5060 Ti 16 GB | vast 54845162 (580.126.09), limit 180.0 W | 30.747 at 82.6 W = 2.69 | class v5 30.843 at 123.9 W = 4.017; 96 samples, SM 2835, memory 13801 | -lgc refused; -lmc refused | USD 0.013 |
| RTX 3060 12 GB | vast 54845154 (580.126.09), limit 170.0 W | 26.534 at 120.4 W = 4.54 | class v5 26.53 at 169.8 W = 6.4; 117 samples, SM 1806, memory 7301 | -lgc refused; -lmc refused | USD 0.008 |
| RTX 4070 12 GB | vast 54845860 (580.126.09), limit 210.0 W | 28.401 at 116.3 W = 4.09 | class v5 28.433 at 165.5 W = 5.821; 108 samples, SM 2820, memory 9801 | -lgc refused; -lmc refused | USD 0.011 |
## 7. Consequences per tier