From e985c1929af7d762c8e9887259f744841069ea64 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 13:08:55 +0000 Subject: [PATCH] Class v6 floor lane 4: section 2 regenerated from the tiers table (the class v5 stock column measured on 14 card classes, the three core figures), section 6 the sweep's rows and its two faults Co-Authored-By: Claude Fable 5.1 --- docs/analysis/class-v6/floor/denominator.md | 132 ++++++++++++-------- 1 file changed, 80 insertions(+), 52 deletions(-) diff --git a/docs/analysis/class-v6/floor/denominator.md b/docs/analysis/class-v6/floor/denominator.md index 63f49e401..7219ab2ed 100644 --- a/docs/analysis/class-v6/floor/denominator.md +++ b/docs/analysis/class-v6/floor/denominator.md @@ -86,43 +86,46 @@ fixed by the class), the occupancy (one warp per block, 24 blocks per SM, fixed ## 2. The table: every card class under class v5, the lowest microjoules per hash with the knobs, and the chip edge -"Floor" is the lowest class v5 energy per hash the card reaches with the knobs available to it today. A measured row -cites its job; an estimated row names its method and its band. The chip columns divide the row's floor by the chip's -class v5 energy of section 1. The Apple rows are at the GPU-plus-DRAM meter (the package adds about 17 W on the M5 +"Floor" is the lowest class v5 energy per hash the card reaches with the knobs available to it today; "stock" is the +card unlocked at 100 percent under class v5 (the sweep's measured rows, 13:4x to 14:0x UK, say "stock, lock owed" in +the table: no rented host allowed the lock). A measured row cites its job; an estimated row names its method and its +band. The chip columns divide the row's floor by the chip's class v5 energy at three core figures: 1.1 pJ per forced op +(the k lane's synthesised sequencer-core floor at N3, the pessimistic column), 3.2 pJ (the record's k = 0.5) and 6.4 +pJ (k = 1, a core as good as a GPU lane); `E_mem` 0.466, 0.321 and 0.14 microjoules. The Apple rows are at the GPU-plus-DRAM meter (the package adds about 17 W on the M5 Max; the wall more); the NVIDIA and AMD rows are whole-card. -| Card | Tier | Class v5 floor, microjoules per hash | Label | GDDR7 board, k = 0.5 / 1 | HBM3 one stack, k = 0.5 / 1 | N2 SRAM die, k = 0.5 / 1 | The point and the band | -|---|---|---|---|---|---|---|---| -| RTX 5090 (32 GB) | 32 GB | 2.33 to 2.37 | measured | 3.0x / 2.1x | 3.6x / 2.4x | 5.0x / 3.0x | 1,300 MHz lock (Balanced): class v4 134.76 MH/s at 312.5 W, +2.0 percent for v5 = 2.37; 1,200 MHz (Efficiency): 133.80 at 305.1, v5 2.33, rate -2.2 percent; the v5lock job read class v5 125.93 at 299.8 W (2.38) with the one-warp worker. Stock 3.48 | -| RTX 5080 (16 GB) | 16 GB | 2.10 | measured (v4 2.06) | 2.7x / 1.9x | 3.3x / 2.2x | 4.5x / 2.7x | 1,100 MHz lock: class v4 71.20 MH/s at 146.6 W (2.06); the draw floors from 1,500 MHz down; the knee between 1,000 and 900. Stock 3.54 (PC 1; the rented host read 18 percent lower watts) | -| RTX 5070 Ti (16 GB) | 16 GB | 1.70 | estimated | 2.2x / 1.5x | 2.6x / 1.8x | 3.7x / 2.2x | stock class v4 78.8 MH/s at 224.0 W (2.84) measured; the Blackwell shape (the v3 draw x0.63, the premium x0.5 at the lock) gives about 131 W at 78.7: 1.66 v4, 1.70 v5; band 1.6 to 2.0 (the rented stock watts may read low against a PC by 15 to 20 percent) | -| RTX 5070 (12 GB) | 12 GB | 1.75 | estimated | 2.2x / 1.6x | 2.7x / 1.8x | 3.8x / 2.2x | stock class v4 58.9 MH/s at 176.1 W (2.99) measured (the paired class v3 run 113 W); the Blackwell shape gives about 103 W: 1.75 v4; band 1.7 to 2.1 | -| RTX 5060 Ti (16 GB) | 16 GB | 2.33 | estimated | 2.9x / 2.1x | 3.6x / 2.4x | 5.0x / 3.0x | stock class v4 30.9 MH/s at 114.8 W (3.72) measured on PC 2 (the Thunderbolt enclosure); the Blackwell shape gives about 72 W: 2.33; band 2.2 to 2.6; the card-in job measures the ladder | -| RTX 5060 (8 GB) | 8 GB | 2.15 | estimated | 2.7x / 1.9x | 3.3x / 2.2x | 4.6x / 2.7x | stock class v3 31.27 MH/s at 75.4 W (2.41) measured; class v4 stock about 110 W (3.5) and the knee about 67 W: 2.15; band 2.0 to 2.5 | -| RTX 4090 (24 GB) | 24 GB | 3.64 | estimated (stock measured, lock owed) | 4.6x / 3.3x | 5.6x / 3.8x | 7.8x / 4.6x | stock measured by floor lane 1: class v3 70.63 MH/s at 253.4 W (3.59), class v4 70.63 at 353.5 W (5.0), idle 34.7 W; the denominator sweep's RunPod 4090 read class v3 70.09 at 222.6 W (3.18, the host spread); the Ada shape (the 4070's tune: v3 x0.72, the premium x0.7) gives about 257 W at the knee: 3.64 v5; band 3.3 to 4.0; the occupancy knob 1 to 2 percent (lane 1) | -| RTX 4080 (16 GB) | 16 GB | 3.33 | estimated | 4.2x / 3.0x | 5.2x / 3.4x | 7.2x / 4.2x | stock class v3 40.71 MH/s at 128.6 W (3.16) and class v4 40.74 at 185.6 W (4.56) measured; the Ada shape gives about 133 W: 3.26 v4, 3.33 v5; the 4080 Super reads the same (42.6 at 134.9 / 195.7 W) | -| RTX 4070 (12 GB) | 12 GB | 3.51 to 3.58 | measured at the tune point | 4.5x / 3.2x | 5.6x / 3.7x | 7.7x / 4.5x | 1,860 MHz lock + the 50 percent cap: class v4 31.08 MH/s at 109.0 W (3.51), v5 3.58; against 3.58 stock on class v3 and 2.57 tuned; the ladder below 1,860 under class v4 is owed (band 3.2 to 3.6) | -| RTX 4070 Ti (12 GB) | 12 GB | 3.60 | estimated | 4.6x / 3.2x | 5.6x / 3.7x | 7.8x / 4.6x | stock class v3 31.24 MH/s at 95.4 W (3.05) and class v4 31.26 at 155.3 W (4.97) measured; the Ada shape gives about 111 W: 3.6; the 4070 Super at a 110 W host cap held 31.26 MH/s at 108.4 W under class v4 (3.47): a cap row measured | -| RTX 4060 Ti (16 GB) | 16 GB | 3.70 | estimated | 4.7x / 3.3x | 5.7x / 3.8x | 8.0x / 4.7x | stock class v3 20.07 MH/s at 79.2 W (3.95) and class v4 20.10 at 102.3 W (5.09) measured; the Ada shape gives about 73 W: 3.7; the 8 GB card reads the same on class v3 (20.10 at 77.5 W) | -| RTX 4060 (8 GB) | 8 GB | 3.77 | estimated | 4.8x / 3.4x | 5.8x / 3.9x | 8.1x / 4.8x | 19.09 MH/s measured on three rented hosts, the watts unread on every one (no power sensor exposed); the 4060 Ti's watts shape: about 95 W stock on class v4 (5.0), 72 W at the knee (3.77); band 3.7 to 4.2 | -| RTX 3090 (24 GB) | 24 GB | 6.40 | estimated | 8.1x / 5.7x | 9.9x / 6.6x | 13.8x / 8.1x | two rented 3090s in the sweep (section 6); the 0.3.12 row 37.79 MH/s at 228.8 W (6.05) and the standing voters 42 to 50 MH/s; the Ampere lever is the power cap and small (the cap knee near 1,800 MHz against a 1,900 to 2,000 stock clock: 10 to 15 percent); the 3090 Ti read class v3 61.95 at 249.5 W (4.03) | -| RTX 3080 (10 GB) | 10 GB | 5.70 | estimated | 7.2x / 5.1x | 8.8x / 5.9x | 12.3x / 7.2x | the 0.3.12 row 40.82 MH/s at 204.9 W (5.02); the sweep's rows pending (section 6); the 3080 Ti measured class v3 58.9 at 293.5 W (4.98) and, at a 220 W host cap with the SM at 749 MHz, class v4 42.6 at 216 W (5.08) with 28 percent of the rate lost: the Ampere cap's known-failed row | -| RTX 3070 (8 GB) | 8 GB | 4.80 | estimated | 6.1x / 4.3x | 7.4x / 5.0x | 10.4x / 6.1x | stock class v3 37.09 MH/s at 142.3 W (3.84) and class v4 37.11 at 196.4 W (5.29) measured; a 10 percent cap: 4.8; the 3070 Ti 38.98 at 177.1 / 267.6 W (4.54 / 6.84) | -| RTX 3060 (12 GB) | 12 GB | 4.90 | estimated | 6.2x / 4.4x | 7.6x / 5.1x | 10.6x / 6.2x | stock class v3 26.89 MH/s at 111.6 W (4.15) measured; class v4 about 5.4 at stock, 4.9 at a mild cap; the 3060 Ti at a 130 W host cap held 33.06 MH/s at 128.7 W under class v4 (3.89): a cap row measured | -| RX 9070 XT (16 GB) | 16 GB | 8.10 | measured (v4 7.9) | 10.3x / 7.3x | 12.6x / 8.4x | 17.5x / 10.3x | the ADLX grid: 18.9 MH/s at 149.3 W (-500 MHz core, -30 percent power), 24 percent under the 202 W stock point (10.7); the shadow costs this card about 3 W (199 W on class v3, 202 on v4): it is the AMD read path, not the ALU, that sets its joule | -| RX 7900 XTX (24 GB) | 24 GB | 8.00 | estimated | 10.1x / 7.2x | 12.4x / 8.3x | 17.3x / 10.2x | not rentable, not owned: by the 9070 XT's dependent-read rate per channel (150 M per second per 32-bit channel) 24 channels give about 28 MH/s at about 300 W stock (10.7), 7.5 to 8.5 with the knob | -| RX 7800 XT (16 GB) | 16 GB | 9.50 | estimated | 12.0x / 8.5x | 14.7x / 9.8x | 20.5x / 12.1x | 16 channels: about 18 MH/s at about 230 W stock (12.8); the knob 9 to 10 | -| RX 7600 XT (16 GB) | 16 GB | 13.0 | estimated | 16.5x / 11.7x | 20.2x / 13.4x | 28.0x / 16.5x | 8 channels: about 9 MH/s at about 160 W stock (18); the knob 12 to 14; the card-in queue holds one for a PC measurement | -| RX 9060 XT (16 GB) | 16 GB | 10.0 | estimated | 12.7x / 9.0x | 15.5x / 10.3x | 21.6x / 12.7x | 8 channels of GDDR6 on RDNA 4: about 9.5 MH/s at about 130 W stock (13.7); the knob 9.5 to 10.5; the card-in queue holds one | -| H100 SXM (80 GB) | DC | 2.00 | estimated (stock measured, lock owed) | 2.5x / 1.8x | 3.1x / 2.1x | 4.3x / 2.5x | stock class v3 248.9 MH/s at 411.7 W (1.65) and class v4 250.6 at 695.1 W (2.77) measured (floor lane 1 the same: v3 254.6 at 450.4 W, v4 251.9 at 699 W with the 700 W cap binding and the clock throttled to 1,750 MHz; class v5 = class v4 on this card, 250.2 at 698.8 W; its -lmc was accepted and the clock stayed 2,619 MHz); the premium 283 W (11.2 pJ per op); a lock on an owned host (Hopper at 1,980 MHz boost, HBM3-bound): about 500 W (2.0); band 1.9 to 2.4; rented hosts refuse `-lgc` | -| L40S (48 GB) | DC | 3.50 | estimated | 4.4x / 3.1x | 5.4x / 3.6x | 7.5x / 4.4x | stock class v3 56.4 MH/s at 220.7 W (3.91) and class v4 56.4 at 277.6 W (4.92) measured; the Ada shape gives about 198 W (3.5) | -| A100 SXM (80 GB) | DC | 3.10 | estimated | 3.9x / 2.8x | 4.8x / 3.2x | 6.7x / 3.9x | stock class v3 138.3 MH/s at 268.6 W (1.94) and class v4 138.0 at 489.3 W (3.54) measured (the premium 221 W = 15.8 pJ per op, the dearest ALU in the record); the card already runs at 1,410 MHz, a cap recovers about 60 W: 3.1 | -| Intel Arc B580 (12 GB) | 12 GB | 10.4 | estimated | 13.2x / 9.4x | 16.2x / 10.8x | 22.5x / 13.3x | 10.6 to 11 MH/s measured on PC 2 (the enclosure), the watts unread; about 110 W by the board's class (OWED); no clock lever in the app for Intel | -| Intel Arc A750 (8 GB) | 8 GB | 18.8 | estimated | 23.8x / 16.9x | 29.1x / 19.4x | 40.5x / 23.9x | unmeasured; Alchemist at 256-bit GDDR6: about 8 MH/s at about 150 W | -| Apple M5 Max | Apple | 1.40 to 1.43 | measured (v4 1.40) | 1.8x / 1.3x | 2.2x / 1.5x | 3.1x / 1.8x | the GPU and DRAM meter: class v3 27.08 MH/s at 21.0 W (0.78), class v4 26.67 at 37.3 W (1.40); the package about 54 W (2.0); no lever | -| Apple M4 Max | Apple | 1.55 | estimated | 2.0x / 1.4x | 2.4x / 1.6x | 3.3x / 2.0x | the same 512-bit LPDDR5X (8,533 against 9,600 MT/s) and a 40-core GPU on N3E: about 26 MH/s; class v3 about 0.85, class v5 about 1.55 at the meter (the M5 Max's rows scaled); a devnet row from an unnamed Apple laptop read 24.3 MH/s on class v4 | -| Apple M4 Pro | Apple | 1.65 | estimated | 2.1x / 1.5x | 2.6x / 1.7x | 3.6x / 2.1x | 256-bit LPDDR5X, a 20-core GPU: about 13.5 MH/s; the shadow's per-hash premium does not shrink with the part, so about 1.65 at the meter | -| Apple M3 Max | Apple | 1.75 | estimated | 2.2x / 1.6x | 2.7x / 1.8x | 3.8x / 2.2x | 512-bit LPDDR5-6400, a 40-core GPU on N3B: about 24 MH/s; about 1.75 at the meter | +| Card | Tier | Class v5 floor, microjoules per hash | Label | Class v5 stock, microjoules per hash (label) | GDDR7 board at 1.1 / 3.2 / 6.4 pJ per forced op | HBM3 one stack | N2 SRAM die | The point, the band and the source | +|---|---|---|---|---|---|---|---|---| +| NVIDIA GeForce RTX 5090 | 32 GB | 2.33 | measured | 3.48 (measured) | 4.0x / 3.0x / 2.1x | 5.4x / 3.6x / 2.4x | 9.3x / 5.0x / 3.0x | 1200 MHz at 100 percent. class v4 best MH per watt (133.80 MH/s at 305.1 W, 2.2 percent of rate given); under class v5 +2.0 percent watts: 311.2 W, 2.33 Stock: the power cap is flat on this hash (0.41 MH/W under every limit), so no power rung is used on any tier | +| NVIDIA GeForce RTX 5080 | 16 GB | 2.06 | measured | 3.48 (measured) | 3.6x / 2.6x / 1.9x | 4.8x / 3.2x / 2.1x | 8.2x / 4.4x / 2.6x | 1100 MHz at 100 percent. class v4 best MH per watt; the draw floors from 1,500 MHz down; the knee is between 1,000 and 900 MHz (900 costs 5.2 percent); class v5 about 2.10 Stock: class v5 measured on a Vast 5080 (13:49 UK): 71.35 MH/s at 248.0 W (3.48, 39 samples, SM 2,769, memory 14,801); class v3 71.16 at 159.4 W on the same host; PC 1 read class v4 71.41 at 253.1 W: the three agree within 3 percent on this card; stock, lock owed on the rented host | +| NVIDIA GeForce RTX 5070 Ti | 16 GB | 1.7 | estimated | 2.84 (measured) | 2.9x / 2.2x / 1.5x | 3.9x / 2.6x / 1.8x | 6.8x / 3.7x / 2.2x | 1100 MHz at 100 percent. the Blackwell shape applied to the measured stock rows (the v3 draw x0.63, the premium x0.5): band 1.6 to 2.0; the rented host refused -lgc, so the knee is the search's to find Stock: rented pod 8 October 10:09Z: class v3 78.69 MH/s at 140.8 W, class v4 78.78 at 224.0 W | +| NVIDIA GeForce RTX 5070 | 12 GB | 1.75 | estimated | 2.99 (measured) | 3.0x / 2.2x / 1.6x | 4.0x / 2.7x / 1.8x | 7.0x / 3.8x / 2.2x | 1100 MHz at 100 percent. the Blackwell shape on the measured stock rows (paired class v3 113 W): band 1.7 to 2.1 Stock: rented v4watts pod 8 October; the 7 October pod read class v3 52.0 MH/s at 102.8 W (the rate differs by host) | +| NVIDIA GeForce RTX 5060 Ti | 16 GB | 2.36 | estimated | 4.02 (measured) | 4.1x / 3.0x / 2.1x | 5.5x / 3.7x / 2.4x | 9.4x / 5.1x / 3.0x | 1100 MHz at 100 percent. the Blackwell shape on the PC 2 stock row: band 2.2 to 2.6; re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: class v5 measured on a Vast 5060 Ti (13:5x UK): 30.84 MH/s at 123.9 W (4.02, 96 samples, SM 2,835); class v3 30.75 at 82.6 W on the same host; PC 2's enclosure row read class v4 30.9 at 114.8 W; stock, lock owed | +| NVIDIA GeForce RTX 5060 | 8 GB | 2.21 | estimated | 3.76 (measured) | 3.8x / 2.8x / 2.0x | 5.1x / 3.4x / 2.3x | 8.8x / 4.8x / 2.8x | 1100 MHz at 100 percent. class v3 measured 31.27 MH/s at 75.4 W; re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: class v5 measured (13:51 UK): 30.80 MH/s at 115.7 W (3.76, 99 samples, SM 2,715, memory 13,801); class v3 30.76 at 78.8 W on the same host (the shadow +37 W = 12 pJ per op); stock, lock owed | +| NVIDIA GeForce RTX 4090 | 24 GB | 3.58 | estimated | 5.0 (measured) | 6.2x / 4.5x / 3.2x | 8.3x / 5.6x / 3.7x | 14.2x / 7.7x / 4.5x | 1860 MHz at 60 percent. the Ada shape (the 4070's measured tune: v3 x0.72, the premium x0.7) on the measured stock rows (class v3 253.4 W, class v4 353.5 W, floor lane 1, 8 October): band 3.3 to 4.0; re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: floor lane 1's rented 4090 at a 450 W limit: class v3 70.63 MH/s at 253.4 W (3.59), class v4 70.63 at 353.5 W (5.0); the denominator sweep's Vast 4090 carried a 250 W host limit: class v5 62.68 at 249.9 W (3.99, the SM at 2,579 MHz) and class v3 62.40 at 200.5 W, so a 250 W cap held 89 percent of the rate at 71 percent of the draw (an Ada cap row measured); stock, lock owed | +| NVIDIA GeForce RTX 4080 | 16 GB | 3.51 | estimated | 4.92 (measured) | 6.1x / 4.4x / 3.2x | 8.1x / 5.4x / 3.6x | 14.0x / 7.6x / 4.5x | 1860 MHz at 60 percent. the Ada shape on the measured stock rows (class v3 128.6 W, class v4 185.6 W): band 3.1 to 3.7; re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: class v5 measured (13:50 UK): 40.80 MH/s at 200.7 W (4.92, 72 samples, SM 2,760); class v3 40.69 at 141.5 W on the same host; the v4watts pod read class v4 40.74 at 185.6 W; stock, lock owed | +| NVIDIA GeForce RTX 4070 | 12 GB | 3.58 | measured | 5.82 (measured) | 6.2x / 4.5x / 3.2x | 8.3x / 5.6x / 3.7x | 14.2x / 7.7x / 4.5x | 1860 MHz at 50 percent. Ember run 6's tune point (1,863 MHz at 50 percent) under class v4 (sh256x27): 31.08 MH/s at 109.0 W, +30 W over class v3 at the same point for no rate; class v5 about 3.58; the ladder below 1,860 under class v4 is owed (band 3.2 to 3.6) Stock: class v5 measured on a Vast 4070 (13:54 UK): 28.43 MH/s at 165.5 W (5.82, 108 samples, SM 2,820, memory 9,801, a 210 W host limit not binding); class v3 28.40 at 116.3 W on the same host (the shadow +49 W = 17 pJ per op at the stock clock); Ember run 6's PC 1 stock on class v3 read 28.72 at 106.0 W; the tune point (1,860 MHz, 50 percent) under class v4 measured 31.08 at 109.0 W (3.51), so the lock takes 34 percent of this card's class v5 draw for no rate | +| NVIDIA GeForce RTX 4070 Ti | 12 GB | 3.6 | estimated | 4.97 (measured) | 6.2x / 4.6x / 3.2x | 8.3x / 5.6x / 3.7x | 14.3x / 7.8x / 4.6x | 1860 MHz at 60 percent. the Ada shape on the measured stock rows (class v3 95.4 W, class v4 155.3 W); the 4070 Super at a 110 W host cap held 31.26 MH/s at 108.4 W under class v4 (3.47): a cap row measured on a sibling Stock: rented pods 8 October | +| NVIDIA GeForce RTX 4060 Ti | 16 GB | 3.7 | estimated | 5.09 (measured) | 6.4x / 4.7x / 3.3x | 8.6x / 5.7x / 3.8x | 14.7x / 8.0x / 4.7x | 1860 MHz at 60 percent. the Ada shape on the measured stock rows (class v3 79.2 W, class v4 102.3 W); the 8 GB card reads the same on class v3 (20.10 MH/s at 77.5 W) Stock: rented pods 7 and 8 October | +| NVIDIA GeForce RTX 4060 | 8 GB | 3.77 | estimated | 4.98 (estimated) | 6.5x / 4.8x / 3.4x | 8.7x / 5.8x / 3.9x | 15.0x / 8.1x / 4.8x | 1860 MHz at 60 percent. 19.09 MH/s measured on three rented hosts, the watts unread on every one (no power sensor): the watts are the 4060 Ti's shape; band 3.7 to 4.2 Stock: watts OWED: every rented 4060 host exposed no power sensor | +| NVIDIA GeForce RTX 3090 | 24 GB | 4.51 | estimated | 5.03 (measured) | 7.8x / 5.7x / 4.1x | 10.4x / 7.0x / 4.7x | 17.9x / 9.7x / 5.7x | 1800 MHz at 80 percent. the Ampere cap lever (10 to 15 percent of the draw; a cap that takes the SM under 1,800 MHz costs rate) on the measured class v5 stock row; the 3090 Ti read class v3 61.95 MH/s at 249.5 W (4.03) Stock: the denominator sweep, RunPod, 8 October 13:25 UK: class v3 60.01 MH/s at 294.6 W (4.91), class v5 61.49 at 309.4 W (5.03, 46 samples, SM 1,673 MHz, memory 9,501); the class v5 premium on this card 5 percent at stock; stock, lock owed (the host refused -lgc and -lmc) | +| NVIDIA GeForce RTX 3080 | 10 GB | 4.2 | estimated | 4.54 (estimated) | 7.3x / 5.3x / 3.8x | 9.7x / 6.5x / 4.3x | 16.7x / 9.1x / 5.3x | 1800 MHz at 80 percent. the Ampere cap lever (10 percent, no lower: a 170 W cap took 34 percent of the class v5 rate on this card, a 180 W cap 11 percent) Stock: the sweep's two 3080 hosts both carried a limit under the class v5 draw: at 170 W class v3 50.71 MH/s at 169.2 W (3.34) held the rate and class v5 fell to 33.37 MH/s at 169.9 W (5.09) with the SM at 689 MHz; at 180 W class v3 49.83 at 177.8 W and class v5 44.31; the uncapped class v5 stock by the 3080 Ti's +30 percent: about 230 W (4.54); the Ampere cap under the shadow costs rate one for one: measured twice | +| NVIDIA GeForce RTX 3070 | 8 GB | 4.8 | estimated | 5.29 (measured) | 8.3x / 6.1x / 4.3x | 11.1x / 7.4x / 5.0x | 19.1x / 10.4x / 6.1x | 1800 MHz at 80 percent. the Ampere cap lever on the measured stock rows (class v3 142.3 W, class v4 196.4 W); the 3070 Ti 38.98 MH/s at 177.1 / 267.6 W Stock: rented pods 8 October | +| NVIDIA GeForce RTX 3060 | 12 GB | 5.77 | estimated | 6.4 (measured) | 10.0x / 7.3x / 5.2x | 13.3x / 8.9x / 6.0x | 23.0x / 12.4x / 7.3x | 1800 MHz at 80 percent. the Ampere cap lever (10 percent) on the measured class v4 stock row; the 3060 Ti at a 130 W host cap held 33.06 MH/s at 128.7 W under class v4 (3.89): a cap row measured on a sibling Stock: class v5 measured on a Vast 3060 (13:53 UK): 26.53 MH/s at 169.8 W (6.40) at the host's 170 W limit (the limit binding: the class v4 pod read 26.89 at 166.1 W, 6.18); class v3 26.53 at 120.4 W on the same host; stock, lock owed | +| AMD Radeon RX 9070 XT | 16 GB | 7.9 | measured | 10.7 (measured) | 13.7x / 10.0x / 7.1x | 18.3x / 12.3x / 8.2x | 31.4x / 17.0x / 10.0x | ADLX -500 MHz, -30 percent. the ADLX grid (24 rows, PC 1): the rate flat at 18.9 MH/s across the grid, the best point -500 MHz core and -30 percent power, 24 percent under stock; class v5 about 8.1 Stock: the app's own power reading at the stock point, 8 October 07:11 UK | +| AMD Radeon RX 7900 XTX | 24 GB | 8.0 | estimated | 10.7 (estimated) | 13.9x / 10.1x / 7.2x | 18.5x / 12.4x / 8.3x | 31.8x / 17.3x / 10.2x | ADLX -500 MHz, -30 percent. not rentable, not owned: the 9070 XT's dependent-read rate per channel (150 M per second per 32-bit channel) on 24 channels; the knob by the 9070 XT grid Stock: modelled | +| AMD Radeon RX 7800 XT | 16 GB | 9.5 | estimated | 12.8 (estimated) | 16.5x / 12.0x / 8.5x | 22.0x / 14.7x / 9.8x | 37.8x / 20.5x / 12.1x | ADLX -500 MHz, -30 percent. 16 channels of GDDR6 at the 9070 XT's per-channel rate; the knob by the 9070 XT grid Stock: modelled | +| AMD Radeon RX 7600 XT | 16 GB | 13.0 | estimated | 17.8 (estimated) | 22.5x / 16.5x / 11.7x | 30.1x / 20.2x / 13.4x | 51.7x / 28.0x / 16.5x | ADLX -500 MHz, -30 percent. 8 channels; the card-in queue holds one for a PC measurement Stock: modelled | +| AMD Radeon RX 9060 XT | 16 GB | 10.0 | estimated | 13.7 (estimated) | 17.3x / 12.7x / 9.0x | 23.1x / 15.5x / 10.3x | 39.8x / 21.6x / 12.7x | ADLX -500 MHz, -30 percent. 8 channels of GDDR6 on RDNA 4; the card-in queue holds one Stock: modelled | +| NVIDIA H100 80GB HBM3 | DC | 2.0 | estimated | 2.58 (measured) | 3.5x / 2.5x / 1.8x | 4.6x / 3.1x / 2.1x | 8.0x / 4.3x / 2.5x | 1400 MHz at 80 percent. the premium 240 to 280 W at stock says half of it is the clock; a lock on an owned host (Hopper at 1,980 MHz boost, HBM3-bound): about 500 W (2.0); band 1.9 to 2.4; rented hosts refuse -lgc Stock: the denominator sweep, RunPod, 13:25 UK: class v3 252.96 MH/s at 382.4 W (1.51), class v5 241.34 at 621.6 W (2.58, 7 samples, the SM throttled to 1,590 MHz under the 700 W cap); the 7 and 8 October pods read class v4 250.6 at 695.1 W (2.77); -lmc answered 'use --lock-memory-clocks-deferred' and the memory clock stayed 2,619; stock, lock owed | +| NVIDIA L40S | DC | 3.78 | estimated | 5.29 (measured) | 6.5x / 4.8x / 3.4x | 8.7x / 5.9x / 3.9x | 15.0x / 8.2x / 4.8x | 1860 MHz at 60 percent. the Ada shape on the measured stock rows (class v3 220.7 W, class v4 277.6 W); re-based on the measured class v5 stock row of 13:5x UK (the architecture's shape on the measured v3 draw and the measured premium) Stock: class v5 measured (13:49 UK): 56.49 MH/s at 298.7 W (5.29, 46 samples, SM 2,520); class v3 56.35 at 221.5 W on the same host; the v4watts pod read class v4 56.4 at 277.6 W; stock, lock owed | +| NVIDIA A100-SXM4-80GB | DC | 2.89 | estimated | 2.99 (measured) | 5.0x / 3.7x / 2.6x | 6.7x / 4.5x / 3.0x | 11.5x / 6.2x / 3.7x | 1200 MHz at 80 percent. the card already runs at 1,410 MHz and the 400 W host limit binds under class v5 (SM clock gives); a cap under it costs rate on Ampere, so the floor is within 5 percent of the measured stock row Stock: the denominator sweep, RunPod, 13:24 UK: class v3 138.06 MH/s at 294.9 W (2.14), class v5 133.11 at 398.1 W (2.99, 19 samples, at the host's 400 W limit, memory 1,593 MHz); the v4watts pod on a 500 W host read class v4 138.0 at 489.3 W (3.54); -lmc 'not supported' on this card; stock, lock owed | +| Intel Arc B580 | 12 GB | 10.4 | estimated | 10.4 (estimated) | 18.0x / 13.2x / 9.3x | 24.1x / 16.1x / 10.7x | 41.4x / 22.4x / 13.2x | stock, no lever. 10.6 to 11 MH/s measured on PC 2 (the enclosure), the watts unread; about 110 W by the board's class (OWED); no clock lever in the app for Intel | +| Intel Arc A750 | 8 GB | 18.8 | estimated | 18.8 (estimated) | 32.6x / 23.8x / 16.9x | 43.5x / 29.2x / 19.4x | 74.8x / 40.5x / 23.9x | stock, no lever. unmeasured; modelled from the B580 and the Alchemist board | +| Apple M5 Max | Apple | 1.4 | measured | 1.4 (measured) | 2.4x / 1.8x / 1.3x | 3.2x / 2.2x / 1.4x | 5.6x / 3.0x / 1.8x | stock, no lever. the GPU and DRAM channels of the IOReport meter (not the wall): class v3 27.08 MH/s at 21.0 W (0.78), class v4 26.67 at 37.3 W (1.40), class v5 about 1.43; the package about 17 W more; no lever (no clock cap on Apple silicon) | +| Apple M4 Max | Apple | 1.55 | estimated | 1.55 (estimated) | 2.7x / 2.0x / 1.4x | 3.6x / 2.4x / 1.6x | 6.2x / 3.3x / 2.0x | stock, no lever. the same 512-bit LPDDR5X (8,533 against 9,600 MT/s) and a 40-core GPU on N3E: about 26 MH/s; class v3 about 0.85 at the meter, class v5 about 1.55 (the M5 Max's rows scaled); a devnet row from an unnamed Apple laptop read 24.3 MH/s on class v4 | +| Apple M4 Pro | Apple | 1.65 | estimated | 1.65 (estimated) | 2.9x / 2.1x / 1.5x | 3.8x / 2.6x / 1.7x | 6.6x / 3.6x / 2.1x | stock, no lever. 256-bit LPDDR5X and a 20-core GPU: about 13.5 MH/s; the shadow's per-hash premium does not shrink with the part, so class v5 reads about 1.65 at the meter | +| Apple M3 Max | Apple | 1.75 | estimated | 1.75 (estimated) | 3.0x / 2.2x / 1.6x | 4.0x / 2.7x / 1.8x | 7.0x / 3.8x / 2.2x | stock, no lever. 512-bit LPDDR5-6400 and a 40-core GPU on N3B: about 24 MH/s; class v5 about 1.75 at the meter | To fold any other chip core figure (the research lane's convention, 13:5x): the edge on a row is the card's class v5 microjoules over `E_mem + 101,170 x c`, with `E_mem` 0.466 (GDDR7 board), 0.321 (one HBM3 stack), 0.14 (the N2 SRAM @@ -264,26 +267,51 @@ The re-measure rule after a class flip, as the table carries it (`remeasure_rule 3. The set also re-measures weekly (`PERIOD_S`) and on a driver major change (the prior key carries driver major and class), and a confirm check that beats its prior by over 1 percent on MH per watt asks for the full plan. -## 6. The denominator sweep (rented, 8 October 13:05 onward, the stock rows the record lacked) +## 6. The denominator sweep (rented, 8 October 13:05 to 14:1x UK, the stock rows the record lacked) -Three drivers of the fleet lane's model sweep (the third, from 13:21, on the coordinator's full-spend order: 16 card classes in parallel on `box-cardbench-v2m.sh`, the class v5 kit's bench run for 200 batches of 2^24 under the sampler so the class v5 watts are a measured mean, and a memory-clock try (`-lmc` to the card's maximum, a second class v5 run when the host accepts it)) (`box-cardbench-v2.sh` for class v3, the v5 kit's fingerprint and the -memprobe ceiling; `box-cardbench-v4watts.sh` for the class v4 shape with the paired class v3 run), one one-shot pod per -card class on RunPod or Vast under the USD 1.00 per hour consumer cap, destroyed at the end of each row; the rows -appended to `~/igneum-fleet/cardbench/rows.jsonl` with `sweep: 2026-10-08-denominator`; spend inside the lane's USD 100. -No rented host allowed `-lgc`, so these are stock rows and the knees above stay estimated. The rows are read into the -table of section 2 as they land; this section lists them. +Three drivers of the fleet lane's model sweep and then a direct ssh runner: `box-cardbench-v2.sh` (class v3, the v5 +kit's fingerprint, the memprobe ceiling), `box-cardbench-v4watts.sh` (the class v4 shape with the paired class v3 run) +and, from 13:21 on the coordinator's full-spend order, `box-cardbench-v2m.sh` on 16 card classes in parallel (the +class v5 kit's bench run for 200 batches of 2^24 under the 1 Hz sampler, so the class v5 watts are a measured mean +over busy samples from 8 s in; then `-lmc` to the card's maximum and a second class v5 run when the host accepted it). +One one-shot pod per card class on RunPod or Vast under the consumer cap, destroyed at the end of each row; the rows +in `~/igneum-fleet/cardbench/rows.jsonl` under `sweep: 2026-10-08-denominator` and `-v5`; the lane's spend about USD 5 +of its USD 100 (the pods ran 3 to 10 minutes each). No rented host allowed `-lgc`; the two that accepted `-lmc` (the +H100 and A100) left the memory clock where it was, so every row here is stock and the knees of section 2 stay +estimated where no founder machine holds the card. -| Card | Class v3 (MH/s at W, microjoules) | Class v4 shape (MH/s at W, microjoules) | v4 over v3 watts | Host, driver, cost | Reading | +Two faults found on the way, for the fleet lane: (1) the sweep driver's ssh wait read every Vast pod of the 13:21 wave +as "never answered ssh in 10 minutes" while a direct `ssh` to the same host and port answered at once (13 pods, all +live; the rows were then taken by a direct runner, `~/igneum-fleet/denom-direct.sh`, which keys on the provider's +own ssh host and port and nothing in the registry); (2) the registry lock `~/igneum-fleet/boxes.json.lock` was held +from 13:32 by three `dn3-q05.py` processes of another lane, so every `Registry.patch` from this lane blocked (a +re-rent sat 12 minutes on a live pod and then destroyed it); the lane's scripts now write the registry best-effort +under a 15 s alarm. The rented stock rows against the PC rows on the one card measured both ways (the 5080): class v5 +71.35 MH/s at 248.0 W rented against class v4 71.41 at 253.1 W on PC 1, within 3 percent, so the 18 percent class v3 +spread of the morning was that host and not the method. + +| Card | Host | Class v3 (MH/s at W = microjoules) | Class v4 shape or class v5 (MH/s at W = microjoules; samples) | The knobs on the host | Cost | |---|---|---|---|---|---| -| RTX 4090 24 GB | 70.09 MH/s at 222.6 W (3.18; RunPod, driver 595.91, 13:07); class v5 kit 70.25 MATCH | pending | | USD 0.74 per hour, 0.03 | floor lane 1's rented 4090 read 70.63 at 253.4 W on the same class: the host spread of 12 percent on the watts at the same rate | -| RTX 5060 8 GB | 24.77 MH/s at 80.7 W (3.26; Vast, 13:11); class v5 kit 31.31 MATCH (a rate the v3 run on this host did not reach) | 16.23 MH/s at 87.8 W (5.41; a second Vast host, 13:10, v4 over v3 +20 percent on that host) | +20 percent | USD 0.02 | the rate varies by host 16 to 31 MH/s on this card class: the 7 October pod read 31.27 on class v3; the per-hash energy at the matched rate stands at 2.4 (v3) | -| RTX 3090 24 GB | pending (two Vast hosts rented 13:05) | pending | | USD 0.14 per hour | | -| RTX 3080 10 GB | pending | pending | | | the first Vast offer vanished at the rent (HTTP 410) | -| RTX 5080 16 GB | the RunPod host failed `cuInit` (a bad host, destroyed); Vast next | pending | | | | -| RTX 5060 Ti 16 GB | pending | pending | | | | -| RTX 5060 8 GB | pending | pending | | | | -| RTX 3060 12 GB | pending | pending | | | | -| RTX 4070 12 GB | pending | pending | | | | +| RTX 4090 24 GB | runpod m1tstl1h1b7r0o (595.91.07), limit 450.0 W | 70.094 at 222.6 W = 3.18 | class v5 70.253 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.032 | +| RTX 5060 8 GB | vast 54840005 (580.126.09), limit 145.0 W | | class v4 16.234 at 87.8 W = 5.41 (v4 over v3 20.4 percent on this host) | | USD 0.009 | +| RTX 5060 8 GB | vast 54840004 (580.126.09), limit 145.0 W | 24.774 at 80.7 W = 3.26 | class v5 31.312 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.011 | +| RTX 5080 16 GB | vast 54840017 (570.133.07), limit 300.0 W | | class v4 70.951 at 239.8 W = 3.38 (v4 over v3 41.1 percent on this host) | | USD 0.052 | +| RTX 3080 10 GB | vast 54840023 (580.159.03), limit 130.0 W | | class v4 24.007 at 129.2 W = 5.38 (v4 over v3 3.0 percent on this host) | | USD 0.025 | +| RTX 3060 12 GB | vast 54840846 (580.126.20), limit 170.0 W | | class v4 26.886 at 166.1 W = 6.18 (v4 over v3 44.6 percent on this host) | | USD 0.007 | +| RTX 4070 12 GB | vast 54840856 (580.173.02), limit 200.0 W | 30.829 at 107.1 W = 3.47 | class v5 30.878 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.023 | +| A100 80 GB | runpod miuoe6gkcb1rhz (580.126.16), limit 400.0 W | 138.062 at 294.9 W = 2.14 | class v5 133.548 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.075 | +| H100 80 GB | runpod zosqj1s802dk3x (580.126.09), limit 700.0 W | 252.959 at 382.4 W = 1.51 | class v5 241.342 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.216 | +| RTX 3090 24 GB | runpod sbrhhunqui25d0 (580.65.06), limit 310.0 W | 60.008 at 294.6 W = 4.91 | class v5 61.486 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.014 | +| RTX 3080 10 GB | vast 54841582 (580.159.03), limit 180.0 W | 49.832 at 177.8 W = 3.57 | class v5 44.308 MH/s (MATCH), watts not sampled on this row | -lgc refused | USD 0.024 | +| RTX 5080 16 GB | vast 54845199 (570.133.07), limit 300.0 W | 71.16 at 159.4 W = 2.24 | class v5 71.349 at 248.0 W = 3.476; 39 samples, SM 2769, memory 14801 | -lgc refused; -lmc refused | USD 0.014 | +| L40S 48 GB | vast 54845608 (595.71.05), limit 350.0 W | 56.354 at 221.5 W = 3.93 | class v5 56.493 at 298.7 W = 5.287; 46 samples, SM 2520, memory 9001 | -lgc refused; -lmc refused | USD 0.04 | +| RTX 4080 16 GB | vast 54845190 (595.84), limit 320.0 W | 40.694 at 141.5 W = 3.48 | class v5 40.8 at 200.7 W = 4.919; 72 samples, SM 2760, memory 10801 | -lgc refused; -lmc refused | USD 0.017 | +| RTX 3080 10 GB | vast 54845161 (580.178.04), limit 170.0 W | 50.708 at 169.2 W = 3.34 | class v5 33.37 at 169.9 W = 5.091; 90 samples, SM 689, memory 9251 | -lgc refused; -lmc refused | USD 0.008 | +| RTX 4090 24 GB | vast 54845886 (595.84), limit 250.0 W | 62.401 at 200.5 W = 3.21 | class v5 62.684 at 249.9 W = 3.987; 46 samples, SM 2579, memory 10251 | -lgc refused; -lmc refused | USD 0.02 | +| RTX 5060 8 GB | vast 54845177 (595.84), limit 140.0 W | 30.76 at 78.8 W = 2.56 | class v5 30.804 at 115.7 W = 3.756; 99 samples, SM 2715, memory 13801 | -lgc refused; -lmc refused | USD 0.011 | +| RTX 5060 Ti 16 GB | vast 54845162 (580.126.09), limit 180.0 W | 30.747 at 82.6 W = 2.69 | class v5 30.843 at 123.9 W = 4.017; 96 samples, SM 2835, memory 13801 | -lgc refused; -lmc refused | USD 0.013 | +| RTX 3060 12 GB | vast 54845154 (580.126.09), limit 170.0 W | 26.534 at 120.4 W = 4.54 | class v5 26.53 at 169.8 W = 6.4; 117 samples, SM 1806, memory 7301 | -lgc refused; -lmc refused | USD 0.008 | +| RTX 4070 12 GB | vast 54845860 (580.126.09), limit 210.0 W | 28.401 at 116.3 W = 4.09 | class v5 28.433 at 165.5 W = 5.821; 108 samples, SM 2820, memory 9801 | -lgc refused; -lmc refused | USD 0.011 | ## 7. Consequences per tier