From 9629fd8dbc4990d0a7f5539169a354e71d954973 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 12:24:39 +0000 Subject: [PATCH] Class v6 10.4: the honest denominator per tier (floor lane 4: the 16 GB Blackwell card at its knee is the honest NVIDIA floor, the 5080 2.06 measured; the only lever the operating point; the Ember tier table); the close table gains the 5080 row Co-Authored-By: Claude Fable 5.1 --- docs/design/class-v6-rotating-family.md | 26 +++++++++++++++++++++++-- 1 file changed, 24 insertions(+), 2 deletions(-) diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index d0d0f8bef..fc9f7c0dd 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -283,7 +283,7 @@ The question the founder set: the stored-dataset chip pays the DRAM's own energy | SM-sparse (`class-v6-floor-sm`) | `E_card` above `E_mem`: the honest card's watts toward the DRAM's own at the activate ceiling, on rented 5090, 4090 and H100 and PC 1 through the hash lane's jobs, with a worker patch | the research file's 20.3b (PC 1, 7 October night, the fourth exe): a quarter of the SMs holds 98.2 percent of the class v4 rate at the SAME draw (460 against 451 W) and 99.8 percent of class v3 at 4 W less; watts minus idle per MH/s never falls below base; the draw follows the work, not the SM count; at the 1,300 lock the sparse shapes collapse. The candidate was closed on that row | the 20.3b reading stands: the honest card's premium-free floor is its idle plus its memory system plus whatever the SMs spend waiting (99 W of 228 at the lock, measured as the residual, not reached by idling SMs); the 5090 at its knee 1.66 microjoules against the GDDR7 chip's 0.466: 3.6x | a measured breakdown of the 99 W that an idle SM does not save (clock tree, L2, fabric), and whether a different occupancy shape (fewer warps per SM at full SM count) moves it | | The shadow's k (`class-v6-floor-k`) | `k` from RTL synthesis (Yosys plus OpenROAD on build-4) replacing every claimed chip-side k; the op mix that maximises k | the GPU side measured (the research file's 15.1a): ARX 11.3 / 6.2 pJ per counted op, mul 13.9 / 8.3, mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, fp32 FMA 9.2 / 5.2, the int8 tile 1.4 to 4.1 per MAC; the chip side was claimed only (a 5 nm SIMD array 2 to 5 pJ per op, approximate) until the lane's first rows (10.2, 13:3x UK): synthesised on ASAP7, k 0.18 at the lock in the absolute convention at N3 (0.35 unscaled), the ARX families 0.17 to 0.18, mad 0.20, mulhi 0.03 | the default stays the record's rows at k 0.5 (the DRAM chip 2.9x at the knee) with the unit floors as the lower bound: the GDDR7 chip with the class v4 shadow at 4.0x at the lock and 5.8x at stock (absolute, floor k) is the worst case the served line must survive, not the reading; the sequencer-core row with the M5 Max column is the headline when it lands; the op mix held at class v4's (15.1b: the shuffle at 4.9x the add on the GPU side; mulhi the worst lever on the card by 6x) | the synthesised pJ per op per family for a 32-lane SIMD core at the chosen node, the resulting k per family, and the mix that maximises k at a fixed GPU premium | | The SRAM full-store die (`class-v6-floor-sram`) | `E_mem` for the chip lane B named strongest: the wire-energy lever, the dataset schedule that keeps the store above USD 5,000 of silicon through 2031, the capex wall recomputed with rotation | lane B's rows (7a, chip-model 5.12): 2 GiB on one N2 reticle, 1.0 nJ per read (0.5 to 2.0), 17x at zero shadow, 2.7x to 4.8x with it, USD 400 to 600 per die; lane A's schedule (3.4): 5.5 / 8 / 11 GiB retiring a quarter of today's consumer cards per step; the M5 Max's measured rate cost of the larger working set (3.3) | lane B's and lane A's rows stand; the schedule decision is the founder's with the per-tier table (0.1) | the schedule in GiB per year that keeps the store at USD 5,000 or more of silicon through 2031 on the SRAM cost curve, and what it costs each tier by year | -| The honest denominator (`class-v6-floor-denominator`) | `E_card` per tier with every software knob (the core lock, the undervolt, the memory clock, the occupancy, the block shape); the Ember tier table | the knee rows (the research file's 20.3a and 20.3b; the cost rows): the 5090 at 1,300 MHz 1.66 to 1.69 microjoules, the 4070 at its tune 2.57, the 5070 Ti stock 1.79, the M5 Max 0.78 (GPU and DRAM channels); the operating point is rank 1 of the research file | those rows stand as the per-tier floor; the Ember knob ships as ordered for 0.3.24 | the per-tier best point with every knob, the Ember table, and the AMD and Apple lines (ADLX or nothing; no lever) | +| The honest denominator (`class-v6-floor-denominator`) | `E_card` per tier with every software knob (the core lock, the undervolt, the memory clock, the occupancy, the block shape); the Ember tier table; FIRST TABLE IN (10.4, 13:4x UK): the honest NVIDIA floor is the 16 GB Blackwell card at its knee (the 5080 2.06 measured), the only lever the operating point | the knee rows (the research file's 20.3a and 20.3b; the cost rows): the 5090 at 1,300 MHz 1.66 to 1.69 microjoules, the 4070 at its tune 2.57, the 5070 Ti stock 1.79, the M5 Max 0.78 (GPU and DRAM channels); the operating point is rank 1 of the research file | those rows stand as the per-tier floor; the Ember knob ships as ordered for 0.3.24 | the per-tier best point with every knob, the Ember table, and the AMD and Apple lines (ADLX or nothing; no lever) | | Invention beyond the four (`class-v6-floor-invention`) | the tensor core's k re-read, the RT core, controller plus PHY cost, proof of latency, era-driven address mapping, the literature since 2023 | the research file: the tensor tile's measured 1.4 to 4.1 pJ per MAC against a 5 nm array's claimed 0.04 to 0.4 (k 0.03 to 0.3, the worse lever); the L2 hit at 1.4 to 2.4 nJ against a chip's SRAM (k 0.1 to 0.3); the texture interpolator excluded (not bit-exact across vendors); the literature table (sections 5 and 8) | those readings stand; no new lever | anything that reads k above 1 on measured rows, which nothing has | The close (20:00): the honest floor per tier against each chip row, the recommended class v6 changes (the op mix, the SM-sparse default, the dataset schedule), the served chip line as a measured range, and what stays for v7, written here from the lanes' rows as they land. @@ -297,10 +297,11 @@ The table, filled with the defaults, each cell replaced as its lane's row lands | RTX 5090 stock (3.36; class v3 2.26) | 3.3x default / 5.8x floor | 3.9x / 7.8x | 5.6x / 21x | no change, MEASURED on three more cards (10.1, 13:25 UK): the honest floor at stock is the base row within 2 percent on the 4090 (best shape saves 4 W), the H100 (1.4 percent) and the 5090; the residual is the clock domain, which only the lock takes off | not applied (default: outcome A, the card -47 percent, dead; the rented-5090 row owed by 14:45) | | RTX 5090 at the 1,300 MHz lock, the record's operating point (2.33 = 1.67 + 0.65) | 2.9x / 4.0x | 3.6x / 5.4x | 3.9x absolute (6.1x in the record's convention, marked) / 14x | no change (the lock pass on PC 1 running; the knee row by 15:00) | not applied (same default) | | Apple M5 Max (1.40; class v3 0.78; the GPU and DRAM channels) | 1.8x / 2.4x | 2.2x / 3.2x | 3.9x / 8.7x | not applicable (no SM lever on Apple) | not applicable (the M5 Max pays 0 at 64 bytes, measured; the chip +33 percent) | +| RTX 5080 at its 1,100 MHz lock, the honest NVIDIA floor (2.06 class v4 measured: 71.20 MH/s at 146.6 W; class v5 2.10; floor lane 4, 10.4) | 2.6x / 3.6x | 3.0x / 4.8x | 3.6x / 13x | no change | not applied | | RTX 4090 stock (5.005; class v3 3.588; rented pod, 10.1) | 4.3x / 8.7x | 4.9x / 11.6x | 6.6x / 31x | no change (measured: 16 of 128 SMs saves 4 W) | not applied | | H100 HBM3 at its 700 W cap (2.776, throttled to 1,750 MHz; class v3 1.769; rented pod, 10.1) | 2.9x / 4.8x | 3.4x / 6.4x | 5.0x / 17x | no change (measured: per-SM throughput-bound, 1.4 percent) | not applied | -The honest floor in one line, on the defaults: against the chip anyone can build (the GDDR7 board) the strongest honest tier by joule, the M5 Max, holds the chip to under 2x and a 5090 at its knee to about 3x; against the chip that needs an N2 project (the SRAM die) the holds are about 4x and 4x; every cell's floor-k figure is the worst case if a chip's core costs no more than its bare units. +The honest floor in one line, on the defaults: against the chip anyone can build (the GDDR7 board) the strongest honest tier by joule, the M5 Max, holds the chip to under 2x, the 16 GB Blackwell card at its knee (the 5080 measured, the 5070 Ti and 5070 modelled; floor lane 4's finding that the honest NVIDIA floor is not the 5090) to about 2.6x and a 5090 at its knee to about 3x; against the chip that needs an N2 project (the SRAM die) the holds are about 4x and 4x; every cell's floor-k figure is the worst case if a chip's core costs no more than its bare units. The capex wall's two thresholds (lane 3, 10.3, re-folded 13:18 UK on lane 5's corrected project floor of USD 20 to 75 M for the cheapest DRAM-board chip): no rational chip project of any kind below about USD 23 M a year of miner revenue (IGN 0.03 at launch emission, USD 62 K a day, a cap of about USD 60 M); every DRAM-board project at a third of the network above about USD 340 M a year (IGN 0.44, USD 0.93 M a day, a cap of about USD 0.9 B), where the SRAM project also starts. The project cost moves the threshold 5x, the chip's edge 1.4x. @@ -416,6 +417,27 @@ W = 16 is the one width that moves the die more than 2x at zero shadow and it co (3) The shadowed SRAM rows on the k lane's synthesised core (0.11 microjoules of shadow at N3, absolute k about 0.1, the shuffle open; the lane's 2.2b): under class v4 21x at W = 4 and 18x at W = 8 at stock, 14x and 12x at the knee; at the honest cards' whole latency shadow 10x and 9.7x at stock, 6.5x and 6.1x at the knee; an N2 core about 1.2x more. On the synthesised core the shadow is not the whole hold (the record's claimed band read 2x to 6x, marked beside): the memory is a third to a half of the die's energy, the width is worth 1.3x with the shadow on, and the lane's line for the served sentence is 6x to 10x at the full shadow (2x to 4x on the claimed band, marked), never under 2x. These are the per-unit-floor figures of 10.2; the k lane's sequencer-core row re-folds them. +### 10.4 The honest denominator per tier (floor lane 4, `docs/analysis/class-v6/floor/denominator.md` on class-v6-floor-denominator, first table 13:4x UK; the rented sweep by 15:00, the final table by 15:15; the Ember tier table `app/igneum-app/tiers/class-v5-tiers.json`, 30 card classes, 5 measured, test 10 of 10 on build-3) + +**The honest NVIDIA floor is the 16 GB Blackwell card at its knee, not the 5090.** Class v5 is class v4's shape plus the state leaves (0.0 percent of rate, +2.0 percent of watts at the 5090's knee, measured, the v5lock job). The chip columns are the chip's class v5 energy at its shadow priced in absolute picojoules per forced op (3.2 pJ, the record's k 0.5 at the 5090's knee; 6.4 pJ, k 1): GDDR7 0.79 / 1.11, HBM3 one stack 0.65 / 0.97, N2 SRAM 0.46 / 0.79 microjoules per hash (the SRAM column at lane B's 64-byte row; at the genesis width, 10.3, the die's memory term is 0.051 and the column reads higher). + +| Tier | Card, the point | Class v5 microjoules per hash | Label | GDDR7 k 0.5 / k 1 | HBM3 | SRAM | +|---|---|---|---|---|---|---| +| 32 GB | 5090, 1,300 MHz lock (Balanced) | 2.37 (v4 2.32 measured at 134.76 MH/s, 312.5 W) | measured | 3.0x / 2.1x | 3.6x / 2.4x | 5.0x / 3.0x | +| 32 GB | 5090, 1,200 MHz (Efficiency, rate -2.2 percent) | 2.33 | measured | 3.0x / 2.1x | 3.6x / 2.4x | 5.0x / 3.0x | +| 16 GB | 5080, 1,100 MHz lock | 2.10 (v4 2.06 measured: 71.20 MH/s at 146.6 W) | measured | 2.7x / 1.9x | 3.3x / 2.2x | 4.5x / 2.7x | +| 16 GB | 5070 Ti at its knee | 1.70 (stock v4 2.84 measured; the Blackwell shape) | modelled | 2.2x / 1.5x | 2.6x / 1.8x | 3.7x / 2.2x | +| 12 GB | 5070 at its knee | 1.75 (stock v4 2.99 measured) | modelled | 2.2x / 1.6x | 2.7x / 1.8x | 3.8x / 2.2x | +| 12 GB | 4070 at its tune (1,860 MHz, 50 percent cap) | 3.58 (v4 3.51 measured: 31.08 at 109.0 W) | measured | 4.5x / 3.2x | 5.6x / 3.7x | 7.7x / 4.5x | +| 24 GB | 4090 at an Ada knee | 3.64 (stock v4 5.0 measured by lane 1; the 4070's shape) | modelled, stock measured | 4.6x / 3.3x | 5.6x / 3.8x | 7.8x / 4.6x | +| 24 GB | 3090 at the Ampere cap | 6.4 (the cap is the only lever, 10 to 15 percent) | modelled | 8.1x / 5.7x | 9.9x / 6.6x | 13.8x / 8.1x | +| 16 GB AMD | 9070 XT, ADLX -500 MHz / -30 percent | 8.1 (v4 7.9 measured: 18.9 MH/s at 149.3 W) | measured | 10.3x / 7.3x | 12.6x / 8.4x | 17.5x / 10.3x | +| DC | H100 at a lock (an owned host) | 2.0 (stock v4 2.77 measured) | modelled | 2.5x / 1.8x | 3.1x / 2.1x | 4.3x / 2.5x | +| Apple | M5 Max, no lever (the GPU and DRAM meter) | 1.43 (v4 1.40 measured) | measured | 1.8x / 1.3x | 2.2x / 1.5x | 3.1x / 1.8x | +| Apple LPDDR6, five years out | a Max-class SoC on JESD209-6 | 1.18 at the meter (0.27 DRAM, 0.35 GPU, 0.56 shadow) | modelled | 1.5x / 1.06x | 1.8x / 1.2x | 2.5x / 1.5x | + +What it means: against the 16 GB Blackwell card at its knee (2.1 measured on the 5080; 1.7 to 2.1 modelled on the 5070 Ti and 5070) the GDDR7 chip is under 2x at k 1 and the SRAM die 2.7x; against the M5 Max 1.3x and 1.8x; against the LPDDR6 Apple part five years out the GDDR7 chip is at parity at k 1 and the SRAM die 1.5x. The only software lever is the operating point (the lock on Blackwell and Ada, the cap on Ampere, the ADLX offsets on AMD, nothing on Apple or Intel); the occupancy is worth 1 to 2 percent (10.1), the memory clock untouched, the block shape and the cache hint closed. The lane's per-tier scoring rule: size the shadow by the Apple tier's 5 percent point as the 2.0 rule already does and keep N at 100,000 (sizing to the Apple ceiling of 130,000 costs every honest miner 8 to 15 percent of electricity for about 0.15x of edge: the 5090's GDDR7 edge at k 1 goes from 2.13x to 1.95x); do not score the acceptance floor per tier (one network-wide number). The re-measure rule after a class flip: a stored set is stale under a new class, the stored point is the provisional start of the re-measure, max is stock and never stale; the measured v4 to v5 flip moved the 5090's knee by nothing. Owed and labelled: every Ada and Ampere knee is modelled because no rented host allows -lgc (14 pods today, every one refused); the rented stock rows are measured; the sweep adds class v5 watts measured on 16 card classes and a memory-clock try on each. + ### 10.5 Invention beyond the four layers (floor lane 5, `docs/analysis/class-v6/floor/invention.md` on class-v6-floor-invention, first KEEP/KILL table 13:4x UK; full by 19:30 UK) **One lever found, conditional on one measurement: the second DRAM sector.** The identity's only `k = 1` work is the memory's own, and the record forced one 32-byte sector per read while the dataset item is 64 bytes. Reading the whole item (W = 16 words; the read-width branch's class w64, fingerprint 836e56e7d496e980, the verifier 0.630 against 0.604 ms) forces a second sector in the same open row: on the chip model's own inputs (909 pJ activate plus 1,150 pJ of movement per sector) the GDDR7 chip's read goes from 2.0 to 3.2 nJ and E_mem from 0.466 to 0.62 microjoules (+33 percent at an unchanged 166 MH/s), one HBM3 stack 0.321 to 0.40 (+24 percent), and on floor lane 3's wire figure the N2 SRAM die 0.036 to 0.126 microjoules (66x to 19x). So the band {1, 4} words is inside one sector and "the chip pays nothing" is right to 32 bytes and wrong at 64; this is the measured reason W = 8 is free for the chip too (section 10.3) and W = 16 is the only width that costs it. The card is the whole question: the 5090 ran w64 at 71.9 MH/s (-47 percent; the exported kernel at full occupancy, four uint4 loads per item, no request hint) while its own 64-byte probe reached 15.7 G reads a second at its best lane count (-13 percent); 589 GB/s is a third of the stream, so the bind is the request path, not the pins. The 9070 XT (-3 percent) and the M5 Max (0) are free at 64 bytes (measured).