diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index db2c3fbc5..a89dd3039 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -258,6 +258,20 @@ The wording this lane proposes for main's word, if the SRAM-store reading stands Rejected by the invention lane with the number (its section 2): per-lane data-dependent branches, reads tied to the shard proof per block, randomised memory topology, the VRAM ratchet as a lever (layer 2's floor stands as written), pool-sampled witnesses, prover-gated eligibility (1.6 kW of proving network-wide at any hash rate: 0.7 percent of the hash at 100 GH/s), time-locked commitments beyond the era VDF, derived reads in the hash, row-straddling reads, scratch draws, the refresh per block. No candidate moves the per-joule identity. +## 10. The floor: the stored-dataset chip's per-joule floor and everything behind it (the five floor lanes of 12:55 UK; their rows by 16:00 to 16:30 UK, their documents under `docs/analysis/class-v6/floor/` by 19:30; this section closes at 20:00 with each lane's row or its default stated and the row owed) + +The question the founder set: the stored-dataset chip pays the DRAM's own energy per random read and nothing else; the honest card pays that plus everything its silicon spends around the read; the floor is the ratio, and every term in it is a lever. The identity (the research file's section 2): at zero shadow `edge = E_card / E_mem`; with the shadow `(E_card + F) / (E_mem + k F)`. The lanes each take one term. + +| Lane (branch) | The term | What this document already holds (measured) | The default if the lane's rows are not in by 19:30 | The row owed | +|---|---|---|---|---| +| SM-sparse (`class-v6-floor-sm`) | `E_card` above `E_mem`: the honest card's watts toward the DRAM's own at the activate ceiling, on rented 5090, 4090 and H100 and PC 1 through the hash lane's jobs, with a worker patch | the research file's 20.3b (PC 1, 7 October night, the fourth exe): a quarter of the SMs holds 98.2 percent of the class v4 rate at the SAME draw (460 against 451 W) and 99.8 percent of class v3 at 4 W less; watts minus idle per MH/s never falls below base; the draw follows the work, not the SM count; at the 1,300 lock the sparse shapes collapse. The candidate was closed on that row | the 20.3b reading stands: the honest card's premium-free floor is its idle plus its memory system plus whatever the SMs spend waiting (99 W of 228 at the lock, measured as the residual, not reached by idling SMs); the 5090 at its knee 1.66 microjoules against the GDDR7 chip's 0.466: 3.6x | a measured breakdown of the 99 W that an idle SM does not save (clock tree, L2, fabric), and whether a different occupancy shape (fewer warps per SM at full SM count) moves it | +| The shadow's k (`class-v6-floor-k`) | `k` from RTL synthesis (Yosys plus OpenROAD on build-4) replacing every claimed chip-side k; the op mix that maximises k | the GPU side measured (the research file's 15.1a): ARX 11.3 / 6.2 pJ per counted op, mul 13.9 / 8.3, mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, fp32 FMA 9.2 / 5.2, the int8 tile 1.4 to 4.1 per MAC; the chip side claimed only (a 5 nm SIMD array 2 to 5 pJ per op, approximate) | the served band stands: k 0.3 to 0.8 for the ALU shadow (2.1x at k = 1, 3.4x at k about 0.33); the op mix held at class v4's (15.1b: the shuffle at 4.9x the add on the GPU side) | the synthesised pJ per op per family for a 32-lane SIMD core at the chosen node, the resulting k per family, and the mix that maximises k at a fixed GPU premium | +| The SRAM full-store die (`class-v6-floor-sram`) | `E_mem` for the chip lane B named strongest: the wire-energy lever, the dataset schedule that keeps the store above USD 5,000 of silicon through 2031, the capex wall recomputed with rotation | lane B's rows (7a, chip-model 5.12): 2 GiB on one N2 reticle, 1.0 nJ per read (0.5 to 2.0), 17x at zero shadow, 2.7x to 4.8x with it, USD 400 to 600 per die; lane A's schedule (3.4): 5.5 / 8 / 11 GiB retiring a quarter of today's consumer cards per step; the M5 Max's measured rate cost of the larger working set (3.3) | lane B's and lane A's rows stand; the schedule decision is the founder's with the per-tier table (0.1) | the schedule in GiB per year that keeps the store at USD 5,000 or more of silicon through 2031 on the SRAM cost curve, and what it costs each tier by year | +| The honest denominator (`class-v6-floor-denominator`) | `E_card` per tier with every software knob (the core lock, the undervolt, the memory clock, the occupancy, the block shape); the Ember tier table | the knee rows (the research file's 20.3a and 20.3b; the cost rows): the 5090 at 1,300 MHz 1.66 to 1.69 microjoules, the 4070 at its tune 2.57, the 5070 Ti stock 1.79, the M5 Max 0.78 (GPU and DRAM channels); the operating point is rank 1 of the research file | those rows stand as the per-tier floor; the Ember knob ships as ordered for 0.3.24 | the per-tier best point with every knob, the Ember table, and the AMD and Apple lines (ADLX or nothing; no lever) | +| Invention beyond the four (`class-v6-floor-invention`) | the tensor core's k re-read, the RT core, controller plus PHY cost, proof of latency, era-driven address mapping, the literature since 2023 | the research file: the tensor tile's measured 1.4 to 4.1 pJ per MAC against a 5 nm array's claimed 0.04 to 0.4 (k 0.03 to 0.3, the worse lever); the L2 hit at 1.4 to 2.4 nJ against a chip's SRAM (k 0.1 to 0.3); the texture interpolator excluded (not bit-exact across vendors); the literature table (sections 5 and 8) | those readings stand; no new lever | anything that reads k above 1 on measured rows, which nothing has | + +The close (20:00): the honest floor per tier against each chip row, the recommended class v6 changes (the op mix, the SM-sparse default, the dataset schedule), the served chip line as a measured range, and what stays for v7, written here from the lanes' rows as they land. + ## 8. Unverified and owed - Lane D's family harness (the family gate's measured coverage: the 10,000-era random stratum and the three corner strata, the lossy cap, width 4, shape 64) lands in `family-gate.md` by 17:00 UK and is layer 4's coverage table; this document's layer 4 cites it where it is named and does not restate it.