From 62ccca7d3b77bbdac9f027bfd64b2396236fa2a3 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 12:15:47 +0000 Subject: [PATCH] Class v6 10.2: the shadow's k from RTL (floor lane 2: k 0.18 at the lock at N3 absolute, a floor; the GDDR7 chip with the class v4 shadow 4.0x at the knee and 5.8x at stock against the served 2.1x; mulhi the worst lever on the card by 6x); the lead and the served-line review carry it Co-Authored-By: Claude Fable 5.1 Documents-only replay of 4e7bac031 (counter-asic-4) for the box mirror master --- docs/design/class-v6-rotating-family.md | 29 ++++++++++++++++++++++--- 1 file changed, 26 insertions(+), 3 deletions(-) diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index fc8636bfb..d7fdcbc78 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -6,7 +6,7 @@ **The lead, as main ordered it (11:2x UK): the strongest chip in the five-year window is not a DRAM chip and no rotating layer reaches it.** Lane B's reading (`docs/analysis/class-v6/hardware-future.md`, master 34f63b3c; carried into the chip model as section 5.12): a 2 GiB SRAM full store on one reticle of merchant N2 (452 mm^2 of macro at 38 Mb/mm^2, claimed) reads about 2,100 MH/s per die at 300 W, 0.14 microjoules per hash, 17x the 5090 at its stock point per joule at zero shadow (8x to 30x on the read-energy band), 5.6x the M5 Max; with the class v4 shadow 4.8x at a core half as costly as a GPU lane and 2.7x at `k = 1`, and at the honest card's whole latency shadow 3.7x and 2.0x; USD 400 to 600 of silicon per die (USD 0.25 to 0.4 per MH/s), an N2 project of USD 100 M to 500 M and 18 to 24 months (claimed), which on the mission lane's model is a break-even cap of about USD 330 M to 1.7 B in the chain's first two years. Beside it a custom HBM4E base die (2028 or later) at 6.5x to 14x, untouched by all four layers; per-bank processing in memory structurally blind to the hash (1.6 percent of reads in-bank at 2 GiB). All modelled on the chip model's method; the GPU side measured. Against the SRAM store at the hash's own width floor lane 3 reads 66x at zero shadow (lane B's 17x was the 64-byte row) and 2.7x to 3.0x at k = 1 and 5.0x to 5.7x at k = 0.5 with the shadow (section 10.3), so the honest range with the shadow is 3x to 6x and the shadow is the whole hold; and the one layer that answers it is layer 2 as a FLOOR that grows faster than SRAM cost falls: each doubling of the floor adds a die (4 GiB two dies at about USD 1,000 and 15x, 8 GiB four dies at USD 2,000 to 2,500 and 13x) while every card tier pays in device memory and, on Apple, in rate. The schedule's cost is priced in section 3 (lane A's per-tier table and main's candidate schedule, 6 GiB at the v6 epoch, 10 two years on, 14 at four, land there by 17:00 UK) and the served chip line is reviewed against this reading in section 9 at 20:00. -The honest frame first, from last night's measured close (`docs/analysis/counter-asic-4-research.md` sections 2 and 15.1a). Two chips exist in the model. A **fixed-function chip** wires one hash: the mixer's rounds, the op mix's lane ratios, the read width, the program length and the block shape are silicon. A **GPU-like chip** stores the dataset in commodity DRAM and runs the hour's program on a programmable integer core (the `f = 1` chip of chip-model-v3 section 5, the chip anyone builds); for it every drawn parameter is firmware. The four layers below render the FIRST kind useless on the day a parameter leaves its wired value, and move nothing for the second kind except the size of the core it must carry and the capex that forces. What the second kind keeps is the identity of section 2: at zero shadow premium the card's whole-card energy over the chip's memory energy (3.6x on a 5090 at its knee, measured card, modelled chip), and with the class v4 shadow 2.1x at a core as good as a GPU lane. No rotation changes that; the layers change which chip can be built and how long its tape-out lives. +The honest frame first, from last night's measured close (`docs/analysis/counter-asic-4-research.md` sections 2 and 15.1a). Two chips exist in the model. A **fixed-function chip** wires one hash: the mixer's rounds, the op mix's lane ratios, the read width, the program length and the block shape are silicon. A **GPU-like chip** stores the dataset in commodity DRAM and runs the hour's program on a programmable integer core (the `f = 1` chip of chip-model-v3 section 5, the chip anyone builds); for it every drawn parameter is firmware. The four layers below render the FIRST kind useless on the day a parameter leaves its wired value, and move nothing for the second kind except the size of the core it must carry and the capex that forces. What the second kind keeps is the identity of section 2: at zero shadow premium the card's whole-card energy over the chip's memory energy (3.6x on a 5090 at its knee, measured card, modelled chip), and with the class v4 shadow 2.1x at a core as good as a GPU lane, which the k lane's first RTL rows (10.2, 13:3x UK) say no chip core is: at N3 a chip pays about 0.18 of the locked 5090's pJ per op (a floor, synthesised on ASAP7 and scaled), so the DRAM chip with the shadow reads 4.0x at the knee and 5.8x at stock. No rotation changes that; the layers change which chip can be built and how long its tape-out lives. | Layer | A fixed-function chip on its release day | A GPU-like chip on its release day | RTX 5090 (measured where stated) | RTX 5070 Ti (MEASURED stock, a rented Vast pod, 10:08 to 10:11 UTC: class v3 78.69 MH/s at 140.8 W, class v4 78.78 at 224.0 W, the premium 83.2 W = 10.3 pJ per counted op, fingerprints equal to the Mac's; no core-lock grid, the host refused -lgc) | Apple M5 Max (measured where stated) | |---|---|---|---|---|---| @@ -253,7 +253,7 @@ The served sentence (`docs/plans/counter-asic-3-public-text-2026-10-07.md`, the |---|---|---|---| | "the strongest chip in our public model" | the GDDR7 stored-dataset chip of chip-model-v3 5.5, the 2027-on chips of 5.12 now in the model | the SRAM store (17x at zero shadow, 2.7x to 4.8x with it) and the custom base die (6.5x to 14x) are modelled in the same file since today | the clause is no longer true of the public model as written: the strongest chip in it is the SRAM store, five years out, with the project cost and the clock beside it | | "2.1x per joule against an RTX 5090 with a core as good as a GPU lane" | the class v4 shadow at the 5090's knee: 82.8 to 90.5 W premium measured four times, 6.4 to 6.6 pJ per counted op; the chip's memory 0.466 microjoules modelled | measured card, modelled chip | stands for the GDDR7 chip; against the SRAM store the same clause reads 2.7x at `k = 1` and 2.0x only at the card's whole latency shadow | -| "3.4x with one three times better" | `k` 0.33 for an ALU-shaped core: an estimate from datapath and wire figures (approximate); the microbench bounds the GPU side (11.3 pJ per ARX op stock, 6.2 at the knee) and the re-weight's hold stands | modelled; the chip side never measured | stands as the pessimistic column for the GDDR7 chip; the SRAM store's pessimistic column is 4.8x (`k` 0.5 on an N2 core, lane B) | +| "3.4x with one three times better" | `k` 0.33 for an ALU-shaped core: an estimate from datapath and wire figures (approximate); the microbench bounds the GPU side (11.3 pJ per ARX op stock, 6.2 at the knee) and the re-weight's hold stands. **The k lane's first RTL rows (10.2, 13:3x UK): k 0.18 at the lock at N3 (0.35 on unscaled ASAP7), every chip figure a floor; the GDDR7 chip with the shadow 4.0x at the knee, 5.8x at stock, so "three times better" is the unscaled-ASAP7 case and a real core at N3 is five to six times better** | modelled; the chip side never measured | stands as the pessimistic column for the GDDR7 chip; the SRAM store's pessimistic column is 4.8x (`k` 0.5 on an N2 core, lane B) | | "a 5090 locked at its knee pays 82 W for that shadow work" | the efficiency pass: 81.8 W at the best points, 82.8 on the packs job, 90.5 on the sparse job's base rows | measured | stands | | "Class v5 then makes the dataset the chain's own state, so a chip that stores it or recomputes it is wrong on every item" | class-v5-stored-state.md: a stateless or stale chip is wrong on every item; a chip that holds the state (one node per farm, the leaves at 16.5 KB/s) is not | designed, measured on the harness | the clause overstates: "a chip that does not follow the chain is wrong on every item" is the true form; a chip that stores the dataset and follows the chain is unmoved (chip-model 5.10) | | "Without class v4 the same chip would reach 5x to 9x" | chip-model 5.4: 5.1x GDDR7 to 9.2x eight HBM3 stacks at zero shadow | modelled | stands for the DRAM chips; the SRAM store reads 17x | @@ -281,13 +281,36 @@ The question the founder set: the stored-dataset chip pays the DRAM's own energy | Lane (branch) | The term | What this document already holds (measured) | The default if the lane's rows are not in by 19:30 | The row owed | |---|---|---|---|---| | SM-sparse (`class-v6-floor-sm`) | `E_card` above `E_mem`: the honest card's watts toward the DRAM's own at the activate ceiling, on rented 5090, 4090 and H100 and PC 1 through the hash lane's jobs, with a worker patch | the research file's 20.3b (PC 1, 7 October night, the fourth exe): a quarter of the SMs holds 98.2 percent of the class v4 rate at the SAME draw (460 against 451 W) and 99.8 percent of class v3 at 4 W less; watts minus idle per MH/s never falls below base; the draw follows the work, not the SM count; at the 1,300 lock the sparse shapes collapse. The candidate was closed on that row | the 20.3b reading stands: the honest card's premium-free floor is its idle plus its memory system plus whatever the SMs spend waiting (99 W of 228 at the lock, measured as the residual, not reached by idling SMs); the 5090 at its knee 1.66 microjoules against the GDDR7 chip's 0.466: 3.6x | a measured breakdown of the 99 W that an idle SM does not save (clock tree, L2, fabric), and whether a different occupancy shape (fewer warps per SM at full SM count) moves it | -| The shadow's k (`class-v6-floor-k`) | `k` from RTL synthesis (Yosys plus OpenROAD on build-4) replacing every claimed chip-side k; the op mix that maximises k | the GPU side measured (the research file's 15.1a): ARX 11.3 / 6.2 pJ per counted op, mul 13.9 / 8.3, mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, fp32 FMA 9.2 / 5.2, the int8 tile 1.4 to 4.1 per MAC; the chip side claimed only (a 5 nm SIMD array 2 to 5 pJ per op, approximate) | the served band stands: k 0.3 to 0.8 for the ALU shadow (2.1x at k = 1, 3.4x at k about 0.33); the op mix held at class v4's (15.1b: the shuffle at 4.9x the add on the GPU side) | the synthesised pJ per op per family for a 32-lane SIMD core at the chosen node, the resulting k per family, and the mix that maximises k at a fixed GPU premium | +| The shadow's k (`class-v6-floor-k`) | `k` from RTL synthesis (Yosys plus OpenROAD on build-4) replacing every claimed chip-side k; the op mix that maximises k | the GPU side measured (the research file's 15.1a): ARX 11.3 / 6.2 pJ per counted op, mul 13.9 / 8.3, mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, fp32 FMA 9.2 / 5.2, the int8 tile 1.4 to 4.1 per MAC; the chip side was claimed only (a 5 nm SIMD array 2 to 5 pJ per op, approximate) until the lane's first rows (10.2, 13:3x UK): synthesised on ASAP7, k 0.18 at the lock in the absolute convention at N3 (0.35 unscaled), the ARX families 0.17 to 0.18, mad 0.20, mulhi 0.03 | superseded by the lane's rows: the GDDR7 chip with the class v4 shadow reads 4.0x at the lock and 5.8x at stock (absolute), not the served 2.1x at k = 1; the op mix held at class v4's (15.1b: the shuffle at 4.9x the add on the GPU side; mulhi the worst lever on the card by 6x) | the synthesised pJ per op per family for a 32-lane SIMD core at the chosen node, the resulting k per family, and the mix that maximises k at a fixed GPU premium | | The SRAM full-store die (`class-v6-floor-sram`) | `E_mem` for the chip lane B named strongest: the wire-energy lever, the dataset schedule that keeps the store above USD 5,000 of silicon through 2031, the capex wall recomputed with rotation | lane B's rows (7a, chip-model 5.12): 2 GiB on one N2 reticle, 1.0 nJ per read (0.5 to 2.0), 17x at zero shadow, 2.7x to 4.8x with it, USD 400 to 600 per die; lane A's schedule (3.4): 5.5 / 8 / 11 GiB retiring a quarter of today's consumer cards per step; the M5 Max's measured rate cost of the larger working set (3.3) | lane B's and lane A's rows stand; the schedule decision is the founder's with the per-tier table (0.1) | the schedule in GiB per year that keeps the store at USD 5,000 or more of silicon through 2031 on the SRAM cost curve, and what it costs each tier by year | | The honest denominator (`class-v6-floor-denominator`) | `E_card` per tier with every software knob (the core lock, the undervolt, the memory clock, the occupancy, the block shape); the Ember tier table | the knee rows (the research file's 20.3a and 20.3b; the cost rows): the 5090 at 1,300 MHz 1.66 to 1.69 microjoules, the 4070 at its tune 2.57, the 5070 Ti stock 1.79, the M5 Max 0.78 (GPU and DRAM channels); the operating point is rank 1 of the research file | those rows stand as the per-tier floor; the Ember knob ships as ordered for 0.3.24 | the per-tier best point with every knob, the Ember table, and the AMD and Apple lines (ADLX or nothing; no lever) | | Invention beyond the four (`class-v6-floor-invention`) | the tensor core's k re-read, the RT core, controller plus PHY cost, proof of latency, era-driven address mapping, the literature since 2023 | the research file: the tensor tile's measured 1.4 to 4.1 pJ per MAC against a 5 nm array's claimed 0.04 to 0.4 (k 0.03 to 0.3, the worse lever); the L2 hit at 1.4 to 2.4 nJ against a chip's SRAM (k 0.1 to 0.3); the texture interpolator excluded (not bit-exact across vendors); the literature table (sections 5 and 8) | those readings stand; no new lever | anything that reads k above 1 on measured rows, which nothing has | The close (20:00): the honest floor per tier against each chip row, the recommended class v6 changes (the op mix, the SM-sparse default, the dataset schedule), the served chip line as a measured range, and what stays for v7, written here from the lanes' rows as they land. +### 10.2 The shadow's k from RTL (floor lane 2, `docs/analysis/class-v6/floor/shadow-k.md` on class-v6-floor-k, build-4 and build-3, first synthesised rows 13:3x UK; RTL and flow under `tools/chip-model/rtl`; full by 19:30 UK) + +**The first measured-by-synthesis k is 0.18 at the 5090's lock in the absolute convention at N3, and nothing reads inside the claimed 0.3 to 0.8 band except the unscaled ASAP7 figure at the lock (0.35).** Method: a minimal lane (instruction register, an 8 x 32-bit register window of flops, read muxes, the unit, one write port; no fetch or decode) in Yosys 0.68 plus OpenROAD on ASAP7, routed, SPEF, power from a random-input gate-level VCD with every pin annotated, the TC corner at 0.70 V; so every chip figure is a FLOOR and every k a floor. The node scaling is claimed from TSMC's headline per-node power reductions (N5 x0.70, N3 x0.50, N2 x0.36 of ASAP7; approximate). The GPU side is the research file's 15.1a, measured. + +| Family (chip RTL) | pJ per op ASAP7 (synthesised) | pJ per op N3 (scaled, claimed) | 5090 pJ per op stock / lock (measured) | k absolute, N3 chip against stock / lock | k at ASAP7 unscaled against the lock | +|---|---|---|---|---|---| +| ARX add | 2.16 | 1.09 | 11.3 / 6.2 | 0.096 / 0.18 | 0.35 | +| ARX xor | 2.09 | 1.05 | 11.3 / 6.2 | 0.093 / 0.17 | 0.34 | +| ARX rotl (imm) / rotr (reg) | 2.15 / 2.13 | 1.08 / 1.07 | 11.3 / 6.2 | 0.095 / 0.17 | 0.35 | +| ARX sub | 2.14 | 1.08 | 11.3 / 6.2 | 0.095 / 0.17 | 0.35 | +| or (lossy; values saturate, activity falls) | 1.45 | 0.73 | 11.3 / 6.2 | 0.065 / 0.12 | 0.23 | +| ARX random mix | 2.24 | 1.13 | 11.3 / 6.2 | 0.10 / 0.18 | 0.36 | +| mul (32 x 32 low) | 1.35 | 0.68 | 13.9 / 8.3 | 0.049 / 0.082 | 0.16 | +| mulhi (high word, the same multiplier) | 1.35 | 0.68 | 39.6 / 21.0 | 0.017 / 0.032 | 0.064 | +| mad (3 reads, mul plus add) | 3.29 | 1.66 | 13.9 / 8.3 | 0.12 / 0.20 | 0.40 | +| index fold (mul by M, rotl R, masks) | 2.37 | 1.19 | 13.9 / 8.3 | 0.086 / 0.14 | 0.29 | +| prmt (byte permute) | 1.27 | 0.64 | 22.3 / 11.5 | 0.029 / 0.056 | 0.11 | +| lop3 (8-bit LUT) | 1.42 | 0.72 | 24.1 / 13.0 | 0.030 / 0.055 | 0.11 | + +The ranking by k, highest first (hardest for a chip; absolute, N3 against the lock): mad 0.20, the ARX families 0.17 to 0.18, the fold 0.14, or 0.12, mul 0.08, prmt and lop3 0.06, mulhi 0.03. mulhi is the worst lever on the card by 6x: the 5090 pays 21 pJ for a high word that costs the chip the same 0.68 pJ as a low word, which is the measured reason the op-mix band keeps mulhi capped at its base (section 2) and the reason the research file's 15.1b hold on a re-weight stands from the chip side as well. The class v4 mix costs the chip about 1.0 to 1.2 pJ per op at N3 (2.0 to 2.4 at ASAP7), the shuffle row pending. + +What it does to the GDDR7 stored-dataset board (E_mem 0.466 microjoules, modelled) with the class v4 shadow (102,100 ops per hash): absolute, the chip's shadow is 102,100 x about 1.1 pJ = 0.11 microjoules at N3, E_chip about 0.58, so the edge reads **4.0x at the lock (the card 2.33) and 5.8x at stock (3.36)**, against the served 2.1x at k = 1 and 3.4x at k 0.33. In the record's convention it is k 0.18 at the lock, 2.33 / (0.466 + 0.18 x 0.652) = 4.0x, the same number there, and at stock k 0.10, 3.36 / (0.466 + 0.10 x 1.10) = 5.8x. **The shadow buys the card about 0.5x to 1x of edge out of the 5.2x and 3.6x zero-shadow figures, not the 1.5x the served k = 1 row implies.** This is the row that moves the served line (section 9): the served "2.1x" is the k = 1 reading, and the first RTL floor says a chip's core is five to six times cheaper than that at N3; the honest shadowed range against the DRAM chip at the knee is 4x by this reading, with every chip figure a floor (no fetch or decode, no clock tree beyond the lane, no node-scaling measurement). Pending from the lane by 19:30 UK: the 32-lane shuffle butterfly and the general crossbar (in place now), the 8 KB scratch as a flop array, the int8 8x8x8 tile, the mix optimiser over the layer-1 band, and the re-fold of lane 3's SRAM rows at W = 4 and 8 in both conventions. + ### 10.3 The SRAM die and the dataset floor (floor lane 3, `docs/analysis/class-v6/floor/sram-and-floor.md` on class-v6-floor-sram at 708d01b4, first reading 12:0x UTC; full by 19:30 UK; everything chip-side modelled on lane B's wire figure, 1.3 pJ per bit plus 0.1 nJ per macro access and 0.05 nJ of controller; the GPU side measured) **The read width is the only wire lever on the SRAM die, and the floor is a ticket, not a joule.** Lane B's 1.0 nJ was the 64-byte row; the die reads what the hash asks for, and the wire scales with the bits moved while the macro access does not: