Class v6: 10.5 invention lane's KEEP/KILL table (the second sector conditional on the w64 job; everything else killed with its number); 10.2 reframed as per-unit floors with the sequencer-core row owed as the headline and k 0.5 the default; lane D's complete lossy curve and ring-C progress in 5.1b

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 650c2ad56 (counter-asic-4) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 12:17:42 +00:00
parent 6ebe399950
commit ace2e5367c

View file

@ -6,7 +6,7 @@
**The lead, as main ordered it (11:2x UK): the strongest chip in the five-year window is not a DRAM chip and no rotating layer reaches it.** Lane B's reading (`docs/analysis/class-v6/hardware-future.md`, master 34f63b3c; carried into the chip model as section 5.12): a 2 GiB SRAM full store on one reticle of merchant N2 (452 mm^2 of macro at 38 Mb/mm^2, claimed) reads about 2,100 MH/s per die at 300 W, 0.14 microjoules per hash, 17x the 5090 at its stock point per joule at zero shadow (8x to 30x on the read-energy band), 5.6x the M5 Max; with the class v4 shadow 4.8x at a core half as costly as a GPU lane and 2.7x at `k = 1`, and at the honest card's whole latency shadow 3.7x and 2.0x; USD 400 to 600 of silicon per die (USD 0.25 to 0.4 per MH/s), an N2 project of USD 100 M to 500 M and 18 to 24 months (claimed), which on the mission lane's model is a break-even cap of about USD 330 M to 1.7 B in the chain's first two years. Beside it a custom HBM4E base die (2028 or later) at 6.5x to 14x, untouched by all four layers; per-bank processing in memory structurally blind to the hash (1.6 percent of reads in-bank at 2 GiB). All modelled on the chip model's method; the GPU side measured. Against the SRAM store at the hash's own width floor lane 3 reads 66x at zero shadow (lane B's 17x was the 64-byte row) and 2.7x to 3.0x at k = 1 and 5.0x to 5.7x at k = 0.5 with the shadow (section 10.3), so the honest range with the shadow is 3x to 6x and the shadow is the whole hold; and the one layer that answers it is layer 2 as a FLOOR that grows faster than SRAM cost falls: each doubling of the floor adds a die (4 GiB two dies at about USD 1,000 and 15x, 8 GiB four dies at USD 2,000 to 2,500 and 13x) while every card tier pays in device memory and, on Apple, in rate. The schedule's cost is priced in section 3 (lane A's per-tier table and main's candidate schedule, 6 GiB at the v6 epoch, 10 two years on, 14 at four, land there by 17:00 UK) and the served chip line is reviewed against this reading in section 9 at 20:00.
The honest frame first, from last night's measured close (`docs/analysis/counter-asic-4-research.md` sections 2 and 15.1a). Two chips exist in the model. A **fixed-function chip** wires one hash: the mixer's rounds, the op mix's lane ratios, the read width, the program length and the block shape are silicon. A **GPU-like chip** stores the dataset in commodity DRAM and runs the hour's program on a programmable integer core (the `f = 1` chip of chip-model-v3 section 5, the chip anyone builds); for it every drawn parameter is firmware. The four layers below render the FIRST kind useless on the day a parameter leaves its wired value, and move nothing for the second kind except the size of the core it must carry and the capex that forces. What the second kind keeps is the identity of section 2: at zero shadow premium the card's whole-card energy over the chip's memory energy (3.6x on a 5090 at its knee, measured card, modelled chip), and with the class v4 shadow 2.1x at a core as good as a GPU lane, which the k lane's first RTL rows (10.2, 13:3x UK) say no chip core is: at N3 a chip pays about 0.18 of the locked 5090's pJ per op (a floor, synthesised on ASAP7 and scaled), so the DRAM chip with the shadow reads 4.0x at the knee and 5.8x at stock. No rotation changes that; the layers change which chip can be built and how long its tape-out lives.
The honest frame first, from last night's measured close (`docs/analysis/counter-asic-4-research.md` sections 2 and 15.1a). Two chips exist in the model. A **fixed-function chip** wires one hash: the mixer's rounds, the op mix's lane ratios, the read width, the program length and the block shape are silicon. A **GPU-like chip** stores the dataset in commodity DRAM and runs the hour's program on a programmable integer core (the `f = 1` chip of chip-model-v3 section 5, the chip anyone builds); for it every drawn parameter is firmware. The four layers below render the FIRST kind useless on the day a parameter leaves its wired value, and move nothing for the second kind except the size of the core it must carry and the capex that forces. What the second kind keeps is the identity of section 2: at zero shadow premium the card's whole-card energy over the chip's memory energy (3.6x on a 5090 at its knee, measured card, modelled chip), and with the class v4 shadow 2.1x at a core as good as a GPU lane, which the k lane's first RTL rows (10.2, 13:3x UK) say no chip's UNITS are: at N3 a bare unit pays about 0.18 of the locked 5090's pJ per op (a per-unit floor, synthesised on ASAP7 and scaled; no fetch, decode or register file, which is the part rotation forces a chip to carry), so the DRAM chip with the shadow reads up to 4.0x at the knee and 5.8x at stock as the worst case, with the sequencer-core row owed as the headline and k 0.5 (2.9x at the knee) the default until it lands. No rotation changes that; the layers change which chip can be built and how long its tape-out lives.
| Layer | A fixed-function chip on its release day | A GPU-like chip on its release day | RTX 5090 (measured where stated) | RTX 5070 Ti (MEASURED stock, a rented Vast pod, 10:08 to 10:11 UTC: class v3 78.69 MH/s at 140.8 W, class v4 78.78 at 224.0 W, the premium 83.2 W = 10.3 pJ per counted op, fingerprints equal to the Mac's; no core-lock grid, the host refused -lgc) | Apple M5 Max (measured where stated) |
|---|---|---|---|---|---|
@ -170,11 +170,11 @@ What the cut settles, each row measured unless marked:
| Finding | The numbers | What it moves in this document |
|---|---|---|
| The op-mix band is settled by measurement | B = 4 with or, mul and mulhi free to rise exhausts the 256-attempt cap in 1.2 percent of eras (61 of 5,000 at the lossy corner); with or, mul and mulhi never raised above their base (section 2's band) r = 0.595, mean attempt 1.47, max 59, 0 exhausted in 3,000. The lossy-share curve (running, about 1,200 eras per point): r = 0.80, 0.88, 0.92, 0.96 at +1 to +4 points; exhaustion 0, 0, 0.2 percent, 1.4 percent; the edge between +2 and +3 points | the layer 1 op-mix row's known-failed test now reads 0 of 3,000 under the band (it owed a 0); the cap on the lossy families stays at the base, with +2 points the most the band could ever open to |
| The op-mix band is settled by measurement | B = 4 with or, mul and mulhi free to rise exhausts the 256-attempt cap in 1.2 percent of eras (61 of 5,000 at the lossy corner); with or, mul and mulhi never raised above their base (section 2's band) r = 0.595, mean attempt 1.47, max 59, 0 exhausted in 3,000. The lossy-share curve, complete at 3,000 eras per point (13:5x UK): r = 0.80, 0.88, 0.92, 0.96 at +1 to +4 points on or, mul and mulhi; exhaustion 0, 0.07, 0.20, 1.10 percent of eras, against independent-attempt estimates of 4e-25, 3e-15, 9e-10, 1e-5, so the per-era correlation is 10^5 to 10^10 above the geometric figure and the band's edge is the measured +2, not the arithmetic's | the layer 1 op-mix row's known-failed test now reads 0 of 3,000 under the band (it owed a 0); the cap on the lossy families stays at the base, with +2 points the most the band could ever open to |
| What breaks at the lossy corner | class v5's last-resort scan passes at its first or second candidate on every exhausted era seen, so the corner costs liveness time, not an unchecked program; the per-era exhaustion is about 1,000x the independent-attempt estimate because one era's attempts share its weight table, which is the independence the scan's 1e-300 assumes | section 5.2's bound: the attempts within an era are not independent draws; the bound is per era from the census, not r^256 |
| The (c''') floor needs a per-width calibration | at width 4 (expectation divided by the width) it refuses 4.5 to 6.5 percent of candidates against 0.6 to 1.0 percent at width 1 | ring B's row: either a floor per width from the clean spread, or the statistic in sigma over its own expectation (the form the bucket statistic now takes); the full report (09:00 UK tomorrow) carries the per-width floor; W pinned at 4 at genesis (section 10.3) makes it one calibration, not a family of them |
| The (c''') floor needs a per-width calibration | at width 4 (expectation divided by the width) it refuses 4.5 to 6.5 percent of candidates against 0.6 to 1.0 percent at width 1 | ring B's row: with W pinned at 4 at genesis (section 10.3) it is one measurement, being taken now (the harness's next build records the pre-floor spread of every candidate at width 4 under the band over 3,000 eras plus the width-1 control); the 09:00 report states either a single width-4 floor at the shipped clean-refusal rate (about 2.4 percent of candidates) or the sigma-over-expectation form, whichever keeps the known-failed hot sets refused at the lower clean cost; the default here is the sigma form |
| The shape axis moves the draw's cost, the mixer axis is invisible | 64 x 108: 1.8 attempts; 256 x 27: 3.8, through (a')'s fixpoint over the block; r 0.714 to 0.731 across m | m is closed per family by adv-mixer-3's ladder at m = 4 and the x16 verifier row, not by drawing eras (section 5.1a's reading stands, now measured); the block-shape row of layer 1 carries the attempt cost |
| The era-stride bias at the bit level | 48 to 58 percent of accepted programs in every stratum carry one site biased at over 6 sigma at 2^20; 73 percent of those at address bit R exactly (12 percent at R+1, 5 at R+2: the product law's bits 0, 1, 2 through `rotl(x*M, R)`); a third over 100 sigma, worst z 1,024; under R in 28..31 the over-100-sigma share falls from 37.7 to 7.5 percent; the bucket statistic sees the same mechanism (p99 +6.8 sigma on bit-clean eras, +44 on biased ones) | one value-level test covers both; the remedy is structural, the index fold of the product's low bits in `load_index` before the rotation (layer 1's rule, section 2), not a per-epoch refusal that would redraw half the epochs; the live-dataset price per site (ring C) is being read now (the F8 census at 2^24 on 64 seeds at two band points) |
| The era-stride bias at the bit level | 48 to 58 percent of accepted programs in every stratum carry one site biased at over 6 sigma at 2^20; 73 percent of those at address bit R exactly (12 percent at R+1, 5 at R+2: the product law's bits 0, 1, 2 through `rotl(x*M, R)`); a third over 100 sigma, worst z 1,024; under R in 28..31 the over-100-sigma share falls from 37.7 to 7.5 percent; the bucket statistic sees the same mechanism (p99 +6.8 sigma on bit-clean eras, +44 on biased ones) | one value-level test covers both; the remedy is structural, the index fold of the product's low bits in `load_index` before the rotation (layer 1's rule, section 2), not a per-epoch refusal that would redraw half the epochs; the live-dataset price per site (ring C) is being read now (the F8 census at 2^24 on 64 seeds at two band points; 31 of 64 seeds PASS so far at the shape-256 point, no test fired) |
| The bound arithmetic on the counts | 10,000 random eras and 3,000 band eras passing bound the failing fraction on the ring-B tests at 3.0e-4 and 1.0e-3 at 95 percent; the floors' miss rates on the live-dataset classes rest on 9 and 5 known-failed cases (under 0.33 and 0.60), tightened by cases from the adversarial tails, not by eras | section 5.1a's sampling bound now has its measured n; the gate-record JSON per era lands with the full report |
Owed from lane D by 09:00 UK tomorrow: the per-width floor, the bucket bound in sigma, the finished lossy curve, the two ring-C live rows, the gate-record JSON per era.
@ -253,7 +253,7 @@ The served sentence (`docs/plans/counter-asic-3-public-text-2026-10-07.md`, the
|---|---|---|---|
| "the strongest chip in our public model" | the GDDR7 stored-dataset chip of chip-model-v3 5.5, the 2027-on chips of 5.12 now in the model | the SRAM store (17x at zero shadow, 2.7x to 4.8x with it) and the custom base die (6.5x to 14x) are modelled in the same file since today | the clause is no longer true of the public model as written: the strongest chip in it is the SRAM store, five years out, with the project cost and the clock beside it |
| "2.1x per joule against an RTX 5090 with a core as good as a GPU lane" | the class v4 shadow at the 5090's knee: 82.8 to 90.5 W premium measured four times, 6.4 to 6.6 pJ per counted op; the chip's memory 0.466 microjoules modelled | measured card, modelled chip | stands for the GDDR7 chip; against the SRAM store the same clause reads 2.7x at `k = 1` and 2.0x only at the card's whole latency shadow |
| "3.4x with one three times better" | `k` 0.33 for an ALU-shaped core: an estimate from datapath and wire figures (approximate); the microbench bounds the GPU side (11.3 pJ per ARX op stock, 6.2 at the knee) and the re-weight's hold stands. **The k lane's first RTL rows (10.2, 13:3x UK): k 0.18 at the lock at N3 (0.35 on unscaled ASAP7), every chip figure a floor; the GDDR7 chip with the shadow 4.0x at the knee, 5.8x at stock, so "three times better" is the unscaled-ASAP7 case and a real core at N3 is five to six times better** | modelled; the chip side never measured | stands as the pessimistic column for the GDDR7 chip; the SRAM store's pessimistic column is 4.8x (`k` 0.5 on an N2 core, lane B) |
| "3.4x with one three times better" | `k` 0.33 for an ALU-shaped core: an estimate from datapath and wire figures (approximate); the microbench bounds the GPU side (11.3 pJ per ARX op stock, 6.2 at the knee) and the re-weight's hold stands. **The k lane's first RTL rows (10.2, 13:3x UK): k 0.18 at the lock at N3 (0.35 on unscaled ASAP7), every chip figure a floor; the GDDR7 chip with the shadow 4.0x at the knee, 5.8x at stock, so "three times better" is the unscaled-ASAP7 unit case and a bare unit at N3 is five to six times better; the sequencer core a rotating family forces sits between, its row owed as the headline; the served figure must survive 4.0x at the knee as the worst case** | modelled; the chip side never measured | stands as the pessimistic column for the GDDR7 chip; the SRAM store's pessimistic column is 4.8x (`k` 0.5 on an N2 core, lane B) |
| "a 5090 locked at its knee pays 82 W for that shadow work" | the efficiency pass: 81.8 W at the best points, 82.8 on the packs job, 90.5 on the sparse job's base rows | measured | stands |
| "Class v5 then makes the dataset the chain's own state, so a chip that stores it or recomputes it is wrong on every item" | class-v5-stored-state.md: a stateless or stale chip is wrong on every item; a chip that holds the state (one node per farm, the leaves at 16.5 KB/s) is not | designed, measured on the harness | the clause overstates: "a chip that does not follow the chain is wrong on every item" is the true form; a chip that stores the dataset and follows the chain is unmoved (chip-model 5.10) |
| "Without class v4 the same chip would reach 5x to 9x" | chip-model 5.4: 5.1x GDDR7 to 9.2x eight HBM3 stacks at zero shadow | modelled | stands for the DRAM chips; the SRAM store reads 17x |
@ -281,7 +281,7 @@ The question the founder set: the stored-dataset chip pays the DRAM's own energy
| Lane (branch) | The term | What this document already holds (measured) | The default if the lane's rows are not in by 19:30 | The row owed |
|---|---|---|---|---|
| SM-sparse (`class-v6-floor-sm`) | `E_card` above `E_mem`: the honest card's watts toward the DRAM's own at the activate ceiling, on rented 5090, 4090 and H100 and PC 1 through the hash lane's jobs, with a worker patch | the research file's 20.3b (PC 1, 7 October night, the fourth exe): a quarter of the SMs holds 98.2 percent of the class v4 rate at the SAME draw (460 against 451 W) and 99.8 percent of class v3 at 4 W less; watts minus idle per MH/s never falls below base; the draw follows the work, not the SM count; at the 1,300 lock the sparse shapes collapse. The candidate was closed on that row | the 20.3b reading stands: the honest card's premium-free floor is its idle plus its memory system plus whatever the SMs spend waiting (99 W of 228 at the lock, measured as the residual, not reached by idling SMs); the 5090 at its knee 1.66 microjoules against the GDDR7 chip's 0.466: 3.6x | a measured breakdown of the 99 W that an idle SM does not save (clock tree, L2, fabric), and whether a different occupancy shape (fewer warps per SM at full SM count) moves it |
| The shadow's k (`class-v6-floor-k`) | `k` from RTL synthesis (Yosys plus OpenROAD on build-4) replacing every claimed chip-side k; the op mix that maximises k | the GPU side measured (the research file's 15.1a): ARX 11.3 / 6.2 pJ per counted op, mul 13.9 / 8.3, mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, fp32 FMA 9.2 / 5.2, the int8 tile 1.4 to 4.1 per MAC; the chip side was claimed only (a 5 nm SIMD array 2 to 5 pJ per op, approximate) until the lane's first rows (10.2, 13:3x UK): synthesised on ASAP7, k 0.18 at the lock in the absolute convention at N3 (0.35 unscaled), the ARX families 0.17 to 0.18, mad 0.20, mulhi 0.03 | superseded by the lane's rows: the GDDR7 chip with the class v4 shadow reads 4.0x at the lock and 5.8x at stock (absolute), not the served 2.1x at k = 1; the op mix held at class v4's (15.1b: the shuffle at 4.9x the add on the GPU side; mulhi the worst lever on the card by 6x) | the synthesised pJ per op per family for a 32-lane SIMD core at the chosen node, the resulting k per family, and the mix that maximises k at a fixed GPU premium |
| The shadow's k (`class-v6-floor-k`) | `k` from RTL synthesis (Yosys plus OpenROAD on build-4) replacing every claimed chip-side k; the op mix that maximises k | the GPU side measured (the research file's 15.1a): ARX 11.3 / 6.2 pJ per counted op, mul 13.9 / 8.3, mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, fp32 FMA 9.2 / 5.2, the int8 tile 1.4 to 4.1 per MAC; the chip side was claimed only (a 5 nm SIMD array 2 to 5 pJ per op, approximate) until the lane's first rows (10.2, 13:3x UK): synthesised on ASAP7, k 0.18 at the lock in the absolute convention at N3 (0.35 unscaled), the ARX families 0.17 to 0.18, mad 0.20, mulhi 0.03 | the default stays the record's rows at k 0.5 (the DRAM chip 2.9x at the knee) with the unit floors as the lower bound: the GDDR7 chip with the class v4 shadow at 4.0x at the lock and 5.8x at stock (absolute, floor k) is the worst case the served line must survive, not the reading; the sequencer-core row with the M5 Max column is the headline when it lands; the op mix held at class v4's (15.1b: the shuffle at 4.9x the add on the GPU side; mulhi the worst lever on the card by 6x) | the synthesised pJ per op per family for a 32-lane SIMD core at the chosen node, the resulting k per family, and the mix that maximises k at a fixed GPU premium |
| The SRAM full-store die (`class-v6-floor-sram`) | `E_mem` for the chip lane B named strongest: the wire-energy lever, the dataset schedule that keeps the store above USD 5,000 of silicon through 2031, the capex wall recomputed with rotation | lane B's rows (7a, chip-model 5.12): 2 GiB on one N2 reticle, 1.0 nJ per read (0.5 to 2.0), 17x at zero shadow, 2.7x to 4.8x with it, USD 400 to 600 per die; lane A's schedule (3.4): 5.5 / 8 / 11 GiB retiring a quarter of today's consumer cards per step; the M5 Max's measured rate cost of the larger working set (3.3) | lane B's and lane A's rows stand; the schedule decision is the founder's with the per-tier table (0.1) | the schedule in GiB per year that keeps the store at USD 5,000 or more of silicon through 2031 on the SRAM cost curve, and what it costs each tier by year |
| The honest denominator (`class-v6-floor-denominator`) | `E_card` per tier with every software knob (the core lock, the undervolt, the memory clock, the occupancy, the block shape); the Ember tier table | the knee rows (the research file's 20.3a and 20.3b; the cost rows): the 5090 at 1,300 MHz 1.66 to 1.69 microjoules, the 4070 at its tune 2.57, the 5070 Ti stock 1.79, the M5 Max 0.78 (GPU and DRAM channels); the operating point is rank 1 of the research file | those rows stand as the per-tier floor; the Ember knob ships as ordered for 0.3.24 | the per-tier best point with every knob, the Ember table, and the AMD and Apple lines (ADLX or nothing; no lever) |
| Invention beyond the four (`class-v6-floor-invention`) | the tensor core's k re-read, the RT core, controller plus PHY cost, proof of latency, era-driven address mapping, the literature since 2023 | the research file: the tensor tile's measured 1.4 to 4.1 pJ per MAC against a 5 nm array's claimed 0.04 to 0.4 (k 0.03 to 0.3, the worse lever); the L2 hit at 1.4 to 2.4 nJ against a chip's SRAM (k 0.1 to 0.3); the texture interpolator excluded (not bit-exact across vendors); the literature table (sections 5 and 8) | those readings stand; no new lever | anything that reads k above 1 on measured rows, which nothing has |
@ -290,7 +290,7 @@ The close (20:00): the honest floor per tier against each chip row, the recommen
### 10.2 The shadow's k from RTL (floor lane 2, `docs/analysis/class-v6/floor/shadow-k.md` on class-v6-floor-k, build-4 and build-3, first synthesised rows 13:3x UK; RTL and flow under `tools/chip-model/rtl`; full by 19:30 UK)
**The first measured-by-synthesis k is 0.18 at the 5090's lock in the absolute convention at N3, and nothing reads inside the claimed 0.3 to 0.8 band except the unscaled ASAP7 figure at the lock (0.35).** Method: a minimal lane (instruction register, an 8 x 32-bit register window of flops, read muxes, the unit, one write port; no fetch or decode) in Yosys 0.68 plus OpenROAD on ASAP7, routed, SPEF, power from a random-input gate-level VCD with every pin annotated, the TC corner at 0.70 V; so every chip figure is a FLOOR and every k a floor. The node scaling is claimed from TSMC's headline per-node power reductions (N5 x0.70, N3 x0.50, N2 x0.36 of ASAP7; approximate). The GPU side is the research file's 15.1a, measured.
**The first synthesised k is 0.18 at the 5090's lock in the absolute convention at N3, and nothing reads inside the claimed 0.3 to 0.8 band except the unscaled ASAP7 figure at the lock (0.35). These are PER-UNIT FLOORS: no fetch, no decode, no register file beyond an 8-entry window, which is the chip rotation kills (section 1). The coordinator's order (13:5x UK): the headline k for the close is the programmable sequencer-core row (fetch, decode, a 32-register file, the per-era registers, the drawn program), asked of the lane with the M5 Max column; until it lands the default is the record's shadowed rows at k 0.5 with these unit floors beside them as the lower bound, and the GDDR7 board at 4.0x at the lock and 5.8x at stock on the floor k is the WORST CASE the served line must survive, not the reading.** Method: a minimal lane (instruction register, an 8 x 32-bit register window of flops, read muxes, the unit, one write port; no fetch or decode) in Yosys 0.68 plus OpenROAD on ASAP7, routed, SPEF, power from a random-input gate-level VCD with every pin annotated, the TC corner at 0.70 V; so every chip figure is a FLOOR and every k a floor. The node scaling is claimed from TSMC's headline per-node power reductions (N5 x0.70, N3 x0.50, N2 x0.36 of ASAP7; approximate). The GPU side is the research file's 15.1a, measured.
| Family (chip RTL) | pJ per op ASAP7 (synthesised) | pJ per op N3 (scaled, claimed) | 5090 pJ per op stock / lock (measured) | k absolute, N3 chip against stock / lock | k at ASAP7 unscaled against the lock |
|---|---|---|---|---|---|
@ -309,7 +309,7 @@ The close (20:00): the honest floor per tier against each chip row, the recommen
The ranking by k, highest first (hardest for a chip; absolute, N3 against the lock): mad 0.20, the ARX families 0.17 to 0.18, the fold 0.14, or 0.12, mul 0.08, prmt and lop3 0.06, mulhi 0.03. mulhi is the worst lever on the card by 6x: the 5090 pays 21 pJ for a high word that costs the chip the same 0.68 pJ as a low word, which is the measured reason the op-mix band keeps mulhi capped at its base (section 2) and the reason the research file's 15.1b hold on a re-weight stands from the chip side as well. The class v4 mix costs the chip about 1.0 to 1.2 pJ per op at N3 (2.0 to 2.4 at ASAP7), the shuffle row pending.
What it does to the GDDR7 stored-dataset board (E_mem 0.466 microjoules, modelled) with the class v4 shadow (102,100 ops per hash): absolute, the chip's shadow is 102,100 x about 1.1 pJ = 0.11 microjoules at N3, E_chip about 0.58, so the edge reads **4.0x at the lock (the card 2.33) and 5.8x at stock (3.36)**, against the served 2.1x at k = 1 and 3.4x at k 0.33. In the record's convention it is k 0.18 at the lock, 2.33 / (0.466 + 0.18 x 0.652) = 4.0x, the same number there, and at stock k 0.10, 3.36 / (0.466 + 0.10 x 1.10) = 5.8x. **The shadow buys the card about 0.5x to 1x of edge out of the 5.2x and 3.6x zero-shadow figures, not the 1.5x the served k = 1 row implies.** This is the row that moves the served line (section 9): the served "2.1x" is the k = 1 reading, and the first RTL floor says a chip's core is five to six times cheaper than that at N3; the honest shadowed range against the DRAM chip at the knee is 4x by this reading, with every chip figure a floor (no fetch or decode, no clock tree beyond the lane, no node-scaling measurement). Pending from the lane by 19:30 UK: the 32-lane shuffle butterfly and the general crossbar (in place now), the 8 KB scratch as a flop array, the int8 8x8x8 tile, the mix optimiser over the layer-1 band, and the re-fold of lane 3's SRAM rows at W = 4 and 8 in both conventions.
What it does to the GDDR7 stored-dataset board (E_mem 0.466 microjoules, modelled) with the class v4 shadow (102,100 ops per hash): absolute, the chip's shadow is 102,100 x about 1.1 pJ = 0.11 microjoules at N3, E_chip about 0.58, so the edge reads **4.0x at the lock (the card 2.33) and 5.8x at stock (3.36)**, against the served 2.1x at k = 1 and 3.4x at k 0.33. In the record's convention it is k 0.18 at the lock, 2.33 / (0.466 + 0.18 x 0.652) = 4.0x, the same number there, and at stock k 0.10, 3.36 / (0.466 + 0.10 x 1.10) = 5.8x. **The shadow buys the card about 0.5x to 1x of edge out of the 5.2x and 3.6x zero-shadow figures, not the 1.5x the served k = 1 row implies.** This is the row the served line must survive (section 9): the served "2.1x" is the k = 1 reading; the unit floor says a core's units are five to six times cheaper than that at N3, and the sequencer core that a rotating family forces (fetch, decode, register file, the drawn program's control) sits between the two, where the lane's next row puts it; until then the default reading is k 0.5 (the DRAM chip 2.9x at the knee in the record's convention) with 4.0x as the floor-k worst case. Pending from the lane by 19:30 UK: the 32-lane shuffle butterfly and the general crossbar (in place now), the 8 KB scratch as a flop array, the int8 8x8x8 tile, the mix optimiser over the layer-1 band, and the re-fold of lane 3's SRAM rows at W = 4 and 8 in both conventions.
### 10.3 The SRAM die and the dataset floor (floor lane 3, `docs/analysis/class-v6/floor/sram-and-floor.md` on class-v6-floor-sram at 708d01b4, first reading 12:0x UTC; full by 19:30 UK; everything chip-side modelled on lane B's wire figure, 1.3 pJ per bit plus 0.1 nJ per macro access and 0.05 nJ of controller; the GPU side measured)
@ -358,6 +358,25 @@ The capex wall on the spec's emission constant (3,168,808,781 base units per DAA
Below about USD 100 M a year of miner revenue (IGN 0.13) no rational SRAM project starts at any edge; above about USD 2.3 B a year (IGN 2.9) every one does; at IGN 0.01 / 0.10 / 1.00 (assumed) a USD 150 M project at 5x and 30 percent reads NPV -148 / -130 / +54 M. **The edge moves p* 1.4x across 3x to 13x; the project cost moves it 5x: the hash does not set the wall, the project cost and the detector do.** Killed with the number: per-era re-fill (0.96 J per era), per-block re-fill (0.96 W on the die against 1.3 to 3.2 percent of GPU rate), straddling atoms (+0.02 nJ on the die, +19 percent of items on the verifier at the gate), a second hot table (the die's own kind of memory, k 0.1 to 0.3). Kept: the per-lane scratch pins the lane (a dependency of the width lever); 3D-stacked SRAM would remove the floor's last effect on joules. The lane's line for the served sentence: about 4x at k 0.5 and 2x at k 1 against a 5090 paying its whole latency shadow, 20x to 60x without it, with the price threshold beside it (about USD 0.4 per IGN at a third of the network, 0.13 for a maker who takes the chain). Owed: the W = 8 rows, the Apple curve past 8 GiB.
### 10.5 Invention beyond the four layers (floor lane 5, `docs/analysis/class-v6/floor/invention.md` on class-v6-floor-invention, first KEEP/KILL table 13:4x UK; full by 19:30 UK)
**One lever found, conditional on one measurement: the second DRAM sector.** The identity's only `k = 1` work is the memory's own, and the record forced one 32-byte sector per read while the dataset item is 64 bytes. Reading the whole item (W = 16 words; the read-width branch's class w64, fingerprint 836e56e7d496e980, the verifier 0.630 against 0.604 ms) forces a second sector in the same open row: on the chip model's own inputs (909 pJ activate plus 1,150 pJ of movement per sector) the GDDR7 chip's read goes from 2.0 to 3.2 nJ and E_mem from 0.466 to 0.62 microjoules (+33 percent at an unchanged 166 MH/s), one HBM3 stack 0.321 to 0.40 (+24 percent), and on floor lane 3's wire figure the N2 SRAM die 0.036 to 0.126 microjoules (66x to 19x). So the band {1, 4} words is inside one sector and "the chip pays nothing" is right to 32 bytes and wrong at 64; this is the measured reason W = 8 is free for the chip too (section 10.3) and W = 16 is the only width that costs it. The card is the whole question: the 5090 ran w64 at 71.9 MH/s (-47 percent; the exported kernel at full occupancy, four uint4 loads per item, no request hint) while its own 64-byte probe reached 15.7 G reads a second at its best lane count (-13 percent); 589 GB/s is a third of the stream, so the bind is the request path, not the pins. The 9070 XT (-3 percent) and the M5 Max (0) are free at 64 bytes (measured).
| Candidate | Verdict | Number (5090 knee, GDDR7 chip) | Label |
|---|---|---|---|
| The second sector, W = 16 words (w64) | KEEP rank 1, conditional | outcome A (the exported form, -47 percent of rate): 5.3x, dead; B (the probe's -13 percent): 3.4x; C (a one-request form within 5 percent of the 4-byte rate): 3.0x to 3.2x at zero shadow, 2.7x at k 0.5, 2.0x at k 1; the chip +33 percent whatever the card does; the M5 Max and the 9070 XT pay nothing | card measured at two occupancies (unlocked, no watts); chip modelled |
| The job that decides it | one PC 1 run, under an hour; asked of the hash lane at 13:5x UK, after kit d | the w64 pack plus the w4 control, the race's occupancy variants, a variant with `ld.global.L2::64B` on the first load (a variantSource anchor rewrite), both clocks, nvidia-smi at 1 Hz; the pass line 95 percent of the 4-byte rate, the kill line under it | owed; default if not run by 19:30 UK: outcome A stands, W = 16 stays `never` |
| Tensor tiles | KILL as content; the k band corrected | merchant silicon at a nominal 0.30 to 0.56 pJ per INT8 MAC (MTIA v2 0.51, AI 100 Ultra 0.34, B200 0.44, MI355X 0.56, M4 ANE 0.30 measured; the H800 loop 0.34 measured); against the class's tile (1.5 pJ at the knee) k 0.2 to 0.4, against the dense tile (0.83) 0.4 to 0.7: never above the ALU band; "3x to 30x" needs an INT4 layer at 0.46 V (the VSQ chip reads 0.052 at its nominal 0.67 V); Apple -35 percent at 1,024 tiles kills it | claimed, measured |
| The RT core | KILL | traversal implementation-defined (Vulkan, DXR, the NVIDIA forum 14 Oct 2024, Blender's Cycles); the BVH opaque; dedicated units 4 to 33 nJ per ray against 290 to 750 measured board-level on an RTX 2080: k 0.01 to 0.2 | claimed, measured |
| Channels per dollar, controller plus PHY | KILL as a lever; KEEP as a model correction (rank 2) | no 28 nm GDDR7 PHY exists (Cadence N3 2024, Innosilicon FinFET; the oldest GDDR6 PHY 12 nm): the record's USD 5 M project and 17 M cap row is unbuildable; the floor is 12 nm GDDR6 at USD 20 to 30 M (cap 70 to 100 M) or N5 GDDR7 at USD 50 to 75 M (cap 170 to 250 M); SK hynix ISSCC 2024 PAM3 I/O 1.10 TX plus 0.76 RX pJ per bit measured, E_mem moves under 0.06 | claimed, modelled |
| Proof of latency | KILL | the chain checks values, never time: a 1 s block against a 0.4 microsecond round trip (2.5 M per block); the farm's node is on its LAN; a sequential prefix favours the chip about 4x; no sub-RTT PoW in the literature | measured latency, modelled |
| A drawn address map | KILL | firmware, a few hundred gates; no prefetch on a dependent chain; 0 to the chip, 0.8 to 3.2 percent of spread to the cards; literature null | measured spread |
| The literature since 2023 | nothing outside the identity | new rows: iPollo V2H 0.14 J per MH, about 13x a 5090 on Ethash (vendor; the precedent's top, not 6.8x); Jasminer X16-Q 5.7x; Pinecone R1X unverified; Qubic pivoted to Doge in April 2026; no random-read joule measured on any current DRAM | claimed |
| Own: video decode, row straddle, scratch in DRAM, independent chains per hash | KILL | a decoder-bound card and a ms verifier; 3.1x at -39 percent of rate (dominated); 465 KB sits in L2; the 5090's read rate plateaus over lanes (expected 0 to 5 percent) | measured, modelled |
| Own: W = 32 (two items) | hold for v7 behind rank 1 | the chip bandwidth-bound at 110 MH/s and 1.02 microjoules; 2.4x only if the card held its rate, which at 128 B the pins forbid (-20 percent at best) | modelled |
Nothing reads k above 1 on a measured GPU figure; the second sector reads k about 0.75 to 0.8 on GDDR7 (1.15 nJ of DRAM movement at k = 1 plus the card's fabric share), the highest on the table. Section 10's default (no new lever) is right unless the w64 job passes.
## 8. Unverified and owed
- Lane D's family harness (the family gate's measured coverage: the 10,000-era random stratum and the three corner strata, the lossy cap, width 4, shape 64) lands in `family-gate.md` by 17:00 UK and is layer 4's coverage table; this document's layer 4 cites it where it is named and does not restate it.