diff --git a/docs/analysis/class-v6/floor/sram-and-floor.md b/docs/analysis/class-v6/floor/sram-and-floor.md index d809e0c90..b9541e1ae 100644 --- a/docs/analysis/class-v6/floor/sram-and-floor.md +++ b/docs/analysis/class-v6/floor/sram-and-floor.md @@ -27,8 +27,11 @@ schedule and which served claim. 2. **The one lever on the wire is the read width, and it is free on the honest cards up to 16 bytes.** `W = 4` words (16 bytes) is measured within 2.7 percent on the 5090 and the 9070 XT and within 1 percent on the M5 Max (the record, 5 October); it moves the die to 176 bits per read, 0.38 nJ, **44x at zero shadow, 5.6x at k = 0.5**. `W = 8` - (32 bytes, exactly the GDDR7 sector the 5090 already fetches) is unmeasured and reads 0.55 nJ, **31x and 5.4x**. - `W = 16` costs the 5090 47 percent of its rate (measured) and is dead. The fold is keyed by the destination + (32 bytes, exactly the GDDR7 sector the 5090 already fetches) has no row of its own but is free by the measured + rows: the 5090's ceiling is about 18 G sectors a second (w16 moved 573 GB/s of sectors at 139.8 MH/s, w64 589 GB/s + at 71.9, read-width.md), an aligned 32-byte read is one sector per load exactly as w16 is, AMD moves its 64-byte line + either way, and the M5 Max read 1.03 of its v2 rate at both w16 and w64; it reads 0.55 nJ, **31x and 5.4x**, and one + PC 1 row confirms it. `W = 16` costs the 5090 47 percent of its rate (two sectors per load, measured) and is dead. The fold is keyed by the destination register through its first word (`verify::fold_words`), so a chip cannot pre-fold a wide read into a narrower table: the raw words must cross the die. A chain of 2, 4 or 8 dependent atoms per step moves nothing: the ratio is per read on both sides. Banking plus a sequencer cannot localise: the chain alternates the 16 load sites and their @@ -42,7 +45,10 @@ schedule and which served claim. under 64 GB. The hop between dies costs 0.04 to 0.15 nJ at the honest widths (UCIe 0.5 pJ per bit, claimed), so a ten-die store still reads 25x to 58x at zero shadow and 5.2x to 5.7x with the shadow; and because every die is powered, the store's USD per MH/s does not rise with the dataset (USD 0.25 to 0.5 at every size). The floor - raises the minimum unit from USD 500 to USD 5,000, which is 2.5 times a 5090's street price. **The schedule that + raises the minimum unit from USD 500 to USD 5,000, which is 2.5 times a 5090's street price. And the floor has a + cost on the honest side that the record read as zero: the 5090 at its knee pays 4, 8 and 10 percent more energy + per hash at 2, 4 and 8 GiB than at 1 GiB (measured, PC 1, 12:58 UK; 1 to 4 percent unlocked), so each step moves + the honest denominator toward the chip by the same share. **The schedule that keeps the most honest cards is lane A's: 5.5 GiB at the v6 epoch (3 reticles, USD 1,500), 8.5 GiB two years on (5 reticles, USD 2,500, the last size inside the 12 GB tier's room), 11.5 GiB at four years (6 reticles, USD 3,000, the last inside the 16 GB tier); each step sits just under a card tier's room and just over a reticle multiple.** @@ -86,6 +92,7 @@ honest card's full latency shadow (2.0x at `k = 1`, 4.0x at `k = 0.5`) is the nu | The card room | 75 percent of card memory, 50 percent of Apple unified; the non-dataset working set 254 to 479 MiB best, 600 to 1,500 worst; the cache 512 MiB from 4 GiB, 1 GiB from 8, 2 GiB from 16, freed after the build in the best case | the standing rule / measured working set | card-lifetime-2026-10-05 section 3; the design document 3.4 | | The card population | by count of the bench table's 32 measured consumer cards: 6 GB 3 percent, 8 GB 22, 10 and 11 GB about 9, 12 GB 22, 16 GB 28, 24 GB and up 16; the fleet's hashrate-weighted census unread | approximate | lane A 4.2a via the design document 3.4 | | The Apple rate cost of the working set | -12 percent at 2 GiB, -20 at 4, -22 at 8 against 1 GiB on the M5 Max | measured, 10:40 UTC | the design document 3.3 | +| The 5090's cost of the working set | against the 1 GiB control (137.65 MH/s at 312.2 W unlocked, 127.39 at 212.6 W at the 1,300 MHz lock): 2 GiB -2.8 percent unlocked and -4.9 at the lock, 4 GiB -3.8 and -11.2, 8 GiB -4.3 and -14.1; MH/W at the lock 0.599, 0.576, 0.553, 0.543 (4, 8 and 10 percent per hash in energy at the knee) | measured, 12:49 to 12:58 UK, PC 1, label ca4-v6 | the hash lane's message of 13:0x UK | | The emission | 3,168,808,781 base units per DAA second, 10^8 base units per IGN (1.0 B IGN a year in years 1 and 2), a 30-day ramp (about 37 M IGN never minted), halving every 63,115,200 DAA s, 80 percent to miners | spec | spec 05 (the economics table, section 2.5) | | The mission lane's break-even model | revenue `s x E2 x p`; cap = `C_proj / s` in years 1 to 2 at `s` = 0.30; its E2 of 4.18 B IGN is not this emission constant's figure (1.57 B to miners), so this file recomputes from the constant and carries both | model | mission/future.md 2.2 | | The re-fill | a 1 GiB rebuild 157 G ops, 13.4 ms on a 5090, 0.16 to 0.5 J on a chip core; class v5 refreshes per epoch (3,600 s) | measured / modelled | the invention lane 2.11, 2.15; the research file row 7 | @@ -106,7 +113,7 @@ rate. With the shadow: `edge = (2.26 + F) / (E_hash0 + k F)`. All modelled; the |---|---|---|---|---|---|---|---|---|---|---|---| | 1 (today, 4 B) | 80 | 0.25 | 8.3 | 0.036 | **66x** | 22x | 5.7x / 3.0x | 4.1x / 2.1x | 4.0x | 0 | modelled | | 4 (16 B) | 176 | 0.38 | 5.6 | 0.054 | **44x** | 14x | 5.6x / 2.9x | 4.0x / 2.1x | 3.9x | 5090 +2.7 percent rate, 9070 XT -1.4, M5 Max within 1 (measured) | modelled chip, measured cards | -| 8 (32 B, the GDDR7 sector) | 304 | 0.55 | 3.9 | 0.078 | **31x** | 10x | 5.4x / 2.9x | 4.0x / 2.0x | 3.6x | unmeasured; the sector is fetched anyway on NVIDIA (32 B) and AMD (64 B line); Apple's LPDDR5X atom 32 B (claimed); a PC 1 and M5 Max job | modelled | +| 8 (32 B, the GDDR7 sector) | 304 | 0.55 | 3.9 | 0.078 | **31x** | 10x | 5.4x / 2.9x | 4.0x / 2.0x | 3.6x | free by the measured rows: one sector per load on NVIDIA as at w16 (the ceiling about 18 G sectors a second: 573 GB/s at w16, 589 at w64), one 64 B line on AMD, and the M5 Max at 1.03 of v2 at both w16 and w64; a PC 1 row confirms (the crate's width set is {1, 4, 16} words, so a pack at 8 needs a generator and emitter line first, the hash lane's caveat) | modelled chip; the card by interpolation of measured rows | | 16 (64 B, the record's atom) | 560 | 0.88 | 2.4 | 0.125 | 19x (the record's 17x at 1.0 nJ) | 6.2x | 5.0x / 2.7x | 3.8x / 2.0x | 3.2x | the 5090 -47 percent of rate (measured): dead | modelled | Reading: the record's 1.0 nJ and 17x are the 64-byte row; at the hash's own width the die is 3.5x stronger at zero @@ -198,6 +205,16 @@ die makes 2,000 to 8,000 MH/s, so the store's silicon is USD 33 to 125 per card- the maker powers every die and the rate scales with the count (section 2.4). The floor does not move USD per MH/s; it moves the minimum ticket. +The honest side pays the floor in joules, which the record read as zero for the NVIDIA tiers: the 5090 at the 1,300 +MHz lock loses 4.9, 11.2 and 14.1 percent of rate at 2, 4 and 8 GiB against 1 GiB (unlocked 2.8, 3.8 and 4.3), which +is 4, 8 and 10 percent more energy per hash at the knee (measured, PC 1, 12:49 to 12:58 UK, the hash lane's ca4-v6 +rows: the page-walk cost the unlocked card hides). So the 5.5 GiB step costs a tuned 5090 about 9 percent per hash and +the 8.5 and 11.5 steps 10 to 11 (the curve past 8 GiB unmeasured, read as flattening), against the chip's ticket rising +from USD 500 to 1,500, 2,500 and 3,000 and its joules not at all; every chip edge in section 2 rises by the same 4 to 11 +percent against a card at its knee. The 9070 XT and the smaller NVIDIA tiers have no size rows (approximate: the +5090's curve). Per step the floor buys the chain USD 1,000 of chip ticket for about 1 percent of the tuned 5090's +energy per hash and 22 percent of today's measured consumer cards by count. + ### 3.3 The growth rule, tied to chain state Layer 2's rule stands (the design document 3.1): the day's item count is the larger of the schedule's floor and the @@ -288,7 +305,7 @@ the detector, not the hash, is the instrument. | Home miner, one 8 GB card | Lever 1 costs it nothing at `W = 4` (measured on larger cards; its own row unmeasured) and under 5 percent at `W = 8` by the sector argument (unmeasured). Lever 2 retires it at the year-2 step (8.5 GiB) on this file's schedule, as on lane A's; a 6 GiB v6 epoch retires it two years earlier. Lever 3 is not its concern: no chip exists, and the first one needs IGN at USD 0.13 to 2.9 | the schedule is the founder's number; the `W = 8` measurement is a PC 1 job | | One 12 GB card | holds to the year-4 step (11.5 GiB) with the cache freed after the build; out at 11.0 to 11.5 with it resident | the same | | One 16 GB card (the 9070 XT, the 5070 Ti) | holds every step to 11.5 GiB; out at the 16 GiB step | the same | -| One 24 or 32 GB card, the 5090 at its knee | holds every step of the schedule; the 24 GB tier leaves at 16 GiB with the cache resident; the 32 GB card at 22 GiB. Against the die with the shadow on at the knee (1.66 microjoules, F 0.65 at the lock): about 3.2x at `k = 0.5`, 1.9x at `k = 1` (modelled, the lock rows) | the Ember knob carries the lock | +| One 24 or 32 GB card, the 5090 at its knee | holds every step of the schedule; the 24 GB tier leaves at 16 GiB with the cache resident; the 32 GB card at 22 GiB. It pays the floor in joules at the knee: 4, 8 and 10 percent per hash at 2, 4 and 8 GiB (measured), about 9 to 11 percent across the schedule; unlocked 3 to 4 percent. Against the die with the shadow on at the lock (the card 1.67 microjoules plus F 0.65, measured; the die at `W = 4`): under the record's convention (`k` against the GPU's op cost at the same operating point, 6.2 pJ at the lock) 6.1x at `k = 0.5` and 3.3x at `k = 1`; under the absolute convention (the die's core costs what it costs, `k` fixed against the stock 11.3 pJ) 3.8x and 2.0x. The record's convention flatters the die at the lock by 1.6x; the absolute one is the physics and is what this file carries for the served line | the Ember knob carries the lock; the research lane picks the convention for its tables | | The unified-memory SoC | a 16 GB Mac leaves at the year-2 step; 32 GB at the 16 GiB step; every Mac pays -21 to -25 percent of rate from the v6 epoch on (measured to 8 GiB). Against the die at class v4 and `k = 0.5` the M5 Max reads 3.6x to 4.0x | finding 3 of lane B: the reference joule | | A rig | nothing changes until a project pays (section 4); the 180-day rotation and the floor do not move the chip's per-MH/s cost | the issuance trigger and the detector, as the record | | A pool user | the chip fleet that matters is one package; the share-pattern detector is the warning | the detector before the public testnet | @@ -297,9 +314,9 @@ the detector, not the hash, is the instrument. ## 7. Recommendation, for the 20:00 close The lever is the read width: pin `W = 4` words at genesis on the measured rows (the 5090 +2.7 percent, the 9070 XT --1.4, the M5 Max within 1) and move to `W = 8` once one PC 1 and one M5 Max job read it inside the 5 percent rule, -which moves the die from 66x to 31x to 44x at zero shadow and by 0.2x to 0.4x with the shadow on, and keep the -per-lane scratch resident so the lane stays pinned. The schedule is 5.5 GiB at the v6 epoch, 8.5 two years on and +-1.4, the M5 Max within 1) and move to `W = 8` once one PC 1 row confirms what the measured rows already say (one +sector per load on NVIDIA, one line on AMD, Apple flat to w64), which moves the die from 66x to 31x to 44x at zero +shadow and by 0.2x to 0.4x with the shadow on, and keep the per-lane scratch resident so the lane stays pinned. The schedule is 5.5 GiB at the v6 epoch, 8.5 two years on and 11.5 at four, each just under a card tier's room and just over a reticle multiple (3, 5 and 6 dies; USD 1,500, 2,500 and 3,000 of silicon), and USD 5,000 is not recommended because it is 20 GiB and retires every card under 32 GB and every Mac under 64 GB whenever it lands. The served claim should say that the strongest chip we can model for 2027 @@ -315,11 +332,13 @@ this file moves a served number; the `W = 8` row and the Apple rate curve past 8 - Every chip-side figure is modelled; the macro access (0.1 nJ, width-independent) and the wire (1.3 pJ per bit at 24 mm) are the record's approximations and move every zero-shadow row by 2x either way; the shadowed rows move under 0.3x. -- `W = 8` on the 5090 (stock and at the lock), the 9070 XT and the M5 Max: unmeasured; the sector argument says free - on NVIDIA and AMD and the Apple atom is claimed. A PC 1 job and a Mac measure-lock job; not run by this lane (no - builds on the Mac; no GPU on build-3 or build-4). +- `W = 8` has no row of its own: its card cost is read by interpolation of the measured w16 and w64 rows (one sector + per load on NVIDIA, one line on AMD, the M5 Max flat at 1.03 through w64). One PC 1 row at stock and at the lock + confirms it; the crate's width set is {1, 4, 16} words, so the pack needs a generator and emitter line first (the + hash lane's caveat, 13:0x UK). Not run by this lane (no builds on the Mac; no GPU on build-3 or build-4). - The Apple rate curve past 8 GiB (the 11.5 and 16 GiB steps): unmeasured; read as -22 to -25 percent. -- The 5090's own size rows at 2, 4 and 8 GiB (the hash lane, about 13:45 UK): read as size-independent here. +- The 5090's size rows at 2, 4 and 8 GiB landed at 12:58 UK (the hash lane, PC 1, label ca4-v6) and are in section 3.2; + the 9070 XT and the smaller NVIDIA tiers have no size rows (their page reach is read as the 5090's, approximate). - The SRAM cost curve past N2 (A16 and A14 density and wafer price): approximate; the flat USD 250 per GiB rests on the density gain and the wafer price cancelling, which a 2x move in either changes by 2x. - The emission figures are the spec's constant read by this lane; the mission lane's E2 of 4.18 B IGN differs and is