Class v6: kit b rows (the mixer draw costs the card nothing, measured; the shuffle-heavy base table 7 W under stock; the 5090 at the knee pays 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB, replacing size-independent); floor lane 3's W = 8 and k-convention update
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
78457959fb
commit
91cb431a92
1 changed files with 7 additions and 7 deletions
|
|
@ -34,7 +34,7 @@ The candidate schedule main named: a floor of 6 GiB at the class v6 epoch, 10 Gi
|
|||
| Apple, 32 GB unified (about 16 GiB) | 32 | 7.4 / 11.9 / 15.9 | holds every step of the candidate schedule, 15.9 of about 16 at the 14 GiB step with no headroom | 28 | the same measured curve | unread | modelled memory, measured rate |
|
||||
| Apple, 64 GB and up (about 32 GiB) | 64 | fits | holds every step | 60 | the same measured curve (the M5 Max itself) | unread | modelled memory, measured rate |
|
||||
|
||||
Lane A's check of these steps under the standing 75 percent rule (section 3.4): 6 GiB does not fit the 8 GB tier (78 to 82 percent of the card), 10 GiB retires the 12 GB tier and the 16 GB Mac at year 2, 14 GiB retires the 16 GB tier at year 4; the schedule that drops the tiers in the order the note described is 5.5 / 8 / 11 GiB, each step retiring about a quarter of today's measured consumer cards by count. The Apple column is the one the founder decides on: the M5 Max is the honest best per joule (0.78 microjoules on its GPU and DRAM channels, 3.1x the 5090) and the Mac tier is a large audience; a 16 GB Mac is out at the 10 GiB step and every Mac pays the measured rate cost of a larger working set (-12 to -22 percent) before any memory limit, which the NVIDIA cards do not pay in this range by the hash lane's reading (their 2 MiB pages; the 5090 size rows at about 14:00 UK measure it). The share of today's measured cards per tier is unread until the fleet's census is sent; the bench table's rows name the cards, not their count.
|
||||
Lane A's check of these steps under the standing 75 percent rule (section 3.4): 6 GiB does not fit the 8 GB tier (78 to 82 percent of the card), 10 GiB retires the 12 GB tier and the 16 GB Mac at year 2, 14 GiB retires the 16 GB tier at year 4; the schedule that drops the tiers in the order the note described is 5.5 / 8 / 11 GiB, each step retiring about a quarter of today's measured consumer cards by count. The Apple column is the one the founder decides on: the M5 Max is the honest best per joule (0.78 microjoules on its GPU and DRAM channels, 3.1x the 5090) and the Mac tier is a large audience; a 16 GB Mac is out at the 10 GiB step and every Mac pays the measured rate cost of a larger working set (-12 to -22 percent) before any memory limit, which the NVIDIA cards pay too at their knee by the 12:5x UK measurement (the 5090 at the 1,300 lock: -4.9 / -11.2 / -14.1 percent of rate at 2 / 4 / 8 GiB, 4 / 8 / 10 percent more energy per hash; unlocked -2.8 / -3.8 / -4.3; section 3.3). The share of today's measured cards per tier is unread until the fleet's census is sent; the bench table's rows name the cards, not their count.
|
||||
|
||||
## 1. Where the four layers sit in the design as it stands
|
||||
|
||||
|
|
@ -54,8 +54,8 @@ The band a card sees across the draw's range, from the hash lane's cost rows (12
|
|||
|
||||
| Parameter | Band (genesis) | Why that band (the measured rows that set it) | Chip rows: fixed-function / GPU-like | Per tier: 5090 / 5070 Ti / M5 Max | Verifier | Open number |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Mixer applications per round `m` | **{4, 8} at genesis; 16 in the list as `admissible: false` (the ladder's rule: a flag in the genesis list, flipped only by the 90 percent upgrade path once the 2019-class core measures it)** | x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED (today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16). Cold alone: build-1 core 4 x8 4.67 ms per warp, x16 8.69 (the hash lane, 10:13 UTC); build-3 core 4 (AX102, a faster core class) x8 3.48, x16 6.53 (the build-server lane, 11:01 UTC). With the SMT sibling loaded for the row's whole length (the method that counts: a sibling bench looped for the row, not run once, which ends in about a second and leaves the verify phase sibling-idle; the hash lane's earlier 5.41 / 10.11 / 9.04 to 9.10 ms rows were of that kind and are withdrawn): build-3 x8 6.24, x16 11.44 ms, about 1.8x on both classes. The multiplier is 1.86x on the verifier (the mixer IS the verifier's cost); chip-model-v3's estimate (9 ms on a 2019-class core) stands within the cold rows** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); **x16 at 11.4 ms loaded is over the 10 ms gate on the measured row, so `admissible: false` stands on the measurement, not only on the chip model's margin; the O-1.14 laptop run can only tighten it** | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one |
|
||||
| Op-mix weights (the ten non-load families) | **B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr); the lossy families (or, mul, mulhi) capped at their base, the ring-A rule `or + mul + mulhi` at most the table's 18 plus B, which keeps the per-candidate rejection r under 0.85 (r^256 under 1e-18)**; the shuffle weight capped at its class v4 value on the energy side | lane D's coverage run (4,900 drawn eras on three boxes, 11:3x UK, measured): at the lossy corner of B = 4 (or, mul and mulhi all at +4, 30 of 75 lossy against 18) r is 0.956 (the shipped 0.681) and 8 of 663 eras exhaust the 256-attempt cap (mean attempt 18.6, max 252), so 1.2 percent of that corner's epochs would take the last-resort program, which fails rule (a) in 9 percent of seeds (adv-accept-3); the random stratum with every weight drawn reads r 0.718, max attempt 107, 0 exhaustions in 1,444. The energy side (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3; a shuffle-heavy table raises the premium per instruction up to 2x (modelled), a multiply-heavy one 1.1x | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix within the capped band (the cost rows: at N = 100,000 the 5090's +146 W stock becomes up to +160 W multiply-heavy; shuffle-heavy excluded by the cap); the 5070 Ti scaled by 83/146, the M5 Max by 16/146; the two packs' measured rows replace this when they land | +0 ms | the known-failed test: the 8 of 663 exhaustions at the uncapped lossy corner, which the capped band must read as 0 |
|
||||
| Mixer applications per round `m` | **{4, 8} at genesis; 16 in the list as `admissible: false` (the ladder's rule: a flag in the genesis list, flipped only by the 90 percent upgrade path once the 2019-class core measures it)** | **The card pays nothing for the draw, MEASURED (the hash lane's kit b on PC 1, 12:49 to 12:58 UK, generator 2 with the era, the multiplier the only difference, 250 batches per row, fingerprints PASS): x8 137.73 MH/s at 320.2 W unlocked and 127.44 at 212.0 W at the 1,300 lock; x16 137.72 at 320.1 and 127.46 at 211.7; within 0.1 MH/s and 0.5 W at either state, item derivation hidden under the read chain, so the verifier's 1.86x per doubling is the draw's whole cost.** x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED (today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16). Cold alone: build-1 core 4 x8 4.67 ms per warp, x16 8.69 (the hash lane, 10:13 UTC); build-3 core 4 (AX102, a faster core class) x8 3.48, x16 6.53 (the build-server lane, 11:01 UTC). With the SMT sibling loaded for the row's whole length (the method that counts: a sibling bench looped for the row, not run once, which ends in about a second and leaves the verify phase sibling-idle; the hash lane's earlier 5.41 / 10.11 / 9.04 to 9.10 ms rows were of that kind and are withdrawn): build-3 x8 6.24, x16 11.44 ms, about 1.8x on both classes. The multiplier is 1.86x on the verifier (the mixer IS the verifier's cost); chip-model-v3's estimate (9 ms on a 2019-class core) stands within the cold rows** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); **x16 at 11.4 ms loaded is over the 10 ms gate on the measured row, so `admissible: false` stands on the measurement, not only on the chip model's margin; the O-1.14 laptop run can only tighten it** | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one |
|
||||
| Op-mix weights (the ten non-load families) | **B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr); the lossy families (or, mul, mulhi) capped at their base, the ring-A rule `or + mul + mulhi` at most the table's 18 plus B, which keeps the per-candidate rejection r under 0.85 (r^256 under 1e-18)**; the shuffle weight capped at its class v4 value on the energy side | lane D's coverage run (4,900 drawn eras on three boxes, 11:3x UK, measured): at the lossy corner of B = 4 (or, mul and mulhi all at +4, 30 of 75 lossy against 18) r is 0.956 (the shipped 0.681) and 8 of 663 eras exhaust the 256-attempt cap (mean attempt 18.6, max 252), so 1.2 percent of that corner's epochs would take the last-resort program, which fails rule (a) in 9 percent of seeds (adv-accept-3); the random stratum with every weight drawn reads r 0.718, max attempt 107, 0 exhaustions in 1,444. The energy side (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3; a shuffle-heavy table raises the premium per instruction up to 2x (modelled), a multiply-heavy one 1.1x | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix within the capped band (the cost rows: at N = 100,000 the 5090's +146 W stock becomes up to +160 W multiply-heavy; shuffle-heavy excluded by the cap); the 5070 Ti scaled by 83/146, the M5 Max by 16/146; the two packs' measured rows replace this when they land. **First row MEASURED (kit b, PC 1, 12:5x UK): the shuffle-heavy table (shfl 13 of 48 against the stock 5) on class mx8 without a shadow block reads 137.65 MH/s at 312.9 W unlocked (7 W under the stock table's 320.2) and 127.33 at 212.4 at the lock (level): on the 48-op base program the weight move is a 2 percent term either way; the row layer 1 needs is the same two tables inside the shadow block (mx8+sh256x27, kit d, about 13:45 UK), and the microbench arithmetic stays the default until it lands** | +0 ms | the known-failed test: the 8 of 663 exhaustions at the uncapped lossy corner, which the capped band must read as 0 |
|
||||
| Read width `W` | **pinned at 4 words (16 bytes) at genesis, not drawn** (floor lane 3, 10.3: the width is the only wire lever on the SRAM die, 66x at 4 bytes against 44x at 16 at zero shadow; 8 words after the owed PC 1 and Mac rows; 16 never) | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15, the M5 Max within 1 percent; w64 bandwidth-bound (71.9 MH/s on the 5090) | the SRAM die's energy per read rises with the bits moved (0.25 to 0.38 nJ); the DRAM chip's toward the card's | 0 / 0 / 0 (measured) | 0 | the W = 8 rows (owed) |
|
||||
| Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none |
|
||||
| **The index fold (a ring-A design rule of layer 1, every era's draw passes through it): `load_index` folds a product's low bits before the stride rotation, so no era's R lands a biased product bit on an address bit** | every era; lane D's coverage (11:3x UK): the class is at HALF the family's epochs, not a corner: 52 to 58 percent of accepted programs in every stratum carry one site whose address bit R (or R+1, R+2) is biased over 6 sigma at 2^20, 33 to 40 percent over 100 sigma, the worst z 1,024, on programs (c''') passes; the F8 tail's bucket excess is the same mechanism at scale (+94 to +128 sigma at the biased bit) | lane D's family harness (build-1, 10:21 UTC, 16 drawn eras, every layer-1 parameter from the era's stream, the chain draw through the real rule, reads at the rule's own 2^20 sample): the index-bit bias fires HARD on 7 of 16 eras, abs z 130 to 511 at one site, every one at address bit R or R+1 (a product's bit 0 at P = 0.25 on era 15, R = 25, z -511; a product's bit 1 at 3/8 on era 13; an or-shaped source at 5/8 on era 6); the other 9 clean under abs z 3.8; the (c''') ratio sees none of it (0.9954 to 1.0000): adv-cache-2's era-stride class measured on class v5 accepted programs at the acceptance's own sample | a chip holding the favoured half of that site's window serves 75 percent of its reads instead of 50: about 1.6 percent of a hash's reads at f = 1/2 for one site, zero at f = 1 (the partial store already costs 1.26x the ops, chip-model 5.4): the f = 1 verdict does not move; the row is an auditor's flag on "uniform random reads", not a chip lever | nothing: the fold is one xor-rotate on the address path, measured as 0 on every card by the era-layout rows (the index form is the era draw's own) | 0 | the fold's form in `load_index` (fold the product's low bits before the rotation) and its vectors; a 6-sigma REFUSAL in layer 4 is not the lever: it would redraw about 40 percent of epochs (7 of 16 eras); the known-failed test is lane D's 7 of 16 eras at the rule's sample, which must read 0 of 16 with the fold |
|
||||
|
|
@ -89,7 +89,7 @@ Reading: layer 2 does not move the chip anyone builds, because a chip buys DRAM
|
|||
|
||||
### 3.3 Per tier
|
||||
|
||||
The hash lane's VRAM rows (12:0x UK, modelled from the measured 0.4 GiB working set plus about 0.5 GiB of driver and app): the dataset needs 3.2, 5.4 and 9.9 GiB of device memory at the floor, 2x and 4x; a 12 GB card falls off at about 9.5 GiB (year 15 on the 1.13.3 schedule), a 16 GB GPU at about 13.5 GiB (year 23), a 16 GB unified Mac at about 8 GiB (year 12), the 5090 at about 29 GiB (year 54). The DRAM-read cost per hash on the NVIDIA cards is size-independent in this range by the lane's reading (2 MiB pages keep the TLB's reach past 8 GiB; the 5090 size rows at 2, 4 and 8 GiB land at about 14:00 UK and are the measurement). **On the M5 Max it is not size-independent, measured by this lane at 10:40 UTC under the Mac's measure lock (Metal packbench, the hash lane's class v3 packs at 2^28 to 2^31 words, the same seed and era, 3 batches of 2^24, vectors 3 of 3 and fingerprints per pack): 26.48 MH/s at 1 GiB (footprint 1,664 MiB, the build 32 ms), 23.26 at 2 GiB (-12.2 percent; 2,688 MiB; 54 ms), 21.31 at 4 GiB (-19.5 percent; 4,736 MiB; 94 ms), 20.61 at 8 GiB (-22.2 percent; 8,832 MiB; 193 ms).** The Apple GPU's dependent random read costs more time as the working set grows past its page reach (approximate reading: a TLB-reach effect on unified LPDDR5X; the power channels were not sampled this run, so the joules per hash move by at least the rate's share), which is a real per-tier cost of layer 2 that the NVIDIA model does not show: at an 8 GiB floor the Apple tier mines 22 percent slower per card than at 1 GiB, before any memory limit. The table below carries it.
|
||||
The hash lane's VRAM rows (12:0x UK, modelled from the measured 0.4 GiB working set plus about 0.5 GiB of driver and app): the dataset needs 3.2, 5.4 and 9.9 GiB of device memory at the floor, 2x and 4x; a 12 GB card falls off at about 9.5 GiB (year 15 on the 1.13.3 schedule), a 16 GB GPU at about 13.5 GiB (year 23), a 16 GB unified Mac at about 8 GiB (year 12), the 5090 at about 29 GiB (year 54). **The DRAM-read cost per hash on the NVIDIA cards is NOT size-independent at the knee, MEASURED (the hash lane's kit b, PC 1, 12:49 to 12:58 UK, the pinned class v3 program 73bcbfe8 at 2, 4 and 8 GiB against the 1 GiB control 137.65 MH/s at 312.2 W unlocked and 127.39 at 212.6 W at the 1,300 MHz lock; 250 batches per row, fingerprints PASS): unlocked 133.86 at 315.7 W (-2.8 percent), 132.42 at 317.2 (-3.8), 131.75 at 319.2 (-4.3); at the lock 121.14 at 210.3 W (-4.9 percent), 113.06 at 204.5 (-11.2), 109.39 at 201.6 (-14.1); MH/W at the lock 0.599, 0.576, 0.553, 0.543, which is 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB.** The lane's earlier reading (2 MiB pages keep the TLB's reach past 8 GiB) holds unlocked, where the card hides most of the page-walk term in its slack; the latency-bound regime at the lock exposes it. So each step of the schedule costs a tuned 5090 about 4 to 5 percent per hash while the chip's joules do not move (floor lane 3: its ticket goes USD 1,500 / 2,500 / 3,000 at 5.5 / 8.5 / 11.5 GiB), and every chip edge against a card at its knee rises by 4 to 11 percent across the schedule; the honest sentence for the schedule decision is USD 1,000 of chip ticket per step for about 1 to 5 percent of the tuned 5090's energy and about a quarter of today's measured cards by count. **On the M5 Max it is not size-independent, measured by this lane at 10:40 UTC under the Mac's measure lock (Metal packbench, the hash lane's class v3 packs at 2^28 to 2^31 words, the same seed and era, 3 batches of 2^24, vectors 3 of 3 and fingerprints per pack): 26.48 MH/s at 1 GiB (footprint 1,664 MiB, the build 32 ms), 23.26 at 2 GiB (-12.2 percent; 2,688 MiB; 54 ms), 21.31 at 4 GiB (-19.5 percent; 4,736 MiB; 94 ms), 20.61 at 8 GiB (-22.2 percent; 8,832 MiB; 193 ms).** The Apple GPU's dependent random read costs more time as the working set grows past its page reach (approximate reading: a TLB-reach effect on unified LPDDR5X; the power channels were not sampled this run, so the joules per hash move by at least the rate's share), which is a real per-tier cost of layer 2 that the NVIDIA model does not show: at an 8 GiB floor the Apple tier mines 22 percent slower per card than at 1 GiB, before any memory limit. The table below carries it.
|
||||
|
||||
| Tier | At the floor (today to year 4) | At a 4 GiB state-driven step | At 16 GiB | Label |
|
||||
|---|---|---|---|---|
|
||||
|
|
@ -280,12 +280,12 @@ The close (20:00): the honest floor per tier against each chip row, the recommen
|
|||
|---|---|---|---|---|---|---|---|---|
|
||||
| 1 (4 bytes, today) | 80 | 0.25 nJ | 8.3 | **66x** | 5.7x / 3.0x | 4.1x / 2.1x | 0 | modelled chip; measured card |
|
||||
| 4 (16 bytes) | 176 | 0.38 | 5.6 | 44x | 5.6x / 2.9x | 4.0x / 2.1x | the 5090 +2.7 percent, the 9070 XT -1.4, the M5 Max within 1 (measured 5 October) | measured card |
|
||||
| 8 (32 bytes, the GDDR7 sector the 5090 fetches anyway) | 304 | 0.55 | 3.9 | 31x | 5.4x / 2.9x | 4.0x / 2.0x | unmeasured (one PC 1 and one Mac job owed) | modelled |
|
||||
| 8 (32 bytes, the GDDR7 sector the 5090 fetches anyway) | 304 | 0.55 | 3.9 | 31x | 5.4x / 2.9x | 4.0x / 2.0x | free by the measured rows (13:06 UK: the 5090's ceiling about 18 G sectors a second, w16 moved 573 GB/s of sectors at 139.8 MH/s and w64 589 at 71.9; an aligned 32-byte read is one sector per load as w16; AMD moves its 64 B line either way; the M5 Max 1.03 at both w16 and w64); one PC 1 row to confirm, and the crate's width set is {1, 4, 16} words, so a pack at 8 needs a generator and emitter line first | measured card (inferred at 8) |
|
||||
| 16 (64 bytes, lane B's record) | 560 | 0.88 | 2.4 | 19x (the record's 17x) | 5.0x / 2.7x | 3.8x / 2.0x | the 5090 -47 percent: dead | measured card |
|
||||
|
||||
Meaning: at the hash's own width the die is 3.5x stronger than the record said at zero shadow; with the shadow on, every row sits at 5.0x to 5.7x (k 0.5) and 2.7x to 3.0x (k 1): **the shadow is the whole hold, the memory moves it 0.2x to 0.7x.** The fold is dst-keyed (`verify::fold_words`), so a wide read cannot be pre-folded; dependent chains per step leave the ratio unchanged (per-read on both sides); banking plus a sequencer cannot localise (uniform on a window of at least 256 MiB, the 16 sites alternating; moving the lane costs about 290 bits, which caps the wire lever at about W = 8); the per-site window shrink buys the die and the DRAM chip nothing; a hop at the honest widths is 0.04 to 0.15 nJ, so a ten-die store still reads 25x to 58x at zero shadow. This moves layer 1's width row: **pin W = 4 (16 bytes) at genesis (measured free on all three vendors; the die from 66x to 44x), W = 8 after the two owed jobs, never 16.**
|
||||
Meaning: at the hash's own width the die is 3.5x stronger than the record said at zero shadow; with the shadow on, every row sits at 5.0x to 5.7x (k 0.5) and 2.7x to 3.0x (k 1): **the shadow is the whole hold, the memory moves it 0.2x to 0.7x.** The fold is dst-keyed (`verify::fold_words`), so a wide read cannot be pre-folded; dependent chains per step leave the ratio unchanged (per-read on both sides); banking plus a sequencer cannot localise (uniform on a window of at least 256 MiB, the 16 sites alternating; moving the lane costs about 290 bits, which caps the wire lever at about W = 8); the per-site window shrink buys the die and the DRAM chip nothing; a hop at the honest widths is 0.04 to 0.15 nJ, so a ten-die store still reads 25x to 58x at zero shadow. This moves layer 1's width row: **pin W = 4 (16 bytes) at genesis (measured free on all three vendors; the die from 66x to 44x), W = 8 (31x; 5.4x at k 0.5) once the generator and emitter carry it and one PC 1 row confirms, never 16.**
|
||||
|
||||
The floor as a ticket: SRAM is flat at about USD 250 per GiB to 2031 (density +6 to 11 percent per node against dearer wafers; claimed, approximate), so USD 5,000 of silicon per store is 20 GiB at N2 and about 21 GiB in 2031, which retires every card under 32 GB and every Mac under 64 GB; the constraint and "fewest cards" cannot both hold. The per-MH/s does not rise with the floor (every die powered: USD 0.25 to 0.5 per MH/s at any size); the floor raises the minimum ticket only.
|
||||
The floor as a ticket: SRAM is flat at about USD 250 per GiB to 2031 (density +6 to 11 percent per node against dearer wafers; claimed, approximate), so USD 5,000 of silicon per store is 20 GiB at N2 and about 21 GiB in 2031, which retires every card under 32 GB and every Mac under 64 GB; the constraint and "fewest cards" cannot both hold. The per-MH/s does not rise with the floor (every die powered: USD 0.25 to 0.5 per MH/s at any size); the floor raises the minimum ticket only. The honest card's side of the floor is measured (the hash lane's kit b, section 3.3): the 5090 at its knee pays 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB, so every chip row's edge against a card at its knee rises by 4 to 11 percent across the schedule while the die's joules do not move; the schedule's sentence is USD 1,000 of chip ticket per step for 1 to 5 percent of the tuned 5090's energy and about a quarter of today's cards by count. The k convention matters at the lock: the record's (k against the GPU's op cost at the same operating point) reads the SRAM die at 6.1x (k 0.5) and 3.3x (k 1); the absolute convention (the die's core costs what it costs, k against the stock 11.3 pJ per op) reads 3.8x and 2.0x; the record's flatters the die by 1.6x at the lock, so the served line in section 9 uses the absolute one.
|
||||
|
||||
| Floor | Reticles 2026 / 2031 | USD of silicon | Tiers out (worst / best working set) | Label |
|
||||
|---|---|---|---|---|
|
||||
|
|
|
|||
Loading…
Reference in a new issue