Class v6: floor lane 4's final table (the class v5 stock column measured on 14 card classes; the lock worth 34 to 41 percent on Blackwell and Ada; the 16 GB Blackwell knee the honest floor); layer 1's N row carries the measured price of a 4x shadow (1.58x the energy per hash on the 5090 and 4090)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
22a3ba3c50
commit
b114f24599
1 changed files with 28 additions and 3 deletions
|
|
@ -59,7 +59,7 @@ The band a card sees across the draw's range, from the hash lane's cost rows (12
|
|||
| Read width `W` | **pinned at 4 words (16 bytes) at genesis, not drawn** (floor lane 3, 10.3: the width is the only wire lever on the SRAM die, 66x at 4 bytes against 44x at 16 at zero shadow; 8 words after the owed PC 1 and Mac rows; 16 never) | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15, the M5 Max within 1 percent; w64 bandwidth-bound (71.9 MH/s on the 5090) | the SRAM die's energy per read rises with the bits moved (0.25 to 0.38 nJ); the DRAM chip's toward the card's | 0 / 0 / 0 (measured) | 0 | the W = 8 rows (owed). The acceptance side of W = 16, for the record (the census lane, `docs/analysis/class-v6/census-w16-mix.md` on class-v6-census at d20eb04bd, 14:1x UK, measured): 256 of 256 accepted without an era at r 0.704 (the control 0.668) and 256 of 256 across eras 0 to 7, 0 instrument refusals, (c''') 1.4 percent, F8-form 0.999 to 1.006 of uniform, the verifier +0.6 percent; and the structural note that at W = 16 the product's low bits 0 to 3 are alignment and never enter the address (the control carries bit 0 biased in 116 of 256 programs, W = 4 moves it to bit 2 in 74 to 75) while the era-stride bit R stays on every width until the index fold. So W = 16 is clean on acceptance and dead on energy (10.5): the width is decided by the card's second sector, not by the generator |
|
||||
| Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none |
|
||||
| **The index fold (a ring-A design rule of layer 1, every era's draw passes through it): `load_index` folds a product's low bits before the stride rotation, so no era's R lands a biased product bit on an address bit** | every era; lane D's coverage (11:3x UK): the class is at HALF the family's epochs, not a corner: 52 to 58 percent of accepted programs in every stratum carry one site whose address bit R (or R+1, R+2) is biased over 6 sigma at 2^20, 33 to 40 percent over 100 sigma, the worst z 1,024, on programs (c''') passes; the F8 tail's bucket excess is the same mechanism at scale (+94 to +128 sigma at the biased bit) | lane D's family harness (build-1, 10:21 UTC, 16 drawn eras, every layer-1 parameter from the era's stream, the chain draw through the real rule, reads at the rule's own 2^20 sample): the index-bit bias fires HARD on 7 of 16 eras, abs z 130 to 511 at one site, every one at address bit R or R+1 (a product's bit 0 at P = 0.25 on era 15, R = 25, z -511; a product's bit 1 at 3/8 on era 13; an or-shaped source at 5/8 on era 6); the other 9 clean under abs z 3.8; the (c''') ratio sees none of it (0.9954 to 1.0000): adv-cache-2's era-stride class measured on class v5 accepted programs at the acceptance's own sample | a chip holding the favoured half of that site's window serves 75 percent of its reads instead of 50: about 1.6 percent of a hash's reads at f = 1/2 for one site, zero at f = 1 (the partial store already costs 1.26x the ops, chip-model 5.4): the f = 1 verdict does not move; the row is an auditor's flag on "uniform random reads", not a chip lever | nothing: the fold is one xor-rotate on the address path, measured as 0 on every card by the era-layout rows (the index form is the era draw's own) | 0 | the fold's form in `load_index` (fold the product's low bits before the rotation) and its vectors; a 6-sigma REFUSAL in layer 4 is not the lever: it would redraw about 40 percent of epochs (7 of 16 eras); the known-failed test is lane D's 7 of 16 eras at the rule's sample, which must read 0 of 16 with the fold |
|
||||
| Program length N | not drawn: the ladder's signal (latency-ladder.md) | an unconditional draw retires the Apple tier at 200,000 (measured -10 percent) | the core sized for the ladder's admissible top (rung 2, 199,600 ops) | the ladder's rows | the ladder's rows | none |
|
||||
| Program length N | not drawn: the ladder's signal (latency-ladder.md) | an unconditional draw retires the Apple tier at 200,000 (measured -10 percent). **The price of a heavier shadow, MEASURED (the founder's 1.5x test, the 1p5x-knobs lane on a rented 5090 and 4090 at stock, 14:28 to 14:39 UK, 250 x 2^24, nvidia-smi 1 Hz, one fingerprint on both cards): class v5 genesis (55,296 shadow ops per hash) against a 1,024-instruction block at 27 passes on the same state (hl-k3-sh1024, generator 5, 221,184 ops): the 5090 140.83 MH/s at 442.7 W (3.14 microjoules) against 111.07 at 551.1 W (4.96; 110.0 at the 575 W limit over 60 s, 5.23); the 4090 62.41 at 279.3 W (4.48) against 62.64 at 439.2 W (7.01): 4x the shadow instructions cost 1.58x the energy per hash on both architectures (0.40x per instruction against the 256-block, the per-pass overhead amortised), power-bound on the 5090 (21 percent fewer MH/s), rate-neutral on the 4090 (+61 percent watts). This replaces the hash lane's modelled +290 W at N = 200,000 with a measured point** | the core sized for the ladder's admissible top (rung 2, 199,600 ops); the chip pays k of the heavier shadow, the card all of it | the ladder's rows; at 4x the shadow the 5090 +1.58x per hash (measured) | the ladder's rows | none |
|
||||
|
||||
## 3. Layer 2: the dataset's size tracks the chain state with a floor, so fixed-memory silicon ages out
|
||||
|
||||
|
|
@ -301,7 +301,7 @@ The table, filled with the defaults, each cell replaced as its lane's row lands
|
|||
| RTX 4090 stock (5.005; class v3 3.588; rented pod, 10.1) | 6.1x / 8.7x | 7.4x / 11.6x | 12x / 31x | no change (measured: 16 of 128 SMs saves 4 W) | not applied |
|
||||
| H100 HBM3 at its 700 W cap (2.776, throttled to 1,750 MHz; class v3 1.769; rented pod, 10.1) | 3.4x / 4.8x | 4.1x / 6.4x | 6.8x / 17x | no change (measured: per-SM throughput-bound, 1.4 percent) | not applied |
|
||||
|
||||
The honest floor in one line, on the synthesised core: against the chip anyone can build (the GDDR7 board) the strongest honest tier by joule, the M5 Max, holds the chip to 1.7x, the 16 GB Blackwell card at its knee (the 5080 measured, the 5070 Ti and 5070 modelled; floor lane 4's finding that the honest NVIDIA floor is not the 5090) to 2.5x and a 5090 at its knee to 2.8x; against the chip that needs an N2 project (the SRAM die at the genesis width) the holds are 3.4x, 5.0x and 5.7x; every cell's floor-k figure is the worst case if a chip's core costs no more than its bare units.
|
||||
The honest floor in one line, on the synthesised core: against the chip anyone can build (the GDDR7 board) the strongest honest tier by joule, the M5 Max, holds the chip to 1.7x, the 16 GB Blackwell card at its knee (the 5080 measured, the 5070 Ti and 5070 modelled; floor lane 4's final table of 14:4x UK, 10.4, the class v5 stock column measured on 14 card classes: the lock is worth 34 to 41 percent of a Blackwell or Ada card's stock draw, so the honest NVIDIA floor is not the 5090) to 2.5x and a 5090 at its knee to 2.8x; against the chip that needs an N2 project (the SRAM die at the genesis width) the holds are 3.4x, 5.0x and 5.7x; every cell's floor-k figure is the worst case if a chip's core costs no more than its bare units.
|
||||
|
||||
The capex wall's two thresholds (lane 3, 10.3, re-folded 13:18 UK on lane 5's corrected project floor of USD 20 to 75 M for the cheapest DRAM-board chip): no rational chip project of any kind below about USD 23 M a year of miner revenue (IGN 0.03 at launch emission, USD 62 K a day, a cap of about USD 60 M); every DRAM-board project at a third of the network above about USD 340 M a year (IGN 0.44, USD 0.93 M a day, a cap of about USD 0.9 B), where the SRAM project also starts. The project cost moves the threshold 5x, the chip's edge 1.4x.
|
||||
|
||||
|
|
@ -512,7 +512,32 @@ The shadow is most of the hold again (the memory a tenth to a fifth of the die's
|
|||
| Apple | M5 Max, no lever (the GPU and DRAM meter) | 1.43 (v4 1.40 measured) | measured | 1.8x / 1.3x | 2.2x / 1.5x | 3.1x / 1.8x |
|
||||
| Apple LPDDR6, five years out | a Max-class SoC on JESD209-6 | 1.18 at the meter (0.27 DRAM, 0.35 GPU, 0.56 shadow) | modelled | 1.5x / 1.06x | 1.8x / 1.2x | 2.5x / 1.5x |
|
||||
|
||||
What it means: against the 16 GB Blackwell card at its knee (2.1 measured on the 5080; 1.7 to 2.1 modelled on the 5070 Ti and 5070) the GDDR7 chip is under 2x at k 1 and the SRAM die 2.7x; against the M5 Max 1.3x and 1.8x; against the LPDDR6 Apple part five years out the GDDR7 chip is at parity at k 1 and the SRAM die 1.5x. The only software lever is the operating point (the lock on Blackwell and Ada, the cap on Ampere, the ADLX offsets on AMD, nothing on Apple or Intel); the occupancy is worth 1 to 2 percent (10.1), the memory clock untouched, the block shape and the cache hint closed. The lane's per-tier scoring rule: size the shadow by the Apple tier's 5 percent point as the 2.0 rule already does and keep N at 100,000 (sizing to the Apple ceiling of 130,000 costs every honest miner 8 to 15 percent of electricity for about 0.15x of edge: the 5090's GDDR7 edge at k 1 goes from 2.13x to 1.95x); do not score the acceptance floor per tier (one network-wide number). The re-measure rule after a class flip: a stored set is stale under a new class, the stored point is the provisional start of the re-measure, max is stock and never stale; the measured v4 to v5 flip moved the 5090's knee by nothing. Two more measured stock rows for the same table (the 1p5x-knobs lane, RunPod secure cloud, 14:28 to 14:39 UK, 250 x 2^24 at one warp per block, nvidia-smi at 1 Hz over the busy window, fingerprints equal to PC 1's; the pods destroyed): class v5 genesis on a rented 5090 (driver 595.91) 140.83 MH/s at 442.7 W = 3.14 microjoules (3.19 over a full 60 s at 448.7 W), on a rented 4090 (driver 570.195) 62.41 at 279.3 W = 4.48; and that lane's heavier-shadow pack hl-k3-sh1024 (not a class v6 value; the 1.5x shadow knob under test there) 111.07 MH/s at 551.1 W = 4.96 on the 5090 (power-bound at the card's 575 W limit over 60 s: 110.0 at 575.0 W = 5.23) and 62.64 at 439.2 W = 7.01 on the 4090: **1.58x and 1.57x the energy per hash of genesis on the two architectures, the same ratio, shown as 21 percent fewer MH/s at the 5090's cap and as 61 percent more watts at a held rate on the 4090.** That is the measured price of a shadow 1.5x heavier on the honest card, which is lane 4's reason for keeping N at 100,000: the card pays the whole of it and the chip pays k of it.
|
||||
What it means: against the 16 GB Blackwell card at its knee (2.1 measured on the 5080; 1.7 to 2.1 modelled on the 5070 Ti and 5070) the GDDR7 chip is under 2x at k 1 and the SRAM die 2.7x; against the M5 Max 1.3x and 1.8x; against the LPDDR6 Apple part five years out the GDDR7 chip is at parity at k 1 and the SRAM die 1.5x. The only software lever is the operating point (the lock on Blackwell and Ada, the cap on Ampere, the ADLX offsets on AMD, nothing on Apple or Intel); the occupancy is worth 1 to 2 percent (10.1), the memory clock untouched, the block shape and the cache hint closed. The lane's per-tier scoring rule: size the shadow by the Apple tier's 5 percent point as the 2.0 rule already does and keep N at 100,000 (sizing to the Apple ceiling of 130,000 costs every honest miner 8 to 15 percent of electricity for about 0.15x of edge: the 5090's GDDR7 edge at k 1 goes from 2.13x to 1.95x); do not score the acceptance floor per tier (one network-wide number). The re-measure rule after a class flip: a stored set is stale under a new class, the stored point is the provisional start of the re-measure, max is stock and never stale; the measured v4 to v5 flip moved the 5090's knee by nothing. **The final table (the lane's fc265d8d, 14:4x UK, ahead of its 15:15 clock; section 2 of its file carries every row with the card's joules and the chip's E_mem separate so any k folds: edge = card microjoules over E_mem + 101,170 x c; E_mem GDDR7 0.466, HBM3 0.321, N2 SRAM 0.14, the HBM4E base die 0.18). The class v5 STOCK column is now measured on 14 card classes (the rented sweep, 13:4x to 14:3x UK, 20 rows, every host refusing -lgc; the two that took -lmc left the memory clock where it was). Floor = the knee with the knobs; stock = unlocked at 100 percent. The chip columns at 1.1 pJ (the k lane's unit floor) / 3.2 pJ (k 0.5, within rounding of the synthesised core's 3.5) / 6.4 pJ (k 1) per forced op:**
|
||||
|
||||
| Card | Floor microjoules (label) | Stock microjoules (label) | GDDR7 | HBM3 | SRAM N2 |
|
||||
|---|---|---|---|---|---|
|
||||
| 5090, 1,200 MHz | 2.33 (measured) | 3.48 (measured) | 4.0 / 3.0 / 2.1x | 5.4 / 3.6 / 2.4x | 9.3 / 5.0 / 3.0x |
|
||||
| 5080, 1,100 MHz | 2.06 (measured) | 3.48 (measured, rented and PC within 3 percent) | 3.6 / 2.6 / 1.9x | 4.8 / 3.2 / 2.1x | 8.2 / 4.4 / 2.6x |
|
||||
| 5070 Ti | 1.70 (modelled) | 2.84 (class v4 measured) | 2.9 / 2.2 / 1.5x | 3.9 / 2.6 / 1.8x | 6.8 / 3.7 / 2.2x |
|
||||
| 5070 | 1.75 (modelled) | 2.99 (class v4 measured) | 3.0 / 2.2 / 1.6x | 4.0 / 2.7 / 1.8x | 7.0 / 3.8 / 2.2x |
|
||||
| 5060 Ti | 2.36 (modelled) | 4.02 (measured) | 4.1 / 3.0 / 2.1x | 5.5 / 3.7 / 2.4x | 9.4 / 5.1 / 3.0x |
|
||||
| 5060 | 2.21 (modelled) | 3.76 (measured) | 3.8 / 2.8 / 2.0x | 5.1 / 3.4 / 2.3x | 8.8 / 4.8 / 2.8x |
|
||||
| 4090 | 3.58 (modelled) | 5.0 (measured, lane 1) | 6.2 / 4.5 / 3.2x | 8.3 / 5.6 / 3.7x | 14.2 / 7.7 / 4.5x |
|
||||
| 4080 | 3.51 (modelled) | 4.92 (measured) | 6.1 / 4.4 / 3.2x | 8.1 / 5.4 / 3.6x | 14.0 / 7.6 / 4.5x |
|
||||
| 4070, 1,860 MHz plus a 50 percent cap | 3.58 (measured) | 5.82 (measured) | 6.2 / 4.5 / 3.2x | 8.3 / 5.6 / 3.7x | 14.2 / 7.7 / 4.5x |
|
||||
| 4060 Ti | 3.81 (modelled) | 5.32 (measured) | 6.6 / 4.8 / 3.4x | 8.8 / 5.9 / 3.9x | 15.1 / 8.2 / 4.8x |
|
||||
| 3090 | 4.51 (modelled) | 5.03 (measured) | 7.8 / 5.7 / 4.1x | 10.4 / 7.0 / 4.7x | 17.9 / 9.7 / 5.7x |
|
||||
| 3080 | 4.2 (modelled) | 4.54 (modelled: both hosts capped; a 170 W cap took a third of the class v5 rate) | 7.3 / 5.3 / 3.8x | 9.7 / 6.5 / 4.3x | 16.7 / 9.1 / 5.3x |
|
||||
| 3060 | 5.77 (modelled) | 6.40 (measured) | 10.0 / 7.3 / 5.2x | 13.3 / 8.9 / 6.0x | 23.0 / 12.4 / 7.3x |
|
||||
| 9070 XT, ADLX -500 MHz, -30 percent | 7.9 (measured) | 10.7 (measured) | 13.7 / 10.0 / 7.1x | 18.3 / 12.3 / 8.2x | 31.4 / 17.0 / 10.0x |
|
||||
| H100 | 2.0 (modelled) | 2.58 (measured) | 3.5 / 2.5 / 1.8x | 4.6 / 3.1 / 2.1x | 8.0 / 4.3 / 2.5x |
|
||||
| A100 | 2.89 (modelled) | 2.99 (measured) | 5.0 / 3.7 / 2.6x | 6.7 / 4.5 / 3.0x | 11.5 / 6.2 / 3.7x |
|
||||
| M5 Max (the meter) | 1.40 (measured) | 1.40 | 2.4 / 1.8 / 1.3x | 3.2 / 2.2 / 1.4x | 5.6 / 3.0 / 1.8x |
|
||||
| Apple LPDDR6 Max, five years out | 1.18 (modelled) | | 2.0 / 1.5 / 1.06x | 2.7 / 1.8 / 1.2x | 4.7 / 2.5 / 1.5x |
|
||||
|
||||
The reading: the lock is worth 34 to 41 percent of a Blackwell or Ada card's class v5 stock draw for under 2 percent of rate (the 5080 3.48 to 2.06 measured, the 4070 5.82 to 3.58 measured), so the stock column and the floor column are different cards; Ampere has only the cap, and a cap under the shadow costs rate one for one (measured twice on the 3080); AMD 24 percent; Apple and Intel nothing; the occupancy 1 to 2 percent (10.1). The honest NVIDIA floor stays the 16 GB Blackwell card at its knee (2.06 measured): under 2x from the GDDR7 chip at k 1 and 2.6x at k 0.5; the pessimistic column at the unit floor reads 3.6x on the 5080 and 2.4x on the M5 Max. Owed and labelled: every Ada and Ampere knee (no rented host allows the lock); the 5070 Ti, 5070 and 3070 class v5 stock rows (their hosts never took the key or dropped mid-bench; the class v4 rows stand within 2 percent).
|
||||
|
||||
Two more measured stock rows beside it (the 1p5x-knobs lane, RunPod secure cloud, 14:28 to 14:39 UK, 250 x 2^24 at one warp per block, nvidia-smi at 1 Hz over the busy window, fingerprints equal to PC 1's; the pods destroyed): class v5 genesis on a rented 5090 (driver 595.91) 140.83 MH/s at 442.7 W = 3.14 microjoules (3.19 over a full 60 s at 448.7 W), on a rented 4090 (driver 570.195) 62.41 at 279.3 W = 4.48; and that lane's heavier-shadow pack hl-k3-sh1024 (not a class v6 value; the 1.5x shadow knob under test there) 111.07 MH/s at 551.1 W = 4.96 on the 5090 (power-bound at the card's 575 W limit over 60 s: 110.0 at 575.0 W = 5.23) and 62.64 at 439.2 W = 7.01 on the 4090: **1.58x and 1.57x the energy per hash of genesis on the two architectures, the same ratio, shown as 21 percent fewer MH/s at the 5090's cap and as 61 percent more watts at a held rate on the 4090.** That is the measured price of a shadow 1.5x heavier on the honest card, which is lane 4's reason for keeping N at 100,000: the card pays the whole of it and the chip pays k of it.
|
||||
|
||||
Owed and labelled: every Ada and Ampere knee is modelled because no rented host allows -lgc (14 pods today, every one refused); the rented stock rows are measured; the sweep adds class v5 watts measured on 16 card classes and a memory-clock try on each.
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue