Class v6: layers 2, 3 and 4 written with their chip rows and per-tier rows, the gate plan (the family analysed as a family) and section 7 (what a fully general chip still gets); layer 1's three open numbers await the hash lane's rows or their defaults at 16:00 UK
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
0467872a63
commit
d94c3e0716
1 changed files with 91 additions and 7 deletions
|
|
@ -39,25 +39,109 @@ PENDING the hash lane's rows (16:00 UK): the x16 mixer verifier and build; the t
|
|||
| Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none |
|
||||
| Program length N | not drawn: the ladder's signal (latency-ladder.md) | an unconditional draw retires the Apple tier at 200,000 (measured -10 percent) | the core sized for the ladder's admissible top (rung 2, 199,600 ops) | the ladder's rows | the ladder's rows | none |
|
||||
|
||||
## 3. Layer 2: the dataset's size tracks the chain state with a floor
|
||||
## 3. Layer 2: the dataset's size tracks the chain state with a floor, so fixed-memory silicon ages out
|
||||
|
||||
PENDING the write-up; the numbers exist: the time-memory curve (chip-model-v3 5.4: every memory system's cheapest point is `f = 1`; a recomputed item costs 6.3 nJ and 9,360 ops against a stored one's 1.2 to 2.0 nJ), the DRAM read cost per the microbench (10.9 nJ per dependent read on the 5090 unlocked, 8.7 at the lock, the whole card's marginal; the memory system's own 2.0 modelled), the state sizes (class-v5-stored-state.md section 3: 93 records today, 61 M on a used chain, the leaves capped at the dataset size), the card-lifetime table (card-lifetime-2026-10-05.md: the 4 GiB step retires 4 GB cards; 8 GB mines to year 12).
|
||||
### 3.1 The rule
|
||||
|
||||
## 4. Layer 3: scheduled family epochs by height
|
||||
Class v5 (`docs/design/class-v5-stored-state.md`) already keys every item to a leaf of the chain's execution state and caps the leaf count at the dataset size (2^24 items at 1 GiB, 2^25 at the designed 2 GiB), sampling the state when it is larger. Layer 2 turns the cap into the size: the day's item count is the larger of the 1.13.3 schedule (2 GiB plus 0.5 GiB a year, the floor) and the state's record count times 64 bytes, rounded up to the power of two the index mapping needs (or the multiply-shift mapping of 1.13.3 option (a), which takes any size), with a genesis-fixed ceiling at the cache growth rule's next doubling so the verifier's lazy derivation stays bounded. Everything below the floor is today's design; everything above it is the chain's own growth deciding the memory a miner must hold.
|
||||
|
||||
PENDING the write-up; the numbers exist: the reserve's chip and card rows (algorithm.md 5.2; counter-asic-3-reserve.md), the family step costs (item 6, measured on the M5 Max, the 5090 and the 9070 XT), the emulation bound of 1.13.2 (8x per op, 5 percent of rate).
|
||||
| State size (records) | Leaves | Dataset under layer 2 | Who holds it | Label |
|
||||
|---|---|---|---|---|
|
||||
| 93 (the devnet today) | 6 KB | the floor: 2 GiB at genesis | every card from 8 GB up | measured state, designed floor |
|
||||
| 10 M (a used chain's accounts alone, class-v5 section 3) | 640 MB | the floor still: 2 GiB (the leaves are a fold into the items, not the items) | the same | arithmetic |
|
||||
| 61 M (10 M accounts, 50 M slots, 1 M code chunks) | 3.9 GB | 4 GiB (the next power of two above the leaves), the year-4 step reached by state instead of by the calendar | 8 GB cards at 75 percent of memory hold 6 GB: the dataset plus the cache (512 MiB) and the scratch fits; 4 GB cards are out | arithmetic on the card-lifetime table |
|
||||
| 250 M (an Ethereum-class state, approximate) | 16 GB | 16 GiB | 24 and 32 GB cards and Apple 64 GB; the 8, 12 and 16 GB tiers out | approximate |
|
||||
|
||||
### 3.2 The chip rows
|
||||
|
||||
The time-memory curve of chip-model-v3 section 5.4 is the whole argument: a stored item costs the chip 1.2 to 2.0 nJ (one random read) and a recomputed item 6.3 nJ and 9,360 ops, so every memory system's cheapest point is `f = 1` and a chip whose DRAM is smaller than the dataset falls back along the curve for the overflow. With the microbench's measured card side beside it:
|
||||
|
||||
| Chip memory | Holds up to | At a 16 GiB dataset | Energy per hash against the 5090's 8.7 nJ per dependent read at the lock (measured, whole card) | Label |
|
||||
|---|---|---|---|---|
|
||||
| GDDR7, 16 devices of 2 GB (the 5090's board without the GPU) | 32 GB | fits | 128 reads x 2.0 nJ = 0.26 microjoules plus static: unchanged | modelled |
|
||||
| One HBM3 stack, 24 GB | 24 GB | fits | 0.32 microjoules: unchanged | modelled |
|
||||
| A 16 GB chip (a cheaper board: 8 devices) | 16 GB | the dataset equals its memory; at 32 GiB it recomputes half: 64 x 2.0 + 64 x 6.3 nJ = 0.53 microjoules against 0.26, the rate halved by the recompute, the chip's capex per MH/s doubled | modelled |
|
||||
| A fixed SRAM mirror (the `f = 0` chip's 256 MiB cache) | the cache only | the cache doubles with the dataset (option C): the 128 mm^2 mirror becomes 255 at year 4, 510 at year 12 (chip-model section 2), at a used chain's state the same by layer 2 in the first year | modelled |
|
||||
|
||||
Reading: layer 2 does not move the chip anyone builds, because a chip buys DRAM a card cannot (a 24 GB HBM3 stack at about USD 200, modelled; a 32 GB GDDR7 board at USD 320) and the card tiers are smaller than the chip's memory at every step. What it does is honest and worth saying plainly: **it retires home cards before chips.** On the card-lifetime table (option (b) steps) the 4 GB tier ends at the 4 GiB step, the 8 GB tier at the 8 GiB step, the 12 and 16 GB tiers at 16 GiB; under layer 2 those steps arrive when the state does, not when the calendar does, so a chain that takes off is a chain whose 8 GB miners leave in its first years. The `f = 0` recompute chip, which nobody builds, is the one chip layer 2 kills outright (its mirror doubles with the state). The floor keeps the schedule's lower bound; a ceiling (the next cache doubling) keeps the verifier's bound.
|
||||
|
||||
### 3.3 Per tier
|
||||
|
||||
| Tier | At the floor (today to year 4) | At a 4 GiB state-driven step | At 16 GiB | Label |
|
||||
|---|---|---|---|---|
|
||||
| RTX 5090, 32 GB | the daily build 13.4 ms at 1 GiB (measured), about 27 ms at 2 GiB; nothing else moves | about 54 ms; nothing else | about 220 ms a day; fits | measured at 1 GiB, scaled |
|
||||
| RTX 5070 Ti, 16 GB | fits | fits (the card-lifetime 16 GB row: room to the 16 GiB step) | OUT (the dataset equals the card) | the table, approximate |
|
||||
| Apple M5 Max, 36 to 128 GB unified | the build 13 to 30 ms (measured) | fits | fits on 64 GB and up at the 50 percent share; a 36 GB machine is out at 16 GiB | measured build, the table's share rule |
|
||||
| 8 and 12 GB cards | fit | 8 GB fits at 75 percent (6 GB usable against 4.5 GB of working set), 12 GB fits | OUT | the table |
|
||||
| A pool user | the pool ships the leaves, 64 B per record (16.5 KB/s to 10,000 members today; the WAN line at about 700,000 records, class-v5 2a.2) | the same | a 16 GB leaf array per member per window: above the WAN line; the pool serves the built dataset or the member builds from the state it holds | class-v5's arithmetic |
|
||||
| A node | one leaf pass per window (100 ns per record: 6 s on one core at 61 M records, 0.2 s on 32 threads) | the same | the same | class-v5's measured rate |
|
||||
|
||||
## 4. Layer 3: scheduled family epochs by height, every 180 days by default, no release
|
||||
|
||||
### 4.1 The rule
|
||||
|
||||
Spec 1.13.2 already unlocks reserve family `n` at the start of era `n` with weight `W_new` taken proportionally from the live families; the eras are the six-month `E_n` of the 1-hour VDF. Layer 3 fixes two things the reserve leaves open: the cadence is a genesis constant (`family_epoch_daa`, 180 days of DAA seconds by default, independent of the era draw so a family can rotate without an era), and the set rotates after the reserve is exhausted (at family epoch `n` past the reserve's end, the live set is the genesis eleven plus the reserve entries whose index is in a fixed cycle over the reserve, so a retired family returns on a fixed schedule and a chip can never wait one out). The weights move by the 1.13.2 rule (proportional). No release carries any of it; the emitter and the verifier hold every family from genesis, with the per-vendor conformance vectors of 1.15 for each.
|
||||
|
||||
### 4.2 The chip rows (the reserve's measured and modelled figures, algorithm.md 5.2 and counter-asic-3-reserve.md)
|
||||
|
||||
| Family | What a chip must add (N5, approximate) | Energy per op at the N5 floor | The honest cards' step cost (measured, ratio to the add-xor-rotate step): Apple / NVIDIA / AMD | What it does to a chip without it on the unlock day |
|
||||
|---|---|---|---|---|
|
||||
| R1 shfla (lane + delta) | a 32-lane crossbar per warp, 3 to 6 adders per lane | 1.00 pJ | 1.91 / 1.53 / 0.75 to 0.84 | at 4 points of 79 the chip without a crossbar emulates through its existing xor-shuffle path or loses 5 percent of the mix's work per op |
|
||||
| R2 perm (byte permute) | a 4x4 byte crossbar, 1 to 2 adders | 0.10 | 1.13 emulated / 1.30 / 1.73 to 1.93 emulated | under 1 percent |
|
||||
| R3 popc and clz | a popcount tree and a priority encoder, 1.5 to 3 adders | 0.10 | 0.87 and 1.01 / 1.50 and 1.63 / 0.91 to 1.30 | under 1 percent |
|
||||
| R4 bfe, R5 shl and shr, R6 sel, R7 andn | 0.1 to 0.3 adders each | 0.02 to 0.06 | 0.75 to 1.54 | 0 |
|
||||
| R8 mm8 (the int8 tile) | a u8 MAC tile per warp, about 100 adders per lane, licensable | 1.60 (the N5 floor); the GPU's own tile measured at 1.5 to 4 pJ per MAC (the research file's 15.1a) | Apple emulated at 1.6x per dot4 (10 steps per tile), NVIDIA 2.43, AMD 1.68 to 1.83 native (layout unverified) | nothing: a chip's MAC array is cheaper than the GPU's (k 0.03 to 0.3, measured GPU side); R8 is kept for datapath diversity, never for joules |
|
||||
| All eight pre-wired | about 8 to 14 adders per lane, about USD 4 of N5 on a 14,000-lane array | | | the GPU-like chip pays USD 4 once and is never surprised |
|
||||
|
||||
Reading: against the `f = 1` chip every family moves the per-joule edge under 10 percent (a family changes 4 points of 79 in the shadow mix at `N x k x 6.4 to 11.3 pJ`), and a chip that pre-wires the reserve pays USD 4; the layer's whole value is against a chip taped out without a family (a fixed-function datapath), which loses the family's share of the work on the unlock day or emulates it at the measured vendor penalties. The rotation after exhaustion closes the one gap the reserve has: a chip that waits for a family to retire.
|
||||
|
||||
### 4.3 Per tier and the known-failed case
|
||||
|
||||
The cost of a live family is its step cost at its weight, checked per vendor at the unlock rehearsal (the 5 percent rule of 1.13.2): at 4 points the worst measured vendor (Apple on shfla, 1.91x per op) pays under 1 percent of its ALU time, which on a latency-bound card is 0 rate; the daily build and the verifier are unmoved (one op per instruction, under 0.01 ms per warp). The 5090 and the 5070 Ti pay the NVIDIA column (1.26 to 1.63x per op on a family's 4 points: 0 rate); the M5 Max pays the Apple column (0 rate at 4 points; the mm8 emulation at 1.6x per dot4 stays under the 8x bound). The known-failed case: a kernel built without the live family refuses at packcheck (the program id carries the family set through the class's allowed list, as `program_id_class` carries the era's `allowed[3]`), and the fast-time harness crosses one family epoch with three nodes and one stale miner, the stale miner's blocks rejected from the first block of the new epoch (the class signal harness's shape, `infra/fast-time/class-v5-signal.mjs`).
|
||||
|
||||
## 5. Layer 4: the acceptance rule and the census as a function of the era draw
|
||||
|
||||
PENDING the write-up; the pieces: (c''') `MIN_DISTINCT_RATIO_V5 = 0.995` (class-v5 section 14: 2.435 percent of candidates under it, attempts +3.4 percent); F8's largest-bucket 6-sigma test on 64-line segments over 2^28 derivations (f8-uniform.md section 7, PASS at +4.84 sigma against the control's +4.18); the per-site bucket bound from this morning's tail attribution (AP-F8-1: p4, p8, p10, p34 as per-site bucket concentration at a narrow-window site); the value-level bias test (adv-cache-2: a product's biased low bits at address bit R; the research file's 20.2b); the duplicate-lane test (the research file's 20.2a, the per-load class's record).
|
||||
### 5.1 The rule
|
||||
|
||||
Every test the generator applies to a candidate program is today a function of the program and the genesis parameters: (a) the stale-source rule, (b) the injecting-write rule, (c) the dynamic test over 64 units on the closed-form stand-in (constant bits, saturation, lane-constant sites, output bias, the distinct-address sum), sub-version 3's (a') freshness fixpoint, (c') saturation per site and (c''') the distinct ratio floor (0.995). Layer 4 makes the rule take the era draw as an input (the mixer multiplier, the weights, the width and the block shape of layer 1; the live family set of layer 3; the dataset size of layer 2) and adds the four tests the night's work named, so that layers 1 and 3 need no per-era cryptanalysis: every era's programs are drawn against the same tests, and a candidate that fails any is redrawn from the next stream values, exactly as today's attempts are.
|
||||
|
||||
| Test | Today | Under layer 4 | The number it rests on | The known-failed case |
|
||||
|---|---|---|---|---|
|
||||
| (c''') the distinct-item ratio floor | 0.995 over the 64 units at the genesis width and mixer | the same floor evaluated with the era's width (a 16-byte load touches one item too) and dataset size; 2.435 percent of candidates under it today, attempts +3.4 percent | class-v5 section 14 (measured census of 4,600 candidates) | a candidate below 0.995 on the exemplar seed 100767 is refused |
|
||||
| F8's uniformity (the largest 64-line bucket within 6 sigma over 2^28 derivations; the top 0.1 percent of items within 1.2x of the window model over 2^24 nonces on 64 seeds) | a gate run by hand per class on the attack board | run by the census tool per era draw at genesis (the band's corners plus 64 random eras) and by the node's acceptance as the per-site version below; a draw whose corner fails is excluded from the band | f8-uniform.md section 7: +4.84 sigma against the control's +4.18, PASS; the tail p4, p8, p10, p34 attributed this morning as per-site bucket concentration at a narrow-window site | the `quarter-lines` and `const-item` plants fire at +75.97 and +92,682 sigma (f8-uniform.md 2.1) |
|
||||
| The per-site largest-bucket bound (this morning's attribution) | named for a next class | per load site, the largest 64-item bucket over the 64 units' addresses within 6 sigma; the narrow-window sites (`k_off` a quarter) judged against their own window's expectation | AP-F8-1's tail: p10 1.50x, p8 1.38x, p34 1.25x, p4 1.22x, each a narrow-window site's bucket | the four tail seeds must be refused; the 60 passing seeds accepted |
|
||||
| The value-level bias test (adv-cache-2; the research file's 20.2b) | named for a next class | per load site, the one-count of every index bit over the 64 units within 6 sigma of n / 2 (the per-load prototype's `BiasedIndexBit`, built last night, measured as the record) | a product's low bits at P(bit 0) = 1/4 placed at address bit R by the stride rotation; 6 of 16 drawn eras over 1.04x, 14 of 17 eras flagged by the instrument on the pre-amendment generator | a program whose site is sourced by a product under an era with R under 28 is refused; the devnet era's R = 29 is not relied on |
|
||||
| The duplicate-lane test (the research file's 20.2a) | built for the per-load class only | kept as a per-load-only test unless a drawn block shape ever places shadow work between loads (layer 1 does not: the block shape is the size, the placement stays after instruction 63) | 1,482 duplicate lanes on the per-load record, 0 to 2 on every sound class | the per-load candidate 0 is refused |
|
||||
|
||||
### 5.2 The cost of the redraw, and its bound
|
||||
|
||||
Attempts per seed are the price. Today's rule under sub-version 3 accepts a candidate at about 30 attempts per seed on the devnet stream (the hash lane's census: 0 of 24,631 seeds exhausted, max attempt 29 against the 256 cap). Each added test raises the rejection rate by its own fraction; the two new tests' fractions on the pre-amendment generator were 4.9 percent (the per-site bound, the tail's four seeds of 64) and up to 14 of 17 eras on the value-level test, which is the one that needs the generator's own fix (the sub-version 3 source rule already refuses a load sourced by a product on the base program; the value-level test catches what the static rule misses). The bound the chain keeps: 256 attempts, the last-resort draw, and the census's exhaustion count per era at genesis (0 of 10^6 is the standing requirement); a band corner whose exhaustion count is not 0 of 10^6 is excluded from the band at genesis, which is the mechanism that lets layer 1 draw without per-era cryptanalysis.
|
||||
|
||||
## 6. The gate plan: the family analysed as a family
|
||||
|
||||
PENDING: the attack board's shape (F1 to F10) and the in-house pass's nine lanes run over the parameter bands, not per era: each row's harness takes the era draw as an input and the census walks the band's corners and 64 random eras; the testnet period as the window; the known-failed test per layer (layer 1: a wired-mixer stand-in at an x16 era reads 0 of 32 lanes; layer 2: a stale-state hasher at a larger state reads 0 of 32; layer 3: a kernel without the live family refuses at packcheck; layer 4: the per-load class's candidate 0 is refused by the generalised rule, and a passing era's census equals the hand census).
|
||||
The attack board (F1 to F10: the shadow's compressibility, the mixer's structure, the cache's recompute, the weak day, the X9 anchor, the verifier, the era draw, uniformity, grinding, the ladder) and the in-house pass's nine lanes (adv-mixer 1 to 3, adv-cache 1 to 3, adv-accept 1 to 3) were each run against one class with its genesis parameters. Under class v6 each harness takes the era draw as an input and the census walks the band:
|
||||
|
||||
| Gate | What runs | The window | The pass line | The known-failed case |
|
||||
|---|---|---|---|---|
|
||||
| G-band (layer 1) | every F row and every adv lane at the band's corners (m in {4, 8, 16} x W in {1, 4} x the weight extremes x the block shapes 64 and 256) plus 64 random eras of the stream | the testnet period, on the pool with --class measure | every row's own line at every corner (F8 1.2x, the 6-sigma buckets, adv-mixer-3's "clean from k = 1", adv-cache-2's hot-set share under 0.1 percent) | the quarter-lines and const-item plants at every corner; a corner that fails any row is out of the band |
|
||||
| G-size (layer 2) | the hot-set and recompute lanes at 2, 4, 8 and 16 GiB with the sample rule; the verifier's lazy bound at each cache doubling | the same | the verifier under 10 ms with the sibling loaded at every size (the ladder's method); the chip curve monotone | a stale-state hasher at a larger state reads 0 of 32 lanes |
|
||||
| G-family (layer 3) | the family-live 5 percent run per vendor at each reserve entry's weight (Metal, CUDA, OpenCL on NVIDIA and AMD), the conformance vectors of 1.15 per family, one family epoch crossed on the fast-time harness with a stale miner | the same | every vendor within 5 percent with the family live; the stale miner's blocks rejected from the first block | the stale kernel refused at packcheck |
|
||||
| G-accept (layer 4) | the generalised rule against the per-load record (candidate 0 refused), the tail's four seeds (refused) and the 60 passing seeds (accepted); the exhaustion census per corner (0 of 10^6) | the same | the hand census equals the rule's verdicts seed for seed | the two plants above |
|
||||
| G-vectors | every vendor's fingerprint equal on one pack per corner (the sub-version 3 pairing method: Metal, CUDA, Apple OpenCL, AMD OpenCL) | before any object | 96 of 96 lanes and the 2^24 fingerprint per corner | a corner's pack refused by a worker on the wrong class |
|
||||
|
||||
Hours (agent, never weeks): the generator's five draws and the band constants 6 to 8; the acceptance rule parameterised plus the two new tests 6 to 8 (the per-load prototype carries the two tests' code); the census tool over the band 4; the family cadence and rotation in the fork 6 to 8 (the class signal's shape); the fast-time harnesses 4 per layer; the attack board re-run over the band is pool time, about 20 corners x the board's 3 to 6 hours each, pipelined on both boxes over the testnet period.
|
||||
|
||||
## 7. What a fully general chip still gets
|
||||
|
||||
PENDING: the identity's numbers per layer (the research file's section 2), the capex rows (section 16), the honest line.
|
||||
The identity of the research file's section 2, with the night's measured rows: `edge = (E_card + F) / (E_mem + k F)`. None of the four layers enters it except through `F` (the premium the card pays) and `k` (the chip core's cost per op over the card's), and the layers move neither for a GPU-like chip:
|
||||
|
||||
| Layer | Its effect on a GPU-like chip's edge | Its effect on that chip's capex | The honest line |
|
||||
|---|---|---|---|
|
||||
| 1 | none on `k` (firmware); `F` moves with the draw and the card pays it first; the core sized for the band's top | +USD 25 to 40 of N5 per chip; the project N5-class from the first tape-out: USD 30 M, a break-even cap of about USD 100 M (the mission lane's model) | a chip that is a GPU's memory system plus a programmable core is drawn against by nothing here |
|
||||
| 2 | none while the dataset fits its DRAM (24 to 32 GB against the card tiers' 8 to 32) | none until 16 GiB; then +USD 200 per HBM3 stack | retires cards before chips |
|
||||
| 3 | under 10 percent per family; USD 4 pre-wired | +USD 4 | a datapath taped out against one family set dies; a general one does not |
|
||||
| 4 | none | none | the safety of 1 and 3, not a lever |
|
||||
| All four, against a 5090 at its knee | **3.6x at zero premium, 2.1x at `k = 1` with the class v4 shadow (measured card, modelled chip), unchanged** | about +USD 30 to 60 per chip; the N5 project forced | "useless as soon as it dropped" is true of a fixed-function ASIC and false of the chip anyone builds; what holds the general chip is the price per joule of the honest card's own operating point and the shadow's premium, as last night's close said |
|
||||
|
||||
## 8. Unverified and owed
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue