Class v6: the 5070 Ti tier measured (a rented pod, stock: class v3 78.69 MH/s at 140.8 W, class v4 78.78 at 224.0 W, the premium 83.2 W, 10.3 pJ per counted op), the per-tier chip edges at stock
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
d94c3e0716
commit
91af4721e2
1 changed files with 4 additions and 2 deletions
|
|
@ -6,13 +6,15 @@
|
|||
|
||||
The honest frame first, from last night's measured close (`docs/analysis/counter-asic-4-research.md` sections 2 and 15.1a). Two chips exist in the model. A **fixed-function chip** wires one hash: the mixer's rounds, the op mix's lane ratios, the read width, the program length and the block shape are silicon. A **GPU-like chip** stores the dataset in commodity DRAM and runs the hour's program on a programmable integer core (the `f = 1` chip of chip-model-v3 section 5, the chip anyone builds); for it every drawn parameter is firmware. The four layers below render the FIRST kind useless on the day a parameter leaves its wired value, and move nothing for the second kind except the size of the core it must carry and the capex that forces. What the second kind keeps is the identity of section 2: at zero shadow premium the card's whole-card energy over the chip's memory energy (3.6x on a 5090 at its knee, measured card, modelled chip), and with the class v4 shadow 2.1x at a core as good as a GPU lane. No rotation changes that; the layers change which chip can be built and how long its tape-out lives.
|
||||
|
||||
| Layer | A fixed-function chip on its release day | A GPU-like chip on its release day | RTX 5090 (measured where stated) | RTX 5070 Ti (DEFAULT until the hash lane's row: the rented 5070's row scaled, approximate) | Apple M5 Max (measured where stated) |
|
||||
| Layer | A fixed-function chip on its release day | A GPU-like chip on its release day | RTX 5090 (measured where stated) | RTX 5070 Ti (MEASURED stock, a rented Vast pod, 10:08 to 10:11 UTC: class v3 78.69 MH/s at 140.8 W, class v4 78.78 at 224.0 W, the premium 83.2 W = 10.3 pJ per counted op, fingerprints equal to the Mac's; no core-lock grid, the host refused -lgc) | Apple M5 Max (measured where stated) |
|
||||
|---|---|---|---|---|---|
|
||||
| 1. Per-era draws of the parameters now fixed by release (mixer round count, op-mix weights, read width, program length, shadow block shape), from chain state | dead at the first era whose draw leaves its wired value: a wired x8 mixer at an x16 era recomputes nothing right; a 4-byte-granularity controller at a 16-byte era moves 4x the bytes; a core sized for 100,000 ops at a 200,000 era runs at half rate; the tape-out lives one era (180 days, layer 3) | firmware; the core sized for the band's top (200,000 ops: about 60 mm^2 of N5 instead of 30, USD 25 to 40 more per chip, modelled); `k` unchanged; capex per MH/s +10 to 20 percent (modelled) | rate: 0 across the band while latency-bound (the ladder's rungs 0 to 2: 0 and -2.7 percent at the 431 W cap, measured); watts: the premium follows N and the mix (6.2 to 11.3 pJ per counted op measured; a shuffle-heavy draw up to 5x per op, so the band excludes it); the verifier +0.2 to +1.5 ms per era draw (measured ladder, estimated mixer) | pending the row; by scaling: holds rate to about 200,000 ops at its cap (the 4070 did), the premium 20 to 30 W | rate -3.3 to -10 points across the N band (measured ladder rungs 0 to 2), 0 for the mixer and the width; the Apple tier sets the band's top (130,000 ops at the 5 percent rule) |
|
||||
| 1. Per-era draws of the parameters now fixed by release (mixer round count, op-mix weights, read width, program length, shadow block shape), from chain state | dead at the first era whose draw leaves its wired value: a wired x8 mixer at an x16 era recomputes nothing right; a 4-byte-granularity controller at a 16-byte era moves 4x the bytes; a core sized for 100,000 ops at a 200,000 era runs at half rate; the tape-out lives one era (180 days, layer 3) | firmware; the core sized for the band's top (200,000 ops: about 60 mm^2 of N5 instead of 30, USD 25 to 40 more per chip, modelled); `k` unchanged; capex per MH/s +10 to 20 percent (modelled) | rate: 0 across the band while latency-bound (the ladder's rungs 0 to 2: 0 and -2.7 percent at the 431 W cap, measured); watts: the premium follows N and the mix (6.2 to 11.3 pJ per counted op measured; a shuffle-heavy draw up to 5x per op, so the band excludes it); the verifier +0.2 to +1.5 ms per era draw (measured ladder, estimated mixer) | rate 0 at class v4 (78.69 to 78.78 MH/s, memory-bound, measured); the class v4 premium 83.2 W at stock (10.3 pJ per counted op, the 5090's 10.8), 224 W under its 300 W limit with no throttle, so the N band's top (200,000 ops) costs about 170 W of premium at stock by the per-op figure and the card's limit binds first (approximate); the knee rows need a host that allows -lgc | rate -3.3 to -10 points across the N band (measured ladder rungs 0 to 2), 0 for the mixer and the width; the Apple tier sets the band's top (130,000 ops at the 5 percent rule) |
|
||||
| 2. The state-derived dataset's size tracks chain-state growth with a floor (class v5's leaves scaled by the state, never below the 1.13.3 schedule) | a chip with fixed memory ages out when the dataset passes it: one HBM3 stack 24 GB, the 5090's board 32 GB; the time-memory curve (chip-model 5.4) says the excess must be recomputed at 6.3 nJ per item against 2.0 per read, so its energy per hash rises with the overflow | the same memory limit; a chip buys DRAM a card cannot (24 GB stacks at USD 200, modelled), so it ages out LAST: the 8, 12 and 16 GB card tiers go first | fine to 32 GB: the 2 GiB genesis dataset plus the state's leaves; at a used chain's 5 GB state (class-v5 section 3) the dataset is capped at the sample size, 2 GiB | 16 GB: fine to the 8 GiB step (year 12 on the schedule); the state floor does not move it | 36 to 128 GB unified: fine to the 16 GiB step; the daily build grows with the size (13 to 30 ms measured at 1 GiB) |
|
||||
| 3. Scheduled family epochs by height, every 180 days by default, no release (the reserve R0 to R8 of spec 1.13.2 unlocking by height, then rotating) | a chip without the family's datapath loses its weight of the mix at the unlock (4 points of 79) or emulates it at the vendor penalty (1.5x to 2.4x per op, measured on the cards); a chip taped out against one family set is a GPU-like chip or dead | pre-wires every family for about USD 4 of N5 silicon (algorithm.md 5.2, modelled); moves the per-joule edge under 10 percent per family | measured family step costs: shfla 1.53x, perm 1.30, mm8 2.43 the add-xor-rotate step; under 1 percent of rate at 4 points | pending; the same families on the same silicon generation | measured: shfla 1.91x, perm 1.13 emulated, mm8 emulated at 1.6x per dot4; under 1 percent of rate at 4 points |
|
||||
| 4. The acceptance floor (c''') and the F8 uniformity test generalised to every era's draw, with a redraw on failure (plus the per-site largest-bucket bound and the value-level bias test as the next class's two tests) | nothing on the chip; it is what makes layers 1 and 3 safe without per-era cryptanalysis | nothing | the generator's attempts per seed (today about 30 at 2.4 percent rejection under (c'''); a redraw costs nothing on a card) | the same | the same |
|
||||
|
||||
Per tier against the stored-dataset chip (GDDR7, 0.466 microjoules per hash, modelled), at stock: the 5090 at 2.26 microjoules (class v3) and 3.36 (class v4) reads 4.9x and 2.2x at `k = 1`; the 5070 Ti at 1.79 and 2.84 microjoules reads 3.8x and 1.9x at `k = 1` (a card whose stock point is nearer its knee); the M5 Max at 0.78 and 1.40 (GPU and DRAM channels) reads 1.7x and 0.9x. The layers do not move these; the operating point and the shadow do.
|
||||
|
||||
What a fully general chip still gets, in one line: the stored-dataset chip with a programmable core at the top of the band keeps 3.6x at zero premium and 2.1x at `k = 1` on a 5090 at its knee (section 2 of the research file), and the four layers cost it about USD 30 to 60 per chip of extra silicon and a project that must be N5-class from the first tape-out (the USD 100 M break-even cap of the mission lane's model).
|
||||
|
||||
## 1. Where the four layers sit in the design as it stands
|
||||
|
|
|
|||
Loading…
Reference in a new issue