Class v6 floor lane 3: section 2.2b, the shadowed SRAM rows re-folded on the k lane's synthesised core (0.11 microjoules of shadow at N3, absolute k about 0.1): 21x at W = 4 and 18x at W = 8 under class v4, 10x at the full shadow, 6x to 7x at the knee; finding 1, the public-claim row and the served sentence carry the range with the record's band marked
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of bed2aebe3 (8e9588de0) for the box mirror master
This commit is contained in:
parent
0478a6c0eb
commit
ee2d3c8606
1 changed files with 41 additions and 10 deletions
|
|
@ -22,8 +22,12 @@ schedule and which served claim.
|
|||
a wordline sensed) does not. At today's read width (`W = 1`, 4 bytes) the die moves about 80 bits per dependent read
|
||||
(address, control, data), not 560: **0.25 nJ per read, 8.3 GH/s per die, 66x the 5090 at zero shadow** (modelled on
|
||||
the record's own wire figure). The record's 17x is the `W = 16` row. With the class v4 shadow at `k = 0.5` the same
|
||||
chip reads 5.7x (the record's 4.8x), at `k = 1` 3.0x: the shadow, not the memory, is the whole hold against this
|
||||
die, and the read-energy band moves the shadowed number by about 0.2x to 0.7x.
|
||||
chip reads 5.7x (the record's 4.8x), at `k = 1` 3.0x. **On the k lane's synthesised core (13:1x UK, section 2.2b:
|
||||
about 0.11 microjoules of shadow per hash at N3, absolute `k` about 0.1, the shuffle row open) the same die reads
|
||||
21x at W = 4 and 18x at W = 8 under class v4, 10x and 9.7x at the honest cards' whole latency shadow, 6x to 7x at
|
||||
the 5090's knee under the full shadow.** On the record's claimed band the shadow is the whole hold and the width
|
||||
moves the shadowed number 0.2x to 0.7x; on the synthesised core the memory is a third to a half of the die's energy
|
||||
and the width is worth 1.3x. The public 2x is not reachable against this die either way.
|
||||
2. **The one lever on the wire is the read width, and it is free on the honest cards up to 16 bytes.** `W = 4` words
|
||||
(16 bytes) is measured within 2.7 percent on the 5090 and the 9070 XT and within 1 percent on the M5 Max (the
|
||||
record, 5 October); it moves the die to 176 bits per read, 0.38 nJ, **44x at zero shadow, 5.6x at k = 0.5**. `W = 8`
|
||||
|
|
@ -70,9 +74,10 @@ schedule and which served claim.
|
|||
stays pinned and the data must travel), and the note that 3D-stacked SRAM (SoIC-class, a vertical hop at about 0.1
|
||||
pJ per bit, approximate) would make a multi-die store read as one die, which weakens the floor on joules further.
|
||||
|
||||
The honest line: against this die the memory lever is worth a factor of two at zero shadow and a tenth with the
|
||||
shadow on; the floor is a ticket; the project cost against the chain's revenue is the wall, and the shadow at the
|
||||
honest card's full latency shadow (2.0x at `k = 1`, 4.0x at `k = 0.5`) is the number the public text can carry.
|
||||
The honest line: against this die the memory lever is worth a factor of two at zero shadow and 1.3x with the shadow
|
||||
on at a synthesised core (a tenth on the record's claimed band); the floor is a ticket; the project cost against the
|
||||
chain's revenue is the wall; and the number the public text can carry is a range, 6x to 10x at the honest cards'
|
||||
whole latency shadow on the synthesised core (2x to 4x on the record's claimed band, marked), never "under 2x".
|
||||
|
||||
## 1. Inputs
|
||||
|
||||
|
|
@ -138,6 +143,32 @@ The `k` lane's RTL rows re-fold these in absolute before 20:00.
|
|||
| 8 | **3.7x / 2.0x** | 5.8x / 3.2x | modelled |
|
||||
| 16 | 3.4x / 1.9x | 5.2x / 3.0x | modelled |
|
||||
|
||||
### 2.2b The re-fold on the k lane's synthesised core (13:1x UK; floor lane 2, `class-v6-floor-k`)
|
||||
|
||||
Floor lane 2's first chip-side figures (ASAP7 routed RTL, Yosys and OpenROAD, a minimal 8-register lane; scaled to
|
||||
N3 by x0.50, approximate): ARX 1.05 to 1.09 pJ per op, mul and mulhi 0.68, mad 1.66, the index fold 1.19, prmt 0.64,
|
||||
lop3 0.72; the shuffle, the crossbar, the scratch and the int8 tile pending. The class v4 mix costs that core about 1.0
|
||||
to 1.2 pJ per op at N3, **0.11 microjoules per hash of shadow at 102,100 ops, independent of the card's operating
|
||||
point**; absolute `k` about 0.10 at stock and 0.18 at the lock (the record's claimed band was 0.3 to 0.8). An N2 core
|
||||
is one node further (x0.7, approximate): 0.08. At the honest cards' whole latency shadow (330,000 ops) the same core
|
||||
pays 0.36 (N3) or 0.25 (N2). The re-fold, `E_chip = E_hash0 + shadow`, against the 5090 at 3.36 (stock, class v4),
|
||||
2.32 (lock), 4.26 (stock, the full shadow), 2.67 (lock, full) and the M5 Max at 1.40 (class v4):
|
||||
|
||||
| `W` | Class v4 shadow, stock / lock, N3 core | The same, N2 core | Full shadow, stock / lock, N3 | The same, N2 | M5 Max class v4, N3 / N2 | Label |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 1 | 23x / 16x | 29x / 20x | 11x / 6.7x | 15x / 9.3x | 9.6x / 12x | modelled on synthesised pJ |
|
||||
| 4 | **21x / 14x** | 25x / 17x | **10x / 6.5x** | 14x / 8.8x | 8.5x / 11x | the same |
|
||||
| 8 | **18x / 12x** | 21x / 15x | **9.7x / 6.1x** | 13x / 8.2x | 7.5x / 8.9x | the same |
|
||||
| 16 | 14x / 9.9x | 16x / 11x | 8.8x / 5.5x | 11x / 7.1x | 6.0x / 6.8x | the same |
|
||||
|
||||
Reading, which replaces "the shadow is the whole hold" of section 0 where the synthesised figure stands: at a
|
||||
measured-class `k` of 0.1 the shadow holds the die to 14x to 23x under class v4 and 6x to 11x at the honest cards'
|
||||
whole latency shadow, not 2x to 6x; the die's memory energy is then a third to a half of its energy per hash, and the
|
||||
read width is worth 1.3x (W = 1 to W = 8) on the shadowed number, not 0.2x. The public 2x is not reachable against
|
||||
this die by any shadow at a synthesised core; what holds it is the project cost against the chain's revenue (section
|
||||
4), which the shadowed edge moves by 1.4x across the whole band. The rows carry the k lane's labels (ASAP7 measured by
|
||||
synthesis, the node scaling approximate, the shuffle open); its own re-fold replaces this table when it lands.
|
||||
|
||||
Sensitivity: the wire at 0.65 pJ per bit halves every wire term (`W = 1` 0.20 nJ, 83x; `W = 8` 0.35 nJ, 49x); at 2.6
|
||||
pJ per bit doubles it (`W = 1` 0.35 nJ, 47x; `W = 8` 0.95 nJ, 18x). The macro at 0.2 nJ adds 0.1 nJ to every row. The
|
||||
shadowed rows move under 0.3x across the whole band.
|
||||
|
|
@ -325,7 +356,7 @@ the detector, not the hash, is the instrument.
|
|||
| The unified-memory SoC | a 16 GB Mac leaves at the year-2 step; 32 GB at the 16 GiB step; every Mac pays -21 to -25 percent of rate from the v6 epoch on (measured to 8 GiB). Against the die at class v4 and `k = 0.5` the M5 Max reads 3.6x to 4.0x | finding 3 of lane B: the reference joule |
|
||||
| A rig | nothing changes until a project pays (section 4); the 180-day rotation and the floor do not move the chip's per-MH/s cost | the issuance trigger and the detector, as the record |
|
||||
| A pool user | the chip fleet that matters is one package; the share-pattern detector is the warning | the detector before the public testnet |
|
||||
| The public claim | "under 2x" is not reachable against the die at any `k` under 1 (2.0x at `k = 1` needs the honest card's whole latency shadow); 4x at `k = 0.5`; 20x to 60x without the shadow. The number a miner can act on is the project's price threshold: USD 0.13 to 2.9 per IGN | section 7's served sentence |
|
||||
| The public claim | "under 2x" is not reachable against the die by any shadow: on the k lane's synthesised core 10x at the honest cards' whole latency shadow at stock, 6x to 7x at the 5090's knee, 18x to 21x under class v4; on the record's claimed band (marked) 2x at `k = 1` and 4x at `k = 0.5`; 20x to 60x without the shadow. The number a miner can act on is the project's price threshold: USD 0.13 to 2.9 per IGN | section 7's served sentence |
|
||||
|
||||
## 7. Recommendation, for the 20:00 close
|
||||
|
||||
|
|
@ -336,10 +367,10 @@ shadow and by 0.2x to 0.4x with the shadow on, and keep the per-lane scratch res
|
|||
11.5 at four, each just under a card tier's room and just over a reticle multiple (3, 5 and 6 dies; USD 1,500, 2,500
|
||||
and 3,000 of silicon), and USD 5,000 is not recommended because it is 20 GiB and retires every card under 32 GB and
|
||||
every Mac under 64 GB whenever it lands. The served claim should say that the strongest chip we can model for 2027
|
||||
to 2028, an SRAM store on 2 nm, would reach about 4x per joule against an RTX 5090 paying its whole latency shadow
|
||||
with a core half as costly as a GPU lane and 2x with one as costly, 20x to 60x without that shadow, and that such a
|
||||
project pays only above about USD 0.4 per IGN at a third of the network or USD 0.13 for a maker who takes the whole
|
||||
chain, which is the pattern the share detector watches for. The floor is a ticket (USD 500 to 3,000 of silicon per
|
||||
to 2028, an SRAM store on 2 nm, would reach about 6x to 10x per joule against an RTX 5090 paying its whole latency
|
||||
shadow with a core at the synthesised cost of a 3 nm lane (2x to 4x on the claimed band the record carried, marked),
|
||||
20x to 60x without that shadow, and that such a project pays only above about USD 0.4 per IGN at a third of the
|
||||
network or USD 0.13 for a maker who takes the whole chain, which is the pattern the share detector watches for. The floor is a ticket (USD 500 to 3,000 of silicon per
|
||||
store across the schedule) and not a per-joule or per-MH/s lever, and the text should not claim otherwise. Nothing in
|
||||
this file moves a served number; the `W = 8` row and the Apple rate curve past 8 GiB are the two measurements owed.
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue