Class v6 floor lane 3: section 2.2b, the shadowed SRAM rows re-folded on the k lane's synthesised core (0.11 microjoules of shadow at N3, absolute k about 0.1): 21x at W = 4 and 18x at W = 8 under class v4, 10x at the full shadow, 6x to 7x at the knee; finding 1, the public-claim row and the served sentence carry the range with the record's band marked

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of bed2aebe3 (8e9588de0) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 12:16:25 +00:00
parent 0478a6c0eb
commit ee2d3c8606

View file

@ -22,8 +22,12 @@ schedule and which served claim.
a wordline sensed) does not. At today's read width (`W = 1`, 4 bytes) the die moves about 80 bits per dependent read
(address, control, data), not 560: **0.25 nJ per read, 8.3 GH/s per die, 66x the 5090 at zero shadow** (modelled on
the record's own wire figure). The record's 17x is the `W = 16` row. With the class v4 shadow at `k = 0.5` the same
chip reads 5.7x (the record's 4.8x), at `k = 1` 3.0x: the shadow, not the memory, is the whole hold against this
die, and the read-energy band moves the shadowed number by about 0.2x to 0.7x.
chip reads 5.7x (the record's 4.8x), at `k = 1` 3.0x. **On the k lane's synthesised core (13:1x UK, section 2.2b:
about 0.11 microjoules of shadow per hash at N3, absolute `k` about 0.1, the shuffle row open) the same die reads
21x at W = 4 and 18x at W = 8 under class v4, 10x and 9.7x at the honest cards' whole latency shadow, 6x to 7x at
the 5090's knee under the full shadow.** On the record's claimed band the shadow is the whole hold and the width
moves the shadowed number 0.2x to 0.7x; on the synthesised core the memory is a third to a half of the die's energy
and the width is worth 1.3x. The public 2x is not reachable against this die either way.
2. **The one lever on the wire is the read width, and it is free on the honest cards up to 16 bytes.** `W = 4` words
(16 bytes) is measured within 2.7 percent on the 5090 and the 9070 XT and within 1 percent on the M5 Max (the
record, 5 October); it moves the die to 176 bits per read, 0.38 nJ, **44x at zero shadow, 5.6x at k = 0.5**. `W = 8`
@ -70,9 +74,10 @@ schedule and which served claim.
stays pinned and the data must travel), and the note that 3D-stacked SRAM (SoIC-class, a vertical hop at about 0.1
pJ per bit, approximate) would make a multi-die store read as one die, which weakens the floor on joules further.
The honest line: against this die the memory lever is worth a factor of two at zero shadow and a tenth with the
shadow on; the floor is a ticket; the project cost against the chain's revenue is the wall, and the shadow at the
honest card's full latency shadow (2.0x at `k = 1`, 4.0x at `k = 0.5`) is the number the public text can carry.
The honest line: against this die the memory lever is worth a factor of two at zero shadow and 1.3x with the shadow
on at a synthesised core (a tenth on the record's claimed band); the floor is a ticket; the project cost against the
chain's revenue is the wall; and the number the public text can carry is a range, 6x to 10x at the honest cards'
whole latency shadow on the synthesised core (2x to 4x on the record's claimed band, marked), never "under 2x".
## 1. Inputs
@ -138,6 +143,32 @@ The `k` lane's RTL rows re-fold these in absolute before 20:00.
| 8 | **3.7x / 2.0x** | 5.8x / 3.2x | modelled |
| 16 | 3.4x / 1.9x | 5.2x / 3.0x | modelled |
### 2.2b The re-fold on the k lane's synthesised core (13:1x UK; floor lane 2, `class-v6-floor-k`)
Floor lane 2's first chip-side figures (ASAP7 routed RTL, Yosys and OpenROAD, a minimal 8-register lane; scaled to
N3 by x0.50, approximate): ARX 1.05 to 1.09 pJ per op, mul and mulhi 0.68, mad 1.66, the index fold 1.19, prmt 0.64,
lop3 0.72; the shuffle, the crossbar, the scratch and the int8 tile pending. The class v4 mix costs that core about 1.0
to 1.2 pJ per op at N3, **0.11 microjoules per hash of shadow at 102,100 ops, independent of the card's operating
point**; absolute `k` about 0.10 at stock and 0.18 at the lock (the record's claimed band was 0.3 to 0.8). An N2 core
is one node further (x0.7, approximate): 0.08. At the honest cards' whole latency shadow (330,000 ops) the same core
pays 0.36 (N3) or 0.25 (N2). The re-fold, `E_chip = E_hash0 + shadow`, against the 5090 at 3.36 (stock, class v4),
2.32 (lock), 4.26 (stock, the full shadow), 2.67 (lock, full) and the M5 Max at 1.40 (class v4):
| `W` | Class v4 shadow, stock / lock, N3 core | The same, N2 core | Full shadow, stock / lock, N3 | The same, N2 | M5 Max class v4, N3 / N2 | Label |
|---|---|---|---|---|---|---|
| 1 | 23x / 16x | 29x / 20x | 11x / 6.7x | 15x / 9.3x | 9.6x / 12x | modelled on synthesised pJ |
| 4 | **21x / 14x** | 25x / 17x | **10x / 6.5x** | 14x / 8.8x | 8.5x / 11x | the same |
| 8 | **18x / 12x** | 21x / 15x | **9.7x / 6.1x** | 13x / 8.2x | 7.5x / 8.9x | the same |
| 16 | 14x / 9.9x | 16x / 11x | 8.8x / 5.5x | 11x / 7.1x | 6.0x / 6.8x | the same |
Reading, which replaces "the shadow is the whole hold" of section 0 where the synthesised figure stands: at a
measured-class `k` of 0.1 the shadow holds the die to 14x to 23x under class v4 and 6x to 11x at the honest cards'
whole latency shadow, not 2x to 6x; the die's memory energy is then a third to a half of its energy per hash, and the
read width is worth 1.3x (W = 1 to W = 8) on the shadowed number, not 0.2x. The public 2x is not reachable against
this die by any shadow at a synthesised core; what holds it is the project cost against the chain's revenue (section
4), which the shadowed edge moves by 1.4x across the whole band. The rows carry the k lane's labels (ASAP7 measured by
synthesis, the node scaling approximate, the shuffle open); its own re-fold replaces this table when it lands.
Sensitivity: the wire at 0.65 pJ per bit halves every wire term (`W = 1` 0.20 nJ, 83x; `W = 8` 0.35 nJ, 49x); at 2.6
pJ per bit doubles it (`W = 1` 0.35 nJ, 47x; `W = 8` 0.95 nJ, 18x). The macro at 0.2 nJ adds 0.1 nJ to every row. The
shadowed rows move under 0.3x across the whole band.
@ -325,7 +356,7 @@ the detector, not the hash, is the instrument.
| The unified-memory SoC | a 16 GB Mac leaves at the year-2 step; 32 GB at the 16 GiB step; every Mac pays -21 to -25 percent of rate from the v6 epoch on (measured to 8 GiB). Against the die at class v4 and `k = 0.5` the M5 Max reads 3.6x to 4.0x | finding 3 of lane B: the reference joule |
| A rig | nothing changes until a project pays (section 4); the 180-day rotation and the floor do not move the chip's per-MH/s cost | the issuance trigger and the detector, as the record |
| A pool user | the chip fleet that matters is one package; the share-pattern detector is the warning | the detector before the public testnet |
| The public claim | "under 2x" is not reachable against the die at any `k` under 1 (2.0x at `k = 1` needs the honest card's whole latency shadow); 4x at `k = 0.5`; 20x to 60x without the shadow. The number a miner can act on is the project's price threshold: USD 0.13 to 2.9 per IGN | section 7's served sentence |
| The public claim | "under 2x" is not reachable against the die by any shadow: on the k lane's synthesised core 10x at the honest cards' whole latency shadow at stock, 6x to 7x at the 5090's knee, 18x to 21x under class v4; on the record's claimed band (marked) 2x at `k = 1` and 4x at `k = 0.5`; 20x to 60x without the shadow. The number a miner can act on is the project's price threshold: USD 0.13 to 2.9 per IGN | section 7's served sentence |
## 7. Recommendation, for the 20:00 close
@ -336,10 +367,10 @@ shadow and by 0.2x to 0.4x with the shadow on, and keep the per-lane scratch res
11.5 at four, each just under a card tier's room and just over a reticle multiple (3, 5 and 6 dies; USD 1,500, 2,500
and 3,000 of silicon), and USD 5,000 is not recommended because it is 20 GiB and retires every card under 32 GB and
every Mac under 64 GB whenever it lands. The served claim should say that the strongest chip we can model for 2027
to 2028, an SRAM store on 2 nm, would reach about 4x per joule against an RTX 5090 paying its whole latency shadow
with a core half as costly as a GPU lane and 2x with one as costly, 20x to 60x without that shadow, and that such a
project pays only above about USD 0.4 per IGN at a third of the network or USD 0.13 for a maker who takes the whole
chain, which is the pattern the share detector watches for. The floor is a ticket (USD 500 to 3,000 of silicon per
to 2028, an SRAM store on 2 nm, would reach about 6x to 10x per joule against an RTX 5090 paying its whole latency
shadow with a core at the synthesised cost of a 3 nm lane (2x to 4x on the claimed band the record carried, marked),
20x to 60x without that shadow, and that such a project pays only above about USD 0.4 per IGN at a third of the
network or USD 0.13 for a maker who takes the whole chain, which is the pattern the share detector watches for. The floor is a ticket (USD 500 to 3,000 of silicon per
store across the schedule) and not a per-joule or per-MH/s lever, and the text should not claim otherwise. Nothing in
this file moves a served number; the `W = 8` row and the Apple rate curve past 8 GiB are the two measurements owed.