Class v6: the last two 20:00 and 19:30 references pulled to the close
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of 086411627 (counter-asic-4) for the box mirror master
This commit is contained in:
parent
af0f82798e
commit
7670ed9f15
1 changed files with 2 additions and 2 deletions
|
|
@ -113,7 +113,7 @@ Main's candidate steps checked first under the standing budget rule: **6 GiB doe
|
|||
|
||||
A step every two years retires about a quarter of today's measured consumer cards each time; on the Apple side every Mac pays the measured rate cost of the larger working set on top (section 3.3: -12 percent at 2 GiB, -20 at 4, -22 at 8 on the M5 Max).
|
||||
|
||||
Against the chips, with the number: (1) a hybrid-bonded or soldered chip sized at launch (the Jasminer X4 shape, the E3 shape) dies at the first step it cannot carry, as the E3 did 20 to 27 months after shipping; a chip sized to the floor with 8 GB dies at the year-2 step on either schedule, one with 16 GB holds through year 4; but the schedule is a consensus field read at genesis, so a maker sizes to the step it wants to survive (USD 160 more of GDDR7 on a USD 470 part; about 3x the 2021 memory die area for the Jasminer shape, approximate), and the E3 died because it was sized to the card fleet's limit, not to a schedule; (2) the f = 1 GDDR7 chip's 32 GB board pays USD 0 through 16 GiB and keeps 5.1x per joule and USD 2.8 per MH/s at every step; (3) the SRAM store pays capex only, at the cache doubling (USD 46 per die at 256 MiB to 111 at 512 MiB, 306 at 1 GiB for the recompute chip's mirror; lane B's full store a die per 2 GiB: USD 400 to 600 each). **The sentence for the 20:00 reading: the schedule is a fleet-retirement rule with a chip tax of USD 0 to 160 per unit on the chips that matter, and it kills only a chip whose maker ignores a public consensus field; if adopted, the 75 percent rule and a 6 GiB floor cannot both stand for the 8 GB tier (5.5 GiB keeps it inside the rule in the best case; about 85 percent is what makes 6 GiB fit).**
|
||||
Against the chips, with the number: (1) a hybrid-bonded or soldered chip sized at launch (the Jasminer X4 shape, the E3 shape) dies at the first step it cannot carry, as the E3 did 20 to 27 months after shipping; a chip sized to the floor with 8 GB dies at the year-2 step on either schedule, one with 16 GB holds through year 4; but the schedule is a consensus field read at genesis, so a maker sizes to the step it wants to survive (USD 160 more of GDDR7 on a USD 470 part; about 3x the 2021 memory die area for the Jasminer shape, approximate), and the E3 died because it was sized to the card fleet's limit, not to a schedule; (2) the f = 1 GDDR7 chip's 32 GB board pays USD 0 through 16 GiB and keeps 5.1x per joule and USD 2.8 per MH/s at every step; (3) the SRAM store pays capex only, at the cache doubling (USD 46 per die at 256 MiB to 111 at 512 MiB, 306 at 1 GiB for the recompute chip's mirror; lane B's full store a die per 2 GiB: USD 400 to 600 each). **The sentence for the close: the schedule is a fleet-retirement rule with a chip tax of USD 0 to 160 per unit on the chips that matter, and it kills only a chip whose maker ignores a public consensus field; if adopted, the 75 percent rule and a 6 GiB floor cannot both stand for the 8 GB tier (5.5 GiB keeps it inside the rule in the best case; about 85 percent is what makes 6 GiB fit).**
|
||||
|
||||
## 4. Layer 3: scheduled family epochs by height, every 180 days by default, no release
|
||||
|
||||
|
|
@ -359,7 +359,7 @@ The decomposition for section 2's term (the H100 microbench, full residency at t
|
|||
|
||||
The ranking by k, highest first (hardest for a chip; absolute, N3 against the lock): mad 0.20, the ARX families 0.17 to 0.18, the fold 0.14, or 0.12, mul 0.08, prmt and lop3 0.06, mulhi 0.03. mulhi is the worst lever on the card by 6x: the 5090 pays 21 pJ for a high word that costs the chip the same 0.68 pJ as a low word, which is the measured reason the op-mix band keeps mulhi capped at its base (section 2) and the reason the research file's 15.1b hold on a re-weight stands from the chip side as well. The class v4 mix costs the chip about 1.0 to 1.2 pJ per op at N3 (2.0 to 2.4 at ASAP7), the shuffle row pending.
|
||||
|
||||
What it does to the GDDR7 stored-dataset board (E_mem 0.466 microjoules, modelled) with the class v4 shadow (102,100 ops per hash): absolute, the chip's shadow is 102,100 x about 1.1 pJ = 0.11 microjoules at N3, E_chip about 0.58, so the edge reads **4.0x at the lock (the card 2.33) and 5.8x at stock (3.36)**, against the served 2.1x at k = 1 and 3.4x at k 0.33. In the record's convention it is k 0.18 at the lock, 2.33 / (0.466 + 0.18 x 0.652) = 4.0x, the same number there, and at stock k 0.10, 3.36 / (0.466 + 0.10 x 1.10) = 5.8x. **The shadow buys the card about 0.5x to 1x of edge out of the 5.2x and 3.6x zero-shadow figures, not the 1.5x the served k = 1 row implies.** This is the row the served line must survive (section 9): the served "2.1x" is the k = 1 reading; the unit floor says a core's units are five to six times cheaper than that at N3, and the sequencer core that a rotating family forces (fetch, decode, register file, the drawn program's control) sits between the two, where the lane's next row puts it; until then the default reading is k 0.5 (the DRAM chip 2.9x at the knee in the record's convention) with 4.0x as the floor-k worst case. Pending from the lane by 19:30 UK: the 32-lane shuffle butterfly and the general crossbar (in place now), the 8 KB scratch as a flop array, the int8 8x8x8 tile, the mix optimiser over the layer-1 band, and the re-fold of lane 3's SRAM rows at W = 4 and 8 in both conventions.
|
||||
What it does to the GDDR7 stored-dataset board (E_mem 0.466 microjoules, modelled) with the class v4 shadow (102,100 ops per hash): absolute, the chip's shadow is 102,100 x about 1.1 pJ = 0.11 microjoules at N3, E_chip about 0.58, so the edge reads **4.0x at the lock (the card 2.33) and 5.8x at stock (3.36)**, against the served 2.1x at k = 1 and 3.4x at k 0.33. In the record's convention it is k 0.18 at the lock, 2.33 / (0.466 + 0.18 x 0.652) = 4.0x, the same number there, and at stock k 0.10, 3.36 / (0.466 + 0.10 x 1.10) = 5.8x. **The shadow buys the card about 0.5x to 1x of edge out of the 5.2x and 3.6x zero-shadow figures, not the 1.5x the served k = 1 row implies.** This is the row the served line must survive (section 9): the served "2.1x" is the k = 1 reading; the unit floor says a core's units are five to six times cheaper than that at N3, and the sequencer core that a rotating family forces (fetch, decode, register file, the drawn program's control) sits between the two, where the lane's next row puts it; until then the default reading is k 0.5 (the DRAM chip 2.9x at the knee in the record's convention) with 4.0x as the floor-k worst case. Pending from the lane by 15:30 UK (later rows as amendments): the 32-lane shuffle butterfly and the general crossbar (in place now), the 8 KB scratch as a flop array, the int8 8x8x8 tile, the mix optimiser over the layer-1 band, and the re-fold of lane 3's SRAM rows at W = 4 and 8 in both conventions.
|
||||
|
||||
### 10.3 The SRAM die and the dataset floor (floor lane 3, `docs/analysis/class-v6/floor/sram-and-floor.md` on class-v6-floor-sram at 708d01b4, first reading 12:0x UTC; full by 15:30 UK (the earlier 19:30 clock pulled by the coordinator at 13:30); everything chip-side modelled on lane B's wire figure, 1.3 pJ per bit plus 0.1 nJ per macro access and 0.05 nJ of controller; the GPU side measured)
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue