diff --git a/docs/analysis/chip-model-v3.md b/docs/analysis/chip-model-v3.md index 8a48ca7a..8de60663 100644 --- a/docs/analysis/chip-model-v3.md +++ b/docs/analysis/chip-model-v3.md @@ -35,10 +35,10 @@ silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does | v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x | | v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x | | v3: mixer x4 | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x | -| v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads) | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.61x or below (owed: the 5090's added-form rate; the hot loads cost it something, the chip nothing) | 1.84x or below | 1.49x | -| v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.61x or below (owed) | 1.84x or below | 1.45x | -| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 576 MiB | 287 / $130 | 0.61x | 1.84x | 1.14x | -| v3 at year 12 (cache 1 GiB, dataset 8 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 1,088 MiB | 542 / $330 | 0.61x | 1.84x | 0.51x | +| v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.66x at the Mac's g = 0.93 (denominator 126.6 MH/s; the 5090's own g is pending the PC rows) | 1.98x | 1.60x | +| v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.71x at the Mac's g = 0.87 (118.4 MH/s; 5090 g pending) | 2.12x | 1.67x | +| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 576 MiB | 287 / $130 | 0.71x | 2.12x | 1.31x | +| v3 at year 12 (cache 1 GiB, dataset 8 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 1,088 MiB | 542 / $330 | 0.71x | 2.12x | 0.59x | | v3 with the mixer at x8 instead (the next lever, not adopted) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x | The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op @@ -49,15 +49,24 @@ chip's cost not at all. Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6 hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53. Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of `sram-mirror.md` -scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. Year 4 and 12 rows: the mirror of +scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. The honest denominator in the added +form is the v2 rate times `g`, the card's measured ratio with the hot loads added: on the M5 Max tonight +`g = 0.93 / 0.87 / 0.83` at 32 / 64 / 96 MiB (the cache agent, relayed by the coordinator at 21:23 UTC; +`docs/plans/hot-table.md` carries the runs); the 5090's and the 9070 XT's `g` are the PC rows, owed, and until they +land the row carries the Mac's `g` against the 5090's rate, which is a mixed figure and is marked so. Year 4 and 12 rows: the mirror of `sram-mirror.md` section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area, not a forecast). ## 3. The margin, plainly -The combined headline row reads 1.84x with the 3x factor at an equal integer budget, 1.5x with the SRAM area -deducted. The claim is "under 2x", and the margin is thin: +The combined headline row without the hot table reads 1.84x with the 3x factor at an equal integer budget, 1.5x +with the SRAM area deducted. With the hot table in the added form it reads 1.98x (32 MiB) and 2.12x (64 MiB) at +an equal integer budget on the Mac's `g`, 1.60x and 1.67x with the SRAM deducted: against THIS chip the added hot +table is a cost to the honest card and none to the chip, so it moves the row the wrong way, by the card's own +`g`; whether it is adopted is the DRAM-only chip's question (`docs/plans/hot-table.md`), not this one's. The claim +is "under 2x" on the equal-silicon convention and on the equal-budget convention without the hot table; with a +64 MiB table on the equal-budget convention it is not, and the margin everywhere is thin: - the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x; - the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart); @@ -76,8 +85,9 @@ order: at the edge of the 10 ms gate; the measured v3 row of `docs/plans/mixer-x4.md` section 6 is what to scale from now, and whether a 2019-class laptop core (unmeasured, O-1.14) passes 10 ms is what decides it. 2. The hot table: adopted or not on the PC rows (`docs/plans/hot-table.md`); in the added form it costs the GPU - nothing it was not already paying in cache misses and the chip die area only, so it is the second lever for the - chip only through area and the first against a DRAM-only chip. + 7 to 17 percent on the Mac and the chip die area only, so against this chip it is a lever in the wrong + direction and against a DRAM-only chip the first lever; if it is adopted, the mixer must carry the extra `1/g` + (x8 at g = 0.87 reads 1.06x at the equal budget, 0.84x with the SRAM deducted). ## 4. What this does not settle