From d2d5f5d6f275cb43dfeb3330f8509ace4afd5e93 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:40:00 +0000 Subject: [PATCH] mixer-x4.md: the x8 packs' Metal and Apple OpenCL rows, the daily-build table per tier (5090, 9070 XT, M5 Max, gfx1036 per prepare, a scaled 8 GB-class row) and the x4 / x8 rule; chip-model-v3.md: the mixer row alone as the headline (layer 5 measured, not adopted, with the 5090 and 9070 XT g beside the Mac's), x8 rows at year 0 and year 4 Co-Authored-By: Claude Fable 5.1 --- docs/analysis/chip-model-v3.md | 26 ++++++++++++++------------ docs/plans/mixer-x4.md | 31 ++++++++++++++++++++++++++++++- 2 files changed, 44 insertions(+), 13 deletions(-) diff --git a/docs/analysis/chip-model-v3.md b/docs/analysis/chip-model-v3.md index 8de606635..0c005d97b 100644 --- a/docs/analysis/chip-model-v3.md +++ b/docs/analysis/chip-model-v3.md @@ -35,11 +35,12 @@ silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does | v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x | | v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x | | v3: mixer x4 | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x | -| v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.66x at the Mac's g = 0.93 (denominator 126.6 MH/s; the 5090's own g is pending the PC rows) | 1.98x | 1.60x | -| v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.71x at the Mac's g = 0.87 (118.4 MH/s; 5090 g pending) | 2.12x | 1.67x | -| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 576 MiB | 287 / $130 | 0.71x | 2.12x | 1.31x | -| v3 at year 12 (cache 1 GiB, dataset 8 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 1,088 MiB | 542 / $330 | 0.71x | 2.12x | 0.59x | -| v3 with the mixer at x8 instead (the next lever, not adopted) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x | +| MEASURED, NOT ADOPTED (layer 5 decided out of v3 on the PC rows, coordinator 21:40 UTC): v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.66x at the Mac's g = 0.93 (126.6 MH/s); 0.70x at the 5090's g = 0.87 (118.4); the 9070 XT's g 0.84 | 1.98x (Mac g), 2.11x (5090 g) | 1.60x, 1.71x | +| MEASURED, NOT ADOPTED: v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.71x at the Mac's g = 0.87 (118.4 MH/s); 0.73x at the 5090's g = 0.84 (114.3); the 9070 XT's g 0.80 | 2.12x (Mac g), 2.19x (5090 g) | 1.67x, 1.73x | +| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), no hot table | x4 | 599,040 | 83.5 | 512 MiB | 255 / $111 | 0.61x | 1.84x | 1.21x | +| v3 at year 12 (cache 1 GiB, dataset 8 GiB) | x4 | 599,040 | 83.5 | 1 GiB | 510 / $306 | 0.61x | 1.84x | 0.59x | +| v3 with the mixer at x8 (the candidate under the coordinator's rule of 21:30 UTC: in if the verify stays under 10 ms per warp on one Mac core and the daily 1 GiB build under 1 s on every discrete card) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x | +| x8 at year 4 | x8 | 1,198,080 | 41.7 | 512 MiB | 255 / $111 | 0.31x | 0.92x | 0.61x | The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The @@ -60,13 +61,14 @@ not a forecast). ## 3. The margin, plainly -The combined headline row without the hot table reads 1.84x with the 3x factor at an equal integer budget, 1.5x -with the SRAM area deducted. With the hot table in the added form it reads 1.98x (32 MiB) and 2.12x (64 MiB) at -an equal integer budget on the Mac's `g`, 1.60x and 1.67x with the SRAM deducted: against THIS chip the added hot -table is a cost to the honest card and none to the chip, so it moves the row the wrong way, by the card's own -`g`; whether it is adopted is the DRAM-only chip's question (`docs/plans/hot-table.md`), not this one's. The claim -is "under 2x" on the equal-silicon convention and on the equal-budget convention without the hot table; with a -64 MiB table on the equal-budget convention it is not, and the margin everywhere is thin: +The combined headline row is the mixer row alone (layer 5 is out: the added form costs the 5090 13 to 16 percent +and the 9070 XT 16 to 20 percent against the 0.97 bar, coordinator 21:40 UTC; the width stays 4 bytes; the era +draws and the cache growth cost this chip nothing at year 0). At x4 it reads 1.84x with the 3x factor at an equal +integer budget, 1.53x with the SRAM area deducted; at x8 0.92x and 0.76x. The hot-table rows above are kept as +measured, not adopted: against THIS chip an added hot table is a cost to the honest card and none to the chip, so +it would have moved the row the wrong way by the card's own `g` (2.11x to 2.19x at the 5090's g). The claim is +"under 2x" at x4 on both conventions, with the margin thin on the equal-budget one; at x8 the chip is under 1x on +both: - the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x; - the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart); diff --git a/docs/plans/mixer-x4.md b/docs/plans/mixer-x4.md index 2d9d013df..f4c92f586 100644 --- a/docs/plans/mixer-x4.md +++ b/docs/plans/mixer-x4.md @@ -208,6 +208,10 @@ base nonce 0. | mx4-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 6f48d5a2aa0dbe5f | 27.7 (wall) | | mx4-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 73caaebb28e808fe | 27.5 | | mx4-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 73caaebb28e808fe | 27.6 (wall) | +| mx8-genesis (the x8 candidate, 21:45 UTC) | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 7c28cfb06c5c65a9 | 27.7 | +| mx8-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 7c28cfb06c5c65a9 | 27.6 (wall) | +| mx8-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | bbb183f72692f840 | 27.7 | +| mx8-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | bbb183f72692f840 | 27.7 (wall) | Reading: the Rust interpreter, Metal and Apple OpenCL agree on the v3 dataset (head, MASK word, 64 samples), on every vector lane and on the 2^24-output fingerprint of each pack; the hash rate is the v2 rate (27.7 MH/s on this @@ -234,7 +238,32 @@ x4 build on the GPU is the row the measure lock will settle. ### 6.4 Timings (`with-lock.sh measure`) -Owed until the measure lock frees. +Owed until the measure lock frees (the session script takes v2, x4 and x8 together: verifier avg of 50 and the +three cold units, the 256 MiB fill on one core, the Metal 1 GiB build, two rounds each). + +### 6.5 The daily build per tier, and the x4 / x8 rule + +The coordinator's rule (21:30 UTC): x8 enters v3 if the per-warp verify stays under 10 ms on one Mac core AND the +daily 1 GiB build stays under 1 s on every discrete card we own; else x4 with the thin margin stated and x8 named +as the next lever. The integrated tier is decided beside it (consequences row C23): its build is per prepare, not +per day, so its consequence is per-day dataset reuse in the workers or a restart per epoch. + +| Card | Build at x1 | x4 | x8 | Source | +|---|---|---|---|---| +| RTX 5090 (PC 2) | 13.4 ms (1 GiB) | owed (PC job, section 8) | owed | bench-log, 3 October, memory-hard entry | +| RX 9070 XT (PC 1, eGPU) | owed | owed | owed | PC job, section 8 | +| M5 Max, Metal | 13 to 30 ms (the two runs of the memory-hard entry; tonight's run-lock figures 21.7 to 30.2 ms at x4 and 22.0 to 30.0 at x8 say the Mac's build is latency-bound, not mixer-bound) | section 6.4 | section 6.4 | this file | +| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | 6.9 / 9.4 / 11.7 s prepare total with the 1 GiB build inside | about 28 to 47 s (approximate: scaled x4; the iGPU's build is arithmetic-bound at x1 already) | about 55 to 94 s (approximate) | `docs/plans/epoch-length.md` section 7 (branch ca2-epoch), M11 table | +| gfx1036 beside WSL build jobs (PC 1) | 55 / 116 / 124 s | about 4 to 8 min (approximate) | about 7 to 17 min (approximate) | same | +| 8 GB-class discrete card (not owned; about a tenth of the 5090's rate, approximate) | about 0.13 s | about 0.5 s | about 1 s, on the edge of the rule | scaled from the 5090 row, approximate | + +Reading, before the measured rows: on the discrete cards the rule is about the 5090 and 9070 XT rows (owed, the PC +job). The integrated tier misses the rule at x4 already: a per-prepare build of 28 to 47 s is a tenth to a quarter +of the 600-DAA-second lead the devnet gives the next program (spec 1.12), and under load it is the whole lead; so +if x4 or x8 goes in, the iGPU tier needs the workers to build the day's dataset once a day and keep it across +epochs (today a prepare rebuilds it: `proto-opencl/host.c` prepareTask builds the pair's cache and dataset per +prepare), or to restart per epoch. The node agent is asked whether the per-day reuse is bounded tonight +(coordinator, 21:45 UTC); until then the iGPU consequence stands as written. ## 7. The chip model