mixer-x4.md: the x8 packs' Metal and Apple OpenCL rows, the daily-build table per tier (5090, 9070 XT, M5 Max, gfx1036 per prepare, a scaled 8 GB-class row) and the x4 / x8 rule; chip-model-v3.md: the mixer row alone as the headline (layer 5 measured, not adopted, with the 5090 and 9070 XT g beside the Mac's), x8 rows at year 0 and year 4

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 21:40:00 +00:00
parent 147db8db44
commit d2d5f5d6f2
2 changed files with 44 additions and 13 deletions

View file

@ -35,11 +35,12 @@ silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does
| v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x |
| v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x |
| v3: mixer x4 | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x |
| v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.66x at the Mac's g = 0.93 (denominator 126.6 MH/s; the 5090's own g is pending the PC rows) | 1.98x | 1.60x |
| v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.71x at the Mac's g = 0.87 (118.4 MH/s; 5090 g pending) | 2.12x | 1.67x |
| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 576 MiB | 287 / $130 | 0.71x | 2.12x | 1.31x |
| v3 at year 12 (cache 1 GiB, dataset 8 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 1,088 MiB | 542 / $330 | 0.71x | 2.12x | 0.59x |
| v3 with the mixer at x8 instead (the next lever, not adopted) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x |
| MEASURED, NOT ADOPTED (layer 5 decided out of v3 on the PC rows, coordinator 21:40 UTC): v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.66x at the Mac's g = 0.93 (126.6 MH/s); 0.70x at the 5090's g = 0.87 (118.4); the 9070 XT's g 0.84 | 1.98x (Mac g), 2.11x (5090 g) | 1.60x, 1.71x |
| MEASURED, NOT ADOPTED: v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.71x at the Mac's g = 0.87 (118.4 MH/s); 0.73x at the 5090's g = 0.84 (114.3); the 9070 XT's g 0.80 | 2.12x (Mac g), 2.19x (5090 g) | 1.67x, 1.73x |
| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), no hot table | x4 | 599,040 | 83.5 | 512 MiB | 255 / $111 | 0.61x | 1.84x | 1.21x |
| v3 at year 12 (cache 1 GiB, dataset 8 GiB) | x4 | 599,040 | 83.5 | 1 GiB | 510 / $306 | 0.61x | 1.84x | 0.59x |
| v3 with the mixer at x8 (the candidate under the coordinator's rule of 21:30 UTC: in if the verify stays under 10 ms per warp on one Mac core and the daily 1 GiB build under 1 s on every discrete card) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x |
| x8 at year 4 | x8 | 1,198,080 | 41.7 | 512 MiB | 255 / $111 | 0.31x | 0.92x | 0.61x |
The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op
weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The
@ -60,13 +61,14 @@ not a forecast).
## 3. The margin, plainly
The combined headline row without the hot table reads 1.84x with the 3x factor at an equal integer budget, 1.5x
with the SRAM area deducted. With the hot table in the added form it reads 1.98x (32 MiB) and 2.12x (64 MiB) at
an equal integer budget on the Mac's `g`, 1.60x and 1.67x with the SRAM deducted: against THIS chip the added hot
table is a cost to the honest card and none to the chip, so it moves the row the wrong way, by the card's own
`g`; whether it is adopted is the DRAM-only chip's question (`docs/plans/hot-table.md`), not this one's. The claim
is "under 2x" on the equal-silicon convention and on the equal-budget convention without the hot table; with a
64 MiB table on the equal-budget convention it is not, and the margin everywhere is thin:
The combined headline row is the mixer row alone (layer 5 is out: the added form costs the 5090 13 to 16 percent
and the 9070 XT 16 to 20 percent against the 0.97 bar, coordinator 21:40 UTC; the width stays 4 bytes; the era
draws and the cache growth cost this chip nothing at year 0). At x4 it reads 1.84x with the 3x factor at an equal
integer budget, 1.53x with the SRAM area deducted; at x8 0.92x and 0.76x. The hot-table rows above are kept as
measured, not adopted: against THIS chip an added hot table is a cost to the honest card and none to the chip, so
it would have moved the row the wrong way by the card's own `g` (2.11x to 2.19x at the 5090's g). The claim is
"under 2x" at x4 on both conventions, with the margin thin on the equal-budget one; at x8 the chip is under 1x on
both:
- the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x;
- the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart);

View file

@ -208,6 +208,10 @@ base nonce 0.
| mx4-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 6f48d5a2aa0dbe5f | 27.7 (wall) |
| mx4-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 73caaebb28e808fe | 27.5 |
| mx4-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 73caaebb28e808fe | 27.6 (wall) |
| mx8-genesis (the x8 candidate, 21:45 UTC) | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 7c28cfb06c5c65a9 | 27.7 |
| mx8-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 7c28cfb06c5c65a9 | 27.6 (wall) |
| mx8-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | bbb183f72692f840 | 27.7 |
| mx8-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | bbb183f72692f840 | 27.7 (wall) |
Reading: the Rust interpreter, Metal and Apple OpenCL agree on the v3 dataset (head, MASK word, 64 samples), on
every vector lane and on the 2^24-output fingerprint of each pack; the hash rate is the v2 rate (27.7 MH/s on this
@ -234,7 +238,32 @@ x4 build on the GPU is the row the measure lock will settle.
### 6.4 Timings (`with-lock.sh measure`)
Owed until the measure lock frees.
Owed until the measure lock frees (the session script takes v2, x4 and x8 together: verifier avg of 50 and the
three cold units, the 256 MiB fill on one core, the Metal 1 GiB build, two rounds each).
### 6.5 The daily build per tier, and the x4 / x8 rule
The coordinator's rule (21:30 UTC): x8 enters v3 if the per-warp verify stays under 10 ms on one Mac core AND the
daily 1 GiB build stays under 1 s on every discrete card we own; else x4 with the thin margin stated and x8 named
as the next lever. The integrated tier is decided beside it (consequences row C23): its build is per prepare, not
per day, so its consequence is per-day dataset reuse in the workers or a restart per epoch.
| Card | Build at x1 | x4 | x8 | Source |
|---|---|---|---|---|
| RTX 5090 (PC 2) | 13.4 ms (1 GiB) | owed (PC job, section 8) | owed | bench-log, 3 October, memory-hard entry |
| RX 9070 XT (PC 1, eGPU) | owed | owed | owed | PC job, section 8 |
| M5 Max, Metal | 13 to 30 ms (the two runs of the memory-hard entry; tonight's run-lock figures 21.7 to 30.2 ms at x4 and 22.0 to 30.0 at x8 say the Mac's build is latency-bound, not mixer-bound) | section 6.4 | section 6.4 | this file |
| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | 6.9 / 9.4 / 11.7 s prepare total with the 1 GiB build inside | about 28 to 47 s (approximate: scaled x4; the iGPU's build is arithmetic-bound at x1 already) | about 55 to 94 s (approximate) | `docs/plans/epoch-length.md` section 7 (branch ca2-epoch), M11 table |
| gfx1036 beside WSL build jobs (PC 1) | 55 / 116 / 124 s | about 4 to 8 min (approximate) | about 7 to 17 min (approximate) | same |
| 8 GB-class discrete card (not owned; about a tenth of the 5090's rate, approximate) | about 0.13 s | about 0.5 s | about 1 s, on the edge of the rule | scaled from the 5090 row, approximate |
Reading, before the measured rows: on the discrete cards the rule is about the 5090 and 9070 XT rows (owed, the PC
job). The integrated tier misses the rule at x4 already: a per-prepare build of 28 to 47 s is a tenth to a quarter
of the 600-DAA-second lead the devnet gives the next program (spec 1.12), and under load it is the whole lead; so
if x4 or x8 goes in, the iGPU tier needs the workers to build the day's dataset once a day and keep it across
epochs (today a prepare rebuilds it: `proto-opencl/host.c` prepareTask builds the pair's cache and dataset per
prepare), or to restart per epoch. The node agent is asked whether the per-day reuse is bounded tonight
(coordinator, 21:45 UTC); until then the iGPU consequence stands as written.
## 7. The chip model