418 lines
35 KiB
Markdown
418 lines
35 KiB
Markdown
# Mixer x4 and the cache growth rule: the class v3 dataset construction
|
|
|
|
5 October 2026 (night). Counter ASIC 2.0, layer 6 (option C) and ledger M16's lever, decided by the coordinator
|
|
under the project lead's delegation at 22:00 UTC (`docs/plans/counter-asic-2-status.md`, "22:00 decided"; the project lead confirms for
|
|
the public testnet genesis). Branch `ca2-mixer`. Worker: ca2-mixer (cryptographer's lane).
|
|
|
|
What this changes, in one line: under program class v3 every mixer application of the dataset item derivation
|
|
becomes four applications with distinct round keys, the eight dependent cache reads per item stay eight, and the
|
|
256 MiB cache doubles on the days the dataset doubles (years 4 and 12). Version 2 is byte-identical: the two
|
|
pinned packs re-export without a changed byte (section 5).
|
|
|
|
## 1. Why this form
|
|
|
|
The recompute attacker of `docs/analysis/m16-recompute-attacker-2026-10-05.md` holds the 256 MiB cache on a die
|
|
and derives every dataset word instead of reading it: 128 items per hash at about 1,170 integer operations and 8
|
|
dependent cache reads each. Its cost is linear in operations per item; the honest miner pays the mixer once a day
|
|
in the dataset build and never per hash; the verifier pays it per item it checks. The multiplier `m` is the one
|
|
parameter that moves the attacker and leaves the honest hash rate untouched.
|
|
|
|
Two shapes give the attacker 4x the operations:
|
|
|
|
| Shape | Mixer applications per item | Dependent cache reads per item | What grows for the verifier | What grows for the chip | What grows for the honest build |
|
|
|---|---|---|---|---|---|
|
|
| A, chosen: `m = 4` applications per round, 8 rounds | 36 | 8 | the ALU part only; the latency part (8 dependent misses per item) unchanged | integer operations 4x; SRAM bandwidth unchanged (1,024 reads per hash) | 4x the mixer arithmetic, same reads |
|
|
| B, alternative: 32 rounds of one application and one read | 33 | 32 | both parts: 4x the dependent misses per item, so about 4x the latency-bound time (spec 1.11: about 8 x 100 ns per item in series without interleaving) | integer operations 3.7x and SRAM bandwidth 4x (4,096 reads per hash, 60 TB/s to match one 5090 at the M16 rate) | 4x the reads too; the GPU build becomes latency-bound at 4x the dependent line fetches |
|
|
|
|
Shape A is chosen because the verifier's latency part is the part the 10 ms gate protects (section 1.11: the
|
|
distinct items of a unit are derived with their chains interleaved so the 8 misses of each item overlap across up
|
|
to 32 items; shape B would make that 32 misses deep). Shape B is implemented nowhere; the rule in the brief: if
|
|
shape A's measured verifier time exceeds 4.8 ms per warp on one M5 Max core, measure both and recommend. Section 6
|
|
has the measurement; it is under that bound, so B stays unimplemented.
|
|
|
|
## 2. Spec text (replaces 1.8.5 and 1.13.3 under class v3; v2 text unchanged)
|
|
|
|
### 1.8.5 Item derivation and dataset mapping
|
|
|
|
Item `t` (16 words) under mixer multiplier `m` (`m = 1` for program class v2, `m = 8` for class v3 as shipped, `LoadClass::MX8`, the x8 decision of section 6.4a at 22:05 UTC; `m = 4` was the candidate this file was written for, corrected 6 October 2026; a class
|
|
parameter, `LoadClass::mixer_mult`):
|
|
|
|
```
|
|
s[i] = K[i] for i in 0..7
|
|
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
|
|
for r in 0..7:
|
|
for j in 0..m-1:
|
|
s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
|
|
a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache
|
|
s[i] = s[i] XOR cache[line a][i] for i in 0..15
|
|
for j in 0..m-1:
|
|
s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
|
|
item(t) = s
|
|
```
|
|
|
|
`M(s, rk)` is the mixer of 1.8.4 with round key `rk`; under `m = 1` the keys are `(r + 1) * 0x9E3779B9` and
|
|
`9 * 0x9E3779B9`, the version 2 text exactly. The round keys of the `9 m` applications are the first `9 m` values
|
|
of the version 2 key sequence, all distinct (the sequence is `k * 0x9E3779B9` for `k = 1 .. 9 m`, and
|
|
`0x9E3779B9` is odd, so no two of the first 2^32 keys coincide). Eight dependent cache reads per item at every
|
|
`m` (`ITEM_ROUNDS = 8`, prototype value): the address of read `r` depends on every earlier read. `9 m` mixer
|
|
applications, 36 under class v3, about 4,700 integer operations per item (130 per application, 1.8.4).
|
|
`dataset[w] = item(w >> 4)[w AND 15]`. A dataset of 2^D words is the prefix of items `0 .. 2^(D-4) - 1`, so an
|
|
item has the same value at every dataset size; and a cache of 2^C words is the prefix of segments of every
|
|
larger cache (1.8.3 fills segments independently of the cache size), but an item's value depends on `C` through
|
|
the line mask, so the item changes on the day the cache doubles.
|
|
|
|
Source: `igneum-pow/src/memhard.rs` (`derive_items`, `round_key_mult`, `Shape`), the emitted `mh_item` of
|
|
`memhard.h`, `memhard.metal` and `kernel.cl` (`emit.rs`, `emit_memhard_core`: the `m` loop is emitted only for
|
|
`m > 1`, so every version 2 pack keeps its text).
|
|
|
|
### 1.13.3 Dataset growth (class v3: option (b) with the cache tied to it, "option C")
|
|
|
|
Designed: 2 GiB at genesis plus 0.5 GiB per year. The linear schedule in bytes, `G x (1 + 86,400 d /
|
|
(4 x 31,536,000)) = G x (1 + d / 1,460)` for the genesis size `G` and the chain day `d` (DAA days since genesis,
|
|
section 1.12), doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at day 10,220
|
|
(year 28). Rule (Designed, decided 5 October 2026 for class v3):
|
|
|
|
```
|
|
doublings(d) = floor(log2(1 + d / 1460)) integer division, then integer log2
|
|
dataset_words(d) = 2^(D_0 + doublings(d)) D_0 = 29 designed (2 GiB), 28 on the devnet (1 GiB); capped at 32
|
|
cache_words(d) = 2^(26 + doublings(d)) 256 MiB, 512 MiB from year 4, 1 GiB from year 12
|
|
```
|
|
|
|
Power-of-two sizes only (option (b)), so every load keeps the `src AND MASK` form of 1.14 item 2 and the cache
|
|
line index keeps `s[0] AND mask`. The cache doubles exactly when the dataset doubles ("option C",
|
|
`docs/analysis/sram-mirror.md` section 7): the cache's job is to stay above any GPU's last-level cache and that
|
|
needs growth; the recompute attacker is priced by the mixer, not by the cache (section 7 below).
|
|
|
|
`d` is `day_index(header.timestamp) - day_index(genesis.timestamp)` with `day_index = timestamp_ms / 86,400,000`
|
|
(`bind::day_index`, the interim day rule), clamped at 0 (`memhard::days_since_genesis`). Under class v2 nothing
|
|
grows: the cache is 2^26 words and the dataset the genesis size on every day.
|
|
|
|
| Chain day `d` | Years | `doublings` | Cache words | Cache | Dataset words (devnet `D_0 = 28`) | Dataset (designed `D_0 = 29`) | Verifier cache fill, one M5 Max core (measured at 256 MiB, section 6, scaled linearly) |
|
|
|---|---|---|---|---|---|---|---|
|
|
| 0 to 1,459 | 0 to 4 | 0 | 2^26 | 256 MiB | 2^28 (1 GiB) | 2 GiB | 0.18 s |
|
|
| 1,460 to 4,379 | 4 to 12 | 1 | 2^27 | 512 MiB | 2^29 (2 GiB) | 4 GiB | 0.36 s |
|
|
| 4,380 to 10,219 | 12 to 28 | 2 | 2^28 | 1 GiB | 2^30 (4 GiB) | 8 GiB | 0.72 s |
|
|
| 10,220 to 21,899 | 28 to 60 | 3 | 2^29 | 2 GiB | 2^31 (8 GiB) | 16 GiB | 1.4 s |
|
|
| 21,900 and on | 60 and on | 4 | 2^30 | 4 GiB | 2^32 (16 GiB, the index cap) | 2^32 words, the cap | 2.9 s |
|
|
|
|
Test: `memhard::tests::growth_schedule_table` pins every row and the day before each step. The devnet pack
|
|
`igneum-devnet-v4-epoch0` is day 20,730 of the Unix count against genesis day 20,729, `d = 1`, so every existing
|
|
size and vector stands.
|
|
|
|
Consequences for the tiers (the rule of 5 October): a verifier (any node, any pool core) holds 512 MiB from year 4
|
|
and 1 GiB from year 12, and fills it once a day in under a second on one 2026 core (the table); a miner's card
|
|
holds the dataset, 4 GiB from year 4 and 8 GiB from year 12 on the designed schedule, so an 8 GB card mines until
|
|
year 12 and a 16 GB card until year 28 (the cache is not in the card's working set at hash time: it is built,
|
|
the dataset built from it, and dropped). Those dates are the design document's own schedule restated as steps;
|
|
option (a) would have faded a 4 GiB card out in year 4 instead of year 4.
|
|
|
|
## 3. Interfaces
|
|
|
|
| Item | Where | Note |
|
|
|---|---|---|
|
|
| `LoadClass { mixer_mult: u8, growth: bool }`, `LoadClass::MX4` ("mx4"), `with_mixer(m, growth)`, `v2_loads()`, `takes_width_roll()` | `generator.rs` | a class with v2 loads takes no width roll: its program stream is version 2's draw for draw, so the v3 program of a seed is the v2 program of that seed, only the dataset differs |
|
|
| `Shape { mixer_mult, cache_log2_words }`, `Shape::for_class_day(class, d)`, `MixParams.shape`, `Cache::fill_log2(key, log2)`, `round_key_mult(r, j, m)` | `memhard.rs` | `Shape::V2` is version 2 |
|
|
| `growth_doublings(d)`, `cache_log2_words(d)`, `dataset_log2_words(D_0, d)`, `days_since_genesis(day, genesis_day)` | `memhard.rs` | the schedule, one function and its two sizes |
|
|
| `DatasetSource::{new_shape, from_key_shape, shape}`, `Epoch::new_class_day`, `Epoch::from_seed_bytes_day(epoch, day, label, class, d, D_0)` | `verify.rs` | the day-sized entries; the v2 entries are unchanged and build the v2 shape |
|
|
| `IGNEUM_MIXER_MULT`, `IGNEUM_CLASS_MIXER_MULT`, `IGNEUM_CACHE_GROWTH` in program.h; `"mixer_mult"`, `"cache_growth"`, the `"item"` string in program.json | `emit.rs` | written only for a class with `m != 1` or growth, so v2 packs do not change |
|
|
| `packfile.h` `mixerMult`; `packbench` and the OpenCL host print the multiplier and the cache size | the three hosts | the kernels carry the construction in their text (one emitter, three dialects); the hosts size the cache from `IGNEUM_CACHE_LOG2_WORDS` already (packbench.swift line 56, host.cu line 65, host.c line 1966) |
|
|
| `igneum-pow --class mx4 [--days d]` on every command | `main.rs` | `--days` sizes the cache for a growth class |
|
|
|
|
Under the ca2-v3 seam (`ProgramClass::V3`, `V3_CLASS`), the integration sets `V3_CLASS = LoadClass::MX4`; the
|
|
chain's day-sized dataset needs the day index, so `Epoch::chain_dataset(day, class)` builds the genesis-size
|
|
cache and a `chain_dataset_day(day_bytes, class, d, D_0)` beside it is the growth entry (section 9, owed to the
|
|
node agent).
|
|
|
|
## 4. Vectors (class v3, `proto-cuda/packs-ca2-mixer/`)
|
|
|
|
Produced by `igneum-pow export --program-class v3` (the Rust CPU interpreter, 5 October 2026, commit 66eeba3) and
|
|
checked on the GPUs in section 6. The program of each pack is the version 2 program of the same seed instruction
|
|
for instruction (`tests/packs.rs`, `v3_packs_are_the_v2_seeds_under_mixer_x4`); the cache is the version 2 cache
|
|
(day 0 of the growth rule); the dataset words and the hashes are new. The 64 sampled indices are those of every
|
|
pack (`emit::sample_indices`).
|
|
|
|
Pack `mx4-genesis` (seed igneum-genesis, day 2026-10-03, generator 3, class mx4, program id e323b9dcaf283a6f, 2^28 words, 2^26-word cache, cache FNV-1a 64 `48c4f5bf24166b2e` as under v2):
|
|
|
|
```
|
|
dataset words 0..15 (item 0):
|
|
61ff2180 0d4c7e6c 2177d443 60df9025 cf8b2e10 63675bfb 25289e58 9c45dc42
|
|
2d271c54 9652369b 2dd77508 5921392c 3afa60ee c640ad68 f2bb56ff cfa46438
|
|
dataset[0x0fffffff] = 5020180e
|
|
sampled words (the first 8 of the 64 in vectors.json):
|
|
dataset[59471966] = de85726d dataset[217795994] = 7cfc31c7 dataset[208353206] = d3cd5289 dataset[42483309] = 6ebbeef1
|
|
dataset[172547758] = 9858413b dataset[148076330] = e786f141 dataset[183853158] = 64f13833 dataset[214389424] = ca229d04
|
|
unit at base nonce 0, lanes 0..31:
|
|
63acd2d273f475ba e929c78b34b80d4b 0b1011cb19982558 1457a0df5497aa11
|
|
957d0f3bb71d98fb ac16901e6e6f6057 8ea1c6279f4b177a f28146e60bd08ba9
|
|
fe5b8cfe87f8e65b 49f87240566ace62 6ef6d6b7bdea8e41 46d9c0dc29a97b9c
|
|
111fe30128db9398 66dc39084f0946d4 8ee11bdfd35fecf2 2861fcfc75db6677
|
|
31c7667d4bde8556 c5989c48858b4ce0 276395e734a9d30d 84217b41e91368ff
|
|
3604861e34d9f697 9f51d8ee16bf3639 c89e47bafa84401c 7ae78c1f10b70e19
|
|
0b8c947157a29a48 d67192e8cfb43842 05a4c6d182c8c675 188e2661f3263f2e
|
|
a2df24238f7fea2e ed69ea7e13ad3a48 2d44ae509bab91b8 adad61931ea4fb70
|
|
unit at base nonce 4096: lane 0 edd508ac57e5699a, lane 1 4eefd56d526cdaeb, lane 31 8892f8604733b1e0
|
|
unit at base nonce 1000000: lane 0 8b3183778a49f59c, lane 1 1831b72a8797e895, lane 31 75eae55eba53a506
|
|
```
|
|
|
|
Pack `mx4-devnet-epoch0` (seed the devnet genesis hash, day bytes of 2026-10-04, generator 3, class mx4, program id 73bcbfe8ccf988f1, 2^28 words, 2^26-word cache, cache FNV-1a 64 `448274a57f508cbc` as under v2):
|
|
|
|
```
|
|
dataset words 0..15 (item 0):
|
|
afe80d67 b9fbd029 6c79f193 95139ad9 96310aff 4609f8b1 75279e63 28235be1
|
|
47b17dcb 718e0ef2 a52588c8 a8bf49d5 19cf243e 5ec8905e a4851f66 af9cd9f3
|
|
dataset[0x0fffffff] = e6a99c7a
|
|
sampled words (the first 8 of the 64 in vectors.json):
|
|
dataset[59471966] = 57642b58 dataset[217795994] = c279badd dataset[208353206] = cbccbaad dataset[42483309] = 32cce392
|
|
dataset[172547758] = 71fdb4c6 dataset[148076330] = f5a268ce dataset[183853158] = 3ca1d676 dataset[214389424] = 7977b03d
|
|
unit at base nonce 0, lanes 0..31:
|
|
212c6442b51e87ae c374795c00839331 b6036a220a98f4b3 eb8b8013e637367b
|
|
db5866e9b73930fd f3f3d01f46e90333 9d913991ab8ed428 7ccb1d8fa100a800
|
|
3cf45ba44f09a0fe 91acf48ef1a63082 6ea46c69fb082f99 581f0218977a9d72
|
|
9a4623a5c62ddf2d ab6eb5e768f0feb4 07b70bdccf8aca12 d666311ae5e4311e
|
|
53114757d669f0a4 bd5d6ace87ce2ce4 fb712015e8189192 a32cec81103e134b
|
|
83f3d18c3289c124 fe29f1984b132b3d c9ffcf4e3774497a ac99c9243dc63809
|
|
d78a7e8217a32f3c ab81ad63d242fc31 0e4c30b7e00024af ce014289fff6778d
|
|
64a1292e2a8b4a91 d5b8c90e681d7e3a 06078117673030fd 51bf77b280173930
|
|
unit at base nonce 4096: lane 0 3c797978566b5950, lane 1 7c759e60185b6411, lane 31 96a903eb9a0ca390
|
|
unit at base nonce 1000000: lane 0 d5a8da0568df8ee7, lane 1 68cb69c04208285a, lane 31 f8ca84a1a5d78cf5
|
|
```
|
|
|
|
## 5. The v2 path is byte-identical
|
|
|
|
`cargo test --test packs` regenerates every file of `igneum-genesis-mh` and `igneum-devnet-v4-epoch0` from
|
|
`program.json` and compares byte for byte (`emitted_sources_match_all_packs`, `export_pack_matches_all_packs`);
|
|
section 6 also records a fresh `igneum-pow export` of both packs diffed against the checked-in directories.
|
|
|
|
## 6. Measurements
|
|
|
|
Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0, 5 October 2026 (night), other agents' builds running beside every
|
|
run; a timing row says which lock it ran under (`measure` is exclusive; `run` and `build` are not timings).
|
|
|
|
### 6.1 The v2 path, byte for byte (no lock needed)
|
|
|
|
`igneum-pow export --seed igneum-genesis --day 2026-10-03` and `igneum-pow export --epoch-hex edc4fa84...fb07
|
|
--day-hex 69676e65756d2d6461792ffa50000000000000` on commit 66eeba3, `diff -r` against
|
|
`proto-cuda/packs/igneum-genesis-mh` and `igneum-devnet-v4-epoch0`: IDENTICAL, both (twelve files each). The
|
|
crate tests regenerate the same files and compare them on every run (`tests/packs.rs`, 12 of 12 pass).
|
|
|
|
### 6.2 Bit-exactness of the class v3 construction on the GPUs (`with-lock.sh run`, 22:05 UTC)
|
|
|
|
`packbench --pack <dir> --batches 1 --batch-log2 24 --group 256` (Metal, built from this branch) and
|
|
`igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 1 --batch-log2 24` (Apple OpenCL, built from
|
|
this branch's host.c). Vectors are the Rust interpreter's; the fingerprint is FNV-1a 64 over the 2^24 outputs at
|
|
base nonce 0.
|
|
|
|
| Pack | Harness | Cache FNV-1a 64 | Dataset head, word MASK, 64 samples | Vectors standalone / in batch | Fingerprint 2^24 | MH/s (GPU time; not a measurement, the run lock) |
|
|
|---|---|---|---|---|---|---|
|
|
| mx4-genesis | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 6f48d5a2aa0dbe5f | 27.6 |
|
|
| mx4-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 6f48d5a2aa0dbe5f | 27.7 (wall) |
|
|
| mx4-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 73caaebb28e808fe | 27.5 |
|
|
| mx4-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 73caaebb28e808fe | 27.6 (wall) |
|
|
| mx8-genesis (the x8 candidate, 21:45 UTC) | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 7c28cfb06c5c65a9 | 27.7 |
|
|
| mx8-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 7c28cfb06c5c65a9 | 27.6 (wall) |
|
|
| mx8-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | bbb183f72692f840 | 27.7 |
|
|
| mx8-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | bbb183f72692f840 | 27.7 (wall) |
|
|
|
|
Reading: the Rust interpreter, Metal and Apple OpenCL agree on the v3 dataset (head, MASK word, 64 samples), on
|
|
every vector lane and on the 2^24-output fingerprint of each pack; the hash rate is the v2 rate (27.7 MH/s on this
|
|
card tonight, readwidth table), as it must be: the hash kernel only loads, the mixer is paid in the build.
|
|
|
|
Dataset build under the run lock (indicative only; the measured rows are in 6.4): Metal 30.2 ms GPU (mx4-genesis)
|
|
and 21.7 ms GPU (mx4-devnet-epoch0) for 1 GiB; Apple OpenCL 55 and 56 ms wall. The readwidth entry's v2 figure on
|
|
this card is 0.6 to 2.1 ms cache fill and a 1 GiB build the bench-log's memory-hard entry puts at 13 to 30 ms; the
|
|
x4 build on the GPU is the row the measure lock will settle.
|
|
|
|
### 6.2a The packs as pinned after the x8 decision (`with-lock.sh run`, 22:12 UTC, the fixed binary's exports)
|
|
|
|
| Pack | How exported | Program id | Unit 0 lane 0 | Fingerprint 2^24 (Metal = Apple OpenCL) | Vectors, self-tests |
|
|
|---|---|---|---|---|---|
|
|
| mx8-genesis | `--program-class v3` (generator 3, V3_CLASS = mx8) | e323b9dcaf283a6f | 19b56348bc85304d | 7c28cfb06c5c65a9 | 3/3 + 3/3, 96 of 96, PASS |
|
|
| mx8-devnet-epoch0 | the chain path, `--program-class v3 --era-hex <genesis>` (the era drawn inside the class, load class `mx8-erad810f22d`) | 73bcbfe8ccf988f1 | d424577fce4a7a60 | 90f794dd556f7a3b | 3/3 + 3/3, 96 of 96, PASS |
|
|
| mx4-genesis (the x4 record) | `--class mx4` (generator 2, the class in the id) | 951b89584750bd75 | 63acd2d273f475ba | 6f48d5a2aa0dbe5f | 3/3 + 3/3, 96 of 96, PASS |
|
|
| mx4-devnet-epoch0 (the x4 record) | `--class mx4` on the chain seeds | (program.json) | 212c6442b51e87ae | 73caaebb28e808fe | 3/3 + 3/3, 96 of 96, PASS |
|
|
|
|
The PC 1 rows of 6.5 ran the earlier mx8-devnet-epoch0 (the load-class export, fingerprint bbb183f72692f840, era
|
|
not inside); the composed pack's PC fingerprints are owed with the next PC round (section 9).
|
|
|
|
### 6.3 Soundness suite on the v3 construction (commit 66eeba3 and after; `with-lock.sh build` for cargo, `run` for the GPU)
|
|
|
|
| Suite | Command | Result |
|
|
|---|---|---|
|
|
| Crate lib tests (memhard schedule table, the by-hand multiplied mixer, the seam, the generator) | `cargo test -j4 --release` | 44 of 44 pass |
|
|
| Pinned packs, v2 and v3 (`tests/packs.rs`: programs, ids, dataset words, 96 vectors per pack, every emitted file byte for byte, 16 masked loads per kernel, the v3 packs as the v2 seeds under mixer x4) | same | 12 of 12 pass |
|
|
| Scratch soundness tests of branch ca2-soundness (cherry-pick 0d8f745, one conflict in packbench.swift's RESULT line resolved by hand, `era_bytes: None` added to the edge program literal) | `cargo test -j4 --release --test scratch` | 7 of 7 pass: bijections, re-hit rates, 56 of 56 edge units, 42 of 42 emitted scr kernels, 200 scratch programs and 800 units on the CPU |
|
|
| v3 fuzz on the CPU (`tests/mixer.rs`): 200 programs through the seam, the contract on every instruction (each is the v2 program of its seed), 4 units each across the 32-bit range with one unit in the top 256 nonces, interpreted twice | `IGNEUM_MIXER_PACKS_OUT=<dir> cargo test -j4 --release --test mixer` | 200 of 200, 800 of 800 units; 200 packs written for the GPU runs (82 s with the three cache fills) |
|
|
| v3 stats beside v2 (`tests/mixer.rs`): 8,192 outputs per seed, bit balance, single-bit avalanche within the unit and across units, duplicates | same | igneum-genesis v3: avalanche 49.99 percent, worst bit z 1.92, 0 duplicates (v2: 49.87, z 2.25); igneum-genesis/stats1 v3: 49.97, z 3.09 (v2: 49.98, z 2.30) |
|
|
| v3 edge (`tests/mixer.rs`): items 0, 1, 2^28 - 1 and 2^32 - 1 by hand at m = 1, 2, 4, 8 on a 2^14-word cache; words 0, 15, 16, 17, MASK - 1, MASK through the interpreter's fetch path; the index wrap at MASK + 1 | same | pass |
|
|
| v3 determinism (`tests/mixer.rs`): two independent epochs, every vector and every emitted file equal, and equal to the pinned pack | same | pass |
|
|
| Metal and Apple OpenCL on the two pinned v3 packs | section 6.2 | 3/3 standalone, 3/3 in batch, 96 of 96 lanes, dataset head, MASK word and 64 samples, one fingerprint per pack across both harnesses |
|
|
| Metal fuzz: the 200 packs, 4 units each standalone and the top-256 unit inside a 512-nonce batch at base 4,294,967,040; every tenth pack on Apple OpenCL as well | `packbench --pack <dir> --batches 1 --batch-log2 9 --batch-base 4294967040` (Metal), `igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 1 --batch-log2 10` (Apple OpenCL), `with-lock.sh run`, 21:22 to 21:24 UTC | Metal 200 of 200 packs PASS (800 of 800 standalone units, 200 of 200 inside the wrapping window, cache and dataset self-tests on every pack); Apple OpenCL 20 of 20 packs PASS (the three vectors.h units, the self-tests); a first run with a packbench built before the `--batch-base` cherry-pick reported 200 of 200 FAIL on an empty RESULT line and was read as such (the watcher rule), the harness rebuilt and the run repeated |
|
|
|
|
### 6.4 Timings (`with-lock.sh measure`, one session, 21:40:12 to 21:40:23 UTC, commit 504cae4)
|
|
|
|
Script `measure-v3.sh` (session scratchpad): `igneum-pow bench --seed igneum-genesis --day 2026-10-03 --warps 50`
|
|
(v2), `... --program-class v3` (x4), `... --class mx8` (x8), two rounds each, then the devnet seeds, then
|
|
`packbench --pack <dir> --batches 2 --batch-log2 22 --group 256` on igneum-genesis-mh, mx4-genesis and mx8-genesis,
|
|
two rounds. The lock was exclusive among the agents' builds and measurements, but the box was not quiet: load
|
|
average 5.6 (one minute) and 26 (fifteen minutes) at the start, from unlocked processes (the devnet node, other
|
|
agents' editors); the v2 row reads 1.31 to 1.36 ms where the quiet readwidth night read 0.604 to 0.626. So the
|
|
absolute numbers below are a loaded-core figure, about 2.2x the quiet one, and the ratios between the rows are the
|
|
measurement (two rounds within 4 percent). A quiet-box re-run is owed (section 9).
|
|
|
|
| Construction | Verifier, ms per 32-lane unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | 256 MiB cache fill, one core | Metal 1 GiB dataset build, GPU ms (round 1 / round 2) |
|
|
|---|---|---|---|---|---|
|
|
| v2 (igneum-genesis) | 1.361 / 1.310 | 1.579 | 1 | 172.1 / 172.6 ms | 29.7 / 21.0 |
|
|
| x4 (mx4, class v3) | 1.956 / 1.923 | 2.043 | 1.45x | 175.3 / 172.3 ms | 20.9 / 21.0 |
|
|
| x8 (mx8) | 2.785 / 2.790 | 2.942 | 2.09x | 172.3 / 172.3 ms | 21.9 / 21.9 |
|
|
| x4, the devnet seeds (mx4-devnet-epoch0) | 1.923 | 2.012 | | 173.9 ms | |
|
|
| x8, the devnet seeds | 2.972 | 2.885 | | 173.6 ms | |
|
|
|
|
Reading. The verifier's ALU part is what grows: x4 adds 0.6 ms per unit for 27 more mixer applications on each of
|
|
4,096 items (110,592 applications, about 5.5 ns each on this core, the lanes' chains interleaved), x8 another
|
|
0.85 ms for 36 more; the latency part (8 dependent misses per item) is the same in every row, which is why the
|
|
measured ratios are 1.45x and 2.1x and not the 4x and 8x of the M16 table's scaling. Shape B (32 rounds of one
|
|
read, section 1) would have multiplied the latency part too; x4 is under the 4.8 ms bar even on the loaded core, so
|
|
B stays unimplemented. The cache fill does not depend on the mixer (it is the ChaCha chain): 172 to 175 ms, the
|
|
spec's 175 to 181 ms of 1.8.3. The Metal 1 GiB build does not move with the mixer at all (21 ms at v2, x4 and x8
|
|
once warm; the 29.7 ms first v2 run is the first-touch cost the hosts fill twice for): on this card the build is
|
|
bound by the 8 dependent cache-line reads per item, not by the arithmetic, so the Mac says nothing about whether
|
|
the 5090's or the 9070 XT's build is arithmetic-bound; that is the PC job (section 8).
|
|
|
|
Against the x4 / x8 rule (section 6.5): the verifier half passes for x8 with 7.1 ms of the 10 ms gate to spare on
|
|
this loaded core (worst cold 2.94 ms; the quiet-core figure would be about 1.3 ms, scaled by the 2.2x of the v2
|
|
row, approximate); x4 leaves 8.0 ms. The build half waits on the PC rows.
|
|
|
|
### 6.5 Verification throughput per tier (consequences row C19), from the loaded-core figures above
|
|
|
|
Warps verified per second on one core = 1,000 / (ms per warp); a pool core verifying members' shares handles that
|
|
many shares per second; a node verifies a block with one unit (plus the header path, under 0.1 ms, not measured
|
|
here); IBD over the 108,000-header pruning window (spec 02) on one core = 108,000 x ms per warp.
|
|
|
|
| Figure | v2 | x4 | x8 | Note |
|
|
|---|---|---|---|---|
|
|
| ms per warp, steady (this session, loaded core) | 1.33 | 1.94 | 2.79 | avg of the two rounds |
|
|
| ms per warp, quiet M5 Max core (scaled by 0.604 / 1.33 = 0.45, approximate) | 0.60 | 0.88 | 1.26 | the readwidth night's v2 figure is measured; x4 and x8 scaled |
|
|
| ms per warp, 2019-class laptop core (approximate: 2.5x the quiet M5 Max figure, the ratio the design document assumes for the gate; unmeasured, O-1.14) | 1.5 | 2.2 | 3.2 | the figure that fixes the gate is a measurement, not this row |
|
|
| Shares per second per core (loaded / quiet, approximate) | 750 / 1,660 | 515 / 1,140 | 358 / 790 | spec 09 section 9.8 item 5 carried 2,270 at v2; re-cut from the quiet row: 1,660 |
|
|
| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second), loaded / quiet | 2.9 / 1.3 | 4.3 / 1.9 | 6.1 / 2.8 | |
|
|
| Node: worst cold single unit (loaded core) | 1.58 ms | 2.04 ms | 2.94 ms | per block |
|
|
| IBD over 108,000 headers on one core, loaded / quiet, minutes | 2.4 / 1.1 | 3.5 / 1.6 | 5.0 / 2.3 | laptop (approximate): 2.7 / 4.0 / 5.8 min; a seed VM core (unmeasured) sits between the laptop and the quiet M5 Max |
|
|
| Margin left under the 10 ms gate for Counter ASIC 3.0 (worst cold, loaded core) | 8.4 ms | 8.0 ms | 7.1 ms | on the 2019-class laptop row (approximate) 7.5 / 6.8 / 5.9 ms steady |
|
|
|
|
Reading: at x4 a pool core serves about 1,100 shares per second on a quiet 2026 core (a 22,000-member pool needs
|
|
two cores); at x8 about 800 (three cores). A node's block verification stays a few milliseconds. The gate's
|
|
remaining margin is what Counter ASIC 3.0 has to spend, and on the unmeasured laptop core it is 6 to 7 ms at x4 and
|
|
about 6 at x8, which is the number the 2019-class measurement (O-1.14) must confirm before x8 is final.
|
|
|
|
### 6.4a The same session on the fixed binary (section 6.6), `with-lock.sh measure`, 22:06:59 to 22:07:06 UTC
|
|
|
|
The verifier of section 6.4 was measured on a binary that carried the inlining regression of section 6.6; after the
|
|
fix, readwidth's binary and the fixed one on the same v2 input in the same minute (checksum 19297e99c7b9a55e), then
|
|
x4 and x8 on the fixed binary; load average 5.5 (the same box, so the ratios of 6.4 stand and the absolute row is
|
|
now the measured one):
|
|
|
|
| Construction | Verifier, ms per unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 |
|
|
|---|---|---|---|
|
|
| v2, readwidth e752fc7's binary | 0.607 / 0.610 | 0.666 | 1 |
|
|
| v2, the fixed binary | 0.609 / 0.611 | 0.657 | 1.00 |
|
|
| x4 (mx4) | 1.238 / 1.237 | 1.396 | 2.03x |
|
|
| x8 (mx8, class v3) | 2.077 / 2.058 | 2.145 | 3.4x |
|
|
| x4, the devnet seeds | 1.240 | 1.289 | |
|
|
| x8, the devnet seeds | 2.058 | 2.144 | |
|
|
|
|
The added cost per unit is the same as on the slow binary (x4 + 0.63 ms, x8 + 1.46 ms: the regression was a
|
|
constant 0.72 ms per unit in the shared item loop), so the ALU reading of 6.4 holds; the ratios against v2 are
|
|
2.0x and 3.4x once v2 is back at 0.61. The 10 ms gate keeps 7.9 ms at x8 (worst cold 2.15 ms) on this core.
|
|
|
|
### 6.5 The daily build per tier, and the x4 / x8 rule
|
|
|
|
The coordinator's rule (21:30 UTC): x8 enters v3 if the per-warp verify stays under 10 ms on one Mac core AND the
|
|
daily 1 GiB build stays under 1 s on every discrete card we own; else x4 with the thin margin stated and x8 named
|
|
as the next lever. The integrated tier is decided beside it (consequences row C23): its build is per prepare, not
|
|
per day, so its consequence is per-day dataset reuse in the workers or a restart per epoch.
|
|
|
|
| Card | Build at x1 | x4 | x8 | Source |
|
|
|---|---|---|---|---|
|
|
| RTX 5090 (PC 1, job run-mixer-x4-pc1-20261005, 22:00 to 22:04 UTC, the worker's `cache ... dataset ... ms` wall line, two dispatches per pack) | 23 to 25 ms | 23 to 25 ms | 23 ms | the PC 1 job (13.4 ms GPU time on 3 October: the wall line carries the launch) |
|
|
| RX 9070 XT (PC 1, gfx1201 on the eGPU, the same job) | 74 ms | 73 to 77 ms | 72 to 76 ms | the PC 1 job |
|
|
| M5 Max, Metal | 13 to 30 ms (the two runs of the memory-hard entry; tonight's run-lock figures 21.7 to 30.2 ms at x4 and 22.0 to 30.0 at x8 say the Mac's build is latency-bound, not mixer-bound) | section 6.4 | section 6.4 | this file |
|
|
| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | 6.9 / 9.4 / 11.7 s prepare total with the 1 GiB build inside | about 28 to 47 s (approximate: scaled x4; the iGPU's build is arithmetic-bound at x1 already) | about 55 to 94 s (approximate) | `docs/plans/epoch-length.md` section 7 (branch ca2-epoch), M11 table |
|
|
| gfx1036 beside WSL build jobs (PC 1) | 55 / 116 / 124 s | about 4 to 8 min (approximate) | about 7 to 17 min (approximate) | same |
|
|
| 8 GB-class discrete card (not owned; about a tenth of the 5090's rate, approximate) | about 0.13 s | about 0.5 s | about 1 s, on the edge of the rule | scaled from the 5090 row, approximate |
|
|
|
|
Decision (coordinator under the delegated rule, 22:05 UTC, on these rows): x8 enters class v3. Both halves pass:
|
|
the verifier at x8 is 2.1 ms per unit on one M5 Max core (6.4a) against the 10 ms gate, and the daily 1 GiB build
|
|
does not move with the mixer on any discrete card we own (5090 23 to 25 ms, 9070 XT 72 to 77 ms, M5 Max 21 ms at
|
|
x1, x4 and x8: latency-bound), 13x to 40x under the 1 s bar. `V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }`;
|
|
the pinned v3 packs are mx8-genesis and mx8-devnet-epoch0 (the latter through the chain path with the era inside the
|
|
class); the x4 packs stay pinned as the candidate's record (generator 2, the class in the id).
|
|
|
|
Reading of the integrated tier: on the discrete cards the rule is settled by the 5090 and 9070 XT rows above. The integrated tier misses the rule at x4 already: a per-prepare build of 28 to 47 s is a tenth to a quarter
|
|
of the 600-DAA-second lead the devnet gives the next program (spec 1.12), and under load it is the whole lead; so
|
|
if x4 or x8 goes in, the iGPU tier needs the workers to build the day's dataset once a day and keep it across
|
|
epochs (today a prepare rebuilds it: `proto-opencl/host.c` prepareTask builds the pair's cache and dataset per
|
|
prepare), or to restart per epoch. The node agent is asked whether the per-day reuse is bounded tonight
|
|
(coordinator, 21:45 UTC); until then the iGPU consequence stands as written.
|
|
|
|
### 6.6 The verifier regression of 0fc0ad1, found and fixed (5 October 2026, 21:59 to 22:07 UTC)
|
|
|
|
The era agent measured the same v2 input with two binaries in one minute: readwidth's 0.604 ms per unit, ca2-v3
|
|
HEAD's 1.33. Bisected under the measure lock (one session, four binaries, two rounds, checksum 19297e99c7b9a55e):
|
|
readwidth e752fc7 0.607 / 0.609; the ca2-v3 seam 6c75dad (before this branch) 0.610 / 0.609; this branch's 0fc0ad1
|
|
1.332 / 1.316; ca2-v3 88dafbc 1.325 / 1.347. So the 2.2x was in 0fc0ad1's `derive_items`, on the version 2 path the
|
|
devnet verifies with, and section 6.4's "loaded box" reading was wrong: the load was real (the same session shows
|
|
it) but the 2x was the code.
|
|
|
|
Variants, each a one-change copy measured against readwidth's binary in the same session:
|
|
|
|
| Variant | ms per unit | Reading |
|
|
|---|---|---|
|
|
| 0fc0ad1 as written (the line mask read from the cache at run time, the loop inlined into `MemhardCpu::fetch`) | 1.33 | the regression |
|
|
| the mask hoisted into a local before the item loop | 1.32 to 1.37 | not the reload |
|
|
| `Cache::line` with the constant mask, on the 0fc0ad1 tree | 0.60 to 0.65 | fixed there |
|
|
| the same constant mask on the merged ca2-v3 tree | 1.33 to 1.46 | not the mask either |
|
|
| one instance per cache size with the mask a constant, `#[inline(always)]` | 1.32 to 1.39 | not the mask |
|
|
| the same instances `#[inline(never)]` | 0.604 / 0.617 / 0.618 / 0.624 | the fix |
|
|
|
|
So it is inlining: the item loop inlined into its callers (`fetch`, `fetch_wide`, `word_at`) runs at 2.2x the
|
|
out-of-line loop, and which small change tips LLVM's decision depends on the rest of the tree (the constant mask
|
|
tipped it on one tree and not on the other). The fix (`memhard::derive_items`): the loop is `derive_items_mask`,
|
|
`#[inline(never)]`, one instance per cache size the growth rule reaches (2^26 to 2^30 words) with the line mask a
|
|
constant, and a run-time-mask instance for every other size (tests). Measured in 6.4a: 0.609 / 0.611 against
|
|
readwidth's 0.607 / 0.610.
|
|
|
|
The class, not the instance: a verifier benchmark with a pinned bound in the crate's CI (the v2 unit at a known
|
|
input against a stored ms-per-unit on a named core, failing on a 1.3x drift) would have caught this at the first
|
|
commit; filed for the next cut (section 9). Until then the era agent's two-binary check (same input, same minute)
|
|
is the rule for every change that touches the item loop.
|
|
|
|
## 7. The chip model
|
|
|
|
`docs/analysis/chip-model-v3.md`.
|
|
|
|
## 8. What is unverified
|
|
|
|
1. The 5090's and the 9070 XT's dataset build at x4 and x8 are measured (6.5, the PC 1 job): both latency-bound,
|
|
under 0.1 s. The gfx1036 is the tier that fails the per-prepare build (6.5), and its consequence (per-day dataset
|
|
reuse in the workers) is with the node agent.
|
|
2. The absolute verifier figures were taken on a loaded core (load average 5.6); the ratios are the measurement
|
|
and the quiet-core figures are scaled. A 2019-class laptop core has not run any construction (O-1.14).
|
|
3. The mixer has had no cryptanalysis (spec 1.8.4); `m` applications with distinct round keys is `m` times the
|
|
work only if no shortcut composes them, which is the same open question as for one application.
|
|
4. The x8 packs are generator 2 with the class in the id (`--class mx8`); if x8 is chosen, the pinned v3 packs are
|
|
re-cut through the seam (`V3_CLASS = MX8`, generator 3) and the tests re-pinned, one commit.
|
|
5. The 5090's rate for the v3 program is the v2 rate by construction (the hash kernel is unchanged, the Mac shows
|
|
27.7 MH/s at v2, x4 and x8); the chip row's denominator stays the readwidth table's 136.1 MH/s until a v3 pack
|
|
runs on the card, which the PC job also gives.
|
|
|
|
## 9. Owed
|
|
|
|
| Item | Owner | When |
|
|
|---|---|---|
|
|
| PC 1 run of the five packs | done 22:04 UTC (run-mixer-x4-pc1-20261005; 6.5 and 6.2a) | |
|
|
| A verifier benchmark with a pinned bound in the crate's CI (section 6.6, the class rule) | ca2-mixer | the next cut |
|
|
| GPU bit-exactness of the re-exported mx8-devnet-epoch0 (the composed class, the era inside) and the mx4 record packs on the PCs (the Mac rows are in 6.2a) | ca2-mixer | the next PC round |
|
|
| The x4 / x8 choice recorded from the rule, then the vectors re-cut once through the seam | done 22:05 UTC: x8 (6.5), V3_CLASS = MX8, mx8 packs re-exported | |
|
|
| `Epoch::chain_dataset_day` wired to the genesis day index in the node (`days_since_genesis(day_index(header), day_index(genesis))`) and `pow_genesis_dataset_log2` in the override | ca2-node | the integration |
|
|
| The spec text of section 2 into `docs/spec/01-lottery-hash.md` 1.8.5 and 1.13.3 (with the v3 vectors into 1.17) | the integration | after the choice |
|
|
| The 2019-class laptop core measurement that fixes the gate (O-1.14) | cryptographer | gate 1 |
|