35 KiB
Mixer x4 and the cache growth rule: the class v3 dataset construction
5 October 2026 (night). Counter ASIC 2.0, layer 6 (option C) and ledger M16's lever, decided by the coordinator
under the project lead's delegation at 22:00 UTC (docs/plans/counter-asic-2-status.md, "22:00 decided"; the project lead confirms for
the public testnet genesis). Branch ca2-mixer. Worker: ca2-mixer (cryptographer's lane).
What this changes, in one line: under program class v3 every mixer application of the dataset item derivation becomes four applications with distinct round keys, the eight dependent cache reads per item stay eight, and the 256 MiB cache doubles on the days the dataset doubles (years 4 and 12). Version 2 is byte-identical: the two pinned packs re-export without a changed byte (section 5).
1. Why this form
The recompute attacker of docs/analysis/m16-recompute-attacker-2026-10-05.md holds the 256 MiB cache on a die
and derives every dataset word instead of reading it: 128 items per hash at about 1,170 integer operations and 8
dependent cache reads each. Its cost is linear in operations per item; the honest miner pays the mixer once a day
in the dataset build and never per hash; the verifier pays it per item it checks. The multiplier m is the one
parameter that moves the attacker and leaves the honest hash rate untouched.
Two shapes give the attacker 4x the operations:
| Shape | Mixer applications per item | Dependent cache reads per item | What grows for the verifier | What grows for the chip | What grows for the honest build |
|---|---|---|---|---|---|
A, chosen: m = 4 applications per round, 8 rounds |
36 | 8 | the ALU part only; the latency part (8 dependent misses per item) unchanged | integer operations 4x; SRAM bandwidth unchanged (1,024 reads per hash) | 4x the mixer arithmetic, same reads |
| B, alternative: 32 rounds of one application and one read | 33 | 32 | both parts: 4x the dependent misses per item, so about 4x the latency-bound time (spec 1.11: about 8 x 100 ns per item in series without interleaving) | integer operations 3.7x and SRAM bandwidth 4x (4,096 reads per hash, 60 TB/s to match one 5090 at the M16 rate) | 4x the reads too; the GPU build becomes latency-bound at 4x the dependent line fetches |
Shape A is chosen because the verifier's latency part is the part the 10 ms gate protects (section 1.11: the distinct items of a unit are derived with their chains interleaved so the 8 misses of each item overlap across up to 32 items; shape B would make that 32 misses deep). Shape B is implemented nowhere; the rule in the brief: if shape A's measured verifier time exceeds 4.8 ms per warp on one M5 Max core, measure both and recommend. Section 6 has the measurement; it is under that bound, so B stays unimplemented.
2. Spec text (replaces 1.8.5 and 1.13.3 under class v3; v2 text unchanged)
1.8.5 Item derivation and dataset mapping
Item t (16 words) under mixer multiplier m (m = 1 for program class v2, m = 4 for class v3; a class
parameter, LoadClass::mixer_mult):
s[i] = K[i] for i in 0..7
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
for r in 0..7:
for j in 0..m-1:
s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache
s[i] = s[i] XOR cache[line a][i] for i in 0..15
for j in 0..m-1:
s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
item(t) = s
M(s, rk) is the mixer of 1.8.4 with round key rk; under m = 1 the keys are (r + 1) * 0x9E3779B9 and
9 * 0x9E3779B9, the version 2 text exactly. The round keys of the 9 m applications are the first 9 m values
of the version 2 key sequence, all distinct (the sequence is k * 0x9E3779B9 for k = 1 .. 9 m, and
0x9E3779B9 is odd, so no two of the first 2^32 keys coincide). Eight dependent cache reads per item at every
m (ITEM_ROUNDS = 8, prototype value): the address of read r depends on every earlier read. 9 m mixer
applications, 36 under class v3, about 4,700 integer operations per item (130 per application, 1.8.4).
dataset[w] = item(w >> 4)[w AND 15]. A dataset of 2^D words is the prefix of items 0 .. 2^(D-4) - 1, so an
item has the same value at every dataset size; and a cache of 2^C words is the prefix of segments of every
larger cache (1.8.3 fills segments independently of the cache size), but an item's value depends on C through
the line mask, so the item changes on the day the cache doubles.
Source: igneum-pow/src/memhard.rs (derive_items, round_key_mult, Shape), the emitted mh_item of
memhard.h, memhard.metal and kernel.cl (emit.rs, emit_memhard_core: the m loop is emitted only for
m > 1, so every version 2 pack keeps its text).
1.13.3 Dataset growth (class v3: option (b) with the cache tied to it, "option C")
Designed: 2 GiB at genesis plus 0.5 GiB per year. The linear schedule in bytes, G x (1 + 86,400 d / (4 x 31,536,000)) = G x (1 + d / 1,460) for the genesis size G and the chain day d (DAA days since genesis,
section 1.12), doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at day 10,220
(year 28). Rule (Designed, decided 5 October 2026 for class v3):
doublings(d) = floor(log2(1 + d / 1460)) integer division, then integer log2
dataset_words(d) = 2^(D_0 + doublings(d)) D_0 = 29 designed (2 GiB), 28 on the devnet (1 GiB); capped at 32
cache_words(d) = 2^(26 + doublings(d)) 256 MiB, 512 MiB from year 4, 1 GiB from year 12
Power-of-two sizes only (option (b)), so every load keeps the src AND MASK form of 1.14 item 2 and the cache
line index keeps s[0] AND mask. The cache doubles exactly when the dataset doubles ("option C",
docs/analysis/sram-mirror.md section 7): the cache's job is to stay above any GPU's last-level cache and that
needs growth; the recompute attacker is priced by the mixer, not by the cache (section 7 below).
d is day_index(header.timestamp) - day_index(genesis.timestamp) with day_index = timestamp_ms / 86,400,000
(bind::day_index, the interim day rule), clamped at 0 (memhard::days_since_genesis). Under class v2 nothing
grows: the cache is 2^26 words and the dataset the genesis size on every day.
Chain day d |
Years | doublings |
Cache words | Cache | Dataset words (devnet D_0 = 28) |
Dataset (designed D_0 = 29) |
Verifier cache fill, one M5 Max core (measured at 256 MiB, section 6, scaled linearly) |
|---|---|---|---|---|---|---|---|
| 0 to 1,459 | 0 to 4 | 0 | 2^26 | 256 MiB | 2^28 (1 GiB) | 2 GiB | 0.18 s |
| 1,460 to 4,379 | 4 to 12 | 1 | 2^27 | 512 MiB | 2^29 (2 GiB) | 4 GiB | 0.36 s |
| 4,380 to 10,219 | 12 to 28 | 2 | 2^28 | 1 GiB | 2^30 (4 GiB) | 8 GiB | 0.72 s |
| 10,220 to 21,899 | 28 to 60 | 3 | 2^29 | 2 GiB | 2^31 (8 GiB) | 16 GiB | 1.4 s |
| 21,900 and on | 60 and on | 4 | 2^30 | 4 GiB | 2^32 (16 GiB, the index cap) | 2^32 words, the cap | 2.9 s |
Test: memhard::tests::growth_schedule_table pins every row and the day before each step. The devnet pack
igneum-devnet-v4-epoch0 is day 20,730 of the Unix count against genesis day 20,729, d = 1, so every existing
size and vector stands.
Consequences for the tiers (the rule of 5 October): a verifier (any node, any pool core) holds 512 MiB from year 4 and 1 GiB from year 12, and fills it once a day in under a second on one 2026 core (the table); a miner's card holds the dataset, 4 GiB from year 4 and 8 GiB from year 12 on the designed schedule, so an 8 GB card mines until year 12 and a 16 GB card until year 28 (the cache is not in the card's working set at hash time: it is built, the dataset built from it, and dropped). Those dates are the design document's own schedule restated as steps; option (a) would have faded a 4 GiB card out in year 4 instead of year 4.
3. Interfaces
| Item | Where | Note |
|---|---|---|
LoadClass { mixer_mult: u8, growth: bool }, LoadClass::MX4 ("mx4"), with_mixer(m, growth), v2_loads(), takes_width_roll() |
generator.rs |
a class with v2 loads takes no width roll: its program stream is version 2's draw for draw, so the v3 program of a seed is the v2 program of that seed, only the dataset differs |
Shape { mixer_mult, cache_log2_words }, Shape::for_class_day(class, d), MixParams.shape, Cache::fill_log2(key, log2), round_key_mult(r, j, m) |
memhard.rs |
Shape::V2 is version 2 |
growth_doublings(d), cache_log2_words(d), dataset_log2_words(D_0, d), days_since_genesis(day, genesis_day) |
memhard.rs |
the schedule, one function and its two sizes |
DatasetSource::{new_shape, from_key_shape, shape}, Epoch::new_class_day, Epoch::from_seed_bytes_day(epoch, day, label, class, d, D_0) |
verify.rs |
the day-sized entries; the v2 entries are unchanged and build the v2 shape |
IGNEUM_MIXER_MULT, IGNEUM_CLASS_MIXER_MULT, IGNEUM_CACHE_GROWTH in program.h; "mixer_mult", "cache_growth", the "item" string in program.json |
emit.rs |
written only for a class with m != 1 or growth, so v2 packs do not change |
packfile.h mixerMult; packbench and the OpenCL host print the multiplier and the cache size |
the three hosts | the kernels carry the construction in their text (one emitter, three dialects); the hosts size the cache from IGNEUM_CACHE_LOG2_WORDS already (packbench.swift line 56, host.cu line 65, host.c line 1966) |
igneum-pow --class mx4 [--days d] on every command |
main.rs |
--days sizes the cache for a growth class |
Under the ca2-v3 seam (ProgramClass::V3, V3_CLASS), the integration sets V3_CLASS = LoadClass::MX4; the
chain's day-sized dataset needs the day index, so Epoch::chain_dataset(day, class) builds the genesis-size
cache and a chain_dataset_day(day_bytes, class, d, D_0) beside it is the growth entry (section 9, owed to the
node agent).
4. Vectors (class v3, proto-cuda/packs-ca2-mixer/)
Produced by igneum-pow export --program-class v3 (the Rust CPU interpreter, 5 October 2026, commit 66eeba3) and
checked on the GPUs in section 6. The program of each pack is the version 2 program of the same seed instruction
for instruction (tests/packs.rs, v3_packs_are_the_v2_seeds_under_mixer_x4); the cache is the version 2 cache
(day 0 of the growth rule); the dataset words and the hashes are new. The 64 sampled indices are those of every
pack (emit::sample_indices).
Pack mx4-genesis (seed igneum-genesis, day 2026-10-03, generator 3, class mx4, program id e323b9dcaf283a6f, 2^28 words, 2^26-word cache, cache FNV-1a 64 48c4f5bf24166b2e as under v2):
dataset words 0..15 (item 0):
61ff2180 0d4c7e6c 2177d443 60df9025 cf8b2e10 63675bfb 25289e58 9c45dc42
2d271c54 9652369b 2dd77508 5921392c 3afa60ee c640ad68 f2bb56ff cfa46438
dataset[0x0fffffff] = 5020180e
sampled words (the first 8 of the 64 in vectors.json):
dataset[59471966] = de85726d dataset[217795994] = 7cfc31c7 dataset[208353206] = d3cd5289 dataset[42483309] = 6ebbeef1
dataset[172547758] = 9858413b dataset[148076330] = e786f141 dataset[183853158] = 64f13833 dataset[214389424] = ca229d04
unit at base nonce 0, lanes 0..31:
63acd2d273f475ba e929c78b34b80d4b 0b1011cb19982558 1457a0df5497aa11
957d0f3bb71d98fb ac16901e6e6f6057 8ea1c6279f4b177a f28146e60bd08ba9
fe5b8cfe87f8e65b 49f87240566ace62 6ef6d6b7bdea8e41 46d9c0dc29a97b9c
111fe30128db9398 66dc39084f0946d4 8ee11bdfd35fecf2 2861fcfc75db6677
31c7667d4bde8556 c5989c48858b4ce0 276395e734a9d30d 84217b41e91368ff
3604861e34d9f697 9f51d8ee16bf3639 c89e47bafa84401c 7ae78c1f10b70e19
0b8c947157a29a48 d67192e8cfb43842 05a4c6d182c8c675 188e2661f3263f2e
a2df24238f7fea2e ed69ea7e13ad3a48 2d44ae509bab91b8 adad61931ea4fb70
unit at base nonce 4096: lane 0 edd508ac57e5699a, lane 1 4eefd56d526cdaeb, lane 31 8892f8604733b1e0
unit at base nonce 1000000: lane 0 8b3183778a49f59c, lane 1 1831b72a8797e895, lane 31 75eae55eba53a506
Pack mx4-devnet-epoch0 (seed the devnet genesis hash, day bytes of 2026-10-04, generator 3, class mx4, program id 73bcbfe8ccf988f1, 2^28 words, 2^26-word cache, cache FNV-1a 64 448274a57f508cbc as under v2):
dataset words 0..15 (item 0):
afe80d67 b9fbd029 6c79f193 95139ad9 96310aff 4609f8b1 75279e63 28235be1
47b17dcb 718e0ef2 a52588c8 a8bf49d5 19cf243e 5ec8905e a4851f66 af9cd9f3
dataset[0x0fffffff] = e6a99c7a
sampled words (the first 8 of the 64 in vectors.json):
dataset[59471966] = 57642b58 dataset[217795994] = c279badd dataset[208353206] = cbccbaad dataset[42483309] = 32cce392
dataset[172547758] = 71fdb4c6 dataset[148076330] = f5a268ce dataset[183853158] = 3ca1d676 dataset[214389424] = 7977b03d
unit at base nonce 0, lanes 0..31:
212c6442b51e87ae c374795c00839331 b6036a220a98f4b3 eb8b8013e637367b
db5866e9b73930fd f3f3d01f46e90333 9d913991ab8ed428 7ccb1d8fa100a800
3cf45ba44f09a0fe 91acf48ef1a63082 6ea46c69fb082f99 581f0218977a9d72
9a4623a5c62ddf2d ab6eb5e768f0feb4 07b70bdccf8aca12 d666311ae5e4311e
53114757d669f0a4 bd5d6ace87ce2ce4 fb712015e8189192 a32cec81103e134b
83f3d18c3289c124 fe29f1984b132b3d c9ffcf4e3774497a ac99c9243dc63809
d78a7e8217a32f3c ab81ad63d242fc31 0e4c30b7e00024af ce014289fff6778d
64a1292e2a8b4a91 d5b8c90e681d7e3a 06078117673030fd 51bf77b280173930
unit at base nonce 4096: lane 0 3c797978566b5950, lane 1 7c759e60185b6411, lane 31 96a903eb9a0ca390
unit at base nonce 1000000: lane 0 d5a8da0568df8ee7, lane 1 68cb69c04208285a, lane 31 f8ca84a1a5d78cf5
5. The v2 path is byte-identical
cargo test --test packs regenerates every file of igneum-genesis-mh and igneum-devnet-v4-epoch0 from
program.json and compares byte for byte (emitted_sources_match_all_packs, export_pack_matches_all_packs);
section 6 also records a fresh igneum-pow export of both packs diffed against the checked-in directories.
6. Measurements
Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0, 5 October 2026 (night), other agents' builds running beside every
run; a timing row says which lock it ran under (measure is exclusive; run and build are not timings).
6.1 The v2 path, byte for byte (no lock needed)
igneum-pow export --seed igneum-genesis --day 2026-10-03 and igneum-pow export --epoch-hex edc4fa84...fb07 --day-hex 69676e65756d2d6461792ffa50000000000000 on commit 66eeba3, diff -r against
proto-cuda/packs/igneum-genesis-mh and igneum-devnet-v4-epoch0: IDENTICAL, both (twelve files each). The
crate tests regenerate the same files and compare them on every run (tests/packs.rs, 12 of 12 pass).
6.2 Bit-exactness of the class v3 construction on the GPUs (with-lock.sh run, 22:05 UTC)
packbench --pack <dir> --batches 1 --batch-log2 24 --group 256 (Metal, built from this branch) and
igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 1 --batch-log2 24 (Apple OpenCL, built from
this branch's host.c). Vectors are the Rust interpreter's; the fingerprint is FNV-1a 64 over the 2^24 outputs at
base nonce 0.
| Pack | Harness | Cache FNV-1a 64 | Dataset head, word MASK, 64 samples | Vectors standalone / in batch | Fingerprint 2^24 | MH/s (GPU time; not a measurement, the run lock) |
|---|---|---|---|---|---|---|
| mx4-genesis | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 6f48d5a2aa0dbe5f | 27.6 |
| mx4-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 6f48d5a2aa0dbe5f | 27.7 (wall) |
| mx4-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 73caaebb28e808fe | 27.5 |
| mx4-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 73caaebb28e808fe | 27.6 (wall) |
| mx8-genesis (the x8 candidate, 21:45 UTC) | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 7c28cfb06c5c65a9 | 27.7 |
| mx8-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 7c28cfb06c5c65a9 | 27.6 (wall) |
| mx8-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | bbb183f72692f840 | 27.7 |
| mx8-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | bbb183f72692f840 | 27.7 (wall) |
Reading: the Rust interpreter, Metal and Apple OpenCL agree on the v3 dataset (head, MASK word, 64 samples), on every vector lane and on the 2^24-output fingerprint of each pack; the hash rate is the v2 rate (27.7 MH/s on this card tonight, readwidth table), as it must be: the hash kernel only loads, the mixer is paid in the build.
Dataset build under the run lock (indicative only; the measured rows are in 6.4): Metal 30.2 ms GPU (mx4-genesis) and 21.7 ms GPU (mx4-devnet-epoch0) for 1 GiB; Apple OpenCL 55 and 56 ms wall. The readwidth entry's v2 figure on this card is 0.6 to 2.1 ms cache fill and a 1 GiB build the bench-log's memory-hard entry puts at 13 to 30 ms; the x4 build on the GPU is the row the measure lock will settle.
6.2a The packs as pinned after the x8 decision (with-lock.sh run, 22:12 UTC, the fixed binary's exports)
| Pack | How exported | Program id | Unit 0 lane 0 | Fingerprint 2^24 (Metal = Apple OpenCL) | Vectors, self-tests |
|---|---|---|---|---|---|
| mx8-genesis | --program-class v3 (generator 3, V3_CLASS = mx8) |
e323b9dcaf283a6f | 19b56348bc85304d | 7c28cfb06c5c65a9 | 3/3 + 3/3, 96 of 96, PASS |
| mx8-devnet-epoch0 | the chain path, --program-class v3 --era-hex <genesis> (the era drawn inside the class, load class mx8-erad810f22d) |
73bcbfe8ccf988f1 | d424577fce4a7a60 | 90f794dd556f7a3b | 3/3 + 3/3, 96 of 96, PASS |
| mx4-genesis (the x4 record) | --class mx4 (generator 2, the class in the id) |
951b89584750bd75 | 63acd2d273f475ba | 6f48d5a2aa0dbe5f | 3/3 + 3/3, 96 of 96, PASS |
| mx4-devnet-epoch0 (the x4 record) | --class mx4 on the chain seeds |
(program.json) | 212c6442b51e87ae | 73caaebb28e808fe | 3/3 + 3/3, 96 of 96, PASS |
The PC 1 rows of 6.5 ran the earlier mx8-devnet-epoch0 (the load-class export, fingerprint bbb183f72692f840, era not inside); the composed pack's PC fingerprints are owed with the next PC round (section 9).
6.3 Soundness suite on the v3 construction (commit 66eeba3 and after; with-lock.sh build for cargo, run for the GPU)
| Suite | Command | Result |
|---|---|---|
| Crate lib tests (memhard schedule table, the by-hand multiplied mixer, the seam, the generator) | cargo test -j4 --release |
44 of 44 pass |
Pinned packs, v2 and v3 (tests/packs.rs: programs, ids, dataset words, 96 vectors per pack, every emitted file byte for byte, 16 masked loads per kernel, the v3 packs as the v2 seeds under mixer x4) |
same | 12 of 12 pass |
Scratch soundness tests of branch ca2-soundness (cherry-pick 0d8f745, one conflict in packbench.swift's RESULT line resolved by hand, era_bytes: None added to the edge program literal) |
cargo test -j4 --release --test scratch |
7 of 7 pass: bijections, re-hit rates, 56 of 56 edge units, 42 of 42 emitted scr kernels, 200 scratch programs and 800 units on the CPU |
v3 fuzz on the CPU (tests/mixer.rs): 200 programs through the seam, the contract on every instruction (each is the v2 program of its seed), 4 units each across the 32-bit range with one unit in the top 256 nonces, interpreted twice |
IGNEUM_MIXER_PACKS_OUT=<dir> cargo test -j4 --release --test mixer |
200 of 200, 800 of 800 units; 200 packs written for the GPU runs (82 s with the three cache fills) |
v3 stats beside v2 (tests/mixer.rs): 8,192 outputs per seed, bit balance, single-bit avalanche within the unit and across units, duplicates |
same | igneum-genesis v3: avalanche 49.99 percent, worst bit z 1.92, 0 duplicates (v2: 49.87, z 2.25); igneum-genesis/stats1 v3: 49.97, z 3.09 (v2: 49.98, z 2.30) |
v3 edge (tests/mixer.rs): items 0, 1, 2^28 - 1 and 2^32 - 1 by hand at m = 1, 2, 4, 8 on a 2^14-word cache; words 0, 15, 16, 17, MASK - 1, MASK through the interpreter's fetch path; the index wrap at MASK + 1 |
same | pass |
v3 determinism (tests/mixer.rs): two independent epochs, every vector and every emitted file equal, and equal to the pinned pack |
same | pass |
| Metal and Apple OpenCL on the two pinned v3 packs | section 6.2 | 3/3 standalone, 3/3 in batch, 96 of 96 lanes, dataset head, MASK word and 64 samples, one fingerprint per pack across both harnesses |
| Metal fuzz: the 200 packs, 4 units each standalone and the top-256 unit inside a 512-nonce batch at base 4,294,967,040; every tenth pack on Apple OpenCL as well | packbench --pack <dir> --batches 1 --batch-log2 9 --batch-base 4294967040 (Metal), igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 1 --batch-log2 10 (Apple OpenCL), with-lock.sh run, 21:22 to 21:24 UTC |
Metal 200 of 200 packs PASS (800 of 800 standalone units, 200 of 200 inside the wrapping window, cache and dataset self-tests on every pack); Apple OpenCL 20 of 20 packs PASS (the three vectors.h units, the self-tests); a first run with a packbench built before the --batch-base cherry-pick reported 200 of 200 FAIL on an empty RESULT line and was read as such (the watcher rule), the harness rebuilt and the run repeated |
6.4 Timings (with-lock.sh measure, one session, 21:40:12 to 21:40:23 UTC, commit 504cae4)
Script measure-v3.sh (session scratchpad): igneum-pow bench --seed igneum-genesis --day 2026-10-03 --warps 50
(v2), ... --program-class v3 (x4), ... --class mx8 (x8), two rounds each, then the devnet seeds, then
packbench --pack <dir> --batches 2 --batch-log2 22 --group 256 on igneum-genesis-mh, mx4-genesis and mx8-genesis,
two rounds. The lock was exclusive among the agents' builds and measurements, but the box was not quiet: load
average 5.6 (one minute) and 26 (fifteen minutes) at the start, from unlocked processes (the devnet node, other
agents' editors); the v2 row reads 1.31 to 1.36 ms where the quiet readwidth night read 0.604 to 0.626. So the
absolute numbers below are a loaded-core figure, about 2.2x the quiet one, and the ratios between the rows are the
measurement (two rounds within 4 percent). A quiet-box re-run is owed (section 9).
| Construction | Verifier, ms per 32-lane unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | 256 MiB cache fill, one core | Metal 1 GiB dataset build, GPU ms (round 1 / round 2) |
|---|---|---|---|---|---|
| v2 (igneum-genesis) | 1.361 / 1.310 | 1.579 | 1 | 172.1 / 172.6 ms | 29.7 / 21.0 |
| x4 (mx4, class v3) | 1.956 / 1.923 | 2.043 | 1.45x | 175.3 / 172.3 ms | 20.9 / 21.0 |
| x8 (mx8) | 2.785 / 2.790 | 2.942 | 2.09x | 172.3 / 172.3 ms | 21.9 / 21.9 |
| x4, the devnet seeds (mx4-devnet-epoch0) | 1.923 | 2.012 | 173.9 ms | ||
| x8, the devnet seeds | 2.972 | 2.885 | 173.6 ms |
Reading. The verifier's ALU part is what grows: x4 adds 0.6 ms per unit for 27 more mixer applications on each of 4,096 items (110,592 applications, about 5.5 ns each on this core, the lanes' chains interleaved), x8 another 0.85 ms for 36 more; the latency part (8 dependent misses per item) is the same in every row, which is why the measured ratios are 1.45x and 2.1x and not the 4x and 8x of the M16 table's scaling. Shape B (32 rounds of one read, section 1) would have multiplied the latency part too; x4 is under the 4.8 ms bar even on the loaded core, so B stays unimplemented. The cache fill does not depend on the mixer (it is the ChaCha chain): 172 to 175 ms, the spec's 175 to 181 ms of 1.8.3. The Metal 1 GiB build does not move with the mixer at all (21 ms at v2, x4 and x8 once warm; the 29.7 ms first v2 run is the first-touch cost the hosts fill twice for): on this card the build is bound by the 8 dependent cache-line reads per item, not by the arithmetic, so the Mac says nothing about whether the 5090's or the 9070 XT's build is arithmetic-bound; that is the PC job (section 8).
Against the x4 / x8 rule (section 6.5): the verifier half passes for x8 with 7.1 ms of the 10 ms gate to spare on this loaded core (worst cold 2.94 ms; the quiet-core figure would be about 1.3 ms, scaled by the 2.2x of the v2 row, approximate); x4 leaves 8.0 ms. The build half waits on the PC rows.
6.5 Verification throughput per tier (consequences row C19), from the loaded-core figures above
Warps verified per second on one core = 1,000 / (ms per warp); a pool core verifying members' shares handles that many shares per second; a node verifies a block with one unit (plus the header path, under 0.1 ms, not measured here); IBD over the 108,000-header pruning window (spec 02) on one core = 108,000 x ms per warp.
| Figure | v2 | x4 | x8 | Note |
|---|---|---|---|---|
| ms per warp, steady (this session, loaded core) | 1.33 | 1.94 | 2.79 | avg of the two rounds |
| ms per warp, quiet M5 Max core (scaled by 0.604 / 1.33 = 0.45, approximate) | 0.60 | 0.88 | 1.26 | the readwidth night's v2 figure is measured; x4 and x8 scaled |
| ms per warp, 2019-class laptop core (approximate: 2.5x the quiet M5 Max figure, the ratio the design document assumes for the gate; unmeasured, O-1.14) | 1.5 | 2.2 | 3.2 | the figure that fixes the gate is a measurement, not this row |
| Shares per second per core (loaded / quiet, approximate) | 750 / 1,660 | 515 / 1,140 | 358 / 790 | spec 09 section 9.8 item 5 carried 2,270 at v2; re-cut from the quiet row: 1,660 |
| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second), loaded / quiet | 2.9 / 1.3 | 4.3 / 1.9 | 6.1 / 2.8 | |
| Node: worst cold single unit (loaded core) | 1.58 ms | 2.04 ms | 2.94 ms | per block |
| IBD over 108,000 headers on one core, loaded / quiet, minutes | 2.4 / 1.1 | 3.5 / 1.6 | 5.0 / 2.3 | laptop (approximate): 2.7 / 4.0 / 5.8 min; a seed VM core (unmeasured) sits between the laptop and the quiet M5 Max |
| Margin left under the 10 ms gate for Counter ASIC 3.0 (worst cold, loaded core) | 8.4 ms | 8.0 ms | 7.1 ms | on the 2019-class laptop row (approximate) 7.5 / 6.8 / 5.9 ms steady |
Reading: at x4 a pool core serves about 1,100 shares per second on a quiet 2026 core (a 22,000-member pool needs two cores); at x8 about 800 (three cores). A node's block verification stays a few milliseconds. The gate's remaining margin is what Counter ASIC 3.0 has to spend, and on the unmeasured laptop core it is 6 to 7 ms at x4 and about 6 at x8, which is the number the 2019-class measurement (O-1.14) must confirm before x8 is final.
6.4a The same session on the fixed binary (section 6.6), with-lock.sh measure, 22:06:59 to 22:07:06 UTC
The verifier of section 6.4 was measured on a binary that carried the inlining regression of section 6.6; after the fix, readwidth's binary and the fixed one on the same v2 input in the same minute (checksum 19297e99c7b9a55e), then x4 and x8 on the fixed binary; load average 5.5 (the same box, so the ratios of 6.4 stand and the absolute row is now the measured one):
| Construction | Verifier, ms per unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 |
|---|---|---|---|
| v2, readwidth e752fc7's binary | 0.607 / 0.610 | 0.666 | 1 |
| v2, the fixed binary | 0.609 / 0.611 | 0.657 | 1.00 |
| x4 (mx4) | 1.238 / 1.237 | 1.396 | 2.03x |
| x8 (mx8, class v3) | 2.077 / 2.058 | 2.145 | 3.4x |
| x4, the devnet seeds | 1.240 | 1.289 | |
| x8, the devnet seeds | 2.058 | 2.144 |
The added cost per unit is the same as on the slow binary (x4 + 0.63 ms, x8 + 1.46 ms: the regression was a constant 0.72 ms per unit in the shared item loop), so the ALU reading of 6.4 holds; the ratios against v2 are 2.0x and 3.4x once v2 is back at 0.61. The 10 ms gate keeps 7.9 ms at x8 (worst cold 2.15 ms) on this core.
6.5 The daily build per tier, and the x4 / x8 rule
The coordinator's rule (21:30 UTC): x8 enters v3 if the per-warp verify stays under 10 ms on one Mac core AND the daily 1 GiB build stays under 1 s on every discrete card we own; else x4 with the thin margin stated and x8 named as the next lever. The integrated tier is decided beside it (consequences row C23): its build is per prepare, not per day, so its consequence is per-day dataset reuse in the workers or a restart per epoch.
| Card | Build at x1 | x4 | x8 | Source |
|---|---|---|---|---|
RTX 5090 (PC 1, job run-mixer-x4-pc1-20261005, 22:00 to 22:04 UTC, the worker's cache ... dataset ... ms wall line, two dispatches per pack) |
23 to 25 ms | 23 to 25 ms | 23 ms | the PC 1 job (13.4 ms GPU time on 3 October: the wall line carries the launch) |
| RX 9070 XT (PC 1, gfx1201 on the eGPU, the same job) | 74 ms | 73 to 77 ms | 72 to 76 ms | the PC 1 job |
| M5 Max, Metal | 13 to 30 ms (the two runs of the memory-hard entry; tonight's run-lock figures 21.7 to 30.2 ms at x4 and 22.0 to 30.0 at x8 say the Mac's build is latency-bound, not mixer-bound) | section 6.4 | section 6.4 | this file |
| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | 6.9 / 9.4 / 11.7 s prepare total with the 1 GiB build inside | about 28 to 47 s (approximate: scaled x4; the iGPU's build is arithmetic-bound at x1 already) | about 55 to 94 s (approximate) | docs/plans/epoch-length.md section 7 (branch ca2-epoch), M11 table |
| gfx1036 beside WSL build jobs (PC 1) | 55 / 116 / 124 s | about 4 to 8 min (approximate) | about 7 to 17 min (approximate) | same |
| 8 GB-class discrete card (not owned; about a tenth of the 5090's rate, approximate) | about 0.13 s | about 0.5 s | about 1 s, on the edge of the rule | scaled from the 5090 row, approximate |
Decision (coordinator under the delegated rule, 22:05 UTC, on these rows): x8 enters class v3. Both halves pass:
the verifier at x8 is 2.1 ms per unit on one M5 Max core (6.4a) against the 10 ms gate, and the daily 1 GiB build
does not move with the mixer on any discrete card we own (5090 23 to 25 ms, 9070 XT 72 to 77 ms, M5 Max 21 ms at
x1, x4 and x8: latency-bound), 13x to 40x under the 1 s bar. V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 };
the pinned v3 packs are mx8-genesis and mx8-devnet-epoch0 (the latter through the chain path with the era inside the
class); the x4 packs stay pinned as the candidate's record (generator 2, the class in the id).
Reading of the integrated tier: on the discrete cards the rule is settled by the 5090 and 9070 XT rows above. The integrated tier misses the rule at x4 already: a per-prepare build of 28 to 47 s is a tenth to a quarter
of the 600-DAA-second lead the devnet gives the next program (spec 1.12), and under load it is the whole lead; so
if x4 or x8 goes in, the iGPU tier needs the workers to build the day's dataset once a day and keep it across
epochs (today a prepare rebuilds it: proto-opencl/host.c prepareTask builds the pair's cache and dataset per
prepare), or to restart per epoch. The node agent is asked whether the per-day reuse is bounded tonight
(coordinator, 21:45 UTC); until then the iGPU consequence stands as written.
6.6 The verifier regression of 0fc0ad1, found and fixed (5 October 2026, 21:59 to 22:07 UTC)
The era agent measured the same v2 input with two binaries in one minute: readwidth's 0.604 ms per unit, ca2-v3
HEAD's 1.33. Bisected under the measure lock (one session, four binaries, two rounds, checksum 19297e99c7b9a55e):
readwidth e752fc7 0.607 / 0.609; the ca2-v3 seam 6c75dad (before this branch) 0.610 / 0.609; this branch's 0fc0ad1
1.332 / 1.316; ca2-v3 88dafbc 1.325 / 1.347. So the 2.2x was in 0fc0ad1's derive_items, on the version 2 path the
devnet verifies with, and section 6.4's "loaded box" reading was wrong: the load was real (the same session shows
it) but the 2x was the code.
Variants, each a one-change copy measured against readwidth's binary in the same session:
| Variant | ms per unit | Reading |
|---|---|---|
0fc0ad1 as written (the line mask read from the cache at run time, the loop inlined into MemhardCpu::fetch) |
1.33 | the regression |
| the mask hoisted into a local before the item loop | 1.32 to 1.37 | not the reload |
Cache::line with the constant mask, on the 0fc0ad1 tree |
0.60 to 0.65 | fixed there |
| the same constant mask on the merged ca2-v3 tree | 1.33 to 1.46 | not the mask either |
one instance per cache size with the mask a constant, #[inline(always)] |
1.32 to 1.39 | not the mask |
the same instances #[inline(never)] |
0.604 / 0.617 / 0.618 / 0.624 | the fix |
So it is inlining: the item loop inlined into its callers (fetch, fetch_wide, word_at) runs at 2.2x the
out-of-line loop, and which small change tips LLVM's decision depends on the rest of the tree (the constant mask
tipped it on one tree and not on the other). The fix (memhard::derive_items): the loop is derive_items_mask,
#[inline(never)], one instance per cache size the growth rule reaches (2^26 to 2^30 words) with the line mask a
constant, and a run-time-mask instance for every other size (tests). Measured in 6.4a: 0.609 / 0.611 against
readwidth's 0.607 / 0.610.
The class, not the instance: a verifier benchmark with a pinned bound in the crate's CI (the v2 unit at a known input against a stored ms-per-unit on a named core, failing on a 1.3x drift) would have caught this at the first commit; filed for the next cut (section 9). Until then the era agent's two-binary check (same input, same minute) is the rule for every change that touches the item loop.
7. The chip model
docs/analysis/chip-model-v3.md.
8. What is unverified
- The 5090's and the 9070 XT's dataset build at x4 and x8 are measured (6.5, the PC 1 job): both latency-bound, under 0.1 s. The gfx1036 is the tier that fails the per-prepare build (6.5), and its consequence (per-day dataset reuse in the workers) is with the node agent.
- The absolute verifier figures were taken on a loaded core (load average 5.6); the ratios are the measurement and the quiet-core figures are scaled. A 2019-class laptop core has not run any construction (O-1.14).
- The mixer has had no cryptanalysis (spec 1.8.4);
mapplications with distinct round keys ismtimes the work only if no shortcut composes them, which is the same open question as for one application. - The x8 packs are generator 2 with the class in the id (
--class mx8); if x8 is chosen, the pinned v3 packs are re-cut through the seam (V3_CLASS = MX8, generator 3) and the tests re-pinned, one commit. - The 5090's rate for the v3 program is the v2 rate by construction (the hash kernel is unchanged, the Mac shows 27.7 MH/s at v2, x4 and x8); the chip row's denominator stays the readwidth table's 136.1 MH/s until a v3 pack runs on the card, which the PC job also gives.
9. Owed
| Item | Owner | When |
|---|---|---|
| PC 1 run of the five packs | done 22:04 UTC (run-mixer-x4-pc1-20261005; 6.5 and 6.2a) | |
| A verifier benchmark with a pinned bound in the crate's CI (section 6.6, the class rule) | ca2-mixer | the next cut |
| GPU bit-exactness of the re-exported mx8-devnet-epoch0 (the composed class, the era inside) and the mx4 record packs on the PCs (the Mac rows are in 6.2a) | ca2-mixer | the next PC round |
| The x4 / x8 choice recorded from the rule, then the vectors re-cut once through the seam | done 22:05 UTC: x8 (6.5), V3_CLASS = MX8, mx8 packs re-exported | |
Epoch::chain_dataset_day wired to the genesis day index in the node (days_since_genesis(day_index(header), day_index(genesis))) and pow_genesis_dataset_log2 in the override |
ca2-node | the integration |
The spec text of section 2 into docs/spec/01-lottery-hash.md 1.8.5 and 1.13.3 (with the v3 vectors into 1.17) |
the integration | after the choice |
| The 2019-class laptop core measurement that fixes the gate (O-1.14) | cryptographer | gate 1 |