igneum/docs/plans/mixer-x4.md

35 KiB

Mixer x4 and the cache growth rule: the class v3 dataset construction

5 October 2026 (night). Counter ASIC 2.0, layer 6 (option C) and ledger M16's lever, decided by the coordinator under the project lead's delegation at 22:00 UTC (docs/plans/counter-asic-2-status.md, "22:00 decided"; the project lead confirms for the public testnet genesis). Branch ca2-mixer. Worker: ca2-mixer (cryptographer's lane).

What this changes, in one line: under program class v3 every mixer application of the dataset item derivation becomes four applications with distinct round keys, the eight dependent cache reads per item stay eight, and the 256 MiB cache doubles on the days the dataset doubles (years 4 and 12). Version 2 is byte-identical: the two pinned packs re-export without a changed byte (section 5).

1. Why this form

The recompute attacker of docs/analysis/m16-recompute-attacker-2026-10-05.md holds the 256 MiB cache on a die and derives every dataset word instead of reading it: 128 items per hash at about 1,170 integer operations and 8 dependent cache reads each. Its cost is linear in operations per item; the honest miner pays the mixer once a day in the dataset build and never per hash; the verifier pays it per item it checks. The multiplier m is the one parameter that moves the attacker and leaves the honest hash rate untouched.

Two shapes give the attacker 4x the operations:

Shape Mixer applications per item Dependent cache reads per item What grows for the verifier What grows for the chip What grows for the honest build
A, chosen: m = 4 applications per round, 8 rounds 36 8 the ALU part only; the latency part (8 dependent misses per item) unchanged integer operations 4x; SRAM bandwidth unchanged (1,024 reads per hash) 4x the mixer arithmetic, same reads
B, alternative: 32 rounds of one application and one read 33 32 both parts: 4x the dependent misses per item, so about 4x the latency-bound time (spec 1.11: about 8 x 100 ns per item in series without interleaving) integer operations 3.7x and SRAM bandwidth 4x (4,096 reads per hash, 60 TB/s to match one 5090 at the M16 rate) 4x the reads too; the GPU build becomes latency-bound at 4x the dependent line fetches

Shape A is chosen because the verifier's latency part is the part the 10 ms gate protects (section 1.11: the distinct items of a unit are derived with their chains interleaved so the 8 misses of each item overlap across up to 32 items; shape B would make that 32 misses deep). Shape B is implemented nowhere; the rule in the brief: if shape A's measured verifier time exceeds 4.8 ms per warp on one M5 Max core, measure both and recommend. Section 6 has the measurement; it is under that bound, so B stays unimplemented.

2. Spec text (replaces 1.8.5 and 1.13.3 under class v3; v2 text unchanged)

1.8.5 Item derivation and dataset mapping

Item t (16 words) under mixer multiplier m (m = 1 for program class v2, m = 4 for class v3; a class parameter, LoadClass::mixer_mult):

s[i]     = K[i]                       for i in 0..7
s[8 + i] = t * MUL[i] + RC[i]         for i in 0..7
for r in 0..7:
    for j in 0..m-1:
        s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
    a = s[0] AND (2^(C - 4) - 1)      cache line index, 2^(C - 4) lines of a 2^C-word cache
    s[i] = s[i] XOR cache[line a][i]  for i in 0..15
for j in 0..m-1:
    s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
item(t) = s

M(s, rk) is the mixer of 1.8.4 with round key rk; under m = 1 the keys are (r + 1) * 0x9E3779B9 and 9 * 0x9E3779B9, the version 2 text exactly. The round keys of the 9 m applications are the first 9 m values of the version 2 key sequence, all distinct (the sequence is k * 0x9E3779B9 for k = 1 .. 9 m, and 0x9E3779B9 is odd, so no two of the first 2^32 keys coincide). Eight dependent cache reads per item at every m (ITEM_ROUNDS = 8, prototype value): the address of read r depends on every earlier read. 9 m mixer applications, 36 under class v3, about 4,700 integer operations per item (130 per application, 1.8.4). dataset[w] = item(w >> 4)[w AND 15]. A dataset of 2^D words is the prefix of items 0 .. 2^(D-4) - 1, so an item has the same value at every dataset size; and a cache of 2^C words is the prefix of segments of every larger cache (1.8.3 fills segments independently of the cache size), but an item's value depends on C through the line mask, so the item changes on the day the cache doubles.

Source: igneum-pow/src/memhard.rs (derive_items, round_key_mult, Shape), the emitted mh_item of memhard.h, memhard.metal and kernel.cl (emit.rs, emit_memhard_core: the m loop is emitted only for m > 1, so every version 2 pack keeps its text).

1.13.3 Dataset growth (class v3: option (b) with the cache tied to it, "option C")

Designed: 2 GiB at genesis plus 0.5 GiB per year. The linear schedule in bytes, G x (1 + 86,400 d / (4 x 31,536,000)) = G x (1 + d / 1,460) for the genesis size G and the chain day d (DAA days since genesis, section 1.12), doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at day 10,220 (year 28). Rule (Designed, decided 5 October 2026 for class v3):

doublings(d)        = floor(log2(1 + d / 1460))        integer division, then integer log2
dataset_words(d)    = 2^(D_0 + doublings(d))            D_0 = 29 designed (2 GiB), 28 on the devnet (1 GiB); capped at 32
cache_words(d)      = 2^(26 + doublings(d))             256 MiB, 512 MiB from year 4, 1 GiB from year 12

Power-of-two sizes only (option (b)), so every load keeps the src AND MASK form of 1.14 item 2 and the cache line index keeps s[0] AND mask. The cache doubles exactly when the dataset doubles ("option C", docs/analysis/sram-mirror.md section 7): the cache's job is to stay above any GPU's last-level cache and that needs growth; the recompute attacker is priced by the mixer, not by the cache (section 7 below).

d is day_index(header.timestamp) - day_index(genesis.timestamp) with day_index = timestamp_ms / 86,400,000 (bind::day_index, the interim day rule), clamped at 0 (memhard::days_since_genesis). Under class v2 nothing grows: the cache is 2^26 words and the dataset the genesis size on every day.

Chain day d Years doublings Cache words Cache Dataset words (devnet D_0 = 28) Dataset (designed D_0 = 29) Verifier cache fill, one M5 Max core (measured at 256 MiB, section 6, scaled linearly)
0 to 1,459 0 to 4 0 2^26 256 MiB 2^28 (1 GiB) 2 GiB 0.18 s
1,460 to 4,379 4 to 12 1 2^27 512 MiB 2^29 (2 GiB) 4 GiB 0.36 s
4,380 to 10,219 12 to 28 2 2^28 1 GiB 2^30 (4 GiB) 8 GiB 0.72 s
10,220 to 21,899 28 to 60 3 2^29 2 GiB 2^31 (8 GiB) 16 GiB 1.4 s
21,900 and on 60 and on 4 2^30 4 GiB 2^32 (16 GiB, the index cap) 2^32 words, the cap 2.9 s

Test: memhard::tests::growth_schedule_table pins every row and the day before each step. The devnet pack igneum-devnet-v4-epoch0 is day 20,730 of the Unix count against genesis day 20,729, d = 1, so every existing size and vector stands.

Consequences for the tiers (the rule of 5 October): a verifier (any node, any pool core) holds 512 MiB from year 4 and 1 GiB from year 12, and fills it once a day in under a second on one 2026 core (the table); a miner's card holds the dataset, 4 GiB from year 4 and 8 GiB from year 12 on the designed schedule, so an 8 GB card mines until year 12 and a 16 GB card until year 28 (the cache is not in the card's working set at hash time: it is built, the dataset built from it, and dropped). Those dates are the design document's own schedule restated as steps; option (a) would have faded a 4 GiB card out in year 4 instead of year 4.

3. Interfaces

Item Where Note
LoadClass { mixer_mult: u8, growth: bool }, LoadClass::MX4 ("mx4"), with_mixer(m, growth), v2_loads(), takes_width_roll() generator.rs a class with v2 loads takes no width roll: its program stream is version 2's draw for draw, so the v3 program of a seed is the v2 program of that seed, only the dataset differs
Shape { mixer_mult, cache_log2_words }, Shape::for_class_day(class, d), MixParams.shape, Cache::fill_log2(key, log2), round_key_mult(r, j, m) memhard.rs Shape::V2 is version 2
growth_doublings(d), cache_log2_words(d), dataset_log2_words(D_0, d), days_since_genesis(day, genesis_day) memhard.rs the schedule, one function and its two sizes
DatasetSource::{new_shape, from_key_shape, shape}, Epoch::new_class_day, Epoch::from_seed_bytes_day(epoch, day, label, class, d, D_0) verify.rs the day-sized entries; the v2 entries are unchanged and build the v2 shape
IGNEUM_MIXER_MULT, IGNEUM_CLASS_MIXER_MULT, IGNEUM_CACHE_GROWTH in program.h; "mixer_mult", "cache_growth", the "item" string in program.json emit.rs written only for a class with m != 1 or growth, so v2 packs do not change
packfile.h mixerMult; packbench and the OpenCL host print the multiplier and the cache size the three hosts the kernels carry the construction in their text (one emitter, three dialects); the hosts size the cache from IGNEUM_CACHE_LOG2_WORDS already (packbench.swift line 56, host.cu line 65, host.c line 1966)
igneum-pow --class mx4 [--days d] on every command main.rs --days sizes the cache for a growth class

Under the ca2-v3 seam (ProgramClass::V3, V3_CLASS), the integration sets V3_CLASS = LoadClass::MX4; the chain's day-sized dataset needs the day index, so Epoch::chain_dataset(day, class) builds the genesis-size cache and a chain_dataset_day(day_bytes, class, d, D_0) beside it is the growth entry (section 9, owed to the node agent).

4. Vectors (class v3, proto-cuda/packs-ca2-mixer/)

Produced by igneum-pow export --program-class v3 (the Rust CPU interpreter, 5 October 2026, commit 66eeba3) and checked on the GPUs in section 6. The program of each pack is the version 2 program of the same seed instruction for instruction (tests/packs.rs, v3_packs_are_the_v2_seeds_under_mixer_x4); the cache is the version 2 cache (day 0 of the growth rule); the dataset words and the hashes are new. The 64 sampled indices are those of every pack (emit::sample_indices).

Pack mx4-genesis (seed igneum-genesis, day 2026-10-03, generator 3, class mx4, program id e323b9dcaf283a6f, 2^28 words, 2^26-word cache, cache FNV-1a 64 48c4f5bf24166b2e as under v2):

dataset words 0..15 (item 0):
  61ff2180 0d4c7e6c 2177d443 60df9025 cf8b2e10 63675bfb 25289e58 9c45dc42
  2d271c54 9652369b 2dd77508 5921392c 3afa60ee c640ad68 f2bb56ff cfa46438
dataset[0x0fffffff] = 5020180e
sampled words (the first 8 of the 64 in vectors.json):
  dataset[59471966] = de85726d   dataset[217795994] = 7cfc31c7   dataset[208353206] = d3cd5289   dataset[42483309] = 6ebbeef1
  dataset[172547758] = 9858413b   dataset[148076330] = e786f141   dataset[183853158] = 64f13833   dataset[214389424] = ca229d04
unit at base nonce 0, lanes 0..31:
  63acd2d273f475ba e929c78b34b80d4b 0b1011cb19982558 1457a0df5497aa11
  957d0f3bb71d98fb ac16901e6e6f6057 8ea1c6279f4b177a f28146e60bd08ba9
  fe5b8cfe87f8e65b 49f87240566ace62 6ef6d6b7bdea8e41 46d9c0dc29a97b9c
  111fe30128db9398 66dc39084f0946d4 8ee11bdfd35fecf2 2861fcfc75db6677
  31c7667d4bde8556 c5989c48858b4ce0 276395e734a9d30d 84217b41e91368ff
  3604861e34d9f697 9f51d8ee16bf3639 c89e47bafa84401c 7ae78c1f10b70e19
  0b8c947157a29a48 d67192e8cfb43842 05a4c6d182c8c675 188e2661f3263f2e
  a2df24238f7fea2e ed69ea7e13ad3a48 2d44ae509bab91b8 adad61931ea4fb70
unit at base nonce 4096: lane 0 edd508ac57e5699a, lane 1 4eefd56d526cdaeb, lane 31 8892f8604733b1e0
unit at base nonce 1000000: lane 0 8b3183778a49f59c, lane 1 1831b72a8797e895, lane 31 75eae55eba53a506

Pack mx4-devnet-epoch0 (seed the devnet genesis hash, day bytes of 2026-10-04, generator 3, class mx4, program id 73bcbfe8ccf988f1, 2^28 words, 2^26-word cache, cache FNV-1a 64 448274a57f508cbc as under v2):

dataset words 0..15 (item 0):
  afe80d67 b9fbd029 6c79f193 95139ad9 96310aff 4609f8b1 75279e63 28235be1
  47b17dcb 718e0ef2 a52588c8 a8bf49d5 19cf243e 5ec8905e a4851f66 af9cd9f3
dataset[0x0fffffff] = e6a99c7a
sampled words (the first 8 of the 64 in vectors.json):
  dataset[59471966] = 57642b58   dataset[217795994] = c279badd   dataset[208353206] = cbccbaad   dataset[42483309] = 32cce392
  dataset[172547758] = 71fdb4c6   dataset[148076330] = f5a268ce   dataset[183853158] = 3ca1d676   dataset[214389424] = 7977b03d
unit at base nonce 0, lanes 0..31:
  212c6442b51e87ae c374795c00839331 b6036a220a98f4b3 eb8b8013e637367b
  db5866e9b73930fd f3f3d01f46e90333 9d913991ab8ed428 7ccb1d8fa100a800
  3cf45ba44f09a0fe 91acf48ef1a63082 6ea46c69fb082f99 581f0218977a9d72
  9a4623a5c62ddf2d ab6eb5e768f0feb4 07b70bdccf8aca12 d666311ae5e4311e
  53114757d669f0a4 bd5d6ace87ce2ce4 fb712015e8189192 a32cec81103e134b
  83f3d18c3289c124 fe29f1984b132b3d c9ffcf4e3774497a ac99c9243dc63809
  d78a7e8217a32f3c ab81ad63d242fc31 0e4c30b7e00024af ce014289fff6778d
  64a1292e2a8b4a91 d5b8c90e681d7e3a 06078117673030fd 51bf77b280173930
unit at base nonce 4096: lane 0 3c797978566b5950, lane 1 7c759e60185b6411, lane 31 96a903eb9a0ca390
unit at base nonce 1000000: lane 0 d5a8da0568df8ee7, lane 1 68cb69c04208285a, lane 31 f8ca84a1a5d78cf5

5. The v2 path is byte-identical

cargo test --test packs regenerates every file of igneum-genesis-mh and igneum-devnet-v4-epoch0 from program.json and compares byte for byte (emitted_sources_match_all_packs, export_pack_matches_all_packs); section 6 also records a fresh igneum-pow export of both packs diffed against the checked-in directories.

6. Measurements

Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0, 5 October 2026 (night), other agents' builds running beside every run; a timing row says which lock it ran under (measure is exclusive; run and build are not timings).

6.1 The v2 path, byte for byte (no lock needed)

igneum-pow export --seed igneum-genesis --day 2026-10-03 and igneum-pow export --epoch-hex edc4fa84...fb07 --day-hex 69676e65756d2d6461792ffa50000000000000 on commit 66eeba3, diff -r against proto-cuda/packs/igneum-genesis-mh and igneum-devnet-v4-epoch0: IDENTICAL, both (twelve files each). The crate tests regenerate the same files and compare them on every run (tests/packs.rs, 12 of 12 pass).

6.2 Bit-exactness of the class v3 construction on the GPUs (with-lock.sh run, 22:05 UTC)

packbench --pack <dir> --batches 1 --batch-log2 24 --group 256 (Metal, built from this branch) and igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 1 --batch-log2 24 (Apple OpenCL, built from this branch's host.c). Vectors are the Rust interpreter's; the fingerprint is FNV-1a 64 over the 2^24 outputs at base nonce 0.

Pack Harness Cache FNV-1a 64 Dataset head, word MASK, 64 samples Vectors standalone / in batch Fingerprint 2^24 MH/s (GPU time; not a measurement, the run lock)
mx4-genesis Metal 48c4f5bf24166b2e PASS PASS 3/3, 3/3 6f48d5a2aa0dbe5f 27.6
mx4-genesis Apple OpenCL PASS PASS 96 of 96 lanes 6f48d5a2aa0dbe5f 27.7 (wall)
mx4-devnet-epoch0 Metal 448274a57f508cbc PASS PASS 3/3, 3/3 73caaebb28e808fe 27.5
mx4-devnet-epoch0 Apple OpenCL PASS PASS 96 of 96 lanes 73caaebb28e808fe 27.6 (wall)
mx8-genesis (the x8 candidate, 21:45 UTC) Metal 48c4f5bf24166b2e PASS PASS 3/3, 3/3 7c28cfb06c5c65a9 27.7
mx8-genesis Apple OpenCL PASS PASS 96 of 96 lanes 7c28cfb06c5c65a9 27.6 (wall)
mx8-devnet-epoch0 Metal 448274a57f508cbc PASS PASS 3/3, 3/3 bbb183f72692f840 27.7
mx8-devnet-epoch0 Apple OpenCL PASS PASS 96 of 96 lanes bbb183f72692f840 27.7 (wall)

Reading: the Rust interpreter, Metal and Apple OpenCL agree on the v3 dataset (head, MASK word, 64 samples), on every vector lane and on the 2^24-output fingerprint of each pack; the hash rate is the v2 rate (27.7 MH/s on this card tonight, readwidth table), as it must be: the hash kernel only loads, the mixer is paid in the build.

Dataset build under the run lock (indicative only; the measured rows are in 6.4): Metal 30.2 ms GPU (mx4-genesis) and 21.7 ms GPU (mx4-devnet-epoch0) for 1 GiB; Apple OpenCL 55 and 56 ms wall. The readwidth entry's v2 figure on this card is 0.6 to 2.1 ms cache fill and a 1 GiB build the bench-log's memory-hard entry puts at 13 to 30 ms; the x4 build on the GPU is the row the measure lock will settle.

6.2a The packs as pinned after the x8 decision (with-lock.sh run, 22:12 UTC, the fixed binary's exports)

Pack How exported Program id Unit 0 lane 0 Fingerprint 2^24 (Metal = Apple OpenCL) Vectors, self-tests
mx8-genesis --program-class v3 (generator 3, V3_CLASS = mx8) e323b9dcaf283a6f 19b56348bc85304d 7c28cfb06c5c65a9 3/3 + 3/3, 96 of 96, PASS
mx8-devnet-epoch0 the chain path, --program-class v3 --era-hex <genesis> (the era drawn inside the class, load class mx8-erad810f22d) 73bcbfe8ccf988f1 d424577fce4a7a60 90f794dd556f7a3b 3/3 + 3/3, 96 of 96, PASS
mx4-genesis (the x4 record) --class mx4 (generator 2, the class in the id) 951b89584750bd75 63acd2d273f475ba 6f48d5a2aa0dbe5f 3/3 + 3/3, 96 of 96, PASS
mx4-devnet-epoch0 (the x4 record) --class mx4 on the chain seeds (program.json) 212c6442b51e87ae 73caaebb28e808fe 3/3 + 3/3, 96 of 96, PASS

The PC 1 rows of 6.5 ran the earlier mx8-devnet-epoch0 (the load-class export, fingerprint bbb183f72692f840, era not inside); the composed pack's PC fingerprints are owed with the next PC round (section 9).

6.3 Soundness suite on the v3 construction (commit 66eeba3 and after; with-lock.sh build for cargo, run for the GPU)

Suite Command Result
Crate lib tests (memhard schedule table, the by-hand multiplied mixer, the seam, the generator) cargo test -j4 --release 44 of 44 pass
Pinned packs, v2 and v3 (tests/packs.rs: programs, ids, dataset words, 96 vectors per pack, every emitted file byte for byte, 16 masked loads per kernel, the v3 packs as the v2 seeds under mixer x4) same 12 of 12 pass
Scratch soundness tests of branch ca2-soundness (cherry-pick 0d8f745, one conflict in packbench.swift's RESULT line resolved by hand, era_bytes: None added to the edge program literal) cargo test -j4 --release --test scratch 7 of 7 pass: bijections, re-hit rates, 56 of 56 edge units, 42 of 42 emitted scr kernels, 200 scratch programs and 800 units on the CPU
v3 fuzz on the CPU (tests/mixer.rs): 200 programs through the seam, the contract on every instruction (each is the v2 program of its seed), 4 units each across the 32-bit range with one unit in the top 256 nonces, interpreted twice IGNEUM_MIXER_PACKS_OUT=<dir> cargo test -j4 --release --test mixer 200 of 200, 800 of 800 units; 200 packs written for the GPU runs (82 s with the three cache fills)
v3 stats beside v2 (tests/mixer.rs): 8,192 outputs per seed, bit balance, single-bit avalanche within the unit and across units, duplicates same igneum-genesis v3: avalanche 49.99 percent, worst bit z 1.92, 0 duplicates (v2: 49.87, z 2.25); igneum-genesis/stats1 v3: 49.97, z 3.09 (v2: 49.98, z 2.30)
v3 edge (tests/mixer.rs): items 0, 1, 2^28 - 1 and 2^32 - 1 by hand at m = 1, 2, 4, 8 on a 2^14-word cache; words 0, 15, 16, 17, MASK - 1, MASK through the interpreter's fetch path; the index wrap at MASK + 1 same pass
v3 determinism (tests/mixer.rs): two independent epochs, every vector and every emitted file equal, and equal to the pinned pack same pass
Metal and Apple OpenCL on the two pinned v3 packs section 6.2 3/3 standalone, 3/3 in batch, 96 of 96 lanes, dataset head, MASK word and 64 samples, one fingerprint per pack across both harnesses
Metal fuzz: the 200 packs, 4 units each standalone and the top-256 unit inside a 512-nonce batch at base 4,294,967,040; every tenth pack on Apple OpenCL as well packbench --pack <dir> --batches 1 --batch-log2 9 --batch-base 4294967040 (Metal), igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 1 --batch-log2 10 (Apple OpenCL), with-lock.sh run, 21:22 to 21:24 UTC Metal 200 of 200 packs PASS (800 of 800 standalone units, 200 of 200 inside the wrapping window, cache and dataset self-tests on every pack); Apple OpenCL 20 of 20 packs PASS (the three vectors.h units, the self-tests); a first run with a packbench built before the --batch-base cherry-pick reported 200 of 200 FAIL on an empty RESULT line and was read as such (the watcher rule), the harness rebuilt and the run repeated

6.4 Timings (with-lock.sh measure, one session, 21:40:12 to 21:40:23 UTC, commit 504cae4)

Script measure-v3.sh (session scratchpad): igneum-pow bench --seed igneum-genesis --day 2026-10-03 --warps 50 (v2), ... --program-class v3 (x4), ... --class mx8 (x8), two rounds each, then the devnet seeds, then packbench --pack <dir> --batches 2 --batch-log2 22 --group 256 on igneum-genesis-mh, mx4-genesis and mx8-genesis, two rounds. The lock was exclusive among the agents' builds and measurements, but the box was not quiet: load average 5.6 (one minute) and 26 (fifteen minutes) at the start, from unlocked processes (the devnet node, other agents' editors); the v2 row reads 1.31 to 1.36 ms where the quiet readwidth night read 0.604 to 0.626. So the absolute numbers below are a loaded-core figure, about 2.2x the quiet one, and the ratios between the rows are the measurement (two rounds within 4 percent). A quiet-box re-run is owed (section 9).

Construction Verifier, ms per 32-lane unit, avg of 50 (round 1 / round 2) Worst cold unit of three Against v2 256 MiB cache fill, one core Metal 1 GiB dataset build, GPU ms (round 1 / round 2)
v2 (igneum-genesis) 1.361 / 1.310 1.579 1 172.1 / 172.6 ms 29.7 / 21.0
x4 (mx4, class v3) 1.956 / 1.923 2.043 1.45x 175.3 / 172.3 ms 20.9 / 21.0
x8 (mx8) 2.785 / 2.790 2.942 2.09x 172.3 / 172.3 ms 21.9 / 21.9
x4, the devnet seeds (mx4-devnet-epoch0) 1.923 2.012 173.9 ms
x8, the devnet seeds 2.972 2.885 173.6 ms

Reading. The verifier's ALU part is what grows: x4 adds 0.6 ms per unit for 27 more mixer applications on each of 4,096 items (110,592 applications, about 5.5 ns each on this core, the lanes' chains interleaved), x8 another 0.85 ms for 36 more; the latency part (8 dependent misses per item) is the same in every row, which is why the measured ratios are 1.45x and 2.1x and not the 4x and 8x of the M16 table's scaling. Shape B (32 rounds of one read, section 1) would have multiplied the latency part too; x4 is under the 4.8 ms bar even on the loaded core, so B stays unimplemented. The cache fill does not depend on the mixer (it is the ChaCha chain): 172 to 175 ms, the spec's 175 to 181 ms of 1.8.3. The Metal 1 GiB build does not move with the mixer at all (21 ms at v2, x4 and x8 once warm; the 29.7 ms first v2 run is the first-touch cost the hosts fill twice for): on this card the build is bound by the 8 dependent cache-line reads per item, not by the arithmetic, so the Mac says nothing about whether the 5090's or the 9070 XT's build is arithmetic-bound; that is the PC job (section 8).

Against the x4 / x8 rule (section 6.5): the verifier half passes for x8 with 7.1 ms of the 10 ms gate to spare on this loaded core (worst cold 2.94 ms; the quiet-core figure would be about 1.3 ms, scaled by the 2.2x of the v2 row, approximate); x4 leaves 8.0 ms. The build half waits on the PC rows.

6.5 Verification throughput per tier (consequences row C19), from the loaded-core figures above

Warps verified per second on one core = 1,000 / (ms per warp); a pool core verifying members' shares handles that many shares per second; a node verifies a block with one unit (plus the header path, under 0.1 ms, not measured here); IBD over the 108,000-header pruning window (spec 02) on one core = 108,000 x ms per warp.

Figure v2 x4 x8 Note
ms per warp, steady (this session, loaded core) 1.33 1.94 2.79 avg of the two rounds
ms per warp, quiet M5 Max core (scaled by 0.604 / 1.33 = 0.45, approximate) 0.60 0.88 1.26 the readwidth night's v2 figure is measured; x4 and x8 scaled
ms per warp, 2019-class laptop core (approximate: 2.5x the quiet M5 Max figure, the ratio the design document assumes for the gate; unmeasured, O-1.14) 1.5 2.2 3.2 the figure that fixes the gate is a measurement, not this row
Shares per second per core (loaded / quiet, approximate) 750 / 1,660 515 / 1,140 358 / 790 spec 09 section 9.8 item 5 carried 2,270 at v2; re-cut from the quiet row: 1,660
Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second), loaded / quiet 2.9 / 1.3 4.3 / 1.9 6.1 / 2.8
Node: worst cold single unit (loaded core) 1.58 ms 2.04 ms 2.94 ms per block
IBD over 108,000 headers on one core, loaded / quiet, minutes 2.4 / 1.1 3.5 / 1.6 5.0 / 2.3 laptop (approximate): 2.7 / 4.0 / 5.8 min; a seed VM core (unmeasured) sits between the laptop and the quiet M5 Max
Margin left under the 10 ms gate for Counter ASIC 3.0 (worst cold, loaded core) 8.4 ms 8.0 ms 7.1 ms on the 2019-class laptop row (approximate) 7.5 / 6.8 / 5.9 ms steady

Reading: at x4 a pool core serves about 1,100 shares per second on a quiet 2026 core (a 22,000-member pool needs two cores); at x8 about 800 (three cores). A node's block verification stays a few milliseconds. The gate's remaining margin is what Counter ASIC 3.0 has to spend, and on the unmeasured laptop core it is 6 to 7 ms at x4 and about 6 at x8, which is the number the 2019-class measurement (O-1.14) must confirm before x8 is final.

6.4a The same session on the fixed binary (section 6.6), with-lock.sh measure, 22:06:59 to 22:07:06 UTC

The verifier of section 6.4 was measured on a binary that carried the inlining regression of section 6.6; after the fix, readwidth's binary and the fixed one on the same v2 input in the same minute (checksum 19297e99c7b9a55e), then x4 and x8 on the fixed binary; load average 5.5 (the same box, so the ratios of 6.4 stand and the absolute row is now the measured one):

Construction Verifier, ms per unit, avg of 50 (round 1 / round 2) Worst cold unit of three Against v2
v2, readwidth e752fc7's binary 0.607 / 0.610 0.666 1
v2, the fixed binary 0.609 / 0.611 0.657 1.00
x4 (mx4) 1.238 / 1.237 1.396 2.03x
x8 (mx8, class v3) 2.077 / 2.058 2.145 3.4x
x4, the devnet seeds 1.240 1.289
x8, the devnet seeds 2.058 2.144

The added cost per unit is the same as on the slow binary (x4 + 0.63 ms, x8 + 1.46 ms: the regression was a constant 0.72 ms per unit in the shared item loop), so the ALU reading of 6.4 holds; the ratios against v2 are 2.0x and 3.4x once v2 is back at 0.61. The 10 ms gate keeps 7.9 ms at x8 (worst cold 2.15 ms) on this core.

6.5 The daily build per tier, and the x4 / x8 rule

The coordinator's rule (21:30 UTC): x8 enters v3 if the per-warp verify stays under 10 ms on one Mac core AND the daily 1 GiB build stays under 1 s on every discrete card we own; else x4 with the thin margin stated and x8 named as the next lever. The integrated tier is decided beside it (consequences row C23): its build is per prepare, not per day, so its consequence is per-day dataset reuse in the workers or a restart per epoch.

Card Build at x1 x4 x8 Source
RTX 5090 (PC 1, job run-mixer-x4-pc1-20261005, 22:00 to 22:04 UTC, the worker's cache ... dataset ... ms wall line, two dispatches per pack) 23 to 25 ms 23 to 25 ms 23 ms the PC 1 job (13.4 ms GPU time on 3 October: the wall line carries the launch)
RX 9070 XT (PC 1, gfx1201 on the eGPU, the same job) 74 ms 73 to 77 ms 72 to 76 ms the PC 1 job
M5 Max, Metal 13 to 30 ms (the two runs of the memory-hard entry; tonight's run-lock figures 21.7 to 30.2 ms at x4 and 22.0 to 30.0 at x8 say the Mac's build is latency-bound, not mixer-bound) section 6.4 section 6.4 this file
Radeon integrated gfx1036 (PC 2), OpenCL, per prepare 6.9 / 9.4 / 11.7 s prepare total with the 1 GiB build inside about 28 to 47 s (approximate: scaled x4; the iGPU's build is arithmetic-bound at x1 already) about 55 to 94 s (approximate) docs/plans/epoch-length.md section 7 (branch ca2-epoch), M11 table
gfx1036 beside WSL build jobs (PC 1) 55 / 116 / 124 s about 4 to 8 min (approximate) about 7 to 17 min (approximate) same
8 GB-class discrete card (not owned; about a tenth of the 5090's rate, approximate) about 0.13 s about 0.5 s about 1 s, on the edge of the rule scaled from the 5090 row, approximate

Decision (coordinator under the delegated rule, 22:05 UTC, on these rows): x8 enters class v3. Both halves pass: the verifier at x8 is 2.1 ms per unit on one M5 Max core (6.4a) against the 10 ms gate, and the daily 1 GiB build does not move with the mixer on any discrete card we own (5090 23 to 25 ms, 9070 XT 72 to 77 ms, M5 Max 21 ms at x1, x4 and x8: latency-bound), 13x to 40x under the 1 s bar. V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }; the pinned v3 packs are mx8-genesis and mx8-devnet-epoch0 (the latter through the chain path with the era inside the class); the x4 packs stay pinned as the candidate's record (generator 2, the class in the id).

Reading of the integrated tier: on the discrete cards the rule is settled by the 5090 and 9070 XT rows above. The integrated tier misses the rule at x4 already: a per-prepare build of 28 to 47 s is a tenth to a quarter of the 600-DAA-second lead the devnet gives the next program (spec 1.12), and under load it is the whole lead; so if x4 or x8 goes in, the iGPU tier needs the workers to build the day's dataset once a day and keep it across epochs (today a prepare rebuilds it: proto-opencl/host.c prepareTask builds the pair's cache and dataset per prepare), or to restart per epoch. The node agent is asked whether the per-day reuse is bounded tonight (coordinator, 21:45 UTC); until then the iGPU consequence stands as written.

6.6 The verifier regression of 0fc0ad1, found and fixed (5 October 2026, 21:59 to 22:07 UTC)

The era agent measured the same v2 input with two binaries in one minute: readwidth's 0.604 ms per unit, ca2-v3 HEAD's 1.33. Bisected under the measure lock (one session, four binaries, two rounds, checksum 19297e99c7b9a55e): readwidth e752fc7 0.607 / 0.609; the ca2-v3 seam 6c75dad (before this branch) 0.610 / 0.609; this branch's 0fc0ad1 1.332 / 1.316; ca2-v3 88dafbc 1.325 / 1.347. So the 2.2x was in 0fc0ad1's derive_items, on the version 2 path the devnet verifies with, and section 6.4's "loaded box" reading was wrong: the load was real (the same session shows it) but the 2x was the code.

Variants, each a one-change copy measured against readwidth's binary in the same session:

Variant ms per unit Reading
0fc0ad1 as written (the line mask read from the cache at run time, the loop inlined into MemhardCpu::fetch) 1.33 the regression
the mask hoisted into a local before the item loop 1.32 to 1.37 not the reload
Cache::line with the constant mask, on the 0fc0ad1 tree 0.60 to 0.65 fixed there
the same constant mask on the merged ca2-v3 tree 1.33 to 1.46 not the mask either
one instance per cache size with the mask a constant, #[inline(always)] 1.32 to 1.39 not the mask
the same instances #[inline(never)] 0.604 / 0.617 / 0.618 / 0.624 the fix

So it is inlining: the item loop inlined into its callers (fetch, fetch_wide, word_at) runs at 2.2x the out-of-line loop, and which small change tips LLVM's decision depends on the rest of the tree (the constant mask tipped it on one tree and not on the other). The fix (memhard::derive_items): the loop is derive_items_mask, #[inline(never)], one instance per cache size the growth rule reaches (2^26 to 2^30 words) with the line mask a constant, and a run-time-mask instance for every other size (tests). Measured in 6.4a: 0.609 / 0.611 against readwidth's 0.607 / 0.610.

The class, not the instance: a verifier benchmark with a pinned bound in the crate's CI (the v2 unit at a known input against a stored ms-per-unit on a named core, failing on a 1.3x drift) would have caught this at the first commit; filed for the next cut (section 9). Until then the era agent's two-binary check (same input, same minute) is the rule for every change that touches the item loop.

7. The chip model

docs/analysis/chip-model-v3.md.

8. What is unverified

  1. The 5090's and the 9070 XT's dataset build at x4 and x8 are measured (6.5, the PC 1 job): both latency-bound, under 0.1 s. The gfx1036 is the tier that fails the per-prepare build (6.5), and its consequence (per-day dataset reuse in the workers) is with the node agent.
  2. The absolute verifier figures were taken on a loaded core (load average 5.6); the ratios are the measurement and the quiet-core figures are scaled. A 2019-class laptop core has not run any construction (O-1.14).
  3. The mixer has had no cryptanalysis (spec 1.8.4); m applications with distinct round keys is m times the work only if no shortcut composes them, which is the same open question as for one application.
  4. The x8 packs are generator 2 with the class in the id (--class mx8); if x8 is chosen, the pinned v3 packs are re-cut through the seam (V3_CLASS = MX8, generator 3) and the tests re-pinned, one commit.
  5. The 5090's rate for the v3 program is the v2 rate by construction (the hash kernel is unchanged, the Mac shows 27.7 MH/s at v2, x4 and x8); the chip row's denominator stays the readwidth table's 136.1 MH/s until a v3 pack runs on the card, which the PC job also gives.

9. Owed

Item Owner When
PC 1 run of the five packs done 22:04 UTC (run-mixer-x4-pc1-20261005; 6.5 and 6.2a)
A verifier benchmark with a pinned bound in the crate's CI (section 6.6, the class rule) ca2-mixer the next cut
GPU bit-exactness of the re-exported mx8-devnet-epoch0 (the composed class, the era inside) and the mx4 record packs on the PCs (the Mac rows are in 6.2a) ca2-mixer the next PC round
The x4 / x8 choice recorded from the rule, then the vectors re-cut once through the seam done 22:05 UTC: x8 (6.5), V3_CLASS = MX8, mx8 packs re-exported
Epoch::chain_dataset_day wired to the genesis day index in the node (days_since_genesis(day_index(header), day_index(genesis))) and pow_genesis_dataset_log2 in the override ca2-node the integration
The spec text of section 2 into docs/spec/01-lottery-hash.md 1.8.5 and 1.13.3 (with the v3 vectors into 1.17) the integration after the choice
The 2019-class laptop core measurement that fixes the gate (O-1.14) cryptographer gate 1