igneum/docs/plans/mixer-x4.md

11 KiB

Mixer x4 and the cache growth rule: the class v3 dataset construction

5 October 2026 (night). Counter ASIC 2.0, layer 6 (option C) and ledger M16's lever, decided by the coordinator under the project lead's delegation at 22:00 UTC (docs/plans/counter-asic-2-status.md, "22:00 decided"; the project lead confirms for the public testnet genesis). Branch ca2-mixer. Worker: ca2-mixer (cryptographer's lane).

What this changes, in one line: under program class v3 every mixer application of the dataset item derivation becomes four applications with distinct round keys, the eight dependent cache reads per item stay eight, and the 256 MiB cache doubles on the days the dataset doubles (years 4 and 12). Version 2 is byte-identical: the two pinned packs re-export without a changed byte (section 5).

1. Why this form

The recompute attacker of docs/analysis/m16-recompute-attacker-2026-10-05.md holds the 256 MiB cache on a die and derives every dataset word instead of reading it: 128 items per hash at about 1,170 integer operations and 8 dependent cache reads each. Its cost is linear in operations per item; the honest miner pays the mixer once a day in the dataset build and never per hash; the verifier pays it per item it checks. The multiplier m is the one parameter that moves the attacker and leaves the honest hash rate untouched.

Two shapes give the attacker 4x the operations:

Shape Mixer applications per item Dependent cache reads per item What grows for the verifier What grows for the chip What grows for the honest build
A, chosen: m = 4 applications per round, 8 rounds 36 8 the ALU part only; the latency part (8 dependent misses per item) unchanged integer operations 4x; SRAM bandwidth unchanged (1,024 reads per hash) 4x the mixer arithmetic, same reads
B, alternative: 32 rounds of one application and one read 33 32 both parts: 4x the dependent misses per item, so about 4x the latency-bound time (spec 1.11: about 8 x 100 ns per item in series without interleaving) integer operations 3.7x and SRAM bandwidth 4x (4,096 reads per hash, 60 TB/s to match one 5090 at the M16 rate) 4x the reads too; the GPU build becomes latency-bound at 4x the dependent line fetches

Shape A is chosen because the verifier's latency part is the part the 10 ms gate protects (section 1.11: the distinct items of a unit are derived with their chains interleaved so the 8 misses of each item overlap across up to 32 items; shape B would make that 32 misses deep). Shape B is implemented nowhere; the rule in the brief: if shape A's measured verifier time exceeds 4.8 ms per warp on one M5 Max core, measure both and recommend. Section 6 has the measurement; it is under that bound, so B stays unimplemented.

2. Spec text (replaces 1.8.5 and 1.13.3 under class v3; v2 text unchanged)

1.8.5 Item derivation and dataset mapping

Item t (16 words) under mixer multiplier m (m = 1 for program class v2, m = 4 for class v3; a class parameter, LoadClass::mixer_mult):

s[i]     = K[i]                       for i in 0..7
s[8 + i] = t * MUL[i] + RC[i]         for i in 0..7
for r in 0..7:
    for j in 0..m-1:
        s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
    a = s[0] AND (2^(C - 4) - 1)      cache line index, 2^(C - 4) lines of a 2^C-word cache
    s[i] = s[i] XOR cache[line a][i]  for i in 0..15
for j in 0..m-1:
    s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
item(t) = s

M(s, rk) is the mixer of 1.8.4 with round key rk; under m = 1 the keys are (r + 1) * 0x9E3779B9 and 9 * 0x9E3779B9, the version 2 text exactly. The round keys of the 9 m applications are the first 9 m values of the version 2 key sequence, all distinct (the sequence is k * 0x9E3779B9 for k = 1 .. 9 m, and 0x9E3779B9 is odd, so no two of the first 2^32 keys coincide). Eight dependent cache reads per item at every m (ITEM_ROUNDS = 8, prototype value): the address of read r depends on every earlier read. 9 m mixer applications, 36 under class v3, about 4,700 integer operations per item (130 per application, 1.8.4). dataset[w] = item(w >> 4)[w AND 15]. A dataset of 2^D words is the prefix of items 0 .. 2^(D-4) - 1, so an item has the same value at every dataset size; and a cache of 2^C words is the prefix of segments of every larger cache (1.8.3 fills segments independently of the cache size), but an item's value depends on C through the line mask, so the item changes on the day the cache doubles.

Source: igneum-pow/src/memhard.rs (derive_items, round_key_mult, Shape), the emitted mh_item of memhard.h, memhard.metal and kernel.cl (emit.rs, emit_memhard_core: the m loop is emitted only for m > 1, so every version 2 pack keeps its text).

1.13.3 Dataset growth (class v3: option (b) with the cache tied to it, "option C")

Designed: 2 GiB at genesis plus 0.5 GiB per year. The linear schedule in bytes, G x (1 + 86,400 d / (4 x 31,536,000)) = G x (1 + d / 1,460) for the genesis size G and the chain day d (DAA days since genesis, section 1.12), doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at day 10,220 (year 28). Rule (Designed, decided 5 October 2026 for class v3):

doublings(d)        = floor(log2(1 + d / 1460))        integer division, then integer log2
dataset_words(d)    = 2^(D_0 + doublings(d))            D_0 = 29 designed (2 GiB), 28 on the devnet (1 GiB); capped at 32
cache_words(d)      = 2^(26 + doublings(d))             256 MiB, 512 MiB from year 4, 1 GiB from year 12

Power-of-two sizes only (option (b)), so every load keeps the src AND MASK form of 1.14 item 2 and the cache line index keeps s[0] AND mask. The cache doubles exactly when the dataset doubles ("option C", docs/analysis/sram-mirror.md section 7): the cache's job is to stay above any GPU's last-level cache and that needs growth; the recompute attacker is priced by the mixer, not by the cache (section 7 below).

d is day_index(header.timestamp) - day_index(genesis.timestamp) with day_index = timestamp_ms / 86,400,000 (bind::day_index, the interim day rule), clamped at 0 (memhard::days_since_genesis). Under class v2 nothing grows: the cache is 2^26 words and the dataset the genesis size on every day.

Chain day d Years doublings Cache words Cache Dataset words (devnet D_0 = 28) Dataset (designed D_0 = 29) Verifier cache fill, one M5 Max core (measured at 256 MiB, section 6, scaled linearly)
0 to 1,459 0 to 4 0 2^26 256 MiB 2^28 (1 GiB) 2 GiB 0.18 s
1,460 to 4,379 4 to 12 1 2^27 512 MiB 2^29 (2 GiB) 4 GiB 0.36 s
4,380 to 10,219 12 to 28 2 2^28 1 GiB 2^30 (4 GiB) 8 GiB 0.72 s
10,220 to 21,899 28 to 60 3 2^29 2 GiB 2^31 (8 GiB) 16 GiB 1.4 s
21,900 and on 60 and on 4 2^30 4 GiB 2^32 (16 GiB, the index cap) 2^32 words, the cap 2.9 s

Test: memhard::tests::growth_schedule_table pins every row and the day before each step. The devnet pack igneum-devnet-v4-epoch0 is day 20,730 of the Unix count against genesis day 20,729, d = 1, so every existing size and vector stands.

Consequences for the tiers (the rule of 5 October): a verifier (any node, any pool core) holds 512 MiB from year 4 and 1 GiB from year 12, and fills it once a day in under a second on one 2026 core (the table); a miner's card holds the dataset, 4 GiB from year 4 and 8 GiB from year 12 on the designed schedule, so an 8 GB card mines until year 12 and a 16 GB card until year 28 (the cache is not in the card's working set at hash time: it is built, the dataset built from it, and dropped). Those dates are the design document's own schedule restated as steps; option (a) would have faded a 4 GiB card out in year 4 instead of year 4.

3. Interfaces

Item Where Note
LoadClass { mixer_mult: u8, growth: bool }, LoadClass::MX4 ("mx4"), with_mixer(m, growth), v2_loads(), takes_width_roll() generator.rs a class with v2 loads takes no width roll: its program stream is version 2's draw for draw, so the v3 program of a seed is the v2 program of that seed, only the dataset differs
Shape { mixer_mult, cache_log2_words }, Shape::for_class_day(class, d), MixParams.shape, Cache::fill_log2(key, log2), round_key_mult(r, j, m) memhard.rs Shape::V2 is version 2
growth_doublings(d), cache_log2_words(d), dataset_log2_words(D_0, d), days_since_genesis(day, genesis_day) memhard.rs the schedule, one function and its two sizes
DatasetSource::{new_shape, from_key_shape, shape}, Epoch::new_class_day, Epoch::from_seed_bytes_day(epoch, day, label, class, d, D_0) verify.rs the day-sized entries; the v2 entries are unchanged and build the v2 shape
IGNEUM_MIXER_MULT, IGNEUM_CLASS_MIXER_MULT, IGNEUM_CACHE_GROWTH in program.h; "mixer_mult", "cache_growth", the "item" string in program.json emit.rs written only for a class with m != 1 or growth, so v2 packs do not change
packfile.h mixerMult; packbench and the OpenCL host print the multiplier and the cache size the three hosts the kernels carry the construction in their text (one emitter, three dialects); the hosts size the cache from IGNEUM_CACHE_LOG2_WORDS already (packbench.swift line 56, host.cu line 65, host.c line 1966)
igneum-pow --class mx4 [--days d] on every command main.rs --days sizes the cache for a growth class

Under the ca2-v3 seam (ProgramClass::V3, V3_CLASS), the integration sets V3_CLASS = LoadClass::MX4; the chain's day-sized dataset needs the day index, so Epoch::chain_dataset(day, class) builds the genesis-size cache and a chain_dataset_day(day_bytes, class, d, D_0) beside it is the growth entry (section 9, owed to the node agent).

4. Vectors (class v3, proto-cuda/packs-ca2-mixer/)

Filled in section 6 from the exported packs: mx4-genesis (seed igneum-genesis, day 2026-10-03, 2^28 words, 2^26-word cache) and mx4-devnet-epoch0 (the devnet genesis hash as the epoch seed, day bytes of 2026-10-04).

5. The v2 path is byte-identical

cargo test --test packs regenerates every file of igneum-genesis-mh and igneum-devnet-v4-epoch0 from program.json and compares byte for byte (emitted_sources_match_all_packs, export_pack_matches_all_packs); section 6 also records a fresh igneum-pow export of both packs diffed against the checked-in directories.

6. Measurements

Filled as they land. Every row names the machine, the date, the command and the lock mode.

7. The chip model

docs/analysis/chip-model-v3.md.

8. What is unverified

Filled at the end.

9. Owed

Filled at the end.