11 KiB
Mixer x4 and the cache growth rule: the class v3 dataset construction
5 October 2026 (night). Counter ASIC 2.0, layer 6 (option C) and ledger M16's lever, decided by the coordinator
under the project lead's delegation at 22:00 UTC (docs/plans/counter-asic-2-status.md, "22:00 decided"; the project lead confirms for
the public testnet genesis). Branch ca2-mixer. Worker: ca2-mixer (cryptographer's lane).
What this changes, in one line: under program class v3 every mixer application of the dataset item derivation becomes four applications with distinct round keys, the eight dependent cache reads per item stay eight, and the 256 MiB cache doubles on the days the dataset doubles (years 4 and 12). Version 2 is byte-identical: the two pinned packs re-export without a changed byte (section 5).
1. Why this form
The recompute attacker of docs/analysis/m16-recompute-attacker-2026-10-05.md holds the 256 MiB cache on a die
and derives every dataset word instead of reading it: 128 items per hash at about 1,170 integer operations and 8
dependent cache reads each. Its cost is linear in operations per item; the honest miner pays the mixer once a day
in the dataset build and never per hash; the verifier pays it per item it checks. The multiplier m is the one
parameter that moves the attacker and leaves the honest hash rate untouched.
Two shapes give the attacker 4x the operations:
| Shape | Mixer applications per item | Dependent cache reads per item | What grows for the verifier | What grows for the chip | What grows for the honest build |
|---|---|---|---|---|---|
A, chosen: m = 4 applications per round, 8 rounds |
36 | 8 | the ALU part only; the latency part (8 dependent misses per item) unchanged | integer operations 4x; SRAM bandwidth unchanged (1,024 reads per hash) | 4x the mixer arithmetic, same reads |
| B, alternative: 32 rounds of one application and one read | 33 | 32 | both parts: 4x the dependent misses per item, so about 4x the latency-bound time (spec 1.11: about 8 x 100 ns per item in series without interleaving) | integer operations 3.7x and SRAM bandwidth 4x (4,096 reads per hash, 60 TB/s to match one 5090 at the M16 rate) | 4x the reads too; the GPU build becomes latency-bound at 4x the dependent line fetches |
Shape A is chosen because the verifier's latency part is the part the 10 ms gate protects (section 1.11: the distinct items of a unit are derived with their chains interleaved so the 8 misses of each item overlap across up to 32 items; shape B would make that 32 misses deep). Shape B is implemented nowhere; the rule in the brief: if shape A's measured verifier time exceeds 4.8 ms per warp on one M5 Max core, measure both and recommend. Section 6 has the measurement; it is under that bound, so B stays unimplemented.
2. Spec text (replaces 1.8.5 and 1.13.3 under class v3; v2 text unchanged)
1.8.5 Item derivation and dataset mapping
Item t (16 words) under mixer multiplier m (m = 1 for program class v2, m = 4 for class v3; a class
parameter, LoadClass::mixer_mult):
s[i] = K[i] for i in 0..7
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
for r in 0..7:
for j in 0..m-1:
s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache
s[i] = s[i] XOR cache[line a][i] for i in 0..15
for j in 0..m-1:
s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
item(t) = s
M(s, rk) is the mixer of 1.8.4 with round key rk; under m = 1 the keys are (r + 1) * 0x9E3779B9 and
9 * 0x9E3779B9, the version 2 text exactly. The round keys of the 9 m applications are the first 9 m values
of the version 2 key sequence, all distinct (the sequence is k * 0x9E3779B9 for k = 1 .. 9 m, and
0x9E3779B9 is odd, so no two of the first 2^32 keys coincide). Eight dependent cache reads per item at every
m (ITEM_ROUNDS = 8, prototype value): the address of read r depends on every earlier read. 9 m mixer
applications, 36 under class v3, about 4,700 integer operations per item (130 per application, 1.8.4).
dataset[w] = item(w >> 4)[w AND 15]. A dataset of 2^D words is the prefix of items 0 .. 2^(D-4) - 1, so an
item has the same value at every dataset size; and a cache of 2^C words is the prefix of segments of every
larger cache (1.8.3 fills segments independently of the cache size), but an item's value depends on C through
the line mask, so the item changes on the day the cache doubles.
Source: igneum-pow/src/memhard.rs (derive_items, round_key_mult, Shape), the emitted mh_item of
memhard.h, memhard.metal and kernel.cl (emit.rs, emit_memhard_core: the m loop is emitted only for
m > 1, so every version 2 pack keeps its text).
1.13.3 Dataset growth (class v3: option (b) with the cache tied to it, "option C")
Designed: 2 GiB at genesis plus 0.5 GiB per year. The linear schedule in bytes, G x (1 + 86,400 d / (4 x 31,536,000)) = G x (1 + d / 1,460) for the genesis size G and the chain day d (DAA days since genesis,
section 1.12), doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at day 10,220
(year 28). Rule (Designed, decided 5 October 2026 for class v3):
doublings(d) = floor(log2(1 + d / 1460)) integer division, then integer log2
dataset_words(d) = 2^(D_0 + doublings(d)) D_0 = 29 designed (2 GiB), 28 on the devnet (1 GiB); capped at 32
cache_words(d) = 2^(26 + doublings(d)) 256 MiB, 512 MiB from year 4, 1 GiB from year 12
Power-of-two sizes only (option (b)), so every load keeps the src AND MASK form of 1.14 item 2 and the cache
line index keeps s[0] AND mask. The cache doubles exactly when the dataset doubles ("option C",
docs/analysis/sram-mirror.md section 7): the cache's job is to stay above any GPU's last-level cache and that
needs growth; the recompute attacker is priced by the mixer, not by the cache (section 7 below).
d is day_index(header.timestamp) - day_index(genesis.timestamp) with day_index = timestamp_ms / 86,400,000
(bind::day_index, the interim day rule), clamped at 0 (memhard::days_since_genesis). Under class v2 nothing
grows: the cache is 2^26 words and the dataset the genesis size on every day.
Chain day d |
Years | doublings |
Cache words | Cache | Dataset words (devnet D_0 = 28) |
Dataset (designed D_0 = 29) |
Verifier cache fill, one M5 Max core (measured at 256 MiB, section 6, scaled linearly) |
|---|---|---|---|---|---|---|---|
| 0 to 1,459 | 0 to 4 | 0 | 2^26 | 256 MiB | 2^28 (1 GiB) | 2 GiB | 0.18 s |
| 1,460 to 4,379 | 4 to 12 | 1 | 2^27 | 512 MiB | 2^29 (2 GiB) | 4 GiB | 0.36 s |
| 4,380 to 10,219 | 12 to 28 | 2 | 2^28 | 1 GiB | 2^30 (4 GiB) | 8 GiB | 0.72 s |
| 10,220 to 21,899 | 28 to 60 | 3 | 2^29 | 2 GiB | 2^31 (8 GiB) | 16 GiB | 1.4 s |
| 21,900 and on | 60 and on | 4 | 2^30 | 4 GiB | 2^32 (16 GiB, the index cap) | 2^32 words, the cap | 2.9 s |
Test: memhard::tests::growth_schedule_table pins every row and the day before each step. The devnet pack
igneum-devnet-v4-epoch0 is day 20,730 of the Unix count against genesis day 20,729, d = 1, so every existing
size and vector stands.
Consequences for the tiers (the rule of 5 October): a verifier (any node, any pool core) holds 512 MiB from year 4 and 1 GiB from year 12, and fills it once a day in under a second on one 2026 core (the table); a miner's card holds the dataset, 4 GiB from year 4 and 8 GiB from year 12 on the designed schedule, so an 8 GB card mines until year 12 and a 16 GB card until year 28 (the cache is not in the card's working set at hash time: it is built, the dataset built from it, and dropped). Those dates are the design document's own schedule restated as steps; option (a) would have faded a 4 GiB card out in year 4 instead of year 4.
3. Interfaces
| Item | Where | Note |
|---|---|---|
LoadClass { mixer_mult: u8, growth: bool }, LoadClass::MX4 ("mx4"), with_mixer(m, growth), v2_loads(), takes_width_roll() |
generator.rs |
a class with v2 loads takes no width roll: its program stream is version 2's draw for draw, so the v3 program of a seed is the v2 program of that seed, only the dataset differs |
Shape { mixer_mult, cache_log2_words }, Shape::for_class_day(class, d), MixParams.shape, Cache::fill_log2(key, log2), round_key_mult(r, j, m) |
memhard.rs |
Shape::V2 is version 2 |
growth_doublings(d), cache_log2_words(d), dataset_log2_words(D_0, d), days_since_genesis(day, genesis_day) |
memhard.rs |
the schedule, one function and its two sizes |
DatasetSource::{new_shape, from_key_shape, shape}, Epoch::new_class_day, Epoch::from_seed_bytes_day(epoch, day, label, class, d, D_0) |
verify.rs |
the day-sized entries; the v2 entries are unchanged and build the v2 shape |
IGNEUM_MIXER_MULT, IGNEUM_CLASS_MIXER_MULT, IGNEUM_CACHE_GROWTH in program.h; "mixer_mult", "cache_growth", the "item" string in program.json |
emit.rs |
written only for a class with m != 1 or growth, so v2 packs do not change |
packfile.h mixerMult; packbench and the OpenCL host print the multiplier and the cache size |
the three hosts | the kernels carry the construction in their text (one emitter, three dialects); the hosts size the cache from IGNEUM_CACHE_LOG2_WORDS already (packbench.swift line 56, host.cu line 65, host.c line 1966) |
igneum-pow --class mx4 [--days d] on every command |
main.rs |
--days sizes the cache for a growth class |
Under the ca2-v3 seam (ProgramClass::V3, V3_CLASS), the integration sets V3_CLASS = LoadClass::MX4; the
chain's day-sized dataset needs the day index, so Epoch::chain_dataset(day, class) builds the genesis-size
cache and a chain_dataset_day(day_bytes, class, d, D_0) beside it is the growth entry (section 9, owed to the
node agent).
4. Vectors (class v3, proto-cuda/packs-ca2-mixer/)
Filled in section 6 from the exported packs: mx4-genesis (seed igneum-genesis, day 2026-10-03, 2^28 words,
2^26-word cache) and mx4-devnet-epoch0 (the devnet genesis hash as the epoch seed, day bytes of 2026-10-04).
5. The v2 path is byte-identical
cargo test --test packs regenerates every file of igneum-genesis-mh and igneum-devnet-v4-epoch0 from
program.json and compares byte for byte (emitted_sources_match_all_packs, export_pack_matches_all_packs);
section 6 also records a fresh igneum-pow export of both packs diffed against the checked-in directories.
6. Measurements
Filled as they land. Every row names the machine, the date, the command and the lock mode.
7. The chip model
docs/analysis/chip-model-v3.md.
8. What is unverified
Filled at the end.
9. Owed
Filled at the end.