mixer x4 and the cache growth rule (Counter ASIC 2.0, class v3 construction): LoadClass mixer_mult and growth, LoadClass::MX4 (v2 loads, no width roll), memhard::Shape in MixParams, m mixer applications per round with keys round_key(r m + j), Cache::fill_log2, the option C schedule (growth_doublings, cache_log2_words, dataset_log2_words, days_since_genesis) with its test table, day-sized Epoch entries, the three emitters (m loop only for m > 1, v2 text unchanged), program.h and program.json fields, packfile.h mixerMult, packbench and OpenCL host prints, --class mx4 and --days on the CLI; docs/plans/mixer-x4.md design and spec text, docs/analysis/chip-model-v3.md, the 5090 and 9070 XT dataset-build playbooks (measurements and vectors to follow)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 20:53:12 +00:00
parent 28635b165f
commit e08909f138
13 changed files with 925 additions and 101 deletions

View file

@ -0,0 +1,86 @@
# The on-die-cache recompute chip against the RTX 5090, class v2 and class v3, everything combined
5 October 2026 (night), Counter ASIC 2.0, worker ca2-mixer. The model is M16's
(`docs/analysis/m16-recompute-attacker-2026-10-05.md`): the strongest chip the plan has priced holds the whole
cache in SRAM and derives every dataset item instead of reading it, so its cost per hash is item derivations,
and its rate at a 50 T op/s integer budget (an RTX 5090's, approximate) is `50 T / (ops per hash)`. Nothing here
is a measurement of a chip; every GPU figure says where it was measured. "Approximate" marks a figure from memory.
## 1. Inputs
| Input | Value | Source |
|---|---|---|
| Items per hash | 128 (one item per load, 128 loads per hash, median 128.00 distinct) | spec 01 sections 1.4.2 and 1.8.5; the 20,000-program census |
| Integer operations per mixer application | about 130 | spec 01 section 1.8.4 |
| Mixer applications per item | 9 under v2; 36 under v3 (`m = 4`, `docs/plans/mixer-x4.md`) | `memhard::Shape::mixers_per_item` |
| Integer operations per item | 1,170 (v2); 4,680 (v3) | 9 x 130; 36 x 130 |
| Integer operations per hash | 149,760 (v2, "150,000"); 599,040 (v3, "600,000") | 128 x the above |
| Chip integer budget | 50 T op/s (approximate: 21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock) | M16 section 3 |
| Fixed-function factor | 3x (approximate, from memory: 2x to 5x is the usual credit for a pipeline with no scheduling or divergence) | M16 section 3 |
| RTX 5090, version 2 programs, measured | 136.1 MH/s (readwidth, tonight, `docs/plans/read-width.md`, pack w4 on PC 2); 139.7 MH/s (M11, 4 October, `docs/bench-log.md`) | this analysis uses tonight's 136.1 as the denominator and quotes both |
| RTX 5090 at w16 (16-byte loads), measured | 139.8 MH/s | readwidth table, tonight (the width stays 4 B: w16 closes nothing) |
| Cache mirror, 256 MiB, N5 headline density | 128 mm^2, $46 per good die (64 mm^2, $21 at the bit-cell lower bound) | `docs/analysis/sram-mirror.md` revision 2, sections 4 and 5 (`ca2-analysis` e6085c6) |
| Cache mirror plus a 96 MB hot table, N5 headline | 175 mm^2, $68 | same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate) |
| 512 MiB and 1 GiB mirrors, N5 headline | 255 mm^2 and 510 mm^2; $111 to $306 | same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node) |
| GPU-class die | 750 mm^2 (the equal-silicon comparison) | M16 section 3 |
## 2. The rows
Chip rate = 50 T op/s / ops per hash. "Bare" = chip rate / 136.1 MH/s. "With the factor" = bare x 3. "Equal
silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does not get, the M16 convention
("minus the area the SRAM takes"). SRAM in mm^2 and dollars at the N5 headline density.
| Row | Mixer | Ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | With the 3x factor | Equal silicon, SRAM deducted, with the factor |
|---|---|---|---|---|---|---|---|---|
| v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x |
| v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x |
| v3: mixer x4 | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x |
| v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads) | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.61x or below (owed: the 5090's added-form rate; the hot loads cost it something, the chip nothing) | 1.84x or below | 1.49x |
| v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.61x or below (owed) | 1.84x or below | 1.45x |
| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 576 MiB | 287 / $130 | 0.61x | 1.84x | 1.14x |
| v3 at year 12 (cache 1 GiB, dataset 8 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 1,088 MiB | 542 / $330 | 0.61x | 1.84x | 0.51x |
| v3 with the mixer at x8 instead (the next lever, not adopted) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x |
The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op
weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The
width rule (4-byte loads kept) changes nothing either: w16 would have moved the honest denominator by 2.7% and the
chip's cost not at all.
Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6
hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53.
Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of `sram-mirror.md`
scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. Year 4 and 12 rows: the mirror of
`sram-mirror.md` section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of
those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area,
not a forecast).
## 3. The margin, plainly
The combined headline row reads 1.84x with the 3x factor at an equal integer budget, 1.5x with the SRAM area
deducted. The claim is "under 2x", and the margin is thin:
- the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x;
- the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart);
- the 50 T op/s budget is approximate; a chip at 55 T op/s reads 2.0x;
- the hot table in the added form lowers the honest denominator by whatever the hot loads cost the GPU (owed from
the PC rows), which raises the chip's gain by the same share, 1.84x or more if the hot loads are free, higher if
not; the hot table's only cost to this chip is 16 to 32 mm^2 of die.
What keeps it under 2x is the mixer, and nothing else in Counter ASIC 2.0 moves this chip (the scratch at any share
gave 2.4x, `docs/analysis/scratch-soundness.md` section 3.4; the hot table taxes the DRAM-only chip, not this one;
the cache growth taxes it only in die area, which is cheap at year 0 and real at year 12). The next levers, in
order:
1. Mixer x8: 0.18x bare and 0.55x with the factor in the M16 table (0.31x and 0.92x against 136.1 here; the M16
table's denominator is 229 MH/s); the CPU verifier at 3.3 to 9.6 ms per warp scaled from the version 1 range,
at the edge of the 10 ms gate; the measured v3 row of `docs/plans/mixer-x4.md` section 6 is what to scale from
now, and whether a 2019-class laptop core (unmeasured, O-1.14) passes 10 ms is what decides it.
2. The hot table: adopted or not on the PC rows (`docs/plans/hot-table.md`); in the added form it costs the GPU
nothing it was not already paying in cache misses and the chip die area only, so it is the second lever for the
chip only through area and the first against a DRAM-only chip.
## 4. What this does not settle
The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point
under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had
no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM.

151
docs/plans/mixer-x4.md Normal file
View file

@ -0,0 +1,151 @@
# Mixer x4 and the cache growth rule: the class v3 dataset construction
5 October 2026 (night). Counter ASIC 2.0, layer 6 (option C) and ledger M16's lever, decided by the coordinator
under the project lead's delegation at 22:00 UTC (`docs/plans/counter-asic-2-status.md`, "22:00 decided"; the project lead confirms for
the public testnet genesis). Branch `ca2-mixer`. Worker: ca2-mixer (cryptographer's lane).
What this changes, in one line: under program class v3 every mixer application of the dataset item derivation
becomes four applications with distinct round keys, the eight dependent cache reads per item stay eight, and the
256 MiB cache doubles on the days the dataset doubles (years 4 and 12). Version 2 is byte-identical: the two
pinned packs re-export without a changed byte (section 5).
## 1. Why this form
The recompute attacker of `docs/analysis/m16-recompute-attacker-2026-10-05.md` holds the 256 MiB cache on a die
and derives every dataset word instead of reading it: 128 items per hash at about 1,170 integer operations and 8
dependent cache reads each. Its cost is linear in operations per item; the honest miner pays the mixer once a day
in the dataset build and never per hash; the verifier pays it per item it checks. The multiplier `m` is the one
parameter that moves the attacker and leaves the honest hash rate untouched.
Two shapes give the attacker 4x the operations:
| Shape | Mixer applications per item | Dependent cache reads per item | What grows for the verifier | What grows for the chip | What grows for the honest build |
|---|---|---|---|---|---|
| A, chosen: `m = 4` applications per round, 8 rounds | 36 | 8 | the ALU part only; the latency part (8 dependent misses per item) unchanged | integer operations 4x; SRAM bandwidth unchanged (1,024 reads per hash) | 4x the mixer arithmetic, same reads |
| B, alternative: 32 rounds of one application and one read | 33 | 32 | both parts: 4x the dependent misses per item, so about 4x the latency-bound time (spec 1.11: about 8 x 100 ns per item in series without interleaving) | integer operations 3.7x and SRAM bandwidth 4x (4,096 reads per hash, 60 TB/s to match one 5090 at the M16 rate) | 4x the reads too; the GPU build becomes latency-bound at 4x the dependent line fetches |
Shape A is chosen because the verifier's latency part is the part the 10 ms gate protects (section 1.11: the
distinct items of a unit are derived with their chains interleaved so the 8 misses of each item overlap across up
to 32 items; shape B would make that 32 misses deep). Shape B is implemented nowhere; the rule in the brief: if
shape A's measured verifier time exceeds 4.8 ms per warp on one M5 Max core, measure both and recommend. Section 6
has the measurement; it is under that bound, so B stays unimplemented.
## 2. Spec text (replaces 1.8.5 and 1.13.3 under class v3; v2 text unchanged)
### 1.8.5 Item derivation and dataset mapping
Item `t` (16 words) under mixer multiplier `m` (`m = 1` for program class v2, `m = 4` for class v3; a class
parameter, `LoadClass::mixer_mult`):
```
s[i] = K[i] for i in 0..7
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
for r in 0..7:
for j in 0..m-1:
s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache
s[i] = s[i] XOR cache[line a][i] for i in 0..15
for j in 0..m-1:
s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
item(t) = s
```
`M(s, rk)` is the mixer of 1.8.4 with round key `rk`; under `m = 1` the keys are `(r + 1) * 0x9E3779B9` and
`9 * 0x9E3779B9`, the version 2 text exactly. The round keys of the `9 m` applications are the first `9 m` values
of the version 2 key sequence, all distinct (the sequence is `k * 0x9E3779B9` for `k = 1 .. 9 m`, and
`0x9E3779B9` is odd, so no two of the first 2^32 keys coincide). Eight dependent cache reads per item at every
`m` (`ITEM_ROUNDS = 8`, prototype value): the address of read `r` depends on every earlier read. `9 m` mixer
applications, 36 under class v3, about 4,700 integer operations per item (130 per application, 1.8.4).
`dataset[w] = item(w >> 4)[w AND 15]`. A dataset of 2^D words is the prefix of items `0 .. 2^(D-4) - 1`, so an
item has the same value at every dataset size; and a cache of 2^C words is the prefix of segments of every
larger cache (1.8.3 fills segments independently of the cache size), but an item's value depends on `C` through
the line mask, so the item changes on the day the cache doubles.
Source: `igneum-pow/src/memhard.rs` (`derive_items`, `round_key_mult`, `Shape`), the emitted `mh_item` of
`memhard.h`, `memhard.metal` and `kernel.cl` (`emit.rs`, `emit_memhard_core`: the `m` loop is emitted only for
`m > 1`, so every version 2 pack keeps its text).
### 1.13.3 Dataset growth (class v3: option (b) with the cache tied to it, "option C")
Designed: 2 GiB at genesis plus 0.5 GiB per year. The linear schedule in bytes, `G x (1 + 86,400 d /
(4 x 31,536,000)) = G x (1 + d / 1,460)` for the genesis size `G` and the chain day `d` (DAA days since genesis,
section 1.12), doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at day 10,220
(year 28). Rule (Designed, decided 5 October 2026 for class v3):
```
doublings(d) = floor(log2(1 + d / 1460)) integer division, then integer log2
dataset_words(d) = 2^(D_0 + doublings(d)) D_0 = 29 designed (2 GiB), 28 on the devnet (1 GiB); capped at 32
cache_words(d) = 2^(26 + doublings(d)) 256 MiB, 512 MiB from year 4, 1 GiB from year 12
```
Power-of-two sizes only (option (b)), so every load keeps the `src AND MASK` form of 1.14 item 2 and the cache
line index keeps `s[0] AND mask`. The cache doubles exactly when the dataset doubles ("option C",
`docs/analysis/sram-mirror.md` section 7): the cache's job is to stay above any GPU's last-level cache and that
needs growth; the recompute attacker is priced by the mixer, not by the cache (section 7 below).
`d` is `day_index(header.timestamp) - day_index(genesis.timestamp)` with `day_index = timestamp_ms / 86,400,000`
(`bind::day_index`, the interim day rule), clamped at 0 (`memhard::days_since_genesis`). Under class v2 nothing
grows: the cache is 2^26 words and the dataset the genesis size on every day.
| Chain day `d` | Years | `doublings` | Cache words | Cache | Dataset words (devnet `D_0 = 28`) | Dataset (designed `D_0 = 29`) | Verifier cache fill, one M5 Max core (measured at 256 MiB, section 6, scaled linearly) |
|---|---|---|---|---|---|---|---|
| 0 to 1,459 | 0 to 4 | 0 | 2^26 | 256 MiB | 2^28 (1 GiB) | 2 GiB | 0.18 s |
| 1,460 to 4,379 | 4 to 12 | 1 | 2^27 | 512 MiB | 2^29 (2 GiB) | 4 GiB | 0.36 s |
| 4,380 to 10,219 | 12 to 28 | 2 | 2^28 | 1 GiB | 2^30 (4 GiB) | 8 GiB | 0.72 s |
| 10,220 to 21,899 | 28 to 60 | 3 | 2^29 | 2 GiB | 2^31 (8 GiB) | 16 GiB | 1.4 s |
| 21,900 and on | 60 and on | 4 | 2^30 | 4 GiB | 2^32 (16 GiB, the index cap) | 2^32 words, the cap | 2.9 s |
Test: `memhard::tests::growth_schedule_table` pins every row and the day before each step. The devnet pack
`igneum-devnet-v4-epoch0` is day 20,730 of the Unix count against genesis day 20,729, `d = 1`, so every existing
size and vector stands.
Consequences for the tiers (the rule of 5 October): a verifier (any node, any pool core) holds 512 MiB from year 4
and 1 GiB from year 12, and fills it once a day in under a second on one 2026 core (the table); a miner's card
holds the dataset, 4 GiB from year 4 and 8 GiB from year 12 on the designed schedule, so an 8 GB card mines until
year 12 and a 16 GB card until year 28 (the cache is not in the card's working set at hash time: it is built,
the dataset built from it, and dropped). Those dates are the design document's own schedule restated as steps;
option (a) would have faded a 4 GiB card out in year 4 instead of year 4.
## 3. Interfaces
| Item | Where | Note |
|---|---|---|
| `LoadClass { mixer_mult: u8, growth: bool }`, `LoadClass::MX4` ("mx4"), `with_mixer(m, growth)`, `v2_loads()`, `takes_width_roll()` | `generator.rs` | a class with v2 loads takes no width roll: its program stream is version 2's draw for draw, so the v3 program of a seed is the v2 program of that seed, only the dataset differs |
| `Shape { mixer_mult, cache_log2_words }`, `Shape::for_class_day(class, d)`, `MixParams.shape`, `Cache::fill_log2(key, log2)`, `round_key_mult(r, j, m)` | `memhard.rs` | `Shape::V2` is version 2 |
| `growth_doublings(d)`, `cache_log2_words(d)`, `dataset_log2_words(D_0, d)`, `days_since_genesis(day, genesis_day)` | `memhard.rs` | the schedule, one function and its two sizes |
| `DatasetSource::{new_shape, from_key_shape, shape}`, `Epoch::new_class_day`, `Epoch::from_seed_bytes_day(epoch, day, label, class, d, D_0)` | `verify.rs` | the day-sized entries; the v2 entries are unchanged and build the v2 shape |
| `IGNEUM_MIXER_MULT`, `IGNEUM_CLASS_MIXER_MULT`, `IGNEUM_CACHE_GROWTH` in program.h; `"mixer_mult"`, `"cache_growth"`, the `"item"` string in program.json | `emit.rs` | written only for a class with `m != 1` or growth, so v2 packs do not change |
| `packfile.h` `mixerMult`; `packbench` and the OpenCL host print the multiplier and the cache size | the three hosts | the kernels carry the construction in their text (one emitter, three dialects); the hosts size the cache from `IGNEUM_CACHE_LOG2_WORDS` already (packbench.swift line 56, host.cu line 65, host.c line 1966) |
| `igneum-pow --class mx4 [--days d]` on every command | `main.rs` | `--days` sizes the cache for a growth class |
Under the ca2-v3 seam (`ProgramClass::V3`, `V3_CLASS`), the integration sets `V3_CLASS = LoadClass::MX4`; the
chain's day-sized dataset needs the day index, so `Epoch::chain_dataset(day, class)` builds the genesis-size
cache and a `chain_dataset_day(day_bytes, class, d, D_0)` beside it is the growth entry (section 9, owed to the
node agent).
## 4. Vectors (class v3, `proto-cuda/packs-ca2-mixer/`)
Filled in section 6 from the exported packs: `mx4-genesis` (seed `igneum-genesis`, day `2026-10-03`, 2^28 words,
2^26-word cache) and `mx4-devnet-epoch0` (the devnet genesis hash as the epoch seed, day bytes of 2026-10-04).
## 5. The v2 path is byte-identical
`cargo test --test packs` regenerates every file of `igneum-genesis-mh` and `igneum-devnet-v4-epoch0` from
`program.json` and compares byte for byte (`emitted_sources_match_all_packs`, `export_pack_matches_all_packs`);
section 6 also records a fresh `igneum-pow export` of both packs diffed against the checked-in directories.
## 6. Measurements
Filled as they land. Every row names the machine, the date, the command and the lock mode.
## 7. The chip model
`docs/analysis/chip-model-v3.md`.
## 8. What is unverified
Filled at the end.
## 9. Owed
Filled at the end.

View file

@ -11,8 +11,8 @@
use crate::generator::{Instr, Op, Program, ProgramClass, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS};
use crate::memhard::{
MixParams, CACHE_LINES_PER_SEGMENT, CACHE_LINE_MASK, CACHE_LOG2_WORDS, CACHE_SEGMENTS, CACHE_SEGMENT_LOG2_LINES,
CACHE_TAG, CACHE_WORDS, CHACHA_ROUNDS, CHACHA_SIGMA, ITEM_ROUNDS,
MixParams, Shape, CACHE_LINES_PER_SEGMENT, CACHE_SEGMENT_LOG2_LINES, CACHE_TAG, CHACHA_ROUNDS, CHACHA_SIGMA,
ITEM_ROUNDS,
};
use crate::seed::SplitMix64;
use crate::verify::{DatasetMode, DatasetSource, Epoch, FOLD_MUL, FOLD_ROT};
@ -100,12 +100,25 @@ fn class_header_lines(p: &Program) -> String {
return String::new();
}
let mut s = String::new();
s.push_str("// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads
if p.class.v2_loads() {
s.push_str("// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
");
s.push_str("// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x.
s.push_str("// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
");
} else {
s.push_str("// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads
");
s.push_str("// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x.
");
}
s.push_str(&format!("#define IGNEUM_LOAD_CLASS {}
", jstr(&p.class.name())));
if p.class.mixer_mult != 1 || p.class.growth {
s.push_str(&format!("#define IGNEUM_CLASS_MIXER_MULT {}
", p.class.mixer_mult));
s.push_str(&format!("#define IGNEUM_CACHE_GROWTH {} // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
", p.class.growth as u8));
}
s.push_str(&format!("#define IGNEUM_LOAD_SLOTS {}
", p.class.load_slots));
s.push_str(&format!("#define IGNEUM_LOAD_MIX {{ {}, {}, {} }}
@ -257,12 +270,19 @@ pub enum LoadSource<'a> {
InlineMemhard(&'a MixParams),
}
fn log2_segments() -> usize {
CACHE_SEGMENTS.trailing_zeros() as usize
fn log2_segments(shape: &Shape) -> usize {
shape.log2_segments() as usize
}
/// The memory-hard core as source text (`emitMemhardCore`). Every parameter is a literal.
/// The memory-hard core as source text (`emitMemhardCore`). Every parameter is a literal: the mixer constants of
/// the day and, from `mp.shape`, the cache size and the mixer multiplier `m`. Under `m = 1` the text is version
/// 2's byte for byte; under `m > 1` the item loop applies `mh_mixer` `m` times per round with the keys
/// `round_key(r m + j)` (spec 01 section 1.8.5 under class v3).
pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
let shape = &mp.shape;
let m = shape.mixer_mult;
let cache_log2_words = shape.cache_log2_words;
let cache_line_mask = shape.cache_line_mask();
let (u, fn_, cptr, wptr, lptr, lcptr) = match dialect {
CoreDialect::Metal => {
("uint", "inline", "device const uint*", "device uint*", "thread uint*", "const thread uint*")
@ -274,18 +294,23 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
};
let k = &mp.key;
let r = &mp.rot;
let m = &mp.mul;
let mul = &mp.mul;
let c = &mp.rc;
let mut s = String::with_capacity(6000);
s.push_str(&format!(
"// Memory-hard dataset core (MEMHARD.md). Cache: 2^{} words in 2^{} segments of {} chained ChaCha{} lines.\n",
CACHE_LOG2_WORDS,
log2_segments(),
cache_log2_words,
log2_segments(shape),
CACHE_LINES_PER_SEGMENT,
CHACHA_ROUNDS
));
s.push_str("// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.\n");
s.push_str(&format!("#define MH_CACHE_LINE_MASK {}\n", hex(CACHE_LINE_MASK)));
if m == 1 {
s.push_str("// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.\n");
} else {
s.push_str(&format!("// Item: 8 rounds of {m} x seed-parameterised mixer + one 64-byte cache read, then {m} x final mixer (class v3, mixer multiplier {m},\n"));
s.push_str("// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.\n");
}
s.push_str(&format!("#define MH_CACHE_LINE_MASK {}\n", hex(cache_line_mask)));
s.push_str(&format!("#define MH_SEGMENT_LINES {}u\n", CACHE_LINES_PER_SEGMENT));
s.push_str("#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }\n");
s.push_str(&format!(
@ -345,7 +370,7 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
s.push_str("// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.\n");
s.push_str(&format!("{fn_} void mh_mixer({lptr} s, {u} rk) {{\n"));
for i in 0..16 {
s.push_str(&format!(" s[{i}] = (s[{i}] ^ ({} + rk)) * {};\n", hex(c[i]), hex(m[i])));
s.push_str(&format!(" s[{i}] = (s[{i}] ^ ({} + rk)) * {};\n", hex(c[i]), hex(mul[i])));
}
let col = (0..4).map(|i| format!("{}u", r[i])).collect::<Vec<_>>().join(", ");
let dia = (4..8).map(|i| format!("{}u", r[i])).collect::<Vec<_>>().join(", ");
@ -355,22 +380,39 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
s.push_str(&format!(" MH_QR(s[2], s[7], s[8], s[13], {dia}) MH_QR(s[3], s[4], s[9], s[14], {dia})\n"));
s.push_str("}\n");
s.push('\n');
s.push_str(&format!(
"// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of mixer + cache line s[0] & mask; final mixer.\n"
));
if m == 1 {
s.push_str(&format!(
"// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of mixer + cache line s[0] & mask; final mixer.\n"
));
} else {
s.push_str(&format!(
"// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of {m} x mixer + cache line s[0] & mask; {m} x final mixer.\n"
));
}
s.push_str(&format!("{fn_} void mh_item({cptr} cache, {u} t, {lptr} s) {{\n"));
for i in 0..8 {
s.push_str(&format!(" s[{i}] = {};\n", hex(k[i])));
}
for i in 0..8 {
s.push_str(&format!(" s[{}] = t * {} + {};\n", 8 + i, hex(m[i]), hex(c[i])));
s.push_str(&format!(" s[{}] = t * {} + {};\n", 8 + i, hex(mul[i]), hex(c[i])));
}
s.push_str(&format!(" for ({u} r = 0u; r < {ITEM_ROUNDS}u; ++r) {{\n"));
s.push_str(" mh_mixer(s, 0x9E3779B9u * (r + 1u));\n");
if m == 1 {
s.push_str(" mh_mixer(s, 0x9E3779B9u * (r + 1u));\n");
} else {
s.push_str(&format!(" for ({u} j = 0u; j < {m}u; ++j) mh_mixer(s, 0x9E3779B9u * (r * {m}u + j + 1u));\n"));
}
s.push_str(&format!(" {cptr} line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);\n"));
s.push_str(&format!(" for ({u} i = 0u; i < 16u; ++i) s[i] ^= line[i];\n"));
s.push_str(" }\n");
s.push_str(&format!(" mh_mixer(s, 0x9E3779B9u * {}u);\n", ITEM_ROUNDS + 1));
if m == 1 {
s.push_str(&format!(" mh_mixer(s, 0x9E3779B9u * {}u);\n", ITEM_ROUNDS + 1));
} else {
s.push_str(&format!(
" for ({u} j = 0u; j < {m}u; ++j) mh_mixer(s, 0x9E3779B9u * ({}u + j + 1u));\n",
ITEM_ROUNDS as u32 * m
));
}
s.push_str("}\n");
s.push_str("// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.\n");
s.push_str(&format!(
@ -386,7 +428,7 @@ pub fn metal_memhard(mp: &MixParams) -> String {
s.push_str("using namespace metal;\n");
s.push_str(&emit_memhard_core(mp, CoreDialect::Metal));
s.push('\n');
s.push_str(&format!("// One thread per segment (2^{} threads).\n", log2_segments()));
s.push_str(&format!("// One thread per segment (2^{} threads).\n", log2_segments(&mp.shape)));
s.push_str(
"kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {\n",
);
@ -1116,10 +1158,13 @@ pub fn program_header(p: &Program, day: &str, ds: &DatasetSource) -> String {
s.push_str(&format!("#define IGNEUM_SEEDW_INIT {{ {} }}\n", join_hex(&p.seed)));
if let Some(mp) = memhard {
s.push_str(&format!("#define IGNEUM_KEY_INIT {{ {} }}\n", join_hex(&mp.key)));
s.push_str(&format!("#define IGNEUM_CACHE_LOG2_WORDS {CACHE_LOG2_WORDS}\n"));
s.push_str(&format!("#define IGNEUM_CACHE_LOG2_WORDS {}\n", mp.shape.cache_log2_words));
s.push_str(&format!("#define IGNEUM_CACHE_SEGMENT_LOG2_LINES {CACHE_SEGMENT_LOG2_LINES}\n"));
s.push_str(&format!("#define IGNEUM_CACHE_SEGMENTS {CACHE_SEGMENTS}u\n"));
s.push_str(&format!("#define IGNEUM_CACHE_SEGMENTS {}u\n", mp.shape.cache_segments()));
s.push_str(&format!("#define IGNEUM_ITEM_ROUNDS {ITEM_ROUNDS}\n"));
if mp.shape.mixer_mult != 1 {
s.push_str(&format!("#define IGNEUM_MIXER_MULT {} // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)\n", mp.shape.mixer_mult));
}
s.push_str(&format!(
"#define IGNEUM_MIX_ROT_INIT {{ {} }}\n",
mp.rot.iter().map(|r| format!("{r}u")).collect::<Vec<_>>().join(", ")
@ -1188,6 +1233,8 @@ pub struct PackVectors {
pub cache_last: Vec<u32>,
/// FNV-1a 64 over the whole cache (memory-hard only)
pub cache_fnv: u64,
/// The cache is 2^cache_log2_words words (memory-hard only; 26 under version 2)
pub cache_log2_words: u32,
}
/// The base nonces of the three vector warps every pack carries.
@ -1249,7 +1296,8 @@ pub fn vectors_header(
s.push_str("};\n");
if memhard {
s.push_str(&format!(
"// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^{CACHE_LOG2_WORDS} words.\n"
"// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^{} words.\n",
v.cache_log2_words
));
s.push_str("static const uint32_t IGNEUM_CACHE_HEAD[16] = {\n");
s.push_str(&format!(" {},\n", join_hex(&v.cache_head[..8])));
@ -1300,6 +1348,11 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
if !p.class.is_v2() {
let c = p.width_counts();
s.push_str(&format!(" \"load_class\": {},\n", jstr(&p.class.name())));
if p.class.mixer_mult != 1 || p.class.growth {
s.push_str(&format!(" \"mixer_mult\": {},\n", p.class.mixer_mult));
s.push_str(&format!(" \"cache_growth\": {},\n", p.class.growth));
s.push_str(&format!(" \"mixer\": \"class v3 (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): every mixer application of the item derivation is {} applications with round keys (r * {} + j + 1) * 0x9E3779B9, the 8 dependent cache reads per item unchanged; cache growth rule option C: cache words = 2^(26 + doublings(day)), dataset words = 2^(genesis_log2 + doublings(day)), doublings(day) = floor(log2(1 + day / 1460)) for day = days since genesis\",\n", p.class.mixer_mult, p.class.mixer_mult));
}
s.push_str(&format!(" \"load_slots\": {},\n", p.class.load_slots));
s.push_str(&format!(" \"load_mix_percent_4_16_64\": [{}, {}, {}],\n", p.class.mix[0], p.class.mix[1], p.class.mix[2]));
s.push_str(&format!(" \"load_width_counts_4_16_64\": [{}, {}, {}],\n", c[0], c[1], c[2]));
@ -1351,9 +1404,12 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
s.push_str(" \"spec\": \"proto-metal/MEMHARD.md\",\n");
s.push_str(&format!(" \"key\": [{}],\n", join_jhex(&mp.key)));
s.push_str(" \"key_derivation\": \"the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]\",\n");
let shape = &mp.shape;
s.push_str(&format!(
" \"cache\": {{\"log2_words\": {CACHE_LOG2_WORDS}, \"bytes\": {}, \"line_words\": 16, \"segment_lines\": {CACHE_LINES_PER_SEGMENT}, \"segments\": {CACHE_SEGMENTS}, \"block\": \"ChaCha{CHACHA_ROUNDS} core + feed-forward, rotations 16 12 8 7\", \"sigma\": [{}], \"tag\": [{}], \"chain\": \"in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0\"}},\n",
CACHE_WORDS as u64 * 4,
" \"cache\": {{\"log2_words\": {}, \"bytes\": {}, \"line_words\": 16, \"segment_lines\": {CACHE_LINES_PER_SEGMENT}, \"segments\": {}, \"block\": \"ChaCha{CHACHA_ROUNDS} core + feed-forward, rotations 16 12 8 7\", \"sigma\": [{}], \"tag\": [{}], \"chain\": \"in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0\"}},\n",
shape.cache_log2_words,
shape.cache_words() as u64 * 4,
shape.cache_segments(),
join_jhex(&CHACHA_SIGMA),
join_jhex(&CACHE_TAG)
));
@ -1364,11 +1420,23 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
join_jhex(&mp.rc)
));
// The Swift writes jhex(cacheLineMask) here, which breaks the JSON. We write the bare literal.
s.push_str(&format!(
" \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = M_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = M_{ITEM_ROUNDS}(s); item(t) = s\",\n",
ITEM_ROUNDS - 1,
CACHE_LINE_MASK
));
if shape.mixer_mult == 1 {
s.push_str(&format!(
" \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = M_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = M_{ITEM_ROUNDS}(s); item(t) = s\",\n",
ITEM_ROUNDS - 1,
shape.cache_line_mask()
));
} else {
let m = shape.mixer_mult;
s.push_str(&format!(
" \"mixer_mult\": {m},\n \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: for j in 0..{}: s = M(s, rk = (r * {m} + j + 1) * 0x9E3779B9); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then for j in 0..{}: s = M(s, rk = ({} + j + 1) * 0x9E3779B9); item(t) = s\",\n",
ITEM_ROUNDS - 1,
m - 1,
shape.cache_line_mask(),
m - 1,
ITEM_ROUNDS as u32 * m
));
}
s.push_str(" \"word\": \"dataset[w] = item(w >> 4)[w & 15]\"\n");
} else {
s.push_str(" \"mode\": \"closed-form\",\n");
@ -1507,6 +1575,7 @@ pub fn export_pack(epoch: &Epoch, day: &str, source: &str) -> Pack {
v.cache_head = w[..16].to_vec();
v.cache_last = w[w.len() - 16..].to_vec();
v.cache_fnv = m.cache.fnv1a64();
v.cache_log2_words = m.shape().cache_log2_words;
}
let is_mh = memhard.is_some();
let mut files = vec![

View file

@ -181,6 +181,13 @@ pub struct LoadClass {
/// Variant 5: the scratch per warp in KiB (32 or 128; the whole working set of a card at full occupancy must
/// stay under 6 GB, coordinator's cap of 5 October 2026). 0 for every other class.
pub scratch_kb: u8,
/// Mixer cost multiplier `m` of the dataset item derivation (Counter ASIC 2.0, M16, decided 5 October 2026 for
/// class v3): every mixer application of spec 01 section 1.8.5 becomes `m` applications with distinct round
/// keys, the 8 dependent cache reads per item unchanged (`memhard::derive_items`). 1 for version 2, 4 for v3.
pub mixer_mult: u8,
/// Cache growth rule, option C (`memhard::growth_doublings`): the cache doubles when the dataset doubles. `false`
/// for version 2 (the cache is 2^26 words on every day), `true` for v3.
pub growth: bool,
}
/// Scratch geometry (variant 5): 16-byte slots, lane-major, 32 lanes per warp; `scratch_kb` KiB per warp gives
@ -205,20 +212,27 @@ impl LoadClass {
impl LoadClass {
/// Generator version 2 as adopted on 4 October 2026: 16 loads of one word. The lottery hash.
pub const V2: LoadClass = LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 };
pub const V2: LoadClass =
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 1, growth: false };
/// The construction decided for program class v3 on 5 October 2026 (Counter ASIC 2.0, `docs/plans/mixer-x4.md`):
/// version 2 loads (16 slots of one word, no scratch, no width roll, so the program stream is version 2's), the
/// mixer applied 4 times per round, and the cache growth rule. Name "mx4".
pub const MX4: LoadClass =
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 4, growth: true };
/// A fixed width (1, 4 or 16 words) with `load_slots` loads per program.
pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass {
let mut mix = [0u8; 3];
let i = WIDTH_WORDS.iter().position(|&w| w == width_words).expect("width must be 1, 4 or 16 words");
mix[i] = 100;
LoadClass { mix, load_slots, scratch: None, scratch_kb: 0 }
LoadClass { mix, load_slots, ..LoadClass::V2 }
}
/// Per-load width drawn from `mix` (percent for 4, 16, 64 bytes), 16 loads per program.
pub fn mixed(mix: [u8; 3]) -> LoadClass {
assert_eq!(mix.iter().map(|&m| m as u32).sum::<u32>(), 100, "the mix must sum to 100");
LoadClass { mix, load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 }
LoadClass { mix, ..LoadClass::V2 }
}
/// Variant 5: version 2 widths, 16 memory operations of which `k` are scratch read-modify-writes into a
@ -226,11 +240,61 @@ impl LoadClass {
pub fn scratch(k: u8, kb: u8) -> LoadClass {
assert!(k as usize <= LOAD_SLOTS);
assert!(kb.is_power_of_two() && kb <= 128, "scratch per warp must be a power of two up to 128 KiB");
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: Some(k), scratch_kb: kb }
LoadClass { scratch: Some(k), scratch_kb: kb, ..LoadClass::V2 }
}
/// Parse "p4,p16,p64" or one of the names of [`LoadClass::name`] ("scr4k32": 4 scratch ops, 32 KiB per warp).
/// This class with the mixer multiplier `m` (1, 2, 4, 8 or 16) and the cache growth rule on or off.
pub fn with_mixer(self, mixer_mult: u8, growth: bool) -> LoadClass {
assert!(mixer_mult >= 1 && mixer_mult <= 16 && mixer_mult.is_power_of_two(), "mixer multiplier must be 1, 2, 4, 8 or 16");
LoadClass { mixer_mult, growth, ..self }
}
/// The mixer multiplier as the item derivation uses it.
pub fn mixer_mult(&self) -> u32 {
self.mixer_mult as u32
}
/// Whether the loads of this class are version 2's: 16 one-word loads, no scratch. Such a class takes no width
/// roll, so its program stream is the version 2 stream draw for draw (the mixer and the cache are properties of
/// the dataset, not of the program).
pub fn v2_loads(&self) -> bool {
self.mix == [100, 0, 0] && self.load_slots as usize == LOAD_SLOTS && self.scratch.is_none()
}
/// Whether every instruction takes the tenth draw (the width roll): every class whose loads are not version 2's.
pub fn takes_width_roll(&self) -> bool {
!self.v2_loads()
}
/// Parse "p4,p16,p64" or one of the names of [`LoadClass::name`] ("scr4k32": 4 scratch ops, 32 KiB per warp;
/// "mx4": the v3 construction; a trailing "m<mult>" and "g" set the mixer multiplier and the growth rule on any
/// load class, "w16m4g" for example).
pub fn parse(s: &str) -> Option<LoadClass> {
if s == "mx4" {
return Some(LoadClass::MX4);
}
// the mixer suffix: "...m<mult>" then an optional "g"
let (s, growth) = match s.strip_suffix('g') {
Some(base) if base.rsplit_once('m').map(|(_, d)| !d.is_empty() && d.bytes().all(|b| b.is_ascii_digit())).unwrap_or(false) => (base, true),
_ => (s, false),
};
if let Some((base, digits)) = s.rsplit_once('m') {
if !digits.is_empty() && digits.bytes().all(|b| b.is_ascii_digit()) && !base.is_empty() && !base.ends_with(',') {
let mult: u8 = digits.parse().ok()?;
if mult == 0 || mult > 16 || !mult.is_power_of_two() {
return None;
}
return Some(LoadClass::parse_loads(base)?.with_mixer(mult, growth));
}
}
if growth {
return None;
}
LoadClass::parse_loads(s)
}
/// The load part of a class name (no mixer suffix).
fn parse_loads(s: &str) -> Option<LoadClass> {
if let Some(rest) = s.strip_prefix("scr") {
let (k, kb) = rest.split_once('k')?;
let k: u8 = k.parse().ok()?;
@ -260,7 +324,7 @@ impl LoadClass {
if slots == 0 || slots as usize >= INSTR_COUNT {
return None;
}
Some(LoadClass { mix, load_slots: slots, scratch: None, scratch_kb: 0 })
Some(LoadClass { mix, load_slots: slots, ..LoadClass::V2 })
}
/// Scratch read-modify-writes per program (0 without a scratch).
@ -272,25 +336,41 @@ impl LoadClass {
*self == LoadClass::V2
}
/// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4".
/// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4k32"; "mx4" for the v3 construction;
/// any other mixer setting appends "m<mult>" and, with the growth rule, "g" ("v2m2", "w16m4g").
pub fn name(&self) -> String {
if self.is_v2() {
return "v2".to_string();
}
if let Some(k) = self.scratch {
return format!("scr{k}k{}", self.scratch_kb);
if *self == LoadClass::MX4 {
return "mx4".to_string();
}
let base = match self.mix {
[100, 0, 0] => "w4".to_string(),
[0, 100, 0] => "w16".to_string(),
[0, 0, 100] => "w64".to_string(),
[a, b, c] => format!("mix{a}-{b}-{c}"),
};
if self.load_slots as usize == LOAD_SLOTS {
base
let loads = LoadClass { mixer_mult: 1, growth: false, ..*self };
let base = if loads.is_v2() {
"v2".to_string()
} else if let Some(k) = self.scratch {
format!("scr{k}k{}", self.scratch_kb)
} else {
format!("{base}x{}", self.load_slots)
let base = match self.mix {
[100, 0, 0] => "w4".to_string(),
[0, 100, 0] => "w16".to_string(),
[0, 0, 100] => "w64".to_string(),
[a, b, c] => format!("mix{a}-{b}-{c}"),
};
if self.load_slots as usize == LOAD_SLOTS {
base
} else {
format!("{base}x{}", self.load_slots)
}
};
let mut s = base;
if self.mixer_mult != 1 || self.growth {
s.push_str(&format!("m{}", self.mixer_mult));
}
if self.growth {
s.push('g');
}
s
}
/// The width in words of a load whose width roll (0..99) is `roll`: the first entry of the mix whose cumulative
@ -333,12 +413,11 @@ pub enum ProgramClass {
V3,
}
/// PLACEHOLDER (ca2-node worker, 5 October 2026): the load class of program class v3 is set to w16 (16 loads of
/// four words, `LoadClass::fixed(4, 16)`) so the node, the miner, the packs and the fast-time gate can be built and
/// run before the Counter ASIC 2.0 measurements decide the width, the per-load mix and the scratch share
/// (`counter-asic-2-rollout.md` section 6). The integration agent on branch ca2-v3 replaces this constant with the
/// decided class; nothing else in the seam names the class, so it is one edit.
pub const V3_CLASS: LoadClass = LoadClass { mix: [0, 100, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 };
/// The load class of program class v3, decided 5 October 2026 (Counter ASIC 2.0, `docs/plans/counter-asic-2-status.md`
/// "22:00 decided", `docs/plans/mixer-x4.md`): [`LoadClass::MX4`], version 2 loads (the width stays 4 bytes, the
/// per-load mix and the scratch share are out), the mixer applied 4 times per round and the cache growth rule. The
/// placeholder of the seam (w16) is replaced here; nothing else in the seam names the class.
pub const V3_CLASS: LoadClass = LoadClass::MX4;
impl ProgramClass {
/// The load class this program class draws from.
@ -494,6 +573,14 @@ pub fn program_id_class(generator: u32, seed: &[u32; 8], attempt: u32, class: &L
b.push(k);
b.push(class.scratch_kb);
}
if class.mixer_mult != 1 || class.growth {
// Counter ASIC 2.0: the mixer multiplier and the growth rule are part of the construction, so a program of
// the same seed under a different mixer carries a different id (under the v3 seam the id is
// program_id(3, seed, attempt) and this branch is not taken)
b.extend_from_slice(b"mixer/");
b.push(class.mixer_mult);
b.push(class.growth as u8);
}
fnv1a64(&b)
}
@ -625,7 +712,8 @@ pub fn candidate_from_words_class(
let rot = 1 + rng.below(31) as u32;
let bit = rng.below(32);
let mask = 1u8 << rng.below(5);
let width = if class.is_v2() { 1 } else { class.width_for_roll(rng.below(100)) };
// Version 2 loads take no width roll, so a mixer class with version 2 loads draws the version 2 program
let width = if class.takes_width_roll() { class.width_for_roll(rng.below(100)) } else { 1 };
let width = if op == Op::Load { width } else { 1 };
if op.is_load() {
fresh[src as usize] = false;

View file

@ -32,7 +32,7 @@ pub mod verify;
pub use bind::{block_init_words, day_bytes, pow256_from_lane, target64_from_le256};
pub use accept::{check as accept_program, AcceptReport, Reject};
pub use generator::{generate, generate_from_seed_bytes, generate_from_seed_bytes_program_class, Instr, Op, Program, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, V3_CLASS};
pub use memhard::{Cache, MemhardCpu, MixParams};
pub use generator::{generate, generate_from_seed_bytes, generate_from_seed_bytes_program_class, Instr, LoadClass, Op, Program, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, V3_CLASS};
pub use memhard::{cache_log2_words, dataset_log2_words, days_since_genesis, growth_doublings, Cache, MemhardCpu, MixParams, Shape};
pub use seed::{fnv1a64, seed_words, SplitMix64};
pub use verify::{hash_warp, interpret_warp_init, verify_block, DatasetMode, DatasetSource, Epoch};

View file

@ -13,7 +13,7 @@
use igneum_pow::emit::export_pack;
use igneum_pow::generator::LoadClass;
use igneum_pow::memhard::Cache;
use igneum_pow::memhard::{Cache, Shape};
use igneum_pow::seed::day_key;
use igneum_pow::verify::{DatasetMode, Epoch, DEFAULT_DATASET_LOG2};
use std::time::Instant;
@ -31,6 +31,8 @@ struct Args {
epoch_hex: Option<String>,
day_hex: Option<String>,
class: LoadClass,
/// Days since genesis for the cache growth rule of a class with `growth` (0: the genesis cache).
days: u64,
}
fn usage() -> ! {
@ -42,7 +44,8 @@ fn usage() -> ! {
\x20 hash-bound --prehash <64 hex> --nonce <u64> print the header-bound hash (bind.rs) of one 64-bit nonce\n\
\x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\
\x20 show the accepted program, one instruction per line\n\
\x20 --class C load class (read-width experiment): v2 (default), w4, w16, w64, w64x4, or p4,p16,p64[xN]"
\x20 --class C load class: v2 (default), mx4 (class v3: mixer x4, cache growth), w4, w16, w64, w64x4, p4,p16,p64[xN], <class>m<mult>[g]\n\
\x20 --days N days since genesis for the cache growth rule of a class with it (default 0: the 2^26-word cache)"
);
std::process::exit(2)
}
@ -61,6 +64,7 @@ fn parse() -> Args {
epoch_hex: None,
day_hex: None,
class: LoadClass::V2,
days: 0,
};
let mut it = std::env::args().skip(1);
a.cmd = it.next().unwrap_or_else(|| usage());
@ -78,6 +82,7 @@ fn parse() -> Args {
"--epoch-hex" => a.epoch_hex = Some(val()),
"--day-hex" => a.day_hex = Some(val()),
"--class" => a.class = LoadClass::parse(&val()).unwrap_or_else(|| usage()),
"--days" => a.days = val().parse().unwrap_or_else(|_| usage()),
_ => usage(),
}
}
@ -93,7 +98,7 @@ fn main() {
"accept" => accept(&a),
"show" => show(&a),
"hash" => {
let e = Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class);
let e = Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days);
println!("{:016x}", e.hash(a.nonce as u32));
}
"hash-bound" => {
@ -106,7 +111,7 @@ fn main() {
let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage());
Epoch::from_seed_bytes_class(&eb, &db, "cli", a.class)
}
_ => Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class),
_ => Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days),
};
let init = igneum_pow::bind::block_init_words(&prehash, a.nonce);
println!("init words {}", init.iter().map(|w| format!("{w:08x}")).collect::<Vec<_>>().join(" "));
@ -124,24 +129,34 @@ fn bench(a: &Args, mode: DatasetMode) {
a.dataset_log2,
mode.name()
);
let shape = Shape::for_class_day(&a.class, a.days);
if mode == DatasetMode::MemoryHard {
// Time the cache fill on its own first (one core), then build the epoch (which fills it again).
let t0 = Instant::now();
let c = Cache::fill(day_key(&a.day));
let c = Cache::fill_log2(day_key(&a.day), shape.cache_log2_words);
let fill_ms = t0.elapsed().as_secs_f64() * 1e3;
println!("cache: fill {fill_ms:.1} ms on one core (2^26 words, 65536 chains of 64 ChaCha12 blocks), FNV-1a 64 {:016x}", c.fnv1a64());
println!(
"cache: fill {fill_ms:.1} ms on one core (2^{} words, {} MiB, {} chains of 64 ChaCha12 blocks), FNV-1a 64 {:016x}",
shape.cache_log2_words,
shape.cache_words() * 4 / (1 << 20),
c.segments(),
c.fnv1a64()
);
drop(c);
}
let t0 = Instant::now();
let e = Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class);
let e = Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days);
let build_ms = t0.elapsed().as_secs_f64() * 1e3;
println!(
"program: class {}, {} loads/hash, {} bytes/hash, widths (1,4,16 words) {:?}, {} items/warp, op mix {}; epoch built in {build_ms:.1} ms",
"program: class {}, {} loads/hash, {} bytes/hash, widths (1,4,16 words) {:?}, {} items/warp, mixer x{} ({} mixers/item), cache 2^{} words, op mix {}; epoch built in {build_ms:.1} ms",
e.program.class.name(),
e.program.loads_per_hash(),
e.program.bytes_per_hash(),
e.program.width_counts(),
e.program.items_per_warp(),
shape.mixer_mult,
shape.mixers_per_item(),
shape.cache_log2_words,
e.program.op_mix()
);
let bases = [0u32, 4096, 1_000_000];
@ -175,7 +190,7 @@ fn export(a: &Args, mode: DatasetMode) {
let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage());
(Epoch::from_seed_bytes_class(&eb, &db, &format!("igneum-epoch/{eh}/day/{dh}"), a.class), format!("bytes:{dh}"))
}
_ => (Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class), a.day.clone()),
_ => (Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days), a.day.clone()),
};
let build_ms = t0.elapsed().as_secs_f64() * 1e3;
println!("igneum-pow export {out}");

View file

@ -3,12 +3,18 @@
//! ARX-multiply mixer. The verifier holds the cache and never the dataset.
//!
//! All arithmetic is on u32 modulo 2^32. Rotations are by 1..31 at every call site.
//!
//! Counter ASIC 2.0 (5 October 2026, `docs/plans/mixer-x4.md`, behind the program class): the construction has a
//! [`Shape`], the mixer multiplier `m` and the cache size. Under `m` every mixer application of an item becomes
//! `m` applications with distinct round keys, the 8 dependent cache reads unchanged; the cache doubles when the
//! dataset doubles ([`growth_doublings`]). [`Shape::V2`] (`m = 1`, 2^26 words) is version 2 bit for bit.
use crate::generator::LoadClass;
use crate::seed::{day_key, fnv1a64_words, SplitMix64};
pub const CACHE_LOG2_WORDS: usize = 26;
pub const CACHE_SEGMENT_LOG2_LINES: usize = 6;
/// 2^26 words = 256 MiB.
/// 2^26 words = 256 MiB (the version 2 cache, and the v3 cache until the first dataset doubling).
pub const CACHE_WORDS: usize = 1 << CACHE_LOG2_WORDS;
/// 2^22 lines of 16 words.
pub const CACHE_LINES: usize = CACHE_WORDS >> 4;
@ -24,6 +30,94 @@ pub const CHACHA_SIGMA: [u32; 4] = [0x61707865, 0x3320646e, 0x79622d32, 0x6b2065
/// "Igne", "umMH".
pub const CACHE_TAG: [u32; 2] = [0x49676e65, 0x756d4d48];
/// The shape of the item derivation and of the cache: the mixer multiplier and the cache size.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
pub struct Shape {
/// Mixer applications per round (and after the last read): 1 under version 2, 4 under class v3.
pub mixer_mult: u32,
/// The cache is 2^cache_log2_words words (26 at genesis; 27 and 28 after the dataset doublings of 1.13.3).
pub cache_log2_words: u32,
}
impl Shape {
/// Version 2: one mixer application per round, a 2^26-word cache.
pub const V2: Shape = Shape { mixer_mult: 1, cache_log2_words: CACHE_LOG2_WORDS as u32 };
/// The shape of a load class on day 0 of the chain (and on every day for a class without the growth rule).
pub fn for_class(class: &LoadClass) -> Shape {
Shape::for_class_day(class, 0)
}
/// The shape of a load class on day `days_since_genesis` of the chain: the class's multiplier, and the cache
/// of [`cache_log2_words`] when the class has the growth rule, else 2^26 words.
pub fn for_class_day(class: &LoadClass, days_since_genesis: u64) -> Shape {
Shape {
mixer_mult: class.mixer_mult(),
cache_log2_words: if class.growth { cache_log2_words(days_since_genesis) } else { CACHE_LOG2_WORDS as u32 },
}
}
pub fn is_v2(&self) -> bool {
*self == Shape::V2
}
pub fn cache_words(&self) -> usize {
1usize << self.cache_log2_words
}
pub fn cache_lines(&self) -> usize {
self.cache_words() >> 4
}
pub fn cache_line_mask(&self) -> u32 {
(self.cache_lines() - 1) as u32
}
pub fn cache_segments(&self) -> usize {
self.cache_lines() >> CACHE_SEGMENT_LOG2_LINES
}
pub fn log2_segments(&self) -> u32 {
self.cache_log2_words - 4 - CACHE_SEGMENT_LOG2_LINES as u32
}
/// Mixer applications per item: `(ITEM_ROUNDS + 1) x m`.
pub fn mixers_per_item(&self) -> u32 {
(ITEM_ROUNDS as u32 + 1) * self.mixer_mult
}
}
// --------------------------------------------------------------------------------------------------------------
// Dataset growth, option C (spec 01 section 1.13.3 option (b) with the cache tied to the dataset's doublings)
// --------------------------------------------------------------------------------------------------------------
/// Days per year of the growth schedule: one year = 31,536,000 DAA seconds of 86,400 (spec 01 section 1.13.3).
pub const GROWTH_DAYS_PER_YEAR: u64 = 365;
/// The linear schedule of 1.13.3, 2 GiB at genesis plus 0.5 GiB per year, is `G x (1 + d / 1460)` for the genesis
/// size `G` and the day `d`: it doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at
/// day 10,220 (year 28) and 16x at day 21,900 (year 60).
pub const GROWTH_DOUBLING_DAYS: u64 = 4 * GROWTH_DAYS_PER_YEAR;
/// The number of dataset doublings reached by day `days_since_genesis` of the chain: `floor(log2(1 + d / 1460))`,
/// in integers (`1 + d / 1460` rounded down, then its integer log2, which equals the real log2's floor because a
/// power of two is an integer). 0 until day 1,459; 1 from day 1,460 (year 4); 2 from day 4,380 (year 12).
pub fn growth_doublings(days_since_genesis: u64) -> u32 {
(1 + days_since_genesis / GROWTH_DOUBLING_DAYS).ilog2()
}
/// The cache size on day `d` under option C: 2^26 words doubled once per dataset doubling (256 MiB, 512 MiB from
/// year 4, 1 GiB from year 12).
pub fn cache_log2_words(days_since_genesis: u64) -> u32 {
CACHE_LOG2_WORDS as u32 + growth_doublings(days_since_genesis)
}
/// The dataset size on day `d` under option (b) of 1.13.3: the genesis size (2^`genesis_log2_words` words: 28 for
/// the 1 GiB packs and the devnet, 29 for the designed 2 GiB) doubled once per doubling of the linear schedule. The
/// result is capped at 32 (the item index is 32 bits, spec 1.13.3).
pub fn dataset_log2_words(genesis_log2_words: u32, days_since_genesis: u64) -> u32 {
(genesis_log2_words + growth_doublings(days_since_genesis)).min(32)
}
/// Days since genesis from two day indices of `bind::day_index` (the header's `timestamp_ms / 86,400,000`): the
/// day of the block and the day of the genesis header. A block before the genesis day (clock skew) is day 0.
pub fn days_since_genesis(day_index: u64, genesis_day_index: u64) -> u64 {
day_index.saturating_sub(genesis_day_index)
}
#[inline(always)]
fn rotl(x: u32, n: u32) -> u32 {
x.rotate_left(n)
@ -66,17 +160,23 @@ pub fn chacha_block(x: &[u32; 16]) -> [u32; 16] {
y
}
/// Mixer parameters drawn from the day key. Draw order: ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15].
/// Mixer parameters drawn from the day key, plus the [`Shape`] the mixer is applied under. Draw order:
/// ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15]. The shape is not drawn: it is the class's.
#[derive(Clone, Debug, PartialEq, Eq)]
pub struct MixParams {
pub key: [u32; 8],
pub rot: [u32; 8],
pub mul: [u32; 16],
pub rc: [u32; 16],
pub shape: Shape,
}
impl MixParams {
/// Version 2 shape.
pub fn new(key: [u32; 8]) -> Self {
Self::with_shape(key, Shape::V2)
}
pub fn with_shape(key: [u32; 8], shape: Shape) -> Self {
let mut rng = SplitMix64::new(key[0] as u64 | ((key[1] as u64) << 32));
let mut rot = [0u32; 8];
let mut mul = [0u32; 16];
@ -90,7 +190,7 @@ impl MixParams {
for c in rc.iter_mut() {
*c = rng.next() as u32;
}
Self { key, rot, mul, rc }
Self { key, rot, mul, rc, shape }
}
/// Parameters for a day string: the key is `seed_words("day/" + day)`.
pub fn for_day(day: &str) -> Self {
@ -104,6 +204,13 @@ pub fn round_key(r: usize) -> u32 {
((r + 1) as u32).wrapping_mul(0x9E3779B9)
}
/// The round key of application `j` (0 <= j < m) of round `r` under multiplier `m`: `round_key(r * m + j)`. For
/// `m = 1` this is `round_key(r)`, version 2's key.
#[inline(always)]
pub fn round_key_mult(r: usize, j: usize, m: usize) -> u32 {
round_key(r * m + j)
}
/// `M_r` on 16 words in place: per word `(s ^ (RC + rk)) * MUL`, then one ChaCha-shaped double round with
/// the four column rotations `ROT[0..3]` and the four diagonal rotations `ROT[4..7]`.
#[inline(always)]
@ -122,9 +229,11 @@ pub fn mixer(s: &mut [u32; 16], rk: u32, mp: &MixParams) {
qr(s, 3, 4, 9, 14, r[4], r[5], r[6], r[7]);
}
/// The 256 MiB cache for one day key.
/// The cache for one day key: 2^log2_words words (256 MiB under version 2).
pub struct Cache {
pub key: [u32; 8],
pub log2_words: u32,
line_mask: u32,
words: Vec<u32>,
}
@ -152,13 +261,22 @@ impl Cache {
}
}
/// The whole cache on the calling thread: 65,536 chains of 64 ChaCha12 blocks, in segment order.
/// The version 2 cache on the calling thread: 65,536 chains of 64 ChaCha12 blocks, in segment order.
pub fn fill(key: [u32; 8]) -> Cache {
let mut words = vec![0u32; CACHE_WORDS];
for seg in 0..CACHE_SEGMENTS {
Self::fill_log2(key, CACHE_LOG2_WORDS as u32)
}
/// A cache of 2^`log2_words` words (26, 27 or 28 under the growth rule; smaller sizes for tests): 2^(log2 - 10)
/// independent chains of 64 lines, the same chain function at every size, so a larger cache's first segments
/// are the smaller cache's segments word for word.
pub fn fill_log2(key: [u32; 8], log2_words: u32) -> Cache {
assert!((10..=30).contains(&log2_words), "cache log2 words must be in 10..=30");
let shape = Shape { mixer_mult: 1, cache_log2_words: log2_words };
let mut words = vec![0u32; shape.cache_words()];
for seg in 0..shape.cache_segments() {
Self::fill_segment(&mut words, seg, &key);
}
Cache { key, words }
Cache { key, log2_words, line_mask: shape.cache_line_mask(), words }
}
pub fn for_day(day: &str) -> Cache {
@ -170,10 +288,20 @@ impl Cache {
&self.words
}
/// Cache line `a` (0 <= a < 2^22) as 16 words.
pub fn lines(&self) -> usize {
self.words.len() >> 4
}
pub fn line_mask(&self) -> u32 {
self.line_mask
}
pub fn segments(&self) -> usize {
self.lines() >> CACHE_SEGMENT_LOG2_LINES
}
/// Cache line `a` (masked to the cache's lines) as 16 words.
#[inline(always)]
pub fn line(&self, a: u32) -> &[u32] {
let o = (a & CACHE_LINE_MASK) as usize * 16;
let o = (a & self.line_mask) as usize * 16;
&self.words[o..o + 16]
}
@ -184,10 +312,13 @@ impl Cache {
}
/// Derive `ts.len()` items into `out`, all chains interleaved round by round so the cache-line misses of
/// independent items overlap in the memory system (`deriveItems` in the Swift).
/// independent items overlap in the memory system (`deriveItems` in the Swift). Under multiplier `m`
/// (`mp.shape.mixer_mult`) round `r` applies `M` with keys `round_key(r m + j)` for `j = 0 .. m - 1` before its
/// one cache read; the final mixer applies `M` with keys `round_key(8 m + j)`. `m = 1` is version 2.
pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; 16]]) {
let n = ts.len();
debug_assert!(out.len() >= n);
let m = mp.shape.mixer_mult as usize;
for k in 0..n {
let s = &mut out[k];
let t = ts[k];
@ -197,9 +328,11 @@ pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32;
}
}
for r in 0..ITEM_ROUNDS {
let rk = round_key(r);
for s in out[..n].iter_mut() {
mixer(s, rk, mp);
for j in 0..m {
let rk = round_key_mult(r, j, m);
for s in out[..n].iter_mut() {
mixer(s, rk, mp);
}
}
for s in out[..n].iter_mut() {
let line = cache.line(s[0]);
@ -208,9 +341,11 @@ pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32;
}
}
}
let rk = round_key(ITEM_ROUNDS);
for s in out[..n].iter_mut() {
mixer(s, rk, mp);
for j in 0..m {
let rk = round_key_mult(ITEM_ROUNDS, j, m);
for s in out[..n].iter_mut() {
mixer(s, rk, mp);
}
}
}
@ -221,7 +356,7 @@ pub fn derive_item(t: u32, mp: &MixParams, cache: &Cache) -> [u32; 16] {
out[0]
}
/// The CPU verifier's view of the memory-hard dataset: the mixer parameters and the 256 MiB cache.
/// The CPU verifier's view of the memory-hard dataset: the mixer parameters (with the shape) and the cache.
pub struct MemhardCpu {
pub params: MixParams,
pub cache: Cache,
@ -231,12 +366,19 @@ pub struct MemhardCpu {
pub const FETCH_MAX: usize = 64;
impl MemhardCpu {
/// Version 2 shape.
pub fn new(key: [u32; 8]) -> Self {
Self { params: MixParams::new(key), cache: Cache::fill(key) }
Self::with_shape(key, Shape::V2)
}
pub fn with_shape(key: [u32; 8], shape: Shape) -> Self {
Self { params: MixParams::with_shape(key, shape), cache: Cache::fill_log2(key, shape.cache_log2_words) }
}
pub fn for_day(day: &str) -> Self {
Self::new(day_key(day))
}
pub fn shape(&self) -> Shape {
self.params.shape
}
/// `dataset[w] = item(w >> 4)[w & 15]`.
pub fn word(&self, w: u32) -> u32 {
derive_item(w >> 4, &self.params, &self.cache)[(w & 15) as usize]
@ -313,6 +455,7 @@ mod tests {
assert_eq!(mp.rc[0], 0xbab68293);
assert_eq!(mp.rc[15], 0x31b49ee2);
assert!(mp.mul.iter().all(|m| m & 1 == 1));
assert_eq!(mp.shape, Shape::V2);
}
#[test]
@ -338,4 +481,96 @@ mod tests {
]
);
}
/// Option C: the schedule table of `docs/plans/mixer-x4.md` (day -> doublings, cache words, dataset words at a
/// 2^28 genesis). The doublings fall at years 4 and 12 exactly, never a day early.
#[test]
fn growth_schedule_table() {
let table: [(u64, u32, u32, u32); 12] = [
(0, 0, 26, 28),
(1, 0, 26, 28),
(365, 0, 26, 28),
(1_459, 0, 26, 28),
(1_460, 1, 27, 29),
(2_920, 1, 27, 29),
(4_379, 1, 27, 29),
(4_380, 2, 28, 30),
(10_219, 2, 28, 30),
(10_220, 3, 29, 31),
(21_900, 4, 30, 32),
(100_000, 6, 32, 32),
];
for (d, k, c, s) in table {
assert_eq!(growth_doublings(d), k, "day {d}");
assert_eq!(cache_log2_words(d), c, "day {d}");
assert_eq!(dataset_log2_words(28, d), s, "day {d}");
}
// the designed 2 GiB genesis: 2^29 words, 2^30 at year 4, 2^31 at year 12
assert_eq!(dataset_log2_words(29, 0), 29);
assert_eq!(dataset_log2_words(29, 1_460), 30);
assert_eq!(dataset_log2_words(29, 4_380), 31);
// the linear schedule itself: 2 GiB x (1 + d / 1460) crosses 4 GiB at day 1,460 and 8 GiB at day 4,380
for d in [1_459u64, 1_460, 4_379, 4_380] {
let bytes = 2u64 * (1 << 30) + (1u64 << 29) * d / 365;
let k = (bytes / (2u64 << 30)).ilog2();
assert_eq!(growth_doublings(d), k, "day {d}: linear {bytes} bytes");
}
assert_eq!(days_since_genesis(20_730, 20_729), 1);
assert_eq!(days_since_genesis(20_729, 20_729), 0);
assert_eq!(days_since_genesis(20_000, 20_729), 0);
let v2 = Shape::for_class_day(&LoadClass::V2, 100_000);
assert_eq!(v2, Shape::V2);
let v3 = Shape::for_class_day(&LoadClass::MX4, 0);
assert_eq!(v3, Shape { mixer_mult: 4, cache_log2_words: 26 });
assert_eq!(Shape::for_class_day(&LoadClass::MX4, 1_460).cache_log2_words, 27);
assert_eq!(v3.mixers_per_item(), 36);
assert_eq!(Shape::V2.mixers_per_item(), 9);
assert_eq!(Shape::V2.cache_segments(), CACHE_SEGMENTS);
assert_eq!(Shape::V2.cache_line_mask(), CACHE_LINE_MASK);
assert_eq!(Shape::V2.log2_segments(), 16);
}
/// The multiplied mixer, restated by hand on a small cache: `m` applications with keys `round_key(r m + j)`
/// before every read, the same 8 reads; `m = 1` is `derive_item` of version 2 word for word; a larger cache's
/// first segments equal the smaller cache's.
#[test]
fn mixer_mult_by_hand() {
let key = day_key("2026-10-03");
let small = Cache::fill_log2(key, 16);
let big = Cache::fill_log2(key, 18);
assert_eq!(&big.words()[..small.words().len()], small.words());
assert_eq!(small.segments(), 64);
assert_eq!(small.line_mask(), 4095);
for m in [1u32, 2, 4] {
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 16 });
for t in [0u32, 1, 12_345, u32::MAX] {
let got = derive_item(t, &mp, &small);
let mut s = [0u32; 16];
s[..8].copy_from_slice(&key);
for i in 0..8 {
s[8 + i] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]);
}
for r in 0..8usize {
for j in 0..m as usize {
mixer(&mut s, round_key(r * m as usize + j), &mp);
}
let line = small.line(s[0]);
for i in 0..16 {
s[i] ^= line[i];
}
}
for j in 0..m as usize {
mixer(&mut s, round_key(8 * m as usize + j), &mp);
}
assert_eq!(got, s, "m {m} t {t}");
}
}
let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16 });
let v3 = MixParams::with_shape(key, Shape { mixer_mult: 4, cache_log2_words: 16 });
assert_ne!(derive_item(0, &v2, &small), derive_item(0, &v3, &small));
assert_eq!(round_key_mult(0, 0, 1), round_key(0));
assert_eq!(round_key_mult(8, 0, 1), round_key(8));
assert_eq!(round_key_mult(2, 3, 4), round_key(11));
assert_eq!(round_key_mult(8, 3, 4), round_key(35));
}
}

View file

@ -2,7 +2,7 @@
//! calls. Dataset words come from the memory-hard cache (default) or from the closed form (old packs).
use crate::generator::{generate, generate_class, Instr, LoadClass, Op, Program, ProgramClass, ITERATIONS, LANES};
use crate::memhard::MemhardCpu;
use crate::memhard::{MemhardCpu, Shape};
use crate::seed::day_key;
/// Read-width experiment (5 October 2026): a `load` of `W` words folds every word into `dst`:
@ -150,21 +150,37 @@ pub struct DatasetSource {
impl DatasetSource {
/// Build the source for a day. Memory-hard mode fills the 256 MiB cache on the calling thread.
pub fn new(day: &str, mode: DatasetMode, log2_words: u32) -> Self {
let mut ds = Self::from_key(day_key(day), mode, log2_words);
Self::new_shape(day, mode, log2_words, Shape::V2)
}
/// [`DatasetSource::new`] with the construction's shape (mixer multiplier, cache size; Counter ASIC 2.0).
pub fn new_shape(day: &str, mode: DatasetMode, log2_words: u32, shape: Shape) -> Self {
let mut ds = Self::from_key_shape(day_key(day), mode, log2_words, shape);
ds.key_bytes = format!("day/{day}").into_bytes();
ds
}
pub fn from_key(key: [u32; 8], mode: DatasetMode, log2_words: u32) -> Self {
Self::from_key_shape(key, mode, log2_words, Shape::V2)
}
/// [`DatasetSource::from_key`] with the construction's shape. Memory-hard mode fills a cache of
/// `2^shape.cache_log2_words` words on the calling thread.
pub fn from_key_shape(key: [u32; 8], mode: DatasetMode, log2_words: u32, shape: Shape) -> Self {
assert!((4..=32).contains(&log2_words), "dataset log2 must be in 4..=32");
let mask = if log2_words == 32 { u32::MAX } else { (1u32 << log2_words) - 1 };
let dataset = match mode {
DatasetMode::ClosedForm => Dataset::ClosedForm { d0: key[0], d1: key[1] },
DatasetMode::MemoryHard => Dataset::MemoryHard(MemhardCpu::new(key)),
DatasetMode::MemoryHard => Dataset::MemoryHard(MemhardCpu::with_shape(key, shape)),
};
Self { log2_words, mask, key, key_bytes: Vec::new(), dataset }
}
/// The shape of the memory-hard construction ([`Shape::V2`] for the closed form, which has none).
pub fn shape(&self) -> Shape {
self.memhard().map(|m| m.shape()).unwrap_or(Shape::V2)
}
pub fn mode(&self) -> DatasetMode {
match self.dataset {
Dataset::ClosedForm { .. } => DatasetMode::ClosedForm,
@ -417,9 +433,10 @@ pub struct Epoch {
/// Default dataset size: 2^28 words = 1 GiB.
pub const DEFAULT_DATASET_LOG2: u32 = 28;
/// Days a day index lies after the network's genesis day (0 for the genesis day and any day before it).
/// Days a day index lies after the network's genesis day (0 for the genesis day and any day before it). The node's
/// entry; the same function as `memhard::days_since_genesis`.
pub fn days_since_genesis(day_index: u64, genesis_day_index: u64) -> u64 {
day_index.saturating_sub(genesis_day_index)
crate::memhard::days_since_genesis(day_index, genesis_day_index)
}
impl Epoch {
@ -427,9 +444,17 @@ impl Epoch {
Self { program: generate(seed), dataset: DatasetSource::new(day, mode, dataset_log2) }
}
/// [`Epoch::new`] with a load class (read-width experiment).
/// [`Epoch::new`] with a load class (read-width experiment; Counter ASIC 2.0: the class's mixer multiplier
/// shapes the dataset, the cache is the genesis size since a string day has no day index).
pub fn new_class(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32, class: LoadClass) -> Self {
Self { program: generate_class(seed, class), dataset: DatasetSource::new(day, mode, dataset_log2) }
Self::new_class_day(seed, day, mode, dataset_log2, class, 0)
}
/// [`Epoch::new_class`] on day `days_since_genesis` of the growth schedule (the cache of
/// `memhard::cache_log2_words` for a class with the growth rule; the dataset size is the caller's).
pub fn new_class_day(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32, class: LoadClass, days_since_genesis: u64) -> Self {
let shape = Shape::for_class_day(&class, days_since_genesis);
Self { program: generate_class(seed, class), dataset: DatasetSource::new_shape(day, mode, dataset_log2, shape) }
}
/// The production shape: memory-hard, 1 GiB dataset.
@ -445,11 +470,24 @@ impl Epoch {
Self::from_seed_bytes_class(epoch_seed, day_bytes, label, LoadClass::V2)
}
/// [`Epoch::from_seed_bytes`] with a load class (read-width experiment).
/// [`Epoch::from_seed_bytes`] with a load class (read-width experiment; Counter ASIC 2.0: the class's mixer
/// multiplier shapes the dataset). Day 0 of the growth schedule: the 2^26-word cache and the 2^28-word dataset,
/// which is every devnet pack and vector. A node past the first doubling calls [`Epoch::from_seed_bytes_day`].
pub fn from_seed_bytes_class(epoch_seed: &[u8], day_bytes: &[u8], label: &str, class: LoadClass) -> Self {
Self::from_seed_bytes_day(epoch_seed, day_bytes, label, class, 0, DEFAULT_DATASET_LOG2)
}
/// The chain's shape on day `days_since_genesis` (`memhard::days_since_genesis(day_index(header), day_index(genesis))`,
/// the node's two day indices): the program of the class, and under the class's growth rule the cache of
/// `memhard::cache_log2_words(d)` and the dataset of `memhard::dataset_log2_words(genesis_dataset_log2, d)`
/// (the genesis size is 28 for the 1 GiB devnet, 29 for the designed 2 GiB). Without the growth rule the cache
/// is 2^26 words and the dataset `2^genesis_dataset_log2` on every day.
pub fn from_seed_bytes_day(epoch_seed: &[u8], day_bytes: &[u8], label: &str, class: LoadClass, days_since_genesis: u64, genesis_dataset_log2: u32) -> Self {
let program = crate::generator::generate_from_seed_bytes_class(label, epoch_seed, class);
let key = crate::seed::seed_words_from_bytes(day_bytes);
let mut dataset = DatasetSource::from_key(key, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2);
let shape = Shape::for_class_day(&class, days_since_genesis);
let dataset_log2 = if class.growth { crate::memhard::dataset_log2_words(genesis_dataset_log2, days_since_genesis) } else { genesis_dataset_log2 };
let mut dataset = DatasetSource::from_key_shape(key, DatasetMode::MemoryHard, dataset_log2, shape);
dataset.key_bytes = day_bytes.to_vec();
Self { program, dataset }
}
@ -477,13 +515,18 @@ impl Epoch {
/// [`Epoch::chain_dataset`] with the day's position since genesis and the network's genesis dataset size: the
/// entry the node's engine and the miner's export build every day cache through, so the cache growth schedule
/// of spec 01 section 1.13.3 has one place to act. SEAM (ca2-node, 5 October 2026): the body here builds the
/// genesis-size cache for every day; the ca2-mixer branch fills the growth rule (`growth_doublings`) and the
/// class v3 item construction, keeping this signature.
/// of spec 01 section 1.13.3 has one place to act (ca2-mixer, 5 October 2026, `docs/plans/mixer-x4.md`): the
/// class's load class gives the mixer multiplier and whether the growth rule applies (`Shape::for_class_day`);
/// under the rule the cache is `2^memhard::cache_log2_words(d)` words and the dataset
/// `2^memhard::dataset_log2_words(genesis_dataset_log2, d)`; without it (class v2) the cache is 2^26 words and
/// the dataset the genesis size on every day. `days_since_genesis` is [`days_since_genesis`] of the block's and
/// the genesis header's day indices.
pub fn chain_dataset_day(day_bytes: &[u8], class: ProgramClass, days_since_genesis: u64, genesis_dataset_log2: u32) -> DatasetSource {
let _ = (class, days_since_genesis);
let lc = class.load_class();
let shape = Shape::for_class_day(&lc, days_since_genesis);
let dataset_log2 = if lc.growth { crate::memhard::dataset_log2_words(genesis_dataset_log2, days_since_genesis) } else { genesis_dataset_log2 };
let key = crate::seed::seed_words_from_bytes(day_bytes);
let mut dataset = DatasetSource::from_key(key, DatasetMode::MemoryHard, genesis_dataset_log2);
let mut dataset = DatasetSource::from_key_shape(key, DatasetMode::MemoryHard, dataset_log2, shape);
dataset.key_bytes = day_bytes.to_vec();
dataset
}

View file

@ -31,6 +31,9 @@ typedef struct {
char loadClass[64];
char programClass[8]; /* IGNEUM_PROGRAM_CLASS: "v2" or "v3" (Counter ASIC 2.0); absent = the generator's class */
char eraHex[65]; /* IGNEUM_ERA_SEED_HEX of a class v3 chain pack; empty otherwise */
// Counter ASIC 2.0 (5 October 2026): the mixer multiplier of the item derivation (IGNEUM_MIXER_MULT, 1 when absent:
// version 2; 4 under class v3). The emitted memhard.h / kernel.cl carry it in their text; this is for the log lines.
uint32_t mixerMult;
// seeds.txt (or program.h): the seeds as the worker protocol carries them
char epochHex[65];
char dayHex[PF_HEX_CAP];
@ -296,6 +299,7 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) {
pk->persistent = 0; pf_define_u32(prog, "IGNEUM_PERSISTENT_WARPS", &pk->persistent);
pk->scratchWordsPerLane = 8192; pf_define_u32(prog, "IGNEUM_SCRATCH_WORDS_PER_LANE", &pk->scratchWordsPerLane);
strcpy(pk->loadClass, "v2"); pf_define_str(prog, "IGNEUM_LOAD_CLASS", pk->loadClass, sizeof(pk->loadClass));
pk->mixerMult = 1; pf_define_u32(prog, "IGNEUM_MIXER_MULT", &pk->mixerMult);
if (pf_define_words(prog, "IGNEUM_KEY_INIT", pk->keyw, 8) != 8) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_KEY_INIT with 8 words"); }
if (!pf_define_str(prog, "IGNEUM_SEED_STRING", pk->seedString, sizeof(pk->seedString))) strncpy(pk->seedString, "(no IGNEUM_SEED_STRING)", sizeof(pk->seedString) - 1);
pf_define_str(prog, "IGNEUM_SEED_BYTES_HEX", ehex, sizeof(ehex));

View file

@ -61,6 +61,8 @@ let scratchOps = Int(defineU32("IGNEUM_SCRATCH_OPS") ?? 0)
let persistent = (defineU32("IGNEUM_PERSISTENT_WARPS") ?? 0) == 1
let scratchWordsPerLane = Int(defineU32("IGNEUM_SCRATCH_WORDS_PER_LANE") ?? 8192)
let className = defineStr("IGNEUM_LOAD_CLASS") ?? "v2"
// Counter ASIC 2.0 (5 October 2026): the mixer multiplier of the item derivation, 1 when absent (version 2), 4 under class v3
let mixerMult = Int(defineU32("IGNEUM_MIXER_MULT") ?? 1)
let seedString = defineStr("IGNEUM_SEED_STRING") ?? "?"
let programId = defineStr("IGNEUM_PROGRAM_ID") ?? ""
@ -200,7 +202,7 @@ for b in 0..<opts.batches {
let hashes = Double(nonces) * Double(opts.batches)
let mhsWall = hashes / wallSum / 1e3, mhsGpu = hashes / gpuSum / 1e3
let packName = (opts.pack as NSString).lastPathComponent
print("pack \(packName) seed \"\(seedString)\" id \(programId) class \(className): loads/hash \(loadsPerHash), dataset bytes/hash \(bytesPerHash), scratch ops/hash \(scratchOps * 8)")
print("pack \(packName) seed \"\(seedString)\" id \(programId) class \(className): loads/hash \(loadsPerHash), dataset bytes/hash \(bytesPerHash), scratch ops/hash \(scratchOps * 8), mixer x\(mixerMult), cache 2^\(cacheLog2) words")
print("device \(device.name); compile \(String(format: "%.0f", compileMs)) ms; cache fill \(String(format: "%.1f", cacheGpu)) ms GPU (\(String(format: "%.1f", cacheWall)) wall); dataset build \(String(format: "%.1f", buildGpu)) ms GPU (\(String(format: "%.1f", buildWall)) wall)")
print("cache FNV-1a 64 \(String(format: "%016llx", cacheFnv)) \(cacheOk ? "PASS" : "FAIL"); dataset head and last \(dsOk ? "PASS" : "FAIL"); vectors standalone \(vecPass)/\(vecBases.count), in batch \(batchVecPass)/\(batchVecN)")
print("warm-up batch \(nonces) hashes: \(String(format: "%.1f", warmGpu)) ms GPU, \(String(format: "%.1f", warmWall)) ms wall")

View file

@ -2092,7 +2092,11 @@ int main(int argc, char** argv) {
printf("day \"%s\" (d0 0x%08x, d1 0x%08x), pack dataset 2^%d words = %d MiB\n",
IGNEUM_DAY_STRING, IGNEUM_DAY0, IGNEUM_DAY1, IGNEUM_DATASET_LOG2, packMib());
#if IGNEUM_DATASET_MODE == 1
printf("dataset construction: memory-hard (256 MiB ChaCha cache, %d dependent cache reads per 64-byte item; proto-metal/MEMHARD.md)\n", IGNEUM_ITEM_ROUNDS);
#ifndef IGNEUM_MIXER_MULT
#define IGNEUM_MIXER_MULT 1
#endif
printf("dataset construction: memory-hard (%u MiB ChaCha cache, %d dependent cache reads per 64-byte item, mixer x%d; proto-metal/MEMHARD.md)\n",
(unsigned)(((uint64_t)CACHE_WORDS_HOST * 4u) >> 20), IGNEUM_ITEM_ROUNDS, IGNEUM_MIXER_MULT);
cachePass = setupCache(&dv, di);
#else
printf("dataset construction: closed-form ds_elem (the original prototype dataset, not memory-hard)\n");

View file

@ -0,0 +1,63 @@
# Igneum run job (bench only): the class v3 dataset construction (mixer x4, docs/plans/mixer-x4.md) on PC 2's RTX 5090 (machine 1ccfe586), 5 October 2026.
# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches
# off ONLY the NVIDIA card in the app through POST <app.url>api/cards, waits for its worker to stop, runs the fetched
# igneum-worker-cuda.exe on the v2 pack and the two v3 packs of the fetched packs folder (the dataset build time per pack is the
# number this job is for: the worker's own `cache ... dataset ... ms` line, the same line the readwidth round printed),
# and switches the card back on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs <job id>` shows them.
# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis and mx4-devnet-epoch0 from branch ca2-mixer, plus
# proto-cuda/packs/igneum-genesis-mh copied in as v2-genesis-mh (the version 2 control, same card, same run).
$ErrorActionPreference = 'Continue'
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
$jobs = Split-Path $env:IGNEUM_JOB_DIR
$fetched = Join-Path $jobs 'fetch-mixer-x4-20261005'
$exe = Join-Path $fetched 'igneum-worker-cuda.exe'
$packs = Join-Path $fetched 'packs-ca2-mixer'
if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 }
if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 }
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe (the NVRTC DLLs come from there)'; exit 2 }
Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination $fetched -Force
Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower()) with $((Get-ChildItem $fetched -Filter 'nvrtc*.dll').Count) NVRTC DLL(s) from $inst"
# the app: switch off the NVIDIA card only, remember its settings
$appDir = $env:IGNEUM_APP_DIR
if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
$urlFile = Join-Path $appDir 'app.url'
$url = $null
if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
$card = $null
if ($url) {
try {
$st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10
$card = $st.mining.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1
if (-not $card) { $card = $st.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 }
} catch { Say ("api/state: " + $_.Exception.Message) }
}
if ($card) {
Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state)
$body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) }
$t = 0
while ($t -lt 90) {
Start-Sleep -Seconds 5; $t += 5
try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { }
}
Write-Output ("RESULT card-off after " + $t + " s")
Start-Sleep -Seconds 5
} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' }
& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" }
# the probe ran in run-readwidth-5090-20261005 (bench-log, 5 October 2026); this second run is the bench only
foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0')) {
$d = Join-Path $packs $pk
Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)"
& $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" }
& $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT $_" }
}
& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" }
if ($card) {
$body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) }
}
exit 0

View file

@ -0,0 +1,64 @@
# Igneum run job (bench only): the class v3 dataset construction (mixer x4, docs/plans/mixer-x4.md) on PC 1's RX 9070 XT on the eGPU (machine ae432dc7), 5 October 2026.
# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches
# off ONLY the NVIDIA card in the app through POST <app.url>api/cards, waits for its worker to stop, runs the fetched
# igneum-worker-opencl.exe on the v2 pack and the two v3 packs of the fetched packs folder (the dataset build time per pack is the
# number this job is for: the worker's own `cache ... dataset ... ms` line, the same line the readwidth round printed),
# and switches the card back on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs <job id>` shows them.
# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis and mx4-devnet-epoch0 from branch ca2-mixer, plus
# proto-cuda/packs/igneum-genesis-mh copied in as v2-genesis-mh (the version 2 control, same card, same run).
$ErrorActionPreference = 'Continue'
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
$jobs = Split-Path $env:IGNEUM_JOB_DIR
$fetched = Join-Path $jobs 'fetch-mixer-x4-20261005'
$exe = Join-Path $fetched 'igneum-worker-opencl.exe'
$packs = Join-Path $fetched 'packs-ca2-mixer'
if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 }
if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 }
Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower())"
# the card's OpenCL device index on the current (3683.0) platform, from the worker's own list (the older platform's duplicate is marked dup)
$list = & $exe --list 2>&1
$list | ForEach-Object { "RESULT list $_" }
$dev = $null
foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1201' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; break } }
if ($null -eq $dev) { Write-Output 'RESULT error no gfx1201 device in --list'; exit 2 }
Write-Output "RESULT device $dev"
# the app: switch off the NVIDIA card only, remember its settings
$appDir = $env:IGNEUM_APP_DIR
if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
$urlFile = Join-Path $appDir 'app.url'
$url = $null
if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
$card = $null
if ($url) {
try {
$st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10
$card = $st.mining.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1
if (-not $card) { $card = $st.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1 }
} catch { Say ("api/state: " + $_.Exception.Message) }
}
if ($card) {
Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state)
$body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) }
$t = 0
while ($t -lt 90) {
Start-Sleep -Seconds 5; $t += 5
try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { }
}
Write-Output ("RESULT card-off after " + $t + " s")
Start-Sleep -Seconds 5
} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' }
# the probe ran in run-readwidth-9070-20261005 (bench-log, 5 October 2026); this second run is the bench only
foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0')) {
$d = Join-Path $packs $pk
Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)"
& $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" }
}
if ($card) {
$body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) }
}
exit 0