diff --git a/docs/analysis/chip-model-v3.md b/docs/analysis/chip-model-v3.md new file mode 100644 index 000000000..8a48ca7af --- /dev/null +++ b/docs/analysis/chip-model-v3.md @@ -0,0 +1,86 @@ +# The on-die-cache recompute chip against the RTX 5090, class v2 and class v3, everything combined + +5 October 2026 (night), Counter ASIC 2.0, worker ca2-mixer. The model is M16's +(`docs/analysis/m16-recompute-attacker-2026-10-05.md`): the strongest chip the plan has priced holds the whole +cache in SRAM and derives every dataset item instead of reading it, so its cost per hash is item derivations, +and its rate at a 50 T op/s integer budget (an RTX 5090's, approximate) is `50 T / (ops per hash)`. Nothing here +is a measurement of a chip; every GPU figure says where it was measured. "Approximate" marks a figure from memory. + +## 1. Inputs + +| Input | Value | Source | +|---|---|---| +| Items per hash | 128 (one item per load, 128 loads per hash, median 128.00 distinct) | spec 01 sections 1.4.2 and 1.8.5; the 20,000-program census | +| Integer operations per mixer application | about 130 | spec 01 section 1.8.4 | +| Mixer applications per item | 9 under v2; 36 under v3 (`m = 4`, `docs/plans/mixer-x4.md`) | `memhard::Shape::mixers_per_item` | +| Integer operations per item | 1,170 (v2); 4,680 (v3) | 9 x 130; 36 x 130 | +| Integer operations per hash | 149,760 (v2, "150,000"); 599,040 (v3, "600,000") | 128 x the above | +| Chip integer budget | 50 T op/s (approximate: 21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock) | M16 section 3 | +| Fixed-function factor | 3x (approximate, from memory: 2x to 5x is the usual credit for a pipeline with no scheduling or divergence) | M16 section 3 | +| RTX 5090, version 2 programs, measured | 136.1 MH/s (readwidth, tonight, `docs/plans/read-width.md`, pack w4 on PC 2); 139.7 MH/s (M11, 4 October, `docs/bench-log.md`) | this analysis uses tonight's 136.1 as the denominator and quotes both | +| RTX 5090 at w16 (16-byte loads), measured | 139.8 MH/s | readwidth table, tonight (the width stays 4 B: w16 closes nothing) | +| Cache mirror, 256 MiB, N5 headline density | 128 mm^2, $46 per good die (64 mm^2, $21 at the bit-cell lower bound) | `docs/analysis/sram-mirror.md` revision 2, sections 4 and 5 (`ca2-analysis` e6085c6) | +| Cache mirror plus a 96 MB hot table, N5 headline | 175 mm^2, $68 | same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate) | +| 512 MiB and 1 GiB mirrors, N5 headline | 255 mm^2 and 510 mm^2; $111 to $306 | same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node) | +| GPU-class die | 750 mm^2 (the equal-silicon comparison) | M16 section 3 | + +## 2. The rows + +Chip rate = 50 T op/s / ops per hash. "Bare" = chip rate / 136.1 MH/s. "With the factor" = bare x 3. "Equal +silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does not get, the M16 convention +("minus the area the SRAM takes"). SRAM in mm^2 and dollars at the N5 headline density. + +| Row | Mixer | Ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | With the 3x factor | Equal silicon, SRAM deducted, with the factor | +|---|---|---|---|---|---|---|---|---| +| v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x | +| v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x | +| v3: mixer x4 | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x | +| v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads) | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.61x or below (owed: the 5090's added-form rate; the hot loads cost it something, the chip nothing) | 1.84x or below | 1.49x | +| v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.61x or below (owed) | 1.84x or below | 1.45x | +| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 576 MiB | 287 / $130 | 0.61x | 1.84x | 1.14x | +| v3 at year 12 (cache 1 GiB, dataset 8 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 1,088 MiB | 542 / $330 | 0.61x | 1.84x | 0.51x | +| v3 with the mixer at x8 instead (the next lever, not adopted) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x | + +The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op +weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The +width rule (4-byte loads kept) changes nothing either: w16 would have moved the honest denominator by 2.7% and the +chip's cost not at all. + +Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6 +hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53. +Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of `sram-mirror.md` +scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. Year 4 and 12 rows: the mirror of +`sram-mirror.md` section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of +those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area, +not a forecast). + +## 3. The margin, plainly + +The combined headline row reads 1.84x with the 3x factor at an equal integer budget, 1.5x with the SRAM area +deducted. The claim is "under 2x", and the margin is thin: + +- the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x; +- the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart); +- the 50 T op/s budget is approximate; a chip at 55 T op/s reads 2.0x; +- the hot table in the added form lowers the honest denominator by whatever the hot loads cost the GPU (owed from + the PC rows), which raises the chip's gain by the same share, 1.84x or more if the hot loads are free, higher if + not; the hot table's only cost to this chip is 16 to 32 mm^2 of die. + +What keeps it under 2x is the mixer, and nothing else in Counter ASIC 2.0 moves this chip (the scratch at any share +gave 2.4x, `docs/analysis/scratch-soundness.md` section 3.4; the hot table taxes the DRAM-only chip, not this one; +the cache growth taxes it only in die area, which is cheap at year 0 and real at year 12). The next levers, in +order: + +1. Mixer x8: 0.18x bare and 0.55x with the factor in the M16 table (0.31x and 0.92x against 136.1 here; the M16 + table's denominator is 229 MH/s); the CPU verifier at 3.3 to 9.6 ms per warp scaled from the version 1 range, + at the edge of the 10 ms gate; the measured v3 row of `docs/plans/mixer-x4.md` section 6 is what to scale from + now, and whether a 2019-class laptop core (unmeasured, O-1.14) passes 10 ms is what decides it. +2. The hot table: adopted or not on the PC rows (`docs/plans/hot-table.md`); in the added form it costs the GPU + nothing it was not already paying in cache misses and the chip die area only, so it is the second lever for the + chip only through area and the first against a DRAM-only chip. + +## 4. What this does not settle + +The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point +under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had +no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM. diff --git a/docs/plans/mixer-x4.md b/docs/plans/mixer-x4.md new file mode 100644 index 000000000..2635da058 --- /dev/null +++ b/docs/plans/mixer-x4.md @@ -0,0 +1,151 @@ +# Mixer x4 and the cache growth rule: the class v3 dataset construction + +5 October 2026 (night). Counter ASIC 2.0, layer 6 (option C) and ledger M16's lever, decided by the coordinator +under the project lead's delegation at 22:00 UTC (`docs/plans/counter-asic-2-status.md`, "22:00 decided"; the project lead confirms for +the public testnet genesis). Branch `ca2-mixer`. Worker: ca2-mixer (cryptographer's lane). + +What this changes, in one line: under program class v3 every mixer application of the dataset item derivation +becomes four applications with distinct round keys, the eight dependent cache reads per item stay eight, and the +256 MiB cache doubles on the days the dataset doubles (years 4 and 12). Version 2 is byte-identical: the two +pinned packs re-export without a changed byte (section 5). + +## 1. Why this form + +The recompute attacker of `docs/analysis/m16-recompute-attacker-2026-10-05.md` holds the 256 MiB cache on a die +and derives every dataset word instead of reading it: 128 items per hash at about 1,170 integer operations and 8 +dependent cache reads each. Its cost is linear in operations per item; the honest miner pays the mixer once a day +in the dataset build and never per hash; the verifier pays it per item it checks. The multiplier `m` is the one +parameter that moves the attacker and leaves the honest hash rate untouched. + +Two shapes give the attacker 4x the operations: + +| Shape | Mixer applications per item | Dependent cache reads per item | What grows for the verifier | What grows for the chip | What grows for the honest build | +|---|---|---|---|---|---| +| A, chosen: `m = 4` applications per round, 8 rounds | 36 | 8 | the ALU part only; the latency part (8 dependent misses per item) unchanged | integer operations 4x; SRAM bandwidth unchanged (1,024 reads per hash) | 4x the mixer arithmetic, same reads | +| B, alternative: 32 rounds of one application and one read | 33 | 32 | both parts: 4x the dependent misses per item, so about 4x the latency-bound time (spec 1.11: about 8 x 100 ns per item in series without interleaving) | integer operations 3.7x and SRAM bandwidth 4x (4,096 reads per hash, 60 TB/s to match one 5090 at the M16 rate) | 4x the reads too; the GPU build becomes latency-bound at 4x the dependent line fetches | + +Shape A is chosen because the verifier's latency part is the part the 10 ms gate protects (section 1.11: the +distinct items of a unit are derived with their chains interleaved so the 8 misses of each item overlap across up +to 32 items; shape B would make that 32 misses deep). Shape B is implemented nowhere; the rule in the brief: if +shape A's measured verifier time exceeds 4.8 ms per warp on one M5 Max core, measure both and recommend. Section 6 +has the measurement; it is under that bound, so B stays unimplemented. + +## 2. Spec text (replaces 1.8.5 and 1.13.3 under class v3; v2 text unchanged) + +### 1.8.5 Item derivation and dataset mapping + +Item `t` (16 words) under mixer multiplier `m` (`m = 1` for program class v2, `m = 4` for class v3; a class +parameter, `LoadClass::mixer_mult`): + +``` +s[i] = K[i] for i in 0..7 +s[8 + i] = t * MUL[i] + RC[i] for i in 0..7 +for r in 0..7: + for j in 0..m-1: + s = M(s, rk = (r * m + j + 1) * 0x9E3779B9) + a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache + s[i] = s[i] XOR cache[line a][i] for i in 0..15 +for j in 0..m-1: + s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9) +item(t) = s +``` + +`M(s, rk)` is the mixer of 1.8.4 with round key `rk`; under `m = 1` the keys are `(r + 1) * 0x9E3779B9` and +`9 * 0x9E3779B9`, the version 2 text exactly. The round keys of the `9 m` applications are the first `9 m` values +of the version 2 key sequence, all distinct (the sequence is `k * 0x9E3779B9` for `k = 1 .. 9 m`, and +`0x9E3779B9` is odd, so no two of the first 2^32 keys coincide). Eight dependent cache reads per item at every +`m` (`ITEM_ROUNDS = 8`, prototype value): the address of read `r` depends on every earlier read. `9 m` mixer +applications, 36 under class v3, about 4,700 integer operations per item (130 per application, 1.8.4). +`dataset[w] = item(w >> 4)[w AND 15]`. A dataset of 2^D words is the prefix of items `0 .. 2^(D-4) - 1`, so an +item has the same value at every dataset size; and a cache of 2^C words is the prefix of segments of every +larger cache (1.8.3 fills segments independently of the cache size), but an item's value depends on `C` through +the line mask, so the item changes on the day the cache doubles. + +Source: `igneum-pow/src/memhard.rs` (`derive_items`, `round_key_mult`, `Shape`), the emitted `mh_item` of +`memhard.h`, `memhard.metal` and `kernel.cl` (`emit.rs`, `emit_memhard_core`: the `m` loop is emitted only for +`m > 1`, so every version 2 pack keeps its text). + +### 1.13.3 Dataset growth (class v3: option (b) with the cache tied to it, "option C") + +Designed: 2 GiB at genesis plus 0.5 GiB per year. The linear schedule in bytes, `G x (1 + 86,400 d / +(4 x 31,536,000)) = G x (1 + d / 1,460)` for the genesis size `G` and the chain day `d` (DAA days since genesis, +section 1.12), doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at day 10,220 +(year 28). Rule (Designed, decided 5 October 2026 for class v3): + +``` +doublings(d) = floor(log2(1 + d / 1460)) integer division, then integer log2 +dataset_words(d) = 2^(D_0 + doublings(d)) D_0 = 29 designed (2 GiB), 28 on the devnet (1 GiB); capped at 32 +cache_words(d) = 2^(26 + doublings(d)) 256 MiB, 512 MiB from year 4, 1 GiB from year 12 +``` + +Power-of-two sizes only (option (b)), so every load keeps the `src AND MASK` form of 1.14 item 2 and the cache +line index keeps `s[0] AND mask`. The cache doubles exactly when the dataset doubles ("option C", +`docs/analysis/sram-mirror.md` section 7): the cache's job is to stay above any GPU's last-level cache and that +needs growth; the recompute attacker is priced by the mixer, not by the cache (section 7 below). + +`d` is `day_index(header.timestamp) - day_index(genesis.timestamp)` with `day_index = timestamp_ms / 86,400,000` +(`bind::day_index`, the interim day rule), clamped at 0 (`memhard::days_since_genesis`). Under class v2 nothing +grows: the cache is 2^26 words and the dataset the genesis size on every day. + +| Chain day `d` | Years | `doublings` | Cache words | Cache | Dataset words (devnet `D_0 = 28`) | Dataset (designed `D_0 = 29`) | Verifier cache fill, one M5 Max core (measured at 256 MiB, section 6, scaled linearly) | +|---|---|---|---|---|---|---|---| +| 0 to 1,459 | 0 to 4 | 0 | 2^26 | 256 MiB | 2^28 (1 GiB) | 2 GiB | 0.18 s | +| 1,460 to 4,379 | 4 to 12 | 1 | 2^27 | 512 MiB | 2^29 (2 GiB) | 4 GiB | 0.36 s | +| 4,380 to 10,219 | 12 to 28 | 2 | 2^28 | 1 GiB | 2^30 (4 GiB) | 8 GiB | 0.72 s | +| 10,220 to 21,899 | 28 to 60 | 3 | 2^29 | 2 GiB | 2^31 (8 GiB) | 16 GiB | 1.4 s | +| 21,900 and on | 60 and on | 4 | 2^30 | 4 GiB | 2^32 (16 GiB, the index cap) | 2^32 words, the cap | 2.9 s | + +Test: `memhard::tests::growth_schedule_table` pins every row and the day before each step. The devnet pack +`igneum-devnet-v4-epoch0` is day 20,730 of the Unix count against genesis day 20,729, `d = 1`, so every existing +size and vector stands. + +Consequences for the tiers (the rule of 5 October): a verifier (any node, any pool core) holds 512 MiB from year 4 +and 1 GiB from year 12, and fills it once a day in under a second on one 2026 core (the table); a miner's card +holds the dataset, 4 GiB from year 4 and 8 GiB from year 12 on the designed schedule, so an 8 GB card mines until +year 12 and a 16 GB card until year 28 (the cache is not in the card's working set at hash time: it is built, +the dataset built from it, and dropped). Those dates are the design document's own schedule restated as steps; +option (a) would have faded a 4 GiB card out in year 4 instead of year 4. + +## 3. Interfaces + +| Item | Where | Note | +|---|---|---| +| `LoadClass { mixer_mult: u8, growth: bool }`, `LoadClass::MX4` ("mx4"), `with_mixer(m, growth)`, `v2_loads()`, `takes_width_roll()` | `generator.rs` | a class with v2 loads takes no width roll: its program stream is version 2's draw for draw, so the v3 program of a seed is the v2 program of that seed, only the dataset differs | +| `Shape { mixer_mult, cache_log2_words }`, `Shape::for_class_day(class, d)`, `MixParams.shape`, `Cache::fill_log2(key, log2)`, `round_key_mult(r, j, m)` | `memhard.rs` | `Shape::V2` is version 2 | +| `growth_doublings(d)`, `cache_log2_words(d)`, `dataset_log2_words(D_0, d)`, `days_since_genesis(day, genesis_day)` | `memhard.rs` | the schedule, one function and its two sizes | +| `DatasetSource::{new_shape, from_key_shape, shape}`, `Epoch::new_class_day`, `Epoch::from_seed_bytes_day(epoch, day, label, class, d, D_0)` | `verify.rs` | the day-sized entries; the v2 entries are unchanged and build the v2 shape | +| `IGNEUM_MIXER_MULT`, `IGNEUM_CLASS_MIXER_MULT`, `IGNEUM_CACHE_GROWTH` in program.h; `"mixer_mult"`, `"cache_growth"`, the `"item"` string in program.json | `emit.rs` | written only for a class with `m != 1` or growth, so v2 packs do not change | +| `packfile.h` `mixerMult`; `packbench` and the OpenCL host print the multiplier and the cache size | the three hosts | the kernels carry the construction in their text (one emitter, three dialects); the hosts size the cache from `IGNEUM_CACHE_LOG2_WORDS` already (packbench.swift line 56, host.cu line 65, host.c line 1966) | +| `igneum-pow --class mx4 [--days d]` on every command | `main.rs` | `--days` sizes the cache for a growth class | + +Under the ca2-v3 seam (`ProgramClass::V3`, `V3_CLASS`), the integration sets `V3_CLASS = LoadClass::MX4`; the +chain's day-sized dataset needs the day index, so `Epoch::chain_dataset(day, class)` builds the genesis-size +cache and a `chain_dataset_day(day_bytes, class, d, D_0)` beside it is the growth entry (section 9, owed to the +node agent). + +## 4. Vectors (class v3, `proto-cuda/packs-ca2-mixer/`) + +Filled in section 6 from the exported packs: `mx4-genesis` (seed `igneum-genesis`, day `2026-10-03`, 2^28 words, +2^26-word cache) and `mx4-devnet-epoch0` (the devnet genesis hash as the epoch seed, day bytes of 2026-10-04). + +## 5. The v2 path is byte-identical + +`cargo test --test packs` regenerates every file of `igneum-genesis-mh` and `igneum-devnet-v4-epoch0` from +`program.json` and compares byte for byte (`emitted_sources_match_all_packs`, `export_pack_matches_all_packs`); +section 6 also records a fresh `igneum-pow export` of both packs diffed against the checked-in directories. + +## 6. Measurements + +Filled as they land. Every row names the machine, the date, the command and the lock mode. + +## 7. The chip model + +`docs/analysis/chip-model-v3.md`. + +## 8. What is unverified + +Filled at the end. + +## 9. Owed + +Filled at the end. diff --git a/igneum-pow/src/emit.rs b/igneum-pow/src/emit.rs index a788c25a2..1b8024819 100644 --- a/igneum-pow/src/emit.rs +++ b/igneum-pow/src/emit.rs @@ -11,8 +11,8 @@ use crate::generator::{Instr, Op, Program, ProgramClass, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS}; use crate::memhard::{ - MixParams, CACHE_LINES_PER_SEGMENT, CACHE_LINE_MASK, CACHE_LOG2_WORDS, CACHE_SEGMENTS, CACHE_SEGMENT_LOG2_LINES, - CACHE_TAG, CACHE_WORDS, CHACHA_ROUNDS, CHACHA_SIGMA, ITEM_ROUNDS, + MixParams, Shape, CACHE_LINES_PER_SEGMENT, CACHE_SEGMENT_LOG2_LINES, CACHE_TAG, CHACHA_ROUNDS, CHACHA_SIGMA, + ITEM_ROUNDS, }; use crate::seed::SplitMix64; use crate::verify::{DatasetMode, DatasetSource, Epoch, FOLD_MUL, FOLD_ROT}; @@ -100,12 +100,25 @@ fn class_header_lines(p: &Program) -> String { return String::new(); } let mut s = String::new(); - s.push_str("// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads + if p.class.v2_loads() { + s.push_str("// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item "); - s.push_str("// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. + s.push_str("// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule. "); + } else { + s.push_str("// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +"); + s.push_str("// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +"); + } s.push_str(&format!("#define IGNEUM_LOAD_CLASS {} ", jstr(&p.class.name()))); + if p.class.mixer_mult != 1 || p.class.growth { + s.push_str(&format!("#define IGNEUM_CLASS_MIXER_MULT {} +", p.class.mixer_mult)); + s.push_str(&format!("#define IGNEUM_CACHE_GROWTH {} // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460)) +", p.class.growth as u8)); + } s.push_str(&format!("#define IGNEUM_LOAD_SLOTS {} ", p.class.load_slots)); s.push_str(&format!("#define IGNEUM_LOAD_MIX {{ {}, {}, {} }} @@ -257,12 +270,19 @@ pub enum LoadSource<'a> { InlineMemhard(&'a MixParams), } -fn log2_segments() -> usize { - CACHE_SEGMENTS.trailing_zeros() as usize +fn log2_segments(shape: &Shape) -> usize { + shape.log2_segments() as usize } -/// The memory-hard core as source text (`emitMemhardCore`). Every parameter is a literal. +/// The memory-hard core as source text (`emitMemhardCore`). Every parameter is a literal: the mixer constants of +/// the day and, from `mp.shape`, the cache size and the mixer multiplier `m`. Under `m = 1` the text is version +/// 2's byte for byte; under `m > 1` the item loop applies `mh_mixer` `m` times per round with the keys +/// `round_key(r m + j)` (spec 01 section 1.8.5 under class v3). pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String { + let shape = &mp.shape; + let m = shape.mixer_mult; + let cache_log2_words = shape.cache_log2_words; + let cache_line_mask = shape.cache_line_mask(); let (u, fn_, cptr, wptr, lptr, lcptr) = match dialect { CoreDialect::Metal => { ("uint", "inline", "device const uint*", "device uint*", "thread uint*", "const thread uint*") @@ -274,18 +294,23 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String { }; let k = &mp.key; let r = &mp.rot; - let m = &mp.mul; + let mul = &mp.mul; let c = &mp.rc; let mut s = String::with_capacity(6000); s.push_str(&format!( "// Memory-hard dataset core (MEMHARD.md). Cache: 2^{} words in 2^{} segments of {} chained ChaCha{} lines.\n", - CACHE_LOG2_WORDS, - log2_segments(), + cache_log2_words, + log2_segments(shape), CACHE_LINES_PER_SEGMENT, CHACHA_ROUNDS )); - s.push_str("// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.\n"); - s.push_str(&format!("#define MH_CACHE_LINE_MASK {}\n", hex(CACHE_LINE_MASK))); + if m == 1 { + s.push_str("// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.\n"); + } else { + s.push_str(&format!("// Item: 8 rounds of {m} x seed-parameterised mixer + one 64-byte cache read, then {m} x final mixer (class v3, mixer multiplier {m},\n")); + s.push_str("// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.\n"); + } + s.push_str(&format!("#define MH_CACHE_LINE_MASK {}\n", hex(cache_line_mask))); s.push_str(&format!("#define MH_SEGMENT_LINES {}u\n", CACHE_LINES_PER_SEGMENT)); s.push_str("#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }\n"); s.push_str(&format!( @@ -345,7 +370,7 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String { s.push_str("// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.\n"); s.push_str(&format!("{fn_} void mh_mixer({lptr} s, {u} rk) {{\n")); for i in 0..16 { - s.push_str(&format!(" s[{i}] = (s[{i}] ^ ({} + rk)) * {};\n", hex(c[i]), hex(m[i]))); + s.push_str(&format!(" s[{i}] = (s[{i}] ^ ({} + rk)) * {};\n", hex(c[i]), hex(mul[i]))); } let col = (0..4).map(|i| format!("{}u", r[i])).collect::>().join(", "); let dia = (4..8).map(|i| format!("{}u", r[i])).collect::>().join(", "); @@ -355,22 +380,39 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String { s.push_str(&format!(" MH_QR(s[2], s[7], s[8], s[13], {dia}) MH_QR(s[3], s[4], s[9], s[14], {dia})\n")); s.push_str("}\n"); s.push('\n'); - s.push_str(&format!( - "// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of mixer + cache line s[0] & mask; final mixer.\n" - )); + if m == 1 { + s.push_str(&format!( + "// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of mixer + cache line s[0] & mask; final mixer.\n" + )); + } else { + s.push_str(&format!( + "// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of {m} x mixer + cache line s[0] & mask; {m} x final mixer.\n" + )); + } s.push_str(&format!("{fn_} void mh_item({cptr} cache, {u} t, {lptr} s) {{\n")); for i in 0..8 { s.push_str(&format!(" s[{i}] = {};\n", hex(k[i]))); } for i in 0..8 { - s.push_str(&format!(" s[{}] = t * {} + {};\n", 8 + i, hex(m[i]), hex(c[i]))); + s.push_str(&format!(" s[{}] = t * {} + {};\n", 8 + i, hex(mul[i]), hex(c[i]))); } s.push_str(&format!(" for ({u} r = 0u; r < {ITEM_ROUNDS}u; ++r) {{\n")); - s.push_str(" mh_mixer(s, 0x9E3779B9u * (r + 1u));\n"); + if m == 1 { + s.push_str(" mh_mixer(s, 0x9E3779B9u * (r + 1u));\n"); + } else { + s.push_str(&format!(" for ({u} j = 0u; j < {m}u; ++j) mh_mixer(s, 0x9E3779B9u * (r * {m}u + j + 1u));\n")); + } s.push_str(&format!(" {cptr} line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);\n")); s.push_str(&format!(" for ({u} i = 0u; i < 16u; ++i) s[i] ^= line[i];\n")); s.push_str(" }\n"); - s.push_str(&format!(" mh_mixer(s, 0x9E3779B9u * {}u);\n", ITEM_ROUNDS + 1)); + if m == 1 { + s.push_str(&format!(" mh_mixer(s, 0x9E3779B9u * {}u);\n", ITEM_ROUNDS + 1)); + } else { + s.push_str(&format!( + " for ({u} j = 0u; j < {m}u; ++j) mh_mixer(s, 0x9E3779B9u * ({}u + j + 1u));\n", + ITEM_ROUNDS as u32 * m + )); + } s.push_str("}\n"); s.push_str("// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.\n"); s.push_str(&format!( @@ -386,7 +428,7 @@ pub fn metal_memhard(mp: &MixParams) -> String { s.push_str("using namespace metal;\n"); s.push_str(&emit_memhard_core(mp, CoreDialect::Metal)); s.push('\n'); - s.push_str(&format!("// One thread per segment (2^{} threads).\n", log2_segments())); + s.push_str(&format!("// One thread per segment (2^{} threads).\n", log2_segments(&mp.shape))); s.push_str( "kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {\n", ); @@ -1116,10 +1158,13 @@ pub fn program_header(p: &Program, day: &str, ds: &DatasetSource) -> String { s.push_str(&format!("#define IGNEUM_SEEDW_INIT {{ {} }}\n", join_hex(&p.seed))); if let Some(mp) = memhard { s.push_str(&format!("#define IGNEUM_KEY_INIT {{ {} }}\n", join_hex(&mp.key))); - s.push_str(&format!("#define IGNEUM_CACHE_LOG2_WORDS {CACHE_LOG2_WORDS}\n")); + s.push_str(&format!("#define IGNEUM_CACHE_LOG2_WORDS {}\n", mp.shape.cache_log2_words)); s.push_str(&format!("#define IGNEUM_CACHE_SEGMENT_LOG2_LINES {CACHE_SEGMENT_LOG2_LINES}\n")); - s.push_str(&format!("#define IGNEUM_CACHE_SEGMENTS {CACHE_SEGMENTS}u\n")); + s.push_str(&format!("#define IGNEUM_CACHE_SEGMENTS {}u\n", mp.shape.cache_segments())); s.push_str(&format!("#define IGNEUM_ITEM_ROUNDS {ITEM_ROUNDS}\n")); + if mp.shape.mixer_mult != 1 { + s.push_str(&format!("#define IGNEUM_MIXER_MULT {} // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)\n", mp.shape.mixer_mult)); + } s.push_str(&format!( "#define IGNEUM_MIX_ROT_INIT {{ {} }}\n", mp.rot.iter().map(|r| format!("{r}u")).collect::>().join(", ") @@ -1188,6 +1233,8 @@ pub struct PackVectors { pub cache_last: Vec, /// FNV-1a 64 over the whole cache (memory-hard only) pub cache_fnv: u64, + /// The cache is 2^cache_log2_words words (memory-hard only; 26 under version 2) + pub cache_log2_words: u32, } /// The base nonces of the three vector warps every pack carries. @@ -1249,7 +1296,8 @@ pub fn vectors_header( s.push_str("};\n"); if memhard { s.push_str(&format!( - "// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^{CACHE_LOG2_WORDS} words.\n" + "// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^{} words.\n", + v.cache_log2_words )); s.push_str("static const uint32_t IGNEUM_CACHE_HEAD[16] = {\n"); s.push_str(&format!(" {},\n", join_hex(&v.cache_head[..8]))); @@ -1300,6 +1348,11 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String { if !p.class.is_v2() { let c = p.width_counts(); s.push_str(&format!(" \"load_class\": {},\n", jstr(&p.class.name()))); + if p.class.mixer_mult != 1 || p.class.growth { + s.push_str(&format!(" \"mixer_mult\": {},\n", p.class.mixer_mult)); + s.push_str(&format!(" \"cache_growth\": {},\n", p.class.growth)); + s.push_str(&format!(" \"mixer\": \"class v3 (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): every mixer application of the item derivation is {} applications with round keys (r * {} + j + 1) * 0x9E3779B9, the 8 dependent cache reads per item unchanged; cache growth rule option C: cache words = 2^(26 + doublings(day)), dataset words = 2^(genesis_log2 + doublings(day)), doublings(day) = floor(log2(1 + day / 1460)) for day = days since genesis\",\n", p.class.mixer_mult, p.class.mixer_mult)); + } s.push_str(&format!(" \"load_slots\": {},\n", p.class.load_slots)); s.push_str(&format!(" \"load_mix_percent_4_16_64\": [{}, {}, {}],\n", p.class.mix[0], p.class.mix[1], p.class.mix[2])); s.push_str(&format!(" \"load_width_counts_4_16_64\": [{}, {}, {}],\n", c[0], c[1], c[2])); @@ -1351,9 +1404,12 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String { s.push_str(" \"spec\": \"proto-metal/MEMHARD.md\",\n"); s.push_str(&format!(" \"key\": [{}],\n", join_jhex(&mp.key))); s.push_str(" \"key_derivation\": \"the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]\",\n"); + let shape = &mp.shape; s.push_str(&format!( - " \"cache\": {{\"log2_words\": {CACHE_LOG2_WORDS}, \"bytes\": {}, \"line_words\": 16, \"segment_lines\": {CACHE_LINES_PER_SEGMENT}, \"segments\": {CACHE_SEGMENTS}, \"block\": \"ChaCha{CHACHA_ROUNDS} core + feed-forward, rotations 16 12 8 7\", \"sigma\": [{}], \"tag\": [{}], \"chain\": \"in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0\"}},\n", - CACHE_WORDS as u64 * 4, + " \"cache\": {{\"log2_words\": {}, \"bytes\": {}, \"line_words\": 16, \"segment_lines\": {CACHE_LINES_PER_SEGMENT}, \"segments\": {}, \"block\": \"ChaCha{CHACHA_ROUNDS} core + feed-forward, rotations 16 12 8 7\", \"sigma\": [{}], \"tag\": [{}], \"chain\": \"in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0\"}},\n", + shape.cache_log2_words, + shape.cache_words() as u64 * 4, + shape.cache_segments(), join_jhex(&CHACHA_SIGMA), join_jhex(&CACHE_TAG) )); @@ -1364,11 +1420,23 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String { join_jhex(&mp.rc) )); // The Swift writes jhex(cacheLineMask) here, which breaks the JSON. We write the bare literal. - s.push_str(&format!( - " \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = M_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = M_{ITEM_ROUNDS}(s); item(t) = s\",\n", - ITEM_ROUNDS - 1, - CACHE_LINE_MASK - )); + if shape.mixer_mult == 1 { + s.push_str(&format!( + " \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = M_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = M_{ITEM_ROUNDS}(s); item(t) = s\",\n", + ITEM_ROUNDS - 1, + shape.cache_line_mask() + )); + } else { + let m = shape.mixer_mult; + s.push_str(&format!( + " \"mixer_mult\": {m},\n \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: for j in 0..{}: s = M(s, rk = (r * {m} + j + 1) * 0x9E3779B9); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then for j in 0..{}: s = M(s, rk = ({} + j + 1) * 0x9E3779B9); item(t) = s\",\n", + ITEM_ROUNDS - 1, + m - 1, + shape.cache_line_mask(), + m - 1, + ITEM_ROUNDS as u32 * m + )); + } s.push_str(" \"word\": \"dataset[w] = item(w >> 4)[w & 15]\"\n"); } else { s.push_str(" \"mode\": \"closed-form\",\n"); @@ -1507,6 +1575,7 @@ pub fn export_pack(epoch: &Epoch, day: &str, source: &str) -> Pack { v.cache_head = w[..16].to_vec(); v.cache_last = w[w.len() - 16..].to_vec(); v.cache_fnv = m.cache.fnv1a64(); + v.cache_log2_words = m.shape().cache_log2_words; } let is_mh = memhard.is_some(); let mut files = vec![ diff --git a/igneum-pow/src/generator.rs b/igneum-pow/src/generator.rs index ce68973ec..4658b3c45 100644 --- a/igneum-pow/src/generator.rs +++ b/igneum-pow/src/generator.rs @@ -181,6 +181,13 @@ pub struct LoadClass { /// Variant 5: the scratch per warp in KiB (32 or 128; the whole working set of a card at full occupancy must /// stay under 6 GB, coordinator's cap of 5 October 2026). 0 for every other class. pub scratch_kb: u8, + /// Mixer cost multiplier `m` of the dataset item derivation (Counter ASIC 2.0, M16, decided 5 October 2026 for + /// class v3): every mixer application of spec 01 section 1.8.5 becomes `m` applications with distinct round + /// keys, the 8 dependent cache reads per item unchanged (`memhard::derive_items`). 1 for version 2, 4 for v3. + pub mixer_mult: u8, + /// Cache growth rule, option C (`memhard::growth_doublings`): the cache doubles when the dataset doubles. `false` + /// for version 2 (the cache is 2^26 words on every day), `true` for v3. + pub growth: bool, } /// Scratch geometry (variant 5): 16-byte slots, lane-major, 32 lanes per warp; `scratch_kb` KiB per warp gives @@ -205,20 +212,27 @@ impl LoadClass { impl LoadClass { /// Generator version 2 as adopted on 4 October 2026: 16 loads of one word. The lottery hash. - pub const V2: LoadClass = LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 }; + pub const V2: LoadClass = + LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 1, growth: false }; + + /// The construction decided for program class v3 on 5 October 2026 (Counter ASIC 2.0, `docs/plans/mixer-x4.md`): + /// version 2 loads (16 slots of one word, no scratch, no width roll, so the program stream is version 2's), the + /// mixer applied 4 times per round, and the cache growth rule. Name "mx4". + pub const MX4: LoadClass = + LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 4, growth: true }; /// A fixed width (1, 4 or 16 words) with `load_slots` loads per program. pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass { let mut mix = [0u8; 3]; let i = WIDTH_WORDS.iter().position(|&w| w == width_words).expect("width must be 1, 4 or 16 words"); mix[i] = 100; - LoadClass { mix, load_slots, scratch: None, scratch_kb: 0 } + LoadClass { mix, load_slots, ..LoadClass::V2 } } /// Per-load width drawn from `mix` (percent for 4, 16, 64 bytes), 16 loads per program. pub fn mixed(mix: [u8; 3]) -> LoadClass { assert_eq!(mix.iter().map(|&m| m as u32).sum::(), 100, "the mix must sum to 100"); - LoadClass { mix, load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 } + LoadClass { mix, ..LoadClass::V2 } } /// Variant 5: version 2 widths, 16 memory operations of which `k` are scratch read-modify-writes into a @@ -226,11 +240,61 @@ impl LoadClass { pub fn scratch(k: u8, kb: u8) -> LoadClass { assert!(k as usize <= LOAD_SLOTS); assert!(kb.is_power_of_two() && kb <= 128, "scratch per warp must be a power of two up to 128 KiB"); - LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: Some(k), scratch_kb: kb } + LoadClass { scratch: Some(k), scratch_kb: kb, ..LoadClass::V2 } } - /// Parse "p4,p16,p64" or one of the names of [`LoadClass::name`] ("scr4k32": 4 scratch ops, 32 KiB per warp). + /// This class with the mixer multiplier `m` (1, 2, 4, 8 or 16) and the cache growth rule on or off. + pub fn with_mixer(self, mixer_mult: u8, growth: bool) -> LoadClass { + assert!(mixer_mult >= 1 && mixer_mult <= 16 && mixer_mult.is_power_of_two(), "mixer multiplier must be 1, 2, 4, 8 or 16"); + LoadClass { mixer_mult, growth, ..self } + } + + /// The mixer multiplier as the item derivation uses it. + pub fn mixer_mult(&self) -> u32 { + self.mixer_mult as u32 + } + + /// Whether the loads of this class are version 2's: 16 one-word loads, no scratch. Such a class takes no width + /// roll, so its program stream is the version 2 stream draw for draw (the mixer and the cache are properties of + /// the dataset, not of the program). + pub fn v2_loads(&self) -> bool { + self.mix == [100, 0, 0] && self.load_slots as usize == LOAD_SLOTS && self.scratch.is_none() + } + + /// Whether every instruction takes the tenth draw (the width roll): every class whose loads are not version 2's. + pub fn takes_width_roll(&self) -> bool { + !self.v2_loads() + } + + /// Parse "p4,p16,p64" or one of the names of [`LoadClass::name`] ("scr4k32": 4 scratch ops, 32 KiB per warp; + /// "mx4": the v3 construction; a trailing "m" and "g" set the mixer multiplier and the growth rule on any + /// load class, "w16m4g" for example). pub fn parse(s: &str) -> Option { + if s == "mx4" { + return Some(LoadClass::MX4); + } + // the mixer suffix: "...m" then an optional "g" + let (s, growth) = match s.strip_suffix('g') { + Some(base) if base.rsplit_once('m').map(|(_, d)| !d.is_empty() && d.bytes().all(|b| b.is_ascii_digit())).unwrap_or(false) => (base, true), + _ => (s, false), + }; + if let Some((base, digits)) = s.rsplit_once('m') { + if !digits.is_empty() && digits.bytes().all(|b| b.is_ascii_digit()) && !base.is_empty() && !base.ends_with(',') { + let mult: u8 = digits.parse().ok()?; + if mult == 0 || mult > 16 || !mult.is_power_of_two() { + return None; + } + return Some(LoadClass::parse_loads(base)?.with_mixer(mult, growth)); + } + } + if growth { + return None; + } + LoadClass::parse_loads(s) + } + + /// The load part of a class name (no mixer suffix). + fn parse_loads(s: &str) -> Option { if let Some(rest) = s.strip_prefix("scr") { let (k, kb) = rest.split_once('k')?; let k: u8 = k.parse().ok()?; @@ -260,7 +324,7 @@ impl LoadClass { if slots == 0 || slots as usize >= INSTR_COUNT { return None; } - Some(LoadClass { mix, load_slots: slots, scratch: None, scratch_kb: 0 }) + Some(LoadClass { mix, load_slots: slots, ..LoadClass::V2 }) } /// Scratch read-modify-writes per program (0 without a scratch). @@ -272,25 +336,41 @@ impl LoadClass { *self == LoadClass::V2 } - /// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4". + /// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4k32"; "mx4" for the v3 construction; + /// any other mixer setting appends "m" and, with the growth rule, "g" ("v2m2", "w16m4g"). pub fn name(&self) -> String { if self.is_v2() { return "v2".to_string(); } - if let Some(k) = self.scratch { - return format!("scr{k}k{}", self.scratch_kb); + if *self == LoadClass::MX4 { + return "mx4".to_string(); } - let base = match self.mix { - [100, 0, 0] => "w4".to_string(), - [0, 100, 0] => "w16".to_string(), - [0, 0, 100] => "w64".to_string(), - [a, b, c] => format!("mix{a}-{b}-{c}"), - }; - if self.load_slots as usize == LOAD_SLOTS { - base + let loads = LoadClass { mixer_mult: 1, growth: false, ..*self }; + let base = if loads.is_v2() { + "v2".to_string() + } else if let Some(k) = self.scratch { + format!("scr{k}k{}", self.scratch_kb) } else { - format!("{base}x{}", self.load_slots) + let base = match self.mix { + [100, 0, 0] => "w4".to_string(), + [0, 100, 0] => "w16".to_string(), + [0, 0, 100] => "w64".to_string(), + [a, b, c] => format!("mix{a}-{b}-{c}"), + }; + if self.load_slots as usize == LOAD_SLOTS { + base + } else { + format!("{base}x{}", self.load_slots) + } + }; + let mut s = base; + if self.mixer_mult != 1 || self.growth { + s.push_str(&format!("m{}", self.mixer_mult)); } + if self.growth { + s.push('g'); + } + s } /// The width in words of a load whose width roll (0..99) is `roll`: the first entry of the mix whose cumulative @@ -333,12 +413,11 @@ pub enum ProgramClass { V3, } -/// PLACEHOLDER (ca2-node worker, 5 October 2026): the load class of program class v3 is set to w16 (16 loads of -/// four words, `LoadClass::fixed(4, 16)`) so the node, the miner, the packs and the fast-time gate can be built and -/// run before the Counter ASIC 2.0 measurements decide the width, the per-load mix and the scratch share -/// (`counter-asic-2-rollout.md` section 6). The integration agent on branch ca2-v3 replaces this constant with the -/// decided class; nothing else in the seam names the class, so it is one edit. -pub const V3_CLASS: LoadClass = LoadClass { mix: [0, 100, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 }; +/// The load class of program class v3, decided 5 October 2026 (Counter ASIC 2.0, `docs/plans/counter-asic-2-status.md` +/// "22:00 decided", `docs/plans/mixer-x4.md`): [`LoadClass::MX4`], version 2 loads (the width stays 4 bytes, the +/// per-load mix and the scratch share are out), the mixer applied 4 times per round and the cache growth rule. The +/// placeholder of the seam (w16) is replaced here; nothing else in the seam names the class. +pub const V3_CLASS: LoadClass = LoadClass::MX4; impl ProgramClass { /// The load class this program class draws from. @@ -494,6 +573,14 @@ pub fn program_id_class(generator: u32, seed: &[u32; 8], attempt: u32, class: &L b.push(k); b.push(class.scratch_kb); } + if class.mixer_mult != 1 || class.growth { + // Counter ASIC 2.0: the mixer multiplier and the growth rule are part of the construction, so a program of + // the same seed under a different mixer carries a different id (under the v3 seam the id is + // program_id(3, seed, attempt) and this branch is not taken) + b.extend_from_slice(b"mixer/"); + b.push(class.mixer_mult); + b.push(class.growth as u8); + } fnv1a64(&b) } @@ -625,7 +712,8 @@ pub fn candidate_from_words_class( let rot = 1 + rng.below(31) as u32; let bit = rng.below(32); let mask = 1u8 << rng.below(5); - let width = if class.is_v2() { 1 } else { class.width_for_roll(rng.below(100)) }; + // Version 2 loads take no width roll, so a mixer class with version 2 loads draws the version 2 program + let width = if class.takes_width_roll() { class.width_for_roll(rng.below(100)) } else { 1 }; let width = if op == Op::Load { width } else { 1 }; if op.is_load() { fresh[src as usize] = false; diff --git a/igneum-pow/src/lib.rs b/igneum-pow/src/lib.rs index 1f95e11a6..7cb094d44 100644 --- a/igneum-pow/src/lib.rs +++ b/igneum-pow/src/lib.rs @@ -32,7 +32,7 @@ pub mod verify; pub use bind::{block_init_words, day_bytes, pow256_from_lane, target64_from_le256}; pub use accept::{check as accept_program, AcceptReport, Reject}; -pub use generator::{generate, generate_from_seed_bytes, generate_from_seed_bytes_program_class, Instr, Op, Program, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, V3_CLASS}; -pub use memhard::{Cache, MemhardCpu, MixParams}; +pub use generator::{generate, generate_from_seed_bytes, generate_from_seed_bytes_program_class, Instr, LoadClass, Op, Program, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, V3_CLASS}; +pub use memhard::{cache_log2_words, dataset_log2_words, days_since_genesis, growth_doublings, Cache, MemhardCpu, MixParams, Shape}; pub use seed::{fnv1a64, seed_words, SplitMix64}; pub use verify::{hash_warp, interpret_warp_init, verify_block, DatasetMode, DatasetSource, Epoch}; diff --git a/igneum-pow/src/main.rs b/igneum-pow/src/main.rs index 1ece86032..a12f225ae 100644 --- a/igneum-pow/src/main.rs +++ b/igneum-pow/src/main.rs @@ -13,7 +13,7 @@ use igneum_pow::emit::export_pack; use igneum_pow::generator::LoadClass; -use igneum_pow::memhard::Cache; +use igneum_pow::memhard::{Cache, Shape}; use igneum_pow::seed::day_key; use igneum_pow::verify::{DatasetMode, Epoch, DEFAULT_DATASET_LOG2}; use std::time::Instant; @@ -31,6 +31,8 @@ struct Args { epoch_hex: Option, day_hex: Option, class: LoadClass, + /// Days since genesis for the cache growth rule of a class with `growth` (0: the genesis cache). + days: u64, } fn usage() -> ! { @@ -42,7 +44,8 @@ fn usage() -> ! { \x20 hash-bound --prehash <64 hex> --nonce print the header-bound hash (bind.rs) of one 64-bit nonce\n\ \x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\ \x20 show the accepted program, one instruction per line\n\ - \x20 --class C load class (read-width experiment): v2 (default), w4, w16, w64, w64x4, or p4,p16,p64[xN]" + \x20 --class C load class: v2 (default), mx4 (class v3: mixer x4, cache growth), w4, w16, w64, w64x4, p4,p16,p64[xN], m[g]\n\ + \x20 --days N days since genesis for the cache growth rule of a class with it (default 0: the 2^26-word cache)" ); std::process::exit(2) } @@ -61,6 +64,7 @@ fn parse() -> Args { epoch_hex: None, day_hex: None, class: LoadClass::V2, + days: 0, }; let mut it = std::env::args().skip(1); a.cmd = it.next().unwrap_or_else(|| usage()); @@ -78,6 +82,7 @@ fn parse() -> Args { "--epoch-hex" => a.epoch_hex = Some(val()), "--day-hex" => a.day_hex = Some(val()), "--class" => a.class = LoadClass::parse(&val()).unwrap_or_else(|| usage()), + "--days" => a.days = val().parse().unwrap_or_else(|_| usage()), _ => usage(), } } @@ -93,7 +98,7 @@ fn main() { "accept" => accept(&a), "show" => show(&a), "hash" => { - let e = Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class); + let e = Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days); println!("{:016x}", e.hash(a.nonce as u32)); } "hash-bound" => { @@ -106,7 +111,7 @@ fn main() { let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage()); Epoch::from_seed_bytes_class(&eb, &db, "cli", a.class) } - _ => Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class), + _ => Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days), }; let init = igneum_pow::bind::block_init_words(&prehash, a.nonce); println!("init words {}", init.iter().map(|w| format!("{w:08x}")).collect::>().join(" ")); @@ -124,24 +129,34 @@ fn bench(a: &Args, mode: DatasetMode) { a.dataset_log2, mode.name() ); + let shape = Shape::for_class_day(&a.class, a.days); if mode == DatasetMode::MemoryHard { // Time the cache fill on its own first (one core), then build the epoch (which fills it again). let t0 = Instant::now(); - let c = Cache::fill(day_key(&a.day)); + let c = Cache::fill_log2(day_key(&a.day), shape.cache_log2_words); let fill_ms = t0.elapsed().as_secs_f64() * 1e3; - println!("cache: fill {fill_ms:.1} ms on one core (2^26 words, 65536 chains of 64 ChaCha12 blocks), FNV-1a 64 {:016x}", c.fnv1a64()); + println!( + "cache: fill {fill_ms:.1} ms on one core (2^{} words, {} MiB, {} chains of 64 ChaCha12 blocks), FNV-1a 64 {:016x}", + shape.cache_log2_words, + shape.cache_words() * 4 / (1 << 20), + c.segments(), + c.fnv1a64() + ); drop(c); } let t0 = Instant::now(); - let e = Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class); + let e = Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days); let build_ms = t0.elapsed().as_secs_f64() * 1e3; println!( - "program: class {}, {} loads/hash, {} bytes/hash, widths (1,4,16 words) {:?}, {} items/warp, op mix {}; epoch built in {build_ms:.1} ms", + "program: class {}, {} loads/hash, {} bytes/hash, widths (1,4,16 words) {:?}, {} items/warp, mixer x{} ({} mixers/item), cache 2^{} words, op mix {}; epoch built in {build_ms:.1} ms", e.program.class.name(), e.program.loads_per_hash(), e.program.bytes_per_hash(), e.program.width_counts(), e.program.items_per_warp(), + shape.mixer_mult, + shape.mixers_per_item(), + shape.cache_log2_words, e.program.op_mix() ); let bases = [0u32, 4096, 1_000_000]; @@ -175,7 +190,7 @@ fn export(a: &Args, mode: DatasetMode) { let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage()); (Epoch::from_seed_bytes_class(&eb, &db, &format!("igneum-epoch/{eh}/day/{dh}"), a.class), format!("bytes:{dh}")) } - _ => (Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class), a.day.clone()), + _ => (Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days), a.day.clone()), }; let build_ms = t0.elapsed().as_secs_f64() * 1e3; println!("igneum-pow export {out}"); diff --git a/igneum-pow/src/memhard.rs b/igneum-pow/src/memhard.rs index c6f0727fd..b38aba64a 100644 --- a/igneum-pow/src/memhard.rs +++ b/igneum-pow/src/memhard.rs @@ -3,12 +3,18 @@ //! ARX-multiply mixer. The verifier holds the cache and never the dataset. //! //! All arithmetic is on u32 modulo 2^32. Rotations are by 1..31 at every call site. +//! +//! Counter ASIC 2.0 (5 October 2026, `docs/plans/mixer-x4.md`, behind the program class): the construction has a +//! [`Shape`], the mixer multiplier `m` and the cache size. Under `m` every mixer application of an item becomes +//! `m` applications with distinct round keys, the 8 dependent cache reads unchanged; the cache doubles when the +//! dataset doubles ([`growth_doublings`]). [`Shape::V2`] (`m = 1`, 2^26 words) is version 2 bit for bit. +use crate::generator::LoadClass; use crate::seed::{day_key, fnv1a64_words, SplitMix64}; pub const CACHE_LOG2_WORDS: usize = 26; pub const CACHE_SEGMENT_LOG2_LINES: usize = 6; -/// 2^26 words = 256 MiB. +/// 2^26 words = 256 MiB (the version 2 cache, and the v3 cache until the first dataset doubling). pub const CACHE_WORDS: usize = 1 << CACHE_LOG2_WORDS; /// 2^22 lines of 16 words. pub const CACHE_LINES: usize = CACHE_WORDS >> 4; @@ -24,6 +30,94 @@ pub const CHACHA_SIGMA: [u32; 4] = [0x61707865, 0x3320646e, 0x79622d32, 0x6b2065 /// "Igne", "umMH". pub const CACHE_TAG: [u32; 2] = [0x49676e65, 0x756d4d48]; +/// The shape of the item derivation and of the cache: the mixer multiplier and the cache size. +#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)] +pub struct Shape { + /// Mixer applications per round (and after the last read): 1 under version 2, 4 under class v3. + pub mixer_mult: u32, + /// The cache is 2^cache_log2_words words (26 at genesis; 27 and 28 after the dataset doublings of 1.13.3). + pub cache_log2_words: u32, +} + +impl Shape { + /// Version 2: one mixer application per round, a 2^26-word cache. + pub const V2: Shape = Shape { mixer_mult: 1, cache_log2_words: CACHE_LOG2_WORDS as u32 }; + + /// The shape of a load class on day 0 of the chain (and on every day for a class without the growth rule). + pub fn for_class(class: &LoadClass) -> Shape { + Shape::for_class_day(class, 0) + } + + /// The shape of a load class on day `days_since_genesis` of the chain: the class's multiplier, and the cache + /// of [`cache_log2_words`] when the class has the growth rule, else 2^26 words. + pub fn for_class_day(class: &LoadClass, days_since_genesis: u64) -> Shape { + Shape { + mixer_mult: class.mixer_mult(), + cache_log2_words: if class.growth { cache_log2_words(days_since_genesis) } else { CACHE_LOG2_WORDS as u32 }, + } + } + + pub fn is_v2(&self) -> bool { + *self == Shape::V2 + } + pub fn cache_words(&self) -> usize { + 1usize << self.cache_log2_words + } + pub fn cache_lines(&self) -> usize { + self.cache_words() >> 4 + } + pub fn cache_line_mask(&self) -> u32 { + (self.cache_lines() - 1) as u32 + } + pub fn cache_segments(&self) -> usize { + self.cache_lines() >> CACHE_SEGMENT_LOG2_LINES + } + pub fn log2_segments(&self) -> u32 { + self.cache_log2_words - 4 - CACHE_SEGMENT_LOG2_LINES as u32 + } + /// Mixer applications per item: `(ITEM_ROUNDS + 1) x m`. + pub fn mixers_per_item(&self) -> u32 { + (ITEM_ROUNDS as u32 + 1) * self.mixer_mult + } +} + +// -------------------------------------------------------------------------------------------------------------- +// Dataset growth, option C (spec 01 section 1.13.3 option (b) with the cache tied to the dataset's doublings) +// -------------------------------------------------------------------------------------------------------------- + +/// Days per year of the growth schedule: one year = 31,536,000 DAA seconds of 86,400 (spec 01 section 1.13.3). +pub const GROWTH_DAYS_PER_YEAR: u64 = 365; +/// The linear schedule of 1.13.3, 2 GiB at genesis plus 0.5 GiB per year, is `G x (1 + d / 1460)` for the genesis +/// size `G` and the day `d`: it doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at +/// day 10,220 (year 28) and 16x at day 21,900 (year 60). +pub const GROWTH_DOUBLING_DAYS: u64 = 4 * GROWTH_DAYS_PER_YEAR; + +/// The number of dataset doublings reached by day `days_since_genesis` of the chain: `floor(log2(1 + d / 1460))`, +/// in integers (`1 + d / 1460` rounded down, then its integer log2, which equals the real log2's floor because a +/// power of two is an integer). 0 until day 1,459; 1 from day 1,460 (year 4); 2 from day 4,380 (year 12). +pub fn growth_doublings(days_since_genesis: u64) -> u32 { + (1 + days_since_genesis / GROWTH_DOUBLING_DAYS).ilog2() +} + +/// The cache size on day `d` under option C: 2^26 words doubled once per dataset doubling (256 MiB, 512 MiB from +/// year 4, 1 GiB from year 12). +pub fn cache_log2_words(days_since_genesis: u64) -> u32 { + CACHE_LOG2_WORDS as u32 + growth_doublings(days_since_genesis) +} + +/// The dataset size on day `d` under option (b) of 1.13.3: the genesis size (2^`genesis_log2_words` words: 28 for +/// the 1 GiB packs and the devnet, 29 for the designed 2 GiB) doubled once per doubling of the linear schedule. The +/// result is capped at 32 (the item index is 32 bits, spec 1.13.3). +pub fn dataset_log2_words(genesis_log2_words: u32, days_since_genesis: u64) -> u32 { + (genesis_log2_words + growth_doublings(days_since_genesis)).min(32) +} + +/// Days since genesis from two day indices of `bind::day_index` (the header's `timestamp_ms / 86,400,000`): the +/// day of the block and the day of the genesis header. A block before the genesis day (clock skew) is day 0. +pub fn days_since_genesis(day_index: u64, genesis_day_index: u64) -> u64 { + day_index.saturating_sub(genesis_day_index) +} + #[inline(always)] fn rotl(x: u32, n: u32) -> u32 { x.rotate_left(n) @@ -66,17 +160,23 @@ pub fn chacha_block(x: &[u32; 16]) -> [u32; 16] { y } -/// Mixer parameters drawn from the day key. Draw order: ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15]. +/// Mixer parameters drawn from the day key, plus the [`Shape`] the mixer is applied under. Draw order: +/// ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15]. The shape is not drawn: it is the class's. #[derive(Clone, Debug, PartialEq, Eq)] pub struct MixParams { pub key: [u32; 8], pub rot: [u32; 8], pub mul: [u32; 16], pub rc: [u32; 16], + pub shape: Shape, } impl MixParams { + /// Version 2 shape. pub fn new(key: [u32; 8]) -> Self { + Self::with_shape(key, Shape::V2) + } + pub fn with_shape(key: [u32; 8], shape: Shape) -> Self { let mut rng = SplitMix64::new(key[0] as u64 | ((key[1] as u64) << 32)); let mut rot = [0u32; 8]; let mut mul = [0u32; 16]; @@ -90,7 +190,7 @@ impl MixParams { for c in rc.iter_mut() { *c = rng.next() as u32; } - Self { key, rot, mul, rc } + Self { key, rot, mul, rc, shape } } /// Parameters for a day string: the key is `seed_words("day/" + day)`. pub fn for_day(day: &str) -> Self { @@ -104,6 +204,13 @@ pub fn round_key(r: usize) -> u32 { ((r + 1) as u32).wrapping_mul(0x9E3779B9) } +/// The round key of application `j` (0 <= j < m) of round `r` under multiplier `m`: `round_key(r * m + j)`. For +/// `m = 1` this is `round_key(r)`, version 2's key. +#[inline(always)] +pub fn round_key_mult(r: usize, j: usize, m: usize) -> u32 { + round_key(r * m + j) +} + /// `M_r` on 16 words in place: per word `(s ^ (RC + rk)) * MUL`, then one ChaCha-shaped double round with /// the four column rotations `ROT[0..3]` and the four diagonal rotations `ROT[4..7]`. #[inline(always)] @@ -122,9 +229,11 @@ pub fn mixer(s: &mut [u32; 16], rk: u32, mp: &MixParams) { qr(s, 3, 4, 9, 14, r[4], r[5], r[6], r[7]); } -/// The 256 MiB cache for one day key. +/// The cache for one day key: 2^log2_words words (256 MiB under version 2). pub struct Cache { pub key: [u32; 8], + pub log2_words: u32, + line_mask: u32, words: Vec, } @@ -152,13 +261,22 @@ impl Cache { } } - /// The whole cache on the calling thread: 65,536 chains of 64 ChaCha12 blocks, in segment order. + /// The version 2 cache on the calling thread: 65,536 chains of 64 ChaCha12 blocks, in segment order. pub fn fill(key: [u32; 8]) -> Cache { - let mut words = vec![0u32; CACHE_WORDS]; - for seg in 0..CACHE_SEGMENTS { + Self::fill_log2(key, CACHE_LOG2_WORDS as u32) + } + + /// A cache of 2^`log2_words` words (26, 27 or 28 under the growth rule; smaller sizes for tests): 2^(log2 - 10) + /// independent chains of 64 lines, the same chain function at every size, so a larger cache's first segments + /// are the smaller cache's segments word for word. + pub fn fill_log2(key: [u32; 8], log2_words: u32) -> Cache { + assert!((10..=30).contains(&log2_words), "cache log2 words must be in 10..=30"); + let shape = Shape { mixer_mult: 1, cache_log2_words: log2_words }; + let mut words = vec![0u32; shape.cache_words()]; + for seg in 0..shape.cache_segments() { Self::fill_segment(&mut words, seg, &key); } - Cache { key, words } + Cache { key, log2_words, line_mask: shape.cache_line_mask(), words } } pub fn for_day(day: &str) -> Cache { @@ -170,10 +288,20 @@ impl Cache { &self.words } - /// Cache line `a` (0 <= a < 2^22) as 16 words. + pub fn lines(&self) -> usize { + self.words.len() >> 4 + } + pub fn line_mask(&self) -> u32 { + self.line_mask + } + pub fn segments(&self) -> usize { + self.lines() >> CACHE_SEGMENT_LOG2_LINES + } + + /// Cache line `a` (masked to the cache's lines) as 16 words. #[inline(always)] pub fn line(&self, a: u32) -> &[u32] { - let o = (a & CACHE_LINE_MASK) as usize * 16; + let o = (a & self.line_mask) as usize * 16; &self.words[o..o + 16] } @@ -184,10 +312,13 @@ impl Cache { } /// Derive `ts.len()` items into `out`, all chains interleaved round by round so the cache-line misses of -/// independent items overlap in the memory system (`deriveItems` in the Swift). +/// independent items overlap in the memory system (`deriveItems` in the Swift). Under multiplier `m` +/// (`mp.shape.mixer_mult`) round `r` applies `M` with keys `round_key(r m + j)` for `j = 0 .. m - 1` before its +/// one cache read; the final mixer applies `M` with keys `round_key(8 m + j)`. `m = 1` is version 2. pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; 16]]) { let n = ts.len(); debug_assert!(out.len() >= n); + let m = mp.shape.mixer_mult as usize; for k in 0..n { let s = &mut out[k]; let t = ts[k]; @@ -197,9 +328,11 @@ pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; } } for r in 0..ITEM_ROUNDS { - let rk = round_key(r); - for s in out[..n].iter_mut() { - mixer(s, rk, mp); + for j in 0..m { + let rk = round_key_mult(r, j, m); + for s in out[..n].iter_mut() { + mixer(s, rk, mp); + } } for s in out[..n].iter_mut() { let line = cache.line(s[0]); @@ -208,9 +341,11 @@ pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; } } } - let rk = round_key(ITEM_ROUNDS); - for s in out[..n].iter_mut() { - mixer(s, rk, mp); + for j in 0..m { + let rk = round_key_mult(ITEM_ROUNDS, j, m); + for s in out[..n].iter_mut() { + mixer(s, rk, mp); + } } } @@ -221,7 +356,7 @@ pub fn derive_item(t: u32, mp: &MixParams, cache: &Cache) -> [u32; 16] { out[0] } -/// The CPU verifier's view of the memory-hard dataset: the mixer parameters and the 256 MiB cache. +/// The CPU verifier's view of the memory-hard dataset: the mixer parameters (with the shape) and the cache. pub struct MemhardCpu { pub params: MixParams, pub cache: Cache, @@ -231,12 +366,19 @@ pub struct MemhardCpu { pub const FETCH_MAX: usize = 64; impl MemhardCpu { + /// Version 2 shape. pub fn new(key: [u32; 8]) -> Self { - Self { params: MixParams::new(key), cache: Cache::fill(key) } + Self::with_shape(key, Shape::V2) + } + pub fn with_shape(key: [u32; 8], shape: Shape) -> Self { + Self { params: MixParams::with_shape(key, shape), cache: Cache::fill_log2(key, shape.cache_log2_words) } } pub fn for_day(day: &str) -> Self { Self::new(day_key(day)) } + pub fn shape(&self) -> Shape { + self.params.shape + } /// `dataset[w] = item(w >> 4)[w & 15]`. pub fn word(&self, w: u32) -> u32 { derive_item(w >> 4, &self.params, &self.cache)[(w & 15) as usize] @@ -313,6 +455,7 @@ mod tests { assert_eq!(mp.rc[0], 0xbab68293); assert_eq!(mp.rc[15], 0x31b49ee2); assert!(mp.mul.iter().all(|m| m & 1 == 1)); + assert_eq!(mp.shape, Shape::V2); } #[test] @@ -338,4 +481,96 @@ mod tests { ] ); } + + /// Option C: the schedule table of `docs/plans/mixer-x4.md` (day -> doublings, cache words, dataset words at a + /// 2^28 genesis). The doublings fall at years 4 and 12 exactly, never a day early. + #[test] + fn growth_schedule_table() { + let table: [(u64, u32, u32, u32); 12] = [ + (0, 0, 26, 28), + (1, 0, 26, 28), + (365, 0, 26, 28), + (1_459, 0, 26, 28), + (1_460, 1, 27, 29), + (2_920, 1, 27, 29), + (4_379, 1, 27, 29), + (4_380, 2, 28, 30), + (10_219, 2, 28, 30), + (10_220, 3, 29, 31), + (21_900, 4, 30, 32), + (100_000, 6, 32, 32), + ]; + for (d, k, c, s) in table { + assert_eq!(growth_doublings(d), k, "day {d}"); + assert_eq!(cache_log2_words(d), c, "day {d}"); + assert_eq!(dataset_log2_words(28, d), s, "day {d}"); + } + // the designed 2 GiB genesis: 2^29 words, 2^30 at year 4, 2^31 at year 12 + assert_eq!(dataset_log2_words(29, 0), 29); + assert_eq!(dataset_log2_words(29, 1_460), 30); + assert_eq!(dataset_log2_words(29, 4_380), 31); + // the linear schedule itself: 2 GiB x (1 + d / 1460) crosses 4 GiB at day 1,460 and 8 GiB at day 4,380 + for d in [1_459u64, 1_460, 4_379, 4_380] { + let bytes = 2u64 * (1 << 30) + (1u64 << 29) * d / 365; + let k = (bytes / (2u64 << 30)).ilog2(); + assert_eq!(growth_doublings(d), k, "day {d}: linear {bytes} bytes"); + } + assert_eq!(days_since_genesis(20_730, 20_729), 1); + assert_eq!(days_since_genesis(20_729, 20_729), 0); + assert_eq!(days_since_genesis(20_000, 20_729), 0); + let v2 = Shape::for_class_day(&LoadClass::V2, 100_000); + assert_eq!(v2, Shape::V2); + let v3 = Shape::for_class_day(&LoadClass::MX4, 0); + assert_eq!(v3, Shape { mixer_mult: 4, cache_log2_words: 26 }); + assert_eq!(Shape::for_class_day(&LoadClass::MX4, 1_460).cache_log2_words, 27); + assert_eq!(v3.mixers_per_item(), 36); + assert_eq!(Shape::V2.mixers_per_item(), 9); + assert_eq!(Shape::V2.cache_segments(), CACHE_SEGMENTS); + assert_eq!(Shape::V2.cache_line_mask(), CACHE_LINE_MASK); + assert_eq!(Shape::V2.log2_segments(), 16); + } + + /// The multiplied mixer, restated by hand on a small cache: `m` applications with keys `round_key(r m + j)` + /// before every read, the same 8 reads; `m = 1` is `derive_item` of version 2 word for word; a larger cache's + /// first segments equal the smaller cache's. + #[test] + fn mixer_mult_by_hand() { + let key = day_key("2026-10-03"); + let small = Cache::fill_log2(key, 16); + let big = Cache::fill_log2(key, 18); + assert_eq!(&big.words()[..small.words().len()], small.words()); + assert_eq!(small.segments(), 64); + assert_eq!(small.line_mask(), 4095); + for m in [1u32, 2, 4] { + let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 16 }); + for t in [0u32, 1, 12_345, u32::MAX] { + let got = derive_item(t, &mp, &small); + let mut s = [0u32; 16]; + s[..8].copy_from_slice(&key); + for i in 0..8 { + s[8 + i] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]); + } + for r in 0..8usize { + for j in 0..m as usize { + mixer(&mut s, round_key(r * m as usize + j), &mp); + } + let line = small.line(s[0]); + for i in 0..16 { + s[i] ^= line[i]; + } + } + for j in 0..m as usize { + mixer(&mut s, round_key(8 * m as usize + j), &mp); + } + assert_eq!(got, s, "m {m} t {t}"); + } + } + let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16 }); + let v3 = MixParams::with_shape(key, Shape { mixer_mult: 4, cache_log2_words: 16 }); + assert_ne!(derive_item(0, &v2, &small), derive_item(0, &v3, &small)); + assert_eq!(round_key_mult(0, 0, 1), round_key(0)); + assert_eq!(round_key_mult(8, 0, 1), round_key(8)); + assert_eq!(round_key_mult(2, 3, 4), round_key(11)); + assert_eq!(round_key_mult(8, 3, 4), round_key(35)); + } } diff --git a/igneum-pow/src/verify.rs b/igneum-pow/src/verify.rs index 3b93e79c0..97fbfe94a 100644 --- a/igneum-pow/src/verify.rs +++ b/igneum-pow/src/verify.rs @@ -2,7 +2,7 @@ //! calls. Dataset words come from the memory-hard cache (default) or from the closed form (old packs). use crate::generator::{generate, generate_class, Instr, LoadClass, Op, Program, ProgramClass, ITERATIONS, LANES}; -use crate::memhard::MemhardCpu; +use crate::memhard::{MemhardCpu, Shape}; use crate::seed::day_key; /// Read-width experiment (5 October 2026): a `load` of `W` words folds every word into `dst`: @@ -150,21 +150,37 @@ pub struct DatasetSource { impl DatasetSource { /// Build the source for a day. Memory-hard mode fills the 256 MiB cache on the calling thread. pub fn new(day: &str, mode: DatasetMode, log2_words: u32) -> Self { - let mut ds = Self::from_key(day_key(day), mode, log2_words); + Self::new_shape(day, mode, log2_words, Shape::V2) + } + + /// [`DatasetSource::new`] with the construction's shape (mixer multiplier, cache size; Counter ASIC 2.0). + pub fn new_shape(day: &str, mode: DatasetMode, log2_words: u32, shape: Shape) -> Self { + let mut ds = Self::from_key_shape(day_key(day), mode, log2_words, shape); ds.key_bytes = format!("day/{day}").into_bytes(); ds } pub fn from_key(key: [u32; 8], mode: DatasetMode, log2_words: u32) -> Self { + Self::from_key_shape(key, mode, log2_words, Shape::V2) + } + + /// [`DatasetSource::from_key`] with the construction's shape. Memory-hard mode fills a cache of + /// `2^shape.cache_log2_words` words on the calling thread. + pub fn from_key_shape(key: [u32; 8], mode: DatasetMode, log2_words: u32, shape: Shape) -> Self { assert!((4..=32).contains(&log2_words), "dataset log2 must be in 4..=32"); let mask = if log2_words == 32 { u32::MAX } else { (1u32 << log2_words) - 1 }; let dataset = match mode { DatasetMode::ClosedForm => Dataset::ClosedForm { d0: key[0], d1: key[1] }, - DatasetMode::MemoryHard => Dataset::MemoryHard(MemhardCpu::new(key)), + DatasetMode::MemoryHard => Dataset::MemoryHard(MemhardCpu::with_shape(key, shape)), }; Self { log2_words, mask, key, key_bytes: Vec::new(), dataset } } + /// The shape of the memory-hard construction ([`Shape::V2`] for the closed form, which has none). + pub fn shape(&self) -> Shape { + self.memhard().map(|m| m.shape()).unwrap_or(Shape::V2) + } + pub fn mode(&self) -> DatasetMode { match self.dataset { Dataset::ClosedForm { .. } => DatasetMode::ClosedForm, @@ -417,9 +433,10 @@ pub struct Epoch { /// Default dataset size: 2^28 words = 1 GiB. pub const DEFAULT_DATASET_LOG2: u32 = 28; -/// Days a day index lies after the network's genesis day (0 for the genesis day and any day before it). +/// Days a day index lies after the network's genesis day (0 for the genesis day and any day before it). The node's +/// entry; the same function as `memhard::days_since_genesis`. pub fn days_since_genesis(day_index: u64, genesis_day_index: u64) -> u64 { - day_index.saturating_sub(genesis_day_index) + crate::memhard::days_since_genesis(day_index, genesis_day_index) } impl Epoch { @@ -427,9 +444,17 @@ impl Epoch { Self { program: generate(seed), dataset: DatasetSource::new(day, mode, dataset_log2) } } - /// [`Epoch::new`] with a load class (read-width experiment). + /// [`Epoch::new`] with a load class (read-width experiment; Counter ASIC 2.0: the class's mixer multiplier + /// shapes the dataset, the cache is the genesis size since a string day has no day index). pub fn new_class(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32, class: LoadClass) -> Self { - Self { program: generate_class(seed, class), dataset: DatasetSource::new(day, mode, dataset_log2) } + Self::new_class_day(seed, day, mode, dataset_log2, class, 0) + } + + /// [`Epoch::new_class`] on day `days_since_genesis` of the growth schedule (the cache of + /// `memhard::cache_log2_words` for a class with the growth rule; the dataset size is the caller's). + pub fn new_class_day(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32, class: LoadClass, days_since_genesis: u64) -> Self { + let shape = Shape::for_class_day(&class, days_since_genesis); + Self { program: generate_class(seed, class), dataset: DatasetSource::new_shape(day, mode, dataset_log2, shape) } } /// The production shape: memory-hard, 1 GiB dataset. @@ -445,11 +470,24 @@ impl Epoch { Self::from_seed_bytes_class(epoch_seed, day_bytes, label, LoadClass::V2) } - /// [`Epoch::from_seed_bytes`] with a load class (read-width experiment). + /// [`Epoch::from_seed_bytes`] with a load class (read-width experiment; Counter ASIC 2.0: the class's mixer + /// multiplier shapes the dataset). Day 0 of the growth schedule: the 2^26-word cache and the 2^28-word dataset, + /// which is every devnet pack and vector. A node past the first doubling calls [`Epoch::from_seed_bytes_day`]. pub fn from_seed_bytes_class(epoch_seed: &[u8], day_bytes: &[u8], label: &str, class: LoadClass) -> Self { + Self::from_seed_bytes_day(epoch_seed, day_bytes, label, class, 0, DEFAULT_DATASET_LOG2) + } + + /// The chain's shape on day `days_since_genesis` (`memhard::days_since_genesis(day_index(header), day_index(genesis))`, + /// the node's two day indices): the program of the class, and under the class's growth rule the cache of + /// `memhard::cache_log2_words(d)` and the dataset of `memhard::dataset_log2_words(genesis_dataset_log2, d)` + /// (the genesis size is 28 for the 1 GiB devnet, 29 for the designed 2 GiB). Without the growth rule the cache + /// is 2^26 words and the dataset `2^genesis_dataset_log2` on every day. + pub fn from_seed_bytes_day(epoch_seed: &[u8], day_bytes: &[u8], label: &str, class: LoadClass, days_since_genesis: u64, genesis_dataset_log2: u32) -> Self { let program = crate::generator::generate_from_seed_bytes_class(label, epoch_seed, class); let key = crate::seed::seed_words_from_bytes(day_bytes); - let mut dataset = DatasetSource::from_key(key, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2); + let shape = Shape::for_class_day(&class, days_since_genesis); + let dataset_log2 = if class.growth { crate::memhard::dataset_log2_words(genesis_dataset_log2, days_since_genesis) } else { genesis_dataset_log2 }; + let mut dataset = DatasetSource::from_key_shape(key, DatasetMode::MemoryHard, dataset_log2, shape); dataset.key_bytes = day_bytes.to_vec(); Self { program, dataset } } @@ -477,13 +515,18 @@ impl Epoch { /// [`Epoch::chain_dataset`] with the day's position since genesis and the network's genesis dataset size: the /// entry the node's engine and the miner's export build every day cache through, so the cache growth schedule - /// of spec 01 section 1.13.3 has one place to act. SEAM (ca2-node, 5 October 2026): the body here builds the - /// genesis-size cache for every day; the ca2-mixer branch fills the growth rule (`growth_doublings`) and the - /// class v3 item construction, keeping this signature. + /// of spec 01 section 1.13.3 has one place to act (ca2-mixer, 5 October 2026, `docs/plans/mixer-x4.md`): the + /// class's load class gives the mixer multiplier and whether the growth rule applies (`Shape::for_class_day`); + /// under the rule the cache is `2^memhard::cache_log2_words(d)` words and the dataset + /// `2^memhard::dataset_log2_words(genesis_dataset_log2, d)`; without it (class v2) the cache is 2^26 words and + /// the dataset the genesis size on every day. `days_since_genesis` is [`days_since_genesis`] of the block's and + /// the genesis header's day indices. pub fn chain_dataset_day(day_bytes: &[u8], class: ProgramClass, days_since_genesis: u64, genesis_dataset_log2: u32) -> DatasetSource { - let _ = (class, days_since_genesis); + let lc = class.load_class(); + let shape = Shape::for_class_day(&lc, days_since_genesis); + let dataset_log2 = if lc.growth { crate::memhard::dataset_log2_words(genesis_dataset_log2, days_since_genesis) } else { genesis_dataset_log2 }; let key = crate::seed::seed_words_from_bytes(day_bytes); - let mut dataset = DatasetSource::from_key(key, DatasetMode::MemoryHard, genesis_dataset_log2); + let mut dataset = DatasetSource::from_key_shape(key, DatasetMode::MemoryHard, dataset_log2, shape); dataset.key_bytes = day_bytes.to_vec(); dataset } diff --git a/proto-cuda/nvrtc/packfile.h b/proto-cuda/nvrtc/packfile.h index ba5587a99..2a740998d 100644 --- a/proto-cuda/nvrtc/packfile.h +++ b/proto-cuda/nvrtc/packfile.h @@ -31,6 +31,9 @@ typedef struct { char loadClass[64]; char programClass[8]; /* IGNEUM_PROGRAM_CLASS: "v2" or "v3" (Counter ASIC 2.0); absent = the generator's class */ char eraHex[65]; /* IGNEUM_ERA_SEED_HEX of a class v3 chain pack; empty otherwise */ + // Counter ASIC 2.0 (5 October 2026): the mixer multiplier of the item derivation (IGNEUM_MIXER_MULT, 1 when absent: + // version 2; 4 under class v3). The emitted memhard.h / kernel.cl carry it in their text; this is for the log lines. + uint32_t mixerMult; // seeds.txt (or program.h): the seeds as the worker protocol carries them char epochHex[65]; char dayHex[PF_HEX_CAP]; @@ -296,6 +299,7 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) { pk->persistent = 0; pf_define_u32(prog, "IGNEUM_PERSISTENT_WARPS", &pk->persistent); pk->scratchWordsPerLane = 8192; pf_define_u32(prog, "IGNEUM_SCRATCH_WORDS_PER_LANE", &pk->scratchWordsPerLane); strcpy(pk->loadClass, "v2"); pf_define_str(prog, "IGNEUM_LOAD_CLASS", pk->loadClass, sizeof(pk->loadClass)); + pk->mixerMult = 1; pf_define_u32(prog, "IGNEUM_MIXER_MULT", &pk->mixerMult); if (pf_define_words(prog, "IGNEUM_KEY_INIT", pk->keyw, 8) != 8) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_KEY_INIT with 8 words"); } if (!pf_define_str(prog, "IGNEUM_SEED_STRING", pk->seedString, sizeof(pk->seedString))) strncpy(pk->seedString, "(no IGNEUM_SEED_STRING)", sizeof(pk->seedString) - 1); pf_define_str(prog, "IGNEUM_SEED_BYTES_HEX", ehex, sizeof(ehex)); diff --git a/proto-metal/packbench.swift b/proto-metal/packbench.swift index 3110bbe12..b7eda2f3c 100644 --- a/proto-metal/packbench.swift +++ b/proto-metal/packbench.swift @@ -61,6 +61,8 @@ let scratchOps = Int(defineU32("IGNEUM_SCRATCH_OPS") ?? 0) let persistent = (defineU32("IGNEUM_PERSISTENT_WARPS") ?? 0) == 1 let scratchWordsPerLane = Int(defineU32("IGNEUM_SCRATCH_WORDS_PER_LANE") ?? 8192) let className = defineStr("IGNEUM_LOAD_CLASS") ?? "v2" +// Counter ASIC 2.0 (5 October 2026): the mixer multiplier of the item derivation, 1 when absent (version 2), 4 under class v3 +let mixerMult = Int(defineU32("IGNEUM_MIXER_MULT") ?? 1) let seedString = defineStr("IGNEUM_SEED_STRING") ?? "?" let programId = defineStr("IGNEUM_PROGRAM_ID") ?? "" @@ -200,7 +202,7 @@ for b in 0..> 20), IGNEUM_ITEM_ROUNDS, IGNEUM_MIXER_MULT); cachePass = setupCache(&dv, di); #else printf("dataset construction: closed-form ds_elem (the original prototype dataset, not memory-hard)\n"); diff --git a/relay/playbooks/mixer-x4-5090-bench.ps1 b/relay/playbooks/mixer-x4-5090-bench.ps1 new file mode 100644 index 000000000..74f4a184a --- /dev/null +++ b/relay/playbooks/mixer-x4-5090-bench.ps1 @@ -0,0 +1,63 @@ +# Igneum run job (bench only): the class v3 dataset construction (mixer x4, docs/plans/mixer-x4.md) on PC 2's RTX 5090 (machine 1ccfe586), 5 October 2026. +# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches +# off ONLY the NVIDIA card in the app through POST api/cards, waits for its worker to stop, runs the fetched +# igneum-worker-cuda.exe on the v2 pack and the two v3 packs of the fetched packs folder (the dataset build time per pack is the +# number this job is for: the worker's own `cache ... dataset ... ms` line, the same line the readwidth round printed), +# and switches the card back on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs ` shows them. +# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis and mx4-devnet-epoch0 from branch ca2-mixer, plus +# proto-cuda/packs/igneum-genesis-mh copied in as v2-genesis-mh (the version 2 control, same card, same run). +$ErrorActionPreference = 'Continue' +function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) } +$jobs = Split-Path $env:IGNEUM_JOB_DIR +$fetched = Join-Path $jobs 'fetch-mixer-x4-20261005' +$exe = Join-Path $fetched 'igneum-worker-cuda.exe' +$packs = Join-Path $fetched 'packs-ca2-mixer' +if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 } +if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 } +$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1 +if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe (the NVRTC DLLs come from there)'; exit 2 } +Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination $fetched -Force +Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower()) with $((Get-ChildItem $fetched -Filter 'nvrtc*.dll').Count) NVRTC DLL(s) from $inst" + +# the app: switch off the NVIDIA card only, remember its settings +$appDir = $env:IGNEUM_APP_DIR +if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' } +$urlFile = Join-Path $appDir 'app.url' +$url = $null +if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() } +$card = $null +if ($url) { + try { + $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10 + $card = $st.mining.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 + if (-not $card) { $card = $st.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 } + } catch { Say ("api/state: " + $_.Exception.Message) } +} +if ($card) { + Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state) + $body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) } + $t = 0 + while ($t -lt 90) { + Start-Sleep -Seconds 5; $t += 5 + try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { } + } + Write-Output ("RESULT card-off after " + $t + " s") + Start-Sleep -Seconds 5 +} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' } + +& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" } +# the probe ran in run-readwidth-5090-20261005 (bench-log, 5 October 2026); this second run is the bench only +foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" } + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT $_" } +} +& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" } + +if ($card) { + $body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) } +} +exit 0 diff --git a/relay/playbooks/mixer-x4-9070-bench.ps1 b/relay/playbooks/mixer-x4-9070-bench.ps1 new file mode 100644 index 000000000..50df5b654 --- /dev/null +++ b/relay/playbooks/mixer-x4-9070-bench.ps1 @@ -0,0 +1,64 @@ +# Igneum run job (bench only): the class v3 dataset construction (mixer x4, docs/plans/mixer-x4.md) on PC 1's RX 9070 XT on the eGPU (machine ae432dc7), 5 October 2026. +# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches +# off ONLY the NVIDIA card in the app through POST api/cards, waits for its worker to stop, runs the fetched +# igneum-worker-opencl.exe on the v2 pack and the two v3 packs of the fetched packs folder (the dataset build time per pack is the +# number this job is for: the worker's own `cache ... dataset ... ms` line, the same line the readwidth round printed), +# and switches the card back on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs ` shows them. +# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis and mx4-devnet-epoch0 from branch ca2-mixer, plus +# proto-cuda/packs/igneum-genesis-mh copied in as v2-genesis-mh (the version 2 control, same card, same run). +$ErrorActionPreference = 'Continue' +function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) } +$jobs = Split-Path $env:IGNEUM_JOB_DIR +$fetched = Join-Path $jobs 'fetch-mixer-x4-20261005' +$exe = Join-Path $fetched 'igneum-worker-opencl.exe' +$packs = Join-Path $fetched 'packs-ca2-mixer' +if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 } +if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 } +Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower())" +# the card's OpenCL device index on the current (3683.0) platform, from the worker's own list (the older platform's duplicate is marked dup) +$list = & $exe --list 2>&1 +$list | ForEach-Object { "RESULT list $_" } +$dev = $null +foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1201' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; break } } +if ($null -eq $dev) { Write-Output 'RESULT error no gfx1201 device in --list'; exit 2 } +Write-Output "RESULT device $dev" + +# the app: switch off the NVIDIA card only, remember its settings +$appDir = $env:IGNEUM_APP_DIR +if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' } +$urlFile = Join-Path $appDir 'app.url' +$url = $null +if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() } +$card = $null +if ($url) { + try { + $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10 + $card = $st.mining.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1 + if (-not $card) { $card = $st.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1 } + } catch { Say ("api/state: " + $_.Exception.Message) } +} +if ($card) { + Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state) + $body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) } + $t = 0 + while ($t -lt 90) { + Start-Sleep -Seconds 5; $t += 5 + try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { } + } + Write-Output ("RESULT card-off after " + $t + " s") + Start-Sleep -Seconds 5 +} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' } + +# the probe ran in run-readwidth-9070-20261005 (bench-log, 5 October 2026); this second run is the bench only +foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" } +} + +if ($card) { + $body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) } +} +exit 0