diff --git a/docs/plans/cryptanalysis/plan-mixer.md b/docs/plans/cryptanalysis/plan-mixer.md new file mode 100644 index 000000000..c9749b7ba --- /dev/null +++ b/docs/plans/cryptanalysis/plan-mixer.md @@ -0,0 +1,215 @@ +# Attack plan: the memory-hard mixer M_r + +Internal adversarial pass, not an independent review. + +Label rule: the phrase "internal adversarial pass, not an independent review" goes on every sentence from this +work that could be quoted in public. This is such a pass. It is not an outside review. + +- Target commit: 017e70376489251e18564c0abce7e466e606c8b3 (class v4 sub-version 3, object byte 7). +- Branch: adv-mixer, from build/master. +- Attacker model: an outsider with the public kit. No defender numbers are read; any defender figure here is + derived from the crate or marked unknown. +- Author identity in the mixer: the attacker has never worked on the hash code. + +## 0. The byte-identity check (the brief's gate 1) + +The brief asks that `git diff --stat 017e7037... HEAD -- igneum-pow` print nothing. It does NOT print nothing. +Six files differ between the frozen commit and build/master HEAD: + +| File | In the target? | +|---|---| +| igneum-pow/src/accept.rs | No (program acceptance, not the mixer) | +| igneum-pow/src/emit.rs | No (kernel emitters) | +| igneum-pow/src/generator.rs | No (program generator; V4_CLASS and Shape present on both) | +| igneum-pow/src/packcheck.rs | No (pack checker) | +| igneum-pow/tests/mixer.rs | No (test harness) | +| igneum-pow/tests/recheck.rs | No (test harness) | + +The mixer itself is byte-identical. `git diff --stat 017e7037... HEAD -- igneum-pow/src/memhard.rs +igneum-pow/src/seed.rs igneum-pow/src/bind.rs igneum-pow/src/derive.rs igneum-pow/Cargo.toml +igneum-pow/Cargo.lock` prints nothing. So the mixer draw code (seed.rs), the mixer function and the item +derivation (memhard.rs), the day rule (bind.rs) and the dependency pin (Cargo.toml, Cargo.lock) are the frozen +ones. The harness builds against build/master's igneum-pow, whose generator.rs and accept.rs differ from frozen; +those files are not M_r. So the numbers this harness produces are the frozen mixer's numbers. Stated plainly: +build/master is NOT byte-identical to the frozen commit over the whole crate, but it IS byte-identical over every +file that defines M_r and its parameters. + +## 1. The target, in the attacker's words + +The dataset item is 16 words of 32 bits. The mixer M is one keyed round made of two layers. + +1. Multiply layer, per word: `s[i] = (s[i] XOR (RC[i] + rk)) * MUL[i]`. MUL[i] is odd, so the multiply is a + bijection on 32 bits. rk is the 32-bit round key, the only thing that changes between applications. +2. Diffusion layer: one ChaCha-shaped double round. Four column quarter rounds with rotations ROT[0..3], then + four diagonal quarter rounds with ROT[4..7]. A quarter round is add, xor, rotate, four times. + +Per item, class v4 (`mixer_mult = 8`): init the state from the day key and `t`, then for each of 8 rounds apply +M eight times (keys `round_key(r*8 + j)`, j = 0..7) and do one dependent cache read; after the last read apply M +eight more times. 9 x 8 = 72 applications per item. Only the round key changes between the 72. ROT (8 values in +1..31), MUL (16 odd values), RC (16 values) are drawn once per day from one SplitMix64 stream seeded with +`K[0] | (K[1] << 32)`, K the day key. + +The day key is `seed_words_from_bytes("igneum-day/" || day_le64)` on the chain (interim rule, `bind::day_bytes`), +or `seed_words("day/" + iso)` in the spec's examples. The day is a pure function of the calendar day, so a weak +day is a public calendar. + +The round key is `round_key(k) = (k + 1) * 0x9E3779B9 mod 2^32`. The 72 keys are the first 72 odd-ish multiples +of 0x9E3779B9. They are fixed, not drawn. + +The cost model (chip-model-v3.md, read sections 1, 2, 5, 6): 130 ops per application hoisted, 9,360 ops per item, +1,198,080 ops per hash at m = 8. A shortcut is priced in ops per item against 9,360. + +## 2. The questions, in attack order + +The order is cheapest-reproducible first, then the structural questions. + +| Rank | Q | What a result looks like | +|---|---|---| +| 1 | Q3 weak parameter draws | counted fraction of days in each class over >= 2^24 day keys, per-day op gain | +| 2 | Q2 round margin | largest K applications a distinguisher reaches, against 8 and 72 | +| 3 | Q1 structural shortcut | an op count below 8x for the 8 keyed applications, or the bound | +| 4 | Q4 anything else | any other measured gain | + +## 3. Method per question + +### Q3, weak parameter draws (harness: f4-weakday, copied read-only into tools/attack/f4-weakday) + +The census walks consecutive chain days through the real draw code (`MixParams::with_shape`) and classifies each +day. Classes: ROT all equal, ROT distinct <= 3 or 4, ROT max multiplicity >= 4, ROT pair sums to 32 (same word +and any), ROT all or mostly in {1,2,30,31} or {8,16,24}; MUL any = 1, any = 2^32-1, any low popcount or low NAF +weight, any < 256, two equal, M2 class (NAF weight <= 3 frees a DSP); RC any = 0, RC + rk = 0 for any of the 72 +keys, RC extreme popcount, two equal. The gain metric M1 is the per-day LUT datapath cost in adder-equivalents, +`32 adds + 32 xors + sum(NAF(MUL_i) - 1)`, gain = census median cost over the day cost. M2 gain = 16/(16-k). + +Tool: `attack-f4 census --from 20729 --count 16777216 --threads 32 [--out file]`. Also `attack-f4 expect +--threads 32` for the exact analytic tail (NAF weight and popcount tables over all 2^31 odd constants, the +16-fold convolution), so the counted census is checked against the closed form. + +Gate: the plan's gain gate is 1.1x. Any class with a per-day gain at or above 1.1x on a non-negligible fraction +of days is a FINDING. A verifier is bit-exact and never skips an application, so ROT and RC values hand a +datapath 0 ops and are reported as structure, not as a wall-time gain. Only MUL weight moves the LUT cost. + +Known-failed shapes (the plant must fire): `attack-f4 plant alleq` (ROT all equal), `plant mul1` (one MUL = 1), +`plant mul1all` (all MUL = 1), `plant rc0` (one RC = 0), `plant rcrk0` (RC + rk = 0). Each prints the day 20729 +draw with the planted field and the detector must flag the matching class. + +### Q2, the round margin (harness: adv-mixer diffusion, new) + +The strict-avalanche census of K consecutive keyed applications, K = 1..12, over N random states. For each state, +flip each of 512 input bits, apply K applications, tally the output bit-flip probability p[in][out]. Report, per +K: dependency holes (p exactly 0 or 1), strong-bias cells (|p - 0.5| above 8 sigma), the worst cell and its +sigma. A distinguisher reaches K if a hole or a strong bias survives at K. The bound is the largest such K +against the 8 between reads and the 72 per item. + +Tool: `attack-adv-mixer diffusion --day 20729 --apps K --states 2000000 --threads 32` for K in 1..12. Also +`--start-app A` to confirm the margin does not depend on where in the 72 the window sits (keys differ). + +Gate: full diffusion (no hole, no strong bias at the census band) at K means the distinguisher does not reach K. +The margin is 8 - K_max between reads and 72 - K_max per item. + +Known-failed shape (the plant must fire): `--plant weak` builds a degenerate day by hand (MUL all 1, RC all 0, +ROT all 16). The detector must report many holes and strong bias at every K. A real day must not. + +The avalanche census is a bound, not a full trail search. It does not prove the absence of a high-order +differential or a linear trail below the census band. Its reach is N states: a bias under 8/sqrt(N) is invisible. +At N = 2e6, 8 sigma is about 0.0057, so a bias below 0.57 percent is not seen. This limit is stated with the +result. A SAT or MILP trail search to tighten Q2 is scoped as owed work, not run tonight. + +### Q1, the structural shortcut (harness: adv-mixer fold, new) + +Three probes on the 8 keyed applications where only rk changes. + +(a) The two multiply layers of adjacent applications do not merge. Test the two-application map g for GF(2) + affinity: for an affine g, `g(a) ^ g(b) ^ g(c) ^ g(a^b^c)` is constant. Count violations over N random + quadruples. Zero violations would mean g is affine and the two applications collapse to one linear map plus a + constant, a BREAK. Many violations is the bound: the diffusion between the two multiply layers is nonlinear, + so the multiplies do not fold. + +(b) Word separability. Flip each input word of the full 8-application block and record which output words move. + A dead (in_word, out_word) pair over all probes is a broken dependency a shortcut could split on. Zero dead + pairs is the bound. + +(c) Key-order commutation. Compare M(M(s,rk1),rk2) with M(M(s,rk2),rk1). Agreement would mean key order does not + matter and the 8 keys could be folded into fewer. Any agreement is a FINDING. + +Tool: `attack-adv-mixer fold --day 20729 --trials 1000000`. + +Gate: (a) at least one violation, (b) zero dead pairs, (c) zero agreements is the bound that the 8 keyed +applications cost 8x. Any breach is priced in ops per item against 9,360 and reported as a BREAK. + +The algebraic view (Q1 candidate 3): the multiply layer is `x -> (x ^ c) * MUL` per word, a bijection but not +GF(2)-linear (the integer multiply carries). The diffusion layer mixes words. So the composition over 8 +applications has rising algebraic degree. Probe (a) is the GF(2)-degree-1 test of the first two applications; a +pass there already rules out the cheapest fold. A full algebraic-degree or integral-distinguisher search is owed +work, scoped not run tonight. + +### Q4, anything else + +Two things to watch while the above runs. First, the round keys are fixed multiples of 0x9E3779B9, not drawn, so +a bad rk is the same every day: `round_key(k)` is checked in the self-test and the RC + rk = 0 class in f4 covers +the one way a fixed rk interacts with a drawn RC. Second, the item init `s[8+i] = t*MUL[i] + RC[i]` reuses MUL +and RC; a MUL[i] = 1 collapses that init word to `t + RC[i]`, which f4's mul1 class already counts. Any further +finding is added here. + +## 4. Known-failed shape per method (the plant each tool must fire on) + +| Method | Tool | Planted weakness | The tool must | +|---|---|---|---| +| Q3 census | attack-f4 plant alleq / mul1 / mul1all / rc0 / rcrk0 | the genesis day with one field forced weak | flag the matching class | +| Q2 diffusion | attack-adv-mixer diffusion --plant weak | MUL all 1, RC all 0, ROT all 16 | report holes and strong bias at every K | +| Q1 fold | (built in) | n/a: the probes are their own control, a real day must pass (a)-(c) | (a) violations > 0, (b) dead = 0, (c) agree = 0 | + +The plant discipline follows the Mac rule: a watcher is trusted only after it fires on a known-failed case. Every +run prints its plant state and its verdict. + +## 5. Box-hours per step + +Build box 2, core band 64-95, nice 10, one slot at a time. Budget 8 box-hours for first results. + +| Step | Command | Estimate | +|---|---|---| +| Build the two harnesses (release) | build-remote.sh --box 2 -- build --release | 0.15 box-hours | +| Q3 census 2^24 days, 32 threads | attack-f4 census --count 16777216 --threads 32 | 0.2 box-hours | +| Q3 analytic tail | attack-f4 expect --threads 32 | 0.3 box-hours | +| Q2 diffusion K = 1..12, 2e6 states each | attack-adv-mixer diffusion per K | 1.5 box-hours total | +| Q2 plant-weak firing check, K = 1..4 | attack-adv-mixer diffusion --plant weak | 0.1 box-hours | +| Q1 fold, 1e6 trials | attack-adv-mixer fold | 0.1 box-hours | +| Headroom for a wider census or a tighter K | | the rest | + +Estimates are first-cut from the op counts (one application is about 130 ops; 2e6 states x 512 flips x K +applications fits a 32-core band in minutes). The report records the box-hours actually spent. + +## 6. Files opened (the outsider read set) + +Only these were read. Nothing else in the repository. + +1. igneum-pow/src/memhard.rs (frozen, via git show at 017e7037). +2. igneum-pow/src/seed.rs (frozen). +3. igneum-pow/src/bind.rs (frozen, the day_bytes and day_index functions). +4. igneum-pow/src/generator.rs (frozen, the LoadClass, V3_CLASS, V4_CLASS, ProgramClass, generator version and + attempt-cap definitions; grepped, not read whole). +5. igneum-pow/tests/mixer.rs (frozen, the head: the test harness contract on the class v3 and v4 construction). +6. igneum-pow/Cargo.toml and Cargo.lock (dependency pin). +7. docs/spec/01-lottery-hash.md section 1.8 (frozen: 1.8.1 to 1.8.5). +8. docs/analysis/chip-model-v3.md sections 1, 2, 5, 6 (HEAD). +9. The public kit packs under proto-cuda/packs-ca3-v4 (frozen, the file listing; the eight packs named in the + brief). The kit zip sha256 is 4f2445c50c58d76a5544023492d8b858d0b07c5e372d31f9c90c4ce51f829154 per the brief, + to be checked when a pack is used as a vector. +10. The Devnet 3 epoch-0 pack v4-devnet3-epoch0 (public-kit class, named by a teammate lane): on build-1 at + /srv/artefacts/packs/v4-devnet3-epoch0/, zip sha256 + e025750f71175ed14d6e2a24e387ebbf1979b1cd0faee9139c41a7671165b334, program id 0xfce15bf61030be57 at attempt 0, + 2^24 fingerprint from base 0 e510ad92b4d24846, day bytes le64(20733). Copied read-only and the sha verified + before any use as a vector. This matches the Devnet 3 epoch-0 program id named in the brief as the check. +11. tools/attack/f4-weakday (build/attack-pass, copied read-only) and tools/attack/f8-uniform (the Mac worktree, + copied with cp -R, original untouched). +12. tools/build-remote.sh, infra/build-server/lib.sh, infra/build-server/remote-run.sh (the operating files). +13. CLAUDE.md (loaded on its own; only its operating rules are followed, not its doc references). + +## 7. What is owed beyond tonight + +- A SAT or MILP differential and linear trail search on M_r reduced to K applications, to tighten Q2 below the + avalanche census band. +- A rotational-XOR search on the drawn double round. +- An algebraic-degree or integral-distinguisher measurement across the 8 applications, to tighten Q1 beyond the + affinity probe. +- The census at more than 2^24 days if any class sits near the gate. diff --git a/tools/attack/adv-mixer/Cargo.toml b/tools/attack/adv-mixer/Cargo.toml new file mode 100644 index 000000000..d84e9727c --- /dev/null +++ b/tools/attack/adv-mixer/Cargo.toml @@ -0,0 +1,21 @@ +[package] +name = "attack-adv-mixer" +version = "0.1.0" +edition = "2021" +description = "Adversarial cryptanalysis of the memory-hard mixer M_r (spec 01 section 1.8.4): diffusion margin across the keyed applications (Q2), the composition-fold probe (Q1), with the planted-weak-day hooks that prove the harness fires" +license = "MIT" +publish = false + +[[bin]] +name = "attack-adv-mixer" +path = "src/main.rs" + +[dependencies] +igneum-pow = { path = "../../../igneum-pow" } + +[workspace] + +[profile.release] +opt-level = 3 +lto = true +codegen-units = 1 diff --git a/tools/attack/adv-mixer/src/main.rs b/tools/attack/adv-mixer/src/main.rs new file mode 100644 index 000000000..229be22cc --- /dev/null +++ b/tools/attack/adv-mixer/src/main.rs @@ -0,0 +1,370 @@ +//! attack-adv-mixer: adversarial cryptanalysis of the memory-hard mixer `M_r` (spec 01 section 1.8.4), the +//! internal adversarial pass, not an independent review. Every number here comes from the real `igneum-pow` +//! mixer (`memhard::mixer`, `round_key_mult`, `MixParams::with_shape`); nothing is re-implemented. +//! +//! The target in the attacker's words. One application `M(s, rk)` on a 16-word state: +//! 1. prologue, per word: `s[i] = (s[i] XOR (RC[i] + rk)) * MUL[i]` (MUL odd, so each is a bijection). +//! 2. one ChaCha-shaped double round: four column quarter rounds with rotations ROT[0..3], four diagonal +//! quarter rounds with ROT[4..7]. +//! Under class v4 (`mixer_mult = 8`) the derivation applies `M` eight times between each pair of the eight +//! dependent cache reads, round keys `round_key(r*8 + j)`, then eight more after the last read: 72 per item. +//! ROT, MUL, RC are drawn once per day from one SplitMix64 stream (the day key's first two words). +//! +//! The commands, each a BOUND or a BREAK with a command and a seed: +//! +//! diffusion --apps K --states N --day D [--start-app A] [--threads T] [--plant weak|none] +//! The strict-avalanche census of K consecutive keyed applications (Q2). For N random states it flips each +//! of the 512 input bits, applies K applications (keys from app A in derive order), and tallies the output +//! bit-flip probability p[in][out] over the N states. Reports, per K, the number of output bits a single +//! input bit never reaches (dependency holes), the number of (in,out) cells with |p - 0.5| above the +//! census bias band, and the worst cell. Full diffusion at K is the bound: the largest K at which a hole +//! or a strong bias survives is the distinguisher reach. The `weak` plant is a hand-built degenerate day +//! (MUL all 1, RC all 0, ROT all 16): the detector must fire on it (holes and strong bias at every K). +//! +//! fold --day D [--trials N] +//! The composition-shortcut probe (Q1). Checks three ways the 8 keyed applications might cost less than 8x: +//! (a) the two multiply layers of adjacent applications do not merge: M has a nonlinear diffusion between +//! them, shown by a GF(2) affinity test of the two-application map on N random probes (an affine map +//! satisfies f(a)+f(b)+f(c)=f(a+b+c); count violations). (b) word separability: does output word w +//! depend on every input word, tested by flipping each input word and checking each output word moves. +//! A shortcut needs a broken dependency. (c) key-only commutation: M(M(s,rk1),rk2) vs M(M(s,rk2),rk1); +//! if they agreed the key order would not matter and keys could be folded. Prints the counts; a zero +//! in (a) or a missing dependency in (b) or an agreement in (c) would be a BREAK, else the BOUND. +//! +//! The genesis test vector (day 2026-10-03) is checked at startup against the spec so the mixer wiring is the +//! library's. + +use igneum_pow::generator::V4_CLASS; +use igneum_pow::memhard::{mixer, round_key_mult, MixParams, Shape, ITEM_ROUNDS}; +use igneum_pow::seed::{day_key, seed_words_from_bytes, SplitMix64}; +use igneum_pow::bind::day_bytes; +use std::sync::atomic::{AtomicU64, Ordering}; +use std::sync::Arc; +use std::thread; + +/// The chain genesis day index (`bind.rs`: day_index(0x1a0ff0f7c00) = 20,729, 3 October 2026). +const GENESIS_DAY: u64 = 20_729; + +fn v4_shape() -> Shape { + let s = Shape::for_class(&V4_CLASS); + assert_eq!(s.mixer_mult, 8, "class v4 is the x8 mixer"); + assert_eq!(s.derive_len, 0, "class v4 has no derivation program"); + s +} + +/// Params for a chain day index (the interim day rule, `bind::day_bytes`). +fn params_of_day(d: u64) -> MixParams { + MixParams::with_shape(seed_words_from_bytes(&day_bytes(d)), v4_shape()) +} + +/// A hand-built degenerate day: identity multiply, zero round constants, one rotation amount everywhere. Not a +/// drawable day (it is the plant); every field is in range (MUL odd, ROT in 1..31). +fn planted_weak(d: u64) -> MixParams { + let mut mp = params_of_day(d); + mp.mul = [1u32; 16]; + mp.rc = [0u32; 16]; + mp.rot = [16u32; 8]; + mp +} + +/// The 72 application round keys of an item under m = 8, in derive order. +fn app_keys() -> [u32; 72] { + let mut k = [0u32; 72]; + for r in 0..=ITEM_ROUNDS { + for j in 0..8 { + k[r * 8 + j] = round_key_mult(r, j, 8); + } + } + k +} + +/// Apply `apps` keyed applications starting at application index `start` (derive order). +#[inline] +fn apply_n(mut s: [u32; 16], mp: &MixParams, keys: &[u32; 72], start: usize, apps: usize) -> [u32; 16] { + for a in 0..apps { + mixer(&mut s, keys[(start + a) % 72], mp); + } + s +} + +/// A per-thread SplitMix64 PRNG for random probe states (the attacker's own randomness, not the mixer's). +struct Rng(SplitMix64); +impl Rng { + fn new(seed: u64) -> Self { + Rng(SplitMix64::new(seed)) + } + fn state(&mut self) -> [u32; 16] { + let mut s = [0u32; 16]; + for w in s.iter_mut() { + *w = self.0.next() as u32; + } + s + } +} + +// -------------------------------------------------------------------------------------------------------------- +// diffusion (Q2): the strict-avalanche census over K keyed applications +// -------------------------------------------------------------------------------------------------------------- + +/// 512 x 512 counters (input bit -> output bit flip count), summed as u64. Flat for cache behaviour. +struct Aval { + n: u64, + counts: Vec, // 512*512 +} +impl Aval { + fn new() -> Self { + Aval { n: 0, counts: vec![0u64; 512 * 512] } + } + fn merge(&mut self, o: &Aval) { + self.n += o.n; + for (a, b) in self.counts.iter_mut().zip(o.counts.iter()) { + *a += b; + } + } +} + +#[inline] +fn bit_of(s: &[u32; 16], b: usize) -> u32 { + (s[b >> 5] >> (b & 31)) & 1 +} +#[inline] +fn flip_bit(s: &mut [u32; 16], b: usize) { + s[b >> 5] ^= 1u32 << (b & 31); +} + +fn diffusion(day: u64, plant_weak: bool, apps: usize, states: u64, start: usize, threads: usize) { + let mp = if plant_weak { planted_weak(day) } else { params_of_day(day) }; + let keys = app_keys(); + let per = states / threads as u64; + let mp = Arc::new(mp); + let keys = Arc::new(keys); + let mut handles = Vec::new(); + for t in 0..threads { + let mp = Arc::clone(&mp); + let keys = Arc::clone(&keys); + let count = if t as u64 == threads as u64 - 1 { states - per * (threads as u64 - 1) } else { per }; + let seed = 0x1234_5678_9abc_def0 ^ ((day as u64).wrapping_mul(0x9E3779B97F4A7C15)) ^ (t as u64 + 1); + handles.push(thread::spawn(move || { + let mut rng = Rng::new(seed); + let mut acc = Aval::new(); + for _ in 0..count { + let base = rng.state(); + let out0 = apply_n(base, &mp, &keys, start, apps); + for ib in 0..512 { + let mut s = base; + flip_bit(&mut s, ib); + let out1 = apply_n(s, &mp, &keys, start, apps); + let row = ib * 512; + for ob in 0..512 { + if bit_of(&out0, ob) != bit_of(&out1, ob) { + acc.counts[row + ob] += 1; + } + } + } + acc.n += 1; + } + acc + })); + } + let mut total = Aval::new(); + for h in handles { + total.merge(&h.join().unwrap()); + } + + // census band: a cell at the ideal 0.5 over N states has stddev 0.5/sqrt(N); call a bias "strong" at 8 sigma. + let n = total.n as f64; + let sigma = 0.5 / n.sqrt(); + let band = 8.0 * sigma; + let mut holes = 0u64; // cells with p == 0 or p == 1 exactly (a dependency hole or a perfect relation) + let mut strong = 0u64; // |p-0.5| > band and not a hole + let mut worst_dev = 0.0f64; + let mut worst = (0usize, 0usize, 0.0f64); + let mut global_flips = 0u64; + for ib in 0..512 { + for ob in 0..512 { + let c = total.counts[ib * 512 + ob]; + global_flips += c; + let p = c as f64 / n; + if c == 0 || c == total.n { + holes += 1; + } else { + let dev = (p - 0.5).abs(); + if dev > band { + strong += 1; + } + if dev > worst_dev { + worst_dev = dev; + worst = (ib, ob, p); + } + } + } + } + let mean_p = global_flips as f64 / (n * 512.0 * 512.0); + println!( + "diffusion day={} plant={} apps={} start={} states={} threads={}", + day, + if plant_weak { "weak" } else { "none" }, + apps, + start, + total.n, + threads + ); + println!(" mean output-flip probability over all 262144 cells: {:.6} (ideal 0.5)", mean_p); + println!(" census band: 8 sigma = {:.6} ({} states)", band, total.n); + println!(" dependency holes (p == 0 or p == 1 exactly): {} of 262144", holes); + println!(" strong-bias cells (|p-0.5| > band, not a hole): {} of 262144", strong); + println!( + " worst cell: in_bit {} -> out_bit {} p = {:.6} dev = {:.6} ({:.1} sigma)", + worst.0, + worst.1, + worst.2, + worst_dev, + worst_dev / sigma + ); + let verdict = if holes > 0 || strong > 0 { "FINDING (holes or strong bias at this K)" } else { "no distinguisher at this K" }; + println!(" VERDICT: {}", verdict); +} + +// -------------------------------------------------------------------------------------------------------------- +// fold (Q1): the composition-shortcut probe +// -------------------------------------------------------------------------------------------------------------- + +#[inline] +fn xor16(a: &[u32; 16], b: &[u32; 16]) -> [u32; 16] { + let mut o = [0u32; 16]; + for i in 0..16 { + o[i] = a[i] ^ b[i]; + } + o +} + +fn fold(day: u64, trials: u64) { + let mp = params_of_day(day); + let keys = app_keys(); + let rk1 = keys[0]; + let rk2 = keys[1]; + let mut rng = Rng::new(0xfeed_face_cafe_babe ^ day.wrapping_mul(0x9E3779B97F4A7C15)); + + // (a) GF(2) affinity of the two-application map g(s) = M(M(s,rk1),rk2). An affine map over GF(2)^512 obeys + // g(a) ^ g(b) ^ g(c) ^ g(a^b^c) = g(0^...) summed; exactly, g(a)^g(b)^g(c)^g(a^b^c) is constant for an + // affine g. We test the 4-point relation g(a)^g(b)^g(c)^g(a^b^c) == g(d0)^g(d0)... using the zero anchor: + // for affine g, g(a)^g(b)^g(c)^g(a^b^c) == g(0) (four points a,b,c,a^b^c vs the origin). Count nonzero. + let g = |s: [u32; 16]| -> [u32; 16] { + let s1 = apply_n(s, &mp, &keys, 0, 1); + apply_n(s1, &mp, &keys, 1, 1) + }; + let g0 = g([0u32; 16]); + let mut affine_violations = 0u64; + for _ in 0..trials { + let a = rng.state(); + let b = rng.state(); + let c = rng.state(); + let abc = xor16(&xor16(&a, &b), &c); + let lhs = xor16(&xor16(&g(a), &g(b)), &xor16(&g(c), &g(abc))); + if lhs != g0 { + affine_violations += 1; + } + } + + // (b) word separability: flip each input word fully (xor 0xffffffff) and see whether every output word of the + // full 8-application block moves on at least one probe. A dead (in_word -> out_word) pair over all probes + // is a broken dependency a shortcut could exploit. + let full = |s: [u32; 16]| apply_n(s, &mp, &keys, 0, 8); + let mut dep = [[false; 16]; 16]; // dep[iw][ow] = out word ow ever changed when in word iw flipped + for _ in 0..trials { + let base = rng.state(); + let o0 = full(base); + for iw in 0..16 { + let mut s = base; + s[iw] ^= 0xffff_ffff; + let o1 = full(s); + for ow in 0..16 { + if o0[ow] != o1[ow] { + dep[iw][ow] = true; + } + } + } + } + let mut dead_pairs = 0u64; + for iw in 0..16 { + for ow in 0..16 { + if !dep[iw][ow] { + dead_pairs += 1; + } + } + } + + // (c) key-order commutation: M(M(s,rk1),rk2) vs M(M(s,rk2),rk1). Agreement would let keys fold. + let mut commute_agree = 0u64; + for _ in 0..trials { + let s = rng.state(); + let ab = apply_n(apply_n(s, &mp, &keys, 0, 1), &mp, &keys, 1, 1); + // apply rk2 then rk1 by temporarily swapping via explicit keys + let mut s1 = s; + mixer(&mut s1, rk2, &mp); + mixer(&mut s1, rk1, &mp); + if ab == s1 { + commute_agree += 1; + } + } + + println!("fold day={} trials={}", day, trials); + println!(" (a) GF(2) affinity violations of the 2-application map: {} of {} (0 would be a BREAK: the map is affine)", affine_violations, trials); + println!(" (b) dead (in_word -> out_word) pairs over the full 8-application block: {} of 256 (any would be a broken dependency)", dead_pairs); + println!(" (c) key-order agreements M(M(.,rk1),rk2) == M(M(.,rk2),rk1): {} of {} (any would let keys fold)", commute_agree, trials); + let verdict = if affine_violations == 0 || dead_pairs > 0 || commute_agree > 0 { "FINDING" } else { "BOUND: no fold on these probes" }; + println!(" VERDICT: {}", verdict); +} + +// -------------------------------------------------------------------------------------------------------------- +// startup self-check: the genesis test vector of the spec +// -------------------------------------------------------------------------------------------------------------- + +fn self_check() { + let mp = MixParams::for_day("2026-10-03"); + assert_eq!(mp.rot, [20, 20, 19, 4, 26, 3, 3, 27], "spec 1.8.4 ROT vector"); + assert_eq!(mp.mul[0], 0x42146205, "spec 1.8.4 MUL[0]"); + assert_eq!(mp.mul[15], 0x99cfb423, "spec 1.8.4 MUL[15]"); + assert_eq!(mp.rc[0], 0xbab68293, "spec 1.8.4 RC[0]"); + assert_eq!(mp.rc[15], 0x31b49ee2, "spec 1.8.4 RC[15]"); + let k = day_key("2026-10-03"); + assert_eq!(k[0], 0x3067619f, "spec 1.8.1 day key K[0]"); + // the 72 application keys are distinct multiples of 0x9E3779B9 + let keys = app_keys(); + for (i, &v) in keys.iter().enumerate() { + assert_eq!(v, ((i as u32) + 1).wrapping_mul(0x9E3779B9), "application key {i}"); + } + eprintln!("self-check: spec 1.8.1 and 1.8.4 genesis vectors OK; 72 application keys are 1..72 times 0x9E3779B9"); +} + +fn arg_u64(args: &[String], flag: &str, default: u64) -> u64 { + args.iter().position(|a| a == flag).and_then(|i| args.get(i + 1)).and_then(|s| s.parse().ok()).unwrap_or(default) +} +fn arg_str<'a>(args: &'a [String], flag: &str, default: &'a str) -> &'a str { + args.iter().position(|a| a == flag).and_then(|i| args.get(i + 1)).map(|s| s.as_str()).unwrap_or(default) +} + +fn main() { + self_check(); + let args: Vec = std::env::args().collect(); + let cmd = args.get(1).map(|s| s.as_str()).unwrap_or("help"); + let day = arg_u64(&args, "--day", GENESIS_DAY); + let threads = arg_u64(&args, "--threads", 16).max(1) as usize; + match cmd { + "diffusion" => { + let apps = arg_u64(&args, "--apps", 8) as usize; + let states = arg_u64(&args, "--states", 100_000); + let start = arg_u64(&args, "--start-app", 0) as usize; + let plant = arg_str(&args, "--plant", "none") == "weak"; + diffusion(day, plant, apps, states, start, threads); + } + "fold" => { + let trials = arg_u64(&args, "--trials", 100_000); + fold(day, trials); + } + _ => { + eprintln!("usage: attack-adv-mixer diffusion|fold [--day D] [--apps K] [--states N] [--start-app A] [--plant weak|none] [--trials N] [--threads T]"); + } + } + let _ = AtomicU64::new(0).fetch_add(0, Ordering::Relaxed); +} diff --git a/tools/attack/f4-weakday/Cargo.lock b/tools/attack/f4-weakday/Cargo.lock new file mode 100644 index 000000000..7921217c2 --- /dev/null +++ b/tools/attack/f4-weakday/Cargo.lock @@ -0,0 +1,14 @@ +# This file is automatically @generated by Cargo. +# It is not intended for manual editing. +version = 4 + +[[package]] +name = "attack-f4-weakday" +version = "0.1.0" +dependencies = [ + "igneum-pow", +] + +[[package]] +name = "igneum-pow" +version = "0.2.0" diff --git a/tools/attack/f4-weakday/Cargo.toml b/tools/attack/f4-weakday/Cargo.toml new file mode 100644 index 000000000..28721f6c7 --- /dev/null +++ b/tools/attack/f4-weakday/Cargo.toml @@ -0,0 +1,21 @@ +[package] +name = "attack-f4-weakday" +version = "0.1.0" +edition = "2021" +description = "Attack pass F4: the weak-day census over 2^24 day keys through MixParams::with_shape (docs/plans/cryptanalysis.md 4.2)" +license = "MIT" +publish = false + +[[bin]] +name = "attack-f4" +path = "src/main.rs" + +[dependencies] +igneum-pow = { path = "../../../igneum-pow" } + +[workspace] + +[profile.release] +opt-level = 3 +lto = true +codegen-units = 1 diff --git a/tools/attack/f4-weakday/src/main.rs b/tools/attack/f4-weakday/src/main.rs new file mode 100644 index 000000000..4d775102a --- /dev/null +++ b/tools/attack/f4-weakday/src/main.rs @@ -0,0 +1,990 @@ +//! Attack pass F4: the weak-day census (`docs/plans/cryptanalysis.md` section 4.2 row F4, `funding.md` B2 rank 3 and +//! B5 rank 3). Every chain day `d` has the key `seed_words_from_bytes("igneum-day/" || d_le64)` (`bind::day_bytes`, +//! the interim day rule) and its mixer constants come from `MixParams::with_shape`, a SplitMix64 stream seeded from +//! `key[0] | key[1] << 32`: `ROT[0..7]` in 1..31, `MUL[0..15]` odd, `RC[0..15]`. This harness walks consecutive chain +//! days through the real draw code (the `igneum-pow` path dependency, nothing re-implemented) and classifies each +//! day's draw. +//! +//! The gain metrics, all exact and structural (what a datapath built for the day pays, in adder-equivalents per +//! mixer application; the verifier and every GPU pay the same ops on every day, so wall time cannot move): +//! +//! * M1, the per-day LUT datapath (an FPGA bitstream synthesised for the day, the only per-day attacker that +//! exists: a constant XOR is absorbed into the next LUT, a rotation by a constant is routing, a 32-bit add or +//! XOR is one 32-bit adder-equivalent, a multiply by a constant is `NAF(MUL) - 1` adders in canonical signed-digit +//! shift-add form): `cost = 32 adds + 32 xors + sum_i (naf(MUL_i) - 1)`. Gain of a day = the census median cost +//! over the day's cost. This is the generous bound: optimal single-constant multiplication is cheaper than NAF for +//! every constant, and a DSP-block multiply does not depend on the value at all. +//! * M2, the DSP-bound datapath (the multiplies in DSP blocks, value-independent): a word whose constant has NAF +//! weight at most 3 (two adders) moves to LUTs and frees its DSP, so gain = 16 / (16 - k) for k such words. +//! * ROT and RC classes: a constant rotation is wiring and a constant XOR is inverters on a per-day datapath, so +//! their value hands a datapath exactly 0 ops. They are censused as structure (counts against the analytic +//! expectation) and the worst members are measured for diffusion (`avalanche`), which bounds the only other +//! thing a rotation draw could move. A bit-exact verifier never lets a chip skip an application, however weak its +//! diffusion, so diffusion is reported and is not a gain. +//! +//! Commands: +//! attack-f4 census --from 20729 --count 16777216 --threads 12 [--out file] [--dedupe] +//! attack-f4 day --index 20729 one day's draw and classification +//! attack-f4 plant alleq|mul1|mul1all|rc0|rcrk0 the known-fail firings: the day 20729 draw with the planted field +//! attack-f4 expect --threads 12 exact per-word tables (NAF weight, popcount) over all 2^31 odd +//! constants, and the 16-fold convolution: the expected M1 tail +//! attack-f4 avalanche --index 20729 [--states 4096] single-application and two-application diffusion of a day + +use igneum_pow::bind::day_bytes; +use igneum_pow::generator::V4_CLASS; +use igneum_pow::memhard::{mixer, round_key_mult, MixParams, Shape, ITEM_ROUNDS}; +use igneum_pow::seed::{seed_words_from_bytes, SplitMix64}; +use std::fmt::Write as _; +use std::io::Write as _; + +/// The chain's genesis day index (`bind.rs` tests: day_index(0x1a0ff0f7c00) = 20,729, 3 October 2026). +const GENESIS_DAY: u64 = 20_729; +/// Mixer applications per item under class v4 (`Shape::mixers_per_item`): 9 x 8. +const APPLICATIONS_PER_ITEM: u64 = (ITEM_ROUNDS as u64 + 1) * 8; +/// Adds and XORs of one application outside the multiply layer: 8 quarter rounds x (4 adds + 4 xors). +const QR_ADDS_XORS: u32 = 64; +/// The gate's gain threshold (plan 1.4 (3), B5 rank 3). +const GAIN_GATE: f64 = 1.1; +/// M2: a constant of NAF weight at most this is cheaper in LUTs than in a DSP block. +const M2_NAF_CEIL: u32 = 3; + +// ------------------------------------------------------------------------------------------------------------ +// The day +// ------------------------------------------------------------------------------------------------------------ + +fn v4_shape() -> Shape { + let s = Shape::for_class(&V4_CLASS); + assert_eq!(s.mixer_mult, 8, "class v4 is the x8 mixer"); + assert_eq!(s.derive_len, 0, "class v4 has no derivation program: the fixed mixer with drawn constants"); + s +} + +/// The chain day's key and its draw through the real code. +fn params_of_day(d: u64) -> MixParams { + MixParams::with_shape(seed_words_from_bytes(&day_bytes(d)), v4_shape()) +} + +fn seed64(key: &[u32; 8]) -> u64 { + key[0] as u64 | ((key[1] as u64) << 32) +} + +/// Non-adjacent-form weight of a 32-bit constant (the number of nonzero signed digits). +fn naf_weight(v: u32) -> u32 { + let mut n = v as u64; + let mut w = 0; + while n != 0 { + if n & 1 == 1 { + // digit +1 when n = 1 mod 4, -1 when n = 3 mod 4 + if n & 3 == 3 { + n += 1; + } else { + n -= 1; + } + w += 1; + } + n >>= 1; + } + w +} + +/// Round keys of the 72 applications of an item under m = 8: `round_key(r * 8 + j)`, r in 0..=8, j in 0..8. +fn round_keys() -> [u32; 72] { + let mut k = [0u32; 72]; + for r in 0..=ITEM_ROUNDS { + for j in 0..8 { + k[r * 8 + j] = round_key_mult(r, j, 8); + } + } + k +} + +#[derive(Clone, Debug, Default)] +struct DayClass { + // ROT + rot_distinct: u32, + rot_all_equal: bool, + rot_max_mult: u32, + rot_comp_same_word: bool, + rot_comp_any: bool, + rot_small: u32, + rot_byte: u32, + // MUL + naf: [u32; 16], + naf_sum: u32, + naf_min: u32, + mul_one: u32, + mul_minus_one: u32, + mul_pop_le2: u32, + mul_pop_le4: u32, + mul_pop_le8: u32, + mul_naf_le2: u32, + mul_naf_le3: u32, + mul_naf_le4: u32, + mul_small: u32, + mul_dup: bool, + // RC + rc_zero: u32, + rc_pop_ext: u32, + rc_rk_zero: u32, + rc_dup: bool, + // gains + cost_m1: u32, + m2_k: u32, +} + +fn classify(mp: &MixParams, rks: &[u32; 72]) -> DayClass { + let mut c = DayClass::default(); + // ROT + let mut seen = [0u32; 32]; + for &r in &mp.rot { + assert!((1..=31).contains(&r)); + seen[r as usize] += 1; + if matches!(r, 1 | 2 | 30 | 31) { + c.rot_small += 1; + } + if matches!(r, 8 | 16 | 24) { + c.rot_byte += 1; + } + } + c.rot_distinct = seen.iter().filter(|&&n| n > 0).count() as u32; + c.rot_max_mult = *seen.iter().max().unwrap(); + c.rot_all_equal = c.rot_distinct == 1; + // the same-word pairs of a quarter round: s[d] takes ROT[0] then ROT[2], s[b] takes ROT[1] then ROT[3]; the + // diagonal round the same with ROT[4..7] + for (a, b) in [(0, 2), (1, 3), (4, 6), (5, 7)] { + if mp.rot[a] + mp.rot[b] == 32 { + c.rot_comp_same_word = true; + } + } + for a in 0..8 { + for b in a + 1..8 { + if mp.rot[a] + mp.rot[b] == 32 { + c.rot_comp_any = true; + } + } + } + // MUL + c.naf_min = u32::MAX; + for i in 0..16 { + let m = mp.mul[i]; + assert!(m & 1 == 1); + let w = naf_weight(m); + c.naf[i] = w; + c.naf_sum += w; + c.naf_min = c.naf_min.min(w); + let p = m.count_ones(); + if m == 1 { + c.mul_one += 1; + } + if m == u32::MAX { + c.mul_minus_one += 1; + } + if p <= 2 { + c.mul_pop_le2 += 1; + } + if p <= 4 { + c.mul_pop_le4 += 1; + } + if p <= 8 { + c.mul_pop_le8 += 1; + } + if w <= 2 { + c.mul_naf_le2 += 1; + } + if w <= 3 { + c.mul_naf_le3 += 1; + } + if w <= 4 { + c.mul_naf_le4 += 1; + } + if m < 256 { + c.mul_small += 1; + } + for j in 0..i { + if mp.mul[j] == m { + c.mul_dup = true; + } + } + } + c.cost_m1 = QR_ADDS_XORS + c.naf_sum - 16; + c.m2_k = c.mul_naf_le3; + // RC + for i in 0..16 { + let r = mp.rc[i]; + if r == 0 { + c.rc_zero += 1; + } + let p = r.count_ones(); + if p <= 4 || p >= 28 { + c.rc_pop_ext += 1; + } + for &rk in rks.iter() { + if r.wrapping_add(rk) == 0 { + c.rc_rk_zero += 1; + } + } + for j in 0..i { + if mp.rc[j] == r { + c.rc_dup = true; + } + } + } + c +} + +fn m2_gain(k: u32) -> f64 { + 16.0 / (16.0 - k as f64) +} + +// ------------------------------------------------------------------------------------------------------------ +// The tally +// ------------------------------------------------------------------------------------------------------------ + +/// The worst member of a class: the lowest M1 cost among the days in it (ties: the earliest day). +#[derive(Clone, Copy, Debug)] +struct Worst { + day: u64, + cost: u32, +} + +impl Worst { + fn none() -> Self { + Worst { day: u64::MAX, cost: u32::MAX } + } + fn offer(&mut self, day: u64, cost: u32) { + if cost < self.cost || (cost == self.cost && day < self.day) { + *self = Worst { day, cost }; + } + } + fn merge(&mut self, o: &Worst) { + if o.day != u64::MAX { + self.offer(o.day, o.cost); + } + } +} + +const CLASSES: &[&str] = &[ + "ROT all equal", + "ROT distinct <= 3", + "ROT distinct <= 4", + "ROT max multiplicity >= 4", + "ROT same-word pair sums to 32", + "ROT any pair sums to 32", + "ROT all 8 in {1,2,30,31}", + "ROT >= 6 in {1,2,30,31}", + "ROT >= 4 in {1,2,30,31}", + "ROT >= 4 in {8,16,24}", + "MUL any = 1", + "MUL any = 2^32-1", + "MUL any popcount <= 2", + "MUL any popcount <= 4", + "MUL any popcount <= 8", + "MUL any NAF weight <= 2", + "MUL any NAF weight <= 3", + "MUL any NAF weight <= 4", + "MUL any < 256", + "MUL two equal", + "MUL M2 k >= 2 (gain >= 1.14x)", + "RC any = 0", + "RC any popcount <= 4 or >= 28", + "RC + rk = 0 for any of the 72 keys", + "RC two equal", +]; + +#[derive(Clone, Debug)] +struct Tally { + n: u64, + class_count: Vec, + class_worst: Vec, + cost_hist: Vec, + naf_sum_hist: Vec, + naf_min_hist: Vec, + rot_distinct_hist: [u64; 9], + rot_small_hist: [u64; 9], + rot_byte_hist: [u64; 9], + rot_mult_hist: [u64; 9], + m2_k_hist: [u64; 17], + /// The 16 lowest-cost days seen (cost, day), sorted ascending. + lowest: Vec<(u32, u64)>, + highest: Vec<(u32, u64)>, + seeds: Vec, +} + +impl Tally { + fn new(dedupe: bool, cap: usize) -> Self { + Tally { + n: 0, + class_count: vec![0; CLASSES.len()], + class_worst: vec![Worst::none(); CLASSES.len()], + cost_hist: vec![0; 400], + naf_sum_hist: vec![0; 400], + naf_min_hist: vec![0; 40], + rot_distinct_hist: [0; 9], + rot_small_hist: [0; 9], + rot_byte_hist: [0; 9], + rot_mult_hist: [0; 9], + m2_k_hist: [0; 17], + lowest: Vec::new(), + highest: Vec::new(), + seeds: if dedupe { Vec::with_capacity(cap) } else { Vec::new() }, + } + } + + fn flags(c: &DayClass) -> [bool; 25] { + [ + c.rot_all_equal, + c.rot_distinct <= 3, + c.rot_distinct <= 4, + c.rot_max_mult >= 4, + c.rot_comp_same_word, + c.rot_comp_any, + c.rot_small == 8, + c.rot_small >= 6, + c.rot_small >= 4, + c.rot_byte >= 4, + c.mul_one > 0, + c.mul_minus_one > 0, + c.mul_pop_le2 > 0, + c.mul_pop_le4 > 0, + c.mul_pop_le8 > 0, + c.mul_naf_le2 > 0, + c.mul_naf_le3 > 0, + c.mul_naf_le4 > 0, + c.mul_small > 0, + c.mul_dup, + c.m2_k >= 2, + c.rc_zero > 0, + c.rc_pop_ext > 0, + c.rc_rk_zero > 0, + c.rc_dup, + ] + } + + fn add(&mut self, day: u64, mp: &MixParams, c: &DayClass, dedupe: bool) { + self.n += 1; + let f = Self::flags(c); + assert_eq!(f.len(), CLASSES.len()); + for (i, &on) in f.iter().enumerate() { + if on { + self.class_count[i] += 1; + self.class_worst[i].offer(day, c.cost_m1); + } + } + self.cost_hist[c.cost_m1 as usize] += 1; + self.naf_sum_hist[c.naf_sum as usize] += 1; + self.naf_min_hist[c.naf_min as usize] += 1; + self.rot_distinct_hist[c.rot_distinct as usize] += 1; + self.rot_small_hist[c.rot_small as usize] += 1; + self.rot_byte_hist[c.rot_byte as usize] += 1; + self.rot_mult_hist[c.rot_max_mult as usize] += 1; + self.m2_k_hist[c.m2_k as usize] += 1; + push_sorted(&mut self.lowest, (c.cost_m1, day), 16, true); + push_sorted(&mut self.highest, (c.cost_m1, day), 4, false); + if dedupe { + self.seeds.push(seed64(&mp.key)); + } + } + + fn merge(&mut self, o: Tally) { + self.n += o.n; + for i in 0..CLASSES.len() { + self.class_count[i] += o.class_count[i]; + self.class_worst[i].merge(&o.class_worst[i]); + } + for (a, b) in self.cost_hist.iter_mut().zip(o.cost_hist.iter()) { + *a += b; + } + for (a, b) in self.naf_sum_hist.iter_mut().zip(o.naf_sum_hist.iter()) { + *a += b; + } + for (a, b) in self.naf_min_hist.iter_mut().zip(o.naf_min_hist.iter()) { + *a += b; + } + for i in 0..9 { + self.rot_distinct_hist[i] += o.rot_distinct_hist[i]; + self.rot_small_hist[i] += o.rot_small_hist[i]; + self.rot_byte_hist[i] += o.rot_byte_hist[i]; + self.rot_mult_hist[i] += o.rot_mult_hist[i]; + } + for i in 0..17 { + self.m2_k_hist[i] += o.m2_k_hist[i]; + } + for e in o.lowest { + push_sorted(&mut self.lowest, e, 16, true); + } + for e in o.highest { + push_sorted(&mut self.highest, e, 4, false); + } + self.seeds.extend(o.seeds); + } +} + +/// Keep the `cap` smallest (ascending) or largest (descending) entries. +fn push_sorted(v: &mut Vec<(u32, u64)>, e: (u32, u64), cap: usize, ascending: bool) { + if v.len() == cap { + let last = *v.last().unwrap(); + let keep = if ascending { e < last } else { e > last }; + if !keep { + return; + } + v.pop(); + } + let pos = if ascending { v.partition_point(|x| *x < e) } else { v.partition_point(|x| *x > e) }; + v.insert(pos, e); +} + +// ------------------------------------------------------------------------------------------------------------ +// Analytic expectations (per day, independent draws; the modulo-31 bias of `below` is 2^-64 per value and ignored) +// ------------------------------------------------------------------------------------------------------------ + +fn ln_choose(n: u64, k: u64) -> f64 { + let lg = |x: u64| -> f64 { (1..=x).map(|i| (i as f64).ln()).sum() }; + lg(n) - lg(k) - lg(n - k) +} + +fn binom_tail(n: u64, p: f64, k_min: u64) -> f64 { + (k_min..=n).map(|k| (ln_choose(n, k) + (k as f64) * p.ln() + ((n - k) as f64) * (1.0 - p).ln()).exp()).sum() +} + +/// P(d distinct values among 8 uniform draws from 31): S(8, d) x 31 falling d / 31^8. +fn p_rot_distinct(d: u32) -> f64 { + // Stirling numbers of the second kind S(8, d) + let mut s = vec![vec![0f64; 9]; 9]; + s[0][0] = 1.0; + for n in 1..=8 { + for k in 1..=n { + s[n][k] = (k as f64) * s[n - 1][k] + s[n - 1][k - 1]; + } + } + let mut falling = 1.0; + for i in 0..d { + falling *= (31 - i) as f64; + } + s[8][d as usize] * falling / 31f64.powi(8) +} + +/// P(max multiplicity >= 4 among 8 draws from 31), by enumeration of the multiplicity patterns is long; the +/// census gives the exact count, so this is the first-order bound: C(8,4) x 31 x 31^-4 (approximate). +fn p_rot_mult4_approx() -> f64 { + 70.0 * 31.0 / 31f64.powi(4) +} + +fn p_any_of(n: u64, p: f64) -> f64 { + 1.0 - (1.0 - p).powi(n as i32) +} + +fn one_in(p: f64) -> String { + if p <= 0.0 { + "0".into() + } else { + format!("1 in {:.3e}", 1.0 / p) + } +} + +// ------------------------------------------------------------------------------------------------------------ +// Printing a day +// ------------------------------------------------------------------------------------------------------------ + +fn hex_bytes(b: &[u8]) -> String { + b.iter().map(|x| format!("{x:02x}")).collect() +} + +fn describe(day: u64, mp: &MixParams, c: &DayClass, median_cost: Option) -> String { + let mut s = String::new(); + let _ = writeln!(s, "day index {day} (genesis + {}), day bytes {} , seed64 {:016x}", day as i64 - GENESIS_DAY as i64, hex_bytes(&day_bytes(day)), seed64(&mp.key)); + let _ = writeln!(s, "key {}", mp.key.iter().map(|w| format!("{w:08x}")).collect::>().join(" ")); + let _ = writeln!(s, "ROT {:?} distinct {} max multiplicity {} small {} byte-aligned {} same-word pair 32 {} any pair 32 {}", mp.rot, c.rot_distinct, c.rot_max_mult, c.rot_small, c.rot_byte, c.rot_comp_same_word, c.rot_comp_any); + let _ = writeln!(s, "MUL {}", mp.mul.iter().map(|w| format!("{w:08x}")).collect::>().join(" ")); + let _ = writeln!(s, "NAF {:?} sum {} min {} =1 {} =-1 {} pop<=4 {} naf<=3 {} <256 {} dup {}", c.naf, c.naf_sum, c.naf_min, c.mul_one, c.mul_minus_one, c.mul_pop_le4, c.mul_naf_le3, c.mul_small, c.mul_dup); + let _ = writeln!(s, "RC {}", mp.rc.iter().map(|w| format!("{w:08x}")).collect::>().join(" ")); + let _ = writeln!(s, "RC zero {} pop<=4|>=28 {} rc+rk=0 {} dup {}", c.rc_zero, c.rc_pop_ext, c.rc_rk_zero, c.rc_dup); + let _ = writeln!(s, "M1 cost {} adder-equivalents per application ({} per item, 72 applications); M2 k {} (gain {:.3}x)", c.cost_m1, c.cost_m1 as u64 * APPLICATIONS_PER_ITEM, c.m2_k, m2_gain(c.m2_k)); + if let Some(m) = median_cost { + let _ = writeln!(s, "M1 gain against the census median {m}: {:.4}x", m as f64 / c.cost_m1 as f64); + } + let flags: Vec<&str> = Tally::flags(c).iter().zip(CLASSES.iter()).filter(|(f, _)| **f).map(|(_, n)| *n).collect(); + let _ = writeln!(s, "classes: {}", if flags.is_empty() { "none".to_string() } else { flags.join("; ") }); + s +} + +// ------------------------------------------------------------------------------------------------------------ +// Commands +// ------------------------------------------------------------------------------------------------------------ + +fn arg(args: &[String], name: &str) -> Option { + args.iter().position(|a| a == name).and_then(|i| args.get(i + 1).cloned()) +} + +fn census(args: &[String]) { + let from: u64 = arg(args, "--from").map(|v| v.parse().unwrap()).unwrap_or(GENESIS_DAY); + let count: u64 = arg(args, "--count").map(|v| v.parse().unwrap()).unwrap_or(1 << 24); + let threads: usize = arg(args, "--threads").map(|v| v.parse().unwrap()).unwrap_or(12); + let out = arg(args, "--out"); + let dedupe = args.iter().any(|a| a == "--dedupe"); + let shape = v4_shape(); + let rks = round_keys(); + let t0 = std::time::Instant::now(); + let chunk = count.div_ceil(threads as u64); + let tallies: Vec = std::thread::scope(|sc| { + let hs: Vec<_> = (0..threads) + .map(|t| { + let rks = &rks; + sc.spawn(move || { + let lo = from + chunk * t as u64; + let hi = (lo + chunk).min(from + count); + let mut tally = Tally::new(dedupe, (hi.saturating_sub(lo)) as usize); + for d in lo..hi { + let mp = params_of_day(d); + debug_assert_eq!(mp.shape, shape); + let c = classify(&mp, rks); + tally.add(d, &mp, &c, dedupe); + } + tally + }) + }) + .collect(); + hs.into_iter().map(|h| h.join().unwrap()).collect() + }); + let mut all = Tally::new(false, 0); + for t in tallies { + all.merge(t); + } + let secs = t0.elapsed().as_secs_f64(); + assert_eq!(all.n, count); + + // the median cost and the gate fractions + let total = all.n; + let mut acc = 0u64; + let mut median = 0u32; + for (c, &n) in all.cost_hist.iter().enumerate() { + acc += n; + if acc * 2 >= total { + median = c as u32; + break; + } + } + let mean = all.cost_hist.iter().enumerate().map(|(c, &n)| c as f64 * n as f64).sum::() / total as f64; + let var = all.cost_hist.iter().enumerate().map(|(c, &n)| (c as f64 - mean).powi(2) * n as f64).sum::() / total as f64; + let gate_cost = |reference: f64| -> u32 { (reference / GAIN_GATE).floor() as u32 }; // cost <= this gives gain >= 1.1x... strictly over: cost < reference/1.1 + let over = |reference: f64| -> u64 { + all.cost_hist.iter().enumerate().filter(|(c, _)| (*c as f64) * GAIN_GATE < reference).map(|(_, &n)| n).sum() + }; + let over_median = over(median as f64); + let over_mean = over(mean); + let m2_over: u64 = all.m2_k_hist.iter().enumerate().filter(|(k, _)| m2_gain(*k as u32) > GAIN_GATE).map(|(_, &n)| n).sum(); + + let mut r = String::new(); + let _ = writeln!(r, "# attack-f4 census: {count} chain days from day index {from} (genesis {GENESIS_DAY}), class v4 shape (mixer x{}, cache 2^{}), {threads} threads, {secs:.1} s", shape.mixer_mult, shape.cache_log2_words); + let _ = writeln!(r, "draw: MixParams::with_shape(seed_words_from_bytes(bind::day_bytes(d)), Shape::for_class(&V4_CLASS)); seed64 = key[0] | key[1] << 32"); + let _ = writeln!(r); + let _ = writeln!(r, "## Gate (plan 1.4 (3), B5 rank 3): fraction of days with any gain over {GAIN_GATE}x under 2^-20 = {:.3e}", 2f64.powi(-20)); + let _ = writeln!(r); + let _ = writeln!(r, "| Metric | Reference | Days over {GAIN_GATE}x | Fraction | Against 2^-20 |"); + let _ = writeln!(r, "|---|---|---|---|---|"); + let frac = |n: u64| n as f64 / total as f64; + let vs = |n: u64| if frac(n) < 2f64.powi(-20) { "under" } else { "OVER" }; + let _ = writeln!(r, "| M1 per-day LUT datapath, adders per application | median cost {median} (gain over {GAIN_GATE}x = cost under {}) | {over_median} | {:.3e} | {} |", gate_cost(median as f64) + 1, frac(over_median), vs(over_median)); + let _ = writeln!(r, "| M1 against the mean cost {mean:.2} (sd {:.2}) | cost under {:.2} | {over_mean} | {:.3e} | {} |", var.sqrt(), mean / GAIN_GATE, frac(over_mean), vs(over_mean)); + let _ = writeln!(r, "| M2 DSP-bound datapath, 16/(16-k) with k = words of NAF weight <= {M2_NAF_CEIL} | k >= 2 | {m2_over} | {:.3e} | {} |", frac(m2_over), vs(m2_over)); + let _ = writeln!(r, "| ROT value (wiring) and RC value (inverters) on a per-day datapath | exact 0 ops moved on every day | 0 | 0 | under |"); + let _ = writeln!(r); + let _ = writeln!(r, "M1 cost = {QR_ADDS_XORS} + sum(NAF(MUL_i) - 1); mean {mean:.3}, sd {:.3}, median {median}, min {} (day {}), max {} (day {})", var.sqrt(), all.lowest[0].0, all.lowest[0].1, all.highest[0].0, all.highest[0].1); + let _ = writeln!(r); + + // the classes + let p_mul1 = p_any_of(16, 2f64.powi(-31)); + let p_small = p_any_of(16, 128.0 / 2f64.powi(31)); + let p_rc0 = p_any_of(16, 2f64.powi(-32)); + let p_rcrk = p_any_of(16, 72.0 / 2f64.powi(32)); + let pop_le = |k: u32| -> f64 { (0..=k).map(|i| ln_choose(32, i as u64).exp()).sum::() }; + let p_rc_pop = p_any_of(16, 2.0 * pop_le(4) / 2f64.powi(32)); + // odd constants with popcount <= k: the low bit is set, so C(31, i) for the other i < k bits + let odd_pop_le = |k: u32| -> f64 { (0..k).map(|i| ln_choose(31, i as u64).exp()).sum::() / 2f64.powi(31) }; + let p_dup16 = |space: f64| -> f64 { 1.0 - (1..16).map(|i| 1.0 - i as f64 / space).product::() }; + let expectations: Vec> = vec![ + Some(31f64.powi(-7)), + Some((1..=3).map(p_rot_distinct).sum()), + Some((1..=4).map(p_rot_distinct).sum()), + Some(p_rot_mult4_approx()), + Some(p_any_of(4, 1.0 / 31.0)), + Some(p_any_of(28, 1.0 / 31.0)), + Some((4f64 / 31.0).powi(8)), + Some(binom_tail(8, 4.0 / 31.0, 6)), + Some(binom_tail(8, 4.0 / 31.0, 4)), + Some(binom_tail(8, 3.0 / 31.0, 4)), + Some(p_mul1), + Some(p_mul1), + Some(p_any_of(16, odd_pop_le(2))), + Some(p_any_of(16, odd_pop_le(4))), + Some(p_any_of(16, odd_pop_le(8))), + None, + None, + None, + Some(p_small), + Some(p_dup16(2f64.powi(31))), + None, + Some(p_rc0), + Some(p_rc_pop), + Some(p_rcrk), + Some(p_dup16(2f64.powi(32))), + ]; + let _ = writeln!(r, "## Classes over {count} days"); + let _ = writeln!(r); + let _ = writeln!(r, "| Class | Count | Fraction | Expected per day (analytic) | Expected count | Worst member (day, M1 cost, M1 gain vs median, M2 gain) |"); + let _ = writeln!(r, "|---|---|---|---|---|---|"); + for (i, name) in CLASSES.iter().enumerate() { + let n = all.class_count[i]; + let w = all.class_worst[i]; + let worst = if w.day == u64::MAX { + "none".to_string() + } else { + let c = classify(¶ms_of_day(w.day), &rks); + format!("day {} , cost {} , {:.4}x , {:.3}x", w.day, w.cost, median as f64 / w.cost as f64, m2_gain(c.m2_k)) + }; + let (ep, ec) = match expectations[i] { + Some(p) => (format!("{p:.3e} ({})", one_in(p)), format!("{:.3}", p * total as f64)), + None => ("see `expect`".to_string(), "see `expect`".to_string()), + }; + let _ = writeln!(r, "| {name} | {n} | {:.3e} | {ep} | {ec} | {worst} |", frac(n)); + } + let _ = writeln!(r); + let _ = writeln!(r, "## Histograms"); + let _ = writeln!(r); + let _ = writeln!(r, "| ROT distinct amounts | Count | Fraction | Expected (S(8,d) 31_d / 31^8) |"); + let _ = writeln!(r, "|---|---|---|---|"); + for d in 1..=8 { + let _ = writeln!(r, "| {d} | {} | {:.4e} | {:.4e} |", all.rot_distinct_hist[d], frac(all.rot_distinct_hist[d]), p_rot_distinct(d as u32)); + } + let _ = writeln!(r); + let _ = writeln!(r, "| ROT amounts in {{1,2,30,31}} | Count | Expected Binomial(8, 4/31) | ROT amounts in {{8,16,24}} | Count | Expected Binomial(8, 3/31) | ROT max multiplicity | Count |"); + let _ = writeln!(r, "|---|---|---|---|---|---|---|---|"); + for k in 0..=8 { + let e1 = binom_tail(8, 4.0 / 31.0, k) - binom_tail(8, 4.0 / 31.0, k + 1); + let e2 = binom_tail(8, 3.0 / 31.0, k) - binom_tail(8, 3.0 / 31.0, k + 1); + let _ = writeln!(r, "| {k} | {} | {:.1} | {k} | {} | {:.1} | {k} | {} |", all.rot_small_hist[k as usize], e1 * total as f64, all.rot_byte_hist[k as usize], e2 * total as f64, all.rot_mult_hist[k as usize]); + } + let _ = writeln!(r); + let _ = writeln!(r, "| M2 k (words of NAF weight <= {M2_NAF_CEIL}) | Count | Gain 16/(16-k) |"); + let _ = writeln!(r, "|---|---|---|"); + for k in 0..=16 { + if all.m2_k_hist[k] > 0 || k <= 3 { + let _ = writeln!(r, "| {k} | {} | {:.3}x |", all.m2_k_hist[k], m2_gain(k as u32)); + } + } + let _ = writeln!(r); + let _ = writeln!(r, "| Minimum NAF weight over the 16 MUL | Count |"); + let _ = writeln!(r, "|---|---|"); + for (w, &n) in all.naf_min_hist.iter().enumerate() { + if n > 0 { + let _ = writeln!(r, "| {w} | {n} |"); + } + } + let _ = writeln!(r); + let _ = writeln!(r, "| M1 cost per application | Count | Cumulative fraction | M1 gain vs median |"); + let _ = writeln!(r, "|---|---|---|---|"); + let mut cum = 0u64; + for (c, &n) in all.cost_hist.iter().enumerate() { + if n > 0 { + cum += n; + let _ = writeln!(r, "| {c} | {n} | {:.4e} | {:.4}x |", frac(cum), median as f64 / c as f64); + } + } + let _ = writeln!(r); + let _ = writeln!(r, "## The 16 lowest-cost days (M1)"); + let _ = writeln!(r); + for &(cost, day) in &all.lowest { + let mp = params_of_day(day); + let c = classify(&mp, &rks); + let _ = writeln!(r, "- day {day}: cost {cost}, gain {:.4}x vs median, NAF sum {}, M2 k {}, day-hex {}", median as f64 / cost as f64, c.naf_sum, c.m2_k, hex_bytes(&day_bytes(day))); + } + if dedupe { + all.seeds.sort_unstable(); + let before = all.seeds.len(); + all.seeds.dedup(); + let _ = writeln!(r); + let _ = writeln!(r, "## 64-bit seeding: {} days, {} distinct 64-bit stream seeds ({} collisions; expected C(n,2)/2^64 = {:.3e})", before, all.seeds.len(), before - all.seeds.len(), (before as f64) * (before as f64 - 1.0) / 2.0 / 2f64.powi(64)); + } + let _ = writeln!(r); + let _ = writeln!(r, "## Worst member of every class, in full"); + let _ = writeln!(r); + let mut shown: Vec = Vec::new(); + for (i, name) in CLASSES.iter().enumerate() { + let w = all.class_worst[i]; + if w.day == u64::MAX || shown.contains(&w.day) { + continue; + } + shown.push(w.day); + let mp = params_of_day(w.day); + let c = classify(&mp, &rks); + let _ = writeln!(r, "### {name}: day {}", w.day); + let _ = writeln!(r, "```"); + let _ = write!(r, "{}", describe(w.day, &mp, &c, Some(median))); + let _ = writeln!(r, "```"); + } + print!("{r}"); + if let Some(p) = out { + std::fs::File::create(&p).unwrap().write_all(r.as_bytes()).unwrap(); + eprintln!("written {p}"); + } +} + +fn day_cmd(args: &[String]) { + let d: u64 = arg(args, "--index").map(|v| v.parse().unwrap()).unwrap_or(GENESIS_DAY); + let median: Option = arg(args, "--median").map(|v| v.parse().unwrap()); + let mp = params_of_day(d); + let c = classify(&mp, &round_keys()); + print!("{}", describe(d, &mp, &c, median)); +} + +/// The known-fail firings: the genesis day's draw with one field forced through this crate's own hook (the +/// `MixParams` fields are public; `igneum-pow` is untouched). Exit 0 when the classifier flags the plant and the +/// gain metric that the plant moves reads over the gate. +fn plant(args: &[String]) { + let what = args.get(2).cloned().unwrap_or_default(); + let median: u32 = arg(args, "--median").map(|v| v.parse().unwrap()).unwrap_or(221); + let rks = round_keys(); + let mut mp = params_of_day(GENESIS_DAY); + let before = classify(&mp, &rks); + println!("before the plant (day {GENESIS_DAY}):"); + print!("{}", describe(GENESIS_DAY, &mp, &before, Some(median))); + let (flag_name, gain_metric): (&str, &str) = match what.as_str() { + "alleq" => { + mp.rot = [7; 8]; + ("ROT all equal", "structure (0 ops on a per-day datapath by construction; diffusion in `avalanche`)") + } + "mul1" => { + mp.mul[5] = 1; + ("MUL any = 1", "M1") + } + "mul1all" => { + mp.mul = [1; 16]; + ("MUL any = 1", "M1") + } + "mulnaf" => { + // the lightest realistic plant: four words at NAF weight 3 (M2 k = 4), the rest untouched + for i in 0..4 { + mp.mul[i] = (1u32 << 20) + (1u32 << 9) + 1; + } + ("MUL any NAF weight <= 3", "M1 and M2") + } + "rc0" => { + mp.rc[3] = 0; + ("RC any = 0", "structure (0 ops on a per-day datapath by construction)") + } + "rcrk0" => { + mp.rc[3] = 0u32.wrapping_sub(round_key_mult(2, 5, 8)); + ("RC + rk = 0 for any of the 72 keys", "structure (one xor of 10,368 ops per item on a generic datapath: 1.0001x)") + } + _ => { + eprintln!("plant alleq|mul1|mul1all|mulnaf|rc0|rcrk0"); + std::process::exit(2); + } + }; + let after = classify(&mp, &rks); + println!("\nafter the plant `{what}`:"); + print!("{}", describe(GENESIS_DAY, &mp, &after, Some(median))); + let idx = CLASSES.iter().position(|n| *n == flag_name).unwrap(); + let flagged = Tally::flags(&after)[idx] && !Tally::flags(&before)[idx]; + let gain_m1 = median as f64 / after.cost_m1 as f64; + let gain_m2 = m2_gain(after.m2_k); + println!("\nplant `{what}`: classifier flag `{flag_name}` {} (was {} before); M1 gain {gain_m1:.4}x, M2 gain {gain_m2:.4}x; gain metric for this plant: {gain_metric}", if flagged { "FIRED" } else { "did NOT fire" }, Tally::flags(&before)[idx]); + let gain_fired = gain_m1 > GAIN_GATE || gain_m2 > GAIN_GATE; + println!("gain over {GAIN_GATE}x: {}", if gain_fired { "FIRED" } else { "not over the gate" }); + if !flagged { + std::process::exit(1); + } +} + +/// Exact per-word tables over every odd 32-bit constant (2^31 of them): the NAF weight distribution and the +/// popcount distribution; then the 16-fold convolution of the NAF-weight distribution gives the expected M1 cost +/// distribution and the expected fraction of days under any cost. +fn expect(args: &[String]) { + let threads: usize = arg(args, "--threads").map(|v| v.parse().unwrap()).unwrap_or(12); + let median: u32 = arg(args, "--median").map(|v| v.parse().unwrap()).unwrap_or(221); + let t0 = std::time::Instant::now(); + let per: Vec<([u64; 40], [u64; 33])> = std::thread::scope(|sc| { + let hs: Vec<_> = (0..threads) + .map(|t| { + sc.spawn(move || { + let mut naf = [0u64; 40]; + let mut pop = [0u64; 33]; + let lo = ((1u64 << 32) * t as u64 / threads as u64) | 1; + let hi = (1u64 << 32) * (t as u64 + 1) / threads as u64; + let mut v = lo; + while v < hi { + naf[naf_weight(v as u32) as usize] += 1; + pop[(v as u32).count_ones() as usize] += 1; + v += 2; + } + (naf, pop) + }) + }) + .collect(); + hs.into_iter().map(|h| h.join().unwrap()).collect() + }); + let mut naf = [0u64; 40]; + let mut pop = [0u64; 33]; + for (a, b) in per { + for i in 0..40 { + naf[i] += a[i]; + } + for i in 0..33 { + pop[i] += b[i]; + } + } + let total: u64 = naf.iter().sum(); + assert_eq!(total, 1 << 31); + println!("# attack-f4 expect: all {total} odd 32-bit constants, {threads} threads, {:.1} s", t0.elapsed().as_secs_f64()); + println!(); + println!("| NAF weight | Odd constants | Fraction | Cumulative | Any of 16 per day | x 2^24 days |"); + println!("|---|---|---|---|---|---|"); + let mut cum = 0u64; + for w in 0..40 { + if naf[w] > 0 { + cum += naf[w]; + let p = cum as f64 / total as f64; + println!("| {w} | {} | {:.4e} | {:.4e} | {:.4e} | {:.3} |", naf[w], naf[w] as f64 / total as f64, p, p_any_of(16, p), p_any_of(16, p) * 2f64.powi(24)); + } + } + println!(); + println!("| Popcount | Odd constants | Cumulative fraction | Any of 16 per day |"); + println!("|---|---|---|---|"); + cum = 0; + for w in 0..33 { + if pop[w] > 0 { + cum += pop[w]; + let p = cum as f64 / total as f64; + println!("| {w} | {} | {:.4e} | {:.4e} |", pop[w], p, p_any_of(16, p)); + } + } + // the 16-fold convolution of the NAF-weight distribution: the exact expected distribution of the NAF sum + let pw: Vec = naf.iter().map(|&n| n as f64 / total as f64).collect(); + let mut dist = vec![0f64; 1]; + dist[0] = 1.0; + for _ in 0..16 { + let mut next = vec![0f64; dist.len() + 39]; + for (i, &a) in dist.iter().enumerate() { + if a == 0.0 { + continue; + } + for (w, &b) in pw.iter().enumerate() { + next[i + w] += a * b; + } + } + dist = next; + } + let mean: f64 = dist.iter().enumerate().map(|(s, &p)| s as f64 * p).sum(); + let var: f64 = dist.iter().enumerate().map(|(s, &p)| (s as f64 - mean).powi(2) * p).sum(); + println!(); + println!("Expected NAF sum over 16 words: mean {mean:.4}, sd {:.4}; expected M1 cost mean {:.4}", var.sqrt(), mean + QR_ADDS_XORS as f64 - 16.0); + println!(); + println!("| M1 cost | NAF sum | Expected fraction of days at this cost | Expected cumulative fraction | Gain vs median {median} |"); + println!("|---|---|---|---|---|"); + let mut c = 0f64; + for (s, &p) in dist.iter().enumerate() { + let cost = s as i64 + QR_ADDS_XORS as i64 - 16; + if cost < 0 { + continue; + } + c += p; + if p > 1e-12 && (cost as f64) <= median as f64 { + println!("| {cost} | {s} | {p:.4e} | {c:.4e} | {:.4}x |", median as f64 / cost as f64); + } + } + let gate: f64 = dist.iter().enumerate().filter(|(s, _)| ((*s as f64) + QR_ADDS_XORS as f64 - 16.0) * GAIN_GATE < median as f64).map(|(_, &p)| p).sum(); + println!(); + println!("Expected fraction of days with M1 gain over {GAIN_GATE}x against median {median}: {gate:.4e} ({} ; x 2^24 = {:.1}); 2^-20 = {:.4e}", one_in(gate), gate * 2f64.powi(24), 2f64.powi(-20)); +} + +/// Diffusion of one day's mixer: for `states` random 16-word states and each of the 512 input bits, the fraction of +/// the 512 output bits that flip after k = 1 and k = 2 applications (keys `round_key_mult(0, 0, 8)` and `(0, 1, 8)`), +/// mean over all, and the minimum per-output-bit flip probability. An ideal mixer reads 0.5 mean and about 0.5 min. +fn avalanche(args: &[String]) { + let d: u64 = arg(args, "--index").map(|v| v.parse().unwrap()).unwrap_or(GENESIS_DAY); + let states: usize = arg(args, "--states").map(|v| v.parse().unwrap()).unwrap_or(4096); + let plant_alleq: Option = arg(args, "--plant-alleq").map(|v| v.parse().unwrap()); + let mut mp = params_of_day(d); + if let Some(r) = plant_alleq { + mp.rot = [r; 8]; + } + let c = classify(&mp, &round_keys()); + println!("avalanche of day {d}{}: ROT {:?}, NAF sum {}, {states} states x 512 input bits", plant_alleq.map(|r| format!(" with ROT planted all {r}")).unwrap_or_default(), mp.rot, c.naf_sum); + let mut rng = SplitMix64::new(0xF4F4_F4F4 ^ d); + for k in 1..=3usize { + let apply = |s: &mut [u32; 16]| { + for j in 0..k { + mixer(s, round_keys()[j], &mp); + } + }; + let mut flips = [0u64; 512]; + let mut total_flips = 0u64; + let mut trials = 0u64; + for _ in 0..states { + let mut base = [0u32; 16]; + for w in base.iter_mut() { + *w = rng.next() as u32; + } + let mut y0 = base; + apply(&mut y0); + for bit in 0..512 { + let mut x = base; + x[bit / 32] ^= 1 << (bit % 32); + apply(&mut x); + for o in 0..512 { + let f = ((x[o / 32] ^ y0[o / 32]) >> (o % 32)) & 1; + flips[o] += f as u64; + total_flips += f as u64; + } + trials += 1; + } + } + let mean = total_flips as f64 / (trials as f64 * 512.0); + let min = flips.iter().map(|&f| f as f64 / trials as f64).fold(1.0, f64::min); + let max = flips.iter().map(|&f| f as f64 / trials as f64).fold(0.0, f64::max); + println!("| {k} application{} | mean flip {mean:.4} | min per output bit {min:.4} | max {max:.4} |", if k > 1 { "s" } else { "" }); + } +} + +fn main() { + let args: Vec = std::env::args().collect(); + match args.get(1).map(|s| s.as_str()) { + Some("census") => census(&args), + Some("day") => day_cmd(&args), + Some("plant") => plant(&args), + Some("expect") => expect(&args), + Some("avalanche") => avalanche(&args), + _ => { + eprintln!("attack-f4 census|day|plant|expect|avalanche (see the module doc)"); + std::process::exit(2); + } + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn naf_weights() { + assert_eq!(naf_weight(0), 0); + assert_eq!(naf_weight(1), 1); + assert_eq!(naf_weight(3), 2); // 4 - 1 + assert_eq!(naf_weight(7), 2); // 8 - 1 + assert_eq!(naf_weight(0xFFFF_FFFF), 2); // 2^32 - 1 + assert_eq!(naf_weight(0xAAAA_AAAB), 17); // alternating odd: the maximum for 32 bits + assert_eq!(naf_weight((1 << 20) + (1 << 9) + 1), 3); + } + + #[test] + fn genesis_day_draw_matches_memhard_md() { + // MEMHARD.md section 1.1 (string day "2026-10-03") is a different key from the chain's day index 20,729; + // what is checked here is that the chain-day path draws through the real code and stays in range. + let mp = params_of_day(GENESIS_DAY); + assert!(mp.rot.iter().all(|r| (1..=31).contains(r))); + assert!(mp.mul.iter().all(|m| m & 1 == 1)); + assert_eq!(mp.shape, v4_shape()); + let s = igneum_pow::seed::day_key("2026-10-03"); + let mp2 = MixParams::with_shape(s, Shape::V2); + assert_eq!(mp2.rot, [20, 20, 19, 4, 26, 3, 3, 27], "MEMHARD.md 1.1 genesis-day ROT"); + } +} diff --git a/tools/attack/f8-uniform/Cargo.lock b/tools/attack/f8-uniform/Cargo.lock new file mode 100644 index 000000000..e79a483fa --- /dev/null +++ b/tools/attack/f8-uniform/Cargo.lock @@ -0,0 +1,14 @@ +# This file is automatically @generated by Cargo. +# It is not intended for manual editing. +version = 4 + +[[package]] +name = "attack-f8" +version = "0.1.0" +dependencies = [ + "igneum-pow", +] + +[[package]] +name = "igneum-pow" +version = "0.2.0" diff --git a/tools/attack/f8-uniform/Cargo.toml b/tools/attack/f8-uniform/Cargo.toml new file mode 100644 index 000000000..bc0ef039b --- /dev/null +++ b/tools/attack/f8-uniform/Cargo.toml @@ -0,0 +1,21 @@ +[package] +name = "attack-f8" +version = "0.1.0" +edition = "2021" +description = "Attack-pass row F8: the uniformity censuses of the class v4 derivation (line index over 2^28 derivations, distinct lines per hash and warp, the cross-hash item histogram), with the plant hooks that prove the harness fires" +license = "MIT" +publish = false + +[[bin]] +name = "attack-f8" +path = "src/main.rs" + +[dependencies] +igneum-pow = { path = "../../../igneum-pow" } + +[workspace] + +[profile.release] +opt-level = 3 +lto = true +codegen-units = 1 diff --git a/tools/attack/f8-uniform/src/main.rs b/tools/attack/f8-uniform/src/main.rs new file mode 100644 index 000000000..cac66dc09 --- /dev/null +++ b/tools/attack/f8-uniform/src/main.rs @@ -0,0 +1,1745 @@ +//! attack-f8: the uniformity censuses of attack-pass row F8 (`docs/plans/cryptanalysis.md` 4.2 F8) on the real +//! class v4 derivation through the `igneum-pow` library. +//! +//! Three censuses, one binary: +//! +//! * `lines`: the cache-line index `s[0] AND line_mask` (MEMHARD.md 1.6, 2^22 lines) of every one of the 8 reads of +//! every item `t < 2^items_log2` on `days` consecutive day keys: the full 2^22-line histogram per round and +//! pooled, the 2^16-bucket (64-line segment) histogram, the largest bucket in sigma, chi-square, and a uniform +//! SplitMix64 control of the same size. Every item is also derived by the library's `derive_items` and compared +//! word for word (the traced mirror is trusted only while that holds). +//! * `warps`: the warp interpreter mirrored with every load's item index and the item's 8 lines recorded (the +//! day's items derived once into a table, as a GPU holds the dataset): distinct lines and items per hash and per +//! warp over `nonces` nonces of one program, the cross-hash item histogram with the hot-set test, and a uniform +//! control. Sampled warps are hashed by `Epoch::hash_warp` as well and must agree bit for bit. +//! +//! Plant hooks (the known-fail firings): `--plant quarter-lines` masks the line index to a quarter of its range, +//! `--plant half-lines` to a half, `--plant const-item` feeds one constant item at the program's first load site. +//! Under a plant the library comparisons are skipped (the plant is not the derivation) and the log says so. +//! +//! Nothing in `igneum-pow` is modified; every constant and the read sequence are the library's. + +use igneum_pow::bind::{day_bytes, hex, unhex}; +use igneum_pow::generator::{EraParams, Instr, Op, Program, ProgramClass, ITERATIONS, LANES}; +use igneum_pow::memhard::{mixer, round_key_mult, Cache, Layout, MixParams, ITEM_ROUNDS}; +use igneum_pow::seed::{seed_words_from_bytes, SplitMix64}; +use igneum_pow::verify::{load_index, splitmix32, DatasetSource, Epoch}; +use std::io::Write; +use std::sync::atomic::{AtomicU32, AtomicUsize, Ordering}; +use std::time::{Instant, SystemTime, UNIX_EPOCH}; + +/// The devnet genesis hash: epoch seed and era seed of the chain's epoch 0 (`igneum-pow/tests/packs.rs`). +const GENESIS_HEX: &str = "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"; +/// The devnet pack's day index (`bind::day_bytes(20730)`, 2026-10-04). +const DEFAULT_DAY: u64 = 20730; +const DATASET_LOG2: u32 = 28; + +#[derive(Clone, Copy, PartialEq, Eq, Debug)] +enum Plant { + None, + QuarterLines, + HalfLines, + ConstItem, +} + +impl Plant { + fn parse(s: &str) -> Plant { + match s { + "none" => Plant::None, + "quarter-lines" => Plant::QuarterLines, + "half-lines" => Plant::HalfLines, + "const-item" => Plant::ConstItem, + _ => panic!("unknown plant {s}"), + } + } + fn line_mask(self) -> u32 { + match self { + Plant::QuarterLines => (1u32 << 20) - 1, + Plant::HalfLines => (1u32 << 21) - 1, + _ => u32::MAX, + } + } + fn name(self) -> &'static str { + match self { + Plant::None => "none", + Plant::QuarterLines => "quarter-lines", + Plant::HalfLines => "half-lines", + Plant::ConstItem => "const-item", + } + } +} + +fn utc_now() -> String { + let s = SystemTime::now().duration_since(UNIX_EPOCH).unwrap().as_secs(); + let (d, t) = (s / 86400, s % 86400); + // civil date from days (Howard Hinnant's algorithm) + let z = d as i64 + 719468; + let era = z.div_euclid(146097); + let doe = z.rem_euclid(146097); + let yoe = (doe - doe / 1460 + doe / 36524 - doe / 146096) / 365; + let y = yoe + era * 400; + let doy = doe - (365 * yoe + yoe / 4 - yoe / 100); + let mp = (5 * doy + 2) / 153; + let dd = doy - (153 * mp + 2) / 5 + 1; + let mm = if mp < 10 { mp + 3 } else { mp - 9 }; + let yy = if mm <= 2 { y + 1 } else { y }; + format!("{yy:04}-{mm:02}-{dd:02}T{:02}:{:02}:{:02}Z", t / 3600, (t / 60) % 60, t % 60) +} + +macro_rules! log { + ($($arg:tt)*) => {{ + println!("[{}] {}", utc_now(), format!($($arg)*)); + std::io::stdout().flush().ok(); + }}; +} + +// -------------------------------------------------------------------------------------------------------------- +// The traced item derivation: `memhard::derive_items_mask` instruction for instruction, with the line index of +// every round recorded. Batched like the library so the 8 dependent misses of independent items overlap. +// -------------------------------------------------------------------------------------------------------------- + +fn derive_traced(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; 16]], lines: &mut [[u32; 8]], plant_mask: u32) { + let n = ts.len(); + let m = mp.shape.mixer_mult as usize; + assert_eq!(mp.shape.derive_len, 0, "class v4 has the fixed mixer"); + let line_mask = cache.line_mask(); + for k in 0..n { + let s = &mut out[k]; + let t = ts[k]; + s[..8].copy_from_slice(&mp.key); + for i in 0..8 { + s[8 + i] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]); + } + } + for r in 0..ITEM_ROUNDS { + for j in 0..m { + let rk = round_key_mult(r, j, m); + for s in out[..n].iter_mut() { + mixer(s, rk, mp); + } + } + for k in 0..n { + let s = &mut out[k]; + // MEMHARD.md 1.6: `a = s[0] AND 0x003fffff`, the cache line index (the mask is the cache's at every size) + let a = s[0] & line_mask & plant_mask; + lines[k][r] = a; + let line = cache.line(a); + for i in 0..16 { + s[i] ^= line[i]; + } + } + } + for j in 0..m { + let rk = round_key_mult(ITEM_ROUNDS, j, m); + for s in out[..n].iter_mut() { + mixer(s, rk, mp); + } + } +} + +// -------------------------------------------------------------------------------------------------------------- +// Histogram statistics +// -------------------------------------------------------------------------------------------------------------- + +struct Stats { + bins: usize, + total: u64, + mean: f64, + sigma: f64, + max: u64, + argmax: usize, + min: u64, + z_max: f64, + z_min: f64, + chi2_per_dof: f64, + chi2_z: f64, + /// (fraction f, top-f share S_f) for f = 1/1000, 1/200, 1/100 + top: [(f64, f64); 3], +} + +fn stats_of(counts: impl Iterator + Clone, bins: usize) -> Stats { + let total: u64 = counts.clone().sum(); + let mean = total as f64 / bins as f64; + let sigma = mean.sqrt(); + let mut max = 0u64; + let mut argmax = 0usize; + let mut min = u64::MAX; + let mut chi2 = 0f64; + for (i, c) in counts.clone().enumerate() { + if c > max { + max = c; + argmax = i; + } + if c < min { + min = c; + } + let d = c as f64 - mean; + chi2 += d * d; + } + chi2 /= mean; + let dof = (bins - 1) as f64; + // count of counts for the top-f shares (counts at or above the cap, rare, are sorted on their own) + let cap = 1usize << 16; + let mut coc = vec![0u64; cap]; + let mut big: Vec = Vec::new(); + for c in counts { + if (c as usize) < cap { + coc[c as usize] += 1; + } else { + big.push(c); + } + } + big.sort_unstable_by(|a, b| b.cmp(a)); + let mut top = [(0f64, 0f64); 3]; + for (k, f) in [1.0 / 1000.0, 1.0 / 200.0, 1.0 / 100.0].into_iter().enumerate() { + let want = (f * bins as f64).round() as u64; + let mut left = want; + let mut reads = 0u64; + for &c in &big { + if left == 0 { + break; + } + reads += c; + left -= 1; + } + let mut v = cap - 1; + while left > 0 { + let n = coc[v].min(left); + reads += n * v as u64; + left -= n; + if v == 0 { + break; + } + v -= 1; + } + top[k] = (f, reads as f64 / total.max(1) as f64); + } + Stats { + bins, + total, + mean, + sigma, + max, + argmax, + min, + z_max: (max as f64 - mean) / sigma, + z_min: (min as f64 - mean) / sigma, + chi2_per_dof: chi2 / dof, + chi2_z: (chi2 - dof) / (2.0 * dof).sqrt(), + top, + } +} + +impl Stats { + fn line(&self, label: &str) -> String { + format!( + "{label}: bins {} reads {} mean {:.3} sigma {:.3} max {} (bin {}) z_max {:+.2} min {} z_min {:+.2} chi2/dof {:.5} chi2_z {:+.2} top0.1% {:.5}% top0.5% {:.5}% top1% {:.5}%", + self.bins, + self.total, + self.mean, + self.sigma, + self.max, + self.argmax, + self.z_max, + self.min, + self.z_min, + self.chi2_per_dof, + self.chi2_z, + self.top[0].1 * 100.0, + self.top[1].1 * 100.0, + self.top[2].1 * 100.0 + ) + } +} + +fn snapshot(counts: &[AtomicU32]) -> Vec { + counts.iter().map(|x| x.load(Ordering::Relaxed) as u64).collect() +} + +fn zero(counts: &[AtomicU32]) { + for c in counts { + c.store(0, Ordering::Relaxed); + } +} + +fn atomic_vec(n: usize) -> Vec { + (0..n).map(|_| AtomicU32::new(0)).collect() +} + +/// A uniform control: `n` SplitMix64 indices into `bins` bins, `threads` threads. +fn control(bins: usize, n: u64, seed: u64, threads: usize, hist: &[AtomicU32]) { + let mask = (bins - 1) as u64; + assert!(bins.is_power_of_two()); + std::thread::scope(|sc| { + for th in 0..threads { + let hist = &hist; + sc.spawn(move || { + let mut s = SplitMix64::new(seed ^ (th as u64).wrapping_mul(0x9E3779B97F4A7C15)); + let per = n / threads as u64 + if (th as u64) < n % threads as u64 { 1 } else { 0 }; + for _ in 0..per { + let i = (s.next() & mask) as usize; + hist[i].fetch_add(1, Ordering::Relaxed); + } + }); + } + }); +} + +/// The hot-set test (F8's definition, reused by F9): for the top-f items by count, the share S_f of all reads +/// they receive, against the same share E_f of a uniform control of the same size; the excess X_f = S_f - E_f. +/// The set is a hot set when X_f >= f for any f in {0.1%, 0.5%, 1%}: after the chance excess is removed, the top +/// f of items capture at least one extra proportional share (at least 2f of the reads above chance, which is +/// what a chip's on-die copy of f of the items would have to win to matter). The statistical sensitivity is +/// printed beside it: X_f in units of f. +fn hot_set_test(label: &str, real: &Stats, ctrl: &Stats) -> bool { + let mut flagged = false; + for k in 0..3 { + let (f, s) = real.top[k]; + let e = ctrl.top[k].1; + let x = s - e; + let fire = x >= f; + flagged |= fire; + log!( + "hot-set {label}: f {:.1}% S_f {:.5}% E_f(control) {:.5}% X_f {:+.5}% X_f/f {:+.4} -> {}", + f * 100.0, + s * 100.0, + e * 100.0, + x * 100.0, + x / f, + if fire { "HOT SET" } else { "no hot set" } + ); + } + log!("hot-set {label}: verdict {}", if flagged { "FLAGGED" } else { "clear" }); + flagged +} + +/// The 6-sigma test: the largest bucket within 6 sigma of uniform. +fn six_sigma_test(label: &str, s: &Stats) -> bool { + let fire = s.z_max > 6.0 || s.z_min < -6.0; + log!( + "6-sigma {label}: largest bucket {} at {:+.2} sigma, smallest {} at {:+.2} sigma -> {}", + s.max, + s.z_max, + s.min, + s.z_min, + if fire { "FLAGGED (beyond 6 sigma)" } else { "within 6 sigma" } + ); + fire +} + +// -------------------------------------------------------------------------------------------------------------- +// Census 1: the line index over all items of `days` day keys +// -------------------------------------------------------------------------------------------------------------- + +fn census_lines(day0: u64, days: u64, items_log2: u32, threads: usize, plant: Plant, validate: &str, out_dir: &str) { + log!("census lines: day0 {day0} days {days} items 2^{items_log2} per day threads {threads} plant {} validate {validate}", plant.name()); + let lines_n = 1usize << 22; + let per_round: Vec> = (0..ITEM_ROUNDS).map(|_| atomic_vec(lines_n)).collect(); + let pooled: Vec> = (0..ITEM_ROUNDS).map(|_| atomic_vec(lines_n)).collect(); + let items = 1u64 << items_log2; + let batch = 64u64; + let mut any_fire = false; + let mut mismatches_total = 0u64; + let mut derivations = 0u64; + let t_all = Instant::now(); + for d in day0..day0 + days { + let t0 = Instant::now(); + let ds = Epoch::chain_dataset_day(&day_bytes(d), ProgramClass::V4, 0, DATASET_LOG2); + let mh = ds.memhard().expect("memory-hard"); + assert_eq!(mh.params.shape.mixer_mult, 8); + assert_eq!(mh.params.shape.cache_log2_words, 26); + assert_eq!(mh.cache.line_mask(), (1u32 << 22) - 1); + log!( + "day {d}: day bytes {} key {} rot {:?} cache 2^26 words fnv {:016x} filled in {:.2} s", + hex(&day_bytes(d)), + mh.params.key.iter().map(|w| format!("{w:08x}")).collect::>().join(""), + mh.params.rot, + mh.cache.fnv1a64(), + t0.elapsed().as_secs_f64() + ); + for h in &per_round { + zero(h); + } + let next = AtomicUsize::new(0); + let mism = AtomicUsize::new(0); + let t1 = Instant::now(); + std::thread::scope(|sc| { + for _ in 0..threads { + let (next, mism, per_round, mp, cache) = (&next, &mism, &per_round, &mh.params, &mh.cache); + sc.spawn(move || { + let mut ts = [0u32; 64]; + let mut out = [[0u32; 16]; 64]; + let mut lines = [[0u32; 8]; 64]; + let mut lib = [[0u32; 16]; 64]; + loop { + let b = next.fetch_add(1, Ordering::Relaxed) as u64; + let start = b * batch; + if start >= items { + break; + } + for k in 0..64 { + ts[k] = (start + k as u64) as u32; + } + derive_traced(&ts, mp, cache, &mut out, &mut lines, plant.line_mask()); + for k in 0..64 { + for r in 0..ITEM_ROUNDS { + per_round[r][lines[k][r] as usize].fetch_add(1, Ordering::Relaxed); + } + } + let check = plant == Plant::None && (validate == "all" || (validate == "sample" && b % 64 == 0)); + if check { + igneum_pow::memhard::derive_items(&ts, mp, cache, &mut lib); + for k in 0..64 { + if lib[k] != out[k] { + mism.fetch_add(1, Ordering::Relaxed); + } + } + } + } + }); + } + }); + let el = t1.elapsed().as_secs_f64(); + derivations += items; + let mm = mism.load(Ordering::Relaxed) as u64; + mismatches_total += mm; + log!( + "day {d}: {items} items ({} line reads) derived in {el:.1} s ({:.2} M items/s); library comparison: {} ({} mismatches)", + items * 8, + items as f64 / el / 1e6, + if plant == Plant::None { validate } else { "skipped under the plant" }, + mm + ); + if mm > 0 { + log!("day {d}: THE TRACED MIRROR DISAGREES WITH THE LIBRARY: {mm} items; nothing below is trusted"); + } + // per-day stats: per round, total over rounds, buckets of 64 lines (one segment) + let mut total = vec![0u64; lines_n]; + for r in 0..ITEM_ROUNDS { + let snap = snapshot(&per_round[r]); + let s = stats_of(snap.iter().copied(), lines_n); + log!("{}", s.line(&format!("day {d} round {r} lines"))); + for (i, c) in snap.iter().enumerate() { + total[i] += c; + pooled[r][i].fetch_add(*c as u32, Ordering::Relaxed); + } + } + let s_full = stats_of(total.iter().copied(), lines_n); + log!("{}", s_full.line(&format!("day {d} all rounds lines"))); + let b: Vec = total.chunks(64).map(|c| c.iter().sum()).collect(); + let s_b = stats_of(b.iter().copied(), b.len()); + log!("{}", s_b.line(&format!("day {d} all rounds buckets64"))); + let fire = six_sigma_test(&format!("day {d} buckets64"), &s_b); + any_fire |= fire; + // control of the same size + let ctrl = atomic_vec(lines_n); + control(lines_n, items * 8, 0xF8_0000 + d, threads, &ctrl); + let cs = snapshot(&ctrl); + let c_full = stats_of(cs.iter().copied(), lines_n); + log!("{}", c_full.line(&format!("day {d} CONTROL lines"))); + let cb: Vec = cs.chunks(64).map(|c| c.iter().sum()).collect(); + let c_b = stats_of(cb.iter().copied(), cb.len()); + log!("{}", c_b.line(&format!("day {d} CONTROL buckets64"))); + let hot = hot_set_test(&format!("day {d} lines"), &s_full, &c_full); + any_fire |= hot; + } + // pooled over the days + let mut total = vec![0u64; lines_n]; + for r in 0..ITEM_ROUNDS { + let snap = snapshot(&pooled[r]); + let s = stats_of(snap.iter().copied(), lines_n); + log!("{}", s.line(&format!("POOLED {days} days round {r} lines"))); + for (i, c) in snap.iter().enumerate() { + total[i] += c; + } + } + let s_full = stats_of(total.iter().copied(), lines_n); + log!("{}", s_full.line(&format!("POOLED {days} days all rounds lines"))); + let b: Vec = total.chunks(64).map(|c| c.iter().sum()).collect(); + let s_b = stats_of(b.iter().copied(), b.len()); + log!("{}", s_b.line(&format!("POOLED {days} days all rounds buckets64"))); + let fire = six_sigma_test(&format!("POOLED {days} days buckets64"), &s_b); + let fire_full = six_sigma_test(&format!("POOLED {days} days full 2^22 lines"), &s_full); + let ctrl = atomic_vec(lines_n); + control(lines_n, derivations * 8, 0xF8_1111, threads, &ctrl); + let cs = snapshot(&ctrl); + let c_full = stats_of(cs.iter().copied(), lines_n); + log!("{}", c_full.line(&format!("POOLED CONTROL lines"))); + let cb: Vec = cs.chunks(64).map(|c| c.iter().sum()).collect(); + let c_b = stats_of(cb.iter().copied(), cb.len()); + log!("{}", c_b.line(&format!("POOLED CONTROL buckets64"))); + let hot = hot_set_test(&format!("POOLED {days} days lines"), &s_full, &c_full); + any_fire |= fire | fire_full | hot; + // files + let tag = format!("lines-d{day0}-n{days}-i{items_log2}-{}", plant.name()); + { + let mut f = std::fs::File::create(format!("{out_dir}/{tag}-buckets64.txt")).unwrap(); + writeln!(f, "# bucket(64 lines = one segment) count ; pooled over {days} days from {day0}, items 2^{items_log2} per day, plant {}", plant.name()).unwrap(); + for (i, c) in b.iter().enumerate() { + writeln!(f, "{i} {c}").unwrap(); + } + let mut g = std::fs::File::create(format!("{out_dir}/{tag}-full.u32le")).unwrap(); + let mut bytes = Vec::with_capacity(lines_n * 4); + for c in &total { + bytes.extend_from_slice(&(*c as u32).to_le_bytes()); + } + g.write_all(&bytes).unwrap(); + } + log!( + "census lines DONE: {derivations} item derivations ({} line reads) over {days} days in {:.1} s; mirror mismatches {mismatches_total}; verdict {}", + derivations * 8, + t_all.elapsed().as_secs_f64(), + if any_fire { "FLAGGED" } else { "PASS (no test fired)" } + ); +} + +// -------------------------------------------------------------------------------------------------------------- +// Census 2 and 3: the warp interpreter mirrored with every item index recorded +// -------------------------------------------------------------------------------------------------------------- + +struct Table { + items: Vec<[u32; 16]>, + lines: Vec<[u32; 8]>, +} + +fn build_table(mp: &MixParams, cache: &Cache, threads: usize, plant: Plant, validate: &str) -> (Table, u64) { + let n = 1usize << (DATASET_LOG2 - 4); + let t0 = Instant::now(); + let mut items = vec![[0u32; 16]; n]; + let mut lines = vec![[0u32; 8]; n]; + let mism = AtomicUsize::new(0); + std::thread::scope(|sc| { + let per = n / threads; + for (th, (ic, lc)) in items.chunks_mut(per).zip(lines.chunks_mut(per)).enumerate() { + let mism = &mism; + sc.spawn(move || { + let mut ts = [0u32; 64]; + let mut lib = [[0u32; 16]; 64]; + let base = th * per; + for (bi, (ib, lb)) in ic.chunks_mut(64).zip(lc.chunks_mut(64)).enumerate() { + let start = base + bi * 64; + for k in 0..ib.len() { + ts[k] = (start + k) as u32; + } + derive_traced(&ts[..ib.len()], mp, cache, ib, lb, plant.line_mask()); + let check = plant == Plant::None && (validate == "all" || (validate == "sample" && bi % 64 == 0)); + if check { + igneum_pow::memhard::derive_items(&ts[..ib.len()], mp, cache, &mut lib); + for k in 0..ib.len() { + if lib[k] != ib[k] { + mism.fetch_add(1, Ordering::Relaxed); + } + } + } + } + }); + } + }); + let mm = mism.load(Ordering::Relaxed) as u64; + log!( + "table: {n} items derived with their 8 lines in {:.1} s; library comparison {} ({mm} mismatches)", + t0.elapsed().as_secs_f64(), + if plant == Plant::None { validate } else { "skipped under the plant" } + ); + (Table { items, lines }, mm) +} + +struct ProgramSpec { + label: String, + epoch_seed: Vec, + era: Vec, +} + +fn program_spec(k: u32) -> ProgramSpec { + let genesis = unhex(GENESIS_HEX).unwrap(); + match k { + 1 => ProgramSpec { label: "p1-devnet-epoch0".into(), epoch_seed: genesis.clone(), era: genesis }, + _ => { + let w = |s: String| -> Vec { seed_words_from_bytes(s.as_bytes()).iter().flat_map(|x| x.to_le_bytes()).collect() }; + ProgramSpec { + label: format!("p{k}-attack-f8"), + epoch_seed: w(format!("igneum-attack-f8/program/{k}")), + era: w(format!("igneum-attack-f8/era/{k}")), + } + } + } +} + +/// The mirror of `verify::interpret_warp_init` for class v4 (the ops a v4 program can hold), reading dataset words +/// from the item table and reporting every load's item index. +/// The mirror of `verify::interpret_warp_init` for class v4 (the ops a v4 program can hold), reading dataset words +/// from the item table and reporting every load's item index and, per load position, how many lanes' source +/// register was saturated (0 or all ones) at the load. +struct Mirror<'a> { + program: &'a Program, + era: Option<&'a EraParams>, + layout: Layout, + mask: u32, + log2: u32, + table: &'a Table, + plant_site: Option, +} + +const POS: usize = 128; + +struct Sink { + /// the item index of every load position, per lane + items: Vec<[u32; LANES]>, + /// the masked word index (the address before the layout split) of every load position, per lane + addrs: Vec<[u32; LANES]>, + /// lanes whose source register read 0 or 2^32 - 1 at the load, per position + sat: [u32; POS], +} + +impl Sink { + fn new() -> Sink { + Sink { items: Vec::with_capacity(POS), addrs: Vec::with_capacity(POS), sat: [0; POS] } + } +} + +impl<'a> Mirror<'a> { + #[inline(always)] + fn step(&self, ins: &Instr, r: &mut [[u32; LANES]; 8], sel: &[u32; LANES], sink: &mut Sink, site: &mut usize) { + let d = ins.dst as usize; + let a = ins.src as usize; + match ins.op { + Op::Add => { + let (imm, imm2, bit) = (ins.imm, ins.imm2, ins.bit as u32); + let src = r[a]; + for lane in 0..LANES { + let s = (sel[lane] >> bit) & 1; + let c = if s != 0 { imm2 } else { imm }; + r[d][lane] = r[d][lane].wrapping_add(src[lane]).wrapping_add(c); + } + } + Op::Sub => { + let src = r[a]; + for lane in 0..LANES { + r[d][lane] = r[d][lane].wrapping_sub(src[lane]); + } + } + Op::Mul => { + let src = r[a]; + for lane in 0..LANES { + r[d][lane] = r[d][lane].wrapping_mul(src[lane]); + } + } + Op::MulHi => { + let src = r[a]; + for lane in 0..LANES { + r[d][lane] = ((r[d][lane] as u64 * src[lane] as u64) >> 32) as u32; + } + } + Op::Xor => { + let src = r[a]; + for lane in 0..LANES { + r[d][lane] ^= src[lane]; + } + } + Op::Or => { + let src = r[a]; + for lane in 0..LANES { + r[d][lane] |= src[lane]; + } + } + Op::Rotl => { + let n = ins.rot; + for lane in 0..LANES { + r[d][lane] = r[d][lane].rotate_left(n); + } + } + Op::Rotr => { + let src = r[a]; + for lane in 0..LANES { + r[d][lane] = r[d][lane].rotate_right(src[lane] & 31); + } + } + Op::Mad => { + let src = r[a]; + let src2 = r[ins.src2 as usize]; + for lane in 0..LANES { + r[d][lane] = src[lane].wrapping_mul(src2[lane]).wrapping_add(r[d][lane]); + } + } + Op::Shfl => { + let src = r[a]; + let m = ins.mask as usize; + for lane in 0..LANES { + r[d][lane] ^= src[lane ^ m]; + } + } + Op::Load => { + assert_eq!(ins.width, 1, "class v4 loads one word"); + let mut ts = [0u32; LANES]; + let mut ws = [0u32; LANES]; + let p = sink.items.len(); + for lane in 0..LANES { + let x = r[a][lane]; + if x == 0 || x == u32::MAX { + sink.sat[p] += 1; + } + let idx = load_index(self.era, ins, x, self.mask, self.log2); + let (t, j) = self.layout.split(idx); + let t = if self.plant_site == Some(*site) { 0x00_1234 } else { t }; + ts[lane] = t; + ws[lane] = idx; + r[d][lane] ^= self.table.items[t as usize][j as usize]; + } + sink.items.push(ts); + sink.addrs.push(ws); + *site += 1; + } + Op::WLoad | Op::Scratch | Op::Hot => panic!("op {:?} is not a class v4 op", ins.op), + } + } + + /// Hashes of the warp at `base_nonce`; the sink carries the item index of every load (128 x 32). + fn warp(&self, base_nonce: u32, sink: &mut Sink) -> [u64; LANES] { + let seed = &self.program.seed; + let mut r = [[0u32; LANES]; 8]; + for lane in 0..LANES { + let nonce = base_nonce.wrapping_add(lane as u32); + for i in 0..8 { + let mut x = nonce ^ seed[i]; + x = x.wrapping_add(0x9e3779b9u32.wrapping_mul(i as u32 + 1)); + x = splitmix32(x); + r[i][lane] = x ^ seed[(i + 1) & 7]; + } + } + sink.items.clear(); + sink.addrs.clear(); + sink.sat = [0; POS]; + for _ in 0..ITERATIONS { + let sel = r[0]; + let mut site = 0usize; + for ins in &self.program.instrs { + self.step(ins, &mut r, &sel, sink, &mut site); + } + for _ in 0..self.program.shadow_reps() { + for ins in &self.program.shadow { + self.step(ins, &mut r, &sel, sink, &mut site); + } + } + } + let mut hashes = [0u64; LANES]; + for lane in 0..LANES { + let lo = r[0][lane] ^ r[1][lane].rotate_left(7) ^ r[2][lane].rotate_left(14) ^ r[3][lane].rotate_left(21); + let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27); + hashes[lane] = ((hi as u64) << 32) | lo as u64; + } + hashes + } +} + +fn pct(h: &[u64], q: f64) -> usize { + let total: u64 = h.iter().sum(); + let want = (total as f64 * q).ceil().max(1.0) as u64; + let mut acc = 0u64; + for (i, c) in h.iter().enumerate() { + acc += c; + if acc >= want { + return i; + } + } + h.len() - 1 +} + +fn dist_line(label: &str, h: &[u64]) -> String { + let total: u64 = h.iter().sum(); + let min = h.iter().position(|&c| c > 0).unwrap_or(0); + let max = h.iter().rposition(|&c| c > 0).unwrap_or(0); + let mean = h.iter().enumerate().map(|(i, c)| i as f64 * *c as f64).sum::() / total.max(1) as f64; + format!( + "{label}: n {total} min {min} p1 {} median {} p99 {} max {max} mean {mean:.4}", + pct(h, 0.01), + pct(h, 0.5), + pct(h, 0.99) + ) +} + +/// The window of every load site as an item range: `(first item, items)`; the era window is the top `k` bits of +/// the 28-bit word index and every layout position lies below 16, so the window is the same aligned range of items. +fn site_item_windows(program: &Program, mask: u32, log2: u32) -> Vec<(u32, u32)> { + program + .instrs + .iter() + .filter(|i| i.op == Op::Load) + .map(|i| { + let (wm, off) = igneum_pow::verify::window(i, mask, log2); + (off >> 4, (wm >> 4) + 1) + }) + .collect() +} + +/// The expected reads per item under the window layer alone (every site uniform on its own window), as a density per +/// quarter of the item space (windows are the dataset, a half or a quarter, aligned). +fn window_density(windows: &[(u32, u32)], items_n: usize, reads_per_site: f64) -> [f64; 4] { + let q = items_n as u32 / 4; + let mut d = [0f64; 4]; + for &(first, n) in windows { + for (k, dq) in d.iter_mut().enumerate() { + let qs = k as u32 * q; + if qs >= first && qs < first + n { + *dq += reads_per_site / n as f64; + } + } + } + d +} + +/// Statistics of `counts` against a per-quarter expected density: chi-square per dof, the largest and smallest +/// bucket in sigma of their own expectation, buckets of `per` items. +fn stats_vs_density(counts: &[u64], dens: &[f64; 4], per: usize) -> (f64, f64, u64, usize, f64) { + let items_n = counts.len(); + let q = items_n / 4; + let mut chi2 = 0f64; + let mut z_max = f64::MIN; + let mut z_min = f64::MAX; + let mut max = 0u64; + let mut argmax = 0usize; + let mut dof = 0usize; + for (b, c) in counts.chunks(per).enumerate() { + let e = dens[(b * per) / q] * per as f64; + if e <= 0.0 { + continue; + } + let s: u64 = c.iter().sum(); + let z = (s as f64 - e) / e.sqrt(); + chi2 += z * z; + dof += 1; + if z > z_max { + z_max = z; + max = s; + argmax = b; + } + if z < z_min { + z_min = z; + } + } + (chi2 / (dof.max(2) - 1) as f64, z_max, max, argmax, z_min) +} + +/// The windowed control: `n` reads, site `i mod 16`, uniform on the site's window. +fn windowed_control(windows: &[(u32, u32)], n: u64, seed: u64, threads: usize, hist: &[AtomicU32]) { + std::thread::scope(|sc| { + for th in 0..threads { + let hist = &hist; + sc.spawn(move || { + let mut s = SplitMix64::new(seed ^ (th as u64).wrapping_mul(0x9E3779B97F4A7C15)); + let per = n / threads as u64 + if (th as u64) < n % threads as u64 { 1 } else { 0 }; + for i in 0..per { + let (first, cnt) = windows[(i % 16) as usize]; + let t = first + (s.next() as u32 & (cnt - 1)); + hist[t as usize].fetch_add(1, Ordering::Relaxed); + } + }); + } + }); +} + +struct ProgramSummary { + label: String, + id: u64, + attempt: u32, + /// X_f against the windowed control at f = 0.1%, 0.5%, 1% + x: [f64; 3], + hot: bool, + /// windowed chi2/dof of the 64-item buckets and the largest bucket in sigma + chi2_w: f64, + z_w: f64, + /// the largest share of one hi16 bucket (256 items) at any load position + max_bucket_share: f64, + /// the largest saturated-source share at any load position + max_sat_share: f64, + /// acceptance-style 2,048-evaluation metrics: the largest count of one address at one position, the largest + /// saturated-source count at one position + acc_max_addr: u32, + acc_max_sat: u32, + /// S_0.1% over the window-model control and over the flat control + ratio_w: f64, + ratio_flat: f64, + /// top-f shares measured and under the window-model control, f = 0.1% and 1% + s01: f64, + e01: f64, + s10: f64, + e10: f64, + /// the hottest item, its count, and the lossy site whose saturated (or one-bit-off) source maps to it, if any + hottest: u32, + hottest_count: u64, + hottest_source: String, + /// share of each site's reads that land on the top 0.1% of items (the attribution pass; zeros without --diag) + by_site_share: [f64; 16], +} + +#[allow(clippy::too_many_arguments)] +fn run_program( + spec: &ProgramSpec, + ds: DatasetSource, + table: &Table, + table_mism: u64, + day: u64, + nonces: u64, + threads: usize, + plant: Plant, + check_every: u64, + diag: bool, + verbose: bool, + out_dir: &str, +) -> (ProgramSummary, DatasetSource) { + let program = Epoch::chain_program(&spec.epoch_seed, Some(&spec.era), ProgramClass::V4, &spec.label); + assert_eq!(program.generator, 4); + assert_eq!(program.class.mixer_mult, 8); + assert_eq!(program.class.shadow.map(|s| (s.instrs, s.reps)), Some((256, 27))); + let era = program.class.era.expect("class v4 draws the era"); + let n_loads = program.instrs.iter().filter(|i| i.op == Op::Load).count(); + assert_eq!(n_loads, 16); + log!( + "program {}: epoch seed {} era seed {} id {:016x} attempt {} class {} op mix {}; era stride mul {:#010x} rot {} interleave {:?} windows {}", + spec.label, + hex(&spec.epoch_seed), + hex(&spec.era), + program.program_id(), + program.attempt, + program.class.name(), + program.op_mix(), + era.stride_mul, + era.stride_rot, + era.pos, + program.instrs.iter().enumerate().filter(|(_, i)| i.op == Op::Load).map(|(k, i)| format!("{k}:{}:{}", i.win, i.off)).collect::>().join(" ") + ); + let sites: Vec = program.instrs.iter().enumerate().filter(|(_, i)| i.op == Op::Load).map(|(k, _)| k).collect(); + if verbose { + // the static shape of every load site: its source register, the last base-program writer of that register + // before the site (cyclic) and whether that writer injects (add, sub, xor, mad, shfl, load), and how often the + // shadow block writes the register (the shadow runs between iteration i's instruction 63 and iteration i+1's + // instruction 0, outside the acceptance rule) + let instrs = &program.instrs; + for (k, ins) in instrs.iter().enumerate() { + if ins.op != Op::Load { + continue; + } + let src = ins.src; + let mut writer = String::from("none"); + let mut chain = Vec::new(); + for back in 1..instrs.len() { + let j = (k + instrs.len() - back) % instrs.len(); + if instrs[j].dst == src { + if writer == "none" { + writer = format!("{} at {j}", instrs[j].op.name()); + } + chain.push(format!("{}@{j}", instrs[j].op.name())); + if instrs[j].op.injects() { + break; + } + } + } + let mut sh: std::collections::BTreeMap<&str, usize> = std::collections::BTreeMap::new(); + for s in program.shadow.iter().filter(|s| s.dst == src) { + *sh.entry(s.op.name()).or_insert(0) += 1; + } + log!( + "load site instr {k}: src r{src} win {} off {}; last base writer {writer}; writers back to the last injecting one: {}; shadow writes of r{src} per rep: {}", + ins.win, + ins.off, + chain.join(" "), + sh.iter().map(|(o, n)| format!("{o}={n}")).collect::>().join(" ") + ); + } + } + let epoch = Epoch { program, dataset: ds }; + let plant_site = if plant == Plant::ConstItem { Some(0) } else { None }; + let mirror = Mirror { + program: &epoch.program, + era: epoch.program.class.era.as_ref(), + layout: epoch.program.class.layout(), + mask: epoch.dataset.mask, + log2: epoch.dataset.log2_words, + table, + plant_site, + }; + let items_n = 1usize << (DATASET_LOG2 - 4); + let windows = site_item_windows(&epoch.program, epoch.dataset.mask, epoch.dataset.log2_words); + let dens = window_density(&windows, items_n, nonces as f64 * 8.0); + log!( + "window layer: site item windows (first, items) {}; expected reads per item by quarter {:.3} {:.3} {:.3} {:.3} (flat uniform {:.3})", + windows.iter().map(|(f, n)| format!("({:#x},2^{})", f, n.trailing_zeros())).collect::>().join(" "), + dens[0], + dens[1], + dens[2], + dens[3], + nonces as f64 * 128.0 / items_n as f64 + ); + let item_hist = atomic_vec(items_n); + let pos_hi: Vec> = (0..POS).map(|_| atomic_vec(1 << 16)).collect(); + let pos_sat: Vec = atomic_vec(POS); + let warps = nonces / LANES as u64; + let next = AtomicUsize::new(0); + let checked = AtomicUsize::new(0); + let hash_mism = AtomicUsize::new(0); + let max_lines_hash = 8 * 128; + let max_lines_warp = max_lines_hash * LANES; + // the acceptance-style metric on the first 64 warps (2,048 evaluations): per position, the count of every + // masked address and of saturated sources + let acc = std::sync::Mutex::new((vec![std::collections::HashMap::::new(); POS], [0u32; POS])); + let t1 = Instant::now(); + let dists: Vec<(Vec, Vec, Vec, Vec)> = std::thread::scope(|sc| { + let mut hs = Vec::new(); + for _ in 0..threads { + let (next, checked, hash_mism, item_hist, mirror, epoch, pos_hi, pos_sat, acc) = + (&next, &checked, &hash_mism, &item_hist, &mirror, &epoch, &pos_hi, &pos_sat, &acc); + hs.push(sc.spawn(move || { + let mut lines_hash = vec![0u64; max_lines_hash + 1]; + let mut items_hash = vec![0u64; 129]; + let mut lines_warp = vec![0u64; max_lines_warp + 1]; + let mut items_warp = vec![0u64; 128 * LANES + 1]; + let mut sink = Sink::new(); + let mut lane_lines: Vec = Vec::with_capacity(max_lines_hash); + let mut lane_items: Vec = Vec::with_capacity(128); + let mut warp_lines: Vec = Vec::with_capacity(max_lines_warp); + let mut warp_items: Vec = Vec::with_capacity(128 * LANES); + loop { + let w = next.fetch_add(1, Ordering::Relaxed) as u64; + if w >= warps { + break; + } + let base = (w * LANES as u64) as u32; + let hashes = mirror.warp(base, &mut sink); + assert_eq!(sink.items.len(), POS); + if plant == Plant::None && (w < 64 || w % check_every == 0) { + let lib = epoch.hash_warp(base); + checked.fetch_add(1, Ordering::Relaxed); + if lib != hashes { + hash_mism.fetch_add(1, Ordering::Relaxed); + } + } + if w < 64 { + let mut g = acc.lock().unwrap(); + for p in 0..POS { + for lane in 0..LANES { + *g.0[p].entry(sink.addrs[p][lane]).or_insert(0) += 1; + } + g.1[p] += sink.sat[p]; + } + } + warp_lines.clear(); + warp_items.clear(); + for lane in 0..LANES { + lane_lines.clear(); + lane_items.clear(); + for load in &sink.items { + let t = load[lane]; + lane_items.push(t); + lane_lines.extend_from_slice(&mirror.table.lines[t as usize]); + } + warp_items.extend_from_slice(&lane_items); + warp_lines.extend_from_slice(&lane_lines); + lane_items.sort_unstable(); + lane_items.dedup(); + lane_lines.sort_unstable(); + lane_lines.dedup(); + items_hash[lane_items.len()] += 1; + lines_hash[lane_lines.len()] += 1; + } + for (p, load) in sink.items.iter().enumerate() { + for lane in 0..LANES { + let t = load[lane]; + item_hist[t as usize].fetch_add(1, Ordering::Relaxed); + pos_hi[p][(t >> 8) as usize].fetch_add(1, Ordering::Relaxed); + } + pos_sat[p].fetch_add(sink.sat[p], Ordering::Relaxed); + } + warp_items.sort_unstable(); + warp_items.dedup(); + warp_lines.sort_unstable(); + warp_lines.dedup(); + items_warp[warp_items.len()] += 1; + lines_warp[warp_lines.len()] += 1; + } + (lines_hash, items_hash, lines_warp, items_warp) + })); + } + hs.into_iter().map(|h| h.join().unwrap()).collect() + }); + let el = t1.elapsed().as_secs_f64(); + let mut lines_hash = vec![0u64; max_lines_hash + 1]; + let mut items_hash = vec![0u64; 129]; + let mut lines_warp = vec![0u64; max_lines_warp + 1]; + let mut items_warp = vec![0u64; 128 * LANES + 1]; + for (a, b, c, d) in &dists { + for (i, v) in a.iter().enumerate() { + lines_hash[i] += v; + } + for (i, v) in b.iter().enumerate() { + items_hash[i] += v; + } + for (i, v) in c.iter().enumerate() { + lines_warp[i] += v; + } + for (i, v) in d.iter().enumerate() { + items_warp[i] += v; + } + } + let ck = checked.load(Ordering::Relaxed); + let hm = hash_mism.load(Ordering::Relaxed); + log!( + "{} warps ({} nonces) interpreted in {el:.1} s ({:.3} ms per warp per thread); Epoch::hash_warp agreement on {ck} warps: {hm} mismatches{}", + warps, + warps * LANES as u64, + el * 1e3 * threads as f64 / warps as f64, + if plant != Plant::None { " (library comparison skipped under the plant)" } else { "" } + ); + if hm > 0 || table_mism > 0 { + log!("THE MIRROR DISAGREES WITH THE LIBRARY (hashes {hm}, items {table_mism}); nothing below is trusted"); + } + let tag = format!("warps-{}-d{day}-n{nonces}-{}", spec.label, plant.name()); + log!("{}", dist_line(&format!("{} distinct lines per hash", spec.label), &lines_hash)); + log!("{}", dist_line(&format!("{} distinct items per hash", spec.label), &items_hash)); + log!("{}", dist_line(&format!("{} distinct lines per warp", spec.label), &lines_warp)); + log!("{}", dist_line(&format!("{} distinct items per warp", spec.label), &items_warp)); + let exp = |k: f64, l: f64| l * (1.0 - (1.0 - 1.0 / l).powf(k)); + log!( + "uniform expectation: distinct lines per hash {:.3} of 1024 reads, per warp {:.1} of 32768 reads (2^22 lines); distinct items per hash {:.4} of 128, per warp {:.2} of 4096 (2^24 items)", + exp(1024.0, 4194304.0), + exp(32768.0, 4194304.0), + exp(128.0, 16777216.0), + exp(4096.0, 16777216.0) + ); + { + let mut f = std::fs::File::create(format!("{out_dir}/{tag}-distinct.txt")).unwrap(); + for (name, h) in [("lines_hash", &lines_hash), ("items_hash", &items_hash), ("lines_warp", &lines_warp), ("items_warp", &items_warp)] { + writeln!(f, "# distinct {name}: value count").unwrap(); + for (i, c) in h.iter().enumerate().filter(|(_, c)| **c > 0) { + writeln!(f, "{name} {i} {c}").unwrap(); + } + } + } + // per-position diagnostics: the largest hi16 bucket share and the saturated-source share + let mut max_bucket_share = 0f64; + let mut max_bucket_pos = 0usize; + let mut max_sat_share = 0f64; + let mut max_sat_pos = 0usize; + { + let mut f = std::fs::File::create(format!("{out_dir}/{tag}-positions.txt")).unwrap(); + writeln!(f, "# position iteration site instr max_hi16_bucket_share saturated_share").unwrap(); + for p in 0..POS { + let mx = pos_hi[p].iter().map(|x| x.load(Ordering::Relaxed)).max().unwrap_or(0) as f64 / nonces as f64; + let sat = pos_sat[p].load(Ordering::Relaxed) as f64 / nonces as f64; + writeln!(f, "{p} {} {} {} {:.6} {:.6}", p / 16, p % 16, sites[p % 16], mx, sat).unwrap(); + if mx > max_bucket_share { + max_bucket_share = mx; + max_bucket_pos = p; + } + if sat > max_sat_share { + max_sat_share = sat; + max_sat_pos = p; + } + } + } + log!( + "positions: largest hi16-bucket (256 items) share {:.4}% at p{} (iteration {}, site {}, instr {}); windowed expectation {:.4}%; largest saturated-source share {:.4}% at p{} (site {}, instr {})", + max_bucket_share * 100.0, + max_bucket_pos, + max_bucket_pos / 16, + max_bucket_pos % 16, + sites[max_bucket_pos % 16], + 100.0 * 256.0 / (windows[max_bucket_pos % 16].1 as f64), + max_sat_share * 100.0, + max_sat_pos, + max_sat_pos % 16, + sites[max_sat_pos % 16] + ); + // the acceptance-style metric over the first 2,048 evaluations + let (acc_max_addr, acc_max_sat, acc_pos, acc_sat_pos) = { + let g = acc.lock().unwrap(); + let mut ma = 0u32; + let mut mp = 0usize; + let mut ms = 0u32; + let mut msp = 0usize; + for p in 0..POS { + let m = g.0[p].values().copied().max().unwrap_or(0); + if m > ma { + ma = m; + mp = p; + } + if g.1[p] > ms { + ms = g.1[p]; + msp = p; + } + } + (ma, ms, mp, msp) + }; + log!( + "acceptance-style (2,048 evaluations, the rule's sample size): the most repeated address at one position {} of 2048 at p{} (site {}, instr {}); saturated sources at one position {} of 2048 at p{} (site {}, instr {}); uniform expectation: repeats 1 to 2, saturated 0", + acc_max_addr, + acc_pos, + acc_pos % 16, + sites[acc_pos % 16], + acc_max_sat, + acc_sat_pos, + acc_sat_pos % 16, + sites[acc_sat_pos % 16] + ); + // the cross-hash item histogram: flat statistics, the windowed null, the windowed control, the hot-set test + let snap = snapshot(&item_hist); + let s_items = stats_of(snap.iter().copied(), items_n); + log!("{}", s_items.line(&format!("{} item histogram (flat)", spec.label))); + let (chi2_w1, z_w1, max_w1, arg_w1, zmin_w1) = stats_vs_density(&snap, &dens, 1); + let (chi2_w, z_w, max_w, arg_w, zmin_w) = stats_vs_density(&snap, &dens, 64); + log!( + "{} item histogram against the window density: full 2^24 chi2/dof {:.5} largest {} (item {:#x}) at {:+.2} sigma smallest at {:+.2}; buckets64 chi2/dof {:.5} largest {} (bucket {}) at {:+.2} sigma smallest at {:+.2}", + spec.label, + chi2_w1, + max_w1, + arg_w1, + z_w1, + zmin_w1, + chi2_w, + max_w, + arg_w, + z_w, + zmin_w + ); + let total_reads = nonces * 128; + let ctrl = atomic_vec(items_n); + windowed_control(&windows, total_reads, 0xF8_3333 ^ program_hash(&spec.label), threads, &ctrl); + let cs = snapshot(&ctrl); + let c_items = stats_of(cs.iter().copied(), items_n); + let (c_chi2_w, c_z_w, _, _, c_zmin_w) = stats_vs_density(&cs, &dens, 64); + let (c_chi2_w1, c_z_w1, _, _, _) = stats_vs_density(&cs, &dens, 1); + log!("{}", c_items.line(&format!("{} WINDOWED CONTROL item histogram (flat)", spec.label))); + log!( + "{} WINDOWED CONTROL against the window density: full chi2/dof {:.5} largest {:+.2} sigma; buckets64 chi2/dof {:.5} largest {:+.2} smallest {:+.2}", + spec.label, + c_chi2_w1, + c_z_w1, + c_chi2_w, + c_z_w, + c_zmin_w + ); + let f1 = z_w > 6.0 || zmin_w < -6.0; + log!( + "6-sigma {} items buckets64 against the window density: largest bucket {:+.2} sigma, smallest {:+.2} -> {}", + spec.label, + z_w, + zmin_w, + if f1 { "FLAGGED (beyond 6 sigma)" } else { "within 6 sigma" } + ); + let f3 = hot_set_test(&format!("{} items (windowed control)", spec.label), &s_items, &c_items); + let flat = atomic_vec(items_n); + control(items_n, total_reads, 0xF8_4444 ^ program_hash(&spec.label), threads, &flat); + let fs = snapshot(&flat); + let f_items = stats_of(fs.iter().copied(), items_n); + log!("{}", f_items.line(&format!("{} FLAT CONTROL item histogram", spec.label))); + let _ = hot_set_test(&format!("{} items (flat control, the auditor's first view)", spec.label), &s_items, &f_items); + let mut x = [0f64; 3]; + let mut ratio_w = [0f64; 3]; + let mut ratio_flat = [0f64; 3]; + for k in 0..3 { + x[k] = s_items.top[k].1 - c_items.top[k].1; + ratio_w[k] = s_items.top[k].1 / c_items.top[k].1; + ratio_flat[k] = s_items.top[k].1 / f_items.top[k].1; + } + log!( + "ratio {}: top 0.1% / 0.5% / 1% share over the WINDOW-MODEL control {:.4}x / {:.4}x / {:.4}x (gate 1.2x at 0.1%: {}); over the FLAT control {:.4}x / {:.4}x / {:.4}x", + spec.label, + ratio_w[0], + ratio_w[1], + ratio_w[2], + if ratio_w[0] <= 1.2 { "within" } else { "BEYOND" }, + ratio_flat[0], + ratio_flat[1], + ratio_flat[2] + ); + { + let mut g = std::fs::File::create(format!("{out_dir}/{tag}-items.u32le")).unwrap(); + let mut bytes = Vec::with_capacity(items_n * 4); + for c in &snap { + bytes.extend_from_slice(&(*c as u32).to_le_bytes()); + } + g.write_all(&bytes).unwrap(); + } + let mut by_site_share = [0f64; 16]; + if diag { + // attribution: the hottest items (the top 0.1% by count, and the top 8) traced back to the load positions + let want = (items_n as f64 / 1000.0).round() as u64; + let mut coc: std::collections::BTreeMap = std::collections::BTreeMap::new(); + for &c in &snap { + *coc.entry(c).or_insert(0) += 1; + } + let mut left = want; + let mut thr = 0u64; + for (&c, &n) in coc.iter().rev() { + thr = c; + if n >= left { + break; + } + left -= n; + } + let mut hot = vec![0u8; items_n]; + let mut n_hot = 0u64; + let mut hot_reads = 0u64; + for (t, &c) in snap.iter().enumerate() { + if c >= thr { + hot[t] = 1; + n_hot += 1; + hot_reads += c; + } + } + let mut top8: Vec<(u64, usize)> = snap.iter().enumerate().map(|(t, &c)| (c, t)).collect(); + top8.sort_unstable_by(|a, b| b.cmp(a)); + top8.truncate(8); + log!( + "attribution: hot threshold count >= {thr} marks {n_hot} items ({:.4}% of items) holding {hot_reads} reads ({:.4}% of reads); top 8 items {}", + n_hot as f64 * 100.0 / items_n as f64, + hot_reads as f64 * 100.0 / total_reads as f64, + top8.iter().map(|(c, t)| format!("t={t:#08x}:{c}")).collect::>().join(" ") + ); + let per_pos: Vec = (0..POS).map(|_| std::sync::atomic::AtomicU64::new(0)).collect(); + let per_top: Vec> = (0..8).map(|_| (0..POS).map(|_| std::sync::atomic::AtomicU64::new(0)).collect()).collect(); + let next2 = AtomicUsize::new(0); + std::thread::scope(|sc| { + for _ in 0..threads { + let (next2, mirror, hot, top8, per_pos, per_top) = (&next2, &mirror, &hot, &top8, &per_pos, &per_top); + sc.spawn(move || { + let mut sink = Sink::new(); + loop { + let w = next2.fetch_add(1, Ordering::Relaxed) as u64; + if w >= warps { + break; + } + mirror.warp((w * LANES as u64) as u32, &mut sink); + for (p, load) in sink.items.iter().enumerate() { + for lane in 0..LANES { + let t = load[lane] as usize; + if hot[t] != 0 { + per_pos[p].fetch_add(1, Ordering::Relaxed); + } + for (k, (_, tt)) in top8.iter().enumerate() { + if *tt == t { + per_top[k][p].fetch_add(1, Ordering::Relaxed); + } + } + } + } + } + }); + } + }); + let expect = n_hot as f64 / items_n as f64; + let mut rows: Vec<(usize, u64)> = per_pos.iter().enumerate().map(|(p, c)| (p, c.load(Ordering::Relaxed))).collect(); + rows.sort_unstable_by(|a, b| b.1.cmp(&a.1)); + log!( + "attribution: reads on hot items per position (flat expectation {:.4}% of each position's {nonces} reads); the 24 largest: {}", + expect * 100.0, + rows.iter().take(24).map(|(p, c)| format!("p{p}(it{} s{})={:.3}%", p / 16, p % 16, *c as f64 * 100.0 / nonces as f64)).collect::>().join(" ") + ); + let mut by_iter = vec![0u64; ITERATIONS]; + let mut by_site = vec![0u64; 16]; + for (p, c) in &rows { + by_iter[p / 16] += c; + by_site[p % 16] += c; + } + for s in 0..16 { + by_site_share[s] = by_site[s] as f64 / (nonces * 8) as f64; + } + log!( + "attribution: hot reads by iteration {} ; by site {}", + by_iter.iter().enumerate().map(|(i, c)| format!("it{i}={:.3}%", *c as f64 * 100.0 / (nonces * 16) as f64)).collect::>().join(" "), + by_site.iter().enumerate().map(|(i, c)| format!("s{i}={:.3}%", *c as f64 * 100.0 / (nonces * 8) as f64)).collect::>().join(" ") + ); + // per-site contribution table: window, share of the site's reads into the top 0.1% set, the entropy of the + // site's item index over nonces (hi16 buckets of 256 items: a window of 2^(24 - k_off) items is 2^(16 - k_off) + // buckets, so uniform on the window = 16 - k_off bits; the runs of 09:04 UTC printed 16 - 2 k_off by mistake), + // saturated-source share, largest bucket share + let load_instrs: Vec<&Instr> = epoch.program.instrs.iter().filter(|i| i.op == Op::Load).collect(); + for s in 0..16 { + let mut acc = vec![0u64; 1 << 16]; + let mut hot_s = 0u64; + let mut sat_s = 0u64; + for it in 0..ITERATIONS { + let p = it * 16 + s; + for (b, c) in pos_hi[p].iter().enumerate() { + acc[b] += c.load(Ordering::Relaxed) as u64; + } + hot_s += per_pos[p].load(Ordering::Relaxed); + sat_s += pos_sat[p].load(Ordering::Relaxed) as u64; + } + let n_s = (nonces * ITERATIONS as u64) as f64; + let mut h = 0f64; + let mut mx = 0u64; + for &c in &acc { + if c > 0 { + let pr = c as f64 / n_s; + h -= pr * pr.log2(); + } + mx = mx.max(c); + } + let ins = load_instrs[s]; + log!( + "site {s} (instr {}, src r{}, k_off {} offset {}, window 2^{} items): share of its reads into the top 0.1% {:.4}% (flat expectation {:.4}%); index entropy {:.3} bits of {} uniform-on-window; saturated source {:.4}%; largest 256-item bucket {:.4}% (window expectation {:.4}%)", + sites[s], + ins.src, + ins.win, + ins.off, + windows[s].1.trailing_zeros(), + hot_s as f64 * 100.0 / n_s, + expect * 100.0, + h, + 16 - ins.win.min(2) as u32, + sat_s as f64 * 100.0 / n_s, + mx as f64 * 100.0 / n_s, + 100.0 * 256.0 / windows[s].1 as f64 + ); + } + for (k, (c, t)) in top8.iter().enumerate() { + let mut pos: Vec<(usize, u64)> = per_top[k].iter().enumerate().map(|(p, x)| (p, x.load(Ordering::Relaxed))).filter(|(_, x)| *x > 0).collect(); + pos.sort_unstable_by(|a, b| b.1.cmp(&a.1)); + log!( + "attribution: item {t:#08x} ({c} reads) read from positions {}", + pos.iter().take(12).map(|(p, x)| format!("p{p}(it{} s{})x{x}", p / 16, p % 16)).collect::>().join(" ") + ); + } + } + // the hottest item and the lossy site whose saturated source maps to it under the era map (all ones, zero, or + // one bit off either), with that site's source register and its last base-program writer + let (hottest, hottest_count) = snap.iter().enumerate().map(|(t, &c)| (t as u32, c)).max_by_key(|&(_, c)| c).unwrap(); + let hottest_source = { + let instrs = &epoch.program.instrs; + let layout = epoch.program.class.layout(); + let era = epoch.program.class.era.as_ref(); + let mut found = String::from("none"); + for (k, ins) in instrs.iter().enumerate() { + if ins.op != Op::Load { + continue; + } + let img = |x: u32| layout.split(load_index(era, ins, x, epoch.dataset.mask, epoch.dataset.log2_words)).0; + let mut cands: Vec<(u32, &str)> = vec![(u32::MAX, "all-ones"), (0, "zero")]; + for b in 0..32 { + cands.push((u32::MAX ^ (1 << b), "one-zero-bit")); + cands.push((1 << b, "one-one-bit")); + } + if let Some((_, what)) = cands.iter().find(|(x, _)| img(*x) == hottest) { + let mut writer = String::from("none"); + for back in 1..instrs.len() { + let j = (k + instrs.len() - back) % instrs.len(); + if instrs[j].dst == ins.src { + writer = format!("{}@{j}", instrs[j].op.name()); + break; + } + } + found = format!("site-instr{k}:r{}:{what}:last-writer-{writer}", ins.src); + break; + } + } + found + }; + log!( + "hottest item {} ({:#08x}): {} reads ({:.4}% of all); predicted source {}", + spec.label, + hottest, + hottest_count, + hottest_count as f64 * 100.0 / total_reads as f64, + hottest_source + ); + let hot = f3; + log!( + "program {} DONE: {} nonces; 6-sigma (windowed) {}; hot-set {}; verdict {}", + spec.label, + nonces, + if f1 { "FLAGGED" } else { "clear" }, + if hot { "FLAGGED" } else { "clear" }, + if f1 || hot { "FLAGGED" } else { "PASS (no test fired)" } + ); + let summary = ProgramSummary { + label: spec.label.clone(), + id: epoch.program.program_id(), + attempt: epoch.program.attempt, + x, + hot, + chi2_w, + z_w, + max_bucket_share, + max_sat_share, + acc_max_addr, + acc_max_sat, + ratio_w: ratio_w[0], + ratio_flat: ratio_flat[0], + s01: s_items.top[0].1, + e01: c_items.top[0].1, + s10: s_items.top[2].1, + e10: c_items.top[2].1, + hottest, + hottest_count, + hottest_source, + by_site_share, + }; + (summary, epoch.dataset) +} + +fn program_hash(label: &str) -> u64 { + igneum_pow::seed::fnv1a64(label.as_bytes()) +} + +#[allow(clippy::too_many_arguments)] +fn census_warps(programs: &[u32], day: u64, nonces: u64, threads: usize, plant: Plant, validate: &str, check_every: u64, diag: bool, out_dir: &str) { + log!( + "census warps: programs {:?} day {day} nonces {nonces} threads {threads} plant {} validate {validate} check_every {check_every} diag {diag}", + programs, + plant.name() + ); + let t0 = Instant::now(); + let mut ds: DatasetSource = Epoch::chain_dataset_day(&day_bytes(day), ProgramClass::V4, 0, DATASET_LOG2); + assert_eq!(ds.log2_words, DATASET_LOG2); + let (table, table_mism) = { + let mh = ds.memhard().unwrap(); + log!("day {day}: cache filled in {:.2} s, fnv {:016x}", t0.elapsed().as_secs_f64(), mh.cache.fnv1a64()); + build_table(&mh.params, &mh.cache, threads, plant, validate) + }; + let verbose = programs.len() <= 3; + let mut summaries = Vec::new(); + for &k in programs { + let spec = program_spec(k); + let (s, d) = run_program(&spec, ds, &table, table_mism, day, nonces, threads, plant, check_every, diag, verbose, out_dir); + ds = d; + summaries.push(s); + } + if programs.len() > 1 { + // the re-gate format (Counter ASIC lane, 7 October 2026): one line per seed, then PASS or FAIL against 1.2x + let mut over = Vec::new(); + for s in &summaries { + log!( + "SEED {} id {:016x} attempt {} S0.1 {:.5}% null0.1 {:.5}% ratio {:.4}x S1.0 {:.5}% null1.0 {:.5}% hottest {:#08x}:{} source {} by-site {}", + s.label, + s.id, + s.attempt, + s.s01 * 100.0, + s.e01 * 100.0, + s.ratio_w, + s.s10 * 100.0, + s.e10 * 100.0, + s.hottest, + s.hottest_count, + s.hottest_source, + if diag { s.by_site_share.iter().enumerate().map(|(i, x)| format!("s{i}={:.3}%", x * 100.0)).collect::>().join(" ") } else { "(no --diag)".to_string() } + ); + if s.ratio_w > 1.2 { + over.push(format!("{}({:.3}x)", s.label, s.ratio_w)); + } + } + log!( + "CENSUS {}: {} seeds at {} nonces, {} over 1.2x of the window model at the top 0.1%{}{}", + if over.is_empty() { "PASS" } else { "FAIL" }, + summaries.len(), + nonces, + over.len(), + if over.is_empty() { "" } else { ": " }, + over.join(" ") + ); + let mut f = std::fs::File::create(format!("{out_dir}/seed-census-d{day}-n{nonces}-p{}-{}.txt", programs[0], programs[programs.len() - 1])).unwrap(); + writeln!(f, "# label id attempt X_0.1% X_0.5% X_1% hot chi2w_buckets64 z_w max_hi16_bucket_share max_sat_share acc_max_addr_of_2048 acc_max_sat_of_2048 ratio_0.1%_window ratio_0.1%_flat").unwrap(); + let mut n_hot = 0usize; + let mut n_acc_addr = [0usize; 3]; + let mut n_acc_sat = [0usize; 3]; + let mut n_sixsig = 0usize; + let mut n_ratio = 0usize; + for s in &summaries { + n_ratio += (s.ratio_w > 1.2) as usize; + writeln!( + f, + "{} {:016x} {} {:.5} {:.5} {:.5} {} {:.4} {:.2} {:.6} {:.6} {} {} {:.4} {:.4}", + s.label, s.id, s.attempt, s.x[0] * 100.0, s.x[1] * 100.0, s.x[2] * 100.0, s.hot as u8, s.chi2_w, s.z_w, s.max_bucket_share, s.max_sat_share, s.acc_max_addr, s.acc_max_sat, s.ratio_w, s.ratio_flat + ) + .unwrap(); + n_hot += s.hot as usize; + n_sixsig += (s.z_w > 6.0) as usize; + for (i, thr) in [3u32, 11, 21].into_iter().enumerate() { + n_acc_addr[i] += (s.acc_max_addr >= thr) as usize; + n_acc_sat[i] += (s.acc_max_sat >= thr) as usize; + } + } + log!( + "SEED CENSUS: {} programs at {} nonces each: hot set (X_f >= f, windowed control) in {}; top 0.1% beyond 1.2x of the window-model control in {}; windowed buckets64 beyond 6 sigma in {}; acceptance-style most-repeated address at one position >= 3 / 11 / 21 of 2048 in {} / {} / {}; saturated sources at one position >= 3 / 11 / 21 of 2048 in {} / {} / {}", + summaries.len(), + nonces, + n_hot, + n_ratio, + n_sixsig, + n_acc_addr[0], + n_acc_addr[1], + n_acc_addr[2], + n_acc_sat[0], + n_acc_sat[1], + n_acc_sat[2] + ); + let mut by_x: Vec<&ProgramSummary> = summaries.iter().collect(); + by_x.sort_by(|a, b| b.x[0].partial_cmp(&a.x[0]).unwrap()); + for s in by_x.iter().take(10) { + log!( + "SEED CENSUS top: {} id {:016x} attempt {} ratio_0.1% window {:.3}x flat {:.3}x X_0.1% {:+.4}% X_1% {:+.4}% hot {} z_w {:+.1} max bucket share {:.4}% max sat share {:.4}% acc addr {} acc sat {}", + s.label, + s.id, + s.attempt, + s.ratio_w, + s.ratio_flat, + s.x[0] * 100.0, + s.x[2] * 100.0, + s.hot as u8, + s.z_w, + s.max_bucket_share * 100.0, + s.max_sat_share * 100.0, + s.acc_max_addr, + s.acc_max_sat + ); + } + } +} +fn usage() -> ! { + eprintln!( + "attack-f8 lines --day0 20730 --days 16 --items-log2 24 --threads 12 --out [--plant none|quarter-lines|half-lines] [--validate all|sample|none]\n\ + attack-f8 warps --program 1|2|3 | --programs a..b --day 20730 --nonces 1000000 --threads 12 --out [--plant none|const-item] [--validate all|sample|none] [--check-every 997] [--diag 1]\n\ + attack-f8 census --seeds 64 --nonces 16777216 --control window --by-site --threads 12 --out the re-gate: one line per seed, PASS/FAIL against 1.2x\n\ + attack-f8 static --programs 1..67 per program, every load site's write chain back to the last injecting write (rule (a'))\n\ + attack-f8 seeds print the three programs' seeds" + ); + std::process::exit(2); +} + +fn main() { + let args: Vec = std::env::args().collect(); + if args.len() < 2 { + usage(); + } + let get = |k: &str, d: &str| -> String { + let mut i = 2; + while i + 1 < args.len() { + if args[i] == k { + return args[i + 1].clone(); + } + i += 1; + } + d.to_string() + }; + let threads: usize = get("--threads", "12").parse().unwrap(); + let out = get("--out", "."); + let plant = Plant::parse(&get("--plant", "none")); + let validate = get("--validate", "all"); + log!("attack-f8 {} (igneum-pow {}); args {:?}", env!("CARGO_PKG_VERSION"), igneum_pow::generator::GENERATOR_VERSION_V4, &args[1..]); + match args[1].as_str() { + "lines" => census_lines( + get("--day0", &DEFAULT_DAY.to_string()).parse().unwrap(), + get("--days", "16").parse().unwrap(), + get("--items-log2", "24").parse().unwrap(), + threads, + plant, + &validate, + &out, + ), + "warps" => census_warps( + &{ + let ps = get("--programs", ""); + if ps.is_empty() { + vec![get("--program", "1").parse::().unwrap()] + } else { + let (a, b) = ps.split_once("..").expect("--programs a..b"); + (a.parse::().unwrap()..=b.parse::().unwrap()).collect::>() + } + }, + get("--day", &DEFAULT_DAY.to_string()).parse().unwrap(), + get("--nonces", "1000000").parse().unwrap(), + threads, + plant, + &validate, + get("--check-every", "997").parse().unwrap(), + get("--diag", "0") == "1", + &out, + ), + "census" => { + // the re-gate entry: `census --seeds 64 --nonces 16777216 --control window --by-site` runs programs 2..=65 + // (the chain-shaped tag seeds; `--programs a..b` overrides) with the window-model control (the only control + // the hot-set test uses; the flat control is printed beside it) and the by-site attribution + let ps = get("--programs", ""); + let programs: Vec = if ps.is_empty() { + let n: u32 = get("--seeds", "64").parse().unwrap(); + (2..=n + 1).collect() + } else { + let (a, b) = ps.split_once("..").expect("--programs a..b"); + (a.parse::().unwrap()..=b.parse::().unwrap()).collect() + }; + let by_site = args.iter().any(|a| a == "--by-site") || get("--diag", "0") == "1"; + census_warps( + &programs, + get("--day", &DEFAULT_DAY.to_string()).parse().unwrap(), + get("--nonces", "16777216").parse().unwrap(), + threads, + plant, + &validate, + get("--check-every", "997").parse().unwrap(), + by_site, + &out, + ); + } + "static" => { + // per program: for every load site, the ops that write its source register between the register's last + // injecting write (add, sub, xor, mad, shfl, load) and the load, walking back cyclically; the proposed + // rule (a') rejects a program with an `or` or a `mul` in any such chain + let ps = get("--programs", "1..3"); + let (a, b) = ps.split_once("..").expect("--programs a..b"); + println!("# label id attempt sites_with_or sites_with_mul sites_with_mulhi sites_with_rot_only sites_clean rule_a_prime chains"); + for k in a.parse::().unwrap()..=b.parse::().unwrap() { + let spec = program_spec(k); + let program = Epoch::chain_program(&spec.epoch_seed, Some(&spec.era), ProgramClass::V4, &spec.label); + let instrs = &program.instrs; + let (mut n_or, mut n_mul, mut n_mulhi, mut n_rot, mut n_clean) = (0, 0, 0, 0, 0); + let mut chains = Vec::new(); + for (k, ins) in instrs.iter().enumerate() { + if ins.op != Op::Load { + continue; + } + let mut ops = Vec::new(); + for back in 1..instrs.len() { + let j = (k + instrs.len() - back) % instrs.len(); + if instrs[j].dst == ins.src { + if instrs[j].op.injects() { + ops.push(format!("{}@{j}", instrs[j].op.name())); + break; + } + ops.push(format!("{}@{j}", instrs[j].op.name())); + } + } + let has = |o: Op| ops.iter().any(|x| x.starts_with(&format!("{}@", o.name()))); + if has(Op::Or) { + n_or += 1; + } else if has(Op::Mul) { + n_mul += 1; + } else if has(Op::MulHi) { + n_mulhi += 1; + } else if has(Op::Rotl) || has(Op::Rotr) { + n_rot += 1; + } else { + n_clean += 1; + } + // the items the saturated sources map to under the era map at the 2^28-word dataset + let mask = (1u32 << DATASET_LOG2) - 1; + let layout = program.class.layout(); + let img = |x: u32| layout.split(load_index(program.class.era.as_ref(), ins, x, mask, DATASET_LOG2)).0; + chains.push(format!("{k}:r{}:{}:ones={:#08x}:zero={:#08x}", ins.src, ops.join("<"), img(u32::MAX), img(0))); + } + println!( + "{} {:016x} {} {n_or} {n_mul} {n_mulhi} {n_rot} {n_clean} {} {}", + spec.label, + program.program_id(), + program.attempt, + if n_or + n_mul > 0 { "REJECT" } else { "accept" }, + chains.join(" ") + ); + } + } + "seeds" => { + for k in 1..=3 { + let s = program_spec(k); + println!("{} epoch_seed {} era {}", s.label, hex(&s.epoch_seed), hex(&s.era)); + } + } + _ => usage(), + } +}