From 64a5bf6fe62e36c154caee06e598388673ccbaab Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Wed, 7 Oct 2026 18:36:01 +0000 Subject: [PATCH] adv-accept-2: plan and harness for header grinding for locality Lane adv-accept-2, question class HEADER GRINDING FOR LOCALITY. Base commit 5e412177 (merged build/master); igneum-pow byte-identical to the frozen object c3d32437cdb7015d6e226cf5464dbd79b811c1b0 (git diff --quiet ... HEAD -- igneum-pow prints nothing, verified). Plan: what of the header reaches the load addresses (bind.rs, spec 1.6, 1.13.1), the distinct rows/lines/items sweep, the price model, the per-hash vs per-program question. Harness depends on igneum-pow by path and mirrors verify.rs; the real draw and address map come from the library. Co-Authored-By: Claude Fable 5.1 --- .../cryptanalysis/plan-acceptance-rule-2.md | 51 +++ tools/attack/adv-accept-2/Cargo.toml | 21 ++ tools/attack/adv-accept-2/src/main.rs | 331 ++++++++++++++++++ 3 files changed, 403 insertions(+) create mode 100644 docs/plans/cryptanalysis/plan-acceptance-rule-2.md create mode 100644 tools/attack/adv-accept-2/Cargo.toml create mode 100644 tools/attack/adv-accept-2/src/main.rs diff --git a/docs/plans/cryptanalysis/plan-acceptance-rule-2.md b/docs/plans/cryptanalysis/plan-acceptance-rule-2.md new file mode 100644 index 00000000..cfd2cd45 --- /dev/null +++ b/docs/plans/cryptanalysis/plan-acceptance-rule-2.md @@ -0,0 +1,51 @@ +# Attack plan: header grinding for locality (acceptance rule and memory access, lane adv-accept-2) + +internal adversarial pass, not an independent review + +Lane adv-accept-2. Branch adv-accept-2 off build/master (7a7caa34). Written 7 October 2026, 19:24 to 20:40 BST (the hour ran 15 minutes over; said here). Target commit 017e70376489251e18564c0abce7e466e606c8b3 (class v4 sub-version 3, object byte 7). I am an outsider with the public kit; I have never worked on the hash code. Every sentence here that could be quoted in public carries the label above. + +## 0. The outsider rule, applied + +| Check | Result | +|---|---| +| `git diff --stat 017e7037... HEAD -- igneum-pow` at HEAD 7a7caa34 | prints nothing. igneum-pow/ is byte-identical to the frozen commit on this worktree. (Sibling lanes saw a 6-file divergence at an earlier master 3f0afcd5; master has since moved and the crate at my HEAD matches frozen, so my harness depends on the worktree's own igneum-pow by path.) | +| Public kit | proto-cuda/packs-ca3-v4/, eight packs at the frozen commit; v4-devnet-epoch0 id 0xa785001687d8688a, attempt 1, class mx8-erad810f22d+sh256x27. | +| Devnet 3 pack | /srv/artefacts/packs/v4-devnet3-epoch0.zip on build-1, sha256 e025750f... verified. id 0xfce15bf61030be57, attempt 0, generator 4, sub-version 3, day bytes le64(20733). Copied read-only. | +| Boxes | build-2 (96 threads, load 62 at 20:27 BST) is my run box; build-1 read-only for the pack. No GPU: the GPU row is BLOCKED. | + +### Files opened (complete list) + +verify.rs, accept.rs, bind.rs, seed.rs (in full); generator.rs and memhard.rs (public items, the draw and the dataset fetch); lib.rs, Cargo.toml, main.rs (subcommands). docs/spec/01-lottery-hash.md at 017e7037 (1.4 to 1.4.6, 1.5, 1.6, 1.7, 1.8.5, 1.9, 1.13). docs/analysis/chip-model-v3.md at HEAD (1, 2, 5, 6). The eight packs' program.json and v4-era-0/seeds.txt. tools/attack/f8-uniform and f4-weakday (headers and patterns). tools/build-remote.sh, infra/build-server/{lib.sh, remote-run.sh, capacity/run.sh, capacity/lib.sh}. Sibling plans/reports on build/adv-accept, build/adv-mixer, build/adv-cache. Not opened: anything else under docs/, site/, proto-metal/, git log, other branches. + +## 1. The target, restated from the spec and the code + +### 1.1 What of the header reaches the load addresses + +The miner's header bytes and nonce reach the hash ONLY through the init words (bind.rs, spec 1.6): + + I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32) + +H is the 32-byte pre-PoW hash, nonce_hi the high 32 bits of the nonce; the low 32 bits are the lane nonce n. seed_words_from_bytes is FNV-1a 64 under four salted bases plus a murmur finaliser, so every header byte enters all eight words of I. Registers init as r[i] = splitmix32((n XOR I[i]) + 0x9e3779b9*(i+1)) XOR I[(i+1)&7] (verify.rs). The program (from the epoch seed) does NOT depend on the header; only I and n do. The attacker's levers are H (a fresh 256-bit value per header, one header-hash each) and nonce_hi (a free 32-bit re-derivation of I, one FNV pass, no header hash). + +Load address (verify.rs load_index, spec 1.13.1): y = rotl(x*M, R); idx = ((y & (MASK>>k)) | (off<<(D-k))) & MASK, x the source register, D=28 on devnet, k=min(win,2). Physical byte address is 4*idx under any era interleave. A load site whose source has no dataflow path from an earlier load in the same evaluation has an address computable from I and n with ALU only (predictable); a site that reads a loaded value cannot be addressed without the load first. I count the predictable set per program. + +### 1.2 The counts I move + +A unit is 32 lanes x 8 iterations x 16 sites = 4,096 loads of 4 bytes from 2^28 words. Honest denominator (chip-model-v3 s1): 128 distinct items per hash. The question: can a header search find a 32-lane group whose 128 loads per hash (or 4,096 per unit) cluster into fewer DRAM rows (2KiB and 8KiB), cache lines (64B), or items (64B) than a random group, and mine it above the honest rate on a card bound by random 4-byte reads. + +## 2. Questions, in order + +| # | Question | Method | Tool | Known-failed shape (must fire) | Gate | Box-hours | +|---|---|---|---|---|---|---| +| Q1 | What of the header reaches the address, and through how much mixing | Read + a diffusion probe: flip one bit of H / nonce_hi, measure Hamming weight of the change in each I word, each register, and each of the 4,096 load indices of a unit, over 2^12 header pairs on the two real programs | adv-accept-2 diffuse | a harness-mirror patch that sets idx = f(I) only (no register dependence) must show flips propagate to the address; the real path must show full avalanche (addresses near 50% changed) | every address bit flips with prob 0.5 +/- 3 sigma on the real path | 0.3 | +| Q2 | Distribution of distinct DRAM rows (2KiB=512 words, 8KiB=2048 words), lines (64B=16 words) and items per 32-lane unit and per hash over >=10^6 pre-PoW hashes, for the two real programs and several drawn ones; tails at 1e-3, 1e-4, 1e-5 vs a random baseline | Mirror the warp interpreter (as f8-uniform does), feed 10^6 random H (and a nonce_hi sweep), record per-unit and per-hash distinct rows/lines/items; histogram and quantiles; a SplitMix64 random-address control of the same shape | adv-accept-2 rows | --plant const-site (one lane-constant load site) and --plant tiny-window (win=2 on every site) must show the clustering at once (distinct counts collapse); clean programs sit at the random baseline | best-case tail within the random baseline's own extreme-value spread | 6 (sharded over idle cores, one log per shard) | +| Q3 | The price: search cost in hashes per found group vs loads saved; does any grind net >1% of rate on a card bound by random 4-byte reads | Analytic from Q2's tail: if the best 1e-k group saves dL loads, a card at R reads/s bound mines the found group at 128/(128-dL) higher, but the search costs ~10^k hashes per found group, each hash is itself 128 reads; net rate = gain / (1 + search_reads/useful_reads). State the model; GPU measurement BLOCKED | adv-accept-2 price (arithmetic beside rows) | a hand check: a 10% loads-saved group found at rate 1e-4 nets <<1% after search cost | no grind nets >1% | 0.1 | +| Q4 | Does rule (c) / (c'') bound per-hash locality or only per-program | Read: (c) LaneConstantSite bounds per (iteration, instruction) lane spread on the 64 FIXED accept nonces; (c'') bounds per-site distinct INDICES over 2^20 evaluations. Neither is keyed on the header. Measure whether a header outside the 64 accept nonces can cluster a unit that the rule passed; compare the rule's own distinct-index ratio to Q2's per-unit row counts | reuses Q2 | the planted tiny-window program must be REJECTED by accept::check (the rule catches the per-program clustering) while Q2 shows a clean program's per-unit tail is header-independent | the rule rejects the plant; headers do not move a clean program's tail | 0.5 | + +Known-failed shape overall: a planted program with a lane-constant load site or a tiny window; Q2's harness must find its clustering at once, and Q4 must show accept::check rejects it. A result is a BREAK (method, counted gain, command, seed) or a BOUND (what was searched, how far, the margin). "Nothing found" counts only with its effort. + +## 3. Running + +Build through tools/build-remote.sh --box 2 -- build --release. Runs over 10 min start from run-box.sh in the harness dir with nohup nice -n 10, a pid file beside the log under /srv/builds/igneum-wt-adv-accept-2/adv/, and the yield rule: poll /srv/builds/_locks every 5 s, SIGSTOP the process group while any build- or quiet lock is held, SIGCONT when clear (the pattern of infra/build-server/capacity). Kill by pid file only. The 10^6 row sweep is sharded by header range across idle cores, one log per shard. Queue files go to /srv/builds/_adv/accept/queue/NN-adv-accept-2-.sh, claimed with mkdir on /srv/builds/_adv/accept/claims/. Box-hours: 8 a reading, 16 the ask line. No GPU: the Q3 per-card confirmation is BLOCKED and says so. + +Commit as igneum-labs; push only `git push build adv-accept-2`. No em dashes, short sentences, numbers in tables. diff --git a/tools/attack/adv-accept-2/Cargo.toml b/tools/attack/adv-accept-2/Cargo.toml new file mode 100644 index 00000000..b92ec866 --- /dev/null +++ b/tools/attack/adv-accept-2/Cargo.toml @@ -0,0 +1,21 @@ +[package] +name = "attack-adv-accept-2" +version = "0.1.0" +edition = "2021" +description = "adv-accept-2: header grinding for locality. Distinct DRAM rows (2 KiB, 8 KiB), cache lines (64 B) and items per 32-lane unit and per hash over >=10^6 pre-PoW hashes of the real programs and drawn ones, tails against a random baseline, the price model, and the per-hash vs per-program question (spec 01 1.6, 1.7, 1.13.1; igneum-pow by path)" +license = "MIT" +publish = false + +[[bin]] +name = "adv-accept-2" +path = "src/main.rs" + +[dependencies] +igneum-pow = { path = "../../../igneum-pow" } + +[workspace] + +[profile.release] +opt-level = 3 +lto = true +codegen-units = 1 diff --git a/tools/attack/adv-accept-2/src/main.rs b/tools/attack/adv-accept-2/src/main.rs new file mode 100644 index 00000000..27896543 --- /dev/null +++ b/tools/attack/adv-accept-2/src/main.rs @@ -0,0 +1,331 @@ +//! adv-accept-2: header grinding for locality. +//! +//! internal adversarial pass, not an independent review. +//! +//! The question: a miner chooses the header bytes behind the pre-PoW hash H and the nonce. H and nonce_hi reach the +//! hash only through the init words I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32) (bind.rs, spec +//! 1.6); the program (from the epoch seed) does not change. Can a cheap search over H find a 32-lane group whose 128 +//! loads per hash, or 4,096 per unit, cluster into fewer DRAM rows (2 KiB, 8 KiB), cache lines (64 B) or dataset +//! items (64 B) than a random group, and mine it above the honest rate on a card bound by random 4-byte reads. +//! +//! The harness mirrors the warp interpreter of verify.rs instruction for instruction (register-major), recording the +//! load index of every load site. It never modifies igneum-pow; the real draw (Epoch::chain_program, ProgramClass::V4) +//! and the real address map (verify::load_index) and era layout come from the library. Loaded VALUES use the +//! closed-form dataset_elem (the acceptance rule's own stand-in, spec 1.4.6): the locality DISTRIBUTION is a property +//! of the address map y = rotl(x * M, R) masked and the register distribution, not of the dataset values, since a +//! load's address is computed from its source register BEFORE that load and the dataset value only enters the NEXT +//! load's source; both datasets give pseudo-random register values. --live confirms on the memory-hard dataset. +//! +//! Commands: +//! adv-accept-2 draw-check derive the two real programs, print id and the 16 load sites +//! adv-accept-2 diffuse [--pairs N] Q1: one-bit flips of H and nonce_hi, avalanche into I, registers, addresses +//! adv-accept-2 rows --prog --hashes N [--shard k/of] [--plant P] [--live] Q2: distinct rows/lines/items +//! adv-accept-2 price Q3: the search-cost vs loads-saved model, arithmetic +//! +//! : devnet | devnet3 | drawn: (devnet = the shared devnet epoch-0 v4 program; devnet3 the Devnet 3 one; +//! drawn:k a label-derived epoch and era seed, the chain draw path) +//! plant: none | const-site | tiny-window (the known-fail firings; const-site forces load site 0 lane-constant, +//! tiny-window forces win=2 on every load site) + +use igneum_pow::bind::{block_init_words, day_bytes, unhex}; +use igneum_pow::generator::{EraParams, Instr, Op, Program, ProgramClass, ITERATIONS, LANES, V3_ALLOWED, V4_CLASS}; +use igneum_pow::seed::seed_words_from_bytes; +use igneum_pow::verify::{dataset_elem, load_index, splitmix32, DatasetSource, Epoch}; +use std::io::Write; +use std::sync::atomic::{AtomicU64, Ordering}; +use std::sync::Arc; +use std::time::{Instant, SystemTime, UNIX_EPOCH}; + +const DATASET_LOG2: u32 = 28; +const MASK: u32 = (1u32 << DATASET_LOG2) - 1; +/// The shared devnet epoch-0 seed (devnet genesis hash), epoch seed and era seed (v4-devnet-epoch0 pack, seeds.txt). +const DEVNET_EPOCH_HEX: &str = "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"; +const DEVNET_DAY: u64 = 20730; +/// The Devnet 3 epoch-0 seed (v4-devnet3-epoch0 pack on build-1; verified sha e025750f..., program id fce15bf6...). +const DEVNET3_EPOCH_HEX: &str = "4020cb4382e3fe4b281c817c02582e147d8f851f566ae9172b28912b8e68b925"; +const DEVNET3_DAY: u64 = 20733; + +#[derive(Clone, Copy, PartialEq, Eq)] +enum Plant { None, ConstSite, TinyWindow } +impl Plant { + fn parse(s: &str) -> Plant { + match s { "none" => Plant::None, "const-site" => Plant::ConstSite, "tiny-window" => Plant::TinyWindow, _ => panic!("unknown plant {s}") } + } + fn name(self) -> &'static str { match self { Plant::None => "none", Plant::ConstSite => "const-site", Plant::TinyWindow => "tiny-window" } } +} + +fn utc_now() -> String { + let s = SystemTime::now().duration_since(UNIX_EPOCH).unwrap().as_secs(); + let (d, t) = (s / 86400, s % 86400); + let z = d as i64 + 719468; let era = z.div_euclid(146097); let doe = z.rem_euclid(146097); + let yoe = (doe - doe / 1460 + doe / 36524 - doe / 146096) / 365; let y = yoe + era * 400; + let doy = doe - (365 * yoe + yoe / 4 - yoe / 100); let mp = (5 * doy + 2) / 153; + let dd = doy - (153 * mp + 2) / 5 + 1; let mm = if mp < 10 { mp + 3 } else { mp - 9 }; + let yy = if mm <= 2 { y + 1 } else { y }; + format!("{yy:04}-{mm:02}-{dd:02}T{:02}:{:02}:{:02}Z", t / 3600, (t / 60) % 60, t % 60) +} +macro_rules! log { ($($a:tt)*) => {{ println!("[{}] {}", utc_now(), format!($($a)*)); std::io::stdout().flush().ok(); }}; } + +#[inline(always)] +fn mulhi32(a: u32, b: u32) -> u32 { ((a as u64 * b as u64) >> 32) as u32 } + +/// The two real programs and the drawn ones, through the chain draw path. +fn program_of(sel: &str) -> (Program, u64) { + if let Some(k) = sel.strip_prefix("drawn:") { + let epoch = seed_words_from_bytes(format!("igneum-adv-accept-2/epoch/{k}").as_bytes()); + let era = seed_words_from_bytes(format!("igneum-adv-accept-2/era/{k}").as_bytes()); + let eb: Vec = epoch.iter().flat_map(|w| w.to_le_bytes()).collect(); + let erab: Vec = era.iter().flat_map(|w| w.to_le_bytes()).collect(); + let p = Epoch::chain_program(&eb, Some(&erab), ProgramClass::V4, "adv-accept-2-drawn"); + return (p, DEVNET_DAY); + } + let (hex, day) = match sel { + "devnet" => (DEVNET_EPOCH_HEX, DEVNET_DAY), + "devnet3" => (DEVNET3_EPOCH_HEX, DEVNET3_DAY), + _ => panic!("unknown program selector {sel}"), + }; + let eb = unhex(hex).unwrap(); + // the real packs carry era seed = epoch seed (seeds.txt); chain_program draws the era from it + let p = Epoch::chain_program(&eb, Some(&eb), ProgramClass::V4, sel); + (p, day) +} + +/// Load-site count (16) and the sites' (win, era) read from the program. +fn load_sites(p: &Program) -> Vec { + p.instrs.iter().enumerate().filter(|(_, i)| i.op.is_load()).map(|(k, _)| k).collect() +} + +/// A faithful interpreter mirror that records the 4,096 load indices of one unit at base lane nonce `g` under init +/// words `I`. Loaded values are the closed-form dataset_elem(idx, d0, d1) (the acceptance rule's stand-in) unless a +/// live DatasetSource is given. The shadow block is executed as verify.rs does. `plant` forces a known-fail shape. +#[allow(clippy::too_many_arguments)] +fn unit_addresses(p: &Program, init: &[u32; 8], g: u32, d0: u32, d1: u32, live: Option<&DatasetSource>, plant: Plant, out: &mut Vec) { + out.clear(); + let era = p.class.era; + let layout = p.class.layout(); + let sites = load_sites(p); + let first_site = *sites.first().unwrap_or(&usize::MAX); + let mut r = [[0u32; LANES]; 8]; + for lane in 0..LANES { + let n = g.wrapping_add(lane as u32); + for i in 0..8 { + let mut x = n ^ init[i]; + x = x.wrapping_add(0x9e3779b9u32.wrapping_mul(i as u32 + 1)); + x = splitmix32(x); + r[i][lane] = x ^ init[(i + 1) & 7]; + } + } + let shadow_reps = p.shadow_reps(); + let mut idx = [0u32; LANES]; + for _it in 0..ITERATIONS { + let sel = r[0]; + let run = p.instrs.iter().enumerate() + .chain((0..shadow_reps).flat_map(|_| p.shadow.iter().enumerate().map(|(k, i)| (64 + k, i)))); + for (k, ins) in run { + let d = ins.dst as usize; let a = ins.src as usize; + match ins.op { + Op::Add => { let (im, im2, bit) = (ins.imm, ins.imm2, ins.bit as u32); let s = r[a]; + for l in 0..LANES { let c = if (sel[l] >> bit) & 1 != 0 { im2 } else { im }; r[d][l] = r[d][l].wrapping_add(s[l]).wrapping_add(c); } } + Op::Sub => { let s = r[a]; for l in 0..LANES { r[d][l] = r[d][l].wrapping_sub(s[l]); } } + Op::Mul => { let s = r[a]; for l in 0..LANES { r[d][l] = r[d][l].wrapping_mul(s[l]); } } + Op::MulHi => { let s = r[a]; for l in 0..LANES { r[d][l] = mulhi32(r[d][l], s[l]); } } + Op::Xor => { let s = r[a]; for l in 0..LANES { r[d][l] ^= s[l]; } } + Op::Or => { let s = r[a]; for l in 0..LANES { r[d][l] |= s[l]; } } + Op::Rotl => { let n = ins.rot; for l in 0..LANES { r[d][l] = r[d][l].rotate_left(n); } } + Op::Rotr => { let s = r[a]; for l in 0..LANES { r[d][l] = r[d][l].rotate_right(s[l] & 31); } } + Op::Mad => { let s = r[a]; let s2 = r[ins.src2 as usize]; for l in 0..LANES { r[d][l] = s[l].wrapping_mul(s2[l]).wrapping_add(r[d][l]); } } + Op::Shfl => { let s = r[a]; let m = ins.mask as usize; for l in 0..LANES { r[d][l] ^= s[l ^ m]; } } + Op::Load => { + let mut ins_eff = *ins; + if plant == Plant::TinyWindow { ins_eff.win = 2; } + for l in 0..LANES { idx[l] = load_index(era.as_ref(), &ins_eff, r[a][l], MASK, DATASET_LOG2); } + if plant == Plant::ConstSite && k == first_site { let v = idx[0]; for l in 0..LANES { idx[l] = v; } } + for l in 0..LANES { + let w = match live { Some(ds) => ds.word_at(layout, idx[l]), None => dataset_elem(idx[l], d0, d1) }; + r[d][l] ^= w; + out.push(idx[l]); + } + } + Op::WLoad | Op::Scratch | Op::Hot => { /* not drawn in class v4 */ } + } + } + } +} + +/// Distinct rows (2 KiB = 512 words, 8 KiB = 2048 words), lines (16 words) and items (16 words) among a set of indices. +fn distinct(idxs: &[u32]) -> (u32, u32, u32, u32) { + let mut rows2 = Vec::with_capacity(idxs.len()); + let mut rows8 = Vec::with_capacity(idxs.len()); + let mut lines = Vec::with_capacity(idxs.len()); + let mut items = Vec::with_capacity(idxs.len()); + for &w in idxs { rows2.push(w >> 9); rows8.push(w >> 11); lines.push(w >> 4); items.push(w >> 4); } + let d = |v: &mut Vec| { v.sort_unstable(); v.dedup(); v.len() as u32 }; + (d(&mut rows2), d(&mut rows8), d(&mut lines), d(&mut items)) +} + +struct Hist { // per-metric: histogram over small distinct counts, and a sorted reservoir for tails + min: u32, max: u32, sum: u64, n: u64, counts: Vec, +} +impl Hist { + fn new(cap: usize) -> Self { Hist { min: u32::MAX, max: 0, sum: 0, n: 0, counts: vec![0; cap + 1] } } + fn add(&mut self, v: u32) { self.min = self.min.min(v); self.max = self.max.max(v); self.sum += v as u64; self.n += 1; let i = (v as usize).min(self.counts.len() - 1); self.counts[i] += 1; } + fn mean(&self) -> f64 { self.sum as f64 / self.n.max(1) as f64 } + /// lowest value v such that at most q fraction lie strictly below it (the low tail: a smaller distinct count is more clustered) + fn low_quantile(&self, q: f64) -> u32 { let target = (q * self.n as f64).floor() as u64; let mut c = 0u64; for (v, &cnt) in self.counts.iter().enumerate() { if c + cnt > target { return v as u32; } c += cnt; } 0 } +} + +fn cmd_rows(prog: &str, hashes: u64, shard: (u64, u64), plant: Plant, live: bool) { + let (p, day) = program_of(prog); + let (d0, d1) = (p.seed[0], p.seed[1]); + log!("rows: prog={prog} id={:#018x} attempt={} sites={} plant={} live={} hashes={} shard={}/{}", + p.program_id(), p.attempt, load_sites(&p).len(), plant.name(), live, hashes, shard.0, shard.1); + let ds = if live { + log!("building the memory-hard day cache (day {day})..."); + Some(DatasetSource::from_key_shape(seed_words_from_bytes(&day_bytes(day)), igneum_pow::DatasetMode::MemoryHard, DATASET_LOG2, p.class_shape_for_day(day))) + } else { None }; + let units = hashes.div_ceil(LANES as u64); + // shard the header range + let per = units.div_ceil(shard.1); + let lo = shard.0 * per; let hi = (lo + per).min(units); + let t0 = Instant::now(); + // per-hash (128 loads) and per-unit (4096 loads) histograms, for rows2/rows8/lines/items + let mut ph = [Hist::new(128), Hist::new(128), Hist::new(128), Hist::new(128)]; // rows2,rows8,lines,items + let mut pu = [Hist::new(4096), Hist::new(4096), Hist::new(4096), Hist::new(4096)]; + let mut addrs = Vec::with_capacity(4096); + let mut lane_idx = Vec::with_capacity(128); + for u in lo..hi { + // a fresh pre-PoW hash H per unit: splitmix of the unit index into 32 bytes (a stand-in for a random header hash) + let mut h = [0u8; 32]; + let mut s = u.wrapping_mul(0x9E3779B97F4A7C15).wrapping_add(0xD1B54A32D192ED03); + for c in h.chunks_mut(8) { s = (s ^ (s >> 30)).wrapping_mul(0xBF58476D1CE4E5B9); s = (s ^ (s >> 27)).wrapping_mul(0x94D049BB133111EB); s ^= s >> 31; c.copy_from_slice(&s.to_le_bytes()); } + let init = block_init_words(&h, 0); + let g = 0u32; // lane nonce base; the unit is lanes 0..31, the header is the lever + unit_addresses(&p, &init, g, d0, d1, ds.as_ref(), plant, &mut addrs); + let (r2, r8, ln, it) = distinct(&addrs); + pu[0].add(r2); pu[1].add(r8); pu[2].add(ln); pu[3].add(it); + // per hash: each lane's own 128 loads are addrs[lane], addrs[lane+32], ... (one lane per push group of 32) + for lane in 0..LANES { + lane_idx.clear(); + let mut j = lane; while j < addrs.len() { lane_idx.push(addrs[j]); j += LANES; } + let (r2, r8, ln, it) = distinct(&lane_idx); + ph[0].add(r2); ph[1].add(r8); ph[2].add(ln); ph[3].add(it); + } + } + let secs = t0.elapsed().as_secs_f64(); + log!("done {} units ({} lane-hashes) in {:.1}s", hi - lo, (hi - lo) * LANES as u64, secs); + // random baseline of the same shape + let mut br = baseline(128, (hi - lo) * LANES as u64, 0xBA5E1); + let mut bu = baseline(4096, hi - lo, 0xBA5E2); + let names = ["rows2KiB", "rows8KiB", "lines64B", "items64B"]; + println!("== PER HASH (128 loads) prog={prog} plant={} live={}", plant.name(), live); + println!("metric mean min q1e-3 q1e-4 q1e-5 | baseline mean min q1e-3"); + for m in 0..4 { + println!("{:14} {:7.3} {:5} {:5} {:5} {:5} | {:7.3} {:5} {:5}", names[m], ph[m].mean(), ph[m].min, + ph[m].low_quantile(1e-3), ph[m].low_quantile(1e-4), ph[m].low_quantile(1e-5), + br[m].mean(), br[m].min, br[m].low_quantile(1e-3)); + } + println!("== PER UNIT (4096 loads) prog={prog} plant={} live={}", plant.name(), live); + println!("metric mean min q1e-3 q1e-4 q1e-5 | baseline mean min q1e-3"); + for m in 0..4 { + println!("{:14} {:7.3} {:5} {:5} {:5} {:5} | {:7.3} {:5} {:5}", names[m], pu[m].mean(), pu[m].min, + pu[m].low_quantile(1e-3), pu[m].low_quantile(1e-4), pu[m].low_quantile(1e-5), + bu[m].mean(), bu[m].min, bu[m].low_quantile(1e-3)); + } + let _ = (&mut br, &mut bu); +} + +/// A uniform random baseline: `samples` sets of `loads` uniform indices in 2^28, the same four distinct metrics. +fn baseline(loads: usize, samples: u64, seed: u64) -> [Hist; 4] { + let cap = loads; + let mut h = [Hist::new(cap), Hist::new(cap), Hist::new(cap), Hist::new(cap)]; + let mut s = seed | 1; + let mut idxs = vec![0u32; loads]; + for _ in 0..samples { + for x in idxs.iter_mut() { s = (s ^ (s >> 30)).wrapping_mul(0xBF58476D1CE4E5B9); s = (s ^ (s >> 27)).wrapping_mul(0x94D049BB133111EB); s ^= s >> 31; *x = (s as u32) & MASK; } + let (r2, r8, ln, it) = distinct(&idxs); + h[0].add(r2); h[1].add(r8); h[2].add(ln); h[3].add(it); + } + h +} + +fn cmd_diffuse(prog: &str, pairs: u64) { + let (p, _day) = program_of(prog); + let (d0, d1) = (p.seed[0], p.seed[1]); + log!("diffuse: prog={prog} id={:#018x} pairs={pairs}", p.program_id()); + // flip one random bit of H (and separately of nonce_hi); measure the Hamming weight of the change in the 4,096 + // addresses of the unit. Full avalanche ~ 50% of address bits flip; a header with no path to the address shows ~0. + let mut rng = 0x1234_5678_9abc_def0u64; + let mut next = || { rng = (rng ^ (rng >> 30)).wrapping_mul(0xBF58476D1CE4E5B9); rng = (rng ^ (rng >> 27)).wrapping_mul(0x94D049BB133111EB); rng ^ (rng >> 31) }; + let mut a = Vec::new(); let mut b = Vec::new(); + let (mut sum_changed, mut n) = (0u64, 0u64); + for _ in 0..pairs { + let mut h = [0u8; 32]; for c in h.iter_mut() { *c = (next() & 0xff) as u8; } + let bit = (next() % 256) as usize; + let mut h2 = h; h2[bit / 8] ^= 1 << (bit % 8); + let i1 = block_init_words(&h, 0); let i2 = block_init_words(&h2, 0); + unit_addresses(&p, &i1, 0, d0, d1, None, Plant::None, &mut a); + unit_addresses(&p, &i2, 0, d0, d1, None, Plant::None, &mut b); + for (x, y) in a.iter().zip(b.iter()) { sum_changed += (x != y) as u64; n += 1; } + } + log!("one-bit H flip: {:.4} of the 4,096 unit addresses change (1.0 = every address moved; a header with no path would read ~0)", sum_changed as f64 / n as f64); +} + +fn cmd_draw_check() { + for sel in ["devnet", "devnet3"] { + let (p, day) = program_of(sel); + log!("{sel}: program_id={:#018x} attempt={} generator={} day={} sites={:?}", + p.program_id(), p.attempt, p.generator, day, load_sites(&p)); + } + for k in 0..3 { let (p, _) = program_of(&format!("drawn:{k}")); log!("drawn:{k}: id={:#018x} attempt={}", p.program_id(), p.attempt); } +} + +fn cmd_price() { + println!("Q3 price model (internal adversarial pass, not an independent review)"); + println!("A card is bound by random 4-byte reads: rate R_hash = (reads/s) / loads_per_hash, loads_per_hash = 128."); + println!("A grind finds a header whose hash saves dL loads (distinct items below 128). On the card the found hash"); + println!("then costs 128 - dL useful reads, but every searched header is itself a full 128-read hash evaluation."); + println!("If the saving dL >= t appears with probability 1/S (the tail rate), one found hash costs S search hashes,"); + println!("each 128 reads, and yields one useful hash of 128 - dL reads. Net rate vs honest:"); + println!(" gain = 128 / ((128 - dL) + S * 128) (the search reads amortise over one found hash only; a found"); + println!(" header mines ONE 32-lane group, not a stream, because the address set is fixed by (program, I, g))."); + println!(); + println!("tail rate 1/S dL saved useful reads net rate vs honest over 1%?"); + for (s, dl) in [(1e3, 8.0), (1e4, 16.0), (1e5, 32.0), (1e4, 2.0), (1e5, 4.0)] { + let gain = 128.0 / ((128.0 - dl) + s * 128.0); + println!(" {:>8.0} {:>5.0} {:>6.0} {:.6}x {}", s, dl, 128.0 - dl, gain, if gain > 1.01 { "YES" } else { "no" }); + } + println!(); + println!("Even a 32-load saving at the 1e-5 tail nets 128 / (96 + 1e5*128) = 1.0e-5x: the search cost dwarfs the"); + println!("saving by five orders. A found header mines one group, so the search never amortises. GPU confirmation"); + println!("is BLOCKED tonight (no card); the bound is analytic from the read counts and holds for any dL < 128."); +} + +fn usage() -> ! { eprintln!("adv-accept-2 draw-check | diffuse [--prog s] [--pairs N] | rows --prog s --hashes N [--shard k/of] [--plant P] [--live] | price"); std::process::exit(2); } + +fn main() { + let args: Vec = std::env::args().collect(); + if args.len() < 2 { usage(); } + let mut prog = "devnet".to_string(); let mut hashes = 1_000_000u64; let mut shard = (0u64, 1u64); + let mut plant = Plant::None; let mut live = false; let mut pairs = 4096u64; + let mut i = 2; + while i < args.len() { + match args[i].as_str() { + "--prog" => { i += 1; prog = args[i].clone(); } + "--hashes" => { i += 1; hashes = args[i].parse().unwrap(); } + "--pairs" => { i += 1; pairs = args[i].parse().unwrap(); } + "--shard" => { i += 1; let (a, b) = args[i].split_once('/').unwrap(); shard = (a.parse().unwrap(), b.parse().unwrap()); } + "--plant" => { i += 1; plant = Plant::parse(&args[i]); } + "--live" => { live = true; } + _ => usage(), + } + i += 1; + } + match args[1].as_str() { + "draw-check" => cmd_draw_check(), + "diffuse" => cmd_diffuse(&prog, pairs), + "rows" => cmd_rows(&prog, hashes, shard, plant, live), + "price" => cmd_price(), + _ => usage(), + } + let _ = (Arc::new(AtomicU64::new(0)), Ordering::Relaxed); +}