adv-accept-2: plan and harness for header grinding for locality
Lane adv-accept-2, question class HEADER GRINDING FOR LOCALITY. Base commit5e412177(merged build/master); igneum-pow byte-identical to the frozen objectc3d32437cd(git diff --quiet ... HEAD -- igneum-pow prints nothing, verified). Plan: what of the header reaches the load addresses (bind.rs, spec 1.6, 1.13.1), the distinct rows/lines/items sweep, the price model, the per-hash vs per-program question. Harness depends on igneum-pow by path and mirrors verify.rs; the real draw and address map come from the library. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
5e41217735
commit
64a5bf6fe6
3 changed files with 403 additions and 0 deletions
51
docs/plans/cryptanalysis/plan-acceptance-rule-2.md
Normal file
51
docs/plans/cryptanalysis/plan-acceptance-rule-2.md
Normal file
|
|
@ -0,0 +1,51 @@
|
|||
# Attack plan: header grinding for locality (acceptance rule and memory access, lane adv-accept-2)
|
||||
|
||||
internal adversarial pass, not an independent review
|
||||
|
||||
Lane adv-accept-2. Branch adv-accept-2 off build/master (7a7caa34). Written 7 October 2026, 19:24 to 20:40 BST (the hour ran 15 minutes over; said here). Target commit 017e70376489251e18564c0abce7e466e606c8b3 (class v4 sub-version 3, object byte 7). I am an outsider with the public kit; I have never worked on the hash code. Every sentence here that could be quoted in public carries the label above.
|
||||
|
||||
## 0. The outsider rule, applied
|
||||
|
||||
| Check | Result |
|
||||
|---|---|
|
||||
| `git diff --stat 017e7037... HEAD -- igneum-pow` at HEAD 7a7caa34 | prints nothing. igneum-pow/ is byte-identical to the frozen commit on this worktree. (Sibling lanes saw a 6-file divergence at an earlier master 3f0afcd5; master has since moved and the crate at my HEAD matches frozen, so my harness depends on the worktree's own igneum-pow by path.) |
|
||||
| Public kit | proto-cuda/packs-ca3-v4/, eight packs at the frozen commit; v4-devnet-epoch0 id 0xa785001687d8688a, attempt 1, class mx8-erad810f22d+sh256x27. |
|
||||
| Devnet 3 pack | /srv/artefacts/packs/v4-devnet3-epoch0.zip on build-1, sha256 e025750f... verified. id 0xfce15bf61030be57, attempt 0, generator 4, sub-version 3, day bytes le64(20733). Copied read-only. |
|
||||
| Boxes | build-2 (96 threads, load 62 at 20:27 BST) is my run box; build-1 read-only for the pack. No GPU: the GPU row is BLOCKED. |
|
||||
|
||||
### Files opened (complete list)
|
||||
|
||||
verify.rs, accept.rs, bind.rs, seed.rs (in full); generator.rs and memhard.rs (public items, the draw and the dataset fetch); lib.rs, Cargo.toml, main.rs (subcommands). docs/spec/01-lottery-hash.md at 017e7037 (1.4 to 1.4.6, 1.5, 1.6, 1.7, 1.8.5, 1.9, 1.13). docs/analysis/chip-model-v3.md at HEAD (1, 2, 5, 6). The eight packs' program.json and v4-era-0/seeds.txt. tools/attack/f8-uniform and f4-weakday (headers and patterns). tools/build-remote.sh, infra/build-server/{lib.sh, remote-run.sh, capacity/run.sh, capacity/lib.sh}. Sibling plans/reports on build/adv-accept, build/adv-mixer, build/adv-cache. Not opened: anything else under docs/, site/, proto-metal/, git log, other branches.
|
||||
|
||||
## 1. The target, restated from the spec and the code
|
||||
|
||||
### 1.1 What of the header reaches the load addresses
|
||||
|
||||
The miner's header bytes and nonce reach the hash ONLY through the init words (bind.rs, spec 1.6):
|
||||
|
||||
I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32)
|
||||
|
||||
H is the 32-byte pre-PoW hash, nonce_hi the high 32 bits of the nonce; the low 32 bits are the lane nonce n. seed_words_from_bytes is FNV-1a 64 under four salted bases plus a murmur finaliser, so every header byte enters all eight words of I. Registers init as r[i] = splitmix32((n XOR I[i]) + 0x9e3779b9*(i+1)) XOR I[(i+1)&7] (verify.rs). The program (from the epoch seed) does NOT depend on the header; only I and n do. The attacker's levers are H (a fresh 256-bit value per header, one header-hash each) and nonce_hi (a free 32-bit re-derivation of I, one FNV pass, no header hash).
|
||||
|
||||
Load address (verify.rs load_index, spec 1.13.1): y = rotl(x*M, R); idx = ((y & (MASK>>k)) | (off<<(D-k))) & MASK, x the source register, D=28 on devnet, k=min(win,2). Physical byte address is 4*idx under any era interleave. A load site whose source has no dataflow path from an earlier load in the same evaluation has an address computable from I and n with ALU only (predictable); a site that reads a loaded value cannot be addressed without the load first. I count the predictable set per program.
|
||||
|
||||
### 1.2 The counts I move
|
||||
|
||||
A unit is 32 lanes x 8 iterations x 16 sites = 4,096 loads of 4 bytes from 2^28 words. Honest denominator (chip-model-v3 s1): 128 distinct items per hash. The question: can a header search find a 32-lane group whose 128 loads per hash (or 4,096 per unit) cluster into fewer DRAM rows (2KiB and 8KiB), cache lines (64B), or items (64B) than a random group, and mine it above the honest rate on a card bound by random 4-byte reads.
|
||||
|
||||
## 2. Questions, in order
|
||||
|
||||
| # | Question | Method | Tool | Known-failed shape (must fire) | Gate | Box-hours |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q1 | What of the header reaches the address, and through how much mixing | Read + a diffusion probe: flip one bit of H / nonce_hi, measure Hamming weight of the change in each I word, each register, and each of the 4,096 load indices of a unit, over 2^12 header pairs on the two real programs | adv-accept-2 diffuse | a harness-mirror patch that sets idx = f(I) only (no register dependence) must show flips propagate to the address; the real path must show full avalanche (addresses near 50% changed) | every address bit flips with prob 0.5 +/- 3 sigma on the real path | 0.3 |
|
||||
| Q2 | Distribution of distinct DRAM rows (2KiB=512 words, 8KiB=2048 words), lines (64B=16 words) and items per 32-lane unit and per hash over >=10^6 pre-PoW hashes, for the two real programs and several drawn ones; tails at 1e-3, 1e-4, 1e-5 vs a random baseline | Mirror the warp interpreter (as f8-uniform does), feed 10^6 random H (and a nonce_hi sweep), record per-unit and per-hash distinct rows/lines/items; histogram and quantiles; a SplitMix64 random-address control of the same shape | adv-accept-2 rows | --plant const-site (one lane-constant load site) and --plant tiny-window (win=2 on every site) must show the clustering at once (distinct counts collapse); clean programs sit at the random baseline | best-case tail within the random baseline's own extreme-value spread | 6 (sharded over idle cores, one log per shard) |
|
||||
| Q3 | The price: search cost in hashes per found group vs loads saved; does any grind net >1% of rate on a card bound by random 4-byte reads | Analytic from Q2's tail: if the best 1e-k group saves dL loads, a card at R reads/s bound mines the found group at 128/(128-dL) higher, but the search costs ~10^k hashes per found group, each hash is itself 128 reads; net rate = gain / (1 + search_reads/useful_reads). State the model; GPU measurement BLOCKED | adv-accept-2 price (arithmetic beside rows) | a hand check: a 10% loads-saved group found at rate 1e-4 nets <<1% after search cost | no grind nets >1% | 0.1 |
|
||||
| Q4 | Does rule (c) / (c'') bound per-hash locality or only per-program | Read: (c) LaneConstantSite bounds per (iteration, instruction) lane spread on the 64 FIXED accept nonces; (c'') bounds per-site distinct INDICES over 2^20 evaluations. Neither is keyed on the header. Measure whether a header outside the 64 accept nonces can cluster a unit that the rule passed; compare the rule's own distinct-index ratio to Q2's per-unit row counts | reuses Q2 | the planted tiny-window program must be REJECTED by accept::check (the rule catches the per-program clustering) while Q2 shows a clean program's per-unit tail is header-independent | the rule rejects the plant; headers do not move a clean program's tail | 0.5 |
|
||||
|
||||
Known-failed shape overall: a planted program with a lane-constant load site or a tiny window; Q2's harness must find its clustering at once, and Q4 must show accept::check rejects it. A result is a BREAK (method, counted gain, command, seed) or a BOUND (what was searched, how far, the margin). "Nothing found" counts only with its effort.
|
||||
|
||||
## 3. Running
|
||||
|
||||
Build through tools/build-remote.sh --box 2 -- build --release. Runs over 10 min start from run-box.sh in the harness dir with nohup nice -n 10, a pid file beside the log under /srv/builds/igneum-wt-adv-accept-2/adv/, and the yield rule: poll /srv/builds/_locks every 5 s, SIGSTOP the process group while any build-<k> or quiet lock is held, SIGCONT when clear (the pattern of infra/build-server/capacity). Kill by pid file only. The 10^6 row sweep is sharded by header range across idle cores, one log per shard. Queue files go to /srv/builds/_adv/accept/queue/NN-adv-accept-2-<name>.sh, claimed with mkdir on /srv/builds/_adv/accept/claims/<name>. Box-hours: 8 a reading, 16 the ask line. No GPU: the Q3 per-card confirmation is BLOCKED and says so.
|
||||
|
||||
Commit as igneum-labs; push only `git push build adv-accept-2`. No em dashes, short sentences, numbers in tables.
|
||||
21
tools/attack/adv-accept-2/Cargo.toml
Normal file
21
tools/attack/adv-accept-2/Cargo.toml
Normal file
|
|
@ -0,0 +1,21 @@
|
|||
[package]
|
||||
name = "attack-adv-accept-2"
|
||||
version = "0.1.0"
|
||||
edition = "2021"
|
||||
description = "adv-accept-2: header grinding for locality. Distinct DRAM rows (2 KiB, 8 KiB), cache lines (64 B) and items per 32-lane unit and per hash over >=10^6 pre-PoW hashes of the real programs and drawn ones, tails against a random baseline, the price model, and the per-hash vs per-program question (spec 01 1.6, 1.7, 1.13.1; igneum-pow by path)"
|
||||
license = "MIT"
|
||||
publish = false
|
||||
|
||||
[[bin]]
|
||||
name = "adv-accept-2"
|
||||
path = "src/main.rs"
|
||||
|
||||
[dependencies]
|
||||
igneum-pow = { path = "../../../igneum-pow" }
|
||||
|
||||
[workspace]
|
||||
|
||||
[profile.release]
|
||||
opt-level = 3
|
||||
lto = true
|
||||
codegen-units = 1
|
||||
331
tools/attack/adv-accept-2/src/main.rs
Normal file
331
tools/attack/adv-accept-2/src/main.rs
Normal file
|
|
@ -0,0 +1,331 @@
|
|||
//! adv-accept-2: header grinding for locality.
|
||||
//!
|
||||
//! internal adversarial pass, not an independent review.
|
||||
//!
|
||||
//! The question: a miner chooses the header bytes behind the pre-PoW hash H and the nonce. H and nonce_hi reach the
|
||||
//! hash only through the init words I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32) (bind.rs, spec
|
||||
//! 1.6); the program (from the epoch seed) does not change. Can a cheap search over H find a 32-lane group whose 128
|
||||
//! loads per hash, or 4,096 per unit, cluster into fewer DRAM rows (2 KiB, 8 KiB), cache lines (64 B) or dataset
|
||||
//! items (64 B) than a random group, and mine it above the honest rate on a card bound by random 4-byte reads.
|
||||
//!
|
||||
//! The harness mirrors the warp interpreter of verify.rs instruction for instruction (register-major), recording the
|
||||
//! load index of every load site. It never modifies igneum-pow; the real draw (Epoch::chain_program, ProgramClass::V4)
|
||||
//! and the real address map (verify::load_index) and era layout come from the library. Loaded VALUES use the
|
||||
//! closed-form dataset_elem (the acceptance rule's own stand-in, spec 1.4.6): the locality DISTRIBUTION is a property
|
||||
//! of the address map y = rotl(x * M, R) masked and the register distribution, not of the dataset values, since a
|
||||
//! load's address is computed from its source register BEFORE that load and the dataset value only enters the NEXT
|
||||
//! load's source; both datasets give pseudo-random register values. --live confirms on the memory-hard dataset.
|
||||
//!
|
||||
//! Commands:
|
||||
//! adv-accept-2 draw-check derive the two real programs, print id and the 16 load sites
|
||||
//! adv-accept-2 diffuse [--pairs N] Q1: one-bit flips of H and nonce_hi, avalanche into I, registers, addresses
|
||||
//! adv-accept-2 rows --prog <sel> --hashes N [--shard k/of] [--plant P] [--live] Q2: distinct rows/lines/items
|
||||
//! adv-accept-2 price Q3: the search-cost vs loads-saved model, arithmetic
|
||||
//!
|
||||
//! <sel>: devnet | devnet3 | drawn:<k> (devnet = the shared devnet epoch-0 v4 program; devnet3 the Devnet 3 one;
|
||||
//! drawn:k a label-derived epoch and era seed, the chain draw path)
|
||||
//! plant: none | const-site | tiny-window (the known-fail firings; const-site forces load site 0 lane-constant,
|
||||
//! tiny-window forces win=2 on every load site)
|
||||
|
||||
use igneum_pow::bind::{block_init_words, day_bytes, unhex};
|
||||
use igneum_pow::generator::{EraParams, Instr, Op, Program, ProgramClass, ITERATIONS, LANES, V3_ALLOWED, V4_CLASS};
|
||||
use igneum_pow::seed::seed_words_from_bytes;
|
||||
use igneum_pow::verify::{dataset_elem, load_index, splitmix32, DatasetSource, Epoch};
|
||||
use std::io::Write;
|
||||
use std::sync::atomic::{AtomicU64, Ordering};
|
||||
use std::sync::Arc;
|
||||
use std::time::{Instant, SystemTime, UNIX_EPOCH};
|
||||
|
||||
const DATASET_LOG2: u32 = 28;
|
||||
const MASK: u32 = (1u32 << DATASET_LOG2) - 1;
|
||||
/// The shared devnet epoch-0 seed (devnet genesis hash), epoch seed and era seed (v4-devnet-epoch0 pack, seeds.txt).
|
||||
const DEVNET_EPOCH_HEX: &str = "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07";
|
||||
const DEVNET_DAY: u64 = 20730;
|
||||
/// The Devnet 3 epoch-0 seed (v4-devnet3-epoch0 pack on build-1; verified sha e025750f..., program id fce15bf6...).
|
||||
const DEVNET3_EPOCH_HEX: &str = "4020cb4382e3fe4b281c817c02582e147d8f851f566ae9172b28912b8e68b925";
|
||||
const DEVNET3_DAY: u64 = 20733;
|
||||
|
||||
#[derive(Clone, Copy, PartialEq, Eq)]
|
||||
enum Plant { None, ConstSite, TinyWindow }
|
||||
impl Plant {
|
||||
fn parse(s: &str) -> Plant {
|
||||
match s { "none" => Plant::None, "const-site" => Plant::ConstSite, "tiny-window" => Plant::TinyWindow, _ => panic!("unknown plant {s}") }
|
||||
}
|
||||
fn name(self) -> &'static str { match self { Plant::None => "none", Plant::ConstSite => "const-site", Plant::TinyWindow => "tiny-window" } }
|
||||
}
|
||||
|
||||
fn utc_now() -> String {
|
||||
let s = SystemTime::now().duration_since(UNIX_EPOCH).unwrap().as_secs();
|
||||
let (d, t) = (s / 86400, s % 86400);
|
||||
let z = d as i64 + 719468; let era = z.div_euclid(146097); let doe = z.rem_euclid(146097);
|
||||
let yoe = (doe - doe / 1460 + doe / 36524 - doe / 146096) / 365; let y = yoe + era * 400;
|
||||
let doy = doe - (365 * yoe + yoe / 4 - yoe / 100); let mp = (5 * doy + 2) / 153;
|
||||
let dd = doy - (153 * mp + 2) / 5 + 1; let mm = if mp < 10 { mp + 3 } else { mp - 9 };
|
||||
let yy = if mm <= 2 { y + 1 } else { y };
|
||||
format!("{yy:04}-{mm:02}-{dd:02}T{:02}:{:02}:{:02}Z", t / 3600, (t / 60) % 60, t % 60)
|
||||
}
|
||||
macro_rules! log { ($($a:tt)*) => {{ println!("[{}] {}", utc_now(), format!($($a)*)); std::io::stdout().flush().ok(); }}; }
|
||||
|
||||
#[inline(always)]
|
||||
fn mulhi32(a: u32, b: u32) -> u32 { ((a as u64 * b as u64) >> 32) as u32 }
|
||||
|
||||
/// The two real programs and the drawn ones, through the chain draw path.
|
||||
fn program_of(sel: &str) -> (Program, u64) {
|
||||
if let Some(k) = sel.strip_prefix("drawn:") {
|
||||
let epoch = seed_words_from_bytes(format!("igneum-adv-accept-2/epoch/{k}").as_bytes());
|
||||
let era = seed_words_from_bytes(format!("igneum-adv-accept-2/era/{k}").as_bytes());
|
||||
let eb: Vec<u8> = epoch.iter().flat_map(|w| w.to_le_bytes()).collect();
|
||||
let erab: Vec<u8> = era.iter().flat_map(|w| w.to_le_bytes()).collect();
|
||||
let p = Epoch::chain_program(&eb, Some(&erab), ProgramClass::V4, "adv-accept-2-drawn");
|
||||
return (p, DEVNET_DAY);
|
||||
}
|
||||
let (hex, day) = match sel {
|
||||
"devnet" => (DEVNET_EPOCH_HEX, DEVNET_DAY),
|
||||
"devnet3" => (DEVNET3_EPOCH_HEX, DEVNET3_DAY),
|
||||
_ => panic!("unknown program selector {sel}"),
|
||||
};
|
||||
let eb = unhex(hex).unwrap();
|
||||
// the real packs carry era seed = epoch seed (seeds.txt); chain_program draws the era from it
|
||||
let p = Epoch::chain_program(&eb, Some(&eb), ProgramClass::V4, sel);
|
||||
(p, day)
|
||||
}
|
||||
|
||||
/// Load-site count (16) and the sites' (win, era) read from the program.
|
||||
fn load_sites(p: &Program) -> Vec<usize> {
|
||||
p.instrs.iter().enumerate().filter(|(_, i)| i.op.is_load()).map(|(k, _)| k).collect()
|
||||
}
|
||||
|
||||
/// A faithful interpreter mirror that records the 4,096 load indices of one unit at base lane nonce `g` under init
|
||||
/// words `I`. Loaded values are the closed-form dataset_elem(idx, d0, d1) (the acceptance rule's stand-in) unless a
|
||||
/// live DatasetSource is given. The shadow block is executed as verify.rs does. `plant` forces a known-fail shape.
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
fn unit_addresses(p: &Program, init: &[u32; 8], g: u32, d0: u32, d1: u32, live: Option<&DatasetSource>, plant: Plant, out: &mut Vec<u32>) {
|
||||
out.clear();
|
||||
let era = p.class.era;
|
||||
let layout = p.class.layout();
|
||||
let sites = load_sites(p);
|
||||
let first_site = *sites.first().unwrap_or(&usize::MAX);
|
||||
let mut r = [[0u32; LANES]; 8];
|
||||
for lane in 0..LANES {
|
||||
let n = g.wrapping_add(lane as u32);
|
||||
for i in 0..8 {
|
||||
let mut x = n ^ init[i];
|
||||
x = x.wrapping_add(0x9e3779b9u32.wrapping_mul(i as u32 + 1));
|
||||
x = splitmix32(x);
|
||||
r[i][lane] = x ^ init[(i + 1) & 7];
|
||||
}
|
||||
}
|
||||
let shadow_reps = p.shadow_reps();
|
||||
let mut idx = [0u32; LANES];
|
||||
for _it in 0..ITERATIONS {
|
||||
let sel = r[0];
|
||||
let run = p.instrs.iter().enumerate()
|
||||
.chain((0..shadow_reps).flat_map(|_| p.shadow.iter().enumerate().map(|(k, i)| (64 + k, i))));
|
||||
for (k, ins) in run {
|
||||
let d = ins.dst as usize; let a = ins.src as usize;
|
||||
match ins.op {
|
||||
Op::Add => { let (im, im2, bit) = (ins.imm, ins.imm2, ins.bit as u32); let s = r[a];
|
||||
for l in 0..LANES { let c = if (sel[l] >> bit) & 1 != 0 { im2 } else { im }; r[d][l] = r[d][l].wrapping_add(s[l]).wrapping_add(c); } }
|
||||
Op::Sub => { let s = r[a]; for l in 0..LANES { r[d][l] = r[d][l].wrapping_sub(s[l]); } }
|
||||
Op::Mul => { let s = r[a]; for l in 0..LANES { r[d][l] = r[d][l].wrapping_mul(s[l]); } }
|
||||
Op::MulHi => { let s = r[a]; for l in 0..LANES { r[d][l] = mulhi32(r[d][l], s[l]); } }
|
||||
Op::Xor => { let s = r[a]; for l in 0..LANES { r[d][l] ^= s[l]; } }
|
||||
Op::Or => { let s = r[a]; for l in 0..LANES { r[d][l] |= s[l]; } }
|
||||
Op::Rotl => { let n = ins.rot; for l in 0..LANES { r[d][l] = r[d][l].rotate_left(n); } }
|
||||
Op::Rotr => { let s = r[a]; for l in 0..LANES { r[d][l] = r[d][l].rotate_right(s[l] & 31); } }
|
||||
Op::Mad => { let s = r[a]; let s2 = r[ins.src2 as usize]; for l in 0..LANES { r[d][l] = s[l].wrapping_mul(s2[l]).wrapping_add(r[d][l]); } }
|
||||
Op::Shfl => { let s = r[a]; let m = ins.mask as usize; for l in 0..LANES { r[d][l] ^= s[l ^ m]; } }
|
||||
Op::Load => {
|
||||
let mut ins_eff = *ins;
|
||||
if plant == Plant::TinyWindow { ins_eff.win = 2; }
|
||||
for l in 0..LANES { idx[l] = load_index(era.as_ref(), &ins_eff, r[a][l], MASK, DATASET_LOG2); }
|
||||
if plant == Plant::ConstSite && k == first_site { let v = idx[0]; for l in 0..LANES { idx[l] = v; } }
|
||||
for l in 0..LANES {
|
||||
let w = match live { Some(ds) => ds.word_at(layout, idx[l]), None => dataset_elem(idx[l], d0, d1) };
|
||||
r[d][l] ^= w;
|
||||
out.push(idx[l]);
|
||||
}
|
||||
}
|
||||
Op::WLoad | Op::Scratch | Op::Hot => { /* not drawn in class v4 */ }
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Distinct rows (2 KiB = 512 words, 8 KiB = 2048 words), lines (16 words) and items (16 words) among a set of indices.
|
||||
fn distinct(idxs: &[u32]) -> (u32, u32, u32, u32) {
|
||||
let mut rows2 = Vec::with_capacity(idxs.len());
|
||||
let mut rows8 = Vec::with_capacity(idxs.len());
|
||||
let mut lines = Vec::with_capacity(idxs.len());
|
||||
let mut items = Vec::with_capacity(idxs.len());
|
||||
for &w in idxs { rows2.push(w >> 9); rows8.push(w >> 11); lines.push(w >> 4); items.push(w >> 4); }
|
||||
let d = |v: &mut Vec<u32>| { v.sort_unstable(); v.dedup(); v.len() as u32 };
|
||||
(d(&mut rows2), d(&mut rows8), d(&mut lines), d(&mut items))
|
||||
}
|
||||
|
||||
struct Hist { // per-metric: histogram over small distinct counts, and a sorted reservoir for tails
|
||||
min: u32, max: u32, sum: u64, n: u64, counts: Vec<u64>,
|
||||
}
|
||||
impl Hist {
|
||||
fn new(cap: usize) -> Self { Hist { min: u32::MAX, max: 0, sum: 0, n: 0, counts: vec![0; cap + 1] } }
|
||||
fn add(&mut self, v: u32) { self.min = self.min.min(v); self.max = self.max.max(v); self.sum += v as u64; self.n += 1; let i = (v as usize).min(self.counts.len() - 1); self.counts[i] += 1; }
|
||||
fn mean(&self) -> f64 { self.sum as f64 / self.n.max(1) as f64 }
|
||||
/// lowest value v such that at most q fraction lie strictly below it (the low tail: a smaller distinct count is more clustered)
|
||||
fn low_quantile(&self, q: f64) -> u32 { let target = (q * self.n as f64).floor() as u64; let mut c = 0u64; for (v, &cnt) in self.counts.iter().enumerate() { if c + cnt > target { return v as u32; } c += cnt; } 0 }
|
||||
}
|
||||
|
||||
fn cmd_rows(prog: &str, hashes: u64, shard: (u64, u64), plant: Plant, live: bool) {
|
||||
let (p, day) = program_of(prog);
|
||||
let (d0, d1) = (p.seed[0], p.seed[1]);
|
||||
log!("rows: prog={prog} id={:#018x} attempt={} sites={} plant={} live={} hashes={} shard={}/{}",
|
||||
p.program_id(), p.attempt, load_sites(&p).len(), plant.name(), live, hashes, shard.0, shard.1);
|
||||
let ds = if live {
|
||||
log!("building the memory-hard day cache (day {day})...");
|
||||
Some(DatasetSource::from_key_shape(seed_words_from_bytes(&day_bytes(day)), igneum_pow::DatasetMode::MemoryHard, DATASET_LOG2, p.class_shape_for_day(day)))
|
||||
} else { None };
|
||||
let units = hashes.div_ceil(LANES as u64);
|
||||
// shard the header range
|
||||
let per = units.div_ceil(shard.1);
|
||||
let lo = shard.0 * per; let hi = (lo + per).min(units);
|
||||
let t0 = Instant::now();
|
||||
// per-hash (128 loads) and per-unit (4096 loads) histograms, for rows2/rows8/lines/items
|
||||
let mut ph = [Hist::new(128), Hist::new(128), Hist::new(128), Hist::new(128)]; // rows2,rows8,lines,items
|
||||
let mut pu = [Hist::new(4096), Hist::new(4096), Hist::new(4096), Hist::new(4096)];
|
||||
let mut addrs = Vec::with_capacity(4096);
|
||||
let mut lane_idx = Vec::with_capacity(128);
|
||||
for u in lo..hi {
|
||||
// a fresh pre-PoW hash H per unit: splitmix of the unit index into 32 bytes (a stand-in for a random header hash)
|
||||
let mut h = [0u8; 32];
|
||||
let mut s = u.wrapping_mul(0x9E3779B97F4A7C15).wrapping_add(0xD1B54A32D192ED03);
|
||||
for c in h.chunks_mut(8) { s = (s ^ (s >> 30)).wrapping_mul(0xBF58476D1CE4E5B9); s = (s ^ (s >> 27)).wrapping_mul(0x94D049BB133111EB); s ^= s >> 31; c.copy_from_slice(&s.to_le_bytes()); }
|
||||
let init = block_init_words(&h, 0);
|
||||
let g = 0u32; // lane nonce base; the unit is lanes 0..31, the header is the lever
|
||||
unit_addresses(&p, &init, g, d0, d1, ds.as_ref(), plant, &mut addrs);
|
||||
let (r2, r8, ln, it) = distinct(&addrs);
|
||||
pu[0].add(r2); pu[1].add(r8); pu[2].add(ln); pu[3].add(it);
|
||||
// per hash: each lane's own 128 loads are addrs[lane], addrs[lane+32], ... (one lane per push group of 32)
|
||||
for lane in 0..LANES {
|
||||
lane_idx.clear();
|
||||
let mut j = lane; while j < addrs.len() { lane_idx.push(addrs[j]); j += LANES; }
|
||||
let (r2, r8, ln, it) = distinct(&lane_idx);
|
||||
ph[0].add(r2); ph[1].add(r8); ph[2].add(ln); ph[3].add(it);
|
||||
}
|
||||
}
|
||||
let secs = t0.elapsed().as_secs_f64();
|
||||
log!("done {} units ({} lane-hashes) in {:.1}s", hi - lo, (hi - lo) * LANES as u64, secs);
|
||||
// random baseline of the same shape
|
||||
let mut br = baseline(128, (hi - lo) * LANES as u64, 0xBA5E1);
|
||||
let mut bu = baseline(4096, hi - lo, 0xBA5E2);
|
||||
let names = ["rows2KiB", "rows8KiB", "lines64B", "items64B"];
|
||||
println!("== PER HASH (128 loads) prog={prog} plant={} live={}", plant.name(), live);
|
||||
println!("metric mean min q1e-3 q1e-4 q1e-5 | baseline mean min q1e-3");
|
||||
for m in 0..4 {
|
||||
println!("{:14} {:7.3} {:5} {:5} {:5} {:5} | {:7.3} {:5} {:5}", names[m], ph[m].mean(), ph[m].min,
|
||||
ph[m].low_quantile(1e-3), ph[m].low_quantile(1e-4), ph[m].low_quantile(1e-5),
|
||||
br[m].mean(), br[m].min, br[m].low_quantile(1e-3));
|
||||
}
|
||||
println!("== PER UNIT (4096 loads) prog={prog} plant={} live={}", plant.name(), live);
|
||||
println!("metric mean min q1e-3 q1e-4 q1e-5 | baseline mean min q1e-3");
|
||||
for m in 0..4 {
|
||||
println!("{:14} {:7.3} {:5} {:5} {:5} {:5} | {:7.3} {:5} {:5}", names[m], pu[m].mean(), pu[m].min,
|
||||
pu[m].low_quantile(1e-3), pu[m].low_quantile(1e-4), pu[m].low_quantile(1e-5),
|
||||
bu[m].mean(), bu[m].min, bu[m].low_quantile(1e-3));
|
||||
}
|
||||
let _ = (&mut br, &mut bu);
|
||||
}
|
||||
|
||||
/// A uniform random baseline: `samples` sets of `loads` uniform indices in 2^28, the same four distinct metrics.
|
||||
fn baseline(loads: usize, samples: u64, seed: u64) -> [Hist; 4] {
|
||||
let cap = loads;
|
||||
let mut h = [Hist::new(cap), Hist::new(cap), Hist::new(cap), Hist::new(cap)];
|
||||
let mut s = seed | 1;
|
||||
let mut idxs = vec![0u32; loads];
|
||||
for _ in 0..samples {
|
||||
for x in idxs.iter_mut() { s = (s ^ (s >> 30)).wrapping_mul(0xBF58476D1CE4E5B9); s = (s ^ (s >> 27)).wrapping_mul(0x94D049BB133111EB); s ^= s >> 31; *x = (s as u32) & MASK; }
|
||||
let (r2, r8, ln, it) = distinct(&idxs);
|
||||
h[0].add(r2); h[1].add(r8); h[2].add(ln); h[3].add(it);
|
||||
}
|
||||
h
|
||||
}
|
||||
|
||||
fn cmd_diffuse(prog: &str, pairs: u64) {
|
||||
let (p, _day) = program_of(prog);
|
||||
let (d0, d1) = (p.seed[0], p.seed[1]);
|
||||
log!("diffuse: prog={prog} id={:#018x} pairs={pairs}", p.program_id());
|
||||
// flip one random bit of H (and separately of nonce_hi); measure the Hamming weight of the change in the 4,096
|
||||
// addresses of the unit. Full avalanche ~ 50% of address bits flip; a header with no path to the address shows ~0.
|
||||
let mut rng = 0x1234_5678_9abc_def0u64;
|
||||
let mut next = || { rng = (rng ^ (rng >> 30)).wrapping_mul(0xBF58476D1CE4E5B9); rng = (rng ^ (rng >> 27)).wrapping_mul(0x94D049BB133111EB); rng ^ (rng >> 31) };
|
||||
let mut a = Vec::new(); let mut b = Vec::new();
|
||||
let (mut sum_changed, mut n) = (0u64, 0u64);
|
||||
for _ in 0..pairs {
|
||||
let mut h = [0u8; 32]; for c in h.iter_mut() { *c = (next() & 0xff) as u8; }
|
||||
let bit = (next() % 256) as usize;
|
||||
let mut h2 = h; h2[bit / 8] ^= 1 << (bit % 8);
|
||||
let i1 = block_init_words(&h, 0); let i2 = block_init_words(&h2, 0);
|
||||
unit_addresses(&p, &i1, 0, d0, d1, None, Plant::None, &mut a);
|
||||
unit_addresses(&p, &i2, 0, d0, d1, None, Plant::None, &mut b);
|
||||
for (x, y) in a.iter().zip(b.iter()) { sum_changed += (x != y) as u64; n += 1; }
|
||||
}
|
||||
log!("one-bit H flip: {:.4} of the 4,096 unit addresses change (1.0 = every address moved; a header with no path would read ~0)", sum_changed as f64 / n as f64);
|
||||
}
|
||||
|
||||
fn cmd_draw_check() {
|
||||
for sel in ["devnet", "devnet3"] {
|
||||
let (p, day) = program_of(sel);
|
||||
log!("{sel}: program_id={:#018x} attempt={} generator={} day={} sites={:?}",
|
||||
p.program_id(), p.attempt, p.generator, day, load_sites(&p));
|
||||
}
|
||||
for k in 0..3 { let (p, _) = program_of(&format!("drawn:{k}")); log!("drawn:{k}: id={:#018x} attempt={}", p.program_id(), p.attempt); }
|
||||
}
|
||||
|
||||
fn cmd_price() {
|
||||
println!("Q3 price model (internal adversarial pass, not an independent review)");
|
||||
println!("A card is bound by random 4-byte reads: rate R_hash = (reads/s) / loads_per_hash, loads_per_hash = 128.");
|
||||
println!("A grind finds a header whose hash saves dL loads (distinct items below 128). On the card the found hash");
|
||||
println!("then costs 128 - dL useful reads, but every searched header is itself a full 128-read hash evaluation.");
|
||||
println!("If the saving dL >= t appears with probability 1/S (the tail rate), one found hash costs S search hashes,");
|
||||
println!("each 128 reads, and yields one useful hash of 128 - dL reads. Net rate vs honest:");
|
||||
println!(" gain = 128 / ((128 - dL) + S * 128) (the search reads amortise over one found hash only; a found");
|
||||
println!(" header mines ONE 32-lane group, not a stream, because the address set is fixed by (program, I, g)).");
|
||||
println!();
|
||||
println!("tail rate 1/S dL saved useful reads net rate vs honest over 1%?");
|
||||
for (s, dl) in [(1e3, 8.0), (1e4, 16.0), (1e5, 32.0), (1e4, 2.0), (1e5, 4.0)] {
|
||||
let gain = 128.0 / ((128.0 - dl) + s * 128.0);
|
||||
println!(" {:>8.0} {:>5.0} {:>6.0} {:.6}x {}", s, dl, 128.0 - dl, gain, if gain > 1.01 { "YES" } else { "no" });
|
||||
}
|
||||
println!();
|
||||
println!("Even a 32-load saving at the 1e-5 tail nets 128 / (96 + 1e5*128) = 1.0e-5x: the search cost dwarfs the");
|
||||
println!("saving by five orders. A found header mines one group, so the search never amortises. GPU confirmation");
|
||||
println!("is BLOCKED tonight (no card); the bound is analytic from the read counts and holds for any dL < 128.");
|
||||
}
|
||||
|
||||
fn usage() -> ! { eprintln!("adv-accept-2 draw-check | diffuse [--prog s] [--pairs N] | rows --prog s --hashes N [--shard k/of] [--plant P] [--live] | price"); std::process::exit(2); }
|
||||
|
||||
fn main() {
|
||||
let args: Vec<String> = std::env::args().collect();
|
||||
if args.len() < 2 { usage(); }
|
||||
let mut prog = "devnet".to_string(); let mut hashes = 1_000_000u64; let mut shard = (0u64, 1u64);
|
||||
let mut plant = Plant::None; let mut live = false; let mut pairs = 4096u64;
|
||||
let mut i = 2;
|
||||
while i < args.len() {
|
||||
match args[i].as_str() {
|
||||
"--prog" => { i += 1; prog = args[i].clone(); }
|
||||
"--hashes" => { i += 1; hashes = args[i].parse().unwrap(); }
|
||||
"--pairs" => { i += 1; pairs = args[i].parse().unwrap(); }
|
||||
"--shard" => { i += 1; let (a, b) = args[i].split_once('/').unwrap(); shard = (a.parse().unwrap(), b.parse().unwrap()); }
|
||||
"--plant" => { i += 1; plant = Plant::parse(&args[i]); }
|
||||
"--live" => { live = true; }
|
||||
_ => usage(),
|
||||
}
|
||||
i += 1;
|
||||
}
|
||||
match args[1].as_str() {
|
||||
"draw-check" => cmd_draw_check(),
|
||||
"diffuse" => cmd_diffuse(&prog, pairs),
|
||||
"rows" => cmd_rows(&prog, hashes, shard, plant, live),
|
||||
"price" => cmd_price(),
|
||||
_ => usage(),
|
||||
}
|
||||
let _ = (Arc::new(AtomicU64::new(0)), Ordering::Relaxed);
|
||||
}
|
||||
Loading…
Reference in a new issue