Counter ASIC 3.0 gates (hash): AP-F8-1 under the windows-union model. The window layer moves the null from 0.115 to 0.160 percent at f = 0.1 percent (1.39x, not 4.05x); the 153x item is the all-ones value of an or-written load source (site 15, or at 61, load at 63), which the era map sends to item 0xca5b92 exactly and whose next seven items are the seven surviving one-zero-bit sources; the popcount model at the measured bias predicts 77,348 reads against 78,479 and the program's S_0.1 percent at 0.58 against 0.52; a static census of 1,024 chain-shaped v4 programs (tools/ca3-v4-uniform, built and run on igneum-build-1): 96.6 percent carry a lossy-sourced load, 48.5 percent an or-sourced one, 4.9 percent an or chain (p3's class); F8's 1.2x gate fails 96.6 percent; the chip ceiling under rule (c) is 1.067x; the two flip options priced; the v5 bound defined

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 09:26:48 +00:00
parent 210084b0e5
commit 4a4bcac0b9
6 changed files with 1259 additions and 0 deletions

View file

@ -0,0 +1,62 @@
# AP-F8-1 under the windows-union model: the hot set is the load source, not the window (7 October 2026)
Branch `ca3-v4-uniform` from master b92a5fd4, worker "v4-hash", on the attack-pass finding AP-F8-1 (`docs/analysis/attack-pass/f8-uniform.md`, branch attack-pass; logs `/srv/builds/igneum-wt-attack/target-attack-f8/log/`). Main's rulings bound this file: no generator change to class v4 on the live devnet; the analysis and its harness only. Every GPU-free number here is arithmetic on F8's logged counts or a run of the static census tool `tools/ca3-v4-uniform/` on igneum-build-1 (built through `tools/build-remote.sh`, rule R1); the chip figures are the terms of `docs/analysis/chip-model-v3.md` and are approximate.
## 1. The null F8's numbers must be read against
Layer 8 (`docs/plans/era-layout.md` 1.4, spec 01 1.13.1 as proposed) gives every load site a window draw `k_off = below(3)`: the site reads the whole dataset, an aligned half or an aligned quarter, at a 2^26-word floor. A quarter-window site concentrates its reads 4x on its quarter and a half-window site 2x on its half, by design; the union of the 16 windows is the whole dataset. For p1 (the devnet epoch-0 program, id c120d7963abdcd96) the 16 draws `0:2:1 1:1:1 2:1:1 3:1:1 4:0:0 5:1:1 6:0:0 7:2:2 8:1:1 9:1:1 10:2:0 11:0:0 12:0:0 13:0:0 14:2:0 15:1:1` (site:k:offset) give an expected read density by quarter of 3.25 : 2.25 : 5.75 : 4.75 sixteenths of the flat mean, which at 2^26 nonces (512 reads per item flat) is 416, 288, 736 and 608 reads per item. The top-f share of a Poisson mixture with those means (`uniform-model.txt`, exact Poisson for p1, a normal approximation for the census):
| Share of all reads on the top f of items, 2^26 nonces | Flat Poisson (F8's control) | Windows-union model, p1 | F8 measured, p1 | Beyond the window model |
|---|---|---|---|---|
| f = 0.1 percent | 0.115 | 0.160 | 0.520 | +0.36 |
| f = 0.5 percent | 0.565 | 0.784 | 1.515 | +0.73 |
| f = 1 percent | 1.120 | 1.553 | 2.495 | +0.94 |
So the window model moves the null from 0.115 to 0.160 percent at f = 0.1 percent (1.39x, not F8's 4.05x) and from 1.12 to 1.55 at f = 1 percent; it explains the 64-item-bucket sigma of p2 that F8 already attributed to the window layer, and it explains every per-site attribution row of p1 except one: sites with a half window over the hot region land 0.20 percent of their reads in the top 0.1 percent (sites 1, 2, 3, 5, 9 at 0.201), quarter-window sites 0 or about 0.4 (sites 0, 10, 14 at 0.000, site 7 at 0.241 straddling), whole-dataset sites 0.10. Site 15 lands 6.374 percent. The excess over the model (+0.36 at f = 0.1 percent) is one site.
## 2. The 153x item is the load source, and the model predicts it to the item
p1's site 15 is the load at instruction 63 (`load dst=3 src=6`, window half 1). Its source r6 was last written at instruction 61: `or dst=6 src=4` (`r6 |= r4`, `verify.rs` Op::Or), after a fresh dataset load into r6 at 47. An OR of two near-uniform registers sets each bit with probability 3/4, so the source takes the all-ones value with probability (3/4)^32 = 1.0e-4 per read and the values of popcount 31, 30, ... with 32, 496, ... times (3/4)^k (1/4)^(32 - k). The era map `y = rotl(x * 0x9ad30d99, 29)`, the half window and the interleave split (`memhard::Layout::split`, positions 0, 2, 12, 13) send x = 0xffffffff to item 0xca5b92: F8's hottest item exactly. F8's next seven items (0x8a5b92, 0xaa5b92, 0xba5b92, 0x825b92, 0xe65b92, 0x985b92, 0xbcc392) are exactly the seven one-zero-bit sources whose zero bit survives the window mask (bits 29, 28, 27, 26, 25, 24 and 15): 7 of 7. The measured count fixes the bit bias: 78,479 reads of 2^26 x 8 site-15 reads is p^32 at p = 0.7585 (r4 is slightly biased itself), and at that p the popcount model predicts 77,348 all-ones reads and 4.86 percent of site 15's reads into the top 0.1 percent of items (measured 6.37; 3.92 at p = 3/4). Per hash that is 4.86 / 16 = 0.30 percent of all reads, and 0.160 + 0.30 = 0.46 against F8's 0.520 at f = 0.1 percent; at f = 1 percent 14.5 / 16 = 0.90, and 1.55 + 0.90 = 2.46 against 2.495.
The same arithmetic for the other lossy writers (`uniform-model.txt`): a `mul` last writer zeroes the low bits by the operands' trailing zeros, so 1.07 percent of the site's reads land on the 0.1 percent of values with 10 or more trailing zeros (p2's `mul`-sourced sites 0 and 14 measured 0.971 and 0.966 percent); a `mulhi` last writer is dense near zero, 0.79 percent on the lowest 0.1 percent of values. An `or` whose operand was itself last written by `or` compounds the bias (3/4 to 7/8 to 15/16): p3's site 15 (`or` at 30, the load at 62) puts 72.4 percent of its reads into the top 0.1 percent, 4.6 percent of all reads on 16,777 items.
This is a fault class, not the window model: the acceptance rule's part (a) (`accept.rs` check_stale_loads) takes any write as a fresh source, and part (c)'s saturation count looks at the 16,384 final register values, not at a load's source mid-program, so an `or`, `mul` or `mulhi` as a load's last writer passes. The per-hash distinct-address check still holds (p1 127.999 items per hash; p3 127.97: a saturated site repeats its item inside a hash), and the acceptance rule's floor of 120 distinct of 128 admits exactly one site repeating its item in all 8 iterations and no more.
## 3. How common it is: the static census (`tools/ca3-v4-uniform`, 1,024 chain-shaped class v4 programs plus F8's p1 to p3)
For every load site, the op that last wrote its source in execution order (base instructions before it, else the shadow block of the previous iteration, else the base instructions after it): injecting (add, sub, xor, mad, shfl, load), bijective (rotl, rotr) or lossy (or, mul, mulhi). Run on igneum-build-1 (`uniform-census.txt`, binary sha256 ce9f83fe... then the narrowed chain rule).
| Census over 1,024 programs | Count | Share |
|---|---|---|
| Load sites by last writer: injecting / bijective / lossy | 11,368 / 2,121 / 2,943 of 16,432 | 69 / 13 / 18 percent; 2.87 lossy sites per program |
| Programs with at least one lossy-sourced load | 992 | 96.6 percent |
| ... with an `or`-sourced load (p1's class, 0.30 percent of all reads per site) | 498 | 48.5 percent |
| ... with an `or`-of-`or` chain (p3's class, about 4.5 percent of all reads per site) | 50 | 4.9 percent |
| ... with a `mul`-sourced load (0.067 percent per site) / a `mulhi`-sourced load (0.049) | 751 / 661 | 73.1 / 64.4 percent |
| Predicted S_0.1 percent (window model plus the lossy sites): median / 90th / 99th / max | 0.45 / 0.88 / 5.29 / 9.82 percent | against the window model's 0.115 to 0.251 |
| p1 / p2 / p3 predicted against F8 measured | 0.579 / 0.323 / 4.72 | 0.520 / 0.272 / 4.60 |
F8's proposed gate (the top 0.1 percent within 1.2x of the window-model control on every one of 64 seeds) fails 96.6 percent of today's programs, because any lossy-sourced site alone exceeds it (0.16 + 0.05 at the least); it is a generator change in a gate's clothing. A 2x bound fails 69.7 percent, 3x 48.9 percent; a bound of S_0.1 percent at or under 1 percent of all reads fails 6.9 percent (the `or` chains and the multi-`or` programs). The static rule "no load whose source's last writer is `or`" fails 48.4 percent; "no lossy last writer" 96.6 percent.
## 4. What the skew is worth to a chip (chip-model-v3.md terms, approximate)
A hot-set cache of the top 0.1 percent of items is 16,777 items x 64 B = 1.07 MB of SRAM, 0.53 mm^2 and $0.25 at 0.49 mm^2 and $0.23 per MB. It serves 0.52 percent of p1's reads (0.16 of them the window model's), 4.6 percent of p3's. The hash is latency-bound on its dependent reads, so a read served on die is time saved: a chip gains at most 1.005x on p1 and 1.048x on p3 from the cache. The ceiling under the live rule: part (c)'s 120-of-128 floor admits one site repeating its item in all 8 iterations and no more (two saturated sites fail it), so at most 8 of 128 reads, 6.25 percent, can sit on a constant item, and a chip's edge from this whole class is at most 1 / (1 - 0.0625) = 1.067x, in 64 bytes of SRAM, on the hours whose program carries such a site. The public claim rests on 2x margins (chip-model-v3.md); 1.067x does not move it, and the union of the windows is still the whole dataset every hour, so no window-level cache exists. What moves: per tier nothing in rate or watts (the honest card reads the hot item from L2 as the chip would), and the 5 percent rule of 2.0 is untouched.
## 5. The two options for the flip, priced (main's ruling 3; nothing ships on this without the project lead's word)
| Option | What changes | Cost | Risk |
|---|---|---|---|
| A. A class amendment in 0.3.19 before the flip: the generator draws a load's source from the registers whose last writer injects (or rule (a) tightened to the same), class v4 re-pinned | a new program stream: new vectors, the seven gate packs re-exported, the six gates again (the hash side G1 to G3 and the verifier re-run here in about an hour of Mac and PC 2 time; G4 to G6 the node lane), every node before the flip by the one-box-at-a-time fleet rule | hours of gate time, a fleet rollout, the 0.3.19 ship on the line | a node that misses the build splits the chain at the flip; the fix itself is small (one draw rule) |
| B. Hold v4 at the floor as it is; the source rule in class v5 | nothing on the devnet; the attack-pass record carries the window null and the bound | a hot set on 48 percent of hours worth up to 1.005x to a chip, on 5 percent of hours up to 1.05x, 1.067x at the rule's ceiling, no chain risk | the public line must state the bound, not "uniform" |
The number that decides it: 1.067x at the ceiling against the 2x margin of the chip claim. Recommendation: B, with the v5 item below, unless the project lead wants the tail tight now.
## 6. The acceptance bound for the next class (main's ruling 4)
Definition: for a program, H = W_0.1(windows) + sum over load sites of h(last writer of the source), with W from the Poisson mixture of the 16 window draws (0.115 to 0.251 percent at 2^26 nonces) and h = 0.30 percent for `or`, 4.5 for an `or` chain, 0.067 for `mul`, 0.049 for `mulhi`, 0 for an injecting or bijective writer (the figures of section 2 at the measured bias). The bound: H at or under 1.2 x W, which is the static rule "every load's source was last written by an injecting op or a rotate" (any lossy writer breaks 1.2x). Its cost as a rejection rule on today's stream: 96.6 percent of candidates, about 30 attempts per seed on average. The cheaper form is a generator draw, not a rejection: draw a load's source from the registers whose last writer injects (today's rule draws from every written register), which costs no attempts and leaves rule (a) as it is. Either way the 64-seed census of F8's phase E is the gate, with the dynamic check extended to count saturated load sources over the 64 units beside the final values.
## 7. What is unverified
- The per-site h figures are the popcount and trailing-zeros models at the biases F8 measured on p1 and p2; p3's chain figure is F8's measurement, not a model. F8's phase E (64 seeds, dynamic) is the test of the whole table.
- The window model's top-f shares for the census use a normal approximation per quarter (p1's exact Poisson 0.160 against 0.159).
- No GPU run and no timing here; every number is a count or arithmetic.

14
tools/ca3-v4-uniform/Cargo.lock generated Normal file
View file

@ -0,0 +1,14 @@
# This file is automatically @generated by Cargo.
# It is not intended for manual editing.
version = 4
[[package]]
name = "ca3-v4-uniform"
version = "0.1.0"
dependencies = [
"igneum-pow",
]
[[package]]
name = "igneum-pow"
version = "0.2.0"

View file

@ -0,0 +1,11 @@
[package]
name = "ca3-v4-uniform"
version = "0.1.0"
edition = "2021"
publish = false
[dependencies]
igneum-pow = { path = "../../igneum-pow" }
[profile.release]
opt-level = 3

View file

@ -0,0 +1,136 @@
//! ca3-v4-uniform: the static census behind the AP-F8-1 analysis (`docs/analysis/ca3-v4-uniform.md`, 7 October 2026).
//! For chain-shaped class v4 programs (generator 4, the era drawn inside the class, the shadow block), one line per
//! program: the 16 load sites' window draws (the windows-union model's input) and, per site, the op that last wrote
//! the load's source register before the load in execution order (the base instructions before it in the iteration,
//! else the shadow block of the previous iteration, else the base instructions after it in the previous iteration),
//! classified as injecting (add, sub, xor, mad, shfl, load: the source is fresh), bijective (rotl, rotr: entropy kept)
//! or lossy (or: bits pinned to 1 with probability 3/4; mul: low bits zeroed by the multiplier's trailing zeros;
//! mulhi: the high word of a product, dense near 0). A lossy last writer is what AP-F8-1 found at p1's site 15
//! (`or` at 61, the load at 63) and p3's site 15 (`or` at 30, the load at 62); it passes the acceptance rule's
//! part (a), which takes any write as fresh. Nothing in `igneum-pow` is touched.
//!
//! ca3-v4-uniform [--n 1024] [--prefix igneum-ca3-v4-uniform] [--f8] (--f8 adds F8's p1, p2, p3)
use igneum_pow::generator::{generate_era, EraParams, Instr, LoadClass, Op, Program, V3_ALLOWED, V4_CLASS};
use igneum_pow::seed::seed_words_from_bytes;
const GENESIS_HEX: &str = "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07";
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum Kind { Inject, Bijective, Lossy, Init }
fn kind(op: Op) -> Kind {
if op.injects() { Kind::Inject } else {
match op {
Op::Rotl | Op::Rotr => Kind::Bijective,
Op::Or | Op::Mul | Op::MulHi => Kind::Lossy,
_ => Kind::Inject,
}
}
}
/// The last writer of `reg` before base instruction `k` in execution order: base 0..k of this iteration, then the
/// shadow block (run at the end of the previous iteration), then base k+1..63 of the previous iteration.
fn last_writer(p: &Program, k: usize, reg: u8) -> (Op, String) {
for j in (0..k).rev() {
if p.instrs[j].dst == reg { return (p.instrs[j].op, format!("base {j}")); }
}
for (j, s) in p.shadow.iter().enumerate().rev() {
if s.dst == reg { return (s.op, format!("shadow {j}")); }
}
for j in (k + 1..p.instrs.len()).rev() {
if p.instrs[j].dst == reg { return (p.instrs[j].op, format!("base {j} (previous iteration)")); }
}
(Op::Add, "init".into())
}
/// For an `or` writer at base `j`: how many of its two operands were themselves last written by `or` (an OR of an OR
/// compounds the bias: 3/4 of the bits set becomes 7/8, then 15/16; p3's site 15 is such a chain, p1's `or` has a
/// `mul`-written operand and is not).
fn or_chain(p: &Program, j: usize, ins: &Instr) -> u32 {
let mut n = 0;
for r in [ins.dst, ins.src] {
let (op, _) = last_writer(p, j, r);
if op == Op::Or { n += 1; }
}
n
}
fn seed32(s: &str) -> Vec<u8> {
seed_words_from_bytes(s.as_bytes()).iter().flat_map(|x| x.to_le_bytes()).collect()
}
fn unhex(s: &str) -> Vec<u8> { (0..s.len()).step_by(2).map(|i| u8::from_str_radix(&s[i..i + 2], 16).unwrap()).collect() }
fn main() {
let args: Vec<String> = std::env::args().collect();
let get = |k: &str, d: &str| -> String { args.windows(2).find(|w| w[0] == k).map(|w| w[1].clone()).unwrap_or(d.into()) };
let n: usize = get("--n", "1024").parse().unwrap();
let prefix = get("--prefix", "igneum-ca3-v4-uniform");
let f8 = args.iter().any(|a| a == "--f8");
let base = LoadClass { era: None, ..V4_CLASS };
let mut specs: Vec<(String, Vec<u8>, Vec<u8>)> = Vec::new();
if f8 {
let g = unhex(GENESIS_HEX);
specs.push(("p1-devnet-epoch0".into(), g.clone(), g));
for k in 2..=3 { specs.push((format!("p{k}-attack-f8"), seed32(&format!("igneum-attack-f8/program/{k}")), seed32(&format!("igneum-attack-f8/era/{k}")))); }
}
for i in 0..n { specs.push((format!("{prefix}/{i}"), seed32(&format!("{prefix}/program/{i}")), seed32(&format!("{prefix}/era/{i}")))); }
let _ = EraParams::test_era_bytes("igneum-era-test/0");
let (mut programs, mut with_lossy, mut with_or, mut with_chain, mut with_mulhi, mut with_mul) = (0usize, 0usize, 0usize, 0usize, 0usize, 0usize);
let mut sites_by = [0usize; 4];
let mut lossy_sites_total = 0usize;
let mut attempts_total = 0u64;
println!("# label\tid\tattempt\twindows(site:k:off)\tlossy_sites\tor_sites\tor_chain_sites\tmul_sites\tmulhi_sites\tper-site last writer (site:instr:src:op@where:kind)");
for (label, epoch, era) in &specs {
let p = generate_era(&format!("igneum-epoch/{}", label), epoch, base, era, &V3_ALLOWED);
programs += 1;
attempts_total += p.attempt as u64;
let mut wins = Vec::new();
let mut rows = Vec::new();
let (mut lossy, mut ors, mut chains, mut muls, mut mulhis) = (0, 0, 0, 0, 0);
let mut site = 0;
for (k, ins) in p.instrs.iter().enumerate() {
if !ins.op.is_load() { continue; }
wins.push(format!("{site}:{}:{}", ins.win, ins.off));
let (op, wh) = last_writer(&p, k, ins.src);
let kd = if wh == "init" { Kind::Init } else { kind(op) };
sites_by[kd as usize] += 1;
let mut tag = format!("{:?}", kd).to_lowercase();
if kd == Kind::Lossy {
lossy += 1;
match op {
Op::Or => {
ors += 1;
if let Some(j) = wh.strip_prefix("base ").and_then(|s| s.split(' ').next()).and_then(|s| s.parse::<usize>().ok()) {
let c = or_chain(&p, j, &p.instrs[j]);
if c > 0 { chains += 1; tag = format!("lossy-or-chain{c}"); }
}
}
Op::Mul => muls += 1,
Op::MulHi => mulhis += 1,
_ => {}
}
}
rows.push(format!("{site}:{k}:r{}:{}@{}:{}", ins.src, op.name(), wh.replace(' ', "_"), tag));
site += 1;
}
lossy_sites_total += lossy;
if lossy > 0 { with_lossy += 1; }
if ors > 0 { with_or += 1; }
if chains > 0 { with_chain += 1; }
if muls > 0 { with_mul += 1; }
if mulhis > 0 { with_mulhi += 1; }
println!("{label}\t{:016x}\t{}\t{}\t{lossy}\t{ors}\t{chains}\t{muls}\t{mulhis}\t{}", p.program_id(), p.attempt, wins.join(" "), rows.join(" "));
}
println!(
"CENSUS: {programs} chain-shaped class v4 programs; sites by last-writer kind inject {} bijective {} lossy {} init {}; lossy sites per program {:.3}; programs with >=1 lossy-sourced load {} ({:.1}%), >=1 or-sourced {} ({:.1}%), >=1 or-chain-sourced {} ({:.1}%), >=1 mul-sourced {} ({:.1}%), >=1 mulhi-sourced {} ({:.1}%); mean attempt {:.3}",
sites_by[0], sites_by[1], sites_by[2], sites_by[3],
lossy_sites_total as f64 / programs as f64,
with_lossy, 100.0 * with_lossy as f64 / programs as f64,
with_or, 100.0 * with_or as f64 / programs as f64,
with_chain, 100.0 * with_chain as f64 / programs as f64,
with_mul, 100.0 * with_mul as f64 / programs as f64,
with_mulhi, 100.0 * with_mulhi as f64 / programs as f64,
attempts_total as f64 / programs as f64
);
}

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,6 @@
flat Poisson top-0.1% share at 2^26: 0.115% (F8 control 0.115%); window model over 1027 programs: min 0.115% median 0.146% max 0.251% (densest quarter 1.00x to 2.31x flat)
p1: window model W_0.1% 0.159%, lossy sites ['mul', 'mulhi', 'or'], predicted S_0.1% 0.579% (F8 measured 0.520% at 2^26)
p2: window model W_0.1% 0.140%, lossy sites ['mul', 'mul', 'mulhi'], predicted S_0.1% 0.323% (F8 measured 0.272% at 10^6, flat control 0.243%)
p3: window model W_0.1% 0.153%, lossy sites ['mul', 'or-chain'], predicted S_0.1% 4.720% (F8 measured 4.595% at 10^6)
gates over the 1,024 census programs: F8's 1.2x-of-window: 96.6% fail; 1.2x-of-flat: 99.3%; 2x-of-window: 69.7%; 3x: 48.9%; S_0.1% <= 1% of reads: 6.9%; no or-sourced load (static): 48.4%; no lossy-sourced load (static): 96.6%
predicted S_0.1% over the census: median 0.45%, 90th pct 0.88%, 99th 5.29%, max 9.82% (or-chain programs 4.9%)