adv-cache-2: Devnet 3 at 2^26 (site 0 attribution), measured line-store and item-store prices; v3 diagnostic (internal adversarial pass, not an independent review)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
3a76a81f46
commit
854773d424
2 changed files with 59 additions and 6 deletions
|
|
@ -20,7 +20,7 @@ Internal adversarial pass, not an independent review. Every sentence in this fil
|
|||
| # | Question | Method | Known-failed shape (fired?) | Gate | Result | Status |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q1 | Line-index distribution over 2^22 lines, per round and pooled; segments; depth | `adv-cache-2 lines`, 16 days x 2^24 items = 2^28 derivations, 2^31 line reads; 256 days (2^32) queued | `--plant quarter-lines`: segments +24.75 sigma, FLAGGED. `--plant half-lines`: +11.67 sigma, FLAGGED. Both fired | Largest segment bucket within 6 sigma per round and pooled; depth flat; top 1 percent of lines under 1 percent plus chance | Pooled 16 days: segments max +4.84 sigma (control +4.24), min -4.25; lines max +5.61 (control +5.35), chi2/dof 0.99937; depth max +1.82 sigma; per round segments max +4.70; top 0.1 / 1 percent of lines 0.11519 / 1.11979 percent against the control's 0.11516 / 1.11958 (1.0003x / 1.0002x); 0 mirror mismatches on 1/64 of the items | PASS (BOUND at 2^31 reads; 2^35 queued) |
|
||||
| Q2a | Hot set of items and lines across the hashes of the two real programs | `adv-cache-2 warps`: item table 2^24 with 8 lines each, the mirrored interpreter with the shadow block; devnet3 at 2^24 nonces (2^31 reads), devnet at 2^26 (2^33 reads); per site against the site's own window | `--plant const-item` (site 0 fed one item): hot set FLAGGED, site 0 outside its window on every read, hottest item 6.25 percent of reads. Fired | Top 0.1 percent of items under 1.2x the window-model control and X_f < f; lines the same; per site the same | Devnet 3: items 1.0032x / 1.0021x of the control at 0.1 / 1 percent; lines 0.9999x. Devnet: items 1.0002x / 1.0001x; lines 1.0002x; verdict PASS. FINDING at one load site of Devnet 3: site 0 (instr 3, window the whole dataset) chi2/dof 1.674 over 2^24 items, its top 0.1 / 1 percent carry 1.29x / 1.25x the control's share, largest item 36 reads at a Poisson mean of 8 (the other 15 sites 1.000x, max 27); the devnet's worst site is 1.011x (site 6, chi2/dof 1.039). Worth 0.03 percent of a hash's reads (section 4.1) | FINDING (small, priced) and BOUND; diagnostic queued (item 25) |
|
||||
| Q2a | Hot set of items and lines across the hashes of the two real programs | `adv-cache-2 warps`: item table 2^24 with 8 lines each, the mirrored interpreter with the shadow block; devnet3 at 2^24 nonces (2^31 reads), devnet at 2^26 (2^33 reads); per site against the site's own window | `--plant const-item` (site 0 fed one item): hot set FLAGGED, site 0 outside its window on every read, hottest item 6.25 percent of reads. Fired | Top 0.1 percent of items under 1.2x the window-model control and X_f < f; lines the same; per site the same | Devnet 3 at 2^24 nonces: items 1.0032x / 1.0021x of the control at 0.1 / 1 percent; lines 0.9999x. Devnet 3 at 2^26: items 1.0071x / 1.0048x, lines 1.0000x. Devnet at 2^26: items 1.0002x / 1.0001x; lines 1.0002x; verdict PASS. FINDING at one load site of Devnet 3: site 0 (instr 3, window the whole dataset) chi2/dof 1.674 at 2^24 nonces and 3.702 at 2^26 (a fixed per-item bias, since the excess grows with the reads), its top 0.1 / 1 percent carry 1.449x / 1.416x the control's share at 2^26, largest item 101 reads at a mean of 32; the other 15 sites 1.000x. Attribution (diagnostic, 2^24 nonces per iteration): iteration 0 uniform (chi2/dof 0.9996), iterations 1 to 7 each at 1.110; the source register's top 16 bits at full entropy (15.997 bits) and never saturated, so the bias sits in its low bits (a multiply as the last writer is the shape; the v3 run prints P(bit b) and the writer). The devnet's worst site is 1.011x (site 6, chi2/dof 1.039). Worth 0.1 percent of a hash's reads to a hottest-items store (section 4.1) | FINDING (small, priced, attributed to one site's source) and BOUND; v3 diagnostic queued (item 25) |
|
||||
| Q2b | Hot set over drawn programs | 16 epochs under the devnet era + 16 drawn eras, 2^24 nonces each | the same plant | the same gates per program; the census line PASS/FAIL | queued (items 23, 24 on box 2) | RUNNING |
|
||||
| Q3(1) | Steering by the item index t | `adv-cache-2 steer`, 2^20 random t x 32 bit flips, rounds 0, 1, 7, days 20730 and 20733 | `--plant t-low` (round-0 address = t AND mask): identity rows at P = 1.0, dead rows for bits 22..31, 40,960 equal pairs at +231,704 sigma. Fired | every cell within 6 sigma of 0.5; equal pairs within 6 sigma of 8 | Day 20730: worst cell 3.44 sigma; equal pairs 13 / 4 / 6 of 2^25 against 8.00 (+1.77 / -1.41 / -0.71 sigma). Day 20733: worst 3.95 sigma; pairs 8 / 11 / 7. 2 x 2,112 cells, none beyond 4 sigma | PASS (BOUND at 2^20 items x 32 flips per day) |
|
||||
| Q3(2) | Steering by the day key: the weak-day scan | `adv-cache-2 days`, 16,384 consecutive chain days x 2^16 items (2^30 derivations), per day the segment and depth histograms and a same-size control | `--plant quarter-lines --plant-day`: fired on one day in a small run before the row counts (queued) | no day with segment chi2_z over 6 or a depth bucket beyond 6 sigma; the per-day max against the control's | queued (item 15 on box 1) | RUNNING |
|
||||
|
|
@ -83,6 +83,12 @@ Window layer of this program (sites instr:win:off): 3:0:0, 8:2:2, 14:1:0, 15:0:0
|
|||
|
||||
Reading. Over all 128 load positions the item reads are the window model and nothing else: the ratio to the windowed control is 1.003x at 0.1 percent. The lines carry chi2/dof 154 under BOTH real and control: a line is referenced by a Poisson(32) number of (item, round) slots, so its read count is about 128 x Poisson(32) and varies by 18 percent by construction; that is the honest construction's variance, not a hot set, and the control reproduces it to 0.9999x. One site is not uniform: site 0 of this program (the load at instruction 3, the first load of the base program, window the whole dataset) concentrates its reads at the item level (chi2/dof 1.67 where every other site reads 1.000). Worth to a chip: site 0 is 1/16 of reads; its top 1 percent of items take 2.58 percent of its reads against 2.06 percent for a uniform site; the excess is 0.52 percent of 1/16 of reads, 0.03 percent of a hash's reads. The diagnostic pass (v2, per iteration, source-register entropy and saturation) is queue item 25; the same statistic over 32 drawn programs (items 23, 24) says how common such a site is.
|
||||
|
||||
### 2.1b Devnet 3 at 2^26 nonces (item 22, v2 binary, 19:55:43 to 20:03:37 UTC, interpretation 315 s)
|
||||
|
||||
Log: `logs/adv-cache-2/warps-devnet3-2e26.log`. 2,167 warps vs `Epoch::hash_warp`, 0 mismatches. Items: top 0.1 percent 0.17409 percent vs the windowed control 0.17286 (1.0071x), top 1 percent 1.69081 vs 1.68270 (1.0048x); hot-set verdict clear (X_f / f = 0.012). Lines: 0.17222 vs 0.17222 (1.0000x); lines-vs-control chi2/dof 614.59 vs 614.39, matches. Windowed 64-item buckets: largest +21.83 sigma against the control's +4.69, FLAGGED: that is site 0's bias seen at the bucket level. Site 0 (instr 3): chi2/dof 3.70165, largest item 101 (+12.20 sigma at mean 32), top 0.1 / 1 percent 0.23843 / 2.12531 percent against its control's 0.16452 / 1.50121 (1.4492x / 1.4157x). Every other site 0.9993x to 1.0017x.
|
||||
|
||||
Diagnostic of site 0, per iteration (2^24 nonces each): iteration 0 chi2/dof 0.9996 (uniform), iterations 1 to 7 at 1.1099 to 1.1105 each, so the bias is the same fixed weight at every iteration after the first; the source register's top 16 bits are at 15.997 of 16 bits of entropy with a largest 2^-16 bucket at 0.00199 percent (uniform 0.00153) and no saturated value in 2^27 reads. A fixed per-item weight with a relative standard deviation of about 0.33 (chi2/dof - 1 = mean x Var(w): 0.11 at mean 1, 2.70 at mean 32, both give Var(w) = 0.084 to 0.11) that leaves the top 16 bits uniform is the signature of a biased low bit of the source: bit 0 of a product is 1 with probability 1/4, and the era stride `rotl(x * M, R)` moves the low bits of `x` to bits R.. of the address, inside the item index. The v3 binary prints P(bit b = 1) for b = 0..7 and the last writer of the source register (item 25); the row is updated when it lands. What a chip gets: with the top 0.1 percent of items held it serves 0.238 instead of 0.165 percent of site 0's reads, 0.0046 percent of a hash's reads; with the top 1 percent, 0.04 percent of a hash's reads. The gain exists and is small. The class (a load whose source register was last written by a non-injecting op) is the acceptance lane's question; this lane hands it the measurement.
|
||||
|
||||
### 2.2 Shared devnet, epoch 0 (program `a785001687d8688a`, day 20730), 2^26 nonces
|
||||
|
||||
Command (box 2, queue item `21-adv-cache-2-warps-devnet-2e26.sh`, started 19:00:50 UTC, 353 s wall on 80 threads at a box load of 360 to 550):
|
||||
|
|
@ -160,9 +166,20 @@ So a partial-store chip that reads the epoch's program and keeps the right half
|
|||
|---|---|---|---|---|---|
|
||||
| Stride (every second line held) | 0.49999 (Q1 pooled); 0.49999 (Devnet 3); 0.50006 (devnet) | 1 evaluation (the held predecessor) | 4.0 | 2,432 | +26 percent, for half the 128 mm^2 (64 mm^2, $23) |
|
||||
| Prefix (lines 0..31 of every segment) | 0.50001 | (65 - L)/2 = 16.5 average | 66 | 40,100 | +4.3x: never the layout to pick |
|
||||
| Hottest half of lines by the program's reference count | 0.5774 (Devnet 3), 0.5782 (devnet), weighted by item reads; 0.5176 in the Q1 day census (selection noise only) | the nearest held predecessor: about 2 evaluations at density 1/2 | the v2 binary measures the exact walk; pending | pending | pending |
|
||||
| Hottest half of lines by the program's reference count | 0.57719 (Devnet 3, 2^26), 0.5782 (devnet); 0.5176 in the Q1 day census (selection noise only) | the nearest held predecessor, measured: 6.655 evaluations per item at 0.423 x 8 = 3.38 misses, 1.97 per miss | 6.655 (measured, Devnet 3; the windowed control 6.654) | 4,046 | +43 percent: worse than the stride's +26 |
|
||||
|
||||
The hottest-lines store is the one place a chip knows more than "uniform": each line is referenced by Poisson(32) item slots, the chip can count them at day start (it derives every item anyway when it builds its table), and the top half by reference count takes 57.8 percent of the line reads under both real programs (the control gives the same number: it is the construction's variance, not a hot set of the mixer). But a hottest-lines store has no chain structure, so a miss walks back to a random held line (about two evaluations instead of the stride's one), and 0.422 x 8 x 2 = 6.8 evaluations per item beats the stride's 4.0 only if the walk is under 1.18; the v2 binary measures the exact walk and the row is updated when item 25 lands. Reading so far: at the measured distributions the cheapest partial cache store is the stride, its price is the honest curve (+26 percent of ops per item at f = 1/2, +4.6x at f = 1/64: 31.5 evaluations per read, 153,000 ops per item), and no measured skew of the line index moves that curve.
|
||||
Measured on Devnet 3 at 2^26 nonces (the v2 `LINE STORE` rows, real and windowed control agreeing to the fourth digit):
|
||||
|
||||
| f | Stride: hit, evaluations per item, ops per item (x of 9,360) | Hottest lines: hit, evaluations per item, ops per item (x of 9,360) |
|
||||
|---|---|---|
|
||||
| 1/2 | 0.50000, 4.000, 2,432 (+0.26x) | 0.57719, 6.655, 4,046 (+0.43x) |
|
||||
| 1/4 | 0.24999, 12.001, 7,296 (+0.78x) | 0.31299, 20.948, 12,737 (+1.36x) |
|
||||
| 1/8 | 0.12499, 28.000, 17,024 (+1.82x) | 0.16648, 47.552, 28,912 (+3.09x) |
|
||||
| 1/16 | 0.06250, 59.998, 36,479 (+3.90x) | 0.08761, 89.993, 54,716 (+5.85x) |
|
||||
| 1/32 | 0.03125, 123.989, 75,385 (+8.05x) | 0.04578, 141.427, 85,988 (+9.19x) |
|
||||
| 1/64 | 0.01563, 251.966, 153,195 (+16.4x) | 0.02380, 187.477, 113,986 (+12.2x) |
|
||||
|
||||
The stride is the cheapest layout at every f down to 1/32; at 1/64 the hottest-lines store wins (the stride then holds line 0 only and walks 31.5 on average, the hottest store holds lines of every depth), and both cost over 12x the item. The hottest-lines store is the one place a chip knows more than "uniform": each line is referenced by Poisson(32) item slots, the chip can count them at day start (it derives every item anyway when it builds its table), and the top half by reference count takes 57.8 percent of the line reads under both real programs (the control gives the same number: it is the construction's variance, not a hot set of the mixer). But a hottest-lines store has no chain structure, so a miss walks back to a random held line: measured 1.97 evaluations per miss, 6.655 per item against the stride's 4.0. Reading: at the measured distributions the cheapest partial cache store is the stride, its price is the honest curve (+26 percent of ops per item at f = 1/2, +16x at f = 1/64), and no measured skew of the line index moves that curve. The items store (section 4.1) at the measured shares, Devnet 3, 2^26 nonces: f = 0.001 serves 0.00174 of reads (uniform 0.00100, control 0.00173), f = 0.01 serves 0.01691 (0.01683), f = 0.1 serves 0.16188 (0.16159), f = 0.25 serves 0.39075, f = 0.5 serves 0.71876, f = 0.75 serves 0.86689 (control 0.86598): ops per hash 0.9993x, 0.9930x, 0.9313x, 0.8124x, 0.5629x and 0.5333x of the model's rows. Every digit of that excess over f is the window layer (the control, which has only the window layer, gives the same shares) except site 0's 0.0001 at f = 0.001 and 0.0009 at f = 0.75.
|
||||
|
||||
### 4.3 The SRAM column
|
||||
|
||||
|
|
@ -179,7 +196,8 @@ The chip model's recompute chip holds 256 MiB (128 mm^2, $46). At f = 1/2 of the
|
|||
| Q2 devnet3 2^24 | 2 | 70 s | 80 | 0.016 |
|
||||
| Q2 devnet 2^26 | 2 | 353 s | 80 | 0.08 |
|
||||
| Q3(3) windows 4,096 (devnet era) | 1 | RUNNING since 18:49 UTC | 80 | |
|
||||
| Q2 devnet3 2^26 (v1b) | 2 | RUNNING since 18:55 UTC | 80 | |
|
||||
| Q2 devnet3 2^26 (v2) | 2 | 474 s | 80 | 0.11 |
|
||||
| Q2b era-fixed 32 programs at 2^23 (v2) | 2 | RUNNING since 19:03 UTC | 80 | |
|
||||
| Pod-hours | none (no GPU, no rented pod) | | | 0 |
|
||||
|
||||
Running total at 19:1x UTC: about 0.2 box-hours plus the two running items. The 8-hour reading is not near.
|
||||
|
|
|
|||
|
|
@ -1040,6 +1040,30 @@ fn run_program(spec: &ProgramSpec, ds: DatasetSource, table: &Table, day: u64, n
|
|||
let era = program.class.era.expect("class v4 draws the era");
|
||||
let sites: Vec<usize> = program.instrs.iter().enumerate().filter(|(_, i)| i.op == Op::Load).map(|(k, _)| k).collect();
|
||||
assert_eq!(sites.len(), 16);
|
||||
{
|
||||
// static shape of every site: its source register, the last base-program writer before it (cyclic) and the
|
||||
// ops of the shadow block that write that register (the shadow runs between iterations)
|
||||
let instrs = &program.instrs;
|
||||
let mut rows = Vec::new();
|
||||
for (si, &k) in sites.iter().enumerate() {
|
||||
let src = instrs[k].src;
|
||||
let mut writer = String::from("none");
|
||||
for back in 1..instrs.len() {
|
||||
let j = (k + instrs.len() - back) % instrs.len();
|
||||
if instrs[j].dst == src {
|
||||
writer = format!("{}@{j}", instrs[j].op.name());
|
||||
break;
|
||||
}
|
||||
}
|
||||
let mut sh: std::collections::BTreeMap<&str, usize> = std::collections::BTreeMap::new();
|
||||
for x in program.shadow.iter().filter(|x| x.dst == src) {
|
||||
*sh.entry(x.op.name()).or_insert(0) += 1;
|
||||
}
|
||||
let last_shadow = program.shadow.iter().rev().find(|x| x.dst == src).map(|x| x.op.name()).unwrap_or("none");
|
||||
rows.push(format!("s{si}:instr{k}:r{src}:base-writer {writer}:shadow-writes {}:last-shadow-writer {last_shadow}", sh.iter().map(|(o, n)| format!("{o}={n}")).collect::<Vec<_>>().join(",")));
|
||||
}
|
||||
log!("static {}: {}", spec.label, rows.join(" | "));
|
||||
}
|
||||
log!("program {}: epoch seed {} era seed {} id {:016x} attempt {} class {} op mix {}; era stride mul {:#010x} rot {} interleave {:?}; sites instr:win:off {}", spec.label, hex(&spec.epoch_seed), hex(&spec.era), program.program_id(), program.attempt, program.class.name(), program.op_mix(), era.stride_mul, era.stride_rot, era.pos, program.instrs.iter().enumerate().filter(|(_, i)| i.op == Op::Load).map(|(k, i)| format!("{k}:{}:{}", i.win, i.off)).collect::<Vec<_>>().join(" "));
|
||||
let epoch = Epoch { program, dataset: ds };
|
||||
let mirror = Mirror { program: &epoch.program, era: epoch.program.class.era.as_ref(), layout: epoch.program.class.layout(), mask: epoch.dataset.mask, log2: epoch.dataset.log2_words, table, plant_site: if plant == Plant::ConstItem { Some(0) } else { None } };
|
||||
|
|
@ -1260,11 +1284,14 @@ fn run_program(spec: &ProgramSpec, ds: DatasetSource, table: &Table, day: u64, n
|
|||
let per_iter: Vec<Vec<AtomicU32>> = (0..ITERATIONS).map(|_| atomic_vec(ITEMS_N)).collect();
|
||||
let hi16: Vec<Vec<AtomicU32>> = (0..ITERATIONS).map(|_| atomic_vec(1 << 16)).collect();
|
||||
let sat: Vec<AtomicU64> = (0..ITERATIONS).map(|_| AtomicU64::new(0)).collect();
|
||||
// P(bit b of the source = 1) for b = 0..7: a multiply as the last writer biases the low bits (bit 0 of a product
|
||||
// is 1 with probability 1/4), which the era stride then carries into the item index
|
||||
let lowbits: Vec<Vec<AtomicU64>> = (0..ITERATIONS).map(|_| (0..8).map(|_| AtomicU64::new(0)).collect()).collect();
|
||||
let next = AtomicUsize::new(0);
|
||||
let diag_warps = warps.min(1 << 19);
|
||||
std::thread::scope(|sc| {
|
||||
for _ in 0..threads {
|
||||
let (next, per_iter, hi16, sat, mirror) = (&next, &per_iter, &hi16, &sat, &mirror);
|
||||
let (next, per_iter, hi16, sat, mirror, lowbits) = (&next, &per_iter, &hi16, &sat, &mirror, &lowbits);
|
||||
sc.spawn(move || {
|
||||
let mut sink = Sink { items: Vec::with_capacity(POS), srcs: Vec::with_capacity(POS) };
|
||||
loop {
|
||||
|
|
@ -1276,6 +1303,7 @@ fn run_program(spec: &ProgramSpec, ds: DatasetSource, table: &Table, day: u64, n
|
|||
for it in 0..ITERATIONS {
|
||||
let p = it * 16 + s;
|
||||
let mut ns = 0u64;
|
||||
let mut lb = [0u64; 8];
|
||||
for lane in 0..LANES {
|
||||
per_iter[it][sink.items[p][lane] as usize].fetch_add(1, Ordering::Relaxed);
|
||||
let x = sink.srcs[p][lane];
|
||||
|
|
@ -1283,8 +1311,14 @@ fn run_program(spec: &ProgramSpec, ds: DatasetSource, table: &Table, day: u64, n
|
|||
if x == 0 || x == u32::MAX {
|
||||
ns += 1;
|
||||
}
|
||||
for b in 0..8 {
|
||||
lb[b] += ((x >> b) & 1) as u64;
|
||||
}
|
||||
}
|
||||
sat[it].fetch_add(ns, Ordering::Relaxed);
|
||||
for b in 0..8 {
|
||||
lowbits[it][b].fetch_add(lb[b], Ordering::Relaxed);
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
|
|
@ -1298,7 +1332,8 @@ fn run_program(spec: &ProgramSpec, ds: DatasetSource, table: &Table, day: u64, n
|
|||
let hh = snapshot(&hi16[it]);
|
||||
let mx = *hh.iter().max().unwrap() as f64 / n as f64;
|
||||
let ent: f64 = hh.iter().filter(|&&c| c > 0).map(|&c| { let p = c as f64 / n as f64; -p * p.log2() }).sum();
|
||||
log!("diag {} worst site {s} (instr {}) iteration {it}: {} nonces; item hist chi2/dof {:.4} z_max {:+.2} top0.1% {:.5}% top1% {:.5}%; source hi16: max bucket share {:.5}% (uniform {:.5}%) entropy {:.3} of 16 bits; saturated source {:.6}%", spec.label, sites[s], n, st.chi2_per_dof, st.z_max, st.top[0].1 * 100.0, st.top[2].1 * 100.0, mx * 100.0, 100.0 / 65536.0, ent, sat[it].load(Ordering::Relaxed) as f64 * 100.0 / n as f64);
|
||||
let lbs: Vec<String> = (0..8).map(|b| format!("{:.4}", lowbits[it][b].load(Ordering::Relaxed) as f64 / n as f64)).collect();
|
||||
log!("diag {} worst site {s} (instr {}) iteration {it}: {} nonces; item hist chi2/dof {:.4} z_max {:+.2} top0.1% {:.5}% top1% {:.5}%; source hi16: max bucket share {:.5}% (uniform {:.5}%) entropy {:.3} of 16 bits; saturated source {:.6}%; P(source bit b = 1) for b = 0..7: {}", spec.label, sites[s], n, st.chi2_per_dof, st.z_max, st.top[0].1 * 100.0, st.top[2].1 * 100.0, mx * 100.0, 100.0 / 65536.0, ent, sat[it].load(Ordering::Relaxed) as f64 * 100.0 / n as f64, lbs.join(" "));
|
||||
}
|
||||
}
|
||||
let (hottest, hottest_count) = snap.iter().enumerate().map(|(t, &c)| (t as u32, c)).max_by_key(|&(_, c)| c).unwrap();
|
||||
|
|
|
|||
Loading…
Reference in a new issue