32 KiB
Attack report: the chained cache, partial-state and hot-set attacks (lane adv-cache-2)
Internal adversarial pass, not an independent review. Every sentence in this file that could be quoted publicly carries that label. Nothing here is a measurement of a chip; every chip figure is arithmetic on the chip model's inputs.
Header
| Field | Value |
|---|---|
| Target commit | 017e70376489251e18564c0abce7e466e606c8b3 (class v4 sub-version 3, object byte 7) |
| Base of this branch | build/master 04c4d9bc (re-merged 18:4x UTC, 7 October 2026); git diff --quiet 017e7037 HEAD -- igneum-pow prints IDENTICAL, so every build here is of the frozen crate |
| Harness | tools/attack/adv-cache-2/ (crate attack-adv-cache-2, binary adv-cache-2, igneum-pow by path), plan docs/plans/cryptanalysis/plan-chained-cache-2.md |
| Binary sha256 (first 16 hex) | v1 937949d03b7a369f (self-test, plants, Q1 2^28, Q3(1), Q3(3) small); v1b 6ccb052b0b2a8e20 (Poisson guard; Q2 devnet3 2^24, devnet 2^26, devnet3 2^26); v2 a4ff595a6799d2b4 (store pricing, control-referenced line tests, worst-site diagnostic; every run from 18:58 UTC) |
| Boxes | build box 1 (96 threads, 125 GB: the CPU-bound censuses) and build box 2 (96 threads, 125 GB: the table-heavy warps censuses); nice 10 on cores 8 to 95; logs under /srv/builds/_adv-adv-cache-2/ on each box, copied into docs/analysis/cryptanalysis/logs/adv-cache-2/ on this branch |
| Public kit | packs-ca3-v4-sub3-20261007.zip sha256 4f2445c5...f829154 verified on box 1; Devnet 3 pack v4-devnet3-epoch0.zip sha256 e025750f...65b334 verified on the read-only copy |
| Validation of the mirrors | Cache FNV-1a 64 of day 20730 448274a57f508cbc and of day 20733 7334fa46e5d972eb match the packs; program ids a785001687d8688a (attempt 1) and fce15bf61030be57 (attempt 0) regenerated from the seeds; warp 0 of both programs equals vectors.json; the Devnet 3 FNV-1a 64 fingerprint over the 2^24 outputs from base 0 reads e510ad92b4d24846 through the traced mirror, the brief's value; every census compares a sample of its items against derive_items and of its warps against Epoch::hash_warp (0 mismatches everywhere, counts per row) |
| Clock | every time in this file is UTC, as the logs stamp it |
Status board
| # | Question | Method | Known-failed shape (fired?) | Gate | Result | Status |
|---|---|---|---|---|---|---|
| Q1 | Line-index distribution over 2^22 lines, per round and pooled; segments; depth | adv-cache-2 lines, 16 days x 2^24 items = 2^28 derivations, 2^31 line reads; 256 days (2^32) queued |
--plant quarter-lines: segments +24.75 sigma, FLAGGED. --plant half-lines: +11.67 sigma, FLAGGED. Both fired |
Largest segment bucket within 6 sigma per round and pooled; depth flat; top 1 percent of lines under 1 percent plus chance | Pooled 16 days: segments max +4.84 sigma (control +4.24), min -4.25; lines max +5.61 (control +5.35), chi2/dof 0.99937; depth max +1.82 sigma; per round segments max +4.70; top 0.1 / 1 percent of lines 0.11519 / 1.11979 percent against the control's 0.11516 / 1.11958 (1.0003x / 1.0002x); 0 mirror mismatches on 1/64 of the items | PASS (BOUND at 2^31 reads; 2^35 queued) |
| Q2a | Hot set of items and lines across the hashes of the two real programs | adv-cache-2 warps: item table 2^24 with 8 lines each, the mirrored interpreter with the shadow block; devnet3 at 2^24 nonces (2^31 reads), devnet at 2^26 (2^33 reads); per site against the site's own window |
--plant const-item (site 0 fed one item): hot set FLAGGED, site 0 outside its window on every read, hottest item 6.25 percent of reads. Fired |
Top 0.1 percent of items under 1.2x the window-model control and X_f < f; lines the same; per site the same | Devnet 3 at 2^24 nonces: items 1.0032x / 1.0021x of the control at 0.1 / 1 percent; lines 0.9999x. Devnet 3 at 2^26: items 1.0071x / 1.0048x, lines 1.0000x. Devnet at 2^26: items 1.0002x / 1.0001x; lines 1.0002x; verdict PASS. FINDING at one load site of Devnet 3: site 0 (instr 3, window the whole dataset) chi2/dof 1.674 at 2^24 nonces and 3.702 at 2^26 (a fixed per-item bias, since the excess grows with the reads), its top 0.1 / 1 percent carry 1.449x / 1.416x the control's share at 2^26, largest item 101 reads at a mean of 32; the other 15 sites 1.000x. Attribution (diagnostic, 2^24 nonces per iteration): iteration 0 uniform (chi2/dof 0.9996), iterations 1 to 7 each at 1.110; the source register's top 16 bits at full entropy (15.997 bits) and never saturated, so the bias sits in its low bits (a multiply as the last writer is the shape; the v3 run prints P(bit b) and the writer). The devnet's worst site is 1.011x (site 6, chi2/dof 1.039). Worth 0.1 percent of a hash's reads to a hottest-items store (section 4.1) | FINDING (small, priced, attributed to one site's source) and BOUND; v3 diagnostic queued (item 25) |
| Q2b | Hot set over drawn programs | 16 epochs under the devnet era + 16 drawn eras, 2^24 nonces each | the same plant | the same gates per program; the census line PASS/FAIL | queued (items 23, 24 on box 2) | RUNNING |
| Q3(1) | Steering by the item index t | adv-cache-2 steer, 2^20 random t x 32 bit flips, rounds 0, 1, 7, days 20730 and 20733 |
--plant t-low (round-0 address = t AND mask): identity rows at P = 1.0, dead rows for bits 22..31, 40,960 equal pairs at +231,704 sigma. Fired |
every cell within 6 sigma of 0.5; equal pairs within 6 sigma of 8 | Day 20730: worst cell 3.44 sigma; equal pairs 13 / 4 / 6 of 2^25 against 8.00 (+1.77 / -1.41 / -0.71 sigma). Day 20733: worst 3.95 sigma; pairs 8 / 11 / 7. 2 x 2,112 cells, none beyond 4 sigma | PASS (BOUND at 2^20 items x 32 flips per day) |
| Q3(2) | Steering by the day key: the weak-day scan | adv-cache-2 days, 16,384 consecutive chain days x 2^16 items (2^30 derivations), per day the segment and depth histograms and a same-size control |
--plant quarter-lines --plant-day: fired on one day in a small run before the row counts (queued) |
no day with segment chi2_z over 6 or a depth bucket beyond 6 sigma; the per-day max against the control's | queued (item 15 on box 1) | RUNNING |
| Q3(3) | Steering by the parameters: the window layer's concentration | adv-cache-2 windows, 4,096 programs under the devnet era and 4,096 under drawn eras, the expected share per quarter from the 16 sites' window draws; measured on the real programs in Q2 |
--plant all-quarter (every site forced to one quarter): top-quarter share 1.0000 on 8 programs. Fired |
none: the distribution is the result (by design above uniform) | Real programs: devnet top quarter 0.42187 of reads (model 0.42188), top aligned half 0.53125; Devnet 3 top quarter 0.39063 (model 0.39062), top half 0.71875. 8 drawn programs: top quarter mean 0.3555, max 0.4531; top half mean 0.6328, max 0.7812. 4,096 + 4,096 running | FINDING (designed, counted; priced in 4.1) |
| Q4 | The price: ops per item and the SRAM column at the measured hit rates | arithmetic in section 4 from Q1, Q2, Q3(3); the v2 binary prints the stride and hottest-lines costs (nearest held line walk) and the hottest-items store per program | hand points: stride f = 1/2 is 0.5 evaluations per read; f = 1/64 is 31.5; the model's f = 0.25 row is 899,072 ops per hash | a BREAK is a store whose measured hit rate puts ops per hash under the model's row for that f by over 10 percent | the only measured excess over f is the window layer's (an aligned half serves 53 to 78 percent of reads: the model's f = 0.5 row overstates the recompute share by up to 1.8x) and the line reference multiplicity (a hottest-lines half store hits 57.8 percent instead of 50, at a higher miss cost than the stride) | FINDING against the model's table, not the hash; v2 rows RUNNING |
| GPU | any GPU measurement | none in this lane | no GPU on either box | BLOCKED |
1. Q1: the line index over 2^31 reads (16 days)
Command (box 1, queue item 10-adv-cache-2-lines-2e28.sh, started 19:23:36 UTC, 63 s wall on 80 threads):
adv-cache-2 lines --day0 20730 --days 16 --items-log2 24 --threads 80 --validate sample --out /srv/builds/_adv-adv-cache-2
Log: logs/adv-cache-2/lines-2e28.log. Seeds: the chain day keys seed_words_from_bytes("igneum-day/" || le64(d)), d = 20730 .. 20745. Control seeds 0xADC20000 + d and 0xADC21111.
| Statistic | Real (pooled, 16 days) | Uniform control, same size |
|---|---|---|
| Line reads | 2,147,483,648 over 2^22 lines, mean 512, sigma 22.63 | the same |
| Largest line | 639 (+5.61 sigma) | 633 (+5.35) |
| Smallest line | 402 (-4.86) | 403 (-4.82) |
| chi2 per dof, lines | 0.99937 (chi2_z -0.91) | 0.99926 (-1.08) |
| Segments (64 lines): largest, smallest | 33,645 (+4.84), 31,998 (-4.25) | 33,535 (+4.24), 32,042 (-4.01) |
| chi2 per dof, segments | 1.00226 | 1.00765 |
| Depth j (64 bins, mean 33,554,432): largest, smallest | +1.82 sigma (j = 6), -3.00 (j = 21) | |
| Top 0.1 / 0.5 / 1 percent of lines, share of reads | 0.11519 / 0.56510 / 1.11979 percent | 0.11516 / 0.56497 / 1.11958 |
| Per round (8 rounds), segments: largest bucket | +4.22 to +4.70 sigma | |
| Per round, depth: largest bucket | +2.19 to +2.84 sigma | |
| Distinct lines per item (of 8) | 8 in 268,433,562 items, 7 in 1,894; P(8) 0.99999294 against 0.99999332 uniform | |
| Distinct segments per item | 8 in 268,321,060, 7 in 114,380, 6 in 16; P(8) 0.999574 against 0.999573 | |
Mirror vs derive_items |
4,194,304 items compared, 0 mismatches |
Reading: no round, no segment and no depth is distinguishable from uniform at 2^31 reads; the top 1 percent of lines carry 1.120 percent of reads, the control 1.120 percent. A chip that learns the hottest lines of a day from a census of this size gains nothing over picking lines at random: the hottest-lines store's hit rate at f = 1/2 reads 0.5176 against 0.5 because the top half by count of a 512-mean Poisson histogram sits 0.8 sigma above the mean by selection, and the control gives the same.
2. Q2: the hot set across the hashes of an epoch
2.1 Devnet 3, epoch 0 (program fce15bf61030be57, day 20733), 2^24 nonces
Command (box 2, queue item 20-adv-cache-2-warps-devnet3-2e24-fp.sh, started 19:23:40 UTC, 70 s wall on 80 threads):
adv-cache-2 warps --programs devnet3 --day 20733 --nonces 16777216 --threads 80 --validate sample --fingerprint 1 --out /srv/builds/_adv-adv-cache-2
Log: logs/adv-cache-2/warps-devnet3-2e24-fp.log. Validation: table 2^24 items vs derive_items on 1/64 of them, 0 mismatches; 589 warps vs Epoch::hash_warp, 0 mismatches; warp 0 equals the pack; fingerprint e510ad92b4d24846 equals the brief's.
Window layer of this program (sites instr:win:off): 3:0:0, 8:2:2, 14:1:0, 15:0:0, 20:1:1, 26:1:1, 28:1:1, 35:0:0, 40:2:3, 43:1:0, 47:1:1, 49:0:0, 52:2:3, 53:0:0, 61:1:1, 62:1:1. Expected share per quarter of the item space 0.1406 / 0.1406 / 0.3281 / 0.3906; measured 0.14062 / 0.14063 / 0.32812 / 0.39063.
| Statistic (2^31 reads over 2^24 items, mean 128) | Real | Windowed control | Flat control |
|---|---|---|---|
| Largest item | 276 (+13.08 sigma flat) | 286 (+13.97) | 193 (+5.75) |
| chi2 per dof (flat) | 26.54 | 26.50 | 1.0003 |
| Top 0.1 percent of items, share | 0.19053 percent | 0.18993 (ratio 1.0032x) | 0.13112 |
| Top 1 percent | 1.80906 | 1.80531 (1.0021x) | 1.24354 |
| Lines (2^34 line reads, mean 4,096): top 0.1 / 1 percent | 0.17243 / 1.56005 | 0.17245 / 1.55989 (0.9999x / 1.0001x) | |
| Lines chi2 per dof | 154.40 | 154.36 | |
| Per site, top 0.1 percent over the site's own uniform-on-window control | 15 sites at 0.998x to 1.0017x; site 0 at 1.2894x (top 1 percent 1.2538x) | ||
| Site 0 (instr 3, win 0): chi2 per dof over 2^24 items, largest item | 1.67431 (chi2_z +1953), 36 at a Poisson mean of 8 | sites 1 to 15: 0.9995 to 1.0010, largest 27 to 65 (the Poisson maxima for their means) |
Reading. Over all 128 load positions the item reads are the window model and nothing else: the ratio to the windowed control is 1.003x at 0.1 percent. The lines carry chi2/dof 154 under BOTH real and control: a line is referenced by a Poisson(32) number of (item, round) slots, so its read count is about 128 x Poisson(32) and varies by 18 percent by construction; that is the honest construction's variance, not a hot set, and the control reproduces it to 0.9999x. One site is not uniform: site 0 of this program (the load at instruction 3, the first load of the base program, window the whole dataset) concentrates its reads at the item level (chi2/dof 1.67 where every other site reads 1.000). Worth to a chip: site 0 is 1/16 of reads; its top 1 percent of items take 2.58 percent of its reads against 2.06 percent for a uniform site; the excess is 0.52 percent of 1/16 of reads, 0.03 percent of a hash's reads. The diagnostic pass (v2, per iteration, source-register entropy and saturation) is queue item 25; the same statistic over 32 drawn programs (items 23, 24) says how common such a site is.
2.1b Devnet 3 at 2^26 nonces (item 22, v2 binary, 19:55:43 to 20:03:37 UTC, interpretation 315 s)
Log: logs/adv-cache-2/warps-devnet3-2e26.log. 2,167 warps vs Epoch::hash_warp, 0 mismatches. Items: top 0.1 percent 0.17409 percent vs the windowed control 0.17286 (1.0071x), top 1 percent 1.69081 vs 1.68270 (1.0048x); hot-set verdict clear (X_f / f = 0.012). Lines: 0.17222 vs 0.17222 (1.0000x); lines-vs-control chi2/dof 614.59 vs 614.39, matches. Windowed 64-item buckets: largest +21.83 sigma against the control's +4.69, FLAGGED: that is site 0's bias seen at the bucket level. Site 0 (instr 3): chi2/dof 3.70165, largest item 101 (+12.20 sigma at mean 32), top 0.1 / 1 percent 0.23843 / 2.12531 percent against its control's 0.16452 / 1.50121 (1.4492x / 1.4157x). Every other site 0.9993x to 1.0017x.
Diagnostic of site 0, per iteration (2^24 nonces each): iteration 0 chi2/dof 0.9996 (uniform), iterations 1 to 7 at 1.1099 to 1.1105 each, so the bias is the same fixed weight at every iteration after the first; the source register's top 16 bits are at 15.997 of 16 bits of entropy with a largest 2^-16 bucket at 0.00199 percent (uniform 0.00153) and no saturated value in 2^27 reads. A fixed per-item weight with a relative standard deviation of about 0.33 (chi2/dof - 1 = mean x Var(w): 0.11 at mean 1, 2.70 at mean 32, both give Var(w) = 0.084 to 0.11) that leaves the top 16 bits uniform is the signature of a biased low bit of the source: bit 0 of a product is 1 with probability 1/4, and the era stride rotl(x * M, R) moves the low bits of x to bits R.. of the address, inside the item index. The v3 binary prints P(bit b = 1) for b = 0..7 and the last writer of the source register (item 25); the row is updated when it lands. What a chip gets: with the top 0.1 percent of items held it serves 0.238 instead of 0.165 percent of site 0's reads, 0.0046 percent of a hash's reads; with the top 1 percent, 0.04 percent of a hash's reads. The gain exists and is small. The class (a load whose source register was last written by a non-injecting op) is the acceptance lane's question; this lane hands it the measurement.
2.2 Shared devnet, epoch 0 (program a785001687d8688a, day 20730), 2^26 nonces
Command (box 2, queue item 21-adv-cache-2-warps-devnet-2e26.sh, started 19:00:50 UTC, 353 s wall on 80 threads at a box load of 360 to 550):
adv-cache-2 warps --programs devnet --day 20730 --nonces 67108864 --threads 80 --validate sample --out /srv/builds/_adv-adv-cache-2
Log: logs/adv-cache-2/warps-devnet-2e26.log. Validation: 2,167 warps vs Epoch::hash_warp, 0 mismatches; warp 0 equals the pack. Sites (instr:win:off): 1:2:3, 4:2:1, 6:2:1, 10:1:1, 12:0:0, 20:0:0, 27:0:0, 30:0:0, 35:2:3, 40:0:0, 41:2:1, 43:0:0, 45:2:1, 52:1:1, 53:0:0, 54:2:1.
| Statistic (2^33 reads over 2^24 items, mean 512) | Real | Windowed control | Flat control |
|---|---|---|---|
| Largest item | 1,038 (+23.25 sigma flat) | 1,011 (+22.05) | 635 (+5.44) |
| chi2 per dof (flat) | 119.005 | 118.994 | 1.0004 |
| Top 0.1 / 1 percent of items | 0.18602 / 1.81245 percent | 0.18599 / 1.81235 (1.0002x / 1.0001x) | 0.11521 / 1.11985 |
| Share per quarter | 0.10937 / 0.42187 / 0.17187 / 0.29688 | model 0.1094 / 0.4219 / 0.1719 / 0.2969 | |
| Lines (2^36 line reads, mean 16,384): top 0.1 / 1 percent | 0.17332 / 1.56635 | 0.17329 / 1.56631 (1.0002x / 1.0000x) | |
| Lines chi2 per dof | 630.31 | 630.31 | |
| Per site, top 0.1 percent over the site's own control | 14 sites at 0.9993x to 1.0013x; site 0 (instr 1) 1.0051x at chi2/dof 1.032; site 6 (instr 27) 1.0111x at chi2/dof 1.039 | ||
| Verdict line | PASS (no test fired) |
Reading: the first real program is the window model to 1.0002x; two of its sites carry a 3 to 4 percent excess variance at the item level (chi2/dof 1.03 to 1.04, far under Devnet 3's 1.67), worth under 0.01 percent of a hash's reads to a store.
3. Q3: steering
3.1 By the item index t (box 1, items 11 and 12, 3 s and 5 s)
adv-cache-2 steer --day 20730 --items-log2 20 --threads 80 log steer-2e20.log
adv-cache-2 steer --day 20733 --items-log2 20 --threads 80 log steer-2e20-d20733.log
2^20 random items (SplitMix64 seed 0xADC27777 ^ batch x golden), each flipped in every one of its 32 bits, the line index of rounds 0, 1 and 7 compared: 3 x 32 x 22 cells of P(address bit flips) against 0.5, sigma 0.00049. Day 20730: worst cell 3.44 sigma; day 20733: 3.95 sigma. Pairs (t, t XOR 2^b) with an equal round-0 line: 13 and 8 of 33,554,432 against 8.00 expected. Plant t-low (the round-0 address replaced by t AND mask) printed identity rows at P = 1.0, dead rows for bits 22 to 31 and 40,960 equal pairs: fired. Reading: after 8 mixer applications no bit of t steers any bit of the first line index, and the later rounds inherit that. The attacker does not choose t in any case (the program's load address does); this row says that even a chooser of t could not aim at a line.
3.2 By the day key: the weak-day scan (box 1, item 15)
RUNNING. Row updated when the log closes.
3.3 By the parameters: the window layer
The one designed non-uniformity. Each load site's window is the dataset, an aligned half or an aligned quarter (win = below(3), off), drawn per instruction from the program stream (spec 1.13.1). Under a uniform window draw, a site contributes 1/4 of its reads to a quarter when win = 0, 1/2 when win = 1 and the quarter is in its half, all of them when win = 2 and it is its quarter. The expected share of the hottest quarter over 16 sites is above 25 percent for almost every program, and the measured shares equal the model to four digits (devnet 0.42187 vs 0.42188; Devnet 3 0.39063 vs 0.39062), so the reads inside a window are uniform and the model is exact.
| Programs | Top-quarter share: min / median / mean / max | Top aligned half: median / mean / max |
|---|---|---|
8 drawn (devnet era, epoch seeds igneum-adv-cache-2/epoch/k, k = 1..8), by the draws |
0.2656 / 0.3594 / 0.3555 / 0.4531 | 0.6562 / 0.6328 / 0.7812 |
| Shared devnet epoch 0 (measured over 2^33 reads) | 0.42187 | 0.53125 |
| Devnet 3 epoch 0 (measured over 2^31 reads) | 0.39063 | 0.71875 |
Exact distribution from the draw (win uniform on {0, 1, 2}, off uniform; 16 sites; enumerated over the 74,613 compositions, scratch python3, reproduced in section 3.3 below) |
0.2500 / 0.3281 / 0.3382 / 1.0000 (p90 0.4062, p99 0.4688; P(over 0.45) 2.8 percent) | 0.5625 / 0.5811 / 1.0000 (p90 0.6562, p99 0.7500; P(at least 0.75) 2.0 percent) |
| 4,096 under the devnet era and 4,096 under drawn eras (the check of the closed form) | RUNNING (items 13, 17) |
The plant all-quarter (every site forced to one quarter) printed 1.0000 on every program: fired.
The closed form. A site's window is the whole dataset with probability 1/3, either aligned half with 1/6 each, any one aligned quarter with 1/12 each; its reads are 1/16 of the hash's. The share of quarter k is (1/16) x (n_W / 4 + n_{H containing k} / 2 + n_{Q = k}). The counts of the seven categories over 16 sites are multinomial, so the distribution of the hottest quarter's share (and of the hottest aligned half's) is a finite sum over the 74,613 compositions of 16 into 7 parts. Result: mean 0.3382 for the top quarter (uniform 0.25) and 0.5811 for the top half (uniform 0.5). At those means the chip model's f = 0.25 row recomputes 128 x (1 - 0.3382) = 84.7 items instead of 96 (ops per hash 793,400 against 899,072, 0.883x; compute-bound 63.0 MH/s against 55.6) and the f = 0.5 row 53.6 items instead of 64 (502,400 against 599,552, 0.838x; 99.5 MH/s against 83.4). The two real programs sit at the 90th and 99th percentiles of the quarter distribution (devnet 0.4219, Devnet 3 0.3906) and Devnet 3's half share 0.71875 at the 97th; the 8,192-program census checks the form.
4. Q4: the price
Nothing here is a chip measurement; the arithmetic is the chip model's (chip-model-v3.md sections 2, 5.2 and 5.4: 9,360 ops per item recomputed, 608 ops per cache block evaluation counted from chacha_block, 128 items per hash, SRAM at 128 mm^2 and $46 per 256 MiB at the N5 headline).
4.1 A chip holding a fraction f of the dataset (items)
The model's row prices 128 (1 - f) items recomputed per hash. Three measured facts move it:
| Store | Measured hit rate | Ops per hash (measured) | Model row at this f | Ratio |
|---|---|---|---|---|
| The hottest 0.1 / 1 percent of items by read count, either real program | the control's share x 1.003 or less: no hot set (Q2) | the model's | 1.00x | |
| Devnet 3's site 0 excess (the FINDING of 2.1) | +0.52 percent of 1/16 of reads | 0.03 percent fewer items recomputed | 1.00x | |
| An aligned quarter of the dataset, chosen per program from the window draws, f = 0.25 | 0.3906 (Devnet 3), 0.4219 (devnet), 0.27 to 0.45 over 8 drawn programs, mean 0.3555 | 128 x (1 - 0.4219) x 9,360 + 512 = 693,000 (devnet); 730,500 (Devnet 3); mean program 772,400 | 899,072 | 0.77x (devnet), 0.81x (Devnet 3), 0.86x (mean) |
| An aligned half, f = 0.5 | 0.53125 (devnet), 0.71875 (Devnet 3), mean 0.6328 over 8 drawn, max 0.7812 | 562,000 (devnet), 337,500 (Devnet 3), 440,400 (mean), 262,700 (max) | 599,552 | 0.94x, 0.56x, 0.73x, 0.44x |
So a partial-store chip that reads the epoch's program and keeps the right half of the dataset in DRAM recomputes 28 to 47 percent of its items instead of 50 percent: the f = 0.5 row's compute-bound rate rises from 83.4 MH/s to 89 (devnet) or 148 MH/s (Devnet 3), 1.07x to 1.78x the model's row, 0.65x to 1.09x the 5090 bare. This is a gain over the chip model's partial rows by design of the window layer, and it is a break of the model's assumption, not of the hash: the model's own verdict (5.4 point 1: every partial row is worse than f = 1 on dollars per MH/s, so nobody builds the partial chip) stands, because the f = 1 chip (store everything, recompute nothing) is still cheaper per hash than any partial row even at a 78 percent hit rate. Per tier: nothing for any honest card (a card holds the whole dataset); for the chip designer the half-store row reads up to 1.8x better than the model's table says, and the model's section 5 table should carry the window hit rate S_f in place of f in its "items recomputed" column. The distribution over 8,192 drawn programs lands with items 13 and 14.
4.2 A chip holding a fraction f of the cache (lines) and recomputing items
| Layout | Hit rate at f = 1/2 (measured) | Miss cost | Block evaluations per item (8 reads) | Ops per item | Against 9,360 |
|---|---|---|---|---|---|
| Stride (every second line held) | 0.49999 (Q1 pooled); 0.49999 (Devnet 3); 0.50006 (devnet) | 1 evaluation (the held predecessor) | 4.0 | 2,432 | +26 percent, for half the 128 mm^2 (64 mm^2, $23) |
| Prefix (lines 0..31 of every segment) | 0.50001 | (65 - L)/2 = 16.5 average | 66 | 40,100 | +4.3x: never the layout to pick |
| Hottest half of lines by the program's reference count | 0.57719 (Devnet 3, 2^26), 0.5782 (devnet); 0.5176 in the Q1 day census (selection noise only) | the nearest held predecessor, measured: 6.655 evaluations per item at 0.423 x 8 = 3.38 misses, 1.97 per miss | 6.655 (measured, Devnet 3; the windowed control 6.654) | 4,046 | +43 percent: worse than the stride's +26 |
Measured on Devnet 3 at 2^26 nonces (the v2 LINE STORE rows, real and windowed control agreeing to the fourth digit):
| f | Stride: hit, evaluations per item, ops per item (x of 9,360) | Hottest lines: hit, evaluations per item, ops per item (x of 9,360) |
|---|---|---|
| 1/2 | 0.50000, 4.000, 2,432 (+0.26x) | 0.57719, 6.655, 4,046 (+0.43x) |
| 1/4 | 0.24999, 12.001, 7,296 (+0.78x) | 0.31299, 20.948, 12,737 (+1.36x) |
| 1/8 | 0.12499, 28.000, 17,024 (+1.82x) | 0.16648, 47.552, 28,912 (+3.09x) |
| 1/16 | 0.06250, 59.998, 36,479 (+3.90x) | 0.08761, 89.993, 54,716 (+5.85x) |
| 1/32 | 0.03125, 123.989, 75,385 (+8.05x) | 0.04578, 141.427, 85,988 (+9.19x) |
| 1/64 | 0.01563, 251.966, 153,195 (+16.4x) | 0.02380, 187.477, 113,986 (+12.2x) |
The stride is the cheapest layout at every f down to 1/32; at 1/64 the hottest-lines store wins (the stride then holds line 0 only and walks 31.5 on average, the hottest store holds lines of every depth), and both cost over 12x the item. The hottest-lines store is the one place a chip knows more than "uniform": each line is referenced by Poisson(32) item slots, the chip can count them at day start (it derives every item anyway when it builds its table), and the top half by reference count takes 57.8 percent of the line reads under both real programs (the control gives the same number: it is the construction's variance, not a hot set of the mixer). But a hottest-lines store has no chain structure, so a miss walks back to a random held line: measured 1.97 evaluations per miss, 6.655 per item against the stride's 4.0. Reading: at the measured distributions the cheapest partial cache store is the stride, its price is the honest curve (+26 percent of ops per item at f = 1/2, +16x at f = 1/64), and no measured skew of the line index moves that curve. The items store (section 4.1) at the measured shares, Devnet 3, 2^26 nonces: f = 0.001 serves 0.00174 of reads (uniform 0.00100, control 0.00173), f = 0.01 serves 0.01691 (0.01683), f = 0.1 serves 0.16188 (0.16159), f = 0.25 serves 0.39075, f = 0.5 serves 0.71876, f = 0.75 serves 0.86689 (control 0.86598): ops per hash 0.9993x, 0.9930x, 0.9313x, 0.8124x, 0.5629x and 0.5333x of the model's rows. Every digit of that excess over f is the window layer (the control, which has only the window layer, gives the same shares) except site 0's 0.0001 at f = 0.001 and 0.0009 at f = 0.75.
4.3 The SRAM column
The chip model's recompute chip holds 256 MiB (128 mm^2, $46). At f = 1/2 of the cache the stride store saves 64 mm^2 and $23 and costs 26 percent more ops per item, 0.31x bare becoming 0.25x; with the SRAM deducted at equal silicon the row moves from 0.76x to 0.74x (the saved area is 8.5 percent of a 750 mm^2 die, the ops cost 26 percent). No measured non-uniformity makes the partial cache store pay.
4.4 Consequences per tier (the standing rule)
| Tier | What these numbers mean | What is being done |
|---|---|---|
| Home miner, one 8 / 12 / 16 GB card, any vendor or OS | Nothing changes for an honest card: it holds the whole 1 GiB dataset and the window layer costs it nothing (the chip model's own row). No hot set exists for a chip to exploit beyond the designed windows | The findings go to the chip model's table (section 4.1) and to the acceptance lane (section 2.1b), not to the miner |
| One 24 or 32 GB card | The same | The same |
| A rig, a pool user | The partial-store chip rows read up to 1.8x better than the model's table at the window hit rate, still under the full-store chip on every metric the model prices; the full-store chip's rows (over 2x per joule, the model's verdict) are unchanged by anything measured here | The verdict on the stored-dataset chip stands as the model wrote it; this lane adds no new chip |
| The public claim | Nothing measured here moves "under 2x" either way | None |
4.5 Q5: the first load and the abandon-on-miss attacker
The first load of a hash (site 0 in iteration 0) depends only on the nonce, the init words and the instructions before it, so a chip can predict it without the dataset; a chip holding a fraction f of items could skip nonces whose first item it does not hold. But a hash is a lottery ticket only when all 128 loads complete, and every later load depends on the dataset words read before it, so abandoning a hash at a miss saves the recompute of that hash and forfeits the ticket; the cost per completed ticket is unchanged, 128 (1 - S_f) items recomputed. The one read a chip can steer is 1 of 128 (0.8 percent of a hash's reads), and only by discarding nonces, which costs it the same fraction of tickets. Bound, by the argument; no run.
5. Box-hours and pod-hours
| Step | Box | Wall | Threads | Box-hours (wall x threads / 96) |
|---|---|---|---|---|
| Builds (6 on box 1, 3 on box 2) | 1, 2 | 9 x about 20 s | 44 to 88 | 0.04 |
| Self-test, five plants, the small real checks | 1, 2 | about 3 min | 40 to 60 | 0.03 |
| Q1 lines 2^28 | 1 | 63 s | 80 | 0.015 |
| Q3(1) steer, two days | 1 | 8 s | 80 | 0.002 |
| Q2 devnet3 2^24 | 2 | 70 s | 80 | 0.016 |
| Q2 devnet 2^26 | 2 | 353 s | 80 | 0.08 |
| Q3(3) windows 4,096 (devnet era) | 1 | RUNNING since 18:49 UTC | 80 | |
| Q2 devnet3 2^26 (v2) | 2 | 474 s | 80 | 0.11 |
| Q2b era-fixed 32 programs at 2^23 (v2) | 2 | RUNNING since 19:03 UTC | 80 | |
| Pod-hours | none (no GPU, no rented pod) | 0 |
Running total at 19:1x UTC: about 0.2 box-hours plus the two running items. The 8-hour reading is not near.
6. The bound reached, honestly
At 2^31 line reads over 16 day keys (Q1) the line index, the segment and the recompute depth are uniform to within the control's own fluctuations (largest excess +5.6 sigma on 2^22 bins against the control's +5.4); the top 1 percent of lines carry 1.1198 percent of reads, the control 1.1196. At 2^31 and 2^33 item reads on the two real programs (Q2) the item histogram is the window model to 1.003x at the top 0.1 percent; one load site of Devnet 3 is non-uniform at the item level (chi2/dof 1.67, 1.29x at the top 0.1 percent of its own reads) and is worth 0.03 percent of a hash's reads to a store; the devnet's worst site is 1.04 and 1.011x. No bit of t steers the line index (Q3(1), 2^20 items x 32 flips x 2 days, worst cell 3.95 sigma). The one counted gain is the designed window layer (Q3(3)): an aligned half of the dataset serves 53 to 78 percent of reads, so the chip model's partial-store rows overstate the recompute share by up to 1.8x; it does not change the model's verdict that the full-store chip is the cheapest. Not searched: the mixer itself (the adv-mixer lane), the chain function (adv-cache), a dataset over 2^28 words, any GPU behaviour.
7. What a longer pass would add
Not a reason to wait: the 2^35-read census (item 16) and the 16,384-day scan (item 15) tighten Q1 and Q3(2) by 4x in sigma and are queued; a census of the per-site chi2 over 8,192 drawn programs (items 23, 24 give 32) would say how often a site like Devnet 3's site 0 is accepted and how large its excess gets; and the window-layer distribution could be stated in closed form from the win draw (1/3 each) and checked against the 8,192-program census.
8. Rule changes during the pass (so the box-hours stay honest)
| Time (UTC, 7 October 2026) | Change |
|---|---|
| 19:3x | The plan's SIGSTOP yield-to-builds loop dropped by the coordinator's order (a build slot is held nearly all the time on both boxes, so a yielding sweep never ran); sweeps run at nice 10 on cores 8 to 95 and the scheduler arbitrates |
| 19:4x | Logs, pid files and results moved out of the worktree mirror to /srv/builds/_adv-adv-cache-2/ (the mirror sync deletes what is not in the tree); the CPU-bound sweeps to box 1 by order |
| 20:03 | Box 1 queue reordered (the weak-day scan and the 2^32 census before the drawn-era windows census, item 14 renamed 17; a small plant item 14 added); items 23 and 24 trimmed from 2^24 to 2^23 nonces per drawn program before they started; item 25 (a v2 repeat of Devnet 3 at 2^26) removed because item 22 had already run under v2, and re-used for the v3 diagnostic at 2^24 |
| 19:55 | One sweep per box at a time across every lane, under flock /srv/builds/_adv/locks/sweep.lock; both drains stopped and restarted at 19:58 to take every further queue item under the lock (the two items already running finish as they were, per the order); the lock file did not exist on either box at 19:58 and is created by the first flock |