From 64bb647e68d1f2367f0faa967bce5ab5101bbb2a Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Wed, 7 Oct 2026 22:15:58 +0000 Subject: [PATCH] adv-cache-2: drawn-era 17 to 20, the rotation rule corrected, sub-class B (internal adversarial pass, not an independent review) Co-Authored-By: Claude Fable 5.1 --- .../cryptanalysis/report-chained-cache-2.md | 16 +++++++++++++--- 1 file changed, 13 insertions(+), 3 deletions(-) diff --git a/docs/analysis/cryptanalysis/report-chained-cache-2.md b/docs/analysis/cryptanalysis/report-chained-cache-2.md index 464fb475e..1a065e0ac 100644 --- a/docs/analysis/cryptanalysis/report-chained-cache-2.md +++ b/docs/analysis/cryptanalysis/report-chained-cache-2.md @@ -21,7 +21,7 @@ Internal adversarial pass, not an independent review. Every sentence in this fil |---|---|---|---|---|---|---| | Q1 | Line-index distribution over 2^22 lines, per round and pooled; segments; depth | `adv-cache-2 lines`, 16 days x 2^24 items = 2^28 derivations, 2^31 line reads; 256 days (2^32) queued | `--plant quarter-lines`: segments +24.75 sigma, FLAGGED. `--plant half-lines`: +11.67 sigma, FLAGGED. Both fired | Largest segment bucket within 6 sigma per round and pooled; depth flat; top 1 percent of lines under 1 percent plus chance | Pooled 16 days: segments max +4.84 sigma (control +4.24), min -4.25; lines max +5.61 (control +5.35), chi2/dof 0.99937; depth max +1.82 sigma; per round segments max +4.70; top 0.1 / 1 percent of lines 0.11519 / 1.11979 percent against the control's 0.11516 / 1.11958 (1.0003x / 1.0002x); 0 mirror mismatches on 1/64 of the items | PASS (BOUND at 2^35 reads: 256 days, section 1.1) | | Q2a | Hot set of items and lines across the hashes of the two real programs | `adv-cache-2 warps`: item table 2^24 with 8 lines each, the mirrored interpreter with the shadow block; devnet3 at 2^24 nonces (2^31 reads), devnet at 2^26 (2^33 reads); per site against the site's own window | `--plant const-item` (site 0 fed one item): hot set FLAGGED, site 0 outside its window on every read, hottest item 6.25 percent of reads. Fired | Top 0.1 percent of items under 1.2x the window-model control and X_f < f; lines the same; per site the same | Devnet 3 at 2^24 nonces: items 1.0032x / 1.0021x of the control at 0.1 / 1 percent; lines 0.9999x. Devnet 3 at 2^26: items 1.0071x / 1.0048x, lines 1.0000x. Devnet at 2^26: items 1.0002x / 1.0001x; lines 1.0002x; verdict PASS. FINDING at one load site of Devnet 3: site 0 (instr 3, window the whole dataset) chi2/dof 1.674 at 2^24 nonces and 3.702 at 2^26 (a fixed per-item bias, since the excess grows with the reads), its top 0.1 / 1 percent carry 1.449x / 1.416x the control's share at 2^26, largest item 101 reads at a mean of 32; the other 15 sites 1.000x. Attributed (v3 diagnostic, 2^24 nonces per iteration): iteration 0 uniform (chi2/dof 0.9996), iterations 1 to 7 each at 1.110; the source register r6's top 16 bits at full entropy and never saturated; P(bit b of r6 = 1) for b = 0..7 reads 0.5000 0.5000 0.5002 ... in iteration 0 and 0.2501 0.3750 0.4375 0.4687 0.4844 0.4920 0.4961 0.4980 in iterations 1 to 7: exactly the low-bit law of a product of two uniform words (bit 0 is 1 with probability 1/4, bit 1 3/8, bit 2 7/16 ...). The static table says why: site 0's source r6 is last written in the BASE program by `sub` at instruction 0, but the shadow block, which runs between iterations, writes r6 last with a `mul`, so from iteration 1 on the load's source is a fresh product. The era stride `rotl(x * M, R)` then carries the biased low bits into the item index. The devnet's worst site is 1.011x (site 6, chi2/dof 1.039). Worth 0.1 percent of a hash's reads to a hottest-items store (section 4.1) | FINDING (small, priced, attributed to one site's source) and BOUND; v3 diagnostic queued (item 25) | -| Q2b | Hot set over drawn programs | 32 epochs under the devnet era (epoch seeds `igneum-adv-cache-2/epoch/k`) at 2^23 nonces each (2^30 reads per program); 32 drawn eras queued | the same plant | the same gates per program; the census line PASS/FAIL | All 32 devnet-era programs and 16 drawn-era programs done (section 2.3): overall every program clear on the hot-set test (items 0.9993x to 1.0075x of the windowed control at the top 0.1 percent, lines 0.9992x to 1.0006x; quarter and half shares the model to four digits). Four more sites of the Devnet 3 kind over 1.2x (era-fixed-20 site 11 1.26x, era-drawn-2 site 13 1.27x, era-drawn-10 site 7 1.23x, era-drawn-15 site 14 1.51x) and four at 1.04x to 1.16x in 48 x 16 = 768 sites; the same base programs under the devnet era (rotation 29) read 1.00x to 1.01x at those sites: the era rotation decides whether a product's biased low bits land inside the item index (section 2.3): era-fixed-20 site 11 (instr 46, r2, last shadow writer `mulhi`: one item read 512 times at a mean of 4, top 0.1 percent 1.26x, source hi16 bucket at 52x uniform) and era-drawn-2 site 13 (instr 55, r6, last shadow writer `sub` after a chain: chi2/dof 2.04, 1.27x). With Devnet 3's site 0 that is 3 sites in 36 programs (576 sites), 0.5 percent of sites, 8 percent of programs; each worth 0.03 to 0.1 percent of a hash's reads to a store (section 4.1) | FINDING (a class with its mechanism: 5 of 800 sites over 1.2x, 9 over 1.04x; a product in the source's last writes plus an era rotation under 28; worst 1.51x, worth 0.05 percent of a hash's reads) and BOUND; 16 drawn-era programs queued at 32 cores | +| Q2b | Hot set over drawn programs | 32 epochs under the devnet era (epoch seeds `igneum-adv-cache-2/epoch/k`) at 2^23 nonces each (2^30 reads per program); 32 drawn eras queued | the same plant | the same gates per program; the census line PASS/FAIL | All 32 devnet-era programs and 20 drawn-era programs done (section 2.3): overall every program clear on the hot-set test (items 0.9993x to 1.0075x of the windowed control at the top 0.1 percent, lines 0.9992x to 1.0006x; quarter and half shares the model to four digits). Four more sites of the Devnet 3 kind over 1.2x (era-fixed-20 site 11 1.26x, era-drawn-2 site 13 1.27x, era-drawn-10 site 7 1.23x, era-drawn-15 site 14 1.51x) and four at 1.04x to 1.16x in 48 x 16 = 768 sites; the same base programs under the devnet era (rotation 29) read 1.00x to 1.01x at those sites: the era rotation decides whether a product's biased low bits land inside the item index (section 2.3): era-fixed-20 site 11 (instr 46, r2, last shadow writer `mulhi`: one item read 512 times at a mean of 4, top 0.1 percent 1.26x, source hi16 bucket at 52x uniform) and era-drawn-2 site 13 (instr 55, r6, last shadow writer `sub` after a chain: chi2/dof 2.04, 1.27x). With Devnet 3's site 0 that is 3 sites in 36 programs (576 sites), 0.5 percent of sites, 8 percent of programs; each worth 0.03 to 0.1 percent of a hash's reads to a store (section 4.1) | FINDING (a class with its mechanism: 6 of 864 sites over 1.2x, 11 over 1.04x; sub-class A a product in the source's last writes whose low bits the era rotation leaves inside the item index, 30 of 31 rotations do; sub-class B a warp-uniform source once in 16,000 warps; worst 1.51x, worth 0.05 percent of a hash's reads) and BOUND; 12 drawn-era programs queued at 32 cores | | Q3(1) | Steering by the item index t | `adv-cache-2 steer`, 2^20 random t x 32 bit flips, rounds 0, 1, 7, days 20730 and 20733 | `--plant t-low` (round-0 address = t AND mask): identity rows at P = 1.0, dead rows for bits 22..31, 40,960 equal pairs at +231,704 sigma. Fired | every cell within 6 sigma of 0.5; equal pairs within 6 sigma of 8 | Day 20730: worst cell 3.44 sigma; equal pairs 13 / 4 / 6 of 2^25 against 8.00 (+1.77 / -1.41 / -0.71 sigma). Day 20733: worst 3.95 sigma; pairs 8 / 11 / 7. 2 x 2,112 cells, none beyond 4 sigma | PASS (BOUND at 2^20 items x 32 flips per day) | | Q3(2) | Steering by the day key: the weak-day scan | `adv-cache-2 days`, 16,384 consecutive chain days x 2^16 items (2^30 derivations), per day the segment and depth histograms and a same-size control | `--plant quarter-lines --plant-day 20733` on 8 days x 2^14 items: day 20733 listed alone at segment chi2_z +1,089.7 and max bucket 22 (+14.1 sigma) against the control's 11; the 7 clean days 10 to 11. Fired (19:23:49Z) | no day with segment chi2_z over 6 or a depth bucket beyond 6 sigma; the per-day max against the control's | 16,384 days (20730 to 37113) x 2^16 items = 2^30 derivations in 707 s on 32 cores; 0 mirror mismatches; worst per-day segment chi2_z +3.96 (day 36139); worst per-day max segment bucket 31 at mean 8 (+8.13 sigma, day 29452) against the control's worst 30 (+7.78); the per-day max-bucket histogram REAL 20:44 21:1797 22:5877 23:5156 24:2335 25:800 26:279 27:66 28:24 29:3 30:2 31:1 against CONTROL 20:30 21:1793 22:5974 23:5034 24:2335 25:825 26:285 27:72 28:29 29:6 30:1; worst depth z_max +4.96; 0 days flagged | PASS (BOUND: no weak day in 16,384, 45 years of chain days) | | Q3(3) | Steering by the parameters: the window layer's concentration | `adv-cache-2 windows`, 4,096 programs under the devnet era and 4,096 under drawn eras, the expected share per quarter from the 16 sites' window draws; measured on the real programs in Q2 | `--plant all-quarter` (every site forced to one quarter): top-quarter share 1.0000 on 8 programs. Fired | none: the distribution is the result (by design above uniform) | Real programs: devnet top quarter 0.42187 of reads (model 0.42188), top aligned half 0.53125; Devnet 3 top quarter 0.39063 (model 0.39062), top half 0.71875. Exact distribution from the draw: top quarter mean 0.3382, median 0.3281, p99 0.4688; top half mean 0.5811, p99 0.75. 4,096 drawn programs: mean 0.3377 / 0.5806, max 0.6250 / 0.8750 (section 3.3) | FINDING (designed, counted; priced in 4.1) | @@ -182,7 +182,16 @@ Programs 12 to 23 (the resumed census, killed by the coordinator at 19:49Z after The controlled comparison the drawn-era seeds give for free: `era-drawn-k` and `era-fixed-k` share the epoch seed, so they are the SAME base program (same id, same sites, same windows) under two era draws. Under the devnet era (stride multiplier 0x9ad30d99, rotation 29, interleave [0, 2, 12, 13]) the sites above read 1.0009x to 1.0112x; under their own eras (rotations 22, 4, 3, 12) the same sites read 1.06x to 1.51x. The mechanism, from `verify::load_index`: `y = rotl(x * M, R)`, then the 28-bit mask (and the window's top `k` bits replaced by the offset). An odd multiplier keeps bit 0 of `x` as bit 0 of the product and leaves the low bits of the product a function of the low bits of `x` alone, so a source register whose low bits are biased (the low-bit law of a product: P(bit 0) = 1/4, bit 1 3/8, bit 2 7/16 ...) has a biased product in the same bits; the rotation then moves those bits to positions R, R + 1, ... of the address. With R = 29 the three most biased bits land at 29, 30, 31, above the 28-bit dataset mask, and are cut; with R in 1..27 they land inside the item index and the bias shows at full strength. Devnet 3's era has R = 9 (its site 0 shows), the shared devnet's R = 29 (its programs are clean), and the four drawn eras with the biased sites have R = 3, 4, 12, 22. Under the 2 GiB genesis dataset (D = 29) R = 29 would land bit 0 inside the mask as well. What a site needs to show the bias: its source register's last writer before the load (counting the shadow block, which runs between iterations) must be a product, a high product or an `or`/`sub` chain on products, and the era rotation must keep the product's low bits inside the masked, un-windowed part of the address. -Five sites over 1.2x and four between 1.04x and 1.2x of 800 examined (50 programs: 2 real, 32 devnet-era, 16 drawn-era) concentrate their item reads beyond their window's uniform: Devnet 3 site 0 (shadow `mul` last writer, the product low-bit law), era-fixed-20 site 11 (shadow `mulhi` last writer; a high half of a product has a low-entropy tail when an operand is small, which shows as a 52x hi16 bucket and one item read 512 times, 16 whole warps), era-drawn-2 site 13 (shadow `sub` last writer with chi2/dof 2.04; its law is not read yet). The common shape: the shadow block's LAST write to the load's source register is a non-injecting or multiplicative op, outside what the acceptance rule's freshness check sees (it checks the base program's dataflow and the shadow's fixpoint, the sibling adv-accept lane's subject). Price to a store: the worst (era-drawn-2 site 13) hands a chip holding the top 1 percent of items 3.23 percent of that site's reads instead of 2.56, so 0.67 percent of 1/16 of a hash's reads: 0.04 percent. Frequency at this sample: 5 of 50 programs (10 percent) carry a site over 1.2x and 9 of 50 (18 percent) a site over 1.04x; among the 16 drawn-era programs it is 2 of 16 over 1.2x and 6 of 16 over 1.04x, against 1 of 32 and 1 of 32 under the devnet era (R = 29, which cuts the three most biased bits). Every flagged site has a product in the last writes of its source register and an era rotation that keeps the product's low bits inside the address. The 16 remaining drawn-era programs are queued. The 32 devnet-era programs are complete: 31 PASS, 1 (era-fixed-20) with the site-11 excess; top-quarter share over the 32: min 0.2656, max 0.5313, mean 0.3613 (the closed form's 0.3382 for a random program; these 32 are a sample of 32 with its own noise, sigma of the mean about 0.012). +| era-drawn-17 (1), era R = 11 | 0.32813 | 1.0003x | 1.0002x | **site 0 (instr 4, src r6, window a quarter): 1.2139x**, chi2/dof 1.72; under the devnet era (era-fixed-17) 1.0003x | windowed 6-sigma FLAGGED | +| era-drawn-18 (2), R = 29 | 0.32813 | 1.0002x | 0.9999x | site 5: 1.0013x | PASS | +| era-drawn-19 (0), R = 12 | 0.29688 | 1.0126x | 1.0096x | site 1 (instr 13, window a quarter): 1.0452x, chi2/dof 1.19 | PASS (hot set clear) | +| era-drawn-20 (1), R = 29 | 0.32811 | 1.0057x | 1.0007x | **site 11 (instr 46, src r2): 1.2234x**, one item read 528 times at a mean of 4: the same site as era-fixed-20's (also R = 29, 1.2624x, 512 reads of one item): sub-class B below | windowed 6-sigma FLAGGED | + +Correction to the rotation rule stated two paragraphs up, after era-drawn-20 and era-drawn-21 (R = 31, site 4 at chi2/dof 1.29): `rotl` by R sends bit b of the product to bit (b + R) mod 32, so the bits cut by the 28-bit mask are those landing in 28..31, that is bits 28 - R .. 31 - R of the product. R = 29 cuts bits 0, 1, 2 (the three most biased) and puts bit 3 (P = 0.469) at address bit 0; R = 30 cuts bits 0 and 1; R = 31 cuts bit 0 alone and puts bit 1 (P = 0.375) at address bit 0; every R from 1 to 27 keeps bit 0 (P = 0.25) inside, and R = 28 cuts bits 0 to 3. So the devnet era's R = 29 is one of the two rotations (29, 30) that remove the strongest bits; 30 of 31 rotations leave at least bit 1 inside the item index. The sub-class A frequency under a random era is therefore the 16-program drawn-era figure (6 of 16 over 1.04x), not the devnet-era figure. + +Sub-class B (era-fixed-20 and era-drawn-20, site 11, R = 29 in both): the bias is not the low-bit law (P(bit 0) 0.476) but a warp-uniform source: one item read 512 and 528 times at a mean of 4 is 16 whole warps whose 32 lanes carried the same value in r2 at that site (the source's top-16-bit bucket at 52x uniform), so the site's items repeat across lanes in about 1 warp in 16,000. The last shadow writer of r2 is `mulhi`; a high product of two words that are both small is 0 for all lanes, and the shfl exchanges then keep the lanes equal. The acceptance rule's lane-constant test (no load site reading one address in all 32 lanes of any of 64 units) samples 64 units and cannot see a 1-in-16,000 event; the chip's gain from it is one item per 16,000 warps, nothing. It is recorded because the same shape at a higher rate would be a lane-constant site the rule would not catch. + +Six sites over 1.2x and five between 1.04x and 1.2x of 864 examined (54 programs: 2 real, 32 devnet-era, 20 drawn-era) concentrate their item reads beyond their window's uniform: Devnet 3 site 0 (shadow `mul` last writer, the product low-bit law), era-fixed-20 site 11 (shadow `mulhi` last writer; a high half of a product has a low-entropy tail when an operand is small, which shows as a 52x hi16 bucket and one item read 512 times, 16 whole warps), era-drawn-2 site 13 (shadow `sub` last writer with chi2/dof 2.04; its law is not read yet). The common shape: the shadow block's LAST write to the load's source register is a non-injecting or multiplicative op, outside what the acceptance rule's freshness check sees (it checks the base program's dataflow and the shadow's fixpoint, the sibling adv-accept lane's subject). Price to a store: the worst (era-drawn-2 site 13) hands a chip holding the top 1 percent of items 3.23 percent of that site's reads instead of 2.56, so 0.67 percent of 1/16 of a hash's reads: 0.04 percent. Frequency at this sample: 6 of 54 programs (11 percent) carry a site over 1.2x and 11 of 54 (20 percent) a site over 1.04x; among the 20 drawn-era programs 4 of 20 over 1.2x and 8 of 20 over 1.04x, against 1 of 32 and 1 of 32 under the devnet era (R = 29, which cuts the three most biased bits). Every sub-class A site has a product in the last writes of its source register and an era rotation that keeps some of the product's low bits inside the address. The 12 remaining drawn-era programs are queued. The 32 devnet-era programs are complete: 31 PASS, 1 (era-fixed-20) with the site-11 excess; top-quarter share over the 32: min 0.2656, max 0.5313, mean 0.3613 (the closed form's 0.3382 for a random program; these 32 are a sample of 32 with its own noise, sigma of the mean about 0.012). ### 2.2 Shared devnet, epoch 0 (program `a785001687d8688a`, day 20730), 2^26 nonces @@ -324,13 +333,14 @@ The first load of a hash (site 0 in iteration 0) depends only on the nonce, the | Q2b era-drawn 2 to 4 (shard s0b, 29 cores) and 5 to 8 (s1, 2,551 s wall mostly waiting, about 8 min of run) | 2 | about 15 min of run | 29 to 32 | 0.08 | | Q1 shard s4 (lease, 32 cores) | 1 | 225 s | 32 | 0.02 | | Q2b era-drawn 9 to 16 (shards s2, s3; about 8 min of run each, the rest lease wait) | 2 | about 16 min of run | 32 | 0.09 | +| Q2b era-drawn 17 to 20 (shard s4, 176 s wall) | 2 | 176 s | 32 | 0.02 | | Pod-hours | none (no GPU, no rented pod) | | | 0 | Running total at 22:15Z: about 2.1 box-hours by wall x threads / 96 (the boxes ran at load 400 to 600 for the first two hours, so the CPU actually consumed is under that). The 8-hour reading is not near. The 8-hour reading is not near. ## 6. The bound reached, honestly -At 2^35 line reads over 272 day keys (Q1) the line index, the segment and the recompute depth are uniform to within the control's own fluctuations (largest excess +4.9 sigma on 2^22 bins against the control's +4.8; one per-day per-round segment at +6.6 among 1.3 x 10^8 buckets); the top 1 percent of lines carry 1.02963 percent of reads, the control 1.02959. Over 16,384 chain days (Q3(2)) no day's cache has a segment or depth skew the control does not also show. At 2^31 and 2^33 item reads on the two real programs (Q2) the item histogram is the window model to 1.003x at the top 0.1 percent; five load sites of 800 (10 percent of programs) are non-uniform at the item level (chi2/dof 1.7 to 3.7, 1.23x to 1.51x at the top 0.1 percent of their own reads), each attributed to a product in the last writes of the load's source register (the shadow block's, which the base program's freedom check does not see) landing its biased low bits inside the item index through an era rotation under 28; each is worth 0.03 to 0.1 percent of a hash's reads to a store. No bit of t steers the line index (Q3(1), 2^20 items x 32 flips x 2 days, worst cell 3.95 sigma). The one counted gain is the designed window layer (Q3(3)): an aligned half of the dataset serves 53 to 78 percent of reads, so the chip model's partial-store rows overstate the recompute share by up to 1.8x; it does not change the model's verdict that the full-store chip is the cheapest. Not searched: the mixer itself (the adv-mixer lane), the chain function (adv-cache), a dataset over 2^28 words, any GPU behaviour. +At 2^35 line reads over 272 day keys (Q1) the line index, the segment and the recompute depth are uniform to within the control's own fluctuations (largest excess +4.9 sigma on 2^22 bins against the control's +4.8; one per-day per-round segment at +6.6 among 1.3 x 10^8 buckets); the top 1 percent of lines carry 1.02963 percent of reads, the control 1.02959. Over 16,384 chain days (Q3(2)) no day's cache has a segment or depth skew the control does not also show. At 2^31 and 2^33 item reads on the two real programs (Q2) the item histogram is the window model to 1.003x at the top 0.1 percent; five load sites of 800 (10 percent of programs) are non-uniform at the item level (chi2/dof 1.7 to 3.7, 1.23x to 1.51x at the top 0.1 percent of their own reads), each attributed to a product in the last writes of the load's source register (the shadow block's, which the base program's freedom check does not see) landing its biased low bits inside the item index through the era rotation (every rotation but 29 and 30 leaves bit 0 or bit 1 inside); each is worth 0.03 to 0.1 percent of a hash's reads to a store. No bit of t steers the line index (Q3(1), 2^20 items x 32 flips x 2 days, worst cell 3.95 sigma). The one counted gain is the designed window layer (Q3(3)): an aligned half of the dataset serves 53 to 78 percent of reads, so the chip model's partial-store rows overstate the recompute share by up to 1.8x; it does not change the model's verdict that the full-store chip is the cheapest. Not searched: the mixer itself (the adv-mixer lane), the chain function (adv-cache), a dataset over 2^28 words, any GPU behaviour. ## 7. What a longer pass would add