adv-cache-2: shard s3b, era-fixed 24 to 27 (internal adversarial pass, not an independent review)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
2157f553cc
commit
e4170b626f
1 changed files with 12 additions and 5 deletions
|
|
@ -21,7 +21,7 @@ Internal adversarial pass, not an independent review. Every sentence in this fil
|
|||
|---|---|---|---|---|---|---|
|
||||
| Q1 | Line-index distribution over 2^22 lines, per round and pooled; segments; depth | `adv-cache-2 lines`, 16 days x 2^24 items = 2^28 derivations, 2^31 line reads; 256 days (2^32) queued | `--plant quarter-lines`: segments +24.75 sigma, FLAGGED. `--plant half-lines`: +11.67 sigma, FLAGGED. Both fired | Largest segment bucket within 6 sigma per round and pooled; depth flat; top 1 percent of lines under 1 percent plus chance | Pooled 16 days: segments max +4.84 sigma (control +4.24), min -4.25; lines max +5.61 (control +5.35), chi2/dof 0.99937; depth max +1.82 sigma; per round segments max +4.70; top 0.1 / 1 percent of lines 0.11519 / 1.11979 percent against the control's 0.11516 / 1.11958 (1.0003x / 1.0002x); 0 mirror mismatches on 1/64 of the items | PASS (BOUND at 2^35 reads: 256 days, section 1.1) |
|
||||
| Q2a | Hot set of items and lines across the hashes of the two real programs | `adv-cache-2 warps`: item table 2^24 with 8 lines each, the mirrored interpreter with the shadow block; devnet3 at 2^24 nonces (2^31 reads), devnet at 2^26 (2^33 reads); per site against the site's own window | `--plant const-item` (site 0 fed one item): hot set FLAGGED, site 0 outside its window on every read, hottest item 6.25 percent of reads. Fired | Top 0.1 percent of items under 1.2x the window-model control and X_f < f; lines the same; per site the same | Devnet 3 at 2^24 nonces: items 1.0032x / 1.0021x of the control at 0.1 / 1 percent; lines 0.9999x. Devnet 3 at 2^26: items 1.0071x / 1.0048x, lines 1.0000x. Devnet at 2^26: items 1.0002x / 1.0001x; lines 1.0002x; verdict PASS. FINDING at one load site of Devnet 3: site 0 (instr 3, window the whole dataset) chi2/dof 1.674 at 2^24 nonces and 3.702 at 2^26 (a fixed per-item bias, since the excess grows with the reads), its top 0.1 / 1 percent carry 1.449x / 1.416x the control's share at 2^26, largest item 101 reads at a mean of 32; the other 15 sites 1.000x. Attributed (v3 diagnostic, 2^24 nonces per iteration): iteration 0 uniform (chi2/dof 0.9996), iterations 1 to 7 each at 1.110; the source register r6's top 16 bits at full entropy and never saturated; P(bit b of r6 = 1) for b = 0..7 reads 0.5000 0.5000 0.5002 ... in iteration 0 and 0.2501 0.3750 0.4375 0.4687 0.4844 0.4920 0.4961 0.4980 in iterations 1 to 7: exactly the low-bit law of a product of two uniform words (bit 0 is 1 with probability 1/4, bit 1 3/8, bit 2 7/16 ...). The static table says why: site 0's source r6 is last written in the BASE program by `sub` at instruction 0, but the shadow block, which runs between iterations, writes r6 last with a `mul`, so from iteration 1 on the load's source is a fresh product. The era stride `rotl(x * M, R)` then carries the biased low bits into the item index. The devnet's worst site is 1.011x (site 6, chi2/dof 1.039). Worth 0.1 percent of a hash's reads to a hottest-items store (section 4.1) | FINDING (small, priced, attributed to one site's source) and BOUND; v3 diagnostic queued (item 25) |
|
||||
| Q2b | Hot set over drawn programs | 32 epochs under the devnet era (epoch seeds `igneum-adv-cache-2/epoch/k`) at 2^23 nonces each (2^30 reads per program); 32 drawn eras queued | the same plant | the same gates per program; the census line PASS/FAIL | 23 of 32 devnet-era programs and 2 drawn-era programs done (section 2.3): overall every program clear on the hot-set test (items 0.9993x to 1.0075x of the windowed control at the top 0.1 percent, lines 0.9992x to 1.0006x; quarter and half shares the model to four digits). Two more sites of the Devnet 3 kind in 25 x 16 = 400 sites: era-fixed-20 site 11 (instr 46, r2, last shadow writer `mulhi`: one item read 512 times at a mean of 4, top 0.1 percent 1.26x, source hi16 bucket at 52x uniform) and era-drawn-2 site 13 (instr 55, r6, last shadow writer `sub` after a chain: chi2/dof 2.04, 1.27x). With Devnet 3's site 0 that is 3 sites in 27 programs (432 sites), 0.7 percent of sites, 11 percent of programs; each worth 0.03 to 0.1 percent of a hash's reads to a store (section 4.1) | FINDING (a class: 3 of 432 sites; small, priced) and BOUND; 9 + 30 programs re-queued at 32 cores |
|
||||
| Q2b | Hot set over drawn programs | 32 epochs under the devnet era (epoch seeds `igneum-adv-cache-2/epoch/k`) at 2^23 nonces each (2^30 reads per program); 32 drawn eras queued | the same plant | the same gates per program; the census line PASS/FAIL | 27 of 32 devnet-era programs and 2 drawn-era programs done (section 2.3): overall every program clear on the hot-set test (items 0.9993x to 1.0075x of the windowed control at the top 0.1 percent, lines 0.9992x to 1.0006x; quarter and half shares the model to four digits). Two more sites of the Devnet 3 kind in 29 x 16 = 464 sites: era-fixed-20 site 11 (instr 46, r2, last shadow writer `mulhi`: one item read 512 times at a mean of 4, top 0.1 percent 1.26x, source hi16 bucket at 52x uniform) and era-drawn-2 site 13 (instr 55, r6, last shadow writer `sub` after a chain: chi2/dof 2.04, 1.27x). With Devnet 3's site 0 that is 3 sites in 31 programs (496 sites), 0.6 percent of sites, 10 percent of programs; each worth 0.03 to 0.1 percent of a hash's reads to a store (section 4.1) | FINDING (a class: 3 of 496 sites; small, priced) and BOUND; 5 + 30 programs queued at 32 cores |
|
||||
| Q3(1) | Steering by the item index t | `adv-cache-2 steer`, 2^20 random t x 32 bit flips, rounds 0, 1, 7, days 20730 and 20733 | `--plant t-low` (round-0 address = t AND mask): identity rows at P = 1.0, dead rows for bits 22..31, 40,960 equal pairs at +231,704 sigma. Fired | every cell within 6 sigma of 0.5; equal pairs within 6 sigma of 8 | Day 20730: worst cell 3.44 sigma; equal pairs 13 / 4 / 6 of 2^25 against 8.00 (+1.77 / -1.41 / -0.71 sigma). Day 20733: worst 3.95 sigma; pairs 8 / 11 / 7. 2 x 2,112 cells, none beyond 4 sigma | PASS (BOUND at 2^20 items x 32 flips per day) |
|
||||
| Q3(2) | Steering by the day key: the weak-day scan | `adv-cache-2 days`, 16,384 consecutive chain days x 2^16 items (2^30 derivations), per day the segment and depth histograms and a same-size control | `--plant quarter-lines --plant-day 20733` on 8 days x 2^14 items: day 20733 listed alone at segment chi2_z +1,089.7 and max bucket 22 (+14.1 sigma) against the control's 11; the 7 clean days 10 to 11. Fired (19:23:49Z) | no day with segment chi2_z over 6 or a depth bucket beyond 6 sigma; the per-day max against the control's | 16,384 days (20730 to 37113) x 2^16 items = 2^30 derivations in 707 s on 32 cores; 0 mirror mismatches; worst per-day segment chi2_z +3.96 (day 36139); worst per-day max segment bucket 31 at mean 8 (+8.13 sigma, day 29452) against the control's worst 30 (+7.78); the per-day max-bucket histogram REAL 20:44 21:1797 22:5877 23:5156 24:2335 25:800 26:279 27:66 28:24 29:3 30:2 31:1 against CONTROL 20:30 21:1793 22:5974 23:5034 24:2335 25:825 26:285 27:72 28:29 29:6 30:1; worst depth z_max +4.96; 0 days flagged | PASS (BOUND: no weak day in 16,384, 45 years of chain days) |
|
||||
| Q3(3) | Steering by the parameters: the window layer's concentration | `adv-cache-2 windows`, 4,096 programs under the devnet era and 4,096 under drawn eras, the expected share per quarter from the 16 sites' window draws; measured on the real programs in Q2 | `--plant all-quarter` (every site forced to one quarter): top-quarter share 1.0000 on 8 programs. Fired | none: the distribution is the result (by design above uniform) | Real programs: devnet top quarter 0.42187 of reads (model 0.42188), top aligned half 0.53125; Devnet 3 top quarter 0.39063 (model 0.39062), top half 0.71875. Exact distribution from the draw: top quarter mean 0.3382, median 0.3281, p99 0.4688; top half mean 0.5811, p99 0.75. 4,096 drawn programs: mean 0.3377 / 0.5806, max 0.6250 / 0.8750 (section 3.3) | FINDING (designed, counted; priced in 4.1) |
|
||||
|
|
@ -81,7 +81,8 @@ The binary's verdict line says FLAGGED on one per-day, per-round test: day 20823
|
|||
| Shard | Days | Reads | Lines: largest / smallest (sigma), chi2/dof | Segments: largest / smallest | Top 1 percent of lines vs control | Verdict |
|
||||
|---|---|---|---|---|---|---|
|
||||
| s0 (20:24:04Z to 20:31:19Z, 261 s of run on 27 to 32 cores) | 20746 to 20809 (a duplicate of 64 days inside the 256-day census above: an oversight in the shard plan, kept as an independent re-read) | 2^33 | +5.26 / -4.97, 0.99924 | +4.06 / -4.55 | 1.05937 vs 1.05938 percent (1.0000x) | PASS |
|
||||
| s2b, s3b, s4 | 21002 to 21193 (fresh) | 3 x 2^33 | queued | | | |
|
||||
| s3b (20:33:27Z to 20:37:11Z, 224 s on 27 cores) | 21066 to 21129 | 2^33 | +4.91 / -5.06, 0.99921 | +4.78 / -3.78 | 1.05942 vs 1.05937 (1.0000x) | PASS |
|
||||
| s4, s2c | 21130 to 21193, 21002 to 21065 | 2 x 2^33 | queued (s2c re-added after my own pkill at 20:33:26Z caught the shard the drain had just started; a wrong kill of mine, 1 s lost) | | | |
|
||||
|
||||
## 2. Q2: the hot set across the hashes of an epoch
|
||||
|
||||
|
|
@ -150,10 +151,14 @@ Programs 12 to 23 (the resumed census, killed by the coordinator at 19:49Z after
|
|||
| era-fixed-21 (0) | 0.32813 | 1.0018x | 1.0003x | under 1.1x | PASS |
|
||||
| era-fixed-22 (4) | 0.35938 | 0.9998x | 0.9999x | under 1.1x | PASS |
|
||||
| era-fixed-23 (0) | 0.37499 | 0.9994x | 0.9999x | under 1.1x | PASS |
|
||||
| era-fixed-24 (0) | 0.31250 | 1.0002x | 1.0004x | site 4: 1.0088x | PASS |
|
||||
| era-fixed-25 (6) | 0.40625 | 1.0000x | 0.9997x | site 5: 1.0014x | PASS |
|
||||
| era-fixed-26 (0) | 0.37501 | 1.0007x | 1.0001x | site 1: 1.0036x | PASS |
|
||||
| era-fixed-27 (7) | 0.32812 | 0.9994x | 0.9999x | site 4: 1.0025x | PASS |
|
||||
| era-drawn-1 (4) | 0.28126 | 1.0007x | 1.0000x | under 1.1x | PASS |
|
||||
| era-drawn-2 (0) | 0.32814 | 1.0075x | 1.0003x | **site 13 (instr 55, src r6): 1.2657x** at the top 0.1 percent, 1.2637x at 1 percent, chi2/dof 2.038 (the Devnet 3 shape); last base writer `sub`, last shadow writer `sub` | hot set clear |
|
||||
|
||||
Three sites of 432 examined (27 programs: 2 real, 23 devnet-era, 2 drawn-era) concentrate their item reads beyond their window's uniform: Devnet 3 site 0 (shadow `mul` last writer, the product low-bit law), era-fixed-20 site 11 (shadow `mulhi` last writer; a high half of a product has a low-entropy tail when an operand is small, which shows as a 52x hi16 bucket and one item read 512 times, 16 whole warps), era-drawn-2 site 13 (shadow `sub` last writer with chi2/dof 2.04; its law is not read yet). The common shape: the shadow block's LAST write to the load's source register is a non-injecting or multiplicative op, outside what the acceptance rule's freshness check sees (it checks the base program's dataflow and the shadow's fixpoint, the sibling adv-accept lane's subject). Price to a store: the worst (era-drawn-2 site 13) hands a chip holding the top 1 percent of items 3.23 percent of that site's reads instead of 2.56, so 0.67 percent of 1/16 of a hash's reads: 0.04 percent. Frequency at this sample: 11 percent of programs carry one such site; the remaining 39 programs are re-queued.
|
||||
Three sites of 496 examined (31 programs: 2 real, 27 devnet-era, 2 drawn-era) concentrate their item reads beyond their window's uniform: Devnet 3 site 0 (shadow `mul` last writer, the product low-bit law), era-fixed-20 site 11 (shadow `mulhi` last writer; a high half of a product has a low-entropy tail when an operand is small, which shows as a 52x hi16 bucket and one item read 512 times, 16 whole warps), era-drawn-2 site 13 (shadow `sub` last writer with chi2/dof 2.04; its law is not read yet). The common shape: the shadow block's LAST write to the load's source register is a non-injecting or multiplicative op, outside what the acceptance rule's freshness check sees (it checks the base program's dataflow and the shadow's fixpoint, the sibling adv-accept lane's subject). Price to a store: the worst (era-drawn-2 site 13) hands a chip holding the top 1 percent of items 3.23 percent of that site's reads instead of 2.56, so 0.67 percent of 1/16 of a hash's reads: 0.04 percent. Frequency at this sample: 3 of 31 programs (10 percent) carry one such site; the remaining 35 programs are queued.
|
||||
|
||||
### 2.2 Shared devnet, epoch 0 (program `a785001687d8688a`, day 20730), 2^26 nonces
|
||||
|
||||
|
|
@ -289,13 +294,15 @@ The first load of a hash (site 0 in iteration 0) depends only on the nonce, the
|
|||
| Q2b era-fixed 12 to 23 (lease, 40 cores, killed) | 2 | 416 s | 40 | 0.05 |
|
||||
| Q2b era-drawn 1 to 2 (lease, 40 cores, killed) | 2 | about 3 min | 40 | 0.02 |
|
||||
| Q1 shard s0 (lease, 27 to 32 cores) | 1 | 435 s | 32 | 0.04 |
|
||||
| Q1 shard s3b (lease, 27 cores) | 1 | 224 s | 27 | 0.02 |
|
||||
| Q2b era-fixed 24 to 27 (lease, 24 to 32 cores) | 2 | 800 s wall (about 6 min of run) | 32 | 0.07 |
|
||||
| Pod-hours | none (no GPU, no rented pod) | | | 0 |
|
||||
|
||||
Running total at 20:3xZ: about 1.7 box-hours by wall x threads / 96 (the boxes ran at load 400 to 600 for the first two hours, so the CPU actually consumed is under that). The 8-hour reading is not near. The 8-hour reading is not near.
|
||||
Running total at 20:40Z: about 1.8 box-hours by wall x threads / 96 (the boxes ran at load 400 to 600 for the first two hours, so the CPU actually consumed is under that). The 8-hour reading is not near. The 8-hour reading is not near.
|
||||
|
||||
## 6. The bound reached, honestly
|
||||
|
||||
At 2^35 line reads over 272 day keys (Q1) the line index, the segment and the recompute depth are uniform to within the control's own fluctuations (largest excess +4.9 sigma on 2^22 bins against the control's +4.8; one per-day per-round segment at +6.6 among 1.3 x 10^8 buckets); the top 1 percent of lines carry 1.02963 percent of reads, the control 1.02959. Over 16,384 chain days (Q3(2)) no day's cache has a segment or depth skew the control does not also show. At 2^31 and 2^33 item reads on the two real programs (Q2) the item histogram is the window model to 1.003x at the top 0.1 percent; three load sites of 432 (Devnet 3's site 0, era-fixed-20's site 11, era-drawn-2's site 13) are non-uniform at the item level (chi2/dof 1.08 to 3.70, 1.26x to 1.45x at the top 0.1 percent of their own reads), each attributed to the shadow block's last write to the load's source register being a product or a non-injecting op, and each worth 0.03 to 0.1 percent of a hash's reads to a store. No bit of t steers the line index (Q3(1), 2^20 items x 32 flips x 2 days, worst cell 3.95 sigma). The one counted gain is the designed window layer (Q3(3)): an aligned half of the dataset serves 53 to 78 percent of reads, so the chip model's partial-store rows overstate the recompute share by up to 1.8x; it does not change the model's verdict that the full-store chip is the cheapest. Not searched: the mixer itself (the adv-mixer lane), the chain function (adv-cache), a dataset over 2^28 words, any GPU behaviour.
|
||||
At 2^35 line reads over 272 day keys (Q1) the line index, the segment and the recompute depth are uniform to within the control's own fluctuations (largest excess +4.9 sigma on 2^22 bins against the control's +4.8; one per-day per-round segment at +6.6 among 1.3 x 10^8 buckets); the top 1 percent of lines carry 1.02963 percent of reads, the control 1.02959. Over 16,384 chain days (Q3(2)) no day's cache has a segment or depth skew the control does not also show. At 2^31 and 2^33 item reads on the two real programs (Q2) the item histogram is the window model to 1.003x at the top 0.1 percent; three load sites of 496 (Devnet 3's site 0, era-fixed-20's site 11, era-drawn-2's site 13) are non-uniform at the item level (chi2/dof 1.08 to 3.70, 1.26x to 1.45x at the top 0.1 percent of their own reads), each attributed to the shadow block's last write to the load's source register being a product or a non-injecting op, and each worth 0.03 to 0.1 percent of a hash's reads to a store. No bit of t steers the line index (Q3(1), 2^20 items x 32 flips x 2 days, worst cell 3.95 sigma). The one counted gain is the designed window layer (Q3(3)): an aligned half of the dataset serves 53 to 78 percent of reads, so the chip model's partial-store rows overstate the recompute share by up to 1.8x; it does not change the model's verdict that the full-store chip is the cheapest. Not searched: the mixer itself (the adv-mixer lane), the chain function (adv-cache), a dataset over 2^28 words, any GPU behaviour.
|
||||
|
||||
## 7. What a longer pass would add
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue