adv-accept report: widened rows at 21 programs, random base rate 1 of 10 beyond, 256-unit selector weak (5 of 11 false positives)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
0bccf4b83a
commit
839109598c
1 changed files with 28 additions and 10 deletions
|
|
@ -292,21 +292,39 @@ sweep's ratio column already locates (rows under 0.9810 at 256 units, then a 2^2
|
|||
| 3664 (lowest of 16,337) | 0.9945 | 1.311x BEYOND | +0.074% (0.74) | clear |
|
||||
| 103378 | 0.9960 | 1.157x within | +0.028% (0.28) | clear |
|
||||
| 105756 | 0.9960 | 0.9996x within | -0.000% (0.00) | clear |
|
||||
| 6610 | 0.9960 | 1.062x within | +0.010% (0.10) | clear |
|
||||
| 106474 | 0.9960 | 0.9996x within | -0.000% (0.00) | clear |
|
||||
| 4765 | 0.9960 | 1.020x within | +0.003% (0.03) | clear |
|
||||
| 4107 (random) | 0.9999 | 1.0004x within | +0.000% | clear |
|
||||
| 101710 (random) | 0.9998 | 1.0034x within | +0.001% | clear |
|
||||
| 101886 (random) | 0.9998 | 1.0000x within | +0.000% | clear |
|
||||
| 6132 (random) | 0.9999 | 0.9996x within | -0.000% | clear |
|
||||
| (17 lowest and 16 random rows land as the runs end, 20:0x BST) | | | | |
|
||||
| 105915 (random, the lowest ratio of the random 20) | 0.9994 | 1.399x BEYOND | +0.065% (0.65) | clear |
|
||||
| 5133 (random) | 0.9999 | 1.0215x within | +0.004% | clear |
|
||||
| 1932 (random) | 0.9999 | 1.0024x within | +0.000% | clear |
|
||||
| 2328 (random) | 0.9999 | 1.0025x within | +0.000% | clear |
|
||||
| 3665 (random) | 0.9998 | 1.0002x within | +0.000% | clear |
|
||||
| 101643 (random) | 0.9999 | 1.0017x within | +0.000% | clear |
|
||||
| (17 lowest and 10 random rows land as the runs end, 20:2x BST) | | | | |
|
||||
|
||||
Partial reading at 20:05 BST, honestly: the selector has false positives. Of the 8 lowest-ratio seeds
|
||||
measured live so far (the five first confirmed plus 3664, 103378, 105756), 6 are beyond the f8 1.2x gate and
|
||||
2 are not (103378 at 1.157x, 105756 at 0.9996x, the latter indistinguishable from a random program). The
|
||||
random control stands at 0 of 4 beyond. So the stand-in ratio at 256 units is a noisy selector: the five
|
||||
seeds confirmed first sat at 0.9986 to 0.9988 and all were beyond; the next three sit LOWER (0.9945 to 0.9960)
|
||||
and only one is beyond, which says the 256-unit read is itself noisy at the 0.996 level (one standard
|
||||
deviation of a 65,536-evaluation distinct count is about 0.001 in the ratio), and a steering attacker would
|
||||
use the 2^20 read (2.8 s per candidate) rather than the 256-unit one. The false-positive rate with its error
|
||||
is the number this table gives at 25 rows.
|
||||
Base rate, corrected at 20:20 BST: at 2^24 nonces the f8 gate fires on random accepted programs too. Of the
|
||||
first 10 random seeds, 1 (105915) is beyond 1.2x at 1.40x, and it is the one with the lowest stand-in ratio
|
||||
of the random set (0.9994 against 0.9998 to 0.9999 for the other nine). So "beyond 1.2x at 2^24" has a base
|
||||
rate near 1 in 10 among accepted programs (the 10^6-nonce census read 0 of 47 beyond because the gate is
|
||||
sharper at 2^24), and the selector's 6 of 8 is an enrichment of about 6x over that base rate, not a clean
|
||||
separation. The hot-set test proper (X_f >= f) has fired on 1 of 18 programs measured at 2^24 (100767).
|
||||
|
||||
Reading at 20:20 BST, honestly: the 256-unit selector is weak. Of the 11 lowest-ratio seeds measured live
|
||||
at 2^24 (the five first confirmed, then 3664, 103378, 105756, 6610, 106474, 4765), 6 are beyond the f8 1.2x
|
||||
gate and 5 are not; the random control is 1 of 10 beyond. The six that are beyond are the five confirmed
|
||||
first (256-unit ratios 0.9986 to 0.9988, 2^20 ratios 0.9832 to 0.9919) and 3664; the five that are not sit
|
||||
LOWER on the 256-unit read (0.9945 to 0.9960), which says the 256-unit read is noise-dominated at that level
|
||||
(one standard deviation of a 65,536-evaluation distinct count is about 0.001 in the ratio) and the real
|
||||
selector is the 2^20 ratio the rule itself computes (2.8 s per candidate: 100064's 0.9832, 100211's 0.9907
|
||||
and 100767's 0.9919 are genuinely low; 3664's 2^20 ratio is in the queued gap read). So: selector enrichment
|
||||
about 5x over the base rate (55 percent beyond against 10 percent), false-positive rate 5 of 11 on the
|
||||
256-unit read; the error on both is binomial at n = 11 and n = 10, about plus or minus 15 points. The
|
||||
hot-set test proper (X_f >= f) fired on 1 of 21 programs at 2^24, and on none of the random 10.
|
||||
|
||||
### Class v5 check of the exemplar (PENDING the lock; the defender's question)
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue