adv-accept report: widened rows at 21 programs, random base rate 1 of 10 beyond, 256-unit selector weak (5 of 11 false positives)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 19:19:41 +00:00
parent 0bccf4b83a
commit 839109598c

View file

@ -292,21 +292,39 @@ sweep's ratio column already locates (rows under 0.9810 at 256 units, then a 2^2
| 3664 (lowest of 16,337) | 0.9945 | 1.311x BEYOND | +0.074% (0.74) | clear |
| 103378 | 0.9960 | 1.157x within | +0.028% (0.28) | clear |
| 105756 | 0.9960 | 0.9996x within | -0.000% (0.00) | clear |
| 6610 | 0.9960 | 1.062x within | +0.010% (0.10) | clear |
| 106474 | 0.9960 | 0.9996x within | -0.000% (0.00) | clear |
| 4765 | 0.9960 | 1.020x within | +0.003% (0.03) | clear |
| 4107 (random) | 0.9999 | 1.0004x within | +0.000% | clear |
| 101710 (random) | 0.9998 | 1.0034x within | +0.001% | clear |
| 101886 (random) | 0.9998 | 1.0000x within | +0.000% | clear |
| 6132 (random) | 0.9999 | 0.9996x within | -0.000% | clear |
| (17 lowest and 16 random rows land as the runs end, 20:0x BST) | | | | |
| 105915 (random, the lowest ratio of the random 20) | 0.9994 | 1.399x BEYOND | +0.065% (0.65) | clear |
| 5133 (random) | 0.9999 | 1.0215x within | +0.004% | clear |
| 1932 (random) | 0.9999 | 1.0024x within | +0.000% | clear |
| 2328 (random) | 0.9999 | 1.0025x within | +0.000% | clear |
| 3665 (random) | 0.9998 | 1.0002x within | +0.000% | clear |
| 101643 (random) | 0.9999 | 1.0017x within | +0.000% | clear |
| (17 lowest and 10 random rows land as the runs end, 20:2x BST) | | | | |
Partial reading at 20:05 BST, honestly: the selector has false positives. Of the 8 lowest-ratio seeds
measured live so far (the five first confirmed plus 3664, 103378, 105756), 6 are beyond the f8 1.2x gate and
2 are not (103378 at 1.157x, 105756 at 0.9996x, the latter indistinguishable from a random program). The
random control stands at 0 of 4 beyond. So the stand-in ratio at 256 units is a noisy selector: the five
seeds confirmed first sat at 0.9986 to 0.9988 and all were beyond; the next three sit LOWER (0.9945 to 0.9960)
and only one is beyond, which says the 256-unit read is itself noisy at the 0.996 level (one standard
deviation of a 65,536-evaluation distinct count is about 0.001 in the ratio), and a steering attacker would
use the 2^20 read (2.8 s per candidate) rather than the 256-unit one. The false-positive rate with its error
is the number this table gives at 25 rows.
Base rate, corrected at 20:20 BST: at 2^24 nonces the f8 gate fires on random accepted programs too. Of the
first 10 random seeds, 1 (105915) is beyond 1.2x at 1.40x, and it is the one with the lowest stand-in ratio
of the random set (0.9994 against 0.9998 to 0.9999 for the other nine). So "beyond 1.2x at 2^24" has a base
rate near 1 in 10 among accepted programs (the 10^6-nonce census read 0 of 47 beyond because the gate is
sharper at 2^24), and the selector's 6 of 8 is an enrichment of about 6x over that base rate, not a clean
separation. The hot-set test proper (X_f >= f) has fired on 1 of 18 programs measured at 2^24 (100767).
Reading at 20:20 BST, honestly: the 256-unit selector is weak. Of the 11 lowest-ratio seeds measured live
at 2^24 (the five first confirmed, then 3664, 103378, 105756, 6610, 106474, 4765), 6 are beyond the f8 1.2x
gate and 5 are not; the random control is 1 of 10 beyond. The six that are beyond are the five confirmed
first (256-unit ratios 0.9986 to 0.9988, 2^20 ratios 0.9832 to 0.9919) and 3664; the five that are not sit
LOWER on the 256-unit read (0.9945 to 0.9960), which says the 256-unit read is noise-dominated at that level
(one standard deviation of a 65,536-evaluation distinct count is about 0.001 in the ratio) and the real
selector is the 2^20 ratio the rule itself computes (2.8 s per candidate: 100064's 0.9832, 100211's 0.9907
and 100767's 0.9919 are genuinely low; 3664's 2^20 ratio is in the queued gap read). So: selector enrichment
about 5x over the base rate (55 percent beyond against 10 percent), false-positive rate 5 of 11 on the
256-unit read; the error on both is binomial at n = 11 and n = 10, about plus or minus 15 points. The
hot-set test proper (X_f >= f) fired on 1 of 21 programs at 2^24, and on none of the random 10.
### Class v5 check of the exemplar (PENDING the lock; the defender's question)