259 lines
25 KiB
Markdown
259 lines
25 KiB
Markdown
# F8. Uniformity censuses of the class v4 derivation
|
|
|
|
Attack-pass row F8 (`docs/plans/cryptanalysis.md` section 4.2; the pass record `docs/analysis/attack-pass-2026-10.md`).
|
|
Run 7 October 2026, 08:13 to 09:3x UTC (09:13 to 10:3x UK) on igneum-build-1. Verdict: **FINDING** (AP-F8-1 below).
|
|
The line-index census is a PASS at its full sample size; the cross-hash item histogram is not uniform, and the cause
|
|
is in the base program, inside the acceptance rule's blind spot.
|
|
|
|
## 1. Target
|
|
|
|
| Item | Value |
|
|
|---|---|
|
|
| Commit | `igneum-pow` at 924288d1 (`attack-pass`); the box built HEAD b2a411d1, whose `igneum-pow` is byte-identical (`git diff --stat 924288d1 HEAD -- igneum-pow` is empty) |
|
|
| Class | `--program-class v4`: `V4_CLASS` = `mx8+sh256x27`, generator 4, mixer x8, the era layout drawn inside the class, the shadow block of 256 instructions x 27 reps |
|
|
| Line index | `proto-metal/MEMHARD.md` section 1.6: `a = s[0] AND 0x003fffff`, 4,194,304 lines of 64 B, 8 dependent reads per item (`memhard.rs` `derive_items_mask`, `cache.line_const(s[0])`) |
|
|
| Item index | `verify.rs` `load_index`: `y = rotl(x * M, R)`, the site's window `(y & (MASK >> k)) \| off`, then `Layout::split` removes the four interleave bits; 2^24 items at the 2^28-word dataset |
|
|
| Reads per hash | 128 loads (16 sites x 8 iterations), so up to 1,024 cache lines per hash and 32,768 per warp. The spec's analytic bound of 832 lines per hash (`docs/spec/01-lottery-hash.md` line 347, 104 loads x 8) predates generator 2's fixed 16 load slots; the current bound is 128 x 8 = 1,024 |
|
|
| Prior figures | `chip-model-v3.md` section 1: "median 128.00 distinct" items per hash (the 20,000-program census); `weak-program-census-2026-10-03.md` line 291: 127.7 distinct addresses per hash under the proposed generator |
|
|
| Day | the devnet pack's day, `bind::day_bytes(20730)` (2026-10-04), day 0 of the growth schedule: a 2^26-word cache, a 2^28-word dataset. Census 1 uses days 20730 to 20745 |
|
|
| Programs | p1 = the devnet epoch-0 derivation (epoch seed and era seed both the genesis hash `edc4fa84...fb07`, program id `c120d7963abdcd96`, attempt 0); p2 and p3 = chain-shaped seeds from tag strings (section 4), attempts 1 and 0 |
|
|
|
|
## 2. Method and harness
|
|
|
|
Harness: `tools/attack/f8-uniform/` (crate `attack-f8`, a path dependency on `igneum-pow`, nothing in the library
|
|
modified). Built on the box through `tools/build-remote.sh`: sha256 `590668...f913` for the firings of 2.1,
|
|
`ef8042...11b2` for sections 3, 4.1 and the first item runs (flat null, `log/c-*`, `log/d-*`), `890955...fbd3` for the
|
|
window-model runs and the seed census (`log/e-*`). The committed source carries one later label fix (the
|
|
"uniform-on-window" entropy reference in the per-site line is 16 - k_off bits; the 09:04 UTC logs print 16 - 2 k_off). Box scratch `/srv/builds/igneum-wt-attack/target-attack-f8/`
|
|
(the name `target-*` is what the box's checkout clean spared at the time; the fix at b92a5fd4 now also spares
|
|
`attack-*`). Every run: `nice -n 10 taskset -c 22-27,70-75`, 12 threads, under `flock -s /srv/builds/_locks/measure`
|
|
in chunks under 3 minutes each (the longest, phase D, under 25 minutes).
|
|
|
|
Two mirrors, each trusted only while it agrees with the library bit for bit:
|
|
|
|
| Mirror | What it records | Agreement check | Result |
|
|
|---|---|---|---|
|
|
| `derive_traced`: `memhard::derive_items_mask` instruction for instruction (`mixer`, `round_key_mult`, `cache.line` from the library), the line index of every round kept | 8 line indices per item | every item of every day also derived by the library's `derive_items` and compared on all 16 words | 0 mismatches on 268,435,456 items (section 3) and on 16,777,216 items per table build (section 4) |
|
|
| `Mirror::warp`: `verify::interpret_warp_init` for the class v4 op set, dataset words from a table of the day's 2^24 items, the item index and the source register of every load kept | 128 item indices per lane, the source value's saturation per position | the 32 hashes of warp 0 to 63 and of every 997th warp compared with `Epoch::hash_warp` | 0 mismatches on 95 warps per program (section 4) |
|
|
|
|
Three censuses:
|
|
|
|
1. `lines`: all 2^24 items of each of 16 consecutive day keys (2^28 item derivations, 2^31 line reads), the full 2^22-line
|
|
histogram per round and pooled, the 2^16-bucket histogram (64 lines, one chained segment per bucket), a uniform
|
|
SplitMix64 control of the same size.
|
|
2. `warps`: 10^6 nonces (31,250 warps) of each of three programs: distinct lines and items per hash and per warp, the
|
|
cross-hash item histogram, per-position diagnostics, an attribution pass from the hottest items back to the load
|
|
positions that read them.
|
|
3. `warps` at 2^26 nonces on p1: the one-epoch cross-hash item histogram at 512 expected reads per item.
|
|
|
|
The tests, defined before the runs:
|
|
|
|
- **6-sigma test**: the largest (and smallest) bucket of a histogram within 6 sigma of its expectation, sigma =
|
|
sqrt(expectation). The gate's bucket is the 64-line segment for lines and the 64-item bucket for items. The
|
|
full-resolution histograms are reported beside a uniform control of the same size, because at a small mean the
|
|
Poisson tail puts the maximum of 4 million bins above 6 sigma by chance (control at mean 8: +6.72 sigma; at mean
|
|
32: +5.83; at mean 512: +5.61).
|
|
- **Hot-set test** (F8's definition, written for F9's reuse): sort items by read count; S_f = the share of all reads
|
|
on the top-f fraction of items, for f in {0.1%, 0.5%, 1%}; E_f = the same share on a control of the same size drawn
|
|
from the design's own null (flat uniform for lines; the window-weighted null for items, section 4.2); the excess
|
|
X_f = S_f - E_f. **A hot set exists at f when X_f >= f**: after the chance excess is removed, the top f of items
|
|
capture at least one extra proportional share, which is what an on-die copy of f of the items would have to win to
|
|
matter. X_f / f is printed as the gain in proportional shares. The acceptance-style form of the same metric (for
|
|
rule (c)'s 2,048 evaluations): per load position, the largest count of one masked address, and the count of
|
|
saturated (0 or 2^32 - 1) source values.
|
|
|
|
### 2.1 The harness fires (known-fail and known-pass)
|
|
|
|
| Plant | What it does | 6-sigma test | Hot-set test | Log |
|
|
|---|---|---|---|---|
|
|
| `quarter-lines` | line index masked to a quarter of its range | buckets64 largest +75.97 sigma (2,231 at mean 512), smallest -22.63: FLAGGED | X_1% = +3.54% (S 5.60% vs control 2.06%), X/f = 3.5 at every f: FLAGGED | `log/a1-lines-quarter.log` |
|
|
| `half-lines` | line index masked to a half | buckets64 largest +29.26 sigma: FLAGGED | X_1% = +1.25%, X/f = 1.25: FLAGGED | `log/a2-lines-half.log` |
|
|
| `const-item` | one constant item at the first load site (1/16 of reads) | items buckets64 largest +92,682 sigma: FLAGGED | X_0.1% = +6.38%, X/f = 63.8: FLAGGED | `log/a4-warps-const-item.log` |
|
|
| none, 2^22 items, one day | the real derivation at a small size | buckets64 largest +4.42 sigma, smallest -4.51: within 6 sigma (control +4.33) | X_f = -0.0004%, -0.0006%, -0.0010%: clear | `log/a3-lines-pass-small.log` |
|
|
|
|
Both tests fire on every plant and neither fires on the real line derivation. Log paths are under
|
|
`/srv/builds/igneum-wt-attack/target-attack-f8/`.
|
|
|
|
## 3. Census 1: the line index over 2^28 derivations (PASS)
|
|
|
|
Sample reached: 16 days x 2^24 items = 268,435,456 item derivations, 2,147,483,648 line reads into 4,194,304 lines
|
|
(512 expected per line, 32,768 per 64-line segment). Mirror mismatches against `derive_items`: 0 of 268,435,456.
|
|
Log: `log/b-lines-16days.log`; histograms `out/lines-d20730-n16-i24-none-buckets64.txt` (65,536 rows) and
|
|
`out/lines-d20730-n16-i24-none-full.u32le` (4,194,304 x u32).
|
|
|
|
| Histogram | Bins | Expected | Largest | Sigma | Smallest | Sigma | chi2/dof | Top 1% share |
|
|
|---|---|---|---|---|---|---|---|---|
|
|
| Pooled, 64-line buckets (the gate) | 65,536 | 32,768 | 33,645 | +4.84 | 31,998 | -4.25 | 1.00226 | 1.01427% |
|
|
| Pooled, full 2^22 lines | 4,194,304 | 512 | 639 | +5.61 | 402 | -4.86 | 0.99937 | 1.11979% |
|
|
| Control, 64-line buckets | 65,536 | 32,768 | 33,524 | +4.18 | 31,960 | -4.46 | 0.99930 | 1.01426% |
|
|
| Control, full 2^22 lines | 4,194,304 | 512 | 639 | +5.61 | 408 | -4.60 | 0.99952 | 1.11956% |
|
|
| Per round 0 to 7, full, pooled (64 per line) | 4,194,304 | 64 | 107 to 113 | +5.38 to +6.12 | 26 to 29 | -4.75 to -4.38 | 0.99855 to 1.00062 | 1.3483% to 1.3488% |
|
|
| One day (20730), 64-line buckets | 65,536 | 2,048 | 2,291 | +5.37 | 1,859 | -4.18 | 0.99985 | 1.05826% |
|
|
| One day, control, 64-line buckets | 65,536 | 2,048 | 2,250 | +4.46 | 1,874 | -3.84 | 1.00377 | 1.05901% |
|
|
|
|
Per day, the gate bucket's largest value ran +4.00 to +5.37 sigma on all 16 days (control +4.46), every day within
|
|
6 sigma. Hot-set test on the pooled lines: X_0.1% = -0.00002%, X_0.5% = +0.00010%, X_1% = +0.00023% (X/f under
|
|
0.0003): clear. Round 0, whose input is the sequential item index through the init `t * MUL[i] + RC[i]` and eight
|
|
mixer applications, is as flat as rounds 1 to 7 (chi2/dof 0.99926; its +5.50 sigma maximum is below the control's
|
|
+5.61 at the pooled size). Round 5's +6.12 sigma at mean 64 is one bin of 4 million at a Poisson tail where the
|
|
control at mean 8 reached +6.72; its chi2/dof is 0.99855.
|
|
|
|
Gate line: the largest bucket is within 6 sigma of uniform (+4.84 on the 64-line buckets, +5.61 on the full 2^22
|
|
lines, both at or below the control), chi2/dof 0.99937, no hot set. **PASS at 2^28 derivations.**
|
|
|
|
## 4. Census 2 and 3: distinct lines per hash and warp, and the cross-hash item histogram
|
|
|
|
Setup per program: the day's 16,777,216 items derived once into a table with their 8 lines (6 to 9 s on 12 threads,
|
|
0 mismatches against `derive_items` on every item), then the warps interpreted from the table at 2.7 to 3.2 ms per
|
|
warp per thread. Logs: `log/c-warps-p{1,2,3}-1e6.log` (first run, flat null) and `log/e-warps-p{1,2,3}-1e6.log`
|
|
(windowed null, section 4.2); distributions `out/warps-<program>-d20730-n1000000-none-distinct.txt`, item histograms
|
|
`...-items.u32le` (16,777,216 x u32), per-position tables `...-positions.txt`.
|
|
|
|
### 4.1 Distinct lines and items per hash and per warp (10^6 nonces each)
|
|
|
|
| Program | Epoch seed / era seed | Lines per hash min / p1 / median / max / mean | Items per hash min / median / mean | Lines per warp min / median / max / mean | Items per warp min / median / mean |
|
|
|---|---|---|---|---|---|
|
|
| p1 `c120d7963abdcd96` (devnet epoch 0) | genesis / genesis | 1,008 / 1,023 / 1,024 / 1,024 / 1,023.867 | 126 / 128 / 127.9989 | 32,579 / 32,636 / 32,680 / 32,635.84 | 4,090 / 4,096 / 4,095.41 |
|
|
| p2 `82f0696f823e9c65` | `59cef1aa...bfdfa` / `9cba001f...1f69` | 1,014 / 1,023 / 1,024 / 1,024 / 1,023.871 | 127 / 128 / 127.9995 | 32,580 / 32,637 / 32,684 / 32,636.08 | 4,091 / 4,096 / 4,095.48 |
|
|
| p3 `e282eed7d47e425e` | `c54e2ddd...c95d` / `1b04f607...b58a` | 999 / 1,016 / 1,024 / 1,024 / 1,023.600 | 125 / 128 / 127.9656 | 32,276 / 32,481 / 32,601 / 32,480.32 | 4,050 / 4,076 / 4,075.85 |
|
|
| Uniform expectation | | 1,023.875 of 1,024 | 127.9995 of 128 | 32,640.3 of 32,768 | 4,095.50 of 4,096 |
|
|
|
|
Per hash, every program reads its 128 items and 1,024 lines as the design intends (p1 and p2 at the uniform
|
|
expectation; p3 a shade under, 127.97 items, which is the same site-15 effect as the finding below: the saturated
|
|
site repeats an item inside a hash 3 times in 100). Per warp, 32 lanes read 32,636 distinct lines of 2^22, a 2 MiB
|
|
working set of cache lines and 256 KiB of dataset items, within 0.01% of uniform on p1 and p2.
|
|
|
|
### 4.2 The cross-hash item histogram and the window layer
|
|
|
|
The era layout's window layer (`docs/plans/era-layout.md` section 1.4, layer 8) makes each load site read an
|
|
aligned half or quarter of the dataset with probability 2/3. The per-site item distribution is therefore not flat by
|
|
design (the diagnostic's "worst bit" reads P(1) = 1.0000 or 0.0000 at every windowed site: the fixed top bits), and
|
|
the summed item histogram has density steps between quarters. For p1 the 16 windows (site:shrink:offset
|
|
`7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1`) give
|
|
expected reads per item by quarter of 3.25 : 2.25 : 5.75 : 4.75 in sixteenths of the flat value. Against a flat
|
|
uniform the 64-item buckets of p2 (a program without the finding) read +10.03 and -9.06 sigma, which is the window
|
|
layer and not a flaw. The item tests are therefore judged against the **window-weighted null**: the expected count of
|
|
every item from the program's 16 windows, and a control that draws each read from a uniformly chosen site's window.
|
|
A chip gains nothing from the window steps: the union of the windows is the whole dataset every hour (era-layout.md
|
|
section 7), the floor window is 2^26 words (256 MiB), and which quarter is dense changes with the program.
|
|
|
|
#### The window model (reproducible by the firms)
|
|
|
|
For load site s with window draw `(k_s, o_s)` at the 2^28-word dataset: `k = min(k_s, 28 - 26)`, the word window is
|
|
`[o_s << (28 - k), (o_s + 1) << (28 - k))`; the item window is `[o_s << (24 - k), (o_s + 1) << (24 - k))` of
|
|
`2^(24 - k)` items (the four interleave positions all lie below bit 16, so the top bits of the word index are the top
|
|
bits of the item index). The expected reads per item is `E[t] = sum over sites s with t in window_s of N x 8 / 2^(24 - k_s)`
|
|
for N nonces (8 iterations per site), a density constant on each quarter of the item space. The windowed control draws
|
|
each of the N x 128 reads as (site = read index mod 16, item uniform on that site's window). Both controls are drawn from
|
|
SplitMix64 with a fixed seed. The tests on items are run against E[t] (chi-square, sigma of the largest and smallest
|
|
64-item bucket) and against the windowed control (the top-f shares); the flat uniform numbers are kept beside them as
|
|
what an auditor sees first.
|
|
|
|
#### Results, 10^6 nonces per program, 128,000,000 reads (`log/e-warps-p{1,2,3}-1e6.log`)
|
|
|
|
| Program | Windows (k_off:offset per site) | Quarter densities (reads per item) | Buckets64 largest sigma, windowed (control) | chi2/dof windowed (control) | Top 0.1% share: real / window control / flat control | Ratio to window control at 0.1% (gate 1.2x) | Ratio to flat control | Hot set (X_f >= f) |
|
|
|---|---|---|---|---|---|---|---|---|
|
|
| p1 devnet epoch 0 | 2:1 1:1 1:1 1:1 0 1:1 0 2:2 1:1 1:1 2:0 0 0 0 2:0 1:1 | 6.20 / 4.29 / 10.97 / 9.06 | +45.77 (+4.95) | 1.2336 (0.9981) | 0.5458% / 0.2891% / 0.2429% | 1.888x BEYOND | 2.247x | yes at 0.1% (X/f 2.57) and 0.5% (1.20); not at 1% (0.46) |
|
|
| p2 | 2:0 1:0 0 0 2:3 2:2 1:0 0 2:3 0 0 0 0 1:1 2:0 0 | 9.54 / 5.72 / 6.68 / 8.58 | +4.59 (+4.64) | 1.0061 (0.9986) | 0.2716% / 0.2639% / 0.2429% | 1.029x within | 1.118x | no (X/f 0.08, 0.05, 0.05) |
|
|
| p3 | 1:0 0 0 2:2 0 0 1:0 1:0 1:0 2:2 2:2 1:1 0 0 0 0 | 7.63 / 7.63 / 10.49 / 4.77 | +12,245.66 (+5.06) | 907.67 (0.9993) | 4.5954% / 0.2792% / 0.2433% | 16.46x BEYOND | 18.92x | yes at every f (X/f 43.2, 9.9, 5.0) |
|
|
|
|
p2 is what the class is designed to be: against the window model its largest bucket is +4.59 sigma (the control +4.64),
|
|
chi2/dof 1.006, the top 0.1% of items hold 1.029x their window-model share, and the flat-control ratio of 1.118x is
|
|
the window layer. p1 and p3 are the finding (section 5). The one-epoch histogram at 2^26 nonces (8,589,934,592 reads,
|
|
512 per item, `log/d-warps-p1-2e26.log`, flat null): p1's top 0.1% hold 0.5199% of reads against 0.1152% flat
|
|
control (X/f 4.05), the top 1% 2.4946% against 1.1198% (X/f 1.37), item 0xca5b92 78,479 reads at a mean of 512, and
|
|
site 15 feeds 6.37% of its reads into the top 0.1% in each of the 8 iterations; the excess grows with N as the
|
|
control's chance excess shrinks, which is the signature of a structural skew. Distinct lines and items per hash and
|
|
per warp at 2^26 nonces: 1,023.866 / 127.9989 / 32,635.6 / 4,095.41, unchanged from 10^6.
|
|
|
|
## 5. AP-F8-1: a saturated load source makes a cross-hash hot set (FINDING)
|
|
|
|
**What**: an accepted class v4 program can read one load site from a register whose last writes after its last
|
|
injecting write are `or` (and, mildly, `mul`), so the site's address has fewer than 32 bits of entropy across nonces
|
|
and the same items are read by many hashes. The per-hash figures (128 distinct items, 1,024 lines) stay intact; the
|
|
cross-hash item histogram does not. It is not the window layer (p2 shows the window layer alone is clean against its
|
|
model) and not the shadow block (iteration 0's load, which runs before any shadow block, is as hot as iterations 1 to
|
|
7: p1 6.372% vs 6.371% to 6.378%; p3 71.9% vs 72.4% to 72.6%).
|
|
|
|
**Where it hides from rule (c)** (`accept.rs`, 2,048 evaluations of the base program): the tests are constant bits
|
|
in FINAL register values, one address in ALL 32 lanes of a unit, saturated FINAL values, output-bit bias, and distinct
|
|
addresses WITHIN a hash. A site whose address is concentrated across hashes but refreshed before the end of the
|
|
iteration passes every one. Rule (a) accepts any write, `or` included, as the refresh between two loads from the same
|
|
register (`check_stale_loads`); `Op::injects` (add, sub, xor, mad, shfl, load) is only used by rule (b), once per
|
|
register per program.
|
|
|
|
**The index derivation at the hot site** (the "writers back to the last injecting one" lines of `log/e-warps-p*.log`):
|
|
|
|
| Program | Hot site | Source | Writes after the last injecting write | Site's reads into the top 0.1% of items (flat expectation) | Index entropy, 256-item buckets (uniform on window) | Saturated source (x = 0 or 2^32 - 1) | Most repeated address at one position in 2,048 evaluations (uniform: 1 to 2) |
|
|
|---|---|---|---|---|---|---|---|
|
|
| p3 | site 15, instr 62 | r5 | `load@17` then `or@19`, `or@30` | 72.43% (0.10%) | 13.411 bits (16) | 1.368% | 44 of 2,048; 32 saturated |
|
|
| p1 | site 15, instr 63 | r6 | `add@51` then `rotl@53`, `or@61` | 6.93% (0.11%) | 14.985 bits (15) | 0.005% | 2 of 2,048; 0 saturated |
|
|
| p2 (clean) | every site | | injecting, or bijective (`rotl`), or `mul`/`mulhi` | 0.41% to 0.97% (0.20%; the window densities) | 13.999 / 14.997 / 15.994 bits (14 / 15 / 16) | 0.000% | 2 of 2,048; 0 saturated |
|
|
|
|
In p3 two `or`s on r5 after its load make the source 1 with probability 7/8 per bit; x = 2^32 - 1 in 1.37% of
|
|
evaluations and the images of the near-saturated values under the stride (`y = rotl(x * M, R)`, 256 x-values per
|
|
item) pile onto a few items: 0xffdf69 takes 213,913 of the site's 8,000,000 reads (2.67%), the top 0.1% of items
|
|
72.4%, and 4.6% of ALL reads of the hash land on 0.1% of the items. In p1 one `or` after `rotl(add)` gives 3/4 per bit
|
|
on the ORed positions: no saturation to speak of (0.005%), but 6.9% of the site's reads on 0.11% of the items (the
|
|
hot items share the low 20 bits `5b92`: 0xca5b92, 0x8a5b92, 0xaa5b92, 0xba5b92, 0x825b92, 0xe65b92, 0x985b92), a
|
|
2.6x proportional excess at f = 0.1%. p2's `mul` sites (10, 13: `mul` after a load or a shuffle) read 0.65% and 0.70%
|
|
into the top 0.1% against 0.41% and 0.48% for their window class (an even multiplier zeroes low bits; the hot items
|
|
0xd6a680, 0xe44400, 0xd25600 end in zero bits), a mild effect that the window-model ratio (1.029x) absorbs.
|
|
|
|
**How common** (the seed census, 64 chain-shaped programs p4 to p67, 262,144 nonces each, `log/e-seed-census-4-67.log`,
|
|
`out/seed-census-d20730-n262144-p4-67.txt`): CENSUS-LINE
|
|
|
|
**Reproduction**: `attack-f8 warps --program 3 --nonces 1000000 --diag 1` (or `--program 1`); the acceptance-style
|
|
numbers come from the same run's "acceptance-style" line. The program is `Epoch::chain_program(epoch_seed, Some(era),
|
|
ProgramClass::V4, label)` with the seeds of section 4.1.
|
|
|
|
**Proposed fix** (not applied; `igneum-pow` untouched, the Counter ASIC lane re-gates on `ca3-v4-uniform` with this
|
|
harness):
|
|
|
|
1. Rule (a'), static: between the last injecting write of a load's source register and the load (cyclically), no
|
|
`or` and no `mul` writes that register; `rotl`, `rotr` and `mulhi` may (bijective, or measured flat: p2 site 15
|
|
reads `mulhi` after `add` at 1.00x). This rejects p1 and p3 at draw time and costs nothing at run time. Programs
|
|
rejected are redrawn as today (`MAX_ATTEMPTS` 32); the census gives the rejection rate.
|
|
2. Rule (c'), dynamic, the same 2,048 evaluations: no load site reads a saturated source (0 or 2^32 - 1) in more
|
|
than 2 evaluations, and no address repeats more than 4 times at one position (uniform expectation 1 to 2; p3 shows
|
|
44 and 32). This catches the strong class only; p1's class needs about 2^16 evaluations to show at a site (65,536
|
|
nonces: largest item count 67 at a mean of 0.5), so (a') is the rule that closes it and (c') is the check that
|
|
fails loudly if (a') is ever loosened.
|
|
3. Packs re-cut for the seeds the new rule rejects (the devnet epoch-0 program p1 is one of them: its site 15 is
|
|
`or@61`), with the gate pack ids re-pinned; the chain's own epochs redraw automatically.
|
|
|
|
**Reuse for F9**: the hot-set metric (section 2) on the per-program item histogram at 2^18 nonces, and the
|
|
acceptance-style pair (most repeated address at a position, saturated sources at a position) at 2,048 evaluations,
|
|
are both emitted by `warps --programs a..b`; a header-grinding search that steers a program to a hot set would show as
|
|
ratio-to-window-model above 1.2x at f = 0.1%.
|
|
|
|
## 6. Consequences per tier
|
|
|
|
| Number | What it means | Per tier |
|
|
|---|---|---|
|
|
| Line index uniform at 2^28 derivations (largest segment +4.84 sigma, chi2/dof 0.99937) | the 256 MiB cache has no hot segment: a chip or a card cannot serve the 8 dependent reads of an item from a cache smaller than the whole 256 MiB (the floor window of era-layout.md) | no change for any card; the verifier's cache stays 256 MiB in RAM on every node |
|
|
| Distinct lines per hash 1,023.87 of 1,024, items 127.999 of 128 (p1, p2); per warp 32,636 lines, 4,095 items | the per-hash working set is 64 KiB of cache lines and 8 KiB of items, per warp 2 MiB of lines and 256 KiB of items; the item-derivation chip's "128 items per hash" input (`chip-model-v3.md`) stands | the 8 GB card and up: unchanged; the recompute chip pays 128 derivations per hash, as modelled |
|
|
| p3-class programs: 4.6% of all dataset reads on 0.1% of items (1 MiB of a 1 GiB dataset); p1-class: 0.59% on 0.11% | a stored-dataset chip with 1 MiB of on-die SRAM serves 4.6% of its reads without touching DRAM on such an epoch; a GPU's L2 (96 MiB on the 5090, 64 MB Infinity Cache on the 9070 XT, vendor figures) holds the same 1 MiB, so both sides gain the same 4.6% of reads and the chip's edge from it is about 0 (the per-joule edge of `evidence.md` row 17 is a DRAM-read figure; a 4.6% read saving on both sides moves it by under 5% on such epochs). The recompute chip (f = 0, SRAM cache) caches the derived hot items and skips up to 4.6% of its 128 derivations per hash on such epochs, a 4.8% rate gain on those epochs only | home cards 8 to 32 GB, rigs, pools: no action; a few percent of epochs run a few percent faster for everyone with an L2. The verifier: `MemhardCpu::fetch` dedupes within a fetch only, so no change. The chip model: the headline 2.1x at k = 1 moves by under 5% on affected epochs and 0 on others; the fix below returns it to 0 everywhere |
|
|
| The acceptance rule's blind spot (cross-hash concentration at one site) | a program class property, not a day or era property: the same seed is hot on every day and under every era, so a chip or a pool that selects epochs cannot gain more than the epoch's own 4.6%; but the public claim "the item map is uniform per program up to the window layer" is false for the affected fraction of seeds until rule (a') lands | the fix is a generator rule plus packs re-cut: a class change under the 95% signalling rule if it lands after the flip, a plain re-cut if it lands in the class v4 cut itself (the lane's call) |
|
|
|
|
## 7. Gate line and verdict
|
|
|
|
| Gate (plan 4.2 F8, the same as 1.4 (4)) | Result | Status |
|
|
|---|---|---|
|
|
| The largest bucket within 6 sigma of uniform on the stated sample sizes (line index, 2^28 derivations) | +4.84 sigma on 64-line segments, +5.61 on 2^22 lines (control +4.18 / +5.61), chi2/dof 0.99937 | PASS |
|
|
| The item distribution within 6 sigma of uniform (against the window model, the design's own null) | p2 +4.59 sigma (control +4.64); p1 +45.77; p3 +12,245.66 | FAIL on p1 and p3 |
|
|
| No hot set under 1% of items among passing seeds (10^6 nonces on three programs; the 64-seed census at 2^18) | p2 none; p1 top 0.1% at 1.888x the window model (2.247x flat), X/f 2.57; p3 16.46x (18.92x flat), X/f 43.2; census: 31 of 64 seeds over 1.2x of the window model, 23 with a hot item | FAIL |
|
|
| The Counter ASIC lane's record gate: top 0.1% within 1.2x of the window-model control on every seed | p2 1.029x; p1 1.888x; p3 16.46x; census: 31 of 64 seeds over 1.2x (median 1.17x, p90 2.70x, max 13.09x on p31) | FAIL |
|
|
|
|
**Verdict: FINDING (AP-F8-1).** The line index passes at 2^28 derivations. The cross-hash item histogram fails
|
|
the hot-set gate on 2 of the 3 named programs (one of them the live devnet epoch-0 program) and on 31 of 64 seeds (48 percent) of of
|
|
the 64-seed census, from `or` (and mildly `mul`) writes on a load's source register after its last injecting write,
|
|
outside every test of rule (c). What it moves: not the mask or the fold (the derivation is uniform) but the
|
|
acceptance rule, (a') and (c') above, and the packs re-cut. Ownership: the Counter ASIC lane (generator and rule),
|
|
re-gated with this harness on the fixed branch; the row reads FIXED-AND-PASSED when every seed of the census passes
|
|
both the hot-set test and the 1.2x gate under the new rule.
|
|
|
|
Sample sizes reached: 2^28 derivations (lines); 10^6 nonces on three programs (distinct lines, hot set); one epoch at
|
|
2^26 nonces (cross-hash histogram); 64 seeds at 2^18 nonces (the census).
|
|
|
|
Times UTC in the logs; the runs ran 08:13 to 09:2x UTC on 7 October 2026 (09:13 to 10:2x UK).
|