igneum/docs/analysis/ca3-v4-uniform.md
igneum-labs 7eed16a29a Pre-public scrub, the text pass (7 October 2026, 19:5x UK): no founder name, personal login, earlier business or personal address in any tracked text file, and a gate check that keeps it so
The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh).

The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge.

Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 18:39:50 +00:00

11 KiB

AP-F8-1 under the windows-union model: the hot set is the load source, not the window (7 October 2026)

Branch ca3-v4-uniform from master b92a5fd4, worker "v4-hash", on the attack-pass finding AP-F8-1 (docs/analysis/attack-pass/f8-uniform.md, branch attack-pass; logs /srv/builds/igneum-wt-attack/target-attack-f8/log/). Main's rulings bound this file: no generator change to class v4 on the live devnet; the analysis and its harness only. Every GPU-free number here is arithmetic on F8's logged counts or a run of the static census tool tools/ca3-v4-uniform/ on igneum-build-1 (built through tools/build-remote.sh, rule R1); the chip figures are the terms of docs/analysis/chip-model-v3.md and are approximate.

1. The null F8's numbers must be read against

Layer 8 (docs/plans/era-layout.md 1.4, spec 01 1.13.1 as proposed) gives every load site a window draw k_off = below(3): the site reads the whole dataset, an aligned half or an aligned quarter, at a 2^26-word floor. A quarter-window site concentrates its reads 4x on its quarter and a half-window site 2x on its half, by design; the union of the 16 windows is the whole dataset. For p1 (the devnet epoch-0 program, id c120d7963abdcd96) the 16 draws 0:2:1 1:1:1 2:1:1 3:1:1 4:0:0 5:1:1 6:0:0 7:2:2 8:1:1 9:1:1 10:2:0 11:0:0 12:0:0 13:0:0 14:2:0 15:1:1 (site:k:offset) give an expected read density by quarter of 3.25 : 2.25 : 5.75 : 4.75 sixteenths of the flat mean, which at 2^26 nonces (512 reads per item flat) is 416, 288, 736 and 608 reads per item. The top-f share of a Poisson mixture with those means (uniform-model.txt, exact Poisson for p1, a normal approximation for the census):

Share of all reads on the top f of items, 2^26 nonces Flat Poisson (F8's control) Windows-union model, p1 F8 measured, p1 Beyond the window model
f = 0.1 percent 0.115 0.160 0.520 +0.36
f = 0.5 percent 0.565 0.784 1.515 +0.73
f = 1 percent 1.120 1.553 2.495 +0.94

So the window model moves the null from 0.115 to 0.160 percent at f = 0.1 percent (1.39x, not F8's 4.05x) and from 1.12 to 1.55 at f = 1 percent; it explains the 64-item-bucket sigma of p2 that F8 already attributed to the window layer, and it explains every per-site attribution row of p1 except one: sites with a half window over the hot region land 0.20 percent of their reads in the top 0.1 percent (sites 1, 2, 3, 5, 9 at 0.201), quarter-window sites 0 or about 0.4 (sites 0, 10, 14 at 0.000, site 7 at 0.241 straddling), whole-dataset sites 0.10. Site 15 lands 6.374 percent. The excess over the model (+0.36 at f = 0.1 percent) is one site.

2. The 153x item is the load source, and the model predicts it to the item

p1's site 15 is the load at instruction 63 (load dst=3 src=6, window half 1). Its source r6 was last written at instruction 61: or dst=6 src=4 (r6 |= r4, verify.rs Op::Or), after a fresh dataset load into r6 at 47. An OR of two near-uniform registers sets each bit with probability 3/4, so the source takes the all-ones value with probability (3/4)^32 = 1.0e-4 per read and the values of popcount 31, 30, ... with 32, 496, ... times (3/4)^k (1/4)^(32 - k). The era map y = rotl(x * 0x9ad30d99, 29), the half window and the interleave split (memhard::Layout::split, positions 0, 2, 12, 13) send x = 0xffffffff to item 0xca5b92: F8's hottest item exactly. F8's next seven items (0x8a5b92, 0xaa5b92, 0xba5b92, 0x825b92, 0xe65b92, 0x985b92, 0xbcc392) are exactly the seven one-zero-bit sources whose zero bit survives the window mask (bits 29, 28, 27, 26, 25, 24 and 15): 7 of 7. The measured count fixes the bit bias: 78,479 reads of 2^26 x 8 site-15 reads is p^32 at p = 0.7585 (r4 is slightly biased itself), and at that p the popcount model predicts 77,348 all-ones reads and 4.86 percent of site 15's reads into the top 0.1 percent of items (measured 6.37; 3.92 at p = 3/4). Per hash that is 4.86 / 16 = 0.30 percent of all reads, and 0.160 + 0.30 = 0.46 against F8's 0.520 at f = 0.1 percent; at f = 1 percent 14.5 / 16 = 0.90, and 1.55 + 0.90 = 2.46 against 2.495.

The same arithmetic for the other lossy writers (uniform-model.txt): a mul last writer zeroes the low bits by the operands' trailing zeros, so 1.07 percent of the site's reads land on the 0.1 percent of values with 10 or more trailing zeros (p2's mul-sourced sites 0 and 14 measured 0.971 and 0.966 percent); a mulhi last writer is dense near zero, 0.79 percent on the lowest 0.1 percent of values. An or whose operand was itself last written by or compounds the bias (3/4 to 7/8 to 15/16): p3's site 15 (or at 30, the load at 62) puts 72.4 percent of its reads into the top 0.1 percent, 4.6 percent of all reads on 16,777 items.

This is a fault class, not the window model: the acceptance rule's part (a) (accept.rs check_stale_loads) takes any write as a fresh source, and part (c)'s saturation count looks at the 16,384 final register values, not at a load's source mid-program, so an or, mul or mulhi as a load's last writer passes. The per-hash distinct-address check still holds (p1 127.999 items per hash; p3 127.97: a saturated site repeats its item inside a hash), and the acceptance rule's floor of 120 distinct of 128 admits exactly one site repeating its item in all 8 iterations and no more.

3. How common it is: the static census (tools/ca3-v4-uniform, 1,024 chain-shaped class v4 programs plus F8's p1 to p3)

For every load site, the op that last wrote its source in execution order (base instructions before it, else the shadow block of the previous iteration, else the base instructions after it): injecting (add, sub, xor, mad, shfl, load), bijective (rotl, rotr) or lossy (or, mul, mulhi). Run on igneum-build-1 (uniform-census.txt, binary sha256 ce9f83fe... then the narrowed chain rule).

Census over 1,024 programs Count Share
Load sites by last writer: injecting / bijective / lossy 11,368 / 2,121 / 2,943 of 16,432 69 / 13 / 18 percent; 2.87 lossy sites per program
Programs with at least one lossy-sourced load 992 96.6 percent
... with an or-sourced load (p1's class, 0.30 percent of all reads per site) 498 48.5 percent
... with an or-of-or chain (p3's class, about 4.5 percent of all reads per site) 50 4.9 percent
... with a mul-sourced load (0.067 percent per site) / a mulhi-sourced load (0.049) 751 / 661 73.1 / 64.4 percent
Predicted S_0.1 percent (window model plus the lossy sites): median / 90th / 99th / max 0.45 / 0.88 / 5.29 / 9.82 percent against the window model's 0.115 to 0.251
p1 / p2 / p3 predicted against F8 measured 0.579 / 0.323 / 4.72 0.520 / 0.272 / 4.60

F8's proposed gate (the top 0.1 percent within 1.2x of the window-model control on every one of 64 seeds) fails 96.6 percent of today's programs, because any lossy-sourced site alone exceeds it (0.16 + 0.05 at the least); it is a generator change in a gate's clothing. A 2x bound fails 69.7 percent, 3x 48.9 percent; a bound of S_0.1 percent at or under 1 percent of all reads fails 6.9 percent (the or chains and the multi-or programs). The static rule "no load whose source's last writer is or" fails 48.4 percent; "no lossy last writer" 96.6 percent.

4. What the skew is worth to a chip (chip-model-v3.md terms, approximate)

A hot-set cache of the top 0.1 percent of items is 16,777 items x 64 B = 1.07 MB of SRAM, 0.53 mm^2 and $0.25 at 0.49 mm^2 and $0.23 per MB. It serves 0.52 percent of p1's reads (0.16 of them the window model's), 4.6 percent of p3's. The hash is latency-bound on its dependent reads, so a read served on die is time saved: a chip gains at most 1.005x on p1 and 1.048x on p3 from the cache. The ceiling under the live rule: part (c)'s 120-of-128 floor admits one site repeating its item in all 8 iterations and no more (two saturated sites fail it), so at most 8 of 128 reads, 6.25 percent, can sit on a constant item, and a chip's edge from this whole class is at most 1 / (1 - 0.0625) = 1.067x, in 64 bytes of SRAM, on the hours whose program carries such a site. The public claim rests on 2x margins (chip-model-v3.md); 1.067x does not move it, and the union of the windows is still the whole dataset every hour, so no window-level cache exists. What moves: per tier nothing in rate or watts (the honest card reads the hot item from L2 as the chip would), and the 5 percent rule of 2.0 is untouched.

5. The two options for the flip, priced (main's ruling 3; nothing ships on this without the founder's word)

Option What changes Cost Risk
A. A class amendment in 0.3.19 before the flip: the generator draws a load's source from the registers whose last writer injects (or rule (a) tightened to the same), class v4 re-pinned a new program stream: new vectors, the seven gate packs re-exported, the six gates again (the hash side G1 to G3 and the verifier re-run here in about an hour of Mac and PC 2 time; G4 to G6 the node lane), every node before the flip by the one-box-at-a-time fleet rule hours of gate time, a fleet rollout, the 0.3.19 ship on the line a node that misses the build splits the chain at the flip; the fix itself is small (one draw rule)
B. Hold v4 at the floor as it is; the source rule in class v5 nothing on the devnet; the attack-pass record carries the window null and the bound a hot set on 48 percent of hours worth up to 1.005x to a chip, on 5 percent of hours up to 1.05x, 1.067x at the rule's ceiling, no chain risk the public line must state the bound, not "uniform"

The number that decides it: 1.067x at the ceiling against the 2x margin of the chip claim. Recommendation: B, with the v5 item below, unless the founder wants the tail tight now.

6. The acceptance bound for the next class (main's ruling 4)

Definition: for a program, H = W_0.1(windows) + sum over load sites of h(last writer of the source), with W from the Poisson mixture of the 16 window draws (0.115 to 0.251 percent at 2^26 nonces) and h = 0.30 percent for or, 4.5 for an or chain, 0.067 for mul, 0.049 for mulhi, 0 for an injecting or bijective writer (the figures of section 2 at the measured bias). The bound: H at or under 1.2 x W, which is the static rule "every load's source was last written by an injecting op or a rotate" (any lossy writer breaks 1.2x). Its cost as a rejection rule on today's stream: 96.6 percent of candidates, about 30 attempts per seed on average. The cheaper form is a generator draw, not a rejection: draw a load's source from the registers whose last writer injects (today's rule draws from every written register), which costs no attempts and leaves rule (a) as it is. Either way the 64-seed census of F8's phase E is the gate, with the dynamic check extended to count saturated load sources over the 64 units beside the final values.

7. What is unverified

  • The per-site h figures are the popcount and trailing-zeros models at the biases F8 measured on p1 and p2; p3's chain figure is F8's measurement, not a model. F8's phase E (64 seeds, dynamic) is the test of the whole table.
  • The window model's top-f shares for the census use a normal approximation per quarter (p1's exact Poisson 0.160 against 0.159).
  • No GPU run and no timing here; every number is a count or arithmetic.