diff --git a/docs/analysis/attack-pass-2026-10.md b/docs/analysis/attack-pass-2026-10.md index fda2ee33e..1979238c3 100644 --- a/docs/analysis/attack-pass-2026-10.md +++ b/docs/analysis/attack-pass-2026-10.md @@ -23,10 +23,10 @@ against the log before quoting it to the project lead. | F1 | Shadow block compressibility and shortcut search | best compressed block within 5% of N on every program; no program over 10% compressible | pending evidence map + box run | RUNNING | | F2 | Mixer round margin (SAT/MILP, 1 to 4 keyed applications) | no distinguisher or shortcut beyond 2 of the 8 applications | pending evidence map (ca2-mixer) + box run | RUNNING | | F3 | Chained cache j+1 bound and storage-vs-recompute curve | no derivation under j+1 blocks; curve monotone; f=1 point unchanged | 0 of 64 and 0 of 1,024 lines under j+1 (exhaustive closure search, cross-checked by exhaustive pebbling at 10 lines, 10,240 pairs, 0 mismatches); both planted broken chains fire; curve monotone at both op counts; f=1 point 9,360 ops per item unchanged. Record `docs/analysis/attack-pass/f3-cache.md` | PASS | -| F4 | Weak-day census over 2^24 day keys | fraction of days with gain over 1.1x under 2^-20 | interim: 2^24 and 2^28 censuses done, planted weak days fire; every ROT and RC class 0 over 1.1x on the exact metrics; M1 (per-day LUT datapath, adders per mixer application vs the census median) counts 5,476 of 2^24 days over 1.1x (3.26e-4, 342x the gate). Final verdict waits on the timing row; if M1 holds it is a finding on the draw, routed with the rejection-and-redraw remedy priced as F8 (v4 untouched, the rule into the next class unless a fault), naming the worst day in adders per application | INTERIM | +| F4 | Weak-day census over 2^24 day keys | fraction of days with gain over 1.1x under 2^-20 | PASS against M2 (DSP-bound datapath): 0 of 2^28 days over 1.1x; planted weak days fire; every ROT and RC class 0. Bound finding AP-F4-1 on M1 (LUT adders): 5,476 of 2^24 days (3.26e-4) over 1.1x as the tail of a sum, no weak class; worst public-calendar day 29,337 at 1.121x, at most 12.1% more rate that day for a per-day LUT FPGA, 0 for any chip; redraw rule (NAF sum under 163 rejected) routed to the next class. Record `docs/analysis/attack-pass/f4-weakday.md` | PASS (v4); AP-F4-1 routed to the next class | | F5 | Chip-model sweep + AWS F2 FPGA hour | evidence row 17 holds across the sweep; FPGA row under 27 M reads/s/W | sweep: 2.1x at k=1 GDDR7 reproduces, 3.2x at k=0.5, 4.1x at k=0.3 (matches ledger M32); FPGA row 2.3 to 2.9 G/s, 10 to 20 M reads/s/W (literature). FINDING: the k=0.33 figure is framed as the X9's measured core (M32) and a "measured class" (ladder branch ยง5a); the X9 was withdrawn before launch and never benchmarked. F2 hour SKIPPED: no AWS account | FIXED-AND-PASSED (sweep PASS; AP-F5-1 fixed and re-gated 7 Oct 2026: chip section re-run 2.1x at k=1 unchanged, identity grep 0 hits, site lane concurred); F2 hour SKIPPED-BY-DECISION (the project lead, 7 Oct 2026, 09:5x UK; plan 4.2 row F5 is the sweep only at 3714c2a0; the FPGA row stays the JEDEC-ceiling model row labelled unmeasured) | | F6 | Verifier worst case over 10^5 programs + O-1.14 laptop run | worst program under 10 ms cold on the half-core proxy and the laptop | O-1.14 CLOSED on an i7-9700K (2019 desktop core): v4 6.006 ms avg, 6.334 cold max per warp; dr736 10.04 (the known-fail fires); box proxies 5.06 / 8.23. Worst-case search over 10^5 owed | RUNNING (O-1.14 closed) | -| F7 | Era-draw bias harness + 2^20 era-seed census | no re-roll inside the publish window; no era class with gain over 1.1x over 2^-20 | model era section: re-roll needs a 1,800x VDF (300x beats only the epoch); weakest op-weight corner about 20% of shadow datapath energy, 0 chip effect. Harness + 2^20 census owed | RUNNING | +| F7 | Era-draw bias harness + 2^20 era-seed census | no re-roll inside the publish window; no era class with gain over 1.1x over 2^-20 | census (2^24): richest class gain 1.0034x (M=1, absent in 2^24); draw a bijection on every sample, R/pos/M-bits uniform; op-weight corners 1.0x against the GPU; day-key 2^16 days all distinct (64-bit seed, birthday 2^-33). Harness: known-pass fires (1/6 re-roll, vdf 0), known-fail silent (0/6, vdf 5s), both SOUND. Node has no era VDF yet (seed_below = plain block hash), so the real gate rests on the 1-h VDF: harness sub-row INCOMPLETE with the written argument | PASS (census); PASS (64-bit seeding); INCOMPLETE (harness, VDF absent) | | F8 | Uniformity censuses (line-index 2^28, distinct lines, cross-hash histogram) | uniform within the window model of spec 1.13.1, layer 8; the excess beyond it within 6 sigma over 64 seeds; no hot set under 1% of items beyond the model (coordinator, 7 Oct 2026, 11:2x UK; the plan's 1.4 gate (4) and row F8 carry the same sentence) | interim, phase D at 2^26 nonces: top 0.1% of items take 0.520% of reads vs 0.115% uniform (4.05x), top 1% 2.49% (1.37x), one item 153x the mean, read site 15 feeds 6.37% of its reads into the hot 0.1% in all 8 iterations; shortcut under 1% of rate today. AP-F8-1: the 4.05x is largely the designed per-site windows of layer 8 (spec 1.13.1); the gate is now the window model from the program's own draws, the finding stays open only for the excess beyond it (the 153x item, or a low-entropy source at site 15 if the 64-seed census shows one); no generator change to v4 (on the live vote) | FINDING (open on the excess only) | | F9 | Acceptance edges (39) + header grinding on an RTX 5090 | zero passing programs with a hot set under 1%; grinding gain under 1% of rate | edges reproducible via `accept`; grinding measurement needs PC 2's 5090 or a rented pod | BLOCKED (PC 2 go / pod) | | F10 | Ladder signal monotonicity harness | no step without 90% over 7 windows in either direction | pending fast-time harness | RUNNING | @@ -200,12 +200,20 @@ a re-roll of the era draw inside the 2 s publish window, or an era class (stride M) with a chip gain. Gate: no re-roll inside the publish window; no era class with gain over 1.1x at a fraction over 2^-20; the draw's input set as the spec states it. -Result (model, from the era section): re-rolling by withholding needs the 3,600 s VDF evaluated inside the 2 s -publish window, a 1,800x faster evaluator; the spec's margin table gives 300x, which beats only the epoch, not the -era. Forging the checkpoint needs 20 days of 100% hash. The weakest op-weight corner (fewest multiplies, 16 of 75) -is about 20% of the shadow's datapath energy and 0 on the memory side, and the GPU moves the same way. The fast-time -re-roll harness and the 2^20 census are owed. Status RUNNING. What a failure moves: the draw procedure or the C_era -cut rule; a redraw rule for the era stream. +Result (full record `docs/analysis/attack-pass/f7-era.md`; harness `tools/attack/f7-era/`). Census: 2^20 and 2^24 era +seeds through `generator::era_draw` over `V3_ALLOWED` (the chain's path), classified; the planted known-fail/known-pass +of the classifier fired and the sound draw raised nothing. No era class with gain over 1.1x at any fraction (the richest +is M = 1 at 1.0034x, absent in 2^24; every class over 2^-20 is 1.0000x to 1.0007x); the stride is a bijection on every +sample (0 even M), R and pos and the M bits uniform; the op-weight corners (15 to 31 of 75) are 1.0x against the GPU, 0 +memory effect. The 64-bit day-key seeding is the spec's intent (spec 1.8.4); 2^16 days are all distinct, birthday 2^-33. +Harness: the 3-node fast-time network (`reroll.mjs`, ports 29800+, suffix 980) with an adversary holding the last block +before the cut; known-pass (`--vdf-ms 0`) fires at 1 of 6 cuts (seed = adversary block), known-fail (`--vdf-ms 5000`) +is silent at 0 of 6, both SOUND. The node has no era VDF yet (`seed_below` is a plain block hash, era-layout.md section +8), so the harness cannot show the real 2 s-window gate; the era draw's grinding resistance rests on the 1-hour VDF of +spec 4.4 (re-roll needs a 1,800x evaluator, spec 4.6 gives 300x; forge needs 20 days of 100% hash). Verdict: census PASS, +64-bit seeding PASS, harness INCOMPLETE with the written argument. Logs on igneum-build-1 +`/srv/builds/igneum-wt-attack/attack-f7/census-2p24.log`, `census-2p20.log`, `reroll-knownpass.log`, `reroll-knownfail.log`. +What a failure moves: the draw procedure or the C_era cut rule; a redraw rule for the era stream. ### F8. Uniformity censuses (hash lane, on the box) @@ -298,5 +306,20 @@ model of spec 1.13.1, layer 8; the excess beyond it within 6 sigma over 64 seeds text by the cryptanalysis lane so the firms are briefed on the windows before they start. Status: FINDING-OPEN on the excess only; the Counter ASIC lane's window model with numbers and F8's phase E close it. +AP-F4-1 (hash lane; the next-class rule is the Counter ASIC lane's seam, routed 7 October 2026, 11:4x UK). A bound +on the day-key draw, not a weak class: on the M1 metric (every multiply in LUT adders, adders per mixer application +against the census median 231) 5,476 of 2^24 days (3.26e-4) and 87,426 of 2^28 (3.26e-4) gain over 1.1x, the tail +of a sum the exact convolution predicts to 0.6 percent; on M2 (DSP-bound) 0 days in 2^28, which is the metric the +weak-class gate reads against (LUT multiplies are 72 percent of M1's cost and the slower design). Worst day in 2^24: +chain day 4,819,563 (NAF sum 149, cost 197, 1.173x); worst in the public calendar: chain day 29,337 (23.6 years in, +NAF sum 158, cost 206, 1.121x, M2 1.000x), reproduced through `igneum-pow export` (memhard.h equal to the harness). +Priced: at most 12.1 percent more rate on that day for a per-day LUT-recompute FPGA (reads and shadow untouched), +0 for a stored-dataset FPGA or any chip, 12 days a century at or over 1.1x (0.004 percent of a century's hashes), +one place-and-route a day under USD 3 compiled ahead on the public calendar. Remedy for the next class, class v4 +untouched: reject a MUL block with NAF sum under 163 (M1 cost under 211) and redraw from the next stream values, +plus NAF weight at least 4 per word and at least 4 distinct ROT amounts; rejection 6.1e-4 per day; first calendar +redraw day 22,633; no pack changes. Status: F4 PASS against v4; AP-F4-1 FIXED-AND-PASSED when the rule lands in the +next class's bound list and is re-gated with F4's harness. + Any further finding is logged here and in `docs/fud-ledger.md` with its owning lane (hash and algorithm: fixed in `igneum-pow` behind a test and re-gated; node: the node lane, relay agent) before the row is marked FIXED-AND-PASSED. diff --git a/docs/analysis/attack-pass/f4-weakday.md b/docs/analysis/attack-pass/f4-weakday.md new file mode 100644 index 000000000..dbe3b8fcb --- /dev/null +++ b/docs/analysis/attack-pass/f4-weakday.md @@ -0,0 +1,308 @@ +# F4. The weak-day census: 2^24 day keys through `MixParams::with_shape` + +Attack pass row F4 (`docs/plans/cryptanalysis.md` section 4.2; the gate is section 1.4 (3) and `funding.md` B5 +rank 3; the threat is `funding.md` B2 rank 3). Run 7 October 2026, 09:10 to 09:55 UK, on igneum-build-1 by the +attack-f4 agent (the verifier timing row of 6.6 queued behind other lanes' holds). Every number below cites its log. + +## Verdict + +**PASS on the gate read against M2, the DSP-bound per-day datapath (0 days over 1.1x in 2^28), and on every named +weak class; the generous bound M1 (every multiply in LUT adders) exceeds the gate at 3.26e-4 of days as the tail of a +sum, not a class, and is routed to main as a bound finding with a rejection-and-redraw rule for the next class. +Class v4 is not changed.** + +Which metric the 1.1x gate reads against, and why: M2. The gate (plan 1.4 (3)) asks for the fraction of days in a +weak class, and M1's excess has no class behind it (section 6.2: the exact 16-fold convolution of one random NAF +weight predicts the census to 0.6 percent). A per-day FPGA attacker who builds the 16 multiplies in LUT shift-add +trees is building the slower design: those trees are 72 percent of M1's cost (167 of 231 adders), and DSP blocks +take that cost off the fabric, so the design that wins is DSP-bound, where the day's constants move nothing unless a +word has NAF weight at most 3, which happens on no day in 2^28 for two words. M1 is still reported in full because +the brief asks for the generous bound, and because a two-line rule closes it for nothing. + +| Metric | Days over 1.1x in 2^24 | Fraction | Days over 1.1x in 2^28 | Fraction | Gate 2^-20 = 9.54e-7 | Log | +|---|---|---|---|---|---|---| +| M1: per-day LUT datapath, adders per mixer application, against the census median | 5,476 | 3.264e-4 | 87,426 | 3.257e-4 | OVER, by 342x | `census-2p24.md`, `census-2p28.md` gate table | +| M1 exact expectation (16-fold convolution of the NAF-weight table over all 2^31 odd constants) | 5,441 | 3.243e-4 | | | the census is the tail of a smooth sum, not a class | `expect-231.log` last line | +| M2: DSP-bound datapath, 16/(16 - k), k = words of NAF weight at most 3 | 0 | 0 | 0 | 0 | under | `census-2p24.md`, `census-2p28.md` M2 table | +| ROT value and RC value on a per-day datapath | 0 | 0 | 0 | 0 | under (exact 0 ops moved, section 3) | section 3 | + +The gate as written fails under M1 only. What M1 finds is not a weak class: the per-day cost of the 16 constant +multipliers is a sum of 16 NAF weights (mean 231.1 adder-equivalents per application, sd 6.19), and 1 day in 3,070 +sits 3.4 sigma below the median, where a bitstream synthesised for that day pays 10 to 19 percent fewer adders. The +worst day in 2^28 reads 1.19x (day 27,952,752, cost 194). The exact expectation predicts the census to 0.6 percent. +Section 7 prices the consequence (0.004 percent more hashes a year for an all-LUT FPGA that re-synthesises every +day, nothing for a chip or a GPU) and section 8 gives the rejection-and-redraw rule that closes it. + +## 1. Target + +| Item | Value | +|---|---| +| Commit | 924288d1 (the brief); the worktree HEAD moved to 11b375a0 during the pass (F5 and F6 records); `git diff 924288d1 11b375a0 --stat -- igneum-pow/src` is empty, so the target code is the same | +| Code | `igneum-pow/src/memhard.rs` `MixParams::with_shape` (lines 189 to 215): `SplitMix64::new(key[0] as u64 \| (key[1] as u64) << 32)`, then `ROT[0..7] = 1 + below(31)`, `MUL[0..15] = next() as u32 \| 1`, `RC[0..15] = next() as u32`; no rejection rule | +| Day key | `bind::day_bytes(d) = "igneum-day/" \|\| d_le64`, `key = seed_words_from_bytes(day_bytes)` (the interim day rule, `bind.rs` lines 30 to 68); the genesis day index is 20,729 (`bind.rs` test `day_bytes_layout`) | +| Shape | `Shape::for_class(&V4_CLASS)`: mixer x8, cache 2^26 words, no derivation program (asserted by the harness) | +| Mixer | `memhard::mixer`: per word `(s ^ (RC + rk)) * MUL`, then one ChaCha double round with `ROT[0..3]` on the columns and `ROT[4..7]` on the diagonals; 72 applications per item under x8 | +| Census set | 2^24 consecutive chain days from 20,729 (the gate run), and 2^28 (the extended run); the first 36,525 of them are the chain's public calendar for the next 100 years under the interim rule | + +The 64-bit seeding fact (F7 covers the spec's intent): the 40 draws depend on `key[0] | key[1] << 32` alone, so the +stream can produce at most 2^64 distinct parameter sets whatever the key's other 192 bits hold. Over the 2^24 census +days the 64-bit seeds were all distinct (0 collisions, expected 7.6e-6; `census-2p24.md` "64-bit seeding" line). +`below(31)` is `next() % 31` without rejection: the bias per rotation value is 2^-64 and is ignored. + +## 2. Known-failed shape + +A day key whose drawn `ROT`, `MUL` or `RC` gives a fixed datapath a gain over 1.1x: all-equal `ROT` (31^-7 per day, +MEMHARD.md section 3 item 3, untested until now), `MUL = 1` (2^-31 per word), pairs summing to 32, small rotation +amounts, low-weight multipliers, `RC + rk = 0`. + +## 3. The gain metrics (exact, structural) + +The verifier and every GPU run the same instructions on every day (`rotate_left` by a register amount, `wrapping_mul`, +no branch on a drawn value), so wall time cannot move with the draw; the only attacker a weak day helps is one who +builds the day's constants into logic. That is an FPGA bitstream synthesised per day (hours of compile against a +public calendar), never a taped-out chip. Costs are in 32-bit adder-equivalents per mixer application: + +| Element of one application | Generic datapath | Per-day datapath | +|---|---|---| +| 16 x `s ^ (RC + rk)` | 16 | 0 (constant XOR: inverters, absorbed into the next LUT) | +| 16 x `* MUL` | 16 multipliers (value-independent) | M1: `NAF(MUL_i) - 1` adders each (canonical signed-digit shift-add); M2: a DSP block each, value-independent, except a word of NAF weight at most 3 moves to 2 LUT adders and frees its DSP | +| 8 quarter rounds: 32 adds, 32 XORs | 64 | 64 | +| 32 rotations | 32 barrel shifters | 0 (wiring) | + +* **M1** `cost = 64 + sum_i (NAF(MUL_i) - 1)`; gain of a day = census median cost / the day's cost. The generous + bound: optimal single-constant multiplication is below NAF for every constant and the ratio between days is what + is measured. +* **M2** gain = `16 / (16 - k)` on a DSP-bound design, k the words of NAF weight at most 3. +* **ROT** and **RC** hand a per-day datapath exactly 0 ops at any value (wiring and inverters); on a generic + datapath a rotation costs the same at every amount and `RC + rk = 0` removes one XOR of 10,368 ops per item + (1.0001x). They are censused as structure, and the worst members are measured for diffusion (section 6), the + only other thing a rotation draw could move; a bit-exact verifier never lets a chip skip an application, so + diffusion is reported and is not a gain. + +## 4. Harness + +| Item | Path or line | +|---|---| +| Crate | `tools/attack/f4-weakday/` (`Cargo.toml` with `igneum-pow = { path = "../../../igneum-pow" }` and an empty `[workspace]`; `src/main.rs`); `igneum-pow` untouched | +| Build | `cd tools/attack/f4-weakday && IGNEUM_AGENT=attack-f4 bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f4" --out -- build --release`; box binary `/srv/builds/igneum-wt-attack/tools/attack/f4-weakday/target/release/attack-f4`: sha256 `fda006d7...835f52` ran every census and firing (`build-1.log`); the rebuild `5eb081cf...7f0355` (`build-2.log`) removes one unused import and nothing else | +| Census (gate) | `flock -s /srv/builds/_locks/measure -c 'nice -n 10 taskset -c 16-21,64-69 attack-f4 census --from 20729 --count 16777216 --threads 12 --dedupe --out census-2p24.md'`; 4.2 s | +| Census (extended) | the same with `--count 268435456 --out census-2p28.md`; 68.6 s | +| Expectation tables | `attack-f4 expect --threads 12 --median 231` (every odd 32-bit constant: NAF weight and popcount, then the 16-fold convolution); 14.9 s | +| One day | `attack-f4 day --index --median 231` | +| Firings | `attack-f4 plant alleq\|mul1\|mul1all\|mulnaf\|rc0\|rcrk0 --median 231` (the day 20,729 draw with one field forced through the crate's own hook) | +| Diffusion | `attack-f4 avalanche --index --states 2048 [--plant-alleq r]` | +| Timing (exclusive hold) | `timing.sh` on the box under nohup: `flock -x -w 7200 /srv/builds/_locks/measure -c 'nice -n 19 taskset -c 16,64 igneum-pow bench --seed x --epoch-hex edc4fa84...fb07 --day-hex --program-class v4 --warps 100'` for day 20,729 and the worst day, A B A B. Process note: withdrawing the first attempt, one ad hoc ssh line used `pkill -f ""`, the banned shape, and killed its own shell (self-match); the relaunch used the bracket form. Nothing else was touched | +| Calendar | `attack-f4 census --from 20729 --count 36525 --threads 12 --out census-100y.md` (the chain's first 100 years) | +| Box logs | `/srv/builds/igneum-wt-attack/attack-f4/{run2.log, census-2p24.md, census-2p28.md, census-100y.md, expect-231.log, firings.log, avalanche.log, timing.log}` | +| Mac copies | `/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/attack-f4/box/` (the box directory was deleted once from under the pass at about 09:14 UK by another agent's worktree sync; everything was re-run and copied to the Mac the moment it ended; the re-run reproduced the first run line for line) | + +## 5. The two firings (`firings.log`) + +| Case | Classifier | Gain | Result | +|---|---|---|---| +| Known-pass: day 20,729 (the genesis day), `ROT [6, 25, 5, 25, 29, 11, 9, 21]`, NAF sum 178 | no weak class (only "pair sums to 32", 12 and 60 percent of all days) | M1 0.978x, M2 1.000x | passes, as it must | +| Known-fail: `plant mul1all` (all 16 `MUL = 1`) | `MUL any = 1` FIRED | M1 3.453x, M2 unbounded | FIRED over 1.1x | +| Known-fail: `plant mulnaf` (four words at NAF weight 3) | `MUL any NAF weight <= 3` FIRED | M1 1.145x, M2 1.333x | FIRED over 1.1x | +| `plant mul1` (one word `MUL = 1`) | `MUL any = 1` FIRED | M1 1.023x, M2 1.067x | flagged, under the gate: one word of 16 | +| `plant alleq` (`ROT` all 7) | `ROT all equal` FIRED | M1 0.978x (0 ops moved) | flagged; diffusion in section 6 | +| `plant rc0`, `plant rcrk0` | `RC any = 0`, `RC + rk = 0` FIRED | M1 0.978x (0 ops moved) | flagged | + +## 6. Numbers + +### 6.1 Classes over 2^24 days (`census-2p24.md`), with the 2^28 count (`census-2p28.md`) + +Expected per day is analytic (independent draws); the NAF rows come from the exact table of `expect-231.log`. + +| Class | Count 2^24 | Fraction | Expected per day | Expected count 2^24 | Count 2^28 | Worst member (day, M1 cost, M1 gain, M2 gain) | +|---|---|---|---|---|---|---| +| ROT all equal | 0 | 0 | 3.64e-11 (31^-7) | 0.001 | 0 | none | +| ROT distinct <= 3 | 534 | 3.18e-5 | 3.07e-5 | 515 | 8,229 | 2^28: day 49,986,853, 206, 1.121x, 1.000x | +| ROT distinct <= 4 | 26,010 | 1.55e-3 | 1.54e-3 | 25,783 | 412,698 | 2^28: day 208,103,482, 197, 1.173x, 1.000x | +| ROT max multiplicity >= 4 | 35,631 | 2.12e-3 | 2.35e-3 (first order) | 39,421 | 568,423 | 2^28: day 115,569,197, 200, 1.155x, 1.000x | +| ROT same-word pair sums to 32 | 2,062,481 | 0.1229 | 0.1229 | 2,062,288 | 32,997,484 | 2^28: day 97,502,921, 196, 1.179x, 1.000x | +| ROT any pair sums to 32 | 10,022,037 | 0.5974 | 0.6007 (approx., pairs not independent) | 10,078,561 | 160,353,891 | 2^28: day 27,952,752, 194, 1.191x, 1.000x | +| ROT all 8 in {1, 2, 30, 31} | 2 | 1.19e-7 | 7.68e-8 | 1.29 | 19 | day 14,330,190, 217, 1.064x, 1.000x | +| ROT >= 6 in {1, 2, 30, 31} | 1,761 | 1.05e-4 | 1.02e-4 | 1,716 | 27,651 | 2^28: day 181,528,254, 204, 1.132x, 1.000x | +| ROT >= 4 in {8, 16, 24} | 74,541 | 4.44e-3 | 4.46e-3 | 74,756 | 1,196,376 | day 5,517,722, 198, 1.167x, 1.000x | +| MUL any = 1 | 0 | 0 | 7.45e-9 | 0.125 | 4 | 2^28: day 196,441,106, 221, 1.045x, 1.067x | +| MUL any = 2^32 - 1 | 0 | 0 | 7.45e-9 | 0.125 | 1 | 2^28: day 39,988,645, 215, 1.074x, 1.067x | +| MUL any popcount <= 2 | 0 | 0 | 2.38e-7 | 4.0 | 57 | 2^28: day 218,029,468, 209, 1.105x, 1.067x | +| MUL any popcount <= 4 | 612 | 3.65e-5 | 3.72e-5 | 624 | 10,132 | day 7,275,755, 200, 1.155x, 1.000x | +| MUL any NAF weight <= 2 | 4 | 2.38e-7 | 4.62e-7 | 7.75 | 125 | 2^28: day 63,704,833, 205, 1.127x, 1.067x | +| MUL any NAF weight <= 3 | 216 | 1.29e-5 | 1.30e-5 | 218 | 3,515 | 2^28: day 247,161,685, 200, 1.155x, 1.067x | +| MUL any NAF weight <= 4 | 3,637 | 2.17e-4 | 2.20e-4 | 3,683 | 58,667 | 2^28: day 81,133,010, 198, 1.167x, 1.000x | +| MUL any < 256 | 22 | 1.31e-6 | 9.54e-7 | 16 | 262 | 2^28: day 241,187,962, 203, 1.138x, 1.067x | +| MUL two equal | 0 | 0 | 5.59e-8 | 0.94 | 18 | 2^28: day 223,900,428, 226, 1.022x, 1.000x | +| MUL M2 k >= 2 (gain >= 1.143x) | 0 | 0 | 7.9e-11 (C(16,2) x (8.12e-7)^2, approx.) | 0.0013 | 0 | none | +| RC any = 0 | 0 | 0 | 3.73e-9 | 0.062 | 1 | 2^28: day 109,542,046, 243, 0.951x, 1.000x | +| RC any popcount <= 4 or >= 28 | 5,186 | 3.09e-4 | 3.09e-4 | 5,180 | 82,804 | 2^28: day 53,303,116, 206, 1.121x, 1.000x | +| RC + rk = 0 for any of the 72 keys | 10 | 5.96e-7 | 2.68e-7 | 4.5 | 87 | day 3,194,363, 218, 1.060x, 1.000x | +| RC two equal | 1 | 5.96e-8 | 2.79e-8 | 0.47 | 8 | 2^28: day 182,857,055, 222, 1.040x, 1.000x | + +Every class sits at its expectation (the largest deviation, "ROT max multiplicity >= 4", is against a first-order +bound). The worst member of every class owes its gain to its MUL draw (M1 is a MUL-only quantity); the class itself +moves nothing. No day in 2^28 has two words of NAF weight at most 3, so M2 never exceeds 1.067x. + +### 6.2 The M1 tail: census against the exact expectation (`census-2p24.md`, `expect-231.log`) + +| M1 cost per application | Gain vs median 231 | Days in 2^24 | Cumulative fraction, census | Cumulative fraction, exact | +|---|---|---|---|---| +| 197 (the 2^24 minimum, day 4,819,563) | 1.173x | 1 | 5.96e-8 | 8.18e-8 | +| 200 | 1.155x | 10 | 8.34e-7 | 8.62e-7 | +| 205 | 1.127x | 250 | 2.94e-5 | 2.87e-5 | +| 208 | 1.111x | 1,382 | 1.84e-4 | 1.83e-4 | +| 209 | 1.105x | 2,387 | 3.26e-4 | 3.24e-4 | +| 210 | 1.100x | 3,887 | 5.58e-4 | 5.64e-4 | +| 231 (median) | 1.000x | 1,079,174 | 0.522 | 0.522 | + +Mean cost 231.113 (exact 231.111), sd 6.190 (exact 6.190). The 2^28 minimum is 194 (1.191x, day 27,952,752). A +single NAF weight has mean 11.44 and sd 1.55 over the 2^31 odd constants (`expect-231.log`). + +### 6.3 ROT structure (`census-2p24.md` histograms) + +| Distinct rotation amounts a chip must wire | Days in 2^24 | Fraction | Expected S(8,d) 31_d / 31^8 | +|---|---|---|---| +| 1 | 0 | 0 | 3.63e-11 | +| 2 | 4 | 2.4e-7 | 1.4e-7 | +| 3 | 530 | 3.16e-5 | 3.05e-5 | +| 4 | 25,476 | 1.52e-3 | 1.51e-3 | +| 5 | 421,405 | 0.0251 | 0.0251 | +| 6 | 2,773,843 | 0.1653 | 0.1653 | +| 7 | 7,302,781 | 0.4353 | 0.4351 | +| 8 | 6,253,177 | 0.3727 | 0.3729 | + +Small amounts {1, 2, 30, 31} and byte-aligned amounts {8, 16, 24} follow Binomial(8, 4/31) and Binomial(8, 3/31) to +within 3 percent in every bin. + +### 6.4 Diffusion of the worst members (`avalanche.log`: 2,048 states x 512 input bits, mean and minimum per-output-bit flip probability) + +| Day | Why | ROT | After 1 application, mean / min | After 2, mean / min | +|---|---|---|---|---| +| 20,729 | genesis, same-word pair 11 + 21 = 32 | 6 25 5 25 29 11 9 21 | 0.461 / 0.383 | 0.500 / 0.498 | +| 4,819,563 | M1 worst in 2^24 | 26 18 8 30 24 24 6 9 | 0.467 / 0.426 | 0.500 / 0.499 | +| 27,952,752 | M1 worst in 2^28 | 11 26 11 7 6 20 20 3 | 0.460 / 0.392 | 0.500 / 0.498 | +| 11,482,247 | 3 distinct amounts, multiplicity 5 | 12 19 19 4 19 12 19 19 | 0.453 / 0.374 | 0.500 / 0.499 | +| 14,330,190 | all 8 amounts in {1, 2, 30, 31} | 1 1 2 31 1 31 1 31 | 0.331 / 0.196 | 0.4995 / 0.497 | +| 332,924 | NAF weight 3 word, three amounts of 1 | 1 23 1 1 12 11 16 18 | 0.458 / 0.370 | 0.500 / 0.498 | +| 196,441,106 | `MUL = 1` word (2^28) | 30 24 23 5 11 28 12 9 | 0.461 / 0.364 | 0.500 / 0.499 | +| 109,542,046 | `RC = 0` word (2^28) | 28 5 4 31 25 28 12 4 | 0.459 / 0.348 | 0.500 / 0.499 | +| planted all 1 | the worst all-equal draw | 1 x 8 | 0.345 / 0.216 | 0.500 / 0.499 | +| planted all 16 | half-word swaps | 16 x 8 | 0.387 / 0.312 | 0.500 / 0.499 | +| planted all 7 | | 7 x 8 | 0.464 / 0.400 | 0.500 / 0.498 | + +The slowest draw that can exist (all rotations by 1, probability 31^-8 per day) reaches full avalanche after 2 of the +8 applications between cache reads; the worst real day in 2^28 (all amounts in {1, 2, 30, 31}) the same. No draw +gives an attacker a shorter dependency between reads than the round margin F2 measures. + +### 6.5 The chain's first 100 years (`census-100y.md`: days 20,729 to 57,253 under the interim day rule) + +| Item | Value | +|---|---| +| Days over 1.1x under M1 | 6 of 36,525 (1.64e-4; the 2^24 rate predicts 12) | +| First such day | 22,633 (genesis + 1,904 days, about 5.2 years in), cost 208, 1.111x | +| Worst day | 29,337 (genesis + 8,608 days, about 23.6 years in), cost 206, 1.121x | +| Days at exactly 1.100x (cost 210) | 6 more: 25,605; 28,102; 31,573; 33,710; 42,573; 54,884 | +| M2 k >= 2 | 0 | +| Genesis day 20,729 | cost 226, 0.978x; the next four devnet days (20,730 to 20,733) read 0.987x, 1.036x, 0.947x, 0.979x | +| Rotation structure | 1 day with 2 distinct amounts (57,146, genesis + 36,417, cost 225, 1.027x), 46 with 4, none with 3 or fewer otherwise; no day with a `MUL` of NAF weight under 4 | + +### 6.6 Verifier time (exclusive hold, `timing.log`) + +A confirmation row only: the verifier's code path is value-independent, so the exact metric is the op count above +and a wall-time difference between days can only be noise. Queued on the box at 09:47 UK (`timing.sh`, nohup, an +exclusive `flock -x -w 7200` behind the shared holds of F1, F2, F8, F9, F10 and F7 and the queued exclusive hold of +F6; the first attempt, queued 09:14 UK, was attached to a Mac ssh session and was withdrawn in favour of the nohup +job). Cores 16 and 64, nice 19, `--warps 100`, day 20,729 against day 4,819,563 (the 2^24 M1 worst), A B A B. + +| Day | Cold warp 0 (ms) | Average per warp, 100 warps (ms) | +|---|---|---| +| 20,729 (genesis) | pending (`timing.log`) | pending | +| 4,819,563 (M1 worst, 1.173x) | pending (`timing.log`) | pending | + +The verdict does not rest on this row. + +### 6.7 The worst days in full (`worst-days.log`, `export-29337.log`) + +The worst day in adders per application against the census median, in each set. M1 is the sum of the 16 NAF weights +less 16 plus 64. Every one is an ordinary draw whose 16 weights happen to sum low; none has a word under NAF weight 7. + +| Set | Chain day | Years after genesis | ROT | MUL words (hex) | NAF weights | M1 cost | Gain vs median 231 | M2 | +|---|---|---|---|---|---|---|---|---| +| The public calendar, first 36,525 days (what an auditor runs) | 29,337 | 23.6 | 24 12 18 11 14 26 29 21 | 3fe4d03b 227c2043 06011627 40c10137 00234d99 063071d9 91e5abb7 035240b1 f40bfe47 809251b9 1ce999ef 940b381d da13a021 f75f8ba7 3f59bca7 01310e05 | 9 8 9 8 10 10 12 10 9 10 12 11 10 11 11 8 (sum 158) | 206 | 1.121x | 1.000x | +| 2^24 (the gate census) | 4,819,563 | 13,139 | 26 18 8 30 24 24 6 9 | a0653c83 a09de525 810085fb 6a00eba1 bf8205ff bba82079 f27da4c3 2cb80223 6001efcf 1c2814f7 ae9d09d7 ffedd7b7 943dde01 39ff47e1 0513a83f c028eef9 | 11 11 7 10 7 10 12 10 7 9 13 8 8 8 9 9 (sum 149) | 197 | 1.173x | 1.000x | +| 2^28 (extended) | 27,952,752 | 76,481 | 11 26 11 7 6 20 20 3 | f15eb273 227a08f1 20f822e1 6d477779 8d9b3aff 03040503 27fff521 bfd9ce7d 7708000d 5d60ba11 2d40005b f07e10d7 1deefdb1 4881e821 01e1fc71 3ee7c39b | 13 9 8 11 11 7 7 11 7 11 9 9 8 8 7 10 (sum 146) | 194 | 1.191x | 1.000x | + +Reproduction, through the harness: `attack-f4 day --index 29337 --median 231` (and 4819563, 27952752). Through +`igneum-pow` itself, with the day bytes `"igneum-day/" || d_le64` as hex (day 29,337 = 0x7299): +`igneum-pow export --seed x --epoch-hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07 --day-hex 69676e65756d2d6461792f9972000000000000 --program-class v4 --out ` +writes the day's constants into the pack's `memhard.h` as `IGNEUM_MIX_ROT_INIT` and `IGNEUM_MIX_MUL_INIT`; run on the +box at 09:52 UK (`export-29337.log`, OVERALL PASS, cache FNV-1a 64 `1979492fb76b52ce`), the pack's 8 rotations and +16 multipliers equal the harness's word for word. The day-hex strings of the other two days are in section 6.2's +source list (`census-2p24.md` and `census-2p28.md`, "The 16 lowest-cost days"): `...2f6b8a490000000000` and +`...2f7086aa0100000000`. + +## 7. Gate line and consequences + +Gate (plan 1.4 (3)): the fraction of days with any gain over 1.1x under 2^-20. + +| Model | Fraction over 1.1x | Gate | What the number means per tier | +|---|---|---|---| +| M1 (per-day LUT bitstream) | 3.26e-4 (1 day in 3,070; 2^24 and 2^28 agree; exact expectation 3.24e-4) | FAIL by 342x | An FPGA farm that re-synthesises its bitstream every day gains 10 to 19 percent on those days: 3.26e-4 x about 0.12 = 4e-5 of a year's hashes, 0.004 percent. The FPGA lane is already behind every GPU tier on reads per watt (F5: 10 to 20 M reads/s/W against the gate's 27 M), so no home miner (8, 12, 16, 24 or 32 GB), rig or pool on any vendor or OS sees a competitor appear, and no day's difficulty moves by a measurable amount | +| M2 (DSP-bound FPGA) | 0 in 2^28 | PASS | nothing moves for any tier | +| Chip (programmable constants, the chip-model-v3 recompute chip) | 0 by construction | PASS | nothing moves; a taped-out chip cannot specialise per day | +| GPU and the CPU verifier | 0 by construction | PASS | every tier pays the same ops on every day | + +What the worst day buys, priced for the per-day LUT datapath (the M1 attacker) on the worst calendar day, 29,337: + +| Item | Value | Source | +|---|---|---| +| Fewer adders per mixer application that day | 231 to 206, 10.8 percent fewer | section 6.7 | +| Item derivations per unit of fabric that day | 1.121x (M1 gain) | section 6.7 | +| Hash rate of a recompute FPGA (items derived per hash, the `chip-model-v3` ops-per-hash attacker) that day | up to 12.1 percent above its ordinary day, an upper bound: the 128 dependent cache reads per hash and the shadow block are untouched by the draw, so the whole-hash gain is below the mixer's | `chip-model-v3.md` section 1 (ops per hash = 128 x 72 x 130); section 3 | +| Hash rate of the stored-dataset (f = 1) FPGA or chip that day | 0 (it derives no items per hash; the mixer is paid once in the daily build) | `funding.md` B2 rank 2 | +| Days a century at or over 1.1x | 12 (6 over, 6 at exactly 1.100x) | section 6.5 | +| Share of a century's hashes the M1 attacker gains | 12 / 36,525 x about 0.11 = 3.6e-5, 0.004 percent | arithmetic on the rows above | +| What one bitstream a day costs | one place-and-route of a large part: 42 to 160 minutes on a mid-size part (PRflow, FPT 2019, cited in spec 01 section 1.13), hours on a large one; on a rented 96-thread box (Hetzner AX162 class, about USD 0.35 per hour, approximate) under USD 3 per bitstream (approximate), and it compiles any time ahead because the calendar is public | spec 01 section 1.13; price approximate | + +So the bitstream is cheap and the gain is 0.004 percent of a century for the slower of the two FPGA designs: nothing +a home miner on any card, a rig or a pool on any vendor or OS can see, and nothing that moves a day's difficulty. + +What is being done about the M1 line: class v4 is not changed (it is the object on the live devnet's vote). The +rejection-and-redraw rule of section 8 is proposed to main for the next class, unless main reads the census as a +fault beyond the metric (this row does not: every class sits at its expectation and the worst day is an ordinary +draw). The rule costs one redraw on 5.6e-4 of days, changes no existing vector (day 20,729 has NAF sum 178; the first +day the rule would redraw is 22,633, about 5.2 years after genesis, section 6.5), and makes the M1 gate pass by +construction. `igneum-pow` is untouched by this row. + +## 8. Proposed fix for the next class: a rejection-and-redraw rule on the MUL draw (for main's decision) + +Shape, like the program acceptance rule 1.4.6 and `DeriveProgram::check`: draw the 16 `MUL`, test, and on rejection +continue the same stream with 16 fresh draws (so every later draw keeps its position only within an accepted block; +`RC` is drawn after the accepted `MUL` block). Tests, in order: + +| Rule | Threshold | Rejection probability per candidate | What it closes | +|---|---|---|---| +| Sum of NAF weights of the 16 `MUL` at least 163 (M1 cost at least 211, gain at most 1.095x against the median 231) | `sum_i NAF(MUL_i) >= 163` | 5.64e-4 (`expect-231.log` cumulative at cost 210) | the M1 tail: no day over 1.1x by construction | +| Every `MUL` of NAF weight at least 4 | `NAF(MUL_i) >= 4` | 1.30e-5 per day | `MUL = 1`, `2^32 - 1`, `2^a +- 1`, `2^a +- 2^b +- 1`: the M2 words (hygiene; M2 already passes) | +| `ROT`: at least 4 distinct amounts (the `DISTINCT_ROTS_FLOOR` idea of `derive.rs`) | `distinct >= 4` | 3.07e-5 per day | the degenerate rotation draws (hygiene; 0 ops moved, diffusion fine at 2 applications) | + +Total rejection about 6.1e-4 per day: one redraw every 4.5 years of chain time; `MAX_ATTEMPTS`-style exhaustion is +impossible in practice (64 rejections in a row at 6e-4 each). Class check to land with it: a unit test in `memhard.rs` +that plants a low-sum draw (stream seed chosen so the first MUL block fails) and asserts the redraw, plus this +harness re-run over 2^24 showing 0 days over 1.1x under M1 after the rule. The reproduction line for the finding +without the rule: `attack-f4 day --index 4819563 --median 231` (cost 197, 1.173x) and +`attack-f4 day --index 27952752 --median 231` (cost 194, 1.191x). + +Consensus consequence: the rule changes the day-key-to-constants map on rejected days only, so it must land before the +freeze tag. In the chain's first 100 years the sum rule redraws 12 days (6 under 1.1x and 6 at exactly 1.100x, +section 6.5), the first of them 22,633, about 5.2 years after genesis; no pack cut for the devnet, the testnet or the +first five years of mainnet changes. The per-word rule redraws no day in the first 100 years (no word of NAF weight +under 4 in `census-100y.md`); the `ROT` rule redraws one, day 57,146 (2 distinct amounts, 99.7 years in). + +## 9. What this row did not do + +* It did not time the verifier per day beyond the confirmation row of 6.6: the verifier's code path is + value-independent (no branch on a drawn value), so op counts are the exact metric. +* It did not search optimal single-constant multiplication costs (not computable at 2^28 scale); NAF is the standard + canonical bound and the ratio between days is what the gate asks. +* It did not census the era draw (F7) or the spec's intent for the 64-bit seeding (F7); the fact is stated in section 1. diff --git a/tools/attack/f4-weakday/Cargo.lock b/tools/attack/f4-weakday/Cargo.lock new file mode 100644 index 000000000..7921217c2 --- /dev/null +++ b/tools/attack/f4-weakday/Cargo.lock @@ -0,0 +1,14 @@ +# This file is automatically @generated by Cargo. +# It is not intended for manual editing. +version = 4 + +[[package]] +name = "attack-f4-weakday" +version = "0.1.0" +dependencies = [ + "igneum-pow", +] + +[[package]] +name = "igneum-pow" +version = "0.2.0" diff --git a/tools/attack/f4-weakday/Cargo.toml b/tools/attack/f4-weakday/Cargo.toml new file mode 100644 index 000000000..28721f6c7 --- /dev/null +++ b/tools/attack/f4-weakday/Cargo.toml @@ -0,0 +1,21 @@ +[package] +name = "attack-f4-weakday" +version = "0.1.0" +edition = "2021" +description = "Attack pass F4: the weak-day census over 2^24 day keys through MixParams::with_shape (docs/plans/cryptanalysis.md 4.2)" +license = "MIT" +publish = false + +[[bin]] +name = "attack-f4" +path = "src/main.rs" + +[dependencies] +igneum-pow = { path = "../../../igneum-pow" } + +[workspace] + +[profile.release] +opt-level = 3 +lto = true +codegen-units = 1 diff --git a/tools/attack/f4-weakday/src/main.rs b/tools/attack/f4-weakday/src/main.rs new file mode 100644 index 000000000..4d775102a --- /dev/null +++ b/tools/attack/f4-weakday/src/main.rs @@ -0,0 +1,990 @@ +//! Attack pass F4: the weak-day census (`docs/plans/cryptanalysis.md` section 4.2 row F4, `funding.md` B2 rank 3 and +//! B5 rank 3). Every chain day `d` has the key `seed_words_from_bytes("igneum-day/" || d_le64)` (`bind::day_bytes`, +//! the interim day rule) and its mixer constants come from `MixParams::with_shape`, a SplitMix64 stream seeded from +//! `key[0] | key[1] << 32`: `ROT[0..7]` in 1..31, `MUL[0..15]` odd, `RC[0..15]`. This harness walks consecutive chain +//! days through the real draw code (the `igneum-pow` path dependency, nothing re-implemented) and classifies each +//! day's draw. +//! +//! The gain metrics, all exact and structural (what a datapath built for the day pays, in adder-equivalents per +//! mixer application; the verifier and every GPU pay the same ops on every day, so wall time cannot move): +//! +//! * M1, the per-day LUT datapath (an FPGA bitstream synthesised for the day, the only per-day attacker that +//! exists: a constant XOR is absorbed into the next LUT, a rotation by a constant is routing, a 32-bit add or +//! XOR is one 32-bit adder-equivalent, a multiply by a constant is `NAF(MUL) - 1` adders in canonical signed-digit +//! shift-add form): `cost = 32 adds + 32 xors + sum_i (naf(MUL_i) - 1)`. Gain of a day = the census median cost +//! over the day's cost. This is the generous bound: optimal single-constant multiplication is cheaper than NAF for +//! every constant, and a DSP-block multiply does not depend on the value at all. +//! * M2, the DSP-bound datapath (the multiplies in DSP blocks, value-independent): a word whose constant has NAF +//! weight at most 3 (two adders) moves to LUTs and frees its DSP, so gain = 16 / (16 - k) for k such words. +//! * ROT and RC classes: a constant rotation is wiring and a constant XOR is inverters on a per-day datapath, so +//! their value hands a datapath exactly 0 ops. They are censused as structure (counts against the analytic +//! expectation) and the worst members are measured for diffusion (`avalanche`), which bounds the only other +//! thing a rotation draw could move. A bit-exact verifier never lets a chip skip an application, however weak its +//! diffusion, so diffusion is reported and is not a gain. +//! +//! Commands: +//! attack-f4 census --from 20729 --count 16777216 --threads 12 [--out file] [--dedupe] +//! attack-f4 day --index 20729 one day's draw and classification +//! attack-f4 plant alleq|mul1|mul1all|rc0|rcrk0 the known-fail firings: the day 20729 draw with the planted field +//! attack-f4 expect --threads 12 exact per-word tables (NAF weight, popcount) over all 2^31 odd +//! constants, and the 16-fold convolution: the expected M1 tail +//! attack-f4 avalanche --index 20729 [--states 4096] single-application and two-application diffusion of a day + +use igneum_pow::bind::day_bytes; +use igneum_pow::generator::V4_CLASS; +use igneum_pow::memhard::{mixer, round_key_mult, MixParams, Shape, ITEM_ROUNDS}; +use igneum_pow::seed::{seed_words_from_bytes, SplitMix64}; +use std::fmt::Write as _; +use std::io::Write as _; + +/// The chain's genesis day index (`bind.rs` tests: day_index(0x1a0ff0f7c00) = 20,729, 3 October 2026). +const GENESIS_DAY: u64 = 20_729; +/// Mixer applications per item under class v4 (`Shape::mixers_per_item`): 9 x 8. +const APPLICATIONS_PER_ITEM: u64 = (ITEM_ROUNDS as u64 + 1) * 8; +/// Adds and XORs of one application outside the multiply layer: 8 quarter rounds x (4 adds + 4 xors). +const QR_ADDS_XORS: u32 = 64; +/// The gate's gain threshold (plan 1.4 (3), B5 rank 3). +const GAIN_GATE: f64 = 1.1; +/// M2: a constant of NAF weight at most this is cheaper in LUTs than in a DSP block. +const M2_NAF_CEIL: u32 = 3; + +// ------------------------------------------------------------------------------------------------------------ +// The day +// ------------------------------------------------------------------------------------------------------------ + +fn v4_shape() -> Shape { + let s = Shape::for_class(&V4_CLASS); + assert_eq!(s.mixer_mult, 8, "class v4 is the x8 mixer"); + assert_eq!(s.derive_len, 0, "class v4 has no derivation program: the fixed mixer with drawn constants"); + s +} + +/// The chain day's key and its draw through the real code. +fn params_of_day(d: u64) -> MixParams { + MixParams::with_shape(seed_words_from_bytes(&day_bytes(d)), v4_shape()) +} + +fn seed64(key: &[u32; 8]) -> u64 { + key[0] as u64 | ((key[1] as u64) << 32) +} + +/// Non-adjacent-form weight of a 32-bit constant (the number of nonzero signed digits). +fn naf_weight(v: u32) -> u32 { + let mut n = v as u64; + let mut w = 0; + while n != 0 { + if n & 1 == 1 { + // digit +1 when n = 1 mod 4, -1 when n = 3 mod 4 + if n & 3 == 3 { + n += 1; + } else { + n -= 1; + } + w += 1; + } + n >>= 1; + } + w +} + +/// Round keys of the 72 applications of an item under m = 8: `round_key(r * 8 + j)`, r in 0..=8, j in 0..8. +fn round_keys() -> [u32; 72] { + let mut k = [0u32; 72]; + for r in 0..=ITEM_ROUNDS { + for j in 0..8 { + k[r * 8 + j] = round_key_mult(r, j, 8); + } + } + k +} + +#[derive(Clone, Debug, Default)] +struct DayClass { + // ROT + rot_distinct: u32, + rot_all_equal: bool, + rot_max_mult: u32, + rot_comp_same_word: bool, + rot_comp_any: bool, + rot_small: u32, + rot_byte: u32, + // MUL + naf: [u32; 16], + naf_sum: u32, + naf_min: u32, + mul_one: u32, + mul_minus_one: u32, + mul_pop_le2: u32, + mul_pop_le4: u32, + mul_pop_le8: u32, + mul_naf_le2: u32, + mul_naf_le3: u32, + mul_naf_le4: u32, + mul_small: u32, + mul_dup: bool, + // RC + rc_zero: u32, + rc_pop_ext: u32, + rc_rk_zero: u32, + rc_dup: bool, + // gains + cost_m1: u32, + m2_k: u32, +} + +fn classify(mp: &MixParams, rks: &[u32; 72]) -> DayClass { + let mut c = DayClass::default(); + // ROT + let mut seen = [0u32; 32]; + for &r in &mp.rot { + assert!((1..=31).contains(&r)); + seen[r as usize] += 1; + if matches!(r, 1 | 2 | 30 | 31) { + c.rot_small += 1; + } + if matches!(r, 8 | 16 | 24) { + c.rot_byte += 1; + } + } + c.rot_distinct = seen.iter().filter(|&&n| n > 0).count() as u32; + c.rot_max_mult = *seen.iter().max().unwrap(); + c.rot_all_equal = c.rot_distinct == 1; + // the same-word pairs of a quarter round: s[d] takes ROT[0] then ROT[2], s[b] takes ROT[1] then ROT[3]; the + // diagonal round the same with ROT[4..7] + for (a, b) in [(0, 2), (1, 3), (4, 6), (5, 7)] { + if mp.rot[a] + mp.rot[b] == 32 { + c.rot_comp_same_word = true; + } + } + for a in 0..8 { + for b in a + 1..8 { + if mp.rot[a] + mp.rot[b] == 32 { + c.rot_comp_any = true; + } + } + } + // MUL + c.naf_min = u32::MAX; + for i in 0..16 { + let m = mp.mul[i]; + assert!(m & 1 == 1); + let w = naf_weight(m); + c.naf[i] = w; + c.naf_sum += w; + c.naf_min = c.naf_min.min(w); + let p = m.count_ones(); + if m == 1 { + c.mul_one += 1; + } + if m == u32::MAX { + c.mul_minus_one += 1; + } + if p <= 2 { + c.mul_pop_le2 += 1; + } + if p <= 4 { + c.mul_pop_le4 += 1; + } + if p <= 8 { + c.mul_pop_le8 += 1; + } + if w <= 2 { + c.mul_naf_le2 += 1; + } + if w <= 3 { + c.mul_naf_le3 += 1; + } + if w <= 4 { + c.mul_naf_le4 += 1; + } + if m < 256 { + c.mul_small += 1; + } + for j in 0..i { + if mp.mul[j] == m { + c.mul_dup = true; + } + } + } + c.cost_m1 = QR_ADDS_XORS + c.naf_sum - 16; + c.m2_k = c.mul_naf_le3; + // RC + for i in 0..16 { + let r = mp.rc[i]; + if r == 0 { + c.rc_zero += 1; + } + let p = r.count_ones(); + if p <= 4 || p >= 28 { + c.rc_pop_ext += 1; + } + for &rk in rks.iter() { + if r.wrapping_add(rk) == 0 { + c.rc_rk_zero += 1; + } + } + for j in 0..i { + if mp.rc[j] == r { + c.rc_dup = true; + } + } + } + c +} + +fn m2_gain(k: u32) -> f64 { + 16.0 / (16.0 - k as f64) +} + +// ------------------------------------------------------------------------------------------------------------ +// The tally +// ------------------------------------------------------------------------------------------------------------ + +/// The worst member of a class: the lowest M1 cost among the days in it (ties: the earliest day). +#[derive(Clone, Copy, Debug)] +struct Worst { + day: u64, + cost: u32, +} + +impl Worst { + fn none() -> Self { + Worst { day: u64::MAX, cost: u32::MAX } + } + fn offer(&mut self, day: u64, cost: u32) { + if cost < self.cost || (cost == self.cost && day < self.day) { + *self = Worst { day, cost }; + } + } + fn merge(&mut self, o: &Worst) { + if o.day != u64::MAX { + self.offer(o.day, o.cost); + } + } +} + +const CLASSES: &[&str] = &[ + "ROT all equal", + "ROT distinct <= 3", + "ROT distinct <= 4", + "ROT max multiplicity >= 4", + "ROT same-word pair sums to 32", + "ROT any pair sums to 32", + "ROT all 8 in {1,2,30,31}", + "ROT >= 6 in {1,2,30,31}", + "ROT >= 4 in {1,2,30,31}", + "ROT >= 4 in {8,16,24}", + "MUL any = 1", + "MUL any = 2^32-1", + "MUL any popcount <= 2", + "MUL any popcount <= 4", + "MUL any popcount <= 8", + "MUL any NAF weight <= 2", + "MUL any NAF weight <= 3", + "MUL any NAF weight <= 4", + "MUL any < 256", + "MUL two equal", + "MUL M2 k >= 2 (gain >= 1.14x)", + "RC any = 0", + "RC any popcount <= 4 or >= 28", + "RC + rk = 0 for any of the 72 keys", + "RC two equal", +]; + +#[derive(Clone, Debug)] +struct Tally { + n: u64, + class_count: Vec, + class_worst: Vec, + cost_hist: Vec, + naf_sum_hist: Vec, + naf_min_hist: Vec, + rot_distinct_hist: [u64; 9], + rot_small_hist: [u64; 9], + rot_byte_hist: [u64; 9], + rot_mult_hist: [u64; 9], + m2_k_hist: [u64; 17], + /// The 16 lowest-cost days seen (cost, day), sorted ascending. + lowest: Vec<(u32, u64)>, + highest: Vec<(u32, u64)>, + seeds: Vec, +} + +impl Tally { + fn new(dedupe: bool, cap: usize) -> Self { + Tally { + n: 0, + class_count: vec![0; CLASSES.len()], + class_worst: vec![Worst::none(); CLASSES.len()], + cost_hist: vec![0; 400], + naf_sum_hist: vec![0; 400], + naf_min_hist: vec![0; 40], + rot_distinct_hist: [0; 9], + rot_small_hist: [0; 9], + rot_byte_hist: [0; 9], + rot_mult_hist: [0; 9], + m2_k_hist: [0; 17], + lowest: Vec::new(), + highest: Vec::new(), + seeds: if dedupe { Vec::with_capacity(cap) } else { Vec::new() }, + } + } + + fn flags(c: &DayClass) -> [bool; 25] { + [ + c.rot_all_equal, + c.rot_distinct <= 3, + c.rot_distinct <= 4, + c.rot_max_mult >= 4, + c.rot_comp_same_word, + c.rot_comp_any, + c.rot_small == 8, + c.rot_small >= 6, + c.rot_small >= 4, + c.rot_byte >= 4, + c.mul_one > 0, + c.mul_minus_one > 0, + c.mul_pop_le2 > 0, + c.mul_pop_le4 > 0, + c.mul_pop_le8 > 0, + c.mul_naf_le2 > 0, + c.mul_naf_le3 > 0, + c.mul_naf_le4 > 0, + c.mul_small > 0, + c.mul_dup, + c.m2_k >= 2, + c.rc_zero > 0, + c.rc_pop_ext > 0, + c.rc_rk_zero > 0, + c.rc_dup, + ] + } + + fn add(&mut self, day: u64, mp: &MixParams, c: &DayClass, dedupe: bool) { + self.n += 1; + let f = Self::flags(c); + assert_eq!(f.len(), CLASSES.len()); + for (i, &on) in f.iter().enumerate() { + if on { + self.class_count[i] += 1; + self.class_worst[i].offer(day, c.cost_m1); + } + } + self.cost_hist[c.cost_m1 as usize] += 1; + self.naf_sum_hist[c.naf_sum as usize] += 1; + self.naf_min_hist[c.naf_min as usize] += 1; + self.rot_distinct_hist[c.rot_distinct as usize] += 1; + self.rot_small_hist[c.rot_small as usize] += 1; + self.rot_byte_hist[c.rot_byte as usize] += 1; + self.rot_mult_hist[c.rot_max_mult as usize] += 1; + self.m2_k_hist[c.m2_k as usize] += 1; + push_sorted(&mut self.lowest, (c.cost_m1, day), 16, true); + push_sorted(&mut self.highest, (c.cost_m1, day), 4, false); + if dedupe { + self.seeds.push(seed64(&mp.key)); + } + } + + fn merge(&mut self, o: Tally) { + self.n += o.n; + for i in 0..CLASSES.len() { + self.class_count[i] += o.class_count[i]; + self.class_worst[i].merge(&o.class_worst[i]); + } + for (a, b) in self.cost_hist.iter_mut().zip(o.cost_hist.iter()) { + *a += b; + } + for (a, b) in self.naf_sum_hist.iter_mut().zip(o.naf_sum_hist.iter()) { + *a += b; + } + for (a, b) in self.naf_min_hist.iter_mut().zip(o.naf_min_hist.iter()) { + *a += b; + } + for i in 0..9 { + self.rot_distinct_hist[i] += o.rot_distinct_hist[i]; + self.rot_small_hist[i] += o.rot_small_hist[i]; + self.rot_byte_hist[i] += o.rot_byte_hist[i]; + self.rot_mult_hist[i] += o.rot_mult_hist[i]; + } + for i in 0..17 { + self.m2_k_hist[i] += o.m2_k_hist[i]; + } + for e in o.lowest { + push_sorted(&mut self.lowest, e, 16, true); + } + for e in o.highest { + push_sorted(&mut self.highest, e, 4, false); + } + self.seeds.extend(o.seeds); + } +} + +/// Keep the `cap` smallest (ascending) or largest (descending) entries. +fn push_sorted(v: &mut Vec<(u32, u64)>, e: (u32, u64), cap: usize, ascending: bool) { + if v.len() == cap { + let last = *v.last().unwrap(); + let keep = if ascending { e < last } else { e > last }; + if !keep { + return; + } + v.pop(); + } + let pos = if ascending { v.partition_point(|x| *x < e) } else { v.partition_point(|x| *x > e) }; + v.insert(pos, e); +} + +// ------------------------------------------------------------------------------------------------------------ +// Analytic expectations (per day, independent draws; the modulo-31 bias of `below` is 2^-64 per value and ignored) +// ------------------------------------------------------------------------------------------------------------ + +fn ln_choose(n: u64, k: u64) -> f64 { + let lg = |x: u64| -> f64 { (1..=x).map(|i| (i as f64).ln()).sum() }; + lg(n) - lg(k) - lg(n - k) +} + +fn binom_tail(n: u64, p: f64, k_min: u64) -> f64 { + (k_min..=n).map(|k| (ln_choose(n, k) + (k as f64) * p.ln() + ((n - k) as f64) * (1.0 - p).ln()).exp()).sum() +} + +/// P(d distinct values among 8 uniform draws from 31): S(8, d) x 31 falling d / 31^8. +fn p_rot_distinct(d: u32) -> f64 { + // Stirling numbers of the second kind S(8, d) + let mut s = vec![vec![0f64; 9]; 9]; + s[0][0] = 1.0; + for n in 1..=8 { + for k in 1..=n { + s[n][k] = (k as f64) * s[n - 1][k] + s[n - 1][k - 1]; + } + } + let mut falling = 1.0; + for i in 0..d { + falling *= (31 - i) as f64; + } + s[8][d as usize] * falling / 31f64.powi(8) +} + +/// P(max multiplicity >= 4 among 8 draws from 31), by enumeration of the multiplicity patterns is long; the +/// census gives the exact count, so this is the first-order bound: C(8,4) x 31 x 31^-4 (approximate). +fn p_rot_mult4_approx() -> f64 { + 70.0 * 31.0 / 31f64.powi(4) +} + +fn p_any_of(n: u64, p: f64) -> f64 { + 1.0 - (1.0 - p).powi(n as i32) +} + +fn one_in(p: f64) -> String { + if p <= 0.0 { + "0".into() + } else { + format!("1 in {:.3e}", 1.0 / p) + } +} + +// ------------------------------------------------------------------------------------------------------------ +// Printing a day +// ------------------------------------------------------------------------------------------------------------ + +fn hex_bytes(b: &[u8]) -> String { + b.iter().map(|x| format!("{x:02x}")).collect() +} + +fn describe(day: u64, mp: &MixParams, c: &DayClass, median_cost: Option) -> String { + let mut s = String::new(); + let _ = writeln!(s, "day index {day} (genesis + {}), day bytes {} , seed64 {:016x}", day as i64 - GENESIS_DAY as i64, hex_bytes(&day_bytes(day)), seed64(&mp.key)); + let _ = writeln!(s, "key {}", mp.key.iter().map(|w| format!("{w:08x}")).collect::>().join(" ")); + let _ = writeln!(s, "ROT {:?} distinct {} max multiplicity {} small {} byte-aligned {} same-word pair 32 {} any pair 32 {}", mp.rot, c.rot_distinct, c.rot_max_mult, c.rot_small, c.rot_byte, c.rot_comp_same_word, c.rot_comp_any); + let _ = writeln!(s, "MUL {}", mp.mul.iter().map(|w| format!("{w:08x}")).collect::>().join(" ")); + let _ = writeln!(s, "NAF {:?} sum {} min {} =1 {} =-1 {} pop<=4 {} naf<=3 {} <256 {} dup {}", c.naf, c.naf_sum, c.naf_min, c.mul_one, c.mul_minus_one, c.mul_pop_le4, c.mul_naf_le3, c.mul_small, c.mul_dup); + let _ = writeln!(s, "RC {}", mp.rc.iter().map(|w| format!("{w:08x}")).collect::>().join(" ")); + let _ = writeln!(s, "RC zero {} pop<=4|>=28 {} rc+rk=0 {} dup {}", c.rc_zero, c.rc_pop_ext, c.rc_rk_zero, c.rc_dup); + let _ = writeln!(s, "M1 cost {} adder-equivalents per application ({} per item, 72 applications); M2 k {} (gain {:.3}x)", c.cost_m1, c.cost_m1 as u64 * APPLICATIONS_PER_ITEM, c.m2_k, m2_gain(c.m2_k)); + if let Some(m) = median_cost { + let _ = writeln!(s, "M1 gain against the census median {m}: {:.4}x", m as f64 / c.cost_m1 as f64); + } + let flags: Vec<&str> = Tally::flags(c).iter().zip(CLASSES.iter()).filter(|(f, _)| **f).map(|(_, n)| *n).collect(); + let _ = writeln!(s, "classes: {}", if flags.is_empty() { "none".to_string() } else { flags.join("; ") }); + s +} + +// ------------------------------------------------------------------------------------------------------------ +// Commands +// ------------------------------------------------------------------------------------------------------------ + +fn arg(args: &[String], name: &str) -> Option { + args.iter().position(|a| a == name).and_then(|i| args.get(i + 1).cloned()) +} + +fn census(args: &[String]) { + let from: u64 = arg(args, "--from").map(|v| v.parse().unwrap()).unwrap_or(GENESIS_DAY); + let count: u64 = arg(args, "--count").map(|v| v.parse().unwrap()).unwrap_or(1 << 24); + let threads: usize = arg(args, "--threads").map(|v| v.parse().unwrap()).unwrap_or(12); + let out = arg(args, "--out"); + let dedupe = args.iter().any(|a| a == "--dedupe"); + let shape = v4_shape(); + let rks = round_keys(); + let t0 = std::time::Instant::now(); + let chunk = count.div_ceil(threads as u64); + let tallies: Vec = std::thread::scope(|sc| { + let hs: Vec<_> = (0..threads) + .map(|t| { + let rks = &rks; + sc.spawn(move || { + let lo = from + chunk * t as u64; + let hi = (lo + chunk).min(from + count); + let mut tally = Tally::new(dedupe, (hi.saturating_sub(lo)) as usize); + for d in lo..hi { + let mp = params_of_day(d); + debug_assert_eq!(mp.shape, shape); + let c = classify(&mp, rks); + tally.add(d, &mp, &c, dedupe); + } + tally + }) + }) + .collect(); + hs.into_iter().map(|h| h.join().unwrap()).collect() + }); + let mut all = Tally::new(false, 0); + for t in tallies { + all.merge(t); + } + let secs = t0.elapsed().as_secs_f64(); + assert_eq!(all.n, count); + + // the median cost and the gate fractions + let total = all.n; + let mut acc = 0u64; + let mut median = 0u32; + for (c, &n) in all.cost_hist.iter().enumerate() { + acc += n; + if acc * 2 >= total { + median = c as u32; + break; + } + } + let mean = all.cost_hist.iter().enumerate().map(|(c, &n)| c as f64 * n as f64).sum::() / total as f64; + let var = all.cost_hist.iter().enumerate().map(|(c, &n)| (c as f64 - mean).powi(2) * n as f64).sum::() / total as f64; + let gate_cost = |reference: f64| -> u32 { (reference / GAIN_GATE).floor() as u32 }; // cost <= this gives gain >= 1.1x... strictly over: cost < reference/1.1 + let over = |reference: f64| -> u64 { + all.cost_hist.iter().enumerate().filter(|(c, _)| (*c as f64) * GAIN_GATE < reference).map(|(_, &n)| n).sum() + }; + let over_median = over(median as f64); + let over_mean = over(mean); + let m2_over: u64 = all.m2_k_hist.iter().enumerate().filter(|(k, _)| m2_gain(*k as u32) > GAIN_GATE).map(|(_, &n)| n).sum(); + + let mut r = String::new(); + let _ = writeln!(r, "# attack-f4 census: {count} chain days from day index {from} (genesis {GENESIS_DAY}), class v4 shape (mixer x{}, cache 2^{}), {threads} threads, {secs:.1} s", shape.mixer_mult, shape.cache_log2_words); + let _ = writeln!(r, "draw: MixParams::with_shape(seed_words_from_bytes(bind::day_bytes(d)), Shape::for_class(&V4_CLASS)); seed64 = key[0] | key[1] << 32"); + let _ = writeln!(r); + let _ = writeln!(r, "## Gate (plan 1.4 (3), B5 rank 3): fraction of days with any gain over {GAIN_GATE}x under 2^-20 = {:.3e}", 2f64.powi(-20)); + let _ = writeln!(r); + let _ = writeln!(r, "| Metric | Reference | Days over {GAIN_GATE}x | Fraction | Against 2^-20 |"); + let _ = writeln!(r, "|---|---|---|---|---|"); + let frac = |n: u64| n as f64 / total as f64; + let vs = |n: u64| if frac(n) < 2f64.powi(-20) { "under" } else { "OVER" }; + let _ = writeln!(r, "| M1 per-day LUT datapath, adders per application | median cost {median} (gain over {GAIN_GATE}x = cost under {}) | {over_median} | {:.3e} | {} |", gate_cost(median as f64) + 1, frac(over_median), vs(over_median)); + let _ = writeln!(r, "| M1 against the mean cost {mean:.2} (sd {:.2}) | cost under {:.2} | {over_mean} | {:.3e} | {} |", var.sqrt(), mean / GAIN_GATE, frac(over_mean), vs(over_mean)); + let _ = writeln!(r, "| M2 DSP-bound datapath, 16/(16-k) with k = words of NAF weight <= {M2_NAF_CEIL} | k >= 2 | {m2_over} | {:.3e} | {} |", frac(m2_over), vs(m2_over)); + let _ = writeln!(r, "| ROT value (wiring) and RC value (inverters) on a per-day datapath | exact 0 ops moved on every day | 0 | 0 | under |"); + let _ = writeln!(r); + let _ = writeln!(r, "M1 cost = {QR_ADDS_XORS} + sum(NAF(MUL_i) - 1); mean {mean:.3}, sd {:.3}, median {median}, min {} (day {}), max {} (day {})", var.sqrt(), all.lowest[0].0, all.lowest[0].1, all.highest[0].0, all.highest[0].1); + let _ = writeln!(r); + + // the classes + let p_mul1 = p_any_of(16, 2f64.powi(-31)); + let p_small = p_any_of(16, 128.0 / 2f64.powi(31)); + let p_rc0 = p_any_of(16, 2f64.powi(-32)); + let p_rcrk = p_any_of(16, 72.0 / 2f64.powi(32)); + let pop_le = |k: u32| -> f64 { (0..=k).map(|i| ln_choose(32, i as u64).exp()).sum::() }; + let p_rc_pop = p_any_of(16, 2.0 * pop_le(4) / 2f64.powi(32)); + // odd constants with popcount <= k: the low bit is set, so C(31, i) for the other i < k bits + let odd_pop_le = |k: u32| -> f64 { (0..k).map(|i| ln_choose(31, i as u64).exp()).sum::() / 2f64.powi(31) }; + let p_dup16 = |space: f64| -> f64 { 1.0 - (1..16).map(|i| 1.0 - i as f64 / space).product::() }; + let expectations: Vec> = vec![ + Some(31f64.powi(-7)), + Some((1..=3).map(p_rot_distinct).sum()), + Some((1..=4).map(p_rot_distinct).sum()), + Some(p_rot_mult4_approx()), + Some(p_any_of(4, 1.0 / 31.0)), + Some(p_any_of(28, 1.0 / 31.0)), + Some((4f64 / 31.0).powi(8)), + Some(binom_tail(8, 4.0 / 31.0, 6)), + Some(binom_tail(8, 4.0 / 31.0, 4)), + Some(binom_tail(8, 3.0 / 31.0, 4)), + Some(p_mul1), + Some(p_mul1), + Some(p_any_of(16, odd_pop_le(2))), + Some(p_any_of(16, odd_pop_le(4))), + Some(p_any_of(16, odd_pop_le(8))), + None, + None, + None, + Some(p_small), + Some(p_dup16(2f64.powi(31))), + None, + Some(p_rc0), + Some(p_rc_pop), + Some(p_rcrk), + Some(p_dup16(2f64.powi(32))), + ]; + let _ = writeln!(r, "## Classes over {count} days"); + let _ = writeln!(r); + let _ = writeln!(r, "| Class | Count | Fraction | Expected per day (analytic) | Expected count | Worst member (day, M1 cost, M1 gain vs median, M2 gain) |"); + let _ = writeln!(r, "|---|---|---|---|---|---|"); + for (i, name) in CLASSES.iter().enumerate() { + let n = all.class_count[i]; + let w = all.class_worst[i]; + let worst = if w.day == u64::MAX { + "none".to_string() + } else { + let c = classify(¶ms_of_day(w.day), &rks); + format!("day {} , cost {} , {:.4}x , {:.3}x", w.day, w.cost, median as f64 / w.cost as f64, m2_gain(c.m2_k)) + }; + let (ep, ec) = match expectations[i] { + Some(p) => (format!("{p:.3e} ({})", one_in(p)), format!("{:.3}", p * total as f64)), + None => ("see `expect`".to_string(), "see `expect`".to_string()), + }; + let _ = writeln!(r, "| {name} | {n} | {:.3e} | {ep} | {ec} | {worst} |", frac(n)); + } + let _ = writeln!(r); + let _ = writeln!(r, "## Histograms"); + let _ = writeln!(r); + let _ = writeln!(r, "| ROT distinct amounts | Count | Fraction | Expected (S(8,d) 31_d / 31^8) |"); + let _ = writeln!(r, "|---|---|---|---|"); + for d in 1..=8 { + let _ = writeln!(r, "| {d} | {} | {:.4e} | {:.4e} |", all.rot_distinct_hist[d], frac(all.rot_distinct_hist[d]), p_rot_distinct(d as u32)); + } + let _ = writeln!(r); + let _ = writeln!(r, "| ROT amounts in {{1,2,30,31}} | Count | Expected Binomial(8, 4/31) | ROT amounts in {{8,16,24}} | Count | Expected Binomial(8, 3/31) | ROT max multiplicity | Count |"); + let _ = writeln!(r, "|---|---|---|---|---|---|---|---|"); + for k in 0..=8 { + let e1 = binom_tail(8, 4.0 / 31.0, k) - binom_tail(8, 4.0 / 31.0, k + 1); + let e2 = binom_tail(8, 3.0 / 31.0, k) - binom_tail(8, 3.0 / 31.0, k + 1); + let _ = writeln!(r, "| {k} | {} | {:.1} | {k} | {} | {:.1} | {k} | {} |", all.rot_small_hist[k as usize], e1 * total as f64, all.rot_byte_hist[k as usize], e2 * total as f64, all.rot_mult_hist[k as usize]); + } + let _ = writeln!(r); + let _ = writeln!(r, "| M2 k (words of NAF weight <= {M2_NAF_CEIL}) | Count | Gain 16/(16-k) |"); + let _ = writeln!(r, "|---|---|---|"); + for k in 0..=16 { + if all.m2_k_hist[k] > 0 || k <= 3 { + let _ = writeln!(r, "| {k} | {} | {:.3}x |", all.m2_k_hist[k], m2_gain(k as u32)); + } + } + let _ = writeln!(r); + let _ = writeln!(r, "| Minimum NAF weight over the 16 MUL | Count |"); + let _ = writeln!(r, "|---|---|"); + for (w, &n) in all.naf_min_hist.iter().enumerate() { + if n > 0 { + let _ = writeln!(r, "| {w} | {n} |"); + } + } + let _ = writeln!(r); + let _ = writeln!(r, "| M1 cost per application | Count | Cumulative fraction | M1 gain vs median |"); + let _ = writeln!(r, "|---|---|---|---|"); + let mut cum = 0u64; + for (c, &n) in all.cost_hist.iter().enumerate() { + if n > 0 { + cum += n; + let _ = writeln!(r, "| {c} | {n} | {:.4e} | {:.4}x |", frac(cum), median as f64 / c as f64); + } + } + let _ = writeln!(r); + let _ = writeln!(r, "## The 16 lowest-cost days (M1)"); + let _ = writeln!(r); + for &(cost, day) in &all.lowest { + let mp = params_of_day(day); + let c = classify(&mp, &rks); + let _ = writeln!(r, "- day {day}: cost {cost}, gain {:.4}x vs median, NAF sum {}, M2 k {}, day-hex {}", median as f64 / cost as f64, c.naf_sum, c.m2_k, hex_bytes(&day_bytes(day))); + } + if dedupe { + all.seeds.sort_unstable(); + let before = all.seeds.len(); + all.seeds.dedup(); + let _ = writeln!(r); + let _ = writeln!(r, "## 64-bit seeding: {} days, {} distinct 64-bit stream seeds ({} collisions; expected C(n,2)/2^64 = {:.3e})", before, all.seeds.len(), before - all.seeds.len(), (before as f64) * (before as f64 - 1.0) / 2.0 / 2f64.powi(64)); + } + let _ = writeln!(r); + let _ = writeln!(r, "## Worst member of every class, in full"); + let _ = writeln!(r); + let mut shown: Vec = Vec::new(); + for (i, name) in CLASSES.iter().enumerate() { + let w = all.class_worst[i]; + if w.day == u64::MAX || shown.contains(&w.day) { + continue; + } + shown.push(w.day); + let mp = params_of_day(w.day); + let c = classify(&mp, &rks); + let _ = writeln!(r, "### {name}: day {}", w.day); + let _ = writeln!(r, "```"); + let _ = write!(r, "{}", describe(w.day, &mp, &c, Some(median))); + let _ = writeln!(r, "```"); + } + print!("{r}"); + if let Some(p) = out { + std::fs::File::create(&p).unwrap().write_all(r.as_bytes()).unwrap(); + eprintln!("written {p}"); + } +} + +fn day_cmd(args: &[String]) { + let d: u64 = arg(args, "--index").map(|v| v.parse().unwrap()).unwrap_or(GENESIS_DAY); + let median: Option = arg(args, "--median").map(|v| v.parse().unwrap()); + let mp = params_of_day(d); + let c = classify(&mp, &round_keys()); + print!("{}", describe(d, &mp, &c, median)); +} + +/// The known-fail firings: the genesis day's draw with one field forced through this crate's own hook (the +/// `MixParams` fields are public; `igneum-pow` is untouched). Exit 0 when the classifier flags the plant and the +/// gain metric that the plant moves reads over the gate. +fn plant(args: &[String]) { + let what = args.get(2).cloned().unwrap_or_default(); + let median: u32 = arg(args, "--median").map(|v| v.parse().unwrap()).unwrap_or(221); + let rks = round_keys(); + let mut mp = params_of_day(GENESIS_DAY); + let before = classify(&mp, &rks); + println!("before the plant (day {GENESIS_DAY}):"); + print!("{}", describe(GENESIS_DAY, &mp, &before, Some(median))); + let (flag_name, gain_metric): (&str, &str) = match what.as_str() { + "alleq" => { + mp.rot = [7; 8]; + ("ROT all equal", "structure (0 ops on a per-day datapath by construction; diffusion in `avalanche`)") + } + "mul1" => { + mp.mul[5] = 1; + ("MUL any = 1", "M1") + } + "mul1all" => { + mp.mul = [1; 16]; + ("MUL any = 1", "M1") + } + "mulnaf" => { + // the lightest realistic plant: four words at NAF weight 3 (M2 k = 4), the rest untouched + for i in 0..4 { + mp.mul[i] = (1u32 << 20) + (1u32 << 9) + 1; + } + ("MUL any NAF weight <= 3", "M1 and M2") + } + "rc0" => { + mp.rc[3] = 0; + ("RC any = 0", "structure (0 ops on a per-day datapath by construction)") + } + "rcrk0" => { + mp.rc[3] = 0u32.wrapping_sub(round_key_mult(2, 5, 8)); + ("RC + rk = 0 for any of the 72 keys", "structure (one xor of 10,368 ops per item on a generic datapath: 1.0001x)") + } + _ => { + eprintln!("plant alleq|mul1|mul1all|mulnaf|rc0|rcrk0"); + std::process::exit(2); + } + }; + let after = classify(&mp, &rks); + println!("\nafter the plant `{what}`:"); + print!("{}", describe(GENESIS_DAY, &mp, &after, Some(median))); + let idx = CLASSES.iter().position(|n| *n == flag_name).unwrap(); + let flagged = Tally::flags(&after)[idx] && !Tally::flags(&before)[idx]; + let gain_m1 = median as f64 / after.cost_m1 as f64; + let gain_m2 = m2_gain(after.m2_k); + println!("\nplant `{what}`: classifier flag `{flag_name}` {} (was {} before); M1 gain {gain_m1:.4}x, M2 gain {gain_m2:.4}x; gain metric for this plant: {gain_metric}", if flagged { "FIRED" } else { "did NOT fire" }, Tally::flags(&before)[idx]); + let gain_fired = gain_m1 > GAIN_GATE || gain_m2 > GAIN_GATE; + println!("gain over {GAIN_GATE}x: {}", if gain_fired { "FIRED" } else { "not over the gate" }); + if !flagged { + std::process::exit(1); + } +} + +/// Exact per-word tables over every odd 32-bit constant (2^31 of them): the NAF weight distribution and the +/// popcount distribution; then the 16-fold convolution of the NAF-weight distribution gives the expected M1 cost +/// distribution and the expected fraction of days under any cost. +fn expect(args: &[String]) { + let threads: usize = arg(args, "--threads").map(|v| v.parse().unwrap()).unwrap_or(12); + let median: u32 = arg(args, "--median").map(|v| v.parse().unwrap()).unwrap_or(221); + let t0 = std::time::Instant::now(); + let per: Vec<([u64; 40], [u64; 33])> = std::thread::scope(|sc| { + let hs: Vec<_> = (0..threads) + .map(|t| { + sc.spawn(move || { + let mut naf = [0u64; 40]; + let mut pop = [0u64; 33]; + let lo = ((1u64 << 32) * t as u64 / threads as u64) | 1; + let hi = (1u64 << 32) * (t as u64 + 1) / threads as u64; + let mut v = lo; + while v < hi { + naf[naf_weight(v as u32) as usize] += 1; + pop[(v as u32).count_ones() as usize] += 1; + v += 2; + } + (naf, pop) + }) + }) + .collect(); + hs.into_iter().map(|h| h.join().unwrap()).collect() + }); + let mut naf = [0u64; 40]; + let mut pop = [0u64; 33]; + for (a, b) in per { + for i in 0..40 { + naf[i] += a[i]; + } + for i in 0..33 { + pop[i] += b[i]; + } + } + let total: u64 = naf.iter().sum(); + assert_eq!(total, 1 << 31); + println!("# attack-f4 expect: all {total} odd 32-bit constants, {threads} threads, {:.1} s", t0.elapsed().as_secs_f64()); + println!(); + println!("| NAF weight | Odd constants | Fraction | Cumulative | Any of 16 per day | x 2^24 days |"); + println!("|---|---|---|---|---|---|"); + let mut cum = 0u64; + for w in 0..40 { + if naf[w] > 0 { + cum += naf[w]; + let p = cum as f64 / total as f64; + println!("| {w} | {} | {:.4e} | {:.4e} | {:.4e} | {:.3} |", naf[w], naf[w] as f64 / total as f64, p, p_any_of(16, p), p_any_of(16, p) * 2f64.powi(24)); + } + } + println!(); + println!("| Popcount | Odd constants | Cumulative fraction | Any of 16 per day |"); + println!("|---|---|---|---|"); + cum = 0; + for w in 0..33 { + if pop[w] > 0 { + cum += pop[w]; + let p = cum as f64 / total as f64; + println!("| {w} | {} | {:.4e} | {:.4e} |", pop[w], p, p_any_of(16, p)); + } + } + // the 16-fold convolution of the NAF-weight distribution: the exact expected distribution of the NAF sum + let pw: Vec = naf.iter().map(|&n| n as f64 / total as f64).collect(); + let mut dist = vec![0f64; 1]; + dist[0] = 1.0; + for _ in 0..16 { + let mut next = vec![0f64; dist.len() + 39]; + for (i, &a) in dist.iter().enumerate() { + if a == 0.0 { + continue; + } + for (w, &b) in pw.iter().enumerate() { + next[i + w] += a * b; + } + } + dist = next; + } + let mean: f64 = dist.iter().enumerate().map(|(s, &p)| s as f64 * p).sum(); + let var: f64 = dist.iter().enumerate().map(|(s, &p)| (s as f64 - mean).powi(2) * p).sum(); + println!(); + println!("Expected NAF sum over 16 words: mean {mean:.4}, sd {:.4}; expected M1 cost mean {:.4}", var.sqrt(), mean + QR_ADDS_XORS as f64 - 16.0); + println!(); + println!("| M1 cost | NAF sum | Expected fraction of days at this cost | Expected cumulative fraction | Gain vs median {median} |"); + println!("|---|---|---|---|---|"); + let mut c = 0f64; + for (s, &p) in dist.iter().enumerate() { + let cost = s as i64 + QR_ADDS_XORS as i64 - 16; + if cost < 0 { + continue; + } + c += p; + if p > 1e-12 && (cost as f64) <= median as f64 { + println!("| {cost} | {s} | {p:.4e} | {c:.4e} | {:.4}x |", median as f64 / cost as f64); + } + } + let gate: f64 = dist.iter().enumerate().filter(|(s, _)| ((*s as f64) + QR_ADDS_XORS as f64 - 16.0) * GAIN_GATE < median as f64).map(|(_, &p)| p).sum(); + println!(); + println!("Expected fraction of days with M1 gain over {GAIN_GATE}x against median {median}: {gate:.4e} ({} ; x 2^24 = {:.1}); 2^-20 = {:.4e}", one_in(gate), gate * 2f64.powi(24), 2f64.powi(-20)); +} + +/// Diffusion of one day's mixer: for `states` random 16-word states and each of the 512 input bits, the fraction of +/// the 512 output bits that flip after k = 1 and k = 2 applications (keys `round_key_mult(0, 0, 8)` and `(0, 1, 8)`), +/// mean over all, and the minimum per-output-bit flip probability. An ideal mixer reads 0.5 mean and about 0.5 min. +fn avalanche(args: &[String]) { + let d: u64 = arg(args, "--index").map(|v| v.parse().unwrap()).unwrap_or(GENESIS_DAY); + let states: usize = arg(args, "--states").map(|v| v.parse().unwrap()).unwrap_or(4096); + let plant_alleq: Option = arg(args, "--plant-alleq").map(|v| v.parse().unwrap()); + let mut mp = params_of_day(d); + if let Some(r) = plant_alleq { + mp.rot = [r; 8]; + } + let c = classify(&mp, &round_keys()); + println!("avalanche of day {d}{}: ROT {:?}, NAF sum {}, {states} states x 512 input bits", plant_alleq.map(|r| format!(" with ROT planted all {r}")).unwrap_or_default(), mp.rot, c.naf_sum); + let mut rng = SplitMix64::new(0xF4F4_F4F4 ^ d); + for k in 1..=3usize { + let apply = |s: &mut [u32; 16]| { + for j in 0..k { + mixer(s, round_keys()[j], &mp); + } + }; + let mut flips = [0u64; 512]; + let mut total_flips = 0u64; + let mut trials = 0u64; + for _ in 0..states { + let mut base = [0u32; 16]; + for w in base.iter_mut() { + *w = rng.next() as u32; + } + let mut y0 = base; + apply(&mut y0); + for bit in 0..512 { + let mut x = base; + x[bit / 32] ^= 1 << (bit % 32); + apply(&mut x); + for o in 0..512 { + let f = ((x[o / 32] ^ y0[o / 32]) >> (o % 32)) & 1; + flips[o] += f as u64; + total_flips += f as u64; + } + trials += 1; + } + } + let mean = total_flips as f64 / (trials as f64 * 512.0); + let min = flips.iter().map(|&f| f as f64 / trials as f64).fold(1.0, f64::min); + let max = flips.iter().map(|&f| f as f64 / trials as f64).fold(0.0, f64::max); + println!("| {k} application{} | mean flip {mean:.4} | min per output bit {min:.4} | max {max:.4} |", if k > 1 { "s" } else { "" }); + } +} + +fn main() { + let args: Vec = std::env::args().collect(); + match args.get(1).map(|s| s.as_str()) { + Some("census") => census(&args), + Some("day") => day_cmd(&args), + Some("plant") => plant(&args), + Some("expect") => expect(&args), + Some("avalanche") => avalanche(&args), + _ => { + eprintln!("attack-f4 census|day|plant|expect|avalanche (see the module doc)"); + std::process::exit(2); + } + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn naf_weights() { + assert_eq!(naf_weight(0), 0); + assert_eq!(naf_weight(1), 1); + assert_eq!(naf_weight(3), 2); // 4 - 1 + assert_eq!(naf_weight(7), 2); // 8 - 1 + assert_eq!(naf_weight(0xFFFF_FFFF), 2); // 2^32 - 1 + assert_eq!(naf_weight(0xAAAA_AAAB), 17); // alternating odd: the maximum for 32 bits + assert_eq!(naf_weight((1 << 20) + (1 << 9) + 1), 3); + } + + #[test] + fn genesis_day_draw_matches_memhard_md() { + // MEMHARD.md section 1.1 (string day "2026-10-03") is a different key from the chain's day index 20,729; + // what is checked here is that the chain-day path draws through the real code and stays in range. + let mp = params_of_day(GENESIS_DAY); + assert!(mp.rot.iter().all(|r| (1..=31).contains(r))); + assert!(mp.mul.iter().all(|m| m & 1 == 1)); + assert_eq!(mp.shape, v4_shape()); + let s = igneum_pow::seed::day_key("2026-10-03"); + let mp2 = MixParams::with_shape(s, Shape::V2); + assert_eq!(mp2.rot, [20, 20, 19, 4, 26, 3, 3, 27], "MEMHARD.md 1.1 genesis-day ROT"); + } +}