The F3 record and its harness crate (tools/attack/f3-cache: the extracted chain, the exhaustive closure search, the pebbling cross-check, the store-set DP and brute force, the two planted broken chains). The optimal-placement observation (3.17 blocks per read at f=1/8, 16.0 at f=1/64) noted against funding.md B2. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
148 lines
17 KiB
Markdown
148 lines
17 KiB
Markdown
# F3: the chained cache's j + 1 bound and the storage-against-recompute curve
|
|
|
|
Attack-pass row F3 of `docs/plans/cryptanalysis.md` section 4.2 (the record is `docs/analysis/attack-pass-2026-10.md`). Run 7 October 2026, 09:10 to 09:12 UK (08:10 to 08:12 UTC in the logs), on igneum-build-1. Verdict: PASS on all three gate clauses. No line (s, j) is derivable in fewer than j + 1 block evaluations without an earlier line, by an exhaustive search over the block dependency graph extracted from the code at 64 and 1,024 lines, cross-checked by an exhaustive pebbling search over every configuration at 10 lines. The storage-against-recompute curve over cache lines is monotone from f = 1/64 to 1. The f = 1 point is unchanged.
|
|
|
|
## Target
|
|
|
|
| Item | Value |
|
|
|---|---|
|
|
| Commit | 924288d1 (the brief); the worktree HEAD moved to 11b375a0 during the run; `igneum-pow/src/memhard.rs` is byte-identical at both (blob ad42470b, `git diff --stat 924288d1 HEAD -- igneum-pow/src/memhard.rs` empty) |
|
|
| Construction, from the code | `Cache::fill_segment`: `in_j = prev XOR (sigma || K || seg || j || tag)`, `line_j = chacha_block(in_j)` where `chacha_block(x) = ChaCha12core(x) + x`, `prev_0 = 0`, `prev_j = line_{j-1}`; 64 lines per segment, 2^16 segments, 2^26 words (256 MiB) |
|
|
| Reads | `derive_items_mask`: 8 dependent reads per item at line index `s[0] AND mask`, so the segment and j of a read are uniform over the 2^22 lines (F8 checks the uniformity) |
|
|
| Known-failed shape | a line (s, j) computable in fewer than j + 1 block evaluations without an earlier line of segment s (the address-steering shape of the MTP break, Dinur and Nadler 2017, needs a data-dependent chain; this chain's inputs are fixed by the key, so the shape to search is a structural shortcut on the dependency graph) |
|
|
| Gate | no derivation under j + 1 blocks; the curve monotone; the f = 1 point unchanged |
|
|
| Prior evidence | none to re-gate: the `ca2-cache` branch named in the status board is the hot-table experiment (`docs/plans/hot-table.md`), not a chain analysis |
|
|
|
|
## Method
|
|
|
|
The model is the code, not the prose. `tools/attack/f3-cache/src/main.rs` runs one chain function, written in the shape of `memhard.rs` (the quarter round, the 6 double rounds, the feed-forward, the prev XOR, the constant block), generically over two word types:
|
|
|
|
| Word type | What it computes | Use |
|
|
|---|---|---|
|
|
| `u32` | the real arithmetic | `verify`: bit-exact against `Cache::fill_segment` on 16 (key, segment) pairs and against `chacha_block` on 100,000 random inputs |
|
|
| taint set | which block outputs a value depends on (add, xor, rotate = union) | `search`: the direct-parent graph of every block, with each computed line relabelled to the single node {j} so parents are direct, not transitive; plus the 16 x 16 (output word, input word) dependency matrix of one block |
|
|
|
|
The exhaustive search: for every target line j, the minimum number of block evaluations with nothing stored is the size of the backward closure of j on the extracted graph (every non-stored block in the closure must be evaluated at least once; once each in dependency order suffices). `pebble` checks that formula against an exhaustive 0-1 BFS over every pebble configuration (place on a node whose parents are pebbled at cost 1, remove at cost 0) for all 2^10 stored sets x 10 targets on each of the three graphs: 10,240 pairs per graph, 0 mismatches. Two deliberately broken chains are the known-fail cases: `skip2` (line j fed from line j - 2) and `nofeed` (no previous line fed in). The curve: for f = 1/64 to 1 (fraction of cache LINES held), the blocks per read on the naive pattern of `funding.md` B2 rank 2 (every L/n-th line from line 0) and on the optimal pattern (exact DP over chunk lengths; brute force over every C(64, n) set for n up to 8, 4,426,165,368 sets at n = 8); ops per item = 9,360 mixer ops (chip-model-v3.md 5.2) + 8 reads x blocks per read x ops per block (608 counted from the code: 48 quarter rounds x 12, 16 feed-forward adds, 16 input XORs; also at MEMHARD.md's approximate 700). Three more checks on the real function: single-bit avalanche and a differential-independence test on `chacha_block`, a census of every line of the real 2^22-line cache for the day key 2026-10-03, and one-core timings of a block, a mixer application, an item and a line recompute.
|
|
|
|
## Harness
|
|
|
|
| Item | Value |
|
|
|---|---|
|
|
| Crate | `/Users/joshm/Projects/igneum-wt-attack/tools/attack/f3-cache/` (`Cargo.toml` with `igneum-pow = { path = "../../../igneum-pow" }` and an empty `[workspace]`; `src/main.rs`; `run-box.sh`) |
|
|
| Build | `cd tools/attack/f3-cache && IGNEUM_AGENT=attack-f3 IGNEUM_TOOLCHAIN_MISMATCH=ok bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f3" --out <scratchpad>/attack-f3 -- build --release` (rc 0, 38 s wall, 0 warnings; the Mac's PATH rustc is 1.69 but `~/.cargo/bin/rustc` is 1.99.0, which the script read as "on both sides"; log `<scratchpad>/attack-f3/build-1.log`) |
|
|
| Binary | box `/srv/builds/igneum-wt-attack/tools/attack/f3-cache/target/release/attack-f3`, sha256 975115a385ca3195...71cc33, 563,104 bytes |
|
|
| Run | on the box: `nohup bash run-box.sh r1 > run-r1.log 2>&1 &` from `/srv/builds/igneum-wt-attack/attack-f3/`; every phase as `flock -s /srv/builds/_locks/measure -c "nice -n 10 taskset -c 12-15,60-63 attack-f3 <cmd>"`, one chunk per phase, the whole run 17 s (08:10:58 to 08:11:15 UTC; box load 54 at start) |
|
|
| Phase lines | `verify`; `search --lines 64|1024 --variant real|skip2|nofeed`; `pebble --lines 10`; `store --lines 64 --brute-max 8`; `store --lines 1024 --brute-max 2`; `curve --lines 64`; `curve --lines 64 --ops-block 700`; `curve --lines 1024`; `avalanche --samples 1048576`; `census --day 2026-10-03`; `bench --n 20000000` |
|
|
| Logs | box `/srv/builds/igneum-wt-attack/attack-f3/run-r1.log` and `r1-<phase>.log`; Mac copies `/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/attack-f3/` |
|
|
|
|
Box hygiene: the box checkout of every `build-remote.sh` run on this worktree executes `git clean -fd` at `/srv/builds/igneum-wt-attack` (remote-run.sh `checkout_tree`), which deletes any untracked scratch directory there. `attack-f3/` and `attack-f3-venv/` are listed in that mirror's `.git/info/exclude` so they survive; nothing in the tree was touched. The F1 lane's `attack-f1-venv/` is untracked and unprotected and will be removed by the next build from any agent on this worktree.
|
|
|
|
## The two firings and the pass
|
|
|
|
| Chain | Direct parents (taint trace) | Lines under j + 1 at 64 lines | Cheapest derivations | Exhaustive pebbling at 10 lines, cost per target | Verdict | Log |
|
|
|---|---|---|---|---|---|---|
|
|
| real (the code) | j - 1 for all 63 lines after line 0 | 0 of 64 | none; every line costs exactly j + 1 (mean 32.5) | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 | PASS | `r1-search-real-64.log`, `r1-pebble-10.log` |
|
|
| real, 1,024-line model | j - 1 for all 1,023 lines after line 0 | 0 of 1,024 | none (mean 512.5) | same graph rule | PASS | `r1-search-real-1024.log` |
|
|
| skip2 (known fail A) | j - 2 for 62 lines, none for 2 | 63 of 64 | j = 1 in 1, j = 63 in 32 (mean 16.5) | 1, 1, 2, 2, 3, 3, 4, 4, 5, 5 | FIRE | `r1-search-skip2-64.log`, `r1-search-skip2-1024.log` |
|
|
| nofeed (known fail B) | none for all 64 | 63 of 64 | every line in 1 block (mean 1.0) | 1 x 10 | FIRE | `r1-search-nofeed-64.log`, `r1-search-nofeed-1024.log` |
|
|
|
|
`verify` (`r1-verify.log`): the model chain equals `Cache::fill_segment` on keys {day 2026-10-03, 3 random} x segments {0, 1, 12345, 65535} (16 of 16), `block == chacha_block` on 100,000 of 100,000 random inputs, the 1,024-line model's first 64 lines equal the 64-line chain, and both broken variants differ from the real chain from line 1 (line 0 equal, as the rule predicts). The block's word dependency matrix is full on every variant (256 of 256 pairs), so the firings come from the chain rule alone.
|
|
|
|
## Derivation cost per line on the model segment (real chain, nothing stored)
|
|
|
|
| j | blocks to derive line j | j + 1 | Log |
|
|
|---|---|---|---|
|
|
| 0 | 1 | 1 | `r1-search-real-1024.log` |
|
|
| 1 | 2 | 2 | |
|
|
| 3 | 4 | 4 | |
|
|
| 7 | 8 | 8 | |
|
|
| 15 | 16 | 16 | |
|
|
| 31 | 32 | 32 | |
|
|
| 63 | 64 | 64 | (the last line of a real segment; `r1-search-real-64.log` lists all 64) |
|
|
| 127 | 128 | 128 | |
|
|
| 255 | 256 | 256 | |
|
|
| 511 | 512 | 512 | |
|
|
| 1,023 | 1,024 | 1,024 | |
|
|
|
|
All 1,024 lines were searched (0 under j + 1, mean 512.5 = (L + 1) / 2); the 64-line table in `r1-search-real-64.log` has every j from 0 to 63 at exactly j + 1.
|
|
|
|
## Store patterns on the real 64-line segment
|
|
|
|
Blocks per read averaged over j uniform in 0..63. "Naive" is `funding.md` B2 rank 2's pattern (every k-th line from line 0). "Optimal" is the exact minimum over store sets of that size (DP; brute force over every set for n up to 8, agreeing with the DP on every row it ran). The gap formula equals the closure cost on the extracted graph on 2,000 of 2,000 random stored sets (`r1-store-64.log`).
|
|
|
|
| f | Stored lines n | SRAM held | Naive blocks per read | Optimal positions | Optimal blocks per read | Brute force over C(64, n) sets |
|
|
|---|---|---|---|---|---|---|
|
|
| 1/64 | 1 | 4 MiB | 31.5 | [32] | 16.0 | 16.0 (64 sets) |
|
|
| 1/32 | 2 | 8 MiB | 15.5 | [21, 43] | 10.5 | 10.5 (2,016 sets) |
|
|
| 1/16 | 4 | 16 MiB | 7.5 | [12, 25, 38, 51] | 6.094 | 6.094 (635,376 sets) |
|
|
| 1/8 | 8 | 32 MiB | 3.5 | [7, 15, 22, 29, 36, 43, 50, 57] | 3.172 | 3.172 (4,426,165,368 sets, 11.2 s) |
|
|
| 1/4 | 16 | 64 MiB | 1.5 | [3, 7, 11, ..., 55, 58, 61] | 1.453 | not run (DP exact) |
|
|
| 1/2 | 32 | 128 MiB | 0.5 | odd lines | 0.5 | not run |
|
|
| 1 | 64 | 256 MiB | 0 | all | 0 | not run |
|
|
|
|
The 1,024-line model (`r1-store-1024.log`) gives 29.68 / 15.05 / 7.40 / 3.48 / 1.50 / 0.5 / 0 at the same f on the optimal pattern: the naive and optimal patterns converge as the chain lengthens, because the wasted stored line 0 and the end effects are a smaller share.
|
|
|
|
## The curve: ops per item against the fraction of cache lines held (real 64-line segment)
|
|
|
|
Ops per item = 9,360 (the 72 mixer applications, hoisted, plus the fold: chip-model-v3.md 5.2) + 8 reads x blocks per read x ops per block. `r1-curve-64.log` (608 ops per block, counted) and `r1-curve-64-memhard.log` (700, MEMHARD.md item 4). Ops per hash = 128 x ops per item + 512. MH/s at the chip model's 50 T op/s budget (approximate, chip-model-v3.md section 1).
|
|
|
|
| f (lines held) | SRAM | Blocks per read, naive / optimal | Ops per item, naive, 608 | Ops per item, optimal, 608 | Ops per item, optimal, 700 | Ops per hash, optimal, 608 | MH/s at 50 T op/s, optimal, 608 |
|
|
|---|---|---|---|---|---|---|---|
|
|
| 1/64 | 4 MiB | 31.5 / 16.0 | 162,576 | 87,184 | 98,960 | 11,160,064 | 4.5 |
|
|
| 1/32 | 8 MiB | 15.5 / 10.5 | 84,752 | 60,432 | 68,160 | 7,735,808 | 6.5 |
|
|
| 1/16 | 16 MiB | 7.5 / 6.094 | 45,840 | 39,000 | 43,485 | 4,992,512 | 10.0 |
|
|
| 1/8 | 32 MiB | 3.5 / 3.172 | 26,384 | 24,788 | 27,122 | 3,173,376 | 15.8 |
|
|
| 1/4 | 64 MiB | 1.5 / 1.453 | 16,656 | 16,428 | 17,498 | 2,103,296 | 23.8 |
|
|
| 1/2 | 128 MiB | 0.5 / 0.5 | 11,792 | 11,792 | 12,160 | 1,509,888 | 33.1 |
|
|
| 1 | 256 MiB | 0 / 0 | 9,360 | 9,360 | 9,360 | 1,198,592 | 41.7 |
|
|
|
|
Monotone: ops per item is non-increasing in f on both patterns at both op counts and on the 1,024-line model (`CURVE ... monotone non-increasing` in all three curve logs). The f = 1 point: 9,360 ops per item, 1,198,592 ops per hash, 41.7 MH/s at 50 T op/s, which is the chip-model-v3.md section 5.4 row "none, f = 0" of the published ITEM curve (the on-die-cache recompute chip of sections 1 to 3). The published item curve stores dataset ITEMS and is a different curve: its f = 1 point (GDDR7, 166.4 MH/s, 0.466 microjoules per hash) contains no cache read and no mixer op, so nothing in this row touches it. `funding.md` B2 rank 2's arithmetic reproduces on the naive pattern at 700 ops per block: 3.5 blocks per read, 2,450 ops per line, 19,600 per item on top of the mixer, 13.5 MH/s (50 T / (128 x 28,960 + 512)).
|
|
|
|
## Measured times, one box core (`r1-bench.log`, `r1-census.log`; nice 10, cores 12-15,60-63, box load 54)
|
|
|
|
| What | Measured | Note |
|
|
|---|---|---|
|
|
| One ChaCha12 block, dependent chain of 20,000,000 | 66.64 ns | |
|
|
| One mixer application (class v4 parameters), dependent chain of 20,000,000 | 17.29 ns | block / application = 3.85 (counted ops 608 / 128 = 4.75) |
|
|
| One item against the 256 MiB cache, batches of 32 | 1,326 ns | 72 applications = 1,245 ns; the 8 dependent reads and the fold add 81 ns because the batch overlaps them |
|
|
| One line recomputed from nothing, 312,500 random (seg, j) | 2,734 ns | 32.5 blocks per line on average, 84.1 ns per block inside the chain |
|
|
| The 256 MiB cache fill, one thread | 0.36 to 0.4 s | 86 ns per block with the writes |
|
|
|
|
In measured time, holding every 8th line at the optimal placement makes an item cost 72 + 8 x 3.172 x 3.85 = 170 mixer-application equivalents against 72, a 2.36x penalty per item (2.65x in counted ops). Holding one line in 64 costs 72 + 8 x 16 x 3.85 = 565, a 7.8x penalty.
|
|
|
|
## Checks on the real function (`r1-avalanche.log`, `r1-census.log`)
|
|
|
|
| Check | Result |
|
|
|---|---|
|
|
| Single-bit avalanche of `chacha_block`, 1,048,576 flips | mean 256.00 of 512 output bits change (ideal 256), min 200, max 312 |
|
|
| (output word, input word) pairs where an output word did not change | worst count 0 of 1,048,576 |
|
|
| Chain step: one bit of line j - 1 flipped | line j changes 255.96 bits, line j + 1 changes 255.94 (131,072 flips) |
|
|
| Differential independence: B(x ^ d) ^ B(x) == B(y ^ d) ^ B(y) over 262,144 (x, y, single-bit d) | 0 cases |
|
|
| Census of the real cache, day key 2026-10-03 | 4,194,304 of 4,194,304 lines distinct, 0 all-zero lines: no two chains merge and no block input repeats |
|
|
|
|
## Gate
|
|
|
|
| Clause | Result | Where |
|
|
|---|---|---|
|
|
| No derivation under j + 1 blocks | 0 of 64 and 0 of 1,024 lines under j + 1 on the extracted graph; the formula exact on 10,240 of 10,240 exhaustive pebbling cases; both known-fail chains fire | `r1-search-real-64.log`, `r1-search-real-1024.log`, `r1-pebble-10.log` |
|
|
| The curve monotone | non-increasing on both patterns, both op counts, both segment lengths | the three `r1-curve-*.log` |
|
|
| The f = 1 point unchanged | 9,360 ops per item = chip-model-v3.md 5.4 "none, f = 0" row; the item curve's GDDR7 f = 1 row (166.4 MH/s, 0.466 microjoules) untouched | `r1-curve-64.log` |
|
|
|
|
Verdict: PASS.
|
|
|
|
## Observations that are not findings
|
|
|
|
| Observation | Number | What it means | What I propose |
|
|
|---|---|---|---|
|
|
| `funding.md` B2 rank 2 prices the honest trade-off at the naive placement | 3.5 blocks per read at f = 1/8 against 3.17 optimal (9.4 percent less); 31.5 against 16.0 at f = 1/64 (2.0x less, because storing line 0 is worthless: it costs 1 block anyway) | the chip at f = 1/8 reads 15.8 MH/s (608 ops per block, optimal placement) or 14.4 (700, optimal) against `funding.md`'s 13.5 (700, naive); still 0.38x of the full SRAM mirror's 41.7 and 0.12x of the 5090's 136.1 (chip-model-v3.md section 2); the curve stays monotone, so the published verdict (the partial chip is not the threat, the full mirror beats it) stands | one sentence in `funding.md` B2 rank 2: "holding every 8th line at the best placement costs 3.2 blocks per read (3.5 for every 8th line from line 0)". Not edited here: outside this row's two files; for main to serialise |
|
|
| The chain's hardness per line is sequential time, not memory | one pebble (64 bytes) over j + 1 steps: the cumulative memory of deriving a line is about 64 x (j + 1) byte-steps | the chain protects the cache by op count, which is exactly what the curve prices in ops; parallel attackers pipeline items and pay E(f) x 608 ops per read in throughput, E(f) block latencies in latency; a chip that holds nothing (f = 0) pays 32.5 x 608 = 19,760 ops per read, 158,080 per item, 167,440 with the mixer (17.9x the mixer alone), 2.3 MH/s at 50 T op/s | nothing to move; the public model should keep quoting ops, never bytes, for this piece |
|
|
| What this row does not cover | a cryptanalytic shortcut inside `chacha_block` in this chaining mode (the differential and avalanche tests are sanity checks, not a bound) | the paid engagement's rank 2 question (`funding.md` B2) stays worth the money; plan 4.2 says the internal pass cannot prove the chain's trade-off curve | none |
|
|
|
|
## Consequences per user tier
|
|
|
|
| Tier | What this row changes |
|
|
|---|---|
|
|
| Home miner, one 8 / 12 / 16 / 24 or 32 GB card, any vendor, any OS | nothing: the honest miner holds the dataset, the verifier holds the 256 MiB cache; no memory, hash rate, or power figure moves |
|
|
| Rig, pool user | nothing |
|
|
| Chip builder | the partial-cache chip is priced 9 percent better at f = 1/8 and 2x better at f = 1/64 than `funding.md` says, and is still worse than the full SRAM mirror at every f below 1; the public per-joule sentence (evidence row 17, 2.1x at k = 1) rests on the item curve's f = 1 point, which this row leaves untouched |
|
|
| The paid review | the firm receives this record and the harness; rank 2's open question is the block function in chaining mode, not the graph |
|