The F3 record and its harness crate (tools/attack/f3-cache: the extracted chain, the exhaustive closure search, the pebbling cross-check, the store-set DP and brute force, the two planted broken chains). The optimal-placement observation (3.17 blocks per read at f=1/8, 16.0 at f=1/64) noted against funding.md B2. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
17 KiB
F3: the chained cache's j + 1 bound and the storage-against-recompute curve
Attack-pass row F3 of docs/plans/cryptanalysis.md section 4.2 (the record is docs/analysis/attack-pass-2026-10.md). Run 7 October 2026, 09:10 to 09:12 UK (08:10 to 08:12 UTC in the logs), on igneum-build-1. Verdict: PASS on all three gate clauses. No line (s, j) is derivable in fewer than j + 1 block evaluations without an earlier line, by an exhaustive search over the block dependency graph extracted from the code at 64 and 1,024 lines, cross-checked by an exhaustive pebbling search over every configuration at 10 lines. The storage-against-recompute curve over cache lines is monotone from f = 1/64 to 1. The f = 1 point is unchanged.
Target
| Item | Value |
|---|---|
| Commit | 924288d1 (the brief); the worktree HEAD moved to 11b375a0 during the run; igneum-pow/src/memhard.rs is byte-identical at both (blob ad42470b, git diff --stat 924288d1 HEAD -- igneum-pow/src/memhard.rs empty) |
| Construction, from the code | Cache::fill_segment: `in_j = prev XOR (sigma |
| Reads | derive_items_mask: 8 dependent reads per item at line index s[0] AND mask, so the segment and j of a read are uniform over the 2^22 lines (F8 checks the uniformity) |
| Known-failed shape | a line (s, j) computable in fewer than j + 1 block evaluations without an earlier line of segment s (the address-steering shape of the MTP break, Dinur and Nadler 2017, needs a data-dependent chain; this chain's inputs are fixed by the key, so the shape to search is a structural shortcut on the dependency graph) |
| Gate | no derivation under j + 1 blocks; the curve monotone; the f = 1 point unchanged |
| Prior evidence | none to re-gate: the ca2-cache branch named in the status board is the hot-table experiment (docs/plans/hot-table.md), not a chain analysis |
Method
The model is the code, not the prose. tools/attack/f3-cache/src/main.rs runs one chain function, written in the shape of memhard.rs (the quarter round, the 6 double rounds, the feed-forward, the prev XOR, the constant block), generically over two word types:
| Word type | What it computes | Use |
|---|---|---|
u32 |
the real arithmetic | verify: bit-exact against Cache::fill_segment on 16 (key, segment) pairs and against chacha_block on 100,000 random inputs |
| taint set | which block outputs a value depends on (add, xor, rotate = union) | search: the direct-parent graph of every block, with each computed line relabelled to the single node {j} so parents are direct, not transitive; plus the 16 x 16 (output word, input word) dependency matrix of one block |
The exhaustive search: for every target line j, the minimum number of block evaluations with nothing stored is the size of the backward closure of j on the extracted graph (every non-stored block in the closure must be evaluated at least once; once each in dependency order suffices). pebble checks that formula against an exhaustive 0-1 BFS over every pebble configuration (place on a node whose parents are pebbled at cost 1, remove at cost 0) for all 2^10 stored sets x 10 targets on each of the three graphs: 10,240 pairs per graph, 0 mismatches. Two deliberately broken chains are the known-fail cases: skip2 (line j fed from line j - 2) and nofeed (no previous line fed in). The curve: for f = 1/64 to 1 (fraction of cache LINES held), the blocks per read on the naive pattern of funding.md B2 rank 2 (every L/n-th line from line 0) and on the optimal pattern (exact DP over chunk lengths; brute force over every C(64, n) set for n up to 8, 4,426,165,368 sets at n = 8); ops per item = 9,360 mixer ops (chip-model-v3.md 5.2) + 8 reads x blocks per read x ops per block (608 counted from the code: 48 quarter rounds x 12, 16 feed-forward adds, 16 input XORs; also at MEMHARD.md's approximate 700). Three more checks on the real function: single-bit avalanche and a differential-independence test on chacha_block, a census of every line of the real 2^22-line cache for the day key 2026-10-03, and one-core timings of a block, a mixer application, an item and a line recompute.
Harness
| Item | Value |
|---|---|
| Crate | /Users/joshm/Projects/igneum-wt-attack/tools/attack/f3-cache/ (Cargo.toml with igneum-pow = { path = "../../../igneum-pow" } and an empty [workspace]; src/main.rs; run-box.sh) |
| Build | cd tools/attack/f3-cache && IGNEUM_AGENT=attack-f3 IGNEUM_TOOLCHAIN_MISMATCH=ok bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f3" --out <scratchpad>/attack-f3 -- build --release (rc 0, 38 s wall, 0 warnings; the Mac's PATH rustc is 1.69 but ~/.cargo/bin/rustc is 1.99.0, which the script read as "on both sides"; log <scratchpad>/attack-f3/build-1.log) |
| Binary | box /srv/builds/igneum-wt-attack/tools/attack/f3-cache/target/release/attack-f3, sha256 975115a385ca3195...71cc33, 563,104 bytes |
| Run | on the box: nohup bash run-box.sh r1 > run-r1.log 2>&1 & from /srv/builds/igneum-wt-attack/attack-f3/; every phase as flock -s /srv/builds/_locks/measure -c "nice -n 10 taskset -c 12-15,60-63 attack-f3 <cmd>", one chunk per phase, the whole run 17 s (08:10:58 to 08:11:15 UTC; box load 54 at start) |
| Phase lines | verify; `search --lines 64 |
| Logs | box /srv/builds/igneum-wt-attack/attack-f3/run-r1.log and r1-<phase>.log; Mac copies /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/attack-f3/ |
Box hygiene: the box checkout of every build-remote.sh run on this worktree executes git clean -fd at /srv/builds/igneum-wt-attack (remote-run.sh checkout_tree), which deletes any untracked scratch directory there. attack-f3/ and attack-f3-venv/ are listed in that mirror's .git/info/exclude so they survive; nothing in the tree was touched. The F1 lane's attack-f1-venv/ is untracked and unprotected and will be removed by the next build from any agent on this worktree.
The two firings and the pass
| Chain | Direct parents (taint trace) | Lines under j + 1 at 64 lines | Cheapest derivations | Exhaustive pebbling at 10 lines, cost per target | Verdict | Log |
|---|---|---|---|---|---|---|
| real (the code) | j - 1 for all 63 lines after line 0 | 0 of 64 | none; every line costs exactly j + 1 (mean 32.5) | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 | PASS | r1-search-real-64.log, r1-pebble-10.log |
| real, 1,024-line model | j - 1 for all 1,023 lines after line 0 | 0 of 1,024 | none (mean 512.5) | same graph rule | PASS | r1-search-real-1024.log |
| skip2 (known fail A) | j - 2 for 62 lines, none for 2 | 63 of 64 | j = 1 in 1, j = 63 in 32 (mean 16.5) | 1, 1, 2, 2, 3, 3, 4, 4, 5, 5 | FIRE | r1-search-skip2-64.log, r1-search-skip2-1024.log |
| nofeed (known fail B) | none for all 64 | 63 of 64 | every line in 1 block (mean 1.0) | 1 x 10 | FIRE | r1-search-nofeed-64.log, r1-search-nofeed-1024.log |
verify (r1-verify.log): the model chain equals Cache::fill_segment on keys {day 2026-10-03, 3 random} x segments {0, 1, 12345, 65535} (16 of 16), block == chacha_block on 100,000 of 100,000 random inputs, the 1,024-line model's first 64 lines equal the 64-line chain, and both broken variants differ from the real chain from line 1 (line 0 equal, as the rule predicts). The block's word dependency matrix is full on every variant (256 of 256 pairs), so the firings come from the chain rule alone.
Derivation cost per line on the model segment (real chain, nothing stored)
| j | blocks to derive line j | j + 1 | Log |
|---|---|---|---|
| 0 | 1 | 1 | r1-search-real-1024.log |
| 1 | 2 | 2 | |
| 3 | 4 | 4 | |
| 7 | 8 | 8 | |
| 15 | 16 | 16 | |
| 31 | 32 | 32 | |
| 63 | 64 | 64 | (the last line of a real segment; r1-search-real-64.log lists all 64) |
| 127 | 128 | 128 | |
| 255 | 256 | 256 | |
| 511 | 512 | 512 | |
| 1,023 | 1,024 | 1,024 |
All 1,024 lines were searched (0 under j + 1, mean 512.5 = (L + 1) / 2); the 64-line table in r1-search-real-64.log has every j from 0 to 63 at exactly j + 1.
Store patterns on the real 64-line segment
Blocks per read averaged over j uniform in 0..63. "Naive" is funding.md B2 rank 2's pattern (every k-th line from line 0). "Optimal" is the exact minimum over store sets of that size (DP; brute force over every set for n up to 8, agreeing with the DP on every row it ran). The gap formula equals the closure cost on the extracted graph on 2,000 of 2,000 random stored sets (r1-store-64.log).
| f | Stored lines n | SRAM held | Naive blocks per read | Optimal positions | Optimal blocks per read | Brute force over C(64, n) sets |
|---|---|---|---|---|---|---|
| 1/64 | 1 | 4 MiB | 31.5 | [32] | 16.0 | 16.0 (64 sets) |
| 1/32 | 2 | 8 MiB | 15.5 | [21, 43] | 10.5 | 10.5 (2,016 sets) |
| 1/16 | 4 | 16 MiB | 7.5 | [12, 25, 38, 51] | 6.094 | 6.094 (635,376 sets) |
| 1/8 | 8 | 32 MiB | 3.5 | [7, 15, 22, 29, 36, 43, 50, 57] | 3.172 | 3.172 (4,426,165,368 sets, 11.2 s) |
| 1/4 | 16 | 64 MiB | 1.5 | [3, 7, 11, ..., 55, 58, 61] | 1.453 | not run (DP exact) |
| 1/2 | 32 | 128 MiB | 0.5 | odd lines | 0.5 | not run |
| 1 | 64 | 256 MiB | 0 | all | 0 | not run |
The 1,024-line model (r1-store-1024.log) gives 29.68 / 15.05 / 7.40 / 3.48 / 1.50 / 0.5 / 0 at the same f on the optimal pattern: the naive and optimal patterns converge as the chain lengthens, because the wasted stored line 0 and the end effects are a smaller share.
The curve: ops per item against the fraction of cache lines held (real 64-line segment)
Ops per item = 9,360 (the 72 mixer applications, hoisted, plus the fold: chip-model-v3.md 5.2) + 8 reads x blocks per read x ops per block. r1-curve-64.log (608 ops per block, counted) and r1-curve-64-memhard.log (700, MEMHARD.md item 4). Ops per hash = 128 x ops per item + 512. MH/s at the chip model's 50 T op/s budget (approximate, chip-model-v3.md section 1).
| f (lines held) | SRAM | Blocks per read, naive / optimal | Ops per item, naive, 608 | Ops per item, optimal, 608 | Ops per item, optimal, 700 | Ops per hash, optimal, 608 | MH/s at 50 T op/s, optimal, 608 |
|---|---|---|---|---|---|---|---|
| 1/64 | 4 MiB | 31.5 / 16.0 | 162,576 | 87,184 | 98,960 | 11,160,064 | 4.5 |
| 1/32 | 8 MiB | 15.5 / 10.5 | 84,752 | 60,432 | 68,160 | 7,735,808 | 6.5 |
| 1/16 | 16 MiB | 7.5 / 6.094 | 45,840 | 39,000 | 43,485 | 4,992,512 | 10.0 |
| 1/8 | 32 MiB | 3.5 / 3.172 | 26,384 | 24,788 | 27,122 | 3,173,376 | 15.8 |
| 1/4 | 64 MiB | 1.5 / 1.453 | 16,656 | 16,428 | 17,498 | 2,103,296 | 23.8 |
| 1/2 | 128 MiB | 0.5 / 0.5 | 11,792 | 11,792 | 12,160 | 1,509,888 | 33.1 |
| 1 | 256 MiB | 0 / 0 | 9,360 | 9,360 | 9,360 | 1,198,592 | 41.7 |
Monotone: ops per item is non-increasing in f on both patterns at both op counts and on the 1,024-line model (CURVE ... monotone non-increasing in all three curve logs). The f = 1 point: 9,360 ops per item, 1,198,592 ops per hash, 41.7 MH/s at 50 T op/s, which is the chip-model-v3.md section 5.4 row "none, f = 0" of the published ITEM curve (the on-die-cache recompute chip of sections 1 to 3). The published item curve stores dataset ITEMS and is a different curve: its f = 1 point (GDDR7, 166.4 MH/s, 0.466 microjoules per hash) contains no cache read and no mixer op, so nothing in this row touches it. funding.md B2 rank 2's arithmetic reproduces on the naive pattern at 700 ops per block: 3.5 blocks per read, 2,450 ops per line, 19,600 per item on top of the mixer, 13.5 MH/s (50 T / (128 x 28,960 + 512)).
Measured times, one box core (r1-bench.log, r1-census.log; nice 10, cores 12-15,60-63, box load 54)
| What | Measured | Note |
|---|---|---|
| One ChaCha12 block, dependent chain of 20,000,000 | 66.64 ns | |
| One mixer application (class v4 parameters), dependent chain of 20,000,000 | 17.29 ns | block / application = 3.85 (counted ops 608 / 128 = 4.75) |
| One item against the 256 MiB cache, batches of 32 | 1,326 ns | 72 applications = 1,245 ns; the 8 dependent reads and the fold add 81 ns because the batch overlaps them |
| One line recomputed from nothing, 312,500 random (seg, j) | 2,734 ns | 32.5 blocks per line on average, 84.1 ns per block inside the chain |
| The 256 MiB cache fill, one thread | 0.36 to 0.4 s | 86 ns per block with the writes |
In measured time, holding every 8th line at the optimal placement makes an item cost 72 + 8 x 3.172 x 3.85 = 170 mixer-application equivalents against 72, a 2.36x penalty per item (2.65x in counted ops). Holding one line in 64 costs 72 + 8 x 16 x 3.85 = 565, a 7.8x penalty.
Checks on the real function (r1-avalanche.log, r1-census.log)
| Check | Result |
|---|---|
Single-bit avalanche of chacha_block, 1,048,576 flips |
mean 256.00 of 512 output bits change (ideal 256), min 200, max 312 |
| (output word, input word) pairs where an output word did not change | worst count 0 of 1,048,576 |
| Chain step: one bit of line j - 1 flipped | line j changes 255.96 bits, line j + 1 changes 255.94 (131,072 flips) |
| Differential independence: B(x ^ d) ^ B(x) == B(y ^ d) ^ B(y) over 262,144 (x, y, single-bit d) | 0 cases |
| Census of the real cache, day key 2026-10-03 | 4,194,304 of 4,194,304 lines distinct, 0 all-zero lines: no two chains merge and no block input repeats |
Gate
| Clause | Result | Where |
|---|---|---|
| No derivation under j + 1 blocks | 0 of 64 and 0 of 1,024 lines under j + 1 on the extracted graph; the formula exact on 10,240 of 10,240 exhaustive pebbling cases; both known-fail chains fire | r1-search-real-64.log, r1-search-real-1024.log, r1-pebble-10.log |
| The curve monotone | non-increasing on both patterns, both op counts, both segment lengths | the three r1-curve-*.log |
| The f = 1 point unchanged | 9,360 ops per item = chip-model-v3.md 5.4 "none, f = 0" row; the item curve's GDDR7 f = 1 row (166.4 MH/s, 0.466 microjoules) untouched | r1-curve-64.log |
Verdict: PASS.
Observations that are not findings
| Observation | Number | What it means | What I propose |
|---|---|---|---|
funding.md B2 rank 2 prices the honest trade-off at the naive placement |
3.5 blocks per read at f = 1/8 against 3.17 optimal (9.4 percent less); 31.5 against 16.0 at f = 1/64 (2.0x less, because storing line 0 is worthless: it costs 1 block anyway) | the chip at f = 1/8 reads 15.8 MH/s (608 ops per block, optimal placement) or 14.4 (700, optimal) against funding.md's 13.5 (700, naive); still 0.38x of the full SRAM mirror's 41.7 and 0.12x of the 5090's 136.1 (chip-model-v3.md section 2); the curve stays monotone, so the published verdict (the partial chip is not the threat, the full mirror beats it) stands |
one sentence in funding.md B2 rank 2: "holding every 8th line at the best placement costs 3.2 blocks per read (3.5 for every 8th line from line 0)". Not edited here: outside this row's two files; for main to serialise |
| The chain's hardness per line is sequential time, not memory | one pebble (64 bytes) over j + 1 steps: the cumulative memory of deriving a line is about 64 x (j + 1) byte-steps | the chain protects the cache by op count, which is exactly what the curve prices in ops; parallel attackers pipeline items and pay E(f) x 608 ops per read in throughput, E(f) block latencies in latency; a chip that holds nothing (f = 0) pays 32.5 x 608 = 19,760 ops per read, 158,080 per item, 167,440 with the mixer (17.9x the mixer alone), 2.3 MH/s at 50 T op/s | nothing to move; the public model should keep quoting ops, never bytes, for this piece |
| What this row does not cover | a cryptanalytic shortcut inside chacha_block in this chaining mode (the differential and avalanche tests are sanity checks, not a bound) |
the paid engagement's rank 2 question (funding.md B2) stays worth the money; plan 4.2 says the internal pass cannot prove the chain's trade-off curve |
none |
Consequences per user tier
| Tier | What this row changes |
|---|---|
| Home miner, one 8 / 12 / 16 / 24 or 32 GB card, any vendor, any OS | nothing: the honest miner holds the dataset, the verifier holds the 256 MiB cache; no memory, hash rate, or power figure moves |
| Rig, pool user | nothing |
| Chip builder | the partial-cache chip is priced 9 percent better at f = 1/8 and 2x better at f = 1/64 than funding.md says, and is still worse than the full SRAM mirror at every f below 1; the public per-joule sentence (evidence row 17, 2.1x at k = 1) rests on the item curve's f = 1 point, which this row leaves untouched |
| The paid review | the firm receives this record and the harness; rank 2's open question is the block function in chaining mode, not the graph |