igneum/docs/analysis/attack-pass/f3-cache.md
igneum-labs b98090649e attack-pass F3: PASS (0 of 64 and 0 of 1,024 lines under j+1, curve monotone, f=1 unchanged); AP-H1 box-clean hazard recorded
The F3 record and its harness crate (tools/attack/f3-cache: the extracted chain,
the exhaustive closure search, the pebbling cross-check, the store-set DP and
brute force, the two planted broken chains). The optimal-placement observation
(3.17 blocks per read at f=1/8, 16.0 at f=1/64) noted against funding.md B2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 08:16:48 +00:00

17 KiB

F3: the chained cache's j + 1 bound and the storage-against-recompute curve

Attack-pass row F3 of docs/plans/cryptanalysis.md section 4.2 (the record is docs/analysis/attack-pass-2026-10.md). Run 7 October 2026, 09:10 to 09:12 UK (08:10 to 08:12 UTC in the logs), on igneum-build-1. Verdict: PASS on all three gate clauses. No line (s, j) is derivable in fewer than j + 1 block evaluations without an earlier line, by an exhaustive search over the block dependency graph extracted from the code at 64 and 1,024 lines, cross-checked by an exhaustive pebbling search over every configuration at 10 lines. The storage-against-recompute curve over cache lines is monotone from f = 1/64 to 1. The f = 1 point is unchanged.

Target

Item Value
Commit 924288d1 (the brief); the worktree HEAD moved to 11b375a0 during the run; igneum-pow/src/memhard.rs is byte-identical at both (blob ad42470b, git diff --stat 924288d1 HEAD -- igneum-pow/src/memhard.rs empty)
Construction, from the code Cache::fill_segment: `in_j = prev XOR (sigma
Reads derive_items_mask: 8 dependent reads per item at line index s[0] AND mask, so the segment and j of a read are uniform over the 2^22 lines (F8 checks the uniformity)
Known-failed shape a line (s, j) computable in fewer than j + 1 block evaluations without an earlier line of segment s (the address-steering shape of the MTP break, Dinur and Nadler 2017, needs a data-dependent chain; this chain's inputs are fixed by the key, so the shape to search is a structural shortcut on the dependency graph)
Gate no derivation under j + 1 blocks; the curve monotone; the f = 1 point unchanged
Prior evidence none to re-gate: the ca2-cache branch named in the status board is the hot-table experiment (docs/plans/hot-table.md), not a chain analysis

Method

The model is the code, not the prose. tools/attack/f3-cache/src/main.rs runs one chain function, written in the shape of memhard.rs (the quarter round, the 6 double rounds, the feed-forward, the prev XOR, the constant block), generically over two word types:

Word type What it computes Use
u32 the real arithmetic verify: bit-exact against Cache::fill_segment on 16 (key, segment) pairs and against chacha_block on 100,000 random inputs
taint set which block outputs a value depends on (add, xor, rotate = union) search: the direct-parent graph of every block, with each computed line relabelled to the single node {j} so parents are direct, not transitive; plus the 16 x 16 (output word, input word) dependency matrix of one block

The exhaustive search: for every target line j, the minimum number of block evaluations with nothing stored is the size of the backward closure of j on the extracted graph (every non-stored block in the closure must be evaluated at least once; once each in dependency order suffices). pebble checks that formula against an exhaustive 0-1 BFS over every pebble configuration (place on a node whose parents are pebbled at cost 1, remove at cost 0) for all 2^10 stored sets x 10 targets on each of the three graphs: 10,240 pairs per graph, 0 mismatches. Two deliberately broken chains are the known-fail cases: skip2 (line j fed from line j - 2) and nofeed (no previous line fed in). The curve: for f = 1/64 to 1 (fraction of cache LINES held), the blocks per read on the naive pattern of funding.md B2 rank 2 (every L/n-th line from line 0) and on the optimal pattern (exact DP over chunk lengths; brute force over every C(64, n) set for n up to 8, 4,426,165,368 sets at n = 8); ops per item = 9,360 mixer ops (chip-model-v3.md 5.2) + 8 reads x blocks per read x ops per block (608 counted from the code: 48 quarter rounds x 12, 16 feed-forward adds, 16 input XORs; also at MEMHARD.md's approximate 700). Three more checks on the real function: single-bit avalanche and a differential-independence test on chacha_block, a census of every line of the real 2^22-line cache for the day key 2026-10-03, and one-core timings of a block, a mixer application, an item and a line recompute.

Harness

Item Value
Crate /Users/joshm/Projects/igneum-wt-attack/tools/attack/f3-cache/ (Cargo.toml with igneum-pow = { path = "../../../igneum-pow" } and an empty [workspace]; src/main.rs; run-box.sh)
Build cd tools/attack/f3-cache && IGNEUM_AGENT=attack-f3 IGNEUM_TOOLCHAIN_MISMATCH=ok bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f3" --out <scratchpad>/attack-f3 -- build --release (rc 0, 38 s wall, 0 warnings; the Mac's PATH rustc is 1.69 but ~/.cargo/bin/rustc is 1.99.0, which the script read as "on both sides"; log <scratchpad>/attack-f3/build-1.log)
Binary box /srv/builds/igneum-wt-attack/tools/attack/f3-cache/target/release/attack-f3, sha256 975115a385ca3195...71cc33, 563,104 bytes
Run on the box: nohup bash run-box.sh r1 > run-r1.log 2>&1 & from /srv/builds/igneum-wt-attack/attack-f3/; every phase as flock -s /srv/builds/_locks/measure -c "nice -n 10 taskset -c 12-15,60-63 attack-f3 <cmd>", one chunk per phase, the whole run 17 s (08:10:58 to 08:11:15 UTC; box load 54 at start)
Phase lines verify; `search --lines 64
Logs box /srv/builds/igneum-wt-attack/attack-f3/run-r1.log and r1-<phase>.log; Mac copies /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/attack-f3/

Box hygiene: the box checkout of every build-remote.sh run on this worktree executes git clean -fd at /srv/builds/igneum-wt-attack (remote-run.sh checkout_tree), which deletes any untracked scratch directory there. attack-f3/ and attack-f3-venv/ are listed in that mirror's .git/info/exclude so they survive; nothing in the tree was touched. The F1 lane's attack-f1-venv/ is untracked and unprotected and will be removed by the next build from any agent on this worktree.

The two firings and the pass

Chain Direct parents (taint trace) Lines under j + 1 at 64 lines Cheapest derivations Exhaustive pebbling at 10 lines, cost per target Verdict Log
real (the code) j - 1 for all 63 lines after line 0 0 of 64 none; every line costs exactly j + 1 (mean 32.5) 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 PASS r1-search-real-64.log, r1-pebble-10.log
real, 1,024-line model j - 1 for all 1,023 lines after line 0 0 of 1,024 none (mean 512.5) same graph rule PASS r1-search-real-1024.log
skip2 (known fail A) j - 2 for 62 lines, none for 2 63 of 64 j = 1 in 1, j = 63 in 32 (mean 16.5) 1, 1, 2, 2, 3, 3, 4, 4, 5, 5 FIRE r1-search-skip2-64.log, r1-search-skip2-1024.log
nofeed (known fail B) none for all 64 63 of 64 every line in 1 block (mean 1.0) 1 x 10 FIRE r1-search-nofeed-64.log, r1-search-nofeed-1024.log

verify (r1-verify.log): the model chain equals Cache::fill_segment on keys {day 2026-10-03, 3 random} x segments {0, 1, 12345, 65535} (16 of 16), block == chacha_block on 100,000 of 100,000 random inputs, the 1,024-line model's first 64 lines equal the 64-line chain, and both broken variants differ from the real chain from line 1 (line 0 equal, as the rule predicts). The block's word dependency matrix is full on every variant (256 of 256 pairs), so the firings come from the chain rule alone.

Derivation cost per line on the model segment (real chain, nothing stored)

j blocks to derive line j j + 1 Log
0 1 1 r1-search-real-1024.log
1 2 2
3 4 4
7 8 8
15 16 16
31 32 32
63 64 64 (the last line of a real segment; r1-search-real-64.log lists all 64)
127 128 128
255 256 256
511 512 512
1,023 1,024 1,024

All 1,024 lines were searched (0 under j + 1, mean 512.5 = (L + 1) / 2); the 64-line table in r1-search-real-64.log has every j from 0 to 63 at exactly j + 1.

Store patterns on the real 64-line segment

Blocks per read averaged over j uniform in 0..63. "Naive" is funding.md B2 rank 2's pattern (every k-th line from line 0). "Optimal" is the exact minimum over store sets of that size (DP; brute force over every set for n up to 8, agreeing with the DP on every row it ran). The gap formula equals the closure cost on the extracted graph on 2,000 of 2,000 random stored sets (r1-store-64.log).

f Stored lines n SRAM held Naive blocks per read Optimal positions Optimal blocks per read Brute force over C(64, n) sets
1/64 1 4 MiB 31.5 [32] 16.0 16.0 (64 sets)
1/32 2 8 MiB 15.5 [21, 43] 10.5 10.5 (2,016 sets)
1/16 4 16 MiB 7.5 [12, 25, 38, 51] 6.094 6.094 (635,376 sets)
1/8 8 32 MiB 3.5 [7, 15, 22, 29, 36, 43, 50, 57] 3.172 3.172 (4,426,165,368 sets, 11.2 s)
1/4 16 64 MiB 1.5 [3, 7, 11, ..., 55, 58, 61] 1.453 not run (DP exact)
1/2 32 128 MiB 0.5 odd lines 0.5 not run
1 64 256 MiB 0 all 0 not run

The 1,024-line model (r1-store-1024.log) gives 29.68 / 15.05 / 7.40 / 3.48 / 1.50 / 0.5 / 0 at the same f on the optimal pattern: the naive and optimal patterns converge as the chain lengthens, because the wasted stored line 0 and the end effects are a smaller share.

The curve: ops per item against the fraction of cache lines held (real 64-line segment)

Ops per item = 9,360 (the 72 mixer applications, hoisted, plus the fold: chip-model-v3.md 5.2) + 8 reads x blocks per read x ops per block. r1-curve-64.log (608 ops per block, counted) and r1-curve-64-memhard.log (700, MEMHARD.md item 4). Ops per hash = 128 x ops per item + 512. MH/s at the chip model's 50 T op/s budget (approximate, chip-model-v3.md section 1).

f (lines held) SRAM Blocks per read, naive / optimal Ops per item, naive, 608 Ops per item, optimal, 608 Ops per item, optimal, 700 Ops per hash, optimal, 608 MH/s at 50 T op/s, optimal, 608
1/64 4 MiB 31.5 / 16.0 162,576 87,184 98,960 11,160,064 4.5
1/32 8 MiB 15.5 / 10.5 84,752 60,432 68,160 7,735,808 6.5
1/16 16 MiB 7.5 / 6.094 45,840 39,000 43,485 4,992,512 10.0
1/8 32 MiB 3.5 / 3.172 26,384 24,788 27,122 3,173,376 15.8
1/4 64 MiB 1.5 / 1.453 16,656 16,428 17,498 2,103,296 23.8
1/2 128 MiB 0.5 / 0.5 11,792 11,792 12,160 1,509,888 33.1
1 256 MiB 0 / 0 9,360 9,360 9,360 1,198,592 41.7

Monotone: ops per item is non-increasing in f on both patterns at both op counts and on the 1,024-line model (CURVE ... monotone non-increasing in all three curve logs). The f = 1 point: 9,360 ops per item, 1,198,592 ops per hash, 41.7 MH/s at 50 T op/s, which is the chip-model-v3.md section 5.4 row "none, f = 0" of the published ITEM curve (the on-die-cache recompute chip of sections 1 to 3). The published item curve stores dataset ITEMS and is a different curve: its f = 1 point (GDDR7, 166.4 MH/s, 0.466 microjoules per hash) contains no cache read and no mixer op, so nothing in this row touches it. funding.md B2 rank 2's arithmetic reproduces on the naive pattern at 700 ops per block: 3.5 blocks per read, 2,450 ops per line, 19,600 per item on top of the mixer, 13.5 MH/s (50 T / (128 x 28,960 + 512)).

Measured times, one box core (r1-bench.log, r1-census.log; nice 10, cores 12-15,60-63, box load 54)

What Measured Note
One ChaCha12 block, dependent chain of 20,000,000 66.64 ns
One mixer application (class v4 parameters), dependent chain of 20,000,000 17.29 ns block / application = 3.85 (counted ops 608 / 128 = 4.75)
One item against the 256 MiB cache, batches of 32 1,326 ns 72 applications = 1,245 ns; the 8 dependent reads and the fold add 81 ns because the batch overlaps them
One line recomputed from nothing, 312,500 random (seg, j) 2,734 ns 32.5 blocks per line on average, 84.1 ns per block inside the chain
The 256 MiB cache fill, one thread 0.36 to 0.4 s 86 ns per block with the writes

In measured time, holding every 8th line at the optimal placement makes an item cost 72 + 8 x 3.172 x 3.85 = 170 mixer-application equivalents against 72, a 2.36x penalty per item (2.65x in counted ops). Holding one line in 64 costs 72 + 8 x 16 x 3.85 = 565, a 7.8x penalty.

Checks on the real function (r1-avalanche.log, r1-census.log)

Check Result
Single-bit avalanche of chacha_block, 1,048,576 flips mean 256.00 of 512 output bits change (ideal 256), min 200, max 312
(output word, input word) pairs where an output word did not change worst count 0 of 1,048,576
Chain step: one bit of line j - 1 flipped line j changes 255.96 bits, line j + 1 changes 255.94 (131,072 flips)
Differential independence: B(x ^ d) ^ B(x) == B(y ^ d) ^ B(y) over 262,144 (x, y, single-bit d) 0 cases
Census of the real cache, day key 2026-10-03 4,194,304 of 4,194,304 lines distinct, 0 all-zero lines: no two chains merge and no block input repeats

Gate

Clause Result Where
No derivation under j + 1 blocks 0 of 64 and 0 of 1,024 lines under j + 1 on the extracted graph; the formula exact on 10,240 of 10,240 exhaustive pebbling cases; both known-fail chains fire r1-search-real-64.log, r1-search-real-1024.log, r1-pebble-10.log
The curve monotone non-increasing on both patterns, both op counts, both segment lengths the three r1-curve-*.log
The f = 1 point unchanged 9,360 ops per item = chip-model-v3.md 5.4 "none, f = 0" row; the item curve's GDDR7 f = 1 row (166.4 MH/s, 0.466 microjoules) untouched r1-curve-64.log

Verdict: PASS.

Observations that are not findings

Observation Number What it means What I propose
funding.md B2 rank 2 prices the honest trade-off at the naive placement 3.5 blocks per read at f = 1/8 against 3.17 optimal (9.4 percent less); 31.5 against 16.0 at f = 1/64 (2.0x less, because storing line 0 is worthless: it costs 1 block anyway) the chip at f = 1/8 reads 15.8 MH/s (608 ops per block, optimal placement) or 14.4 (700, optimal) against funding.md's 13.5 (700, naive); still 0.38x of the full SRAM mirror's 41.7 and 0.12x of the 5090's 136.1 (chip-model-v3.md section 2); the curve stays monotone, so the published verdict (the partial chip is not the threat, the full mirror beats it) stands one sentence in funding.md B2 rank 2: "holding every 8th line at the best placement costs 3.2 blocks per read (3.5 for every 8th line from line 0)". Not edited here: outside this row's two files; for main to serialise
The chain's hardness per line is sequential time, not memory one pebble (64 bytes) over j + 1 steps: the cumulative memory of deriving a line is about 64 x (j + 1) byte-steps the chain protects the cache by op count, which is exactly what the curve prices in ops; parallel attackers pipeline items and pay E(f) x 608 ops per read in throughput, E(f) block latencies in latency; a chip that holds nothing (f = 0) pays 32.5 x 608 = 19,760 ops per read, 158,080 per item, 167,440 with the mixer (17.9x the mixer alone), 2.3 MH/s at 50 T op/s nothing to move; the public model should keep quoting ops, never bytes, for this piece
What this row does not cover a cryptanalytic shortcut inside chacha_block in this chaining mode (the differential and avalanche tests are sanity checks, not a bound) the paid engagement's rank 2 question (funding.md B2) stays worth the money; plan 4.2 says the internal pass cannot prove the chain's trade-off curve none

Consequences per user tier

Tier What this row changes
Home miner, one 8 / 12 / 16 / 24 or 32 GB card, any vendor, any OS nothing: the honest miner holds the dataset, the verifier holds the 256 MiB cache; no memory, hash rate, or power figure moves
Rig, pool user nothing
Chip builder the partial-cache chip is priced 9 percent better at f = 1/8 and 2x better at f = 1/64 than funding.md says, and is still worse than the full SRAM mirror at every f below 1; the public per-joule sentence (evidence row 17, 2.1x at k = 1) rests on the item curve's f = 1 point, which this row leaves untouched
The paid review the firm receives this record and the harness; rank 2's open question is the block function in chaining mode, not the graph