igneum/docs/analysis/attack-pass/f9-grind.md

19 KiB

F9: acceptance edges, the hot-set search, header grinding

Attack-pass row F9 of docs/plans/cryptanalysis.md section 4.2 (record: docs/analysis/attack-pass-2026-10.md). Sub-agent attack-f9, 7 October 2026. Status: IN PROGRESS (rewritten as each run lands; the numbers below are the ones already final, each with its log).

Target

Item Value
Commit 924288d1 (branch attack-pass, worktree igneum-wt-attack)
Generator 4, class v4 mx8+sh256x27 composed with the era draw (LoadClass::era(V4_CLASS, E, [4 bytes])), era seed E = the devnet epoch-0 seed edc4fa84...fb07 (proto-cuda/packs-ca3-v4/v4-devnet-epoch0/seeds.txt)
Rule igneum-pow/src/accept.rs: (a) stale load sources, (b) injecting writes, (c) the 2,048-evaluation dynamic test on the closed-form stand-in dataset_elem at 2^28 words with init words = seed words; redraw on rejection up to 32 attempts
Header binding igneum-pow/src/bind.rs: init words I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32), one I per 32-lane warp, the lane nonce in the low 32 bits
Memory-hard dataset for the edges the devnet day 20730 (day_seed_hex 69676e65756d2d6461792ffa50000000000000), class v4 shape (mixer x8, cache 2^26 words, dataset 2^28 words), Epoch::chain_dataset_day
Card RunPod RTX 5090 (170 SMs, 32,120 MiB, driver 570.195.03, CUDA 12.8.1), pack v4-devnet-epoch0 built there with nvcc -O3 -arch=sm_120

Known-failed shape

A seed grind that steers a program to a hot cache set for DRAM locality, or an edge where the closed-form stand-in disagrees with the live verifier in the attacker's favour.

Gate

Zero passing programs with a hot set under 1 percent of items among 10^6 seeds; the grinding gain under 1 percent of rate at any search cost. What a failure moves: the closed-form stand-in replaced by the live verdict at the edges; a locality term in rule (c).

Harness

tools/attack/f9-grind/ (crate attack-f9, igneum-pow as a path dependency, nothing in igneum-pow edited; built on igneum-build-1 through tools/build-remote.sh; the binary on the box at /srv/builds/igneum-wt-attack/tools/attack/f9-grind/target/release/attack-f9):

Sub-command What it does
selftest the firings listed below
edges sub-row (a): every candidate of every seed through the re-implemented dynamic test twice, closed form and memory-hard, every metric of rule (c) measured to the end (no early abort) with its margin; the attempts continue until both stand-ins have accepted, so the chosen program under each is known
hotset sub-row (b): the seed's accepted program, its 2,048 x 128 address record under the acceptance init and under a block init; per site the nonce-independent address bits, the distinct addresses and the most-read address; the histogram at bucket scales 2^8 to 2^24 words against a window-aware Poisson expectation (the era windows send a site to the dataset, a half or a quarter of it) with a Bonferroni tail; the taint count of init-determined loads
inspect one seed's program with every site that repeats an address
grind, table-random sub-row (c), CPU side: the init-determined load sites of the devnet epoch-0 program, the per-warp search over K nonce_hi values for the fewest distinct 128 B lines inside those load instructions (--mode intra, what the coalescer merges) or across them (lines, pages), the search cost per hit, and the per-warp init tables the card reads
reference the 64 bound hashes the card's known-pass compares against; the file used on the pod came from the pre-built igneum-pow hash-bound instead (ref.txt, sha256 e831458a...)
summarise.py the census summaries quoted below

tools/attack/f9-grind/pod/ (the card): make-variants.py copies the pack's kernel_bound.cu into five kernels (per-warp init table; the five init-determined loads broadcast to lane 0's address; every load broadcast; the first such load broadcast; lane 1 reading lane 0's address at the first such load), f9-host.cu fills the cache and dataset with the pack's own kernels, checks them against vectors.h, checks the bound hash against the reference, checks the per-warp kernel on an all-equal table against the honest kernel, then times the eight variants in interleaved rounds; run.sh builds on the pod, samples nvidia-smi once a second and joins the samples to the phases (join-power.py).

Hot-set metric (F8's record docs/analysis/attack-pass/f8-uniform.md did not exist when this harness was written, so the metric is defined here). Strict reading, "any hot bucket": a site with 7 or more nonce-independent address bits (support at most 2^21 of 2^28 words, 0.78 percent of items; the window's own fixed bits not counted), or any bucket of at most 2^20 words (0.39 percent of the dataset) at scales 2^8, 2^12, 2^16, 2^20 whose count has a Poisson tail against its window-aware expectation under 10^-6 after the Bonferroni correction. Gate reading, "flagged": the reads above expectation in those hot buckets (the hot share, what a cache of the hot set saves at most) reach 1 percent of the program's reads, or a site has 7 constant bits.

Firings (the harness is trusted only after these)

Log: /srv/builds/igneum-wt-attack/attack-f9/selftest.log (copy in the Mac scratchpad f9-box/selftest.log).

Check Known-pass Known-fail Result
1 the re-implemented dynamic test against accept::check on 300 class v4 candidates: every verdict, the first failing condition, and distinct, saturated and bias of every accepted report equal PASS (300 candidates, 16 rejected by accept, all equal, 1.5 s)
2 both stand-ins forced equal (closed form twice) on 200 candidates PASS, 0 disagreements
3 the memory-hard stand-in gives different words: distinct 262,117 against 262,106, bias 56 against 64 on one program PASS (they differ)
4 50 accepted programs, none with 7 constant bits (worst 0) 50 plants (an accepted program rewritten to xor a,a; add a,a,256; mulhi a,b before a load from a, 48 of 50 still pass rule (c)) all flagged (support 256 words, site distinct 255 or 256) PASS for the plant; 4 of the 50 clean programs have a hot bucket (sub-row (b): that is the finding, not a harness fault)
5 taint on the devnet epoch-0 program against the hand reading of kernel_bound.cu: loads 7, 8, 9, 10 read r7, r4, r2, r0 (no load before them); load 31 reads r5 = r5 x r4 from instruction 12, both untouched by any load; loads 11, 13, 29, 30 read r1, r6, r4, r3, each written by an earlier load or by mad from r7 after load 10 PASS: sites (0,7) (0,8) (0,9) (0,10) (0,31)
card 1 cache FNV 448274a57f508cbc, dataset head, last word and 64 samples, the 96 pack vectors PASS (pod log/host.log)
card 2 the bound hash against the 64 reference lines of igneum-pow hash-bound (nonce_hi 0, prehash 000102..1f) PASS, 0 wrong
card 3 the per-warp kernel on an all-equal table equals the honest kernel on 64 lanes the two-init table: warp 0 equal, warp 1 differs in 32 of 32 lanes PASS both
card 4 the forced kernels change the hash: forced4 64 of 64 lanes, forcedall 64, forced1 64, pair 64 PASS (they fire)
card 5 a deliberately locality-maximising choice shows a measurable change: pair (one line of 4,096 saved per warp) +0.21 percent, forced1 (31 lines) +6.8 percent, forced4 (155 lines) +43.3 percent, forcedall +181.9 percent in the smoke run PASS (measurable from one saved line up)

Sub-row (a): the edges

Run: attack-f9 edges over seeds igneum-f9/0 to igneum-f9/99999, 8 threads on cores 28-31,76-79 in 10,000-seed chunks under a shared hold of the box measure lock (run-census.sh); output edges.part*.tsv, summary by summarise.py edges. The first 20,000 seeds (parts 0 and 1) are summarised here; the full 10^5 replaces this table when the run ends.

Quantity First 20,000 seeds
Candidates evaluated 21,020
Verdicts agreeing 21,007
Disagreements 13 (0.062 percent of candidates; the 3 October census had 39 in 100,000 on its generator)
Seeds whose chosen program differs 13 (every disagreement moves the chosen attempt, because the next attempt was accepted by both)
Exhausted seeds 0 on either stand-in
Rejected by the closed form / by the memory-hard dataset 1,013 / 1,014 (850 static, the rest (c))
First failing (c) condition, closed form const_bit 80, saturated 54, distinct 24, lane_const 4, bias 1
Accepted margins, closed form saturated at most 81 of 164, bias at most 120 of 136, distinct sum at least 247,335 (bound 245,760), nearly constant final bits up to 2,047 of 2,048
Accepted margins, memory-hard saturated at most 82, bias at most 115, distinct at least 247,678

The 13 disagreements: 12 are const_bit, a final register bit equal in all 2,048 evaluations on one dataset and in 2,044 to 2,047 of them on the other (7 where the closed form accepts, 5 where the memory-hard dataset accepts); 1 is bias, output bias 155 against 93 (6.9 against 4.1 sigma, the two draws' difference 2.7 sigma of sampling noise), where the memory-hard dataset accepts. No disagreement on saturation, lane-constant sites or the distinct count: those metrics are the same to within 20 on both datasets. What an attacker gains from a program the closed form accepts and the live dataset would reject: a register whose final bit is pinned in 2,047 of 2,048 hashes instead of 2,048, which no test downstream of the fold can see (the 64 output bits stay within 120 of 1,024 on every accepted program) and which no chip can turn into skipped work; the reverse direction loses the chain a program with one pinned bit. Either way the chosen program moves to the next attempt, which both stand-ins accept. Nothing in the attacker's favour: the verdict's dependence on the stand-in is a 0.06 percent coin flip on a one-bit property.

Run: attack-f9 hotset over seeds igneum-f9/0 to igneum-f9/999999, 8 threads on cores 32-35,80-83 in 100,000-seed chunks; output hotset.part*.tsv. A 2,000-seed timing sample (hotset-timing.tsv, seeds 5,000,000 to 5,001,999) is summarised here; the 10^6 census replaces it when the run ends.

Quantity 2,000-seed sample
Programs with any hot bucket (strict) 161 of 2,000 (8.1 percent)
Programs flagged at the gate reading (hot share at least 1 percent of reads) 21 of 2,000 (1.05 percent)
Worst hot share 10.3 percent of the program's reads (seed 5,000,968)
Sites with 7 or more constant address bits 0 (max 0)
Fewest distinct addresses at a site in 2,048 evaluations 434
Most evaluations reading one address at a site 767 of 2,048
Init-determined loads in iteration 0 (programs by count) 1: 234, 2: 573, 3: 594, 4: 368, 5: 179, 6: 42, 7: 10; none after iteration 0

FINDING F9-1 (or-saturation hot words). inspect --seed 5000968 (inspect-5000968.log): the load at instruction 33 reads r4; r4 is written by or r4 |= r0 (20), mulhi (27) and or r4 |= r6 (28), and r6 itself by or r6 |= r7 (7). or is absorbing toward all ones: after two or writes from independent words every bit is set with probability 7/8 and the whole register with probability (7/8)^32 = 1.4 percent; chained across iterations the mass grows, and on this program r4 is 0xffffffff at that site in 731 to 739 of 2,048 evaluations in iterations 1 to 7 (36 percent). The site then reads one word, rotl(0xffffffff x M, R) & window | offset = 0x0ca59e4c for a full window and 0x04a59e4c, 0x08a59e4c, ... for the windowed sites; near-all-ones values add a few hundred more. The same word family appears in every flagged program (seeds 4,000,001, 4,000,037, 4,000,040 in the selftest: or writes at 54 and 55 before the load at 60, or at 1 before the load at 3). Rule (c) does not see it: the saturation test counts final register values only (the register is overwritten before the end), the lane-constant test needs all 32 lanes equal, the distinct test counts per lane per hash (the hot word repeats across iterations, so it costs one distinct of 128), and the output bias stays within tolerance. The 3 October census measured an or_sat_frac per program (section 7.3, max 0.0102) but the adopted rule kept only the final-value count.

What it is worth to an attacker: nothing asymmetric. The hot words are the same for every lane that saturates, so the GPU's coalescer and L1 already serve them without a DRAM transaction, and a chip gets exactly the same. What it costs the design: those programs do fewer memory-hard reads than rule (c) promises (up to 10 percent fewer on the worst program in 2,000, at least 1 percent fewer on about 1 program in 100), so the per-hash memory work of class v4 is not the uniform 128 random reads the chip model assumes on every epoch. Gate reading: FAIL in the strict reading (zero passing programs with a hot set under 1 percent of items), FAIL in the share reading too (programs with a hot set capturing at least 1 percent of reads exist at about 1 percent of epochs). Proposed fix, for the hash lane (not applied here): a per-site line in rule (c), "every load site reads at least 2,000 distinct addresses over the 2,048 evaluations" (uniform gives 2,048 minus 0.008 expected repeats; the saturated sites read 434 to 1,855), computed from the addresses the test already collects (one sort of 2,048 per site, 128 sites, under a millisecond); the redraw rate rises by about the strict-reading fraction (8 percent of candidates) unless the threshold is placed at the share reading. The alternative, dropping the or family from the draw table, changes the frozen weights and is for the lane to weigh. Class check: a chip gains nothing today, but a stand-in that lets 1 percent of epochs run with a 1 to 10 percent lighter memory side is a published-number problem (evidence row 17's per-hash reads).

(the 10^6 numbers and the hot-share distribution replace the sample when the census ends)

Sub-row (c): header grinding

What an attacker can steer

Only a load whose address register has not yet absorbed a dataset word is a function of the init words and the nonce alone (taint analysis, init_determined_sites). On the devnet epoch-0 program these are the loads at instructions 7, 8, 9, 10 and 31 of iteration 0; from iteration 1 every register is tainted. Over the 2,000-seed sample the count is 1 to 7 per program, median 3, always in iteration 0 only. Everything after depends on dataset words the miner must fetch first. The init words themselves are an FNV hash of the header and nonce_hi, so the attacker cannot choose them, only draw them; and one draw serves a whole warp (the shuffles couple the 32 lanes), so a per-lane draw costs 32 hashes per lane.

The search (CPU)

grind draws K init words per warp (nonce_hi 0 to K-1 under the fixed prehash) and keeps the one with the fewest distinct 128 B lines among the init-determined loads. Each try costs 32 lanes x (8 init + 32 prefix instructions) = 1,280 lane-instructions; the warp's hash costs 32 x (512 + 55,296) = 1,785,856 lane-instructions, the derivation not counted. Logs: grind-k10-crosssite.log (lines counted across the five sites: coincidences that are at best an L2 hit), grind-k10-pages.log (2 KB pages across the sites), and the intra runs (lines inside one load instruction, what the coalescer merges into one transaction) that feed the card.

Metric K Warps Mean lines or pages saved per warp (of 4,096 loads) Warps improved Search per warp in hashes
lines across the sites 2^10 524,288 0.895 (0.022 percent) 89 percent 0.73
2 KB pages across the sites 2^10 65,536 1.455 (0.036 percent) 98 percent 0.73
lines inside one instruction (intra) 2^10 2^17 (pending) 0.73
lines inside one instruction (intra) 2^14 2^17 (pending) 11.7

The card (RTX 5090)

Smoke run (1 round of 2 s per variant, pod smoke logs): the calibration of what one saved line is worth.

Variant What changes MH/s Against honest W MH/J against honest
honest the pack kernel, one init per dispatch 141.76 482
perwarp-random per-warp init table, no search 141.73 -0.02 percent 486 -0.7 percent
perwarp-k10 per-warp table, best of 2^10 (cross-site table) 141.74 -0.01 percent 487 -1.0 percent
perwarp-k14 (the same table in the smoke run) 141.74 -0.01 percent 488 -1.1 percent
pair lane 1 reads lane 0's address at load 7: 1 line of 4,096 saved 142.05 +0.21 percent 490 -1.4 percent
forced1 load 7 broadcast: 31 lines saved 151.39 +6.8 percent 506 +1.8 percent
forced4 loads 7, 8, 9, 10, 31 broadcast: 155 lines saved 203.11 +43.3 percent 530 +30 percent
forcedall every load broadcast: 3,968 lines saved 399.61 +182 percent 520 +161 percent

Reading: the class v4 kernel on the 5090 is bound by its random reads (141.8 MH/s x 128 = 18.1 G reads per second, the card's measured random-read ceiling in docs/bench-log.md), and a load instruction completes when its slowest lane's transaction returns, so one saved line is worth about 0.2 percent of rate, 31 lines 6.8 percent, and the five init-determined loads fully coalesced 43 percent. That ceiling is unreachable by search: it needs the 32 lanes' 28-bit addresses to fall in one line at five sites, probability 2^-115 per draw. What a draw can reach is one coalesced pair at one site (probability 5 x C(32,2) / 2^23 = 3 x 10^-4 per try, about 3,400 tries per pair); two pairs need about 6 million tries, m pairs about 3,400^m / m! tries. One pair is worth 0.2 percent of one warp's hash and costs 3,400 x 1,280 lane-instructions = 2.4 hashes of search. The measured per-warp tables (5 rounds of 8 s, pending) are the direct check.

(the full run's table replaces the smoke run when it ends)

Consequences per tier

(filled in with the verdict)

Logs

Log Path
selftest, inspect, grind, census parts and drivers /srv/builds/igneum-wt-attack/attack-f9/ on igneum-build-1 (selftest.log, inspect-*.log, grind-*.log, edges.part*.tsv, edges.driver.log, hotset.part*.tsv, hotset.driver.log, hotset-timing.tsv, ref.txt, table-*.bin)
card smoke run and full run the pod's /workspace/f9/podjob/smoke/log/ and log/ (run.log, nvcc.log, host.log, power.csv, power-by-variant.txt, sha256.txt), copied to /srv/builds/igneum-wt-attack/attack-f9/pod/ at the end

The card, the measurement (RTX 5090 pod ap-f9, pod/host.log, sha256 83b1372c..., copied to

/srv/builds/igneum-wt-attack/attack-f9/pod/)

Per-warp header grinding at K = 2^14 draws per warp against the honest kernel, 5 interleaved rounds of 8 s each: 141.62 against 141.61 MH/s, +0.004 percent of rate, sd 0.003, at 11.7 hashes of search per hash. The unreachable ceiling (the five init-determined loads fully coalesced, forced1 extended) is +43 percent. Gate: the grinding gain under 1 percent of rate at any search cost. Sub-row (c): PASS. Pod time about 1 h 50 min from 09:02 UK; destroyed on the lane's done line.