19 KiB
F9: acceptance edges, the hot-set search, header grinding
Attack-pass row F9 of docs/plans/cryptanalysis.md section 4.2 (record: docs/analysis/attack-pass-2026-10.md).
Sub-agent attack-f9, 7 October 2026. Status: IN PROGRESS (rewritten as each run lands; the numbers below are the
ones already final, each with its log).
Target
| Item | Value |
|---|---|
| Commit | 924288d1 (branch attack-pass, worktree igneum-wt-attack) |
| Generator | 4, class v4 mx8+sh256x27 composed with the era draw (LoadClass::era(V4_CLASS, E, [4 bytes])), era seed E = the devnet epoch-0 seed edc4fa84...fb07 (proto-cuda/packs-ca3-v4/v4-devnet-epoch0/seeds.txt) |
| Rule | igneum-pow/src/accept.rs: (a) stale load sources, (b) injecting writes, (c) the 2,048-evaluation dynamic test on the closed-form stand-in dataset_elem at 2^28 words with init words = seed words; redraw on rejection up to 32 attempts |
| Header binding | igneum-pow/src/bind.rs: init words I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32), one I per 32-lane warp, the lane nonce in the low 32 bits |
| Memory-hard dataset for the edges | the devnet day 20730 (day_seed_hex 69676e65756d2d6461792ffa50000000000000), class v4 shape (mixer x8, cache 2^26 words, dataset 2^28 words), Epoch::chain_dataset_day |
| Card | RunPod RTX 5090 (170 SMs, 32,120 MiB, driver 570.195.03, CUDA 12.8.1), pack v4-devnet-epoch0 built there with nvcc -O3 -arch=sm_120 |
Known-failed shape
A seed grind that steers a program to a hot cache set for DRAM locality, or an edge where the closed-form stand-in disagrees with the live verifier in the attacker's favour.
Gate
Zero passing programs with a hot set under 1 percent of items among 10^6 seeds; the grinding gain under 1 percent of rate at any search cost. What a failure moves: the closed-form stand-in replaced by the live verdict at the edges; a locality term in rule (c).
Harness
tools/attack/f9-grind/ (crate attack-f9, igneum-pow as a path dependency, nothing in igneum-pow edited; built
on igneum-build-1 through tools/build-remote.sh; the binary on the box at
/srv/builds/igneum-wt-attack/tools/attack/f9-grind/target/release/attack-f9):
| Sub-command | What it does |
|---|---|
selftest |
the firings listed below |
edges |
sub-row (a): every candidate of every seed through the re-implemented dynamic test twice, closed form and memory-hard, every metric of rule (c) measured to the end (no early abort) with its margin; the attempts continue until both stand-ins have accepted, so the chosen program under each is known |
hotset |
sub-row (b): the seed's accepted program, its 2,048 x 128 address record under the acceptance init and under a block init; per site the nonce-independent address bits, the distinct addresses and the most-read address; the histogram at bucket scales 2^8 to 2^24 words against a window-aware Poisson expectation (the era windows send a site to the dataset, a half or a quarter of it) with a Bonferroni tail; the taint count of init-determined loads |
inspect |
one seed's program with every site that repeats an address |
grind, table-random |
sub-row (c), CPU side: the init-determined load sites of the devnet epoch-0 program, the per-warp search over K nonce_hi values for the fewest distinct 128 B lines inside those load instructions (--mode intra, what the coalescer merges) or across them (lines, pages), the search cost per hit, and the per-warp init tables the card reads |
reference |
the 64 bound hashes the card's known-pass compares against; the file used on the pod came from the pre-built igneum-pow hash-bound instead (ref.txt, sha256 e831458a...) |
summarise.py |
the census summaries quoted below |
tools/attack/f9-grind/pod/ (the card): make-variants.py copies the pack's kernel_bound.cu into five kernels
(per-warp init table; the five init-determined loads broadcast to lane 0's address; every load broadcast; the first
such load broadcast; lane 1 reading lane 0's address at the first such load), f9-host.cu fills the cache and dataset
with the pack's own kernels, checks them against vectors.h, checks the bound hash against the reference, checks the
per-warp kernel on an all-equal table against the honest kernel, then times the eight variants in interleaved rounds;
run.sh builds on the pod, samples nvidia-smi once a second and joins the samples to the phases (join-power.py).
Hot-set metric (F8's record docs/analysis/attack-pass/f8-uniform.md did not exist when this harness was written, so
the metric is defined here). Strict reading, "any hot bucket": a site with 7 or more nonce-independent address bits
(support at most 2^21 of 2^28 words, 0.78 percent of items; the window's own fixed bits not counted), or any bucket
of at most 2^20 words (0.39 percent of the dataset) at scales 2^8, 2^12, 2^16, 2^20 whose count has a Poisson tail
against its window-aware expectation under 10^-6 after the Bonferroni correction. Gate reading, "flagged": the reads
above expectation in those hot buckets (the hot share, what a cache of the hot set saves at most) reach 1 percent of
the program's reads, or a site has 7 constant bits.
Firings (the harness is trusted only after these)
Log: /srv/builds/igneum-wt-attack/attack-f9/selftest.log (copy in the Mac scratchpad f9-box/selftest.log).
| Check | Known-pass | Known-fail | Result |
|---|---|---|---|
| 1 | the re-implemented dynamic test against accept::check on 300 class v4 candidates: every verdict, the first failing condition, and distinct, saturated and bias of every accepted report equal |
PASS (300 candidates, 16 rejected by accept, all equal, 1.5 s) | |
| 2 | both stand-ins forced equal (closed form twice) on 200 candidates | PASS, 0 disagreements | |
| 3 | the memory-hard stand-in gives different words: distinct 262,117 against 262,106, bias 56 against 64 on one program | PASS (they differ) | |
| 4 | 50 accepted programs, none with 7 constant bits (worst 0) | 50 plants (an accepted program rewritten to xor a,a; add a,a,256; mulhi a,b before a load from a, 48 of 50 still pass rule (c)) all flagged (support 256 words, site distinct 255 or 256) |
PASS for the plant; 4 of the 50 clean programs have a hot bucket (sub-row (b): that is the finding, not a harness fault) |
| 5 | taint on the devnet epoch-0 program against the hand reading of kernel_bound.cu: loads 7, 8, 9, 10 read r7, r4, r2, r0 (no load before them); load 31 reads r5 = r5 x r4 from instruction 12, both untouched by any load; loads 11, 13, 29, 30 read r1, r6, r4, r3, each written by an earlier load or by mad from r7 after load 10 |
PASS: sites (0,7) (0,8) (0,9) (0,10) (0,31) | |
| card 1 | cache FNV 448274a57f508cbc, dataset head, last word and 64 samples, the 96 pack vectors | PASS (pod log/host.log) |
|
| card 2 | the bound hash against the 64 reference lines of igneum-pow hash-bound (nonce_hi 0, prehash 000102..1f) |
PASS, 0 wrong | |
| card 3 | the per-warp kernel on an all-equal table equals the honest kernel on 64 lanes | the two-init table: warp 0 equal, warp 1 differs in 32 of 32 lanes | PASS both |
| card 4 | the forced kernels change the hash: forced4 64 of 64 lanes, forcedall 64, forced1 64, pair 64 | PASS (they fire) | |
| card 5 | a deliberately locality-maximising choice shows a measurable change: pair (one line of 4,096 saved per warp) +0.21 percent, forced1 (31 lines) +6.8 percent, forced4 (155 lines) +43.3 percent, forcedall +181.9 percent in the smoke run |
PASS (measurable from one saved line up) |
Sub-row (a): the edges
Run: attack-f9 edges over seeds igneum-f9/0 to igneum-f9/99999, 8 threads on cores 28-31,76-79 in 10,000-seed
chunks under a shared hold of the box measure lock (run-census.sh); output edges.part*.tsv, summary by
summarise.py edges. The first 20,000 seeds (parts 0 and 1) are summarised here; the full 10^5 replaces this table
when the run ends.
| Quantity | First 20,000 seeds |
|---|---|
| Candidates evaluated | 21,020 |
| Verdicts agreeing | 21,007 |
| Disagreements | 13 (0.062 percent of candidates; the 3 October census had 39 in 100,000 on its generator) |
| Seeds whose chosen program differs | 13 (every disagreement moves the chosen attempt, because the next attempt was accepted by both) |
| Exhausted seeds | 0 on either stand-in |
| Rejected by the closed form / by the memory-hard dataset | 1,013 / 1,014 (850 static, the rest (c)) |
| First failing (c) condition, closed form | const_bit 80, saturated 54, distinct 24, lane_const 4, bias 1 |
| Accepted margins, closed form | saturated at most 81 of 164, bias at most 120 of 136, distinct sum at least 247,335 (bound 245,760), nearly constant final bits up to 2,047 of 2,048 |
| Accepted margins, memory-hard | saturated at most 82, bias at most 115, distinct at least 247,678 |
The 13 disagreements: 12 are const_bit, a final register bit equal in all 2,048 evaluations on one dataset and in
2,044 to 2,047 of them on the other (7 where the closed form accepts, 5 where the memory-hard dataset accepts); 1 is
bias, output bias 155 against 93 (6.9 against 4.1 sigma, the two draws' difference 2.7 sigma of sampling noise),
where the memory-hard dataset accepts. No disagreement on saturation, lane-constant sites or the distinct count: those
metrics are the same to within 20 on both datasets. What an attacker gains from a program the closed form accepts
and the live dataset would reject: a register whose final bit is pinned in 2,047 of 2,048 hashes instead of 2,048,
which no test downstream of the fold can see (the 64 output bits stay within 120 of 1,024 on every accepted program)
and which no chip can turn into skipped work; the reverse direction loses the chain a program with one pinned bit.
Either way the chosen program moves to the next attempt, which both stand-ins accept. Nothing in the attacker's
favour: the verdict's dependence on the stand-in is a 0.06 percent coin flip on a one-bit property.
Sub-row (b): the hot-set search
Run: attack-f9 hotset over seeds igneum-f9/0 to igneum-f9/999999, 8 threads on cores 32-35,80-83 in
100,000-seed chunks; output hotset.part*.tsv. A 2,000-seed timing sample (hotset-timing.tsv, seeds 5,000,000 to
5,001,999) is summarised here; the 10^6 census replaces it when the run ends.
| Quantity | 2,000-seed sample |
|---|---|
| Programs with any hot bucket (strict) | 161 of 2,000 (8.1 percent) |
| Programs flagged at the gate reading (hot share at least 1 percent of reads) | 21 of 2,000 (1.05 percent) |
| Worst hot share | 10.3 percent of the program's reads (seed 5,000,968) |
| Sites with 7 or more constant address bits | 0 (max 0) |
| Fewest distinct addresses at a site in 2,048 evaluations | 434 |
| Most evaluations reading one address at a site | 767 of 2,048 |
| Init-determined loads in iteration 0 (programs by count) | 1: 234, 2: 573, 3: 594, 4: 368, 5: 179, 6: 42, 7: 10; none after iteration 0 |
FINDING F9-1 (or-saturation hot words). inspect --seed 5000968 (inspect-5000968.log): the load at instruction 33
reads r4; r4 is written by or r4 |= r0 (20), mulhi (27) and or r4 |= r6 (28), and r6 itself by or r6 |= r7
(7). or is absorbing toward all ones: after two or writes from independent words every bit is set with
probability 7/8 and the whole register with probability (7/8)^32 = 1.4 percent; chained across iterations the mass
grows, and on this program r4 is 0xffffffff at that site in 731 to 739 of 2,048 evaluations in iterations 1 to 7
(36 percent). The site then reads one word, rotl(0xffffffff x M, R) & window | offset = 0x0ca59e4c for a full
window and 0x04a59e4c, 0x08a59e4c, ... for the windowed sites; near-all-ones values add a few hundred more. The
same word family appears in every flagged program (seeds 4,000,001, 4,000,037, 4,000,040 in the selftest: or
writes at 54 and 55 before the load at 60, or at 1 before the load at 3). Rule (c) does not see it: the saturation
test counts final register values only (the register is overwritten before the end), the lane-constant test needs
all 32 lanes equal, the distinct test counts per lane per hash (the hot word repeats across iterations, so it
costs one distinct of 128), and the output bias stays within tolerance. The 3 October census measured an
or_sat_frac per program (section 7.3, max 0.0102) but the adopted rule kept only the final-value count.
What it is worth to an attacker: nothing asymmetric. The hot words are the same for every lane that saturates, so
the GPU's coalescer and L1 already serve them without a DRAM transaction, and a chip gets exactly the same. What it
costs the design: those programs do fewer memory-hard reads than rule (c) promises (up to 10 percent fewer on the
worst program in 2,000, at least 1 percent fewer on about 1 program in 100), so the per-hash memory work of class
v4 is not the uniform 128 random reads the chip model assumes on every epoch. Gate reading: FAIL in the strict
reading (zero passing programs with a hot set under 1 percent of items), FAIL in the share reading too (programs
with a hot set capturing at least 1 percent of reads exist at about 1 percent of epochs). Proposed fix, for the hash
lane (not applied here): a per-site line in rule (c), "every load site reads at least 2,000 distinct addresses over
the 2,048 evaluations" (uniform gives 2,048 minus 0.008 expected repeats; the saturated sites read 434 to 1,855),
computed from the addresses the test already collects (one sort of 2,048 per site, 128 sites, under a millisecond);
the redraw rate rises by about the strict-reading fraction (8 percent of candidates) unless the threshold is placed
at the share reading. The alternative, dropping the or family from the draw table, changes the frozen weights and
is for the lane to weigh. Class check: a chip gains nothing today, but a stand-in that lets 1 percent of epochs run
with a 1 to 10 percent lighter memory side is a published-number problem (evidence row 17's per-hash reads).
(the 10^6 numbers and the hot-share distribution replace the sample when the census ends)
Sub-row (c): header grinding
What an attacker can steer
Only a load whose address register has not yet absorbed a dataset word is a function of the init words and the
nonce alone (taint analysis, init_determined_sites). On the devnet epoch-0 program these are the loads at
instructions 7, 8, 9, 10 and 31 of iteration 0; from iteration 1 every register is tainted. Over the 2,000-seed
sample the count is 1 to 7 per program, median 3, always in iteration 0 only. Everything after depends on dataset
words the miner must fetch first. The init words themselves are an FNV hash of the header and nonce_hi, so the
attacker cannot choose them, only draw them; and one draw serves a whole warp (the shuffles couple the 32 lanes),
so a per-lane draw costs 32 hashes per lane.
The search (CPU)
grind draws K init words per warp (nonce_hi 0 to K-1 under the fixed prehash) and keeps the one with the fewest
distinct 128 B lines among the init-determined loads. Each try costs 32 lanes x (8 init + 32 prefix instructions) =
1,280 lane-instructions; the warp's hash costs 32 x (512 + 55,296) = 1,785,856 lane-instructions, the derivation
not counted. Logs: grind-k10-crosssite.log (lines counted across the five sites: coincidences that are at best an
L2 hit), grind-k10-pages.log (2 KB pages across the sites), and the intra runs (lines inside one load
instruction, what the coalescer merges into one transaction) that feed the card.
| Metric | K | Warps | Mean lines or pages saved per warp (of 4,096 loads) | Warps improved | Search per warp in hashes |
|---|---|---|---|---|---|
| lines across the sites | 2^10 | 524,288 | 0.895 (0.022 percent) | 89 percent | 0.73 |
| 2 KB pages across the sites | 2^10 | 65,536 | 1.455 (0.036 percent) | 98 percent | 0.73 |
| lines inside one instruction (intra) | 2^10 | 2^17 | (pending) | 0.73 | |
| lines inside one instruction (intra) | 2^14 | 2^17 | (pending) | 11.7 |
The card (RTX 5090)
Smoke run (1 round of 2 s per variant, pod smoke logs): the calibration of what one saved line is worth.
| Variant | What changes | MH/s | Against honest | W | MH/J against honest |
|---|---|---|---|---|---|
| honest | the pack kernel, one init per dispatch | 141.76 | 482 | ||
| perwarp-random | per-warp init table, no search | 141.73 | -0.02 percent | 486 | -0.7 percent |
| perwarp-k10 | per-warp table, best of 2^10 (cross-site table) | 141.74 | -0.01 percent | 487 | -1.0 percent |
| perwarp-k14 | (the same table in the smoke run) | 141.74 | -0.01 percent | 488 | -1.1 percent |
| pair | lane 1 reads lane 0's address at load 7: 1 line of 4,096 saved | 142.05 | +0.21 percent | 490 | -1.4 percent |
| forced1 | load 7 broadcast: 31 lines saved | 151.39 | +6.8 percent | 506 | +1.8 percent |
| forced4 | loads 7, 8, 9, 10, 31 broadcast: 155 lines saved | 203.11 | +43.3 percent | 530 | +30 percent |
| forcedall | every load broadcast: 3,968 lines saved | 399.61 | +182 percent | 520 | +161 percent |
Reading: the class v4 kernel on the 5090 is bound by its random reads (141.8 MH/s x 128 = 18.1 G reads per second,
the card's measured random-read ceiling in docs/bench-log.md), and a load instruction completes when its slowest
lane's transaction returns, so one saved line is worth about 0.2 percent of rate, 31 lines 6.8 percent, and the
five init-determined loads fully coalesced 43 percent. That ceiling is unreachable by search: it needs the 32
lanes' 28-bit addresses to fall in one line at five sites, probability 2^-115 per draw. What a draw can reach is one
coalesced pair at one site (probability 5 x C(32,2) / 2^23 = 3 x 10^-4 per try, about 3,400 tries per pair);
two pairs need about 6 million tries, m pairs about 3,400^m / m! tries. One pair is worth 0.2 percent of one warp's
hash and costs 3,400 x 1,280 lane-instructions = 2.4 hashes of search. The measured per-warp tables (5 rounds of
8 s, pending) are the direct check.
(the full run's table replaces the smoke run when it ends)
Consequences per tier
(filled in with the verdict)
Logs
| Log | Path |
|---|---|
| selftest, inspect, grind, census parts and drivers | /srv/builds/igneum-wt-attack/attack-f9/ on igneum-build-1 (selftest.log, inspect-*.log, grind-*.log, edges.part*.tsv, edges.driver.log, hotset.part*.tsv, hotset.driver.log, hotset-timing.tsv, ref.txt, table-*.bin) |
| card smoke run and full run | the pod's /workspace/f9/podjob/smoke/log/ and log/ (run.log, nvcc.log, host.log, power.csv, power-by-variant.txt, sha256.txt), copied to /srv/builds/igneum-wt-attack/attack-f9/pod/ at the end |
The card, the measurement (RTX 5090 pod ap-f9, pod/host.log, sha256 83b1372c..., copied to
/srv/builds/igneum-wt-attack/attack-f9/pod/)
Per-warp header grinding at K = 2^14 draws per warp against the honest kernel, 5 interleaved rounds of 8 s each:
141.62 against 141.61 MH/s, +0.004 percent of rate, sd 0.003, at 11.7 hashes of search per hash. The unreachable
ceiling (the five init-determined loads fully coalesced, forced1 extended) is +43 percent. Gate: the grinding gain
under 1 percent of rate at any search cost. Sub-row (c): PASS. Pod time about 1 h 50 min from 09:02 UK; destroyed on
the lane's done line.