igneum/docs/plans/cryptanalysis/plan-acceptance-rule.md
igneum-labs caa8564ec1 adv-accept: re-scope to the bypass lane, yield run script, sharded sweep, f8 harness as adv-live
Plan section 2.4 records the re-scope (Q1, Q2, Q5 here; Q3 and Q4 to the sibling lanes) and
the operating changes (both boxes, nice 10, the capacity layer's yield, queued sweeps). The
sweep is shardable (--seed-start), writes every row as it lands, counts the Q5 structure
and draws from the F8 label space so the unmodified f8 harness (src/live.rs, built as
adv-live) measures the live hot set of the same program. Devnet 3 epoch-0 program.json
copied read-only as an input.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 18:28:19 +00:00

306 lines
20 KiB
Markdown

# Attack plan: the program acceptance rule and its generator (class v4 sub-version 3)
internal adversarial pass, not an independent review
Author lane: adv-accept. Date: 7 October 2026, 19:xx BST. Branch: adv-accept off build/master.
Target commit: 017e70376489251e18564c0abce7e466e606c8b3 (class v4 sub-version 3, object byte 7).
Box: igneum-build-2, core band 0-31, nice 10, one slot at a time.
I am an outside attacker with the public kit. I have never worked on the hash code. Everything
below is priced against the chip model of docs/analysis/chip-model-v3.md.
## 0. Frozen-crate check (the outsider rule)
The brief said build/master's igneum-pow/ is byte-identical to the frozen commit and that
`git diff --stat 017e7037 HEAD -- igneum-pow` must print nothing. It does NOT print nothing.
git diff --stat 017e70376489251e18564c0abce7e466e606c8b3 3f0afcd5 -- igneum-pow
6 files changed, 30 insertions(+), 635 deletions(-)
(accept.rs 292 lines changed, generator.rs 320, emit.rs, packcheck.rs, tests/mixer.rs, tests/recheck.rs)
So build/master (HEAD 3f0afcd5) has moved past the frozen commit and has stripped most of the
class v4 sub-version-3 acceptance and generator code. To obey the outsider rule I pinned my
worktree to the frozen crate and packs:
git checkout 017e70376489251e18564c0abce7e466e606c8b3 -- igneum-pow proto-cuda/packs-ca3-v4
git diff --stat 017e7037 -- igneum-pow proto-cuda/packs-ca3-v4 # prints nothing: my tree == frozen
Every build and run in this pass is against the frozen crate. My harness crate depends on it by
path (../../../igneum-pow), as f8-uniform does.
## 1. Files I opened
Only the public kit, read at the frozen commit unless marked HEAD:
- igneum-pow/src/accept.rs, generator.rs, verify.rs, seed.rs, memhard.rs, bind.rs, main.rs, lib.rs
- igneum-pow/Cargo.toml, Cargo.lock
- docs/spec/01-lottery-hash.md (sections 1.3, 1.4, 1.4.6, 1.5, 1.6, 1.7, 1.8.5, 1.11, 1.12, 1.13.1, 1.13.3, 1.14-1.17)
- docs/analysis/chip-model-v3.md (HEAD; sections 1, 2, 5, 6)
- proto-cuda/packs-ca3-v4/ program.json of all eight packs; v4-devnet-epoch0 and v4-era-0 headers and seeds.txt
- tools/attack/f8-uniform/src/main.rs, Cargo.toml (copied from the Mac worktree, originals untouched)
- tools/attack/f4-weakday/src/main.rs, Cargo.toml (from build/attack-pass)
- tools/build-remote.sh, infra/build-server/lib.sh, infra/build-server/remote-run.sh
I did not open anything else under docs/, no site/, no proto-metal/, no other branch's source, no
git log. I extracted emit.rs to scratch but did not need to read it.
Kit zip check owed: the brief gives sha256 4f2445c5...f829154 for packs-ca3-v4-sub3; I verify the
on-disk packs match the frozen tree by git, not the zip (I was not handed the zip). Devnet class v4
epoch-0 program id in the pack is a785001687d8688a, which matches the brief. Devnet 3's epoch-0 id
fce15bf61030be57 is my derive check in Q3.
Second public-kit input (relayed by the adv-cache lane, 7 October 2026, public-kit class): the
Devnet 3 epoch-0 pack, v4-devnet3-epoch0.zip sha256
e025750f71175ed14d6e2a24e387ebbf1979b1cd0faee9139c41a7671165b334, program_id 0xfce15bf61030be57 at
attempt 0 (sub-version 3), day bytes le64(20733), 2^24 fingerprint from base 0 e510ad92b4d24846. A
second REAL program accepted at attempt 0 under the frozen rule, beside the shared devnet's (which is
attempt 1). I treat it as data, verify the sha before any use, and use it only as a known-passing
program for Q1 (hot set) and Q3 (derive check: my draw of Devnet 3 epoch 0 must give id
fce15bf61030be57 and fingerprint e510ad92b4d24846, else my draw path is wrong). It lives on build-1;
I copy the file read-only, I do not build or run on box 1.
## 2. The target, in my own words
### 2.1 The generator (spec 1.4, generator.rs candidate_from_words_class with V4_CLASS)
An epoch seed (32 bytes) gives seed words S[0..7] = seed_words_from_bytes(seed) for attempt 0, or
seed_words_from_bytes(seed || k_le32) for attempt k. The program stream is SplitMix64 seeded with
(S0 | S1<<32) XOR ((S2 | S3<<32) * 0x9E3779B97F4A7C15).
Class v4 = class v3 (mx8, era draw) plus a 256-instruction shadow block run 27 times per iteration.
Draw order:
1. 16 load slots: partial Fisher-Yates over instruction indices 1..63 (slot 0 is never a load).
2. 64 instructions, 9 draws each (op by weight, dst, src, src2, imm, imm2, rot, bit, mask), plus the
class v3/v4 width roll (consumed, width pinned to 4 bytes) and two era window draws (win=below(3),
off). A load's src is drawn only from registers that are FRESH by dataflow: written by an earlier
instruction, no load has read since, and (v4) the last writer kept entropy. Freshness value rule:
load/scratch/hot keep fresh iff source fresh; add/sub/xor/mad/shfl fresh iff dst or src fresh;
rotl/rotr fresh iff dst fresh; or/mul/mulhi never fresh. A shared-operand idiom (or-then-xor,
or-then-sub, xor-then-or on one operand) marks the result not fresh.
3. 256 shadow instructions, all ALU, drawn from the same stream after the base program.
### 2.2 The acceptance rule (spec 1.4.6, accept.rs check)
A candidate passes iff:
- (a) check_stale_loads: no load reads a register unwritten since the previous load from it, cyclic
over the 64 base instructions (two passes to close the wrap).
- (b) check_injecting_writes: every r0..r7 is the dst of an injecting op (add, sub, xor, mad, shfl,
load/variants).
- (a') check_fresh_sources_v4 (v4 shape only): run the freshness fixpoint over base||shadow to its
fixpoint (<=9 passes), then one checking pass; every load's source must be fresh.
- (c) check_dynamic: interpret 64 units x 32 lanes = 2,048 evaluations on the closed-form dataset
dataset_elem(idx, S0, S1) at 2^28 words, init words = S, base nonces low32(next()) & ~31 from
SplitMix64(FNV-1a("igneum-accept/" || S as LE bytes)). The shadow block IS executed (sub-version 3
fix). Tests: no register bit constant across all 2,048 finals; no load site reads one address in
all 32 lanes of any unit; <164 of 16,384 final regs saturated (0 or 2^32-1); every output bit ones
within 136 of 1,024; distinct masked addresses per lane per eval, summed, > 245,760 (mean >120/128).
- (c') SaturatedSource (v4): per load site, count of evaluations whose source value was 0/all-ones
< 164 over 16,384.
- (c'') (v4): LowEntropySite = per-site distinct dataset word indices over 2^20 evaluations
(ACCEPT_UNITS_DISTINCT_V4=4096 units) against uniform expectation N - N^2/2W on the site window,
ratio must reach MIN_DISTINCT_RATIO_V4 = 0.98. The RepeatedSource / most_repeated path (limit 8)
is present but NOT wired into check() (the finding is caught by the ratio instead).
Redraw: attempt 0,1,2,... up to MAX_ATTEMPTS_V4 = 256 for the v4 shape (32 otherwise). If all 256
fail, last_resort_v4 rewrites every or/mul/mulhi in base and shadow to xor and returns it, accepted
as drawn (no further check). Reached with probability about (2/3)^256.
### 2.3 What the rule is defending (chip model)
The live dataset is memory-hard (256 MiB cache, 8 dependent reads per item, mixer x8). The stand-in
is a pure closed form. The lottery is fair only if an accepted program reads a near-uniform, wide set
of dataset items on the LIVE dataset, so no on-die SRAM copy of a hot subset lets a chip skip the
memory. chip-model-v3.md section 5 prices the on-die-cache recompute chip at 0.92x and the stored
(f=1) chip over 2x per joule; a hot-set that concentrates reads would hand a chip exactly what
section 5.7's "hot set" lever asks about. Items per hash target: 128 distinct of 128 loads.
## 2.4 Re-scope (main, 7 October 2026, 19:2x BST) and the operating changes
The acceptance rule now has three lanes. This lane (adv-accept) is THE BYPASS: a program that passes
(a) to (c'') on the closed-form stand-in and has exploitable locality on the LIVE dataset. That is Q1
(the 10^6-seed passing-program search with the hot-set measure), Q2 (the stand-in gap reproduced and
bounded) and Q5 (the generator distinguishers that pass the rule). Q4 (header grinding) is lane
adv-accept-2's and Q3 (exhaustion and steering of the draw, the 256 cap, the last resort, the program
id) is lane adv-accept-3's; their methods below stay as written for the record and are handed over,
not run here. Shared queue for spill-over: /srv/builds/_adv/accept/queue/ on build-2, claimed by
mkdir under claims/. Operating changes from the coordinator: both boxes, nice 10 on every idle core
(no core band), the capacity layer's yield (SIGSTOP the run's process group while any build slot or
hold is taken, SIGCONT when clear; copied into run-box.sh), sweeps queued back to back, 8 box-hours
is a reading not a stop (ask line 16), first results by 00:00 BST tonight.
The seed space of the sweep is the attack-pass F8 harness's own label space
("igneum-attack-f8/program/k", ".../era/k"), so the unmodified f8 harness (built here as adv-live)
measures the live hot set of exactly the program my sweep reports for seed k.
## 3. The questions, in my order
Q1 (rank 1). Steering: a program that passes every part of the rule but concentrates its 128 live
loads on a small hot set of items/lines. The rule checks the CLOSED-FORM stand-in; the live set can
differ. This is the break that matters most, because it is the one the chip model prices.
Q2. The stand-in gap: closed-form verdict vs live-dataset verdict disagree. Bound the count and
whether any disagreement is exploitable (an accept that should be a reject on the live set).
Q3. Attempt grinding: whoever influences or predicts the epoch seed grinds attempt k from seed||k
for a favourable program. What the 256 cap and last_resort_v4 hand an attacker.
Q4. Header grinding for locality: the miner picks header bytes behind the pre-PoW hash, which only
enter the init words I (bind.rs), not the program. Can a cheap header search cluster a 32-lane
group's loads into fewer DRAM rows/lines.
Q5. Generator distinguishers: structural weaknesses of drawn programs that pass the rule (a lossy
last write, a contraction, register-set collapse, low-entropy sites the 0.98 floor still admits).
Q6. Anything else.
## 4. Method per question, with the known-failed shape each tool must fire on
All harness code: tools/attack/adv-accept/ (own crate, depends on igneum-pow by path). I reuse the
real draw and the real accept::check through the library; I never re-implement the rule. Where I need
the live dataset I build it through Epoch::chain_dataset_day(ProgramClass::V4), as f8-uniform does.
### Q1. Steering / hot set on the live dataset
Method. Over >= 10^6 epoch seeds, draw the chain's class v4 program (generate_era over V4_CLASS with
V3_ALLOWED, the era seed = a second label-derived seed, as f8 does), keep only accepted programs
(accept::check Ok). For each accepted program, run its 128 loads per hash over many nonces on the
LIVE memory-hard dataset (one day's cache, items derived into a table as a GPU holds them, the
f8-uniform warps census path), build the cross-hash item histogram and the per-site histogram, and
measure the hot set: the share S_f of reads taken by the top f of items for f in {0.1%, 1%} against a
uniform (and window-model) control, excess X_f = S_f - E_f, flag when X_f >= f. Also the per-site
distinct-index entropy and the distinct items/lines per hash and per warp. Rank seeds by X_0.1% and
by hottest-item share. f8-uniform already does most of this for its 64 labelled seeds; I extend the
seed sweep to >= 10^6 and keep the accept filter and the hot-set metric.
Known-failed shape. f8-uniform's --plant const-item feeds one constant item at the first load site;
--plant quarter-lines / half-lines mask the line index to a quarter/half. My sweep must FLAG the
planted program (X_f large, hot-set FLAGGED) and must leave the clean control within noise. If the
plant does not fire, the metric is wrong and nothing else counts.
Gain. The hot set measured as the fraction of items taking 0.1% and 1% of reads vs uniform, turned
into the chip's on-die copy size (section 5.7): a hot set holding share p of reads in q of the items
lets a chip store q items in SRAM and serve p of loads on die. I report p and q and the implied
SRAM MB and the gain row it lands on.
Box-hours. Seed sweep at 10^6 seeds, accept-filtered, ratio pass is the dear part: accept::check on a
v4 program is about 2.8 s per chosen candidate (the 2^20 ratio pass dominates) per the code comment,
but the cheap filter is (a)(b)(a')(c)(c') first; most seeds die before the ratio. I run the draw+
static+cheap-dynamic filter at ~10^6 seeds (a few core-hours on 32 cores) and the full live-dataset
histogram only on the top few hundred by a cheap proxy (per-site closed-form distinct count). Budget:
3 box-hours for Q1 (one build, one long nohup run, poll the log).
### Q2. The stand-in gap (closed form vs live)
Method. For accepted programs, compare the (c) verdict metrics computed on the closed form against
the same metrics computed on the live memory-hard dataset at 2^28: the distinct-address sum, the
saturation counts, the output bias, and the per-site distinct-index ratio. Count programs where the
live metric would reject but the closed-form accepts (a false accept) or vice versa. The spec claims
the two agree on all but 39 of 100,000 threshold-edge cases; I re-measure on my sweep and look for a
systematic (not threshold) disagreement I can steer into.
Known-failed shape. Construct a program by hand whose closed-form distinct count sits just above the
245,760 floor but whose live distinct count I force low by planting a source that is near-constant on
the live dataset only (e.g. a load whose source on the live set collapses because the live item map
folds it). The tool must report that planted program as a closed-form accept / live reject.
Gain / bound. Either a BREAK (a seed family that is a false accept and reads a hot set on the live
set — folds into Q1) or a BOUND: the measured disagreement count and margin over my sweep, stated
with the seed count.
Box-hours. Rides on Q1's live-dataset build (the cache fill is the cost, shared). 1 box-hour.
### Q3. Attempt grinding and the last resort
Method. Analytic plus a measured sweep. (a) The attempt stream: attempt k uses
seed_words_from_bytes(seed || k_le32); the program is a deterministic function of (seed, k). An
attacker who sets the epoch seed picks the (seed, k=first-accept) pair; an attacker who only predicts
it cannot change it. I quantify: over many seeds, how much does the hot-set metric of Q1 vary across
the first N accepted attempts of ONE seed vs across seeds, i.e. can grinding k (without changing the
seed) reach a worse program than attempt 0. Since k changes the whole seed-word set, each attempt is
an independent draw, so grinding k is just grinding seeds at a fixed epoch seed only if the attacker
controls nothing; if the attacker controls the epoch seed, k adds no freedom beyond the seed. (b)
The 256 cap and last_resort_v4: I drive a seed that exhausts (the test names igneum-f9/331672) and
measure what last_resort_v4 produces — is it predictable (yes, a deterministic function of attempt
256's draw with or/mul/mulhi->xor) and is it weak (every register fresh by construction, but does it
read a hot set or fail (c) on the live set though (a') passes). last_resort has NO (c)/(c'')/(c')
check, so it is the one accepted program the hot-set rule never sees.
Known-failed shape. I force the last-resort path on a planted exhausting seed (or by calling
last_resort_v4 directly on a rejected candidate, as the crate's own test does) and run its live hot
set through the Q1 metric. The tool fires if the last-resort program's hot set beats the floor the
rule enforces on normal programs.
Derive check. I derive Devnet 3's epoch-0 program id and confirm it equals fce15bf61030be57 (the
brief's check); a mismatch means my draw path is wrong and Q3 is void until fixed.
Box-hours. 1 box-hour (mostly the exhausting-seed search and the last-resort live run).
### Q4. Header grinding for locality
Method. The header bytes enter only the init words I = seed_words_from_bytes("igneum-block/" || H ||
nonce_hi) (bind.rs), not the program and not the dataset. A 32-lane group's 128 loads per lane are a
function of (program, I, base lane nonce). I measure, for one fixed accepted program and one live
dataset, how the distinct DRAM rows / cache lines touched by a 32-lane group vary over header
choices (vary H and nonce_hi), and whether a cheap search finds a header whose group clusters into
materially fewer lines than the mean. Cost vs gain: the search is hashing headers and running the
warp; the gain is fewer lines per group. I bound the best clustering found over a header budget and
the search cost in hashes per unit of clustering.
Known-failed shape. Plant a program whose loads are already clustered (reuse Q1's const-item plant)
so the per-group line count is low for every header; the tool must show the low count and show that
header choice does not move it (clustering is a program property, not a header one) — the control
that header grinding gives nothing. The positive known-fail: artificially make the address a function
of I alone (a one-line patch in my harness mirror, never in the library) and show the search then
finds a clustering header, so the search itself works.
Box-hours. 1 box-hour (one program, one cache, a header sweep).
### Q5. Generator distinguishers
Method. Over the Q1 sweep of accepted programs, count structural shapes: a lossy op (or/mul/mulhi)
as the last write to an output register; a contraction (mul/mulhi) feeding the fold; the register-set
reachability (does the program collapse to fewer than 8 live registers at the fold); low-entropy load
sites that pass the 0.98 ratio floor but sit in its tail (the open F8 seeds p4/p8/p10/p34 at 0.9927
to 0.9963). For each shape I ask whether it lowers the live distinct-item count or raises the hot
set. The 0.98 floor itself: I measure the distribution of the per-site ratio over accepted programs
and the gap between the floor and the clean median, to see how much low-entropy headroom the rule
leaves.
Known-failed shape. The crate's own diag tests name the band: F8's p23 (shared-operand andnot),
p15/p18/p19/p56 (fail the ratio), p4/p8/p10/p34 (open tail). My tool must reproduce the ratio verdict
on p15/p18/p19/p56 (rejected) and p23 attempt 1 (accepted), matching the crate's test, before I
trust any new finding.
Box-hours. Rides on Q1. 0.5 box-hours.
### Q6. Anything else
Free lane: the (c) base-nonce stream is keyed only on the seed words, not the day or the header, so
every node runs the SAME 2,048 acceptance evaluations for a seed forever; I check whether that fixed
sample is itself grindable (a program tuned to pass the 2,048 fixed nonces but fail off-sample). The
closed-form dataset is fixed at 2^28 whatever the live size; I check whether the live size change
(growth doublings) moves any accepted program off its (c) verdict. Bounds only unless something fires.
## 5. Box-hour budget
| Step | Box-hours | What |
|---|---|---|
| Build the harness (release) | 0.3 | one build-remote.sh --box 2 build |
| Q1 seed sweep + live hot set | 3.0 | nohup long run, poll |
| Q2 stand-in gap | 1.0 | rides on Q1 cache |
| Q3 attempt grind + last resort | 1.0 | |
| Q4 header locality | 1.0 | |
| Q5 distinguishers | 0.5 | rides on Q1 |
| Q6 + re-runs | 1.0 | |
| Total | 7.8 | within the 8 box-hour budget for first results |
GPU is not available tonight, so any per-card hash-rate confirmation of a hot-set gain is BLOCKED and
said so; I report the hot set and the implied SRAM size analytically against chip-model-v3.md.
## 6. Running rules I follow
Build: from tools/attack/adv-accept, `tools/build-remote.sh --box 2 -- build --release`. Long runs:
build first, then nohup the built binary on the box under nice 10 taskset 0-31, pid file beside the
log, logs under /srv/builds/igneum-wt-adv-accept/adv/, poll the log, kill by pid file only. One box
slot at a time. Commit as igneum-labs, push only to the build mirror (git push build adv-accept),
never origin. No cargo on this Mac.