Plan section 2.4 records the re-scope (Q1, Q2, Q5 here; Q3 and Q4 to the sibling lanes) and the operating changes (both boxes, nice 10, the capacity layer's yield, queued sweeps). The sweep is shardable (--seed-start), writes every row as it lands, counts the Q5 structure and draws from the F8 label space so the unmodified f8 harness (src/live.rs, built as adv-live) measures the live hot set of the same program. Devnet 3 epoch-0 program.json copied read-only as an input. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
306 lines
20 KiB
Markdown
306 lines
20 KiB
Markdown
# Attack plan: the program acceptance rule and its generator (class v4 sub-version 3)
|
|
|
|
internal adversarial pass, not an independent review
|
|
|
|
Author lane: adv-accept. Date: 7 October 2026, 19:xx BST. Branch: adv-accept off build/master.
|
|
Target commit: 017e70376489251e18564c0abce7e466e606c8b3 (class v4 sub-version 3, object byte 7).
|
|
Box: igneum-build-2, core band 0-31, nice 10, one slot at a time.
|
|
|
|
I am an outside attacker with the public kit. I have never worked on the hash code. Everything
|
|
below is priced against the chip model of docs/analysis/chip-model-v3.md.
|
|
|
|
## 0. Frozen-crate check (the outsider rule)
|
|
|
|
The brief said build/master's igneum-pow/ is byte-identical to the frozen commit and that
|
|
`git diff --stat 017e7037 HEAD -- igneum-pow` must print nothing. It does NOT print nothing.
|
|
|
|
git diff --stat 017e70376489251e18564c0abce7e466e606c8b3 3f0afcd5 -- igneum-pow
|
|
6 files changed, 30 insertions(+), 635 deletions(-)
|
|
(accept.rs 292 lines changed, generator.rs 320, emit.rs, packcheck.rs, tests/mixer.rs, tests/recheck.rs)
|
|
|
|
So build/master (HEAD 3f0afcd5) has moved past the frozen commit and has stripped most of the
|
|
class v4 sub-version-3 acceptance and generator code. To obey the outsider rule I pinned my
|
|
worktree to the frozen crate and packs:
|
|
|
|
git checkout 017e70376489251e18564c0abce7e466e606c8b3 -- igneum-pow proto-cuda/packs-ca3-v4
|
|
git diff --stat 017e7037 -- igneum-pow proto-cuda/packs-ca3-v4 # prints nothing: my tree == frozen
|
|
|
|
Every build and run in this pass is against the frozen crate. My harness crate depends on it by
|
|
path (../../../igneum-pow), as f8-uniform does.
|
|
|
|
## 1. Files I opened
|
|
|
|
Only the public kit, read at the frozen commit unless marked HEAD:
|
|
|
|
- igneum-pow/src/accept.rs, generator.rs, verify.rs, seed.rs, memhard.rs, bind.rs, main.rs, lib.rs
|
|
- igneum-pow/Cargo.toml, Cargo.lock
|
|
- docs/spec/01-lottery-hash.md (sections 1.3, 1.4, 1.4.6, 1.5, 1.6, 1.7, 1.8.5, 1.11, 1.12, 1.13.1, 1.13.3, 1.14-1.17)
|
|
- docs/analysis/chip-model-v3.md (HEAD; sections 1, 2, 5, 6)
|
|
- proto-cuda/packs-ca3-v4/ program.json of all eight packs; v4-devnet-epoch0 and v4-era-0 headers and seeds.txt
|
|
- tools/attack/f8-uniform/src/main.rs, Cargo.toml (copied from the Mac worktree, originals untouched)
|
|
- tools/attack/f4-weakday/src/main.rs, Cargo.toml (from build/attack-pass)
|
|
- tools/build-remote.sh, infra/build-server/lib.sh, infra/build-server/remote-run.sh
|
|
|
|
I did not open anything else under docs/, no site/, no proto-metal/, no other branch's source, no
|
|
git log. I extracted emit.rs to scratch but did not need to read it.
|
|
|
|
Kit zip check owed: the brief gives sha256 4f2445c5...f829154 for packs-ca3-v4-sub3; I verify the
|
|
on-disk packs match the frozen tree by git, not the zip (I was not handed the zip). Devnet class v4
|
|
epoch-0 program id in the pack is a785001687d8688a, which matches the brief. Devnet 3's epoch-0 id
|
|
fce15bf61030be57 is my derive check in Q3.
|
|
|
|
Second public-kit input (relayed by the adv-cache lane, 7 October 2026, public-kit class): the
|
|
Devnet 3 epoch-0 pack, v4-devnet3-epoch0.zip sha256
|
|
e025750f71175ed14d6e2a24e387ebbf1979b1cd0faee9139c41a7671165b334, program_id 0xfce15bf61030be57 at
|
|
attempt 0 (sub-version 3), day bytes le64(20733), 2^24 fingerprint from base 0 e510ad92b4d24846. A
|
|
second REAL program accepted at attempt 0 under the frozen rule, beside the shared devnet's (which is
|
|
attempt 1). I treat it as data, verify the sha before any use, and use it only as a known-passing
|
|
program for Q1 (hot set) and Q3 (derive check: my draw of Devnet 3 epoch 0 must give id
|
|
fce15bf61030be57 and fingerprint e510ad92b4d24846, else my draw path is wrong). It lives on build-1;
|
|
I copy the file read-only, I do not build or run on box 1.
|
|
|
|
## 2. The target, in my own words
|
|
|
|
### 2.1 The generator (spec 1.4, generator.rs candidate_from_words_class with V4_CLASS)
|
|
|
|
An epoch seed (32 bytes) gives seed words S[0..7] = seed_words_from_bytes(seed) for attempt 0, or
|
|
seed_words_from_bytes(seed || k_le32) for attempt k. The program stream is SplitMix64 seeded with
|
|
(S0 | S1<<32) XOR ((S2 | S3<<32) * 0x9E3779B97F4A7C15).
|
|
|
|
Class v4 = class v3 (mx8, era draw) plus a 256-instruction shadow block run 27 times per iteration.
|
|
Draw order:
|
|
1. 16 load slots: partial Fisher-Yates over instruction indices 1..63 (slot 0 is never a load).
|
|
2. 64 instructions, 9 draws each (op by weight, dst, src, src2, imm, imm2, rot, bit, mask), plus the
|
|
class v3/v4 width roll (consumed, width pinned to 4 bytes) and two era window draws (win=below(3),
|
|
off). A load's src is drawn only from registers that are FRESH by dataflow: written by an earlier
|
|
instruction, no load has read since, and (v4) the last writer kept entropy. Freshness value rule:
|
|
load/scratch/hot keep fresh iff source fresh; add/sub/xor/mad/shfl fresh iff dst or src fresh;
|
|
rotl/rotr fresh iff dst fresh; or/mul/mulhi never fresh. A shared-operand idiom (or-then-xor,
|
|
or-then-sub, xor-then-or on one operand) marks the result not fresh.
|
|
3. 256 shadow instructions, all ALU, drawn from the same stream after the base program.
|
|
|
|
### 2.2 The acceptance rule (spec 1.4.6, accept.rs check)
|
|
|
|
A candidate passes iff:
|
|
- (a) check_stale_loads: no load reads a register unwritten since the previous load from it, cyclic
|
|
over the 64 base instructions (two passes to close the wrap).
|
|
- (b) check_injecting_writes: every r0..r7 is the dst of an injecting op (add, sub, xor, mad, shfl,
|
|
load/variants).
|
|
- (a') check_fresh_sources_v4 (v4 shape only): run the freshness fixpoint over base||shadow to its
|
|
fixpoint (<=9 passes), then one checking pass; every load's source must be fresh.
|
|
- (c) check_dynamic: interpret 64 units x 32 lanes = 2,048 evaluations on the closed-form dataset
|
|
dataset_elem(idx, S0, S1) at 2^28 words, init words = S, base nonces low32(next()) & ~31 from
|
|
SplitMix64(FNV-1a("igneum-accept/" || S as LE bytes)). The shadow block IS executed (sub-version 3
|
|
fix). Tests: no register bit constant across all 2,048 finals; no load site reads one address in
|
|
all 32 lanes of any unit; <164 of 16,384 final regs saturated (0 or 2^32-1); every output bit ones
|
|
within 136 of 1,024; distinct masked addresses per lane per eval, summed, > 245,760 (mean >120/128).
|
|
- (c') SaturatedSource (v4): per load site, count of evaluations whose source value was 0/all-ones
|
|
< 164 over 16,384.
|
|
- (c'') (v4): LowEntropySite = per-site distinct dataset word indices over 2^20 evaluations
|
|
(ACCEPT_UNITS_DISTINCT_V4=4096 units) against uniform expectation N - N^2/2W on the site window,
|
|
ratio must reach MIN_DISTINCT_RATIO_V4 = 0.98. The RepeatedSource / most_repeated path (limit 8)
|
|
is present but NOT wired into check() (the finding is caught by the ratio instead).
|
|
|
|
Redraw: attempt 0,1,2,... up to MAX_ATTEMPTS_V4 = 256 for the v4 shape (32 otherwise). If all 256
|
|
fail, last_resort_v4 rewrites every or/mul/mulhi in base and shadow to xor and returns it, accepted
|
|
as drawn (no further check). Reached with probability about (2/3)^256.
|
|
|
|
### 2.3 What the rule is defending (chip model)
|
|
|
|
The live dataset is memory-hard (256 MiB cache, 8 dependent reads per item, mixer x8). The stand-in
|
|
is a pure closed form. The lottery is fair only if an accepted program reads a near-uniform, wide set
|
|
of dataset items on the LIVE dataset, so no on-die SRAM copy of a hot subset lets a chip skip the
|
|
memory. chip-model-v3.md section 5 prices the on-die-cache recompute chip at 0.92x and the stored
|
|
(f=1) chip over 2x per joule; a hot-set that concentrates reads would hand a chip exactly what
|
|
section 5.7's "hot set" lever asks about. Items per hash target: 128 distinct of 128 loads.
|
|
|
|
## 2.4 Re-scope (main, 7 October 2026, 19:2x BST) and the operating changes
|
|
|
|
The acceptance rule now has three lanes. This lane (adv-accept) is THE BYPASS: a program that passes
|
|
(a) to (c'') on the closed-form stand-in and has exploitable locality on the LIVE dataset. That is Q1
|
|
(the 10^6-seed passing-program search with the hot-set measure), Q2 (the stand-in gap reproduced and
|
|
bounded) and Q5 (the generator distinguishers that pass the rule). Q4 (header grinding) is lane
|
|
adv-accept-2's and Q3 (exhaustion and steering of the draw, the 256 cap, the last resort, the program
|
|
id) is lane adv-accept-3's; their methods below stay as written for the record and are handed over,
|
|
not run here. Shared queue for spill-over: /srv/builds/_adv/accept/queue/ on build-2, claimed by
|
|
mkdir under claims/. Operating changes from the coordinator: both boxes, nice 10 on every idle core
|
|
(no core band), the capacity layer's yield (SIGSTOP the run's process group while any build slot or
|
|
hold is taken, SIGCONT when clear; copied into run-box.sh), sweeps queued back to back, 8 box-hours
|
|
is a reading not a stop (ask line 16), first results by 00:00 BST tonight.
|
|
|
|
The seed space of the sweep is the attack-pass F8 harness's own label space
|
|
("igneum-attack-f8/program/k", ".../era/k"), so the unmodified f8 harness (built here as adv-live)
|
|
measures the live hot set of exactly the program my sweep reports for seed k.
|
|
|
|
## 3. The questions, in my order
|
|
|
|
Q1 (rank 1). Steering: a program that passes every part of the rule but concentrates its 128 live
|
|
loads on a small hot set of items/lines. The rule checks the CLOSED-FORM stand-in; the live set can
|
|
differ. This is the break that matters most, because it is the one the chip model prices.
|
|
|
|
Q2. The stand-in gap: closed-form verdict vs live-dataset verdict disagree. Bound the count and
|
|
whether any disagreement is exploitable (an accept that should be a reject on the live set).
|
|
|
|
Q3. Attempt grinding: whoever influences or predicts the epoch seed grinds attempt k from seed||k
|
|
for a favourable program. What the 256 cap and last_resort_v4 hand an attacker.
|
|
|
|
Q4. Header grinding for locality: the miner picks header bytes behind the pre-PoW hash, which only
|
|
enter the init words I (bind.rs), not the program. Can a cheap header search cluster a 32-lane
|
|
group's loads into fewer DRAM rows/lines.
|
|
|
|
Q5. Generator distinguishers: structural weaknesses of drawn programs that pass the rule (a lossy
|
|
last write, a contraction, register-set collapse, low-entropy sites the 0.98 floor still admits).
|
|
|
|
Q6. Anything else.
|
|
|
|
## 4. Method per question, with the known-failed shape each tool must fire on
|
|
|
|
All harness code: tools/attack/adv-accept/ (own crate, depends on igneum-pow by path). I reuse the
|
|
real draw and the real accept::check through the library; I never re-implement the rule. Where I need
|
|
the live dataset I build it through Epoch::chain_dataset_day(ProgramClass::V4), as f8-uniform does.
|
|
|
|
### Q1. Steering / hot set on the live dataset
|
|
|
|
Method. Over >= 10^6 epoch seeds, draw the chain's class v4 program (generate_era over V4_CLASS with
|
|
V3_ALLOWED, the era seed = a second label-derived seed, as f8 does), keep only accepted programs
|
|
(accept::check Ok). For each accepted program, run its 128 loads per hash over many nonces on the
|
|
LIVE memory-hard dataset (one day's cache, items derived into a table as a GPU holds them, the
|
|
f8-uniform warps census path), build the cross-hash item histogram and the per-site histogram, and
|
|
measure the hot set: the share S_f of reads taken by the top f of items for f in {0.1%, 1%} against a
|
|
uniform (and window-model) control, excess X_f = S_f - E_f, flag when X_f >= f. Also the per-site
|
|
distinct-index entropy and the distinct items/lines per hash and per warp. Rank seeds by X_0.1% and
|
|
by hottest-item share. f8-uniform already does most of this for its 64 labelled seeds; I extend the
|
|
seed sweep to >= 10^6 and keep the accept filter and the hot-set metric.
|
|
|
|
Known-failed shape. f8-uniform's --plant const-item feeds one constant item at the first load site;
|
|
--plant quarter-lines / half-lines mask the line index to a quarter/half. My sweep must FLAG the
|
|
planted program (X_f large, hot-set FLAGGED) and must leave the clean control within noise. If the
|
|
plant does not fire, the metric is wrong and nothing else counts.
|
|
|
|
Gain. The hot set measured as the fraction of items taking 0.1% and 1% of reads vs uniform, turned
|
|
into the chip's on-die copy size (section 5.7): a hot set holding share p of reads in q of the items
|
|
lets a chip store q items in SRAM and serve p of loads on die. I report p and q and the implied
|
|
SRAM MB and the gain row it lands on.
|
|
|
|
Box-hours. Seed sweep at 10^6 seeds, accept-filtered, ratio pass is the dear part: accept::check on a
|
|
v4 program is about 2.8 s per chosen candidate (the 2^20 ratio pass dominates) per the code comment,
|
|
but the cheap filter is (a)(b)(a')(c)(c') first; most seeds die before the ratio. I run the draw+
|
|
static+cheap-dynamic filter at ~10^6 seeds (a few core-hours on 32 cores) and the full live-dataset
|
|
histogram only on the top few hundred by a cheap proxy (per-site closed-form distinct count). Budget:
|
|
3 box-hours for Q1 (one build, one long nohup run, poll the log).
|
|
|
|
### Q2. The stand-in gap (closed form vs live)
|
|
|
|
Method. For accepted programs, compare the (c) verdict metrics computed on the closed form against
|
|
the same metrics computed on the live memory-hard dataset at 2^28: the distinct-address sum, the
|
|
saturation counts, the output bias, and the per-site distinct-index ratio. Count programs where the
|
|
live metric would reject but the closed-form accepts (a false accept) or vice versa. The spec claims
|
|
the two agree on all but 39 of 100,000 threshold-edge cases; I re-measure on my sweep and look for a
|
|
systematic (not threshold) disagreement I can steer into.
|
|
|
|
Known-failed shape. Construct a program by hand whose closed-form distinct count sits just above the
|
|
245,760 floor but whose live distinct count I force low by planting a source that is near-constant on
|
|
the live dataset only (e.g. a load whose source on the live set collapses because the live item map
|
|
folds it). The tool must report that planted program as a closed-form accept / live reject.
|
|
|
|
Gain / bound. Either a BREAK (a seed family that is a false accept and reads a hot set on the live
|
|
set — folds into Q1) or a BOUND: the measured disagreement count and margin over my sweep, stated
|
|
with the seed count.
|
|
|
|
Box-hours. Rides on Q1's live-dataset build (the cache fill is the cost, shared). 1 box-hour.
|
|
|
|
### Q3. Attempt grinding and the last resort
|
|
|
|
Method. Analytic plus a measured sweep. (a) The attempt stream: attempt k uses
|
|
seed_words_from_bytes(seed || k_le32); the program is a deterministic function of (seed, k). An
|
|
attacker who sets the epoch seed picks the (seed, k=first-accept) pair; an attacker who only predicts
|
|
it cannot change it. I quantify: over many seeds, how much does the hot-set metric of Q1 vary across
|
|
the first N accepted attempts of ONE seed vs across seeds, i.e. can grinding k (without changing the
|
|
seed) reach a worse program than attempt 0. Since k changes the whole seed-word set, each attempt is
|
|
an independent draw, so grinding k is just grinding seeds at a fixed epoch seed only if the attacker
|
|
controls nothing; if the attacker controls the epoch seed, k adds no freedom beyond the seed. (b)
|
|
The 256 cap and last_resort_v4: I drive a seed that exhausts (the test names igneum-f9/331672) and
|
|
measure what last_resort_v4 produces — is it predictable (yes, a deterministic function of attempt
|
|
256's draw with or/mul/mulhi->xor) and is it weak (every register fresh by construction, but does it
|
|
read a hot set or fail (c) on the live set though (a') passes). last_resort has NO (c)/(c'')/(c')
|
|
check, so it is the one accepted program the hot-set rule never sees.
|
|
|
|
Known-failed shape. I force the last-resort path on a planted exhausting seed (or by calling
|
|
last_resort_v4 directly on a rejected candidate, as the crate's own test does) and run its live hot
|
|
set through the Q1 metric. The tool fires if the last-resort program's hot set beats the floor the
|
|
rule enforces on normal programs.
|
|
|
|
Derive check. I derive Devnet 3's epoch-0 program id and confirm it equals fce15bf61030be57 (the
|
|
brief's check); a mismatch means my draw path is wrong and Q3 is void until fixed.
|
|
|
|
Box-hours. 1 box-hour (mostly the exhausting-seed search and the last-resort live run).
|
|
|
|
### Q4. Header grinding for locality
|
|
|
|
Method. The header bytes enter only the init words I = seed_words_from_bytes("igneum-block/" || H ||
|
|
nonce_hi) (bind.rs), not the program and not the dataset. A 32-lane group's 128 loads per lane are a
|
|
function of (program, I, base lane nonce). I measure, for one fixed accepted program and one live
|
|
dataset, how the distinct DRAM rows / cache lines touched by a 32-lane group vary over header
|
|
choices (vary H and nonce_hi), and whether a cheap search finds a header whose group clusters into
|
|
materially fewer lines than the mean. Cost vs gain: the search is hashing headers and running the
|
|
warp; the gain is fewer lines per group. I bound the best clustering found over a header budget and
|
|
the search cost in hashes per unit of clustering.
|
|
|
|
Known-failed shape. Plant a program whose loads are already clustered (reuse Q1's const-item plant)
|
|
so the per-group line count is low for every header; the tool must show the low count and show that
|
|
header choice does not move it (clustering is a program property, not a header one) — the control
|
|
that header grinding gives nothing. The positive known-fail: artificially make the address a function
|
|
of I alone (a one-line patch in my harness mirror, never in the library) and show the search then
|
|
finds a clustering header, so the search itself works.
|
|
|
|
Box-hours. 1 box-hour (one program, one cache, a header sweep).
|
|
|
|
### Q5. Generator distinguishers
|
|
|
|
Method. Over the Q1 sweep of accepted programs, count structural shapes: a lossy op (or/mul/mulhi)
|
|
as the last write to an output register; a contraction (mul/mulhi) feeding the fold; the register-set
|
|
reachability (does the program collapse to fewer than 8 live registers at the fold); low-entropy load
|
|
sites that pass the 0.98 ratio floor but sit in its tail (the open F8 seeds p4/p8/p10/p34 at 0.9927
|
|
to 0.9963). For each shape I ask whether it lowers the live distinct-item count or raises the hot
|
|
set. The 0.98 floor itself: I measure the distribution of the per-site ratio over accepted programs
|
|
and the gap between the floor and the clean median, to see how much low-entropy headroom the rule
|
|
leaves.
|
|
|
|
Known-failed shape. The crate's own diag tests name the band: F8's p23 (shared-operand andnot),
|
|
p15/p18/p19/p56 (fail the ratio), p4/p8/p10/p34 (open tail). My tool must reproduce the ratio verdict
|
|
on p15/p18/p19/p56 (rejected) and p23 attempt 1 (accepted), matching the crate's test, before I
|
|
trust any new finding.
|
|
|
|
Box-hours. Rides on Q1. 0.5 box-hours.
|
|
|
|
### Q6. Anything else
|
|
|
|
Free lane: the (c) base-nonce stream is keyed only on the seed words, not the day or the header, so
|
|
every node runs the SAME 2,048 acceptance evaluations for a seed forever; I check whether that fixed
|
|
sample is itself grindable (a program tuned to pass the 2,048 fixed nonces but fail off-sample). The
|
|
closed-form dataset is fixed at 2^28 whatever the live size; I check whether the live size change
|
|
(growth doublings) moves any accepted program off its (c) verdict. Bounds only unless something fires.
|
|
|
|
## 5. Box-hour budget
|
|
|
|
| Step | Box-hours | What |
|
|
|---|---|---|
|
|
| Build the harness (release) | 0.3 | one build-remote.sh --box 2 build |
|
|
| Q1 seed sweep + live hot set | 3.0 | nohup long run, poll |
|
|
| Q2 stand-in gap | 1.0 | rides on Q1 cache |
|
|
| Q3 attempt grind + last resort | 1.0 | |
|
|
| Q4 header locality | 1.0 | |
|
|
| Q5 distinguishers | 0.5 | rides on Q1 |
|
|
| Q6 + re-runs | 1.0 | |
|
|
| Total | 7.8 | within the 8 box-hour budget for first results |
|
|
|
|
GPU is not available tonight, so any per-card hash-rate confirmation of a hot-set gain is BLOCKED and
|
|
said so; I report the hot set and the implied SRAM size analytically against chip-model-v3.md.
|
|
|
|
## 6. Running rules I follow
|
|
|
|
Build: from tools/attack/adv-accept, `tools/build-remote.sh --box 2 -- build --release`. Long runs:
|
|
build first, then nohup the built binary on the box under nice 10 taskset 0-31, pid file beside the
|
|
log, logs under /srv/builds/igneum-wt-adv-accept/adv/, poll the log, kill by pid file only. One box
|
|
slot at a time. Commit as igneum-labs, push only to the build mirror (git push build adv-accept),
|
|
never origin. No cargo on this Mac.
|