igneum-pow at HEAD is IDENTICAL to the frozen object c3d32437 (and so is build/master's).
Shards 00 (box 2) and 01 (box 1) running under the yield script; shards 02..09 in the shared
queue; the adv-live const-item plant and its control running on box 1.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
20 KiB
Attack plan: the program acceptance rule and its generator (class v4 sub-version 3)
internal adversarial pass, not an independent review
Author lane: adv-accept. Date: 7 October 2026, 19:xx BST. Branch: adv-accept off build/master.
Target commit: 017e703764 (class v4 sub-version 3, object byte 7).
Box: igneum-build-2, core band 0-31, nice 10, one slot at a time.
I am an outside attacker with the public kit. I have never worked on the hash code. Everything below is priced against the chip model of docs/analysis/chip-model-v3.md.
0. Frozen-crate check (the outsider rule)
The brief said build/master's igneum-pow/ is byte-identical to the frozen commit and that
git diff --stat 017e7037 HEAD -- igneum-pow must print nothing. It does NOT print nothing.
git diff --stat 017e70376489251e18564c0abce7e466e606c8b3 3f0afcd5 -- igneum-pow
6 files changed, 30 insertions(+), 635 deletions(-)
(accept.rs 292 lines changed, generator.rs 320, emit.rs, packcheck.rs, tests/mixer.rs, tests/recheck.rs)
So build/master (HEAD 3f0afcd5) has moved past the frozen commit and has stripped most of the
class v4 sub-version-3 acceptance and generator code. To obey the outsider rule I pinned my
worktree to the frozen crate and packs:
git checkout 017e70376489251e18564c0abce7e466e606c8b3 -- igneum-pow proto-cuda/packs-ca3-v4
git diff --stat 017e7037 -- igneum-pow proto-cuda/packs-ca3-v4 # prints nothing: my tree == frozen
Every build and run in this pass is against the frozen crate. My harness crate depends on it by path (../../../igneum-pow), as f8-uniform does.
Re-base (coordinator's order, 19:3x BST): build/master moved to 7a7caa34, whose igneum-pow/ IS the
frozen object. Merged into this branch at 8748e955ceb48be5a6cbf7f8718a4884d3de9828;
git diff --quiet 017e7037 HEAD -- igneum-pow prints IDENTICAL, and the same check against
build/master 7a7caa34 prints IDENTICAL. Nothing I built or measured was on the stale tree: this
worktree's igneum-pow/ had already been reset to the frozen object (byte-identical) before the first
build, so the harness binaries are of the frozen crate; they are rebuilt at the merged HEAD anyway.
1. Files I opened
Only the public kit, read at the frozen commit unless marked HEAD:
- igneum-pow/src/accept.rs, generator.rs, verify.rs, seed.rs, memhard.rs, bind.rs, main.rs, lib.rs
- igneum-pow/Cargo.toml, Cargo.lock
- docs/spec/01-lottery-hash.md (sections 1.3, 1.4, 1.4.6, 1.5, 1.6, 1.7, 1.8.5, 1.11, 1.12, 1.13.1, 1.13.3, 1.14-1.17)
- docs/analysis/chip-model-v3.md (HEAD; sections 1, 2, 5, 6)
- proto-cuda/packs-ca3-v4/ program.json of all eight packs; v4-devnet-epoch0 and v4-era-0 headers and seeds.txt
- tools/attack/f8-uniform/src/main.rs, Cargo.toml (copied from the Mac worktree, originals untouched)
- tools/attack/f4-weakday/src/main.rs, Cargo.toml (from build/attack-pass)
- tools/build-remote.sh, infra/build-server/lib.sh, infra/build-server/remote-run.sh
I did not open anything else under docs/, no site/, no proto-metal/, no other branch's source, no git log. I extracted emit.rs to scratch but did not need to read it.
Kit zip check owed: the brief gives sha256 4f2445c5...f829154 for packs-ca3-v4-sub3; I verify the on-disk packs match the frozen tree by git, not the zip (I was not handed the zip). Devnet class v4 epoch-0 program id in the pack is a785001687d8688a, which matches the brief. Devnet 3's epoch-0 id fce15bf61030be57 is my derive check in Q3.
Second public-kit input (relayed by the adv-cache lane, 7 October 2026, public-kit class): the Devnet 3 epoch-0 pack, v4-devnet3-epoch0.zip sha256 e025750f71175ed14d6e2a24e387ebbf1979b1cd0faee9139c41a7671165b334, program_id 0xfce15bf61030be57 at attempt 0 (sub-version 3), day bytes le64(20733), 2^24 fingerprint from base 0 e510ad92b4d24846. A second REAL program accepted at attempt 0 under the frozen rule, beside the shared devnet's (which is attempt 1). I treat it as data, verify the sha before any use, and use it only as a known-passing program for Q1 (hot set) and Q3 (derive check: my draw of Devnet 3 epoch 0 must give id fce15bf61030be57 and fingerprint e510ad92b4d24846, else my draw path is wrong). It lives on build-1; I copy the file read-only, I do not build or run on box 1.
2. The target, in my own words
2.1 The generator (spec 1.4, generator.rs candidate_from_words_class with V4_CLASS)
An epoch seed (32 bytes) gives seed words S[0..7] = seed_words_from_bytes(seed) for attempt 0, or seed_words_from_bytes(seed || k_le32) for attempt k. The program stream is SplitMix64 seeded with (S0 | S1<<32) XOR ((S2 | S3<<32) * 0x9E3779B97F4A7C15).
Class v4 = class v3 (mx8, era draw) plus a 256-instruction shadow block run 27 times per iteration. Draw order:
- 16 load slots: partial Fisher-Yates over instruction indices 1..63 (slot 0 is never a load).
- 64 instructions, 9 draws each (op by weight, dst, src, src2, imm, imm2, rot, bit, mask), plus the class v3/v4 width roll (consumed, width pinned to 4 bytes) and two era window draws (win=below(3), off). A load's src is drawn only from registers that are FRESH by dataflow: written by an earlier instruction, no load has read since, and (v4) the last writer kept entropy. Freshness value rule: load/scratch/hot keep fresh iff source fresh; add/sub/xor/mad/shfl fresh iff dst or src fresh; rotl/rotr fresh iff dst fresh; or/mul/mulhi never fresh. A shared-operand idiom (or-then-xor, or-then-sub, xor-then-or on one operand) marks the result not fresh.
- 256 shadow instructions, all ALU, drawn from the same stream after the base program.
2.2 The acceptance rule (spec 1.4.6, accept.rs check)
A candidate passes iff:
- (a) check_stale_loads: no load reads a register unwritten since the previous load from it, cyclic over the 64 base instructions (two passes to close the wrap).
- (b) check_injecting_writes: every r0..r7 is the dst of an injecting op (add, sub, xor, mad, shfl, load/variants).
- (a') check_fresh_sources_v4 (v4 shape only): run the freshness fixpoint over base||shadow to its fixpoint (<=9 passes), then one checking pass; every load's source must be fresh.
- (c) check_dynamic: interpret 64 units x 32 lanes = 2,048 evaluations on the closed-form dataset dataset_elem(idx, S0, S1) at 2^28 words, init words = S, base nonces low32(next()) & ~31 from SplitMix64(FNV-1a("igneum-accept/" || S as LE bytes)). The shadow block IS executed (sub-version 3 fix). Tests: no register bit constant across all 2,048 finals; no load site reads one address in all 32 lanes of any unit; <164 of 16,384 final regs saturated (0 or 2^32-1); every output bit ones within 136 of 1,024; distinct masked addresses per lane per eval, summed, > 245,760 (mean >120/128).
- (c') SaturatedSource (v4): per load site, count of evaluations whose source value was 0/all-ones < 164 over 16,384.
- (c'') (v4): LowEntropySite = per-site distinct dataset word indices over 2^20 evaluations (ACCEPT_UNITS_DISTINCT_V4=4096 units) against uniform expectation N - N^2/2W on the site window, ratio must reach MIN_DISTINCT_RATIO_V4 = 0.98. The RepeatedSource / most_repeated path (limit 8) is present but NOT wired into check() (the finding is caught by the ratio instead).
Redraw: attempt 0,1,2,... up to MAX_ATTEMPTS_V4 = 256 for the v4 shape (32 otherwise). If all 256 fail, last_resort_v4 rewrites every or/mul/mulhi in base and shadow to xor and returns it, accepted as drawn (no further check). Reached with probability about (2/3)^256.
2.3 What the rule is defending (chip model)
The live dataset is memory-hard (256 MiB cache, 8 dependent reads per item, mixer x8). The stand-in is a pure closed form. The lottery is fair only if an accepted program reads a near-uniform, wide set of dataset items on the LIVE dataset, so no on-die SRAM copy of a hot subset lets a chip skip the memory. chip-model-v3.md section 5 prices the on-die-cache recompute chip at 0.92x and the stored (f=1) chip over 2x per joule; a hot-set that concentrates reads would hand a chip exactly what section 5.7's "hot set" lever asks about. Items per hash target: 128 distinct of 128 loads.
2.4 Re-scope (main, 7 October 2026, 19:2x BST) and the operating changes
The acceptance rule now has three lanes. This lane (adv-accept) is THE BYPASS: a program that passes (a) to (c'') on the closed-form stand-in and has exploitable locality on the LIVE dataset. That is Q1 (the 10^6-seed passing-program search with the hot-set measure), Q2 (the stand-in gap reproduced and bounded) and Q5 (the generator distinguishers that pass the rule). Q4 (header grinding) is lane adv-accept-2's and Q3 (exhaustion and steering of the draw, the 256 cap, the last resort, the program id) is lane adv-accept-3's; their methods below stay as written for the record and are handed over, not run here. Shared queue for spill-over: /srv/builds/_adv/accept/queue/ on build-2, claimed by mkdir under claims/. Operating changes from the coordinator: both boxes, nice 10 on every idle core (no core band), the capacity layer's yield (SIGSTOP the run's process group while any build slot or hold is taken, SIGCONT when clear; copied into run-box.sh), sweeps queued back to back, 8 box-hours is a reading not a stop (ask line 16), first results by 00:00 BST tonight.
The seed space of the sweep is the attack-pass F8 harness's own label space ("igneum-attack-f8/program/k", ".../era/k"), so the unmodified f8 harness (built here as adv-live) measures the live hot set of exactly the program my sweep reports for seed k.
3. The questions, in my order
Q1 (rank 1). Steering: a program that passes every part of the rule but concentrates its 128 live loads on a small hot set of items/lines. The rule checks the CLOSED-FORM stand-in; the live set can differ. This is the break that matters most, because it is the one the chip model prices.
Q2. The stand-in gap: closed-form verdict vs live-dataset verdict disagree. Bound the count and whether any disagreement is exploitable (an accept that should be a reject on the live set).
Q3. Attempt grinding: whoever influences or predicts the epoch seed grinds attempt k from seed||k for a favourable program. What the 256 cap and last_resort_v4 hand an attacker.
Q4. Header grinding for locality: the miner picks header bytes behind the pre-PoW hash, which only enter the init words I (bind.rs), not the program. Can a cheap header search cluster a 32-lane group's loads into fewer DRAM rows/lines.
Q5. Generator distinguishers: structural weaknesses of drawn programs that pass the rule (a lossy last write, a contraction, register-set collapse, low-entropy sites the 0.98 floor still admits).
Q6. Anything else.
4. Method per question, with the known-failed shape each tool must fire on
All harness code: tools/attack/adv-accept/ (own crate, depends on igneum-pow by path). I reuse the real draw and the real accept::check through the library; I never re-implement the rule. Where I need the live dataset I build it through Epoch::chain_dataset_day(ProgramClass::V4), as f8-uniform does.
Q1. Steering / hot set on the live dataset
Method. Over >= 10^6 epoch seeds, draw the chain's class v4 program (generate_era over V4_CLASS with V3_ALLOWED, the era seed = a second label-derived seed, as f8 does), keep only accepted programs (accept::check Ok). For each accepted program, run its 128 loads per hash over many nonces on the LIVE memory-hard dataset (one day's cache, items derived into a table as a GPU holds them, the f8-uniform warps census path), build the cross-hash item histogram and the per-site histogram, and measure the hot set: the share S_f of reads taken by the top f of items for f in {0.1%, 1%} against a uniform (and window-model) control, excess X_f = S_f - E_f, flag when X_f >= f. Also the per-site distinct-index entropy and the distinct items/lines per hash and per warp. Rank seeds by X_0.1% and by hottest-item share. f8-uniform already does most of this for its 64 labelled seeds; I extend the seed sweep to >= 10^6 and keep the accept filter and the hot-set metric.
Known-failed shape. f8-uniform's --plant const-item feeds one constant item at the first load site; --plant quarter-lines / half-lines mask the line index to a quarter/half. My sweep must FLAG the planted program (X_f large, hot-set FLAGGED) and must leave the clean control within noise. If the plant does not fire, the metric is wrong and nothing else counts.
Gain. The hot set measured as the fraction of items taking 0.1% and 1% of reads vs uniform, turned into the chip's on-die copy size (section 5.7): a hot set holding share p of reads in q of the items lets a chip store q items in SRAM and serve p of loads on die. I report p and q and the implied SRAM MB and the gain row it lands on.
Box-hours. Seed sweep at 10^6 seeds, accept-filtered, ratio pass is the dear part: accept::check on a v4 program is about 2.8 s per chosen candidate (the 2^20 ratio pass dominates) per the code comment, but the cheap filter is (a)(b)(a')(c)(c') first; most seeds die before the ratio. I run the draw+ static+cheap-dynamic filter at ~10^6 seeds (a few core-hours on 32 cores) and the full live-dataset histogram only on the top few hundred by a cheap proxy (per-site closed-form distinct count). Budget: 3 box-hours for Q1 (one build, one long nohup run, poll the log).
Q2. The stand-in gap (closed form vs live)
Method. For accepted programs, compare the (c) verdict metrics computed on the closed form against the same metrics computed on the live memory-hard dataset at 2^28: the distinct-address sum, the saturation counts, the output bias, and the per-site distinct-index ratio. Count programs where the live metric would reject but the closed-form accepts (a false accept) or vice versa. The spec claims the two agree on all but 39 of 100,000 threshold-edge cases; I re-measure on my sweep and look for a systematic (not threshold) disagreement I can steer into.
Known-failed shape. Construct a program by hand whose closed-form distinct count sits just above the 245,760 floor but whose live distinct count I force low by planting a source that is near-constant on the live dataset only (e.g. a load whose source on the live set collapses because the live item map folds it). The tool must report that planted program as a closed-form accept / live reject.
Gain / bound. Either a BREAK (a seed family that is a false accept and reads a hot set on the live set — folds into Q1) or a BOUND: the measured disagreement count and margin over my sweep, stated with the seed count.
Box-hours. Rides on Q1's live-dataset build (the cache fill is the cost, shared). 1 box-hour.
Q3. Attempt grinding and the last resort
Method. Analytic plus a measured sweep. (a) The attempt stream: attempt k uses seed_words_from_bytes(seed || k_le32); the program is a deterministic function of (seed, k). An attacker who sets the epoch seed picks the (seed, k=first-accept) pair; an attacker who only predicts it cannot change it. I quantify: over many seeds, how much does the hot-set metric of Q1 vary across the first N accepted attempts of ONE seed vs across seeds, i.e. can grinding k (without changing the seed) reach a worse program than attempt 0. Since k changes the whole seed-word set, each attempt is an independent draw, so grinding k is just grinding seeds at a fixed epoch seed only if the attacker controls nothing; if the attacker controls the epoch seed, k adds no freedom beyond the seed. (b) The 256 cap and last_resort_v4: I drive a seed that exhausts (the test names igneum-f9/331672) and measure what last_resort_v4 produces — is it predictable (yes, a deterministic function of attempt 256's draw with or/mul/mulhi->xor) and is it weak (every register fresh by construction, but does it read a hot set or fail (c) on the live set though (a') passes). last_resort has NO (c)/(c'')/(c') check, so it is the one accepted program the hot-set rule never sees.
Known-failed shape. I force the last-resort path on a planted exhausting seed (or by calling last_resort_v4 directly on a rejected candidate, as the crate's own test does) and run its live hot set through the Q1 metric. The tool fires if the last-resort program's hot set beats the floor the rule enforces on normal programs.
Derive check. I derive Devnet 3's epoch-0 program id and confirm it equals fce15bf61030be57 (the brief's check); a mismatch means my draw path is wrong and Q3 is void until fixed.
Box-hours. 1 box-hour (mostly the exhausting-seed search and the last-resort live run).
Q4. Header grinding for locality
Method. The header bytes enter only the init words I = seed_words_from_bytes("igneum-block/" || H || nonce_hi) (bind.rs), not the program and not the dataset. A 32-lane group's 128 loads per lane are a function of (program, I, base lane nonce). I measure, for one fixed accepted program and one live dataset, how the distinct DRAM rows / cache lines touched by a 32-lane group vary over header choices (vary H and nonce_hi), and whether a cheap search finds a header whose group clusters into materially fewer lines than the mean. Cost vs gain: the search is hashing headers and running the warp; the gain is fewer lines per group. I bound the best clustering found over a header budget and the search cost in hashes per unit of clustering.
Known-failed shape. Plant a program whose loads are already clustered (reuse Q1's const-item plant) so the per-group line count is low for every header; the tool must show the low count and show that header choice does not move it (clustering is a program property, not a header one) — the control that header grinding gives nothing. The positive known-fail: artificially make the address a function of I alone (a one-line patch in my harness mirror, never in the library) and show the search then finds a clustering header, so the search itself works.
Box-hours. 1 box-hour (one program, one cache, a header sweep).
Q5. Generator distinguishers
Method. Over the Q1 sweep of accepted programs, count structural shapes: a lossy op (or/mul/mulhi) as the last write to an output register; a contraction (mul/mulhi) feeding the fold; the register-set reachability (does the program collapse to fewer than 8 live registers at the fold); low-entropy load sites that pass the 0.98 ratio floor but sit in its tail (the open F8 seeds p4/p8/p10/p34 at 0.9927 to 0.9963). For each shape I ask whether it lowers the live distinct-item count or raises the hot set. The 0.98 floor itself: I measure the distribution of the per-site ratio over accepted programs and the gap between the floor and the clean median, to see how much low-entropy headroom the rule leaves.
Known-failed shape. The crate's own diag tests name the band: F8's p23 (shared-operand andnot), p15/p18/p19/p56 (fail the ratio), p4/p8/p10/p34 (open tail). My tool must reproduce the ratio verdict on p15/p18/p19/p56 (rejected) and p23 attempt 1 (accepted), matching the crate's test, before I trust any new finding.
Box-hours. Rides on Q1. 0.5 box-hours.
Q6. Anything else
Free lane: the (c) base-nonce stream is keyed only on the seed words, not the day or the header, so every node runs the SAME 2,048 acceptance evaluations for a seed forever; I check whether that fixed sample is itself grindable (a program tuned to pass the 2,048 fixed nonces but fail off-sample). The closed-form dataset is fixed at 2^28 whatever the live size; I check whether the live size change (growth doublings) moves any accepted program off its (c) verdict. Bounds only unless something fires.
5. Box-hour budget
| Step | Box-hours | What |
|---|---|---|
| Build the harness (release) | 0.3 | one build-remote.sh --box 2 build |
| Q1 seed sweep + live hot set | 3.0 | nohup long run, poll |
| Q2 stand-in gap | 1.0 | rides on Q1 cache |
| Q3 attempt grind + last resort | 1.0 | |
| Q4 header locality | 1.0 | |
| Q5 distinguishers | 0.5 | rides on Q1 |
| Q6 + re-runs | 1.0 | |
| Total | 7.8 | within the 8 box-hour budget for first results |
GPU is not available tonight, so any per-card hash-rate confirmation of a hot-set gain is BLOCKED and said so; I report the hot set and the implied SRAM size analytically against chip-model-v3.md.
6. Running rules I follow
Build: from tools/attack/adv-accept, tools/build-remote.sh --box 2 -- build --release. Long runs:
build first, then nohup the built binary on the box under nice 10 taskset 0-31, pid file beside the
log, logs under /srv/builds/igneum-wt-adv-accept/adv/, poll the log, kill by pid file only. One box
slot at a time. Commit as igneum-labs, push only to the build mirror (git push build adv-accept),
never origin. No cargo on this Mac.