Lane adv-accept-2, question class HEADER GRINDING FOR LOCALITY. Base commit5e412177(merged build/master); igneum-pow byte-identical to the frozen objectc3d32437cd(git diff --quiet ... HEAD -- igneum-pow prints nothing, verified). Plan: what of the header reaches the load addresses (bind.rs, spec 1.6, 1.13.1), the distinct rows/lines/items sweep, the price model, the per-hash vs per-program question. Harness depends on igneum-pow by path and mirrors verify.rs; the real draw and address map come from the library. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
8 KiB
Attack plan: header grinding for locality (acceptance rule and memory access, lane adv-accept-2)
internal adversarial pass, not an independent review
Lane adv-accept-2. Branch adv-accept-2 off build/master (7a7caa34). Written 7 October 2026, 19:24 to 20:40 BST (the hour ran 15 minutes over; said here). Target commit 017e703764 (class v4 sub-version 3, object byte 7). I am an outsider with the public kit; I have never worked on the hash code. Every sentence here that could be quoted in public carries the label above.
0. The outsider rule, applied
| Check | Result |
|---|---|
git diff --stat 017e7037... HEAD -- igneum-pow at HEAD 7a7caa34 |
prints nothing. igneum-pow/ is byte-identical to the frozen commit on this worktree. (Sibling lanes saw a 6-file divergence at an earlier master 3f0afcd5; master has since moved and the crate at my HEAD matches frozen, so my harness depends on the worktree's own igneum-pow by path.) |
| Public kit | proto-cuda/packs-ca3-v4/, eight packs at the frozen commit; v4-devnet-epoch0 id 0xa785001687d8688a, attempt 1, class mx8-erad810f22d+sh256x27. |
| Devnet 3 pack | /srv/artefacts/packs/v4-devnet3-epoch0.zip on build-1, sha256 e025750f... verified. id 0xfce15bf61030be57, attempt 0, generator 4, sub-version 3, day bytes le64(20733). Copied read-only. |
| Boxes | build-2 (96 threads, load 62 at 20:27 BST) is my run box; build-1 read-only for the pack. No GPU: the GPU row is BLOCKED. |
Files opened (complete list)
verify.rs, accept.rs, bind.rs, seed.rs (in full); generator.rs and memhard.rs (public items, the draw and the dataset fetch); lib.rs, Cargo.toml, main.rs (subcommands). docs/spec/01-lottery-hash.md at 017e7037 (1.4 to 1.4.6, 1.5, 1.6, 1.7, 1.8.5, 1.9, 1.13). docs/analysis/chip-model-v3.md at HEAD (1, 2, 5, 6). The eight packs' program.json and v4-era-0/seeds.txt. tools/attack/f8-uniform and f4-weakday (headers and patterns). tools/build-remote.sh, infra/build-server/{lib.sh, remote-run.sh, capacity/run.sh, capacity/lib.sh}. Sibling plans/reports on build/adv-accept, build/adv-mixer, build/adv-cache. Not opened: anything else under docs/, site/, proto-metal/, git log, other branches.
1. The target, restated from the spec and the code
1.1 What of the header reaches the load addresses
The miner's header bytes and nonce reach the hash ONLY through the init words (bind.rs, spec 1.6):
I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32)
H is the 32-byte pre-PoW hash, nonce_hi the high 32 bits of the nonce; the low 32 bits are the lane nonce n. seed_words_from_bytes is FNV-1a 64 under four salted bases plus a murmur finaliser, so every header byte enters all eight words of I. Registers init as r[i] = splitmix32((n XOR I[i]) + 0x9e3779b9*(i+1)) XOR I[(i+1)&7] (verify.rs). The program (from the epoch seed) does NOT depend on the header; only I and n do. The attacker's levers are H (a fresh 256-bit value per header, one header-hash each) and nonce_hi (a free 32-bit re-derivation of I, one FNV pass, no header hash).
Load address (verify.rs load_index, spec 1.13.1): y = rotl(xM, R); idx = ((y & (MASK>>k)) | (off<<(D-k))) & MASK, x the source register, D=28 on devnet, k=min(win,2). Physical byte address is 4idx under any era interleave. A load site whose source has no dataflow path from an earlier load in the same evaluation has an address computable from I and n with ALU only (predictable); a site that reads a loaded value cannot be addressed without the load first. I count the predictable set per program.
1.2 The counts I move
A unit is 32 lanes x 8 iterations x 16 sites = 4,096 loads of 4 bytes from 2^28 words. Honest denominator (chip-model-v3 s1): 128 distinct items per hash. The question: can a header search find a 32-lane group whose 128 loads per hash (or 4,096 per unit) cluster into fewer DRAM rows (2KiB and 8KiB), cache lines (64B), or items (64B) than a random group, and mine it above the honest rate on a card bound by random 4-byte reads.
2. Questions, in order
| # | Question | Method | Tool | Known-failed shape (must fire) | Gate | Box-hours |
|---|---|---|---|---|---|---|
| Q1 | What of the header reaches the address, and through how much mixing | Read + a diffusion probe: flip one bit of H / nonce_hi, measure Hamming weight of the change in each I word, each register, and each of the 4,096 load indices of a unit, over 2^12 header pairs on the two real programs | adv-accept-2 diffuse | a harness-mirror patch that sets idx = f(I) only (no register dependence) must show flips propagate to the address; the real path must show full avalanche (addresses near 50% changed) | every address bit flips with prob 0.5 +/- 3 sigma on the real path | 0.3 |
| Q2 | Distribution of distinct DRAM rows (2KiB=512 words, 8KiB=2048 words), lines (64B=16 words) and items per 32-lane unit and per hash over >=10^6 pre-PoW hashes, for the two real programs and several drawn ones; tails at 1e-3, 1e-4, 1e-5 vs a random baseline | Mirror the warp interpreter (as f8-uniform does), feed 10^6 random H (and a nonce_hi sweep), record per-unit and per-hash distinct rows/lines/items; histogram and quantiles; a SplitMix64 random-address control of the same shape | adv-accept-2 rows | --plant const-site (one lane-constant load site) and --plant tiny-window (win=2 on every site) must show the clustering at once (distinct counts collapse); clean programs sit at the random baseline | best-case tail within the random baseline's own extreme-value spread | 6 (sharded over idle cores, one log per shard) |
| Q3 | The price: search cost in hashes per found group vs loads saved; does any grind net >1% of rate on a card bound by random 4-byte reads | Analytic from Q2's tail: if the best 1e-k group saves dL loads, a card at R reads/s bound mines the found group at 128/(128-dL) higher, but the search costs ~10^k hashes per found group, each hash is itself 128 reads; net rate = gain / (1 + search_reads/useful_reads). State the model; GPU measurement BLOCKED | adv-accept-2 price (arithmetic beside rows) | a hand check: a 10% loads-saved group found at rate 1e-4 nets <<1% after search cost | no grind nets >1% | 0.1 |
| Q4 | Does rule (c) / (c'') bound per-hash locality or only per-program | Read: (c) LaneConstantSite bounds per (iteration, instruction) lane spread on the 64 FIXED accept nonces; (c'') bounds per-site distinct INDICES over 2^20 evaluations. Neither is keyed on the header. Measure whether a header outside the 64 accept nonces can cluster a unit that the rule passed; compare the rule's own distinct-index ratio to Q2's per-unit row counts | reuses Q2 | the planted tiny-window program must be REJECTED by accept::check (the rule catches the per-program clustering) while Q2 shows a clean program's per-unit tail is header-independent | the rule rejects the plant; headers do not move a clean program's tail | 0.5 |
Known-failed shape overall: a planted program with a lane-constant load site or a tiny window; Q2's harness must find its clustering at once, and Q4 must show accept::check rejects it. A result is a BREAK (method, counted gain, command, seed) or a BOUND (what was searched, how far, the margin). "Nothing found" counts only with its effort.
3. Running
Build through tools/build-remote.sh --box 2 -- build --release. Runs over 10 min start from run-box.sh in the harness dir with nohup nice -n 10, a pid file beside the log under /srv/builds/igneum-wt-adv-accept-2/adv/, and the yield rule: poll /srv/builds/_locks every 5 s, SIGSTOP the process group while any build- or quiet lock is held, SIGCONT when clear (the pattern of infra/build-server/capacity). Kill by pid file only. The 10^6 row sweep is sharded by header range across idle cores, one log per shard. Queue files go to /srv/builds/_adv/accept/queue/NN-adv-accept-2-.sh, claimed with mkdir on /srv/builds/_adv/accept/claims/. Box-hours: 8 a reading, 16 the ask line. No GPU: the Q3 per-card confirmation is BLOCKED and says so.
Commit as igneum-labs; push only git push build adv-accept-2. No em dashes, short sentences, numbers in tables.