17 KiB
F6: the verifier's worst case over 10^5 class v4 programs
Attack-pass row F6 (docs/plans/cryptanalysis.md 4.2; the record docs/analysis/attack-pass-2026-10.md), the box search.
The O-1.14 laptop relay run is not in this record (main runs it separately). Written 7 October 2026.
Target
| Item | Value |
|---|---|
| Commit | 924288d1 on branch attack-pass (igneum-pow is byte-identical at the worktree HEAD 8e36faf6: git diff --stat 924288d1..HEAD -- igneum-pow is empty) |
| Class | --program-class v4: generator 4 on V4_CLASS = mx8+sh256x27 (LoadClass::MX8 plus ShadowClass { instrs: 256, reps: 27 }), no era bytes (the same draw igneum-pow bench --program-class v4 --seed S makes) |
| Dataset | day 2026-10-03, Shape::for_class_day(V4_CLASS, 0): cache 2^26 words (256 MiB), mixer x8, dataset 2^28 words, memory-hard |
| Work per hash | 64 base instructions x 8 iterations (16 loads) plus 256 shadow instructions x 27 passes x 8 iterations = 55,296 shadow instructions, 101,192 counted ops at the 1.83 convention |
| Gate | 10 ms per 32-lane warp, cold, on the half-core proxy (plan 1.4 item 6; spec 01 section 1.9 and 1.11; algorithm.md 3.3 and 5.5) |
| Programs | 10^5 deterministic string seeds attack-f6/0 to attack-f6/99999 through the class v4 chain draw with its acceptance rule (5.22 percent needed a second or third attempt, max attempt 3) |
The era draw is not in the search: generator.rs draws the shadow block from NONLOAD_WEIGHTS with no era perturbation
(no perturb path exists in the code at this commit), and the era parameters change only the load addressing, not the op
counts. Every drawn program has exactly 48 non-load base instructions and 256 shadow instructions, so the verifier's cost
differs between programs only through the family mix (the per-family cost on the CPU) and the data.
Known-failed shape
A drawn program whose verifier warp exceeds 10 ms cold on the half-core proxy. The acceptance rule (spec 1.4.6) bounds the
miner's side (distinctness, bias, saturation); nothing bounds the verifier's cost per program, and the half-core headroom
of the average program is 1.8 ms (algorithm.md 5.5), so a family mix that costs the CPU interpreter more than the average
could cross the gate.
Harness
tools/attack/f6-verifier/ (its own cargo crate, igneum-pow as a path dependency, the same release profile as the CLI:
opt-level 3, LTO, one codegen unit). It builds the day's dataset ONCE (DatasetSource::new_shape, 0.58 s on the box) and
swaps programs under it: Epoch { program, dataset } is only the pair, and verify::hash_warp(&program, base, &dataset)
takes both, so one dataset serves every program. The naive path (igneum-pow bench per seed) refills the cache every
time (370 ms) and would take 10 hours per core.
| Command | What it does |
|---|---|
attack-f6 scan --program-class v4 --count N --start S --threads T --cold-reps R --flush swap --out F |
program i = seed attack-f6/<S+i>; one CSV line per program: generation time, R timed warps (each after a flush), their min, the op counts by family for the base and the shadow block |
attack-f6 time (--program-class v4 | --class dr736) --seeds-file F --cold-reps R --steady W --flush sweep |
the deep re-time: R cold warps (each after a 256 MiB write sweep, what the cache fill does before bench's "single cold run"), max, median, min, a steady average of W warps, and GATE 10 ms PASS/FAIL on the max |
attack-f6 micro --program-class v4 |
the genesis program with its shadow block rewritten to one family at a time against the same program with no shadow: the per-family cost of a shadow instruction (ranking weights only) |
attack-f6 load --program-class v4 --seconds 0 |
hashes class v4 warps on the calling core until killed: the SMT sibling's load for the half-core proxy, the same class the 3.3 proxy ran on both siblings |
rank.py --scan ... --weights ... --column ... --top 50 --out-prefix P |
the proxy ranking (sum over families of weight x (8 x base count + 216 x shadow count)), the distribution (min, median, p99, p99.9, max with the seed), the worst-N lists, a no-intercept regression of time on the family counts |
Build line (from the crate directory):
IGNEUM_AGENT=attack-f6 bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f6" --out <scratch> -- build --release
(build 08:04 to 08:06 UTC, rustc 1.99.0, binary sha256 89c35674...17f3f9, on the box at
/srv/builds/igneum-wt-attack/tools/attack/f6-verifier/target/release/attack-f6).
Box scratch: /srv/builds/igneum-wt-attack/target-attack-f6/ (not the attack-f6/ the brief named: that path is an
untracked directory in the worktree mirror and build-remote.sh's checkout step runs git clean -fd on the mirror
before every build of this worktree by any agent, which removed it once at 08:04 UTC; target-* is on the clean's keep
list, so this name survives). Logs there: phase1.log, scan-0.csv, scan-50000.csv (phase 1), phase2a.log,
clock-2a.log, full-0.csv, phase2b.log, clock-2b.log, full-50000.csv, phase2c.log, clock-2c.log (phase 2),
smoke.log (the functional check). Copies on the Mac under the session scratchpad attack-f6/.
Run lines:
| Phase | Hold | Cores | Line |
|---|---|---|---|
| 1, pre-screen (counts and a coarse time, no timing claim) | flock -s per 50,000-program chunk (2.6 min each) |
nice -n 10 taskset -c 42-47,90-95, 12 threads |
attack-f6 scan --program-class v4 --count 50000 --start {0,50000} --threads 12 --cold-reps 2 --flush swap |
| 2A, firings, micro, full pass first half | flock -x |
nice -n 19 taskset -c 40; half-core: core 88 running attack-f6 load |
phase2a.sh |
| 2B, full pass second half, worst 50 by proxy and worst 50 by coarse time re-timed on both proxies | flock -x |
same | phase2b.sh |
| 2C, worst 1,000 by the full one-core pass on the half-core; worst 10 deep re-timed on both proxies | flock -x |
same | phase2c.sh |
The clock of cores 40 and 88 (scaling_cur_freq) and the load average were read every 5 s during every exclusive hold
(clock-2*.log).
The two firings (batch A, exclusive hold taken 08:15:20 UTC, core 40 at 3,799.9 MHz throughout, clock-2a.log)
attack-f6 time ... --cold-reps 5 --steady 20 --flush sweep, the gate applied to the max of the 5 cold warps
(phase2a.log). Lane-0 vectors equal the known ones (dr736 e23d389f3eea0c83, the Mac's).
| Case | Proxy | Cold max / median / min (ms) | Steady, avg of 20 | Known value (algorithm.md 3.3) |
Verdict |
|---|---|---|---|---|---|
known-fail, --class dr736, genesis seed |
one core | 10.284 / 9.982 / 9.952 | 9.562 | 10.51 cold, 9.76 steady | FAIL (fired) |
known-fail, --class dr736, genesis seed |
half-core | 14.135 / 13.795 / 13.400 | 13.225 | 15.49 | FAIL (fired) |
known-pass, --program-class v4, genesis seed |
one core | 5.156 / 5.140 / 5.135 | 4.909 | 5.06 cold, 4.90 steady | PASS (fired) |
known-pass, --program-class v4, genesis seed |
half-core | 8.624 / 8.510 / 8.389 | 8.268 | 8.23 | PASS (fired) |
The harness reads the known-fail class over the gate and the known-pass class under it on both proxies, within 2 percent
of the 3.3 one-core numbers and within 9 percent on the half-core (the earlier half-core run loaded the sibling with
igneum-pow bench of the same class; this one hashes class v4 warps on it continuously).
What one program can move (batch A, micro, core 40 solo)
The genesis class v4 program with no shadow block: 4.511 ms steady; with its drawn shadow: 4.920 ms. So the whole 55,296-
instruction shadow block costs 0.41 ms per warp on this core (7.4 us per 1,000 shadow instructions; model.py carries
7.0) and the base program with its 128 loads and 4,096 item derivations costs the other 4.5 ms. The verifier's cost is
92 percent dataset derivation (the x8 mixer, 8 dependent cache reads per item), which no drawn program changes: every
class v4 program has 16 loads and the acceptance rule's distinctness test keeps the items per warp near 4,096. The family
mix of the shadow can move at most a fraction of 0.41 ms. The per-family rewrite (the 256 shadow instructions all one
family) reads add 4.798, sub 4.703, xor 4.683, rotl 4.818, mad 4.805, shfl 5.042, rotr 5.160 ms per warp; the mul,
mulhi and or rows (0.97, 1.55, 1.13 ms) are degenerate (the registers collapse to 0 or all-ones, every lane then loads
the same item and the memory side vanishes) and are not ALU costs. Batch A's micro line printed its per-1,000 column
1,000x too small (ns per instruction); fixed in the source, the numbers above are the ms-per-warp column, which is right.
Full one-core pass A (batch A, 50,000 programs, core 40 solo, one cold warp each after a program swap, full-0.csv)
617 s for 50,000 programs (12.3 ms each, 2.5 ms of it generation and acceptance). The per-family regression on the 50,000 timings (no intercept) gives 86 to 98 us per 1,000 executed instructions by family, R^2 0.001: the family mix explains none of the program-to-program variation. Those regression weights (add 89.1, sub 87.7, mul 90.6, mulhi 98.4, xor 89.6, or 86.0, rotl 89.4, rotr 86.3, mad 91.2, shfl 95.6) are the proxy weights used for the worst-50-by-proxy list.
| Pass A | n | min | median | p99 | p99.9 | max |
|---|---|---|---|---|---|---|
| all rows | 50,000 | 4.610 (attack-f6/1278) |
4.950 | 6.753 | 7.834 | 10.710 (attack-f6/26705) |
| rows outside the three disturbed blocks | 44,000 | 4.610 | 4.948 | 5.606 | 5.662 | 6.194 (attack-f6/48484) |
The rows over 6 ms sit in three 2,000-program blocks (26,000 to 27,999: 337 rows; 36,000 to 37,999: 300; 38,000 to
39,999: 294) and nowhere else (0 in each of the other 22 blocks); the neighbours of the 10.71 ms seed, unrelated programs,
all read 7.6 to 8.5 ms. clock-2a.log shows core 88 (the idle sibling of a solo run) at 3.8 GHz for stretches in those
minutes (08:21:25, 08:21:45, 08:22:05 to 08:22:25 UTC) and core 40 dipping to 3.68 GHz at 08:21:00: a foreign process
(the box's hands run outside the measure lock) sat on the sibling, which is the half-core condition, and those rows read
half-core numbers. They are not program properties and not numbers; the three blocks are re-scanned in batch B, and from
batch B on a per-second sampler logs every process whose last CPU was 40 or 88 (clock-2b.log, clock-2c.log) so a
disturbed row can be named.
Batches B and C: queued, starved of the exclusive hold (state at 09:46 UTC)
Batch B (the second 50,000 of the full one-core pass, the re-scan of the three disturbed blocks, the worst 50 by proxy
and the worst 50 by phase-1 coarse time re-timed 10 cold reps each on both proxies) was queued with flock -x -w 7200
at 08:33:56 UTC and had not taken the file by 09:46 UTC. Linux flock gives a pending exclusive waiter no priority over
new shared takers; eight lanes re-take the file in chunks (17 to 38 shared holders at every reading, F2 spawning many
short ones, one F9 hold 33 minutes old at 09:06, over the 30-minute cap), so the file is never free. Left on the box:
run2b-retry.sh re-queues batch B up to four more times (2 h each); run2c-auto.sh waits for BATCH B DONE, builds the
batch C lists on the box (mklists.py: the worst 1,000 and worst 10 by the full one-core pass, pass A's three disturbed
blocks replaced by their re-scan) and queues batch C (the worst 1,000 on the half-core at 2 cold reps each, the worst 10
at 20 cold reps on both proxies). A marker sits beside the lock (/srv/builds/_locks/measure.wanted-by-attack-f6). When
they land, timeparse.py --log phase2b.log --clock clock-2b.log (and 2c) prints the per-seed tables with the sampler's
foreign-process column, and this record is completed.
Numbers so far, the gate, the verdict
| Quantity | Value | Where |
|---|---|---|
| Programs drawn and counted (phase 1) | 100,000 | scan-0.csv, scan-50000.csv |
| Programs timed cold on core 40 alone (one-core proxy) | 50,000 (44,000 clean, 6,000 in disturbed blocks awaiting the re-scan) | full-0.csv |
| One-core cold, clean rows: min / median / p99 / p99.9 / max | 4.610 / 4.948 / 5.606 / 5.662 / 6.194 ms (attack-f6/48484), single warps, not yet re-timed |
full-0.csv |
| Genesis class v4, half-core, max of 5 cold | 8.624 ms | phase2a.log |
| Half-core over one-core, genesis class v4 | 1.67x (8.624 / 5.156) | phase2a.log |
| Programs timed on the half-core proxy | 1 (the genesis seed) | phase2a.log |
| Core 40 clock during every exclusive timing | 3,799.9 MHz (dips to 3,680 MHz only in the disturbed minutes) | clock-2a.log |
Gate line: the worst program under 10 ms cold on the half-core proxy. Not yet measured: no drawn program other than the
genesis seed has a half-core number, and the one-core worst (6.194 ms, a single warp with no sampler running) has not been
re-timed. Carried to the half-core at the genesis ratio it would read 6.19 x 1.67 = 10.3 ms, over the gate; carried at the
additive half-core cost of the genesis program (8.624 - 5.156 = 3.47 ms) it would read 9.66 ms, under it by 0.34 ms. The
2.5x bracket of algorithm.md 5.5 is between. Whether attack-f6/48484 (and the other clean rows over 5.6 ms: 1 in 100
of the pass) is a program property or a short disturbance is what batch C's 20-rep re-time with the sampler decides;
the record says FINDING if its half-core cold max reads 10 ms or more.
Verdict: INCOMPLETE. 100,000 programs drawn and ranked, 50,000 timed on the one-core proxy, 0 of the worst re-timed on the half-core proxy; the two firings fired; the box search's second half and the half-core re-times are queued and starved of the exclusive hold.
The ladder ceiling implied so far
algorithm.md 5.5 and model.py --section ladder set the ceiling from the half-core headroom at 12.1 us per 1,000 shadow
instructions, N = 101,192 + instructions x 1.83.
| Worst program on the half-core | Headroom to 10 ms | Shadow instructions it buys | Ceiling N (counted ops) |
|---|---|---|---|
| 8.23 ms (3.3's average, the published figure) | 1.77 ms | 146,000 | about 370,000 |
| 8.62 ms (genesis seed, this run's max of 5) | 1.38 ms | 114,000 | about 310,000 |
| 9.66 ms (48484 if the additive carry holds) | 0.34 ms | 28,000 | about 152,000 |
| 10.3 ms (48484 if the 1.67x carry holds) | none | 0 | below today's 101,192: the floor rung 100,000 is the ceiling |
The proposed genesis ladder {100,000; 130,000; 200,000; 330,000; 650,000; 1,000,000} already exceeds the 310,000 ceiling at its fourth rung on the genesis program alone; on the worst program the ceiling could be the floor. This is the ladder's own open question (plan 1.1, "ceiling set by the verifier"), and the number that sets it is the half-core worst case still queued.
Consequences per user tier (at the numbers measured so far; the model's table of 5.5 at 10 ms beside them)
| At 8.62 ms (genesis, half-core max) | At 9.66 ms (worst, additive carry, unverified) | At 10.3 ms (worst, 1.67x carry, unverified) | Model at 10 ms | |
|---|---|---|---|---|
| A node on a 2019-class laptop core at 1 bps | 0.9 percent of one core | 1.0 | 1.0 | 1 |
| At 10 bps (the Devnet 2 experiment) | 8.6 percent of one core | 9.7 | 10.3 | 10 |
| IBD over the 108,000-header pruning window, one core | 15.5 min | 17.4 | 18.5 | 18 |
| Header flood: invalid headers per second that saturate one core | 116 | 104 | 97 | 100 |
| A pool core verifying shares, shares per second per core | 116 | 104 | 97 | 100 |
What each tier does with it: a home miner (8, 12, 16, 24 or 32 GB card, any vendor, any OS) runs a node that spends about 1 percent of one CPU core on the hash at 1 bps whatever the drawn program, and 9 to 10 percent at 10 bps; the card is not involved. A rig is the same per node. A pool verifying shares at 100 per second per core needs one core per 100 shares per second at the worst program, 116 at the average: a pool that sized its share verification at the average loses 14 percent of its per-core headroom on the worst program, so pools size at 97 shares per second per core (the 10 ms figure) and never at the average. A node under header flood holds at about 100 invalid headers per second per core on any program, the M15 figure. The 2019-class core itself is still the half-core proxy until the O-1.14 laptop run lands (main's lane).
What this lane does about it: completes batches B and C when the hold comes (automatic, on the box); if the worst
program reads 10 ms or more on the half-core proxy, the finding goes to main with the seed, the reproduction line
igneum-pow bench --program-class v4 --seed <seed> --day 2026-10-03 --warps 50 on core 40 and the half-core, and the
proposed fix: an acceptance-rule bound on verifier cost (a per-program cost model over the family counts checked at
draw time, a redraw when it exceeds the bound, exactly as rule (c) redraws on bias) and the ladder's ceiling set from
the measured worst, not the average; igneum-pow is not edited by this lane.