229 lines
19 KiB
Markdown
229 lines
19 KiB
Markdown
# F6: the verifier's worst case over 10^5 class v4 programs
|
|
|
|
Attack-pass row F6 (`docs/plans/cryptanalysis.md` 4.2; the record `docs/analysis/attack-pass-2026-10.md`), the box search.
|
|
The O-1.14 laptop relay run is not in this record (main runs it separately). Written 7 October 2026.
|
|
|
|
## Target
|
|
|
|
| Item | Value |
|
|
|---|---|
|
|
| Commit | `924288d1` on branch `attack-pass` (`igneum-pow` is byte-identical at the worktree HEAD `8e36faf6`: `git diff --stat 924288d1..HEAD -- igneum-pow` is empty) |
|
|
| Class | `--program-class v4`: generator 4 on `V4_CLASS` = `mx8+sh256x27` (`LoadClass::MX8` plus `ShadowClass { instrs: 256, reps: 27 }`), no era bytes (the same draw `igneum-pow bench --program-class v4 --seed S` makes) |
|
|
| Dataset | day `2026-10-03`, `Shape::for_class_day(V4_CLASS, 0)`: cache 2^26 words (256 MiB), mixer x8, dataset 2^28 words, memory-hard |
|
|
| Work per hash | 64 base instructions x 8 iterations (16 loads) plus 256 shadow instructions x 27 passes x 8 iterations = 55,296 shadow instructions, 101,192 counted ops at the 1.83 convention |
|
|
| Gate | 10 ms per 32-lane warp, cold, on the half-core proxy (plan 1.4 item 6; spec 01 section 1.9 and 1.11; `algorithm.md` 3.3 and 5.5) |
|
|
| Programs | 10^5 deterministic string seeds `attack-f6/0` to `attack-f6/99999` through the class v4 chain draw with its acceptance rule (5.22 percent needed a second or third attempt, max attempt 3) |
|
|
|
|
The era draw is not in the search: `generator.rs` draws the shadow block from `NONLOAD_WEIGHTS` with no era perturbation
|
|
(no `perturb` path exists in the code at this commit), and the era parameters change only the load addressing, not the op
|
|
counts. Every drawn program has exactly 48 non-load base instructions and 256 shadow instructions, so the verifier's cost
|
|
differs between programs only through the family mix (the per-family cost on the CPU) and the data.
|
|
|
|
## Known-failed shape
|
|
|
|
A drawn program whose verifier warp exceeds 10 ms cold on the half-core proxy. The acceptance rule (spec 1.4.6) bounds the
|
|
miner's side (distinctness, bias, saturation); nothing bounds the verifier's cost per program, and the half-core headroom
|
|
of the average program is 1.8 ms (`algorithm.md` 5.5), so a family mix that costs the CPU interpreter more than the average
|
|
could cross the gate.
|
|
|
|
## Harness
|
|
|
|
`tools/attack/f6-verifier/` (its own cargo crate, `igneum-pow` as a path dependency, the same release profile as the CLI:
|
|
opt-level 3, LTO, one codegen unit). It builds the day's dataset ONCE (`DatasetSource::new_shape`, 0.58 s on the box) and
|
|
swaps programs under it: `Epoch { program, dataset }` is only the pair, and `verify::hash_warp(&program, base, &dataset)`
|
|
takes both, so one dataset serves every program. The naive path (`igneum-pow bench` per seed) refills the cache every
|
|
time (370 ms) and would take 10 hours per core.
|
|
|
|
| Command | What it does |
|
|
|---|---|
|
|
| `attack-f6 scan --program-class v4 --count N --start S --threads T --cold-reps R --flush swap --out F` | program i = seed `attack-f6/<S+i>`; one CSV line per program: generation time, R timed warps (each after a flush), their min, the op counts by family for the base and the shadow block |
|
|
| `attack-f6 time (--program-class v4 \| --class dr736) --seeds-file F --cold-reps R --steady W --flush sweep` | the deep re-time: R cold warps (each after a 256 MiB write sweep, what the cache fill does before `bench`'s "single cold run"), max, median, min, a steady average of W warps, and `GATE 10 ms PASS/FAIL` on the max |
|
|
| `attack-f6 micro --program-class v4` | the genesis program with its shadow block rewritten to one family at a time against the same program with no shadow: the per-family cost of a shadow instruction (ranking weights only) |
|
|
| `attack-f6 load --program-class v4 --seconds 0` | hashes class v4 warps on the calling core until killed: the SMT sibling's load for the half-core proxy, the same class the 3.3 proxy ran on both siblings |
|
|
| `rank.py --scan ... --weights ... --column ... --top 50 --out-prefix P` | the proxy ranking (sum over families of weight x (8 x base count + 216 x shadow count)), the distribution (min, median, p99, p99.9, max with the seed), the worst-N lists, a no-intercept regression of time on the family counts |
|
|
|
|
Build line (from the crate directory):
|
|
`IGNEUM_AGENT=attack-f6 bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f6" --out <scratch> -- build --release`
|
|
(build 08:04 to 08:06 UTC, rustc 1.99.0, binary sha256 `89c35674...17f3f9`, on the box at
|
|
`/srv/builds/igneum-wt-attack/tools/attack/f6-verifier/target/release/attack-f6`).
|
|
|
|
Box scratch: `/srv/builds/igneum-wt-attack/target-attack-f6/` (not the `attack-f6/` the brief named: that path is an
|
|
untracked directory in the worktree mirror and `build-remote.sh`'s checkout step runs `git clean -fd` on the mirror
|
|
before every build of this worktree by any agent, which removed it once at 08:04 UTC; `target-*` is on the clean's keep
|
|
list, so this name survives). Logs there: `phase1.log`, `scan-0.csv`, `scan-50000.csv` (phase 1), `phase2a.log`,
|
|
`clock-2a.log`, `full-0.csv`, `phase2b.log`, `clock-2b.log`, `full-50000.csv`, `phase2c.log`, `clock-2c.log` (phase 2),
|
|
`smoke.log` (the functional check). Copies on the Mac under the session scratchpad `attack-f6/`.
|
|
|
|
Run lines:
|
|
|
|
| Phase | Hold | Cores | Line |
|
|
|---|---|---|---|
|
|
| 1, pre-screen (counts and a coarse time, no timing claim) | `flock -s` per 50,000-program chunk (2.6 min each) | `nice -n 10 taskset -c 42-47,90-95`, 12 threads | `attack-f6 scan --program-class v4 --count 50000 --start {0,50000} --threads 12 --cold-reps 2 --flush swap` |
|
|
| 2A, firings, micro, full pass first half | `flock -x` | `nice -n 19 taskset -c 40`; half-core: core 88 running `attack-f6 load` | `phase2a.sh` |
|
|
| 2B, full pass second half, worst 50 by proxy and worst 50 by coarse time re-timed on both proxies | `flock -x` | same | `phase2b.sh` |
|
|
| 2C, worst 1,000 by the full one-core pass on the half-core; worst 10 deep re-timed on both proxies | `flock -x` | same | `phase2c.sh` |
|
|
|
|
The clock of cores 40 and 88 (`scaling_cur_freq`) and the load average were read every 5 s during every exclusive hold
|
|
(`clock-2*.log`).
|
|
|
|
## The two firings (batch A, exclusive hold taken 08:15:20 UTC, core 40 at 3,799.9 MHz throughout, `clock-2a.log`)
|
|
|
|
`attack-f6 time ... --cold-reps 5 --steady 20 --flush sweep`, the gate applied to the max of the 5 cold warps
|
|
(`phase2a.log`). Lane-0 vectors equal the known ones (dr736 `e23d389f3eea0c83`, the Mac's).
|
|
|
|
| Case | Proxy | Cold max / median / min (ms) | Steady, avg of 20 | Known value (`algorithm.md` 3.3) | Verdict |
|
|
|---|---|---|---|---|---|
|
|
| known-fail, `--class dr736`, genesis seed | one core | 10.284 / 9.982 / 9.952 | 9.562 | 10.51 cold, 9.76 steady | FAIL (fired) |
|
|
| known-fail, `--class dr736`, genesis seed | half-core | 14.135 / 13.795 / 13.400 | 13.225 | 15.49 | FAIL (fired) |
|
|
| known-pass, `--program-class v4`, genesis seed | one core | 5.156 / 5.140 / 5.135 | 4.909 | 5.06 cold, 4.90 steady | PASS (fired) |
|
|
| known-pass, `--program-class v4`, genesis seed | half-core | 8.624 / 8.510 / 8.389 | 8.268 | 8.23 | PASS (fired) |
|
|
|
|
The harness reads the known-fail class over the gate and the known-pass class under it on both proxies, within 2 percent
|
|
of the 3.3 one-core numbers and within 9 percent on the half-core (the earlier half-core run loaded the sibling with
|
|
`igneum-pow bench` of the same class; this one hashes class v4 warps on it continuously).
|
|
|
|
## What one program can move (batch A, `micro`, core 40 solo)
|
|
|
|
The genesis class v4 program with no shadow block: 4.511 ms steady; with its drawn shadow: 4.920 ms. So the whole 55,296-
|
|
instruction shadow block costs 0.41 ms per warp on this core (7.4 us per 1,000 shadow instructions; `model.py` carries
|
|
7.0) and the base program with its 128 loads and 4,096 item derivations costs the other 4.5 ms. The verifier's cost is
|
|
92 percent dataset derivation (the x8 mixer, 8 dependent cache reads per item), which no drawn program changes: every
|
|
class v4 program has 16 loads and the acceptance rule's distinctness test keeps the items per warp near 4,096. The family
|
|
mix of the shadow can move at most a fraction of 0.41 ms. The per-family rewrite (the 256 shadow instructions all one
|
|
family) reads add 4.798, sub 4.703, xor 4.683, rotl 4.818, mad 4.805, shfl 5.042, rotr 5.160 ms per warp; the mul,
|
|
mulhi and or rows (0.97, 1.55, 1.13 ms) are degenerate (the registers collapse to 0 or all-ones, every lane then loads
|
|
the same item and the memory side vanishes) and are not ALU costs. Batch A's micro line printed its per-1,000 column
|
|
1,000x too small (ns per instruction); fixed in the source, the numbers above are the ms-per-warp column, which is right.
|
|
|
|
## Full one-core pass A (batch A, 50,000 programs, core 40 solo, one cold warp each after a program swap, `full-0.csv`)
|
|
|
|
617 s for 50,000 programs (12.3 ms each, 2.5 ms of it generation and acceptance). The per-family regression on the
|
|
50,000 timings (no intercept) gives 86 to 98 us per 1,000 executed instructions by family, R^2 0.001: the family mix
|
|
explains none of the program-to-program variation. Those regression weights (add 89.1, sub 87.7, mul 90.6, mulhi 98.4,
|
|
xor 89.6, or 86.0, rotl 89.4, rotr 86.3, mad 91.2, shfl 95.6) are the proxy weights used for the worst-50-by-proxy list.
|
|
|
|
| Pass A | n | min | median | p99 | p99.9 | max |
|
|
|---|---|---|---|---|---|---|
|
|
| all rows | 50,000 | 4.610 (`attack-f6/1278`) | 4.950 | 6.753 | 7.834 | 10.710 (`attack-f6/26705`) |
|
|
| rows outside the three disturbed blocks | 44,000 | 4.610 | 4.948 | 5.606 | 5.662 | 6.194 (`attack-f6/48484`) |
|
|
|
|
The rows over 6 ms sit in three 2,000-program blocks (26,000 to 27,999: 337 rows; 36,000 to 37,999: 300; 38,000 to
|
|
39,999: 294) and nowhere else (0 in each of the other 22 blocks); the neighbours of the 10.71 ms seed, unrelated programs,
|
|
all read 7.6 to 8.5 ms. `clock-2a.log` shows core 88 (the idle sibling of a solo run) at 3.8 GHz for stretches in those
|
|
minutes (08:21:25, 08:21:45, 08:22:05 to 08:22:25 UTC) and core 40 dipping to 3.68 GHz at 08:21:00: a foreign process
|
|
(the box's hands run outside the measure lock) sat on the sibling, which is the half-core condition, and those rows read
|
|
half-core numbers. They are not program properties and not numbers; the three blocks are re-scanned in batch B, and from
|
|
batch B on a per-second sampler logs every process whose last CPU was 40 or 88 (`clock-2b.log`, `clock-2c.log`) so a
|
|
disturbed row can be named.
|
|
|
|
## Batches B and C: queued, starved of the exclusive hold (state at 09:46 UTC)
|
|
|
|
Batch B (the second 50,000 of the full one-core pass, the re-scan of the three disturbed blocks, the worst 50 by proxy
|
|
and the worst 50 by phase-1 coarse time re-timed 10 cold reps each on both proxies) was queued with `flock -x -w 7200`
|
|
at 08:33:56 UTC and had not taken the file by 09:46 UTC. Linux `flock` gives a pending exclusive waiter no priority over
|
|
new shared takers; eight lanes re-take the file in chunks (17 to 38 shared holders at every reading, F2 spawning many
|
|
short ones, one F9 hold 33 minutes old at 09:06, over the 30-minute cap), so the file is never free. Left on the box:
|
|
`run2b-retry.sh` re-queues batch B up to four more times (2 h each); `run2c-auto.sh` waits for `BATCH B DONE`, builds the
|
|
batch C lists on the box (`mklists.py`: the worst 1,000 and worst 10 by the full one-core pass, pass A's three disturbed
|
|
blocks replaced by their re-scan) and queues batch C (the worst 1,000 on the half-core at 2 cold reps each, the worst 10
|
|
at 20 cold reps on both proxies). A marker sits beside the lock (`/srv/builds/_locks/measure.wanted-by-attack-f6`). When
|
|
they land, `timeparse.py --log phase2b.log --clock clock-2b.log` (and `2c`) prints the per-seed tables with the sampler's
|
|
foreign-process column, and this record is completed.
|
|
|
|
## Numbers so far, the gate, the verdict
|
|
|
|
| Quantity | Value | Where |
|
|
|---|---|---|
|
|
| Programs drawn and counted (phase 1) | 100,000 | `scan-0.csv`, `scan-50000.csv` |
|
|
| Programs timed cold on core 40 alone (one-core proxy) | 50,000 (44,000 clean, 6,000 in disturbed blocks awaiting the re-scan) | `full-0.csv` |
|
|
| One-core cold, clean rows: min / median / p99 / p99.9 / max | 4.610 / 4.948 / 5.606 / 5.662 / 6.194 ms (`attack-f6/48484`), single warps, not yet re-timed | `full-0.csv` |
|
|
| Genesis class v4, half-core, max of 5 cold | 8.624 ms | `phase2a.log` |
|
|
| Half-core over one-core, genesis class v4 | 1.67x (8.624 / 5.156) | `phase2a.log` |
|
|
| Programs timed on the half-core proxy | 1 (the genesis seed) | `phase2a.log` |
|
|
| Core 40 clock during every exclusive timing | 3,799.9 MHz (dips to 3,680 MHz only in the disturbed minutes) | `clock-2a.log` |
|
|
|
|
Gate line: the worst program under 10 ms cold on the half-core proxy. Not yet measured: no drawn program other than the
|
|
genesis seed has a half-core number, and the one-core worst (6.194 ms, a single warp with no sampler running) has not been
|
|
re-timed. Carried to the half-core at the genesis ratio it would read 6.19 x 1.67 = 10.3 ms, over the gate; carried at the
|
|
additive half-core cost of the genesis program (8.624 - 5.156 = 3.47 ms) it would read 9.66 ms, under it by 0.34 ms. The
|
|
2.5x bracket of `algorithm.md` 5.5 is between. Whether `attack-f6/48484` (and the other clean rows over 5.6 ms: 1 in 100
|
|
of the pass) is a program property or a short disturbance is what batch C's 20-rep re-time with the sampler decides;
|
|
the record says FINDING if its half-core cold max reads 10 ms or more.
|
|
|
|
Verdict: INCOMPLETE. 100,000 programs drawn and ranked, 50,000 timed on the one-core proxy, 0 of the worst re-timed on
|
|
the half-core proxy; the two firings fired; the box search's second half and the half-core re-times are queued and
|
|
starved of the exclusive hold.
|
|
|
|
## The ladder ceiling implied so far
|
|
|
|
`algorithm.md` 5.5 and `model.py --section ladder` set the ceiling from the half-core headroom at 12.1 us per 1,000 shadow
|
|
instructions, N = 101,192 + instructions x 1.83.
|
|
|
|
| Worst program on the half-core | Headroom to 10 ms | Shadow instructions it buys | Ceiling N (counted ops) |
|
|
|---|---|---|---|
|
|
| 8.23 ms (3.3's average, the published figure) | 1.77 ms | 146,000 | about 370,000 |
|
|
| 8.62 ms (genesis seed, this run's max of 5) | 1.38 ms | 114,000 | about 310,000 |
|
|
| 9.66 ms (48484 if the additive carry holds) | 0.34 ms | 28,000 | about 152,000 |
|
|
| 10.3 ms (48484 if the 1.67x carry holds) | none | 0 | below today's 101,192: the floor rung 100,000 is the ceiling |
|
|
|
|
The proposed genesis ladder {100,000; 130,000; 200,000; 330,000; 650,000; 1,000,000} already exceeds the 310,000 ceiling
|
|
at its fourth rung on the genesis program alone; on the worst program the ceiling could be the floor. This is the
|
|
ladder's own open question (plan 1.1, "ceiling set by the verifier"), and the number that sets it is the half-core
|
|
worst case still queued.
|
|
|
|
## Consequences per user tier (at the numbers measured so far; the model's table of 5.5 at 10 ms beside them)
|
|
|
|
| | At 8.62 ms (genesis, half-core max) | At 9.66 ms (worst, additive carry, unverified) | At 10.3 ms (worst, 1.67x carry, unverified) | Model at 10 ms |
|
|
|---|---|---|---|---|
|
|
| A node on a 2019-class laptop core at 1 bps | 0.9 percent of one core | 1.0 | 1.0 | 1 |
|
|
| At 10 bps (the Devnet 2 experiment) | 8.6 percent of one core | 9.7 | 10.3 | 10 |
|
|
| IBD over the 108,000-header pruning window, one core | 15.5 min | 17.4 | 18.5 | 18 |
|
|
| Header flood: invalid headers per second that saturate one core | 116 | 104 | 97 | 100 |
|
|
| A pool core verifying shares, shares per second per core | 116 | 104 | 97 | 100 |
|
|
|
|
What each tier does with it: a home miner (8, 12, 16, 24 or 32 GB card, any vendor, any OS) runs a node that spends about
|
|
1 percent of one CPU core on the hash at 1 bps whatever the drawn program, and 9 to 10 percent at 10 bps; the card is not
|
|
involved. A rig is the same per node. A pool verifying shares at 100 per second per core needs one core per 100 shares
|
|
per second at the worst program, 116 at the average: a pool that sized its share verification at the average loses 14
|
|
percent of its per-core headroom on the worst program, so pools size at 97 shares per second per core (the 10 ms figure)
|
|
and never at the average. A node under header flood holds at about 100 invalid headers per second per core on any
|
|
program, the M15 figure. The 2019-class core itself is still the half-core proxy until the O-1.14 laptop run lands
|
|
(main's lane).
|
|
|
|
What this lane does about it: completes batches B and C when the hold comes (automatic, on the box); if the worst
|
|
program reads 10 ms or more on the half-core proxy, the finding goes to main with the seed, the reproduction line
|
|
`igneum-pow bench --program-class v4 --seed <seed> --day 2026-10-03 --warps 50` on core 40 and the half-core, and the
|
|
proposed fix: an acceptance-rule bound on verifier cost (a per-program cost model over the family counts checked at
|
|
draw time, a redraw when it exceeds the bound, exactly as rule (c) redraws on bias) and the ladder's ceiling set from
|
|
the measured worst, not the average; `igneum-pow` is not edited by this lane.
|
|
|
|
## Batches B and C landed (7 October 2026, 13:5x UTC, cores 40 and 88 under the per-core lease)
|
|
|
|
The measure file was retired at 13:3x UTC and replaced by per-core leases; cores 40 and 88 were leased to this row
|
|
(`/srv/builds/_bin/lease cores 40,88 --owner attack-pass`), so the batches ran with nothing else on those cores while
|
|
builds continued on the rest of the box. Core 40's clock (`clock-2b.log`, `clock-2c.log`): median 3,799.9 MHz in both
|
|
batches, 954 of 958 and 57 of 63 samples at 3.7 GHz or more, the 1.5 GHz readings between runs.
|
|
|
|
| Batch | What | Programs | Reps | Worst program | Half-core cold max | Log |
|
|
|---|---|---|---|---|---|---|
|
|
| B | the second 50,000 of the one-core pass, then the worst 50 by proxy and the worst 50 by coarse time re-timed cold on both proxies | 200 re-timed | 10 | `attack-f6/87142` | 8.708 ms | `phase2b.log`, done 13:52:11Z |
|
|
| C | the worst 1,000 by the full one-core pass on the half-core, then the worst 10 at 20 cold reps on both proxies | 1,000 + 10 | 2, then 20 | `attack-f6/88521` (8.629), `attack-f6/15781` (8.414) | 8.629 ms | `phase2c.log`, done 13:57:45Z |
|
|
|
|
The one-core worst of the full pass (`attack-f6/48484`, 6.194 ms single warp) does not reach the half-core top ten:
|
|
its one-core reading was a short disturbance, as the batch C re-scan shows. Every program timed on the half-core
|
|
proxy reads under 9 ms.
|
|
|
|
## Gate line and verdict
|
|
|
|
Gate: the worst program under 10 ms cold on the half-core proxy (and on a 2019-class core: O-1.14, the i7-9700K row,
|
|
class v4 6.334 ms cold max). Result: the worst of 100,000 class v4 programs on the half-core proxy is 8.708 ms,
|
|
1.29 ms under the gate; the genesis program reads 8.624 on the same proxy, so the worst drawn program costs 1 percent
|
|
more than the genesis one and the distribution is tight (one-core clean rows 4.610 to 6.194 ms, p99 5.606).
|
|
Verdict: PASS. The two firings fired (dr736 FAIL at 15.49 ms half-core; class v4 genesis PASS at 8.62). Consequences
|
|
per tier: a 2019-class node verifying the worst class v4 program spends 0.9 percent of one core at 1 bps and 9 percent
|
|
at 10 bps on the pessimistic proxy; a header flood needs about 115 invalid headers a second to saturate one such
|
|
core; a pool verifies about 115 shares a second per core; IBD of 108,000 headers is about 16 minutes of one core.
|
|
Implied ladder ceiling on the half-core proxy from the worst program: 10 - 8.708 = 1.29 ms of headroom buys about
|
|
106,000 shadow instructions, N about 300,000 counted ops at the 1.83 convention (approximate), against 370,000
|
|
from the genesis program's headroom; the ladder's ceiling should be read from the worst program, not the genesis
|
|
one, so rung 2 (199,600) stays admissible and rung 3 (330,700) does not on this proxy.
|