251 lines
16 KiB
Markdown
251 lines
16 KiB
Markdown
# Finality v2 attack harness
|
|
|
|
Adversarial testing of Igneum sustained-mining finality (rule v2, `docs/spec/03-finality.md`) against a private
|
|
network of our own `igneumd` nodes. Each scenario has a pass criterion taken from spec section 3 and a measured
|
|
result. This is robustness and conformance testing of our own software, the practice upstream Kaspa and the
|
|
Ethereum clients follow. It is gate 3 work (the cryptographer owns gate 3).
|
|
|
|
The hostile behaviour lives only in test-only flags of `igneum-miner` (worktree `vendor/igneum-node-fin-attacks`,
|
|
branch `fin-attacks`), never in honest node or consensus code:
|
|
|
|
- `vmine <url> <secs> [--share f] [--bps f] [--label s] [--no-vote] [--equivocate] [--drop-votes] [--sybil b:bb:a:ab] [--pulse burst:on:period]`
|
|
drives the node over the same `getBlockTemplate` / `submitBlock` / `submitFinalityVote` RPCs a real miner uses.
|
|
The network runs with `skip_proof_of_work`, so a block is "found" on a Poisson clock at a chosen hash share and
|
|
the miner controls its share of blocks exactly. Difficulty still retargets from block cadence, so `--pulse`
|
|
exercises the DAA controller in the loop.
|
|
- `--equivocate` the miner signs a second, wrong hash at every index (S1).
|
|
- `--sybil b:bb:a:ab` one miner mints `b` keys with `bb` blocks each (below dust) and `a` keys with `ab`
|
|
blocks each (above dust), from a single process (S2, ledger F17).
|
|
- `--drop-votes` the miner strips the node's finality section from its coinbase, so its blocks carry no votes
|
|
or certificates, while it still votes over RPC (S4, ledger F3).
|
|
- `--pulse burst:on:period` the submit rate is multiplied by `burst` for `on` seconds of every `period` (S5,
|
|
ledger F14).
|
|
- `fin-rpc-attack <url>` submits malformed, mis-signed and replayed votes over `submitFinalityVote` (S8).
|
|
|
|
## Isolation (never touch the live devnet)
|
|
|
|
Ports 27800 and up, data under `/tmp/igneum-fin-attacks`, network id `igneum-devnet-800` (own handshake magic and
|
|
data directory). The live devnet (gRPC 26610, p2p 26611, observer 26640/26641/28640) and ports other agents use
|
|
(up to 27799) are never touched. Loopback addresses are never gossiped, so no link forms that a scenario did not
|
|
ask for. Everything started is stopped at the end and on SIGINT.
|
|
|
|
## Build
|
|
|
|
```
|
|
cd vendor/igneum-node/ && git worktree add -b fin-attacks ../igneum-node-fin-attacks master # once
|
|
cd ../igneum-node-fin-attacks
|
|
CARGO_TARGET_DIR=target nice -n 19 ~/.cargo/bin/cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow
|
|
```
|
|
|
|
This produces `target/release/igneumd` and `target/release/igneum-miner`. Override the location with `IGNEUMD`
|
|
and `IGNEUM_MINER`.
|
|
|
|
## Run
|
|
|
|
```
|
|
node tools/finality-attacks/run.mjs # the achievable catalogue, priority order 3,2,1,6,4,8,5
|
|
node tools/finality-attacks/run.mjs s1 s6 # named scenarios
|
|
node tools/finality-attacks/run.mjs --quick # short durations (smoke)
|
|
SCALE=0.7 node tools/finality-attacks/run.mjs # scale every duration
|
|
node tools/finality-attacks/run.mjs s3 --fast-time # the 60x fast-time profile (infra/fast-time/README.md): the weight
|
|
# window and min_daa are 120 DAA, so the first lock comes at about
|
|
# DAA 150 instead of 7,200; the node is the devnet-v4 integration build
|
|
```
|
|
|
|
Under the devnet parameters of 4 October 2026 (`min_daa` = window = 7,200 DAA) a scenario cannot see a lock before
|
|
DAA 7,200, which at the six miners' 6 blocks/s is 20 minutes of warm-up per scenario; `--fast-time` brings that to
|
|
20 s (s3 measured: 16 locks per node and PASS in 113 s wall, 4 Oct 2026).
|
|
|
|
Results print as a table and are written to `/tmp/igneum-fin-attacks/results.{txt,json}`; the S8 transcript is at
|
|
`/tmp/igneum-fin-attacks/s8-fin-rpc-attack.out`. Exit code is non-zero if any scenario failed.
|
|
|
|
## Rule v3 runner (4 October 2026, evening): `v3.mjs`
|
|
|
|
`node tools/finality-attacks/v3.mjs fold split50 split70` drives the fast-time 3-node network (ports 29700 and up, suffix 970, `/tmp/igneum-fin-v3`) with an emulated one-way delay on both proxied links (`DELAY_MS`, default 300; the proxy holds every byte, in place of tc/netem which macOS lacks) against the `finality-fixes` node build (`vendor/igneum-node/target-finality/release`), rule v3 on (`finality_v3_activation_daa` 0 in the merged override) or `--v2` for the control. `fold` counts the signers of the certificate each node holds per locked index (ledger F22), `split50` is the 3/3 split longer than the old bound W / (3 R) with the heal (F21), `split70` the 4/2 split with the 4 side at 70% of the weight (6B; exactly 4/6 is a knife edge under both rules). Run it through `tools/lock/with-lock.sh run` (the measure-only mode of 4 October 2026 evening: a functional run whose outputs are counts, locks and seconds does not block builds). Results in `/tmp/igneum-fin-v3/results-<rule>.md`; the record is `docs/bench-log.md`, "finality rule v3". `vote-timing.py` is the cloud-log analysis behind F22 (`infra/cloud-devnet/results/2026-10-04/f22-vote-timing.md`). The library now reads `IGNEUM_FIN_BASE_PORT`, `IGNEUM_FIN_SUFFIX`, `IGNEUM_FIN_TMP` and `IGNEUM_FIN_OVERRIDE_JSON`, so two harness networks can run side by side; `Proxy` takes `{ delayMs }`.
|
|
|
|
## The catalogue
|
|
|
|
| # | Scenario | Criterion (spec) |
|
|
|---|---|---|
|
|
| s3 | dishonest aggregators | other aggregators' certs still lock; a sub-quorum cert cannot lock (Q3); block-carried votes give participation (F3); lock latency < 2 s median |
|
|
| s2 | Sybil dust | dust keys zero weight; above-dust weight = blocks; total weight = voters' blue blocks; sortition by weight not key count (F17) |
|
|
| s1 | equivocation at scale | both keys stripped within one checkpoint on every node; no conflicting certificate; honest locks continue (3.6) |
|
|
| s6 | partition with the floor | 3/3: zero locks either side, resume on heal; 4/2: the 4 side locks, the 2 side does not (Q3 floor, 3.3.1) |
|
|
| s4 | vote-dropping producer | participation and locks unaffected; delay measured (F3) |
|
|
| s8 | malformed votes over RPC | rejected without a crash; node stays up (C2) |
|
|
| s5 | pulsed rental | burst weight proportional to its block share; cannot lock alone (F14) |
|
|
|
|
## What this harness does NOT cover (needs a finality-aware p2p probe, not built this session)
|
|
|
|
The `submitFinalityVote` RPC is the only way to inject a vote from outside a node; there is no submit-certificate
|
|
RPC, and certificates only enter a node over p2p message 70 (`protocol/flows/src/v10/finality.rs`) or by a node
|
|
building one from votes. So three things are out of reach without a p2p probe that speaks the fork's handshake and
|
|
message 70, which the harness worktree `vendor/igneum-node-harness` has for blocks but not for finality:
|
|
|
|
- Injecting a hand-crafted **sub-quorum certificate** to confirm it is rejected on the wire (S3). The in-node
|
|
path is covered by code review: `FinalityManager::ingest_certificate` verifies the aggregate signature and the
|
|
voter count, and `lock_test` requires both `3 x signed x P >= 2 x active_num` and `30 x signed >= 17 x total`
|
|
before a certificate locks, so a certificate below either threshold is recorded as `Certified`, never `Locked`.
|
|
- The **oversized-bitmap and replayed-certificate** half of S8 over message 70. `Certificate::read` bounds the
|
|
bitmap at `1 << 20` bytes and the relay flow (`v10/finality.rs`) bounds a message at `1 << 20` bytes and
|
|
returns `ProtocolError` on a malformed certificate (which disconnects the peer); this is code review, not a
|
|
live wire test.
|
|
- A true **vote-dropping relay** (S3/S4). Votes ride in blocks, not only in p2p messages, so a relay that drops
|
|
message 70 cannot suppress a vote that any block carries (the F3 design). The adversary that can withhold votes
|
|
is a block producer, which S4 covers with `--drop-votes`.
|
|
|
|
Scenario 7 (eclipse of one voter with an adversarial side chain) also needs the p2p probe to feed a private fork
|
|
and is not run this session.
|
|
|
|
## FAIL: S2, aggregator sortition is per key, not per weight (ledger F17)
|
|
|
|
`consensus/core/src/finality.rs` `is_aggregator(output, voters, aggregators)` returns
|
|
`output * voters < aggregators << 64`, where `voters` is the **count** of keys above dust, not their weight. So
|
|
the 8 aggregators per checkpoint are drawn uniformly over keys, exactly the per-key draw F17 flags for the shard
|
|
sortition of spec 7.2. Measured: one miner minting 199 above-dust keys made the voter count 199; a real signer's
|
|
chance of being sortition-eligible fell to about `8 / 199`, so almost no certificate named a legitimate aggregator
|
|
and the network fell back to "anyone MAY aggregate" (zero-aggregator certificates).
|
|
|
|
Blast radius is smaller than the shard case: the aggregator role grants no reward and no power, because any node
|
|
may aggregate and a certificate must still meet Q3 by **weight**, which a Sybil split does not change (S2 confirms
|
|
dust keys carry zero weight and the total equals the voters' blue blocks). So this is a liveness and tidiness
|
|
defect, not a safety one. It should still be fixed so the sortition means what the spec says and S2 never produces
|
|
8,192-voter certificates with inflated bitmaps (O-3.12).
|
|
|
|
Proposed fix (draw the 8 aggregators by weight, mirroring the spec 7.2 step-2 fix for the shard sortition; the
|
|
rule stays the cryptographer's to ratify, this is a diff for review, not a change made to the rule):
|
|
|
|
```diff
|
|
--- a/consensus/core/src/finality.rs
|
|
+++ b/consensus/core/src/finality.rs
|
|
@@ pub fn is_aggregator(output: u64, voters: u64, aggregators: u64) -> bool
|
|
-/// S1: a voter is an aggregator when its VRF output falls below `aggregators / voters` of the range. With 8 or
|
|
-/// fewer voters every voter is an aggregator.
|
|
-pub fn is_aggregator(output: u64, voters: u64, aggregators: u64) -> bool {
|
|
- if voters <= aggregators {
|
|
- return true;
|
|
- }
|
|
- // output / 2^64 < aggregators / voters <=> output * voters < aggregators * 2^64
|
|
- (output as u128) * (voters as u128) < (aggregators as u128) << 64
|
|
-}
|
|
+/// S1 (F17 fix): a voter is an aggregator when its VRF output falls below `aggregators * weight / total` of the
|
|
+/// range, so splitting weight across many keys does not buy aggregator tickets (each sub-key's threshold shrinks
|
|
+/// in proportion). With total weight at or below `aggregators`, every voter is eligible.
|
|
+pub fn is_aggregator_weighted(output: u64, weight: u64, total_weight: u64, aggregators: u64) -> bool {
|
|
+ if total_weight <= aggregators || weight == 0 {
|
|
+ return weight > 0;
|
|
+ }
|
|
+ // output / 2^64 < aggregators * weight / total <=> output * total < aggregators * weight * 2^64
|
|
+ (output as u128) * (total_weight as u128) < (aggregators as u128) * (weight as u128) << 64
|
|
+}
|
|
```
|
|
|
|
The call site in `consensus/src/processes/finality.rs` (`ingest_vote`, where `is_aggregator` is called with
|
|
`table.voters.len()`) passes the key's weight and the table total instead of the voter count. A voter with weight
|
|
`w` then holds an aggregator chance of `aggregators * w / total`, so a split into `n` keys of weight `w/n` holds
|
|
the same total chance it had as one key. This is the same shape as the spec 7.2 step-2 fix and O-5.1's one-key
|
|
versus 1,000-keys acceptance test applies unchanged.
|
|
|
|
## FAIL: S6A, the floor is time-bounded on a real DAG (spec 3.3.1, 3.7 item 8)
|
|
|
|
Measured: a 3/3 split after a 252-s shared warmup (window weight F = 1,439 DAA at the cut). One side crossed the
|
|
56.7% floor 84 s into a 90-s split and locked 8 checkpoints alone; the other side never did (0 conflicting
|
|
certificates only because the first side locked indices the second never reached). Both healed cleanly.
|
|
|
|
Why: the floor denominator is total weight in the window of C_i, and after the cut only the active side keeps
|
|
adding blocks to its own window while the other side's blocks are frozen. With share(T) = (F/2 + R T) / (F + R T)
|
|
(R the side's block rate), the floor is crossed at
|
|
|
|
T* = 2F / (13R)
|
|
|
|
74 s predicted at F = 1,439 and R = 3 blocks/s, 84 s measured (sibling losses lower R). `sim/finality_v2.py`
|
|
holds weights fixed during a partition, so it could not see this; spec 3.7 item 8 lists it as outside the model.
|
|
|
|
Mainnet scale, approximate (formula only, window slide and DAA lag ignored): full 30-day window F = 2,592,000,
|
|
1 block/s total:
|
|
|
|
| Split | Majority R | T* |
|
|
|---|---|---|
|
|
| 50/50 | 0.5 | 9.2 days |
|
|
| 55/45 | 0.55 | 2.1 days |
|
|
| 60/40 | 0.6 | locks at once (3.3.1 already says so) |
|
|
|
|
Minimum fix: state the bound in spec 3.3.1 and in the exchange guidance of 3.9 (a partition longer than T* can
|
|
end with one side holding a lock the other never saw).
|
|
|
|
Rule option for gate 3 (a diff for review, not applied; it trades liveness for the bound). Evaluate the floor
|
|
against the weight table of the last locked checkpoint while no newer lock exists, so a stalled side cannot lift
|
|
its own share of total by mining alone:
|
|
|
|
```diff
|
|
--- a/consensus/src/processes/finality.rs
|
|
+++ b/consensus/src/processes/finality.rs
|
|
@@ fn evaluate(&self, state: &mut FinalityState, index: u64)
|
|
let table = self.voters_at(cp.hash, state);
|
|
+ // Floor anchor (gate 3 option): while no lock is newer than the last one, the 17/30 test uses the
|
|
+ // weight table of the last locked checkpoint, so a side that keeps mining during a stall cannot raise
|
|
+ // its own share of "total". Signers that had no weight at the anchor contribute 0 to the floor.
|
|
+ let floor_table = match state.locks.iter().next_back() {
|
|
+ Some((&l, &h)) if l < index => self.voters_at(h, state),
|
|
+ _ => table.clone(),
|
|
+ };
|
|
@@
|
|
- let passes = self.lock_test(signed, active_num, p, table.total);
|
|
+ let signed_floor: u64 = signers.iter().map(|k| floor_table.weight(k)).sum();
|
|
+ let passes = self.lock_test_split(signed, active_num, p, signed_floor, floor_table.total);
|
|
```
|
|
|
|
with `lock_test_split` applying `3 x signed x P >= 2 x active_num` at C_i and `30 x signed_floor >= 17 x
|
|
floor_total` at the anchor. Cost: after a permanent loss of weight (3.3.1 D, 50% churn) the floor can no longer be
|
|
met by the survivors or by new miners until an operator moves the anchor, where today it clears in 4.1 days.
|
|
That is the decision for gate 3; the bound above is the fact either way.
|
|
|
|
## FAIL: S5, a burst locks alone while the window is young (ledger F1, spec 3.8 not implemented)
|
|
|
|
Measured: a miner with base share 1/6 pulsing 10x for 20 s of every 120 s. Over the run its weight share equalled
|
|
its block share (35.3% vs 35.3%, ratio 0.999), so W2 counts blocks and the retarget lag bought nothing: the
|
|
amplification half of F14 passes. But checkpoints 1 to 10 were locked by the burster alone: its first burst gave
|
|
one key 66.7% to 71.7% of a window that held under 300 blocks (cp 5: 98 of 147 signed by 1 of 6 voters; cp 10:
|
|
201 of 297), above both Q3 tests. From cp 11 every lock needed 3 or 4 signers as the share decayed.
|
|
|
|
This is ledger F1 measured on a real node: with no first-month gate (`FinalityParams::DEVNET.min_daa` 0,
|
|
`MAINNET.min_daa` 3,600 DAA, one hour against a 30-day window; spec 3.8's rule is "Proposed, Open O-3.1" and 3.10
|
|
says it is not implemented) a short burst owns a young window and certifies alone.
|
|
|
|
Proposed fix (spec 3.8's recommended rule, one line per network):
|
|
|
|
```diff
|
|
--- a/consensus/core/src/finality.rs
|
|
+++ b/consensus/core/src/finality.rs
|
|
@@ impl FinalityParams {
|
|
pub const MAINNET: FinalityParams = FinalityParams {
|
|
...
|
|
- min_daa: 3_600,
|
|
+ // C5 / 3.8: no certificate until the window holds a full window of history (ledger F1)
|
|
+ min_daa: 2_592_000,
|
|
};
|
|
pub const DEVNET: FinalityParams = FinalityParams {
|
|
...
|
|
- min_daa: 0,
|
|
+ min_daa: 7_200,
|
|
};
|
|
```
|
|
|
|
`evaluate` already gates certificate formation on `cp.daa_score >= self.params.min_daa`, so no other code moves.
|
|
The litepaper then states the first month runs plain GHOSTDAG under the 12-hour finality depth (F15).
|
|
|
|
## Summary of the run (2026-10-04, SCALE 0.6, six voters)
|
|
|
|
| # | Scenario | Verdict |
|
|
|---|---|---|
|
|
| s3 | dishonest aggregators | PASS (wire injection not run) |
|
|
| s2 | Sybil dust | weights PASS, sortition FAIL (F17) |
|
|
| s1 | equivocation at scale | PASS |
|
|
| s6A | partition 3/3 | FAIL (floor time-bounded, T* = 2F/13R) |
|
|
| s6B | partition 4/2 | PASS |
|
|
| s4 | vote-dropping producer | PASS, 0 ms added |
|
|
| s8 | malformed votes over RPC | PASS (RPC half) |
|
|
| s5 | pulsed miner | amplification PASS, lock-alone FAIL (F1) |
|
|
| s7 | eclipse | not run (needs the p2p probe) |
|
|
|
|
Full numbers: `docs/bench-log.md`, entry "finality v2 attack harness" of 2026-10-04.
|