diff --git a/docs/analysis/m16-recompute-attacker-2026-10-05.md b/docs/analysis/m16-recompute-attacker-2026-10-05.md new file mode 100644 index 000000000..a6329d0f6 --- /dev/null +++ b/docs/analysis/m16-recompute-attacker-2026-10-05.md @@ -0,0 +1,115 @@ +# M16: the recompute attacker with the 256 MiB cache on a die, a cost model + +5 October 2026 (evening), FUD ledger sweep round 6. Ledger M16, review R3.5, the chip designer's attack 3. + +The claim: "put 256 MiB of SRAM on a die and the dataset is never needed: 128 items per hash at about 1,170 integer +operations and 8 near-free reads each; integer operations per dollar is where silicon beats a GPU." + +This file prices that device from the rules in the specification and the rates measured so far. Nothing here is a +measurement of a chip. Every figure says where it comes from; "approximate" marks a figure from memory. + +## 1. What the honest miner pays per hash + +| Quantity | Value | Source | +|---|---|---| +| Loads per hash | 128 (16 load slots x 8 iterations), every accepted program | `igneum-pow/src/generator.rs` (`LOAD_SLOTS`, `ITERATIONS`), spec 01 section 1.4.2 | +| Distinct addresses per hash | 120.05 to 128, median 128.00, over 20,000 accepted programs | 20,000-program census, `docs/bench-log.md` "generator version 2"; `docs/analysis/weak-program-census-2026-10-03.md` | +| Bytes per load | 4 (one dataset word) | spec 01 section 1.8.5 | +| Dataset | 1 GiB (2^28 words), built once a day from the 256 MiB cache | spec 01 section 1.8.5; CLAUDE.md | +| Honest rate, RTX 5090, 1 GiB dataset | 228.95 Mhash/s, 23.8 G random loads/s, 95.2 GB/s useful | `docs/bench-log.md`, "3 October 2026, RTX 5090, memory-hard dataset" | +| Honest rate, RTX 5090, 64 MiB dataset (fits the 96 MiB L2) | 1,352 Mhash/s, 5.8x the 1 GiB rate | `docs/bench-log.md`, RTX 5090 dataset sweep (cited in ledger M1) | +| Dataset build, RTX 5090 | 13.4 ms for 1 GiB (1,253 M items/s); cache fill 0.67 ms | the same entry | +| Projected honest rate for version 2 programs, RTX 5090 | 141 Mhash/s (approximate: 18.0 G distinct loads/s at 128 distinct loads per hash; not yet run) | `docs/bench-log.md`, weak-program census, "Hash-rate spread" | + +The honest hash is 128 dependent random 4-byte reads over a buffer larger than any on-chip cache. The ALU work of +the program (64 instructions x 8 iterations per lane) is not what bounds it: the 5.8x step between the 64 MiB and +1 GiB datasets on the same card is the memory system, not the arithmetic. + +## 2. What the recompute attacker pays per hash + +The attacker holds the 256 MiB cache (on a die, the premise) and derives each dataset word on demand instead of +reading it. + +| Quantity | Value | Source | +|---|---|---| +| Items per hash | 128 (one dataset item of 16 words per load; two loads in one item share it, so at most 128 and in the census's accepted population about 128) | spec 01 section 1.8.5, `dataset[w] = item(w >> 4)[w AND 15]` | +| Mixer applications per item | 9 (`ITEM_ROUNDS = 8` dependent cache reads, nine mixer applications) | spec 01 sections 1.8.4 and 1.8.5 | +| Operations per mixer application | about 130 integer operations (16 xor-add-multiply steps and 8 ChaCha quarter rounds on a 16-word state) | spec 01 section 1.8.4 | +| Operations per item | about 1,170 | 9 x 130 | +| Cache reads per item | 8 dependent 64-byte lines (each address depends on every earlier read) | spec 01 section 1.8.5 | +| Operations per hash | about 150,000 (128 x 1,170) | arithmetic | +| Cache bytes per hash | 65,536 (128 x 8 x 64) in 1,024 dependent reads | arithmetic | + +Measured on Apple silicon (the only inline kernel run so far): the inline kernel does 9.48 Mhash/s on the M5 Max at +both a 256 MiB and a 1 GiB dataset, against 94.8 honest at 256 MiB and 45.2 at 1 GiB (`proto-metal/MEMHARD.md` +section 2.2: 10x and 4.8x slower). The same inline kernel under heavy load on 3 October ran 17x slower than honest at +a 256 MiB dataset (`docs/bench-log.md`, "R3.26 / M15", M16 note). The Mac's 256 MiB cache sits in DRAM, so its +inline kernel is bound by the 1,024 dependent cache-line reads per hash and does not price a die; it only shows +the kernel exists and is bit-exact. + +## 3. The die, priced in the GPU's own units + +To match ONE RTX 5090 at its honest 1 GiB rate the attacker's chip must deliver, per second: + +| Need | Value | Arithmetic | +|---|---|---| +| Integer operations | 34 T op/s | 228.95 M hash/s x 150,000 op/hash | +| Cache-line reads | 234 G reads/s, 15 TB/s of SRAM bandwidth | 228.95 M x 1,024 reads x 64 B | +| SRAM | 256 MiB, with the next day's cache under construction beside it | spec 01 (the cache is per day) | + +What the GPU itself has (approximate, from memory, for scale): the RTX 5090's integer throughput is about 50 T +op/s (21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock), the figure the ledger entry already +carries; on-die SRAM at 256 MiB costs about 100 to 300 mm^2 on a current node (the low end from a 0.02 um^2 bit +cell with array overhead, the high end from wafer-scale parts at about 1 MB per mm^2), against a 750 mm^2 +class GPU die; 15 TB/s of on-die SRAM bandwidth is within what wafer-scale parts quote and is not the bound. + +So, at equal silicon and equal integer throughput, the recompute attacker reaches 50 T / 150,000 = about 0.33 +Ghash/s: + +| Against | Honest 5090 rate | Attacker gain at equal integer budget | +|---|---|---| +| Closed-form and version 1 programs as measured | 229 Mhash/s | 1.5x | +| Version 2 programs as projected (distinct-load bound) | 141 Mhash/s | 2.4x (the figure in the ledger entry) | + +Before any chip-versus-GPU efficiency factor. A fixed-function pipeline with no instruction scheduling and no +warp divergence is usually credited with 2x to 5x over a GPU on integer work (approximate, from memory; the +honest target in the ledger is "under 2x"). Taking 3x: 4.5x to 7x over a 5090 at equal die area, minus the area +the SRAM takes (13% to 40% of the die), so about 3x to 6x. That is the exposure as the parameters stand, and it is +arithmetic, not a measurement. + +## 4. The lever, and why it costs the honest miner nothing + +The attacker's cost is linear in operations per item. The honest miner pays the mixer once per day in the +dataset build (2^26 items x 1,170 op = 78 G op, 13.4 ms on the 5090 as measured) and never per hash. The +verifier pays it per load it checks (spec 01 section 1.11: the CPU verifier derives the distinct items of a warp +from the cache, 0.41 to 1.2 ms per warp with the cache on one M5 Max core, `docs/bench-log.md` 3 October). + +| Mixer cost multiplier m | Operations per hash | Attacker rate at 50 T op/s | Gain against 229 Mhash/s (equal silicon, no efficiency factor) | Gain with a 3x fixed-function factor | Honest daily dataset build, 5090 | CPU verify per warp (scaled from 0.41 to 1.2 ms) | +|---|---|---|---|---|---|---| +| 1 (today, 1,170 op/item) | 150,000 | 0.33 Ghash/s | 1.5x | 4.4x | 13.4 ms | 0.4 to 1.2 ms | +| 2 | 300,000 | 0.17 Ghash/s | 0.73x | 2.2x | 27 ms | 0.8 to 2.4 ms | +| 4 | 600,000 | 0.083 Ghash/s | 0.36x | 1.1x | 54 ms | 1.6 to 4.8 ms | +| 8 | 1,200,000 | 0.042 Ghash/s | 0.18x | 0.55x | 107 ms | 3.3 to 9.6 ms | +| 16 | 2,400,000 | 0.021 Ghash/s | 0.09x | 0.27x | 214 ms | 6.6 to 19 ms | + +Reading. The mixer cost is the one parameter that moves the recompute attacker and leaves the honest hash rate +untouched; it is bounded by the 10 ms CPU verification gate (ledger M9, spec 1.16), which at today's verify time +allows about 8x before the slow core of the gate (a 2019-class laptop core, unmeasured, O-1.14) is at risk. The +cache size is the other lever and is linear in SRAM area, which is the cheaper side for the attacker: 256 MiB to +1 GiB moves the die from about 100 to 300 mm^2 to 400 to 1,200 mm^2, which is a multi-die part. Both are +prototype values of spec 1.16, fixed at gate 1. + +## 5. What this does not settle + +1. The inline kernel on NVIDIA at a 64 MiB cache inside the 5090's 96 MiB L2 (the on-die SRAM emulation the + ledger entry names) has not run; it is a PC job. It would put a measured point under the "50 T op/s" row: the + rate a real integer engine reaches on the actual mixer with the cache in SRAM-class memory. +2. The time-memory curve (O-1.6: store a fraction f of the dataset, recompute the rest) is not drawn; the model + above is the f = 0 endpoint. A partial-store attacker with HBM instead of SRAM is a different device and may be + the cheaper one. +3. The mixer has had no cryptanalysis (spec 1.8.4, `MEMHARD.md` section 3); a shortcut inside the mixer would cut + the 1,170 directly. +4. Monero's seven years do not price this device (ledger C13). + +Decision at gate 1 (owner: the project lead): the cache size rule "exceeds what one die can hold, and grows", and the mixer +cost multiplier, against the CPU verify gate. diff --git a/docs/fud-ledger.md b/docs/fud-ledger.md index 11e99a90a..657d6c946 100644 --- a/docs/fud-ledger.md +++ b/docs/fud-ledger.md @@ -1109,7 +1109,7 @@ Evidence: `docs/analysis/weak-program-census-2026-10-03.md` sections 7 and 9. Re ### M20. Pruning proofs are checked with the kHeavyHash stub "Wait thirty hours, start a fresh node, and watch it reject the honest pruning proof: `validate.rs:192` runs kHeavyHash on headers mined under the lottery, which pass with probability 2^-28. And if you loosen that, I forge levels with an ASIC that already exists." -Status: Open, acknowledged, fix named. Sweep (5 October 2026): still the stub on `devnet-v4` (`consensus/src/processes/pruning_proof/validate.rs:192` calls `calc_block_level_check_pow`; `apply.rs:74, 200` and `mod.rs:207` call `calc_block_level`). The test, a fresh node syncing a chain past its pruning depth, needs a network older than that depth; a fresh node against the live devnet is outside this sweep's rules, so it was not run. When it bites: the devnet's pruning depth is `PRUNING_DURATION` 108,000 DAA (`consensus/core/src/config/constants.rs:94`; the derived lower bound is 63,398), and the pruning point first leaves genesis once a finality point sits a full pruning depth below the tip, DAA 108,000 to 151,200, which at 1.05 DAA/s from DAA 33,000 at 17:37 UTC on 4 October falls between about 14:00 UTC on 5 October and 01:00 UTC on 6 October (approximate). From then on a fresh node receives a pruning proof and the stub rejects the honest headers with probability about 1 - 2^-28 each. Fix row in `docs/fud-fixes.md` section 2.5. +Status: Fixed in the node (rolled out 5 October 2026, 0.3.5: `m20-pruning` d35b00cf merged into fork 20139145 as its last merge, `cargo test -p kaspa-consensus --features igneum-pow -- pruning_proof` 4 passed at 03:15 UTC, `docs/plans/release-0.3.5.md` 1b and 3b: pruning proofs are checked with the Igneum lottery hash and the chain seeds; `pruning_proof/validate.rs`, the IBD proof flow and the p2p proof messages carry the seeds). Sweep (5 October 2026, evening): the live test is still owed and now possible: the Mac node's pruning point is still genesis at DAA 113,289 (`getBlockDagInfo` at 16:00 UTC, pruning point `edc4fa84...` with DAA score 0), so no node has yet served or checked a lottery-hashed pruning proof on the live devnet; the first fresh node to sync after the pruning point moves (expected between 14:00 UTC on 5 October and 01:00 UTC on 6 October by the entry's own arithmetic, approximate) is the measurement, and a fresh `igneumd` on this Mac against the live seed is read-only for the network and should be run then. Was: Open, acknowledged, fix named. Sweep (5 October 2026): still the stub on `devnet-v4` (`consensus/src/processes/pruning_proof/validate.rs:192` calls `calc_block_level_check_pow`; `apply.rs:74, 200` and `mod.rs:207` call `calc_block_level`). The test, a fresh node syncing a chain past its pruning depth, needs a network older than that depth; a fresh node against the live devnet is outside this sweep's rules, so it was not run. When it bites: the devnet's pruning depth is `PRUNING_DURATION` 108,000 DAA (`consensus/core/src/config/constants.rs:94`; the derived lower bound is 63,398), and the pruning point first leaves genesis once a finality point sits a full pruning depth below the tip, DAA 108,000 to 151,200, which at 1.05 DAA/s from DAA 33,000 at 17:37 UTC on 4 October falls between about 14:00 UTC on 5 October and 01:00 UTC on 6 October (approximate). From then on a fresh node receives a pruning proof and the stub rejects the honest headers with probability about 1 - 2^-28 each. Fix row in `docs/fud-fixes.md` section 2.5. Answer: Correct. `consensus/src/processes/pruning_proof/validate.rs:192` calls `calc_block_level_check_pow`, which runs the stub, and `apply.rs` and `mod.rs` call `calc_block_level` the same way; `docs/fork-divergence.md` records that seeds must be threaded through pruning-proof validation before a pruning network. The devnet will pass its pruning depth (108,000 blocks, sooner after tonight's overshoot) and a fresh node will show it. Fix: derive the epoch and day for proof headers from the proof's own headers (fork map a4, O-2.5) and remove the stub from the proof path. @@ -1333,7 +1333,9 @@ Evidence: spec 4.3 (O-4.3), 3.3.1, 3.5, 3.7 item 2, 3.9. Experiment: O-3.16. Rev ### F22. Certificates carry 8 to 10 of 12 votes on a healthy network, so locks sit a hair above the floor "Two of your cloud checkpoints locked at 66.8% against a 66.7% floor with every node connected. One slow aggregator and finality pauses." -Status: Fix built, pending rollout (4 October 2026, evening; branch `finality-fixes` of the node, behind `finality_v3_activation_daa`, `docs/plans/finality-v3-rollout-devnet.md`). Was: Open, found by the team (4 October 2026, 12-node cloud devnet, second partition run). Sweep (5 October 2026): nothing runnable without the devnet rollout; rule v3 is measured on the fast-time 3-node network (held certificates 6 of 6 at 11 of 11 indices, 0 conflicts, lock latency unchanged; bench-log "finality rule v3"). The rollout is an operator step. +Status: Fixed in the node and shipped, rule not yet activated on the live devnet (5 October 2026, evening sweep). Was: Fix built, pending rollout (4 October 2026, evening; branch `finality-fixes` of the node, behind `finality_v3_activation_daa`, `docs/plans/finality-v3-rollout-devnet.md`). Was: Open, found by the team (4 October 2026, 12-node cloud devnet, second partition run). + +Sweep (5 October 2026, evening): the code is on every node since 0.3.4 (the 0.3.4 fork was `finality-fixes` 6aa69a45, which carries the fold and the frozen table; 0.3.5 to 0.3.8 keep it, `docs/plans/release-0.3.5.md` 1b) but the switch is not thrown: the live override file read by the 0.3.6 digest check is `{"difficulty_v2_activation_daa": 33000, "proving_v0_activation_daa": 84100}` (`docs/plans/release-0.3.6.md` 8g) and the default is never on every network (`consensus/core/src/config/params.rs`, `finality_v3_activation_daa: u64::MAX` on devnet), so the live devnet runs rule v2 and today's locks still carry 67.0% to 71.9% of active weight (node logs, 5 October 2026: PC 2 `checkpoint 2612 LOCKED ... 67.0% of active`, PC 1 `3080 LOCKED ... 71.9%`). Rollout is the operator step N3 of the plan (publish the switch in the manifest override and every hand node's file at once, as the proving activation did). Decision owner: the project lead (N3). Sweep (5 October 2026): nothing runnable without the devnet rollout; rule v3 is measured on the fast-time 3-node network (held certificates 6 of 6 at 11 of 11 indices, 0 conflicts, lock latency unchanged; bench-log "finality rule v3"). The rollout is an operator step. Answer: True as measured, and the cause is not a cut-off at all: the node builds the certificate the instant the votes it holds meet Q3, and carries that one. Measured on the cloud logs of the healthy stretch 11:45 to 14:00 UTC (212 indices, `infra/cloud-devnet/results/2026-10-04/f22-vote-timing.md`, script `tools/finality-attacks/vote-timing.py`): the first certificate was built median 1.24 s (p99 1.71 s) after the first node determined the checkpoint, with 7 to 10 of 12 signers (mean 8.27); by then 10.24 votes had been issued on average, so about two were in flight (the miner's 1-s poll, a 250-ms gossip pump per hop, inter-region RTT up to 289 ms) and about two were issued later; the last of the 12 votes was issued median 1.45 s, p90 2.36 s after the first determination. Holding for 1 s after the first build would have carried all 12 votes at 192 of 212 indices; the other 20 are one event, miners 01, 06 and 11 down together for 10 minutes across indices 377 to 396 (the afternoon `hop.sh` restarts), not relay lag. The fix (spec Q4, rule v3): the first certificate still forms at quorum, so lock latency is unchanged; once every voter has signed, or `certificate_fold` DAA seconds after the determination (3 on devnet, 6 on mainnet), a node rebuilds the certificate from every vote it has seen and gossips the heavier one, and every node replaces a held certificate with a verified heavier one over the same block. Presence needs nothing: under the block reading of Q2 a late vote already counts once any block carries it. Unit test `fold_round_carries_late_votes_and_heavier_certificates_replace` (node, `processes::finality`). Network figures: `docs/bench-log.md`, "finality rule v3". @@ -1516,7 +1518,7 @@ the project lead's decision of 4 October 2026: the total-weight floor of Q3 is 2 - **F16** (a lock can become uncertified after a heal). Status: rule unchanged (spec 3.11.4: a verified certificate is never withdrawn, O-3.17 implements the conflict report). What the floor changes is how the state F16 describes arises: two certificates at one index now need equivocators holding at least one third of total weight in every scenario (two certificates need 4/3 of weight in signatures), not 13.3% across a partition that outlasts the presence decay (`sim/results_v2.md` H at 2/3: 0 conflicts and no lock on either side through a 33% equivocator, conflicts from minute 0 at 34%). Ten days of 100% hashrate in public, or the long-partition case of F21. - **F18** ("a silent minority cannot freeze finality" is false under the floor). Status: Fixed again (4 October 2026). The litepaper now says a lock needs two thirds of all 30-day weight and that finality pauses whenever less than two thirds is connected and signing. The pause threshold moved from about 42% of weight silent to one third: in the model, whose keys are in outage 2.2% of the time, 30% silent locks every checkpoint, 32% locks 88%, 33% locks 11% and 34% locks none for as long as it stays silent (`sim/results_v2.md` L1). Was: Fixed (56.7% sentence). - **F9** (half the hashrate leaves and finality stalls for ten days). Status: Conceded, stated (4 October 2026). Under the 2/3 floor the critic's number is back: 50% churn pauses finality for 10 days and 35% churn for 1.4 days, until the departed weight ages out of the window (`sim/results_v2.md` D's total column, which the 2/3 floor equals arithmetically, and L2). The chain runs on proof of work meanwhile, the node reports the pause, and the litepaper says so. This is the price of the one-third safety bound and was taken knowingly. Was: Answered by design (the active denominator recovered in two hours). -- **F21** (new, from attack scenario 6A). "Your floor is a fraction of a table each side computes for itself. Cut the network in half and leave it cut: after a while each half's window is full of its own blocks, each half holds two thirds of its own table, and both lock without any attacker at all. Your simulation never saw it because it kept the weights global." Status: Fix built, pending rollout (4 October 2026, evening): rule v3, spec 3.3 Q5, the frozen weight table, on the node's `finality-fixes` branch behind `finality_v3_activation_daa` (`docs/plans/finality-v3-rollout-devnet.md`). The rule: a certificate also needs its signers to hold two thirds of the weight table at the last certified checkpoint on C_i's chain, at that table's weights, while that checkpoint is less than one window old. Both sides of a partition share that table and neither can fill it, so no side under two thirds locks until 30 days have passed without a certified checkpoint; at the heal locking resumes on one chain. Proved first in `sim/finality_v2.py` scenario M (`sim/results_v2.md`, "Rule v3"): 50/50, 60/40 and 55/45 splits never lock in 12 days (v2: days 10.2, 5.2, 7.9), both sides of a 31-day split lock alone at day 30.00 when the frozen table expires, the 70/30 majority locks at once under both rules, every pre-heal lock is kept and the first lock after the heal comes 0 minutes in, the 34% equivocator still conflicts (the one-third bound of 3.11.2 is untouched). The price, stated in 3.7 item 2: a set of one third or more that stops mining and signing at once pauses finality for 30 days (v2: 1.7 days at 35%, 10.1 at 50%); a gradual departure costs nothing because every certified checkpoint re-freezes the table. A view still cannot count blocks it has never seen, so a partition longer than a window forks as before; the fix moves the bound from a third of the window to the whole of it. Node: `processes::finality` (`frozen_table`, the Q5 test in `evaluate`), unit test `frozen_table_holds_a_side_without_the_other_keys_for_one_window`; network figures in `docs/bench-log.md`, "finality rule v3". Was: Conceded, stated (4 October 2026): spec 3.3.1, 3.7 item 9, 3.9 guidance; `sim/results_v2.md` L4 (view-local weights). Correct. A side with pre-split share s holds s + (1 - s) t / 30 of its own table on day t and two thirds of it from day 30 (2/3 - s) / (1 - s): day 10 at 50/50 (day 4 under the old floor), day 5 for the 60 side of 60/40 (at once under the old floor). On the devnet the old floor fell at 84 s of a young 1,439-DAA window (`docs/bench-log.md`, "finality v2 attack harness", S6A); at 2/3 the same cut on a full 1,800-DAA window held for the whole 150-s split and fell at 205 s against a predicted W / (3R) = 200 s (`docs/bench-log.md`, "finality floor 2/3", 6A and 6A long heal). What the devnet adds to the simulation: after the heal the other side's blocks are merged red, so each side's own share jumps rather than drifts, both sides of a 50/50 split certify their own checkpoints within 10 s of each other, and F1 then pins each node to its own certified chain: 26 conflicting certificates and 23 disagreeing locked indices across three nodes, no equivocation, a finality fork that the network heal did not undo and that only an operator's trusted certificate (F5, not implemented) can resolve. No rule removes it, because a view cannot count blocks it has never seen; the floor at two thirds moved the day from 4 to 10, and the exchange guidance treats a node partitioned for more than a day as proof of work until it has rejoined. A rule option for gate 3, not adopted: evaluate the floor against the table of the last locked checkpoint while no newer lock exists, which trades F9's 10-day recovery for a manual override. +- **F21** (new, from attack scenario 6A). "Your floor is a fraction of a table each side computes for itself. Cut the network in half and leave it cut: after a while each half's window is full of its own blocks, each half holds two thirds of its own table, and both lock without any attacker at all. Your simulation never saw it because it kept the weights global." Status: Fixed in the node and shipped, rule not yet activated on the live devnet (5 October 2026, evening sweep: the code has shipped in every node since 0.3.4 and the switch `finality_v3_activation_daa` is absent from the live override file, default never; the F22 entry above carries the evidence; the overlay-against-GHOSTDAG run of C4 exercised rule v3 on the live node line on the fast-time harness tonight; decision owner for N3: the project lead). Was: Fix built, pending rollout (4 October 2026, evening): rule v3, spec 3.3 Q5, the frozen weight table, on the node's `finality-fixes` branch behind `finality_v3_activation_daa` (`docs/plans/finality-v3-rollout-devnet.md`). The rule: a certificate also needs its signers to hold two thirds of the weight table at the last certified checkpoint on C_i's chain, at that table's weights, while that checkpoint is less than one window old. Both sides of a partition share that table and neither can fill it, so no side under two thirds locks until 30 days have passed without a certified checkpoint; at the heal locking resumes on one chain. Proved first in `sim/finality_v2.py` scenario M (`sim/results_v2.md`, "Rule v3"): 50/50, 60/40 and 55/45 splits never lock in 12 days (v2: days 10.2, 5.2, 7.9), both sides of a 31-day split lock alone at day 30.00 when the frozen table expires, the 70/30 majority locks at once under both rules, every pre-heal lock is kept and the first lock after the heal comes 0 minutes in, the 34% equivocator still conflicts (the one-third bound of 3.11.2 is untouched). The price, stated in 3.7 item 2: a set of one third or more that stops mining and signing at once pauses finality for 30 days (v2: 1.7 days at 35%, 10.1 at 50%); a gradual departure costs nothing because every certified checkpoint re-freezes the table. A view still cannot count blocks it has never seen, so a partition longer than a window forks as before; the fix moves the bound from a third of the window to the whole of it. Node: `processes::finality` (`frozen_table`, the Q5 test in `evaluate`), unit test `frozen_table_holds_a_side_without_the_other_keys_for_one_window`; network figures in `docs/bench-log.md`, "finality rule v3". Was: Conceded, stated (4 October 2026): spec 3.3.1, 3.7 item 9, 3.9 guidance; `sim/results_v2.md` L4 (view-local weights). Correct. A side with pre-split share s holds s + (1 - s) t / 30 of its own table on day t and two thirds of it from day 30 (2/3 - s) / (1 - s): day 10 at 50/50 (day 4 under the old floor), day 5 for the 60 side of 60/40 (at once under the old floor). On the devnet the old floor fell at 84 s of a young 1,439-DAA window (`docs/bench-log.md`, "finality v2 attack harness", S6A); at 2/3 the same cut on a full 1,800-DAA window held for the whole 150-s split and fell at 205 s against a predicted W / (3R) = 200 s (`docs/bench-log.md`, "finality floor 2/3", 6A and 6A long heal). What the devnet adds to the simulation: after the heal the other side's blocks are merged red, so each side's own share jumps rather than drifts, both sides of a 50/50 split certify their own checkpoints within 10 s of each other, and F1 then pins each node to its own certified chain: 26 conflicting certificates and 23 disagreeing locked indices across three nodes, no equivocation, a finality fork that the network heal did not undo and that only an operator's trusted certificate (F5, not implemented) can resolve. No rule removes it, because a view cannot count blocks it has never seen; the floor at two thirds moved the day from 4 to 10, and the exchange guidance treats a node partitioned for more than a day as proof of work until it has rejoined. A rule option for gate 3, not adopted: evaluate the floor against the table of the last locked checkpoint while no newer lock exists, which trades F9's 10-day recovery for a manual override. ### M24. Your two-lane controller oscillates for an hour when a second miner joins mid-epoch "Watched your devnet this morning. The second 5090 came in at 10:12 and the difficulty never settled: 102M to 164M for forty minutes, 54 to 81 blocks a minute, three or four clamp steps stacked inside a second. Your fast lane and your slow lane disagree by a hair under the trigger and the rule flips between them every two minutes. One card joining is the mildest event a chain can see." @@ -1649,7 +1651,9 @@ Evidence: the files above. ### G12. The PoW schedule comes from the environment on every network, including mainnet "Your mainnet gate refuses the override file. It does not refuse `IGNEUM_POW_EPOCH_BLOCKS`. A node without a file installs the schedule from the environment and `Params.pow_epoch_blocks` is never consulted." -Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, local worktree `vendor/igneum-node-fud`, no remote; main repo branch `fud-consensus`). Was: Open (4 October 2026). +Status: Fixed (rolled out 5 October 2026, 0.3.5: fork 20139145 = `finality-fixes` 6aa69a45 + `fud-consensus` 977db931 + `m20-pruning` d35b00cf + the miner branches, cut 07:33 BST, `docs/plans/release-0.3.5.md` 1b and 4b; node line 2b6d23ef from 0.3.6 at 09:34 UTC carries the same consensus code; every reachable app machine on the line by 09:32 UTC and the hand nodes and the seed restarted on it for the proving activation at DAA 84,100 the same morning, `docs/plans/release-0.3.6.md` 8j, bench-log "live devnet: the first shards proven"). Was: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, local worktree `vendor/igneum-node-fud`, no remote; main repo branch `fud-consensus`). Was: Open (4 October 2026). + +Sweep (5 October 2026, evening): measurement line: every node on the line prints `Consensus params digest: f10a4eab...` at start (node logs of PC 1, PC 2, the Mac, Sam's Mac and the US laptop, 5 October 2026, the digest with difficulty 33,000 and proving 84,100), and the environment variables have no reader left in the miner (`IGNEUM_POW_DAY_MS` moved into the template, M25 confirmed the day-mismatch rejection on 5 October). No mainnet node has been started; the mainnet refusal line is covered by the unit test `env_pow_schedule_is_devnet_and_simnet_only` only. The fix: `PowSchedule::from_env()` is gone. `PowSchedule::from_env_for(network, base)` applies the three variables on devnet and simnet only and returns `None` elsewhere; the daemon calls `Params::apply_env_pow_schedule()` after the override file, prints "PoW schedule from the environment (...)" when applied and "Ignoring IGNEUM_POW_... on igneum-mainnet: the environment never sets a consensus parameter outside devnet and simnet" when not, then installs the network's schedule from `Params` on every network (`install_pow_schedule`), so `Params.pow_epoch_blocks` is what runs; the lazy fallback in `pow_schedule()` installs the devnet constants, never the environment. The miner takes all three schedule values from the template (`pow_epoch.day_ms` joined the two epoch fields in `PowEpochInfo`, the RPC model and the gRPC proto), so `IGNEUM_POW_DAY_MS` has no reader left in the miner (R4.1.9's day split is closed with it). The effective schedule is part of the params digest of X18, so a devnet node with the variable set cannot connect to one without it. Unit test `env_pow_schedule_is_devnet_and_simnet_only` (consensus-core, `config::params::tests`): the variable moves the devnet and simnet schedule and digest, leaves mainnet's and testnet's untouched, and the caller can tell "ignored" from "nothing set". Spec 2.8 and 8.7 state the rule. @@ -1678,7 +1682,9 @@ Evidence: `git ls-files | xargs grep -lF ` counts, `git log -S`. Experime ### X18. Two nodes with two override files connect, and only some mismatches fork "Your handshake compares the network name and nothing else. A PoW or difficulty mismatch forks and bans; a `finality` mismatch is a WARN; `rollout-v2.sh` throws the finality block away when it writes the file; the app rewrites the packaged file on every start." -Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026). +Status: Fixed (rolled out 5 October 2026, 0.3.5: fork 20139145 = `finality-fixes` 6aa69a45 + `fud-consensus` 977db931 + `m20-pruning` d35b00cf + the miner branches, cut 07:33 BST, `docs/plans/release-0.3.5.md` 1b and 4b; node line 2b6d23ef from 0.3.6 at 09:34 UTC carries the same consensus code; every reachable app machine on the line by 09:32 UTC and the hand nodes and the seed restarted on it for the proving activation at DAA 84,100 the same morning, `docs/plans/release-0.3.6.md` 8j, bench-log "live devnet: the first shards proven"). Was: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026). + +Sweep (5 October 2026, evening): measurement line, live: the digest refused a real mismatch the morning it shipped. A hand node restarted early with another `proving_v0_activation_daa` was refused by every peer for 20 minutes (bench-log "live devnet: the first shards proven"); the node logs carry 115 `consensus params digest mismatch` lines on PC 1, 46 on PC 2 and 56 on the Mac between 08:15 and 08:35 BST on 5 October 2026 (`Refusing peer ...: consensus params digest mismatch, local f10a4eab... remote a6da35e8...`), none since. Still open from this entry: the digest-less allowance on devnet and simnet (to remove), `rollout-v2.sh` still writes the file as two fields (R4.1.12). The fix: `Params::consensus_digest()` (BLAKE2b-256, domain `IgneumParamsDigest`, every consensus field in a fixed tagged order: genesis, difficulty, mass and lane limits, blockrate, crescendo, the ten finality fields, the PoW schedule, the three activation heights; not the dead `timestamp_deviation_tolerance`, not seeders or ports; spec 2.8 lists it). The version message carries it (`paramsDigest`, field 11); `initialize_connection` refuses a peer whose digest differs with `ProtocolError::ParamsDigestMismatch` and one WARN naming both digests before any flow is registered, so a finality-only mismatch, which used to connect and WARN "names N voters" after the fact, never connects. A peer with no digest (an older build) is refused on mainnet and testnet and let in with a WARN on devnet and simnet while the devnet rolls (an allowance to remove afterwards). The node prints its digest at start. Unit test `consensus_digest_covers_every_consensus_field_and_nothing_else`. Measured (`docs/bench-log.md`, "round-4 consensus items", digest run; `tools/finality-attacks/fud.mjs digest`, fast time, ports 29400+): a listener on the shared fast-time override and a dialler whose finality block differs by one DAA second of window: the dialler was refused at the handshake on both sides ("Refusing peer ...: consensus params digest mismatch, local 4bf7... remote 7a40..." on the listener, the reject message on the dialler), 0 peers after 25 s on both; a third node on the shared override connected in 1 s. Control on the finality-fixes build 6aa69a45: the mismatched dialler connected (1 peer, no line). `rollout-v2.sh` still rewrites the file as two fields and the app still rewrites the packaged file (R4.1.12): with the digest both now fail loudly at the handshake instead of forking; the merge fix for the script is not done tonight. @@ -1689,7 +1695,9 @@ Evidence: the files above. Experiment: two nodes on different files; the handsha ### F23. The equivocation ban is node-local, so honest nodes refuse each other's certificates "Evidence detected from an RPC vote stamps the sink's DAA; evidence carried in a block stamps the carrier's DAA. Two honest nodes hold different `until` for the same key, their voter lists differ by one at every checkpoint between the two expiries, and `voter_count` refuses the other's certificate for good." -Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026); reproduced by the red team the same evening on the stock s1 scenario (n0 kept the ban until DAA 726 against 239 on n1 and n2, 9 and 3 to 4 certificates refused "names N voters"). +Status: Fixed (rolled out 5 October 2026, 0.3.5: fork 20139145 = `finality-fixes` 6aa69a45 + `fud-consensus` 977db931 + `m20-pruning` d35b00cf + the miner branches, cut 07:33 BST, `docs/plans/release-0.3.5.md` 1b and 4b; node line 2b6d23ef from 0.3.6 at 09:34 UTC carries the same consensus code; every reachable app machine on the line by 09:32 UTC and the hand nodes and the seed restarted on it for the proving activation at DAA 84,100 the same morning, `docs/plans/release-0.3.6.md` 8j, bench-log "live devnet: the first shards proven"). Was: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026); reproduced by the red team the same evening on the stock s1 scenario (n0 kept the ban until DAA 726 against 239 on n1 and n2, 9 and 3 to 4 certificates refused "names N voters"). + +Sweep (5 October 2026, evening): measurement line, live: the persisted finality state converted on the upgrade on every machine (`Finality: state blob of layout 1 read and converted (0 evidence records)` once on PC 1, PC 2, the Mac, Sam's Mac and the US laptop, 5 October 2026), no node lost its locks, and in the 10 hours to 15:45 UTC the node logs of the five app machines carry 0 certificates refused "names N voters", 0 CONFLICTING and 0 EQUIVOCATION lines against 5,243 LOCKED lines (intake query over every `nodelog-*` upload, split per line server-side). No equivocation has happened on the live devnet, so the ban path itself is exercised only by the fast-time run below (`fud.mjs ban`: voter lists agree on three nodes at every index, 0 refusals). The fix: the node-local `stripped` map is gone. Evidence is kept as `EvidenceRecord` (the two votes, the carriers with their DAA scores) and the ban at a checkpoint C is a function of C's past (`bans_at`): the key is stripped at C when some carrier lies in C's past and `daa(C) < daa(lowest carrier in C's past) + ban`. Evidence detected over RPC or gossip strips nothing until a block carries it; the node puts it in its next templates. The weight table cache stays ban-free and `voters_at` applies the checkpoint's own bans, so a certificate built before the carrier existed verifies on a node that saw the evidence later, and a node that saw it over RPC counts the same voters as one that saw it in the block. Evidence records are bounded (4,096; dropped once the ban ended two windows below the sink or never carried within one ban of being seen; 16 carriers per record), the same for vote and certificate carriers. The persisted state is layout 2; a layout-1 blob is read and converted on start, so no devnet node loses its locks on the upgrade. Unit test `ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list` (three `TestConsensus` nodes on one chain: the voter list agrees on all three at every checkpoint, the key is a voter before the carrier and after the ban and nowhere in between, the third node verifies the first two's certificates at every locked index). Measured (`docs/bench-log.md`, "round-4 consensus items", ban run; `fud.mjs ban`, 480 s, six voters, one equivocation at index 9 by a voter on n0 over RPC, n2 cut off 45 s around it and healed, so it saw the evidence late from the carrier block): on the `fud-consensus` build every node names 5 voters at the same indices (10 to 12 in the final pass, 10 to 13 in the first), 0 certificates refused "names N voters", 0 conflicting certificates, 0 locked indices disagreeing, voter counts agree at every index with lines on two or more nodes, locks continue to index 14 or 15 on all three. Control on the finality-fixes build: n0 (the RPC detector) refused 2 certificates "names N voters, this node counts M", the rest agreed because each node built its own; the red team's stock s1 scenario the same evening gave 9 / 3 / 4 refusals. Spec 3.6 and the 3.10 row state the rule. @@ -1702,7 +1710,9 @@ Red-team run, 4 October 2026 (evening, the 0.3.4 finality-fixes build with rule ### F24. A checkpoint determination is never revisited "After a reorg deeper than `checkpoint_depth`, the node's record for that index names a block off its chain. Every certificate the network forms for that index is refused as conflicting, with no equivocation anywhere, and the node voted for a block that is not on its chain." -Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026); extends F7 and C4. +Status: Fixed (rolled out 5 October 2026, 0.3.5: fork 20139145 = `finality-fixes` 6aa69a45 + `fud-consensus` 977db931 + `m20-pruning` d35b00cf + the miner branches, cut 07:33 BST, `docs/plans/release-0.3.5.md` 1b and 4b; node line 2b6d23ef from 0.3.6 at 09:34 UTC carries the same consensus code; every reachable app machine on the line by 09:32 UTC and the hand nodes and the seed restarted on it for the proving activation at DAA 84,100 the same morning, `docs/plans/release-0.3.6.md` 8j, bench-log "live devnet: the first shards proven"). Was: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026); extends F7 and C4. + +Sweep (5 October 2026, evening): measurement line, live: the re-determination fired on the real devnet and left no hole. Checkpoints 2970 and 2971 were re-determined at 10:56:44 to 10:56:45 BST on 5 October 2026 on PC 1, PC 2 and the Mac at once (`Finality: checkpoint 2971 re-determined: block bb4aa204...` on all three, 7 / 8 / 4 re-determined lines per machine over the day including the first-start ones at indices 704 to 722), with 0 CONFLICTING lines and 0 refusals anywhere in the 10 hours to 15:45 UTC, and the three nodes' later locks agree (the console's last lock #3646 at 15:46 UTC). Before this fix the same event would have logged CONFLICTING for every certificate at the moved index and left the index unlocked on the losing side. The fix: after every virtual change, every unlocked checkpoint record whose block is no longer a chain ancestor of the sink is determined again on the new chain ("re-determined" log line); the certificate held over the old block is dropped, the fold clock restarts. A certificate over a block other than the node's determination at an unlocked index (or at an index not yet determined, up to 64 ahead) is no longer logged CONFLICTING and discarded: it is held pending (4 per index) and verified when a re-determination names its block; a block that cannot be the index's checkpoint on any chain (C1 as a function of the DAG: blue score under the target, or the selected parent's not) is refused outright. CONFLICTING now means what 3.11.4 says: a certificate against a LOCK. The certificate verification itself is unchanged; a locked index is never re-determined (fork choice keeps the chain through it). The node's own keys do not re-vote at a re-determined index (that would be equivocation); the lock there comes from the network's certificate, which is the point. Unit tests `reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate` and `a_locked_checkpoint_pins_the_chain_and_a_certificate_against_it_conflicts`. Measured (`docs/bench-log.md`, "round-4 consensus items", reorg run; `fud.mjs reorg`, n0 with 30% of the weight cut off for 180 s while the 70% side kept locking, then healed): on the `fud-consensus` build n0 re-determined its own 1 to 2 split indices on the majority chain, verified the pending certificates at determination (2 in the final pass), logged 0 CONFLICTING and 0 refusals, and ended with every index the majority locked locked on the same block (0 disagreeing). Control on the finality-fixes build: n0 refused 3 certificates "is for X, this node's checkpoint is Y" and never locked indices 8 and 9 that the majority locked (the permanent hole of this entry), 0 disagreeing because the hole is not a lock. The final pass also found that a chain which becomes the sink at a lower blue score than the old one (more blue work per block after the split's difficulty drift) re-determined index 9 at a sink below its depth, to a block below the target: fixed the same night (an index the new chain has not reached is un-determined and determined again when it has; unit test `a_shallower_sink_un_determines_the_indices_it_cannot_reach`) and re-run (`reorg-final2`, fork 977db931): the majority locked nothing during the split this time (Poisson again), n0 re-determined its one split index, verified 1 pending certificate at determination, 0 CONFLICTING, 0 refused, 0 disagreeing, no record below its target, every index the majority locked locked on n0 too. The red team's own reproductions on the new build (`tools/finality-attacks/redteam/rtfin.mjs`, fin-attacks miner at 6 blocks/s): `f24c` (3/3 split, 16 s cut) PASS with 0 refusals, 0 CONFLICTING, 0 stuck indices, 0 disagreeing; `f24b` (4/2 split, 24 s cut) FAIL on both builds for a reason outside F24: at 6 blocks/s the DAA score advances 6 a second, so the 24-s cut is 144 DAA, longer than the 120-DAA weight window and the 60-DAA merge depth; the two chains cannot merge at the heal and each side's table holds only its own keys (n1's lone locks read "100.0% of total, 100.0% of the table frozen at lock 39"), which is the partition longer than a window that spec 3.7 item 9 states and F21 conceded, not a reorg. The scenario's "under merge depth" assumes 1 DAA a second. Spec 3.2 C1 and C4, and the 3.10 rows, state the rule. @@ -1715,7 +1725,7 @@ Red-team run, 5 October 2026, 00:56 (the 0.3.4 finality-fixes build, rule v3 on, ### F25. The fast-time harnesses cannot start a node, and the timestamp probe tests the old rule "Both attack harnesses rebuild each node's override with `JSON.parse` and `JSON.stringify` of `infra/fast-time/override-60x.json`. That file now carries two `u64::MAX` sentinels (`difficulty_v2_activation_daa`, `proving_v0_activation_daa`); a JavaScript number cannot hold them, the round-trip writes `18446744073709552000`, and `igneumd` refuses the file as a floating point where a u64 is expected. Every `--fast-time` run of `tools/finality-attacks` and `tools/harness` fails at the first node. Scenario 2 of `tools/harness` still probes the 132 s future bound and reports FAIL against the 10 s rule the node has carried since the timestamp fix." -Status: Fixed in both harnesses (4 October 2026, night: `tools/finality-attacks/lib/net.mjs` kept the sentinels as BigInt through the merge since the v3 runner of the evening; `tools/harness/lib/net.mjs` got the same reviver on branch `fud-memory`, merged into `fud-consensus`); the stale timestamp criterion of scenario 2 is not touched. Was: Open (4 October 2026, red-team run). Tooling, low severity: no consensus effect, but every fast-time attack run is blind until it is fixed. +Status: Fixed (rolled out 5 October 2026, 0.3.5: both harness libraries reached master with the `fud-consensus` merge 7abce72 of the 0.3.5 cut; the sweep of 5 October ran `run.mjs s4` and `s5 --fast-time`, `fud.mjs` and tonight's `c4.mjs` through them, every node started). Was: Fixed in both harnesses (4 October 2026, night: `tools/finality-attacks/lib/net.mjs` kept the sentinels as BigInt through the merge since the v3 runner of the evening; `tools/harness/lib/net.mjs` got the same reviver on branch `fud-memory`, merged into `fud-consensus`); the stale timestamp criterion of scenario 2 is not touched. Was: Open (4 October 2026, red-team run). Tooling, low severity: no consensus effect, but every fast-time attack run is blind until it is fixed. Answer: Correct, measured. The red-team run's first scenario errored on it (`docs/review/redteam-2026-10-04.md`, "Tooling defect"); `tools/proving-v0/run.mjs` already edits the file as text for this reason. Scenario 2 live probe: past floor pmt+1, future flip between +130.00 and +130.01 s of the probe's own offsets, every stamp from +10 s rejected; the node is right, the criterion is stale. Smallest fix: in `tools/finality-attacks/lib/net.mjs` and `tools/harness/lib/net.mjs` `overrideParams`, drop the two sentinel fields before `stringify` (absent means never) or splice the extra fields into the file text; in `tools/harness/scenarios/s2-timestamp.mjs`, probe `max(pmt + 1, parent - 10 s)` and the +10 s bound. Also stale: `tools/exec-attacks/scenario3_pgas.mjs` waits for an over-budget transaction to be included and skipped with `BlockProvingBudget`; since F-exec-B the mempool refuses it with the metered pgas, so the check should accept `ProvingGasAboveBlockLimit` from the pool (`docs/review/redteam-2026-10-04.md` row 28). And `tools/finality-attacks` scenario 5 compares the burster's share of the weight window with its share of the whole run, which only agree when the run is shorter than the window (row 17). @@ -1751,10 +1761,10 @@ Evidence: the files above. Experiment: `igneum-miner` with `IGNEUM_POW_DAY_MS=14 ### M26. The interval fault guard freezes its baseline and loops "On a trip you skip the STATUS print, so the baseline it would have updated stays frozen, and you roll the counters back to it. Any healthy rate over ten times a slow first interval trips again every interval, forever, with no STATUS line and no `faults=` for the app to read." -Status: Fixed (4 October 2026, evening), fork commits `aea5ac6d`, `501363e0` and `945153ab` on `miner-reliability` (`igneum/miner/src/guard.rs`; the second fixes a fill loop that spun forever after a guard kill, found by the test network). Replaced: Open (4 October 2026). +Status: Fixed (rolled out 5 October 2026, 0.3.5: `miner-reliability` 945153ab merged into fork 20139145, `docs/plans/release-0.3.5.md` 1b; the guard tests in the 0.3.5 `igneum-miner` suite, 12 passed). Was: Fixed (4 October 2026, evening), fork commits `aea5ac6d`, `501363e0` and `945153ab` on `miner-reliability` (`igneum/miner/src/guard.rs`; the second fixes a fill loop that spun forever after a guard kill, found by the test network). Replaced: Open (4 October 2026). Fix: `IntervalGuard` builds its baseline from the healthy intervals of the current worker process (a moving average, two intervals before it can trip) and forgets it when the worker restarts; the STATUS line is printed on a trip, with `faults=`; `RestartPolicy` restarts a guard-killed worker after 2 s, doubling per trip inside ten minutes, and the third trip ends the miner with exit 43 so the app shows a faulted card instead of a loop. The 60 s no-line timeout no longer skips the guards and the STATUS line. Measured against a fake worker on a private test network: `docs/bench-log.md`, "4 October 2026, miner fault guards and the app watchdog measured against a fake worker". -Status: Fix built (4 October 2026, branch `miner-reliability` of the node, `aea5ac6d`), pending merge and a run on a PC. `IntervalGuard`: the baseline is a moving average of the worker process's healthy intervals, forgotten on a restart; the STATUS line is printed on a trip with `faults=`. `RestartPolicy`: 2 s, doubling inside ten minutes, the third trip exits 43 so a supervisor can mark the card. Unit tests in `guard.rs` replay the slow-first-interval case and the back-off. The GPU stress run (`--status-secs 10` with a stress tool for 15 s) needs a PC (hardware). Was: Open (4 October 2026). +Sweep (5 October 2026, evening): measurement line, live, from the five app machines' miner logs over the 20 hours to 15:30 UTC: on the 0.3.5 to 0.3.8 miners every PC fault line is of one kind and rare (PC 1: 2 `WORKER FAULT` lines in the 07:00 hour and 4 in the 10:00 hour, PC 2: none after 03:00 UTC), the `no job completed ... killing the worker, it restarts` guard fired 43 times on PC 1 and 40 on PC 2 in the whole window, every time with a restart and a STATUS line afterwards, and the console shows 0 faults and 0 restarts on every card at 15:46 UTC. The loop this entry describes (a trip every interval for ever) does not appear on any machine since the 0.3.5 start. Earlier line, kept for the record: Fix built (4 October 2026, branch `miner-reliability` of the node, `aea5ac6d`), pending merge and a run on a PC. `IntervalGuard`: the baseline is a moving average of the worker process's healthy intervals, forgotten on a restart; the STATUS line is printed on a trip with `faults=`. `RestartPolicy`: 2 s, doubling inside ten minutes, the third trip exits 43 so a supervisor can mark the card. Unit tests in `guard.rs` replay the slow-first-interval case and the back-off. The GPU stress run (`--status-secs 10` with a stress tool for 15 s) needs a PC (hardware). Was: Open (4 October 2026). Answer: Correct. `igneum/miner/src/main.rs:1404-1418` with the update at `:1452-1454` skipped by `continue`; the restart has no cap and no growing back-off (`:1213-1216`). A slow first interval (a game on the GPU, a foreground self-heal build on a slow card) is enough. Fix: update the baseline on a trip, or compare to the previous interval; cap restarts with a growing back-off. Review id R4.2.1. @@ -1763,8 +1773,8 @@ Evidence: the file above. Experiment: `--status-secs 10` with a GPU stress tool ### M27. A flapping node makes the worker rebuild once per template "A prepare goes out whenever the wanted pair differs from the prepared one. No count, no interval, no once-per-epoch. Each one writes a pack on the CPU with the job loop stalled and costs the worker a full build; `prepare-failed` resends on the next fill." -Status: Fixed (4 October 2026, evening), fork commit `aea5ac6d` on `miner-reliability` (`guard::PrepareLimiter`): one `prepare` per pair per epoch, one retry after `prepare-failed`, none within 30 s of the last; a held prepare is printed once (`PREPARE held ...`). Measured with a worker that refused every prepare across three epochs: `docs/bench-log.md`, the same entry as M26. Replaced: Open (4 October 2026). -Status: Fix built (4 October 2026, branch `miner-reliability`, `aea5ac6d`), pending merge. `PrepareLimiter`: one prepare per pair per epoch, one retry after prepare-failed, none within 30 s of the last, a held prepare said once; a unit test replays a flapping node. The fast-time simnet with a flipping `next_epoch_seed` was not run (it needs a patched node and a GPU worker). Was: Open (4 October 2026). +Status: Fixed (rolled out 5 October 2026, 0.3.5, as M26). Was: Fixed (4 October 2026, evening), fork commit `aea5ac6d` on `miner-reliability` (`guard::PrepareLimiter`): one `prepare` per pair per epoch, one retry after `prepare-failed`, none within 30 s of the last; a held prepare is printed once (`PREPARE held ...`). Measured with a worker that refused every prepare across three epochs: `docs/bench-log.md`, the same entry as M26. Replaced: Open (4 October 2026). +Sweep (5 October 2026, evening): measurement line, live, the exact failure this entry predicts happened on the 0.3.4 miners and stopped with 0.3.5. Between 01:24 and 02:59 UTC on 5 October 2026 both PCs' NVIDIA workers refused the prepared pack for epoch 1130e9ea... (`prepare-failed ...: the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT`) and the 0.3.4 miner re-sent the prepare every 0.7 s: 4,299 `prepare-failed` lines on PC 1 and 4,233 on PC 2 in two hours, with 3,609 and 3,494 `WORKER FAULT seed mismatch` lines, two boundaries (DAA 61,200 and 64,800) crossed by inline compile instead of a swap on both PCs. From the 0.3.5 start (07:15 UTC) to 15:30 UTC: 0 `prepare-failed` lines on any machine, 4 and 4 `WORKER FAULT` lines, and every boundary from DAA 68,400 to 111,600 swapped with no pause on both PCs (intake query over the `miner-*` uploads, `docs/bench-log.md`, "FUD ledger sweep round 6", M11). The flapping-node simnet was not run; the live storm stands in for it. Earlier line, kept for the record: Fix built (4 October 2026, branch `miner-reliability`, `aea5ac6d`), pending merge. `PrepareLimiter`: one prepare per pair per epoch, one retry after prepare-failed, none within 30 s of the last, a held prepare said once; a unit test replays a flapping node. The fast-time simnet with a flipping `next_epoch_seed` was not run (it needs a patched node and a GPU worker). Was: Open (4 October 2026). Answer: Correct. `igneum/miner/src/main.rs:1113-1165, 1337-1339`; `worker.cpp:644-658`; `host.c:1208-1225`. Estimated loss 30 to 60 percent against a node that alternates seeds per template; a stale home node on a fork is the realistic trigger. Fix: at most one prepare per pair per epoch and none within 30 s of the last. Review id R4.2.2. @@ -1782,8 +1792,8 @@ Evidence: the files above. Experiment: a tampered pack with matching vectors mus ### X21. A wrong program burns power with a green rate "The CPU re-check counts mismatches and does nothing: no threshold, no stop. The app reads `hash`, `now`, `template_age` and `synced` from STATUS and nothing else, so `mismatched=` and `WORKER FAULT` never reach the card." -Status: Fixed (4 October 2026, evening), fork commit `aea5ac6d` (`guard::MismatchGuard`: three consecutive CPU re-check mismatches kill the worker, `WORKER FAULT cpu re-check ...`) and app commit `f39e240` on `miner-reliability` (`app/igneum-app/src/watchdog.rs`: the app reads `mismatched=`, `faults=` and the `WORKER FAULT` lines, shows them on the card, and its own watchdog restarts a miner once for no status in 90 s or a zero rate for 60 s while synced, then marks the card faulted; a silent node is restarted in-process). Measured: `docs/bench-log.md`, the same entry as M26. Replaced: Open (4 October 2026). -Status: Fix built (4 October 2026): `MismatchGuard` on branch `miner-reliability` (`aea5ac6d`) kills the worker after three consecutive CPU re-check mismatches (`WORKER FAULT cpu re-check`), and the app reads `mismatched=` and `WORKER FAULT` on branch `release-0.3.5` (`app/igneum-app/src/engine.rs`, 3 matches; 0 on master). Pending merge and release; the edited-kernel test needs a GPU worker (hardware). Was: Open (4 October 2026). +Status: Fixed (rolled out 5 October 2026, 0.3.5: the miner half with `miner-reliability` in fork 20139145, the app half with app commit f7d2af7 of the 0.3.5 branch, `docs/plans/release-0.3.5.md` 1a). Was: Fixed (4 October 2026, evening), fork commit `aea5ac6d` (`guard::MismatchGuard`: three consecutive CPU re-check mismatches kill the worker, `WORKER FAULT cpu re-check ...`) and app commit `f39e240` on `miner-reliability` (`app/igneum-app/src/watchdog.rs`: the app reads `mismatched=`, `faults=` and the `WORKER FAULT` lines, shows them on the card, and its own watchdog restarts a miner once for no status in 90 s or a zero rate for 60 s while synced, then marks the card faulted; a silent node is restarted in-process). Measured: `docs/bench-log.md`, the same entry as M26. Replaced: Open (4 October 2026). +Sweep (5 October 2026, evening): measurement line, live: every STATUS line of every card on the five machines carries `mismatched=0` and `faults=0` and the console reads the fields (Machines card at 15:46 UTC on 5 October 2026: `0 mismatched, 0 restarts` per card, `0 faults` per machine), so the path from the miner's counter to the card exists on the fleet; the edited-kernel test that would turn a card red was not run (it needs a hand-edited pack on a PC). Earlier line, kept for the record: Fix built (4 October 2026): `MismatchGuard` on branch `miner-reliability` (`aea5ac6d`) kills the worker after three consecutive CPU re-check mismatches (`WORKER FAULT cpu re-check`), and the app reads `mismatched=` and `WORKER FAULT` on branch `release-0.3.5` (`app/igneum-app/src/engine.rs`, 3 matches; 0 on master). Pending merge and release; the edited-kernel test needs a GPU worker (hardware). Was: Open (4 October 2026). Answer: Correct. `igneum/miner/src/main.rs:1250-1261`; `app/igneum-app/src/engine.rs:~1762-1780` (zero matches for either string). The PowerShell launcher matches them (`igneum-common.ps1:843`), which is what README.txt and TEST.md describe. Fix: stop the worker after 3 consecutive mismatches and show it on the card; the app reads both fields. Review id R4.2.3. @@ -1828,7 +1838,9 @@ Evidence: the files above. ### M30. A block or transaction flood grows the 0.3.4 node by hundreds of megabytes in a minute "On the 3 October ordering-layer node (no execution layer) the resource-exhaustion scenario grew RSS by 4, 11 and 14 MB and the 50x block flood by 30 MB. On the 0.3.4 build, same harness, same scenarios, same 60 s: template flood +6 MB, submit flood +269 MB, mempool flood +270 MB, block flood 302 to 1,082 MB on both nodes (567 MB at 10 s, 824 MB at 20 s). The harness calls it a pass because its bound is baseline + 512 MB; a peer that keeps going is not bounded by the harness." -Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-memory`, merged into `fud-consensus`; main repo branch `fud-memory`, merged into `fud-consensus`). Was: Open (4 October 2026, red-team run). Serious: a single peer at 50 blocks/s or 500 transactions/s is the devnet's own fast-miner event, and a node that grows 13 MB/s under it runs out of memory in minutes on the 2 to 4 GB cloud nodes. +Status: Fixed (rolled out 5 October 2026, 0.3.5: fork 20139145 = `finality-fixes` 6aa69a45 + `fud-consensus` 977db931 + `m20-pruning` d35b00cf + the miner branches, cut 07:33 BST, `docs/plans/release-0.3.5.md` 1b and 4b; node line 2b6d23ef from 0.3.6 at 09:34 UTC carries the same consensus code; every reachable app machine on the line by 09:32 UTC and the hand nodes and the seed restarted on it for the proving activation at DAA 84,100 the same morning, `docs/plans/release-0.3.6.md` 8j, bench-log "live devnet: the first shards proven"). Was: Fix built, pending rollout (4 October 2026, night; node branch `fud-memory`, merged into `fud-consensus`; main repo branch `fud-memory`, merged into `fud-consensus`). Was: Open (4 October 2026, red-team run). + +Sweep (5 October 2026, evening): measurement line, live: cache builds now equal node restarts, not epoch rolls. In the 10 hours to 15:45 UTC on 5 October 2026 the node logs show 5 `PoW cache built` lines on PC 1, 5 on PC 2, 8 on the Mac and 1 on Sam's Mac, each at a node start (the 0.3.5, 0.3.6, 0.3.7 and 0.3.8 restarts and the proving-activation restart), and none at the hourly boundaries: the swaps at DAA 108,000 (14:29 UTC) and 111,600 (15:29 UTC) built no cache on any machine (the last build on PC 1 is the 0.3.8 restart at 13:42 UTC). The 0.3.4 engine built one 256 MiB cache per hourly roll (about 270 MB an hour of growth on the app node). The 30 MB per 1,000 blocks of non-cache growth and the unbounded `ExecState.records` stay as stated. Serious: a single peer at 50 blocks/s or 500 transactions/s is the devnet's own fast-miner event, and a node that grows 13 MB/s under it runs out of memory in minutes on the 2 to 4 GB cloud nodes. The cause, measured (not the one guessed below): every RSS step in both floods is one `PoW cache built` line, 256 MiB each. The lottery engine (`consensus/pow/src/igneum.rs`, `IgneumEngine`) keyed its resident entries by `(epoch seed, day)` and built a full 256 MiB cache per entry, `KEEP = 4`, although the cache depends on the day seed alone (`igneum_pow::Epoch::from_seed_bytes`: cache from the day bytes, program from the epoch seed). The floods ran on the 60x profile, where an epoch rolls every 60 DAA, so the 50x block flood rolled it every 10 to 20 s and paid a cache each time; the 3 October run was on the devnet profile (3,600-DAA epochs, no roll in 60 s), which is what differed, not the execution layer. The s6 figures were cumulative from one starting RSS: the mempool flood's "+270 MB" was the submit flood's growth carried forward (its own cost is 1 MB), and 197 chain blocks cost under 1 MB in `ExecState.records`. On the live devnet the same engine costs 256 MiB per hourly epoch roll up to `KEEP`, which is the steady-state growth the coordinator saw on the 0.3.4 app node (1,081 MB at 27 min, 2,258 MB at 4 h 14 min, about 270 MB an hour). The fix: caches keyed by day, `KEEP_DAYS = 3` (3 x 256 MiB resident, plus at most 2 in-flight builds), programs keyed by `(epoch seed, day)` at a few KB each (`KEEP = 8`, LRU); an epoch roll on the same day builds no cache; node and miner compute the identical hash through `EpochRef` (unit test `epoch_rolls_share_the_day_cache`). No consensus rule changed. Measured (`docs/bench-log.md`, "ledger M30", two-node fast-time floods, before on the shipping finality-fixes build, after on `fud-memory` 796f758d): s6 submit flood +263 / +257 MB with 1 cache build each, after +3 / +2 MB and 0 builds; s7 50x block flood 302 to 1,085 MB with 3 builds, after 302 to 318 MB and 0 builds; template and mempool floods unchanged at +6 and +1 MB. Steady state, measured on two nodes at 1 block/s with no flood for 1,500 blocks on the 60x profile (`tools/harness/scenarios/s8-steady.mjs`, bench-log "ledger M30", steady-state paragraph): the shipping build reached 1,342 MB by block 514 after 9 cache builds (one per epoch roll; five 256 MiB chunks, the fifth an evicted one the allocator keeps) and 1,371 MB at 1,529 blocks; the fixed build held 319 MB at 510 blocks with 1 build and 603 MB at 1,526 blocks with 2 (the fast-time day rolled once), so on the devnet profile it is 256 MiB flat, 512 MiB around midnight UTC, 768 MiB worst case. Both builds then climb 30 MB per 1,000 blocks (30.7 before, 30.2 after), which is not the PoW cache (72 MB of non-cache footprint at 1,526 blocks against 23 MB at 0; the reading, not a measurement, is the consensus database's buffers and rusty-kaspa's entry-sized caches filling); the live node's 36 MB per 1,000 blocks between epoch rolls fits it. The app node's 2,258 MB at 4 h 14 min exceeds what this node build can reach (about 1.7 GB at 15,000 blocks) and includes its GPU worker, which needs its own `vmmap -summary`. Still growing by design and not bounded tonight: `ExecState.records` (1 to 2 KB per chain block, approximate, 100 to 170 MB a day at 1 block/s; a window must cover the proving sortition window and the by-number RPC history), `SNAPSHOT_RING = 64` state clones scaling with state size, the finality key registry (grows with distinct vote keys; votes, certificates, locks and evidence are trimmed). The harness now reports per-load RSS deltas and cache-build counts and takes `IGNEUM_HARNESS_BASE_PORT` and `IGNEUM_HARNESS_TMP`. @@ -1839,7 +1851,9 @@ Evidence: `/tmp/igneum-redteam-ord/results/s6-exhaustion.json` and `s7-flood.jso ### M31. The 0.3.4 node cannot produce a block template on mainnet, testnet or simnet parameters "`getBlockTemplate` on a `--simnet` node from the finality-fixes build answers every call with `Coinbase payload is above max length (204). Try to shorten the extra data.` and the network never makes a block. The coinbase of this build carries the vote-key reveal, the proof-record section and the finality section; only `DEVNET_PARAMS` was raised to `MAX_COINBASE_PAYLOAD_LEN_WITH_FINALITY` (16,384). `MAINNET_PARAMS`, `TESTNET_PARAMS` and `SIMNET_PARAMS` still carry Kaspa's 204 (`consensus/core/src/config/params.rs:705, 766, 828` against `:900`)." -Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`). Was: Open (4 October 2026, red-team run). Serious for anything that is not the devnet: a mainnet or testnet genesis on these parameters cannot be mined by a voting miner at all; harmless on the live devnet, whose parameters carry the raise. +Status: Fixed (rolled out 5 October 2026, 0.3.5: fork 20139145 = `finality-fixes` 6aa69a45 + `fud-consensus` 977db931 + `m20-pruning` d35b00cf + the miner branches, cut 07:33 BST, `docs/plans/release-0.3.5.md` 1b and 4b; node line 2b6d23ef from 0.3.6 at 09:34 UTC carries the same consensus code; every reachable app machine on the line by 09:32 UTC and the hand nodes and the seed restarted on it for the proving activation at DAA 84,100 the same morning, `docs/plans/release-0.3.6.md` 8j, bench-log "live devnet: the first shards proven"). Was: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`). Was: Open (4 October 2026, red-team run). Serious for anything that is not the devnet: a mainnet or testnet genesis on these parameters cannot be mined by a voting miner at all; harmless on the live devnet, whose parameters carry the raise. + +Sweep (5 October 2026, evening): measurement line: the unit test `largest_coinbase_fits_on_every_network` passed in every 0.3.5 suite run (`cargo test -p kaspa-consensus-core`, 101 passed on the 0.3.6 fork tip a11455e7 on 5 October 2026, `docs/plans/release-0.3.6.md` 3b) and the testnet identity `igneum-testnet-1` adopted the same morning carries the raised limit with every switch at 0, so the first testnet genesis is mineable by a voting miner; no simnet network has been started since the cut, so the template line on simnet is covered by the test, not a run. The fix: `max_coinbase_payload_len` is `MAX_COINBASE_PAYLOAD_LEN_WITH_FINALITY` (16,384) on mainnet, testnet and simnet as on devnet, and the template builder fits the finality section into what the payload has left (`finality::encode_section_within`, certificates first, then evidence, then votes; the RPC computes the budget from the fixed part, the script, the node version, the miner's extra data and the record section), so a coinbase can never exceed the limit on any network whatever the voter count. The worst case did not fit even on devnet before: 8 certificates over 8,192 voters, 8 pieces of evidence and 48 votes come to about 30 KB, so the per-item bounds alone were not a bound. Unit test `largest_coinbase_fits_on_every_network` (consensus-core): the fixed part, the version, a key reveal, a full record section (8 records of 274 bytes) and the cut finality section fit under every network's limit, the cut keeps every certificate and every piece of evidence and at least 8 votes, the uncut section would not fit, and a budget under one certificate yields an empty section. The exec-attacks network (`tools/exec-attacks/net.sh`, `--simnet`) can drop its override once this build ships; not re-run tonight. @@ -1898,7 +1912,7 @@ Evidence: the commits above. Experiment: `curl https://igneum.network/api/live` - **Moved to Fixed or Rolled out, from evidence that already existed and had not reached the ledger:** M15 (cheap checks before the PoW engine, merged and live), M17 (hot swap, measured on the live devnet across three vendors), M19 (the census cells filled, 16 loads a Definition), M24 (rule v2 activated at DAA 33,000), P20 (the buffered save confirmed on the third run), X16 (the evidence page exists), G11 (the spec is public). - **Moved to Answered with evidence:** F1 (the first-month gate implemented, measured on 4 October and re-confirmed tonight on the finality-fixes build: the burster locks nothing alone; the launch month is arithmetic now), F7 (reorg depth p50 1, p99 3, max 5 on a 12-node, 5-region network), F19 (bought keys are worth their blocks; scenario K at both floors), E12 (the simulation half), E15 (the security-budget model: the floor is crossed in year 7, 11 or never by price, and no fee level moves it). -- **Fix built on a branch, pending merge or rollout:** M26, M27, X21 (`miner-reliability` `aea5ac6d`, with the app half on `release-0.3.5`), F22 (rule v3, fast-time network), part of X22. +- **Fix built on a branch, pending merge or rollout:** M26, M27, X21 (`miner-reliability` `aea5ac6d`, with the app half on `release-0.3.5`), F22 (rule v3, fast-time network), part of X22. Sweep (5 October 2026, evening): M26, M27, X21, F23, F24, G12, X18, M30, M31, F25 and M20 are now "Fixed (rolled out 5 October 2026, 0.3.5)" with their live measurement lines; F21 and F22 are shipped in every node since 0.3.4 and wait only for the switch N3 (decision owner the project lead). - **Open with new evidence and a fix row (fud-fixes section 2.5):** M20 (the stub is still in the pruning-proof path on `devnet-v4`; it goes live the moment the devnet passes its pruning depth), M21 (the k table from the fork's own function), M28, X20, P15, X29 (half fixed: file modes), X14 (hashing concentration measured), M25 (confirmed live in the sweep: a mismatched day length is rejected as `BlockInvalid` with no reason named, 0 of 4 against 7 of 7), M14 and F14 (no amplification in the chain model under rule v2 or Kaspa's rule, nor on the finality-fixes node under fast time, s5 ratio 0.864; the finality-with-DAA run still owed). - **Open, needs hardware, a person or the devnet:** M1, M11 (a rig), M16 (the 5090), M22, X19 (the node line; `faketime`), X23 to X28 (the relay owner and PC 1; the token is rotated, the run-task binding is not), E13, E14, E16, E17 (the draw lines), F16, F20 (the devnet test), C4 (module off), P3, P14, P16, P17, P21, P22, X13, X15, X17, G10, L1 to L5, L8, D5, D6. - **Nothing became worse.** Two things are closer than they look: M20 becomes a live failure for every fresh node once the devnet's pruning point leaves genesis, which at the devnet's 1.05 DAA/s (DAA 33,000 at 17:37 UTC on 4 October) is between DAA 108,000 (`PRUNING_DURATION`, about 14:00 UTC on 5 October) and DAA 151,200 (the first finality point a full pruning depth below the tip, about 01:00 UTC on 6 October), approximate; X23's operational half (a run task on the PCs needs only the relay token) is unchanged after the rotation. diff --git a/docs/spec/05-fees-and-economics.md b/docs/spec/05-fees-and-economics.md index cf6ed1970..9695fbf2f 100644 --- a/docs/spec/05-fees-and-economics.md +++ b/docs/spec/05-fees-and-economics.md @@ -11,7 +11,7 @@ Designed. Every transaction pays a base fee in both gas dimensions: | Dimension | What it meters | Who sets it | |---|---|---| | Execution gas | EVM execution, Ethereum's rule | Ethereum's EIP-1559-style base fee over the ordered sequence | -| Proving-cost gas | Proving cycles the transaction will cost the provers | A second base fee adjusted per block from the unproven backlog, smoothed over the difficulty window (section 2.3) so cards do not flip between hashing and proving every block | +| Proving-cost gas | Proving cycles the transaction will cost the provers | A second base fee `f_p`, adjusted per chain block by the same EIP-1559 step as `f_e`: toward a target of `B_p / 2` of proving gas used, denominator 8, never below the floor of section 5.11 (one definition, 5 October 2026, ledger P14; `next_base_fee` in `igneum/exec/src/executor.rs`, applied per chain block in `service.rs`). The unproven backlog does not move `f_p`; it halves `B_p` (design 4.3, the backlog rule), which raises `f_p` through the step. No smoothing over the difficulty window is implemented or specified: the two-dimension step is per chain block | The base fee in both dimensions is **burned in full**. A miner cannot stuff blocks with its own transactions for free; wash gas loses its whole base fee (ledger E3). The proving-cost budget per block is a consensus constant set from measured prover throughput (phase 2 gate: one shard on a 12 GB card in about 20 s, Target, unmeasured, ledger P1), so a transaction that is cheap to run and brutal to prove cannot stall the provers for everyone. How the node folds the proving-cost dimension into the quoted gas price so `eth_estimateGas` keeps working is fixed in section 7.1 (ledger P5, closed 3 October 2026). diff --git a/tools/finality-attacks/c4.mjs b/tools/finality-attacks/c4.mjs new file mode 100644 index 000000000..840b4b8d2 --- /dev/null +++ b/tools/finality-attacks/c4.mjs @@ -0,0 +1,170 @@ +// Ledger C4 (5 October 2026, evening): does the finality overlay change GHOSTDAG's fork choice, measured. The same +// split is run with the module ON (rule v3 from checkpoint DAA 0) and OFF (`finality.min_daa` never, so no +// certificate can ever form and fork choice is bare GHOSTDAG on the same binary). Fast-time 3-node network on +// ports 29800+, network igneum-devnet-980, data under /tmp/igneum-fin-c4; the live devnet is never touched. +// +// node tools/finality-attacks/c4.mjs on; node tools/finality-attacks/c4.mjs off # one mode per process +// SPLIT=150 WARM=230 HEAL=200 node tools/finality-attacks/c4.mjs on +// +// Topology (as v3.mjs): n1 listens; n0 dials n1 through proxy P0, n2 dials n1 through proxy P2; cutting P0 isolates +// n0 (side A) from n1 and n2 (side B). +// +// The scenario, "weight against work": before the cut side B holds 70% of the weight table (p0..p3 at share 0.175 +// on n1 and n2) and side A 30% (q0, q1 at 0.15 on n0), one block per second in all. At the cut every miner is +// restarted with the rates swapped: side A mines at RA blocks/s (0.6) and side B at RB (0.4), so during the split +// side A builds the heavier chain by blue work while side B, under the frozen weight table of rule v3, is the only +// side that can certify a checkpoint (A holds 30% of the frozen table and cannot lock until the table expires, +// 120 DAA of its own blocks later; B needs 50 DAA of its own blocks for its first new lock: RB / RA must exceed +// 50 / 120 and RA must exceed RB, hence 0.6 / 0.4). At the heal GHOSTDAG alone follows A's heavier chain; the +// overlay requires every candidate tip to pass through B's certified checkpoint. The measurement is which chain +// the three nodes converge to, whether they converge at all, and what each node had to reorganise. + +const ROOT = new URL('../../', import.meta.url).pathname; +const NODE_ROOT = process.env.IGNEUM_NODE_ROOT || '/Users/joshm/Projects/igneum/'; +process.env.IGNEUM_FIN_BASE_PORT ||= '29800'; +process.env.IGNEUM_FIN_SUFFIX ||= '980'; +process.env.IGNEUM_FIN_TMP ||= '/tmp/igneum-fin-c4'; +process.env.IGNEUM_FAST_TIME ||= '1'; +// the live node line (fork 2b6d23ef, the 0.3.6 to 0.3.8 node) built on the Mac on 5 October 2026 +process.env.IGNEUMD ||= `${NODE_ROOT}vendor/igneum-node/target-036/release/igneumd`; +process.env.IGNEUM_MINER ||= `${NODE_ROOT}vendor/igneum-node/target-036/release/igneum-miner`; +const DELAY_MS = +(process.env.DELAY_MS || 100); +const WARM = +(process.env.WARM || 230), SPLIT = +(process.env.SPLIT || 150), HEAL = +(process.env.HEAL || 200); +const RA = +(process.env.RA || 0.6), RB = +(process.env.RB || 0.4); +// ONE mode per process: lib/net.mjs reads IGNEUM_FIN_OVERRIDE_JSON when it is imported, so the override must be in +// the environment before the import (the first draft set it inside network() and ran rule v2 twice; 5 October 2026). +const MODE = process.argv.slice(2).filter(a => !a.startsWith('--'))[0] || 'on'; +if (MODE !== 'on' && MODE !== 'off') { console.error(`mode must be on or off, got ${MODE}`); process.exit(2); } +process.env.IGNEUM_FIN_OVERRIDE_JSON = JSON.stringify(MODE === 'off' ? { finality: { min_daa: 9007199254740991 } } : { finality_v3_activation_daa: 0 }); + +const { Node, Miner, Proxy, stopAll, sleep, log, assertBinaries, TMP, IGNEUMD } = await import('./lib/net.mjs'); +const { mkdirSync, writeFileSync, appendFileSync } = await import('node:fs'); +mkdirSync(TMP, { recursive: true }); +const results = []; +const out = (line) => { console.log(line); appendFileSync(`${TMP}/results-${MODE}.md`, line + '\n'); }; + +const lockedMap = (cp) => new Map((cp?.checkpoints || []).filter(c => c.state === 'locked').map(c => [c.index, c.hash])); +const maxLocked = (cp) => Math.max(0, ...lockedMap(cp).keys()); +async function checkpoints(node, last = 800) { return node.rpc.call('getFinalityCheckpoints', { last }).catch(() => null); } +async function peers(node) { const r = await node.rpc.call('getConnectedPeerInfo', {}).catch(() => null); return (r?.peerInfo || r?.infos || []).length; } +async function dag(node) { return node.rpc.call('getBlockDagInfo', {}).catch(() => null); } +async function blueScore(node, hash) { const r = await node.rpc.call('getBlock', { hash, includeTransactions: false }).catch(() => null); return Number(r?.block?.verboseData?.blueScore ?? NaN); } +async function isChainAncestor(node, hash) { + // the block is on the node's selected chain when the node's own checkpoint record for its index names it; cheaper and + // exact: ask the chain from the pruning point to the sink and look for it + const d = await dag(node); if (!d) return null; + const r = await node.rpc.call('getVirtualChainFromBlock', { startHash: d.pruningPointHash, includeAcceptedTransactionIds: false }).catch(() => null); + if (!r) return null; + return (r.addedChainBlockHashes || []).includes(hash); +} + +async function network(mode) { + const n1 = new Node(1, { name: 'n1' }); + await n1.start(); + const p0 = new Proxy(0, n1.p2pPort, { delayMs: DELAY_MS }); await p0.start(); + const p2 = new Proxy(2, n1.p2pPort, { delayMs: DELAY_MS }); await p2.start(); + const n0 = new Node(0, { name: 'n0', connect: [p0.addr] }); + const n2 = new Node(2, { name: 'n2', connect: [p2.addr] }); + await n0.start(); await n2.start(); + await sleep(3000); + log(`network up (${mode}): peers n0 ${await peers(n0)} n1 ${await peers(n1)} n2 ${await peers(n2)}; override ${process.env.IGNEUM_FIN_OVERRIDE_JSON}`); + return { n0, n1, n2, p0, p2 }; +} + +async function run(mode) { + const name = `weight-vs-work-${mode}`; + const { n0, n1, n2, p0 } = await network(mode); + const secsWarm = WARM + 30, secsSplit = SPLIT + HEAL + 60; + const warmPlan = [[n0, 'q0', 0.15], [n0, 'q1', 0.15], [n1, 'p0', 0.175], [n1, 'p1', 0.175], [n2, 'p2', 0.175], [n2, 'p3', 0.175]]; + let miners = warmPlan.map(([node, label, share]) => new Miner(node, { label, share, bps: 1, secs: secsWarm }).start()); + await sleep(WARM * 1000); + const w = await n1.rpc.call('getFinalityWeights', {}).catch(() => ({})); + const before = await Promise.all([n0, n1, n2].map(n => checkpoints(n))); + const beforeMax = before.map(maxLocked); + const preMax = Math.max(...beforeMax); + const dagsCut = await Promise.all([n0, n1, n2].map(dag)); + log(`${name}: cut at warm ${WARM} s: window daa ~${w.daaScore}, voters ${w.voters}, max locked ${beforeMax.join('/')}, blocks ${dagsCut.map(d => d?.blockCount).join('/')}`); + // the cut, and the rates swapped: A at RA (two keys), B at RB (four keys) + for (const m of miners) await m.stop(); + const tCut = Date.now(); + p0.cut(); + const splitPlan = [[n0, 'q0', RA / 2], [n0, 'q1', RA / 2], [n1, 'p0', RB / 4], [n1, 'p1', RB / 4], [n2, 'p2', RB / 4], [n2, 'p3', RB / 4]]; + miners = splitPlan.map(([node, label, share]) => new Miner(node, { label, share, bps: 1, secs: secsSplit }).start()); + const firstNew = [null, null, null], maxNew = [...beforeMax]; + while (Date.now() - tCut < SPLIT * 1000) { + const cps = await Promise.all([n0, n1, n2].map(n => checkpoints(n))); + cps.forEach((cp, i) => { + const m = maxLocked(cp); + if (m > maxNew[i]) maxNew[i] = m; + if (firstNew[i] == null && m > preMax) firstNew[i] = Math.round((Date.now() - tCut) / 1000); + }); + await sleep(3000); + } + const newLocks = maxNew.map((m, i) => Math.max(0, m - preMax)); + // the two chains at the end of the split + const dagsEnd = await Promise.all([n0, n1, n2].map(dag)); + const sinkA = dagsEnd[0]?.sink, sinkB = dagsEnd[1]?.sink; + const bsA = await blueScore(n0, sinkA), bsB = await blueScore(n1, sinkB); + const cpsEnd = await Promise.all([n0, n1, n2].map(n => checkpoints(n))); + const bLockedDuring = [...lockedMap(cpsEnd[1]).entries()].filter(([i]) => i > preMax); + log(`${name}: end of split: A sink blue score ${bsA} (${dagsEnd[0]?.blockCount} blocks), B sink blue score ${bsB} (${dagsEnd[1]?.blockCount} blocks); B locked ${bLockedDuring.length} new index(es) ${bLockedDuring.map(([i]) => i).join(',')}; A locked ${newLocks[0]}`); + p0.heal(); + const tHeal = Date.now(); + let reconnected = null; + while (Date.now() - tHeal < HEAL * 1000) { + if (reconnected == null && (await peers(n0)) > 0) reconnected = Math.round((Date.now() - tHeal) / 1000); + await sleep(3000); + } + for (const m of miners) await m.stop(); + await sleep(4000); + const after = await Promise.all([n0, n1, n2].map(n => checkpoints(n))); + const afterMax = after.map(maxLocked); + const dagsAfter = await Promise.all([n0, n1, n2].map(dag)); + const sinks = dagsAfter.map(d => d?.sink); + const converged = new Set(sinks).size === 1; + // where did the network end: on A's split chain, on B's, or on neither (a merge of both is still "through" one) + const onA = await Promise.all([n0, n1, n2].map(n => isChainAncestor(n, sinkA))); + const onB = await Promise.all([n0, n1, n2].map(n => isChainAncestor(n, sinkB))); + const maps = after.map(lockedMap); + let disagree = 0; + const common = new Set([...maps[0].keys()].filter(k => maps[1].has(k) && maps[2].has(k))); + for (const k of common) if (new Set(maps.map(m => m.get(k))).size > 1) disagree++; + const bAdoptedByA = bLockedDuring.every(([i, h]) => maps[0].get(i) === h); + const conflicts = [n0, n1, n2].map(n => n.grepLog(/CONFLICTING certificate/).length); + const redetermined = [n0, n1, n2].map(n => n.grepLog(/re-determined/).length); + const reorgs = [n0, n1, n2].map(n => n.grepLog(/reorg deeper than the snapshot ring|Reorg|reorg/i).length); + await stopAll(); + const heavier = bsA > bsB ? 'A' : 'B'; + const ended = converged ? (onB[0] && !onA[0] ? 'B' : onA[0] && !onB[0] ? 'A' : onA[0] && onB[0] ? 'both merged' : 'neither') : 'not converged'; + const pass = mode === 'on' + ? (newLocks[0] === 0 && bLockedDuring.length > 0 && converged && ended === 'B' && conflicts.every(c => c === 0) && disagree === 0) + : (newLocks.every(x => x === 0) && converged && ended === heavier); + out(`\n### ${name}: warm ${WARM} s at 1 block/s (B 70% of weight, A 30%), split ${SPLIT} s with A at ${RA} and B at ${RB} blocks/s, heal window ${HEAL} s, link delay ${DELAY_MS} ms, module ${mode} (${mode === 'on' ? 'rule v3 from checkpoint DAA 0' : 'min_daa never: no certificate can form'}), node ${IGNEUMD.split('/').slice(-3).join('/')}\n`); + out('| measure | n0 (side A, work majority) | n1 (side B, weight majority) | n2 (side B) |'); + out('|---|---|---|---|'); + out(`| max locked index at the cut | ${beforeMax.join(' | ')} |`); + out(`| new locks during the split (index above ${preMax}) | ${newLocks.join(' | ')} |`); + out(`| first new lock, s after the cut | ${firstNew.map(x => x ?? 'none').join(' | ')} |`); + out(`| max locked index at the end of the heal window | ${afterMax.join(' | ')} |`); + out(`| sink at the end of the heal window | ${sinks.map(s => String(s).slice(0, 10)).join(' | ')} |`); + out(`| A's split tip on the final chain / B's split tip on the final chain | ${onA.map((a, i) => `${a} / ${onB[i]}`).join(' | ')} |`); + out(`| conflicting certificates logged | ${conflicts.join(' | ')} |`); + out(`| re-determined lines (F24) | ${redetermined.join(' | ')} |`); + out(`| reorg lines in the node log | ${reorgs.join(' | ')} |`); + out(`\nAt the end of the split: A's sink blue score ${bsA} against B's ${bsB} (the heavier chain by blue work is ${heavier}'s); B locked ${bLockedDuring.length} new checkpoint(s) during the split${bLockedDuring.length ? ' at index ' + bLockedDuring.map(([i]) => i).join(', ') : ''}. After the heal: n0 reconnected ${reconnected == null ? 'not within the heal window' : reconnected + ' s after the gate reopened'}; the three sinks ${converged ? 'agree' : 'DISAGREE'}; the network ended on ${ended}'s chain; A adopted B's split-time locks: ${bAdoptedByA}; locked indices disagreeing across the three nodes: ${disagree}. ${pass ? 'PASS' : 'FAIL'} against the expectation for module ${mode} (${mode === 'on' ? "B's certified chain wins although A's is heavier" : 'the heavier chain wins'}).`); + results.push({ name, pass, heavier, ended, converged, bsA, bsB, newLocks, bLocked: bLockedDuring.length, conflicts, disagree }); +} + +async function main() { + assertBinaries(); + for (const mode of [MODE]) { + log(`=== module ${mode} starting ===`); + try { await run(mode); } catch (e) { log(`${mode} threw: ${e.stack || e}`); results.push({ name: mode, pass: false }); await stopAll(); } + log(`=== module ${mode} done ===`); + } + out('\n' + results.map(r => `[${r.pass ? 'PASS' : 'FAIL'}] ${r.name}`).join('\n')); + writeFileSync(`${TMP}/results-${MODE}.json`, JSON.stringify(results, null, 2)); + await stopAll(); + process.exit(results.some(r => !r.pass) ? 1 : 0); +} +main();