From 7dade79bcbc3772243ed079d5f3ddf44043b96df Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 01:30:25 +0000 Subject: [PATCH] Aggregation cost on the 5090 (5 October, night): the chained aggregation is 2.1 s alone and 9.7 s beside the miner, the batch-log2 curve (2^16 buys 1.6x for a fifth of the hash rate), batch and tree folds estimated, two streams and SP1 knobs closed; the host times the stdin build and names the knobs, --save-shards; the PC 2 job scripts and the readers The statement and the pinned guests are unchanged; every fixture proof verifies as before. The defaults stay (batch-log2 22, SP1 defaults): the one knob that moves a mining card's prover costs a fifth of the hash rate; the plan carries the trade for the project lead and the batch fold for the next pin. Measured: docs/bench-log.md "aggregation cost on the RTX 5090"; the plan line: docs/plans/proving-v1.md "Aggregation cost (5 October, night)". Also: make-package's gate skips the exporter's .node-plan.json side files and takes the run lock for its execute step; the state-reply class (/api/state answering {} once paid_wei passes u64::MAX) found on the way and fixed on the app branch at 42f36b3. Co-Authored-By: Claude Fable 5.1 (cherry picked from commit ea38ece9eaa5949dd657cbfea5308c4948177ff8) --- docs/bench-log.md | 87 ++++ docs/plans/proving-v1.md | 19 + proving/igneum-prove/host/src/main.rs | 30 +- proving/igneum-prove/host/src/proof_system.rs | 14 + proving/windows-wsl2/make-package.sh | 6 +- tools/proving-v1/agg-cost-table.mjs | 50 ++ tools/proving-v1/miner-rate.mjs | 35 ++ tools/proving-v1/pc2-agg-cost-restore.ps1 | 57 +++ tools/proving-v1/pc2-agg-cost.ps1 | 433 ++++++++++++++++++ 9 files changed, 723 insertions(+), 8 deletions(-) create mode 100644 tools/proving-v1/agg-cost-table.mjs create mode 100644 tools/proving-v1/miner-rate.mjs create mode 100644 tools/proving-v1/pc2-agg-cost-restore.ps1 create mode 100644 tools/proving-v1/pc2-agg-cost.ps1 diff --git a/docs/bench-log.md b/docs/bench-log.md index a0ebfa9e..81974853 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1709,3 +1709,90 @@ Reported by the aggregation-cost agent: PC 2's `/api/state` answered `{}` (2 byt Cause: `ProvingState.paid_wei: u128` and serde_json `to_value` (1.0.151, `value/ser.rs` `serialize_u128`: u64 range or an error); the error became `json!({})`. Fix: the field serialises as a decimal string; `state_json` logs the error once. Test `a_paid_total_over_u64_max_still_serialises_the_whole_state` (`cargo test --offline -q paid_wei`, 1 passed). +| The fast-time 3-node harness (`tools/proving-v1/net.mjs`, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac) | run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (`tools/proving-v1/report-2026-10-05.json`). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; `segmentsInWindow` proven 3, unproven 1. The shard side: a v1 shard's `shardWei` = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed | + +## 5 October 2026 (night), aggregation cost on the RTX 5090: what a per-block aggregation spends and what each lever gives (proving engineer, agg-cost) + +the project lead, 5 October 2026: "fix everything else in the numbers tonight". The number under test: the chained segment aggregation cost 9.6 to 9.7 s a block on PC 2's 5090 while the card mined (`chain-pc2-pv1c`, the entry above), 2.2 s on 4 October with the card to itself. Target: under 3 s a block, the miner's slowdown of the prover under 1.5x, the proof statement unchanged. Branch `agg-cost` (worktree `igneum-wt-agg-cost`, from `proving-v1` 219517f). Host changes (statement untouched, `elf/` untouched): the aggregation's stdin build timed apart from the prove call, the deferred-proof count and the SP1 knobs in the RESULT lines, `--mode chain --save-shards` (every shard's compressed proof written next to the results, so `--mode aggregate` re-runs the same proofs under other settings). Jobs: `agg-cost-pc2-1` (21:01:20Z to 21:25:11Z, `tools/proving-v1/pc2-agg-cost.ps1`, the package `igneum-prove-wsl2-aggcost.zip` fetched by `fetch-prove-aggcost` 20:55:39Z, built in WSL2 against the live target dir in 5 s, installed to `/opt/igneum-aggcost`, the live `/opt/igneum` untouched, `--mode id` the pinned pair) and `agg-cost-pc2-2` (21:34:00Z, the same script). The live prover was switched OFF for the runs (its sp1-gpu-server would otherwise be shared through `/tmp/sp1-cuda-0.sock` and carry its own environment; `gpu_server_before running=0`) and ON again at the end. Fixtures: four consecutive live blocks cut from PC 2's own node (86165..86168 at tip 86195, one empty shard each, every one MATCHES natively), the same four for every phase of job 1. App 0.3.9 on PC 2 throughout. + +Known-finished case of the host changes before the GPU (this Mac, CPU, run lock, 20:41Z to 20:44Z): `--mode chain` over `fixtures/chain/block-81046.json` with `--save-shards` (shard 38.5 s, aggregate 43.4 s, the proof file written), then `--mode aggregate` over that saved shard proof with `SP1_WORKER_VERIFY_INTERMEDIATES=false` (46.6 s, the same statement `0x3dedb8ea...`), `--mode verify-segment` VERIFIED in 0.027 s; known-failed: a wrong statement NOT VERIFIED in 0.027 s. Unit tests: `cargo test --release -p igneum-prove-core -p igneum-prove-host`: core 8 passed, host 9 passed and 1 ignored (build lock, 20:53Z). + +### Lever 1, the profile: where a per-block aggregation goes + +| What | Measured (job `agg-cost-pc2-1`) | +|---|---| +| The host's own share of an aggregation (the stdin build: the AggInput, the proof clones into the request) | 0.000 s on every block, mining or idle (the `stdin` field of every `RESULT chain block` line): everything is inside the one `prove().compressed()` call to the GPU server | +| The GPU server's log at `RUST_LOG=info` (phase A0, the same chain of 1, stderr captured) | 1 line: sp1-gpu-server 6.8.1 prints no spans and no timings, so the step costs below are read from the deferred-proof count, not from a profiler | +| Aggregation with 1 deferred proof (the first block, no previous proof) against 2 (every chained block), the card mining | 7.9 s against 9.6, 9.6, 9.8 s: the second deferred proof costs 1.7 to 1.9 s under the miner | +| The same, the miners paused (phase C, the same fixtures, 21:05:51Z) | 1.7 s against 2.1, 2.1, 2.2 s: the second deferred proof costs 0.4 to 0.5 s alone | +| The shard proof of an empty shard | 7.4 to 7.8 s mining, 1.9 to 2.2 s alone | +| A whole block (one empty shard plus its aggregation) | 17.1 to 17.3 s mining (end to end 67.4 s for 4 blocks), 4.1 s alone (16.4 s for 4) | +| GPU utilisation over the phase (1-s `nvidia-smi` samples) | 93.9% mining (80 samples, the miner's), 15.8% alone (32 samples): the prover alone keeps the card busy a sixth of the time. Its work is short GPU bursts between CPU phases (the executor, the witness and recursion-program generation run on the CPU inside the server), and the miner's kernels fill the gaps | +| GPU memory peak | 16,195 MiB mining (the miner's 3.4 GB resident), 14,483 MiB alone | +| The slowdown by the miner, same fixtures, same host, 2 min apart | shards 3.6x, the first aggregation 4.6x, a chained aggregation 4.5x, a block 4.2x | +| Setup per host process (client plus two key setups) | 13.0 to 15.7 s, mining or not | + +Reading. An aggregation is three or four recursion steps on the card (the aggregator guest's one core shard, its lift, one deferred program per verified proof, the compose), each a burst of under half a second when the card is free. The chained aggregation's extra deferred proof is the only part that grows with the chain rule, 0.4 to 0.5 s alone. Everything else the 9.7 s holds is the miner: with the card at 94% from the lottery kernels, every prover burst waits for a time slice, and a 2.1-s aggregation becomes 9.7 s. The 4 October 2.2 s (two shards, no previous proof, the card to itself) and tonight's 1.7 s (one shard) and 2.1 s (one shard plus the previous proof) agree within the deferred count. + +### Lever 2, batch and tree folds (estimate from the measured step costs; the statement is pinned, no guest was changed tonight) + +A fold of K blocks' shard proofs plus the previous segment proof in ONE aggregator call would cost one core shard, one lift, K + 1 deferred programs and the compose tree in place of K chained aggregations. From the measured rows (alone: a 1-deferred aggregation 1.7 s, each further deferred proof 0.45 s; mining: 7.9 s and 1.8 s): + +| Fold | Deferred proofs per call | Per block, card alone (estimate) | Per block, card mining (estimate) | Rule | +|---|---|---|---|---| +| chained, as pinned (measured) | 2 | 2.1 s | 9.7 s | one call per block | +| batch of 4 | 5 | (1.7 + 4 x 0.45) / 4 = 0.9 s | (7.9 + 4 x 1.8) / 4 = 3.8 s | one call per 4 blocks | +| batch of 8 | 9 | (1.7 + 8 x 0.45) / 8 = 0.7 s | (7.9 + 8 x 1.8) / 8 = 2.8 s | one call per 8 blocks | +| tree of 4 (2 + 2, then the pair) | 3 per call, 3 calls | 3 x (1.7 + 2 x 0.45) / 4 = 1.9 s | 3 x (7.9 + 2 x 1.8) / 4 = 8.6 s | no gain over the chain: every call pays the fixed part | + +Reading. A batch fold halves to quarters the per-block aggregation but changes the aggregator's statement (`AggInput` carries one block's shards and the guest asserts one block hash), so it is a new pinned guest and a new program id: a provers-off drain and a rollout (proving/README.md, pinned guests). It does not reach 3 s on a mining card by itself (2.8 s at K = 8 is on the line), and the shard proof beside it stays 7.4 s a block on a mining card. The lever that moves both is the card's other job, lever 4. A tree fold gains nothing here because the fixed part of a call (the core shard and the lift) dominates the per-proof part 4 to 1. + +### Levers 3 and 4, two streams and the miner's kernels (job `agg-cost-pc2-2` and the re-run) + +Job `agg-cost-pc2-2` (21:34:00Z to 21:49:22Z) ran with the 5090 idle throughout: job 1's `/api/resume` had left the worker off (below), so the rows that needed the miner (the batch-log2 curve, the two streams beside the miner, the time-slice policy, the chosen combination) are void and wait for a re-run; the idle rows are measured. + +| What | Measured (job `agg-cost-pc2-2`, card idle) | +|---|---| +| Aggregate-only over job 1's four saved shard proofs (`--mode aggregate --proofs b1;b2;b3;b4 --parent ...`, one process, the same statement `0x3a995f24...` as the chain run), default knobs (phase B0, then C1) | 1.7, 2.0, 2.0, 2.0 s (1, 2, 2, 2 deferred proofs), 8.1 s for four; C1: 1.8, 2.1, 2.1, 2.1 s, 8.3 s | +| The same with `SP1_WORKER_VERIFY_INTERMEDIATES=false` (phase B; the server inherits the host's environment, the knob printed in the `sp1 knobs` line) | 1.7, 2.0, 2.0, 2.0 s, 7.8 s for four: no gain (0.3 s over four, inside the run-to-run spread of 0.2 s). The knobs that change the recursion shape (`SP1_WORKER_MAX_COMPOSE_ARITY`, `MAX_REDUCE_ARITY`) were not tried: a different shape is a different recursion key set and the pinned verifier would refuse the proof | +| A 4-deferred aggregation (block-344-shards4, four prototype shards of 6.75 M pgas, phase C2) | shards 42.8 s (10.7 s each, the 4 October 10.2 to 10.7 s), aggregation 2.4 s with 4 deferred proofs; GPU peak 28,402 MiB (the prototype shard's 28.3 GB), utilisation 27.7% over the phase. With 1.7 s at one deferred proof and 2.0 to 2.1 s at two: 0.25 s per further deferred proof alone, so a batch of 8 would cost about 3.5 s a call, 0.45 s a block (estimate, the pinned statement forbids it) | +| Two host processes at once on the one card (phase G0: chains of 2 on disjoint blocks, started 2 s apart) | both connected to ONE sp1-gpu-server (the first process's child; the socket is per device, `/tmp/sp1-cuda-0.sock`): process 1 shard 2.2 and 3.5 s, aggregation 3.0 and 4.0 s (12.9 s for 2 blocks against 8.2 s alone); process 2 shard 3.3 s, aggregation 3.6 s, then its second block died with `CudaClientError: Failed to read the response: early eof` when process 1 finished and its server exited. GPU 24,911 MiB, utilisation 12.4% and 13.1%. Two streams through SP1 6.8.1's server are serialised on one socket and the second dies with the first: no throughput gain (3 blocks in 33 s against 4 in 16.4 s) and a failure mode; lever 3 is closed on this SP1 version | +| Job 3 (`agg-cost-pc2-3`, 22:41:15Z, app 0.3.10, the same script with the socket rule and a card switch): phase A, the app's 5090 miner at 117.0 MH/s mean (n 3, STATUS lines 22:44:45Z to 22:46:11Z), four fresh live blocks 90896..90899 | shards 8.0, 7.8, 7.6, 7.8 s; aggregations 8.0 s (1 deferred), 10.0, 10.0, 10.0 s (2 deferred); 69.5 s for four, 17.8 s a block; GPU 93.8%, peak 16,245 MiB: the job-1 baseline reproduced 100 min later on other blocks | +| Job 3's own-miner phases | void: the state reads came back empty (the class below), the card switch did nothing, phase D launched my miner beside the app's (the app's dropped to 62.2 MH/s, mine read 60.6 MH/s), then PC 2's app restarted at 23:03:30Z and the job died with it; no curve point | +| The GPU time-slice policy (`nvidia-smi compute-policy --set-timeslice`, the restore job `agg-cost-restore-1`, 23:16:53Z) | "Not Supported" on PC 2 (RTX 5090, driver 13.3, the Windows nvidia-smi, not elevated): the lever is closed on this driver; an elevated try is not worth a slot, the error is the driver's, not a permission's | +| The own-miner phases of job 2 | void: no 5090 miner was running to copy the command line from (the worker off since 21:25Z) | + +The curve, job `agg-cost-pc2-6` (01:12:09Z to 01:24:14Z, app 0.3.11, PC 2 to itself; every phase closed before the next job landed on PC 2 at 01:24:21Z). The app's 5090 miner switched off through `/api/cards` (the keys from `settings.json`; the worker was still alive after 120 s, `/api/pause` as the fallback stopped it in 5 s), then the job's OWN miner on the 5090 with the app's command line (`igneum-miner mine ... --worker igneum-worker-cuda.exe --identities 8 --worker-args "--device 0 --pack packs\devnet --race off [--batch-log2 B]"`, the base variant, its STATUS line every 10 s), the same four live blocks 96556..96559 (one empty shard each) proven by `--mode chain` under it, the miner's rate from its own `now=` field (the first two lines skipped). `--batch-log2 B` sets the worker's nonces per kernel launch (2^B; 22 is the worker's default, 4,194,304 nonces, about 35 ms a launch at 120 MH/s; `proto-cuda/nvrtc/worker.cpp`). + +| batch-log2 | Shard proof (4, s) | Aggregation (1 deferred, then 2) (s) | A block (s) | GPU util. (%) | GPU peak (MiB) | Own miner (MH/s wall, n) | Against the card alone (4.1 s a block) | +|---|---|---|---|---|---|---|---| +| 22 (the default), phase D | 8.1, 7.8, 7.9, 7.8 | 8.4; 10.3, 10.0, 10.4 | 18.1 | 95.5 | 16,580 | 103.9 (9) | 4.4x | +| 20, E20 | 8.1, 7.8, 7.8, 7.8 | 8.4; 10.3, 10.1, 10.1 | 18.0 | 94.9 | 16,461 | 103.7 (8) | 4.4x | +| 18, E18 | 7.0, 6.7, 6.7, 6.7 | 7.2; 8.9, 8.8, 8.8 | 15.6 | 91.5 | 16,487 | 99.3 (7), minus 4.4% | 3.8x | +| 16, E16 | 5.1, 4.9, 4.9, 4.9 | 5.0; 6.1, 6.2, 6.2 | 11.1 | 85.3 | 16,519 | 83.8 (6), minus 19% | 2.7x | +| 16 again, phase H (the job's own choice: the shortest chain) | 5.0, 4.9, 4.8, 4.9 | 4.9; 6.1, 6.2, 6.2 | 11.1 | 85.7 | 16,487 | 84.0 (6) | 2.7x | + +Reading. Between 2^22 and 2^20 nothing moves: the card's time-slice scheduler alternates the two contexts whatever the kernel length above a few milliseconds. From 2^18 down the miner's launches get short enough (about 2 ms at 2^18, 0.5 ms at 2^16) that the prover's bursts find the card sooner, and the miner pays in launch overhead and idle gaps: at 2^16 the prover runs 1.6x faster (18.1 to 11.1 s a block, the chained aggregation 10.2 to 6.2 s) for a fifth of the hash rate, and it is still 2.7x slower than on a card to itself. The trade is about 1 MH/s per 0.37 s of block time at the 2^16 point, and the 3-s aggregation and the 1.5x slowdown are not reachable on a mining card by the kernel length; a 2^14 point (approximate, extrapolated) would be about 8 s a block at about 65 MH/s. The phase E0 (a 4-deferred aggregation under the miner) failed in 0.1 s: its proof paths pointed at `/` where job 1 had left its shard proofs, but job 2's block-344 proofs sit in job 2's own folder (`$JOB` was exported from job 2 on); the 4-deferred cost under the miner stays an estimate (lever 2 above). The app's own 5090 miner ran at 117 MH/s (job 3, 22:44Z) and 110 to 129 MH/s (its STATUS lines at 01:10Z) with the prover beside it, against my miner's 104 MH/s at the default batch: my miner runs the base variant with `--race off` (no tuning file on PC 2), so the curve's rates are relative to each other, not to the app's. + +### Lever 5, the host side under WSL2 (what the chain-mode numbers leave out) + +| What | Measured | +|---|---| +| The export (`igneum_exportSegments` 0..tip, 75 to 77 MB over curl.exe to a file on `C:`) | 1.1 to 1.5 s | +| The cut (`igneum-prove-export` replaying from genesis, then `--mode native`), four blocks | 18 s for four including the native checks (21:01:28Z to 21:01:46Z), about 4 s a block; the export's file sits on `/mnt/c` | +| The key setup per host process | 13.0 to 15.7 s on PC 2 (8.0 to 8.5 s on the Mac CPU): `--mode chain` and `--mode aggregate` pay it once per process, the app's loop pays it per shard | +| The proof file write through the WSL2 bridge | the 4 October entry ("shard proving on the RTX 5090"): 24 min of unbuffered `save` across `/mnt/c`, fixed by the 4 MB buffer; tonight `--save-shards` wrote the four 1.27 MB proofs inside the chain phase with no visible gap (the A phase's 80.4 s wall against 67.4 s of proving plus 13.0 s of setup) | +| Native Linux | not measured: no native Linux machine with an NVIDIA card exists in the project tonight, and the 4 October numbers were also WSL2 (Ubuntu 24.04 under PC 2's Windows). The WSL2 cost inside a `prove()` call is not separable from here; the host-side pieces above are what a native box would also skip or keep | + +### What went wrong, measured + +| What | Fixed | +|---|---| +| Job 1's per-phase command ran with `$JOB` empty (the bash variables of `vars.sh` were set, not exported, and the command runs in a child bash): `--out /results-A.json`, the saved shard proofs in `/` on the WSL root, so the aggregate-only phases B0, B, C1 and the prototype-shard phase C2 failed in 0.0 s ("No such file") | `export` in `vars.sh`; job 2 reads the proofs from `/` | +| Job 1's own-miner phases launched the iGPU miner (the first `igneum-miner mine` process matched; the 5090's is the second) and `if (StartMiner ...)` was always true (PowerShell: a function's emitted RESULT strings are part of its output), so D and E ran with the 5090 idle and the AMD iGPU at 3.4 MH/s: three more idle replicates of the chain (2.0 to 2.2 s shards, 1.8 and 2.2 s aggregations), no curve | the miner matched on `igneum-worker-cuda`, the outcome in a script-scope flag, `--race off` for the own miner (no tuning file on PC 2; a race costs up to 120 s a start) | +| Job 1's `/api/resume` at 21:25:11Z answered ok and the 5090 miner stayed off (card state `off`, hash 0.0, 1,760 MiB on the card) until the 0.3.10 restart; job 2 waited its full 600 s for a hash rate and ran its mining phases void | the restore job `tools/proving-v1/pc2-agg-cost-restore.ps1` also posts `/api/start`; the Counter ASIC coordinator opened a task chip for the resume defect | +| Jobs 3 and 4 (`agg-cost-pc2-3` 22:41Z on app 0.3.10, `agg-cost-pc2-4` 00:18Z on 0.3.11): every `/api/state` read came back as the two bytes `{}` (job 4's raw-body print: `raw_len=2`; the same reads gave the full state on 0.3.9 at 21:01Z and the AMD agent saw the empty reply at 22:22Z), so the card switch found no card, the app's 5090 miner kept mining, and job 3 ran a second miner beside it (two miners at about 60 MH/s each) while job 4's double-mining guard voided its own-miner phases. The class is the app's, not the reader's: `state_json()` (engine.rs:180) does `serde_json::to_value(st).unwrap_or(json!({}))`, and the value that fails is `ProvingState.paid_wei: u128` (serde_json 1.0.151 refuses a u128 over u64::MAX, 18.45 IGN; the proving-v1 agent's diagnosis): a paid shard averages 1.23 IGN, so the reply empties about 15 paid shards after every app start and comes back at the next restart, which matches the times (full at 21:01Z with paid_wei 0, empty from 22:22Z after the prover had paid from 22:02Z). Fixed on the app branch proving-v1 at 6714a45 (paid_wei as a decimal string, the error logged, an `{"error":...}` reply on any future failure) | job 5 reads the card keys from the app's `settings.json` (`cards`: key to enabled and identities), restores the 5090's 8 identities first (the restore job of 23:16:53Z had set 2: its parser read the next card's value), refuses before any pause when it cannot name the card, waits on the CUDA worker process count for the card to stop, and checks the worker is back at the end | +| Job 5 (`agg-cost-pc2-5`, 01:10:44Z) failed at PowerShell's parse in 1 s: `$RestoreIdentities:` inside a double-quoted string (a drive-qualified variable); no card or miner touched | `${RestoreIdentities}:`; the other `$name:` shapes are inside single-quoted bash here-strings | +| Job 6's identities step found `settings.json` already at 8 identities under the active key `nvidia:0:NVIDIA GeForce RTX 5090` (a stale key `nvidia:NVIDIA GeForce RTX 5090` carries 2), so no change was sent; job 6's `/api/cards` with the 5090 disabled answered ok but the worker ran on for 120 s, `/api/pause` stopped it in 5 s, and at the end `/api/resume` brought it back in 5 s on 0.3.11 | the card switch keeps the pause as its fallback; the resume path works on 0.3.11 | +| PC 2 ran three jobs at once from 01:24Z (`run-prover-on-pc2-20261006` at 01:24:21Z, the ledger suites build at 01:26:15Z, while agg-cost-pc2-6's closing report was still being uploaded): the app does not serialise jobs, "one job per machine at a time" holds only by the coordinator's word; job 6 had closed at 01:24:14Z, so its rows are clean | nothing of mine to fix; a rule for the job runner | +| The make-package gate ran the exporter's side files (`block-N.json.node-plan.json`) as fixtures and failed; its execute step took the exclusive `measure` lock for a cycle count and queued 25 min behind a packbench run | the glob skips `.node-plan.json`; the execute step runs under the `run` lock (a count, not a time) | diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index bedff564..50b64d14 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -154,3 +154,22 @@ The aggregation-cost agent's jobs read the two bytes `{}` from `/api/state` on P What it means: the dashboard on a proving machine shows nothing within about 12 minutes of its prover's first payout; every PC playbook that reads a card from `/api/state` fails the same way (the agent's job 5 reads settings.json instead). Mining, proving and payouts are untouched; it is the status page only. Fix on the app branch: `paid_wei` serialises as a decimal string (the dashboard already reads it with `Number()`), `state_json` logs `[error] state_json: ...` once instead of answering `{}`, and the reply on any future serialisation error carries `error` and `version` rather than nothing; unit test `a_paid_total_over_u64_max_still_serialises_the_whole_state`. Not in 0.3.11 (that tree closed at 22c2363, master 630da6b, published); 6714a45 heads 0.3.12, the morning's first cut, app only, before PC 1's relaunch (coordinator, counter-asic-2-rollout.md 8a); until then the workaround is settings.json for the card keys. + +## Aggregation cost (5 October, night) + +the project lead, 5 October 2026: "fix everything else in the numbers tonight". Branch `agg-cost`; every number in `docs/bench-log.md`, "aggregation cost on the RTX 5090", with its job id. The proof statement and the pinned guests are unchanged: every existing fixture proof still verifies (`verify-segment` 0.027 s on the Mac, 0.036 to 0.041 s on PC 2). + +| What | Before (5 October evening, `chain-pc2-pv1c`) | After (5 October night) | +|---|---|---| +| Chained aggregation, the card mining | 9.6 to 9.7 s a block | 9.6 to 9.8 s a block, the same (job `agg-cost-pc2-1`, phase A); the miner's presence is the whole cost: 2.1 to 2.2 s a block with the card to itself, 1.7 s unchained | +| Shard proof (empty shard), the card mining | 7.3 to 7.7 s | 7.4 to 7.8 s; 1.9 to 2.2 s with the card to itself | +| The miner's slowdown of the prover | 3 to 4x (against 4 October) | measured on the same fixtures 2 min apart: shards 3.6x, chained aggregation 4.5x, a block 4.2x | +| Where the time goes | not profiled | the host's share 0.000 s (the prove call is everything); the GPU server prints no timings; the second deferred proof (the chain rule) costs 0.4 to 0.5 s alone and 1.7 to 1.9 s under the miner; the prover alone keeps the card busy 15.8% of the time, the miner 93.9% | +| SP1 knobs (`SP1_WORKER_VERIFY_INTERMEDIATES=false`) | not tried | no gain: 7.8 s against 8.1 s over four aggregations, inside the spread; the shape knobs would change the recursion keys the pinned verifier accepts | +| Batch fold (K blocks in one aggregator call) | not estimated | estimate from the measured step costs: 0.9 s a block alone and 3.8 s mining at K = 4, 0.7 and 2.8 s at K = 8 (0.25 s per further deferred proof alone, 1.8 s mining); a tree fold gains nothing. A new pinned guest and program id either way, so not tonight | +| Two prover processes on one card | not tried | closed on SP1 6.8.1: both share one GPU server socket, run slower together (6.9 s a block against 4.1) and the second dies with the first (`early eof`) | +| The miner's kernel length (`--batch-log2` of the CUDA worker, 2^B nonces a launch; job `agg-cost-pc2-6`, the 5090 alone with the job's own miner) | not tried | 2^22 (the default) and 2^20: 18.1 and 18.0 s a block, 10.0 to 10.4 s a chained aggregation, 104 MH/s; 2^18: 15.6 s, 8.8 s, 99 MH/s (minus 4%); 2^16: 11.1 s, 6.1 to 6.2 s, 84 MH/s (minus 19%), reproduced | +| The GPU time-slice policy (`nvidia-smi compute-policy --set-timeslice`) | not tried | "Not Supported" on PC 2 (driver 13.3, Windows): closed | +| The chosen combination | the defaults | the defaults stay: batch-log2 22 and SP1's default knobs. The one knob that moves the prover (2^16) costs a fifth of the hash rate all the time for a prover that is busy a few seconds a minute on the devnet; it is the project lead's trade, not a default (below) | + +Reading. The per-block aggregation is 2.1 s and a block 4.1 s on a 5090 that only proves, 9.7 and 17.5 s on one that also mines; no knob, fold or stream on tonight's SP1 changes the first pair, and only the miner's kernel length changes the second, at 1 MH/s per 0.37 s of block time. So "under 3 s a block" and "under 1.5x" are met on a card that is not mining and are not reachable on one that is. What that means per tier: a 5090 that mines and proves delivers a proven empty block every 17.5 s (6 cards for 1 block/s), the same card proving only every 4.1 s (2 cards, plus the shard work of full blocks: the fleet table above), and a batch fold of the aggregator (a new pinned guest) would bring the proving-only card to about 2.7 s a block and the mining one to about 12 s. What is being done: the app and host defaults are left as measured; the plan's open decision for the project lead is whether a card that holds a shard assignment should drop to 2^16 for the proof's minute (1.6x faster proof, 19% of its hash rate for that minute) or whether proving-only cards carry the aggregation (the clean 2.1 s), and the batch fold goes on the next pin's list. The state class found on the way (`/api/state` answering `{}` once `paid_wei` passes u64::MAX, fixed on the app branch at 6714a45) is in the bench log with the rest. diff --git a/proving/igneum-prove/host/src/main.rs b/proving/igneum-prove/host/src/main.rs index d7ed375b..005182c2 100644 --- a/proving/igneum-prove/host/src/main.rs +++ b/proving/igneum-prove/host/src/main.rs @@ -93,9 +93,12 @@ fn run() -> Result<()> { // proof (the chain rule of design 5.3), the measurement of docs/plans/proving-v1.md step 2 let list = arg("--chain").or_else(|| args.get(1).filter(|a| !a.starts_with("--")).cloned()).context("--chain (consecutive fixtures)")?; let fixtures: Vec = list.split(',').map(|s| s.trim().to_string()).filter(|s| !s.is_empty()).collect(); - return run_chain(&pinned, &fixtures, prover, out_path.as_deref()); + // --save-shards writes every shard's compressed proof next to the results (block-N-shard-i-compressed.bin), + // so `--mode aggregate` can re-run the aggregation of the same proofs under other settings + let save_shards = args.iter().any(|a| a == "--save-shards"); + return run_chain(&pinned, &fixtures, prover, out_path.as_deref(), save_shards); } - let path = args.get(1).filter(|a| !a.starts_with("--")).context("usage: igneum-prove-host [--mode native|execute|shard|compressed|block|all] [--shard N] [--budget ] [--prover 0x..] [--out results.json]; --mode chain --chain [--prover 0x..] [--out results.json]; --mode aggregate --proofs --parent 0x.. [--prev prev.bin] [--out results.json]; --mode verify --proof --statement 0x..; --mode verify-segment --proof --statement 0x..; --mode id")?; + let path = args.get(1).filter(|a| !a.starts_with("--")).context("usage: igneum-prove-host [--mode native|execute|shard|compressed|block|all] [--shard N] [--budget ] [--prover 0x..] [--out results.json]; --mode chain --chain [--prover 0x..] [--out results.json] [--save-shards]; --mode aggregate --proofs --parent 0x.. [--prev prev.bin] [--out results.json]; --mode verify --proof --statement 0x..; --mode verify-segment --proof --statement 0x..; --mode id")?; let shard_index: usize = arg("--shard").map(|s| s.parse()).transpose()?.unwrap_or(0); // `--budget `: re-plan the fixture's block at a TEST budget (the S_p curve of 5 October 2026); the fixture's // own per-shard plan is then not compared (the chain and the sums still are), and `--out` records the cut @@ -577,6 +580,8 @@ fn setup_sp1(pinned: &pinned::Pinned, results: &mut serde_json::Map) -> std::path::PathBuf { /// `--mode chain`: every fixture in order, consecutive on the chain (number and parent hash), each block's shards /// proven compressed and aggregated with the previous block's aggregated proof (`AggInput.prev`, the chain rule), /// every proof verified. One RESULT line per shard, per block (with the running totals) and for the chain. -fn run_chain(pinned: &pinned::Pinned, fixtures: &[String], prover: Address, out_path: Option<&str>) -> Result<()> { +fn run_chain(pinned: &pinned::Pinned, fixtures: &[String], prover: Address, out_path: Option<&str>, save_shards: bool) -> Result<()> { if fixtures.is_empty() { bail!("--chain needs at least one fixture"); } @@ -660,6 +665,7 @@ fn run_chain(pinned: &pinned::Pinned, fixtures: &[String], prover: Address, out_ let mut blocks_json = Vec::with_capacity(built.len()); let (mut shard_total, mut agg_total, mut shards_total) = (0.0f64, 0.0f64, 0usize); let out_dir = out_dir_of(out_path); + let mut shard_files: Vec = Vec::new(); for (number, _hash, shards, _) in &built { let block_t = Instant::now(); let mut proofs: Vec = Vec::with_capacity(shards.len()); @@ -676,12 +682,18 @@ fn run_chain(pinned: &pinned::Pinned, fixtures: &[String], prover: Address, out_ } shard_secs.push(dt); shard_total += dt; + if save_shards { + let file = out_dir.join(format!("block-{number}-shard-{i}-compressed.bin")); + save_proof(&p.proof, file.clone()); + shard_files.push(file.display().to_string()); + } proofs.push(p); } shards_total += proofs.len(); stage(&format!("chain block {number} aggregate {} shards{}", proofs.len(), if prev.is_some() { " with the previous block proof" } else { "" })); let seg = sp1.aggregate(prev.as_ref(), &proofs)?; let adt = sp1.last_timing("aggregate").unwrap_or_default().as_secs_f64(); + let sdt = sp1.last_timing("aggregate-stdin").unwrap_or_default().as_secs_f64(); agg_total += adt; let claim = SegmentClaim::from_block(&seg.output); let ok = sp1.verify_segment(&seg, &claim); @@ -690,9 +702,10 @@ fn run_chain(pinned: &pinned::Pinned, fixtures: &[String], prover: Address, out_ let block_s = block_t.elapsed().as_secs_f64(); let cumulative = chain_t.elapsed().as_secs_f64(); println!( - "RESULT chain block {number}: {} shards ({:.1} s of shard proofs), aggregate prove {adt:.1} s, proof {bytes} bytes, verify {vdt:.3} s, {}; chain_len {}, agg_vk {}; this block {block_s:.1} s, cumulative {cumulative:.1} s over {} block(s) at {}", + "RESULT chain block {number}: {} shards ({:.1} s of shard proofs), aggregate prove {adt:.1} s (stdin {sdt:.3} s, {} deferred proofs), proof {bytes} bytes, verify {vdt:.3} s, {}; chain_len {}, agg_vk {}; this block {block_s:.1} s, cumulative {cumulative:.1} s over {} block(s) at {}", seg.output.shard_count, shard_secs.iter().sum::(), + proofs.len() + usize::from(prev.is_some()), if ok { "VERIFIED (shard program id, aggregator id and claim checked)" } else { "VERIFY FAILED" }, seg.output.chain_len, seg.output.agg_vk, @@ -707,7 +720,7 @@ fn run_chain(pinned: &pinned::Pinned, fixtures: &[String], prover: Address, out_ bail!("block {number}: chain_len {} is not {expected_len}", seg.output.chain_len); } blocks_json.push(serde_json::json!({ - "number": number, "shards": seg.output.shard_count, "shard_prove_seconds": shard_secs, "aggregate_prove_seconds": adt, + "number": number, "shards": seg.output.shard_count, "shard_prove_seconds": shard_secs, "aggregate_prove_seconds": adt, "aggregate_stdin_seconds": sdt, "aggregate_verify_seconds": vdt, "proof_bytes": bytes, "chain_len": seg.output.chain_len, "block_seconds": block_s, "cumulative_seconds": cumulative, "post_root": seg.output.post_root.to_string(), "statement": alloy_primitives::keccak256(seg.output.to_bytes()).to_string(), })); @@ -732,6 +745,7 @@ fn run_chain(pinned: &pinned::Pinned, fixtures: &[String], prover: Address, out_ results.insert("shard_prove_seconds_total".into(), shard_total.into()); results.insert("aggregate_prove_seconds_total".into(), agg_total.into()); results.insert("chain_seconds".into(), total.into()); + results.insert("shard_proof_files".into(), shard_files.into_iter().map(serde_json::Value::from).collect::>().into()); drop(sp1); finish(results, out_path.map(|s| s.to_string())) } @@ -794,16 +808,18 @@ fn run_aggregate(pinned: &pinned::Pinned, proofs: &str, parent: &str, prev_path: for shards in &blocks { let number = shards[0].output.number; stage(&format!("aggregate block {number}, {} shards", shards.len())); + let deferred = shards.len() + usize::from(prev.is_some()); let seg = sp1.aggregate(prev.as_ref(), shards)?; let adt = sp1.last_timing("aggregate").unwrap_or_default().as_secs_f64(); + let sdt = sp1.last_timing("aggregate-stdin").unwrap_or_default().as_secs_f64(); let claim = SegmentClaim::from_block(&seg.output); let ok = sp1.verify_segment(&seg, &claim); let vdt = sp1.last_timing("verify-block").unwrap_or_default().as_secs_f64(); - println!("RESULT aggregate block {number}: {} shards, prove {adt:.1} s, proof {} bytes, verify {vdt:.3} s, {}; chain_len {}, post {} at {}", seg.output.shard_count, bincode::serialize(&seg.proof)?.len(), if ok { "VERIFIED" } else { "VERIFY FAILED" }, seg.output.chain_len, seg.output.post_root, now()); + println!("RESULT aggregate block {number}: {} shards, prove {adt:.1} s (stdin {sdt:.3} s, {deferred} deferred proofs), proof {} bytes, verify {vdt:.3} s, {}; chain_len {}, post {} at {}", seg.output.shard_count, bincode::serialize(&seg.proof)?.len(), if ok { "VERIFIED" } else { "VERIFY FAILED" }, seg.output.chain_len, seg.output.post_root, now()); if !ok { bail!("the aggregated proof of block {number} did not verify"); } - per_block.push(serde_json::json!({ "number": number, "shards": seg.output.shard_count, "aggregate_prove_seconds": adt, "aggregate_verify_seconds": vdt, "chain_len": seg.output.chain_len })); + per_block.push(serde_json::json!({ "number": number, "shards": seg.output.shard_count, "aggregate_prove_seconds": adt, "aggregate_stdin_seconds": sdt, "deferred_proofs": deferred, "aggregate_verify_seconds": vdt, "chain_len": seg.output.chain_len })); prev = Some(seg); } let seg = prev.unwrap(); diff --git a/proving/igneum-prove/host/src/proof_system.rs b/proving/igneum-prove/host/src/proof_system.rs index 840fae0f..bfc60e2e 100644 --- a/proving/igneum-prove/host/src/proof_system.rs +++ b/proving/igneum-prove/host/src/proof_system.rs @@ -259,6 +259,16 @@ impl Sp1ProofSystem { self.timings.lock().unwrap().push((what.to_string(), dt)); } + /// The SP1 prover knobs set in this process's environment (the GPU server inherits them; the names from + /// sp1-core-executor 6.8.1 `opts.rs` and sp1-prover 6.8.1 `worker/config.rs`), for the RESULT lines, so a + /// measurement names the settings it ran under. "none" when the defaults apply. + pub fn env_knobs() -> String { + let fixed = ["SHARD_SIZE", "ELEMENT_THRESHOLD", "HEIGHT_THRESHOLD", "FULL_SIZE_SHARDS", "MINIMAL_TRACE_CHUNK_THRESHOLD", "TRACE_CHUNK_SLOTS", "MEMORY_LIMIT", "WITHOUT_VK_VERIFICATION", "RUST_LOG"]; + let mut out: Vec = std::env::vars().filter(|(k, _)| k.starts_with("SP1_WORKER_") || fixed.contains(&k.as_str())).map(|(k, v)| format!("{k}={v}")).collect(); + out.sort(); + if out.is_empty() { "none".into() } else { out.join(" ") } + } + pub fn last_timing(&self, what: &str) -> Option { self.timings.lock().unwrap().iter().rev().find(|(k, _)| k == what).map(|(_, d)| *d) } @@ -286,6 +296,9 @@ impl ProofSystem for Sp1ProofSystem { /// The aggregator guest over the shard proofs (and the previous segment's proof when given), by recursion. fn aggregate(&self, prev: Option<&Sp1SegmentProof>, shards: &[Sp1ShardProof]) -> Result { let first = shards.first().ok_or_else(|| anyhow!("no shards"))?; + // 5 October 2026 (aggregation cost): the stdin build (the proof clones into the request) is timed apart + // from the prove call, so the host's own share of an aggregation is visible next to the GPU's. + let t_stdin = Instant::now(); let mut stdin = SP1Stdin::new(); let input = AggInput { shard_vk: self.shard_vk_hash(), @@ -302,6 +315,7 @@ impl ProofSystem for Sp1ProofSystem { let SP1Proof::Compressed(proof) = p.proof.proof.clone() else { return Err(anyhow!("the previous block proof is not a compressed proof")) }; stdin.write_proof(*proof, self.agg_vk.vk.clone()); } + self.record("aggregate-stdin", t_stdin.elapsed()); let t = Instant::now(); let proof = self.client.prove(&self.agg_pk, stdin).compressed().run()?; self.record("aggregate", t.elapsed()); diff --git a/proving/windows-wsl2/make-package.sh b/proving/windows-wsl2/make-package.sh index 4cb16deb..7298dde3 100755 --- a/proving/windows-wsl2/make-package.sh +++ b/proving/windows-wsl2/make-package.sh @@ -34,9 +34,13 @@ if [ "${SKIP_GATE:-0}" != "1" ]; then fi H="$ROOT/proving/igneum-prove/target/release/igneum-prove-host" for f in "$ROOT"/proving/fixtures/block-*.json; do + # the exporter's side files (block-N.json.node-plan.json, 5 October 2026) are not fixtures + case "$f" in *.node-plan.json) continue ;; esac if ! "$H" "$f" --mode native >>"$GATE_LOG" 2>&1; then echo "GATE FAILED: native run of $(basename "$f"); see $GATE_LOG"; exit 1; fi done - if ! "$ROOT/tools/lock/with-lock.sh" measure "$H" "$ROOT/proving/fixtures/block-338-shard1.json" --mode execute --shard 0 >>"$GATE_LOG" 2>&1; then + # the execute step reports a cycle count, not a time: the `run` lock (tools/lock/with-lock.sh: counts, not ms), so the + # gate does not queue behind every build and measurement on the Mac (5 October 2026: 25 min behind a packbench run) + if ! "$ROOT/tools/lock/with-lock.sh" run "$H" "$ROOT/proving/fixtures/block-338-shard1.json" --mode execute --shard 0 >>"$GATE_LOG" 2>&1; then echo "GATE FAILED: the guest did not execute the shard fixture (the exact failure the PC hit on 4 October); see $GATE_LOG"; exit 1 fi "$H" --mode id | tee -a "$GATE_LOG" # the pinned program ids this package carries (the PC's build embeds the same elf/ files) diff --git a/tools/proving-v1/agg-cost-table.mjs b/tools/proving-v1/agg-cost-table.mjs new file mode 100644 index 00000000..e1eefca3 --- /dev/null +++ b/tools/proving-v1/agg-cost-table.mjs @@ -0,0 +1,50 @@ +#!/usr/bin/env node +// The aggregation-cost curve from one or more pc2-agg-cost.ps1 jobs (5 October 2026 night): per phase, the shard and +// aggregation seconds per block (the chain and aggregate-only RESULT lines), the GPU utilisation and memory peak of +// the phase, the own miner's rate (MH/s wall) and its batch-log2. Reads the job uploads from the log intake like +// tools/jobs.mjs (DATABASE_URL in ~/.config/igneum/env). Prints a markdown table. +// node tools/proving-v1/agg-cost-table.mjs agg-cost-pc2-1 agg-cost-pc2-2 ... +import { readFileSync } from 'node:fs'; +import { homedir } from 'node:os'; +const ids = process.argv.slice(2); +if (!ids.length) { console.error('usage: agg-cost-table.mjs [...]'); process.exit(2); } +const m = /^DATABASE_URL=(.*)$/m.exec(readFileSync(`${homedir()}/.config/igneum/env`, 'utf8')); +const url = m[1].trim().replace(/^['"]|['"]$/g, ''); +const host = new URL(url).hostname.replace('-pooler', ''); +const sql = async (query, params = []) => { + const r = await fetch(`https://${host}/sql`, { method: 'POST', headers: { 'Neon-Connection-String': url, 'Content-Type': 'application/json' }, body: JSON.stringify({ query, params }) }); + const j = await r.json(); + if (!r.ok) throw new Error(j.message || JSON.stringify(j)); + return j.rows; +}; +const f1 = x => (Math.round(x * 10) / 10).toFixed(1); +const rows = []; +for (const id of ids) { + const ups = await sql('SELECT lines FROM miner_logs WHERE run_id LIKE $1 ORDER BY received_at DESC LIMIT 1', [`job-${id}-%`]); + if (!ups.length) { console.error(`${id}: no upload`); continue; } + const lines = ups[0].lines.split('\n'); + const phases = new Map(); + const ph = l => { if (!phases.has(l)) phases.set(l, { job: id, label: l, shard: [], agg: [], deferred: [], util: '', mem: '', rate: '', batch: '', note: '' }); return phases.get(l); }; + for (const raw of lines) { + const l = raw.replace(/^\d+(\.\d+)? job \S+: /, ''); + let x; + if ((x = /^([A-Z0-9]+): RESULT chain block \d+ shard \d+: compressed prove ([\d.]+) s/.exec(l))) ph(x[1]).shard.push(Number(x[2])); + else if ((x = /^([A-Z0-9]+): RESULT (?:chain|aggregate) block \d+: .*?(?:aggregate )?prove ([\d.]+) s \(stdin [\d.]+ s, (\d+) deferred/.exec(l))) { ph(x[1]).agg.push(Number(x[2])); ph(x[1]).deferred.push(Number(x[3])); } + else if ((x = /^RESULT phase_gpu ([A-Z0-9]+) samples=\d+ memory_used_max_mib=(\d+) util_mean_pct=([\d.]+)/.exec(l))) { ph(x[1]).mem = x[2]; ph(x[1]).util = x[3]; } + else if ((x = /^RESULT ([A-Z0-9]+) miner rate (n=\d+ mean=[\d.]+)/.exec(l))) ph(x[1]).rate = x[2].replace('n=', 'n ').replace(' mean=', ', mean '); + else if ((x = /^RESULT ([A-Z0-9]+) own miner started .* batch_log2=(\d+)/.exec(l))) ph(x[1]).batch = x[2]; + else if ((x = /^RESULT ([A-Z0-9]+) own miner (not hashing|EXITED|: no app miner)/.exec(l))) ph(x[1]).note = 'own miner ' + x[2]; + else if ((x = /^RESULT phase ([A-Z0-9]+) end .* exit (\d+) wall ([\d.]+) s/.exec(l))) { const p = ph(x[1]); p.exit = x[2]; p.wall = x[3]; } + else if ((x = /^RESULT phase ([A-Z0-9]+) skipped/.exec(l))) ph(x[1]).note = 'skipped'; + else if ((x = /^RESULT (H (?:choice|knobs)): (.*)$/.exec(l))) console.error(`${id}: ${x[1]}: ${x[2]}`); + } + for (const p of phases.values()) rows.push(p); +} +const mean = a => a.length ? a.reduce((s, v) => s + v, 0) / a.length : NaN; +console.log('| Job | Phase | batch-log2 | Shard s (each) | Aggregation s (each, deferred proofs) | Block s (shard + aggregation, mean) | GPU util % | GPU peak MiB | Miner MH/s | Note |'); +console.log('|---|---|---|---|---|---|---|---|---|---|'); +for (const p of rows) { + const chained = p.agg.filter((_, i) => p.deferred[i] >= 2); + const block = p.shard.length && p.agg.length ? f1(mean(p.shard) + mean(chained.length ? chained : p.agg)) : ''; + console.log(`| ${p.job} | ${p.label} | ${p.batch || (p.label.startsWith('E') && /^E\d+$/.test(p.label) ? p.label.slice(1) : '')} | ${p.shard.map(f1).join(', ')} | ${p.agg.map((a, i) => `${f1(a)} (${p.deferred[i]})`).join(', ')} | ${block} | ${p.util} | ${p.mem} | ${p.rate} | ${[p.note, p.exit && p.exit !== '0' ? `exit ${p.exit}` : ''].filter(Boolean).join('; ')} |`); +} diff --git a/tools/proving-v1/miner-rate.mjs b/tools/proving-v1/miner-rate.mjs new file mode 100644 index 00000000..051c919c --- /dev/null +++ b/tools/proving-v1/miner-rate.mjs @@ -0,0 +1,35 @@ +#!/usr/bin/env node +// The hash rate of one miner over a UTC window, from its STATUS lines in the log intake (the app uploads the miner's +// log every minute; each STATUS line starts with a unix timestamp and carries `now= MH/s wall`, the last +// interval's rate). For the aggregation-cost phases that keep the app's own miner running (docs/bench-log.md, +// 5 October 2026 night), where the job cannot read the app's state from PowerShell 5.1. +// node tools/proving-v1/miner-rate.mjs