Proving v1: step 1 measured with the fleet mining (the prover costs the 5090 4.0%, GPU peak 15.6 GB with the miner), the fleet coverage window (2.4%, p50 44 s)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 19:53:31 +00:00
parent ae3b681b3d
commit 251ea75b23
2 changed files with 10 additions and 5 deletions

View file

@ -1595,8 +1595,10 @@ Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `v
| What | Measured |
|---|---|
| The default rule (`app/igneum-app/src/provedefault.rs`) | `cargo test --release -p igneum-app provedefault` on this Mac (the app crate, build lock, 19:05Z): 5 passed (a 5090 with WSL2 on Windows is on; Windows without WSL2 off with the Set up hint; Linux needs no WSL2 and the 12 GB gate holds, a 10 GB 3080 and a 16 GB AMD card stay off; Apple silicon off; the biggest qualifying card is named) |
| Mining alone against mining with the prover (PC 2 job `prover-cost-pc2-pv1`, `tools/proving-v1/pc2-prover-cost.ps1`, 5 + 5 min, published 18:40:48Z, ran 18:41:13Z) | VOID: the job waited 20 min for the 5090 worker to hash and it never did (the pack fault above); "mining alone" was 0 MH/s. Re-run after the coordinator's go. The job's `/api/state` read also came back empty on PC 2 (`cards=::0`, every sample) while `igneum_getProvingStatus` on 127.0.0.1:26790 answered; the re-run script prints the URL file and the error |
| GPU memory during proving (the same job, phase B: the prover on for 5 min, 298 one-second `nvidia-smi --query-gpu=memory.used` samples, the 5090 worker dead so the card held nothing else) | memory.used min 1,654 MiB, max 13,816 MiB, utilisation mean 2.7%, power max 190.6 W. The shards were the devnet's empty shards (0 pgas); the peak on a full shard at `S_p` is the chain job's row (held). Read: a 16 GB card clears tonight's peak, a 12 GB card does not, and the 24 GB of a 4090 has 10 GB of headroom |
| Mining alone against mining with the prover, first try (PC 2 job `prover-cost-pc2-pv1`, `tools/proving-v1/pc2-prover-cost.ps1`, published 18:40:48Z, ran 18:41:13Z) | VOID: the job waited 20 min for the 5090 worker to hash and it never did (the pack fault above); "mining alone" was 0 MH/s |
| Mining alone against mining with the prover, the re-run after the coordinator's go (job `prover-cost-pc2-pv1b`, ran 19:20:36Z to 19:51:17Z; the 5090 worker restored at 19:16Z and hashing throughout; prover OFF by `POST /api/prove {"on":false}` 19:40:39Z, back ON 19:46:09Z, left on). The job's own `/api/state` samples stayed empty on PC 2 (`Invoke-RestMethod` returns an object PowerShell 5.1 cannot walk, `cards=0`, the fix is for the next run), so the hash rate is read from the miner's own STATUS lines (`miner-nvidia-1ccfe586-1` uploads, `now=... MH/s wall`, one every 30 s, the intake table `miner_logs`) | prover OFF, 19:41:09 to 19:46:09Z: n 10, mean 124.72 MH/s, p50 124.81, min 124.10, max 125.38. Prover ON, 19:46:39 to 19:51:13Z: n 9, mean 119.74, p50 118.87, min 118.08, max 123.42. The 15 min before the job with the prover on (19:25 to 19:40Z): n 30, mean 119.88, p50 118.79. So the prover costs the 5090 5.0 MH/s, 4.0% of its hash rate, while it proves the devnet's empty shards one after another (1.4 a minute here: the node's paidShards 559 -> 566 over the 5-min phase). A full shard at `S_p` keeps the card busier (the 4 October run proved one in 10.9 s); the cost at that load is the chain job's row |
| GPU memory during proving, first try (phase B of the first job: the prover on for 5 min, 298 one-second `nvidia-smi --query-gpu=memory.used` samples, the 5090 worker dead so the card held nothing else) | memory.used min 1,654 MiB, max 13,816 MiB, utilisation mean 2.7%, power max 190.6 W: the prover alone on empty shards |
| GPU memory with the miner AND the prover on the card (the re-run's phase B, 298 one-second samples, 19:46 to 19:51Z) | memory.used min 3,396 MiB (the miner's dataset and program resident), max 15,590 MiB, utilisation mean 92.9%, power max 328.6 W. So the prover's own peak is about 12.2 GB on an empty shard (15,590 minus the miner's 3,396), and the two together need 15.6 GB: a 16 GB card (5080, 9070 XT class, if it had a CUDA path) sits 0.4 GB under tonight's peak with no room for a full shard, a 24 GB 4090 has 8.4 GB of headroom, a 12 GB card cannot mine and prove at once on this build. The full-shard peak is the chain job's row |
| Shards per minute with the mining worker dead | the node's `paidShards` 510 -> 518 over the 5-min phase: 1.6 shards a minute from one 5090 through the app's loop (export, cut, prove, sign, submit) |
| Host RAM (Windows `Win32_OperatingSystem` and the `vmmem` working set, sampled every 15 s) | host used 25,550 MB of 63,132 MB at the end; the WSL2 VM's working set 7,915 MB (2,334 MB used of 30,914 MB inside the distribution) |
| The SP1 GPU server's compiled targets (`cuobjdump --list-elf /root/.sp1/bin/sp1-gpu-server` inside PC 2's Ubuntu-24.04, CUDA 12.8, driver 610.47) | `sp1-gpu-server` 6.8.1 (251,306,680 bytes, sha256 c2642ad1c42e85d8525159cf0c7cd5200d8766c9be1283f452a1f9bf9fea725c, the asset `sp1_gpu_server_v6.8.1_x86_64.tar.gz` the SDK downloads, `sp1-cuda-6.8.1/src/server.rs`): one ELF each for sm_80, sm_86, sm_89, sm_90, sm_100 and sm_120; `strings` finds compute_120 PTX as well. So sm_89 (Ada: RTX 4090, 4080) is compiled in natively, no JIT; so are Ampere (3090, 3060), Hopper, Blackwell datacentre (sm_100) and consumer (sm_120, the 5090). Nothing for AMD (no HIP path in SP1) |
@ -1616,7 +1618,8 @@ Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `v
| What | Measured |
|---|---|
| A 3-minute window at 18:57Z on node 1 (`node tools/proving-v1/coverage.mjs --minutes 3`, chain blocks 80754..80839, 86 blocks) | 4 blocks with a paid shard (4.7%), 4 fully proven, 4 of 86 shards; on-chain latency (carrier timestamp minus block timestamp) n 4: min 36 s, p50 39 s, max 44 s; 0 content blocks. One prover (PC 2), the Mac verifier node down, PC 2 producing few blocks (3.4 MH/s): the degraded state above, not the fleet's number |
| The 30-minute window | waits for the fleet (the coordinator's go) |
| A 30-minute window, 19:13 to 19:43Z, the degraded fleet (PC 2 the only prover, its 5090 worker restored at 19:16Z, the Mac app node down by decision: the Mac app is attached to node 1) | `node tools/proving-v1/coverage.mjs --minutes 30 --watch` on node 1: chain blocks 81236..82668, 1,433 blocks; 38 with a paid shard (2.7%), all 38 fully proven (one shard a block, 0 content blocks); on-chain latency n 38: min 36, p50 44, p90 52, p99 62, max 65 s. The live page's 10-minute proving object read 0 shards and 0 provers at 19:42Z (it counts what its own node verified; that node is the Mac app node, down), so the chain's own count is the number |
| A 30-minute window with the fleet mining (PC 2 at 119 MH/s from 19:16Z, PC 1 at 128.8 from 19:18Z; PC 2 still the only prover, its prover OFF for the 5 min of the cost job's phase A inside this window; the Mac app node down by decision) | `coverage.mjs --minutes 30 --watch`, 19:21 to 19:51Z on node 1: chain blocks 81644..83069, 1,426 blocks; 34 with a paid shard (2.4%), all fully proven (one shard a block, no content); on-chain latency n 34: min 38, p50 44, p90 51, p99 52, max 53 s. One 5090 through the app's loop as it is covers 2.4 to 2.7% of the blocks; the latency from block to carried record is 44 s at the median, under the litepaper's minute, and would be the same for every block if the fleet were 40 cards (the table below) |
### Step 3, the fleet size (arithmetic from measured inputs; every input names its entry)

View file

@ -30,7 +30,9 @@ that say how many cards cover the chain.
| What | Number |
|---|---|
| GPU memory during proving on the 5090 (empty shards, 298 one-second samples) | min 1,654 MiB, max 13,816 MiB; so a 16 GB card clears it and a 12 GB card does not; the full-shard peak is the held chain job's row |
| The prover's cost to a mining 5090 (the re-run with the fleet mining, hash rate from the miner's own STATUS lines) | 124.7 MH/s alone, 119.7 MH/s with the prover on: 5.0 MH/s, 4.0%, on empty shards at 1.4 a minute |
| GPU memory on the 5090: the prover alone (empty shards) / the miner and the prover together | max 13,816 MiB / max 15,590 MiB (the miner holds 3,396 MiB); a 24 GB 4090 has 8.4 GB of headroom, a 16 GB card 0.4 GB, a 12 GB card cannot do both on this build; the full-shard peak is the chain job's row |
| Coverage, 30-min window with the fleet mining, one prover | 2.4% of blocks, latency p50 44 s, p99 52 s |
| Host RAM | host used 25.6 GB of 63 GB; the WSL2 VM 7.9 GB working set |
| sp1-gpu-server 6.8.1 compiled targets (cuobjdump) | sm_80, sm_86, sm_89, sm_90, sm_100, sm_120 and compute_120 PTX: Ada (4090) is native, no JIT; nothing for AMD |
| Shards a minute, one 5090 through the app's loop (empty shards) | 1.6 |
@ -39,7 +41,7 @@ that say how many cards cover the chain.
| The harness (3 nodes, fast time, trust mode) | PASSED, 21 checks in 197 s: paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90% |
| Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) | 4.7% of blocks proven, on-chain latency p50 39 s |
| 5090-class cards for 100% at 1 block/s | 40 with the loop as it is (empty blocks), 14 at one full shard a block, 45 at `B_p` (four full shards), about 8 at the adopted v1 budgets (approximate): the table in the bench log |
| HELD (coordinator, 19:00Z): mining alone against mining with the prover; the chain of 8 on the 5090 (N = 2, 4, 8); the 30-min coverage with the fleet mining | re-run after the go: jobs `prover-cost-pc2-pv1` (script fixed to print its state error) and `pc2-chain.ps1` with the package `igneum-prove-wsl2-pv1.zip` |
| The chain of 8 on the 5090 (N = 2, 4, 8), job `chain-pc2-pv1` (published 19:52Z after the go) | (the bench log row when it ends) |
## The rule, in one paragraph (spec 7.8)