From bb0453ba49f38f357e8301d141be40c3dea2a60a Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 12:01:54 +0000 Subject: [PATCH] bench-log: the real RTX 4070 alone (7,553 MiB compressed in 7.7 s at 2^26; core-only 5,761 MiB at 2^25): a 12 GB card proves alone, measured on the card Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 30 ++++++++++++++++++++++++++++++ 1 file changed, 30 insertions(+) diff --git a/docs/bench-log.md b/docs/bench-log.md index f0860286a..0faeb695a 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1657,6 +1657,36 @@ a pool rule, 20x the bytes a shard on the relay; nothing on chain. Every number the real 12 GB card (route 1, `tools/prover-floor/pc-12gb-card.ps1`, kit `igneum-prove-wsl2-floor.zip` e377c7df...) is the measurement that decides the public line. +Route 1, THE REAL 12 GB CARD (6 October 2026, the project lead's word at 11:25Z through main; an ASUS Dual RTX 4070 OC 12 GB, +12,282 MiB, nvidia-smi index 1, PCIe 4.0 x4 through a USB4 enclosure, mining beside the 5090 in PC 1). PC 1 already +had WSL2 Ubuntu-24.04 with nvcc 12.8 and cargo 1.99 (jobs `floor-pc1-wsl2-enable`, elevated, changed nothing, and +`floor-pc1-toolchain`, 11:21 to 11:23Z); the v4 server built cold in 353 s (`floor-pc1-build`, 11:25:33 to 11:31:34Z, +480 crates, `ed7b1e4b862c451b4a2cefa0e544f7b6a9e04279690c3f8c9881666bd7f8f23a`, sm_86/89/120, sm_89 native for the +4070); the kit `igneum-prove-wsl2-floor.zip` (e377c7df...) fetched (`floor-pc1-kit`, 11:32:49Z) and the host built +from it on PC 1 (adcad5ea..., the pinned ids matching). Every proof steered to the 4070 by `IGNEUM_CUDA_DEVICE=1` +(the host change: the SDK starts the server with `CUDA_VISIBLE_DEVICES=1`); `CUDA_DEVICE_ORDER=PCI_BUS_ID`. + +Alone (job `card12-alone`, 11:53:01 to 12:00:53Z, the miners stopped by the job, the prover off then on; the 4070's +own idle **0 MiB** read before and after, so every peak is the server's own; `--mode shard` = the core proof, its +CPU verify, then the compressed proof, the sampler split at the core RESULT; every proof VERIFIED): + +| Threshold | Fixture | Cycles | Core-only peak MiB | Full compressed peak MiB | Core s | Compressed s | Core proof bytes | CPU verify s | +|---|---|---|---|---|---|---|---|---| +| 2^26 | v1 shard | 4,717,439 | 7,265 | **7,553** | 3.7 | **7.7** | 14,379,043 | 0.432 | +| 2^26 | block-72854 (the "empty block first" fixture is NOT empty: a 5x bigger shard) | 23,594,943 | 7,041 | 7,617 | 17.2 | 35.9 | 69,847,451 | 2.204 | +| 2^27 | v1 shard | 4.7 M | 10,177 | 10,177 | 3.1 | 5.5 | 10,147,581 | 0.308 | +| 2^27 | block-72854 | 23.6 M | 10,017 | 10,241 | 12.9 | 22.8 | 35,910,731 | 1.120 | +| 2^25 | v1 shard | 4.7 M | **5,761** | 7,617 | 5.2 | 12.1 | 25,621,187 | 0.793 | +| 2^25 | block-72854 | 23.6 M | 5,761 | 7,649 | 31.0 | 74.3 | 164,004,701 | 5.248 | + +Reading. On the card itself a 12 GB card PROVES ALONE: the compressed v1 shard at 2^26 holds 7.55 GB of 12.28 GB +(4.7 GB spare) in 7.7 s, at 2^27 10.2 GB (2.1 GB spare) in 5.5 s, and a 23.6 M-cycle shard fits at 2^26 in 36 s. +The 5090's allocation pattern overstated the 4070 by about 0.65 GB (8.2 GB "own" against 7.55 measured). The +4070 proves the v1 shard 1.35x slower than the 5090 alone (7.7 against 5.7 s at 2^26) through a PCIe 4.0 x4 link. +Core-only at 2^25 is 5.76 GB on the card. Consequence for the public line: "12 GB proves (prove-only)" is now +measured on a 12 GB card, so the line moves from 24 GB once the patched server ships (the packaging row before +0.3.12); mine-and-prove is the beside-the-miner job's row (below, `card12-miner`). + Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1 patch takes the shard's term out (20.5 to 12.7 GB on the v1 shard) at a 1.26x time cost (5.3 s against 4.2 s,