Bench log: shard proving on the RTX 5090 measured (10.9 s per shard, 2.2 s aggregation); evidence rows 15 and 16; journey line

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-04 18:17:16 +00:00
parent b7442a549a
commit b1d702bcf2
6 changed files with 83 additions and 59 deletions

View file

@ -854,7 +854,7 @@ worker with jobs queued and no job done for 60 s), workers emit `need <epoch> <d
here: the swap time of the forced prepare on the PC; the OpenCL run's "2 jobs in 154 s" were the two jobs before the
first mismatch and are not a rate.
## 4 October 2026, shard proving on the RTX 5090 (skeleton: the GPU column fills from the first `shard-benchmark` job)
## 4 October 2026, shard proving on the RTX 5090: a full shard compressed in 10.9 s, a two-shard block aggregated in 2.2 s, all verified
Machine: the project lead's PC 2 (RTX 5090, 32,607 MiB; WSL2 Ubuntu 24.04, SP1 v6.8.1 with the cuda feature, the toolchain of the
4 October morning run in `~/igneum-prove`), the Igneum Miner app (machine id 1ccfe586) stopping its miners for the
@ -866,17 +866,37 @@ executor alone. `S_p` = 7,500,000 pgas provisional; the fixtures carry 6.75 M pg
| Stage | Apple M5 Max CPU (4 October, loaded) | RTX 5090 (job id, UTC) |
|---|---|---|
| shard at `S_p` (block-338-shard1, 6.75 M pgas): execute | 60.0 M cycles, 9 per pgas, 6.9 s (block-344 shard 0, the same size) | |
| shard at `S_p`: core proof (prove s, bytes, verify s) | not run at `S_p` (200-pgas shard: 83.1 s, 7,310,257 B, 0.368 s) | |
| shard at `S_p`: compressed proof (prove s, bytes, verify s) | not run at `S_p` (200-pgas shard: 272.3 s, 1,272,897 B, 0.075 s) | |
| two-shard block (block-341-shards2): compressed proof per shard | not run | |
| two-shard block: aggregation (prove s, bytes, verify s) | not run (three 200-pgas shards: 244.5 s, 1,272,909 B, 0.084 s) | |
| two-shard block: end to end, first shard proof to the verified block proof | not run (three 200-pgas shards: 1,139 s) | |
| four-shard block (block-344-shards4, near `B_p`): compressed proof per shard | not run | |
| four-shard block: aggregation (prove s, bytes, verify s) | not run (aggregator statement 1.66 M cycles) | |
| four-shard block: end to end | not run | |
| setup (one per program id) | 39.4 to 60.6 s (client plus two key setups) | |
| GPU idle wait before the run, package download and extract, build (incremental) | | |
| shard at `S_p` (block-338-shard1, 6.75 M pgas): execute | 60.0 M cycles, 9 per pgas, 6.9 s (block-344 shard 0, the same size) | 60,759,590 cycles, 9 per pgas, 44 per EVM gas, 1.63 s (run-20261004-173115, 17:40:20) |
| shard at `S_p`: core proof (prove s, bytes, verify s) | not run at `S_p` (200-pgas shard: 83.1 s, 7,310,257 B, 0.368 s) | 8.3 s, 18,116,295 B, 0.564 s, VERIFIED (run-20261004-173115, 17:40:29) |
| shard at `S_p`: compressed proof (prove s, bytes, verify s) | not run at `S_p` (200-pgas shard: 272.3 s, 1,272,897 B, 0.075 s) | 10.9 s, 1,272,897 B, 0.040 s, VERIFIED (run-20261004-173115, 18:04:44) |
| two-shard block (block-341-shards2): compressed proof per shard | not run | 11.7 s and 10.0 s, 1,272,897 B each, verify 0.039 and 0.038 s (run-20261004-173115, 18:06:50 and 18:07:00) |
| two-shard block: aggregation (prove s, bytes, verify s) | not run (three 200-pgas shards: 244.5 s, 1,272,909 B, 0.084 s) | 2.2 s, 1,272,909 B, 0.039 s, VERIFIED, shard program id and claim checked (run-20261004-173115) |
| two-shard block: end to end, first shard proof to the verified block proof | not run (three 200-pgas shards: 1,139 s) | 24 s of GPU stages (setup 12.6 s, two compressed proofs, aggregation); 2 min 18 s wall with the proof saves (run-20261004-173115, 18:06:26 to 18:08:44) |
| four-shard block (block-344-shards4, near `B_p`): compressed proof per shard | not run | not run: the script read only its second argument (fixed 18:10 UTC); next job |
| four-shard block: aggregation (prove s, bytes, verify s) | not run (aggregator statement 1.66 M cycles) | next job |
| four-shard block: end to end | not run | next job |
| setup (one per program id) | 39.4 to 60.6 s (client plus two key setups) | 21.3 s first process (client 6.7, shard keys 14.6, aggregator keys 0.03); 12.6 s second process |
| GPU idle wait before the run, package download and extract, build (incremental) | | miners stopped 17:38:56; build 58 s (sources already compiled once); job 1,789 s wall, of which 25 min 46 s was saving proofs (below); mining resumed by itself at 127 MH/s |
Job run-20261004-173115 (the `run` kind, prove only, as root inside the app's own WSL2 instance), log intake run
`job-run-20261004-173115-1ccfe586`, exit 0 after 1,789 s. Fixtures captured from the devnet: block 338 (one shard at
S_p, 11 transactions) and block 341 (two shards, 14 transactions). Every proof verified on the PC; the three tampered
witnesses per fixture (balance, storage or code, dropped account) were rejected before any proving.
What the numbers say. One RTX 5090 turns a full shard into the 1.27 MB compressed proof the chain carries in about
11 s, and folds a block's shards into one proof in about 2 s more. Against the launch target of 20 to 60 s behind the
tip, a single card has 9 s of slack on a one-shard block; a two-shard block needs two cards or two rounds. The 44
cycles per EVM gas and 9 cycles per prover gas are the first measured constants for the prover-gas schedule (spec 7,
provisional S_p).
What went wrong, measured. The job was silent for 24 min 4 s between the core proof (17:40:29 UTC) and the
compressed stage (18:04:33 UTC), and 1 min 42 s after the compressed proof: SP1's `save` writes a proof straight into
an unbuffered file, and with the results folder under `/mnt/c` every field element was one round trip across the WSL2
file bridge. The host now saves through a 4 MB buffer and prints a timed `saved` line, and the script keeps results
on the Linux side and copies them once per stage (ledger P20). The four-shard fixture (block 344) was not run: the
script read only its second argument, also fixed. Both fixes are in the package rebuilt at 18:10 UTC; the next job
measures them.
Reading: pending the run. What it decides: the first `S_p` point (the shard stage time on the card against the
20 to 60 s proof lag of the litepaper), whether the P20 Drop fix holds on the GPU path (exit 0, no abort after the

View file

@ -39,8 +39,8 @@ Versions in the table: `igneum-pow` is the Rust crate at `igneum-pow/Cargo.toml`
| 12 | The difficulty rule recovers from a hashrate step within minutes, where Kaspa's sampled rule never settles. A step inside an epoch set the rule oscillating on the live devnet on 4 October 2026; rule v2 removes it in the simulator and on a test network and is built but not yet rolled out | Spec 2.3; litepaper Speed (implied); bench page | tested by the team | repo `e9328c6`, `abb5a5d` (attacks), `67bf226` (rule v2); fork `difficulty` branch (timestamp fix) and `devnet-v4` `a21ff239` (`difficulty_v2_activation_daa`, `REF_WINDOW_V2 = 600`); `sim/difficulty/sim.py --live` | The live record `sim/difficulty/records/live-2026-10-04.csv` (8,090 headers, `pull_live.py`) and the hash-rate record beside it; `sim/difficulty/sim.py` on the synthetic set and the DAG replay; `sim/difficulty/attacks/attacks.py`; `sim/difficulty/testnet_v2.py` (3 nodes, activation at DAA 900); `cargo test --release -p kaspa-consensus --lib difficulty` (15 pass); bench-log "difficulty controller", "difficulty rule under attack", "timestamp attack fixed", "difficulty rule v2" | Live devnet v4, 4 October 2026 (UTC): a second RTX 5090 joining 7 minutes into an epoch (about 152 to 280 MH/s) hardened the difficulty 70M to 144M in 90 s and then swung by about a third for 40 minutes around the true level of 139M while the epoch-long reference lane carried the join; that card leaving for 4 minutes eased 116M to 67M and back to 106M; the epoch boundary with both PCs restarting took 152M to 77M in 3 minutes, after which the rule held within 1.3% per minute with no flips. Cause: the reference lane covered the whole epoch, so a mid-epoch step polluted it for the hour and the 25% trigger flipped on the short lane's noise. The DAG replay reproduces the record (std of log difficulty 0.115 against 0.134, 4.3 peaks against 4). Rule v2 (reference window 600 DAA) on the replay: std 0.026, 0 flips, mean 142.6M against 139M true; on a 3-node test network the v2 nodes eased a leave with no peak and held a rejoin within 3% after 60 s, and a node without the activation height forked off at it as designed. Rule v2 rolled onto the 12-node cloud network on 4 October (all nodes crossed the height on one chain; a hash-rate step then settled in 160 to 270 s with no swing) and activates on the devnet at DAA 33,000 the same evening. Timestamp forging (ledger M23) fixed the same day: a 50% forger drifts the rate under 1.1% where the 3 October rule gave it a 9.9x difficulty. Simulator, settled seconds: x50 step 62 to 66 (Kaspa 1,542), /50 step 657 to 753 (Kaspa 12,296). Apple M5 Max under load 7 to 442; the DAG model is fitted on one scale; the pool hopper's 0.7-point excess over Kaspa's rule stays open | none yet |
| 13 | Every node executes the ordered transactions natively and reaches the same state root | Litepaper Proving ("Every node executes ... natively"), Building ("runs on Igneum unchanged") | tested by the team | repo `f5f8c80`, `8dae48b`; fork `devnet-v4` `dc749905`; revm 43.0.3 | `node tools/evm-smoke/smoke.mjs` against a 3-node `igneumd`; `igneum-exec-diff seq.json`; bench-log "execution layer devnet v3" and "devnet-v4 integration" | Simnet, 3 October 2026: 87 of 87 viem checks, state roots identical on 3 nodes at four heights, 57 executed and 19 skipped transactions agree with plain revm, 0 mismatches. Merged node on real proof of work, 4 October 2026: 84 of 85 checks (the miss needs parallel blocks the network did not produce in 36 s), 59 transfers in 10 chain blocks, state roots identical on 3 nodes, `igneum-exec-diff` 0 mismatches over 59 transactions; the live devnet v4 runs this execution layer. Apple M5 Max. The prover is a stub; state is rebuilt from genesis at start; no EVM transaction relay between nodes | none yet |
| 14 | Ethereum bytecode runs unchanged, with the documented differences of spec 7.1 | Homepage Build card; litepaper Building | tested by the team | as row 13; fixes `F-exec-A`, `F-exec-B` (spec 7.5) | `tools/evm-smoke/smoke.mjs`: deploy via viem, `increment`, `hashLoop`, `eth_estimateGas`, `eth_getLogs`; `tools/exec-attacks` scenarios 1 and 3; bench-log "execution layer attack fixes" | Deployment, calls, reverts, logs and gas estimates behave as viem expects; chain id 4463; the prototype pgas table gives 0.0095 to 0.028 pgas per gas, below the design's band before calibration, 3 October 2026. 4 October 2026: a transaction that would cross the block's proving budget is refused by the mempool and, if forced in, aborted and charged with its nonce advanced (25 of 25 checks; 30 of 30 malformed cases). Apple M5 Max. The `Prover` precompile, proof records and the shard planner are not in the node | none yet |
| 15 | Every block is proven, with the proof landing within about a minute at launch | Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate | implemented | repo `d7e1f89` (GPU proof), `e01a3cc`, `292e800`, `eedd136` (`proving/igneum-prove`: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6 | `proving/windows-wsl2` (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; `igneum-prove-host --mode block` on `proving/fixtures/`; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards" | First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture `block-78-increment` (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in `docs/benchmarks/proving-e2e.md`. Two host defects (an abort after the upload, a 10-minute idle wait) fixed, to be confirmed on the PC (ledger P20) | none yet |
| 16 | A 12 GB card proves one shard in about 20 s | Litepaper Proving ("The proving budget"); roadmap gate 2 | designed | spec 5.1 (Target), 7.6 (`S_p` provisional, 7,500,000 pgas = `B_p` / 4) | `PROVE-SHARD.bat` on the RTX 5090 (pending); the end-to-end standard in `docs/benchmarks/proving-e2e.md`; bench-log "proving: devnet v4 shards" | Unmeasured on any GPU. A shard at the provisional `S_p` is 60 M SP1 cycles on the prototype pgas table (9 cycles per pgas; the modexp entry about 100x its SP1 cost), executed in 4.6 to 7.7 s on the Mac CPU and not yet proven; the GPU row of the bench-log table is empty until the PC runs it, 4 October 2026. A per-shard time can be met by shrinking the shard, so the project does not use it as a pass mark | none yet |
| 15 | Every block is proven, with the proof landing within about a minute at launch | Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate | implemented | repo `d7e1f89` (GPU proof), `e01a3cc`, `292e800`, `eedd136` (`proving/igneum-prove`: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6 | `proving/windows-wsl2` (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; `igneum-prove-host --mode block` on `proving/fixtures/`; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards" | First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture `block-78-increment` (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in `docs/benchmarks/proving-e2e.md`. Second RTX 5090 run, 4 October 2026 evening (job run-20261004-173115): a full shard at the provisional S_p (6.75 M pgas, 60.8 M cycles) executed in 1.63 s, core proof 8.3 s (18.1 MB), compressed proof 10.9 s (1.27 MB, verify 0.040 s); a two-shard block (13.5 M pgas) proved shard by shard (11.7 s and 10.0 s) and aggregated in 2.2 s, 24 s of GPU stages end to end, every proof verified, six tampered witnesses rejected. The two host defects (an abort after the upload, an idle wait that turned out to be an unbuffered 18 MB proof save through the WSL2 file bridge, 24 minutes) are fixed (ledger P20) | none yet |
| 16 | A 12 GB card proves one shard in about 20 s | Litepaper Proving ("The proving budget"); roadmap gate 2 | designed | spec 5.1 (Target), 7.6 (`S_p` provisional, 7,500,000 pgas = `B_p` / 4) | `PROVE-SHARD.bat` on the RTX 5090 (pending); the end-to-end standard in `docs/benchmarks/proving-e2e.md`; bench-log "proving: devnet v4 shards" | Measured on a 32 GB card, not yet on a 12 GB card. A shard at the provisional `S_p` is 60.8 M SP1 cycles on the prototype pgas table (9 cycles per pgas, 44 per EVM gas; the modexp entry about 100x its SP1 cost); on an RTX 5090 (4 October 2026 evening, job run-20261004-173115) it executed in 1.63 s and its compressed proof took 10.9 s, verified in 0.040 s, so the 32 GB card is inside the 20 s target with margin. Whether a 12 GB card proves it at all, and in what time, is the next measurement (an RTX 3060 and an RTX 5060 Ti 16 GB are on order). A per-shard time can be met by shrinking the shard, so the project does not use it as a pass mark | none yet |
| 17 | The chip resistance target: a chip gains under 2x over a GPU | Litepaper Mining, "What Igneum does not claim"; homepage "no chip can be built for it" | designed | spec 0.2 (Target); O-1.17 | Public benchmark with a leaderboard by card model and a standing bounty, January 2027 (O-1.17); the on-die-SRAM test on the RTX 5090 (R3.5) | A target, not a measurement. Review round 3 priced a recompute chip with the 256 MiB cache on die at about 2.4x, approximate, before the usual chip-versus-GPU integer gain; the design answer (cache larger than any die) is open (spec 1.16) | none yet |
| 18 | The chip resistance measurements: the program is random-access bound, not bandwidth bound, and sits beyond a card's on-chip cache | Litepaper Mining ("bound by memory bandwidth", to be corrected), vs RandomX "Measured so far" | tested by the team | repo `aba248d`, `f2a1a64`, `4b95c5e` | RTX 5090 dataset sweep 4 MiB to 1 GiB with `proto-cuda/host.cu`; bench-log "RTX 5090 first run" and "dataset sweep" | At 1 GiB: 228.1 Mhash/s, 23.7 G random loads/s, 94.9 GB/s useful against a 1,638 GB/s dataset fill; inside the 96 MiB L2 (4 and 64 MiB) 1,340 to 1,353 Mhash/s, about 5.8x faster; 104 against 128 loads per hash gives 228 against 185 Mhash/s, proportional. 3 October 2026, RTX 5090, Windows, CUDA 12.8, version 1 programs. Prototype dataset 1 GiB against 2 GB at genesis; a pure random-read microbenchmark (R3 chip designer, attack 2) has not run; the sweep has not been repeated on version 2 | none yet |
| 19 | The lottery hash is sound as a hash: uniform output, deterministic, no out-of-bounds read, fuzzed | Litepaper vs RandomX ("Every number above is measured and logged") | tested by the team | repo `c52307e`, `58a5a63`, `b27da39`; `proto-metal/TESTS.md` | `proto-metal/igneum-bench --fuzz --edge --stats --determinism --memcheck`; `--fuzz 2000` on the version 2 generator; `igneum-census`; bench-log "hardening tests", the re-run on the memory-hard dataset, "generator version 2 adopted" | Version 1: 10,200 random programs, 1,305,600 hashes, 0 mismatches; 14 of 14 edge cases; bit frequency within 2.90 sigma, avalanche mean 31.99 to 32.04 of 32; deterministic fingerprint across 5 runs; every dataset read masked, 3 October 2026. Version 2, 4 October 2026: 2,000 random programs through the Metal cross-check, 8,000 warps, 0 mismatches, 128 loads per hash on every program; 20,000-program census, 5.2% rejected (4.1% static, 1.1% dynamic). Apple M5 Max. Statistics are not a security proof; the edge, stats and memcheck sections were not re-run on version 2 (they do not depend on the generator); the seed derivation review (O-1.4) is open; the fuzz set has run on Metal and the CPU only | none yet |

File diff suppressed because one or more lines are too long

View file

@ -154,6 +154,7 @@ if (existsSync(join(docs, 'bench-log.md'))) {
['Igneum-node devnet v2', 'Finality rule v2 live on a four-miner test network'],
['Execution layer devnet v3', 'EVM execution layer: identical state on three nodes'],
['Weak-program census', 'Census of 400,000 programs: redundant loads found'],
['Shard proving on the RTX 5090', 'Shard layer on the RTX 5090: a full shard proven in 10.9 s, a block aggregated in 2.2 s'],
['Proving v0 on the RTX 5090', 'First GPU proof of an Igneum block: 1.4 s on an RTX 5090'],
['Proving v0', 'First SP1 proof of an Igneum block, on a laptop CPU'],
['Windows node package', 'Windows node package: cross-compiled, two-peer sync'],

View file

@ -97,8 +97,8 @@ footer{border-top:1px solid var(--line);padding-block:32px 48px;font-size:13px;c
<tr data-status="tested by the team"><td class="n">12</td><td class="claim">The difficulty rule recovers from a hashrate step within minutes, where Kaspa's sampled rule never settles. A step inside an epoch set the rule oscillating on the live devnet on 4 October 2026; rule v2 removes it in the simulator and on a test network and is built but not yet rolled out<div class="where">Spec 2.3; litepaper Speed (implied); bench page</div></td><td><span class="st st-2">tested by the team</span></td><td class="mono">repo <code>e9328c6</code>, <code>abb5a5d</code> (attacks), <code>67bf226</code> (rule v2); fork <code>difficulty</code> branch (timestamp fix) and <code>devnet-v4</code> <code>a21ff239</code> (<code>difficulty_v2_activation_daa</code>, <code>REF_WINDOW_V2 = 600</code>); <code>sim/difficulty/sim.py --live</code></td><td>The live record <code>sim/difficulty/records/live-2026-10-04.csv</code> (8,090 headers, <code>pull_live.py</code>) and the hash-rate record beside it; <code>sim/difficulty/sim.py</code> on the synthetic set and the DAG replay; <code>sim/difficulty/attacks/attacks.py</code>; <code>sim/difficulty/testnet_v2.py</code> (3 nodes, activation at DAA 900); <code>cargo test --release -p kaspa-consensus --lib difficulty</code> (15 pass); bench-log "difficulty controller", "difficulty rule under attack", "timestamp attack fixed", "difficulty rule v2"</td><td>Live devnet v4, 4 October 2026 (UTC): a second RTX 5090 joining 7 minutes into an epoch (about 152 to 280 MH/s) hardened the difficulty 70M to 144M in 90 s and then swung by about a third for 40 minutes around the true level of 139M while the epoch-long reference lane carried the join; that card leaving for 4 minutes eased 116M to 67M and back to 106M; the epoch boundary with both PCs restarting took 152M to 77M in 3 minutes, after which the rule held within 1.3% per minute with no flips. Cause: the reference lane covered the whole epoch, so a mid-epoch step polluted it for the hour and the 25% trigger flipped on the short lane's noise. The DAG replay reproduces the record (std of log difficulty 0.115 against 0.134, 4.3 peaks against 4). Rule v2 (reference window 600 DAA) on the replay: std 0.026, 0 flips, mean 142.6M against 139M true; on a 3-node test network the v2 nodes eased a leave with no peak and held a rejoin within 3% after 60 s, and a node without the activation height forked off at it as designed. Rule v2 rolled onto the 12-node cloud network on 4 October (all nodes crossed the height on one chain; a hash-rate step then settled in 160 to 270 s with no swing) and activates on the devnet at DAA 33,000 the same evening. Timestamp forging (ledger M23) fixed the same day: a 50% forger drifts the rate under 1.1% where the 3 October rule gave it a 9.9x difficulty. Simulator, settled seconds: x50 step 62 to 66 (Kaspa 1,542), /50 step 657 to 753 (Kaspa 12,296). Apple M5 Max under load 7 to 442; the DAG model is fitted on one scale; the pool hopper's 0.7-point excess over Kaspa's rule stays open</td><td class="iv">none yet</td></tr>
<tr data-status="tested by the team"><td class="n">13</td><td class="claim">Every node executes the ordered transactions natively and reaches the same state root<div class="where">Litepaper Proving ("Every node executes ... natively"), Building ("runs on Igneum unchanged")</div></td><td><span class="st st-2">tested by the team</span></td><td class="mono">repo <code>f5f8c80</code>, <code>8dae48b</code>; fork <code>devnet-v4</code> <code>dc749905</code>; revm 43.0.3</td><td><code>node tools/evm-smoke/smoke.mjs</code> against a 3-node <code>igneumd</code>; <code>igneum-exec-diff seq.json</code>; bench-log "execution layer devnet v3" and "devnet-v4 integration"</td><td>Simnet, 3 October 2026: 87 of 87 viem checks, state roots identical on 3 nodes at four heights, 57 executed and 19 skipped transactions agree with plain revm, 0 mismatches. Merged node on real proof of work, 4 October 2026: 84 of 85 checks (the miss needs parallel blocks the network did not produce in 36 s), 59 transfers in 10 chain blocks, state roots identical on 3 nodes, <code>igneum-exec-diff</code> 0 mismatches over 59 transactions; the live devnet v4 runs this execution layer. Apple M5 Max. The prover is a stub; state is rebuilt from genesis at start; no EVM transaction relay between nodes</td><td class="iv">none yet</td></tr>
<tr data-status="tested by the team"><td class="n">14</td><td class="claim">Ethereum bytecode runs unchanged, with the documented differences of spec 7.1<div class="where">Homepage Build card; litepaper Building</div></td><td><span class="st st-2">tested by the team</span></td><td class="mono">as row 13; fixes <code>F-exec-A</code>, <code>F-exec-B</code> (spec 7.5)</td><td><code>tools/evm-smoke/smoke.mjs</code>: deploy via viem, <code>increment</code>, <code>hashLoop</code>, <code>eth_estimateGas</code>, <code>eth_getLogs</code>; <code>tools/exec-attacks</code> scenarios 1 and 3; bench-log "execution layer attack fixes"</td><td>Deployment, calls, reverts, logs and gas estimates behave as viem expects; chain id 4463; the prototype pgas table gives 0.0095 to 0.028 pgas per gas, below the design's band before calibration, 3 October 2026. 4 October 2026: a transaction that would cross the block's proving budget is refused by the mempool and, if forced in, aborted and charged with its nonce advanced (25 of 25 checks; 30 of 30 malformed cases). Apple M5 Max. The <code>Prover</code> precompile, proof records and the shard planner are not in the node</td><td class="iv">none yet</td></tr>
<tr data-status="implemented"><td class="n">15</td><td class="claim">Every block is proven, with the proof landing within about a minute at launch<div class="where">Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate</div></td><td><span class="st st-1">implemented</span></td><td class="mono">repo <code>d7e1f89</code> (GPU proof), <code>e01a3cc</code>, <code>292e800</code>, <code>eedd136</code> (<code>proving/igneum-prove</code>: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6</td><td><code>proving/windows-wsl2</code> (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; <code>igneum-prove-host --mode block</code> on <code>proving/fixtures/</code>; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards"</td><td>First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture <code>block-78-increment</code> (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in <code>docs/benchmarks/proving-e2e.md</code>. Two host defects (an abort after the upload, a 10-minute idle wait) fixed, to be confirmed on the PC (ledger P20)</td><td class="iv">none yet</td></tr>
<tr data-status="designed"><td class="n">16</td><td class="claim">A 12 GB card proves one shard in about 20 s<div class="where">Litepaper Proving ("The proving budget"); roadmap gate 2</div></td><td><span class="st st-0">designed</span></td><td class="mono">spec 5.1 (Target), 7.6 (<code>S_p</code> provisional, 7,500,000 pgas = <code>B_p</code> / 4)</td><td><code>PROVE-SHARD.bat</code> on the RTX 5090 (pending); the end-to-end standard in <code>docs/benchmarks/proving-e2e.md</code>; bench-log "proving: devnet v4 shards"</td><td>Unmeasured on any GPU. A shard at the provisional <code>S_p</code> is 60 M SP1 cycles on the prototype pgas table (9 cycles per pgas; the modexp entry about 100x its SP1 cost), executed in 4.6 to 7.7 s on the Mac CPU and not yet proven; the GPU row of the bench-log table is empty until the PC runs it, 4 October 2026. A per-shard time can be met by shrinking the shard, so the project does not use it as a pass mark</td><td class="iv">none yet</td></tr>
<tr data-status="implemented"><td class="n">15</td><td class="claim">Every block is proven, with the proof landing within about a minute at launch<div class="where">Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate</div></td><td><span class="st st-1">implemented</span></td><td class="mono">repo <code>d7e1f89</code> (GPU proof), <code>e01a3cc</code>, <code>292e800</code>, <code>eedd136</code> (<code>proving/igneum-prove</code>: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6</td><td><code>proving/windows-wsl2</code> (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; <code>igneum-prove-host --mode block</code> on <code>proving/fixtures/</code>; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards"</td><td>First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture <code>block-78-increment</code> (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in <code>docs/benchmarks/proving-e2e.md</code>. Second RTX 5090 run, 4 October 2026 evening (job run-20261004-173115): a full shard at the provisional S_p (6.75 M pgas, 60.8 M cycles) executed in 1.63 s, core proof 8.3 s (18.1 MB), compressed proof 10.9 s (1.27 MB, verify 0.040 s); a two-shard block (13.5 M pgas) proved shard by shard (11.7 s and 10.0 s) and aggregated in 2.2 s, 24 s of GPU stages end to end, every proof verified, six tampered witnesses rejected. The two host defects (an abort after the upload, an idle wait that turned out to be an unbuffered 18 MB proof save through the WSL2 file bridge, 24 minutes) are fixed (ledger P20)</td><td class="iv">none yet</td></tr>
<tr data-status="designed"><td class="n">16</td><td class="claim">A 12 GB card proves one shard in about 20 s<div class="where">Litepaper Proving ("The proving budget"); roadmap gate 2</div></td><td><span class="st st-0">designed</span></td><td class="mono">spec 5.1 (Target), 7.6 (<code>S_p</code> provisional, 7,500,000 pgas = <code>B_p</code> / 4)</td><td><code>PROVE-SHARD.bat</code> on the RTX 5090 (pending); the end-to-end standard in <code>docs/benchmarks/proving-e2e.md</code>; bench-log "proving: devnet v4 shards"</td><td>Measured on a 32 GB card, not yet on a 12 GB card. A shard at the provisional <code>S_p</code> is 60.8 M SP1 cycles on the prototype pgas table (9 cycles per pgas, 44 per EVM gas; the modexp entry about 100x its SP1 cost); on an RTX 5090 (4 October 2026 evening, job run-20261004-173115) it executed in 1.63 s and its compressed proof took 10.9 s, verified in 0.040 s, so the 32 GB card is inside the 20 s target with margin. Whether a 12 GB card proves it at all, and in what time, is the next measurement (an RTX 3060 and an RTX 5060 Ti 16 GB are on order). A per-shard time can be met by shrinking the shard, so the project does not use it as a pass mark</td><td class="iv">none yet</td></tr>
<tr data-status="designed"><td class="n">17</td><td class="claim">The chip resistance target: a chip gains under 2x over a GPU<div class="where">Litepaper Mining, "What Igneum does not claim"; homepage "no chip can be built for it"</div></td><td><span class="st st-0">designed</span></td><td class="mono">spec 0.2 (Target); O-1.17</td><td>Public benchmark with a leaderboard by card model and a standing bounty, January 2027 (O-1.17); the on-die-SRAM test on the RTX 5090 (R3.5)</td><td>A target, not a measurement. Review round 3 priced a recompute chip with the 256 MiB cache on die at about 2.4x, approximate, before the usual chip-versus-GPU integer gain; the design answer (cache larger than any die) is open (spec 1.16)</td><td class="iv">none yet</td></tr>
<tr data-status="tested by the team"><td class="n">18</td><td class="claim">The chip resistance measurements: the program is random-access bound, not bandwidth bound, and sits beyond a card's on-chip cache<div class="where">Litepaper Mining ("bound by memory bandwidth", to be corrected), vs RandomX "Measured so far"</div></td><td><span class="st st-2">tested by the team</span></td><td class="mono">repo <code>aba248d</code>, <code>f2a1a64</code>, <code>4b95c5e</code></td><td>RTX 5090 dataset sweep 4 MiB to 1 GiB with <code>proto-cuda/host.cu</code>; bench-log "RTX 5090 first run" and "dataset sweep"</td><td>At 1 GiB: 228.1 Mhash/s, 23.7 G random loads/s, 94.9 GB/s useful against a 1,638 GB/s dataset fill; inside the 96 MiB L2 (4 and 64 MiB) 1,340 to 1,353 Mhash/s, about 5.8x faster; 104 against 128 loads per hash gives 228 against 185 Mhash/s, proportional. 3 October 2026, RTX 5090, Windows, CUDA 12.8, version 1 programs. Prototype dataset 1 GiB against 2 GB at genesis; a pure random-read microbenchmark (R3 chip designer, attack 2) has not run; the sweep has not been repeated on version 2</td><td class="iv">none yet</td></tr>
<tr data-status="tested by the team"><td class="n">19</td><td class="claim">The lottery hash is sound as a hash: uniform output, deterministic, no out-of-bounds read, fuzzed<div class="where">Litepaper vs RandomX ("Every number above is measured and logged")</div></td><td><span class="st st-2">tested by the team</span></td><td class="mono">repo <code>c52307e</code>, <code>58a5a63</code>, <code>b27da39</code>; <code>proto-metal/TESTS.md</code></td><td><code>proto-metal/igneum-bench --fuzz --edge --stats --determinism --memcheck</code>; <code>--fuzz 2000</code> on the version 2 generator; <code>igneum-census</code>; bench-log "hardening tests", the re-run on the memory-hard dataset, "generator version 2 adopted"</td><td>Version 1: 10,200 random programs, 1,305,600 hashes, 0 mismatches; 14 of 14 edge cases; bit frequency within 2.90 sigma, avalanche mean 31.99 to 32.04 of 32; deterministic fingerprint across 5 runs; every dataset read masked, 3 October 2026. Version 2, 4 October 2026: 2,000 random programs through the Metal cross-check, 8,000 warps, 0 mismatches, 128 loads per hash on every program; 20,000-program census, 5.2% rejected (4.1% static, 1.1% dynamic). Apple M5 Max. Statistics are not a security proof; the edge, stats and memcheck sections were not re-run on version 2 (they do not depend on the generator); the seed derivation review (O-1.4) is open; the fuzz set has run on Metal and the CPU only</td><td class="iv">none yet</td></tr>

View file

@ -50,6 +50,11 @@
}
],
"log": [
{
"date": "2026-10-04",
"text": "Shard proving on the RTX 5090: a full shard compressed in 10.9 s, a two-shard block aggregated in 2.2 s, all verified",
"short": "Shard layer on the RTX 5090: a full shard proven in 10.9 s, a block aggregated in 2.2 s"
},
{
"date": "2026-10-04",
"text": "Cloud devnet: 12 igneumd nodes in 5 locations, inter-region latency, two Singapore partitions, hash-rate steps",
@ -70,11 +75,6 @@
"text": "Difficulty rule v2 activated on the live devnet at DAA 33,000 by height switch, no fresh chain",
"short": "Difficulty rule v2 switched on the live devnet by height, no fresh chain"
},
{
"date": "2026-10-04",
"text": "Shard proving on the RTX 5090",
"short": "Shard proving on the RTX 5090"
},
{
"date": "2026-10-04",
"text": "First outside machine on the devnet: an Apple silicon laptop through the Igneum Miner app",
@ -175,6 +175,41 @@
"text": "Devnet-v4 integration: nine branches merged, 3-node test network on the merged node, Windows cross-build",
"short": "Devnet v4: nine branches merged into one node"
},
{
"date": "2026-10-03",
"text": "First devnet blocks on the real lottery hash: CPU, then Metal GPU, three worker implementations",
"short": "First devnet blocks on the real lottery hash"
},
{
"date": "2026-10-03",
"text": "Proto-opencl: OpenCL path built and proven without AMD silicon",
"short": "OpenCL worker built and proven without AMD silicon"
},
{
"date": "2026-10-03",
"text": "Igneum-node devnet v0: 3-node igneum-devnet at 1 BPS with the 80/20 coinbase and vote_key_hash",
"short": "Devnet v0: three nodes at one block per second"
},
{
"date": "2026-10-03",
"text": "AMD gfx1036 , AMD OpenCL 2.1 driver 3652.0",
"short": "AMD integrated GPU runs the hash through OpenCL"
},
{
"date": "2026-10-03",
"text": "RTX 5090 through NVIDIA OpenCL",
"short": "RTX 5090 through NVIDIA OpenCL"
},
{
"date": "2026-10-03",
"text": "RTX 5090 first run",
"short": "RTX 5090 first run"
},
{
"date": "2026-10-03",
"text": "RTX 5090, memory-hard dataset",
"short": "RTX 5090 on the memory-hard dataset"
},
{
"date": "2026-10-03",
"text": "Per-identity hash rate \"decay\" on the RTX 5090: diagnosis and Metal reproduction",
@ -214,41 +249,6 @@
"date": "2026-10-03",
"text": "Difficulty controller: devnet record, simulator, Igneum dual-lane rule, 3-node CPU test network",
"short": "Igneum dual-lane difficulty rule built and simulated"
},
{
"date": "2026-10-03",
"text": "Devnet started four months ahead of plan. Finality, the EVM layer and the Igneum difficulty controller are in build.",
"short": "Devnet started four months ahead of plan"
},
{
"date": "2026-10-03",
"text": "First blocks mined over the network from a Windows PC: an RTX 5090 and an AMD chip, 16 identities, with the live devnet page drawing every block.",
"short": "First blocks from a Windows PC: RTX 5090 and AMD"
},
{
"date": "2026-10-03",
"text": "First devnet blocks on the real lottery hash: CPU, then Metal GPU, three worker implementations",
"short": "First devnet blocks on the real lottery hash"
},
{
"date": "2026-10-03",
"text": "Igneum-node devnet v0: 3-node igneum-devnet at 1 BPS with the 80/20 coinbase and vote_key_hash",
"short": "Devnet v0: three nodes at one block per second"
},
{
"date": "2026-10-03",
"text": "Proto-opencl: OpenCL path built and proven without AMD silicon",
"short": "OpenCL worker built and proven without AMD silicon"
},
{
"date": "2026-10-03",
"text": "AMD gfx1036 , AMD OpenCL 2.1 driver 3652.0",
"short": "AMD integrated GPU runs the hash through OpenCL"
},
{
"date": "2026-10-03",
"text": "RTX 5090 through NVIDIA OpenCL",
"short": "RTX 5090 through NVIDIA OpenCL"
}
]
}