igneum/docs/bench-log.md
igneum-labs ff263245ac Igneum: design docs, Metal lottery-hash prototype, CUDA test pack, finality simulation
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 15:06:01 +00:00

4.9 KiB

Igneum bench log

Append-only. Every number here was measured on the machine named, on the date given.

2026-10-03 proto-metal / igneum-bench, first run

Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, macOS Darwin 25.6.0, Swift 5.8.1, Metal 4. Build: swiftc -O -o igneum-bench main.swift -framework Metal. Source: proto-metal/main.swift. Setup: 1 GiB dataset (2^28 uint32), 64 instructions x 8 iterations, threadgroup 32 (threadExecutionWidth 32), 4 timed batches x 2^22 nonces after one warm-up batch. Verification: 3 warps per program, CPU interpreter vs GPU.

Seed Loads/hash Compile ms Mhash/s GB/s useful CPU verify ms/warp Verify
igneum-genesis 104 49.6 (cold) 45.2 18.8 0.015 PASS
igneum-genesis/epoch1 104 20.3 48.4 20.1 0.019 PASS
igneum-second-seed 104 46.6 (cold) 35.5 14.8 0.016 PASS
igneum-second-seed/epoch1 144 18.7 35.4 20.4 0.017 PASS
igneum-hourly 128 52.0 (cold) 36.6 18.7 0.021 PASS
igneum-hourly/epoch1 128 21.6 37.5 19.2 0.017 PASS
igneum-hourly/epoch2 120 23.7 36.6 17.6 0.016 PASS

Dataset size sweep (seed igneum-genesis): 4 MiB 569 Mhash/s, 64 MiB 183, 256 MiB 94, 512 MiB 69, 1 GiB 44. Dataset fill 1 GiB: 2.34 ms GPU time (427 GB/s) warm, 5.57 ms on first run of a process. Result: 21 warps, 672 hashes, zero mismatches. OVERALL PASS. Reading: memory bound at 1 GiB (12.8x drop from cache-resident), limited by random access rather than bandwidth, CPU verify roughly 250x under the 10 ms gate with a cheap dataset element. Apple silicon only. Details in proto-metal/README.md.

2026-10-03 proto-cuda / program pack export (Mac side only; RTX 5090 run pending)

Machine: the same Apple M5 Max. No CUDA toolchain exists on it, so nothing below is an NVIDIA measurement. Added --export-pack <dir> to proto-metal/main.swift. It writes, per seed, the CUDA kernel (kernel.cu), C headers (program.h, vectors.h), JSON twins, and the Metal source, into proto-cuda/packs/<seed>/. Vectors: 3 warps (base nonces 0, 4096, 1000000), 96 x 64-bit outputs from the CPU interpreter, plus dataset words 0..15 and word [MASK]. The exporter runs the Metal kernel for the same warps and refuses to write unless all 96 match.

Pack Loads/hash Op mix CPU interpreter vs Metal GPU (3 warps) CUDA text in CPU emulation (clang, 32 threads/warp)
igneum-genesis 104 load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1 PASS 3/3 PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 and 2 warps/block
igneum-hourly 128 load=16 add=8 xor=7 mad=5 mul=5 mulhi=5 shfl=5 rotl=4 or=3 rotr=3 sub=3 PASS 3/3 PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 warp/block

Sweep sizes 4, 64, 256, 512 MiB in the emulation: dataset self-test PASS at each (vectors only apply at 1 GiB). The regular Metal bench was re-run after the change: igneum-genesis 44.8 Mhash/s and igneum-hourly 36.3 Mhash/s at 1 GiB with 2 x 2^20 batches, PASS 3/3 warps each (consistent with the first-run table above, not a new figure). Harness: proto-cuda/host.cu, build.sh, build.bat, README.md, CHECKLIST.md, emu/. Pending: build and run on the project lead's RTX 5090 (CUDA 12.8 or newer, -arch=sm_120). No NVIDIA hash rate exists yet. Emulation rates are not recorded because they measure the Mac's CPU, not a GPU.

2026-10-03 sim/finality_sim.py, sustained-mining finality vote weight (model, not hardware)

Model: 1 day steps, 86,400 Poisson blocks/day, 1,000 Pareto honest keys (top key 17%), perfect retarget, every block blue, no latency, no VRF noise. Seed 7, seed 11 agrees. Rule: weight = 30-day sum of counted blocks, counted = min(actual, 2 x yesterday + f). Lock at 2/3 of total. Floors f in {1, 10, 100, 1000}. A: weight tracks hashrate at steady state (corr 1.00000); full weight from zero history on day 41 (f=1) to 32 (f=1000). B: 60% renter on one key takes 59.9% of rewards on day 1, crosses 1/3 of weight on day 20 to 26, 50% on day 27 to 34, never 2/3. 75% renter reaches 2/3 on day 29 to 34, day 27 with the cap removed. C: splitting defeats the cap. 10,000 fresh keys at f=1 move the 50% crossing from day 34 to 27 (no-cap figure 26); at f=10 they match no-cap exactly. The 30-day trickle buys 2 more days for 11.6% of the network. D: honest doubling is under-weighted 28 to 31 days; old miners lock alone for 20 to 23 days. E: 30% churn leaves 71% live, lock never lost. Threshold is 1/3: 35% stalls 1 day, 50% stalls 10 to 11 days. Active-24h total removes every stall. F: 51% patient owner holds 51.0% weight from day 45, vetoes from day 7 to 20, never locks alone. 67% owner locks alone from day 34 to 44. Recommend: f=1, cap 2x, window 30 (the window is the defence, the cap is worth 1 to 8 days), total = all keys with weight in the window. Details in sim/results.md.