Igneum bench log
Append-only. Every number here was measured on the machine named, on the date given.
2026-10-03 proto-metal / igneum-bench, first run
@@ -51,12 +51,12 @@ footer{border-top:1px solid var(--line);padding-block:24px 48px;font-size:13px;c2026-10-03 proto-cuda / program pack export (Mac side only; RTX 5090 run pending)
Machine: the same Apple M5 Max. No CUDA toolchain exists on it, so nothing below is an NVIDIA measurement. Added --export-pack <dir> to proto-metal/main.swift. It writes, per seed, the CUDA kernel (kernel.cu), C headers (program.h, vectors.h), JSON twins, and the Metal source, into proto-cuda/packs/<seed>/. Vectors: 3 warps (base nonces 0, 4096, 1000000), 96 x 64-bit outputs from the CPU interpreter, plus dataset words 0..15 and word [MASK]. The exporter runs the Metal kernel for the same warps and refuses to write unless all 96 match.
| Pack | Loads/hash | Op mix | CPU interpreter vs Metal GPU (3 warps) | CUDA text in CPU emulation (clang, 32 threads/warp) |
|---|---|---|---|---|
| igneum-genesis | 104 | load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1 | PASS 3/3 | PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 and 2 warps/block |
| igneum-hourly | 128 | load=16 add=8 xor=7 mad=5 mul=5 mulhi=5 shfl=5 rotl=4 or=3 rotr=3 sub=3 | PASS 3/3 | PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 warp/block |
Sweep sizes 4, 64, 256, 512 MiB in the emulation: dataset self-test PASS at each (vectors only apply at 1 GiB). The regular Metal bench was re-run after the change: igneum-genesis 44.8 Mhash/s and igneum-hourly 36.3 Mhash/s at 1 GiB with 2 x 2^20 batches, PASS 3/3 warps each (consistent with the first-run table above, not a new figure). Harness: proto-cuda/host.cu, build.sh, build.bat, README.md, CHECKLIST.md, emu/. Pending: build and run on the project lead's RTX 5090 (CUDA 12.8 or newer, -arch=sm_120). No NVIDIA hash rate exists yet. Emulation rates are not recorded because they measure the Mac's CPU, not a GPU.
Sweep sizes 4, 64, 256, 512 MiB in the emulation: dataset self-test PASS at each (vectors only apply at 1 GiB). The regular Metal bench was re-run after the change: igneum-genesis 44.8 Mhash/s and igneum-hourly 36.3 Mhash/s at 1 GiB with 2 x 2^20 batches, PASS 3/3 warps each (consistent with the first-run table above, not a new figure). Harness: proto-cuda/host.cu, build.sh, build.bat, README.md, CHECKLIST.md, emu/. Pending: build and run on the project's RTX 5090 (CUDA 12.8 or newer, -arch=sm_120). No NVIDIA hash rate exists yet. Emulation rates are not recorded because they measure the Mac's CPU, not a GPU.
2026-10-03 sim/finality_sim.py, sustained-mining finality vote weight (model, not hardware)
Model: 1 day steps, 86,400 Poisson blocks/day, 1,000 Pareto honest keys (top key 17%), perfect retarget, every block blue, no latency, no VRF noise. Seed 7, seed 11 agrees. Rule: weight = 30-day sum of counted blocks, counted = min(actual, 2 x yesterday + f). Lock at 2/3 of total. Floors f in {1, 10, 100, 1000}. A: weight tracks hashrate at steady state (corr 1.00000); full weight from zero history on day 41 (f=1) to 32 (f=1000). B: 60% renter on one key takes 59.9% of rewards on day 1, crosses 1/3 of weight on day 20 to 26, 50% on day 27 to 34, never 2/3. 75% renter reaches 2/3 on day 29 to 34, day 27 with the cap removed. C: splitting defeats the cap. 10,000 fresh keys at f=1 move the 50% crossing from day 34 to 27 (no-cap figure 26); at f=10 they match no-cap exactly. The 30-day trickle buys 2 more days for 11.6% of the network. D: honest doubling is under-weighted 28 to 31 days; old miners lock alone for 20 to 23 days. E: 30% churn leaves 71% live, lock never lost. Threshold is 1/3: 35% stalls 1 day, 50% stalls 10 to 11 days. Active-24h total removes every stall. F: 51% patient owner holds 51.0% weight from day 45, vetoes from day 7 to 20, never locks alone. 67% owner locks alone from day 34 to 44. Recommend: f=1, cap 2x, window 30 (the window is the defence, the cap is worth 1 to 8 days), total = all keys with weight in the window. Details in sim/results.md.
2026-10-03 proto-metal hardening tests (correctness and soundness of the lottery hash, Metal only)
Machine: the same Apple M5 Max. Added --fuzz, --edge, --stats, --determinism, --memcheck, --inline-dataset to proto-metal/main.swift. Full tables and commands in proto-metal/TESTS.md. Fuzz: 10,200 random programs (200 + 10,000, cold compiles 13.6 to 48.5 ms), 4 random full-range warps each, dataset drawn from 64 MiB / 256 MiB / 1 GiB: 40,800 warps, 1,305,600 hashes, 0 mismatches, 0 compile failures, 0 static mask failures, generator contract (rotl 1..31, mask in {1,2,4,8,16}, src != dst) held on every instruction. 196 s for the 10,000 run. Edge: 14 hand-built cases (rotr by register 0 / 32 / -32 / 31 / 63, rotl 1 and 31, mulhi max operands, shfl masks 1..16, loads at index 0 and MASK via in-range and out-of-range registers, add/sub/mul/mad wraparound, zero loads, 64 loads) with operand values proven by a traced interpreter: 14/14 PASS, 128/128 lanes each. rotl by 0 (never generated) agreed too, recorded as informational only. Stats (3 seeds, 2^20 nonces each): bit frequency max deviation 2.90 sigma over 192 bit positions; avalanche 16,000 flips mean 31.99 to 32.04 (expect 32), std 3.98 to 4.01 (expect 4), every output bit flips with probability 0.490 to 0.508; chi-square on four 16-bit windows all within 2.3 sigma; 0 duplicates. Looks uniform. Not a security proof. Determinism: 5 runs and 3 compiles (one forced cold, 30 ms) of 2^20 hashes gave fingerprint 933787e8cfefccb7 every time; dataset fill deterministic (0b1a77899ee60493 twice) and 4,096 sampled words incl. 0 and MASK match the CPU closed form. Memcheck: every dataset[ in the MSL is dataset[rN & MASK] (13/13 at 3 sizes), CUDA twin 13/13 plus one guarded fill write; 4 MiB run with nonces up to 0xffffffff completed and 4 wrapping warps matched the CPU; 416/416 load indices exceeded MASK before masking. Bench re-run after the changes: igneum-genesis 44.56 Mhash/s, epoch1 47.74 Mhash/s at 1 GiB, PASS 3/3 warps each (within 2 percent of the first-run table). --export-pack igneum-genesis re-run is byte-identical to the existing pack. SHORTCUT MEASURED: --inline-dataset replaces every load with the six-op closed form ds_elem and never reads memory: 4,888 Mhash/s wall (6,274 GPU time) vs 44.6 honest at 1 GiB, about 110x, and about 9x the cache-resident honest rate. With a closed-form dataset the hash is not memory-hard; an expensive dataset derivation is required, not optional. Not demonstrated: cryptographic strength, weak-program frequency and rejection, NVIDIA/AMD bit-exactness (CUDA run still pending), CPU verify gate with an expensive dataset element. Next three tests for the cryptographer are listed in TESTS.md section 8.
3 October 2026, RTX 5090 first run (the project lead's PC, Windows, CUDA 12.8 runtime, driver 13.4, Visual Studio 2026 with the 14.30 toolset selected via vcvarsall -vcvars_ver=14.30)
+3 October 2026, RTX 5090 first run (Windows PC, CUDA 12.8 runtime, driver 13.4, Visual Studio 2026 with the 14.30 toolset selected via vcvarsall -vcvars_ver=14.30)
Pack igneum-genesis, dataset 1024 MiB, 5 batches x 2^24 hashes, 1 warp per block.
| Card | Mhash/s at 1 GiB | GB/s useful | random loads/s | dataset fill | vectors |
|---|---|---|---|---|---|
| NVIDIA RTX 5090 (170 SMs, 32 GB) | 228.1 | 94.9 | 23.7 G | 0.66 ms, 1638 GB/s | 96/96 PASS, standalone and in batch |
| Apple M5 Max (40 GPU cores, same program, same day) | 45.2 | 18.8 | 4.6 G | 2.34 ms, 427 GB/s | 96/96 PASS |
Result: the same hourly program, generated on the Mac, compiled by Apple's Metal and NVIDIA's CUDA, produced identical hashes on both vendors. Vendor independence of the lottery program is demonstrated for one program; igneum-hourly and the dataset sweep are the next runs. The ratio 5090 to M5 Max is about 5x on hashes and on random loads per second, approximate, consistent with a memory-bound program (random-access bound, not bandwidth bound: the 5090 writes the dataset at 1638 GB/s but hashes at 95 GB/s of useful 4-byte loads). Caveat unchanged: the prototype dataset is closed-form and not yet memory-hard (see TESTS.md), so these are prototype numbers, not mining numbers.
@@ -87,7 +87,9 @@ footer{border-top:1px solid var(--line);padding-block:24px 48px;font-size:13px;c3 October 2026, RTX 5090 through NVIDIA OpenCL (fourth compiler path on the same card)
Pack igneum-genesis-mh, 1024 MiB, local-memory exchange (NVIDIA's OpenCL lists no sub-group shuffle extension).
| Check | Result |
|---|---|
| Cache check and dataset self-test | PASS |
| Vectors | 96/96 PASS, batch fingerprint 98af644e993239e2, identical to the AMD gfx1036 run |
| Hash rate | 219.6 Mhash/s via OpenCL against 229.0 via CUDA, about 4% apart, approximate |
Reading: NVIDIA's OpenCL compiler and NVIDIA's CUDA compiler agree with each other, with AMD's OpenCL, with Apple's Metal and OpenCL, and with the CPU reference. The batch fingerprint over 16.7 million consecutive nonces is identical on the AMD integrated chip and the 5090, which is a far stronger statement than the 96 vectors alone. The local-memory exchange costs about 4% against CUDA's warp shuffle on this card, approximate.