igneum/proto-metal
igneum-labs 0b3cf78bb0 Site: live chain scene with miner sparks, shard fills and lock ripples; OS and GitHub logos
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 15:59:18 +00:00
..
main.swift Site: live chain scene with miner sparks, shard fills and lock ripples; OS and GitHub logos 2026-10-03 15:59:18 +00:00
README.md Prototype: fuzz, edge, determinism, memcheck and statistics tests; inline-dataset shortcut measured 2026-10-03 15:14:18 +00:00
TESTS.md Prototype: fuzz, edge, determinism, memcheck and statistics tests; inline-dataset shortcut measured 2026-10-03 15:14:18 +00:00

igneum-bench (proto-metal)

First prototype of Igneum's random-program GPU proof-of-work, on Apple Metal. One Swift file, no packages, no Xcode. The Metal kernel is generated as text from a seed and compiled at runtime with MTLDevice.makeLibrary(source:options:).

What it does

  1. Derives a 32-byte seed from a string (FNV-1a 64, four salts).
  2. Generates a program of 64 integer instructions over 8 x uint32 lane registers, run for 8 iterations. Ops: add, sub, mul, mulhi, xor, or, rotl (immediate), rotr (register), mad, shfl_xor (simd_shuffle_xor across the 32-lane SIMD group, masks 1 to 16), load (dst ^= dataset[src & MASK]). Load weight 25 percent. Each add picks one of two immediates from a bit of r0 sampled at the top of the iteration, branchless select.
  3. Emits Metal Shading Language, compiles it, runs it with threadgroup size 32 (one SIMD group per threadgroup).
  4. Fills a 1 GiB dataset (2^28 uint32) on the GPU from a closed-form function of (daySeed, index). The CPU computes any element on demand and never holds the dataset.
  5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit.
  6. --hours N regenerates and recompiles N programs in sequence (the hourly epoch model). Default N is 2 so verification always covers two seeds.

Build

cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal

Tested with Swift 5.8.1 from Command Line Tools on macOS (Darwin 25.6.0), no Xcode.

Run

./igneum-bench                                    # default seed, 1 GiB dataset, 4 batches of 2^22 nonces, 2 epochs
./igneum-bench --seed "my-seed"                    # a different program
./igneum-bench --hours 3                           # three seeds in sequence
./igneum-bench --dataset-log2 20 --batch-log2 16   # small, fast validation
./igneum-bench --dump ./generated                  # also write the generated .metal source
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis   # CUDA program pack

Flags: --seed, --day, --hours, --batch-log2 (default 22), --batches (default 4), --dataset-log2 (default 28), --verify-warps (default 3), --dump <dir>, --export-pack <dir>. Exit code 0 means every verified warp matched.

Hardening tests (added 3 October 2026, results and commands in TESTS.md): --fuzz N [--fuzz-seed <string>], --edge, --stats, --determinism, --memcheck. Any of these runs instead of the bench; several may be combined in one invocation; exit code 0 only if every selected test passed. --inline-dataset is a bench variant that computes every dataset element inline instead of loading it (the shortcut measurement in TESTS.md section 7).

--export-pack <dir> (added 3 October 2026) does not run the bench. It generates the program for --seed, emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (base nonces 0, 4096, 1000000) with the CPU interpreter, cross-checks them on the Metal GPU, and writes kernel.cu, program.h, vectors.h, program.json, vectors.json and program.metal into the directory. See ../proto-cuda/README.md.

Measured on this machine, 3 October 2026

Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory. threadExecutionWidth reported as 32. Rates are wall-clock over 4 batches x 4,194,304 hashes after one warm-up batch. GPU-timestamp rates agreed with wall-clock to within 0.1 percent in every run. "GB/s useful" is loads per hash x 4 bytes x hashes per second; it counts only the 4 bytes the program consumed per load, not the cache line moved.

Seed Loads/hash Compile ms (library + pipeline) Mhash/s GB/s useful CPU verify ms/warp (avg of 20) Verify
igneum-genesis 104 49.6 (cold) 45.2 18.8 0.015 PASS 3/3 warps
igneum-genesis/epoch1 104 20.3 48.4 20.1 0.019 PASS 3/3 warps
igneum-second-seed 104 46.6 (cold) 35.5 14.8 0.016 PASS 3/3 warps
igneum-second-seed/epoch1 144 18.7 35.4 20.4 0.017 PASS 3/3 warps
igneum-hourly 128 52.0 (cold) 36.6 18.7 0.021 PASS 3/3 warps
igneum-hourly/epoch1 128 21.6 37.5 19.2 0.017 PASS 3/3 warps
igneum-hourly/epoch2 120 23.7 36.6 17.6 0.016 PASS 3/3 warps

Single-run CPU verify times (no repetition) ranged 0.015 to 0.041 ms per warp. Verified warps per program: warp 0, warp 65537, warp 131071 of batch 0 (nonces 0 to 31, 2097184 to 2097215, 4194272 to 4194303). 21 warps, 672 hashes, zero mismatches.

Dataset fill (1 GiB, GPU timestamps): 5.57 ms on the first run of a process (179 GB/s), 2.34 to 2.35 ms on later runs (427 GB/s). The gap is probably first-touch page mapping of the private buffer; not measured further. Compile of the fixed fill kernel: 227 ms the very first time on this Mac, 0.7 to 0.9 ms afterwards, which looks like the system shader cache.

Dataset size sweep, seed igneum-genesis, same program throughout

Dataset Mhash/s GB/s useful Random loads/s
4 MiB (2^20) 569 237 59.2 G
64 MiB (2^24) 183 76 19.0 G
256 MiB (2^26) 94 39 9.8 G
512 MiB (2^27) 69 29 7.2 G
1 GiB (2^28) 44 18 4.6 G

Observations

  1. The hash is memory bound at 1 GiB. The same program with identical arithmetic runs 12.8x faster when the dataset fits in cache (4 MiB) than at 1 GiB. Nothing but the dataset size changed.
  2. It is bound by random access, not by raw bandwidth. The GPU wrote the dataset at 427 GB/s in the same process, yet the hash consumed 15 to 20 GB/s of useful bytes. Each load uses 4 bytes of whatever line the memory system moved. The real traffic is some multiple of the useful figure; the line size was not measured.
  3. Across six of seven programs the GPU sustained 4.4 to 5.1 G random 4-byte loads per second at 1 GiB. The outlier (igneum-second-seed, 3.7 G/s) has the same 13 loads per iteration as igneum-genesis, so load count alone does not set the rate. The shape of the address dependency chain is the likely cause. Not measured.
  4. The 10 ms CPU gate passes with a wide margin: 0.015 to 0.041 ms per 32-lane warp, roughly 250x under the gate. Two caveats. The dataset element is a cheap closed form, so on-demand lookup costs nothing; a dataset that is expensive to derive (the anti-ASIC direction) would move this number. And this is one fast M5 core.
  5. Hourly regeneration is cheap: 19 to 24 ms of compile when warm, 47 to 52 ms for the first program of a process. That is noise against a one-hour epoch.
  6. The generator's load share came out at 13 to 18 of 64 instructions (20 to 28 percent) against a 25 percent weight. The op mix is printed per run.
  7. The GPU timestamps and the CPU wall clock agree, so the dispatch itself (not command overhead) is what is being measured. Batches of 2^22 nonces take 86 to 118 ms each.
  8. Only Apple silicon was measured. Nothing here says anything about NVIDIA or AMD, where SIMD width, cache line size and memory latency differ.
  9. Measured later on 3 October 2026 (TESTS.md section 7): because the dataset element is a six-operation closed form, a kernel that computes it inline instead of loading it runs at about 4,900 Mhash/s against 44.6 for the honest kernel at 1 GiB, roughly 110x. Observation 1 describes the honest kernel only. The prototype is not memory-hard until the dataset element costs more to derive than to load.

What to try next

  • Wider loads (uint4, 16 bytes per load) so the useful bytes approach what the memory system actually moves.
  • Two or more independent address chains per lane, to see whether the rate is latency or throughput limited.
  • A dataset element that costs real work to derive, then re-measure the CPU verify time against the 10 ms gate.
  • A program-quality filter in the generator (for example, reject programs where OR saturates a register).
  • Output distribution tests on the 64-bit results before this hash is used for leader election.
  • The same kernel on a discrete GPU, through a CUDA or Vulkan port of the generator.

Files

  • main.swift: generator, MSL emitter, GPU driver, CPU interpreter, CLI.
  • igneum-bench: the built binary (not checked in by intent; rebuild with the command above).