7.3 KiB
igneum-bench (proto-metal)
First prototype of Igneum's random-program GPU proof-of-work, on Apple Metal.
One Swift file, no packages, no Xcode. The Metal kernel is generated as text from a seed
and compiled at runtime with MTLDevice.makeLibrary(source:options:).
What it does
- Derives a 32-byte seed from a string (FNV-1a 64, four salts).
- Generates a program of 64 integer instructions over 8 x uint32 lane registers, run for 8 iterations.
Ops: add, sub, mul, mulhi, xor, or, rotl (immediate), rotr (register), mad, shfl_xor (simd_shuffle_xor
across the 32-lane SIMD group, masks 1 to 16), load (
dst ^= dataset[src & MASK]). Load weight 25 percent. Each add picks one of two immediates from a bit of r0 sampled at the top of the iteration, branchlessselect. - Emits Metal Shading Language, compiles it, runs it with threadgroup size 32 (one SIMD group per threadgroup).
- Fills a 1 GiB dataset (2^28 uint32) on the GPU from a closed-form function of (daySeed, index). The CPU computes any element on demand and never holds the dataset.
- Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit.
--hours Nregenerates and recompiles N programs in sequence (the hourly epoch model). Default N is 2 so verification always covers two seeds.
Build
cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal
Tested with Swift 5.8.1 from Command Line Tools on macOS (Darwin 25.6.0), no Xcode.
Run
./igneum-bench # default seed, 1 GiB dataset, 4 batches of 2^22 nonces, 2 epochs
./igneum-bench --seed "my-seed" # a different program
./igneum-bench --hours 3 # three seeds in sequence
./igneum-bench --dataset-log2 20 --batch-log2 16 # small, fast validation
./igneum-bench --dump ./generated # also write the generated .metal source
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis # CUDA program pack
Flags: --seed, --day, --hours, --batch-log2 (default 22), --batches (default 4),
--dataset-log2 (default 28), --verify-warps (default 3), --dump <dir>, --export-pack <dir>.
Exit code 0 means every verified warp matched.
--export-pack <dir> (added 3 October 2026) does not run the bench. It generates the program for --seed,
emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (base nonces 0, 4096, 1000000)
with the CPU interpreter, cross-checks them on the Metal GPU, and writes kernel.cu, program.h, vectors.h,
program.json, vectors.json and program.metal into the directory. See ../proto-cuda/README.md.
Measured on this machine, 3 October 2026
Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory. threadExecutionWidth reported as 32.
Rates are wall-clock over 4 batches x 4,194,304 hashes after one warm-up batch. GPU-timestamp rates
agreed with wall-clock to within 0.1 percent in every run. "GB/s useful" is loads per hash x 4 bytes x
hashes per second; it counts only the 4 bytes the program consumed per load, not the cache line moved.
| Seed | Loads/hash | Compile ms (library + pipeline) | Mhash/s | GB/s useful | CPU verify ms/warp (avg of 20) | Verify |
|---|---|---|---|---|---|---|
| igneum-genesis | 104 | 49.6 (cold) | 45.2 | 18.8 | 0.015 | PASS 3/3 warps |
| igneum-genesis/epoch1 | 104 | 20.3 | 48.4 | 20.1 | 0.019 | PASS 3/3 warps |
| igneum-second-seed | 104 | 46.6 (cold) | 35.5 | 14.8 | 0.016 | PASS 3/3 warps |
| igneum-second-seed/epoch1 | 144 | 18.7 | 35.4 | 20.4 | 0.017 | PASS 3/3 warps |
| igneum-hourly | 128 | 52.0 (cold) | 36.6 | 18.7 | 0.021 | PASS 3/3 warps |
| igneum-hourly/epoch1 | 128 | 21.6 | 37.5 | 19.2 | 0.017 | PASS 3/3 warps |
| igneum-hourly/epoch2 | 120 | 23.7 | 36.6 | 17.6 | 0.016 | PASS 3/3 warps |
Single-run CPU verify times (no repetition) ranged 0.015 to 0.041 ms per warp. Verified warps per program: warp 0, warp 65537, warp 131071 of batch 0 (nonces 0 to 31, 2097184 to 2097215, 4194272 to 4194303). 21 warps, 672 hashes, zero mismatches.
Dataset fill (1 GiB, GPU timestamps): 5.57 ms on the first run of a process (179 GB/s), 2.34 to 2.35 ms on later runs (427 GB/s). The gap is probably first-touch page mapping of the private buffer; not measured further. Compile of the fixed fill kernel: 227 ms the very first time on this Mac, 0.7 to 0.9 ms afterwards, which looks like the system shader cache.
Dataset size sweep, seed igneum-genesis, same program throughout
| Dataset | Mhash/s | GB/s useful | Random loads/s |
|---|---|---|---|
| 4 MiB (2^20) | 569 | 237 | 59.2 G |
| 64 MiB (2^24) | 183 | 76 | 19.0 G |
| 256 MiB (2^26) | 94 | 39 | 9.8 G |
| 512 MiB (2^27) | 69 | 29 | 7.2 G |
| 1 GiB (2^28) | 44 | 18 | 4.6 G |
Observations
- The hash is memory bound at 1 GiB. The same program with identical arithmetic runs 12.8x faster when the dataset fits in cache (4 MiB) than at 1 GiB. Nothing but the dataset size changed.
- It is bound by random access, not by raw bandwidth. The GPU wrote the dataset at 427 GB/s in the same process, yet the hash consumed 15 to 20 GB/s of useful bytes. Each load uses 4 bytes of whatever line the memory system moved. The real traffic is some multiple of the useful figure; the line size was not measured.
- Across six of seven programs the GPU sustained 4.4 to 5.1 G random 4-byte loads per second at 1 GiB. The outlier (igneum-second-seed, 3.7 G/s) has the same 13 loads per iteration as igneum-genesis, so load count alone does not set the rate. The shape of the address dependency chain is the likely cause. Not measured.
- The 10 ms CPU gate passes with a wide margin: 0.015 to 0.041 ms per 32-lane warp, roughly 250x under the gate. Two caveats. The dataset element is a cheap closed form, so on-demand lookup costs nothing; a dataset that is expensive to derive (the anti-ASIC direction) would move this number. And this is one fast M5 core.
- Hourly regeneration is cheap: 19 to 24 ms of compile when warm, 47 to 52 ms for the first program of a process. That is noise against a one-hour epoch.
- The generator's load share came out at 13 to 18 of 64 instructions (20 to 28 percent) against a 25 percent weight. The op mix is printed per run.
- The GPU timestamps and the CPU wall clock agree, so the dispatch itself (not command overhead) is what is being measured. Batches of 2^22 nonces take 86 to 118 ms each.
- Only Apple silicon was measured. Nothing here says anything about NVIDIA or AMD, where SIMD width, cache line size and memory latency differ.
What to try next
- Wider loads (uint4, 16 bytes per load) so the useful bytes approach what the memory system actually moves.
- Two or more independent address chains per lane, to see whether the rate is latency or throughput limited.
- A dataset element that costs real work to derive, then re-measure the CPU verify time against the 10 ms gate.
- A program-quality filter in the generator (for example, reject programs where OR saturates a register).
- Output distribution tests on the 64-bit results before this hash is used for leader election.
- The same kernel on a discrete GPU, through a CUDA or Vulkan port of the generator.
Files
main.swift: generator, MSL emitter, GPU driver, CPU interpreter, CLI.igneum-bench: the built binary (not checked in by intent; rebuild with the command above).