# igneum-bench (proto-metal)
First prototype of Igneum's random-program GPU proof-of-work, on Apple Metal.
One Swift file, no packages, no Xcode. The Metal kernel is generated as text from a seed
and compiled at runtime with `MTLDevice.makeLibrary(source:options:)`.
## What it does
1. Derives a 32-byte seed from a string (FNV-1a 64, four salts).
2. Generates a program of 64 integer instructions over 8 x uint32 lane registers, run for 8 iterations.
Ops: add, sub, mul, mulhi, xor, or, rotl (immediate), rotr (register), mad, shfl_xor (simd_shuffle_xor
across the 32-lane SIMD group, masks 1 to 16), load (`dst ^= dataset[src & MASK]`). Load weight 25 percent.
Each add picks one of two immediates from a bit of r0 sampled at the top of the iteration, branchless `select`.
3. Emits Metal Shading Language, compiles it, runs it with threadgroup size 32 (one SIMD group per threadgroup).
4. Fills a 1 GiB dataset (2^28 uint32) on the GPU from a closed-form function of (daySeed, index).
The CPU computes any element on demand and never holds the dataset.
5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit.
6. `--hours N` regenerates and recompiles N programs in sequence (the hourly epoch model). Default N is 2 so
verification always covers two seeds.
## Build
```
cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal
```
Tested with Swift 5.8.1 from Command Line Tools on macOS (Darwin 25.6.0), no Xcode.
## Run
```
./igneum-bench # default seed, 1 GiB dataset, 4 batches of 2^22 nonces, 2 epochs
./igneum-bench --seed "my-seed" # a different program
./igneum-bench --hours 3 # three seeds in sequence
./igneum-bench --dataset-log2 20 --batch-log2 16 # small, fast validation
./igneum-bench --dump ./generated # also write the generated .metal source
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis # CUDA program pack
```
Flags: `--seed`, `--day`, `--hours`, `--batch-log2` (default 22), `--batches` (default 4),
`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump
`, `--export-pack `.
Exit code 0 means every verified warp matched.
`--export-pack ` (added 3 October 2026) does not run the bench. It generates the program for `--seed`,
emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (base nonces 0, 4096, 1000000)
with the CPU interpreter, cross-checks them on the Metal GPU, and writes `kernel.cu`, `program.h`, `vectors.h`,
`program.json`, `vectors.json` and `program.metal` into the directory. See `../proto-cuda/README.md`.
## Measured on this machine, 3 October 2026
Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory. `threadExecutionWidth` reported as 32.
Rates are wall-clock over 4 batches x 4,194,304 hashes after one warm-up batch. GPU-timestamp rates
agreed with wall-clock to within 0.1 percent in every run. "GB/s useful" is loads per hash x 4 bytes x
hashes per second; it counts only the 4 bytes the program consumed per load, not the cache line moved.
| Seed | Loads/hash | Compile ms (library + pipeline) | Mhash/s | GB/s useful | CPU verify ms/warp (avg of 20) | Verify |
|---|---|---|---|---|---|---|
| igneum-genesis | 104 | 49.6 (cold) | 45.2 | 18.8 | 0.015 | PASS 3/3 warps |
| igneum-genesis/epoch1 | 104 | 20.3 | 48.4 | 20.1 | 0.019 | PASS 3/3 warps |
| igneum-second-seed | 104 | 46.6 (cold) | 35.5 | 14.8 | 0.016 | PASS 3/3 warps |
| igneum-second-seed/epoch1 | 144 | 18.7 | 35.4 | 20.4 | 0.017 | PASS 3/3 warps |
| igneum-hourly | 128 | 52.0 (cold) | 36.6 | 18.7 | 0.021 | PASS 3/3 warps |
| igneum-hourly/epoch1 | 128 | 21.6 | 37.5 | 19.2 | 0.017 | PASS 3/3 warps |
| igneum-hourly/epoch2 | 120 | 23.7 | 36.6 | 17.6 | 0.016 | PASS 3/3 warps |
Single-run CPU verify times (no repetition) ranged 0.015 to 0.041 ms per warp. Verified warps per
program: warp 0, warp 65537, warp 131071 of batch 0 (nonces 0 to 31, 2097184 to 2097215, 4194272 to 4194303).
21 warps, 672 hashes, zero mismatches.
Dataset fill (1 GiB, GPU timestamps): 5.57 ms on the first run of a process (179 GB/s), 2.34 to 2.35 ms on later
runs (427 GB/s). The gap is probably first-touch page mapping of the private buffer; not measured further.
Compile of the fixed fill kernel: 227 ms the very first time on this Mac, 0.7 to 0.9 ms afterwards, which looks
like the system shader cache.
### Dataset size sweep, seed igneum-genesis, same program throughout
| Dataset | Mhash/s | GB/s useful | Random loads/s |
|---|---|---|---|
| 4 MiB (2^20) | 569 | 237 | 59.2 G |
| 64 MiB (2^24) | 183 | 76 | 19.0 G |
| 256 MiB (2^26) | 94 | 39 | 9.8 G |
| 512 MiB (2^27) | 69 | 29 | 7.2 G |
| 1 GiB (2^28) | 44 | 18 | 4.6 G |
## Observations
1. The hash is memory bound at 1 GiB. The same program with identical arithmetic runs 12.8x faster when the
dataset fits in cache (4 MiB) than at 1 GiB. Nothing but the dataset size changed.
2. It is bound by random access, not by raw bandwidth. The GPU wrote the dataset at 427 GB/s in the same
process, yet the hash consumed 15 to 20 GB/s of useful bytes. Each load uses 4 bytes of whatever line the
memory system moved. The real traffic is some multiple of the useful figure; the line size was not measured.
3. Across six of seven programs the GPU sustained 4.4 to 5.1 G random 4-byte loads per second at 1 GiB.
The outlier (igneum-second-seed, 3.7 G/s) has the same 13 loads per iteration as igneum-genesis, so load
count alone does not set the rate. The shape of the address dependency chain is the likely cause. Not measured.
4. The 10 ms CPU gate passes with a wide margin: 0.015 to 0.041 ms per 32-lane warp, roughly 250x under the
gate. Two caveats. The dataset element is a cheap closed form, so on-demand lookup costs nothing; a dataset
that is expensive to derive (the anti-ASIC direction) would move this number. And this is one fast M5 core.
5. Hourly regeneration is cheap: 19 to 24 ms of compile when warm, 47 to 52 ms for the first program of a
process. That is noise against a one-hour epoch.
6. The generator's load share came out at 13 to 18 of 64 instructions (20 to 28 percent) against a 25 percent
weight. The op mix is printed per run.
7. The GPU timestamps and the CPU wall clock agree, so the dispatch itself (not command overhead) is what is
being measured. Batches of 2^22 nonces take 86 to 118 ms each.
8. Only Apple silicon was measured. Nothing here says anything about NVIDIA or AMD, where SIMD width, cache
line size and memory latency differ.
## What to try next
- Wider loads (uint4, 16 bytes per load) so the useful bytes approach what the memory system actually moves.
- Two or more independent address chains per lane, to see whether the rate is latency or throughput limited.
- A dataset element that costs real work to derive, then re-measure the CPU verify time against the 10 ms gate.
- A program-quality filter in the generator (for example, reject programs where OR saturates a register).
- Output distribution tests on the 64-bit results before this hash is used for leader election.
- The same kernel on a discrete GPU, through a CUDA or Vulkan port of the generator.
## Files
- `main.swift`: generator, MSL emitter, GPU driver, CPU interpreter, CLI.
- `igneum-bench`: the built binary (not checked in by intent; rebuild with the command above).