|
|
||
|---|---|---|
| .. | ||
| dot4-probe.swift | ||
| family-probe.swift | ||
| main.swift | ||
| MEMHARD.md | ||
| packbench.swift | ||
| README.md | ||
| TESTS.md | ||
igneum-bench (proto-metal)
First prototype of Igneum's random-program GPU proof-of-work, on Apple Metal.
One Swift file, no packages, no Xcode. The Metal kernel is generated as text from a seed
and compiled at runtime with MTLDevice.makeLibrary(source:options:).
What it does
- Derives a 32-byte seed from a string (FNV-1a 64, four salts).
- Generates a program of 64 integer instructions over 8 x uint32 lane registers, run for 8 iterations.
Ops: add, sub, mul, mulhi, xor, or, rotl (immediate), rotr (register), mad, shfl_xor (simd_shuffle_xor
across the 32-lane SIMD group, masks 1 to 16), load (
dst ^= dataset[src & MASK]). Load weight 25 percent. Each add picks one of two immediates from a bit of r0 sampled at the top of the iteration, branchlessselect. - Emits Metal Shading Language, compiles it, runs it with threadgroup size 32 (one SIMD group per threadgroup).
- Builds a 1 GiB dataset (2^28 uint32) on the GPU. Since 3 October 2026 (later the same day) the default is the
memory-hard construction of
MEMHARD.md: a 256 MiB cache of chained ChaCha12 blocks filled from the day seed, and each 64-byte item derived by 8 dependent cache reads through a seed-parameterised ARX mixer. The original closed-form element is still available with--closed-form. The hash kernel is the same in both modes. - Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit. The CPU verifier holds the 256 MiB cache (computed on one core, compared word for word with the GPU's) and derives every dataset word on demand; it never holds the dataset.
--hours Nregenerates and recompiles N programs in sequence (the hourly epoch model). Default N is 2 so verification always covers two seeds.
Build
cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal
Tested with Swift 5.8.1 from Command Line Tools on macOS (Darwin 25.6.0), no Xcode.
Run
./igneum-bench # default seed, 1 GiB dataset, 4 batches of 2^22 nonces, 2 epochs
./igneum-bench --seed "my-seed" # a different program
./igneum-bench --hours 3 # three seeds in sequence
./igneum-bench --dataset-log2 20 --batch-log2 16 # small, fast validation
./igneum-bench --dump ./generated # also write the generated .metal source
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis # CUDA program pack
Flags: --seed, --day, --hours, --batch-log2 (default 22), --batches (default 4),
--dataset-log2 (default 28), --verify-warps (default 3), --dump <dir>, --export-pack <dir>,
--closed-form (original dataset), --load-weight W and --wide-frac P (generator levers, defaults 25 and 0
reproduce the default generator exactly; see MEMHARD.md section 2.4).
Exit code 0 means every verified warp matched.
Hardening tests (added 3 October 2026, results and commands in TESTS.md): --fuzz N [--fuzz-seed <string>],
--edge, --stats, --determinism, --memcheck. Any of these runs instead of the bench; several may be combined
in one invocation; exit code 0 only if every selected test passed. --inline-dataset is a bench variant that
computes every dataset element inline instead of loading it (the shortcut measurement in TESTS.md section 7).
--export-pack <dir> (added 3 October 2026) does not run the bench. It generates the program for --seed,
emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (base nonces 0, 4096, 1000000)
with the CPU interpreter, cross-checks them on the Metal GPU, and writes kernel.cu, program.h, vectors.h,
program.json, vectors.json and program.metal into the directory. See ../proto-cuda/README.md.
Memory-hard dataset (3 October 2026, later the same day)
MEMHARD.md specifies the construction and holds every measurement. Headline, Apple M5 Max, 1 GiB, seed
igneum-genesis: honest kernel 45.2 Mhash/s in both constructions; the inline shortcut kernel went from 5,014 Mhash/s
(closed form, 111x faster than honest) to 9.49 Mhash/s (memory-hard, 4.8x slower than honest); CPU verification
0.63 to 0.80 ms per 32-lane warp at 104 loads per hash and 1.21 ms at 144 loads, against the 10 ms gate; cache fill
2 ms on the GPU and 185 ms on one CPU core; dataset build 20.6 ms. Fuzz, edge, determinism, memcheck and stats were
re-run on the new dataset and pass. The tables below are the original closed-form measurements and still reproduce
with --closed-form.
Measured on this machine, 3 October 2026
Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory. threadExecutionWidth reported as 32.
Rates are wall-clock over 4 batches x 4,194,304 hashes after one warm-up batch. GPU-timestamp rates
agreed with wall-clock to within 0.1 percent in every run. "GB/s useful" is loads per hash x 4 bytes x
hashes per second; it counts only the 4 bytes the program consumed per load, not the cache line moved.
| Seed | Loads/hash | Compile ms (library + pipeline) | Mhash/s | GB/s useful | CPU verify ms/warp (avg of 20) | Verify |
|---|---|---|---|---|---|---|
| igneum-genesis | 104 | 49.6 (cold) | 45.2 | 18.8 | 0.015 | PASS 3/3 warps |
| igneum-genesis/epoch1 | 104 | 20.3 | 48.4 | 20.1 | 0.019 | PASS 3/3 warps |
| igneum-second-seed | 104 | 46.6 (cold) | 35.5 | 14.8 | 0.016 | PASS 3/3 warps |
| igneum-second-seed/epoch1 | 144 | 18.7 | 35.4 | 20.4 | 0.017 | PASS 3/3 warps |
| igneum-hourly | 128 | 52.0 (cold) | 36.6 | 18.7 | 0.021 | PASS 3/3 warps |
| igneum-hourly/epoch1 | 128 | 21.6 | 37.5 | 19.2 | 0.017 | PASS 3/3 warps |
| igneum-hourly/epoch2 | 120 | 23.7 | 36.6 | 17.6 | 0.016 | PASS 3/3 warps |
Single-run CPU verify times (no repetition) ranged 0.015 to 0.041 ms per warp. Verified warps per program: warp 0, warp 65537, warp 131071 of batch 0 (nonces 0 to 31, 2097184 to 2097215, 4194272 to 4194303). 21 warps, 672 hashes, zero mismatches.
Dataset fill (1 GiB, GPU timestamps): 5.57 ms on the first run of a process (179 GB/s), 2.34 to 2.35 ms on later runs (427 GB/s). The gap is probably first-touch page mapping of the private buffer; not measured further. Compile of the fixed fill kernel: 227 ms the very first time on this Mac, 0.7 to 0.9 ms afterwards, which looks like the system shader cache.
Dataset size sweep, seed igneum-genesis, same program throughout
| Dataset | Mhash/s | GB/s useful | Random loads/s |
|---|---|---|---|
| 4 MiB (2^20) | 569 | 237 | 59.2 G |
| 64 MiB (2^24) | 183 | 76 | 19.0 G |
| 256 MiB (2^26) | 94 | 39 | 9.8 G |
| 512 MiB (2^27) | 69 | 29 | 7.2 G |
| 1 GiB (2^28) | 44 | 18 | 4.6 G |
Observations
- The hash is memory bound at 1 GiB. The same program with identical arithmetic runs 12.8x faster when the dataset fits in cache (4 MiB) than at 1 GiB. Nothing but the dataset size changed.
- It is bound by random access, not by raw bandwidth. The GPU wrote the dataset at 427 GB/s in the same process, yet the hash consumed 15 to 20 GB/s of useful bytes. Each load uses 4 bytes of whatever line the memory system moved. The real traffic is some multiple of the useful figure; the line size was not measured.
- Across six of seven programs the GPU sustained 4.4 to 5.1 G random 4-byte loads per second at 1 GiB. The outlier (igneum-second-seed, 3.7 G/s) has the same 13 loads per iteration as igneum-genesis, so load count alone does not set the rate. The shape of the address dependency chain is the likely cause. Not measured.
- The 10 ms CPU gate passes with a wide margin: 0.015 to 0.041 ms per 32-lane warp, roughly 250x under the gate. Two caveats. The dataset element is a cheap closed form, so on-demand lookup costs nothing; a dataset that is expensive to derive (the anti-ASIC direction) would move this number. And this is one fast M5 core.
- Hourly regeneration is cheap: 19 to 24 ms of compile when warm, 47 to 52 ms for the first program of a process. That is noise against a one-hour epoch.
- The generator's load share came out at 13 to 18 of 64 instructions (20 to 28 percent) against a 25 percent weight. The op mix is printed per run.
- The GPU timestamps and the CPU wall clock agree, so the dispatch itself (not command overhead) is what is being measured. Batches of 2^22 nonces take 86 to 118 ms each.
- Only Apple silicon was measured. Nothing here says anything about NVIDIA or AMD, where SIMD width, cache line size and memory latency differ.
- Measured later on 3 October 2026 (
TESTS.mdsection 7): because the dataset element is a six-operation closed form, a kernel that computes it inline instead of loading it runs at about 4,900 Mhash/s against 44.6 for the honest kernel at 1 GiB, roughly 110x. Observation 1 describes the honest kernel only. The prototype is not memory-hard until the dataset element costs more to derive than to load. Resolved later the same day: with the memory-hard dataset the inline kernel runs at 0.21 of the honest rate (MEMHARD.md).
What to try next
- Wider loads (uint4, 16 bytes per load) so the useful bytes approach what the memory system actually moves.
- Two or more independent address chains per lane, to see whether the rate is latency or throughput limited.
- Done 3 October 2026: a dataset element that costs real work to derive (
MEMHARD.md); CPU verify re-measured at 0.63 to 1.21 ms per warp. - A program-quality filter in the generator (for example, reject programs where OR saturates a register).
- Output distribution tests on the 64-bit results before this hash is used for leader election.
- The same kernel on a discrete GPU, through a CUDA or Vulkan port of the generator.
Files
main.swift: generator, MSL emitter, GPU driver, CPU interpreter, memory-hard dataset (cache, items, verifier), CLI.MEMHARD.md: the memory-hard dataset construction and its measurements.igneum-bench: the built binary (not checked in by intent; rebuild with the command above).