147 lines
9.8 KiB
Markdown
147 lines
9.8 KiB
Markdown
# igneum-bench (proto-metal)
|
|
|
|
First prototype of Igneum's random-program GPU proof-of-work, on Apple Metal.
|
|
One Swift file, no packages, no Xcode. The Metal kernel is generated as text from a seed
|
|
and compiled at runtime with `MTLDevice.makeLibrary(source:options:)`.
|
|
|
|
## What it does
|
|
|
|
1. Derives a 32-byte seed from a string (FNV-1a 64, four salts).
|
|
2. Generates a program of 64 integer instructions over 8 x uint32 lane registers, run for 8 iterations.
|
|
Ops: add, sub, mul, mulhi, xor, or, rotl (immediate), rotr (register), mad, shfl_xor (simd_shuffle_xor
|
|
across the 32-lane SIMD group, masks 1 to 16), load (`dst ^= dataset[src & MASK]`). Load weight 25 percent.
|
|
Each add picks one of two immediates from a bit of r0 sampled at the top of the iteration, branchless `select`.
|
|
3. Emits Metal Shading Language, compiles it, runs it with threadgroup size 32 (one SIMD group per threadgroup).
|
|
4. Builds a 1 GiB dataset (2^28 uint32) on the GPU. Since 3 October 2026 (later the same day) the default is the
|
|
memory-hard construction of `MEMHARD.md`: a 256 MiB cache of chained ChaCha12 blocks filled from the day seed,
|
|
and each 64-byte item derived by 8 dependent cache reads through a seed-parameterised ARX mixer. The original
|
|
closed-form element is still available with `--closed-form`. The hash kernel is the same in both modes.
|
|
5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit. The CPU
|
|
verifier holds the 256 MiB cache (computed on one core, compared word for word with the GPU's) and derives every
|
|
dataset word on demand; it never holds the dataset.
|
|
6. `--hours N` regenerates and recompiles N programs in sequence (the hourly epoch model). Default N is 2 so
|
|
verification always covers two seeds.
|
|
|
|
## Build
|
|
|
|
```
|
|
cd proto-metal
|
|
swiftc -O -o igneum-bench main.swift -framework Metal
|
|
```
|
|
|
|
Tested with Swift 5.8.1 from Command Line Tools on macOS (Darwin 25.6.0), no Xcode.
|
|
|
|
## Run
|
|
|
|
```
|
|
./igneum-bench # default seed, 1 GiB dataset, 4 batches of 2^22 nonces, 2 epochs
|
|
./igneum-bench --seed "my-seed" # a different program
|
|
./igneum-bench --hours 3 # three seeds in sequence
|
|
./igneum-bench --dataset-log2 20 --batch-log2 16 # small, fast validation
|
|
./igneum-bench --dump ./generated # also write the generated .metal source
|
|
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis # CUDA program pack
|
|
```
|
|
|
|
Flags: `--seed`, `--day`, `--hours`, `--batch-log2` (default 22), `--batches` (default 4),
|
|
`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump <dir>`, `--export-pack <dir>`,
|
|
`--closed-form` (original dataset), `--load-weight W` and `--wide-frac P` (generator levers, defaults 25 and 0
|
|
reproduce the default generator exactly; see `MEMHARD.md` section 2.4).
|
|
Exit code 0 means every verified warp matched.
|
|
|
|
Hardening tests (added 3 October 2026, results and commands in `TESTS.md`): `--fuzz N [--fuzz-seed <string>]`,
|
|
`--edge`, `--stats`, `--determinism`, `--memcheck`. Any of these runs instead of the bench; several may be combined
|
|
in one invocation; exit code 0 only if every selected test passed. `--inline-dataset` is a bench variant that
|
|
computes every dataset element inline instead of loading it (the shortcut measurement in `TESTS.md` section 7).
|
|
|
|
`--export-pack <dir>` (added 3 October 2026) does not run the bench. It generates the program for `--seed`,
|
|
emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (base nonces 0, 4096, 1000000)
|
|
with the CPU interpreter, cross-checks them on the Metal GPU, and writes `kernel.cu`, `program.h`, `vectors.h`,
|
|
`program.json`, `vectors.json` and `program.metal` into the directory. See `../proto-cuda/README.md`.
|
|
|
|
## Memory-hard dataset (3 October 2026, later the same day)
|
|
|
|
`MEMHARD.md` specifies the construction and holds every measurement. Headline, Apple M5 Max, 1 GiB, seed
|
|
igneum-genesis: honest kernel 45.2 Mhash/s in both constructions; the inline shortcut kernel went from 5,014 Mhash/s
|
|
(closed form, 111x faster than honest) to 9.49 Mhash/s (memory-hard, 4.8x slower than honest); CPU verification
|
|
0.63 to 0.80 ms per 32-lane warp at 104 loads per hash and 1.21 ms at 144 loads, against the 10 ms gate; cache fill
|
|
2 ms on the GPU and 185 ms on one CPU core; dataset build 20.6 ms. Fuzz, edge, determinism, memcheck and stats were
|
|
re-run on the new dataset and pass. The tables below are the original closed-form measurements and still reproduce
|
|
with `--closed-form`.
|
|
|
|
## Measured on this machine, 3 October 2026
|
|
|
|
Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory. `threadExecutionWidth` reported as 32.
|
|
Rates are wall-clock over 4 batches x 4,194,304 hashes after one warm-up batch. GPU-timestamp rates
|
|
agreed with wall-clock to within 0.1 percent in every run. "GB/s useful" is loads per hash x 4 bytes x
|
|
hashes per second; it counts only the 4 bytes the program consumed per load, not the cache line moved.
|
|
|
|
| Seed | Loads/hash | Compile ms (library + pipeline) | Mhash/s | GB/s useful | CPU verify ms/warp (avg of 20) | Verify |
|
|
|---|---|---|---|---|---|---|
|
|
| igneum-genesis | 104 | 49.6 (cold) | 45.2 | 18.8 | 0.015 | PASS 3/3 warps |
|
|
| igneum-genesis/epoch1 | 104 | 20.3 | 48.4 | 20.1 | 0.019 | PASS 3/3 warps |
|
|
| igneum-second-seed | 104 | 46.6 (cold) | 35.5 | 14.8 | 0.016 | PASS 3/3 warps |
|
|
| igneum-second-seed/epoch1 | 144 | 18.7 | 35.4 | 20.4 | 0.017 | PASS 3/3 warps |
|
|
| igneum-hourly | 128 | 52.0 (cold) | 36.6 | 18.7 | 0.021 | PASS 3/3 warps |
|
|
| igneum-hourly/epoch1 | 128 | 21.6 | 37.5 | 19.2 | 0.017 | PASS 3/3 warps |
|
|
| igneum-hourly/epoch2 | 120 | 23.7 | 36.6 | 17.6 | 0.016 | PASS 3/3 warps |
|
|
|
|
Single-run CPU verify times (no repetition) ranged 0.015 to 0.041 ms per warp. Verified warps per
|
|
program: warp 0, warp 65537, warp 131071 of batch 0 (nonces 0 to 31, 2097184 to 2097215, 4194272 to 4194303).
|
|
21 warps, 672 hashes, zero mismatches.
|
|
|
|
Dataset fill (1 GiB, GPU timestamps): 5.57 ms on the first run of a process (179 GB/s), 2.34 to 2.35 ms on later
|
|
runs (427 GB/s). The gap is probably first-touch page mapping of the private buffer; not measured further.
|
|
Compile of the fixed fill kernel: 227 ms the very first time on this Mac, 0.7 to 0.9 ms afterwards, which looks
|
|
like the system shader cache.
|
|
|
|
### Dataset size sweep, seed igneum-genesis, same program throughout
|
|
|
|
| Dataset | Mhash/s | GB/s useful | Random loads/s |
|
|
|---|---|---|---|
|
|
| 4 MiB (2^20) | 569 | 237 | 59.2 G |
|
|
| 64 MiB (2^24) | 183 | 76 | 19.0 G |
|
|
| 256 MiB (2^26) | 94 | 39 | 9.8 G |
|
|
| 512 MiB (2^27) | 69 | 29 | 7.2 G |
|
|
| 1 GiB (2^28) | 44 | 18 | 4.6 G |
|
|
|
|
## Observations
|
|
|
|
1. The hash is memory bound at 1 GiB. The same program with identical arithmetic runs 12.8x faster when the
|
|
dataset fits in cache (4 MiB) than at 1 GiB. Nothing but the dataset size changed.
|
|
2. It is bound by random access, not by raw bandwidth. The GPU wrote the dataset at 427 GB/s in the same
|
|
process, yet the hash consumed 15 to 20 GB/s of useful bytes. Each load uses 4 bytes of whatever line the
|
|
memory system moved. The real traffic is some multiple of the useful figure; the line size was not measured.
|
|
3. Across six of seven programs the GPU sustained 4.4 to 5.1 G random 4-byte loads per second at 1 GiB.
|
|
The outlier (igneum-second-seed, 3.7 G/s) has the same 13 loads per iteration as igneum-genesis, so load
|
|
count alone does not set the rate. The shape of the address dependency chain is the likely cause. Not measured.
|
|
4. The 10 ms CPU gate passes with a wide margin: 0.015 to 0.041 ms per 32-lane warp, roughly 250x under the
|
|
gate. Two caveats. The dataset element is a cheap closed form, so on-demand lookup costs nothing; a dataset
|
|
that is expensive to derive (the anti-ASIC direction) would move this number. And this is one fast M5 core.
|
|
5. Hourly regeneration is cheap: 19 to 24 ms of compile when warm, 47 to 52 ms for the first program of a
|
|
process. That is noise against a one-hour epoch.
|
|
6. The generator's load share came out at 13 to 18 of 64 instructions (20 to 28 percent) against a 25 percent
|
|
weight. The op mix is printed per run.
|
|
7. The GPU timestamps and the CPU wall clock agree, so the dispatch itself (not command overhead) is what is
|
|
being measured. Batches of 2^22 nonces take 86 to 118 ms each.
|
|
8. Only Apple silicon was measured. Nothing here says anything about NVIDIA or AMD, where SIMD width, cache
|
|
line size and memory latency differ.
|
|
9. Measured later on 3 October 2026 (`TESTS.md` section 7): because the dataset element is a six-operation
|
|
closed form, a kernel that computes it inline instead of loading it runs at about 4,900 Mhash/s against
|
|
44.6 for the honest kernel at 1 GiB, roughly 110x. Observation 1 describes the honest kernel only. The
|
|
prototype is not memory-hard until the dataset element costs more to derive than to load. Resolved later the
|
|
same day: with the memory-hard dataset the inline kernel runs at 0.21 of the honest rate (`MEMHARD.md`).
|
|
|
|
## What to try next
|
|
|
|
- Wider loads (uint4, 16 bytes per load) so the useful bytes approach what the memory system actually moves.
|
|
- Two or more independent address chains per lane, to see whether the rate is latency or throughput limited.
|
|
- Done 3 October 2026: a dataset element that costs real work to derive (`MEMHARD.md`); CPU verify re-measured at 0.63 to 1.21 ms per warp.
|
|
- A program-quality filter in the generator (for example, reject programs where OR saturates a register).
|
|
- Output distribution tests on the 64-bit results before this hash is used for leader election.
|
|
- The same kernel on a discrete GPU, through a CUDA or Vulkan port of the generator.
|
|
|
|
## Files
|
|
|
|
- `main.swift`: generator, MSL emitter, GPU driver, CPU interpreter, memory-hard dataset (cache, items, verifier), CLI.
|
|
- `MEMHARD.md`: the memory-hard dataset construction and its measurements.
|
|
- `igneum-bench`: the built binary (not checked in by intent; rebuild with the command above).
|