igneum-pow 0.2.0: generator v2 draws exactly 16 load slots from instructions 1..63, a load's source from the registers written earlier and not read by a load since, the other 48 ops from the ten non-load weights; accept.rs is spec 01 section 1.4.6 (static: no stale load source, every register injected; dynamic: 64 units on the seed-keyed closed-form dataset, no constant bit, no lane-constant site, under 164 saturated, bias within 136 of 1024, distinct addresses above 245,760); a rejected candidate is replaced by the next attempt of the seed (seed || k_le32), 32 a consensus fault. Packs carry the generator version, attempt and program id. Version 1 kept as generate_v1 for the census. Packs: igneum-genesis, igneum-hourly, igneum-genesis-mh regenerated by igneum-pow export; new igneum-devnet-v4-epoch0 (devnet genesis hash, day bytes 20730). Checks: Rust 39 of 39 tests; Metal natively via the Swift port (export cross-check 3 of 3 warps, identical programs and vectors on five seeds incl. three with attempt 1, fuzz 2,000 of 2,000); CUDA emu 4 of 4 packs; OpenCL emu 2 packs x 2 configurations; Apple OpenCL 4 of 4 packs at 27.9 Mhash/s. Census 20,000: 5.225 percent rejected, accepted distinct mean 127.887. Spec 01 0.2 (1.4.2, 1.4.3, 1.4.6, 1.11, 1.15, 1.16, 1.17), igneum-pow README, the CUDA, OpenCL and Metal test notes, bench-log entry, ledger M5 and M6 Fixed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
22 KiB
proto-metal hardening tests
Date: 3 October 2026. Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, macOS Darwin 25.6.0,
Swift 5.8.1 from Command Line Tools, no Xcode, Metal shaders compiled at runtime.
Build: swiftc -O -o igneum-bench main.swift -framework Metal (4 s).
Every number below was produced on this machine on this date by the commands shown.
Later on 3 October 2026 the default dataset became the memory-hard construction of MEMHARD.md. The tables in this
file are from the original closed-form dataset and reproduce with --closed-form added to each command. Every test
here was re-run on the new dataset with the same commands and passed; those results are in MEMHARD.md section 2.5.
The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one.
On 4 October 2026 the generator became version 2 (exactly 16 loads per program, fresh-source loads, the acceptance
rule of spec 01 section 1.4.6; igneum-pow/src/{generator,accept}.rs, ported into main.swift as
generateProgramV2, candidateProgram, acceptProgram). Sections 1 to 7 below describe the version 1 generator and
its vectors, which are retired; the loads-per-hash ranges they quote no longer occur. Section 9 holds the version 2
run. The --load-weight and --wide-frac levers select the retired version 1 generator and exist only so
MEMHARD.md section 2.4 reproduces.
The tests live in main.swift next to the bench and share its generator, MSL emitter, dataset build and CPU
interpreter without modification. The CPU interpreter gained an optional trace hook (cpuWarpTraced) that the bench
does not use; cpuWarp calls it with nil.
What the GPU-vs-CPU comparison means: the GPU runs the generated Metal kernel, the CPU runs a plain Swift interpreter over the same instruction list, and every 64-bit output is compared bit for bit. The interpreter never reads the dataset buffer; it recomputes each element from the closed form. So agreement also covers the dataset fill kernel and the address masking, not only the arithmetic.
1. Fuzz: random programs, random warps, three dataset sizes
Command (required run):
./igneum-bench --fuzz 200
For each of N seeds: generate the program, emit MSL, static mask check, compile on the GPU, draw a dataset size
from {64 MiB, 256 MiB, 1 GiB}, draw 4 base nonces uniformly from the full 32-bit range (so warps straddle
2^31 and wrap past 2^32), run the GPU, run the CPU interpreter, compare all 128 outputs. The master seed fixes
the whole sequence, so any failure is reproducible with --fuzz-seed. The run also asserts the generator's
contract that the emitter relies on: rotl immediate in 1..31, shuffle mask in {1,2,4,8,16}, source register
never the destination.
| Dataset | Programs | Pass | Fail |
|---|---|---|---|
| 2^24 words (64 MiB) | 63 | 63 | 0 |
| 2^26 words (256 MiB) | 64 | 64 | 0 |
| 2^28 words (1 GiB) | 73 | 73 | 0 |
| all | 200 | 200 | 0 |
200 programs: pass 200, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. 800 warps compared (25,600 hashes). Loads per hash ranged 64 to 208. Op totals across the 200 programs: load 3204, add 1472, xor 1276, mul 1087, mad 1062, shfl 1038, rotl 892, mulhi 762, sub 742, rotr 739, or 526. Compile (library + pipeline) min 0.0 ms, avg 18.3 ms, max 46.4 ms; the zero is the system shader cache answering for the first 20 seeds, which an earlier smoke run with the same master seed had already compiled. GPU dispatch 156 ms total, CPU interpreter 16 ms total, wall 3.9 s.
Large run, separate master seed so every compile is cold:
./igneum-bench --fuzz 10000 --fuzz-seed igneum-fuzz-large-2026-10-03
| Dataset | Programs | Pass | Fail |
|---|---|---|---|
| 2^24 words (64 MiB) | 3346 | 3346 | 0 |
| 2^26 words (256 MiB) | 3334 | 3334 | 0 |
| 2^28 words (1 GiB) | 3320 | 3320 | 0 |
| all | 10000 | 10000 | 0 |
10,000 programs: pass 10,000, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. 40,000 warps compared (1,280,000 hashes). Loads per hash ranged 40 to 232. Op totals: load 160,123, add 76,698, xor 64,138, mad 51,577, shfl 51,298, mul 50,919, rotl 44,945, rotr 38,333, mulhi 38,284, sub 38,234, or 25,451. Compile min 13.6 ms, avg 18.5 ms, max 48.5 ms (all cold). GPU dispatch 8.2 s total, CPU interpreter 0.8 s total, wall 196 s. FUZZ: PASS.
Combined with the 200-program run: 10,200 programs, 40,800 warps, 1,305,600 hashes, zero mismatches. The op totals show every one of the 11 instruction families was exercised tens of thousands of times.
2. Edge cases
Command:
./igneum-bench --edge
Hand-built programs, dataset 1 GiB (MASK 0x0fffffff), 4 warps at base nonces 0x00000000, 0x00100000,
0x7ffffff0 and 0xffffffe0, 128 lanes per case. Operand values are forced with r = r - r (zero) followed by a
constant add, then proven: the traced CPU interpreter checks the named register value just before the
instruction in question, on lane 0 of all 4 warps in all 8 iterations. "Held" means 32 of 32 checks passed.
| Case | Instrs | Loads/hash | Preconditions | GPU vs CPU | Result |
|---|---|---|---|---|---|
| rotl immediate by 1 and by 31 | 6 | 0 | none needed | 128/128 lanes | PASS |
| rotr by register == 0 | 3 | 0 | held | 128/128 lanes | PASS |
| rotr by register == 32 (32 mod 32 = 0) | 5 | 0 | held | 128/128 lanes | PASS |
| rotr by register == 0xFFFFFFE0 (-32, 0 mod 32) | 5 | 0 | held | 128/128 lanes | PASS |
| rotr by register == 31, 63 and 1 | 10 | 0 | held | 128/128 lanes | PASS |
| mulhi 0xFFFFFFFF x 0xFFFFFFFF (result 0xFFFFFFFE), 0x80000000 x 2 (result 1), x 0 (result 0) | 15 | 0 | held, results checked | 128/128 lanes | PASS |
| shfl_xor every mask 1..16 in sequence | 16 | 0 | none needed | 128/128 lanes | PASS |
| load at index 0 via register 0, and via register MASK+1 | 6 | 16 | held | 128/128 lanes | PASS |
| load at index MASK via register MASK, and via register 0xFFFFFFFF | 8 | 16 | held | 128/128 lanes | PASS |
| add wraparound 0xFFFFFFFF + 1 (result 0) | 5 | 0 | held, result checked | 128/128 lanes | PASS |
| sub wraparound 0 - 1 (result 0xFFFFFFFF) | 6 | 0 | held, result checked | 128/128 lanes | PASS |
| mul 0xFFFFFFFF x 0xFFFFFFFF (low result 1), mad same + 5 (result 6) | 15 | 0 | held, results checked | 128/128 lanes | PASS |
| generated program with every load replaced by xor (zero loads) | 64 | 0 | none needed | 128/128 lanes | PASS |
| 64 loads and nothing else | 64 | 512 | none needed | 128/128 lanes | PASS |
| rotl immediate by 0 (outside the generator's 1..31 contract) | 2 | 0 | none needed | 128/128 lanes | info only: agrees |
EDGE: PASS, 14 of 14 counted cases. The last row is informational: the generator never emits a rotate by 0
(the fuzz run asserts this on every instruction), and the MSL rotl_imm would shift by 32 for it. On this GPU
and compiler the result happened to equal the CPU's. Nothing may rely on that; the contract stays 1..31.
The shuffle case covers masks 3, 5, 6, 7, 9 to 15 that the generator does not emit; simd_shuffle_xor and the
interpreter's lane ^ mask agreed for all 16.
3. Output statistics
Command:
./igneum-bench --stats
Three seeds: igneum-genesis, igneum-genesis/stats1, igneum-genesis/stats2. For each, 2^20 consecutive
nonces from 0 hashed on the GPU at 1 GiB (warps 0 and 32767 spot-checked against the CPU, both matched), then
on the CPU side: (a) ones count per output bit, (b) avalanche from single-bit nonce flips, run on the GPU as
pairs of one-warp dispatches, (c) chi-square over 65,536 buckets for each 16-bit window of the output,
(d) duplicate count after sorting.
| Seed | Loads/hash | Bit freq min..max | Max bit deviation (sigma) | Avalanche 1k mean / std | Avalanche 16k mean / std | Per-output-bit flip prob (16k) | Worst chi2 z of 4 windows | Dups |
|---|---|---|---|---|---|---|---|---|
| igneum-genesis | 104 | 0.4990..0.5009 | 2.07 (bit 38) | 32.07 / 4.09 | 32.035 / 3.99 | 0.492..0.507 | 0.42 | 0 |
| igneum-genesis/stats1 | 128 | 0.4986..0.5011 | 2.90 (bit 3) | 32.14 / 3.94 | 31.987 / 3.98 | 0.490..0.508 | 2.26 | 0 |
| igneum-genesis/stats2 | 176 | 0.4989..0.5010 | 2.29 (bit 9) | 31.90 / 3.97 | 32.035 / 4.01 | 0.492..0.508 | 1.65 | 0 |
Chi-square detail (df 65535, expected 65535, sigma 362):
| Seed | bits 0..15 | bits 16..31 | bits 32..47 | bits 48..63 |
|---|---|---|---|---|
| igneum-genesis | 65681 (z 0.40) | 65686 (z 0.42) | 65513 (z -0.06) | 65478 (z -0.16) |
| igneum-genesis/stats1 | 65163 (z -1.03) | 65768 (z 0.64) | 65820 (z 0.79) | 66353 (z 2.26) |
| igneum-genesis/stats2 | 66132 (z 1.65) | 65447 (z -0.24) | 65942 (z 1.12) | 65960 (z 1.17) |
Reading the numbers. (a) With 2^20 samples the sigma on a bit count is 512; the largest deviation over 192 bit positions (3 seeds x 64) was 2.90 sigma, which is what 192 draws from a fair coin produce. (b) An ideal function changes 32 of 64 output bits on average with standard deviation 4. The 1,000-trial means are within 0.14 of 32 (standard error 0.13); the 16,000-trial means are within 0.035 of 32 (standard error 0.03), the standard deviations are 3.98 to 4.01, and every one of the 64 output bits flipped with probability 0.490 to 0.508 (sigma 0.004). The per-input-bit means over 500 flips each ranged 31.5 to 32.5 (sigma 0.18). An earlier 1,000-trial sample for igneum-genesis, drawn with a different RNG salt before the 16,000-trial pass was added, gave mean 31.72; the 16,000-trial figure of 32.035 shows that was sampling noise. (c) All 12 chi-square values are within 2.3 sigma of their expectation. (d) No duplicate among 2^20 outputs; the birthday expectation at 64 bits is 3e-8.
Verdict: the output looks uniform on every measure tried, for all three seeds. STATS: PASS.
This is a sanity check for obvious structural bias. It is not a proof of cryptographic strength, it says nothing about adversarially chosen inputs, and 3 seeds is not a statement about the population of programs.
4. Determinism
Command:
./igneum-bench --determinism
Seed igneum-genesis, 2^20 nonces from 0, 1 GiB dataset. The output buffer is filled with a sentinel before every run so an unwritten lane would show.
| Check | Result |
|---|---|
| Two independent generations of the program give identical MSL text (4,770 bytes) | yes |
| 5 GPU runs of the same pipeline, fingerprint FNV-1a over 8 MiB | 933787e8cfefccb7 all 5 runs, 0 outputs differ, 0 unwritten lanes |
| Second compile of identical source (0.1 ms, served by the system shader cache) | fingerprint identical |
| Third compile with a comment tag appended so the cache misses (30.0 ms, a real recompile) | fingerprint identical |
| CPU interpreter on 8 warps of the reference run (first, last, 6 random) | all match |
| Dataset filled twice, both blitted to shared memory and fingerprinted | 0b1a77899ee60493 both times |
4,096 sampled dataset words incl. indices 0, 1, MASK-1, MASK vs CPU datasetElem |
all match |
GPU time per 2^20 batch: 33.0 ms first, 23.2 to 23.8 ms after. DETERMINISM: PASS.
5. Memory safety of dataset indexing
Command:
./igneum-bench --memcheck
Static: the generated MSL for igneum-genesis at 2^20, 2^24 and 2^28 words has 13 dataset[ accesses, all 13
of the exact form dataset[rN & MASK], and the identifier dataset appears 14 times (13 accesses plus the
kernel parameter). The CUDA twin from --export-pack has 14 ds[ accesses: 13 of the form ds[rN & mask]
and the fill kernel's one write, guarded by if (i < n). The same static check ran on all fuzz programs
(section 1) with zero failures.
Dynamic, 4 MiB dataset (MASK 0x000fffff): four full batches of 2^20 nonces from bases 0xfff00000 (last nonce 0xffffffff), 0xffffffe0 (wraps to 0 inside the batch), 0x80000000 and 0 all completed; 4 verification warps at bases 0xffffffe0, 0xffffffff, 0x80000000, 0xfff00000 matched the CPU. In those warps 416 of 416 load indices were above MASK before masking (a random 32-bit value is below 2^20 with probability 2^-12), so the mask was exercised on every load.
Metal does not bounds-check device buffers, so "it did not crash" is weak evidence on its own. The static check is the guarantee: the emitter has exactly one load template and it masks. MEMCHECK: PASS.
6. The bench still works
Command, after all the changes above:
./igneum-bench
| Seed | Compile ms | Mhash/s (wall) | Mhash/s (GPU time) | GB/s useful | Loads/hash | CPU verify ms/warp | Verify |
|---|---|---|---|---|---|---|---|
| igneum-genesis | 0.2 (shader cache) | 44.56 | 44.59 | 18.54 | 104 | 0.017 | PASS 3/3 warps |
| igneum-genesis/epoch1 | 0.3 (shader cache) | 47.74 | 47.77 | 19.86 | 104 | 0.017 | PASS 3/3 warps |
Dataset fill 5.05 ms GPU (first fill of the process). OVERALL: PASS. The rates match the 3 October table in README.md (45.2 and 48.4) to within 2 percent, so the bench path is unchanged. The compile figures are the system shader cache answering for source compiled earlier today; they are not new compile measurements.
Pack export re-run to a scratch directory and compared byte for byte with the pack written before these
changes (../proto-cuda/packs/igneum-genesis):
./igneum-bench --seed igneum-genesis --export-pack <scratch>/pack-genesis
diff -r <scratch>/pack-genesis ../proto-cuda/packs/igneum-genesis
Metal GPU cross-check PASS 3/3 warps, pack written, and diff -r reported no differences: all six files
(kernel.cu, program.h, vectors.h, program.json, vectors.json, program.metal) are byte-identical
to the pack exported before these changes.
7. Shortcut measurement: the closed-form dataset is not memory-hard
Command:
./igneum-bench --inline-dataset --hours 1
The dataset element is ds_elem(i, d0, d1), six integer operations. A miner can replace every
dataset[a & MASK] with ds_elem(a & MASK, d0, d1) and never read memory. --inline-dataset emits exactly
that kernel; the CPU verification still passes because the function is unchanged.
| Kernel | Dataset | Mhash/s (wall) | Mhash/s (GPU time) | Verify |
|---|---|---|---|---|
| honest (loads from the 1 GiB buffer), this session | 1 GiB | 44.6 | 44.6 | PASS 3/3 |
| honest, README sweep of 3 October | 4 MiB (cache resident) | 569 | not recorded | PASS |
| inline ds_elem, no memory read | 1 GiB mask | 4,888 | 6,274 | PASS 3/3 |
| inline ds_elem, no memory read | 4 MiB mask | 5,646 | 6,024 | PASS 3/3 |
The inline kernel runs about 110x faster than the honest kernel by wall clock (140x by GPU timestamps) and about 9x faster than the honest kernel with a cache-resident dataset. The inline figures are approximate: the 4 timed batches took 2.7 to 3.4 ms in total, so command overhead is visible in the wall figure. The GPU-time figure is consistent with the ALU peak: roughly 1,500 integer instructions per hash at 6.3 Ghash/s is about 9.4 T instructions/s across 40 cores, in the range of this part's compute throughput.
Reading: with a closed-form dataset the prototype hash is compute-bound for anyone who skips the buffer, and
the honest kernel's memory traffic is voluntary. Memory-hardness has to come from a dataset element that
costs more to derive than to load. RandomX gets this from a cache of roughly 256 MiB expanded by SuperscalarHash into a
dataset of roughly 2 GiB (approximate, from memory; vendor/RandomX is not cloned on this machine, so there
is no file citation yet). The README already listed an expensive dataset element under "What to try next"; this
measurement is why it is not optional.
8. Conclusions
Demonstrated on this machine
- Bit-exact GPU-versus-CPU agreement of the random-program hash over 10,200 random programs (200 plus the 10,000 run), 4 random warps each across the full 32-bit nonce range, at 64 MiB, 256 MiB and 1 GiB, plus the 7 programs in README.md, plus 14 hand-built edge cases with their operand values proven. Zero mismatches, zero compile failures.
- Deterministic GPU execution: 5 runs and 3 compiles (one forced cold) of the same program give the same 8 MiB of output; the dataset fill is deterministic and matches the CPU closed form at sampled indices including 0 and MASK.
- Every dataset index in the emitted MSL and CUDA is masked, by static check on every program tested, and the mask is exercised by essentially every load.
- No obvious structural bias in the 64-bit output for 3 seeds at 2^20 nonces: bit frequencies, avalanche (mean 32.0, std 4.0, every output bit flips with probability 0.49 to 0.51), chi-square on four 16-bit windows, zero duplicates.
- The emitter's contract on the generator (rotl 1..31, masks in {1,2,4,8,16}, src != dst) held on every instruction of every fuzzed program.
Not demonstrated
- Cryptographic security of the construction. Nothing here speaks to preimage, second-preimage or collision
resistance, or to an adversary who chooses nonces or influences the epoch seed. The per-register
initialisation is
splitmix32(nonce ^ seed) ^ seed, which is a bijection of the nonce per register, and the body is add-rotate-xor-multiply withor(which destroys information) and no non-linear table. None of that has been analysed. - Resistance to algebraic or structural attacks on weak programs. 3 seeds were measured for bias; the
population of programs was not. A program whose
orchain saturates a register, or whose load addresses collapse to few values, would be both biased and shortcut-able, and nothing yet rejects such programs or bounds how often they occur. - Memory-hardness. Section 7 measures the shortcut directly: with a closed-form dataset the honest kernel's memory traffic is optional. The README's "memory bound at 1 GiB" describes the honest kernel, not the fastest kernel. This is a design gap of the prototype, not a bug in the tests.
- Vendor independence. Only Apple Metal was run. The CUDA twin exists as text and passed a CPU emulation;
it has not run on NVIDIA hardware, where
__shfl_xor_sync,__umulhiand shift semantics must be shown to agree bit for bit with these vectors. - The 10 ms CPU verification gate against a dataset that is expensive to derive. The current 0.02 ms per warp is with a six-operation element.
The next three tests the real cryptographer must do before the spec is final
- Weak-program census and rejection filter. Generate at least 10^5 programs. For each, measure on the CPU
interpreter (no GPU needed): output bit bias and avalanche on 2^12 nonces, distinct load addresses per
hash, fraction of registers saturated by
or, and whether any register's final value is independent of the nonce. Report the distribution, define the rejection rule, and bound the advantage of a miner who can grind the epoch seed over candidate blocks. - Replace the closed-form dataset with an expensive derivation (RandomX style: a cache of hundreds of MiB
expanded by a slow hash, or Argon2-based) and re-run three measurements: the
--inline-datasetshortcut ratio (must fall to about 1), the GPU rate at 1 GiB, and the CPU verify time against the 10 ms gate with on-demand element derivation. Also measure distinct cache lines touched per hash so a cache-resident shortcut is ruled out, not assumed. - Cross-vendor bit-exactness and a written spec. Run the exported pack and the fuzz vectors on NVIDIA (RTX 5090, CUDA) and on at least one AMD GPU, with the same 4-warp-per-program comparison, including the edge-case programs. Then fix the hash as a specification with test vectors and replace the ad-hoc seed derivation (FNV-1a plus SplitMix) with a standard hash so the seed-to-program mapping is auditable.
9. Generator version 2 (4 October 2026)
Build: swiftc -O -o igneum-bench main.swift -framework Metal (7 s), same machine, idle (load 3).
Fuzz, the required run at a reduced size (the version 1 run was 10,000):
./igneum-bench --fuzz 2000 --fuzz-seed igneum-fuzz-gen2-2026-10-04
| Dataset | Programs | Pass | Fail |
|---|---|---|---|
| 2^24 words (64 MiB) | 647 | 647 | 0 |
| 2^26 words (256 MiB) | 645 | 645 | 0 |
| 2^28 words (1 GiB) | 708 | 708 | 0 |
| all | 2000 | 2000 | 0 |
2,000 programs: pass 2,000, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. 8,000 warps compared (256,000 hashes), memory-hard dataset. Loads per hash 128 to 128 (every program; version 1 ranged 40 to 232). Op totals: load 32,000, add 15,350, xor 12,802, mad 10,380, mul 10,208, shfl 10,147, rotl 8,983, sub 7,704, rotr 7,632, mulhi 7,596, or 5,198. Compile min 15.7 ms, avg 21.8 ms, max 67.3 ms (all cold). GPU dispatch 1.7 s total, CPU interpreter 17.7 s total (the memory-hard dataset), wall 91.4 s. Every fuzzed program is the accepted attempt of its seed, so the acceptance rule ran 2,000 times inside this run as well. FUZZ: PASS.
Cross-implementation, the Swift generator against the Rust crate (igneum-pow export for the same seeds, closed
form unless stated): instruction lists, seed words and all 96 vectors identical on igneum-genesis,
igneum-hourly, and the three seeds whose attempt 0 is rejected and attempt 1 accepted
(igneum-census-2026-10-03/22, /37, /51); igneum-genesis memory-hard: identical instruction list and 96
vectors between the Swift export (Metal GPU cross-check PASS 3 of 3 warps, cache FNV 48c4f5bf24166b2e) and the
checked-in Rust pack proto-cuda/packs/igneum-genesis-mh. The Swift exporter is no longer the pack source; it is
the Metal cross-check of the Rust packs (proto-cuda/README.md, "Regenerating a pack"), and its program.json
still writes format igneum-program-pack-2 without the generator fields.
Not re-run under version 2: sections 2 (edge, hand-built programs that bypass the generator), 3 (stats), 4 (determinism), 5 (memcheck) and 7 (shortcut). Nothing in them depends on how the instruction list is drawn.
Files
main.swift: tests under// MARK: - Hardening tests, flags--fuzz,--fuzz-seed,--edge,--stats,--determinism,--memcheck,--inline-dataset; generator version 2 and the acceptance rule under// MARK: - Generator version 2 and the acceptance rule.TESTS.md: this file.- Raw logs of the runs quoted above were kept in the session scratchpad and are not checked in; every table is reproducible with the command above it.