igneum/proto-metal/TESTS.md
igneum-labs fdcab858e3 Lottery hash: generator version 2 (16 load slots, fresh sources, acceptance rule), every vector re-cut, packs regenerated, three workers re-checked, 20,000-program census
igneum-pow 0.2.0: generator v2 draws exactly 16 load slots from instructions 1..63, a
load's source from the registers written earlier and not read by a load since, the other
48 ops from the ten non-load weights; accept.rs is spec 01 section 1.4.6 (static: no
stale load source, every register injected; dynamic: 64 units on the seed-keyed
closed-form dataset, no constant bit, no lane-constant site, under 164 saturated, bias
within 136 of 1024, distinct addresses above 245,760); a rejected candidate is replaced
by the next attempt of the seed (seed || k_le32), 32 a consensus fault. Packs carry the
generator version, attempt and program id. Version 1 kept as generate_v1 for the census.

Packs: igneum-genesis, igneum-hourly, igneum-genesis-mh regenerated by igneum-pow export;
new igneum-devnet-v4-epoch0 (devnet genesis hash, day bytes 20730). Checks: Rust 39 of
39 tests; Metal natively via the Swift port (export cross-check 3 of 3 warps, identical
programs and vectors on five seeds incl. three with attempt 1, fuzz 2,000 of 2,000);
CUDA emu 4 of 4 packs; OpenCL emu 2 packs x 2 configurations; Apple OpenCL 4 of 4 packs
at 27.9 Mhash/s. Census 20,000: 5.225 percent rejected, accepted distinct mean 127.887.

Spec 01 0.2 (1.4.2, 1.4.3, 1.4.6, 1.11, 1.15, 1.16, 1.17), igneum-pow README, the CUDA,
OpenCL and Metal test notes, bench-log entry, ledger M5 and M6 Fixed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 07:52:40 +00:00

22 KiB

proto-metal hardening tests

Date: 3 October 2026. Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, macOS Darwin 25.6.0, Swift 5.8.1 from Command Line Tools, no Xcode, Metal shaders compiled at runtime. Build: swiftc -O -o igneum-bench main.swift -framework Metal (4 s).

Every number below was produced on this machine on this date by the commands shown.

Later on 3 October 2026 the default dataset became the memory-hard construction of MEMHARD.md. The tables in this file are from the original closed-form dataset and reproduce with --closed-form added to each command. Every test here was re-run on the new dataset with the same commands and passed; those results are in MEMHARD.md section 2.5. The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one.

On 4 October 2026 the generator became version 2 (exactly 16 loads per program, fresh-source loads, the acceptance rule of spec 01 section 1.4.6; igneum-pow/src/{generator,accept}.rs, ported into main.swift as generateProgramV2, candidateProgram, acceptProgram). Sections 1 to 7 below describe the version 1 generator and its vectors, which are retired; the loads-per-hash ranges they quote no longer occur. Section 9 holds the version 2 run. The --load-weight and --wide-frac levers select the retired version 1 generator and exist only so MEMHARD.md section 2.4 reproduces.

The tests live in main.swift next to the bench and share its generator, MSL emitter, dataset build and CPU interpreter without modification. The CPU interpreter gained an optional trace hook (cpuWarpTraced) that the bench does not use; cpuWarp calls it with nil.

What the GPU-vs-CPU comparison means: the GPU runs the generated Metal kernel, the CPU runs a plain Swift interpreter over the same instruction list, and every 64-bit output is compared bit for bit. The interpreter never reads the dataset buffer; it recomputes each element from the closed form. So agreement also covers the dataset fill kernel and the address masking, not only the arithmetic.

1. Fuzz: random programs, random warps, three dataset sizes

Command (required run):

./igneum-bench --fuzz 200

For each of N seeds: generate the program, emit MSL, static mask check, compile on the GPU, draw a dataset size from {64 MiB, 256 MiB, 1 GiB}, draw 4 base nonces uniformly from the full 32-bit range (so warps straddle 2^31 and wrap past 2^32), run the GPU, run the CPU interpreter, compare all 128 outputs. The master seed fixes the whole sequence, so any failure is reproducible with --fuzz-seed. The run also asserts the generator's contract that the emitter relies on: rotl immediate in 1..31, shuffle mask in {1,2,4,8,16}, source register never the destination.

Dataset Programs Pass Fail
2^24 words (64 MiB) 63 63 0
2^26 words (256 MiB) 64 64 0
2^28 words (1 GiB) 73 73 0
all 200 200 0

200 programs: pass 200, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. 800 warps compared (25,600 hashes). Loads per hash ranged 64 to 208. Op totals across the 200 programs: load 3204, add 1472, xor 1276, mul 1087, mad 1062, shfl 1038, rotl 892, mulhi 762, sub 742, rotr 739, or 526. Compile (library + pipeline) min 0.0 ms, avg 18.3 ms, max 46.4 ms; the zero is the system shader cache answering for the first 20 seeds, which an earlier smoke run with the same master seed had already compiled. GPU dispatch 156 ms total, CPU interpreter 16 ms total, wall 3.9 s.

Large run, separate master seed so every compile is cold:

./igneum-bench --fuzz 10000 --fuzz-seed igneum-fuzz-large-2026-10-03
Dataset Programs Pass Fail
2^24 words (64 MiB) 3346 3346 0
2^26 words (256 MiB) 3334 3334 0
2^28 words (1 GiB) 3320 3320 0
all 10000 10000 0

10,000 programs: pass 10,000, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. 40,000 warps compared (1,280,000 hashes). Loads per hash ranged 40 to 232. Op totals: load 160,123, add 76,698, xor 64,138, mad 51,577, shfl 51,298, mul 50,919, rotl 44,945, rotr 38,333, mulhi 38,284, sub 38,234, or 25,451. Compile min 13.6 ms, avg 18.5 ms, max 48.5 ms (all cold). GPU dispatch 8.2 s total, CPU interpreter 0.8 s total, wall 196 s. FUZZ: PASS.

Combined with the 200-program run: 10,200 programs, 40,800 warps, 1,305,600 hashes, zero mismatches. The op totals show every one of the 11 instruction families was exercised tens of thousands of times.

2. Edge cases

Command:

./igneum-bench --edge

Hand-built programs, dataset 1 GiB (MASK 0x0fffffff), 4 warps at base nonces 0x00000000, 0x00100000, 0x7ffffff0 and 0xffffffe0, 128 lanes per case. Operand values are forced with r = r - r (zero) followed by a constant add, then proven: the traced CPU interpreter checks the named register value just before the instruction in question, on lane 0 of all 4 warps in all 8 iterations. "Held" means 32 of 32 checks passed.

Case Instrs Loads/hash Preconditions GPU vs CPU Result
rotl immediate by 1 and by 31 6 0 none needed 128/128 lanes PASS
rotr by register == 0 3 0 held 128/128 lanes PASS
rotr by register == 32 (32 mod 32 = 0) 5 0 held 128/128 lanes PASS
rotr by register == 0xFFFFFFE0 (-32, 0 mod 32) 5 0 held 128/128 lanes PASS
rotr by register == 31, 63 and 1 10 0 held 128/128 lanes PASS
mulhi 0xFFFFFFFF x 0xFFFFFFFF (result 0xFFFFFFFE), 0x80000000 x 2 (result 1), x 0 (result 0) 15 0 held, results checked 128/128 lanes PASS
shfl_xor every mask 1..16 in sequence 16 0 none needed 128/128 lanes PASS
load at index 0 via register 0, and via register MASK+1 6 16 held 128/128 lanes PASS
load at index MASK via register MASK, and via register 0xFFFFFFFF 8 16 held 128/128 lanes PASS
add wraparound 0xFFFFFFFF + 1 (result 0) 5 0 held, result checked 128/128 lanes PASS
sub wraparound 0 - 1 (result 0xFFFFFFFF) 6 0 held, result checked 128/128 lanes PASS
mul 0xFFFFFFFF x 0xFFFFFFFF (low result 1), mad same + 5 (result 6) 15 0 held, results checked 128/128 lanes PASS
generated program with every load replaced by xor (zero loads) 64 0 none needed 128/128 lanes PASS
64 loads and nothing else 64 512 none needed 128/128 lanes PASS
rotl immediate by 0 (outside the generator's 1..31 contract) 2 0 none needed 128/128 lanes info only: agrees

EDGE: PASS, 14 of 14 counted cases. The last row is informational: the generator never emits a rotate by 0 (the fuzz run asserts this on every instruction), and the MSL rotl_imm would shift by 32 for it. On this GPU and compiler the result happened to equal the CPU's. Nothing may rely on that; the contract stays 1..31. The shuffle case covers masks 3, 5, 6, 7, 9 to 15 that the generator does not emit; simd_shuffle_xor and the interpreter's lane ^ mask agreed for all 16.

3. Output statistics

Command:

./igneum-bench --stats

Three seeds: igneum-genesis, igneum-genesis/stats1, igneum-genesis/stats2. For each, 2^20 consecutive nonces from 0 hashed on the GPU at 1 GiB (warps 0 and 32767 spot-checked against the CPU, both matched), then on the CPU side: (a) ones count per output bit, (b) avalanche from single-bit nonce flips, run on the GPU as pairs of one-warp dispatches, (c) chi-square over 65,536 buckets for each 16-bit window of the output, (d) duplicate count after sorting.

Seed Loads/hash Bit freq min..max Max bit deviation (sigma) Avalanche 1k mean / std Avalanche 16k mean / std Per-output-bit flip prob (16k) Worst chi2 z of 4 windows Dups
igneum-genesis 104 0.4990..0.5009 2.07 (bit 38) 32.07 / 4.09 32.035 / 3.99 0.492..0.507 0.42 0
igneum-genesis/stats1 128 0.4986..0.5011 2.90 (bit 3) 32.14 / 3.94 31.987 / 3.98 0.490..0.508 2.26 0
igneum-genesis/stats2 176 0.4989..0.5010 2.29 (bit 9) 31.90 / 3.97 32.035 / 4.01 0.492..0.508 1.65 0

Chi-square detail (df 65535, expected 65535, sigma 362):

Seed bits 0..15 bits 16..31 bits 32..47 bits 48..63
igneum-genesis 65681 (z 0.40) 65686 (z 0.42) 65513 (z -0.06) 65478 (z -0.16)
igneum-genesis/stats1 65163 (z -1.03) 65768 (z 0.64) 65820 (z 0.79) 66353 (z 2.26)
igneum-genesis/stats2 66132 (z 1.65) 65447 (z -0.24) 65942 (z 1.12) 65960 (z 1.17)

Reading the numbers. (a) With 2^20 samples the sigma on a bit count is 512; the largest deviation over 192 bit positions (3 seeds x 64) was 2.90 sigma, which is what 192 draws from a fair coin produce. (b) An ideal function changes 32 of 64 output bits on average with standard deviation 4. The 1,000-trial means are within 0.14 of 32 (standard error 0.13); the 16,000-trial means are within 0.035 of 32 (standard error 0.03), the standard deviations are 3.98 to 4.01, and every one of the 64 output bits flipped with probability 0.490 to 0.508 (sigma 0.004). The per-input-bit means over 500 flips each ranged 31.5 to 32.5 (sigma 0.18). An earlier 1,000-trial sample for igneum-genesis, drawn with a different RNG salt before the 16,000-trial pass was added, gave mean 31.72; the 16,000-trial figure of 32.035 shows that was sampling noise. (c) All 12 chi-square values are within 2.3 sigma of their expectation. (d) No duplicate among 2^20 outputs; the birthday expectation at 64 bits is 3e-8.

Verdict: the output looks uniform on every measure tried, for all three seeds. STATS: PASS.

This is a sanity check for obvious structural bias. It is not a proof of cryptographic strength, it says nothing about adversarially chosen inputs, and 3 seeds is not a statement about the population of programs.

4. Determinism

Command:

./igneum-bench --determinism

Seed igneum-genesis, 2^20 nonces from 0, 1 GiB dataset. The output buffer is filled with a sentinel before every run so an unwritten lane would show.

Check Result
Two independent generations of the program give identical MSL text (4,770 bytes) yes
5 GPU runs of the same pipeline, fingerprint FNV-1a over 8 MiB 933787e8cfefccb7 all 5 runs, 0 outputs differ, 0 unwritten lanes
Second compile of identical source (0.1 ms, served by the system shader cache) fingerprint identical
Third compile with a comment tag appended so the cache misses (30.0 ms, a real recompile) fingerprint identical
CPU interpreter on 8 warps of the reference run (first, last, 6 random) all match
Dataset filled twice, both blitted to shared memory and fingerprinted 0b1a77899ee60493 both times
4,096 sampled dataset words incl. indices 0, 1, MASK-1, MASK vs CPU datasetElem all match

GPU time per 2^20 batch: 33.0 ms first, 23.2 to 23.8 ms after. DETERMINISM: PASS.

5. Memory safety of dataset indexing

Command:

./igneum-bench --memcheck

Static: the generated MSL for igneum-genesis at 2^20, 2^24 and 2^28 words has 13 dataset[ accesses, all 13 of the exact form dataset[rN & MASK], and the identifier dataset appears 14 times (13 accesses plus the kernel parameter). The CUDA twin from --export-pack has 14 ds[ accesses: 13 of the form ds[rN & mask] and the fill kernel's one write, guarded by if (i < n). The same static check ran on all fuzz programs (section 1) with zero failures.

Dynamic, 4 MiB dataset (MASK 0x000fffff): four full batches of 2^20 nonces from bases 0xfff00000 (last nonce 0xffffffff), 0xffffffe0 (wraps to 0 inside the batch), 0x80000000 and 0 all completed; 4 verification warps at bases 0xffffffe0, 0xffffffff, 0x80000000, 0xfff00000 matched the CPU. In those warps 416 of 416 load indices were above MASK before masking (a random 32-bit value is below 2^20 with probability 2^-12), so the mask was exercised on every load.

Metal does not bounds-check device buffers, so "it did not crash" is weak evidence on its own. The static check is the guarantee: the emitter has exactly one load template and it masks. MEMCHECK: PASS.

6. The bench still works

Command, after all the changes above:

./igneum-bench
Seed Compile ms Mhash/s (wall) Mhash/s (GPU time) GB/s useful Loads/hash CPU verify ms/warp Verify
igneum-genesis 0.2 (shader cache) 44.56 44.59 18.54 104 0.017 PASS 3/3 warps
igneum-genesis/epoch1 0.3 (shader cache) 47.74 47.77 19.86 104 0.017 PASS 3/3 warps

Dataset fill 5.05 ms GPU (first fill of the process). OVERALL: PASS. The rates match the 3 October table in README.md (45.2 and 48.4) to within 2 percent, so the bench path is unchanged. The compile figures are the system shader cache answering for source compiled earlier today; they are not new compile measurements.

Pack export re-run to a scratch directory and compared byte for byte with the pack written before these changes (../proto-cuda/packs/igneum-genesis):

./igneum-bench --seed igneum-genesis --export-pack <scratch>/pack-genesis
diff -r <scratch>/pack-genesis ../proto-cuda/packs/igneum-genesis

Metal GPU cross-check PASS 3/3 warps, pack written, and diff -r reported no differences: all six files (kernel.cu, program.h, vectors.h, program.json, vectors.json, program.metal) are byte-identical to the pack exported before these changes.

7. Shortcut measurement: the closed-form dataset is not memory-hard

Command:

./igneum-bench --inline-dataset --hours 1

The dataset element is ds_elem(i, d0, d1), six integer operations. A miner can replace every dataset[a & MASK] with ds_elem(a & MASK, d0, d1) and never read memory. --inline-dataset emits exactly that kernel; the CPU verification still passes because the function is unchanged.

Kernel Dataset Mhash/s (wall) Mhash/s (GPU time) Verify
honest (loads from the 1 GiB buffer), this session 1 GiB 44.6 44.6 PASS 3/3
honest, README sweep of 3 October 4 MiB (cache resident) 569 not recorded PASS
inline ds_elem, no memory read 1 GiB mask 4,888 6,274 PASS 3/3
inline ds_elem, no memory read 4 MiB mask 5,646 6,024 PASS 3/3

The inline kernel runs about 110x faster than the honest kernel by wall clock (140x by GPU timestamps) and about 9x faster than the honest kernel with a cache-resident dataset. The inline figures are approximate: the 4 timed batches took 2.7 to 3.4 ms in total, so command overhead is visible in the wall figure. The GPU-time figure is consistent with the ALU peak: roughly 1,500 integer instructions per hash at 6.3 Ghash/s is about 9.4 T instructions/s across 40 cores, in the range of this part's compute throughput.

Reading: with a closed-form dataset the prototype hash is compute-bound for anyone who skips the buffer, and the honest kernel's memory traffic is voluntary. Memory-hardness has to come from a dataset element that costs more to derive than to load. RandomX gets this from a cache of roughly 256 MiB expanded by SuperscalarHash into a dataset of roughly 2 GiB (approximate, from memory; vendor/RandomX is not cloned on this machine, so there is no file citation yet). The README already listed an expensive dataset element under "What to try next"; this measurement is why it is not optional.

8. Conclusions

Demonstrated on this machine

  1. Bit-exact GPU-versus-CPU agreement of the random-program hash over 10,200 random programs (200 plus the 10,000 run), 4 random warps each across the full 32-bit nonce range, at 64 MiB, 256 MiB and 1 GiB, plus the 7 programs in README.md, plus 14 hand-built edge cases with their operand values proven. Zero mismatches, zero compile failures.
  2. Deterministic GPU execution: 5 runs and 3 compiles (one forced cold) of the same program give the same 8 MiB of output; the dataset fill is deterministic and matches the CPU closed form at sampled indices including 0 and MASK.
  3. Every dataset index in the emitted MSL and CUDA is masked, by static check on every program tested, and the mask is exercised by essentially every load.
  4. No obvious structural bias in the 64-bit output for 3 seeds at 2^20 nonces: bit frequencies, avalanche (mean 32.0, std 4.0, every output bit flips with probability 0.49 to 0.51), chi-square on four 16-bit windows, zero duplicates.
  5. The emitter's contract on the generator (rotl 1..31, masks in {1,2,4,8,16}, src != dst) held on every instruction of every fuzzed program.

Not demonstrated

  1. Cryptographic security of the construction. Nothing here speaks to preimage, second-preimage or collision resistance, or to an adversary who chooses nonces or influences the epoch seed. The per-register initialisation is splitmix32(nonce ^ seed) ^ seed, which is a bijection of the nonce per register, and the body is add-rotate-xor-multiply with or (which destroys information) and no non-linear table. None of that has been analysed.
  2. Resistance to algebraic or structural attacks on weak programs. 3 seeds were measured for bias; the population of programs was not. A program whose or chain saturates a register, or whose load addresses collapse to few values, would be both biased and shortcut-able, and nothing yet rejects such programs or bounds how often they occur.
  3. Memory-hardness. Section 7 measures the shortcut directly: with a closed-form dataset the honest kernel's memory traffic is optional. The README's "memory bound at 1 GiB" describes the honest kernel, not the fastest kernel. This is a design gap of the prototype, not a bug in the tests.
  4. Vendor independence. Only Apple Metal was run. The CUDA twin exists as text and passed a CPU emulation; it has not run on NVIDIA hardware, where __shfl_xor_sync, __umulhi and shift semantics must be shown to agree bit for bit with these vectors.
  5. The 10 ms CPU verification gate against a dataset that is expensive to derive. The current 0.02 ms per warp is with a six-operation element.

The next three tests the real cryptographer must do before the spec is final

  1. Weak-program census and rejection filter. Generate at least 10^5 programs. For each, measure on the CPU interpreter (no GPU needed): output bit bias and avalanche on 2^12 nonces, distinct load addresses per hash, fraction of registers saturated by or, and whether any register's final value is independent of the nonce. Report the distribution, define the rejection rule, and bound the advantage of a miner who can grind the epoch seed over candidate blocks.
  2. Replace the closed-form dataset with an expensive derivation (RandomX style: a cache of hundreds of MiB expanded by a slow hash, or Argon2-based) and re-run three measurements: the --inline-dataset shortcut ratio (must fall to about 1), the GPU rate at 1 GiB, and the CPU verify time against the 10 ms gate with on-demand element derivation. Also measure distinct cache lines touched per hash so a cache-resident shortcut is ruled out, not assumed.
  3. Cross-vendor bit-exactness and a written spec. Run the exported pack and the fuzz vectors on NVIDIA (RTX 5090, CUDA) and on at least one AMD GPU, with the same 4-warp-per-program comparison, including the edge-case programs. Then fix the hash as a specification with test vectors and replace the ad-hoc seed derivation (FNV-1a plus SplitMix) with a standard hash so the seed-to-program mapping is auditable.

9. Generator version 2 (4 October 2026)

Build: swiftc -O -o igneum-bench main.swift -framework Metal (7 s), same machine, idle (load 3).

Fuzz, the required run at a reduced size (the version 1 run was 10,000):

./igneum-bench --fuzz 2000 --fuzz-seed igneum-fuzz-gen2-2026-10-04
Dataset Programs Pass Fail
2^24 words (64 MiB) 647 647 0
2^26 words (256 MiB) 645 645 0
2^28 words (1 GiB) 708 708 0
all 2000 2000 0

2,000 programs: pass 2,000, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. 8,000 warps compared (256,000 hashes), memory-hard dataset. Loads per hash 128 to 128 (every program; version 1 ranged 40 to 232). Op totals: load 32,000, add 15,350, xor 12,802, mad 10,380, mul 10,208, shfl 10,147, rotl 8,983, sub 7,704, rotr 7,632, mulhi 7,596, or 5,198. Compile min 15.7 ms, avg 21.8 ms, max 67.3 ms (all cold). GPU dispatch 1.7 s total, CPU interpreter 17.7 s total (the memory-hard dataset), wall 91.4 s. Every fuzzed program is the accepted attempt of its seed, so the acceptance rule ran 2,000 times inside this run as well. FUZZ: PASS.

Cross-implementation, the Swift generator against the Rust crate (igneum-pow export for the same seeds, closed form unless stated): instruction lists, seed words and all 96 vectors identical on igneum-genesis, igneum-hourly, and the three seeds whose attempt 0 is rejected and attempt 1 accepted (igneum-census-2026-10-03/22, /37, /51); igneum-genesis memory-hard: identical instruction list and 96 vectors between the Swift export (Metal GPU cross-check PASS 3 of 3 warps, cache FNV 48c4f5bf24166b2e) and the checked-in Rust pack proto-cuda/packs/igneum-genesis-mh. The Swift exporter is no longer the pack source; it is the Metal cross-check of the Rust packs (proto-cuda/README.md, "Regenerating a pack"), and its program.json still writes format igneum-program-pack-2 without the generator fields.

Not re-run under version 2: sections 2 (edge, hand-built programs that bypass the generator), 3 (stats), 4 (determinism), 5 (memcheck) and 7 (shortcut). Nothing in them depends on how the instruction list is drawn.

Files

  • main.swift: tests under // MARK: - Hardening tests, flags --fuzz, --fuzz-seed, --edge, --stats, --determinism, --memcheck, --inline-dataset; generator version 2 and the acceptance rule under // MARK: - Generator version 2 and the acceptance rule.
  • TESTS.md: this file.
  • Raw logs of the runs quoted above were kept in the session scratchpad and are not checked in; every table is reproducible with the command above it.