igneum-pow 0.2.0: generator v2 draws exactly 16 load slots from instructions 1..63, a load's source from the registers written earlier and not read by a load since, the other 48 ops from the ten non-load weights; accept.rs is spec 01 section 1.4.6 (static: no stale load source, every register injected; dynamic: 64 units on the seed-keyed closed-form dataset, no constant bit, no lane-constant site, under 164 saturated, bias within 136 of 1024, distinct addresses above 245,760); a rejected candidate is replaced by the next attempt of the seed (seed || k_le32), 32 a consensus fault. Packs carry the generator version, attempt and program id. Version 1 kept as generate_v1 for the census. Packs: igneum-genesis, igneum-hourly, igneum-genesis-mh regenerated by igneum-pow export; new igneum-devnet-v4-epoch0 (devnet genesis hash, day bytes 20730). Checks: Rust 39 of 39 tests; Metal natively via the Swift port (export cross-check 3 of 3 warps, identical programs and vectors on five seeds incl. three with attempt 1, fuzz 2,000 of 2,000); CUDA emu 4 of 4 packs; OpenCL emu 2 packs x 2 configurations; Apple OpenCL 4 of 4 packs at 27.9 Mhash/s. Census 20,000: 5.225 percent rejected, accepted distinct mean 127.887. Spec 01 0.2 (1.4.2, 1.4.3, 1.4.6, 1.11, 1.15, 1.16, 1.17), igneum-pow README, the CUDA, OpenCL and Metal test notes, bench-log entry, ledger M5 and M6 Fixed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
368 lines
22 KiB
Markdown
368 lines
22 KiB
Markdown
# proto-metal hardening tests
|
|
|
|
Date: 3 October 2026. Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, macOS Darwin 25.6.0,
|
|
Swift 5.8.1 from Command Line Tools, no Xcode, Metal shaders compiled at runtime.
|
|
Build: `swiftc -O -o igneum-bench main.swift -framework Metal` (4 s).
|
|
|
|
Every number below was produced on this machine on this date by the commands shown.
|
|
|
|
Later on 3 October 2026 the default dataset became the memory-hard construction of `MEMHARD.md`. The tables in this
|
|
file are from the original closed-form dataset and reproduce with `--closed-form` added to each command. Every test
|
|
here was re-run on the new dataset with the same commands and passed; those results are in `MEMHARD.md` section 2.5.
|
|
The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one.
|
|
|
|
On 4 October 2026 the generator became version 2 (exactly 16 loads per program, fresh-source loads, the acceptance
|
|
rule of spec 01 section 1.4.6; `igneum-pow/src/{generator,accept}.rs`, ported into `main.swift` as
|
|
`generateProgramV2`, `candidateProgram`, `acceptProgram`). Sections 1 to 7 below describe the version 1 generator and
|
|
its vectors, which are retired; the loads-per-hash ranges they quote no longer occur. Section 9 holds the version 2
|
|
run. The `--load-weight` and `--wide-frac` levers select the retired version 1 generator and exist only so
|
|
`MEMHARD.md` section 2.4 reproduces.
|
|
|
|
The tests live in `main.swift` next to the bench and share its generator, MSL emitter, dataset build and CPU
|
|
interpreter without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench
|
|
does not use; `cpuWarp` calls it with `nil`.
|
|
|
|
What the GPU-vs-CPU comparison means: the GPU runs the generated Metal kernel, the CPU runs a plain Swift
|
|
interpreter over the same instruction list, and every 64-bit output is compared bit for bit. The interpreter
|
|
never reads the dataset buffer; it recomputes each element from the closed form. So agreement also covers the
|
|
dataset fill kernel and the address masking, not only the arithmetic.
|
|
|
|
## 1. Fuzz: random programs, random warps, three dataset sizes
|
|
|
|
Command (required run):
|
|
|
|
```
|
|
./igneum-bench --fuzz 200
|
|
```
|
|
|
|
For each of N seeds: generate the program, emit MSL, static mask check, compile on the GPU, draw a dataset size
|
|
from {64 MiB, 256 MiB, 1 GiB}, draw 4 base nonces uniformly from the full 32-bit range (so warps straddle
|
|
2^31 and wrap past 2^32), run the GPU, run the CPU interpreter, compare all 128 outputs. The master seed fixes
|
|
the whole sequence, so any failure is reproducible with `--fuzz-seed`. The run also asserts the generator's
|
|
contract that the emitter relies on: rotl immediate in 1..31, shuffle mask in {1,2,4,8,16}, source register
|
|
never the destination.
|
|
|
|
| Dataset | Programs | Pass | Fail |
|
|
|---|---|---|---|
|
|
| 2^24 words (64 MiB) | 63 | 63 | 0 |
|
|
| 2^26 words (256 MiB) | 64 | 64 | 0 |
|
|
| 2^28 words (1 GiB) | 73 | 73 | 0 |
|
|
| all | 200 | 200 | 0 |
|
|
|
|
200 programs: pass 200, mismatch 0, compile failures 0, static mask failures 0, contract failures 0.
|
|
800 warps compared (25,600 hashes). Loads per hash ranged 64 to 208. Op totals across the 200 programs:
|
|
load 3204, add 1472, xor 1276, mul 1087, mad 1062, shfl 1038, rotl 892, mulhi 762, sub 742, rotr 739, or 526.
|
|
Compile (library + pipeline) min 0.0 ms, avg 18.3 ms, max 46.4 ms; the zero is the system shader cache
|
|
answering for the first 20 seeds, which an earlier smoke run with the same master seed had already compiled.
|
|
GPU dispatch 156 ms total, CPU interpreter 16 ms total, wall 3.9 s.
|
|
|
|
Large run, separate master seed so every compile is cold:
|
|
|
|
```
|
|
./igneum-bench --fuzz 10000 --fuzz-seed igneum-fuzz-large-2026-10-03
|
|
```
|
|
|
|
| Dataset | Programs | Pass | Fail |
|
|
|---|---|---|---|
|
|
| 2^24 words (64 MiB) | 3346 | 3346 | 0 |
|
|
| 2^26 words (256 MiB) | 3334 | 3334 | 0 |
|
|
| 2^28 words (1 GiB) | 3320 | 3320 | 0 |
|
|
| all | 10000 | 10000 | 0 |
|
|
|
|
10,000 programs: pass 10,000, mismatch 0, compile failures 0, static mask failures 0, contract failures 0.
|
|
40,000 warps compared (1,280,000 hashes). Loads per hash ranged 40 to 232. Op totals: load 160,123,
|
|
add 76,698, xor 64,138, mad 51,577, shfl 51,298, mul 50,919, rotl 44,945, rotr 38,333, mulhi 38,284,
|
|
sub 38,234, or 25,451. Compile min 13.6 ms, avg 18.5 ms, max 48.5 ms (all cold). GPU dispatch 8.2 s total,
|
|
CPU interpreter 0.8 s total, wall 196 s. FUZZ: PASS.
|
|
|
|
Combined with the 200-program run: 10,200 programs, 40,800 warps, 1,305,600 hashes, zero mismatches.
|
|
The op totals show every one of the 11 instruction families was exercised tens of thousands of times.
|
|
|
|
## 2. Edge cases
|
|
|
|
Command:
|
|
|
|
```
|
|
./igneum-bench --edge
|
|
```
|
|
|
|
Hand-built programs, dataset 1 GiB (MASK 0x0fffffff), 4 warps at base nonces 0x00000000, 0x00100000,
|
|
0x7ffffff0 and 0xffffffe0, 128 lanes per case. Operand values are forced with `r = r - r` (zero) followed by a
|
|
constant add, then proven: the traced CPU interpreter checks the named register value just before the
|
|
instruction in question, on lane 0 of all 4 warps in all 8 iterations. "Held" means 32 of 32 checks passed.
|
|
|
|
| Case | Instrs | Loads/hash | Preconditions | GPU vs CPU | Result |
|
|
|---|---|---|---|---|---|
|
|
| rotl immediate by 1 and by 31 | 6 | 0 | none needed | 128/128 lanes | PASS |
|
|
| rotr by register == 0 | 3 | 0 | held | 128/128 lanes | PASS |
|
|
| rotr by register == 32 (32 mod 32 = 0) | 5 | 0 | held | 128/128 lanes | PASS |
|
|
| rotr by register == 0xFFFFFFE0 (-32, 0 mod 32) | 5 | 0 | held | 128/128 lanes | PASS |
|
|
| rotr by register == 31, 63 and 1 | 10 | 0 | held | 128/128 lanes | PASS |
|
|
| mulhi 0xFFFFFFFF x 0xFFFFFFFF (result 0xFFFFFFFE), 0x80000000 x 2 (result 1), x 0 (result 0) | 15 | 0 | held, results checked | 128/128 lanes | PASS |
|
|
| shfl_xor every mask 1..16 in sequence | 16 | 0 | none needed | 128/128 lanes | PASS |
|
|
| load at index 0 via register 0, and via register MASK+1 | 6 | 16 | held | 128/128 lanes | PASS |
|
|
| load at index MASK via register MASK, and via register 0xFFFFFFFF | 8 | 16 | held | 128/128 lanes | PASS |
|
|
| add wraparound 0xFFFFFFFF + 1 (result 0) | 5 | 0 | held, result checked | 128/128 lanes | PASS |
|
|
| sub wraparound 0 - 1 (result 0xFFFFFFFF) | 6 | 0 | held, result checked | 128/128 lanes | PASS |
|
|
| mul 0xFFFFFFFF x 0xFFFFFFFF (low result 1), mad same + 5 (result 6) | 15 | 0 | held, results checked | 128/128 lanes | PASS |
|
|
| generated program with every load replaced by xor (zero loads) | 64 | 0 | none needed | 128/128 lanes | PASS |
|
|
| 64 loads and nothing else | 64 | 512 | none needed | 128/128 lanes | PASS |
|
|
| rotl immediate by 0 (outside the generator's 1..31 contract) | 2 | 0 | none needed | 128/128 lanes | info only: agrees |
|
|
|
|
EDGE: PASS, 14 of 14 counted cases. The last row is informational: the generator never emits a rotate by 0
|
|
(the fuzz run asserts this on every instruction), and the MSL `rotl_imm` would shift by 32 for it. On this GPU
|
|
and compiler the result happened to equal the CPU's. Nothing may rely on that; the contract stays 1..31.
|
|
The shuffle case covers masks 3, 5, 6, 7, 9 to 15 that the generator does not emit; `simd_shuffle_xor` and the
|
|
interpreter's `lane ^ mask` agreed for all 16.
|
|
|
|
## 3. Output statistics
|
|
|
|
Command:
|
|
|
|
```
|
|
./igneum-bench --stats
|
|
```
|
|
|
|
Three seeds: `igneum-genesis`, `igneum-genesis/stats1`, `igneum-genesis/stats2`. For each, 2^20 consecutive
|
|
nonces from 0 hashed on the GPU at 1 GiB (warps 0 and 32767 spot-checked against the CPU, both matched), then
|
|
on the CPU side: (a) ones count per output bit, (b) avalanche from single-bit nonce flips, run on the GPU as
|
|
pairs of one-warp dispatches, (c) chi-square over 65,536 buckets for each 16-bit window of the output,
|
|
(d) duplicate count after sorting.
|
|
|
|
| Seed | Loads/hash | Bit freq min..max | Max bit deviation (sigma) | Avalanche 1k mean / std | Avalanche 16k mean / std | Per-output-bit flip prob (16k) | Worst chi2 z of 4 windows | Dups |
|
|
|---|---|---|---|---|---|---|---|---|
|
|
| igneum-genesis | 104 | 0.4990..0.5009 | 2.07 (bit 38) | 32.07 / 4.09 | 32.035 / 3.99 | 0.492..0.507 | 0.42 | 0 |
|
|
| igneum-genesis/stats1 | 128 | 0.4986..0.5011 | 2.90 (bit 3) | 32.14 / 3.94 | 31.987 / 3.98 | 0.490..0.508 | 2.26 | 0 |
|
|
| igneum-genesis/stats2 | 176 | 0.4989..0.5010 | 2.29 (bit 9) | 31.90 / 3.97 | 32.035 / 4.01 | 0.492..0.508 | 1.65 | 0 |
|
|
|
|
Chi-square detail (df 65535, expected 65535, sigma 362):
|
|
|
|
| Seed | bits 0..15 | bits 16..31 | bits 32..47 | bits 48..63 |
|
|
|---|---|---|---|---|
|
|
| igneum-genesis | 65681 (z 0.40) | 65686 (z 0.42) | 65513 (z -0.06) | 65478 (z -0.16) |
|
|
| igneum-genesis/stats1 | 65163 (z -1.03) | 65768 (z 0.64) | 65820 (z 0.79) | 66353 (z 2.26) |
|
|
| igneum-genesis/stats2 | 66132 (z 1.65) | 65447 (z -0.24) | 65942 (z 1.12) | 65960 (z 1.17) |
|
|
|
|
Reading the numbers. (a) With 2^20 samples the sigma on a bit count is 512; the largest deviation over 192
|
|
bit positions (3 seeds x 64) was 2.90 sigma, which is what 192 draws from a fair coin produce. (b) An ideal
|
|
function changes 32 of 64 output bits on average with standard deviation 4. The 1,000-trial means are within
|
|
0.14 of 32 (standard error 0.13); the 16,000-trial means are within 0.035 of 32 (standard error 0.03), the
|
|
standard deviations are 3.98 to 4.01, and every one of the 64 output bits flipped with probability 0.490 to
|
|
0.508 (sigma 0.004). The per-input-bit means over 500 flips each ranged 31.5 to 32.5 (sigma 0.18). An
|
|
earlier 1,000-trial sample for igneum-genesis, drawn with a different RNG salt before the 16,000-trial pass
|
|
was added, gave mean 31.72; the 16,000-trial figure of 32.035 shows that was sampling noise. (c) All 12
|
|
chi-square values are within 2.3 sigma of their expectation. (d) No duplicate among 2^20 outputs; the birthday
|
|
expectation at 64 bits is 3e-8.
|
|
|
|
Verdict: the output looks uniform on every measure tried, for all three seeds. STATS: PASS.
|
|
|
|
This is a sanity check for obvious structural bias. It is not a proof of cryptographic strength, it says
|
|
nothing about adversarially chosen inputs, and 3 seeds is not a statement about the population of programs.
|
|
|
|
## 4. Determinism
|
|
|
|
Command:
|
|
|
|
```
|
|
./igneum-bench --determinism
|
|
```
|
|
|
|
Seed igneum-genesis, 2^20 nonces from 0, 1 GiB dataset. The output buffer is filled with a sentinel before
|
|
every run so an unwritten lane would show.
|
|
|
|
| Check | Result |
|
|
|---|---|
|
|
| Two independent generations of the program give identical MSL text (4,770 bytes) | yes |
|
|
| 5 GPU runs of the same pipeline, fingerprint FNV-1a over 8 MiB | 933787e8cfefccb7 all 5 runs, 0 outputs differ, 0 unwritten lanes |
|
|
| Second compile of identical source (0.1 ms, served by the system shader cache) | fingerprint identical |
|
|
| Third compile with a comment tag appended so the cache misses (30.0 ms, a real recompile) | fingerprint identical |
|
|
| CPU interpreter on 8 warps of the reference run (first, last, 6 random) | all match |
|
|
| Dataset filled twice, both blitted to shared memory and fingerprinted | 0b1a77899ee60493 both times |
|
|
| 4,096 sampled dataset words incl. indices 0, 1, MASK-1, MASK vs CPU `datasetElem` | all match |
|
|
|
|
GPU time per 2^20 batch: 33.0 ms first, 23.2 to 23.8 ms after. DETERMINISM: PASS.
|
|
|
|
## 5. Memory safety of dataset indexing
|
|
|
|
Command:
|
|
|
|
```
|
|
./igneum-bench --memcheck
|
|
```
|
|
|
|
Static: the generated MSL for igneum-genesis at 2^20, 2^24 and 2^28 words has 13 `dataset[` accesses, all 13
|
|
of the exact form `dataset[rN & MASK]`, and the identifier `dataset` appears 14 times (13 accesses plus the
|
|
kernel parameter). The CUDA twin from `--export-pack` has 14 `ds[` accesses: 13 of the form `ds[rN & mask]`
|
|
and the fill kernel's one write, guarded by `if (i < n)`. The same static check ran on all fuzz programs
|
|
(section 1) with zero failures.
|
|
|
|
Dynamic, 4 MiB dataset (MASK 0x000fffff): four full batches of 2^20 nonces from bases 0xfff00000 (last nonce
|
|
0xffffffff), 0xffffffe0 (wraps to 0 inside the batch), 0x80000000 and 0 all completed; 4 verification warps at
|
|
bases 0xffffffe0, 0xffffffff, 0x80000000, 0xfff00000 matched the CPU. In those warps 416 of 416 load indices
|
|
were above MASK before masking (a random 32-bit value is below 2^20 with probability 2^-12), so the mask was
|
|
exercised on every load.
|
|
|
|
Metal does not bounds-check device buffers, so "it did not crash" is weak evidence on its own. The static check
|
|
is the guarantee: the emitter has exactly one load template and it masks. MEMCHECK: PASS.
|
|
|
|
## 6. The bench still works
|
|
|
|
Command, after all the changes above:
|
|
|
|
```
|
|
./igneum-bench
|
|
```
|
|
|
|
| Seed | Compile ms | Mhash/s (wall) | Mhash/s (GPU time) | GB/s useful | Loads/hash | CPU verify ms/warp | Verify |
|
|
|---|---|---|---|---|---|---|---|
|
|
| igneum-genesis | 0.2 (shader cache) | 44.56 | 44.59 | 18.54 | 104 | 0.017 | PASS 3/3 warps |
|
|
| igneum-genesis/epoch1 | 0.3 (shader cache) | 47.74 | 47.77 | 19.86 | 104 | 0.017 | PASS 3/3 warps |
|
|
|
|
Dataset fill 5.05 ms GPU (first fill of the process). OVERALL: PASS. The rates match the 3 October table in
|
|
README.md (45.2 and 48.4) to within 2 percent, so the bench path is unchanged. The compile figures are the
|
|
system shader cache answering for source compiled earlier today; they are not new compile measurements.
|
|
|
|
Pack export re-run to a scratch directory and compared byte for byte with the pack written before these
|
|
changes (`../proto-cuda/packs/igneum-genesis`):
|
|
|
|
```
|
|
./igneum-bench --seed igneum-genesis --export-pack <scratch>/pack-genesis
|
|
diff -r <scratch>/pack-genesis ../proto-cuda/packs/igneum-genesis
|
|
```
|
|
|
|
Metal GPU cross-check PASS 3/3 warps, pack written, and `diff -r` reported no differences: all six files
|
|
(`kernel.cu`, `program.h`, `vectors.h`, `program.json`, `vectors.json`, `program.metal`) are byte-identical
|
|
to the pack exported before these changes.
|
|
|
|
## 7. Shortcut measurement: the closed-form dataset is not memory-hard
|
|
|
|
Command:
|
|
|
|
```
|
|
./igneum-bench --inline-dataset --hours 1
|
|
```
|
|
|
|
The dataset element is `ds_elem(i, d0, d1)`, six integer operations. A miner can replace every
|
|
`dataset[a & MASK]` with `ds_elem(a & MASK, d0, d1)` and never read memory. `--inline-dataset` emits exactly
|
|
that kernel; the CPU verification still passes because the function is unchanged.
|
|
|
|
| Kernel | Dataset | Mhash/s (wall) | Mhash/s (GPU time) | Verify |
|
|
|---|---|---|---|---|
|
|
| honest (loads from the 1 GiB buffer), this session | 1 GiB | 44.6 | 44.6 | PASS 3/3 |
|
|
| honest, README sweep of 3 October | 4 MiB (cache resident) | 569 | not recorded | PASS |
|
|
| inline ds_elem, no memory read | 1 GiB mask | 4,888 | 6,274 | PASS 3/3 |
|
|
| inline ds_elem, no memory read | 4 MiB mask | 5,646 | 6,024 | PASS 3/3 |
|
|
|
|
The inline kernel runs about 110x faster than the honest kernel by wall clock (140x by GPU timestamps) and
|
|
about 9x faster than the honest kernel with a cache-resident dataset. The inline figures are approximate: the
|
|
4 timed batches took 2.7 to 3.4 ms in total, so command overhead is visible in the wall figure. The GPU-time
|
|
figure is consistent with the ALU peak: roughly 1,500 integer instructions per hash at 6.3 Ghash/s is about
|
|
9.4 T instructions/s across 40 cores, in the range of this part's compute throughput.
|
|
|
|
Reading: with a closed-form dataset the prototype hash is compute-bound for anyone who skips the buffer, and
|
|
the honest kernel's memory traffic is voluntary. Memory-hardness has to come from a dataset element that
|
|
costs more to derive than to load. RandomX gets this from a cache of roughly 256 MiB expanded by SuperscalarHash into a
|
|
dataset of roughly 2 GiB (approximate, from memory; `vendor/RandomX` is not cloned on this machine, so there
|
|
is no file citation yet). The README already listed an expensive dataset element under "What to try next"; this
|
|
measurement is why it is not optional.
|
|
|
|
## 8. Conclusions
|
|
|
|
### Demonstrated on this machine
|
|
|
|
1. Bit-exact GPU-versus-CPU agreement of the random-program hash over 10,200 random programs (200 plus
|
|
the 10,000 run), 4 random warps each across the full 32-bit nonce range, at 64 MiB, 256 MiB and 1 GiB, plus
|
|
the 7 programs in README.md, plus 14 hand-built edge cases with their operand values proven. Zero mismatches,
|
|
zero compile failures.
|
|
2. Deterministic GPU execution: 5 runs and 3 compiles (one forced cold) of the same program give the same
|
|
8 MiB of output; the dataset fill is deterministic and matches the CPU closed form at sampled indices
|
|
including 0 and MASK.
|
|
3. Every dataset index in the emitted MSL and CUDA is masked, by static check on every program tested, and
|
|
the mask is exercised by essentially every load.
|
|
4. No obvious structural bias in the 64-bit output for 3 seeds at 2^20 nonces: bit frequencies, avalanche
|
|
(mean 32.0, std 4.0, every output bit flips with probability 0.49 to 0.51), chi-square on four 16-bit
|
|
windows, zero duplicates.
|
|
5. The emitter's contract on the generator (rotl 1..31, masks in {1,2,4,8,16}, src != dst) held on every
|
|
instruction of every fuzzed program.
|
|
|
|
### Not demonstrated
|
|
|
|
1. Cryptographic security of the construction. Nothing here speaks to preimage, second-preimage or collision
|
|
resistance, or to an adversary who chooses nonces or influences the epoch seed. The per-register
|
|
initialisation is `splitmix32(nonce ^ seed) ^ seed`, which is a bijection of the nonce per register, and
|
|
the body is add-rotate-xor-multiply with `or` (which destroys information) and no non-linear table. None
|
|
of that has been analysed.
|
|
2. Resistance to algebraic or structural attacks on weak programs. 3 seeds were measured for bias; the
|
|
population of programs was not. A program whose `or` chain saturates a register, or whose load addresses
|
|
collapse to few values, would be both biased and shortcut-able, and nothing yet rejects such programs or
|
|
bounds how often they occur.
|
|
3. Memory-hardness. Section 7 measures the shortcut directly: with a closed-form dataset the honest kernel's
|
|
memory traffic is optional. The README's "memory bound at 1 GiB" describes the honest kernel, not the
|
|
fastest kernel. This is a design gap of the prototype, not a bug in the tests.
|
|
4. Vendor independence. Only Apple Metal was run. The CUDA twin exists as text and passed a CPU emulation;
|
|
it has not run on NVIDIA hardware, where `__shfl_xor_sync`, `__umulhi` and shift semantics must be shown
|
|
to agree bit for bit with these vectors.
|
|
5. The 10 ms CPU verification gate against a dataset that is expensive to derive. The current 0.02 ms per warp
|
|
is with a six-operation element.
|
|
|
|
### The next three tests the real cryptographer must do before the spec is final
|
|
|
|
1. Weak-program census and rejection filter. Generate at least 10^5 programs. For each, measure on the CPU
|
|
interpreter (no GPU needed): output bit bias and avalanche on 2^12 nonces, distinct load addresses per
|
|
hash, fraction of registers saturated by `or`, and whether any register's final value is independent of
|
|
the nonce. Report the distribution, define the rejection rule, and bound the advantage of a miner who can
|
|
grind the epoch seed over candidate blocks.
|
|
2. Replace the closed-form dataset with an expensive derivation (RandomX style: a cache of hundreds of MiB
|
|
expanded by a slow hash, or Argon2-based) and re-run three measurements: the `--inline-dataset` shortcut
|
|
ratio (must fall to about 1), the GPU rate at 1 GiB, and the CPU verify time against the 10 ms gate with
|
|
on-demand element derivation. Also measure distinct cache lines touched per hash so a cache-resident
|
|
shortcut is ruled out, not assumed.
|
|
3. Cross-vendor bit-exactness and a written spec. Run the exported pack and the fuzz vectors on NVIDIA
|
|
(RTX 5090, CUDA) and on at least one AMD GPU, with the same 4-warp-per-program comparison, including the
|
|
edge-case programs. Then fix the hash as a specification with test vectors and replace the ad-hoc seed
|
|
derivation (FNV-1a plus SplitMix) with a standard hash so the seed-to-program mapping is auditable.
|
|
|
|
## 9. Generator version 2 (4 October 2026)
|
|
|
|
Build: `swiftc -O -o igneum-bench main.swift -framework Metal` (7 s), same machine, idle (load 3).
|
|
|
|
Fuzz, the required run at a reduced size (the version 1 run was 10,000):
|
|
|
|
```
|
|
./igneum-bench --fuzz 2000 --fuzz-seed igneum-fuzz-gen2-2026-10-04
|
|
```
|
|
|
|
| Dataset | Programs | Pass | Fail |
|
|
|---|---|---|---|
|
|
| 2^24 words (64 MiB) | 647 | 647 | 0 |
|
|
| 2^26 words (256 MiB) | 645 | 645 | 0 |
|
|
| 2^28 words (1 GiB) | 708 | 708 | 0 |
|
|
| all | 2000 | 2000 | 0 |
|
|
|
|
2,000 programs: pass 2,000, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. 8,000 warps
|
|
compared (256,000 hashes), memory-hard dataset. Loads per hash 128 to 128 (every program; version 1 ranged 40 to
|
|
232). Op totals: load 32,000, add 15,350, xor 12,802, mad 10,380, mul 10,208, shfl 10,147, rotl 8,983, sub 7,704,
|
|
rotr 7,632, mulhi 7,596, or 5,198. Compile min 15.7 ms, avg 21.8 ms, max 67.3 ms (all cold). GPU dispatch 1.7 s
|
|
total, CPU interpreter 17.7 s total (the memory-hard dataset), wall 91.4 s. Every fuzzed program is the accepted
|
|
attempt of its seed, so the acceptance rule ran 2,000 times inside this run as well. FUZZ: PASS.
|
|
|
|
Cross-implementation, the Swift generator against the Rust crate (`igneum-pow export` for the same seeds, closed
|
|
form unless stated): instruction lists, seed words and all 96 vectors identical on `igneum-genesis`,
|
|
`igneum-hourly`, and the three seeds whose attempt 0 is rejected and attempt 1 accepted
|
|
(`igneum-census-2026-10-03/22`, `/37`, `/51`); `igneum-genesis` memory-hard: identical instruction list and 96
|
|
vectors between the Swift export (Metal GPU cross-check PASS 3 of 3 warps, cache FNV `48c4f5bf24166b2e`) and the
|
|
checked-in Rust pack `proto-cuda/packs/igneum-genesis-mh`. The Swift exporter is no longer the pack source; it is
|
|
the Metal cross-check of the Rust packs (`proto-cuda/README.md`, "Regenerating a pack"), and its `program.json`
|
|
still writes format `igneum-program-pack-2` without the generator fields.
|
|
|
|
Not re-run under version 2: sections 2 (edge, hand-built programs that bypass the generator), 3 (stats), 4
|
|
(determinism), 5 (memcheck) and 7 (shortcut). Nothing in them depends on how the instruction list is drawn.
|
|
|
|
## Files
|
|
|
|
- `main.swift`: tests under `// MARK: - Hardening tests`, flags `--fuzz`, `--fuzz-seed`, `--edge`, `--stats`,
|
|
`--determinism`, `--memcheck`, `--inline-dataset`; generator version 2 and the acceptance rule under
|
|
`// MARK: - Generator version 2 and the acceptance rule`.
|
|
- `TESTS.md`: this file.
|
|
- Raw logs of the runs quoted above were kept in the session scratchpad and are not checked in; every table
|
|
is reproducible with the command above it.
|