diff --git a/docs/bench-log.md b/docs/bench-log.md index 82e634cb..fca790e4 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -58,3 +58,15 @@ D: honest doubling is under-weighted 28 to 31 days; old miners lock alone for 20 E: 30% churn leaves 71% live, lock never lost. Threshold is 1/3: 35% stalls 1 day, 50% stalls 10 to 11 days. Active-24h total removes every stall. F: 51% patient owner holds 51.0% weight from day 45, vetoes from day 7 to 20, never locks alone. 67% owner locks alone from day 34 to 44. Recommend: f=1, cap 2x, window 30 (the window is the defence, the cap is worth 1 to 8 days), total = all keys with weight in the window. Details in sim/results.md. + +## 2026-10-03 proto-metal hardening tests (correctness and soundness of the lottery hash, Metal only) + +Machine: the same Apple M5 Max. Added `--fuzz`, `--edge`, `--stats`, `--determinism`, `--memcheck`, `--inline-dataset` to `proto-metal/main.swift`. Full tables and commands in `proto-metal/TESTS.md`. +Fuzz: 10,200 random programs (200 + 10,000, cold compiles 13.6 to 48.5 ms), 4 random full-range warps each, dataset drawn from 64 MiB / 256 MiB / 1 GiB: 40,800 warps, 1,305,600 hashes, 0 mismatches, 0 compile failures, 0 static mask failures, generator contract (rotl 1..31, mask in {1,2,4,8,16}, src != dst) held on every instruction. 196 s for the 10,000 run. +Edge: 14 hand-built cases (rotr by register 0 / 32 / -32 / 31 / 63, rotl 1 and 31, mulhi max operands, shfl masks 1..16, loads at index 0 and MASK via in-range and out-of-range registers, add/sub/mul/mad wraparound, zero loads, 64 loads) with operand values proven by a traced interpreter: 14/14 PASS, 128/128 lanes each. rotl by 0 (never generated) agreed too, recorded as informational only. +Stats (3 seeds, 2^20 nonces each): bit frequency max deviation 2.90 sigma over 192 bit positions; avalanche 16,000 flips mean 31.99 to 32.04 (expect 32), std 3.98 to 4.01 (expect 4), every output bit flips with probability 0.490 to 0.508; chi-square on four 16-bit windows all within 2.3 sigma; 0 duplicates. Looks uniform. Not a security proof. +Determinism: 5 runs and 3 compiles (one forced cold, 30 ms) of 2^20 hashes gave fingerprint 933787e8cfefccb7 every time; dataset fill deterministic (0b1a77899ee60493 twice) and 4,096 sampled words incl. 0 and MASK match the CPU closed form. +Memcheck: every `dataset[` in the MSL is `dataset[rN & MASK]` (13/13 at 3 sizes), CUDA twin 13/13 plus one guarded fill write; 4 MiB run with nonces up to 0xffffffff completed and 4 wrapping warps matched the CPU; 416/416 load indices exceeded MASK before masking. +Bench re-run after the changes: igneum-genesis 44.56 Mhash/s, epoch1 47.74 Mhash/s at 1 GiB, PASS 3/3 warps each (within 2 percent of the first-run table). `--export-pack igneum-genesis` re-run is byte-identical to the existing pack. +SHORTCUT MEASURED: `--inline-dataset` replaces every load with the six-op closed form ds_elem and never reads memory: 4,888 Mhash/s wall (6,274 GPU time) vs 44.6 honest at 1 GiB, about 110x, and about 9x the cache-resident honest rate. With a closed-form dataset the hash is not memory-hard; an expensive dataset derivation is required, not optional. +Not demonstrated: cryptographic strength, weak-program frequency and rejection, NVIDIA/AMD bit-exactness (CUDA run still pending), CPU verify gate with an expensive dataset element. Next three tests for the cryptographer are listed in TESTS.md section 8. diff --git a/proto-metal/README.md b/proto-metal/README.md index e1c5c6ab..f9089e38 100644 --- a/proto-metal/README.md +++ b/proto-metal/README.md @@ -42,6 +42,11 @@ Flags: `--seed`, `--day`, `--hours`, `--batch-log2` (default 22), `--batches` (d `--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump `, `--export-pack `. Exit code 0 means every verified warp matched. +Hardening tests (added 3 October 2026, results and commands in `TESTS.md`): `--fuzz N [--fuzz-seed ]`, +`--edge`, `--stats`, `--determinism`, `--memcheck`. Any of these runs instead of the bench; several may be combined +in one invocation; exit code 0 only if every selected test passed. `--inline-dataset` is a bench variant that +computes every dataset element inline instead of loading it (the shortcut measurement in `TESTS.md` section 7). + `--export-pack ` (added 3 October 2026) does not run the bench. It generates the program for `--seed`, emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (base nonces 0, 4096, 1000000) with the CPU interpreter, cross-checks them on the Metal GPU, and writes `kernel.cu`, `program.h`, `vectors.h`, @@ -104,6 +109,10 @@ like the system shader cache. being measured. Batches of 2^22 nonces take 86 to 118 ms each. 8. Only Apple silicon was measured. Nothing here says anything about NVIDIA or AMD, where SIMD width, cache line size and memory latency differ. +9. Measured later on 3 October 2026 (`TESTS.md` section 7): because the dataset element is a six-operation + closed form, a kernel that computes it inline instead of loading it runs at about 4,900 Mhash/s against + 44.6 for the honest kernel at 1 GiB, roughly 110x. Observation 1 describes the honest kernel only. The + prototype is not memory-hard until the dataset element costs more to derive than to load. ## What to try next diff --git a/proto-metal/TESTS.md b/proto-metal/TESTS.md new file mode 100644 index 00000000..7c4ef0ae --- /dev/null +++ b/proto-metal/TESTS.md @@ -0,0 +1,318 @@ +# proto-metal hardening tests + +Date: 3 October 2026. Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, macOS Darwin 25.6.0, +Swift 5.8.1 from Command Line Tools, no Xcode, Metal shaders compiled at runtime. +Build: `swiftc -O -o igneum-bench main.swift -framework Metal` (4 s). + +Every number below was produced on this machine on this date by the commands shown. The tests live in +`main.swift` next to the bench and share its generator, MSL emitter, dataset fill and CPU interpreter +without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench +does not use; `cpuWarp` calls it with `nil`. + +What the GPU-vs-CPU comparison means: the GPU runs the generated Metal kernel, the CPU runs a plain Swift +interpreter over the same instruction list, and every 64-bit output is compared bit for bit. The interpreter +never reads the dataset buffer; it recomputes each element from the closed form. So agreement also covers the +dataset fill kernel and the address masking, not only the arithmetic. + +## 1. Fuzz: random programs, random warps, three dataset sizes + +Command (required run): + +``` +./igneum-bench --fuzz 200 +``` + +For each of N seeds: generate the program, emit MSL, static mask check, compile on the GPU, draw a dataset size +from {64 MiB, 256 MiB, 1 GiB}, draw 4 base nonces uniformly from the full 32-bit range (so warps straddle +2^31 and wrap past 2^32), run the GPU, run the CPU interpreter, compare all 128 outputs. The master seed fixes +the whole sequence, so any failure is reproducible with `--fuzz-seed`. The run also asserts the generator's +contract that the emitter relies on: rotl immediate in 1..31, shuffle mask in {1,2,4,8,16}, source register +never the destination. + +| Dataset | Programs | Pass | Fail | +|---|---|---|---| +| 2^24 words (64 MiB) | 63 | 63 | 0 | +| 2^26 words (256 MiB) | 64 | 64 | 0 | +| 2^28 words (1 GiB) | 73 | 73 | 0 | +| all | 200 | 200 | 0 | + +200 programs: pass 200, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. +800 warps compared (25,600 hashes). Loads per hash ranged 64 to 208. Op totals across the 200 programs: +load 3204, add 1472, xor 1276, mul 1087, mad 1062, shfl 1038, rotl 892, mulhi 762, sub 742, rotr 739, or 526. +Compile (library + pipeline) min 0.0 ms, avg 18.3 ms, max 46.4 ms; the zero is the system shader cache +answering for the first 20 seeds, which an earlier smoke run with the same master seed had already compiled. +GPU dispatch 156 ms total, CPU interpreter 16 ms total, wall 3.9 s. + +Large run, separate master seed so every compile is cold: + +``` +./igneum-bench --fuzz 10000 --fuzz-seed igneum-fuzz-large-2026-10-03 +``` + +| Dataset | Programs | Pass | Fail | +|---|---|---|---| +| 2^24 words (64 MiB) | 3346 | 3346 | 0 | +| 2^26 words (256 MiB) | 3334 | 3334 | 0 | +| 2^28 words (1 GiB) | 3320 | 3320 | 0 | +| all | 10000 | 10000 | 0 | + +10,000 programs: pass 10,000, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. +40,000 warps compared (1,280,000 hashes). Loads per hash ranged 40 to 232. Op totals: load 160,123, +add 76,698, xor 64,138, mad 51,577, shfl 51,298, mul 50,919, rotl 44,945, rotr 38,333, mulhi 38,284, +sub 38,234, or 25,451. Compile min 13.6 ms, avg 18.5 ms, max 48.5 ms (all cold). GPU dispatch 8.2 s total, +CPU interpreter 0.8 s total, wall 196 s. FUZZ: PASS. + +Combined with the 200-program run: 10,200 programs, 40,800 warps, 1,305,600 hashes, zero mismatches. +The op totals show every one of the 11 instruction families was exercised tens of thousands of times. + +## 2. Edge cases + +Command: + +``` +./igneum-bench --edge +``` + +Hand-built programs, dataset 1 GiB (MASK 0x0fffffff), 4 warps at base nonces 0x00000000, 0x00100000, +0x7ffffff0 and 0xffffffe0, 128 lanes per case. Operand values are forced with `r = r - r` (zero) followed by a +constant add, then proven: the traced CPU interpreter checks the named register value just before the +instruction in question, on lane 0 of all 4 warps in all 8 iterations. "Held" means 32 of 32 checks passed. + +| Case | Instrs | Loads/hash | Preconditions | GPU vs CPU | Result | +|---|---|---|---|---|---| +| rotl immediate by 1 and by 31 | 6 | 0 | none needed | 128/128 lanes | PASS | +| rotr by register == 0 | 3 | 0 | held | 128/128 lanes | PASS | +| rotr by register == 32 (32 mod 32 = 0) | 5 | 0 | held | 128/128 lanes | PASS | +| rotr by register == 0xFFFFFFE0 (-32, 0 mod 32) | 5 | 0 | held | 128/128 lanes | PASS | +| rotr by register == 31, 63 and 1 | 10 | 0 | held | 128/128 lanes | PASS | +| mulhi 0xFFFFFFFF x 0xFFFFFFFF (result 0xFFFFFFFE), 0x80000000 x 2 (result 1), x 0 (result 0) | 15 | 0 | held, results checked | 128/128 lanes | PASS | +| shfl_xor every mask 1..16 in sequence | 16 | 0 | none needed | 128/128 lanes | PASS | +| load at index 0 via register 0, and via register MASK+1 | 6 | 16 | held | 128/128 lanes | PASS | +| load at index MASK via register MASK, and via register 0xFFFFFFFF | 8 | 16 | held | 128/128 lanes | PASS | +| add wraparound 0xFFFFFFFF + 1 (result 0) | 5 | 0 | held, result checked | 128/128 lanes | PASS | +| sub wraparound 0 - 1 (result 0xFFFFFFFF) | 6 | 0 | held, result checked | 128/128 lanes | PASS | +| mul 0xFFFFFFFF x 0xFFFFFFFF (low result 1), mad same + 5 (result 6) | 15 | 0 | held, results checked | 128/128 lanes | PASS | +| generated program with every load replaced by xor (zero loads) | 64 | 0 | none needed | 128/128 lanes | PASS | +| 64 loads and nothing else | 64 | 512 | none needed | 128/128 lanes | PASS | +| rotl immediate by 0 (outside the generator's 1..31 contract) | 2 | 0 | none needed | 128/128 lanes | info only: agrees | + +EDGE: PASS, 14 of 14 counted cases. The last row is informational: the generator never emits a rotate by 0 +(the fuzz run asserts this on every instruction), and the MSL `rotl_imm` would shift by 32 for it. On this GPU +and compiler the result happened to equal the CPU's. Nothing may rely on that; the contract stays 1..31. +The shuffle case covers masks 3, 5, 6, 7, 9 to 15 that the generator does not emit; `simd_shuffle_xor` and the +interpreter's `lane ^ mask` agreed for all 16. + +## 3. Output statistics + +Command: + +``` +./igneum-bench --stats +``` + +Three seeds: `igneum-genesis`, `igneum-genesis/stats1`, `igneum-genesis/stats2`. For each, 2^20 consecutive +nonces from 0 hashed on the GPU at 1 GiB (warps 0 and 32767 spot-checked against the CPU, both matched), then +on the CPU side: (a) ones count per output bit, (b) avalanche from single-bit nonce flips, run on the GPU as +pairs of one-warp dispatches, (c) chi-square over 65,536 buckets for each 16-bit window of the output, +(d) duplicate count after sorting. + +| Seed | Loads/hash | Bit freq min..max | Max bit deviation (sigma) | Avalanche 1k mean / std | Avalanche 16k mean / std | Per-output-bit flip prob (16k) | Worst chi2 z of 4 windows | Dups | +|---|---|---|---|---|---|---|---|---| +| igneum-genesis | 104 | 0.4990..0.5009 | 2.07 (bit 38) | 32.07 / 4.09 | 32.035 / 3.99 | 0.492..0.507 | 0.42 | 0 | +| igneum-genesis/stats1 | 128 | 0.4986..0.5011 | 2.90 (bit 3) | 32.14 / 3.94 | 31.987 / 3.98 | 0.490..0.508 | 2.26 | 0 | +| igneum-genesis/stats2 | 176 | 0.4989..0.5010 | 2.29 (bit 9) | 31.90 / 3.97 | 32.035 / 4.01 | 0.492..0.508 | 1.65 | 0 | + +Chi-square detail (df 65535, expected 65535, sigma 362): + +| Seed | bits 0..15 | bits 16..31 | bits 32..47 | bits 48..63 | +|---|---|---|---|---| +| igneum-genesis | 65681 (z 0.40) | 65686 (z 0.42) | 65513 (z -0.06) | 65478 (z -0.16) | +| igneum-genesis/stats1 | 65163 (z -1.03) | 65768 (z 0.64) | 65820 (z 0.79) | 66353 (z 2.26) | +| igneum-genesis/stats2 | 66132 (z 1.65) | 65447 (z -0.24) | 65942 (z 1.12) | 65960 (z 1.17) | + +Reading the numbers. (a) With 2^20 samples the sigma on a bit count is 512; the largest deviation over 192 +bit positions (3 seeds x 64) was 2.90 sigma, which is what 192 draws from a fair coin produce. (b) An ideal +function changes 32 of 64 output bits on average with standard deviation 4. The 1,000-trial means are within +0.14 of 32 (standard error 0.13); the 16,000-trial means are within 0.035 of 32 (standard error 0.03), the +standard deviations are 3.98 to 4.01, and every one of the 64 output bits flipped with probability 0.490 to +0.508 (sigma 0.004). The per-input-bit means over 500 flips each ranged 31.5 to 32.5 (sigma 0.18). An +earlier 1,000-trial sample for igneum-genesis, drawn with a different RNG salt before the 16,000-trial pass +was added, gave mean 31.72; the 16,000-trial figure of 32.035 shows that was sampling noise. (c) All 12 +chi-square values are within 2.3 sigma of their expectation. (d) No duplicate among 2^20 outputs; the birthday +expectation at 64 bits is 3e-8. + +Verdict: the output looks uniform on every measure tried, for all three seeds. STATS: PASS. + +This is a sanity check for obvious structural bias. It is not a proof of cryptographic strength, it says +nothing about adversarially chosen inputs, and 3 seeds is not a statement about the population of programs. + +## 4. Determinism + +Command: + +``` +./igneum-bench --determinism +``` + +Seed igneum-genesis, 2^20 nonces from 0, 1 GiB dataset. The output buffer is filled with a sentinel before +every run so an unwritten lane would show. + +| Check | Result | +|---|---| +| Two independent generations of the program give identical MSL text (4,770 bytes) | yes | +| 5 GPU runs of the same pipeline, fingerprint FNV-1a over 8 MiB | 933787e8cfefccb7 all 5 runs, 0 outputs differ, 0 unwritten lanes | +| Second compile of identical source (0.1 ms, served by the system shader cache) | fingerprint identical | +| Third compile with a comment tag appended so the cache misses (30.0 ms, a real recompile) | fingerprint identical | +| CPU interpreter on 8 warps of the reference run (first, last, 6 random) | all match | +| Dataset filled twice, both blitted to shared memory and fingerprinted | 0b1a77899ee60493 both times | +| 4,096 sampled dataset words incl. indices 0, 1, MASK-1, MASK vs CPU `datasetElem` | all match | + +GPU time per 2^20 batch: 33.0 ms first, 23.2 to 23.8 ms after. DETERMINISM: PASS. + +## 5. Memory safety of dataset indexing + +Command: + +``` +./igneum-bench --memcheck +``` + +Static: the generated MSL for igneum-genesis at 2^20, 2^24 and 2^28 words has 13 `dataset[` accesses, all 13 +of the exact form `dataset[rN & MASK]`, and the identifier `dataset` appears 14 times (13 accesses plus the +kernel parameter). The CUDA twin from `--export-pack` has 14 `ds[` accesses: 13 of the form `ds[rN & mask]` +and the fill kernel's one write, guarded by `if (i < n)`. The same static check ran on all fuzz programs +(section 1) with zero failures. + +Dynamic, 4 MiB dataset (MASK 0x000fffff): four full batches of 2^20 nonces from bases 0xfff00000 (last nonce +0xffffffff), 0xffffffe0 (wraps to 0 inside the batch), 0x80000000 and 0 all completed; 4 verification warps at +bases 0xffffffe0, 0xffffffff, 0x80000000, 0xfff00000 matched the CPU. In those warps 416 of 416 load indices +were above MASK before masking (a random 32-bit value is below 2^20 with probability 2^-12), so the mask was +exercised on every load. + +Metal does not bounds-check device buffers, so "it did not crash" is weak evidence on its own. The static check +is the guarantee: the emitter has exactly one load template and it masks. MEMCHECK: PASS. + +## 6. The bench still works + +Command, after all the changes above: + +``` +./igneum-bench +``` + +| Seed | Compile ms | Mhash/s (wall) | Mhash/s (GPU time) | GB/s useful | Loads/hash | CPU verify ms/warp | Verify | +|---|---|---|---|---|---|---|---| +| igneum-genesis | 0.2 (shader cache) | 44.56 | 44.59 | 18.54 | 104 | 0.017 | PASS 3/3 warps | +| igneum-genesis/epoch1 | 0.3 (shader cache) | 47.74 | 47.77 | 19.86 | 104 | 0.017 | PASS 3/3 warps | + +Dataset fill 5.05 ms GPU (first fill of the process). OVERALL: PASS. The rates match the 3 October table in +README.md (45.2 and 48.4) to within 2 percent, so the bench path is unchanged. The compile figures are the +system shader cache answering for source compiled earlier today; they are not new compile measurements. + +Pack export re-run to a scratch directory and compared byte for byte with the pack written before these +changes (`../proto-cuda/packs/igneum-genesis`): + +``` +./igneum-bench --seed igneum-genesis --export-pack /pack-genesis +diff -r /pack-genesis ../proto-cuda/packs/igneum-genesis +``` + +Metal GPU cross-check PASS 3/3 warps, pack written, and `diff -r` reported no differences: all six files +(`kernel.cu`, `program.h`, `vectors.h`, `program.json`, `vectors.json`, `program.metal`) are byte-identical +to the pack exported before these changes. + +## 7. Shortcut measurement: the closed-form dataset is not memory-hard + +Command: + +``` +./igneum-bench --inline-dataset --hours 1 +``` + +The dataset element is `ds_elem(i, d0, d1)`, six integer operations. A miner can replace every +`dataset[a & MASK]` with `ds_elem(a & MASK, d0, d1)` and never read memory. `--inline-dataset` emits exactly +that kernel; the CPU verification still passes because the function is unchanged. + +| Kernel | Dataset | Mhash/s (wall) | Mhash/s (GPU time) | Verify | +|---|---|---|---|---| +| honest (loads from the 1 GiB buffer), this session | 1 GiB | 44.6 | 44.6 | PASS 3/3 | +| honest, README sweep of 3 October | 4 MiB (cache resident) | 569 | not recorded | PASS | +| inline ds_elem, no memory read | 1 GiB mask | 4,888 | 6,274 | PASS 3/3 | +| inline ds_elem, no memory read | 4 MiB mask | 5,646 | 6,024 | PASS 3/3 | + +The inline kernel runs about 110x faster than the honest kernel by wall clock (140x by GPU timestamps) and +about 9x faster than the honest kernel with a cache-resident dataset. The inline figures are approximate: the +4 timed batches took 2.7 to 3.4 ms in total, so command overhead is visible in the wall figure. The GPU-time +figure is consistent with the ALU peak: roughly 1,500 integer instructions per hash at 6.3 Ghash/s is about +9.4 T instructions/s across 40 cores, in the range of this part's compute throughput. + +Reading: with a closed-form dataset the prototype hash is compute-bound for anyone who skips the buffer, and +the honest kernel's memory traffic is voluntary. Memory-hardness has to come from a dataset element that +costs more to derive than to load. RandomX gets this from a cache of roughly 256 MiB expanded by SuperscalarHash into a +dataset of roughly 2 GiB (approximate, from memory; `vendor/RandomX` is not cloned on this machine, so there +is no file citation yet). The README already listed an expensive dataset element under "What to try next"; this +measurement is why it is not optional. + +## 8. Conclusions + +### Demonstrated on this machine + +1. Bit-exact GPU-versus-CPU agreement of the random-program hash over 10,200 random programs (200 plus + the 10,000 run), 4 random warps each across the full 32-bit nonce range, at 64 MiB, 256 MiB and 1 GiB, plus + the 7 programs in README.md, plus 14 hand-built edge cases with their operand values proven. Zero mismatches, + zero compile failures. +2. Deterministic GPU execution: 5 runs and 3 compiles (one forced cold) of the same program give the same + 8 MiB of output; the dataset fill is deterministic and matches the CPU closed form at sampled indices + including 0 and MASK. +3. Every dataset index in the emitted MSL and CUDA is masked, by static check on every program tested, and + the mask is exercised by essentially every load. +4. No obvious structural bias in the 64-bit output for 3 seeds at 2^20 nonces: bit frequencies, avalanche + (mean 32.0, std 4.0, every output bit flips with probability 0.49 to 0.51), chi-square on four 16-bit + windows, zero duplicates. +5. The emitter's contract on the generator (rotl 1..31, masks in {1,2,4,8,16}, src != dst) held on every + instruction of every fuzzed program. + +### Not demonstrated + +1. Cryptographic security of the construction. Nothing here speaks to preimage, second-preimage or collision + resistance, or to an adversary who chooses nonces or influences the epoch seed. The per-register + initialisation is `splitmix32(nonce ^ seed) ^ seed`, which is a bijection of the nonce per register, and + the body is add-rotate-xor-multiply with `or` (which destroys information) and no non-linear table. None + of that has been analysed. +2. Resistance to algebraic or structural attacks on weak programs. 3 seeds were measured for bias; the + population of programs was not. A program whose `or` chain saturates a register, or whose load addresses + collapse to few values, would be both biased and shortcut-able, and nothing yet rejects such programs or + bounds how often they occur. +3. Memory-hardness. Section 7 measures the shortcut directly: with a closed-form dataset the honest kernel's + memory traffic is optional. The README's "memory bound at 1 GiB" describes the honest kernel, not the + fastest kernel. This is a design gap of the prototype, not a bug in the tests. +4. Vendor independence. Only Apple Metal was run. The CUDA twin exists as text and passed a CPU emulation; + it has not run on NVIDIA hardware, where `__shfl_xor_sync`, `__umulhi` and shift semantics must be shown + to agree bit for bit with these vectors. +5. The 10 ms CPU verification gate against a dataset that is expensive to derive. The current 0.02 ms per warp + is with a six-operation element. + +### The next three tests the real cryptographer must do before the spec is final + +1. Weak-program census and rejection filter. Generate at least 10^5 programs. For each, measure on the CPU + interpreter (no GPU needed): output bit bias and avalanche on 2^12 nonces, distinct load addresses per + hash, fraction of registers saturated by `or`, and whether any register's final value is independent of + the nonce. Report the distribution, define the rejection rule, and bound the advantage of a miner who can + grind the epoch seed over candidate blocks. +2. Replace the closed-form dataset with an expensive derivation (RandomX style: a cache of hundreds of MiB + expanded by a slow hash, or Argon2-based) and re-run three measurements: the `--inline-dataset` shortcut + ratio (must fall to about 1), the GPU rate at 1 GiB, and the CPU verify time against the 10 ms gate with + on-demand element derivation. Also measure distinct cache lines touched per hash so a cache-resident + shortcut is ruled out, not assumed. +3. Cross-vendor bit-exactness and a written spec. Run the exported pack and the fuzz vectors on NVIDIA + (RTX 5090, CUDA) and on at least one AMD GPU, with the same 4-warp-per-program comparison, including the + edge-case programs. Then fix the hash as a specification with test vectors and replace the ad-hoc seed + derivation (FNV-1a plus SplitMix) with a standard hash so the seed-to-program mapping is auditable. + +## Files + +- `main.swift`: tests under `// MARK: - Hardening tests`, flags `--fuzz`, `--fuzz-seed`, `--edge`, `--stats`, + `--determinism`, `--memcheck`, `--inline-dataset`. +- `TESTS.md`: this file. +- Raw logs of the runs quoted above were kept in the session scratchpad and are not checked in; every table + is reproducible with the command above it. diff --git a/proto-metal/main.swift b/proto-metal/main.swift index a9faca9b..7fb1958f 100644 --- a/proto-metal/main.swift +++ b/proto-metal/main.swift @@ -24,6 +24,8 @@ struct Options { var stats = false // --stats: output distribution sanity checks on 2^20 nonces var determinism = false // --determinism: 5 identical GPU runs + double compile var memcheck = false // --memcheck: static mask check + 4 MiB run with wrapping nonces + // Shortcut measurement (bench variant, not a test): compute dataset elements inline instead of loading them. + var inlineDataset = false var anyTest: Bool { fuzz != nil || edge || stats || determinism || memcheck } } @@ -49,6 +51,7 @@ func parseArgs() -> Options { case "--stats": o.stats = true case "--determinism": o.determinism = true case "--memcheck": o.memcheck = true + case "--inline-dataset": o.inlineDataset = true case "-h", "--help": print(""" igneum-bench [--seed ] [--hours N] [--batch-log2 22] [--batches 4] @@ -61,6 +64,8 @@ func parseArgs() -> Options { [--stats] output distribution sanity checks on 2^20 nonces, 3 seeds [--determinism] 5 identical GPU runs of 2^20 nonces, double compile, dataset fill check [--memcheck] static dataset-index mask check, 4 MiB run with wrapping nonces + shortcut measurement: + [--inline-dataset] bench variant: every load computes ds_elem(index) inline, no memory read """) exit(0) default: @@ -184,7 +189,9 @@ func generateProgram(seedString: String) -> Program { func hex(_ v: UInt32) -> String { String(format: "0x%08xu", v) } -func generateMSL(_ p: Program, datasetLog2: Int) -> String { +// inlineDay: when set, every load computes ds_elem(index, day) in registers instead of reading the buffer. +// Same function, no memory traffic. Used only by --inline-dataset to measure the closed-form shortcut. +func generateMSL(_ p: Program, datasetLog2: Int, inlineDay: (UInt32, UInt32)? = nil) -> String { let mask = UInt32((1 << datasetLog2) - 1) var s = """ #include @@ -236,7 +243,9 @@ func generateMSL(_ p: Program, datasetLog2: Int) -> String { case .rotr: line = "\(d) = rotr_var(\(d), \(a));" case .mad: line = "\(d) = \(a) * \(b) + \(d);" case .shfl: line = "\(d) = \(d) ^ simd_shuffle_xor(\(a), (ushort)\(ins.mask));" - case .load: line = "\(d) = \(d) ^ dataset[\(a) & MASK];" + case .load: + if let dd = inlineDay { line = "\(d) = \(d) ^ ds_elem(\(a) & MASK, \(hex(dd.0)), \(hex(dd.1)));" } + else { line = "\(d) = \(d) ^ dataset[\(a) & MASK];" } } s += " \(line) // \(k)\n" } @@ -728,7 +737,8 @@ struct EpochResult { func runEpoch(gpu: GPU, opts: Options, seedString: String, dataset: MTLBuffer, day: (UInt32, UInt32)) -> EpochResult { let program = generateProgram(seedString: seedString) - let msl = generateMSL(program, datasetLog2: opts.datasetLog2) + let msl = generateMSL(program, datasetLog2: opts.datasetLog2, inlineDay: opts.inlineDataset ? day : nil) + if opts.inlineDataset { print("\nNOTE: --inline-dataset: loads compute ds_elem inline, the dataset buffer is never read") } if let dir = opts.dumpDir { try? FileManager.default.createDirectory(atPath: dir, withIntermediateDirectories: true) let safe = seedString.replacingOccurrences(of: "/", with: "_") @@ -1308,35 +1318,44 @@ func runStats(_ opts: Options, gpu: GPU, day: (UInt32, UInt32)) -> Bool { var dups = 0 for i in 1.. (mean: Double, std: Double, minD: Int, maxD: Int, perBitMin: Double, perBitMax: Double, perOutMin: Double, perOutMax: Double)? { + let sw = seedWords("avalanche/\(salt)/" + seedString) + var rng = SplitMix64(s: UInt64(sw[0]) | (UInt64(sw[1]) << 32)) + var bases = [UInt32]() + var flipped = [Int]() + for t in 0..> UInt64(b)) & 1 == 1 { perOut[b] += 1 } + minDiff = min(minDiff, d); maxDiff = max(maxDiff, d) + } + let mean = Double(diffs.reduce(0, +)) / Double(trials) + let variance = diffs.reduce(0.0) { $0 + (Double($1) - mean) * (Double($1) - mean) } / Double(trials - 1) + var pbMin = 64.0, pbMax = 0.0 + for b in 0..<32 where perBitN[b] > 0 { let m = Double(perBitSum[b]) / Double(perBitN[b]); pbMin = min(pbMin, m); pbMax = max(pbMax, m) } + let poMin = Double(perOut.min()!) / Double(trials), poMax = Double(perOut.max()!) / Double(trials) + return (mean, variance.squareRoot(), minDiff, maxDiff, pbMin, pbMax, poMin, poMax) } - guard let av = gpuWarps(gpu, k, dataset: dataset, bases: bases) else { print("FAIL: avalanche GPU run"); return false } - var diffs = [Int]() - var perBitSum = [Int](repeating: 0, count: 32), perBitN = [Int](repeating: 0, count: 32) - var minDiff = 64, maxDiff = 0 - for t in 0.. 0 { let m = Double(perBitSum[b]) / Double(perBitN[b]); perBitMin = min(perBitMin, m); perBitMax = max(perBitMax, m) } - // Expected for an ideal function: mean 32, std 4 (binomial 64 x 0.5). Standard error of the mean over 1000 trials is 0.13. - let avOk = abs(mean - 32) < 0.6 && std > 3.3 && std < 4.7 + guard let a1 = avalanche(1000, salt: "small"), let a2 = avalanche(16000, salt: "large") else { print("FAIL: avalanche GPU run"); return false } + // Expected for an ideal function: mean 32, std 4 (binomial 64 x 0.5). Standard error of the mean: + // 0.13 bits at 1,000 trials, 0.03 bits at 16,000. Thresholds are about 4.5 standard errors. + let avOk = abs(a1.mean - 32) < 0.6 && a1.std > 3.3 && a1.std < 4.7 && abs(a2.mean - 32) < 0.15 && a2.std > 3.6 && a2.std < 4.4 let freqOk = maxZ < 4.5 let chiOk = chiWorstZ < 4.5 let dupOk = dups == 0 @@ -1345,14 +1364,15 @@ func runStats(_ opts: Options, gpu: GPU, day: (UInt32, UInt32)) -> Bool { print("\nseed \"\(seedString)\": loads/hash \(program.loadsPerHash), GPU \(fmt(gms, 1)) ms for 2^20 hashes, CPU spot check 2 warps \(spot ? "PASS" : "FAIL")") print(" (a) bit frequency: min \(fmt(minFreq, 4)) max \(fmt(maxFreq, 4)); largest deviation \(fmt(maxDev, 0)) counts at bit \(maxBit) = \(fmt(maxZ, 2)) sigma (sigma \(fmt(sigma, 0)), 64 bits, expect max under about 3.5)") - print(" (b) avalanche over \(trials) single-bit nonce flips: mean \(fmt(mean, 2)) std \(fmt(std, 2)) min \(minDiff) max \(maxDiff) of 64 bits (expect mean 32, std 4); per-input-bit mean range \(fmt(perBitMin, 1))..\(fmt(perBitMax, 1))") + print(" (b) avalanche, 1000 single-bit nonce flips: mean \(fmt(a1.mean, 2)) std \(fmt(a1.std, 2)) min \(a1.minD) max \(a1.maxD) of 64 bits (expect mean 32, std 4); per-input-bit mean range \(fmt(a1.perBitMin, 1))..\(fmt(a1.perBitMax, 1))") + print(" (b) avalanche, 16000 flips (500 per input bit): mean \(fmt(a2.mean, 3)) std \(fmt(a2.std, 2)) min \(a2.minD) max \(a2.maxD); per-input-bit mean range \(fmt(a2.perBitMin, 2))..\(fmt(a2.perBitMax, 2)); per-output-bit flip probability range \(fmt(a2.perOutMin, 3))..\(fmt(a2.perOutMax, 3)) (expect 0.5, sigma 0.004)") for r in chiRows { print(" (c) \(r)") } print(" (d) duplicate 64-bit outputs among 2^20: \(dups) (expected about 3e-8)") print(" verdict: \(ok ? "no obvious bias" : "SUSPECT")") - rows.append("| \(seedString) | \(program.loadsPerHash) | \(fmt(minFreq, 4))..\(fmt(maxFreq, 4)) | \(fmt(maxZ, 2)) | \(fmt(mean, 2)) | \(fmt(std, 2)) | \(fmt(chiWorstZ, 2)) | \(dups) | \(ok ? "uniform-looking" : "SUSPECT") |") + rows.append("| \(seedString) | \(program.loadsPerHash) | \(fmt(minFreq, 4))..\(fmt(maxFreq, 4)) | \(fmt(maxZ, 2)) | \(fmt(a1.mean, 2)) / \(fmt(a1.std, 2)) | \(fmt(a2.mean, 3)) / \(fmt(a2.std, 2)) | \(fmt(a2.perOutMin, 3))..\(fmt(a2.perOutMax, 3)) | \(fmt(chiWorstZ, 2)) | \(dups) | \(ok ? "uniform-looking" : "SUSPECT") |") } - print("\n| Seed | Loads/hash | Bit freq min..max | Max bit z | Avalanche mean | Avalanche std | Worst chi2 z (4 windows) | Dups | Verdict |") - print("|---|---|---|---|---|---|---|---|---|") + print("\n| Seed | Loads/hash | Bit freq min..max | Max bit z | Avalanche 1k mean / std | Avalanche 16k mean / std | Per-output-bit flip prob | Worst chi2 z (4 windows) | Dups | Verdict |") + print("|---|---|---|---|---|---|---|---|---|---|") for r in rows { print(r) } print("STATS: \(allOk ? "PASS (no obvious structural bias; not a security proof)" : "FAIL (something looks biased)")") return allOk @@ -1376,10 +1396,14 @@ func runDeterminism(_ opts: Options, gpu: GPU, day: (UInt32, UInt32)) -> Bool { print("generator: two generations of the program give identical MSL source: \(sameSource ? "yes" : "NO") (\(msl1.utf8.count) bytes)") if !sameSource { ok = false } - // Two separate compiles of the same source. - let k1: CompiledHash, k2: CompiledHash - do { k1 = try compileHash(gpu, msl: msl1); k2 = try compileHash(gpu, msl: msl2) } catch { print("FAIL: compile \(error)"); return false } - print("compiled twice: \(fmt(k1.totalMs, 1)) ms and \(fmt(k2.totalMs, 1)) ms") + // Three compiles: the same source twice (Metal's shader cache may serve the second), and once more with a + // comment tag appended so the cache misses and the compiler really runs again on an identical kernel. + let tag = "\n// recompile tag \(h64(nowNs()))\n" + let k1: CompiledHash, k2: CompiledHash, k3: CompiledHash + do { + k1 = try compileHash(gpu, msl: msl1); k2 = try compileHash(gpu, msl: msl2); k3 = try compileHash(gpu, msl: msl2 + tag) + } catch { print("FAIL: compile \(error)"); return false } + print("compiled three times: identical source \(fmt(k1.totalMs, 1)) ms and \(fmt(k2.totalMs, 1)) ms (a sub-millisecond second figure means the system shader cache answered), tagged source \(fmt(k3.totalMs, 1)) ms (forced recompile)") func runOnce(_ k: CompiledHash) -> (fp: UInt64, sentinels: Int, ms: Double)? { memset(outBuf.contents(), 0xAA, n * 8) @@ -1407,7 +1431,12 @@ func runDeterminism(_ opts: Options, gpu: GPU, day: (UInt32, UInt32)) -> Bool { var differ2 = 0 for i in 0..