Prototype: fuzz, edge, determinism, memcheck and statistics tests; inline-dataset shortcut measured
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
ff263245ac
commit
8f179aaaaf
4 changed files with 408 additions and 40 deletions
|
|
@ -58,3 +58,15 @@ D: honest doubling is under-weighted 28 to 31 days; old miners lock alone for 20
|
|||
E: 30% churn leaves 71% live, lock never lost. Threshold is 1/3: 35% stalls 1 day, 50% stalls 10 to 11 days. Active-24h total removes every stall.
|
||||
F: 51% patient owner holds 51.0% weight from day 45, vetoes from day 7 to 20, never locks alone. 67% owner locks alone from day 34 to 44.
|
||||
Recommend: f=1, cap 2x, window 30 (the window is the defence, the cap is worth 1 to 8 days), total = all keys with weight in the window. Details in sim/results.md.
|
||||
|
||||
## 2026-10-03 proto-metal hardening tests (correctness and soundness of the lottery hash, Metal only)
|
||||
|
||||
Machine: the same Apple M5 Max. Added `--fuzz`, `--edge`, `--stats`, `--determinism`, `--memcheck`, `--inline-dataset` to `proto-metal/main.swift`. Full tables and commands in `proto-metal/TESTS.md`.
|
||||
Fuzz: 10,200 random programs (200 + 10,000, cold compiles 13.6 to 48.5 ms), 4 random full-range warps each, dataset drawn from 64 MiB / 256 MiB / 1 GiB: 40,800 warps, 1,305,600 hashes, 0 mismatches, 0 compile failures, 0 static mask failures, generator contract (rotl 1..31, mask in {1,2,4,8,16}, src != dst) held on every instruction. 196 s for the 10,000 run.
|
||||
Edge: 14 hand-built cases (rotr by register 0 / 32 / -32 / 31 / 63, rotl 1 and 31, mulhi max operands, shfl masks 1..16, loads at index 0 and MASK via in-range and out-of-range registers, add/sub/mul/mad wraparound, zero loads, 64 loads) with operand values proven by a traced interpreter: 14/14 PASS, 128/128 lanes each. rotl by 0 (never generated) agreed too, recorded as informational only.
|
||||
Stats (3 seeds, 2^20 nonces each): bit frequency max deviation 2.90 sigma over 192 bit positions; avalanche 16,000 flips mean 31.99 to 32.04 (expect 32), std 3.98 to 4.01 (expect 4), every output bit flips with probability 0.490 to 0.508; chi-square on four 16-bit windows all within 2.3 sigma; 0 duplicates. Looks uniform. Not a security proof.
|
||||
Determinism: 5 runs and 3 compiles (one forced cold, 30 ms) of 2^20 hashes gave fingerprint 933787e8cfefccb7 every time; dataset fill deterministic (0b1a77899ee60493 twice) and 4,096 sampled words incl. 0 and MASK match the CPU closed form.
|
||||
Memcheck: every `dataset[` in the MSL is `dataset[rN & MASK]` (13/13 at 3 sizes), CUDA twin 13/13 plus one guarded fill write; 4 MiB run with nonces up to 0xffffffff completed and 4 wrapping warps matched the CPU; 416/416 load indices exceeded MASK before masking.
|
||||
Bench re-run after the changes: igneum-genesis 44.56 Mhash/s, epoch1 47.74 Mhash/s at 1 GiB, PASS 3/3 warps each (within 2 percent of the first-run table). `--export-pack igneum-genesis` re-run is byte-identical to the existing pack.
|
||||
SHORTCUT MEASURED: `--inline-dataset` replaces every load with the six-op closed form ds_elem and never reads memory: 4,888 Mhash/s wall (6,274 GPU time) vs 44.6 honest at 1 GiB, about 110x, and about 9x the cache-resident honest rate. With a closed-form dataset the hash is not memory-hard; an expensive dataset derivation is required, not optional.
|
||||
Not demonstrated: cryptographic strength, weak-program frequency and rejection, NVIDIA/AMD bit-exactness (CUDA run still pending), CPU verify gate with an expensive dataset element. Next three tests for the cryptographer are listed in TESTS.md section 8.
|
||||
|
|
|
|||
|
|
@ -42,6 +42,11 @@ Flags: `--seed`, `--day`, `--hours`, `--batch-log2` (default 22), `--batches` (d
|
|||
`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump <dir>`, `--export-pack <dir>`.
|
||||
Exit code 0 means every verified warp matched.
|
||||
|
||||
Hardening tests (added 3 October 2026, results and commands in `TESTS.md`): `--fuzz N [--fuzz-seed <string>]`,
|
||||
`--edge`, `--stats`, `--determinism`, `--memcheck`. Any of these runs instead of the bench; several may be combined
|
||||
in one invocation; exit code 0 only if every selected test passed. `--inline-dataset` is a bench variant that
|
||||
computes every dataset element inline instead of loading it (the shortcut measurement in `TESTS.md` section 7).
|
||||
|
||||
`--export-pack <dir>` (added 3 October 2026) does not run the bench. It generates the program for `--seed`,
|
||||
emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (base nonces 0, 4096, 1000000)
|
||||
with the CPU interpreter, cross-checks them on the Metal GPU, and writes `kernel.cu`, `program.h`, `vectors.h`,
|
||||
|
|
@ -104,6 +109,10 @@ like the system shader cache.
|
|||
being measured. Batches of 2^22 nonces take 86 to 118 ms each.
|
||||
8. Only Apple silicon was measured. Nothing here says anything about NVIDIA or AMD, where SIMD width, cache
|
||||
line size and memory latency differ.
|
||||
9. Measured later on 3 October 2026 (`TESTS.md` section 7): because the dataset element is a six-operation
|
||||
closed form, a kernel that computes it inline instead of loading it runs at about 4,900 Mhash/s against
|
||||
44.6 for the honest kernel at 1 GiB, roughly 110x. Observation 1 describes the honest kernel only. The
|
||||
prototype is not memory-hard until the dataset element costs more to derive than to load.
|
||||
|
||||
## What to try next
|
||||
|
||||
|
|
|
|||
318
proto-metal/TESTS.md
Normal file
318
proto-metal/TESTS.md
Normal file
|
|
@ -0,0 +1,318 @@
|
|||
# proto-metal hardening tests
|
||||
|
||||
Date: 3 October 2026. Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, macOS Darwin 25.6.0,
|
||||
Swift 5.8.1 from Command Line Tools, no Xcode, Metal shaders compiled at runtime.
|
||||
Build: `swiftc -O -o igneum-bench main.swift -framework Metal` (4 s).
|
||||
|
||||
Every number below was produced on this machine on this date by the commands shown. The tests live in
|
||||
`main.swift` next to the bench and share its generator, MSL emitter, dataset fill and CPU interpreter
|
||||
without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench
|
||||
does not use; `cpuWarp` calls it with `nil`.
|
||||
|
||||
What the GPU-vs-CPU comparison means: the GPU runs the generated Metal kernel, the CPU runs a plain Swift
|
||||
interpreter over the same instruction list, and every 64-bit output is compared bit for bit. The interpreter
|
||||
never reads the dataset buffer; it recomputes each element from the closed form. So agreement also covers the
|
||||
dataset fill kernel and the address masking, not only the arithmetic.
|
||||
|
||||
## 1. Fuzz: random programs, random warps, three dataset sizes
|
||||
|
||||
Command (required run):
|
||||
|
||||
```
|
||||
./igneum-bench --fuzz 200
|
||||
```
|
||||
|
||||
For each of N seeds: generate the program, emit MSL, static mask check, compile on the GPU, draw a dataset size
|
||||
from {64 MiB, 256 MiB, 1 GiB}, draw 4 base nonces uniformly from the full 32-bit range (so warps straddle
|
||||
2^31 and wrap past 2^32), run the GPU, run the CPU interpreter, compare all 128 outputs. The master seed fixes
|
||||
the whole sequence, so any failure is reproducible with `--fuzz-seed`. The run also asserts the generator's
|
||||
contract that the emitter relies on: rotl immediate in 1..31, shuffle mask in {1,2,4,8,16}, source register
|
||||
never the destination.
|
||||
|
||||
| Dataset | Programs | Pass | Fail |
|
||||
|---|---|---|---|
|
||||
| 2^24 words (64 MiB) | 63 | 63 | 0 |
|
||||
| 2^26 words (256 MiB) | 64 | 64 | 0 |
|
||||
| 2^28 words (1 GiB) | 73 | 73 | 0 |
|
||||
| all | 200 | 200 | 0 |
|
||||
|
||||
200 programs: pass 200, mismatch 0, compile failures 0, static mask failures 0, contract failures 0.
|
||||
800 warps compared (25,600 hashes). Loads per hash ranged 64 to 208. Op totals across the 200 programs:
|
||||
load 3204, add 1472, xor 1276, mul 1087, mad 1062, shfl 1038, rotl 892, mulhi 762, sub 742, rotr 739, or 526.
|
||||
Compile (library + pipeline) min 0.0 ms, avg 18.3 ms, max 46.4 ms; the zero is the system shader cache
|
||||
answering for the first 20 seeds, which an earlier smoke run with the same master seed had already compiled.
|
||||
GPU dispatch 156 ms total, CPU interpreter 16 ms total, wall 3.9 s.
|
||||
|
||||
Large run, separate master seed so every compile is cold:
|
||||
|
||||
```
|
||||
./igneum-bench --fuzz 10000 --fuzz-seed igneum-fuzz-large-2026-10-03
|
||||
```
|
||||
|
||||
| Dataset | Programs | Pass | Fail |
|
||||
|---|---|---|---|
|
||||
| 2^24 words (64 MiB) | 3346 | 3346 | 0 |
|
||||
| 2^26 words (256 MiB) | 3334 | 3334 | 0 |
|
||||
| 2^28 words (1 GiB) | 3320 | 3320 | 0 |
|
||||
| all | 10000 | 10000 | 0 |
|
||||
|
||||
10,000 programs: pass 10,000, mismatch 0, compile failures 0, static mask failures 0, contract failures 0.
|
||||
40,000 warps compared (1,280,000 hashes). Loads per hash ranged 40 to 232. Op totals: load 160,123,
|
||||
add 76,698, xor 64,138, mad 51,577, shfl 51,298, mul 50,919, rotl 44,945, rotr 38,333, mulhi 38,284,
|
||||
sub 38,234, or 25,451. Compile min 13.6 ms, avg 18.5 ms, max 48.5 ms (all cold). GPU dispatch 8.2 s total,
|
||||
CPU interpreter 0.8 s total, wall 196 s. FUZZ: PASS.
|
||||
|
||||
Combined with the 200-program run: 10,200 programs, 40,800 warps, 1,305,600 hashes, zero mismatches.
|
||||
The op totals show every one of the 11 instruction families was exercised tens of thousands of times.
|
||||
|
||||
## 2. Edge cases
|
||||
|
||||
Command:
|
||||
|
||||
```
|
||||
./igneum-bench --edge
|
||||
```
|
||||
|
||||
Hand-built programs, dataset 1 GiB (MASK 0x0fffffff), 4 warps at base nonces 0x00000000, 0x00100000,
|
||||
0x7ffffff0 and 0xffffffe0, 128 lanes per case. Operand values are forced with `r = r - r` (zero) followed by a
|
||||
constant add, then proven: the traced CPU interpreter checks the named register value just before the
|
||||
instruction in question, on lane 0 of all 4 warps in all 8 iterations. "Held" means 32 of 32 checks passed.
|
||||
|
||||
| Case | Instrs | Loads/hash | Preconditions | GPU vs CPU | Result |
|
||||
|---|---|---|---|---|---|
|
||||
| rotl immediate by 1 and by 31 | 6 | 0 | none needed | 128/128 lanes | PASS |
|
||||
| rotr by register == 0 | 3 | 0 | held | 128/128 lanes | PASS |
|
||||
| rotr by register == 32 (32 mod 32 = 0) | 5 | 0 | held | 128/128 lanes | PASS |
|
||||
| rotr by register == 0xFFFFFFE0 (-32, 0 mod 32) | 5 | 0 | held | 128/128 lanes | PASS |
|
||||
| rotr by register == 31, 63 and 1 | 10 | 0 | held | 128/128 lanes | PASS |
|
||||
| mulhi 0xFFFFFFFF x 0xFFFFFFFF (result 0xFFFFFFFE), 0x80000000 x 2 (result 1), x 0 (result 0) | 15 | 0 | held, results checked | 128/128 lanes | PASS |
|
||||
| shfl_xor every mask 1..16 in sequence | 16 | 0 | none needed | 128/128 lanes | PASS |
|
||||
| load at index 0 via register 0, and via register MASK+1 | 6 | 16 | held | 128/128 lanes | PASS |
|
||||
| load at index MASK via register MASK, and via register 0xFFFFFFFF | 8 | 16 | held | 128/128 lanes | PASS |
|
||||
| add wraparound 0xFFFFFFFF + 1 (result 0) | 5 | 0 | held, result checked | 128/128 lanes | PASS |
|
||||
| sub wraparound 0 - 1 (result 0xFFFFFFFF) | 6 | 0 | held, result checked | 128/128 lanes | PASS |
|
||||
| mul 0xFFFFFFFF x 0xFFFFFFFF (low result 1), mad same + 5 (result 6) | 15 | 0 | held, results checked | 128/128 lanes | PASS |
|
||||
| generated program with every load replaced by xor (zero loads) | 64 | 0 | none needed | 128/128 lanes | PASS |
|
||||
| 64 loads and nothing else | 64 | 512 | none needed | 128/128 lanes | PASS |
|
||||
| rotl immediate by 0 (outside the generator's 1..31 contract) | 2 | 0 | none needed | 128/128 lanes | info only: agrees |
|
||||
|
||||
EDGE: PASS, 14 of 14 counted cases. The last row is informational: the generator never emits a rotate by 0
|
||||
(the fuzz run asserts this on every instruction), and the MSL `rotl_imm` would shift by 32 for it. On this GPU
|
||||
and compiler the result happened to equal the CPU's. Nothing may rely on that; the contract stays 1..31.
|
||||
The shuffle case covers masks 3, 5, 6, 7, 9 to 15 that the generator does not emit; `simd_shuffle_xor` and the
|
||||
interpreter's `lane ^ mask` agreed for all 16.
|
||||
|
||||
## 3. Output statistics
|
||||
|
||||
Command:
|
||||
|
||||
```
|
||||
./igneum-bench --stats
|
||||
```
|
||||
|
||||
Three seeds: `igneum-genesis`, `igneum-genesis/stats1`, `igneum-genesis/stats2`. For each, 2^20 consecutive
|
||||
nonces from 0 hashed on the GPU at 1 GiB (warps 0 and 32767 spot-checked against the CPU, both matched), then
|
||||
on the CPU side: (a) ones count per output bit, (b) avalanche from single-bit nonce flips, run on the GPU as
|
||||
pairs of one-warp dispatches, (c) chi-square over 65,536 buckets for each 16-bit window of the output,
|
||||
(d) duplicate count after sorting.
|
||||
|
||||
| Seed | Loads/hash | Bit freq min..max | Max bit deviation (sigma) | Avalanche 1k mean / std | Avalanche 16k mean / std | Per-output-bit flip prob (16k) | Worst chi2 z of 4 windows | Dups |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| igneum-genesis | 104 | 0.4990..0.5009 | 2.07 (bit 38) | 32.07 / 4.09 | 32.035 / 3.99 | 0.492..0.507 | 0.42 | 0 |
|
||||
| igneum-genesis/stats1 | 128 | 0.4986..0.5011 | 2.90 (bit 3) | 32.14 / 3.94 | 31.987 / 3.98 | 0.490..0.508 | 2.26 | 0 |
|
||||
| igneum-genesis/stats2 | 176 | 0.4989..0.5010 | 2.29 (bit 9) | 31.90 / 3.97 | 32.035 / 4.01 | 0.492..0.508 | 1.65 | 0 |
|
||||
|
||||
Chi-square detail (df 65535, expected 65535, sigma 362):
|
||||
|
||||
| Seed | bits 0..15 | bits 16..31 | bits 32..47 | bits 48..63 |
|
||||
|---|---|---|---|---|
|
||||
| igneum-genesis | 65681 (z 0.40) | 65686 (z 0.42) | 65513 (z -0.06) | 65478 (z -0.16) |
|
||||
| igneum-genesis/stats1 | 65163 (z -1.03) | 65768 (z 0.64) | 65820 (z 0.79) | 66353 (z 2.26) |
|
||||
| igneum-genesis/stats2 | 66132 (z 1.65) | 65447 (z -0.24) | 65942 (z 1.12) | 65960 (z 1.17) |
|
||||
|
||||
Reading the numbers. (a) With 2^20 samples the sigma on a bit count is 512; the largest deviation over 192
|
||||
bit positions (3 seeds x 64) was 2.90 sigma, which is what 192 draws from a fair coin produce. (b) An ideal
|
||||
function changes 32 of 64 output bits on average with standard deviation 4. The 1,000-trial means are within
|
||||
0.14 of 32 (standard error 0.13); the 16,000-trial means are within 0.035 of 32 (standard error 0.03), the
|
||||
standard deviations are 3.98 to 4.01, and every one of the 64 output bits flipped with probability 0.490 to
|
||||
0.508 (sigma 0.004). The per-input-bit means over 500 flips each ranged 31.5 to 32.5 (sigma 0.18). An
|
||||
earlier 1,000-trial sample for igneum-genesis, drawn with a different RNG salt before the 16,000-trial pass
|
||||
was added, gave mean 31.72; the 16,000-trial figure of 32.035 shows that was sampling noise. (c) All 12
|
||||
chi-square values are within 2.3 sigma of their expectation. (d) No duplicate among 2^20 outputs; the birthday
|
||||
expectation at 64 bits is 3e-8.
|
||||
|
||||
Verdict: the output looks uniform on every measure tried, for all three seeds. STATS: PASS.
|
||||
|
||||
This is a sanity check for obvious structural bias. It is not a proof of cryptographic strength, it says
|
||||
nothing about adversarially chosen inputs, and 3 seeds is not a statement about the population of programs.
|
||||
|
||||
## 4. Determinism
|
||||
|
||||
Command:
|
||||
|
||||
```
|
||||
./igneum-bench --determinism
|
||||
```
|
||||
|
||||
Seed igneum-genesis, 2^20 nonces from 0, 1 GiB dataset. The output buffer is filled with a sentinel before
|
||||
every run so an unwritten lane would show.
|
||||
|
||||
| Check | Result |
|
||||
|---|---|
|
||||
| Two independent generations of the program give identical MSL text (4,770 bytes) | yes |
|
||||
| 5 GPU runs of the same pipeline, fingerprint FNV-1a over 8 MiB | 933787e8cfefccb7 all 5 runs, 0 outputs differ, 0 unwritten lanes |
|
||||
| Second compile of identical source (0.1 ms, served by the system shader cache) | fingerprint identical |
|
||||
| Third compile with a comment tag appended so the cache misses (30.0 ms, a real recompile) | fingerprint identical |
|
||||
| CPU interpreter on 8 warps of the reference run (first, last, 6 random) | all match |
|
||||
| Dataset filled twice, both blitted to shared memory and fingerprinted | 0b1a77899ee60493 both times |
|
||||
| 4,096 sampled dataset words incl. indices 0, 1, MASK-1, MASK vs CPU `datasetElem` | all match |
|
||||
|
||||
GPU time per 2^20 batch: 33.0 ms first, 23.2 to 23.8 ms after. DETERMINISM: PASS.
|
||||
|
||||
## 5. Memory safety of dataset indexing
|
||||
|
||||
Command:
|
||||
|
||||
```
|
||||
./igneum-bench --memcheck
|
||||
```
|
||||
|
||||
Static: the generated MSL for igneum-genesis at 2^20, 2^24 and 2^28 words has 13 `dataset[` accesses, all 13
|
||||
of the exact form `dataset[rN & MASK]`, and the identifier `dataset` appears 14 times (13 accesses plus the
|
||||
kernel parameter). The CUDA twin from `--export-pack` has 14 `ds[` accesses: 13 of the form `ds[rN & mask]`
|
||||
and the fill kernel's one write, guarded by `if (i < n)`. The same static check ran on all fuzz programs
|
||||
(section 1) with zero failures.
|
||||
|
||||
Dynamic, 4 MiB dataset (MASK 0x000fffff): four full batches of 2^20 nonces from bases 0xfff00000 (last nonce
|
||||
0xffffffff), 0xffffffe0 (wraps to 0 inside the batch), 0x80000000 and 0 all completed; 4 verification warps at
|
||||
bases 0xffffffe0, 0xffffffff, 0x80000000, 0xfff00000 matched the CPU. In those warps 416 of 416 load indices
|
||||
were above MASK before masking (a random 32-bit value is below 2^20 with probability 2^-12), so the mask was
|
||||
exercised on every load.
|
||||
|
||||
Metal does not bounds-check device buffers, so "it did not crash" is weak evidence on its own. The static check
|
||||
is the guarantee: the emitter has exactly one load template and it masks. MEMCHECK: PASS.
|
||||
|
||||
## 6. The bench still works
|
||||
|
||||
Command, after all the changes above:
|
||||
|
||||
```
|
||||
./igneum-bench
|
||||
```
|
||||
|
||||
| Seed | Compile ms | Mhash/s (wall) | Mhash/s (GPU time) | GB/s useful | Loads/hash | CPU verify ms/warp | Verify |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| igneum-genesis | 0.2 (shader cache) | 44.56 | 44.59 | 18.54 | 104 | 0.017 | PASS 3/3 warps |
|
||||
| igneum-genesis/epoch1 | 0.3 (shader cache) | 47.74 | 47.77 | 19.86 | 104 | 0.017 | PASS 3/3 warps |
|
||||
|
||||
Dataset fill 5.05 ms GPU (first fill of the process). OVERALL: PASS. The rates match the 3 October table in
|
||||
README.md (45.2 and 48.4) to within 2 percent, so the bench path is unchanged. The compile figures are the
|
||||
system shader cache answering for source compiled earlier today; they are not new compile measurements.
|
||||
|
||||
Pack export re-run to a scratch directory and compared byte for byte with the pack written before these
|
||||
changes (`../proto-cuda/packs/igneum-genesis`):
|
||||
|
||||
```
|
||||
./igneum-bench --seed igneum-genesis --export-pack <scratch>/pack-genesis
|
||||
diff -r <scratch>/pack-genesis ../proto-cuda/packs/igneum-genesis
|
||||
```
|
||||
|
||||
Metal GPU cross-check PASS 3/3 warps, pack written, and `diff -r` reported no differences: all six files
|
||||
(`kernel.cu`, `program.h`, `vectors.h`, `program.json`, `vectors.json`, `program.metal`) are byte-identical
|
||||
to the pack exported before these changes.
|
||||
|
||||
## 7. Shortcut measurement: the closed-form dataset is not memory-hard
|
||||
|
||||
Command:
|
||||
|
||||
```
|
||||
./igneum-bench --inline-dataset --hours 1
|
||||
```
|
||||
|
||||
The dataset element is `ds_elem(i, d0, d1)`, six integer operations. A miner can replace every
|
||||
`dataset[a & MASK]` with `ds_elem(a & MASK, d0, d1)` and never read memory. `--inline-dataset` emits exactly
|
||||
that kernel; the CPU verification still passes because the function is unchanged.
|
||||
|
||||
| Kernel | Dataset | Mhash/s (wall) | Mhash/s (GPU time) | Verify |
|
||||
|---|---|---|---|---|
|
||||
| honest (loads from the 1 GiB buffer), this session | 1 GiB | 44.6 | 44.6 | PASS 3/3 |
|
||||
| honest, README sweep of 3 October | 4 MiB (cache resident) | 569 | not recorded | PASS |
|
||||
| inline ds_elem, no memory read | 1 GiB mask | 4,888 | 6,274 | PASS 3/3 |
|
||||
| inline ds_elem, no memory read | 4 MiB mask | 5,646 | 6,024 | PASS 3/3 |
|
||||
|
||||
The inline kernel runs about 110x faster than the honest kernel by wall clock (140x by GPU timestamps) and
|
||||
about 9x faster than the honest kernel with a cache-resident dataset. The inline figures are approximate: the
|
||||
4 timed batches took 2.7 to 3.4 ms in total, so command overhead is visible in the wall figure. The GPU-time
|
||||
figure is consistent with the ALU peak: roughly 1,500 integer instructions per hash at 6.3 Ghash/s is about
|
||||
9.4 T instructions/s across 40 cores, in the range of this part's compute throughput.
|
||||
|
||||
Reading: with a closed-form dataset the prototype hash is compute-bound for anyone who skips the buffer, and
|
||||
the honest kernel's memory traffic is voluntary. Memory-hardness has to come from a dataset element that
|
||||
costs more to derive than to load. RandomX gets this from a cache of roughly 256 MiB expanded by SuperscalarHash into a
|
||||
dataset of roughly 2 GiB (approximate, from memory; `vendor/RandomX` is not cloned on this machine, so there
|
||||
is no file citation yet). The README already listed an expensive dataset element under "What to try next"; this
|
||||
measurement is why it is not optional.
|
||||
|
||||
## 8. Conclusions
|
||||
|
||||
### Demonstrated on this machine
|
||||
|
||||
1. Bit-exact GPU-versus-CPU agreement of the random-program hash over 10,200 random programs (200 plus
|
||||
the 10,000 run), 4 random warps each across the full 32-bit nonce range, at 64 MiB, 256 MiB and 1 GiB, plus
|
||||
the 7 programs in README.md, plus 14 hand-built edge cases with their operand values proven. Zero mismatches,
|
||||
zero compile failures.
|
||||
2. Deterministic GPU execution: 5 runs and 3 compiles (one forced cold) of the same program give the same
|
||||
8 MiB of output; the dataset fill is deterministic and matches the CPU closed form at sampled indices
|
||||
including 0 and MASK.
|
||||
3. Every dataset index in the emitted MSL and CUDA is masked, by static check on every program tested, and
|
||||
the mask is exercised by essentially every load.
|
||||
4. No obvious structural bias in the 64-bit output for 3 seeds at 2^20 nonces: bit frequencies, avalanche
|
||||
(mean 32.0, std 4.0, every output bit flips with probability 0.49 to 0.51), chi-square on four 16-bit
|
||||
windows, zero duplicates.
|
||||
5. The emitter's contract on the generator (rotl 1..31, masks in {1,2,4,8,16}, src != dst) held on every
|
||||
instruction of every fuzzed program.
|
||||
|
||||
### Not demonstrated
|
||||
|
||||
1. Cryptographic security of the construction. Nothing here speaks to preimage, second-preimage or collision
|
||||
resistance, or to an adversary who chooses nonces or influences the epoch seed. The per-register
|
||||
initialisation is `splitmix32(nonce ^ seed) ^ seed`, which is a bijection of the nonce per register, and
|
||||
the body is add-rotate-xor-multiply with `or` (which destroys information) and no non-linear table. None
|
||||
of that has been analysed.
|
||||
2. Resistance to algebraic or structural attacks on weak programs. 3 seeds were measured for bias; the
|
||||
population of programs was not. A program whose `or` chain saturates a register, or whose load addresses
|
||||
collapse to few values, would be both biased and shortcut-able, and nothing yet rejects such programs or
|
||||
bounds how often they occur.
|
||||
3. Memory-hardness. Section 7 measures the shortcut directly: with a closed-form dataset the honest kernel's
|
||||
memory traffic is optional. The README's "memory bound at 1 GiB" describes the honest kernel, not the
|
||||
fastest kernel. This is a design gap of the prototype, not a bug in the tests.
|
||||
4. Vendor independence. Only Apple Metal was run. The CUDA twin exists as text and passed a CPU emulation;
|
||||
it has not run on NVIDIA hardware, where `__shfl_xor_sync`, `__umulhi` and shift semantics must be shown
|
||||
to agree bit for bit with these vectors.
|
||||
5. The 10 ms CPU verification gate against a dataset that is expensive to derive. The current 0.02 ms per warp
|
||||
is with a six-operation element.
|
||||
|
||||
### The next three tests the real cryptographer must do before the spec is final
|
||||
|
||||
1. Weak-program census and rejection filter. Generate at least 10^5 programs. For each, measure on the CPU
|
||||
interpreter (no GPU needed): output bit bias and avalanche on 2^12 nonces, distinct load addresses per
|
||||
hash, fraction of registers saturated by `or`, and whether any register's final value is independent of
|
||||
the nonce. Report the distribution, define the rejection rule, and bound the advantage of a miner who can
|
||||
grind the epoch seed over candidate blocks.
|
||||
2. Replace the closed-form dataset with an expensive derivation (RandomX style: a cache of hundreds of MiB
|
||||
expanded by a slow hash, or Argon2-based) and re-run three measurements: the `--inline-dataset` shortcut
|
||||
ratio (must fall to about 1), the GPU rate at 1 GiB, and the CPU verify time against the 10 ms gate with
|
||||
on-demand element derivation. Also measure distinct cache lines touched per hash so a cache-resident
|
||||
shortcut is ruled out, not assumed.
|
||||
3. Cross-vendor bit-exactness and a written spec. Run the exported pack and the fuzz vectors on NVIDIA
|
||||
(RTX 5090, CUDA) and on at least one AMD GPU, with the same 4-warp-per-program comparison, including the
|
||||
edge-case programs. Then fix the hash as a specification with test vectors and replace the ad-hoc seed
|
||||
derivation (FNV-1a plus SplitMix) with a standard hash so the seed-to-program mapping is auditable.
|
||||
|
||||
## Files
|
||||
|
||||
- `main.swift`: tests under `// MARK: - Hardening tests`, flags `--fuzz`, `--fuzz-seed`, `--edge`, `--stats`,
|
||||
`--determinism`, `--memcheck`, `--inline-dataset`.
|
||||
- `TESTS.md`: this file.
|
||||
- Raw logs of the runs quoted above were kept in the session scratchpad and are not checked in; every table
|
||||
is reproducible with the command above it.
|
||||
|
|
@ -24,6 +24,8 @@ struct Options {
|
|||
var stats = false // --stats: output distribution sanity checks on 2^20 nonces
|
||||
var determinism = false // --determinism: 5 identical GPU runs + double compile
|
||||
var memcheck = false // --memcheck: static mask check + 4 MiB run with wrapping nonces
|
||||
// Shortcut measurement (bench variant, not a test): compute dataset elements inline instead of loading them.
|
||||
var inlineDataset = false
|
||||
var anyTest: Bool { fuzz != nil || edge || stats || determinism || memcheck }
|
||||
}
|
||||
|
||||
|
|
@ -49,6 +51,7 @@ func parseArgs() -> Options {
|
|||
case "--stats": o.stats = true
|
||||
case "--determinism": o.determinism = true
|
||||
case "--memcheck": o.memcheck = true
|
||||
case "--inline-dataset": o.inlineDataset = true
|
||||
case "-h", "--help":
|
||||
print("""
|
||||
igneum-bench [--seed <string>] [--hours N] [--batch-log2 22] [--batches 4]
|
||||
|
|
@ -61,6 +64,8 @@ func parseArgs() -> Options {
|
|||
[--stats] output distribution sanity checks on 2^20 nonces, 3 seeds
|
||||
[--determinism] 5 identical GPU runs of 2^20 nonces, double compile, dataset fill check
|
||||
[--memcheck] static dataset-index mask check, 4 MiB run with wrapping nonces
|
||||
shortcut measurement:
|
||||
[--inline-dataset] bench variant: every load computes ds_elem(index) inline, no memory read
|
||||
""")
|
||||
exit(0)
|
||||
default:
|
||||
|
|
@ -184,7 +189,9 @@ func generateProgram(seedString: String) -> Program {
|
|||
|
||||
func hex(_ v: UInt32) -> String { String(format: "0x%08xu", v) }
|
||||
|
||||
func generateMSL(_ p: Program, datasetLog2: Int) -> String {
|
||||
// inlineDay: when set, every load computes ds_elem(index, day) in registers instead of reading the buffer.
|
||||
// Same function, no memory traffic. Used only by --inline-dataset to measure the closed-form shortcut.
|
||||
func generateMSL(_ p: Program, datasetLog2: Int, inlineDay: (UInt32, UInt32)? = nil) -> String {
|
||||
let mask = UInt32((1 << datasetLog2) - 1)
|
||||
var s = """
|
||||
#include <metal_stdlib>
|
||||
|
|
@ -236,7 +243,9 @@ func generateMSL(_ p: Program, datasetLog2: Int) -> String {
|
|||
case .rotr: line = "\(d) = rotr_var(\(d), \(a));"
|
||||
case .mad: line = "\(d) = \(a) * \(b) + \(d);"
|
||||
case .shfl: line = "\(d) = \(d) ^ simd_shuffle_xor(\(a), (ushort)\(ins.mask));"
|
||||
case .load: line = "\(d) = \(d) ^ dataset[\(a) & MASK];"
|
||||
case .load:
|
||||
if let dd = inlineDay { line = "\(d) = \(d) ^ ds_elem(\(a) & MASK, \(hex(dd.0)), \(hex(dd.1)));" }
|
||||
else { line = "\(d) = \(d) ^ dataset[\(a) & MASK];" }
|
||||
}
|
||||
s += " \(line) // \(k)\n"
|
||||
}
|
||||
|
|
@ -728,7 +737,8 @@ struct EpochResult {
|
|||
|
||||
func runEpoch(gpu: GPU, opts: Options, seedString: String, dataset: MTLBuffer, day: (UInt32, UInt32)) -> EpochResult {
|
||||
let program = generateProgram(seedString: seedString)
|
||||
let msl = generateMSL(program, datasetLog2: opts.datasetLog2)
|
||||
let msl = generateMSL(program, datasetLog2: opts.datasetLog2, inlineDay: opts.inlineDataset ? day : nil)
|
||||
if opts.inlineDataset { print("\nNOTE: --inline-dataset: loads compute ds_elem inline, the dataset buffer is never read") }
|
||||
if let dir = opts.dumpDir {
|
||||
try? FileManager.default.createDirectory(atPath: dir, withIntermediateDirectories: true)
|
||||
let safe = seedString.replacingOccurrences(of: "/", with: "_")
|
||||
|
|
@ -1308,35 +1318,44 @@ func runStats(_ opts: Options, gpu: GPU, day: (UInt32, UInt32)) -> Bool {
|
|||
var dups = 0
|
||||
for i in 1..<n where sorted[i] == sorted[i - 1] { dups += 1 }
|
||||
|
||||
// (b) avalanche: 1000 random nonces, flip bit (i mod 32), count changed output bits. Run on the GPU.
|
||||
let sw = seedWords("avalanche/" + seedString)
|
||||
var rng = SplitMix64(s: UInt64(sw[0]) | (UInt64(sw[1]) << 32))
|
||||
let trials = 1000
|
||||
var bases = [UInt32]()
|
||||
var flipped = [Int]()
|
||||
for t in 0..<trials {
|
||||
let nonce = UInt32(truncatingIfNeeded: rng.next())
|
||||
let bit = t % 32
|
||||
bases.append(nonce); bases.append(nonce ^ (1 << UInt32(bit)))
|
||||
flipped.append(bit)
|
||||
// (b) avalanche: random nonces, flip bit (t mod 32), count changed output bits. Run on the GPU.
|
||||
// The brief asks for 1,000 trials; a 16,000-trial pass (500 per input bit) is added because the
|
||||
// standard error of the mean at 1,000 trials is 0.13 bits, too coarse to see a small bias.
|
||||
func avalanche(_ trials: Int, salt: String) -> (mean: Double, std: Double, minD: Int, maxD: Int, perBitMin: Double, perBitMax: Double, perOutMin: Double, perOutMax: Double)? {
|
||||
let sw = seedWords("avalanche/\(salt)/" + seedString)
|
||||
var rng = SplitMix64(s: UInt64(sw[0]) | (UInt64(sw[1]) << 32))
|
||||
var bases = [UInt32]()
|
||||
var flipped = [Int]()
|
||||
for t in 0..<trials {
|
||||
let nonce = UInt32(truncatingIfNeeded: rng.next())
|
||||
let bit = t % 32
|
||||
bases.append(nonce); bases.append(nonce ^ (1 << UInt32(bit)))
|
||||
flipped.append(bit)
|
||||
}
|
||||
guard let av = gpuWarps(gpu, k, dataset: dataset, bases: bases) else { return nil }
|
||||
var diffs = [Int]()
|
||||
var perBitSum = [Int](repeating: 0, count: 32), perBitN = [Int](repeating: 0, count: 32)
|
||||
var perOut = [Int](repeating: 0, count: 64)
|
||||
var minDiff = 64, maxDiff = 0
|
||||
for t in 0..<trials {
|
||||
let x = av[2 * t][0] ^ av[2 * t + 1][0]
|
||||
let d = popcount64(x)
|
||||
diffs.append(d)
|
||||
perBitSum[flipped[t]] += d; perBitN[flipped[t]] += 1
|
||||
for b in 0..<64 where (x >> UInt64(b)) & 1 == 1 { perOut[b] += 1 }
|
||||
minDiff = min(minDiff, d); maxDiff = max(maxDiff, d)
|
||||
}
|
||||
let mean = Double(diffs.reduce(0, +)) / Double(trials)
|
||||
let variance = diffs.reduce(0.0) { $0 + (Double($1) - mean) * (Double($1) - mean) } / Double(trials - 1)
|
||||
var pbMin = 64.0, pbMax = 0.0
|
||||
for b in 0..<32 where perBitN[b] > 0 { let m = Double(perBitSum[b]) / Double(perBitN[b]); pbMin = min(pbMin, m); pbMax = max(pbMax, m) }
|
||||
let poMin = Double(perOut.min()!) / Double(trials), poMax = Double(perOut.max()!) / Double(trials)
|
||||
return (mean, variance.squareRoot(), minDiff, maxDiff, pbMin, pbMax, poMin, poMax)
|
||||
}
|
||||
guard let av = gpuWarps(gpu, k, dataset: dataset, bases: bases) else { print("FAIL: avalanche GPU run"); return false }
|
||||
var diffs = [Int]()
|
||||
var perBitSum = [Int](repeating: 0, count: 32), perBitN = [Int](repeating: 0, count: 32)
|
||||
var minDiff = 64, maxDiff = 0
|
||||
for t in 0..<trials {
|
||||
let d = popcount64(av[2 * t][0] ^ av[2 * t + 1][0])
|
||||
diffs.append(d)
|
||||
perBitSum[flipped[t]] += d; perBitN[flipped[t]] += 1
|
||||
minDiff = min(minDiff, d); maxDiff = max(maxDiff, d)
|
||||
}
|
||||
let mean = Double(diffs.reduce(0, +)) / Double(trials)
|
||||
let variance = diffs.reduce(0.0) { $0 + (Double($1) - mean) * (Double($1) - mean) } / Double(trials - 1)
|
||||
let std = variance.squareRoot()
|
||||
var perBitMin = 64.0, perBitMax = 0.0
|
||||
for b in 0..<32 where perBitN[b] > 0 { let m = Double(perBitSum[b]) / Double(perBitN[b]); perBitMin = min(perBitMin, m); perBitMax = max(perBitMax, m) }
|
||||
// Expected for an ideal function: mean 32, std 4 (binomial 64 x 0.5). Standard error of the mean over 1000 trials is 0.13.
|
||||
let avOk = abs(mean - 32) < 0.6 && std > 3.3 && std < 4.7
|
||||
guard let a1 = avalanche(1000, salt: "small"), let a2 = avalanche(16000, salt: "large") else { print("FAIL: avalanche GPU run"); return false }
|
||||
// Expected for an ideal function: mean 32, std 4 (binomial 64 x 0.5). Standard error of the mean:
|
||||
// 0.13 bits at 1,000 trials, 0.03 bits at 16,000. Thresholds are about 4.5 standard errors.
|
||||
let avOk = abs(a1.mean - 32) < 0.6 && a1.std > 3.3 && a1.std < 4.7 && abs(a2.mean - 32) < 0.15 && a2.std > 3.6 && a2.std < 4.4
|
||||
let freqOk = maxZ < 4.5
|
||||
let chiOk = chiWorstZ < 4.5
|
||||
let dupOk = dups == 0
|
||||
|
|
@ -1345,14 +1364,15 @@ func runStats(_ opts: Options, gpu: GPU, day: (UInt32, UInt32)) -> Bool {
|
|||
|
||||
print("\nseed \"\(seedString)\": loads/hash \(program.loadsPerHash), GPU \(fmt(gms, 1)) ms for 2^20 hashes, CPU spot check 2 warps \(spot ? "PASS" : "FAIL")")
|
||||
print(" (a) bit frequency: min \(fmt(minFreq, 4)) max \(fmt(maxFreq, 4)); largest deviation \(fmt(maxDev, 0)) counts at bit \(maxBit) = \(fmt(maxZ, 2)) sigma (sigma \(fmt(sigma, 0)), 64 bits, expect max under about 3.5)")
|
||||
print(" (b) avalanche over \(trials) single-bit nonce flips: mean \(fmt(mean, 2)) std \(fmt(std, 2)) min \(minDiff) max \(maxDiff) of 64 bits (expect mean 32, std 4); per-input-bit mean range \(fmt(perBitMin, 1))..\(fmt(perBitMax, 1))")
|
||||
print(" (b) avalanche, 1000 single-bit nonce flips: mean \(fmt(a1.mean, 2)) std \(fmt(a1.std, 2)) min \(a1.minD) max \(a1.maxD) of 64 bits (expect mean 32, std 4); per-input-bit mean range \(fmt(a1.perBitMin, 1))..\(fmt(a1.perBitMax, 1))")
|
||||
print(" (b) avalanche, 16000 flips (500 per input bit): mean \(fmt(a2.mean, 3)) std \(fmt(a2.std, 2)) min \(a2.minD) max \(a2.maxD); per-input-bit mean range \(fmt(a2.perBitMin, 2))..\(fmt(a2.perBitMax, 2)); per-output-bit flip probability range \(fmt(a2.perOutMin, 3))..\(fmt(a2.perOutMax, 3)) (expect 0.5, sigma 0.004)")
|
||||
for r in chiRows { print(" (c) \(r)") }
|
||||
print(" (d) duplicate 64-bit outputs among 2^20: \(dups) (expected about 3e-8)")
|
||||
print(" verdict: \(ok ? "no obvious bias" : "SUSPECT")")
|
||||
rows.append("| \(seedString) | \(program.loadsPerHash) | \(fmt(minFreq, 4))..\(fmt(maxFreq, 4)) | \(fmt(maxZ, 2)) | \(fmt(mean, 2)) | \(fmt(std, 2)) | \(fmt(chiWorstZ, 2)) | \(dups) | \(ok ? "uniform-looking" : "SUSPECT") |")
|
||||
rows.append("| \(seedString) | \(program.loadsPerHash) | \(fmt(minFreq, 4))..\(fmt(maxFreq, 4)) | \(fmt(maxZ, 2)) | \(fmt(a1.mean, 2)) / \(fmt(a1.std, 2)) | \(fmt(a2.mean, 3)) / \(fmt(a2.std, 2)) | \(fmt(a2.perOutMin, 3))..\(fmt(a2.perOutMax, 3)) | \(fmt(chiWorstZ, 2)) | \(dups) | \(ok ? "uniform-looking" : "SUSPECT") |")
|
||||
}
|
||||
print("\n| Seed | Loads/hash | Bit freq min..max | Max bit z | Avalanche mean | Avalanche std | Worst chi2 z (4 windows) | Dups | Verdict |")
|
||||
print("|---|---|---|---|---|---|---|---|---|")
|
||||
print("\n| Seed | Loads/hash | Bit freq min..max | Max bit z | Avalanche 1k mean / std | Avalanche 16k mean / std | Per-output-bit flip prob | Worst chi2 z (4 windows) | Dups | Verdict |")
|
||||
print("|---|---|---|---|---|---|---|---|---|---|")
|
||||
for r in rows { print(r) }
|
||||
print("STATS: \(allOk ? "PASS (no obvious structural bias; not a security proof)" : "FAIL (something looks biased)")")
|
||||
return allOk
|
||||
|
|
@ -1376,10 +1396,14 @@ func runDeterminism(_ opts: Options, gpu: GPU, day: (UInt32, UInt32)) -> Bool {
|
|||
print("generator: two generations of the program give identical MSL source: \(sameSource ? "yes" : "NO") (\(msl1.utf8.count) bytes)")
|
||||
if !sameSource { ok = false }
|
||||
|
||||
// Two separate compiles of the same source.
|
||||
let k1: CompiledHash, k2: CompiledHash
|
||||
do { k1 = try compileHash(gpu, msl: msl1); k2 = try compileHash(gpu, msl: msl2) } catch { print("FAIL: compile \(error)"); return false }
|
||||
print("compiled twice: \(fmt(k1.totalMs, 1)) ms and \(fmt(k2.totalMs, 1)) ms")
|
||||
// Three compiles: the same source twice (Metal's shader cache may serve the second), and once more with a
|
||||
// comment tag appended so the cache misses and the compiler really runs again on an identical kernel.
|
||||
let tag = "\n// recompile tag \(h64(nowNs()))\n"
|
||||
let k1: CompiledHash, k2: CompiledHash, k3: CompiledHash
|
||||
do {
|
||||
k1 = try compileHash(gpu, msl: msl1); k2 = try compileHash(gpu, msl: msl2); k3 = try compileHash(gpu, msl: msl2 + tag)
|
||||
} catch { print("FAIL: compile \(error)"); return false }
|
||||
print("compiled three times: identical source \(fmt(k1.totalMs, 1)) ms and \(fmt(k2.totalMs, 1)) ms (a sub-millisecond second figure means the system shader cache answered), tagged source \(fmt(k3.totalMs, 1)) ms (forced recompile)")
|
||||
|
||||
func runOnce(_ k: CompiledHash) -> (fp: UInt64, sentinels: Int, ms: Double)? {
|
||||
memset(outBuf.contents(), 0xAA, n * 8)
|
||||
|
|
@ -1407,7 +1431,12 @@ func runDeterminism(_ opts: Options, gpu: GPU, day: (UInt32, UInt32)) -> Bool {
|
|||
var differ2 = 0
|
||||
for i in 0..<n where p[i] != reference[i] { differ2 += 1 }
|
||||
if differ2 != 0 || r2.sentinels != 0 { ok = false }
|
||||
print("run on compile 2: fingerprint \(h64(r2.fp)), \(differ2) outputs differ from compile 1 run 1, \(r2.sentinels) unwritten lanes: \(differ2 == 0 ? "identical" : "DIFFERENT")")
|
||||
print("run on compile 2 (identical source): fingerprint \(h64(r2.fp)), \(differ2) outputs differ from compile 1 run 1, \(r2.sentinels) unwritten lanes: \(differ2 == 0 ? "identical" : "DIFFERENT")")
|
||||
guard let r3 = runOnce(k3) else { print("FAIL: GPU run on compile 3"); return false }
|
||||
var differ3 = 0
|
||||
for i in 0..<n where p[i] != reference[i] { differ3 += 1 }
|
||||
if differ3 != 0 || r3.sentinels != 0 { ok = false }
|
||||
print("run on compile 3 (forced recompile): fingerprint \(h64(r3.fp)), \(differ3) outputs differ from compile 1 run 1, \(r3.sentinels) unwritten lanes: \(differ3 == 0 ? "identical" : "DIFFERENT")")
|
||||
|
||||
// CPU reference on the first and last warp and 6 others, so the fingerprint is tied to the verified function.
|
||||
var cpuBad = 0
|
||||
|
|
|
|||
Loading…
Reference in a new issue