diff --git a/proto-metal/README.md b/proto-metal/README.md index f9089e388..3517c1539 100644 --- a/proto-metal/README.md +++ b/proto-metal/README.md @@ -12,9 +12,13 @@ and compiled at runtime with `MTLDevice.makeLibrary(source:options:)`. across the 32-lane SIMD group, masks 1 to 16), load (`dst ^= dataset[src & MASK]`). Load weight 25 percent. Each add picks one of two immediates from a bit of r0 sampled at the top of the iteration, branchless `select`. 3. Emits Metal Shading Language, compiles it, runs it with threadgroup size 32 (one SIMD group per threadgroup). -4. Fills a 1 GiB dataset (2^28 uint32) on the GPU from a closed-form function of (daySeed, index). - The CPU computes any element on demand and never holds the dataset. -5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit. +4. Builds a 1 GiB dataset (2^28 uint32) on the GPU. Since 3 October 2026 (later the same day) the default is the + memory-hard construction of `MEMHARD.md`: a 256 MiB cache of chained ChaCha12 blocks filled from the day seed, + and each 64-byte item derived by 8 dependent cache reads through a seed-parameterised ARX mixer. The original + closed-form element is still available with `--closed-form`. The hash kernel is the same in both modes. +5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit. The CPU + verifier holds the 256 MiB cache (computed on one core, compared word for word with the GPU's) and derives every + dataset word on demand; it never holds the dataset. 6. `--hours N` regenerates and recompiles N programs in sequence (the hourly epoch model). Default N is 2 so verification always covers two seeds. @@ -39,7 +43,9 @@ Tested with Swift 5.8.1 from Command Line Tools on macOS (Darwin 25.6.0), no Xco ``` Flags: `--seed`, `--day`, `--hours`, `--batch-log2` (default 22), `--batches` (default 4), -`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump `, `--export-pack `. +`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump `, `--export-pack `, +`--closed-form` (original dataset), `--load-weight W` and `--wide-frac P` (generator levers, defaults 25 and 0 +reproduce the default generator exactly; see `MEMHARD.md` section 2.4). Exit code 0 means every verified warp matched. Hardening tests (added 3 October 2026, results and commands in `TESTS.md`): `--fuzz N [--fuzz-seed ]`, @@ -52,6 +58,16 @@ emits it a second time as a CUDA kernel, computes expected outputs for 3 warps ( with the CPU interpreter, cross-checks them on the Metal GPU, and writes `kernel.cu`, `program.h`, `vectors.h`, `program.json`, `vectors.json` and `program.metal` into the directory. See `../proto-cuda/README.md`. +## Memory-hard dataset (3 October 2026, later the same day) + +`MEMHARD.md` specifies the construction and holds every measurement. Headline, Apple M5 Max, 1 GiB, seed +igneum-genesis: honest kernel 45.2 Mhash/s in both constructions; the inline shortcut kernel went from 5,014 Mhash/s +(closed form, 111x faster than honest) to 9.49 Mhash/s (memory-hard, 4.8x slower than honest); CPU verification +0.63 to 0.80 ms per 32-lane warp at 104 loads per hash and 1.21 ms at 144 loads, against the 10 ms gate; cache fill +2 ms on the GPU and 185 ms on one CPU core; dataset build 20.6 ms. Fuzz, edge, determinism, memcheck and stats were +re-run on the new dataset and pass. The tables below are the original closed-form measurements and still reproduce +with `--closed-form`. + ## Measured on this machine, 3 October 2026 Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory. `threadExecutionWidth` reported as 32. @@ -112,18 +128,20 @@ like the system shader cache. 9. Measured later on 3 October 2026 (`TESTS.md` section 7): because the dataset element is a six-operation closed form, a kernel that computes it inline instead of loading it runs at about 4,900 Mhash/s against 44.6 for the honest kernel at 1 GiB, roughly 110x. Observation 1 describes the honest kernel only. The - prototype is not memory-hard until the dataset element costs more to derive than to load. + prototype is not memory-hard until the dataset element costs more to derive than to load. Resolved later the + same day: with the memory-hard dataset the inline kernel runs at 0.21 of the honest rate (`MEMHARD.md`). ## What to try next - Wider loads (uint4, 16 bytes per load) so the useful bytes approach what the memory system actually moves. - Two or more independent address chains per lane, to see whether the rate is latency or throughput limited. -- A dataset element that costs real work to derive, then re-measure the CPU verify time against the 10 ms gate. +- Done 3 October 2026: a dataset element that costs real work to derive (`MEMHARD.md`); CPU verify re-measured at 0.63 to 1.21 ms per warp. - A program-quality filter in the generator (for example, reject programs where OR saturates a register). - Output distribution tests on the 64-bit results before this hash is used for leader election. - The same kernel on a discrete GPU, through a CUDA or Vulkan port of the generator. ## Files -- `main.swift`: generator, MSL emitter, GPU driver, CPU interpreter, CLI. +- `main.swift`: generator, MSL emitter, GPU driver, CPU interpreter, memory-hard dataset (cache, items, verifier), CLI. +- `MEMHARD.md`: the memory-hard dataset construction and its measurements. - `igneum-bench`: the built binary (not checked in by intent; rebuild with the command above). diff --git a/proto-metal/TESTS.md b/proto-metal/TESTS.md index 7c4ef0ae6..f2717fbbe 100644 --- a/proto-metal/TESTS.md +++ b/proto-metal/TESTS.md @@ -4,7 +4,12 @@ Date: 3 October 2026. Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, Swift 5.8.1 from Command Line Tools, no Xcode, Metal shaders compiled at runtime. Build: `swiftc -O -o igneum-bench main.swift -framework Metal` (4 s). -Every number below was produced on this machine on this date by the commands shown. The tests live in +Every number below was produced on this machine on this date by the commands shown. + +Later on 3 October 2026 the default dataset became the memory-hard construction of `MEMHARD.md`. The tables in this +file are from the original closed-form dataset and reproduce with `--closed-form` added to each command. Every test +here was re-run on the new dataset with the same commands and passed; those results are in `MEMHARD.md` section 2.5. +The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one. The tests live in `main.swift` next to the bench and share its generator, MSL emitter, dataset fill and CPU interpreter without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench does not use; `cpuWarp` calls it with `nil`. diff --git a/site/index.html b/site/index.html index 44b87a4db..a16668d19 100644 --- a/site/index.html +++ b/site/index.html @@ -249,7 +249,7 @@ footer .wrap{padding-block:48px 32px}
-
simulated preview ยท how testnet will look
+
simulated preview
blocks 0proven 0locked 0