Memory-hard dataset: 256 MiB ChaCha cache, 8 dependent reads per item, CPU verifier on the cache, levers, CUDA pack igneum-genesis-mh

proto-metal: default dataset is now the memory-hard construction (MEMHARD.md), --closed-form keeps the original.
Cache fill 2 ms GPU / 185 ms one CPU core; dataset build 20.6 ms; GPU cache == CPU cache on all 2^26 words.
Shortcut ratio: inline kernel 111x faster than honest (closed form) to 4.8x slower (memory-hard), 1 GiB.
CPU verify 0.63 to 0.80 ms per warp at 104 loads, 1.21 ms at 144 loads (4,608 items): 10 ms gate met.
Levers --load-weight and --wide-frac implemented and measured, both off; default generator unchanged.
Fuzz 200/200, edge, determinism, memcheck, stats re-run on the new dataset, all PASS.
proto-cuda: host.cu handles both dataset modes; new pack igneum-genesis-mh with memhard.h; clang emulation PASS
including the three-way cache check. Old packs unchanged; closed-form export is byte-identical to them.
docs/bench-log.md: dated summary.

Note: a concurrent session running git commit -a swept earlier states of these files into its site commits
(7b28d5e through d6539fa); this commit carries the remainder.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-03 16:12:00 +00:00
parent 3fc61b0dff
commit 3b00e8b480
4 changed files with 73 additions and 18 deletions

View file

@ -111,3 +111,15 @@ E: active FAILS the partition test: 50/50 honest split, no attacker, both sides
F: delayed eclipse of a 20% pool is harmless (participation 0 after 2 h, back in 2 h, 0 conflicts); a 34% attacker poisoning that pool finalises a private fork in 49 min under active, never under total.
Floor hybrid: active denominator never below 0.85 x total (lock needs 56.7% of total) gives 0 conflicts in every partition and eclipse, recovers in 0 / 13 min at 34 / 40% silent and 2 min at 35% churn; costs 4.1 days at 50% churn and liveness ends near 42% silent. Floor 0.80 does not stop the eclipse (54% > 53.3%).
Recommend: active/cert + floor 0.85, presence 240, dust 100, quorum 2/3, grace at least 3x worst delay. Not modelled: real GHOSTDAG merge and post-heal fork choice, DAA lag, VRF aggregators, certificate revocation.
## 2026-10-03 proto-metal memory-hard dataset (cache + 8 dependent reads), Metal only; CUDA pack emulated
Machine: the same Apple M5 Max (one performance core for the CPU figures). Construction, every table and the commands are in `proto-metal/MEMHARD.md`. Default dataset is now memory-hard; `--closed-form` keeps the original for comparison.
Construction: 256 MiB cache = 2^22 lines of 64 B in 2^16 chains of 64 ChaCha12 blocks with feed-forward (in_j = prev ^ (sigma || K[8] || seg || j || tag)); item t = 16 words, 8 rounds of (seed-parameterised ARX-multiply mixer, read cache line s[0] & (2^22-1), xor) plus a final mixer; dataset[w] = item(w >> 4)[w & 15]. Hash kernel unchanged.
Cache fill: 2.0 ms GPU (0.6 to 2.1 across runs), 185 ms one CPU core (Swift), 162 ms C++ host reference. Dataset build 1 GiB: 20.6 ms GPU (29.4 first in process), 814 M items/s, 6.5 G cache-line reads/s. GPU cache == CPU cache on all 2^26 words every run (FNV-1a 64 48c4f5bf24166b2e for day 2026-10-03).
Shortcut ratio, seed igneum-genesis, 1 GiB: honest 45.2 Mhash/s in both constructions. Inline kernel (never reads the dataset): closed form 5,014 Mhash/s (111x FASTER than honest); memory-hard 9.49 Mhash/s (0.21 of honest, 4.8x SLOWER). At a 256 MiB dataset: honest 94.8, inline 9.48 (0.10).
CPU verify per 32-lane warp (holds only the cache, derives every word on demand, 32 lanes interleaved): 0.649 / 0.631 / 0.701 ms for igneum-genesis, /epoch1, /epoch2 (104, 104, 112 loads; 3,328 to 3,584 items); 0.801 ms igneum-second-seed (104 loads); 1.205 ms igneum-second-seed/epoch1 (144 loads, 4,608 items). Cold single warps 1.16 to 2.11 ms. Closed form was 0.017 ms. 10 ms GATE MET, margin about 8x steady.
Levers (implemented, measured, OFF by default; default generator unchanged): (a) --load-weight 17: 72 to 80 loads/hash, CPU 0.457 to 0.512 ms/warp, GPU 55.0 to 73.4 Mhash/s. (b) --wide-frac 50 (warp-coalesced 128 B loads): CPU 0.233 to 0.489 ms/warp, GPU 56.1 to 135.2 Mhash/s and useful bandwidth up to 56 GB/s, so (b) erodes the random-access bound. (a)+(b): CPU 0.223 to 0.276, GPU 106 to 139. Recommendation: no lever; (a) is the fallback if a slower verifier ever threatens the gate; (b) not recommended.
Tests re-run on the new dataset: fuzz 200/200 (800 warps, 25,600 hashes, 0 mismatches, CPU interpreter 1.23 s), edge 14/14, determinism PASS (fingerprint 62a4f0eb018df273), memcheck PASS, stats PASS (3 seeds, no obvious bias). 3 warps x 3 seeds bit-exact in the bench run.
CUDA: new pack proto-cuda/packs/igneum-genesis-mh (kernel.cu with cache-fill and build kernels, memhard.h shared by device and host, vectors incl. cache head/last/FNV and 64 sampled words). host.cu handles both modes; old packs unchanged (closed-form export re-run is byte-identical in kernel.cu and program.metal). clang emulation (emu/emu.sh igneum-genesis-mh): cache check PASS (all words, FNV == Mac), dataset self-test PASS at 1 GiB, 3/3 vectors standalone and 2/2 in batch at 2 warps/block. RTX 5090 and AMD runs of this pack PENDING; no NVIDIA figure for the memory-hard dataset exists.
Not demonstrated: cross-vendor results for the new dataset; the shortcut ratio on a discrete GPU; time-memory trade-offs between the two measured points; cryptographic strength of the mixer and the chained cache; distinct-lines-per-hash census.

View file

@ -63,6 +63,21 @@ defined identically in Metal Shading Language, CUDA C++ and Swift's wrapping ope
| Fill | one thread per word, threadgroup 256 | one thread per word, block 256, with an `i < n` guard (a no-op for power-of-two sizes) |
| Self-test | not needed on the Mac (CPU and GPU share one process) | head 16 words and word `[MASK]` against values the Mac wrote into `vectors.h`; 64 pseudo-random words against `host_ds_elem` |
## Dataset, memory-hard pack (added later on 3 October 2026)
Pack `igneum-genesis-mh`, `IGNEUM_DATASET_MODE 1`. Construction in `../proto-metal/MEMHARD.md`.
| Item | Metal (Mac) | CUDA pack | Host reference (host.cu) | CPU verifier (main.swift) |
|---|---|---|---|---|
| Core text (ChaCha12 block, cache segment chain, mixer M_r, item derivation) | `emitMemhardCore(cuda: false)`, constants as literals | `memhard.h`: `emitMemhardCore(cuda: true)`, the same emitter, `IGNEUM_HD` = `__host__ __device__` | the same `memhard.h`, compiled as plain C++ (`IGNEUM_HD` = `static inline`) | separate Swift implementation (`chachaBlock`, `cpuFillSegment`, `mixer`, `deriveItems`) |
| Cache fill | `igneum_cache_fill`, one thread per segment | `igneum_cache_fill<<<256, 256>>>`, `if (seg < nSegments)` | `mh_cache_segment` for all 65,536 segments on one thread | `cpuFillCache`, one core |
| Dataset build | `igneum_build`, one thread per item, writes 16 words | `igneum_build<<<items/256, 256>>>`, `if (t < nItems)`, writes through a `d` pointer (no `ds[` write, so the static mask check expects 0 fill writes) | not built; words derived on demand with `mh_word` | not built; `MemhardCPU.fetch` |
| Agreement | GPU cache == CPU cache on all 2^26 words in every run; 1,024 sampled GPU dataset words == CPU derivation; 21 bench warps, 800 fuzz warps, edge, determinism, memcheck PASS | emulation: emulated-GPU cache == host cache on all words, host FNV == Mac FNV, dataset self-test PASS at 1 GiB, 3/3 vectors standalone, 2/2 in batch at 2 warps per block | | |
Residual risk specific to this pack: `mh_item` reads a cache line as 16 scalar `uint32_t` loads through a `const uint32_t*`;
nvcc may or may not vectorise them. That affects speed on NVIDIA, not results. The 256 MiB cache plus 1 GiB dataset
need 1.25 GiB of device memory plus the output buffer.
## How each claim above was verified
1. Mac, Metal GPU against the CPU interpreter: the regular bench (`proto-metal/igneum-bench`) passed 3 warps per

View file

@ -6,8 +6,10 @@ them bit for bit against results produced on the Mac.
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset,
checks the GPU against known answers, and times the kernel. Nothing here earns anything.
Status on 3 October 2026: packs exported and cross-checked on the Mac, CUDA run on the RTX 5090 pending.
No NVIDIA hash rate has been measured. Any figure you see for NVIDIA in this repo before the 5090 run is wrong.
Status on 3 October 2026: the two closed-form packs ran on the RTX 5090 (96/96 vectors PASS each, see
`docs/bench-log.md`). The memory-hard pack `igneum-genesis-mh` (added later the same day, construction in
`../proto-metal/MEMHARD.md`) has passed only the clang emulation on the Mac; its 5090 and AMD runs are pending, and no
NVIDIA figure for the memory-hard dataset exists yet.
## Layout
@ -24,11 +26,18 @@ proto-cuda/
program.json the instruction list and all constants, for any other implementation
vectors.json the same vectors as JSON
program.metal the Metal source the Mac ran, for diffing by eye
memhard.h memory-hard packs only: the cache fill and item derivation core, compiled for device and host
memhard.metal memory-hard packs only: the Metal cache-fill and build kernels the Mac ran
emu/ CPU emulation shim: compile and check a pack with plain clang++/g++, no GPU
```
Two packs are checked in: `igneum-genesis` (104 loads per hash) and `igneum-hourly` (128 loads per hash).
Both were cross-checked on the Mac's Metal GPU before being written.
Three packs are checked in. `igneum-genesis` (104 loads per hash) and `igneum-hourly` (128 loads per hash) use the
original closed-form dataset (`IGNEUM_DATASET_MODE 0`, implied when the macro is absent). `igneum-genesis-mh` is the
same program as `igneum-genesis` over the memory-hard dataset (`IGNEUM_DATASET_MODE 1`): a 256 MiB cache of chained
ChaCha12 blocks filled on the GPU from the day key, and every 64-byte dataset item derived from 8 dependent cache
reads through a seed-parameterised mixer (`../proto-metal/MEMHARD.md`). The hash kernel text is identical in both
packs; only the dataset contents differ, so the 96 expected outputs differ. All three were cross-checked on the
Mac's Metal GPU before being written.
## Prerequisites
@ -52,6 +61,7 @@ Linux:
cd proto-cuda
./build.sh # pack igneum-genesis, -arch=sm_120
./build.sh igneum-hourly # the second pack
./build.sh igneum-genesis-mh # the memory-hard pack (needs 256 MiB more device memory for the cache)
./build.sh igneum-genesis native # if sm_120 is refused, let nvcc pick the installed GPU
```
@ -61,6 +71,7 @@ Windows (x64 Native Tools Command Prompt):
cd proto-cuda
build.bat
build.bat igneum-hourly
build.bat igneum-genesis-mh
build.bat igneum-genesis native
```
@ -85,6 +96,8 @@ Notes
./igneum-bench-cuda-igneum-genesis --sweep # 4, 64, 256, 512, 1024 MiB in sequence (the Mac's sweep)
./igneum-bench-cuda-igneum-genesis --block-warps 4 # 4 warps per block instead of 1 (still bit-exact)
./igneum-bench-cuda-igneum-hourly # the second program
./igneum-bench-cuda-igneum-genesis-mh # memory-hard dataset: cache fill + build, cache check, vectors, bench
./igneum-bench-cuda-igneum-genesis-mh --sweep # the same sweep over the memory-hard dataset
```
On Windows the binaries are `igneum-bench-cuda-igneum-genesis.exe` and so on.
@ -95,23 +108,33 @@ Flags: `--dataset-mib N` (power of two, default 1024), `--sweep`, `--batch-log2
What it prints, in order:
1. GPU name, SM count, memory, clocks, L2, warp size, driver and runtime versions, registers per thread and
resident warps per SM for the kernel.
2. Dataset fill time (twice; the Mac saw a first-touch cost on the first fill) and write GB/s.
3. Dataset self-test: 16 head words and word `[MASK]` against values from the Mac, 64 random words against the
host formula.
4. Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again
resident warps per SM for the kernel, and which dataset construction the pack uses.
2. Memory-hard packs only: cache fill time on the GPU (twice), cache fill time on the host (one thread, the same
`memhard.h` text), then the cache check: every one of the 2^26 words GPU versus host, the host FNV-1a 64 against
the Mac's, and the head and last line against the Mac's. A FAIL here stops nothing but fails OVERALL.
3. Dataset fill time (closed form, twice, with write GB/s) or dataset build time (memory-hard, twice, with items/s
and cache-line reads/s).
4. Dataset self-test: 16 head words and word `[MASK]` against values from the Mac, 64 random words against the
host formula (closed form) or the host derivation from the host cache (memory-hard), and the Mac's 64 sampled
words (memory-hard packs; those inside the current dataset size).
5. Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again
read out of the warm-up batch so the bench configuration itself is checked. PASS or FAIL per warp, with the
first differing lane printed on FAIL.
5. Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s,
6. Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s,
GB/s useful (loads per hash x 4 bytes x hashes/s, the same definition as the Mac's table).
6. A summary table in Markdown and `OVERALL: PASS` or `FAIL`. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.
7. A summary table in Markdown and `OVERALL: PASS` or `FAIL`. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.
Vectors are only checked when the dataset is the pack's size (1024 MiB), because the outputs depend on the
address mask. At other sizes the table says "skipped (not pack size)" and only the dataset self-test counts.
## What PASS means
- The CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not every word).
- Closed-form packs: the CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not
every word).
- Memory-hard pack: the CUDA cache-fill kernel produced, word for word, the same 256 MiB cache as the host and as the
Mac (FNV-1a 64, head, last line), and the CUDA build kernel produced the same dataset words as the host derivation
and the Mac's samples (sampled, not every word). The whole chain from day key to dataset word agrees across three
compilers (Apple Metal, host C++, NVIDIA CUDA).
- For 96 nonces spread across the nonce space, the RTX 5090 produced the same 64-bit outputs as the Mac's CPU
interpreter, which had itself matched the Mac's Metal GPU. The random program, the register init, the warp
shuffles, the multiply-high and rotates, and the dataset addressing all agree between Apple and NVIDIA.
@ -132,7 +155,9 @@ the Mac's entry. Do not edit the numbers; if a run looks odd, run it again and l
`emu/emu.sh <pack> [flags]` compiles `host.cu` and the pack's `kernel.cu` as plain C++17 against a shim
`cuda_runtime.h` and runs the kernels on host threads (32 per warp, a barrier inside `__shfl_xor_sync`). Use
small batches (`--batch-log2 13 --batches 1`). Only PASS/FAIL matters; the rates it prints are noise.
This is how the CUDA text was checked on the Mac on 3 October 2026 (both packs PASS, see `CHECKLIST.md`).
This is how the CUDA text was checked on the Mac on 3 October 2026 (all three packs PASS, see `CHECKLIST.md`). The
memory-hard pack builds the 1 GiB dataset on host threads, which takes about a second on the M5 Max; the cache fill
on one host thread took 161 ms.
It is not an nvcc build and says nothing about NVIDIA hardware.
## Regenerating a pack
@ -142,9 +167,11 @@ On the Mac:
```
cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh # memory-hard (default)
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
```
The exporter runs the CPU interpreter for the three vector warps, runs the Metal kernel for the same warps,
and refuses to write anything unless all 96 outputs match. `--day` and `--dataset-log2` change the dataset
and refuses to write anything unless all 96 outputs match. For a memory-hard pack it also refuses unless the GPU
cache equals the CPU cache on every word and the sampled GPU dataset words equal the CPU derivation. `--day` and `--dataset-log2` change the dataset
constants and are recorded in the pack.

View file

@ -9,9 +9,10 @@ Every number below was produced on this machine on this date by the commands sho
Later on 3 October 2026 the default dataset became the memory-hard construction of `MEMHARD.md`. The tables in this
file are from the original closed-form dataset and reproduce with `--closed-form` added to each command. Every test
here was re-run on the new dataset with the same commands and passed; those results are in `MEMHARD.md` section 2.5.
The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one. The tests live in
`main.swift` next to the bench and share its generator, MSL emitter, dataset fill and CPU interpreter
without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench
The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one.
The tests live in `main.swift` next to the bench and share its generator, MSL emitter, dataset build and CPU
interpreter without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench
does not use; `cpuWarp` calls it with `nil`.
What the GPU-vs-CPU comparison means: the GPU runs the generated Metal kernel, the CPU runs a plain Swift