Memory-hard dataset: 256 MiB ChaCha cache, 8 dependent reads per item, CPU verifier on the cache, levers, CUDA pack igneum-genesis-mh
proto-metal: default dataset is now the memory-hard construction (MEMHARD.md), --closed-form keeps the original. Cache fill 2 ms GPU / 185 ms one CPU core; dataset build 20.6 ms; GPU cache == CPU cache on all 2^26 words. Shortcut ratio: inline kernel 111x faster than honest (closed form) to 4.8x slower (memory-hard), 1 GiB. CPU verify 0.63 to 0.80 ms per warp at 104 loads, 1.21 ms at 144 loads (4,608 items): 10 ms gate met. Levers --load-weight and --wide-frac implemented and measured, both off; default generator unchanged. Fuzz 200/200, edge, determinism, memcheck, stats re-run on the new dataset, all PASS. proto-cuda: host.cu handles both dataset modes; new pack igneum-genesis-mh with memhard.h; clang emulation PASS including the three-way cache check. Old packs unchanged; closed-form export is byte-identical to them. docs/bench-log.md: dated summary. Note: a concurrent session running git commit -a swept earlier states of these files into its site commits (7b28d5e through d6539fa); this commit carries the remainder. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
bbb264a1b1
commit
58a5a636b7
4 changed files with 73 additions and 18 deletions
|
|
@ -111,3 +111,15 @@ E: active FAILS the partition test: 50/50 honest split, no attacker, both sides
|
|||
F: delayed eclipse of a 20% pool is harmless (participation 0 after 2 h, back in 2 h, 0 conflicts); a 34% attacker poisoning that pool finalises a private fork in 49 min under active, never under total.
|
||||
Floor hybrid: active denominator never below 0.85 x total (lock needs 56.7% of total) gives 0 conflicts in every partition and eclipse, recovers in 0 / 13 min at 34 / 40% silent and 2 min at 35% churn; costs 4.1 days at 50% churn and liveness ends near 42% silent. Floor 0.80 does not stop the eclipse (54% > 53.3%).
|
||||
Recommend: active/cert + floor 0.85, presence 240, dust 100, quorum 2/3, grace at least 3x worst delay. Not modelled: real GHOSTDAG merge and post-heal fork choice, DAA lag, VRF aggregators, certificate revocation.
|
||||
|
||||
## 2026-10-03 proto-metal memory-hard dataset (cache + 8 dependent reads), Metal only; CUDA pack emulated
|
||||
|
||||
Machine: the same Apple M5 Max (one performance core for the CPU figures). Construction, every table and the commands are in `proto-metal/MEMHARD.md`. Default dataset is now memory-hard; `--closed-form` keeps the original for comparison.
|
||||
Construction: 256 MiB cache = 2^22 lines of 64 B in 2^16 chains of 64 ChaCha12 blocks with feed-forward (in_j = prev ^ (sigma || K[8] || seg || j || tag)); item t = 16 words, 8 rounds of (seed-parameterised ARX-multiply mixer, read cache line s[0] & (2^22-1), xor) plus a final mixer; dataset[w] = item(w >> 4)[w & 15]. Hash kernel unchanged.
|
||||
Cache fill: 2.0 ms GPU (0.6 to 2.1 across runs), 185 ms one CPU core (Swift), 162 ms C++ host reference. Dataset build 1 GiB: 20.6 ms GPU (29.4 first in process), 814 M items/s, 6.5 G cache-line reads/s. GPU cache == CPU cache on all 2^26 words every run (FNV-1a 64 48c4f5bf24166b2e for day 2026-10-03).
|
||||
Shortcut ratio, seed igneum-genesis, 1 GiB: honest 45.2 Mhash/s in both constructions. Inline kernel (never reads the dataset): closed form 5,014 Mhash/s (111x FASTER than honest); memory-hard 9.49 Mhash/s (0.21 of honest, 4.8x SLOWER). At a 256 MiB dataset: honest 94.8, inline 9.48 (0.10).
|
||||
CPU verify per 32-lane warp (holds only the cache, derives every word on demand, 32 lanes interleaved): 0.649 / 0.631 / 0.701 ms for igneum-genesis, /epoch1, /epoch2 (104, 104, 112 loads; 3,328 to 3,584 items); 0.801 ms igneum-second-seed (104 loads); 1.205 ms igneum-second-seed/epoch1 (144 loads, 4,608 items). Cold single warps 1.16 to 2.11 ms. Closed form was 0.017 ms. 10 ms GATE MET, margin about 8x steady.
|
||||
Levers (implemented, measured, OFF by default; default generator unchanged): (a) --load-weight 17: 72 to 80 loads/hash, CPU 0.457 to 0.512 ms/warp, GPU 55.0 to 73.4 Mhash/s. (b) --wide-frac 50 (warp-coalesced 128 B loads): CPU 0.233 to 0.489 ms/warp, GPU 56.1 to 135.2 Mhash/s and useful bandwidth up to 56 GB/s, so (b) erodes the random-access bound. (a)+(b): CPU 0.223 to 0.276, GPU 106 to 139. Recommendation: no lever; (a) is the fallback if a slower verifier ever threatens the gate; (b) not recommended.
|
||||
Tests re-run on the new dataset: fuzz 200/200 (800 warps, 25,600 hashes, 0 mismatches, CPU interpreter 1.23 s), edge 14/14, determinism PASS (fingerprint 62a4f0eb018df273), memcheck PASS, stats PASS (3 seeds, no obvious bias). 3 warps x 3 seeds bit-exact in the bench run.
|
||||
CUDA: new pack proto-cuda/packs/igneum-genesis-mh (kernel.cu with cache-fill and build kernels, memhard.h shared by device and host, vectors incl. cache head/last/FNV and 64 sampled words). host.cu handles both modes; old packs unchanged (closed-form export re-run is byte-identical in kernel.cu and program.metal). clang emulation (emu/emu.sh igneum-genesis-mh): cache check PASS (all words, FNV == Mac), dataset self-test PASS at 1 GiB, 3/3 vectors standalone and 2/2 in batch at 2 warps/block. RTX 5090 and AMD runs of this pack PENDING; no NVIDIA figure for the memory-hard dataset exists.
|
||||
Not demonstrated: cross-vendor results for the new dataset; the shortcut ratio on a discrete GPU; time-memory trade-offs between the two measured points; cryptographic strength of the mixer and the chained cache; distinct-lines-per-hash census.
|
||||
|
|
|
|||
|
|
@ -63,6 +63,21 @@ defined identically in Metal Shading Language, CUDA C++ and Swift's wrapping ope
|
|||
| Fill | one thread per word, threadgroup 256 | one thread per word, block 256, with an `i < n` guard (a no-op for power-of-two sizes) |
|
||||
| Self-test | not needed on the Mac (CPU and GPU share one process) | head 16 words and word `[MASK]` against values the Mac wrote into `vectors.h`; 64 pseudo-random words against `host_ds_elem` |
|
||||
|
||||
## Dataset, memory-hard pack (added later on 3 October 2026)
|
||||
|
||||
Pack `igneum-genesis-mh`, `IGNEUM_DATASET_MODE 1`. Construction in `../proto-metal/MEMHARD.md`.
|
||||
|
||||
| Item | Metal (Mac) | CUDA pack | Host reference (host.cu) | CPU verifier (main.swift) |
|
||||
|---|---|---|---|---|
|
||||
| Core text (ChaCha12 block, cache segment chain, mixer M_r, item derivation) | `emitMemhardCore(cuda: false)`, constants as literals | `memhard.h`: `emitMemhardCore(cuda: true)`, the same emitter, `IGNEUM_HD` = `__host__ __device__` | the same `memhard.h`, compiled as plain C++ (`IGNEUM_HD` = `static inline`) | separate Swift implementation (`chachaBlock`, `cpuFillSegment`, `mixer`, `deriveItems`) |
|
||||
| Cache fill | `igneum_cache_fill`, one thread per segment | `igneum_cache_fill<<<256, 256>>>`, `if (seg < nSegments)` | `mh_cache_segment` for all 65,536 segments on one thread | `cpuFillCache`, one core |
|
||||
| Dataset build | `igneum_build`, one thread per item, writes 16 words | `igneum_build<<<items/256, 256>>>`, `if (t < nItems)`, writes through a `d` pointer (no `ds[` write, so the static mask check expects 0 fill writes) | not built; words derived on demand with `mh_word` | not built; `MemhardCPU.fetch` |
|
||||
| Agreement | GPU cache == CPU cache on all 2^26 words in every run; 1,024 sampled GPU dataset words == CPU derivation; 21 bench warps, 800 fuzz warps, edge, determinism, memcheck PASS | emulation: emulated-GPU cache == host cache on all words, host FNV == Mac FNV, dataset self-test PASS at 1 GiB, 3/3 vectors standalone, 2/2 in batch at 2 warps per block | | |
|
||||
|
||||
Residual risk specific to this pack: `mh_item` reads a cache line as 16 scalar `uint32_t` loads through a `const uint32_t*`;
|
||||
nvcc may or may not vectorise them. That affects speed on NVIDIA, not results. The 256 MiB cache plus 1 GiB dataset
|
||||
need 1.25 GiB of device memory plus the output buffer.
|
||||
|
||||
## How each claim above was verified
|
||||
|
||||
1. Mac, Metal GPU against the CPU interpreter: the regular bench (`proto-metal/igneum-bench`) passed 3 warps per
|
||||
|
|
|
|||
|
|
@ -6,8 +6,10 @@ them bit for bit against results produced on the Mac.
|
|||
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset,
|
||||
checks the GPU against known answers, and times the kernel. Nothing here earns anything.
|
||||
|
||||
Status on 3 October 2026: packs exported and cross-checked on the Mac, CUDA run on the RTX 5090 pending.
|
||||
No NVIDIA hash rate has been measured. Any figure you see for NVIDIA in this repo before the 5090 run is wrong.
|
||||
Status on 3 October 2026: the two closed-form packs ran on the RTX 5090 (96/96 vectors PASS each, see
|
||||
`docs/bench-log.md`). The memory-hard pack `igneum-genesis-mh` (added later the same day, construction in
|
||||
`../proto-metal/MEMHARD.md`) has passed only the clang emulation on the Mac; its 5090 and AMD runs are pending, and no
|
||||
NVIDIA figure for the memory-hard dataset exists yet.
|
||||
|
||||
## Layout
|
||||
|
||||
|
|
@ -24,11 +26,18 @@ proto-cuda/
|
|||
program.json the instruction list and all constants, for any other implementation
|
||||
vectors.json the same vectors as JSON
|
||||
program.metal the Metal source the Mac ran, for diffing by eye
|
||||
memhard.h memory-hard packs only: the cache fill and item derivation core, compiled for device and host
|
||||
memhard.metal memory-hard packs only: the Metal cache-fill and build kernels the Mac ran
|
||||
emu/ CPU emulation shim: compile and check a pack with plain clang++/g++, no GPU
|
||||
```
|
||||
|
||||
Two packs are checked in: `igneum-genesis` (104 loads per hash) and `igneum-hourly` (128 loads per hash).
|
||||
Both were cross-checked on the Mac's Metal GPU before being written.
|
||||
Three packs are checked in. `igneum-genesis` (104 loads per hash) and `igneum-hourly` (128 loads per hash) use the
|
||||
original closed-form dataset (`IGNEUM_DATASET_MODE 0`, implied when the macro is absent). `igneum-genesis-mh` is the
|
||||
same program as `igneum-genesis` over the memory-hard dataset (`IGNEUM_DATASET_MODE 1`): a 256 MiB cache of chained
|
||||
ChaCha12 blocks filled on the GPU from the day key, and every 64-byte dataset item derived from 8 dependent cache
|
||||
reads through a seed-parameterised mixer (`../proto-metal/MEMHARD.md`). The hash kernel text is identical in both
|
||||
packs; only the dataset contents differ, so the 96 expected outputs differ. All three were cross-checked on the
|
||||
Mac's Metal GPU before being written.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
|
|
@ -52,6 +61,7 @@ Linux:
|
|||
cd proto-cuda
|
||||
./build.sh # pack igneum-genesis, -arch=sm_120
|
||||
./build.sh igneum-hourly # the second pack
|
||||
./build.sh igneum-genesis-mh # the memory-hard pack (needs 256 MiB more device memory for the cache)
|
||||
./build.sh igneum-genesis native # if sm_120 is refused, let nvcc pick the installed GPU
|
||||
```
|
||||
|
||||
|
|
@ -61,6 +71,7 @@ Windows (x64 Native Tools Command Prompt):
|
|||
cd proto-cuda
|
||||
build.bat
|
||||
build.bat igneum-hourly
|
||||
build.bat igneum-genesis-mh
|
||||
build.bat igneum-genesis native
|
||||
```
|
||||
|
||||
|
|
@ -85,6 +96,8 @@ Notes
|
|||
./igneum-bench-cuda-igneum-genesis --sweep # 4, 64, 256, 512, 1024 MiB in sequence (the Mac's sweep)
|
||||
./igneum-bench-cuda-igneum-genesis --block-warps 4 # 4 warps per block instead of 1 (still bit-exact)
|
||||
./igneum-bench-cuda-igneum-hourly # the second program
|
||||
./igneum-bench-cuda-igneum-genesis-mh # memory-hard dataset: cache fill + build, cache check, vectors, bench
|
||||
./igneum-bench-cuda-igneum-genesis-mh --sweep # the same sweep over the memory-hard dataset
|
||||
```
|
||||
|
||||
On Windows the binaries are `igneum-bench-cuda-igneum-genesis.exe` and so on.
|
||||
|
|
@ -95,23 +108,33 @@ Flags: `--dataset-mib N` (power of two, default 1024), `--sweep`, `--batch-log2
|
|||
|
||||
What it prints, in order:
|
||||
1. GPU name, SM count, memory, clocks, L2, warp size, driver and runtime versions, registers per thread and
|
||||
resident warps per SM for the kernel.
|
||||
2. Dataset fill time (twice; the Mac saw a first-touch cost on the first fill) and write GB/s.
|
||||
3. Dataset self-test: 16 head words and word `[MASK]` against values from the Mac, 64 random words against the
|
||||
host formula.
|
||||
4. Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again
|
||||
resident warps per SM for the kernel, and which dataset construction the pack uses.
|
||||
2. Memory-hard packs only: cache fill time on the GPU (twice), cache fill time on the host (one thread, the same
|
||||
`memhard.h` text), then the cache check: every one of the 2^26 words GPU versus host, the host FNV-1a 64 against
|
||||
the Mac's, and the head and last line against the Mac's. A FAIL here stops nothing but fails OVERALL.
|
||||
3. Dataset fill time (closed form, twice, with write GB/s) or dataset build time (memory-hard, twice, with items/s
|
||||
and cache-line reads/s).
|
||||
4. Dataset self-test: 16 head words and word `[MASK]` against values from the Mac, 64 random words against the
|
||||
host formula (closed form) or the host derivation from the host cache (memory-hard), and the Mac's 64 sampled
|
||||
words (memory-hard packs; those inside the current dataset size).
|
||||
5. Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again
|
||||
read out of the warm-up batch so the bench configuration itself is checked. PASS or FAIL per warp, with the
|
||||
first differing lane printed on FAIL.
|
||||
5. Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s,
|
||||
6. Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s,
|
||||
GB/s useful (loads per hash x 4 bytes x hashes/s, the same definition as the Mac's table).
|
||||
6. A summary table in Markdown and `OVERALL: PASS` or `FAIL`. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.
|
||||
7. A summary table in Markdown and `OVERALL: PASS` or `FAIL`. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.
|
||||
|
||||
Vectors are only checked when the dataset is the pack's size (1024 MiB), because the outputs depend on the
|
||||
address mask. At other sizes the table says "skipped (not pack size)" and only the dataset self-test counts.
|
||||
|
||||
## What PASS means
|
||||
|
||||
- The CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not every word).
|
||||
- Closed-form packs: the CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not
|
||||
every word).
|
||||
- Memory-hard pack: the CUDA cache-fill kernel produced, word for word, the same 256 MiB cache as the host and as the
|
||||
Mac (FNV-1a 64, head, last line), and the CUDA build kernel produced the same dataset words as the host derivation
|
||||
and the Mac's samples (sampled, not every word). The whole chain from day key to dataset word agrees across three
|
||||
compilers (Apple Metal, host C++, NVIDIA CUDA).
|
||||
- For 96 nonces spread across the nonce space, the RTX 5090 produced the same 64-bit outputs as the Mac's CPU
|
||||
interpreter, which had itself matched the Mac's Metal GPU. The random program, the register init, the warp
|
||||
shuffles, the multiply-high and rotates, and the dataset addressing all agree between Apple and NVIDIA.
|
||||
|
|
@ -132,7 +155,9 @@ the Mac's entry. Do not edit the numbers; if a run looks odd, run it again and l
|
|||
`emu/emu.sh <pack> [flags]` compiles `host.cu` and the pack's `kernel.cu` as plain C++17 against a shim
|
||||
`cuda_runtime.h` and runs the kernels on host threads (32 per warp, a barrier inside `__shfl_xor_sync`). Use
|
||||
small batches (`--batch-log2 13 --batches 1`). Only PASS/FAIL matters; the rates it prints are noise.
|
||||
This is how the CUDA text was checked on the Mac on 3 October 2026 (both packs PASS, see `CHECKLIST.md`).
|
||||
This is how the CUDA text was checked on the Mac on 3 October 2026 (all three packs PASS, see `CHECKLIST.md`). The
|
||||
memory-hard pack builds the 1 GiB dataset on host threads, which takes about a second on the M5 Max; the cache fill
|
||||
on one host thread took 161 ms.
|
||||
It is not an nvcc build and says nothing about NVIDIA hardware.
|
||||
|
||||
## Regenerating a pack
|
||||
|
|
@ -142,9 +167,11 @@ On the Mac:
|
|||
```
|
||||
cd proto-metal
|
||||
swiftc -O -o igneum-bench main.swift -framework Metal
|
||||
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
|
||||
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh # memory-hard (default)
|
||||
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
|
||||
```
|
||||
|
||||
The exporter runs the CPU interpreter for the three vector warps, runs the Metal kernel for the same warps,
|
||||
and refuses to write anything unless all 96 outputs match. `--day` and `--dataset-log2` change the dataset
|
||||
and refuses to write anything unless all 96 outputs match. For a memory-hard pack it also refuses unless the GPU
|
||||
cache equals the CPU cache on every word and the sampled GPU dataset words equal the CPU derivation. `--day` and `--dataset-log2` change the dataset
|
||||
constants and are recorded in the pack.
|
||||
|
|
|
|||
|
|
@ -9,9 +9,10 @@ Every number below was produced on this machine on this date by the commands sho
|
|||
Later on 3 October 2026 the default dataset became the memory-hard construction of `MEMHARD.md`. The tables in this
|
||||
file are from the original closed-form dataset and reproduce with `--closed-form` added to each command. Every test
|
||||
here was re-run on the new dataset with the same commands and passed; those results are in `MEMHARD.md` section 2.5.
|
||||
The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one. The tests live in
|
||||
`main.swift` next to the bench and share its generator, MSL emitter, dataset fill and CPU interpreter
|
||||
without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench
|
||||
The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one.
|
||||
|
||||
The tests live in `main.swift` next to the bench and share its generator, MSL emitter, dataset build and CPU
|
||||
interpreter without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench
|
||||
does not use; `cpuWarp` calls it with `nil`.
|
||||
|
||||
What the GPU-vs-CPU comparison means: the GPU runs the generated Metal kernel, the CPU runs a plain Swift
|
||||
|
|
|
|||
Loading…
Reference in a new issue