Chain scene: header reads simulated preview

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-03 16:09:44 +00:00
parent 8117c988e3
commit 2dc2408433
3 changed files with 32 additions and 9 deletions

View file

@ -12,9 +12,13 @@ and compiled at runtime with `MTLDevice.makeLibrary(source:options:)`.
across the 32-lane SIMD group, masks 1 to 16), load (`dst ^= dataset[src & MASK]`). Load weight 25 percent.
Each add picks one of two immediates from a bit of r0 sampled at the top of the iteration, branchless `select`.
3. Emits Metal Shading Language, compiles it, runs it with threadgroup size 32 (one SIMD group per threadgroup).
4. Fills a 1 GiB dataset (2^28 uint32) on the GPU from a closed-form function of (daySeed, index).
The CPU computes any element on demand and never holds the dataset.
5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit.
4. Builds a 1 GiB dataset (2^28 uint32) on the GPU. Since 3 October 2026 (later the same day) the default is the
memory-hard construction of `MEMHARD.md`: a 256 MiB cache of chained ChaCha12 blocks filled from the day seed,
and each 64-byte item derived by 8 dependent cache reads through a seed-parameterised ARX mixer. The original
closed-form element is still available with `--closed-form`. The hash kernel is the same in both modes.
5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit. The CPU
verifier holds the 256 MiB cache (computed on one core, compared word for word with the GPU's) and derives every
dataset word on demand; it never holds the dataset.
6. `--hours N` regenerates and recompiles N programs in sequence (the hourly epoch model). Default N is 2 so
verification always covers two seeds.
@ -39,7 +43,9 @@ Tested with Swift 5.8.1 from Command Line Tools on macOS (Darwin 25.6.0), no Xco
```
Flags: `--seed`, `--day`, `--hours`, `--batch-log2` (default 22), `--batches` (default 4),
`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump <dir>`, `--export-pack <dir>`.
`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump <dir>`, `--export-pack <dir>`,
`--closed-form` (original dataset), `--load-weight W` and `--wide-frac P` (generator levers, defaults 25 and 0
reproduce the default generator exactly; see `MEMHARD.md` section 2.4).
Exit code 0 means every verified warp matched.
Hardening tests (added 3 October 2026, results and commands in `TESTS.md`): `--fuzz N [--fuzz-seed <string>]`,
@ -52,6 +58,16 @@ emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (
with the CPU interpreter, cross-checks them on the Metal GPU, and writes `kernel.cu`, `program.h`, `vectors.h`,
`program.json`, `vectors.json` and `program.metal` into the directory. See `../proto-cuda/README.md`.
## Memory-hard dataset (3 October 2026, later the same day)
`MEMHARD.md` specifies the construction and holds every measurement. Headline, Apple M5 Max, 1 GiB, seed
igneum-genesis: honest kernel 45.2 Mhash/s in both constructions; the inline shortcut kernel went from 5,014 Mhash/s
(closed form, 111x faster than honest) to 9.49 Mhash/s (memory-hard, 4.8x slower than honest); CPU verification
0.63 to 0.80 ms per 32-lane warp at 104 loads per hash and 1.21 ms at 144 loads, against the 10 ms gate; cache fill
2 ms on the GPU and 185 ms on one CPU core; dataset build 20.6 ms. Fuzz, edge, determinism, memcheck and stats were
re-run on the new dataset and pass. The tables below are the original closed-form measurements and still reproduce
with `--closed-form`.
## Measured on this machine, 3 October 2026
Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory. `threadExecutionWidth` reported as 32.
@ -112,18 +128,20 @@ like the system shader cache.
9. Measured later on 3 October 2026 (`TESTS.md` section 7): because the dataset element is a six-operation
closed form, a kernel that computes it inline instead of loading it runs at about 4,900 Mhash/s against
44.6 for the honest kernel at 1 GiB, roughly 110x. Observation 1 describes the honest kernel only. The
prototype is not memory-hard until the dataset element costs more to derive than to load.
prototype is not memory-hard until the dataset element costs more to derive than to load. Resolved later the
same day: with the memory-hard dataset the inline kernel runs at 0.21 of the honest rate (`MEMHARD.md`).
## What to try next
- Wider loads (uint4, 16 bytes per load) so the useful bytes approach what the memory system actually moves.
- Two or more independent address chains per lane, to see whether the rate is latency or throughput limited.
- A dataset element that costs real work to derive, then re-measure the CPU verify time against the 10 ms gate.
- Done 3 October 2026: a dataset element that costs real work to derive (`MEMHARD.md`); CPU verify re-measured at 0.63 to 1.21 ms per warp.
- A program-quality filter in the generator (for example, reject programs where OR saturates a register).
- Output distribution tests on the 64-bit results before this hash is used for leader election.
- The same kernel on a discrete GPU, through a CUDA or Vulkan port of the generator.
## Files
- `main.swift`: generator, MSL emitter, GPU driver, CPU interpreter, CLI.
- `main.swift`: generator, MSL emitter, GPU driver, CPU interpreter, memory-hard dataset (cache, items, verifier), CLI.
- `MEMHARD.md`: the memory-hard dataset construction and its measurements.
- `igneum-bench`: the built binary (not checked in by intent; rebuild with the command above).

View file

@ -4,7 +4,12 @@ Date: 3 October 2026. Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory,
Swift 5.8.1 from Command Line Tools, no Xcode, Metal shaders compiled at runtime.
Build: `swiftc -O -o igneum-bench main.swift -framework Metal` (4 s).
Every number below was produced on this machine on this date by the commands shown. The tests live in
Every number below was produced on this machine on this date by the commands shown.
Later on 3 October 2026 the default dataset became the memory-hard construction of `MEMHARD.md`. The tables in this
file are from the original closed-form dataset and reproduce with `--closed-form` added to each command. Every test
here was re-run on the new dataset with the same commands and passed; those results are in `MEMHARD.md` section 2.5.
The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one. The tests live in
`main.swift` next to the bench and share its generator, MSL emitter, dataset fill and CPU interpreter
without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench
does not use; `cpuWarp` calls it with `nil`.

View file

@ -249,7 +249,7 @@ footer .wrap{padding-block:48px 32px}
</div>
<div class="card viz reveal" aria-label="Animated block DAG, simulated preview">
<div class="viz-head">
<div class="eyebrow"><span class="dot" style="background:var(--ember);margin-right:8px"></span>simulated preview · how testnet will look</div>
<div class="eyebrow"><span class="dot" style="background:var(--ember);margin-right:8px"></span>simulated preview</div>
<div class="viz-stats mono"><span>blocks <b id="c-blocks">0</b></span><span>proven <b id="c-proven">0</b></span><span>locked <b id="c-locked">0</b></span></div>
</div>
<canvas id="chain" aria-hidden="true"></canvas>