Chain scene: header reads simulated preview
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
8117c988e3
commit
2dc2408433
3 changed files with 32 additions and 9 deletions
|
|
@ -12,9 +12,13 @@ and compiled at runtime with `MTLDevice.makeLibrary(source:options:)`.
|
|||
across the 32-lane SIMD group, masks 1 to 16), load (`dst ^= dataset[src & MASK]`). Load weight 25 percent.
|
||||
Each add picks one of two immediates from a bit of r0 sampled at the top of the iteration, branchless `select`.
|
||||
3. Emits Metal Shading Language, compiles it, runs it with threadgroup size 32 (one SIMD group per threadgroup).
|
||||
4. Fills a 1 GiB dataset (2^28 uint32) on the GPU from a closed-form function of (daySeed, index).
|
||||
The CPU computes any element on demand and never holds the dataset.
|
||||
5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit.
|
||||
4. Builds a 1 GiB dataset (2^28 uint32) on the GPU. Since 3 October 2026 (later the same day) the default is the
|
||||
memory-hard construction of `MEMHARD.md`: a 256 MiB cache of chained ChaCha12 blocks filled from the day seed,
|
||||
and each 64-byte item derived by 8 dependent cache reads through a seed-parameterised ARX mixer. The original
|
||||
closed-form element is still available with `--closed-form`. The hash kernel is the same in both modes.
|
||||
5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit. The CPU
|
||||
verifier holds the 256 MiB cache (computed on one core, compared word for word with the GPU's) and derives every
|
||||
dataset word on demand; it never holds the dataset.
|
||||
6. `--hours N` regenerates and recompiles N programs in sequence (the hourly epoch model). Default N is 2 so
|
||||
verification always covers two seeds.
|
||||
|
||||
|
|
@ -39,7 +43,9 @@ Tested with Swift 5.8.1 from Command Line Tools on macOS (Darwin 25.6.0), no Xco
|
|||
```
|
||||
|
||||
Flags: `--seed`, `--day`, `--hours`, `--batch-log2` (default 22), `--batches` (default 4),
|
||||
`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump <dir>`, `--export-pack <dir>`.
|
||||
`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump <dir>`, `--export-pack <dir>`,
|
||||
`--closed-form` (original dataset), `--load-weight W` and `--wide-frac P` (generator levers, defaults 25 and 0
|
||||
reproduce the default generator exactly; see `MEMHARD.md` section 2.4).
|
||||
Exit code 0 means every verified warp matched.
|
||||
|
||||
Hardening tests (added 3 October 2026, results and commands in `TESTS.md`): `--fuzz N [--fuzz-seed <string>]`,
|
||||
|
|
@ -52,6 +58,16 @@ emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (
|
|||
with the CPU interpreter, cross-checks them on the Metal GPU, and writes `kernel.cu`, `program.h`, `vectors.h`,
|
||||
`program.json`, `vectors.json` and `program.metal` into the directory. See `../proto-cuda/README.md`.
|
||||
|
||||
## Memory-hard dataset (3 October 2026, later the same day)
|
||||
|
||||
`MEMHARD.md` specifies the construction and holds every measurement. Headline, Apple M5 Max, 1 GiB, seed
|
||||
igneum-genesis: honest kernel 45.2 Mhash/s in both constructions; the inline shortcut kernel went from 5,014 Mhash/s
|
||||
(closed form, 111x faster than honest) to 9.49 Mhash/s (memory-hard, 4.8x slower than honest); CPU verification
|
||||
0.63 to 0.80 ms per 32-lane warp at 104 loads per hash and 1.21 ms at 144 loads, against the 10 ms gate; cache fill
|
||||
2 ms on the GPU and 185 ms on one CPU core; dataset build 20.6 ms. Fuzz, edge, determinism, memcheck and stats were
|
||||
re-run on the new dataset and pass. The tables below are the original closed-form measurements and still reproduce
|
||||
with `--closed-form`.
|
||||
|
||||
## Measured on this machine, 3 October 2026
|
||||
|
||||
Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory. `threadExecutionWidth` reported as 32.
|
||||
|
|
@ -112,18 +128,20 @@ like the system shader cache.
|
|||
9. Measured later on 3 October 2026 (`TESTS.md` section 7): because the dataset element is a six-operation
|
||||
closed form, a kernel that computes it inline instead of loading it runs at about 4,900 Mhash/s against
|
||||
44.6 for the honest kernel at 1 GiB, roughly 110x. Observation 1 describes the honest kernel only. The
|
||||
prototype is not memory-hard until the dataset element costs more to derive than to load.
|
||||
prototype is not memory-hard until the dataset element costs more to derive than to load. Resolved later the
|
||||
same day: with the memory-hard dataset the inline kernel runs at 0.21 of the honest rate (`MEMHARD.md`).
|
||||
|
||||
## What to try next
|
||||
|
||||
- Wider loads (uint4, 16 bytes per load) so the useful bytes approach what the memory system actually moves.
|
||||
- Two or more independent address chains per lane, to see whether the rate is latency or throughput limited.
|
||||
- A dataset element that costs real work to derive, then re-measure the CPU verify time against the 10 ms gate.
|
||||
- Done 3 October 2026: a dataset element that costs real work to derive (`MEMHARD.md`); CPU verify re-measured at 0.63 to 1.21 ms per warp.
|
||||
- A program-quality filter in the generator (for example, reject programs where OR saturates a register).
|
||||
- Output distribution tests on the 64-bit results before this hash is used for leader election.
|
||||
- The same kernel on a discrete GPU, through a CUDA or Vulkan port of the generator.
|
||||
|
||||
## Files
|
||||
|
||||
- `main.swift`: generator, MSL emitter, GPU driver, CPU interpreter, CLI.
|
||||
- `main.swift`: generator, MSL emitter, GPU driver, CPU interpreter, memory-hard dataset (cache, items, verifier), CLI.
|
||||
- `MEMHARD.md`: the memory-hard dataset construction and its measurements.
|
||||
- `igneum-bench`: the built binary (not checked in by intent; rebuild with the command above).
|
||||
|
|
|
|||
|
|
@ -4,7 +4,12 @@ Date: 3 October 2026. Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory,
|
|||
Swift 5.8.1 from Command Line Tools, no Xcode, Metal shaders compiled at runtime.
|
||||
Build: `swiftc -O -o igneum-bench main.swift -framework Metal` (4 s).
|
||||
|
||||
Every number below was produced on this machine on this date by the commands shown. The tests live in
|
||||
Every number below was produced on this machine on this date by the commands shown.
|
||||
|
||||
Later on 3 October 2026 the default dataset became the memory-hard construction of `MEMHARD.md`. The tables in this
|
||||
file are from the original closed-form dataset and reproduce with `--closed-form` added to each command. Every test
|
||||
here was re-run on the new dataset with the same commands and passed; those results are in `MEMHARD.md` section 2.5.
|
||||
The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one. The tests live in
|
||||
`main.swift` next to the bench and share its generator, MSL emitter, dataset fill and CPU interpreter
|
||||
without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench
|
||||
does not use; `cpuWarp` calls it with `nil`.
|
||||
|
|
|
|||
|
|
@ -249,7 +249,7 @@ footer .wrap{padding-block:48px 32px}
|
|||
</div>
|
||||
<div class="card viz reveal" aria-label="Animated block DAG, simulated preview">
|
||||
<div class="viz-head">
|
||||
<div class="eyebrow"><span class="dot" style="background:var(--ember);margin-right:8px"></span>simulated preview · how testnet will look</div>
|
||||
<div class="eyebrow"><span class="dot" style="background:var(--ember);margin-right:8px"></span>simulated preview</div>
|
||||
<div class="viz-stats mono"><span>blocks <b id="c-blocks">0</b></span><span>proven <b id="c-proven">0</b></span><span>locked <b id="c-locked">0</b></span></div>
|
||||
</div>
|
||||
<canvas id="chain" aria-hidden="true"></canvas>
|
||||
|
|
|
|||
Loading…
Reference in a new issue