docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes, the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8 two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The lane's standing rule: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is least efficient per op. Chip rows in sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by placeholders in the READMEs and the run script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| results | ||
| bench.cu | ||
| cpu_rows.c | ||
| kernel.cu | ||
| kernel_sd.cu | ||
| memhard.h | ||
| program.h | ||
| README.md | ||
| run_cpu.sh | ||
| run_gpu.sh | ||
| sd.h | ||
| vectors.h | ||
| verify_sd.c | ||
state-dataset: the dataset commits to chain state (class "sd1")
Horizon lane 8, new proof of work. Prototype and measurements, 6 October 2026, 19:34 to 19:50 UTC. Everything here is a TEST HARNESS: no pool, no network, no wallet, nothing touches the devnet.
The scheme in one paragraph
The hash kernel is unchanged. What changes is the daily dataset build. Today item t of the 1 GiB dataset is
mh_item(cache, t): the 16-word state starts as s[0..7] = K (day key) and s[8..15] = t * MUL[i] + RC[i], then
8 rounds of (8 mixers + one dependent 64-byte cache read XORed in), then 8 final mixers. Under sd1 a 64-byte STATE
LEAF for item t is XORed into those 16 initial words before the first mixer: s[i] ^= leaf(t)[i]. In the real design
leaf(t) is the t-th 64-byte leaf of a canonical serialisation of the chain's execution state at a certified checkpoint
20 minutes before the day boundary, zero padded where the state is shorter than the dataset. A miner therefore cannot
build the day's dataset without the state, and a verifier that holds no dataset needs leaf(t) for every item it
derives. The prototype measures what that costs (build time, device memory, verifier time) and buys (which leaves a
hash touches, and what a light client would have to be shown).
Leaf stand-in (synthetic, prototype only)
leaf(t) = mh_chacha_block(x) (the pack's own ChaCha12 block from memhard.h) with
x = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574, S[0..7], t, 0, 0x49676e65, 0x53746174) and
S[i] = K[i] ^ 0x5a5a5a5a (stand-in state root). For the pack's day key
S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102. The leaf array is 2^24 items x 64 B = 1 GiB
for the pack's 1 GiB dataset (2^28 words = 2^24 items). Definition in sd.h.
Files
| file | what |
|---|---|
kernel.cu, memhard.h, program.h, vectors.h |
the mx8-genesis pack, copied unmodified from proto-cuda/packs-ca2-mixer/mx8-genesis/ |
sd.h |
mh_leaf, mh_item_sd, mh_word_sd (host and device, plain C compatible) |
kernel_sd.cu |
the pack's kernel.cu plus igneum_leaves, igneum_build_sd and their launch wrappers appended; igneum_hash byte for byte the pack's (checked with diff) |
bench.cu |
the harness: `--mode control |
verify_sd.c |
plain C: host cache fill, item bit-exact check from the GPU's items file, 32-lane register-major interpreter of the program against the GPU's dump, lane-0 load trace, per-unit rows |
cpu_rows.c |
plain C: the igneum-build-1 rows B.1 and B.2 (OpenMP for the 32-thread row) |
run_gpu.sh, run_cpu.sh |
the exact commands, as run |
results/gpu/, results/cpu/ |
every log, the dumps and the items files, copied back from the boxes |
Boxes
- GPU box 2: NVIDIA GeForce RTX 4090, 24 GB (24083 MiB reported), 128 SMs, driver 595.91.07 (CUDA driver 13.2), nvcc 12.8.93,
-arch=sm_89, power limit 450 W, Ubuntu 24.04.1. Host CPU AMD EPYC 7352 (48 threads). A CPU-only pool daemon shares the host (ports 4463/4480); the plain C rows there were pinned to core 2. Working directory/root/horizon-newpow/state-dataset. - CPU box igneum-build-1: AMD EPYC 9454P (48 cores, 96 threads), 128 GB, gcc 13.3.0, L3 256 MiB. Shared with other agents'
builds: load average 19 to 32 during the readings (printed in each log). Rows pinned to core 4 under
nice -n 19; the 32-thread row on cores 4 to 35. Working directory/srv/builds/horizon-newpow/state-dataset.
Commands
From the Mac (zsh; the GPU box has no rsync, tar over ssh instead):
cd /Users/joshm/Projects/igneum-wt-horizon/proto-newpow/state-dataset
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.cu *.h *.c *.sh | ssh -i ~/.ssh/igneum-fleet -p <box-2-port> root@<box-2-ip> 'cd /root/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_gpu.sh > run_gpu.log 2>&1'
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.c *.h *.sh | ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 'cd /srv/builds/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_cpu.sh > run_cpu.log 2>&1'
On GPU box 2 (run_gpu.sh):
export PATH=/usr/local/cuda/bin:$PATH
nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
gcc -O2 -std=c11 -o verify_sd verify_sd.c
./bench --mode control --batches 10 --items-out items_control.txt --dump dump_control.txt 4 --power-seconds 20
./bench --mode sd1 --batches 10 --items-out items_sd1.txt --dump dump_sd1.txt 4 --power-seconds 20
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 64
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 256
./bench --mode control --batches 10 --power-seconds 20 # second pass
./bench --mode sd1 --batches 10 --power-seconds 20 # second pass
taskset -c 2 ./verify_sd --mode control --items items_control.txt --dump dump_control.txt
taskset -c 2 ./verify_sd --mode sd1 --items items_sd1.txt --dump dump_sd1.txt
On igneum-build-1 (run_cpu.sh):
gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
gcc -O2 -std=c11 -o verify_sd verify_sd.c
taskset -c 4 nice -n 19 ./cpu_rows --rows # B.1 (i)(ii)(iii) and the one-core B.2 rows, twice
taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 # B.2 leaf array on 32 threads, twice
taskset -c 4 nice -n 19 ./verify_sd --mode sd1 # the per-unit rows of verify_sd.c on this CPU
Timing method: cache fill, leaf array and dataset build are CUDA-event times of the second of two launches (as
host.cu does). MH/s is CUDA-event time over 10 batches of 2^24 after a warm-up batch at base 0. Watts and SM MHz are
the mean of nvidia-smi -l 1 samples taken after the first 10 s of a 20 s window in which the hash kernel runs back
to back (10 samples used of 22). Fingerprint = FNV-1a 64 over the 2^24 little-endian u64 outputs at base nonce 0,
the project's definition.
RESULTS
A. GPU, RTX 4090 (control vs sd1, two passes each)
| row | control | sd1 | note |
|---|---|---|---|
| cache fill, GPU, second pass (ms) | 1.81, 1.85 | 1.85, 1.85 | 256 MiB, unchanged kernel |
| leaf array, GPU, second pass (ms) | none | 5.49, 5.49 | 2^24 ChaCha12 blocks, 1 GiB written at 195 GB/s |
| dataset build, second pass (ms) | 30.55, 30.55 | 31.98, 31.98 | +1.43 ms (+4.7%): one coalesced 64 B read per item |
| build + leaves (ms) | 30.55 | 37.47 | synthetic leaves; a real snapshot arrives from the node instead (see chunked rows) |
| the 3 Mac vectors, standalone and in batch | PASS, PASS | n/a (new values) | sd1 base 0 lane 0 = b600edbed969becc, base 4096 = a533e89c78bb6b74, base 1000000 = beb4cb0c6bab8163 |
| dataset head, word [MASK], 64 Mac samples | PASS | n/a (new values) | sd1 dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf |
| 64 random words vs host derivation | PASS (mh_word) | PASS (mh_word_sd) | |
| 2^24 fingerprint at base 0 | 7c28cfb06c5c65a9 (= the pack's, both passes) | d5b0c16390cad0e8 (new, four runs) | |
| hash rate, 10 batches of 2^24 (MH/s) | 63.083, 63.083 | 63.088, 63.087 | equal within 0.01%; the hash kernel is the same binary over a dataset of the same shape |
| hash rate in the 20 s power window (MH/s) | 63.078, 63.079 | 63.083, 63.083 | |
| watts, mean after 10 s | 207.6, 205.2 | 207.9, 206.5 | run to run noise 2 W |
| SM MHz, mem MHz | 2745, 10251 | 2745, 10251 | |
| MH/s per W | 0.304, 0.307 | 0.303, 0.305 | |
| hash kernel registers | 29 | 29 | 24 resident blocks/SM at 1 warp/block |
Bit-exactness of the build (A.3): 1,024 random item indices (SplitMix64 seeded 0x5d1, t = z & (2^24 - 1)), the 16 GPU
words of each item against the host derivation in plain C on the box's CPU with the host-filled cache (65536
mh_cache_segment calls, 380 ms in the bench, 573 to 587 ms pinned to core 2 in verify_sd):
| check | control | sd1 |
|---|---|---|
| items equal, in-process (bench.cu, host mh_item / mh_item_sd) | 1024 of 1024 | 1024 of 1024 |
| items equal, plain C (verify_sd.c from the items file) | 1024 of 1024 | 1024 of 1024 |
| items equal after the chunked rebuild (64 MiB chunks; 256 MiB chunks) | n/a | 1024 of 1024; 1024 of 1024 |
| 64 device leaves vs host mh_leaf (incl. t = 0 and 2^24 - 1) | n/a | PASS |
| host cache FNV-1a 64 | 48c4f5bf24166b2e = Mac | 48c4f5bf24166b2e = Mac |
GPU hash outputs vs the C interpreter (4 warps dumped, bases 0, 32, 64, 96; loads through mh_word / mh_word_sd):
| control | sd1 | |
|---|---|---|
| lanes equal | 128 of 128 | 128 of 128 |
| interpreter vs the Mac vector at base 0 | 32 of 32 | n/a |
| interpreting time (4,096 item derivations per warp) | 37 ms | 42 ms |
Device memory (A.4), cudaMemGetInfo, MiB used (context 395 included):
| phase | control | sd1 resident leaves | sd1 chunked 64 MiB | sd1 chunked 256 MiB |
|---|---|---|---|---|
| during the build | 1675 | 2699 | 1741 | 1933 |
| while hashing (leaves freed) | 1803 | 1803 | 1803 | 1803 |
| chunked rebuild total, copies + builds, one stream (ms) | 75.80 | 75.47 |
Whole-array host to device copy of the 1 GiB leaf array from pinned memory: 62 ms = 17.2 GB/s (this box's PCIe link
under load; device to host 55 ms). So the chunked build is PCIe-bound: 62 ms of copy plus the 32 ms build, partly
serialised on one stream, gives 76 ms. Two chunk buffers on two streams would hide most of the build under the copy
(not done; the kernel already takes an item range [t0, t0 + n) with the chunk's leaves at leaves - 16 t0, so
chunking needed no kernel change at all, only the loop in bench.cu).
B. CPU, igneum-build-1 (EPYC 9454P, one core pinned, nice 19), ms per unit of 4,096 items, 100 units
Two readings, because the box is shared: reading 1 at load 19, reading 2 at load 31.
| row | reading 1 | reading 2 | per item |
|---|---|---|---|
| (i) 4,096 random 64 B reads from a 2 GiB resident leaf array | 0.108 ms (min 0.092, max 0.172) | 0.163 ms | 26 to 40 ns |
| (i) the same from an 8 GiB array (the year-12 size) | 0.139 ms (min 0.127, max 0.183) | 0.209 ms | 34 to 51 ns |
| (ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array) | 0.284 ms | 0.285 ms | 69 ns |
| (iii) 4,096 x mh_item on the host cache, naive, no interleaving | 9.01 ms (min 8.91, max 9.60) | 11.24 ms (min 9.38, max 13.51) | 2.2 to 2.7 us |
| 4,096 x (leaf + mh_item_sd), naive | 9.31 ms | 11.64 ms |
The same rows on GPU box 2's EPYC 7352, core 2 (verify_sd.c): (ii) 0.369 to 0.374 ms, (iii) 10.14 to 10.21 ms, leaf + mh_item_sd 10.56 to 10.58 ms.
The project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive (iii) figure here is an upper bound. The sd1 increment is the row that matters: 0.11 to 0.21 ms per unit when the verifier holds the 2 to 8 GiB state in RAM, or 0.28 ms when it derives the stand-in leaf (in the real design it cannot derive a state leaf, it must hold the state or be shown it). Against the 2.06 ms interleaved verifier that is +5% to +14%; against the naive 9 to 11 ms it is +1% to +3%.
B.2, the daily snapshot build on CPU (igneum-build-1):
| row | time |
|---|---|
| 1 GiB leaf array, 2^24 ChaCha12 blocks, one core | 1.670 s, 1.677 s (100 ns per leaf) |
| the same on 32 threads (OpenMP), second pass (first pass 0.172 s with page faults) | 0.073 s, 0.083 s |
| host cache fill, 256 MiB, one core | 0.447 s, 0.452 s (0.380 s on the 4090 box's host, unpinned; 0.57 s pinned beside the pool daemon) |
B.3, the state sample a light client would check (lane 0 of warp base 0, from the instrumented interpreter):
| control | sd1 | |
|---|---|---|
| distinct words touched by the 128 loads | 128 | 128 |
| distinct items touched | 128 of 2^24 | 128 of 2^24 |
| first 8 load words | 8528400 43034175 255822847 136709537 ... | 8528400 85793927 255822847 180392080 ... |
Loads 1 and 3 share their address in both modes because their addresses are fixed by the nonce before any load's value feeds back; from load 2 on the sd1 trail diverges. Merkle arithmetic for a 2 GiB state (2^25 leaves of 64 B, depth 25, 32-byte hashes): 128 openings x 25 x 32 B = 102,400 B of siblings + 128 x 64 B = 8,192 B of leaves = 110,592 B (108 KiB) per lane, uncompressed; a 32-lane unit is 32 x 108 KiB = 3.4 MiB, before deduplicating shared upper levels (128 random paths in a depth-25 tree share only their top seven levels, so dedup saves about 20% per lane, approximate; across the 32 lanes of a unit the same seven levels are shared again). Every one of the 128 loads is a distinct item, so there is no saving from repeated items.
What the numbers mean, per tier (the consequences rule)
- Hash rate, watts, MH/s per W: unchanged for every tier on every vendor, because the hash kernel is the pack's and the dataset has the same shape. Nothing to do.
- Daily build: +1.4 ms on the 4090 (30.6 to 32.0 ms) when the leaves are resident, 76 ms when streamed from the host over PCIe in chunks. At the designed 2 GiB dataset the leaf array is 2 GiB and the streamed build is about 150 ms (approximate, scaling the 17 GB/s copy); on a PCIe 3 x8 slot in a rig, about 4x that (approximate). All far inside the daily window on every tier.
- Device memory: with resident leaves the build peak at 2 GiB is 2 GiB + 256 MiB + 2 GiB + context, about 4.6 GiB
(approximate), which fits an 8 GB card but not beside a prover. Streamed in 64 MiB chunks the peak is dataset + cache
- 64 MiB + context, measured 1741 MiB at 1 GiB here, about 2.7 GiB at 2 GiB (approximate): no tier loses memory it has today. Recommendation: ship chunked only; never hold the leaf array on the device.
- The real cost is delivery: every miner needs the 1 to 2 GiB state serialisation once a day. A miner beside its own node reads it from disk or loopback (seconds). A pool user without a node must get it from the pool (2 GiB per day per miner, or a shared download), and a light verifier either holds the state (0.11 to 0.21 ms per unit extra, measured) or is shown 3.4 MiB of Merkle openings per unit (arithmetic above), which is not a light client any more. The design decision this prototype leaves open is which of those two the protocol asks of a header verifier.
- CPU snapshot build: 1.7 s on one core, 0.07 s on 32, for the synthetic leaves; the real serialisation is bounded by the node's state read, not by this.
What failed, what was cut
- Nothing on the list was cut. verify_sd.c's lane-0 trace printed nothing on the first run (the trace counter was reset on every warp, so it read 0 after warp 3); fixed, the verifiers re-run, the fix is in the file.
- The GPU box has no rsync; the sources went over with tar through ssh (
COPYFILE_DISABLE=1 --no-xattrs, the Mac's tar otherwise writes Apple xattr headers that GNU tar warns about). -std=c11hidesclock_gettime; both C files define_POSIX_C_SOURCE 200809L.- igneum-build-1 was under other agents' load (19 to 32) for every reading; both readings are given. The 4090 box's CPU rows ran beside the pool daemon, pinned to core 2. The GPU itself was idle apart from this bench.
- Not measured: a two-stream chunked build (copy and build overlapped), AMD and Apple builds, the 2 GiB dataset itself (the pack is 1 GiB; the 2 GiB figures above are labelled approximate).