igneum/proto-newpow/state-dataset
igneum-labs a664af6fc9 Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s
docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes,
the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as
proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8
two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the
execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's
CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The
lane's standing rule: a shadow lever only works through joules the honest card is forced to
spend, so shadow work goes where the GPU is least efficient per op. Chip rows in
sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by
placeholders in the READMEs and the run script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 20:18:39 +00:00
..
results Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
bench.cu Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
cpu_rows.c Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
kernel.cu Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
kernel_sd.cu Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
memhard.h Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
program.h Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
README.md Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
run_cpu.sh Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
run_gpu.sh Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
sd.h Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
vectors.h Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00
verify_sd.c Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s 2026-10-06 20:18:39 +00:00

state-dataset: the dataset commits to chain state (class "sd1")

Horizon lane 8, new proof of work. Prototype and measurements, 6 October 2026, 19:34 to 19:50 UTC. Everything here is a TEST HARNESS: no pool, no network, no wallet, nothing touches the devnet.

The scheme in one paragraph

The hash kernel is unchanged. What changes is the daily dataset build. Today item t of the 1 GiB dataset is mh_item(cache, t): the 16-word state starts as s[0..7] = K (day key) and s[8..15] = t * MUL[i] + RC[i], then 8 rounds of (8 mixers + one dependent 64-byte cache read XORed in), then 8 final mixers. Under sd1 a 64-byte STATE LEAF for item t is XORed into those 16 initial words before the first mixer: s[i] ^= leaf(t)[i]. In the real design leaf(t) is the t-th 64-byte leaf of a canonical serialisation of the chain's execution state at a certified checkpoint 20 minutes before the day boundary, zero padded where the state is shorter than the dataset. A miner therefore cannot build the day's dataset without the state, and a verifier that holds no dataset needs leaf(t) for every item it derives. The prototype measures what that costs (build time, device memory, verifier time) and buys (which leaves a hash touches, and what a light client would have to be shown).

Leaf stand-in (synthetic, prototype only)

leaf(t) = mh_chacha_block(x) (the pack's own ChaCha12 block from memhard.h) with x = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574, S[0..7], t, 0, 0x49676e65, 0x53746174) and S[i] = K[i] ^ 0x5a5a5a5a (stand-in state root). For the pack's day key S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102. The leaf array is 2^24 items x 64 B = 1 GiB for the pack's 1 GiB dataset (2^28 words = 2^24 items). Definition in sd.h.

Files

file what
kernel.cu, memhard.h, program.h, vectors.h the mx8-genesis pack, copied unmodified from proto-cuda/packs-ca2-mixer/mx8-genesis/
sd.h mh_leaf, mh_item_sd, mh_word_sd (host and device, plain C compatible)
kernel_sd.cu the pack's kernel.cu plus igneum_leaves, igneum_build_sd and their launch wrappers appended; igneum_hash byte for byte the pack's (checked with diff)
bench.cu the harness: `--mode control
verify_sd.c plain C: host cache fill, item bit-exact check from the GPU's items file, 32-lane register-major interpreter of the program against the GPU's dump, lane-0 load trace, per-unit rows
cpu_rows.c plain C: the igneum-build-1 rows B.1 and B.2 (OpenMP for the 32-thread row)
run_gpu.sh, run_cpu.sh the exact commands, as run
results/gpu/, results/cpu/ every log, the dumps and the items files, copied back from the boxes

Boxes

  • GPU box 2: NVIDIA GeForce RTX 4090, 24 GB (24083 MiB reported), 128 SMs, driver 595.91.07 (CUDA driver 13.2), nvcc 12.8.93, -arch=sm_89, power limit 450 W, Ubuntu 24.04.1. Host CPU AMD EPYC 7352 (48 threads). A CPU-only pool daemon shares the host (ports 4463/4480); the plain C rows there were pinned to core 2. Working directory /root/horizon-newpow/state-dataset.
  • CPU box igneum-build-1: AMD EPYC 9454P (48 cores, 96 threads), 128 GB, gcc 13.3.0, L3 256 MiB. Shared with other agents' builds: load average 19 to 32 during the readings (printed in each log). Rows pinned to core 4 under nice -n 19; the 32-thread row on cores 4 to 35. Working directory /srv/builds/horizon-newpow/state-dataset.

Commands

From the Mac (zsh; the GPU box has no rsync, tar over ssh instead):

cd /Users/joshm/Projects/igneum-wt-horizon/proto-newpow/state-dataset
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.cu *.h *.c *.sh | ssh -i ~/.ssh/igneum-fleet -p <box-2-port> root@<box-2-ip> 'cd /root/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_gpu.sh > run_gpu.log 2>&1'
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.c *.h *.sh    | ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 'cd /srv/builds/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_cpu.sh > run_cpu.log 2>&1'

On GPU box 2 (run_gpu.sh):

export PATH=/usr/local/cuda/bin:$PATH
nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
gcc -O2 -std=c11 -o verify_sd verify_sd.c
./bench --mode control --batches 10 --items-out items_control.txt --dump dump_control.txt 4 --power-seconds 20
./bench --mode sd1     --batches 10 --items-out items_sd1.txt     --dump dump_sd1.txt 4     --power-seconds 20
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 64
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 256
./bench --mode control --batches 10 --power-seconds 20        # second pass
./bench --mode sd1     --batches 10 --power-seconds 20        # second pass
taskset -c 2 ./verify_sd --mode control --items items_control.txt --dump dump_control.txt
taskset -c 2 ./verify_sd --mode sd1     --items items_sd1.txt     --dump dump_sd1.txt

On igneum-build-1 (run_cpu.sh):

gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
gcc -O2 -std=c11 -o verify_sd verify_sd.c
taskset -c 4    nice -n 19 ./cpu_rows --rows          # B.1 (i)(ii)(iii) and the one-core B.2 rows, twice
taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32     # B.2 leaf array on 32 threads, twice
taskset -c 4    nice -n 19 ./verify_sd --mode sd1     # the per-unit rows of verify_sd.c on this CPU

Timing method: cache fill, leaf array and dataset build are CUDA-event times of the second of two launches (as host.cu does). MH/s is CUDA-event time over 10 batches of 2^24 after a warm-up batch at base 0. Watts and SM MHz are the mean of nvidia-smi -l 1 samples taken after the first 10 s of a 20 s window in which the hash kernel runs back to back (10 samples used of 22). Fingerprint = FNV-1a 64 over the 2^24 little-endian u64 outputs at base nonce 0, the project's definition.

RESULTS

A. GPU, RTX 4090 (control vs sd1, two passes each)

row control sd1 note
cache fill, GPU, second pass (ms) 1.81, 1.85 1.85, 1.85 256 MiB, unchanged kernel
leaf array, GPU, second pass (ms) none 5.49, 5.49 2^24 ChaCha12 blocks, 1 GiB written at 195 GB/s
dataset build, second pass (ms) 30.55, 30.55 31.98, 31.98 +1.43 ms (+4.7%): one coalesced 64 B read per item
build + leaves (ms) 30.55 37.47 synthetic leaves; a real snapshot arrives from the node instead (see chunked rows)
the 3 Mac vectors, standalone and in batch PASS, PASS n/a (new values) sd1 base 0 lane 0 = b600edbed969becc, base 4096 = a533e89c78bb6b74, base 1000000 = beb4cb0c6bab8163
dataset head, word [MASK], 64 Mac samples PASS n/a (new values) sd1 dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf
64 random words vs host derivation PASS (mh_word) PASS (mh_word_sd)
2^24 fingerprint at base 0 7c28cfb06c5c65a9 (= the pack's, both passes) d5b0c16390cad0e8 (new, four runs)
hash rate, 10 batches of 2^24 (MH/s) 63.083, 63.083 63.088, 63.087 equal within 0.01%; the hash kernel is the same binary over a dataset of the same shape
hash rate in the 20 s power window (MH/s) 63.078, 63.079 63.083, 63.083
watts, mean after 10 s 207.6, 205.2 207.9, 206.5 run to run noise 2 W
SM MHz, mem MHz 2745, 10251 2745, 10251
MH/s per W 0.304, 0.307 0.303, 0.305
hash kernel registers 29 29 24 resident blocks/SM at 1 warp/block

Bit-exactness of the build (A.3): 1,024 random item indices (SplitMix64 seeded 0x5d1, t = z & (2^24 - 1)), the 16 GPU words of each item against the host derivation in plain C on the box's CPU with the host-filled cache (65536 mh_cache_segment calls, 380 ms in the bench, 573 to 587 ms pinned to core 2 in verify_sd):

check control sd1
items equal, in-process (bench.cu, host mh_item / mh_item_sd) 1024 of 1024 1024 of 1024
items equal, plain C (verify_sd.c from the items file) 1024 of 1024 1024 of 1024
items equal after the chunked rebuild (64 MiB chunks; 256 MiB chunks) n/a 1024 of 1024; 1024 of 1024
64 device leaves vs host mh_leaf (incl. t = 0 and 2^24 - 1) n/a PASS
host cache FNV-1a 64 48c4f5bf24166b2e = Mac 48c4f5bf24166b2e = Mac

GPU hash outputs vs the C interpreter (4 warps dumped, bases 0, 32, 64, 96; loads through mh_word / mh_word_sd):

control sd1
lanes equal 128 of 128 128 of 128
interpreter vs the Mac vector at base 0 32 of 32 n/a
interpreting time (4,096 item derivations per warp) 37 ms 42 ms

Device memory (A.4), cudaMemGetInfo, MiB used (context 395 included):

phase control sd1 resident leaves sd1 chunked 64 MiB sd1 chunked 256 MiB
during the build 1675 2699 1741 1933
while hashing (leaves freed) 1803 1803 1803 1803
chunked rebuild total, copies + builds, one stream (ms) 75.80 75.47

Whole-array host to device copy of the 1 GiB leaf array from pinned memory: 62 ms = 17.2 GB/s (this box's PCIe link under load; device to host 55 ms). So the chunked build is PCIe-bound: 62 ms of copy plus the 32 ms build, partly serialised on one stream, gives 76 ms. Two chunk buffers on two streams would hide most of the build under the copy (not done; the kernel already takes an item range [t0, t0 + n) with the chunk's leaves at leaves - 16 t0, so chunking needed no kernel change at all, only the loop in bench.cu).

B. CPU, igneum-build-1 (EPYC 9454P, one core pinned, nice 19), ms per unit of 4,096 items, 100 units

Two readings, because the box is shared: reading 1 at load 19, reading 2 at load 31.

row reading 1 reading 2 per item
(i) 4,096 random 64 B reads from a 2 GiB resident leaf array 0.108 ms (min 0.092, max 0.172) 0.163 ms 26 to 40 ns
(i) the same from an 8 GiB array (the year-12 size) 0.139 ms (min 0.127, max 0.183) 0.209 ms 34 to 51 ns
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array) 0.284 ms 0.285 ms 69 ns
(iii) 4,096 x mh_item on the host cache, naive, no interleaving 9.01 ms (min 8.91, max 9.60) 11.24 ms (min 9.38, max 13.51) 2.2 to 2.7 us
4,096 x (leaf + mh_item_sd), naive 9.31 ms 11.64 ms

The same rows on GPU box 2's EPYC 7352, core 2 (verify_sd.c): (ii) 0.369 to 0.374 ms, (iii) 10.14 to 10.21 ms, leaf + mh_item_sd 10.56 to 10.58 ms.

The project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive (iii) figure here is an upper bound. The sd1 increment is the row that matters: 0.11 to 0.21 ms per unit when the verifier holds the 2 to 8 GiB state in RAM, or 0.28 ms when it derives the stand-in leaf (in the real design it cannot derive a state leaf, it must hold the state or be shown it). Against the 2.06 ms interleaved verifier that is +5% to +14%; against the naive 9 to 11 ms it is +1% to +3%.

B.2, the daily snapshot build on CPU (igneum-build-1):

row time
1 GiB leaf array, 2^24 ChaCha12 blocks, one core 1.670 s, 1.677 s (100 ns per leaf)
the same on 32 threads (OpenMP), second pass (first pass 0.172 s with page faults) 0.073 s, 0.083 s
host cache fill, 256 MiB, one core 0.447 s, 0.452 s (0.380 s on the 4090 box's host, unpinned; 0.57 s pinned beside the pool daemon)

B.3, the state sample a light client would check (lane 0 of warp base 0, from the instrumented interpreter):

control sd1
distinct words touched by the 128 loads 128 128
distinct items touched 128 of 2^24 128 of 2^24
first 8 load words 8528400 43034175 255822847 136709537 ... 8528400 85793927 255822847 180392080 ...

Loads 1 and 3 share their address in both modes because their addresses are fixed by the nonce before any load's value feeds back; from load 2 on the sd1 trail diverges. Merkle arithmetic for a 2 GiB state (2^25 leaves of 64 B, depth 25, 32-byte hashes): 128 openings x 25 x 32 B = 102,400 B of siblings + 128 x 64 B = 8,192 B of leaves = 110,592 B (108 KiB) per lane, uncompressed; a 32-lane unit is 32 x 108 KiB = 3.4 MiB, before deduplicating shared upper levels (128 random paths in a depth-25 tree share only their top seven levels, so dedup saves about 20% per lane, approximate; across the 32 lanes of a unit the same seven levels are shared again). Every one of the 128 loads is a distinct item, so there is no saving from repeated items.

What the numbers mean, per tier (the consequences rule)

  • Hash rate, watts, MH/s per W: unchanged for every tier on every vendor, because the hash kernel is the pack's and the dataset has the same shape. Nothing to do.
  • Daily build: +1.4 ms on the 4090 (30.6 to 32.0 ms) when the leaves are resident, 76 ms when streamed from the host over PCIe in chunks. At the designed 2 GiB dataset the leaf array is 2 GiB and the streamed build is about 150 ms (approximate, scaling the 17 GB/s copy); on a PCIe 3 x8 slot in a rig, about 4x that (approximate). All far inside the daily window on every tier.
  • Device memory: with resident leaves the build peak at 2 GiB is 2 GiB + 256 MiB + 2 GiB + context, about 4.6 GiB (approximate), which fits an 8 GB card but not beside a prover. Streamed in 64 MiB chunks the peak is dataset + cache
    • 64 MiB + context, measured 1741 MiB at 1 GiB here, about 2.7 GiB at 2 GiB (approximate): no tier loses memory it has today. Recommendation: ship chunked only; never hold the leaf array on the device.
  • The real cost is delivery: every miner needs the 1 to 2 GiB state serialisation once a day. A miner beside its own node reads it from disk or loopback (seconds). A pool user without a node must get it from the pool (2 GiB per day per miner, or a shared download), and a light verifier either holds the state (0.11 to 0.21 ms per unit extra, measured) or is shown 3.4 MiB of Merkle openings per unit (arithmetic above), which is not a light client any more. The design decision this prototype leaves open is which of those two the protocol asks of a header verifier.
  • CPU snapshot build: 1.7 s on one core, 0.07 s on 32, for the synthetic leaves; the real serialisation is bounded by the node's state read, not by this.

What failed, what was cut

  • Nothing on the list was cut. verify_sd.c's lane-0 trace printed nothing on the first run (the trace counter was reset on every warp, so it read 0 after warp 3); fixed, the verifiers re-run, the fix is in the file.
  • The GPU box has no rsync; the sources went over with tar through ssh (COPYFILE_DISABLE=1 --no-xattrs, the Mac's tar otherwise writes Apple xattr headers that GNU tar warns about).
  • -std=c11 hides clock_gettime; both C files define _POSIX_C_SOURCE 200809L.
  • igneum-build-1 was under other agents' load (19 to 32) for every reading; both readings are given. The 4090 box's CPU rows ran beside the pool daemon, pinned to core 2. The GPU itself was idle apart from this bench.
  • Not measured: a two-stream chunked build (copy and build overlapped), AMD and Apple builds, the 2 GiB dataset itself (the pack is 1 GiB; the 2 GiB figures above are labelled approximate).