docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes, the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8 two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The lane's standing rule: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is least efficient per op. Chip rows in sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by placeholders in the READMEs and the run script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
24 lines
2.1 KiB
Text
24 lines
2.1 KiB
Text
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
|
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
|
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
|
device memory at start: 395 MiB used of 24083 MiB (context)
|
|
cache fill (GPU): 1.86 ms first, 1.81 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
|
cache fill (host, one thread): 380.5 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
|
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.1 M items/s
|
|
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
|
|
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
|
|
dataset self-test: PASS (64 random words vs host mh_word: PASS)
|
|
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
|
|
vector warp base 0: PASS (0 of 32 lanes differ)
|
|
vector warp base 4096: PASS (0 of 32 lanes differ)
|
|
vector warp base 1000000: PASS (0 of 32 lanes differ)
|
|
warm-up batch: 2^24 hashes in 266.02 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
|
|
vector warp base 0 in batch: PASS
|
|
vector warp base 4096 in batch: PASS
|
|
vector warp base 1000000 in batch: PASS
|
|
dump: 4 warps (bases 0, 32, ...) written to dump_control.txt
|
|
timed: 10 batches x 2^24 hashes, GPU 2659.56 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
|
|
power window: 76 batches in 20.2 s -> 63.078 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.6 W, SM 2745 MHz, mem 10251 MHz
|
|
-> 0.304 MH/s per W (window rate / mean W)
|
|
device memory while hashing: 1803 MiB used (leaves freed)
|
|
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.81 host_cache_ms=380.5 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.078 watts=207.6 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS
|