docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes, the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8 two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The lane's standing rule: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is least efficient per op. Chip rows in sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by placeholders in the READMEs and the run script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
85 lines
7.3 KiB
Text
85 lines
7.3 KiB
Text
== Tue Oct 6 19:43:12 UTC 2026 on a4cac49bb840
|
|
NVIDIA GeForce RTX 4090, 595.91.07, 3135 MHz, 24564 MiB
|
|
Build cuda_12.8.r12.8/compiler.35583870_0
|
|
== build
|
|
== control
|
|
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
|
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
|
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
|
device memory at start: 395 MiB used of 24083 MiB (context)
|
|
cache fill (GPU): 1.86 ms first, 1.81 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
|
cache fill (host, one thread): 380.5 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
|
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.1 M items/s
|
|
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
|
|
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
|
|
dataset self-test: PASS (64 random words vs host mh_word: PASS)
|
|
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
|
|
vector warp base 0: PASS (0 of 32 lanes differ)
|
|
vector warp base 4096: PASS (0 of 32 lanes differ)
|
|
vector warp base 1000000: PASS (0 of 32 lanes differ)
|
|
warm-up batch: 2^24 hashes in 266.02 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
|
|
vector warp base 0 in batch: PASS
|
|
vector warp base 4096 in batch: PASS
|
|
vector warp base 1000000 in batch: PASS
|
|
dump: 4 warps (bases 0, 32, ...) written to dump_control.txt
|
|
timed: 10 batches x 2^24 hashes, GPU 2659.56 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
|
|
power window: 76 batches in 20.2 s -> 63.078 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.6 W, SM 2745 MHz, mem 10251 MHz
|
|
-> 0.304 MH/s per W (window rate / mean W)
|
|
device memory while hashing: 1803 MiB used (leaves freed)
|
|
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.81 host_cache_ms=380.5 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.078 watts=207.6 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS
|
|
== sd1
|
|
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
|
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
|
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
|
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
|
device memory at start: 395 MiB used of 24083 MiB (context)
|
|
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
|
cache fill (host, one thread): 383.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
|
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.56 ms first, 5.49 ms second -> 195.4 GB/s written
|
|
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
|
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
|
|
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
|
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
|
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
|
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
|
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
|
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
|
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
|
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
|
dump: 4 warps (bases 0, 32, ...) written to dump_sd1.txt
|
|
timed: 10 batches x 2^24 hashes, GPU 2659.35 ms -> 63.088 MH/s (32.30 GB/s useful, loads x 4 B)
|
|
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.9 W, SM 2745 MHz, mem 10251 MHz
|
|
-> 0.303 MH/s per W (window rate / mean W)
|
|
device memory while hashing: 1803 MiB used (leaves freed)
|
|
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=383.9 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.088 mhs_window=63.083 watts=207.9 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
|
== verify (plain C, CPU core 2)
|
|
verify_sd mode control
|
|
host cache fill: 587.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
|
interpreter vs Mac vector base 0: 32 of 32 lanes equal PASS
|
|
item bit-exactness (plain C, mh_item on the host cache): 1024 of 1024 items equal PASS
|
|
warp base 0: 32 of 32 lanes equal (lane 0 gpu 19b56348bc85304d interp 19b56348bc85304d)
|
|
warp base 32: 32 of 32 lanes equal (lane 0 gpu bbea9a410fe7d5b7 interp bbea9a410fe7d5b7)
|
|
warp base 64: 32 of 32 lanes equal (lane 0 gpu 68b73c212b47e6f2 interp 68b73c212b47e6f2)
|
|
warp base 96: 32 of 32 lanes equal (lane 0 gpu 3a0fd3ada6b5a797 interp 3a0fd3ada6b5a797)
|
|
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (37 ms interpreting, 4096 derivations per warp)
|
|
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
|
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.369 ms
|
|
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 10.14 ms (min 9.98, max 10.31)
|
|
4,096 x (leaf + mh_item_sd), naive: 10.56 ms
|
|
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
|
RESULT verify mode=control host_cache_ms=587.4 items_equal=1024 lanes_equal=128/128
|
|
verify_sd mode sd1
|
|
host cache fill: 572.9 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
|
item bit-exactness (plain C, mh_item_sd with mh_leaf on the host cache): 1024 of 1024 items equal PASS
|
|
warp base 0: 32 of 32 lanes equal (lane 0 gpu b600edbed969becc interp b600edbed969becc)
|
|
warp base 32: 32 of 32 lanes equal (lane 0 gpu e451566071a2a0ba interp e451566071a2a0ba)
|
|
warp base 64: 32 of 32 lanes equal (lane 0 gpu 34548f5056f27d78 interp 34548f5056f27d78)
|
|
warp base 96: 32 of 32 lanes equal (lane 0 gpu c13d6ea3e05c5e46 interp c13d6ea3e05c5e46)
|
|
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (42 ms interpreting, 4096 derivations per warp)
|
|
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
|
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.374 ms
|
|
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 10.21 ms (min 10.09, max 12.24)
|
|
4,096 x (leaf + mh_item_sd), naive: 10.58 ms
|
|
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
|
RESULT verify mode=sd1 host_cache_ms=572.9 items_equal=1024 lanes_equal=128/128
|
|
== done Tue Oct 6 19:44:17 UTC 2026
|