docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes, the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8 two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The lane's standing rule: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is least efficient per op. Chip rows in sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by placeholders in the READMEs and the run script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
27 lines
2.3 KiB
Text
27 lines
2.3 KiB
Text
== Tue Oct 6 07:44:26 PM UTC 2026 on igneum-build-1, load 18.98 28.50 18.66
|
|
Model name: AMD EPYC 9454P 48-Core Processor
|
|
gcc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0
|
|
== build
|
|
== B.1 and one-core B.2 rows, core 4, nice 19
|
|
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
|
|
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.108 ms per unit (min 0.092, max 0.172; 26 ns per read; array write pass 1022 ms)
|
|
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.139 ms per unit (min 0.127, max 0.183; 34 ns per read; array write pass 3783 ms)
|
|
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.284 ms per unit (69 ns per leaf)
|
|
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.447 s
|
|
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms per unit (min 8.91, max 9.60)
|
|
4,096 x (leaf + mh_item_sd), naive: 9.31 ms per unit
|
|
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
|
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.670 s (100 ns per leaf)
|
|
== B.2 leaf array on 32 threads (cores 4-35), nice 19
|
|
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.172 s (first pass, includes page faults)
|
|
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.073 s
|
|
== verify_sd per-unit rows on this CPU (core 4), for comparison
|
|
verify_sd mode sd1
|
|
host cache fill: 480.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
|
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
|
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.285 ms
|
|
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms (min 8.92, max 9.44)
|
|
4,096 x (leaf + mh_item_sd), naive: 9.27 ms
|
|
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
|
RESULT verify mode=sd1 host_cache_ms=480.4 items_equal=-1 lanes_equal=-1/0
|
|
== done Tue Oct 6 07:44:39 PM UTC 2026, load 16.75 27.55 18.50
|