igneum/proto-newpow/state-dataset/results/cpu/rows2.log
igneum-labs a664af6fc9 Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s
docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes,
the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as
proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8
two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the
execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's
CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The
lane's standing rule: a shadow lever only works through joules the honest card is forced to
spend, so shadow work goes where the GPU is least efficient per op. Chip rows in
sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by
placeholders in the READMEs and the run script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 20:18:39 +00:00

9 lines
1,012 B
Text

B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.163 ms per unit (min 0.144, max 0.201; 40 ns per read; array write pass 1252 ms)
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.209 ms per unit (min 0.203, max 0.229; 51 ns per read; array write pass 4325 ms)
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.285 ms per unit (70 ns per leaf)
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.452 s
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 11.24 ms per unit (min 9.38, max 13.51)
4,096 x (leaf + mh_item_sd), naive: 11.64 ms per unit
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.677 s (100 ns per leaf)