docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes, the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8 two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The lane's standing rule: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is least efficient per op. Chip rows in sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by placeholders in the READMEs and the run script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
19 lines
1.5 KiB
Text
19 lines
1.5 KiB
Text
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 512 mm8 per hash = 4096 path = PTX mma.sync.m8n8k16.u8
|
|
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
|
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
|
CUDA: driver 12.8, runtime 12.8
|
|
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
|
cache fill (GPU): 1.93 ms first, 1.89 ms second (256 MiB)
|
|
cache fill (host, one thread): 348.1 ms, GPU == host all words: PASS
|
|
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
|
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.54 ms second
|
|
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
|
pack vectors: not applicable at R = 512 (checked at R = 0 only)
|
|
warm-up batch: 16777216 hashes in 265.98 ms wall
|
|
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
|
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
|
|
sustain start epoch 1791316287.337
|
|
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
|
|
dump: 32 warps (1024 lanes) written to out/dump_512.txt
|
|
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
|
OVERALL: PASS
|