docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes, the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8 two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The lane's standing rule: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is least efficient per op. Chip rows in sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by placeholders in the READMEs and the run script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| out | ||
| bench.cu | ||
| gen_block.py | ||
| gen_ref_program.py | ||
| kernel.cu | ||
| kernel_mm8.cu | ||
| memhard.h | ||
| mm8_block.h | ||
| program.h | ||
| README.md | ||
| ref_program.inc | ||
| run.sh | ||
| summarise.py | ||
| vectors.h | ||
| verify_ref.c | ||
mma-shadow: prototype kernel for class "mx8+mm8xR" (Horizon lane 8, new proof of work)
Measured 6 October 2026 on GPU box 1 (RTX 4090 24 GB). Everything here lives in this directory; the pack files (kernel.cu, memhard.h, program.h, vectors.h) are verbatim copies of proto-cuda/packs-ca2-mixer/mx8-genesis.
The scheme
The shipped mx8 hash (8 iterations of 64 straight-line instructions over r0..r7, 16 dataset loads, 8 shuffles, the
fold) plus a block of R mm8 steps at the end of every iteration, after instruction 63 and before the next
iteration samples sel, so 8 x R mm8 steps per hash. Nothing else changes. One mm8 step k with drawn registers
(a_k, b_k, c_k, c2_k), a_k != b_k, c_k != c2_k: the 32 lanes' r[a] form A (8 x 16 u8, lane l holds
A[l >> 2][4 (l & 3) .. +3], byte 0 = lowest k), the 32 lanes' r[b] form B (16 x 8 u8, lane l holds
B[4 (l & 3) .. +3][l >> 2]), C = A x B exact in int32, and lane l does r[c] += C[l >> 2][2 (l & 3)] and
r[c2] += C[l >> 2][2 (l & 3) + 1], both modulo 2^32. That is the PTX mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32
fragment layout with a zero accumulator (d0 into r[c], d1 into r[c2]), the same arithmetic as the mm8 warp_ref in
proto-cuda/family-probe.cu. Both tile outputs are consumed per step (design correction received 6 Oct 2026 before
anything was measured, so the single-output bit form was never built). The draws come from a SplitMix64 stream
seeded with FNV-1a-64 of "igneum-mm8/igneum-genesis" (seed 0x79f1fc5b6ed6112e): a = below(8); b = below(7),
b += (b >= a); c = below(8); c2 = below(7), c2 += (c2 >= c). R is the compile-time macro IGNEUM_MM8_R; R = 0 is
the control and is bit-exact with the pack (3 vector warps and the 2^24 fingerprint 7c28cfb06c5c65a9).
Files
| File | What |
|---|---|
gen_block.py |
draws the 512-step table, writes mm8_block.h (packed uint16 table + X-macro step list) |
kernel_mm8.cu |
the pack's kernel.cu with r0..r7 as uint32_t r[8] (mechanical rewrite, same arithmetic) and the mm8 block; PTX path by default, -DIGNEUM_MM8_REF for the shuffle-and-byte-product reference path (12 shuffles + 32 byte products per lane per step). Cache fill, build and launch wrappers unchanged |
bench.cu |
harness derived from proto-cuda/host.cu (serve mode stripped): device info, igneum_hash_info registers and occupancy, GPU cache fill + host fill + FNV check, GPU dataset build + self-test, pack vectors at R = 0, 2^24 fingerprint at base nonce 0, 1 warm-up + 10 timed batches (CUDA events), --sustain S for the power meter, --dump file n |
gen_ref_program.py |
turns the 64 instruction lines of kernel.cu into the 32-lane C interpreter body ref_program.inc (nothing transcribed by hand) |
verify_ref.c |
plain C CPU reference: host cache (65536 segments), lazy mh_word loads, register-major interpreter, mm8 block in the spec layout, fold; compares a dump, then times the verifier per unit |
run.sh |
the ladder R in {0, 8, 32, 128, 512}: builds, power sampling, fingerprints, dumps, CPU check, summary |
summarise.py |
builds the RESULTS table from out/ |
Exact commands (on the box, under /root/horizon-newpow/mma-shadow)
export PATH=/usr/local/cuda/bin:$PATH
python3 gen_block.py mm8_block.h
python3 gen_ref_program.py kernel.cu
gcc -O2 -o verify_ref verify_ref.c
for R in 0 8 32 128 512; do
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -Xptxas -v -o bench_$R bench.cu kernel_mm8.cu
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -DIGNEUM_MM8_REF -o bench_${R}_ref bench.cu kernel_mm8.cu
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu,timestamp --format=csv,noheader -l 1 > out/power_$R.csv &
./bench_$R --batches 10 --sustain 25 --dump out/dump_$R.txt 32 > out/bench_$R.log
kill %1
./bench_${R}_ref --fingerprint-only --no-host-cache > out/bench_${R}_ref.log
taskset -c 2 ./verify_ref out/dump_$R.txt $R --time > out/verify_$R.log # add --self-test at R = 0
done
python3 summarise.py > out/results.md
./run.sh is exactly that sequence (plus the idle baseline and the PTX == ref fingerprint comparison). From the
Mac: rsync -az -e "ssh -i ~/.ssh/igneum-fleet -p <box-1-port>" proto-newpow/mma-shadow/ root@<box-1-ip>:/root/horizon-newpow/mma-shadow/
then touch * on the far side before building.
Card and driver
NVIDIA GeForce RTX 4090, 128 SMs, cc 8.9, 24 GB. Driver 570.172.08 (the brief said 595; nvidia-smi reports 570.172.08), CUDA 12.8 (nvcc V12.8.93), g++ 13.3.0, Ubuntu 24.04, 96 CPU threads. Idle baseline before the run: 15.0 to 15.3 W at 210 MHz SM, 45 C, 0 % utilisation. The host carried a CPU load average of about 13 from other tenants during the run (the GPU itself was idle and ours alone); that is why the verifier timings were pinned to one core.
Method notes
- Timing: one warm-up batch (also the fingerprint batch) then 10 timed batches of 2^24 hashes between CUDA events, 1 warp per block (the pack bench default, 24 resident warps per SM at 29 registers). Power: nvidia-smi at 1 Hz during a 25 s sustained phase after the timed batches; the mean takes samples from 10 s after the sustain start to its end (the timestamps are in the csv). Microjoules per hash = mean watts / (MH/s x 10^6) x 10^6.
- Fingerprint = FNV-1a 64 over the 2^24 little-endian u64 outputs at base nonce 0. The PTX build and the reference build must agree (the reference is the plain-integer byte-product form of the same fragment layout).
- CPU == GPU: 32 warps at SplitMix64(0x1234) 32-aligned bases, 1,024 lanes, recomputed by
verify_ref. - The verifier here is a naive register-major interpreter with no interleaving and a lazy
mh_wordper load (72 mixer applications and 8 cache reads per word, 4,096 words per unit). It is slower than the project's Rust verifier (2.06 ms per unit on an M5 Max core) on this box's core, so the number that matters is the mm8 block delta (ms per unit at R minus ms per unit at R = 0), not the total.
RESULTS (6 Oct 2026, RTX 4090, driver 570.172.08, CUDA 12.8, sm_89, 1 warp/block, batch 2^24 x 10)
| R | mm8 per hash (8R) | MH/s (GPU time) | ratio to R = 0 | watts mean (sustain, after first 10 s) | SM MHz | uJ per hash | max C | fingerprint (2^24 at base 0) | PTX == ref | CPU == GPU lanes | verifier ms per unit on the box core: R = 0, R, block delta | regs per thread (PTX build) | regs (ref build) | blocks per SM (cudaOccupancy) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 (matches the pack) | yes | 1024 of 1024 | 10.18, 10.08, -0.10 (noise) | 29 | 29 | 24 (24 warps/SM, 50 %) |
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.93, 9.97, +0.05 | 30 | 77 | 24 |
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.08, 10.24, +0.17 | 29 | 151 | 24 |
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.18, 11.33, +1.14 | 32 | 175 | 24 |
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.89, 14.28, +4.39 | 36 | 213 | 24 |
Sustained MH/s over the 25 s power window: 63.075, 63.074, 63.072, 63.034, 63.073 for R = 0, 8, 32, 128, 512 (wall,
including a sync per batch). Wall and GPU-event rates agree to 0.01 MH/s at every R. Idle baseline 15.0 to 15.3 W.
Registers per thread at R = 0, 8, 32, 128, 512: 29, 30, 29, 32, 36 (PTX path), no spills, no stack, at every R.
R = 0 self-checks: cache FNV 48c4f5bf24166b2e PASS (GPU == host fill word for word), dataset head, [MASK], 64 random
points and 64 Mac samples PASS, the pack's 3 vector warps PASS standalone and in batch, fingerprint 7c28cfb06c5c65a9.
The C interpreter also reproduces the 3 pack vectors at R = 0 (--self-test).
What the numbers say
- On the RTX 4090 the mm8 block is free in hash rate up to R = 512: 4,096 tensor instructions per hash leave the rate at 63.08 MH/s, identical to the control to the third decimal. The kernel is latency-bound on the 16 dependent random loads per iteration; the tensor work fills stalls that were already there. At R = 512 the card issues about 2.6 x 10^11 mma.m8n8k16 per second, roughly 2.7 x 10^14 u8 multiply-adds per second (1,024 per instruction), which is in the region of 40 % of the card's dense int8 tensor peak (approximate, from the published TOPS figure). So the next doublings would start to cost hash rate; R = 512 is near the top of the free band, not in the middle of it.
- Power is the only GPU cost that moves: 201 W to 216 W (+7.3 %) and 3.19 to 3.42 uJ per hash (+7.2 %) from R = 0 to R = 512. SM clock stayed pinned at 2670 MHz at every R, temperature peaked at 64 C, no throttling seen.
- Correctness chain holds at every rung: the PTX fragment read and the plain-integer reference path agree on all
2^24 lanes of the fingerprint batch at every R, and the CPU interpreter matches the GPU on all 1,024 dumped lanes.
The probe's fragment layout (family-probe.cu
warp_ref, mm8) was used as written and needed no correction. - Verifier cost of the block on the box's core: about 1.1 us per mm8 step per unit (32 lanes x 32 byte products, naive loops), so +1.14 ms per unit at R = 128 and +4.39 ms per unit at R = 512. Against the project's Rust verifier at 2.06 ms per unit, R = 512 would roughly triple verification time unless the block is vectorised (the 8 x 8 x 16 tile is 1,024 MACs, a few hundred ns with SIMD, approximate); R = 32 adds 0.17 ms per unit (+8 % of 2.06 ms) and R = 128 adds 1.14 ms (+55 %). The totals in the table (about 10 ms per unit) are this naive interpreter's cost with a lazy mh_word per load and are not comparable to the Rust verifier; the delta column is the number to use.
- The reference-path build is only a correctness oracle: fully unrolled it reaches 213 registers at R = 512 (8 blocks per SM) and was never timed.
Raw outputs
out/ holds everything the box produced: run.log (the whole run), bench_R.log and bench_R_ref.log,
power_R.csv (1 Hz: W, SM MHz, C, util, timestamp), power_idle.csv, ptxas_R.txt and ptxas_R_ref.txt,
dump_R.txt (32 warps, 1,024 lanes), verify_R.log, results.md (summarise.py).
What was cut
Nothing from the brief. The host 1 GiB dataset was never built on the CPU (lazy mh_word, as allowed). The 32-lane
dump bases are 32-bit, so base + 31 cannot overflow (base = low32(next()) & ~31).