mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 32 mm8 per hash = 256 path = reference (shuffles + byte products) mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h) GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM CUDA: driver 12.8, runtime 12.8 igneum_hash_info: 151 registers/thread, 12 resident blocks/SM at 1 warp(s)/block = 12 resident warps/SM (25.0% of 48) cache fill (GPU): 1.80 ms first, 1.77 ms second (256 MiB) cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS) dataset build (GPU, 1024 MiB): 30.57 ms first, 30.52 ms second dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS) pack vectors: not applicable at R = 32 (checked at R = 0 only) warm-up batch: 16777216 hashes in 265.96 ms wall fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a OVERALL: PASS