Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s

docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes,
the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as
proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8
two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the
execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's
CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The
lane's standing rule: a shadow lever only works through joules the honest card is forced to
spend, so shadow work goes where the GPU is least efficient per op. Chip rows in
sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by
placeholders in the READMEs and the run script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-06 21:18:39 +01:00
parent 4617a01e60
commit cbcff478f8
87 changed files with 13059 additions and 13 deletions

View file

@ -10,9 +10,10 @@ Josh's mandate, verbatim: "if we create a new way of hashing or a new way of pro
|---|---|
| 19:05 | Lane started. Read: preamble, CLAUDE.md, the two personas, spec 01 (whole), 04, 07, chip-model-v3 (whole), asic-resistance-history (sections 0 to 3 and 4.3, 5), latency-shadow-2026-10-06 (whole), counter-asic-3-status sections 1 to 5, counter-asic-3-node section 6 (the P2 signalling rule), int8-matrix-family sections 1 to 3, scratch-soundness verdict, proving-methods (whole), fud-ledger M1, M7, M16, M22, M28, P2, F13, proto-cuda host.cu and the mx8-genesis pack (kernel.cu, memhard.h, program.h, vectors.h), proto-cuda/emu, family-probe.cu, the fleet's prover-tiers-real-cards.md, bench-log line 2582 (rental cost) |
| 19:25 | Fleet agent asked for two boxes; answered at 19:29: two quiet RTX 4090s (RunPod, nvcc 12.8 at /usr/local/cuda/bin, directory /root/horizon-newpow, until 22:30Z). No quiet Ampere card exists tonight; a loaded 3090 is offered. Main's note: cost rows use bench-log 2582 (USD 0.0117 per MH/s-hour) |
| 19:35 | File skeleton written. Two prototype sub-agents launched (budget two at once): `mma-shadow` on box 1 (47.47.180.77), `state-dataset` on box 2 (213.173.98.36) plus CPU rows on igneum-build-1. Designs being written in this file meanwhile |
| 19:35 | File skeleton written. Two prototype sub-agents launched (budget two at once): `mma-shadow` on box 1 (<box-1-ip>), `state-dataset` on box 2 (<box-2-ip>) plus CPU rows on igneum-build-1. Designs being written in this file meanwhile |
| 19:41 | Section 3 complete: the three designs, the one-table comparison, the migration path. Scheme A's verdict is already visible in its own numbers (A1 dead on 2.9 MB of openings per block, A2 dead on sampleability and a 32 to 40 ms proof verify; A0 is scheme C with the trace as state). Prototypes running: `mma-shadow` (box 1) and `state-dataset` (box 2 and igneum-build-1). mm8 two-output correction sent to the prototype |
| (next) | Section 4 reviews (cryptographer, consensus engineer, two per scheme); the pick; section 5 measured rows as they land; section 6 verdicts; section 7 ranked next steps |
| 19:50 | Section 4 (reviews, the pick: B and C), section 7 (ranked next steps) and section 8 (open questions) written. Scheme C prototype complete and measured on box 2 and igneum-build-1: hash rate and watts unchanged (63.08 vs 63.09 MH/s, 207 W), build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact 1,024 of 1,024 items and 128 of 128 lanes; rows in 5.2. Scheme B ladder running on box 1 (R = 0, 8, 32, 128 measured, 512 in progress) |
| 20:20 | Scheme B ladder complete on box 1 (R = 0 to 512, rate flat at 63.08 MH/s, 201 to 216 W, every fingerprint PTX = reference, 1,024 of 1,024 lanes at every R, verifier delta 0.05 to 4.39 ms). Sections 5.1, 5.3, 6 and 9 written. Box 2 released 19:53Z; box 1 released on the mma agent's report. File complete |
## 1. What was read and the facts this lane stands on
@ -133,8 +134,8 @@ So the fleet splits by generation: Turing-and-later NVIDIA and RDNA 3-and-later
| Scheme | What it is | Chip edge per joule vs the 5090 (f = 1 chip, chip-model-v3 method) | Verifier ms per unit (model, then measured in section 5) | Vendor bit-exactness | Ships as class v5? |
|---|---|---|---|---|---|
| A, mining is proving | A1 committed codeword, A2 proving steps: dead on bytes and on sampleability; A0 trace-as-dataset survives and is C with the trace as state | A0 unchanged (5.1x GDDR7); A1, A2 not priced | A0: 2.06 + the leaf-read row; A1: 2.9 MB of openings; A2: 32 to 40 ms | A0 adds the zkVM executor to consensus | NEVER as mining = proving; A0 folds into C |
| B, tensor-shaped shadow | class v3 plus `8 R` int8 8x8x16 tile steps per hash in the PTX fragment layout, two outputs per lane | memory + `8 R x 1,024 x e_mac x k_mma`; `k_mma` near 1 with a floor near 0.5 (approximate); rows from the measured `e_mac` in 5.3 | 2.06 + about 0.9 microseconds per step per unit scalar (R = 128: +0.9 ms; R = 512: +3.7 ms); VNNI 4x to 16x less | native on sm_75+ and RDNA 3+ (AMD layout unverified); emulated on Pascal, RDNA 2, Apple | prototype further; a class v5 candidate after the AMD gate |
| C, stored state | the daily dataset derives from the execution state snapshot; kernel unchanged; a state sample per block | unchanged (5.1x GDDR7, 7.5x to 9.2x HBM3); the f = 0 chip disappears | 2.06 + 4,096 leaf reads (section 5.2) | nothing vendor-specific; the serialisation is the consensus risk | prototype further; a class v5 candidate on its own or beside B |
| B, tensor-shaped shadow | class v3 plus `8 R` int8 8x8x16 tile steps per hash in the PTX fragment layout, two outputs per lane | memory + `8 R x 1,024 x e_mac x k_mma`; `k_mma` near 1 with a floor near 0.5 (approximate); rows from the measured `e_mac` in 5.3 | 2.06 + about 0.9 microseconds per step per unit scalar (R = 128: +0.9 ms; R = 512: +3.7 ms); VNNI 4x to 16x less | native on sm_75+ and RDNA 3+ (AMD layout unverified); emulated on Pascal, RDNA 2, Apple | before measurement: prototype further. After section 5: NEVER as class content (the block costs the honest card 0.056 to 0.70 pJ per multiply-add, so it forces no joules on a chip); kept as reserve R8 evidence |
| C, stored state | the daily dataset derives from the execution state snapshot; kernel unchanged; a state sample per block | unchanged (5.1x GDDR7, 7.5x to 9.2x HBM3); the f = 0 chip disappears | 2.06 + 4,096 leaf reads (section 5.2) | nothing vendor-specific; the serialisation is the consensus risk | before measurement: prototype further. After section 5: SHIP as the class v5 candidate |
### 3.5 Migration through the class system (common to B and C)
@ -165,7 +166,7 @@ Written by this lane in the persona files' voices (`.claude/agents/cryptographer
### 4.4 The pick
B and C are the two to prototype: both are class objects on the shipped hash, both leave the dataset's memory bound untouched, and they compose (B is kernel text, C is the daily build). A is not prototyped: A1 and A2 fail on bytes and on sampleability before any kernel, and A0 is C with a worse data source. The order of merit at this point, before measurement: C first (no vendor risk, unchanged hash rate by construction, a real new property per block, the pool caveat stated), B second (a real lever against the f = 1 chip with a bounded `k`, a vendor split and an unverified AMD layout). Section 6 revisits the order on the measured rows.
B and C are the two to prototype: both are class objects on the shipped hash, both leave the dataset's memory bound untouched, and they compose (B is kernel text, C is the daily build). A is not prototyped: A1 and A2 fail on bytes and on sampleability before any kernel, and A0 is C with a worse data source. The order of merit at this point, before measurement: C first (no vendor risk, unchanged hash rate by construction, a real new property per block, the pool caveat stated), B second (a real lever against the f = 1 chip with a bounded `k`, a vendor split and an unverified AMD layout). Section 6 revisits the order on the measured rows: C holds, B retires into the reserve.
## 5. The prototypes and the measured rows
@ -173,31 +174,115 @@ Both prototypes live under `proto-newpow/` with a README carrying the exact comm
### 5.1 `mma-shadow` (scheme B), box 1
ROWS_B
Prototype: `proto-newpow/mma-shadow/` (README with every command, `kernel_mm8.cu` with the PTX path and the `IGNEUM_MM8_REF` reference path, `bench.cu`, `verify_ref.c`, `gen_block.py`, `gen_ref_program.py`, `run.sh`, `out/` with every log and csv). Box 1: RTX 4090 24 GB (128 SMs), driver 570.172 (the box reports 570, not the 595 of box 2), nvcc 12.8.93, `-arch=sm_89`, 1 warp per block, 10 timed batches of 2^24 after a warm-up, power from `nvidia-smi` at 1 Hz over a 25-s sustained phase (mean after its first 10 s), idle 15.0 to 15.3 W. The two-output tile form of section 3.2 was built (the single-output form never was). Run 19:36 to 19:52 UTC.
| R (steps per iteration) | mm8 per hash (8 R) | MH/s (GPU time) | Ratio to R = 0 | Watts, mean | Block watts over R = 0 | SM MHz | Microjoules per hash | Marginal pJ per multiply-add | Fingerprint (2^24 at base 0) | PTX = reference path | CPU = GPU (1,024 lanes) | Verifier block delta, ms per unit (box core, scalar C) | Registers per thread |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 (control) | 0 | 63.08 | 1.000 | 201.2 | 0 | 2,670 | 3.19 | | 7c28cfb06c5c65a9 (the pack's; the 3 pack vectors PASS standalone and in batch) | yes | 1,024 of 1,024 | 0 (the R = 0 run read -0.10, noise) | 29 |
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2.9 | 2,670 | 3.24 | 0.70 | 06fc2593bfb94b4f | yes | 1,024 of 1,024 | +0.05 | 30 |
| 32 | 256 | 63.08 | 1.000 | 207.7 | 6.5 | 2,670 | 3.29 | 0.39 | 26e83a65f519c865 | yes | 1,024 of 1,024 | +0.17 | 29 |
| 128 | 1,024 | 63.08 | 1.000 | 212.7 | 11.5 | 2,670 | 3.37 | 0.17 | 42223c2113188335 | yes | 1,024 of 1,024 | +1.14 | 32 |
| 512 | 4,096 | 63.08 | 1.000 | 215.9 | 14.7 | 2,670 | 3.42 | 0.056 | 02b7002d747f3711 | yes | 1,024 of 1,024 | +4.39 | 36 |
Sustained rates over the power window 63.07 to 63.03 MH/s at every R; wall and event rates agree to 0.01 MH/s; no spills, no stack, temperature 58 to 64 C, no throttling. The verifier totals in the box's logs (about 10 ms per unit) are a naive scalar interpreter with a lazy `mh_word` per load on a shared Zen 4 core at load 13 and are not comparable to the project's Rust verifier (2.06 ms per unit on an M5 Max core, interleaved); the block delta is the number: about 1.1 microseconds per `mm8` step per unit, 32 lanes x 32 byte products in plain loops. (At R = 512 the card issues about 258 G tiles per second, 2.6 x 10^14 multiply-adds per second, in the region of 40 percent of the 4090's dense int8 tensor peak, approximate from the published TOPS figure, so R = 512 is near the top of the free band, not its middle.)
What the rows say:
1. The tensor block is free in hash rate to R = 512 on the 4090: 4,096 tile instructions per hash leave the rate at 63.08 MH/s to the third decimal. The kernel is latency-bound on its 128 dependent loads and the tensor work fills stalls that were already there, as the ALU shadow did on the 5090 to 150,800 ops (`latency-shadow-2026-10-06.md` 5).
2. The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised), 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op, `latency-shadow-2026-10-06.md` 5, item 4); the tensor block at its free-band ceiling costs a third of that.
3. That is the finding, and it is negative for the scheme's purpose (section 5.3): a shadow lever moves the chip's edge only by the joules it makes the HONEST card spend on work the chip cannot do more cheaply. The tensor path is so efficient on the GPU that the block adds 0.23 microjoules at most, so at `k = 1` the chip's edge falls from 6.9x to 4.9x on GDDR7 against this 4090, where the ALU shadow took the 5090 from 5.6x to 2.1x, and the tensor block costs the verifier 26x more per unit of chip-forcing energy (4.39 ms scalar per 0.23 microjoules against 0.17 ms per 0.6 microjoules). The property the design hoped for (`k_mma` near 1 because the GPU's tensor engine is near the floor) is real and is exactly why the lever is weak: there are no joules in it to force.
4. The correctness chain holds at every rung: the PTX fragment read and the plain-integer reference agree on all 2^24 lanes at every R, and the CPU interpreter matches the GPU on 1,024 lanes at every R; the probe's fragment layout (`family-probe.cu` mm8 `warp_ref`) was used as written and needed no correction. Registers 29 to 36, occupancy unchanged. This is the first class-shaped evidence that an `mm8` family is cheap and exact for the honest NVIDIA card, which is what the reserve entry R8 needs; it is not evidence for a class v5.
Consequences per tier (B at R = 128, the largest R whose scalar verifier delta stays near 1 ms):
| Tier | Meaning | What to do |
|---|---|---|
| NVIDIA Turing and later (RTX 20 to 50 series), any memory size | rate unchanged, +11.5 W on a 4090 (+5.7 percent), per-pound unchanged, per-watt 0.95x | nothing; the block would not be adopted on this evidence |
| NVIDIA Pascal (GTX 10 series) | emulation at about 16 steps per `mm8` (approximate): 1,024 per hash is about 16,000 counted steps, inside the latency shadow of a 1080-class card by the counted-ops rule, approximate; unmeasured | the emulation path exists in the kernel text (`IGNEUM_MM8_REF`) and is bit-exact; a GTX 1080 row is owed if B ever proceeds |
| AMD RDNA 3 and 4 | the WMMA layout is unverified (status item 6); emulation otherwise at about 12 to 16 steps per `mm8` | the gate of section 7 rank 2 stays open; not worth running unless B proceeds |
| Apple | emulation at about 10 steps per `mm8` (approximate), 10,000 counted steps at R = 128 inside the M5 Max's 130,000 ceiling | owed; not worth running unless B proceeds |
| Rig, pool user | nothing changes in shares; a rig pays the block's watts per card | nothing |
| Node operator (verifier) | +1.14 ms per unit scalar at R = 128 (+55 percent of 2.06 ms), +4.39 ms at R = 512; a SIMD byte-dot path would cut it 4x to 16x (approximate) | the verifier cost per joule forced is the reason the scheme loses to the ALU shadow |
| A chip | must carry a tensor array per 32 lanes in flight; at the measured block energy it pays at most 0.23 microjoules per hash more than today at `k = 1`, 0.07 at `k = 0.3` | the chip's edge barely moves (5.3) |
### 5.2 `state-dataset` (scheme C), box 2 and igneum-build-1
ROWS_C
Prototype: `proto-newpow/state-dataset/` (README with every command and log; `kernel_sd.cu` = the pack's kernel plus `igneum_leaves` and `igneum_build_sd`, the hash kernel byte for byte the pack's; `verify_sd.c` the 32-lane C interpreter and the item check; `cpu_rows.c` the igneum-build-1 rows). Box 2: RTX 4090 24 GB, driver 595.91, nvcc 12.8, host EPYC 7352; igneum-build-1 EPYC 9454P at load 19 to 32 (shared), one core pinned, nice 19. Run 19:34 to 19:50 UTC. The leaf stand-in: one ChaCha12 block of (sigma, S, t, 0, tag) with `S = K XOR 0x5a5a5a5a` (the README defines it).
| Row | Control (mx8-genesis) | sd1 | Reading |
|---|---|---|---|
| Hash rate, 10 batches of 2^24, two passes | 63.083, 63.083 MH/s | 63.088, 63.087 MH/s | equal within 0.01 percent: the kernel is the same binary over a dataset of the same shape |
| Watts, mean after 10 s of a 20-s window; SM MHz | 207.6, 205.2 W; 2,745 | 207.9, 206.5 W; 2,745 | equal within the 2 W run-to-run noise; 3.27 microjoules per hash on this 4090 (the 5090 is 2.34 to 2.65) |
| MH/s per W | 0.304, 0.307 | 0.303, 0.305 | unchanged |
| Cache fill, GPU | 1.81, 1.85 ms | 1.85, 1.85 ms | unchanged |
| Leaf array (2^24 ChaCha12 blocks, 1 GiB) on the GPU | none | 5.49 ms | the synthetic stand-in; a real snapshot comes from the node |
| Dataset build, second pass | 30.55 ms | 31.98 ms | +1.43 ms (+4.7 percent): one coalesced 64 B read per item |
| Chunked build (leaves streamed from pinned host memory in 64 MiB or 256 MiB chunks, one stream) | | 75.8 ms, 75.5 ms | PCIe-bound: 62 ms of copy (17.2 GB/s on this box) plus the build, partly serialised; two streams would hide most of the build (not done) |
| Device memory during the build (context 395 MiB included) | 1,675 MiB | 2,699 MiB resident leaves; 1,741 MiB chunked at 64 MiB; 1,933 at 256 MiB | while hashing 1,803 MiB in every mode (leaves freed) |
| 2^24 fingerprint at base 0 | 7c28cfb06c5c65a9 (the pack's, both passes) | d5b0c16390cad0e8 (four runs) | |
| Pack vectors (3 warps, standalone and in batch); dataset head, word [MASK], 64 Mac samples | PASS | n/a (new values; sd1 base 0 lane 0 = b600edbed969becc) | the harness is the pack's |
| Build bit-exact: 1,024 random items, GPU words against the plain C host derivation | 1,024 of 1,024 | 1,024 of 1,024 (also after both chunked rebuilds); 64 device leaves = host leaves | |
| Hash bit-exact: 4 dumped warps (128 lanes) against the C interpreter with lazy `mh_word` / `mh_word_sd` | 128 of 128 (and 32 of 32 against the Mac vector at base 0) | 128 of 128 | |
| Host cache FNV-1a 64 | 48c4f5bf24166b2e = the Mac's | same | |
The CPU rows (igneum-build-1, ms per unit of 4,096 items, 100 units; two readings at load 19 and 31):
| Row | Reading 1 | Reading 2 | Per item |
|---|---|---|---|
| (i) 4,096 random 64 B reads from a 2 GiB resident leaf array | 0.108 ms (0.092 to 0.172) | 0.163 ms | 26 to 40 ns |
| (i) the same from an 8 GiB array (the year-12 size) | 0.139 ms | 0.209 ms | 34 to 51 ns |
| (ii) 4,096 leaves derived on the fly (one ChaCha12 block each) | 0.284 ms | 0.285 ms | 69 ns |
| (iii) 4,096 x `mh_item` on the host cache, naive, no interleaving | 9.01 ms | 11.24 ms | 2.2 to 2.7 microseconds |
| 4,096 x (leaf read + `mh_item_sd`), naive | 9.31 ms | 11.64 ms | |
| Daily leaf array on CPU (2^24 blocks): one core / 32 threads | 1.67 s / 0.073 s | 1.68 s / 0.083 s | 100 ns per leaf |
| Host cache fill (256 MiB), one core | 0.447 s | 0.452 s | |
| Light-client sample: distinct items touched by lane 0's 128 loads | 128 of 2^24 | | 128 openings x 25 levels x 32 B + 128 x 64 B = 108 KiB per lane; 3.4 MiB per 32-lane unit before dedup (arithmetic) |
Reading, against the gate: the project's Rust verifier does the 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items (the naive C figure here, 9 to 11 ms, is an upper bound and is not the production verifier); sd1 adds 0.11 to 0.21 ms per unit when the verifier holds the state in RAM (+5 to +10 percent of 2.06 ms; about +0.3 to +0.5 ms on a 2019-class core by the 2.5x rule), so the class v3 verifier plus sd1 sits at about 2.2 ms on the M5 Max core and about 5.7 ms on the 2019-class core, inside the 10 ms gate with the x8 margin intact. The snapshot's own cost is the state read on the node, not the leaf hashing (1.7 s on one core for 2^24 leaves).
Consequences per tier (sd1):
| Tier | Meaning | What to do |
|---|---|---|
| Home miner, one 8 GB card | hash rate, watts and MH/s per W unchanged (measured equal); the build must stream the leaves (resident leaves at the designed 2 GiB dataset would peak at about 4.6 GiB, approximate; chunked 64 MiB about 2.7 GiB, approximate, measured 1,741 MiB at 1 GiB), so the 8 GB mine-and-prove profiles of `prover-tiers-real-cards.md` keep their memory | ship chunked only; the kernel already takes an item range, so chunking needed no kernel change |
| 12, 16, 24, 32 GB cards | unchanged in rate and watts; the daily build +1.4 ms resident or about 76 ms streamed (about 150 ms at 2 GiB, approximate) | nothing |
| AMD and Apple | the hash kernel is untouched, so the measured class v3 rates stand (18.9 MH/s on the 9070 XT, 27.1 on the M5 Max); the build kernels gain one 64 B read per item in OpenCL and Metal | the two build kernels in the emitter's other dialects |
| Rig | per card as above; one 2 GiB snapshot per day per rig, shared by its cards | the node beside the rig serves it over loopback |
| Pool user | a 2 GiB download per day from the pool, or the pool ships the built dataset; today a pool user needs nothing but the day key | the pool protocol (spec 09) gains a daily snapshot or dataset fetch |
| Solo miner | must run a node with execution (it already does for templates) and hold the state | nothing new beyond RAM |
| Node operator | +2 GiB RAM for the snapshot (+4 GiB at year 4), a daily serialisation pass, the verifier +0.11 to 0.21 ms per unit | the node RAM budget stays inside a 16 GB home node at launch sizes |
| Light client, holder | a verifiable sample of 128 state leaves per lane per block exists; checking it costs 108 KiB per lane fetched on demand, so it is an availability spot check, not a light verification of PoW | rank 6 of section 7 |
| Prover, rollup customer | nothing | |
### 5.3 The chip rows on the measured numbers
CHIP_ROWS
Method: `sim/horizon/new-pow/chip_rows.py` (the chip-model-v3 section 5.4 f = 1 chip plus the block's measured energy times `k`; the card figures are 5.1's; every chip figure is arithmetic on `chip-model-v3.md`'s cited and approximate memory figures). The denominator here is the 4090 measured tonight (3.19 microjoules per hash at the control, a card whose memory system is half the 5090's), not the 5090, so the R = 0 column reads 6.9x where `chip-model-v3.md` reads 5.1x for the 5090; the movement with R is the row that matters.
| Scheme and setting | Honest card microjoules per hash | Chip, GDDR7, k = 1 / 0.5 / 0.3 (microjoules) | Edge over the card, GDDR7 | Chip, one HBM3 stack, k = 1 / 0.5 / 0.3 | Edge, HBM3 |
|---|---|---|---|---|---|
| Class v3 today (R = 0, the 4090 control) | 3.19 | 0.466 | 6.9x | 0.321 | 9.9x |
| B at R = 128 (1,024 tiles per hash, block 0.18 microjoules) | 3.37 | 0.646 / 0.556 / 0.520 | 5.2x / 6.1x / 6.5x | 0.501 / 0.411 / 0.375 | 6.7x / 8.2x / 9.0x |
| B at R = 512 (4,096 tiles per hash, block 0.23 microjoules, the free band's top) | 3.42 | 0.696 / 0.581 / 0.535 | 4.9x / 5.9x / 6.4x | 0.551 / 0.436 / 0.390 | 6.2x / 7.8x / 8.8x |
| For comparison, the ALU shadow at N = 100,000 on the 5090 (`latency-shadow-2026-10-06.md` 6, measured 3.27 microjoules at the 431 W cap) | 3.27 | 1.59 / 1.03 / 0.80 | 2.1x / 3.2x / 4.1x | 1.44 / 0.88 / 0.66 | 2.3x / 3.7x / 5.0x |
| C (sd1) at any size | unchanged (measured equal) | 0.466 | unchanged | 0.321 | unchanged; the f = 0 recompute chip must add 2 GiB of DRAM and becomes this chip |
Reading: B moves the f = 1 chip's edge by 1.4x to 2x at `k = 1` and by 1.1x at `k = 0.3`, at a verifier cost of 1.1 to 4.4 ms per unit (scalar); the ALU shadow moves it by 2.7x at `k = 1` and 1.4x at `k = 0.3` at 0.17 ms per unit. On every axis the measured tensor block is the weaker lever, for the reason stated in 5.1 item 3. C moves nothing and claims nothing against the chip; its merit is elsewhere.
## 6. Verdicts
| Scheme | Verdict | Why, in one line |
|---|---|---|
| A, mining is proving | NEVER (A1, A2); A0 folds into C | one proof per segment is not a distribution of puzzles; the bytes (2.9 MB of openings per block) or the verify (32 to 40 ms) kill every form that is not "hold the trace", and holding the trace is C with a worse data source |
| B, tensor-shaped shadow | VERDICT_B |
| C, stored state | VERDICT_C |
| B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joules |
| C, stored state | SHIP AS CLASS v5 CANDIDATE (through the spec items of 4.3 and the Devnet 2 gate): hash rate and watts unchanged by construction and measured equal, build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact on 1,024 items and 128 lanes; a new property per block (a random sample of state) and a new requirement per mining operation (hold the state); the open decision is what a header verifier is asked to hold |
**A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.
VERDICT_BC_PROSE
NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joulesC_PROSE
## 7. Ranked next steps
Hours are agent hours (Josh's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own.
Hours are agent hours (Josh's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own. Ranks 2, 3 and 4 were written as B's gates before the ladder landed; after section 5 they are WITHDRAWN (B is not carried forward as class content; the rows stay so the reasoning is visible) and the live order is 1, 5, 6, 7, 8.
| Rank | Proposal | Evidence | Model | Hours | Consequence per tier | Gate |
|---|---|---|---|---|---|---|
@ -236,4 +321,8 @@ One paragraph each.
## 9. Summary for the coordinator
SUMMARY
This lane wrote three candidate proofs of work for Igneum, reviewed each in two personas, prototyped the two that survived on real cards, and measured. (1) Scheme C, the dataset derived from the chain's execution state, is the one to carry forward as the class v5 candidate: measured on an RTX 4090 the hash rate and watts are unchanged to the third decimal (63.08 MH/s, 207 W, the kernel is byte for byte the shipped one), the daily build grows by 1.4 ms resident or about 76 ms streamed, the verifier by 0.11 to 0.21 ms per unit with the state in RAM, every GPU item and lane agrees with the plain C derivation (1,024 of 1,024, 128 of 128); it forces every mining operation to hold state, names 128 random state leaves per lane per block, and removes the f = 0 recompute chip as a category; its costs are a 2 GiB daily delivery for pool miners and three spec items (canonical serialisation, a pruning-proof witness, the pause rule). (2) Scheme B, a tensor-shaped shadow of int8 8x8x16 tiles, is bit-exact on NVIDIA at every rung (2^24 lanes PTX = reference, 1,024 lanes CPU = GPU, R = 8 to 512) and free in hash rate to 4,096 tiles per hash, and that is exactly why it fails as an anti-chip lever: it costs the honest card 0.056 to 0.70 pJ per multiply-add and at most 0.23 microjoules per hash, so the f = 1 chip's edge moves from 6.9x to 4.9x at k = 1 against the 4090 where the class v4 ALU shadow moved the 5090's from 5.6x to 2.1x, at 26x the verifier cost per joule forced; never as class content, kept as the evidence and the two-output correction for reserve family R8. (3) Scheme A, mining is proving, is never: one proof per segment is not a distribution of puzzles; the useful fraction is bounded by proving work per segment over network hashes per segment (8 percent at 1 GH/s, 0.08 percent at 100 GH/s on this chain's measured figures); every form that is not "hold the trace" dies on 2.9 MB of openings per block or a 32 to 40 ms proof verify, and "hold the trace" is scheme C with a worse data source.
1. Scheme C costs nothing in hash rate or watts (63.083 against 63.088 MH/s, 205 to 208 W on the 4090, measured) and 0.11 to 0.21 ms per unit of verifier time (measured on igneum-build-1), so class v5 can carry it; the model is `item_sd(t) = item(t) with s ^= leaf(t)` over the day's certified state root.
2. The tensor shadow moves the f = 1 chip's per-joule edge by at most 1.4x (6.9x to 4.9x at k = 1, R = 512) for 4.39 ms of scalar verifier per unit, against 2.7x for 0.17 ms from the ALU shadow; the model is `chip_rows.py` on the measured block energy of 0.23 microjoules per hash.
3. Mining cannot be proving on a lottery with a 10 ms verifier: the useful fraction is bounded by 5 GPU-seconds of proving per 8-second segment over the network's hashes in that segment, 8 percent at 1 GH/s and falling with the hash rate; the model is that ratio on `prover-tiers-real-cards.md` and bench-log 2582.

View file

@ -0,0 +1,125 @@
# mma-shadow: prototype kernel for class "mx8+mm8xR" (Horizon lane 8, new proof of work)
Measured 6 October 2026 on GPU box 1 (RTX 4090 24 GB). Everything here lives in this directory; the pack files
(kernel.cu, memhard.h, program.h, vectors.h) are verbatim copies of proto-cuda/packs-ca2-mixer/mx8-genesis.
## The scheme
The shipped mx8 hash (8 iterations of 64 straight-line instructions over r0..r7, 16 dataset loads, 8 shuffles, the
fold) plus a block of R `mm8` steps at the end of every iteration, after instruction 63 and before the next
iteration samples `sel`, so 8 x R mm8 steps per hash. Nothing else changes. One mm8 step k with drawn registers
(a_k, b_k, c_k, c2_k), a_k != b_k, c_k != c2_k: the 32 lanes' r[a] form A (8 x 16 u8, lane l holds
A[l >> 2][4 (l & 3) .. +3], byte 0 = lowest k), the 32 lanes' r[b] form B (16 x 8 u8, lane l holds
B[4 (l & 3) .. +3][l >> 2]), C = A x B exact in int32, and lane l does r[c] += C[l >> 2][2 (l & 3)] and
r[c2] += C[l >> 2][2 (l & 3) + 1], both modulo 2^32. That is the PTX `mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32`
fragment layout with a zero accumulator (d0 into r[c], d1 into r[c2]), the same arithmetic as the mm8 `warp_ref` in
proto-cuda/family-probe.cu. Both tile outputs are consumed per step (design correction received 6 Oct 2026 before
anything was measured, so the single-output `bit` form was never built). The draws come from a SplitMix64 stream
seeded with FNV-1a-64 of "igneum-mm8/igneum-genesis" (seed 0x79f1fc5b6ed6112e): a = below(8); b = below(7),
b += (b >= a); c = below(8); c2 = below(7), c2 += (c2 >= c). R is the compile-time macro `IGNEUM_MM8_R`; R = 0 is
the control and is bit-exact with the pack (3 vector warps and the 2^24 fingerprint 7c28cfb06c5c65a9).
## Files
| File | What |
|---|---|
| `gen_block.py` | draws the 512-step table, writes `mm8_block.h` (packed uint16 table + X-macro step list) |
| `kernel_mm8.cu` | the pack's kernel.cu with r0..r7 as `uint32_t r[8]` (mechanical rewrite, same arithmetic) and the mm8 block; PTX path by default, `-DIGNEUM_MM8_REF` for the shuffle-and-byte-product reference path (12 shuffles + 32 byte products per lane per step). Cache fill, build and launch wrappers unchanged |
| `bench.cu` | harness derived from proto-cuda/host.cu (serve mode stripped): device info, `igneum_hash_info` registers and occupancy, GPU cache fill + host fill + FNV check, GPU dataset build + self-test, pack vectors at R = 0, 2^24 fingerprint at base nonce 0, 1 warm-up + 10 timed batches (CUDA events), `--sustain S` for the power meter, `--dump file n` |
| `gen_ref_program.py` | turns the 64 instruction lines of kernel.cu into the 32-lane C interpreter body `ref_program.inc` (nothing transcribed by hand) |
| `verify_ref.c` | plain C CPU reference: host cache (65536 segments), lazy `mh_word` loads, register-major interpreter, mm8 block in the spec layout, fold; compares a dump, then times the verifier per unit |
| `run.sh` | the ladder R in {0, 8, 32, 128, 512}: builds, power sampling, fingerprints, dumps, CPU check, summary |
| `summarise.py` | builds the RESULTS table from `out/` |
## Exact commands (on the box, under /root/horizon-newpow/mma-shadow)
```
export PATH=/usr/local/cuda/bin:$PATH
python3 gen_block.py mm8_block.h
python3 gen_ref_program.py kernel.cu
gcc -O2 -o verify_ref verify_ref.c
for R in 0 8 32 128 512; do
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -Xptxas -v -o bench_$R bench.cu kernel_mm8.cu
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -DIGNEUM_MM8_REF -o bench_${R}_ref bench.cu kernel_mm8.cu
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu,timestamp --format=csv,noheader -l 1 > out/power_$R.csv &
./bench_$R --batches 10 --sustain 25 --dump out/dump_$R.txt 32 > out/bench_$R.log
kill %1
./bench_${R}_ref --fingerprint-only --no-host-cache > out/bench_${R}_ref.log
taskset -c 2 ./verify_ref out/dump_$R.txt $R --time > out/verify_$R.log # add --self-test at R = 0
done
python3 summarise.py > out/results.md
```
`./run.sh` is exactly that sequence (plus the idle baseline and the PTX == ref fingerprint comparison). From the
Mac: `rsync -az -e "ssh -i ~/.ssh/igneum-fleet -p <box-1-port>" proto-newpow/mma-shadow/ root@<box-1-ip>:/root/horizon-newpow/mma-shadow/`
then `touch *` on the far side before building.
## Card and driver
NVIDIA GeForce RTX 4090, 128 SMs, cc 8.9, 24 GB. Driver 570.172.08 (the brief said 595; nvidia-smi reports 570.172.08),
CUDA 12.8 (nvcc V12.8.93), g++ 13.3.0, Ubuntu 24.04, 96 CPU threads. Idle baseline before the run: 15.0 to 15.3 W at
210 MHz SM, 45 C, 0 % utilisation. The host carried a CPU load average of about 13 from other tenants during the run
(the GPU itself was idle and ours alone); that is why the verifier timings were pinned to one core.
## Method notes
- Timing: one warm-up batch (also the fingerprint batch) then 10 timed batches of 2^24 hashes between CUDA events,
1 warp per block (the pack bench default, 24 resident warps per SM at 29 registers). Power: nvidia-smi at 1 Hz
during a 25 s sustained phase after the timed batches; the mean takes samples from 10 s after the sustain start to
its end (the timestamps are in the csv). Microjoules per hash = mean watts / (MH/s x 10^6) x 10^6.
- Fingerprint = FNV-1a 64 over the 2^24 little-endian u64 outputs at base nonce 0. The PTX build and the reference
build must agree (the reference is the plain-integer byte-product form of the same fragment layout).
- CPU == GPU: 32 warps at SplitMix64(0x1234) 32-aligned bases, 1,024 lanes, recomputed by `verify_ref`.
- The verifier here is a naive register-major interpreter with no interleaving and a lazy `mh_word` per load
(72 mixer applications and 8 cache reads per word, 4,096 words per unit). It is slower than the project's Rust
verifier (2.06 ms per unit on an M5 Max core) on this box's core, so the number that matters is the mm8 block
delta (ms per unit at R minus ms per unit at R = 0), not the total.
## RESULTS (6 Oct 2026, RTX 4090, driver 570.172.08, CUDA 12.8, sm_89, 1 warp/block, batch 2^24 x 10)
| R | mm8 per hash (8R) | MH/s (GPU time) | ratio to R = 0 | watts mean (sustain, after first 10 s) | SM MHz | uJ per hash | max C | fingerprint (2^24 at base 0) | PTX == ref | CPU == GPU lanes | verifier ms per unit on the box core: R = 0, R, block delta | regs per thread (PTX build) | regs (ref build) | blocks per SM (cudaOccupancy) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 (matches the pack) | yes | 1024 of 1024 | 10.18, 10.08, -0.10 (noise) | 29 | 29 | 24 (24 warps/SM, 50 %) |
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.93, 9.97, +0.05 | 30 | 77 | 24 |
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.08, 10.24, +0.17 | 29 | 151 | 24 |
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.18, 11.33, +1.14 | 32 | 175 | 24 |
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.89, 14.28, +4.39 | 36 | 213 | 24 |
Sustained MH/s over the 25 s power window: 63.075, 63.074, 63.072, 63.034, 63.073 for R = 0, 8, 32, 128, 512 (wall,
including a sync per batch). Wall and GPU-event rates agree to 0.01 MH/s at every R. Idle baseline 15.0 to 15.3 W.
Registers per thread at R = 0, 8, 32, 128, 512: 29, 30, 29, 32, 36 (PTX path), no spills, no stack, at every R.
R = 0 self-checks: cache FNV 48c4f5bf24166b2e PASS (GPU == host fill word for word), dataset head, [MASK], 64 random
points and 64 Mac samples PASS, the pack's 3 vector warps PASS standalone and in batch, fingerprint 7c28cfb06c5c65a9.
The C interpreter also reproduces the 3 pack vectors at R = 0 (`--self-test`).
## What the numbers say
- On the RTX 4090 the mm8 block is free in hash rate up to R = 512: 4,096 tensor instructions per hash leave the
rate at 63.08 MH/s, identical to the control to the third decimal. The kernel is latency-bound on the 16 dependent
random loads per iteration; the tensor work fills stalls that were already there. At R = 512 the card issues about
2.6 x 10^11 mma.m8n8k16 per second, roughly 2.7 x 10^14 u8 multiply-adds per second (1,024 per instruction), which is
in the region of 40 % of the card's dense int8 tensor peak (approximate, from the published TOPS figure). So the
next doublings would start to cost hash rate; R = 512 is near the top of the free band, not in the middle of it.
- Power is the only GPU cost that moves: 201 W to 216 W (+7.3 %) and 3.19 to 3.42 uJ per hash (+7.2 %) from R = 0
to R = 512. SM clock stayed pinned at 2670 MHz at every R, temperature peaked at 64 C, no throttling seen.
- Correctness chain holds at every rung: the PTX fragment read and the plain-integer reference path agree on all
2^24 lanes of the fingerprint batch at every R, and the CPU interpreter matches the GPU on all 1,024 dumped lanes.
The probe's fragment layout (family-probe.cu `warp_ref`, mm8) was used as written and needed no correction.
- Verifier cost of the block on the box's core: about 1.1 us per mm8 step per unit (32 lanes x 32 byte products,
naive loops), so +1.14 ms per unit at R = 128 and +4.39 ms per unit at R = 512. Against the project's Rust verifier
at 2.06 ms per unit, R = 512 would roughly triple verification time unless the block is vectorised (the 8 x 8 x 16
tile is 1,024 MACs, a few hundred ns with SIMD, approximate); R = 32 adds 0.17 ms per unit (+8 % of 2.06 ms) and
R = 128 adds 1.14 ms (+55 %). The totals in the table (about 10 ms per unit) are this naive interpreter's cost
with a lazy mh_word per load and are not comparable to the Rust verifier; the delta column is the number to use.
- The reference-path build is only a correctness oracle: fully unrolled it reaches 213 registers at R = 512 (8 blocks
per SM) and was never timed.
## Raw outputs
`out/` holds everything the box produced: `run.log` (the whole run), `bench_R.log` and `bench_R_ref.log`,
`power_R.csv` (1 Hz: W, SM MHz, C, util, timestamp), `power_idle.csv`, `ptxas_R.txt` and `ptxas_R_ref.txt`,
`dump_R.txt` (32 warps, 1,024 lanes), `verify_R.log`, `results.md` (summarise.py).
## What was cut
Nothing from the brief. The host 1 GiB dataset was never built on the CPU (lazy `mh_word`, as allowed). The 32-lane
dump bases are 32-bit, so base + 31 cannot overflow (base = low32(next()) & ~31).

View file

@ -0,0 +1,308 @@
// bench.cu (proto-newpow/mma-shadow): host harness for the mx8+mm8xR prototype, derived from proto-cuda/host.cu
// with serve mode stripped. TEST HARNESS ONLY: no pool, no network, no wallet.
//
// Build: nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=<R> [-DIGNEUM_MM8_REF] -o bench_<R> bench.cu kernel_mm8.cu
// Steps: device info, kernel registers and occupancy, GPU cache fill + host cache fill + check (FNV, head, last),
// GPU dataset build + self-test (head, [MASK], 64 random points vs host mh_word, Mac samples), the pack's 3 vector warps
// (R == 0 only; they must PASS), the 2^24 batch fingerprint at base nonce 0 (FNV-1a 64 over the output bytes), timing
// (1 warm-up + N timed batches with CUDA events), an optional sustain phase for the power meter, and --dump.
#include <cuda_runtime.h>
#include <cstdint>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <chrono>
#include <string>
#include <vector>
#include "program.h"
#include "vectors.h"
#include "memhard.h"
#include "mm8_block.h"
#ifndef IGNEUM_MM8_R
#define IGNEUM_MM8_R 0
#endif
#define CUDA_CHECK(call) do { cudaError_t err_ = (call); if (err_ != cudaSuccess) { \
std::fprintf(stderr, "CUDA error: %s (%d)\n at %s:%d\n in %s\n", cudaGetErrorString(err_), (int)err_, __FILE__, __LINE__, #call); \
std::exit(2); } } while (0)
static double wallMs() {
using namespace std::chrono;
return duration<double, std::milli>(steady_clock::now().time_since_epoch()).count();
}
static double epochS() {
using namespace std::chrono;
return duration<double>(system_clock::now().time_since_epoch()).count();
}
static uint64_t fnv1a64(const void* p, size_t n) {
const uint8_t* b = (const uint8_t*)p;
uint64_t h = 0xcbf29ce484222325ull;
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
return h;
}
static uint64_t splitmix64(uint64_t& s) {
s += 0x9E3779B97F4A7C15ull;
uint64_t z = s;
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull;
z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
return z ^ (z >> 31);
}
static const uint32_t CACHE_WORDS = 1u << IGNEUM_CACHE_LOG2_WORDS;
static uint32_t* gCache = nullptr;
static std::vector<uint32_t> hCache;
static bool setupCache(bool hostFill) {
size_t bytes = (size_t)CACHE_WORDS * 4u;
CUDA_CHECK(cudaMalloc((void**)&gCache, bytes));
cudaEvent_t e0, e1;
CUDA_CHECK(cudaEventCreate(&e0)); CUDA_CHECK(cudaEventCreate(&e1));
float ms[2] = {0.f, 0.f};
for (int pass = 0; pass < 2; ++pass) {
CUDA_CHECK(cudaEventRecord(e0));
CUDA_CHECK(igneum_launch_cache_fill(gCache, IGNEUM_CACHE_SEGMENTS));
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
CUDA_CHECK(cudaEventElapsedTime(&ms[pass], e0, e1));
}
std::printf("cache fill (GPU): %.2f ms first, %.2f ms second (%u MiB)\n", ms[0], ms[1], (unsigned)(bytes >> 20));
std::vector<uint32_t> dev(CACHE_WORDS);
CUDA_CHECK(cudaMemcpy(dev.data(), gCache, bytes, cudaMemcpyDeviceToHost));
uint64_t fnvDev = fnv1a64(dev.data(), bytes);
bool fnvOk = fnvDev == IGNEUM_CACHE_FNV64;
bool headOk = std::memcmp(dev.data(), IGNEUM_CACHE_HEAD, 64) == 0;
bool lastOk = std::memcmp(dev.data() + CACHE_WORDS - 16u, IGNEUM_CACHE_LAST, 64) == 0;
bool same = true;
double hostMs = 0;
if (hostFill) {
hCache.assign(CACHE_WORDS, 0u);
double h0 = wallMs();
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(hCache.data(), seg);
hostMs = wallMs() - h0;
same = std::memcmp(dev.data(), hCache.data(), bytes) == 0;
std::printf("cache fill (host, one thread): %.1f ms, GPU == host all words: %s\n", hostMs, same ? "PASS" : "FAIL");
} else {
hCache.swap(dev); // the device cache (FNV-checked) serves the host-side dataset derivation
}
bool pass = fnvOk && headOk && lastOk && same;
std::printf("cache check: %s (device FNV-1a 64 %016llx vs Mac %016llx %s, head %s, last line %s)\n",
pass ? "PASS" : "FAIL", (unsigned long long)fnvDev, (unsigned long long)IGNEUM_CACHE_FNV64,
fnvOk ? "PASS" : "FAIL", headOk ? "PASS" : "FAIL", lastOk ? "PASS" : "FAIL");
CUDA_CHECK(cudaEventDestroy(e0)); CUDA_CHECK(cudaEventDestroy(e1));
return pass;
}
static bool compareWarp(const uint64_t* got, const uint64_t* want, uint32_t base, const char* how) {
int bad = 0, first = -1;
for (int l = 0; l < 32; ++l) if (got[l] != want[l]) { if (first < 0) first = l; ++bad; }
if (bad == 0) std::printf("verify warp base %u %s: PASS\n", base, how);
else std::printf("verify warp base %u %s: FAIL %d of 32 lanes differ, first lane %d: gpu=%016llx expected=%016llx\n",
base, how, bad, first, (unsigned long long)got[first], (unsigned long long)want[first]);
return bad == 0;
}
static void usage() {
std::printf("bench_R [--batches 10] [--block-warps 1] [--batch-log2 24] [--fingerprint-only] [--sustain S] [--no-host-cache]\n"
" [--dump <file> <n>] [--device 0]\n"
" --fingerprint-only skip the timed batches (for the reference build)\n"
" --sustain S after the timed batches keep launching batches for S seconds (power meter window)\n"
" --dump file n write n whole warps at SplitMix64(0x1234) 32-aligned bases as lines: base lane value_hex\n"
" --no-host-cache skip the one-thread host cache fill (the FNV-checked device cache then feeds the host derivation)\n");
}
int main(int argc, char** argv) {
int batches = 10, blockWarps = 1, batchLog2 = 24, device = 0, dumpN = 0;
double sustain = 0;
bool fpOnly = false, hostCache = true;
std::string dumpFile;
for (int i = 1; i < argc; ++i) {
std::string a = argv[i];
auto nextInt = [&](int& dst) { if (i + 1 >= argc) { usage(); std::exit(2); } dst = std::atoi(argv[++i]); };
if (a == "--batches") nextInt(batches);
else if (a == "--block-warps") nextInt(blockWarps);
else if (a == "--batch-log2") nextInt(batchLog2);
else if (a == "--device") nextInt(device);
else if (a == "--fingerprint-only") fpOnly = true;
else if (a == "--no-host-cache") hostCache = false;
else if (a == "--sustain") { if (i + 1 >= argc) { usage(); return 2; } sustain = std::atof(argv[++i]); }
else if (a == "--dump") { if (i + 2 >= argc) { usage(); return 2; } dumpFile = argv[++i]; dumpN = std::atoi(argv[++i]); }
else if (a == "-h" || a == "--help") { usage(); return 0; }
else { std::printf("unknown argument %s\n", argv[i]); usage(); return 2; }
}
if (blockWarps < 1 || blockWarps > 32 || batchLog2 < 10 || batchLog2 > 28 || batches < 1) { usage(); return 2; }
std::printf("mma-shadow bench pack \"%s\" class mx8+mm8xR R = %d mm8 per hash = %d path = %s\n",
IGNEUM_SEED_STRING, IGNEUM_MM8_R, 8 * IGNEUM_MM8_R,
#ifdef IGNEUM_MM8_REF
"reference (shuffles + byte products)"
#else
"PTX mma.sync.m8n8k16.u8"
#endif
);
std::printf("mm8 table seed 0x%016llx (first steps: %s)\n", (unsigned long long)IGNEUM_MM8_SEED, "see mm8_block.h");
int count = 0;
CUDA_CHECK(cudaGetDeviceCount(&count));
if (count == 0 || device >= count) { std::printf("FAIL: no CUDA device %d\n", device); return 2; }
CUDA_CHECK(cudaSetDevice(device));
cudaDeviceProp prop; std::memset(&prop, 0, sizeof(prop));
CUDA_CHECK(cudaGetDeviceProperties(&prop, device));
int drv = 0, rt = 0; CUDA_CHECK(cudaDriverGetVersion(&drv)); CUDA_CHECK(cudaRuntimeGetVersion(&rt));
int clk = 0, memclk = 0, bus = 0, l2 = 0, thrSM = 0;
cudaDeviceGetAttribute(&clk, cudaDevAttrClockRate, device);
cudaDeviceGetAttribute(&memclk, cudaDevAttrMemoryClockRate, device);
cudaDeviceGetAttribute(&bus, cudaDevAttrGlobalMemoryBusWidth, device);
cudaDeviceGetAttribute(&l2, cudaDevAttrL2CacheSize, device);
cudaDeviceGetAttribute(&thrSM, cudaDevAttrMaxThreadsPerMultiProcessor, device);
cudaGetLastError();
std::printf("GPU: %s (%d SMs, cc %d.%d, %.0f MiB) SM clock %d MHz, mem clock %d MHz, bus %d bits, L2 %d MiB, max %d threads/SM\n",
prop.name, prop.multiProcessorCount, prop.major, prop.minor, (double)prop.totalGlobalMem / 1048576.0,
clk / 1000, memclk / 1000, bus, l2 / 1048576, thrSM);
std::printf("CUDA: driver %d.%d, runtime %d.%d\n", drv / 1000, (drv % 100) / 10, rt / 1000, (rt % 100) / 10);
int regs = 0, blocksPerSM = 0;
CUDA_CHECK(igneum_hash_info(&regs, &blocksPerSM, (uint32_t)blockWarps));
std::printf("igneum_hash_info: %d registers/thread, %d resident blocks/SM at %d warp(s)/block = %d resident warps/SM (%.1f%% of %d)\n",
regs, blocksPerSM, blockWarps, blocksPerSM * blockWarps, 100.0 * blocksPerSM * blockWarps * 32 / thrSM, thrSM / 32);
bool cachePass = setupCache(hostCache);
uint32_t nonces = 1u << batchLog2;
uint32_t mask = IGNEUM_MASK;
uint64_t dsBytes = (uint64_t)(mask + 1u) * 4ull;
uint32_t* dDs = nullptr;
CUDA_CHECK(cudaMalloc((void**)&dDs, (size_t)dsBytes));
uint64_t* dOut = nullptr;
CUDA_CHECK(cudaMalloc((void**)&dOut, (size_t)nonces * sizeof(uint64_t)));
cudaEvent_t e0, e1;
CUDA_CHECK(cudaEventCreate(&e0)); CUDA_CHECK(cudaEventCreate(&e1));
float buildMs[2] = {0.f, 0.f};
for (int pass = 0; pass < 2; ++pass) {
CUDA_CHECK(cudaEventRecord(e0));
CUDA_CHECK(igneum_launch_build(dDs, gCache, (mask + 1u) / 16u));
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
CUDA_CHECK(cudaEventElapsedTime(&buildMs[pass], e0, e1));
}
std::printf("dataset build (GPU, %llu MiB): %.2f ms first, %.2f ms second\n", (unsigned long long)(dsBytes >> 20), buildMs[0], buildMs[1]);
// Dataset self-test
bool dsPass;
{
uint32_t head[16];
CUDA_CHECK(cudaMemcpy(head, dDs, sizeof(head), cudaMemcpyDeviceToHost));
int badHead = 0;
for (int i = 0; i < 16; ++i) if (head[i] != IGNEUM_DS_HEAD[i]) { if (!badHead) std::printf(" dataset[%d] = %08x, Mac %08x\n", i, head[i], IGNEUM_DS_HEAD[i]); ++badHead; }
uint32_t last = 0;
CUDA_CHECK(cudaMemcpy(&last, dDs + IGNEUM_DS_LAST_INDEX, 4, cudaMemcpyDeviceToHost));
bool lastOk = last == IGNEUM_DS_LAST;
int badRnd = 0;
uint64_t s = 0x9E3779B97F4A7C15ull ^ (uint64_t)(mask + 1u);
for (int k = 0; k < 64; ++k) {
uint32_t idx = (uint32_t)splitmix64(s) & mask, v = 0;
CUDA_CHECK(cudaMemcpy(&v, dDs + idx, 4, cudaMemcpyDeviceToHost));
uint32_t want = mh_word(hCache.data(), idx);
if (v != want) { if (!badRnd) std::printf(" dataset[%u] = %08x, host derivation %08x\n", idx, v, want); ++badRnd; }
}
int badSample = 0;
for (int k = 0; k < IGNEUM_DS_SAMPLES; ++k) {
uint32_t v = 0;
CUDA_CHECK(cudaMemcpy(&v, dDs + IGNEUM_DS_SAMPLE_INDEX[k], 4, cudaMemcpyDeviceToHost));
if (v != IGNEUM_DS_SAMPLE_VALUE[k]) { if (!badSample) std::printf(" dataset[%u] = %08x, Mac %08x\n", IGNEUM_DS_SAMPLE_INDEX[k], v, IGNEUM_DS_SAMPLE_VALUE[k]); ++badSample; }
}
dsPass = badHead == 0 && lastOk && badRnd == 0 && badSample == 0;
std::printf("dataset self-test: %s (head 16 %s, [MASK] %s, 64 random points vs host derivation %s, %d Mac samples %s)\n",
dsPass ? "PASS" : "FAIL", badHead == 0 ? "PASS" : "FAIL", lastOk ? "PASS" : "FAIL",
badRnd == 0 ? "PASS" : "FAIL", (int)IGNEUM_DS_SAMPLES, badSample == 0 ? "PASS" : "FAIL");
}
// Pack vectors (R == 0 only: at R > 0 the function is different by design)
bool vecPass = true;
uint64_t got[32];
if (IGNEUM_MM8_R == 0) {
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
CUDA_CHECK(igneum_launch_hash(dDs, dOut, IGNEUM_VEC_BASE[w], mask, 32u, 1u));
CUDA_CHECK(cudaDeviceSynchronize());
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
vecPass = compareWarp(got, IGNEUM_VEC_OUT[w], IGNEUM_VEC_BASE[w], "(pack vector, standalone)") && vecPass;
}
} else {
std::printf("pack vectors: not applicable at R = %d (checked at R = 0 only)\n", IGNEUM_MM8_R);
}
// Warm-up batch at base 0 = the fingerprint batch
double w0 = wallMs();
CUDA_CHECK(igneum_launch_hash(dDs, dOut, 0u, mask, nonces, (uint32_t)blockWarps));
CUDA_CHECK(cudaDeviceSynchronize());
double w1 = wallMs();
std::vector<uint64_t> hOut(nonces);
CUDA_CHECK(cudaMemcpy(hOut.data(), dOut, (size_t)nonces * 8u, cudaMemcpyDeviceToHost));
uint64_t fp = fnv1a64(hOut.data(), (size_t)nonces * 8u);
std::printf("warm-up batch: %u hashes in %.2f ms wall\n", nonces, w1 - w0);
std::printf("fingerprint: %016llx (FNV-1a 64 over the 2^%d outputs at base nonce 0, little-endian u64 bytes)%s\n",
(unsigned long long)fp, batchLog2,
IGNEUM_MM8_R == 0 ? (fp == 0x7c28cfb06c5c65a9ull ? " == 7c28cfb06c5c65a9 PASS" : " != 7c28cfb06c5c65a9 FAIL") : "");
bool fpPass = (IGNEUM_MM8_R != 0) || fp == 0x7c28cfb06c5c65a9ull;
if (IGNEUM_MM8_R == 0 && batchLog2 >= 20) {
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
if ((uint64_t)IGNEUM_VEC_BASE[w] + 32ull > nonces) continue;
vecPass = compareWarp(hOut.data() + IGNEUM_VEC_BASE[w], IGNEUM_VEC_OUT[w], IGNEUM_VEC_BASE[w], "(pack vector, in batch)") && vecPass;
}
}
double mhs = 0, mhsWall = 0, mhsSustain = 0;
if (!fpOnly) {
CUDA_CHECK(cudaEventRecord(e0));
double t0 = wallMs();
for (int b = 1; b <= batches; ++b) {
uint32_t base = (uint32_t)((uint64_t)b * (uint64_t)nonces);
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, (uint32_t)blockWarps));
}
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
double t1 = wallMs();
float gpuMs = 0.f;
CUDA_CHECK(cudaEventElapsedTime(&gpuMs, e0, e1));
double total = (double)nonces * batches;
mhs = total / (gpuMs / 1000.0) / 1e6;
mhsWall = total / ((t1 - t0) / 1000.0) / 1e6;
std::printf("timed: %d batches x %u hashes: GPU %.2f ms -> %.3f MH/s (GPU time), wall %.2f ms -> %.3f MH/s\n",
batches, nonces, gpuMs, mhs, t1 - t0, mhsWall);
if (sustain > 0) {
double s0 = epochS(), sw0 = wallMs();
long long n = 0;
std::printf("sustain start epoch %.3f\n", s0); std::fflush(stdout);
while (wallMs() - sw0 < sustain * 1000.0) {
uint32_t base = (uint32_t)((uint64_t)(n + batches + 1) * (uint64_t)nonces);
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, (uint32_t)blockWarps));
CUDA_CHECK(cudaDeviceSynchronize());
++n;
}
double s1 = epochS();
mhsSustain = (double)n * nonces / (s1 - s0) / 1e6;
std::printf("sustain end epoch %.3f: %lld batches in %.2f s -> %.3f MH/s (wall, incl. sync)\n", s1, n, s1 - s0, mhsSustain);
}
}
if (dumpN > 0) {
FILE* f = std::fopen(dumpFile.c_str(), "w");
if (!f) { std::printf("FAIL: cannot open %s\n", dumpFile.c_str()); return 2; }
uint64_t s = 0x1234ull;
for (int i = 0; i < dumpN; ++i) {
uint32_t base = (uint32_t)splitmix64(s) & ~31u;
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, 32u, 1u));
CUDA_CHECK(cudaDeviceSynchronize());
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
for (int l = 0; l < 32; ++l) std::fprintf(f, "%u %d %016llx\n", base, l, (unsigned long long)got[l]);
}
std::fclose(f);
std::printf("dump: %d warps (%d lanes) written to %s\n", dumpN, 32 * dumpN, dumpFile.c_str());
}
bool overall = cachePass && dsPass && vecPass && fpPass;
std::printf("SUMMARY R=%d regs=%d blocksPerSM=%d mhs=%.3f mhs_wall=%.3f mhs_sustain=%.3f fingerprint=%016llx cache=%s dataset=%s vectors=%s\n",
IGNEUM_MM8_R, regs, blocksPerSM, mhs, mhsWall, mhsSustain, (unsigned long long)fp,
cachePass ? "PASS" : "FAIL", dsPass ? "PASS" : "FAIL", IGNEUM_MM8_R == 0 ? (vecPass ? "PASS" : "FAIL") : "n/a");
std::printf("OVERALL: %s\n", overall ? "PASS" : "FAIL");
CUDA_CHECK(cudaFree(dOut)); CUDA_CHECK(cudaFree(dDs)); CUDA_CHECK(cudaFree(gCache));
return overall ? 0 : 1;
}

View file

@ -0,0 +1,75 @@
#!/usr/bin/env python3
# gen_block.py: draws the mm8 block table for the mx8+mm8xR prototype and writes mm8_block.h.
# Stream: SplitMix64 seeded with FNV-1a-64 of the bytes "igneum-mm8/igneum-genesis".
# Per step k: a = below(8); b = below(7), b += (b >= a); c = below(8); c2 = below(7), c2 += (c2 >= c).
# Step semantics: r[c] += C[l >> 2][2 * (l & 3)] (PTX d0), r[c2] += C[l >> 2][2 * (l & 3) + 1] (PTX d1), both mod 2^32.
import sys
MASK64 = (1 << 64) - 1
R_MAX = 512
def fnv1a64(data: bytes) -> int:
h = 0xcbf29ce484222325
for byte in data:
h ^= byte
h = (h * 0x100000001b3) & MASK64
return h
class SplitMix64:
def __init__(self, seed: int):
self.s = seed & MASK64
def next(self) -> int:
self.s = (self.s + 0x9E3779B97F4A7C15) & MASK64
z = self.s
z = ((z ^ (z >> 30)) * 0xBF58476D1CE4E5B9) & MASK64
z = ((z ^ (z >> 27)) * 0x94D049BB133111EB) & MASK64
return z ^ (z >> 31)
def below(self, n: int) -> int:
return self.next() % n
def main():
seed_text = b"igneum-mm8/igneum-genesis"
seed = fnv1a64(seed_text)
rng = SplitMix64(seed)
rows = []
for k in range(R_MAX):
a = rng.below(8)
b = rng.below(7)
b += 1 if b >= a else 0
c = rng.below(8)
c2 = rng.below(7)
c2 += 1 if c2 >= c else 0
assert a != b and 0 <= b < 8 and c != c2 and 0 <= c2 < 8
rows.append((a, b, c, c2))
out = []
out.append("// Generated by gen_block.py. mm8 block draws for class mx8+mm8xR, seed text \"%s\"," % seed_text.decode())
out.append("// FNV-1a-64 seed 0x%016x, SplitMix64 stream. Step k uses (a, b, c, c2) = row k: r[c] += d0, r[c2] += d1. Do not edit by hand." % seed)
out.append("#pragma once")
out.append("#ifdef __cplusplus")
out.append("#include <cstdint>")
out.append("#else")
out.append("#include <stdint.h>")
out.append("#endif")
out.append("#define IGNEUM_MM8_R_MAX %d" % R_MAX)
out.append("#define IGNEUM_MM8_SEED 0x%016xull" % seed)
out.append("// Packed uint16: a = v & 7, b = (v >> 3) & 7, c = (v >> 6) & 7, c2 = (v >> 9) & 7.")
out.append("#define IGNEUM_MM8_TABLE_INIT { \\")
for i in range(0, R_MAX, 16):
chunk = rows[i:i+16]
vals = ["0x%03xu" % (a | (b << 3) | (c << 6) | (c2 << 9)) for (a, b, c, c2) in chunk]
out.append(" " + ", ".join(vals) + (", \\" if i + 16 < R_MAX else " }"))
out.append("// X-macro list: X(k, a, b, c, c2) for every step k in 0..R_MAX-1. The kernel guards each with k < IGNEUM_MM8_R,")
out.append("// so the register indices are compile-time constants (r[] stays in registers, no local memory).")
out.append("#define IGNEUM_MM8_STEPS(X) \\")
for k, (a, b, c, c2) in enumerate(rows):
out.append(" X(%d, %d, %d, %d, %d)%s" % (k, a, b, c, c2, " \\" if k + 1 < R_MAX else ""))
out.append("// Readable form, step: a b c c2")
for k, (a, b, c, c2) in enumerate(rows):
out.append("// %3d: %d %d %d %d" % (k, a, b, c, c2))
path = sys.argv[1] if len(sys.argv) > 1 else "mm8_block.h"
with open(path, "w") as f:
f.write("\n".join(out) + "\n")
print("seed 0x%016x, %d steps, first 4: %s -> %s" % (seed, R_MAX, rows[:4], path))
if __name__ == "__main__":
main()

View file

@ -0,0 +1,23 @@
#!/usr/bin/env python3
# gen_ref_program.py: turns the 64 instruction lines of the pack's kernel.cu into a 32-lane C interpreter body
# (ref_program.inc, included by verify_ref.c). Register-major: R[8][32]. A shuffle copies its source register
# into T before the lane loop and reads T[l ^ mask]. Loads become mh_word(cache, x & mask) (lazy dataset).
import re, sys
src = open(sys.argv[1] if len(sys.argv) > 1 else "kernel.cu").read().split("\n")
lines = [l for l in src if re.search(r"//\s*\d+ (mad|xor|load|shfl|mulhi|rotr|or|add|sub|mul|rotl)\s*$", l)]
assert len(lines) == 64, len(lines)
out = ["// Generated by gen_ref_program.py from the pack's kernel.cu. Do not edit by hand.",
"// Expects: uint32_t R[8][32], T[32], SEL[32]; const uint32_t* cache; uint32_t mask; macro L = for (int l = 0; l < 32; ++l)."]
for i, l in enumerate(lines):
body, comment = l.strip().rsplit("//", 1)
body = body.strip()
m = re.search(r"__shfl_xor_sync\(0xffffffffu, r(\d), (\d+)\)", body)
if m:
out.append(" memcpy(T, R[%s], sizeof T);" % m.group(1))
body = body.replace(m.group(0), "T[l ^ %s]" % m.group(2))
body = re.sub(r"ds\[(r\d) & mask\]", r"mh_word(cache, \1 & mask)", body)
body = re.sub(r"\br([0-7])\b", r"R[\1][l]", body)
body = body.replace("__umulhi(", "umulhi(").replace("sel ", "SEL[l] ")
out.append(" L { %s } //%s" % (body, comment))
open("ref_program.inc", "w").write("\n".join(out) + "\n")
print("wrote ref_program.inc with %d instructions" % len(lines))

View file

@ -0,0 +1,164 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
uint32_t* d = ds + (size_t)t * 16u;
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r2 = r3 * r4 + r2; // 0 mad
r2 = r1 * r1 + r2; // 1 mad
r2 = r3 * r2 + r2; // 2 mad
r3 = r3 ^ r5; // 3 xor
r7 = r7 ^ ds[r2 & mask]; // 4 load
r5 = r5 ^ ds[r7 & mask]; // 5 load
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
r1 = __umulhi(r1, r5); // 8 mulhi
r6 = rotr_var(r6, r3); // 9 rotr
r3 = r3 | r4; // 10 or
r4 = r4 ^ ds[r3 & mask]; // 11 load
r0 = __umulhi(r0, r4); // 12 mulhi
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
r0 = r0 ^ ds[r4 & mask]; // 14 load
r2 = r2 - r4; // 15 sub
r2 = r2 ^ ds[r0 & mask]; // 16 load
r7 = r7 ^ ds[r2 & mask]; // 17 load
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
r5 = r5 * r0; // 19 mul
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
r6 = __umulhi(r6, r2); // 22 mulhi
r6 = r6 ^ ds[r1 & mask]; // 23 load
r5 = r5 * r0; // 24 mul
r5 = rotl_imm(r5, 19u); // 25 rotl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
r0 = r0 ^ r5; // 27 xor
r0 = r0 ^ r4; // 28 xor
r3 = r3 - r0; // 29 sub
r5 = r5 * r1; // 30 mul
r7 = r7 ^ ds[r2 & mask]; // 31 load
r1 = r1 ^ ds[r0 & mask]; // 32 load
r5 = r5 ^ r6; // 33 xor
r5 = r5 ^ ds[r1 & mask]; // 34 load
r0 = __umulhi(r0, r5); // 35 mulhi
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
r7 = r7 ^ ds[r0 & mask]; // 37 load
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
r2 = r2 ^ r5; // 40 xor
r3 = r6 * r3 + r3; // 41 mad
r6 = r6 - r7; // 42 sub
r7 = r7 ^ r0; // 43 xor
r1 = r1 ^ ds[r7 & mask]; // 44 load
r2 = r2 * r3; // 45 mul
r1 = __umulhi(r1, r5); // 46 mulhi
r4 = r4 - r3; // 47 sub
r2 = rotr_var(r2, r6); // 48 rotr
r3 = r3 ^ ds[r5 & mask]; // 49 load
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
r0 = r0 * r2; // 51 mul
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
r7 = rotl_imm(r7, 14u); // 54 rotl
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
r6 = r6 ^ ds[r7 & mask]; // 56 load
r1 = rotr_var(r1, r5); // 57 rotr
r5 = r5 ^ ds[r4 & mask]; // 58 load
r6 = r6 ^ ds[r2 & mask]; // 59 load
r3 = r5 * r0 + r3; // 60 mad
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
r5 = rotl_imm(r5, 19u); // 63 rotl
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,217 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// kernel_mm8.cu (proto-newpow/mma-shadow): the pack's kernel.cu with r0..r7 rewritten as uint32_t r[8] (same arithmetic)
// and a block of IGNEUM_MM8_R mm8 steps at the end of every iteration (class "mx8+mm8xR"). IGNEUM_MM8_R 0 is the control
// and is bit-exact with the pack. Define IGNEUM_MM8_REF for the shuffle-and-byte-product reference path instead of the PTX mma.
// Step semantics (6 Oct 2026 correction): both tile outputs are consumed, r[c] += d0 and r[c2] += d1.
// Cache fill, dataset build and the launch wrappers are the pack's text, unchanged.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
#include "mm8_block.h"
#ifndef IGNEUM_MM8_R
#define IGNEUM_MM8_R 0
#endif
#if IGNEUM_MM8_R < 0 || IGNEUM_MM8_R > IGNEUM_MM8_R_MAX
#error "IGNEUM_MM8_R out of range"
#endif
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
uint32_t* d = ds + (size_t)t * 16u;
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// One mm8 step. A = the 32 lanes' r[a] (8 x 16 u8, lane l holds A[l >> 2][4 * (l & 3) .. +3], byte 0 = lowest k),
// B = the 32 lanes' r[b] (16 x 8 u8, lane l holds B[4 * (l & 3) .. +3][l >> 2]). C = A * B exact in int32.
// Lane l adds C[l >> 2][2 * (l & 3)] into r[c] and C[l >> 2][2 * (l & 3) + 1] into r[c2], both modulo 2^32 (c != c2),
// so every step consumes both outputs of the tile. This is the PTX mma.m8n8k16 .u8 fragment layout (d0, d1), the same
// arithmetic as proto-cuda/family-probe.cu's mm8 warp_ref.
#ifdef IGNEUM_MM8_REF
// Reference path: gather A's row (l >> 2) and B's two columns 2 * (l & 3) and 2 * (l & 3) + 1 with 12 shuffles,
// then 32 byte products per lane in plain integer code.
__device__ __forceinline__ void mm8_pair(uint32_t av, uint32_t bv, uint32_t lane, uint32_t& d0, uint32_t& d1) {
uint32_t row = lane >> 2, col0 = 2u * (lane & 3u);
uint32_t acc0 = 0u, acc1 = 0u;
#pragma unroll
for (uint32_t j = 0u; j < 4u; ++j) {
uint32_t aw = __shfl_sync(0xffffffffu, av, (int)(4u * row + j));
uint32_t bw0 = __shfl_sync(0xffffffffu, bv, (int)(4u * col0 + j));
uint32_t bw1 = __shfl_sync(0xffffffffu, bv, (int)(4u * (col0 + 1u) + j));
#pragma unroll
for (uint32_t t = 0u; t < 4u; ++t) {
uint32_t ab = (aw >> (8u * t)) & 0xffu;
acc0 += ab * ((bw0 >> (8u * t)) & 0xffu);
acc1 += ab * ((bw1 >> (8u * t)) & 0xffu);
}
}
d0 = acc0; d1 = acc1;
}
#else
// Native path: one tensor instruction with a zero accumulator, d0 and d1 straight out of the fragment.
__device__ __forceinline__ void mm8_pair(uint32_t av, uint32_t bv, uint32_t lane, uint32_t& d0, uint32_t& d1) {
(void)lane;
asm("mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32 {%0,%1}, {%2}, {%3}, {%4,%5};"
: "=r"(d0), "=r"(d1) : "r"(av), "r"(bv), "r"(0u), "r"(0u));
}
#endif
#define IGNEUM_MM8_STEP(k, a, b, c, c2) if ((k) < IGNEUM_MM8_R) { uint32_t d0_, d1_; mm8_pair(r[a], r[b], lane, d0_, d1_); r[c] += d0_; r[c2] += d1_; }
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r[8];
uint32_t lane = threadIdx.x & 31u;
(void)lane;
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r[0] = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r[1] = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r[2] = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r[3] = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r[4] = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r[5] = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r[6] = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r[7] = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r[0];
r[2] = r[3] * r[4] + r[2]; // 0 mad
r[2] = r[1] * r[1] + r[2]; // 1 mad
r[2] = r[3] * r[2] + r[2]; // 2 mad
r[3] = r[3] ^ r[5]; // 3 xor
r[7] = r[7] ^ ds[r[2] & mask]; // 4 load
r[5] = r[5] ^ ds[r[7] & mask]; // 5 load
r[1] = r[1] ^ __shfl_xor_sync(0xffffffffu, r[4], 8); // 6 shfl
r[7] = r[7] ^ __shfl_xor_sync(0xffffffffu, r[3], 8); // 7 shfl
r[1] = __umulhi(r[1], r[5]); // 8 mulhi
r[6] = rotr_var(r[6], r[3]); // 9 rotr
r[3] = r[3] | r[4]; // 10 or
r[4] = r[4] ^ ds[r[3] & mask]; // 11 load
r[0] = __umulhi(r[0], r[4]); // 12 mulhi
r[5] = r[5] + r[1] + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
r[0] = r[0] ^ ds[r[4] & mask]; // 14 load
r[2] = r[2] - r[4]; // 15 sub
r[2] = r[2] ^ ds[r[0] & mask]; // 16 load
r[7] = r[7] ^ ds[r[2] & mask]; // 17 load
r[7] = r[7] ^ __shfl_xor_sync(0xffffffffu, r[3], 4); // 18 shfl
r[5] = r[5] * r[0]; // 19 mul
r[3] = r[3] ^ __shfl_xor_sync(0xffffffffu, r[4], 2); // 20 shfl
r[2] = r[2] ^ __shfl_xor_sync(0xffffffffu, r[4], 16); // 21 shfl
r[6] = __umulhi(r[6], r[2]); // 22 mulhi
r[6] = r[6] ^ ds[r[1] & mask]; // 23 load
r[5] = r[5] * r[0]; // 24 mul
r[5] = rotl_imm(r[5], 19u); // 25 rotl
r[7] = r[7] ^ __shfl_xor_sync(0xffffffffu, r[6], 2); // 26 shfl
r[0] = r[0] ^ r[5]; // 27 xor
r[0] = r[0] ^ r[4]; // 28 xor
r[3] = r[3] - r[0]; // 29 sub
r[5] = r[5] * r[1]; // 30 mul
r[7] = r[7] ^ ds[r[2] & mask]; // 31 load
r[1] = r[1] ^ ds[r[0] & mask]; // 32 load
r[5] = r[5] ^ r[6]; // 33 xor
r[5] = r[5] ^ ds[r[1] & mask]; // 34 load
r[0] = __umulhi(r[0], r[5]); // 35 mulhi
r[5] = r[5] ^ __shfl_xor_sync(0xffffffffu, r[2], 4); // 36 shfl
r[7] = r[7] ^ ds[r[0] & mask]; // 37 load
r[3] = r[3] + r[1] + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
r[1] = r[1] ^ __shfl_xor_sync(0xffffffffu, r[5], 4); // 39 shfl
r[2] = r[2] ^ r[5]; // 40 xor
r[3] = r[6] * r[3] + r[3]; // 41 mad
r[6] = r[6] - r[7]; // 42 sub
r[7] = r[7] ^ r[0]; // 43 xor
r[1] = r[1] ^ ds[r[7] & mask]; // 44 load
r[2] = r[2] * r[3]; // 45 mul
r[1] = __umulhi(r[1], r[5]); // 46 mulhi
r[4] = r[4] - r[3]; // 47 sub
r[2] = rotr_var(r[2], r[6]); // 48 rotr
r[3] = r[3] ^ ds[r[5] & mask]; // 49 load
r[1] = r[1] + r[5] + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
r[0] = r[0] * r[2]; // 51 mul
r[0] = r[0] + r[2] + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
r[1] = r[1] + r[0] + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
r[7] = rotl_imm(r[7], 14u); // 54 rotl
r[3] = r[3] + r[7] + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
r[6] = r[6] ^ ds[r[7] & mask]; // 56 load
r[1] = rotr_var(r[1], r[5]); // 57 rotr
r[5] = r[5] ^ ds[r[4] & mask]; // 58 load
r[6] = r[6] ^ ds[r[2] & mask]; // 59 load
r[3] = r[5] * r[0] + r[3]; // 60 mad
r[5] = r[5] + r[7] + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
r[4] = r[4] + r[6] + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
r[5] = rotl_imm(r[5], 19u); // 63 rotl
#if IGNEUM_MM8_R > 0
// mm8 block: IGNEUM_MM8_R steps after instruction 63, before the next iteration samples sel.
IGNEUM_MM8_STEPS(IGNEUM_MM8_STEP)
#endif
}
uint32_t lo = r[0] ^ rotl_imm(r[1], 7u) ^ rotl_imm(r[2], 14u) ^ rotl_imm(r[3], 21u);
uint32_t hi = r[4] ^ rotl_imm(r[5], 9u) ^ rotl_imm(r[6], 18u) ^ rotl_imm(r[7], 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,109 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#if defined(__CUDACC__)
#define IGNEUM_HD __host__ __device__ __forceinline__
#elif defined(_MSC_VER) && !defined(__cplusplus)
#define IGNEUM_HD static __inline
#else
#define IGNEUM_HD static inline
#endif
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint32_t r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0x3067619fu ^ prev[4];
x[5] = 0x3c269176u ^ prev[5];
x[6] = 0x84a03b03u ^ prev[6];
x[7] = 0xf8c63294u ^ prev[7];
x[8] = 0xff977c5bu ^ prev[8];
x[9] = 0xe60def3eu ^ prev[9];
x[10] = 0x63630141u ^ prev[10];
x[11] = 0xb8fbcb58u ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
s[0] = 0x3067619fu;
s[1] = 0x3c269176u;
s[2] = 0x84a03b03u;
s[3] = 0xf8c63294u;
s[4] = 0xff977c5bu;
s[5] = 0xe60def3eu;
s[6] = 0x63630141u;
s[7] = 0xb8fbcb58u;
s[8] = t * 0x42146205u + 0xbab68293u;
s[9] = t * 0x52cbe0fbu + 0xcc162340u;
s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
s[11] = t * 0x6728907fu + 0xe62b8997u;
s[12] = t * 0xd81d9751u + 0xc9c80297u;
s[13] = t * 0x132952c3u + 0xf74a1654u;
s[14] = t * 0xf60de277u + 0x3d704af5u;
s[15] = t * 0x05358035u + 0x3cf522b7u;
for (uint32_t r = 0u; r < 8u; ++r) {
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
}
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,24 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 0 mm8 per hash = 0 path = PTX mma.sync.m8n8k16.u8
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache fill (GPU): 1.93 ms first, 1.89 ms second (256 MiB)
cache fill (host, one thread): 375.8 ms, GPU == host all words: PASS
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.55 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
verify warp base 0 (pack vector, standalone): PASS
verify warp base 4096 (pack vector, standalone): PASS
verify warp base 1000000 (pack vector, standalone): PASS
warm-up batch: 16777216 hashes in 266.01 ms wall
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
verify warp base 0 (pack vector, in batch): PASS
verify warp base 4096 (pack vector, in batch): PASS
verify warp base 1000000 (pack vector, in batch): PASS
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.74 ms -> 63.078 MH/s
sustain start epoch 1791315749.655
sustain end epoch 1791315774.658: 94 batches in 25.00 s -> 63.075 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_0.txt
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.075 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
OVERALL: PASS

View file

@ -0,0 +1,19 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 0 mm8 per hash = 0 path = reference (shuffles + byte products)
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache fill (GPU): 1.80 ms first, 1.77 ms second (256 MiB)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.58 ms first, 30.51 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
verify warp base 0 (pack vector, standalone): PASS
verify warp base 4096 (pack vector, standalone): PASS
verify warp base 1000000 (pack vector, standalone): PASS
warm-up batch: 16777216 hashes in 266.00 ms wall
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
verify warp base 0 (pack vector, in batch): PASS
verify warp base 4096 (pack vector, in batch): PASS
verify warp base 1000000 (pack vector, in batch): PASS
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
OVERALL: PASS

View file

@ -0,0 +1,19 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 128 mm8 per hash = 1024 path = PTX mma.sync.m8n8k16.u8
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 32 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache fill (GPU): 1.91 ms first, 1.87 ms second (256 MiB)
cache fill (host, one thread): 347.9 ms, GPU == host all words: PASS
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.61 ms first, 30.56 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
pack vectors: not applicable at R = 128 (checked at R = 0 only)
warm-up batch: 16777216 hashes in 265.96 ms wall
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.82 ms -> 63.077 MH/s (GPU time), wall 2659.83 ms -> 63.076 MH/s
sustain start epoch 1791315890.874
sustain end epoch 1791315915.893: 94 batches in 25.02 s -> 63.034 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_128.txt
SUMMARY R=128 regs=32 blocksPerSM=24 mhs=63.077 mhs_wall=63.076 mhs_sustain=63.034 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS

View file

@ -0,0 +1,14 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 128 mm8 per hash = 1024 path = reference (shuffles + byte products)
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 175 registers/thread, 8 resident blocks/SM at 1 warp(s)/block = 8 resident warps/SM (16.7% of 48)
cache fill (GPU): 1.78 ms first, 1.79 ms second (256 MiB)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.58 ms first, 30.52 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
pack vectors: not applicable at R = 128 (checked at R = 0 only)
warm-up batch: 16777216 hashes in 447.05 ms wall
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=128 regs=175 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS

View file

@ -0,0 +1,19 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 32 mm8 per hash = 256 path = PTX mma.sync.m8n8k16.u8
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache fill (GPU): 1.90 ms first, 1.87 ms second (256 MiB)
cache fill (host, one thread): 349.6 ms, GPU == host all words: PASS
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.56 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
pack vectors: not applicable at R = 32 (checked at R = 0 only)
warm-up batch: 16777216 hashes in 266.00 ms wall
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.75 ms -> 63.078 MH/s (GPU time), wall 2659.77 ms -> 63.078 MH/s
sustain start epoch 1791315833.292
sustain end epoch 1791315858.296: 94 batches in 25.00 s -> 63.072 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_32.txt
SUMMARY R=32 regs=29 blocksPerSM=24 mhs=63.078 mhs_wall=63.078 mhs_sustain=63.072 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS

View file

@ -0,0 +1,14 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 32 mm8 per hash = 256 path = reference (shuffles + byte products)
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 151 registers/thread, 12 resident blocks/SM at 1 warp(s)/block = 12 resident warps/SM (25.0% of 48)
cache fill (GPU): 1.80 ms first, 1.77 ms second (256 MiB)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.57 ms first, 30.52 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
pack vectors: not applicable at R = 32 (checked at R = 0 only)
warm-up batch: 16777216 hashes in 265.96 ms wall
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS

View file

@ -0,0 +1,19 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 512 mm8 per hash = 4096 path = PTX mma.sync.m8n8k16.u8
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache fill (GPU): 1.93 ms first, 1.89 ms second (256 MiB)
cache fill (host, one thread): 348.1 ms, GPU == host all words: PASS
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.54 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
pack vectors: not applicable at R = 512 (checked at R = 0 only)
warm-up batch: 16777216 hashes in 265.98 ms wall
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
sustain start epoch 1791316287.337
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_512.txt
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS

View file

@ -0,0 +1,14 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 512 mm8 per hash = 4096 path = reference (shuffles + byte products)
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 213 registers/thread, 8 resident blocks/SM at 1 warp(s)/block = 8 resident warps/SM (16.7% of 48)
cache fill (GPU): 1.82 ms first, 1.78 ms second (256 MiB)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.58 ms first, 30.52 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
pack vectors: not applicable at R = 512 (checked at R = 0 only)
warm-up batch: 16777216 hashes in 994.87 ms wall
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=512 regs=213 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS

View file

@ -0,0 +1,19 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 8 mm8 per hash = 64 path = PTX mma.sync.m8n8k16.u8
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 30 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache fill (GPU): 1.92 ms first, 1.86 ms second (256 MiB)
cache fill (host, one thread): 349.8 ms, GPU == host all words: PASS
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.61 ms first, 30.55 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
pack vectors: not applicable at R = 8 (checked at R = 0 only)
warm-up batch: 16777216 hashes in 266.01 ms wall
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.75 ms -> 63.078 MH/s
sustain start epoch 1791315790.867
sustain end epoch 1791315815.870: 94 batches in 25.00 s -> 63.074 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_8.txt
SUMMARY R=8 regs=30 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.074 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS

View file

@ -0,0 +1,14 @@
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 8 mm8 per hash = 64 path = reference (shuffles + byte products)
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
CUDA: driver 12.8, runtime 12.8
igneum_hash_info: 77 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache fill (GPU): 1.81 ms first, 1.79 ms second (256 MiB)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset build (GPU, 1024 MiB): 30.59 ms first, 30.52 ms second
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
pack vectors: not applicable at R = 8 (checked at R = 0 only)
warm-up batch: 16777216 hashes in 265.99 ms wall
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=8 regs=77 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,30 @@
15.15 W, 210 MHz, 45, 0 %, 2026/10/06 19:42:24.816
48.43 W, 2520 MHz, 47, 0 %, 2026/10/06 19:42:25.847
106.67 W, 2670 MHz, 51, 0 %, 2026/10/06 19:42:26.847
156.77 W, 2670 MHz, 52, 0 %, 2026/10/06 19:42:27.848
200.24 W, 2670 MHz, 53, 100 %, 2026/10/06 19:42:28.848
200.29 W, 2670 MHz, 53, 100 %, 2026/10/06 19:42:29.848
202.82 W, 2670 MHz, 54, 100 %, 2026/10/06 19:42:30.848
201.27 W, 2670 MHz, 54, 100 %, 2026/10/06 19:42:31.849
201.02 W, 2670 MHz, 54, 100 %, 2026/10/06 19:42:32.849
201.02 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:33.849
201.04 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:34.849
200.96 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:35.850
200.94 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:36.850
200.96 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:37.850
201.00 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:38.850
201.02 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:39.851
201.02 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:40.851
201.12 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:41.851
201.07 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:42.851
201.02 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:43.852
201.03 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:44.852
201.03 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:45.852
201.21 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:46.852
201.31 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:47.853
201.33 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:48.853
201.33 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:49.853
201.52 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:50.853
201.41 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:51.854
201.51 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:52.854
201.49 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:53.854
1 15.15 W 210 MHz 45 0 % 2026/10/06 19:42:24.816
2 48.43 W 2520 MHz 47 0 % 2026/10/06 19:42:25.847
3 106.67 W 2670 MHz 51 0 % 2026/10/06 19:42:26.847
4 156.77 W 2670 MHz 52 0 % 2026/10/06 19:42:27.848
5 200.24 W 2670 MHz 53 100 % 2026/10/06 19:42:28.848
6 200.29 W 2670 MHz 53 100 % 2026/10/06 19:42:29.848
7 202.82 W 2670 MHz 54 100 % 2026/10/06 19:42:30.848
8 201.27 W 2670 MHz 54 100 % 2026/10/06 19:42:31.849
9 201.02 W 2670 MHz 54 100 % 2026/10/06 19:42:32.849
10 201.02 W 2670 MHz 55 100 % 2026/10/06 19:42:33.849
11 201.04 W 2670 MHz 55 100 % 2026/10/06 19:42:34.849
12 200.96 W 2670 MHz 55 100 % 2026/10/06 19:42:35.850
13 200.94 W 2670 MHz 55 100 % 2026/10/06 19:42:36.850
14 200.96 W 2670 MHz 55 100 % 2026/10/06 19:42:37.850
15 201.00 W 2670 MHz 56 100 % 2026/10/06 19:42:38.850
16 201.02 W 2670 MHz 56 100 % 2026/10/06 19:42:39.851
17 201.02 W 2670 MHz 56 100 % 2026/10/06 19:42:40.851
18 201.12 W 2670 MHz 56 100 % 2026/10/06 19:42:41.851
19 201.07 W 2670 MHz 56 100 % 2026/10/06 19:42:42.851
20 201.02 W 2670 MHz 57 100 % 2026/10/06 19:42:43.852
21 201.03 W 2670 MHz 57 100 % 2026/10/06 19:42:44.852
22 201.03 W 2670 MHz 57 100 % 2026/10/06 19:42:45.852
23 201.21 W 2670 MHz 57 100 % 2026/10/06 19:42:46.852
24 201.31 W 2670 MHz 57 100 % 2026/10/06 19:42:47.853
25 201.33 W 2670 MHz 58 100 % 2026/10/06 19:42:48.853
26 201.33 W 2670 MHz 58 100 % 2026/10/06 19:42:49.853
27 201.52 W 2670 MHz 58 100 % 2026/10/06 19:42:50.853
28 201.41 W 2670 MHz 58 100 % 2026/10/06 19:42:51.854
29 201.51 W 2670 MHz 58 100 % 2026/10/06 19:42:52.854
30 201.49 W 2670 MHz 58 100 % 2026/10/06 19:42:53.854

View file

@ -0,0 +1,30 @@
18.57 W, 210 MHz, 53, 0 %, 2026/10/06 19:44:46.034
38.21 W, 2520 MHz, 54, 0 %, 2026/10/06 19:44:47.037
108.02 W, 2670 MHz, 55, 0 %, 2026/10/06 19:44:48.037
143.43 W, 2670 MHz, 59, 100 %, 2026/10/06 19:44:49.037
207.98 W, 2670 MHz, 60, 100 %, 2026/10/06 19:44:50.038
208.85 W, 2670 MHz, 60, 100 %, 2026/10/06 19:44:51.038
209.59 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:52.038
210.00 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:53.038
210.08 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:54.039
210.10 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:55.039
209.96 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:56.039
210.40 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:57.039
210.88 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:58.040
211.38 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:59.040
211.56 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:00.040
211.76 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:01.040
212.68 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:02.040
212.12 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:03.041
211.51 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:04.041
212.43 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:05.041
212.79 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:06.041
212.98 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:07.042
212.89 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:08.042
212.51 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:09.042
212.93 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:10.042
212.91 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:11.042
213.05 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:12.043
213.00 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:13.043
213.23 W, 2670 MHz, 64, 100 %, 2026/10/06 19:45:14.043
213.14 W, 2670 MHz, 64, 100 %, 2026/10/06 19:45:15.043
1 18.57 W 210 MHz 53 0 % 2026/10/06 19:44:46.034
2 38.21 W 2520 MHz 54 0 % 2026/10/06 19:44:47.037
3 108.02 W 2670 MHz 55 0 % 2026/10/06 19:44:48.037
4 143.43 W 2670 MHz 59 100 % 2026/10/06 19:44:49.037
5 207.98 W 2670 MHz 60 100 % 2026/10/06 19:44:50.038
6 208.85 W 2670 MHz 60 100 % 2026/10/06 19:44:51.038
7 209.59 W 2670 MHz 61 100 % 2026/10/06 19:44:52.038
8 210.00 W 2670 MHz 61 100 % 2026/10/06 19:44:53.038
9 210.08 W 2670 MHz 61 100 % 2026/10/06 19:44:54.039
10 210.10 W 2670 MHz 61 100 % 2026/10/06 19:44:55.039
11 209.96 W 2670 MHz 61 100 % 2026/10/06 19:44:56.039
12 210.40 W 2670 MHz 62 100 % 2026/10/06 19:44:57.039
13 210.88 W 2670 MHz 62 100 % 2026/10/06 19:44:58.040
14 211.38 W 2670 MHz 62 100 % 2026/10/06 19:44:59.040
15 211.56 W 2670 MHz 62 100 % 2026/10/06 19:45:00.040
16 211.76 W 2670 MHz 62 100 % 2026/10/06 19:45:01.040
17 212.68 W 2670 MHz 62 100 % 2026/10/06 19:45:02.040
18 212.12 W 2670 MHz 62 100 % 2026/10/06 19:45:03.041
19 211.51 W 2670 MHz 62 100 % 2026/10/06 19:45:04.041
20 212.43 W 2670 MHz 63 100 % 2026/10/06 19:45:05.041
21 212.79 W 2670 MHz 63 100 % 2026/10/06 19:45:06.041
22 212.98 W 2670 MHz 63 100 % 2026/10/06 19:45:07.042
23 212.89 W 2670 MHz 63 100 % 2026/10/06 19:45:08.042
24 212.51 W 2670 MHz 63 100 % 2026/10/06 19:45:09.042
25 212.93 W 2670 MHz 63 100 % 2026/10/06 19:45:10.042
26 212.91 W 2670 MHz 63 100 % 2026/10/06 19:45:11.042
27 213.05 W 2670 MHz 63 100 % 2026/10/06 19:45:12.043
28 213.00 W 2670 MHz 63 100 % 2026/10/06 19:45:13.043
29 213.23 W 2670 MHz 64 100 % 2026/10/06 19:45:14.043
30 213.14 W 2670 MHz 64 100 % 2026/10/06 19:45:15.043

View file

@ -0,0 +1,30 @@
18.08 W, 210 MHz, 53, 0 %, 2026/10/06 19:43:48.466
40.68 W, 2520 MHz, 55, 0 %, 2026/10/06 19:43:49.498
106.86 W, 2670 MHz, 56, 52 %, 2026/10/06 19:43:50.498
141.51 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:51.498
203.49 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:52.499
203.95 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:53.499
204.46 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:54.499
204.84 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:55.499
204.88 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:56.500
205.53 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:57.500
205.87 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:58.500
205.71 W, 2670 MHz, 62, 100 %, 2026/10/06 19:43:59.500
206.17 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:00.500
206.54 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:01.501
206.21 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:02.501
206.09 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:03.504
206.43 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:04.504
206.97 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:05.504
207.32 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:06.504
207.58 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:07.505
207.28 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:08.505
207.92 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:09.505
208.31 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:10.505
208.17 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:11.505
207.89 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:12.506
208.03 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:13.506
208.60 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:14.506
208.41 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:15.506
208.39 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:16.508
208.19 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:17.508
1 18.08 W 210 MHz 53 0 % 2026/10/06 19:43:48.466
2 40.68 W 2520 MHz 55 0 % 2026/10/06 19:43:49.498
3 106.86 W 2670 MHz 56 52 % 2026/10/06 19:43:50.498
4 141.51 W 2670 MHz 60 100 % 2026/10/06 19:43:51.498
5 203.49 W 2670 MHz 60 100 % 2026/10/06 19:43:52.499
6 203.95 W 2670 MHz 60 100 % 2026/10/06 19:43:53.499
7 204.46 W 2670 MHz 61 100 % 2026/10/06 19:43:54.499
8 204.84 W 2670 MHz 61 100 % 2026/10/06 19:43:55.499
9 204.88 W 2670 MHz 61 100 % 2026/10/06 19:43:56.500
10 205.53 W 2670 MHz 61 100 % 2026/10/06 19:43:57.500
11 205.87 W 2670 MHz 61 100 % 2026/10/06 19:43:58.500
12 205.71 W 2670 MHz 62 100 % 2026/10/06 19:43:59.500
13 206.17 W 2670 MHz 62 100 % 2026/10/06 19:44:00.500
14 206.54 W 2670 MHz 62 100 % 2026/10/06 19:44:01.501
15 206.21 W 2670 MHz 61 100 % 2026/10/06 19:44:02.501
16 206.09 W 2670 MHz 62 100 % 2026/10/06 19:44:03.504
17 206.43 W 2670 MHz 62 100 % 2026/10/06 19:44:04.504
18 206.97 W 2670 MHz 62 100 % 2026/10/06 19:44:05.504
19 207.32 W 2670 MHz 62 100 % 2026/10/06 19:44:06.504
20 207.58 W 2670 MHz 62 100 % 2026/10/06 19:44:07.505
21 207.28 W 2670 MHz 63 100 % 2026/10/06 19:44:08.505
22 207.92 W 2670 MHz 63 100 % 2026/10/06 19:44:09.505
23 208.31 W 2670 MHz 63 100 % 2026/10/06 19:44:10.505
24 208.17 W 2670 MHz 63 100 % 2026/10/06 19:44:11.505
25 207.89 W 2670 MHz 63 100 % 2026/10/06 19:44:12.506
26 208.03 W 2670 MHz 63 100 % 2026/10/06 19:44:13.506
27 208.60 W 2670 MHz 63 100 % 2026/10/06 19:44:14.506
28 208.41 W 2670 MHz 63 100 % 2026/10/06 19:44:15.506
29 208.39 W 2670 MHz 63 100 % 2026/10/06 19:44:16.508
30 208.19 W 2670 MHz 63 100 % 2026/10/06 19:44:17.508

View file

@ -0,0 +1,30 @@
15.26 W, 210 MHz, 46, 0 %, 2026/10/06 19:51:22.484
32.16 W, 2520 MHz, 48, 0 %, 2026/10/06 19:51:23.516
84.22 W, 2670 MHz, 50, 100 %, 2026/10/06 19:51:24.517
138.29 W, 2670 MHz, 54, 100 %, 2026/10/06 19:51:25.517
214.64 W, 2670 MHz, 55, 100 %, 2026/10/06 19:51:26.517
215.01 W, 2670 MHz, 55, 100 %, 2026/10/06 19:51:27.517
215.23 W, 2670 MHz, 56, 100 %, 2026/10/06 19:51:28.517
215.43 W, 2670 MHz, 56, 100 %, 2026/10/06 19:51:29.518
215.27 W, 2670 MHz, 56, 100 %, 2026/10/06 19:51:30.518
215.35 W, 2670 MHz, 57, 100 %, 2026/10/06 19:51:31.518
215.48 W, 2670 MHz, 57, 100 %, 2026/10/06 19:51:32.518
215.41 W, 2670 MHz, 57, 100 %, 2026/10/06 19:51:33.518
215.39 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:34.518
215.49 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:35.519
215.67 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:36.519
215.71 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:37.519
215.68 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:38.519
215.70 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:39.519
215.75 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:40.519
215.72 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:41.520
215.80 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:42.520
215.90 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:43.520
215.95 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:44.520
216.10 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:45.520
216.21 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:46.521
216.13 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:47.521
216.06 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:48.521
216.07 W, 2670 MHz, 61, 100 %, 2026/10/06 19:51:49.521
216.16 W, 2670 MHz, 61, 100 %, 2026/10/06 19:51:50.521
216.10 W, 2670 MHz, 61, 100 %, 2026/10/06 19:51:51.521
1 15.26 W 210 MHz 46 0 % 2026/10/06 19:51:22.484
2 32.16 W 2520 MHz 48 0 % 2026/10/06 19:51:23.516
3 84.22 W 2670 MHz 50 100 % 2026/10/06 19:51:24.517
4 138.29 W 2670 MHz 54 100 % 2026/10/06 19:51:25.517
5 214.64 W 2670 MHz 55 100 % 2026/10/06 19:51:26.517
6 215.01 W 2670 MHz 55 100 % 2026/10/06 19:51:27.517
7 215.23 W 2670 MHz 56 100 % 2026/10/06 19:51:28.517
8 215.43 W 2670 MHz 56 100 % 2026/10/06 19:51:29.518
9 215.27 W 2670 MHz 56 100 % 2026/10/06 19:51:30.518
10 215.35 W 2670 MHz 57 100 % 2026/10/06 19:51:31.518
11 215.48 W 2670 MHz 57 100 % 2026/10/06 19:51:32.518
12 215.41 W 2670 MHz 57 100 % 2026/10/06 19:51:33.518
13 215.39 W 2670 MHz 58 100 % 2026/10/06 19:51:34.518
14 215.49 W 2670 MHz 58 100 % 2026/10/06 19:51:35.519
15 215.67 W 2670 MHz 58 100 % 2026/10/06 19:51:36.519
16 215.71 W 2670 MHz 58 100 % 2026/10/06 19:51:37.519
17 215.68 W 2670 MHz 59 100 % 2026/10/06 19:51:38.519
18 215.70 W 2670 MHz 59 100 % 2026/10/06 19:51:39.519
19 215.75 W 2670 MHz 59 100 % 2026/10/06 19:51:40.519
20 215.72 W 2670 MHz 59 100 % 2026/10/06 19:51:41.520
21 215.80 W 2670 MHz 59 100 % 2026/10/06 19:51:42.520
22 215.90 W 2670 MHz 60 100 % 2026/10/06 19:51:43.520
23 215.95 W 2670 MHz 60 100 % 2026/10/06 19:51:44.520
24 216.10 W 2670 MHz 60 100 % 2026/10/06 19:51:45.520
25 216.21 W 2670 MHz 60 100 % 2026/10/06 19:51:46.521
26 216.13 W 2670 MHz 60 100 % 2026/10/06 19:51:47.521
27 216.06 W 2670 MHz 60 100 % 2026/10/06 19:51:48.521
28 216.07 W 2670 MHz 61 100 % 2026/10/06 19:51:49.521
29 216.16 W 2670 MHz 61 100 % 2026/10/06 19:51:50.521
30 216.10 W 2670 MHz 61 100 % 2026/10/06 19:51:51.521

View file

@ -0,0 +1,30 @@
16.99 W, 210 MHz, 51, 0 %, 2026/10/06 19:43:06.035
40.60 W, 2520 MHz, 52, 0 %, 2026/10/06 19:43:07.067
106.97 W, 2670 MHz, 53, 10 %, 2026/10/06 19:43:08.067
137.95 W, 2670 MHz, 57, 100 %, 2026/10/06 19:43:09.068
201.17 W, 2670 MHz, 58, 100 %, 2026/10/06 19:43:10.068
201.47 W, 2670 MHz, 58, 100 %, 2026/10/06 19:43:11.068
201.65 W, 2670 MHz, 58, 100 %, 2026/10/06 19:43:12.068
201.97 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:13.068
202.25 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:14.069
202.06 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:15.069
202.28 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:16.069
202.47 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:17.069
202.62 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:18.069
202.62 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:19.070
203.00 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:20.070
203.19 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:21.070
203.83 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:22.070
203.66 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:23.070
203.32 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:24.070
203.21 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:25.071
203.42 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:26.071
203.90 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:27.071
203.79 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:28.071
203.96 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:29.071
204.32 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:30.072
204.51 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:31.072
204.61 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:32.072
204.72 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:33.072
205.39 W, 2670 MHz, 62, 100 %, 2026/10/06 19:43:34.073
205.90 W, 2670 MHz, 62, 100 %, 2026/10/06 19:43:35.073
1 16.99 W 210 MHz 51 0 % 2026/10/06 19:43:06.035
2 40.60 W 2520 MHz 52 0 % 2026/10/06 19:43:07.067
3 106.97 W 2670 MHz 53 10 % 2026/10/06 19:43:08.067
4 137.95 W 2670 MHz 57 100 % 2026/10/06 19:43:09.068
5 201.17 W 2670 MHz 58 100 % 2026/10/06 19:43:10.068
6 201.47 W 2670 MHz 58 100 % 2026/10/06 19:43:11.068
7 201.65 W 2670 MHz 58 100 % 2026/10/06 19:43:12.068
8 201.97 W 2670 MHz 59 100 % 2026/10/06 19:43:13.068
9 202.25 W 2670 MHz 59 100 % 2026/10/06 19:43:14.069
10 202.06 W 2670 MHz 59 100 % 2026/10/06 19:43:15.069
11 202.28 W 2670 MHz 59 100 % 2026/10/06 19:43:16.069
12 202.47 W 2670 MHz 59 100 % 2026/10/06 19:43:17.069
13 202.62 W 2670 MHz 59 100 % 2026/10/06 19:43:18.069
14 202.62 W 2670 MHz 60 100 % 2026/10/06 19:43:19.070
15 203.00 W 2670 MHz 60 100 % 2026/10/06 19:43:20.070
16 203.19 W 2670 MHz 60 100 % 2026/10/06 19:43:21.070
17 203.83 W 2670 MHz 60 100 % 2026/10/06 19:43:22.070
18 203.66 W 2670 MHz 60 100 % 2026/10/06 19:43:23.070
19 203.32 W 2670 MHz 60 100 % 2026/10/06 19:43:24.070
20 203.21 W 2670 MHz 60 100 % 2026/10/06 19:43:25.071
21 203.42 W 2670 MHz 61 100 % 2026/10/06 19:43:26.071
22 203.90 W 2670 MHz 61 100 % 2026/10/06 19:43:27.071
23 203.79 W 2670 MHz 61 100 % 2026/10/06 19:43:28.071
24 203.96 W 2670 MHz 61 100 % 2026/10/06 19:43:29.071
25 204.32 W 2670 MHz 61 100 % 2026/10/06 19:43:30.072
26 204.51 W 2670 MHz 61 100 % 2026/10/06 19:43:31.072
27 204.61 W 2670 MHz 61 100 % 2026/10/06 19:43:32.072
28 204.72 W 2670 MHz 61 100 % 2026/10/06 19:43:33.072
29 205.39 W 2670 MHz 62 100 % 2026/10/06 19:43:34.073
30 205.90 W 2670 MHz 62 100 % 2026/10/06 19:43:35.073

View file

@ -0,0 +1,5 @@
15.34 W, 210 MHz, 45, 0 %
15.21 W, 210 MHz, 45, 0 %
15.12 W, 210 MHz, 45, 0 %
15.05 W, 210 MHz, 45, 0 %
15.02 W, 210 MHz, 45, 0 %
1 15.34 W 210 MHz 45 0 %
2 15.21 W 210 MHz 45 0 %
3 15.12 W 210 MHz 45 0 %
4 15.05 W 210 MHz 45 0 %
5 15.02 W 210 MHz 45 0 %

View file

@ -0,0 +1,23 @@
ptxas info : 0 bytes gmem
kernel_mm8.cu(98): warning #550-D: variable "lane" was set but never used
uint32_t lane = threadIdx.x & 31u;
^
Remark: The warnings can be suppressed with "-diag-suppress <warning-number>"
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 12.262 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 12.536 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.384 ms

View file

@ -0,0 +1,23 @@
ptxas info : 0 bytes gmem
kernel_mm8.cu(98): warning #550-D: variable "lane" was set but never used
uint32_t lane = threadIdx.x & 31u;
^
Remark: The warnings can be suppressed with "-diag-suppress <warning-number>"
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 12.413 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 12.749 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.256 ms

View file

@ -0,0 +1,17 @@
ptxas info : 0 bytes gmem
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 32 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 27.671 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 11.908 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.275 ms

View file

@ -0,0 +1,17 @@
ptxas info : 0 bytes gmem
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 175 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 9367.163 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 12.495 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.300 ms

View file

@ -0,0 +1,17 @@
ptxas info : 0 bytes gmem
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 15.442 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 12.497 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.272 ms

View file

@ -0,0 +1,17 @@
ptxas info : 0 bytes gmem
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 151 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 520.236 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 12.442 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.444 ms

View file

@ -0,0 +1,17 @@
ptxas info : 0 bytes gmem
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 36 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 83.096 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 11.882 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.282 ms

View file

@ -0,0 +1,17 @@
ptxas info : 0 bytes gmem
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 213 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 273654.000 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 12.613 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.467 ms

View file

@ -0,0 +1,17 @@
ptxas info : 0 bytes gmem
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 30 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 13.112 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 12.461 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.257 ms

View file

@ -0,0 +1,17 @@
ptxas info : 0 bytes gmem
ptxas info : 0 bytes gmem
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 77 registers, used 0 barriers, 376 bytes cmem[0]
ptxas info : Compile time = 78.215 ms
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
ptxas info : Function properties for _Z12igneum_buildPjPKjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
ptxas info : Compile time = 11.877 ms
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
ptxas info : Function properties for _Z17igneum_cache_fillPjj
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
ptxas info : Compile time = 6.236 ms

View file

@ -0,0 +1,7 @@
| R | mm8/hash | MH/s (GPU) | ratio to R=0 | W mean | SM MHz | uJ/hash | max C | fingerprint | PTX == ref | CPU == GPU | verifier ms/unit (R=0, R, delta) | regs | blocks/SM |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 | yes | 1024 of 1024 | 10.179, 10.076, -0.103 | 29 | 24 |
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.926, 9.974, 0.048 | 30 | 24 |
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.075, 10.240, 0.165 | 29 | 24 |
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.183, 11.326, 1.142 | 32 | 24 |
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.887, 14.276, 4.389 | 36 | 24 |

View file

@ -0,0 +1,130 @@
== Tue Oct 6 19:42:12 UTC 2026 host 5c32e87a3fb6 ==
name, driver_version, power.draw [W], clocks.current.sm [MHz], temperature.gpu
NVIDIA GeForce RTX 4090, 570.172.08, 15.34 W, 210 MHz, 45
idle baseline (5 samples):
15.34 W, 210 MHz, 45, 0 %
15.21 W, 210 MHz, 45, 0 %
15.12 W, 210 MHz, 45, 0 %
15.05 W, 210 MHz, 45, 0 %
15.02 W, 210 MHz, 45, 0 %
== build verify_ref (gcc -O2)
== R=0: build bench_0 (PTX path) and bench_0_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
== R=0: timed run with power sampling
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
verify warp base 0 (pack vector, standalone): PASS
verify warp base 4096 (pack vector, standalone): PASS
verify warp base 1000000 (pack vector, standalone): PASS
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
verify warp base 0 (pack vector, in batch): PASS
verify warp base 4096 (pack vector, in batch): PASS
verify warp base 1000000 (pack vector, in batch): PASS
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.74 ms -> 63.078 MH/s
sustain end epoch 1791315774.658: 94 batches in 25.00 s -> 63.075 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_0.txt
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.075 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
OVERALL: PASS
== R=0: reference path fingerprint
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
OVERALL: PASS
FPCHECK R=0 PTX == ref: yes (7c28cfb06c5c65a9)
== R=0: CPU reference on the 32-warp dump (taskset -c 2)
self-test R=0 vector base 0: PASS (0 lanes differ)
self-test R=0 vector base 4096: PASS (0 lanes differ)
self-test R=0 vector base 1000000: PASS (0 lanes differ)
1024 of 1024 lanes equal (PASS) [10.152 ms per unit during the check]
verifier: 10.179 ms per unit at R=0, 10.076 ms per unit at R=0, mm8 block delta -0.103 ms per unit (0.00 us per mm8 step), averaged over 32 units, one core
VERIFY R=0 equal=1024 of=1024 ms_r0=10.179 ms_r=10.076 delta=-0.103
== R=8: build bench_8 (PTX path) and bench_8_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 30 registers, used 0 barriers, 376 bytes cmem[0]
== R=8: timed run with power sampling
igneum_hash_info: 30 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.75 ms -> 63.078 MH/s
sustain end epoch 1791315815.870: 94 batches in 25.00 s -> 63.074 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_8.txt
SUMMARY R=8 regs=30 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.074 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
== R=8: reference path fingerprint
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=8 regs=77 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
FPCHECK R=8 PTX == ref: yes (06fc2593bfb94b4f)
== R=8: CPU reference on the 32-warp dump (taskset -c 2)
1024 of 1024 lanes equal (PASS) [10.106 ms per unit during the check]
verifier: 9.926 ms per unit at R=0, 9.974 ms per unit at R=8, mm8 block delta 0.048 ms per unit (0.74 us per mm8 step), averaged over 32 units, one core
VERIFY R=8 equal=1024 of=1024 ms_r0=9.926 ms_r=9.974 delta=0.048
== R=32: build bench_32 (PTX path) and bench_32_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
== R=32: timed run with power sampling
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.75 ms -> 63.078 MH/s (GPU time), wall 2659.77 ms -> 63.078 MH/s
sustain end epoch 1791315858.296: 94 batches in 25.00 s -> 63.072 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_32.txt
SUMMARY R=32 regs=29 blocksPerSM=24 mhs=63.078 mhs_wall=63.078 mhs_sustain=63.072 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
== R=32: reference path fingerprint
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
FPCHECK R=32 PTX == ref: yes (26e83a65f519c865)
== R=32: CPU reference on the 32-warp dump (taskset -c 2)
1024 of 1024 lanes equal (PASS) [10.432 ms per unit during the check]
verifier: 10.075 ms per unit at R=0, 10.240 ms per unit at R=32, mm8 block delta 0.165 ms per unit (0.64 us per mm8 step), averaged over 32 units, one core
VERIFY R=32 equal=1024 of=1024 ms_r0=10.075 ms_r=10.240 delta=0.165
== R=128: build bench_128 (PTX path) and bench_128_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 32 registers, used 0 barriers, 376 bytes cmem[0]
== R=128: timed run with power sampling
igneum_hash_info: 32 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.82 ms -> 63.077 MH/s (GPU time), wall 2659.83 ms -> 63.076 MH/s
sustain end epoch 1791315915.893: 94 batches in 25.02 s -> 63.034 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_128.txt
SUMMARY R=128 regs=32 blocksPerSM=24 mhs=63.077 mhs_wall=63.076 mhs_sustain=63.034 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
== R=128: reference path fingerprint
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=128 regs=175 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
FPCHECK R=128 PTX == ref: yes (42223c2113188335)
== R=128: CPU reference on the 32-warp dump (taskset -c 2)
1024 of 1024 lanes equal (PASS) [11.349 ms per unit during the check]
verifier: 10.183 ms per unit at R=0, 11.326 ms per unit at R=128, mm8 block delta 1.142 ms per unit (1.12 us per mm8 step), averaged over 32 units, one core
VERIFY R=128 equal=1024 of=1024 ms_r0=10.183 ms_r=11.326 delta=1.142
== R=512: build bench_512 (PTX path) and bench_512_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 36 registers, used 0 barriers, 376 bytes cmem[0]
== R=512: timed run with power sampling
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_512.txt
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
== R=512: reference path fingerprint
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=512 regs=213 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
FPCHECK R=512 PTX == ref: yes (02b7002d747f3711)
== R=512: CPU reference on the 32-warp dump (taskset -c 2)
1024 of 1024 lanes equal (PASS) [14.440 ms per unit during the check]
verifier: 9.887 ms per unit at R=0, 14.276 ms per unit at R=512, mm8 block delta 4.389 ms per unit (1.07 us per mm8 step), averaged over 32 units, one core
VERIFY R=512 equal=1024 of=1024 ms_r0=9.887 ms_r=14.276 delta=4.389
== Tue Oct 6 19:51:57 UTC 2026 done

View file

@ -0,0 +1,137 @@
== Tue Oct 6 19:42:12 UTC 2026 host 5c32e87a3fb6 ==
name, driver_version, power.draw [W], clocks.current.sm [MHz], temperature.gpu
NVIDIA GeForce RTX 4090, 570.172.08, 15.34 W, 210 MHz, 45
idle baseline (5 samples):
15.34 W, 210 MHz, 45, 0 %
15.21 W, 210 MHz, 45, 0 %
15.12 W, 210 MHz, 45, 0 %
15.05 W, 210 MHz, 45, 0 %
15.02 W, 210 MHz, 45, 0 %
== build verify_ref (gcc -O2)
== R=0: build bench_0 (PTX path) and bench_0_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
== R=0: timed run with power sampling
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
verify warp base 0 (pack vector, standalone): PASS
verify warp base 4096 (pack vector, standalone): PASS
verify warp base 1000000 (pack vector, standalone): PASS
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
verify warp base 0 (pack vector, in batch): PASS
verify warp base 4096 (pack vector, in batch): PASS
verify warp base 1000000 (pack vector, in batch): PASS
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.74 ms -> 63.078 MH/s
sustain end epoch 1791315774.658: 94 batches in 25.00 s -> 63.075 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_0.txt
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.075 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
OVERALL: PASS
== R=0: reference path fingerprint
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
OVERALL: PASS
FPCHECK R=0 PTX == ref: yes (7c28cfb06c5c65a9)
== R=0: CPU reference on the 32-warp dump (taskset -c 2)
self-test R=0 vector base 0: PASS (0 lanes differ)
self-test R=0 vector base 4096: PASS (0 lanes differ)
self-test R=0 vector base 1000000: PASS (0 lanes differ)
1024 of 1024 lanes equal (PASS) [10.152 ms per unit during the check]
verifier: 10.179 ms per unit at R=0, 10.076 ms per unit at R=0, mm8 block delta -0.103 ms per unit (0.00 us per mm8 step), averaged over 32 units, one core
VERIFY R=0 equal=1024 of=1024 ms_r0=10.179 ms_r=10.076 delta=-0.103
== R=8: build bench_8 (PTX path) and bench_8_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 30 registers, used 0 barriers, 376 bytes cmem[0]
== R=8: timed run with power sampling
igneum_hash_info: 30 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.75 ms -> 63.078 MH/s
sustain end epoch 1791315815.870: 94 batches in 25.00 s -> 63.074 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_8.txt
SUMMARY R=8 regs=30 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.074 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
== R=8: reference path fingerprint
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=8 regs=77 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
FPCHECK R=8 PTX == ref: yes (06fc2593bfb94b4f)
== R=8: CPU reference on the 32-warp dump (taskset -c 2)
1024 of 1024 lanes equal (PASS) [10.106 ms per unit during the check]
verifier: 9.926 ms per unit at R=0, 9.974 ms per unit at R=8, mm8 block delta 0.048 ms per unit (0.74 us per mm8 step), averaged over 32 units, one core
VERIFY R=8 equal=1024 of=1024 ms_r0=9.926 ms_r=9.974 delta=0.048
== R=32: build bench_32 (PTX path) and bench_32_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
== R=32: timed run with power sampling
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.75 ms -> 63.078 MH/s (GPU time), wall 2659.77 ms -> 63.078 MH/s
sustain end epoch 1791315858.296: 94 batches in 25.00 s -> 63.072 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_32.txt
SUMMARY R=32 regs=29 blocksPerSM=24 mhs=63.078 mhs_wall=63.078 mhs_sustain=63.072 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
== R=32: reference path fingerprint
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
FPCHECK R=32 PTX == ref: yes (26e83a65f519c865)
== R=32: CPU reference on the 32-warp dump (taskset -c 2)
1024 of 1024 lanes equal (PASS) [10.432 ms per unit during the check]
verifier: 10.075 ms per unit at R=0, 10.240 ms per unit at R=32, mm8 block delta 0.165 ms per unit (0.64 us per mm8 step), averaged over 32 units, one core
VERIFY R=32 equal=1024 of=1024 ms_r0=10.075 ms_r=10.240 delta=0.165
== R=128: build bench_128 (PTX path) and bench_128_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 32 registers, used 0 barriers, 376 bytes cmem[0]
== R=128: timed run with power sampling
igneum_hash_info: 32 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.82 ms -> 63.077 MH/s (GPU time), wall 2659.83 ms -> 63.076 MH/s
sustain end epoch 1791315915.893: 94 batches in 25.02 s -> 63.034 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_128.txt
SUMMARY R=128 regs=32 blocksPerSM=24 mhs=63.077 mhs_wall=63.076 mhs_sustain=63.034 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
== R=128: reference path fingerprint
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=128 regs=175 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
FPCHECK R=128 PTX == ref: yes (42223c2113188335)
== R=128: CPU reference on the 32-warp dump (taskset -c 2)
1024 of 1024 lanes equal (PASS) [11.349 ms per unit during the check]
verifier: 10.183 ms per unit at R=0, 11.326 ms per unit at R=128, mm8 block delta 1.142 ms per unit (1.12 us per mm8 step), averaged over 32 units, one core
VERIFY R=128 equal=1024 of=1024 ms_r0=10.183 ms_r=11.326 delta=1.142
== R=512: build bench_512 (PTX path) and bench_512_ref
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
ptxas info : Used 36 registers, used 0 barriers, 376 bytes cmem[0]
== R=512: timed run with power sampling
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
dump: 32 warps (1024 lanes) written to out/dump_512.txt
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
== R=512: reference path fingerprint
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
SUMMARY R=512 regs=213 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
OVERALL: PASS
FPCHECK R=512 PTX == ref: yes (02b7002d747f3711)
== R=512: CPU reference on the 32-warp dump (taskset -c 2)
1024 of 1024 lanes equal (PASS) [14.440 ms per unit during the check]
verifier: 9.887 ms per unit at R=0, 14.276 ms per unit at R=512, mm8 block delta 4.389 ms per unit (1.07 us per mm8 step), averaged over 32 units, one core
VERIFY R=512 equal=1024 of=1024 ms_r0=9.887 ms_r=14.276 delta=4.389
== Tue Oct 6 19:51:57 UTC 2026 done
| R | mm8/hash | MH/s (GPU) | ratio to R=0 | W mean | SM MHz | uJ/hash | max C | fingerprint | PTX == ref | CPU == GPU | verifier ms/unit (R=0, R, delta) | regs | blocks/SM |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 | yes | 1024 of 1024 | 10.179, 10.076, -0.103 | 29 | 24 |
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.926, 9.974, 0.048 | 30 | 24 |
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.075, 10.240, 0.165 | 29 | 24 |
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.183, 11.326, 1.142 | 32 | 24 |
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.887, 14.276, 4.389 | 36 | 24 |

View file

@ -0,0 +1,8 @@
cache: host fill 528.5 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
self-test R=0 vector base 0: PASS (0 lanes differ)
self-test R=0 vector base 4096: PASS (0 lanes differ)
self-test R=0 vector base 1000000: PASS (0 lanes differ)
dump: 1024 lines, 32 units, R = 0 (0 mm8 per hash)
1024 of 1024 lanes equal (PASS) [10.152 ms per unit during the check]
verifier: 10.179 ms per unit at R=0, 10.076 ms per unit at R=0, mm8 block delta -0.103 ms per unit (0.00 us per mm8 step), averaged over 32 units, one core
VERIFY R=0 equal=1024 of=1024 ms_r0=10.179 ms_r=10.076 delta=-0.103

View file

@ -0,0 +1,5 @@
cache: host fill 533.2 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
dump: 1024 lines, 32 units, R = 128 (1024 mm8 per hash)
1024 of 1024 lanes equal (PASS) [11.349 ms per unit during the check]
verifier: 10.183 ms per unit at R=0, 11.326 ms per unit at R=128, mm8 block delta 1.142 ms per unit (1.12 us per mm8 step), averaged over 32 units, one core
VERIFY R=128 equal=1024 of=1024 ms_r0=10.183 ms_r=11.326 delta=1.142

View file

@ -0,0 +1,5 @@
cache: host fill 535.4 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
dump: 1024 lines, 32 units, R = 32 (256 mm8 per hash)
1024 of 1024 lanes equal (PASS) [10.432 ms per unit during the check]
verifier: 10.075 ms per unit at R=0, 10.240 ms per unit at R=32, mm8 block delta 0.165 ms per unit (0.64 us per mm8 step), averaged over 32 units, one core
VERIFY R=32 equal=1024 of=1024 ms_r0=10.075 ms_r=10.240 delta=0.165

View file

@ -0,0 +1,5 @@
cache: host fill 530.6 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
dump: 1024 lines, 32 units, R = 512 (4096 mm8 per hash)
1024 of 1024 lanes equal (PASS) [14.440 ms per unit during the check]
verifier: 9.887 ms per unit at R=0, 14.276 ms per unit at R=512, mm8 block delta 4.389 ms per unit (1.07 us per mm8 step), averaged over 32 units, one core
VERIFY R=512 equal=1024 of=1024 ms_r0=9.887 ms_r=14.276 delta=4.389

View file

@ -0,0 +1,5 @@
cache: host fill 558.2 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
dump: 1024 lines, 32 units, R = 8 (64 mm8 per hash)
1024 of 1024 lanes equal (PASS) [10.106 ms per unit during the check]
verifier: 9.926 ms per unit at R=0, 9.974 ms per unit at R=8, mm8 block delta 0.048 ms per unit (0.74 us per mm8 step), averaged over 32 units, one core
VERIFY R=8 equal=1024 of=1024 ms_r0=9.926 ms_r=9.974 delta=0.048

View file

@ -0,0 +1,66 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#ifndef IGNEUM_NO_CUDA
#include <cuda_runtime.h>
#endif
#define IGNEUM_SEED_STRING "igneum-genesis"
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
#define IGNEUM_GENERATOR 3
#define IGNEUM_PROGRAM_ATTEMPT 0
#define IGNEUM_PROGRAM_ID 0xe323b9dcaf283a6full
#define IGNEUM_DAY_STRING "2026-10-03"
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
#define IGNEUM_DAY0 0x3067619fu
#define IGNEUM_DAY1 0x3c269176u
#define IGNEUM_DATASET_LOG2 28
#define IGNEUM_MASK 0x0fffffffu
#define IGNEUM_LANES 32
#define IGNEUM_ITERATIONS 8
#define IGNEUM_INSTR_COUNT 64
#define IGNEUM_LOADS_PER_HASH 128
#define IGNEUM_WIDE_LOADS_PER_HASH 0
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
#define IGNEUM_PROGRAM_CLASS "v3"
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
#define IGNEUM_LOAD_CLASS "mx8"
#define IGNEUM_CLASS_MIXER_MULT 8
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
#define IGNEUM_LOAD_SLOTS 16
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
#define IGNEUM_BYTES_PER_HASH 512
#define IGNEUM_FOLD_ROT 11
#define IGNEUM_FOLD_MUL 0x9e3779b1u
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
#define IGNEUM_DATASET_MODE 1
#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u }
#define IGNEUM_CACHE_LOG2_WORDS 26
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
#define IGNEUM_CACHE_SEGMENTS 65536u
#define IGNEUM_ITEM_ROUNDS 8
#define IGNEUM_MIXER_MULT 8 // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)
#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u }
#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u }
#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u }
#ifndef IGNEUM_NO_CUDA
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif

View file

@ -0,0 +1,74 @@
// Generated by gen_ref_program.py from the pack's kernel.cu. Do not edit by hand.
// Expects: uint32_t R[8][32], T[32], SEL[32]; const uint32_t* cache; uint32_t mask; macro L = for (int l = 0; l < 32; ++l).
L { R[2][l] = R[3][l] * R[4][l] + R[2][l]; } // 0 mad
L { R[2][l] = R[1][l] * R[1][l] + R[2][l]; } // 1 mad
L { R[2][l] = R[3][l] * R[2][l] + R[2][l]; } // 2 mad
L { R[3][l] = R[3][l] ^ R[5][l]; } // 3 xor
L { R[7][l] = R[7][l] ^ mh_word(cache, R[2][l] & mask); } // 4 load
L { R[5][l] = R[5][l] ^ mh_word(cache, R[7][l] & mask); } // 5 load
memcpy(T, R[4], sizeof T);
L { R[1][l] = R[1][l] ^ T[l ^ 8]; } // 6 shfl
memcpy(T, R[3], sizeof T);
L { R[7][l] = R[7][l] ^ T[l ^ 8]; } // 7 shfl
L { R[1][l] = umulhi(R[1][l], R[5][l]); } // 8 mulhi
L { R[6][l] = rotr_var(R[6][l], R[3][l]); } // 9 rotr
L { R[3][l] = R[3][l] | R[4][l]; } // 10 or
L { R[4][l] = R[4][l] ^ mh_word(cache, R[3][l] & mask); } // 11 load
L { R[0][l] = umulhi(R[0][l], R[4][l]); } // 12 mulhi
L { R[5][l] = R[5][l] + R[1][l] + ((((SEL[l] >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); } // 13 add
L { R[0][l] = R[0][l] ^ mh_word(cache, R[4][l] & mask); } // 14 load
L { R[2][l] = R[2][l] - R[4][l]; } // 15 sub
L { R[2][l] = R[2][l] ^ mh_word(cache, R[0][l] & mask); } // 16 load
L { R[7][l] = R[7][l] ^ mh_word(cache, R[2][l] & mask); } // 17 load
memcpy(T, R[3], sizeof T);
L { R[7][l] = R[7][l] ^ T[l ^ 4]; } // 18 shfl
L { R[5][l] = R[5][l] * R[0][l]; } // 19 mul
memcpy(T, R[4], sizeof T);
L { R[3][l] = R[3][l] ^ T[l ^ 2]; } // 20 shfl
memcpy(T, R[4], sizeof T);
L { R[2][l] = R[2][l] ^ T[l ^ 16]; } // 21 shfl
L { R[6][l] = umulhi(R[6][l], R[2][l]); } // 22 mulhi
L { R[6][l] = R[6][l] ^ mh_word(cache, R[1][l] & mask); } // 23 load
L { R[5][l] = R[5][l] * R[0][l]; } // 24 mul
L { R[5][l] = rotl_imm(R[5][l], 19u); } // 25 rotl
memcpy(T, R[6], sizeof T);
L { R[7][l] = R[7][l] ^ T[l ^ 2]; } // 26 shfl
L { R[0][l] = R[0][l] ^ R[5][l]; } // 27 xor
L { R[0][l] = R[0][l] ^ R[4][l]; } // 28 xor
L { R[3][l] = R[3][l] - R[0][l]; } // 29 sub
L { R[5][l] = R[5][l] * R[1][l]; } // 30 mul
L { R[7][l] = R[7][l] ^ mh_word(cache, R[2][l] & mask); } // 31 load
L { R[1][l] = R[1][l] ^ mh_word(cache, R[0][l] & mask); } // 32 load
L { R[5][l] = R[5][l] ^ R[6][l]; } // 33 xor
L { R[5][l] = R[5][l] ^ mh_word(cache, R[1][l] & mask); } // 34 load
L { R[0][l] = umulhi(R[0][l], R[5][l]); } // 35 mulhi
memcpy(T, R[2], sizeof T);
L { R[5][l] = R[5][l] ^ T[l ^ 4]; } // 36 shfl
L { R[7][l] = R[7][l] ^ mh_word(cache, R[0][l] & mask); } // 37 load
L { R[3][l] = R[3][l] + R[1][l] + ((((SEL[l] >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); } // 38 add
memcpy(T, R[5], sizeof T);
L { R[1][l] = R[1][l] ^ T[l ^ 4]; } // 39 shfl
L { R[2][l] = R[2][l] ^ R[5][l]; } // 40 xor
L { R[3][l] = R[6][l] * R[3][l] + R[3][l]; } // 41 mad
L { R[6][l] = R[6][l] - R[7][l]; } // 42 sub
L { R[7][l] = R[7][l] ^ R[0][l]; } // 43 xor
L { R[1][l] = R[1][l] ^ mh_word(cache, R[7][l] & mask); } // 44 load
L { R[2][l] = R[2][l] * R[3][l]; } // 45 mul
L { R[1][l] = umulhi(R[1][l], R[5][l]); } // 46 mulhi
L { R[4][l] = R[4][l] - R[3][l]; } // 47 sub
L { R[2][l] = rotr_var(R[2][l], R[6][l]); } // 48 rotr
L { R[3][l] = R[3][l] ^ mh_word(cache, R[5][l] & mask); } // 49 load
L { R[1][l] = R[1][l] + R[5][l] + ((((SEL[l] >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); } // 50 add
L { R[0][l] = R[0][l] * R[2][l]; } // 51 mul
L { R[0][l] = R[0][l] + R[2][l] + ((((SEL[l] >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); } // 52 add
L { R[1][l] = R[1][l] + R[0][l] + ((((SEL[l] >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); } // 53 add
L { R[7][l] = rotl_imm(R[7][l], 14u); } // 54 rotl
L { R[3][l] = R[3][l] + R[7][l] + ((((SEL[l] >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); } // 55 add
L { R[6][l] = R[6][l] ^ mh_word(cache, R[7][l] & mask); } // 56 load
L { R[1][l] = rotr_var(R[1][l], R[5][l]); } // 57 rotr
L { R[5][l] = R[5][l] ^ mh_word(cache, R[4][l] & mask); } // 58 load
L { R[6][l] = R[6][l] ^ mh_word(cache, R[2][l] & mask); } // 59 load
L { R[3][l] = R[5][l] * R[0][l] + R[3][l]; } // 60 mad
L { R[5][l] = R[5][l] + R[7][l] + ((((SEL[l] >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); } // 61 add
L { R[4][l] = R[4][l] + R[6][l] + ((((SEL[l] >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); } // 62 add
L { R[5][l] = rotl_imm(R[5][l], 19u); } // 63 rotl

43
proto-newpow/mma-shadow/run.sh Executable file
View file

@ -0,0 +1,43 @@
#!/bin/bash
# run.sh (proto-newpow/mma-shadow): build and measure the ladder R in {0, 8, 32, 128, 512} on the GPU box.
# Runs under /root/horizon-newpow/mma-shadow. Every artefact lands in ./out/.
set -u
export PATH=/usr/local/cuda/bin:$PATH
cd "$(dirname "$0")"
LADDER="${LADDER:-0 8 32 128 512}"
SUSTAIN="${SUSTAIN:-25}" # seconds of sustained hashing for the power meter (mean taken after the first 10 s)
mkdir -p out
echo "== $(date -u) host $(hostname) ==" | tee out/run.log
nvidia-smi --query-gpu=name,driver_version,power.draw,clocks.sm,temperature.gpu --format=csv | tee -a out/run.log
echo "idle baseline (5 samples):" | tee -a out/run.log
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu --format=csv,noheader -l 1 > out/power_idle.csv & P=$!; sleep 5; kill $P 2>/dev/null; wait $P 2>/dev/null; cat out/power_idle.csv | tee -a out/run.log
echo "== build verify_ref (gcc -O2)" | tee -a out/run.log
gcc -O2 -o verify_ref verify_ref.c 2>&1 | tee -a out/run.log
for R in $LADDER; do
echo "== R=$R: build bench_$R (PTX path) and bench_${R}_ref" | tee -a out/run.log
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -Xptxas -v -o bench_$R bench.cu kernel_mm8.cu 2> out/ptxas_$R.txt || { echo "BUILD FAILED bench_$R"; cat out/ptxas_$R.txt; continue; }
grep -A2 "igneum_hash" out/ptxas_$R.txt | grep -i "registers\|spill" | tee -a out/run.log
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -DIGNEUM_MM8_REF -Xptxas -v -o bench_${R}_ref bench.cu kernel_mm8.cu 2> out/ptxas_${R}_ref.txt || { echo "BUILD FAILED bench_${R}_ref"; cat out/ptxas_${R}_ref.txt; }
echo "== R=$R: timed run with power sampling" | tee -a out/run.log
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu,timestamp --format=csv,noheader -l 1 > out/power_$R.csv &
SMI=$!
./bench_$R --batches 10 --sustain $SUSTAIN --dump out/dump_$R.txt 32 > out/bench_$R.log 2>&1
kill $SMI 2>/dev/null; wait $SMI 2>/dev/null
grep -E "igneum_hash_info|cache check|dataset self-test|verify warp|fingerprint|timed:|sustain end|SUMMARY|OVERALL|dump:" out/bench_$R.log | tee -a out/run.log
echo "== R=$R: reference path fingerprint" | tee -a out/run.log
./bench_${R}_ref --fingerprint-only --no-host-cache > out/bench_${R}_ref.log 2>&1
grep -E "fingerprint|OVERALL" out/bench_${R}_ref.log | tee -a out/run.log
FP=$(grep -oE "fingerprint: [0-9a-f]{16}" out/bench_$R.log | cut -d' ' -f2)
FPR=$(grep -oE "fingerprint: [0-9a-f]{16}" out/bench_${R}_ref.log | cut -d' ' -f2)
if [ -n "$FP" ] && [ "$FP" = "$FPR" ]; then echo "FPCHECK R=$R PTX == ref: yes ($FP)" | tee -a out/run.log; else echo "FPCHECK R=$R PTX == ref: NO (ptx $FP ref $FPR)" | tee -a out/run.log; fi
echo "== R=$R: CPU reference on the 32-warp dump (taskset -c 2)" | tee -a out/run.log
taskset -c 2 ./verify_ref out/dump_$R.txt $R --time $( [ "$R" = "0" ] && echo --self-test ) > out/verify_$R.log 2>&1
grep -E "self-test|lanes equal|first mismatch|verifier:|VERIFY" out/verify_$R.log | tee -a out/run.log
done
echo "== $(date -u) done" | tee -a out/run.log
python3 summarise.py | tee out/results.md

View file

@ -0,0 +1,45 @@
#!/usr/bin/env python3
# summarise.py: builds the RESULTS table from out/*.log and out/power_*.csv (run by run.sh on the box).
import re, glob, os
from datetime import datetime
rows = []
def grab(pat, text, default=None):
m = re.search(pat, text); return m.group(1) if m else default
def power_mean(R, log):
try:
s0 = float(grab(r"sustain start epoch ([0-9.]+)", log)); s1 = float(grab(r"sustain end epoch ([0-9.]+)", log))
except Exception:
return None, None, None
pw, clk, tmp = [], [], []
for line in open("out/power_%s.csv" % R):
parts = [p.strip() for p in line.split(",")]
if len(parts) < 5: continue
try:
ts = datetime.strptime(parts[4], "%Y/%m/%d %H:%M:%S.%f").timestamp()
except Exception:
continue
if ts >= s0 + 10 and ts <= s1:
pw.append(float(parts[0].split()[0])); clk.append(float(parts[1].split()[0])); tmp.append(float(parts[2].split()[0]))
if not pw: return None, None, None
return sum(pw) / len(pw), sum(clk) / len(clk), max(tmp)
base_mhs = None
out = ["| R | mm8/hash | MH/s (GPU) | ratio to R=0 | W mean | SM MHz | uJ/hash | max C | fingerprint | PTX == ref | CPU == GPU | verifier ms/unit (R=0, R, delta) | regs | blocks/SM |", "|---|---|---|---|---|---|---|---|---|---|---|---|---|---|"]
for Rs in sorted([int(os.path.basename(p)[6:-4]) for p in glob.glob("out/bench_*.log") if "_ref" not in p]):
log = open("out/bench_%d.log" % Rs).read()
summ = grab(r"(SUMMARY.*)", log, "")
mhs = float(grab(r"mhs=([0-9.]+)", summ, "0")); regs = grab(r"regs=(\d+)", summ, "?"); bps = grab(r"blocksPerSM=(\d+)", summ, "?")
fp = grab(r"fingerprint=([0-9a-f]{16})", summ, "?")
if Rs == 0: base_mhs = mhs
ratio = ("%.3f" % (mhs / base_mhs)) if base_mhs else "?"
runlog = open("out/run.log").read()
fpc = grab(r"FPCHECK R=%d PTX == ref: (\S+)" % Rs, runlog, "?")
ver = open("out/verify_%d.log" % Rs).read() if os.path.exists("out/verify_%d.log" % Rs) else ""
vs = grab(r"(VERIFY.*)", ver, "")
eq = grab(r"equal=(\d+)", vs, "?"); of = grab(r"of=(\d+)", vs, "?")
ms0 = grab(r"ms_r0=([0-9.]+)", vs, "?"); msr = grab(r"ms_r=([0-9.]+)", vs, "?"); dl = grab(r"delta=([0-9.-]+)", vs, "?")
w, clk, tmx = power_mean(Rs, log)
uj = ("%.2f" % (w / (mhs * 1e6) * 1e6)) if (w and mhs) else "?"
out.append("| %d | %d | %.2f | %s | %s | %s | %s | %s | %s | %s | %s of %s | %s, %s, %s | %s | %s |" % (
Rs, 8 * Rs, mhs, ratio, ("%.1f" % w) if w else "?", ("%.0f" % clk) if clk else "?", uj, ("%.0f" % tmx) if tmx else "?",
fp, fpc, eq, of, ms0, msr, dl, regs, bps))
print("\n".join(out))

View file

@ -0,0 +1,57 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#define IGNEUM_VEC_WARPS 3
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
{ // base nonce 0
0x19b56348bc85304dull, 0xb08a9cfb44aa720full, 0xe1f8f627f780eff7ull, 0x2ff3e86ff1696161ull, 0xb65e0578c257e8acull, 0xf37b5705c2bebaaeull, 0xb19fef670982389eull, 0x2374331b28827a11ull,
0xf492cdd05dda7f88ull, 0x37700f1385e19885ull, 0xa27524f3010b3e87ull, 0x6c8a24de8b938c43ull, 0x3dc017c8820cfd16ull, 0x5c80d37146657b68ull, 0x214082cac03f734aull, 0x1d67c665145d72f3ull,
0x9fb0ae736eb87ff5ull, 0xcd7d416ae3a6be61ull, 0x0069cda85c6de58full, 0x0b7d82e717224baaull, 0x600960f6d7be30f0ull, 0xeb29032663b0c4d2ull, 0xf29583a5766408f1ull, 0x6c9532a99ad7314dull,
0x9a949fd0959cc80full, 0xa20258520e7f6c25ull, 0xf2ab11f9bb032e38ull, 0xcc967bcd0c8d07c1ull, 0x37745267bb3231f2ull, 0x35a046048c2b69b3ull, 0xaa51834cd3f364f3ull, 0x359192708e4f754aull
},
{ // base nonce 4096
0x62fb132a9943127aull, 0x0b703e577e7f4ecaull, 0xf9f24f5522ce7593ull, 0x3cf5c516abc4332aull, 0xd25523f5f6d7a127ull, 0xd2081a002f983682ull, 0xbf46c54e9b3c4254ull, 0xca362e291e5e5f4dull,
0x6039712f10f457a3ull, 0x8a34b7cabf97c23bull, 0xa473c6a2e0bf59bcull, 0x6cf3926513a4b069ull, 0x297ec2998376a40dull, 0x8efd7f601a8f28dbull, 0x8e72532dfdc1e544ull, 0x917c2b2ebe2a7e00ull,
0x923fbb2d2f635c25ull, 0xce864ea5c0dedad9ull, 0x4b8ec7e874e446efull, 0x1b69b69465449196ull, 0x5ef3a8a6edb369cfull, 0x06c263ef9ce63fc4ull, 0x9c2048fd9d9e2639ull, 0x457fdd96ca4a138eull,
0xd1904018b8d7b6e3ull, 0x8682312fb2e96ab8ull, 0xdc3257e0d0f979a5ull, 0xa51b0a8519d87db5ull, 0x334f08ec056e618bull, 0x3464ce71dc65119dull, 0x6a4d6df066332e04ull, 0x7d7866cb9cfca8ffull
},
{ // base nonce 1000000
0x86b6cb0e13d89b03ull, 0x96299a3f19d7ef15ull, 0x67d2c55100d2f876ull, 0x0a4dfe97d671b728ull, 0x41e4489014d42595ull, 0xf11cb1958c0c0e82ull, 0xf8b70b0c0a03175full, 0x632299df87d5063eull,
0xe198417776130492ull, 0x8ffc5449290d7be2ull, 0x5f2e264eb1311f1bull, 0x988376463ac88586ull, 0x83969eadda489c26ull, 0xbed0a2c3f255d306ull, 0x1a949d271961a819ull, 0x5bce06eb6984725cull,
0x94d5d6a1b0ffd4e9ull, 0xf3c78bae6c2182b4ull, 0xb97e9fe1bbfcdd55ull, 0x70262d1d4c0eccb2ull, 0x1fc93b427dba28d9ull, 0x02b2e3c4317f2a2dull, 0x54d3d42a588edcb9ull, 0x79998677846e7cceull,
0x486522a5425f821aull, 0x95fa88e933360e52ull, 0xc8bae2da2b883f6cull, 0xbe3eb610ad33614full, 0x20efb3c4de82907full, 0xd6b650cfedfb26b7ull, 0x8c24447a646dba26ull, 0x9c004678515e44ecull
}
};
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
static const uint32_t IGNEUM_DS_HEAD[16] = {
0xfdad4319u, 0x1a7b68e1u, 0xde6db608u, 0x13d73892u, 0xd17f447au, 0xb2221ccfu, 0x9db004bdu, 0x57d7d367u,
0xdbc4cf34u, 0x697c009au, 0xc43af1d4u, 0x97f12b2eu, 0x74c37cd0u, 0xc651ea15u, 0x665a6d29u, 0x22330a2du
};
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
static const uint32_t IGNEUM_DS_LAST = 0xa83e7aa6u;
// 64 sampled dataset words (index, value) computed on the Mac.
#define IGNEUM_DS_SAMPLES 64
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
};
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
0x3230bc7bu, 0x7fbfe2c9u, 0xb2690991u, 0x1745c7c5u, 0x0ab0ccafu, 0x1bf87d6bu, 0x160139fdu, 0x719817acu, 0x0155df4bu, 0xbe1e86c3u, 0x680bcd6cu, 0x79c3dc6cu, 0x181e7e5fu, 0x0713a109u, 0xc705dd9fu, 0x3933b7a8u, 0xdd1c0431u, 0x50522b30u, 0xa0020b38u, 0xbff39e96u, 0x21b67e18u, 0x740f8db3u, 0x2baba568u, 0x2c9bef83u, 0x0ad9b671u, 0xc4327869u, 0x7b4fd7d0u, 0x2c29965fu, 0xec56f15fu, 0x61111746u, 0x303a1d6eu, 0xbddcfd1au, 0xf829a355u, 0x6d5df2a9u, 0x01ab8e44u, 0x06d13507u, 0xda8dcfc6u, 0x01a703e1u, 0xafe7d2c1u, 0xc091c3a2u, 0xac1814feu, 0x6e6ff62au, 0x8fdf01bau, 0xdd3f7159u, 0xdfa0d75cu, 0x26684c35u, 0x7f441e63u, 0x88df2570u, 0x8aa4d5ebu, 0xcc816c05u, 0x434df890u, 0xcd392ad6u, 0x1ab4cb63u, 0x595926fau, 0x7cd76b41u, 0x20cb95c4u, 0x13cf823fu, 0xf9daf901u, 0xff9af40au, 0x2c7dfa51u, 0x871206dbu, 0x938c116cu, 0xb64bf199u, 0x5751f874u
};
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u,
0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u
};
static const uint32_t IGNEUM_CACHE_LAST[16] = {
0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du,
0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu
};
static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull;

View file

@ -0,0 +1,152 @@
// verify_ref.c (proto-newpow/mma-shadow): plain C CPU reference for the mx8+mm8xR prototype.
// Reads a bench --dump file (lines: base lane value_hex, 32 lanes per warp), recomputes every warp with a
// register-major 32-lane interpreter of the pack's program (ref_program.inc, generated from kernel.cu), lazy
// dataset words through memhard.h's mh_word over a host-filled 256 MiB cache, the mm8 block in the spec layout,
// then the fold. Prints "N of 1024 lanes equal" and the first mismatch, then times the verifier per 32-lane unit.
// Build: gcc -O2 -o verify_ref verify_ref.c Run: taskset -c 2 ./verify_ref dump.txt <R> [--time]
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
#define IGNEUM_NO_CUDA
#include "program.h"
#include "vectors.h"
#include "memhard.h"
#include "mm8_block.h"
static const uint16_t MM8_TABLE[IGNEUM_MM8_R_MAX] = IGNEUM_MM8_TABLE_INIT;
static uint32_t splitmix32(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }
static uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
static uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
static uint32_t umulhi(uint32_t a, uint32_t b) { return (uint32_t)(((uint64_t)a * (uint64_t)b) >> 32); }
static double nowMs(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1e3 + ts.tv_nsec / 1e6; }
#define L for (int l = 0; l < 32; ++l)
// One mm8 step on the 32-lane register file (spec layout, identical to family-probe.cu's mm8 warp_ref):
// r[c] += C[l >> 2][2 * (l & 3)] (d0), r[c2] += C[l >> 2][2 * (l & 3) + 1] (d1).
static void mm8_step(uint32_t R[8][32], int a, int b, int c, int c2) {
uint32_t A[32], B[32];
memcpy(A, R[a], sizeof A);
memcpy(B, R[b], sizeof B);
L {
int row = l >> 2, col0 = 2 * (l & 3);
uint32_t acc0 = 0u, acc1 = 0u;
for (int k = 0; k < 16; ++k) {
uint32_t av = (A[row * 4 + k / 4] >> (8 * (k % 4))) & 0xffu;
acc0 += av * ((B[col0 * 4 + k / 4] >> (8 * (k % 4))) & 0xffu);
acc1 += av * ((B[(col0 + 1) * 4 + k / 4] >> (8 * (k % 4))) & 0xffu);
}
R[c][l] += acc0;
R[c2][l] += acc1;
}
}
static void hash_unit(const uint32_t* cache, uint32_t base, int Rsteps, uint64_t out[32]) {
static const uint32_t SEEDW[8] = IGNEUM_SEEDW_INIT;
uint32_t R[8][32], T[32], SEL[32];
const uint32_t mask = IGNEUM_MASK;
L {
uint32_t nonce = base + (uint32_t)l;
for (int i = 0; i < 8; ++i) {
uint32_t x = nonce ^ SEEDW[i]; x += 0x9e3779b9u * (uint32_t)(i + 1); x = splitmix32(x);
R[i][l] = x ^ SEEDW[(i + 1) & 7];
}
}
for (int it = 0; it < 8; ++it) {
L { SEL[l] = R[0][l]; }
#include "ref_program.inc"
for (int k = 0; k < Rsteps; ++k) {
uint16_t v = MM8_TABLE[k];
mm8_step(R, v & 7, (v >> 3) & 7, (v >> 6) & 7, (v >> 9) & 7);
}
}
L {
uint32_t lo = R[0][l] ^ rotl_imm(R[1][l], 7u) ^ rotl_imm(R[2][l], 14u) ^ rotl_imm(R[3][l], 21u);
uint32_t hi = R[4][l] ^ rotl_imm(R[5][l], 9u) ^ rotl_imm(R[6][l], 18u) ^ rotl_imm(R[7][l], 27u);
out[l] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
}
static uint64_t fnv1a64(const void* p, size_t n) {
const uint8_t* b = (const uint8_t*)p; uint64_t h = 0xcbf29ce484222325ull;
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
return h;
}
int main(int argc, char** argv) {
if (argc < 3) { printf("usage: verify_ref <dump file> <R> [--time] [--self-test]\n"); return 2; }
const char* path = argv[1];
int Rsteps = atoi(argv[2]);
int doTime = 0, selfTest = 0;
for (int i = 3; i < argc; ++i) { if (!strcmp(argv[i], "--time")) doTime = 1; if (!strcmp(argv[i], "--self-test")) selfTest = 1; }
if (Rsteps < 0 || Rsteps > IGNEUM_MM8_R_MAX) { printf("R out of range\n"); return 2; }
size_t words = (size_t)1u << IGNEUM_CACHE_LOG2_WORDS;
uint32_t* cache = (uint32_t*)malloc(words * 4u);
if (!cache) { printf("no memory for the cache\n"); return 2; }
double c0 = nowMs();
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(cache, seg);
double c1 = nowMs();
uint64_t fnv = fnv1a64(cache, words * 4u);
printf("cache: host fill %.1f ms, FNV-1a 64 %016llx vs Mac %016llx %s\n", c1 - c0, (unsigned long long)fnv,
(unsigned long long)IGNEUM_CACHE_FNV64, fnv == IGNEUM_CACHE_FNV64 ? "PASS" : "FAIL");
if (fnv != IGNEUM_CACHE_FNV64) return 1;
if (selfTest) { // R = 0 interpreter against the pack's vectors
int ok = 1;
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
uint64_t out[32]; hash_unit(cache, IGNEUM_VEC_BASE[w], 0, out);
int bad = 0; L { if (out[l] != IGNEUM_VEC_OUT[w][l]) ++bad; }
printf("self-test R=0 vector base %u: %s (%d lanes differ)\n", IGNEUM_VEC_BASE[w], bad ? "FAIL" : "PASS", bad);
ok = ok && !bad;
}
if (!ok) return 1;
}
// Read the dump: 32 consecutive lines per warp, same base
FILE* f = fopen(path, "r");
if (!f) { printf("cannot open %s\n", path); return 2; }
uint32_t* bases = NULL; uint64_t* vals = NULL; size_t n = 0, cap = 0;
unsigned base; int lane; unsigned long long v;
while (fscanf(f, "%u %d %llx", &base, &lane, &v) == 3) {
if (n == cap) { cap = cap ? cap * 2 : 1024; bases = realloc(bases, cap * 4); vals = realloc(vals, cap * 8); }
bases[n] = base; vals[n] = v; ++n;
}
fclose(f);
if (n % 32 != 0) { printf("dump has %zu lines, not a multiple of 32\n", n); return 2; }
size_t units = n / 32;
printf("dump: %zu lines, %zu units, R = %d (%d mm8 per hash)\n", n, units, Rsteps, 8 * Rsteps);
size_t equal = 0; int firstPrinted = 0;
double v0 = nowMs();
for (size_t u = 0; u < units; ++u) {
uint64_t out[32]; hash_unit(cache, bases[u * 32], Rsteps, out);
L {
if (out[l] == vals[u * 32 + l]) ++equal;
else if (!firstPrinted) { firstPrinted = 1; printf("first mismatch: base %u lane %d: cpu %016llx gpu %016llx\n", bases[u * 32], l, (unsigned long long)out[l], (unsigned long long)vals[u * 32 + l]); }
}
}
double v1 = nowMs();
printf("%zu of %zu lanes equal (%s) [%.3f ms per unit during the check]\n", equal, n, equal == n ? "PASS" : "FAIL", (v1 - v0) / units);
if (doTime) {
int Rs[2] = { 0, Rsteps }; double ms[2] = { 0, 0 };
for (int t = 0; t < 2; ++t) {
uint64_t out[32]; volatile uint64_t sink = 0;
hash_unit(cache, bases[0], Rs[t], out); // warm
double t0 = nowMs();
for (size_t u = 0; u < units; ++u) { hash_unit(cache, bases[u * 32], Rs[t], out); sink ^= out[0]; }
double t1 = nowMs();
ms[t] = (t1 - t0) / units;
(void)sink;
}
printf("verifier: %.3f ms per unit at R=0, %.3f ms per unit at R=%d, mm8 block delta %.3f ms per unit (%.2f us per mm8 step), averaged over %zu units, one core\n",
ms[0], ms[1], Rsteps, ms[1] - ms[0], Rsteps ? (ms[1] - ms[0]) * 1000.0 / (8.0 * Rsteps) : 0.0, units);
printf("VERIFY R=%d equal=%zu of=%zu ms_r0=%.3f ms_r=%.3f delta=%.3f\n", Rsteps, equal, n, ms[0], ms[1], ms[1] - ms[0]);
}
free(cache); free(bases); free(vals);
return equal == n ? 0 : 1;
}

View file

@ -0,0 +1,219 @@
# state-dataset: the dataset commits to chain state (class "sd1")
Horizon lane 8, new proof of work. Prototype and measurements, 6 October 2026, 19:34 to 19:50 UTC.
Everything here is a TEST HARNESS: no pool, no network, no wallet, nothing touches the devnet.
## The scheme in one paragraph
The hash kernel is unchanged. What changes is the daily dataset build. Today item t of the 1 GiB dataset is
`mh_item(cache, t)`: the 16-word state starts as `s[0..7] = K` (day key) and `s[8..15] = t * MUL[i] + RC[i]`, then
8 rounds of (8 mixers + one dependent 64-byte cache read XORed in), then 8 final mixers. Under sd1 a 64-byte STATE
LEAF for item t is XORed into those 16 initial words before the first mixer: `s[i] ^= leaf(t)[i]`. In the real design
leaf(t) is the t-th 64-byte leaf of a canonical serialisation of the chain's execution state at a certified checkpoint
20 minutes before the day boundary, zero padded where the state is shorter than the dataset. A miner therefore cannot
build the day's dataset without the state, and a verifier that holds no dataset needs leaf(t) for every item it
derives. The prototype measures what that costs (build time, device memory, verifier time) and buys (which leaves a
hash touches, and what a light client would have to be shown).
### Leaf stand-in (synthetic, prototype only)
`leaf(t) = mh_chacha_block(x)` (the pack's own ChaCha12 block from memhard.h) with
`x = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574, S[0..7], t, 0, 0x49676e65, 0x53746174)` and
`S[i] = K[i] ^ 0x5a5a5a5a` (stand-in state root). For the pack's day key
`S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102`. The leaf array is 2^24 items x 64 B = 1 GiB
for the pack's 1 GiB dataset (2^28 words = 2^24 items). Definition in `sd.h`.
## Files
| file | what |
|---|---|
| `kernel.cu`, `memhard.h`, `program.h`, `vectors.h` | the mx8-genesis pack, copied unmodified from `proto-cuda/packs-ca2-mixer/mx8-genesis/` |
| `sd.h` | `mh_leaf`, `mh_item_sd`, `mh_word_sd` (host and device, plain C compatible) |
| `kernel_sd.cu` | the pack's kernel.cu plus `igneum_leaves`, `igneum_build_sd` and their launch wrappers appended; `igneum_hash` byte for byte the pack's (checked with diff) |
| `bench.cu` | the harness: `--mode control|sd1`, `--dump <file> <nwarps>`, `--items-out <file>`, `--chunk-mib M`, `--power-seconds S` |
| `verify_sd.c` | plain C: host cache fill, item bit-exact check from the GPU's items file, 32-lane register-major interpreter of the program against the GPU's dump, lane-0 load trace, per-unit rows |
| `cpu_rows.c` | plain C: the igneum-build-1 rows B.1 and B.2 (OpenMP for the 32-thread row) |
| `run_gpu.sh`, `run_cpu.sh` | the exact commands, as run |
| `results/gpu/`, `results/cpu/` | every log, the dumps and the items files, copied back from the boxes |
## Boxes
- GPU box 2: NVIDIA GeForce RTX 4090, 24 GB (24083 MiB reported), 128 SMs, driver 595.91.07 (CUDA driver 13.2), nvcc 12.8.93,
`-arch=sm_89`, power limit 450 W, Ubuntu 24.04.1. Host CPU AMD EPYC 7352 (48 threads). A CPU-only pool daemon shares the
host (ports 4463/4480); the plain C rows there were pinned to core 2. Working directory `/root/horizon-newpow/state-dataset`.
- CPU box igneum-build-1: AMD EPYC 9454P (48 cores, 96 threads), 128 GB, gcc 13.3.0, L3 256 MiB. Shared with other agents'
builds: load average 19 to 32 during the readings (printed in each log). Rows pinned to core 4 under `nice -n 19`; the
32-thread row on cores 4 to 35. Working directory `/srv/builds/horizon-newpow/state-dataset`.
## Commands
From the Mac (zsh; the GPU box has no rsync, tar over ssh instead):
```
cd /Users/joshm/Projects/igneum-wt-horizon/proto-newpow/state-dataset
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.cu *.h *.c *.sh | ssh -i ~/.ssh/igneum-fleet -p <box-2-port> root@<box-2-ip> 'cd /root/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_gpu.sh > run_gpu.log 2>&1'
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.c *.h *.sh | ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 'cd /srv/builds/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_cpu.sh > run_cpu.log 2>&1'
```
On GPU box 2 (`run_gpu.sh`):
```
export PATH=/usr/local/cuda/bin:$PATH
nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
gcc -O2 -std=c11 -o verify_sd verify_sd.c
./bench --mode control --batches 10 --items-out items_control.txt --dump dump_control.txt 4 --power-seconds 20
./bench --mode sd1 --batches 10 --items-out items_sd1.txt --dump dump_sd1.txt 4 --power-seconds 20
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 64
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 256
./bench --mode control --batches 10 --power-seconds 20 # second pass
./bench --mode sd1 --batches 10 --power-seconds 20 # second pass
taskset -c 2 ./verify_sd --mode control --items items_control.txt --dump dump_control.txt
taskset -c 2 ./verify_sd --mode sd1 --items items_sd1.txt --dump dump_sd1.txt
```
On igneum-build-1 (`run_cpu.sh`):
```
gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
gcc -O2 -std=c11 -o verify_sd verify_sd.c
taskset -c 4 nice -n 19 ./cpu_rows --rows # B.1 (i)(ii)(iii) and the one-core B.2 rows, twice
taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 # B.2 leaf array on 32 threads, twice
taskset -c 4 nice -n 19 ./verify_sd --mode sd1 # the per-unit rows of verify_sd.c on this CPU
```
Timing method: cache fill, leaf array and dataset build are CUDA-event times of the second of two launches (as
host.cu does). MH/s is CUDA-event time over 10 batches of 2^24 after a warm-up batch at base 0. Watts and SM MHz are
the mean of `nvidia-smi -l 1` samples taken after the first 10 s of a 20 s window in which the hash kernel runs back
to back (10 samples used of 22). Fingerprint = FNV-1a 64 over the 2^24 little-endian u64 outputs at base nonce 0,
the project's definition.
## RESULTS
### A. GPU, RTX 4090 (control vs sd1, two passes each)
| row | control | sd1 | note |
|---|---|---|---|
| cache fill, GPU, second pass (ms) | 1.81, 1.85 | 1.85, 1.85 | 256 MiB, unchanged kernel |
| leaf array, GPU, second pass (ms) | none | 5.49, 5.49 | 2^24 ChaCha12 blocks, 1 GiB written at 195 GB/s |
| dataset build, second pass (ms) | 30.55, 30.55 | 31.98, 31.98 | +1.43 ms (+4.7%): one coalesced 64 B read per item |
| build + leaves (ms) | 30.55 | 37.47 | synthetic leaves; a real snapshot arrives from the node instead (see chunked rows) |
| the 3 Mac vectors, standalone and in batch | PASS, PASS | n/a (new values) | sd1 base 0 lane 0 = b600edbed969becc, base 4096 = a533e89c78bb6b74, base 1000000 = beb4cb0c6bab8163 |
| dataset head, word [MASK], 64 Mac samples | PASS | n/a (new values) | sd1 dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf |
| 64 random words vs host derivation | PASS (mh_word) | PASS (mh_word_sd) | |
| 2^24 fingerprint at base 0 | 7c28cfb06c5c65a9 (= the pack's, both passes) | d5b0c16390cad0e8 (new, four runs) | |
| hash rate, 10 batches of 2^24 (MH/s) | 63.083, 63.083 | 63.088, 63.087 | equal within 0.01%; the hash kernel is the same binary over a dataset of the same shape |
| hash rate in the 20 s power window (MH/s) | 63.078, 63.079 | 63.083, 63.083 | |
| watts, mean after 10 s | 207.6, 205.2 | 207.9, 206.5 | run to run noise 2 W |
| SM MHz, mem MHz | 2745, 10251 | 2745, 10251 | |
| MH/s per W | 0.304, 0.307 | 0.303, 0.305 | |
| hash kernel registers | 29 | 29 | 24 resident blocks/SM at 1 warp/block |
Bit-exactness of the build (A.3): 1,024 random item indices (SplitMix64 seeded 0x5d1, `t = z & (2^24 - 1)`), the 16 GPU
words of each item against the host derivation in plain C on the box's CPU with the host-filled cache (65536
`mh_cache_segment` calls, 380 ms in the bench, 573 to 587 ms pinned to core 2 in verify_sd):
| check | control | sd1 |
|---|---|---|
| items equal, in-process (bench.cu, host mh_item / mh_item_sd) | 1024 of 1024 | 1024 of 1024 |
| items equal, plain C (verify_sd.c from the items file) | 1024 of 1024 | 1024 of 1024 |
| items equal after the chunked rebuild (64 MiB chunks; 256 MiB chunks) | n/a | 1024 of 1024; 1024 of 1024 |
| 64 device leaves vs host mh_leaf (incl. t = 0 and 2^24 - 1) | n/a | PASS |
| host cache FNV-1a 64 | 48c4f5bf24166b2e = Mac | 48c4f5bf24166b2e = Mac |
GPU hash outputs vs the C interpreter (4 warps dumped, bases 0, 32, 64, 96; loads through mh_word / mh_word_sd):
| | control | sd1 |
|---|---|---|
| lanes equal | 128 of 128 | 128 of 128 |
| interpreter vs the Mac vector at base 0 | 32 of 32 | n/a |
| interpreting time (4,096 item derivations per warp) | 37 ms | 42 ms |
Device memory (A.4), cudaMemGetInfo, MiB used (context 395 included):
| phase | control | sd1 resident leaves | sd1 chunked 64 MiB | sd1 chunked 256 MiB |
|---|---|---|---|---|
| during the build | 1675 | 2699 | 1741 | 1933 |
| while hashing (leaves freed) | 1803 | 1803 | 1803 | 1803 |
| chunked rebuild total, copies + builds, one stream (ms) | | | 75.80 | 75.47 |
Whole-array host to device copy of the 1 GiB leaf array from pinned memory: 62 ms = 17.2 GB/s (this box's PCIe link
under load; device to host 55 ms). So the chunked build is PCIe-bound: 62 ms of copy plus the 32 ms build, partly
serialised on one stream, gives 76 ms. Two chunk buffers on two streams would hide most of the build under the copy
(not done; the kernel already takes an item range `[t0, t0 + n)` with the chunk's leaves at `leaves - 16 t0`, so
chunking needed no kernel change at all, only the loop in bench.cu).
### B. CPU, igneum-build-1 (EPYC 9454P, one core pinned, nice 19), ms per unit of 4,096 items, 100 units
Two readings, because the box is shared: reading 1 at load 19, reading 2 at load 31.
| row | reading 1 | reading 2 | per item |
|---|---|---|---|
| (i) 4,096 random 64 B reads from a 2 GiB resident leaf array | 0.108 ms (min 0.092, max 0.172) | 0.163 ms | 26 to 40 ns |
| (i) the same from an 8 GiB array (the year-12 size) | 0.139 ms (min 0.127, max 0.183) | 0.209 ms | 34 to 51 ns |
| (ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array) | 0.284 ms | 0.285 ms | 69 ns |
| (iii) 4,096 x mh_item on the host cache, naive, no interleaving | 9.01 ms (min 8.91, max 9.60) | 11.24 ms (min 9.38, max 13.51) | 2.2 to 2.7 us |
| 4,096 x (leaf + mh_item_sd), naive | 9.31 ms | 11.64 ms | |
The same rows on GPU box 2's EPYC 7352, core 2 (verify_sd.c): (ii) 0.369 to 0.374 ms, (iii) 10.14 to 10.21 ms, leaf +
mh_item_sd 10.56 to 10.58 ms.
The project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent
cache misses across items; the naive (iii) figure here is an upper bound. The sd1 increment is the row that matters:
0.11 to 0.21 ms per unit when the verifier holds the 2 to 8 GiB state in RAM, or 0.28 ms when it derives the stand-in
leaf (in the real design it cannot derive a state leaf, it must hold the state or be shown it). Against the 2.06 ms
interleaved verifier that is +5% to +14%; against the naive 9 to 11 ms it is +1% to +3%.
B.2, the daily snapshot build on CPU (igneum-build-1):
| row | time |
|---|---|
| 1 GiB leaf array, 2^24 ChaCha12 blocks, one core | 1.670 s, 1.677 s (100 ns per leaf) |
| the same on 32 threads (OpenMP), second pass (first pass 0.172 s with page faults) | 0.073 s, 0.083 s |
| host cache fill, 256 MiB, one core | 0.447 s, 0.452 s (0.380 s on the 4090 box's host, unpinned; 0.57 s pinned beside the pool daemon) |
B.3, the state sample a light client would check (lane 0 of warp base 0, from the instrumented interpreter):
| | control | sd1 |
|---|---|---|
| distinct words touched by the 128 loads | 128 | 128 |
| distinct items touched | 128 of 2^24 | 128 of 2^24 |
| first 8 load words | 8528400 43034175 255822847 136709537 ... | 8528400 85793927 255822847 180392080 ... |
Loads 1 and 3 share their address in both modes because their addresses are fixed by the nonce before any load's
value feeds back; from load 2 on the sd1 trail diverges. Merkle arithmetic for a 2 GiB state (2^25 leaves of 64 B,
depth 25, 32-byte hashes): 128 openings x 25 x 32 B = 102,400 B of siblings + 128 x 64 B = 8,192 B of leaves =
110,592 B (108 KiB) per lane, uncompressed; a 32-lane unit is 32 x 108 KiB = 3.4 MiB, before deduplicating shared
upper levels (128 random paths in a depth-25 tree share only their top seven levels, so dedup saves about 20% per lane,
approximate; across the 32 lanes of a unit the same seven levels are shared again). Every one of the 128 loads is a distinct item, so there is no saving from repeated items.
## What the numbers mean, per tier (the consequences rule)
- Hash rate, watts, MH/s per W: unchanged for every tier on every vendor, because the hash kernel is the pack's and the
dataset has the same shape. Nothing to do.
- Daily build: +1.4 ms on the 4090 (30.6 to 32.0 ms) when the leaves are resident, 76 ms when streamed from the host
over PCIe in chunks. At the designed 2 GiB dataset the leaf array is 2 GiB and the streamed build is about 150 ms
(approximate, scaling the 17 GB/s copy); on a PCIe 3 x8 slot in a rig, about 4x that (approximate). All far inside the
daily window on every tier.
- Device memory: with resident leaves the build peak at 2 GiB is 2 GiB + 256 MiB + 2 GiB + context, about 4.6 GiB
(approximate), which fits an 8 GB card but not beside a prover. Streamed in 64 MiB chunks the peak is dataset + cache
+ 64 MiB + context, measured 1741 MiB at 1 GiB here, about 2.7 GiB at 2 GiB (approximate): no tier loses memory it has
today. Recommendation: ship chunked only; never hold the leaf array on the device.
- The real cost is delivery: every miner needs the 1 to 2 GiB state serialisation once a day. A miner
beside its own node reads it from disk or loopback (seconds). A pool user without a node must get it from the pool
(2 GiB per day per miner, or a shared download), and a light verifier either holds the state (0.11 to 0.21 ms per unit
extra, measured) or is shown 3.4 MiB of Merkle openings per unit (arithmetic above), which is not a light client any
more. The design decision this prototype leaves open is which of those two the protocol asks of a header verifier.
- CPU snapshot build: 1.7 s on one core, 0.07 s on 32, for the synthetic leaves; the real serialisation is bounded by the
node's state read, not by this.
## What failed, what was cut
- Nothing on the list was cut. verify_sd.c's lane-0 trace printed nothing on the first run (the trace counter was reset
on every warp, so it read 0 after warp 3); fixed, the verifiers re-run, the fix is in the file.
- The GPU box has no rsync; the sources went over with tar through ssh (`COPYFILE_DISABLE=1 --no-xattrs`, the Mac's
tar otherwise writes Apple xattr headers that GNU tar warns about).
- `-std=c11` hides `clock_gettime`; both C files define `_POSIX_C_SOURCE 200809L`.
- igneum-build-1 was under other agents' load (19 to 32) for every reading; both readings are given. The 4090 box's
CPU rows ran beside the pool daemon, pinned to core 2. The GPU itself was idle apart from this bench.
- Not measured: a two-stream chunked build (copy and build overlapped), AMD and Apple builds, the 2 GiB dataset
itself (the pack is 1 GiB; the 2 GiB figures above are labelled approximate).

View file

@ -0,0 +1,464 @@
// state-dataset prototype bench (Horizon lane 8, class sd1). TEST HARNESS ONLY: no pool, no network, no wallet.
//
// Shape follows proto-cuda/host.cu: cache fill (GPU, twice, CUDA events), host cache fill and check, dataset build
// (twice, second pass reported), dataset self-test, the 3 Mac vectors, a warm-up batch of 2^24 at base nonce 0
// (fingerprinted: FNV-1a 64 over the 2^24 little-endian u64 outputs), N timed batches (CUDA events), then a power
// window where nvidia-smi samples at 1 Hz while the hash kernel runs back to back.
//
// --mode control the unmodified pack (igneum_build)
// --mode sd1 leaf array (igneum_leaves) + igneum_build_sd; the hash kernel is the pack's
// --dump <file> <n> write the outputs of the first n warps of the base-0 batch (for verify_sd.c)
// --items-out <file> write the 16 words of 1,024 random items (SplitMix64 seeded 0x5d1) read back from the GPU
// --batches N timed batches after the warm-up (default 10)
// --batch-log2 B nonces per batch (default 24)
// --power-seconds S length of the nvidia-smi window (default 20; 0 skips it)
// --chunk-mib M sd1 only: after the resident build, rebuild with the leaf array streamed from pinned host
// memory in M MiB chunks (the shape a miner uses when the leaves do not fit beside the dataset)
// --device D
//
// Build: nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
#include <cuda_runtime.h>
#include <cstdint>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <chrono>
#include <string>
#include <vector>
#include <thread>
#include <mutex>
#include "program.h"
#include "vectors.h"
#include "memhard.h"
#include "sd.h"
cudaError_t igneum_launch_leaves(uint32_t* leaves, uint32_t nItems);
cudaError_t igneum_launch_build_sd(uint32_t* ds, const uint32_t* cache, const uint32_t* leaves, uint32_t t0, uint32_t nItems);
#define CUDA_CHECK(call) do { cudaError_t err_ = (call); if (err_ != cudaSuccess) { \
std::fprintf(stderr, "CUDA error: %s (%d)\n at %s:%d\n in %s\n", cudaGetErrorString(err_), (int)err_, __FILE__, __LINE__, #call); \
std::exit(2); } } while (0)
static double wallMs() {
using namespace std::chrono;
return duration<double, std::milli>(steady_clock::now().time_since_epoch()).count();
}
static uint64_t fnv1a64(const void* p, size_t n) {
const uint8_t* b = (const uint8_t*)p;
uint64_t h = 0xcbf29ce484222325ull;
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
return h;
}
static uint64_t splitmix64(uint64_t& s) {
s += 0x9E3779B97F4A7C15ull;
uint64_t z = s;
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull;
z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
return z ^ (z >> 31);
}
static float eventMs(cudaEvent_t a, cudaEvent_t b) { float ms = 0.f; CUDA_CHECK(cudaEventElapsedTime(&ms, a, b)); return ms; }
struct Opts {
bool sd1 = false;
int batches = 10, batchLog2 = 24, device = 0, powerSeconds = 20, dumpWarps = 0, chunkMib = 0;
std::string dump, itemsOut;
};
static void usage() {
std::printf("bench --mode control|sd1 [--dump <file> <nwarps>] [--items-out <file>] [--batches 10] [--batch-log2 24] [--power-seconds 20] [--chunk-mib M] [--device 0]\n");
}
static Opts parse(int argc, char** argv) {
Opts o;
for (int i = 1; i < argc; ++i) {
std::string a = argv[i];
auto need = [&](int k) { if (i + k >= argc) { usage(); std::exit(2); } };
if (a == "--mode") { need(1); std::string m = argv[++i]; if (m == "sd1") o.sd1 = true; else if (m == "control") o.sd1 = false; else { usage(); std::exit(2); } }
else if (a == "--dump") { need(2); o.dump = argv[++i]; o.dumpWarps = std::atoi(argv[++i]); }
else if (a == "--items-out") { need(1); o.itemsOut = argv[++i]; }
else if (a == "--batches") { need(1); o.batches = std::atoi(argv[++i]); }
else if (a == "--batch-log2") { need(1); o.batchLog2 = std::atoi(argv[++i]); }
else if (a == "--power-seconds") { need(1); o.powerSeconds = std::atoi(argv[++i]); }
else if (a == "--device") { need(1); o.device = std::atoi(argv[++i]); }
else if (a == "--chunk-mib") { need(1); o.chunkMib = std::atoi(argv[++i]); }
else if (a == "-h" || a == "--help") { usage(); std::exit(0); }
else { std::printf("unknown argument %s\n", argv[i]); usage(); std::exit(2); }
}
if (o.batchLog2 < 10 || o.batchLog2 > 28 || o.batches < 1) { usage(); std::exit(2); }
return o;
}
// nvidia-smi sampler: one line per second, read on its own thread until `timeout` ends the process.
struct Sample { double t; double watts; double smMHz; double memMHz; };
struct Sampler {
std::vector<Sample> samples;
std::mutex mu;
std::thread th;
double t0 = 0;
void start(int device, int seconds) {
t0 = wallMs();
char cmd[512];
std::snprintf(cmd, sizeof(cmd), "timeout %d nvidia-smi -i %d --query-gpu=power.draw,clocks.sm,clocks.mem --format=csv,noheader,nounits -l 1 2>/dev/null", seconds + 2, device);
std::string c = cmd;
th = std::thread([this, c]() {
FILE* f = popen(c.c_str(), "r");
if (!f) return;
char line[256];
while (std::fgets(line, sizeof(line), f)) {
double w = 0, sm = 0, mem = 0;
if (std::sscanf(line, "%lf , %lf , %lf", &w, &sm, &mem) == 3) {
std::lock_guard<std::mutex> g(mu);
samples.push_back({ wallMs() - t0, w, sm, mem });
}
}
pclose(f);
});
}
void join() { if (th.joinable()) th.join(); }
};
static const uint32_t KEYW[8] = IGNEUM_KEY_INIT;
int main(int argc, char** argv) {
Opts o = parse(argc, argv);
const char* modeName = o.sd1 ? "sd1" : "control";
std::printf("state-dataset bench mode %s pack \"%s\" (test harness: no pool, no network, no wallet)\n", modeName, IGNEUM_SEED_STRING);
int count = 0;
CUDA_CHECK(cudaGetDeviceCount(&count));
if (count == 0) { std::printf("FAIL: no CUDA device\n"); return 2; }
CUDA_CHECK(cudaSetDevice(o.device));
cudaDeviceProp prop;
std::memset(&prop, 0, sizeof(prop));
CUDA_CHECK(cudaGetDeviceProperties(&prop, o.device));
int drv = 0, rt = 0;
CUDA_CHECK(cudaDriverGetVersion(&drv));
CUDA_CHECK(cudaRuntimeGetVersion(&rt));
std::printf("GPU: %s (%d SMs, cc %d.%d, %.0f MiB), CUDA driver %d.%d runtime %d.%d\n", prop.name, prop.multiProcessorCount, prop.major, prop.minor,
(double)prop.totalGlobalMem / 1048576.0, drv / 1000, (drv % 100) / 10, rt / 1000, (rt % 100) / 10);
int regs = 0, bps = 0;
CUDA_CHECK(igneum_hash_info(&regs, &bps, 1u));
std::printf("hash kernel: %d registers/thread, %d resident blocks/SM at 1 warp/block\n", regs, bps);
if (o.sd1) {
std::printf("sd1 leaf stand-in: S[i] = K[i] ^ 0x%08x -> S =", SD_STATE_ROOT_XOR);
for (int i = 0; i < 8; ++i) std::printf(" %08x", KEYW[i] ^ SD_STATE_ROOT_XOR);
std::printf("; leaf(t) = ChaCha12 block of (sigma, S, t, 0, %08x, %08x)\n", SD_TAG0, SD_TAG1);
}
size_t free0 = 0, total = 0;
CUDA_CHECK(cudaMemGetInfo(&free0, &total));
std::printf("device memory at start: %.0f MiB used of %.0f MiB (context)\n", (double)(total - free0) / 1048576.0, (double)total / 1048576.0);
cudaEvent_t e0, e1;
CUDA_CHECK(cudaEventCreate(&e0));
CUDA_CHECK(cudaEventCreate(&e1));
// ---- cache: GPU fill twice, host fill, check
const uint32_t cacheWords = 1u << IGNEUM_CACHE_LOG2_WORDS;
const size_t cacheBytes = (size_t)cacheWords * 4u;
uint32_t* dCache = nullptr;
CUDA_CHECK(cudaMalloc((void**)&dCache, cacheBytes));
double cacheFill[2] = { 0, 0 };
for (int pass = 0; pass < 2; ++pass) {
CUDA_CHECK(cudaEventRecord(e0));
CUDA_CHECK(igneum_launch_cache_fill(dCache, IGNEUM_CACHE_SEGMENTS));
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
cacheFill[pass] = eventMs(e0, e1);
}
std::printf("cache fill (GPU): %.2f ms first, %.2f ms second (%u segments x 64 ChaCha12 blocks, %u MiB)\n", cacheFill[0], cacheFill[1], (unsigned)IGNEUM_CACHE_SEGMENTS, (unsigned)(cacheBytes >> 20));
std::vector<uint32_t> hCache(cacheWords);
double h0 = wallMs();
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(hCache.data(), seg);
double hostCacheMs = wallMs() - h0;
bool cachePass = false;
{
std::vector<uint32_t> dev(cacheWords);
CUDA_CHECK(cudaMemcpy(dev.data(), dCache, cacheBytes, cudaMemcpyDeviceToHost));
bool same = std::memcmp(dev.data(), hCache.data(), cacheBytes) == 0;
uint64_t fnv = fnv1a64(hCache.data(), cacheBytes);
cachePass = same && fnv == IGNEUM_CACHE_FNV64;
std::printf("cache fill (host, one thread): %.1f ms; cache check: %s (GPU == host %s, host FNV-1a 64 %016llx vs Mac %016llx)\n",
hostCacheMs, cachePass ? "PASS" : "FAIL", same ? "PASS" : "FAIL", (unsigned long long)fnv, (unsigned long long)IGNEUM_CACHE_FNV64);
}
// ---- sd1: leaf array
const uint32_t words = 1u << IGNEUM_DATASET_LOG2;
const uint32_t mask = words - 1u;
const uint32_t nItems = words / 16u;
uint32_t* dLeaves = nullptr;
double leavesMs[2] = { 0, 0 };
bool leafPass = true;
if (o.sd1) {
CUDA_CHECK(cudaMalloc((void**)&dLeaves, (size_t)nItems * 64u));
for (int pass = 0; pass < 2; ++pass) {
CUDA_CHECK(cudaEventRecord(e0));
CUDA_CHECK(igneum_launch_leaves(dLeaves, nItems));
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
leavesMs[pass] = eventMs(e0, e1);
}
std::printf("leaf array (GPU, %u x 64 B = %u MiB): %.2f ms first, %.2f ms second -> %.1f GB/s written\n", nItems, (unsigned)(((size_t)nItems * 64u) >> 20),
leavesMs[0], leavesMs[1], (double)nItems * 64.0 / 1e9 / (leavesMs[1] / 1000.0));
// 64 leaves against the host derivation
uint64_t s = 0x1ea7ull;
int bad = 0;
for (int k = 0; k < 64; ++k) {
uint32_t t = (k == 0) ? 0u : (k == 1) ? nItems - 1u : (uint32_t)splitmix64(s) & (nItems - 1u);
uint32_t got[16], want[16];
CUDA_CHECK(cudaMemcpy(got, dLeaves + (size_t)t * 16u, 64, cudaMemcpyDeviceToHost));
mh_leaf(t, want);
if (std::memcmp(got, want, 64) != 0) { if (bad == 0) std::printf(" leaf[%u] differs: gpu %08x host %08x\n", t, got[0], want[0]); ++bad; }
}
leafPass = (bad == 0);
std::printf("leaf check: %s (64 leaves incl. 0 and %u vs host mh_leaf)\n", leafPass ? "PASS" : "FAIL", nItems - 1u);
}
// ---- dataset build, twice
uint32_t* dDs = nullptr;
CUDA_CHECK(cudaMalloc((void**)&dDs, (size_t)words * 4u));
double buildMs[2] = { 0, 0 };
for (int pass = 0; pass < 2; ++pass) {
CUDA_CHECK(cudaEventRecord(e0));
if (o.sd1) CUDA_CHECK(igneum_launch_build_sd(dDs, dCache, dLeaves, 0u, nItems));
else CUDA_CHECK(igneum_launch_build(dDs, dCache, nItems));
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
buildMs[pass] = eventMs(e0, e1);
}
std::printf("dataset build (%s): %.2f ms first, %.2f ms second -> %.1f M items/s\n", o.sd1 ? "igneum_build_sd, leaf XOR before the first mixer" : "igneum_build, the pack's",
buildMs[0], buildMs[1], (double)nItems / 1e6 / (buildMs[1] / 1000.0));
size_t freeB = 0;
CUDA_CHECK(cudaMemGetInfo(&freeB, &total));
double usedBuildMiB = (double)(total - freeB) / 1048576.0;
std::printf("device memory after the build: %.0f MiB used (context %.0f + cache %u + %sdataset %u MiB)\n", usedBuildMiB, (double)(total - free0) / 1048576.0,
(unsigned)(cacheBytes >> 20), o.sd1 ? "leaves 1024 + " : "", (unsigned)(((size_t)words * 4u) >> 20));
// ---- dataset self-test
bool dsPass = true;
{
uint32_t head[16];
CUDA_CHECK(cudaMemcpy(head, dDs, sizeof(head), cudaMemcpyDeviceToHost));
std::printf("dataset[0..3] = %08x %08x %08x %08x", head[0], head[1], head[2], head[3]);
if (!o.sd1) {
bool headOk = std::memcmp(head, IGNEUM_DS_HEAD, 64) == 0;
uint32_t last = 0;
CUDA_CHECK(cudaMemcpy(&last, dDs + IGNEUM_DS_LAST_INDEX, 4, cudaMemcpyDeviceToHost));
bool lastOk = (last == IGNEUM_DS_LAST);
int badSample = 0;
for (int k = 0; k < IGNEUM_DS_SAMPLES; ++k) {
uint32_t v = 0;
CUDA_CHECK(cudaMemcpy(&v, dDs + IGNEUM_DS_SAMPLE_INDEX[k], 4, cudaMemcpyDeviceToHost));
if (v != IGNEUM_DS_SAMPLE_VALUE[k]) ++badSample;
}
dsPass = headOk && lastOk && badSample == 0;
std::printf(" head 16 vs Mac %s, word [MASK] vs Mac %s, 64 Mac samples %s\n", headOk ? "PASS" : "FAIL", lastOk ? "PASS" : "FAIL", badSample == 0 ? "PASS" : "FAIL");
} else {
std::printf(" (sd1: new values, no Mac expectation)\n");
}
// 64 random words vs the host derivation (host.cu's points)
int badRnd = 0;
uint64_t s = 0x9E3779B97F4A7C15ull ^ (uint64_t)words;
for (int k = 0; k < 64; ++k) {
uint32_t idx = (uint32_t)splitmix64(s) & mask;
uint32_t v = 0;
CUDA_CHECK(cudaMemcpy(&v, dDs + idx, 4, cudaMemcpyDeviceToHost));
uint32_t want = o.sd1 ? mh_word_sd(hCache.data(), idx) : mh_word(hCache.data(), idx);
if (v != want) ++badRnd;
}
dsPass = dsPass && badRnd == 0;
std::printf("dataset self-test: %s (64 random words vs host %s: %s)\n", dsPass ? "PASS" : "FAIL", o.sd1 ? "mh_word_sd" : "mh_word", badRnd == 0 ? "PASS" : "FAIL");
}
// ---- 1,024 random items (SplitMix64 seeded 0x5d1) read back and compared with the host item derivation
int itemsEqual = 0;
{
FILE* f = o.itemsOut.empty() ? nullptr : std::fopen(o.itemsOut.c_str(), "w");
if (f) std::fprintf(f, "mode %s\nnitems 1024\n", modeName);
uint64_t s = 0x5d1ull;
for (int k = 0; k < 1024; ++k) {
uint32_t t = (uint32_t)splitmix64(s) & (nItems - 1u);
uint32_t got[16], want[16];
CUDA_CHECK(cudaMemcpy(got, dDs + (size_t)t * 16u, 64, cudaMemcpyDeviceToHost));
if (o.sd1) { uint32_t leaf[16]; mh_leaf(t, leaf); mh_item_sd(hCache.data(), leaf, t, want); }
else mh_item(hCache.data(), t, want);
if (std::memcmp(got, want, 64) == 0) ++itemsEqual;
else if (itemsEqual == k) std::printf(" item %u differs: gpu %08x host %08x (first difference)\n", t, got[0], want[0]);
if (f) { std::fprintf(f, "%u", t); for (int i = 0; i < 16; ++i) std::fprintf(f, " %08x", got[i]); std::fprintf(f, "\n"); }
}
if (f) std::fclose(f);
std::printf("item bit-exactness (in-process, host %s on the host cache): %d of 1024 items equal\n", o.sd1 ? "mh_item_sd with host mh_leaf" : "mh_item", itemsEqual);
}
// ---- sd1, chunked: the leaf array lives in pinned host memory (as a real snapshot would, delivered by the node)
// and is streamed to the device in chunks; the dataset is rebuilt chunk by chunk and re-checked. Peak device memory
// is then cache + dataset + one chunk. One stream, so copy and build serialise; two chunk buffers on two streams
// would overlap them (not done here).
double chunkedMs = 0, h2dMs = 0, usedChunkMiB = 0;
int chunkedEqual = -1;
if (o.sd1 && o.chunkMib > 0) {
const size_t leafBytes = (size_t)nItems * 64u;
const size_t chunkBytes = (size_t)o.chunkMib << 20;
const uint32_t chunkItems = (uint32_t)(chunkBytes / 64u);
if (chunkBytes > leafBytes || (leafBytes % chunkBytes) != 0) { std::printf("FAIL: --chunk-mib must divide 1024\n"); return 2; }
uint32_t* hLeaves = nullptr;
CUDA_CHECK(cudaMallocHost((void**)&hLeaves, leafBytes));
CUDA_CHECK(cudaEventRecord(e0));
CUDA_CHECK(cudaMemcpy(hLeaves, dLeaves, leafBytes, cudaMemcpyDeviceToHost));
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
double d2h = eventMs(e0, e1);
CUDA_CHECK(cudaFree(dLeaves)); dLeaves = nullptr;
// one-shot H2D of the whole array into the dataset buffer (overwritten by the rebuild anyway): the PCIe rate
CUDA_CHECK(cudaEventRecord(e0));
CUDA_CHECK(cudaMemcpy(dDs, hLeaves, leafBytes, cudaMemcpyHostToDevice));
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
h2dMs = eventMs(e0, e1);
uint32_t* dChunk = nullptr;
CUDA_CHECK(cudaMalloc((void**)&dChunk, chunkBytes));
CUDA_CHECK(cudaMemGetInfo(&freeB, &total));
usedChunkMiB = (double)(total - freeB) / 1048576.0;
CUDA_CHECK(cudaEventRecord(e0));
for (uint32_t t0 = 0; t0 < nItems; t0 += chunkItems) {
CUDA_CHECK(cudaMemcpyAsync(dChunk, hLeaves + (size_t)t0 * 16u, chunkBytes, cudaMemcpyHostToDevice, 0));
CUDA_CHECK(igneum_launch_build_sd(dDs, dCache, dChunk, t0, chunkItems));
}
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
chunkedMs = eventMs(e0, e1);
// re-check the same 1,024 items
uint64_t s = 0x5d1ull; chunkedEqual = 0;
for (int k = 0; k < 1024; ++k) {
uint32_t t = (uint32_t)splitmix64(s) & (nItems - 1u);
uint32_t got[16], want[16], leaf[16];
CUDA_CHECK(cudaMemcpy(got, dDs + (size_t)t * 16u, 64, cudaMemcpyDeviceToHost));
mh_leaf(t, leaf); mh_item_sd(hCache.data(), leaf, t, want);
if (std::memcmp(got, want, 64) == 0) ++chunkedEqual;
}
std::printf("chunked build (leaves from pinned host memory in %d MiB chunks, %u chunks, one stream): %.2f ms total (copies + builds); whole-array H2D alone %.2f ms = %.1f GB/s, D2H %.2f ms\n",
o.chunkMib, nItems / chunkItems, chunkedMs, h2dMs, (double)leafBytes / 1e9 / (h2dMs / 1000.0), d2h);
std::printf("device memory during the chunked build: %.0f MiB used (context + cache 256 + dataset 1024 + chunk %d MiB); items after the chunked rebuild: %d of 1024 equal\n", usedChunkMiB, o.chunkMib, chunkedEqual);
CUDA_CHECK(cudaFree(dChunk));
CUDA_CHECK(cudaFreeHost(hLeaves));
}
// The leaf array is only needed for the build; a miner frees it before hashing.
if (dLeaves) { CUDA_CHECK(cudaFree(dLeaves)); dLeaves = nullptr; }
// ---- vectors (control: against the Mac; sd1: printed)
const uint32_t nonces = 1u << o.batchLog2;
uint64_t* dOut = nullptr;
CUDA_CHECK(cudaMalloc((void**)&dOut, (size_t)nonces * 8u));
bool vecPass = true;
{
uint64_t got[32];
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
CUDA_CHECK(igneum_launch_hash(dDs, dOut, IGNEUM_VEC_BASE[w], mask, 32u, 1u));
CUDA_CHECK(cudaDeviceSynchronize());
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
if (!o.sd1) {
int bad = 0;
for (int l = 0; l < 32; ++l) if (got[l] != IGNEUM_VEC_OUT[w][l]) ++bad;
vecPass = vecPass && bad == 0;
std::printf("vector warp base %u: %s (%d of 32 lanes differ)\n", IGNEUM_VEC_BASE[w], bad == 0 ? "PASS" : "FAIL", bad);
} else {
std::printf("sd1 warp base %u lane 0: %016llx (new value; checked by verify_sd.c through the dump)\n", IGNEUM_VEC_BASE[w], (unsigned long long)got[0]);
}
}
}
// ---- warm-up batch at base 0: fingerprint, in-batch vectors, dump
double w0 = wallMs();
CUDA_CHECK(igneum_launch_hash(dDs, dOut, 0u, mask, nonces, 1u));
CUDA_CHECK(cudaDeviceSynchronize());
double warmMs = wallMs() - w0;
std::vector<uint64_t> hOut(nonces);
CUDA_CHECK(cudaMemcpy(hOut.data(), dOut, (size_t)nonces * 8u, cudaMemcpyDeviceToHost));
uint64_t fp = fnv1a64(hOut.data(), (size_t)nonces * 8u);
std::printf("warm-up batch: 2^%d hashes in %.2f ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): %016llx%s\n", o.batchLog2, warmMs,
(unsigned long long)fp, o.sd1 ? " (sd1, new value)" : (fp == 0x7c28cfb06c5c65a9ull ? " = 7c28cfb06c5c65a9 (the pack's)" : " DIFFERS from 7c28cfb06c5c65a9"));
bool fpPass = o.sd1 || fp == 0x7c28cfb06c5c65a9ull;
if (!o.sd1) {
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
if ((uint64_t)IGNEUM_VEC_BASE[w] + 32ull > nonces) continue;
int bad = 0;
for (int l = 0; l < 32; ++l) if (hOut[IGNEUM_VEC_BASE[w] + l] != IGNEUM_VEC_OUT[w][l]) ++bad;
vecPass = vecPass && bad == 0;
std::printf("vector warp base %u in batch: %s\n", IGNEUM_VEC_BASE[w], bad == 0 ? "PASS" : "FAIL");
}
}
if (!o.dump.empty() && o.dumpWarps > 0) {
FILE* f = std::fopen(o.dump.c_str(), "w");
if (!f) { std::printf("FAIL: cannot write %s\n", o.dump.c_str()); return 2; }
std::fprintf(f, "mode %s\nwarps %d\n", modeName, o.dumpWarps);
for (int w = 0; w < o.dumpWarps; ++w) {
std::fprintf(f, "base %u\n", 32u * (uint32_t)w);
for (int l = 0; l < 32; ++l) std::fprintf(f, "%016llx\n", (unsigned long long)hOut[32u * (uint32_t)w + l]);
}
std::fclose(f);
std::printf("dump: %d warps (bases 0, 32, ...) written to %s\n", o.dumpWarps, o.dump.c_str());
}
// ---- timed batches
CUDA_CHECK(cudaEventRecord(e0));
for (int b = 1; b <= o.batches; ++b) {
uint32_t base = (uint32_t)((uint64_t)b * (uint64_t)nonces);
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, 1u));
}
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
double gpuMs = eventMs(e0, e1);
double mhs = (double)nonces * (double)o.batches / (gpuMs / 1000.0) / 1e6;
std::printf("timed: %d batches x 2^%d hashes, GPU %.2f ms -> %.3f MH/s (%.2f GB/s useful, loads x 4 B)\n", o.batches, o.batchLog2, gpuMs, mhs, mhs * 1e6 * IGNEUM_LOADS_PER_HASH * 4.0 / 1e9);
// ---- power window: hash back to back for powerSeconds while nvidia-smi samples at 1 Hz
double powerW = 0, smMHz = 0, memMHz = 0, mhsWindow = 0;
int nSamples = 0, nUsed = 0;
if (o.powerSeconds > 0) {
Sampler smp;
smp.start(o.device, o.powerSeconds);
double start = wallMs();
int b = o.batches + 1;
int done = 0;
CUDA_CHECK(cudaEventRecord(e0));
while (wallMs() - start < o.powerSeconds * 1000.0) {
uint32_t base = (uint32_t)((uint64_t)b++ * (uint64_t)nonces);
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, 1u));
CUDA_CHECK(cudaDeviceSynchronize());
++done;
}
CUDA_CHECK(cudaEventRecord(e1));
CUDA_CHECK(cudaEventSynchronize(e1));
double wms = eventMs(e0, e1);
mhsWindow = (double)nonces * (double)done / (wms / 1000.0) / 1e6;
smp.join();
std::lock_guard<std::mutex> g(smp.mu);
nSamples = (int)smp.samples.size();
for (const Sample& s : smp.samples) {
if (s.t < 10000.0 || s.t > (double)o.powerSeconds * 1000.0) continue; // mean after 10 s, inside the load window
powerW += s.watts; smMHz += s.smMHz; memMHz += s.memMHz; ++nUsed;
}
if (nUsed) { powerW /= nUsed; smMHz /= nUsed; memMHz /= nUsed; }
std::printf("power window: %d batches in %.1f s -> %.3f MH/s (events, incl. sync gaps); nvidia-smi %d samples, %d after 10 s: mean %.1f W, SM %.0f MHz, mem %.0f MHz\n",
done, wms / 1000.0, mhsWindow, nSamples, nUsed, powerW, smMHz, memMHz);
if (nUsed) std::printf(" -> %.3f MH/s per W (window rate / mean W)\n", mhsWindow / powerW);
else std::printf(" WARNING: no nvidia-smi samples inside the window (is nvidia-smi on PATH?)\n");
}
size_t freeH = 0;
CUDA_CHECK(cudaMemGetInfo(&freeH, &total));
std::printf("device memory while hashing: %.0f MiB used (leaves freed)\n", (double)(total - freeH) / 1048576.0);
bool overall = cachePass && leafPass && dsPass && vecPass && fpPass && itemsEqual == 1024 && (chunkedEqual < 0 || chunkedEqual == 1024);
std::printf("RESULT mode=%s gpu=%s cache_fill_ms=%.2f host_cache_ms=%.1f leaves_ms=%.2f build_ms=%.2f mem_build_mib=%.0f chunk_mib=%d chunked_ms=%.2f mem_chunked_mib=%.0f chunked_equal=%d items_equal=%d/1024 fingerprint=%016llx mhs=%.3f mhs_window=%.3f watts=%.1f sm_mhz=%.0f mem_mhz=%.0f cache=%s dataset=%s vectors=%s overall=%s\n",
modeName, prop.name, cacheFill[1], hostCacheMs, leavesMs[1], buildMs[1], usedBuildMiB, o.chunkMib, chunkedMs, usedChunkMiB, chunkedEqual, itemsEqual, (unsigned long long)fp, mhs, mhsWindow, powerW, smMHz, memMHz,
cachePass ? "PASS" : "FAIL", dsPass ? "PASS" : "FAIL", o.sd1 ? "n/a" : (vecPass ? "PASS" : "FAIL"), overall ? "PASS" : "FAIL");
CUDA_CHECK(cudaFree(dOut));
CUDA_CHECK(cudaFree(dDs));
CUDA_CHECK(cudaFree(dCache));
return overall ? 0 : 1;
}

View file

@ -0,0 +1,139 @@
/* state-dataset prototype: CPU rows for igneum-build-1 (Horizon lane 8, class sd1). Plain C, -O2.
*
* B.1 the verifier's extra cost per 32-lane unit (4,096 items per unit), ms per unit over 100 units, one core:
* (i) 4,096 random 64-byte reads from a host-resident leaf array of 2 GiB and of 8 GiB (random t, independent,
* the array written once so every page is resident)
* (ii) 4,096 leaves derived on the fly from the state root (one ChaCha12 block each, no array)
* (iii) 4,096 x mh_item on the host cache, naive (no interleaving); plus leaf + mh_item_sd
* B.2 the daily snapshot build on CPU: the 1 GiB leaf array (2^24 ChaCha12 blocks) on one core and on N threads
* (OpenMP), and the host cache fill on one core.
*
* Build: gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
* Run: taskset -c 4 nice -n 19 ./cpu_rows --rows (B.1 and the one-core B.2 rows)
* taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 (the N-thread leaf build)
*/
#define _POSIX_C_SOURCE 200809L
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
#include <omp.h>
#define IGNEUM_NO_CUDA
#include "program.h"
#include "memhard.h"
#include "sd.h"
static double nowMs(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1e3 + ts.tv_nsec / 1e6; }
static uint64_t splitmix64(uint64_t* s) {
*s += 0x9E3779B97F4A7C15ull; uint64_t z = *s;
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull; z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
return z ^ (z >> 31);
}
static volatile uint32_t sink;
/* (i): random 64-byte reads from an array of gib GiB */
static void row_reads(int gib, int units) {
size_t items = ((size_t)gib << 30) / 64u;
uint32_t* arr = (uint32_t*)malloc((size_t)gib << 30);
if (!arr) { printf(" (i) %d GiB: malloc failed\n", gib); return; }
double t0 = nowMs();
/* write every 64-byte line once (resident pages); the content is a cheap pattern, the read cost does not depend on it */
for (size_t i = 0; i < items; ++i) { uint32_t* l = arr + i * 16u; l[0] = (uint32_t)i; l[15] = (uint32_t)(i >> 32) ^ 0x5d1u; }
double fillMs = nowMs() - t0;
uint64_t s = 0x5d1ull; uint32_t acc[16] = { 0 };
double sum = 0, mn = 1e9, mx = 0;
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
for (int u = 0; u < units; ++u) {
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)(splitmix64(&s) % items);
double a = nowMs();
for (int k = 0; k < 4096; ++k) { const uint32_t* l = arr + (size_t)ts[k] * 16u; for (int i = 0; i < 16; ++i) acc[i] ^= l[i]; }
double b = nowMs();
sum += b - a; if (b - a < mn) mn = b - a; if (b - a > mx) mx = b - a;
}
for (int i = 0; i < 16; ++i) sink ^= acc[i];
printf(" (i) 4,096 random 64-byte reads from a %d GiB resident array: %.3f ms per unit (min %.3f, max %.3f; %.0f ns per read; array write pass %.0f ms)\n",
gib, sum / units, mn, mx, sum / units * 1e6 / 4096.0, fillMs);
free(ts); free(arr);
}
int main(int argc, char** argv) {
int rows = 0, leavesThreads = 0, units = 100;
for (int i = 1; i < argc; ++i) {
if (!strcmp(argv[i], "--rows")) rows = 1;
else if (!strcmp(argv[i], "--leaves") && i + 1 < argc) leavesThreads = atoi(argv[++i]);
else if (!strcmp(argv[i], "--units") && i + 1 < argc) units = atoi(argv[++i]);
else { printf("usage: cpu_rows [--rows] [--leaves N] [--units 100]\n"); return 2; }
}
const uint32_t words = 1u << IGNEUM_DATASET_LOG2, nItems = words / 16u;
const uint32_t cacheWords = 1u << IGNEUM_CACHE_LOG2_WORDS;
if (rows) {
printf("B.1 per-unit rows (4,096 per unit, %d units, one thread, ms per unit)\n", units);
row_reads(2, units);
row_reads(8, units);
/* (ii) */
{
uint64_t s = 0x5d1ull; double sum = 0;
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
for (int u = 0; u < units; ++u) {
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)splitmix64(&s) & (nItems - 1u);
double a = nowMs();
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16]; mh_leaf(ts[k], leaf); sink ^= leaf[0] ^ leaf[15]; }
sum += nowMs() - a;
}
printf(" (ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): %.3f ms per unit (%.0f ns per leaf)\n", sum / units, sum / units * 1e6 / 4096.0);
free(ts);
}
/* B.2 host cache fill, one core; then (iii) */
uint32_t* cache = (uint32_t*)malloc((size_t)cacheWords * 4u);
double c0 = nowMs();
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(cache, seg);
double cacheMs = nowMs() - c0;
printf("B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: %.3f s\n", cacheMs / 1000.0);
{
uint64_t s = 0x5d1ull; double sumItem = 0, sumSd = 0, mn = 1e9, mx = 0;
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
for (int u = 0; u < units; ++u) {
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)splitmix64(&s) & (nItems - 1u);
double a = nowMs();
for (int k = 0; k < 4096; ++k) { uint32_t it[16]; mh_item(cache, ts[k], it); sink ^= it[0]; }
double b = nowMs();
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16], it[16]; mh_leaf(ts[k], leaf); mh_item_sd(cache, leaf, ts[k], it); sink ^= it[0]; }
double c = nowMs();
sumItem += b - a; sumSd += c - b; if (b - a < mn) mn = b - a; if (b - a > mx) mx = b - a;
}
printf(" (iii) 4,096 x mh_item on the host cache, naive (no interleaving): %.2f ms per unit (min %.2f, max %.2f)\n", sumItem / units, mn, mx);
printf(" 4,096 x (leaf + mh_item_sd), naive: %.2f ms per unit\n", sumSd / units);
printf(" reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.\n");
free(ts);
}
free(cache);
/* B.2 leaf array on one core */
{
uint32_t* leaves = (uint32_t*)malloc((size_t)nItems * 64u);
double a = nowMs();
for (uint32_t t = 0; t < nItems; ++t) mh_leaf(t, leaves + (size_t)t * 16u);
double ms = nowMs() - a;
sink ^= leaves[0] ^ leaves[(size_t)nItems * 16u - 1u];
printf("B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: %.3f s (%.0f ns per leaf)\n", ms / 1000.0, ms * 1e6 / nItems);
free(leaves);
}
}
if (leavesThreads > 0) {
omp_set_num_threads(leavesThreads);
uint32_t* leaves = (uint32_t*)malloc((size_t)nItems * 64u);
/* first touch in parallel too, then time a second full build so page faults are not in the number */
for (int pass = 0; pass < 2; ++pass) {
double a = nowMs();
#pragma omp parallel for schedule(static)
for (int64_t t = 0; t < (int64_t)nItems; ++t) mh_leaf((uint32_t)t, leaves + (size_t)t * 16u);
double ms = nowMs() - a;
printf("B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), %d threads (OpenMP, %d actual): %.3f s%s\n", leavesThreads, omp_get_max_threads(), ms / 1000.0, pass == 0 ? " (first pass, includes page faults)" : "");
}
sink ^= leaves[0];
free(leaves);
}
return 0;
}

View file

@ -0,0 +1,164 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
uint32_t* d = ds + (size_t)t * 16u;
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r2 = r3 * r4 + r2; // 0 mad
r2 = r1 * r1 + r2; // 1 mad
r2 = r3 * r2 + r2; // 2 mad
r3 = r3 ^ r5; // 3 xor
r7 = r7 ^ ds[r2 & mask]; // 4 load
r5 = r5 ^ ds[r7 & mask]; // 5 load
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
r1 = __umulhi(r1, r5); // 8 mulhi
r6 = rotr_var(r6, r3); // 9 rotr
r3 = r3 | r4; // 10 or
r4 = r4 ^ ds[r3 & mask]; // 11 load
r0 = __umulhi(r0, r4); // 12 mulhi
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
r0 = r0 ^ ds[r4 & mask]; // 14 load
r2 = r2 - r4; // 15 sub
r2 = r2 ^ ds[r0 & mask]; // 16 load
r7 = r7 ^ ds[r2 & mask]; // 17 load
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
r5 = r5 * r0; // 19 mul
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
r6 = __umulhi(r6, r2); // 22 mulhi
r6 = r6 ^ ds[r1 & mask]; // 23 load
r5 = r5 * r0; // 24 mul
r5 = rotl_imm(r5, 19u); // 25 rotl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
r0 = r0 ^ r5; // 27 xor
r0 = r0 ^ r4; // 28 xor
r3 = r3 - r0; // 29 sub
r5 = r5 * r1; // 30 mul
r7 = r7 ^ ds[r2 & mask]; // 31 load
r1 = r1 ^ ds[r0 & mask]; // 32 load
r5 = r5 ^ r6; // 33 xor
r5 = r5 ^ ds[r1 & mask]; // 34 load
r0 = __umulhi(r0, r5); // 35 mulhi
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
r7 = r7 ^ ds[r0 & mask]; // 37 load
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
r2 = r2 ^ r5; // 40 xor
r3 = r6 * r3 + r3; // 41 mad
r6 = r6 - r7; // 42 sub
r7 = r7 ^ r0; // 43 xor
r1 = r1 ^ ds[r7 & mask]; // 44 load
r2 = r2 * r3; // 45 mul
r1 = __umulhi(r1, r5); // 46 mulhi
r4 = r4 - r3; // 47 sub
r2 = rotr_var(r2, r6); // 48 rotr
r3 = r3 ^ ds[r5 & mask]; // 49 load
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
r0 = r0 * r2; // 51 mul
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
r7 = rotl_imm(r7, 14u); // 54 rotl
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
r6 = r6 ^ ds[r7 & mask]; // 56 load
r1 = rotr_var(r1, r5); // 57 rotr
r5 = r5 ^ ds[r4 & mask]; // 58 load
r6 = r6 ^ ds[r2 & mask]; // 59 load
r3 = r5 * r0 + r3; // 60 mad
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
r5 = rotl_imm(r5, 19u); // 63 rotl
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,213 @@
// state-dataset prototype: the mx8-genesis pack kernel plus igneum_leaves and igneum_build_sd (appended at the end).
// The hash kernel igneum_hash and every other line of the pack are byte for byte the pack's kernel.cu.
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
#include "sd.h" // state-dataset prototype: mh_leaf, mh_item_sd (lane 8, class sd1)
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
uint32_t* d = ds + (size_t)t * 16u;
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r2 = r3 * r4 + r2; // 0 mad
r2 = r1 * r1 + r2; // 1 mad
r2 = r3 * r2 + r2; // 2 mad
r3 = r3 ^ r5; // 3 xor
r7 = r7 ^ ds[r2 & mask]; // 4 load
r5 = r5 ^ ds[r7 & mask]; // 5 load
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
r1 = __umulhi(r1, r5); // 8 mulhi
r6 = rotr_var(r6, r3); // 9 rotr
r3 = r3 | r4; // 10 or
r4 = r4 ^ ds[r3 & mask]; // 11 load
r0 = __umulhi(r0, r4); // 12 mulhi
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
r0 = r0 ^ ds[r4 & mask]; // 14 load
r2 = r2 - r4; // 15 sub
r2 = r2 ^ ds[r0 & mask]; // 16 load
r7 = r7 ^ ds[r2 & mask]; // 17 load
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
r5 = r5 * r0; // 19 mul
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
r6 = __umulhi(r6, r2); // 22 mulhi
r6 = r6 ^ ds[r1 & mask]; // 23 load
r5 = r5 * r0; // 24 mul
r5 = rotl_imm(r5, 19u); // 25 rotl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
r0 = r0 ^ r5; // 27 xor
r0 = r0 ^ r4; // 28 xor
r3 = r3 - r0; // 29 sub
r5 = r5 * r1; // 30 mul
r7 = r7 ^ ds[r2 & mask]; // 31 load
r1 = r1 ^ ds[r0 & mask]; // 32 load
r5 = r5 ^ r6; // 33 xor
r5 = r5 ^ ds[r1 & mask]; // 34 load
r0 = __umulhi(r0, r5); // 35 mulhi
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
r7 = r7 ^ ds[r0 & mask]; // 37 load
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
r2 = r2 ^ r5; // 40 xor
r3 = r6 * r3 + r3; // 41 mad
r6 = r6 - r7; // 42 sub
r7 = r7 ^ r0; // 43 xor
r1 = r1 ^ ds[r7 & mask]; // 44 load
r2 = r2 * r3; // 45 mul
r1 = __umulhi(r1, r5); // 46 mulhi
r4 = r4 - r3; // 47 sub
r2 = rotr_var(r2, r6); // 48 rotr
r3 = r3 ^ ds[r5 & mask]; // 49 load
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
r0 = r0 * r2; // 51 mul
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
r7 = rotl_imm(r7, 14u); // 54 rotl
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
r6 = r6 ^ ds[r7 & mask]; // 56 load
r1 = rotr_var(r1, r5); // 57 rotr
r5 = r5 ^ ds[r4 & mask]; // 58 load
r6 = r6 ^ ds[r2 & mask]; // 59 load
r3 = r5 * r0 + r3; // 60 mad
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
r5 = rotl_imm(r5, 19u); // 63 rotl
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}
// ---------------------------------------------------------------------------------------------
// state-dataset prototype (class sd1). One thread per item.
// igneum_leaves: leaves[t * 16 .. t * 16 + 15] = mh_leaf(t) (the synthetic 64-byte state leaf, sd.h).
// igneum_build_sd: item t = mh_item_sd(cache, leaves + 16 t, t). The leaf read is one coalesced 64-byte read per
// thread (consecutive threads read consecutive leaves), so streaming the leaf array in chunks is a matter of
// launching over an item range [t0, t0 + n) with the chunk's leaves at leaves - 16 t0: nothing in the kernel
// depends on the whole array being resident.
__global__ void igneum_leaves(uint32_t* leaves, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t leaf[16];
mh_leaf(t, leaf);
uint32_t* d = leaves + (size_t)t * 16u;
for (uint32_t i = 0u; i < 16u; ++i) d[i] = leaf[i];
}
}
__global__ void igneum_build_sd(uint32_t* ds, const uint32_t* cache, const uint32_t* leaves, uint32_t t0, uint32_t nItems) {
uint32_t t = t0 + blockIdx.x * blockDim.x + threadIdx.x;
if (t < t0 + nItems) {
uint32_t leaf[16];
const uint32_t* l = leaves + (size_t)(t - t0) * 16u;
for (uint32_t i = 0u; i < 16u; ++i) leaf[i] = l[i];
uint32_t s[16];
mh_item_sd(cache, leaf, t, s);
uint32_t* d = ds + (size_t)t * 16u;
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
}
}
cudaError_t igneum_launch_leaves(uint32_t* leaves, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_leaves<<<grid, block>>>(leaves, nItems);
return cudaGetLastError();
}
// Builds items [t0, t0 + nItems) from the leaves of that range (leaves points at leaf t0).
cudaError_t igneum_launch_build_sd(uint32_t* ds, const uint32_t* cache, const uint32_t* leaves, uint32_t t0, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build_sd<<<grid, block>>>(ds, cache, leaves, t0, nItems);
return cudaGetLastError();
}

View file

@ -0,0 +1,109 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#if defined(__CUDACC__)
#define IGNEUM_HD __host__ __device__ __forceinline__
#elif defined(_MSC_VER) && !defined(__cplusplus)
#define IGNEUM_HD static __inline
#else
#define IGNEUM_HD static inline
#endif
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint32_t r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0x3067619fu ^ prev[4];
x[5] = 0x3c269176u ^ prev[5];
x[6] = 0x84a03b03u ^ prev[6];
x[7] = 0xf8c63294u ^ prev[7];
x[8] = 0xff977c5bu ^ prev[8];
x[9] = 0xe60def3eu ^ prev[9];
x[10] = 0x63630141u ^ prev[10];
x[11] = 0xb8fbcb58u ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
s[0] = 0x3067619fu;
s[1] = 0x3c269176u;
s[2] = 0x84a03b03u;
s[3] = 0xf8c63294u;
s[4] = 0xff977c5bu;
s[5] = 0xe60def3eu;
s[6] = 0x63630141u;
s[7] = 0xb8fbcb58u;
s[8] = t * 0x42146205u + 0xbab68293u;
s[9] = t * 0x52cbe0fbu + 0xcc162340u;
s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
s[11] = t * 0x6728907fu + 0xe62b8997u;
s[12] = t * 0xd81d9751u + 0xc9c80297u;
s[13] = t * 0x132952c3u + 0xf74a1654u;
s[14] = t * 0xf60de277u + 0x3d704af5u;
s[15] = t * 0x05358035u + 0x3cf522b7u;
for (uint32_t r = 0u; r < 8u; ++r) {
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
}
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }

View file

@ -0,0 +1,66 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#ifndef IGNEUM_NO_CUDA
#include <cuda_runtime.h>
#endif
#define IGNEUM_SEED_STRING "igneum-genesis"
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
#define IGNEUM_GENERATOR 3
#define IGNEUM_PROGRAM_ATTEMPT 0
#define IGNEUM_PROGRAM_ID 0xe323b9dcaf283a6full
#define IGNEUM_DAY_STRING "2026-10-03"
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
#define IGNEUM_DAY0 0x3067619fu
#define IGNEUM_DAY1 0x3c269176u
#define IGNEUM_DATASET_LOG2 28
#define IGNEUM_MASK 0x0fffffffu
#define IGNEUM_LANES 32
#define IGNEUM_ITERATIONS 8
#define IGNEUM_INSTR_COUNT 64
#define IGNEUM_LOADS_PER_HASH 128
#define IGNEUM_WIDE_LOADS_PER_HASH 0
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
#define IGNEUM_PROGRAM_CLASS "v3"
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
#define IGNEUM_LOAD_CLASS "mx8"
#define IGNEUM_CLASS_MIXER_MULT 8
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
#define IGNEUM_LOAD_SLOTS 16
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
#define IGNEUM_BYTES_PER_HASH 512
#define IGNEUM_FOLD_ROT 11
#define IGNEUM_FOLD_MUL 0x9e3779b1u
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
#define IGNEUM_DATASET_MODE 1
#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u }
#define IGNEUM_CACHE_LOG2_WORDS 26
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
#define IGNEUM_CACHE_SEGMENTS 65536u
#define IGNEUM_ITEM_ROUNDS 8
#define IGNEUM_MIXER_MULT 8 // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)
#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u }
#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u }
#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u }
#ifndef IGNEUM_NO_CUDA
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif

View file

@ -0,0 +1,2 @@
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.179 s (first pass, includes page faults)
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.083 s

View file

@ -0,0 +1,2 @@
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.172 s (first pass, includes page faults)
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.073 s

View file

@ -0,0 +1,9 @@
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.108 ms per unit (min 0.092, max 0.172; 26 ns per read; array write pass 1022 ms)
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.139 ms per unit (min 0.127, max 0.183; 34 ns per read; array write pass 3783 ms)
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.284 ms per unit (69 ns per leaf)
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.447 s
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms per unit (min 8.91, max 9.60)
4,096 x (leaf + mh_item_sd), naive: 9.31 ms per unit
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.670 s (100 ns per leaf)

View file

@ -0,0 +1,9 @@
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.163 ms per unit (min 0.144, max 0.201; 40 ns per read; array write pass 1252 ms)
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.209 ms per unit (min 0.203, max 0.229; 51 ns per read; array write pass 4325 ms)
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.285 ms per unit (70 ns per leaf)
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.452 s
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 11.24 ms per unit (min 9.38, max 13.51)
4,096 x (leaf + mh_item_sd), naive: 11.64 ms per unit
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.677 s (100 ns per leaf)

View file

@ -0,0 +1,27 @@
== Tue Oct 6 07:44:26 PM UTC 2026 on igneum-build-1, load 18.98 28.50 18.66
Model name: AMD EPYC 9454P 48-Core Processor
gcc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0
== build
== B.1 and one-core B.2 rows, core 4, nice 19
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.108 ms per unit (min 0.092, max 0.172; 26 ns per read; array write pass 1022 ms)
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.139 ms per unit (min 0.127, max 0.183; 34 ns per read; array write pass 3783 ms)
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.284 ms per unit (69 ns per leaf)
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.447 s
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms per unit (min 8.91, max 9.60)
4,096 x (leaf + mh_item_sd), naive: 9.31 ms per unit
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.670 s (100 ns per leaf)
== B.2 leaf array on 32 threads (cores 4-35), nice 19
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.172 s (first pass, includes page faults)
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.073 s
== verify_sd per-unit rows on this CPU (core 4), for comparison
verify_sd mode sd1
host cache fill: 480.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.285 ms
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms (min 8.92, max 9.44)
4,096 x (leaf + mh_item_sd), naive: 9.27 ms
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
RESULT verify mode=sd1 host_cache_ms=480.4 items_equal=-1 lanes_equal=-1/0
== done Tue Oct 6 07:44:39 PM UTC 2026, load 16.75 27.55 18.50

View file

@ -0,0 +1,8 @@
verify_sd mode sd1
host cache fill: 480.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.285 ms
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms (min 8.92, max 9.44)
4,096 x (leaf + mh_item_sd), naive: 9.27 ms
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
RESULT verify mode=sd1 host_cache_ms=480.4 items_equal=-1 lanes_equal=-1/0

View file

@ -0,0 +1,24 @@
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
device memory at start: 395 MiB used of 24083 MiB (context)
cache fill (GPU): 1.86 ms first, 1.81 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
cache fill (host, one thread): 380.5 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.1 M items/s
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
dataset self-test: PASS (64 random words vs host mh_word: PASS)
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
vector warp base 0: PASS (0 of 32 lanes differ)
vector warp base 4096: PASS (0 of 32 lanes differ)
vector warp base 1000000: PASS (0 of 32 lanes differ)
warm-up batch: 2^24 hashes in 266.02 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
vector warp base 0 in batch: PASS
vector warp base 4096 in batch: PASS
vector warp base 1000000 in batch: PASS
dump: 4 warps (bases 0, 32, ...) written to dump_control.txt
timed: 10 batches x 2^24 hashes, GPU 2659.56 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
power window: 76 batches in 20.2 s -> 63.078 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.6 W, SM 2745 MHz, mem 10251 MHz
-> 0.304 MH/s per W (window rate / mean W)
device memory while hashing: 1803 MiB used (leaves freed)
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.81 host_cache_ms=380.5 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.078 watts=207.6 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS

View file

@ -0,0 +1,23 @@
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
device memory at start: 395 MiB used of 24083 MiB (context)
cache fill (GPU): 1.86 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
cache fill (host, one thread): 380.0 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.2 M items/s
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
dataset self-test: PASS (64 random words vs host mh_word: PASS)
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
vector warp base 0: PASS (0 of 32 lanes differ)
vector warp base 4096: PASS (0 of 32 lanes differ)
vector warp base 1000000: PASS (0 of 32 lanes differ)
warm-up batch: 2^24 hashes in 265.99 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
vector warp base 0 in batch: PASS
vector warp base 4096 in batch: PASS
vector warp base 1000000 in batch: PASS
timed: 10 batches x 2^24 hashes, GPU 2659.55 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
power window: 76 batches in 20.2 s -> 63.079 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 205.2 W, SM 2745 MHz, mem 10251 MHz
-> 0.307 MH/s per W (window rate / mean W)
device memory while hashing: 1803 MiB used (leaves freed)
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=380.0 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.079 watts=205.2 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS

View file

@ -0,0 +1,134 @@
mode control
warps 4
base 0
19b56348bc85304d
b08a9cfb44aa720f
e1f8f627f780eff7
2ff3e86ff1696161
b65e0578c257e8ac
f37b5705c2bebaae
b19fef670982389e
2374331b28827a11
f492cdd05dda7f88
37700f1385e19885
a27524f3010b3e87
6c8a24de8b938c43
3dc017c8820cfd16
5c80d37146657b68
214082cac03f734a
1d67c665145d72f3
9fb0ae736eb87ff5
cd7d416ae3a6be61
0069cda85c6de58f
0b7d82e717224baa
600960f6d7be30f0
eb29032663b0c4d2
f29583a5766408f1
6c9532a99ad7314d
9a949fd0959cc80f
a20258520e7f6c25
f2ab11f9bb032e38
cc967bcd0c8d07c1
37745267bb3231f2
35a046048c2b69b3
aa51834cd3f364f3
359192708e4f754a
base 32
bbea9a410fe7d5b7
11bf76519a1414d4
59c52e47a752a4fa
3964f1251e8b90e8
75230c9369b790af
8dec6e557dc03a87
0b819148666b3f13
7707a6b8d499e40f
445de8aa5b9699d5
6becbf9e66d7e9de
7a46928b6ed6861f
0cbfca3b2c769872
04e97a5b926b55d5
2a7ba12634ae81bd
96cb4eaf4e25600a
2826bd0a4355a91d
d7ba781350c80652
e87c0820f6136930
e8864f498a00a8e5
c1b445f3b5748847
72def57f2e515d11
f4708f2d80c98b8d
da99416aa6245b67
90f6e6d21f11fc66
a9e3c5f36c715c31
e5e2a586d61e4869
98e1ce97884663da
36e8af738edddc85
8530526bbf3d85bb
22128f72173887c6
30b7aa6ab94449e2
2ea3ca54249bd934
base 64
68b73c212b47e6f2
55acf1fa7bbac2dd
45001afaa68359f7
3271315def7632fb
8af0fe0f8b16148b
2e0c7f7bac0276e1
fe178e4364b66e31
f39df092fabe08ad
e6d7ad0a5d97a7eb
0013aa74d3b41bba
f57f5241a626b717
747067ae7e06c706
50325d1d39c18b09
1fb5aa02196bd962
78eaa19e7dd002f4
8fe0600a8ff1c365
021d685d64796246
281446f54c93f2a5
ba2cc4bef235b006
bfb4bc312bd962a4
cecc08bdb679ca9c
1f079cb2d03b55a1
a8cc9f0647e103a0
2615618b9cb4877f
853f0dd6f03607f2
bf54084008cd8ca3
21986116f4e97e6c
42dbca5da95d1b1c
d7669ed19dea6986
17e407aa53422e43
8c7d4f5bfde2f58e
13adadad05f4540a
base 96
3a0fd3ada6b5a797
aa5a637cb9caa3b7
728b73955e0ad47d
5092aa36c478581c
a463e4220676d004
ce92dfc65d3af26e
2943fd8967c7f215
71acd3ccc8054eaa
2b93c6cd8c22c051
00f10e1b004bab4d
02e2733b82de9bce
1c9d5cd0ed5350ce
dcd13404e4ab0b21
d4afc0e3a2814f63
f2b37d4b0322ff2f
0212964cea5688de
546351f9a3ec9957
d1fe39207bff2a50
a9a38d54a9f85090
e039372a0cc1aa9e
008a4289a796e05e
a9a1dca5b9de1fbc
751a1769f1203a6d
5b66d9c952febc43
47ac39abda0dd053
c7b86597a7ef8d83
d87824ea2fd452c4
82a39552b84b113c
562aa1cb671a8056
4dc55d7e9e0829dd
bb0cf26fbb506ddc
353c62d9a3b6aaed

View file

@ -0,0 +1,134 @@
mode sd1
warps 4
base 0
b600edbed969becc
2d988317678e9099
29609900b3a81764
2506a510ec612b1c
cdc516143acae7eb
dd2901cc2335449d
41614b8fbf48a689
317f5df1eb7c0cbe
2da7edd46d732703
b89b6f78568676a9
e93e96d589f70d8b
197bd4558fe42bf7
911c2ed1bca014bc
448984ee31e0f576
6a9af9585e6ce6ba
446ba1d41100da43
d206864c5393aef9
46b93bf7a7b89196
85ce18e332a13c31
34f1d32a0f153659
d2b8703f431f1237
3bda4edaa92faa6f
d8e5b9f18abf75fd
234a323a6d619bfd
9558b3c4e42cd2f7
016372acc63f6372
7ebac0c7e5fc87f7
81f40d3ef4b5508b
aa6ec7854aacde98
692a050230de2fe8
18b7df938406f9a1
2412df7ada5a3202
base 32
e451566071a2a0ba
d63578a9312e5b32
6bdaecc6d5f7158b
694c691e7dc1de0a
d98a0709679797b3
17a7c68f4ff5e95f
58222d4156171349
af4bcd67686aad89
fdad4d96931b9f2a
5e16c79bd1f2c688
7437b2913e288981
7b4485a218901ccf
540694fa54e15619
e42600d66ac6e0f5
3b2b6322a93124d2
b474e79e977515ae
85c356e1a7f560fe
47961522447a02fc
5cf31818bea49c4a
a517ba58dbed530c
d1b37bf54bd68236
27905e28697a6ef7
42a743ba99c22bf9
5ec48e4245545d4a
38ecfefd174944d8
3dd1891a53a9b415
44120379678fcd2a
e2b43434e9e5b97f
db448577e2229d66
41eb61946c60e9a7
53db71d828b8d7e6
ea19aabc66d1b11a
base 64
34548f5056f27d78
e7fea9a49cf2556e
ebdaf245ef41ec6d
3c77731e46a78bd3
e985324128b13372
c13540b76c4dfc4a
9050d43f68bf20e9
4aa7762e9a08d90a
d106b0610790817b
5c9c91398c9b5e7c
56b71969e73c466e
2d61dea47bf479bb
2ddb7cc8f81fd24c
f422f7adbb7a230e
794858ac6a6bada1
c7a37f1c3cebeea4
241f6632d74d699e
f8e536083c72f300
3234e942615b9891
573677792f6bfe53
6838b021f005dee5
7e1042640bdf5609
4a0399c7b6ca97fd
37ad996ce5fb721a
f633ed2dc43bc76d
e569e4a049054982
a9c5fa29d0ab9603
ca0dcf5025e589f4
17687fa3872dd204
965e7bea52b291a7
d3aec969f424c835
44453d33abd93ef6
base 96
c13d6ea3e05c5e46
1c291747d4676990
349040c82d1e4ad8
f05a14026be82f03
5ee5f52611690c2b
60c99fb5c988e677
3d0899be8af84174
74dad192f88cca2c
005480c1f5e84d6f
a38fa3da9db77186
9495da0d026c51fc
54f4ed7062e92211
6f910c4ad2774774
55752a05b9fa5b71
883b5367142927bf
96449de7c472e375
f8722d3292d2948b
21d0485f77308767
010a6d05ede7793e
56cbfe0810728799
e136e570c4ed04a9
848c020abe56db44
71aa3cddb1bc9408
58757e0f19ddb3ef
f2e4273de46c5968
8664ce7ab76eceac
881db7223c9a750b
788a1553b16895ea
c84a69728aab9555
9019c8d87a1da486
4ab485aa2deb7ca5
b2795a944cca750d

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,85 @@
== Tue Oct 6 19:43:12 UTC 2026 on a4cac49bb840
NVIDIA GeForce RTX 4090, 595.91.07, 3135 MHz, 24564 MiB
Build cuda_12.8.r12.8/compiler.35583870_0
== build
== control
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
device memory at start: 395 MiB used of 24083 MiB (context)
cache fill (GPU): 1.86 ms first, 1.81 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
cache fill (host, one thread): 380.5 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.1 M items/s
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
dataset self-test: PASS (64 random words vs host mh_word: PASS)
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
vector warp base 0: PASS (0 of 32 lanes differ)
vector warp base 4096: PASS (0 of 32 lanes differ)
vector warp base 1000000: PASS (0 of 32 lanes differ)
warm-up batch: 2^24 hashes in 266.02 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
vector warp base 0 in batch: PASS
vector warp base 4096 in batch: PASS
vector warp base 1000000 in batch: PASS
dump: 4 warps (bases 0, 32, ...) written to dump_control.txt
timed: 10 batches x 2^24 hashes, GPU 2659.56 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
power window: 76 batches in 20.2 s -> 63.078 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.6 W, SM 2745 MHz, mem 10251 MHz
-> 0.304 MH/s per W (window rate / mean W)
device memory while hashing: 1803 MiB used (leaves freed)
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.81 host_cache_ms=380.5 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.078 watts=207.6 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS
== sd1
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
device memory at start: 395 MiB used of 24083 MiB (context)
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
cache fill (host, one thread): 383.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.56 ms first, 5.49 ms second -> 195.4 GB/s written
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
dump: 4 warps (bases 0, 32, ...) written to dump_sd1.txt
timed: 10 batches x 2^24 hashes, GPU 2659.35 ms -> 63.088 MH/s (32.30 GB/s useful, loads x 4 B)
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.9 W, SM 2745 MHz, mem 10251 MHz
-> 0.303 MH/s per W (window rate / mean W)
device memory while hashing: 1803 MiB used (leaves freed)
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=383.9 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.088 mhs_window=63.083 watts=207.9 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS
== verify (plain C, CPU core 2)
verify_sd mode control
host cache fill: 587.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
interpreter vs Mac vector base 0: 32 of 32 lanes equal PASS
item bit-exactness (plain C, mh_item on the host cache): 1024 of 1024 items equal PASS
warp base 0: 32 of 32 lanes equal (lane 0 gpu 19b56348bc85304d interp 19b56348bc85304d)
warp base 32: 32 of 32 lanes equal (lane 0 gpu bbea9a410fe7d5b7 interp bbea9a410fe7d5b7)
warp base 64: 32 of 32 lanes equal (lane 0 gpu 68b73c212b47e6f2 interp 68b73c212b47e6f2)
warp base 96: 32 of 32 lanes equal (lane 0 gpu 3a0fd3ada6b5a797 interp 3a0fd3ada6b5a797)
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (37 ms interpreting, 4096 derivations per warp)
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.369 ms
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 10.14 ms (min 9.98, max 10.31)
4,096 x (leaf + mh_item_sd), naive: 10.56 ms
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
RESULT verify mode=control host_cache_ms=587.4 items_equal=1024 lanes_equal=128/128
verify_sd mode sd1
host cache fill: 572.9 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
item bit-exactness (plain C, mh_item_sd with mh_leaf on the host cache): 1024 of 1024 items equal PASS
warp base 0: 32 of 32 lanes equal (lane 0 gpu b600edbed969becc interp b600edbed969becc)
warp base 32: 32 of 32 lanes equal (lane 0 gpu e451566071a2a0ba interp e451566071a2a0ba)
warp base 64: 32 of 32 lanes equal (lane 0 gpu 34548f5056f27d78 interp 34548f5056f27d78)
warp base 96: 32 of 32 lanes equal (lane 0 gpu c13d6ea3e05c5e46 interp c13d6ea3e05c5e46)
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (42 ms interpreting, 4096 derivations per warp)
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.374 ms
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 10.21 ms (min 10.09, max 12.24)
4,096 x (leaf + mh_item_sd), naive: 10.58 ms
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
RESULT verify mode=sd1 host_cache_ms=572.9 items_equal=1024 lanes_equal=128/128
== done Tue Oct 6 19:44:17 UTC 2026

View file

@ -0,0 +1,23 @@
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
device memory at start: 395 MiB used of 24083 MiB (context)
cache fill (GPU): 1.73 ms first, 1.70 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
cache fill (host, one thread): 380.7 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.16 ms first, 5.10 ms second -> 210.7 GB/s written
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.00 ms first, 31.92 ms second -> 525.6 M items/s
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
chunked build (leaves from pinned host memory in 256 MiB chunks, 4 chunks, one stream): 75.47 ms total (copies + builds); whole-array H2D alone 61.92 ms = 17.3 GB/s, D2H 54.61 ms
device memory during the chunked build: 1933 MiB used (context + cache 256 + dataset 1024 + chunk 256 MiB); items after the chunked rebuild: 1024 of 1024 equal
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
warm-up batch: 2^24 hashes in 265.97 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
timed: 10 batches x 2^24 hashes, GPU 2659.40 ms -> 63.086 MH/s (32.30 GB/s useful, loads x 4 B)
device memory while hashing: 1803 MiB used (leaves freed)
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.70 host_cache_ms=380.7 leaves_ms=5.10 build_ms=31.92 mem_build_mib=2699 chunk_mib=256 chunked_ms=75.47 mem_chunked_mib=1933 chunked_equal=1024 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.086 mhs_window=0.000 watts=0.0 sm_mhz=0 mem_mhz=0 cache=PASS dataset=PASS vectors=n/a overall=PASS

View file

@ -0,0 +1,23 @@
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
device memory at start: 395 MiB used of 24083 MiB (context)
cache fill (GPU): 1.85 ms first, 1.82 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
cache fill (host, one thread): 382.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.53 ms first, 5.46 ms second -> 196.5 GB/s written
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
chunked build (leaves from pinned host memory in 64 MiB chunks, 16 chunks, one stream): 75.80 ms total (copies + builds); whole-array H2D alone 62.41 ms = 17.2 GB/s, D2H 54.85 ms
device memory during the chunked build: 1741 MiB used (context + cache 256 + dataset 1024 + chunk 64 MiB); items after the chunked rebuild: 1024 of 1024 equal
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
warm-up batch: 2^24 hashes in 265.98 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
timed: 10 batches x 2^24 hashes, GPU 2659.43 ms -> 63.086 MH/s (32.30 GB/s useful, loads x 4 B)
device memory while hashing: 1803 MiB used (leaves freed)
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.82 host_cache_ms=382.9 leaves_ms=5.46 build_ms=31.98 mem_build_mib=2699 chunk_mib=64 chunked_ms=75.80 mem_chunked_mib=1741 chunked_equal=1024 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.086 mhs_window=0.000 watts=0.0 sm_mhz=0 mem_mhz=0 cache=PASS dataset=PASS vectors=n/a overall=PASS

View file

@ -0,0 +1,24 @@
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
device memory at start: 395 MiB used of 24083 MiB (context)
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
cache fill (host, one thread): 383.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.56 ms first, 5.49 ms second -> 195.4 GB/s written
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
dump: 4 warps (bases 0, 32, ...) written to dump_sd1.txt
timed: 10 batches x 2^24 hashes, GPU 2659.35 ms -> 63.088 MH/s (32.30 GB/s useful, loads x 4 B)
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.9 W, SM 2745 MHz, mem 10251 MHz
-> 0.303 MH/s per W (window rate / mean W)
device memory while hashing: 1803 MiB used (leaves freed)
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=383.9 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.088 mhs_window=63.083 watts=207.9 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS

View file

@ -0,0 +1,23 @@
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
device memory at start: 395 MiB used of 24083 MiB (context)
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
cache fill (host, one thread): 382.6 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.55 ms first, 5.49 ms second -> 195.4 GB/s written
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.07 ms first, 31.98 ms second -> 524.6 M items/s
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
timed: 10 batches x 2^24 hashes, GPU 2659.36 ms -> 63.087 MH/s (32.30 GB/s useful, loads x 4 B)
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 206.5 W, SM 2745 MHz, mem 10251 MHz
-> 0.305 MH/s per W (window rate / mean W)
device memory while hashing: 1803 MiB used (leaves freed)
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=382.6 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.087 mhs_window=63.083 watts=206.5 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS

View file

@ -0,0 +1,13 @@
verify_sd mode control
host cache fill: 574.3 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
interpreter vs Mac vector base 0: 32 of 32 lanes equal PASS
item bit-exactness (plain C, mh_item on the host cache): 1024 of 1024 items equal PASS
warp base 0: 32 of 32 lanes equal (lane 0 gpu 19b56348bc85304d interp 19b56348bc85304d)
warp base 32: 32 of 32 lanes equal (lane 0 gpu bbea9a410fe7d5b7 interp bbea9a410fe7d5b7)
warp base 64: 32 of 32 lanes equal (lane 0 gpu 68b73c212b47e6f2 interp 68b73c212b47e6f2)
warp base 96: 32 of 32 lanes equal (lane 0 gpu 3a0fd3ada6b5a797 interp 3a0fd3ada6b5a797)
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (37 ms interpreting, 4096 derivations per warp)
lane 0 of warp base 0: 128 loads touch 128 distinct words in 128 distinct items (of 2^24 items)
first 8 load words: 8528400 43034175 255822847 136709537 36562395 83203117 165232476 263214792
Merkle openings for a 2 GiB state (2^25 leaves of 64 B, depth 25, 32-byte hashes): 128 openings x 25 x 32 = 102400 bytes of siblings + 128 x 64 = 8192 bytes of leaves = 110592 bytes (108.0 KiB) per lane, uncompressed; 128 distinct items need only 128 openings = 110592 bytes
RESULT verify mode=control host_cache_ms=574.3 items_equal=1024 lanes_equal=128/128

View file

@ -0,0 +1,12 @@
verify_sd mode sd1
host cache fill: 565.8 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
item bit-exactness (plain C, mh_item_sd with mh_leaf on the host cache): 1024 of 1024 items equal PASS
warp base 0: 32 of 32 lanes equal (lane 0 gpu b600edbed969becc interp b600edbed969becc)
warp base 32: 32 of 32 lanes equal (lane 0 gpu e451566071a2a0ba interp e451566071a2a0ba)
warp base 64: 32 of 32 lanes equal (lane 0 gpu 34548f5056f27d78 interp 34548f5056f27d78)
warp base 96: 32 of 32 lanes equal (lane 0 gpu c13d6ea3e05c5e46 interp c13d6ea3e05c5e46)
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (42 ms interpreting, 4096 derivations per warp)
lane 0 of warp base 0: 128 loads touch 128 distinct words in 128 distinct items (of 2^24 items)
first 8 load words: 8528400 85793927 255822847 180392080 172515305 211564358 34853938 259149705
Merkle openings for a 2 GiB state (2^25 leaves of 64 B, depth 25, 32-byte hashes): 128 openings x 25 x 32 = 102400 bytes of siblings + 128 x 64 = 8192 bytes of leaves = 110592 bytes (108.0 KiB) per lane, uncompressed; 128 distinct items need only 128 openings = 110592 bytes
RESULT verify mode=sd1 host_cache_ms=565.8 items_equal=1024 lanes_equal=128/128

View file

@ -0,0 +1,23 @@
#!/bin/bash
# state-dataset prototype: CPU rows on igneum-build-1 (EPYC 9454P). Runs inside /srv/builds/horizon-newpow/state-dataset.
# From the Mac:
# rsync -az -e "ssh -i ~/.ssh/igneum_ed25519" proto-newpow/state-dataset/ build@188.40.146.49:/srv/builds/horizon-newpow/state-dataset/
# ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 'cd /srv/builds/horizon-newpow/state-dataset && touch * && bash run_cpu.sh'
set -euo pipefail
cd "$(dirname "$0")"
touch *.c *.h
echo "== $(date -u) on $(hostname), load $(cut -d' ' -f1-3 /proc/loadavg)"
lscpu | grep "Model name"; gcc --version | head -1
echo "== build"
gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
gcc -O2 -std=c11 -o verify_sd verify_sd.c
echo "== B.1 and one-core B.2 rows, core 4, nice 19"
taskset -c 4 nice -n 19 ./cpu_rows --rows | tee rows.log
echo "== B.2 leaf array on 32 threads (cores 4-35), nice 19"
taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 | tee leaves32.log
echo "== second reading of the B.1 rows (the box is shared; the load average is printed with each reading)"
taskset -c 4 nice -n 19 ./cpu_rows --rows | tee rows2.log
taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 | tee leaves32-2.log
echo "== verify_sd per-unit rows on this CPU (core 4), for comparison"
taskset -c 4 nice -n 19 ./verify_sd --mode sd1 | tee verify_rows.log
echo "== done $(date -u), load $(cut -d' ' -f1-3 /proc/loadavg)"

View file

@ -0,0 +1,30 @@
#!/bin/bash
# state-dataset prototype: GPU box 2 (RTX 4090). Runs inside /root/horizon-newpow/state-dataset after the rsync.
# From the Mac:
# rsync -az -e "ssh -i ~/.ssh/igneum-fleet -p <box-2-port>" proto-newpow/state-dataset/ root@<box-2-ip>:/root/horizon-newpow/state-dataset/
# ssh -i ~/.ssh/igneum-fleet -p <box-2-port> root@<box-2-ip> 'cd /root/horizon-newpow/state-dataset && touch * && bash run_gpu.sh'
set -euo pipefail
export PATH=/usr/local/cuda/bin:$PATH
cd "$(dirname "$0")"
touch *.cu *.h *.c
echo "== $(date -u) on $(hostname)"
nvidia-smi --query-gpu=name,driver_version,clocks.max.sm,memory.total --format=csv,noheader
nvcc --version | tail -1
echo "== build"
nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
gcc -O2 -std=c11 -o verify_sd verify_sd.c
echo "== control"
./bench --mode control --batches 10 --items-out items_control.txt --dump dump_control.txt 4 --power-seconds 20 | tee control.log
echo "== sd1"
./bench --mode sd1 --batches 10 --items-out items_sd1.txt --dump dump_sd1.txt 4 --power-seconds 20 | tee sd1.log
echo "== sd1 chunked (leaves streamed from pinned host memory in 64 and 256 MiB chunks)"
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 64 | tee sd1-chunk64.log
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 256 | tee sd1-chunk256.log
echo "== second pass of the timed rows (repeatability)"
./bench --mode control --batches 10 --power-seconds 20 | tee control2.log
./bench --mode sd1 --batches 10 --power-seconds 20 | tee sd12.log
echo "== verify (plain C, CPU core 2)"
taskset -c 2 ./verify_sd --mode control --items items_control.txt --dump dump_control.txt | tee verify_control.log
taskset -c 2 ./verify_sd --mode sd1 --items items_sd1.txt --dump dump_sd1.txt | tee verify_sd1.log
echo "== logs to collect: run_gpu.log control*.log sd1*.log verify_*.log dump_*.txt items_*.txt"
echo "== done $(date -u)"

View file

@ -0,0 +1,73 @@
// state-dataset prototype (class "sd1", Horizon lane 8): the dataset commits to chain state.
// Shared by kernel_sd.cu (device), bench.cu (host reference), verify_sd.c and cpu_rows.c (plain C).
//
// Scheme: item t is derived exactly as mh_item in memhard.h, except that a 64-byte STATE LEAF for item t is
// XORed into the 16 initial state words before the first mixer:
// s[0..7] = K, s[8..15] = t * MUL + RC (as today)
// s[i] ^= leaf(t)[i] for i in 0..15 (new)
// 8 rounds of (8 mixers + one dependent cache line), 8 final mixers (as today)
// The hash kernel is unchanged.
//
// Leaf stand-in (SYNTHETIC, for the prototype only). In the real design leaf(t) is the t-th 64-byte leaf of the
// canonical serialisation of the chain's execution state at the certified checkpoint 20 minutes before the day
// boundary, zero padded. Here:
// leaf(t) = mh_chacha_block(x) with
// x = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574, S[0..7], t, 0, 0x49676e65, 0x53746174)
// S[i] = K[i] ^ 0x5a5a5a5a (stand-in state root; K = the pack's day key words IGNEUM_KEY_INIT)
// For the pack's day key this gives S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102.
#pragma once
#include "memhard.h"
#define SD_STATE_ROOT_XOR 0x5a5a5a5au
#define SD_TAG0 0x49676e65u /* "Igne" */
#define SD_TAG1 0x53746174u /* "Stat" */
IGNEUM_HD void mh_leaf(uint32_t t, uint32_t* leaf) {
uint32_t x[16];
x[0] = 0x61707865u; x[1] = 0x3320646eu; x[2] = 0x79622d32u; x[3] = 0x6b206574u;
x[4] = 0x3067619fu ^ SD_STATE_ROOT_XOR;
x[5] = 0x3c269176u ^ SD_STATE_ROOT_XOR;
x[6] = 0x84a03b03u ^ SD_STATE_ROOT_XOR;
x[7] = 0xf8c63294u ^ SD_STATE_ROOT_XOR;
x[8] = 0xff977c5bu ^ SD_STATE_ROOT_XOR;
x[9] = 0xe60def3eu ^ SD_STATE_ROOT_XOR;
x[10] = 0x63630141u ^ SD_STATE_ROOT_XOR;
x[11] = 0xb8fbcb58u ^ SD_STATE_ROOT_XOR;
x[12] = t; x[13] = 0u; x[14] = SD_TAG0; x[15] = SD_TAG1;
mh_chacha_block(x, leaf);
}
// Item t under sd1, given its leaf (16 words). Same text as mh_item plus the one XOR line.
IGNEUM_HD void mh_item_sd(const uint32_t* cache, const uint32_t* leaf, uint32_t t, uint32_t* s) {
s[0] = 0x3067619fu;
s[1] = 0x3c269176u;
s[2] = 0x84a03b03u;
s[3] = 0xf8c63294u;
s[4] = 0xff977c5bu;
s[5] = 0xe60def3eu;
s[6] = 0x63630141u;
s[7] = 0xb8fbcb58u;
s[8] = t * 0x42146205u + 0xbab68293u;
s[9] = t * 0x52cbe0fbu + 0xcc162340u;
s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
s[11] = t * 0x6728907fu + 0xe62b8997u;
s[12] = t * 0xd81d9751u + 0xc9c80297u;
s[13] = t * 0x132952c3u + 0xf74a1654u;
s[14] = t * 0xf60de277u + 0x3d704af5u;
s[15] = t * 0x05358035u + 0x3cf522b7u;
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= leaf[i];
for (uint32_t r = 0u; r < 8u; ++r) {
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
}
// dataset[w] under sd1 without the dataset or the leaf array: derive the leaf and the item, take word w & 15.
IGNEUM_HD uint32_t mh_word_sd(const uint32_t* cache, uint32_t w) {
uint32_t leaf[16]; uint32_t s[16];
mh_leaf(w >> 4u, leaf);
mh_item_sd(cache, leaf, w >> 4u, s);
return s[w & 15u];
}

View file

@ -0,0 +1,57 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#define IGNEUM_VEC_WARPS 3
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
{ // base nonce 0
0x19b56348bc85304dull, 0xb08a9cfb44aa720full, 0xe1f8f627f780eff7ull, 0x2ff3e86ff1696161ull, 0xb65e0578c257e8acull, 0xf37b5705c2bebaaeull, 0xb19fef670982389eull, 0x2374331b28827a11ull,
0xf492cdd05dda7f88ull, 0x37700f1385e19885ull, 0xa27524f3010b3e87ull, 0x6c8a24de8b938c43ull, 0x3dc017c8820cfd16ull, 0x5c80d37146657b68ull, 0x214082cac03f734aull, 0x1d67c665145d72f3ull,
0x9fb0ae736eb87ff5ull, 0xcd7d416ae3a6be61ull, 0x0069cda85c6de58full, 0x0b7d82e717224baaull, 0x600960f6d7be30f0ull, 0xeb29032663b0c4d2ull, 0xf29583a5766408f1ull, 0x6c9532a99ad7314dull,
0x9a949fd0959cc80full, 0xa20258520e7f6c25ull, 0xf2ab11f9bb032e38ull, 0xcc967bcd0c8d07c1ull, 0x37745267bb3231f2ull, 0x35a046048c2b69b3ull, 0xaa51834cd3f364f3ull, 0x359192708e4f754aull
},
{ // base nonce 4096
0x62fb132a9943127aull, 0x0b703e577e7f4ecaull, 0xf9f24f5522ce7593ull, 0x3cf5c516abc4332aull, 0xd25523f5f6d7a127ull, 0xd2081a002f983682ull, 0xbf46c54e9b3c4254ull, 0xca362e291e5e5f4dull,
0x6039712f10f457a3ull, 0x8a34b7cabf97c23bull, 0xa473c6a2e0bf59bcull, 0x6cf3926513a4b069ull, 0x297ec2998376a40dull, 0x8efd7f601a8f28dbull, 0x8e72532dfdc1e544ull, 0x917c2b2ebe2a7e00ull,
0x923fbb2d2f635c25ull, 0xce864ea5c0dedad9ull, 0x4b8ec7e874e446efull, 0x1b69b69465449196ull, 0x5ef3a8a6edb369cfull, 0x06c263ef9ce63fc4ull, 0x9c2048fd9d9e2639ull, 0x457fdd96ca4a138eull,
0xd1904018b8d7b6e3ull, 0x8682312fb2e96ab8ull, 0xdc3257e0d0f979a5ull, 0xa51b0a8519d87db5ull, 0x334f08ec056e618bull, 0x3464ce71dc65119dull, 0x6a4d6df066332e04ull, 0x7d7866cb9cfca8ffull
},
{ // base nonce 1000000
0x86b6cb0e13d89b03ull, 0x96299a3f19d7ef15ull, 0x67d2c55100d2f876ull, 0x0a4dfe97d671b728ull, 0x41e4489014d42595ull, 0xf11cb1958c0c0e82ull, 0xf8b70b0c0a03175full, 0x632299df87d5063eull,
0xe198417776130492ull, 0x8ffc5449290d7be2ull, 0x5f2e264eb1311f1bull, 0x988376463ac88586ull, 0x83969eadda489c26ull, 0xbed0a2c3f255d306ull, 0x1a949d271961a819ull, 0x5bce06eb6984725cull,
0x94d5d6a1b0ffd4e9ull, 0xf3c78bae6c2182b4ull, 0xb97e9fe1bbfcdd55ull, 0x70262d1d4c0eccb2ull, 0x1fc93b427dba28d9ull, 0x02b2e3c4317f2a2dull, 0x54d3d42a588edcb9ull, 0x79998677846e7cceull,
0x486522a5425f821aull, 0x95fa88e933360e52ull, 0xc8bae2da2b883f6cull, 0xbe3eb610ad33614full, 0x20efb3c4de82907full, 0xd6b650cfedfb26b7ull, 0x8c24447a646dba26ull, 0x9c004678515e44ecull
}
};
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
static const uint32_t IGNEUM_DS_HEAD[16] = {
0xfdad4319u, 0x1a7b68e1u, 0xde6db608u, 0x13d73892u, 0xd17f447au, 0xb2221ccfu, 0x9db004bdu, 0x57d7d367u,
0xdbc4cf34u, 0x697c009au, 0xc43af1d4u, 0x97f12b2eu, 0x74c37cd0u, 0xc651ea15u, 0x665a6d29u, 0x22330a2du
};
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
static const uint32_t IGNEUM_DS_LAST = 0xa83e7aa6u;
// 64 sampled dataset words (index, value) computed on the Mac.
#define IGNEUM_DS_SAMPLES 64
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
};
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
0x3230bc7bu, 0x7fbfe2c9u, 0xb2690991u, 0x1745c7c5u, 0x0ab0ccafu, 0x1bf87d6bu, 0x160139fdu, 0x719817acu, 0x0155df4bu, 0xbe1e86c3u, 0x680bcd6cu, 0x79c3dc6cu, 0x181e7e5fu, 0x0713a109u, 0xc705dd9fu, 0x3933b7a8u, 0xdd1c0431u, 0x50522b30u, 0xa0020b38u, 0xbff39e96u, 0x21b67e18u, 0x740f8db3u, 0x2baba568u, 0x2c9bef83u, 0x0ad9b671u, 0xc4327869u, 0x7b4fd7d0u, 0x2c29965fu, 0xec56f15fu, 0x61111746u, 0x303a1d6eu, 0xbddcfd1au, 0xf829a355u, 0x6d5df2a9u, 0x01ab8e44u, 0x06d13507u, 0xda8dcfc6u, 0x01a703e1u, 0xafe7d2c1u, 0xc091c3a2u, 0xac1814feu, 0x6e6ff62au, 0x8fdf01bau, 0xdd3f7159u, 0xdfa0d75cu, 0x26684c35u, 0x7f441e63u, 0x88df2570u, 0x8aa4d5ebu, 0xcc816c05u, 0x434df890u, 0xcd392ad6u, 0x1ab4cb63u, 0x595926fau, 0x7cd76b41u, 0x20cb95c4u, 0x13cf823fu, 0xf9daf901u, 0xff9af40au, 0x2c7dfa51u, 0x871206dbu, 0x938c116cu, 0xb64bf199u, 0x5751f874u
};
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u,
0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u
};
static const uint32_t IGNEUM_CACHE_LAST[16] = {
0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du,
0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu
};
static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull;

View file

@ -0,0 +1,270 @@
/* state-dataset prototype: plain C verifier (Horizon lane 8, class sd1).
*
* 1. Fills the 256 MiB cache on the host (mh_cache_segment over all 65536 segments), times it, checks the FNV.
* 2. --items <file>: the 16 words of 1,024 random items the bench read back from the GPU, against
* mh_item (control) or mh_item_sd(hostCache, hostLeaf(t), t) (sd1). Reports "n of 1024 items equal".
* 3. --dump <file>: a 32-lane register-major interpreter of the pack's 64-instruction program, loads through
* mh_word (control) or mh_word_sd (sd1) on the host cache, against the GPU's warps. Reports "n of N lanes equal".
* Lane 0 of the first warp is instrumented: its 128 load addresses, distinct items, and the Merkle arithmetic.
* 4. Per-unit rows on this CPU (ms per unit, 100 units): (ii) 4,096 leaf derivations, (iii) 4,096 x mh_item
* naive, and 4,096 x (leaf + mh_item_sd). The full B.1 rows with the 2 GiB and 8 GiB leaf arrays are cpu_rows.c.
*
* Build: gcc -O2 -std=c11 -o verify_sd verify_sd.c
* Run: taskset -c 2 ./verify_sd --mode sd1 --items items_sd1.txt --dump dump_sd1.txt
*/
#define _POSIX_C_SOURCE 200809L
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
#define IGNEUM_NO_CUDA
#include "program.h"
#include "vectors.h"
#include "memhard.h"
#include "sd.h"
static double nowMs(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1e3 + ts.tv_nsec / 1e6; }
static uint64_t fnv1a64(const void* p, size_t n) {
const uint8_t* b = (const uint8_t*)p; uint64_t h = 0xcbf29ce484222325ull;
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
return h;
}
static uint64_t splitmix64(uint64_t* s) {
*s += 0x9E3779B97F4A7C15ull; uint64_t z = *s;
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull; z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
return z ^ (z >> 31);
}
static uint32_t* hCache;
static int gSd1 = 0;
/* ---- the program, ported from kernel.cu (igneum_hash), register-major over 32 lanes ---- */
static uint32_t splitmix32(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }
static uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
static uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
static uint32_t umulhi(uint32_t a, uint32_t b) { return (uint32_t)(((uint64_t)a * (uint64_t)b) >> 32); }
/* load instrumentation for one lane */
static uint32_t gTraceLane = 0xffffffffu;
static uint32_t gTrace[128];
static int gTraceN = 0;
static uint32_t LOAD(uint32_t lane, uint32_t r, uint32_t mask) {
uint32_t w = r & mask;
if (lane == gTraceLane && gTraceN < 128) gTrace[gTraceN++] = w;
return gSd1 ? mh_word_sd(hCache, w) : mh_word(hCache, w);
}
/* rX[l] ^= rY[l ^ m] for all lanes, through a copy of rY */
#define SHFL_XOR_INTO(rX, rY, m) do { uint32_t tmp_[32]; for (int l_ = 0; l_ < 32; ++l_) tmp_[l_] = rY[l_ ^ (m)]; for (int l_ = 0; l_ < 32; ++l_) rX[l_] ^= tmp_[l_]; } while (0)
static void interpret_warp(uint32_t base, uint32_t mask, uint64_t out[32]) {
uint32_t r0[32], r1[32], r2[32], r3[32], r4[32], r5[32], r6[32], r7[32], sel[32];
for (int l = 0; l < 32; ++l) {
uint32_t nonce = base + (uint32_t)l; uint32_t x;
x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0[l] = x ^ 0x1a155b25u;
x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1[l] = x ^ 0xfddfb732u;
x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2[l] = x ^ 0x4b5af2e8u;
x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3[l] = x ^ 0xc55caf33u;
x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4[l] = x ^ 0xa27c13b7u;
x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5[l] = x ^ 0x06628a48u;
x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6[l] = x ^ 0x03852469u;
x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7[l] = x ^ 0x67a9a7beu;
}
#define L for (int l = 0; l < 32; ++l)
for (uint32_t it = 0u; it < 8u; ++it) {
L sel[l] = r0[l];
L r2[l] = r3[l] * r4[l] + r2[l]; /* 0 mad */
L r2[l] = r1[l] * r1[l] + r2[l]; /* 1 mad */
L r2[l] = r3[l] * r2[l] + r2[l]; /* 2 mad */
L r3[l] = r3[l] ^ r5[l]; /* 3 xor */
L r7[l] = r7[l] ^ LOAD(l, r2[l], mask); /* 4 load */
L r5[l] = r5[l] ^ LOAD(l, r7[l], mask); /* 5 load */
SHFL_XOR_INTO(r1, r4, 8); /* 6 shfl */
SHFL_XOR_INTO(r7, r3, 8); /* 7 shfl */
L r1[l] = umulhi(r1[l], r5[l]); /* 8 mulhi */
L r6[l] = rotr_var(r6[l], r3[l]); /* 9 rotr */
L r3[l] = r3[l] | r4[l]; /* 10 or */
L r4[l] = r4[l] ^ LOAD(l, r3[l], mask); /* 11 load */
L r0[l] = umulhi(r0[l], r4[l]); /* 12 mulhi */
L r5[l] = r5[l] + r1[l] + ((((sel[l] >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); /* 13 add */
L r0[l] = r0[l] ^ LOAD(l, r4[l], mask); /* 14 load */
L r2[l] = r2[l] - r4[l]; /* 15 sub */
L r2[l] = r2[l] ^ LOAD(l, r0[l], mask); /* 16 load */
L r7[l] = r7[l] ^ LOAD(l, r2[l], mask); /* 17 load */
SHFL_XOR_INTO(r7, r3, 4); /* 18 shfl */
L r5[l] = r5[l] * r0[l]; /* 19 mul */
SHFL_XOR_INTO(r3, r4, 2); /* 20 shfl */
SHFL_XOR_INTO(r2, r4, 16); /* 21 shfl */
L r6[l] = umulhi(r6[l], r2[l]); /* 22 mulhi */
L r6[l] = r6[l] ^ LOAD(l, r1[l], mask); /* 23 load */
L r5[l] = r5[l] * r0[l]; /* 24 mul */
L r5[l] = rotl_imm(r5[l], 19u); /* 25 rotl */
SHFL_XOR_INTO(r7, r6, 2); /* 26 shfl */
L r0[l] = r0[l] ^ r5[l]; /* 27 xor */
L r0[l] = r0[l] ^ r4[l]; /* 28 xor */
L r3[l] = r3[l] - r0[l]; /* 29 sub */
L r5[l] = r5[l] * r1[l]; /* 30 mul */
L r7[l] = r7[l] ^ LOAD(l, r2[l], mask); /* 31 load */
L r1[l] = r1[l] ^ LOAD(l, r0[l], mask); /* 32 load */
L r5[l] = r5[l] ^ r6[l]; /* 33 xor */
L r5[l] = r5[l] ^ LOAD(l, r1[l], mask); /* 34 load */
L r0[l] = umulhi(r0[l], r5[l]); /* 35 mulhi */
SHFL_XOR_INTO(r5, r2, 4); /* 36 shfl */
L r7[l] = r7[l] ^ LOAD(l, r0[l], mask); /* 37 load */
L r3[l] = r3[l] + r1[l] + ((((sel[l] >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); /* 38 add */
SHFL_XOR_INTO(r1, r5, 4); /* 39 shfl */
L r2[l] = r2[l] ^ r5[l]; /* 40 xor */
L r3[l] = r6[l] * r3[l] + r3[l]; /* 41 mad */
L r6[l] = r6[l] - r7[l]; /* 42 sub */
L r7[l] = r7[l] ^ r0[l]; /* 43 xor */
L r1[l] = r1[l] ^ LOAD(l, r7[l], mask); /* 44 load */
L r2[l] = r2[l] * r3[l]; /* 45 mul */
L r1[l] = umulhi(r1[l], r5[l]); /* 46 mulhi */
L r4[l] = r4[l] - r3[l]; /* 47 sub */
L r2[l] = rotr_var(r2[l], r6[l]); /* 48 rotr */
L r3[l] = r3[l] ^ LOAD(l, r5[l], mask); /* 49 load */
L r1[l] = r1[l] + r5[l] + ((((sel[l] >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); /* 50 add */
L r0[l] = r0[l] * r2[l]; /* 51 mul */
L r0[l] = r0[l] + r2[l] + ((((sel[l] >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); /* 52 add */
L r1[l] = r1[l] + r0[l] + ((((sel[l] >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); /* 53 add */
L r7[l] = rotl_imm(r7[l], 14u); /* 54 rotl */
L r3[l] = r3[l] + r7[l] + ((((sel[l] >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); /* 55 add */
L r6[l] = r6[l] ^ LOAD(l, r7[l], mask); /* 56 load */
L r1[l] = rotr_var(r1[l], r5[l]); /* 57 rotr */
L r5[l] = r5[l] ^ LOAD(l, r4[l], mask); /* 58 load */
L r6[l] = r6[l] ^ LOAD(l, r2[l], mask); /* 59 load */
L r3[l] = r5[l] * r0[l] + r3[l]; /* 60 mad */
L r5[l] = r5[l] + r7[l] + ((((sel[l] >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); /* 61 add */
L r4[l] = r4[l] + r6[l] + ((((sel[l] >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); /* 62 add */
L r5[l] = rotl_imm(r5[l], 19u); /* 63 rotl */
}
L {
uint32_t lo = r0[l] ^ rotl_imm(r1[l], 7u) ^ rotl_imm(r2[l], 14u) ^ rotl_imm(r3[l], 21u);
uint32_t hi = r4[l] ^ rotl_imm(r5[l], 9u) ^ rotl_imm(r6[l], 18u) ^ rotl_imm(r7[l], 27u);
out[l] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
#undef L
}
static int cmp_u32(const void* a, const void* b) { uint32_t x = *(const uint32_t*)a, y = *(const uint32_t*)b; return x < y ? -1 : x > y; }
int main(int argc, char** argv) {
const char* itemsFile = NULL; const char* dumpFile = NULL; int units = 100; int rows = 1;
for (int i = 1; i < argc; ++i) {
if (!strcmp(argv[i], "--mode") && i + 1 < argc) { ++i; gSd1 = !strcmp(argv[i], "sd1"); }
else if (!strcmp(argv[i], "--items") && i + 1 < argc) itemsFile = argv[++i];
else if (!strcmp(argv[i], "--dump") && i + 1 < argc) dumpFile = argv[++i];
else if (!strcmp(argv[i], "--units") && i + 1 < argc) units = atoi(argv[++i]);
else if (!strcmp(argv[i], "--no-rows")) rows = 0;
else { printf("usage: verify_sd --mode control|sd1 [--items f] [--dump f] [--units 100] [--no-rows]\n"); return 2; }
}
const uint32_t words = 1u << IGNEUM_DATASET_LOG2, mask = words - 1u, nItems = words / 16u;
const uint32_t cacheWords = 1u << IGNEUM_CACHE_LOG2_WORDS;
printf("verify_sd mode %s\n", gSd1 ? "sd1" : "control");
hCache = (uint32_t*)malloc((size_t)cacheWords * 4u);
double t0 = nowMs();
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(hCache, seg);
double cacheMs = nowMs() - t0;
uint64_t fnv = fnv1a64(hCache, (size_t)cacheWords * 4u);
printf("host cache fill: %.1f ms one thread; FNV-1a 64 %016llx vs Mac %016llx %s\n", cacheMs, (unsigned long long)fnv, (unsigned long long)IGNEUM_CACHE_FNV64, fnv == IGNEUM_CACHE_FNV64 ? "PASS" : "FAIL");
/* control sanity: the first Mac vector through the interpreter */
if (!gSd1) {
uint64_t out[32]; interpret_warp(0u, mask, out);
int bad = 0; for (int l = 0; l < 32; ++l) if (out[l] != IGNEUM_VEC_OUT[0][l]) ++bad;
printf("interpreter vs Mac vector base 0: %d of 32 lanes equal %s\n", 32 - bad, bad == 0 ? "PASS" : "FAIL");
}
int itemsEqual = -1;
if (itemsFile) {
FILE* f = fopen(itemsFile, "r");
if (!f) { printf("FAIL: cannot open %s\n", itemsFile); return 2; }
char mode[16]; int n = 0;
if (fscanf(f, "mode %15s\nnitems %d\n", mode, &n) != 2) { printf("FAIL: bad items file header\n"); return 2; }
if (strcmp(mode, gSd1 ? "sd1" : "control") != 0) printf("WARNING: items file mode %s, verifier mode %s\n", mode, gSd1 ? "sd1" : "control");
itemsEqual = 0; int read = 0;
for (int k = 0; k < n; ++k) {
uint32_t t, got[16], want[16];
if (fscanf(f, "%u", &t) != 1) break;
for (int i = 0; i < 16; ++i) if (fscanf(f, " %x", &got[i]) != 1) break;
++read;
if (t >= nItems) continue;
if (gSd1) { uint32_t leaf[16]; mh_leaf(t, leaf); mh_item_sd(hCache, leaf, t, want); } else mh_item(hCache, t, want);
if (memcmp(got, want, 64) == 0) ++itemsEqual;
else if (itemsEqual == k) printf(" item %u: gpu %08x host %08x (first difference)\n", t, got[0], want[0]);
}
fclose(f);
printf("item bit-exactness (plain C, %s on the host cache): %d of %d items equal %s\n", gSd1 ? "mh_item_sd with mh_leaf" : "mh_item", itemsEqual, read, itemsEqual == read && read > 0 ? "PASS" : "FAIL");
}
int lanesEqual = -1, lanesTotal = 0;
if (dumpFile) {
FILE* f = fopen(dumpFile, "r");
if (!f) { printf("FAIL: cannot open %s\n", dumpFile); return 2; }
char mode[16]; int nw = 0;
if (fscanf(f, "mode %15s\nwarps %d\n", mode, &nw) != 2) { printf("FAIL: bad dump header\n"); return 2; }
if (strcmp(mode, gSd1 ? "sd1" : "control") != 0) printf("WARNING: dump mode %s, verifier mode %s\n", mode, gSd1 ? "sd1" : "control");
lanesEqual = 0;
double tI = nowMs();
for (int w = 0; w < nw; ++w) {
uint32_t base; uint64_t got[32], out[32];
if (fscanf(f, "base %u\n", &base) != 1) break;
for (int l = 0; l < 32; ++l) { unsigned long long v; if (fscanf(f, "%llx\n", &v) != 1) break; got[l] = v; }
if (w == 0) { gTraceLane = 0u; gTraceN = 0; } else gTraceLane = 0xffffffffu; /* trace lane 0 of the first warp only */
interpret_warp(base, mask, out);
int eq = 0; for (int l = 0; l < 32; ++l) if (out[l] == got[l]) ++eq;
lanesEqual += eq; lanesTotal += 32;
printf("warp base %u: %d of 32 lanes equal (lane 0 gpu %016llx interp %016llx)\n", base, eq, (unsigned long long)got[0], (unsigned long long)out[0]);
}
fclose(f);
printf("GPU hash outputs vs interpreter: %d of %d lanes equal %s (%.0f ms interpreting, %d derivations per warp)\n", lanesEqual, lanesTotal, lanesEqual == lanesTotal && lanesTotal > 0 ? "PASS" : "FAIL", nowMs() - tI, 32 * IGNEUM_LOADS_PER_HASH);
/* B.3: the state sample of lane 0 of the first warp */
if (gTraceN == 128) {
uint32_t items[128]; for (int i = 0; i < 128; ++i) items[i] = gTrace[i] >> 4;
qsort(items, 128, sizeof(uint32_t), cmp_u32);
int distinctItems = 1; for (int i = 1; i < 128; ++i) if (items[i] != items[i - 1]) ++distinctItems;
uint32_t ws[128]; memcpy(ws, gTrace, sizeof(ws)); qsort(ws, 128, sizeof(uint32_t), cmp_u32);
int distinctWords = 1; for (int i = 1; i < 128; ++i) if (ws[i] != ws[i - 1]) ++distinctWords;
printf("lane 0 of warp base 0: 128 loads touch %d distinct words in %d distinct items (of 2^%d items)\n", distinctWords, distinctItems, IGNEUM_DATASET_LOG2 - 4);
printf(" first 8 load words: %u %u %u %u %u %u %u %u\n", gTrace[0], gTrace[1], gTrace[2], gTrace[3], gTrace[4], gTrace[5], gTrace[6], gTrace[7]);
long openings = 128, depth = 25, hash = 32;
printf(" Merkle openings for a 2 GiB state (2^25 leaves of 64 B, depth %ld, %ld-byte hashes): %ld openings x %ld x %ld = %ld bytes of siblings + %ld x 64 = %ld bytes of leaves = %ld bytes (%.1f KiB) per lane, uncompressed; %d distinct items need only %d openings = %ld bytes\n",
depth, hash, openings, depth, hash, openings * depth * hash, openings, openings * 64, openings * depth * hash + openings * 64, (openings * depth * hash + openings * 64) / 1024.0,
distinctItems, distinctItems, (long)distinctItems * (depth * hash + 64));
}
}
if (rows) {
/* (ii) 4,096 leaf derivations per unit; (iii) 4,096 x mh_item naive; and leaf + mh_item_sd */
uint64_t s = 0x5d1ull; volatile uint32_t sink = 0;
double sumLeaf = 0, sumItem = 0, sumSd = 0, minItem = 1e9, maxItem = 0;
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
for (int u = 0; u < units; ++u) {
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)splitmix64(&s) & (nItems - 1u);
double a = nowMs();
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16]; mh_leaf(ts[k], leaf); sink ^= leaf[0]; }
double b = nowMs();
for (int k = 0; k < 4096; ++k) { uint32_t it[16]; mh_item(hCache, ts[k], it); sink ^= it[0]; }
double c = nowMs();
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16], it[16]; mh_leaf(ts[k], leaf); mh_item_sd(hCache, leaf, ts[k], it); sink ^= it[0]; }
double d = nowMs();
sumLeaf += b - a; sumItem += c - b; sumSd += d - c;
if (c - b < minItem) minItem = c - b;
if (c - b > maxItem) maxItem = c - b;
}
free(ts);
printf("per-unit rows on this CPU (4,096 per unit, %d units, one thread, ms per unit):\n", units);
printf(" (ii) 4,096 leaf derivations (ChaCha12 block each, no array): %.3f ms\n", sumLeaf / units);
printf(" (iii) 4,096 x mh_item on the host cache, naive (no interleaving): %.2f ms (min %.2f, max %.2f)\n", sumItem / units, minItem, maxItem);
printf(" 4,096 x (leaf + mh_item_sd), naive: %.2f ms\n", sumSd / units);
printf(" reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.\n");
}
printf("RESULT verify mode=%s host_cache_ms=%.1f items_equal=%d lanes_equal=%d/%d\n", gSd1 ? "sd1" : "control", cacheMs, itemsEqual, lanesEqual, lanesTotal);
free(hCache);
return 0;
}

View file

@ -0,0 +1,3 @@
# sim/horizon/new-pow
`chip_rows.py`: the scheme B chip rows (f = 1 chip plus a tensor array at k times the card's block energy) from the measured card figures. Run: `python3 chip_rows.py --card-uj 2.56 --card-uj-r 3.0 --r 128`. Inputs and outputs are explained in `docs/analysis/horizon/new-pow.md` section 5.3.

View file

@ -0,0 +1,36 @@
#!/usr/bin/env python3
"""Chip rows for the Horizon new-pow lane (scheme B, the tensor-shaped shadow), by the chip-model-v3 section 5 method.
Usage:
python3 chip_rows.py --card-uj <microjoules per hash of the honest card at R=0> \
--card-uj-r <microjoules per hash at the measured R> --r <R> \
[--mem-uj 0.466] [--k 1.0 0.5 0.3]
Model (docs/analysis/chip-model-v3.md 5.4 and docs/analysis/latency-shadow-2026-10-06.md 6):
chip energy per hash = memory system energy (f = 1 chip: 0.466 uJ GDDR7, 0.321 uJ one HBM3 stack)
+ block energy on the chip = (card block energy) x k
card block energy = card_uj_r - card_uj (the measured marginal of the block on the honest card)
gain per joule = card_uj_r / chip_uj
Every chip figure is arithmetic on cited figures and approximate; the card figures are measured and named in the lane file.
"""
import argparse
p = argparse.ArgumentParser()
p.add_argument("--card-uj", type=float, required=True, help="honest card microjoules per hash at R = 0 (measured)")
p.add_argument("--card-uj-r", type=float, required=True, help="honest card microjoules per hash at the measured R")
p.add_argument("--r", type=int, required=True, help="mm8 steps per iteration (8 R per hash)")
p.add_argument("--mem-uj", type=float, nargs="+", default=[0.466, 0.321], help="f = 1 chip memory energy per hash, uJ (GDDR7, HBM3 one stack)")
p.add_argument("--k", type=float, nargs="+", default=[1.0, 0.5, 0.3], help="chip block energy per op over the card's")
a = p.parse_args()
block = a.card_uj_r - a.card_uj
macs = 8 * a.r * 1024
print(f"R = {a.r}: {8*a.r} mm8 per hash, {macs} multiply-adds per hash")
print(f"card: {a.card_uj:.3f} uJ at R = 0, {a.card_uj_r:.3f} uJ at R = {a.r}; block {block:.3f} uJ = {block*1e6/macs if macs else 0:.3f} pJ per multiply-add")
print("| memory | k | chip uJ per hash | gain per joule (card over chip) | gain at R = 0 for comparison |")
print("|---|---|---|---|---|")
names = ["GDDR7 (16 devices)", "HBM3 (one stack)", "HBM3 (eight stacks)"]
for i, m in enumerate(a.mem_uj):
for k in a.k:
chip = m + block * k
print(f"| {names[i] if i < len(names) else m} | {k} | {chip:.3f} | {a.card_uj_r/chip:.2f}x | {a.card_uj/m:.2f}x |")