Horizon: lane 8 (new-proof-of-work) lands: three schemes, two prototypes measured on rented 4090s
docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes, the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8 two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The lane's standing rule: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is least efficient per op. Chip rows in sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by placeholders in the READMEs and the run script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
62941ae627
commit
a664af6fc9
87 changed files with 13059 additions and 13 deletions
|
|
@ -10,9 +10,10 @@ the project lead's mandate, verbatim: "if we create a new way of hashing or a ne
|
|||
|---|---|
|
||||
| 19:05 | Lane started. Read: preamble, CLAUDE.md, the two personas, spec 01 (whole), 04, 07, chip-model-v3 (whole), asic-resistance-history (sections 0 to 3 and 4.3, 5), latency-shadow-2026-10-06 (whole), counter-asic-3-status sections 1 to 5, counter-asic-3-node section 6 (the P2 signalling rule), int8-matrix-family sections 1 to 3, scratch-soundness verdict, proving-methods (whole), fud-ledger M1, M7, M16, M22, M28, P2, F13, proto-cuda host.cu and the mx8-genesis pack (kernel.cu, memhard.h, program.h, vectors.h), proto-cuda/emu, family-probe.cu, the fleet's prover-tiers-real-cards.md, bench-log line 2582 (rental cost) |
|
||||
| 19:25 | Fleet agent asked for two boxes; answered at 19:29: two quiet RTX 4090s (RunPod, nvcc 12.8 at /usr/local/cuda/bin, directory /root/horizon-newpow, until 22:30Z). No quiet Ampere card exists tonight; a loaded 3090 is offered. Main's note: cost rows use bench-log 2582 (USD 0.0117 per MH/s-hour) |
|
||||
| 19:35 | File skeleton written. Two prototype sub-agents launched (budget two at once): `mma-shadow` on box 1 (47.47.180.77), `state-dataset` on box 2 (213.173.98.36) plus CPU rows on igneum-build-1. Designs being written in this file meanwhile |
|
||||
| 19:35 | File skeleton written. Two prototype sub-agents launched (budget two at once): `mma-shadow` on box 1 (<box-1-ip>), `state-dataset` on box 2 (<box-2-ip>) plus CPU rows on igneum-build-1. Designs being written in this file meanwhile |
|
||||
| 19:41 | Section 3 complete: the three designs, the one-table comparison, the migration path. Scheme A's verdict is already visible in its own numbers (A1 dead on 2.9 MB of openings per block, A2 dead on sampleability and a 32 to 40 ms proof verify; A0 is scheme C with the trace as state). Prototypes running: `mma-shadow` (box 1) and `state-dataset` (box 2 and igneum-build-1). mm8 two-output correction sent to the prototype |
|
||||
| (next) | Section 4 reviews (cryptographer, consensus engineer, two per scheme); the pick; section 5 measured rows as they land; section 6 verdicts; section 7 ranked next steps |
|
||||
| 19:50 | Section 4 (reviews, the pick: B and C), section 7 (ranked next steps) and section 8 (open questions) written. Scheme C prototype complete and measured on box 2 and igneum-build-1: hash rate and watts unchanged (63.08 vs 63.09 MH/s, 207 W), build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact 1,024 of 1,024 items and 128 of 128 lanes; rows in 5.2. Scheme B ladder running on box 1 (R = 0, 8, 32, 128 measured, 512 in progress) |
|
||||
| 20:20 | Scheme B ladder complete on box 1 (R = 0 to 512, rate flat at 63.08 MH/s, 201 to 216 W, every fingerprint PTX = reference, 1,024 of 1,024 lanes at every R, verifier delta 0.05 to 4.39 ms). Sections 5.1, 5.3, 6 and 9 written. Box 2 released 19:53Z; box 1 released on the mma agent's report. File complete |
|
||||
|
||||
## 1. What was read and the facts this lane stands on
|
||||
|
||||
|
|
@ -133,8 +134,8 @@ So the fleet splits by generation: Turing-and-later NVIDIA and RDNA 3-and-later
|
|||
| Scheme | What it is | Chip edge per joule vs the 5090 (f = 1 chip, chip-model-v3 method) | Verifier ms per unit (model, then measured in section 5) | Vendor bit-exactness | Ships as class v5? |
|
||||
|---|---|---|---|---|---|
|
||||
| A, mining is proving | A1 committed codeword, A2 proving steps: dead on bytes and on sampleability; A0 trace-as-dataset survives and is C with the trace as state | A0 unchanged (5.1x GDDR7); A1, A2 not priced | A0: 2.06 + the leaf-read row; A1: 2.9 MB of openings; A2: 32 to 40 ms | A0 adds the zkVM executor to consensus | NEVER as mining = proving; A0 folds into C |
|
||||
| B, tensor-shaped shadow | class v3 plus `8 R` int8 8x8x16 tile steps per hash in the PTX fragment layout, two outputs per lane | memory + `8 R x 1,024 x e_mac x k_mma`; `k_mma` near 1 with a floor near 0.5 (approximate); rows from the measured `e_mac` in 5.3 | 2.06 + about 0.9 microseconds per step per unit scalar (R = 128: +0.9 ms; R = 512: +3.7 ms); VNNI 4x to 16x less | native on sm_75+ and RDNA 3+ (AMD layout unverified); emulated on Pascal, RDNA 2, Apple | prototype further; a class v5 candidate after the AMD gate |
|
||||
| C, stored state | the daily dataset derives from the execution state snapshot; kernel unchanged; a state sample per block | unchanged (5.1x GDDR7, 7.5x to 9.2x HBM3); the f = 0 chip disappears | 2.06 + 4,096 leaf reads (section 5.2) | nothing vendor-specific; the serialisation is the consensus risk | prototype further; a class v5 candidate on its own or beside B |
|
||||
| B, tensor-shaped shadow | class v3 plus `8 R` int8 8x8x16 tile steps per hash in the PTX fragment layout, two outputs per lane | memory + `8 R x 1,024 x e_mac x k_mma`; `k_mma` near 1 with a floor near 0.5 (approximate); rows from the measured `e_mac` in 5.3 | 2.06 + about 0.9 microseconds per step per unit scalar (R = 128: +0.9 ms; R = 512: +3.7 ms); VNNI 4x to 16x less | native on sm_75+ and RDNA 3+ (AMD layout unverified); emulated on Pascal, RDNA 2, Apple | before measurement: prototype further. After section 5: NEVER as class content (the block costs the honest card 0.056 to 0.70 pJ per multiply-add, so it forces no joules on a chip); kept as reserve R8 evidence |
|
||||
| C, stored state | the daily dataset derives from the execution state snapshot; kernel unchanged; a state sample per block | unchanged (5.1x GDDR7, 7.5x to 9.2x HBM3); the f = 0 chip disappears | 2.06 + 4,096 leaf reads (section 5.2) | nothing vendor-specific; the serialisation is the consensus risk | before measurement: prototype further. After section 5: SHIP as the class v5 candidate |
|
||||
|
||||
### 3.5 Migration through the class system (common to B and C)
|
||||
|
||||
|
|
@ -165,7 +166,7 @@ Written by this lane in the persona files' voices (`.claude/agents/cryptographer
|
|||
|
||||
### 4.4 The pick
|
||||
|
||||
B and C are the two to prototype: both are class objects on the shipped hash, both leave the dataset's memory bound untouched, and they compose (B is kernel text, C is the daily build). A is not prototyped: A1 and A2 fail on bytes and on sampleability before any kernel, and A0 is C with a worse data source. The order of merit at this point, before measurement: C first (no vendor risk, unchanged hash rate by construction, a real new property per block, the pool caveat stated), B second (a real lever against the f = 1 chip with a bounded `k`, a vendor split and an unverified AMD layout). Section 6 revisits the order on the measured rows.
|
||||
B and C are the two to prototype: both are class objects on the shipped hash, both leave the dataset's memory bound untouched, and they compose (B is kernel text, C is the daily build). A is not prototyped: A1 and A2 fail on bytes and on sampleability before any kernel, and A0 is C with a worse data source. The order of merit at this point, before measurement: C first (no vendor risk, unchanged hash rate by construction, a real new property per block, the pool caveat stated), B second (a real lever against the f = 1 chip with a bounded `k`, a vendor split and an unverified AMD layout). Section 6 revisits the order on the measured rows: C holds, B retires into the reserve.
|
||||
|
||||
## 5. The prototypes and the measured rows
|
||||
|
||||
|
|
@ -173,31 +174,115 @@ Both prototypes live under `proto-newpow/` with a README carrying the exact comm
|
|||
|
||||
### 5.1 `mma-shadow` (scheme B), box 1
|
||||
|
||||
ROWS_B
|
||||
Prototype: `proto-newpow/mma-shadow/` (README with every command, `kernel_mm8.cu` with the PTX path and the `IGNEUM_MM8_REF` reference path, `bench.cu`, `verify_ref.c`, `gen_block.py`, `gen_ref_program.py`, `run.sh`, `out/` with every log and csv). Box 1: RTX 4090 24 GB (128 SMs), driver 570.172 (the box reports 570, not the 595 of box 2), nvcc 12.8.93, `-arch=sm_89`, 1 warp per block, 10 timed batches of 2^24 after a warm-up, power from `nvidia-smi` at 1 Hz over a 25-s sustained phase (mean after its first 10 s), idle 15.0 to 15.3 W. The two-output tile form of section 3.2 was built (the single-output form never was). Run 19:36 to 19:52 UTC.
|
||||
|
||||
| R (steps per iteration) | mm8 per hash (8 R) | MH/s (GPU time) | Ratio to R = 0 | Watts, mean | Block watts over R = 0 | SM MHz | Microjoules per hash | Marginal pJ per multiply-add | Fingerprint (2^24 at base 0) | PTX = reference path | CPU = GPU (1,024 lanes) | Verifier block delta, ms per unit (box core, scalar C) | Registers per thread |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| 0 (control) | 0 | 63.08 | 1.000 | 201.2 | 0 | 2,670 | 3.19 | | 7c28cfb06c5c65a9 (the pack's; the 3 pack vectors PASS standalone and in batch) | yes | 1,024 of 1,024 | 0 (the R = 0 run read -0.10, noise) | 29 |
|
||||
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2.9 | 2,670 | 3.24 | 0.70 | 06fc2593bfb94b4f | yes | 1,024 of 1,024 | +0.05 | 30 |
|
||||
| 32 | 256 | 63.08 | 1.000 | 207.7 | 6.5 | 2,670 | 3.29 | 0.39 | 26e83a65f519c865 | yes | 1,024 of 1,024 | +0.17 | 29 |
|
||||
| 128 | 1,024 | 63.08 | 1.000 | 212.7 | 11.5 | 2,670 | 3.37 | 0.17 | 42223c2113188335 | yes | 1,024 of 1,024 | +1.14 | 32 |
|
||||
| 512 | 4,096 | 63.08 | 1.000 | 215.9 | 14.7 | 2,670 | 3.42 | 0.056 | 02b7002d747f3711 | yes | 1,024 of 1,024 | +4.39 | 36 |
|
||||
|
||||
Sustained rates over the power window 63.07 to 63.03 MH/s at every R; wall and event rates agree to 0.01 MH/s; no spills, no stack, temperature 58 to 64 C, no throttling. The verifier totals in the box's logs (about 10 ms per unit) are a naive scalar interpreter with a lazy `mh_word` per load on a shared Zen 4 core at load 13 and are not comparable to the project's Rust verifier (2.06 ms per unit on an M5 Max core, interleaved); the block delta is the number: about 1.1 microseconds per `mm8` step per unit, 32 lanes x 32 byte products in plain loops. (At R = 512 the card issues about 258 G tiles per second, 2.6 x 10^14 multiply-adds per second, in the region of 40 percent of the 4090's dense int8 tensor peak, approximate from the published TOPS figure, so R = 512 is near the top of the free band, not its middle.)
|
||||
|
||||
What the rows say:
|
||||
|
||||
1. The tensor block is free in hash rate to R = 512 on the 4090: 4,096 tile instructions per hash leave the rate at 63.08 MH/s to the third decimal. The kernel is latency-bound on its 128 dependent loads and the tensor work fills stalls that were already there, as the ALU shadow did on the 5090 to 150,800 ops (`latency-shadow-2026-10-06.md` 5).
|
||||
2. The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised), 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op, `latency-shadow-2026-10-06.md` 5, item 4); the tensor block at its free-band ceiling costs a third of that.
|
||||
3. That is the finding, and it is negative for the scheme's purpose (section 5.3): a shadow lever moves the chip's edge only by the joules it makes the HONEST card spend on work the chip cannot do more cheaply. The tensor path is so efficient on the GPU that the block adds 0.23 microjoules at most, so at `k = 1` the chip's edge falls from 6.9x to 4.9x on GDDR7 against this 4090, where the ALU shadow took the 5090 from 5.6x to 2.1x, and the tensor block costs the verifier 26x more per unit of chip-forcing energy (4.39 ms scalar per 0.23 microjoules against 0.17 ms per 0.6 microjoules). The property the design hoped for (`k_mma` near 1 because the GPU's tensor engine is near the floor) is real and is exactly why the lever is weak: there are no joules in it to force.
|
||||
4. The correctness chain holds at every rung: the PTX fragment read and the plain-integer reference agree on all 2^24 lanes at every R, and the CPU interpreter matches the GPU on 1,024 lanes at every R; the probe's fragment layout (`family-probe.cu` mm8 `warp_ref`) was used as written and needed no correction. Registers 29 to 36, occupancy unchanged. This is the first class-shaped evidence that an `mm8` family is cheap and exact for the honest NVIDIA card, which is what the reserve entry R8 needs; it is not evidence for a class v5.
|
||||
|
||||
Consequences per tier (B at R = 128, the largest R whose scalar verifier delta stays near 1 ms):
|
||||
|
||||
| Tier | Meaning | What to do |
|
||||
|---|---|---|
|
||||
| NVIDIA Turing and later (RTX 20 to 50 series), any memory size | rate unchanged, +11.5 W on a 4090 (+5.7 percent), per-pound unchanged, per-watt 0.95x | nothing; the block would not be adopted on this evidence |
|
||||
| NVIDIA Pascal (GTX 10 series) | emulation at about 16 steps per `mm8` (approximate): 1,024 per hash is about 16,000 counted steps, inside the latency shadow of a 1080-class card by the counted-ops rule, approximate; unmeasured | the emulation path exists in the kernel text (`IGNEUM_MM8_REF`) and is bit-exact; a GTX 1080 row is owed if B ever proceeds |
|
||||
| AMD RDNA 3 and 4 | the WMMA layout is unverified (status item 6); emulation otherwise at about 12 to 16 steps per `mm8` | the gate of section 7 rank 2 stays open; not worth running unless B proceeds |
|
||||
| Apple | emulation at about 10 steps per `mm8` (approximate), 10,000 counted steps at R = 128 inside the M5 Max's 130,000 ceiling | owed; not worth running unless B proceeds |
|
||||
| Rig, pool user | nothing changes in shares; a rig pays the block's watts per card | nothing |
|
||||
| Node operator (verifier) | +1.14 ms per unit scalar at R = 128 (+55 percent of 2.06 ms), +4.39 ms at R = 512; a SIMD byte-dot path would cut it 4x to 16x (approximate) | the verifier cost per joule forced is the reason the scheme loses to the ALU shadow |
|
||||
| A chip | must carry a tensor array per 32 lanes in flight; at the measured block energy it pays at most 0.23 microjoules per hash more than today at `k = 1`, 0.07 at `k = 0.3` | the chip's edge barely moves (5.3) |
|
||||
|
||||
### 5.2 `state-dataset` (scheme C), box 2 and igneum-build-1
|
||||
|
||||
ROWS_C
|
||||
Prototype: `proto-newpow/state-dataset/` (README with every command and log; `kernel_sd.cu` = the pack's kernel plus `igneum_leaves` and `igneum_build_sd`, the hash kernel byte for byte the pack's; `verify_sd.c` the 32-lane C interpreter and the item check; `cpu_rows.c` the igneum-build-1 rows). Box 2: RTX 4090 24 GB, driver 595.91, nvcc 12.8, host EPYC 7352; igneum-build-1 EPYC 9454P at load 19 to 32 (shared), one core pinned, nice 19. Run 19:34 to 19:50 UTC. The leaf stand-in: one ChaCha12 block of (sigma, S, t, 0, tag) with `S = K XOR 0x5a5a5a5a` (the README defines it).
|
||||
|
||||
| Row | Control (mx8-genesis) | sd1 | Reading |
|
||||
|---|---|---|---|
|
||||
| Hash rate, 10 batches of 2^24, two passes | 63.083, 63.083 MH/s | 63.088, 63.087 MH/s | equal within 0.01 percent: the kernel is the same binary over a dataset of the same shape |
|
||||
| Watts, mean after 10 s of a 20-s window; SM MHz | 207.6, 205.2 W; 2,745 | 207.9, 206.5 W; 2,745 | equal within the 2 W run-to-run noise; 3.27 microjoules per hash on this 4090 (the 5090 is 2.34 to 2.65) |
|
||||
| MH/s per W | 0.304, 0.307 | 0.303, 0.305 | unchanged |
|
||||
| Cache fill, GPU | 1.81, 1.85 ms | 1.85, 1.85 ms | unchanged |
|
||||
| Leaf array (2^24 ChaCha12 blocks, 1 GiB) on the GPU | none | 5.49 ms | the synthetic stand-in; a real snapshot comes from the node |
|
||||
| Dataset build, second pass | 30.55 ms | 31.98 ms | +1.43 ms (+4.7 percent): one coalesced 64 B read per item |
|
||||
| Chunked build (leaves streamed from pinned host memory in 64 MiB or 256 MiB chunks, one stream) | | 75.8 ms, 75.5 ms | PCIe-bound: 62 ms of copy (17.2 GB/s on this box) plus the build, partly serialised; two streams would hide most of the build (not done) |
|
||||
| Device memory during the build (context 395 MiB included) | 1,675 MiB | 2,699 MiB resident leaves; 1,741 MiB chunked at 64 MiB; 1,933 at 256 MiB | while hashing 1,803 MiB in every mode (leaves freed) |
|
||||
| 2^24 fingerprint at base 0 | 7c28cfb06c5c65a9 (the pack's, both passes) | d5b0c16390cad0e8 (four runs) | |
|
||||
| Pack vectors (3 warps, standalone and in batch); dataset head, word [MASK], 64 Mac samples | PASS | n/a (new values; sd1 base 0 lane 0 = b600edbed969becc) | the harness is the pack's |
|
||||
| Build bit-exact: 1,024 random items, GPU words against the plain C host derivation | 1,024 of 1,024 | 1,024 of 1,024 (also after both chunked rebuilds); 64 device leaves = host leaves | |
|
||||
| Hash bit-exact: 4 dumped warps (128 lanes) against the C interpreter with lazy `mh_word` / `mh_word_sd` | 128 of 128 (and 32 of 32 against the Mac vector at base 0) | 128 of 128 | |
|
||||
| Host cache FNV-1a 64 | 48c4f5bf24166b2e = the Mac's | same | |
|
||||
|
||||
The CPU rows (igneum-build-1, ms per unit of 4,096 items, 100 units; two readings at load 19 and 31):
|
||||
|
||||
| Row | Reading 1 | Reading 2 | Per item |
|
||||
|---|---|---|---|
|
||||
| (i) 4,096 random 64 B reads from a 2 GiB resident leaf array | 0.108 ms (0.092 to 0.172) | 0.163 ms | 26 to 40 ns |
|
||||
| (i) the same from an 8 GiB array (the year-12 size) | 0.139 ms | 0.209 ms | 34 to 51 ns |
|
||||
| (ii) 4,096 leaves derived on the fly (one ChaCha12 block each) | 0.284 ms | 0.285 ms | 69 ns |
|
||||
| (iii) 4,096 x `mh_item` on the host cache, naive, no interleaving | 9.01 ms | 11.24 ms | 2.2 to 2.7 microseconds |
|
||||
| 4,096 x (leaf read + `mh_item_sd`), naive | 9.31 ms | 11.64 ms | |
|
||||
| Daily leaf array on CPU (2^24 blocks): one core / 32 threads | 1.67 s / 0.073 s | 1.68 s / 0.083 s | 100 ns per leaf |
|
||||
| Host cache fill (256 MiB), one core | 0.447 s | 0.452 s | |
|
||||
| Light-client sample: distinct items touched by lane 0's 128 loads | 128 of 2^24 | | 128 openings x 25 levels x 32 B + 128 x 64 B = 108 KiB per lane; 3.4 MiB per 32-lane unit before dedup (arithmetic) |
|
||||
|
||||
Reading, against the gate: the project's Rust verifier does the 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items (the naive C figure here, 9 to 11 ms, is an upper bound and is not the production verifier); sd1 adds 0.11 to 0.21 ms per unit when the verifier holds the state in RAM (+5 to +10 percent of 2.06 ms; about +0.3 to +0.5 ms on a 2019-class core by the 2.5x rule), so the class v3 verifier plus sd1 sits at about 2.2 ms on the M5 Max core and about 5.7 ms on the 2019-class core, inside the 10 ms gate with the x8 margin intact. The snapshot's own cost is the state read on the node, not the leaf hashing (1.7 s on one core for 2^24 leaves).
|
||||
|
||||
Consequences per tier (sd1):
|
||||
|
||||
| Tier | Meaning | What to do |
|
||||
|---|---|---|
|
||||
| Home miner, one 8 GB card | hash rate, watts and MH/s per W unchanged (measured equal); the build must stream the leaves (resident leaves at the designed 2 GiB dataset would peak at about 4.6 GiB, approximate; chunked 64 MiB about 2.7 GiB, approximate, measured 1,741 MiB at 1 GiB), so the 8 GB mine-and-prove profiles of `prover-tiers-real-cards.md` keep their memory | ship chunked only; the kernel already takes an item range, so chunking needed no kernel change |
|
||||
| 12, 16, 24, 32 GB cards | unchanged in rate and watts; the daily build +1.4 ms resident or about 76 ms streamed (about 150 ms at 2 GiB, approximate) | nothing |
|
||||
| AMD and Apple | the hash kernel is untouched, so the measured class v3 rates stand (18.9 MH/s on the 9070 XT, 27.1 on the M5 Max); the build kernels gain one 64 B read per item in OpenCL and Metal | the two build kernels in the emitter's other dialects |
|
||||
| Rig | per card as above; one 2 GiB snapshot per day per rig, shared by its cards | the node beside the rig serves it over loopback |
|
||||
| Pool user | a 2 GiB download per day from the pool, or the pool ships the built dataset; today a pool user needs nothing but the day key | the pool protocol (spec 09) gains a daily snapshot or dataset fetch |
|
||||
| Solo miner | must run a node with execution (it already does for templates) and hold the state | nothing new beyond RAM |
|
||||
| Node operator | +2 GiB RAM for the snapshot (+4 GiB at year 4), a daily serialisation pass, the verifier +0.11 to 0.21 ms per unit | the node RAM budget stays inside a 16 GB home node at launch sizes |
|
||||
| Light client, holder | a verifiable sample of 128 state leaves per lane per block exists; checking it costs 108 KiB per lane fetched on demand, so it is an availability spot check, not a light verification of PoW | rank 6 of section 7 |
|
||||
| Prover, rollup customer | nothing | |
|
||||
|
||||
### 5.3 The chip rows on the measured numbers
|
||||
|
||||
CHIP_ROWS
|
||||
Method: `sim/horizon/new-pow/chip_rows.py` (the chip-model-v3 section 5.4 f = 1 chip plus the block's measured energy times `k`; the card figures are 5.1's; every chip figure is arithmetic on `chip-model-v3.md`'s cited and approximate memory figures). The denominator here is the 4090 measured tonight (3.19 microjoules per hash at the control, a card whose memory system is half the 5090's), not the 5090, so the R = 0 column reads 6.9x where `chip-model-v3.md` reads 5.1x for the 5090; the movement with R is the row that matters.
|
||||
|
||||
| Scheme and setting | Honest card microjoules per hash | Chip, GDDR7, k = 1 / 0.5 / 0.3 (microjoules) | Edge over the card, GDDR7 | Chip, one HBM3 stack, k = 1 / 0.5 / 0.3 | Edge, HBM3 |
|
||||
|---|---|---|---|---|---|
|
||||
| Class v3 today (R = 0, the 4090 control) | 3.19 | 0.466 | 6.9x | 0.321 | 9.9x |
|
||||
| B at R = 128 (1,024 tiles per hash, block 0.18 microjoules) | 3.37 | 0.646 / 0.556 / 0.520 | 5.2x / 6.1x / 6.5x | 0.501 / 0.411 / 0.375 | 6.7x / 8.2x / 9.0x |
|
||||
| B at R = 512 (4,096 tiles per hash, block 0.23 microjoules, the free band's top) | 3.42 | 0.696 / 0.581 / 0.535 | 4.9x / 5.9x / 6.4x | 0.551 / 0.436 / 0.390 | 6.2x / 7.8x / 8.8x |
|
||||
| For comparison, the ALU shadow at N = 100,000 on the 5090 (`latency-shadow-2026-10-06.md` 6, measured 3.27 microjoules at the 431 W cap) | 3.27 | 1.59 / 1.03 / 0.80 | 2.1x / 3.2x / 4.1x | 1.44 / 0.88 / 0.66 | 2.3x / 3.7x / 5.0x |
|
||||
| C (sd1) at any size | unchanged (measured equal) | 0.466 | unchanged | 0.321 | unchanged; the f = 0 recompute chip must add 2 GiB of DRAM and becomes this chip |
|
||||
|
||||
Reading: B moves the f = 1 chip's edge by 1.4x to 2x at `k = 1` and by 1.1x at `k = 0.3`, at a verifier cost of 1.1 to 4.4 ms per unit (scalar); the ALU shadow moves it by 2.7x at `k = 1` and 1.4x at `k = 0.3` at 0.17 ms per unit. On every axis the measured tensor block is the weaker lever, for the reason stated in 5.1 item 3. C moves nothing and claims nothing against the chip; its merit is elsewhere.
|
||||
|
||||
## 6. Verdicts
|
||||
|
||||
| Scheme | Verdict | Why, in one line |
|
||||
|---|---|---|
|
||||
| A, mining is proving | NEVER (A1, A2); A0 folds into C | one proof per segment is not a distribution of puzzles; the bytes (2.9 MB of openings per block) or the verify (32 to 40 ms) kill every form that is not "hold the trace", and holding the trace is C with a worse data source |
|
||||
| B, tensor-shaped shadow | VERDICT_B |
|
||||
| C, stored state | VERDICT_C |
|
||||
| B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joules |
|
||||
| C, stored state | SHIP AS CLASS v5 CANDIDATE (through the spec items of 4.3 and the Devnet 2 gate): hash rate and watts unchanged by construction and measured equal, build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact on 1,024 items and 128 lanes; a new property per block (a random sample of state) and a new requirement per mining operation (hold the state); the open decision is what a header verifier is asked to hold |
|
||||
|
||||
**A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.
|
||||
|
||||
VERDICT_BC_PROSE
|
||||
NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joulesC_PROSE
|
||||
|
||||
## 7. Ranked next steps
|
||||
|
||||
Hours are agent hours (the project lead's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own.
|
||||
Hours are agent hours (the project lead's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own. Ranks 2, 3 and 4 were written as B's gates before the ladder landed; after section 5 they are WITHDRAWN (B is not carried forward as class content; the rows stay so the reasoning is visible) and the live order is 1, 5, 6, 7, 8.
|
||||
|
||||
| Rank | Proposal | Evidence | Model | Hours | Consequence per tier | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
|
|
@ -236,4 +321,8 @@ One paragraph each.
|
|||
|
||||
## 9. Summary for the coordinator
|
||||
|
||||
SUMMARY
|
||||
This lane wrote three candidate proofs of work for Igneum, reviewed each in two personas, prototyped the two that survived on real cards, and measured. (1) Scheme C, the dataset derived from the chain's execution state, is the one to carry forward as the class v5 candidate: measured on an RTX 4090 the hash rate and watts are unchanged to the third decimal (63.08 MH/s, 207 W, the kernel is byte for byte the shipped one), the daily build grows by 1.4 ms resident or about 76 ms streamed, the verifier by 0.11 to 0.21 ms per unit with the state in RAM, every GPU item and lane agrees with the plain C derivation (1,024 of 1,024, 128 of 128); it forces every mining operation to hold state, names 128 random state leaves per lane per block, and removes the f = 0 recompute chip as a category; its costs are a 2 GiB daily delivery for pool miners and three spec items (canonical serialisation, a pruning-proof witness, the pause rule). (2) Scheme B, a tensor-shaped shadow of int8 8x8x16 tiles, is bit-exact on NVIDIA at every rung (2^24 lanes PTX = reference, 1,024 lanes CPU = GPU, R = 8 to 512) and free in hash rate to 4,096 tiles per hash, and that is exactly why it fails as an anti-chip lever: it costs the honest card 0.056 to 0.70 pJ per multiply-add and at most 0.23 microjoules per hash, so the f = 1 chip's edge moves from 6.9x to 4.9x at k = 1 against the 4090 where the class v4 ALU shadow moved the 5090's from 5.6x to 2.1x, at 26x the verifier cost per joule forced; never as class content, kept as the evidence and the two-output correction for reserve family R8. (3) Scheme A, mining is proving, is never: one proof per segment is not a distribution of puzzles; the useful fraction is bounded by proving work per segment over network hashes per segment (8 percent at 1 GH/s, 0.08 percent at 100 GH/s on this chain's measured figures); every form that is not "hold the trace" dies on 2.9 MB of openings per block or a 32 to 40 ms proof verify, and "hold the trace" is scheme C with a worse data source.
|
||||
|
||||
1. Scheme C costs nothing in hash rate or watts (63.083 against 63.088 MH/s, 205 to 208 W on the 4090, measured) and 0.11 to 0.21 ms per unit of verifier time (measured on igneum-build-1), so class v5 can carry it; the model is `item_sd(t) = item(t) with s ^= leaf(t)` over the day's certified state root.
|
||||
2. The tensor shadow moves the f = 1 chip's per-joule edge by at most 1.4x (6.9x to 4.9x at k = 1, R = 512) for 4.39 ms of scalar verifier per unit, against 2.7x for 0.17 ms from the ALU shadow; the model is `chip_rows.py` on the measured block energy of 0.23 microjoules per hash.
|
||||
3. Mining cannot be proving on a lottery with a 10 ms verifier: the useful fraction is bounded by 5 GPU-seconds of proving per 8-second segment over the network's hashes in that segment, 8 percent at 1 GH/s and falling with the hash rate; the model is that ratio on `prover-tiers-real-cards.md` and bench-log 2582.
|
||||
|
|
|
|||
125
proto-newpow/mma-shadow/README.md
Normal file
125
proto-newpow/mma-shadow/README.md
Normal file
|
|
@ -0,0 +1,125 @@
|
|||
# mma-shadow: prototype kernel for class "mx8+mm8xR" (Horizon lane 8, new proof of work)
|
||||
|
||||
Measured 6 October 2026 on GPU box 1 (RTX 4090 24 GB). Everything here lives in this directory; the pack files
|
||||
(kernel.cu, memhard.h, program.h, vectors.h) are verbatim copies of proto-cuda/packs-ca2-mixer/mx8-genesis.
|
||||
|
||||
## The scheme
|
||||
|
||||
The shipped mx8 hash (8 iterations of 64 straight-line instructions over r0..r7, 16 dataset loads, 8 shuffles, the
|
||||
fold) plus a block of R `mm8` steps at the end of every iteration, after instruction 63 and before the next
|
||||
iteration samples `sel`, so 8 x R mm8 steps per hash. Nothing else changes. One mm8 step k with drawn registers
|
||||
(a_k, b_k, c_k, c2_k), a_k != b_k, c_k != c2_k: the 32 lanes' r[a] form A (8 x 16 u8, lane l holds
|
||||
A[l >> 2][4 (l & 3) .. +3], byte 0 = lowest k), the 32 lanes' r[b] form B (16 x 8 u8, lane l holds
|
||||
B[4 (l & 3) .. +3][l >> 2]), C = A x B exact in int32, and lane l does r[c] += C[l >> 2][2 (l & 3)] and
|
||||
r[c2] += C[l >> 2][2 (l & 3) + 1], both modulo 2^32. That is the PTX `mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32`
|
||||
fragment layout with a zero accumulator (d0 into r[c], d1 into r[c2]), the same arithmetic as the mm8 `warp_ref` in
|
||||
proto-cuda/family-probe.cu. Both tile outputs are consumed per step (design correction received 6 Oct 2026 before
|
||||
anything was measured, so the single-output `bit` form was never built). The draws come from a SplitMix64 stream
|
||||
seeded with FNV-1a-64 of "igneum-mm8/igneum-genesis" (seed 0x79f1fc5b6ed6112e): a = below(8); b = below(7),
|
||||
b += (b >= a); c = below(8); c2 = below(7), c2 += (c2 >= c). R is the compile-time macro `IGNEUM_MM8_R`; R = 0 is
|
||||
the control and is bit-exact with the pack (3 vector warps and the 2^24 fingerprint 7c28cfb06c5c65a9).
|
||||
|
||||
## Files
|
||||
|
||||
| File | What |
|
||||
|---|---|
|
||||
| `gen_block.py` | draws the 512-step table, writes `mm8_block.h` (packed uint16 table + X-macro step list) |
|
||||
| `kernel_mm8.cu` | the pack's kernel.cu with r0..r7 as `uint32_t r[8]` (mechanical rewrite, same arithmetic) and the mm8 block; PTX path by default, `-DIGNEUM_MM8_REF` for the shuffle-and-byte-product reference path (12 shuffles + 32 byte products per lane per step). Cache fill, build and launch wrappers unchanged |
|
||||
| `bench.cu` | harness derived from proto-cuda/host.cu (serve mode stripped): device info, `igneum_hash_info` registers and occupancy, GPU cache fill + host fill + FNV check, GPU dataset build + self-test, pack vectors at R = 0, 2^24 fingerprint at base nonce 0, 1 warm-up + 10 timed batches (CUDA events), `--sustain S` for the power meter, `--dump file n` |
|
||||
| `gen_ref_program.py` | turns the 64 instruction lines of kernel.cu into the 32-lane C interpreter body `ref_program.inc` (nothing transcribed by hand) |
|
||||
| `verify_ref.c` | plain C CPU reference: host cache (65536 segments), lazy `mh_word` loads, register-major interpreter, mm8 block in the spec layout, fold; compares a dump, then times the verifier per unit |
|
||||
| `run.sh` | the ladder R in {0, 8, 32, 128, 512}: builds, power sampling, fingerprints, dumps, CPU check, summary |
|
||||
| `summarise.py` | builds the RESULTS table from `out/` |
|
||||
|
||||
## Exact commands (on the box, under /root/horizon-newpow/mma-shadow)
|
||||
|
||||
```
|
||||
export PATH=/usr/local/cuda/bin:$PATH
|
||||
python3 gen_block.py mm8_block.h
|
||||
python3 gen_ref_program.py kernel.cu
|
||||
gcc -O2 -o verify_ref verify_ref.c
|
||||
for R in 0 8 32 128 512; do
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -Xptxas -v -o bench_$R bench.cu kernel_mm8.cu
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -DIGNEUM_MM8_REF -o bench_${R}_ref bench.cu kernel_mm8.cu
|
||||
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu,timestamp --format=csv,noheader -l 1 > out/power_$R.csv &
|
||||
./bench_$R --batches 10 --sustain 25 --dump out/dump_$R.txt 32 > out/bench_$R.log
|
||||
kill %1
|
||||
./bench_${R}_ref --fingerprint-only --no-host-cache > out/bench_${R}_ref.log
|
||||
taskset -c 2 ./verify_ref out/dump_$R.txt $R --time > out/verify_$R.log # add --self-test at R = 0
|
||||
done
|
||||
python3 summarise.py > out/results.md
|
||||
```
|
||||
`./run.sh` is exactly that sequence (plus the idle baseline and the PTX == ref fingerprint comparison). From the
|
||||
Mac: `rsync -az -e "ssh -i ~/.ssh/igneum-fleet -p <box-1-port>" proto-newpow/mma-shadow/ root@<box-1-ip>:/root/horizon-newpow/mma-shadow/`
|
||||
then `touch *` on the far side before building.
|
||||
|
||||
## Card and driver
|
||||
|
||||
NVIDIA GeForce RTX 4090, 128 SMs, cc 8.9, 24 GB. Driver 570.172.08 (the brief said 595; nvidia-smi reports 570.172.08),
|
||||
CUDA 12.8 (nvcc V12.8.93), g++ 13.3.0, Ubuntu 24.04, 96 CPU threads. Idle baseline before the run: 15.0 to 15.3 W at
|
||||
210 MHz SM, 45 C, 0 % utilisation. The host carried a CPU load average of about 13 from other tenants during the run
|
||||
(the GPU itself was idle and ours alone); that is why the verifier timings were pinned to one core.
|
||||
|
||||
## Method notes
|
||||
|
||||
- Timing: one warm-up batch (also the fingerprint batch) then 10 timed batches of 2^24 hashes between CUDA events,
|
||||
1 warp per block (the pack bench default, 24 resident warps per SM at 29 registers). Power: nvidia-smi at 1 Hz
|
||||
during a 25 s sustained phase after the timed batches; the mean takes samples from 10 s after the sustain start to
|
||||
its end (the timestamps are in the csv). Microjoules per hash = mean watts / (MH/s x 10^6) x 10^6.
|
||||
- Fingerprint = FNV-1a 64 over the 2^24 little-endian u64 outputs at base nonce 0. The PTX build and the reference
|
||||
build must agree (the reference is the plain-integer byte-product form of the same fragment layout).
|
||||
- CPU == GPU: 32 warps at SplitMix64(0x1234) 32-aligned bases, 1,024 lanes, recomputed by `verify_ref`.
|
||||
- The verifier here is a naive register-major interpreter with no interleaving and a lazy `mh_word` per load
|
||||
(72 mixer applications and 8 cache reads per word, 4,096 words per unit). It is slower than the project's Rust
|
||||
verifier (2.06 ms per unit on an M5 Max core) on this box's core, so the number that matters is the mm8 block
|
||||
delta (ms per unit at R minus ms per unit at R = 0), not the total.
|
||||
|
||||
## RESULTS (6 Oct 2026, RTX 4090, driver 570.172.08, CUDA 12.8, sm_89, 1 warp/block, batch 2^24 x 10)
|
||||
|
||||
| R | mm8 per hash (8R) | MH/s (GPU time) | ratio to R = 0 | watts mean (sustain, after first 10 s) | SM MHz | uJ per hash | max C | fingerprint (2^24 at base 0) | PTX == ref | CPU == GPU lanes | verifier ms per unit on the box core: R = 0, R, block delta | regs per thread (PTX build) | regs (ref build) | blocks per SM (cudaOccupancy) |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 (matches the pack) | yes | 1024 of 1024 | 10.18, 10.08, -0.10 (noise) | 29 | 29 | 24 (24 warps/SM, 50 %) |
|
||||
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.93, 9.97, +0.05 | 30 | 77 | 24 |
|
||||
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.08, 10.24, +0.17 | 29 | 151 | 24 |
|
||||
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.18, 11.33, +1.14 | 32 | 175 | 24 |
|
||||
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.89, 14.28, +4.39 | 36 | 213 | 24 |
|
||||
|
||||
Sustained MH/s over the 25 s power window: 63.075, 63.074, 63.072, 63.034, 63.073 for R = 0, 8, 32, 128, 512 (wall,
|
||||
including a sync per batch). Wall and GPU-event rates agree to 0.01 MH/s at every R. Idle baseline 15.0 to 15.3 W.
|
||||
Registers per thread at R = 0, 8, 32, 128, 512: 29, 30, 29, 32, 36 (PTX path), no spills, no stack, at every R.
|
||||
R = 0 self-checks: cache FNV 48c4f5bf24166b2e PASS (GPU == host fill word for word), dataset head, [MASK], 64 random
|
||||
points and 64 Mac samples PASS, the pack's 3 vector warps PASS standalone and in batch, fingerprint 7c28cfb06c5c65a9.
|
||||
The C interpreter also reproduces the 3 pack vectors at R = 0 (`--self-test`).
|
||||
|
||||
## What the numbers say
|
||||
|
||||
- On the RTX 4090 the mm8 block is free in hash rate up to R = 512: 4,096 tensor instructions per hash leave the
|
||||
rate at 63.08 MH/s, identical to the control to the third decimal. The kernel is latency-bound on the 16 dependent
|
||||
random loads per iteration; the tensor work fills stalls that were already there. At R = 512 the card issues about
|
||||
2.6 x 10^11 mma.m8n8k16 per second, roughly 2.7 x 10^14 u8 multiply-adds per second (1,024 per instruction), which is
|
||||
in the region of 40 % of the card's dense int8 tensor peak (approximate, from the published TOPS figure). So the
|
||||
next doublings would start to cost hash rate; R = 512 is near the top of the free band, not in the middle of it.
|
||||
- Power is the only GPU cost that moves: 201 W to 216 W (+7.3 %) and 3.19 to 3.42 uJ per hash (+7.2 %) from R = 0
|
||||
to R = 512. SM clock stayed pinned at 2670 MHz at every R, temperature peaked at 64 C, no throttling seen.
|
||||
- Correctness chain holds at every rung: the PTX fragment read and the plain-integer reference path agree on all
|
||||
2^24 lanes of the fingerprint batch at every R, and the CPU interpreter matches the GPU on all 1,024 dumped lanes.
|
||||
The probe's fragment layout (family-probe.cu `warp_ref`, mm8) was used as written and needed no correction.
|
||||
- Verifier cost of the block on the box's core: about 1.1 us per mm8 step per unit (32 lanes x 32 byte products,
|
||||
naive loops), so +1.14 ms per unit at R = 128 and +4.39 ms per unit at R = 512. Against the project's Rust verifier
|
||||
at 2.06 ms per unit, R = 512 would roughly triple verification time unless the block is vectorised (the 8 x 8 x 16
|
||||
tile is 1,024 MACs, a few hundred ns with SIMD, approximate); R = 32 adds 0.17 ms per unit (+8 % of 2.06 ms) and
|
||||
R = 128 adds 1.14 ms (+55 %). The totals in the table (about 10 ms per unit) are this naive interpreter's cost
|
||||
with a lazy mh_word per load and are not comparable to the Rust verifier; the delta column is the number to use.
|
||||
- The reference-path build is only a correctness oracle: fully unrolled it reaches 213 registers at R = 512 (8 blocks
|
||||
per SM) and was never timed.
|
||||
|
||||
## Raw outputs
|
||||
|
||||
`out/` holds everything the box produced: `run.log` (the whole run), `bench_R.log` and `bench_R_ref.log`,
|
||||
`power_R.csv` (1 Hz: W, SM MHz, C, util, timestamp), `power_idle.csv`, `ptxas_R.txt` and `ptxas_R_ref.txt`,
|
||||
`dump_R.txt` (32 warps, 1,024 lanes), `verify_R.log`, `results.md` (summarise.py).
|
||||
|
||||
## What was cut
|
||||
|
||||
Nothing from the brief. The host 1 GiB dataset was never built on the CPU (lazy `mh_word`, as allowed). The 32-lane
|
||||
dump bases are 32-bit, so base + 31 cannot overflow (base = low32(next()) & ~31).
|
||||
308
proto-newpow/mma-shadow/bench.cu
Normal file
308
proto-newpow/mma-shadow/bench.cu
Normal file
|
|
@ -0,0 +1,308 @@
|
|||
// bench.cu (proto-newpow/mma-shadow): host harness for the mx8+mm8xR prototype, derived from proto-cuda/host.cu
|
||||
// with serve mode stripped. TEST HARNESS ONLY: no pool, no network, no wallet.
|
||||
//
|
||||
// Build: nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=<R> [-DIGNEUM_MM8_REF] -o bench_<R> bench.cu kernel_mm8.cu
|
||||
// Steps: device info, kernel registers and occupancy, GPU cache fill + host cache fill + check (FNV, head, last),
|
||||
// GPU dataset build + self-test (head, [MASK], 64 random points vs host mh_word, Mac samples), the pack's 3 vector warps
|
||||
// (R == 0 only; they must PASS), the 2^24 batch fingerprint at base nonce 0 (FNV-1a 64 over the output bytes), timing
|
||||
// (1 warm-up + N timed batches with CUDA events), an optional sustain phase for the power meter, and --dump.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <chrono>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
#include "program.h"
|
||||
#include "vectors.h"
|
||||
#include "memhard.h"
|
||||
#include "mm8_block.h"
|
||||
#ifndef IGNEUM_MM8_R
|
||||
#define IGNEUM_MM8_R 0
|
||||
#endif
|
||||
|
||||
#define CUDA_CHECK(call) do { cudaError_t err_ = (call); if (err_ != cudaSuccess) { \
|
||||
std::fprintf(stderr, "CUDA error: %s (%d)\n at %s:%d\n in %s\n", cudaGetErrorString(err_), (int)err_, __FILE__, __LINE__, #call); \
|
||||
std::exit(2); } } while (0)
|
||||
|
||||
static double wallMs() {
|
||||
using namespace std::chrono;
|
||||
return duration<double, std::milli>(steady_clock::now().time_since_epoch()).count();
|
||||
}
|
||||
static double epochS() {
|
||||
using namespace std::chrono;
|
||||
return duration<double>(system_clock::now().time_since_epoch()).count();
|
||||
}
|
||||
static uint64_t fnv1a64(const void* p, size_t n) {
|
||||
const uint8_t* b = (const uint8_t*)p;
|
||||
uint64_t h = 0xcbf29ce484222325ull;
|
||||
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
|
||||
return h;
|
||||
}
|
||||
static uint64_t splitmix64(uint64_t& s) {
|
||||
s += 0x9E3779B97F4A7C15ull;
|
||||
uint64_t z = s;
|
||||
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull;
|
||||
z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
|
||||
return z ^ (z >> 31);
|
||||
}
|
||||
|
||||
static const uint32_t CACHE_WORDS = 1u << IGNEUM_CACHE_LOG2_WORDS;
|
||||
static uint32_t* gCache = nullptr;
|
||||
static std::vector<uint32_t> hCache;
|
||||
|
||||
static bool setupCache(bool hostFill) {
|
||||
size_t bytes = (size_t)CACHE_WORDS * 4u;
|
||||
CUDA_CHECK(cudaMalloc((void**)&gCache, bytes));
|
||||
cudaEvent_t e0, e1;
|
||||
CUDA_CHECK(cudaEventCreate(&e0)); CUDA_CHECK(cudaEventCreate(&e1));
|
||||
float ms[2] = {0.f, 0.f};
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(igneum_launch_cache_fill(gCache, IGNEUM_CACHE_SEGMENTS));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
CUDA_CHECK(cudaEventElapsedTime(&ms[pass], e0, e1));
|
||||
}
|
||||
std::printf("cache fill (GPU): %.2f ms first, %.2f ms second (%u MiB)\n", ms[0], ms[1], (unsigned)(bytes >> 20));
|
||||
std::vector<uint32_t> dev(CACHE_WORDS);
|
||||
CUDA_CHECK(cudaMemcpy(dev.data(), gCache, bytes, cudaMemcpyDeviceToHost));
|
||||
uint64_t fnvDev = fnv1a64(dev.data(), bytes);
|
||||
bool fnvOk = fnvDev == IGNEUM_CACHE_FNV64;
|
||||
bool headOk = std::memcmp(dev.data(), IGNEUM_CACHE_HEAD, 64) == 0;
|
||||
bool lastOk = std::memcmp(dev.data() + CACHE_WORDS - 16u, IGNEUM_CACHE_LAST, 64) == 0;
|
||||
bool same = true;
|
||||
double hostMs = 0;
|
||||
if (hostFill) {
|
||||
hCache.assign(CACHE_WORDS, 0u);
|
||||
double h0 = wallMs();
|
||||
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(hCache.data(), seg);
|
||||
hostMs = wallMs() - h0;
|
||||
same = std::memcmp(dev.data(), hCache.data(), bytes) == 0;
|
||||
std::printf("cache fill (host, one thread): %.1f ms, GPU == host all words: %s\n", hostMs, same ? "PASS" : "FAIL");
|
||||
} else {
|
||||
hCache.swap(dev); // the device cache (FNV-checked) serves the host-side dataset derivation
|
||||
}
|
||||
bool pass = fnvOk && headOk && lastOk && same;
|
||||
std::printf("cache check: %s (device FNV-1a 64 %016llx vs Mac %016llx %s, head %s, last line %s)\n",
|
||||
pass ? "PASS" : "FAIL", (unsigned long long)fnvDev, (unsigned long long)IGNEUM_CACHE_FNV64,
|
||||
fnvOk ? "PASS" : "FAIL", headOk ? "PASS" : "FAIL", lastOk ? "PASS" : "FAIL");
|
||||
CUDA_CHECK(cudaEventDestroy(e0)); CUDA_CHECK(cudaEventDestroy(e1));
|
||||
return pass;
|
||||
}
|
||||
|
||||
static bool compareWarp(const uint64_t* got, const uint64_t* want, uint32_t base, const char* how) {
|
||||
int bad = 0, first = -1;
|
||||
for (int l = 0; l < 32; ++l) if (got[l] != want[l]) { if (first < 0) first = l; ++bad; }
|
||||
if (bad == 0) std::printf("verify warp base %u %s: PASS\n", base, how);
|
||||
else std::printf("verify warp base %u %s: FAIL %d of 32 lanes differ, first lane %d: gpu=%016llx expected=%016llx\n",
|
||||
base, how, bad, first, (unsigned long long)got[first], (unsigned long long)want[first]);
|
||||
return bad == 0;
|
||||
}
|
||||
|
||||
static void usage() {
|
||||
std::printf("bench_R [--batches 10] [--block-warps 1] [--batch-log2 24] [--fingerprint-only] [--sustain S] [--no-host-cache]\n"
|
||||
" [--dump <file> <n>] [--device 0]\n"
|
||||
" --fingerprint-only skip the timed batches (for the reference build)\n"
|
||||
" --sustain S after the timed batches keep launching batches for S seconds (power meter window)\n"
|
||||
" --dump file n write n whole warps at SplitMix64(0x1234) 32-aligned bases as lines: base lane value_hex\n"
|
||||
" --no-host-cache skip the one-thread host cache fill (the FNV-checked device cache then feeds the host derivation)\n");
|
||||
}
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
int batches = 10, blockWarps = 1, batchLog2 = 24, device = 0, dumpN = 0;
|
||||
double sustain = 0;
|
||||
bool fpOnly = false, hostCache = true;
|
||||
std::string dumpFile;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
std::string a = argv[i];
|
||||
auto nextInt = [&](int& dst) { if (i + 1 >= argc) { usage(); std::exit(2); } dst = std::atoi(argv[++i]); };
|
||||
if (a == "--batches") nextInt(batches);
|
||||
else if (a == "--block-warps") nextInt(blockWarps);
|
||||
else if (a == "--batch-log2") nextInt(batchLog2);
|
||||
else if (a == "--device") nextInt(device);
|
||||
else if (a == "--fingerprint-only") fpOnly = true;
|
||||
else if (a == "--no-host-cache") hostCache = false;
|
||||
else if (a == "--sustain") { if (i + 1 >= argc) { usage(); return 2; } sustain = std::atof(argv[++i]); }
|
||||
else if (a == "--dump") { if (i + 2 >= argc) { usage(); return 2; } dumpFile = argv[++i]; dumpN = std::atoi(argv[++i]); }
|
||||
else if (a == "-h" || a == "--help") { usage(); return 0; }
|
||||
else { std::printf("unknown argument %s\n", argv[i]); usage(); return 2; }
|
||||
}
|
||||
if (blockWarps < 1 || blockWarps > 32 || batchLog2 < 10 || batchLog2 > 28 || batches < 1) { usage(); return 2; }
|
||||
|
||||
std::printf("mma-shadow bench pack \"%s\" class mx8+mm8xR R = %d mm8 per hash = %d path = %s\n",
|
||||
IGNEUM_SEED_STRING, IGNEUM_MM8_R, 8 * IGNEUM_MM8_R,
|
||||
#ifdef IGNEUM_MM8_REF
|
||||
"reference (shuffles + byte products)"
|
||||
#else
|
||||
"PTX mma.sync.m8n8k16.u8"
|
||||
#endif
|
||||
);
|
||||
std::printf("mm8 table seed 0x%016llx (first steps: %s)\n", (unsigned long long)IGNEUM_MM8_SEED, "see mm8_block.h");
|
||||
|
||||
int count = 0;
|
||||
CUDA_CHECK(cudaGetDeviceCount(&count));
|
||||
if (count == 0 || device >= count) { std::printf("FAIL: no CUDA device %d\n", device); return 2; }
|
||||
CUDA_CHECK(cudaSetDevice(device));
|
||||
cudaDeviceProp prop; std::memset(&prop, 0, sizeof(prop));
|
||||
CUDA_CHECK(cudaGetDeviceProperties(&prop, device));
|
||||
int drv = 0, rt = 0; CUDA_CHECK(cudaDriverGetVersion(&drv)); CUDA_CHECK(cudaRuntimeGetVersion(&rt));
|
||||
int clk = 0, memclk = 0, bus = 0, l2 = 0, thrSM = 0;
|
||||
cudaDeviceGetAttribute(&clk, cudaDevAttrClockRate, device);
|
||||
cudaDeviceGetAttribute(&memclk, cudaDevAttrMemoryClockRate, device);
|
||||
cudaDeviceGetAttribute(&bus, cudaDevAttrGlobalMemoryBusWidth, device);
|
||||
cudaDeviceGetAttribute(&l2, cudaDevAttrL2CacheSize, device);
|
||||
cudaDeviceGetAttribute(&thrSM, cudaDevAttrMaxThreadsPerMultiProcessor, device);
|
||||
cudaGetLastError();
|
||||
std::printf("GPU: %s (%d SMs, cc %d.%d, %.0f MiB) SM clock %d MHz, mem clock %d MHz, bus %d bits, L2 %d MiB, max %d threads/SM\n",
|
||||
prop.name, prop.multiProcessorCount, prop.major, prop.minor, (double)prop.totalGlobalMem / 1048576.0,
|
||||
clk / 1000, memclk / 1000, bus, l2 / 1048576, thrSM);
|
||||
std::printf("CUDA: driver %d.%d, runtime %d.%d\n", drv / 1000, (drv % 100) / 10, rt / 1000, (rt % 100) / 10);
|
||||
int regs = 0, blocksPerSM = 0;
|
||||
CUDA_CHECK(igneum_hash_info(®s, &blocksPerSM, (uint32_t)blockWarps));
|
||||
std::printf("igneum_hash_info: %d registers/thread, %d resident blocks/SM at %d warp(s)/block = %d resident warps/SM (%.1f%% of %d)\n",
|
||||
regs, blocksPerSM, blockWarps, blocksPerSM * blockWarps, 100.0 * blocksPerSM * blockWarps * 32 / thrSM, thrSM / 32);
|
||||
|
||||
bool cachePass = setupCache(hostCache);
|
||||
|
||||
uint32_t nonces = 1u << batchLog2;
|
||||
uint32_t mask = IGNEUM_MASK;
|
||||
uint64_t dsBytes = (uint64_t)(mask + 1u) * 4ull;
|
||||
uint32_t* dDs = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dDs, (size_t)dsBytes));
|
||||
uint64_t* dOut = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dOut, (size_t)nonces * sizeof(uint64_t)));
|
||||
cudaEvent_t e0, e1;
|
||||
CUDA_CHECK(cudaEventCreate(&e0)); CUDA_CHECK(cudaEventCreate(&e1));
|
||||
float buildMs[2] = {0.f, 0.f};
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(igneum_launch_build(dDs, gCache, (mask + 1u) / 16u));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
CUDA_CHECK(cudaEventElapsedTime(&buildMs[pass], e0, e1));
|
||||
}
|
||||
std::printf("dataset build (GPU, %llu MiB): %.2f ms first, %.2f ms second\n", (unsigned long long)(dsBytes >> 20), buildMs[0], buildMs[1]);
|
||||
|
||||
// Dataset self-test
|
||||
bool dsPass;
|
||||
{
|
||||
uint32_t head[16];
|
||||
CUDA_CHECK(cudaMemcpy(head, dDs, sizeof(head), cudaMemcpyDeviceToHost));
|
||||
int badHead = 0;
|
||||
for (int i = 0; i < 16; ++i) if (head[i] != IGNEUM_DS_HEAD[i]) { if (!badHead) std::printf(" dataset[%d] = %08x, Mac %08x\n", i, head[i], IGNEUM_DS_HEAD[i]); ++badHead; }
|
||||
uint32_t last = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&last, dDs + IGNEUM_DS_LAST_INDEX, 4, cudaMemcpyDeviceToHost));
|
||||
bool lastOk = last == IGNEUM_DS_LAST;
|
||||
int badRnd = 0;
|
||||
uint64_t s = 0x9E3779B97F4A7C15ull ^ (uint64_t)(mask + 1u);
|
||||
for (int k = 0; k < 64; ++k) {
|
||||
uint32_t idx = (uint32_t)splitmix64(s) & mask, v = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&v, dDs + idx, 4, cudaMemcpyDeviceToHost));
|
||||
uint32_t want = mh_word(hCache.data(), idx);
|
||||
if (v != want) { if (!badRnd) std::printf(" dataset[%u] = %08x, host derivation %08x\n", idx, v, want); ++badRnd; }
|
||||
}
|
||||
int badSample = 0;
|
||||
for (int k = 0; k < IGNEUM_DS_SAMPLES; ++k) {
|
||||
uint32_t v = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&v, dDs + IGNEUM_DS_SAMPLE_INDEX[k], 4, cudaMemcpyDeviceToHost));
|
||||
if (v != IGNEUM_DS_SAMPLE_VALUE[k]) { if (!badSample) std::printf(" dataset[%u] = %08x, Mac %08x\n", IGNEUM_DS_SAMPLE_INDEX[k], v, IGNEUM_DS_SAMPLE_VALUE[k]); ++badSample; }
|
||||
}
|
||||
dsPass = badHead == 0 && lastOk && badRnd == 0 && badSample == 0;
|
||||
std::printf("dataset self-test: %s (head 16 %s, [MASK] %s, 64 random points vs host derivation %s, %d Mac samples %s)\n",
|
||||
dsPass ? "PASS" : "FAIL", badHead == 0 ? "PASS" : "FAIL", lastOk ? "PASS" : "FAIL",
|
||||
badRnd == 0 ? "PASS" : "FAIL", (int)IGNEUM_DS_SAMPLES, badSample == 0 ? "PASS" : "FAIL");
|
||||
}
|
||||
|
||||
// Pack vectors (R == 0 only: at R > 0 the function is different by design)
|
||||
bool vecPass = true;
|
||||
uint64_t got[32];
|
||||
if (IGNEUM_MM8_R == 0) {
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, IGNEUM_VEC_BASE[w], mask, 32u, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
|
||||
vecPass = compareWarp(got, IGNEUM_VEC_OUT[w], IGNEUM_VEC_BASE[w], "(pack vector, standalone)") && vecPass;
|
||||
}
|
||||
} else {
|
||||
std::printf("pack vectors: not applicable at R = %d (checked at R = 0 only)\n", IGNEUM_MM8_R);
|
||||
}
|
||||
|
||||
// Warm-up batch at base 0 = the fingerprint batch
|
||||
double w0 = wallMs();
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, 0u, mask, nonces, (uint32_t)blockWarps));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
double w1 = wallMs();
|
||||
std::vector<uint64_t> hOut(nonces);
|
||||
CUDA_CHECK(cudaMemcpy(hOut.data(), dOut, (size_t)nonces * 8u, cudaMemcpyDeviceToHost));
|
||||
uint64_t fp = fnv1a64(hOut.data(), (size_t)nonces * 8u);
|
||||
std::printf("warm-up batch: %u hashes in %.2f ms wall\n", nonces, w1 - w0);
|
||||
std::printf("fingerprint: %016llx (FNV-1a 64 over the 2^%d outputs at base nonce 0, little-endian u64 bytes)%s\n",
|
||||
(unsigned long long)fp, batchLog2,
|
||||
IGNEUM_MM8_R == 0 ? (fp == 0x7c28cfb06c5c65a9ull ? " == 7c28cfb06c5c65a9 PASS" : " != 7c28cfb06c5c65a9 FAIL") : "");
|
||||
bool fpPass = (IGNEUM_MM8_R != 0) || fp == 0x7c28cfb06c5c65a9ull;
|
||||
if (IGNEUM_MM8_R == 0 && batchLog2 >= 20) {
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
if ((uint64_t)IGNEUM_VEC_BASE[w] + 32ull > nonces) continue;
|
||||
vecPass = compareWarp(hOut.data() + IGNEUM_VEC_BASE[w], IGNEUM_VEC_OUT[w], IGNEUM_VEC_BASE[w], "(pack vector, in batch)") && vecPass;
|
||||
}
|
||||
}
|
||||
|
||||
double mhs = 0, mhsWall = 0, mhsSustain = 0;
|
||||
if (!fpOnly) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
double t0 = wallMs();
|
||||
for (int b = 1; b <= batches; ++b) {
|
||||
uint32_t base = (uint32_t)((uint64_t)b * (uint64_t)nonces);
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, (uint32_t)blockWarps));
|
||||
}
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
double t1 = wallMs();
|
||||
float gpuMs = 0.f;
|
||||
CUDA_CHECK(cudaEventElapsedTime(&gpuMs, e0, e1));
|
||||
double total = (double)nonces * batches;
|
||||
mhs = total / (gpuMs / 1000.0) / 1e6;
|
||||
mhsWall = total / ((t1 - t0) / 1000.0) / 1e6;
|
||||
std::printf("timed: %d batches x %u hashes: GPU %.2f ms -> %.3f MH/s (GPU time), wall %.2f ms -> %.3f MH/s\n",
|
||||
batches, nonces, gpuMs, mhs, t1 - t0, mhsWall);
|
||||
if (sustain > 0) {
|
||||
double s0 = epochS(), sw0 = wallMs();
|
||||
long long n = 0;
|
||||
std::printf("sustain start epoch %.3f\n", s0); std::fflush(stdout);
|
||||
while (wallMs() - sw0 < sustain * 1000.0) {
|
||||
uint32_t base = (uint32_t)((uint64_t)(n + batches + 1) * (uint64_t)nonces);
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, (uint32_t)blockWarps));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
++n;
|
||||
}
|
||||
double s1 = epochS();
|
||||
mhsSustain = (double)n * nonces / (s1 - s0) / 1e6;
|
||||
std::printf("sustain end epoch %.3f: %lld batches in %.2f s -> %.3f MH/s (wall, incl. sync)\n", s1, n, s1 - s0, mhsSustain);
|
||||
}
|
||||
}
|
||||
|
||||
if (dumpN > 0) {
|
||||
FILE* f = std::fopen(dumpFile.c_str(), "w");
|
||||
if (!f) { std::printf("FAIL: cannot open %s\n", dumpFile.c_str()); return 2; }
|
||||
uint64_t s = 0x1234ull;
|
||||
for (int i = 0; i < dumpN; ++i) {
|
||||
uint32_t base = (uint32_t)splitmix64(s) & ~31u;
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, 32u, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
|
||||
for (int l = 0; l < 32; ++l) std::fprintf(f, "%u %d %016llx\n", base, l, (unsigned long long)got[l]);
|
||||
}
|
||||
std::fclose(f);
|
||||
std::printf("dump: %d warps (%d lanes) written to %s\n", dumpN, 32 * dumpN, dumpFile.c_str());
|
||||
}
|
||||
|
||||
bool overall = cachePass && dsPass && vecPass && fpPass;
|
||||
std::printf("SUMMARY R=%d regs=%d blocksPerSM=%d mhs=%.3f mhs_wall=%.3f mhs_sustain=%.3f fingerprint=%016llx cache=%s dataset=%s vectors=%s\n",
|
||||
IGNEUM_MM8_R, regs, blocksPerSM, mhs, mhsWall, mhsSustain, (unsigned long long)fp,
|
||||
cachePass ? "PASS" : "FAIL", dsPass ? "PASS" : "FAIL", IGNEUM_MM8_R == 0 ? (vecPass ? "PASS" : "FAIL") : "n/a");
|
||||
std::printf("OVERALL: %s\n", overall ? "PASS" : "FAIL");
|
||||
CUDA_CHECK(cudaFree(dOut)); CUDA_CHECK(cudaFree(dDs)); CUDA_CHECK(cudaFree(gCache));
|
||||
return overall ? 0 : 1;
|
||||
}
|
||||
75
proto-newpow/mma-shadow/gen_block.py
Normal file
75
proto-newpow/mma-shadow/gen_block.py
Normal file
|
|
@ -0,0 +1,75 @@
|
|||
#!/usr/bin/env python3
|
||||
# gen_block.py: draws the mm8 block table for the mx8+mm8xR prototype and writes mm8_block.h.
|
||||
# Stream: SplitMix64 seeded with FNV-1a-64 of the bytes "igneum-mm8/igneum-genesis".
|
||||
# Per step k: a = below(8); b = below(7), b += (b >= a); c = below(8); c2 = below(7), c2 += (c2 >= c).
|
||||
# Step semantics: r[c] += C[l >> 2][2 * (l & 3)] (PTX d0), r[c2] += C[l >> 2][2 * (l & 3) + 1] (PTX d1), both mod 2^32.
|
||||
import sys
|
||||
|
||||
MASK64 = (1 << 64) - 1
|
||||
R_MAX = 512
|
||||
|
||||
def fnv1a64(data: bytes) -> int:
|
||||
h = 0xcbf29ce484222325
|
||||
for byte in data:
|
||||
h ^= byte
|
||||
h = (h * 0x100000001b3) & MASK64
|
||||
return h
|
||||
|
||||
class SplitMix64:
|
||||
def __init__(self, seed: int):
|
||||
self.s = seed & MASK64
|
||||
def next(self) -> int:
|
||||
self.s = (self.s + 0x9E3779B97F4A7C15) & MASK64
|
||||
z = self.s
|
||||
z = ((z ^ (z >> 30)) * 0xBF58476D1CE4E5B9) & MASK64
|
||||
z = ((z ^ (z >> 27)) * 0x94D049BB133111EB) & MASK64
|
||||
return z ^ (z >> 31)
|
||||
def below(self, n: int) -> int:
|
||||
return self.next() % n
|
||||
|
||||
def main():
|
||||
seed_text = b"igneum-mm8/igneum-genesis"
|
||||
seed = fnv1a64(seed_text)
|
||||
rng = SplitMix64(seed)
|
||||
rows = []
|
||||
for k in range(R_MAX):
|
||||
a = rng.below(8)
|
||||
b = rng.below(7)
|
||||
b += 1 if b >= a else 0
|
||||
c = rng.below(8)
|
||||
c2 = rng.below(7)
|
||||
c2 += 1 if c2 >= c else 0
|
||||
assert a != b and 0 <= b < 8 and c != c2 and 0 <= c2 < 8
|
||||
rows.append((a, b, c, c2))
|
||||
out = []
|
||||
out.append("// Generated by gen_block.py. mm8 block draws for class mx8+mm8xR, seed text \"%s\"," % seed_text.decode())
|
||||
out.append("// FNV-1a-64 seed 0x%016x, SplitMix64 stream. Step k uses (a, b, c, c2) = row k: r[c] += d0, r[c2] += d1. Do not edit by hand." % seed)
|
||||
out.append("#pragma once")
|
||||
out.append("#ifdef __cplusplus")
|
||||
out.append("#include <cstdint>")
|
||||
out.append("#else")
|
||||
out.append("#include <stdint.h>")
|
||||
out.append("#endif")
|
||||
out.append("#define IGNEUM_MM8_R_MAX %d" % R_MAX)
|
||||
out.append("#define IGNEUM_MM8_SEED 0x%016xull" % seed)
|
||||
out.append("// Packed uint16: a = v & 7, b = (v >> 3) & 7, c = (v >> 6) & 7, c2 = (v >> 9) & 7.")
|
||||
out.append("#define IGNEUM_MM8_TABLE_INIT { \\")
|
||||
for i in range(0, R_MAX, 16):
|
||||
chunk = rows[i:i+16]
|
||||
vals = ["0x%03xu" % (a | (b << 3) | (c << 6) | (c2 << 9)) for (a, b, c, c2) in chunk]
|
||||
out.append(" " + ", ".join(vals) + (", \\" if i + 16 < R_MAX else " }"))
|
||||
out.append("// X-macro list: X(k, a, b, c, c2) for every step k in 0..R_MAX-1. The kernel guards each with k < IGNEUM_MM8_R,")
|
||||
out.append("// so the register indices are compile-time constants (r[] stays in registers, no local memory).")
|
||||
out.append("#define IGNEUM_MM8_STEPS(X) \\")
|
||||
for k, (a, b, c, c2) in enumerate(rows):
|
||||
out.append(" X(%d, %d, %d, %d, %d)%s" % (k, a, b, c, c2, " \\" if k + 1 < R_MAX else ""))
|
||||
out.append("// Readable form, step: a b c c2")
|
||||
for k, (a, b, c, c2) in enumerate(rows):
|
||||
out.append("// %3d: %d %d %d %d" % (k, a, b, c, c2))
|
||||
path = sys.argv[1] if len(sys.argv) > 1 else "mm8_block.h"
|
||||
with open(path, "w") as f:
|
||||
f.write("\n".join(out) + "\n")
|
||||
print("seed 0x%016x, %d steps, first 4: %s -> %s" % (seed, R_MAX, rows[:4], path))
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
23
proto-newpow/mma-shadow/gen_ref_program.py
Normal file
23
proto-newpow/mma-shadow/gen_ref_program.py
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
#!/usr/bin/env python3
|
||||
# gen_ref_program.py: turns the 64 instruction lines of the pack's kernel.cu into a 32-lane C interpreter body
|
||||
# (ref_program.inc, included by verify_ref.c). Register-major: R[8][32]. A shuffle copies its source register
|
||||
# into T before the lane loop and reads T[l ^ mask]. Loads become mh_word(cache, x & mask) (lazy dataset).
|
||||
import re, sys
|
||||
src = open(sys.argv[1] if len(sys.argv) > 1 else "kernel.cu").read().split("\n")
|
||||
lines = [l for l in src if re.search(r"//\s*\d+ (mad|xor|load|shfl|mulhi|rotr|or|add|sub|mul|rotl)\s*$", l)]
|
||||
assert len(lines) == 64, len(lines)
|
||||
out = ["// Generated by gen_ref_program.py from the pack's kernel.cu. Do not edit by hand.",
|
||||
"// Expects: uint32_t R[8][32], T[32], SEL[32]; const uint32_t* cache; uint32_t mask; macro L = for (int l = 0; l < 32; ++l)."]
|
||||
for i, l in enumerate(lines):
|
||||
body, comment = l.strip().rsplit("//", 1)
|
||||
body = body.strip()
|
||||
m = re.search(r"__shfl_xor_sync\(0xffffffffu, r(\d), (\d+)\)", body)
|
||||
if m:
|
||||
out.append(" memcpy(T, R[%s], sizeof T);" % m.group(1))
|
||||
body = body.replace(m.group(0), "T[l ^ %s]" % m.group(2))
|
||||
body = re.sub(r"ds\[(r\d) & mask\]", r"mh_word(cache, \1 & mask)", body)
|
||||
body = re.sub(r"\br([0-7])\b", r"R[\1][l]", body)
|
||||
body = body.replace("__umulhi(", "umulhi(").replace("sel ", "SEL[l] ")
|
||||
out.append(" L { %s } //%s" % (body, comment))
|
||||
open("ref_program.inc", "w").write("\n".join(out) + "\n")
|
||||
print("wrote ref_program.inc with %d instructions" % len(lines))
|
||||
164
proto-newpow/mma-shadow/kernel.cu
Normal file
164
proto-newpow/mma-shadow/kernel.cu
Normal file
|
|
@ -0,0 +1,164 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
217
proto-newpow/mma-shadow/kernel_mm8.cu
Normal file
217
proto-newpow/mma-shadow/kernel_mm8.cu
Normal file
|
|
@ -0,0 +1,217 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// kernel_mm8.cu (proto-newpow/mma-shadow): the pack's kernel.cu with r0..r7 rewritten as uint32_t r[8] (same arithmetic)
|
||||
// and a block of IGNEUM_MM8_R mm8 steps at the end of every iteration (class "mx8+mm8xR"). IGNEUM_MM8_R 0 is the control
|
||||
// and is bit-exact with the pack. Define IGNEUM_MM8_REF for the shuffle-and-byte-product reference path instead of the PTX mma.
|
||||
// Step semantics (6 Oct 2026 correction): both tile outputs are consumed, r[c] += d0 and r[c2] += d1.
|
||||
// Cache fill, dataset build and the launch wrappers are the pack's text, unchanged.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
#include "mm8_block.h"
|
||||
#ifndef IGNEUM_MM8_R
|
||||
#define IGNEUM_MM8_R 0
|
||||
#endif
|
||||
#if IGNEUM_MM8_R < 0 || IGNEUM_MM8_R > IGNEUM_MM8_R_MAX
|
||||
#error "IGNEUM_MM8_R out of range"
|
||||
#endif
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
|
||||
// One mm8 step. A = the 32 lanes' r[a] (8 x 16 u8, lane l holds A[l >> 2][4 * (l & 3) .. +3], byte 0 = lowest k),
|
||||
// B = the 32 lanes' r[b] (16 x 8 u8, lane l holds B[4 * (l & 3) .. +3][l >> 2]). C = A * B exact in int32.
|
||||
// Lane l adds C[l >> 2][2 * (l & 3)] into r[c] and C[l >> 2][2 * (l & 3) + 1] into r[c2], both modulo 2^32 (c != c2),
|
||||
// so every step consumes both outputs of the tile. This is the PTX mma.m8n8k16 .u8 fragment layout (d0, d1), the same
|
||||
// arithmetic as proto-cuda/family-probe.cu's mm8 warp_ref.
|
||||
#ifdef IGNEUM_MM8_REF
|
||||
// Reference path: gather A's row (l >> 2) and B's two columns 2 * (l & 3) and 2 * (l & 3) + 1 with 12 shuffles,
|
||||
// then 32 byte products per lane in plain integer code.
|
||||
__device__ __forceinline__ void mm8_pair(uint32_t av, uint32_t bv, uint32_t lane, uint32_t& d0, uint32_t& d1) {
|
||||
uint32_t row = lane >> 2, col0 = 2u * (lane & 3u);
|
||||
uint32_t acc0 = 0u, acc1 = 0u;
|
||||
#pragma unroll
|
||||
for (uint32_t j = 0u; j < 4u; ++j) {
|
||||
uint32_t aw = __shfl_sync(0xffffffffu, av, (int)(4u * row + j));
|
||||
uint32_t bw0 = __shfl_sync(0xffffffffu, bv, (int)(4u * col0 + j));
|
||||
uint32_t bw1 = __shfl_sync(0xffffffffu, bv, (int)(4u * (col0 + 1u) + j));
|
||||
#pragma unroll
|
||||
for (uint32_t t = 0u; t < 4u; ++t) {
|
||||
uint32_t ab = (aw >> (8u * t)) & 0xffu;
|
||||
acc0 += ab * ((bw0 >> (8u * t)) & 0xffu);
|
||||
acc1 += ab * ((bw1 >> (8u * t)) & 0xffu);
|
||||
}
|
||||
}
|
||||
d0 = acc0; d1 = acc1;
|
||||
}
|
||||
#else
|
||||
// Native path: one tensor instruction with a zero accumulator, d0 and d1 straight out of the fragment.
|
||||
__device__ __forceinline__ void mm8_pair(uint32_t av, uint32_t bv, uint32_t lane, uint32_t& d0, uint32_t& d1) {
|
||||
(void)lane;
|
||||
asm("mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32 {%0,%1}, {%2}, {%3}, {%4,%5};"
|
||||
: "=r"(d0), "=r"(d1) : "r"(av), "r"(bv), "r"(0u), "r"(0u));
|
||||
}
|
||||
#endif
|
||||
#define IGNEUM_MM8_STEP(k, a, b, c, c2) if ((k) < IGNEUM_MM8_R) { uint32_t d0_, d1_; mm8_pair(r[a], r[b], lane, d0_, d1_); r[c] += d0_; r[c2] += d1_; }
|
||||
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r[8];
|
||||
uint32_t lane = threadIdx.x & 31u;
|
||||
(void)lane;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r[0] = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r[1] = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r[2] = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r[3] = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r[4] = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r[5] = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r[6] = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r[7] = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r[0];
|
||||
r[2] = r[3] * r[4] + r[2]; // 0 mad
|
||||
r[2] = r[1] * r[1] + r[2]; // 1 mad
|
||||
r[2] = r[3] * r[2] + r[2]; // 2 mad
|
||||
r[3] = r[3] ^ r[5]; // 3 xor
|
||||
r[7] = r[7] ^ ds[r[2] & mask]; // 4 load
|
||||
r[5] = r[5] ^ ds[r[7] & mask]; // 5 load
|
||||
r[1] = r[1] ^ __shfl_xor_sync(0xffffffffu, r[4], 8); // 6 shfl
|
||||
r[7] = r[7] ^ __shfl_xor_sync(0xffffffffu, r[3], 8); // 7 shfl
|
||||
r[1] = __umulhi(r[1], r[5]); // 8 mulhi
|
||||
r[6] = rotr_var(r[6], r[3]); // 9 rotr
|
||||
r[3] = r[3] | r[4]; // 10 or
|
||||
r[4] = r[4] ^ ds[r[3] & mask]; // 11 load
|
||||
r[0] = __umulhi(r[0], r[4]); // 12 mulhi
|
||||
r[5] = r[5] + r[1] + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r[0] = r[0] ^ ds[r[4] & mask]; // 14 load
|
||||
r[2] = r[2] - r[4]; // 15 sub
|
||||
r[2] = r[2] ^ ds[r[0] & mask]; // 16 load
|
||||
r[7] = r[7] ^ ds[r[2] & mask]; // 17 load
|
||||
r[7] = r[7] ^ __shfl_xor_sync(0xffffffffu, r[3], 4); // 18 shfl
|
||||
r[5] = r[5] * r[0]; // 19 mul
|
||||
r[3] = r[3] ^ __shfl_xor_sync(0xffffffffu, r[4], 2); // 20 shfl
|
||||
r[2] = r[2] ^ __shfl_xor_sync(0xffffffffu, r[4], 16); // 21 shfl
|
||||
r[6] = __umulhi(r[6], r[2]); // 22 mulhi
|
||||
r[6] = r[6] ^ ds[r[1] & mask]; // 23 load
|
||||
r[5] = r[5] * r[0]; // 24 mul
|
||||
r[5] = rotl_imm(r[5], 19u); // 25 rotl
|
||||
r[7] = r[7] ^ __shfl_xor_sync(0xffffffffu, r[6], 2); // 26 shfl
|
||||
r[0] = r[0] ^ r[5]; // 27 xor
|
||||
r[0] = r[0] ^ r[4]; // 28 xor
|
||||
r[3] = r[3] - r[0]; // 29 sub
|
||||
r[5] = r[5] * r[1]; // 30 mul
|
||||
r[7] = r[7] ^ ds[r[2] & mask]; // 31 load
|
||||
r[1] = r[1] ^ ds[r[0] & mask]; // 32 load
|
||||
r[5] = r[5] ^ r[6]; // 33 xor
|
||||
r[5] = r[5] ^ ds[r[1] & mask]; // 34 load
|
||||
r[0] = __umulhi(r[0], r[5]); // 35 mulhi
|
||||
r[5] = r[5] ^ __shfl_xor_sync(0xffffffffu, r[2], 4); // 36 shfl
|
||||
r[7] = r[7] ^ ds[r[0] & mask]; // 37 load
|
||||
r[3] = r[3] + r[1] + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r[1] = r[1] ^ __shfl_xor_sync(0xffffffffu, r[5], 4); // 39 shfl
|
||||
r[2] = r[2] ^ r[5]; // 40 xor
|
||||
r[3] = r[6] * r[3] + r[3]; // 41 mad
|
||||
r[6] = r[6] - r[7]; // 42 sub
|
||||
r[7] = r[7] ^ r[0]; // 43 xor
|
||||
r[1] = r[1] ^ ds[r[7] & mask]; // 44 load
|
||||
r[2] = r[2] * r[3]; // 45 mul
|
||||
r[1] = __umulhi(r[1], r[5]); // 46 mulhi
|
||||
r[4] = r[4] - r[3]; // 47 sub
|
||||
r[2] = rotr_var(r[2], r[6]); // 48 rotr
|
||||
r[3] = r[3] ^ ds[r[5] & mask]; // 49 load
|
||||
r[1] = r[1] + r[5] + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r[0] = r[0] * r[2]; // 51 mul
|
||||
r[0] = r[0] + r[2] + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r[1] = r[1] + r[0] + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r[7] = rotl_imm(r[7], 14u); // 54 rotl
|
||||
r[3] = r[3] + r[7] + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r[6] = r[6] ^ ds[r[7] & mask]; // 56 load
|
||||
r[1] = rotr_var(r[1], r[5]); // 57 rotr
|
||||
r[5] = r[5] ^ ds[r[4] & mask]; // 58 load
|
||||
r[6] = r[6] ^ ds[r[2] & mask]; // 59 load
|
||||
r[3] = r[5] * r[0] + r[3]; // 60 mad
|
||||
r[5] = r[5] + r[7] + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r[4] = r[4] + r[6] + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r[5] = rotl_imm(r[5], 19u); // 63 rotl
|
||||
#if IGNEUM_MM8_R > 0
|
||||
// mm8 block: IGNEUM_MM8_R steps after instruction 63, before the next iteration samples sel.
|
||||
IGNEUM_MM8_STEPS(IGNEUM_MM8_STEP)
|
||||
#endif
|
||||
}
|
||||
uint32_t lo = r[0] ^ rotl_imm(r[1], 7u) ^ rotl_imm(r[2], 14u) ^ rotl_imm(r[3], 21u);
|
||||
uint32_t hi = r[4] ^ rotl_imm(r[5], 9u) ^ rotl_imm(r[6], 18u) ^ rotl_imm(r[7], 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
109
proto-newpow/mma-shadow/memhard.h
Normal file
109
proto-newpow/mma-shadow/memhard.h
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
|
||||
// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0x3067619fu ^ prev[4];
|
||||
x[5] = 0x3c269176u ^ prev[5];
|
||||
x[6] = 0x84a03b03u ^ prev[6];
|
||||
x[7] = 0xf8c63294u ^ prev[7];
|
||||
x[8] = 0xff977c5bu ^ prev[8];
|
||||
x[9] = 0xe60def3eu ^ prev[9];
|
||||
x[10] = 0x63630141u ^ prev[10];
|
||||
x[11] = 0xb8fbcb58u ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
|
||||
s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
|
||||
s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
|
||||
s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
|
||||
s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
|
||||
s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
|
||||
s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
|
||||
s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
|
||||
s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
|
||||
s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
|
||||
s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
|
||||
s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
|
||||
s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
|
||||
s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
|
||||
s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
|
||||
s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0x3067619fu;
|
||||
s[1] = 0x3c269176u;
|
||||
s[2] = 0x84a03b03u;
|
||||
s[3] = 0xf8c63294u;
|
||||
s[4] = 0xff977c5bu;
|
||||
s[5] = 0xe60def3eu;
|
||||
s[6] = 0x63630141u;
|
||||
s[7] = 0xb8fbcb58u;
|
||||
s[8] = t * 0x42146205u + 0xbab68293u;
|
||||
s[9] = t * 0x52cbe0fbu + 0xcc162340u;
|
||||
s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
|
||||
s[11] = t * 0x6728907fu + 0xe62b8997u;
|
||||
s[12] = t * 0xd81d9751u + 0xc9c80297u;
|
||||
s[13] = t * 0x132952c3u + 0xf74a1654u;
|
||||
s[14] = t * 0xf60de277u + 0x3d704af5u;
|
||||
s[15] = t * 0x05358035u + 0x3cf522b7u;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
1072
proto-newpow/mma-shadow/mm8_block.h
Normal file
1072
proto-newpow/mma-shadow/mm8_block.h
Normal file
File diff suppressed because it is too large
Load diff
24
proto-newpow/mma-shadow/out/bench_0.log
Normal file
24
proto-newpow/mma-shadow/out/bench_0.log
Normal file
|
|
@ -0,0 +1,24 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 0 mm8 per hash = 0 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.93 ms first, 1.89 ms second (256 MiB)
|
||||
cache fill (host, one thread): 375.8 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.55 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
verify warp base 0 (pack vector, standalone): PASS
|
||||
verify warp base 4096 (pack vector, standalone): PASS
|
||||
verify warp base 1000000 (pack vector, standalone): PASS
|
||||
warm-up batch: 16777216 hashes in 266.01 ms wall
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
verify warp base 0 (pack vector, in batch): PASS
|
||||
verify warp base 4096 (pack vector, in batch): PASS
|
||||
verify warp base 1000000 (pack vector, in batch): PASS
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.74 ms -> 63.078 MH/s
|
||||
sustain start epoch 1791315749.655
|
||||
sustain end epoch 1791315774.658: 94 batches in 25.00 s -> 63.075 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_0.txt
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.075 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_0_ref.log
Normal file
19
proto-newpow/mma-shadow/out/bench_0_ref.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 0 mm8 per hash = 0 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.80 ms first, 1.77 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.58 ms first, 30.51 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
verify warp base 0 (pack vector, standalone): PASS
|
||||
verify warp base 4096 (pack vector, standalone): PASS
|
||||
verify warp base 1000000 (pack vector, standalone): PASS
|
||||
warm-up batch: 16777216 hashes in 266.00 ms wall
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
verify warp base 0 (pack vector, in batch): PASS
|
||||
verify warp base 4096 (pack vector, in batch): PASS
|
||||
verify warp base 1000000 (pack vector, in batch): PASS
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_128.log
Normal file
19
proto-newpow/mma-shadow/out/bench_128.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 128 mm8 per hash = 1024 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 32 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.91 ms first, 1.87 ms second (256 MiB)
|
||||
cache fill (host, one thread): 347.9 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.61 ms first, 30.56 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 128 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 265.96 ms wall
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.82 ms -> 63.077 MH/s (GPU time), wall 2659.83 ms -> 63.076 MH/s
|
||||
sustain start epoch 1791315890.874
|
||||
sustain end epoch 1791315915.893: 94 batches in 25.02 s -> 63.034 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_128.txt
|
||||
SUMMARY R=128 regs=32 blocksPerSM=24 mhs=63.077 mhs_wall=63.076 mhs_sustain=63.034 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
14
proto-newpow/mma-shadow/out/bench_128_ref.log
Normal file
14
proto-newpow/mma-shadow/out/bench_128_ref.log
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 128 mm8 per hash = 1024 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 175 registers/thread, 8 resident blocks/SM at 1 warp(s)/block = 8 resident warps/SM (16.7% of 48)
|
||||
cache fill (GPU): 1.78 ms first, 1.79 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.58 ms first, 30.52 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 128 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 447.05 ms wall
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=128 regs=175 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_32.log
Normal file
19
proto-newpow/mma-shadow/out/bench_32.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 32 mm8 per hash = 256 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.90 ms first, 1.87 ms second (256 MiB)
|
||||
cache fill (host, one thread): 349.6 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.56 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 32 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 266.00 ms wall
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.75 ms -> 63.078 MH/s (GPU time), wall 2659.77 ms -> 63.078 MH/s
|
||||
sustain start epoch 1791315833.292
|
||||
sustain end epoch 1791315858.296: 94 batches in 25.00 s -> 63.072 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_32.txt
|
||||
SUMMARY R=32 regs=29 blocksPerSM=24 mhs=63.078 mhs_wall=63.078 mhs_sustain=63.072 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
14
proto-newpow/mma-shadow/out/bench_32_ref.log
Normal file
14
proto-newpow/mma-shadow/out/bench_32_ref.log
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 32 mm8 per hash = 256 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 151 registers/thread, 12 resident blocks/SM at 1 warp(s)/block = 12 resident warps/SM (25.0% of 48)
|
||||
cache fill (GPU): 1.80 ms first, 1.77 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.57 ms first, 30.52 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 32 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 265.96 ms wall
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_512.log
Normal file
19
proto-newpow/mma-shadow/out/bench_512.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 512 mm8 per hash = 4096 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.93 ms first, 1.89 ms second (256 MiB)
|
||||
cache fill (host, one thread): 348.1 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.54 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 512 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 265.98 ms wall
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
|
||||
sustain start epoch 1791316287.337
|
||||
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_512.txt
|
||||
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
14
proto-newpow/mma-shadow/out/bench_512_ref.log
Normal file
14
proto-newpow/mma-shadow/out/bench_512_ref.log
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 512 mm8 per hash = 4096 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 213 registers/thread, 8 resident blocks/SM at 1 warp(s)/block = 8 resident warps/SM (16.7% of 48)
|
||||
cache fill (GPU): 1.82 ms first, 1.78 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.58 ms first, 30.52 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 512 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 994.87 ms wall
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=512 regs=213 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_8.log
Normal file
19
proto-newpow/mma-shadow/out/bench_8.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 8 mm8 per hash = 64 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 30 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.92 ms first, 1.86 ms second (256 MiB)
|
||||
cache fill (host, one thread): 349.8 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.61 ms first, 30.55 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 8 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 266.01 ms wall
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.75 ms -> 63.078 MH/s
|
||||
sustain start epoch 1791315790.867
|
||||
sustain end epoch 1791315815.870: 94 batches in 25.00 s -> 63.074 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_8.txt
|
||||
SUMMARY R=8 regs=30 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.074 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
14
proto-newpow/mma-shadow/out/bench_8_ref.log
Normal file
14
proto-newpow/mma-shadow/out/bench_8_ref.log
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 8 mm8 per hash = 64 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 77 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.81 ms first, 1.79 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.59 ms first, 30.52 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 8 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 265.99 ms wall
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=8 regs=77 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
1024
proto-newpow/mma-shadow/out/dump_0.txt
Normal file
1024
proto-newpow/mma-shadow/out/dump_0.txt
Normal file
File diff suppressed because it is too large
Load diff
1024
proto-newpow/mma-shadow/out/dump_128.txt
Normal file
1024
proto-newpow/mma-shadow/out/dump_128.txt
Normal file
File diff suppressed because it is too large
Load diff
1024
proto-newpow/mma-shadow/out/dump_32.txt
Normal file
1024
proto-newpow/mma-shadow/out/dump_32.txt
Normal file
File diff suppressed because it is too large
Load diff
1024
proto-newpow/mma-shadow/out/dump_512.txt
Normal file
1024
proto-newpow/mma-shadow/out/dump_512.txt
Normal file
File diff suppressed because it is too large
Load diff
1024
proto-newpow/mma-shadow/out/dump_8.txt
Normal file
1024
proto-newpow/mma-shadow/out/dump_8.txt
Normal file
File diff suppressed because it is too large
Load diff
30
proto-newpow/mma-shadow/out/power_0.csv
Normal file
30
proto-newpow/mma-shadow/out/power_0.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
15.15 W, 210 MHz, 45, 0 %, 2026/10/06 19:42:24.816
|
||||
48.43 W, 2520 MHz, 47, 0 %, 2026/10/06 19:42:25.847
|
||||
106.67 W, 2670 MHz, 51, 0 %, 2026/10/06 19:42:26.847
|
||||
156.77 W, 2670 MHz, 52, 0 %, 2026/10/06 19:42:27.848
|
||||
200.24 W, 2670 MHz, 53, 100 %, 2026/10/06 19:42:28.848
|
||||
200.29 W, 2670 MHz, 53, 100 %, 2026/10/06 19:42:29.848
|
||||
202.82 W, 2670 MHz, 54, 100 %, 2026/10/06 19:42:30.848
|
||||
201.27 W, 2670 MHz, 54, 100 %, 2026/10/06 19:42:31.849
|
||||
201.02 W, 2670 MHz, 54, 100 %, 2026/10/06 19:42:32.849
|
||||
201.02 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:33.849
|
||||
201.04 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:34.849
|
||||
200.96 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:35.850
|
||||
200.94 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:36.850
|
||||
200.96 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:37.850
|
||||
201.00 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:38.850
|
||||
201.02 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:39.851
|
||||
201.02 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:40.851
|
||||
201.12 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:41.851
|
||||
201.07 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:42.851
|
||||
201.02 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:43.852
|
||||
201.03 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:44.852
|
||||
201.03 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:45.852
|
||||
201.21 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:46.852
|
||||
201.31 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:47.853
|
||||
201.33 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:48.853
|
||||
201.33 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:49.853
|
||||
201.52 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:50.853
|
||||
201.41 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:51.854
|
||||
201.51 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:52.854
|
||||
201.49 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:53.854
|
||||
|
30
proto-newpow/mma-shadow/out/power_128.csv
Normal file
30
proto-newpow/mma-shadow/out/power_128.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
18.57 W, 210 MHz, 53, 0 %, 2026/10/06 19:44:46.034
|
||||
38.21 W, 2520 MHz, 54, 0 %, 2026/10/06 19:44:47.037
|
||||
108.02 W, 2670 MHz, 55, 0 %, 2026/10/06 19:44:48.037
|
||||
143.43 W, 2670 MHz, 59, 100 %, 2026/10/06 19:44:49.037
|
||||
207.98 W, 2670 MHz, 60, 100 %, 2026/10/06 19:44:50.038
|
||||
208.85 W, 2670 MHz, 60, 100 %, 2026/10/06 19:44:51.038
|
||||
209.59 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:52.038
|
||||
210.00 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:53.038
|
||||
210.08 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:54.039
|
||||
210.10 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:55.039
|
||||
209.96 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:56.039
|
||||
210.40 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:57.039
|
||||
210.88 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:58.040
|
||||
211.38 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:59.040
|
||||
211.56 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:00.040
|
||||
211.76 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:01.040
|
||||
212.68 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:02.040
|
||||
212.12 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:03.041
|
||||
211.51 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:04.041
|
||||
212.43 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:05.041
|
||||
212.79 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:06.041
|
||||
212.98 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:07.042
|
||||
212.89 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:08.042
|
||||
212.51 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:09.042
|
||||
212.93 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:10.042
|
||||
212.91 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:11.042
|
||||
213.05 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:12.043
|
||||
213.00 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:13.043
|
||||
213.23 W, 2670 MHz, 64, 100 %, 2026/10/06 19:45:14.043
|
||||
213.14 W, 2670 MHz, 64, 100 %, 2026/10/06 19:45:15.043
|
||||
|
30
proto-newpow/mma-shadow/out/power_32.csv
Normal file
30
proto-newpow/mma-shadow/out/power_32.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
18.08 W, 210 MHz, 53, 0 %, 2026/10/06 19:43:48.466
|
||||
40.68 W, 2520 MHz, 55, 0 %, 2026/10/06 19:43:49.498
|
||||
106.86 W, 2670 MHz, 56, 52 %, 2026/10/06 19:43:50.498
|
||||
141.51 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:51.498
|
||||
203.49 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:52.499
|
||||
203.95 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:53.499
|
||||
204.46 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:54.499
|
||||
204.84 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:55.499
|
||||
204.88 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:56.500
|
||||
205.53 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:57.500
|
||||
205.87 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:58.500
|
||||
205.71 W, 2670 MHz, 62, 100 %, 2026/10/06 19:43:59.500
|
||||
206.17 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:00.500
|
||||
206.54 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:01.501
|
||||
206.21 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:02.501
|
||||
206.09 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:03.504
|
||||
206.43 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:04.504
|
||||
206.97 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:05.504
|
||||
207.32 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:06.504
|
||||
207.58 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:07.505
|
||||
207.28 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:08.505
|
||||
207.92 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:09.505
|
||||
208.31 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:10.505
|
||||
208.17 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:11.505
|
||||
207.89 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:12.506
|
||||
208.03 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:13.506
|
||||
208.60 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:14.506
|
||||
208.41 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:15.506
|
||||
208.39 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:16.508
|
||||
208.19 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:17.508
|
||||
|
30
proto-newpow/mma-shadow/out/power_512.csv
Normal file
30
proto-newpow/mma-shadow/out/power_512.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
15.26 W, 210 MHz, 46, 0 %, 2026/10/06 19:51:22.484
|
||||
32.16 W, 2520 MHz, 48, 0 %, 2026/10/06 19:51:23.516
|
||||
84.22 W, 2670 MHz, 50, 100 %, 2026/10/06 19:51:24.517
|
||||
138.29 W, 2670 MHz, 54, 100 %, 2026/10/06 19:51:25.517
|
||||
214.64 W, 2670 MHz, 55, 100 %, 2026/10/06 19:51:26.517
|
||||
215.01 W, 2670 MHz, 55, 100 %, 2026/10/06 19:51:27.517
|
||||
215.23 W, 2670 MHz, 56, 100 %, 2026/10/06 19:51:28.517
|
||||
215.43 W, 2670 MHz, 56, 100 %, 2026/10/06 19:51:29.518
|
||||
215.27 W, 2670 MHz, 56, 100 %, 2026/10/06 19:51:30.518
|
||||
215.35 W, 2670 MHz, 57, 100 %, 2026/10/06 19:51:31.518
|
||||
215.48 W, 2670 MHz, 57, 100 %, 2026/10/06 19:51:32.518
|
||||
215.41 W, 2670 MHz, 57, 100 %, 2026/10/06 19:51:33.518
|
||||
215.39 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:34.518
|
||||
215.49 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:35.519
|
||||
215.67 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:36.519
|
||||
215.71 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:37.519
|
||||
215.68 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:38.519
|
||||
215.70 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:39.519
|
||||
215.75 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:40.519
|
||||
215.72 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:41.520
|
||||
215.80 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:42.520
|
||||
215.90 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:43.520
|
||||
215.95 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:44.520
|
||||
216.10 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:45.520
|
||||
216.21 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:46.521
|
||||
216.13 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:47.521
|
||||
216.06 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:48.521
|
||||
216.07 W, 2670 MHz, 61, 100 %, 2026/10/06 19:51:49.521
|
||||
216.16 W, 2670 MHz, 61, 100 %, 2026/10/06 19:51:50.521
|
||||
216.10 W, 2670 MHz, 61, 100 %, 2026/10/06 19:51:51.521
|
||||
|
30
proto-newpow/mma-shadow/out/power_8.csv
Normal file
30
proto-newpow/mma-shadow/out/power_8.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
16.99 W, 210 MHz, 51, 0 %, 2026/10/06 19:43:06.035
|
||||
40.60 W, 2520 MHz, 52, 0 %, 2026/10/06 19:43:07.067
|
||||
106.97 W, 2670 MHz, 53, 10 %, 2026/10/06 19:43:08.067
|
||||
137.95 W, 2670 MHz, 57, 100 %, 2026/10/06 19:43:09.068
|
||||
201.17 W, 2670 MHz, 58, 100 %, 2026/10/06 19:43:10.068
|
||||
201.47 W, 2670 MHz, 58, 100 %, 2026/10/06 19:43:11.068
|
||||
201.65 W, 2670 MHz, 58, 100 %, 2026/10/06 19:43:12.068
|
||||
201.97 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:13.068
|
||||
202.25 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:14.069
|
||||
202.06 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:15.069
|
||||
202.28 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:16.069
|
||||
202.47 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:17.069
|
||||
202.62 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:18.069
|
||||
202.62 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:19.070
|
||||
203.00 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:20.070
|
||||
203.19 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:21.070
|
||||
203.83 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:22.070
|
||||
203.66 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:23.070
|
||||
203.32 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:24.070
|
||||
203.21 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:25.071
|
||||
203.42 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:26.071
|
||||
203.90 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:27.071
|
||||
203.79 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:28.071
|
||||
203.96 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:29.071
|
||||
204.32 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:30.072
|
||||
204.51 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:31.072
|
||||
204.61 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:32.072
|
||||
204.72 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:33.072
|
||||
205.39 W, 2670 MHz, 62, 100 %, 2026/10/06 19:43:34.073
|
||||
205.90 W, 2670 MHz, 62, 100 %, 2026/10/06 19:43:35.073
|
||||
|
5
proto-newpow/mma-shadow/out/power_idle.csv
Normal file
5
proto-newpow/mma-shadow/out/power_idle.csv
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
15.34 W, 210 MHz, 45, 0 %
|
||||
15.21 W, 210 MHz, 45, 0 %
|
||||
15.12 W, 210 MHz, 45, 0 %
|
||||
15.05 W, 210 MHz, 45, 0 %
|
||||
15.02 W, 210 MHz, 45, 0 %
|
||||
|
23
proto-newpow/mma-shadow/out/ptxas_0.txt
Normal file
23
proto-newpow/mma-shadow/out/ptxas_0.txt
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
kernel_mm8.cu(98): warning #550-D: variable "lane" was set but never used
|
||||
uint32_t lane = threadIdx.x & 31u;
|
||||
^
|
||||
|
||||
Remark: The warnings can be suppressed with "-diag-suppress <warning-number>"
|
||||
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.262 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.536 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.384 ms
|
||||
23
proto-newpow/mma-shadow/out/ptxas_0_ref.txt
Normal file
23
proto-newpow/mma-shadow/out/ptxas_0_ref.txt
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
kernel_mm8.cu(98): warning #550-D: variable "lane" was set but never used
|
||||
uint32_t lane = threadIdx.x & 31u;
|
||||
^
|
||||
|
||||
Remark: The warnings can be suppressed with "-diag-suppress <warning-number>"
|
||||
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.413 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.749 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.256 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_128.txt
Normal file
17
proto-newpow/mma-shadow/out/ptxas_128.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 32 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 27.671 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 11.908 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.275 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_128_ref.txt
Normal file
17
proto-newpow/mma-shadow/out/ptxas_128_ref.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 175 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 9367.163 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.495 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.300 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_32.txt
Normal file
17
proto-newpow/mma-shadow/out/ptxas_32.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 15.442 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.497 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.272 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_32_ref.txt
Normal file
17
proto-newpow/mma-shadow/out/ptxas_32_ref.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 151 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 520.236 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.442 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.444 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_512.txt
Normal file
17
proto-newpow/mma-shadow/out/ptxas_512.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 36 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 83.096 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 11.882 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.282 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_512_ref.txt
Normal file
17
proto-newpow/mma-shadow/out/ptxas_512_ref.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 213 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 273654.000 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.613 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.467 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_8.txt
Normal file
17
proto-newpow/mma-shadow/out/ptxas_8.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 30 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 13.112 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.461 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.257 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_8_ref.txt
Normal file
17
proto-newpow/mma-shadow/out/ptxas_8_ref.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 77 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 78.215 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 11.877 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.236 ms
|
||||
7
proto-newpow/mma-shadow/out/results.md
Normal file
7
proto-newpow/mma-shadow/out/results.md
Normal file
|
|
@ -0,0 +1,7 @@
|
|||
| R | mm8/hash | MH/s (GPU) | ratio to R=0 | W mean | SM MHz | uJ/hash | max C | fingerprint | PTX == ref | CPU == GPU | verifier ms/unit (R=0, R, delta) | regs | blocks/SM |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 | yes | 1024 of 1024 | 10.179, 10.076, -0.103 | 29 | 24 |
|
||||
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.926, 9.974, 0.048 | 30 | 24 |
|
||||
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.075, 10.240, 0.165 | 29 | 24 |
|
||||
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.183, 11.326, 1.142 | 32 | 24 |
|
||||
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.887, 14.276, 4.389 | 36 | 24 |
|
||||
130
proto-newpow/mma-shadow/out/run.log
Normal file
130
proto-newpow/mma-shadow/out/run.log
Normal file
|
|
@ -0,0 +1,130 @@
|
|||
== Tue Oct 6 19:42:12 UTC 2026 host 5c32e87a3fb6 ==
|
||||
name, driver_version, power.draw [W], clocks.current.sm [MHz], temperature.gpu
|
||||
NVIDIA GeForce RTX 4090, 570.172.08, 15.34 W, 210 MHz, 45
|
||||
idle baseline (5 samples):
|
||||
15.34 W, 210 MHz, 45, 0 %
|
||||
15.21 W, 210 MHz, 45, 0 %
|
||||
15.12 W, 210 MHz, 45, 0 %
|
||||
15.05 W, 210 MHz, 45, 0 %
|
||||
15.02 W, 210 MHz, 45, 0 %
|
||||
== build verify_ref (gcc -O2)
|
||||
== R=0: build bench_0 (PTX path) and bench_0_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=0: timed run with power sampling
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
verify warp base 0 (pack vector, standalone): PASS
|
||||
verify warp base 4096 (pack vector, standalone): PASS
|
||||
verify warp base 1000000 (pack vector, standalone): PASS
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
verify warp base 0 (pack vector, in batch): PASS
|
||||
verify warp base 4096 (pack vector, in batch): PASS
|
||||
verify warp base 1000000 (pack vector, in batch): PASS
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.74 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315774.658: 94 batches in 25.00 s -> 63.075 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_0.txt
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.075 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
== R=0: reference path fingerprint
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
FPCHECK R=0 PTX == ref: yes (7c28cfb06c5c65a9)
|
||||
== R=0: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
self-test R=0 vector base 0: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 4096: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 1000000: PASS (0 lanes differ)
|
||||
1024 of 1024 lanes equal (PASS) [10.152 ms per unit during the check]
|
||||
verifier: 10.179 ms per unit at R=0, 10.076 ms per unit at R=0, mm8 block delta -0.103 ms per unit (0.00 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=0 equal=1024 of=1024 ms_r0=10.179 ms_r=10.076 delta=-0.103
|
||||
== R=8: build bench_8 (PTX path) and bench_8_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 30 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=8: timed run with power sampling
|
||||
igneum_hash_info: 30 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.75 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315815.870: 94 batches in 25.00 s -> 63.074 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_8.txt
|
||||
SUMMARY R=8 regs=30 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.074 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=8: reference path fingerprint
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=8 regs=77 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=8 PTX == ref: yes (06fc2593bfb94b4f)
|
||||
== R=8: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [10.106 ms per unit during the check]
|
||||
verifier: 9.926 ms per unit at R=0, 9.974 ms per unit at R=8, mm8 block delta 0.048 ms per unit (0.74 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=8 equal=1024 of=1024 ms_r0=9.926 ms_r=9.974 delta=0.048
|
||||
== R=32: build bench_32 (PTX path) and bench_32_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=32: timed run with power sampling
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.75 ms -> 63.078 MH/s (GPU time), wall 2659.77 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315858.296: 94 batches in 25.00 s -> 63.072 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_32.txt
|
||||
SUMMARY R=32 regs=29 blocksPerSM=24 mhs=63.078 mhs_wall=63.078 mhs_sustain=63.072 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=32: reference path fingerprint
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=32 PTX == ref: yes (26e83a65f519c865)
|
||||
== R=32: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [10.432 ms per unit during the check]
|
||||
verifier: 10.075 ms per unit at R=0, 10.240 ms per unit at R=32, mm8 block delta 0.165 ms per unit (0.64 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=32 equal=1024 of=1024 ms_r0=10.075 ms_r=10.240 delta=0.165
|
||||
== R=128: build bench_128 (PTX path) and bench_128_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 32 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=128: timed run with power sampling
|
||||
igneum_hash_info: 32 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.82 ms -> 63.077 MH/s (GPU time), wall 2659.83 ms -> 63.076 MH/s
|
||||
sustain end epoch 1791315915.893: 94 batches in 25.02 s -> 63.034 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_128.txt
|
||||
SUMMARY R=128 regs=32 blocksPerSM=24 mhs=63.077 mhs_wall=63.076 mhs_sustain=63.034 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=128: reference path fingerprint
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=128 regs=175 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=128 PTX == ref: yes (42223c2113188335)
|
||||
== R=128: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [11.349 ms per unit during the check]
|
||||
verifier: 10.183 ms per unit at R=0, 11.326 ms per unit at R=128, mm8 block delta 1.142 ms per unit (1.12 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=128 equal=1024 of=1024 ms_r0=10.183 ms_r=11.326 delta=1.142
|
||||
== R=512: build bench_512 (PTX path) and bench_512_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 36 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=512: timed run with power sampling
|
||||
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
|
||||
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_512.txt
|
||||
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=512: reference path fingerprint
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=512 regs=213 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=512 PTX == ref: yes (02b7002d747f3711)
|
||||
== R=512: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [14.440 ms per unit during the check]
|
||||
verifier: 9.887 ms per unit at R=0, 14.276 ms per unit at R=512, mm8 block delta 4.389 ms per unit (1.07 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=512 equal=1024 of=1024 ms_r0=9.887 ms_r=14.276 delta=4.389
|
||||
== Tue Oct 6 19:51:57 UTC 2026 done
|
||||
137
proto-newpow/mma-shadow/out/run_stdout.log
Normal file
137
proto-newpow/mma-shadow/out/run_stdout.log
Normal file
|
|
@ -0,0 +1,137 @@
|
|||
== Tue Oct 6 19:42:12 UTC 2026 host 5c32e87a3fb6 ==
|
||||
name, driver_version, power.draw [W], clocks.current.sm [MHz], temperature.gpu
|
||||
NVIDIA GeForce RTX 4090, 570.172.08, 15.34 W, 210 MHz, 45
|
||||
idle baseline (5 samples):
|
||||
15.34 W, 210 MHz, 45, 0 %
|
||||
15.21 W, 210 MHz, 45, 0 %
|
||||
15.12 W, 210 MHz, 45, 0 %
|
||||
15.05 W, 210 MHz, 45, 0 %
|
||||
15.02 W, 210 MHz, 45, 0 %
|
||||
== build verify_ref (gcc -O2)
|
||||
== R=0: build bench_0 (PTX path) and bench_0_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=0: timed run with power sampling
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
verify warp base 0 (pack vector, standalone): PASS
|
||||
verify warp base 4096 (pack vector, standalone): PASS
|
||||
verify warp base 1000000 (pack vector, standalone): PASS
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
verify warp base 0 (pack vector, in batch): PASS
|
||||
verify warp base 4096 (pack vector, in batch): PASS
|
||||
verify warp base 1000000 (pack vector, in batch): PASS
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.74 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315774.658: 94 batches in 25.00 s -> 63.075 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_0.txt
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.075 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
== R=0: reference path fingerprint
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
FPCHECK R=0 PTX == ref: yes (7c28cfb06c5c65a9)
|
||||
== R=0: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
self-test R=0 vector base 0: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 4096: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 1000000: PASS (0 lanes differ)
|
||||
1024 of 1024 lanes equal (PASS) [10.152 ms per unit during the check]
|
||||
verifier: 10.179 ms per unit at R=0, 10.076 ms per unit at R=0, mm8 block delta -0.103 ms per unit (0.00 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=0 equal=1024 of=1024 ms_r0=10.179 ms_r=10.076 delta=-0.103
|
||||
== R=8: build bench_8 (PTX path) and bench_8_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 30 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=8: timed run with power sampling
|
||||
igneum_hash_info: 30 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.75 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315815.870: 94 batches in 25.00 s -> 63.074 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_8.txt
|
||||
SUMMARY R=8 regs=30 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.074 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=8: reference path fingerprint
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=8 regs=77 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=8 PTX == ref: yes (06fc2593bfb94b4f)
|
||||
== R=8: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [10.106 ms per unit during the check]
|
||||
verifier: 9.926 ms per unit at R=0, 9.974 ms per unit at R=8, mm8 block delta 0.048 ms per unit (0.74 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=8 equal=1024 of=1024 ms_r0=9.926 ms_r=9.974 delta=0.048
|
||||
== R=32: build bench_32 (PTX path) and bench_32_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=32: timed run with power sampling
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.75 ms -> 63.078 MH/s (GPU time), wall 2659.77 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315858.296: 94 batches in 25.00 s -> 63.072 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_32.txt
|
||||
SUMMARY R=32 regs=29 blocksPerSM=24 mhs=63.078 mhs_wall=63.078 mhs_sustain=63.072 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=32: reference path fingerprint
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=32 PTX == ref: yes (26e83a65f519c865)
|
||||
== R=32: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [10.432 ms per unit during the check]
|
||||
verifier: 10.075 ms per unit at R=0, 10.240 ms per unit at R=32, mm8 block delta 0.165 ms per unit (0.64 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=32 equal=1024 of=1024 ms_r0=10.075 ms_r=10.240 delta=0.165
|
||||
== R=128: build bench_128 (PTX path) and bench_128_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 32 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=128: timed run with power sampling
|
||||
igneum_hash_info: 32 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.82 ms -> 63.077 MH/s (GPU time), wall 2659.83 ms -> 63.076 MH/s
|
||||
sustain end epoch 1791315915.893: 94 batches in 25.02 s -> 63.034 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_128.txt
|
||||
SUMMARY R=128 regs=32 blocksPerSM=24 mhs=63.077 mhs_wall=63.076 mhs_sustain=63.034 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=128: reference path fingerprint
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=128 regs=175 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=128 PTX == ref: yes (42223c2113188335)
|
||||
== R=128: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [11.349 ms per unit during the check]
|
||||
verifier: 10.183 ms per unit at R=0, 11.326 ms per unit at R=128, mm8 block delta 1.142 ms per unit (1.12 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=128 equal=1024 of=1024 ms_r0=10.183 ms_r=11.326 delta=1.142
|
||||
== R=512: build bench_512 (PTX path) and bench_512_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 36 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=512: timed run with power sampling
|
||||
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
|
||||
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_512.txt
|
||||
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=512: reference path fingerprint
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=512 regs=213 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=512 PTX == ref: yes (02b7002d747f3711)
|
||||
== R=512: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [14.440 ms per unit during the check]
|
||||
verifier: 9.887 ms per unit at R=0, 14.276 ms per unit at R=512, mm8 block delta 4.389 ms per unit (1.07 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=512 equal=1024 of=1024 ms_r0=9.887 ms_r=14.276 delta=4.389
|
||||
== Tue Oct 6 19:51:57 UTC 2026 done
|
||||
| R | mm8/hash | MH/s (GPU) | ratio to R=0 | W mean | SM MHz | uJ/hash | max C | fingerprint | PTX == ref | CPU == GPU | verifier ms/unit (R=0, R, delta) | regs | blocks/SM |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 | yes | 1024 of 1024 | 10.179, 10.076, -0.103 | 29 | 24 |
|
||||
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.926, 9.974, 0.048 | 30 | 24 |
|
||||
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.075, 10.240, 0.165 | 29 | 24 |
|
||||
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.183, 11.326, 1.142 | 32 | 24 |
|
||||
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.887, 14.276, 4.389 | 36 | 24 |
|
||||
8
proto-newpow/mma-shadow/out/verify_0.log
Normal file
8
proto-newpow/mma-shadow/out/verify_0.log
Normal file
|
|
@ -0,0 +1,8 @@
|
|||
cache: host fill 528.5 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
self-test R=0 vector base 0: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 4096: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 1000000: PASS (0 lanes differ)
|
||||
dump: 1024 lines, 32 units, R = 0 (0 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [10.152 ms per unit during the check]
|
||||
verifier: 10.179 ms per unit at R=0, 10.076 ms per unit at R=0, mm8 block delta -0.103 ms per unit (0.00 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=0 equal=1024 of=1024 ms_r0=10.179 ms_r=10.076 delta=-0.103
|
||||
5
proto-newpow/mma-shadow/out/verify_128.log
Normal file
5
proto-newpow/mma-shadow/out/verify_128.log
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
cache: host fill 533.2 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
dump: 1024 lines, 32 units, R = 128 (1024 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [11.349 ms per unit during the check]
|
||||
verifier: 10.183 ms per unit at R=0, 11.326 ms per unit at R=128, mm8 block delta 1.142 ms per unit (1.12 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=128 equal=1024 of=1024 ms_r0=10.183 ms_r=11.326 delta=1.142
|
||||
5
proto-newpow/mma-shadow/out/verify_32.log
Normal file
5
proto-newpow/mma-shadow/out/verify_32.log
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
cache: host fill 535.4 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
dump: 1024 lines, 32 units, R = 32 (256 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [10.432 ms per unit during the check]
|
||||
verifier: 10.075 ms per unit at R=0, 10.240 ms per unit at R=32, mm8 block delta 0.165 ms per unit (0.64 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=32 equal=1024 of=1024 ms_r0=10.075 ms_r=10.240 delta=0.165
|
||||
5
proto-newpow/mma-shadow/out/verify_512.log
Normal file
5
proto-newpow/mma-shadow/out/verify_512.log
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
cache: host fill 530.6 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
dump: 1024 lines, 32 units, R = 512 (4096 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [14.440 ms per unit during the check]
|
||||
verifier: 9.887 ms per unit at R=0, 14.276 ms per unit at R=512, mm8 block delta 4.389 ms per unit (1.07 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=512 equal=1024 of=1024 ms_r0=9.887 ms_r=14.276 delta=4.389
|
||||
5
proto-newpow/mma-shadow/out/verify_8.log
Normal file
5
proto-newpow/mma-shadow/out/verify_8.log
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
cache: host fill 558.2 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
dump: 1024 lines, 32 units, R = 8 (64 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [10.106 ms per unit during the check]
|
||||
verifier: 9.926 ms per unit at R=0, 9.974 ms per unit at R=8, mm8 block delta 0.048 ms per unit (0.74 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=8 equal=1024 of=1024 ms_r0=9.926 ms_r=9.974 delta=0.048
|
||||
66
proto-newpow/mma-shadow/program.h
Normal file
66
proto-newpow/mma-shadow/program.h
Normal file
|
|
@ -0,0 +1,66 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-genesis"
|
||||
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0xe323b9dcaf283a6full
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "mx8"
|
||||
#define IGNEUM_CLASS_MIXER_MULT 8
|
||||
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
|
||||
#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIXER_MULT 8 // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)
|
||||
#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
74
proto-newpow/mma-shadow/ref_program.inc
Normal file
74
proto-newpow/mma-shadow/ref_program.inc
Normal file
|
|
@ -0,0 +1,74 @@
|
|||
// Generated by gen_ref_program.py from the pack's kernel.cu. Do not edit by hand.
|
||||
// Expects: uint32_t R[8][32], T[32], SEL[32]; const uint32_t* cache; uint32_t mask; macro L = for (int l = 0; l < 32; ++l).
|
||||
L { R[2][l] = R[3][l] * R[4][l] + R[2][l]; } // 0 mad
|
||||
L { R[2][l] = R[1][l] * R[1][l] + R[2][l]; } // 1 mad
|
||||
L { R[2][l] = R[3][l] * R[2][l] + R[2][l]; } // 2 mad
|
||||
L { R[3][l] = R[3][l] ^ R[5][l]; } // 3 xor
|
||||
L { R[7][l] = R[7][l] ^ mh_word(cache, R[2][l] & mask); } // 4 load
|
||||
L { R[5][l] = R[5][l] ^ mh_word(cache, R[7][l] & mask); } // 5 load
|
||||
memcpy(T, R[4], sizeof T);
|
||||
L { R[1][l] = R[1][l] ^ T[l ^ 8]; } // 6 shfl
|
||||
memcpy(T, R[3], sizeof T);
|
||||
L { R[7][l] = R[7][l] ^ T[l ^ 8]; } // 7 shfl
|
||||
L { R[1][l] = umulhi(R[1][l], R[5][l]); } // 8 mulhi
|
||||
L { R[6][l] = rotr_var(R[6][l], R[3][l]); } // 9 rotr
|
||||
L { R[3][l] = R[3][l] | R[4][l]; } // 10 or
|
||||
L { R[4][l] = R[4][l] ^ mh_word(cache, R[3][l] & mask); } // 11 load
|
||||
L { R[0][l] = umulhi(R[0][l], R[4][l]); } // 12 mulhi
|
||||
L { R[5][l] = R[5][l] + R[1][l] + ((((SEL[l] >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); } // 13 add
|
||||
L { R[0][l] = R[0][l] ^ mh_word(cache, R[4][l] & mask); } // 14 load
|
||||
L { R[2][l] = R[2][l] - R[4][l]; } // 15 sub
|
||||
L { R[2][l] = R[2][l] ^ mh_word(cache, R[0][l] & mask); } // 16 load
|
||||
L { R[7][l] = R[7][l] ^ mh_word(cache, R[2][l] & mask); } // 17 load
|
||||
memcpy(T, R[3], sizeof T);
|
||||
L { R[7][l] = R[7][l] ^ T[l ^ 4]; } // 18 shfl
|
||||
L { R[5][l] = R[5][l] * R[0][l]; } // 19 mul
|
||||
memcpy(T, R[4], sizeof T);
|
||||
L { R[3][l] = R[3][l] ^ T[l ^ 2]; } // 20 shfl
|
||||
memcpy(T, R[4], sizeof T);
|
||||
L { R[2][l] = R[2][l] ^ T[l ^ 16]; } // 21 shfl
|
||||
L { R[6][l] = umulhi(R[6][l], R[2][l]); } // 22 mulhi
|
||||
L { R[6][l] = R[6][l] ^ mh_word(cache, R[1][l] & mask); } // 23 load
|
||||
L { R[5][l] = R[5][l] * R[0][l]; } // 24 mul
|
||||
L { R[5][l] = rotl_imm(R[5][l], 19u); } // 25 rotl
|
||||
memcpy(T, R[6], sizeof T);
|
||||
L { R[7][l] = R[7][l] ^ T[l ^ 2]; } // 26 shfl
|
||||
L { R[0][l] = R[0][l] ^ R[5][l]; } // 27 xor
|
||||
L { R[0][l] = R[0][l] ^ R[4][l]; } // 28 xor
|
||||
L { R[3][l] = R[3][l] - R[0][l]; } // 29 sub
|
||||
L { R[5][l] = R[5][l] * R[1][l]; } // 30 mul
|
||||
L { R[7][l] = R[7][l] ^ mh_word(cache, R[2][l] & mask); } // 31 load
|
||||
L { R[1][l] = R[1][l] ^ mh_word(cache, R[0][l] & mask); } // 32 load
|
||||
L { R[5][l] = R[5][l] ^ R[6][l]; } // 33 xor
|
||||
L { R[5][l] = R[5][l] ^ mh_word(cache, R[1][l] & mask); } // 34 load
|
||||
L { R[0][l] = umulhi(R[0][l], R[5][l]); } // 35 mulhi
|
||||
memcpy(T, R[2], sizeof T);
|
||||
L { R[5][l] = R[5][l] ^ T[l ^ 4]; } // 36 shfl
|
||||
L { R[7][l] = R[7][l] ^ mh_word(cache, R[0][l] & mask); } // 37 load
|
||||
L { R[3][l] = R[3][l] + R[1][l] + ((((SEL[l] >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); } // 38 add
|
||||
memcpy(T, R[5], sizeof T);
|
||||
L { R[1][l] = R[1][l] ^ T[l ^ 4]; } // 39 shfl
|
||||
L { R[2][l] = R[2][l] ^ R[5][l]; } // 40 xor
|
||||
L { R[3][l] = R[6][l] * R[3][l] + R[3][l]; } // 41 mad
|
||||
L { R[6][l] = R[6][l] - R[7][l]; } // 42 sub
|
||||
L { R[7][l] = R[7][l] ^ R[0][l]; } // 43 xor
|
||||
L { R[1][l] = R[1][l] ^ mh_word(cache, R[7][l] & mask); } // 44 load
|
||||
L { R[2][l] = R[2][l] * R[3][l]; } // 45 mul
|
||||
L { R[1][l] = umulhi(R[1][l], R[5][l]); } // 46 mulhi
|
||||
L { R[4][l] = R[4][l] - R[3][l]; } // 47 sub
|
||||
L { R[2][l] = rotr_var(R[2][l], R[6][l]); } // 48 rotr
|
||||
L { R[3][l] = R[3][l] ^ mh_word(cache, R[5][l] & mask); } // 49 load
|
||||
L { R[1][l] = R[1][l] + R[5][l] + ((((SEL[l] >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); } // 50 add
|
||||
L { R[0][l] = R[0][l] * R[2][l]; } // 51 mul
|
||||
L { R[0][l] = R[0][l] + R[2][l] + ((((SEL[l] >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); } // 52 add
|
||||
L { R[1][l] = R[1][l] + R[0][l] + ((((SEL[l] >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); } // 53 add
|
||||
L { R[7][l] = rotl_imm(R[7][l], 14u); } // 54 rotl
|
||||
L { R[3][l] = R[3][l] + R[7][l] + ((((SEL[l] >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); } // 55 add
|
||||
L { R[6][l] = R[6][l] ^ mh_word(cache, R[7][l] & mask); } // 56 load
|
||||
L { R[1][l] = rotr_var(R[1][l], R[5][l]); } // 57 rotr
|
||||
L { R[5][l] = R[5][l] ^ mh_word(cache, R[4][l] & mask); } // 58 load
|
||||
L { R[6][l] = R[6][l] ^ mh_word(cache, R[2][l] & mask); } // 59 load
|
||||
L { R[3][l] = R[5][l] * R[0][l] + R[3][l]; } // 60 mad
|
||||
L { R[5][l] = R[5][l] + R[7][l] + ((((SEL[l] >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); } // 61 add
|
||||
L { R[4][l] = R[4][l] + R[6][l] + ((((SEL[l] >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); } // 62 add
|
||||
L { R[5][l] = rotl_imm(R[5][l], 19u); } // 63 rotl
|
||||
43
proto-newpow/mma-shadow/run.sh
Executable file
43
proto-newpow/mma-shadow/run.sh
Executable file
|
|
@ -0,0 +1,43 @@
|
|||
#!/bin/bash
|
||||
# run.sh (proto-newpow/mma-shadow): build and measure the ladder R in {0, 8, 32, 128, 512} on the GPU box.
|
||||
# Runs under /root/horizon-newpow/mma-shadow. Every artefact lands in ./out/.
|
||||
set -u
|
||||
export PATH=/usr/local/cuda/bin:$PATH
|
||||
cd "$(dirname "$0")"
|
||||
LADDER="${LADDER:-0 8 32 128 512}"
|
||||
SUSTAIN="${SUSTAIN:-25}" # seconds of sustained hashing for the power meter (mean taken after the first 10 s)
|
||||
mkdir -p out
|
||||
echo "== $(date -u) host $(hostname) ==" | tee out/run.log
|
||||
nvidia-smi --query-gpu=name,driver_version,power.draw,clocks.sm,temperature.gpu --format=csv | tee -a out/run.log
|
||||
echo "idle baseline (5 samples):" | tee -a out/run.log
|
||||
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu --format=csv,noheader -l 1 > out/power_idle.csv & P=$!; sleep 5; kill $P 2>/dev/null; wait $P 2>/dev/null; cat out/power_idle.csv | tee -a out/run.log
|
||||
|
||||
echo "== build verify_ref (gcc -O2)" | tee -a out/run.log
|
||||
gcc -O2 -o verify_ref verify_ref.c 2>&1 | tee -a out/run.log
|
||||
|
||||
for R in $LADDER; do
|
||||
echo "== R=$R: build bench_$R (PTX path) and bench_${R}_ref" | tee -a out/run.log
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -Xptxas -v -o bench_$R bench.cu kernel_mm8.cu 2> out/ptxas_$R.txt || { echo "BUILD FAILED bench_$R"; cat out/ptxas_$R.txt; continue; }
|
||||
grep -A2 "igneum_hash" out/ptxas_$R.txt | grep -i "registers\|spill" | tee -a out/run.log
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -DIGNEUM_MM8_REF -Xptxas -v -o bench_${R}_ref bench.cu kernel_mm8.cu 2> out/ptxas_${R}_ref.txt || { echo "BUILD FAILED bench_${R}_ref"; cat out/ptxas_${R}_ref.txt; }
|
||||
|
||||
echo "== R=$R: timed run with power sampling" | tee -a out/run.log
|
||||
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu,timestamp --format=csv,noheader -l 1 > out/power_$R.csv &
|
||||
SMI=$!
|
||||
./bench_$R --batches 10 --sustain $SUSTAIN --dump out/dump_$R.txt 32 > out/bench_$R.log 2>&1
|
||||
kill $SMI 2>/dev/null; wait $SMI 2>/dev/null
|
||||
grep -E "igneum_hash_info|cache check|dataset self-test|verify warp|fingerprint|timed:|sustain end|SUMMARY|OVERALL|dump:" out/bench_$R.log | tee -a out/run.log
|
||||
|
||||
echo "== R=$R: reference path fingerprint" | tee -a out/run.log
|
||||
./bench_${R}_ref --fingerprint-only --no-host-cache > out/bench_${R}_ref.log 2>&1
|
||||
grep -E "fingerprint|OVERALL" out/bench_${R}_ref.log | tee -a out/run.log
|
||||
FP=$(grep -oE "fingerprint: [0-9a-f]{16}" out/bench_$R.log | cut -d' ' -f2)
|
||||
FPR=$(grep -oE "fingerprint: [0-9a-f]{16}" out/bench_${R}_ref.log | cut -d' ' -f2)
|
||||
if [ -n "$FP" ] && [ "$FP" = "$FPR" ]; then echo "FPCHECK R=$R PTX == ref: yes ($FP)" | tee -a out/run.log; else echo "FPCHECK R=$R PTX == ref: NO (ptx $FP ref $FPR)" | tee -a out/run.log; fi
|
||||
|
||||
echo "== R=$R: CPU reference on the 32-warp dump (taskset -c 2)" | tee -a out/run.log
|
||||
taskset -c 2 ./verify_ref out/dump_$R.txt $R --time $( [ "$R" = "0" ] && echo --self-test ) > out/verify_$R.log 2>&1
|
||||
grep -E "self-test|lanes equal|first mismatch|verifier:|VERIFY" out/verify_$R.log | tee -a out/run.log
|
||||
done
|
||||
echo "== $(date -u) done" | tee -a out/run.log
|
||||
python3 summarise.py | tee out/results.md
|
||||
45
proto-newpow/mma-shadow/summarise.py
Normal file
45
proto-newpow/mma-shadow/summarise.py
Normal file
|
|
@ -0,0 +1,45 @@
|
|||
#!/usr/bin/env python3
|
||||
# summarise.py: builds the RESULTS table from out/*.log and out/power_*.csv (run by run.sh on the box).
|
||||
import re, glob, os
|
||||
from datetime import datetime
|
||||
rows = []
|
||||
def grab(pat, text, default=None):
|
||||
m = re.search(pat, text); return m.group(1) if m else default
|
||||
def power_mean(R, log):
|
||||
try:
|
||||
s0 = float(grab(r"sustain start epoch ([0-9.]+)", log)); s1 = float(grab(r"sustain end epoch ([0-9.]+)", log))
|
||||
except Exception:
|
||||
return None, None, None
|
||||
pw, clk, tmp = [], [], []
|
||||
for line in open("out/power_%s.csv" % R):
|
||||
parts = [p.strip() for p in line.split(",")]
|
||||
if len(parts) < 5: continue
|
||||
try:
|
||||
ts = datetime.strptime(parts[4], "%Y/%m/%d %H:%M:%S.%f").timestamp()
|
||||
except Exception:
|
||||
continue
|
||||
if ts >= s0 + 10 and ts <= s1:
|
||||
pw.append(float(parts[0].split()[0])); clk.append(float(parts[1].split()[0])); tmp.append(float(parts[2].split()[0]))
|
||||
if not pw: return None, None, None
|
||||
return sum(pw) / len(pw), sum(clk) / len(clk), max(tmp)
|
||||
base_mhs = None
|
||||
out = ["| R | mm8/hash | MH/s (GPU) | ratio to R=0 | W mean | SM MHz | uJ/hash | max C | fingerprint | PTX == ref | CPU == GPU | verifier ms/unit (R=0, R, delta) | regs | blocks/SM |", "|---|---|---|---|---|---|---|---|---|---|---|---|---|---|"]
|
||||
for Rs in sorted([int(os.path.basename(p)[6:-4]) for p in glob.glob("out/bench_*.log") if "_ref" not in p]):
|
||||
log = open("out/bench_%d.log" % Rs).read()
|
||||
summ = grab(r"(SUMMARY.*)", log, "")
|
||||
mhs = float(grab(r"mhs=([0-9.]+)", summ, "0")); regs = grab(r"regs=(\d+)", summ, "?"); bps = grab(r"blocksPerSM=(\d+)", summ, "?")
|
||||
fp = grab(r"fingerprint=([0-9a-f]{16})", summ, "?")
|
||||
if Rs == 0: base_mhs = mhs
|
||||
ratio = ("%.3f" % (mhs / base_mhs)) if base_mhs else "?"
|
||||
runlog = open("out/run.log").read()
|
||||
fpc = grab(r"FPCHECK R=%d PTX == ref: (\S+)" % Rs, runlog, "?")
|
||||
ver = open("out/verify_%d.log" % Rs).read() if os.path.exists("out/verify_%d.log" % Rs) else ""
|
||||
vs = grab(r"(VERIFY.*)", ver, "")
|
||||
eq = grab(r"equal=(\d+)", vs, "?"); of = grab(r"of=(\d+)", vs, "?")
|
||||
ms0 = grab(r"ms_r0=([0-9.]+)", vs, "?"); msr = grab(r"ms_r=([0-9.]+)", vs, "?"); dl = grab(r"delta=([0-9.-]+)", vs, "?")
|
||||
w, clk, tmx = power_mean(Rs, log)
|
||||
uj = ("%.2f" % (w / (mhs * 1e6) * 1e6)) if (w and mhs) else "?"
|
||||
out.append("| %d | %d | %.2f | %s | %s | %s | %s | %s | %s | %s | %s of %s | %s, %s, %s | %s | %s |" % (
|
||||
Rs, 8 * Rs, mhs, ratio, ("%.1f" % w) if w else "?", ("%.0f" % clk) if clk else "?", uj, ("%.0f" % tmx) if tmx else "?",
|
||||
fp, fpc, eq, of, ms0, msr, dl, regs, bps))
|
||||
print("\n".join(out))
|
||||
57
proto-newpow/mma-shadow/vectors.h
Normal file
57
proto-newpow/mma-shadow/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x19b56348bc85304dull, 0xb08a9cfb44aa720full, 0xe1f8f627f780eff7ull, 0x2ff3e86ff1696161ull, 0xb65e0578c257e8acull, 0xf37b5705c2bebaaeull, 0xb19fef670982389eull, 0x2374331b28827a11ull,
|
||||
0xf492cdd05dda7f88ull, 0x37700f1385e19885ull, 0xa27524f3010b3e87ull, 0x6c8a24de8b938c43ull, 0x3dc017c8820cfd16ull, 0x5c80d37146657b68ull, 0x214082cac03f734aull, 0x1d67c665145d72f3ull,
|
||||
0x9fb0ae736eb87ff5ull, 0xcd7d416ae3a6be61ull, 0x0069cda85c6de58full, 0x0b7d82e717224baaull, 0x600960f6d7be30f0ull, 0xeb29032663b0c4d2ull, 0xf29583a5766408f1ull, 0x6c9532a99ad7314dull,
|
||||
0x9a949fd0959cc80full, 0xa20258520e7f6c25ull, 0xf2ab11f9bb032e38ull, 0xcc967bcd0c8d07c1ull, 0x37745267bb3231f2ull, 0x35a046048c2b69b3ull, 0xaa51834cd3f364f3ull, 0x359192708e4f754aull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x62fb132a9943127aull, 0x0b703e577e7f4ecaull, 0xf9f24f5522ce7593ull, 0x3cf5c516abc4332aull, 0xd25523f5f6d7a127ull, 0xd2081a002f983682ull, 0xbf46c54e9b3c4254ull, 0xca362e291e5e5f4dull,
|
||||
0x6039712f10f457a3ull, 0x8a34b7cabf97c23bull, 0xa473c6a2e0bf59bcull, 0x6cf3926513a4b069ull, 0x297ec2998376a40dull, 0x8efd7f601a8f28dbull, 0x8e72532dfdc1e544ull, 0x917c2b2ebe2a7e00ull,
|
||||
0x923fbb2d2f635c25ull, 0xce864ea5c0dedad9ull, 0x4b8ec7e874e446efull, 0x1b69b69465449196ull, 0x5ef3a8a6edb369cfull, 0x06c263ef9ce63fc4ull, 0x9c2048fd9d9e2639ull, 0x457fdd96ca4a138eull,
|
||||
0xd1904018b8d7b6e3ull, 0x8682312fb2e96ab8ull, 0xdc3257e0d0f979a5ull, 0xa51b0a8519d87db5ull, 0x334f08ec056e618bull, 0x3464ce71dc65119dull, 0x6a4d6df066332e04ull, 0x7d7866cb9cfca8ffull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x86b6cb0e13d89b03ull, 0x96299a3f19d7ef15ull, 0x67d2c55100d2f876ull, 0x0a4dfe97d671b728ull, 0x41e4489014d42595ull, 0xf11cb1958c0c0e82ull, 0xf8b70b0c0a03175full, 0x632299df87d5063eull,
|
||||
0xe198417776130492ull, 0x8ffc5449290d7be2ull, 0x5f2e264eb1311f1bull, 0x988376463ac88586ull, 0x83969eadda489c26ull, 0xbed0a2c3f255d306ull, 0x1a949d271961a819ull, 0x5bce06eb6984725cull,
|
||||
0x94d5d6a1b0ffd4e9ull, 0xf3c78bae6c2182b4ull, 0xb97e9fe1bbfcdd55ull, 0x70262d1d4c0eccb2ull, 0x1fc93b427dba28d9ull, 0x02b2e3c4317f2a2dull, 0x54d3d42a588edcb9ull, 0x79998677846e7cceull,
|
||||
0x486522a5425f821aull, 0x95fa88e933360e52ull, 0xc8bae2da2b883f6cull, 0xbe3eb610ad33614full, 0x20efb3c4de82907full, 0xd6b650cfedfb26b7ull, 0x8c24447a646dba26ull, 0x9c004678515e44ecull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0xfdad4319u, 0x1a7b68e1u, 0xde6db608u, 0x13d73892u, 0xd17f447au, 0xb2221ccfu, 0x9db004bdu, 0x57d7d367u,
|
||||
0xdbc4cf34u, 0x697c009au, 0xc43af1d4u, 0x97f12b2eu, 0x74c37cd0u, 0xc651ea15u, 0x665a6d29u, 0x22330a2du
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xa83e7aa6u;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0x3230bc7bu, 0x7fbfe2c9u, 0xb2690991u, 0x1745c7c5u, 0x0ab0ccafu, 0x1bf87d6bu, 0x160139fdu, 0x719817acu, 0x0155df4bu, 0xbe1e86c3u, 0x680bcd6cu, 0x79c3dc6cu, 0x181e7e5fu, 0x0713a109u, 0xc705dd9fu, 0x3933b7a8u, 0xdd1c0431u, 0x50522b30u, 0xa0020b38u, 0xbff39e96u, 0x21b67e18u, 0x740f8db3u, 0x2baba568u, 0x2c9bef83u, 0x0ad9b671u, 0xc4327869u, 0x7b4fd7d0u, 0x2c29965fu, 0xec56f15fu, 0x61111746u, 0x303a1d6eu, 0xbddcfd1au, 0xf829a355u, 0x6d5df2a9u, 0x01ab8e44u, 0x06d13507u, 0xda8dcfc6u, 0x01a703e1u, 0xafe7d2c1u, 0xc091c3a2u, 0xac1814feu, 0x6e6ff62au, 0x8fdf01bau, 0xdd3f7159u, 0xdfa0d75cu, 0x26684c35u, 0x7f441e63u, 0x88df2570u, 0x8aa4d5ebu, 0xcc816c05u, 0x434df890u, 0xcd392ad6u, 0x1ab4cb63u, 0x595926fau, 0x7cd76b41u, 0x20cb95c4u, 0x13cf823fu, 0xf9daf901u, 0xff9af40au, 0x2c7dfa51u, 0x871206dbu, 0x938c116cu, 0xb64bf199u, 0x5751f874u
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u,
|
||||
0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du,
|
||||
0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull;
|
||||
152
proto-newpow/mma-shadow/verify_ref.c
Normal file
152
proto-newpow/mma-shadow/verify_ref.c
Normal file
|
|
@ -0,0 +1,152 @@
|
|||
// verify_ref.c (proto-newpow/mma-shadow): plain C CPU reference for the mx8+mm8xR prototype.
|
||||
// Reads a bench --dump file (lines: base lane value_hex, 32 lanes per warp), recomputes every warp with a
|
||||
// register-major 32-lane interpreter of the pack's program (ref_program.inc, generated from kernel.cu), lazy
|
||||
// dataset words through memhard.h's mh_word over a host-filled 256 MiB cache, the mm8 block in the spec layout,
|
||||
// then the fold. Prints "N of 1024 lanes equal" and the first mismatch, then times the verifier per 32-lane unit.
|
||||
// Build: gcc -O2 -o verify_ref verify_ref.c Run: taskset -c 2 ./verify_ref dump.txt <R> [--time]
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <time.h>
|
||||
#define IGNEUM_NO_CUDA
|
||||
#include "program.h"
|
||||
#include "vectors.h"
|
||||
#include "memhard.h"
|
||||
#include "mm8_block.h"
|
||||
|
||||
static const uint16_t MM8_TABLE[IGNEUM_MM8_R_MAX] = IGNEUM_MM8_TABLE_INIT;
|
||||
|
||||
static uint32_t splitmix32(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }
|
||||
static uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
static uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
static uint32_t umulhi(uint32_t a, uint32_t b) { return (uint32_t)(((uint64_t)a * (uint64_t)b) >> 32); }
|
||||
static double nowMs(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1e3 + ts.tv_nsec / 1e6; }
|
||||
|
||||
#define L for (int l = 0; l < 32; ++l)
|
||||
|
||||
// One mm8 step on the 32-lane register file (spec layout, identical to family-probe.cu's mm8 warp_ref):
|
||||
// r[c] += C[l >> 2][2 * (l & 3)] (d0), r[c2] += C[l >> 2][2 * (l & 3) + 1] (d1).
|
||||
static void mm8_step(uint32_t R[8][32], int a, int b, int c, int c2) {
|
||||
uint32_t A[32], B[32];
|
||||
memcpy(A, R[a], sizeof A);
|
||||
memcpy(B, R[b], sizeof B);
|
||||
L {
|
||||
int row = l >> 2, col0 = 2 * (l & 3);
|
||||
uint32_t acc0 = 0u, acc1 = 0u;
|
||||
for (int k = 0; k < 16; ++k) {
|
||||
uint32_t av = (A[row * 4 + k / 4] >> (8 * (k % 4))) & 0xffu;
|
||||
acc0 += av * ((B[col0 * 4 + k / 4] >> (8 * (k % 4))) & 0xffu);
|
||||
acc1 += av * ((B[(col0 + 1) * 4 + k / 4] >> (8 * (k % 4))) & 0xffu);
|
||||
}
|
||||
R[c][l] += acc0;
|
||||
R[c2][l] += acc1;
|
||||
}
|
||||
}
|
||||
|
||||
static void hash_unit(const uint32_t* cache, uint32_t base, int Rsteps, uint64_t out[32]) {
|
||||
static const uint32_t SEEDW[8] = IGNEUM_SEEDW_INIT;
|
||||
uint32_t R[8][32], T[32], SEL[32];
|
||||
const uint32_t mask = IGNEUM_MASK;
|
||||
L {
|
||||
uint32_t nonce = base + (uint32_t)l;
|
||||
for (int i = 0; i < 8; ++i) {
|
||||
uint32_t x = nonce ^ SEEDW[i]; x += 0x9e3779b9u * (uint32_t)(i + 1); x = splitmix32(x);
|
||||
R[i][l] = x ^ SEEDW[(i + 1) & 7];
|
||||
}
|
||||
}
|
||||
for (int it = 0; it < 8; ++it) {
|
||||
L { SEL[l] = R[0][l]; }
|
||||
#include "ref_program.inc"
|
||||
for (int k = 0; k < Rsteps; ++k) {
|
||||
uint16_t v = MM8_TABLE[k];
|
||||
mm8_step(R, v & 7, (v >> 3) & 7, (v >> 6) & 7, (v >> 9) & 7);
|
||||
}
|
||||
}
|
||||
L {
|
||||
uint32_t lo = R[0][l] ^ rotl_imm(R[1][l], 7u) ^ rotl_imm(R[2][l], 14u) ^ rotl_imm(R[3][l], 21u);
|
||||
uint32_t hi = R[4][l] ^ rotl_imm(R[5][l], 9u) ^ rotl_imm(R[6][l], 18u) ^ rotl_imm(R[7][l], 27u);
|
||||
out[l] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
}
|
||||
|
||||
static uint64_t fnv1a64(const void* p, size_t n) {
|
||||
const uint8_t* b = (const uint8_t*)p; uint64_t h = 0xcbf29ce484222325ull;
|
||||
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
|
||||
return h;
|
||||
}
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
if (argc < 3) { printf("usage: verify_ref <dump file> <R> [--time] [--self-test]\n"); return 2; }
|
||||
const char* path = argv[1];
|
||||
int Rsteps = atoi(argv[2]);
|
||||
int doTime = 0, selfTest = 0;
|
||||
for (int i = 3; i < argc; ++i) { if (!strcmp(argv[i], "--time")) doTime = 1; if (!strcmp(argv[i], "--self-test")) selfTest = 1; }
|
||||
if (Rsteps < 0 || Rsteps > IGNEUM_MM8_R_MAX) { printf("R out of range\n"); return 2; }
|
||||
|
||||
size_t words = (size_t)1u << IGNEUM_CACHE_LOG2_WORDS;
|
||||
uint32_t* cache = (uint32_t*)malloc(words * 4u);
|
||||
if (!cache) { printf("no memory for the cache\n"); return 2; }
|
||||
double c0 = nowMs();
|
||||
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(cache, seg);
|
||||
double c1 = nowMs();
|
||||
uint64_t fnv = fnv1a64(cache, words * 4u);
|
||||
printf("cache: host fill %.1f ms, FNV-1a 64 %016llx vs Mac %016llx %s\n", c1 - c0, (unsigned long long)fnv,
|
||||
(unsigned long long)IGNEUM_CACHE_FNV64, fnv == IGNEUM_CACHE_FNV64 ? "PASS" : "FAIL");
|
||||
if (fnv != IGNEUM_CACHE_FNV64) return 1;
|
||||
|
||||
if (selfTest) { // R = 0 interpreter against the pack's vectors
|
||||
int ok = 1;
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
uint64_t out[32]; hash_unit(cache, IGNEUM_VEC_BASE[w], 0, out);
|
||||
int bad = 0; L { if (out[l] != IGNEUM_VEC_OUT[w][l]) ++bad; }
|
||||
printf("self-test R=0 vector base %u: %s (%d lanes differ)\n", IGNEUM_VEC_BASE[w], bad ? "FAIL" : "PASS", bad);
|
||||
ok = ok && !bad;
|
||||
}
|
||||
if (!ok) return 1;
|
||||
}
|
||||
|
||||
// Read the dump: 32 consecutive lines per warp, same base
|
||||
FILE* f = fopen(path, "r");
|
||||
if (!f) { printf("cannot open %s\n", path); return 2; }
|
||||
uint32_t* bases = NULL; uint64_t* vals = NULL; size_t n = 0, cap = 0;
|
||||
unsigned base; int lane; unsigned long long v;
|
||||
while (fscanf(f, "%u %d %llx", &base, &lane, &v) == 3) {
|
||||
if (n == cap) { cap = cap ? cap * 2 : 1024; bases = realloc(bases, cap * 4); vals = realloc(vals, cap * 8); }
|
||||
bases[n] = base; vals[n] = v; ++n;
|
||||
}
|
||||
fclose(f);
|
||||
if (n % 32 != 0) { printf("dump has %zu lines, not a multiple of 32\n", n); return 2; }
|
||||
size_t units = n / 32;
|
||||
printf("dump: %zu lines, %zu units, R = %d (%d mm8 per hash)\n", n, units, Rsteps, 8 * Rsteps);
|
||||
|
||||
size_t equal = 0; int firstPrinted = 0;
|
||||
double v0 = nowMs();
|
||||
for (size_t u = 0; u < units; ++u) {
|
||||
uint64_t out[32]; hash_unit(cache, bases[u * 32], Rsteps, out);
|
||||
L {
|
||||
if (out[l] == vals[u * 32 + l]) ++equal;
|
||||
else if (!firstPrinted) { firstPrinted = 1; printf("first mismatch: base %u lane %d: cpu %016llx gpu %016llx\n", bases[u * 32], l, (unsigned long long)out[l], (unsigned long long)vals[u * 32 + l]); }
|
||||
}
|
||||
}
|
||||
double v1 = nowMs();
|
||||
printf("%zu of %zu lanes equal (%s) [%.3f ms per unit during the check]\n", equal, n, equal == n ? "PASS" : "FAIL", (v1 - v0) / units);
|
||||
|
||||
if (doTime) {
|
||||
int Rs[2] = { 0, Rsteps }; double ms[2] = { 0, 0 };
|
||||
for (int t = 0; t < 2; ++t) {
|
||||
uint64_t out[32]; volatile uint64_t sink = 0;
|
||||
hash_unit(cache, bases[0], Rs[t], out); // warm
|
||||
double t0 = nowMs();
|
||||
for (size_t u = 0; u < units; ++u) { hash_unit(cache, bases[u * 32], Rs[t], out); sink ^= out[0]; }
|
||||
double t1 = nowMs();
|
||||
ms[t] = (t1 - t0) / units;
|
||||
(void)sink;
|
||||
}
|
||||
printf("verifier: %.3f ms per unit at R=0, %.3f ms per unit at R=%d, mm8 block delta %.3f ms per unit (%.2f us per mm8 step), averaged over %zu units, one core\n",
|
||||
ms[0], ms[1], Rsteps, ms[1] - ms[0], Rsteps ? (ms[1] - ms[0]) * 1000.0 / (8.0 * Rsteps) : 0.0, units);
|
||||
printf("VERIFY R=%d equal=%zu of=%zu ms_r0=%.3f ms_r=%.3f delta=%.3f\n", Rsteps, equal, n, ms[0], ms[1], ms[1] - ms[0]);
|
||||
}
|
||||
free(cache); free(bases); free(vals);
|
||||
return equal == n ? 0 : 1;
|
||||
}
|
||||
219
proto-newpow/state-dataset/README.md
Normal file
219
proto-newpow/state-dataset/README.md
Normal file
|
|
@ -0,0 +1,219 @@
|
|||
# state-dataset: the dataset commits to chain state (class "sd1")
|
||||
|
||||
Horizon lane 8, new proof of work. Prototype and measurements, 6 October 2026, 19:34 to 19:50 UTC.
|
||||
Everything here is a TEST HARNESS: no pool, no network, no wallet, nothing touches the devnet.
|
||||
|
||||
## The scheme in one paragraph
|
||||
|
||||
The hash kernel is unchanged. What changes is the daily dataset build. Today item t of the 1 GiB dataset is
|
||||
`mh_item(cache, t)`: the 16-word state starts as `s[0..7] = K` (day key) and `s[8..15] = t * MUL[i] + RC[i]`, then
|
||||
8 rounds of (8 mixers + one dependent 64-byte cache read XORed in), then 8 final mixers. Under sd1 a 64-byte STATE
|
||||
LEAF for item t is XORed into those 16 initial words before the first mixer: `s[i] ^= leaf(t)[i]`. In the real design
|
||||
leaf(t) is the t-th 64-byte leaf of a canonical serialisation of the chain's execution state at a certified checkpoint
|
||||
20 minutes before the day boundary, zero padded where the state is shorter than the dataset. A miner therefore cannot
|
||||
build the day's dataset without the state, and a verifier that holds no dataset needs leaf(t) for every item it
|
||||
derives. The prototype measures what that costs (build time, device memory, verifier time) and buys (which leaves a
|
||||
hash touches, and what a light client would have to be shown).
|
||||
|
||||
### Leaf stand-in (synthetic, prototype only)
|
||||
|
||||
`leaf(t) = mh_chacha_block(x)` (the pack's own ChaCha12 block from memhard.h) with
|
||||
`x = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574, S[0..7], t, 0, 0x49676e65, 0x53746174)` and
|
||||
`S[i] = K[i] ^ 0x5a5a5a5a` (stand-in state root). For the pack's day key
|
||||
`S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102`. The leaf array is 2^24 items x 64 B = 1 GiB
|
||||
for the pack's 1 GiB dataset (2^28 words = 2^24 items). Definition in `sd.h`.
|
||||
|
||||
## Files
|
||||
|
||||
| file | what |
|
||||
|---|---|
|
||||
| `kernel.cu`, `memhard.h`, `program.h`, `vectors.h` | the mx8-genesis pack, copied unmodified from `proto-cuda/packs-ca2-mixer/mx8-genesis/` |
|
||||
| `sd.h` | `mh_leaf`, `mh_item_sd`, `mh_word_sd` (host and device, plain C compatible) |
|
||||
| `kernel_sd.cu` | the pack's kernel.cu plus `igneum_leaves`, `igneum_build_sd` and their launch wrappers appended; `igneum_hash` byte for byte the pack's (checked with diff) |
|
||||
| `bench.cu` | the harness: `--mode control|sd1`, `--dump <file> <nwarps>`, `--items-out <file>`, `--chunk-mib M`, `--power-seconds S` |
|
||||
| `verify_sd.c` | plain C: host cache fill, item bit-exact check from the GPU's items file, 32-lane register-major interpreter of the program against the GPU's dump, lane-0 load trace, per-unit rows |
|
||||
| `cpu_rows.c` | plain C: the igneum-build-1 rows B.1 and B.2 (OpenMP for the 32-thread row) |
|
||||
| `run_gpu.sh`, `run_cpu.sh` | the exact commands, as run |
|
||||
| `results/gpu/`, `results/cpu/` | every log, the dumps and the items files, copied back from the boxes |
|
||||
|
||||
## Boxes
|
||||
|
||||
- GPU box 2: NVIDIA GeForce RTX 4090, 24 GB (24083 MiB reported), 128 SMs, driver 595.91.07 (CUDA driver 13.2), nvcc 12.8.93,
|
||||
`-arch=sm_89`, power limit 450 W, Ubuntu 24.04.1. Host CPU AMD EPYC 7352 (48 threads). A CPU-only pool daemon shares the
|
||||
host (ports 4463/4480); the plain C rows there were pinned to core 2. Working directory `/root/horizon-newpow/state-dataset`.
|
||||
- CPU box igneum-build-1: AMD EPYC 9454P (48 cores, 96 threads), 128 GB, gcc 13.3.0, L3 256 MiB. Shared with other agents'
|
||||
builds: load average 19 to 32 during the readings (printed in each log). Rows pinned to core 4 under `nice -n 19`; the
|
||||
32-thread row on cores 4 to 35. Working directory `/srv/builds/horizon-newpow/state-dataset`.
|
||||
|
||||
## Commands
|
||||
|
||||
From the Mac (zsh; the GPU box has no rsync, tar over ssh instead):
|
||||
|
||||
```
|
||||
cd /Users/joshm/Projects/igneum-wt-horizon/proto-newpow/state-dataset
|
||||
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.cu *.h *.c *.sh | ssh -i ~/.ssh/igneum-fleet -p <box-2-port> root@<box-2-ip> 'cd /root/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_gpu.sh > run_gpu.log 2>&1'
|
||||
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.c *.h *.sh | ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 'cd /srv/builds/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_cpu.sh > run_cpu.log 2>&1'
|
||||
```
|
||||
|
||||
On GPU box 2 (`run_gpu.sh`):
|
||||
|
||||
```
|
||||
export PATH=/usr/local/cuda/bin:$PATH
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
|
||||
gcc -O2 -std=c11 -o verify_sd verify_sd.c
|
||||
./bench --mode control --batches 10 --items-out items_control.txt --dump dump_control.txt 4 --power-seconds 20
|
||||
./bench --mode sd1 --batches 10 --items-out items_sd1.txt --dump dump_sd1.txt 4 --power-seconds 20
|
||||
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 64
|
||||
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 256
|
||||
./bench --mode control --batches 10 --power-seconds 20 # second pass
|
||||
./bench --mode sd1 --batches 10 --power-seconds 20 # second pass
|
||||
taskset -c 2 ./verify_sd --mode control --items items_control.txt --dump dump_control.txt
|
||||
taskset -c 2 ./verify_sd --mode sd1 --items items_sd1.txt --dump dump_sd1.txt
|
||||
```
|
||||
|
||||
On igneum-build-1 (`run_cpu.sh`):
|
||||
|
||||
```
|
||||
gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
|
||||
gcc -O2 -std=c11 -o verify_sd verify_sd.c
|
||||
taskset -c 4 nice -n 19 ./cpu_rows --rows # B.1 (i)(ii)(iii) and the one-core B.2 rows, twice
|
||||
taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 # B.2 leaf array on 32 threads, twice
|
||||
taskset -c 4 nice -n 19 ./verify_sd --mode sd1 # the per-unit rows of verify_sd.c on this CPU
|
||||
```
|
||||
|
||||
Timing method: cache fill, leaf array and dataset build are CUDA-event times of the second of two launches (as
|
||||
host.cu does). MH/s is CUDA-event time over 10 batches of 2^24 after a warm-up batch at base 0. Watts and SM MHz are
|
||||
the mean of `nvidia-smi -l 1` samples taken after the first 10 s of a 20 s window in which the hash kernel runs back
|
||||
to back (10 samples used of 22). Fingerprint = FNV-1a 64 over the 2^24 little-endian u64 outputs at base nonce 0,
|
||||
the project's definition.
|
||||
|
||||
## RESULTS
|
||||
|
||||
### A. GPU, RTX 4090 (control vs sd1, two passes each)
|
||||
|
||||
| row | control | sd1 | note |
|
||||
|---|---|---|---|
|
||||
| cache fill, GPU, second pass (ms) | 1.81, 1.85 | 1.85, 1.85 | 256 MiB, unchanged kernel |
|
||||
| leaf array, GPU, second pass (ms) | none | 5.49, 5.49 | 2^24 ChaCha12 blocks, 1 GiB written at 195 GB/s |
|
||||
| dataset build, second pass (ms) | 30.55, 30.55 | 31.98, 31.98 | +1.43 ms (+4.7%): one coalesced 64 B read per item |
|
||||
| build + leaves (ms) | 30.55 | 37.47 | synthetic leaves; a real snapshot arrives from the node instead (see chunked rows) |
|
||||
| the 3 Mac vectors, standalone and in batch | PASS, PASS | n/a (new values) | sd1 base 0 lane 0 = b600edbed969becc, base 4096 = a533e89c78bb6b74, base 1000000 = beb4cb0c6bab8163 |
|
||||
| dataset head, word [MASK], 64 Mac samples | PASS | n/a (new values) | sd1 dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf |
|
||||
| 64 random words vs host derivation | PASS (mh_word) | PASS (mh_word_sd) | |
|
||||
| 2^24 fingerprint at base 0 | 7c28cfb06c5c65a9 (= the pack's, both passes) | d5b0c16390cad0e8 (new, four runs) | |
|
||||
| hash rate, 10 batches of 2^24 (MH/s) | 63.083, 63.083 | 63.088, 63.087 | equal within 0.01%; the hash kernel is the same binary over a dataset of the same shape |
|
||||
| hash rate in the 20 s power window (MH/s) | 63.078, 63.079 | 63.083, 63.083 | |
|
||||
| watts, mean after 10 s | 207.6, 205.2 | 207.9, 206.5 | run to run noise 2 W |
|
||||
| SM MHz, mem MHz | 2745, 10251 | 2745, 10251 | |
|
||||
| MH/s per W | 0.304, 0.307 | 0.303, 0.305 | |
|
||||
| hash kernel registers | 29 | 29 | 24 resident blocks/SM at 1 warp/block |
|
||||
|
||||
Bit-exactness of the build (A.3): 1,024 random item indices (SplitMix64 seeded 0x5d1, `t = z & (2^24 - 1)`), the 16 GPU
|
||||
words of each item against the host derivation in plain C on the box's CPU with the host-filled cache (65536
|
||||
`mh_cache_segment` calls, 380 ms in the bench, 573 to 587 ms pinned to core 2 in verify_sd):
|
||||
|
||||
| check | control | sd1 |
|
||||
|---|---|---|
|
||||
| items equal, in-process (bench.cu, host mh_item / mh_item_sd) | 1024 of 1024 | 1024 of 1024 |
|
||||
| items equal, plain C (verify_sd.c from the items file) | 1024 of 1024 | 1024 of 1024 |
|
||||
| items equal after the chunked rebuild (64 MiB chunks; 256 MiB chunks) | n/a | 1024 of 1024; 1024 of 1024 |
|
||||
| 64 device leaves vs host mh_leaf (incl. t = 0 and 2^24 - 1) | n/a | PASS |
|
||||
| host cache FNV-1a 64 | 48c4f5bf24166b2e = Mac | 48c4f5bf24166b2e = Mac |
|
||||
|
||||
GPU hash outputs vs the C interpreter (4 warps dumped, bases 0, 32, 64, 96; loads through mh_word / mh_word_sd):
|
||||
|
||||
| | control | sd1 |
|
||||
|---|---|---|
|
||||
| lanes equal | 128 of 128 | 128 of 128 |
|
||||
| interpreter vs the Mac vector at base 0 | 32 of 32 | n/a |
|
||||
| interpreting time (4,096 item derivations per warp) | 37 ms | 42 ms |
|
||||
|
||||
Device memory (A.4), cudaMemGetInfo, MiB used (context 395 included):
|
||||
|
||||
| phase | control | sd1 resident leaves | sd1 chunked 64 MiB | sd1 chunked 256 MiB |
|
||||
|---|---|---|---|---|
|
||||
| during the build | 1675 | 2699 | 1741 | 1933 |
|
||||
| while hashing (leaves freed) | 1803 | 1803 | 1803 | 1803 |
|
||||
| chunked rebuild total, copies + builds, one stream (ms) | | | 75.80 | 75.47 |
|
||||
|
||||
Whole-array host to device copy of the 1 GiB leaf array from pinned memory: 62 ms = 17.2 GB/s (this box's PCIe link
|
||||
under load; device to host 55 ms). So the chunked build is PCIe-bound: 62 ms of copy plus the 32 ms build, partly
|
||||
serialised on one stream, gives 76 ms. Two chunk buffers on two streams would hide most of the build under the copy
|
||||
(not done; the kernel already takes an item range `[t0, t0 + n)` with the chunk's leaves at `leaves - 16 t0`, so
|
||||
chunking needed no kernel change at all, only the loop in bench.cu).
|
||||
|
||||
### B. CPU, igneum-build-1 (EPYC 9454P, one core pinned, nice 19), ms per unit of 4,096 items, 100 units
|
||||
|
||||
Two readings, because the box is shared: reading 1 at load 19, reading 2 at load 31.
|
||||
|
||||
| row | reading 1 | reading 2 | per item |
|
||||
|---|---|---|---|
|
||||
| (i) 4,096 random 64 B reads from a 2 GiB resident leaf array | 0.108 ms (min 0.092, max 0.172) | 0.163 ms | 26 to 40 ns |
|
||||
| (i) the same from an 8 GiB array (the year-12 size) | 0.139 ms (min 0.127, max 0.183) | 0.209 ms | 34 to 51 ns |
|
||||
| (ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array) | 0.284 ms | 0.285 ms | 69 ns |
|
||||
| (iii) 4,096 x mh_item on the host cache, naive, no interleaving | 9.01 ms (min 8.91, max 9.60) | 11.24 ms (min 9.38, max 13.51) | 2.2 to 2.7 us |
|
||||
| 4,096 x (leaf + mh_item_sd), naive | 9.31 ms | 11.64 ms | |
|
||||
|
||||
The same rows on GPU box 2's EPYC 7352, core 2 (verify_sd.c): (ii) 0.369 to 0.374 ms, (iii) 10.14 to 10.21 ms, leaf +
|
||||
mh_item_sd 10.56 to 10.58 ms.
|
||||
|
||||
The project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent
|
||||
cache misses across items; the naive (iii) figure here is an upper bound. The sd1 increment is the row that matters:
|
||||
0.11 to 0.21 ms per unit when the verifier holds the 2 to 8 GiB state in RAM, or 0.28 ms when it derives the stand-in
|
||||
leaf (in the real design it cannot derive a state leaf, it must hold the state or be shown it). Against the 2.06 ms
|
||||
interleaved verifier that is +5% to +14%; against the naive 9 to 11 ms it is +1% to +3%.
|
||||
|
||||
B.2, the daily snapshot build on CPU (igneum-build-1):
|
||||
|
||||
| row | time |
|
||||
|---|---|
|
||||
| 1 GiB leaf array, 2^24 ChaCha12 blocks, one core | 1.670 s, 1.677 s (100 ns per leaf) |
|
||||
| the same on 32 threads (OpenMP), second pass (first pass 0.172 s with page faults) | 0.073 s, 0.083 s |
|
||||
| host cache fill, 256 MiB, one core | 0.447 s, 0.452 s (0.380 s on the 4090 box's host, unpinned; 0.57 s pinned beside the pool daemon) |
|
||||
|
||||
B.3, the state sample a light client would check (lane 0 of warp base 0, from the instrumented interpreter):
|
||||
|
||||
| | control | sd1 |
|
||||
|---|---|---|
|
||||
| distinct words touched by the 128 loads | 128 | 128 |
|
||||
| distinct items touched | 128 of 2^24 | 128 of 2^24 |
|
||||
| first 8 load words | 8528400 43034175 255822847 136709537 ... | 8528400 85793927 255822847 180392080 ... |
|
||||
|
||||
Loads 1 and 3 share their address in both modes because their addresses are fixed by the nonce before any load's
|
||||
value feeds back; from load 2 on the sd1 trail diverges. Merkle arithmetic for a 2 GiB state (2^25 leaves of 64 B,
|
||||
depth 25, 32-byte hashes): 128 openings x 25 x 32 B = 102,400 B of siblings + 128 x 64 B = 8,192 B of leaves =
|
||||
110,592 B (108 KiB) per lane, uncompressed; a 32-lane unit is 32 x 108 KiB = 3.4 MiB, before deduplicating shared
|
||||
upper levels (128 random paths in a depth-25 tree share only their top seven levels, so dedup saves about 20% per lane,
|
||||
approximate; across the 32 lanes of a unit the same seven levels are shared again). Every one of the 128 loads is a distinct item, so there is no saving from repeated items.
|
||||
|
||||
## What the numbers mean, per tier (the consequences rule)
|
||||
|
||||
- Hash rate, watts, MH/s per W: unchanged for every tier on every vendor, because the hash kernel is the pack's and the
|
||||
dataset has the same shape. Nothing to do.
|
||||
- Daily build: +1.4 ms on the 4090 (30.6 to 32.0 ms) when the leaves are resident, 76 ms when streamed from the host
|
||||
over PCIe in chunks. At the designed 2 GiB dataset the leaf array is 2 GiB and the streamed build is about 150 ms
|
||||
(approximate, scaling the 17 GB/s copy); on a PCIe 3 x8 slot in a rig, about 4x that (approximate). All far inside the
|
||||
daily window on every tier.
|
||||
- Device memory: with resident leaves the build peak at 2 GiB is 2 GiB + 256 MiB + 2 GiB + context, about 4.6 GiB
|
||||
(approximate), which fits an 8 GB card but not beside a prover. Streamed in 64 MiB chunks the peak is dataset + cache
|
||||
+ 64 MiB + context, measured 1741 MiB at 1 GiB here, about 2.7 GiB at 2 GiB (approximate): no tier loses memory it has
|
||||
today. Recommendation: ship chunked only; never hold the leaf array on the device.
|
||||
- The real cost is delivery: every miner needs the 1 to 2 GiB state serialisation once a day. A miner
|
||||
beside its own node reads it from disk or loopback (seconds). A pool user without a node must get it from the pool
|
||||
(2 GiB per day per miner, or a shared download), and a light verifier either holds the state (0.11 to 0.21 ms per unit
|
||||
extra, measured) or is shown 3.4 MiB of Merkle openings per unit (arithmetic above), which is not a light client any
|
||||
more. The design decision this prototype leaves open is which of those two the protocol asks of a header verifier.
|
||||
- CPU snapshot build: 1.7 s on one core, 0.07 s on 32, for the synthetic leaves; the real serialisation is bounded by the
|
||||
node's state read, not by this.
|
||||
|
||||
## What failed, what was cut
|
||||
|
||||
- Nothing on the list was cut. verify_sd.c's lane-0 trace printed nothing on the first run (the trace counter was reset
|
||||
on every warp, so it read 0 after warp 3); fixed, the verifiers re-run, the fix is in the file.
|
||||
- The GPU box has no rsync; the sources went over with tar through ssh (`COPYFILE_DISABLE=1 --no-xattrs`, the Mac's
|
||||
tar otherwise writes Apple xattr headers that GNU tar warns about).
|
||||
- `-std=c11` hides `clock_gettime`; both C files define `_POSIX_C_SOURCE 200809L`.
|
||||
- igneum-build-1 was under other agents' load (19 to 32) for every reading; both readings are given. The 4090 box's
|
||||
CPU rows ran beside the pool daemon, pinned to core 2. The GPU itself was idle apart from this bench.
|
||||
- Not measured: a two-stream chunked build (copy and build overlapped), AMD and Apple builds, the 2 GiB dataset
|
||||
itself (the pack is 1 GiB; the 2 GiB figures above are labelled approximate).
|
||||
464
proto-newpow/state-dataset/bench.cu
Normal file
464
proto-newpow/state-dataset/bench.cu
Normal file
|
|
@ -0,0 +1,464 @@
|
|||
// state-dataset prototype bench (Horizon lane 8, class sd1). TEST HARNESS ONLY: no pool, no network, no wallet.
|
||||
//
|
||||
// Shape follows proto-cuda/host.cu: cache fill (GPU, twice, CUDA events), host cache fill and check, dataset build
|
||||
// (twice, second pass reported), dataset self-test, the 3 Mac vectors, a warm-up batch of 2^24 at base nonce 0
|
||||
// (fingerprinted: FNV-1a 64 over the 2^24 little-endian u64 outputs), N timed batches (CUDA events), then a power
|
||||
// window where nvidia-smi samples at 1 Hz while the hash kernel runs back to back.
|
||||
//
|
||||
// --mode control the unmodified pack (igneum_build)
|
||||
// --mode sd1 leaf array (igneum_leaves) + igneum_build_sd; the hash kernel is the pack's
|
||||
// --dump <file> <n> write the outputs of the first n warps of the base-0 batch (for verify_sd.c)
|
||||
// --items-out <file> write the 16 words of 1,024 random items (SplitMix64 seeded 0x5d1) read back from the GPU
|
||||
// --batches N timed batches after the warm-up (default 10)
|
||||
// --batch-log2 B nonces per batch (default 24)
|
||||
// --power-seconds S length of the nvidia-smi window (default 20; 0 skips it)
|
||||
// --chunk-mib M sd1 only: after the resident build, rebuild with the leaf array streamed from pinned host
|
||||
// memory in M MiB chunks (the shape a miner uses when the leaves do not fit beside the dataset)
|
||||
// --device D
|
||||
//
|
||||
// Build: nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
|
||||
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <chrono>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
#include <thread>
|
||||
#include <mutex>
|
||||
|
||||
#include "program.h"
|
||||
#include "vectors.h"
|
||||
#include "memhard.h"
|
||||
#include "sd.h"
|
||||
|
||||
cudaError_t igneum_launch_leaves(uint32_t* leaves, uint32_t nItems);
|
||||
cudaError_t igneum_launch_build_sd(uint32_t* ds, const uint32_t* cache, const uint32_t* leaves, uint32_t t0, uint32_t nItems);
|
||||
|
||||
#define CUDA_CHECK(call) do { cudaError_t err_ = (call); if (err_ != cudaSuccess) { \
|
||||
std::fprintf(stderr, "CUDA error: %s (%d)\n at %s:%d\n in %s\n", cudaGetErrorString(err_), (int)err_, __FILE__, __LINE__, #call); \
|
||||
std::exit(2); } } while (0)
|
||||
|
||||
static double wallMs() {
|
||||
using namespace std::chrono;
|
||||
return duration<double, std::milli>(steady_clock::now().time_since_epoch()).count();
|
||||
}
|
||||
static uint64_t fnv1a64(const void* p, size_t n) {
|
||||
const uint8_t* b = (const uint8_t*)p;
|
||||
uint64_t h = 0xcbf29ce484222325ull;
|
||||
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
|
||||
return h;
|
||||
}
|
||||
static uint64_t splitmix64(uint64_t& s) {
|
||||
s += 0x9E3779B97F4A7C15ull;
|
||||
uint64_t z = s;
|
||||
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull;
|
||||
z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
|
||||
return z ^ (z >> 31);
|
||||
}
|
||||
static float eventMs(cudaEvent_t a, cudaEvent_t b) { float ms = 0.f; CUDA_CHECK(cudaEventElapsedTime(&ms, a, b)); return ms; }
|
||||
|
||||
struct Opts {
|
||||
bool sd1 = false;
|
||||
int batches = 10, batchLog2 = 24, device = 0, powerSeconds = 20, dumpWarps = 0, chunkMib = 0;
|
||||
std::string dump, itemsOut;
|
||||
};
|
||||
|
||||
static void usage() {
|
||||
std::printf("bench --mode control|sd1 [--dump <file> <nwarps>] [--items-out <file>] [--batches 10] [--batch-log2 24] [--power-seconds 20] [--chunk-mib M] [--device 0]\n");
|
||||
}
|
||||
|
||||
static Opts parse(int argc, char** argv) {
|
||||
Opts o;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
std::string a = argv[i];
|
||||
auto need = [&](int k) { if (i + k >= argc) { usage(); std::exit(2); } };
|
||||
if (a == "--mode") { need(1); std::string m = argv[++i]; if (m == "sd1") o.sd1 = true; else if (m == "control") o.sd1 = false; else { usage(); std::exit(2); } }
|
||||
else if (a == "--dump") { need(2); o.dump = argv[++i]; o.dumpWarps = std::atoi(argv[++i]); }
|
||||
else if (a == "--items-out") { need(1); o.itemsOut = argv[++i]; }
|
||||
else if (a == "--batches") { need(1); o.batches = std::atoi(argv[++i]); }
|
||||
else if (a == "--batch-log2") { need(1); o.batchLog2 = std::atoi(argv[++i]); }
|
||||
else if (a == "--power-seconds") { need(1); o.powerSeconds = std::atoi(argv[++i]); }
|
||||
else if (a == "--device") { need(1); o.device = std::atoi(argv[++i]); }
|
||||
else if (a == "--chunk-mib") { need(1); o.chunkMib = std::atoi(argv[++i]); }
|
||||
else if (a == "-h" || a == "--help") { usage(); std::exit(0); }
|
||||
else { std::printf("unknown argument %s\n", argv[i]); usage(); std::exit(2); }
|
||||
}
|
||||
if (o.batchLog2 < 10 || o.batchLog2 > 28 || o.batches < 1) { usage(); std::exit(2); }
|
||||
return o;
|
||||
}
|
||||
|
||||
// nvidia-smi sampler: one line per second, read on its own thread until `timeout` ends the process.
|
||||
struct Sample { double t; double watts; double smMHz; double memMHz; };
|
||||
struct Sampler {
|
||||
std::vector<Sample> samples;
|
||||
std::mutex mu;
|
||||
std::thread th;
|
||||
double t0 = 0;
|
||||
void start(int device, int seconds) {
|
||||
t0 = wallMs();
|
||||
char cmd[512];
|
||||
std::snprintf(cmd, sizeof(cmd), "timeout %d nvidia-smi -i %d --query-gpu=power.draw,clocks.sm,clocks.mem --format=csv,noheader,nounits -l 1 2>/dev/null", seconds + 2, device);
|
||||
std::string c = cmd;
|
||||
th = std::thread([this, c]() {
|
||||
FILE* f = popen(c.c_str(), "r");
|
||||
if (!f) return;
|
||||
char line[256];
|
||||
while (std::fgets(line, sizeof(line), f)) {
|
||||
double w = 0, sm = 0, mem = 0;
|
||||
if (std::sscanf(line, "%lf , %lf , %lf", &w, &sm, &mem) == 3) {
|
||||
std::lock_guard<std::mutex> g(mu);
|
||||
samples.push_back({ wallMs() - t0, w, sm, mem });
|
||||
}
|
||||
}
|
||||
pclose(f);
|
||||
});
|
||||
}
|
||||
void join() { if (th.joinable()) th.join(); }
|
||||
};
|
||||
|
||||
static const uint32_t KEYW[8] = IGNEUM_KEY_INIT;
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
Opts o = parse(argc, argv);
|
||||
const char* modeName = o.sd1 ? "sd1" : "control";
|
||||
std::printf("state-dataset bench mode %s pack \"%s\" (test harness: no pool, no network, no wallet)\n", modeName, IGNEUM_SEED_STRING);
|
||||
|
||||
int count = 0;
|
||||
CUDA_CHECK(cudaGetDeviceCount(&count));
|
||||
if (count == 0) { std::printf("FAIL: no CUDA device\n"); return 2; }
|
||||
CUDA_CHECK(cudaSetDevice(o.device));
|
||||
cudaDeviceProp prop;
|
||||
std::memset(&prop, 0, sizeof(prop));
|
||||
CUDA_CHECK(cudaGetDeviceProperties(&prop, o.device));
|
||||
int drv = 0, rt = 0;
|
||||
CUDA_CHECK(cudaDriverGetVersion(&drv));
|
||||
CUDA_CHECK(cudaRuntimeGetVersion(&rt));
|
||||
std::printf("GPU: %s (%d SMs, cc %d.%d, %.0f MiB), CUDA driver %d.%d runtime %d.%d\n", prop.name, prop.multiProcessorCount, prop.major, prop.minor,
|
||||
(double)prop.totalGlobalMem / 1048576.0, drv / 1000, (drv % 100) / 10, rt / 1000, (rt % 100) / 10);
|
||||
int regs = 0, bps = 0;
|
||||
CUDA_CHECK(igneum_hash_info(®s, &bps, 1u));
|
||||
std::printf("hash kernel: %d registers/thread, %d resident blocks/SM at 1 warp/block\n", regs, bps);
|
||||
if (o.sd1) {
|
||||
std::printf("sd1 leaf stand-in: S[i] = K[i] ^ 0x%08x -> S =", SD_STATE_ROOT_XOR);
|
||||
for (int i = 0; i < 8; ++i) std::printf(" %08x", KEYW[i] ^ SD_STATE_ROOT_XOR);
|
||||
std::printf("; leaf(t) = ChaCha12 block of (sigma, S, t, 0, %08x, %08x)\n", SD_TAG0, SD_TAG1);
|
||||
}
|
||||
|
||||
size_t free0 = 0, total = 0;
|
||||
CUDA_CHECK(cudaMemGetInfo(&free0, &total));
|
||||
std::printf("device memory at start: %.0f MiB used of %.0f MiB (context)\n", (double)(total - free0) / 1048576.0, (double)total / 1048576.0);
|
||||
|
||||
cudaEvent_t e0, e1;
|
||||
CUDA_CHECK(cudaEventCreate(&e0));
|
||||
CUDA_CHECK(cudaEventCreate(&e1));
|
||||
|
||||
// ---- cache: GPU fill twice, host fill, check
|
||||
const uint32_t cacheWords = 1u << IGNEUM_CACHE_LOG2_WORDS;
|
||||
const size_t cacheBytes = (size_t)cacheWords * 4u;
|
||||
uint32_t* dCache = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dCache, cacheBytes));
|
||||
double cacheFill[2] = { 0, 0 };
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(igneum_launch_cache_fill(dCache, IGNEUM_CACHE_SEGMENTS));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
cacheFill[pass] = eventMs(e0, e1);
|
||||
}
|
||||
std::printf("cache fill (GPU): %.2f ms first, %.2f ms second (%u segments x 64 ChaCha12 blocks, %u MiB)\n", cacheFill[0], cacheFill[1], (unsigned)IGNEUM_CACHE_SEGMENTS, (unsigned)(cacheBytes >> 20));
|
||||
std::vector<uint32_t> hCache(cacheWords);
|
||||
double h0 = wallMs();
|
||||
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(hCache.data(), seg);
|
||||
double hostCacheMs = wallMs() - h0;
|
||||
bool cachePass = false;
|
||||
{
|
||||
std::vector<uint32_t> dev(cacheWords);
|
||||
CUDA_CHECK(cudaMemcpy(dev.data(), dCache, cacheBytes, cudaMemcpyDeviceToHost));
|
||||
bool same = std::memcmp(dev.data(), hCache.data(), cacheBytes) == 0;
|
||||
uint64_t fnv = fnv1a64(hCache.data(), cacheBytes);
|
||||
cachePass = same && fnv == IGNEUM_CACHE_FNV64;
|
||||
std::printf("cache fill (host, one thread): %.1f ms; cache check: %s (GPU == host %s, host FNV-1a 64 %016llx vs Mac %016llx)\n",
|
||||
hostCacheMs, cachePass ? "PASS" : "FAIL", same ? "PASS" : "FAIL", (unsigned long long)fnv, (unsigned long long)IGNEUM_CACHE_FNV64);
|
||||
}
|
||||
|
||||
// ---- sd1: leaf array
|
||||
const uint32_t words = 1u << IGNEUM_DATASET_LOG2;
|
||||
const uint32_t mask = words - 1u;
|
||||
const uint32_t nItems = words / 16u;
|
||||
uint32_t* dLeaves = nullptr;
|
||||
double leavesMs[2] = { 0, 0 };
|
||||
bool leafPass = true;
|
||||
if (o.sd1) {
|
||||
CUDA_CHECK(cudaMalloc((void**)&dLeaves, (size_t)nItems * 64u));
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(igneum_launch_leaves(dLeaves, nItems));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
leavesMs[pass] = eventMs(e0, e1);
|
||||
}
|
||||
std::printf("leaf array (GPU, %u x 64 B = %u MiB): %.2f ms first, %.2f ms second -> %.1f GB/s written\n", nItems, (unsigned)(((size_t)nItems * 64u) >> 20),
|
||||
leavesMs[0], leavesMs[1], (double)nItems * 64.0 / 1e9 / (leavesMs[1] / 1000.0));
|
||||
// 64 leaves against the host derivation
|
||||
uint64_t s = 0x1ea7ull;
|
||||
int bad = 0;
|
||||
for (int k = 0; k < 64; ++k) {
|
||||
uint32_t t = (k == 0) ? 0u : (k == 1) ? nItems - 1u : (uint32_t)splitmix64(s) & (nItems - 1u);
|
||||
uint32_t got[16], want[16];
|
||||
CUDA_CHECK(cudaMemcpy(got, dLeaves + (size_t)t * 16u, 64, cudaMemcpyDeviceToHost));
|
||||
mh_leaf(t, want);
|
||||
if (std::memcmp(got, want, 64) != 0) { if (bad == 0) std::printf(" leaf[%u] differs: gpu %08x host %08x\n", t, got[0], want[0]); ++bad; }
|
||||
}
|
||||
leafPass = (bad == 0);
|
||||
std::printf("leaf check: %s (64 leaves incl. 0 and %u vs host mh_leaf)\n", leafPass ? "PASS" : "FAIL", nItems - 1u);
|
||||
}
|
||||
|
||||
// ---- dataset build, twice
|
||||
uint32_t* dDs = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dDs, (size_t)words * 4u));
|
||||
double buildMs[2] = { 0, 0 };
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
if (o.sd1) CUDA_CHECK(igneum_launch_build_sd(dDs, dCache, dLeaves, 0u, nItems));
|
||||
else CUDA_CHECK(igneum_launch_build(dDs, dCache, nItems));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
buildMs[pass] = eventMs(e0, e1);
|
||||
}
|
||||
std::printf("dataset build (%s): %.2f ms first, %.2f ms second -> %.1f M items/s\n", o.sd1 ? "igneum_build_sd, leaf XOR before the first mixer" : "igneum_build, the pack's",
|
||||
buildMs[0], buildMs[1], (double)nItems / 1e6 / (buildMs[1] / 1000.0));
|
||||
size_t freeB = 0;
|
||||
CUDA_CHECK(cudaMemGetInfo(&freeB, &total));
|
||||
double usedBuildMiB = (double)(total - freeB) / 1048576.0;
|
||||
std::printf("device memory after the build: %.0f MiB used (context %.0f + cache %u + %sdataset %u MiB)\n", usedBuildMiB, (double)(total - free0) / 1048576.0,
|
||||
(unsigned)(cacheBytes >> 20), o.sd1 ? "leaves 1024 + " : "", (unsigned)(((size_t)words * 4u) >> 20));
|
||||
|
||||
// ---- dataset self-test
|
||||
bool dsPass = true;
|
||||
{
|
||||
uint32_t head[16];
|
||||
CUDA_CHECK(cudaMemcpy(head, dDs, sizeof(head), cudaMemcpyDeviceToHost));
|
||||
std::printf("dataset[0..3] = %08x %08x %08x %08x", head[0], head[1], head[2], head[3]);
|
||||
if (!o.sd1) {
|
||||
bool headOk = std::memcmp(head, IGNEUM_DS_HEAD, 64) == 0;
|
||||
uint32_t last = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&last, dDs + IGNEUM_DS_LAST_INDEX, 4, cudaMemcpyDeviceToHost));
|
||||
bool lastOk = (last == IGNEUM_DS_LAST);
|
||||
int badSample = 0;
|
||||
for (int k = 0; k < IGNEUM_DS_SAMPLES; ++k) {
|
||||
uint32_t v = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&v, dDs + IGNEUM_DS_SAMPLE_INDEX[k], 4, cudaMemcpyDeviceToHost));
|
||||
if (v != IGNEUM_DS_SAMPLE_VALUE[k]) ++badSample;
|
||||
}
|
||||
dsPass = headOk && lastOk && badSample == 0;
|
||||
std::printf(" head 16 vs Mac %s, word [MASK] vs Mac %s, 64 Mac samples %s\n", headOk ? "PASS" : "FAIL", lastOk ? "PASS" : "FAIL", badSample == 0 ? "PASS" : "FAIL");
|
||||
} else {
|
||||
std::printf(" (sd1: new values, no Mac expectation)\n");
|
||||
}
|
||||
// 64 random words vs the host derivation (host.cu's points)
|
||||
int badRnd = 0;
|
||||
uint64_t s = 0x9E3779B97F4A7C15ull ^ (uint64_t)words;
|
||||
for (int k = 0; k < 64; ++k) {
|
||||
uint32_t idx = (uint32_t)splitmix64(s) & mask;
|
||||
uint32_t v = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&v, dDs + idx, 4, cudaMemcpyDeviceToHost));
|
||||
uint32_t want = o.sd1 ? mh_word_sd(hCache.data(), idx) : mh_word(hCache.data(), idx);
|
||||
if (v != want) ++badRnd;
|
||||
}
|
||||
dsPass = dsPass && badRnd == 0;
|
||||
std::printf("dataset self-test: %s (64 random words vs host %s: %s)\n", dsPass ? "PASS" : "FAIL", o.sd1 ? "mh_word_sd" : "mh_word", badRnd == 0 ? "PASS" : "FAIL");
|
||||
}
|
||||
|
||||
// ---- 1,024 random items (SplitMix64 seeded 0x5d1) read back and compared with the host item derivation
|
||||
int itemsEqual = 0;
|
||||
{
|
||||
FILE* f = o.itemsOut.empty() ? nullptr : std::fopen(o.itemsOut.c_str(), "w");
|
||||
if (f) std::fprintf(f, "mode %s\nnitems 1024\n", modeName);
|
||||
uint64_t s = 0x5d1ull;
|
||||
for (int k = 0; k < 1024; ++k) {
|
||||
uint32_t t = (uint32_t)splitmix64(s) & (nItems - 1u);
|
||||
uint32_t got[16], want[16];
|
||||
CUDA_CHECK(cudaMemcpy(got, dDs + (size_t)t * 16u, 64, cudaMemcpyDeviceToHost));
|
||||
if (o.sd1) { uint32_t leaf[16]; mh_leaf(t, leaf); mh_item_sd(hCache.data(), leaf, t, want); }
|
||||
else mh_item(hCache.data(), t, want);
|
||||
if (std::memcmp(got, want, 64) == 0) ++itemsEqual;
|
||||
else if (itemsEqual == k) std::printf(" item %u differs: gpu %08x host %08x (first difference)\n", t, got[0], want[0]);
|
||||
if (f) { std::fprintf(f, "%u", t); for (int i = 0; i < 16; ++i) std::fprintf(f, " %08x", got[i]); std::fprintf(f, "\n"); }
|
||||
}
|
||||
if (f) std::fclose(f);
|
||||
std::printf("item bit-exactness (in-process, host %s on the host cache): %d of 1024 items equal\n", o.sd1 ? "mh_item_sd with host mh_leaf" : "mh_item", itemsEqual);
|
||||
}
|
||||
|
||||
// ---- sd1, chunked: the leaf array lives in pinned host memory (as a real snapshot would, delivered by the node)
|
||||
// and is streamed to the device in chunks; the dataset is rebuilt chunk by chunk and re-checked. Peak device memory
|
||||
// is then cache + dataset + one chunk. One stream, so copy and build serialise; two chunk buffers on two streams
|
||||
// would overlap them (not done here).
|
||||
double chunkedMs = 0, h2dMs = 0, usedChunkMiB = 0;
|
||||
int chunkedEqual = -1;
|
||||
if (o.sd1 && o.chunkMib > 0) {
|
||||
const size_t leafBytes = (size_t)nItems * 64u;
|
||||
const size_t chunkBytes = (size_t)o.chunkMib << 20;
|
||||
const uint32_t chunkItems = (uint32_t)(chunkBytes / 64u);
|
||||
if (chunkBytes > leafBytes || (leafBytes % chunkBytes) != 0) { std::printf("FAIL: --chunk-mib must divide 1024\n"); return 2; }
|
||||
uint32_t* hLeaves = nullptr;
|
||||
CUDA_CHECK(cudaMallocHost((void**)&hLeaves, leafBytes));
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(cudaMemcpy(hLeaves, dLeaves, leafBytes, cudaMemcpyDeviceToHost));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
double d2h = eventMs(e0, e1);
|
||||
CUDA_CHECK(cudaFree(dLeaves)); dLeaves = nullptr;
|
||||
// one-shot H2D of the whole array into the dataset buffer (overwritten by the rebuild anyway): the PCIe rate
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(cudaMemcpy(dDs, hLeaves, leafBytes, cudaMemcpyHostToDevice));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
h2dMs = eventMs(e0, e1);
|
||||
uint32_t* dChunk = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dChunk, chunkBytes));
|
||||
CUDA_CHECK(cudaMemGetInfo(&freeB, &total));
|
||||
usedChunkMiB = (double)(total - freeB) / 1048576.0;
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
for (uint32_t t0 = 0; t0 < nItems; t0 += chunkItems) {
|
||||
CUDA_CHECK(cudaMemcpyAsync(dChunk, hLeaves + (size_t)t0 * 16u, chunkBytes, cudaMemcpyHostToDevice, 0));
|
||||
CUDA_CHECK(igneum_launch_build_sd(dDs, dCache, dChunk, t0, chunkItems));
|
||||
}
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
chunkedMs = eventMs(e0, e1);
|
||||
// re-check the same 1,024 items
|
||||
uint64_t s = 0x5d1ull; chunkedEqual = 0;
|
||||
for (int k = 0; k < 1024; ++k) {
|
||||
uint32_t t = (uint32_t)splitmix64(s) & (nItems - 1u);
|
||||
uint32_t got[16], want[16], leaf[16];
|
||||
CUDA_CHECK(cudaMemcpy(got, dDs + (size_t)t * 16u, 64, cudaMemcpyDeviceToHost));
|
||||
mh_leaf(t, leaf); mh_item_sd(hCache.data(), leaf, t, want);
|
||||
if (std::memcmp(got, want, 64) == 0) ++chunkedEqual;
|
||||
}
|
||||
std::printf("chunked build (leaves from pinned host memory in %d MiB chunks, %u chunks, one stream): %.2f ms total (copies + builds); whole-array H2D alone %.2f ms = %.1f GB/s, D2H %.2f ms\n",
|
||||
o.chunkMib, nItems / chunkItems, chunkedMs, h2dMs, (double)leafBytes / 1e9 / (h2dMs / 1000.0), d2h);
|
||||
std::printf("device memory during the chunked build: %.0f MiB used (context + cache 256 + dataset 1024 + chunk %d MiB); items after the chunked rebuild: %d of 1024 equal\n", usedChunkMiB, o.chunkMib, chunkedEqual);
|
||||
CUDA_CHECK(cudaFree(dChunk));
|
||||
CUDA_CHECK(cudaFreeHost(hLeaves));
|
||||
}
|
||||
|
||||
// The leaf array is only needed for the build; a miner frees it before hashing.
|
||||
if (dLeaves) { CUDA_CHECK(cudaFree(dLeaves)); dLeaves = nullptr; }
|
||||
|
||||
// ---- vectors (control: against the Mac; sd1: printed)
|
||||
const uint32_t nonces = 1u << o.batchLog2;
|
||||
uint64_t* dOut = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dOut, (size_t)nonces * 8u));
|
||||
bool vecPass = true;
|
||||
{
|
||||
uint64_t got[32];
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, IGNEUM_VEC_BASE[w], mask, 32u, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
|
||||
if (!o.sd1) {
|
||||
int bad = 0;
|
||||
for (int l = 0; l < 32; ++l) if (got[l] != IGNEUM_VEC_OUT[w][l]) ++bad;
|
||||
vecPass = vecPass && bad == 0;
|
||||
std::printf("vector warp base %u: %s (%d of 32 lanes differ)\n", IGNEUM_VEC_BASE[w], bad == 0 ? "PASS" : "FAIL", bad);
|
||||
} else {
|
||||
std::printf("sd1 warp base %u lane 0: %016llx (new value; checked by verify_sd.c through the dump)\n", IGNEUM_VEC_BASE[w], (unsigned long long)got[0]);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ---- warm-up batch at base 0: fingerprint, in-batch vectors, dump
|
||||
double w0 = wallMs();
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, 0u, mask, nonces, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
double warmMs = wallMs() - w0;
|
||||
std::vector<uint64_t> hOut(nonces);
|
||||
CUDA_CHECK(cudaMemcpy(hOut.data(), dOut, (size_t)nonces * 8u, cudaMemcpyDeviceToHost));
|
||||
uint64_t fp = fnv1a64(hOut.data(), (size_t)nonces * 8u);
|
||||
std::printf("warm-up batch: 2^%d hashes in %.2f ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): %016llx%s\n", o.batchLog2, warmMs,
|
||||
(unsigned long long)fp, o.sd1 ? " (sd1, new value)" : (fp == 0x7c28cfb06c5c65a9ull ? " = 7c28cfb06c5c65a9 (the pack's)" : " DIFFERS from 7c28cfb06c5c65a9"));
|
||||
bool fpPass = o.sd1 || fp == 0x7c28cfb06c5c65a9ull;
|
||||
if (!o.sd1) {
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
if ((uint64_t)IGNEUM_VEC_BASE[w] + 32ull > nonces) continue;
|
||||
int bad = 0;
|
||||
for (int l = 0; l < 32; ++l) if (hOut[IGNEUM_VEC_BASE[w] + l] != IGNEUM_VEC_OUT[w][l]) ++bad;
|
||||
vecPass = vecPass && bad == 0;
|
||||
std::printf("vector warp base %u in batch: %s\n", IGNEUM_VEC_BASE[w], bad == 0 ? "PASS" : "FAIL");
|
||||
}
|
||||
}
|
||||
if (!o.dump.empty() && o.dumpWarps > 0) {
|
||||
FILE* f = std::fopen(o.dump.c_str(), "w");
|
||||
if (!f) { std::printf("FAIL: cannot write %s\n", o.dump.c_str()); return 2; }
|
||||
std::fprintf(f, "mode %s\nwarps %d\n", modeName, o.dumpWarps);
|
||||
for (int w = 0; w < o.dumpWarps; ++w) {
|
||||
std::fprintf(f, "base %u\n", 32u * (uint32_t)w);
|
||||
for (int l = 0; l < 32; ++l) std::fprintf(f, "%016llx\n", (unsigned long long)hOut[32u * (uint32_t)w + l]);
|
||||
}
|
||||
std::fclose(f);
|
||||
std::printf("dump: %d warps (bases 0, 32, ...) written to %s\n", o.dumpWarps, o.dump.c_str());
|
||||
}
|
||||
|
||||
// ---- timed batches
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
for (int b = 1; b <= o.batches; ++b) {
|
||||
uint32_t base = (uint32_t)((uint64_t)b * (uint64_t)nonces);
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, 1u));
|
||||
}
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
double gpuMs = eventMs(e0, e1);
|
||||
double mhs = (double)nonces * (double)o.batches / (gpuMs / 1000.0) / 1e6;
|
||||
std::printf("timed: %d batches x 2^%d hashes, GPU %.2f ms -> %.3f MH/s (%.2f GB/s useful, loads x 4 B)\n", o.batches, o.batchLog2, gpuMs, mhs, mhs * 1e6 * IGNEUM_LOADS_PER_HASH * 4.0 / 1e9);
|
||||
|
||||
// ---- power window: hash back to back for powerSeconds while nvidia-smi samples at 1 Hz
|
||||
double powerW = 0, smMHz = 0, memMHz = 0, mhsWindow = 0;
|
||||
int nSamples = 0, nUsed = 0;
|
||||
if (o.powerSeconds > 0) {
|
||||
Sampler smp;
|
||||
smp.start(o.device, o.powerSeconds);
|
||||
double start = wallMs();
|
||||
int b = o.batches + 1;
|
||||
int done = 0;
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
while (wallMs() - start < o.powerSeconds * 1000.0) {
|
||||
uint32_t base = (uint32_t)((uint64_t)b++ * (uint64_t)nonces);
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
++done;
|
||||
}
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
double wms = eventMs(e0, e1);
|
||||
mhsWindow = (double)nonces * (double)done / (wms / 1000.0) / 1e6;
|
||||
smp.join();
|
||||
std::lock_guard<std::mutex> g(smp.mu);
|
||||
nSamples = (int)smp.samples.size();
|
||||
for (const Sample& s : smp.samples) {
|
||||
if (s.t < 10000.0 || s.t > (double)o.powerSeconds * 1000.0) continue; // mean after 10 s, inside the load window
|
||||
powerW += s.watts; smMHz += s.smMHz; memMHz += s.memMHz; ++nUsed;
|
||||
}
|
||||
if (nUsed) { powerW /= nUsed; smMHz /= nUsed; memMHz /= nUsed; }
|
||||
std::printf("power window: %d batches in %.1f s -> %.3f MH/s (events, incl. sync gaps); nvidia-smi %d samples, %d after 10 s: mean %.1f W, SM %.0f MHz, mem %.0f MHz\n",
|
||||
done, wms / 1000.0, mhsWindow, nSamples, nUsed, powerW, smMHz, memMHz);
|
||||
if (nUsed) std::printf(" -> %.3f MH/s per W (window rate / mean W)\n", mhsWindow / powerW);
|
||||
else std::printf(" WARNING: no nvidia-smi samples inside the window (is nvidia-smi on PATH?)\n");
|
||||
}
|
||||
|
||||
size_t freeH = 0;
|
||||
CUDA_CHECK(cudaMemGetInfo(&freeH, &total));
|
||||
std::printf("device memory while hashing: %.0f MiB used (leaves freed)\n", (double)(total - freeH) / 1048576.0);
|
||||
|
||||
bool overall = cachePass && leafPass && dsPass && vecPass && fpPass && itemsEqual == 1024 && (chunkedEqual < 0 || chunkedEqual == 1024);
|
||||
std::printf("RESULT mode=%s gpu=%s cache_fill_ms=%.2f host_cache_ms=%.1f leaves_ms=%.2f build_ms=%.2f mem_build_mib=%.0f chunk_mib=%d chunked_ms=%.2f mem_chunked_mib=%.0f chunked_equal=%d items_equal=%d/1024 fingerprint=%016llx mhs=%.3f mhs_window=%.3f watts=%.1f sm_mhz=%.0f mem_mhz=%.0f cache=%s dataset=%s vectors=%s overall=%s\n",
|
||||
modeName, prop.name, cacheFill[1], hostCacheMs, leavesMs[1], buildMs[1], usedBuildMiB, o.chunkMib, chunkedMs, usedChunkMiB, chunkedEqual, itemsEqual, (unsigned long long)fp, mhs, mhsWindow, powerW, smMHz, memMHz,
|
||||
cachePass ? "PASS" : "FAIL", dsPass ? "PASS" : "FAIL", o.sd1 ? "n/a" : (vecPass ? "PASS" : "FAIL"), overall ? "PASS" : "FAIL");
|
||||
|
||||
CUDA_CHECK(cudaFree(dOut));
|
||||
CUDA_CHECK(cudaFree(dDs));
|
||||
CUDA_CHECK(cudaFree(dCache));
|
||||
return overall ? 0 : 1;
|
||||
}
|
||||
139
proto-newpow/state-dataset/cpu_rows.c
Normal file
139
proto-newpow/state-dataset/cpu_rows.c
Normal file
|
|
@ -0,0 +1,139 @@
|
|||
/* state-dataset prototype: CPU rows for igneum-build-1 (Horizon lane 8, class sd1). Plain C, -O2.
|
||||
*
|
||||
* B.1 the verifier's extra cost per 32-lane unit (4,096 items per unit), ms per unit over 100 units, one core:
|
||||
* (i) 4,096 random 64-byte reads from a host-resident leaf array of 2 GiB and of 8 GiB (random t, independent,
|
||||
* the array written once so every page is resident)
|
||||
* (ii) 4,096 leaves derived on the fly from the state root (one ChaCha12 block each, no array)
|
||||
* (iii) 4,096 x mh_item on the host cache, naive (no interleaving); plus leaf + mh_item_sd
|
||||
* B.2 the daily snapshot build on CPU: the 1 GiB leaf array (2^24 ChaCha12 blocks) on one core and on N threads
|
||||
* (OpenMP), and the host cache fill on one core.
|
||||
*
|
||||
* Build: gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
|
||||
* Run: taskset -c 4 nice -n 19 ./cpu_rows --rows (B.1 and the one-core B.2 rows)
|
||||
* taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 (the N-thread leaf build)
|
||||
*/
|
||||
#define _POSIX_C_SOURCE 200809L
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <time.h>
|
||||
#include <omp.h>
|
||||
#define IGNEUM_NO_CUDA
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
#include "sd.h"
|
||||
|
||||
static double nowMs(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1e3 + ts.tv_nsec / 1e6; }
|
||||
static uint64_t splitmix64(uint64_t* s) {
|
||||
*s += 0x9E3779B97F4A7C15ull; uint64_t z = *s;
|
||||
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull; z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
|
||||
return z ^ (z >> 31);
|
||||
}
|
||||
|
||||
static volatile uint32_t sink;
|
||||
|
||||
/* (i): random 64-byte reads from an array of gib GiB */
|
||||
static void row_reads(int gib, int units) {
|
||||
size_t items = ((size_t)gib << 30) / 64u;
|
||||
uint32_t* arr = (uint32_t*)malloc((size_t)gib << 30);
|
||||
if (!arr) { printf(" (i) %d GiB: malloc failed\n", gib); return; }
|
||||
double t0 = nowMs();
|
||||
/* write every 64-byte line once (resident pages); the content is a cheap pattern, the read cost does not depend on it */
|
||||
for (size_t i = 0; i < items; ++i) { uint32_t* l = arr + i * 16u; l[0] = (uint32_t)i; l[15] = (uint32_t)(i >> 32) ^ 0x5d1u; }
|
||||
double fillMs = nowMs() - t0;
|
||||
uint64_t s = 0x5d1ull; uint32_t acc[16] = { 0 };
|
||||
double sum = 0, mn = 1e9, mx = 0;
|
||||
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
|
||||
for (int u = 0; u < units; ++u) {
|
||||
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)(splitmix64(&s) % items);
|
||||
double a = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { const uint32_t* l = arr + (size_t)ts[k] * 16u; for (int i = 0; i < 16; ++i) acc[i] ^= l[i]; }
|
||||
double b = nowMs();
|
||||
sum += b - a; if (b - a < mn) mn = b - a; if (b - a > mx) mx = b - a;
|
||||
}
|
||||
for (int i = 0; i < 16; ++i) sink ^= acc[i];
|
||||
printf(" (i) 4,096 random 64-byte reads from a %d GiB resident array: %.3f ms per unit (min %.3f, max %.3f; %.0f ns per read; array write pass %.0f ms)\n",
|
||||
gib, sum / units, mn, mx, sum / units * 1e6 / 4096.0, fillMs);
|
||||
free(ts); free(arr);
|
||||
}
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
int rows = 0, leavesThreads = 0, units = 100;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
if (!strcmp(argv[i], "--rows")) rows = 1;
|
||||
else if (!strcmp(argv[i], "--leaves") && i + 1 < argc) leavesThreads = atoi(argv[++i]);
|
||||
else if (!strcmp(argv[i], "--units") && i + 1 < argc) units = atoi(argv[++i]);
|
||||
else { printf("usage: cpu_rows [--rows] [--leaves N] [--units 100]\n"); return 2; }
|
||||
}
|
||||
const uint32_t words = 1u << IGNEUM_DATASET_LOG2, nItems = words / 16u;
|
||||
const uint32_t cacheWords = 1u << IGNEUM_CACHE_LOG2_WORDS;
|
||||
|
||||
if (rows) {
|
||||
printf("B.1 per-unit rows (4,096 per unit, %d units, one thread, ms per unit)\n", units);
|
||||
row_reads(2, units);
|
||||
row_reads(8, units);
|
||||
/* (ii) */
|
||||
{
|
||||
uint64_t s = 0x5d1ull; double sum = 0;
|
||||
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
|
||||
for (int u = 0; u < units; ++u) {
|
||||
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)splitmix64(&s) & (nItems - 1u);
|
||||
double a = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16]; mh_leaf(ts[k], leaf); sink ^= leaf[0] ^ leaf[15]; }
|
||||
sum += nowMs() - a;
|
||||
}
|
||||
printf(" (ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): %.3f ms per unit (%.0f ns per leaf)\n", sum / units, sum / units * 1e6 / 4096.0);
|
||||
free(ts);
|
||||
}
|
||||
/* B.2 host cache fill, one core; then (iii) */
|
||||
uint32_t* cache = (uint32_t*)malloc((size_t)cacheWords * 4u);
|
||||
double c0 = nowMs();
|
||||
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(cache, seg);
|
||||
double cacheMs = nowMs() - c0;
|
||||
printf("B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: %.3f s\n", cacheMs / 1000.0);
|
||||
{
|
||||
uint64_t s = 0x5d1ull; double sumItem = 0, sumSd = 0, mn = 1e9, mx = 0;
|
||||
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
|
||||
for (int u = 0; u < units; ++u) {
|
||||
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)splitmix64(&s) & (nItems - 1u);
|
||||
double a = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { uint32_t it[16]; mh_item(cache, ts[k], it); sink ^= it[0]; }
|
||||
double b = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16], it[16]; mh_leaf(ts[k], leaf); mh_item_sd(cache, leaf, ts[k], it); sink ^= it[0]; }
|
||||
double c = nowMs();
|
||||
sumItem += b - a; sumSd += c - b; if (b - a < mn) mn = b - a; if (b - a > mx) mx = b - a;
|
||||
}
|
||||
printf(" (iii) 4,096 x mh_item on the host cache, naive (no interleaving): %.2f ms per unit (min %.2f, max %.2f)\n", sumItem / units, mn, mx);
|
||||
printf(" 4,096 x (leaf + mh_item_sd), naive: %.2f ms per unit\n", sumSd / units);
|
||||
printf(" reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.\n");
|
||||
free(ts);
|
||||
}
|
||||
free(cache);
|
||||
/* B.2 leaf array on one core */
|
||||
{
|
||||
uint32_t* leaves = (uint32_t*)malloc((size_t)nItems * 64u);
|
||||
double a = nowMs();
|
||||
for (uint32_t t = 0; t < nItems; ++t) mh_leaf(t, leaves + (size_t)t * 16u);
|
||||
double ms = nowMs() - a;
|
||||
sink ^= leaves[0] ^ leaves[(size_t)nItems * 16u - 1u];
|
||||
printf("B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: %.3f s (%.0f ns per leaf)\n", ms / 1000.0, ms * 1e6 / nItems);
|
||||
free(leaves);
|
||||
}
|
||||
}
|
||||
if (leavesThreads > 0) {
|
||||
omp_set_num_threads(leavesThreads);
|
||||
uint32_t* leaves = (uint32_t*)malloc((size_t)nItems * 64u);
|
||||
/* first touch in parallel too, then time a second full build so page faults are not in the number */
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
double a = nowMs();
|
||||
#pragma omp parallel for schedule(static)
|
||||
for (int64_t t = 0; t < (int64_t)nItems; ++t) mh_leaf((uint32_t)t, leaves + (size_t)t * 16u);
|
||||
double ms = nowMs() - a;
|
||||
printf("B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), %d threads (OpenMP, %d actual): %.3f s%s\n", leavesThreads, omp_get_max_threads(), ms / 1000.0, pass == 0 ? " (first pass, includes page faults)" : "");
|
||||
}
|
||||
sink ^= leaves[0];
|
||||
free(leaves);
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
164
proto-newpow/state-dataset/kernel.cu
Normal file
164
proto-newpow/state-dataset/kernel.cu
Normal file
|
|
@ -0,0 +1,164 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
213
proto-newpow/state-dataset/kernel_sd.cu
Normal file
213
proto-newpow/state-dataset/kernel_sd.cu
Normal file
|
|
@ -0,0 +1,213 @@
|
|||
// state-dataset prototype: the mx8-genesis pack kernel plus igneum_leaves and igneum_build_sd (appended at the end).
|
||||
// The hash kernel igneum_hash and every other line of the pack are byte for byte the pack's kernel.cu.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
#include "sd.h" // state-dataset prototype: mh_leaf, mh_item_sd (lane 8, class sd1)
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// state-dataset prototype (class sd1). One thread per item.
|
||||
// igneum_leaves: leaves[t * 16 .. t * 16 + 15] = mh_leaf(t) (the synthetic 64-byte state leaf, sd.h).
|
||||
// igneum_build_sd: item t = mh_item_sd(cache, leaves + 16 t, t). The leaf read is one coalesced 64-byte read per
|
||||
// thread (consecutive threads read consecutive leaves), so streaming the leaf array in chunks is a matter of
|
||||
// launching over an item range [t0, t0 + n) with the chunk's leaves at leaves - 16 t0: nothing in the kernel
|
||||
// depends on the whole array being resident.
|
||||
__global__ void igneum_leaves(uint32_t* leaves, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t leaf[16];
|
||||
mh_leaf(t, leaf);
|
||||
uint32_t* d = leaves + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = leaf[i];
|
||||
}
|
||||
}
|
||||
__global__ void igneum_build_sd(uint32_t* ds, const uint32_t* cache, const uint32_t* leaves, uint32_t t0, uint32_t nItems) {
|
||||
uint32_t t = t0 + blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < t0 + nItems) {
|
||||
uint32_t leaf[16];
|
||||
const uint32_t* l = leaves + (size_t)(t - t0) * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) leaf[i] = l[i];
|
||||
uint32_t s[16];
|
||||
mh_item_sd(cache, leaf, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_leaves(uint32_t* leaves, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_leaves<<<grid, block>>>(leaves, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
// Builds items [t0, t0 + nItems) from the leaves of that range (leaves points at leaf t0).
|
||||
cudaError_t igneum_launch_build_sd(uint32_t* ds, const uint32_t* cache, const uint32_t* leaves, uint32_t t0, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build_sd<<<grid, block>>>(ds, cache, leaves, t0, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
109
proto-newpow/state-dataset/memhard.h
Normal file
109
proto-newpow/state-dataset/memhard.h
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
|
||||
// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0x3067619fu ^ prev[4];
|
||||
x[5] = 0x3c269176u ^ prev[5];
|
||||
x[6] = 0x84a03b03u ^ prev[6];
|
||||
x[7] = 0xf8c63294u ^ prev[7];
|
||||
x[8] = 0xff977c5bu ^ prev[8];
|
||||
x[9] = 0xe60def3eu ^ prev[9];
|
||||
x[10] = 0x63630141u ^ prev[10];
|
||||
x[11] = 0xb8fbcb58u ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
|
||||
s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
|
||||
s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
|
||||
s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
|
||||
s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
|
||||
s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
|
||||
s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
|
||||
s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
|
||||
s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
|
||||
s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
|
||||
s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
|
||||
s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
|
||||
s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
|
||||
s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
|
||||
s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
|
||||
s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0x3067619fu;
|
||||
s[1] = 0x3c269176u;
|
||||
s[2] = 0x84a03b03u;
|
||||
s[3] = 0xf8c63294u;
|
||||
s[4] = 0xff977c5bu;
|
||||
s[5] = 0xe60def3eu;
|
||||
s[6] = 0x63630141u;
|
||||
s[7] = 0xb8fbcb58u;
|
||||
s[8] = t * 0x42146205u + 0xbab68293u;
|
||||
s[9] = t * 0x52cbe0fbu + 0xcc162340u;
|
||||
s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
|
||||
s[11] = t * 0x6728907fu + 0xe62b8997u;
|
||||
s[12] = t * 0xd81d9751u + 0xc9c80297u;
|
||||
s[13] = t * 0x132952c3u + 0xf74a1654u;
|
||||
s[14] = t * 0xf60de277u + 0x3d704af5u;
|
||||
s[15] = t * 0x05358035u + 0x3cf522b7u;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
66
proto-newpow/state-dataset/program.h
Normal file
66
proto-newpow/state-dataset/program.h
Normal file
|
|
@ -0,0 +1,66 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-genesis"
|
||||
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0xe323b9dcaf283a6full
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "mx8"
|
||||
#define IGNEUM_CLASS_MIXER_MULT 8
|
||||
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
|
||||
#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIXER_MULT 8 // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)
|
||||
#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
2
proto-newpow/state-dataset/results/cpu/leaves32-2.log
Normal file
2
proto-newpow/state-dataset/results/cpu/leaves32-2.log
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.179 s (first pass, includes page faults)
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.083 s
|
||||
2
proto-newpow/state-dataset/results/cpu/leaves32.log
Normal file
2
proto-newpow/state-dataset/results/cpu/leaves32.log
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.172 s (first pass, includes page faults)
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.073 s
|
||||
9
proto-newpow/state-dataset/results/cpu/rows.log
Normal file
9
proto-newpow/state-dataset/results/cpu/rows.log
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
|
||||
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.108 ms per unit (min 0.092, max 0.172; 26 ns per read; array write pass 1022 ms)
|
||||
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.139 ms per unit (min 0.127, max 0.183; 34 ns per read; array write pass 3783 ms)
|
||||
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.284 ms per unit (69 ns per leaf)
|
||||
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.447 s
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms per unit (min 8.91, max 9.60)
|
||||
4,096 x (leaf + mh_item_sd), naive: 9.31 ms per unit
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.670 s (100 ns per leaf)
|
||||
9
proto-newpow/state-dataset/results/cpu/rows2.log
Normal file
9
proto-newpow/state-dataset/results/cpu/rows2.log
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
|
||||
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.163 ms per unit (min 0.144, max 0.201; 40 ns per read; array write pass 1252 ms)
|
||||
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.209 ms per unit (min 0.203, max 0.229; 51 ns per read; array write pass 4325 ms)
|
||||
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.285 ms per unit (70 ns per leaf)
|
||||
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.452 s
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 11.24 ms per unit (min 9.38, max 13.51)
|
||||
4,096 x (leaf + mh_item_sd), naive: 11.64 ms per unit
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.677 s (100 ns per leaf)
|
||||
27
proto-newpow/state-dataset/results/cpu/run_cpu.log
Normal file
27
proto-newpow/state-dataset/results/cpu/run_cpu.log
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
== Tue Oct 6 07:44:26 PM UTC 2026 on igneum-build-1, load 18.98 28.50 18.66
|
||||
Model name: AMD EPYC 9454P 48-Core Processor
|
||||
gcc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0
|
||||
== build
|
||||
== B.1 and one-core B.2 rows, core 4, nice 19
|
||||
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
|
||||
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.108 ms per unit (min 0.092, max 0.172; 26 ns per read; array write pass 1022 ms)
|
||||
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.139 ms per unit (min 0.127, max 0.183; 34 ns per read; array write pass 3783 ms)
|
||||
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.284 ms per unit (69 ns per leaf)
|
||||
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.447 s
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms per unit (min 8.91, max 9.60)
|
||||
4,096 x (leaf + mh_item_sd), naive: 9.31 ms per unit
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.670 s (100 ns per leaf)
|
||||
== B.2 leaf array on 32 threads (cores 4-35), nice 19
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.172 s (first pass, includes page faults)
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.073 s
|
||||
== verify_sd per-unit rows on this CPU (core 4), for comparison
|
||||
verify_sd mode sd1
|
||||
host cache fill: 480.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
||||
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.285 ms
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms (min 8.92, max 9.44)
|
||||
4,096 x (leaf + mh_item_sd), naive: 9.27 ms
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
RESULT verify mode=sd1 host_cache_ms=480.4 items_equal=-1 lanes_equal=-1/0
|
||||
== done Tue Oct 6 07:44:39 PM UTC 2026, load 16.75 27.55 18.50
|
||||
8
proto-newpow/state-dataset/results/cpu/verify_rows.log
Normal file
8
proto-newpow/state-dataset/results/cpu/verify_rows.log
Normal file
|
|
@ -0,0 +1,8 @@
|
|||
verify_sd mode sd1
|
||||
host cache fill: 480.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
||||
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.285 ms
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms (min 8.92, max 9.44)
|
||||
4,096 x (leaf + mh_item_sd), naive: 9.27 ms
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
RESULT verify mode=sd1 host_cache_ms=480.4 items_equal=-1 lanes_equal=-1/0
|
||||
24
proto-newpow/state-dataset/results/gpu/control.log
Normal file
24
proto-newpow/state-dataset/results/gpu/control.log
Normal file
|
|
@ -0,0 +1,24 @@
|
|||
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.86 ms first, 1.81 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 380.5 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.1 M items/s
|
||||
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
|
||||
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
|
||||
dataset self-test: PASS (64 random words vs host mh_word: PASS)
|
||||
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
|
||||
vector warp base 0: PASS (0 of 32 lanes differ)
|
||||
vector warp base 4096: PASS (0 of 32 lanes differ)
|
||||
vector warp base 1000000: PASS (0 of 32 lanes differ)
|
||||
warm-up batch: 2^24 hashes in 266.02 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
|
||||
vector warp base 0 in batch: PASS
|
||||
vector warp base 4096 in batch: PASS
|
||||
vector warp base 1000000 in batch: PASS
|
||||
dump: 4 warps (bases 0, 32, ...) written to dump_control.txt
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.56 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.078 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.6 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.304 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.81 host_cache_ms=380.5 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.078 watts=207.6 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS
|
||||
23
proto-newpow/state-dataset/results/gpu/control2.log
Normal file
23
proto-newpow/state-dataset/results/gpu/control2.log
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.86 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 380.0 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.2 M items/s
|
||||
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
|
||||
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
|
||||
dataset self-test: PASS (64 random words vs host mh_word: PASS)
|
||||
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
|
||||
vector warp base 0: PASS (0 of 32 lanes differ)
|
||||
vector warp base 4096: PASS (0 of 32 lanes differ)
|
||||
vector warp base 1000000: PASS (0 of 32 lanes differ)
|
||||
warm-up batch: 2^24 hashes in 265.99 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
|
||||
vector warp base 0 in batch: PASS
|
||||
vector warp base 4096 in batch: PASS
|
||||
vector warp base 1000000 in batch: PASS
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.55 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.079 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 205.2 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.307 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=380.0 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.079 watts=205.2 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS
|
||||
134
proto-newpow/state-dataset/results/gpu/dump_control.txt
Normal file
134
proto-newpow/state-dataset/results/gpu/dump_control.txt
Normal file
|
|
@ -0,0 +1,134 @@
|
|||
mode control
|
||||
warps 4
|
||||
base 0
|
||||
19b56348bc85304d
|
||||
b08a9cfb44aa720f
|
||||
e1f8f627f780eff7
|
||||
2ff3e86ff1696161
|
||||
b65e0578c257e8ac
|
||||
f37b5705c2bebaae
|
||||
b19fef670982389e
|
||||
2374331b28827a11
|
||||
f492cdd05dda7f88
|
||||
37700f1385e19885
|
||||
a27524f3010b3e87
|
||||
6c8a24de8b938c43
|
||||
3dc017c8820cfd16
|
||||
5c80d37146657b68
|
||||
214082cac03f734a
|
||||
1d67c665145d72f3
|
||||
9fb0ae736eb87ff5
|
||||
cd7d416ae3a6be61
|
||||
0069cda85c6de58f
|
||||
0b7d82e717224baa
|
||||
600960f6d7be30f0
|
||||
eb29032663b0c4d2
|
||||
f29583a5766408f1
|
||||
6c9532a99ad7314d
|
||||
9a949fd0959cc80f
|
||||
a20258520e7f6c25
|
||||
f2ab11f9bb032e38
|
||||
cc967bcd0c8d07c1
|
||||
37745267bb3231f2
|
||||
35a046048c2b69b3
|
||||
aa51834cd3f364f3
|
||||
359192708e4f754a
|
||||
base 32
|
||||
bbea9a410fe7d5b7
|
||||
11bf76519a1414d4
|
||||
59c52e47a752a4fa
|
||||
3964f1251e8b90e8
|
||||
75230c9369b790af
|
||||
8dec6e557dc03a87
|
||||
0b819148666b3f13
|
||||
7707a6b8d499e40f
|
||||
445de8aa5b9699d5
|
||||
6becbf9e66d7e9de
|
||||
7a46928b6ed6861f
|
||||
0cbfca3b2c769872
|
||||
04e97a5b926b55d5
|
||||
2a7ba12634ae81bd
|
||||
96cb4eaf4e25600a
|
||||
2826bd0a4355a91d
|
||||
d7ba781350c80652
|
||||
e87c0820f6136930
|
||||
e8864f498a00a8e5
|
||||
c1b445f3b5748847
|
||||
72def57f2e515d11
|
||||
f4708f2d80c98b8d
|
||||
da99416aa6245b67
|
||||
90f6e6d21f11fc66
|
||||
a9e3c5f36c715c31
|
||||
e5e2a586d61e4869
|
||||
98e1ce97884663da
|
||||
36e8af738edddc85
|
||||
8530526bbf3d85bb
|
||||
22128f72173887c6
|
||||
30b7aa6ab94449e2
|
||||
2ea3ca54249bd934
|
||||
base 64
|
||||
68b73c212b47e6f2
|
||||
55acf1fa7bbac2dd
|
||||
45001afaa68359f7
|
||||
3271315def7632fb
|
||||
8af0fe0f8b16148b
|
||||
2e0c7f7bac0276e1
|
||||
fe178e4364b66e31
|
||||
f39df092fabe08ad
|
||||
e6d7ad0a5d97a7eb
|
||||
0013aa74d3b41bba
|
||||
f57f5241a626b717
|
||||
747067ae7e06c706
|
||||
50325d1d39c18b09
|
||||
1fb5aa02196bd962
|
||||
78eaa19e7dd002f4
|
||||
8fe0600a8ff1c365
|
||||
021d685d64796246
|
||||
281446f54c93f2a5
|
||||
ba2cc4bef235b006
|
||||
bfb4bc312bd962a4
|
||||
cecc08bdb679ca9c
|
||||
1f079cb2d03b55a1
|
||||
a8cc9f0647e103a0
|
||||
2615618b9cb4877f
|
||||
853f0dd6f03607f2
|
||||
bf54084008cd8ca3
|
||||
21986116f4e97e6c
|
||||
42dbca5da95d1b1c
|
||||
d7669ed19dea6986
|
||||
17e407aa53422e43
|
||||
8c7d4f5bfde2f58e
|
||||
13adadad05f4540a
|
||||
base 96
|
||||
3a0fd3ada6b5a797
|
||||
aa5a637cb9caa3b7
|
||||
728b73955e0ad47d
|
||||
5092aa36c478581c
|
||||
a463e4220676d004
|
||||
ce92dfc65d3af26e
|
||||
2943fd8967c7f215
|
||||
71acd3ccc8054eaa
|
||||
2b93c6cd8c22c051
|
||||
00f10e1b004bab4d
|
||||
02e2733b82de9bce
|
||||
1c9d5cd0ed5350ce
|
||||
dcd13404e4ab0b21
|
||||
d4afc0e3a2814f63
|
||||
f2b37d4b0322ff2f
|
||||
0212964cea5688de
|
||||
546351f9a3ec9957
|
||||
d1fe39207bff2a50
|
||||
a9a38d54a9f85090
|
||||
e039372a0cc1aa9e
|
||||
008a4289a796e05e
|
||||
a9a1dca5b9de1fbc
|
||||
751a1769f1203a6d
|
||||
5b66d9c952febc43
|
||||
47ac39abda0dd053
|
||||
c7b86597a7ef8d83
|
||||
d87824ea2fd452c4
|
||||
82a39552b84b113c
|
||||
562aa1cb671a8056
|
||||
4dc55d7e9e0829dd
|
||||
bb0cf26fbb506ddc
|
||||
353c62d9a3b6aaed
|
||||
134
proto-newpow/state-dataset/results/gpu/dump_sd1.txt
Normal file
134
proto-newpow/state-dataset/results/gpu/dump_sd1.txt
Normal file
|
|
@ -0,0 +1,134 @@
|
|||
mode sd1
|
||||
warps 4
|
||||
base 0
|
||||
b600edbed969becc
|
||||
2d988317678e9099
|
||||
29609900b3a81764
|
||||
2506a510ec612b1c
|
||||
cdc516143acae7eb
|
||||
dd2901cc2335449d
|
||||
41614b8fbf48a689
|
||||
317f5df1eb7c0cbe
|
||||
2da7edd46d732703
|
||||
b89b6f78568676a9
|
||||
e93e96d589f70d8b
|
||||
197bd4558fe42bf7
|
||||
911c2ed1bca014bc
|
||||
448984ee31e0f576
|
||||
6a9af9585e6ce6ba
|
||||
446ba1d41100da43
|
||||
d206864c5393aef9
|
||||
46b93bf7a7b89196
|
||||
85ce18e332a13c31
|
||||
34f1d32a0f153659
|
||||
d2b8703f431f1237
|
||||
3bda4edaa92faa6f
|
||||
d8e5b9f18abf75fd
|
||||
234a323a6d619bfd
|
||||
9558b3c4e42cd2f7
|
||||
016372acc63f6372
|
||||
7ebac0c7e5fc87f7
|
||||
81f40d3ef4b5508b
|
||||
aa6ec7854aacde98
|
||||
692a050230de2fe8
|
||||
18b7df938406f9a1
|
||||
2412df7ada5a3202
|
||||
base 32
|
||||
e451566071a2a0ba
|
||||
d63578a9312e5b32
|
||||
6bdaecc6d5f7158b
|
||||
694c691e7dc1de0a
|
||||
d98a0709679797b3
|
||||
17a7c68f4ff5e95f
|
||||
58222d4156171349
|
||||
af4bcd67686aad89
|
||||
fdad4d96931b9f2a
|
||||
5e16c79bd1f2c688
|
||||
7437b2913e288981
|
||||
7b4485a218901ccf
|
||||
540694fa54e15619
|
||||
e42600d66ac6e0f5
|
||||
3b2b6322a93124d2
|
||||
b474e79e977515ae
|
||||
85c356e1a7f560fe
|
||||
47961522447a02fc
|
||||
5cf31818bea49c4a
|
||||
a517ba58dbed530c
|
||||
d1b37bf54bd68236
|
||||
27905e28697a6ef7
|
||||
42a743ba99c22bf9
|
||||
5ec48e4245545d4a
|
||||
38ecfefd174944d8
|
||||
3dd1891a53a9b415
|
||||
44120379678fcd2a
|
||||
e2b43434e9e5b97f
|
||||
db448577e2229d66
|
||||
41eb61946c60e9a7
|
||||
53db71d828b8d7e6
|
||||
ea19aabc66d1b11a
|
||||
base 64
|
||||
34548f5056f27d78
|
||||
e7fea9a49cf2556e
|
||||
ebdaf245ef41ec6d
|
||||
3c77731e46a78bd3
|
||||
e985324128b13372
|
||||
c13540b76c4dfc4a
|
||||
9050d43f68bf20e9
|
||||
4aa7762e9a08d90a
|
||||
d106b0610790817b
|
||||
5c9c91398c9b5e7c
|
||||
56b71969e73c466e
|
||||
2d61dea47bf479bb
|
||||
2ddb7cc8f81fd24c
|
||||
f422f7adbb7a230e
|
||||
794858ac6a6bada1
|
||||
c7a37f1c3cebeea4
|
||||
241f6632d74d699e
|
||||
f8e536083c72f300
|
||||
3234e942615b9891
|
||||
573677792f6bfe53
|
||||
6838b021f005dee5
|
||||
7e1042640bdf5609
|
||||
4a0399c7b6ca97fd
|
||||
37ad996ce5fb721a
|
||||
f633ed2dc43bc76d
|
||||
e569e4a049054982
|
||||
a9c5fa29d0ab9603
|
||||
ca0dcf5025e589f4
|
||||
17687fa3872dd204
|
||||
965e7bea52b291a7
|
||||
d3aec969f424c835
|
||||
44453d33abd93ef6
|
||||
base 96
|
||||
c13d6ea3e05c5e46
|
||||
1c291747d4676990
|
||||
349040c82d1e4ad8
|
||||
f05a14026be82f03
|
||||
5ee5f52611690c2b
|
||||
60c99fb5c988e677
|
||||
3d0899be8af84174
|
||||
74dad192f88cca2c
|
||||
005480c1f5e84d6f
|
||||
a38fa3da9db77186
|
||||
9495da0d026c51fc
|
||||
54f4ed7062e92211
|
||||
6f910c4ad2774774
|
||||
55752a05b9fa5b71
|
||||
883b5367142927bf
|
||||
96449de7c472e375
|
||||
f8722d3292d2948b
|
||||
21d0485f77308767
|
||||
010a6d05ede7793e
|
||||
56cbfe0810728799
|
||||
e136e570c4ed04a9
|
||||
848c020abe56db44
|
||||
71aa3cddb1bc9408
|
||||
58757e0f19ddb3ef
|
||||
f2e4273de46c5968
|
||||
8664ce7ab76eceac
|
||||
881db7223c9a750b
|
||||
788a1553b16895ea
|
||||
c84a69728aab9555
|
||||
9019c8d87a1da486
|
||||
4ab485aa2deb7ca5
|
||||
b2795a944cca750d
|
||||
1026
proto-newpow/state-dataset/results/gpu/items_control.txt
Normal file
1026
proto-newpow/state-dataset/results/gpu/items_control.txt
Normal file
File diff suppressed because it is too large
Load diff
1026
proto-newpow/state-dataset/results/gpu/items_sd1.txt
Normal file
1026
proto-newpow/state-dataset/results/gpu/items_sd1.txt
Normal file
File diff suppressed because it is too large
Load diff
85
proto-newpow/state-dataset/results/gpu/run_gpu.log
Normal file
85
proto-newpow/state-dataset/results/gpu/run_gpu.log
Normal file
|
|
@ -0,0 +1,85 @@
|
|||
== Tue Oct 6 19:43:12 UTC 2026 on a4cac49bb840
|
||||
NVIDIA GeForce RTX 4090, 595.91.07, 3135 MHz, 24564 MiB
|
||||
Build cuda_12.8.r12.8/compiler.35583870_0
|
||||
== build
|
||||
== control
|
||||
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.86 ms first, 1.81 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 380.5 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.1 M items/s
|
||||
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
|
||||
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
|
||||
dataset self-test: PASS (64 random words vs host mh_word: PASS)
|
||||
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
|
||||
vector warp base 0: PASS (0 of 32 lanes differ)
|
||||
vector warp base 4096: PASS (0 of 32 lanes differ)
|
||||
vector warp base 1000000: PASS (0 of 32 lanes differ)
|
||||
warm-up batch: 2^24 hashes in 266.02 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
|
||||
vector warp base 0 in batch: PASS
|
||||
vector warp base 4096 in batch: PASS
|
||||
vector warp base 1000000 in batch: PASS
|
||||
dump: 4 warps (bases 0, 32, ...) written to dump_control.txt
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.56 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.078 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.6 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.304 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.81 host_cache_ms=380.5 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.078 watts=207.6 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS
|
||||
== sd1
|
||||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 383.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.56 ms first, 5.49 ms second -> 195.4 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
dump: 4 warps (bases 0, 32, ...) written to dump_sd1.txt
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.35 ms -> 63.088 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.9 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.303 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=383.9 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.088 mhs_window=63.083 watts=207.9 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||
== verify (plain C, CPU core 2)
|
||||
verify_sd mode control
|
||||
host cache fill: 587.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
interpreter vs Mac vector base 0: 32 of 32 lanes equal PASS
|
||||
item bit-exactness (plain C, mh_item on the host cache): 1024 of 1024 items equal PASS
|
||||
warp base 0: 32 of 32 lanes equal (lane 0 gpu 19b56348bc85304d interp 19b56348bc85304d)
|
||||
warp base 32: 32 of 32 lanes equal (lane 0 gpu bbea9a410fe7d5b7 interp bbea9a410fe7d5b7)
|
||||
warp base 64: 32 of 32 lanes equal (lane 0 gpu 68b73c212b47e6f2 interp 68b73c212b47e6f2)
|
||||
warp base 96: 32 of 32 lanes equal (lane 0 gpu 3a0fd3ada6b5a797 interp 3a0fd3ada6b5a797)
|
||||
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (37 ms interpreting, 4096 derivations per warp)
|
||||
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
||||
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.369 ms
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 10.14 ms (min 9.98, max 10.31)
|
||||
4,096 x (leaf + mh_item_sd), naive: 10.56 ms
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
RESULT verify mode=control host_cache_ms=587.4 items_equal=1024 lanes_equal=128/128
|
||||
verify_sd mode sd1
|
||||
host cache fill: 572.9 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
item bit-exactness (plain C, mh_item_sd with mh_leaf on the host cache): 1024 of 1024 items equal PASS
|
||||
warp base 0: 32 of 32 lanes equal (lane 0 gpu b600edbed969becc interp b600edbed969becc)
|
||||
warp base 32: 32 of 32 lanes equal (lane 0 gpu e451566071a2a0ba interp e451566071a2a0ba)
|
||||
warp base 64: 32 of 32 lanes equal (lane 0 gpu 34548f5056f27d78 interp 34548f5056f27d78)
|
||||
warp base 96: 32 of 32 lanes equal (lane 0 gpu c13d6ea3e05c5e46 interp c13d6ea3e05c5e46)
|
||||
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (42 ms interpreting, 4096 derivations per warp)
|
||||
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
||||
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.374 ms
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 10.21 ms (min 10.09, max 12.24)
|
||||
4,096 x (leaf + mh_item_sd), naive: 10.58 ms
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
RESULT verify mode=sd1 host_cache_ms=572.9 items_equal=1024 lanes_equal=128/128
|
||||
== done Tue Oct 6 19:44:17 UTC 2026
|
||||
23
proto-newpow/state-dataset/results/gpu/sd1-chunk256.log
Normal file
23
proto-newpow/state-dataset/results/gpu/sd1-chunk256.log
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.73 ms first, 1.70 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 380.7 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.16 ms first, 5.10 ms second -> 210.7 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.00 ms first, 31.92 ms second -> 525.6 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
chunked build (leaves from pinned host memory in 256 MiB chunks, 4 chunks, one stream): 75.47 ms total (copies + builds); whole-array H2D alone 61.92 ms = 17.3 GB/s, D2H 54.61 ms
|
||||
device memory during the chunked build: 1933 MiB used (context + cache 256 + dataset 1024 + chunk 256 MiB); items after the chunked rebuild: 1024 of 1024 equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 265.97 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.40 ms -> 63.086 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.70 host_cache_ms=380.7 leaves_ms=5.10 build_ms=31.92 mem_build_mib=2699 chunk_mib=256 chunked_ms=75.47 mem_chunked_mib=1933 chunked_equal=1024 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.086 mhs_window=0.000 watts=0.0 sm_mhz=0 mem_mhz=0 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||
23
proto-newpow/state-dataset/results/gpu/sd1-chunk64.log
Normal file
23
proto-newpow/state-dataset/results/gpu/sd1-chunk64.log
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.85 ms first, 1.82 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 382.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.53 ms first, 5.46 ms second -> 196.5 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
chunked build (leaves from pinned host memory in 64 MiB chunks, 16 chunks, one stream): 75.80 ms total (copies + builds); whole-array H2D alone 62.41 ms = 17.2 GB/s, D2H 54.85 ms
|
||||
device memory during the chunked build: 1741 MiB used (context + cache 256 + dataset 1024 + chunk 64 MiB); items after the chunked rebuild: 1024 of 1024 equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 265.98 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.43 ms -> 63.086 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.82 host_cache_ms=382.9 leaves_ms=5.46 build_ms=31.98 mem_build_mib=2699 chunk_mib=64 chunked_ms=75.80 mem_chunked_mib=1741 chunked_equal=1024 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.086 mhs_window=0.000 watts=0.0 sm_mhz=0 mem_mhz=0 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||
24
proto-newpow/state-dataset/results/gpu/sd1.log
Normal file
24
proto-newpow/state-dataset/results/gpu/sd1.log
Normal file
|
|
@ -0,0 +1,24 @@
|
|||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 383.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.56 ms first, 5.49 ms second -> 195.4 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
dump: 4 warps (bases 0, 32, ...) written to dump_sd1.txt
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.35 ms -> 63.088 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.9 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.303 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=383.9 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.088 mhs_window=63.083 watts=207.9 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||
23
proto-newpow/state-dataset/results/gpu/sd12.log
Normal file
23
proto-newpow/state-dataset/results/gpu/sd12.log
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 382.6 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.55 ms first, 5.49 ms second -> 195.4 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.07 ms first, 31.98 ms second -> 524.6 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.36 ms -> 63.087 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 206.5 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.305 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=382.6 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.087 mhs_window=63.083 watts=206.5 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||
13
proto-newpow/state-dataset/results/gpu/verify_control.log
Normal file
13
proto-newpow/state-dataset/results/gpu/verify_control.log
Normal file
|
|
@ -0,0 +1,13 @@
|
|||
verify_sd mode control
|
||||
host cache fill: 574.3 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
interpreter vs Mac vector base 0: 32 of 32 lanes equal PASS
|
||||
item bit-exactness (plain C, mh_item on the host cache): 1024 of 1024 items equal PASS
|
||||
warp base 0: 32 of 32 lanes equal (lane 0 gpu 19b56348bc85304d interp 19b56348bc85304d)
|
||||
warp base 32: 32 of 32 lanes equal (lane 0 gpu bbea9a410fe7d5b7 interp bbea9a410fe7d5b7)
|
||||
warp base 64: 32 of 32 lanes equal (lane 0 gpu 68b73c212b47e6f2 interp 68b73c212b47e6f2)
|
||||
warp base 96: 32 of 32 lanes equal (lane 0 gpu 3a0fd3ada6b5a797 interp 3a0fd3ada6b5a797)
|
||||
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (37 ms interpreting, 4096 derivations per warp)
|
||||
lane 0 of warp base 0: 128 loads touch 128 distinct words in 128 distinct items (of 2^24 items)
|
||||
first 8 load words: 8528400 43034175 255822847 136709537 36562395 83203117 165232476 263214792
|
||||
Merkle openings for a 2 GiB state (2^25 leaves of 64 B, depth 25, 32-byte hashes): 128 openings x 25 x 32 = 102400 bytes of siblings + 128 x 64 = 8192 bytes of leaves = 110592 bytes (108.0 KiB) per lane, uncompressed; 128 distinct items need only 128 openings = 110592 bytes
|
||||
RESULT verify mode=control host_cache_ms=574.3 items_equal=1024 lanes_equal=128/128
|
||||
12
proto-newpow/state-dataset/results/gpu/verify_sd1.log
Normal file
12
proto-newpow/state-dataset/results/gpu/verify_sd1.log
Normal file
|
|
@ -0,0 +1,12 @@
|
|||
verify_sd mode sd1
|
||||
host cache fill: 565.8 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
item bit-exactness (plain C, mh_item_sd with mh_leaf on the host cache): 1024 of 1024 items equal PASS
|
||||
warp base 0: 32 of 32 lanes equal (lane 0 gpu b600edbed969becc interp b600edbed969becc)
|
||||
warp base 32: 32 of 32 lanes equal (lane 0 gpu e451566071a2a0ba interp e451566071a2a0ba)
|
||||
warp base 64: 32 of 32 lanes equal (lane 0 gpu 34548f5056f27d78 interp 34548f5056f27d78)
|
||||
warp base 96: 32 of 32 lanes equal (lane 0 gpu c13d6ea3e05c5e46 interp c13d6ea3e05c5e46)
|
||||
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (42 ms interpreting, 4096 derivations per warp)
|
||||
lane 0 of warp base 0: 128 loads touch 128 distinct words in 128 distinct items (of 2^24 items)
|
||||
first 8 load words: 8528400 85793927 255822847 180392080 172515305 211564358 34853938 259149705
|
||||
Merkle openings for a 2 GiB state (2^25 leaves of 64 B, depth 25, 32-byte hashes): 128 openings x 25 x 32 = 102400 bytes of siblings + 128 x 64 = 8192 bytes of leaves = 110592 bytes (108.0 KiB) per lane, uncompressed; 128 distinct items need only 128 openings = 110592 bytes
|
||||
RESULT verify mode=sd1 host_cache_ms=565.8 items_equal=1024 lanes_equal=128/128
|
||||
23
proto-newpow/state-dataset/run_cpu.sh
Executable file
23
proto-newpow/state-dataset/run_cpu.sh
Executable file
|
|
@ -0,0 +1,23 @@
|
|||
#!/bin/bash
|
||||
# state-dataset prototype: CPU rows on igneum-build-1 (EPYC 9454P). Runs inside /srv/builds/horizon-newpow/state-dataset.
|
||||
# From the Mac:
|
||||
# rsync -az -e "ssh -i ~/.ssh/igneum_ed25519" proto-newpow/state-dataset/ build@188.40.146.49:/srv/builds/horizon-newpow/state-dataset/
|
||||
# ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 'cd /srv/builds/horizon-newpow/state-dataset && touch * && bash run_cpu.sh'
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")"
|
||||
touch *.c *.h
|
||||
echo "== $(date -u) on $(hostname), load $(cut -d' ' -f1-3 /proc/loadavg)"
|
||||
lscpu | grep "Model name"; gcc --version | head -1
|
||||
echo "== build"
|
||||
gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
|
||||
gcc -O2 -std=c11 -o verify_sd verify_sd.c
|
||||
echo "== B.1 and one-core B.2 rows, core 4, nice 19"
|
||||
taskset -c 4 nice -n 19 ./cpu_rows --rows | tee rows.log
|
||||
echo "== B.2 leaf array on 32 threads (cores 4-35), nice 19"
|
||||
taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 | tee leaves32.log
|
||||
echo "== second reading of the B.1 rows (the box is shared; the load average is printed with each reading)"
|
||||
taskset -c 4 nice -n 19 ./cpu_rows --rows | tee rows2.log
|
||||
taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 | tee leaves32-2.log
|
||||
echo "== verify_sd per-unit rows on this CPU (core 4), for comparison"
|
||||
taskset -c 4 nice -n 19 ./verify_sd --mode sd1 | tee verify_rows.log
|
||||
echo "== done $(date -u), load $(cut -d' ' -f1-3 /proc/loadavg)"
|
||||
30
proto-newpow/state-dataset/run_gpu.sh
Executable file
30
proto-newpow/state-dataset/run_gpu.sh
Executable file
|
|
@ -0,0 +1,30 @@
|
|||
#!/bin/bash
|
||||
# state-dataset prototype: GPU box 2 (RTX 4090). Runs inside /root/horizon-newpow/state-dataset after the rsync.
|
||||
# From the Mac:
|
||||
# rsync -az -e "ssh -i ~/.ssh/igneum-fleet -p <box-2-port>" proto-newpow/state-dataset/ root@<box-2-ip>:/root/horizon-newpow/state-dataset/
|
||||
# ssh -i ~/.ssh/igneum-fleet -p <box-2-port> root@<box-2-ip> 'cd /root/horizon-newpow/state-dataset && touch * && bash run_gpu.sh'
|
||||
set -euo pipefail
|
||||
export PATH=/usr/local/cuda/bin:$PATH
|
||||
cd "$(dirname "$0")"
|
||||
touch *.cu *.h *.c
|
||||
echo "== $(date -u) on $(hostname)"
|
||||
nvidia-smi --query-gpu=name,driver_version,clocks.max.sm,memory.total --format=csv,noheader
|
||||
nvcc --version | tail -1
|
||||
echo "== build"
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
|
||||
gcc -O2 -std=c11 -o verify_sd verify_sd.c
|
||||
echo "== control"
|
||||
./bench --mode control --batches 10 --items-out items_control.txt --dump dump_control.txt 4 --power-seconds 20 | tee control.log
|
||||
echo "== sd1"
|
||||
./bench --mode sd1 --batches 10 --items-out items_sd1.txt --dump dump_sd1.txt 4 --power-seconds 20 | tee sd1.log
|
||||
echo "== sd1 chunked (leaves streamed from pinned host memory in 64 and 256 MiB chunks)"
|
||||
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 64 | tee sd1-chunk64.log
|
||||
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 256 | tee sd1-chunk256.log
|
||||
echo "== second pass of the timed rows (repeatability)"
|
||||
./bench --mode control --batches 10 --power-seconds 20 | tee control2.log
|
||||
./bench --mode sd1 --batches 10 --power-seconds 20 | tee sd12.log
|
||||
echo "== verify (plain C, CPU core 2)"
|
||||
taskset -c 2 ./verify_sd --mode control --items items_control.txt --dump dump_control.txt | tee verify_control.log
|
||||
taskset -c 2 ./verify_sd --mode sd1 --items items_sd1.txt --dump dump_sd1.txt | tee verify_sd1.log
|
||||
echo "== logs to collect: run_gpu.log control*.log sd1*.log verify_*.log dump_*.txt items_*.txt"
|
||||
echo "== done $(date -u)"
|
||||
73
proto-newpow/state-dataset/sd.h
Normal file
73
proto-newpow/state-dataset/sd.h
Normal file
|
|
@ -0,0 +1,73 @@
|
|||
// state-dataset prototype (class "sd1", Horizon lane 8): the dataset commits to chain state.
|
||||
// Shared by kernel_sd.cu (device), bench.cu (host reference), verify_sd.c and cpu_rows.c (plain C).
|
||||
//
|
||||
// Scheme: item t is derived exactly as mh_item in memhard.h, except that a 64-byte STATE LEAF for item t is
|
||||
// XORed into the 16 initial state words before the first mixer:
|
||||
// s[0..7] = K, s[8..15] = t * MUL + RC (as today)
|
||||
// s[i] ^= leaf(t)[i] for i in 0..15 (new)
|
||||
// 8 rounds of (8 mixers + one dependent cache line), 8 final mixers (as today)
|
||||
// The hash kernel is unchanged.
|
||||
//
|
||||
// Leaf stand-in (SYNTHETIC, for the prototype only). In the real design leaf(t) is the t-th 64-byte leaf of the
|
||||
// canonical serialisation of the chain's execution state at the certified checkpoint 20 minutes before the day
|
||||
// boundary, zero padded. Here:
|
||||
// leaf(t) = mh_chacha_block(x) with
|
||||
// x = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574, S[0..7], t, 0, 0x49676e65, 0x53746174)
|
||||
// S[i] = K[i] ^ 0x5a5a5a5a (stand-in state root; K = the pack's day key words IGNEUM_KEY_INIT)
|
||||
// For the pack's day key this gives S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102.
|
||||
#pragma once
|
||||
#include "memhard.h"
|
||||
|
||||
#define SD_STATE_ROOT_XOR 0x5a5a5a5au
|
||||
#define SD_TAG0 0x49676e65u /* "Igne" */
|
||||
#define SD_TAG1 0x53746174u /* "Stat" */
|
||||
|
||||
IGNEUM_HD void mh_leaf(uint32_t t, uint32_t* leaf) {
|
||||
uint32_t x[16];
|
||||
x[0] = 0x61707865u; x[1] = 0x3320646eu; x[2] = 0x79622d32u; x[3] = 0x6b206574u;
|
||||
x[4] = 0x3067619fu ^ SD_STATE_ROOT_XOR;
|
||||
x[5] = 0x3c269176u ^ SD_STATE_ROOT_XOR;
|
||||
x[6] = 0x84a03b03u ^ SD_STATE_ROOT_XOR;
|
||||
x[7] = 0xf8c63294u ^ SD_STATE_ROOT_XOR;
|
||||
x[8] = 0xff977c5bu ^ SD_STATE_ROOT_XOR;
|
||||
x[9] = 0xe60def3eu ^ SD_STATE_ROOT_XOR;
|
||||
x[10] = 0x63630141u ^ SD_STATE_ROOT_XOR;
|
||||
x[11] = 0xb8fbcb58u ^ SD_STATE_ROOT_XOR;
|
||||
x[12] = t; x[13] = 0u; x[14] = SD_TAG0; x[15] = SD_TAG1;
|
||||
mh_chacha_block(x, leaf);
|
||||
}
|
||||
|
||||
// Item t under sd1, given its leaf (16 words). Same text as mh_item plus the one XOR line.
|
||||
IGNEUM_HD void mh_item_sd(const uint32_t* cache, const uint32_t* leaf, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0x3067619fu;
|
||||
s[1] = 0x3c269176u;
|
||||
s[2] = 0x84a03b03u;
|
||||
s[3] = 0xf8c63294u;
|
||||
s[4] = 0xff977c5bu;
|
||||
s[5] = 0xe60def3eu;
|
||||
s[6] = 0x63630141u;
|
||||
s[7] = 0xb8fbcb58u;
|
||||
s[8] = t * 0x42146205u + 0xbab68293u;
|
||||
s[9] = t * 0x52cbe0fbu + 0xcc162340u;
|
||||
s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
|
||||
s[11] = t * 0x6728907fu + 0xe62b8997u;
|
||||
s[12] = t * 0xd81d9751u + 0xc9c80297u;
|
||||
s[13] = t * 0x132952c3u + 0xf74a1654u;
|
||||
s[14] = t * 0xf60de277u + 0x3d704af5u;
|
||||
s[15] = t * 0x05358035u + 0x3cf522b7u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= leaf[i];
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
|
||||
}
|
||||
|
||||
// dataset[w] under sd1 without the dataset or the leaf array: derive the leaf and the item, take word w & 15.
|
||||
IGNEUM_HD uint32_t mh_word_sd(const uint32_t* cache, uint32_t w) {
|
||||
uint32_t leaf[16]; uint32_t s[16];
|
||||
mh_leaf(w >> 4u, leaf);
|
||||
mh_item_sd(cache, leaf, w >> 4u, s);
|
||||
return s[w & 15u];
|
||||
}
|
||||
57
proto-newpow/state-dataset/vectors.h
Normal file
57
proto-newpow/state-dataset/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x19b56348bc85304dull, 0xb08a9cfb44aa720full, 0xe1f8f627f780eff7ull, 0x2ff3e86ff1696161ull, 0xb65e0578c257e8acull, 0xf37b5705c2bebaaeull, 0xb19fef670982389eull, 0x2374331b28827a11ull,
|
||||
0xf492cdd05dda7f88ull, 0x37700f1385e19885ull, 0xa27524f3010b3e87ull, 0x6c8a24de8b938c43ull, 0x3dc017c8820cfd16ull, 0x5c80d37146657b68ull, 0x214082cac03f734aull, 0x1d67c665145d72f3ull,
|
||||
0x9fb0ae736eb87ff5ull, 0xcd7d416ae3a6be61ull, 0x0069cda85c6de58full, 0x0b7d82e717224baaull, 0x600960f6d7be30f0ull, 0xeb29032663b0c4d2ull, 0xf29583a5766408f1ull, 0x6c9532a99ad7314dull,
|
||||
0x9a949fd0959cc80full, 0xa20258520e7f6c25ull, 0xf2ab11f9bb032e38ull, 0xcc967bcd0c8d07c1ull, 0x37745267bb3231f2ull, 0x35a046048c2b69b3ull, 0xaa51834cd3f364f3ull, 0x359192708e4f754aull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x62fb132a9943127aull, 0x0b703e577e7f4ecaull, 0xf9f24f5522ce7593ull, 0x3cf5c516abc4332aull, 0xd25523f5f6d7a127ull, 0xd2081a002f983682ull, 0xbf46c54e9b3c4254ull, 0xca362e291e5e5f4dull,
|
||||
0x6039712f10f457a3ull, 0x8a34b7cabf97c23bull, 0xa473c6a2e0bf59bcull, 0x6cf3926513a4b069ull, 0x297ec2998376a40dull, 0x8efd7f601a8f28dbull, 0x8e72532dfdc1e544ull, 0x917c2b2ebe2a7e00ull,
|
||||
0x923fbb2d2f635c25ull, 0xce864ea5c0dedad9ull, 0x4b8ec7e874e446efull, 0x1b69b69465449196ull, 0x5ef3a8a6edb369cfull, 0x06c263ef9ce63fc4ull, 0x9c2048fd9d9e2639ull, 0x457fdd96ca4a138eull,
|
||||
0xd1904018b8d7b6e3ull, 0x8682312fb2e96ab8ull, 0xdc3257e0d0f979a5ull, 0xa51b0a8519d87db5ull, 0x334f08ec056e618bull, 0x3464ce71dc65119dull, 0x6a4d6df066332e04ull, 0x7d7866cb9cfca8ffull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x86b6cb0e13d89b03ull, 0x96299a3f19d7ef15ull, 0x67d2c55100d2f876ull, 0x0a4dfe97d671b728ull, 0x41e4489014d42595ull, 0xf11cb1958c0c0e82ull, 0xf8b70b0c0a03175full, 0x632299df87d5063eull,
|
||||
0xe198417776130492ull, 0x8ffc5449290d7be2ull, 0x5f2e264eb1311f1bull, 0x988376463ac88586ull, 0x83969eadda489c26ull, 0xbed0a2c3f255d306ull, 0x1a949d271961a819ull, 0x5bce06eb6984725cull,
|
||||
0x94d5d6a1b0ffd4e9ull, 0xf3c78bae6c2182b4ull, 0xb97e9fe1bbfcdd55ull, 0x70262d1d4c0eccb2ull, 0x1fc93b427dba28d9ull, 0x02b2e3c4317f2a2dull, 0x54d3d42a588edcb9ull, 0x79998677846e7cceull,
|
||||
0x486522a5425f821aull, 0x95fa88e933360e52ull, 0xc8bae2da2b883f6cull, 0xbe3eb610ad33614full, 0x20efb3c4de82907full, 0xd6b650cfedfb26b7ull, 0x8c24447a646dba26ull, 0x9c004678515e44ecull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0xfdad4319u, 0x1a7b68e1u, 0xde6db608u, 0x13d73892u, 0xd17f447au, 0xb2221ccfu, 0x9db004bdu, 0x57d7d367u,
|
||||
0xdbc4cf34u, 0x697c009au, 0xc43af1d4u, 0x97f12b2eu, 0x74c37cd0u, 0xc651ea15u, 0x665a6d29u, 0x22330a2du
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xa83e7aa6u;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0x3230bc7bu, 0x7fbfe2c9u, 0xb2690991u, 0x1745c7c5u, 0x0ab0ccafu, 0x1bf87d6bu, 0x160139fdu, 0x719817acu, 0x0155df4bu, 0xbe1e86c3u, 0x680bcd6cu, 0x79c3dc6cu, 0x181e7e5fu, 0x0713a109u, 0xc705dd9fu, 0x3933b7a8u, 0xdd1c0431u, 0x50522b30u, 0xa0020b38u, 0xbff39e96u, 0x21b67e18u, 0x740f8db3u, 0x2baba568u, 0x2c9bef83u, 0x0ad9b671u, 0xc4327869u, 0x7b4fd7d0u, 0x2c29965fu, 0xec56f15fu, 0x61111746u, 0x303a1d6eu, 0xbddcfd1au, 0xf829a355u, 0x6d5df2a9u, 0x01ab8e44u, 0x06d13507u, 0xda8dcfc6u, 0x01a703e1u, 0xafe7d2c1u, 0xc091c3a2u, 0xac1814feu, 0x6e6ff62au, 0x8fdf01bau, 0xdd3f7159u, 0xdfa0d75cu, 0x26684c35u, 0x7f441e63u, 0x88df2570u, 0x8aa4d5ebu, 0xcc816c05u, 0x434df890u, 0xcd392ad6u, 0x1ab4cb63u, 0x595926fau, 0x7cd76b41u, 0x20cb95c4u, 0x13cf823fu, 0xf9daf901u, 0xff9af40au, 0x2c7dfa51u, 0x871206dbu, 0x938c116cu, 0xb64bf199u, 0x5751f874u
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u,
|
||||
0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du,
|
||||
0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull;
|
||||
270
proto-newpow/state-dataset/verify_sd.c
Normal file
270
proto-newpow/state-dataset/verify_sd.c
Normal file
|
|
@ -0,0 +1,270 @@
|
|||
/* state-dataset prototype: plain C verifier (Horizon lane 8, class sd1).
|
||||
*
|
||||
* 1. Fills the 256 MiB cache on the host (mh_cache_segment over all 65536 segments), times it, checks the FNV.
|
||||
* 2. --items <file>: the 16 words of 1,024 random items the bench read back from the GPU, against
|
||||
* mh_item (control) or mh_item_sd(hostCache, hostLeaf(t), t) (sd1). Reports "n of 1024 items equal".
|
||||
* 3. --dump <file>: a 32-lane register-major interpreter of the pack's 64-instruction program, loads through
|
||||
* mh_word (control) or mh_word_sd (sd1) on the host cache, against the GPU's warps. Reports "n of N lanes equal".
|
||||
* Lane 0 of the first warp is instrumented: its 128 load addresses, distinct items, and the Merkle arithmetic.
|
||||
* 4. Per-unit rows on this CPU (ms per unit, 100 units): (ii) 4,096 leaf derivations, (iii) 4,096 x mh_item
|
||||
* naive, and 4,096 x (leaf + mh_item_sd). The full B.1 rows with the 2 GiB and 8 GiB leaf arrays are cpu_rows.c.
|
||||
*
|
||||
* Build: gcc -O2 -std=c11 -o verify_sd verify_sd.c
|
||||
* Run: taskset -c 2 ./verify_sd --mode sd1 --items items_sd1.txt --dump dump_sd1.txt
|
||||
*/
|
||||
#define _POSIX_C_SOURCE 200809L
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <time.h>
|
||||
#define IGNEUM_NO_CUDA
|
||||
#include "program.h"
|
||||
#include "vectors.h"
|
||||
#include "memhard.h"
|
||||
#include "sd.h"
|
||||
|
||||
static double nowMs(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1e3 + ts.tv_nsec / 1e6; }
|
||||
static uint64_t fnv1a64(const void* p, size_t n) {
|
||||
const uint8_t* b = (const uint8_t*)p; uint64_t h = 0xcbf29ce484222325ull;
|
||||
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
|
||||
return h;
|
||||
}
|
||||
static uint64_t splitmix64(uint64_t* s) {
|
||||
*s += 0x9E3779B97F4A7C15ull; uint64_t z = *s;
|
||||
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull; z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
|
||||
return z ^ (z >> 31);
|
||||
}
|
||||
|
||||
static uint32_t* hCache;
|
||||
static int gSd1 = 0;
|
||||
|
||||
/* ---- the program, ported from kernel.cu (igneum_hash), register-major over 32 lanes ---- */
|
||||
static uint32_t splitmix32(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }
|
||||
static uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
static uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
static uint32_t umulhi(uint32_t a, uint32_t b) { return (uint32_t)(((uint64_t)a * (uint64_t)b) >> 32); }
|
||||
|
||||
/* load instrumentation for one lane */
|
||||
static uint32_t gTraceLane = 0xffffffffu;
|
||||
static uint32_t gTrace[128];
|
||||
static int gTraceN = 0;
|
||||
|
||||
static uint32_t LOAD(uint32_t lane, uint32_t r, uint32_t mask) {
|
||||
uint32_t w = r & mask;
|
||||
if (lane == gTraceLane && gTraceN < 128) gTrace[gTraceN++] = w;
|
||||
return gSd1 ? mh_word_sd(hCache, w) : mh_word(hCache, w);
|
||||
}
|
||||
|
||||
/* rX[l] ^= rY[l ^ m] for all lanes, through a copy of rY */
|
||||
#define SHFL_XOR_INTO(rX, rY, m) do { uint32_t tmp_[32]; for (int l_ = 0; l_ < 32; ++l_) tmp_[l_] = rY[l_ ^ (m)]; for (int l_ = 0; l_ < 32; ++l_) rX[l_] ^= tmp_[l_]; } while (0)
|
||||
|
||||
static void interpret_warp(uint32_t base, uint32_t mask, uint64_t out[32]) {
|
||||
uint32_t r0[32], r1[32], r2[32], r3[32], r4[32], r5[32], r6[32], r7[32], sel[32];
|
||||
for (int l = 0; l < 32; ++l) {
|
||||
uint32_t nonce = base + (uint32_t)l; uint32_t x;
|
||||
x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0[l] = x ^ 0x1a155b25u;
|
||||
x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1[l] = x ^ 0xfddfb732u;
|
||||
x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2[l] = x ^ 0x4b5af2e8u;
|
||||
x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3[l] = x ^ 0xc55caf33u;
|
||||
x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4[l] = x ^ 0xa27c13b7u;
|
||||
x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5[l] = x ^ 0x06628a48u;
|
||||
x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6[l] = x ^ 0x03852469u;
|
||||
x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7[l] = x ^ 0x67a9a7beu;
|
||||
}
|
||||
#define L for (int l = 0; l < 32; ++l)
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
L sel[l] = r0[l];
|
||||
L r2[l] = r3[l] * r4[l] + r2[l]; /* 0 mad */
|
||||
L r2[l] = r1[l] * r1[l] + r2[l]; /* 1 mad */
|
||||
L r2[l] = r3[l] * r2[l] + r2[l]; /* 2 mad */
|
||||
L r3[l] = r3[l] ^ r5[l]; /* 3 xor */
|
||||
L r7[l] = r7[l] ^ LOAD(l, r2[l], mask); /* 4 load */
|
||||
L r5[l] = r5[l] ^ LOAD(l, r7[l], mask); /* 5 load */
|
||||
SHFL_XOR_INTO(r1, r4, 8); /* 6 shfl */
|
||||
SHFL_XOR_INTO(r7, r3, 8); /* 7 shfl */
|
||||
L r1[l] = umulhi(r1[l], r5[l]); /* 8 mulhi */
|
||||
L r6[l] = rotr_var(r6[l], r3[l]); /* 9 rotr */
|
||||
L r3[l] = r3[l] | r4[l]; /* 10 or */
|
||||
L r4[l] = r4[l] ^ LOAD(l, r3[l], mask); /* 11 load */
|
||||
L r0[l] = umulhi(r0[l], r4[l]); /* 12 mulhi */
|
||||
L r5[l] = r5[l] + r1[l] + ((((sel[l] >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); /* 13 add */
|
||||
L r0[l] = r0[l] ^ LOAD(l, r4[l], mask); /* 14 load */
|
||||
L r2[l] = r2[l] - r4[l]; /* 15 sub */
|
||||
L r2[l] = r2[l] ^ LOAD(l, r0[l], mask); /* 16 load */
|
||||
L r7[l] = r7[l] ^ LOAD(l, r2[l], mask); /* 17 load */
|
||||
SHFL_XOR_INTO(r7, r3, 4); /* 18 shfl */
|
||||
L r5[l] = r5[l] * r0[l]; /* 19 mul */
|
||||
SHFL_XOR_INTO(r3, r4, 2); /* 20 shfl */
|
||||
SHFL_XOR_INTO(r2, r4, 16); /* 21 shfl */
|
||||
L r6[l] = umulhi(r6[l], r2[l]); /* 22 mulhi */
|
||||
L r6[l] = r6[l] ^ LOAD(l, r1[l], mask); /* 23 load */
|
||||
L r5[l] = r5[l] * r0[l]; /* 24 mul */
|
||||
L r5[l] = rotl_imm(r5[l], 19u); /* 25 rotl */
|
||||
SHFL_XOR_INTO(r7, r6, 2); /* 26 shfl */
|
||||
L r0[l] = r0[l] ^ r5[l]; /* 27 xor */
|
||||
L r0[l] = r0[l] ^ r4[l]; /* 28 xor */
|
||||
L r3[l] = r3[l] - r0[l]; /* 29 sub */
|
||||
L r5[l] = r5[l] * r1[l]; /* 30 mul */
|
||||
L r7[l] = r7[l] ^ LOAD(l, r2[l], mask); /* 31 load */
|
||||
L r1[l] = r1[l] ^ LOAD(l, r0[l], mask); /* 32 load */
|
||||
L r5[l] = r5[l] ^ r6[l]; /* 33 xor */
|
||||
L r5[l] = r5[l] ^ LOAD(l, r1[l], mask); /* 34 load */
|
||||
L r0[l] = umulhi(r0[l], r5[l]); /* 35 mulhi */
|
||||
SHFL_XOR_INTO(r5, r2, 4); /* 36 shfl */
|
||||
L r7[l] = r7[l] ^ LOAD(l, r0[l], mask); /* 37 load */
|
||||
L r3[l] = r3[l] + r1[l] + ((((sel[l] >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); /* 38 add */
|
||||
SHFL_XOR_INTO(r1, r5, 4); /* 39 shfl */
|
||||
L r2[l] = r2[l] ^ r5[l]; /* 40 xor */
|
||||
L r3[l] = r6[l] * r3[l] + r3[l]; /* 41 mad */
|
||||
L r6[l] = r6[l] - r7[l]; /* 42 sub */
|
||||
L r7[l] = r7[l] ^ r0[l]; /* 43 xor */
|
||||
L r1[l] = r1[l] ^ LOAD(l, r7[l], mask); /* 44 load */
|
||||
L r2[l] = r2[l] * r3[l]; /* 45 mul */
|
||||
L r1[l] = umulhi(r1[l], r5[l]); /* 46 mulhi */
|
||||
L r4[l] = r4[l] - r3[l]; /* 47 sub */
|
||||
L r2[l] = rotr_var(r2[l], r6[l]); /* 48 rotr */
|
||||
L r3[l] = r3[l] ^ LOAD(l, r5[l], mask); /* 49 load */
|
||||
L r1[l] = r1[l] + r5[l] + ((((sel[l] >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); /* 50 add */
|
||||
L r0[l] = r0[l] * r2[l]; /* 51 mul */
|
||||
L r0[l] = r0[l] + r2[l] + ((((sel[l] >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); /* 52 add */
|
||||
L r1[l] = r1[l] + r0[l] + ((((sel[l] >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); /* 53 add */
|
||||
L r7[l] = rotl_imm(r7[l], 14u); /* 54 rotl */
|
||||
L r3[l] = r3[l] + r7[l] + ((((sel[l] >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); /* 55 add */
|
||||
L r6[l] = r6[l] ^ LOAD(l, r7[l], mask); /* 56 load */
|
||||
L r1[l] = rotr_var(r1[l], r5[l]); /* 57 rotr */
|
||||
L r5[l] = r5[l] ^ LOAD(l, r4[l], mask); /* 58 load */
|
||||
L r6[l] = r6[l] ^ LOAD(l, r2[l], mask); /* 59 load */
|
||||
L r3[l] = r5[l] * r0[l] + r3[l]; /* 60 mad */
|
||||
L r5[l] = r5[l] + r7[l] + ((((sel[l] >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); /* 61 add */
|
||||
L r4[l] = r4[l] + r6[l] + ((((sel[l] >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); /* 62 add */
|
||||
L r5[l] = rotl_imm(r5[l], 19u); /* 63 rotl */
|
||||
}
|
||||
L {
|
||||
uint32_t lo = r0[l] ^ rotl_imm(r1[l], 7u) ^ rotl_imm(r2[l], 14u) ^ rotl_imm(r3[l], 21u);
|
||||
uint32_t hi = r4[l] ^ rotl_imm(r5[l], 9u) ^ rotl_imm(r6[l], 18u) ^ rotl_imm(r7[l], 27u);
|
||||
out[l] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
#undef L
|
||||
}
|
||||
|
||||
static int cmp_u32(const void* a, const void* b) { uint32_t x = *(const uint32_t*)a, y = *(const uint32_t*)b; return x < y ? -1 : x > y; }
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
const char* itemsFile = NULL; const char* dumpFile = NULL; int units = 100; int rows = 1;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
if (!strcmp(argv[i], "--mode") && i + 1 < argc) { ++i; gSd1 = !strcmp(argv[i], "sd1"); }
|
||||
else if (!strcmp(argv[i], "--items") && i + 1 < argc) itemsFile = argv[++i];
|
||||
else if (!strcmp(argv[i], "--dump") && i + 1 < argc) dumpFile = argv[++i];
|
||||
else if (!strcmp(argv[i], "--units") && i + 1 < argc) units = atoi(argv[++i]);
|
||||
else if (!strcmp(argv[i], "--no-rows")) rows = 0;
|
||||
else { printf("usage: verify_sd --mode control|sd1 [--items f] [--dump f] [--units 100] [--no-rows]\n"); return 2; }
|
||||
}
|
||||
const uint32_t words = 1u << IGNEUM_DATASET_LOG2, mask = words - 1u, nItems = words / 16u;
|
||||
const uint32_t cacheWords = 1u << IGNEUM_CACHE_LOG2_WORDS;
|
||||
printf("verify_sd mode %s\n", gSd1 ? "sd1" : "control");
|
||||
|
||||
hCache = (uint32_t*)malloc((size_t)cacheWords * 4u);
|
||||
double t0 = nowMs();
|
||||
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(hCache, seg);
|
||||
double cacheMs = nowMs() - t0;
|
||||
uint64_t fnv = fnv1a64(hCache, (size_t)cacheWords * 4u);
|
||||
printf("host cache fill: %.1f ms one thread; FNV-1a 64 %016llx vs Mac %016llx %s\n", cacheMs, (unsigned long long)fnv, (unsigned long long)IGNEUM_CACHE_FNV64, fnv == IGNEUM_CACHE_FNV64 ? "PASS" : "FAIL");
|
||||
|
||||
/* control sanity: the first Mac vector through the interpreter */
|
||||
if (!gSd1) {
|
||||
uint64_t out[32]; interpret_warp(0u, mask, out);
|
||||
int bad = 0; for (int l = 0; l < 32; ++l) if (out[l] != IGNEUM_VEC_OUT[0][l]) ++bad;
|
||||
printf("interpreter vs Mac vector base 0: %d of 32 lanes equal %s\n", 32 - bad, bad == 0 ? "PASS" : "FAIL");
|
||||
}
|
||||
|
||||
int itemsEqual = -1;
|
||||
if (itemsFile) {
|
||||
FILE* f = fopen(itemsFile, "r");
|
||||
if (!f) { printf("FAIL: cannot open %s\n", itemsFile); return 2; }
|
||||
char mode[16]; int n = 0;
|
||||
if (fscanf(f, "mode %15s\nnitems %d\n", mode, &n) != 2) { printf("FAIL: bad items file header\n"); return 2; }
|
||||
if (strcmp(mode, gSd1 ? "sd1" : "control") != 0) printf("WARNING: items file mode %s, verifier mode %s\n", mode, gSd1 ? "sd1" : "control");
|
||||
itemsEqual = 0; int read = 0;
|
||||
for (int k = 0; k < n; ++k) {
|
||||
uint32_t t, got[16], want[16];
|
||||
if (fscanf(f, "%u", &t) != 1) break;
|
||||
for (int i = 0; i < 16; ++i) if (fscanf(f, " %x", &got[i]) != 1) break;
|
||||
++read;
|
||||
if (t >= nItems) continue;
|
||||
if (gSd1) { uint32_t leaf[16]; mh_leaf(t, leaf); mh_item_sd(hCache, leaf, t, want); } else mh_item(hCache, t, want);
|
||||
if (memcmp(got, want, 64) == 0) ++itemsEqual;
|
||||
else if (itemsEqual == k) printf(" item %u: gpu %08x host %08x (first difference)\n", t, got[0], want[0]);
|
||||
}
|
||||
fclose(f);
|
||||
printf("item bit-exactness (plain C, %s on the host cache): %d of %d items equal %s\n", gSd1 ? "mh_item_sd with mh_leaf" : "mh_item", itemsEqual, read, itemsEqual == read && read > 0 ? "PASS" : "FAIL");
|
||||
}
|
||||
|
||||
int lanesEqual = -1, lanesTotal = 0;
|
||||
if (dumpFile) {
|
||||
FILE* f = fopen(dumpFile, "r");
|
||||
if (!f) { printf("FAIL: cannot open %s\n", dumpFile); return 2; }
|
||||
char mode[16]; int nw = 0;
|
||||
if (fscanf(f, "mode %15s\nwarps %d\n", mode, &nw) != 2) { printf("FAIL: bad dump header\n"); return 2; }
|
||||
if (strcmp(mode, gSd1 ? "sd1" : "control") != 0) printf("WARNING: dump mode %s, verifier mode %s\n", mode, gSd1 ? "sd1" : "control");
|
||||
lanesEqual = 0;
|
||||
double tI = nowMs();
|
||||
for (int w = 0; w < nw; ++w) {
|
||||
uint32_t base; uint64_t got[32], out[32];
|
||||
if (fscanf(f, "base %u\n", &base) != 1) break;
|
||||
for (int l = 0; l < 32; ++l) { unsigned long long v; if (fscanf(f, "%llx\n", &v) != 1) break; got[l] = v; }
|
||||
if (w == 0) { gTraceLane = 0u; gTraceN = 0; } else gTraceLane = 0xffffffffu; /* trace lane 0 of the first warp only */
|
||||
interpret_warp(base, mask, out);
|
||||
int eq = 0; for (int l = 0; l < 32; ++l) if (out[l] == got[l]) ++eq;
|
||||
lanesEqual += eq; lanesTotal += 32;
|
||||
printf("warp base %u: %d of 32 lanes equal (lane 0 gpu %016llx interp %016llx)\n", base, eq, (unsigned long long)got[0], (unsigned long long)out[0]);
|
||||
}
|
||||
fclose(f);
|
||||
printf("GPU hash outputs vs interpreter: %d of %d lanes equal %s (%.0f ms interpreting, %d derivations per warp)\n", lanesEqual, lanesTotal, lanesEqual == lanesTotal && lanesTotal > 0 ? "PASS" : "FAIL", nowMs() - tI, 32 * IGNEUM_LOADS_PER_HASH);
|
||||
|
||||
/* B.3: the state sample of lane 0 of the first warp */
|
||||
if (gTraceN == 128) {
|
||||
uint32_t items[128]; for (int i = 0; i < 128; ++i) items[i] = gTrace[i] >> 4;
|
||||
qsort(items, 128, sizeof(uint32_t), cmp_u32);
|
||||
int distinctItems = 1; for (int i = 1; i < 128; ++i) if (items[i] != items[i - 1]) ++distinctItems;
|
||||
uint32_t ws[128]; memcpy(ws, gTrace, sizeof(ws)); qsort(ws, 128, sizeof(uint32_t), cmp_u32);
|
||||
int distinctWords = 1; for (int i = 1; i < 128; ++i) if (ws[i] != ws[i - 1]) ++distinctWords;
|
||||
printf("lane 0 of warp base 0: 128 loads touch %d distinct words in %d distinct items (of 2^%d items)\n", distinctWords, distinctItems, IGNEUM_DATASET_LOG2 - 4);
|
||||
printf(" first 8 load words: %u %u %u %u %u %u %u %u\n", gTrace[0], gTrace[1], gTrace[2], gTrace[3], gTrace[4], gTrace[5], gTrace[6], gTrace[7]);
|
||||
long openings = 128, depth = 25, hash = 32;
|
||||
printf(" Merkle openings for a 2 GiB state (2^25 leaves of 64 B, depth %ld, %ld-byte hashes): %ld openings x %ld x %ld = %ld bytes of siblings + %ld x 64 = %ld bytes of leaves = %ld bytes (%.1f KiB) per lane, uncompressed; %d distinct items need only %d openings = %ld bytes\n",
|
||||
depth, hash, openings, depth, hash, openings * depth * hash, openings, openings * 64, openings * depth * hash + openings * 64, (openings * depth * hash + openings * 64) / 1024.0,
|
||||
distinctItems, distinctItems, (long)distinctItems * (depth * hash + 64));
|
||||
}
|
||||
}
|
||||
|
||||
if (rows) {
|
||||
/* (ii) 4,096 leaf derivations per unit; (iii) 4,096 x mh_item naive; and leaf + mh_item_sd */
|
||||
uint64_t s = 0x5d1ull; volatile uint32_t sink = 0;
|
||||
double sumLeaf = 0, sumItem = 0, sumSd = 0, minItem = 1e9, maxItem = 0;
|
||||
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
|
||||
for (int u = 0; u < units; ++u) {
|
||||
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)splitmix64(&s) & (nItems - 1u);
|
||||
double a = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16]; mh_leaf(ts[k], leaf); sink ^= leaf[0]; }
|
||||
double b = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { uint32_t it[16]; mh_item(hCache, ts[k], it); sink ^= it[0]; }
|
||||
double c = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16], it[16]; mh_leaf(ts[k], leaf); mh_item_sd(hCache, leaf, ts[k], it); sink ^= it[0]; }
|
||||
double d = nowMs();
|
||||
sumLeaf += b - a; sumItem += c - b; sumSd += d - c;
|
||||
if (c - b < minItem) minItem = c - b;
|
||||
if (c - b > maxItem) maxItem = c - b;
|
||||
}
|
||||
free(ts);
|
||||
printf("per-unit rows on this CPU (4,096 per unit, %d units, one thread, ms per unit):\n", units);
|
||||
printf(" (ii) 4,096 leaf derivations (ChaCha12 block each, no array): %.3f ms\n", sumLeaf / units);
|
||||
printf(" (iii) 4,096 x mh_item on the host cache, naive (no interleaving): %.2f ms (min %.2f, max %.2f)\n", sumItem / units, minItem, maxItem);
|
||||
printf(" 4,096 x (leaf + mh_item_sd), naive: %.2f ms\n", sumSd / units);
|
||||
printf(" reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.\n");
|
||||
}
|
||||
|
||||
printf("RESULT verify mode=%s host_cache_ms=%.1f items_equal=%d lanes_equal=%d/%d\n", gSd1 ? "sd1" : "control", cacheMs, itemsEqual, lanesEqual, lanesTotal);
|
||||
free(hCache);
|
||||
return 0;
|
||||
}
|
||||
3
sim/horizon/new-pow/README.md
Normal file
3
sim/horizon/new-pow/README.md
Normal file
|
|
@ -0,0 +1,3 @@
|
|||
# sim/horizon/new-pow
|
||||
|
||||
`chip_rows.py`: the scheme B chip rows (f = 1 chip plus a tensor array at k times the card's block energy) from the measured card figures. Run: `python3 chip_rows.py --card-uj 2.56 --card-uj-r 3.0 --r 128`. Inputs and outputs are explained in `docs/analysis/horizon/new-pow.md` section 5.3.
|
||||
36
sim/horizon/new-pow/chip_rows.py
Normal file
36
sim/horizon/new-pow/chip_rows.py
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Chip rows for the Horizon new-pow lane (scheme B, the tensor-shaped shadow), by the chip-model-v3 section 5 method.
|
||||
|
||||
Usage:
|
||||
python3 chip_rows.py --card-uj <microjoules per hash of the honest card at R=0> \
|
||||
--card-uj-r <microjoules per hash at the measured R> --r <R> \
|
||||
[--mem-uj 0.466] [--k 1.0 0.5 0.3]
|
||||
|
||||
Model (docs/analysis/chip-model-v3.md 5.4 and docs/analysis/latency-shadow-2026-10-06.md 6):
|
||||
chip energy per hash = memory system energy (f = 1 chip: 0.466 uJ GDDR7, 0.321 uJ one HBM3 stack)
|
||||
+ block energy on the chip = (card block energy) x k
|
||||
card block energy = card_uj_r - card_uj (the measured marginal of the block on the honest card)
|
||||
gain per joule = card_uj_r / chip_uj
|
||||
Every chip figure is arithmetic on cited figures and approximate; the card figures are measured and named in the lane file.
|
||||
"""
|
||||
import argparse
|
||||
|
||||
p = argparse.ArgumentParser()
|
||||
p.add_argument("--card-uj", type=float, required=True, help="honest card microjoules per hash at R = 0 (measured)")
|
||||
p.add_argument("--card-uj-r", type=float, required=True, help="honest card microjoules per hash at the measured R")
|
||||
p.add_argument("--r", type=int, required=True, help="mm8 steps per iteration (8 R per hash)")
|
||||
p.add_argument("--mem-uj", type=float, nargs="+", default=[0.466, 0.321], help="f = 1 chip memory energy per hash, uJ (GDDR7, HBM3 one stack)")
|
||||
p.add_argument("--k", type=float, nargs="+", default=[1.0, 0.5, 0.3], help="chip block energy per op over the card's")
|
||||
a = p.parse_args()
|
||||
|
||||
block = a.card_uj_r - a.card_uj
|
||||
macs = 8 * a.r * 1024
|
||||
print(f"R = {a.r}: {8*a.r} mm8 per hash, {macs} multiply-adds per hash")
|
||||
print(f"card: {a.card_uj:.3f} uJ at R = 0, {a.card_uj_r:.3f} uJ at R = {a.r}; block {block:.3f} uJ = {block*1e6/macs if macs else 0:.3f} pJ per multiply-add")
|
||||
print("| memory | k | chip uJ per hash | gain per joule (card over chip) | gain at R = 0 for comparison |")
|
||||
print("|---|---|---|---|---|")
|
||||
names = ["GDDR7 (16 devices)", "HBM3 (one stack)", "HBM3 (eight stacks)"]
|
||||
for i, m in enumerate(a.mem_uj):
|
||||
for k in a.k:
|
||||
chip = m + block * k
|
||||
print(f"| {names[i] if i < len(names) else m} | {k} | {chip:.3f} | {a.card_uj_r/chip:.2f}x | {a.card_uj/m:.2f}x |")
|
||||
Loading…
Reference in a new issue