GPU workers on the real hash: serve protocol in Metal, CUDA and OpenCL, Windows mining package, devnet v1 log

proto-metal --serve compiles igneum_hash_bound at runtime and mines jobs from stdin (init words in buffer 3, program
and dataset cached per seed); proto-cuda and proto-opencl host serve modes from the pack's kernel_bound.cu / .cl with a
seed guard; --vendor device filter for OpenCL. windows-miner/: START-MINING.bat + start-mining.ps1 (GPU and tool
detection, pack export, cached builds, MINERS identities per vendor, status every 30 s, uploads every 60 s, rebuild on
seed change, Ctrl+C summary), README.txt, make-package.sh (igneum-mine-test.zip with a cross-compiled miner).
docs: fork-divergence devnet v1, bench-log entry with the CPU, Metal and overnight numbers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-03 19:18:19 +00:00
parent b02b87dfd0
commit 27e944d8eb
13 changed files with 1041 additions and 27 deletions

View file

@ -222,3 +222,14 @@ Emission: 80/20 exact on 39 of 39 single-payee coinbases (the 40th merged two bl
Unit tests, release profile: kaspa-consensus-core igneum 8 pass (subsidy table, ramp, split, cap), params window test 1 pass, kaspa-pow 5 pass (stub) and 6 pass with `--features igneum-pow` (engine smoke, one 256 MiB cache, 0.39 s), kaspa-consensus coinbase 8 pass.
Per-second subsidy, 8 decimals: 3,168,808,781 units (31.68808781 coins) for years 0 to 2, 1,584,404,390 for years 2 to 4, 792,202,195 for years 4 to 6, 1 unit in period 31, 0 from period 32; ramp day 0 is 316,880,878; the sum is under the 4,000,000,000-coin cap by less than 100 coins.
Not done: the devnet ran on the kHeavyHash stub, not the lottery engine (the miner has no igneum-pow path yet and the lane hash does not absorb the header); no VDF seed, no finality, no prover payout; 8 versus 18 decimals open (docs/fork-divergence.md).
## 3 October 2026, first devnet blocks on the real lottery hash: CPU, then Metal GPU, three worker implementations (consensus-engineer)
Machine: Apple M5 Max (18 logical cores, 40 GPU cores), rustc 1.99.0, Swift 5.8.1, fork `vendor/igneum-node` at "PoW: header-bound lottery engine, real-hash miner modes, genesis bits 0x1e400000" plus the overnight genesis commit; `igneum-pow` at "igneum-pow: header binding ...". Release builds with `--features igneum-pow`.
Binding (spec 01 section 1.6, O-1.9, `igneum-pow/src/bind.rs`): I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32), lane nonce = low 32 bits of the 64-bit header nonce, H = header hash with the nonce zeroed and the timestamp kept, pow256 = lane hash in the top 64 bits with zero low bits (pow <= target is exactly lane <= target >> 192). Interim day seed "igneum-day/" || day_le64 (day = timestamp_ms / 86,400,000); epoch seed = the 32 bytes of the epoch block hash (genesis for epoch 0). 8 bound vectors in `igneum-pow/README.md`; 29 crate tests pass; the 288 pack vectors and the pack files are unchanged, `program_bound.metal`, `kernel_bound.cu` and `kernel_bound.cl` are new pack files.
CPU rate on the real hash: 0.44 ms per 32-lane warp on one core (crate bench); 0.142 MH/s with one 6-thread miner (1.35 ms per warp per thread); 0.294 MH/s with three 6-thread miners at once (2.0 ms per warp per thread, memory-latency bound), 0.37 MH/s later in the run. Genesis bits for the CPU devnet: 0x1e400000 = 2^18 expected hashes per block.
CPU devnet, 3 nodes (node 1 RPC and p2p on 0.0.0.0, ports 26610/26611, 26620/26621, 26630/26631), three 6-thread `igneum-miner mine --engine igneum-pow`, 660 s: found 252 + 309 + 272 = 833 blocks, 0 rejected; watch window 641 s: 825 blocks on all three nodes = 1.29 blocks/s; sink identical on all nodes at 61 of 64 samples (tips 1 or 2, twice 3). DAA: 1.43 blocks/s over the first 600 blocks at difficulty 131,072, then difficulty 184,000 to 187,000 (1.41x) and 1.03 blocks/s over blocks 610 to 816. Each node logged "PoW accepted <hash> by igneum-lottery-v1-bound" 833 times, 0 rejections during the run. `igneum-miner bad-nonce`: Reject(BlockInvalid) and a "PoW rejected ... by igneum-lottery-v1-bound" line. `inspect 30`: vote_key_hash identical on all 3 nodes for 30 of 30, 80/20 exact on 19 of 19 single-payee coinbases. Epoch 0 seed = devnet genesis hash 03115da0...d86b, day 20729, program 104 loads/hash.
Metal GPU worker (`proto-metal/igneum-bench --serve`, runtime-compiled `igneum_hash_bound`, init words in buffer 3, 2^22 nonces per dispatch, CPU-side scan of the 64-bit outputs) driven by `igneum-miner --worker` on node 1 of the same devnet, 300 s: 506 jobs of 2^24 nonces, 5,636 blocks found and accepted, 0 rejected, 0 CPU/GPU mismatches (every found nonce is re-hashed on the CPU before submit); nodes 833 -> 5,073 blocks on all three (14.57 blocks/s over the 291 s watch window), sink identical at 28 of 29 samples; difficulty 180,562 -> 7,134,312 (39x) and still rising at the end. Epoch change crossed live at DAA 3,600: new seed 76c39fcf..., 128-load program compiled by the worker in 129 ms (first program 6.5 ms), no rejected blocks across the change. Rate through the worker: 32.4 MH/s inside jobs, 28.2 MH/s wall, against 45.2 MH/s raw bench (`igneum-bench` 2^22 x 4 batches): the gap is the read-back and CPU scan per dispatch, one command buffer per dispatch, and the template round trip per job. Early in the run one job found about 45 sibling blocks of one template; the miner now submits one block per job.
Overnight devnet (started 20:07 BST): genesis bits 0x1d100000 = 2^28 expected hashes per block, sized for RTX 5090 229 + M5 Max 45 + gfx1036 4 = 278 MH/s (1.04 blocks/s); genesis hash edc4fa84...fb07; epoch 0 program 136 loads/hash compiled in 69.5 ms on Metal; node 1 alone (0.0.0.0:26610/26611, `caffeinate -dims`, logs /tmp/igneum-devnet/node1.log) with the Metal worker (`/tmp/igneum-devnet/metal-worker.log`): 31 blocks in 273 s = 0.11 blocks/s at 31.1 MH/s, as expected for 2^28 until the PC joins.
Worker implementations, all bit-exact with `igneum-pow hash-bound` for a 96-nonce job across the 2^32 lane boundary (nonces 4294967264 to 4294967359, prehash ab x 32, devnet seeds): Metal (M5 Max GPU), OpenCL (`proto-opencl/host.c --serve`, Apple OpenCL 1.2 on the M5 Max, local-memory exchange, `--vendor` device filter added), CUDA (`proto-cuda/host.cu --serve` with the pack's `kernel_bound.cu`, through the clang emulation shim: bench PASS on the Rust-exported devnet pack, serve job bit-exact, seed-mismatch error exercised). CUDA and OpenCL are built ahead of time per pack; the Windows launcher re-exports and rebuilds at each epoch or day change (miner exit 42). `igneum-miner.exe` cross-compiled for x86_64-pc-windows-gnu with Homebrew mingw-w64 (9.4 MB; no rocksdb in the miner's closure).
Not done: no NVIDIA or AMD run of the bound kernels yet (the Windows package `~/Desktop/igneum-mine-test.zip` is the next step); NVRTC and runtime OpenCL rebuilds inside the workers; the block level for pruning proofs on a lane-only pow value; 8 vs 18 decimals; the VDF seeds and finality.

View file

@ -2,14 +2,14 @@
Fork: `vendor/igneum-node`, a git clone of `vendor/rusty-kaspa` at v2.1.0 commit `01b532e8` (22 Sep 2026). Every Igneum change is a commit on top of that base, one per subject, so `git log 01b532e8..HEAD` in the fork is the full list. This file is the reading guide: what each change touched, why, how risky it is, and what to do when upstream moves. `docs/fork-map.md` is the plan this implements; row ids there (a1 to f2) are cited below.
Status: devnet v0, mining layer only (3 Oct 2026). Finality, VDF seeds, the zkEVM and the proving layer are not in the fork yet.
Status: devnet v1, mining layer on the real header-bound lottery hash (3 Oct 2026, later the same day as v0): CPU and GPU miners, three workers (Metal, CUDA, OpenCL). Finality, VDF seeds, the zkEVM and the proving layer are not in the fork yet.
## The table
| File (vendor/igneum-node/) | What changed | Why | Risk | Upstream-merge note |
|---|---|---|---|---|
| `consensus/core/src/header.rs`, `consensus/core/src/hashing/header.rs` | `Header` gains `vote_key_hash: Hash` (32 bytes) as the last field; `new_finalized` takes it; `hash_override_nonce_time` writes it after `pruning_point` so the header hash and the PoW commit to it | Finality rule v2: vote weight is blue blocks per BLS vote key. The hash of the key rides in every header now so the weight table can be built from headers alone later (fork-map d) | Low by depth, high by spread: every header hash and every genesis hash moved | Any upstream change to `Header` or to the hashing order conflicts here. Re-derive the four genesis hashes after merging (`consensus/core/src/config/genesis.rs` tests print them). |
| `consensus/core/src/config/genesis.rs` | All four genesis hashes recomputed; devnet genesis has payload `igneum-devnet`, timestamp 2026-10-03T00:00Z, bits `0x1e020000` (2^23 expected hashes per block, approximate), nonce 0 | Hash field change above; a devnet that three CPU miners on one Mac can hold at about 1 block per second | Low | Mainnet, testnet and simnet genesis blocks are still Kaspa's content with new hashes. Replace them with Igneum genesis blocks before any public network. |
| `consensus/core/src/config/genesis.rs` | All four genesis hashes recomputed; devnet genesis has payload `igneum-devnet`, timestamp 2026-10-03T00:00Z, nonce 0; bits `0x1d100000` (2^28 expected hashes per block, hash `edc4fa84...fb07`) for the GPU devnet (RTX 5090 229 MH/s + M5 Max 45 MH/s + gfx1036 4 MH/s measured, about 1.04 blocks/s); earlier the same day `0x1e400000` (2^18) for three 6-thread CPU miners on the real hash (0.294 MH/s: 825 blocks in 641 s) and `0x1e020000` (2^23) for the kHeavyHash stub | A devnet the miners at hand hold near 1 block per second until the DAA takes over at 600 blocks | Low | Mainnet, testnet and simnet genesis blocks are still Kaspa's content with new hashes. Replace them with Igneum genesis blocks before any public network. Every bits change moves the devnet genesis hash (print it with the ignored test). |
| `consensus/core/src/errors/block.rs`, `consensus/src/pipeline/header_processor/pre_ghostdag_validation.rs` | `RuleError::MissingVoteKeyHash`; `check_vote_key_hash_present` rejects an all-zero `vote_key_hash` | Nodes only check presence until the finality layer exists | Low | Keep. The presence check becomes a key-registry check later. |
| `protocol/p2p/proto/p2p.proto`, `protocol/p2p/src/convert/{header,block,messages}.rs`, `protocol/flows/src/v10/request_headers.rs` | `BlockHeader` wire message gains `Hash voteKeyHash = 15`; converters copy it both ways | p2p round trip of the field | Low | Field number 15 must stay unique if upstream adds header fields. Protocol version is still Kaspa's 11; bump it when the fork gets its own peers. |
| `rpc/grpc/core/proto/rpc.proto`, `rpc/core/src/model/header.rs`, `rpc/core/src/model/optional/header.rs`, `rpc/core/src/model/verbosity.rs`, `rpc/core/src/convert/verbosity.rs`, `rpc/grpc/core/src/convert/{header,optional/header}.rs`, `rpc/service/src/converter/consensus.rs`, `consensus/client/src/header.rs` | `RpcHeader` and `RpcOptionalHeader` carry `vote_key_hash`; gRPC proto fields (`string voteKeyHash = 16` on the header, `optional string voteKeyHash = 15` on the optional header) and converters; wasm client header | RPC round trip, template to miner and block back | Low | Mechanical. Borsh wire for wRPC changed shape: old wRPC clients cannot decode headers. |
@ -20,11 +20,11 @@ Status: devnet v0, mining layer only (3 Oct 2026). Finality, VDF seeds, the zkEV
| `consensus/core/src/igneum.rs` (new), `consensus/core/src/lib.rs`, `consensus/core/src/constants.rs` | Emission constants and functions: 1,000,000,000 coins in year one, per-second rate halving every two years (`SUBSIDY_PER_SECOND_BY_PERIOD[i] = 3,168,808,781 >> i`, 33 periods), `block_subsidy(daa_score, bps)`, 30-day linear launch ramp from 10%, `proving_pool_share` 20% and `producer_share` 80%, `proving_pool_script_public_key` (OP_RETURN tagged `igneum-proving-pool-v0`), `POW_EPOCH_BLOCKS = 3,600` | The design's schedule, hard cap 4,000,000,000 (the geometric series sums to it; rounding leaves under 100 coins unminted), no emission treasury (fork-map b1, b2) | Medium: consensus money | Pure addition. Keep as the one source of the schedule. |
| `consensus/src/processes/coinbase.rs` | Kaspa's pre-deflationary phase, 426-month table and Crescendo rescaling removed; `CoinbaseManager::new(max_spk_len, max_payload_len, bps)`; `calc_block_subsidy` reads the period table; `expected_coinbase_transaction` pays 80% plus fees per rewarded block to its declared script and pools 20% of every subsidy (blues and reds) into one output to the proving pool script, placed after the blue outputs and before any red reward; tests rewritten | 80/20 split in consensus; the 20% is burned on devnet v0 and becomes the prover payout when the proving layer records prover sets (fork-map b3) | High: changes what every node accepts as a valid coinbase | Upstream edits to `coinbase.rs` will conflict. The payload format is unchanged (full subsidy in the payload, split derived from it), so Kaspa's payload parsing merges cleanly. |
| `consensus/src/consensus/services.rs`, `consensus/src/consensus/test_consensus.rs`, `consensus/src/pipeline/body_processor/body_validation_in_context.rs`, `consensus/src/processes/parents_builder.rs`, `consensus/src/processes/transaction_validator/tx_validation_in_isolation.rs` | Call sites of the new `CoinbaseManager` constructor; subsidy expectations in tests use `igneum::block_subsidy` | Wiring | Low | Mechanical. |
| `consensus/pow/src/igneum.rs` (new), `consensus/pow/src/lib.rs`, `consensus/pow/Cargo.toml`, `Cargo.lock` | `PowEngine` trait (`check_header(header, &EpochSeeds) -> (passed, pow)`), `HeavyHashEngine` stub (default), `IgneumEngine` behind feature `igneum-pow` calling the `igneum-pow` crate (`Epoch::memory_hard`, `Epoch::hash`), caching three `(epoch seed, day seed)` entries of program plus 256 MiB cache; `igneum-pow` as an optional path dependency `../../../../igneum-pow`; doc note on the existing `calc_block_level` (pruning proofs still use the stub for block levels) | The hash swap behind a trait so the devnet runs on the stub while the real path is wired (fork-map a1 to a3) | Medium | Pure addition in the pow crate. `State` (kHeavyHash) is untouched, so upstream pow changes merge. The path dependency must become a workspace or git dependency when the fork gets its own repository. |
| `consensus/src/pipeline/header_processor/processor.rs`, `consensus/src/pipeline/header_processor/pre_ghostdag_validation.rs` | PoW check moved from `validate_header_in_isolation` to after GHOSTDAG (`check_pow_and_calc_block_level(header, selected_parent)`); `epoch_seed` walks the selected-parent chain to the last block below the epoch's start DAA score (genesis for epoch 0) with a memo; `pow_engine: Arc<dyn PowEngine>` on the processor | The epoch seed is chain state, so PoW cannot be checked in isolation any more (fork-map a4). Temporary seed rule until the 10-minute VDF over a certified checkpoint exists | High: a header now reaches GHOSTDAG before its PoW is checked, so an attacker can make a node run GHOSTDAG on headers with bad nonces (bounded by the per-peer header rate; the stub engine ignores the seeds so devnet v0 is not exposed) | Upstream rarely touches this ordering, but any refactor of `process_header` conflicts. Pruning-proof validation (`processes/pruning_proof/validate.rs`) still uses the stub for block levels; thread seeds through it before enabling the real engine on a pruning network. |
| `consensus/pow/src/igneum.rs` (new), `consensus/pow/src/lib.rs`, `consensus/pow/Cargo.toml`, `Cargo.lock` | `PowEngine` trait (`check_header(header, &EpochSeeds) -> (passed, pow)`), `HeavyHashEngine` stub (default), `IgneumEngine` behind feature `igneum-pow` calling the `igneum-pow` crate in its header-bound form (`Epoch::from_seed_bytes`, `Epoch::pow_bound`): `H` = header hash with the nonce zeroed (timestamp kept), init words `seed_words_from_bytes("igneum-block/" \|\| H \|\| nonce_hi_le32)`, lane nonce = low 32 bits, pow value = lane hash in the top 64 bits with zero low bits; epoch seed bytes = the epoch block hash, day seed bytes = `"igneum-day/" \|\| day_le64`; caches three `(epoch seed, day seed)` entries of program plus 256 MiB cache and exposes them to the miner (`epoch_for`, `header_prehash`, `target64`); `igneum-pow` as an optional path dependency `../../../../igneum-pow`; doc note on the existing `calc_block_level` (pruning proofs still use the stub for block levels) | The hash swap behind a trait (fork-map a1 to a3); the binding closes the "lane hash does not absorb the header" gap of v0 (spec 01 O-1.9) | Medium | Pure addition in the pow crate. `State` (kHeavyHash) is untouched, so upstream pow changes merge. The path dependency must become a workspace or git dependency when the fork gets its own repository. |
| `consensus/src/pipeline/header_processor/processor.rs`, `consensus/src/pipeline/header_processor/pre_ghostdag_validation.rs` | PoW check moved from `validate_header_in_isolation` to after GHOSTDAG (`check_pow_and_calc_block_level(header, selected_parent)`); `epoch_seed` walks the selected-parent chain to the last block below the epoch's start DAA score (genesis for epoch 0) with a memo; `pow_engine: Arc<dyn PowEngine>` on the processor; one `info` line per accepted or rejected PoW naming the engine (`PoW accepted <hash> by igneum-lottery-v1-bound ...`) | The epoch seed is chain state, so PoW cannot be checked in isolation any more (fork-map a4). Temporary seed rule until the 10-minute VDF over a certified checkpoint exists | High: a header now reaches GHOSTDAG before its PoW is checked, so an attacker can make a node run GHOSTDAG on headers with bad nonces (bounded by the per-peer header rate; the stub engine ignores the seeds so devnet v0 is not exposed) | Upstream rarely touches this ordering, but any refactor of `process_header` conflicts. Pruning-proof validation (`processes/pruning_proof/validate.rs`) still uses the stub for block levels; thread seeds through it before enabling the real engine on a pruning network. |
| `consensus/Cargo.toml`, `kaspad/Cargo.toml` | Feature `igneum-pow` forwarded (`kaspad -> kaspa-consensus -> kaspa-pow`) | `cargo build -p kaspad --features igneum-pow` selects the real engine | Low | Keep. |
| `consensus/core/src/config/constants.rs` | Comment block only: the DAA constants kept at 1 BPS (sample every 4 blocks, 661 samples, 2,644-block window, min window 150 samples) and the two timestamp rules kept (132 s future tolerance in isolation, strictly above the sampled past median time of 27 samples in context) | Difficulty step verified rather than changed (fork-map c1, c2); the known gap (per-epoch hash-speed step vs a 44-minute window) is recorded there | None | Comment only. |
| `igneum/miner/` (new crate `igneum-miner`), `Cargo.toml` (workspace member), `Cargo.lock` | CPU devnet miner on `kaspa_pow::State` (the stub), sets `vote_key_hash` from a label; `watch` prints block counts, DAA, tips, peers, difficulty and sink per node and a blocks-per-second summary; `inspect` walks the selected chain and checks `vote_key_hash` equality across nodes and the 80/20 coinbase split | Kaspa ships no miner; the devnet needs one that follows the fork's own PoW crate | None to consensus | Internal tool. Switch it to `PowEngine` when the real engine is the default. |
| `igneum/miner/` (new crate `igneum-miner`), `Cargo.toml` (workspace member), `Cargo.lock` | Devnet miner: `--engine stub` (kHeavyHash, v0) or `--engine igneum-pow` (CPU warps through the node's own `IgneumEngine` cache, 64-bit nonces, random start per thread), `--worker <exe>` (GPU serve protocol: `job`/`found`/`done` lines, CPU re-check of every found nonce, one submit per template, `--exit-on-seed-change` exits 42 for ahead-of-time workers, `--worker-args`, `--status-secs`), `export-pack` (the pack for the node's current epoch and day plus `seeds.txt`), `bad-nonce` (submits an unmined nonce, expects a rejection); the epoch seed is derived as the node does it, walking from the sink over gRPC; `watch` and `inspect` as in v0 | Kaspa ships no miner; the devnet needs one that follows the fork's own PoW crate, and the GPU workers need a driver | None to consensus | Internal tool. Depends on `igneum-pow` by path and on `kaspa-pow` with the `igneum-pow` feature. Cross-compiles for `x86_64-pc-windows-gnu` with mingw-w64 (no rocksdb in its closure). |
## Decisions recorded as open
@ -32,8 +32,9 @@ Status: devnet v0, mining layer only (3 Oct 2026). Finality, VDF seeds, the zkEV
|---|---|---|
| 8 or 18 decimals | Kaspa's 8 (`SOMPI_PER_KASPA`), so one coin is 100,000,000 units and the cap is 4e17 units, inside u64 | The zkEVM side expects 18 decimals (wei). 18 decimals put the cap at 4e27, which does not fit u64, so the UTXO amount type, mass rules and every RPC amount would change. Decide with the execution engineer before the EVM bridge; a fixed 1e10 scaling at the bridge is the alternative. |
| Epoch seed | Hash of the last selected-chain block of the previous 3,600-block epoch; genesis for epoch 0 | The design uses a 10-minute class-group VDF over a certified checkpoint (bench-log, proto-vdf). The v0 rule is grindable in principle (a miner choosing which block ends an epoch) and needs the VDF and checkpoints to close. |
| Day seed for the 256 MiB cache | `header.timestamp / 86,400,000` | Timestamps are miner-chosen inside the two timestamp rules, so a day boundary can be straddled by a few blocks; harmless for a cache seed, but the exact rule is not final. |
| Lane hash to 256-bit target | Lane hash (64 bits) in the top 64 bits, cSHAKE of the header folded with the lane in the low 192 bits | The lane hash does not absorb the header (see the TODO in `consensus/pow/src/igneum.rs`): a miner could tabulate an epoch's 2^32 lane hashes once. The cryptographer owns the fix before the real engine is the default. |
| Day seed for the 256 MiB cache | `"igneum-day/" \|\| day_le64` with `day = header.timestamp / 86,400,000` | Timestamps are miner-chosen inside the two timestamp rules, so a day boundary can be straddled by a few blocks; harmless for a cache seed. Spec 01 O-1.10 proposes the first epoch seed of the day instead, which waits for the VDF schedule. |
| Lane hash to 256-bit target | Lane hash (64 bits) in the top 64 bits, low 192 bits zero; `pow <= target` is exactly `lane <= target >> 192` | Closed for the binding (3 Oct 2026, `igneum-pow/src/bind.rs`: the init words commit to the nonce-zeroed header hash and the high nonce word). Still open: the block level for pruning proofs reads `calc_level_from_pow` on a value whose low 192 bits are zero (`leading_zeros(lane)` shifted), and pruning-proof validation itself still uses the stub. |
| GPU workers and the hourly program | Metal recompiles at runtime; CUDA and OpenCL are built ahead of time per pack and are rebuilt by the launcher at each epoch or day change (miner exit 42) | NVRTC and a runtime OpenCL rebuild inside the worker would remove the rebuild gap (about 30 s for CUDA). |
| Proving pool payee | OP_RETURN burn tagged `igneum-proving-pool-v0` | Becomes a payout to the prover set of the proven block once the proving layer records prover sets (20 to 60 s behind the tip). |
| Finality and pruning depths | Kaspa's 12 h and 30 h at 1 BPS | Upper bounds. Live finality is the 30-s certified checkpoint; pruning depth must stay above the longest checkpoint gap. |
| PoW before or after GHOSTDAG | After | Needed for the chain-derived seed; costs GHOSTDAG work on invalid headers. A header-only seed (for example the VDF output carried in the header and verified against the checkpoint) would move it back. |

105
proto-cuda/WINDOWS-MINER.md Normal file
View file

@ -0,0 +1,105 @@
# Mining the Igneum devnet from a Windows PC
Written 3 October 2026 for the project lead's PC (RTX 5090 at about 229 MH/s on the memory-hard hash, Ryzen 9800X3D with the
gfx1036 integrated AMD GPU at about 4 MH/s; `docs/bench-log.md`). Two ways: the packaged test (one double-click), or by
hand from this repository. Both mine against node 1 on the Mac (LAN 192.168.68.64, gRPC 26610, p2p 26611), which runs
`kaspad --devnet` built with `--features igneum-pow`, so every block is checked with the real header-bound lottery hash.
## The packaged test (igneum-mine-test.zip)
Unzip anywhere, double-click `START-MINING.bat`. The package is built by `windows-miner/make-package.sh` on the Mac
from `windows-miner/` (START-MINING.bat, start-mining.ps1, README.txt, upload-log.bat), `proto-cuda/` (host.cu,
build.bat), `proto-opencl/` (host.c, build.bat) and a cross-compiled `igneum-miner.exe`
(`cargo build --release -p igneum-miner --target x86_64-pc-windows-gnu` with Homebrew mingw-w64, so no Rust on the PC).
What the launcher does: lists GPUs and tools; exports the pack for the node's current epoch and day
(`igneum-miner.exe export-pack`); builds `proto-cuda\igneum-bench-cuda-devnet.exe` and
`proto-opencl\igneum-bench-cl-devnet.exe` from it (cached by the pack's `seeds.txt`); starts `MINERS` (default 8)
`igneum-miner mine ... --worker <exe> --exit-on-seed-change` identities per GPU vendor, each with its own worker process
(about 1.3 GiB of GPU memory: 256 MiB cache plus 1 GiB dataset), its own vote key and payout address derived from the
PC name and the index (`--payout-label nvidia-<PC>-<i>`, the same on a rerun) and its own random 64-bit nonce start per
job; prints a status line per identity and a total per vendor every 30 s; uploads the logs every 60 s under
`<vendor>-<PC>-<index>`; on miner exit code 42 (the epoch or day seed changed) stops that vendor's identities,
re-exports, rebuilds and restarts them all, logging the time (the exe cannot be rebuilt while an instance runs it); on
any other exit restarts that identity after 10 s; Ctrl+C prints a summary and uploads once more. One shared worker per
vendor would need a multiplexing protocol and is not built; with 8 identities the 5090 holds about 10 GiB of worker
memory, the integrated AMD GPU takes the same from system RAM.
Settings at the top of START-MINING.bat: `NODE_HOST`, `NODE_PORT`, `PAYOUT_ADDRESS` (one address for every identity;
empty = derived per identity), `MINERS`, `NVIDIA_DEVICE` (CUDA index),
`CUDA_ARCH` (sm_120, or native), `AMD_VENDOR` (the OpenCL vendor string to match, default "Advanced Micro Devices", so
NVIDIA's OpenCL device is never picked for the AMD worker), `AMD_DEVICE` (flat index override), `USE_NVIDIA`, `USE_AMD`,
`VCVARS` and `VCVARS_VER` (14.30, the MSVC v143 toolset CUDA 12.8 accepts; Visual Studio 2026's own 14.5x crashes
cudafe++, see `build.bat`).
## By hand, from the repository
Prerequisites on the PC:
1. CUDA Toolkit 12.8 or newer (nvcc on PATH; it also ships `CL\cl.h` and `OpenCL.lib` for the AMD worker).
2. Visual Studio 2026 with the "MSVC v143 - VS 2022 C++ x64 build tools" component (toolset 14.30).
3. Rust, only if you build the miner on the PC: `winget install Rustlang.Rustup` then open a new prompt.
4. The repository, for example `git clone` of igneum-network/igneum into `C:\igneum` (the fork is `vendor\igneum-node`,
the hash crate `igneum-pow`; the miner depends on both by relative path, so keep the layout).
Open a plain command prompt and select the toolset first (every build command below assumes this):
```
"C:\Program Files\Microsoft Visual Studio\18\Community\VC\Auxiliary\Build\vcvarsall.bat" x64 -vcvars_ver=14.30
```
Build the miner (10 to 15 minutes the first time; `rand`, `ring`, `secp256k1-sys` and the gRPC stack build with MSVC;
`rocksdb` is not in the miner's dependency closure, so no CMake or clang is needed; a full `cargo build -p kaspad` on
Windows would need them and is not required for mining):
```
cd C:\igneum\vendor\igneum-node
cargo build --release -p igneum-miner
```
Export the pack for the node's current epoch and day, then build the workers from it:
```
cd C:\igneum\proto-cuda
..\vendor\igneum-node\target\release\igneum-miner.exe export-pack grpc://192.168.68.64:26610 packs\devnet
build.bat devnet sm_120 (adds -allow-unsupported-compiler and kernel_bound.cu itself)
cd ..\proto-opencl
build.bat devnet (MSVC, OpenCL.lib from %CUDA_PATH%)
```
Run (one window per worker; each miner starts every job at its own random 64-bit nonce):
```
cd C:\igneum\vendor\igneum-node
target\release\igneum-miner.exe mine grpc://192.168.68.64:26610 1 100000000 rtx5090 --worker ..\..\proto-cuda\igneum-bench-cuda-devnet.exe --worker-args "--device 0" --status-secs 30 --exit-on-seed-change
target\release\igneum-miner.exe mine grpc://192.168.68.64:26610 1 100000000 gfx1036 --worker ..\..\proto-opencl\igneum-bench-cl-devnet.exe --worker-args "--vendor Advanced_Micro_Devices" --status-secs 30 --exit-on-seed-change
```
`--exit-on-seed-change` makes the miner exit with code 42 when the epoch (3,600 DAA blocks) or the day changes, because
the CUDA and OpenCL workers are built ahead of time for one pack: export again, rebuild, restart (the launcher does
this). The Metal worker on the Mac recompiles at runtime instead. Every `found` nonce is re-checked on the CPU by the
miner before it is submitted; a `WORKER MISMATCH` line means the GPU and the CPU disagree and nothing is submitted.
## The worker protocol (shared by proto-metal, proto-cuda and proto-opencl)
```
stdin job <job_id> <header_prehash_hex 64> <target_hex 16> <nonce_start u64> <nonce_count u64> <epoch_seed_hex 64> <day_seed_hex>
quit
stdout ready <metal|cuda|opencl> <device> ...
found <job_id> <nonce u64> <hash_hex 16> every nonce with hash <= target (the 64-bit lane hash against target256 >> 192)
done <job_id> <hashes> <ms>
error <job_id> <text>
```
`nonce_start` is 32-aligned, `nonce_count` a multiple of 32; lane nonces are the low 32 bits, the high word enters the
init words `seed_words_from_bytes("igneum-block/" || prehash || nonce_hi_le32)` (`igneum-pow/src/bind.rs`). The CUDA
and OpenCL workers answer `error ... seed mismatch` for a job whose seeds are not the pack's.
## What was and was not tested on the Mac (3 October 2026)
Tested: the Metal worker against the devnet (5,636 blocks accepted in 300 s, 0 mismatches, an epoch change crossed
live); the OpenCL worker built with clang against Apple's OpenCL runtime (same bound hashes as Rust); the CUDA host
with `kernel_bound.cu` through the clang emulation shim (`emu/emu.sh igneum-devnet-emu`, bench PASS and a serve job
bit-exact with Rust, seed-mismatch error exercised); `igneum-miner.exe` cross-compiled for x86_64-pc-windows-gnu.
Not tested here: nvcc and cl.exe builds, the WMI GPU listing, vcvarsall import, winget, taskkill and Ctrl+C handling in
start-mining.ps1 (no PowerShell on this Mac), and running the .exe files on Windows.

View file

@ -31,12 +31,16 @@ if not exist "packs\%PACK%\kernel.cu" (
exit /b 1
)
rem Packs written by igneum-pow or igneum-miner export-pack carry kernel_bound.cu (the header-bound kernel for --serve).
set BOUND=
if exist "packs\%PACK%\kernel_bound.cu" set BOUND=-DIGNEUM_BOUND "packs\%PACK%\kernel_bound.cu"
nvcc --version
echo nvcc -O3 -std=c++17 -arch=%ARCH% -allow-unsupported-compiler -I "packs\%PACK%" -o "igneum-bench-cuda-%PACK%.exe" host.cu "packs\%PACK%\kernel.cu"
nvcc -O3 -std=c++17 -arch=%ARCH% -allow-unsupported-compiler -I "packs\%PACK%" -o "igneum-bench-cuda-%PACK%.exe" host.cu "packs\%PACK%\kernel.cu"
echo nvcc -O3 -std=c++17 -arch=%ARCH% -allow-unsupported-compiler -I "packs\%PACK%" -o "igneum-bench-cuda-%PACK%.exe" host.cu "packs\%PACK%\kernel.cu" %BOUND%
nvcc -O3 -std=c++17 -arch=%ARCH% -allow-unsupported-compiler -I "packs\%PACK%" -o "igneum-bench-cuda-%PACK%.exe" host.cu "packs\%PACK%\kernel.cu" %BOUND%
if errorlevel 1 (
echo build failed. The script already passes -allow-unsupported-compiler for Visual Studio versions newer than CUDA expects. If the errors are inside CUDA headers, install the "MSVC v143 - VS 2022 C++ x64 build tools" component from the Visual Studio Installer and build again, or install a newer CUDA Toolkit.
exit /b 1
)
echo built igneum-bench-cuda-%PACK%.exe
if not "%BOUND%"=="" echo with --serve support (kernel_bound.cu)
endlocal

View file

@ -19,8 +19,10 @@ if [ ! -f "packs/$PACK/kernel.cu" ]; then
exit 1
fi
BOUND=()
if [ -f "packs/$PACK/kernel_bound.cu" ]; then BOUND=(-DIGNEUM_BOUND "packs/$PACK/kernel_bound.cu"); fi
nvcc --version | tail -n 2
set -x
nvcc -O3 -std=c++17 -arch="$ARCH" -I "packs/$PACK" -o "igneum-bench-cuda-$PACK" host.cu "packs/$PACK/kernel.cu"
nvcc -O3 -std=c++17 -arch="$ARCH" -I "packs/$PACK" -o "igneum-bench-cuda-$PACK" host.cu "packs/$PACK/kernel.cu" "${BOUND[@]}"
set +x
echo "built ./igneum-bench-cuda-$PACK"

View file

@ -13,8 +13,14 @@ mkdir -p "$BUILD"
cp "$CUDA_DIR/host.cu" "$BUILD/host_emu.cpp"
# Rewrite only the launch syntax; everything else is the exact generated text.
sed -E 's/([A-Za-z_0-9]+)<<<([^,]+), ([^>]+)>>>\(/emu_launch(\1, \2, \3, /' "$CUDA_DIR/packs/$PACK/kernel.cu" > "$BUILD/kernel_emu.cpp"
# A pack with kernel_bound.cu (igneum-pow / igneum-miner export-pack) also gets the --serve worker compiled in.
BOUND=()
if [ -f "$CUDA_DIR/packs/$PACK/kernel_bound.cu" ]; then
sed -E 's/([A-Za-z_0-9]+)<<<([^,]+), ([^>]+)>>>\(/emu_launch(\1, \2, \3, /' "$CUDA_DIR/packs/$PACK/kernel_bound.cu" > "$BUILD/kernel_bound_emu.cpp"
BOUND=(-DIGNEUM_BOUND "$BUILD/kernel_bound_emu.cpp")
fi
CXX="${CXX:-c++}"
"$CXX" -std=c++17 -O2 -Wall -Wextra -I "$HERE" -I "$CUDA_DIR/packs/$PACK" \
-o "$BUILD/igneum-emu" "$BUILD/host_emu.cpp" "$BUILD/kernel_emu.cpp" "$HERE/shim.cpp" -pthread
-o "$BUILD/igneum-emu" "$BUILD/host_emu.cpp" "$BUILD/kernel_emu.cpp" "${BOUND[@]}" "$HERE/shim.cpp" -pthread
echo "compiled $BUILD/igneum-emu (CPU emulation, not a GPU build)"
exec "$BUILD/igneum-emu" "$@"

View file

@ -18,6 +18,7 @@
#include <chrono>
#include <string>
#include <vector>
#include <iostream>
#include "program.h"
#include "vectors.h"
@ -132,6 +133,7 @@ struct Options {
int blockWarps = 1;
bool sweep = false;
int device = 0;
bool serve = false; // --serve: GPU worker for igneum-miner --worker (jobs on stdin), 3 October 2026
};
static int packMib() { return (int)(((1ull << IGNEUM_DATASET_LOG2) * 4ull) >> 20); }
@ -144,7 +146,9 @@ static void usage() {
" --batch-log2 B nonces per batch = 2^B (default 24)\n"
" --batches N timed batches after one warm-up batch (default 5)\n"
" --block-warps W warps per thread block, 1..32 (default 1 = one warp per block, like the Metal run)\n"
" --device D CUDA device index (default 0)\n", packMib());
" --device D CUDA device index (default 0)\n"
" --serve GPU worker for igneum-miner --worker: reads \"job ...\" lines on stdin, prints found/done lines\n"
" (needs the pack's kernel_bound.cu compiled in: build.bat adds it when the pack has one)\n", packMib());
}
static bool isPow2(long long v) { return v > 0 && (v & (v - 1)) == 0; }
@ -163,6 +167,7 @@ static Options parseArgs(int argc, char** argv) {
else if (a == "--block-warps") next(o.blockWarps);
else if (a == "--device") next(o.device);
else if (a == "--sweep") o.sweep = true;
else if (a == "--serve") o.serve = true;
else if (a == "-h" || a == "--help") { usage(); std::exit(0); }
else { std::printf("unknown argument %s\n", argv[i]); usage(); std::exit(2); }
}
@ -386,6 +391,154 @@ static SizeResult runSize(const Options& o, int mib, uint64_t* dOut, uint32_t no
return r;
}
// ---------------------------------------------------------------------------------------------
// Serve mode: GPU worker for igneum-miner --worker (3 October 2026)
//
// Protocol, one line each. Only "found", "done" and "error" are parsed by the miner; every other line is logged.
// stdin: job <job_id> <header_prehash_hex 64> <target_hex 16> <nonce_start u64> <nonce_count u64> <epoch_seed_hex 64> <day_seed_hex>
// quit
// stdout: ready cuda <device>
// found <job_id> <nonce u64> <hash_hex 16> every nonce whose 64-bit hash is <= target
// done <job_id> <hashes> <ms> end of the job (wall ms)
// error <job_id> <text>
// The program is compiled ahead of time from the pack (no NVRTC), so this worker serves exactly one epoch seed and
// one day seed: the pack's. A job for other seeds is answered with an error naming both; re-export the pack with
// `igneum-miner export-pack <node> <dir>` and rebuild. The init words of a dispatch are
// seed_words_from_bytes("igneum-block/" || prehash || nonce_hi_le32), passed by value to igneum_hash_bound
// (kernel_bound.cu); the lane nonce is baseNonce + gid as in the bench kernel.
#ifdef IGNEUM_BOUND
struct IgneumInitWords { uint32_t w[8]; };
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif
static void emitLine(const char* s) { std::fputs(s, stdout); std::fputc('\n', stdout); std::fflush(stdout); }
// seed_words_from_bytes of igneum-pow/src/seed.rs: FNV-1a 64 with four salts, each finalised.
static void seedWordsFromBytes(const uint8_t* b, size_t n, uint32_t out[8]) {
for (uint64_t salt = 0; salt < 4; ++salt) {
uint64_t h = 0xcbf29ce484222325ull ^ (salt * 0x9E3779B97F4A7C15ull);
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
h ^= h >> 33; h *= 0xff51afd7ed558ccdull; h ^= h >> 33;
out[2 * salt] = (uint32_t)h;
out[2 * salt + 1] = (uint32_t)(h >> 32);
}
}
static bool unhexStr(const std::string& s, std::vector<uint8_t>& out) {
if (s.size() % 2) return false;
out.clear();
for (size_t i = 0; i < s.size(); i += 2) {
unsigned v = 0;
if (std::sscanf(s.substr(i, 2).c_str(), "%2x", &v) != 1) return false;
out.push_back((uint8_t)v);
}
return true;
}
static int runServe(const Options& o) {
#if !defined(IGNEUM_BOUND) || IGNEUM_DATASET_MODE != 1
(void)o;
emitLine("error 0 this binary was built without the pack's kernel_bound.cu (IGNEUM_BOUND) or from a closed-form pack; rebuild with build.bat from a memory-hard pack");
return 2;
#else
cudaDeviceProp prop;
std::memset(&prop, 0, sizeof(prop));
CUDA_CHECK(cudaGetDeviceProperties(&prop, o.device));
std::string devName = prop.name;
for (char& c : devName) if (c == ' ') c = '_';
// The pack's seeds: the program seed words and the day key words
static const uint32_t KEYW[8] = IGNEUM_KEY_INIT;
// Cache and dataset, once
if (!setupCache()) { emitLine("error 0 cache check failed (GPU cache differs from the host cache or the pack's FNV)"); return 1; }
const uint32_t words = 1u << IGNEUM_DATASET_LOG2;
const uint32_t mask = words - 1u;
uint32_t* dDs = nullptr;
CUDA_CHECK(cudaMalloc((void**)&dDs, (size_t)words * 4u));
CUDA_CHECK(igneum_launch_build(dDs, gCache, words / 16u));
CUDA_CHECK(cudaDeviceSynchronize());
const uint32_t batch = 1u << o.batchLog2;
uint64_t* dOut = nullptr;
CUDA_CHECK(cudaMalloc((void**)&dOut, (size_t)batch * sizeof(uint64_t)));
std::vector<uint64_t> hOut(batch);
int regs = 0, bps = 0;
igneum_hash_bound_info(&regs, &bps, (uint32_t)o.blockWarps);
std::printf("ready cuda %s pack %s dataset-log2 %d batch %u regs %d\n", devName.c_str(), IGNEUM_SEED_STRING, IGNEUM_DATASET_LOG2, batch, regs);
std::fflush(stdout);
std::string line;
while (std::getline(std::cin, line)) {
if (line == "quit") break;
std::vector<std::string> f;
{ size_t i = 0; while (i < line.size()) { while (i < line.size() && line[i] == ' ') ++i; size_t j = i; while (j < line.size() && line[j] != ' ') ++j; if (j > i) f.push_back(line.substr(i, j - i)); i = j; } }
if (f.empty()) continue;
if (f[0] != "job") { std::printf("info ignored: %s\n", line.c_str()); std::fflush(stdout); continue; }
std::string jobId = f.size() > 1 ? f[1] : "0";
if (f.size() < 8) { std::printf("error %s malformed job line (need 7 fields after job)\n", jobId.c_str()); std::fflush(stdout); continue; }
std::vector<uint8_t> prehash, epochSeed, daySeed;
uint64_t target = 0, nonceStart = 0, nonceCount = 0;
if (!unhexStr(f[2], prehash) || prehash.size() != 32 || std::sscanf(f[3].c_str(), "%llx", (unsigned long long*)&target) != 1 ||
std::sscanf(f[4].c_str(), "%llu", (unsigned long long*)&nonceStart) != 1 || std::sscanf(f[5].c_str(), "%llu", (unsigned long long*)&nonceCount) != 1 ||
!unhexStr(f[6], epochSeed) || epochSeed.size() != 32 || !unhexStr(f[7], daySeed)) {
std::printf("error %s bad field (prehash 64 hex, target 16 hex, nonce_start u64, nonce_count u64, epoch_seed 64 hex, day_seed hex)\n", jobId.c_str()); std::fflush(stdout); continue;
}
if (nonceCount == 0 || nonceCount % 32 != 0 || (nonceStart & 31) != 0) { std::printf("error %s nonce_start must be 32-aligned and nonce_count a non-zero multiple of 32\n", jobId.c_str()); std::fflush(stdout); continue; }
uint32_t sw[8], kw[8];
seedWordsFromBytes(epochSeed.data(), epochSeed.size(), sw);
seedWordsFromBytes(daySeed.data(), daySeed.size(), kw);
if (std::memcmp(sw, SEEDW, 32) != 0) {
std::printf("error %s epoch seed mismatch: this worker was built for pack \"%s\" (seed words %08x %08x ...), the job's epoch seed %s gives %08x %08x ...; run igneum-miner export-pack and rebuild\n",
jobId.c_str(), IGNEUM_SEED_STRING, SEEDW[0], SEEDW[1], f[6].substr(0, 16).c_str(), sw[0], sw[1]);
std::fflush(stdout); continue;
}
if (std::memcmp(kw, KEYW, 32) != 0) {
std::printf("error %s day seed mismatch: this worker's cache is for key %08x %08x ..., the job's day seed %s gives %08x %08x ...; run igneum-miner export-pack and rebuild\n",
jobId.c_str(), KEYW[0], KEYW[1], f[7].c_str(), kw[0], kw[1]);
std::fflush(stdout); continue;
}
double t0 = wallMs();
uint64_t remaining = nonceCount, hashes = 0;
uint32_t hi = (uint32_t)(nonceStart >> 32), lo = (uint32_t)nonceStart;
bool failed = false;
while (remaining > 0) {
uint64_t room = (uint64_t)(0xffffffffu - lo) + 1ull;
uint64_t chunk64 = remaining < batch ? remaining : batch;
if (chunk64 > room) chunk64 = room;
uint32_t chunk = (uint32_t)chunk64;
IgneumInitWords iw;
{
uint8_t b[49];
std::memcpy(b, "igneum-block/", 13);
std::memcpy(b + 13, prehash.data(), 32);
b[45] = (uint8_t)hi; b[46] = (uint8_t)(hi >> 8); b[47] = (uint8_t)(hi >> 16); b[48] = (uint8_t)(hi >> 24);
seedWordsFromBytes(b, 49, iw.w);
}
cudaError_t e = igneum_launch_hash_bound(dDs, dOut, lo, mask, iw, chunk, (uint32_t)o.blockWarps);
if (e == cudaSuccess) e = cudaDeviceSynchronize();
if (e == cudaSuccess) e = cudaMemcpy(hOut.data(), dOut, (size_t)chunk * sizeof(uint64_t), cudaMemcpyDeviceToHost);
if (e != cudaSuccess) { std::printf("error %s dispatch failed: %s\n", jobId.c_str(), cudaGetErrorString(e)); std::fflush(stdout); failed = true; break; }
for (uint32_t i = 0; i < chunk; ++i) if (hOut[i] <= target) {
uint64_t nonce = ((uint64_t)hi << 32) | (uint64_t)(uint32_t)(lo + i);
std::printf("found %s %llu %016llx\n", jobId.c_str(), (unsigned long long)nonce, (unsigned long long)hOut[i]);
}
std::fflush(stdout);
hashes += chunk;
remaining -= chunk;
if (chunk64 == room) { hi += 1u; lo = 0u; } else lo += chunk;
}
if (failed) continue;
std::printf("done %s %llu %.2f\n", jobId.c_str(), (unsigned long long)hashes, wallMs() - t0);
std::fflush(stdout);
}
cudaFree(dOut);
cudaFree(dDs);
cudaFree(gCache);
return 0;
#endif
}
// ---------------------------------------------------------------------------------------------
// Main
@ -398,6 +551,10 @@ int main(int argc, char** argv) {
if (count == 0) { std::printf("FAIL: no CUDA device\n"); return 2; }
if (o.device < 0 || o.device >= count) { std::printf("FAIL: device %d out of range (%d devices)\n", o.device, count); return 2; }
CUDA_CHECK(cudaSetDevice(o.device));
if (o.serve) {
if (o.batchLog2 == 24) { Options s2 = o; s2.batchLog2 = 22; return runServe(s2); } // 2^22 nonces per dispatch by default
return runServe(o);
}
cudaDeviceProp prop;
std::memset(&prop, 0, sizeof(prop));

View file

@ -0,0 +1,5 @@
Igneum devnet mining test. Double-click START-MINING.bat; it finds your GPUs, builds the NVIDIA (CUDA) and AMD (OpenCL) workers on the first run, and mines against the Mac node at 192.168.68.64 port 26610 (gRPC; p2p is 26611). Change the address, port, payout address or device choice at the top of START-MINING.bat.
It needs CUDA Toolkit 12.8 or newer and Visual Studio with the MSVC v143 (14.30) x64 component; the AMD worker uses the OpenCL headers the CUDA Toolkit ships. The miner (igneum-miner.exe) is prebuilt, no Rust needed.
MINERS=8 (top of the bat) starts 8 miner identities per GPU vendor, each with its own worker (about 1.3 GiB of GPU memory each), its own vote key and payout address derived from the PC name and the index (the same on a rerun), and its own random 64-bit nonce start per job, so no two miners repeat work; set MINERS=1 for a single-identity run. Logs are written next to the bat as nvidia-<index>-<stamp>.log, amd-<index>-<stamp>.log and launcher-<stamp>.log, and uploaded every 60 seconds under the labels nvidia-<PC name>-<index> and amd-<PC name>-<index>, so the log is already with us.
Every hour the program changes (epoch of 3,600 blocks) and once a day the cache seed changes: the miner stops with code 42, the launcher re-exports the pack from the node, rebuilds the worker (about 30 s for CUDA, a few seconds for OpenCL) and restarts it; a worker that dies for any other reason is restarted after 10 s. A status line per worker prints every 30 s.
Stop it in the morning with Ctrl+C in the window: the miners and workers are killed, a summary prints, the logs are uploaded one last time and the window stays open. The Mac side restarts with: cd ~/Projects/igneum/vendor/igneum-node && caffeinate -dims ./target/release/kaspad --devnet --nodnsseed --disable-upnp --enable-unsynced-mining --appdir=/tmp/igneum-devnet/node1 --rpclisten=0.0.0.0:26610 --listen=0.0.0.0:26611 --nologfiles >> /tmp/igneum-devnet/node1.log 2>&1 &

View file

@ -0,0 +1,40 @@
@echo off
rem Igneum devnet mining test for Windows. Double-click this file.
rem It finds the GPUs in this PC, builds the matching GPU workers on first run (NVIDIA: CUDA, AMD: OpenCL), and runs
rem one igneum-miner per worker against the Mac node. Logs land next to this file and are uploaded every 60 seconds.
rem Stop with Ctrl+C; the window stays open at the end.
rem ---- settings: edit these lines only ----------------------------------------------------------
set "NODE_HOST=192.168.68.64"
set "NODE_PORT=26610"
rem Devnet payout address (igneumdev:...). Leave empty for the miner's fixed test address.
set "PAYOUT_ADDRESS="
rem NVIDIA: CUDA device index (0 = first card) and the nvcc architecture (sm_120 = RTX 50 series; "native" lets nvcc pick).
set "NVIDIA_DEVICE=0"
set "CUDA_ARCH=sm_120"
rem AMD: by default the first OpenCL GPU whose vendor string contains AMD_VENDOR is used, never an NVIDIA device.
rem Set AMD_DEVICE to a flat device index from the worker's --list output to override.
set "AMD_VENDOR=Advanced Micro Devices"
set "AMD_DEVICE="
rem Miner identities per GPU vendor. Each identity runs its own worker process (about 1.3 GiB of GPU memory each) with its
rem own vote key and payout address derived from this PC's name and the index, and its own random nonce start.
set "MINERS=8"
rem Set to 0 to skip a vendor even if the hardware is present.
set "USE_NVIDIA=1"
set "USE_AMD=1"
rem Visual Studio: the MSVC toolset CUDA 12.8 accepts (14.30 = MSVC v143). Adjust the path if Visual Studio is elsewhere.
set "VCVARS=C:\Program Files\Microsoft Visual Studio\18\Community\VC\Auxiliary\Build\vcvarsall.bat"
set "VCVARS_VER=14.30"
rem -------------------------------------------------------------------------------------------------
cd /d "%~dp0"
where powershell >nul 2>nul
if errorlevel 1 (
echo PowerShell was not found. Windows 10 and 11 ship it; nothing else can be done here.
pause
exit /b 1
)
powershell -NoProfile -ExecutionPolicy Bypass -File "%~dp0start-mining.ps1"
echo.
echo The mining test has stopped. The log files are next to this file and were uploaded if the upload tool is present.
pause

View file

@ -0,0 +1,26 @@
#!/usr/bin/env bash
# Builds igneum-mine-test.zip (one double-click mining test for a Windows PC) on the Mac.
# Needs: a cross-compiled igneum-miner.exe (cargo build --release -p igneum-miner --target x86_64-pc-windows-gnu in
# vendor/igneum-node, with Homebrew mingw-w64) and the sources in proto-cuda/, proto-opencl/ and this directory.
# Usage: windows-miner/make-package.sh [output zip, default ~/Desktop/igneum-mine-test.zip]
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
ROOT="$(cd "$HERE/../.." && pwd)"
OUT="${1:-$HOME/Desktop/igneum-mine-test.zip}"
EXE="$ROOT/vendor/igneum-node/target/x86_64-pc-windows-gnu/release/igneum-miner.exe"
[ -f "$EXE" ] || { echo "no $EXE; cross-compile the miner first" >&2; exit 1; }
STAGE="$(mktemp -d)/igneum-mine-test"
mkdir -p "$STAGE/proto-cuda/packs" "$STAGE/proto-opencl"
cp "$HERE/START-MINING.bat" "$HERE/start-mining.ps1" "$HERE/README.txt" "$HERE/upload-log.bat" "$HERE/UPLOAD.md" "$STAGE/"
cp "$ROOT/proto-cuda/WINDOWS-MINER.md" "$STAGE/"
cp "$EXE" "$STAGE/igneum-miner.exe"
cp "$ROOT/proto-cuda/host.cu" "$ROOT/proto-cuda/build.bat" "$ROOT/proto-cuda/README.md" "$STAGE/proto-cuda/"
cp "$ROOT/proto-opencl/host.c" "$ROOT/proto-opencl/build.bat" "$ROOT/proto-opencl/README.md" "$ROOT/proto-opencl/WAVEFRONT.md" "$STAGE/proto-opencl/"
# CRLF for the files Windows tools read as text
for f in "$STAGE/START-MINING.bat" "$STAGE/README.txt" "$STAGE/proto-cuda/build.bat" "$STAGE/proto-opencl/build.bat"; do
perl -pi -e 's/\r?\n/\r\n/' "$f"
done
rm -f "$OUT"
(cd "$(dirname "$STAGE")" && zip -qr "$OUT" "$(basename "$STAGE")")
ls -la "$OUT"
unzip -l "$OUT" | tail -n +4 | awk '{print $4}' | sed '/^$/d'

View file

@ -0,0 +1,295 @@
# Igneum devnet mining test, Windows launcher. Started by START-MINING.bat, which sets the settings as environment
# variables (NODE_HOST, NODE_PORT, PAYOUT_ADDRESS, NVIDIA_DEVICE, CUDA_ARCH, AMD_VENDOR, AMD_DEVICE, USE_NVIDIA, USE_AMD,
# VCVARS, VCVARS_VER). 3 October 2026.
#
# What it does, in order:
# 1. Lists the GPUs and the tools (nvcc, MSVC through vcvarsall, an OpenCL SDK) and prints which workers will start.
# 2. Makes sure igneum-miner.exe exists (shipped prebuilt; built with cargo only if missing).
# 3. Exports the pack for the node's current epoch and day (igneum-miner export-pack) and builds the workers from it
# on first run; the build is cached by the pack's seeds, so later starts are instant.
# 4. Starts MINERS igneum-miner identities per vendor (nvidia, amd), each with its own worker process (256 MiB cache
# plus 1 GiB dataset of GPU memory per instance), its own deterministic vote key and payout address
# (<vendor>-<PC name>-<index>), and its own random 64-bit nonce start per job. Logs: <vendor>-<index>-<stamp>.log
# (no index when MINERS=1), uploaded under the label <vendor>-<PC name>-<index>.
# 5. Every 30 s prints a status line per worker; every 60 s uploads the logs with upload-log.bat (if present);
# restarts a miner that exits (code 42 = the epoch or day seed changed: re-export, rebuild, restart, with the
# rebuild time logged); Ctrl+C stops everything and prints a summary.
$ErrorActionPreference = 'Continue'
$root = Split-Path -Parent $MyInvocation.MyCommand.Path
Set-Location $root
$nodeHost = if ($env:NODE_HOST) { $env:NODE_HOST } else { '192.168.68.64' }
$nodePort = if ($env:NODE_PORT) { $env:NODE_PORT } else { '26610' }
$nodeUrl = "grpc://${nodeHost}:${nodePort}"
$payout = $env:PAYOUT_ADDRESS
$nvidiaDevice = if ($env:NVIDIA_DEVICE) { $env:NVIDIA_DEVICE } else { '0' }
$cudaArch = if ($env:CUDA_ARCH) { $env:CUDA_ARCH } else { 'sm_120' }
$amdVendor = if ($env:AMD_VENDOR) { $env:AMD_VENDOR } else { 'Advanced Micro Devices' }
$amdDevice = $env:AMD_DEVICE
$useNvidia = ($env:USE_NVIDIA -ne '0')
$useAmd = ($env:USE_AMD -ne '0')
$vcvars = if ($env:VCVARS) { $env:VCVARS } else { 'C:\Program Files\Microsoft Visual Studio\18\Community\VC\Auxiliary\Build\vcvarsall.bat' }
$vcvarsVer = if ($env:VCVARS_VER) { $env:VCVARS_VER } else { '14.30' }
$minersPerVendor = 1
if ($env:MINERS -and [int]$env:MINERS -ge 1) { $minersPerVendor = [int]$env:MINERS }
$stamp = Get-Date -Format 'yyyyMMdd-HHmmss'
$machine = $env:COMPUTERNAME
$launcherLog = Join-Path $root "launcher-$stamp.log"
$uploader = Join-Path $root 'upload-log.bat'
$started = Get-Date
function Log([string]$msg) {
$line = "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') $msg"
Write-Host $line
Add-Content -Path $launcherLog -Value $line
}
Log "igneum-mine-test launcher $machine node $nodeUrl (run $stamp)"
# ---- 1. hardware and tools ---------------------------------------------------------------------------------------
$gpuNames = @()
try { $gpuNames = @(Get-CimInstance Win32_VideoController | ForEach-Object { $_.Name }) } catch { Log "could not list GPUs through WMI: $_" }
Log ("GPUs found: " + ($(if ($gpuNames.Count) { $gpuNames -join '; ' } else { 'none reported' })))
$hasNvidia = ($gpuNames | Where-Object { $_ -match 'NVIDIA' }).Count -gt 0
$hasAmd = ($gpuNames | Where-Object { $_ -match 'AMD|Radeon' }).Count -gt 0
function Import-VcVars {
if (-not (Test-Path $vcvars)) { return $false }
$out = cmd /c "`"$vcvars`" x64 -vcvars_ver=$vcvarsVer >nul 2>&1 && set" 2>$null
if (-not $out) { return $false }
foreach ($l in $out) {
$i = $l.IndexOf('=')
if ($i -gt 0) { [Environment]::SetEnvironmentVariable($l.Substring(0, $i), $l.Substring($i + 1), 'Process') }
}
return [bool](Get-Command cl.exe -ErrorAction SilentlyContinue)
}
$haveCl = [bool](Get-Command cl.exe -ErrorAction SilentlyContinue)
if (-not $haveCl) { $haveCl = Import-VcVars }
$haveNvcc = [bool](Get-Command nvcc.exe -ErrorAction SilentlyContinue)
$openclSdk = $null
if ($env:OPENCL_SDK -and (Test-Path (Join-Path $env:OPENCL_SDK 'include\CL\cl.h'))) { $openclSdk = $env:OPENCL_SDK }
elseif ($env:CUDA_PATH -and (Test-Path (Join-Path $env:CUDA_PATH 'include\CL\cl.h'))) { $openclSdk = $env:CUDA_PATH }
elseif ($env:OCL_ROOT -and (Test-Path (Join-Path $env:OCL_ROOT 'include\CL\cl.h'))) { $openclSdk = $env:OCL_ROOT }
Log ("tools: nvcc " + $(if ($haveNvcc) { 'found' } else { 'MISSING (install CUDA Toolkit 12.8 or newer)' }) +
", MSVC $vcvarsVer " + $(if ($haveCl) { 'found' } else { "MISSING (install Visual Studio with the MSVC v143 x64 component; looked at $vcvars)" }) +
", OpenCL SDK " + $(if ($openclSdk) { "found at $openclSdk" } else { 'MISSING (the CUDA Toolkit ships one, or set OPENCL_SDK)' }))
$plan = @()
$skipped = @()
if ($hasNvidia -and $useNvidia) {
if ($haveNvcc -and $haveCl) { $plan += 'nvidia' } else { $skipped += "nvidia: GPU present but " + $(if (-not $haveNvcc) { 'nvcc' } else { 'MSVC' }) + " is missing" }
} elseif ($hasNvidia) { $skipped += 'nvidia: disabled by USE_NVIDIA=0' } else { $skipped += 'nvidia: no NVIDIA GPU reported' }
if ($hasAmd -and $useAmd) {
if ($haveCl -and $openclSdk) { $plan += 'amd' } else { $skipped += "amd: GPU present but " + $(if (-not $haveCl) { 'MSVC' } else { 'an OpenCL SDK' }) + " is missing" }
} elseif ($hasAmd) { $skipped += 'amd: disabled by USE_AMD=0' } else { $skipped += 'amd: no AMD GPU reported' }
Log ("workers to start: " + $(if ($plan.Count) { $plan -join ', ' } else { 'none' }))
foreach ($s in $skipped) { Log "skipped $s" }
if ($plan.Count -eq 0) { Log 'nothing to run. Fix the missing tool above and start again.'; exit 1 }
# ---- 2. the miner ----------------------------------------------------------------------------------------------------
$miner = Join-Path $root 'igneum-miner.exe'
if (-not (Test-Path $miner)) {
Log 'igneum-miner.exe is not in the package; building it from src\igneum-node (needs Rust)'
if (-not (Get-Command cargo -ErrorAction SilentlyContinue)) {
Log 'cargo not found: installing rustup with winget (silent)'
winget install --id Rustlang.Rustup -e --silent --accept-package-agreements --accept-source-agreements | Out-Null
$env:PATH = "$env:USERPROFILE\.cargo\bin;$env:PATH"
}
if (-not (Test-Path (Join-Path $root 'src\igneum-node\Cargo.toml'))) { Log 'no src\igneum-node in the package either; see WINDOWS-MINER.md'; exit 1 }
Push-Location (Join-Path $root 'src\igneum-node')
cargo build --release -p igneum-miner 2>&1 | ForEach-Object { Add-Content -Path $launcherLog -Value $_ }
Pop-Location
Copy-Item (Join-Path $root 'src\igneum-node\target\release\igneum-miner.exe') $miner -ErrorAction SilentlyContinue
if (-not (Test-Path $miner)) { Log 'the miner did not build; see the launcher log'; exit 1 }
}
try { $tnc = Test-NetConnection -ComputerName $nodeHost -Port ([int]$nodePort) -WarningAction SilentlyContinue -InformationLevel Quiet } catch { $tnc = $false }
if (-not $tnc) { Log "cannot reach the node at ${nodeHost}:${nodePort}. Is the Mac node running and on the same network?"; exit 1 }
Log "node reachable at ${nodeHost}:${nodePort}"
# ---- 3. pack and worker builds -------------------------------------------------------------------------------------------
$packDir = Join-Path $root 'proto-cuda\packs\devnet'
function Export-Pack {
$t = Get-Date
New-Item -ItemType Directory -Force -Path $packDir | Out-Null
$out = & $miner export-pack $nodeUrl $packDir 2>&1
$out | ForEach-Object { Add-Content -Path $launcherLog -Value " export-pack: $_" }
$seeds = Join-Path $packDir 'seeds.txt'
if (-not (Test-Path $seeds)) { Log 'export-pack wrote no seeds.txt (is the node reachable?)'; return $null }
$txt = (Get-Content $seeds) -join ' | '
Log ("pack exported in {0:N1} s: {1}" -f ((Get-Date) - $t).TotalSeconds, $txt)
return (Get-Content $seeds -Raw)
}
function Build-Worker([string]$name, [string]$seeds) {
if ($name -eq 'nvidia') { $exe = Join-Path $root 'proto-cuda\igneum-bench-cuda-devnet.exe'; $dir = Join-Path $root 'proto-cuda'; $cmd = "build.bat devnet $cudaArch" }
else { $exe = Join-Path $root 'proto-opencl\igneum-bench-cl-devnet.exe'; $dir = Join-Path $root 'proto-opencl'; $cmd = 'build.bat devnet' }
$stampFile = "$exe.seeds"
if ((Test-Path $exe) -and (Test-Path $stampFile) -and ((Get-Content $stampFile -Raw) -eq $seeds)) {
Log "$name worker already built for this epoch and day (cached)"
return $exe
}
$t = Get-Date
Log "building the $name worker from the pack ($cmd)"
Push-Location $dir
$out = cmd /c $cmd 2>&1
$rc = $LASTEXITCODE
Pop-Location
$out | ForEach-Object { Add-Content -Path $launcherLog -Value " $name build: $_" }
if ($rc -ne 0 -or -not (Test-Path $exe)) { Log "$name worker build FAILED (exit $rc); the full output is in $launcherLog"; return $null }
Set-Content -Path $stampFile -Value $seeds -NoNewline
Log ("{0} worker built in {1:N1} s ({2})" -f $name, ((Get-Date) - $t).TotalSeconds, $exe)
return $exe
}
$seeds = Export-Pack
if (-not $seeds) { exit 1 }
# ---- 4. start miners ------------------------------------------------------------------------------------------------------------
# $vendors[name] = @{ exe; instances = @(...) }; an instance = @{ vendor; index; label; log; runId; proc; restarts; startedAt }
$vendors = @{}
foreach ($name in $plan) {
$exe = Build-Worker $name $seeds
if (-not $exe) { Log "$name skipped: no worker binary"; continue }
$instances = @()
for ($i = 1; $i -le $minersPerVendor; $i++) {
$suffix = if ($minersPerVendor -gt 1) { "-$i" } else { '' }
$instances += @{ vendor = $name; index = $i; label = "$name-$machine$suffix"; log = (Join-Path $root "$name$suffix-$stamp.log"); runId = "$name-$machine$suffix-$stamp"; proc = $null; restarts = 0; startedAt = $null }
}
$vendors[$name] = @{ name = $name; exe = $exe; instances = $instances; rebuilds = 0 }
}
if ($vendors.Count -eq 0) { Log 'no worker could be built; see the launcher log'; exit 1 }
Log ("identities per vendor: $minersPerVendor (each identity runs its own worker: about 1.3 GiB of GPU memory per instance)")
function Start-Miner($v, $inst) {
# --worker-args takes ONE argument (quoted); the miner splits it on spaces and turns underscores into spaces
$margs = @('mine', $nodeUrl, '1', '100000000', $inst.label, '--worker', "`"$($v.exe)`"", '--status-secs', '30', '--exit-on-seed-change')
if ($payout) { $margs += @('--address', $payout) } else { $margs += @('--payout-label', $inst.label) }
if ($v.name -eq 'nvidia') { $margs += @('--worker-args', "`"--device $nvidiaDevice`"") }
else {
if ($amdDevice) { $margs += @('--worker-args', "`"--device $amdDevice`"") }
else { $margs += @('--worker-args', ('"--vendor ' + $amdVendor.Replace(' ', '_') + '"')) }
}
if (-not (Test-Path $inst.log)) { Set-Content -Path $inst.log -Value "igneum-mine-test $($inst.label) $machine $(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') node $nodeUrl run $($inst.runId)" }
$cmdLine = "`"$miner`" " + ($margs -join ' ') + " >> `"$($inst.log)`" 2>&1"
$inst.proc = Start-Process -FilePath 'cmd.exe' -ArgumentList @('/c', $cmdLine) -NoNewWindow -PassThru
$inst.startedAt = Get-Date
Log "$($inst.label) started (pid $($inst.proc.Id)): $cmdLine"
}
function Stop-Vendor($v) {
foreach ($inst in $v.instances) {
if ($inst.proc -and -not $inst.proc.HasExited) { cmd /c "taskkill /T /F /PID $($inst.proc.Id) >nul 2>&1" | Out-Null }
}
}
foreach ($v in $vendors.Values) { foreach ($inst in $v.instances) { Start-Miner $v $inst; Start-Sleep -Milliseconds 500 } }
Start-Sleep -Seconds 10
foreach ($v in $vendors.Values) {
$first = $v.instances[0]
$ready = Select-String -Path $first.log -Pattern 'worker: ready' | Select-Object -Last 1
if ($ready) { Log "$($v.name): $($ready.Line)" } else { Log "$($v.name): no ready line yet (see $($first.log))" }
}
# ---- 5. loop: status, uploads, restarts, Ctrl+C --------------------------------------------------------------------------
function Upload-Logs {
if (-not (Test-Path $uploader)) { return }
foreach ($v in $vendors.Values) {
foreach ($inst in $v.instances) {
$o = cmd /c "`"$uploader`" `"$($inst.log)`" $($inst.label) $($inst.runId)" 2>&1
Add-Content -Path $launcherLog -Value (" upload " + $inst.label + ": " + ($o -join ' '))
}
}
$o = cmd /c "`"$uploader`" `"$launcherLog`" launcher-$machine launcher-$machine-$stamp" 2>&1
}
function Instance-Stats($inst) {
$accepted = (Select-String -Path $inst.log -Pattern 'ACCEPTED block').Count
$last = Select-String -Path $inst.log -Pattern 'STATUS' | Select-Object -Last 1
$wall = 0.0; $age = 'n/a'
if ($last -and $last.Line -match 'hash=([0-9.]+) MH/s wall \(([0-9.]+) MH/s inside jobs\) template_age=([0-9.]+)s') { $wall = [double]$Matches[1]; $age = $Matches[3] + 's' }
return @{ accepted = $accepted; wall = $wall; age = $age }
}
function Print-Status {
$up = ((Get-Date) - $started).ToString('hh\:mm\:ss')
foreach ($v in $vendors.Values) {
$total = 0.0; $acc = 0
foreach ($inst in $v.instances) {
$st = Instance-Stats $inst
$total += $st.wall; $acc += $st.accepted
Log ("{0}: accepted {1}, {2:N2} MH/s, template age {3}, restarts {4}" -f $inst.label, $st.accepted, $st.wall, $st.age, $inst.restarts)
}
Log ("{0} TOTAL: {1} identities, accepted {2} blocks, {3:N2} MH/s, rebuilds {4}, up {5}" -f $v.name, $v.instances.Count, $acc, $total, $v.rebuilds, $up)
}
}
function Stop-All { foreach ($v in $vendors.Values) { Stop-Vendor $v } }
[Console]::TreatControlCAsInput = $true
Log 'running. Press Ctrl+C to stop (the summary prints and the logs are uploaded one last time).'
$lastStatus = Get-Date
$lastUpload = Get-Date
$stop = $false
try {
while (-not $stop) {
Start-Sleep -Milliseconds 1000
while ([Console]::KeyAvailable) {
$k = [Console]::ReadKey($true)
if ($k.Key -eq 'C' -and ($k.Modifiers -band [ConsoleModifiers]::Control)) { $stop = $true }
}
if ($stop) { break }
foreach ($v in @($vendors.Values)) {
$seedChange = $false
foreach ($inst in $v.instances) {
if ($inst.proc -and $inst.proc.HasExited -and $inst.proc.ExitCode -eq 42) { $seedChange = $true }
}
if ($seedChange) {
# The whole vendor group restarts: the worker exe cannot be rebuilt while another identity still runs it
$v.rebuilds += 1
Log "$($v.name): the epoch or day seed changed (miner exit 42). Stopping its $($v.instances.Count) identities, re-exporting the pack, rebuilding the worker."
Stop-Vendor $v
$t = Get-Date
$seeds = Export-Pack
if ($seeds) {
$exe = Build-Worker $v.name $seeds
if ($exe) { $v.exe = $exe }
}
Log ("{0}: rebuild for the new seeds took {1:N1} s in total" -f $v.name, ((Get-Date) - $t).TotalSeconds)
foreach ($inst in $v.instances) { Start-Miner $v $inst; Start-Sleep -Milliseconds 500 }
continue
}
foreach ($inst in $v.instances) {
if ($inst.proc -and $inst.proc.HasExited) {
$inst.restarts += 1
Log "$($inst.label): miner exited unexpectedly with code $($inst.proc.ExitCode); restarting in 10 s (restart $($inst.restarts))"
Start-Sleep -Seconds 10
Start-Miner $v $inst
}
}
}
if (((Get-Date) - $lastStatus).TotalSeconds -ge 30) { Print-Status; $lastStatus = Get-Date }
if (((Get-Date) - $lastUpload).TotalSeconds -ge 60) { Upload-Logs; $lastUpload = Get-Date }
}
} finally {
[Console]::TreatControlCAsInput = $false
Log 'stopping the miners and workers'
Stop-All
Start-Sleep -Seconds 2
Log ('SUMMARY after ' + ((Get-Date) - $started).ToString('hh\:mm\:ss'))
Print-Status
Upload-Logs
Log 'logs uploaded (if upload-log.bat is present); they are also next to START-MINING.bat'
}

View file

@ -17,6 +17,7 @@ struct Options {
var verifyWarps = 3
var dumpDir: String? = nil
var exportPack: String? = nil // write a CUDA program pack for --seed into this directory and exit
var serve = false // --serve: GPU worker for igneum-miner (jobs on stdin, results on stdout), 3 October 2026
// Hardening tests (added 3 October 2026). Any of these runs instead of the bench.
var fuzz: Int? = nil // --fuzz N: N random programs, GPU vs CPU on 4 random warps each
var fuzzSeed = "igneum-fuzz-2026-10-03"
@ -52,6 +53,7 @@ func parseArgs() -> Options {
case "--verify-warps": o.verifyWarps = Int(take()) ?? o.verifyWarps
case "--dump": o.dumpDir = take()
case "--export-pack": o.exportPack = take()
case "--serve": o.serve = true
case "--fuzz": o.fuzz = Int(take()) ?? 200
case "--fuzz-seed": o.fuzzSeed = take()
case "--edge": o.edge = true
@ -70,6 +72,7 @@ func parseArgs() -> Options {
[--load-weight W] generator lever (a): percent weight of the load op (default 25)
[--wide-frac P] generator lever (b): percent of loads emitted as warp-coalesced 128-byte loads (default 0)
[--export-pack <dir>] write the CUDA + OpenCL program pack for --seed, then exit
[--serve] GPU worker for igneum-miner --worker: reads "job ..." lines on stdin (see runServe)
hardening tests (run instead of the bench; several may be combined; exit 0 only if all pass):
[--fuzz N [--fuzz-seed <string>]] N random programs, GPU vs CPU, 4 random warps each,
dataset size drawn from 64 MiB, 256 MiB, 1 GiB
@ -118,18 +121,20 @@ func parseArgs() -> Options {
return x
}
// 32-byte seed (8 x uint32) from a string: FNV-1a 64 with four salts, each finalised.
func seedWords(_ s: String) -> [UInt32] {
// 32-byte seed (8 x uint32) from bytes: FNV-1a 64 with four salts, each finalised (seed_words_from_bytes in igneum-pow).
func seedWordsBytes(_ bytes: [UInt8]) -> [UInt32] {
var words = [UInt32]()
for salt in 0..<4 {
var h: UInt64 = 0xcbf29ce484222325 ^ (UInt64(salt) &* 0x9E3779B97F4A7C15)
for b in s.utf8 { h ^= UInt64(b); h &*= 0x100000001b3 }
for b in bytes { h ^= UInt64(b); h &*= 0x100000001b3 }
h ^= h >> 33; h &*= 0xff51afd7ed558ccd; h ^= h >> 33
words.append(UInt32(truncatingIfNeeded: h))
words.append(UInt32(truncatingIfNeeded: h >> 32))
}
return words
}
// The same from a string (its UTF-8 bytes).
func seedWords(_ s: String) -> [UInt32] { seedWordsBytes(Array(s.utf8)) }
struct SplitMix64 {
var s: UInt64
@ -392,8 +397,10 @@ struct GeneratorConfig {
}
var generatorConfig = GeneratorConfig()
func generateProgram(seedString: String) -> Program {
let sw = seedWords(seedString)
func generateProgram(seedString: String) -> Program { generateProgram(seedString: seedString, words: seedWords(seedString)) }
// The generator from already-derived seed words (what the chain feeds: the epoch block hash on devnet v0, the VDF output later).
func generateProgram(seedString: String, words sw: [UInt32]) -> Program {
var rng = SplitMix64(s: (UInt64(sw[0]) | (UInt64(sw[1]) << 32)) ^ ((UInt64(sw[2]) | (UInt64(sw[3]) << 32)) &* 0x9E3779B97F4A7C15))
var instrs = [Instr]()
let weights = generatorConfig.weights
@ -537,7 +544,7 @@ enum LoadSource {
case inlineMemhard(MixParams) // memory-hard: mh_word(cache, index), 8 dependent cache reads per word
}
func generateMSL(_ p: Program, datasetLog2: Int, source: LoadSource = .stored) -> String {
func generateMSL(_ p: Program, datasetLog2: Int, source: LoadSource = .stored, bound: Bool = false) -> String {
let mask = UInt32((1 << datasetLog2) - 1)
var s = """
#include <metal_stdlib>
@ -571,18 +578,22 @@ func generateMSL(_ p: Program, datasetLog2: Int, source: LoadSource = .stored) -
s += emitMemhardCore(mp, cuda: false) + "\n"
buffer0 = "device const uint* cache [[buffer(0)]]"
}
// Header-bound variant (3 October 2026, igneum-pow/src/bind.rs): same body, the init words come from buffer 3.
let kernelName = bound ? "igneum_hash_bound" : "igneum_hash"
let initArg = bound ? " constant uint* initw [[buffer(3)]],\n" : ""
let iw = bound ? "initw" : "SEEDW"
s += """
kernel void igneum_hash(\(buffer0),
kernel void \(kernelName)(\(buffer0),
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
uint gid [[thread_position_in_grid]]) {
\(initArg) uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
"""
if p.hasWide { s += " uint lane = gid & 31u;\n" }
for i in 0..<8 {
s += " { uint x = nonce ^ SEEDW[\(i)]; x += 0x9e3779b9u * \(i + 1)u; x = splitmix32(x); r\(i) = x ^ SEEDW[\((i + 1) & 7)]; }\n"
s += " { uint x = nonce ^ \(iw)[\(i)]; x += 0x9e3779b9u * \(i + 1)u; x = splitmix32(x); r\(i) = x ^ \(iw)[\((i + 1) & 7)]; }\n"
}
s += "\n for (uint it = 0u; it < \(Program.iterations)u; ++it) {\n uint sel = r0;\n"
// The word index expression for a load: plain = a & MASK; wide = lane 0's a, aligned to 32 words, plus lane.
@ -1463,9 +1474,14 @@ final class DatasetContext {
var lastBuildGPUms = 0.0, lastBuildWallMs = 0.0
private var cpu: MemhardCPU?
init(gpu: GPU, closedForm: Bool, dayString: String) {
convenience init(gpu: GPU, closedForm: Bool, dayString: String) {
self.init(gpu: gpu, closedForm: closedForm, dayString: dayString, key: seedWords("day/" + dayString))
}
// From already-derived key words (serve mode: seed_words_from_bytes of the day seed bytes the miner sends).
init(gpu: GPU, closedForm: Bool, dayString: String, key: [UInt32]) {
self.gpu = gpu; closed = closedForm; self.dayString = dayString
key = seedWords("day/" + dayString)
self.key = key
mp = closedForm ? nil : MixParams(key: key)
let t0 = nowNs()
do {
@ -2433,11 +2449,159 @@ func runTests(_ opts: Options) -> Never {
exit(all ? 0 : 1)
}
// MARK: - Serve mode (GPU worker for igneum-miner --worker), 3 October 2026
//
// Protocol, one line each. Only "found", "done" and "error" are parsed by the miner; every other line is logged.
// stdin: job <job_id> <header_prehash_hex 64> <target_hex 16> <nonce_start u64> <nonce_count u64> <epoch_seed_hex 64> <day_seed_hex>
// quit
// stdout: ready metal <device>
// found <job_id> <nonce u64> <hash_hex 16> every nonce whose 64-bit hash is <= target (hash <= target64)
// done <job_id> <hashes> <ms> end of the job (wall ms, dispatch plus scan)
// error <job_id> <text>
// The program for an epoch seed is generateProgram(words: seedWordsBytes(epoch_seed)), compiled once and cached; the
// cache and 1 GiB dataset for a day seed come from seedWordsBytes(day_seed_bytes), built once and cached (two of each).
// The init words of a dispatch are seedWordsBytes("igneum-block/" || prehash || nonce_hi_le32) and go to buffer 3 of
// igneum_hash_bound; the lane nonce is baseNonce + gid as in the bench kernel.
func emit(_ line: String) { print(line); fflush(stdout) }
func unhex(_ s: String) -> [UInt8]? {
let chars = Array(s.utf8)
if chars.count % 2 != 0 { return nil }
var out = [UInt8](); out.reserveCapacity(chars.count / 2)
var i = 0
while i < chars.count {
guard let hi = UInt8(String(UnicodeScalar(chars[i])), radix: 16), let lo = UInt8(String(UnicodeScalar(chars[i + 1])), radix: 16) else { return nil }
out.append(hi << 4 | lo); i += 2
}
return out
}
func blockInitWords(prehash: [UInt8], nonceHi: UInt32) -> [UInt32] {
var b = Array("igneum-block/".utf8)
b += prehash
b += [UInt8(nonceHi & 0xff), UInt8((nonceHi >> 8) & 0xff), UInt8((nonceHi >> 16) & 0xff), UInt8((nonceHi >> 24) & 0xff)]
return seedWordsBytes(b)
}
final class ServeProgram {
let seedHex: String
let program: Program
let compiled: CompiledHash
init(seedHex: String, program: Program, compiled: CompiledHash) { self.seedHex = seedHex; self.program = program; self.compiled = compiled }
}
final class ServeDataset {
let dayHex: String
let ctx: DatasetContext
let buffer: MTLBuffer
init(dayHex: String, ctx: DatasetContext, buffer: MTLBuffer) { self.dayHex = dayHex; self.ctx = ctx; self.buffer = buffer }
}
func compileBound(_ gpu: GPU, msl: String) throws -> CompiledHash {
let t0 = nowNs()
let lib = try gpu.device.makeLibrary(source: msl, options: MTLCompileOptions())
let t1 = nowNs()
guard let fn = lib.makeFunction(name: "igneum_hash_bound") else { throw IgneumError("no igneum_hash_bound function in library") }
let pipe = try gpu.device.makeComputePipelineState(function: fn)
let t2 = nowNs()
return CompiledHash(pipeline: pipe, libraryMs: ms(t0, t1), pipelineMs: ms(t1, t2))
}
func runServe(_ opts: Options) -> Never {
let gpu = GPU()
let datasetLog2 = opts.datasetLog2
let batch = 1 << opts.batchLog2 // nonces per dispatch
var programs = [ServeProgram]()
var datasets = [ServeDataset]()
guard let outBuf = gpu.device.makeBuffer(length: batch * 8, options: .storageModeShared) else { emit("error 0 cannot allocate the output buffer"); exit(1) }
emit("ready metal \(gpu.device.name.replacingOccurrences(of: " ", with: "_")) dataset-log2 \(datasetLog2) batch \(batch)")
while let line = readLine(strippingNewline: true) {
let f = line.split(separator: " ").map(String.init)
if f.isEmpty { continue }
if f[0] == "quit" { break }
if f[0] != "job" { emit("info ignored: \(line)"); continue }
if f.count < 8 { emit("error \(f.count > 1 ? f[1] : "0") malformed job line (need 7 fields after job)"); continue }
let jobId = f[1]
guard let prehash = unhex(f[2]), prehash.count == 32, let target = UInt64(f[3], radix: 16),
let nonceStart = UInt64(f[4]), let nonceCount = UInt64(f[5]),
let epochSeed = unhex(f[6]), epochSeed.count == 32, let daySeed = unhex(f[7]) else {
emit("error \(jobId) bad field (prehash 64 hex, target 16 hex, nonce_start u64, nonce_count u64, epoch_seed 64 hex, day_seed hex)"); continue
}
if nonceCount == 0 || nonceCount % 32 != 0 || (nonceStart & 31) != 0 { emit("error \(jobId) nonce_start must be 32-aligned and nonce_count a non-zero multiple of 32"); continue }
let t0 = nowNs()
// Program for the epoch seed
var prog = programs.first { $0.seedHex == f[6] }
if prog == nil {
let p = generateProgram(seedString: "epoch/" + f[6], words: seedWordsBytes(epochSeed))
let msl = generateMSL(p, datasetLog2: datasetLog2, source: .stored, bound: true)
do {
let c = try compileBound(gpu, msl: msl)
emit("info program epoch \(f[6].prefix(16)) loads/hash \(p.loadsPerHash) compiled in \(fmt(c.totalMs, 1)) ms")
let sp = ServeProgram(seedHex: f[6], program: p, compiled: c)
programs.append(sp); if programs.count > 2 { programs.removeFirst() }
prog = sp
} catch { emit("error \(jobId) Metal compile failed: \(error)"); continue }
}
// Cache and dataset for the day seed
var ds = datasets.first { $0.dayHex == f[7] }
if ds == nil {
let key = seedWordsBytes(daySeed)
let ctx = DatasetContext(gpu: gpu, closedForm: false, dayString: "day/" + f[7], key: key)
let buf = ctx.makeDataset(log2: datasetLog2)
emit("info dataset day \(f[7]) key \(key.map { String(format: "%08x", $0) }.joined(separator: " ")) cache fill \(fmt(ctx.cacheFillGPUms, 1)) ms build \(fmt(ctx.lastBuildGPUms, 1)) ms GPU")
let sd = ServeDataset(dayHex: f[7], ctx: ctx, buffer: buf)
datasets.append(sd); if datasets.count > 2 { datasets.removeFirst() }
ds = sd
}
guard let program = prog, let dataset = ds else { continue }
// Mine: chunks of at most `batch` lane nonces that share one high word
var remaining = nonceCount
var hi = UInt32(truncatingIfNeeded: nonceStart >> 32)
var lo = UInt32(truncatingIfNeeded: nonceStart)
var hashes: UInt64 = 0
var found = 0
var failed = false
let outPtr = outBuf.contents().bindMemory(to: UInt64.self, capacity: batch)
while remaining > 0 {
let room = UInt64(UInt32.max - lo) + 1 // lane nonces left before the high word steps
let chunk = Int(min(min(remaining, UInt64(batch)), room))
var initw = blockInitWords(prehash: prehash, nonceHi: hi)
let cb = gpu.queue.makeCommandBuffer()!
let enc = cb.makeComputeCommandEncoder()!
enc.setComputePipelineState(program.compiled.pipeline)
enc.setBuffer(dataset.buffer, offset: 0, index: 0)
enc.setBuffer(outBuf, offset: 0, index: 1)
var b = lo
enc.setBytes(&b, length: 4, index: 2)
enc.setBytes(&initw, length: 32, index: 3)
enc.dispatchThreadgroups(MTLSize(width: chunk / 32, height: 1, depth: 1), threadsPerThreadgroup: MTLSize(width: 32, height: 1, depth: 1))
enc.endEncoding()
cb.commit(); cb.waitUntilCompleted()
if let e = cb.error { emit("error \(jobId) dispatch failed: \(e)"); failed = true; break }
for i in 0..<chunk where outPtr[i] <= target {
let nonce = (UInt64(hi) << 32) | UInt64(lo &+ UInt32(i))
emit("found \(jobId) \(nonce) \(String(format: "%016llx", outPtr[i]))")
found += 1
}
hashes += UInt64(chunk)
remaining -= UInt64(chunk)
let (newLo, wrapped) = lo.addingReportingOverflow(UInt32(truncatingIfNeeded: chunk))
lo = newLo
if wrapped || (chunk == Int(room)) { hi &+= 1; lo = 0 }
}
if failed { continue }
emit("done \(jobId) \(hashes) \(fmt(ms(t0, nowNs()), 2))")
}
exit(0)
}
// MARK: - Main
let opts = parseArgs()
generatorConfig = GeneratorConfig(loadWeight: opts.loadWeight, wideFrac: opts.wideFrac)
if opts.exportPack != nil { exportPack(opts) }
if opts.serve { runServe(opts) }
if opts.anyTest { runTests(opts) }
let gpu = GPU()
print("igneum-bench")

View file

@ -27,6 +27,7 @@
#ifdef _WIN32
#define WIN32_LEAN_AND_MEAN
#include <windows.h>
#define strtok_r strtok_s
#else
#include <time.h>
#include <dlfcn.h>
@ -190,6 +191,9 @@ typedef struct {
int timeWall; // 1 = rate from wall time, 0 = from device event profiling, -1 = auto (wall on the Apple platform)
const char* kernelPath;
const char* extraOpts;
int serve; // --serve: GPU worker for igneum-miner --worker (jobs on stdin), 3 October 2026
int kernelGiven; // --kernel was passed
const char* vendor; // --vendor S: pick the first GPU whose vendor string contains S (default: first GPU of any vendor)
} Options;
static int packMib(void) { return (int)(((1ull << IGNEUM_DATASET_LOG2) * 4ull) >> 20); }
@ -210,7 +214,10 @@ static void usage(void) {
" --kernel P path to the pack's kernel.cl (default: the path compiled in, %s)\n"
" --build-opts S extra options appended to clBuildProgram (for example \"-cl-std=CL2.0\")\n"
" --time T event (default): hashes/s from device event profiling, like cudaEvent time; wall: from host wall time.\n"
" Apple's OpenCL runtime reports unusable event timestamps, so wall is the default on the Apple platform.\n", packMib(), IGNEUM_KERNEL_PATH);
" Apple's OpenCL runtime reports unusable event timestamps, so wall is the default on the Apple platform.\n"
" --vendor S choose the first GPU whose vendor string contains S (for example \"Advanced Micro Devices\"); fails if none\n"
" --serve GPU worker for igneum-miner --worker: reads \"job ...\" lines on stdin, prints found/done lines.\n"
" Builds the pack's kernel_bound.cl (next to the compiled-in kernel.cl) unless --kernel says otherwise.\n", packMib(), IGNEUM_KERNEL_PATH);
}
static int isPow2(long long v) { return v > 0 && (v & (v - 1)) == 0; }
@ -220,7 +227,7 @@ static Options parseArgs(int argc, char** argv) {
Options o;
int i;
o.datasetMib = 1024; o.batchLog2 = 24; o.batches = 5; o.groupWarps = 1; o.sweep = 0; o.device = -1;
o.exchange = 0; o.list = 0; o.timeWall = -1; o.kernelPath = IGNEUM_KERNEL_PATH; o.extraOpts = "";
o.exchange = 0; o.list = 0; o.timeWall = -1; o.kernelPath = IGNEUM_KERNEL_PATH; o.extraOpts = ""; o.serve = 0; o.kernelGiven = 0; o.vendor = NULL;
for (i = 1; i < argc; ++i) {
const char* a = argv[i];
int needs = (strcmp(a, "--dataset-mib") == 0 || strcmp(a, "--batch-log2") == 0 || strcmp(a, "--batches") == 0 ||
@ -232,7 +239,9 @@ static Options parseArgs(int argc, char** argv) {
else if (strcmp(a, "--batches") == 0) o.batches = atoi(argv[++i]);
else if (strcmp(a, "--group-warps") == 0) o.groupWarps = atoi(argv[++i]);
else if (strcmp(a, "--device") == 0) o.device = atoi(argv[++i]);
else if (strcmp(a, "--kernel") == 0) o.kernelPath = argv[++i];
else if (strcmp(a, "--kernel") == 0) { o.kernelPath = argv[++i]; o.kernelGiven = 1; }
else if (strcmp(a, "--serve") == 0) o.serve = 1;
else if (strcmp(a, "--vendor") == 0) { if (i + 1 >= argc) { usage(); exit(2); } o.vendor = argv[++i]; }
else if (strcmp(a, "--build-opts") == 0) o.extraOpts = argv[++i];
else if (strcmp(a, "--time") == 0) {
const char* m = argv[++i];
@ -376,6 +385,7 @@ typedef struct {
cl_command_queue q;
cl_program prog;
cl_kernel kHash, kCacheFill, kBuild, kFill;
cl_kernel kHashBound; // igneum_hash_bound (serve mode; NULL when the source has none)
int exchange; // 0 local memory, 1 khr sub-group shuffle, 2 intel
size_t subGroupSize; // as queried for a 32-item work-group, 0 if not queried
char exchangeNote[512];
@ -412,6 +422,8 @@ static int buildProgram(Device* dv, const DeviceInfo* di, const char* src, size_
return 1;
}
dv->kHash = clCreateKernel(dv->prog, "igneum_hash", &err); CL_CHECK_ERR(err, "clCreateKernel igneum_hash");
dv->kHashBound = clCreateKernel(dv->prog, "igneum_hash_bound", &err);
if (err != CL_SUCCESS) dv->kHashBound = NULL; /* kernel.cl without the bound kernel: fine outside --serve */
#if IGNEUM_DATASET_MODE == 1
dv->kCacheFill = clCreateKernel(dv->prog, "igneum_cache_fill", &err); CL_CHECK_ERR(err, "clCreateKernel igneum_cache_fill");
dv->kBuild = clCreateKernel(dv->prog, "igneum_build", &err); CL_CHECK_ERR(err, "clCreateKernel igneum_build");
@ -423,11 +435,12 @@ static int buildProgram(Device* dv, const DeviceInfo* di, const char* src, size_
static void releaseProgram(Device* dv) {
if (dv->kHash) clReleaseKernel(dv->kHash);
if (dv->kHashBound) clReleaseKernel(dv->kHashBound);
if (dv->kCacheFill) clReleaseKernel(dv->kCacheFill);
if (dv->kBuild) clReleaseKernel(dv->kBuild);
if (dv->kFill) clReleaseKernel(dv->kFill);
if (dv->prog) clReleaseProgram(dv->prog);
dv->kHash = dv->kCacheFill = dv->kBuild = dv->kFill = NULL; dv->prog = NULL;
dv->kHash = dv->kCacheFill = dv->kBuild = dv->kFill = dv->kHashBound = NULL; dv->prog = NULL;
}
// Sub-group size of igneum_hash for a work-group of `local` items. 0 if the query is unavailable (reason in *why).
@ -824,6 +837,173 @@ static SizeResult runSize(Device* dv, const DeviceInfo* di, const Options* o, in
// ---------------------------------------------------------------------------------------------
// Main
/* ---------------------------------------------------------------------------------------------
* Serve mode: GPU worker for igneum-miner --worker (3 October 2026)
*
* Protocol, one line each. Only "found", "done" and "error" are parsed by the miner; every other line is logged.
* stdin: job <job_id> <header_prehash_hex 64> <target_hex 16> <nonce_start u64> <nonce_count u64> <epoch_seed_hex 64> <day_seed_hex>
* quit
* stdout: ready opencl <device>
* found <job_id> <nonce u64> <hash_hex 16> every nonce whose 64-bit hash is <= target
* done <job_id> <hashes> <ms> end of the job (wall ms)
* error <job_id> <text>
* The pack's program.h is compiled in and kernel_bound.cl is built at runtime, so this worker serves exactly one
* epoch seed and one day seed: the pack's. A job for other seeds is answered with an error naming both; re-export the
* pack with `igneum-miner export-pack <node> <dir>` and rebuild. The init words of a dispatch are
* seed_words_from_bytes("igneum-block/" || prehash || nonce_hi_le32), written to a small buffer that is the fifth
* argument of igneum_hash_bound; the lane nonce is baseNonce + gid as in the bench kernel. The exchange rule of
* WAVEFRONT.md applies unchanged (the bound kernel has the same body and the same IGNEUM_EXCHANGE build).
*/
static void seedWordsFromBytes(const uint8_t* b, size_t n, uint32_t out[8]) {
uint64_t salt;
for (salt = 0; salt < 4; ++salt) {
uint64_t h = 0xcbf29ce484222325ull ^ (salt * 0x9E3779B97F4A7C15ull);
size_t i;
for (i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
h ^= h >> 33; h *= 0xff51afd7ed558ccdull; h ^= h >> 33;
out[2 * salt] = (uint32_t)h;
out[2 * salt + 1] = (uint32_t)(h >> 32);
}
}
static int unhexBuf(const char* s, uint8_t* out, size_t cap, size_t* len) {
size_t n = strlen(s), i;
if (n % 2 || n / 2 > cap) return 0;
for (i = 0; i < n; i += 2) {
unsigned v = 0;
char two[3] = { s[i], s[i + 1], 0 };
if (sscanf(two, "%2x", &v) != 1) return 0;
out[i / 2] = (uint8_t)v;
}
*len = n / 2;
return 1;
}
static int runServe(Device* dv, const DeviceInfo* di, const Options* o) {
#if IGNEUM_DATASET_MODE != 1
(void)dv; (void)di; (void)o;
printf("error 0 --serve needs a memory-hard pack (IGNEUM_DATASET_MODE 1)\n"); fflush(stdout);
return 2;
#else
static const uint32_t KEYW[8] = IGNEUM_KEY_INIT;
const uint32_t words = 1u << IGNEUM_DATASET_LOG2;
const uint32_t mask = words - 1u;
const uint32_t batch = 1u << (o->batchLog2 == 24 ? 22 : o->batchLog2); /* 2^22 nonces per dispatch by default */
size_t groupSize = 32 * (size_t)o->groupWarps;
cl_int err = 0;
cl_mem dDs, dOut, dInit;
cl_uint nItems = words / 16u;
uint64_t* hOut;
char devName[256];
char line[1024];
size_t k;
if (!dv->kHashBound) { printf("error 0 the kernel source has no igneum_hash_bound (build from the pack's kernel_bound.cl, or pass --kernel)\n"); fflush(stdout); return 2; }
if (!setupCache(dv, di)) { printf("error 0 cache check failed (device cache differs from the host cache or the pack's FNV)\n"); fflush(stdout); return 1; }
dDs = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, (size_t)words * 4u, NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer dataset");
CL_CHECK(clSetKernelArg(dv->kBuild, 0, sizeof(cl_mem), &dDs));
CL_CHECK(clSetKernelArg(dv->kBuild, 1, sizeof(cl_mem), &gCache));
CL_CHECK(clSetKernelArg(dv->kBuild, 2, sizeof(cl_uint), &nItems));
clReleaseEvent(launch1D(dv, dv->kBuild, nItems, kernelMaxLocal(dv, dv->kBuild, di, 256)));
CL_CHECK(clFinish(dv->q));
dOut = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, (size_t)batch * sizeof(uint64_t), NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer out");
dInit = clCreateBuffer(dv->ctx, CL_MEM_READ_ONLY, 32, NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer init words");
hOut = (uint64_t*)malloc((size_t)batch * sizeof(uint64_t));
strncpy(devName, di->name, 255); devName[255] = 0;
for (k = 0; devName[k]; ++k) if (devName[k] == ' ') devName[k] = '_';
printf("ready opencl %s platform %s pack %s dataset-log2 %d batch %u exchange %d\n", devName, di->platformName, IGNEUM_SEED_STRING, IGNEUM_DATASET_LOG2, batch, dv->exchange);
fflush(stdout);
while (fgets(line, sizeof(line), stdin)) {
char* f[9];
int nf = 0;
char* tok;
char* save = NULL;
char jobId[64];
uint8_t prehash[32], epochSeed[32], daySeed[256];
size_t prehashLen = 0, epochLen = 0, dayLen = 0;
unsigned long long target = 0, nonceStart = 0, nonceCount = 0;
uint32_t sw[8], kw[8];
uint64_t remaining, hashes = 0;
uint32_t hi, lo;
double t0;
int failed = 0;
line[strcspn(line, "\r\n")] = 0;
for (tok = strtok_r(line, " ", &save); tok && nf < 9; tok = strtok_r(NULL, " ", &save)) f[nf++] = tok;
if (nf == 0) continue;
if (strcmp(f[0], "quit") == 0) break;
if (strcmp(f[0], "job") != 0) { printf("info ignored line\n"); fflush(stdout); continue; }
strncpy(jobId, nf > 1 ? f[1] : "0", 63); jobId[63] = 0;
if (nf < 8) { printf("error %s malformed job line (need 7 fields after job)\n", jobId); fflush(stdout); continue; }
if (!unhexBuf(f[2], prehash, 32, &prehashLen) || prehashLen != 32 || sscanf(f[3], "%llx", &target) != 1 ||
sscanf(f[4], "%llu", &nonceStart) != 1 || sscanf(f[5], "%llu", &nonceCount) != 1 ||
!unhexBuf(f[6], epochSeed, 32, &epochLen) || epochLen != 32 || !unhexBuf(f[7], daySeed, sizeof(daySeed), &dayLen)) {
printf("error %s bad field (prehash 64 hex, target 16 hex, nonce_start u64, nonce_count u64, epoch_seed 64 hex, day_seed hex)\n", jobId); fflush(stdout); continue;
}
if (nonceCount == 0 || nonceCount % 32 != 0 || (nonceStart & 31) != 0) { printf("error %s nonce_start must be 32-aligned and nonce_count a non-zero multiple of 32\n", jobId); fflush(stdout); continue; }
seedWordsFromBytes(epochSeed, 32, sw);
seedWordsFromBytes(daySeed, dayLen, kw);
if (memcmp(sw, SEEDW, 32) != 0) {
printf("error %s epoch seed mismatch: this worker was built for pack \"%s\" (seed words %08x %08x ...), the job's epoch seed %.16s gives %08x %08x ...; run igneum-miner export-pack and rebuild\n",
jobId, IGNEUM_SEED_STRING, SEEDW[0], SEEDW[1], f[6], sw[0], sw[1]);
fflush(stdout); continue;
}
if (memcmp(kw, KEYW, 32) != 0) {
printf("error %s day seed mismatch: this worker's cache is for key %08x %08x ..., the job's day seed %s gives %08x %08x ...; run igneum-miner export-pack and rebuild\n",
jobId, KEYW[0], KEYW[1], f[7], kw[0], kw[1]);
fflush(stdout); continue;
}
t0 = wallMs();
remaining = nonceCount;
hi = (uint32_t)(nonceStart >> 32); lo = (uint32_t)nonceStart;
while (remaining > 0) {
uint64_t room = (uint64_t)(0xffffffffu - lo) + 1ull;
uint64_t chunk64 = remaining < batch ? remaining : batch;
uint32_t chunk, i;
uint32_t iw[8];
uint8_t b[49];
cl_uint baseNonce, maskArg = mask;
cl_event ev;
if (chunk64 > room) chunk64 = room;
chunk = (uint32_t)chunk64;
memcpy(b, "igneum-block/", 13);
memcpy(b + 13, prehash, 32);
b[45] = (uint8_t)hi; b[46] = (uint8_t)(hi >> 8); b[47] = (uint8_t)(hi >> 16); b[48] = (uint8_t)(hi >> 24);
seedWordsFromBytes(b, 49, iw);
CL_CHECK(clEnqueueWriteBuffer(dv->q, dInit, CL_TRUE, 0, 32, iw, 0, NULL, NULL));
baseNonce = lo;
CL_CHECK(clSetKernelArg(dv->kHashBound, 0, sizeof(cl_mem), &dDs));
CL_CHECK(clSetKernelArg(dv->kHashBound, 1, sizeof(cl_mem), &dOut));
CL_CHECK(clSetKernelArg(dv->kHashBound, 2, sizeof(cl_uint), &baseNonce));
CL_CHECK(clSetKernelArg(dv->kHashBound, 3, sizeof(cl_uint), &maskArg));
CL_CHECK(clSetKernelArg(dv->kHashBound, 4, sizeof(cl_mem), &dInit));
ev = launch1D(dv, dv->kHashBound, chunk, groupSize);
err = clFinish(dv->q);
clReleaseEvent(ev);
if (err != CL_SUCCESS) { printf("error %s dispatch failed: %s\n", jobId, clErrName(err)); fflush(stdout); failed = 1; break; }
CL_CHECK(clEnqueueReadBuffer(dv->q, dOut, CL_TRUE, 0, (size_t)chunk * sizeof(uint64_t), hOut, 0, NULL, NULL));
for (i = 0; i < chunk; ++i) if (hOut[i] <= target) {
unsigned long long nonce = ((unsigned long long)hi << 32) | (unsigned long long)(uint32_t)(lo + i);
printf("found %s %llu %016llx\n", jobId, nonce, (unsigned long long)hOut[i]);
}
fflush(stdout);
hashes += chunk;
remaining -= chunk;
if (chunk64 == room) { hi += 1u; lo = 0u; } else lo += chunk;
}
if (failed) continue;
printf("done %s %llu %.2f\n", jobId, (unsigned long long)hashes, wallMs() - t0);
fflush(stdout);
}
free(hOut);
clReleaseMemObject(dInit);
clReleaseMemObject(dOut);
clReleaseMemObject(dDs);
return 0;
#endif
}
int main(int argc, char** argv) {
Options o = parseArgs(argc, argv);
DeviceInfo* devs = NULL;
@ -845,6 +1025,14 @@ int main(int argc, char** argv) {
if (o.device >= 0) {
if (o.device >= nDev) { printf("FAIL: --device %d out of range (%d devices)\n", o.device, nDev); return 2; }
chosen = o.device;
} else if (o.vendor) {
for (i = 0; i < nDev; ++i) if ((devs[i].type & CL_DEVICE_TYPE_GPU) && strstr(devs[i].vendor, o.vendor)) { chosen = i; break; }
if (chosen < 0) {
printf("OpenCL devices (%d):\n", nDev);
for (i = 0; i < nDev; ++i) printDevice(i, &devs[i], 0);
printf("FAIL: no GPU whose vendor contains \"%s\"\n", o.vendor);
return 2;
}
} else {
for (i = 0; i < nDev; ++i) if (devs[i].type & CL_DEVICE_TYPE_GPU) { chosen = i; break; }
if (chosen < 0) chosen = 0;
@ -864,6 +1052,15 @@ int main(int argc, char** argv) {
dv.q = clCreateCommandQueue(dv.ctx, di->device, CL_QUEUE_PROFILING_ENABLE, &err);
CL_CHECK_ERR(err, "clCreateCommandQueue");
if (o.serve && !o.kernelGiven) {
/* The bound kernel lives next to the compiled-in kernel.cl as kernel_bound.cl (packs from igneum-pow or igneum-miner export-pack). */
static char boundPath[1024];
size_t n = strlen(o.kernelPath);
if (n >= 9 && strcmp(o.kernelPath + n - 9, "kernel.cl") == 0) {
snprintf(boundPath, sizeof(boundPath), "%.*skernel_bound.cl", (int)(n - 9), o.kernelPath);
o.kernelPath = boundPath;
}
}
src = readFile(o.kernelPath, &srcLen);
if (!src) { printf("FAIL: cannot read kernel source %s (run from proto-opencl/ or pass --kernel)\n", o.kernelPath); return 2; }
printf("kernel source: %s (%llu bytes)\n", o.kernelPath, (unsigned long long)srcLen);
@ -871,6 +1068,7 @@ int main(int argc, char** argv) {
free(src);
printf("build options: %s\n", dv.buildOptions);
printf("exchange: %s\n", dv.exchangeNote);
if (o.serve) return runServe(&dv, di, &o);
{
size_t wg = 0;
cl_ulong lmem = 0;