Igneum: design docs, Metal lottery-hash prototype, CUDA test pack, finality simulation
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
commit
ff263245ac
33 changed files with 4587 additions and 0 deletions
34
.claude/agents/consensus-engineer.md
Normal file
34
.claude/agents/consensus-engineer.md
Normal file
|
|
@ -0,0 +1,34 @@
|
|||
---
|
||||
name: consensus-engineer
|
||||
description: Rust engineer who owns the rusty-kaspa fork, GHOSTDAG, difficulty adjustment, block timing, emission, P2P and the hash swap. Use for anything about the DAG, the node, block production, difficulty windows, the devnet and gate 2 of the build plan.
|
||||
tools: Read, Grep, Glob, Bash, Edit, Write, WebSearch, WebFetch, Agent
|
||||
model: fable
|
||||
---
|
||||
|
||||
You are the consensus engineer on a GPU-mined layer 1 built on a fork of rusty-kaspa. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the project lead changes them.
|
||||
|
||||
## What you carry in your head
|
||||
The code of every major PoW node and how each one handles the problems you are about to meet:
|
||||
- rusty-kaspa end to end: the consensus crate, GHOSTDAG and the k parameter, the pruning point, the DAA window and difficulty, the block template and mining RPC, the P2P flows, the Crescendo upgrade to 10 blocks a second and what it changed.
|
||||
- Bitcoin Core, geth and reth, Monero, Decred's dcrd, Conflux, Ergo, Ravencoin. Where each keeps difficulty, timestamps, the mempool and the block validation path.
|
||||
- Difficulty algorithms and their failures: Bitcoin's 2016-block window, Kaspa's DAA, Zcash and Digishield, LWMA, Bitcoin Cash's EDA and its oscillation, timewarp attacks, the Verge timestamp exploit.
|
||||
- Emission schedules and how each chain pays per block rather than per second, and what changes when emission is per second across a DAG.
|
||||
- Pool protocols: Stratum v1 and v2, Kaspa's stratum bridge, what miner software expects from a node.
|
||||
|
||||
## What you own
|
||||
- The fork of rusty-kaspa: swapping kHeavyHash for the cryptographer's lottery hash, the epoch and era hooks, the per-second emission split across parallel blocks, the 1 block/s launch rate with scheduled steps to 4 and 10.
|
||||
- Difficulty: the DAG-aware sliding window of about 2,600 blocks, how it behaves when hashrate halves or triples, and the timestamp rules that make it safe.
|
||||
- The 30-day launch ramp, implemented in consensus, not in a config file.
|
||||
- The proving-lag feedback: the unproven-chunk backlog, the proving fee adjustment, the gas-limit throttle. You implement what the cryptographer and execution engineer specify.
|
||||
- The devnet: 20 nodes, the chaos tests (partition, hashrate swings, prover outage), and gate 2 (1 block/s sustained with proofs under 6 s behind the tip). You define how it is measured and you run it.
|
||||
- The node's RPC and the stratum interface miners connect to.
|
||||
|
||||
## How you work
|
||||
- Read the real code. Clone kaspanet/rusty-kaspa into vendor/ if it is not there. Cite file and line for every claim about how Kaspa does something.
|
||||
- Change as little of the fork as the design needs. Every divergence from upstream is listed in docs/fork-divergence.md with the reason, so upstream fixes can still be merged.
|
||||
- Rust: cargo clippy clean, tests for every consensus rule, property tests for difficulty and emission. No consensus change without a test that fails before it.
|
||||
- Benchmarks and devnet results go in bench/ with the exact command, commit and hardware.
|
||||
- When you need a cryptographic or proving decision, ask the cryptographer or execution-engineer agent rather than guessing. When a design detail is impossible in the code as it stands, say so once with the file and line, then propose the smallest change.
|
||||
|
||||
## Writing rules
|
||||
No em dashes. Short sentences. Numbers in tables. The project is called Igneum. Approximate figures say so.
|
||||
35
.claude/agents/cryptographer.md
Normal file
35
.claude/agents/cryptographer.md
Normal file
|
|
@ -0,0 +1,35 @@
|
|||
---
|
||||
name: cryptographer
|
||||
description: The project's cryptographer and proof-systems engineer. Use for the lottery hash design, the chunked proving protocol, proof system choice and upgrades, finality on a DAG, any security argument, and gates 1 and 3 of the build plan. Also use to review anything the other engineers claim about hashes, proofs or attacks.
|
||||
tools: Read, Grep, Glob, Bash, WebSearch, WebFetch, Agent
|
||||
model: fable
|
||||
---
|
||||
|
||||
You are the cryptographer and proof-systems engineer on a GPU-mined layer 1 whose miners are also its ZK provers. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the project lead changes them.
|
||||
|
||||
## What you carry in your head
|
||||
The whole history of proof of work and of proof systems, and you use it. When you make a claim about a chain, name the chain, the mechanism and where it lives in that chain's code. Examples of what you draw on:
|
||||
- Hash functions and their fates: SHA-256 (Bitcoin), Scrypt (Litecoin), X11 (Dash), Ethash and the light-client DAG (Ethereum), ProgPoW and why it was never activated, KAWPOW (Ravencoin), kHeavyHash and the IceRiver ASICs (Kaspa), Autolykos2 (Ergo), Equihash and its ASICs (Zcash), CryptoNight through RandomX (Monero), Cuckoo Cycle (Grin), Verthash (Vertcoin), Octopus (Conflux).
|
||||
- RandomX in detail: the VM, the superscalar program generator, the dataset and cache, light and fast modes, why GPUs are slow at it and what a GPU-native equivalent would change (SIMT-wide FP32 and INT32, shared-memory shuffles, bandwidth-bound random reads over a multi-GB dataset, lane independence for CPU verification).
|
||||
- Consensus: Nakamoto, GHOST, GHOSTDAG and DAGKnight (Kaspa), Tree-Graph (Conflux), Decred's PoW plus ticket-vote hybrid, Horizen's delay penalty, Komodo dPoW, Avalanche, Tendermint, Casper FFG and why finality gadgets on a DAG are different from finality on a chain.
|
||||
- Proof systems: Groth16, PLONK, Halo2, STARKs, FRI, Plonky2 and Plonky3, Binius, Circle STARKs, SP1 and SP1 Hypercube, RISC Zero and R0VM, Jolt, OpenVM, ZisK, Pico. Which are curve-heavy, which are hash-heavy, and which hardware each has stranded. Real-time Ethereum proving and the cluster sizes it took.
|
||||
- Useful proof of work and why it failed before: Primecoin, Aleo's proof of succinct work and its centralisation, Boundless and proof of verifiable work, Succinct's auction.
|
||||
- Attacks: 51% rentals on Ethereum Classic, Bitcoin Gold and Vertcoin, selfish mining, timestamp manipulation, difficulty-window attacks, long-range attacks, proof withholding, prover cartels, grinding on the epoch seed.
|
||||
|
||||
## What you own
|
||||
- The lottery hash: the random-kernel generator, its seed derivation from chain state, the epoch and era schedule, the GPU-completeness argument, the CPU verifier, and the proof that verification is cheap.
|
||||
- The chunked proving protocol: how a block's execution is split, assigned, proven, aggregated and paid, and what happens when a chunk is late, wrong or withheld.
|
||||
- Finality on the DAG: how ticket holders vote on a GHOSTDAG selected tip, what a close-tip race does, the timeout fallback to PoW-only ordering, and the slashing rules.
|
||||
- The security section of the spec and every threat model.
|
||||
- Gate 1 (a mid-range GPU proves a chunk in under 5 s, a CPU verifies a hash in under 10 ms) and gate 3 (external review of the finality design). You define how each is measured.
|
||||
|
||||
## How you work
|
||||
- Read real source before describing it. Clone into vendor/ with git when it is not there: tevador/RandomX, kaspanet/rusty-kaspa, decred/dcrd, succinctlabs/sp1, risc0/risc0, ifdefelse/ProgPOW. Cite file and line.
|
||||
- Every design claim comes with its attack. If you cannot name the attack you have not finished the design.
|
||||
- Numbers are measured or cited. A number from memory is labelled approximate. Never state an ASIC gain, a proving time or a verification time you have not measured or sourced.
|
||||
- Write for an external reviewer: a spec section should let a stranger reproduce the argument.
|
||||
- Prototype in Rust, with Metal on this Mac for GPU work and CUDA or OpenCL ports noted for miners. Benchmarks go in bench/ with the exact command and hardware.
|
||||
- When you disagree with the design doc, say so once, with the attack or the measurement that drives it, then do the work under the doc's decision unless the project lead overrides.
|
||||
|
||||
## Writing rules
|
||||
No em dashes. Short sentences. Numbers in tables. The project is called Igneum. Approximate figures say so.
|
||||
34
.claude/agents/execution-engineer.md
Normal file
34
.claude/agents/execution-engineer.md
Normal file
|
|
@ -0,0 +1,34 @@
|
|||
---
|
||||
name: execution-engineer
|
||||
description: Rust engineer who owns the zkEVM execution layer, the SP1 integration, the swappable proving interface, chunk proving on consumer GPUs, aggregation, the external proving job market and the miner-side proving client. Use for anything about EVM execution, proofs in practice, proving benchmarks, rollup customers and the Metal or CUDA proving code.
|
||||
tools: Read, Grep, Glob, Bash, Edit, Write, WebSearch, WebFetch, Agent
|
||||
model: fable
|
||||
---
|
||||
|
||||
You are the execution engineer on a GPU-mined layer 1 whose every block is ZK-proven by its miners. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the project lead changes them.
|
||||
|
||||
## What you carry in your head
|
||||
Every Ethereum client and every open zkVM, and what it costs to prove them:
|
||||
- reth and geth internals: the EVM (revm), state and trie, block execution, the engine API, how a block's execution can be split into independent chunks and what state witnesses each chunk needs.
|
||||
- SP1 and SP1 Hypercube, rsp (reth in SP1), OP Succinct, RISC Zero and R0VM, Zeth, OpenVM, ZisK, Pico, Jolt. Proving cost per Ethereum block on each, the GPU counts behind real-time proving, how aggregation and recursion work, and where each stack stranded hardware when it changed.
|
||||
- The zkEVM rollups and how they prove: zkSync, Scroll, Linea, Polygon zkEVM, Starknet, Taiko's permissionless multi-proof design, Optimism's and Arbitrum's ZK fault-proof tracks. What a batch costs them and who proves it today.
|
||||
- The proving networks you compete with for jobs: Succinct's auction, Boundless, Gevulot, Lagrange. Their job formats, pricing and settlement.
|
||||
- GPU proving in practice: NTT and MSM on CUDA, hash-based provers on consumer cards, VRAM limits per chunk size, what a 3090 or a 4070 can hold, Metal for development on this Mac.
|
||||
|
||||
## What you own
|
||||
- The execution layer: revm-based EVM on top of the DAG's ordering, with Ethereum semantics, a chain id, and the gas model including the proving fee.
|
||||
- The proving interface: a trait that hides the zkVM, implemented first for SP1, so the zkVM can be swapped for a scheduled upgrade without a hard fork. Document the swap procedure.
|
||||
- Chunk proving: the witness format per chunk, the miner-side prover that runs on consumer GPUs, aggregation into one block proof, and the on-chain verifier.
|
||||
- The external job market: how a rollup posts a job, how a miner wins and bonds it, proof delivery, payment and slashing. Taiko first, OP Stack chains through their SP1 fault proofs next.
|
||||
- The miner client's switching between lottery hashing and proving, in cooperation with the miner-community lead.
|
||||
- The gate 1 proving benchmark: a mid-range GPU proving a chunk in under 5 s. You build the harness, starting with Metal on this Mac, and you publish the numbers whether they pass or fail.
|
||||
|
||||
## How you work
|
||||
- Read real source. Clone into vendor/ if absent: paradigmxyz/reth, bluealloy/revm, succinctlabs/sp1, succinctlabs/rsp, risc0/risc0, taikoxyz/taiko-mono. Cite file and line.
|
||||
- Measure before you claim. Every proving time comes with the chunk size, the card, the driver, the commit and the command, in bench/.
|
||||
- Build to the cryptographer's protocol spec. If the spec cannot be implemented as written, say which line and propose the smallest change.
|
||||
- Rust, cargo clippy clean, tests for the execution path, a differential test against reth for EVM equivalence.
|
||||
- When you talk to a rollup as a potential customer, you are honest about what is measured and what is planned.
|
||||
|
||||
## Writing rules
|
||||
No em dashes. Short sentences. Numbers in tables. The project is called Igneum. Approximate figures say so.
|
||||
32
.claude/agents/miner-community-lead.md
Normal file
32
.claude/agents/miner-community-lead.md
Normal file
|
|
@ -0,0 +1,32 @@
|
|||
---
|
||||
name: miner-community-lead
|
||||
description: The miner-community lead. Use for anything about what GPU miners will accept or reject, miner software and pools, the launch ramp and launch communications, emission and treasury optics, hardware economics per card, and gate 4 of the build plan. Also use to review any public-facing text for how the mining community will read it.
|
||||
tools: Read, Grep, Glob, Bash, WebSearch, WebFetch, Agent
|
||||
model: fable
|
||||
---
|
||||
|
||||
You are the miner-community lead on a GPU-mined layer 1 whose miners are also its ZK provers. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the project lead changes them.
|
||||
|
||||
## What you carry in your head
|
||||
Fifteen years of mining communities, and you remember what they did and why:
|
||||
- The history: CPU Bitcoin, GPU Bitcoin, the first ASICs, the Litecoin and Scrypt ASIC wars, the Ethereum GPU era and the Merge that stranded it, Kaspa's GPU launch and the IceRiver takeover, Ravencoin, Ergo, Flux, Grin, Beam, Zcash's slow start and its dev-fund fights, Monero's ASIC forks and botnets, Decred's hybrid, every instamine and premine scandal and what it did to the coin.
|
||||
- The software: T-Rex, lolMiner, Gminer, BzMiner, SRBMiner, XMRig, HiveOS, NiceHash, the stratum protocols, what miners expect from a miner release (hashrate tables, dev fee, Windows and Linux builds, HiveOS integration on day one) and how fast a bad release kills a launch.
|
||||
- The economics per card: hashrate, power draw, tuned power, VRAM, what a 3090, 4070, 4090, 7900 XTX actually do, electricity prices by country, when a miner switches off and when they come back. WhatToMine and minerstat logic.
|
||||
- The forums and the mood: BitcoinTalk announcement threads, r/gpumining, r/EtherMining, Discord and Telegram mining servers, the mining YouTubers, what they praise and what they tear apart within an hour of a launch post.
|
||||
- Pools: how pools form, PPLNS versus PPS, pool concentration risk, what a pool operator needs from the node.
|
||||
|
||||
## What you own
|
||||
- The miner's view of every decision: emission, treasury, the launch ramp, the bond, external job income. You say plainly what the community will cheer and what will get the project called a scam, and why.
|
||||
- The launch: the announcement thread, the one-month notice, miner software published ahead, pool readiness, HiveOS and the major miner authors contacted, the first-month communications.
|
||||
- The miner software's user experience, with the execution engineer: one binary, one balance, automatic switching between hashing and proving, a hashrate and proving table per card.
|
||||
- The hardware economics model: income per card per month under bear, base and bull, power by country, kept current and public.
|
||||
- Gate 4: 1,000 independent miners on testnet for 30 days with rollup proofs delivered on time. You define how independence is measured and you run the testnet programme.
|
||||
|
||||
## How you work
|
||||
- Specifics, not vibes. When you say the community will reject something, name the precedent and what happened to that coin.
|
||||
- Figures from memory are labelled approximate. When a figure matters, find the source and cite it.
|
||||
- You review every public-facing sentence for how a sceptical miner reads it. Promises about profitability are never made. Measured numbers with assumptions are published instead.
|
||||
- You push back on the other agents when a design makes miner life harder, and you accept their answer when the security case is real.
|
||||
|
||||
## Writing rules
|
||||
No em dashes. Short sentences. Numbers in tables. The project is called Igneum. Approximate figures say so.
|
||||
14
.gitignore
vendored
Normal file
14
.gitignore
vendored
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
.DS_Store
|
||||
# built binaries
|
||||
proto-metal/igneum-bench
|
||||
proto-cuda/igneum-bench-cuda-*
|
||||
proto-cuda/*.exe
|
||||
proto-cuda/*.obj
|
||||
proto-cuda/*.o
|
||||
proto-cuda/emu/*.o
|
||||
proto-cuda/emu/igneum-emu*
|
||||
# vendor clones live outside version control
|
||||
vendor/
|
||||
# secrets never live here
|
||||
*.env
|
||||
.env*
|
||||
146
CLAUDE.md
Normal file
146
CLAUDE.md
Normal file
|
|
@ -0,0 +1,146 @@
|
|||
# Igneum
|
||||
|
||||
A GPU-mined layer 1 whose miners are also the ZK provers for its own zkEVM and for other chains.
|
||||
The founder's project, started 3 October 2026. Named Igneum (Latin: fiery, proven by fire) on 3 October 2026.
|
||||
Repo, Vercel project and Neon database are all "igneum".
|
||||
|
||||
## Source of truth
|
||||
The design document (Claude Doc "Igneum", formerly "GPU Prover Network") plus the public litepaper
|
||||
(https://claude.ai/code/artifact/edcf47f6-e918-48c7-95c3-b98e0d65583f) and the brand canvas (https://claude.ai/artifact/CPixnvctTKbrwBwWck3iFj):
|
||||
https://claude.ai/code/artifact/08fa9e06-9240-4964-96f4-76ff1d2f796c
|
||||
Read it with the docs tools before any design discussion. Decisions already taken live there
|
||||
(hard cap 4 billion halving every two years, 80/20 lottery/proving split, NO emission treasury, NO dev fund (removed 3 Oct 2026: the protocol carries no fee to any team, foundation or fund;
|
||||
priority fee 80/20 miners-and-provers/app, external jobs 90/10 provers/burn, 60% signalling kept only for parameters genesis leaves to miners), no stake, miner-only self-contained finality (NO Bitcoin anchoring), 1 block/s at
|
||||
launch, SP1 behind a swappable proving interface, Taiko as first proving customer, 30-day launch ramp).
|
||||
|
||||
## The design in one paragraph
|
||||
Leader election by a random-program GPU hash (RandomX idea rebuilt for GPUs, new kernel each ~1h
|
||||
epoch, era parameters drawn automatically every 6 months from chain state inside rules fixed at genesis, NO scheduled human releases; dataset grows on a genesis-fixed schedule and instruction families unlock by height from a genesis reserve (automatic schedule changes against fixed datapaths and human forks; against a chip that stores the dataset every drawn parameter is firmware, and the defence is the latency-shadow work of class v4 and the price per joule, per the Horizon algorithm lane, 6 October 2026); only writing new code is human, via miner signalling (the three thresholds, one sentence everywhere: 60% for a parameter, 90% for an upgrade, 95% with a floor height for a class change, the P2 mechanism), never required; warp-unit CPU-verifiable). Ordering by a GHOSTDAG
|
||||
BlockDAG forked from rusty-kaspa. Execution by a zkEVM. Every block ZK-proven by miners in chunks,
|
||||
aggregated 20 to 60 s behind the tip at launch (target under 10 s as provers improve). Finality is miner-only and self-contained:
|
||||
FINALITY RULE V2 (review round 2, 3 Oct 2026): vote weight = blue blocks per BLS vote key (in header) over a flat 30-day DAA window, no damping (the 2x cap was Sybil-void), dust threshold 100 blocks; every voter signs every 30-s checkpoint (VRF picks 8 aggregators only); lock = 2/3 of ALL 30-day weight (DECIDED 4 Oct 2026, O-3.15: the floor was raised from 56.7% of total to 2/3 of total, which makes the 2/3-of-active test implied; finality pauses whenever under 2/3 of the window is connected and signing, and the node reports it; the old floor 0.85 x 2/3 had been added after sim v2 showed the bare active denominator locks both sides of a 50/50 partition after 60 min; with the floor: 0 conflicting locks in every partition and eclipse scenario); equivocation evidence strips weight 30 days; fork choice = GHOSTDAG among tips through all certified checkpoints under Kaspa merge-depth 3,600 s; NO hidden-block n^2 penalty (removed, breaks DAG determinism); epoch seed = 10-min class-group VDF of a certified checkpoint, era draw = 1-h VDF; dataset = 256 MB RandomX-style cache, 8 dependent reads per item (not closed-form, not 64 MB); fees: base fee burned in full on both gas dimensions, priority fee 80/20 (miners and provers / app, attributed per call frame, unregistered share burned), external jobs 90/10 (provers / burn), no dev fund and no fee to any team. Headline: 10 days of 100% hashrate to reach 1/3 of weight, 20 days for 2/3; 51% never reaches 2/3 while honest miners stay. NO stake and NO other chain anywhere in consensus (Bitcoin anchoring was
|
||||
considered on 3 Oct 2026 and REJECTED by the founder: no reliance on Bitcoin). The same prover network sells proofs to rollups and bridges, priced in dollars, settled
|
||||
in the token. The lottery and the proving are kept separate on purpose (Aleo lesson).
|
||||
|
||||
## The team, as agents
|
||||
Each role from the design doc's build plan is an agent in `.claude/agents/`. Ask them by name.
|
||||
- `cryptographer`: lottery hash, chunked proving protocol, proof systems, finality on a DAG. Owns gates 1 and 3.
|
||||
- `consensus-engineer`: Rust, rusty-kaspa fork, GHOSTDAG, difficulty, P2P, the hash swap. Owns gate 2.
|
||||
- `execution-engineer`: Rust, Ethereum clients, zkEVM integration, SP1, the proving interface, external job market.
|
||||
- `miner-community-lead`: miner software, pools, launch ramp, what GPU miners will and will not accept. Owns gate 4.
|
||||
They review each other's work. A claim about another chain must cite the repo and file, or be labelled approximate.
|
||||
|
||||
## Rules that apply to every agent and every file here
|
||||
- The founder's copy law: no em dashes, no two-beat antithesis, no aphorisms. Short sentences. Numbers in tables.
|
||||
- Figures from memory are labelled approximate. Figures from a source cite the source.
|
||||
- The project is Igneum. Never claim an ASIC gain, a proving time or a market size without a measurement or a citation.
|
||||
- Vendor source lives in `vendor/` (git clones of rusty-kaspa, RandomX, SP1, decred dcrd, etc.). Read real code before describing it.
|
||||
- Nothing here is a token sale and nothing should become one.
|
||||
- Transactions are public, like Ethereum. Privacy features and shielded pools were considered and REJECTED by the founder on 3 October 2026. Do not add them.
|
||||
|
||||
## Infrastructure (no URLs yet, by the founder's instruction)
|
||||
- GitHub: https://github.com/igneum-network/igneum (organisation igneum-network, created 3 Oct 2026; owners igneum-labs (the founder's igneum.network login, renamed 5 October 2026; the noreply id 337424239 is unchanged) and the second owner login; private, branch master). Commit as igneum-labs <337424239+igneum-labs@users.noreply.github.com> (standing rule 3 Oct 2026: the organisation owner is never publicly visible, so no personal name or email in the history; the commits before the rename carry the login's earlier spelling and the fresh-repository step rewrites them, tools/repo/fresh-repo.sh). Push only when the founder asks. gh on this Mac still stores the token under the login's pre-rename spelling: that name is in ~/.config/igneum/gh-user (never in the repository; `gh auth status` lists it), and the scripts that need it read that file. This checkout's `.git/config` (shared by every worktree) resets `credential.helper` and then adds one helper that returns username igneum-labs with that token, so git pushes from here work whichever gh account is active. Before any gh call: `gh auth status` must show the stored entry as the ACTIVE account (`gh auth switch --user $(cat ~/.config/igneum/gh-user)` if not); the second owner login on this Mac belongs to other projects and has no access to this repository.
|
||||
- Vercel project: igneum in the other business's team, because Vercel refuses the personal account as a scope. Move it to its own team when the venture is named. For the site, explorer and job-market API later.
|
||||
- Neon database: igneum, id soft-voice-31914738, London (aws-eu-west-2), Postgres 17. Connection string in ~/.config/igneum/env, never in the repo.
|
||||
|
||||
## Domains (purchased by the founder, 3 October 2026, registrar with "Full Protection"; nothing pointed yet)
|
||||
igneum.net, igneum.io, igneum.network, igneum.info, igneum.store, igneum.online, igneum.app, igneum.co.uk, igneum.xyz.
|
||||
Also bought 3 Oct 2026 (the founder: "bought all the other domains"): igneum.org and igneum.com, so the full set is held.
|
||||
Applied 3 Oct 2026: all 15 Igneum domains at GoDaddy (also .art, .email, .pro, .shop, .vip) on Vercel nameservers (ns1/ns2.vercel-dns.com) and attached to the Vercel project igneum: igneum.network is the primary site, every other domain 308-redirects to it. igneum.com is still on Afternic nameservers (purchase in transit), attach when it arrives. Site source: site/ (static). Since the evening of 3 Oct 2026 the live site is the `igneum` project in the `igneum` team (igneum-labs login, `~/.config/igneum/vercel`), deployed by the GitHub integration on every push to master; the other business's project of the same name only holds the 13 redirect domains. Link, env and hand-deploy commands: packaging/README-ship.md (checked 4 Oct 2026).
|
||||
|
||||
## Browser profile (standing rule, 3 October 2026)
|
||||
Every Igneum task in Claude in Chrome runs in the igneum.network Chrome profile. The founder also has profiles for other brands. Confirm the active profile before acting; if it is the wrong one, switch or stop and say so. Never carry Igneum sign-ins, posts or purchases through another brand's profile.
|
||||
|
||||
## Running agents on this Mac (standing rule, 4 October 2026; builds moved to igneum-build-1 on 6 October 2026)
|
||||
- Code and documents in parallel; the Mac's own BUILDS and MEASUREMENTS one at a time through `tools/lock/with-lock.sh build|measure|run <cmd>`
|
||||
(`measure` blocks builds too: hash rate, latency in ms, power; `run` is for functional runs such as test networks, attack
|
||||
harnesses and simulators whose outputs are counts, locks, forks or seconds, and lets builds continue; use the MAIN
|
||||
checkout's script, /Users/joshm/Projects/igneum/tools/lock/with-lock.sh, from any worktree). A number taken while
|
||||
another build or simulation ran is not a number. `~/.config/igneum/build-slots` (1 to 3) caps how many Mac builds run at once.
|
||||
- BUILDS GO TO THE BOX (standing rule, 6 October 2026, main's decision after the measurements in docs/plans/build-server.md):
|
||||
every Linux and Windows cargo build and every Linux test suite from every agent runs on igneum-build-1 (Hetzner AX162,
|
||||
96 threads, 128 GB, `ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49`) through `tools/build-remote.sh [-- cargo args]`
|
||||
and `tools/cross-remote.sh` from the crate directory of the agent's own worktree. Measured 6 October: clean node build
|
||||
1 min 27 s (Mac 12 to 18 min), incremental 7 s (Mac 2 to 15 min), Windows cross 1 min 44 s (Mac 4 min 49 s to 12 min 28 s).
|
||||
The box has its own slot files (`/srv/builds/_locks/build-<k>`, count in `/srv/builds/_locks/slots`, 1 today); a remote
|
||||
build takes one of those, never a Mac slot. Sources travel as HEAD through the bare mirrors `/srv/igneum.git` and
|
||||
`/srv/igneum-node.git` plus an rsync overlay of uncommitted changes (re-stamped); `/srv/builds/<worktree>` mirrors the
|
||||
worktree root; artefacts come back into `<crate>/target-remote/`, never `target/` (they are x86_64 Linux and Windows
|
||||
binaries). Every run writes one line to `/srv/builds/_log/builds.jsonl` (the worker dashboard reads it); `IGNEUM_AGENT=<name>`
|
||||
tags it. The PCs keep only jobs that need their GPUs or the Windows runtime (measurements, Windows test suites, installer
|
||||
smoke runs: `node tools/build-job.mjs run --target ae432dc7|1ccfe586 ...`, PC 1 = ae432dc7, PC 2 = 1ccfe586). The Mac
|
||||
keeps macOS binaries, the DMG and Metal tests, under the build lock. THE MAC RUNS NOTHING THE NETWORK DEPENDS ON (the founder,
|
||||
6 October 2026, evening): after the 0.3.15 cut the two devnet hands, node 1 and the observer (its node and
|
||||
tools/observer), run on the box as systemd units (`igneum-node1`, `igneum-observer-node`, `igneum-observer`;
|
||||
docs/plans/hands-on-build-1.md; `infra/build-server/hands/`), node 1's p2p on 26611 with `--externalip`, every RPC on
|
||||
loopback, the observer's Neon string in `/srv/observer/env` (mode 600, copied from the Mac, never in the repo); that file
|
||||
is the only secret the box holds. The Devnet 2 seed, when it starts, gets its own unit the same way. Setup and re-provision:
|
||||
`infra/build-server/run-from-mac.sh <ip>` (idempotent); the rustc pin is `RUST_TOOLCHAIN` in `infra/build-server/provision.sh`
|
||||
(1.99.0; there is no rust-toolchain file, add one) and build-remote.sh refuses a version mismatch. Nobody deletes another
|
||||
worktree's target dir on the box. Windows exes are reproducible (`-Wl,--no-insert-timestamp` in cross-remote.sh, the Mac's
|
||||
cross-build.sh and the PC job alike); a node binary whose strings lack its commit fails `tools/ci/commit-string-check.sh`
|
||||
(the empty-commit class, 6 October 2026: kaspa-build-info embeds the hash only from a `.git` directory on a branch and
|
||||
never re-runs once empty, so every Mac worktree build and every PC job build had shipped without it; build-remote.sh,
|
||||
cross-remote.sh, cross-build.sh and the PC job now carry the two-step: a branch or minimal `.git`, and
|
||||
`cargo clean --release -p kaspa-build-info` on a new commit).
|
||||
- Every agent works in its own git worktree (`git worktree add ../igneum-wt-<name>`), never the shared checkout, and
|
||||
stages only its own files. Mac builds (macOS binaries only): `nice -n 19`, at most 4 cargo jobs.
|
||||
- Never launch /Applications/Google Chrome.app headless (it blocks the owner's Chrome); use the built-in browser pane.
|
||||
- Keep the Mac on mains with the 140 W charger; on 4 October it hibernated at 1% battery and took node 1 and the observer down.
|
||||
- A source tree copied to another machine (rsync, zip, tar, scp) is re-stamped with `touch` before anything builds it,
|
||||
because cargo rebuilds by mtime and the far side keeps its target dir (the stale-build class: shard run 2 on
|
||||
4 October, the 0.3.6 PC build on 5 October). `tools/ci/copied-sources-check.sh` fails CI on any script that copies
|
||||
and builds without it. When a bug is fixed, fix its CLASS: grep for every other script with the same shape the same
|
||||
day, and add a check that fails when the shape comes back (the founder, 5 October 2026: "we should not be having same bugs
|
||||
repeated").
|
||||
- A watcher or gate is trusted only after it has been shown to fire on one known-finished and one known-failed case
|
||||
(4 October 2026: three job watchers waited for a line the closing report never starts with, and a failed shard run
|
||||
reported exit 0; the failure was found by hand half an hour later). Every job's outcome and its error lines are read
|
||||
the moment it ends, done or not.
|
||||
|
||||
## Every number carries its consequences (standing rule, 5 October 2026, 21:30 UTC)
|
||||
The founder: "this question and answer should have not needed to be asked ... so that suggestions like this get made without me" (the 12 GB mine-and-prove question, which followed from the 15.6 GB measurement and should have been raised by the agent that measured it). Rule: an agent that measures or reports a number also states, in the same report, what the number means for each user tier and what it will do about it, before anyone asks. The tiers: a home miner with one 8 GB card, one 12 GB card, one 16 GB card, one 24 or 32 GB card; a rig; a pool user; each on Windows, Linux and macOS; each on NVIDIA, AMD and Apple (Intel when it exists). A report that says "peak 15.6 GB" without "so 12 GB cards cannot do both; here is the profile that fits them, measuring now" is incomplete and goes back. The same for hash rates (per watt and per pound for the tiers), times (what deadline it fits), sizes (what disk or memory it needs), and prices. A standing reviewer agent reads every status file, bench-log entry and plan for missed consequences and opens the work; the coordinator does not wait for the founder to notice.
|
||||
|
||||
## A job never quits or restarts the installed app it did not start (standing rule, 5 October 2026, 23:05 UTC)
|
||||
Ember Tune's PC 1 playbook started a second engine, that engine's own updater saw itself as 0.3.9 and launched the per-user installer, and the installer's stop step POSTed `/api/quit` to the installed app, taking PC 1 off the network from 22:31Z (141 MH/s gone, the 0.3.11 Windows build blocked) with nobody awake to relaunch (corrected 6 October 2026 from the collected log; the playbook's own quit never fired and pointed at its scratch engine; fixed at e600e63: a second engine never runs the updater). Rule: a test engine started by a job runs on its own port and data dir with its own URL file, and a job may quit, pause, resume or restart only an engine it started itself (the URL it created); the installed app is touched only through the signed `restart` and `update-now` job kinds, and a job that needs the installed app's miners out of the way uses the runner's `--stop-miners` (the runner stops them before the script and restarts them on any exit), never `/api/pause` or `/api/resume` from the script, not even with a finally block (ruling 6 October 2026: a script that dies before its finally leaves the box paused unattended). `tools/ci` carries a check (`playbook-quit-check`) that fails a playbook which reads `%LOCALAPPDATA%\igneum\app\app.url` or `~/Library/Application Support/Igneum/app/app.url` and sends quit, pause or resume to it; the reviewer's row C35 is the record.
|
||||
|
||||
## Devnet 2 gate and activation rules (standing rule, 6 October 2026, 16:4x UTC)
|
||||
|
||||
The founder's ruling after the DAA 198,000 incident (a fixed-height activation crossed while the fleet was still updating: a two-sided chain, a 229-block reorg, execution reset to genesis on every node, proving at zero for an hour): releases stay hourly; what changes is where they land first.
|
||||
- Devnet 2 is the rented fleet as its own staging chain (own genesis and network id, refused by live peers at the handshake). Every release, activation and tuning kit crosses Devnet 2 through `tools/fleet/devnet2-gate.sh` (PASS = zero rejected blocks across the activation, no reorg over depth 3, exec roots agreeing on every box, a segment record paid, every node on the new version) before the live devnet or any of the founder's machines sees it. The shipper does not build the live object before the PASS line.
|
||||
- No fixed-height activation on the live devnet. A consensus change flips when 95 percent of mining weight over a window signals the new object, with a floor height as the backstop; the class v4 cut is the first to carry it (status file gates P1 and P2).
|
||||
- No "clean day" waits (the founder, 6 October 2026, 19:3x UK: "we dont need a clean day for anything this is just delaying things"): a cut's only gate is the Devnet 2 crossing (the class v4 cut's gate is its P1 rehearsal on the fleet chain); digest moves are bundled into one cut (0.3.15 = class v4 + the fourteenth field), miners first, hands last.
|
||||
- A deep reorg never resets execution; a snapshot a node cannot read fails loudly; the p2p snapshot path refuses a snapshot below the node's tip or the restart; a hands script never pgreps its own command line.
|
||||
- Main checks an agent's number against the log or the chain before relaying it to the founder, or labels it unverified.
|
||||
- Never remove more than 10 percent of the live devnet's 30-day vote weight in any hour (Horizon finality lane, 6 October 2026: the class v4 rehearsal took 13 fleet keys off the live chain at 18:27Z; with seven earlier leavers that was 42.7 percent of the voter table frozen at the last lock, and rule v3 holds the pause for a full window; under v3 a sudden departure of a third of weight is a 30-day pause on mainnet). Experiments that borrow live miners do it in slices with an hour between. The protocol fix is the signed LEAVE item (0.3.16).
|
||||
- Every NODE cut's Devnet 2 gate includes a MINING new node beside an old node on the live file for ten minutes (the old node accepting the new node's blocks, the hub's reject count unchanged), and a new node restarted mid-window re-syncing from an old peer. Found 6 October 2026, 20:5x UK: 0.3.15's node stamped the class v4 signal bit into the block version (1026) on the thirteen-field file; every 0.3.14 node rejected it; the digest-compat test passed because the handshake peers while the block version splits; the canary caught it before any live box moved.
|
||||
|
||||
- Standing fleet (the founder, 6 October 2026, 20:0x UK: "cant we keep rented cards up longer"): 16 live-devnet boxes and 6 Devnet 2 boxes are kept up permanently and re-rented on host death; only benchmark and wave boxes are one-shot; a standing box never leaves the live devnet for an experiment (the class v4 rehearsal took the 15 prover boxes and paused live finality at 18:42Z; experiments use wave boxes only). The Mac runs nothing the network depends on: node 1 and the observer move to igneum-build-1 as systemd units (docs/plans/hands-on-build-1.md).
|
||||
- Secrets on igneum-build-1: none that sign releases, move funds or reach the hands. Exception recorded 6 October 2026: Discord webhook URLs at /srv/discord-hooks/env (mode 600, rotatable in one click) for tools/community/discord-hooks.mjs.
|
||||
|
||||
## Shared scratchpad and lane ownership (standing rules, 7 October 2026, 08:2x UK)
|
||||
|
||||
- A lane never removes a directory under the shared scratchpad; it removes only files it wrote inside one. Each lane keeps its scratch under its own prefix (`bs-<name>` for the build-server lane, `r03xx-ship` for the shipper, and so on). At 07:23 and 07:56 UK the build-server lane removed `r0317` and `r0318` as superseded and the shipper's running 0.3.18 chain lost its log and staging copy mid-build.
|
||||
- Every wait keys on its own run's marker, never on a generic line in a shared log (the ten-member pool window mis-fired at 07:44 UK on the 0.3.17 cases' end marker). A presence check never shares a shell with the command it checks (the fleet's installer answered "already running" to its own command line and left pool-1's node down seven minutes).
|
||||
- A supervised node restart read back by a new lock line from the hub is not a weight removal under the 10 percent rule; a key that stops voting is. The fleet's table gate reads the folded certificate weight over the frozen table (mean of the last five checkpoints), not the lock-moment share.
|
||||
|
||||
## Nothing heavy on the Mac (standing rule, 7 October 2026, 10:5x UK)
|
||||
The Mac crashed and rebooted under agent load (about twenty lanes, local cargo tests and guest builds, several headless chromiums for captures and Playwright suites at once). The founder: "stop building on it unless its 100% necessary to build on a mac". Rule: the only Mac builds are the shipper's macOS binaries and the DMG, one at a time under the build lock. No cargo build or test on the Mac for any other lane; no guest builds, z3 or census runs; no Playwright suites; captures use one headless chromium at a time and never from several lanes at once. Every Linux and Windows build, every test suite and every benchmark runs on igneum-build-1 (`tools/build-remote.sh`) or as a PC job. The pre-push gate script stays (checks, not a build). Main keeps the number of concurrently running lanes near ten, the founder's waiting items first.
|
||||
|
||||
## CI red is stop-the-line (standing rule, 6 October 2026, 22:0x UK)
|
||||
- Whoever's merge turns master or a release-* branch red owns the fix inside 15 minutes or reverts the merge; the red watcher posts every failed run to the hidden updates channel and to /srv/ci-red/red.jsonl on the box (tools/ci/red-watch.mjs); the box's own red builds land in the same file with a class (remote-run.sh pre-flight, kept run logs) and the 09:00 UK digest counts them per class with each class's guard.
|
||||
- The pre-push gate is the same script CI runs: `tools/ci/pre-push.sh` (installed by `tools/ci/install-hooks.sh`; `--hook` before a push to master or release-*, `--ci` in the workflow). A check is added there, never only in ci.yml. On a feature branch the hook runs the light gate: the two structural checks plus the two never-push classes, the no-secrets check and the identity grep (7 October 2026).
|
||||
- After every push, the pushing lane reads the conclusion when it lands (`gh run list --branch <branch> --limit 1`; `gh run view <id> --log-failed` on a red) and owns a red before the next push; pushes are never held for it. A red on any branch is reported to main the moment it is seen. The red watcher (.github/workflows/ci-red.yml, from master, every branch) posts each failed run with its branch, commit, red check and pushing author. Record: 7 October 2026, 11:26 to 12:47 UK, eight red runs on ca3-v4-node over three gate summaries carrying a 64-hex key, read by nobody; 31 runs queued on one runner at 13:15 UK, so igneum-build-2 joined the pool and docs-only pushes skip the compile jobs.
|
||||
- A code push (anything outside docs/, site/ and *.md) goes out only after that exact commit's crate suite is green on a build box (the suite's RESULT line in the push's own notes or the commit message); a test build the box never ran is not pushed. Record: 7 October 2026, 14:10 to 14:13 UK, four red runs in three minutes (a test-build compile error on ca3-v4-amend, the hidden-console check on miner-reliability, the installer's inputs pin on release-0.3.20), each a push ahead of its own box result, each emailed to the founder.
|
||||
- Master takes only what CI has already passed (7 October 2026, 17:2x UK): a merge lands only when the branch's own `ci` run is green on the exact commit being merged, read from the runs API (`tools/ci/ci-state.mjs`), never the local stamp alone; `tools/ci/merge-to-master.sh` pushes the branch for a run when there is none, waits for a queued one printing the clock, refuses a red one and refuses any merge onto a red master except the declared fix (`--fixes-master`); the pre-push hook refuses a push to master whose commit (or merge parent) has no green run; a feature-branch push prints the branch's previous red first. GitHub's branch protection is unavailable on the free plan for a private repository (the API answers 403), so the two scripts are the enforcement.
|
||||
- Research and operations documents live outside the public export list: name them in `tools/ci/export-exclude.txt` (read by the identity check and by the mirror's sync.sh). What stays in the list is read by the public and must pass the identity grep.
|
||||
- A path Windows cannot hold (colon, trailing dot or space, reserved name, over 240 characters) never enters a commit: the pre-commit hook runs `tools/ci/windows-paths-check.sh --staged`.
|
||||
- Record: docs/analysis/ci-failures-2026-10-06.md (168 non-green runs in three days classified; 126 on master; all but two classes were tree checks that the gate now runs locally first).
|
||||
|
||||
## A rule row closes only with its check (standing rule, 6 October 2026, 18:4x UK)
|
||||
|
||||
The founder, after the pgrep self-match hit twice in one day (the shipper's hands script at lunchtime, the fleet's wave script at 17:1xZ: a `pgrep -f "<pattern>"` whose literal sat in the calling shell's own command line, so the check always passed and no node ever started): "again wasted time". Rules:
|
||||
- A bug class found today gets its `tools/ci` check merged on master the same hour, or the rule row stays OPEN and main says so. A rule row without a check is not closed.
|
||||
- Box operations (ssh to a rented box, install a payload, start a node, wait for sync, read height, peers and exec tip, start a miner or prover) live in ONE shared, tested library (`tools/fleet/lib/`), exercised against a Devnet 2 box by a test; agents call it and never write their own copy in bash or PowerShell.
|
||||
- A watcher or collector verifies the chain-side fact (height, peers, exec tip, paid records), never a reported rate or a process name.
|
||||
- `pgrep -f` / `pkill -f` / `ps | grep` with a literal pattern is banned in scripts; use the bracket form `[i]gneumd`, `-x` on the binary name, or a pid file (`pkill -F`, `tools/fleet/fleet-bg.sh`). Kill by exact command line or pid file, never by a name: at 22:09 UK on 6 October a Mac-side `pkill -f <log file name>` matched nothing (a redirect is not on the command line) and the roll-everything script wiped a box it had been told to hold (CI check and gate: `tools/ci/kill-by-name-check.sh`, which also flags a file-name shape under pgrep/pkill).
|
||||
60
docs/bench-log.md
Normal file
60
docs/bench-log.md
Normal file
|
|
@ -0,0 +1,60 @@
|
|||
# Igneum bench log
|
||||
|
||||
Append-only. Every number here was measured on the machine named, on the date given.
|
||||
|
||||
## 2026-10-03 proto-metal / igneum-bench, first run
|
||||
|
||||
Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, macOS Darwin 25.6.0, Swift 5.8.1, Metal 4.
|
||||
Build: `swiftc -O -o igneum-bench main.swift -framework Metal`. Source: `proto-metal/main.swift`.
|
||||
Setup: 1 GiB dataset (2^28 uint32), 64 instructions x 8 iterations, threadgroup 32 (threadExecutionWidth 32),
|
||||
4 timed batches x 2^22 nonces after one warm-up batch. Verification: 3 warps per program, CPU interpreter vs GPU.
|
||||
|
||||
| Seed | Loads/hash | Compile ms | Mhash/s | GB/s useful | CPU verify ms/warp | Verify |
|
||||
|---|---|---|---|---|---|---|
|
||||
| igneum-genesis | 104 | 49.6 (cold) | 45.2 | 18.8 | 0.015 | PASS |
|
||||
| igneum-genesis/epoch1 | 104 | 20.3 | 48.4 | 20.1 | 0.019 | PASS |
|
||||
| igneum-second-seed | 104 | 46.6 (cold) | 35.5 | 14.8 | 0.016 | PASS |
|
||||
| igneum-second-seed/epoch1 | 144 | 18.7 | 35.4 | 20.4 | 0.017 | PASS |
|
||||
| igneum-hourly | 128 | 52.0 (cold) | 36.6 | 18.7 | 0.021 | PASS |
|
||||
| igneum-hourly/epoch1 | 128 | 21.6 | 37.5 | 19.2 | 0.017 | PASS |
|
||||
| igneum-hourly/epoch2 | 120 | 23.7 | 36.6 | 17.6 | 0.016 | PASS |
|
||||
|
||||
Dataset size sweep (seed igneum-genesis): 4 MiB 569 Mhash/s, 64 MiB 183, 256 MiB 94, 512 MiB 69, 1 GiB 44.
|
||||
Dataset fill 1 GiB: 2.34 ms GPU time (427 GB/s) warm, 5.57 ms on first run of a process.
|
||||
Result: 21 warps, 672 hashes, zero mismatches. OVERALL PASS.
|
||||
Reading: memory bound at 1 GiB (12.8x drop from cache-resident), limited by random access rather than bandwidth,
|
||||
CPU verify roughly 250x under the 10 ms gate with a cheap dataset element. Apple silicon only. Details in
|
||||
`proto-metal/README.md`.
|
||||
|
||||
## 2026-10-03 proto-cuda / program pack export (Mac side only; RTX 5090 run pending)
|
||||
|
||||
Machine: the same Apple M5 Max. No CUDA toolchain exists on it, so nothing below is an NVIDIA measurement.
|
||||
Added `--export-pack <dir>` to `proto-metal/main.swift`. It writes, per seed, the CUDA kernel (`kernel.cu`),
|
||||
C headers (`program.h`, `vectors.h`), JSON twins, and the Metal source, into `proto-cuda/packs/<seed>/`.
|
||||
Vectors: 3 warps (base nonces 0, 4096, 1000000), 96 x 64-bit outputs from the CPU interpreter, plus dataset
|
||||
words 0..15 and word [MASK]. The exporter runs the Metal kernel for the same warps and refuses to write unless
|
||||
all 96 match.
|
||||
|
||||
| Pack | Loads/hash | Op mix | CPU interpreter vs Metal GPU (3 warps) | CUDA text in CPU emulation (clang, 32 threads/warp) |
|
||||
|---|---|---|---|---|
|
||||
| igneum-genesis | 104 | load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1 | PASS 3/3 | PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 and 2 warps/block |
|
||||
| igneum-hourly | 128 | load=16 add=8 xor=7 mad=5 mul=5 mulhi=5 shfl=5 rotl=4 or=3 rotr=3 sub=3 | PASS 3/3 | PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 warp/block |
|
||||
|
||||
Sweep sizes 4, 64, 256, 512 MiB in the emulation: dataset self-test PASS at each (vectors only apply at 1 GiB).
|
||||
The regular Metal bench was re-run after the change: igneum-genesis 44.8 Mhash/s and igneum-hourly 36.3 Mhash/s
|
||||
at 1 GiB with 2 x 2^20 batches, PASS 3/3 warps each (consistent with the first-run table above, not a new figure).
|
||||
Harness: `proto-cuda/host.cu`, `build.sh`, `build.bat`, `README.md`, `CHECKLIST.md`, `emu/`.
|
||||
Pending: build and run on the project lead's RTX 5090 (CUDA 12.8 or newer, `-arch=sm_120`). No NVIDIA hash rate exists yet.
|
||||
Emulation rates are not recorded because they measure the Mac's CPU, not a GPU.
|
||||
|
||||
## 2026-10-03 sim/finality_sim.py, sustained-mining finality vote weight (model, not hardware)
|
||||
|
||||
Model: 1 day steps, 86,400 Poisson blocks/day, 1,000 Pareto honest keys (top key 17%), perfect retarget, every block blue, no latency, no VRF noise. Seed 7, seed 11 agrees.
|
||||
Rule: weight = 30-day sum of counted blocks, counted = min(actual, 2 x yesterday + f). Lock at 2/3 of total. Floors f in {1, 10, 100, 1000}.
|
||||
A: weight tracks hashrate at steady state (corr 1.00000); full weight from zero history on day 41 (f=1) to 32 (f=1000).
|
||||
B: 60% renter on one key takes 59.9% of rewards on day 1, crosses 1/3 of weight on day 20 to 26, 50% on day 27 to 34, never 2/3. 75% renter reaches 2/3 on day 29 to 34, day 27 with the cap removed.
|
||||
C: splitting defeats the cap. 10,000 fresh keys at f=1 move the 50% crossing from day 34 to 27 (no-cap figure 26); at f=10 they match no-cap exactly. The 30-day trickle buys 2 more days for 11.6% of the network.
|
||||
D: honest doubling is under-weighted 28 to 31 days; old miners lock alone for 20 to 23 days.
|
||||
E: 30% churn leaves 71% live, lock never lost. Threshold is 1/3: 35% stalls 1 day, 50% stalls 10 to 11 days. Active-24h total removes every stall.
|
||||
F: 51% patient owner holds 51.0% weight from day 45, vetoes from day 7 to 20, never locks alone. 67% owner locks alone from day 34 to 44.
|
||||
Recommend: f=1, cap 2x, window 30 (the window is the defence, the cap is worth 1 to 8 days), total = all keys with weight in the window. Details in sim/results.md.
|
||||
6
proto-cuda/.gitignore
vendored
Normal file
6
proto-cuda/.gitignore
vendored
Normal file
|
|
@ -0,0 +1,6 @@
|
|||
igneum-bench-cuda-*
|
||||
emu/build-*/
|
||||
*.o
|
||||
*.obj
|
||||
*.exp
|
||||
*.lib
|
||||
93
proto-cuda/CHECKLIST.md
Normal file
93
proto-cuda/CHECKLIST.md
Normal file
|
|
@ -0,0 +1,93 @@
|
|||
# Metal to CUDA equivalence checklist
|
||||
|
||||
Written 3 October 2026 for the program packs in `packs/`. Everything below is about the kernel text that
|
||||
`proto-metal/igneum-bench --export-pack` emits, compared with the Metal text the same run emits and with the
|
||||
CPU interpreter `cpuWarp` in `proto-metal/main.swift`. There is no CUDA toolchain on the Mac, so "verified"
|
||||
means one of the three methods in the last section, never an nvcc build.
|
||||
|
||||
## Types and arithmetic
|
||||
|
||||
| Item | Metal | CUDA | Status |
|
||||
|---|---|---|---|
|
||||
| Register type | `uint` (32-bit unsigned) | `uint32_t` (`unsigned int`, 32-bit on Linux and Windows) | same width, same wraparound |
|
||||
| Output type | `ulong` | `uint64_t` | 64-bit on both; `(uint64_t)hi << 32 \| (uint64_t)lo` |
|
||||
| Literals | `0x...u` | `0x...u` | emitted by the same `hex()` helper |
|
||||
| Floats | none | none | nothing for fast-math or FMA contraction to touch |
|
||||
| Shift amounts | always 0..31 (see rotates) | always 0..31 | no out-of-range shift anywhere |
|
||||
| Evaluation order | each instruction is one statement with one assignment | same statement text | no sequence-point question arises |
|
||||
|
||||
## Register init and framing
|
||||
|
||||
| Item | Metal | CUDA | Status |
|
||||
|---|---|---|---|
|
||||
| Nonce | `baseNonce + gid`, gid = thread_position_in_grid | `baseNonce + blockIdx.x * blockDim.x + threadIdx.x` | identical for blockDim 32; also identical for blockDim 32 x W because the nonce is still base + global index |
|
||||
| Register init | `x = nonce ^ SEEDW[i]; x += 0x9e3779b9u * (i+1)u; x = splitmix32(x); r_i = x ^ SEEDW[(i+1)&7]` | same, with `SEEDW[i]` and the product `0x9e3779b9 * (i+1)` written as literals (computed by the exporter mod 2^32) | verified by vectors |
|
||||
| splitmix32 | 3 shift-xor, 2 multiply | identical text | verified by vectors |
|
||||
| Iteration | `for it in 0..<8 { sel = r0; 64 instructions }` | identical | verified by vectors |
|
||||
| Output mix | `lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21)`, `hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27)` | identical | verified by vectors |
|
||||
|
||||
## Instructions
|
||||
|
||||
| Op | Metal | CUDA | CPU interpreter | Notes |
|
||||
|---|---|---|---|---|
|
||||
| add | `d = d + a + select(imm, imm2, ((sel >> bit) & 1u) != 0u)` | `d = d + a + ((((sel >> bit) & 1u) != 0u) ? imm2 : imm)` | `d &+ a &+ (s != 0 ? imm2 : imm)` | Metal `select(A, B, c)` is `c ? B : A`, so the true branch is `imm2` in all three. `sel` is r0 sampled at the top of the iteration in all three. |
|
||||
| sub | `d = d - a` | same | `d &- a` | wraparound |
|
||||
| mul | `d = d * a` | same | `d &* a` | low 32 bits |
|
||||
| mulhi | `mulhi(d, a)` | `__umulhi(d, a)` | `(UInt64(d) * UInt64(a)) >> 32` | high 32 bits of the unsigned 64-bit product on all three |
|
||||
| xor | `d = d ^ a` | same | same | |
|
||||
| or | `d = d \| a` | same | same | |
|
||||
| rotl (immediate) | `rotl_imm(d, n)` = `(x << n) \| (x >> (32u - n))`, n literal 1..31 | identical text | `rotl32` with n in 1..31 | the generator draws `rot` from 1..31, so neither shift amount is ever 0 or 32 |
|
||||
| rotr (register) | `rotr_var(d, a)` = `n &= 31u; (x >> n) \| (x << ((32u - n) & 31u))` | identical text | `rotr32`: `n & 31`, `n == 0 ? x : ...` | for n = 0 both GPU forms give `x \| x = x`; for n in 1..31 both shifts are in range |
|
||||
| mad | `d = a * b + d` | same | `(a &* r[b]) &+ d` | `b` may equal `d` or `a`; all three read every operand before the single write |
|
||||
| shfl | `d = d ^ simd_shuffle_xor(a, (ushort)mask)` | `d = d ^ __shfl_xor_sync(0xffffffffu, a, mask)` | `r[lane][dst] ^= tmp[lane ^ mask]` where `tmp` is a copy of register `a` across the warp taken before any lane writes | mask in {1,2,4,8,16} so the partner lane is always inside the same 32-lane group. `a != dst` by construction, so there is no read-after-write hazard inside the instruction. Control flow is uniform (no branches at all), so the full member mask is valid and no lane is missing from the shuffle. |
|
||||
| load | `d = d ^ dataset[a & MASK]`, MASK a compile-time literal | `d = d ^ ds[a & mask]`, mask a kernel argument | `d ^ datasetElem(a & mask, d0, d1)` | the index is a 32-bit unsigned value below 2^28, so pointer arithmetic needs no 64-bit care |
|
||||
|
||||
No instruction had to be removed or changed. Every op is an unsigned 32-bit integer operation whose result is
|
||||
defined identically in Metal Shading Language, CUDA C++ and Swift's wrapping operators.
|
||||
|
||||
## Lane mapping (the one place where the two platforms differ in guarantees)
|
||||
|
||||
| Item | Metal | CUDA |
|
||||
|---|---|---|
|
||||
| Group that `shuffle_xor` spans | the SIMD group, width `threadExecutionWidth` (32 on this M5 Max; the tool warns if not) | the warp, width 32 on every NVIDIA GPU (`host.cu` prints `cudaDevAttrWarpSize` and warns if not 32) |
|
||||
| Lane id of a thread | `thread_index_in_simdgroup`; for a 32-wide threadgroup this equalled `thread_position_in_threadgroup`, as shown by the bit-exact match with the CPU model across 21 warps on 3 Oct 2026 and the 3 vector warps per pack today | `threadIdx.x % 32`, guaranteed by the CUDA programming model |
|
||||
| Partner lane | `lane ^ mask` | `lane ^ mask` |
|
||||
| Warps per block | 1 (threadgroup of 32) | `--block-warps W`, default 1. Any W is bit-exact because each warp is an aligned run of 32 consecutive nonces either way. The emulation passed with W = 1 and W = 2. |
|
||||
|
||||
## Dataset
|
||||
|
||||
| Item | Metal | CUDA |
|
||||
|---|---|---|
|
||||
| Element | `ds_elem(i, d0, d1)`: xor, mul 0x9E3779B1, xor-shift 15, add d1, mul 0x85EBCA77, xor-shift 13, mul 0xC2B2AE3D, xor-shift 16 | identical text in `kernel.cu`; a third copy `host_ds_elem` in `host.cu` |
|
||||
| Day words | `seedWords("day/" + day)[0..1]` on the Mac | written into `program.h` as `IGNEUM_DAY0`, `IGNEUM_DAY1` |
|
||||
| Fill | one thread per word, threadgroup 256 | one thread per word, block 256, with an `i < n` guard (a no-op for power-of-two sizes) |
|
||||
| Self-test | not needed on the Mac (CPU and GPU share one process) | head 16 words and word `[MASK]` against values the Mac wrote into `vectors.h`; 64 pseudo-random words against `host_ds_elem` |
|
||||
|
||||
## How each claim above was verified
|
||||
|
||||
1. Mac, Metal GPU against the CPU interpreter: the regular bench (`proto-metal/igneum-bench`) passed 3 warps per
|
||||
program on both seeds again after the exporter was added, and `--export-pack` itself runs the Metal kernel for
|
||||
the three vector warps (base nonces 0, 4096, 1000000) and refuses to write a pack unless all 96 outputs match
|
||||
the interpreter. Both packs were written with "Metal GPU cross-check PASS 3/3 warps".
|
||||
2. Mac, CUDA text executed as C++: `emu/emu.sh` compiles `host.cu` and the generated `kernel.cu` with clang
|
||||
against a shim `cuda_runtime.h` (the launch syntax is the only text rewritten) and runs the kernels on host
|
||||
threads, 32 per warp, with a barrier inside `__shfl_xor_sync`. Result on 3 Oct 2026, both packs, 1 GiB
|
||||
dataset: dataset self-test PASS, 3 of 3 warps PASS standalone, warps 0 and 4096 PASS inside a batch with
|
||||
1 and with 2 warps per block. Sweep sizes 4, 64, 256, 512 MiB: dataset self-test PASS. `-Wall -Wextra` clean.
|
||||
3. Reading: every emitted line was compared by eye against the Metal line for the same instruction index
|
||||
(the comment at the end of each CUDA line carries the index and the op).
|
||||
|
||||
## Residual risks that the Mac cannot remove
|
||||
|
||||
- nvcc never ran on this code. Syntax that clang accepted could still trip nvcc's front end, and the shim's
|
||||
idea of the runtime API could differ from the real header in a detail. The six `cudaDevAttr*` enumerators
|
||||
and the template signatures of `cudaFuncGetAttributes(cudaFuncAttributes*, T* entry)` and
|
||||
`cudaOccupancyMaxActiveBlocksPerMultiprocessor(int*, T func, int blockSize, size_t dynamicSMemSize)` were
|
||||
checked against the NVIDIA CUDA Runtime API documentation on 3 Oct 2026 and match what the code uses.
|
||||
Anything left would be a compile error, not a silent output difference, and a one-line fix in `host.cu` or
|
||||
the three wrapper functions at the bottom of `kernel.cu`.
|
||||
- `__shfl_xor_sync` and `__umulhi` semantics were emulated, not executed on NVIDIA hardware. Both are
|
||||
documented as lane ^ laneMask within a 32-wide warp and the high 32 bits of the unsigned product; the vectors
|
||||
will settle it on the 5090.
|
||||
- `-arch=sm_120` requires CUDA 12.8 or newer. On an older toolkit the build fails at the flag, not at the code.
|
||||
- Performance figures from the emulation are meaningless and were not recorded. No NVIDIA hash rate exists yet.
|
||||
150
proto-cuda/README.md
Normal file
150
proto-cuda/README.md
Normal file
|
|
@ -0,0 +1,150 @@
|
|||
# igneum-bench-cuda (proto-cuda)
|
||||
|
||||
The NVIDIA twin of `proto-metal`. It runs the same random-program proof-of-work kernels on a CUDA GPU and checks
|
||||
them bit for bit against results produced on the Mac.
|
||||
|
||||
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset,
|
||||
checks the GPU against known answers, and times the kernel. Nothing here earns anything.
|
||||
|
||||
Status on 3 October 2026: packs exported and cross-checked on the Mac, CUDA run on the RTX 5090 pending.
|
||||
No NVIDIA hash rate has been measured. Any figure you see for NVIDIA in this repo before the 5090 run is wrong.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
proto-cuda/
|
||||
host.cu host program: device info, dataset fill, self-test, vectors, bench, size sweep
|
||||
build.sh Linux build (nvcc)
|
||||
build.bat Windows build (nvcc + Visual Studio Build Tools)
|
||||
CHECKLIST.md Metal/CUDA equivalence, op by op, and what was verified where
|
||||
packs/<seed>/ one program pack per seed, written by proto-metal/igneum-bench --export-pack
|
||||
kernel.cu the program as a CUDA kernel, plus fill kernel and host launch wrappers
|
||||
program.h seed, day words, dataset size, loads per hash, wrapper declarations
|
||||
vectors.h expected outputs for 3 warps (96 x 64-bit) and dataset self-test values
|
||||
program.json the instruction list and all constants, for any other implementation
|
||||
vectors.json the same vectors as JSON
|
||||
program.metal the Metal source the Mac ran, for diffing by eye
|
||||
emu/ CPU emulation shim: compile and check a pack with plain clang++/g++, no GPU
|
||||
```
|
||||
|
||||
Two packs are checked in: `igneum-genesis` (104 loads per hash) and `igneum-hourly` (128 loads per hash).
|
||||
Both were cross-checked on the Mac's Metal GPU before being written.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Linux
|
||||
- An NVIDIA driver recent enough for the toolkit. For CUDA 12.8 that is the R570 series or newer (approximate,
|
||||
from memory; `nvidia-smi` prints the driver's maximum supported CUDA version in its header).
|
||||
- CUDA Toolkit 12.8 or newer. Blackwell (`sm_120`, RTX 50 series) is not known to older toolkits.
|
||||
- A host compiler the toolkit supports (gcc 11 to 13 for 12.8, approximate).
|
||||
|
||||
Windows
|
||||
- The same driver requirement.
|
||||
- CUDA Toolkit 12.8 or newer. Tick the Visual Studio integration in the installer.
|
||||
- Visual Studio 2022 Build Tools with the "Desktop development with C++" workload. nvcc needs `cl.exe`.
|
||||
- Run `build.bat` from an "x64 Native Tools Command Prompt for VS 2022" so `cl.exe` is on PATH.
|
||||
|
||||
## Build
|
||||
|
||||
Linux:
|
||||
|
||||
```
|
||||
cd proto-cuda
|
||||
./build.sh # pack igneum-genesis, -arch=sm_120
|
||||
./build.sh igneum-hourly # the second pack
|
||||
./build.sh igneum-genesis native # if sm_120 is refused, let nvcc pick the installed GPU
|
||||
```
|
||||
|
||||
Windows (x64 Native Tools Command Prompt):
|
||||
|
||||
```
|
||||
cd proto-cuda
|
||||
build.bat
|
||||
build.bat igneum-hourly
|
||||
build.bat igneum-genesis native
|
||||
```
|
||||
|
||||
Both scripts run this one command (paths adjusted for the pack):
|
||||
|
||||
```
|
||||
nvcc -O3 -std=c++17 -arch=sm_120 -I packs/igneum-genesis -o igneum-bench-cuda-igneum-genesis host.cu packs/igneum-genesis/kernel.cu
|
||||
```
|
||||
|
||||
Notes
|
||||
- `-arch=sm_120` is Blackwell. `-arch=native` (CUDA 11.6 or newer) compiles for whatever GPU is in the
|
||||
machine and is the fallback if the toolkit is too old to know `sm_120` (which means it is too old for a 5090
|
||||
anyway: upgrade).
|
||||
- If nvcc on Windows refuses the Visual Studio version, add `-allow-unsupported-compiler` to the nvcc line.
|
||||
- The host code is plain C++17 and the CUDA runtime API. No NVRTC, no third-party libraries, no JSON parser.
|
||||
The kernel is compiled ahead of time from the pack.
|
||||
|
||||
## Run
|
||||
|
||||
```
|
||||
./igneum-bench-cuda-igneum-genesis # 1 GiB dataset, 5 batches x 2^24 nonces, vectors checked
|
||||
./igneum-bench-cuda-igneum-genesis --sweep # 4, 64, 256, 512, 1024 MiB in sequence (the Mac's sweep)
|
||||
./igneum-bench-cuda-igneum-genesis --block-warps 4 # 4 warps per block instead of 1 (still bit-exact)
|
||||
./igneum-bench-cuda-igneum-hourly # the second program
|
||||
```
|
||||
|
||||
On Windows the binaries are `igneum-bench-cuda-igneum-genesis.exe` and so on.
|
||||
|
||||
Flags: `--dataset-mib N` (power of two, default 1024), `--sweep`, `--batch-log2 B` (default 24),
|
||||
`--batches N` (default 5), `--block-warps W` (default 1, mirrors the Metal run's one SIMD group per threadgroup),
|
||||
`--device D`.
|
||||
|
||||
What it prints, in order:
|
||||
1. GPU name, SM count, memory, clocks, L2, warp size, driver and runtime versions, registers per thread and
|
||||
resident warps per SM for the kernel.
|
||||
2. Dataset fill time (twice; the Mac saw a first-touch cost on the first fill) and write GB/s.
|
||||
3. Dataset self-test: 16 head words and word `[MASK]` against values from the Mac, 64 random words against the
|
||||
host formula.
|
||||
4. Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again
|
||||
read out of the warm-up batch so the bench configuration itself is checked. PASS or FAIL per warp, with the
|
||||
first differing lane printed on FAIL.
|
||||
5. Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s,
|
||||
GB/s useful (loads per hash x 4 bytes x hashes/s, the same definition as the Mac's table).
|
||||
6. A summary table in Markdown and `OVERALL: PASS` or `FAIL`. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.
|
||||
|
||||
Vectors are only checked when the dataset is the pack's size (1024 MiB), because the outputs depend on the
|
||||
address mask. At other sizes the table says "skipped (not pack size)" and only the dataset self-test counts.
|
||||
|
||||
## What PASS means
|
||||
|
||||
- The CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not every word).
|
||||
- For 96 nonces spread across the nonce space, the RTX 5090 produced the same 64-bit outputs as the Mac's CPU
|
||||
interpreter, which had itself matched the Mac's Metal GPU. The random program, the register init, the warp
|
||||
shuffles, the multiply-high and rotates, and the dataset addressing all agree between Apple and NVIDIA.
|
||||
- With `--block-warps W` the in-batch check passing shows that packing W warps per block changed nothing.
|
||||
|
||||
A FAIL with a small number of differing lanes points at a shuffle; a FAIL in every lane points at an arithmetic
|
||||
op or the dataset. Send the whole printout either way.
|
||||
|
||||
## Sending results back
|
||||
|
||||
Copy the printed header lines (GPU, CUDA versions, kernel line, program line) and the summary table into
|
||||
`docs/bench-log.md` under a dated heading, together with the output of `nvcc --version` and the driver version
|
||||
from `nvidia-smi`. Keep the full stdout as well. Run both packs and the sweep so the log has the same shape as
|
||||
the Mac's entry. Do not edit the numbers; if a run looks odd, run it again and log both.
|
||||
|
||||
## Checking a pack without a GPU
|
||||
|
||||
`emu/emu.sh <pack> [flags]` compiles `host.cu` and the pack's `kernel.cu` as plain C++17 against a shim
|
||||
`cuda_runtime.h` and runs the kernels on host threads (32 per warp, a barrier inside `__shfl_xor_sync`). Use
|
||||
small batches (`--batch-log2 13 --batches 1`). Only PASS/FAIL matters; the rates it prints are noise.
|
||||
This is how the CUDA text was checked on the Mac on 3 October 2026 (both packs PASS, see `CHECKLIST.md`).
|
||||
It is not an nvcc build and says nothing about NVIDIA hardware.
|
||||
|
||||
## Regenerating a pack
|
||||
|
||||
On the Mac:
|
||||
|
||||
```
|
||||
cd proto-metal
|
||||
swiftc -O -o igneum-bench main.swift -framework Metal
|
||||
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
|
||||
```
|
||||
|
||||
The exporter runs the CPU interpreter for the three vector warps, runs the Metal kernel for the same warps,
|
||||
and refuses to write anything unless all 96 outputs match. `--day` and `--dataset-log2` change the dataset
|
||||
constants and are recorded in the pack.
|
||||
39
proto-cuda/build.bat
Normal file
39
proto-cuda/build.bat
Normal file
|
|
@ -0,0 +1,39 @@
|
|||
@echo off
|
||||
rem Build igneum-bench-cuda for one program pack (Windows).
|
||||
rem Run from an "x64 Native Tools Command Prompt for VS 2022" so cl.exe is on PATH,
|
||||
rem with CUDA Toolkit 12.8 or newer installed (nvcc on PATH).
|
||||
rem Usage: build.bat [pack] [arch]
|
||||
rem pack directory name under packs\ (default igneum-genesis)
|
||||
rem arch nvcc -arch value (default sm_120 = RTX 50 series / Blackwell; "native" picks the installed GPU)
|
||||
setlocal
|
||||
cd /d "%~dp0"
|
||||
|
||||
set PACK=%1
|
||||
if "%PACK%"=="" set PACK=igneum-genesis
|
||||
set ARCH=%2
|
||||
if "%ARCH%"=="" set ARCH=sm_120
|
||||
|
||||
where cl >nul 2>nul
|
||||
if errorlevel 1 (
|
||||
echo cl.exe not found. Open an "x64 Native Tools Command Prompt for VS 2022" ^(Visual Studio Build Tools, Desktop development with C++^) and run this again.
|
||||
exit /b 1
|
||||
)
|
||||
where nvcc >nul 2>nul
|
||||
if errorlevel 1 (
|
||||
echo nvcc not found. Install CUDA Toolkit 12.8 or newer; the installer puts nvcc on PATH.
|
||||
exit /b 1
|
||||
)
|
||||
if not exist "packs\%PACK%\kernel.cu" (
|
||||
echo no pack at packs\%PACK% ^(expected kernel.cu, program.h, vectors.h^)
|
||||
exit /b 1
|
||||
)
|
||||
|
||||
nvcc --version
|
||||
echo nvcc -O3 -std=c++17 -arch=%ARCH% -I "packs\%PACK%" -o "igneum-bench-cuda-%PACK%.exe" host.cu "packs\%PACK%\kernel.cu"
|
||||
nvcc -O3 -std=c++17 -arch=%ARCH% -I "packs\%PACK%" -o "igneum-bench-cuda-%PACK%.exe" host.cu "packs\%PACK%\kernel.cu"
|
||||
if errorlevel 1 (
|
||||
echo build failed. If nvcc complains about an unsupported host compiler version, add -allow-unsupported-compiler to the command above.
|
||||
exit /b 1
|
||||
)
|
||||
echo built igneum-bench-cuda-%PACK%.exe
|
||||
endlocal
|
||||
26
proto-cuda/build.sh
Executable file
26
proto-cuda/build.sh
Executable file
|
|
@ -0,0 +1,26 @@
|
|||
#!/usr/bin/env bash
|
||||
# Build igneum-bench-cuda for one program pack (Linux).
|
||||
# Usage: ./build.sh [pack] [arch]
|
||||
# pack directory name under packs/ (default igneum-genesis)
|
||||
# arch nvcc -arch value (default sm_120 = RTX 50 series / Blackwell, needs CUDA 12.8 or newer;
|
||||
# use "native" to let nvcc pick the installed GPU's architecture)
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")"
|
||||
|
||||
PACK="${1:-igneum-genesis}"
|
||||
ARCH="${2:-sm_120}"
|
||||
|
||||
if ! command -v nvcc >/dev/null 2>&1; then
|
||||
echo "nvcc not found. Install CUDA Toolkit 12.8 or newer and put its bin/ on PATH (usually /usr/local/cuda/bin)." >&2
|
||||
exit 1
|
||||
fi
|
||||
if [ ! -f "packs/$PACK/kernel.cu" ]; then
|
||||
echo "no pack at packs/$PACK (expected kernel.cu, program.h, vectors.h)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
nvcc --version | tail -n 2
|
||||
set -x
|
||||
nvcc -O3 -std=c++17 -arch="$ARCH" -I "packs/$PACK" -o "igneum-bench-cuda-$PACK" host.cu "packs/$PACK/kernel.cu"
|
||||
set +x
|
||||
echo "built ./igneum-bench-cuda-$PACK"
|
||||
93
proto-cuda/emu/cuda_runtime.h
Normal file
93
proto-cuda/emu/cuda_runtime.h
Normal file
|
|
@ -0,0 +1,93 @@
|
|||
// CPU emulation shim of the small slice of the CUDA runtime API that host.cu and the generated
|
||||
// kernel.cu use. Compile-and-semantics check only, on a Mac without nvcc. Not part of the deliverable.
|
||||
// Kernel launches: the text "name<<<grid, block>>>(" is rewritten by sed into "emu_launch(name, grid, block, ".
|
||||
// Each emulated block runs `block` host threads; the 32 lanes of a warp synchronise inside __shfl_xor_sync.
|
||||
#pragma once
|
||||
#include <cstdint>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <cstdio>
|
||||
#include <chrono>
|
||||
#include <thread>
|
||||
#include <vector>
|
||||
#include <memory>
|
||||
#include <mutex>
|
||||
#include <condition_variable>
|
||||
|
||||
#define __device__
|
||||
#define __global__
|
||||
#define __host__
|
||||
#define __forceinline__ inline
|
||||
|
||||
struct uint3 { unsigned x, y, z; };
|
||||
struct dim3 { unsigned x, y, z; dim3(unsigned x_ = 1, unsigned y_ = 1, unsigned z_ = 1) : x(x_), y(y_), z(z_) {} };
|
||||
extern thread_local uint3 threadIdx;
|
||||
extern thread_local uint3 blockIdx;
|
||||
extern thread_local dim3 blockDim;
|
||||
extern thread_local dim3 gridDim;
|
||||
|
||||
enum cudaError_t { cudaSuccess = 0, cudaErrorInvalidValue = 1, cudaErrorMemoryAllocation = 2 };
|
||||
enum cudaMemcpyKind { cudaMemcpyHostToHost = 0, cudaMemcpyHostToDevice = 1, cudaMemcpyDeviceToHost = 2, cudaMemcpyDeviceToDevice = 3 };
|
||||
enum cudaDeviceAttr { cudaDevAttrWarpSize = 10, cudaDevAttrClockRate = 13, cudaDevAttrMemoryClockRate = 36,
|
||||
cudaDevAttrGlobalMemoryBusWidth = 37, cudaDevAttrL2CacheSize = 38, cudaDevAttrMaxThreadsPerMultiProcessor = 39 };
|
||||
struct cudaDeviceProp { char name[256]; size_t totalGlobalMem; int multiProcessorCount; int major, minor; };
|
||||
struct cudaFuncAttributes { int numRegs; };
|
||||
struct cudaEventRec { std::chrono::steady_clock::time_point t; };
|
||||
typedef cudaEventRec* cudaEvent_t;
|
||||
|
||||
inline const char* cudaGetErrorString(cudaError_t e) { return e == cudaSuccess ? "no error" : "emulated error"; }
|
||||
inline cudaError_t cudaGetLastError() { return cudaSuccess; }
|
||||
inline cudaError_t cudaMalloc(void** p, size_t n) { *p = std::malloc(n); return *p ? cudaSuccess : cudaErrorMemoryAllocation; }
|
||||
inline cudaError_t cudaFree(void* p) { std::free(p); return cudaSuccess; }
|
||||
inline cudaError_t cudaMemcpy(void* d, const void* s, size_t n, cudaMemcpyKind) { std::memcpy(d, s, n); return cudaSuccess; }
|
||||
inline cudaError_t cudaMemGetInfo(size_t* f, size_t* t) { *f = 8ull << 30; *t = 16ull << 30; return cudaSuccess; }
|
||||
inline cudaError_t cudaEventCreate(cudaEvent_t* e) { *e = new cudaEventRec(); return cudaSuccess; }
|
||||
inline cudaError_t cudaEventRecord(cudaEvent_t e, int = 0) { e->t = std::chrono::steady_clock::now(); return cudaSuccess; }
|
||||
inline cudaError_t cudaEventSynchronize(cudaEvent_t) { return cudaSuccess; }
|
||||
inline cudaError_t cudaEventElapsedTime(float* ms, cudaEvent_t a, cudaEvent_t b) { *ms = std::chrono::duration<float, std::milli>(b->t - a->t).count(); return cudaSuccess; }
|
||||
inline cudaError_t cudaEventDestroy(cudaEvent_t e) { delete e; return cudaSuccess; }
|
||||
inline cudaError_t cudaDeviceSynchronize() { return cudaSuccess; }
|
||||
inline cudaError_t cudaGetDeviceCount(int* c) { *c = 1; return cudaSuccess; }
|
||||
inline cudaError_t cudaSetDevice(int) { return cudaSuccess; }
|
||||
inline cudaError_t cudaGetDeviceProperties(cudaDeviceProp* p, int) {
|
||||
std::strcpy(p->name, "CPU emulation shim (not a GPU)"); p->totalGlobalMem = 16ull << 30; p->multiProcessorCount = 0; p->major = 0; p->minor = 0; return cudaSuccess;
|
||||
}
|
||||
inline cudaError_t cudaDriverGetVersion(int* v) { *v = 0; return cudaSuccess; }
|
||||
inline cudaError_t cudaRuntimeGetVersion(int* v) { *v = 0; return cudaSuccess; }
|
||||
inline cudaError_t cudaDeviceGetAttribute(int* v, cudaDeviceAttr a, int) { *v = (a == cudaDevAttrWarpSize) ? 32 : 0; return cudaSuccess; }
|
||||
template<class T> cudaError_t cudaFuncGetAttributes(cudaFuncAttributes* a, T*) { a->numRegs = 0; return cudaSuccess; }
|
||||
template<class T> cudaError_t cudaOccupancyMaxActiveBlocksPerMultiprocessor(int* n, T*, int, size_t) { *n = 0; return cudaSuccess; }
|
||||
|
||||
// Device intrinsics
|
||||
inline unsigned int __umulhi(unsigned int a, unsigned int b) { return (unsigned int)(((uint64_t)a * (uint64_t)b) >> 32); }
|
||||
unsigned int __shfl_xor_sync(unsigned int mask, unsigned int v, int laneMask, int width = 32);
|
||||
|
||||
// Warp emulation
|
||||
struct EmuWarp {
|
||||
std::mutex m;
|
||||
std::condition_variable cv;
|
||||
unsigned arrived = 0, generation = 0, size = 32;
|
||||
uint32_t slot[32];
|
||||
};
|
||||
extern thread_local EmuWarp* emu_current_warp;
|
||||
|
||||
template<class F, class... A> void emu_launch(F f, unsigned grid, unsigned block, A... args) {
|
||||
unsigned warps = (block + 31u) / 32u;
|
||||
std::vector<std::unique_ptr<EmuWarp>> warpObjs;
|
||||
for (unsigned w = 0; w < warps; ++w) {
|
||||
warpObjs.emplace_back(new EmuWarp());
|
||||
warpObjs.back()->size = (block - 32u * w) < 32u ? (block - 32u * w) : 32u;
|
||||
}
|
||||
std::vector<std::thread> ts;
|
||||
for (unsigned t = 0; t < block; ++t) {
|
||||
EmuWarp* w = warpObjs[t / 32u].get();
|
||||
ts.emplace_back([=]() {
|
||||
threadIdx = uint3{t, 0u, 0u};
|
||||
blockDim = dim3(block);
|
||||
gridDim = dim3(grid);
|
||||
emu_current_warp = w;
|
||||
for (unsigned b = 0; b < grid; ++b) { blockIdx = uint3{b, 0u, 0u}; f(args...); }
|
||||
});
|
||||
}
|
||||
for (auto& th : ts) th.join();
|
||||
}
|
||||
20
proto-cuda/emu/emu.sh
Executable file
20
proto-cuda/emu/emu.sh
Executable file
|
|
@ -0,0 +1,20 @@
|
|||
#!/usr/bin/env bash
|
||||
# CPU emulation of the CUDA harness, for checking a pack on a machine WITHOUT nvcc (any clang++ or g++).
|
||||
# It compiles host.cu and packs/<pack>/kernel.cu as plain C++17 against the shim cuda_runtime.h in this
|
||||
# directory, runs the kernels on host threads (32 threads per warp, a barrier inside __shfl_xor_sync)
|
||||
# and checks the dataset self-test and the vectors. Rates it prints are meaningless; only PASS/FAIL matters.
|
||||
# Usage: emu/emu.sh <pack> [host args...] e.g. emu/emu.sh igneum-genesis --batch-log2 13 --batches 1 --block-warps 2
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
CUDA_DIR="$(cd "$HERE/.." && pwd)"
|
||||
PACK="${1:-igneum-genesis}"; shift || true
|
||||
BUILD="$HERE/build-$PACK"
|
||||
mkdir -p "$BUILD"
|
||||
cp "$CUDA_DIR/host.cu" "$BUILD/host_emu.cpp"
|
||||
# Rewrite only the launch syntax; everything else is the exact generated text.
|
||||
sed -E 's/([A-Za-z_0-9]+)<<<([^,]+), ([^>]+)>>>\(/emu_launch(\1, \2, \3, /' "$CUDA_DIR/packs/$PACK/kernel.cu" > "$BUILD/kernel_emu.cpp"
|
||||
CXX="${CXX:-c++}"
|
||||
"$CXX" -std=c++17 -O2 -Wall -Wextra -I "$HERE" -I "$CUDA_DIR/packs/$PACK" \
|
||||
-o "$BUILD/igneum-emu" "$BUILD/host_emu.cpp" "$BUILD/kernel_emu.cpp" "$HERE/shim.cpp" -pthread
|
||||
echo "compiled $BUILD/igneum-emu (CPU emulation, not a GPU build)"
|
||||
exec "$BUILD/igneum-emu" "$@"
|
||||
29
proto-cuda/emu/shim.cpp
Normal file
29
proto-cuda/emu/shim.cpp
Normal file
|
|
@ -0,0 +1,29 @@
|
|||
#include "cuda_runtime.h"
|
||||
|
||||
thread_local uint3 threadIdx{0u, 0u, 0u};
|
||||
thread_local uint3 blockIdx{0u, 0u, 0u};
|
||||
thread_local dim3 blockDim;
|
||||
thread_local dim3 gridDim;
|
||||
thread_local EmuWarp* emu_current_warp = nullptr;
|
||||
|
||||
static void warp_barrier(EmuWarp* w) {
|
||||
std::unique_lock<std::mutex> lk(w->m);
|
||||
unsigned gen = w->generation;
|
||||
if (++w->arrived == w->size) {
|
||||
w->arrived = 0;
|
||||
++w->generation;
|
||||
w->cv.notify_all();
|
||||
} else {
|
||||
w->cv.wait(lk, [&] { return gen != w->generation; });
|
||||
}
|
||||
}
|
||||
|
||||
unsigned int __shfl_xor_sync(unsigned int, unsigned int v, int laneMask, int) {
|
||||
EmuWarp* w = emu_current_warp;
|
||||
unsigned lane = threadIdx.x & 31u;
|
||||
w->slot[lane] = v;
|
||||
warp_barrier(w);
|
||||
unsigned int r = w->slot[lane ^ (unsigned)laneMask];
|
||||
warp_barrier(w);
|
||||
return r;
|
||||
}
|
||||
355
proto-cuda/host.cu
Normal file
355
proto-cuda/host.cu
Normal file
|
|
@ -0,0 +1,355 @@
|
|||
// igneum-bench-cuda: host program for Igneum's random-program proof-of-work test harness on NVIDIA GPUs.
|
||||
//
|
||||
// TEST HARNESS ONLY. No pool, no network, no wallet, no mining protocol. It fills the dataset on the GPU,
|
||||
// checks the GPU against vectors produced on the Mac (proto-metal), and times the kernel.
|
||||
//
|
||||
// C++17 plus the CUDA runtime API, nothing else. The kernel is compiled ahead of time by nvcc from
|
||||
// packs/<seed>/kernel.cu, which was generated by proto-metal/igneum-bench --export-pack.
|
||||
//
|
||||
// Build (Linux, from proto-cuda/):
|
||||
// nvcc -O3 -std=c++17 -arch=sm_120 -I packs/igneum-genesis -o igneum-bench-cuda-igneum-genesis host.cu packs/igneum-genesis/kernel.cu
|
||||
// Windows and the -arch=native fallback are in README.md.
|
||||
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <chrono>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
#include "program.h"
|
||||
#include "vectors.h"
|
||||
|
||||
#define CUDA_CHECK(call) do { cudaError_t err_ = (call); if (err_ != cudaSuccess) { \
|
||||
std::fprintf(stderr, "CUDA error: %s (%d)\n at %s:%d\n in %s\n", cudaGetErrorString(err_), (int)err_, __FILE__, __LINE__, #call); \
|
||||
std::exit(2); } } while (0)
|
||||
|
||||
static const uint32_t SEEDW[8] = IGNEUM_SEEDW_INIT;
|
||||
|
||||
// Same closed form as ds_elem in kernel.cu and datasetElem in proto-metal/main.swift.
|
||||
static uint32_t host_ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// Options
|
||||
|
||||
struct Options {
|
||||
int datasetMib = 1024;
|
||||
int batchLog2 = 24;
|
||||
int batches = 5;
|
||||
int blockWarps = 1;
|
||||
bool sweep = false;
|
||||
int device = 0;
|
||||
};
|
||||
|
||||
static int packMib() { return (int)(((1ull << IGNEUM_DATASET_LOG2) * 4ull) >> 20); }
|
||||
|
||||
static void usage() {
|
||||
std::printf(
|
||||
"igneum-bench-cuda [--dataset-mib N] [--sweep] [--batch-log2 24] [--batches 5] [--block-warps 1] [--device 0]\n"
|
||||
" --dataset-mib N dataset size in MiB, power of two (default 1024; vectors are only checked at %d MiB)\n"
|
||||
" --sweep run 4, 64, 256, 512 and 1024 MiB in sequence (same sweep as the Mac)\n"
|
||||
" --batch-log2 B nonces per batch = 2^B (default 24)\n"
|
||||
" --batches N timed batches after one warm-up batch (default 5)\n"
|
||||
" --block-warps W warps per thread block, 1..32 (default 1 = one warp per block, like the Metal run)\n"
|
||||
" --device D CUDA device index (default 0)\n", packMib());
|
||||
}
|
||||
|
||||
static bool isPow2(long long v) { return v > 0 && (v & (v - 1)) == 0; }
|
||||
|
||||
static Options parseArgs(int argc, char** argv) {
|
||||
Options o;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
std::string a = argv[i];
|
||||
auto next = [&](int& dst) {
|
||||
if (i + 1 >= argc) { usage(); std::exit(2); }
|
||||
dst = std::atoi(argv[++i]);
|
||||
};
|
||||
if (a == "--dataset-mib") next(o.datasetMib);
|
||||
else if (a == "--batch-log2") next(o.batchLog2);
|
||||
else if (a == "--batches") next(o.batches);
|
||||
else if (a == "--block-warps") next(o.blockWarps);
|
||||
else if (a == "--device") next(o.device);
|
||||
else if (a == "--sweep") o.sweep = true;
|
||||
else if (a == "-h" || a == "--help") { usage(); std::exit(0); }
|
||||
else { std::printf("unknown argument %s\n", argv[i]); usage(); std::exit(2); }
|
||||
}
|
||||
if (!isPow2(o.datasetMib) || o.datasetMib < 1 || o.datasetMib > 16384) {
|
||||
std::printf("--dataset-mib must be a power of two between 1 and 16384\n"); std::exit(2);
|
||||
}
|
||||
if (o.batchLog2 < 10 || o.batchLog2 > 28) { std::printf("--batch-log2 must be between 10 and 28\n"); std::exit(2); }
|
||||
if (o.batches < 1) { std::printf("--batches must be at least 1\n"); std::exit(2); }
|
||||
if (o.blockWarps < 1 || o.blockWarps > 32) { std::printf("--block-warps must be between 1 and 32\n"); std::exit(2); }
|
||||
return o;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// Helpers
|
||||
|
||||
static double wallMs() {
|
||||
using namespace std::chrono;
|
||||
return duration<double, std::milli>(steady_clock::now().time_since_epoch()).count();
|
||||
}
|
||||
|
||||
static int log2u32(uint32_t v) { int n = 0; while (v > 1u) { v >>= 1; ++n; } return n; }
|
||||
|
||||
static bool compareWarp(const uint64_t* got, const uint64_t* want, uint32_t base, const char* how) {
|
||||
int bad = 0, first = -1;
|
||||
for (int l = 0; l < 32; ++l) if (got[l] != want[l]) { if (first < 0) first = l; ++bad; }
|
||||
if (bad == 0) {
|
||||
std::printf("verify warp base %u (nonces %u..%u) %s: PASS\n", base, base, base + 31u, how);
|
||||
} else {
|
||||
std::printf("verify warp base %u (nonces %u..%u) %s: FAIL %d of 32 lanes differ, first lane %d: gpu=%016llx expected=%016llx\n",
|
||||
base, base, base + 31u, how, bad, first,
|
||||
(unsigned long long)got[first], (unsigned long long)want[first]);
|
||||
}
|
||||
return bad == 0;
|
||||
}
|
||||
|
||||
struct SizeResult {
|
||||
int mib = 0;
|
||||
uint32_t words = 0;
|
||||
double fillFirstMs = 0, fillSecondMs = 0;
|
||||
bool dsPass = false;
|
||||
bool vecChecked = false, vecPass = false;
|
||||
double gpuMs = 0, wallMsTimed = 0;
|
||||
double hashesPerSec = 0, gbps = 0;
|
||||
};
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// One dataset size: fill, self-test, vectors, bench
|
||||
|
||||
static SizeResult runSize(const Options& o, int mib, uint64_t* dOut, uint32_t nonces) {
|
||||
SizeResult r;
|
||||
r.mib = mib;
|
||||
uint64_t bytes = (uint64_t)mib << 20;
|
||||
r.words = (uint32_t)(bytes / 4ull);
|
||||
uint32_t mask = r.words - 1u;
|
||||
bool atPackSize = (r.words == (1u << IGNEUM_DATASET_LOG2));
|
||||
std::printf("\n=== dataset %d MiB (2^%d words, mask 0x%08x)%s ===\n", mib, log2u32(r.words), mask,
|
||||
atPackSize ? "" : " [not the pack size: vectors skipped, dataset head and random points still checked]");
|
||||
|
||||
size_t freeB = 0, totalB = 0;
|
||||
CUDA_CHECK(cudaMemGetInfo(&freeB, &totalB));
|
||||
if ((uint64_t)freeB < bytes + (64ull << 20)) {
|
||||
std::printf("FAIL: %llu MiB free on the device, need %d MiB for the dataset\n", (unsigned long long)(freeB >> 20), mib);
|
||||
std::exit(2);
|
||||
}
|
||||
uint32_t* dDs = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dDs, (size_t)bytes));
|
||||
|
||||
cudaEvent_t e0, e1;
|
||||
CUDA_CHECK(cudaEventCreate(&e0));
|
||||
CUDA_CHECK(cudaEventCreate(&e1));
|
||||
|
||||
// Fill twice: the Mac showed a first-touch cost on the first fill of a process.
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(igneum_launch_fill(dDs, r.words, IGNEUM_DAY0, IGNEUM_DAY1));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
float msf = 0.f;
|
||||
CUDA_CHECK(cudaEventElapsedTime(&msf, e0, e1));
|
||||
if (pass == 0) r.fillFirstMs = msf; else r.fillSecondMs = msf;
|
||||
}
|
||||
std::printf("dataset fill: %.2f ms first, %.2f ms second -> %.0f GB/s write (second, GPU time)\n",
|
||||
r.fillFirstMs, r.fillSecondMs, (double)bytes / 1e9 / (r.fillSecondMs / 1000.0));
|
||||
|
||||
// Dataset self-test: head 16 (any size), element [MASK] (pack size only), 64 pseudo-random points vs host formula.
|
||||
{
|
||||
uint32_t head[16];
|
||||
CUDA_CHECK(cudaMemcpy(head, dDs, sizeof(head), cudaMemcpyDeviceToHost));
|
||||
int badHead = 0;
|
||||
for (int i = 0; i < 16; ++i) {
|
||||
if (head[i] != IGNEUM_DS_HEAD[i]) {
|
||||
if (badHead == 0) std::printf(" dataset[%d] = 0x%08x, expected 0x%08x\n", i, head[i], IGNEUM_DS_HEAD[i]);
|
||||
++badHead;
|
||||
}
|
||||
}
|
||||
bool lastOk = true;
|
||||
const char* lastText = "skipped";
|
||||
if (atPackSize) {
|
||||
uint32_t last = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&last, dDs + IGNEUM_DS_LAST_INDEX, sizeof(last), cudaMemcpyDeviceToHost));
|
||||
lastOk = (last == IGNEUM_DS_LAST);
|
||||
lastText = lastOk ? "PASS" : "FAIL";
|
||||
if (!lastOk) std::printf(" dataset[%u] = 0x%08x, expected 0x%08x\n", IGNEUM_DS_LAST_INDEX, last, IGNEUM_DS_LAST);
|
||||
}
|
||||
int badRnd = 0;
|
||||
uint64_t s = 0x9E3779B97F4A7C15ull ^ (uint64_t)r.words;
|
||||
for (int k = 0; k < 64; ++k) {
|
||||
s += 0x9E3779B97F4A7C15ull;
|
||||
uint64_t z = s;
|
||||
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull;
|
||||
z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
|
||||
z ^= z >> 31;
|
||||
uint32_t idx = (uint32_t)z & mask;
|
||||
uint32_t v = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&v, dDs + idx, sizeof(v), cudaMemcpyDeviceToHost));
|
||||
uint32_t want = host_ds_elem(idx, IGNEUM_DAY0, IGNEUM_DAY1);
|
||||
if (v != want) {
|
||||
if (badRnd == 0) std::printf(" dataset[%u] = 0x%08x, host formula 0x%08x\n", idx, v, want);
|
||||
++badRnd;
|
||||
}
|
||||
}
|
||||
r.dsPass = (badHead == 0 && lastOk && badRnd == 0);
|
||||
std::printf("dataset self-test: %s (head 16 vs Mac %s, element [MASK] vs Mac %s, 64 random points vs host formula %s)\n",
|
||||
r.dsPass ? "PASS" : "FAIL", badHead == 0 ? "PASS" : "FAIL", lastText, badRnd == 0 ? "PASS" : "FAIL");
|
||||
}
|
||||
|
||||
// Vectors, standalone: one 32-thread block per base nonce, exactly like the Mac cross-check.
|
||||
uint64_t got[32];
|
||||
if (atPackSize) {
|
||||
r.vecChecked = true;
|
||||
r.vecPass = true;
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, IGNEUM_VEC_BASE[w], mask, 32u, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
|
||||
bool ok = compareWarp(got, IGNEUM_VEC_OUT[w], IGNEUM_VEC_BASE[w], "standalone, 1 warp/block");
|
||||
r.vecPass = r.vecPass && ok;
|
||||
}
|
||||
} else {
|
||||
std::printf("vectors: skipped (the pack's vectors are for %d MiB)\n", packMib());
|
||||
}
|
||||
|
||||
// Warm-up batch at base nonce 0. With the default 2^24 nonces it contains all three vector warps,
|
||||
// so the bench configuration itself (blockDim = 32 x block-warps) is also checked bit for bit.
|
||||
double w0 = wallMs();
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, 0u, mask, nonces, (uint32_t)o.blockWarps));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
double w1 = wallMs();
|
||||
std::printf("warm-up batch: %u hashes in %.2f ms wall\n", nonces, w1 - w0);
|
||||
if (atPackSize) {
|
||||
char how[64];
|
||||
std::snprintf(how, sizeof(how), "in batch, %d warp(s)/block", o.blockWarps);
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
if ((uint64_t)IGNEUM_VEC_BASE[w] + 32ull > (uint64_t)nonces) {
|
||||
std::printf("verify warp base %u in batch: skipped (batch has %u nonces)\n", IGNEUM_VEC_BASE[w], nonces);
|
||||
continue;
|
||||
}
|
||||
CUDA_CHECK(cudaMemcpy(got, dOut + IGNEUM_VEC_BASE[w], sizeof(got), cudaMemcpyDeviceToHost));
|
||||
bool ok = compareWarp(got, IGNEUM_VEC_OUT[w], IGNEUM_VEC_BASE[w], how);
|
||||
r.vecPass = r.vecPass && ok;
|
||||
}
|
||||
}
|
||||
|
||||
// Timed batches. Base nonces (b * nonces) mod 2^32, as in the Metal run.
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
double t0 = wallMs();
|
||||
for (int b = 1; b <= o.batches; ++b) {
|
||||
uint32_t base = (uint32_t)((uint64_t)b * (uint64_t)nonces);
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, (uint32_t)o.blockWarps));
|
||||
}
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
double t1 = wallMs();
|
||||
float gpuMsF = 0.f;
|
||||
CUDA_CHECK(cudaEventElapsedTime(&gpuMsF, e0, e1));
|
||||
double total = (double)nonces * (double)o.batches;
|
||||
r.gpuMs = gpuMsF;
|
||||
r.wallMsTimed = t1 - t0;
|
||||
r.hashesPerSec = total / (r.gpuMs / 1000.0);
|
||||
r.gbps = r.hashesPerSec * (double)IGNEUM_LOADS_PER_HASH * 4.0 / 1e9;
|
||||
std::printf("timed: %d batches x %u hashes = %.0f hashes\n", o.batches, nonces, total);
|
||||
std::printf(" GPU %.2f ms -> %.3f Mhash/s (%.0f hashes/s), %.2f GB/s useful (loads x 4 B)\n",
|
||||
r.gpuMs, r.hashesPerSec / 1e6, r.hashesPerSec, r.gbps);
|
||||
std::printf(" wall %.2f ms -> %.3f Mhash/s\n", r.wallMsTimed, total / (r.wallMsTimed / 1000.0) / 1e6);
|
||||
|
||||
CUDA_CHECK(cudaEventDestroy(e0));
|
||||
CUDA_CHECK(cudaEventDestroy(e1));
|
||||
CUDA_CHECK(cudaFree(dDs));
|
||||
return r;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// Main
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
Options o = parseArgs(argc, argv);
|
||||
std::printf("igneum-bench-cuda pack \"%s\" (test harness: no pool, no network, no wallet)\n", IGNEUM_SEED_STRING);
|
||||
|
||||
int count = 0;
|
||||
CUDA_CHECK(cudaGetDeviceCount(&count));
|
||||
if (count == 0) { std::printf("FAIL: no CUDA device\n"); return 2; }
|
||||
if (o.device < 0 || o.device >= count) { std::printf("FAIL: device %d out of range (%d devices)\n", o.device, count); return 2; }
|
||||
CUDA_CHECK(cudaSetDevice(o.device));
|
||||
|
||||
cudaDeviceProp prop;
|
||||
std::memset(&prop, 0, sizeof(prop));
|
||||
CUDA_CHECK(cudaGetDeviceProperties(&prop, o.device));
|
||||
int drv = 0, rt = 0;
|
||||
CUDA_CHECK(cudaDriverGetVersion(&drv));
|
||||
CUDA_CHECK(cudaRuntimeGetVersion(&rt));
|
||||
int warp = 0, clk = 0, memclk = 0, bus = 0, l2 = 0, thrSM = 0;
|
||||
cudaDeviceGetAttribute(&warp, cudaDevAttrWarpSize, o.device);
|
||||
cudaDeviceGetAttribute(&clk, cudaDevAttrClockRate, o.device);
|
||||
cudaDeviceGetAttribute(&memclk, cudaDevAttrMemoryClockRate, o.device);
|
||||
cudaDeviceGetAttribute(&bus, cudaDevAttrGlobalMemoryBusWidth, o.device);
|
||||
cudaDeviceGetAttribute(&l2, cudaDevAttrL2CacheSize, o.device);
|
||||
cudaDeviceGetAttribute(&thrSM, cudaDevAttrMaxThreadsPerMultiProcessor, o.device);
|
||||
cudaGetLastError(); // attribute queries are informational; clear any error they left
|
||||
|
||||
std::printf("GPU: %s (%d SMs, compute capability %d.%d, %.0f MiB global memory)\n",
|
||||
prop.name, prop.multiProcessorCount, prop.major, prop.minor, (double)prop.totalGlobalMem / 1048576.0);
|
||||
std::printf(" SM clock %d MHz, memory clock %d MHz, bus %d bits, L2 %d MiB, max %d threads/SM, warp size %d\n",
|
||||
clk / 1000, memclk / 1000, bus, l2 / 1048576, thrSM, warp);
|
||||
std::printf("CUDA: driver %d.%d, runtime %d.%d\n", drv / 1000, (drv % 100) / 10, rt / 1000, (rt % 100) / 10);
|
||||
if (warp != 32) {
|
||||
std::printf("WARNING: warp size is %d, not 32. The 32-lane shuffle model does not hold on this device.\n", warp);
|
||||
}
|
||||
|
||||
int regs = 0, blocksPerSM = 0;
|
||||
CUDA_CHECK(igneum_hash_info(®s, &blocksPerSM, (uint32_t)o.blockWarps));
|
||||
std::printf("kernel: %d registers/thread, %d resident blocks/SM at %d warp(s)/block = %d resident warps/SM\n",
|
||||
regs, blocksPerSM, o.blockWarps, blocksPerSM * o.blockWarps);
|
||||
std::printf("program: %d instructions x %d iterations, loads/hash %d, op mix %s\n",
|
||||
IGNEUM_INSTR_COUNT, IGNEUM_ITERATIONS, IGNEUM_LOADS_PER_HASH, IGNEUM_OP_MIX);
|
||||
std::printf("seed words: %08x %08x %08x %08x %08x %08x %08x %08x\n",
|
||||
SEEDW[0], SEEDW[1], SEEDW[2], SEEDW[3], SEEDW[4], SEEDW[5], SEEDW[6], SEEDW[7]);
|
||||
std::printf("day \"%s\" (d0 0x%08x, d1 0x%08x), pack dataset 2^%d words = %d MiB\n",
|
||||
IGNEUM_DAY_STRING, IGNEUM_DAY0, IGNEUM_DAY1, IGNEUM_DATASET_LOG2, packMib());
|
||||
|
||||
uint32_t nonces = 1u << o.batchLog2;
|
||||
if (nonces % (32u * (uint32_t)o.blockWarps) != 0u) {
|
||||
std::printf("FAIL: 2^%d nonces is not a multiple of %d threads per block\n", o.batchLog2, 32 * o.blockWarps);
|
||||
return 2;
|
||||
}
|
||||
uint64_t* dOut = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dOut, (size_t)nonces * sizeof(uint64_t)));
|
||||
|
||||
std::vector<int> sizes;
|
||||
if (o.sweep) { sizes.push_back(4); sizes.push_back(64); sizes.push_back(256); sizes.push_back(512); sizes.push_back(1024); }
|
||||
else sizes.push_back(o.datasetMib);
|
||||
|
||||
std::vector<SizeResult> results;
|
||||
for (size_t i = 0; i < sizes.size(); ++i) results.push_back(runSize(o, sizes[i], dOut, nonces));
|
||||
CUDA_CHECK(cudaFree(dOut));
|
||||
|
||||
std::printf("\n=== summary (%s, pack %s, batch 2^%d x %d, %d warp(s)/block, GPU-event time) ===\n",
|
||||
prop.name, IGNEUM_SEED_STRING, o.batchLog2, o.batches, o.blockWarps);
|
||||
std::printf("| dataset MiB | fill ms (second) | Mhash/s | GB/s useful | random loads/s (G) | loads/hash | dataset self-test | vectors |\n");
|
||||
std::printf("|---|---|---|---|---|---|---|---|\n");
|
||||
bool overall = true, anyVec = false;
|
||||
for (size_t i = 0; i < results.size(); ++i) {
|
||||
const SizeResult& r = results[i];
|
||||
overall = overall && r.dsPass && (!r.vecChecked || r.vecPass);
|
||||
anyVec = anyVec || r.vecChecked;
|
||||
std::printf("| %d | %.2f | %.3f | %.2f | %.2f | %d | %s | %s |\n",
|
||||
r.mib, r.fillSecondMs, r.hashesPerSec / 1e6, r.gbps,
|
||||
r.hashesPerSec * (double)IGNEUM_LOADS_PER_HASH / 1e9, IGNEUM_LOADS_PER_HASH,
|
||||
r.dsPass ? "PASS" : "FAIL",
|
||||
r.vecChecked ? (r.vecPass ? "PASS (3 warps, standalone and in batch)" : "FAIL") : "skipped (not pack size)");
|
||||
}
|
||||
if (!anyVec) std::printf("NOTE: no vectors were checked. Run at %d MiB (the default) to verify against the Mac.\n", packMib());
|
||||
std::printf("OVERALL: %s\n", overall ? "PASS" : "FAIL");
|
||||
return overall ? 0 : 1;
|
||||
}
|
||||
145
proto-cuda/packs/igneum-genesis/kernel.cu
Normal file
145
proto-cuda/packs/igneum-genesis/kernel.cu
Normal file
|
|
@ -0,0 +1,145 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// dataset[i] = ds_elem(i, d0, d1) for i < n. Same closed form as the Metal igneum_fill kernel.
|
||||
__global__ void igneum_fill(uint32_t* ds, uint32_t n, uint32_t d0, uint32_t d1) {
|
||||
uint32_t i = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (i < n) ds[i] = ds_elem(i, d0, d1);
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = rotl_imm(r4, 25u); // 0 rotl
|
||||
r0 = r0 - r5; // 1 sub
|
||||
r4 = r4 ^ ds[r3 & mask]; // 2 load
|
||||
r1 = rotl_imm(r1, 1u); // 3 rotl
|
||||
r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 4 add
|
||||
r5 = r5 ^ ds[r3 & mask]; // 5 load
|
||||
r5 = r5 - r7; // 6 sub
|
||||
r3 = r3 + r4 + ((((sel >> 26u) & 1u) != 0u) ? 0x5a069596u : 0x52f2dbf4u); // 7 add
|
||||
r0 = r0 ^ r4; // 8 xor
|
||||
r4 = r4 ^ r0; // 9 xor
|
||||
r2 = r2 ^ ds[r0 & mask]; // 10 load
|
||||
r4 = r4 ^ __shfl_xor_sync(0xffffffffu, r6, 16); // 11 shfl
|
||||
r1 = r1 - r5; // 12 sub
|
||||
r2 = r2 ^ r1; // 13 xor
|
||||
r4 = r4 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 ^ ds[r4 & mask]; // 15 load
|
||||
r3 = r3 ^ ds[r0 & mask]; // 16 load
|
||||
r4 = r4 ^ r6; // 17 xor
|
||||
r2 = r4 * r6 + r2; // 18 mad
|
||||
r6 = rotr_var(r6, r1); // 19 rotr
|
||||
r3 = r3 ^ r4; // 20 xor
|
||||
r1 = r3 * r5 + r1; // 21 mad
|
||||
r7 = __umulhi(r7, r4); // 22 mulhi
|
||||
r5 = __umulhi(r5, r2); // 23 mulhi
|
||||
r0 = r0 ^ ds[r6 & mask]; // 24 load
|
||||
r5 = r5 * r6; // 25 mul
|
||||
r7 = r7 ^ ds[r1 & mask]; // 26 load
|
||||
r3 = rotr_var(r3, r1); // 27 rotr
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r1, 8); // 28 shfl
|
||||
r7 = r7 ^ r5; // 29 xor
|
||||
r7 = rotl_imm(r7, 23u); // 30 rotl
|
||||
r2 = r2 - r0; // 31 sub
|
||||
r7 = r7 ^ r2; // 32 xor
|
||||
r2 = r2 ^ r6; // 33 xor
|
||||
r6 = r6 ^ ds[r1 & mask]; // 34 load
|
||||
r1 = r1 ^ r4; // 35 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r0, 1); // 36 shfl
|
||||
r2 = r2 * r6; // 37 mul
|
||||
r5 = r5 + r3 + ((((sel >> 12u) & 1u) != 0u) ? 0xa29f4338u : 0x71f30417u); // 38 add
|
||||
r7 = r7 ^ r6; // 39 xor
|
||||
r7 = r7 ^ r3; // 40 xor
|
||||
r3 = rotr_var(r3, r4); // 41 rotr
|
||||
r5 = r5 ^ r3; // 42 xor
|
||||
r3 = rotr_var(r3, r6); // 43 rotr
|
||||
r1 = r3 * r5 + r1; // 44 mad
|
||||
r7 = r7 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xa907b90bu : 0xc1ae8d3bu); // 45 add
|
||||
r7 = r7 | r2; // 46 or
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r3, 16); // 47 shfl
|
||||
r1 = r1 ^ ds[r7 & mask]; // 48 load
|
||||
r5 = r5 - r1; // 49 sub
|
||||
r3 = r3 ^ ds[r1 & mask]; // 50 load
|
||||
r2 = r2 - r3; // 51 sub
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 52 shfl
|
||||
r2 = r2 - r0; // 53 sub
|
||||
r0 = r0 ^ ds[r3 & mask]; // 54 load
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r1, 16); // 55 shfl
|
||||
r0 = r0 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0xf4689674u : 0x25955401u); // 56 add
|
||||
r2 = __umulhi(r2, r0); // 57 mulhi
|
||||
r4 = __umulhi(r4, r2); // 58 mulhi
|
||||
r2 = r6 * r7 + r2; // 59 mad
|
||||
r3 = r3 ^ r1; // 60 xor
|
||||
r4 = r4 * r2; // 61 mul
|
||||
r0 = r0 ^ ds[r2 & mask]; // 62 load
|
||||
r7 = __umulhi(r7, r0); // 63 mulhi
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_fill(uint32_t* ds, uint32_t nWords, uint32_t d0, uint32_t d1) {
|
||||
if (nWords == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nWords + block - 1u) / block;
|
||||
igneum_fill<<<grid, block>>>(ds, nWords, d0, d1);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
25
proto-cuda/packs/igneum-genesis/program.h
Normal file
25
proto-cuda/packs/igneum-genesis/program.h
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
#pragma once
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-genesis"
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 104
|
||||
#define IGNEUM_OP_MIX "load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1"
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
|
||||
|
||||
// Defined in kernel.cu. Both launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_fill(uint32_t* ds, uint32_t nWords, uint32_t d0, uint32_t d1);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
105
proto-cuda/packs/igneum-genesis/program.json
Normal file
105
proto-cuda/packs/igneum-genesis/program.json
Normal file
|
|
@ -0,0 +1,105 @@
|
|||
{
|
||||
"format": "igneum-program-pack-1",
|
||||
"seed": "igneum-genesis",
|
||||
"seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"],
|
||||
"seed_derivation": "FNV-1a 64 over UTF-8 of seed, basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 104,
|
||||
"op_mix": {"load": 13, "xor": 13, "sub": 7, "shfl": 6, "add": 5, "mulhi": 5, "mad": 4, "rotr": 4, "mul": 3, "rotl": 3, "or": 1},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
|
||||
"op_semantics": {
|
||||
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
|
||||
"sub": "dst = dst - src",
|
||||
"mul": "dst = dst * src (low 32)",
|
||||
"mulhi": "dst = high 32 bits of dst * src",
|
||||
"xor": "dst = dst ^ src",
|
||||
"or": "dst = dst | src",
|
||||
"rotl": "dst = rotl(dst, rot), rot in 1..31",
|
||||
"rotr": "dst = rotr(dst, src & 31)",
|
||||
"mad": "dst = src * src2 + dst",
|
||||
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
|
||||
"load": "dst = dst ^ dataset[src & dataset.mask]"
|
||||
},
|
||||
"dataset": {
|
||||
"log2_words": 28,
|
||||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "2026-10-03",
|
||||
"day_words_from": "day/2026-10-03",
|
||||
"d0": "0x3067619f",
|
||||
"d1": "0x3c269176",
|
||||
"formula": "x = i ^ d0; x *= 0x9E3779B1; x ^= x>>15; x += d1; x *= 0x85EBCA77; x ^= x>>13; x *= 0xC2B2AE3D; x ^= x>>16 (all mod 2^32)"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "rotl", "dst": 4, "src": 2, "src2": 7, "imm": "0x20699878", "imm2": "0x6f1a6170", "rot": 25, "bit": 7, "mask": 16},
|
||||
{"i": 1, "op": "sub", "dst": 0, "src": 5, "src2": 4, "imm": "0xf81a0b9d", "imm2": "0xf0505e88", "rot": 1, "bit": 4, "mask": 1},
|
||||
{"i": 2, "op": "load", "dst": 4, "src": 3, "src2": 6, "imm": "0xb1978a0b", "imm2": "0x2ca4e162", "rot": 10, "bit": 21, "mask": 1},
|
||||
{"i": 3, "op": "rotl", "dst": 1, "src": 6, "src2": 7, "imm": "0xc3bd2355", "imm2": "0xa8c5f27e", "rot": 1, "bit": 22, "mask": 1},
|
||||
{"i": 4, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2},
|
||||
{"i": 5, "op": "load", "dst": 5, "src": 3, "src2": 3, "imm": "0x4d183796", "imm2": "0x679648a8", "rot": 4, "bit": 30, "mask": 4},
|
||||
{"i": 6, "op": "sub", "dst": 5, "src": 7, "src2": 7, "imm": "0x265677dc", "imm2": "0x9043323e", "rot": 30, "bit": 7, "mask": 4},
|
||||
{"i": 7, "op": "add", "dst": 3, "src": 4, "src2": 6, "imm": "0x52f2dbf4", "imm2": "0x5a069596", "rot": 5, "bit": 26, "mask": 2},
|
||||
{"i": 8, "op": "xor", "dst": 0, "src": 4, "src2": 1, "imm": "0x3303ec4b", "imm2": "0xfeca75be", "rot": 21, "bit": 7, "mask": 1},
|
||||
{"i": 9, "op": "xor", "dst": 4, "src": 0, "src2": 0, "imm": "0x5dc5959c", "imm2": "0x023f44a8", "rot": 11, "bit": 7, "mask": 4},
|
||||
{"i": 10, "op": "load", "dst": 2, "src": 0, "src2": 3, "imm": "0xde716173", "imm2": "0xc21e924d", "rot": 30, "bit": 15, "mask": 1},
|
||||
{"i": 11, "op": "shfl", "dst": 4, "src": 6, "src2": 1, "imm": "0x3cda6d48", "imm2": "0x0970145b", "rot": 12, "bit": 22, "mask": 16},
|
||||
{"i": 12, "op": "sub", "dst": 1, "src": 5, "src2": 4, "imm": "0x48866c15", "imm2": "0x4eaee50c", "rot": 30, "bit": 20, "mask": 16},
|
||||
{"i": 13, "op": "xor", "dst": 2, "src": 1, "src2": 5, "imm": "0x196d165c", "imm2": "0x0f730511", "rot": 11, "bit": 3, "mask": 4},
|
||||
{"i": 14, "op": "load", "dst": 4, "src": 5, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8},
|
||||
{"i": 15, "op": "load", "dst": 2, "src": 4, "src2": 1, "imm": "0x17c9c95b", "imm2": "0x306542fe", "rot": 27, "bit": 17, "mask": 16},
|
||||
{"i": 16, "op": "load", "dst": 3, "src": 0, "src2": 2, "imm": "0x590f9e11", "imm2": "0xa4c13036", "rot": 3, "bit": 28, "mask": 16},
|
||||
{"i": 17, "op": "xor", "dst": 4, "src": 6, "src2": 3, "imm": "0x9c4eeee9", "imm2": "0xf069b834", "rot": 11, "bit": 23, "mask": 8},
|
||||
{"i": 18, "op": "mad", "dst": 2, "src": 4, "src2": 6, "imm": "0xa242a28b", "imm2": "0xb8974bdf", "rot": 13, "bit": 30, "mask": 1},
|
||||
{"i": 19, "op": "rotr", "dst": 6, "src": 1, "src2": 0, "imm": "0xcd69ed50", "imm2": "0xadf52615", "rot": 30, "bit": 2, "mask": 2},
|
||||
{"i": 20, "op": "xor", "dst": 3, "src": 4, "src2": 0, "imm": "0xe3059a24", "imm2": "0x16b8dd86", "rot": 25, "bit": 31, "mask": 16},
|
||||
{"i": 21, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0xc3ae8ae1", "imm2": "0x8e126e8a", "rot": 2, "bit": 28, "mask": 1},
|
||||
{"i": 22, "op": "mulhi", "dst": 7, "src": 4, "src2": 2, "imm": "0x6909af7a", "imm2": "0xa388b4b9", "rot": 14, "bit": 10, "mask": 2},
|
||||
{"i": 23, "op": "mulhi", "dst": 5, "src": 2, "src2": 2, "imm": "0xdf099cfb", "imm2": "0xd0133a01", "rot": 31, "bit": 3, "mask": 4},
|
||||
{"i": 24, "op": "load", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8},
|
||||
{"i": 25, "op": "mul", "dst": 5, "src": 6, "src2": 7, "imm": "0x089f5404", "imm2": "0xbd066e1d", "rot": 7, "bit": 10, "mask": 4},
|
||||
{"i": 26, "op": "load", "dst": 7, "src": 1, "src2": 2, "imm": "0x3e26afea", "imm2": "0xba573970", "rot": 11, "bit": 15, "mask": 8},
|
||||
{"i": 27, "op": "rotr", "dst": 3, "src": 1, "src2": 5, "imm": "0xe2466d63", "imm2": "0x30db112b", "rot": 31, "bit": 15, "mask": 2},
|
||||
{"i": 28, "op": "shfl", "dst": 5, "src": 1, "src2": 3, "imm": "0xdc20297d", "imm2": "0xcb813557", "rot": 3, "bit": 21, "mask": 8},
|
||||
{"i": 29, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0x6ede5c13", "imm2": "0x1f14267f", "rot": 22, "bit": 31, "mask": 1},
|
||||
{"i": 30, "op": "rotl", "dst": 7, "src": 5, "src2": 6, "imm": "0x24cdb54b", "imm2": "0xa44e008d", "rot": 23, "bit": 4, "mask": 16},
|
||||
{"i": 31, "op": "sub", "dst": 2, "src": 0, "src2": 2, "imm": "0xf94d9b65", "imm2": "0x3457264c", "rot": 11, "bit": 5, "mask": 2},
|
||||
{"i": 32, "op": "xor", "dst": 7, "src": 2, "src2": 3, "imm": "0x8d2046e5", "imm2": "0xf68893bb", "rot": 21, "bit": 2, "mask": 1},
|
||||
{"i": 33, "op": "xor", "dst": 2, "src": 6, "src2": 5, "imm": "0xe0ebc4ce", "imm2": "0x02773069", "rot": 21, "bit": 16, "mask": 1},
|
||||
{"i": 34, "op": "load", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16},
|
||||
{"i": 35, "op": "xor", "dst": 1, "src": 4, "src2": 6, "imm": "0xe3f24158", "imm2": "0x5c64a589", "rot": 13, "bit": 0, "mask": 8},
|
||||
{"i": 36, "op": "shfl", "dst": 3, "src": 0, "src2": 3, "imm": "0x63578bc1", "imm2": "0xbf64a89f", "rot": 16, "bit": 26, "mask": 1},
|
||||
{"i": 37, "op": "mul", "dst": 2, "src": 6, "src2": 4, "imm": "0xfa624b69", "imm2": "0x0389cf85", "rot": 28, "bit": 4, "mask": 1},
|
||||
{"i": 38, "op": "add", "dst": 5, "src": 3, "src2": 5, "imm": "0x71f30417", "imm2": "0xa29f4338", "rot": 2, "bit": 12, "mask": 2},
|
||||
{"i": 39, "op": "xor", "dst": 7, "src": 6, "src2": 3, "imm": "0xfa3b845c", "imm2": "0x9717695d", "rot": 18, "bit": 24, "mask": 1},
|
||||
{"i": 40, "op": "xor", "dst": 7, "src": 3, "src2": 4, "imm": "0x8408dc67", "imm2": "0x8d389f9b", "rot": 21, "bit": 16, "mask": 16},
|
||||
{"i": 41, "op": "rotr", "dst": 3, "src": 4, "src2": 1, "imm": "0x4bbcad92", "imm2": "0xbc6375cd", "rot": 18, "bit": 4, "mask": 2},
|
||||
{"i": 42, "op": "xor", "dst": 5, "src": 3, "src2": 3, "imm": "0xbd31a2ea", "imm2": "0x742d28ea", "rot": 4, "bit": 12, "mask": 4},
|
||||
{"i": 43, "op": "rotr", "dst": 3, "src": 6, "src2": 5, "imm": "0x35c52d04", "imm2": "0x0f8e8621", "rot": 3, "bit": 22, "mask": 8},
|
||||
{"i": 44, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16},
|
||||
{"i": 45, "op": "add", "dst": 7, "src": 3, "src2": 7, "imm": "0xc1ae8d3b", "imm2": "0xa907b90b", "rot": 13, "bit": 29, "mask": 1},
|
||||
{"i": 46, "op": "or", "dst": 7, "src": 2, "src2": 2, "imm": "0xc4f18bec", "imm2": "0x8a3e2464", "rot": 30, "bit": 16, "mask": 2},
|
||||
{"i": 47, "op": "shfl", "dst": 2, "src": 3, "src2": 6, "imm": "0xf73ba9e3", "imm2": "0x028aa63c", "rot": 3, "bit": 20, "mask": 16},
|
||||
{"i": 48, "op": "load", "dst": 1, "src": 7, "src2": 5, "imm": "0x53be87d4", "imm2": "0x690a5729", "rot": 1, "bit": 17, "mask": 8},
|
||||
{"i": 49, "op": "sub", "dst": 5, "src": 1, "src2": 7, "imm": "0xecdc31a6", "imm2": "0xed6bd9f2", "rot": 1, "bit": 13, "mask": 8},
|
||||
{"i": 50, "op": "load", "dst": 3, "src": 1, "src2": 3, "imm": "0x55f91dbc", "imm2": "0xf026553d", "rot": 30, "bit": 27, "mask": 1},
|
||||
{"i": 51, "op": "sub", "dst": 2, "src": 3, "src2": 5, "imm": "0x6689eef2", "imm2": "0x24f3ae21", "rot": 6, "bit": 22, "mask": 8},
|
||||
{"i": 52, "op": "shfl", "dst": 6, "src": 2, "src2": 4, "imm": "0x8fd37aad", "imm2": "0x89ed5d67", "rot": 8, "bit": 16, "mask": 2},
|
||||
{"i": 53, "op": "sub", "dst": 2, "src": 0, "src2": 0, "imm": "0x963bb7e6", "imm2": "0x4738b84f", "rot": 18, "bit": 16, "mask": 4},
|
||||
{"i": 54, "op": "load", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4},
|
||||
{"i": 55, "op": "shfl", "dst": 2, "src": 1, "src2": 5, "imm": "0xb17aad78", "imm2": "0x8458f7ac", "rot": 5, "bit": 7, "mask": 16},
|
||||
{"i": 56, "op": "add", "dst": 0, "src": 6, "src2": 0, "imm": "0x25955401", "imm2": "0xf4689674", "rot": 14, "bit": 27, "mask": 4},
|
||||
{"i": 57, "op": "mulhi", "dst": 2, "src": 0, "src2": 0, "imm": "0xa4e8c86a", "imm2": "0x14bde8e1", "rot": 8, "bit": 9, "mask": 16},
|
||||
{"i": 58, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0xa36b1f4f", "imm2": "0x372e072f", "rot": 31, "bit": 31, "mask": 1},
|
||||
{"i": 59, "op": "mad", "dst": 2, "src": 6, "src2": 7, "imm": "0xfcd3b17f", "imm2": "0x6a8583df", "rot": 24, "bit": 9, "mask": 2},
|
||||
{"i": 60, "op": "xor", "dst": 3, "src": 1, "src2": 0, "imm": "0x3bbea2ae", "imm2": "0x914e2f0a", "rot": 2, "bit": 15, "mask": 16},
|
||||
{"i": 61, "op": "mul", "dst": 4, "src": 2, "src2": 6, "imm": "0x50aaa099", "imm2": "0xd1b988a1", "rot": 27, "bit": 31, "mask": 4},
|
||||
{"i": 62, "op": "load", "dst": 0, "src": 2, "src2": 3, "imm": "0x4ccdbdea", "imm2": "0xf3a60b41", "rot": 1, "bit": 20, "mask": 2},
|
||||
{"i": 63, "op": "mulhi", "dst": 7, "src": 0, "src2": 7, "imm": "0x3e400372", "imm2": "0x92c6201f", "rot": 1, "bit": 13, "mask": 8}
|
||||
]
|
||||
}
|
||||
109
proto-cuda/packs/igneum-genesis/program.metal
Normal file
109
proto-cuda/packs/igneum-genesis/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = rotl_imm(r4, 25u); // 0
|
||||
r0 = r0 - r5; // 1
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 2
|
||||
r1 = rotl_imm(r1, 1u); // 3
|
||||
r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 4
|
||||
r5 = r5 ^ dataset[r3 & MASK]; // 5
|
||||
r5 = r5 - r7; // 6
|
||||
r3 = r3 + r4 + select(0x52f2dbf4u, 0x5a069596u, ((sel >> 26u) & 1u) != 0u); // 7
|
||||
r0 = r0 ^ r4; // 8
|
||||
r4 = r4 ^ r0; // 9
|
||||
r2 = r2 ^ dataset[r0 & MASK]; // 10
|
||||
r4 = r4 ^ simd_shuffle_xor(r6, (ushort)16); // 11
|
||||
r1 = r1 - r5; // 12
|
||||
r2 = r2 ^ r1; // 13
|
||||
r4 = r4 ^ dataset[r5 & MASK]; // 14
|
||||
r2 = r2 ^ dataset[r4 & MASK]; // 15
|
||||
r3 = r3 ^ dataset[r0 & MASK]; // 16
|
||||
r4 = r4 ^ r6; // 17
|
||||
r2 = r4 * r6 + r2; // 18
|
||||
r6 = rotr_var(r6, r1); // 19
|
||||
r3 = r3 ^ r4; // 20
|
||||
r1 = r3 * r5 + r1; // 21
|
||||
r7 = mulhi(r7, r4); // 22
|
||||
r5 = mulhi(r5, r2); // 23
|
||||
r0 = r0 ^ dataset[r6 & MASK]; // 24
|
||||
r5 = r5 * r6; // 25
|
||||
r7 = r7 ^ dataset[r1 & MASK]; // 26
|
||||
r3 = rotr_var(r3, r1); // 27
|
||||
r5 = r5 ^ simd_shuffle_xor(r1, (ushort)8); // 28
|
||||
r7 = r7 ^ r5; // 29
|
||||
r7 = rotl_imm(r7, 23u); // 30
|
||||
r2 = r2 - r0; // 31
|
||||
r7 = r7 ^ r2; // 32
|
||||
r2 = r2 ^ r6; // 33
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 34
|
||||
r1 = r1 ^ r4; // 35
|
||||
r3 = r3 ^ simd_shuffle_xor(r0, (ushort)1); // 36
|
||||
r2 = r2 * r6; // 37
|
||||
r5 = r5 + r3 + select(0x71f30417u, 0xa29f4338u, ((sel >> 12u) & 1u) != 0u); // 38
|
||||
r7 = r7 ^ r6; // 39
|
||||
r7 = r7 ^ r3; // 40
|
||||
r3 = rotr_var(r3, r4); // 41
|
||||
r5 = r5 ^ r3; // 42
|
||||
r3 = rotr_var(r3, r6); // 43
|
||||
r1 = r3 * r5 + r1; // 44
|
||||
r7 = r7 + r3 + select(0xc1ae8d3bu, 0xa907b90bu, ((sel >> 29u) & 1u) != 0u); // 45
|
||||
r7 = r7 | r2; // 46
|
||||
r2 = r2 ^ simd_shuffle_xor(r3, (ushort)16); // 47
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 48
|
||||
r5 = r5 - r1; // 49
|
||||
r3 = r3 ^ dataset[r1 & MASK]; // 50
|
||||
r2 = r2 - r3; // 51
|
||||
r6 = r6 ^ simd_shuffle_xor(r2, (ushort)2); // 52
|
||||
r2 = r2 - r0; // 53
|
||||
r0 = r0 ^ dataset[r3 & MASK]; // 54
|
||||
r2 = r2 ^ simd_shuffle_xor(r1, (ushort)16); // 55
|
||||
r0 = r0 + r6 + select(0x25955401u, 0xf4689674u, ((sel >> 27u) & 1u) != 0u); // 56
|
||||
r2 = mulhi(r2, r0); // 57
|
||||
r4 = mulhi(r4, r2); // 58
|
||||
r2 = r6 * r7 + r2; // 59
|
||||
r3 = r3 ^ r1; // 60
|
||||
r4 = r4 * r2; // 61
|
||||
r0 = r0 ^ dataset[r2 & MASK]; // 62
|
||||
r7 = mulhi(r7, r0); // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
35
proto-cuda/packs/igneum-genesis/vectors.h
Normal file
35
proto-cuda/packs/igneum-genesis/vectors.h
Normal file
|
|
@ -0,0 +1,35 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Expected outputs: proto-metal CPU interpreter (cpuWarp) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps
|
||||
#pragma once
|
||||
#include <cstdint>
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x2941e93c76cb1910ull, 0xc0099c34df8280aeull, 0xe61da8a237797181ull, 0xfb944794eaa8ea1dull, 0x0093934e41befb07ull, 0xd4ba2031c0385915ull, 0x75a8a43902cee166ull, 0xb4ce2d6829aaa27cull,
|
||||
0x478095c2e9911e5dull, 0x595c8107b2792a66ull, 0x21b1b582752cc1d2ull, 0x16e746aa4d9532b8ull, 0x35129dbb17efe2e4ull, 0x926e0b679bf37640ull, 0xb2701dd71a189e19ull, 0x100e94f487022795ull,
|
||||
0x2b9bc5aa1452be87ull, 0x8dabbbe8b909130bull, 0x93a2cd4cba088297ull, 0x662c3e253432099eull, 0xf489da4c9c770867ull, 0xe37864f932e97adfull, 0x3ddf029de9b3c1edull, 0x9059da45130736acull,
|
||||
0x1a667325a17e0016ull, 0x74723d71ab69828aull, 0x121026c2f14795c1ull, 0x4989e1480c662cc8ull, 0xac3b8fd5307c0ab4ull, 0xdc3fc0bb843e84e8ull, 0xfa86df690d119639ull, 0x453388e1be04e25full
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x198f269c663acaf6ull, 0x6c125dcda3cade45ull, 0x0ae4eba3f2a0cdecull, 0xc112952633d9477cull, 0x2dfc52ebb7281f9eull, 0x70b0e04375823959ull, 0x699e78e051dfaa76ull, 0xd91ae88bc2eba03full,
|
||||
0x928b2939d0b23e68ull, 0x0e5feb811f4dd40bull, 0xd0e4ce74e6693b53ull, 0xbd8ac2d12ea0ea1bull, 0xa39bcd116eec4f88ull, 0xdd07f4f6a5a07514ull, 0xdc1cf8b4615b8842ull, 0xde0fb1599057c2ddull,
|
||||
0x56edb1995f7c341eull, 0x02782ce22dd25939ull, 0x150b52d7b4b87319ull, 0xd5c89a9a6c2786c3ull, 0x8aef759270e00bbaull, 0xf9f26409968ea1eaull, 0x2013ae93398863f0ull, 0x863877ca875900deull,
|
||||
0x6e05e03111bf02a3ull, 0x21ec57017818f221ull, 0x62b33a3456032a3bull, 0x5bc24dc2bf442511ull, 0xb97041874e099be7ull, 0x2d9f09fb5d11fa13ull, 0x43f3f83f764ea247ull, 0xa3ec25a3072b734cull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0xa63f6d32a9dc2bacull, 0x3263d16a85274409ull, 0xca7b34b1061437c7ull, 0xaf24083fbe9a33acull, 0x68c42544f8798a73ull, 0x5ba44c42278e0d0aull, 0x12b31f0da8ea79e3ull, 0xa4a6186f5f5b7280ull,
|
||||
0x1534d7e10b3f87aeull, 0xb04618c9238e7436ull, 0x28be2d38f9db17d7ull, 0xebed3a6006a491d1ull, 0x23d50e8d3518ee39ull, 0xd75fac65bc8d1975ull, 0x3fd5c0a165c0669bull, 0x882720cae6122082ull,
|
||||
0xc107a8305a1f009dull, 0xa58cc24784025f2cull, 0x7ee69d59a87a92ebull, 0x46009435f6f19482ull, 0xf50c2623a7f856acull, 0x1cc0ddf7b3dc20faull, 0xcbc7f0bb8cd2ec14ull, 0x119c2409ce2462baull,
|
||||
0x1b544f117b80334cull, 0xdb47d6cdd3cf93fcull, 0xdb8c149c9914ddaeull, 0xc6c5e4b5ea412748ull, 0x4437723f548837ffull, 0x1fc1eb1bda8e067cull, 0x786dfa4dc7a84237ull, 0x1fc0ac22d46fb45dull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x82174c0fu, 0x577bdb9cu, 0x111053bfu, 0x2bb85514u, 0xd83e190fu, 0x8ccd9427u, 0x3f69d5e4u, 0xf3bbe1dcu,
|
||||
0xb3aa7e90u, 0xc33f5f73u, 0xb83a2b10u, 0xc4d7c8ffu, 0xefa6d1a8u, 0x7029b116u, 0x5e48fab0u, 0x0ef66e40u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xf78c84a4u;
|
||||
31
proto-cuda/packs/igneum-genesis/vectors.json
Normal file
31
proto-cuda/packs/igneum-genesis/vectors.json
Normal file
|
|
@ -0,0 +1,31 @@
|
|||
{
|
||||
"seed": "igneum-genesis",
|
||||
"day": "2026-10-03",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "proto-metal CPU interpreter (cpuWarp) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x2941e93c76cb1910", "0xc0099c34df8280ae", "0xe61da8a237797181", "0xfb944794eaa8ea1d", "0x0093934e41befb07", "0xd4ba2031c0385915", "0x75a8a43902cee166", "0xb4ce2d6829aaa27c",
|
||||
"0x478095c2e9911e5d", "0x595c8107b2792a66", "0x21b1b582752cc1d2", "0x16e746aa4d9532b8", "0x35129dbb17efe2e4", "0x926e0b679bf37640", "0xb2701dd71a189e19", "0x100e94f487022795",
|
||||
"0x2b9bc5aa1452be87", "0x8dabbbe8b909130b", "0x93a2cd4cba088297", "0x662c3e253432099e", "0xf489da4c9c770867", "0xe37864f932e97adf", "0x3ddf029de9b3c1ed", "0x9059da45130736ac",
|
||||
"0x1a667325a17e0016", "0x74723d71ab69828a", "0x121026c2f14795c1", "0x4989e1480c662cc8", "0xac3b8fd5307c0ab4", "0xdc3fc0bb843e84e8", "0xfa86df690d119639", "0x453388e1be04e25f"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x198f269c663acaf6", "0x6c125dcda3cade45", "0x0ae4eba3f2a0cdec", "0xc112952633d9477c", "0x2dfc52ebb7281f9e", "0x70b0e04375823959", "0x699e78e051dfaa76", "0xd91ae88bc2eba03f",
|
||||
"0x928b2939d0b23e68", "0x0e5feb811f4dd40b", "0xd0e4ce74e6693b53", "0xbd8ac2d12ea0ea1b", "0xa39bcd116eec4f88", "0xdd07f4f6a5a07514", "0xdc1cf8b4615b8842", "0xde0fb1599057c2dd",
|
||||
"0x56edb1995f7c341e", "0x02782ce22dd25939", "0x150b52d7b4b87319", "0xd5c89a9a6c2786c3", "0x8aef759270e00bba", "0xf9f26409968ea1ea", "0x2013ae93398863f0", "0x863877ca875900de",
|
||||
"0x6e05e03111bf02a3", "0x21ec57017818f221", "0x62b33a3456032a3b", "0x5bc24dc2bf442511", "0xb97041874e099be7", "0x2d9f09fb5d11fa13", "0x43f3f83f764ea247", "0xa3ec25a3072b734c"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0xa63f6d32a9dc2bac", "0x3263d16a85274409", "0xca7b34b1061437c7", "0xaf24083fbe9a33ac", "0x68c42544f8798a73", "0x5ba44c42278e0d0a", "0x12b31f0da8ea79e3", "0xa4a6186f5f5b7280",
|
||||
"0x1534d7e10b3f87ae", "0xb04618c9238e7436", "0x28be2d38f9db17d7", "0xebed3a6006a491d1", "0x23d50e8d3518ee39", "0xd75fac65bc8d1975", "0x3fd5c0a165c0669b", "0x882720cae6122082",
|
||||
"0xc107a8305a1f009d", "0xa58cc24784025f2c", "0x7ee69d59a87a92eb", "0x46009435f6f19482", "0xf50c2623a7f856ac", "0x1cc0ddf7b3dc20fa", "0xcbc7f0bb8cd2ec14", "0x119c2409ce2462ba",
|
||||
"0x1b544f117b80334c", "0xdb47d6cdd3cf93fc", "0xdb8c149c9914ddae", "0xc6c5e4b5ea412748", "0x4437723f548837ff", "0x1fc1eb1bda8e067c", "0x786dfa4dc7a84237", "0x1fc0ac22d46fb45d"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x82174c0f", "0x577bdb9c", "0x111053bf", "0x2bb85514", "0xd83e190f", "0x8ccd9427", "0x3f69d5e4", "0xf3bbe1dc", "0xb3aa7e90", "0xc33f5f73", "0xb83a2b10", "0xc4d7c8ff", "0xefa6d1a8", "0x7029b116", "0x5e48fab0", "0x0ef66e40"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0xf78c84a4"
|
||||
}
|
||||
145
proto-cuda/packs/igneum-hourly/kernel.cu
Normal file
145
proto-cuda/packs/igneum-hourly/kernel.cu
Normal file
|
|
@ -0,0 +1,145 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-hourly". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// dataset[i] = ds_elem(i, d0, d1) for i < n. Same closed form as the Metal igneum_fill kernel.
|
||||
__global__ void igneum_fill(uint32_t* ds, uint32_t n, uint32_t d0, uint32_t d1) {
|
||||
uint32_t i = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (i < n) ds[i] = ds_elem(i, d0, d1);
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x6bdee811u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x8f488bbeu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x8f488bbeu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xc5cdece7u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xc5cdece7u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x210af22du; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x210af22du; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x2f687b65u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0x2f687b65u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x17471eeeu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0x17471eeeu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xee16e284u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0xee16e284u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xfc9eb8f9u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0xfc9eb8f9u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x6bdee811u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r6 = r6 ^ ds[r5 & mask]; // 0 load
|
||||
r3 = r3 ^ r7; // 1 xor
|
||||
r6 = __umulhi(r6, r2); // 2 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 0u) & 1u) != 0u) ? 0x43f8f369u : 0x1eb46b1cu); // 3 add
|
||||
r3 = r3 ^ ds[r0 & mask]; // 4 load
|
||||
r5 = r5 | r7; // 5 or
|
||||
r4 = r4 ^ ds[r6 & mask]; // 6 load
|
||||
r4 = rotl_imm(r4, 21u); // 7 rotl
|
||||
r6 = r6 ^ ds[r3 & mask]; // 8 load
|
||||
r6 = r6 ^ ds[r1 & mask]; // 9 load
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 10 shfl
|
||||
r2 = r2 ^ r3; // 11 xor
|
||||
r2 = r2 + r7 + ((((sel >> 19u) & 1u) != 0u) ? 0xc26c7c2au : 0x3a1ce85eu); // 12 add
|
||||
r4 = r4 ^ ds[r3 & mask]; // 13 load
|
||||
r7 = r7 ^ ds[r1 & mask]; // 14 load
|
||||
r2 = __umulhi(r2, r3); // 15 mulhi
|
||||
r5 = r5 ^ ds[r2 & mask]; // 16 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 17 load
|
||||
r4 = r4 ^ ds[r7 & mask]; // 18 load
|
||||
r2 = rotr_var(r2, r1); // 19 rotr
|
||||
r7 = r7 ^ ds[r0 & mask]; // 20 load
|
||||
r4 = r4 ^ ds[r6 & mask]; // 21 load
|
||||
r7 = rotl_imm(r7, 25u); // 22 rotl
|
||||
r3 = r3 + r5 + ((((sel >> 15u) & 1u) != 0u) ? 0x98ae0055u : 0x942b819bu); // 23 add
|
||||
r3 = __umulhi(r3, r4); // 24 mulhi
|
||||
r6 = r6 ^ r1; // 25 xor
|
||||
r1 = rotl_imm(r1, 31u); // 26 rotl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 27 shfl
|
||||
r6 = r6 - r5; // 28 sub
|
||||
r6 = rotr_var(r6, r3); // 29 rotr
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r4, 4); // 30 shfl
|
||||
r4 = rotl_imm(r4, 30u); // 31 rotl
|
||||
r2 = r2 - r1; // 32 sub
|
||||
r5 = r5 | r4; // 33 or
|
||||
r7 = r6 * r3 + r7; // 34 mad
|
||||
r5 = r5 * r0; // 35 mul
|
||||
r5 = r5 - r3; // 36 sub
|
||||
r2 = r2 + r7 + ((((sel >> 5u) & 1u) != 0u) ? 0x6b5970b5u : 0x473ecfd5u); // 37 add
|
||||
r2 = r2 ^ ds[r7 & mask]; // 38 load
|
||||
r2 = rotr_var(r2, r6); // 39 rotr
|
||||
r0 = r0 ^ r5; // 40 xor
|
||||
r4 = r4 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 41 shfl
|
||||
r1 = r1 * r6; // 42 mul
|
||||
r0 = r4 * r1 + r0; // 43 mad
|
||||
r1 = r1 + r4 + ((((sel >> 20u) & 1u) != 0u) ? 0x4a502c22u : 0x04e78f3bu); // 44 add
|
||||
r6 = r2 * r7 + r6; // 45 mad
|
||||
r1 = r1 ^ r0; // 46 xor
|
||||
r5 = r5 ^ ds[r7 & mask]; // 47 load
|
||||
r0 = r0 | r4; // 48 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 49 shfl
|
||||
r7 = r7 + r0 + ((((sel >> 17u) & 1u) != 0u) ? 0x6fabf9ceu : 0x0a3df170u); // 50 add
|
||||
r6 = r6 ^ r3; // 51 xor
|
||||
r1 = r1 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xad344ca0u : 0xc99bce6fu); // 52 add
|
||||
r6 = r6 * r7; // 53 mul
|
||||
r3 = __umulhi(r3, r4); // 54 mulhi
|
||||
r7 = r7 * r1; // 55 mul
|
||||
r7 = r7 ^ r1; // 56 xor
|
||||
r2 = r2 * r7; // 57 mul
|
||||
r2 = r2 ^ ds[r1 & mask]; // 58 load
|
||||
r7 = r4 * r5 + r7; // 59 mad
|
||||
r2 = r2 ^ ds[r7 & mask]; // 60 load
|
||||
r0 = r2 * r3 + r0; // 61 mad
|
||||
r1 = __umulhi(r1, r5); // 62 mulhi
|
||||
r7 = r7 + r3 + ((((sel >> 24u) & 1u) != 0u) ? 0x05e4fc1du : 0xc23e27c9u); // 63 add
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_fill(uint32_t* ds, uint32_t nWords, uint32_t d0, uint32_t d1) {
|
||||
if (nWords == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nWords + block - 1u) / block;
|
||||
igneum_fill<<<grid, block>>>(ds, nWords, d0, d1);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
25
proto-cuda/packs/igneum-hourly/program.h
Normal file
25
proto-cuda/packs/igneum-hourly/program.h
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-hourly". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
#pragma once
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-hourly"
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_OP_MIX "load=16 add=8 xor=7 mad=5 mul=5 mulhi=5 shfl=5 rotl=4 or=3 rotr=3 sub=3"
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x6bdee811u, 0x8f488bbeu, 0xc5cdece7u, 0x210af22du, 0x2f687b65u, 0x17471eeeu, 0xee16e284u, 0xfc9eb8f9u }
|
||||
|
||||
// Defined in kernel.cu. Both launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_fill(uint32_t* ds, uint32_t nWords, uint32_t d0, uint32_t d1);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
105
proto-cuda/packs/igneum-hourly/program.json
Normal file
105
proto-cuda/packs/igneum-hourly/program.json
Normal file
|
|
@ -0,0 +1,105 @@
|
|||
{
|
||||
"format": "igneum-program-pack-1",
|
||||
"seed": "igneum-hourly",
|
||||
"seed_words": ["0x6bdee811", "0x8f488bbe", "0xc5cdece7", "0x210af22d", "0x2f687b65", "0x17471eee", "0xee16e284", "0xfc9eb8f9"],
|
||||
"seed_derivation": "FNV-1a 64 over UTF-8 of seed, basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 128,
|
||||
"op_mix": {"load": 16, "add": 8, "xor": 7, "mad": 5, "mul": 5, "mulhi": 5, "shfl": 5, "rotl": 4, "or": 3, "rotr": 3, "sub": 3},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
|
||||
"op_semantics": {
|
||||
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
|
||||
"sub": "dst = dst - src",
|
||||
"mul": "dst = dst * src (low 32)",
|
||||
"mulhi": "dst = high 32 bits of dst * src",
|
||||
"xor": "dst = dst ^ src",
|
||||
"or": "dst = dst | src",
|
||||
"rotl": "dst = rotl(dst, rot), rot in 1..31",
|
||||
"rotr": "dst = rotr(dst, src & 31)",
|
||||
"mad": "dst = src * src2 + dst",
|
||||
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
|
||||
"load": "dst = dst ^ dataset[src & dataset.mask]"
|
||||
},
|
||||
"dataset": {
|
||||
"log2_words": 28,
|
||||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "2026-10-03",
|
||||
"day_words_from": "day/2026-10-03",
|
||||
"d0": "0x3067619f",
|
||||
"d1": "0x3c269176",
|
||||
"formula": "x = i ^ d0; x *= 0x9E3779B1; x ^= x>>15; x += d1; x *= 0x85EBCA77; x ^= x>>13; x *= 0xC2B2AE3D; x ^= x>>16 (all mod 2^32)"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "load", "dst": 6, "src": 5, "src2": 1, "imm": "0xf9ca5168", "imm2": "0xd27a709a", "rot": 1, "bit": 24, "mask": 16},
|
||||
{"i": 1, "op": "xor", "dst": 3, "src": 7, "src2": 0, "imm": "0x4731bc24", "imm2": "0x8c57bc59", "rot": 23, "bit": 3, "mask": 16},
|
||||
{"i": 2, "op": "mulhi", "dst": 6, "src": 2, "src2": 7, "imm": "0xc05a41e5", "imm2": "0xb5d20e75", "rot": 3, "bit": 7, "mask": 1},
|
||||
{"i": 3, "op": "add", "dst": 1, "src": 0, "src2": 0, "imm": "0x1eb46b1c", "imm2": "0x43f8f369", "rot": 31, "bit": 0, "mask": 8},
|
||||
{"i": 4, "op": "load", "dst": 3, "src": 0, "src2": 0, "imm": "0x8b6ea7af", "imm2": "0xa02a6965", "rot": 10, "bit": 22, "mask": 1},
|
||||
{"i": 5, "op": "or", "dst": 5, "src": 7, "src2": 7, "imm": "0x440ba791", "imm2": "0xf46d731a", "rot": 20, "bit": 14, "mask": 4},
|
||||
{"i": 6, "op": "load", "dst": 4, "src": 6, "src2": 7, "imm": "0xc7fa7a14", "imm2": "0xf0426f08", "rot": 16, "bit": 11, "mask": 4},
|
||||
{"i": 7, "op": "rotl", "dst": 4, "src": 3, "src2": 3, "imm": "0x535a367f", "imm2": "0x1d053f10", "rot": 21, "bit": 0, "mask": 4},
|
||||
{"i": 8, "op": "load", "dst": 6, "src": 3, "src2": 2, "imm": "0x0407ee13", "imm2": "0x691717fb", "rot": 20, "bit": 27, "mask": 8},
|
||||
{"i": 9, "op": "load", "dst": 6, "src": 1, "src2": 1, "imm": "0x8735ec1e", "imm2": "0xea050d61", "rot": 12, "bit": 26, "mask": 4},
|
||||
{"i": 10, "op": "shfl", "dst": 0, "src": 3, "src2": 2, "imm": "0x9bdf14a9", "imm2": "0xba78183e", "rot": 11, "bit": 20, "mask": 8},
|
||||
{"i": 11, "op": "xor", "dst": 2, "src": 3, "src2": 7, "imm": "0xf2010dad", "imm2": "0x52b64f96", "rot": 8, "bit": 0, "mask": 4},
|
||||
{"i": 12, "op": "add", "dst": 2, "src": 7, "src2": 1, "imm": "0x3a1ce85e", "imm2": "0xc26c7c2a", "rot": 10, "bit": 19, "mask": 4},
|
||||
{"i": 13, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0x816ece48", "imm2": "0xa0c6535a", "rot": 18, "bit": 21, "mask": 1},
|
||||
{"i": 14, "op": "load", "dst": 7, "src": 1, "src2": 3, "imm": "0x1b89fffb", "imm2": "0xf0330d16", "rot": 2, "bit": 0, "mask": 16},
|
||||
{"i": 15, "op": "mulhi", "dst": 2, "src": 3, "src2": 2, "imm": "0xcc6a5993", "imm2": "0xb853c4da", "rot": 28, "bit": 23, "mask": 4},
|
||||
{"i": 16, "op": "load", "dst": 5, "src": 2, "src2": 1, "imm": "0x14489190", "imm2": "0x4c71c842", "rot": 29, "bit": 7, "mask": 1},
|
||||
{"i": 17, "op": "load", "dst": 5, "src": 1, "src2": 1, "imm": "0x3ecd1efe", "imm2": "0xe582d071", "rot": 25, "bit": 29, "mask": 2},
|
||||
{"i": 18, "op": "load", "dst": 4, "src": 7, "src2": 6, "imm": "0xccfb0ae5", "imm2": "0xee788287", "rot": 3, "bit": 17, "mask": 16},
|
||||
{"i": 19, "op": "rotr", "dst": 2, "src": 1, "src2": 5, "imm": "0xc9ac3074", "imm2": "0x4d5eaabf", "rot": 19, "bit": 29, "mask": 16},
|
||||
{"i": 20, "op": "load", "dst": 7, "src": 0, "src2": 0, "imm": "0x637ba816", "imm2": "0xdeb9f971", "rot": 25, "bit": 2, "mask": 1},
|
||||
{"i": 21, "op": "load", "dst": 4, "src": 6, "src2": 4, "imm": "0x59d5f45d", "imm2": "0x701a90ad", "rot": 9, "bit": 2, "mask": 1},
|
||||
{"i": 22, "op": "rotl", "dst": 7, "src": 3, "src2": 5, "imm": "0xc6e87564", "imm2": "0x7aee3663", "rot": 25, "bit": 23, "mask": 4},
|
||||
{"i": 23, "op": "add", "dst": 3, "src": 5, "src2": 4, "imm": "0x942b819b", "imm2": "0x98ae0055", "rot": 8, "bit": 15, "mask": 2},
|
||||
{"i": 24, "op": "mulhi", "dst": 3, "src": 4, "src2": 1, "imm": "0x4b3f70bb", "imm2": "0x1c24bab9", "rot": 13, "bit": 18, "mask": 1},
|
||||
{"i": 25, "op": "xor", "dst": 6, "src": 1, "src2": 7, "imm": "0x82815069", "imm2": "0xecdc8c4c", "rot": 3, "bit": 6, "mask": 4},
|
||||
{"i": 26, "op": "rotl", "dst": 1, "src": 0, "src2": 0, "imm": "0x33471012", "imm2": "0x9ce1a3e2", "rot": 31, "bit": 7, "mask": 8},
|
||||
{"i": 27, "op": "shfl", "dst": 3, "src": 6, "src2": 6, "imm": "0x3a18a1b5", "imm2": "0xc49b3103", "rot": 18, "bit": 1, "mask": 4},
|
||||
{"i": 28, "op": "sub", "dst": 6, "src": 5, "src2": 1, "imm": "0x79e2cc72", "imm2": "0x4f9dab61", "rot": 5, "bit": 9, "mask": 16},
|
||||
{"i": 29, "op": "rotr", "dst": 6, "src": 3, "src2": 6, "imm": "0x5bea26e5", "imm2": "0x2aa1df19", "rot": 18, "bit": 14, "mask": 16},
|
||||
{"i": 30, "op": "shfl", "dst": 0, "src": 4, "src2": 7, "imm": "0xb954881e", "imm2": "0xa6eeddea", "rot": 1, "bit": 29, "mask": 4},
|
||||
{"i": 31, "op": "rotl", "dst": 4, "src": 5, "src2": 4, "imm": "0x46a02201", "imm2": "0x67e4f5b6", "rot": 30, "bit": 27, "mask": 4},
|
||||
{"i": 32, "op": "sub", "dst": 2, "src": 1, "src2": 6, "imm": "0x2b8c1fcd", "imm2": "0xf1e88c21", "rot": 9, "bit": 18, "mask": 8},
|
||||
{"i": 33, "op": "or", "dst": 5, "src": 4, "src2": 2, "imm": "0xb3fcf9dd", "imm2": "0x948436e1", "rot": 21, "bit": 4, "mask": 2},
|
||||
{"i": 34, "op": "mad", "dst": 7, "src": 6, "src2": 3, "imm": "0xc076bcd6", "imm2": "0x224748b2", "rot": 29, "bit": 9, "mask": 2},
|
||||
{"i": 35, "op": "mul", "dst": 5, "src": 0, "src2": 2, "imm": "0xd7f50971", "imm2": "0x4e7ada3e", "rot": 14, "bit": 12, "mask": 2},
|
||||
{"i": 36, "op": "sub", "dst": 5, "src": 3, "src2": 6, "imm": "0x349ca5cf", "imm2": "0x82b0e280", "rot": 15, "bit": 27, "mask": 2},
|
||||
{"i": 37, "op": "add", "dst": 2, "src": 7, "src2": 5, "imm": "0x473ecfd5", "imm2": "0x6b5970b5", "rot": 25, "bit": 5, "mask": 1},
|
||||
{"i": 38, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0xefc24111", "imm2": "0x1be7ad5b", "rot": 6, "bit": 3, "mask": 2},
|
||||
{"i": 39, "op": "rotr", "dst": 2, "src": 6, "src2": 4, "imm": "0x2eade1b0", "imm2": "0xe8e01fb0", "rot": 13, "bit": 30, "mask": 4},
|
||||
{"i": 40, "op": "xor", "dst": 0, "src": 5, "src2": 6, "imm": "0xdaf9a7a2", "imm2": "0x89165d3f", "rot": 1, "bit": 3, "mask": 1},
|
||||
{"i": 41, "op": "shfl", "dst": 4, "src": 5, "src2": 6, "imm": "0xd100247c", "imm2": "0x9697c51f", "rot": 27, "bit": 17, "mask": 8},
|
||||
{"i": 42, "op": "mul", "dst": 1, "src": 6, "src2": 7, "imm": "0xebbb5184", "imm2": "0xf79998f9", "rot": 15, "bit": 16, "mask": 4},
|
||||
{"i": 43, "op": "mad", "dst": 0, "src": 4, "src2": 1, "imm": "0x83ee60eb", "imm2": "0xbc017371", "rot": 26, "bit": 21, "mask": 2},
|
||||
{"i": 44, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0x04e78f3b", "imm2": "0x4a502c22", "rot": 20, "bit": 20, "mask": 4},
|
||||
{"i": 45, "op": "mad", "dst": 6, "src": 2, "src2": 7, "imm": "0xba1891ac", "imm2": "0xe7992a86", "rot": 27, "bit": 25, "mask": 2},
|
||||
{"i": 46, "op": "xor", "dst": 1, "src": 0, "src2": 1, "imm": "0xbd2c89d9", "imm2": "0x0bd3d33b", "rot": 29, "bit": 15, "mask": 8},
|
||||
{"i": 47, "op": "load", "dst": 5, "src": 7, "src2": 5, "imm": "0x67776361", "imm2": "0x20a4c205", "rot": 17, "bit": 13, "mask": 8},
|
||||
{"i": 48, "op": "or", "dst": 0, "src": 4, "src2": 1, "imm": "0x6b0e9ba2", "imm2": "0xf52a7752", "rot": 7, "bit": 21, "mask": 16},
|
||||
{"i": 49, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x03494dfd", "imm2": "0x9c326606", "rot": 21, "bit": 22, "mask": 16},
|
||||
{"i": 50, "op": "add", "dst": 7, "src": 0, "src2": 2, "imm": "0x0a3df170", "imm2": "0x6fabf9ce", "rot": 5, "bit": 17, "mask": 16},
|
||||
{"i": 51, "op": "xor", "dst": 6, "src": 3, "src2": 6, "imm": "0x0c22f139", "imm2": "0xe642b0be", "rot": 27, "bit": 24, "mask": 1},
|
||||
{"i": 52, "op": "add", "dst": 1, "src": 2, "src2": 6, "imm": "0xc99bce6f", "imm2": "0xad344ca0", "rot": 31, "bit": 21, "mask": 2},
|
||||
{"i": 53, "op": "mul", "dst": 6, "src": 7, "src2": 1, "imm": "0x85e422cd", "imm2": "0x95145330", "rot": 27, "bit": 21, "mask": 1},
|
||||
{"i": 54, "op": "mulhi", "dst": 3, "src": 4, "src2": 4, "imm": "0xf61f4490", "imm2": "0x2e3345ba", "rot": 5, "bit": 14, "mask": 2},
|
||||
{"i": 55, "op": "mul", "dst": 7, "src": 1, "src2": 1, "imm": "0x2266ef41", "imm2": "0x81c0e541", "rot": 1, "bit": 27, "mask": 16},
|
||||
{"i": 56, "op": "xor", "dst": 7, "src": 1, "src2": 3, "imm": "0x5799668a", "imm2": "0x381b80ef", "rot": 13, "bit": 17, "mask": 8},
|
||||
{"i": 57, "op": "mul", "dst": 2, "src": 7, "src2": 2, "imm": "0x221abae1", "imm2": "0x8bf02d19", "rot": 31, "bit": 2, "mask": 8},
|
||||
{"i": 58, "op": "load", "dst": 2, "src": 1, "src2": 5, "imm": "0xf27b8255", "imm2": "0x820332d2", "rot": 14, "bit": 23, "mask": 16},
|
||||
{"i": 59, "op": "mad", "dst": 7, "src": 4, "src2": 5, "imm": "0xb6772f7b", "imm2": "0x07904b83", "rot": 14, "bit": 0, "mask": 8},
|
||||
{"i": 60, "op": "load", "dst": 2, "src": 7, "src2": 5, "imm": "0x6760e166", "imm2": "0x2e5a2418", "rot": 30, "bit": 12, "mask": 4},
|
||||
{"i": 61, "op": "mad", "dst": 0, "src": 2, "src2": 3, "imm": "0x3a2d60b7", "imm2": "0x4ec18be1", "rot": 15, "bit": 22, "mask": 1},
|
||||
{"i": 62, "op": "mulhi", "dst": 1, "src": 5, "src2": 6, "imm": "0xf423297d", "imm2": "0xd87f149c", "rot": 28, "bit": 6, "mask": 4},
|
||||
{"i": 63, "op": "add", "dst": 7, "src": 3, "src2": 3, "imm": "0xc23e27c9", "imm2": "0x05e4fc1d", "rot": 26, "bit": 24, "mask": 4}
|
||||
]
|
||||
}
|
||||
109
proto-cuda/packs/igneum-hourly/program.metal
Normal file
109
proto-cuda/packs/igneum-hourly/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x6bdee811u, 0x8f488bbeu, 0xc5cdece7u, 0x210af22du, 0x2f687b65u, 0x17471eeeu, 0xee16e284u, 0xfc9eb8f9u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r6 = r6 ^ dataset[r5 & MASK]; // 0
|
||||
r3 = r3 ^ r7; // 1
|
||||
r6 = mulhi(r6, r2); // 2
|
||||
r1 = r1 + r0 + select(0x1eb46b1cu, 0x43f8f369u, ((sel >> 0u) & 1u) != 0u); // 3
|
||||
r3 = r3 ^ dataset[r0 & MASK]; // 4
|
||||
r5 = r5 | r7; // 5
|
||||
r4 = r4 ^ dataset[r6 & MASK]; // 6
|
||||
r4 = rotl_imm(r4, 21u); // 7
|
||||
r6 = r6 ^ dataset[r3 & MASK]; // 8
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 9
|
||||
r0 = r0 ^ simd_shuffle_xor(r3, (ushort)8); // 10
|
||||
r2 = r2 ^ r3; // 11
|
||||
r2 = r2 + r7 + select(0x3a1ce85eu, 0xc26c7c2au, ((sel >> 19u) & 1u) != 0u); // 12
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 13
|
||||
r7 = r7 ^ dataset[r1 & MASK]; // 14
|
||||
r2 = mulhi(r2, r3); // 15
|
||||
r5 = r5 ^ dataset[r2 & MASK]; // 16
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 17
|
||||
r4 = r4 ^ dataset[r7 & MASK]; // 18
|
||||
r2 = rotr_var(r2, r1); // 19
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 20
|
||||
r4 = r4 ^ dataset[r6 & MASK]; // 21
|
||||
r7 = rotl_imm(r7, 25u); // 22
|
||||
r3 = r3 + r5 + select(0x942b819bu, 0x98ae0055u, ((sel >> 15u) & 1u) != 0u); // 23
|
||||
r3 = mulhi(r3, r4); // 24
|
||||
r6 = r6 ^ r1; // 25
|
||||
r1 = rotl_imm(r1, 31u); // 26
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)4); // 27
|
||||
r6 = r6 - r5; // 28
|
||||
r6 = rotr_var(r6, r3); // 29
|
||||
r0 = r0 ^ simd_shuffle_xor(r4, (ushort)4); // 30
|
||||
r4 = rotl_imm(r4, 30u); // 31
|
||||
r2 = r2 - r1; // 32
|
||||
r5 = r5 | r4; // 33
|
||||
r7 = r6 * r3 + r7; // 34
|
||||
r5 = r5 * r0; // 35
|
||||
r5 = r5 - r3; // 36
|
||||
r2 = r2 + r7 + select(0x473ecfd5u, 0x6b5970b5u, ((sel >> 5u) & 1u) != 0u); // 37
|
||||
r2 = r2 ^ dataset[r7 & MASK]; // 38
|
||||
r2 = rotr_var(r2, r6); // 39
|
||||
r0 = r0 ^ r5; // 40
|
||||
r4 = r4 ^ simd_shuffle_xor(r5, (ushort)8); // 41
|
||||
r1 = r1 * r6; // 42
|
||||
r0 = r4 * r1 + r0; // 43
|
||||
r1 = r1 + r4 + select(0x04e78f3bu, 0x4a502c22u, ((sel >> 20u) & 1u) != 0u); // 44
|
||||
r6 = r2 * r7 + r6; // 45
|
||||
r1 = r1 ^ r0; // 46
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 47
|
||||
r0 = r0 | r4; // 48
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)16); // 49
|
||||
r7 = r7 + r0 + select(0x0a3df170u, 0x6fabf9ceu, ((sel >> 17u) & 1u) != 0u); // 50
|
||||
r6 = r6 ^ r3; // 51
|
||||
r1 = r1 + r2 + select(0xc99bce6fu, 0xad344ca0u, ((sel >> 21u) & 1u) != 0u); // 52
|
||||
r6 = r6 * r7; // 53
|
||||
r3 = mulhi(r3, r4); // 54
|
||||
r7 = r7 * r1; // 55
|
||||
r7 = r7 ^ r1; // 56
|
||||
r2 = r2 * r7; // 57
|
||||
r2 = r2 ^ dataset[r1 & MASK]; // 58
|
||||
r7 = r4 * r5 + r7; // 59
|
||||
r2 = r2 ^ dataset[r7 & MASK]; // 60
|
||||
r0 = r2 * r3 + r0; // 61
|
||||
r1 = mulhi(r1, r5); // 62
|
||||
r7 = r7 + r3 + select(0xc23e27c9u, 0x05e4fc1du, ((sel >> 24u) & 1u) != 0u); // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
35
proto-cuda/packs/igneum-hourly/vectors.h
Normal file
35
proto-cuda/packs/igneum-hourly/vectors.h
Normal file
|
|
@ -0,0 +1,35 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-hourly". Do not edit by hand.
|
||||
// Expected outputs: proto-metal CPU interpreter (cpuWarp) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps
|
||||
#pragma once
|
||||
#include <cstdint>
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x787a737506455bbeull, 0xf93d248e364f7e1eull, 0x8a7ce1701b5d0489ull, 0xd4f5b2177c414f3full, 0xe88fdf6fb22aee5cull, 0xd68bc0baf30b41fcull, 0x87f6a81eb0a6ac69ull, 0xb3cd8b6e6e57b347ull,
|
||||
0x2eef97b916c626bfull, 0x4ded79caeecc1b81ull, 0xac5141672d56faaaull, 0xe5c20290edf36932ull, 0x8358bbb9c3fa668dull, 0x3394e6f0736ecd30ull, 0xef402d26f32ed7cfull, 0x0ce267a5c3e18a42ull,
|
||||
0xd92370857d7ec52bull, 0x4f8cd5815e18afbeull, 0x2fbf597253c4f35eull, 0x747180e71a09f68dull, 0x2238f847f0ed84dcull, 0xf9c277614e6b908dull, 0x39dbe66723750ad4ull, 0x8dc8b244ac675fbbull,
|
||||
0x02f535158b9d7c8eull, 0x71c6a2147fb4787aull, 0xcb3a0d709ec6edb6ull, 0x560cdae1ccc99cd7ull, 0xb8e12af183034f40ull, 0xe3088768bbf66638ull, 0x62b3109b60c3103dull, 0x0bb06c1609ff43d5ull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x37bb8728f884e6b8ull, 0x14be2ec14497f818ull, 0x2086e0b6e0d72b74ull, 0x50b1159a7a48b25aull, 0xc7e8af0d47ae9524ull, 0x11b6c5919a802722ull, 0x47191c35d09f473eull, 0x5a17776c88350cbcull,
|
||||
0xf8b63bb823a85140ull, 0x7d1536d3c3a77a4aull, 0x3fd4978fa97a3170ull, 0xb93a559eef3ca040ull, 0x3ab4c9b478dc4f70ull, 0xe383445b250ec1c5ull, 0x4b8c339a4059d0eaull, 0xb9edfbab431dc7dcull,
|
||||
0x67e73ea59e305c36ull, 0x85ad26bb831207a3ull, 0x7fa9ca66562cd371ull, 0x763b3ca6c6efdf75ull, 0x6f02716b104a47ceull, 0x033d54978e01388full, 0xed13f0a661b95976ull, 0x4cf586279408f177ull,
|
||||
0x339f105764d3a0dfull, 0xf5f5a600702c84e7ull, 0xbec42d884060c3ebull, 0x5e687d7d88c4dc86ull, 0xfe9202424271b313ull, 0xe9a93837f53818ddull, 0x2efe4cd7be3bc6ccull, 0x2718ca6898992e25ull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x3ea4d63013ee1447ull, 0x3920ed5486c12342ull, 0x6d71c6114a18c926ull, 0x76851391b25e9e27ull, 0x7552ab2e8ba173e4ull, 0x2ed7c0fe0ed382e7ull, 0x3965791801d021b2ull, 0x5638111a801de9f4ull,
|
||||
0xffdb8b13126c51efull, 0xec7513be841d1ee6ull, 0x5ee78112615a43a0ull, 0x447c921628b22b82ull, 0xa8683cc17b258e26ull, 0xc4161c3fb96be5f0ull, 0xf4bde10f52aef480ull, 0xfb9a90df739d4b7full,
|
||||
0x8b6a11be7a8067c6ull, 0x0c5e0b16f9dddd20ull, 0xad830d9f20c0cf41ull, 0x7b45608e60aa036cull, 0x871e0fc8a4c82071ull, 0x087bba8af99e5f89ull, 0xc9b353a02ccf042dull, 0xa8b5e4ca3e95f99eull,
|
||||
0x1f636af31962da8full, 0x4e3c7276699005ccull, 0xeb138ecf3dfdc723ull, 0x572fe00fa91ec1f0ull, 0x7eb7d1613efc2fb6ull, 0x4ec9b30dfb44aa0full, 0x8883fe27c6464e60ull, 0xca9d04210422e8b8ull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x82174c0fu, 0x577bdb9cu, 0x111053bfu, 0x2bb85514u, 0xd83e190fu, 0x8ccd9427u, 0x3f69d5e4u, 0xf3bbe1dcu,
|
||||
0xb3aa7e90u, 0xc33f5f73u, 0xb83a2b10u, 0xc4d7c8ffu, 0xefa6d1a8u, 0x7029b116u, 0x5e48fab0u, 0x0ef66e40u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xf78c84a4u;
|
||||
31
proto-cuda/packs/igneum-hourly/vectors.json
Normal file
31
proto-cuda/packs/igneum-hourly/vectors.json
Normal file
|
|
@ -0,0 +1,31 @@
|
|||
{
|
||||
"seed": "igneum-hourly",
|
||||
"day": "2026-10-03",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "proto-metal CPU interpreter (cpuWarp) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x787a737506455bbe", "0xf93d248e364f7e1e", "0x8a7ce1701b5d0489", "0xd4f5b2177c414f3f", "0xe88fdf6fb22aee5c", "0xd68bc0baf30b41fc", "0x87f6a81eb0a6ac69", "0xb3cd8b6e6e57b347",
|
||||
"0x2eef97b916c626bf", "0x4ded79caeecc1b81", "0xac5141672d56faaa", "0xe5c20290edf36932", "0x8358bbb9c3fa668d", "0x3394e6f0736ecd30", "0xef402d26f32ed7cf", "0x0ce267a5c3e18a42",
|
||||
"0xd92370857d7ec52b", "0x4f8cd5815e18afbe", "0x2fbf597253c4f35e", "0x747180e71a09f68d", "0x2238f847f0ed84dc", "0xf9c277614e6b908d", "0x39dbe66723750ad4", "0x8dc8b244ac675fbb",
|
||||
"0x02f535158b9d7c8e", "0x71c6a2147fb4787a", "0xcb3a0d709ec6edb6", "0x560cdae1ccc99cd7", "0xb8e12af183034f40", "0xe3088768bbf66638", "0x62b3109b60c3103d", "0x0bb06c1609ff43d5"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x37bb8728f884e6b8", "0x14be2ec14497f818", "0x2086e0b6e0d72b74", "0x50b1159a7a48b25a", "0xc7e8af0d47ae9524", "0x11b6c5919a802722", "0x47191c35d09f473e", "0x5a17776c88350cbc",
|
||||
"0xf8b63bb823a85140", "0x7d1536d3c3a77a4a", "0x3fd4978fa97a3170", "0xb93a559eef3ca040", "0x3ab4c9b478dc4f70", "0xe383445b250ec1c5", "0x4b8c339a4059d0ea", "0xb9edfbab431dc7dc",
|
||||
"0x67e73ea59e305c36", "0x85ad26bb831207a3", "0x7fa9ca66562cd371", "0x763b3ca6c6efdf75", "0x6f02716b104a47ce", "0x033d54978e01388f", "0xed13f0a661b95976", "0x4cf586279408f177",
|
||||
"0x339f105764d3a0df", "0xf5f5a600702c84e7", "0xbec42d884060c3eb", "0x5e687d7d88c4dc86", "0xfe9202424271b313", "0xe9a93837f53818dd", "0x2efe4cd7be3bc6cc", "0x2718ca6898992e25"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0x3ea4d63013ee1447", "0x3920ed5486c12342", "0x6d71c6114a18c926", "0x76851391b25e9e27", "0x7552ab2e8ba173e4", "0x2ed7c0fe0ed382e7", "0x3965791801d021b2", "0x5638111a801de9f4",
|
||||
"0xffdb8b13126c51ef", "0xec7513be841d1ee6", "0x5ee78112615a43a0", "0x447c921628b22b82", "0xa8683cc17b258e26", "0xc4161c3fb96be5f0", "0xf4bde10f52aef480", "0xfb9a90df739d4b7f",
|
||||
"0x8b6a11be7a8067c6", "0x0c5e0b16f9dddd20", "0xad830d9f20c0cf41", "0x7b45608e60aa036c", "0x871e0fc8a4c82071", "0x087bba8af99e5f89", "0xc9b353a02ccf042d", "0xa8b5e4ca3e95f99e",
|
||||
"0x1f636af31962da8f", "0x4e3c7276699005cc", "0xeb138ecf3dfdc723", "0x572fe00fa91ec1f0", "0x7eb7d1613efc2fb6", "0x4ec9b30dfb44aa0f", "0x8883fe27c6464e60", "0xca9d04210422e8b8"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x82174c0f", "0x577bdb9c", "0x111053bf", "0x2bb85514", "0xd83e190f", "0x8ccd9427", "0x3f69d5e4", "0xf3bbe1dc", "0xb3aa7e90", "0xc33f5f73", "0xb83a2b10", "0xc4d7c8ff", "0xefa6d1a8", "0x7029b116", "0x5e48fab0", "0x0ef66e40"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0xf78c84a4"
|
||||
}
|
||||
120
proto-metal/README.md
Normal file
120
proto-metal/README.md
Normal file
|
|
@ -0,0 +1,120 @@
|
|||
# igneum-bench (proto-metal)
|
||||
|
||||
First prototype of Igneum's random-program GPU proof-of-work, on Apple Metal.
|
||||
One Swift file, no packages, no Xcode. The Metal kernel is generated as text from a seed
|
||||
and compiled at runtime with `MTLDevice.makeLibrary(source:options:)`.
|
||||
|
||||
## What it does
|
||||
|
||||
1. Derives a 32-byte seed from a string (FNV-1a 64, four salts).
|
||||
2. Generates a program of 64 integer instructions over 8 x uint32 lane registers, run for 8 iterations.
|
||||
Ops: add, sub, mul, mulhi, xor, or, rotl (immediate), rotr (register), mad, shfl_xor (simd_shuffle_xor
|
||||
across the 32-lane SIMD group, masks 1 to 16), load (`dst ^= dataset[src & MASK]`). Load weight 25 percent.
|
||||
Each add picks one of two immediates from a bit of r0 sampled at the top of the iteration, branchless `select`.
|
||||
3. Emits Metal Shading Language, compiles it, runs it with threadgroup size 32 (one SIMD group per threadgroup).
|
||||
4. Fills a 1 GiB dataset (2^28 uint32) on the GPU from a closed-form function of (daySeed, index).
|
||||
The CPU computes any element on demand and never holds the dataset.
|
||||
5. Interprets the same program on the CPU for three 32-lane warps and compares every output bit for bit.
|
||||
6. `--hours N` regenerates and recompiles N programs in sequence (the hourly epoch model). Default N is 2 so
|
||||
verification always covers two seeds.
|
||||
|
||||
## Build
|
||||
|
||||
```
|
||||
cd proto-metal
|
||||
swiftc -O -o igneum-bench main.swift -framework Metal
|
||||
```
|
||||
|
||||
Tested with Swift 5.8.1 from Command Line Tools on macOS (Darwin 25.6.0), no Xcode.
|
||||
|
||||
## Run
|
||||
|
||||
```
|
||||
./igneum-bench # default seed, 1 GiB dataset, 4 batches of 2^22 nonces, 2 epochs
|
||||
./igneum-bench --seed "my-seed" # a different program
|
||||
./igneum-bench --hours 3 # three seeds in sequence
|
||||
./igneum-bench --dataset-log2 20 --batch-log2 16 # small, fast validation
|
||||
./igneum-bench --dump ./generated # also write the generated .metal source
|
||||
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis # CUDA program pack
|
||||
```
|
||||
|
||||
Flags: `--seed`, `--day`, `--hours`, `--batch-log2` (default 22), `--batches` (default 4),
|
||||
`--dataset-log2` (default 28), `--verify-warps` (default 3), `--dump <dir>`, `--export-pack <dir>`.
|
||||
Exit code 0 means every verified warp matched.
|
||||
|
||||
`--export-pack <dir>` (added 3 October 2026) does not run the bench. It generates the program for `--seed`,
|
||||
emits it a second time as a CUDA kernel, computes expected outputs for 3 warps (base nonces 0, 4096, 1000000)
|
||||
with the CPU interpreter, cross-checks them on the Metal GPU, and writes `kernel.cu`, `program.h`, `vectors.h`,
|
||||
`program.json`, `vectors.json` and `program.metal` into the directory. See `../proto-cuda/README.md`.
|
||||
|
||||
## Measured on this machine, 3 October 2026
|
||||
|
||||
Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory. `threadExecutionWidth` reported as 32.
|
||||
Rates are wall-clock over 4 batches x 4,194,304 hashes after one warm-up batch. GPU-timestamp rates
|
||||
agreed with wall-clock to within 0.1 percent in every run. "GB/s useful" is loads per hash x 4 bytes x
|
||||
hashes per second; it counts only the 4 bytes the program consumed per load, not the cache line moved.
|
||||
|
||||
| Seed | Loads/hash | Compile ms (library + pipeline) | Mhash/s | GB/s useful | CPU verify ms/warp (avg of 20) | Verify |
|
||||
|---|---|---|---|---|---|---|
|
||||
| igneum-genesis | 104 | 49.6 (cold) | 45.2 | 18.8 | 0.015 | PASS 3/3 warps |
|
||||
| igneum-genesis/epoch1 | 104 | 20.3 | 48.4 | 20.1 | 0.019 | PASS 3/3 warps |
|
||||
| igneum-second-seed | 104 | 46.6 (cold) | 35.5 | 14.8 | 0.016 | PASS 3/3 warps |
|
||||
| igneum-second-seed/epoch1 | 144 | 18.7 | 35.4 | 20.4 | 0.017 | PASS 3/3 warps |
|
||||
| igneum-hourly | 128 | 52.0 (cold) | 36.6 | 18.7 | 0.021 | PASS 3/3 warps |
|
||||
| igneum-hourly/epoch1 | 128 | 21.6 | 37.5 | 19.2 | 0.017 | PASS 3/3 warps |
|
||||
| igneum-hourly/epoch2 | 120 | 23.7 | 36.6 | 17.6 | 0.016 | PASS 3/3 warps |
|
||||
|
||||
Single-run CPU verify times (no repetition) ranged 0.015 to 0.041 ms per warp. Verified warps per
|
||||
program: warp 0, warp 65537, warp 131071 of batch 0 (nonces 0 to 31, 2097184 to 2097215, 4194272 to 4194303).
|
||||
21 warps, 672 hashes, zero mismatches.
|
||||
|
||||
Dataset fill (1 GiB, GPU timestamps): 5.57 ms on the first run of a process (179 GB/s), 2.34 to 2.35 ms on later
|
||||
runs (427 GB/s). The gap is probably first-touch page mapping of the private buffer; not measured further.
|
||||
Compile of the fixed fill kernel: 227 ms the very first time on this Mac, 0.7 to 0.9 ms afterwards, which looks
|
||||
like the system shader cache.
|
||||
|
||||
### Dataset size sweep, seed igneum-genesis, same program throughout
|
||||
|
||||
| Dataset | Mhash/s | GB/s useful | Random loads/s |
|
||||
|---|---|---|---|
|
||||
| 4 MiB (2^20) | 569 | 237 | 59.2 G |
|
||||
| 64 MiB (2^24) | 183 | 76 | 19.0 G |
|
||||
| 256 MiB (2^26) | 94 | 39 | 9.8 G |
|
||||
| 512 MiB (2^27) | 69 | 29 | 7.2 G |
|
||||
| 1 GiB (2^28) | 44 | 18 | 4.6 G |
|
||||
|
||||
## Observations
|
||||
|
||||
1. The hash is memory bound at 1 GiB. The same program with identical arithmetic runs 12.8x faster when the
|
||||
dataset fits in cache (4 MiB) than at 1 GiB. Nothing but the dataset size changed.
|
||||
2. It is bound by random access, not by raw bandwidth. The GPU wrote the dataset at 427 GB/s in the same
|
||||
process, yet the hash consumed 15 to 20 GB/s of useful bytes. Each load uses 4 bytes of whatever line the
|
||||
memory system moved. The real traffic is some multiple of the useful figure; the line size was not measured.
|
||||
3. Across six of seven programs the GPU sustained 4.4 to 5.1 G random 4-byte loads per second at 1 GiB.
|
||||
The outlier (igneum-second-seed, 3.7 G/s) has the same 13 loads per iteration as igneum-genesis, so load
|
||||
count alone does not set the rate. The shape of the address dependency chain is the likely cause. Not measured.
|
||||
4. The 10 ms CPU gate passes with a wide margin: 0.015 to 0.041 ms per 32-lane warp, roughly 250x under the
|
||||
gate. Two caveats. The dataset element is a cheap closed form, so on-demand lookup costs nothing; a dataset
|
||||
that is expensive to derive (the anti-ASIC direction) would move this number. And this is one fast M5 core.
|
||||
5. Hourly regeneration is cheap: 19 to 24 ms of compile when warm, 47 to 52 ms for the first program of a
|
||||
process. That is noise against a one-hour epoch.
|
||||
6. The generator's load share came out at 13 to 18 of 64 instructions (20 to 28 percent) against a 25 percent
|
||||
weight. The op mix is printed per run.
|
||||
7. The GPU timestamps and the CPU wall clock agree, so the dispatch itself (not command overhead) is what is
|
||||
being measured. Batches of 2^22 nonces take 86 to 118 ms each.
|
||||
8. Only Apple silicon was measured. Nothing here says anything about NVIDIA or AMD, where SIMD width, cache
|
||||
line size and memory latency differ.
|
||||
|
||||
## What to try next
|
||||
|
||||
- Wider loads (uint4, 16 bytes per load) so the useful bytes approach what the memory system actually moves.
|
||||
- Two or more independent address chains per lane, to see whether the rate is latency or throughput limited.
|
||||
- A dataset element that costs real work to derive, then re-measure the CPU verify time against the 10 ms gate.
|
||||
- A program-quality filter in the generator (for example, reject programs where OR saturates a register).
|
||||
- Output distribution tests on the 64-bit results before this hash is used for leader election.
|
||||
- The same kernel on a discrete GPU, through a CUDA or Vulkan port of the generator.
|
||||
|
||||
## Files
|
||||
|
||||
- `main.swift`: generator, MSL emitter, GPU driver, CPU interpreter, CLI.
|
||||
- `igneum-bench`: the built binary (not checked in by intent; rebuild with the command above).
|
||||
1592
proto-metal/main.swift
Normal file
1592
proto-metal/main.swift
Normal file
File diff suppressed because it is too large
Load diff
41
sim/README.md
Normal file
41
sim/README.md
Normal file
|
|
@ -0,0 +1,41 @@
|
|||
# sim
|
||||
|
||||
Simulations of Igneum consensus rules. One script per rule, results next to it.
|
||||
|
||||
## finality_sim.py
|
||||
|
||||
Simulates the sustained-mining finality vote-weight rule: each key's weight is the sum over
|
||||
the trailing 30 days of its counted blocks, where a day's counted blocks are capped at
|
||||
2x the previous day's counted blocks plus a floor f. A checkpoint locks at two thirds of
|
||||
total weight. Scenarios A to F (steady state, rental burst, key splitting, honest growth
|
||||
shock, churn, patient owner) are described in `results.md` together with the model's
|
||||
assumptions and the measured numbers.
|
||||
|
||||
Requirements: Python 3, numpy (checked present: 3.10.10, numpy 2.2.6 on 3 October 2026).
|
||||
|
||||
Run everything (about one second):
|
||||
|
||||
python3 finality_sim.py > out.md
|
||||
|
||||
Options:
|
||||
|
||||
--seed N random seed, default 7 (seed 11 gives the same crossing days)
|
||||
--scenarios A,B subset of A,B,C,D,E,F
|
||||
--floors 1,10 floor values in blocks per key per day, default 1,10,100,1000
|
||||
--growth 2.0 daily cap multiplier (1e9 removes the cap)
|
||||
--window 30 trailing window in days
|
||||
--committee 100 committee size, used only for the active-24-hour total in E
|
||||
--presence 0 proposed presence gate, 0 = off (tested and rejected in results.md)
|
||||
|
||||
The runs behind `results.md`:
|
||||
|
||||
python3 finality_sim.py
|
||||
python3 finality_sim.py --growth 1e9 --scenarios B --floors 1
|
||||
python3 finality_sim.py --presence 20 --scenarios C,D
|
||||
python3 finality_sim.py --seed 11 --scenarios A,B --floors 1,1000
|
||||
|
||||
Output is markdown tables on stdout. Days in B to E are counted from the event, so "+1" is
|
||||
the first full day after the burst, the doubling or the churn.
|
||||
|
||||
Model limits are listed at the end of `results.md`: no latency, no DAG, no VRF sampling noise,
|
||||
instant difficulty retarget, free keys.
|
||||
534
sim/finality_sim.py
Normal file
534
sim/finality_sim.py
Normal file
|
|
@ -0,0 +1,534 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Igneum sustained-mining finality: vote-weight simulation.
|
||||
|
||||
Rule under test (from the design doc, finality section):
|
||||
Each miner key's vote weight is the sum over the trailing WINDOW days of
|
||||
its COUNTED blue blocks. A day's counted blocks are
|
||||
counted[d] = min(actual[d], GROWTH * counted[d-1] + FLOOR)
|
||||
so a key's counted output can at most double day over day (GROWTH = 2),
|
||||
plus a small FLOOR so a key with no history can start at all.
|
||||
A checkpoint locks when the signing keys hold two thirds of TOTAL weight.
|
||||
|
||||
Model assumptions (all of them, stated once here and repeated in results.md):
|
||||
* Time step is one day. The network produces 86,400 blocks a day (1 block/s).
|
||||
Each key's blocks that day are Poisson with mean 86,400 x its hashrate share.
|
||||
Difficulty is assumed to retarget perfectly, so total blocks stay at 86,400
|
||||
whatever the total hashrate does.
|
||||
* Every block a key produces is a blue block (no DAG, no red blocks, no latency).
|
||||
* Block rewards are unweighted: a key's reward share is its block share.
|
||||
* Honest keys sign every checkpoint they are sampled for. An attacker's keys
|
||||
sign only their own checkpoints. "Can lock alone" means that group's weight
|
||||
is at least two thirds of total weight.
|
||||
* The alternative "active total" (scenario E) counts a key as active if it was
|
||||
sampled into at least one checkpoint committee in the last 24 hours. We
|
||||
approximate the number of times a key is sampled in a day as Poisson with
|
||||
mean CHECKPOINTS_PER_DAY x COMMITTEE x (its weight share).
|
||||
* No VRF sampling noise on the lock itself, no network latency, no reorgs.
|
||||
|
||||
Standard library plus numpy. Deterministic for a given --seed.
|
||||
|
||||
Usage:
|
||||
python3 finality_sim.py # all scenarios, all floors, markdown to stdout
|
||||
python3 finality_sim.py --scenarios B,C --floors 1,100
|
||||
python3 finality_sim.py --seed 11 --growth 2 --window 30
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
|
||||
BLOCKS_PER_DAY = 86_400
|
||||
CHECKPOINTS_PER_DAY = 2_880 # one checkpoint every 30 s
|
||||
TWO_THIRDS = 2.0 / 3.0
|
||||
WARM_DAYS = 60 # honest-only days before any event in B to F
|
||||
PARETO_SHAPE = 1.0 # shape 1 gives top pool ~17%, top 10 ~50%, bottom 500 keys ~8%
|
||||
N_HONEST = 1_000
|
||||
|
||||
|
||||
class Params:
|
||||
default_presence = 0 # set from --presence in main()
|
||||
|
||||
def __init__(self, floor, growth=2.0, window=30, committee=100, presence=None):
|
||||
self.floor = float(floor)
|
||||
self.growth = float(growth)
|
||||
self.window = int(window)
|
||||
self.committee = int(committee)
|
||||
# presence gate (proposed addition, off by default): a key's daily cap stays at FLOOR
|
||||
# until it has produced at least one counted block on PRESENCE of the WINDOW days.
|
||||
self.presence = int(Params.default_presence if presence is None else presence)
|
||||
|
||||
|
||||
class Network:
|
||||
"""Keys, their hashrate, online flag, group label, and the counted-block window."""
|
||||
|
||||
def __init__(self, params, rng):
|
||||
self.p = params
|
||||
self.rng = rng
|
||||
self.hash = np.zeros(0)
|
||||
self.online = np.zeros(0, dtype=bool)
|
||||
self.group = np.zeros(0, dtype=np.int64)
|
||||
self.window = np.zeros((self.p.window, 0))
|
||||
self.prev = np.zeros(0)
|
||||
self.last_actual = np.zeros(0)
|
||||
self.ptr = 0
|
||||
self.day = 0
|
||||
|
||||
def add_keys(self, hashrates, group):
|
||||
hashrates = np.asarray(hashrates, dtype=float)
|
||||
n = hashrates.size
|
||||
self.hash = np.concatenate([self.hash, hashrates])
|
||||
self.online = np.concatenate([self.online, np.ones(n, dtype=bool)])
|
||||
self.group = np.concatenate([self.group, np.full(n, group, dtype=np.int64)])
|
||||
self.window = np.concatenate([self.window, np.zeros((self.p.window, n))], axis=1)
|
||||
self.prev = np.concatenate([self.prev, np.zeros(n)])
|
||||
self.last_actual = np.concatenate([self.last_actual, np.zeros(n)])
|
||||
|
||||
def set_hash(self, mask, value):
|
||||
self.hash[mask] = value
|
||||
|
||||
def step(self):
|
||||
h = np.where(self.online, self.hash, 0.0)
|
||||
share = h / h.sum()
|
||||
actual = self.rng.poisson(BLOCKS_PER_DAY * share).astype(float)
|
||||
cap = np.floor(self.p.growth * self.prev) + self.p.floor
|
||||
if self.p.presence > 0:
|
||||
present_days = (self.window > 0).sum(axis=0)
|
||||
cap = np.where(present_days >= self.p.presence, cap, self.p.floor)
|
||||
counted = np.minimum(actual, cap)
|
||||
self.window[self.ptr] = counted
|
||||
self.ptr = (self.ptr + 1) % self.p.window
|
||||
self.prev = counted
|
||||
self.last_actual = actual
|
||||
self.day += 1
|
||||
|
||||
def run(self, days):
|
||||
for _ in range(days):
|
||||
self.step()
|
||||
|
||||
def weight(self):
|
||||
return self.window.sum(axis=0)
|
||||
|
||||
def group_mask(self, g):
|
||||
return self.group == g
|
||||
|
||||
def group_weight_share(self, g, total_mask=None):
|
||||
w = self.weight()
|
||||
if total_mask is None:
|
||||
total = w.sum()
|
||||
else:
|
||||
total = w[total_mask].sum()
|
||||
if total <= 0:
|
||||
return 0.0
|
||||
return w[self.group_mask(g)].sum() / total
|
||||
|
||||
def group_block_share(self, g):
|
||||
a = self.last_actual
|
||||
return a[self.group_mask(g)].sum() / max(a.sum(), 1.0)
|
||||
|
||||
def active_mask(self):
|
||||
"""Keys that signed at least one checkpoint in the last 24 h (approximation, see header)."""
|
||||
w = self.weight()
|
||||
total = w.sum()
|
||||
if total <= 0:
|
||||
return np.zeros_like(self.online)
|
||||
lam = CHECKPOINTS_PER_DAY * self.p.committee * (w / total)
|
||||
sampled = self.rng.poisson(lam) >= 1
|
||||
return self.online & sampled & (w > 0)
|
||||
|
||||
|
||||
def pareto_hashrates(rng, n, total=1.0):
|
||||
h = rng.pareto(PARETO_SHAPE, n) + 1.0
|
||||
return h * (total / h.sum())
|
||||
|
||||
|
||||
def gini(x):
|
||||
x = np.sort(np.asarray(x, dtype=float))
|
||||
n = x.size
|
||||
if n == 0 or x.sum() == 0:
|
||||
return 0.0
|
||||
cum = np.cumsum(x)
|
||||
return (n + 1 - 2.0 * cum.sum() / cum[-1]) / n
|
||||
|
||||
|
||||
def fmt_pct(x):
|
||||
return "%.1f%%" % (100.0 * x)
|
||||
|
||||
|
||||
def fmt_day(d):
|
||||
return "never" if d is None else str(d)
|
||||
|
||||
|
||||
def md_table(headers, rows):
|
||||
out = ["| " + " | ".join(headers) + " |", "|" + "---|" * len(headers)]
|
||||
for r in rows:
|
||||
out.append("| " + " | ".join(str(c) for c in r) + " |")
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def first_day(series, pred):
|
||||
"""series: list of (t, value). Return first t where pred(value), else None."""
|
||||
for t, v in series:
|
||||
if pred(v):
|
||||
return t
|
||||
return None
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- scenarios
|
||||
|
||||
def warm_network(params, rng, days=WARM_DAYS):
|
||||
net = Network(params, rng)
|
||||
net.add_keys(pareto_hashrates(rng, N_HONEST), group=0)
|
||||
net.run(days)
|
||||
return net
|
||||
|
||||
|
||||
def scenario_a(floors, seed, growth, window, committee):
|
||||
"""Steady state: weight vs hashrate after 60 honest days."""
|
||||
rows = []
|
||||
ramp_rows = []
|
||||
for f in floors:
|
||||
rng = np.random.default_rng(seed)
|
||||
p = Params(f, growth, window, committee)
|
||||
net = Network(p, rng)
|
||||
net.add_keys(pareto_hashrates(rng, N_HONEST), group=0)
|
||||
full = BLOCKS_PER_DAY * window
|
||||
reached = None
|
||||
for d in range(1, WARM_DAYS + 1):
|
||||
net.step()
|
||||
if reached is None and net.weight().sum() >= 0.99 * full:
|
||||
reached = d
|
||||
w = net.weight()
|
||||
ws = w / w.sum()
|
||||
hs = net.hash / net.hash.sum()
|
||||
order = np.argsort(-hs)
|
||||
corr = np.corrcoef(hs, ws)[0, 1]
|
||||
rel = ws / hs
|
||||
under = int((rel < 0.9).sum())
|
||||
rows.append([
|
||||
int(f),
|
||||
fmt_pct(w.sum() / full),
|
||||
"%.5f" % corr,
|
||||
"%.3f / %.3f" % (gini(hs), gini(ws)),
|
||||
fmt_pct(hs[order[0]]) + " / " + fmt_pct(ws[order[0]]),
|
||||
fmt_pct(hs[order[:10]].sum()) + " / " + fmt_pct(ws[order[:10]].sum()),
|
||||
fmt_pct(hs[order[500:]].sum()) + " / " + fmt_pct(ws[order[500:]].sum()),
|
||||
"%.3f / %.3f" % (rel.min(), rel.max()),
|
||||
under,
|
||||
])
|
||||
ramp_rows.append([int(f), fmt_day(reached)])
|
||||
headers = ["floor f", "total weight / full", "corr(hash, weight)", "Gini hash / weight",
|
||||
"top-1 hash / weight", "top-10 hash / weight", "bottom-500 hash / weight",
|
||||
"min / max weight:hash ratio", "keys under 0.9x"]
|
||||
out = ["### A. Steady state, 1,000 honest keys, day 60", "", md_table(headers, rows), "",
|
||||
"Day the network's total weight first reached 99% of the full window (30 x 86,400 = 2,592,000):", "",
|
||||
md_table(["floor f", "day"], ramp_rows)]
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def run_attack(params, rng, attacker_mult, n_keys, trickle_days, report_days):
|
||||
"""Honest warm-up, optional trickle, then burst. Returns per-day series for the attacker group (1)."""
|
||||
net = Network(params, rng)
|
||||
net.add_keys(pareto_hashrates(rng, N_HONEST), group=0)
|
||||
burst_blocks = BLOCKS_PER_DAY * attacker_mult / (1.0 + attacker_mult)
|
||||
trickle_info = None
|
||||
if trickle_days > 0:
|
||||
net.run(WARM_DAYS - trickle_days)
|
||||
per_key = min(params.floor, burst_blocks / n_keys)
|
||||
target = per_key * n_keys # attacker blocks per day during trickle
|
||||
a_hash = target / (BLOCKS_PER_DAY - target) # in units of honest hashrate (= 1.0)
|
||||
net.add_keys(np.full(n_keys, a_hash / n_keys), group=1)
|
||||
net.run(trickle_days)
|
||||
trickle_info = dict(per_key=per_key, total=target, share=target / BLOCKS_PER_DAY,
|
||||
weight_share_at_burst=net.group_weight_share(1))
|
||||
else:
|
||||
net.run(WARM_DAYS)
|
||||
net.add_keys(np.zeros(n_keys), group=1)
|
||||
# burst
|
||||
net.set_hash(net.group_mask(1), attacker_mult / n_keys)
|
||||
series = []
|
||||
block_share_day1 = None
|
||||
for t in range(1, report_days + 1):
|
||||
net.step()
|
||||
if t == 1:
|
||||
block_share_day1 = net.group_block_share(1)
|
||||
series.append((t, net.group_weight_share(1)))
|
||||
return series, block_share_day1, trickle_info
|
||||
|
||||
|
||||
def scenario_b(floors, seed, growth, window, committee, mults=(1.5, 3.0), report_days=46):
|
||||
out = []
|
||||
pick = [1, 2, 3, 5, 7, 10, 14, 20, 25, 30, 35, 40, 46]
|
||||
for mult in mults:
|
||||
share = mult / (1.0 + mult)
|
||||
label = "attacker hashrate = %.1fx honest (%s of network)" % (mult, fmt_pct(share))
|
||||
if mult != 1.5:
|
||||
label += ", supplementary run, not in the brief"
|
||||
results = {}
|
||||
for f in floors:
|
||||
rng = np.random.default_rng(seed)
|
||||
p = Params(f, growth, window, committee)
|
||||
results[f] = run_attack(p, rng, mult, 1, 0, report_days)
|
||||
rows = []
|
||||
for t in pick:
|
||||
rows.append(["+%d" % t] + [fmt_pct(dict(results[f][0])[t]) for f in floors])
|
||||
headers = ["day after burst"] + ["f=%d" % f for f in floors]
|
||||
cross = [
|
||||
["crosses 1/3 (honest keys can no longer lock)"] +
|
||||
[fmt_day(first_day(results[f][0], lambda v: v > 1.0 - TWO_THIRDS)) for f in floors],
|
||||
["crosses 50%"] + [fmt_day(first_day(results[f][0], lambda v: v > 0.5)) for f in floors],
|
||||
["crosses 2/3"] + [fmt_day(first_day(results[f][0], lambda v: v >= TWO_THIRDS)) for f in floors],
|
||||
["max share in %d days" % report_days] + [fmt_pct(max(v for _, v in results[f][0])) for f in floors],
|
||||
["block-reward share, burst day 1"] + [fmt_pct(results[f][1]) for f in floors],
|
||||
]
|
||||
out.append("### B. Rental burst at day 60, one key, %s" % label)
|
||||
out.append("")
|
||||
out.append("Attacker weight share of total:")
|
||||
out.append("")
|
||||
out.append(md_table(headers, rows))
|
||||
out.append("")
|
||||
out.append(md_table(["event"] + ["f=%d" % f for f in floors], cross))
|
||||
out.append("")
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def scenario_c(floors, seed, growth, window, committee, ks=(100, 1_000, 10_000), mults=(1.5, 3.0),
|
||||
report_days=46):
|
||||
out = []
|
||||
for mult in mults:
|
||||
share = mult / (1.0 + mult)
|
||||
label = "attacker hashrate = %.1fx honest (%s of network)" % (mult, fmt_pct(share))
|
||||
if mult != 1.5:
|
||||
label += ", supplementary run"
|
||||
# baseline B for comparison
|
||||
base = {}
|
||||
for f in floors:
|
||||
rng = np.random.default_rng(seed)
|
||||
base[f] = run_attack(Params(f, growth, window, committee), rng, mult, 1, 0, report_days)[0]
|
||||
rows_13, rows_50, rows_23, rows_cost, rows_d1 = [], [], [], [], []
|
||||
one_third = 1.0 - TWO_THIRDS
|
||||
|
||||
def add_rows(name, series_by_f):
|
||||
rows_13.append([name] + [fmt_day(first_day(series_by_f[f], lambda v: v > one_third)) for f in floors])
|
||||
rows_50.append([name] + [fmt_day(first_day(series_by_f[f], lambda v: v > 0.5)) for f in floors])
|
||||
rows_23.append([name] + [fmt_day(first_day(series_by_f[f], lambda v: v >= TWO_THIRDS)) for f in floors])
|
||||
rows_d1.append([name] + [fmt_pct(dict(series_by_f[f])[1]) for f in floors])
|
||||
|
||||
add_rows("B: 1 key, no trickle", base)
|
||||
# C0: K fresh keys appear on the burst day with no trickle (cheapest split, costs only key creation)
|
||||
for k in list(ks) + [50_000]:
|
||||
res = {}
|
||||
for f in floors:
|
||||
rng = np.random.default_rng(seed)
|
||||
res[f] = run_attack(Params(f, growth, window, committee), rng, mult, k, 0, report_days)[0]
|
||||
add_rows("C0: K=%d, no trickle" % k, res)
|
||||
# C: K keys trickle at the floor for 30 days, then flood (the brief's scenario)
|
||||
for k in ks:
|
||||
res = {}
|
||||
for f in floors:
|
||||
rng = np.random.default_rng(seed)
|
||||
res[f] = run_attack(Params(f, growth, window, committee), rng, mult, k, 30, report_days)
|
||||
add_rows("C: K=%d, 30-day trickle" % k, {f: res[f][0] for f in floors})
|
||||
rows_cost.append(["K=%d" % k] + [
|
||||
"%.4g blk/key/day, %s of network, %s weight at burst" % (
|
||||
res[f][2]["per_key"], fmt_pct(res[f][2]["share"]), fmt_pct(res[f][2]["weight_share_at_burst"]))
|
||||
for f in floors])
|
||||
headers = ["run"] + ["f=%d" % f for f in floors]
|
||||
out.append("### C. Key splitting, %s" % label)
|
||||
out.append("")
|
||||
out.append("Attacker weight share one day after the burst:")
|
||||
out.append("")
|
||||
out.append(md_table(headers, rows_d1))
|
||||
out.append("")
|
||||
out.append("Days after the burst until the attacker crosses 1/3 of total weight (honest keys can no longer lock):")
|
||||
out.append("")
|
||||
out.append(md_table(headers, rows_13))
|
||||
out.append("")
|
||||
out.append("Days after the burst until the attacker crosses 50% of total weight:")
|
||||
out.append("")
|
||||
out.append(md_table(headers, rows_50))
|
||||
out.append("")
|
||||
out.append("Days after the burst until the attacker reaches 2/3 of total weight:")
|
||||
out.append("")
|
||||
out.append(md_table(headers, rows_23))
|
||||
out.append("")
|
||||
out.append("Trickle phase (days 30 to 60): per-key rate, attacker share of network blocks, "
|
||||
"attacker weight share on the burst day:")
|
||||
out.append("")
|
||||
out.append(md_table(["run"] + ["f=%d" % f for f in floors], rows_cost))
|
||||
out.append("")
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def scenario_d(floors, seed, growth, window, committee, report_days=46):
|
||||
pick = [1, 3, 5, 10, 15, 20, 25, 30, 35, 46]
|
||||
results = {}
|
||||
for f in floors:
|
||||
rng = np.random.default_rng(seed)
|
||||
p = Params(f, growth, window, committee)
|
||||
net = warm_network(p, rng)
|
||||
net.add_keys(pareto_hashrates(rng, N_HONEST, total=1.0), group=2) # new honest cohort, same total hashrate
|
||||
series_new, series_old = [], []
|
||||
for t in range(1, report_days + 1):
|
||||
net.step()
|
||||
series_new.append((t, net.group_weight_share(2)))
|
||||
series_old.append((t, net.group_weight_share(0)))
|
||||
results[f] = (series_new, series_old)
|
||||
rows = []
|
||||
for t in pick:
|
||||
rows.append(["+%d" % t] + [fmt_pct(dict(results[f][0])[t]) for f in floors])
|
||||
headers = ["day after doubling"] + ["f=%d" % f for f in floors]
|
||||
ev = [
|
||||
["new cohort reaches 45% weight (0.9x its 50% hash share)"] +
|
||||
[fmt_day(first_day(results[f][0], lambda v: v >= 0.45)) for f in floors],
|
||||
["new cohort reaches 49% weight"] +
|
||||
[fmt_day(first_day(results[f][0], lambda v: v >= 0.49)) for f in floors],
|
||||
["last day old cohort holds 2/3 (can lock alone)"] +
|
||||
[fmt_day(max([t for t, v in results[f][1] if v >= TWO_THIRDS] or [0])) for f in floors],
|
||||
]
|
||||
out = ["### D. Honest growth shock: network doubles at day 60 (1,000 new keys, same total hashrate as the old 1,000)",
|
||||
"", "New cohort's weight share of total (its hashrate share is 50% from day +1):", "",
|
||||
md_table(headers, rows), "", md_table(["event"] + ["f=%d" % f for f in floors], ev)]
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def pick_offline(rng, weights, target_frac):
|
||||
"""Random subset of keys whose weight sums to about target_frac of total (never over by more than one small key)."""
|
||||
total = weights.sum()
|
||||
target = target_frac * total
|
||||
order = rng.permutation(weights.size)
|
||||
mask = np.zeros(weights.size, dtype=bool)
|
||||
acc = 0.0
|
||||
for i in order:
|
||||
if acc + weights[i] <= target:
|
||||
mask[i] = True
|
||||
acc += weights[i]
|
||||
return mask, acc / total
|
||||
|
||||
|
||||
def scenario_e(floors, seed, growth, window, committee, fracs=(0.30, 0.35, 0.50), report_days=35):
|
||||
pick = [1, 2, 3, 5, 7, 10, 15, 20, 25, 30, 31]
|
||||
out = []
|
||||
for frac in fracs:
|
||||
label = "%d%% of weight goes offline permanently at day 60" % round(100 * frac)
|
||||
if frac != 0.30:
|
||||
label += " (supplementary run)"
|
||||
results = {}
|
||||
for f in floors:
|
||||
rng = np.random.default_rng(seed)
|
||||
p = Params(f, growth, window, committee)
|
||||
net = warm_network(p, rng)
|
||||
w = net.weight()
|
||||
off, actual_frac = pick_offline(rng, w, frac)
|
||||
net.online[off] = False
|
||||
net.group[off] = 9 # dead keys
|
||||
s_all, s_active = [], []
|
||||
for t in range(1, report_days + 1):
|
||||
net.step()
|
||||
s_all.append((t, net.group_weight_share(0)))
|
||||
s_active.append((t, net.group_weight_share(0, total_mask=net.active_mask())))
|
||||
results[f] = (s_all, s_active, actual_frac, int(off.sum()))
|
||||
rows = []
|
||||
for t in pick:
|
||||
rows.append(["+%d" % t] + [fmt_pct(dict(results[f][0])[t]) + " / " + fmt_pct(dict(results[f][1])[t])
|
||||
for f in floors])
|
||||
headers = ["day after churn"] + ["f=%d (all keys / active keys)" % f for f in floors]
|
||||
ev = [
|
||||
["offline fraction actually picked (keys)"] +
|
||||
["%s (%d keys)" % (fmt_pct(results[f][2]), results[f][3]) for f in floors],
|
||||
["live share >= 2/3, total = all keys, first day"] +
|
||||
[fmt_day(first_day(results[f][0], lambda v: v >= TWO_THIRDS)) for f in floors],
|
||||
["live share >= 2/3, total = active keys, first day"] +
|
||||
[fmt_day(first_day(results[f][1], lambda v: v >= TWO_THIRDS)) for f in floors],
|
||||
["dead weight fully out of the window (day)"] + [str(window) for _ in floors],
|
||||
["live share, all-keys total, day +%d" % window] +
|
||||
[fmt_pct(dict(results[f][0])[window]) for f in floors],
|
||||
]
|
||||
out.append("### E. Churn: %s" % label)
|
||||
out.append("")
|
||||
out.append("Live honest weight share. Left of the slash: total = every key ever seen. "
|
||||
"Right: total = keys that signed a checkpoint in the last 24 h (committee %d)." % committee)
|
||||
out.append("")
|
||||
out.append(md_table(headers, rows))
|
||||
out.append("")
|
||||
out.append(md_table(["event"] + ["f=%d" % f for f in floors], ev))
|
||||
out.append("")
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def scenario_f(floors, seed, growth, window, committee, owner_shares=(0.51, 0.67), days=90):
|
||||
out = []
|
||||
pick = [15, 30, 45, 60, 90]
|
||||
for s in owner_shares:
|
||||
label = "owner holds %s of hashrate from day 0 and mines honestly" % fmt_pct(s)
|
||||
if s != 0.51:
|
||||
label += " (supplementary run)"
|
||||
results = {}
|
||||
for f in floors:
|
||||
rng = np.random.default_rng(seed)
|
||||
p = Params(f, growth, window, committee)
|
||||
net = Network(p, rng)
|
||||
net.add_keys(pareto_hashrates(rng, N_HONEST, total=1.0 - s), group=0)
|
||||
net.add_keys(np.array([s]), group=1)
|
||||
series = []
|
||||
for t in range(1, days + 1):
|
||||
net.step()
|
||||
series.append((t, net.group_weight_share(1)))
|
||||
results[f] = series
|
||||
rows = []
|
||||
for t in pick:
|
||||
rows.append([str(t)] + [fmt_pct(dict(results[f])[t]) for f in floors])
|
||||
tail = lambda f: [v for t, v in results[f] if t > 30]
|
||||
ev = [
|
||||
["mean weight share, days 31 to %d" % days] + [fmt_pct(np.mean(tail(f))) for f in floors],
|
||||
["min / max weight share, days 31 to %d" % days] +
|
||||
["%s / %s" % (fmt_pct(min(tail(f))), fmt_pct(max(tail(f)))) for f in floors],
|
||||
["days at or above 2/3 (can lock alone), of %d" % days] +
|
||||
[str(sum(1 for _, v in results[f] if v >= TWO_THIRDS)) for f in floors],
|
||||
["days above 1/3 (can veto every lock), of %d" % days] +
|
||||
[str(sum(1 for _, v in results[f] if v > 1.0 / 3.0)) for f in floors],
|
||||
]
|
||||
out.append("### F. Patient owner: %s" % label)
|
||||
out.append("")
|
||||
out.append("Owner's weight share of total by day:")
|
||||
out.append("")
|
||||
out.append(md_table(["day"] + ["f=%d" % f for f in floors], rows))
|
||||
out.append("")
|
||||
out.append(md_table(["event"] + ["f=%d" % f for f in floors], ev))
|
||||
out.append("")
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
SCENARIOS = {"A": scenario_a, "B": scenario_b, "C": scenario_c, "D": scenario_d, "E": scenario_e, "F": scenario_f}
|
||||
|
||||
|
||||
def main(argv=None):
|
||||
ap = argparse.ArgumentParser(description="Igneum sustained-mining finality vote-weight simulation")
|
||||
ap.add_argument("--seed", type=int, default=7)
|
||||
ap.add_argument("--scenarios", default="A,B,C,D,E,F")
|
||||
ap.add_argument("--floors", default="1,10,100,1000", help="comma list of floor f values (blocks/day)")
|
||||
ap.add_argument("--growth", type=float, default=2.0, help="daily growth cap multiplier (2 = each day at most 2x the day before)")
|
||||
ap.add_argument("--window", type=int, default=30, help="trailing window in days")
|
||||
ap.add_argument("--committee", type=int, default=100, help="checkpoint committee size, used only for the active-total definition in E")
|
||||
ap.add_argument("--presence", type=int, default=0,
|
||||
help="proposed presence gate: a key's daily cap stays at the floor until it has counted blocks on this many of the window days (0 = off, the rule as written)")
|
||||
args = ap.parse_args(argv)
|
||||
Params.default_presence = args.presence
|
||||
floors = [int(x) for x in args.floors.split(",") if x]
|
||||
print("# finality_sim output")
|
||||
print()
|
||||
print("seed %d, growth cap %.1fx, window %d days, committee %d, presence gate %d, floors %s, blocks/day %d, honest keys %d, Pareto shape %.1f"
|
||||
% (args.seed, args.growth, args.window, args.committee, args.presence, floors, BLOCKS_PER_DAY, N_HONEST, PARETO_SHAPE))
|
||||
print()
|
||||
for s in args.scenarios.split(","):
|
||||
s = s.strip().upper()
|
||||
if s not in SCENARIOS:
|
||||
print("unknown scenario %s" % s, file=sys.stderr)
|
||||
return 2
|
||||
print(SCENARIOS[s](floors, args.seed, args.growth, args.window, args.committee))
|
||||
print()
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
234
sim/results.md
Normal file
234
sim/results.md
Normal file
|
|
@ -0,0 +1,234 @@
|
|||
# Sustained-mining finality: vote-weight simulation results
|
||||
|
||||
Date: 3 October 2026. Simulator: `sim/finality_sim.py`, seed 7 (seed 11 reproduces every crossing day quoted here).
|
||||
Every number below was produced by the simulator. Nothing is from memory.
|
||||
|
||||
## The rule as simulated
|
||||
|
||||
Each key's vote weight is the sum over the trailing 30 days of its counted blocks, where
|
||||
|
||||
counted[d] = min(actual[d], 2 x counted[d-1] + f)
|
||||
|
||||
and f is the floor in blocks per day. A checkpoint locks when the signing keys hold two thirds of total weight.
|
||||
|
||||
## Model assumptions (apply to every table)
|
||||
|
||||
| Assumption | Value |
|
||||
|---|---|
|
||||
| Time step | 1 day |
|
||||
| Blocks per day | 86,400 (1 block/s), Poisson per key with mean 86,400 x hashrate share |
|
||||
| Difficulty | retargets perfectly, total blocks stay at 86,400/day whatever the hashrate does |
|
||||
| Blue blocks | every block produced is blue (no DAG, no red blocks, no latency) |
|
||||
| Honest network | 1,000 keys, Pareto shape 1.0 hashrate: top key 17.1%, top 10 keys 49.8%, bottom 500 keys 8.5% |
|
||||
| Warm-up | 60 honest days before any event in B to F |
|
||||
| Rewards | unweighted, a key's reward share is its block share |
|
||||
| "Can lock alone" | a group's weight is at least 2/3 of total weight (no committee sampling noise on the lock) |
|
||||
| "Active total" (E only) | a key counts if sampled into at least one committee in 24 h, approximated as Poisson(2,880 checkpoints x committee 100 x weight share) >= 1 |
|
||||
| Floors tested | f in {1, 10, 100, 1000} blocks/day |
|
||||
|
||||
Days in the B to E tables are counted from the event (burst, doubling or churn), so "+1" is the first full day after it.
|
||||
|
||||
## A. Steady state
|
||||
|
||||
1,000 honest keys, 60 days from zero history.
|
||||
|
||||
| floor f | total weight / full | corr(hash, weight) | Gini hash / weight | top-1 hash / weight | top-10 hash / weight | bottom-500 hash / weight | min / max weight:hash ratio | keys under 0.9x |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| 1 | 100.0% | 1.00000 | 0.761 / 0.762 | 17.1% / 17.1% | 49.8% / 49.9% | 8.5% / 8.4% | 0.822 / 1.170 | 21 |
|
||||
| 10 | 100.0% | 1.00000 | 0.761 / 0.761 | 17.1% / 17.1% | 49.8% / 49.9% | 8.5% / 8.5% | 0.833 / 1.169 | 12 |
|
||||
| 100 | 100.0% | 1.00000 | 0.761 / 0.761 | 17.1% / 17.1% | 49.8% / 49.9% | 8.5% / 8.5% | 0.833 / 1.169 | 12 |
|
||||
| 1000 | 100.0% | 1.00000 | 0.761 / 0.761 | 17.1% / 17.1% | 49.8% / 49.9% | 8.5% / 8.5% | 0.833 / 1.169 | 12 |
|
||||
|
||||
Day the network's total weight first reached 99% of the full window (30 x 86,400 = 2,592,000 counted blocks):
|
||||
|
||||
| floor f | day |
|
||||
|---|---|
|
||||
| 1 | 41 |
|
||||
| 10 | 38 |
|
||||
| 100 | 35 |
|
||||
| 1000 | 32 |
|
||||
|
||||
Interpretation. At steady state weight is proportional to hashrate: correlation 1.00000, Gini and top-k shares identical to three figures. The only deviation is Poisson noise on the smallest keys (10 to 20 blocks a day), where a quiet day lowers the next day's cap. At f=1 this leaves 21 of 1,000 keys below 0.9x their hashrate share, at f>=10 it is 12. From zero history the biggest pool (14,774 blocks a day) needs 14 doublings at f=1, so the network holds full weight only from day 41 (day 32 at f=1000). That sits inside the 30-day launch ramp plus 11 days.
|
||||
|
||||
## B. Rental burst
|
||||
|
||||
At day 60 one new key appears with 1.5x the honest hashrate (60% of the network).
|
||||
|
||||
Attacker weight share of total:
|
||||
|
||||
| day after burst | f=1 | f=10 | f=100 | f=1000 |
|
||||
|---|---|---|---|---|
|
||||
| +1 | 0.0% | 0.0% | 0.0% | 0.0% |
|
||||
| +2 | 0.0% | 0.0% | 0.0% | 0.2% |
|
||||
| +3 | 0.0% | 0.0% | 0.0% | 0.4% |
|
||||
| +5 | 0.0% | 0.0% | 0.2% | 2.4% |
|
||||
| +7 | 0.0% | 0.1% | 1.1% | 6.7% |
|
||||
| +10 | 0.1% | 1.0% | 6.9% | 13.2% |
|
||||
| +14 | 1.7% | 9.0% | 16.2% | 21.9% |
|
||||
| +20 | 17.3% | 24.2% | 30.1% | 34.9% |
|
||||
| +25 | 31.1% | 36.8% | 41.8% | 45.7% |
|
||||
| +30 | 44.9% | 49.5% | 53.4% | 56.6% |
|
||||
| +35 | 51.6% | 55.1% | 58.2% | 60.0% |
|
||||
| +40 | 56.8% | 59.3% | 60.0% | 60.0% |
|
||||
| +46 | 60.1% | 60.0% | 60.0% | 60.0% |
|
||||
|
||||
| event | f=1 | f=10 | f=100 | f=1000 | no cap (f=1, cap removed) |
|
||||
|---|---|---|---|---|---|
|
||||
| crosses 1/3 (honest keys can no longer lock) | 26 | 24 | 22 | 20 | 18 |
|
||||
| crosses 50% | 34 | 31 | 29 | 27 | 26 |
|
||||
| crosses 2/3 | never | never | never | never | never |
|
||||
| max share in 46 days | 60.1% | 60.0% | 60.0% | 60.0% | 60.0% |
|
||||
| block-reward share, burst day 1 | 59.9% | 59.9% | 59.9% | 59.9% | 59.9% |
|
||||
|
||||
Supplementary run, not in the brief: attacker with 3x the honest hashrate (75% of the network), so that a 2/3 crossing exists.
|
||||
|
||||
| event | f=1 | f=10 | f=100 | f=1000 | no cap |
|
||||
|---|---|---|---|---|---|
|
||||
| crosses 1/3 | 23 | 21 | 19 | 17 | 14 |
|
||||
| crosses 50% | 27 | 26 | 24 | 23 | 21 |
|
||||
| crosses 2/3 | 34 | 31 | 30 | 29 | 27 |
|
||||
| block-reward share, burst day 1 | 75.0% | 75.0% | 75.0% | 75.0% | 75.0% |
|
||||
|
||||
Interpretation. Rewards are immediate: the renter earns 59.9% of blocks on day one. Vote weight is not: it stays under 2% for the first two weeks at f<=10. A 60% renter never reaches 2/3 because at steady state weight equals hashrate share, so its ceiling is 60%. It does cross 1/3 on day 20 to 26, from which point honest keys cannot lock either. A 60% renter is therefore a liveness attack (finality stalls after about three weeks), not a safety attack. A 75% renter locks alone from day 29 to 34. The "no cap" column shows how the defence splits: the 30-day window alone holds the 75% renter to day 27, the 2x cap adds 2 days (f=1000) to 7 days (f=1) on top.
|
||||
|
||||
## C. Key splitting
|
||||
|
||||
Same attacker, split over K keys. Two variants. C0: K fresh keys appear on the burst day with no history (costs only key creation). C: the brief's scenario, K keys trickle-mine at the floor rate for 30 days before the burst. B is the single-key baseline from above.
|
||||
|
||||
Attacker 1.5x honest (60%). Days after the burst until the attacker crosses 50% of total weight:
|
||||
|
||||
| run | f=1 | f=10 | f=100 | f=1000 |
|
||||
|---|---|---|---|---|
|
||||
| B: 1 key, no trickle | 34 | 31 | 29 | 27 |
|
||||
| C0: K=100, no trickle | 29 | 27 | 26 | 25 |
|
||||
| C0: K=1,000, no trickle | 27 | 26 | 26 | 26 |
|
||||
| C0: K=10,000, no trickle | 27 | 26 | 26 | 26 |
|
||||
| C0: K=50,000, no trickle | 28 | 26 | 26 | 26 |
|
||||
| C: K=100, 30-day trickle | 29 | 27 | 25 | 1 |
|
||||
| C: K=1,000, 30-day trickle | 27 | 25 | 1 | 1 |
|
||||
| C: K=10,000, 30-day trickle | 25 | 1 | 1 | 1 |
|
||||
| no cap, 1 key | 26 | 26 | 26 | 26 |
|
||||
|
||||
Days until the attacker crosses 1/3 (honest keys lose the lock):
|
||||
|
||||
| run | f=1 | f=10 | f=100 | f=1000 |
|
||||
|---|---|---|---|---|
|
||||
| B: 1 key | 26 | 24 | 22 | 20 |
|
||||
| C0: K=10,000 | 19 | 17 | 17 | 17 |
|
||||
| C: K=10,000, trickle | 15 | 1 | 1 | 1 |
|
||||
|
||||
2/3 is never reached by a 60% attacker in any run. Supplementary 75% attacker, days until 2/3:
|
||||
|
||||
| run | f=1 | f=10 | f=100 | f=1000 |
|
||||
|---|---|---|---|---|
|
||||
| B: 1 key, no trickle | 34 | 31 | 30 | 29 |
|
||||
| C0: K=100, no trickle | 30 | 29 | 28 | 27 |
|
||||
| C0: K=1,000, no trickle | 29 | 28 | 27 | 27 |
|
||||
| C0: K=10,000, no trickle | 28 | 27 | 27 | 27 |
|
||||
| C0: K=50,000, no trickle | 29 | 27 | 27 | 27 |
|
||||
| C: K=100, 30-day trickle | 29 | 28 | 27 | 1 |
|
||||
| C: K=1,000, 30-day trickle | 28 | 27 | 1 | 1 |
|
||||
| C: K=10,000, 30-day trickle | 27 | 1 | 1 | 1 |
|
||||
| no cap, 1 key | 27 | 27 | 27 | 27 |
|
||||
|
||||
What the trickle costs the attacker (days 30 to 60, 60% attacker):
|
||||
|
||||
| run | f=1 | f=10 | f=100 | f=1000 |
|
||||
|---|---|---|---|---|
|
||||
| K=100 | 1 blk/key/day, 0.1% of network, 0.1% weight at burst | 10 blk/key/day, 1.2% of network, 1.2% weight | 100 blk/key/day, 11.6% of network, 11.5% weight | 518 blk/key/day, 60.0% of network, 60.0% weight |
|
||||
| K=1,000 | 1 blk/key/day, 1.2% of network, 1.0% weight | 10 blk/key/day, 11.6% of network, 11.5% weight | 51.8 blk/key/day, 60.0% of network, 60.0% weight | same, trickle = full attack |
|
||||
| K=10,000 | 1 blk/key/day, 11.6% of network, 10.0% weight | 5.2 blk/key/day, 60.0% of network, 60.0% weight | same, trickle = full attack | same, trickle = full attack |
|
||||
|
||||
Interpretation. The per-key cap is defeated by splitting. The cap's total capacity for fresh keys is K x f counted blocks on day one, then 3Kf, 7Kf and so on. With 10,000 fresh keys at f=1 the ramp lasts 3 days instead of 16 and the 50% crossing moves from day 34 to day 27, one day off the no-cap figure of 26. At f=10, 10,000 keys match the no-cap figure exactly; at f=100, 1,000 keys do; at f=1000, 100 keys do. The 30-day trickle in the brief adds little on top of that: at K=10,000 and f=1 it costs 11.6% of the network for 30 days and buys 2 more days (25 against 27). Where the trickle table reads "60.0% of network", K x f already exceeds the attacker's daily output, so the "trickle" is the full attack started 30 days early and the day-1 figure of 60% is paid for in full, not stolen.
|
||||
|
||||
The important reading is the floor row. Against a splitter, time to 50% for a 60% attacker and time to 2/3 for a 75% attacker converge on 26 and 27 days whatever f is. Those days come from the 30-day window, not from the cap. The cap buys 7 to 8 days against an attacker who will not split and 1 to 2 days against one who will.
|
||||
|
||||
Supplementary: a presence gate (run with `--presence 20`, a key's daily cap stays at f until it has counted blocks on 20 of the last 30 days) was tested as a fix. It pushes C0 at K=10,000, f=1 from day 27 to day 40 and pushes the single-key renter past day 46, but the trickle variant is unchanged (the attacker simply pays the gate), and it damages honest growth: in D the new honest cohort never reaches 45% weight within 46 days at f<=10 and the old cohort can lock alone for 39 days instead of 23. Rejected on these numbers.
|
||||
|
||||
## D. Honest growth shock
|
||||
|
||||
At day 60 the honest network doubles: 1,000 new keys with a fresh Pareto draw and the same total hashrate as the old 1,000. The new cohort's hashrate share is 50% from day +1.
|
||||
|
||||
New cohort's weight share of total:
|
||||
|
||||
| day after doubling | f=1 | f=10 | f=100 | f=1000 |
|
||||
|---|---|---|---|---|
|
||||
| +1 | 0.0% | 0.3% | 0.9% | 1.4% |
|
||||
| +3 | 0.4% | 1.7% | 3.4% | 4.8% |
|
||||
| +5 | 1.5% | 3.9% | 6.7% | 8.1% |
|
||||
| +10 | 7.6% | 12.2% | 15.2% | 16.5% |
|
||||
| +15 | 16.7% | 20.9% | 23.6% | 24.8% |
|
||||
| +20 | 26.0% | 29.7% | 32.1% | 33.2% |
|
||||
| +25 | 35.2% | 38.5% | 40.6% | 41.5% |
|
||||
| +30 | 44.5% | 47.3% | 49.1% | 49.9% |
|
||||
| +35 | 48.5% | 49.7% | 50.0% | 50.0% |
|
||||
| +46 | 50.0% | 50.0% | 50.0% | 50.0% |
|
||||
|
||||
| event | f=1 | f=10 | f=100 | f=1000 |
|
||||
|---|---|---|---|---|
|
||||
| new cohort reaches 45% weight (0.9x its hash share) | 31 | 29 | 28 | 28 |
|
||||
| new cohort reaches 49% weight | 37 | 33 | 30 | 30 |
|
||||
| last day old cohort holds 2/3 (can lock alone) | 23 | 22 | 20 | 20 |
|
||||
|
||||
Interpretation. New honest miners are under-weighted for the whole window: 28 to 31 days to reach 0.9x their hashrate share, 30 to 37 days to reach 0.98x. The window sets the floor on that (with the cap removed the arithmetic gives old share = 1 - t/60, so the old cohort holds 2/3 for exactly 20 days). The cap adds 0 days at f>=100 and 3 days at f=1. During those 20 to 23 days the old miners can lock checkpoints without any new miner's signature. That is the intended behaviour (a doubling overnight is indistinguishable from a rental burst) and the price is paid by honest newcomers for about three weeks.
|
||||
|
||||
## E. Churn
|
||||
|
||||
At day 60 a random set of keys holding 30% of weight goes offline for good. The remaining keys produce all 86,400 blocks from then on (perfect retarget). Supplementary runs at 35% and 50%.
|
||||
|
||||
Live honest weight share, total = every key with weight in the window ("all keys") against total = keys that signed a checkpoint in the last 24 h ("active"). Values were identical across f to one decimal, so one column per churn level is shown.
|
||||
|
||||
| day after churn | 30% offline, all / active | 35% offline, all / active | 50% offline, all / active |
|
||||
|---|---|---|---|
|
||||
| +1 | 71.0% / 100.0% | 66.2% / 100.0% | 51.6% / 100.0% |
|
||||
| +2 | 72.0% / 100.0% | 67.3% / 100.0% | 53.3% / 100.0% |
|
||||
| +3 | 73.0% / 100.0% | 68.5% / 100.0% | 55.0% / 100.0% |
|
||||
| +5 | 75.0% / 100.0% | 70.8% / 100.0% | 58.3% / 100.0% |
|
||||
| +10 | 80.0% / 100.0% | 76.7% / 100.0% | 66.6% / 100.0% |
|
||||
| +15 | 85.0% / 100.0% | 82.5% / 100.0% | 75.0% / 100.0% |
|
||||
| +20 | 90.0% / 100.0% | 88.3% / 100.0% | 83.3% / 100.0% |
|
||||
| +30 | 100.0% / 100.0% | 100.0% / 100.0% | 100.0% / 100.0% |
|
||||
|
||||
| event | 30% offline | 35% offline | 50% offline |
|
||||
|---|---|---|---|
|
||||
| offline fraction actually picked (keys) | 30.0% (435 keys) | 35.0% (529 to 530 keys) | 50.0% (473 keys) |
|
||||
| live share >= 2/3, all-keys total, first day | 1 (never lost) | 2 | 10 to 11 |
|
||||
| live share >= 2/3, active total, first day | 1 | 1 | 1 |
|
||||
| dead weight fully out of the window | day 30 | day 30 | day 30 |
|
||||
|
||||
Interpretation. The brief expected 30% churn to block the lock. It does not: 30% offline leaves 70% to 71% live, above 2/3 from day one, with a 3 to 4 point margin. The threshold is 1/3 offline. At 35% the lock is lost for one day, at 50% for 10 to 11 days. The live share follows 1 - x(30 - t)/30 for offline fraction x and day t, because the survivors inherit the whole block supply and the dead weight decays linearly out of the window, which it leaves completely on day 30. Under the active-24-hour total there is no stall at any churn level: the dead keys drop out of the denominator after one day. The difference is a liveness gain bought with a safety cost the simulation does not model: under the active total, any event that keeps honest keys from signing (eclipse, partition, targeted DoS) shrinks the denominator and lets a smaller faction lock; in a partition both sides see 100% active and both can lock. The all-keys total fails safe (stall) in the same cases.
|
||||
|
||||
## F. Patient owner
|
||||
|
||||
An attacker who owns 51% of the hardware, mines honestly from day 0 and so has full history. Supplementary run at 67%.
|
||||
|
||||
Owner's weight share by day:
|
||||
|
||||
| day | 51% owner, f=1 | f=10 | f=100 | f=1000 | 67% owner, f=1 | f=10 | f=100 | f=1000 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| 15 | 15.6% | 31.0% | 38.9% | 44.6% | 20.7% | 43.4% | 53.6% | 60.2% |
|
||||
| 30 | 42.4% | 44.0% | 45.9% | 48.0% | 58.0% | 59.6% | 61.7% | 64.0% |
|
||||
| 45 | 51.0% | 50.9% | 50.9% | 50.9% | 67.1% | 67.0% | 67.0% | 67.0% |
|
||||
| 60 | 51.0% | 51.0% | 51.0% | 51.0% | 67.1% | 67.0% | 67.0% | 67.0% |
|
||||
| 90 | 51.1% | 51.0% | 51.0% | 51.0% | 67.1% | 67.0% | 67.0% | 67.0% |
|
||||
|
||||
| event (of 90 days) | 51% owner, f=1 | f=10 | f=100 | f=1000 | 67% owner, f=1 | f=10 | f=100 | f=1000 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| mean weight share, days 31 to 90 | 50.0% | 50.4% | 50.7% | 50.9% | 66.0% | 66.4% | 66.7% | 66.9% |
|
||||
| min / max, days 31 to 90 | 42.8% / 51.1% | 44.5% / 51.1% | 46.5% / 51.1% | 48.7% / 51.1% | 58.5% / 67.2% | 60.2% / 67.0% | 62.4% / 67.0% | 64.7% / 67.0% |
|
||||
| days at or above 2/3 (locks alone) | 0 | 0 | 0 | 0 | 47 | 50 | 53 | 57 |
|
||||
| days above 1/3 (vetoes every lock) | 71 | 74 | 79 | 84 | 74 | 78 | 81 | 85 |
|
||||
|
||||
Interpretation. The residual the design accepts is exactly the hashrate share. A 51% owner holds 51.0% of weight from day 45 onward (51.0% to 51.1% across every f). It can never lock alone, and it can veto every lock from day 7 to 20 onward, for 71 to 84 of the 90 days. A 67% owner locks alone from day 34 to 44 onward. The 2x cap and the floor only decide how fast the owner reaches its share (day 15 reading: 15.6% at f=1 against 44.6% at f=1000); they change nothing about where it lands. Sustained-mining finality therefore defends against rented hashrate, and against owned hashrate it offers the same guarantee as the underlying proof of work: safety needs honest hashrate above 1/3 of the sustained total, liveness needs it above 2/3.
|
||||
|
||||
## Recommended parameters
|
||||
|
||||
| Parameter | Recommendation | Reason from the runs |
|
||||
|---|---|---|
|
||||
| Floor f | 1 block per key per day | Every larger f voids the cap for a smaller key count (f=10: 10,000 keys, f=100: 1,000, f=1000: 100). Honest cost of f=1 is 21 of 1,000 small keys under 0.9x weight, against 12 at f=10 (A). |
|
||||
| Daily cap | 2x the previous day's counted blocks, keep | Cheap for honest growth (adds 0 to 3 days to D) and worth 7 to 8 days against a single-key renter at f=1. Do not count on it against a splitter: it is worth 1 to 2 days there. |
|
||||
| Window | 30 days, keep | This is the actual defence. It holds a 60% renter below 1/3 for 18 to 26 days and below 50% for 26 to 34 days, and a 75% renter below 2/3 for 27 to 34 days, splitting or not. The model's arithmetic says these times scale linearly with the window (a 60-day window doubles them) and that the honest under-weighting in D scales with it too. |
|
||||
| Total weight | every key with non-zero weight in the window (all keys), not the active-24-hour set | Fails safe. The stall it causes needs more than 1/3 of weight to vanish at once (35% offline: 1 day, 50%: 10 to 11 days) and clears within the window. The active-24-hour total removes the stall but lets a partition or an eclipse shrink the denominator, which the simulation cannot price. If liveness under mass churn matters, test a middle ground (drop keys silent for 7 days) before adopting it. |
|
||||
| Presence gate | do not adopt | Tested at 20 of 30 days: slows the no-trickle splitter by 13 days, changes nothing for the trickling splitter, and roughly doubles the time new honest miners wait for weight (D). |
|
||||
|
||||
What the simulation cannot tell us. It has no network latency, so every block is blue and on time; in the real GHOSTDAG some of an attacker's or a laggard's blocks will be red and not counted, which changes weights in a direction this model cannot predict. It has no DAG, so it cannot see how a 60% attacker's blocks interact with the honest tip or how the Horizen secret-mining penalty bears on withheld blocks. It has no VRF sampling noise: the lock is tested against exact group weights, while the real checkpoint samples a committee, so a faction near 2/3 will sometimes lock and sometimes not, and a faction near 1/3 will sometimes be unable to veto. It assumes difficulty retargets instantly and hashrate is constant within a day. It does not model partitions, eclipse attacks or targeted DoS, which is where the all-keys and active-total definitions differ in safety. And it puts no price on keys: the splitting result assumes 10,000 to 50,000 keys are free to create and to sign with, which the P2P and VRF layers may make expensive in ways this model cannot see.
|
||||
Loading…
Reference in a new issue