Merge remote-tracking branch 'origin/master' into release-0.3.15
2
.github/workflows/ci.yml
vendored
|
|
@ -97,6 +97,8 @@ jobs:
|
|||
run: bash tools/ci/commit-string-check.sh --self-test
|
||||
- name: build server remote checkout self-test (the stale-overlay class of 6 October 2026)
|
||||
run: bash infra/build-server/remote-run.sh --self-test
|
||||
- name: no shell assignment hides behind a trailing comment (the swallowed-defaults class of 6 October 2026)
|
||||
run: bash tools/ci/defaults-line-check.sh --self-test && bash tools/ci/defaults-line-check.sh
|
||||
- name: no secret file names and no 64-hex secrets in the tree (self-test first, then the tree)
|
||||
run: bash tools/ci/no-secrets-check.sh --self-test && bash tools/ci/no-secrets-check.sh
|
||||
- name: faucet unit tests (validation, the daily limits, the signed transaction; keccak, RLP and secp256k1 vectors)
|
||||
|
|
|
|||
|
|
@ -2,36 +2,160 @@
|
|||
|
||||
Written 6 to 7 October 2026 by the Horizon coordinator (branch `horizon`, worktree `igneum-wt-horizon`) on the project lead's ask of 6 October 2026, 20:4x UK: "deep backward and forward predictive research and modelling for our algo and all of our systems: is there room for improvement, room for more coin utility, algo improvements, security improvements, anything we can do to make a 51% attack impossible, basically creating a level of polish that has not been seen before." Later the same evening: "research anything else that we can research too, predictions, forward thinking, what we can actually do that has not been done or applied, think outside the box", "be revolutionary", and "if we create a new way of hashing or a new way of proof of work to revolutionise the space then that's absolutely fine, I want you to deploy everything to create something that has not ever been done before."
|
||||
|
||||
The bar main set, and the bar this document holds every claim to: "impossible" is not available to any proof-of-work chain. The bar is that a majority of hash buys nothing: it cannot reverse what finality locked, cannot forge a proof the nodes re-execute, cannot change a rule without 95 percent signalling, and loses more than it earns. Every claim here is a number with a model or a simulation behind it, priced in rented hash at the measured 6 October rate (`docs/bench-log.md`, "Rental cost of hash, 6 October 2026": USD 0.0117 per MH/s-hour, the live devnet at 1.16 GH/s), with its consequence per tier and what to build.
|
||||
The bar main set, and the bar this document holds every claim to: "impossible" is not available to any proof-of-work chain. The bar is that a majority of hash buys nothing: it cannot reverse what finality locked, cannot forge a proof the nodes re-execute, cannot change a rule without 95 percent signalling, and loses more than it earns. Every claim here is a number with a model or a simulation behind it, priced in rented hash at the measured 6 October rate (`docs/bench-log.md`, "Rental cost of hash, 6 October 2026": USD 0.0117 per MH/s-hour, so USD 11.7 per GH/s-hour and about USD 281 per GH/s-day; the live devnet at 1.16 GH/s), with its consequence per tier and what to build. Hours are agent hours (the project lead's rule: Claude-side work takes hours, never weeks).
|
||||
|
||||
## The lanes
|
||||
|
||||
| Lane | File | State |
|
||||
|---|---|---|
|
||||
| 1 consensus-security | `docs/analysis/horizon/consensus-security.md`, and the paper `docs/analysis/51-percent.md` | running |
|
||||
| 2 algorithm | `docs/analysis/horizon/algorithm.md` | running |
|
||||
| 3 finality-and-weight | `docs/analysis/horizon/finality-and-weight.md` | running |
|
||||
| 4 economy-and-utility | `docs/analysis/horizon/economy-and-utility.md` | running |
|
||||
| 5 network | `docs/analysis/horizon/network.md` | queued (waits for the 10 bps rows of `block-rate-devnet2.md`) |
|
||||
| 6 polish | `docs/analysis/horizon/polish.md` | queued |
|
||||
| 7 frontier | `docs/analysis/horizon/frontier.md`, model `sim/horizon/frontier/frontier_model.py` | landed 6 Oct 2026, 21:1x UK |
|
||||
| 8 new-proof-of-work | `docs/analysis/horizon/new-pow.md`, prototypes in `proto-newpow/` | running (designs tonight, measured rows tomorrow afternoon UK) |
|
||||
| 1 consensus-security | `docs/analysis/horizon/consensus-security.md`, and the paper `docs/analysis/51-percent.md`; models `sim/horizon/consensus-security/` | landed, commit eec2cd7 |
|
||||
| 2 algorithm | `docs/analysis/horizon/algorithm.md`; model `sim/horizon/algorithm/model.py` | landed, commit 5ff7393 |
|
||||
| 3 finality-and-weight | `docs/analysis/horizon/finality-and-weight.md`; models `sim/horizon/finality-and-weight/` | landed, commit c3aa502 |
|
||||
| 4 economy-and-utility | `docs/analysis/horizon/economy-and-utility.md`; models `sim/horizon/economy-and-utility/` | landed, commit 4617a01 |
|
||||
| 5 network | `docs/analysis/horizon/network.md` | running (takes the 10 bps runs A, A2 and the 1 bps control B) |
|
||||
| 6 polish | `docs/analysis/horizon/polish.md` | running |
|
||||
| 7 frontier | `docs/analysis/horizon/frontier.md`; model `sim/horizon/frontier/frontier_model.py` | landed, commit 5ff7393 |
|
||||
| 8 new-proof-of-work | `docs/analysis/horizon/new-pow.md`; prototypes `proto-newpow/` | designs landed (5ff7393); measured rows by the afternoon of 7 October UK |
|
||||
|
||||
## 1. One page for the project lead
|
||||
|
||||
(Written last, from the lanes.)
|
||||
What the night found, in the order it matters.
|
||||
|
||||
1. **The one line where a majority earns more than it spends is the proving pool, not the chain.** Consensus checks a carried proof record's signature and its statement, not the proof (spec 07, 7.7 items 3 and 4; ledger P21, decided as v0 through the public testnet). Any block producer can therefore carry its own fake-proof records and be paid its block share of the 20 percent pool: 11,636 IGN an hour at 51 percent of blocks (lane 1, `cost_model.py`). Every other attack line costs more than it earns. Fix: verify the aggregated segment proof in consensus, 8 to 12 hours, no liveness cost. Until it lands, the public text must not say the 20 percent pool is "paid to provers" without the caveat.
|
||||
2. **Tonight's two-hour finality pause was the rule working, then the frozen table holding it, and the fix is a signed exit.** Twenty keys holding 42.7 percent of the frozen voter table left in three minutes (the class v4 rehearsal job); checkpoint 6843 saw 53.1 percent of total and the 2/3 rule paused as designed; locks had formed without the hub, so topology is refuted (lane 3, from the observer rows). Rule v2 would have re-locked after 35 minutes; rule v3's frozen table holds it for a window: 2 hours on the devnet, 30 days on mainnet for the same event. Of five candidate rules simulated, only the departure announcement (a `leave` item in blocks, the key out of every denominator one hour later) keeps zero conflicting locks in every partition and eclipse AND ends the pause in under an hour (6 hours). The operational rule costs nothing: a standing box never leaves the live chain for an experiment, and any orchestrated departure over 10 percent of weight goes in slices under 10 percent an hour.
|
||||
3. **A 51 percent attacker on Igneum buys a 90-second reorder window and nothing past a certificate.** A 45 to 51 percent withholder wins the selected-chain race over a 90-s hold 70 to 85 percent of the time (32 to 46 blocks at 1 bps); the lock lands 63 to 93 s after the checkpoint block and bounds what a majority can reorder; sustained withholding lifts weight share to about 56 percent, never two thirds (lane 1, GHOSTDAG simulator mirroring `protocol.rs`, 20 seeds). The veto (one third of weight) costs 20 days at 51 percent of hash: USD 6k at 1 GH/s, USD 5.8 M at 1 TH/s, half earned back as subsidy; once held, a pause is free and a pause-time 12-hour double spend costs USD 146 to 146k. Weight-gated deep fork choice (a tip forked more than 10 minutes back is a candidate only if its builders hold a third of the weight table at the fork) and vote-or-burn (a silent key's blocks burn 20 percent of their producer share) close those two lines, 16 to 24 hours together, no liveness cost. The paper is `docs/analysis/51-percent.md`.
|
||||
4. **The chip question is settled in kind and open in degree.** The chip that matters is the stored-dataset memory-controller chip: 5.7x per joule against the 5090 at class v3, 2.1x at class v4 with a chip core as efficient as the GPU's (k = 1), 0.9x against the M5 Max (lane 2, `chip-model-v3` method). The reserve and the era draw buy about nothing against a chip (every drawn parameter is firmware; a reserve block is about USD 4 of silicon) and the public text should say what they do buy. The lever is the latency-shadow size N, and lane 2 and lane 7 agree it belongs in the era draw at genesis as a verifier-bounded ladder {100k, 130k, 200k, 330k, 650k, 1.0M}, each step by 90 percent miner signal, never unconditional (an unconditional doubling retires the M5 Max at era 1). First measured verifier proxies tonight: class v4 5.06 ms cold on a box core, 8.23 with the SMT sibling loaded (passes); dr736 10.51 and 15.49 (out); R0 is dr368.
|
||||
5. **Fees are not a security budget for a decade, and the proving price is a function of network hash.** All utility curves together pay miners and provers USD 450 a day at launch and USD 4,200 in year 5 against USD 54,800 and 13,700 of daily emission (lane 4). The price a prover must charge is the subsidy it forgoes, 1 / network hash: 100 to 300x Boundless's rate at 1.16 GH/s, 0.2 to 0.4x at 100 GH/s beside its miner. The adopted job floor overprices the market above about USD 0.014 per IGN. A ten-day prover refusal strands 547,570 IGN a day of pool credit in an escrow with no rule to return it. Fixes: decouple the job price from `f_p` (16 hours), roll unproven credit forward (8 hours), publish the price as a formula, never a number (3 hours).
|
||||
6. **The signalling window can be bought for a day.** The P2 one-day 95 percent window costs about 19 N of hash for 24 hours (USD 534k at 100 GH/s) and would force a class flip onto a fleet not yet on the object; a 6 percent holdout buys delay to the floor for nothing. Seven consecutive daily windows with the floor a week past publish fixes it, 3 hours. The three signalling thresholds (60 parameter, 90 upgrade, 95 class with floor) are stated inconsistently across spec 5.7, CLAUDE.md and the litepaper and must become one sentence.
|
||||
7. **New proof of work.** Scheme A (mining is proving) is ruled out twice over, by lane 8's design pass (2.9 MB of openings per block, a 32 to 40 ms proof verify against the 10 ms gate, sampleability) and by lane 7's prior-art pass (Ball et al. 2017, Ofelimos 2022, Aleo). Schemes B (a tensor-shaped integer shadow that forces a chip to carry a GPU-class datapath) and C (the dataset derived from the execution state, so every hash proves the miner holds the chain) are in prototype on two rented 4090s; measured rows land by the afternoon. (Section 4, lane 8, when it lands.)
|
||||
8. **The frontier list, honestly cut.** Do now: the finality weight table carried inside the recursive segment proof (a consensus proof at mergeset cost, a browser that trusts no node for the voter set; 60 hours, measure the in-guest BLS and colouring cycles first), reproducible-build attestations with Ember refusing a release under N of M (12 hours), a WASM verifier of the wrapped block proof in the tab (16 hours), Ember as node, wallet and light client for everyone (20 hours). Never: proving others' chains as the main income (all of Ethereum L1's proving is about USD 36 a day at the Sep 2026 tracker cost against USD 13,700 a day of year-1 emission at USD 0.005), the hash partly a proof, proof verify on a hardware wallet's secure element, a general unverifiable compute market, burn-redirect audit bounties (a dev fund with a veto).
|
||||
|
||||
Rows from lanes 5 (network: block rate, the controller's red-block feedback seen in run A, node and bandwidth tiers), 6 (polish: the ten things the project lead would notice) and 8 (the two prototypes' measured rows) join this page and the table below when they land.
|
||||
|
||||
## 2. The ranked list: top 25 across every lane
|
||||
|
||||
Rank is payoff over cost across lanes, with safety first, then liveness, then money, then text. "L1 r3" means lane 1's own rank 3; the lane file holds the full evidence row.
|
||||
|
||||
| Rank | Item | Lane | Evidence | Model | Hours | Consequence per tier | What to build | Gate |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| 1 | Verify the aggregated segment proof in consensus; a record whose proof fails against the pinned aggregator key is invalid | L1 r1 | spec 07 7.7 items 3 and 4; P21; a producer captures its block share of the 20 percent pool with fake records | capture = pool x (H inside the window + most outside): 11,636 IGN/h at 51 percent; SP1 light-verifier cost per record 1.8 to 2.4 s measured (bench-log) | 8 to 12 | prover: paid only for proofs; holder: the pool is real; node operator: one verify per record on the proof pool thread; miners: nothing | the verify call in the record check of `consensus/core` proving, the v0 per-shard record payout-only and capped at the exclusive window | fast-time: a fake-proof record is rejected by every node and the producer loses the block; a true record pays |
|
||||
| 2 | The departure announcement: a `leave` item (key, DAA score, signature) in blocks; D = 1 h later the key is in no denominator, sliding or frozen, and its votes are invalid; app Stop and the fleet library send it | L3 r1, L1 r8 | tonight's pause (L3 3.1): 42.7 percent left in 3 min, the frozen table held a window; sim T: first lock 1 h after a 34, 45 or 50 percent departure, 0 conflicts in every partition, eclipse and equivocator row | `finality_horizon.py` seeds 7, 11, 13 | 6 (+4 for F5's trusted certificate) | holder: a 30-day mainnet pause becomes 1 h when leavers are honest; a silent leaver still costs a window; pool and rig: one message on a clean stop; attacker: buying keys to leave them gains nothing (w + L must still reach 2/3) | the item in the coinbase finality section, the denominator rule in `finality.rs`, the send in Ember and `tools/fleet/lib` | harness: 45 percent leaves with leaves, lock within D + 1 checkpoint; 0 conflicts in 50/50, 60/40 and the 34 percent eclipse |
|
||||
| 3 | Signalling over 7 consecutive daily windows at 95 percent, the floor no nearer than 7 days past the publish, the stale-box list empty before the floor | L1 r6, L4 4.5 | the P2 one-day window is buyable: 19 N of hash for 24 h (USD 534k at 100 GH/s) forces a flip onto an unready fleet; a 6 percent holdout buys delay for nothing | `signalling.py`, `signal_game.py` | 3 | pool and rig: a week more before a class change and a week of visible share; home miner: a week to update; attacker: the bill x7 in public | the window count in the P2 rule, spec 5.7 text, one sentence carrying 60 / 90 / 95 in spec, CLAUDE.md and the litepaper | fast-time: 6 of 7 days at 95 percent does not flip; 7 does; the harness's failed case |
|
||||
| 4 | Weight-gated deep fork choice: a tip whose fork point is older than D (10 min of past-median time) is a candidate only if the blocks on it since the fork were produced by keys holding at least 1/3 of the weight table at the fork point | L1 r2 | during a pause or the first 20 days a renter with fresh keys double-spends at the 12-h depth for USD 146 at 1 GH/s | `cost_model.py`; deterministic over the block's past | 10 to 16 | holder and exchange: rented hash cannot reorg past 10 min even while nothing locks; honest miners: nothing (their keys hold the weight); new chain: the first 20 days gain a bound they lack today | the candidate filter in GHOSTDAG tip selection over the certified weight table | fast-time: a 51 percent fresh-key fork from 15 min back is refused by every honest node; a 1/3-weight fork is accepted; partition heal unchanged |
|
||||
| 5 | Vote-or-burn: a block whose producer key has participation under 0.5 over the presence window in its own past burns 20 percent of its producer share | L1 r3 | the pause as a liveness attack has zero marginal cost once the veto is held | 0.2 x A x 0.8 x subsidy: 7,757 IGN/h at 34 percent (USD 155/h at 0.02) | 6 to 8 | attacker: a pause costs; honest home miner: nothing while signing (the official client signs every checkpoint); pool user: the pool's participation; partition-safe because the test is the block's own past | the participation read in coinbase validation, the burn in the subsidy split | fast-time: a 34 percent silent set's blocks pay 80 percent; a 50/50 partition burns nothing on either side |
|
||||
| 6 | Roll unproven shard credit into the next proven segment's pool instead of the escrow for ever | L4 r2 | a ten-day prover refusal strands 547,570 IGN a day (5.5 M IGN) with no rule; spec 5.3, 7.7 item 3, 7.8 item 7 are silent | `stress.py` refuse scenario, `stranded_share` | 8 | prover: credit is never lost to the market's slow days; holder: no undecided burn; rollup customer: nothing | the roll-forward in the pool credit split, spec 5.3 text | fast-time: 100 unproven then 10 proven segments return the escrow to zero |
|
||||
| 7 | Decouple the job price from `f_p`: reserve = measured proving electricity per pgas at the published settlement rate (spec 5.10.3), the 1.5 premium becomes a bid | L4 r1 | at the adopted floor a billion-cycle job is 15 IGN = USD 0.075 / 0.30 / 1.50 against Boundless's 0.21 (approximate); above USD 0.014 per IGN the chain overprices by rule | `utility.py` sections 1 to 2 | 16 | rollup customer: a price that can clear; prover: still paid above electricity; holder: more jobs, more burn | the reserve rule in spec 5.4 and 5.11, the bid field on the job | a simulated job book clears within 20 percent of Boundless's median at all three prices |
|
||||
| 8 | Operational rule, no protocol change: a standing box never leaves the live chain for an experiment; any orchestrated departure over 10 percent of weight is staged in slices under 10 percent an hour; the fleet library refuses to swap a standing box's chain | L3 r2 | the rehearsal took 36.5 percent at once; sim L2 and M4: a gradual departure costs nothing because every lock re-freezes the table | spec 3.7 item 2 | 1 | fleet operator: finality stays on; every tier: no pause from our own experiments | a refusal in `tools/fleet/lib` plus a `--staged` departure | the next rehearsal keeps `finality_active` true |
|
||||
| 9 | Peer floor and mesh, and no `unwrap` on a peer-driven path with a sync-request fuzz gate (tonight's pruned-node crash: `consensus/src/processes/sync/mod.rs:87`, plus 27 sibling sites listed in L1) | L1 r4 | run A's star (321 tips, 77 percent red); the hub crash at 19:57Z: any peer can crash any pruned node at zero hash cost | a star with its hub down is n islands; the sibling list by file:line | 9 to 11 | node operator: no crash from a peer's request; home miner and rig: the app alarms under 3 outbound peers or 2 checkpoints without a vote; fleet: no star | the fleet lib dials 3 boxes beside the hands; the node alarm and app line; the unwrap sweep; the fuzz in `tools/ci` | the fuzz runs the sync request space against a pruned node with no panic; the alarm fires on a known-cut case and stays quiet on a healthy one |
|
||||
| 10 | N (the latency-shadow size) as a genesis ladder per era {100k, 130k, 200k, 330k, 650k, 1.0M}, floor and ceiling fixed at genesis (10x verifier headroom on the 2.5x rule; O-1.14 sets the ceiling), one era draw consumed as for `epoch_len`, each step by 90 percent miner signal over 7 days, never unconditional | L2 r5, L7 finding 1 | HBM4 raises the stored-dataset chip's bare edge (4.9x to 12x depending on tFAW, L2 and L7 disagree and both are labelled); N = 100k holds the chip at 2.1x on GDDR7 at k = 1, N = 330k at 1.3x; an unconditional doubling retires the M5 Max at era 1 (-10.5 percent at 200k) | `model.py --section ladder` (L2 5.3a, the reconciled table) | 6 to 8 | 100k to 130k: the M5 Max -3.3 points, nobody else; 130k to 200k: M5 Max -6 more, the 5090 -2.7 at its cap, the 4070 +21 W; 200k to 330k: compute-bound on every capped NVIDIA card; pool users nothing at any step; verifier +0.17 to 0.56 ms per warp | the ladder in the era-draw table of spec 1.13, the step rule on P2's mechanism | per step: every public-benchmark card within 5 percent of its previous-step rate, bit-exact on three vendors, verifier under 10 ms cold on the O-1.14 core |
|
||||
| 11 | Aggregate-first vote verification and aggregated in-block carriage as the mainnet default (spec 3.4.2 item 2, decided) | L3 r3 | per-vote verification is 1.6 s per checkpoint per node at 1,000 voters and about 13 s at 8,192 (approximate); measured lock delay p50 0.86 to 1.26 s at 4 to 93 voters on node 1 | L3 4.3 table; blst 0.3.17 `fast_aggregate_verify` already in the fork | 8 | node operator on a laptop: a checkpoint costs one pairing, not a thousand; every tier: lock delay stays under 3 s at 8,192 voters | batch votes per (index, hash), verify once, bisect on failure; the in-block aggregate | fast-time with 1,000 and 8,000 synthetic voters: p50 lock delay under 3 s, CPU under 25 percent of a core, 0 conflicts |
|
||||
| 12 | The finality weight table carried inside the recursive segment proof, updated one mergeset per segment: a consensus proof at mergeset cost | L7 r1 | phase two's hardest item (P4) becomes incremental on code that exists; about one BLS verify per 30 s (one shard's budget, approximate, unmeasured) | `frontier_model.py` | 60 (first 10: measure) | phone wallet and bridge: finality from one proof, no node trusted for the voter set; rollup customer: one object; prover: one more public-value block per segment | the W2 transition in the aggregator guest, the GHOSTDAG colouring check in-guest | first: BLS verify and colouring cycle counts inside the SP1 guest on a 24 GB card within 2x of the estimate; then a proof per checkpoint under 30 s on a proving-only 5090 |
|
||||
| 13 | Reproducible-build attestations in a registry contract; Ember refuses a release under N of M attestations | L7 r3 | ledger G7, G13: the update key is an operator channel; Bitcoin's guix.sigs is the precedent with the chain as the sigs repo | text and contract | 12 | every tier: an update needs N independent builders, not one key; fleet: the box's reproducible Windows and Linux builds (CLAUDE.md 6 Oct) are the inputs | the contract, the builder tool, the Ember check | a release with 1 of 3 attestations is refused by Ember; 3 of 3 installs; the known-failed case |
|
||||
| 14 | Close O-1.14 with a real 2019 laptop run; adopt the half-core proxy as the standing stand-in; dr736 out, dr368 in as R0 | L2 r1 | box proxy: dr736 10.51 ms cold and 15.49 half-core; class v4 5.06 and 8.23 | L2 5.5 | 2 | miners 0; pool verifier cores x1.3 at dr368 instead of x2.4; a 2019 node keeps 1.8 ms under the gate on class v4 | Windows cross build on the box, relay to the laptop, bench, log | ms per warp under 10 cold on the real core for class v4 and dr368 |
|
||||
| 15 | The vote signs the execution root too (chain id, index, block hash, post_root); a certificate pins the state; snapshots check against the last certificate | L1 r5 | snapshot poisoning: a wrong state above the pin is undetectable today | every voter executes natively (spec 07) | 8 to 12 | holder: a lock is a state lock; node joining from a snapshot: cannot be poisoned; cost: exec lag enters lock latency | the vote message, the pin rule, the snapshot check | fast-time: a poisoned snapshot node cannot join the quorum; honest lock delay rises by the measured exec lag only |
|
||||
| 16 | `finality_provisional` reported beside `finality_active`, never as a lock, plus a detector-driven `security_alert` (correlated group over 1/3 of a day's blue blocks; pause; conflict) shown by wallets and the explorer as "confirm at 12 h" | L3 r4, L1 r7 | sim P: provisional conflicts in every partition (516 to 1,062 per 360 min), final 0; the exchange guidance exists only as text today | L3 4.1; the detector of counter-asic-3 | 3 + 4 to 6 | exchanges: a third row in the guidance; holder: told when to wait; home miner: a line in Ember | the RPC fields, the explorer and wallet copy, spec 3.9 row | forced pause: explorer and wallet show provisional and the alert; a healthy day shows neither |
|
||||
| 17 | Every miner-signalled execution parameter enters the consensus digest the same release, with a `tools/ci` check that fails a `Params` field marked signalled and absent from the digest | L4 r5 | signal-then-defect on an execution parameter is a silent state fork unless the handshake refuses the defector; `Params.fees` is in the digest, nothing guarantees the next | `signal_game.py` section 3 | 6 | node operator: no quiet fork; every tier: a defector is refused at the handshake | the check, the digest rule in spec 5.8 | the check fails on a planted field and passes on the live set |
|
||||
| 18 | The 80/20 split as a 60 percent parameter inside a hard band [10, 30] percent | L4 r4 | not load-bearing at launch traffic; 30 percent buys backlog relief at 100 shards a block (economy-2026-10-04 5.3) | `stress.py` | 10 | prover and miner: a split the market can move inside a band nobody can break; holder: the cap and the halving untouched | the band in spec 5.3 and 5.7 | harness: a 60 percent vote reaches 30 percent; a 100 percent vote cannot pass the band |
|
||||
| 19 | A WebAssembly verifier of the wrapped block proof in the tab, with the millisecond count shown | L7 r4 | three working precedents (Helios WASM, ziren-wasm-verifier, wasm-groth16-verifier); the certificate half already runs at 58 to 68 ms warm, 139 to 155 ms cold (bench-log round 6, P3) | text | 16 | holder and rollup customer: verify in the browser; the P3 wrapper measurement comes forward | the Groth16 or Plonk wrapper, the WASM verifier on the site | a block proof verifies in the tab under 500 ms on a laptop; the measurement labelled on the page |
|
||||
| 20 | Ember as node, wallet and light client for everyone | L7 r5 | 1,000 testnet miners become 1,000 verifying nodes; Monero shows about 5,000 peers (monero.fail), Ethereum 8,136 execution nodes (ethernodes), approximate | text | 20 | home miner: one app; node count = miner count; Sybil counts of nodes irrelevant by design; node cost per lane 5's tiers | the node and wallet surfaces inside Ember | a fresh Windows install mines, verifies and shows a balance in the first 60 s with no second download |
|
||||
| 21 | Work-stake: vote weight as the external-job bond, with the spec sentence "no coin stake; the only thing at stake is 30 days of public work" written first | L7 r2 | a 0.1 percent key has about 16,427 IGN of 30-day pool income plus its vote at risk against a designed 0.0015 IGN coin bond per job | `frontier_model.py` section 3 | 24 (prototype) | prover: a bond nobody can buy; rollup customer: a griefing cost that scales with the prover's standing; holder: "no stake" stays true as "no coin stake" | the bond rule in the job claim | a griefed job costs the griefer its weight for 30 days in the fast-time harness; the spec sentence lands before the prototype |
|
||||
| 22 | Client-shipped certified checkpoint (index, hash, voter-table digest) refused if missing from the DAG; exec generations spaced geometrically to the finality depth (about 16, 1.8 GB today) | L1 r9, r10 | long-range and seed attacks; a pause-time deep reorg needs a peer's snapshot today | assumevalid's shape; 114.8 MB per snapshot measured | 3 + 3 | every tier: a fresh install cannot be bootstrapped onto a private DAG; node operator: disk for generations, no blocked executor after a deep reorg | the checkpoint in the release, the generation schedule | a cold node refuses a DAG missing the checkpoint; a 5,000-block reorg re-executes from a generation with no snapshot request |
|
||||
| 23 | Public-text corrections, one bundle: the prover's price as a formula with network hash as the input, never a number; the era draw and the reserve described as schedule changes against fixed datapaths, not unpredictability against a chip, with the chip's USD per MH/s-hour beside the honest cards'; F19's "day 19 to 20" is a v2 number (v3: a full window); funding.md section 4's dev-fee ceiling is 1 percent of the producer share (38,520 / 154,080 / 770,400); the dev fee is "default-on, switchable", not "optional"; "proofs at the cost of power" conditioned on hash; the three signalling thresholds in one sentence | L4 r3, r7; L2 r3; L3; L1 | each row cites its lane section | text | 1 to 3 each | holders and critics read claims that survive review | the litepaper, the customer brief, spec 5.7, funding.md, the ledger F19 | the site's forbidden-strings check and a reviewer's read |
|
||||
| 24 | Measure the FPGA lane on AWS F2 (one VU47P, HBM2, about USD 1.98 an hour) and replace the ceiling row | L2 r2 | the measured 2.4 G reads/s equals the JEDEC tFAW ceiling; the old 1.9x row rests on a 12 ns tFAW the JEDEC HBM2 table does not give (28 ns); the soft overlay reads 0.30x to 0.47x of the 5090 per watt | L2 5.1 | 8 to 10 plus USD 2 to 8 | none today; the public FPGA claim becomes a measured number | the overlay bench on F2 | reads per second per watt at 1 GiB; the row replaced |
|
||||
| 25 | Ember tune as the shipped default per card model, and the two fleet measurements every price rests on: the miner's hash loss while each tier proves, and a full 30 M-cycle shard beside the miner on 12 and 16 GB cards | L2 r7, L4 r8 | the 4070 at 3.65 uJ untuned and 2.57 tuned (-30 percent); the hybrid row is the only one that undercuts the market and rests on one 5090 measurement; the 4.7 M fixture is 16 percent of a full shard (linear scaling says 240 s on a 3060, outside the 120-s claim timeout) | L2 5.1; `utility.py hybrid_hash_loss` | 2 + 6 (fleet) | every NVIDIA tier gains 10 to 30 percent per joule; the 12 GB tier learns whether it can claim a full shard in time | the defaults table in the app from the fleet priors; eleven rows with both numbers in `prover-tiers-real-cards.md` | the rows land; the claim timeout is set from them |
|
||||
|
||||
Rows from lanes 5, 6 and 8 are inserted and the table re-ranked when they land (expected: the controller's red-aware correction and the block-rate recommendation from lane 5; the ten the project lead-visible polish rows from lane 6; the class v5 verdicts from lane 8).
|
||||
|
||||
### Not recommended, with the reason (as valuable as the list above)
|
||||
|
||||
| Item | Lanes | Why not |
|
||||
|---|---|---|
|
||||
| Prover attestations as a second finality leg | L1 r12, L3 r8 | provers are the miners (same vote keys), so no new party; proof coverage is 2.4 to 4.7 percent of blocks tonight with lag p99 62 s, so every lock would wait on proofs; revisit only after 99 percent of blocks are proven within 60 s for 7 days with 3 provers per block, and even then it adds about a minute of lock delay |
|
||||
| Time-locked (vesting) weight | L1 r13 | a bought key transfers vested weight, so the acquired-keys bound is unchanged; honest new cohorts wait longer |
|
||||
| Any automatic re-lock after an abrupt departure | L1 r14, L3 r7 | departure and partition are the same observation in one view; the decaying denominator produces 467 to 473 conflicting locks in a 360-minute 50/50 honest split, the hysteresis floor brings the 13.3 percent equivocator bound back (565 to 597 conflicts at 20 percent); the two-tier rule is a report, never a lock |
|
||||
| Mining is proving (the lottery's work as a proving step) | L8 scheme A, L7 3.10 | dead on bytes (2.9 MB of openings per block), on the verifier (32 to 40 ms against 10), and on sampleability (Ball et al. 2017, Ofelimos 2022); re-opens Aleo's fastest-prover-wins |
|
||||
| Proving others' chains as the main income | L7 3.11 | all of Ethereum L1's proving is about USD 36 a day at the Sep 2026 tracker cost against USD 13,700 a day of year-1 emission at USD 0.005; demand must grow 1,000x against a falling cost curve |
|
||||
| Burn-redirect audit bounties, a review escrow by 60 percent signal | L7 3.6, L4 task 3 | a dev fund with a veto, the switch spec 5.5 removed; the honest payer for a second audit is the entity's own provers |
|
||||
| Proof verification on a hardware wallet's secure element | L7 3.9 | a bn254 pairing on a Cortex-M-class element is seconds to minutes (approximate); do the companion verify |
|
||||
| A general compute market for unverifiable work (rendering, inference) | L7 3.12 | an escrow without a verifier is a trust-me payment with lower fees; the verifiable subsets are named and kept |
|
||||
|
||||
## 3. What a 51 percent attacker can and cannot do
|
||||
|
||||
(Summary of `docs/analysis/51-percent.md`.)
|
||||
The paper is `docs/analysis/51-percent.md` (lane 1). In one table, at the measured USD 11.7 per GH/s-hour:
|
||||
|
||||
| Hash share | What it buys | For how long | Cost at 1 GH/s / 1 TH/s of network | What it earns | Net |
|
||||
|---|---|---|---|---|---|
|
||||
| 20 percent | reorders only the last k = 18 blocks; nothing past a certificate | per attempt | subsidy forgone while withholding | its block share minus reds | loss |
|
||||
| 34 percent | wins the selected-chain race only inside 60 s (40 percent of the time); holds the veto (blocks every lock) after 20 days of mining at 51 percent, or 30 days at 52 percent of the network's hash | 20 to 30 days to acquire, then free to hold while silent | about USD 6k / 5.8 M to acquire (half earned back as subsidy); a pause then costs nothing | subsidy; nothing from the pause itself unless it double-spends at the 12-h depth during the pause (USD 146 to 146k) | the one line vote-or-burn (rank 5) and weight-gated fork choice (rank 4) close |
|
||||
| 45 to 51 percent | wins the selected-chain race over a 90-s hold 70 to 85 percent of the time (32 to 46 blocks at 1 bps); turns 26 to 28 percent of honest blocks red; lifts its weight share to about 56 percent, never 2/3 | the 63 to 93 s between a checkpoint block and its lock | USD 11.7 / 11,700 per hour of rented hash; the market could not supply a TH/s on 6 October (RunPod: 0 pods for 20 asks) | a quarter of honest subsidy while withholding; the pool capture of rank 1 (11,636 IGN/h) until the in-consensus verifier lands | loss on the chain; gain only through the pool, which rank 1 closes |
|
||||
| 67 percent of weight | locks alone | 2.03 x N of hash for 30 days: USD 17,100 per GH/s of network | | subsidy | at rental equilibrium about 33 percent of 30 days of subsidy net |
|
||||
| any share | a rule change | 95 percent of blue-block weight over the window, or the floor | 19 N for 24 h today (USD 534k at 100 GH/s); x7 under rank 3 | | |
|
||||
|
||||
What no share buys: a block the nodes do not re-execute (the native-execution veto), a state the aggregated proof chain does not commit to, a lock past a certificate under two thirds of weight, a parameter change without the signal.
|
||||
|
||||
The residual risks stated plainly: the first 20 days (no weight table yet); a pause after a sudden departure of a third of weight (30 days under v3 until rank 2 lands); the Sybil count of keys (X5's definition adopted, the measurement scheduled); the proving pool until rank 1; the 2/3-of-total rule under long churn (the window bound: a 50/50 honest partition locks on both sides from day 10).
|
||||
|
||||
## 4. Per lane: the three biggest findings
|
||||
|
||||
### Lane 1, consensus-security
|
||||
1. The lock bounds a majority; the k-cluster does not (45 to 51 percent wins a 90-s race 70 to 85 percent of the time; the lock at 63 to 93 s is the bound).
|
||||
2. The veto is cheap (20 days of 51 percent: USD 6k at 1 GH/s, 5.8 M at 1 TH/s, half earned back) and the pause is then free.
|
||||
3. The proving pool is capturable today (11,636 IGN/h at 51 percent) with a correct-statement fake-proof record; the only attack that earns more than it costs.
|
||||
|
||||
### Lane 2, algorithm
|
||||
1. Class v4's verifier measured on a 2022 server core: 4.90 / 5.06 / 8.23 ms (steady / cold / half-core); dr736 9.76 / 10.51 / 15.49, out; R0 is dr368.
|
||||
2. The f = 1 GDDR7 chip: 5.7x per joule against the 5090 at v3, 2.1x at v4 with k = 1; USD 0.00021 per MH/s-hour against USD 0.0117 rented; the FPGA soft overlay 0.30x to 0.47x per watt.
|
||||
3. The reserve and the era draw buy about nothing against a chip; the N ladder in the era draw at genesis is the lever, and lane 7's HBM4 column is reconciled with three named disagreements.
|
||||
|
||||
### Lane 3, finality-and-weight
|
||||
1. Tonight's pause: the 2/3 rule at checkpoint 6843 (53.1 percent of total after a 42.7 percent departure), then the frozen table (Q5) holding it for a window; locks formed without the hub; topology refuted.
|
||||
2. Only the departure announcement passes the line (0 conflicting locks everywhere, the pause under an hour); the decaying denominator, the hysteresis floor and the two-tier rule fail or are reports.
|
||||
3. Weight capture costs USD 8,424 x N x W/(1 - W) to rent for the window: the veto 0.52 N for 30 days (USD 4,300 per GH/s of network), a lock alone 2.03 N (USD 17,100); lock delay measured p50 0.86 to 1.26 s at 4 to 93 voters; the path breaks near 1,000 voters on per-vote verification.
|
||||
|
||||
### Lane 4, economy-and-utility
|
||||
1. The proving price is h/N: 100 to 300x Boundless at 1.16 GH/s, 0.2 to 0.4x at 100 GH/s beside the miner; the adopted floor overprices above USD 0.014 per IGN.
|
||||
2. Fees are not a security budget for a decade (USD 450 a day at launch, 4,200 in year 5, against 54,800 and 13,700 of emission); sustained honest hash costs USD 24.8 per GH/s-day against USD 281 rented, so the 20-day veto costs 11.8x the honest fleet at every price and year.
|
||||
3. The 80/20 survives every stress but a ten-day prover refusal, which strands 547,570 IGN a day of pool credit with no rule to return it.
|
||||
|
||||
### Lane 7, frontier
|
||||
1. HBM4 raises the stored-dataset chip's edge (lane 2 bounds the figure at 4.9x to 12x bare by tFAW); the N schedule belongs in the era draw at genesis.
|
||||
2. Vote weight is already a slashable, non-purchasable bond: work-stake for external jobs, with "no coin stake" written into the spec first.
|
||||
3. The consensus proof can be incremental: the weight table inside the recursive segment proof, one mergeset per segment; the first measurement is the in-guest BLS verify and colouring cycle counts.
|
||||
|
||||
### Lanes 5, 6, 8
|
||||
(Filled when they land.)
|
||||
|
||||
## 5. What was not run, and why
|
||||
|
||||
| Item | Lane | Why |
|
||||
|---|---|---|
|
||||
| The in-guest BLS verify and GHOSTDAG colouring cycle counts (rank 12's first gate) | 7, 3 | no 24 GB card free tonight: PC 2 and the fleet were on class v4 and Devnet 2 |
|
||||
| The real 2019-class core (O-1.14) | 2 | the US laptop was not on the relay; the box's half-core proxy stands in |
|
||||
| The FPGA soft overlay on real HBM2 | 2 | no FPGA in the fleet; AWS F2 plan written |
|
||||
| The AMD watts at every N | 2 | the ADLX sampler row is owed on the runner's `--cards-off` mechanism (counter-asic-3-status 6a) |
|
||||
| A `cargo bench` of `fast_aggregate_verify` at 93, 1,000 and 8,192 keys | 3 | the BLS figures are arithmetic on blst's published timings, approximate |
|
||||
| The 10 bps runs A2 (mesh) and B (control) | 5 | landing during the night; lane 5 reads them before it closes |
|
||||
| The full 30 M-cycle shard beside the miner on any tier; the hash loss while proving on Ampere and Ada | 4 | the fleet measured the 4.7 M fixture only |
|
||||
| Market prices for Taiko-class batch proving and Bonsai | 4 | not published; marked approximate |
|
||||
| The observed end of tonight's pause | 3 | expected at DAA 216,402, about 20:40Z; the timeline closes when the observer rows show the lock |
|
||||
|
||||
## 6. Rules and corrections for main
|
||||
|
||||
| # | Rule or correction | Lane | State |
|
||||
|---|---|---|---|
|
||||
| 1 | P21 is a priced economic hole until the in-consensus verifier lands; the 20 percent pool is not "paid to provers" without the caveat | 1 | relayed |
|
||||
| 2 | The one-day P2 signalling window is buyable for a day; seven consecutive windows | 1, 4 | relayed |
|
||||
| 3 | The DAG controller reads blue work only, so a withholder or a star eases difficulty 16 to 32 percent (run A's loop) | 1, 5 | relayed; lane 5 owns the fix |
|
||||
| 4 | No `unwrap` on a peer-driven path; the sibling list; a sync-request fuzz in `tools/ci` | 1 | relayed |
|
||||
| 5 | A standing box never leaves the live chain for an experiment; orchestrated departures over 10 percent staged under 10 percent an hour | 3 | relayed; the fleet library refusal is rank 8 |
|
||||
| 6 | Under rule v3 any sudden departure of a third of weight is a 30-day mainnet pause; the leave item is the one safe way to shorten it | 3 | relayed |
|
||||
| 7 | F19's "stalls until day 19 to 20" is a v2 number; under v3 a full window | 3 | relayed, ledger text owed |
|
||||
| 8 | "No coin stake; the only thing at stake is 30 days of public work" into spec 03/05 and the litepaper before any work-stake prototype | 7 | relayed; routed to the site-miner agent |
|
||||
| 9 | The litepaper's "proving: a second income" carries the USD 36 a day arithmetic | 7 | relayed; routed |
|
||||
| 10 | Ember's updater installs nothing while finality is paused | 7 | relayed; sent to the Ember agent for 0.3.16 |
|
||||
| 11 | The customer brief's "priced in dollars per proof" and the litepaper's "proofs at the cost of power" are conditioned on network hash | 4 | relayed |
|
||||
| 12 | The three signalling thresholds (60 / 90 / 95) in one sentence across spec 5.7, CLAUDE.md and the litepaper | 4 | relayed |
|
||||
| 13 | The spec is silent on pool credit nobody claims; today it is a burn nobody decided | 4 | relayed |
|
||||
| 14 | funding.md section 4 overstates the dev-fee ceiling by a quarter | 4 | relayed |
|
||||
| 15 | Any future miner-signalled execution parameter outside the consensus digest makes signal-then-defect a quiet state fork | 4 | relayed |
|
||||
| 16 | The era draw and the reserve are not unpredictability against a chip; the public text should say what they buy | 2 | relayed |
|
||||
|
|
|
|||
244
docs/analysis/horizon/network.md
Normal file
|
|
@ -0,0 +1,244 @@
|
|||
# Horizon lane 5: network. Block rate, propagation, node cost, bandwidth, pruning and the snapshot path
|
||||
|
||||
6 October 2026, evening UK. Lane 5 of the Horizon programme. Worktree `/Users/joshm/Projects/igneum-wt-horizon`, branch `horizon`. Scripts in `sim/horizon/network/` (README there). Nothing live was touched: the Devnet 2 seed log on igneum-build-1 and `~/Desktop/fleet/bps/A.jsonl` were read, never written.
|
||||
|
||||
What was read: `docs/spec/02-consensus.md` (2.1 parameters, 2.3 the difficulty rule, 2.4 header, 2.5 emission), `08-client-security.md`, `10-light-client.md`; the fud-close worktree's `docs/spec/03-finality.md` C1 and 3.4.2 (vote item 281 bytes, the bitmap and per-block bounds); `docs/fud-ledger.md` M20, M21 (block sizes, k re-derived, the 490 KB body run), X20, M30; `docs/bench-log.md` entries "4 October 2026, cloud devnet" (inter-region RTT and propagation), "ledger M30" (RSS, the 256 MiB cache, the s8 steady slope), "6 October 2026, 12:25 to 13:20Z, the finality route" (the seed as the only peer, the route overflow), "Rental cost of hash, 6 October 2026"; `docs/plans/hands-on-build-1.md` (node 1 and the observer: data dirs 856 and 822 MB, the 127.6 MB exec snapshot), `seed-nodes.md`, `cloud-devnet.md`, `release-0.3.14.md` (the snapshot path and the restart pin), `release-0.3.13.md` 4a; the fleet worktree's `docs/analysis/block-rate-devnet2.md` (read at 19:5xZ and again at 20:1xZ: RUN_A, RUN_B, TIER_TABLE and RECOMMENDATION are still placeholders), `tools/fleet/bps-collect.py`, `box-dn2.sh` (one `--addpeer=<seed>`: the star), `devnet2-override.json`, `docs/bench-log.md` of that branch; the live record of run A: `~/Desktop/fleet/bps/A.jsonl` (43 rows, 19:18 to 19:51Z), `collect-A.log`, `runA-start2.log`, `runB.sh` (run B starts at 20:25Z) and the seed's log `/home/build/dn2seed-A.log` on igneum-build-1 (20,529 lines, 19:17 to 19:59Z, read over ssh); `vendor/rusty-kaspa` `consensus/core/src/config/bps.rs` (the k table, `calculate_ghostdag_k`, parents, mergeset, pruning depth), `constants.rs` (delay bound 5 s, delta 0.01, DAA window), `params.rs` (mainnet `BlockrateParams::new::<10>()`, Crescendo activation 110,165,000), `consensus/src/processes/difficulty.rs` and `window.rs` (the window holds every mergeset block, blue and red), `docs/crescendo-guide.md` (10-bps node requirements); the fork `vendor/igneum-node` (read-only) at `release-0.3.14-node`: `consensus/core/src/igneum.rs` `difficulty` (CAP_BLOCKS 20, the sanitised clock, the clamps), `consensus/src/processes/difficulty.rs` `igneum_difficulty_bits` (blue-work steps on the selected chain), `consensus/src/processes/finality.rs` (MAX_VOTES_PER_BLOCK 48, KEEP_CHECKPOINTS 2,000), `igneum/exec/src/proving.rs` `p2p_snapshot_gate` and `on_exec_snapshot`, `igneum/exec/src/snapshot.rs`; the fork branch `devnet2-bps` 1279a1d6 (`IGNEUMD_DEVNET_BPS`); lane 3's `finality-and-weight.md` sections 5.5 and 8 (the aggregation path tonight), lane 4's `sim/horizon/consensus-security/ghostdag_results_{1,10}bps.md`.
|
||||
|
||||
## 1. The question and the answer in one paragraph
|
||||
|
||||
the project lead asked for Kaspa's answer to solo-miner variance: a higher block rate. Run A ran Devnet 2 at 10 blocks per second through one seed and produced 77 percent red blocks, 321 tips and a 55-block reorg. The propagation model says the links and the star did not do that: with the measured latencies it predicts under 0.1 percent red at 10 bps in a star and in a mesh. What did it is the seed's CPU per block, measured at 61 ms (narrow DAG) to 345 ms (mergeset 150 to 200), against a budget of 100 ms per block at 10 bps; with that cost in the model the star gives 47 to 88 percent red and queueing waits of 26 to 1,769 s, which are the "Accepted 100 blocks via relay" batches in the log. The difficulty rule then read blue work over a chain step capped at 2 s and hardened until the DAG ran at 2 s / (chain-step spacing) of target (model 4 blocks/s at a 5-s spacing, record 3.3 to 3.6), while a narrower-but-still-wide DAG would have made it ease (the direction main reported); counting every mergeset block over the real span, as Kaspa's window does, is unbiased in both regimes. The block rate for the public testnet is 1 bps; 10 bps is a gated step that needs the per-block node cost under 50 ms on a laptop core at a mergeset of 248, the checkpoint interval and the clock cap re-denominated in DAA seconds, and vote aggregation, because with C1 in blue blocks 8,192 voters at 10 bps are 66 GB per node per day of votes.
|
||||
|
||||
## 2. Method
|
||||
|
||||
| Step | What | Where | Machine, lock |
|
||||
|---|---|---|---|
|
||||
| Record | run A's per-minute rows (blocks, blue score, tips, peers, exec tip, CPU, RSS, bytes), the seed log's `PoW accepted` per minute, `Processed` lines (parents, mergeset per 10 s), `Finality: checkpoint N determined` (blue score against DAA), reorg lines, route drops | `~/Desktop/fleet/bps/A.jsonl`; `/home/build/dn2seed-A.log` on igneum-build-1 | read only |
|
||||
| Propagation | event simulation: Poisson production, star or mesh, lognormal links, inv/request/block hops, a hub (or every node) as a single server with a per-block cost, GHOSTDAG colouring with the first k-cluster condition, parent and mergeset caps, reorg depth at miner 0 | `sim/horizon/network/propagation.py`, `results.md`, `results-2.md` | Mac, `with-lock.sh run nice -n 19`, seed 7 |
|
||||
| Controller | the fork's estimator against a wide DAG in closed loop with the 3% / 10% clamps, against the whole-DAG estimator and a red-corrected blue estimator | `controller.py`, `controller-10bps.md`, `controller-1bps.md` | Mac, seed-free (deterministic) |
|
||||
| Arithmetic | k and parameters per rate, votes and bounds, bytes per node per day, CPU budget, RSS and disk, pruned and archival growth, subsidy and payout intervals, finality timing, light-client bytes | `cost.py`, `cost-tables.md` | Mac |
|
||||
|
||||
Nothing was built. No node ran. The fleet's run B (1 bps control, starts 20:25Z) and the mesh variant A2 (not scripted in the fleet worktree at 20:1xZ: no `mesh` or `A2` in `tools/fleet/` or its plans) had not landed when this file closed; section 7 names what they owe.
|
||||
|
||||
## 3. Evidence
|
||||
|
||||
### 3.1 Run A, measured (igneum-devnet-2, 10 bps, 41 miners through the seed on igneum-build-1, genesis bits 505413632)
|
||||
|
||||
Phases from the seed log's checkpoint series (blue score against DAA score; red share = 1 - blue / DAA over the interval) and `Processed` lines; CPU per block from `A.jsonl` `cpu_rss` (ps lifetime %CPU times elapsed, differenced) over the accepted-block deltas.
|
||||
|
||||
| Window (Z) | Production at the seed, blocks/s | Blue rate, blocks/s | Red share over the window | Mergeset per block (Processed) | Parents | Tips at the seed | Hub CPU per accepted block | Hub RSS | What the log shows |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| 19:23 to 19:29 | pods join (peers=1 each, blocks=0 synced=false on some), 26 blocks/s peak at 19:30 | | 33% (cp 1 to 29) | 1 to 35 | 1 to 16 | 6 to 30 (pods) | 61 ms | 0.45 to 1.9 GB | the 55-block reorg at 19:27:41Z "unwinding to height 1" (a late-joining pod's chain from genesis) |
|
||||
| 19:29 to 19:31 | 19 to 26 | 1.3 | 77% (cp 29 to 37) | 29 to 36 | 15 | | 115 ms | 1.9 to 2.5 GB | |
|
||||
| 19:31 to 19:37 | 13 to 19 | 0.7 | 94% (cp 37 to 45) | 53 to 197 | 12 to 15 | 295 | | 2.5 to 4.0 GB | `Accepted 97 / 100 / 31 blocks ... via relay` batches; headers ahead of blocks |
|
||||
| 19:37 to 19:41 | 8 to 14 | 0.5 | 95% (cp 45 to 49) | 150 to 190 | 12 | 218 to 295 | 345 ms (19:37 to 19:51 mean) | 4.0 to 4.5 GB | window filled at 19:37Z; first lock cp 47 at 19:40:08Z |
|
||||
| 19:41 to 19:47 | 3 to 7 | 1.0 | 50 to 88% (cp 49 to 61) | 80 to 125 | 13 to 15 | 230 to 350 | | 4.6 to 5.2 GB | `incoming route for IgneumFinality is full, message dropped ... the peer stays`: 109 to 2,547 drops per peer by 19:56Z |
|
||||
| 19:47 to 19:59 | 3.3 to 3.7 | 1.4 to 1.8 | 41 to 56% (cp 61 to 92) | 47 to 90 | 15 to 16 | 348 to 396 | | 5.4 GB at 19:51 | 16 of 47 checkpoints after the window filled LOCKED (cp 47 to 79); cp 61 determined 19:47:00, locked 19:47:18 |
|
||||
| Cumulative to 19:59:32Z | 6.1 (DAA 12,233 in 2,012 s) | 1.38 (blue 2,786) | 77.2% | | | | | | main's 12.4 blocks/s, 77 percent, 321 tips, max reorg 55 are the 19:3xZ to 19:45Z readings of the same record |
|
||||
|
||||
Other measured rows of the run: 421 selected-chain reorgs at the seed (132 of depth 1, 62 of 2, 42 of 3, 25 of 4, 19 of 5; 2 of 43, 1 of 44, 1 of 55); the seed's `rx_tx` counter (netns-wide, so an upper bound) 325 MB in and 2,133 MB out between 19:37:05 and 19:51:15Z (850 s, 4,284 accepted blocks): 2.5 MB/s out, 61 KB/s per peer, about 12 KB per block per peer (a block with its finality section of up to 48 votes at 281 bytes is about 14 KB); exec tip 154 chain blocks at 19:51Z against 11,239 blocks (the executor follows the selected chain, which advanced about one chain block per 10 s at the widest). The run's difficulty values are not in the record: `bps-collect.py` strips `difficulty=` from the watch line before storing it and the node log carries no bits; main's statement that the rule "lowered difficulty" is therefore unverified here, and the production curve (26 to 3.5 blocks/s at a hash the pods' peers=41 say stayed connected) is the measured fact section 5.4 reads.
|
||||
|
||||
### 3.2 Measured propagation and block sizes (the inputs the model takes)
|
||||
|
||||
| Quantity | Value | Source | Label |
|
||||
|---|---|---|---|
|
||||
| Inter-region RTT | 35 ms (hel1-fsn1) to 289 ms (sin-ash); 0.4 to 0.7 inside a location | bench-log, cloud devnet 4 Oct | measured |
|
||||
| Block propagation to 80% of 12 nodes over about 3 hops, 723-byte bodies | p50 343, p90 497, p99 666, max 2,313 ms | same | measured |
|
||||
| One relay hop on 100-ms proxied links (inv, request, block) | p50 318 to 329 ms; own-node processing 6 to 16 ms | fud-ledger M21 run, 5 Oct | measured |
|
||||
| Two hops with 490 KB bodies | p50 641, p99 812, max 857 ms (59-byte bodies: 626 / 1,093 / 2,099) | same | measured |
|
||||
| Live devnet block, 1 coinbase, 22 keys' votes partly | p50 723 B, p90 1,022, max 6,908; coinbase payload p50 395, max 6,580 | fud-ledger M21 sweep | measured |
|
||||
| Vote item, certificate, proof record | 281 B; 273 B + V/8; 274 B | spec 03 3.4.2 (fud-close), fork | cited |
|
||||
| k from the measured delays at 1 bps | p99 0.67 s gives k 5, max 2.3 s gives k 10; k 18 is the 5-s bound | fud-ledger M21, `calculate_ghostdag_k` | cited |
|
||||
| Kaspa 10-bps node requirements | 8 cores, 16 GB RAM, 256 GB SSD, 40 Mbit/s minimum; 12 to 16 cores, 32 GB preferred | `vendor/rusty-kaspa/docs/crescendo-guide.md` | cited |
|
||||
|
||||
### 3.3 Measured node figures
|
||||
|
||||
| Quantity | Value | Source | Label |
|
||||
|---|---|---|---|
|
||||
| PoW cache | 256 MiB per day key, KEEP_DAYS 3, 768 MiB worst, 512 MiB around midnight UTC | bench-log M30 entry | measured |
|
||||
| RSS slope, narrow DAG, 60x fast time, 1 block/s | 30.2 MB per 1,000 blocks (s8 steady, 1,500 blocks) | same | measured |
|
||||
| The live app node on the devnet profile | 1,081 MB at 27 min, 2,258 MB at 4 h 14 min (about 80 KB per block, derived) | same | measured, derived |
|
||||
| Run A's seed | 153 MB before the chain, 5.42 GB at 12,000 blocks (about 440 KB per block; 41 peers, mergeset up to 197) | A.jsonl | measured |
|
||||
| Exec snapshot | 127,564,588 B at tip 141,700 chain blocks (0.9 KB per chain block); `ExecState.records` 1 to 2 KB per chain block | hands-on-build-1.md; M30 note | measured; approximate |
|
||||
| Node 1 and observer data dirs after 3 days | 856 MB and 822 MB | hands-on-build-1.md | measured |
|
||||
| Lottery verify per header | class v3 2.79 ms on a loaded M5 Max core; class v4 4.90 steady, 5.06 cold, 8.23 ms on a half core (box proxy); gate 10 ms | consequences C29; lane 2 section 5.5; spec 01 | measured |
|
||||
| Hub CPU per accepted block at 10 bps | 61 ms (mergeset about 8), 115 ms (about 30), 345 ms (150 to 200) | A.jsonl, section 3.1 | measured |
|
||||
|
||||
### 3.4 What the fleet still owes (read `block-rate-devnet2.md` at 20:1xZ)
|
||||
|
||||
RUN_A, RUN_B, TIER_TABLE and RECOMMENDATION are placeholders. Run B (1 bps, same boxes, fresh genesis, 30 min) starts at 20:25Z by `runB.sh` and its rows land about 21:00Z; the mesh variant A2 is not scripted (`box-dn2.sh` takes one `--addpeer`). The collector's `red` field is 0 in every row (its grep pattern matches no log line), so the fleet's red share must come from blue score against DAA as section 3.1 does; its `--report` payout arithmetic uses 28 MH/s for a 4070 where the brief uses 25. Section 5.2 states this lane's prediction for A2 before its rows land.
|
||||
|
||||
## 4. Model
|
||||
|
||||
Inputs are labelled measured (M), cited (C), simulated (S) or approximate (A).
|
||||
|
||||
1. **k and the parameters per rate** (C, `bps.rs`): k = min k with P(Poisson(2 D lambda) > k) < 0.01 at D = 5 s: 18, 124, 362 at 1, 10, 32 bps; 1,074 at 100 bps (the table stops at 32; `calculate_ghostdag_k` in f64 underflows at x = 1,000, the lane's `cost.py` does it in log space). Parents = clamp(k/2, 10, 16); mergeset limit = clamp(2k, 180, 512); merge depth 3,600 bps; finality 43,200 bps; pruning max(108,000 bps, the Prunality lower bound); maturity 100 bps.
|
||||
2. **Red share from topology and delay** (S, `propagation.py`): blocks arrive as Poisson(bps) split over miners; a block reaches a peer after one hop = 3 lognormal link latencies (median `--link`, sigma 0.29 from the cloud devnet's p50/p90) plus 10 ms processing; a star hub (or every node, `--node-s0`) is a single server with service s0 + s1 x mergeset ms and a FIFO queue; each block's colour is GHOSTDAG's first k-cluster condition over its own past (deterministic given parents); red share = reds in the mergesets of the final selected chain over merged blocks. The abstract rule behind it: reds appear when blocks in flight 2 d lambda exceed k, where d is the effective delay including queueing.
|
||||
3. **Hub queue** (A, M/M/1 reading): utilisation rho = lambda x s; the knee is rho = 1: s = 100 ms at 10 bps, 1 s at 1 bps, 31 ms at 32 bps, 10 ms at 100 bps. Above the knee the wait grows without bound and d becomes the wait, not the link.
|
||||
4. **The controller against a wide DAG** (C for the rule, `controller.py` for the loop): the fork walks the selected chain; a step carries work = blue_work(b) - blue_work(selected parent) (the mergeset blues only) and solvetime = min(clock step, 20 T) (`igneum.rs` CAP_BLOCKS, `difficulty.rs` `igneum_difficulty_bits`). With chain-step spacing sigma = max(1/lambda, d), mergeset m = lambda sigma and blues = min(m, k + 1): estimate / true hash = (blues / m) x (sigma / min(sigma, cap)). Two biases: when sigma > cap the step is capped and the rule over-reads by sigma / cap and hardens to a DAG rate of target x cap / sigma; when the DAG is wide (m > k + 1) and sigma <= cap it under-reads by (k + 1) / m and eases. Kaspa's window (`window.rs` `push_mergeset`: every mergeset block above the blue-score floor, blue and red; `difficulty.rs` `calculate_difficulty_bits`: average target x measured span / expected span) reads the whole-DAG rate over the real span: estimate / true = 1 in both regimes. A blue estimator divided by (1 - observed red share) is the same quantity.
|
||||
5. **Bytes per node per day** (C+S+A, `cost.py` section 3): blocks per day x (header 286 + 32 x parents + body 300 + 274 / 8) + relay overhead (degree x 40 B inv + 40 B request per block, Kaspa's blockrelay flow, A) + votes V x checkpoints per day x 281 + one certificate (273 + V / 8) per block. Checkpoints per day = 2,880 x bps if C1 stays "every 30 blue blocks", 2,880 if it is re-denominated in DAA seconds.
|
||||
6. **CPU budget** (A): the header pipeline validates in order, so per-block cost x bps must stay under about 0.8 core: 800 ms at 1 bps, 80 at 10, 25 at 32, 8 at 100. Cost = lottery verify (M) + BLS per carried vote (A 1.5 ms, unmeasured on this stack, O-10.3) + GHOSTDAG and reachability (O(k x mergeset) store reads; M at the hub only as a total) + exec and record checks.
|
||||
7. **RSS and disk** (M slopes, A steady state): caches 768 MiB worst plus the DAG store's growth per block (30, 80 or 440 KB measured in three settings) over the pruning window; disk per day = blocks x bytes + votes + exec records; archival = the same per year without pruning.
|
||||
8. **Payout and subsidy** (C): subsidy per block = 31.688 IGN / bps at full ramp, 80% to the producer; interval = 1 / (bps x miner hash / network hash).
|
||||
9. **Finality timing** (C, spec 03 C1): checkpoint every 30 blue blocks, determination d = 60 blue (placeholder, 20 on the devnet), lock about 3 s after determination: cadence 30 / bps s, lock (60 / bps + 3) s if C1 and d stay in blue blocks.
|
||||
|
||||
## 5. Results
|
||||
|
||||
### 5.1 Propagation alone: red share and reorg depth per rate and delay (mesh of degree 8, 42 nodes, ideal nodes; `results.md` section D)
|
||||
|
||||
| bps | k | hop delay median ms | red share | mergeset mean / max | tips mean | reorg max / p99 at miner 0 | delay to 90% of nodes, s |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 1 | 18 | 50 / 100 / 300 / 1,000 | 0.0% in every cell | 1.1 / 2 to 2.3 / 7 | 1.1 to 2.4 | 1 / 1 to 2 / 2 | 0.09 to 1.82 |
|
||||
| 10 | 124 | 50 / 100 / 300 / 1,000 | 0.0% in every cell | 1.7 / 6 to 9.2 / 22 | 1.7 to 8.1 | 2 / 1 to 4 / 3 | 0.09 to 1.82 |
|
||||
| 32 | 362 | 50 / 100 / 300 / 1,000 | 0.0% in every cell | 2.9 / 9 to 18.8 / 42 | 2.9 to 14.4 | 3 / 2 to 6 / 5 | 0.09 to 1.82 |
|
||||
| 100 | 1,074 | 50 / 100 / 300 / 1,000 | 0.0% in every cell | 6.0 / 17 to 31.1 / 104 | 5.6 to 25.3 | 4 / 3 to 8 / 8 | 0.09 to 1.82 |
|
||||
|
||||
Reading. Kaspa's k is derived for a 5-s delay bound; at the measured delays (under 1 s per hop, under 2 s to 90 percent of nodes) blocks in flight stay far under k at every rate and no block turns red. The natural reorg depth is the DAG width: 2 to 4 chain blocks at 10 bps, 8 at 100 bps with 1-s hops. Lane 4's simulator (uniform one-way delay to every node, 8 miners) reads p99 11 at 10 bps with d = 0.35 s and 99 with d = 2 s (`ghostdag_results_10bps.md` section 1); the two agree in shape (depth grows with bps x d) and differ in the delay model, so the fleet's measured reorg distribution at 10 bps (section 3.1: 421 reorgs, p50 1, 4 over 40) is the arbiter: outside the start artefact and the saturated phase it sits at 1 to 8, inside this lane's model. Kaspa's published 10-bps red rates are not in the clone (`docs/crescendo-guide.md` carries requirements only); from the k derivation the design red rate is under 1 percent of blocks (delta 0.01 on anticones, approximate), which both simulators reproduce.
|
||||
|
||||
### 5.2 Run A predicted from topology, then from the hub's cost (`results.md` A to C, `results-2.md` E, F, I)
|
||||
|
||||
| Model input | Red share | Hub wait | Hub utilisation | Mergeset mean / max | Reorg max | What it says |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Star, 41 miners, link 40 ms, ideal hub, run A's measured production schedule, join spread 360 s | 0.0% | 0 | 0 | 3.5 / 14 | 4 | topology and latency alone predict no reds at 10 bps |
|
||||
| Star, hub cost 6 + 2 x mergeset ms (a linear GHOSTDAG cost), same schedule | 0.0% | 0.00 s | 0.14 | 3.7 / 15 | 3 | a hub under 20 ms per block keeps up |
|
||||
| Star, hub cost 61 ms per block (M, mergeset 8), same schedule, 1,200 s | 46.8% | 26 s | 0.65 mean (1.6 during the 26 blocks/s burst) | 8.8 / 194 | 43 | the genesis burst alone saturates a 61-ms hub for 5 minutes |
|
||||
| Star, hub cost 115 ms (M, mergeset 30) | 88.1% | 321 s | 1.23 | 20.7 / 199 | 2 | the measured mid-run regime |
|
||||
| Star, hub cost 345 ms (M, wide DAG) | 82.6% | 1,769 s | 3.66 | 9.8 / 45 | 3 | relay in batches, most blocks unmerged at the end |
|
||||
| Measured run A, cumulative | 77.2% | relay batches of 100 blocks | 180% of one core sustained (section 3.1) | 8 to 197 | 55 (start artefact) | |
|
||||
| Star at a constant 10 bps, hub cost 20 / 40 / 60 / 80 ms | 0.0% | 0.00 / 0.01 / 0.05 / 0.18 s | 0.20 / 0.39 / 0.61 / 0.81 | 3.5 to 4.9 | 2 to 4 | under the knee |
|
||||
| the same, 100 / 150 ms | 15.2% / 75.2% | 8.5 / 157 s | 1.01 / 1.51 | 20 / 116 and 27 / 66 | 8 / 3 | the knee is 100 ms per block at 10 bps |
|
||||
| Star at 10 bps, ideal hub, link 100 / 200 / 300 ms | 0.0% | 0 | 0 | 5.8 / 15, 10 / 20, 12.4 / 23 | 3 to 4 | pure latency up to a 2.2-s delay to 90% of nodes gives no reds |
|
||||
| Star with the 26 blocks/s burst for 5 min then 10 bps, hub 6 + 2 m | 0.0% | 0.01 s | 0.33 | 5.4 / 17 | 3 | the burst alone is harmless to a cheap hub |
|
||||
|
||||
So the model predicts run A's red share from the hub's measured per-block CPU and not from its topology: 47 to 88 percent against 77 percent measured, with the waits that the log's relay batches show. The prediction for the mesh variant A2 (`results-2.md` G): a mesh of degree 8 whose nodes each pay 40 or 80 ms per block runs 10 bps at 0 percent red (node wait 0.01 and 0.12 s); at 115 ms per block every node is its own hub and the mesh goes to 52 percent red with 26-s waits and reorgs of 16. The pods run the same binary as the seed on smaller CPUs, so A2 with tonight's genesis bits (a 26 blocks/s burst) is predicted red again; A2 with genesis bits set for 10 bps at 2 GH/s and every node started synced is predicted under 5 percent red if the per-block cost on a pod is under 80 ms, and 50 percent or more if it is 115 ms. The control at 1 bps (`results-2.md` H): 115 or 345 ms per block gives 0 percent red and waits of 0.01 to 0.07 s, which is the live devnet's experience.
|
||||
|
||||
### 5.3 What the hub's 61 to 345 ms per block is made of (the breakdown is not measured; the candidates and their bounds)
|
||||
|
||||
| Component | Per block at 10 bps | Label | Note |
|
||||
|---|---|---|---|
|
||||
| Lottery verify, class v3 (the Devnet 2 override activates v3 at epoch 1) | 2.8 ms | M | one warp per header |
|
||||
| BLS verification of carried votes, up to 48 per block | up to 72 ms at 1.5 ms per verify | A | run A's blocks carried up to 48 of 42 voters' votes; the route drops say the finality path was the hot one |
|
||||
| GHOSTDAG and reachability at mergeset m, k 124 | O(k x m) store reads: 1,000 at m 8, 25,000 at m 200 | C (protocol.rs shape) | the 61 to 345 ms rise tracks m |
|
||||
| Relay to 41 peers (serialise 14 KB x 41, inv handling) | 2.5 MB/s out measured | M | a mesh node of degree 8 does one fifth of it |
|
||||
| Exec of chain blocks, record checks | small: one chain block per 10 s at the widest | M | |
|
||||
|
||||
The gate that settles it is a profile (proposal 5). Whatever the split, the serial budget rule of section 4.6 is the design constraint for every block-rate step: at 10 bps the whole per-block path must stay under 80 ms on the slowest node the network wants to keep, at the mergeset limit 248, not at the narrow-DAG average.
|
||||
|
||||
### 5.4 The controller: what the record shows and what the model says (`controller-10bps.md`, `controller-1bps.md`)
|
||||
|
||||
| Regime (10 bps, k 124, cap 2 s) | Fork's estimator (blue work over the capped chain step) | Whole-DAG estimator (Kaspa's window shape) | Blue estimator corrected by (1 - r) |
|
||||
|---|---|---|---|
|
||||
| d = 0.3 / 1 / 2 s | DAG rate 10.00, red 0%, estimate 1.00x | 10.00, 1.00x | 10.00, 1.00x |
|
||||
| d = 3 s | settles at 6.67 blocks/s, estimate 1.50x, difficulty 1.50x the correct value | 10.00 | 10.00 |
|
||||
| d = 5 s | 4.00 blocks/s, 2.50x | 10.00 | 10.00 |
|
||||
| d = 10 s | 2.00 blocks/s, 5.00x | 10.00 (red 0%, mergeset 100) | 10.00 |
|
||||
| d = 30 s (a saturated hub) | 5.21 blocks/s, red 19.5%, estimate 12.1x, difficulty 1.93x | 10.00, red 58%, mergeset 300 | 10.00 |
|
||||
| 1 bps, k 18, cap 20 s: d = 20 / 30 / 60 s | runs away upward: 179 / 43 / 10 blocks/s, red 97 to 99%, estimate 0.01 to 0.09x (the under-read, main's direction) | 1.00 / 1.00 / 1.17 blocks/s | 1.00 / 1.00 / 1.17 |
|
||||
|
||||
Reading against the record. Run A's production fell from 26 blocks/s to 3.3 to 3.6 while blue stayed 1.4: the DAG hardened. The fork's rule at a chain-step spacing of 5 to 6 s (the hub's queue made chain blocks 10 to 60 s apart at the widest, 3 to 6 s late in the run) settles at target x 2 s / spacing = 3.3 to 4 blocks/s, which is the record's late plateau. The ease direction main reported is the other bias (blue work only) and the model shows it where chain steps stay under the cap while the DAG is wider than k + 1: at 1 bps with d of 20 s or more, or at 10 bps when production exceeds k / (2 d). Either way the rule is reading the wrong quantity: a controller that counts every mergeset block's work over the real span (what Kaspa's `calculate_difficulty_bits` does over its sampled window of blue and red mergeset blocks) holds 10.00 in every cell. A second inconsistency at 10 bps: CAP_BLOCKS 20 is 2 s while FUTURE_TOLERANCE_MS and BACK_TOLERANCE_MS stay 10 s, so a forged stamp (10 s) no longer fits inside half a cap; spec 2.3 derived the 10 s as half the cap at 1 bps. The cap, the tolerances and the 60 T lag bound should be denominated in DAA seconds (20 s, 10 s, 60 s) at every rate, and the step of a chain block that merges m blocks should be allowed m x 20 T before clamping.
|
||||
|
||||
Cost of the bug tonight per tier: a home miner's blocks were 77 percent red (a red inside the DAA window pays its 80% to the merging miner, spec 2.5, so the solo miner lost the subsidy of 3 blocks in 4); a rig the same; a pool user nothing (Devnet 2 is a staging chain); the node operator saw 5.4 GB RSS and 180% CPU on a 96-thread box; a prover saw the exec tip at 154 chain blocks against 11,239 blocks; a holder nothing (no value on Devnet 2).
|
||||
|
||||
### 5.5 Node tiers per block rate (`cost-tables.md` 4 and 5)
|
||||
|
||||
| Rate | Per-block CPU budget (0.8 core) | Class v4 verify share of it (box proxy 4.9 ms; a 2019 laptop core about 2.5x, lane 2's rule: 12 ms) | Votes per block (V/30 if C1 stays in blue blocks) and their BLS cost at V = 1,000 | RSS: caches + DAG store over the 30-h window at 30 / 80 KB per block | Disk per day (blocks, votes at V = 1,000, exec records) | Verdict: laptop 2019-class 8 GB SATA | Raspberry-class 8 GB USB SSD |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 1 bps | 800 ms | 0.6% (laptop 1.5%) | 33 votes, 50 ms | 0.77 + 3.3 to 8.6 GB | 0.92 GB | runs if the DAG store plateaus under about 4 GB (owed: the 30-h measurement); marginal at 80 KB per block | the same question; CPU fine (class v4 verify under 15 ms, GHOSTDAG at mergeset 1 to 2) |
|
||||
| 10 bps | 80 ms | 6% (laptop 15%) | 33 votes, 50 ms: 62% of the budget on its own | 0.77 + 33 to 86 GB | 9.0 GB | no: the hub's measured 61 ms at a narrow DAG already uses 76% of the budget on a server core; RSS over 8 GB within hours unless the per-block footprint falls 10x | no |
|
||||
| 32 bps | 25 ms | 20% (laptop 48%) | 33 votes, 50 ms: over budget | 0.77 + 104 to 276 GB | 29 GB | no | no |
|
||||
| 100 bps | 8 ms | 61% (laptop 150%) | over budget | 0.77 + 326 to 864 GB | 91 GB | no: the verifier gate alone forbids it | no |
|
||||
|
||||
Pruning changes the disk, not the RSS and CPU: the pruned node keeps the 108,000 DAA-s window (136 MB of headers and bodies at 1 bps, 1.55 GB at 10 bps, `cost-tables.md` 6) plus the UTXO set, the exec state (128 MB measured plus 1.5 KB per chain block) and the finality store's 2,000 indices (KEEP_CHECKPOINTS). What must change for 10 bps is the per-block footprint in RAM (30 to 440 KB measured; Kaspa's 16 GB minimum at 10 bps says their footprint is near 10 KB per block, derived from 16 GB over 1,080,000 blocks, approximate) and the per-block CPU (under 50 ms at mergeset 248 on a laptop core).
|
||||
|
||||
### 5.6 Bandwidth per node per day (`cost-tables.md` 2 and 3)
|
||||
|
||||
| Rate | Blocks (header + body + records) | Relay overhead (degree 8) | Votes at V = 12 / 100 / 1,000 / 8,192, C1 in blue blocks | Certificates at V = 12 / 8,192 | Total at V = 100 | Total at V = 8,192 | Mean Mbit/s at 8,192 |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 1 | 57 MB | 31 MB | 10 MB / 81 MB / 809 MB / 6.6 GB | 24 MB / 112 MB | 194 MB | 6.8 GB | 0.63 |
|
||||
| 10 | 649 MB | 311 MB | 97 MB / 809 MB / 8.1 GB / 66 GB | 237 MB / 1.1 GB | 2.0 GB | 68 GB | 6.3 |
|
||||
| 32 | 2.4 GB | 1.0 GB | 311 MB / 2.6 GB / 26 GB / 212 GB | 759 MB / 3.6 GB | 6.8 GB | 219 GB | 20 |
|
||||
| 100 | 9.3 GB | 3.1 GB | 971 MB / 8.1 GB / 81 GB / 663 GB | 2.4 GB / 11 GB | 23 GB | 687 GB | 64 |
|
||||
|
||||
With C1 in DAA seconds the vote column is 81 MB / 809 MB / 6.6 GB per day at every rate; with lane 3's aggregated certificate (1.2 KB per checkpoint) in place of carried votes it is 3.5 MB per day. What carries it: a home connection (10 Mbit/s up, 50 down, approximate) carries 1 bps at any voter count and 10 bps at under 1,000 voters or with C1 in seconds; a Raspberry-class box on ethernet the same, bounded by its CPU not its link; a mobile node is a light client: 3.4 MB per day in checkpoint mode at 1 bps and 1,000 voters, 34 MB at 10 bps if C1 stays in blue blocks (`cost-tables.md` 10). A star hub pays its peer count times the block bytes in upload: run A's seed sent 2.5 MB/s (20 Mbit/s) to 41 peers; the three testnet seeds (cx23, 20 TB per month included, `seed-nodes.md`) would spend 6.5 TB per month each at that rate, inside the allowance and outside good sense; the peer floor of proposal 4 spreads it.
|
||||
|
||||
### 5.7 Votes and the 3.4.2 arithmetic per rate (`cost-tables.md` 2)
|
||||
|
||||
| Rate | Checkpoints per day (C1 in blue blocks) | Votes per block at V = 8,192 (V / 30) | Drain capacity per checkpoint at the cap 48 / 384 | Vote bytes per day at 8,192, single votes | Archival votes per year at 1,000 / 8,192 |
|
||||
|---|---|---|---|---|---|
|
||||
| 1 | 2,880 | 273 | 1,440 / 11,520 | 6.6 GB | 296 GB / 2.4 TB |
|
||||
| 10 | 28,800 | 273 | 1,440 / 11,520 | 66 GB | 3.0 TB / 24 TB |
|
||||
| 32 | 92,160 | 273 | 1,440 / 11,520 | 212 GB | 9.5 TB / 77 TB |
|
||||
| 100 | 288,000 | 273 | 1,440 / 11,520 | 663 GB | 30 TB / 242 TB |
|
||||
|
||||
Because C1 counts blue blocks, the votes per block and the drain capacity per checkpoint are the same at every rate (the 3.4.2 arithmetic holds: 384 per block drains 8,192 in 21.3 blocks), while the bytes per day and the checkpoint cadence scale with bps: 3-s checkpoints at 10 bps, 0.3 s at 100. Tonight at 42 voters and 10 bps the seed's inbound IgneumFinality route (4,096 deep since the fin-route fix) overflowed at 109 to 2,547 drops per peer, and 16 of 47 determinable checkpoints locked; at 1 bps the same voters cost one tenth. Re-denominating C1 and d in DAA seconds (300 blue blocks and 600 at 10 bps) keeps the finality cost, the lock delay (63 s) and the light-client bytes at their 1-bps values across every rate step; it is a one-line spec change and a parameter in the fork.
|
||||
|
||||
### 5.8 Pruning, archival and the snapshot path
|
||||
|
||||
What a pruned node keeps (spec 02 2.1, `cost-tables.md` 6): headers and bodies back to the pruning point (108,000 DAA s, never past the latest certified checkpoint, F3), with their GHOSTDAG and reachability data (the 2x index factor is approximate); the pruning proof (levels of headers, `pruning_proof`); the UTXO set; the finality store's last 2,000 indices; the exec state (`exec-snapshot.bin` 128 MB measured at chain block 141,700, growing 0.9 KB per chain block, plus `ExecState.records` 1.5 KB per chain block until a window bounds it, M30 note); the proof-record window (`RECORD_WINDOW_CHAIN_BLOCKS`). An archival node keeps every block and its finality section: 40 GB per year of blocks at 1 bps (453 GB at 10 bps) plus votes as table 5.7, so the archival cost is the votes, not the blocks, until certificates replace carried votes (1.3 GB per year at 1 bps).
|
||||
|
||||
The p2p snapshot path as shipped in 0.3.14 (`proving.rs` `on_exec_snapshot`, `p2p_snapshot_gate`, read on `exec-sync-0313`): a peer's `ExecSnapshot` (version, chain id, genesis, tip number and hash, state root, records, the account and storage dump, fees, epoch, paid shards) is accepted only when its tip is a chain block known to this node's consensus, its tip is not 0, not below this node's exec restart, above this node's own executed tip, and only while the executor is blocked or has no state; and, when the restart pin is configured, the snapshot's record at the restart block must carry the pinned root (`exec_restart_state_root`, the fourteenth override field). The file is then written as the node's own resume point and served onward. Three things it does not check: the snapshot's state root at its tip against anything the chain commits to; the records between the restart block and the tip; agreement among peers. The poisoning attack: a peer of a blocked node (every node is blocked after a reorg deeper than its ring or after a restart below its retention root, the 6 October incident class) serves a snapshot whose tip is a real chain block above the victim's tip, whose record at the restart block is correct, and whose state at the tip is wrong. The victim loads it, executes forward from a false state, persists it, serves it to its own peers, and from then on refuses every honest proof record (`check_record`: the statement over its own records disagrees) while its own prover's records are refused by the network. Bound: it is a liveness attack on the victim and its downstream peers, not a consensus or a proof break (full nodes execute natively and a proof over the false state is a proof of the wrong statement; a light client verifies proofs against the chain's records, not against a node's state), and it needs the attacker among the victim's peers at the moment it is blocked, which the peer floor makes a 1-in-(peer count) race unless the attacker runs most of the victim's peers (an eclipse). The hardening: accept a snapshot only when its state root at the newest proven segment at or below its tip equals the `post_root` of the proof record the chain carries for that segment (the chain already carries 274-byte records in coinbase payloads, so this is a lookup, not a protocol change), and take the (tip, root) pair from N of M peers (the light client's 3 of 5, spec 10.6) before loading; a snapshot that fails either is refused and the peer is dropped for the session. Then poisoning needs a false proof record in the chain, which needs the proving key and a block that carries it, and the bound is the proof system's. Sizes per tier: the snapshot is the state (128 MB today, growing with accounts), so a home node's recovery is a 128 MB download and a sha256; an archival node's history is the table above; a light client never holds one.
|
||||
|
||||
### 5.9 Finality through this lane's eyes (cross-reference to lane 3)
|
||||
|
||||
Lane 3 (`finality-and-weight.md` 1 and 5.5) refutes the topology hypothesis for tonight's pause on the live devnet: certificates 6824 to 6842 formed while node 1 and the observer were down, the zero-aggregator fallback carried a quarter of the certificates, and the pause began at the checkpoint where signing weight fell to 53.1 percent of the frozen table, which is the two-thirds rule; the fleet's star is around the seed (the finality route entry: "on a Vast box the seed is the only peer"), not the Mac. This lane adds the measured shape of that star under load: 41 pods each with peers=1, every block and every vote through one process at 180 percent of one core, 2.5 MB/s of upload, the IgneumFinality route dropping thousands of messages per peer, 16 of 47 checkpoints locked. Lane 3's hub-cut harness case (its proposal 6) and this lane's peer floor (proposal 4) are one piece of work: the floor is what makes the harness case pass (a node with 4 outbound peers keeps blocks and votes when any one peer dies), and the fleet library is where both land.
|
||||
|
||||
### 5.10 Block rate: payout intervals, subsidy, lock delay and the solo-miner answer (`cost-tables.md` 7 to 9)
|
||||
|
||||
| Rate | Subsidy per block (full ramp), producer's 80% | 4070 (25 MH/s) at 1.16 GH/s / 100 GH/s / 1 TH/s / 10 TH/s | 5090 (128 MH/s) at the same | 8x 4090 rig (459 MH/s) at the same | Lock after a checkpoint if d stays 60 blue |
|
||||
|---|---|---|---|---|---|
|
||||
| 1 | 31.69 IGN, 25.35 | 46 s / 1.1 h / 11.1 h / 4.6 d | 9.1 s / 13 min / 2.2 h / 21.7 h | 2.5 s / 3.6 min / 36 min / 6.1 h | 63 s |
|
||||
| 10 | 3.17, 2.54 | 4.6 s / 6.7 min / 1.1 h / 11.1 h | 0.9 s / 1.3 min / 13 min / 2.2 h | 0.3 s / 22 s / 3.6 min / 36 min | 9 s |
|
||||
| 32 | 0.99, 0.79 | 1.4 s / 2.1 min / 21 min / 3.5 h | 0.3 s / 24 s / 4.1 min / 41 min | 0.1 s / 6.8 s / 1.1 min / 11 min | 4.9 s |
|
||||
| 100 | 0.32, 0.25 | 0.5 s / 40 s / 6.7 min / 1.1 h | 0.1 s / 7.8 s / 1.3 min / 13 min | 0.0 s / 2.2 s / 22 s / 3.6 min | 3.6 s |
|
||||
|
||||
The solo-miner question: a 4070 at 25 MH/s sees 21.6 blocks a day at 1 bps on a 100 GH/s network (548 IGN a day at full ramp) and 2.16 a day at 1 TH/s (55 IGN); one block a day needs 0.046 bps at 100 GH/s, 0.46 bps at 1 TH/s and 4.6 bps at 10 TH/s. Kaspa's 10 bps buys the solo miner a daily block up to about 20 TH/s; Igneum at 1 bps gives it up to about 2 TH/s, which is 1,700x tonight's hash and USD 562,000 a day of rented hash at the bench entry's USD 281 per GH/s-day. What the rate costs the node, from 5.5 and 5.6: 10x the bytes (2 GB a day at 100 voters), a per-block CPU budget of 80 ms that the measured 61 to 345 ms does not meet, an RSS that leaves 8 GB within hours at the measured footprints, 3-s checkpoints and 66 GB a day of votes at 8,192 voters unless C1 moves to seconds, and the controller's cap inconsistency. Per tier: a home miner on one card gains variance relief it does not yet need (a 4070 at 1 bps sees a block an hour at 100 GH/s); a rig gains nothing (36 min per block at 1 TH/s already); a pool user nothing (pools remove variance); a prover gains nothing and pays the exec follower's 10x chain blocks; a holder gains faster locks (9 s) only if d is left in blue blocks, which costs the finality bytes of 5.7; a rollup customer the same lock question; the node operator pays all of it.
|
||||
|
||||
## 6. Ranked proposals
|
||||
|
||||
| Rank | Proposal | Evidence | Model | Hours | Consequence per tier | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 1 | Block rate: 1 bps for the public testnet and for mainnet launch; 10 bps stays a planned step (spec 2.1) behind three gates: per-block node CPU under 50 ms at mergeset 248 on a 2019 laptop core, C1 / d / the clock cap in DAA seconds, vote aggregation; no 32 or 100 bps step is proposed | 5.2 (reds are node throughput), 5.5 (budget 80 ms against 61 to 345 measured), 5.7 (66 GB a day of votes), 5.10 (1 bps already gives a 4070 a block an hour at 100 GH/s) | `propagation.py` F and G, `cost.py` 4 and 7 | 1 (the spec line) plus the gates below | home miner: no node change, a block an hour at 100 GH/s; rig and pool: unchanged; prover: the exec follower stays at 1 chain block per second; holder and rollup: 63-s locks; node operator: today's node | the three gates pass on Devnet 2 at 10 bps with zero red over the knee, before any rate step |
|
||||
| 2 | Controller: every lane counts every mergeset block's work (blue and red) over the real span; the clock cap, tolerances and lag bound in DAA seconds (20, 10, 60 s) at every rate; a chain step that merges m blocks may span m x 20 T before clamping | 3.1 (26 to 3.5 blocks/s with blue at 1.4), 5.4 (the fork over-reads by spacing / cap and under-reads by (k + 1) / m; Kaspa's window reads the whole DAG) | `controller.py`: 10.00 in every cell for the whole-DAG estimator | 8 (rule and unit tests in `difficulty.rs`, `igneum.rs`) + 3 (the harness case: fast-time 3 nodes, a 3-s proxy delay at 10 bps; pass = DAG rate within 10% of target and difficulty within 10% of the hash-implied value after 10 min, the same at 1 bps with a 30-s delay) | home miner: the subsidy stops tracking the controller's error (tonight 77% of blocks unpaid to their miner); node operator: no runaway DAG from a slow hub; holder: emission on schedule (spec 2.5's "when the controller lets the rate run" clause stops firing) | the harness case passes; `sim/difficulty/sim.py --dag-delay` replay of run A's production curve settles at 10 bps |
|
||||
| 3 | Finality interval and determination in DAA seconds (C1: checkpoint every 30 DAA s; d in DAA s), one carriage per vote (a block carries a vote only if no block in its past does, which is the rule as written; the relay dedups by (key, index)), and lane 3's aggregated certificate as the archival object | 5.7 (votes and route load scale with bps under the blue-block rule), 3.1 (2,547 route drops per peer, 16 of 47 locks) | `cost.py` 2 and 10 | 4 (spec 03 C1, 3.10 table) + 8 (fork: `finality.rs` interval in DAA s, relay dedup) | home miner and rig: finality bytes stay at 81 MB a day at 100 voters whatever the rate; node operator: the route stays inside 4,096; light client: 3.4 MB a day at every rate; holder: 63-s locks at every rate | a Devnet 2 run at 10 bps with 42 voters shows 0 route drops and every determinable checkpoint locked |
|
||||
| 4 | Peer floor and topology: a node reports synced and mines only with at least 4 outbound peers (8 on seeds); the fleet library gives every box the seed plus 2 random boxes; the seed list carries 3 seeds; the hub-cut harness case of lane 3's proposal 6 with the floor as its pass condition | 3.1 (peers=1 on every pod, 2.5 MB/s out of one process), lane 3 section 5.5 ("on a Vast box the seed is the only peer") | `propagation.py` B (mesh of degree 4 and 8 against the star: the same red share at an ideal hub, one fifth of the upload per node, no single queue) | 4 (fleet library and the synced rule in `tools/fleet/lib/`) + 3 (the harness case) | node operator and seed: upload spread 41 ways to 8 ways; home miner: a dead seed no longer stops blocks and votes; finality: votes reach aggregators by two paths | the hub-cut case passes: a lock within 2 checkpoints of the cut, 0 conflicting certificates at the hub's return, every box at 4 or more peers throughout |
|
||||
| 5 | Measure the per-block node cost and its split (lottery verify, BLS per vote, GHOSTDAG and reachability, relay, exec) with `perf` on a Devnet 2 box and on a 2019-class laptop at 1 and 10 bps, at a narrow DAG and at mergeset 248; set the budget rule (cost x bps under 0.8 core) as a CI check on the fast-time network | 3.1 and 5.3 (61 to 345 ms measured as a total only) | section 4.6 | 3 (profile) + 4 (the CI check in `tools/ci`) | every node tier: a number per machine class instead of the hub's; the 8 GB laptop and the Raspberry-class verdicts of 5.5 become measurements | the profile lands in the bench log with per-component ms; the check fails a build whose per-block cost exceeds the budget at the network's rate |
|
||||
| 6 | Snapshot path hardening: a p2p snapshot is accepted only when its root at the newest proven segment at or below its tip equals the chain's proof record for that segment, and (tip, root) agree on 3 of 5 peers; a failing peer is dropped for the session; the gate's unit test gains the two cases | 5.8 (the gate checks tips and the restart pin, never the state at the tip; the 6 October seed loaded a tip-0 snapshot from a fleet box before the gate) | section 5.8 | 6 (fork `proving.rs`, `snapshot.rs`; the record lookup exists in `check_record`) | node operator: recovery from a peer cannot be poisoned below an eclipse; prover: no false-state records from a poisoned node; home miner: the app's node recovers by itself after a deep reorg | `tools/exec-sync/reorg.mjs` gains a poisoned-snapshot case: refused, peer dropped, honest snapshot loaded after it |
|
||||
| 7 | Devnet 2 genesis bits set for the fleet's hash at the run's rate (no 26 blocks/s burst), and the fleet rule: a box mines only after synced with the peer floor (run A's pods mined from genesis for up to 6 minutes before they saw the chain) | 3.1 (the 55-block reorg "unwinding to height 1", the 33% red first phase), 5.2 (the burst alone saturates a 61-ms hub for 5 minutes) | `propagation.py` C and E | 2 (`box-dn2.sh`, `start-seed.sh`, `devnet2-override.json` per rate) | fleet operator: run A2 and run B measure the rate, not the start; every other tier: none | the next Devnet 2 run shows no reorg over 5 in its first 10 minutes and a first-minute production within 2x of target |
|
||||
| 8 | RSS steady state: run a node through a full 30-h pruning window at 1 bps (fast time) and at 10 bps on Devnet 2, read RSS every 10 minutes, bound the DAG store (reachability and GHOSTDAG per block) and `ExecState.records`; decide the 8 GB tier from the plateau | 3.3 (slopes 30, 80, 440 KB per block over short runs; none is a steady state), 5.5 (8 GB is marginal at 1 bps at 80 KB per block) | `cost.py` 5 | 3 (the run) + 8 (the bounds, if the plateau is over 4 GB) | home miner on an 8 GB laptop: a yes or no at 1 bps; Raspberry-class: the same; node operator: a RAM line on the requirements page | RSS plateau under 4 GB at 1 bps over 30 h; the requirements page carries the measured line |
|
||||
| 9 | Node and client tiers published with sizes: full pruned node (1 bps: 0.9 GB a day, 136 MB window plus state, 4 GB RAM target), archival node (40 GB a year of blocks plus 1.3 GB of certificates once aggregated; 296 GB a year of votes until then at 1,000 voters), light client checkpoint mode (3.4 MB a day), phase two (2.3 MB a day) | 5.6, 5.8, `cost-tables.md` 6 and 10 | `cost.py` | 2 (the requirements page and spec 10.5's table at 1 bps, with the rate rows) | every tier knows its cost before the testnet | the page's numbers match `cost-tables.md` and the first testnet week's measured bytes within 30% |
|
||||
| 10 | Bandwidth and route bounds per peer: at most 2 x V / 30 votes per peer per block interval accepted (a second belt behind the dedup), block relay to at most 16 peers per node, the IgneumFinality route sized from V and the rate | 3.1 (2,547 drops per peer), 5.6 | `cost.py` 3 | 4 (fork p2p flows) | seed and node operator: a bound on what one peer can make a node do; home miner: none | the s7 flood scenario gains a vote flood: 0 disconnects, bounded CPU |
|
||||
|
||||
**1. Block rate.** Everything measured tonight says 1 bps is the rate the current node can run and 10 bps is not: the hub spent 61 ms per block at a narrow DAG against a budget of 100, and the model turns that cost into tonight's red share (5.2). Nothing in propagation forbids 10 bps (5.1: zero reds at every measured delay with k 124), so the step stays on the plan (spec 2.1) as Kaspa took it, after a test campaign, with three gates: the per-block cost (proposal 5), the time denomination of finality and the controller (proposals 2 and 3), and vote aggregation (lane 3). The variance argument does not need the step yet: a 4070 sees a block an hour at 100 GH/s and two a day at 1 TH/s at 1 bps (5.10); a pool user never sees variance. Consequences: no node tier changes for the testnet; the testnet gives the measured bytes and RSS that fill proposals 8 and 9.
|
||||
|
||||
**2. Controller.** The rule walks the selected chain and sums blue-work increments over capped clock steps (`igneum_difficulty_bits`, section 4.4). In a wide DAG both parts mislead it: the cap turns a 5-s chain step into 2 s (over-read, harden, the record's fall to 3.5 blocks/s) and the blue-only work hides the reds (under-read, ease, the direction main reported). Kaspa's window pushes every mergeset block, red included, and divides by the real span (`window.rs`, `difficulty.rs`), which the model holds at target in every cell. The change is inside `igneum_difficulty_bits` and `igneum_target`: work = the mergeset's total work per chain step, the step allowed m x 20 T before the cap, the cap and tolerances in DAA seconds. The harness case is the proxy-delayed fast-time network of `sim/difficulty/attacks/README.md` at 10 bps with a 3-s hold, and `sim.py --dag-delay` replaying run A's production curve (the schedule file is in `sim/horizon/network/`).
|
||||
|
||||
**3. Finality interval in seconds.** C1 says 30 blue blocks, d says 60 blue blocks; at 10 bps that is a checkpoint every 3 s and ten times the votes, certificates, route load and light-client bytes per day (5.7), and tonight's seed showed the route overflowing at 42 voters. Written in DAA seconds the whole finality cost is rate-invariant and 3.4.2's arithmetic (384 votes per block drain 8,192 in 21 blocks) gets 10x more headroom at 10 bps. The one-carriage rule is already the spec's text ("not already in its past"); the relay dedup by (key, index) makes the wire match it.
|
||||
|
||||
**4. Peer floor.** Run A's pods had one peer each; the live fleet's boxes have the seed as their only peer; the hub paid 41x the upload and ran one queue for every block and vote. Four outbound peers before a node calls itself synced, the seed plus two random boxes in the fleet library, three seeds in the list: the mesh runs of 5.2 show the same zero red at an ideal hub with one fifth of the per-node upload and no single queue, and lane 3's hub-cut case becomes passable. This is the proposal that serves both lanes.
|
||||
|
||||
**5. The per-block profile.** The 61 to 345 ms is a total; the split decides which fix buys 10 bps: if BLS of carried votes dominates, aggregation and the dedup buy it; if GHOSTDAG at k 124 dominates, the mergeset limit and a cheaper reachability do; if relay dominates, the peer floor does. Three hours with `perf` on a Devnet 2 box, then a CI check that fails a build whose cost exceeds the budget at the network's rate, so the class is closed the way CLAUDE.md asks.
|
||||
|
||||
**6. Snapshot hardening.** The gate refuses tips, not states (5.8). The chain carries the proof records' roots, so a snapshot's root at its newest proven segment can be checked against them without a protocol change, and 3-of-5 peer agreement makes poisoning an eclipse. The attack is a liveness attack on the victim, never a consensus break, but a poisoned seed serving its state onward is the fleet-wide class the 6 October gate was written for.
|
||||
|
||||
**7 to 10.** Devnet 2's genesis bits and the start-synced rule make the next runs measure the rate instead of the start (the 55-block reorg and the first-phase reds were the start); the 30-h RSS run decides the 8 GB tier with a measurement instead of three slopes; the tier page and the per-peer bounds are two hours each and close open rows in spec 10.5 and the finality route entry.
|
||||
|
||||
## 7. Open questions and what could not be run
|
||||
|
||||
| Question | Why not tonight | What closes it |
|
||||
|---|---|---|
|
||||
| Run B (1 bps control) and A2 (mesh) rows | B starts at 20:25Z, rows about 21:00Z; A2 is not scripted | the fleet's RUN_B row; A2 built on `box-dn2.sh` with 3 `--addpeer` entries (seed plus two boxes) and genesis bits for 10 bps; this lane's prediction for A2 is in 5.2 |
|
||||
| The difficulty trajectory of run A | the collector strips `difficulty=`; the node log has no bits | `bps-collect.py` keeps the field; or `getBlockDagInfo` difficulty per minute on the next run |
|
||||
| The per-block CPU split | no profile; the hub's figure is a total that includes relay to 41 peers | proposal 5 |
|
||||
| RSS steady state at 1 and 10 bps | three slopes over 1,500 to 15,000 blocks, none over a pruning window | proposal 8 |
|
||||
| Kaspa's measured 10-bps red rate | not in the clone (`crescendo-guide.md` has requirements only) | the kaspanet KIP-14 text and the Crescendo testnet-10 reports, not cloned; approximate under 1 percent by the k derivation |
|
||||
| BLS verify per vote on this stack | O-10.3 open; 1.5 ms is approximate | `fast_aggregate_verify` timing with the forked `blst` build |
|
||||
| The pods' per-block cost (the A2 prediction's fork) | the pods' CPUs and their node logs were not read (fleet boxes are out of this lane's reach) | the fleet's collector reads `cpu_rss` on every box, not only the seed |
|
||||
| GHOSTDAG's second k-cluster condition | `propagation.py` applies the candidate's own anticone bound only; lane 4's `ghostdag_sim.py` carries both | re-run the grid with lane 4's colouring if a cell is contested (every cell here is 0 percent, so the undercount cannot change the reading) |
|
||||
| The 55-block reorg's cause | read from the log as a late-joining pod's chain from genesis ("unwinding to height 1") | the start-synced rule of proposal 7 removes the class; the next run's reorg distribution confirms |
|
||||
|
||||
## 8. Summary for the coordinator
|
||||
|
||||
Run A at 10 bps did not measure propagation; it measured a node that needs 61 to 345 ms per block through one hub whose budget at that rate is 100 ms. The propagation model with the measured links predicts under 0.1 percent red in the star and in the mesh; with the hub's measured per-block CPU it predicts 47 to 88 percent against the 77.2 percent measured, with the queueing waits the log's 100-block relay batches show, and it predicts that the mesh variant A2 goes red too unless every pod's per-block cost is under 80 ms and the genesis burst is removed. The controller then hardened (production 26 to 3.5 blocks/s with blue at 1.4) because its estimator reads blue work over a chain step capped at 2 s; the whole-DAG estimator Kaspa uses is unbiased in every regime the model covers. The block rate for the public testnet and mainnet launch is 1 bps; 10 bps stays a gated step, and the gates are a per-block cost under 50 ms on a laptop core at mergeset 248, the finality interval and the clock cap in DAA seconds, and vote aggregation. Lane 3's finding stands (the pause was the rule, the fleet's star is around the seed); the peer floor is the piece of work both lanes need.
|
||||
|
||||
1. Reds from node cost, not topology: red share 0.0 percent in every cell of the bps x delay grid (1 to 100 bps, 50 to 1,000 ms hops, k from Kaspa's table) and in the run A replay with an ideal hub; 47 / 88 / 83 percent with the hub at 61 / 115 / 345 ms per block (measured); the knee at 10 bps is 100 ms per block (`sim/horizon/network/propagation.py`, `results.md`, `results-2.md`).
|
||||
2. The controller's two biases: estimate / true hash = (blues / m) x (spacing / cap); at a 5-s chain-step spacing the fork settles the DAG at 4.0 blocks/s against 10 (record 3.3 to 3.6), at a 20-s spacing at 1 bps it runs away to 179 blocks/s; the whole-DAG estimator holds 10.00 and 1.00 in every cell (`controller.py`).
|
||||
3. Finality cost scales with the rate under C1 as written: 8,192 voters cost 6.6 GB a day per node at 1 bps and 66 GB at 10 bps (24 TB a year archival); in DAA seconds 6.6 GB at every rate, 3.5 MB a day with aggregated certificates; tonight's seed dropped up to 2,547 finality messages per peer and locked 16 of 47 checkpoints at 42 voters (`cost.py`, section 3.1).
|
||||
|
|
@ -10,9 +10,10 @@ the project lead's mandate, verbatim: "if we create a new way of hashing or a ne
|
|||
|---|---|
|
||||
| 19:05 | Lane started. Read: preamble, CLAUDE.md, the two personas, spec 01 (whole), 04, 07, chip-model-v3 (whole), asic-resistance-history (sections 0 to 3 and 4.3, 5), latency-shadow-2026-10-06 (whole), counter-asic-3-status sections 1 to 5, counter-asic-3-node section 6 (the P2 signalling rule), int8-matrix-family sections 1 to 3, scratch-soundness verdict, proving-methods (whole), fud-ledger M1, M7, M16, M22, M28, P2, F13, proto-cuda host.cu and the mx8-genesis pack (kernel.cu, memhard.h, program.h, vectors.h), proto-cuda/emu, family-probe.cu, the fleet's prover-tiers-real-cards.md, bench-log line 2582 (rental cost) |
|
||||
| 19:25 | Fleet agent asked for two boxes; answered at 19:29: two quiet RTX 4090s (RunPod, nvcc 12.8 at /usr/local/cuda/bin, directory /root/horizon-newpow, until 22:30Z). No quiet Ampere card exists tonight; a loaded 3090 is offered. Main's note: cost rows use bench-log 2582 (USD 0.0117 per MH/s-hour) |
|
||||
| 19:35 | File skeleton written. Two prototype sub-agents launched (budget two at once): `mma-shadow` on box 1 (47.47.180.77), `state-dataset` on box 2 (213.173.98.36) plus CPU rows on igneum-build-1. Designs being written in this file meanwhile |
|
||||
| 19:35 | File skeleton written. Two prototype sub-agents launched (budget two at once): `mma-shadow` on box 1 (<box-1-ip>), `state-dataset` on box 2 (<box-2-ip>) plus CPU rows on igneum-build-1. Designs being written in this file meanwhile |
|
||||
| 19:41 | Section 3 complete: the three designs, the one-table comparison, the migration path. Scheme A's verdict is already visible in its own numbers (A1 dead on 2.9 MB of openings per block, A2 dead on sampleability and a 32 to 40 ms proof verify; A0 is scheme C with the trace as state). Prototypes running: `mma-shadow` (box 1) and `state-dataset` (box 2 and igneum-build-1). mm8 two-output correction sent to the prototype |
|
||||
| (next) | Section 4 reviews (cryptographer, consensus engineer, two per scheme); the pick; section 5 measured rows as they land; section 6 verdicts; section 7 ranked next steps |
|
||||
| 19:50 | Section 4 (reviews, the pick: B and C), section 7 (ranked next steps) and section 8 (open questions) written. Scheme C prototype complete and measured on box 2 and igneum-build-1: hash rate and watts unchanged (63.08 vs 63.09 MH/s, 207 W), build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact 1,024 of 1,024 items and 128 of 128 lanes; rows in 5.2. Scheme B ladder running on box 1 (R = 0, 8, 32, 128 measured, 512 in progress) |
|
||||
| 20:20 | Scheme B ladder complete on box 1 (R = 0 to 512, rate flat at 63.08 MH/s, 201 to 216 W, every fingerprint PTX = reference, 1,024 of 1,024 lanes at every R, verifier delta 0.05 to 4.39 ms). Sections 5.1, 5.3, 6 and 9 written. Box 2 released 19:53Z; box 1 released on the mma agent's report. File complete |
|
||||
|
||||
## 1. What was read and the facts this lane stands on
|
||||
|
||||
|
|
@ -133,8 +134,8 @@ So the fleet splits by generation: Turing-and-later NVIDIA and RDNA 3-and-later
|
|||
| Scheme | What it is | Chip edge per joule vs the 5090 (f = 1 chip, chip-model-v3 method) | Verifier ms per unit (model, then measured in section 5) | Vendor bit-exactness | Ships as class v5? |
|
||||
|---|---|---|---|---|---|
|
||||
| A, mining is proving | A1 committed codeword, A2 proving steps: dead on bytes and on sampleability; A0 trace-as-dataset survives and is C with the trace as state | A0 unchanged (5.1x GDDR7); A1, A2 not priced | A0: 2.06 + the leaf-read row; A1: 2.9 MB of openings; A2: 32 to 40 ms | A0 adds the zkVM executor to consensus | NEVER as mining = proving; A0 folds into C |
|
||||
| B, tensor-shaped shadow | class v3 plus `8 R` int8 8x8x16 tile steps per hash in the PTX fragment layout, two outputs per lane | memory + `8 R x 1,024 x e_mac x k_mma`; `k_mma` near 1 with a floor near 0.5 (approximate); rows from the measured `e_mac` in 5.3 | 2.06 + about 0.9 microseconds per step per unit scalar (R = 128: +0.9 ms; R = 512: +3.7 ms); VNNI 4x to 16x less | native on sm_75+ and RDNA 3+ (AMD layout unverified); emulated on Pascal, RDNA 2, Apple | prototype further; a class v5 candidate after the AMD gate |
|
||||
| C, stored state | the daily dataset derives from the execution state snapshot; kernel unchanged; a state sample per block | unchanged (5.1x GDDR7, 7.5x to 9.2x HBM3); the f = 0 chip disappears | 2.06 + 4,096 leaf reads (section 5.2) | nothing vendor-specific; the serialisation is the consensus risk | prototype further; a class v5 candidate on its own or beside B |
|
||||
| B, tensor-shaped shadow | class v3 plus `8 R` int8 8x8x16 tile steps per hash in the PTX fragment layout, two outputs per lane | memory + `8 R x 1,024 x e_mac x k_mma`; `k_mma` near 1 with a floor near 0.5 (approximate); rows from the measured `e_mac` in 5.3 | 2.06 + about 0.9 microseconds per step per unit scalar (R = 128: +0.9 ms; R = 512: +3.7 ms); VNNI 4x to 16x less | native on sm_75+ and RDNA 3+ (AMD layout unverified); emulated on Pascal, RDNA 2, Apple | before measurement: prototype further. After section 5: NEVER as class content (the block costs the honest card 0.056 to 0.70 pJ per multiply-add, so it forces no joules on a chip); kept as reserve R8 evidence |
|
||||
| C, stored state | the daily dataset derives from the execution state snapshot; kernel unchanged; a state sample per block | unchanged (5.1x GDDR7, 7.5x to 9.2x HBM3); the f = 0 chip disappears | 2.06 + 4,096 leaf reads (section 5.2) | nothing vendor-specific; the serialisation is the consensus risk | before measurement: prototype further. After section 5: SHIP as the class v5 candidate |
|
||||
|
||||
### 3.5 Migration through the class system (common to B and C)
|
||||
|
||||
|
|
@ -165,7 +166,7 @@ Written by this lane in the persona files' voices (`.claude/agents/cryptographer
|
|||
|
||||
### 4.4 The pick
|
||||
|
||||
B and C are the two to prototype: both are class objects on the shipped hash, both leave the dataset's memory bound untouched, and they compose (B is kernel text, C is the daily build). A is not prototyped: A1 and A2 fail on bytes and on sampleability before any kernel, and A0 is C with a worse data source. The order of merit at this point, before measurement: C first (no vendor risk, unchanged hash rate by construction, a real new property per block, the pool caveat stated), B second (a real lever against the f = 1 chip with a bounded `k`, a vendor split and an unverified AMD layout). Section 6 revisits the order on the measured rows.
|
||||
B and C are the two to prototype: both are class objects on the shipped hash, both leave the dataset's memory bound untouched, and they compose (B is kernel text, C is the daily build). A is not prototyped: A1 and A2 fail on bytes and on sampleability before any kernel, and A0 is C with a worse data source. The order of merit at this point, before measurement: C first (no vendor risk, unchanged hash rate by construction, a real new property per block, the pool caveat stated), B second (a real lever against the f = 1 chip with a bounded `k`, a vendor split and an unverified AMD layout). Section 6 revisits the order on the measured rows: C holds, B retires into the reserve.
|
||||
|
||||
## 5. The prototypes and the measured rows
|
||||
|
||||
|
|
@ -173,31 +174,119 @@ Both prototypes live under `proto-newpow/` with a README carrying the exact comm
|
|||
|
||||
### 5.1 `mma-shadow` (scheme B), box 1
|
||||
|
||||
ROWS_B
|
||||
Prototype: `proto-newpow/mma-shadow/` (README with every command, `kernel_mm8.cu` with the PTX path and the `IGNEUM_MM8_REF` reference path, `bench.cu`, `verify_ref.c`, `gen_block.py`, `gen_ref_program.py`, `run.sh`, `out/` with every log and csv). Box 1: RTX 4090 24 GB (128 SMs), driver 570.172 (the box reports 570, not the 595 of box 2), nvcc 12.8.93, `-arch=sm_89`, 1 warp per block, 10 timed batches of 2^24 after a warm-up, power from `nvidia-smi` at 1 Hz over a 25-s sustained phase (mean after its first 10 s), idle 15.0 to 15.3 W. The two-output tile form of section 3.2 was built (the single-output form never was). Run 19:36 to 19:52 UTC.
|
||||
|
||||
| R (steps per iteration) | mm8 per hash (8 R) | MH/s (GPU time) | Ratio to R = 0 | Watts, mean | Block watts over R = 0 | SM MHz | Microjoules per hash | Marginal pJ per multiply-add | Fingerprint (2^24 at base 0) | PTX = reference path | CPU = GPU (1,024 lanes) | Verifier block delta, ms per unit (box core, scalar C) | Registers per thread |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| 0 (control) | 0 | 63.08 | 1.000 | 201.2 | 0 | 2,670 | 3.19 | | 7c28cfb06c5c65a9 (the pack's; the 3 pack vectors PASS standalone and in batch) | yes | 1,024 of 1,024 | 0 (the R = 0 run read -0.10, noise) | 29 |
|
||||
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2.9 | 2,670 | 3.24 | 0.70 | 06fc2593bfb94b4f | yes | 1,024 of 1,024 | +0.05 | 30 |
|
||||
| 32 | 256 | 63.08 | 1.000 | 207.7 | 6.5 | 2,670 | 3.29 | 0.39 | 26e83a65f519c865 | yes | 1,024 of 1,024 | +0.17 | 29 |
|
||||
| 128 | 1,024 | 63.08 | 1.000 | 212.7 | 11.5 | 2,670 | 3.37 | 0.17 | 42223c2113188335 | yes | 1,024 of 1,024 | +1.14 | 32 |
|
||||
| 512 | 4,096 | 63.08 | 1.000 | 215.9 | 14.7 | 2,670 | 3.42 | 0.056 | 02b7002d747f3711 | yes | 1,024 of 1,024 | +4.39 | 36 |
|
||||
|
||||
Sustained rates over the power window 63.07 to 63.03 MH/s at every R; wall and event rates agree to 0.01 MH/s; no spills, no stack, temperature 58 to 64 C, no throttling. The verifier totals in the box's logs (about 10 ms per unit) are a naive scalar interpreter with a lazy `mh_word` per load on a shared Zen 4 core at load 13 and are not comparable to the project's Rust verifier (2.06 ms per unit on an M5 Max core, interleaved); the block delta is the number: about 1.1 microseconds per `mm8` step per unit, 32 lanes x 32 byte products in plain loops. (At R = 512 the card issues about 258 G tiles per second, 2.6 x 10^14 multiply-adds per second, in the region of 40 percent of the 4090's dense int8 tensor peak, approximate from the published TOPS figure, so R = 512 is near the top of the free band, not its middle.)
|
||||
|
||||
What the rows say:
|
||||
|
||||
1. The tensor block is free in hash rate to R = 512 on the 4090: 4,096 tile instructions per hash leave the rate at 63.08 MH/s to the third decimal. The kernel is latency-bound on its 128 dependent loads and the tensor work fills stalls that were already there, as the ALU shadow did on the 5090 to 150,800 ops (`latency-shadow-2026-10-06.md` 5).
|
||||
2. The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised), 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op, `latency-shadow-2026-10-06.md` 5, item 4); the tensor block at its free-band ceiling costs a third of that.
|
||||
3. That is the finding, and it is negative for the scheme's purpose (section 5.3): a shadow lever moves the chip's edge only by the joules it makes the HONEST card spend on work the chip cannot do more cheaply. The tensor path is so efficient on the GPU that the block adds 0.23 microjoules at most, so at `k = 1` the chip's edge falls from 6.9x to 4.9x on GDDR7 against this 4090, where the ALU shadow took the 5090 from 5.6x to 2.1x, and the tensor block costs the verifier 26x more per unit of chip-forcing energy (4.39 ms scalar per 0.23 microjoules against 0.17 ms per 0.6 microjoules). The property the design hoped for (`k_mma` near 1 because the GPU's tensor engine is near the floor) is real and is exactly why the lever is weak: there are no joules in it to force.
|
||||
4. The correctness chain holds at every rung: the PTX fragment read and the plain-integer reference agree on all 2^24 lanes at every R, and the CPU interpreter matches the GPU on 1,024 lanes at every R; the probe's fragment layout (`family-probe.cu` mm8 `warp_ref`) was used as written and needed no correction. Registers 29 to 36, occupancy unchanged. This is the first class-shaped evidence that an `mm8` family is cheap and exact for the honest NVIDIA card, which is what the reserve entry R8 needs; it is not evidence for a class v5.
|
||||
|
||||
Consequences per tier (B at R = 128, the largest R whose scalar verifier delta stays near 1 ms):
|
||||
|
||||
| Tier | Meaning | What to do |
|
||||
|---|---|---|
|
||||
| NVIDIA Turing and later (RTX 20 to 50 series), any memory size | rate unchanged, +11.5 W on a 4090 (+5.7 percent), per-pound unchanged, per-watt 0.95x | nothing; the block would not be adopted on this evidence |
|
||||
| NVIDIA Pascal (GTX 10 series) | emulation at about 16 steps per `mm8` (approximate): 1,024 per hash is about 16,000 counted steps, inside the latency shadow of a 1080-class card by the counted-ops rule, approximate; unmeasured | the emulation path exists in the kernel text (`IGNEUM_MM8_REF`) and is bit-exact; a GTX 1080 row is owed if B ever proceeds |
|
||||
| AMD RDNA 3 and 4 | the WMMA layout is unverified (status item 6); emulation otherwise at about 12 to 16 steps per `mm8` | the gate of section 7 rank 2 stays open; not worth running unless B proceeds |
|
||||
| Apple | emulation at about 10 steps per `mm8` (approximate), 10,000 counted steps at R = 128 inside the M5 Max's 130,000 ceiling | owed; not worth running unless B proceeds |
|
||||
| Rig, pool user | nothing changes in shares; a rig pays the block's watts per card | nothing |
|
||||
| Node operator (verifier) | +1.14 ms per unit scalar at R = 128 (+55 percent of 2.06 ms), +4.39 ms at R = 512; a SIMD byte-dot path would cut it 4x to 16x (approximate) | the verifier cost per joule forced is the reason the scheme loses to the ALU shadow |
|
||||
| A chip | must carry a tensor array per 32 lanes in flight; at the measured block energy it pays at most 0.23 microjoules per hash more than today at `k = 1`, 0.07 at `k = 0.3` | the chip's edge barely moves (5.3) |
|
||||
|
||||
### 5.2 `state-dataset` (scheme C), box 2 and igneum-build-1
|
||||
|
||||
ROWS_C
|
||||
Prototype: `proto-newpow/state-dataset/` (README with every command and log; `kernel_sd.cu` = the pack's kernel plus `igneum_leaves` and `igneum_build_sd`, the hash kernel byte for byte the pack's; `verify_sd.c` the 32-lane C interpreter and the item check; `cpu_rows.c` the igneum-build-1 rows). Box 2: RTX 4090 24 GB, driver 595.91, nvcc 12.8, host EPYC 7352; igneum-build-1 EPYC 9454P at load 19 to 32 (shared), one core pinned, nice 19. Run 19:34 to 19:50 UTC. The leaf stand-in: one ChaCha12 block of (sigma, S, t, 0, tag) with `S = K XOR 0x5a5a5a5a` (the README defines it).
|
||||
|
||||
| Row | Control (mx8-genesis) | sd1 | Reading |
|
||||
|---|---|---|---|
|
||||
| Hash rate, 10 batches of 2^24, two passes | 63.083, 63.083 MH/s | 63.088, 63.087 MH/s | equal within 0.01 percent: the kernel is the same binary over a dataset of the same shape |
|
||||
| Watts, mean after 10 s of a 20-s window; SM MHz | 207.6, 205.2 W; 2,745 | 207.9, 206.5 W; 2,745 | equal within the 2 W run-to-run noise; 3.27 microjoules per hash on this 4090 (the 5090 is 2.34 to 2.65) |
|
||||
| MH/s per W | 0.304, 0.307 | 0.303, 0.305 | unchanged |
|
||||
| Cache fill, GPU | 1.81, 1.85 ms | 1.85, 1.85 ms | unchanged |
|
||||
| Leaf array (2^24 ChaCha12 blocks, 1 GiB) on the GPU | none | 5.49 ms | the synthetic stand-in; a real snapshot comes from the node |
|
||||
| Dataset build, second pass | 30.55 ms | 31.98 ms | +1.43 ms (+4.7 percent): one coalesced 64 B read per item |
|
||||
| Chunked build (leaves streamed from pinned host memory in 64 MiB or 256 MiB chunks, one stream) | | 75.8 ms, 75.5 ms | PCIe-bound: 62 ms of copy (17.2 GB/s on this box) plus the build, partly serialised; two streams would hide most of the build (not done) |
|
||||
| Device memory during the build (context 395 MiB included) | 1,675 MiB | 2,699 MiB resident leaves; 1,741 MiB chunked at 64 MiB; 1,933 at 256 MiB | while hashing 1,803 MiB in every mode (leaves freed) |
|
||||
| 2^24 fingerprint at base 0 | 7c28cfb06c5c65a9 (the pack's, both passes) | d5b0c16390cad0e8 (four runs) | |
|
||||
| Pack vectors (3 warps, standalone and in batch); dataset head, word [MASK], 64 Mac samples | PASS | n/a (new values; sd1 base 0 lane 0 = b600edbed969becc) | the harness is the pack's |
|
||||
| Build bit-exact: 1,024 random items, GPU words against the plain C host derivation | 1,024 of 1,024 | 1,024 of 1,024 (also after both chunked rebuilds); 64 device leaves = host leaves | |
|
||||
| Hash bit-exact: 4 dumped warps (128 lanes) against the C interpreter with lazy `mh_word` / `mh_word_sd` | 128 of 128 (and 32 of 32 against the Mac vector at base 0) | 128 of 128 | |
|
||||
| Host cache FNV-1a 64 | 48c4f5bf24166b2e = the Mac's | same | |
|
||||
|
||||
The CPU rows (igneum-build-1, ms per unit of 4,096 items, 100 units; two readings at load 19 and 31):
|
||||
|
||||
| Row | Reading 1 | Reading 2 | Per item |
|
||||
|---|---|---|---|
|
||||
| (i) 4,096 random 64 B reads from a 2 GiB resident leaf array | 0.108 ms (0.092 to 0.172) | 0.163 ms | 26 to 40 ns |
|
||||
| (i) the same from an 8 GiB array (the year-12 size) | 0.139 ms | 0.209 ms | 34 to 51 ns |
|
||||
| (ii) 4,096 leaves derived on the fly (one ChaCha12 block each) | 0.284 ms | 0.285 ms | 69 ns |
|
||||
| (iii) 4,096 x `mh_item` on the host cache, naive, no interleaving | 9.01 ms | 11.24 ms | 2.2 to 2.7 microseconds |
|
||||
| 4,096 x (leaf read + `mh_item_sd`), naive | 9.31 ms | 11.64 ms | |
|
||||
| Daily leaf array on CPU (2^24 blocks): one core / 32 threads | 1.67 s / 0.073 s | 1.68 s / 0.083 s | 100 ns per leaf |
|
||||
| Host cache fill (256 MiB), one core | 0.447 s | 0.452 s | |
|
||||
| Light-client sample: distinct items touched by lane 0's 128 loads | 128 of 2^24 | | 128 openings x 25 levels x 32 B + 128 x 64 B = 108 KiB per lane; 3.4 MiB per 32-lane unit before dedup (arithmetic) |
|
||||
|
||||
Reading, against the gate: the project's Rust verifier does the 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items (the naive C figure here, 9 to 11 ms, is an upper bound and is not the production verifier); sd1 adds 0.11 to 0.21 ms per unit when the verifier holds the state in RAM (+5 to +10 percent of 2.06 ms; about +0.3 to +0.5 ms on a 2019-class core by the 2.5x rule), so the class v3 verifier plus sd1 sits at about 2.2 ms on the M5 Max core and about 5.7 ms on the 2019-class core, inside the 10 ms gate with the x8 margin intact. The snapshot's own cost is the state read on the node, not the leaf hashing (1.7 s on one core for 2^24 leaves).
|
||||
|
||||
Consequences per tier (sd1):
|
||||
|
||||
| Tier | Meaning | What to do |
|
||||
|---|---|---|
|
||||
| Home miner, one 8 GB card | hash rate, watts and MH/s per W unchanged (measured equal); the build must stream the leaves (resident leaves at the designed 2 GiB dataset would peak at about 4.6 GiB, approximate; chunked 64 MiB about 2.7 GiB, approximate, measured 1,741 MiB at 1 GiB), so the 8 GB mine-and-prove profiles of `prover-tiers-real-cards.md` keep their memory | ship chunked only; the kernel already takes an item range, so chunking needed no kernel change |
|
||||
| 12, 16, 24, 32 GB cards | unchanged in rate and watts; the daily build +1.4 ms resident or about 76 ms streamed (about 150 ms at 2 GiB, approximate) | nothing |
|
||||
| AMD and Apple | the hash kernel is untouched, so the measured class v3 rates stand (18.9 MH/s on the 9070 XT, 27.1 on the M5 Max); the build kernels gain one 64 B read per item in OpenCL and Metal | the two build kernels in the emitter's other dialects |
|
||||
| Rig | per card as above; one 2 GiB snapshot per day per rig, shared by its cards | the node beside the rig serves it over loopback |
|
||||
| Pool user | a 2 GiB download per day from the pool, or the pool ships the built dataset; today a pool user needs nothing but the day key | the pool protocol (spec 09) gains a daily snapshot or dataset fetch |
|
||||
| Solo miner | must run a node with execution (it already does for templates) and hold the state | nothing new beyond RAM |
|
||||
| Node operator | +2 GiB RAM for the snapshot (+4 GiB at year 4), a daily serialisation pass, the verifier +0.11 to 0.21 ms per unit | the node RAM budget stays inside a 16 GB home node at launch sizes |
|
||||
| Light client, holder | a verifiable sample of 128 state leaves per lane per block exists; checking it costs 108 KiB per lane fetched on demand, so it is an availability spot check, not a light verification of PoW | rank 6 of section 7 |
|
||||
| Prover, rollup customer | nothing | |
|
||||
|
||||
### 5.3 The chip rows on the measured numbers
|
||||
|
||||
CHIP_ROWS
|
||||
Method: `sim/horizon/new-pow/chip_rows.py` (the chip-model-v3 section 5.4 f = 1 chip plus the block's measured energy times `k`; the card figures are 5.1's; every chip figure is arithmetic on `chip-model-v3.md`'s cited and approximate memory figures). The denominator here is the 4090 measured tonight (3.19 microjoules per hash at the control, a card whose memory system is half the 5090's), not the 5090, so the R = 0 column reads 6.9x where `chip-model-v3.md` reads 5.1x for the 5090; the movement with R is the row that matters.
|
||||
|
||||
| Scheme and setting | Honest card microjoules per hash | Chip, GDDR7, k = 1 / 0.5 / 0.3 (microjoules) | Edge over the card, GDDR7 | Chip, one HBM3 stack, k = 1 / 0.5 / 0.3 | Edge, HBM3 |
|
||||
|---|---|---|---|---|---|
|
||||
| Class v3 today (R = 0, the 4090 control) | 3.19 | 0.466 | 6.9x | 0.321 | 9.9x |
|
||||
| B at R = 128 (1,024 tiles per hash, block 0.18 microjoules) | 3.37 | 0.646 / 0.556 / 0.520 | 5.2x / 6.1x / 6.5x | 0.501 / 0.411 / 0.375 | 6.7x / 8.2x / 9.0x |
|
||||
| B at R = 512 (4,096 tiles per hash, block 0.23 microjoules, the free band's top) | 3.42 | 0.696 / 0.581 / 0.535 | 4.9x / 5.9x / 6.4x | 0.551 / 0.436 / 0.390 | 6.2x / 7.8x / 8.8x |
|
||||
| For comparison, the ALU shadow at N = 100,000 on the 5090 (`latency-shadow-2026-10-06.md` 6, measured 3.27 microjoules at the 431 W cap) | 3.27 | 1.59 / 1.03 / 0.80 | 2.1x / 3.2x / 4.1x | 1.44 / 0.88 / 0.66 | 2.3x / 3.7x / 5.0x |
|
||||
| C (sd1) at any size | unchanged (measured equal) | 0.466 | unchanged | 0.321 | unchanged; the f = 0 recompute chip must add 2 GiB of DRAM and becomes this chip |
|
||||
|
||||
Reading: B moves the f = 1 chip's edge by 1.4x to 2x at `k = 1` and by 1.1x at `k = 0.3`, at a verifier cost of 1.1 to 4.4 ms per unit (scalar); the ALU shadow moves it by 2.7x at `k = 1` and 1.4x at `k = 0.3` at 0.17 ms per unit. On every axis the measured tensor block is the weaker lever, for the reason stated in 5.1 item 3. C moves nothing and claims nothing against the chip; its merit is elsewhere.
|
||||
|
||||
## 6. Verdicts
|
||||
|
||||
| Scheme | Verdict | Why, in one line |
|
||||
|---|---|---|
|
||||
| A, mining is proving | NEVER (A1, A2); A0 folds into C | one proof per segment is not a distribution of puzzles; the bytes (2.9 MB of openings per block) or the verify (32 to 40 ms) kill every form that is not "hold the trace", and holding the trace is C with a worse data source |
|
||||
| B, tensor-shaped shadow | VERDICT_B |
|
||||
| C, stored state | VERDICT_C |
|
||||
| B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joules |
|
||||
| C, stored state | SHIP AS CLASS v5 CANDIDATE (through the spec items of 4.3 and the Devnet 2 gate): hash rate and watts unchanged by construction and measured equal, build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact on 1,024 items and 128 lanes; a new property per block (a random sample of state) and a new requirement per mining operation (hold the state); the open decision is what a header verifier is asked to hold |
|
||||
|
||||
**A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.
|
||||
|
||||
VERDICT_BC_PROSE
|
||||
**B, in full.** The scheme asked whether a GPU-shaped puzzle could make a chip into a GPU. The answer the 4090 gives is that the tensor-shaped work is free for the honest card (rate unchanged to R = 512, 15 W at most for 4,096 tiles per hash) and therefore nearly free for the chip too: a lever against the f = 1 memory-controller chip works only through joules the honest card is forced to spend on work the chip cannot do more cheaply, and the ALU shadow of class v4 spends 11 pJ per op where the tensor path spends 0.056 to 0.70 pJ per multiply-add. The chip's edge moves from 6.9x to 4.9x at k = 1 (4090 denominator) where the ALU shadow moved the 5090's by 2.7x (5.6x to 2.1x), and the verifier pays 26x more per unit of forcing energy (4.39 ms per 0.23 microjoules against 0.17 ms per 0.6). There is also the vendor split (native on Turing and later and on RDNA 3 and later, emulated on Pascal, RDNA 2 and every Apple card) and the unverified AMD layout. So: never as class v5 content for this purpose. What the measurement does establish, and what should be kept: an mm8 family is exact on NVIDIA at every R tried (2^24 lanes PTX = reference, 1,024 lanes CPU = GPU at five rungs), cheap for the honest card, and bounded in verifier cost by a known law (1.1 microseconds per step per unit, scalar); that is the evidence the reserve entry R8 lacked, and the two-output form (both tile outputs consumed) is a correction R8 needs so a chip cannot halve the tile. The standing rule this leaves for the programme: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is LEAST efficient per op among the operations a chip cannot do much better (RandomX's argument), never where it is most efficient.
|
||||
|
||||
**C, in full.** The scheme does what it says at a cost that is measured and small: the hash kernel is byte for byte the shipped one, the 4090's rate and watts are equal within noise (63.083 against 63.088 MH/s, 205 to 208 W), the daily build grows by 1.4 ms resident or about 76 ms streamed over PCIe (about 150 ms at the designed 2 GiB, approximate), device memory during the build stays at dataset plus cache plus one 64 MiB chunk, the verifier grows by 0.11 to 0.21 ms per unit with the state in RAM (about 2.2 ms on the M5 Max core, about 5.7 ms on a 2019-class core by the 2.5x rule, inside the 10 ms gate), and every GPU word and lane agrees with the plain C derivation (1,024 of 1,024 items, 128 of 128 lanes). It is never weaker than today's dataset against any chip and it removes the f = 0 recompute chip as a category. What it adds has not been done on a shipped chain in this form: the dataset IS the keyed execution state, so every mining operation holds the chain, and every block names 128 random state leaves per lane that any full node can open against the day's state root. Its limits are stated: a pool can ship the dataset (a 2 GiB daily delivery per pool miner, where today a pool miner needs only the day key), a header verifier must hold the state or be shown 3.4 MiB of openings per unit (so light clients still rely on certificates, as spec 10 says), and the canonical serialisation is new consensus-critical code that must cross a day boundary and a finality pause on the fast-time harness before Devnet 2. Verdict: ship as the class v5 candidate, with the three spec items of 4.3 (the serialisation, the pruning-proof witness, the pause rule), the Devnet 2 gate, and the open decision on what a header verifier is asked to hold in front of it.
|
||||
|
||||
**The ranking after measurement.** C, then nothing else from this lane as class content; B retires into the reserve entry it improves; A is recorded as a bound. The brief's hope of "mining is proving" is answered with the reason it cannot be, which is worth more to the project than a design that pretended otherwise.
|
||||
|
||||
## 7. Ranked next steps
|
||||
|
||||
Hours are agent hours (the project lead's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own.
|
||||
Hours are agent hours (the project lead's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own. Ranks 2, 3 and 4 were written as B's gates before the ladder landed; after section 5 they are WITHDRAWN (B is not carried forward as class content; the rows stay so the reasoning is visible) and the live order is 1, 5, 6, 7, 8.
|
||||
|
||||
| Rank | Proposal | Evidence | Model | Hours | Consequence per tier | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
|
|
@ -236,4 +325,8 @@ One paragraph each.
|
|||
|
||||
## 9. Summary for the coordinator
|
||||
|
||||
SUMMARY
|
||||
This lane wrote three candidate proofs of work for Igneum, reviewed each in two personas, prototyped the two that survived on real cards, and measured. (1) Scheme C, the dataset derived from the chain's execution state, is the one to carry forward as the class v5 candidate: measured on an RTX 4090 the hash rate and watts are unchanged to the third decimal (63.08 MH/s, 207 W, the kernel is byte for byte the shipped one), the daily build grows by 1.4 ms resident or about 76 ms streamed, the verifier by 0.11 to 0.21 ms per unit with the state in RAM, every GPU item and lane agrees with the plain C derivation (1,024 of 1,024, 128 of 128); it forces every mining operation to hold state, names 128 random state leaves per lane per block, and removes the f = 0 recompute chip as a category; its costs are a 2 GiB daily delivery for pool miners and three spec items (canonical serialisation, a pruning-proof witness, the pause rule). (2) Scheme B, a tensor-shaped shadow of int8 8x8x16 tiles, is bit-exact on NVIDIA at every rung (2^24 lanes PTX = reference, 1,024 lanes CPU = GPU, R = 8 to 512) and free in hash rate to 4,096 tiles per hash, and that is exactly why it fails as an anti-chip lever: it costs the honest card 0.056 to 0.70 pJ per multiply-add and at most 0.23 microjoules per hash, so the f = 1 chip's edge moves from 6.9x to 4.9x at k = 1 against the 4090 where the class v4 ALU shadow moved the 5090's from 5.6x to 2.1x, at 26x the verifier cost per joule forced; never as class content, kept as the evidence and the two-output correction for reserve family R8. (3) Scheme A, mining is proving, is never: one proof per segment is not a distribution of puzzles; the useful fraction is bounded by proving work per segment over network hashes per segment (8 percent at 1 GH/s, 0.08 percent at 100 GH/s on this chain's measured figures); every form that is not "hold the trace" dies on 2.9 MB of openings per block or a 32 to 40 ms proof verify, and "hold the trace" is scheme C with a worse data source.
|
||||
|
||||
1. Scheme C costs nothing in hash rate or watts (63.083 against 63.088 MH/s, 205 to 208 W on the 4090, measured) and 0.11 to 0.21 ms per unit of verifier time (measured on igneum-build-1), so class v5 can carry it; the model is `item_sd(t) = item(t) with s ^= leaf(t)` over the day's certified state root.
|
||||
2. The tensor shadow moves the f = 1 chip's per-joule edge by at most 1.4x (6.9x to 4.9x at k = 1, R = 512) for 4.39 ms of scalar verifier per unit, against 2.7x for 0.17 ms from the ALU shadow; the model is `chip_rows.py` on the measured block energy of 0.23 microjoules per hash.
|
||||
3. Mining cannot be proving on a lottery with a 10 ms verifier: the useful fraction is bounded by 5 GPU-seconds of proving per 8-second segment over the network's hashes in that segment, 8 percent at 1 GH/s and falling with the hash rate; the model is that ratio on `prover-tiers-real-cards.md` and bench-log 2582.
|
||||
|
|
|
|||
491
docs/analysis/horizon/polish.md
Normal file
|
|
@ -0,0 +1,491 @@
|
|||
# Horizon lane 6: polish. Every shipped system against the best in its class
|
||||
|
||||
Date: 6 October 2026, evening UK (written 20:00 to 21:00Z, while the live devnet's finality was still paused). Lane 6 of the Horizon programme. Worktree `/Users/joshm/Projects/igneum-wt-horizon` (branch `horizon`, HEAD c3aa502). Output: this file only. No file outside it was edited; nothing was built, deployed, posted or started; the installed app and every live service were read, never touched.
|
||||
|
||||
the project lead's bar: "a level of polish that has not been seen before." This file is the honest audit: what each shipped system does, the named comparator and the exact screen or feature it has that we lack or do worse, the ledger rows that close the gap (95 rows, ids Q1 to Q105 with gaps, each with a file or screen, a severity, hours of agent work, an owner and a gate), and what is already better than the comparator. The first ten rows are what the project lead will notice first when he opens each product tomorrow.
|
||||
|
||||
## 0. What was read and run
|
||||
|
||||
Read, in the worktree: `app/igneum-app` (src, ui, tests), `app/mac/IgneumMiner.swift`, `app/windows/host.cpp`, `docs/plans/miner-ui-2.md`, `miner-ui-3.md`, `miner-ui-3-audit.md`, `ember-tune.md`, `release-0.3.14.md`, `docs/design/app-screens/`, `docs/design/miner-tuning.md`, `docs/fud-ledger.md` (X17, X19, X21 to X28, M25 to M29, G7, G10, G13), `site/` (every page, `api/`, `lib/`, `verify/`, `build.mjs` read not run, `scrub.mjs`, `forbidden-strings.txt`, `downloads.json`, `journey.json`, `vercel.json`, `sitemap.xml`), `docs/design/site-polish-2026-10-04.md`, `docs/plans/site-ui-3-shots/`, `docs/plans/site-miner-2026-10-06.png`, `docs/review/round-4-reddit-2026-10-06.md`, `docs/plans/explorer.md` and `explorer/`, `relay/` (ui, api, lib, clients, playbooks, tests), `tools/relay.mjs`, `tools/console.mjs`, `tools/jobs.mjs`, `docs/design/relay/`, `docs/design/console/`, `pool/` (src, web, tests, README), `docs/plans/pool.md`, `docs/spec/09-pool-protocol.md`, `tools/community/discord-hooks.mjs` and its tests and fixtures, `docs/community/discord-hooks.md`, `tools/fleet/`, `tools/build-job.mjs`, `tools/build-remote.sh`, `tools/ci/`, `docs/plans/build-job.md`, `build-server.md`, `hands-on-build-1.md`, `node-changes.md`, `packaging/README-ship.md`, `packaging/ota/README.md`, `packaging/mac/README.md`, `packaging/windows/README.md`, `tools/observer/` (observer.mjs, README), `docs/spec/03` section 3.9, and this lane's siblings `finality-and-weight.md` (3.1, 6) and `frontier.md` (3.7).
|
||||
|
||||
Read, read-only, in other agents' worktrees: `/Users/joshm/Projects/igneum-wt-ship0315/docs/plans/release-0.3.15.md` (the newest release notes; `release-0.3.15.md` does not exist on `horizon` or `master`), `/Users/joshm/Projects/igneum-wt-wallet*` (wallet 0.1.4 source, UI 3 branch), `/Users/joshm/Projects/igneum-wt-gpu-fleet/tools/fleet/`, `/Users/joshm/Projects/igneum/vendor/igneum-node-0310/consensus/src/processes/finality.rs` (the node's `finality_reason`).
|
||||
|
||||
Run (nothing live, nothing built): `node --test app/igneum-app/ui/*.test.mjs` (35 pass), `node --test relay/test/*.test.mjs` per file (24 pass), `node --test tools/community/discord-hooks.test.mjs` (31 pass, every live poster stubbed), one `curl -s https://igneum.network/` (the allowed single fetch; saved to the scratchpad), and one read-only SQL pass over the observer's Neon tables (`live_checkpoints`, `live_events`, `live_state`) through the same HTTPS SQL endpoint `site/api/_neon.mjs` uses, to see what every surface had to render during tonight's pause. No browser automation.
|
||||
|
||||
Comparator claims are from memory unless a repository or page is named, and are labelled approximate.
|
||||
|
||||
## 1. The ten rows the project lead will notice first
|
||||
|
||||
| Rank | Id | Row | File or screen | Severity | Hours | Owner | Who it hits | Gate |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| 1 | Q1 | The pause has no cause anywhere. The node reports `finality_reason` (`active`, `window filling, N of M`, `paused`) but not the frozen-table share that held tonight's pause; the observer drops even the reason; the API has no reason field; `/live` computes "N% of weight silent" from the sliding table, which read 11 percent silent at 20:05Z while finality was paused (held by the table frozen at lock 6842, 57.3 percent signing). Add `finality_reason`, `held_by` (frozen index, its signing share, its expiry DAA) and lane 3's `finality_provisional` end to end | `vendor/igneum-node-0310/consensus/src/processes/finality.rs:1794-1807` (reason exists, no frozen share), `tools/observer/observer.mjs:720` (copies `finalityActive` only), `site/api/live.mjs:114-122` (no reason), `site/live.html:520` (the silent percentage) | the project lead-visible | 4 (node 1.5, observer and API 1, `/live` and `/api/stats` copy 1, spec 3.9 row 0.5) | consensus-engineer (node), site owner (observer, API, page) | holder, exchange: the one line that says whether a pause is a silent third, a split, or a table waiting to expire; miner: nothing lost, but the app can finally explain itself; operator: a pause with a cause and an end time | A forced pause on Devnet 2 (one slice over a third stopped): `/api/live.finality` carries `reason`, `held_by`, `provisional`; `/live` reads "finality paused since 18:40Z: 89% of the sliding weight is signing, held by the table frozen at lock 6842 (57% signing) until DAA 216,402 (about 20:40Z)"; the exchange guidance in spec 3.9 carries the provisional row (lane 3 proposal 4) |
|
||||
| 2 | Q2 | Ember never says "paused". The engine keeps only `last_lock`, `last_lock_at`, `age_s`, `votes`; the node line showed "#6842 · 1 h ago" under "a point the miners agreed can never be undone" for the whole two-hour pause. No `finality_active` anywhere in the app | `app/igneum-app/src/state.rs:247-248`, `src/engine.rs:3495-3497, 3896-3902` (lock parsed from the miner's `LOCK checkpoint` line only), `ui/app.js:1350-1351` | the project lead-visible | 3 | app owner (engine reads `getFinalityCheckpoints` or the miner's `FINALITY` line, spec 3.9 row; UI state and view test) | home miner and rig: the app tells them finality is paused and why, instead of a lock age that climbs; pool user: the same on the pool page later | Mock state with `finality_active:false, reason:"paused"` renders "Finality paused since 18:40Z, 89% of the 30-day weight signing, held by the frozen table until about 20:40Z" on the node line and in the pill; `view.test.mjs` case; seen on the Devnet 2 forced pause |
|
||||
| 3 | Q3 | A fresh machine's first 60 seconds start with a warning. macOS: ad hoc signature, no Developer ID, no notarisation; the engine strips `com.apple.quarantine` from its own bundle on start; the README's "Right-click > Open" bypass is gone on macOS 15 (approximate). Windows: installer and exes unsigned, SmartScreen "Windows protected your PC", then a UAC prompt for the firewall rule 20 to 50 s in before any screen explains it | `packaging/mac/build-dmg.sh:118-122`, `packaging/mac/README.md:51-53`, `app/igneum-app/src/main.rs:78`, `packaging/windows/README.md:76, 111`, `packaging/windows/build-installer.ps1:188`, `app/igneum-app/src/ota.rs:981-1015` | the project lead-visible | 6 (Developer ID signing plus `notarytool` and stapling in `build-dmg.sh` 3; Authenticode `signtool` step in `windows.yml` and `build-installer.ps1` 2; firewall step after `setup_done` with a sentence on the Cards screen 1), plus the project lead: an Apple Developer account and an OV or EV code-signing certificate (purchases, hours of the project lead's time, approximate) | app owner; the project lead (the certificates) | every new miner on every tier, Windows and macOS: Signal's and Tailscale's installers open with no warning (approximate); today ours open with two | Fresh macOS 15 VM: the DMG's app opens on double-click, `spctl -a -vv` says accepted and notarised; fresh Windows 11 VM: no SmartScreen interstitial, `signtool verify /pa` passes; the first UAC prompt appears after the Cards screen names it |
|
||||
| 4 | Q4 | Every update is "urgent" once an activation height is behind the node: `fork_is_close(Some(198000), 209000)` is true, so the manifest's stale `activation 198000` turned the 0.3.14 update into "the node is 0 blocks away. Installing now", stripped Later and skipped every safe-moment guard (the PC 1 install under a measurement job at 17:52:54Z). The same path has no finality input: an update applies through a pause | `app/igneum-app/src/manifest.rs:350-355` (`daa + 1800 >= h`), `:322-347` (`safe_to_apply`, no finality), `src/ota.rs:484`, `ui/app.js:217`, `docs/plans/release-0.3.14.md:79`, `release-0.3.15.md` section 6 | the project lead-visible | 3 (close means within 1,800 blocks ahead and not behind 1; the publisher drops a passed `activation_height` 0.5; `Moment.finality_paused` holds a non-urgent update 1.5) | app owner | home miner, rig: no surprise restart mid-measurement or mid-pause; operator: an activation that has passed is not an emergency | Unit tests: behind by any amount is not close; `publish-manifest.sh` refuses a passed height; a paused mock holds the update with the words "finality is paused; installing when it resumes"; the 0.3.16 manifest carries no stale height |
|
||||
| 5 | Q5 | The homepage says "final" while finality is paused, in a 90-word hero, under a share card that renders as a thumbnail. The chain scene takes `src.locked` and draws the dashed line labelled "final" at the newest locked block with no `finality.active` check (R4.6.3 was fixed on `/live` only); the hero is one 90-word paragraph with two bench deep links and "ships when the packaging row lands"; the miner section still says "a 24 GB NVIDIA card proves as well" while the hero and litepaper say 8 GB; home, litepaper, live and explorer share a 256 px `og-small.png` with `summary` cards | `site/index.html:810, 838` (final label), `:345` (hero), `:494` and `site/miner.html:7, 13, 21, 265, 382` (24 GB), `site/index.html:15-19` (OG) | the project lead-visible | 3 (final label 0.5, hero 1, 24 GB sentence 0.5, four 1200x630 cards 1) | site owner | every visitor; a holder or exchange reading "final" during a pause is the worst of them | No "final" label while `finality.active` is false (the `/live` rule); hero under 40 words with one link; one proving-tier sentence on every page; every page's shared card is 1200x630 |
|
||||
| 6 | Q6 | The hub shows stale data as live, said "active" for the first 20 unlocked checkpoints, says "paused" with no since or cause, and buried the pause under 401 miner_quiet and miner_back events. After the first load an API failure only changes the header to "offline: ..."; tiles, cards and "Refreshed" keep the old values. `finality_active` flips only after `presence_window` (20 on the devnet) indices without a lock, so 18:42 to 18:50Z read "active" in white with no lock forming. No `finality_paused` or `finality_resumed` event exists; `live_events` between 18:25Z and 21:30Z holds 203 `miner_quiet`, 198 `miner_back`, 55 `difficulty`, 30 `checkpoint_locked` (the last at 18:42:10Z) and no pause row. The checkpoint table also showed "% of active" above 100 (102.4 to 113.4 percent on indices 6900 to 6930) | `relay/ui.html:564` (stale), `:367-368` (tile), `relay/api/console.mjs:214`, `vendor/igneum-node-0310/consensus/src/processes/finality.rs:1794` and `vendor/igneum-node/consensus/core/src/finality.rs:63, 75` (`presence_window` 240 mainnet, 20 devnet), `tools/observer/observer.mjs:652` (the only finality event), `live_checkpoints.fraction_active` rows 6900 to 6930 | the project lead-visible | 4 (dim every panel with "last data N min ago" 1; amber on the first under-2/3 checkpoint 0.5; `finality_paused` and `finality_resumed` events, and collapse quiet/back pairs under 5 min 1.5; clamp or explain the active fraction 1) | fleet agent (hub), consensus-engineer (observer events, the fraction) | operator: the hub is the project lead's first screen; a holder reading the public API gets the same events | Cut the network: every hub panel dims within 15 s; a forced Devnet 2 pause turns the tile amber on the first checkpoint under two thirds, posts one `finality_paused` event with the cause and one `finality_resumed` with the duration; no fraction above 100 percent on any row for 24 h |
|
||||
| 7 | Q7 | Discord said nothing. The bot has no credentials file and its timer is not installed, so tonight's pause produced zero posts; had it been live, a condition already firing at the watcher's first look is marked `preexisting` and never opens an incident; the pulse hides the pause in its description with no colour or title change | `docs/community/discord-hooks.md:81-82`, `tools/community/discord-hooks.mjs:481-485`, `:252-255`, `:158`, `infra/build-server/discord-hooks/install.sh` | the project lead-visible | 3 (install and one live pulse 1; a pre-existing condition opens with "since at least <first look>" 1; paused pulse in a distinct colour with a title suffix 1) | miner-community-lead (community owner) | every Discord reader, which on launch day is every miner | `check` prints three "set"; one live pulse in #numbers; a fixture where the pause predates the first tick opens an incident; the paused pulse renders amber with "(finality paused)" in the title |
|
||||
| 8 | Q8 | The relay's security fixes are not on master. Commit 28c028b (X23 to X28: `relay/lib/guard.mjs`, `handler.mjs`, run signatures, retention) is on `fud-close` and the `ledger-*` branches only; `git merge-base --is-ancestor 28c028b master` is false. On this tree any of the three intake keys (one ships in every miner app) can post, sync and delete console items, read the whole feed, and every client puts the token in the URL path | `relay/api/console.mjs:249, 283-301`, `relay/lib/relay.mjs:34-46`, `relay/clients/agent.sh:10`, `igneum-agent.ps1:13, 72, 154`, `tools/relay.mjs:24` | operator-visible (security) | 2 (merge or cherry-pick with the 47 relay tests 1; handler test that an intake key gets 403 on post, sync, delete 1) | fleet agent (relay owner) | operator: the console is the control plane for every PC and the fleet; a miner app's intake key is in every install | `28c028b` is an ancestor of the deploying branch; 47 relay tests pass; the handler test above passes; `relay/README.md` names the deployed commit |
|
||||
| 9 | Q86 | The fleet page is blind to the standing fleet, to death and to finality. `page.py:43-47` publishes `standing` and `devnet2` and the page references neither (grep 0); a dead box reads "running" for up to 20 minutes because state comes from `boxes.json` and `standing.jsonl` is never read; no fleet or console script reads `finality_active`, so tonight's pause was found by hand. The night the project lead ruled that 22 boxes stay up permanently, the page that shows them cannot show them | `igneum-wt-gpu-fleet/tools/fleet/page.py:38-47`, dlsite `index.html:243`, `lib/standing.py:54-56`, `tools/console.mjs:124` | the project lead-visible | 6 (standing block with USD per day and the Devnet 2 gate line 2; "unreachable since HH:MM" within 2 minutes 2; finality on the page and in `console.mjs chain` 2) | fleet agent | operator (the project lead reads this page every evening); every tier indirectly: a dead standing box is weight that left silently, the class of tonight's pause | the page shows the standing count and spend; a box killed by hand reads "unreachable since" within 2 minutes; a forced Devnet 2 pause reads on the page and in the console |
|
||||
| 10 | Q9 | The wallet page describes 0.1.4 while UI 3 (five state words, light mode, pounds line) sits unmerged on `wallet-ui-3`; the MetaMask guide's testnet RPC is still labelled a placeholder, its devnet RPC port is the miner's node (26790) while a wallet-only Mac runs 26800, and `wallet_addEthereumChain` carries `blockExplorerUrls: []` and no `iconUrls`, so MetaMask shows no explorer link and a blank icon | `site/wallet.html:261-266, 312`, `igneum-wt-wallet-ui/docs/plans/wallet-ui-3.md:93, 112`, `site/metamask.html:197, 225, 233, 285-286`, `igneum-wt-wallet/app/igneum-wallet/src/node.rs:24-26` | the project lead-visible | 6 (wallet 0.1.5 with UI 3 published and the page rewritten 4; guide: both ports, the placeholder line removed, explorer URL and icon in the request 2) | app owner (wallet), site owner (guide) | holder: the first non-miner product; a wrong port or a placeholder RPC is a dead end at the first step | `downloads.json` wallet-mac 0.1.5; the page names pending, included, executed, proven, finalised; `eth_chainId` returns 0x116e from the guide's URL; MetaMask shows the explorer link after the add-chain button |
|
||||
|
||||
Next in line, outside the ten: Q10 (the explorer knows nothing about finality and swallows failures after first load, 3 hours), Q80 (the Windows installer pipeline is dead on GitHub billing) and Q12 (Ember's light mode fails contrast on the private key and every live state).
|
||||
|
||||
Hours for the ten: 40 agent hours, plus the project lead's certificate purchases for Q3.
|
||||
|
||||
## 2. Tonight's pause on every surface (the ledger row the project lead asked for)
|
||||
|
||||
Times UTC. The chain-side facts are lane 3's (`finality-and-weight.md` 3.1) and the observer's own rows, read tonight: last lock 6842 at about 18:39:40Z; 6843 proposed at 53.1 percent of total; `finality_active` false from about 18:50Z (index 6862, twenty indices without a lock, `presence_window` 20); no lock after 6842 as of 20:05:03Z (`live_state.updated_at`), `total_weight` 7,167, `active_weight` 6,346 (88.5 percent signing on the sliding table), `voters` 85, and no `finality_reason` key in the stored JSON. Sampled `live_checkpoints` rows:
|
||||
|
||||
| Index | First seen | State | Signed, percent of total | Signed, percent of active | Votes |
|
||||
|---|---|---|---|---|---|
|
||||
| 6820 | 18:28:30 | locked | 88.8 | 100.0 | 80 |
|
||||
| 6840 | 18:42:10 (ingested after the hub came back) | locked | 79.5 | 98.6 | 83 |
|
||||
| 6850 | 18:43:41 | proposed | 55.2 | 76.8 | 75 |
|
||||
| 6880 | 19:00:22 | proposed | 50.1 | 81.2 | 49 |
|
||||
| 6910 | 19:13:55 | proposed | 67.4 | 107.4 | 55 |
|
||||
| 6920 | 19:19:07 | proposed | 74.8 | 113.4 | 78 |
|
||||
| 6960 | 19:40:19 | proposed | 86.0 | 100.0 | 78 |
|
||||
| 7010 | 20:04:32 | proposed | 87.7 | 99.1 | 73 |
|
||||
|
||||
From 6910 (19:13:55Z) every proposed checkpoint carried over two thirds of the sliding total and did not lock: the frozen table of Q5 held it (lane 3, 3.1). No surface could say so, because no field carries it. The "percent of active" above 100 on 6900 to 6930 is a display bug on the hub's checkpoint table (the active denominator lags the signers).
|
||||
|
||||
| Surface | What it showed 18:40Z to 20:05Z (from the code path and the rows above) | Honest? | Fix row |
|
||||
|---|---|---|---|
|
||||
| Ember (`app.js:1350`) | "last lock #6842 · 1 h 25 min ago" with "a point the miners agreed can never be undone; locks at 2/3 of the 30-day weight". No pause word anywhere | No: the age climbed and the note contradicted the state | Q2 |
|
||||
| Site `/live` (`live.html:512-520`) | Status LIVE, blocks flowing, finality bar grey, "finality paused: 45% of weight silent" at 18:50Z falling to "11% of weight silent" by 20:05Z, no final region, no checkpoint ticks; "last lock #6842, 1 h ago" in the stats (`:404`, no active check); a hovered checkpoint block still read "locked" (`:450`) | Half: the pause is named (R4.6.3) but the percentage says the opposite of why it is paused once the stayers passed two thirds | Q1 |
|
||||
| Homepage (`index.html:810, 838, 879`) | Counter "last lock: paused"; the chain scene still drew the dashed line labelled "final" at the newest locked block | No: "final" and "paused" on one screen | Q5 |
|
||||
| Explorer (`explorer.html`, `block.html`) | Identical to a normal hour; stars on locked checkpoint rows; block pages of ordinary blocks "not a checkpoint" | No: nothing marked the pause | Q10 |
|
||||
| Hub (`relay/ui.html:367-368`) | 18:30 to 18:42Z: "observer STALE N s" in red, tiles frozen, finality "active" (the hub itself was down with the Mac). 18:42 to about 18:50Z: "active" in white, no lock forming. About 18:50Z onward: "paused" in amber, "last lock #6842" in amber, eight rows "proposed" at 53 to 89 percent, Events card full of `miner_quiet` and `miner_back` | Half: paused, but no since, no cause, no end; active for ten minutes of no lock | Q6 |
|
||||
| Discord | Nothing: the bot is not live. Had it been: the 20:00Z pulse would have read "Finality paused since 2026-10-06 18:4x UTC: nothing is final until two thirds of the 30-day weight signs again; mining and proving are paid as usual" in the description, and the watcher would have marked the pause `preexisting` and never opened an incident | No | Q7 |
|
||||
| Wallet 0.1.4 (`igneum-wt-wallet` `app.js:185, 340`) | Node card "paused" when a local node reports `finality_active` false, or "not checkpoint-checkable without a node" on the public RPC; transaction rows stay "in block N" with no mention | Half | Q35 |
|
||||
| Pool page (`pool/web/index.html`) | Not deployed; the code has no finality concept; it would have confirmed and paid on blueness throughout | n/a (a design choice, stated in Q50) | Q50 |
|
||||
|
||||
The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operator surface either hid the pause, named it without a cause, or contradicted it, for two hours, during the first finality pause the project has had in public. The gate across all of them is one forced pause on Devnet 2 with a screenshot of each surface, before the public testnet opens.
|
||||
|
||||
## 3. Per system
|
||||
|
||||
### 3.1 Miner app (Ember)
|
||||
|
||||
**What it does today.** A Rust engine (`app/igneum-app/src/`) runs `igneumd` and one GPU worker per card, serves a tokenised dashboard on 127.0.0.1 (`src/server.rs:11-33`, routes `:184-359`: `api/state`, `api/log`, `detect`, `setup`, `key/saved`, `key/reveal`, `start`, `pause`, `resume`, `cards`, `tune/goal`, `settings`, `prove`, `update/check|install|auto|open`, `jobs/allow|check`, `power/apply|control`, `sweep/*`, `clock/*`, `quit`). Screens (`ui/index.html`): Welcome (line 71), Cards (102), Address (127), the one-time key sheet (438), then a rail with Mine, Earnings, Prove, Settings (164 to 413), a log drawer (460), a status strip (`Notices`, `ui/app.js:7-140`) and an update card (`UpdateCard`, `app.js:154-243`). Ember Tune (`src/ember.rs`, `src/sweep.rs`, `docs/plans/ember-tune.md`) tunes power and clocks per card against a goal. Updates are Ed25519-signed manifests (`src/ota.rs`, `src/manifest.rs`). Hosts: `app/mac/IgneumMiner.swift` (WKWebView, menu bar), `app/windows/host.cpp` (WebView2, tray). The shipped build is 0.3.14 with miner-ui-3; 0.3.15 (miner-ui-4: the big-number hero and the chain scene) is staged, its canary failed on block version 1026 and was rolled back (`release-0.3.15.md` section 5). UI tests: 35 pass.
|
||||
|
||||
**Comparators and the exact gap** (all approximate, from memory of the products).
|
||||
|
||||
| Comparator | Their screen or feature | Ours, and where it falls short |
|
||||
|---|---|---|
|
||||
| NiceHash QuickMiner | Profitability per day in fiat on the main screen; Rig Manager web view of your own rigs | Earnings reads "£0.00 earned: nothing is bought or sold on devnet" (`index.html:250`); mined IGN is never shown, only proving `paid_wei` (`app.js:1364`); no owner-facing remote view (the console is the operator's) |
|
||||
| HiveOS | Hash rate and temperature charts over hours; Telegram or Discord alert on a GPU fault or a rig offline; per-GPU fan and memory clock | A 10-minute blocks strip only (`app.js:1052-1127`); no outbound alert; manual control is the power slider (`app.js:1255`); the memory knob exists in the engine with no control |
|
||||
| lolMiner, T-Rex consoles | Per-GPU accepted, rejected, invalid and faults on every status line | Per card only "N blocks" (`app.js:306`); `rejected_session` a total in the hero sub-line (`app.js:1290`); `mismatched=` and `faults=`, which X21 made the engine read, never reach the row |
|
||||
| Nanominer | Restart-on-hash-drop watchdog with a visible threshold | `src/watchdog.rs` exists; the row says "stopped after repeated faults" (`app.js:296`) with no count or threshold |
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q2 | Finality pause invisible (section 1) | `state.rs:247`, `engine.rs:3495`, `app.js:1350` | the project lead-visible | 3 | app owner | section 1 |
|
||||
| Q4 | Every update urgent once the activation is behind; no finality hold (section 1) | `manifest.rs:322-355` | the project lead-visible | 3 | app owner | section 1 |
|
||||
| Q11 | Offline start says "The update failed." A check error with no held manifest calls `set_error`; the cached manifest is loaded for the floor but never into `self.manifest`; the notice offers "Try again", which POSTs install | `ota.rs:649-654, 195-200`, `app.js:29, 1198` | user-visible | 1 | app owner | notices test: a check-class error renders "Could not check for updates" with no install action |
|
||||
| Q12 | Light mode fails contrast on every live text: molten `#B8731F` on white 3.81:1, on bone 3.38:1; ember `#E04A14` on bone 3.61:1; molten carries the private key, the address and the live states. `miner-ui-3.md` section 9 claims 4.5:1 everywhere | `app.css:29, 263, 338, 431` | user-visible | 1 | app owner | a node script over the token pairs asserts 4.5:1 for every text token in both schemes; CI runs it |
|
||||
| Q13 | Per-card rejects and faults never reach the row (the lolMiner line) | `app.js:291-309, 1290` | the project lead-visible on PC 1 | 2 | miner-community-lead | row meta "N rejected · M faults" when nonzero; view test |
|
||||
| Q14 | Windows first run raises a UAC prompt (firewall rule) before any screen explains it (part of Q3) | `ota.rs:981-1015`, `index.html:71-99` | user-visible | 1 | app owner | the firewall step runs after `setup_done`; the Cards screen names it |
|
||||
| Q15 | Mac first-run instruction stale: "Right-click > Open" (part of Q3) | `packaging/mac/README.md:51`, `site/miner.html` download copy | the project lead-visible on a fresh Mac | 1 | app owner | the download page and README name the Privacy and Security step until the certificate lands |
|
||||
| Q16 | No rate history: a 10-minute strip against HiveOS's hours | `app.js:1052-1127` | user-visible | 4 | app owner | one-hour per-card sparkline from an engine ring buffer; view test |
|
||||
| Q17 | Hill climb hidden until `tune_climb` is on the state | `index.html:341`, `app.js:1407` | operator-visible | 1 | app owner | the row shows whenever the engine answers `tune/goal` |
|
||||
| Q18 | Earnings never shows mined IGN; "£0.00 earned" is the first number on the tab | `index.html:250`, `app.js:1364` | user-visible | 2 | miner-community-lead | the tab shows blocks mined and the subsidy they earned in IGN, with the devnet line under it |
|
||||
| Q19 | 10 px type in five places, 9 px ruler | `app.css:196, 271, 272, 302, 491, 531` | cosmetic | 1 | app owner | no `font-size` under 11 px except the ruler |
|
||||
| Q20 | Design screens in `docs/design/app-screens/*.png` are a UI 1 app (v0.3.0 tiles, a Finality card the shipped UI lacks) | `docs/design/app-screens/` | cosmetic | 1 | app owner | screens regenerated from `?screen=` on the current build |
|
||||
| Q21 | AMD step line prints a percent as watts ("owed: a unit word for AMD") | `ember-tune.md:278` | operator-visible | 1 | miner-community-lead | the line reads "70% (an offset)"; unit test |
|
||||
| Q22 | A code comment naming the founder ships in the UI bundle (`// The list is ordered by performance (the project lead, 6 October 2026)`), and the ui bundle is outside every forbidden-string check | `app/igneum-app/ui/app.js:270`, `tools/ci/identity-check.sh` | cosmetic (identity rule) | 0.5 | app owner | `tools/ci` forbidden-string check covers `app/igneum-app/ui/`; the comment reads "(ruling of 6 October 2026)" |
|
||||
| Q23 | Two copy-law slips: "Votes lock the chain; leave it on." (aphorism), "The window can close. The miner keeps going" (two-beat) | `index.html:404, 81` | cosmetic | 0.5 | app owner | reworded; the grep below stays at 0 em dashes |
|
||||
|
||||
**Error states, exact strings.** Engine down (four failed polls): pill "engine away" (`app.js:1437`); node line "Node: no answer", sub "the engine is not answering; the window reconnects by itself"; Mac host "The engine stopped. Quit and open Igneum Miner again." (`IgneumMiner.swift:233`); Windows "The engine stopped. Close this window and open Igneum Miner again." (`host.cpp:301`). Node down: "the node could not start" or the engine message, "the node is not running", "the external node went away" (`app.js:452-453`, `engine.rs:3163`); pill "node failed" (`app.js:1428`); digest box "the node is not running" (`app.js:1341`); canvas "waiting for the node to sync" (`app.js:1099`). Finality paused: nothing (Q2). Update server unreachable: "The update failed. Manifest signature: <curl error>." (`app.js:1360`, Q11). GPU lost: "Card removed: <name>. Its worker stopped." (`app.js:95`), row word "removed", sub "unplugged; its worker stopped. The row goes in five minutes." (`app.js:303`); Code 43: "<name>: not usable (Code 43). No worker runs on it." with "reboot with the card attached; if it persists, reinstall the driver with the card attached" (`app.js:100`, `detect.rs:55`).
|
||||
|
||||
**Already better than the comparators.** Automatic efficiency tuning against a goal with the money consequence on screen (`app.js:408-419`; 37 to 41 percent more hashes per watt measured on PC 1, `ember-tune.md:254, 271`). Signed, hash-checked, rolled-back updates with a safe-moment rule and machine slots (`ota.rs`, `manifest.rs`). A reason word on every zero rate and hot-plug words (`ember-tune.md` 6a). A virtualised log drawer with time jump and error marks (`app.js:880-1050`). Confirmations in place, no dialogs. A tokenised local API with origin checks (`server.rs:102-125`). Dark-scheme contrast is good: bone 17.4:1, ash 7.0:1, ember 5.6:1, molten 11.0:1 on obsidian (`app.css:19`). Focus rings, `aria-live` strip, labelled switches and sliders, reduced motion honoured on every animation (`app.css:67, 113, 142, 148, 157, 420, 437, 466`; `app.js:1137`). None of the four comparators ship the first two.
|
||||
|
||||
### 3.2 Wallet
|
||||
|
||||
**What it does today.** One desktop wallet (Igneum Wallet, Rust engine plus a local window; `igneum-wt-wallet/app/igneum-wallet/`: `engine.rs`, `vault.rs`, `hd.rs`, `tx.rs`, `finality.rs`, `node.rs`, `updater.rs`, `server.rs`; ui `index.html`, `app.js`, `lock-screen.js`, `update-card.js`), shipped as 0.1.4 (`site/downloads.json` wallet-mac, 19.6 MB; `release-0.3.15.md:90`). 0.1.5 (bundled 0.3.14 node, RPC on localhost) has a DMG and is unpublished (`release-0.3.14.md:82`). UI 3 is commit ab8dfdc on `wallet-ui-3` and `wallet-0.1.5`, not on master. No wallet inside the miner app (it holds the payout key, `wallet.json`, and a button "Open the wallet" to the site, `app/igneum-app/ui/app.js:826`). No web wallet. A MetaMask guide (`site/metamask.html`), an address page (`site/address.html`, noindex, balance plus blocks mined), a marketing page (`site/wallet.html`). Keys: BIP-39 (12 to 24 words, `hd.rs:23`), BIP-44 Ethereum path; vault Argon2id 64 MiB, 3 passes, XChaCha20-Poly1305 (`vault.rs:14-81`); key zeroised on lock (`engine.rs:392`); 24 words shown once, three typed back (`engine.rs:321`); raw key export behind password or Touch ID (`ui/index.html:271-275`). Chain ids 4461, 4462, 4463 (`metamask.html:233`). Signing: EIP-1559 value transfers only; no EIP-712, no `personal_sign`, no contract calls, no dapp provider, no hardware wallet (spec 8.5 item 2 decided, not built; `docs/fud-ledger.md:2058`).
|
||||
|
||||
**Comparators and the exact gap** (approximate unless a standard is named).
|
||||
|
||||
| Comparator | Their feature | Ours |
|
||||
|---|---|---|
|
||||
| MetaMask | Account list and HD derivation; address book; speed up and cancel a pending tx; `window.ethereum` provider; `eth_signTypedData_v4` | One account (`engine.rs:343` refuses a second import); an address field only (`ui/index.html:177`); no replace-by-fee; no provider; no EIP-712 |
|
||||
| Rabby | Pre-sign simulation and recipient risk flags | Review shows amount, fee and total (`app.js:293-295`) |
|
||||
| Frame | Ledger, Trezor, GridPlus signers; injected provider for any browser | Neither (spec 8.5 "MUST offer", no code) |
|
||||
| Monero GUI | View-only wallet; subaddresses; a settable node address with status | Node source engine-chosen, no settable endpoint (`wallet-ui-3.md:112`) |
|
||||
| KDX, Kaspium | Sync bar with peer count; per-transaction DAA score | "reading the chain, N of M blocks" (`app.js:175`), no peer count |
|
||||
| EIP-1193 | `request` provider | Used on the guide (`metamask.html:294`) |
|
||||
| EIP-6963 | Multi-provider discovery | Absent; the guide falls back to whichever extension owns `window.ethereum` (`:292`) |
|
||||
| EIP-3085 | `wallet_addEthereumChain` with `blockExplorerUrls`, `iconUrls` | `blockExplorerUrls: []`, no icon (`:225, 285-286`) |
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q9 | Page vs UI 3; guide placeholders; empty explorer URL (section 1) | `site/wallet.html`, `site/metamask.html` | the project lead-visible | 6 | app owner, site owner | section 1 |
|
||||
| Q30 | Private key export copies to the clipboard with no clear and no timer | `igneum-wt-wallet/.../ui/app.js:129, 463` | user-visible | 2 | app owner | clipboard cleared after 60 s by the host; unit test on the timer |
|
||||
| Q31 | Wallet OTA trusts one key; no `revoked_keys`, unlike the miner's updater | `docs/plans/consequences-2026-10-05.md:36`, `updater.rs:7, 729` | operator-visible | 3 | app owner | updater carries `OTA_PUBLIC_KEYS [K1, K2]` and honours `revoked_keys`; the three-key test of the rig installer |
|
||||
| Q32 | No hardware wallet path despite spec 8.5 "MUST offer" | `docs/spec/08-client-security.md:36` | the project lead-visible | 24, or 0.5 to relabel | app owner (cryptographer reviews) | a Ledger signs one devnet transfer through the Ethereum app at chain id 4463; or the spec row reads "Designed, not shipped" |
|
||||
| Q33 | Single account, no second address, no address book | `engine.rs:343`, `ui/index.html:177` | user-visible | 6 | app owner | two derived accounts, a saved recipient reused in Send |
|
||||
| Q34 | `confirm()` on Remove wallet never renders in the Mac host | `app.js:598`, `wallet-ui-3-audit.md:189` | user-visible | 1 | app owner | in-page two-step card; snapshot test |
|
||||
| Q35 | Finality pause: node card "paused" or "not checkpoint-checkable without a node", rows stay "in block N"; public-RPC mode never says what to do | wallet `app.js:185, 321, 340` | user-visible | 2 | app owner | the row reads "in block N; finality is paused on the network" and the public-RPC line names the fix with a button |
|
||||
| Q36 | Reorged-out transaction: "in a block" then the row vanishes; no resend hint | `app.js:340`, `docs/fud-ledger.md:2154-2158` | user-visible | 3 | execution-engineer (node side), app owner (the row) | a transfer unwound in the P23 conformance run shows "reorged out, resending" and ends executed |
|
||||
| Q37 | The address page shows no finality state and "n/a" with a bare reason for balance | `site/address.html:205-209, 242-252` | cosmetic | 2 | site owner | the page shows the latest locked checkpoint and whether finality is active |
|
||||
| Q38 | The permanent phishing sentence differs between surfaces ("words or your key" on the wallet, "your key" in the miner) against spec 8.5's one verbatim sentence | `site/wallet.html:366`, wallet `ui/index.html:54`, `app/igneum-app/ui/index.html:97`, `08-client-security.md:39` | cosmetic | 0.5 | site owner, app owner | one sentence, grep-identical on all three |
|
||||
| Q39 | Vault file permissions not set explicitly (no `0o600` in `vault.rs`) | `vault.rs` | operator-visible | 0.5 | app owner | the file is created 0600; test |
|
||||
|
||||
**Error states.** Node down: pill "no node", balance note "waiting for a node" (`app.js:160, 176`), the node line's raw engine message (`:186`; UI 3 replaces it with "the node stopped answering; trying again", `wallet-ui-3.md:199`). Pending versus final: "PENDING, waiting for a block", "IN A BLOCK, chain block N; final once a verified checkpoint covers it", "FINAL, under checkpoint N, certificate verified by this wallet", "FAILED, the execution failed; the fee was still paid" (`app.js:340`). Address page: "The API did not answer: <message>" (`address.html:270`), "Not found." (`:247`). MetaMask page: "No Ethereum wallet found", "Waiting for the wallet...", "Cancelled in the wallet.", "The wallet refused: <message>" (`metamask.html:292-296`).
|
||||
|
||||
**Security notes.** Plaintext key only in engine memory, zeroised on lock and after each sign (`engine.rs:392, 563`); the window never holds the key (nonce-gated reveal, `server.rs:108`); local API on 127.0.0.1 with a random port, per-run token path and a host token (`server.rs:1, 39, 71`); chain id checked at quote and send (`engine.rs:549`). Testnet RPC is https (`testnet-go.md:13`); devnet is http to 127.0.0.1. Chain ids 4461 to 4463 were absent from chainid.network on 3 October (`docs/design/execution-layer.md:354`), not yet registered (`frontier.md:645`); from memory no listed chain uses them (approximate).
|
||||
|
||||
**Already better.** The wallet verifies finality certificates itself with the node's consensus code (`finality.rs:101-121`) and never trusts the RPC's "locked"; MetaMask and Rabby trust the RPC's block tag (approximate). Touch ID is bound to the exact quote (address, value, nonce, chain id) with a 30 s one-use nonce (`engine.rs:539`). Argon2id at 64 MiB is above MetaMask's PBKDF2 default (approximate). The QR is drawn locally as SVG with the `ethereum:` URI and a checksummed address (`server.rs:140`). The MetaMask page says plainly that nothing on it asks for a seed (`metamask.html:172`).
|
||||
|
||||
### 3.3 Hub and relay
|
||||
|
||||
**What it does today.** One Vercel project (`igneum-relay`, Neon database, Blob store; `relay/README.md:3-5`). `relay/ui.html` (583 lines) is the console at `/r/<token>/` (`relay/vercel.json:5-8`); `relay/index.html` is a placeholder. Functions: `relay/api/relay.mjs` (feed, items, files, inbox, register, name, role), `relay/api/console.mjs` (machines, jobs, builds, chain, log, results, tuning), `relay/api/wake.mjs` (public long-poll for the apps' jobs file, 30 a minute per IP, `relay/lib/wake.mjs:15, 42-54`). Auth: the 20-character token in the path or `x-relay-token`, or any of `RELAY_KEY`, `LOG_INTAKE_KEY`, `LOG_INTAKE_KEY_NEXT` in `x-igneum-key` (`relay/lib/relay.mjs:34-46`); only `task`, a `run` or `task` drop, `name`, `role`, `delete` demand the token (`relay/api/relay.mjs:114-115`). Two job models: relay `run` items polled every 20 s by `relay/clients/igneum-agent.ps1:165` and executed as administrator; and signed app jobs in `igneum-jobs.json`, whose status the console derives from `miner_logs` rows (`relay/api/console.mjs:172-188`). Seven tabs (`relay/ui.html:186-192`): Machines, Jobs, Builds, Chain, Work log, Results, Relay; refreshed every 15 s when visible (`:577`), server cache 10 s (`console.mjs:21`). Chain tiles: blocks, identities, hash estimate, difficulty, last lock, finality, peers, blocks per second (`:363-370`). Mac tools: `tools/relay.mjs`, `tools/console.mjs`, `tools/jobs.mjs`. Tests: 24 pass per file (`node --test relay/test` as a bare directory fails under Node 22.23.2 with MODULE_NOT_FOUND; the README's command needs the glob).
|
||||
|
||||
**Comparators and the exact gap** (approximate).
|
||||
|
||||
| Comparator | Their screen | Ours |
|
||||
|---|---|---|
|
||||
| Tailscale admin console, Machines | Last seen, OS and version, key expiry with "Disable key expiry", tags; an ACL editor with tests | Last seen and a red card after 180 s (`relay/ui.html:270, 273`); no credential state (one static token, `relay/lib/relay.mjs:36`), no per-machine key, no revoke; roles are labels never checked (`relay/api/relay.mjs:191-195`) |
|
||||
| GitHub Actions job view | Live streaming log with line anchors, per-step timing, re-run, artefacts, annotations at the top | One text blob in a modal (`relay/ui.html:319`); `duration_s` and the error lines that `tools/jobs.mjs:41-44` already parses are discarded by `console.mjs:184-187`; no re-run; results are separate items |
|
||||
| Vercel deploy log | Build steps, deployment list, instant rollback, diff between deployments | Builds tab shows the current manifest, the last CI fetch, 40 posts and a folder listing (`relay/ui.html:323-341`); no rollback, no manifest history, no diff |
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q6 | Stale shown as live; "active" for 20 unlocked indices; "paused" with no since or cause; event flood (section 1) | `relay/ui.html:564, 367-368` | the project lead-visible | 4 | fleet agent, consensus-engineer | section 1 |
|
||||
| Q8 | X23 to X28 fixes off master; intake key can write and delete; token in URL (section 1) | `relay/api/console.mjs:249`, `relay/lib/relay.mjs:44` | operator-visible | 2 | fleet agent | section 1 |
|
||||
| Q40 | A hung job reads "running" forever: no elapsed time, no timeout, no "last line N min ago" | `relay/api/console.mjs:182-187`, `relay/ui.html:317` | the project lead-visible | 2 | fleet agent | a killed job reads "no report for 12 min" in red within one refresh |
|
||||
| Q41 | Jobs tab trusts the unsigned `igneum-jobs.json` while the apps verify the signed file | `relay/api/console.mjs:170` vs `tools/jobs.mjs:57-60` | operator-visible | 1.5 | fleet agent | the tab shows the envelope's signature state per file |
|
||||
| Q42 | Job result modal: no error lines first, no duration, no anchors (the Actions view) | `relay/ui.html:319`, `console.mjs:184-187` | the project lead-visible | 3 | fleet agent | a failed run opens with its five error lines and its duration first |
|
||||
| Q43 | Machines: no per-machine key state, no revoke, roles never enforced (the Tailscale view) | `relay/lib/relay.mjs:14`, `relay/api/relay.mjs:191-195` | operator-visible | 3 (after Q8) | fleet agent | a revoked PC's register is 403 and its card says "key revoked" |
|
||||
| Q44 | Builds: no rollback, no manifest history (the Vercel view) | `relay/ui.html:328-341` | operator-visible | 4 | fleet agent | one tap republishes the previous manifest and the fleet's OTA state shows it |
|
||||
| Q45 | "live feed unreachable: live feed unreachable" doubled; the whole Chain tab empties, losing Events and the infra card | `relay/api/console.mjs:210`, `relay/ui.html:357` | cosmetic | 0.5 | fleet agent | live feed down: tiles grey, Events and infra card still render |
|
||||
| Q46 | Replay: `register` and `task` have no nonce; a captured `task` POST re-queues a run | `relay/api/relay.mjs:153-167` | operator-visible (security) | 1 | fleet agent | a replayed `task` is refused; test |
|
||||
|
||||
**Error states.** Relay API unreachable: header "offline: Failed to fetch" and a grey dot (`relay/ui.html:564, :32`); the panel is replaced only while it still shows "loading". Node down: a machine card goes red after 180 s with "silent 4h" (`:270, 273`; `SILENT_S`, `console.mjs:35`); a deliberate stop reads "stopped (quit) 2h ago" (`relay/lib/parse.mjs:134-140`); the live API down reads "live feed unreachable: <error>" (`:357`); the observer stale reads "observer STALE N s" in red after 30 s (`:361`, `site/api/live.mjs:8`) while every tile keeps its last value. Finality paused: `tile('finality', f.active ? 'active' : 'paused', ...)` in amber `#FFB35C` (`:368, :17, :65`); the "last lock" tile amber too (`:367`); eight checkpoint rows with state and fraction (`:374`). Job hang: "running" until a `SUMMARY` with `finished_at` arrives (`console.mjs:179-182`); relay `run` items show "read" and never "done" (`tools/relay.mjs:58`).
|
||||
|
||||
**Security, X23 to X28 on this tree.** X23: the tier split exists (`relay/api/relay.mjs:114-115`), no run signature, no per-machine HMAC, the intake key still writes the console (`console.mjs:249`). X24: every client builds `/r/$TOKEN/api` (`agent.sh:10`, `send.sh:11`, `igneum-agent.ps1:13`, `tools/relay.mjs:24`, `tools/console.mjs:27`). X25: `Arm-Restart` at top level and `/RL HIGHEST` (`igneum-agent.ps1:154, 72`). X26: feed to 500 with no retention; `delete` leaves blobs (`relay.mjs:48, 148-152`); `__DL_BASE__` substituted into bodies (`tools/relay.mjs:79`). X27: `from` is free text, `register` for any hostname (`relay.mjs:30, 153-167`). X28: constant-time compare and HSTS present (`relay/lib/auth.mjs:5-11`, `relay/vercel.json:38-41`); still a GET that acks (`relay.mjs:85`), `RELAY-REBOOT` anywhere in output (`igneum-agent.ps1:133`), no rate limit, `NOPASSWD:ALL` (`wsl-setup.ps1:40`), `sudo -S` with the password (`prover-setup.ps1:21`). All of this is what Q8 merges.
|
||||
|
||||
**Already better.** One phone-first page, no login flow, seven tabs, 15 s refresh (`relay/ui.html:5, 577`); Tailscale, Actions and Vercel each take several screens on a phone (approximate). The machine card fuses miner STATUS, node log, app telemetry and OTA state and remembers an OTA line after it scrolls out (`console.mjs:61-84`); no comparator shows workload telemetry. Stale marking at 180 s drops the card from the total; stopped and silent are distinct words (`parse.mjs:16-21, 134-140`). Drop-anywhere upload direct to Blob with a progress bar and a 50 MB cap (`ui.html:506-534`). The wake long-poll is dependency-free and tested with a fake clock (`relay/lib/wake.mjs`). Headers: nosniff, frame DENY, noindex, no-referrer, no-store, HSTS on every path (`relay/vercel.json:21-43`).
|
||||
|
||||
### 3.4 Site
|
||||
|
||||
**What it does today.** Thirteen static pages plus 404 (`site/index.html`, `litepaper.html`, `live.html`, `bench.html` 611 KB rendered from `docs/bench-log.md`, `evidence.html`, `ledger.html` 312 KB generated, `miner.html`, `miners.html` from `miner-bench.json`, `wallet.html`, `metamask.html`, `faucet.html`, `explorer.html`, `block.html`, `address.html`). `site/build.mjs:49-50, 357-372` injects `partials/head.html` (self-hosted fonts, tokens), `nav.html`, `footer.html`, stamps download links from `downloads.json`, inlines `journey.json`. `scrub.mjs` runs on bench and miners only (`build.mjs:171, 401`). Live data: index polls `/api/live` every 2 s (`index.html:704`; 10 s while off); the light client fetches `/api/checkpoint` and imports two noble libraries from cdn.jsdelivr (`verify/verify.js:5-6`); `live.html` polls every 2 s (`:665-666`); faucet POSTs `/api/faucet`. APIs (`site/api/`): `live`, `stats`, `supply`, `explorer`, `checkpoint`, `faucet`, `log`, all reading Neon tables written by `tools/observer`. Edge cache: live 1 s, stats and supply 10 s, explorer 5 s (`vercel.json`). The live HTML (one curl, 20:02Z) is 60,009 bytes, served with HSTS (preload), `X-Frame-Options: DENY`, nosniff, `Referrer-Policy: strict-origin-when-cross-origin`, no CSP; download buttons stamped "v0.3.14 · 45.4 MB"; "Public testnet: not yet open; the devnet build is here for people who want to look."
|
||||
|
||||
**Comparators and the exact gap** (approximate unless a page is named).
|
||||
|
||||
| Comparator | Their page or element | Ours |
|
||||
|---|---|---|
|
||||
| kaspa.org, ethereum.org | A Developers or Docs nav item | Nav is Litepaper, Live devnet, Engineering log, Miner, Wallet, Evidence, GitHub (`partials/nav.html:14-21`); builders land on a litepaper section (`index.html:335`) |
|
||||
| getmonero.org/downloads | Every platform, hashes and signing keys on one page | Split across `index.html:363-365`, `/miner#get`, `/wallet`; the homepage's Linux and HiveOS button links `/miner#get`, not a file; sha256 shown only for the HiveOS package (`miner.html:558`), none beside the Windows and Mac buttons, no signing key, no verify line |
|
||||
| z.cash, aztec.network | Named people, a transparency page | "one founder, pseudonymous" once in the litepaper; the Reddit review's key-powers table (round 4, section 3 row 8) absent |
|
||||
| kaspa.org, getmonero.org | Media kit, press kit, language switch, newsletter | None (`site/` has no `/press` or `/brand`) |
|
||||
| ethereum.org, kaspa.org | Light and dark | Dark only; only the litepaper honours `prefers-color-scheme` |
|
||||
| succinct.xyz, aztec.network | 1200x630 share cards | Home, litepaper, live, explorer use 256 px `og-small.png` with `summary` (`index.html:15-19`); miner, wallet, bench, evidence use the large card |
|
||||
|
||||
**Lighthouse, estimated from source (approximate; no run).** Common: `lang="en"`, a skip link (`partials/nav.html:1`), one `:focus-visible` ring (`head.html:22`), self-hosted fonts with `font-display:swap`, two preloads and fallback metrics (`head.html:1-19`), no render-blocking third-party CSS, canonical, description, OG and twitter tags on every page.
|
||||
|
||||
| Page | Performance | Accessibility | Best practices | SEO | What holds it down |
|
||||
|---|---|---|---|---|---|
|
||||
| index.html | 78 to 85 | 88 | 95 | 92 | 89 KB HTML, 35.0 KB inline JS, 24.0 KB CSS; two rAF canvas loops at display rate with no visibility or intersection stop (`:790, 853, 857`); 2,880 px webp with no srcset; h3 before the first h2 (`:344, 359`); ember on bone 3.08:1 at 15 px in the light section (`:119-120`), eyebrow `#7A776F` on bone 3.97:1 at 12 px (`:130`); 256 px OG card |
|
||||
| explorer.html | 88 to 93 | 90 | 100 | 85 | 32 KB, no images, no third party; `aria-live="polite"` on an 8-tile strip refreshed every 10 s (`:222`); not in the sitemap |
|
||||
| live.html | 85 to 90 | 88 | 100 | 85 | 64 KB, 35.6 KB JS; the canvas loop is capped at 30 fps and stops off screen (`:605-612`), good; `aria-live="polite"` on the 8-cell strip every 2 s (`:225`); canvas `aria-hidden` with no text alternative (`:241`) |
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q5 | "final" during a pause; 90-word hero; 24 GB contradiction; thumbnail OG (section 1) | `index.html:345, 494, 810, 838, 15-19` | the project lead-visible | 3 | site owner | section 1 |
|
||||
| Q50 | No downloads page: hashes only for HiveOS, no signing key, no verify line, the Linux button goes to an anchor | `index.html:363-365, 516`, `miner.html:549-558` | user-visible | 2 | site owner | `/download` lists every artefact with version, size, sha256, the OTA key fingerprint and a verify command per OS; the Linux button links the file |
|
||||
| Q51 | `/miners` has 6 rows, two card models, "not measured" MH/W on every row; the eleven-card fleet table is not ingested | `miners.html:183`, `miner-bench.json` | user-visible | 2 | miner-community-lead | rows carry W and MH/W for every measured card |
|
||||
| Q52 | `dl\.igneum` in `forbidden-strings.txt` matches the public download host on three built pages; the scrub guards only bench and miners | `forbidden-strings.txt:25`, `build.mjs:171, 401` | operator-visible | 0.5 | site owner | the pattern becomes `dl\.igneum\.network/dl/(?!public/)` or `/dl/[0-9a-f]{12,}`, and the check runs on every built page |
|
||||
| Q53 | Sitemap misses `/miner`, `/explorer`, `/faucet`, `/ledger`; every lastmod is 4 or 5 Oct; 404 says "These four pages are everything the site has" | `sitemap.xml`, `404.html:138` | cosmetic | 0.5 | site owner | the sitemap is generated by the build from the page list |
|
||||
| Q54 | `aria-live="polite"` on strips refreshed every 2 s and 10 s (a screen reader hears eight values every refresh) | `live.html:225`, `explorer.html:222` | user-visible | 0.5 | site owner | `aria-live="off"` on the strips plus one polite status line on state change |
|
||||
| Q55 | Homepage canvases never stop off screen or cap their rate | `index.html:790, 853, 857` | user-visible | 1 | site owner | 30 fps cap and IntersectionObserver as `live.html:605-612` |
|
||||
| Q56 | Light-section contrast fails (3.08:1 links, 3.97:1 eyebrow) | `index.html:119-120, 130` | cosmetic | 0.5 | site owner | AA on every token pair; the same node script as Q12 |
|
||||
| Q57 | "devnet v0" eyebrow on the live page; the chain is v4 | `live.html:219` | cosmetic | 0.1 | site owner | matches the app's network word |
|
||||
| Q58 | Heading order (h3 before the first h2) | `index.html:359` | cosmetic | 0.2 | site owner | h2 or a styled div |
|
||||
| Q59 | No team, custody or key-powers page (Reddit round 4 artefact 8); "0 admin keys in consensus" tile unchanged | `index.html:611`; no file | user-visible | 1 (plus the project lead's policy call) | site owner | the key-powers table published; the tile links it |
|
||||
| Q60 | Roadmap dates versus the testnet: the journey says "Public testnet, Aug to Oct 2027" and the litepaper "Pools and the public testnet are August 2027", while `docs/plans/testnet-go.md` has `igneum-testnet-1` seeds up and a go checklist dated 5 October 2026 | `index.html:679`, `litepaper.html:719`, `docs/plans/testnet-go.md:1-4` | the project lead-visible | 0.5 (after the project lead decides which is true) | site owner | one date for the public testnet on the journey, the litepaper and the download notice |
|
||||
| Q61 | No Content-Security-Policy header on the site (the relay has none either); inline scripts throughout, two pinned CDN modules on the homepage | `site/vercel.json:7`, `relay/vercel.json:31-43`, `verify/verify.js:5-6` | operator-visible (security) | 2 | site owner | a CSP with hashes or nonces for the inline scripts and `script-src` limited to self and cdn.jsdelivr; every page renders with no console violation |
|
||||
| Q62 | Copy-law borderline headlines: "Mined by GPUs. Proven by fire." (the brand line, the project lead's call), "GPUs are back · for good", "Last hour's chip is already obsolete.", "Your coins. Final means final.", "Dates slip. Gates do not.", "Install. Start. The card mines and proves.", "Nothing is mined here." | `index.html:343, 344, 424`, `wallet.html` h1, `litepaper.html:678`, `miner.html` h1, `404.html` h1 | cosmetic | 1 | site owner (the project lead rules on the brand line) | each either kept by the project lead's word or reworded |
|
||||
|
||||
**Reddit round 4, still open in the HTML.** Finding 6 (24 GB): half fixed (Q5). Finding 7 (eleven-card table): open (Q51). Finding 8 (admin keys): sentence added to the litepaper (`:659`), tile unchanged, no table (Q59). Finding 15 ("Monero's idea, finished for GPUs"): open (`index.html:444`, `litepaper.html:429`). Finding 16 ("devnet v0"): open (Q57). Finding 22: `/ledger` present, no nav item. Findings 2, 3, 5, 9, 10, 12, 13, 17, 19: fixed in the current files. "0% anyone else in the protocol" legend still live (`index.html:604`).
|
||||
|
||||
**Error states, exact strings.** API unreachable: `/live` after three failed polls shows "OFFLINE", "api unreachable", the canvas frozen with "api unreachable, scene frozen" (`live.html:665, 600`); `/explorer` writes the fetch error into the table on first load and is silent after (`:342-345`); `/block` "The API did not answer: <message>" (`:328`); the homepage "simulated preview · live feed unavailable" and "Live feed unavailable. This is a simulation and its counters count simulated blocks." (`index.html:870-872`). Observer stale over 30 s: "OFFLINE", "observer updated N min ago", "observer offline, last update N min ago" (`live.html:663`); explorer eyebrow "observer offline, last known" (`:318`); homepage "simulated preview · observer offline, last update N s ago" (`:708`). Finality paused: section 2.
|
||||
|
||||
**Already better.** Self-hosted latin subsets with preloads and computed fallback metrics (`head.html:1-19`); ethereum.org and kaspa.org load heavier bundles (approximate). Zero third-party scripts except the two pinned noble modules; no analytics, no cookie banner. Honest live states: "simulated preview · observer offline", "finality paused: N% of weight silent", "api unreachable, scene frozen"; Etherscan shows nothing when its indexer lags (approximate). A browser-side BLS certificate verifier on the homepage (`verify/core.js`), which no comparator homepage has. A public API with a CI field contract (`api/stats.mjs:10-14`) and supply derived from the emission rule with an hourly coinbase check. HSTS with preload, frame DENY, nosniff, immutable font caching, skip link, reduced-motion rule on every page. A dated journey with pass or fail gates inlined at build, and a ledger page of every criticism; none of the six site comparators publish the equivalent.
|
||||
|
||||
### 3.5 Explorer
|
||||
|
||||
**What it does today.** `explorer.html` polls `/api/explorer?blocks=50` every 5 s and `/api/stats` plus `/api/supply` every 10 s (`:342-343`); `block.html` fetches `?block=` or `?height=` once (`:276`); `address.html` fetches `?address=` once; `/block/:id` and `/address/:addr` are Vercel rewrites. Search classification is client-side (`lib/explorer.mjs:19-27`: hash, 0x address, bech32, height) with a server fallback for tx hashes (`api/explorer.mjs:49-59`). The data window is 24 hours of observer rows; balance needs `EXPLORER_EVM_RPC`, unset on the devnet deployment (`docs/plans/explorer.md` section 7).
|
||||
|
||||
**Comparators and the exact gap** (approximate).
|
||||
|
||||
| Comparator | Their feature | Ours |
|
||||
|---|---|---|
|
||||
| Etherscan, Blockscout | A transaction page: decoded input, logs, status, fee | None; a tx hash lands on `/block/<hash>#tx-<hash>` (`api/explorer.mjs:57`); `block.html:310` lists hash, from, to, value, bytes, with from and to "needs an EVM RPC" when unset |
|
||||
| Etherscan, Blockscout | Address page with transactions, token transfers, nonce, contract tab | Blocks mined, 24-hour earnings, balance only with an RPC (`address.html:209, 281`) |
|
||||
| Blockscout | Hosted API reference (Swagger), rate-limit statement | A card on `/explorer` (`explorer.html:232-240`) and `docs/api/public-stats.md` in the repo |
|
||||
| explorer.kaspa.org | DAG view inside the explorer; merge-set order per block | DAG view lives on `/live`; block page has blue work, parents, children, mergeset blues and reds (`block.html:285-325`), no prev and next navigation |
|
||||
| mempool.space | Mempool as the product, fee charts | Mempool count is one tile on `/live` (`live.html:233`); no charts on the explorer |
|
||||
| All four | Per-block status (final, confirmations) | None for an ordinary block (Q10) |
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q10 | No finality state; failures swallowed (named under the ten in section 1) | `api/explorer.mjs:16-25`, `explorer.html:222, 309, 342-345` | user-visible | 3 | site owner | during a forced pause every block row carries "paused" and the strip tile reads the Q1 reason; cut the API: the table dims with "api unreachable, last update N s ago" within 10 s |
|
||||
| Q63 | No tx page, no hosted API reference | `api/explorer.mjs:57`, `explorer.html:232-240` | user-visible | 4 | site owner | `/tx/<hash>` with status, fee, from, to, input bytes, shard and proof state; `/api` page with every field and the cache times |
|
||||
| Q64 | Address page: no transaction list, no finality (Q37 covers the finality half) | `address.html` | user-visible | 3 | site owner | the last 50 transfers in and out, each with its block and state word |
|
||||
| Q65 | No prev and next block navigation; no block-by-height neighbours | `block.html` | cosmetic | 1 | site owner | two links from `selected_parent` and the first child |
|
||||
| Q66 | Blocks per second inverted against the other pages ("1.2 s per block" vs "0.9 blocks / s") and difficulty "126.81 M" with a space against "126.81M" on `/live` | `explorer.html:330-331`, `live.html:329` | cosmetic | 0.5 | site owner | section 4.3's one formatter |
|
||||
|
||||
**Already better.** Block page depth for a DAG: parents, children, mergeset blues and reds, selected parent, shards with prover and lag, checkpoint weight fractions and certificates (`block.html:285-325`); explorer.kaspa.org shows less of the mergeset (approximate). Client-side search that accepts a height, a hash, an 0x or bech32 address with no round trip (`lib/explorer.mjs:19-27`). Five-second edge cache with a stale flag from the observer, so a stale page says so on first load.
|
||||
|
||||
### 3.6 Pool
|
||||
|
||||
**What it does today.** `igneum-pool`, a Rust crate (`pool/Cargo.toml`, 12 source files) speaking the spec 09 protocol in its v0 form: newline JSON over plain TCP (`pool/src/server.rs:1-2`, `protocol.rs:1-9`), the project's own message set (`hello`, `welcome`, `authorize`, `seeds`, `template`, `job`, `share`, `share_result`, `solution`, `stats`; `protocol.rs:21-170`). Every share is CPU-verified on the node's warp verifier (`verify.rs:52-64`; 1.35 ms isolated, 2.1 ms under load, `docs/plans/pool.md` section 5). PPLNS only, window in blocks of weight, default 2 (`pplns.rs:1-5`, `config.rs:67`); fee 1 percent, minimum payout 1 IGN (`config.rs:65-66`); EIP-1559 transfers from the pool's coinbase key, at most 16 per round, failed receipts re-credited (`payout.rs:197-299`). One page, `pool/web/index.html`, baked into the binary (`api.rs:166-174`): a stat strip, connect card, address lookup, blocks, payments, API card. Stats API `/api/stats`, `/api/blocks`, `/api/payments`, `/api/miners/<addr>`, `/api/pool-stats`, `/health` (`api.rs:185-200`) with fixture-pinned field contracts (`api.rs:255-272`). Operator surface: a `STATUS` line every 30 s on stdout (`main.rs:152-174`). Deployment: a written, unexecuted systemd unit behind Caddy (`pool/README.md:75-116`; "not executed", `docs/plans/pool.md` section 8). Nothing is deployed. The horizon tree equals `igneum-wt-pool-rebase` HEAD 269364a.
|
||||
|
||||
**Comparators and the exact gap.**
|
||||
|
||||
| Comparator | Their feature | Ours |
|
||||
|---|---|---|
|
||||
| 2miners | Per-worker hash rate charts over time; payout history per miner; worker-offline alerts by email or Telegram; estimated earnings per day; luck | A single 10-minute and 1-hour number, samples kept 15 minutes and not persisted (`state.rs:704-705, 779-781`); the API returns 50 payments and the page renders none of them (`api.rs:148`, `web/index.html:181-194`); no contact field (`protocol.rs:50-60`); no earnings estimate; luck defined inverted against MiningPoolStats' convention (`api.rs:72-74`) |
|
||||
| WoolyPooly | Solo mode; worker offline alerts | `scheme` hard-coded `"pplns"` (`server.rs:53`, `api.rs:80`); no alerts |
|
||||
| Kaspa acc-pool (approximate) | Prometheus and Grafana | No `/metrics`; stdout only |
|
||||
| Stratum v2 reference (github.com/stratum-mining/stratum, not cloned, approximate) | Noise-encrypted transport, job negotiation, binary framing, header-only mining | Plain TCP while spec 9.3 demands TLS 1.3 and `binding` is sent empty (`server.rs:1-2`, `protocol.rs:58-60`, `09-pool-protocol.md:37, 43, 49`); mode A only; JSON lines, and the `template` ships the whole `RpcRawBlock` to every member every second (`node.rs:147-164`) |
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q67 | Member transport is plaintext TCP against spec 9.3's TLS 1.3 with session binding; `binding` ignored | `server.rs:1-2, 115`, `protocol.rs:58-60` | user-visible (hijack risk) | 8 | consensus-engineer | rustls listener; a replayed `authorize` on a second connection is refused (O-9.7) |
|
||||
| Q68 | Node down shows "difficulty 0, DAA 0" and a ramp-day-0 reward; no "node unreachable" state | `node.rs:350-368`, `web/index.html:163-164` | the project lead-visible (once deployed) | 2 | miner-community-lead | the page renders "Node unreachable since <time>" against a stopped node |
|
||||
| Q69 | Raw integers (DAA, block count, shares) with no separators; difficulty uses `toLocaleString()` with no locale, so it varies by browser; hash rate 2 decimals against the site's 1 | `web/index.html:152, 163, 172, 190` | the project lead-visible | 1 | miner-community-lead | every number formatted as section 4.3 says |
|
||||
| Q70 | Ledger is a JSON file rewritten every 15 s; hash-rate samples and check cost not persisted, so a restart zeroes every rate tile and the 24-hour luck | `state.rs:704-710`, `main.rs:131-144` | user-visible | 4 | miner-community-lead | restart test: tiles recover within one sample window; README names the loss window |
|
||||
| Q71 | Address lookup never shows the payment history the API returns | `api.rs:148`, `web/index.html:181-194` | user-visible | 1 | miner-community-lead | payments table under the lookup |
|
||||
| Q72 | No per-miner history, no charts, no worker-offline notice, no `/metrics`; `/health` always ok | `state.rs:779-781`, `web/index.html:187`, `main.rs:152-174`, `api.rs:195` | user-visible, operator-visible | 4 + 4 + 2 + 1 | miner-community-lead | hourly buckets per address kept 7 days with a sparkline; opt-in webhook after 10 min silence; Prometheus text endpoint; health reflects `net.synced` and template age |
|
||||
| Q73 | Finality: the pool confirms and pays on blueness and has no finality concept; the page says nothing during a pause (a design choice to state) | `node.rs:279-345`, `server.rs:9` | user-visible | 1 | miner-community-lead | the page carries the network's finality state and the sentence "payouts follow blue confirmation, not finality" |
|
||||
|
||||
**Already better.** Every share CPU-verified at the node's own engine, no sampling (`verify.rs:52`; spec 9.8 item 5); most pools trust the miner's claim for low-difficulty shares (approximate). The member's vote key rides in every header it hashes (`node.rs:47`), so pool concentration does not become vote concentration; no Stratum pool does this. PPLNS split snapshotted at find time, credited only on blue confirmation, orphans pay nobody (`state.rs:821-852`). Failed payout receipts re-credit the balance (`payout.rs:288-292`). Exact binary share weights (`vardiff.rs:138-140`). Field contracts pinned by fixture tests.
|
||||
|
||||
### 3.7 Discord
|
||||
|
||||
**What it does today.** `tools/community/discord-hooks.mjs` (798 lines, Node 22, no dependencies). Webhook URLs in `~/.config/igneum/discord` (`:18-19, 41, 520`); state in `~/.config/igneum/discord-hooks-state.json`. Posts: network pulse to #numbers at 03:00, 09:00, 15:00, 21:00 UK (`shapePulse` `:236-273`); daily digest 09:00 (`:276-299`); weekly numbers Monday 09:00 (`:302-352`); release to #announcements by the shipper (`:394-409`); incident open and resolve to #incidents by hand or by the watcher (`:412-433`). Watcher conditions: `finality_paused` (5-minute hold), `proof_lag` (900 s), `observer_silent` (180 s) (`WATCH` `:37`, `:437-472`); one incident per condition, resolved after 120 s clear (`:490-503`). Idempotency keys per post (`:562, 569`); 429 backoff (`:534-551`); a forbidden-string guard before every post (`:47-96`); `allowed_mentions: {parse: []}`. Runner: a systemd timer on igneum-build-1, `tick --live` every minute (`infra/build-server/discord-hooks/`). Per `docs/community/discord-hooks.md:81-82`, no live post has been made and the install has not run. Tests: 31 pass; every live poster is stubbed.
|
||||
|
||||
**Comparators and the exact gap** (approximate).
|
||||
|
||||
| Comparator | Their feature | Ours |
|
||||
|---|---|---|
|
||||
| Ethereum, Kaspa, Monero Discords | Ticker bots renaming channels with height, hash rate, price | None; a webhook cannot rename a channel, a bot token is needed |
|
||||
| The same | Embed colour by severity; a pinned current-status message; "resolved" as an edit of the open message | Single ember colour for every kind (`:33, 158`); no pin; resolve posts a second message (`:706-712`) |
|
||||
| The same | Explorer links per block or lock | Only `/live` and `/miner` links (`:271, 297, 350, 420, 432`) |
|
||||
| The same | Reorg or stall alerts | No condition for a block-production stall (the fixture shows 11 zero minutes at 18:30 to 18:41Z) or a reorg (`:437-472`) |
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q7 | Bot not live; pre-existing never opens; pulse hides the pause (section 1) | `discord-hooks.mjs:481-485, 252-255, 158` | the project lead-visible | 3 | miner-community-lead | section 1 |
|
||||
| Q74 | No block-production stall condition, the actual first signal of tonight's 18:30Z gap | `:437-472`, `fixtures/discord/live.json` | user-visible | 2 | miner-community-lead | `blocks_stalled` condition, 3 minutes, fixture test |
|
||||
| Q75 | Resolve posts a new message instead of editing the open one | `:706-712` | cosmetic | 2 | miner-community-lead | PATCH `/messages/<id>` with the resolve fields |
|
||||
| Q76 | Embeds link only `/live` and `/miner`; the last lock could link `/block/<hash>` | `:271, 297, 350` | cosmetic | 1 | miner-community-lead | last lock links its block page |
|
||||
| Q77 | "Last lock" names 93 voters while the hash-rate field says 56 vote keys active; readers will ask | `:259, 262` | user-visible | 1 | miner-community-lead | one clause explains the two counts, or one number |
|
||||
| Q78 | Four semicolon-balanced two-beat sentences in the shipped incident texts ("Blocks keep being produced; the chain is not locking checkpoints", "finality pauses rather than locks on a minority", "Vote keys are not machines; the machine count is owed", "Blocks are produced and final as usual; their proofs arrive late") | `:347, 448, 452, 458` | cosmetic | 0.5 | miner-community-lead | reworded as single statements |
|
||||
|
||||
**Already better.** Release embeds carry the full sha256 of every artefact plus the exact verify commands (`:399-406`); most project bots post a link. Every post passes a guard that refuses the founder's name, hosts, IPs, paths, webhook URLs and mentions (`:47-96`). Every number names its API field and the posts link the raw JSON (`:314-341`). Incident texts state who is affected per tier (`:449, 458, 467`). A dry-run HTML preview renders the exact card before anything goes live (`:602-634`).
|
||||
|
||||
### 3.8 Fleet tooling
|
||||
|
||||
**What it does today.** On the `gpu-fleet` branch (`/Users/joshm/Projects/igneum-wt-gpu-fleet/tools/fleet/`, read-only; the `horizon` tree holds only `publish-fleet.sh`): providers `vast.py` (console.vast.ai v0: filtered bundle search `:47-54`, rent with optional onstart `:61-72`, instances, destroy, a price-stamped `ledger.jsonl` `:88-102`) and `runpod.py` (`:27-44`); the orchestrator `fleet.py` (rent-card, wait, setup, status, run, pull, destroy, tail, sh, dn2-version, list; `:5-13, 183-187`); the shared library the 6 October rule demands, `lib/box.py` (`Box.run/put/alive`, `install_payload` sha-checked, `stop_node/start_node`, `height/daa/peers/synced/wait_synced`, `exec_status/exec_tip/proving_status/paid_segments/paid_shards/state_root`, `node_version`, `rejects_since`, `max_reorg_since`, `start_miner/stop_miners/start_prover/stop_all`, `destroy`; `Registry`) and `lib/standing.py` (roster, install, check, update, rerent, loop, `weight_check`); `lib/test_box.py` runs against dn2-3 with one known-failed case (`:10-12`) and a PASS line (`:37`), by hand, not in CI. Box side: 20 `box-*.sh` scripts (setup, matrix, ember, prover, dn2, standing supervisor, node-swap, exec-snapshot, wave, wave-pool, pool, rig, rig-prover, rehearsal, kill, datadir, floor-v5, segal-host) plus `dn2-kill.sh`, `dn400-worker.sh`, `rehearsal-kill.sh`. Collectors: `collect.py`, `bps-collect.py`, `night.py` (60 s probe, 15-minute rows, exec-reset re-run `:45-55`), `autorun.py`, `restage.py`, `swap.py`. Gates: `devnet2-gate.sh` (PASS = zero rejects, max reorg under 4, version equal, exec roots equal at a common height, one paid segment; `:92-100`), `canary.sh`, `canary-next.sh`. Page: `page.py` builds `fleet.json`; `publish-fleet.sh` copies it and runs a full Vercel production deploy (`:10`); `publish-0315.py` is the first standing publish. PC job runner `tools/build-job.mjs` (run, publish, watch, fetch, verify; targets ae432dc7 and 1ccfe586; signed kinds `run`, `fetch`, `collect`, `restart`, `update-now`, `shard-benchmark`, `build`; `--stop-miners` becomes `stop_miners_first`, consumed at `app/igneum-app/src/jobrun.rs:563`). Console `tools/console.mjs` (post, log, machines, chain, jobs, builds, tuning, results, sync, url). CI checks in `tools/ci` (22): bash-body, check-workflow-shell, commit-string, copied-sources, identity-check, install-hooks, kit-path, link-check, no-conflict-markers, no-foreign-tree-writes, no-secrets, override-json, pinned-guests, playbook-quit, prover-socket, ps-drive-ref, public-api-check, second-engine, signer-pipe, windows-spawn, plus fixtures; `pgrep-self-match-check.sh` exists on `gpu-fleet` only (its `ci.yml:82-83`), not on `horizon` or `master`.
|
||||
|
||||
**Comparators and the exact gap** (approximate unless cited).
|
||||
|
||||
| Comparator | Their feature | Ours |
|
||||
|---|---|---|
|
||||
| Vast CLI | `search offers` with any filter expression; `create instance --onstart`; `logs` without ssh | A fixed filter set (`vast.py:47-52`); `rent-card` never passes onstart (`fleet.py:79`), so install is a second ssh step (`:102-109`); `tail` needs ssh (`:180`) |
|
||||
| SkyPilot | Declarative task YAML, `autostop -i`, spot recovery, `sky status`, `sky logs` | Imperative; one-shot boxes have no autostop (the caps at `page.py:45` are display only; 59 boxes running at 20:02Z); recovery is `standing.loop` after two dead checks at 600 s (`lib/standing.py:126-140`); `check` prints raw dicts (`:148`) |
|
||||
| Nomad | Restart policy: attempts, interval, delay, mode; allocation health in the UI | `box-standing.sh` restarts every 60 s with `pkill -9`, no cap, no backoff, no failure count (`:33, 46-49, 70`) |
|
||||
| HiveOS farm view | Per-rig cards: hash, temps, fan, accepted and rejected, last seen, flight sheets, bulk actions | The fleet page card shows card, state pill, "doing", USD/h, hours only (dlsite `index.html:176`); `standing.py:56` strips the `miner=` field the supervisor writes (`box-standing.sh:64` has MH/s, accepted and rejected, GPU util, watts); `console.mjs:114-116` has the HiveOS-style row for the app machines, not for rented boxes |
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q86 | The fleet page is blind to the standing fleet, to death and to finality. `page.py:43-47` publishes `standing` and `devnet2`, and the page's `index.html` references neither (grep 0; it reads boxes, results, phases, night, log, `:243`); a dead box reads "running" for up to 20 minutes because state comes from `boxes.json` (`page.py:38`) and `standing.jsonl` is never read; no fleet or console script reads `finality_active` (only `bps-collect.py:30` tails "Finality:" log lines), so tonight's pause was found by hand (`CLAUDE.md`, the standing-fleet rule) | `page.py:38-47`, dlsite `index.html:243`, `lib/standing.py`, `tools/console.mjs:124` | the project lead-visible | 6 (standing block, USD per day, Devnet 2 height and last gate line 2; "unreachable since HH:MM" from `standing.jsonl` within 2 minutes 2; `getFinalityCheckpoints` in the standing check and "finality paused since, reason" on the page and the console 2) | fleet agent | the page shows the standing count and spend; a box killed by hand reads "unreachable since" within 2 minutes; a forced Devnet 2 pause reads on the page and in `console.mjs chain` |
|
||||
| Q87 | `rerent` does not set up or supervise the new box: the docstring promises both (`lib/standing.py:14-16`), the code rents and patches rows (`:78-91`); a re-rented card bills idle | `lib/standing.py:78-91` | the project lead-visible (spend) | 2 | fleet agent | the new label shows synced and mining on the page within the install time |
|
||||
| Q88 | One-shot boxes have no autostop; `cap_vast 1000` and `cap_runpod 500` enforce nothing (`page.py:45`; `vast.py:61-72` has no cap check) | `page.py:45`, `vast.py:61-72`, `autorun.py` | the project lead-visible (spend) | 1.5 | fleet agent | `autorun` destroys a one-shot box N minutes after its done line; `rent` refuses above the cap |
|
||||
| Q89 | Supervisor restart loop without backoff or a failure line | `box-standing.sh:46-49` | operator-visible | 1.5 | fleet agent | after N failures a `standing_node_failed` line and a page flag |
|
||||
| Q90 | Dropped miner and GPU facts: `standing.py:56` strips `miner=`; `gpu=` never parsed | `lib/standing.py:56`, `box-standing.sh:64` | operator-visible | 1 | fleet agent | MH/s, accepted and rejected, watts per box on the page (the HiveOS card) |
|
||||
| Q91 | A transient ssh drop counts as death: `Box.run` returns 124 on timeout (`lib/box.py:49`), `check_one` marks dead on any rc (`standing.py:54`), the loop never asks the provider (`fleet.py:48-60` `refresh_ssh` unused); Vast proxy drops are known (`canary.sh:41`) | `lib/box.py:49`, `lib/standing.py:54` | operator-visible | 1 | fleet agent | provider status consulted before a dead check counts |
|
||||
| Q92 | `lib/test_box.py` pins a stale binary and digest (`:31` igneumd-0313, `:34` digest 4a0b8726; 0.3.15 is 06211d55, `publish-0315.py:15`) and is not scheduled | `lib/test_box.py:31-34` | operator-visible | 1.5 | fleet agent | a nightly PASS line on the page |
|
||||
| Q93 | The Devnet 2 gate reads no peers, no finality, no exec `blocked` (`read_box`, `devnet2-gate.sh:51-53`; `night.py:15` reads `blocked`, the gate does not) | `devnet2-gate.sh:51-53` | operator-visible | 2 | fleet agent | FAIL on 0 peers or finality paused at step 4 |
|
||||
| Q94 | `publish-0315.py` hard-codes one worktree and release (`:14-18`); `page.py:13` points at the master checkout's publish script; every publish is a full Vercel deploy | `publish-0315.py:14-18`, `page.py:13`, `publish-fleet.sh:10` | cosmetic | 2 | fleet agent | `publish.py` reads paths and sha from the release manifest; the page deploys as one file |
|
||||
| Q95 | Five 6 October rule rows have no `tools/ci` check and are therefore OPEN by the project's own rule (table below) | `CLAUDE.md` 6 October rules, `tools/ci/` | operator-visible | 5 (1 each) | fleet agent (consensus-engineer for the gate check) | each check fires on a known-failed fixture and passes a known-good one |
|
||||
|
||||
**The 6 October rules against their checks.**
|
||||
|
||||
| Rule (CLAUDE.md, 6 October 2026) | Check today | What the check is |
|
||||
|---|---|---|
|
||||
| `pgrep -f` with a literal pattern is banned | `pgrep-self-match-check.sh` on `gpu-fleet` only; absent from `horizon` and `master` | merge it; wire it in `ci.yml` on master |
|
||||
| Box operations live in one tested library | none; `devnet2-gate.sh:40-41` and `canary.sh:32-33` define their own `SSH()` and `SCP()` past `lib.box` | fail any file outside `tools/fleet/lib` that runs `ssh -i ~/.ssh/igneum-fleet` |
|
||||
| A watcher verifies the chain-side fact and is trusted only after a known-finished and a known-failed case | none; `devnet2-gate.sh` and `canary.sh` have no `--self-test` | a `--self-test` flag on every gate with both fixtures |
|
||||
| A standing box never leaves the live chain for an experiment; never remove over 10 percent of weight in an hour | `weight_check` exists (`standing.py:108-125`) but nothing enforces its call; `Box.destroy` (`box.py:136`) has no standing guard | destroy refuses standing rows; CI greps that scripts calling `stop_miners` on a label call `weight_check` first |
|
||||
| No live build before the Devnet 2 PASS line | none; `devnet2-status.json last_gate` reads "not run during the block-rate experiment" | the ship script requires `PASS <sha16>` in `devnet2-status.json`, as `commit-string-check` runs inside `build-remote.sh` |
|
||||
| Have checks | commit-string, override-json, ps-drive-ref, no-foreign-tree-writes, playbook-quit, second-engine | |
|
||||
|
||||
**Error states, exact strings.** Console machines (`tools/console.mjs:114`): `STOPPED (<reason>) <ago> ago`, `SILENT`, `live`, then `seen <ago> ago`, `synced` or `not synced`; cards `(stale, not in the total)` and `| <fault>` (`:115`). Console chain (`:124`): `last lock #N (blue N) | lag Ns STALE`; no "paused" word in `console.mjs` (the web tile has it). `fleet.py`: `no ssh` (`:116`), `ssh timeout` (`:118`), `scp failed` (`:106, 133`), `pull failed: ...` (`:145`). `lib/box.py`: `(124, "", "timeout")` (`:49`), `SshError("<label>: rc N: <err>")` (`:53`). `standing.py`: `re-renting <label> -> <new>` (`:134`), `behind: <label> <sha> wanted <sha> (a publish script moves it; the loop only reports)` (`:137`; "behind" is the binary sha, never height), `standing check: N boxes, N alive, N synced, N behind` (`:138`). Fleet page: the state pill (renting, installing, running, done, failed) and `doing` text only; no dead, behind or finality string exists. `devnet2-gate.sh`: `FAIL <box>: N rejected blocks`, `FAIL <box>: a selected-chain reorg of depth N`, `FAIL <box>: version V`, `FAIL <box>: exec root at height H ... differs`, `FAIL no segment record paid on any box during the run` (`:92-100`).
|
||||
|
||||
**Already better.** Every rent and destroy writes a price-stamped ledger line (`vast.py:43-45`, `runpod.py:25-26`), which Vast's CLI does not (approximate). The Devnet 2 gate compares exec state roots across every box at one height (`devnet2-gate.sh:98-99`), beyond any SkyPilot or Nomad health check. Gates read chain facts (height, paid segments, reorg depth), not process names. The supervisor runs the exec-recovery recipe by itself (`box-standing.sh:66-67`). The 10 percent weight rule is a real function (`standing.py:108-125`). The console distinguishes stopped-on-purpose from silent (`parse.mjs:134-140`).
|
||||
|
||||
### 3.9 The node's operator surface
|
||||
|
||||
**What it does today.** Fork-added flags (`igneum-wt-ship0315/vendor/igneum-node-0315/kaspad/src/args.rs`): `--devnet-suffix`, `--evm-rpclisten` (default 26790), `--evm-disable`, `--override-params-file`, `--igneum-exec-snapshot` (`:297`), `--ua-rule` (`:440`), `--rocksdb-wal-dir` (`:502`), `--devnet` alias `--igneum-devnet`. Kept from kaspad: `--loglevel` with per-subsystem `<subsystem>=<level>` (`kaspad/src/args.rs:262-271`), `--perf-metrics` (`:446`, debug lines only), `--utxoindex`, `--rpclisten-borsh`, `--rpclisten-json`, `--externalip`, `--appdir`. Env knobs: `IGNEUM_PROOF_VERIFIER`, `IGNEUM_PROOF_VERIFY`, `IGNEUM_DEVNET_GENESIS_BITS`, `IGNEUMD_DEVNET_BPS`, `IGNEUM_POW_STRIKES` and `IGNEUM_ATTACK_TS_OFFSET_MS` (feature-gated after X19). Finality logs (`consensus/src/processes/finality.rs`): `Finality: checkpoint {i} determined: block {h} (blue score, daa)` (`:625`), `Finality: certificate at index N received: X of Y voters, weight A of B` (`:1127`), `Finality: CONFLICTING certificate at index N` (`:1023`), `Finality: no body tip passes through locked checkpoint ...; fork choice falls back to depth finality` (`:1448`). No pause or resume log line; "paused" exists only as the polled RPC reason (`window filling, N of M`, `active`, `paused`; `igneum-node-0310 finality.rs:1794-1807`); the conflict reason is "O-3.17, not reported yet". RPC `getFinalityCheckpoints` returns `finality_active`, `finality_reason`, `window_filled_daa`, `window_full_daa`, latest lock (`rpc/core/src/model/finality.rs:293-299`); no provisional field, no frozen-table share. Exec JSON-RPC on one listener: `igneum_getExecStatus` (executedTip, blocked, startedFrom), `igneum_getProvingStatus`, `igneum_getFinalityView`, `igneum_submitProofRecord` (`igneum/exec/src/rpc.rs:858-979`). Metrics: no Prometheus endpoint (no hit in the fork's Cargo files). X19 (`docs/fud-ledger.md:1936-1946`): a slow clock disconnecting peers silently, `time_offset` unused, attack env compiled in, a dead override field; fixed on branch `ledger-fixes-0311` fbb0082a (one WARN per minute clock-skew line, attack env behind `attack-switches`, override refuses the field), pending merge; not done: the local-ahead mirror case, the faketime run, the skew field in `getBlockDagInfo` (`docs/plans/node-changes.md:18-19`).
|
||||
|
||||
**Comparators and the exact gap.** kaspad parity holds (`vendor/rusty-kaspa/kaspad/src/args.rs:241` loglevel, `:259` rpclisten-borsh, `:348` utxoindex, `:392` externalip, `:401` perf-metrics) but the fork's own subsystems (finality, exec, proving) got no new knobs. geth (approximate): `--metrics` with Prometheus or InfluxDB, `--log.format json`, `--http.api` allowlist; the fork has none, and writes on the EVM listener share it with reads (`rpc.rs:920`). reth (approximate): `--metrics`, `RUST_LOG` env filter, `reth db stats`, Grafana dashboards in the repo; the fork has no DB tool or dashboards. monerod (approximate): `status`, runtime `set_log`, `limit_rate`, `sync_info`; the fork's log level is start-only, no bandwidth limit, status is `igneum-miner watch` plus `getConnectedPeerInfo`.
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q96 | No log line on a finality flip: an operator tailing the log never sees the pause begin or end | `consensus/src/processes/finality.rs` (no pause line) | operator-visible | 1.5 | consensus-engineer | fast-time simnet prints `Finality: paused (<reason>)` and `Finality: resumed at index N`, one line per flip |
|
||||
| Q97 | `finality_reason` says `paused` and nothing else: no share that left, no frozen-table hold, no expiry, no conflict reason (O-3.17) (the node half of Q1) | `finality.rs:1794-1807` | the project lead-visible (through every surface) | 1.5 (counted in Q1) | consensus-engineer | tonight's pause reads "paused: 42.7 percent of weight left the window; held by the table frozen at lock 6842 until DAA 216,402" |
|
||||
| Q98 | No `/metrics`: finality_active, latest lock, exec tip, blocked, paid segments, peers, DAA as Prometheus gauges (the axum listener exists) | `igneum/exec/src/rpc.rs` | operator-visible | 3 | execution-engineer | `curl :26790/metrics` scraped by a Grafana dashboard committed under `infra/` |
|
||||
| Q99 | Log level is start-only; no runtime change (monerod `set_log`, reth env filter) | `kaspad/src/args.rs:262-271` | operator-visible | 2 | consensus-engineer | an RPC or SIGHUP changes the level without a restart |
|
||||
| Q100 | Clock-skew field in `getBlockDagInfo` and the local-ahead mirror case still owed after X19 | `docs/plans/node-changes.md:18-19` | operator-visible | 2 (after the `ledger-fixes-0311` merge) | consensus-engineer | `faketime -60s` gives one WARN and the field |
|
||||
| Q101 | No `--evm-rpc-api` allowlist: submit and export share the listener with `eth_` reads | `rpc.rs:920` | operator-visible (security) | 2 | execution-engineer | a public listener answers `eth_` and refuses `igneum_submitProofRecord` |
|
||||
| Q102 | No JSON log format: collectors grep four reject patterns (`lib/box.py:119`) | `kaspad` logger | operator-visible | 2 | consensus-engineer | `--log-format json`; the gate reads a field, not a pattern |
|
||||
| Q103 | No `igneumd db` subcommand (column sizes, exec snapshot tip, finality blob layout) | none | operator-visible | 3 | execution-engineer | `igneumd db stats` on a devnet datadir |
|
||||
| Q104 | Merge `ledger-fixes-0311` (X19) into 0.3.16 | `docs/fud-ledger.md` X19 | operator-visible | 1 | consensus-engineer | the faketime WARN on the shipped binary |
|
||||
| Q105 | Eleven fork `.rs` files carry an em dash, one inside a log or help string (upstream origin not verified, approximate) | `igneum-node-0315` (a grep for the character) | cosmetic | 0.5 | consensus-engineer | 0 in any string the operator can see |
|
||||
|
||||
**Already better.** `finality_active` ships with a reason string, which kaspad, geth and reth have no analog for. `--ua-rule` version admission (`args.rs:440-445`) is beyond kaspad. A fixed-height activation is refused by rule; the override file is refused on a dead or duplicate field. Every build binary carries its commit or fails CI.
|
||||
|
||||
### 3.10 Downloads, install and the update flow
|
||||
|
||||
**What it does today.** `site/downloads.json` (updated 2026-10-06T17:48:20Z) lists four artefacts: Windows installer 0.3.14 (45.4 MB), Mac DMG 0.3.14 (41.9 MB), HiveOS package 0.3.14 (24.7 MB), Mac wallet 0.1.4 (19.6 MB), with sha256 and size, under `https://dl.igneum.network/dl/public/` and four stable aliases under `/public/` (`packaging/README-ship.md`, "The public downloads path"). `tools/ship-app.mjs` cuts a version in eleven checked, resumable steps (preflight, bump, inputs, commit, ci, fetch, dmg, copy, manifest, deploy, verify, console; `README-ship.md`). The Windows installer is built only by `windows.yml` on GitHub-hosted runners; since 19:24Z tonight every hosted job dies at start ("recent account payments have failed or your spending limit needs to be increased", `release-0.3.15.md` 6a), so no Windows installer can be cut until billing is fixed. 0.3.15 is staged with `--until manifest`; its Devnet 2 canary failed (block version 1026 on the thirteen-field file) and was rolled back. The ship tool's commit step pushed the release tree to master although the run had `--branch release-0.3.15` (`release-0.3.15.md` 6a).
|
||||
|
||||
**Signing and notarisation state.** macOS: `codesign -s -` (ad hoc) on every binary and the bundle (`build-dmg.sh:118-122`); no Developer ID, no notarisation, no stapling; the engine strips `com.apple.quarantine` from its own bundle on start (`main.rs:78`, `packaging/mac/README.md:51-53`). Windows: unsigned installer and exes; SmartScreen "Windows protected your PC" (`packaging/windows/README.md:76, 111`; `build-installer.ps1:188`); the installer is per-user (`PrivilegesRequired=lowest`) so no UAC for install; the firewall rule asks once, 20 to 50 s in (`ota.rs:981`). Comparators: Signal and Tailscale ship Developer ID-signed, notarised DMGs and Authenticode-signed installers; both open with no interstitial (approximate).
|
||||
|
||||
**The update flow today, precisely** (the baseline for frontier 3.7, reproducible-build attestations with N of M).
|
||||
|
||||
| Step | What happens | Where |
|
||||
|---|---|---|
|
||||
| 1. Build | Node fork and workers built on igneum-build-1 (Linux, Windows cross) and the Mac (macOS); Windows exes reproducible (`-Wl,--no-insert-timestamp`); a binary whose strings lack its commit fails `tools/ci/commit-string-check.sh`. The Windows installer and payload zip are built by `windows.yml` on GitHub-hosted `windows-latest` from `payload-inputs.zip`, whose `payload-inputs.json` (sha256 and size of the zip and every file, fork commit, repo commit) is signed on the Mac with the OTA key and verified in CI against the key compiled into the app and the pin `packaging/windows/node-source.pin` before anything is built (G13, fixed 5 October; 16 local tests pass) | `tools/build-remote.sh`, `cross-remote.sh`, `packaging/windows/push-inputs.sh`, `.github/workflows/windows.yml`, `docs/fud-ledger.md` G13 |
|
||||
| 2. Sign | One Ed25519 key, generated once on the Mac (4 October 2026), seed in `~/.config/igneum/ota-signing-key` (0600), never in the repo or CI; public key the constant `OTA_PUBLIC_KEY_HEX` in `app/igneum-app/src/manifest.rs:28`; `igneum-ota-sign sign` over the canonical manifest bytes (64-byte detached signature, `igneum-app-latest.json.sig`); `publish-manifest.sh` refuses to sign when the embedded key is not the one in `~/.config/igneum`. Rotation = a bridge build with the new constant signed by the old key. The wallet's updater trusts one key with no revocation list (Q31). Key not in hardware (G7 says it will be) | `packaging/ota/README.md` "Keys", `publish-manifest.sh` |
|
||||
| 3. Publish | Manifest (version, per-platform url, sha256, size, kind, `min_supported_version`, notes, `activation_height`) and `.sig` to `dl.igneum.network/dl/<token>/igneum-app-latest.json`, public copy under `dl/public/`; the apps' long-poll on the relay's `/wake` fetches it within seconds; the `igneum-jobs.json` channel uses the same key | `packaging/ota/publish-manifest.sh`, `publish-public.sh`, `relay/api/wake.mjs` |
|
||||
| 4. Check | On start (20 to 50 s in) and hourly with jitter; both files through curl; signature verified over the manifest bytes before parsing; a bad signature reads "manifest signature does not verify" and the manifest is never parsed; a version not newer, or a manifest without this platform, ends the round | `ota.rs:144, 628, 1038-1048`, `manifest.rs:114-120, 175-178` |
|
||||
| 5. Download and verify | Into `<app data>/app/updates/` with resume and `--retry 3`; size and sha256 from the manifest ("sha256 mismatch: the file is not what the manifest signed"); the Mac re-hashes before mounting | `ota.rs:1054-1079, 1124` |
|
||||
| 6. Stage | macOS: mount, copy to `.Igneum Miner.app.new`, run its engine with `--version` and demand the manifest's version; an unwritable folder gives `manual` and "Open the download". Windows: the installer is the staged file | `ota.rs:1096`, `packaging/ota/README.md` "Stage" |
|
||||
| 7. Safe moment | Not urgent: network has not lost over 30 percent of identities in 10 minutes, this machine's minute of the hour, node synced, no hourly boundary within 180 s, no worker starting; a ready update applies anyway after 6 hours. Urgent (fork within 1,800 blocks, unsupported version, Install now) skips every guard. Q4: a passed activation counts as close; no finality input; 0.3.15 adds "never while a remote job is active" | `manifest.rs:322-355` |
|
||||
| 8. Apply | The engine writes `update-pending.json`, starts the helper and exits through its quit path (miners 8 s, node 30 s, last log upload). macOS helper swaps bundles, strips quarantine, reopens, restores `.previous` on failure. Windows helper runs the per-user installer `/VERYSILENT ... /IGNOTA=1` first, with the engine still mining, then the installer stops the engine and relaunches | `packaging/ota/README.md` "Apply", `ota-apply.sh`, `ota-apply.ps1` |
|
||||
| 9. Rollback | The new engine counts starts in `update-pending.json`; a third start without 90 healthy seconds restores the previous version ("rolled back" in Settings); the first Windows update from 0.3.0 has no rollback target | `ota.rs:371` |
|
||||
| 10. What the user sees | "Igneum Miner X is available.", "Downloading X: N%.", "X is ready.", "Installing X. The app restarts itself. Mining continues until then.", "X did not stay up and was rolled back.", "updated to X from Y" in the event feed; Settings: Check now, Install now, automatic switch, the key fingerprint under "Allow remote jobs from Igneum (signed)" | `ui/app.js:29-52`, `index.html` Settings |
|
||||
| 11. Second engine, paused finality | A sweep or `IGNEUM_APP_NO_OTA=1` engine never runs the updater (`engine.rs:601, 778`; the 5 October rule). Nothing in the update path reads finality (Q4) | |
|
||||
|
||||
So the trust today is one key, one builder (the box and the Mac, both the project's), one manifest, no attestation by anyone else, no on-chain record, and the client installs whatever that one key signs at the moment the safe-moment rule allows. That is the baseline frontier 3.7 moves from: N of M attestations in a registry contract, the updater refusing under N, and no install while finality is paused.
|
||||
|
||||
**Rows.**
|
||||
|
||||
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Q3 | Signing and notarisation; first-run prompts (section 1) | `build-dmg.sh:118-122`, `windows.yml` | the project lead-visible | 6 plus certificates | app owner, the project lead | section 1 |
|
||||
| Q80 | The Windows installer pipeline is dead: every GitHub-hosted job fails at start on billing; the installer and payload zip have no other builder, so 0.3.15 and every hotfix are Mac and HiveOS only until the project lead fixes Billing & plans, or the installer step moves to igneum-build-1 (Inno Setup under Wine, or a self-hosted Windows runner on PC 1) | `release-0.3.15.md` 6a, `.github/workflows/windows.yml`, `packaging/windows/build-installer.ps1` | the project lead-visible | 0 (the project lead: billing) or 4 (installer step on the box or PC 1 as a signed job) | app owner; the project lead | a green `windows.yml` run, or `tools/ship-app.mjs --dry-run` showing the fetch step satisfied from the new builder |
|
||||
| Q81 | The ship tool pushes master whatever `--branch` says (the 19:30Z push of the release tree to master) | `tools/ship-app.mjs` commit step | operator-visible | 1 | app owner | `--branch X` pushes X; a test on a scratch repo |
|
||||
| Q82 | The download page shows no hash for Windows and Mac, no OTA key fingerprint, no verify command; the stable aliases are not explained (Q50 covers the page; this row is the artefact side) | `miner.html:549-558`, `packaging/ota/publish-public.sh` | user-visible | 1 | site owner | `igneum-downloads.json` carries the fingerprint; the page prints it with `shasum -a 256` and `certutil -hashfile` lines |
|
||||
| Q83 | The update key is one software key on one Mac, not in hardware, against G7's "held in hardware and its policy published" | `packaging/ota/README.md` "Keys", `docs/fud-ledger.md` G7 | operator-visible (security) | 3 (a YubiKey or Secure Enclave signer behind `igneum-ota-sign`, the policy page) | cryptographer, app owner | a release signed with the hardware key installs; the key policy is a public page; the software seed is retired |
|
||||
| Q84 | An activation height in the manifest outlives its activation and turns every later update urgent (the Q4 class, publisher side) | `publish-manifest.sh --activation-height` | operator-visible | 0.5 | app owner | the publisher refuses a height at or below the live DAA and drops a stale carried-over one |
|
||||
|
||||
**First 60 seconds, fresh Windows PC** (from the code and the READMEs): download `Igneum-Miner-Setup-0.3.14.exe` from `/miner`, SmartScreen interstitial "Windows protected your PC" (More info, Run anyway), per-user install with no UAC, "Start Igneum Miner now" (`Igneum-Miner.iss:86`), tray "Igneum Miner: starting", Welcome, Cards (detect runs on entry, `app.js:641`), Address, the one-time key sheet, "Start mining"; between 20 and 50 s in, a UAC prompt for the `igneumd.exe` inbound firewall rule arrives on top of whatever screen is up (`ota.rs:981`; declined = the node dials out, never asked again); closing the window gives "Igneum Miner keeps mining" (`host.cpp:354`). Fresh Mac: DMG, drag to Applications, Gatekeeper refuses the ad hoc app (macOS 15 has no Right-click > Open; the user must find Privacy and Security > Open Anyway, approximate), "Starting the engine" (`IgneumMiner.swift:146`), quarantine cleared from the bundle (`main.rs:78`), the same three screens, a menu-bar item with "Pause mining". Both tiers of card see the same screens; an 8 GB card gets two vote identities and a 24 GB card eight (`detect.rs`), which no screen explains (the "a card runs several" line lives on the site, not the app).
|
||||
|
||||
## 4. Cross-cutting
|
||||
|
||||
### 4.1 Accessibility (estimated from HTML and CSS)
|
||||
|
||||
| Surface | Contrast | Keyboard | Screen reader | Motion |
|
||||
|---|---|---|---|---|
|
||||
| Ember dark | bone 17.4:1, ash 7.0:1, ember 5.6:1, molten 11.0:1 on obsidian: passes | `:focus-visible` ring (`app.css:66`), rail arrow keys, Escape on the card and quit (`app.js:681-688, 979`) | `aria-live="polite"` strip (`index.html:50`), `role="dialog" aria-modal` (`:419`), `radiogroup` with `aria-checked` (`:194`), labelled switches and sliders; gaps: the blocks canvas `aria-hidden` with no text alternative (`:241`), the drawer a `div` with no landmark, `user-select:none` on body (`app.css:36`) | reduced motion honoured everywhere (`app.css:67, 113, 142, 148, 157, 420, 437, 466`; `app.js:1137`) |
|
||||
| Ember light | molten `#B8731F` on white 3.81:1 and on bone 3.38:1; ember `#E04A14` on bone 3.61:1: fails AA on the key, the address and every live state (Q12) | same | same | same |
|
||||
| Site | dark passes (ash 6.97:1, ember 5.64:1 on obsidian); the homepage light section fails on links 3.08:1 and the eyebrow 3.97:1 (Q56) | skip link and one focus ring on every page (`nav.html:1`, `head.html:22`); explorer rows clickable with the hash link as the keyboard target | `aria-live="polite"` on strips refreshed every 2 and 10 s (Q54); `aria-live="off"` on the homepage live strip, correct; the DAG canvases `aria-hidden` with no alternative | reduced-motion rule on every page (`head.html:62`); `live.html` caps 30 fps and stops off screen; the homepage does not (Q55) |
|
||||
| Hub | amber `#FFB35C` and red on graphite pass (approximate from the tokens, `relay/ui.html:17, 65`) | buttons and tabs are native elements; no skip link | no live regions; the console is an operator tool | none needed |
|
||||
| Pool page | default light tokens, not audited beyond the formats | native | none | none |
|
||||
|
||||
### 4.2 Copy law
|
||||
|
||||
The em dash character, counted per shipped file: 0 in every `site/*.html` and `site/partials/*.html`, 0 in `app/igneum-app/ui/index.html` and `app.js`, 0 in `relay/index.html` and `relay/ui.html`, 0 in `pool/web/index.html`, 0 in `tools/community/discord-hooks.mjs`, 0 in `app/mac/*.swift` and `app/windows/host.cpp`, 0 in every `app/igneum-app/src/*.rs` user-facing string, 0 in the wallet's `ui/`. The rule holds everywhere that ships.
|
||||
|
||||
Forbidden strings (`site/forbidden-strings.txt`, 25 patterns) per shipped file, with what each hit is:
|
||||
|
||||
| File | Pattern | Count | What it is |
|
||||
|---|---|---|---|
|
||||
| `site/index.html` | `dl\.igneum` | 2 | the public download aliases (`:514-515`); the pattern is meant for the token path (Q52) |
|
||||
| `site/miner.html` | `dl\.igneum` | 5 | the same aliases and the HiveOS installation URL (`:549-557`) |
|
||||
| `site/wallet.html` | `dl\.igneum` | 1 | the wallet alias (`:371`) |
|
||||
| `site/downloads.json` | `dl\.igneum` | 1 | the `base` field |
|
||||
| `relay/ui.html` | `Hetzner` | 1 | the "Hetzner network" card title (`:377`); private console, rule does not apply, but one word |
|
||||
| `app/igneum-app/ui/app.js` | `the project lead` | 1 | a code comment (`:270`); the ui bundle is outside every check (Q22) |
|
||||
| `tools/community/discord-hooks.mjs` | `MacBook`, `\+0100`, `tailscale`, `ts\.net`, `dl\.igneum` | 1, 1, 1, 1, 2 | the guard's own regexes (`:51, 59, 61`), the read-only fleet URL (`:31`), and `ts\.net` matching "stats.net" in a template (`:255`): false positives, none posted |
|
||||
| `pool/README.md` | `Hetzner`, `/opt/igneum`, `CLAUDE\.md` | 5, 2, 1 | the deploy section; not a site page, but it would fail the scrub if copied |
|
||||
|
||||
Every other pattern is 0 on every file. Two-beat antithesis and aphorisms in shipped copy: the headlines listed in Q62, the two Ember lines in Q23, the four Discord incident sentences in Q78.
|
||||
|
||||
### 4.3 Consistency of numbers across surfaces
|
||||
|
||||
| Quantity | Site home | Site live | Explorer | Ember | Hub | Discord | Pool page |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| Hash rate | `fmtHash`: 1 decimal, MH/s or GH/s only, "pending" when null (`index.html:719`) | 1 decimal, kH to TH, unit in `<small>` (`live.html:329, 652`) | 1 decimal, kH to PH, integer under 1 kH/s (`lib/explorer.mjs:62-67`) | 1 decimal MH/s (`app.js:1450`) | 2 decimals GH/s, 1 decimal MH/s (`relay/ui.html:246`) | 1 decimal MH/s, 2 decimals GH/s (`discord-hooks.mjs:101-109`) | 2 decimals once scaled (`web/index.html:152`) |
|
||||
| Blocks per second | `toFixed(1)` "blocks / s" (`index.html:375, 722`) | `toFixed(2)` (`live.html:228, 650`) | inverted: "1.2 s per block" (`explorer.html:331`) | not shown | `toFixed(2)` (`relay/ui.html:370`) | 6-hour window | not shown |
|
||||
| Height, counts | `toLocaleString('en-GB')` "chain block" (`index.html:377`) | `compact()` "110.0k blocks" (`live.html:229, 651`) | full digits "Height" (`explorer.html:212, 328`) | `withCommas` | `nf()` separators | `en-GB` separators (`:99`) | raw integers, no separators (`web/index.html:163, 172`) |
|
||||
| Difficulty | not shown | `compact()` "126.81M" | "126.81 M" with a space | `compact()` | `nf()` | 6-hour delta | `toLocaleString()` with no locale |
|
||||
| Miner count | "vote keys active in 10 min (a card runs several)" (`index.html:374`) | "Identities" (`live.html:232`), table heading "Miner" (`:254`) | not shown | "ids" per card | "identities 10 min" | "56 vote keys active" and "93 voters" in one post (Q77) | "Miners / workers" |
|
||||
| IGN | not shown | not shown | reward 4 decimals, supply integer | 4 decimals (`app.js:1364`) | | 2 decimals (`fmtIgn`) | 4 on tiles, 6 in tables |
|
||||
| Price | none anywhere; "pounds a day" once (`index.html:510`); Ember has a `priceNow()` with a devnet zero | | | | | | |
|
||||
| Finality | "#N" or "paused" (`index.html:879`) | "#N, 3 min ago" plus the bar (`live.html:404, 512`) | none | "#N · 1 h ago" (`app.js:1350`) | "active" or "paused" | "checkpoint 6,842 at 55.x% of weight, 1 h ago, 93 voters" | none |
|
||||
|
||||
Row **Q85** (cross-cutting, 2 hours, site owner with the app owner): one shared formatter module (`site/lib/format.mjs`) for hash rate (1 decimal, unit ladder kH to PH), counts (`en-GB` separators, never `compact()` for the same quantity shown in full elsewhere), blocks per second (2 decimals, never inverted), difficulty (one spelling), IGN (4 decimals on tiles, 6 in tables), and one word for a vote identity ("vote key" on every surface); the app, hub, Discord and pool copy the same table as a Rust and a JS constant, with a test that renders the fixture values identically. Gate: the fixture renders byte-identical on all seven surfaces. Severity: the project lead-visible (he reads three of these side by side every evening).
|
||||
|
||||
### 4.4 Error states, one table
|
||||
|
||||
| Condition | Ember | Site home | Site live | Explorer | Hub | Discord | Wallet | Pool |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| Node down | "the node is not running", pill "node failed", "waiting for the node to sync" | "simulated preview · observer offline" after 30 s | "OFFLINE", "observer offline, last update N" | "observer offline, last known" | card red "silent 4h"; "observer STALE N s" | `observer_silent` after 180 s (not live) | "no node", "waiting for a node" | "difficulty 0, DAA 0", no word (Q68) |
|
||||
| API or relay unreachable | "The update failed." (Q11) | "live feed unavailable" | "api unreachable, scene frozen" | error on first load, silent after (Q10) | "offline: Failed to fetch", stale panels (Q6) | n/a | "The API did not answer: ..." (address page) | "stats unavailable" |
|
||||
| Finality paused | nothing (Q2) | "paused" counter, "final" label (Q5) | bar "finality paused: N% silent" (Q1) | nothing (Q10) | tile "paused", no cause (Q6) | nothing (not live), `preexisting` (Q7) | "paused" or "not checkpoint-checkable" | nothing (Q73) |
|
||||
| GPU lost | "Card removed: <name>. Its worker stopped." | | | | card telemetry drops | | | member drops from the count |
|
||||
| Job hung | "job running" notice | | | | "running" forever (Q40) | | | |
|
||||
|
||||
### 4.5 First-run and update flow
|
||||
|
||||
Section 3.10 carries both in full. The two facts to carry forward: a fresh machine on either OS meets two warnings before the first screen (Q3), and the update flow is one software key, one builder, one manifest, with a stale activation height that makes every update urgent (Q4, Q84) and no attestation by anyone outside the project (the frontier 3.7 baseline).
|
||||
|
||||
## 5. Already best in class (honest both ways)
|
||||
|
||||
1. Ember Tune: automatic per-card efficiency tuning with the money consequence on screen, measured 37 to 41 percent more hashes per watt on PC 1 (`ember-tune.md:254, 271`). No miner console ships it.
|
||||
2. The update path: Ed25519-signed manifest verified before parsing, sha256 of the artefact, staged swap, safe-moment rule, machine slots, automatic rollback after two failed starts (`ota.rs`, `manifest.rs`). HiveOS and NiceHash update without a safe-moment rule (approximate).
|
||||
3. The wallet verifies finality certificates itself with the node's consensus code (`finality.rs:101-121`); Touch ID is bound to the exact quote (`engine.rs:539`). No mainstream EVM wallet does either (approximate).
|
||||
4. The site: self-hosted fonts with fallback metrics, zero third-party scripts beyond two pinned modules, no analytics, no banner, HSTS preload, honest live states, a browser-side BLS verifier on the homepage (`verify/core.js`), a dated journey with gates and a public ledger of every criticism.
|
||||
5. The block page: parents, children, mergeset blues and reds, selected parent, shards with prover and lag, checkpoint fractions and certificates (`block.html:285-325`).
|
||||
6. The pool: every share CPU-verified at the node's own engine; the member's vote key in every header it hashes (`node.rs:47`), so pool share never becomes vote share.
|
||||
7. Discord: full sha256 and verify commands in every release post; a forbidden-string guard on every post; incident texts that say who is affected per tier.
|
||||
8. The hub: one phone-first page with every PC's miner, node, app and OTA state fused on one card, stopped and silent as different words, drop-anywhere uploads.
|
||||
9. The relay's wake long-poll: dependency-free, fake-clock tested, 30 a minute per IP.
|
||||
10. The ship pipeline: eleven checked, resumable steps with secrets scrubbed from every line and a signed inputs manifest verified in CI before any build (G13).
|
||||
|
||||
## 6. Open questions and what could not be run
|
||||
|
||||
- The pause's end was not observed: at 20:05:03Z (`live_state.updated_at`) no lock after 6842 had formed. Lane 3 expects the frozen table to expire at DAA 216,402, about 20:40Z. Whoever reads this after 21:00Z should add the resume time and whether any surface showed it.
|
||||
- Which commit the deployed relay runs is unknown from this tree (`relay/README.md` does not name it); Q8 assumes master. The deployed site is master by the GitHub integration.
|
||||
- Lighthouse numbers are estimates from source; no browser run was allowed. The gate for every site row is a real run.
|
||||
- The comparator claims are from memory; where a repository or page is named it is cited, otherwise approximate. The Stratum v2 reference was not cloned.
|
||||
- The fleet and node sections depend on the gpu-fleet worktree and the vendor forks, read tonight; the fleet's own page was not opened.
|
||||
- Hours are agent hours (the project lead's rule); Q3 and Q80 also need the project lead (certificates, GitHub billing).
|
||||
|
||||
## 7. Summary for the coordinator
|
||||
|
||||
Every shipped system was audited against named comparators from the code and the design docs, with 95 ledger rows (ids Q1 to Q105), each with a file or screen, a severity, hours, an owner and a gate; the top ten total 40 agent hours; every row summed is about 230 agent hours. The three findings that matter most: (1) during tonight's two-hour finality pause no public or operator surface could say why, because the node's `finality_reason` is dropped at `tools/observer/observer.mjs:720`, the frozen-table share that held the pause is reported by nothing, Ember has no pause state at all, the homepage drew "final", the explorer showed nothing and Discord was not live; one forced pause on Devnet 2 with a screenshot per surface is the gate (Q1, Q2, Q5, Q6, Q7, Q10). (2) A fresh machine meets two warnings before the first screen (ad hoc Mac signature with a quarantine strip, unsigned Windows installer plus a UAC prompt 20 to 50 s in), and the Windows installer cannot be rebuilt at all tonight because GitHub-hosted runners are blocked on billing (Q3, Q80). (3) The update path treats any passed activation height as "within 1,800 blocks", so every update since 0.3.14's manifest has been urgent, skipping the safe-moment rule (the PC 1 install under a measurement job); the same path has no finality input and is one software key, one builder, one manifest, which is the baseline frontier 3.7 moves from (Q4, Q83, Q84). Rules for main: the observer and every API carry `finality_reason`, `held_by` and `finality_provisional` before the public testnet; the `tools/ci` forbidden-string check covers `app/igneum-app/ui/`; one formatter table for every number on every surface (Q85).
|
||||
|
|
@ -11,7 +11,7 @@ A proof-of-work chain mined on consumer GPUs, where the same cards prove every I
|
|||
| Item | What it is | Status today |
|
||||
|---|---|---|
|
||||
| Proofs for your chain | Your batches or blocks proven by Igneum's GPU prover network and returned to the address you name | Designed. No shard has been proven on a card yet |
|
||||
| Price in dollars | Jobs are priced in dollars per proof. At launch the fee is paid on your own chain, in your currency, to a payout contract keyed by miner address, because Igneum cannot yet see your chain. Settlement in the coin follows when the proof bridge exists | Spec section 5.4 |
|
||||
| Price in dollars | Jobs are priced in dollars per proof, at or above the subsidy the prover forgoes while it proves, which is a formula with network hash as the input, never a fixed number: per shard, (card hash ÷ network hash) × 0.8 × 31.688 IGN × shard seconds, plus electricity (under a cent per billion cycles on every card). The price falls as one over network hash: at the devnet's 1.16 GH/s a quote is 100 to 300x the published market rate, and a card proving beside its miner is competitive near 100 GH/s (the Horizon economy lane, `docs/analysis/horizon/economy-and-utility.md` sections 3.1 and 4.1, 6 October 2026; approximate beyond the one card measured; ledger E20). At launch the fee is paid on your own chain, in your currency, to a payout contract keyed by miner address, because Igneum cannot yet see your chain. Settlement in the coin follows when the proof bridge exists | Spec section 5.4 |
|
||||
| Delivery rule | A job is claimed with a bond that is slashed on a late or bad proof; a job nobody proves by its deadline expires and refunds in full. At launch your chain's own bond and slashing apply to the miner who claimed the job | Design document, execution layer, sections 5.2 and 6. Bond size and timeout are open (O-5.6) |
|
||||
| A versioned interface | Jobs run against the `ProofSystem` trait, version 1 of which is SP1. A later version is a release with its own test-vector set and a three-month overlap, so your integration survives a prover swap | Design document, execution layer, section 5.6 |
|
||||
| Verification you can run | A job proof is a single proof your contract verifies on your own chain; Igneum's own segment proofs recursively verify it, so no relayer or committee is in the path | Designed |
|
||||
|
|
@ -25,7 +25,7 @@ One row per route, so operator income and protocol income never blur. Rows 1 to
|
|||
| 1. Emission, per block | IGN, new coins on the published schedule | 80% the block's miner, 20% the proving pool for the provers of that block | None | None. Implemented in consensus on the devnet |
|
||||
| 2. Base fee, both gas dimensions | IGN | Nobody | The base fee the chain sets per block | All of it. Implemented on the devnet |
|
||||
| 3. Priority fee | IGN | 80% the block's miner and provers; 20% the apps whose code ran, per call frame | The tip the sender sets | The share of any frame in an unregistered contract. Implemented on the devnet |
|
||||
| 4. External job, at launch | Your currency, on your chain | The miner who delivered, through a payout contract keyed by miner address | Priced in dollars per proof; your chain's own bond and slashing apply | None; Igneum cannot see the payment. Designed |
|
||||
| 4. External job, at launch | Your currency, on your chain | The miner who delivered, through a payout contract keyed by miner address | Priced in dollars per proof, at or above the subsidy the prover forgoes (the formula in network hash in the "Price in dollars" row, never a fixed number); your chain's own bond and slashing apply | None; Igneum cannot see the payment. Designed |
|
||||
| 5. External job, after the proof bridge | IGN, on Igneum | 90% the provers who delivered | The job fee | 10%. Designed, phase two |
|
||||
| 6. The official client's dev fee | IGN | The project, as operator income, never the protocol | 1 block template in 100 requested with the dev address; off with one flag | None. Implemented, measured on a test network 4 October 2026 |
|
||||
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
# Reproduced: Igneum Miner 0.3.14 (node 4c6b129d75c3d77a3689d22f1e1dc721b556aebb, app a90f6a5371ed19c62511255a4713eb36191848b9)
|
||||
|
||||
06 October 2026, 20:03 UTC on igneum-build-1 by `infra/build-server/repro/rebuild-on-box.sh` (driven by `tools/repro/rebuild-release.sh`): a clean clone of the fork at the node commit on branch `release-0.3.14-node` under a clean clone of the repo at the app commit, 2 independent clean passes per target in one target path each (no sccache, SOURCE_DATE_EPOCH 1791305478, TZ UTC), rustc 1.99.0, x86_64-w64-mingw32-gcc-posix (GCC) 13-posix, clang 18.1.3, glibc 2.39. Shipped hashes read from the public downloads (no token) where marked. Whole run 3 s; log `/srv/builds/_repro/0.3.14/rebuild.log`.
|
||||
06 October 2026, 20:20 UTC on igneum-build-1 by `infra/build-server/repro/rebuild-on-box.sh` (driven by `tools/repro/rebuild-release.sh`): a clean clone of the fork at the node commit on branch `release-0.3.14-node` under a clean clone of the repo at the app commit, 2 independent clean passes per target in one target path each (no sccache, SOURCE_DATE_EPOCH 1791305478, TZ UTC), rustc 1.99.0, x86_64-w64-mingw32-gcc-posix (GCC) 13-posix, clang 18.1.3, glibc 2.39. Shipped hashes read from the public downloads (no token) where marked. Whole run 382 s; log `/srv/builds/_repro/0.3.14/rebuild.log`.
|
||||
|
||||
| Artefact | Shipped sha256 (source) | Box pass A | Box pass B | A vs shipped | A vs B | Reason for a DIFFER |
|
||||
|---|---|---|---|---|---|---|
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
# Reproduced: Igneum Miner 0.3.15 (node 713ef876073d3661e9b48d2ead9a515afd1b2156, app 563485b769868ee34a530f1c40f7419109cc493f)
|
||||
|
||||
06 October 2026, 20:03 UTC on igneum-build-1 by `infra/build-server/repro/rebuild-on-box.sh` (driven by `tools/repro/rebuild-release.sh`): a clean clone of the fork at the node commit on branch `release-0.3.15-node` under a clean clone of the repo at the app commit, 2 independent clean passes per target in one target path each (no sccache, SOURCE_DATE_EPOCH 1791312828, TZ UTC), rustc 1.99.0, x86_64-w64-mingw32-gcc-posix (GCC) 13-posix, clang 18.1.3, glibc 2.39. Shipped hashes read from the public downloads (no token) where marked. Whole run 2 s; log `/srv/builds/_repro/0.3.15/rebuild.log`.
|
||||
06 October 2026, 20:21 UTC on igneum-build-1 by `infra/build-server/repro/rebuild-on-box.sh` (driven by `tools/repro/rebuild-release.sh`): a clean clone of the fork at the node commit on branch `release-0.3.15-node` under a clean clone of the repo at the app commit, 2 independent clean passes per target in one target path each (no sccache, SOURCE_DATE_EPOCH 1791312828, TZ UTC), rustc 1.99.0, x86_64-w64-mingw32-gcc-posix (GCC) 13-posix, clang 18.1.3, glibc 2.39. Shipped hashes read from the public downloads (no token) where marked. Whole run 480 s; log `/srv/builds/_repro/0.3.15/rebuild.log`.
|
||||
|
||||
| Artefact | Shipped sha256 (source) | Box pass A | Box pass B | A vs shipped | A vs B | Reason for a DIFFER |
|
||||
|---|---|---|---|---|---|---|
|
||||
|
|
|
|||
|
|
@ -449,6 +449,24 @@ Sweep (5 October 2026, evening): stated. `site/litepaper.html`, Economics first
|
|||
|
||||
---
|
||||
|
||||
### P24. "20 percent of emission to provers" without the caveat that consensus does not verify the proof
|
||||
"Your 20 percent pays whoever submits a proof record whose statement matches the node's own execution. The node never checks the SP1 proof behind it. So the block producer, who writes the record, can claim shard pay with a false proof today, in proportion to its hash. Say so next to the 20 percent."
|
||||
|
||||
Status: Conceded, stated (6 October 2026, evening; the Horizon security lane `docs/analysis/horizon/consensus-security.md` finding 3 and its proposal 1): `site/litepaper.html`, Economics, the 20% proving-pool row carries the caveat: consensus does not yet verify the carried proof, it checks the record's statement against native execution and its signature, so today a block producer could claim shard pay with a false proof (ledger P21; the in-consensus verifier is the 0.3.16 fix).
|
||||
|
||||
Answer: Correct; it is P21's finding read from the payer's side. The fix is the security lane's proposal 1: a record whose aggregated segment proof does not verify against the pinned aggregator key is invalid in consensus; the per-shard v0 record stays payout-only until then and is capped at the exclusive window. Consequence per tier: no honest miner loses anything today, since the pool is paid per valid record and the devnet's producers are the project's; the risk is a dishonest producer at the public testnet, which is why the fix lands in 0.3.16, before it.
|
||||
|
||||
Evidence: `docs/analysis/horizon/consensus-security.md` finding 3 and proposal 1, 6 October 2026; spec 07 7.7 item 4 and 7.8 item 8; ledger P21.
|
||||
|
||||
### P25. Unclaimed pool credit is stranded in the escrow
|
||||
"Spec 5.3 pays the first valid proof included in a block; 7.7 refuses a record older than 600 chain blocks; 7.8 says an unproven segment's aggregator share stays in the escrow. Nothing says what happens to the shard credit nobody claims. It sits there for ever. A ten-day refusal strands millions of IGN."
|
||||
|
||||
Status: Conceded, stated (6 October 2026, evening; the Horizon economy lane `docs/analysis/horizon/economy-and-utility.md` section 4.2 and proposal 2): `site/litepaper.html`, Economics, the 20% proving-pool row carries the note: unclaimed pool credit is today stranded in the escrow, no rule returns it; the fix rolls an unproven shard's credit into the next proven segment's pool (0.3.16).
|
||||
|
||||
Answer: Correct. The lane's simulation puts a ten-day refusal at 5.5 M IGN stranded (section 4.2: the pool share paid falls to 0.67 with 0.33 stranded during the refusal, 182,387 IGN a day averaged); the devnet already burns the coinbase's 20% output at an unspendable script, so nothing is lost that was ever claimable today. The rule: an unproven shard's credit rolls forward into the next proven segment's pool instead of sitting in the escrow. Consequence per tier: a prover sees a larger pool after a refusal instead of a smaller one; nothing changes for a miner or a pool user.
|
||||
|
||||
Evidence: `docs/analysis/horizon/economy-and-utility.md` section 4.2 (the stranded share) and proposal 2, 6 October 2026; spec 5.3, 7.7 item 3, 7.8 item 7.
|
||||
|
||||
## 4. Economics and the coin
|
||||
|
||||
### E1. Hard cap plus burn is a security budget cliff
|
||||
|
|
@ -538,6 +556,24 @@ Answer: Correct. The whole public proving market is three to four orders of magn
|
|||
|
||||
Evidence: `docs/analysis/horizon/frontier.md` section 3.11 and the summary (section 7), 6 October 2026; the tracker (ethproofs, "sub-half-cent" fields, September 2026, secondary); the emission schedule (31.688 IGN a block in year 1).
|
||||
|
||||
### E20. "Proofs at the cost of power" is the electricity, not the price
|
||||
"Your cards' electricity is cheap, fine. The price a prover must charge is the lottery income it gives up while it proves, and that scales as one over the network's hash. At your devnet's 1.16 GH/s every quote is a hundred times the market. 'Marginal cost close to power' and 'priced in dollars per proof' are both unconditioned."
|
||||
|
||||
Status: Conceded, stated (6 October 2026, evening; the Horizon economy lane `docs/analysis/horizon/economy-and-utility.md` sections 3.1 and 4.1, proposal 3): `site/litepaper.html`, The problem ("a supplier whose electricity cost is close to power and whose price is the subsidy it forgoes, which falls as the network's hash grows"), Building on Igneum ("Proofs priced by the subsidy forgone", with the price as a formula in network hash: per shard, card hash over network hash x 0.8 x 31.688 IGN x shard seconds, plus electricity; 100 to 300x the published market rate at 1.16 GH/s, competitive near 100 GH/s beside the miner, approximate beyond the one card measured), the payment-routes row 4 ("at or above the subsidy the prover forgoes ... never a fixed number"), For miners, Economics and the questions list. The price is published as a formula, never a number. Supersedes P6's stated sentence ("a supplier whose marginal cost is close to power"), which the text check now carries in the conditioned form.
|
||||
|
||||
Answer: Correct. Electricity is under a cent per billion cycles on every card; the price is `h/N` x the subsidy per shard. At 100 GH/s a card proving alone is at 1 to 3x the published market and a card beside its miner at 0.2 to 0.4x, the only row where Igneum undercuts the market, and it rests on the 4% hash loss measured on one card. Consequence per tier: at launch a home miner earns more hashing than proving for outsiders at any card size; a rig the same; the proving market is upside for the fleet as a whole only as network hash grows.
|
||||
|
||||
Evidence: `docs/analysis/horizon/economy-and-utility.md` sections 3.1 (the cost model), 4.1 (the table per card and scale) and proposal 3, 6 October 2026; the Boundless median of USD 0.21 per billion cycles (`developer-adoption.md` 2b, approximate).
|
||||
|
||||
### E21. The dev fee is 1 percent of the producer share, and the funding plan's ceiling took all rewards
|
||||
"Your funding plan says the 1 percent fee could be USD 48,000 to 963,000 a year on year-one rewards of 963 million IGN. The fee template only moves the producer payout. The pool is paid per record. Your ceiling is a quarter too high, and the litepaper never says which share the fee is of."
|
||||
|
||||
Status: Conceded, stated (6 October 2026, evening; the Horizon economy lane `docs/analysis/horizon/economy-and-utility.md` section 4.4 and proposal 7): `site/litepaper.html`, the payment-routes row 6 and the Ember section read "default-on, switchable, 1 percent of the producer share"; `docs/plans/funding.md` section 4's ceiling is 1 percent of the producer share, USD 38,520, 154,080 and 770,400 a year at USD 0.005, 0.02 and 0.10 per IGN with every miner on the official client, corrected from 48,000, 193,000 and 963,000.
|
||||
|
||||
Answer: Correct. The fee block carries the dev address in the producer output only; the 20% pool output and the per-record escrow are untouched, so the fee's base is 80% of emission. Consequence per tier: a miner on Ember with the fee on gives up 1 in 100 of its own block rewards and nothing of its proving pay; off with one flag, the same on every tier.
|
||||
|
||||
Evidence: `docs/analysis/horizon/economy-and-utility.md` section 4.4 (`devfee_out.md`) and proposal 7, 6 October 2026; the fee measured on a test network, 4 October 2026 (bench-log: 9 fee blocks in 785).
|
||||
|
||||
## 5. Governance and the founders
|
||||
|
||||
### G1. No cryptography team
|
||||
|
|
@ -578,6 +614,8 @@ Evidence: design doc, decisions table row "Who are you?". Journey: `site/journey
|
|||
|
||||
Status: Conceded, stated (5 October 2026, night): `site/litepaper.html`, Governance, "There are no admin keys in consensus"; `site/index.html`, tile "admin keys in consensus". Was: Conceded, wording fix needed.
|
||||
|
||||
Status: Conceded, stated (6 October 2026, night): the home page was redrawn as one statement, the live scene, three facts and the downloads, so the tile "admin keys in consensus" is no longer on `site/index.html`; the sentence stands on `site/litepaper.html`, Governance ("There are no admin keys in consensus"). The 5 October line below is the history.
|
||||
|
||||
Answer: Correct. Consensus has no admin keys. The genesis apps are contracts and each will have an upgrade policy; the bridge's is the one that matters, and if it starts as a committee bridge (E7) it has keys by definition. The development fund contract named in the quote no longer exists: the fund was removed on 3 October 2026. The litepaper's sentence must be scoped to consensus and each genesis contract must publish its key policy before launch.
|
||||
|
||||
Evidence: litepaper "Governance". Fix: overclaims list, item 46.
|
||||
|
|
@ -626,6 +664,15 @@ Evidence: design doc Finality v2, Residual risks bullet 3.
|
|||
|
||||
---
|
||||
|
||||
### G15. Three signalling thresholds, four numbers across the documents
|
||||
"Spec 5.7 says 90 percent for an upgrade, 5.5 says 60 percent for a parameter, the class-change rule says 95 with a floor, the project rules file says 90 and 60, and the litepaper's Mining section says a 90 percent signal turns a spare defence on, which is a class change your own rule sets at 95. Pick one sentence and put it everywhere."
|
||||
|
||||
Status: Conceded, stated (6 October 2026, evening; the Horizon economy lane `docs/analysis/horizon/economy-and-utility.md` proposal 4, lane 3's signalling results): one sentence in `site/litepaper.html`, Governance ("Miners set what genesis leaves open") and Mining ("Miners hold the switch"): miners signal three things at three thresholds, 60 percent of blue blocks over two weeks for a parameter genesis leaves open, 90 percent for an upgrade (new code), and 95 percent with a floor height for a class change. The project rules file carries the same sentence (the coordinator's commit of 6 October 2026). Spec 5.5 (60) and 5.7 (90) agree with it; the 95-with-floor rule is the class v4 cut's P2 rule.
|
||||
|
||||
Answer: Correct. The three numbers are three different things: a parameter is a dial inside rules genesis fixed, an upgrade is new code every node must run, a class change moves the hash itself and so takes the highest bar with a floor height as the backstop against a holdout. The lane's game (section 4.3): a 6 percent holdout costs near zero and buys only delay to the floor; a 30 percent pool holds a veto over upgrades at 90 and over class changes until the floor.
|
||||
|
||||
Evidence: `docs/analysis/horizon/economy-and-utility.md` section 4.3 and proposal 4, 6 October 2026; spec 5.5, 5.7, 5.8; the class v4 cut's status file (gates P1 and P2).
|
||||
|
||||
## 6. Comparisons
|
||||
|
||||
### C1. vs Monero: GPUs were excluded on purpose
|
||||
|
|
@ -642,6 +689,8 @@ Evidence: litepaper "For miners", hardware paragraph. Wording: overclaims list,
|
|||
|
||||
Status: Conceded, stated (5 October 2026, night): `site/litepaper.html`, Mining, "no chip publicly shipped, approximate" and "Monero is precedent, not proof"; `site/index.html`, hero, "a chip gains too little to take your place". Was: Conceded, label needed.
|
||||
|
||||
Status: Conceded, stated (6 October 2026, night): the home page no longer carries the RandomX paragraph; "since 2019 (approximate)" and "precedent, not proof" stand on `site/litepaper.html` (vs RandomX, Mining). The 5 October line below is the history.
|
||||
|
||||
Answer: True. "No chip publicly shipped, approximate" is the defensible phrasing. The argument from Monero is precedent, not proof, and the bounty exists because precedent is not proof.
|
||||
|
||||
Evidence: none. Fix: overclaims list, item 17.
|
||||
|
|
@ -853,6 +902,8 @@ Sweep (5 October 2026, evening): stated. `site/index.html`: hero "See the miner"
|
|||
|
||||
Status: Conceded in part, labelled, stated (5 October 2026, night): `site/index.html`, the sentence "All of it will be on this page, live" is no longer on the page; the proofs feed now ends "Live rows arrive with the public testnet, August 2027" (overclaim 61). Was: Conceded in part, labelled.
|
||||
|
||||
Status: Conceded in part, labelled, stated (6 October 2026, night): the proofs feed moved to the litepaper's proving section with the home-page redesign and reads "Live rows arrive with the public testnet. The public testnet is weeks away: three seed nodes and the public RPC are up, and it opens when the go checklist closes." (X31). The 5 October line below is the history.
|
||||
|
||||
Answer: The tagline plays on "proven" as in ZK proofs and "cupel". The homepage marks every live panel "PREVIEW", "prototype" or "at testnet", which is honest. The sentence "All of it will be on this page, live" is a promise about a future testnet and should say when. The tagline stays; nothing in the ledger depends on it.
|
||||
|
||||
Evidence: `site/index.html`. Fix: overclaims list, item 62.
|
||||
|
|
@ -907,6 +958,8 @@ Sweep (5 October 2026, evening): stated in part. `site/partials/footer.html` on
|
|||
|
||||
Status: Conceded, stated (5 October 2026, night): `site/litepaper.html`, Roadmap phase 6, and `site/journey.json` phase 6, "No listing is arranged, promised or sought by the project". Was: Conceded, fix now.
|
||||
|
||||
Status: Conceded, stated (6 October 2026, night): the home page's journey is no longer shown (the inlined feed remains in the page source); the sentence "No listing is arranged, promised or sought by the project" stands on `site/litepaper.html` Roadmap phase 6 and in `site/journey.json`. The 5 October line below is the history.
|
||||
|
||||
Answer: Correct. No listing is arranged, promised or sought by the project. The phrase comes off the roadmap and the homepage journey.
|
||||
|
||||
Evidence: litepaper "Roadmap", `site/journey.json`. Fix: overclaims list, item 75.
|
||||
|
|
@ -1142,7 +1195,7 @@ Replace with: "A chain built so a chip gains too little to take your place."
|
|||
Replace with: "Benchmark: January 2027"
|
||||
|
||||
61. HP: "All of it will be on this page, live."
|
||||
Replace with: "All of it will be on this page, live, from public testnet in August 2027."
|
||||
Replace with: "All of it will be on this page, live, from the public testnet." (the month was removed on 6 October 2026, X31)
|
||||
|
||||
62. HP economics: "Not one coin to a founder, a fund or a stake."
|
||||
Replace with: "Not one coin of emission to a founder, a fund or a stake. The official client carries a 1% dev fee to the founder's company, as every GPU miner does; any client without it is welcome."
|
||||
|
|
@ -1633,7 +1686,7 @@ Round 2 (5 October 2026, night): the signing half, from block payloads. No RPC e
|
|||
### X15. Remove the founders from a test network and show what continues
|
||||
"'The chain runs without its founders' is a sentence. Take the team's miners, provers, aggregators, seed nodes, observer and site off a running testnet and show what keeps producing blocks, proofs and locks."
|
||||
|
||||
Status: Open, blocked on the public testnet (August 2027 per the litepaper roadmap): next step O-X.2 run at a published time on that testnet, with the protocol already in the Answer below (every project-run node, miner, prover, aggregator and seed stopped, the observer and live page down, 24 hours of blocks per second, proof lag, certificates per hour and a fresh sync from the shipped seed list). Was: Open, experiment scheduled (O-X.2). Sweep (5 October 2026): a public-testnet experiment; not runnable before it exists.
|
||||
Status: Open, blocked on the public testnet (weeks away, when the go checklist closes; see X31): next step O-X.2 run at a published time on that testnet, with the protocol already in the Answer below (every project-run node, miner, prover, aggregator and seed stopped, the observer and live page down, 24 hours of blocks per second, proof lag, certificates per hour and a fresh sync from the shipped seed list). Was: Open, experiment scheduled (O-X.2). Sweep (5 October 2026): a public-testnet experiment; not runnable before it exists.
|
||||
|
||||
Answer: Correct, and it is the right test for the litepaper's sentence (Governance). The test: on the public testnet, at a published time, stop every node, miner, prover, aggregator and seed node the project runs, take the observer feed and the live page down, and record for 24 hours: blocks per second, proof lag, certificates per hour, and a fresh node syncing from the seed list in the client (spec 10.6). What continues is what the sentence may claim. Dependencies the test will expose: the seed list, the release key (G7), the reference pool, the VDF evaluators (every node ships one, 4.5) and the founders' own hashrate share (E2).
|
||||
|
||||
|
|
@ -2166,6 +2219,33 @@ Answer: Correct at discovery; fixed before this document was written. The eviden
|
|||
|
||||
Evidence: the commits above. Experiment: `curl https://igneum.network/api/live` holds no address, no 64-hex key hash and no payout address.
|
||||
|
||||
### X31. The public testnet dated "August 2027" on the site
|
||||
"The litepaper's For miners section said 'Pools and the public testnet are August 2027', the proving section said 'Live rows arrive with the public testnet, August 2027', the roadmap's phase 5 read 'Aug to Oct 2027' and the home page's journey carried the same row. igneum-testnet-1's genesis is final, three seed nodes and the public RPC are up, and the testnet opens when the go checklist (docs/plans/testnet-go.md) closes, which is weeks away."
|
||||
|
||||
Status: Fixed, stated (6 October 2026, night, the owner's decision): every mention of the month is gone from the site. The sentence everywhere is "The public testnet is weeks away: three seed nodes and the public RPC are up, and it opens when the go checklist closes." (`site/litepaper.html` For miners and the proving section, the roadmap row 5 reads "Weeks away: when the go checklist closes", `site/journey.json` phase 5 and the home page's inlined journey carry the same row). No calendar month is given for the testnet; the owner gives one if he wants one.
|
||||
|
||||
Answer: The date was the plan of 3 October 2026 and the chain overtook it: the testnet genesis was fixed on 5 October, the three seeds and rpc.testnet.igneum.network are up, and the remaining work is the go checklist. Rows that quoted the month (X3, O-X.2's blocker note, overclaim item 75's replacement text) read the new sentence by reference to this row.
|
||||
|
||||
Evidence: `docs/plans/testnet-go.md`; `docs/igneum-testnet` notes (genesis 87617621..., seeds seed1 to seed3.testnet.igneum.network, public RPC). Checked by `tools/ci/ledger-text-check.mjs` (the X3 and X31 rows).
|
||||
|
||||
### X32. The roadmap carried calendar months beside a testnet that is weeks away
|
||||
"After X31 the roadmap read phase 4 'Apr to Jul 2027' and phase 6 'Nov 2027' with phase 5 'weeks away' between them, and phases 1 to 3 carried 'Oct to Nov 2026', 'Nov 2026 to Jan 2027' and '20 nodes by Mar 2027'. A reader spots the contradiction at once."
|
||||
|
||||
Status: Fixed, stated (6 October 2026, night, the owner's ruling): every calendar month is out of the roadmap. Each phase is worded by its gate in the shape phase 5 has: 1 "Under way; closes when the specification is out for external review", 2 "Under way; closes at its gate", 3 "Live since 3 October 2026; closes at its gate", 4 "Closes when the finality design passes external review and one rollup signs for the testnet", 5 "Weeks away: when the go checklist closes", 6 "After the testnet has passed its gate: 1,000 independent miners for 30 days and rollup proofs on time". The order is unchanged. The roadmap's lead reads "Six phases from specification to a fair launch, with four public gates and no calendar dates: each phase closes at its gate." (`site/litepaper.html` Roadmap, `site/journey.json`, the home page's inlined journey). The 3 October 2026 start of the devnet is a fact, not a target, and stays.
|
||||
|
||||
Answer: Dates were the plan of 3 October 2026; the gates are the plan. "Dates slip. Gates do not." was already the roadmap's own sentence, and the roadmap now says only the gates. The owner gives a month if he wants one.
|
||||
|
||||
Evidence: `site/litepaper.html` Roadmap; `site/journey.json`; X31.
|
||||
|
||||
### X33. The public benchmark dated "January 2027"
|
||||
"After X31 and X32 the litepaper still said the public benchmark with a leaderboard 'is January 2027' (For miners) and 'ships in January 2027' (Questions miners ask), a calendar month beside a roadmap that names none."
|
||||
|
||||
Status: Fixed, stated (6 October 2026, night, the owner's ruling): both sentences read "The public benchmark with a leaderboard ships with the public testnet." (`site/litepaper.html`, For miners and Questions miners ask). The gate is one the reader already knows from the roadmap's phase 5. The 5 October answers in M1, X1 and the overclaims list that name January 2027 are the history of the plan and stay as written.
|
||||
|
||||
Answer: The month was the plan of 3 October 2026. The benchmark tool is part of the testnet's go checklist, so it ships with the testnet; no month is given for either.
|
||||
|
||||
Evidence: `site/litepaper.html`; `docs/plans/testnet-go.md`; X31, X32.
|
||||
|
||||
## Status updates, 4 October 2026 (round 4)
|
||||
|
||||
- **F21** (the long-partition fork). Extended: a side locks alone when its own share of its own table reaches two thirds, at `t = W (2/3 - s) / (1 - s)`: 50/50 at 2,400 DAA s on the devnet (about 40 minutes), 10 days on mainnet; the 60 side of 60/40 at 1,200 DAA s (20 minutes), 5 days; the ledger's measured `W/(3R)` = 200 s is this formula at s = 1/2. At HEAD a second certificate at an index is kept, logged and ignored (`processes/finality.rs:650-655, 661-666`) and `fork_choice_lock` (`:886-901`) pins the node. Public text: `site/litepaper.html:511` says a third of the blocks is needed to split finality in a partition; the partition alone does it. Replacement sentence in `docs/review/round-4-2026-10-04.md` section 1 (b). Review id R4.1.5.
|
||||
|
|
|
|||
|
|
@ -241,6 +241,12 @@ igneum-miner c5b48910...): DIFFER, and the binaries say why: the shipped pair ne
|
|||
box's 06211d55... of 19:00Z (the box build needs GLIBC_2.39, which HiveOS cannot run; the zig lane is right for that
|
||||
package). The Windows exes have no shipped hash yet (the public installer is still 0.3.14). `docs/evidence/reproduced/0.3.15.md`.
|
||||
|
||||
Re-run 20:13 to 20:22Z with sccache REALLY off (the build-server agent's finding: an empty `RUSTC_WRAPPER=` is read by cargo as
|
||||
unset and falls back to the box's config, so the first runs' passes could take hits; now `RUSTC_WRAPPER=/usr/bin/env`, a
|
||||
pass-through, and `SOURCE_DATE_EPOCH` = the author time of lib.sh `bs_sde`): both versions, all four artefacts, A vs B
|
||||
**MATCH** with the same hashes as above (0.3.14: 03f35e05, 900c1f0b, 166e604e, fefd266c; 0.3.15: 1f1b6eee, a34e0a56,
|
||||
9b377455, 65b30edd); passes of 73 to 177 s with the two repros side by side on the two slots. The MATCH rows stand on their own.
|
||||
|
||||
What it means: the box is deterministic for a given commit and path, so a release built on it can be checked by anyone with the
|
||||
same toolchain by rebuilding and comparing; the shipped 0.3.14 and 0.3.15 Linux bytes came from the Mac's zig lane with a
|
||||
build clock inside and cannot be reproduced anywhere, and the Windows 0.3.14 exes carried a PE timestamp. The reason to ship
|
||||
|
|
|
|||
|
|
@ -55,7 +55,7 @@ The unfunded half is the half that comes after the chain exists and before and j
|
|||
|
||||
## 4. What the 1% fee could be, and why it is not counted
|
||||
|
||||
The fee is 1% of rewards on the official client. Rewards in year one are 963 million IGN (ramp included, `docs/analysis/security-budget.md`). At the low, base and high price inputs of that analysis the fee's ceiling, with every miner on the official client, is USD 48,000, 193,000 and 963,000 a year. Those are inputs, not expectations, and the share of miners on the official client is unknown. Nothing in section 2 is funded against them.
|
||||
The fee is 1% of the producer share on the official client, default on and switchable: the fee template moves only the producer payout (80% of emission), and the proving pool is paid per record and carries none of it. Rewards in year one are 963 million IGN (ramp included, `docs/analysis/security-budget.md`), so the producer share is 770 million. At USD 0.005, 0.02 and 0.10 per IGN the fee's ceiling, with every miner on the official client, is USD 38,520, 154,080 and 770,400 a year (corrected 6 October 2026 from 48,000, 193,000 and 963,000, which took 1% of all rewards and overstated the ceiling by a quarter; the Horizon economy lane, `docs/analysis/horizon/economy-and-utility.md` section 4.4). Those are inputs, not expectations, and the share of miners on the official client is unknown. Nothing in section 2 is funded against them.
|
||||
|
||||
## 5. Rules
|
||||
|
||||
|
|
|
|||
BIN
docs/plans/site-ui-3-shots/after/home-1440-dark-fold.jpg
Normal file
|
After Width: | Height: | Size: 65 KiB |
BIN
docs/plans/site-ui-3-shots/after/home-1440-dark.jpg
Normal file
|
After Width: | Height: | Size: 185 KiB |
BIN
docs/plans/site-ui-3-shots/after/home-1440-light-fold.jpg
Normal file
|
After Width: | Height: | Size: 60 KiB |
BIN
docs/plans/site-ui-3-shots/after/home-1440-light.jpg
Normal file
|
After Width: | Height: | Size: 180 KiB |
BIN
docs/plans/site-ui-3-shots/after/home-390-dark-fold.jpg
Normal file
|
After Width: | Height: | Size: 30 KiB |
BIN
docs/plans/site-ui-3-shots/after/home-390-dark.jpg
Normal file
|
After Width: | Height: | Size: 127 KiB |
BIN
docs/plans/site-ui-3-shots/after/home-390-light-fold.jpg
Normal file
|
After Width: | Height: | Size: 31 KiB |
BIN
docs/plans/site-ui-3-shots/after/home-390-light.jpg
Normal file
|
After Width: | Height: | Size: 129 KiB |
BIN
docs/plans/site-ui-3-shots/after/index-steps-1440-dark-clip.jpg
Normal file
|
After Width: | Height: | Size: 58 KiB |
BIN
docs/plans/site-ui-3-shots/after/index-steps-1440-dark.webm
Normal file
BIN
docs/plans/site-ui-3-shots/after/index-steps-390-dark-clip.jpg
Normal file
|
After Width: | Height: | Size: 42 KiB |
|
|
@ -12,7 +12,10 @@
|
|||
# before the next pass wipes the dir), WITHOUT sccache (a hit would hand pass B pass A's object and hide a
|
||||
# non-determinism), each under a build slot through remote-run.sh (one JSONL line per pass, kind node-linux or
|
||||
# node-windows, tool repro). The Windows environment is cross-remote.sh's, flag for flag (static libgcc and libstdc++,
|
||||
# -Wl,--no-insert-timestamp). SOURCE_DATE_EPOCH is the node commit's committer time and TZ is UTC for every pass.
|
||||
# -Wl,--no-insert-timestamp). SOURCE_DATE_EPOCH is the node commit's author time (lib.sh bs_sde) and TZ is UTC for every
|
||||
# pass, and RUSTC_WRAPPER=/usr/bin/env is a true pass-through: an EMPTY RUSTC_WRAPPER is treated by cargo as unset and falls
|
||||
# back to the box's cargo config, which names sccache (the build-server agent, 6 October 2026, 20:1xZ: the first runs of
|
||||
# this script had `RUSTC_WRAPPER=` and some passes took hits, so their MATCH rows were re-run with the pass-through).
|
||||
# Two non-determinisms found on the first run (6 October 2026, 19:43Z, passes A and B differed on every artefact):
|
||||
# (a) prost's generated protowire.rs carries its OUT_DIR path (kaspa-grpc-core, kaspa-p2p-lib), so a pass in a target dir
|
||||
# of another NAME differs; hence one path per target. (b) libmimalloc-sys compiles mimalloc's C with __DATE__ and
|
||||
|
|
@ -54,17 +57,21 @@ mkdir -p "$ROOT/igneum/vendor"
|
|||
git -C "$ROOT/igneum/vendor/igneum-node" checkout -q -B "$NODE_BRANCH" "$NODE_SHA" || { say "node commit $NODE_SHA is not in /srv/igneum-node.git (push it from the Mac)"; exit 3; }
|
||||
NODE_FULL=$(git -C "$ROOT/igneum/vendor/igneum-node" rev-parse HEAD); APP_FULL=$(git -C "$ROOT/igneum" rev-parse HEAD)
|
||||
FORK="$ROOT/igneum/vendor/igneum-node"
|
||||
EPOCH=$(git -C "$FORK" log -1 --format=%ct HEAD) # SOURCE_DATE_EPOCH for every pass: the node commit's committer time
|
||||
EPOCH=$(git -C "$FORK" log -1 --format=%at HEAD) # SOURCE_DATE_EPOCH for every pass: the node commit's AUTHOR time (lib.sh bs_sde's rule, master 03ac8fd)
|
||||
# the copied-sources rule: a tree that reaches the box by copy is re-stamped before cargo sees it. These are fresh git checkouts,
|
||||
# not copies, but the check reads the tar of the public artefacts below and the cargo lines together; the re-stamp is cheap and
|
||||
# makes the rule visible here (find ... touch over both clean clones, .git and target dirs excluded)
|
||||
find "$ROOT/igneum" -type f -not -path '*/.git/*' -not -path '*/target*' -exec touch {} +
|
||||
|
||||
# 2. the builds: one remote-run.sh invocation per pass and target; no sccache
|
||||
run_pass() { # <kind: linux|windows> <pass letter>
|
||||
local kind="$1" pass="$2" tdir="target-repro-$1" cmd envb="" arts keep="$ROOT/pass-$1-$2"
|
||||
if [ "$kind" = linux ]; then
|
||||
cmd="RUSTC_WRAPPER= CARGO_TARGET_DIR='$tdir' cargo build --release -p kaspad -p igneum-miner --features kaspad/igneum-pow 2>&1 | tail -3; ( exit \${PIPESTATUS[0]} )"
|
||||
cmd="RUSTC_WRAPPER=/usr/bin/env CARGO_TARGET_DIR='$tdir' cargo build --release -p kaspad -p igneum-miner --features kaspad/igneum-pow 2>&1 | tail -3; ( exit \${PIPESTATUS[0]} )"
|
||||
arts="$tdir/release/igneumd $tdir/release/igneum-miner"; BR_KIND=node-linux; BR_TARGET=x86_64-unknown-linux-gnu
|
||||
else
|
||||
envb='LLVM_LIB=$(ls -d /usr/lib/llvm-*/lib 2>/dev/null | sort -V | tail -1); export CC_x86_64_pc_windows_gnu=x86_64-w64-mingw32-gcc-posix CXX_x86_64_pc_windows_gnu=x86_64-w64-mingw32-g++-posix AR_x86_64_pc_windows_gnu=x86_64-w64-mingw32-ar CARGO_TARGET_X86_64_PC_WINDOWS_GNU_LINKER=x86_64-w64-mingw32-gcc-posix; export CARGO_TARGET_X86_64_PC_WINDOWS_GNU_RUSTFLAGS="-C link-arg=-static -C link-arg=-static-libgcc -C link-arg=-static-libstdc++ -C link-arg=-Wl,--no-insert-timestamp"; export IGNEUM_WINDRES=x86_64-w64-mingw32-windres LIBCLANG_PATH="$LLVM_LIB" BINDGEN_EXTRA_CLANG_ARGS_x86_64_pc_windows_gnu="--target=x86_64-w64-mingw32 --sysroot=/usr/x86_64-w64-mingw32 -I/usr/x86_64-w64-mingw32/include"; '
|
||||
cmd="${envb}RUSTC_WRAPPER= CARGO_TARGET_DIR='$tdir' cargo build --release -p kaspad -p igneum-miner --features igneum-pow --target $WIN 2>&1 | tail -3; ( exit \${PIPESTATUS[0]} )"
|
||||
cmd="${envb}RUSTC_WRAPPER=/usr/bin/env CARGO_TARGET_DIR='$tdir' cargo build --release -p kaspad -p igneum-miner --features igneum-pow --target $WIN 2>&1 | tail -3; ( exit \${PIPESTATUS[0]} )"
|
||||
arts="$tdir/$WIN/release/igneumd.exe $tdir/$WIN/release/igneum-miner.exe"; BR_KIND=node-windows; BR_TARGET=$WIN
|
||||
fi
|
||||
if [ "$REUSE" = 1 ]; then local all=1 a; for a in $arts; do [ -f "$keep/$(basename "$a")" ] || all=0; done; [ "$all" = 1 ] && { say "$kind pass $pass reused from $keep (--reuse)"; return 0; }; fi
|
||||
|
|
|
|||
125
proto-newpow/mma-shadow/README.md
Normal file
|
|
@ -0,0 +1,125 @@
|
|||
# mma-shadow: prototype kernel for class "mx8+mm8xR" (Horizon lane 8, new proof of work)
|
||||
|
||||
Measured 6 October 2026 on GPU box 1 (RTX 4090 24 GB). Everything here lives in this directory; the pack files
|
||||
(kernel.cu, memhard.h, program.h, vectors.h) are verbatim copies of proto-cuda/packs-ca2-mixer/mx8-genesis.
|
||||
|
||||
## The scheme
|
||||
|
||||
The shipped mx8 hash (8 iterations of 64 straight-line instructions over r0..r7, 16 dataset loads, 8 shuffles, the
|
||||
fold) plus a block of R `mm8` steps at the end of every iteration, after instruction 63 and before the next
|
||||
iteration samples `sel`, so 8 x R mm8 steps per hash. Nothing else changes. One mm8 step k with drawn registers
|
||||
(a_k, b_k, c_k, c2_k), a_k != b_k, c_k != c2_k: the 32 lanes' r[a] form A (8 x 16 u8, lane l holds
|
||||
A[l >> 2][4 (l & 3) .. +3], byte 0 = lowest k), the 32 lanes' r[b] form B (16 x 8 u8, lane l holds
|
||||
B[4 (l & 3) .. +3][l >> 2]), C = A x B exact in int32, and lane l does r[c] += C[l >> 2][2 (l & 3)] and
|
||||
r[c2] += C[l >> 2][2 (l & 3) + 1], both modulo 2^32. That is the PTX `mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32`
|
||||
fragment layout with a zero accumulator (d0 into r[c], d1 into r[c2]), the same arithmetic as the mm8 `warp_ref` in
|
||||
proto-cuda/family-probe.cu. Both tile outputs are consumed per step (design correction received 6 Oct 2026 before
|
||||
anything was measured, so the single-output `bit` form was never built). The draws come from a SplitMix64 stream
|
||||
seeded with FNV-1a-64 of "igneum-mm8/igneum-genesis" (seed 0x79f1fc5b6ed6112e): a = below(8); b = below(7),
|
||||
b += (b >= a); c = below(8); c2 = below(7), c2 += (c2 >= c). R is the compile-time macro `IGNEUM_MM8_R`; R = 0 is
|
||||
the control and is bit-exact with the pack (3 vector warps and the 2^24 fingerprint 7c28cfb06c5c65a9).
|
||||
|
||||
## Files
|
||||
|
||||
| File | What |
|
||||
|---|---|
|
||||
| `gen_block.py` | draws the 512-step table, writes `mm8_block.h` (packed uint16 table + X-macro step list) |
|
||||
| `kernel_mm8.cu` | the pack's kernel.cu with r0..r7 as `uint32_t r[8]` (mechanical rewrite, same arithmetic) and the mm8 block; PTX path by default, `-DIGNEUM_MM8_REF` for the shuffle-and-byte-product reference path (12 shuffles + 32 byte products per lane per step). Cache fill, build and launch wrappers unchanged |
|
||||
| `bench.cu` | harness derived from proto-cuda/host.cu (serve mode stripped): device info, `igneum_hash_info` registers and occupancy, GPU cache fill + host fill + FNV check, GPU dataset build + self-test, pack vectors at R = 0, 2^24 fingerprint at base nonce 0, 1 warm-up + 10 timed batches (CUDA events), `--sustain S` for the power meter, `--dump file n` |
|
||||
| `gen_ref_program.py` | turns the 64 instruction lines of kernel.cu into the 32-lane C interpreter body `ref_program.inc` (nothing transcribed by hand) |
|
||||
| `verify_ref.c` | plain C CPU reference: host cache (65536 segments), lazy `mh_word` loads, register-major interpreter, mm8 block in the spec layout, fold; compares a dump, then times the verifier per unit |
|
||||
| `run.sh` | the ladder R in {0, 8, 32, 128, 512}: builds, power sampling, fingerprints, dumps, CPU check, summary |
|
||||
| `summarise.py` | builds the RESULTS table from `out/` |
|
||||
|
||||
## Exact commands (on the box, under /root/horizon-newpow/mma-shadow)
|
||||
|
||||
```
|
||||
export PATH=/usr/local/cuda/bin:$PATH
|
||||
python3 gen_block.py mm8_block.h
|
||||
python3 gen_ref_program.py kernel.cu
|
||||
gcc -O2 -o verify_ref verify_ref.c
|
||||
for R in 0 8 32 128 512; do
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -Xptxas -v -o bench_$R bench.cu kernel_mm8.cu
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -DIGNEUM_MM8_REF -o bench_${R}_ref bench.cu kernel_mm8.cu
|
||||
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu,timestamp --format=csv,noheader -l 1 > out/power_$R.csv &
|
||||
./bench_$R --batches 10 --sustain 25 --dump out/dump_$R.txt 32 > out/bench_$R.log
|
||||
kill %1
|
||||
./bench_${R}_ref --fingerprint-only --no-host-cache > out/bench_${R}_ref.log
|
||||
taskset -c 2 ./verify_ref out/dump_$R.txt $R --time > out/verify_$R.log # add --self-test at R = 0
|
||||
done
|
||||
python3 summarise.py > out/results.md
|
||||
```
|
||||
`./run.sh` is exactly that sequence (plus the idle baseline and the PTX == ref fingerprint comparison). From the
|
||||
Mac: `rsync -az -e "ssh -i ~/.ssh/igneum-fleet -p <box-1-port>" proto-newpow/mma-shadow/ root@<box-1-ip>:/root/horizon-newpow/mma-shadow/`
|
||||
then `touch *` on the far side before building.
|
||||
|
||||
## Card and driver
|
||||
|
||||
NVIDIA GeForce RTX 4090, 128 SMs, cc 8.9, 24 GB. Driver 570.172.08 (the brief said 595; nvidia-smi reports 570.172.08),
|
||||
CUDA 12.8 (nvcc V12.8.93), g++ 13.3.0, Ubuntu 24.04, 96 CPU threads. Idle baseline before the run: 15.0 to 15.3 W at
|
||||
210 MHz SM, 45 C, 0 % utilisation. The host carried a CPU load average of about 13 from other tenants during the run
|
||||
(the GPU itself was idle and ours alone); that is why the verifier timings were pinned to one core.
|
||||
|
||||
## Method notes
|
||||
|
||||
- Timing: one warm-up batch (also the fingerprint batch) then 10 timed batches of 2^24 hashes between CUDA events,
|
||||
1 warp per block (the pack bench default, 24 resident warps per SM at 29 registers). Power: nvidia-smi at 1 Hz
|
||||
during a 25 s sustained phase after the timed batches; the mean takes samples from 10 s after the sustain start to
|
||||
its end (the timestamps are in the csv). Microjoules per hash = mean watts / (MH/s x 10^6) x 10^6.
|
||||
- Fingerprint = FNV-1a 64 over the 2^24 little-endian u64 outputs at base nonce 0. The PTX build and the reference
|
||||
build must agree (the reference is the plain-integer byte-product form of the same fragment layout).
|
||||
- CPU == GPU: 32 warps at SplitMix64(0x1234) 32-aligned bases, 1,024 lanes, recomputed by `verify_ref`.
|
||||
- The verifier here is a naive register-major interpreter with no interleaving and a lazy `mh_word` per load
|
||||
(72 mixer applications and 8 cache reads per word, 4,096 words per unit). It is slower than the project's Rust
|
||||
verifier (2.06 ms per unit on an M5 Max core) on this box's core, so the number that matters is the mm8 block
|
||||
delta (ms per unit at R minus ms per unit at R = 0), not the total.
|
||||
|
||||
## RESULTS (6 Oct 2026, RTX 4090, driver 570.172.08, CUDA 12.8, sm_89, 1 warp/block, batch 2^24 x 10)
|
||||
|
||||
| R | mm8 per hash (8R) | MH/s (GPU time) | ratio to R = 0 | watts mean (sustain, after first 10 s) | SM MHz | uJ per hash | max C | fingerprint (2^24 at base 0) | PTX == ref | CPU == GPU lanes | verifier ms per unit on the box core: R = 0, R, block delta | regs per thread (PTX build) | regs (ref build) | blocks per SM (cudaOccupancy) |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 (matches the pack) | yes | 1024 of 1024 | 10.18, 10.08, -0.10 (noise) | 29 | 29 | 24 (24 warps/SM, 50 %) |
|
||||
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.93, 9.97, +0.05 | 30 | 77 | 24 |
|
||||
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.08, 10.24, +0.17 | 29 | 151 | 24 |
|
||||
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.18, 11.33, +1.14 | 32 | 175 | 24 |
|
||||
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.89, 14.28, +4.39 | 36 | 213 | 24 |
|
||||
|
||||
Sustained MH/s over the 25 s power window: 63.075, 63.074, 63.072, 63.034, 63.073 for R = 0, 8, 32, 128, 512 (wall,
|
||||
including a sync per batch). Wall and GPU-event rates agree to 0.01 MH/s at every R. Idle baseline 15.0 to 15.3 W.
|
||||
Registers per thread at R = 0, 8, 32, 128, 512: 29, 30, 29, 32, 36 (PTX path), no spills, no stack, at every R.
|
||||
R = 0 self-checks: cache FNV 48c4f5bf24166b2e PASS (GPU == host fill word for word), dataset head, [MASK], 64 random
|
||||
points and 64 Mac samples PASS, the pack's 3 vector warps PASS standalone and in batch, fingerprint 7c28cfb06c5c65a9.
|
||||
The C interpreter also reproduces the 3 pack vectors at R = 0 (`--self-test`).
|
||||
|
||||
## What the numbers say
|
||||
|
||||
- On the RTX 4090 the mm8 block is free in hash rate up to R = 512: 4,096 tensor instructions per hash leave the
|
||||
rate at 63.08 MH/s, identical to the control to the third decimal. The kernel is latency-bound on the 16 dependent
|
||||
random loads per iteration; the tensor work fills stalls that were already there. At R = 512 the card issues about
|
||||
2.6 x 10^11 mma.m8n8k16 per second, roughly 2.7 x 10^14 u8 multiply-adds per second (1,024 per instruction), which is
|
||||
in the region of 40 % of the card's dense int8 tensor peak (approximate, from the published TOPS figure). So the
|
||||
next doublings would start to cost hash rate; R = 512 is near the top of the free band, not in the middle of it.
|
||||
- Power is the only GPU cost that moves: 201 W to 216 W (+7.3 %) and 3.19 to 3.42 uJ per hash (+7.2 %) from R = 0
|
||||
to R = 512. SM clock stayed pinned at 2670 MHz at every R, temperature peaked at 64 C, no throttling seen.
|
||||
- Correctness chain holds at every rung: the PTX fragment read and the plain-integer reference path agree on all
|
||||
2^24 lanes of the fingerprint batch at every R, and the CPU interpreter matches the GPU on all 1,024 dumped lanes.
|
||||
The probe's fragment layout (family-probe.cu `warp_ref`, mm8) was used as written and needed no correction.
|
||||
- Verifier cost of the block on the box's core: about 1.1 us per mm8 step per unit (32 lanes x 32 byte products,
|
||||
naive loops), so +1.14 ms per unit at R = 128 and +4.39 ms per unit at R = 512. Against the project's Rust verifier
|
||||
at 2.06 ms per unit, R = 512 would roughly triple verification time unless the block is vectorised (the 8 x 8 x 16
|
||||
tile is 1,024 MACs, a few hundred ns with SIMD, approximate); R = 32 adds 0.17 ms per unit (+8 % of 2.06 ms) and
|
||||
R = 128 adds 1.14 ms (+55 %). The totals in the table (about 10 ms per unit) are this naive interpreter's cost
|
||||
with a lazy mh_word per load and are not comparable to the Rust verifier; the delta column is the number to use.
|
||||
- The reference-path build is only a correctness oracle: fully unrolled it reaches 213 registers at R = 512 (8 blocks
|
||||
per SM) and was never timed.
|
||||
|
||||
## Raw outputs
|
||||
|
||||
`out/` holds everything the box produced: `run.log` (the whole run), `bench_R.log` and `bench_R_ref.log`,
|
||||
`power_R.csv` (1 Hz: W, SM MHz, C, util, timestamp), `power_idle.csv`, `ptxas_R.txt` and `ptxas_R_ref.txt`,
|
||||
`dump_R.txt` (32 warps, 1,024 lanes), `verify_R.log`, `results.md` (summarise.py).
|
||||
|
||||
## What was cut
|
||||
|
||||
Nothing from the brief. The host 1 GiB dataset was never built on the CPU (lazy `mh_word`, as allowed). The 32-lane
|
||||
dump bases are 32-bit, so base + 31 cannot overflow (base = low32(next()) & ~31).
|
||||
308
proto-newpow/mma-shadow/bench.cu
Normal file
|
|
@ -0,0 +1,308 @@
|
|||
// bench.cu (proto-newpow/mma-shadow): host harness for the mx8+mm8xR prototype, derived from proto-cuda/host.cu
|
||||
// with serve mode stripped. TEST HARNESS ONLY: no pool, no network, no wallet.
|
||||
//
|
||||
// Build: nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=<R> [-DIGNEUM_MM8_REF] -o bench_<R> bench.cu kernel_mm8.cu
|
||||
// Steps: device info, kernel registers and occupancy, GPU cache fill + host cache fill + check (FNV, head, last),
|
||||
// GPU dataset build + self-test (head, [MASK], 64 random points vs host mh_word, Mac samples), the pack's 3 vector warps
|
||||
// (R == 0 only; they must PASS), the 2^24 batch fingerprint at base nonce 0 (FNV-1a 64 over the output bytes), timing
|
||||
// (1 warm-up + N timed batches with CUDA events), an optional sustain phase for the power meter, and --dump.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <chrono>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
#include "program.h"
|
||||
#include "vectors.h"
|
||||
#include "memhard.h"
|
||||
#include "mm8_block.h"
|
||||
#ifndef IGNEUM_MM8_R
|
||||
#define IGNEUM_MM8_R 0
|
||||
#endif
|
||||
|
||||
#define CUDA_CHECK(call) do { cudaError_t err_ = (call); if (err_ != cudaSuccess) { \
|
||||
std::fprintf(stderr, "CUDA error: %s (%d)\n at %s:%d\n in %s\n", cudaGetErrorString(err_), (int)err_, __FILE__, __LINE__, #call); \
|
||||
std::exit(2); } } while (0)
|
||||
|
||||
static double wallMs() {
|
||||
using namespace std::chrono;
|
||||
return duration<double, std::milli>(steady_clock::now().time_since_epoch()).count();
|
||||
}
|
||||
static double epochS() {
|
||||
using namespace std::chrono;
|
||||
return duration<double>(system_clock::now().time_since_epoch()).count();
|
||||
}
|
||||
static uint64_t fnv1a64(const void* p, size_t n) {
|
||||
const uint8_t* b = (const uint8_t*)p;
|
||||
uint64_t h = 0xcbf29ce484222325ull;
|
||||
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
|
||||
return h;
|
||||
}
|
||||
static uint64_t splitmix64(uint64_t& s) {
|
||||
s += 0x9E3779B97F4A7C15ull;
|
||||
uint64_t z = s;
|
||||
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull;
|
||||
z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
|
||||
return z ^ (z >> 31);
|
||||
}
|
||||
|
||||
static const uint32_t CACHE_WORDS = 1u << IGNEUM_CACHE_LOG2_WORDS;
|
||||
static uint32_t* gCache = nullptr;
|
||||
static std::vector<uint32_t> hCache;
|
||||
|
||||
static bool setupCache(bool hostFill) {
|
||||
size_t bytes = (size_t)CACHE_WORDS * 4u;
|
||||
CUDA_CHECK(cudaMalloc((void**)&gCache, bytes));
|
||||
cudaEvent_t e0, e1;
|
||||
CUDA_CHECK(cudaEventCreate(&e0)); CUDA_CHECK(cudaEventCreate(&e1));
|
||||
float ms[2] = {0.f, 0.f};
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(igneum_launch_cache_fill(gCache, IGNEUM_CACHE_SEGMENTS));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
CUDA_CHECK(cudaEventElapsedTime(&ms[pass], e0, e1));
|
||||
}
|
||||
std::printf("cache fill (GPU): %.2f ms first, %.2f ms second (%u MiB)\n", ms[0], ms[1], (unsigned)(bytes >> 20));
|
||||
std::vector<uint32_t> dev(CACHE_WORDS);
|
||||
CUDA_CHECK(cudaMemcpy(dev.data(), gCache, bytes, cudaMemcpyDeviceToHost));
|
||||
uint64_t fnvDev = fnv1a64(dev.data(), bytes);
|
||||
bool fnvOk = fnvDev == IGNEUM_CACHE_FNV64;
|
||||
bool headOk = std::memcmp(dev.data(), IGNEUM_CACHE_HEAD, 64) == 0;
|
||||
bool lastOk = std::memcmp(dev.data() + CACHE_WORDS - 16u, IGNEUM_CACHE_LAST, 64) == 0;
|
||||
bool same = true;
|
||||
double hostMs = 0;
|
||||
if (hostFill) {
|
||||
hCache.assign(CACHE_WORDS, 0u);
|
||||
double h0 = wallMs();
|
||||
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(hCache.data(), seg);
|
||||
hostMs = wallMs() - h0;
|
||||
same = std::memcmp(dev.data(), hCache.data(), bytes) == 0;
|
||||
std::printf("cache fill (host, one thread): %.1f ms, GPU == host all words: %s\n", hostMs, same ? "PASS" : "FAIL");
|
||||
} else {
|
||||
hCache.swap(dev); // the device cache (FNV-checked) serves the host-side dataset derivation
|
||||
}
|
||||
bool pass = fnvOk && headOk && lastOk && same;
|
||||
std::printf("cache check: %s (device FNV-1a 64 %016llx vs Mac %016llx %s, head %s, last line %s)\n",
|
||||
pass ? "PASS" : "FAIL", (unsigned long long)fnvDev, (unsigned long long)IGNEUM_CACHE_FNV64,
|
||||
fnvOk ? "PASS" : "FAIL", headOk ? "PASS" : "FAIL", lastOk ? "PASS" : "FAIL");
|
||||
CUDA_CHECK(cudaEventDestroy(e0)); CUDA_CHECK(cudaEventDestroy(e1));
|
||||
return pass;
|
||||
}
|
||||
|
||||
static bool compareWarp(const uint64_t* got, const uint64_t* want, uint32_t base, const char* how) {
|
||||
int bad = 0, first = -1;
|
||||
for (int l = 0; l < 32; ++l) if (got[l] != want[l]) { if (first < 0) first = l; ++bad; }
|
||||
if (bad == 0) std::printf("verify warp base %u %s: PASS\n", base, how);
|
||||
else std::printf("verify warp base %u %s: FAIL %d of 32 lanes differ, first lane %d: gpu=%016llx expected=%016llx\n",
|
||||
base, how, bad, first, (unsigned long long)got[first], (unsigned long long)want[first]);
|
||||
return bad == 0;
|
||||
}
|
||||
|
||||
static void usage() {
|
||||
std::printf("bench_R [--batches 10] [--block-warps 1] [--batch-log2 24] [--fingerprint-only] [--sustain S] [--no-host-cache]\n"
|
||||
" [--dump <file> <n>] [--device 0]\n"
|
||||
" --fingerprint-only skip the timed batches (for the reference build)\n"
|
||||
" --sustain S after the timed batches keep launching batches for S seconds (power meter window)\n"
|
||||
" --dump file n write n whole warps at SplitMix64(0x1234) 32-aligned bases as lines: base lane value_hex\n"
|
||||
" --no-host-cache skip the one-thread host cache fill (the FNV-checked device cache then feeds the host derivation)\n");
|
||||
}
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
int batches = 10, blockWarps = 1, batchLog2 = 24, device = 0, dumpN = 0;
|
||||
double sustain = 0;
|
||||
bool fpOnly = false, hostCache = true;
|
||||
std::string dumpFile;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
std::string a = argv[i];
|
||||
auto nextInt = [&](int& dst) { if (i + 1 >= argc) { usage(); std::exit(2); } dst = std::atoi(argv[++i]); };
|
||||
if (a == "--batches") nextInt(batches);
|
||||
else if (a == "--block-warps") nextInt(blockWarps);
|
||||
else if (a == "--batch-log2") nextInt(batchLog2);
|
||||
else if (a == "--device") nextInt(device);
|
||||
else if (a == "--fingerprint-only") fpOnly = true;
|
||||
else if (a == "--no-host-cache") hostCache = false;
|
||||
else if (a == "--sustain") { if (i + 1 >= argc) { usage(); return 2; } sustain = std::atof(argv[++i]); }
|
||||
else if (a == "--dump") { if (i + 2 >= argc) { usage(); return 2; } dumpFile = argv[++i]; dumpN = std::atoi(argv[++i]); }
|
||||
else if (a == "-h" || a == "--help") { usage(); return 0; }
|
||||
else { std::printf("unknown argument %s\n", argv[i]); usage(); return 2; }
|
||||
}
|
||||
if (blockWarps < 1 || blockWarps > 32 || batchLog2 < 10 || batchLog2 > 28 || batches < 1) { usage(); return 2; }
|
||||
|
||||
std::printf("mma-shadow bench pack \"%s\" class mx8+mm8xR R = %d mm8 per hash = %d path = %s\n",
|
||||
IGNEUM_SEED_STRING, IGNEUM_MM8_R, 8 * IGNEUM_MM8_R,
|
||||
#ifdef IGNEUM_MM8_REF
|
||||
"reference (shuffles + byte products)"
|
||||
#else
|
||||
"PTX mma.sync.m8n8k16.u8"
|
||||
#endif
|
||||
);
|
||||
std::printf("mm8 table seed 0x%016llx (first steps: %s)\n", (unsigned long long)IGNEUM_MM8_SEED, "see mm8_block.h");
|
||||
|
||||
int count = 0;
|
||||
CUDA_CHECK(cudaGetDeviceCount(&count));
|
||||
if (count == 0 || device >= count) { std::printf("FAIL: no CUDA device %d\n", device); return 2; }
|
||||
CUDA_CHECK(cudaSetDevice(device));
|
||||
cudaDeviceProp prop; std::memset(&prop, 0, sizeof(prop));
|
||||
CUDA_CHECK(cudaGetDeviceProperties(&prop, device));
|
||||
int drv = 0, rt = 0; CUDA_CHECK(cudaDriverGetVersion(&drv)); CUDA_CHECK(cudaRuntimeGetVersion(&rt));
|
||||
int clk = 0, memclk = 0, bus = 0, l2 = 0, thrSM = 0;
|
||||
cudaDeviceGetAttribute(&clk, cudaDevAttrClockRate, device);
|
||||
cudaDeviceGetAttribute(&memclk, cudaDevAttrMemoryClockRate, device);
|
||||
cudaDeviceGetAttribute(&bus, cudaDevAttrGlobalMemoryBusWidth, device);
|
||||
cudaDeviceGetAttribute(&l2, cudaDevAttrL2CacheSize, device);
|
||||
cudaDeviceGetAttribute(&thrSM, cudaDevAttrMaxThreadsPerMultiProcessor, device);
|
||||
cudaGetLastError();
|
||||
std::printf("GPU: %s (%d SMs, cc %d.%d, %.0f MiB) SM clock %d MHz, mem clock %d MHz, bus %d bits, L2 %d MiB, max %d threads/SM\n",
|
||||
prop.name, prop.multiProcessorCount, prop.major, prop.minor, (double)prop.totalGlobalMem / 1048576.0,
|
||||
clk / 1000, memclk / 1000, bus, l2 / 1048576, thrSM);
|
||||
std::printf("CUDA: driver %d.%d, runtime %d.%d\n", drv / 1000, (drv % 100) / 10, rt / 1000, (rt % 100) / 10);
|
||||
int regs = 0, blocksPerSM = 0;
|
||||
CUDA_CHECK(igneum_hash_info(®s, &blocksPerSM, (uint32_t)blockWarps));
|
||||
std::printf("igneum_hash_info: %d registers/thread, %d resident blocks/SM at %d warp(s)/block = %d resident warps/SM (%.1f%% of %d)\n",
|
||||
regs, blocksPerSM, blockWarps, blocksPerSM * blockWarps, 100.0 * blocksPerSM * blockWarps * 32 / thrSM, thrSM / 32);
|
||||
|
||||
bool cachePass = setupCache(hostCache);
|
||||
|
||||
uint32_t nonces = 1u << batchLog2;
|
||||
uint32_t mask = IGNEUM_MASK;
|
||||
uint64_t dsBytes = (uint64_t)(mask + 1u) * 4ull;
|
||||
uint32_t* dDs = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dDs, (size_t)dsBytes));
|
||||
uint64_t* dOut = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dOut, (size_t)nonces * sizeof(uint64_t)));
|
||||
cudaEvent_t e0, e1;
|
||||
CUDA_CHECK(cudaEventCreate(&e0)); CUDA_CHECK(cudaEventCreate(&e1));
|
||||
float buildMs[2] = {0.f, 0.f};
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(igneum_launch_build(dDs, gCache, (mask + 1u) / 16u));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
CUDA_CHECK(cudaEventElapsedTime(&buildMs[pass], e0, e1));
|
||||
}
|
||||
std::printf("dataset build (GPU, %llu MiB): %.2f ms first, %.2f ms second\n", (unsigned long long)(dsBytes >> 20), buildMs[0], buildMs[1]);
|
||||
|
||||
// Dataset self-test
|
||||
bool dsPass;
|
||||
{
|
||||
uint32_t head[16];
|
||||
CUDA_CHECK(cudaMemcpy(head, dDs, sizeof(head), cudaMemcpyDeviceToHost));
|
||||
int badHead = 0;
|
||||
for (int i = 0; i < 16; ++i) if (head[i] != IGNEUM_DS_HEAD[i]) { if (!badHead) std::printf(" dataset[%d] = %08x, Mac %08x\n", i, head[i], IGNEUM_DS_HEAD[i]); ++badHead; }
|
||||
uint32_t last = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&last, dDs + IGNEUM_DS_LAST_INDEX, 4, cudaMemcpyDeviceToHost));
|
||||
bool lastOk = last == IGNEUM_DS_LAST;
|
||||
int badRnd = 0;
|
||||
uint64_t s = 0x9E3779B97F4A7C15ull ^ (uint64_t)(mask + 1u);
|
||||
for (int k = 0; k < 64; ++k) {
|
||||
uint32_t idx = (uint32_t)splitmix64(s) & mask, v = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&v, dDs + idx, 4, cudaMemcpyDeviceToHost));
|
||||
uint32_t want = mh_word(hCache.data(), idx);
|
||||
if (v != want) { if (!badRnd) std::printf(" dataset[%u] = %08x, host derivation %08x\n", idx, v, want); ++badRnd; }
|
||||
}
|
||||
int badSample = 0;
|
||||
for (int k = 0; k < IGNEUM_DS_SAMPLES; ++k) {
|
||||
uint32_t v = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&v, dDs + IGNEUM_DS_SAMPLE_INDEX[k], 4, cudaMemcpyDeviceToHost));
|
||||
if (v != IGNEUM_DS_SAMPLE_VALUE[k]) { if (!badSample) std::printf(" dataset[%u] = %08x, Mac %08x\n", IGNEUM_DS_SAMPLE_INDEX[k], v, IGNEUM_DS_SAMPLE_VALUE[k]); ++badSample; }
|
||||
}
|
||||
dsPass = badHead == 0 && lastOk && badRnd == 0 && badSample == 0;
|
||||
std::printf("dataset self-test: %s (head 16 %s, [MASK] %s, 64 random points vs host derivation %s, %d Mac samples %s)\n",
|
||||
dsPass ? "PASS" : "FAIL", badHead == 0 ? "PASS" : "FAIL", lastOk ? "PASS" : "FAIL",
|
||||
badRnd == 0 ? "PASS" : "FAIL", (int)IGNEUM_DS_SAMPLES, badSample == 0 ? "PASS" : "FAIL");
|
||||
}
|
||||
|
||||
// Pack vectors (R == 0 only: at R > 0 the function is different by design)
|
||||
bool vecPass = true;
|
||||
uint64_t got[32];
|
||||
if (IGNEUM_MM8_R == 0) {
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, IGNEUM_VEC_BASE[w], mask, 32u, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
|
||||
vecPass = compareWarp(got, IGNEUM_VEC_OUT[w], IGNEUM_VEC_BASE[w], "(pack vector, standalone)") && vecPass;
|
||||
}
|
||||
} else {
|
||||
std::printf("pack vectors: not applicable at R = %d (checked at R = 0 only)\n", IGNEUM_MM8_R);
|
||||
}
|
||||
|
||||
// Warm-up batch at base 0 = the fingerprint batch
|
||||
double w0 = wallMs();
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, 0u, mask, nonces, (uint32_t)blockWarps));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
double w1 = wallMs();
|
||||
std::vector<uint64_t> hOut(nonces);
|
||||
CUDA_CHECK(cudaMemcpy(hOut.data(), dOut, (size_t)nonces * 8u, cudaMemcpyDeviceToHost));
|
||||
uint64_t fp = fnv1a64(hOut.data(), (size_t)nonces * 8u);
|
||||
std::printf("warm-up batch: %u hashes in %.2f ms wall\n", nonces, w1 - w0);
|
||||
std::printf("fingerprint: %016llx (FNV-1a 64 over the 2^%d outputs at base nonce 0, little-endian u64 bytes)%s\n",
|
||||
(unsigned long long)fp, batchLog2,
|
||||
IGNEUM_MM8_R == 0 ? (fp == 0x7c28cfb06c5c65a9ull ? " == 7c28cfb06c5c65a9 PASS" : " != 7c28cfb06c5c65a9 FAIL") : "");
|
||||
bool fpPass = (IGNEUM_MM8_R != 0) || fp == 0x7c28cfb06c5c65a9ull;
|
||||
if (IGNEUM_MM8_R == 0 && batchLog2 >= 20) {
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
if ((uint64_t)IGNEUM_VEC_BASE[w] + 32ull > nonces) continue;
|
||||
vecPass = compareWarp(hOut.data() + IGNEUM_VEC_BASE[w], IGNEUM_VEC_OUT[w], IGNEUM_VEC_BASE[w], "(pack vector, in batch)") && vecPass;
|
||||
}
|
||||
}
|
||||
|
||||
double mhs = 0, mhsWall = 0, mhsSustain = 0;
|
||||
if (!fpOnly) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
double t0 = wallMs();
|
||||
for (int b = 1; b <= batches; ++b) {
|
||||
uint32_t base = (uint32_t)((uint64_t)b * (uint64_t)nonces);
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, (uint32_t)blockWarps));
|
||||
}
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
double t1 = wallMs();
|
||||
float gpuMs = 0.f;
|
||||
CUDA_CHECK(cudaEventElapsedTime(&gpuMs, e0, e1));
|
||||
double total = (double)nonces * batches;
|
||||
mhs = total / (gpuMs / 1000.0) / 1e6;
|
||||
mhsWall = total / ((t1 - t0) / 1000.0) / 1e6;
|
||||
std::printf("timed: %d batches x %u hashes: GPU %.2f ms -> %.3f MH/s (GPU time), wall %.2f ms -> %.3f MH/s\n",
|
||||
batches, nonces, gpuMs, mhs, t1 - t0, mhsWall);
|
||||
if (sustain > 0) {
|
||||
double s0 = epochS(), sw0 = wallMs();
|
||||
long long n = 0;
|
||||
std::printf("sustain start epoch %.3f\n", s0); std::fflush(stdout);
|
||||
while (wallMs() - sw0 < sustain * 1000.0) {
|
||||
uint32_t base = (uint32_t)((uint64_t)(n + batches + 1) * (uint64_t)nonces);
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, (uint32_t)blockWarps));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
++n;
|
||||
}
|
||||
double s1 = epochS();
|
||||
mhsSustain = (double)n * nonces / (s1 - s0) / 1e6;
|
||||
std::printf("sustain end epoch %.3f: %lld batches in %.2f s -> %.3f MH/s (wall, incl. sync)\n", s1, n, s1 - s0, mhsSustain);
|
||||
}
|
||||
}
|
||||
|
||||
if (dumpN > 0) {
|
||||
FILE* f = std::fopen(dumpFile.c_str(), "w");
|
||||
if (!f) { std::printf("FAIL: cannot open %s\n", dumpFile.c_str()); return 2; }
|
||||
uint64_t s = 0x1234ull;
|
||||
for (int i = 0; i < dumpN; ++i) {
|
||||
uint32_t base = (uint32_t)splitmix64(s) & ~31u;
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, 32u, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
|
||||
for (int l = 0; l < 32; ++l) std::fprintf(f, "%u %d %016llx\n", base, l, (unsigned long long)got[l]);
|
||||
}
|
||||
std::fclose(f);
|
||||
std::printf("dump: %d warps (%d lanes) written to %s\n", dumpN, 32 * dumpN, dumpFile.c_str());
|
||||
}
|
||||
|
||||
bool overall = cachePass && dsPass && vecPass && fpPass;
|
||||
std::printf("SUMMARY R=%d regs=%d blocksPerSM=%d mhs=%.3f mhs_wall=%.3f mhs_sustain=%.3f fingerprint=%016llx cache=%s dataset=%s vectors=%s\n",
|
||||
IGNEUM_MM8_R, regs, blocksPerSM, mhs, mhsWall, mhsSustain, (unsigned long long)fp,
|
||||
cachePass ? "PASS" : "FAIL", dsPass ? "PASS" : "FAIL", IGNEUM_MM8_R == 0 ? (vecPass ? "PASS" : "FAIL") : "n/a");
|
||||
std::printf("OVERALL: %s\n", overall ? "PASS" : "FAIL");
|
||||
CUDA_CHECK(cudaFree(dOut)); CUDA_CHECK(cudaFree(dDs)); CUDA_CHECK(cudaFree(gCache));
|
||||
return overall ? 0 : 1;
|
||||
}
|
||||
75
proto-newpow/mma-shadow/gen_block.py
Normal file
|
|
@ -0,0 +1,75 @@
|
|||
#!/usr/bin/env python3
|
||||
# gen_block.py: draws the mm8 block table for the mx8+mm8xR prototype and writes mm8_block.h.
|
||||
# Stream: SplitMix64 seeded with FNV-1a-64 of the bytes "igneum-mm8/igneum-genesis".
|
||||
# Per step k: a = below(8); b = below(7), b += (b >= a); c = below(8); c2 = below(7), c2 += (c2 >= c).
|
||||
# Step semantics: r[c] += C[l >> 2][2 * (l & 3)] (PTX d0), r[c2] += C[l >> 2][2 * (l & 3) + 1] (PTX d1), both mod 2^32.
|
||||
import sys
|
||||
|
||||
MASK64 = (1 << 64) - 1
|
||||
R_MAX = 512
|
||||
|
||||
def fnv1a64(data: bytes) -> int:
|
||||
h = 0xcbf29ce484222325
|
||||
for byte in data:
|
||||
h ^= byte
|
||||
h = (h * 0x100000001b3) & MASK64
|
||||
return h
|
||||
|
||||
class SplitMix64:
|
||||
def __init__(self, seed: int):
|
||||
self.s = seed & MASK64
|
||||
def next(self) -> int:
|
||||
self.s = (self.s + 0x9E3779B97F4A7C15) & MASK64
|
||||
z = self.s
|
||||
z = ((z ^ (z >> 30)) * 0xBF58476D1CE4E5B9) & MASK64
|
||||
z = ((z ^ (z >> 27)) * 0x94D049BB133111EB) & MASK64
|
||||
return z ^ (z >> 31)
|
||||
def below(self, n: int) -> int:
|
||||
return self.next() % n
|
||||
|
||||
def main():
|
||||
seed_text = b"igneum-mm8/igneum-genesis"
|
||||
seed = fnv1a64(seed_text)
|
||||
rng = SplitMix64(seed)
|
||||
rows = []
|
||||
for k in range(R_MAX):
|
||||
a = rng.below(8)
|
||||
b = rng.below(7)
|
||||
b += 1 if b >= a else 0
|
||||
c = rng.below(8)
|
||||
c2 = rng.below(7)
|
||||
c2 += 1 if c2 >= c else 0
|
||||
assert a != b and 0 <= b < 8 and c != c2 and 0 <= c2 < 8
|
||||
rows.append((a, b, c, c2))
|
||||
out = []
|
||||
out.append("// Generated by gen_block.py. mm8 block draws for class mx8+mm8xR, seed text \"%s\"," % seed_text.decode())
|
||||
out.append("// FNV-1a-64 seed 0x%016x, SplitMix64 stream. Step k uses (a, b, c, c2) = row k: r[c] += d0, r[c2] += d1. Do not edit by hand." % seed)
|
||||
out.append("#pragma once")
|
||||
out.append("#ifdef __cplusplus")
|
||||
out.append("#include <cstdint>")
|
||||
out.append("#else")
|
||||
out.append("#include <stdint.h>")
|
||||
out.append("#endif")
|
||||
out.append("#define IGNEUM_MM8_R_MAX %d" % R_MAX)
|
||||
out.append("#define IGNEUM_MM8_SEED 0x%016xull" % seed)
|
||||
out.append("// Packed uint16: a = v & 7, b = (v >> 3) & 7, c = (v >> 6) & 7, c2 = (v >> 9) & 7.")
|
||||
out.append("#define IGNEUM_MM8_TABLE_INIT { \\")
|
||||
for i in range(0, R_MAX, 16):
|
||||
chunk = rows[i:i+16]
|
||||
vals = ["0x%03xu" % (a | (b << 3) | (c << 6) | (c2 << 9)) for (a, b, c, c2) in chunk]
|
||||
out.append(" " + ", ".join(vals) + (", \\" if i + 16 < R_MAX else " }"))
|
||||
out.append("// X-macro list: X(k, a, b, c, c2) for every step k in 0..R_MAX-1. The kernel guards each with k < IGNEUM_MM8_R,")
|
||||
out.append("// so the register indices are compile-time constants (r[] stays in registers, no local memory).")
|
||||
out.append("#define IGNEUM_MM8_STEPS(X) \\")
|
||||
for k, (a, b, c, c2) in enumerate(rows):
|
||||
out.append(" X(%d, %d, %d, %d, %d)%s" % (k, a, b, c, c2, " \\" if k + 1 < R_MAX else ""))
|
||||
out.append("// Readable form, step: a b c c2")
|
||||
for k, (a, b, c, c2) in enumerate(rows):
|
||||
out.append("// %3d: %d %d %d %d" % (k, a, b, c, c2))
|
||||
path = sys.argv[1] if len(sys.argv) > 1 else "mm8_block.h"
|
||||
with open(path, "w") as f:
|
||||
f.write("\n".join(out) + "\n")
|
||||
print("seed 0x%016x, %d steps, first 4: %s -> %s" % (seed, R_MAX, rows[:4], path))
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
23
proto-newpow/mma-shadow/gen_ref_program.py
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
#!/usr/bin/env python3
|
||||
# gen_ref_program.py: turns the 64 instruction lines of the pack's kernel.cu into a 32-lane C interpreter body
|
||||
# (ref_program.inc, included by verify_ref.c). Register-major: R[8][32]. A shuffle copies its source register
|
||||
# into T before the lane loop and reads T[l ^ mask]. Loads become mh_word(cache, x & mask) (lazy dataset).
|
||||
import re, sys
|
||||
src = open(sys.argv[1] if len(sys.argv) > 1 else "kernel.cu").read().split("\n")
|
||||
lines = [l for l in src if re.search(r"//\s*\d+ (mad|xor|load|shfl|mulhi|rotr|or|add|sub|mul|rotl)\s*$", l)]
|
||||
assert len(lines) == 64, len(lines)
|
||||
out = ["// Generated by gen_ref_program.py from the pack's kernel.cu. Do not edit by hand.",
|
||||
"// Expects: uint32_t R[8][32], T[32], SEL[32]; const uint32_t* cache; uint32_t mask; macro L = for (int l = 0; l < 32; ++l)."]
|
||||
for i, l in enumerate(lines):
|
||||
body, comment = l.strip().rsplit("//", 1)
|
||||
body = body.strip()
|
||||
m = re.search(r"__shfl_xor_sync\(0xffffffffu, r(\d), (\d+)\)", body)
|
||||
if m:
|
||||
out.append(" memcpy(T, R[%s], sizeof T);" % m.group(1))
|
||||
body = body.replace(m.group(0), "T[l ^ %s]" % m.group(2))
|
||||
body = re.sub(r"ds\[(r\d) & mask\]", r"mh_word(cache, \1 & mask)", body)
|
||||
body = re.sub(r"\br([0-7])\b", r"R[\1][l]", body)
|
||||
body = body.replace("__umulhi(", "umulhi(").replace("sel ", "SEL[l] ")
|
||||
out.append(" L { %s } //%s" % (body, comment))
|
||||
open("ref_program.inc", "w").write("\n".join(out) + "\n")
|
||||
print("wrote ref_program.inc with %d instructions" % len(lines))
|
||||
164
proto-newpow/mma-shadow/kernel.cu
Normal file
|
|
@ -0,0 +1,164 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
217
proto-newpow/mma-shadow/kernel_mm8.cu
Normal file
|
|
@ -0,0 +1,217 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// kernel_mm8.cu (proto-newpow/mma-shadow): the pack's kernel.cu with r0..r7 rewritten as uint32_t r[8] (same arithmetic)
|
||||
// and a block of IGNEUM_MM8_R mm8 steps at the end of every iteration (class "mx8+mm8xR"). IGNEUM_MM8_R 0 is the control
|
||||
// and is bit-exact with the pack. Define IGNEUM_MM8_REF for the shuffle-and-byte-product reference path instead of the PTX mma.
|
||||
// Step semantics (6 Oct 2026 correction): both tile outputs are consumed, r[c] += d0 and r[c2] += d1.
|
||||
// Cache fill, dataset build and the launch wrappers are the pack's text, unchanged.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
#include "mm8_block.h"
|
||||
#ifndef IGNEUM_MM8_R
|
||||
#define IGNEUM_MM8_R 0
|
||||
#endif
|
||||
#if IGNEUM_MM8_R < 0 || IGNEUM_MM8_R > IGNEUM_MM8_R_MAX
|
||||
#error "IGNEUM_MM8_R out of range"
|
||||
#endif
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
|
||||
// One mm8 step. A = the 32 lanes' r[a] (8 x 16 u8, lane l holds A[l >> 2][4 * (l & 3) .. +3], byte 0 = lowest k),
|
||||
// B = the 32 lanes' r[b] (16 x 8 u8, lane l holds B[4 * (l & 3) .. +3][l >> 2]). C = A * B exact in int32.
|
||||
// Lane l adds C[l >> 2][2 * (l & 3)] into r[c] and C[l >> 2][2 * (l & 3) + 1] into r[c2], both modulo 2^32 (c != c2),
|
||||
// so every step consumes both outputs of the tile. This is the PTX mma.m8n8k16 .u8 fragment layout (d0, d1), the same
|
||||
// arithmetic as proto-cuda/family-probe.cu's mm8 warp_ref.
|
||||
#ifdef IGNEUM_MM8_REF
|
||||
// Reference path: gather A's row (l >> 2) and B's two columns 2 * (l & 3) and 2 * (l & 3) + 1 with 12 shuffles,
|
||||
// then 32 byte products per lane in plain integer code.
|
||||
__device__ __forceinline__ void mm8_pair(uint32_t av, uint32_t bv, uint32_t lane, uint32_t& d0, uint32_t& d1) {
|
||||
uint32_t row = lane >> 2, col0 = 2u * (lane & 3u);
|
||||
uint32_t acc0 = 0u, acc1 = 0u;
|
||||
#pragma unroll
|
||||
for (uint32_t j = 0u; j < 4u; ++j) {
|
||||
uint32_t aw = __shfl_sync(0xffffffffu, av, (int)(4u * row + j));
|
||||
uint32_t bw0 = __shfl_sync(0xffffffffu, bv, (int)(4u * col0 + j));
|
||||
uint32_t bw1 = __shfl_sync(0xffffffffu, bv, (int)(4u * (col0 + 1u) + j));
|
||||
#pragma unroll
|
||||
for (uint32_t t = 0u; t < 4u; ++t) {
|
||||
uint32_t ab = (aw >> (8u * t)) & 0xffu;
|
||||
acc0 += ab * ((bw0 >> (8u * t)) & 0xffu);
|
||||
acc1 += ab * ((bw1 >> (8u * t)) & 0xffu);
|
||||
}
|
||||
}
|
||||
d0 = acc0; d1 = acc1;
|
||||
}
|
||||
#else
|
||||
// Native path: one tensor instruction with a zero accumulator, d0 and d1 straight out of the fragment.
|
||||
__device__ __forceinline__ void mm8_pair(uint32_t av, uint32_t bv, uint32_t lane, uint32_t& d0, uint32_t& d1) {
|
||||
(void)lane;
|
||||
asm("mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32 {%0,%1}, {%2}, {%3}, {%4,%5};"
|
||||
: "=r"(d0), "=r"(d1) : "r"(av), "r"(bv), "r"(0u), "r"(0u));
|
||||
}
|
||||
#endif
|
||||
#define IGNEUM_MM8_STEP(k, a, b, c, c2) if ((k) < IGNEUM_MM8_R) { uint32_t d0_, d1_; mm8_pair(r[a], r[b], lane, d0_, d1_); r[c] += d0_; r[c2] += d1_; }
|
||||
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r[8];
|
||||
uint32_t lane = threadIdx.x & 31u;
|
||||
(void)lane;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r[0] = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r[1] = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r[2] = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r[3] = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r[4] = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r[5] = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r[6] = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r[7] = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r[0];
|
||||
r[2] = r[3] * r[4] + r[2]; // 0 mad
|
||||
r[2] = r[1] * r[1] + r[2]; // 1 mad
|
||||
r[2] = r[3] * r[2] + r[2]; // 2 mad
|
||||
r[3] = r[3] ^ r[5]; // 3 xor
|
||||
r[7] = r[7] ^ ds[r[2] & mask]; // 4 load
|
||||
r[5] = r[5] ^ ds[r[7] & mask]; // 5 load
|
||||
r[1] = r[1] ^ __shfl_xor_sync(0xffffffffu, r[4], 8); // 6 shfl
|
||||
r[7] = r[7] ^ __shfl_xor_sync(0xffffffffu, r[3], 8); // 7 shfl
|
||||
r[1] = __umulhi(r[1], r[5]); // 8 mulhi
|
||||
r[6] = rotr_var(r[6], r[3]); // 9 rotr
|
||||
r[3] = r[3] | r[4]; // 10 or
|
||||
r[4] = r[4] ^ ds[r[3] & mask]; // 11 load
|
||||
r[0] = __umulhi(r[0], r[4]); // 12 mulhi
|
||||
r[5] = r[5] + r[1] + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r[0] = r[0] ^ ds[r[4] & mask]; // 14 load
|
||||
r[2] = r[2] - r[4]; // 15 sub
|
||||
r[2] = r[2] ^ ds[r[0] & mask]; // 16 load
|
||||
r[7] = r[7] ^ ds[r[2] & mask]; // 17 load
|
||||
r[7] = r[7] ^ __shfl_xor_sync(0xffffffffu, r[3], 4); // 18 shfl
|
||||
r[5] = r[5] * r[0]; // 19 mul
|
||||
r[3] = r[3] ^ __shfl_xor_sync(0xffffffffu, r[4], 2); // 20 shfl
|
||||
r[2] = r[2] ^ __shfl_xor_sync(0xffffffffu, r[4], 16); // 21 shfl
|
||||
r[6] = __umulhi(r[6], r[2]); // 22 mulhi
|
||||
r[6] = r[6] ^ ds[r[1] & mask]; // 23 load
|
||||
r[5] = r[5] * r[0]; // 24 mul
|
||||
r[5] = rotl_imm(r[5], 19u); // 25 rotl
|
||||
r[7] = r[7] ^ __shfl_xor_sync(0xffffffffu, r[6], 2); // 26 shfl
|
||||
r[0] = r[0] ^ r[5]; // 27 xor
|
||||
r[0] = r[0] ^ r[4]; // 28 xor
|
||||
r[3] = r[3] - r[0]; // 29 sub
|
||||
r[5] = r[5] * r[1]; // 30 mul
|
||||
r[7] = r[7] ^ ds[r[2] & mask]; // 31 load
|
||||
r[1] = r[1] ^ ds[r[0] & mask]; // 32 load
|
||||
r[5] = r[5] ^ r[6]; // 33 xor
|
||||
r[5] = r[5] ^ ds[r[1] & mask]; // 34 load
|
||||
r[0] = __umulhi(r[0], r[5]); // 35 mulhi
|
||||
r[5] = r[5] ^ __shfl_xor_sync(0xffffffffu, r[2], 4); // 36 shfl
|
||||
r[7] = r[7] ^ ds[r[0] & mask]; // 37 load
|
||||
r[3] = r[3] + r[1] + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r[1] = r[1] ^ __shfl_xor_sync(0xffffffffu, r[5], 4); // 39 shfl
|
||||
r[2] = r[2] ^ r[5]; // 40 xor
|
||||
r[3] = r[6] * r[3] + r[3]; // 41 mad
|
||||
r[6] = r[6] - r[7]; // 42 sub
|
||||
r[7] = r[7] ^ r[0]; // 43 xor
|
||||
r[1] = r[1] ^ ds[r[7] & mask]; // 44 load
|
||||
r[2] = r[2] * r[3]; // 45 mul
|
||||
r[1] = __umulhi(r[1], r[5]); // 46 mulhi
|
||||
r[4] = r[4] - r[3]; // 47 sub
|
||||
r[2] = rotr_var(r[2], r[6]); // 48 rotr
|
||||
r[3] = r[3] ^ ds[r[5] & mask]; // 49 load
|
||||
r[1] = r[1] + r[5] + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r[0] = r[0] * r[2]; // 51 mul
|
||||
r[0] = r[0] + r[2] + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r[1] = r[1] + r[0] + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r[7] = rotl_imm(r[7], 14u); // 54 rotl
|
||||
r[3] = r[3] + r[7] + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r[6] = r[6] ^ ds[r[7] & mask]; // 56 load
|
||||
r[1] = rotr_var(r[1], r[5]); // 57 rotr
|
||||
r[5] = r[5] ^ ds[r[4] & mask]; // 58 load
|
||||
r[6] = r[6] ^ ds[r[2] & mask]; // 59 load
|
||||
r[3] = r[5] * r[0] + r[3]; // 60 mad
|
||||
r[5] = r[5] + r[7] + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r[4] = r[4] + r[6] + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r[5] = rotl_imm(r[5], 19u); // 63 rotl
|
||||
#if IGNEUM_MM8_R > 0
|
||||
// mm8 block: IGNEUM_MM8_R steps after instruction 63, before the next iteration samples sel.
|
||||
IGNEUM_MM8_STEPS(IGNEUM_MM8_STEP)
|
||||
#endif
|
||||
}
|
||||
uint32_t lo = r[0] ^ rotl_imm(r[1], 7u) ^ rotl_imm(r[2], 14u) ^ rotl_imm(r[3], 21u);
|
||||
uint32_t hi = r[4] ^ rotl_imm(r[5], 9u) ^ rotl_imm(r[6], 18u) ^ rotl_imm(r[7], 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
109
proto-newpow/mma-shadow/memhard.h
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
|
||||
// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0x3067619fu ^ prev[4];
|
||||
x[5] = 0x3c269176u ^ prev[5];
|
||||
x[6] = 0x84a03b03u ^ prev[6];
|
||||
x[7] = 0xf8c63294u ^ prev[7];
|
||||
x[8] = 0xff977c5bu ^ prev[8];
|
||||
x[9] = 0xe60def3eu ^ prev[9];
|
||||
x[10] = 0x63630141u ^ prev[10];
|
||||
x[11] = 0xb8fbcb58u ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
|
||||
s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
|
||||
s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
|
||||
s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
|
||||
s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
|
||||
s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
|
||||
s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
|
||||
s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
|
||||
s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
|
||||
s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
|
||||
s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
|
||||
s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
|
||||
s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
|
||||
s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
|
||||
s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
|
||||
s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0x3067619fu;
|
||||
s[1] = 0x3c269176u;
|
||||
s[2] = 0x84a03b03u;
|
||||
s[3] = 0xf8c63294u;
|
||||
s[4] = 0xff977c5bu;
|
||||
s[5] = 0xe60def3eu;
|
||||
s[6] = 0x63630141u;
|
||||
s[7] = 0xb8fbcb58u;
|
||||
s[8] = t * 0x42146205u + 0xbab68293u;
|
||||
s[9] = t * 0x52cbe0fbu + 0xcc162340u;
|
||||
s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
|
||||
s[11] = t * 0x6728907fu + 0xe62b8997u;
|
||||
s[12] = t * 0xd81d9751u + 0xc9c80297u;
|
||||
s[13] = t * 0x132952c3u + 0xf74a1654u;
|
||||
s[14] = t * 0xf60de277u + 0x3d704af5u;
|
||||
s[15] = t * 0x05358035u + 0x3cf522b7u;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
1072
proto-newpow/mma-shadow/mm8_block.h
Normal file
24
proto-newpow/mma-shadow/out/bench_0.log
Normal file
|
|
@ -0,0 +1,24 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 0 mm8 per hash = 0 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.93 ms first, 1.89 ms second (256 MiB)
|
||||
cache fill (host, one thread): 375.8 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.55 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
verify warp base 0 (pack vector, standalone): PASS
|
||||
verify warp base 4096 (pack vector, standalone): PASS
|
||||
verify warp base 1000000 (pack vector, standalone): PASS
|
||||
warm-up batch: 16777216 hashes in 266.01 ms wall
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
verify warp base 0 (pack vector, in batch): PASS
|
||||
verify warp base 4096 (pack vector, in batch): PASS
|
||||
verify warp base 1000000 (pack vector, in batch): PASS
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.74 ms -> 63.078 MH/s
|
||||
sustain start epoch 1791315749.655
|
||||
sustain end epoch 1791315774.658: 94 batches in 25.00 s -> 63.075 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_0.txt
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.075 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_0_ref.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 0 mm8 per hash = 0 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.80 ms first, 1.77 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.58 ms first, 30.51 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
verify warp base 0 (pack vector, standalone): PASS
|
||||
verify warp base 4096 (pack vector, standalone): PASS
|
||||
verify warp base 1000000 (pack vector, standalone): PASS
|
||||
warm-up batch: 16777216 hashes in 266.00 ms wall
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
verify warp base 0 (pack vector, in batch): PASS
|
||||
verify warp base 4096 (pack vector, in batch): PASS
|
||||
verify warp base 1000000 (pack vector, in batch): PASS
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_128.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 128 mm8 per hash = 1024 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 32 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.91 ms first, 1.87 ms second (256 MiB)
|
||||
cache fill (host, one thread): 347.9 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.61 ms first, 30.56 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 128 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 265.96 ms wall
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.82 ms -> 63.077 MH/s (GPU time), wall 2659.83 ms -> 63.076 MH/s
|
||||
sustain start epoch 1791315890.874
|
||||
sustain end epoch 1791315915.893: 94 batches in 25.02 s -> 63.034 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_128.txt
|
||||
SUMMARY R=128 regs=32 blocksPerSM=24 mhs=63.077 mhs_wall=63.076 mhs_sustain=63.034 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
14
proto-newpow/mma-shadow/out/bench_128_ref.log
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 128 mm8 per hash = 1024 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 175 registers/thread, 8 resident blocks/SM at 1 warp(s)/block = 8 resident warps/SM (16.7% of 48)
|
||||
cache fill (GPU): 1.78 ms first, 1.79 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.58 ms first, 30.52 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 128 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 447.05 ms wall
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=128 regs=175 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_32.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 32 mm8 per hash = 256 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.90 ms first, 1.87 ms second (256 MiB)
|
||||
cache fill (host, one thread): 349.6 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.56 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 32 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 266.00 ms wall
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.75 ms -> 63.078 MH/s (GPU time), wall 2659.77 ms -> 63.078 MH/s
|
||||
sustain start epoch 1791315833.292
|
||||
sustain end epoch 1791315858.296: 94 batches in 25.00 s -> 63.072 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_32.txt
|
||||
SUMMARY R=32 regs=29 blocksPerSM=24 mhs=63.078 mhs_wall=63.078 mhs_sustain=63.072 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
14
proto-newpow/mma-shadow/out/bench_32_ref.log
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 32 mm8 per hash = 256 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 151 registers/thread, 12 resident blocks/SM at 1 warp(s)/block = 12 resident warps/SM (25.0% of 48)
|
||||
cache fill (GPU): 1.80 ms first, 1.77 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.57 ms first, 30.52 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 32 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 265.96 ms wall
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_512.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 512 mm8 per hash = 4096 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.93 ms first, 1.89 ms second (256 MiB)
|
||||
cache fill (host, one thread): 348.1 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.62 ms first, 30.54 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 512 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 265.98 ms wall
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
|
||||
sustain start epoch 1791316287.337
|
||||
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_512.txt
|
||||
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
14
proto-newpow/mma-shadow/out/bench_512_ref.log
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 512 mm8 per hash = 4096 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 213 registers/thread, 8 resident blocks/SM at 1 warp(s)/block = 8 resident warps/SM (16.7% of 48)
|
||||
cache fill (GPU): 1.82 ms first, 1.78 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.58 ms first, 30.52 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 512 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 994.87 ms wall
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=512 regs=213 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
19
proto-newpow/mma-shadow/out/bench_8.log
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 8 mm8 per hash = 64 path = PTX mma.sync.m8n8k16.u8
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 30 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.92 ms first, 1.86 ms second (256 MiB)
|
||||
cache fill (host, one thread): 349.8 ms, GPU == host all words: PASS
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.61 ms first, 30.55 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 8 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 266.01 ms wall
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.75 ms -> 63.078 MH/s
|
||||
sustain start epoch 1791315790.867
|
||||
sustain end epoch 1791315815.870: 94 batches in 25.00 s -> 63.074 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_8.txt
|
||||
SUMMARY R=8 regs=30 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.074 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
14
proto-newpow/mma-shadow/out/bench_8_ref.log
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
mma-shadow bench pack "igneum-genesis" class mx8+mm8xR R = 8 mm8 per hash = 64 path = reference (shuffles + byte products)
|
||||
mm8 table seed 0x79f1fc5b6ed6112e (first steps: see mm8_block.h)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24092 MiB) SM clock 2520 MHz, mem clock 10501 MHz, bus 384 bits, L2 72 MiB, max 1536 threads/SM
|
||||
CUDA: driver 12.8, runtime 12.8
|
||||
igneum_hash_info: 77 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache fill (GPU): 1.81 ms first, 1.79 ms second (256 MiB)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset build (GPU, 1024 MiB): 30.59 ms first, 30.52 ms second
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
pack vectors: not applicable at R = 8 (checked at R = 0 only)
|
||||
warm-up batch: 16777216 hashes in 265.99 ms wall
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=8 regs=77 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
1024
proto-newpow/mma-shadow/out/dump_0.txt
Normal file
1024
proto-newpow/mma-shadow/out/dump_128.txt
Normal file
1024
proto-newpow/mma-shadow/out/dump_32.txt
Normal file
1024
proto-newpow/mma-shadow/out/dump_512.txt
Normal file
1024
proto-newpow/mma-shadow/out/dump_8.txt
Normal file
30
proto-newpow/mma-shadow/out/power_0.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
15.15 W, 210 MHz, 45, 0 %, 2026/10/06 19:42:24.816
|
||||
48.43 W, 2520 MHz, 47, 0 %, 2026/10/06 19:42:25.847
|
||||
106.67 W, 2670 MHz, 51, 0 %, 2026/10/06 19:42:26.847
|
||||
156.77 W, 2670 MHz, 52, 0 %, 2026/10/06 19:42:27.848
|
||||
200.24 W, 2670 MHz, 53, 100 %, 2026/10/06 19:42:28.848
|
||||
200.29 W, 2670 MHz, 53, 100 %, 2026/10/06 19:42:29.848
|
||||
202.82 W, 2670 MHz, 54, 100 %, 2026/10/06 19:42:30.848
|
||||
201.27 W, 2670 MHz, 54, 100 %, 2026/10/06 19:42:31.849
|
||||
201.02 W, 2670 MHz, 54, 100 %, 2026/10/06 19:42:32.849
|
||||
201.02 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:33.849
|
||||
201.04 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:34.849
|
||||
200.96 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:35.850
|
||||
200.94 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:36.850
|
||||
200.96 W, 2670 MHz, 55, 100 %, 2026/10/06 19:42:37.850
|
||||
201.00 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:38.850
|
||||
201.02 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:39.851
|
||||
201.02 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:40.851
|
||||
201.12 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:41.851
|
||||
201.07 W, 2670 MHz, 56, 100 %, 2026/10/06 19:42:42.851
|
||||
201.02 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:43.852
|
||||
201.03 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:44.852
|
||||
201.03 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:45.852
|
||||
201.21 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:46.852
|
||||
201.31 W, 2670 MHz, 57, 100 %, 2026/10/06 19:42:47.853
|
||||
201.33 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:48.853
|
||||
201.33 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:49.853
|
||||
201.52 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:50.853
|
||||
201.41 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:51.854
|
||||
201.51 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:52.854
|
||||
201.49 W, 2670 MHz, 58, 100 %, 2026/10/06 19:42:53.854
|
||||
|
30
proto-newpow/mma-shadow/out/power_128.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
18.57 W, 210 MHz, 53, 0 %, 2026/10/06 19:44:46.034
|
||||
38.21 W, 2520 MHz, 54, 0 %, 2026/10/06 19:44:47.037
|
||||
108.02 W, 2670 MHz, 55, 0 %, 2026/10/06 19:44:48.037
|
||||
143.43 W, 2670 MHz, 59, 100 %, 2026/10/06 19:44:49.037
|
||||
207.98 W, 2670 MHz, 60, 100 %, 2026/10/06 19:44:50.038
|
||||
208.85 W, 2670 MHz, 60, 100 %, 2026/10/06 19:44:51.038
|
||||
209.59 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:52.038
|
||||
210.00 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:53.038
|
||||
210.08 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:54.039
|
||||
210.10 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:55.039
|
||||
209.96 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:56.039
|
||||
210.40 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:57.039
|
||||
210.88 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:58.040
|
||||
211.38 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:59.040
|
||||
211.56 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:00.040
|
||||
211.76 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:01.040
|
||||
212.68 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:02.040
|
||||
212.12 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:03.041
|
||||
211.51 W, 2670 MHz, 62, 100 %, 2026/10/06 19:45:04.041
|
||||
212.43 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:05.041
|
||||
212.79 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:06.041
|
||||
212.98 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:07.042
|
||||
212.89 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:08.042
|
||||
212.51 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:09.042
|
||||
212.93 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:10.042
|
||||
212.91 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:11.042
|
||||
213.05 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:12.043
|
||||
213.00 W, 2670 MHz, 63, 100 %, 2026/10/06 19:45:13.043
|
||||
213.23 W, 2670 MHz, 64, 100 %, 2026/10/06 19:45:14.043
|
||||
213.14 W, 2670 MHz, 64, 100 %, 2026/10/06 19:45:15.043
|
||||
|
30
proto-newpow/mma-shadow/out/power_32.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
18.08 W, 210 MHz, 53, 0 %, 2026/10/06 19:43:48.466
|
||||
40.68 W, 2520 MHz, 55, 0 %, 2026/10/06 19:43:49.498
|
||||
106.86 W, 2670 MHz, 56, 52 %, 2026/10/06 19:43:50.498
|
||||
141.51 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:51.498
|
||||
203.49 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:52.499
|
||||
203.95 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:53.499
|
||||
204.46 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:54.499
|
||||
204.84 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:55.499
|
||||
204.88 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:56.500
|
||||
205.53 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:57.500
|
||||
205.87 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:58.500
|
||||
205.71 W, 2670 MHz, 62, 100 %, 2026/10/06 19:43:59.500
|
||||
206.17 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:00.500
|
||||
206.54 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:01.501
|
||||
206.21 W, 2670 MHz, 61, 100 %, 2026/10/06 19:44:02.501
|
||||
206.09 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:03.504
|
||||
206.43 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:04.504
|
||||
206.97 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:05.504
|
||||
207.32 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:06.504
|
||||
207.58 W, 2670 MHz, 62, 100 %, 2026/10/06 19:44:07.505
|
||||
207.28 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:08.505
|
||||
207.92 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:09.505
|
||||
208.31 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:10.505
|
||||
208.17 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:11.505
|
||||
207.89 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:12.506
|
||||
208.03 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:13.506
|
||||
208.60 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:14.506
|
||||
208.41 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:15.506
|
||||
208.39 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:16.508
|
||||
208.19 W, 2670 MHz, 63, 100 %, 2026/10/06 19:44:17.508
|
||||
|
30
proto-newpow/mma-shadow/out/power_512.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
15.26 W, 210 MHz, 46, 0 %, 2026/10/06 19:51:22.484
|
||||
32.16 W, 2520 MHz, 48, 0 %, 2026/10/06 19:51:23.516
|
||||
84.22 W, 2670 MHz, 50, 100 %, 2026/10/06 19:51:24.517
|
||||
138.29 W, 2670 MHz, 54, 100 %, 2026/10/06 19:51:25.517
|
||||
214.64 W, 2670 MHz, 55, 100 %, 2026/10/06 19:51:26.517
|
||||
215.01 W, 2670 MHz, 55, 100 %, 2026/10/06 19:51:27.517
|
||||
215.23 W, 2670 MHz, 56, 100 %, 2026/10/06 19:51:28.517
|
||||
215.43 W, 2670 MHz, 56, 100 %, 2026/10/06 19:51:29.518
|
||||
215.27 W, 2670 MHz, 56, 100 %, 2026/10/06 19:51:30.518
|
||||
215.35 W, 2670 MHz, 57, 100 %, 2026/10/06 19:51:31.518
|
||||
215.48 W, 2670 MHz, 57, 100 %, 2026/10/06 19:51:32.518
|
||||
215.41 W, 2670 MHz, 57, 100 %, 2026/10/06 19:51:33.518
|
||||
215.39 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:34.518
|
||||
215.49 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:35.519
|
||||
215.67 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:36.519
|
||||
215.71 W, 2670 MHz, 58, 100 %, 2026/10/06 19:51:37.519
|
||||
215.68 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:38.519
|
||||
215.70 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:39.519
|
||||
215.75 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:40.519
|
||||
215.72 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:41.520
|
||||
215.80 W, 2670 MHz, 59, 100 %, 2026/10/06 19:51:42.520
|
||||
215.90 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:43.520
|
||||
215.95 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:44.520
|
||||
216.10 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:45.520
|
||||
216.21 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:46.521
|
||||
216.13 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:47.521
|
||||
216.06 W, 2670 MHz, 60, 100 %, 2026/10/06 19:51:48.521
|
||||
216.07 W, 2670 MHz, 61, 100 %, 2026/10/06 19:51:49.521
|
||||
216.16 W, 2670 MHz, 61, 100 %, 2026/10/06 19:51:50.521
|
||||
216.10 W, 2670 MHz, 61, 100 %, 2026/10/06 19:51:51.521
|
||||
|
30
proto-newpow/mma-shadow/out/power_8.csv
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
16.99 W, 210 MHz, 51, 0 %, 2026/10/06 19:43:06.035
|
||||
40.60 W, 2520 MHz, 52, 0 %, 2026/10/06 19:43:07.067
|
||||
106.97 W, 2670 MHz, 53, 10 %, 2026/10/06 19:43:08.067
|
||||
137.95 W, 2670 MHz, 57, 100 %, 2026/10/06 19:43:09.068
|
||||
201.17 W, 2670 MHz, 58, 100 %, 2026/10/06 19:43:10.068
|
||||
201.47 W, 2670 MHz, 58, 100 %, 2026/10/06 19:43:11.068
|
||||
201.65 W, 2670 MHz, 58, 100 %, 2026/10/06 19:43:12.068
|
||||
201.97 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:13.068
|
||||
202.25 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:14.069
|
||||
202.06 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:15.069
|
||||
202.28 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:16.069
|
||||
202.47 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:17.069
|
||||
202.62 W, 2670 MHz, 59, 100 %, 2026/10/06 19:43:18.069
|
||||
202.62 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:19.070
|
||||
203.00 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:20.070
|
||||
203.19 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:21.070
|
||||
203.83 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:22.070
|
||||
203.66 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:23.070
|
||||
203.32 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:24.070
|
||||
203.21 W, 2670 MHz, 60, 100 %, 2026/10/06 19:43:25.071
|
||||
203.42 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:26.071
|
||||
203.90 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:27.071
|
||||
203.79 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:28.071
|
||||
203.96 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:29.071
|
||||
204.32 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:30.072
|
||||
204.51 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:31.072
|
||||
204.61 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:32.072
|
||||
204.72 W, 2670 MHz, 61, 100 %, 2026/10/06 19:43:33.072
|
||||
205.39 W, 2670 MHz, 62, 100 %, 2026/10/06 19:43:34.073
|
||||
205.90 W, 2670 MHz, 62, 100 %, 2026/10/06 19:43:35.073
|
||||
|
5
proto-newpow/mma-shadow/out/power_idle.csv
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
15.34 W, 210 MHz, 45, 0 %
|
||||
15.21 W, 210 MHz, 45, 0 %
|
||||
15.12 W, 210 MHz, 45, 0 %
|
||||
15.05 W, 210 MHz, 45, 0 %
|
||||
15.02 W, 210 MHz, 45, 0 %
|
||||
|
23
proto-newpow/mma-shadow/out/ptxas_0.txt
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
kernel_mm8.cu(98): warning #550-D: variable "lane" was set but never used
|
||||
uint32_t lane = threadIdx.x & 31u;
|
||||
^
|
||||
|
||||
Remark: The warnings can be suppressed with "-diag-suppress <warning-number>"
|
||||
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.262 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.536 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.384 ms
|
||||
23
proto-newpow/mma-shadow/out/ptxas_0_ref.txt
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
kernel_mm8.cu(98): warning #550-D: variable "lane" was set but never used
|
||||
uint32_t lane = threadIdx.x & 31u;
|
||||
^
|
||||
|
||||
Remark: The warnings can be suppressed with "-diag-suppress <warning-number>"
|
||||
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.413 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.749 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.256 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_128.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 32 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 27.671 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 11.908 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.275 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_128_ref.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 175 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 9367.163 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.495 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.300 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_32.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 15.442 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.497 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.272 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_32_ref.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 151 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 520.236 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.442 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.444 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_512.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 36 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 83.096 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 11.882 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.282 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_512_ref.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 213 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 273654.000 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.613 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.467 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_8.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 30 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 13.112 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 12.461 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.257 ms
|
||||
17
proto-newpow/mma-shadow/out/ptxas_8_ref.txt
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ptxas info : 0 bytes gmem
|
||||
ptxas info : 0 bytes gmem
|
||||
ptxas info : Compiling entry function '_Z11igneum_hashPKjPmjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z11igneum_hashPKjPmjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 77 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
ptxas info : Compile time = 78.215 ms
|
||||
ptxas info : Compiling entry function '_Z12igneum_buildPjPKjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z12igneum_buildPjPKjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 40 registers, used 0 barriers, 372 bytes cmem[0]
|
||||
ptxas info : Compile time = 11.877 ms
|
||||
ptxas info : Compiling entry function '_Z17igneum_cache_fillPjj' for 'sm_89'
|
||||
ptxas info : Function properties for _Z17igneum_cache_fillPjj
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 43 registers, used 0 barriers, 364 bytes cmem[0]
|
||||
ptxas info : Compile time = 6.236 ms
|
||||
7
proto-newpow/mma-shadow/out/results.md
Normal file
|
|
@ -0,0 +1,7 @@
|
|||
| R | mm8/hash | MH/s (GPU) | ratio to R=0 | W mean | SM MHz | uJ/hash | max C | fingerprint | PTX == ref | CPU == GPU | verifier ms/unit (R=0, R, delta) | regs | blocks/SM |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 | yes | 1024 of 1024 | 10.179, 10.076, -0.103 | 29 | 24 |
|
||||
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.926, 9.974, 0.048 | 30 | 24 |
|
||||
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.075, 10.240, 0.165 | 29 | 24 |
|
||||
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.183, 11.326, 1.142 | 32 | 24 |
|
||||
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.887, 14.276, 4.389 | 36 | 24 |
|
||||
130
proto-newpow/mma-shadow/out/run.log
Normal file
|
|
@ -0,0 +1,130 @@
|
|||
== Tue Oct 6 19:42:12 UTC 2026 host 5c32e87a3fb6 ==
|
||||
name, driver_version, power.draw [W], clocks.current.sm [MHz], temperature.gpu
|
||||
NVIDIA GeForce RTX 4090, 570.172.08, 15.34 W, 210 MHz, 45
|
||||
idle baseline (5 samples):
|
||||
15.34 W, 210 MHz, 45, 0 %
|
||||
15.21 W, 210 MHz, 45, 0 %
|
||||
15.12 W, 210 MHz, 45, 0 %
|
||||
15.05 W, 210 MHz, 45, 0 %
|
||||
15.02 W, 210 MHz, 45, 0 %
|
||||
== build verify_ref (gcc -O2)
|
||||
== R=0: build bench_0 (PTX path) and bench_0_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=0: timed run with power sampling
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
verify warp base 0 (pack vector, standalone): PASS
|
||||
verify warp base 4096 (pack vector, standalone): PASS
|
||||
verify warp base 1000000 (pack vector, standalone): PASS
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
verify warp base 0 (pack vector, in batch): PASS
|
||||
verify warp base 4096 (pack vector, in batch): PASS
|
||||
verify warp base 1000000 (pack vector, in batch): PASS
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.74 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315774.658: 94 batches in 25.00 s -> 63.075 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_0.txt
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.075 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
== R=0: reference path fingerprint
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
FPCHECK R=0 PTX == ref: yes (7c28cfb06c5c65a9)
|
||||
== R=0: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
self-test R=0 vector base 0: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 4096: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 1000000: PASS (0 lanes differ)
|
||||
1024 of 1024 lanes equal (PASS) [10.152 ms per unit during the check]
|
||||
verifier: 10.179 ms per unit at R=0, 10.076 ms per unit at R=0, mm8 block delta -0.103 ms per unit (0.00 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=0 equal=1024 of=1024 ms_r0=10.179 ms_r=10.076 delta=-0.103
|
||||
== R=8: build bench_8 (PTX path) and bench_8_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 30 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=8: timed run with power sampling
|
||||
igneum_hash_info: 30 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.75 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315815.870: 94 batches in 25.00 s -> 63.074 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_8.txt
|
||||
SUMMARY R=8 regs=30 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.074 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=8: reference path fingerprint
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=8 regs=77 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=8 PTX == ref: yes (06fc2593bfb94b4f)
|
||||
== R=8: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [10.106 ms per unit during the check]
|
||||
verifier: 9.926 ms per unit at R=0, 9.974 ms per unit at R=8, mm8 block delta 0.048 ms per unit (0.74 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=8 equal=1024 of=1024 ms_r0=9.926 ms_r=9.974 delta=0.048
|
||||
== R=32: build bench_32 (PTX path) and bench_32_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=32: timed run with power sampling
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.75 ms -> 63.078 MH/s (GPU time), wall 2659.77 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315858.296: 94 batches in 25.00 s -> 63.072 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_32.txt
|
||||
SUMMARY R=32 regs=29 blocksPerSM=24 mhs=63.078 mhs_wall=63.078 mhs_sustain=63.072 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=32: reference path fingerprint
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=32 PTX == ref: yes (26e83a65f519c865)
|
||||
== R=32: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [10.432 ms per unit during the check]
|
||||
verifier: 10.075 ms per unit at R=0, 10.240 ms per unit at R=32, mm8 block delta 0.165 ms per unit (0.64 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=32 equal=1024 of=1024 ms_r0=10.075 ms_r=10.240 delta=0.165
|
||||
== R=128: build bench_128 (PTX path) and bench_128_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 32 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=128: timed run with power sampling
|
||||
igneum_hash_info: 32 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.82 ms -> 63.077 MH/s (GPU time), wall 2659.83 ms -> 63.076 MH/s
|
||||
sustain end epoch 1791315915.893: 94 batches in 25.02 s -> 63.034 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_128.txt
|
||||
SUMMARY R=128 regs=32 blocksPerSM=24 mhs=63.077 mhs_wall=63.076 mhs_sustain=63.034 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=128: reference path fingerprint
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=128 regs=175 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=128 PTX == ref: yes (42223c2113188335)
|
||||
== R=128: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [11.349 ms per unit during the check]
|
||||
verifier: 10.183 ms per unit at R=0, 11.326 ms per unit at R=128, mm8 block delta 1.142 ms per unit (1.12 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=128 equal=1024 of=1024 ms_r0=10.183 ms_r=11.326 delta=1.142
|
||||
== R=512: build bench_512 (PTX path) and bench_512_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 36 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=512: timed run with power sampling
|
||||
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
|
||||
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_512.txt
|
||||
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=512: reference path fingerprint
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=512 regs=213 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=512 PTX == ref: yes (02b7002d747f3711)
|
||||
== R=512: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [14.440 ms per unit during the check]
|
||||
verifier: 9.887 ms per unit at R=0, 14.276 ms per unit at R=512, mm8 block delta 4.389 ms per unit (1.07 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=512 equal=1024 of=1024 ms_r0=9.887 ms_r=14.276 delta=4.389
|
||||
== Tue Oct 6 19:51:57 UTC 2026 done
|
||||
137
proto-newpow/mma-shadow/out/run_stdout.log
Normal file
|
|
@ -0,0 +1,137 @@
|
|||
== Tue Oct 6 19:42:12 UTC 2026 host 5c32e87a3fb6 ==
|
||||
name, driver_version, power.draw [W], clocks.current.sm [MHz], temperature.gpu
|
||||
NVIDIA GeForce RTX 4090, 570.172.08, 15.34 W, 210 MHz, 45
|
||||
idle baseline (5 samples):
|
||||
15.34 W, 210 MHz, 45, 0 %
|
||||
15.21 W, 210 MHz, 45, 0 %
|
||||
15.12 W, 210 MHz, 45, 0 %
|
||||
15.05 W, 210 MHz, 45, 0 %
|
||||
15.02 W, 210 MHz, 45, 0 %
|
||||
== build verify_ref (gcc -O2)
|
||||
== R=0: build bench_0 (PTX path) and bench_0_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=0: timed run with power sampling
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
verify warp base 0 (pack vector, standalone): PASS
|
||||
verify warp base 4096 (pack vector, standalone): PASS
|
||||
verify warp base 1000000 (pack vector, standalone): PASS
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
verify warp base 0 (pack vector, in batch): PASS
|
||||
verify warp base 4096 (pack vector, in batch): PASS
|
||||
verify warp base 1000000 (pack vector, in batch): PASS
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.74 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315774.658: 94 batches in 25.00 s -> 63.075 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_0.txt
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.075 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
== R=0: reference path fingerprint
|
||||
fingerprint: 7c28cfb06c5c65a9 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes) == 7c28cfb06c5c65a9 PASS
|
||||
SUMMARY R=0 regs=29 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=7c28cfb06c5c65a9 cache=PASS dataset=PASS vectors=PASS
|
||||
OVERALL: PASS
|
||||
FPCHECK R=0 PTX == ref: yes (7c28cfb06c5c65a9)
|
||||
== R=0: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
self-test R=0 vector base 0: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 4096: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 1000000: PASS (0 lanes differ)
|
||||
1024 of 1024 lanes equal (PASS) [10.152 ms per unit during the check]
|
||||
verifier: 10.179 ms per unit at R=0, 10.076 ms per unit at R=0, mm8 block delta -0.103 ms per unit (0.00 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=0 equal=1024 of=1024 ms_r0=10.179 ms_r=10.076 delta=-0.103
|
||||
== R=8: build bench_8 (PTX path) and bench_8_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 30 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=8: timed run with power sampling
|
||||
igneum_hash_info: 30 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.73 ms -> 63.079 MH/s (GPU time), wall 2659.75 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315815.870: 94 batches in 25.00 s -> 63.074 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_8.txt
|
||||
SUMMARY R=8 regs=30 blocksPerSM=24 mhs=63.079 mhs_wall=63.078 mhs_sustain=63.074 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=8: reference path fingerprint
|
||||
fingerprint: 06fc2593bfb94b4f (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=8 regs=77 blocksPerSM=24 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=06fc2593bfb94b4f cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=8 PTX == ref: yes (06fc2593bfb94b4f)
|
||||
== R=8: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [10.106 ms per unit during the check]
|
||||
verifier: 9.926 ms per unit at R=0, 9.974 ms per unit at R=8, mm8 block delta 0.048 ms per unit (0.74 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=8 equal=1024 of=1024 ms_r0=9.926 ms_r=9.974 delta=0.048
|
||||
== R=32: build bench_32 (PTX path) and bench_32_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 29 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=32: timed run with power sampling
|
||||
igneum_hash_info: 29 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.75 ms -> 63.078 MH/s (GPU time), wall 2659.77 ms -> 63.078 MH/s
|
||||
sustain end epoch 1791315858.296: 94 batches in 25.00 s -> 63.072 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_32.txt
|
||||
SUMMARY R=32 regs=29 blocksPerSM=24 mhs=63.078 mhs_wall=63.078 mhs_sustain=63.072 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=32: reference path fingerprint
|
||||
fingerprint: 26e83a65f519c865 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=32 regs=151 blocksPerSM=12 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=26e83a65f519c865 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=32 PTX == ref: yes (26e83a65f519c865)
|
||||
== R=32: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [10.432 ms per unit during the check]
|
||||
verifier: 10.075 ms per unit at R=0, 10.240 ms per unit at R=32, mm8 block delta 0.165 ms per unit (0.64 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=32 equal=1024 of=1024 ms_r0=10.075 ms_r=10.240 delta=0.165
|
||||
== R=128: build bench_128 (PTX path) and bench_128_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 32 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=128: timed run with power sampling
|
||||
igneum_hash_info: 32 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.82 ms -> 63.077 MH/s (GPU time), wall 2659.83 ms -> 63.076 MH/s
|
||||
sustain end epoch 1791315915.893: 94 batches in 25.02 s -> 63.034 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_128.txt
|
||||
SUMMARY R=128 regs=32 blocksPerSM=24 mhs=63.077 mhs_wall=63.076 mhs_sustain=63.034 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=128: reference path fingerprint
|
||||
fingerprint: 42223c2113188335 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=128 regs=175 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=42223c2113188335 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=128 PTX == ref: yes (42223c2113188335)
|
||||
== R=128: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [11.349 ms per unit during the check]
|
||||
verifier: 10.183 ms per unit at R=0, 11.326 ms per unit at R=128, mm8 block delta 1.142 ms per unit (1.12 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=128 equal=1024 of=1024 ms_r0=10.183 ms_r=11.326 delta=1.142
|
||||
== R=512: build bench_512 (PTX path) and bench_512_ref
|
||||
0 bytes stack frame, 0 bytes spill stores, 0 bytes spill loads
|
||||
ptxas info : Used 36 registers, used 0 barriers, 376 bytes cmem[0]
|
||||
== R=512: timed run with power sampling
|
||||
igneum_hash_info: 36 registers/thread, 24 resident blocks/SM at 1 warp(s)/block = 24 resident warps/SM (50.0% of 48)
|
||||
cache check: PASS (device FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head PASS, last line PASS)
|
||||
dataset self-test: PASS (head 16 PASS, [MASK] PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
timed: 10 batches x 16777216 hashes: GPU 2659.85 ms -> 63.076 MH/s (GPU time), wall 2659.86 ms -> 63.075 MH/s
|
||||
sustain end epoch 1791316312.340: 94 batches in 25.00 s -> 63.073 MH/s (wall, incl. sync)
|
||||
dump: 32 warps (1024 lanes) written to out/dump_512.txt
|
||||
SUMMARY R=512 regs=36 blocksPerSM=24 mhs=63.076 mhs_wall=63.075 mhs_sustain=63.073 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
== R=512: reference path fingerprint
|
||||
fingerprint: 02b7002d747f3711 (FNV-1a 64 over the 2^24 outputs at base nonce 0, little-endian u64 bytes)
|
||||
SUMMARY R=512 regs=213 blocksPerSM=8 mhs=0.000 mhs_wall=0.000 mhs_sustain=0.000 fingerprint=02b7002d747f3711 cache=PASS dataset=PASS vectors=n/a
|
||||
OVERALL: PASS
|
||||
FPCHECK R=512 PTX == ref: yes (02b7002d747f3711)
|
||||
== R=512: CPU reference on the 32-warp dump (taskset -c 2)
|
||||
1024 of 1024 lanes equal (PASS) [14.440 ms per unit during the check]
|
||||
verifier: 9.887 ms per unit at R=0, 14.276 ms per unit at R=512, mm8 block delta 4.389 ms per unit (1.07 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=512 equal=1024 of=1024 ms_r0=9.887 ms_r=14.276 delta=4.389
|
||||
== Tue Oct 6 19:51:57 UTC 2026 done
|
||||
| R | mm8/hash | MH/s (GPU) | ratio to R=0 | W mean | SM MHz | uJ/hash | max C | fingerprint | PTX == ref | CPU == GPU | verifier ms/unit (R=0, R, delta) | regs | blocks/SM |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| 0 | 0 | 63.08 | 1.000 | 201.2 | 2670 | 3.19 | 58 | 7c28cfb06c5c65a9 | yes | 1024 of 1024 | 10.179, 10.076, -0.103 | 29 | 24 |
|
||||
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2670 | 3.24 | 62 | 06fc2593bfb94b4f | yes | 1024 of 1024 | 9.926, 9.974, 0.048 | 30 | 24 |
|
||||
| 32 | 256 | 63.08 | 1.000 | 207.7 | 2670 | 3.29 | 63 | 26e83a65f519c865 | yes | 1024 of 1024 | 10.075, 10.240, 0.165 | 29 | 24 |
|
||||
| 128 | 1024 | 63.08 | 1.000 | 212.7 | 2670 | 3.37 | 64 | 42223c2113188335 | yes | 1024 of 1024 | 10.183, 11.326, 1.142 | 32 | 24 |
|
||||
| 512 | 4096 | 63.08 | 1.000 | 215.9 | 2670 | 3.42 | 61 | 02b7002d747f3711 | yes | 1024 of 1024 | 9.887, 14.276, 4.389 | 36 | 24 |
|
||||
8
proto-newpow/mma-shadow/out/verify_0.log
Normal file
|
|
@ -0,0 +1,8 @@
|
|||
cache: host fill 528.5 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
self-test R=0 vector base 0: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 4096: PASS (0 lanes differ)
|
||||
self-test R=0 vector base 1000000: PASS (0 lanes differ)
|
||||
dump: 1024 lines, 32 units, R = 0 (0 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [10.152 ms per unit during the check]
|
||||
verifier: 10.179 ms per unit at R=0, 10.076 ms per unit at R=0, mm8 block delta -0.103 ms per unit (0.00 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=0 equal=1024 of=1024 ms_r0=10.179 ms_r=10.076 delta=-0.103
|
||||
5
proto-newpow/mma-shadow/out/verify_128.log
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
cache: host fill 533.2 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
dump: 1024 lines, 32 units, R = 128 (1024 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [11.349 ms per unit during the check]
|
||||
verifier: 10.183 ms per unit at R=0, 11.326 ms per unit at R=128, mm8 block delta 1.142 ms per unit (1.12 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=128 equal=1024 of=1024 ms_r0=10.183 ms_r=11.326 delta=1.142
|
||||
5
proto-newpow/mma-shadow/out/verify_32.log
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
cache: host fill 535.4 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
dump: 1024 lines, 32 units, R = 32 (256 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [10.432 ms per unit during the check]
|
||||
verifier: 10.075 ms per unit at R=0, 10.240 ms per unit at R=32, mm8 block delta 0.165 ms per unit (0.64 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=32 equal=1024 of=1024 ms_r0=10.075 ms_r=10.240 delta=0.165
|
||||
5
proto-newpow/mma-shadow/out/verify_512.log
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
cache: host fill 530.6 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
dump: 1024 lines, 32 units, R = 512 (4096 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [14.440 ms per unit during the check]
|
||||
verifier: 9.887 ms per unit at R=0, 14.276 ms per unit at R=512, mm8 block delta 4.389 ms per unit (1.07 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=512 equal=1024 of=1024 ms_r0=9.887 ms_r=14.276 delta=4.389
|
||||
5
proto-newpow/mma-shadow/out/verify_8.log
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
cache: host fill 558.2 ms, FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
dump: 1024 lines, 32 units, R = 8 (64 mm8 per hash)
|
||||
1024 of 1024 lanes equal (PASS) [10.106 ms per unit during the check]
|
||||
verifier: 9.926 ms per unit at R=0, 9.974 ms per unit at R=8, mm8 block delta 0.048 ms per unit (0.74 us per mm8 step), averaged over 32 units, one core
|
||||
VERIFY R=8 equal=1024 of=1024 ms_r0=9.926 ms_r=9.974 delta=0.048
|
||||
66
proto-newpow/mma-shadow/program.h
Normal file
|
|
@ -0,0 +1,66 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-genesis"
|
||||
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0xe323b9dcaf283a6full
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "mx8"
|
||||
#define IGNEUM_CLASS_MIXER_MULT 8
|
||||
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
|
||||
#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIXER_MULT 8 // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)
|
||||
#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
74
proto-newpow/mma-shadow/ref_program.inc
Normal file
|
|
@ -0,0 +1,74 @@
|
|||
// Generated by gen_ref_program.py from the pack's kernel.cu. Do not edit by hand.
|
||||
// Expects: uint32_t R[8][32], T[32], SEL[32]; const uint32_t* cache; uint32_t mask; macro L = for (int l = 0; l < 32; ++l).
|
||||
L { R[2][l] = R[3][l] * R[4][l] + R[2][l]; } // 0 mad
|
||||
L { R[2][l] = R[1][l] * R[1][l] + R[2][l]; } // 1 mad
|
||||
L { R[2][l] = R[3][l] * R[2][l] + R[2][l]; } // 2 mad
|
||||
L { R[3][l] = R[3][l] ^ R[5][l]; } // 3 xor
|
||||
L { R[7][l] = R[7][l] ^ mh_word(cache, R[2][l] & mask); } // 4 load
|
||||
L { R[5][l] = R[5][l] ^ mh_word(cache, R[7][l] & mask); } // 5 load
|
||||
memcpy(T, R[4], sizeof T);
|
||||
L { R[1][l] = R[1][l] ^ T[l ^ 8]; } // 6 shfl
|
||||
memcpy(T, R[3], sizeof T);
|
||||
L { R[7][l] = R[7][l] ^ T[l ^ 8]; } // 7 shfl
|
||||
L { R[1][l] = umulhi(R[1][l], R[5][l]); } // 8 mulhi
|
||||
L { R[6][l] = rotr_var(R[6][l], R[3][l]); } // 9 rotr
|
||||
L { R[3][l] = R[3][l] | R[4][l]; } // 10 or
|
||||
L { R[4][l] = R[4][l] ^ mh_word(cache, R[3][l] & mask); } // 11 load
|
||||
L { R[0][l] = umulhi(R[0][l], R[4][l]); } // 12 mulhi
|
||||
L { R[5][l] = R[5][l] + R[1][l] + ((((SEL[l] >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); } // 13 add
|
||||
L { R[0][l] = R[0][l] ^ mh_word(cache, R[4][l] & mask); } // 14 load
|
||||
L { R[2][l] = R[2][l] - R[4][l]; } // 15 sub
|
||||
L { R[2][l] = R[2][l] ^ mh_word(cache, R[0][l] & mask); } // 16 load
|
||||
L { R[7][l] = R[7][l] ^ mh_word(cache, R[2][l] & mask); } // 17 load
|
||||
memcpy(T, R[3], sizeof T);
|
||||
L { R[7][l] = R[7][l] ^ T[l ^ 4]; } // 18 shfl
|
||||
L { R[5][l] = R[5][l] * R[0][l]; } // 19 mul
|
||||
memcpy(T, R[4], sizeof T);
|
||||
L { R[3][l] = R[3][l] ^ T[l ^ 2]; } // 20 shfl
|
||||
memcpy(T, R[4], sizeof T);
|
||||
L { R[2][l] = R[2][l] ^ T[l ^ 16]; } // 21 shfl
|
||||
L { R[6][l] = umulhi(R[6][l], R[2][l]); } // 22 mulhi
|
||||
L { R[6][l] = R[6][l] ^ mh_word(cache, R[1][l] & mask); } // 23 load
|
||||
L { R[5][l] = R[5][l] * R[0][l]; } // 24 mul
|
||||
L { R[5][l] = rotl_imm(R[5][l], 19u); } // 25 rotl
|
||||
memcpy(T, R[6], sizeof T);
|
||||
L { R[7][l] = R[7][l] ^ T[l ^ 2]; } // 26 shfl
|
||||
L { R[0][l] = R[0][l] ^ R[5][l]; } // 27 xor
|
||||
L { R[0][l] = R[0][l] ^ R[4][l]; } // 28 xor
|
||||
L { R[3][l] = R[3][l] - R[0][l]; } // 29 sub
|
||||
L { R[5][l] = R[5][l] * R[1][l]; } // 30 mul
|
||||
L { R[7][l] = R[7][l] ^ mh_word(cache, R[2][l] & mask); } // 31 load
|
||||
L { R[1][l] = R[1][l] ^ mh_word(cache, R[0][l] & mask); } // 32 load
|
||||
L { R[5][l] = R[5][l] ^ R[6][l]; } // 33 xor
|
||||
L { R[5][l] = R[5][l] ^ mh_word(cache, R[1][l] & mask); } // 34 load
|
||||
L { R[0][l] = umulhi(R[0][l], R[5][l]); } // 35 mulhi
|
||||
memcpy(T, R[2], sizeof T);
|
||||
L { R[5][l] = R[5][l] ^ T[l ^ 4]; } // 36 shfl
|
||||
L { R[7][l] = R[7][l] ^ mh_word(cache, R[0][l] & mask); } // 37 load
|
||||
L { R[3][l] = R[3][l] + R[1][l] + ((((SEL[l] >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); } // 38 add
|
||||
memcpy(T, R[5], sizeof T);
|
||||
L { R[1][l] = R[1][l] ^ T[l ^ 4]; } // 39 shfl
|
||||
L { R[2][l] = R[2][l] ^ R[5][l]; } // 40 xor
|
||||
L { R[3][l] = R[6][l] * R[3][l] + R[3][l]; } // 41 mad
|
||||
L { R[6][l] = R[6][l] - R[7][l]; } // 42 sub
|
||||
L { R[7][l] = R[7][l] ^ R[0][l]; } // 43 xor
|
||||
L { R[1][l] = R[1][l] ^ mh_word(cache, R[7][l] & mask); } // 44 load
|
||||
L { R[2][l] = R[2][l] * R[3][l]; } // 45 mul
|
||||
L { R[1][l] = umulhi(R[1][l], R[5][l]); } // 46 mulhi
|
||||
L { R[4][l] = R[4][l] - R[3][l]; } // 47 sub
|
||||
L { R[2][l] = rotr_var(R[2][l], R[6][l]); } // 48 rotr
|
||||
L { R[3][l] = R[3][l] ^ mh_word(cache, R[5][l] & mask); } // 49 load
|
||||
L { R[1][l] = R[1][l] + R[5][l] + ((((SEL[l] >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); } // 50 add
|
||||
L { R[0][l] = R[0][l] * R[2][l]; } // 51 mul
|
||||
L { R[0][l] = R[0][l] + R[2][l] + ((((SEL[l] >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); } // 52 add
|
||||
L { R[1][l] = R[1][l] + R[0][l] + ((((SEL[l] >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); } // 53 add
|
||||
L { R[7][l] = rotl_imm(R[7][l], 14u); } // 54 rotl
|
||||
L { R[3][l] = R[3][l] + R[7][l] + ((((SEL[l] >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); } // 55 add
|
||||
L { R[6][l] = R[6][l] ^ mh_word(cache, R[7][l] & mask); } // 56 load
|
||||
L { R[1][l] = rotr_var(R[1][l], R[5][l]); } // 57 rotr
|
||||
L { R[5][l] = R[5][l] ^ mh_word(cache, R[4][l] & mask); } // 58 load
|
||||
L { R[6][l] = R[6][l] ^ mh_word(cache, R[2][l] & mask); } // 59 load
|
||||
L { R[3][l] = R[5][l] * R[0][l] + R[3][l]; } // 60 mad
|
||||
L { R[5][l] = R[5][l] + R[7][l] + ((((SEL[l] >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); } // 61 add
|
||||
L { R[4][l] = R[4][l] + R[6][l] + ((((SEL[l] >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); } // 62 add
|
||||
L { R[5][l] = rotl_imm(R[5][l], 19u); } // 63 rotl
|
||||
43
proto-newpow/mma-shadow/run.sh
Executable file
|
|
@ -0,0 +1,43 @@
|
|||
#!/bin/bash
|
||||
# run.sh (proto-newpow/mma-shadow): build and measure the ladder R in {0, 8, 32, 128, 512} on the GPU box.
|
||||
# Runs under /root/horizon-newpow/mma-shadow. Every artefact lands in ./out/.
|
||||
set -u
|
||||
export PATH=/usr/local/cuda/bin:$PATH
|
||||
cd "$(dirname "$0")"
|
||||
LADDER="${LADDER:-0 8 32 128 512}"
|
||||
SUSTAIN="${SUSTAIN:-25}" # seconds of sustained hashing for the power meter (mean taken after the first 10 s)
|
||||
mkdir -p out
|
||||
echo "== $(date -u) host $(hostname) ==" | tee out/run.log
|
||||
nvidia-smi --query-gpu=name,driver_version,power.draw,clocks.sm,temperature.gpu --format=csv | tee -a out/run.log
|
||||
echo "idle baseline (5 samples):" | tee -a out/run.log
|
||||
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu --format=csv,noheader -l 1 > out/power_idle.csv & P=$!; sleep 5; kill $P 2>/dev/null; wait $P 2>/dev/null; cat out/power_idle.csv | tee -a out/run.log
|
||||
|
||||
echo "== build verify_ref (gcc -O2)" | tee -a out/run.log
|
||||
gcc -O2 -o verify_ref verify_ref.c 2>&1 | tee -a out/run.log
|
||||
|
||||
for R in $LADDER; do
|
||||
echo "== R=$R: build bench_$R (PTX path) and bench_${R}_ref" | tee -a out/run.log
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -Xptxas -v -o bench_$R bench.cu kernel_mm8.cu 2> out/ptxas_$R.txt || { echo "BUILD FAILED bench_$R"; cat out/ptxas_$R.txt; continue; }
|
||||
grep -A2 "igneum_hash" out/ptxas_$R.txt | grep -i "registers\|spill" | tee -a out/run.log
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -DIGNEUM_MM8_R=$R -DIGNEUM_MM8_REF -Xptxas -v -o bench_${R}_ref bench.cu kernel_mm8.cu 2> out/ptxas_${R}_ref.txt || { echo "BUILD FAILED bench_${R}_ref"; cat out/ptxas_${R}_ref.txt; }
|
||||
|
||||
echo "== R=$R: timed run with power sampling" | tee -a out/run.log
|
||||
nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu,timestamp --format=csv,noheader -l 1 > out/power_$R.csv &
|
||||
SMI=$!
|
||||
./bench_$R --batches 10 --sustain $SUSTAIN --dump out/dump_$R.txt 32 > out/bench_$R.log 2>&1
|
||||
kill $SMI 2>/dev/null; wait $SMI 2>/dev/null
|
||||
grep -E "igneum_hash_info|cache check|dataset self-test|verify warp|fingerprint|timed:|sustain end|SUMMARY|OVERALL|dump:" out/bench_$R.log | tee -a out/run.log
|
||||
|
||||
echo "== R=$R: reference path fingerprint" | tee -a out/run.log
|
||||
./bench_${R}_ref --fingerprint-only --no-host-cache > out/bench_${R}_ref.log 2>&1
|
||||
grep -E "fingerprint|OVERALL" out/bench_${R}_ref.log | tee -a out/run.log
|
||||
FP=$(grep -oE "fingerprint: [0-9a-f]{16}" out/bench_$R.log | cut -d' ' -f2)
|
||||
FPR=$(grep -oE "fingerprint: [0-9a-f]{16}" out/bench_${R}_ref.log | cut -d' ' -f2)
|
||||
if [ -n "$FP" ] && [ "$FP" = "$FPR" ]; then echo "FPCHECK R=$R PTX == ref: yes ($FP)" | tee -a out/run.log; else echo "FPCHECK R=$R PTX == ref: NO (ptx $FP ref $FPR)" | tee -a out/run.log; fi
|
||||
|
||||
echo "== R=$R: CPU reference on the 32-warp dump (taskset -c 2)" | tee -a out/run.log
|
||||
taskset -c 2 ./verify_ref out/dump_$R.txt $R --time $( [ "$R" = "0" ] && echo --self-test ) > out/verify_$R.log 2>&1
|
||||
grep -E "self-test|lanes equal|first mismatch|verifier:|VERIFY" out/verify_$R.log | tee -a out/run.log
|
||||
done
|
||||
echo "== $(date -u) done" | tee -a out/run.log
|
||||
python3 summarise.py | tee out/results.md
|
||||
45
proto-newpow/mma-shadow/summarise.py
Normal file
|
|
@ -0,0 +1,45 @@
|
|||
#!/usr/bin/env python3
|
||||
# summarise.py: builds the RESULTS table from out/*.log and out/power_*.csv (run by run.sh on the box).
|
||||
import re, glob, os
|
||||
from datetime import datetime
|
||||
rows = []
|
||||
def grab(pat, text, default=None):
|
||||
m = re.search(pat, text); return m.group(1) if m else default
|
||||
def power_mean(R, log):
|
||||
try:
|
||||
s0 = float(grab(r"sustain start epoch ([0-9.]+)", log)); s1 = float(grab(r"sustain end epoch ([0-9.]+)", log))
|
||||
except Exception:
|
||||
return None, None, None
|
||||
pw, clk, tmp = [], [], []
|
||||
for line in open("out/power_%s.csv" % R):
|
||||
parts = [p.strip() for p in line.split(",")]
|
||||
if len(parts) < 5: continue
|
||||
try:
|
||||
ts = datetime.strptime(parts[4], "%Y/%m/%d %H:%M:%S.%f").timestamp()
|
||||
except Exception:
|
||||
continue
|
||||
if ts >= s0 + 10 and ts <= s1:
|
||||
pw.append(float(parts[0].split()[0])); clk.append(float(parts[1].split()[0])); tmp.append(float(parts[2].split()[0]))
|
||||
if not pw: return None, None, None
|
||||
return sum(pw) / len(pw), sum(clk) / len(clk), max(tmp)
|
||||
base_mhs = None
|
||||
out = ["| R | mm8/hash | MH/s (GPU) | ratio to R=0 | W mean | SM MHz | uJ/hash | max C | fingerprint | PTX == ref | CPU == GPU | verifier ms/unit (R=0, R, delta) | regs | blocks/SM |", "|---|---|---|---|---|---|---|---|---|---|---|---|---|---|"]
|
||||
for Rs in sorted([int(os.path.basename(p)[6:-4]) for p in glob.glob("out/bench_*.log") if "_ref" not in p]):
|
||||
log = open("out/bench_%d.log" % Rs).read()
|
||||
summ = grab(r"(SUMMARY.*)", log, "")
|
||||
mhs = float(grab(r"mhs=([0-9.]+)", summ, "0")); regs = grab(r"regs=(\d+)", summ, "?"); bps = grab(r"blocksPerSM=(\d+)", summ, "?")
|
||||
fp = grab(r"fingerprint=([0-9a-f]{16})", summ, "?")
|
||||
if Rs == 0: base_mhs = mhs
|
||||
ratio = ("%.3f" % (mhs / base_mhs)) if base_mhs else "?"
|
||||
runlog = open("out/run.log").read()
|
||||
fpc = grab(r"FPCHECK R=%d PTX == ref: (\S+)" % Rs, runlog, "?")
|
||||
ver = open("out/verify_%d.log" % Rs).read() if os.path.exists("out/verify_%d.log" % Rs) else ""
|
||||
vs = grab(r"(VERIFY.*)", ver, "")
|
||||
eq = grab(r"equal=(\d+)", vs, "?"); of = grab(r"of=(\d+)", vs, "?")
|
||||
ms0 = grab(r"ms_r0=([0-9.]+)", vs, "?"); msr = grab(r"ms_r=([0-9.]+)", vs, "?"); dl = grab(r"delta=([0-9.-]+)", vs, "?")
|
||||
w, clk, tmx = power_mean(Rs, log)
|
||||
uj = ("%.2f" % (w / (mhs * 1e6) * 1e6)) if (w and mhs) else "?"
|
||||
out.append("| %d | %d | %.2f | %s | %s | %s | %s | %s | %s | %s | %s of %s | %s, %s, %s | %s | %s |" % (
|
||||
Rs, 8 * Rs, mhs, ratio, ("%.1f" % w) if w else "?", ("%.0f" % clk) if clk else "?", uj, ("%.0f" % tmx) if tmx else "?",
|
||||
fp, fpc, eq, of, ms0, msr, dl, regs, bps))
|
||||
print("\n".join(out))
|
||||
57
proto-newpow/mma-shadow/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x19b56348bc85304dull, 0xb08a9cfb44aa720full, 0xe1f8f627f780eff7ull, 0x2ff3e86ff1696161ull, 0xb65e0578c257e8acull, 0xf37b5705c2bebaaeull, 0xb19fef670982389eull, 0x2374331b28827a11ull,
|
||||
0xf492cdd05dda7f88ull, 0x37700f1385e19885ull, 0xa27524f3010b3e87ull, 0x6c8a24de8b938c43ull, 0x3dc017c8820cfd16ull, 0x5c80d37146657b68ull, 0x214082cac03f734aull, 0x1d67c665145d72f3ull,
|
||||
0x9fb0ae736eb87ff5ull, 0xcd7d416ae3a6be61ull, 0x0069cda85c6de58full, 0x0b7d82e717224baaull, 0x600960f6d7be30f0ull, 0xeb29032663b0c4d2ull, 0xf29583a5766408f1ull, 0x6c9532a99ad7314dull,
|
||||
0x9a949fd0959cc80full, 0xa20258520e7f6c25ull, 0xf2ab11f9bb032e38ull, 0xcc967bcd0c8d07c1ull, 0x37745267bb3231f2ull, 0x35a046048c2b69b3ull, 0xaa51834cd3f364f3ull, 0x359192708e4f754aull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x62fb132a9943127aull, 0x0b703e577e7f4ecaull, 0xf9f24f5522ce7593ull, 0x3cf5c516abc4332aull, 0xd25523f5f6d7a127ull, 0xd2081a002f983682ull, 0xbf46c54e9b3c4254ull, 0xca362e291e5e5f4dull,
|
||||
0x6039712f10f457a3ull, 0x8a34b7cabf97c23bull, 0xa473c6a2e0bf59bcull, 0x6cf3926513a4b069ull, 0x297ec2998376a40dull, 0x8efd7f601a8f28dbull, 0x8e72532dfdc1e544ull, 0x917c2b2ebe2a7e00ull,
|
||||
0x923fbb2d2f635c25ull, 0xce864ea5c0dedad9ull, 0x4b8ec7e874e446efull, 0x1b69b69465449196ull, 0x5ef3a8a6edb369cfull, 0x06c263ef9ce63fc4ull, 0x9c2048fd9d9e2639ull, 0x457fdd96ca4a138eull,
|
||||
0xd1904018b8d7b6e3ull, 0x8682312fb2e96ab8ull, 0xdc3257e0d0f979a5ull, 0xa51b0a8519d87db5ull, 0x334f08ec056e618bull, 0x3464ce71dc65119dull, 0x6a4d6df066332e04ull, 0x7d7866cb9cfca8ffull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x86b6cb0e13d89b03ull, 0x96299a3f19d7ef15ull, 0x67d2c55100d2f876ull, 0x0a4dfe97d671b728ull, 0x41e4489014d42595ull, 0xf11cb1958c0c0e82ull, 0xf8b70b0c0a03175full, 0x632299df87d5063eull,
|
||||
0xe198417776130492ull, 0x8ffc5449290d7be2ull, 0x5f2e264eb1311f1bull, 0x988376463ac88586ull, 0x83969eadda489c26ull, 0xbed0a2c3f255d306ull, 0x1a949d271961a819ull, 0x5bce06eb6984725cull,
|
||||
0x94d5d6a1b0ffd4e9ull, 0xf3c78bae6c2182b4ull, 0xb97e9fe1bbfcdd55ull, 0x70262d1d4c0eccb2ull, 0x1fc93b427dba28d9ull, 0x02b2e3c4317f2a2dull, 0x54d3d42a588edcb9ull, 0x79998677846e7cceull,
|
||||
0x486522a5425f821aull, 0x95fa88e933360e52ull, 0xc8bae2da2b883f6cull, 0xbe3eb610ad33614full, 0x20efb3c4de82907full, 0xd6b650cfedfb26b7ull, 0x8c24447a646dba26ull, 0x9c004678515e44ecull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0xfdad4319u, 0x1a7b68e1u, 0xde6db608u, 0x13d73892u, 0xd17f447au, 0xb2221ccfu, 0x9db004bdu, 0x57d7d367u,
|
||||
0xdbc4cf34u, 0x697c009au, 0xc43af1d4u, 0x97f12b2eu, 0x74c37cd0u, 0xc651ea15u, 0x665a6d29u, 0x22330a2du
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xa83e7aa6u;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0x3230bc7bu, 0x7fbfe2c9u, 0xb2690991u, 0x1745c7c5u, 0x0ab0ccafu, 0x1bf87d6bu, 0x160139fdu, 0x719817acu, 0x0155df4bu, 0xbe1e86c3u, 0x680bcd6cu, 0x79c3dc6cu, 0x181e7e5fu, 0x0713a109u, 0xc705dd9fu, 0x3933b7a8u, 0xdd1c0431u, 0x50522b30u, 0xa0020b38u, 0xbff39e96u, 0x21b67e18u, 0x740f8db3u, 0x2baba568u, 0x2c9bef83u, 0x0ad9b671u, 0xc4327869u, 0x7b4fd7d0u, 0x2c29965fu, 0xec56f15fu, 0x61111746u, 0x303a1d6eu, 0xbddcfd1au, 0xf829a355u, 0x6d5df2a9u, 0x01ab8e44u, 0x06d13507u, 0xda8dcfc6u, 0x01a703e1u, 0xafe7d2c1u, 0xc091c3a2u, 0xac1814feu, 0x6e6ff62au, 0x8fdf01bau, 0xdd3f7159u, 0xdfa0d75cu, 0x26684c35u, 0x7f441e63u, 0x88df2570u, 0x8aa4d5ebu, 0xcc816c05u, 0x434df890u, 0xcd392ad6u, 0x1ab4cb63u, 0x595926fau, 0x7cd76b41u, 0x20cb95c4u, 0x13cf823fu, 0xf9daf901u, 0xff9af40au, 0x2c7dfa51u, 0x871206dbu, 0x938c116cu, 0xb64bf199u, 0x5751f874u
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u,
|
||||
0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du,
|
||||
0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull;
|
||||
152
proto-newpow/mma-shadow/verify_ref.c
Normal file
|
|
@ -0,0 +1,152 @@
|
|||
// verify_ref.c (proto-newpow/mma-shadow): plain C CPU reference for the mx8+mm8xR prototype.
|
||||
// Reads a bench --dump file (lines: base lane value_hex, 32 lanes per warp), recomputes every warp with a
|
||||
// register-major 32-lane interpreter of the pack's program (ref_program.inc, generated from kernel.cu), lazy
|
||||
// dataset words through memhard.h's mh_word over a host-filled 256 MiB cache, the mm8 block in the spec layout,
|
||||
// then the fold. Prints "N of 1024 lanes equal" and the first mismatch, then times the verifier per 32-lane unit.
|
||||
// Build: gcc -O2 -o verify_ref verify_ref.c Run: taskset -c 2 ./verify_ref dump.txt <R> [--time]
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <time.h>
|
||||
#define IGNEUM_NO_CUDA
|
||||
#include "program.h"
|
||||
#include "vectors.h"
|
||||
#include "memhard.h"
|
||||
#include "mm8_block.h"
|
||||
|
||||
static const uint16_t MM8_TABLE[IGNEUM_MM8_R_MAX] = IGNEUM_MM8_TABLE_INIT;
|
||||
|
||||
static uint32_t splitmix32(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }
|
||||
static uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
static uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
static uint32_t umulhi(uint32_t a, uint32_t b) { return (uint32_t)(((uint64_t)a * (uint64_t)b) >> 32); }
|
||||
static double nowMs(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1e3 + ts.tv_nsec / 1e6; }
|
||||
|
||||
#define L for (int l = 0; l < 32; ++l)
|
||||
|
||||
// One mm8 step on the 32-lane register file (spec layout, identical to family-probe.cu's mm8 warp_ref):
|
||||
// r[c] += C[l >> 2][2 * (l & 3)] (d0), r[c2] += C[l >> 2][2 * (l & 3) + 1] (d1).
|
||||
static void mm8_step(uint32_t R[8][32], int a, int b, int c, int c2) {
|
||||
uint32_t A[32], B[32];
|
||||
memcpy(A, R[a], sizeof A);
|
||||
memcpy(B, R[b], sizeof B);
|
||||
L {
|
||||
int row = l >> 2, col0 = 2 * (l & 3);
|
||||
uint32_t acc0 = 0u, acc1 = 0u;
|
||||
for (int k = 0; k < 16; ++k) {
|
||||
uint32_t av = (A[row * 4 + k / 4] >> (8 * (k % 4))) & 0xffu;
|
||||
acc0 += av * ((B[col0 * 4 + k / 4] >> (8 * (k % 4))) & 0xffu);
|
||||
acc1 += av * ((B[(col0 + 1) * 4 + k / 4] >> (8 * (k % 4))) & 0xffu);
|
||||
}
|
||||
R[c][l] += acc0;
|
||||
R[c2][l] += acc1;
|
||||
}
|
||||
}
|
||||
|
||||
static void hash_unit(const uint32_t* cache, uint32_t base, int Rsteps, uint64_t out[32]) {
|
||||
static const uint32_t SEEDW[8] = IGNEUM_SEEDW_INIT;
|
||||
uint32_t R[8][32], T[32], SEL[32];
|
||||
const uint32_t mask = IGNEUM_MASK;
|
||||
L {
|
||||
uint32_t nonce = base + (uint32_t)l;
|
||||
for (int i = 0; i < 8; ++i) {
|
||||
uint32_t x = nonce ^ SEEDW[i]; x += 0x9e3779b9u * (uint32_t)(i + 1); x = splitmix32(x);
|
||||
R[i][l] = x ^ SEEDW[(i + 1) & 7];
|
||||
}
|
||||
}
|
||||
for (int it = 0; it < 8; ++it) {
|
||||
L { SEL[l] = R[0][l]; }
|
||||
#include "ref_program.inc"
|
||||
for (int k = 0; k < Rsteps; ++k) {
|
||||
uint16_t v = MM8_TABLE[k];
|
||||
mm8_step(R, v & 7, (v >> 3) & 7, (v >> 6) & 7, (v >> 9) & 7);
|
||||
}
|
||||
}
|
||||
L {
|
||||
uint32_t lo = R[0][l] ^ rotl_imm(R[1][l], 7u) ^ rotl_imm(R[2][l], 14u) ^ rotl_imm(R[3][l], 21u);
|
||||
uint32_t hi = R[4][l] ^ rotl_imm(R[5][l], 9u) ^ rotl_imm(R[6][l], 18u) ^ rotl_imm(R[7][l], 27u);
|
||||
out[l] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
}
|
||||
|
||||
static uint64_t fnv1a64(const void* p, size_t n) {
|
||||
const uint8_t* b = (const uint8_t*)p; uint64_t h = 0xcbf29ce484222325ull;
|
||||
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
|
||||
return h;
|
||||
}
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
if (argc < 3) { printf("usage: verify_ref <dump file> <R> [--time] [--self-test]\n"); return 2; }
|
||||
const char* path = argv[1];
|
||||
int Rsteps = atoi(argv[2]);
|
||||
int doTime = 0, selfTest = 0;
|
||||
for (int i = 3; i < argc; ++i) { if (!strcmp(argv[i], "--time")) doTime = 1; if (!strcmp(argv[i], "--self-test")) selfTest = 1; }
|
||||
if (Rsteps < 0 || Rsteps > IGNEUM_MM8_R_MAX) { printf("R out of range\n"); return 2; }
|
||||
|
||||
size_t words = (size_t)1u << IGNEUM_CACHE_LOG2_WORDS;
|
||||
uint32_t* cache = (uint32_t*)malloc(words * 4u);
|
||||
if (!cache) { printf("no memory for the cache\n"); return 2; }
|
||||
double c0 = nowMs();
|
||||
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(cache, seg);
|
||||
double c1 = nowMs();
|
||||
uint64_t fnv = fnv1a64(cache, words * 4u);
|
||||
printf("cache: host fill %.1f ms, FNV-1a 64 %016llx vs Mac %016llx %s\n", c1 - c0, (unsigned long long)fnv,
|
||||
(unsigned long long)IGNEUM_CACHE_FNV64, fnv == IGNEUM_CACHE_FNV64 ? "PASS" : "FAIL");
|
||||
if (fnv != IGNEUM_CACHE_FNV64) return 1;
|
||||
|
||||
if (selfTest) { // R = 0 interpreter against the pack's vectors
|
||||
int ok = 1;
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
uint64_t out[32]; hash_unit(cache, IGNEUM_VEC_BASE[w], 0, out);
|
||||
int bad = 0; L { if (out[l] != IGNEUM_VEC_OUT[w][l]) ++bad; }
|
||||
printf("self-test R=0 vector base %u: %s (%d lanes differ)\n", IGNEUM_VEC_BASE[w], bad ? "FAIL" : "PASS", bad);
|
||||
ok = ok && !bad;
|
||||
}
|
||||
if (!ok) return 1;
|
||||
}
|
||||
|
||||
// Read the dump: 32 consecutive lines per warp, same base
|
||||
FILE* f = fopen(path, "r");
|
||||
if (!f) { printf("cannot open %s\n", path); return 2; }
|
||||
uint32_t* bases = NULL; uint64_t* vals = NULL; size_t n = 0, cap = 0;
|
||||
unsigned base; int lane; unsigned long long v;
|
||||
while (fscanf(f, "%u %d %llx", &base, &lane, &v) == 3) {
|
||||
if (n == cap) { cap = cap ? cap * 2 : 1024; bases = realloc(bases, cap * 4); vals = realloc(vals, cap * 8); }
|
||||
bases[n] = base; vals[n] = v; ++n;
|
||||
}
|
||||
fclose(f);
|
||||
if (n % 32 != 0) { printf("dump has %zu lines, not a multiple of 32\n", n); return 2; }
|
||||
size_t units = n / 32;
|
||||
printf("dump: %zu lines, %zu units, R = %d (%d mm8 per hash)\n", n, units, Rsteps, 8 * Rsteps);
|
||||
|
||||
size_t equal = 0; int firstPrinted = 0;
|
||||
double v0 = nowMs();
|
||||
for (size_t u = 0; u < units; ++u) {
|
||||
uint64_t out[32]; hash_unit(cache, bases[u * 32], Rsteps, out);
|
||||
L {
|
||||
if (out[l] == vals[u * 32 + l]) ++equal;
|
||||
else if (!firstPrinted) { firstPrinted = 1; printf("first mismatch: base %u lane %d: cpu %016llx gpu %016llx\n", bases[u * 32], l, (unsigned long long)out[l], (unsigned long long)vals[u * 32 + l]); }
|
||||
}
|
||||
}
|
||||
double v1 = nowMs();
|
||||
printf("%zu of %zu lanes equal (%s) [%.3f ms per unit during the check]\n", equal, n, equal == n ? "PASS" : "FAIL", (v1 - v0) / units);
|
||||
|
||||
if (doTime) {
|
||||
int Rs[2] = { 0, Rsteps }; double ms[2] = { 0, 0 };
|
||||
for (int t = 0; t < 2; ++t) {
|
||||
uint64_t out[32]; volatile uint64_t sink = 0;
|
||||
hash_unit(cache, bases[0], Rs[t], out); // warm
|
||||
double t0 = nowMs();
|
||||
for (size_t u = 0; u < units; ++u) { hash_unit(cache, bases[u * 32], Rs[t], out); sink ^= out[0]; }
|
||||
double t1 = nowMs();
|
||||
ms[t] = (t1 - t0) / units;
|
||||
(void)sink;
|
||||
}
|
||||
printf("verifier: %.3f ms per unit at R=0, %.3f ms per unit at R=%d, mm8 block delta %.3f ms per unit (%.2f us per mm8 step), averaged over %zu units, one core\n",
|
||||
ms[0], ms[1], Rsteps, ms[1] - ms[0], Rsteps ? (ms[1] - ms[0]) * 1000.0 / (8.0 * Rsteps) : 0.0, units);
|
||||
printf("VERIFY R=%d equal=%zu of=%zu ms_r0=%.3f ms_r=%.3f delta=%.3f\n", Rsteps, equal, n, ms[0], ms[1], ms[1] - ms[0]);
|
||||
}
|
||||
free(cache); free(bases); free(vals);
|
||||
return equal == n ? 0 : 1;
|
||||
}
|
||||
219
proto-newpow/state-dataset/README.md
Normal file
|
|
@ -0,0 +1,219 @@
|
|||
# state-dataset: the dataset commits to chain state (class "sd1")
|
||||
|
||||
Horizon lane 8, new proof of work. Prototype and measurements, 6 October 2026, 19:34 to 19:50 UTC.
|
||||
Everything here is a TEST HARNESS: no pool, no network, no wallet, nothing touches the devnet.
|
||||
|
||||
## The scheme in one paragraph
|
||||
|
||||
The hash kernel is unchanged. What changes is the daily dataset build. Today item t of the 1 GiB dataset is
|
||||
`mh_item(cache, t)`: the 16-word state starts as `s[0..7] = K` (day key) and `s[8..15] = t * MUL[i] + RC[i]`, then
|
||||
8 rounds of (8 mixers + one dependent 64-byte cache read XORed in), then 8 final mixers. Under sd1 a 64-byte STATE
|
||||
LEAF for item t is XORed into those 16 initial words before the first mixer: `s[i] ^= leaf(t)[i]`. In the real design
|
||||
leaf(t) is the t-th 64-byte leaf of a canonical serialisation of the chain's execution state at a certified checkpoint
|
||||
20 minutes before the day boundary, zero padded where the state is shorter than the dataset. A miner therefore cannot
|
||||
build the day's dataset without the state, and a verifier that holds no dataset needs leaf(t) for every item it
|
||||
derives. The prototype measures what that costs (build time, device memory, verifier time) and buys (which leaves a
|
||||
hash touches, and what a light client would have to be shown).
|
||||
|
||||
### Leaf stand-in (synthetic, prototype only)
|
||||
|
||||
`leaf(t) = mh_chacha_block(x)` (the pack's own ChaCha12 block from memhard.h) with
|
||||
`x = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574, S[0..7], t, 0, 0x49676e65, 0x53746174)` and
|
||||
`S[i] = K[i] ^ 0x5a5a5a5a` (stand-in state root). For the pack's day key
|
||||
`S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102`. The leaf array is 2^24 items x 64 B = 1 GiB
|
||||
for the pack's 1 GiB dataset (2^28 words = 2^24 items). Definition in `sd.h`.
|
||||
|
||||
## Files
|
||||
|
||||
| file | what |
|
||||
|---|---|
|
||||
| `kernel.cu`, `memhard.h`, `program.h`, `vectors.h` | the mx8-genesis pack, copied unmodified from `proto-cuda/packs-ca2-mixer/mx8-genesis/` |
|
||||
| `sd.h` | `mh_leaf`, `mh_item_sd`, `mh_word_sd` (host and device, plain C compatible) |
|
||||
| `kernel_sd.cu` | the pack's kernel.cu plus `igneum_leaves`, `igneum_build_sd` and their launch wrappers appended; `igneum_hash` byte for byte the pack's (checked with diff) |
|
||||
| `bench.cu` | the harness: `--mode control|sd1`, `--dump <file> <nwarps>`, `--items-out <file>`, `--chunk-mib M`, `--power-seconds S` |
|
||||
| `verify_sd.c` | plain C: host cache fill, item bit-exact check from the GPU's items file, 32-lane register-major interpreter of the program against the GPU's dump, lane-0 load trace, per-unit rows |
|
||||
| `cpu_rows.c` | plain C: the igneum-build-1 rows B.1 and B.2 (OpenMP for the 32-thread row) |
|
||||
| `run_gpu.sh`, `run_cpu.sh` | the exact commands, as run |
|
||||
| `results/gpu/`, `results/cpu/` | every log, the dumps and the items files, copied back from the boxes |
|
||||
|
||||
## Boxes
|
||||
|
||||
- GPU box 2: NVIDIA GeForce RTX 4090, 24 GB (24083 MiB reported), 128 SMs, driver 595.91.07 (CUDA driver 13.2), nvcc 12.8.93,
|
||||
`-arch=sm_89`, power limit 450 W, Ubuntu 24.04.1. Host CPU AMD EPYC 7352 (48 threads). A CPU-only pool daemon shares the
|
||||
host (ports 4463/4480); the plain C rows there were pinned to core 2. Working directory `/root/horizon-newpow/state-dataset`.
|
||||
- CPU box igneum-build-1: AMD EPYC 9454P (48 cores, 96 threads), 128 GB, gcc 13.3.0, L3 256 MiB. Shared with other agents'
|
||||
builds: load average 19 to 32 during the readings (printed in each log). Rows pinned to core 4 under `nice -n 19`; the
|
||||
32-thread row on cores 4 to 35. Working directory `/srv/builds/horizon-newpow/state-dataset`.
|
||||
|
||||
## Commands
|
||||
|
||||
From the Mac (zsh; the GPU box has no rsync, tar over ssh instead):
|
||||
|
||||
```
|
||||
cd /Users/joshm/Projects/igneum-wt-horizon/proto-newpow/state-dataset
|
||||
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.cu *.h *.c *.sh | ssh -i ~/.ssh/igneum-fleet -p <box-2-port> root@<box-2-ip> 'cd /root/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_gpu.sh > run_gpu.log 2>&1'
|
||||
COPYFILE_DISABLE=1 tar czf - --no-xattrs *.c *.h *.sh | ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 'cd /srv/builds/horizon-newpow/state-dataset && tar xzf - && touch * && bash run_cpu.sh > run_cpu.log 2>&1'
|
||||
```
|
||||
|
||||
On GPU box 2 (`run_gpu.sh`):
|
||||
|
||||
```
|
||||
export PATH=/usr/local/cuda/bin:$PATH
|
||||
nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
|
||||
gcc -O2 -std=c11 -o verify_sd verify_sd.c
|
||||
./bench --mode control --batches 10 --items-out items_control.txt --dump dump_control.txt 4 --power-seconds 20
|
||||
./bench --mode sd1 --batches 10 --items-out items_sd1.txt --dump dump_sd1.txt 4 --power-seconds 20
|
||||
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 64
|
||||
./bench --mode sd1 --batches 10 --power-seconds 0 --chunk-mib 256
|
||||
./bench --mode control --batches 10 --power-seconds 20 # second pass
|
||||
./bench --mode sd1 --batches 10 --power-seconds 20 # second pass
|
||||
taskset -c 2 ./verify_sd --mode control --items items_control.txt --dump dump_control.txt
|
||||
taskset -c 2 ./verify_sd --mode sd1 --items items_sd1.txt --dump dump_sd1.txt
|
||||
```
|
||||
|
||||
On igneum-build-1 (`run_cpu.sh`):
|
||||
|
||||
```
|
||||
gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
|
||||
gcc -O2 -std=c11 -o verify_sd verify_sd.c
|
||||
taskset -c 4 nice -n 19 ./cpu_rows --rows # B.1 (i)(ii)(iii) and the one-core B.2 rows, twice
|
||||
taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 # B.2 leaf array on 32 threads, twice
|
||||
taskset -c 4 nice -n 19 ./verify_sd --mode sd1 # the per-unit rows of verify_sd.c on this CPU
|
||||
```
|
||||
|
||||
Timing method: cache fill, leaf array and dataset build are CUDA-event times of the second of two launches (as
|
||||
host.cu does). MH/s is CUDA-event time over 10 batches of 2^24 after a warm-up batch at base 0. Watts and SM MHz are
|
||||
the mean of `nvidia-smi -l 1` samples taken after the first 10 s of a 20 s window in which the hash kernel runs back
|
||||
to back (10 samples used of 22). Fingerprint = FNV-1a 64 over the 2^24 little-endian u64 outputs at base nonce 0,
|
||||
the project's definition.
|
||||
|
||||
## RESULTS
|
||||
|
||||
### A. GPU, RTX 4090 (control vs sd1, two passes each)
|
||||
|
||||
| row | control | sd1 | note |
|
||||
|---|---|---|---|
|
||||
| cache fill, GPU, second pass (ms) | 1.81, 1.85 | 1.85, 1.85 | 256 MiB, unchanged kernel |
|
||||
| leaf array, GPU, second pass (ms) | none | 5.49, 5.49 | 2^24 ChaCha12 blocks, 1 GiB written at 195 GB/s |
|
||||
| dataset build, second pass (ms) | 30.55, 30.55 | 31.98, 31.98 | +1.43 ms (+4.7%): one coalesced 64 B read per item |
|
||||
| build + leaves (ms) | 30.55 | 37.47 | synthetic leaves; a real snapshot arrives from the node instead (see chunked rows) |
|
||||
| the 3 Mac vectors, standalone and in batch | PASS, PASS | n/a (new values) | sd1 base 0 lane 0 = b600edbed969becc, base 4096 = a533e89c78bb6b74, base 1000000 = beb4cb0c6bab8163 |
|
||||
| dataset head, word [MASK], 64 Mac samples | PASS | n/a (new values) | sd1 dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf |
|
||||
| 64 random words vs host derivation | PASS (mh_word) | PASS (mh_word_sd) | |
|
||||
| 2^24 fingerprint at base 0 | 7c28cfb06c5c65a9 (= the pack's, both passes) | d5b0c16390cad0e8 (new, four runs) | |
|
||||
| hash rate, 10 batches of 2^24 (MH/s) | 63.083, 63.083 | 63.088, 63.087 | equal within 0.01%; the hash kernel is the same binary over a dataset of the same shape |
|
||||
| hash rate in the 20 s power window (MH/s) | 63.078, 63.079 | 63.083, 63.083 | |
|
||||
| watts, mean after 10 s | 207.6, 205.2 | 207.9, 206.5 | run to run noise 2 W |
|
||||
| SM MHz, mem MHz | 2745, 10251 | 2745, 10251 | |
|
||||
| MH/s per W | 0.304, 0.307 | 0.303, 0.305 | |
|
||||
| hash kernel registers | 29 | 29 | 24 resident blocks/SM at 1 warp/block |
|
||||
|
||||
Bit-exactness of the build (A.3): 1,024 random item indices (SplitMix64 seeded 0x5d1, `t = z & (2^24 - 1)`), the 16 GPU
|
||||
words of each item against the host derivation in plain C on the box's CPU with the host-filled cache (65536
|
||||
`mh_cache_segment` calls, 380 ms in the bench, 573 to 587 ms pinned to core 2 in verify_sd):
|
||||
|
||||
| check | control | sd1 |
|
||||
|---|---|---|
|
||||
| items equal, in-process (bench.cu, host mh_item / mh_item_sd) | 1024 of 1024 | 1024 of 1024 |
|
||||
| items equal, plain C (verify_sd.c from the items file) | 1024 of 1024 | 1024 of 1024 |
|
||||
| items equal after the chunked rebuild (64 MiB chunks; 256 MiB chunks) | n/a | 1024 of 1024; 1024 of 1024 |
|
||||
| 64 device leaves vs host mh_leaf (incl. t = 0 and 2^24 - 1) | n/a | PASS |
|
||||
| host cache FNV-1a 64 | 48c4f5bf24166b2e = Mac | 48c4f5bf24166b2e = Mac |
|
||||
|
||||
GPU hash outputs vs the C interpreter (4 warps dumped, bases 0, 32, 64, 96; loads through mh_word / mh_word_sd):
|
||||
|
||||
| | control | sd1 |
|
||||
|---|---|---|
|
||||
| lanes equal | 128 of 128 | 128 of 128 |
|
||||
| interpreter vs the Mac vector at base 0 | 32 of 32 | n/a |
|
||||
| interpreting time (4,096 item derivations per warp) | 37 ms | 42 ms |
|
||||
|
||||
Device memory (A.4), cudaMemGetInfo, MiB used (context 395 included):
|
||||
|
||||
| phase | control | sd1 resident leaves | sd1 chunked 64 MiB | sd1 chunked 256 MiB |
|
||||
|---|---|---|---|---|
|
||||
| during the build | 1675 | 2699 | 1741 | 1933 |
|
||||
| while hashing (leaves freed) | 1803 | 1803 | 1803 | 1803 |
|
||||
| chunked rebuild total, copies + builds, one stream (ms) | | | 75.80 | 75.47 |
|
||||
|
||||
Whole-array host to device copy of the 1 GiB leaf array from pinned memory: 62 ms = 17.2 GB/s (this box's PCIe link
|
||||
under load; device to host 55 ms). So the chunked build is PCIe-bound: 62 ms of copy plus the 32 ms build, partly
|
||||
serialised on one stream, gives 76 ms. Two chunk buffers on two streams would hide most of the build under the copy
|
||||
(not done; the kernel already takes an item range `[t0, t0 + n)` with the chunk's leaves at `leaves - 16 t0`, so
|
||||
chunking needed no kernel change at all, only the loop in bench.cu).
|
||||
|
||||
### B. CPU, igneum-build-1 (EPYC 9454P, one core pinned, nice 19), ms per unit of 4,096 items, 100 units
|
||||
|
||||
Two readings, because the box is shared: reading 1 at load 19, reading 2 at load 31.
|
||||
|
||||
| row | reading 1 | reading 2 | per item |
|
||||
|---|---|---|---|
|
||||
| (i) 4,096 random 64 B reads from a 2 GiB resident leaf array | 0.108 ms (min 0.092, max 0.172) | 0.163 ms | 26 to 40 ns |
|
||||
| (i) the same from an 8 GiB array (the year-12 size) | 0.139 ms (min 0.127, max 0.183) | 0.209 ms | 34 to 51 ns |
|
||||
| (ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array) | 0.284 ms | 0.285 ms | 69 ns |
|
||||
| (iii) 4,096 x mh_item on the host cache, naive, no interleaving | 9.01 ms (min 8.91, max 9.60) | 11.24 ms (min 9.38, max 13.51) | 2.2 to 2.7 us |
|
||||
| 4,096 x (leaf + mh_item_sd), naive | 9.31 ms | 11.64 ms | |
|
||||
|
||||
The same rows on GPU box 2's EPYC 7352, core 2 (verify_sd.c): (ii) 0.369 to 0.374 ms, (iii) 10.14 to 10.21 ms, leaf +
|
||||
mh_item_sd 10.56 to 10.58 ms.
|
||||
|
||||
The project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent
|
||||
cache misses across items; the naive (iii) figure here is an upper bound. The sd1 increment is the row that matters:
|
||||
0.11 to 0.21 ms per unit when the verifier holds the 2 to 8 GiB state in RAM, or 0.28 ms when it derives the stand-in
|
||||
leaf (in the real design it cannot derive a state leaf, it must hold the state or be shown it). Against the 2.06 ms
|
||||
interleaved verifier that is +5% to +14%; against the naive 9 to 11 ms it is +1% to +3%.
|
||||
|
||||
B.2, the daily snapshot build on CPU (igneum-build-1):
|
||||
|
||||
| row | time |
|
||||
|---|---|
|
||||
| 1 GiB leaf array, 2^24 ChaCha12 blocks, one core | 1.670 s, 1.677 s (100 ns per leaf) |
|
||||
| the same on 32 threads (OpenMP), second pass (first pass 0.172 s with page faults) | 0.073 s, 0.083 s |
|
||||
| host cache fill, 256 MiB, one core | 0.447 s, 0.452 s (0.380 s on the 4090 box's host, unpinned; 0.57 s pinned beside the pool daemon) |
|
||||
|
||||
B.3, the state sample a light client would check (lane 0 of warp base 0, from the instrumented interpreter):
|
||||
|
||||
| | control | sd1 |
|
||||
|---|---|---|
|
||||
| distinct words touched by the 128 loads | 128 | 128 |
|
||||
| distinct items touched | 128 of 2^24 | 128 of 2^24 |
|
||||
| first 8 load words | 8528400 43034175 255822847 136709537 ... | 8528400 85793927 255822847 180392080 ... |
|
||||
|
||||
Loads 1 and 3 share their address in both modes because their addresses are fixed by the nonce before any load's
|
||||
value feeds back; from load 2 on the sd1 trail diverges. Merkle arithmetic for a 2 GiB state (2^25 leaves of 64 B,
|
||||
depth 25, 32-byte hashes): 128 openings x 25 x 32 B = 102,400 B of siblings + 128 x 64 B = 8,192 B of leaves =
|
||||
110,592 B (108 KiB) per lane, uncompressed; a 32-lane unit is 32 x 108 KiB = 3.4 MiB, before deduplicating shared
|
||||
upper levels (128 random paths in a depth-25 tree share only their top seven levels, so dedup saves about 20% per lane,
|
||||
approximate; across the 32 lanes of a unit the same seven levels are shared again). Every one of the 128 loads is a distinct item, so there is no saving from repeated items.
|
||||
|
||||
## What the numbers mean, per tier (the consequences rule)
|
||||
|
||||
- Hash rate, watts, MH/s per W: unchanged for every tier on every vendor, because the hash kernel is the pack's and the
|
||||
dataset has the same shape. Nothing to do.
|
||||
- Daily build: +1.4 ms on the 4090 (30.6 to 32.0 ms) when the leaves are resident, 76 ms when streamed from the host
|
||||
over PCIe in chunks. At the designed 2 GiB dataset the leaf array is 2 GiB and the streamed build is about 150 ms
|
||||
(approximate, scaling the 17 GB/s copy); on a PCIe 3 x8 slot in a rig, about 4x that (approximate). All far inside the
|
||||
daily window on every tier.
|
||||
- Device memory: with resident leaves the build peak at 2 GiB is 2 GiB + 256 MiB + 2 GiB + context, about 4.6 GiB
|
||||
(approximate), which fits an 8 GB card but not beside a prover. Streamed in 64 MiB chunks the peak is dataset + cache
|
||||
+ 64 MiB + context, measured 1741 MiB at 1 GiB here, about 2.7 GiB at 2 GiB (approximate): no tier loses memory it has
|
||||
today. Recommendation: ship chunked only; never hold the leaf array on the device.
|
||||
- The real cost is delivery: every miner needs the 1 to 2 GiB state serialisation once a day. A miner
|
||||
beside its own node reads it from disk or loopback (seconds). A pool user without a node must get it from the pool
|
||||
(2 GiB per day per miner, or a shared download), and a light verifier either holds the state (0.11 to 0.21 ms per unit
|
||||
extra, measured) or is shown 3.4 MiB of Merkle openings per unit (arithmetic above), which is not a light client any
|
||||
more. The design decision this prototype leaves open is which of those two the protocol asks of a header verifier.
|
||||
- CPU snapshot build: 1.7 s on one core, 0.07 s on 32, for the synthetic leaves; the real serialisation is bounded by the
|
||||
node's state read, not by this.
|
||||
|
||||
## What failed, what was cut
|
||||
|
||||
- Nothing on the list was cut. verify_sd.c's lane-0 trace printed nothing on the first run (the trace counter was reset
|
||||
on every warp, so it read 0 after warp 3); fixed, the verifiers re-run, the fix is in the file.
|
||||
- The GPU box has no rsync; the sources went over with tar through ssh (`COPYFILE_DISABLE=1 --no-xattrs`, the Mac's
|
||||
tar otherwise writes Apple xattr headers that GNU tar warns about).
|
||||
- `-std=c11` hides `clock_gettime`; both C files define `_POSIX_C_SOURCE 200809L`.
|
||||
- igneum-build-1 was under other agents' load (19 to 32) for every reading; both readings are given. The 4090 box's
|
||||
CPU rows ran beside the pool daemon, pinned to core 2. The GPU itself was idle apart from this bench.
|
||||
- Not measured: a two-stream chunked build (copy and build overlapped), AMD and Apple builds, the 2 GiB dataset
|
||||
itself (the pack is 1 GiB; the 2 GiB figures above are labelled approximate).
|
||||
464
proto-newpow/state-dataset/bench.cu
Normal file
|
|
@ -0,0 +1,464 @@
|
|||
// state-dataset prototype bench (Horizon lane 8, class sd1). TEST HARNESS ONLY: no pool, no network, no wallet.
|
||||
//
|
||||
// Shape follows proto-cuda/host.cu: cache fill (GPU, twice, CUDA events), host cache fill and check, dataset build
|
||||
// (twice, second pass reported), dataset self-test, the 3 Mac vectors, a warm-up batch of 2^24 at base nonce 0
|
||||
// (fingerprinted: FNV-1a 64 over the 2^24 little-endian u64 outputs), N timed batches (CUDA events), then a power
|
||||
// window where nvidia-smi samples at 1 Hz while the hash kernel runs back to back.
|
||||
//
|
||||
// --mode control the unmodified pack (igneum_build)
|
||||
// --mode sd1 leaf array (igneum_leaves) + igneum_build_sd; the hash kernel is the pack's
|
||||
// --dump <file> <n> write the outputs of the first n warps of the base-0 batch (for verify_sd.c)
|
||||
// --items-out <file> write the 16 words of 1,024 random items (SplitMix64 seeded 0x5d1) read back from the GPU
|
||||
// --batches N timed batches after the warm-up (default 10)
|
||||
// --batch-log2 B nonces per batch (default 24)
|
||||
// --power-seconds S length of the nvidia-smi window (default 20; 0 skips it)
|
||||
// --chunk-mib M sd1 only: after the resident build, rebuild with the leaf array streamed from pinned host
|
||||
// memory in M MiB chunks (the shape a miner uses when the leaves do not fit beside the dataset)
|
||||
// --device D
|
||||
//
|
||||
// Build: nvcc -O3 -std=c++17 -arch=sm_89 -Xcompiler -pthread -o bench bench.cu kernel_sd.cu
|
||||
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <chrono>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
#include <thread>
|
||||
#include <mutex>
|
||||
|
||||
#include "program.h"
|
||||
#include "vectors.h"
|
||||
#include "memhard.h"
|
||||
#include "sd.h"
|
||||
|
||||
cudaError_t igneum_launch_leaves(uint32_t* leaves, uint32_t nItems);
|
||||
cudaError_t igneum_launch_build_sd(uint32_t* ds, const uint32_t* cache, const uint32_t* leaves, uint32_t t0, uint32_t nItems);
|
||||
|
||||
#define CUDA_CHECK(call) do { cudaError_t err_ = (call); if (err_ != cudaSuccess) { \
|
||||
std::fprintf(stderr, "CUDA error: %s (%d)\n at %s:%d\n in %s\n", cudaGetErrorString(err_), (int)err_, __FILE__, __LINE__, #call); \
|
||||
std::exit(2); } } while (0)
|
||||
|
||||
static double wallMs() {
|
||||
using namespace std::chrono;
|
||||
return duration<double, std::milli>(steady_clock::now().time_since_epoch()).count();
|
||||
}
|
||||
static uint64_t fnv1a64(const void* p, size_t n) {
|
||||
const uint8_t* b = (const uint8_t*)p;
|
||||
uint64_t h = 0xcbf29ce484222325ull;
|
||||
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
|
||||
return h;
|
||||
}
|
||||
static uint64_t splitmix64(uint64_t& s) {
|
||||
s += 0x9E3779B97F4A7C15ull;
|
||||
uint64_t z = s;
|
||||
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull;
|
||||
z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
|
||||
return z ^ (z >> 31);
|
||||
}
|
||||
static float eventMs(cudaEvent_t a, cudaEvent_t b) { float ms = 0.f; CUDA_CHECK(cudaEventElapsedTime(&ms, a, b)); return ms; }
|
||||
|
||||
struct Opts {
|
||||
bool sd1 = false;
|
||||
int batches = 10, batchLog2 = 24, device = 0, powerSeconds = 20, dumpWarps = 0, chunkMib = 0;
|
||||
std::string dump, itemsOut;
|
||||
};
|
||||
|
||||
static void usage() {
|
||||
std::printf("bench --mode control|sd1 [--dump <file> <nwarps>] [--items-out <file>] [--batches 10] [--batch-log2 24] [--power-seconds 20] [--chunk-mib M] [--device 0]\n");
|
||||
}
|
||||
|
||||
static Opts parse(int argc, char** argv) {
|
||||
Opts o;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
std::string a = argv[i];
|
||||
auto need = [&](int k) { if (i + k >= argc) { usage(); std::exit(2); } };
|
||||
if (a == "--mode") { need(1); std::string m = argv[++i]; if (m == "sd1") o.sd1 = true; else if (m == "control") o.sd1 = false; else { usage(); std::exit(2); } }
|
||||
else if (a == "--dump") { need(2); o.dump = argv[++i]; o.dumpWarps = std::atoi(argv[++i]); }
|
||||
else if (a == "--items-out") { need(1); o.itemsOut = argv[++i]; }
|
||||
else if (a == "--batches") { need(1); o.batches = std::atoi(argv[++i]); }
|
||||
else if (a == "--batch-log2") { need(1); o.batchLog2 = std::atoi(argv[++i]); }
|
||||
else if (a == "--power-seconds") { need(1); o.powerSeconds = std::atoi(argv[++i]); }
|
||||
else if (a == "--device") { need(1); o.device = std::atoi(argv[++i]); }
|
||||
else if (a == "--chunk-mib") { need(1); o.chunkMib = std::atoi(argv[++i]); }
|
||||
else if (a == "-h" || a == "--help") { usage(); std::exit(0); }
|
||||
else { std::printf("unknown argument %s\n", argv[i]); usage(); std::exit(2); }
|
||||
}
|
||||
if (o.batchLog2 < 10 || o.batchLog2 > 28 || o.batches < 1) { usage(); std::exit(2); }
|
||||
return o;
|
||||
}
|
||||
|
||||
// nvidia-smi sampler: one line per second, read on its own thread until `timeout` ends the process.
|
||||
struct Sample { double t; double watts; double smMHz; double memMHz; };
|
||||
struct Sampler {
|
||||
std::vector<Sample> samples;
|
||||
std::mutex mu;
|
||||
std::thread th;
|
||||
double t0 = 0;
|
||||
void start(int device, int seconds) {
|
||||
t0 = wallMs();
|
||||
char cmd[512];
|
||||
std::snprintf(cmd, sizeof(cmd), "timeout %d nvidia-smi -i %d --query-gpu=power.draw,clocks.sm,clocks.mem --format=csv,noheader,nounits -l 1 2>/dev/null", seconds + 2, device);
|
||||
std::string c = cmd;
|
||||
th = std::thread([this, c]() {
|
||||
FILE* f = popen(c.c_str(), "r");
|
||||
if (!f) return;
|
||||
char line[256];
|
||||
while (std::fgets(line, sizeof(line), f)) {
|
||||
double w = 0, sm = 0, mem = 0;
|
||||
if (std::sscanf(line, "%lf , %lf , %lf", &w, &sm, &mem) == 3) {
|
||||
std::lock_guard<std::mutex> g(mu);
|
||||
samples.push_back({ wallMs() - t0, w, sm, mem });
|
||||
}
|
||||
}
|
||||
pclose(f);
|
||||
});
|
||||
}
|
||||
void join() { if (th.joinable()) th.join(); }
|
||||
};
|
||||
|
||||
static const uint32_t KEYW[8] = IGNEUM_KEY_INIT;
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
Opts o = parse(argc, argv);
|
||||
const char* modeName = o.sd1 ? "sd1" : "control";
|
||||
std::printf("state-dataset bench mode %s pack \"%s\" (test harness: no pool, no network, no wallet)\n", modeName, IGNEUM_SEED_STRING);
|
||||
|
||||
int count = 0;
|
||||
CUDA_CHECK(cudaGetDeviceCount(&count));
|
||||
if (count == 0) { std::printf("FAIL: no CUDA device\n"); return 2; }
|
||||
CUDA_CHECK(cudaSetDevice(o.device));
|
||||
cudaDeviceProp prop;
|
||||
std::memset(&prop, 0, sizeof(prop));
|
||||
CUDA_CHECK(cudaGetDeviceProperties(&prop, o.device));
|
||||
int drv = 0, rt = 0;
|
||||
CUDA_CHECK(cudaDriverGetVersion(&drv));
|
||||
CUDA_CHECK(cudaRuntimeGetVersion(&rt));
|
||||
std::printf("GPU: %s (%d SMs, cc %d.%d, %.0f MiB), CUDA driver %d.%d runtime %d.%d\n", prop.name, prop.multiProcessorCount, prop.major, prop.minor,
|
||||
(double)prop.totalGlobalMem / 1048576.0, drv / 1000, (drv % 100) / 10, rt / 1000, (rt % 100) / 10);
|
||||
int regs = 0, bps = 0;
|
||||
CUDA_CHECK(igneum_hash_info(®s, &bps, 1u));
|
||||
std::printf("hash kernel: %d registers/thread, %d resident blocks/SM at 1 warp/block\n", regs, bps);
|
||||
if (o.sd1) {
|
||||
std::printf("sd1 leaf stand-in: S[i] = K[i] ^ 0x%08x -> S =", SD_STATE_ROOT_XOR);
|
||||
for (int i = 0; i < 8; ++i) std::printf(" %08x", KEYW[i] ^ SD_STATE_ROOT_XOR);
|
||||
std::printf("; leaf(t) = ChaCha12 block of (sigma, S, t, 0, %08x, %08x)\n", SD_TAG0, SD_TAG1);
|
||||
}
|
||||
|
||||
size_t free0 = 0, total = 0;
|
||||
CUDA_CHECK(cudaMemGetInfo(&free0, &total));
|
||||
std::printf("device memory at start: %.0f MiB used of %.0f MiB (context)\n", (double)(total - free0) / 1048576.0, (double)total / 1048576.0);
|
||||
|
||||
cudaEvent_t e0, e1;
|
||||
CUDA_CHECK(cudaEventCreate(&e0));
|
||||
CUDA_CHECK(cudaEventCreate(&e1));
|
||||
|
||||
// ---- cache: GPU fill twice, host fill, check
|
||||
const uint32_t cacheWords = 1u << IGNEUM_CACHE_LOG2_WORDS;
|
||||
const size_t cacheBytes = (size_t)cacheWords * 4u;
|
||||
uint32_t* dCache = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dCache, cacheBytes));
|
||||
double cacheFill[2] = { 0, 0 };
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(igneum_launch_cache_fill(dCache, IGNEUM_CACHE_SEGMENTS));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
cacheFill[pass] = eventMs(e0, e1);
|
||||
}
|
||||
std::printf("cache fill (GPU): %.2f ms first, %.2f ms second (%u segments x 64 ChaCha12 blocks, %u MiB)\n", cacheFill[0], cacheFill[1], (unsigned)IGNEUM_CACHE_SEGMENTS, (unsigned)(cacheBytes >> 20));
|
||||
std::vector<uint32_t> hCache(cacheWords);
|
||||
double h0 = wallMs();
|
||||
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(hCache.data(), seg);
|
||||
double hostCacheMs = wallMs() - h0;
|
||||
bool cachePass = false;
|
||||
{
|
||||
std::vector<uint32_t> dev(cacheWords);
|
||||
CUDA_CHECK(cudaMemcpy(dev.data(), dCache, cacheBytes, cudaMemcpyDeviceToHost));
|
||||
bool same = std::memcmp(dev.data(), hCache.data(), cacheBytes) == 0;
|
||||
uint64_t fnv = fnv1a64(hCache.data(), cacheBytes);
|
||||
cachePass = same && fnv == IGNEUM_CACHE_FNV64;
|
||||
std::printf("cache fill (host, one thread): %.1f ms; cache check: %s (GPU == host %s, host FNV-1a 64 %016llx vs Mac %016llx)\n",
|
||||
hostCacheMs, cachePass ? "PASS" : "FAIL", same ? "PASS" : "FAIL", (unsigned long long)fnv, (unsigned long long)IGNEUM_CACHE_FNV64);
|
||||
}
|
||||
|
||||
// ---- sd1: leaf array
|
||||
const uint32_t words = 1u << IGNEUM_DATASET_LOG2;
|
||||
const uint32_t mask = words - 1u;
|
||||
const uint32_t nItems = words / 16u;
|
||||
uint32_t* dLeaves = nullptr;
|
||||
double leavesMs[2] = { 0, 0 };
|
||||
bool leafPass = true;
|
||||
if (o.sd1) {
|
||||
CUDA_CHECK(cudaMalloc((void**)&dLeaves, (size_t)nItems * 64u));
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(igneum_launch_leaves(dLeaves, nItems));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
leavesMs[pass] = eventMs(e0, e1);
|
||||
}
|
||||
std::printf("leaf array (GPU, %u x 64 B = %u MiB): %.2f ms first, %.2f ms second -> %.1f GB/s written\n", nItems, (unsigned)(((size_t)nItems * 64u) >> 20),
|
||||
leavesMs[0], leavesMs[1], (double)nItems * 64.0 / 1e9 / (leavesMs[1] / 1000.0));
|
||||
// 64 leaves against the host derivation
|
||||
uint64_t s = 0x1ea7ull;
|
||||
int bad = 0;
|
||||
for (int k = 0; k < 64; ++k) {
|
||||
uint32_t t = (k == 0) ? 0u : (k == 1) ? nItems - 1u : (uint32_t)splitmix64(s) & (nItems - 1u);
|
||||
uint32_t got[16], want[16];
|
||||
CUDA_CHECK(cudaMemcpy(got, dLeaves + (size_t)t * 16u, 64, cudaMemcpyDeviceToHost));
|
||||
mh_leaf(t, want);
|
||||
if (std::memcmp(got, want, 64) != 0) { if (bad == 0) std::printf(" leaf[%u] differs: gpu %08x host %08x\n", t, got[0], want[0]); ++bad; }
|
||||
}
|
||||
leafPass = (bad == 0);
|
||||
std::printf("leaf check: %s (64 leaves incl. 0 and %u vs host mh_leaf)\n", leafPass ? "PASS" : "FAIL", nItems - 1u);
|
||||
}
|
||||
|
||||
// ---- dataset build, twice
|
||||
uint32_t* dDs = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dDs, (size_t)words * 4u));
|
||||
double buildMs[2] = { 0, 0 };
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
if (o.sd1) CUDA_CHECK(igneum_launch_build_sd(dDs, dCache, dLeaves, 0u, nItems));
|
||||
else CUDA_CHECK(igneum_launch_build(dDs, dCache, nItems));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
buildMs[pass] = eventMs(e0, e1);
|
||||
}
|
||||
std::printf("dataset build (%s): %.2f ms first, %.2f ms second -> %.1f M items/s\n", o.sd1 ? "igneum_build_sd, leaf XOR before the first mixer" : "igneum_build, the pack's",
|
||||
buildMs[0], buildMs[1], (double)nItems / 1e6 / (buildMs[1] / 1000.0));
|
||||
size_t freeB = 0;
|
||||
CUDA_CHECK(cudaMemGetInfo(&freeB, &total));
|
||||
double usedBuildMiB = (double)(total - freeB) / 1048576.0;
|
||||
std::printf("device memory after the build: %.0f MiB used (context %.0f + cache %u + %sdataset %u MiB)\n", usedBuildMiB, (double)(total - free0) / 1048576.0,
|
||||
(unsigned)(cacheBytes >> 20), o.sd1 ? "leaves 1024 + " : "", (unsigned)(((size_t)words * 4u) >> 20));
|
||||
|
||||
// ---- dataset self-test
|
||||
bool dsPass = true;
|
||||
{
|
||||
uint32_t head[16];
|
||||
CUDA_CHECK(cudaMemcpy(head, dDs, sizeof(head), cudaMemcpyDeviceToHost));
|
||||
std::printf("dataset[0..3] = %08x %08x %08x %08x", head[0], head[1], head[2], head[3]);
|
||||
if (!o.sd1) {
|
||||
bool headOk = std::memcmp(head, IGNEUM_DS_HEAD, 64) == 0;
|
||||
uint32_t last = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&last, dDs + IGNEUM_DS_LAST_INDEX, 4, cudaMemcpyDeviceToHost));
|
||||
bool lastOk = (last == IGNEUM_DS_LAST);
|
||||
int badSample = 0;
|
||||
for (int k = 0; k < IGNEUM_DS_SAMPLES; ++k) {
|
||||
uint32_t v = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&v, dDs + IGNEUM_DS_SAMPLE_INDEX[k], 4, cudaMemcpyDeviceToHost));
|
||||
if (v != IGNEUM_DS_SAMPLE_VALUE[k]) ++badSample;
|
||||
}
|
||||
dsPass = headOk && lastOk && badSample == 0;
|
||||
std::printf(" head 16 vs Mac %s, word [MASK] vs Mac %s, 64 Mac samples %s\n", headOk ? "PASS" : "FAIL", lastOk ? "PASS" : "FAIL", badSample == 0 ? "PASS" : "FAIL");
|
||||
} else {
|
||||
std::printf(" (sd1: new values, no Mac expectation)\n");
|
||||
}
|
||||
// 64 random words vs the host derivation (host.cu's points)
|
||||
int badRnd = 0;
|
||||
uint64_t s = 0x9E3779B97F4A7C15ull ^ (uint64_t)words;
|
||||
for (int k = 0; k < 64; ++k) {
|
||||
uint32_t idx = (uint32_t)splitmix64(s) & mask;
|
||||
uint32_t v = 0;
|
||||
CUDA_CHECK(cudaMemcpy(&v, dDs + idx, 4, cudaMemcpyDeviceToHost));
|
||||
uint32_t want = o.sd1 ? mh_word_sd(hCache.data(), idx) : mh_word(hCache.data(), idx);
|
||||
if (v != want) ++badRnd;
|
||||
}
|
||||
dsPass = dsPass && badRnd == 0;
|
||||
std::printf("dataset self-test: %s (64 random words vs host %s: %s)\n", dsPass ? "PASS" : "FAIL", o.sd1 ? "mh_word_sd" : "mh_word", badRnd == 0 ? "PASS" : "FAIL");
|
||||
}
|
||||
|
||||
// ---- 1,024 random items (SplitMix64 seeded 0x5d1) read back and compared with the host item derivation
|
||||
int itemsEqual = 0;
|
||||
{
|
||||
FILE* f = o.itemsOut.empty() ? nullptr : std::fopen(o.itemsOut.c_str(), "w");
|
||||
if (f) std::fprintf(f, "mode %s\nnitems 1024\n", modeName);
|
||||
uint64_t s = 0x5d1ull;
|
||||
for (int k = 0; k < 1024; ++k) {
|
||||
uint32_t t = (uint32_t)splitmix64(s) & (nItems - 1u);
|
||||
uint32_t got[16], want[16];
|
||||
CUDA_CHECK(cudaMemcpy(got, dDs + (size_t)t * 16u, 64, cudaMemcpyDeviceToHost));
|
||||
if (o.sd1) { uint32_t leaf[16]; mh_leaf(t, leaf); mh_item_sd(hCache.data(), leaf, t, want); }
|
||||
else mh_item(hCache.data(), t, want);
|
||||
if (std::memcmp(got, want, 64) == 0) ++itemsEqual;
|
||||
else if (itemsEqual == k) std::printf(" item %u differs: gpu %08x host %08x (first difference)\n", t, got[0], want[0]);
|
||||
if (f) { std::fprintf(f, "%u", t); for (int i = 0; i < 16; ++i) std::fprintf(f, " %08x", got[i]); std::fprintf(f, "\n"); }
|
||||
}
|
||||
if (f) std::fclose(f);
|
||||
std::printf("item bit-exactness (in-process, host %s on the host cache): %d of 1024 items equal\n", o.sd1 ? "mh_item_sd with host mh_leaf" : "mh_item", itemsEqual);
|
||||
}
|
||||
|
||||
// ---- sd1, chunked: the leaf array lives in pinned host memory (as a real snapshot would, delivered by the node)
|
||||
// and is streamed to the device in chunks; the dataset is rebuilt chunk by chunk and re-checked. Peak device memory
|
||||
// is then cache + dataset + one chunk. One stream, so copy and build serialise; two chunk buffers on two streams
|
||||
// would overlap them (not done here).
|
||||
double chunkedMs = 0, h2dMs = 0, usedChunkMiB = 0;
|
||||
int chunkedEqual = -1;
|
||||
if (o.sd1 && o.chunkMib > 0) {
|
||||
const size_t leafBytes = (size_t)nItems * 64u;
|
||||
const size_t chunkBytes = (size_t)o.chunkMib << 20;
|
||||
const uint32_t chunkItems = (uint32_t)(chunkBytes / 64u);
|
||||
if (chunkBytes > leafBytes || (leafBytes % chunkBytes) != 0) { std::printf("FAIL: --chunk-mib must divide 1024\n"); return 2; }
|
||||
uint32_t* hLeaves = nullptr;
|
||||
CUDA_CHECK(cudaMallocHost((void**)&hLeaves, leafBytes));
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(cudaMemcpy(hLeaves, dLeaves, leafBytes, cudaMemcpyDeviceToHost));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
double d2h = eventMs(e0, e1);
|
||||
CUDA_CHECK(cudaFree(dLeaves)); dLeaves = nullptr;
|
||||
// one-shot H2D of the whole array into the dataset buffer (overwritten by the rebuild anyway): the PCIe rate
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
CUDA_CHECK(cudaMemcpy(dDs, hLeaves, leafBytes, cudaMemcpyHostToDevice));
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
h2dMs = eventMs(e0, e1);
|
||||
uint32_t* dChunk = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dChunk, chunkBytes));
|
||||
CUDA_CHECK(cudaMemGetInfo(&freeB, &total));
|
||||
usedChunkMiB = (double)(total - freeB) / 1048576.0;
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
for (uint32_t t0 = 0; t0 < nItems; t0 += chunkItems) {
|
||||
CUDA_CHECK(cudaMemcpyAsync(dChunk, hLeaves + (size_t)t0 * 16u, chunkBytes, cudaMemcpyHostToDevice, 0));
|
||||
CUDA_CHECK(igneum_launch_build_sd(dDs, dCache, dChunk, t0, chunkItems));
|
||||
}
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
chunkedMs = eventMs(e0, e1);
|
||||
// re-check the same 1,024 items
|
||||
uint64_t s = 0x5d1ull; chunkedEqual = 0;
|
||||
for (int k = 0; k < 1024; ++k) {
|
||||
uint32_t t = (uint32_t)splitmix64(s) & (nItems - 1u);
|
||||
uint32_t got[16], want[16], leaf[16];
|
||||
CUDA_CHECK(cudaMemcpy(got, dDs + (size_t)t * 16u, 64, cudaMemcpyDeviceToHost));
|
||||
mh_leaf(t, leaf); mh_item_sd(hCache.data(), leaf, t, want);
|
||||
if (std::memcmp(got, want, 64) == 0) ++chunkedEqual;
|
||||
}
|
||||
std::printf("chunked build (leaves from pinned host memory in %d MiB chunks, %u chunks, one stream): %.2f ms total (copies + builds); whole-array H2D alone %.2f ms = %.1f GB/s, D2H %.2f ms\n",
|
||||
o.chunkMib, nItems / chunkItems, chunkedMs, h2dMs, (double)leafBytes / 1e9 / (h2dMs / 1000.0), d2h);
|
||||
std::printf("device memory during the chunked build: %.0f MiB used (context + cache 256 + dataset 1024 + chunk %d MiB); items after the chunked rebuild: %d of 1024 equal\n", usedChunkMiB, o.chunkMib, chunkedEqual);
|
||||
CUDA_CHECK(cudaFree(dChunk));
|
||||
CUDA_CHECK(cudaFreeHost(hLeaves));
|
||||
}
|
||||
|
||||
// The leaf array is only needed for the build; a miner frees it before hashing.
|
||||
if (dLeaves) { CUDA_CHECK(cudaFree(dLeaves)); dLeaves = nullptr; }
|
||||
|
||||
// ---- vectors (control: against the Mac; sd1: printed)
|
||||
const uint32_t nonces = 1u << o.batchLog2;
|
||||
uint64_t* dOut = nullptr;
|
||||
CUDA_CHECK(cudaMalloc((void**)&dOut, (size_t)nonces * 8u));
|
||||
bool vecPass = true;
|
||||
{
|
||||
uint64_t got[32];
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, IGNEUM_VEC_BASE[w], mask, 32u, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
CUDA_CHECK(cudaMemcpy(got, dOut, sizeof(got), cudaMemcpyDeviceToHost));
|
||||
if (!o.sd1) {
|
||||
int bad = 0;
|
||||
for (int l = 0; l < 32; ++l) if (got[l] != IGNEUM_VEC_OUT[w][l]) ++bad;
|
||||
vecPass = vecPass && bad == 0;
|
||||
std::printf("vector warp base %u: %s (%d of 32 lanes differ)\n", IGNEUM_VEC_BASE[w], bad == 0 ? "PASS" : "FAIL", bad);
|
||||
} else {
|
||||
std::printf("sd1 warp base %u lane 0: %016llx (new value; checked by verify_sd.c through the dump)\n", IGNEUM_VEC_BASE[w], (unsigned long long)got[0]);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ---- warm-up batch at base 0: fingerprint, in-batch vectors, dump
|
||||
double w0 = wallMs();
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, 0u, mask, nonces, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
double warmMs = wallMs() - w0;
|
||||
std::vector<uint64_t> hOut(nonces);
|
||||
CUDA_CHECK(cudaMemcpy(hOut.data(), dOut, (size_t)nonces * 8u, cudaMemcpyDeviceToHost));
|
||||
uint64_t fp = fnv1a64(hOut.data(), (size_t)nonces * 8u);
|
||||
std::printf("warm-up batch: 2^%d hashes in %.2f ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): %016llx%s\n", o.batchLog2, warmMs,
|
||||
(unsigned long long)fp, o.sd1 ? " (sd1, new value)" : (fp == 0x7c28cfb06c5c65a9ull ? " = 7c28cfb06c5c65a9 (the pack's)" : " DIFFERS from 7c28cfb06c5c65a9"));
|
||||
bool fpPass = o.sd1 || fp == 0x7c28cfb06c5c65a9ull;
|
||||
if (!o.sd1) {
|
||||
for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) {
|
||||
if ((uint64_t)IGNEUM_VEC_BASE[w] + 32ull > nonces) continue;
|
||||
int bad = 0;
|
||||
for (int l = 0; l < 32; ++l) if (hOut[IGNEUM_VEC_BASE[w] + l] != IGNEUM_VEC_OUT[w][l]) ++bad;
|
||||
vecPass = vecPass && bad == 0;
|
||||
std::printf("vector warp base %u in batch: %s\n", IGNEUM_VEC_BASE[w], bad == 0 ? "PASS" : "FAIL");
|
||||
}
|
||||
}
|
||||
if (!o.dump.empty() && o.dumpWarps > 0) {
|
||||
FILE* f = std::fopen(o.dump.c_str(), "w");
|
||||
if (!f) { std::printf("FAIL: cannot write %s\n", o.dump.c_str()); return 2; }
|
||||
std::fprintf(f, "mode %s\nwarps %d\n", modeName, o.dumpWarps);
|
||||
for (int w = 0; w < o.dumpWarps; ++w) {
|
||||
std::fprintf(f, "base %u\n", 32u * (uint32_t)w);
|
||||
for (int l = 0; l < 32; ++l) std::fprintf(f, "%016llx\n", (unsigned long long)hOut[32u * (uint32_t)w + l]);
|
||||
}
|
||||
std::fclose(f);
|
||||
std::printf("dump: %d warps (bases 0, 32, ...) written to %s\n", o.dumpWarps, o.dump.c_str());
|
||||
}
|
||||
|
||||
// ---- timed batches
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
for (int b = 1; b <= o.batches; ++b) {
|
||||
uint32_t base = (uint32_t)((uint64_t)b * (uint64_t)nonces);
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, 1u));
|
||||
}
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
double gpuMs = eventMs(e0, e1);
|
||||
double mhs = (double)nonces * (double)o.batches / (gpuMs / 1000.0) / 1e6;
|
||||
std::printf("timed: %d batches x 2^%d hashes, GPU %.2f ms -> %.3f MH/s (%.2f GB/s useful, loads x 4 B)\n", o.batches, o.batchLog2, gpuMs, mhs, mhs * 1e6 * IGNEUM_LOADS_PER_HASH * 4.0 / 1e9);
|
||||
|
||||
// ---- power window: hash back to back for powerSeconds while nvidia-smi samples at 1 Hz
|
||||
double powerW = 0, smMHz = 0, memMHz = 0, mhsWindow = 0;
|
||||
int nSamples = 0, nUsed = 0;
|
||||
if (o.powerSeconds > 0) {
|
||||
Sampler smp;
|
||||
smp.start(o.device, o.powerSeconds);
|
||||
double start = wallMs();
|
||||
int b = o.batches + 1;
|
||||
int done = 0;
|
||||
CUDA_CHECK(cudaEventRecord(e0));
|
||||
while (wallMs() - start < o.powerSeconds * 1000.0) {
|
||||
uint32_t base = (uint32_t)((uint64_t)b++ * (uint64_t)nonces);
|
||||
CUDA_CHECK(igneum_launch_hash(dDs, dOut, base, mask, nonces, 1u));
|
||||
CUDA_CHECK(cudaDeviceSynchronize());
|
||||
++done;
|
||||
}
|
||||
CUDA_CHECK(cudaEventRecord(e1));
|
||||
CUDA_CHECK(cudaEventSynchronize(e1));
|
||||
double wms = eventMs(e0, e1);
|
||||
mhsWindow = (double)nonces * (double)done / (wms / 1000.0) / 1e6;
|
||||
smp.join();
|
||||
std::lock_guard<std::mutex> g(smp.mu);
|
||||
nSamples = (int)smp.samples.size();
|
||||
for (const Sample& s : smp.samples) {
|
||||
if (s.t < 10000.0 || s.t > (double)o.powerSeconds * 1000.0) continue; // mean after 10 s, inside the load window
|
||||
powerW += s.watts; smMHz += s.smMHz; memMHz += s.memMHz; ++nUsed;
|
||||
}
|
||||
if (nUsed) { powerW /= nUsed; smMHz /= nUsed; memMHz /= nUsed; }
|
||||
std::printf("power window: %d batches in %.1f s -> %.3f MH/s (events, incl. sync gaps); nvidia-smi %d samples, %d after 10 s: mean %.1f W, SM %.0f MHz, mem %.0f MHz\n",
|
||||
done, wms / 1000.0, mhsWindow, nSamples, nUsed, powerW, smMHz, memMHz);
|
||||
if (nUsed) std::printf(" -> %.3f MH/s per W (window rate / mean W)\n", mhsWindow / powerW);
|
||||
else std::printf(" WARNING: no nvidia-smi samples inside the window (is nvidia-smi on PATH?)\n");
|
||||
}
|
||||
|
||||
size_t freeH = 0;
|
||||
CUDA_CHECK(cudaMemGetInfo(&freeH, &total));
|
||||
std::printf("device memory while hashing: %.0f MiB used (leaves freed)\n", (double)(total - freeH) / 1048576.0);
|
||||
|
||||
bool overall = cachePass && leafPass && dsPass && vecPass && fpPass && itemsEqual == 1024 && (chunkedEqual < 0 || chunkedEqual == 1024);
|
||||
std::printf("RESULT mode=%s gpu=%s cache_fill_ms=%.2f host_cache_ms=%.1f leaves_ms=%.2f build_ms=%.2f mem_build_mib=%.0f chunk_mib=%d chunked_ms=%.2f mem_chunked_mib=%.0f chunked_equal=%d items_equal=%d/1024 fingerprint=%016llx mhs=%.3f mhs_window=%.3f watts=%.1f sm_mhz=%.0f mem_mhz=%.0f cache=%s dataset=%s vectors=%s overall=%s\n",
|
||||
modeName, prop.name, cacheFill[1], hostCacheMs, leavesMs[1], buildMs[1], usedBuildMiB, o.chunkMib, chunkedMs, usedChunkMiB, chunkedEqual, itemsEqual, (unsigned long long)fp, mhs, mhsWindow, powerW, smMHz, memMHz,
|
||||
cachePass ? "PASS" : "FAIL", dsPass ? "PASS" : "FAIL", o.sd1 ? "n/a" : (vecPass ? "PASS" : "FAIL"), overall ? "PASS" : "FAIL");
|
||||
|
||||
CUDA_CHECK(cudaFree(dOut));
|
||||
CUDA_CHECK(cudaFree(dDs));
|
||||
CUDA_CHECK(cudaFree(dCache));
|
||||
return overall ? 0 : 1;
|
||||
}
|
||||
139
proto-newpow/state-dataset/cpu_rows.c
Normal file
|
|
@ -0,0 +1,139 @@
|
|||
/* state-dataset prototype: CPU rows for igneum-build-1 (Horizon lane 8, class sd1). Plain C, -O2.
|
||||
*
|
||||
* B.1 the verifier's extra cost per 32-lane unit (4,096 items per unit), ms per unit over 100 units, one core:
|
||||
* (i) 4,096 random 64-byte reads from a host-resident leaf array of 2 GiB and of 8 GiB (random t, independent,
|
||||
* the array written once so every page is resident)
|
||||
* (ii) 4,096 leaves derived on the fly from the state root (one ChaCha12 block each, no array)
|
||||
* (iii) 4,096 x mh_item on the host cache, naive (no interleaving); plus leaf + mh_item_sd
|
||||
* B.2 the daily snapshot build on CPU: the 1 GiB leaf array (2^24 ChaCha12 blocks) on one core and on N threads
|
||||
* (OpenMP), and the host cache fill on one core.
|
||||
*
|
||||
* Build: gcc -O2 -std=c11 -fopenmp -o cpu_rows cpu_rows.c
|
||||
* Run: taskset -c 4 nice -n 19 ./cpu_rows --rows (B.1 and the one-core B.2 rows)
|
||||
* taskset -c 4-35 nice -n 19 ./cpu_rows --leaves 32 (the N-thread leaf build)
|
||||
*/
|
||||
#define _POSIX_C_SOURCE 200809L
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <time.h>
|
||||
#include <omp.h>
|
||||
#define IGNEUM_NO_CUDA
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
#include "sd.h"
|
||||
|
||||
static double nowMs(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1e3 + ts.tv_nsec / 1e6; }
|
||||
static uint64_t splitmix64(uint64_t* s) {
|
||||
*s += 0x9E3779B97F4A7C15ull; uint64_t z = *s;
|
||||
z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9ull; z = (z ^ (z >> 27)) * 0x94D049BB133111EBull;
|
||||
return z ^ (z >> 31);
|
||||
}
|
||||
|
||||
static volatile uint32_t sink;
|
||||
|
||||
/* (i): random 64-byte reads from an array of gib GiB */
|
||||
static void row_reads(int gib, int units) {
|
||||
size_t items = ((size_t)gib << 30) / 64u;
|
||||
uint32_t* arr = (uint32_t*)malloc((size_t)gib << 30);
|
||||
if (!arr) { printf(" (i) %d GiB: malloc failed\n", gib); return; }
|
||||
double t0 = nowMs();
|
||||
/* write every 64-byte line once (resident pages); the content is a cheap pattern, the read cost does not depend on it */
|
||||
for (size_t i = 0; i < items; ++i) { uint32_t* l = arr + i * 16u; l[0] = (uint32_t)i; l[15] = (uint32_t)(i >> 32) ^ 0x5d1u; }
|
||||
double fillMs = nowMs() - t0;
|
||||
uint64_t s = 0x5d1ull; uint32_t acc[16] = { 0 };
|
||||
double sum = 0, mn = 1e9, mx = 0;
|
||||
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
|
||||
for (int u = 0; u < units; ++u) {
|
||||
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)(splitmix64(&s) % items);
|
||||
double a = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { const uint32_t* l = arr + (size_t)ts[k] * 16u; for (int i = 0; i < 16; ++i) acc[i] ^= l[i]; }
|
||||
double b = nowMs();
|
||||
sum += b - a; if (b - a < mn) mn = b - a; if (b - a > mx) mx = b - a;
|
||||
}
|
||||
for (int i = 0; i < 16; ++i) sink ^= acc[i];
|
||||
printf(" (i) 4,096 random 64-byte reads from a %d GiB resident array: %.3f ms per unit (min %.3f, max %.3f; %.0f ns per read; array write pass %.0f ms)\n",
|
||||
gib, sum / units, mn, mx, sum / units * 1e6 / 4096.0, fillMs);
|
||||
free(ts); free(arr);
|
||||
}
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
int rows = 0, leavesThreads = 0, units = 100;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
if (!strcmp(argv[i], "--rows")) rows = 1;
|
||||
else if (!strcmp(argv[i], "--leaves") && i + 1 < argc) leavesThreads = atoi(argv[++i]);
|
||||
else if (!strcmp(argv[i], "--units") && i + 1 < argc) units = atoi(argv[++i]);
|
||||
else { printf("usage: cpu_rows [--rows] [--leaves N] [--units 100]\n"); return 2; }
|
||||
}
|
||||
const uint32_t words = 1u << IGNEUM_DATASET_LOG2, nItems = words / 16u;
|
||||
const uint32_t cacheWords = 1u << IGNEUM_CACHE_LOG2_WORDS;
|
||||
|
||||
if (rows) {
|
||||
printf("B.1 per-unit rows (4,096 per unit, %d units, one thread, ms per unit)\n", units);
|
||||
row_reads(2, units);
|
||||
row_reads(8, units);
|
||||
/* (ii) */
|
||||
{
|
||||
uint64_t s = 0x5d1ull; double sum = 0;
|
||||
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
|
||||
for (int u = 0; u < units; ++u) {
|
||||
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)splitmix64(&s) & (nItems - 1u);
|
||||
double a = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16]; mh_leaf(ts[k], leaf); sink ^= leaf[0] ^ leaf[15]; }
|
||||
sum += nowMs() - a;
|
||||
}
|
||||
printf(" (ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): %.3f ms per unit (%.0f ns per leaf)\n", sum / units, sum / units * 1e6 / 4096.0);
|
||||
free(ts);
|
||||
}
|
||||
/* B.2 host cache fill, one core; then (iii) */
|
||||
uint32_t* cache = (uint32_t*)malloc((size_t)cacheWords * 4u);
|
||||
double c0 = nowMs();
|
||||
for (uint32_t seg = 0; seg < IGNEUM_CACHE_SEGMENTS; ++seg) mh_cache_segment(cache, seg);
|
||||
double cacheMs = nowMs() - c0;
|
||||
printf("B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: %.3f s\n", cacheMs / 1000.0);
|
||||
{
|
||||
uint64_t s = 0x5d1ull; double sumItem = 0, sumSd = 0, mn = 1e9, mx = 0;
|
||||
uint32_t* ts = (uint32_t*)malloc(4096 * 4);
|
||||
for (int u = 0; u < units; ++u) {
|
||||
for (int k = 0; k < 4096; ++k) ts[k] = (uint32_t)splitmix64(&s) & (nItems - 1u);
|
||||
double a = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { uint32_t it[16]; mh_item(cache, ts[k], it); sink ^= it[0]; }
|
||||
double b = nowMs();
|
||||
for (int k = 0; k < 4096; ++k) { uint32_t leaf[16], it[16]; mh_leaf(ts[k], leaf); mh_item_sd(cache, leaf, ts[k], it); sink ^= it[0]; }
|
||||
double c = nowMs();
|
||||
sumItem += b - a; sumSd += c - b; if (b - a < mn) mn = b - a; if (b - a > mx) mx = b - a;
|
||||
}
|
||||
printf(" (iii) 4,096 x mh_item on the host cache, naive (no interleaving): %.2f ms per unit (min %.2f, max %.2f)\n", sumItem / units, mn, mx);
|
||||
printf(" 4,096 x (leaf + mh_item_sd), naive: %.2f ms per unit\n", sumSd / units);
|
||||
printf(" reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.\n");
|
||||
free(ts);
|
||||
}
|
||||
free(cache);
|
||||
/* B.2 leaf array on one core */
|
||||
{
|
||||
uint32_t* leaves = (uint32_t*)malloc((size_t)nItems * 64u);
|
||||
double a = nowMs();
|
||||
for (uint32_t t = 0; t < nItems; ++t) mh_leaf(t, leaves + (size_t)t * 16u);
|
||||
double ms = nowMs() - a;
|
||||
sink ^= leaves[0] ^ leaves[(size_t)nItems * 16u - 1u];
|
||||
printf("B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: %.3f s (%.0f ns per leaf)\n", ms / 1000.0, ms * 1e6 / nItems);
|
||||
free(leaves);
|
||||
}
|
||||
}
|
||||
if (leavesThreads > 0) {
|
||||
omp_set_num_threads(leavesThreads);
|
||||
uint32_t* leaves = (uint32_t*)malloc((size_t)nItems * 64u);
|
||||
/* first touch in parallel too, then time a second full build so page faults are not in the number */
|
||||
for (int pass = 0; pass < 2; ++pass) {
|
||||
double a = nowMs();
|
||||
#pragma omp parallel for schedule(static)
|
||||
for (int64_t t = 0; t < (int64_t)nItems; ++t) mh_leaf((uint32_t)t, leaves + (size_t)t * 16u);
|
||||
double ms = nowMs() - a;
|
||||
printf("B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), %d threads (OpenMP, %d actual): %.3f s%s\n", leavesThreads, omp_get_max_threads(), ms / 1000.0, pass == 0 ? " (first pass, includes page faults)" : "");
|
||||
}
|
||||
sink ^= leaves[0];
|
||||
free(leaves);
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
164
proto-newpow/state-dataset/kernel.cu
Normal file
|
|
@ -0,0 +1,164 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
213
proto-newpow/state-dataset/kernel_sd.cu
Normal file
|
|
@ -0,0 +1,213 @@
|
|||
// state-dataset prototype: the mx8-genesis pack kernel plus igneum_leaves and igneum_build_sd (appended at the end).
|
||||
// The hash kernel igneum_hash and every other line of the pack are byte for byte the pack's kernel.cu.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
#include "sd.h" // state-dataset prototype: mh_leaf, mh_item_sd (lane 8, class sd1)
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// state-dataset prototype (class sd1). One thread per item.
|
||||
// igneum_leaves: leaves[t * 16 .. t * 16 + 15] = mh_leaf(t) (the synthetic 64-byte state leaf, sd.h).
|
||||
// igneum_build_sd: item t = mh_item_sd(cache, leaves + 16 t, t). The leaf read is one coalesced 64-byte read per
|
||||
// thread (consecutive threads read consecutive leaves), so streaming the leaf array in chunks is a matter of
|
||||
// launching over an item range [t0, t0 + n) with the chunk's leaves at leaves - 16 t0: nothing in the kernel
|
||||
// depends on the whole array being resident.
|
||||
__global__ void igneum_leaves(uint32_t* leaves, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t leaf[16];
|
||||
mh_leaf(t, leaf);
|
||||
uint32_t* d = leaves + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = leaf[i];
|
||||
}
|
||||
}
|
||||
__global__ void igneum_build_sd(uint32_t* ds, const uint32_t* cache, const uint32_t* leaves, uint32_t t0, uint32_t nItems) {
|
||||
uint32_t t = t0 + blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < t0 + nItems) {
|
||||
uint32_t leaf[16];
|
||||
const uint32_t* l = leaves + (size_t)(t - t0) * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) leaf[i] = l[i];
|
||||
uint32_t s[16];
|
||||
mh_item_sd(cache, leaf, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_leaves(uint32_t* leaves, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_leaves<<<grid, block>>>(leaves, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
// Builds items [t0, t0 + nItems) from the leaves of that range (leaves points at leaf t0).
|
||||
cudaError_t igneum_launch_build_sd(uint32_t* ds, const uint32_t* cache, const uint32_t* leaves, uint32_t t0, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build_sd<<<grid, block>>>(ds, cache, leaves, t0, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
109
proto-newpow/state-dataset/memhard.h
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
|
||||
// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0x3067619fu ^ prev[4];
|
||||
x[5] = 0x3c269176u ^ prev[5];
|
||||
x[6] = 0x84a03b03u ^ prev[6];
|
||||
x[7] = 0xf8c63294u ^ prev[7];
|
||||
x[8] = 0xff977c5bu ^ prev[8];
|
||||
x[9] = 0xe60def3eu ^ prev[9];
|
||||
x[10] = 0x63630141u ^ prev[10];
|
||||
x[11] = 0xb8fbcb58u ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
|
||||
s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
|
||||
s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
|
||||
s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
|
||||
s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
|
||||
s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
|
||||
s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
|
||||
s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
|
||||
s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
|
||||
s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
|
||||
s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
|
||||
s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
|
||||
s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
|
||||
s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
|
||||
s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
|
||||
s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0x3067619fu;
|
||||
s[1] = 0x3c269176u;
|
||||
s[2] = 0x84a03b03u;
|
||||
s[3] = 0xf8c63294u;
|
||||
s[4] = 0xff977c5bu;
|
||||
s[5] = 0xe60def3eu;
|
||||
s[6] = 0x63630141u;
|
||||
s[7] = 0xb8fbcb58u;
|
||||
s[8] = t * 0x42146205u + 0xbab68293u;
|
||||
s[9] = t * 0x52cbe0fbu + 0xcc162340u;
|
||||
s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
|
||||
s[11] = t * 0x6728907fu + 0xe62b8997u;
|
||||
s[12] = t * 0xd81d9751u + 0xc9c80297u;
|
||||
s[13] = t * 0x132952c3u + 0xf74a1654u;
|
||||
s[14] = t * 0xf60de277u + 0x3d704af5u;
|
||||
s[15] = t * 0x05358035u + 0x3cf522b7u;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
66
proto-newpow/state-dataset/program.h
Normal file
|
|
@ -0,0 +1,66 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-genesis"
|
||||
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0xe323b9dcaf283a6full
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "mx8"
|
||||
#define IGNEUM_CLASS_MIXER_MULT 8
|
||||
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
|
||||
#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIXER_MULT 8 // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)
|
||||
#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
2
proto-newpow/state-dataset/results/cpu/leaves32-2.log
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.179 s (first pass, includes page faults)
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.083 s
|
||||
2
proto-newpow/state-dataset/results/cpu/leaves32.log
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.172 s (first pass, includes page faults)
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.073 s
|
||||
9
proto-newpow/state-dataset/results/cpu/rows.log
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
|
||||
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.108 ms per unit (min 0.092, max 0.172; 26 ns per read; array write pass 1022 ms)
|
||||
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.139 ms per unit (min 0.127, max 0.183; 34 ns per read; array write pass 3783 ms)
|
||||
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.284 ms per unit (69 ns per leaf)
|
||||
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.447 s
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms per unit (min 8.91, max 9.60)
|
||||
4,096 x (leaf + mh_item_sd), naive: 9.31 ms per unit
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.670 s (100 ns per leaf)
|
||||
9
proto-newpow/state-dataset/results/cpu/rows2.log
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
|
||||
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.163 ms per unit (min 0.144, max 0.201; 40 ns per read; array write pass 1252 ms)
|
||||
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.209 ms per unit (min 0.203, max 0.229; 51 ns per read; array write pass 4325 ms)
|
||||
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.285 ms per unit (70 ns per leaf)
|
||||
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.452 s
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 11.24 ms per unit (min 9.38, max 13.51)
|
||||
4,096 x (leaf + mh_item_sd), naive: 11.64 ms per unit
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.677 s (100 ns per leaf)
|
||||
27
proto-newpow/state-dataset/results/cpu/run_cpu.log
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
== Tue Oct 6 07:44:26 PM UTC 2026 on igneum-build-1, load 18.98 28.50 18.66
|
||||
Model name: AMD EPYC 9454P 48-Core Processor
|
||||
gcc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0
|
||||
== build
|
||||
== B.1 and one-core B.2 rows, core 4, nice 19
|
||||
B.1 per-unit rows (4,096 per unit, 100 units, one thread, ms per unit)
|
||||
(i) 4,096 random 64-byte reads from a 2 GiB resident array: 0.108 ms per unit (min 0.092, max 0.172; 26 ns per read; array write pass 1022 ms)
|
||||
(i) 4,096 random 64-byte reads from a 8 GiB resident array: 0.139 ms per unit (min 0.127, max 0.183; 34 ns per read; array write pass 3783 ms)
|
||||
(ii) 4,096 leaves derived on the fly (one ChaCha12 block each, no array): 0.284 ms per unit (69 ns per leaf)
|
||||
B.2 host cache fill (256 MiB, 65536 segments x 64 ChaCha12 blocks), one core: 0.447 s
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms per unit (min 8.91, max 9.60)
|
||||
4,096 x (leaf + mh_item_sd), naive: 9.31 ms per unit
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), one core: 1.670 s (100 ns per leaf)
|
||||
== B.2 leaf array on 32 threads (cores 4-35), nice 19
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.172 s (first pass, includes page faults)
|
||||
B.2 leaf array (1 GiB, 2^24 ChaCha12 blocks), 32 threads (OpenMP, 32 actual): 0.073 s
|
||||
== verify_sd per-unit rows on this CPU (core 4), for comparison
|
||||
verify_sd mode sd1
|
||||
host cache fill: 480.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
||||
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.285 ms
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms (min 8.92, max 9.44)
|
||||
4,096 x (leaf + mh_item_sd), naive: 9.27 ms
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
RESULT verify mode=sd1 host_cache_ms=480.4 items_equal=-1 lanes_equal=-1/0
|
||||
== done Tue Oct 6 07:44:39 PM UTC 2026, load 16.75 27.55 18.50
|
||||
8
proto-newpow/state-dataset/results/cpu/verify_rows.log
Normal file
|
|
@ -0,0 +1,8 @@
|
|||
verify_sd mode sd1
|
||||
host cache fill: 480.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
||||
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.285 ms
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 9.01 ms (min 8.92, max 9.44)
|
||||
4,096 x (leaf + mh_item_sd), naive: 9.27 ms
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
RESULT verify mode=sd1 host_cache_ms=480.4 items_equal=-1 lanes_equal=-1/0
|
||||
24
proto-newpow/state-dataset/results/gpu/control.log
Normal file
|
|
@ -0,0 +1,24 @@
|
|||
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.86 ms first, 1.81 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 380.5 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.1 M items/s
|
||||
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
|
||||
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
|
||||
dataset self-test: PASS (64 random words vs host mh_word: PASS)
|
||||
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
|
||||
vector warp base 0: PASS (0 of 32 lanes differ)
|
||||
vector warp base 4096: PASS (0 of 32 lanes differ)
|
||||
vector warp base 1000000: PASS (0 of 32 lanes differ)
|
||||
warm-up batch: 2^24 hashes in 266.02 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
|
||||
vector warp base 0 in batch: PASS
|
||||
vector warp base 4096 in batch: PASS
|
||||
vector warp base 1000000 in batch: PASS
|
||||
dump: 4 warps (bases 0, 32, ...) written to dump_control.txt
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.56 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.078 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.6 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.304 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.81 host_cache_ms=380.5 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.078 watts=207.6 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS
|
||||
23
proto-newpow/state-dataset/results/gpu/control2.log
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.86 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 380.0 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.2 M items/s
|
||||
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
|
||||
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
|
||||
dataset self-test: PASS (64 random words vs host mh_word: PASS)
|
||||
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
|
||||
vector warp base 0: PASS (0 of 32 lanes differ)
|
||||
vector warp base 4096: PASS (0 of 32 lanes differ)
|
||||
vector warp base 1000000: PASS (0 of 32 lanes differ)
|
||||
warm-up batch: 2^24 hashes in 265.99 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
|
||||
vector warp base 0 in batch: PASS
|
||||
vector warp base 4096 in batch: PASS
|
||||
vector warp base 1000000 in batch: PASS
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.55 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.079 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 205.2 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.307 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=380.0 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.079 watts=205.2 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS
|
||||
134
proto-newpow/state-dataset/results/gpu/dump_control.txt
Normal file
|
|
@ -0,0 +1,134 @@
|
|||
mode control
|
||||
warps 4
|
||||
base 0
|
||||
19b56348bc85304d
|
||||
b08a9cfb44aa720f
|
||||
e1f8f627f780eff7
|
||||
2ff3e86ff1696161
|
||||
b65e0578c257e8ac
|
||||
f37b5705c2bebaae
|
||||
b19fef670982389e
|
||||
2374331b28827a11
|
||||
f492cdd05dda7f88
|
||||
37700f1385e19885
|
||||
a27524f3010b3e87
|
||||
6c8a24de8b938c43
|
||||
3dc017c8820cfd16
|
||||
5c80d37146657b68
|
||||
214082cac03f734a
|
||||
1d67c665145d72f3
|
||||
9fb0ae736eb87ff5
|
||||
cd7d416ae3a6be61
|
||||
0069cda85c6de58f
|
||||
0b7d82e717224baa
|
||||
600960f6d7be30f0
|
||||
eb29032663b0c4d2
|
||||
f29583a5766408f1
|
||||
6c9532a99ad7314d
|
||||
9a949fd0959cc80f
|
||||
a20258520e7f6c25
|
||||
f2ab11f9bb032e38
|
||||
cc967bcd0c8d07c1
|
||||
37745267bb3231f2
|
||||
35a046048c2b69b3
|
||||
aa51834cd3f364f3
|
||||
359192708e4f754a
|
||||
base 32
|
||||
bbea9a410fe7d5b7
|
||||
11bf76519a1414d4
|
||||
59c52e47a752a4fa
|
||||
3964f1251e8b90e8
|
||||
75230c9369b790af
|
||||
8dec6e557dc03a87
|
||||
0b819148666b3f13
|
||||
7707a6b8d499e40f
|
||||
445de8aa5b9699d5
|
||||
6becbf9e66d7e9de
|
||||
7a46928b6ed6861f
|
||||
0cbfca3b2c769872
|
||||
04e97a5b926b55d5
|
||||
2a7ba12634ae81bd
|
||||
96cb4eaf4e25600a
|
||||
2826bd0a4355a91d
|
||||
d7ba781350c80652
|
||||
e87c0820f6136930
|
||||
e8864f498a00a8e5
|
||||
c1b445f3b5748847
|
||||
72def57f2e515d11
|
||||
f4708f2d80c98b8d
|
||||
da99416aa6245b67
|
||||
90f6e6d21f11fc66
|
||||
a9e3c5f36c715c31
|
||||
e5e2a586d61e4869
|
||||
98e1ce97884663da
|
||||
36e8af738edddc85
|
||||
8530526bbf3d85bb
|
||||
22128f72173887c6
|
||||
30b7aa6ab94449e2
|
||||
2ea3ca54249bd934
|
||||
base 64
|
||||
68b73c212b47e6f2
|
||||
55acf1fa7bbac2dd
|
||||
45001afaa68359f7
|
||||
3271315def7632fb
|
||||
8af0fe0f8b16148b
|
||||
2e0c7f7bac0276e1
|
||||
fe178e4364b66e31
|
||||
f39df092fabe08ad
|
||||
e6d7ad0a5d97a7eb
|
||||
0013aa74d3b41bba
|
||||
f57f5241a626b717
|
||||
747067ae7e06c706
|
||||
50325d1d39c18b09
|
||||
1fb5aa02196bd962
|
||||
78eaa19e7dd002f4
|
||||
8fe0600a8ff1c365
|
||||
021d685d64796246
|
||||
281446f54c93f2a5
|
||||
ba2cc4bef235b006
|
||||
bfb4bc312bd962a4
|
||||
cecc08bdb679ca9c
|
||||
1f079cb2d03b55a1
|
||||
a8cc9f0647e103a0
|
||||
2615618b9cb4877f
|
||||
853f0dd6f03607f2
|
||||
bf54084008cd8ca3
|
||||
21986116f4e97e6c
|
||||
42dbca5da95d1b1c
|
||||
d7669ed19dea6986
|
||||
17e407aa53422e43
|
||||
8c7d4f5bfde2f58e
|
||||
13adadad05f4540a
|
||||
base 96
|
||||
3a0fd3ada6b5a797
|
||||
aa5a637cb9caa3b7
|
||||
728b73955e0ad47d
|
||||
5092aa36c478581c
|
||||
a463e4220676d004
|
||||
ce92dfc65d3af26e
|
||||
2943fd8967c7f215
|
||||
71acd3ccc8054eaa
|
||||
2b93c6cd8c22c051
|
||||
00f10e1b004bab4d
|
||||
02e2733b82de9bce
|
||||
1c9d5cd0ed5350ce
|
||||
dcd13404e4ab0b21
|
||||
d4afc0e3a2814f63
|
||||
f2b37d4b0322ff2f
|
||||
0212964cea5688de
|
||||
546351f9a3ec9957
|
||||
d1fe39207bff2a50
|
||||
a9a38d54a9f85090
|
||||
e039372a0cc1aa9e
|
||||
008a4289a796e05e
|
||||
a9a1dca5b9de1fbc
|
||||
751a1769f1203a6d
|
||||
5b66d9c952febc43
|
||||
47ac39abda0dd053
|
||||
c7b86597a7ef8d83
|
||||
d87824ea2fd452c4
|
||||
82a39552b84b113c
|
||||
562aa1cb671a8056
|
||||
4dc55d7e9e0829dd
|
||||
bb0cf26fbb506ddc
|
||||
353c62d9a3b6aaed
|
||||
134
proto-newpow/state-dataset/results/gpu/dump_sd1.txt
Normal file
|
|
@ -0,0 +1,134 @@
|
|||
mode sd1
|
||||
warps 4
|
||||
base 0
|
||||
b600edbed969becc
|
||||
2d988317678e9099
|
||||
29609900b3a81764
|
||||
2506a510ec612b1c
|
||||
cdc516143acae7eb
|
||||
dd2901cc2335449d
|
||||
41614b8fbf48a689
|
||||
317f5df1eb7c0cbe
|
||||
2da7edd46d732703
|
||||
b89b6f78568676a9
|
||||
e93e96d589f70d8b
|
||||
197bd4558fe42bf7
|
||||
911c2ed1bca014bc
|
||||
448984ee31e0f576
|
||||
6a9af9585e6ce6ba
|
||||
446ba1d41100da43
|
||||
d206864c5393aef9
|
||||
46b93bf7a7b89196
|
||||
85ce18e332a13c31
|
||||
34f1d32a0f153659
|
||||
d2b8703f431f1237
|
||||
3bda4edaa92faa6f
|
||||
d8e5b9f18abf75fd
|
||||
234a323a6d619bfd
|
||||
9558b3c4e42cd2f7
|
||||
016372acc63f6372
|
||||
7ebac0c7e5fc87f7
|
||||
81f40d3ef4b5508b
|
||||
aa6ec7854aacde98
|
||||
692a050230de2fe8
|
||||
18b7df938406f9a1
|
||||
2412df7ada5a3202
|
||||
base 32
|
||||
e451566071a2a0ba
|
||||
d63578a9312e5b32
|
||||
6bdaecc6d5f7158b
|
||||
694c691e7dc1de0a
|
||||
d98a0709679797b3
|
||||
17a7c68f4ff5e95f
|
||||
58222d4156171349
|
||||
af4bcd67686aad89
|
||||
fdad4d96931b9f2a
|
||||
5e16c79bd1f2c688
|
||||
7437b2913e288981
|
||||
7b4485a218901ccf
|
||||
540694fa54e15619
|
||||
e42600d66ac6e0f5
|
||||
3b2b6322a93124d2
|
||||
b474e79e977515ae
|
||||
85c356e1a7f560fe
|
||||
47961522447a02fc
|
||||
5cf31818bea49c4a
|
||||
a517ba58dbed530c
|
||||
d1b37bf54bd68236
|
||||
27905e28697a6ef7
|
||||
42a743ba99c22bf9
|
||||
5ec48e4245545d4a
|
||||
38ecfefd174944d8
|
||||
3dd1891a53a9b415
|
||||
44120379678fcd2a
|
||||
e2b43434e9e5b97f
|
||||
db448577e2229d66
|
||||
41eb61946c60e9a7
|
||||
53db71d828b8d7e6
|
||||
ea19aabc66d1b11a
|
||||
base 64
|
||||
34548f5056f27d78
|
||||
e7fea9a49cf2556e
|
||||
ebdaf245ef41ec6d
|
||||
3c77731e46a78bd3
|
||||
e985324128b13372
|
||||
c13540b76c4dfc4a
|
||||
9050d43f68bf20e9
|
||||
4aa7762e9a08d90a
|
||||
d106b0610790817b
|
||||
5c9c91398c9b5e7c
|
||||
56b71969e73c466e
|
||||
2d61dea47bf479bb
|
||||
2ddb7cc8f81fd24c
|
||||
f422f7adbb7a230e
|
||||
794858ac6a6bada1
|
||||
c7a37f1c3cebeea4
|
||||
241f6632d74d699e
|
||||
f8e536083c72f300
|
||||
3234e942615b9891
|
||||
573677792f6bfe53
|
||||
6838b021f005dee5
|
||||
7e1042640bdf5609
|
||||
4a0399c7b6ca97fd
|
||||
37ad996ce5fb721a
|
||||
f633ed2dc43bc76d
|
||||
e569e4a049054982
|
||||
a9c5fa29d0ab9603
|
||||
ca0dcf5025e589f4
|
||||
17687fa3872dd204
|
||||
965e7bea52b291a7
|
||||
d3aec969f424c835
|
||||
44453d33abd93ef6
|
||||
base 96
|
||||
c13d6ea3e05c5e46
|
||||
1c291747d4676990
|
||||
349040c82d1e4ad8
|
||||
f05a14026be82f03
|
||||
5ee5f52611690c2b
|
||||
60c99fb5c988e677
|
||||
3d0899be8af84174
|
||||
74dad192f88cca2c
|
||||
005480c1f5e84d6f
|
||||
a38fa3da9db77186
|
||||
9495da0d026c51fc
|
||||
54f4ed7062e92211
|
||||
6f910c4ad2774774
|
||||
55752a05b9fa5b71
|
||||
883b5367142927bf
|
||||
96449de7c472e375
|
||||
f8722d3292d2948b
|
||||
21d0485f77308767
|
||||
010a6d05ede7793e
|
||||
56cbfe0810728799
|
||||
e136e570c4ed04a9
|
||||
848c020abe56db44
|
||||
71aa3cddb1bc9408
|
||||
58757e0f19ddb3ef
|
||||
f2e4273de46c5968
|
||||
8664ce7ab76eceac
|
||||
881db7223c9a750b
|
||||
788a1553b16895ea
|
||||
c84a69728aab9555
|
||||
9019c8d87a1da486
|
||||
4ab485aa2deb7ca5
|
||||
b2795a944cca750d
|
||||
1026
proto-newpow/state-dataset/results/gpu/items_control.txt
Normal file
1026
proto-newpow/state-dataset/results/gpu/items_sd1.txt
Normal file
85
proto-newpow/state-dataset/results/gpu/run_gpu.log
Normal file
|
|
@ -0,0 +1,85 @@
|
|||
== Tue Oct 6 19:43:12 UTC 2026 on a4cac49bb840
|
||||
NVIDIA GeForce RTX 4090, 595.91.07, 3135 MHz, 24564 MiB
|
||||
Build cuda_12.8.r12.8/compiler.35583870_0
|
||||
== build
|
||||
== control
|
||||
state-dataset bench mode control pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.86 ms first, 1.81 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 380.5 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
dataset build (igneum_build, the pack's): 30.62 ms first, 30.55 ms second -> 549.1 M items/s
|
||||
device memory after the build: 1675 MiB used (context 395 + cache 256 + dataset 1024 MiB)
|
||||
dataset[0..3] = fdad4319 1a7b68e1 de6db608 13d73892 head 16 vs Mac PASS, word [MASK] vs Mac PASS, 64 Mac samples PASS
|
||||
dataset self-test: PASS (64 random words vs host mh_word: PASS)
|
||||
item bit-exactness (in-process, host mh_item on the host cache): 1024 of 1024 items equal
|
||||
vector warp base 0: PASS (0 of 32 lanes differ)
|
||||
vector warp base 4096: PASS (0 of 32 lanes differ)
|
||||
vector warp base 1000000: PASS (0 of 32 lanes differ)
|
||||
warm-up batch: 2^24 hashes in 266.02 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): 7c28cfb06c5c65a9 = 7c28cfb06c5c65a9 (the pack's)
|
||||
vector warp base 0 in batch: PASS
|
||||
vector warp base 4096 in batch: PASS
|
||||
vector warp base 1000000 in batch: PASS
|
||||
dump: 4 warps (bases 0, 32, ...) written to dump_control.txt
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.56 ms -> 63.083 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.078 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.6 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.304 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=control gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.81 host_cache_ms=380.5 leaves_ms=0.00 build_ms=30.55 mem_build_mib=1675 items_equal=1024/1024 fingerprint=7c28cfb06c5c65a9 mhs=63.083 mhs_window=63.078 watts=207.6 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=PASS overall=PASS
|
||||
== sd1
|
||||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 383.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.56 ms first, 5.49 ms second -> 195.4 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
dump: 4 warps (bases 0, 32, ...) written to dump_sd1.txt
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.35 ms -> 63.088 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.9 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.303 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=383.9 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.088 mhs_window=63.083 watts=207.9 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||
== verify (plain C, CPU core 2)
|
||||
verify_sd mode control
|
||||
host cache fill: 587.4 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
interpreter vs Mac vector base 0: 32 of 32 lanes equal PASS
|
||||
item bit-exactness (plain C, mh_item on the host cache): 1024 of 1024 items equal PASS
|
||||
warp base 0: 32 of 32 lanes equal (lane 0 gpu 19b56348bc85304d interp 19b56348bc85304d)
|
||||
warp base 32: 32 of 32 lanes equal (lane 0 gpu bbea9a410fe7d5b7 interp bbea9a410fe7d5b7)
|
||||
warp base 64: 32 of 32 lanes equal (lane 0 gpu 68b73c212b47e6f2 interp 68b73c212b47e6f2)
|
||||
warp base 96: 32 of 32 lanes equal (lane 0 gpu 3a0fd3ada6b5a797 interp 3a0fd3ada6b5a797)
|
||||
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (37 ms interpreting, 4096 derivations per warp)
|
||||
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
||||
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.369 ms
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 10.14 ms (min 9.98, max 10.31)
|
||||
4,096 x (leaf + mh_item_sd), naive: 10.56 ms
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
RESULT verify mode=control host_cache_ms=587.4 items_equal=1024 lanes_equal=128/128
|
||||
verify_sd mode sd1
|
||||
host cache fill: 572.9 ms one thread; FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS
|
||||
item bit-exactness (plain C, mh_item_sd with mh_leaf on the host cache): 1024 of 1024 items equal PASS
|
||||
warp base 0: 32 of 32 lanes equal (lane 0 gpu b600edbed969becc interp b600edbed969becc)
|
||||
warp base 32: 32 of 32 lanes equal (lane 0 gpu e451566071a2a0ba interp e451566071a2a0ba)
|
||||
warp base 64: 32 of 32 lanes equal (lane 0 gpu 34548f5056f27d78 interp 34548f5056f27d78)
|
||||
warp base 96: 32 of 32 lanes equal (lane 0 gpu c13d6ea3e05c5e46 interp c13d6ea3e05c5e46)
|
||||
GPU hash outputs vs interpreter: 128 of 128 lanes equal PASS (42 ms interpreting, 4096 derivations per warp)
|
||||
per-unit rows on this CPU (4,096 per unit, 100 units, one thread, ms per unit):
|
||||
(ii) 4,096 leaf derivations (ChaCha12 block each, no array): 0.374 ms
|
||||
(iii) 4,096 x mh_item on the host cache, naive (no interleaving): 10.21 ms (min 10.09, max 12.24)
|
||||
4,096 x (leaf + mh_item_sd), naive: 10.58 ms
|
||||
reference: the project's Rust verifier does 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items; the naive figure here is an upper bound.
|
||||
RESULT verify mode=sd1 host_cache_ms=572.9 items_equal=1024 lanes_equal=128/128
|
||||
== done Tue Oct 6 19:44:17 UTC 2026
|
||||
23
proto-newpow/state-dataset/results/gpu/sd1-chunk256.log
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.73 ms first, 1.70 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 380.7 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.16 ms first, 5.10 ms second -> 210.7 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.00 ms first, 31.92 ms second -> 525.6 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
chunked build (leaves from pinned host memory in 256 MiB chunks, 4 chunks, one stream): 75.47 ms total (copies + builds); whole-array H2D alone 61.92 ms = 17.3 GB/s, D2H 54.61 ms
|
||||
device memory during the chunked build: 1933 MiB used (context + cache 256 + dataset 1024 + chunk 256 MiB); items after the chunked rebuild: 1024 of 1024 equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 265.97 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.40 ms -> 63.086 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.70 host_cache_ms=380.7 leaves_ms=5.10 build_ms=31.92 mem_build_mib=2699 chunk_mib=256 chunked_ms=75.47 mem_chunked_mib=1933 chunked_equal=1024 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.086 mhs_window=0.000 watts=0.0 sm_mhz=0 mem_mhz=0 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||
23
proto-newpow/state-dataset/results/gpu/sd1-chunk64.log
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.85 ms first, 1.82 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 382.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.53 ms first, 5.46 ms second -> 196.5 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
chunked build (leaves from pinned host memory in 64 MiB chunks, 16 chunks, one stream): 75.80 ms total (copies + builds); whole-array H2D alone 62.41 ms = 17.2 GB/s, D2H 54.85 ms
|
||||
device memory during the chunked build: 1741 MiB used (context + cache 256 + dataset 1024 + chunk 64 MiB); items after the chunked rebuild: 1024 of 1024 equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 265.98 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.43 ms -> 63.086 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.82 host_cache_ms=382.9 leaves_ms=5.46 build_ms=31.98 mem_build_mib=2699 chunk_mib=64 chunked_ms=75.80 mem_chunked_mib=1741 chunked_equal=1024 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.086 mhs_window=0.000 watts=0.0 sm_mhz=0 mem_mhz=0 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||
24
proto-newpow/state-dataset/results/gpu/sd1.log
Normal file
|
|
@ -0,0 +1,24 @@
|
|||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 383.9 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.56 ms first, 5.49 ms second -> 195.4 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.05 ms first, 31.98 ms second -> 524.7 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
dump: 4 warps (bases 0, 32, ...) written to dump_sd1.txt
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.35 ms -> 63.088 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 207.9 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.303 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=383.9 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.088 mhs_window=63.083 watts=207.9 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||
23
proto-newpow/state-dataset/results/gpu/sd12.log
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
state-dataset bench mode sd1 pack "igneum-genesis" (test harness: no pool, no network, no wallet)
|
||||
GPU: NVIDIA GeForce RTX 4090 (128 SMs, cc 8.9, 24083 MiB), CUDA driver 13.2 runtime 12.8
|
||||
hash kernel: 29 registers/thread, 24 resident blocks/SM at 1 warp/block
|
||||
sd1 leaf stand-in: S[i] = K[i] ^ 0x5a5a5a5a -> S = 6a3d3bc5 667ccb2c defa6159 a29c68ce a5cd2601 bc57b564 39395b1b e2a19102; leaf(t) = ChaCha12 block of (sigma, S, t, 0, 49676e65, 53746174)
|
||||
device memory at start: 395 MiB used of 24083 MiB (context)
|
||||
cache fill (GPU): 1.87 ms first, 1.85 ms second (65536 segments x 64 ChaCha12 blocks, 256 MiB)
|
||||
cache fill (host, one thread): 382.6 ms; cache check: PASS (GPU == host PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e)
|
||||
leaf array (GPU, 16777216 x 64 B = 1024 MiB): 5.55 ms first, 5.49 ms second -> 195.4 GB/s written
|
||||
leaf check: PASS (64 leaves incl. 0 and 16777215 vs host mh_leaf)
|
||||
dataset build (igneum_build_sd, leaf XOR before the first mixer): 32.07 ms first, 31.98 ms second -> 524.6 M items/s
|
||||
device memory after the build: 2699 MiB used (context 395 + cache 256 + leaves 1024 + dataset 1024 MiB)
|
||||
dataset[0..3] = 29d7b07a edc07cd7 fd248836 d1e4d9cf (sd1: new values, no Mac expectation)
|
||||
dataset self-test: PASS (64 random words vs host mh_word_sd: PASS)
|
||||
item bit-exactness (in-process, host mh_item_sd with host mh_leaf on the host cache): 1024 of 1024 items equal
|
||||
sd1 warp base 0 lane 0: b600edbed969becc (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 4096 lane 0: a533e89c78bb6b74 (new value; checked by verify_sd.c through the dump)
|
||||
sd1 warp base 1000000 lane 0: beb4cb0c6bab8163 (new value; checked by verify_sd.c through the dump)
|
||||
warm-up batch: 2^24 hashes in 267.38 ms wall; fingerprint (FNV-1a 64 over the outputs at base 0): d5b0c16390cad0e8 (sd1, new value)
|
||||
timed: 10 batches x 2^24 hashes, GPU 2659.36 ms -> 63.087 MH/s (32.30 GB/s useful, loads x 4 B)
|
||||
power window: 76 batches in 20.2 s -> 63.083 MH/s (events, incl. sync gaps); nvidia-smi 22 samples, 10 after 10 s: mean 206.5 W, SM 2745 MHz, mem 10251 MHz
|
||||
-> 0.305 MH/s per W (window rate / mean W)
|
||||
device memory while hashing: 1803 MiB used (leaves freed)
|
||||
RESULT mode=sd1 gpu=NVIDIA GeForce RTX 4090 cache_fill_ms=1.85 host_cache_ms=382.6 leaves_ms=5.49 build_ms=31.98 mem_build_mib=2699 items_equal=1024/1024 fingerprint=d5b0c16390cad0e8 mhs=63.087 mhs_window=63.083 watts=206.5 sm_mhz=2745 mem_mhz=10251 cache=PASS dataset=PASS vectors=n/a overall=PASS
|
||||