Ledger close round 2: M16 the inline-cache bench beside the CUDA worker (gen.py, inline_bench.cpp, the Mac emulation check PASS, the PC 2 kit and playbook), E17 the draw-line sampler, P17 the conformance driver, its run log, the bench-log entries, the design 8.2 row, the three ledger entries
This commit is contained in:
parent
a4b3a5d980
commit
bb5bbccf08
14 changed files with 1477 additions and 4 deletions
|
|
@ -1717,3 +1717,42 @@ Owed: the tampered pack against the real worker on PC 2's RTX 5090 (the job tool
|
|||
|
||||
**Not run.** The X20 fast-time simnet timing (a simnet chain is a few hundred blocks; the unit test's step count is the evidence, the 10^6-block chain is still the owed experiment) and the X19 two-node skew run (needs an `attack-switches` build of the node).
|
||||
|
||||
|
||||
## 6 October 2026 (night), ledger close round 2: M16 the inline-cache kernel on the RTX 5090 (the on-die SRAM emulation) and E17 the draw line per setting
|
||||
|
||||
Owner: the ledger-pc2 agent, docs worktree `igneum-wt-ledger-pc2` (branch `ledger-pc2` from `fud-close`), fork worktree `vendor/igneum-node-ledger2` (branch `ledger-fixes-2` from `ledger-fixes` bd1b676a). The experiment the ledger names for M16: the honest 1 GiB-dataset kernel against the same program with every dataset read replaced by its derivation from the cache (spec 01 section 1.8: 8 dependent cache-line reads and 9 mixer applications per item), at the 256 MiB cache and at a 64 MiB cache-line mask that fits inside the RTX 5090's 96 MiB L2. The 64 MiB setting is NOT the construction; it is the recompute attacker's device modelled on the honest card (the cache in SRAM-class memory, the card's own integer engine on the mixer), the "50 T op/s" row of `docs/analysis/m16-recompute-attacker-2026-10-05.md` section 3 measured instead of assumed.
|
||||
|
||||
**The kernel (Mac, build lock, `proto-cuda/inline-bench/`).** `gen.py` derives `kernel_inline.cu` and `memhard_inline.h` from the pack's `kernel_bound.cu` and `memhard.h` by counted text substitution (the 16 `ds[x & mask]` loads of the devnet-v4 epoch-0 pack become `mhi_word(cache, (x) & mask, lineMask)`; the cache-line mask becomes a kernel argument; refused on any other count). `inline_bench.cpp` is a benchmark beside the shipped worker, never inside it: the same run-time driver API and NVRTC loading as `nvrtc/worker.cpp`, three programs compiled from the pack's texts, three bit-exact checks before any timing (the run is void if one fails), then fixed-time windows per setting with an `nvidia-smi` sampler thread for E17. `check-mac.sh` builds it against `emu_backend.cpp` (the two libraries as host functions, the real kernel texts compiled by clang and run on host threads through `proto-cuda/emu`, as the worker's own emulation test does).
|
||||
|
||||
| Check (Mac, host threads, pack `igneum-devnet-v4-epoch0`, 12.3 s wall) | Result |
|
||||
|---|---|
|
||||
| 1. the pack's self-test through the honest kernel (cache head, last line, FNV-1a 64 `448274a57f508cbc`; dataset head, word [268435455], 64 samples; 96 vector lanes) | PASS |
|
||||
| 2. the inline kernel at the 256 MiB mask (0x003fffff) on the 96 vector lanes against the pack's vectors | PASS, every lane equal; the dataset was never read |
|
||||
| 3. inline64 (mask 0x000fffff): a 64 MiB dataset built by `igneum_build_inline` at that mask and read by the honest kernel, against `igneum_hash_inline` at the same mask, 8,288 lanes | PASS, stored twin == recomputed; differs from the pack's vectors, as a smaller cache must |
|
||||
| 3. inline32 (mask 0x0007ffff), the same | PASS, 8,288 lanes |
|
||||
|
||||
Texts (sha256, first 16): kernel.cu `0a75576f74ddb9a2`, kernel_bound.cu `2a076d717d481df5`, program.h `4420cc0f103b7332`, memhard.h `88dfa13835df0cce`, kernel_inline.cu `5aaabaa0b01e4fea` (16 inline loads), memhard_inline.h `dd86cd2610e5540e`. The Windows exe: `build-windows.sh` (mingw, static, KERNEL32 and the Universal CRT only), 355,840 bytes, sha256 `2e6a21de6d9e752e5e22b325d4d6de4fe2b5c88d88f01204715aa3758d22ef87`. The PC kit `igneum-inline-bench-kit.zip` (the exe, the pack with the inline texts and the M28 stamp, the playbook `job-pc2.ps1`): 202,138 bytes, sha256 `80ce0290a456b5c221ffde02bb35f74a25e48394e4cbccf53335a83e8aeb08ce`, on the downloads host, verified live. The PC 2 job is queued behind the 0.3.11 rollout (the scheduler's order); its table follows below when the RESULT lines are in.
|
||||
|
||||
## 6 October 2026 (night), ledger close round 2: P17 the four-state word in the RPC, unit tests and the conformance run on a fast-time 3-node network
|
||||
|
||||
Owner: the ledger-pc2 agent. Fork branch `ledger-fixes-2` (worktree `vendor/igneum-node-ledger2`, from `ledger-fixes` bd1b676a, itself from `release-0.3.6`), commit b1e98b79: `igneum_getTransactionStatus` returns one `state` word (design 2.4: `included`, `executed`, `proven`, `finalised`, and `finality not active` in place of finalised while `finality_active` is false; `pending` for the mempool, `unknown` for nothing held), a `failure` field for the three paths (`skipped`, `reorged out` with the height, `finality paused`), the `finalized` and `safe` block tags bound to the latest locked checkpoint's chain block with the fallback of spec 3.9 flagged in `finalizedSource` (the finality depth of spec 02 section 2.1 below the tip, genesis while the chain is younger; never the tip), a new `igneum_getFinalityView`, the chain follower reading the node's finality report after every pass, and a bounded memory of unwound transactions. Design 8.2's RPC table updated on `ledger-pc2`. Every number here is a count, a pass or a second; nothing is a rate.
|
||||
|
||||
**Suite (Mac, build lock, debug profile, `cargo test -j 4 -p igneum-exec`, target dir seeded from the round-1 worktree; a Mac run because PC 2 was queued behind the 0.3.11 rollout for over 30 minutes).** First attempt: 4 compile errors, `RpcErr` without `Debug` (the same class as round 1's first attempt; derived). Second: 15 of 16, the fallback test's own expectation wrong at a tip younger than the depth (the label; fixed in `finalized_height`). Third: 16 passed, 0 failed (13 of them the suite's, 3 new: `one_word_per_transaction_and_the_failure_paths`, `finalized_tag_resolves_to_the_lock_or_the_documented_fallback`, `a_reorg_marks_the_transaction_reorged_out_until_it_is_included_again`). Release build of `igneumd` and `igneum-miner` for the network: 2 min 25 s.
|
||||
|
||||
**Conformance run (Mac, run lock, `tools/p17-conformance/run.mjs`, ports 29950 to 29973, network `igneum-devnet-1950`, the 60x fast-time profile of this tree with `skip_proof_of_work` true: checkpoints every 30 DAA s, `min_daa` 120, presence window 1, finality depth 720; three nodes, n0 and n2 dialling n1 through 50 ms proxies; three voters at 1/3 each; log `tools/p17-conformance/run-2026-10-06.log`).** Three starts: the first refused the main checkout's fast-time file (the dead `timestamp_deviation_tolerance` field; this tree's copy is clean and is now the one read), the second collided with the first's nodes still exiting (a port gate added), the third ran to the end: PASSED in 437.5 s.
|
||||
|
||||
| Observation | What the RPC said | When (s from start) |
|
||||
|---|---|---|
|
||||
| Fallback before the first lock | tip 8, `finalized` tag 0, source "genesis: the chain is younger than the finality depth", `latest` 8 | 12.7 |
|
||||
| First lock | index 5 at chain block 129; `finalized` tag 129 against tip 146, source "locked checkpoint" | 168.9 |
|
||||
| tx1, the happy path (n0) | pending at 168.9; executed in chain block 147 at 172.2; a routine 1-block selected-chain reorg at 172.4 (`unknown`, failure `reorged out`, from 147); executed again in chain block 149 at 174.4 (node log 00:25:56.5; the watcher logs the first time of each word, so this repeat is in the node log, not the step line); finalised at 198.3 under checkpoint 6 (locked 00:26:20.1); `finalized` tag 153 while `latest` was 173 | 168.9 to 198.3 |
|
||||
| Skipped: one nonce, two copies, two nodes in one instant | copy A executed in chain block 176 at 200.1, unwound by a 1-block reorg at 200.3, then `included` with failure `skipped` (NonceTooLow) at 201.1; copy B executed in chain block 176 | 198.3 to 204.2 |
|
||||
| Reorged out: n2 alone behind a cut link | tx3 sent only to n2: executed in n2's chain block 183 at 238.8; at 258.5 n2 reports finality paused (its lock is one index behind, presence window 1); link healed at 238.8, the heavier 2/3 chain won at 303.7: n2 has 11 reorg log lines, tx3 reports `unknown` with failure `reorged out` from 183 | 238.8 to 363.8 |
|
||||
| Not re-included after the reorg | tx3 stayed `unknown` for 60 s after the heal: the EVM pool has `on_chain_block` only and no reorg hook, so an unwound transaction leaves the node's view unless the sender sends it again (finding; the word the RPC gives is the design's) | 303.7 to 363.8 |
|
||||
| Finality paused: 2/3 of the weight stops signing (v1 and v2 restarted with `--no-vote`) | n0's flag false with reason `paused` at 399.0; tx1 flips to `finality not active`, lockedCovered true; tx4 (sent at 399.0) executed in chain block 342 at 402.4 with failure `finality paused`; `finalized` tag held at 294 (the last lock) against tip 339 | 399.0 to 402.9 |
|
||||
| Resumed (the two voters sign again) | the flag true at 413.2 (lock at 320), tx4's failure cleared, finalised at 437.3 under the lock at 348; tx1 back to `finalised` | 413.0 to 437.5 |
|
||||
| Proven | not exercised: no shard is proved on this network (the SP1 CPU prove takes minutes on a loaded Mac and the proving loop is not in the harness); the word is covered by the unit test on the paid-shard rule | |
|
||||
|
||||
Two artefacts of the fast profile worth one sentence each. The flag flickers to `paused` for 0.6 s at every checkpoint determination (00:26:19.5 determined, 00:26:20.1 locked: with presence window 1 the test `latest lock + window >= next index` fails for the gap), which showed on tx1 at 197.7 as a one-pass `finality paused`; at mainnet's window of 240 the gap is invisible. The 1-block selected-chain reorg is routine at 1 block/s with 50 ms links (20 on n0 in 7 minutes), so `reorged out` is a transient of about 2 s whenever the unwound block is merged again by the next chain block; a wallet treats it as final only after the merge depth, as design 2.3 says.
|
||||
|
||||
Consequences. For a wallet or exchange: the one word is now in the RPC and `finalized` never names the tip; the pause shows as "finality not active" with the certificate still binding. For the pool: a reorged-out transaction is dropped from the node's view and must be resent (the pool owns no reorg hook; an item for the execution engineer, not changed here). Owed: the `proven` transition on a network with the proving loop; the phone app and the explorer still show the three words of phone-app 3 and 9 and need the `state` field (ledger X-side, not this round).
|
||||
|
|
|
|||
|
|
@ -363,7 +363,8 @@ Standard namespace, unchanged semantics where the table says so; RPC "blocks" ar
|
|||
| `eth_getBlockReceipts`, `eth_getUncleBy*` | Supported; uncles always empty | Merged blocks are not uncles; see `igneum_*` |
|
||||
| `igneum_getDagBlock(hash)` | New | Header, parents, blue or red, merging chain block, transactions with executed or skipped and the skip reason |
|
||||
| `igneum_getSegment(number)` | New | Mergeset in order, executed set, shard plan, proof record if any |
|
||||
| `igneum_getTransactionStatus(hash)` | New | `{included_in: [hashes], executing_copy, executed, proven, locked, skip_reason}` |
|
||||
| `igneum_getTransactionStatus(hash)` | New; the one-word `state` of 2.4 landed on the fork branch `ledger-fixes-2` (6 October 2026, ledger P17) | `{state, failure, finality, finalityActive, finalityReason, lockedCovered, finalizedHeight, finalizedSource, chainBlockNumber, shard, reorgedFrom, includedIn: [...], executingCopy, executed, proven, locked, inMempool}`. `state` is exactly one of `included`, `executed`, `proven`, `finalised`, or `finality not active` in place of finalised while `finality_active` is false (`pending` for a mempool-only transaction, `unknown` when no block and no mempool holds it). `failure` names the path: `skipped` (every copy skipped by rule), `reorged out` (unwound by a selected-chain reorg, `reorgedFrom` the height), `finality paused` (executed, no lock covers it, finality not active). `proven` = the shard holding the executed copy has a paid proof record carried by a chain block |
|
||||
| `igneum_getFinalityView()` | New (ledger P17) | `{finalityActive, finalityReason, latestLockedIndex, latestLockedHash, latestLockedHeight, finalizedHeight, finalizedSource, finalityDepthDaa, tip}`: what the `finalized` and `safe` block tags resolve to. Rule: the latest locked checkpoint's chain block when the executor holds it (a certificate stays binding through a pause); else the fallback of spec 3.9 for `finality_active` false, the highest chain block at least the consensus finality depth (spec 02 section 2.1, 43,200 DAA s on mainnet) below the tip, which is genesis while the chain is younger than that depth; never the tip. `finalizedSource` says which |
|
||||
| `igneum_getProvingFee`, `igneum_getBudgets` | New | `f_p`, `B_e`, `B_p`, backlog depth |
|
||||
|
||||
### 8.3 Hardhat and Foundry
|
||||
|
|
|
|||
|
|
@ -1179,7 +1179,7 @@ Evidence: the files and lines above. Fix: review's first of five. Review id R3.2
|
|||
### M16. The 256 MiB cache fits on a die, so the recompute attacker is compute bound
|
||||
"Your 4.8x-slower shortcut ran with the cache in DRAM behind a chip that cannot hold it. Put 256 MiB of SRAM on a die and the dataset is never needed: 128 items per hash at about 1,170 integer operations and 8 near-free reads each. That is 150,000 operations per hash, and integer operations per dollar is where silicon beats a GPU."
|
||||
|
||||
Status: Open, cost model written, the on-die emulation still a PC job (5 October 2026, evening sweep). Was: Open, experiment scheduled. Sweep (5 October 2026): not runnable on this Mac; the on-die emulation needs the RTX 5090 (an inline kernel at a 64 MiB cache inside its 96 MiB L2). The only newer number is the loaded-Mac reconfirmation of 3 October (inline 17x slower at a 256 MiB cache, bench-log "R3.26 / M15", M16 note), which does not price a die. Hardware item.
|
||||
Status: Open, the kernel written and bit-exact on the Mac, the PC 2 run queued behind the 0.3.11 rollout (6 October 2026, night, ledger close round 2; the table lands in this entry and in the bench-log "ledger close round 2: M16" when the job's RESULT lines are in). Was: Open, cost model written, the on-die emulation still a PC job (5 October 2026, evening sweep). Was: Open, experiment scheduled. Sweep (5 October 2026): not runnable on this Mac; the on-die emulation needs the RTX 5090 (an inline kernel at a 64 MiB cache inside its 96 MiB L2). The only newer number is the loaded-Mac reconfirmation of 3 October (inline 17x slower at a 256 MiB cache, bench-log "R3.26 / M15", M16 note), which does not price a die. Hardware item.
|
||||
|
||||
Sweep (5 October 2026, evening): `docs/analysis/m16-recompute-attacker-2026-10-05.md` prices the device from the specification and the measured rates. Per hash the attacker recomputes 128 items at 9 mixer applications of about 130 operations and 8 dependent 64-byte cache reads each: about 150,000 integer operations and 1,024 dependent SRAM reads (65 KB). To match one RTX 5090 at its measured 229 Mhash/s the chip needs 34 T integer op/s and 15 TB/s of SRAM bandwidth beside 256 MiB of SRAM (100 to 300 mm^2 on a current node, approximate). At the 5090's own integer budget (about 50 T op/s, approximate) that is 0.33 Ghash/s: 1.5x the measured closed-form rate, 2.4x the 141 Mhash/s projected for version 2 programs, before any fixed-function factor; with a 3x factor (approximate) 3x to 6x at equal die area. The lever: the mixer cost is paid by the honest miner once a day (13.4 ms per 1 GiB on the 5090, measured) and by the attacker per hash, so doubling it halves the attacker's rate at zero honest cost, 4x puts the equal-silicon gain at 0.36x and the factored gain near 1x, bounded by the CPU verify gate (0.41 to 1.2 ms per warp today, 10 ms the gate, so about 8x of headroom on the M5 Max core). What is still unmeasured: the inline kernel on the 5090 with a 64 MiB cache inside its L2 (the SRAM emulation, a PC job), the time-memory curve of O-1.6, and any cryptanalysis of the mixer. Decision at gate 1 (owner the project lead): cache size and mixer cost against the verify gate.
|
||||
|
||||
|
|
@ -1187,6 +1187,8 @@ Answer: Correct as arithmetic, unmeasured as a device. From spec 1.8.4 and 1.8.5
|
|||
|
||||
Evidence: spec 1.8, 1.16; `proto-metal/MEMHARD.md` 2.2 (Apple only). Experiment: `--inline-dataset` on the RTX 5090 at a 64 MiB cache (inside its 96 MiB L2, the SRAM emulation) and at 256 MiB against the honest 1 GiB kernel; the O-1.6 time-memory curve; a CPU fill and verify time at a 1 GiB cache. Decision at gate 1: cache size "exceeds what one die can hold, and grows". Review ids R3.5 and the chip designer's pricing.
|
||||
|
||||
Round 2 (6 October 2026, night): the CUDA inline path did not exist (the shipped worker reads the dataset only; `--inline-dataset` was Metal), so it was written as a benchmark beside the worker, never inside it: `proto-cuda/inline-bench/` (`gen.py` derives `kernel_inline.cu` and `memhard_inline.h` from the pack by counted text substitution, the 16 dataset loads of the devnet-v4 epoch-0 program becoming `mhi_word(cache, x & mask, lineMask)` with the cache-line mask a kernel argument; `inline_bench.cpp` loads the driver API and NVRTC at run time as the worker does, compiles the pack's texts, runs three bit-exact checks and then fixed-time windows for honest 1 GiB, inline at 256 MiB, inline at 64 MiB (the first quarter of the cache, inside the 5090's 96 MiB L2: the on-die SRAM emulation) and inline at 32 MiB, with an `nvidia-smi` sampler per window for E17). Mac check (host threads through the project's CUDA emulation shim, build lock, 12.3 s): check 1 the pack's self-test PASS (96 of 96 lanes); check 2 the inline kernel at the 256 MiB mask equals the pack's vectors on all 96 lanes, the dataset never read (this is what proves the substitution and the derivation); check 3 at the 64 MiB and 32 MiB masks a stored dataset built at that mask and read by the honest kernel equals the recomputed path on 8,288 lanes each, and differs from the pack's vectors as a smaller cache must. The Windows exe (mingw, 355,840 bytes, sha256 2e6a21de...) and the pack with the inline texts went to the downloads host as `igneum-inline-bench-kit.zip` (202,138 bytes, sha256 80ce0290...); the PC 2 job (one `run` job, the app's miners paused and the live prover off for about 4 minutes on the card, both restored, two passes at 1 and 8 warps per block, 15 s per setting) waits on the scheduler's go behind the 0.3.11 rollout. What the number will test: the analysis's "50 T op/s" row, 1.5x the measured 229 Mhash/s and 2.4x the projected 141 at equal integer budget; the inline64 rate IS the attacker's rate on this silicon with the cache in L2, so the gain is inline64 over honest, before any fixed-function factor. The CPU fill and verify at a 1 GiB cache and the O-1.6 curve are not part of this job.
|
||||
|
||||
### M17. Every ahead-of-time miner stops at the epoch boundary; whoever compiles in process mines alone
|
||||
"Your log: last block of epoch 0 at 21:12:14, nothing for the next 91 seconds while the Windows launcher killed eight identities, re-exported, rebuilt two binaries and restarted them. The Metal worker compiles in 129 ms. The first 2.5% of every hour goes to whoever does not use your launcher."
|
||||
|
||||
|
|
@ -1492,7 +1494,7 @@ Evidence: P1; execution-layer 9.1 R2 and R4; `docs/bench-log.md` (no shard has b
|
|||
### P17. Interfaces must show four states, and the design shows three
|
||||
"Included, executed, proven, finalised are four different facts. Your status call returns executed, proven, locked and a list of blocks. A wallet that shows a balance at 'included' is lying by omission."
|
||||
|
||||
Status: Answered with evidence for what the RPC returns today (5 October 2026, night, ledger close round 1: report only, no code change); the four-state word in the response and the conformance run (O-7.2) are round 2. Was: Open, rule written (3 October 2026, night) in `docs/design/execution-layer.md` 2.4; experiment scheduled (O-7.2). Sweep (5 October 2026): the conformance set is devnet work; not run.
|
||||
Status: Fixed on a branch, pending merge (6 October 2026, night, ledger close round 2): fork `ledger-fixes-2` b1e98b79, suite igneum-exec 16 of 16 on the Mac (labelled a Mac run), the O-7.2 conformance run PASSED on a fast-time 3-node network (bench-log "ledger close round 2: P17"). Was: Answered with evidence for what the RPC returns today (5 October 2026, night, ledger close round 1: report only, no code change); the four-state word in the response and the conformance run (O-7.2) are round 2. Was: Open, rule written (3 October 2026, night) in `docs/design/execution-layer.md` 2.4; experiment scheduled (O-7.2). Sweep (5 October 2026): the conformance set is devnet work; not run.
|
||||
|
||||
Answer: Correct. Design 2.3 defines executed, proven and locked; `igneum_getTransactionStatus` carries `included_in` as a list and never as a state; the phone app (phone-app 3 and 9) shows three words. A transaction in a block not yet on a selected chain, or in a merged block waiting its turn in the sequence, is included and nothing more, and on a DAG that gap is routine. The rule now in 2.4: every RPC, wallet, explorer and app reports exactly one of included, executed, proven, finalised (the user-facing word for locked), never a stronger word than the chain's own state, and shows "finality not active" in place of finalised while `finality_active` is false (3.9). Proven says nothing about availability: full blocks are what every full node holds (F13) and a light client takes data from nodes under spec 10.1. Test: a wallet and RPC conformance set on the devnet, one transaction through each state and the three failure paths (skipped, reorged out, finality paused).
|
||||
|
||||
|
|
@ -1500,6 +1502,8 @@ Evidence: execution-layer 2.3, 2.4, 8.2; phone-app 3 and 9; spec 3.9, 10.1. Expe
|
|||
|
||||
Run (5 October 2026, night, report only, no bench-log entry because nothing ran): on the running fork line (`vendor/igneum-node-036`, release-0.3.6 at `a24ab01a`, the binary node 1 runs) `igneum_getTransactionStatus` (`igneum/exec/src/rpc.rs:739-751`) returns one object: `includedIn`, a list with one entry per block that carries the transaction (block hash, `chainBlockNumber`, `executed`, `skipReason`); `executingCopy`, the block whose copy executed, or null; `executed`, true once the hash is in the executor's index; `proven: false` and `locked: false`, both constants in the source, never true on this line whatever proof records or certificates the chain holds; and `inMempool`. The states a caller can tell apart are therefore: in the mempool only (`inMempool` true, `includedIn` empty); included and not executed (`includedIn` non-empty, `executed` false: a block off the selected chain, or a merged block waiting its turn in the sequence); executed (`executed` true, `executingCopy` set); skipped (an `includedIn` entry carrying `skipReason`, `executed` false). Proven and finalised cannot be read from this call at all. The block tags of the EVM calls resolve in `resolve_block` (`rpc.rs:251-262`): the arm at line 257 maps `latest`, `pending`, `safe` and `finalized` all to the tip, and the comment at lines 226 to 228 says no certified checkpoint exists yet and the RPC must not pretend otherwise. That comment predates the first live lock (bench-log 4 October 2026, checkpoint 242), so today `eth_getBlockByNumber("finalized")` returns the tip of a chain that holds locked checkpoints below it: the tag overstates. Round 2 (code): the one-word state (included, executed, proven, finalised, or `finality not active`) in the response, `finalized` bound to the last locked checkpoint's chain block, then the O-7.2 conformance run.
|
||||
|
||||
Round 2 (6 October 2026, night): implemented on the fork branch `ledger-fixes-2` (b1e98b79, from `ledger-fixes` bd1b676a). `igneum_getTransactionStatus` now carries one `state` word, exactly one of `included`, `executed`, `proven`, `finalised`, with `finality not active` in place of finalised while `finality_active` is false (design 2.4; `pending` for a mempool-only transaction, `unknown` when no block and no mempool holds it), and a `failure` field naming the path: `skipped` (every copy skipped by rule), `reorged out` (unwound by a selected-chain reorg, `reorgedFrom` the height; a bounded memory of 10,000 unwound hashes, cleared on re-inclusion), `finality paused` (executed, no lock covers it, the flag false). `proven` is true when the shard holding the executed copy has a paid proof record carried by a chain block (the first true value this call has ever returned; it was a constant). The `finalized` and `safe` block tags resolve through one rule (`finalized_height`): the latest locked checkpoint's chain block when the executor holds it (a certificate stays binding through a pause, spec 3.9), else the fallback spec 3.9 gives for `finality_active` false, the highest chain block at least the consensus finality depth (spec 02 section 2.1, 43,200 DAA s on mainnet, 720 blocks on the fast-time profile) below the tip, genesis while the chain is younger; never the tip, and `finalizedSource` says which. A new `igneum_getFinalityView` reports the flag, its reason, the latest lock and what the tag resolves to; the chain follower reads the node's finality report after every pass, so the RPC never touches consensus. Unit tests (3 new, 16 of 16 in the crate on the Mac under the build lock, labelled a Mac run because PC 2 was queued behind the 0.3.11 rollout). Conformance (O-7.2, `tools/p17-conformance/run.mjs`, 3 nodes at 1 block/s behind 50 ms proxies, three voters, 437.5 s, PASSED): before the first lock the tag resolved to genesis by the depth rule while `latest` was 8; the first lock at 168.9 s (index 5, chain block 129) moved the tag to 129 against tip 146; one transfer went pending, executed (147), through a routine 1-block reorg (`reorged out` for about 2 s, then executed again in 149) to finalised at 198.3 s with the tag at 153 under a tip of 173; two copies of one nonce sent to two nodes in one instant gave one `executed` and one `included` with failure `skipped` (NonceTooLow); a node cut off alone executed a transfer in its own chain block 183, and when the link healed and the heavier 2/3 chain won (11 reorg lines) the transfer reported `unknown` with failure `reorged out` from 183; 2/3 of the weight silenced paused finality at 399 s (reason `paused`), the finalised transfer flipped to `finality not active` with `lockedCovered` true, a fresh transfer executed with failure `finality paused` and the tag held at the last lock (294) against tip 339, and both reached `finalised` within 25 s of the voters signing again. Not exercised: `proven` on the network (no shard is proved on it; the unit test covers the paid-shard rule). Findings beside the fix: (1) an unwound transaction does not return to the EVM mempool (the pool has `on_chain_block` and no reorg hook), so after a reorg it must be resent; the RPC says `reorged out` and a wallet knows to resend, but the pool half is the execution engineer's (not changed here); (2) with the fast profile's presence window of 1 the flag flickers to `paused` for 0.6 s at every checkpoint determination (240 on mainnet hides it); (3) the phone app and the explorer still show three words and need the `state` field (phone-app 3 and 9; a text-and-app item for the next round). Design 8.2's RPC row updated on `ledger-pc2`.
|
||||
|
||||
### E12. Selfish operators under a price shock
|
||||
"Outside proving pays ten times more and IGN halves in a week. Every rational operator leaves hashing for jobs. Do your internal proofs go unproven, do queues grow, does anything bring them back? Assume nobody runs your client's scheduler."
|
||||
|
||||
|
|
@ -2039,12 +2043,14 @@ Evidence: session scratchpad `rt/logs/exec_b/miner_node1.log` (the template erro
|
|||
### E17. Unlogged inputs behind the economics, minor
|
||||
"The cap's 110 MH/s and its draw are not in the bench-log; the Mac's draw is not logged; the economy sim's one measured input is a 229 MH/s card against today's 124; mining-versus-pool flips from 4.9x for mining on today's devnet to 930x for proving at 10,000 cards and no document says it depends on fleet size; the app-share text omits '100,000-gas calls' and the open base unit; emission ran at up to 2x schedule; shard-scale prover cost is unmeasured on any GPU; the iGPU default off is right."
|
||||
|
||||
Status: Open, minor (the draw lines need the machines); text half stated (5 October 2026, night): `site/litepaper.html`, Building, Why build here, "a million 100,000-gas calls a day at a 1 gwei tip pays about 7,300 IGN a year, with 1 gwei taken as a billionth of an IGN (the base unit is Open)"; For miners, "Which of the two in-chain streams pays more per GPU-second depends on the size of the fleet" with the 4.9x and 930x arithmetic marked approximate. Was: Open, minor; the simulator half re-run (5 October 2026 sweep). `sim/economy/sim.py` with the 5090 class at 124 MH/s (the devnet's measured rate) in place of 229, scenarios a and b, 2 seeds, 111 s: baseline profit USD 5.59 per 5090 card-day (6.37 at 229), 3090 2.73 (1.66), 3060 1.34 (0.78), small 0.74 (0.39); under the b shock 4.52 / 2.25 / 1.46 / 0.30; no backlog, every block proven within 60 s, hash trough 75% of pre-event under b (82% at 229), 5% of cards off (10%), oscillation flag in 1 of 2 seeds. The direction is as the review said: a slower flagship card shifts income toward the smaller classes and makes the fleet more sensitive to the price shock. The draw lines (`nvidia-smi` on both PCs, `powermetrics` on the Mac) need the machines (person). Was: Open, minor (4 October 2026).
|
||||
Status: Open, minor: the PC 2 draw lines are in the queued M16 job (one `nvidia-smi` line per setting beside its rate, sampled every 2 s inside each timed window; 6 October 2026, night, ledger close round 2), the PC 1 line and the Mac's `powermetrics` line (sudo) are owed. Was: Open, minor (the draw lines need the machines); text half stated (5 October 2026, night): `site/litepaper.html`, Building, Why build here, "a million 100,000-gas calls a day at a 1 gwei tip pays about 7,300 IGN a year, with 1 gwei taken as a billionth of an IGN (the base unit is Open)"; For miners, "Which of the two in-chain streams pays more per GPU-second depends on the size of the fleet" with the 4.9x and 930x arithmetic marked approximate. Was: Open, minor; the simulator half re-run (5 October 2026 sweep). `sim/economy/sim.py` with the 5090 class at 124 MH/s (the devnet's measured rate) in place of 229, scenarios a and b, 2 seeds, 111 s: baseline profit USD 5.59 per 5090 card-day (6.37 at 229), 3090 2.73 (1.66), 3060 1.34 (0.78), small 0.74 (0.39); under the b shock 4.52 / 2.25 / 1.46 / 0.30; no backlog, every block proven within 60 s, hash trough 75% of pre-event under b (82% at 229), 5% of cards off (10%), oscillation flag in 1 of 2 seeds. The direction is as the review said: a slower flagship card shifts income toward the smaller classes and makes the fleet more sensitive to the price shock. The draw lines (`nvidia-smi` on both PCs, `powermetrics` on the Mac) need the machines (person). Was: Open, minor (4 October 2026).
|
||||
|
||||
Answer: Correct on each point; `docs/review/round-4-2026-10-04.md` section 7 holds the arithmetic. Fix: one `nvidia-smi` draw line per setting with the rate beside it on both PCs; one `powermetrics` line on the Mac; the sim rerun at 124 MH/s with the comparison stated as a function of fleet size; "a million 100,000-gas calls a day" in the litepaper. Review ids R4.7.7 to R4.7.13.
|
||||
|
||||
Evidence: the files above. Experiment: the three draw lines.
|
||||
|
||||
Round 2 (6 October 2026, night): the draw line rides inside the M16 job on PC 2 (`proto-cuda/inline-bench/inline_bench.cpp`, a sampler thread running `nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu --format=csv,noheader` every 2 s during each timed window, every sample printed with the setting's name, the median of the first field in the row beside the hash rate; the playbook pauses the app's miners and switches the prover off first and prints the compute apps left on the card, so a shared-card run is labelled as such). The lines land here and in the bench-log when the job's RESULT lines are in. PC 1 is off limits tonight (the 0.3.10 and 0.3.11 rollout and its app outage), so its line is owed; the Mac's `powermetrics` line needs sudo and is owed to a person.
|
||||
|
||||
### E18. The dev fee is a protocol fee with better PR
|
||||
"A 1% dev fee in the official miner is 1% of the chain's hashrate paid to one company for as long as miners run it. Call it what it is: a founder allocation, hidden in the client, and nobody can tell how often it really fires."
|
||||
|
||||
|
|
|
|||
2
proto-cuda/inline-bench/.gitignore
vendored
Normal file
2
proto-cuda/inline-bench/.gitignore
vendored
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
build-mac/
|
||||
build-win/
|
||||
21
proto-cuda/inline-bench/build-windows.sh
Executable file
21
proto-cuda/inline-bench/build-windows.sh
Executable file
|
|
@ -0,0 +1,21 @@
|
|||
#!/usr/bin/env bash
|
||||
# Cross-compiles igneum-inline-bench.exe on the Mac with Homebrew mingw-w64, as nvrtc/build-windows.sh does for the
|
||||
# shipped worker: the driver API and NVRTC are loaded at run time (cuda_api.h), no import library; static
|
||||
# libgcc/libstdc++/winpthread, so the exe needs only KERNEL32 and the Universal CRT. The NVRTC DLLs are copied next
|
||||
# to it on the PC from the installed Igneum Miner (the run job does that). Needs nvrtc/fetch-redist.sh's headers
|
||||
# (IGNEUM_CUDA_INC overrides the include dir). Under the build lock: tools/lock/with-lock.sh build inline-bench/build-windows.sh
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
CUDA="$(cd "$HERE/.." && pwd)"
|
||||
INC="${IGNEUM_CUDA_INC:-$CUDA/nvrtc/redist/include}"
|
||||
[ -f "$INC/cuda.h" ] && [ -f "$INC/nvrtc.h" ] || { echo "no cuda.h / nvrtc.h under $INC" >&2; exit 1; }
|
||||
CXX=x86_64-w64-mingw32-g++
|
||||
command -v "$CXX" >/dev/null || { echo "$CXX not found (brew install mingw-w64)" >&2; exit 1; }
|
||||
OUT="${1:-$HERE/build-win}"
|
||||
mkdir -p "$OUT"
|
||||
"$CXX" -std=c++17 -O2 -Wall -Wextra -static -I "$INC" -I "$HERE" -o "$OUT/igneum-inline-bench.exe" "$HERE/inline_bench.cpp"
|
||||
x86_64-w64-mingw32-strip "$OUT/igneum-inline-bench.exe"
|
||||
printf '%s: %d bytes, imports:' "$OUT/igneum-inline-bench.exe" "$(stat -f %z "$OUT/igneum-inline-bench.exe")"
|
||||
x86_64-w64-mingw32-objdump -p "$OUT/igneum-inline-bench.exe" | sed -n 's/^[[:space:]]*DLL Name: //p' | tr '\n' ' '
|
||||
echo
|
||||
shasum -a 256 "$OUT/igneum-inline-bench.exe"
|
||||
39
proto-cuda/inline-bench/check-mac.sh
Executable file
39
proto-cuda/inline-bench/check-mac.sh
Executable file
|
|
@ -0,0 +1,39 @@
|
|||
#!/usr/bin/env bash
|
||||
# The Mac check of igneum-inline-bench (ledger M16): no GPU. gen.py derives the inline texts from the pack into a
|
||||
# scratch copy of the pack; the pack's kernel.cu and kernel_bound.cu and the generated kernel_inline.cu are compiled
|
||||
# by clang as host code through the proto-cuda/emu shim (launch syntax rewritten by sed, each in its namespace); the
|
||||
# bench is built with -DIGNEUM_EMU against emu_backend.cpp and run with --check, so the three bit-exact checks run
|
||||
# on the real kernel texts (the 1 GiB dataset is built on host threads; about a minute on an idle M5 Max). Rates
|
||||
# from this build are not numbers. Under the build lock: tools/lock/with-lock.sh build inline-bench/check-mac.sh.
|
||||
# Usage: check-mac.sh [pack dir] [out dir] (default pack proto-cuda/packs/igneum-devnet-v4-epoch0; needs nvrtc/fetch-redist.sh's headers:
|
||||
# IGNEUM_CUDA_INC overrides the include dir, default proto-cuda/nvrtc/redist/include)
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
CUDA="$(cd "$HERE/.." && pwd)"
|
||||
ROOT="$(cd "$CUDA/.." && pwd)"
|
||||
EMU="$CUDA/emu"
|
||||
INC="${IGNEUM_CUDA_INC:-$CUDA/nvrtc/redist/include}"
|
||||
PACK_SRC="${1:-$CUDA/packs/igneum-devnet-v4-epoch0}"
|
||||
OUT="${2:-$HERE/build-mac}"
|
||||
[ -f "$INC/cuda.h" ] && [ -f "$INC/nvrtc.h" ] || { echo "no cuda.h / nvrtc.h under $INC (nvrtc/fetch-redist.sh, or IGNEUM_CUDA_INC=...)" >&2; exit 1; }
|
||||
CXX="${CXX:-c++}"
|
||||
rm -rf "$OUT"; mkdir -p "$OUT/pack"
|
||||
cp "$PACK_SRC"/* "$OUT/pack/"
|
||||
python3 "$HERE/gen.py" "$OUT/pack" "$OUT/pack"
|
||||
python3 "$HERE/stamp-pack.py" "$OUT/pack"
|
||||
# the kernels as the emulation runs them: the launch syntax rewritten, each unit in its namespace
|
||||
emu_kernel() { # file stem, namespace
|
||||
{ echo '#include <cuda_runtime.h>'; echo '#include <cstdint>'; echo "namespace $2 {"
|
||||
sed -E 's/([A-Za-z_0-9]+)<<<([^,]+), ([^>]+)>>>\(/emu_launch(\1, \2, \3, /' "$OUT/pack/$1.cu"
|
||||
echo "}"; } > "$OUT/$1_emu.cpp"
|
||||
"$CXX" -std=c++17 -O2 -w -I "$EMU" -I "$OUT/pack" -c "$OUT/$1_emu.cpp" -o "$OUT/$1_emu.o"
|
||||
}
|
||||
emu_kernel kernel emu_pack
|
||||
emu_kernel kernel_bound emu_pack
|
||||
emu_kernel kernel_inline emu_inline
|
||||
"$CXX" -std=c++17 -O2 -Wall -Wextra -DIGNEUM_EMU -I "$INC" -I "$EMU" -I "$HERE" -c "$HERE/inline_bench.cpp" -o "$OUT/inline_bench.o"
|
||||
"$CXX" -std=c++17 -O2 -w -DIGNEUM_EMU -I "$INC" -I "$EMU" -c "$HERE/emu_backend.cpp" -o "$OUT/emu_backend.o"
|
||||
"$CXX" -std=c++17 -O2 -w -I "$EMU" -c "$EMU/shim.cpp" -o "$OUT/shim.o"
|
||||
"$CXX" -o "$OUT/igneum-inline-bench-emu" "$OUT/inline_bench.o" "$OUT/emu_backend.o" "$OUT/shim.o" "$OUT/kernel_emu.o" "$OUT/kernel_bound_emu.o" "$OUT/kernel_inline_emu.o" -pthread
|
||||
echo "built $OUT/igneum-inline-bench-emu (CPU emulation, not a GPU build)"
|
||||
exec "$OUT/igneum-inline-bench-emu" --pack "$OUT/pack" --check --check-mib "${CHECK_MIB:-64}" --settings honest,inline256,inline64,inline32 --smi ""
|
||||
150
proto-cuda/inline-bench/emu_backend.cpp
Normal file
150
proto-cuda/inline-bench/emu_backend.cpp
Normal file
|
|
@ -0,0 +1,150 @@
|
|||
// emu_backend.cpp: the CUDA driver API and NVRTC as host functions for igneum-inline-bench on the Mac (no NVIDIA
|
||||
// GPU). The same shape as proto-cuda/nvrtc/emu/emu_backend.cpp: the bench's own code path is unchanged, the two
|
||||
// library tables are filled with stand-ins, and the five kernels are the real texts compiled by clang and run on
|
||||
// host threads through proto-cuda/emu (32 threads per warp, a barrier inside __shfl_xor_sync). What it proves: the
|
||||
// three bit-exact checks of inline_bench.cpp on the real kernel texts. What it cannot prove: that NVRTC accepts the
|
||||
// texts and that the driver runs them; that is the RTX 5090 run. Not part of any deliverable.
|
||||
//
|
||||
// check-mac.sh compiles the pack's kernel.cu and kernel_bound.cu into namespace emu_pack and the generated
|
||||
// kernel_inline.cu into namespace emu_inline (launch syntax rewritten by sed, as emu/emu.sh does), then links them
|
||||
// with this file and inline_bench.cpp built with -DIGNEUM_EMU.
|
||||
#include <cstdint>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
#include <map>
|
||||
|
||||
#include "../nvrtc/cuda_api.h"
|
||||
#include "cuda_runtime.h" // the shim (proto-cuda/emu)
|
||||
|
||||
namespace emu_pack {
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
void igneum_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw);
|
||||
}
|
||||
namespace emu_inline {
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
void igneum_build_inline(uint32_t* ds, const uint32_t* cache, uint32_t nItems, uint32_t lineMask);
|
||||
void igneum_hash_inline(const uint32_t* cache, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw, uint32_t lineMask);
|
||||
}
|
||||
|
||||
// ---- NVRTC stand-in: records the source and its headers, hands back a tag the module loader maps to the compiled-in kernels
|
||||
struct EmuProg { std::string src, name, log; std::map<std::string, std::string> headers; std::vector<std::string> exprs; int which = 0; };
|
||||
static nvrtcResult e_version(int* major, int* minor) { *major = 12; *minor = 8; return NVRTC_SUCCESS; }
|
||||
static int ARCHS[] = { 75, 80, 86, 89, 90, 100, 120 };
|
||||
static nvrtcResult e_numArchs(int* n) { *n = (int)(sizeof(ARCHS) / sizeof(ARCHS[0])); return NVRTC_SUCCESS; }
|
||||
static nvrtcResult e_archs(int* out) { for (size_t i = 0; i < sizeof(ARCHS) / sizeof(ARCHS[0]); ++i) out[i] = ARCHS[i]; return NVRTC_SUCCESS; }
|
||||
static nvrtcResult e_create(nvrtcProgram* prog, const char* src, const char* name, int numHeaders, const char* const* headers, const char* const* includeNames) {
|
||||
EmuProg* p = new EmuProg(); p->src = src ? src : ""; p->name = name ? name : "";
|
||||
for (int i = 0; i < numHeaders; ++i) p->headers[includeNames[i]] = headers[i];
|
||||
*prog = (nvrtcProgram)p; return NVRTC_SUCCESS;
|
||||
}
|
||||
static nvrtcResult e_destroy(nvrtcProgram* prog) { delete (EmuProg*)*prog; *prog = nullptr; return NVRTC_SUCCESS; }
|
||||
static nvrtcResult e_compile(nvrtcProgram prog, int numOptions, const char* const* options) {
|
||||
EmuProg* p = (EmuProg*)prog;
|
||||
bool arch = false, std17 = false;
|
||||
for (int i = 0; i < numOptions; ++i) { std::string o = options[i]; if (o.rfind("--gpu-architecture=", 0) == 0) arch = true; if (o == "--std=c++17") std17 = true; }
|
||||
if (!arch || !std17) { p->log = "emu-nvrtc: expected --gpu-architecture=... and --std=c++17"; return NVRTC_ERROR_COMPILATION; }
|
||||
if (p->name == "kernel.cu") p->which = 1; else if (p->name == "kernel_bound.cu") p->which = 2; else if (p->name == "kernel_inline.cu") p->which = 3;
|
||||
else { p->log = "emu-nvrtc: unknown program " + p->name; return NVRTC_ERROR_COMPILATION; }
|
||||
// the texts handed over must be the ones check-mac.sh compiled: the same bytes are read from the same files, so
|
||||
// the sizes recorded here are what the bench prints; the hashes are printed by the bench itself
|
||||
std::printf("info emu-nvrtc: %s %zu bytes, headers %zu, compiled in as unit %d\n", p->name.c_str(), p->src.size(), p->headers.size(), p->which);
|
||||
return NVRTC_SUCCESS;
|
||||
}
|
||||
static nvrtcResult e_logSize(nvrtcProgram prog, size_t* n) { *n = ((EmuProg*)prog)->log.size() + 1; return NVRTC_SUCCESS; }
|
||||
static nvrtcResult e_log(nvrtcProgram prog, char* out) { std::strcpy(out, ((EmuProg*)prog)->log.c_str()); return NVRTC_SUCCESS; }
|
||||
static std::string image(EmuProg* p) { return "EMU-IMAGE:" + std::to_string(p->which); }
|
||||
static nvrtcResult e_imgSize(nvrtcProgram prog, size_t* n) { *n = image((EmuProg*)prog).size() + 1; return NVRTC_SUCCESS; }
|
||||
static nvrtcResult e_img(nvrtcProgram prog, char* out) { std::strcpy(out, image((EmuProg*)prog).c_str()); return NVRTC_SUCCESS; }
|
||||
static nvrtcResult e_addName(nvrtcProgram prog, const char* e) { ((EmuProg*)prog)->exprs.push_back(e); return NVRTC_SUCCESS; }
|
||||
static nvrtcResult e_lowered(nvrtcProgram prog, const char* e, const char** out) {
|
||||
EmuProg* p = (EmuProg*)prog;
|
||||
for (const std::string& s : p->exprs) if (s == e) { *out = s.c_str(); return NVRTC_SUCCESS; }
|
||||
return NVRTC_ERROR_NAME_EXPRESSION_NOT_VALID;
|
||||
}
|
||||
static const char* e_errstr(nvrtcResult r) { return r == NVRTC_SUCCESS ? "NVRTC_SUCCESS" : r == NVRTC_ERROR_COMPILATION ? "NVRTC_ERROR_COMPILATION" : "NVRTC_ERROR (emulated)"; }
|
||||
|
||||
void emu_fill_nvrtc(Rtc& r) {
|
||||
r.version = e_version; r.getNumSupportedArchs = e_numArchs; r.getSupportedArchs = e_archs;
|
||||
r.createProgram = e_create; r.destroyProgram = e_destroy; r.compileProgram = e_compile;
|
||||
r.getProgramLogSize = e_logSize; r.getProgramLog = e_log;
|
||||
r.getPTXSize = e_imgSize; r.getPTX = e_img; r.getCUBINSize = e_imgSize; r.getCUBIN = e_img;
|
||||
r.addNameExpression = e_addName; r.getLoweredName = e_lowered; r.getErrorString = e_errstr;
|
||||
}
|
||||
|
||||
// ---- Driver API stand-in
|
||||
static CUresult d_init(unsigned) { return CUDA_SUCCESS; }
|
||||
static CUresult d_driverVersion(int* v) { *v = 12080; return CUDA_SUCCESS; }
|
||||
static CUresult d_count(int* c) { *c = 1; return CUDA_SUCCESS; }
|
||||
static CUresult d_get(CUdevice* d, int i) { *d = i; return CUDA_SUCCESS; }
|
||||
static CUresult d_name(char* out, int len, CUdevice) { std::snprintf(out, (size_t)len, "CPU emulation shim (not a GPU)"); return CUDA_SUCCESS; }
|
||||
static CUresult d_attr(int* v, CUdevice_attribute a, CUdevice) {
|
||||
switch (a) {
|
||||
case CU_DEVICE_ATTRIBUTE_COMPUTE_CAPABILITY_MAJOR: *v = 12; break;
|
||||
case CU_DEVICE_ATTRIBUTE_COMPUTE_CAPABILITY_MINOR: *v = 0; break;
|
||||
case CU_DEVICE_ATTRIBUTE_WARP_SIZE: *v = 32; break;
|
||||
default: *v = 0; break;
|
||||
}
|
||||
return CUDA_SUCCESS;
|
||||
}
|
||||
static CUresult d_totalMem(size_t* b, CUdevice) { *b = 16ull << 30; return CUDA_SUCCESS; }
|
||||
static CUresult d_ctxFlags(CUdevice, unsigned) { return CUDA_SUCCESS; }
|
||||
static int gCtx = 0;
|
||||
static CUresult d_ctxRetain(CUcontext* c, CUdevice) { *c = (CUcontext)&gCtx; return CUDA_SUCCESS; }
|
||||
static CUresult d_ctxRelease(CUdevice) { return CUDA_SUCCESS; }
|
||||
static CUresult d_ctxSet(CUcontext) { return CUDA_SUCCESS; }
|
||||
static CUresult d_ctxSync() { return CUDA_SUCCESS; }
|
||||
static CUresult d_memInfo(size_t* f, size_t* t) { *f = 8ull << 30; *t = 16ull << 30; return CUDA_SUCCESS; }
|
||||
static CUresult d_alloc(CUdeviceptr* p, size_t n) { void* m = std::malloc(n); if (!m) return CUDA_ERROR_OUT_OF_MEMORY; *p = (CUdeviceptr)(uintptr_t)m; return CUDA_SUCCESS; }
|
||||
static CUresult d_free(CUdeviceptr p) { std::free((void*)(uintptr_t)p); return CUDA_SUCCESS; }
|
||||
static CUresult d_dtoh(void* dst, CUdeviceptr src, size_t n) { std::memcpy(dst, (const void*)(uintptr_t)src, n); return CUDA_SUCCESS; }
|
||||
static CUresult d_modLoad(CUmodule* m, const void* img) {
|
||||
const char* s = (const char*)img;
|
||||
if (std::strncmp(s, "EMU-IMAGE:", 10) != 0) return CUDA_ERROR_INVALID_IMAGE;
|
||||
*m = (CUmodule)(uintptr_t)(s[10] - '0');
|
||||
return CUDA_SUCCESS;
|
||||
}
|
||||
static CUresult d_modUnload(CUmodule) { return CUDA_SUCCESS; }
|
||||
static CUresult d_getFn(CUfunction* f, CUmodule m, const char* name) {
|
||||
int unit = (int)(uintptr_t)m, idx = 0;
|
||||
if (unit == 1 && std::strcmp(name, "igneum_cache_fill") == 0) idx = 1;
|
||||
else if (unit == 1 && std::strcmp(name, "igneum_build") == 0) idx = 2;
|
||||
else if (unit == 2 && std::strcmp(name, "igneum_hash_bound") == 0) idx = 3;
|
||||
else if (unit == 3 && std::strcmp(name, "igneum_build_inline") == 0) idx = 4;
|
||||
else if (unit == 3 && std::strcmp(name, "igneum_hash_inline") == 0) idx = 5;
|
||||
else return CUDA_ERROR_NOT_FOUND;
|
||||
*f = (CUfunction)(uintptr_t)idx;
|
||||
return CUDA_SUCCESS;
|
||||
}
|
||||
template <class T> static T arg(void** params, int i) { return *(T*)params[i]; }
|
||||
template <class T> static T* dptr(void** params, int i) { return (T*)(uintptr_t)(*(CUdeviceptr*)params[i]); }
|
||||
static CUresult d_launch(CUfunction f, unsigned gx, unsigned, unsigned, unsigned bx, unsigned, unsigned, unsigned, CUstream, void** params, void**) {
|
||||
switch ((int)(uintptr_t)f) {
|
||||
case 1: emu_launch(emu_pack::igneum_cache_fill, gx, bx, dptr<uint32_t>(params, 0), arg<uint32_t>(params, 1)); return CUDA_SUCCESS;
|
||||
case 2: emu_launch(emu_pack::igneum_build, gx, bx, dptr<uint32_t>(params, 0), dptr<const uint32_t>(params, 1), arg<uint32_t>(params, 2)); return CUDA_SUCCESS;
|
||||
case 3: emu_launch(emu_pack::igneum_hash_bound, gx, bx, dptr<const uint32_t>(params, 0), dptr<uint64_t>(params, 1), arg<uint32_t>(params, 2), arg<uint32_t>(params, 3), arg<emu_pack::IgneumInitWords>(params, 4)); return CUDA_SUCCESS;
|
||||
case 4: emu_launch(emu_inline::igneum_build_inline, gx, bx, dptr<uint32_t>(params, 0), dptr<const uint32_t>(params, 1), arg<uint32_t>(params, 2), arg<uint32_t>(params, 3)); return CUDA_SUCCESS;
|
||||
case 5: emu_launch(emu_inline::igneum_hash_inline, gx, bx, dptr<const uint32_t>(params, 0), dptr<uint64_t>(params, 1), arg<uint32_t>(params, 2), arg<uint32_t>(params, 3), arg<emu_inline::IgneumInitWords>(params, 4), arg<uint32_t>(params, 5)); return CUDA_SUCCESS;
|
||||
default: return CUDA_ERROR_INVALID_HANDLE;
|
||||
}
|
||||
}
|
||||
static CUresult d_streamCreate(CUstream* s, unsigned) { *s = (CUstream)&gCtx; return CUDA_SUCCESS; }
|
||||
static CUresult d_streamSync(CUstream) { return CUDA_SUCCESS; }
|
||||
static CUresult d_streamDestroy(CUstream) { return CUDA_SUCCESS; }
|
||||
static CUresult d_funcAttr(int* v, CUfunction_attribute, CUfunction) { *v = 0; return CUDA_SUCCESS; }
|
||||
static CUresult d_occ(int* n, CUfunction, int, size_t) { *n = 0; return CUDA_SUCCESS; }
|
||||
static CUresult d_errStr(CUresult r, const char** s) { *s = r == CUDA_SUCCESS ? "no error" : "emulated driver error"; return CUDA_SUCCESS; }
|
||||
static CUresult d_errName(CUresult r, const char** s) { *s = r == CUDA_SUCCESS ? "CUDA_SUCCESS" : "CUDA_ERROR (emulated)"; return CUDA_SUCCESS; }
|
||||
|
||||
void emu_fill_driver(Drv& d) {
|
||||
d.init = d_init; d.driverGetVersion = d_driverVersion; d.deviceGetCount = d_count; d.deviceGet = d_get; d.deviceGetName = d_name;
|
||||
d.deviceGetAttribute = d_attr; d.deviceTotalMem = d_totalMem; d.primaryCtxSetFlags = d_ctxFlags; d.primaryCtxRetain = d_ctxRetain;
|
||||
d.primaryCtxRelease = d_ctxRelease; d.ctxSetCurrent = d_ctxSet; d.ctxSynchronize = d_ctxSync; d.memGetInfo = d_memInfo;
|
||||
d.memAlloc = d_alloc; d.memFree = d_free; d.memcpyDtoH = d_dtoh; d.moduleLoadData = d_modLoad; d.moduleUnload = d_modUnload;
|
||||
d.moduleGetFunction = d_getFn; d.launchKernel = d_launch; d.streamCreate = d_streamCreate; d.streamSynchronize = d_streamSync;
|
||||
d.streamDestroy = d_streamDestroy; d.funcGetAttribute = d_funcAttr; d.occupancy = d_occ; d.getErrorString = d_errStr; d.getErrorName = d_errName;
|
||||
}
|
||||
93
proto-cuda/inline-bench/gen.py
Executable file
93
proto-cuda/inline-bench/gen.py
Executable file
|
|
@ -0,0 +1,93 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Ledger M16: the inline-cache kernel texts, derived from a pack by text substitution (never by hand).
|
||||
|
||||
gen.py <pack dir> <out dir>
|
||||
|
||||
Writes <out dir>/memhard_inline.h and <out dir>/kernel_inline.cu from the pack's memhard.h and kernel_bound.cu:
|
||||
|
||||
memhard_inline.h the pack's memory-hard core with the cache-line mask as a run-time argument (`lineMask`) instead
|
||||
of the literal MH_CACHE_LINE_MASK, every identifier renamed mh_ -> mhi_ and MH_ -> MHI_ so the
|
||||
two headers can share one translation unit
|
||||
kernel_inline.cu the pack's bound hash kernel with every `ds[x & mask]` replaced by `mhi_word(cache, (x) & mask,
|
||||
lineMask)`: the dataset is never read, each loaded word is derived from the cache on the spot
|
||||
(8 dependent cache-line reads and 9 mixer applications per item, spec 01 section 1.8); plus
|
||||
`igneum_build_inline`, the pack's dataset build kernel on the same parameterised core, which is
|
||||
the stored-dataset cross-check for a smaller cache
|
||||
|
||||
Every substitution is counted and the script refuses a pack whose text does not match the expected count (one mask
|
||||
literal, one mh_item signature, one mh_word, IGNEUM_LOADS_PER_HASH / IGNEUM_ITERATIONS loads), so a generator change
|
||||
that moves the load syntax fails here and not on the card. The inline kernel's output at the full 256 MiB mask must
|
||||
equal the pack's vectors bit for bit (inline_bench.cpp check 2); that is what proves the substitution.
|
||||
"""
|
||||
import re, sys, os, hashlib
|
||||
|
||||
def main():
|
||||
if len(sys.argv) != 3:
|
||||
print(__doc__); sys.exit(2)
|
||||
pack, out = sys.argv[1], sys.argv[2]
|
||||
os.makedirs(out, exist_ok=True)
|
||||
program = open(os.path.join(pack, 'program.h')).read()
|
||||
memhard = open(os.path.join(pack, 'memhard.h')).read()
|
||||
bound = open(os.path.join(pack, 'kernel_bound.cu')).read()
|
||||
loads = int(re.search(r'#define IGNEUM_LOADS_PER_HASH (\d+)', program).group(1))
|
||||
iters = int(re.search(r'#define IGNEUM_ITERATIONS (\d+)', program).group(1))
|
||||
cache_log2 = int(re.search(r'#define IGNEUM_CACHE_LOG2_WORDS (\d+)', program).group(1))
|
||||
expected_loads = loads // iters
|
||||
mask_literal = '0x%08xu' % ((1 << (cache_log2 - 4)) - 1)
|
||||
|
||||
# ---- memhard_inline.h
|
||||
h, n = re.subn(r'#define MH_CACHE_LINE_MASK ' + re.escape(mask_literal) + r'\n', '', memhard)
|
||||
if n != 1: fail('MH_CACHE_LINE_MASK literal %s' % mask_literal, n, 1)
|
||||
h, n = re.subn(r'\(s\[0\] & MH_CACHE_LINE_MASK\)', '(s[0] & lineMask)', h)
|
||||
if n != 1: fail('cache-line read mask', n, 1)
|
||||
h, n = re.subn(r'IGNEUM_HD void mh_item\(const uint32_t\* cache, uint32_t t, uint32_t\* s\)',
|
||||
'IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s, uint32_t lineMask)', h)
|
||||
if n != 1: fail('mh_item signature', n, 1)
|
||||
h, n = re.subn(r'IGNEUM_HD uint32_t mh_word\(const uint32_t\* cache, uint32_t w\) \{ uint32_t s\[16\]; mh_item\(cache, w >> 4u, s\); return s\[w & 15u\]; \}',
|
||||
'IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w, uint32_t lineMask) { uint32_t s[16]; mh_item(cache, w >> 4u, s, lineMask); return s[w & 15u]; }', h)
|
||||
if n != 1: fail('mh_word', n, 1)
|
||||
if 'MH_CACHE_LINE_MASK' in h: fail('MH_CACHE_LINE_MASK still referenced', 1, 0)
|
||||
h = re.sub(r'\bmh_', 'mhi_', h)
|
||||
h = re.sub(r'\bMH_', 'MHI_', h)
|
||||
h = h.replace('// Generated by', '// Derived by proto-cuda/inline-bench/gen.py (ledger M16) from memhard.h, itself generated by', 1)
|
||||
h = h.replace('#pragma once\n', '#pragma once\n// INLINE VARIANT: the cache-line mask is the run-time argument lineMask (0x%08x for the pack\'s %d MiB cache; a\n// smaller mask emulates a smaller cache that sits inside the GPU\'s L2). Identifiers carry the mhi_ / MHI_ prefix.\n' % ((1 << (cache_log2 - 4)) - 1, (1 << cache_log2) * 4 >> 20), 1)
|
||||
|
||||
# ---- kernel_inline.cu
|
||||
cut = bound.find('\ncudaError_t ')
|
||||
if cut < 0: fail('host launch wrappers in kernel_bound.cu', 0, 1)
|
||||
k = bound[:cut + 1]
|
||||
k, n = re.subn(r'#include "program.h"\n', '#include "program.h"\n#include "memhard_inline.h"\n', k)
|
||||
if n != 1: fail('program.h include', n, 1)
|
||||
k, n = re.subn(r'__global__ void igneum_hash_bound\(const uint32_t\* ds, uint64_t\* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw\)',
|
||||
'__global__ void igneum_hash_inline(const uint32_t* cache, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw, uint32_t lineMask)', k)
|
||||
if n != 1: fail('igneum_hash_bound signature', n, 1)
|
||||
k, n = re.subn(r'ds\[([^\]]+?) & mask\]', r'mhi_word(cache, (\1) & mask, lineMask)', k)
|
||||
if n != expected_loads: fail('ds[x & mask] loads', n, expected_loads)
|
||||
code = '\n'.join(l for l in k.split('\n') if not l.lstrip().startswith('//'))
|
||||
if re.search(r'\bds\b', code): fail('a dataset reference survived outside the comments', 1, 0)
|
||||
k = k.replace('// Generated by', '// Derived by proto-cuda/inline-bench/gen.py (ledger M16) from kernel_bound.cu, itself generated by', 1)
|
||||
k += '''
|
||||
// The pack's dataset build on the parameterised core: dataset item t derived with the cache-line mask `lineMask`.
|
||||
// With lineMask = the pack's mask this is igneum_build of kernel.cu; with a smaller mask it builds the dataset a
|
||||
// smaller cache implies, which igneum_hash_bound then reads: the stored twin of igneum_hash_inline at that mask.
|
||||
__global__ void igneum_build_inline(uint32_t* ds, const uint32_t* cache, uint32_t nItems, uint32_t lineMask) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mhi_item(cache, t, s, lineMask);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
'''
|
||||
open(os.path.join(out, 'memhard_inline.h'), 'w').write(h)
|
||||
open(os.path.join(out, 'kernel_inline.cu'), 'w').write(k)
|
||||
for name, text in (('memhard_inline.h', h), ('kernel_inline.cu', k)):
|
||||
print('%s %d bytes sha256 %s' % (name, len(text), hashlib.sha256(text.encode()).hexdigest()))
|
||||
print('loads replaced %d (IGNEUM_LOADS_PER_HASH %d / IGNEUM_ITERATIONS %d), pack mask %s' % (expected_loads, loads, iters, mask_literal))
|
||||
|
||||
def fail(what, got, want):
|
||||
print('gen.py: %s: %d substitution(s), expected %d; the pack layout is not the one this script knows' % (what, got, want), file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
main()
|
||||
605
proto-cuda/inline-bench/inline_bench.cpp
Normal file
605
proto-cuda/inline-bench/inline_bench.cpp
Normal file
|
|
@ -0,0 +1,605 @@
|
|||
// igneum-inline-bench: ledger M16 (the 256 MiB cache on a die) and E17 (the draw line per setting).
|
||||
//
|
||||
// BENCHMARK ONLY, beside the shipped worker (proto-cuda/nvrtc/worker.cpp), never inside it. No pool, no network,
|
||||
// no wallet. It measures, on one NVIDIA card and one pack, the honest kernel against the recompute attacker's:
|
||||
//
|
||||
// honest the pack's bound hash kernel reading the 1 GiB dataset (what every miner runs)
|
||||
// inline256 the same program with every dataset read replaced by its derivation from the 256 MiB cache
|
||||
// (8 dependent cache-line reads and 9 mixer applications per item, spec 01 section 1.8); the
|
||||
// cache sits in VRAM, so this is the Mac's MEMHARD.md 2.2 number on NVIDIA
|
||||
// inline64 the same derivation with a 64 MiB cache-line mask: the first quarter of the cache is all the kernel
|
||||
// touches, which fits inside the RTX 5090's 96 MiB L2. That is the on-die SRAM emulation the ledger
|
||||
// entry names: the attacker's rate with the cache in SRAM-class memory and the GPU's own integer
|
||||
// engine doing the mixer work. A 64 MiB cache is NOT the construction; it is the attacker's device
|
||||
// modelled on the honest card, so its outputs equal no pack vector and are checked another way
|
||||
// inline32 the same at 32 MiB (a second point inside the L2)
|
||||
//
|
||||
// Three bit-exact checks run before anything is timed, and the run is void if one fails:
|
||||
// 1. the pack's own self-test (cache head, last line, FNV; dataset head, last word, samples; the 96 vector lanes
|
||||
// through the honest kernel)
|
||||
// 2. the inline kernel at the pack's 256 MiB mask on the 96 vector lanes == the pack's vectors (this proves the
|
||||
// text substitution of gen.py and the derivation: the dataset is never read and the hashes are the same)
|
||||
// 3. at the 64 MiB mask (and 32 MiB): a dataset of --check-mib MiB built by igneum_build_inline at that mask,
|
||||
// read by the honest kernel, against igneum_hash_inline at the same mask on the same nonces: the stored and
|
||||
// the recomputed path of one construction agree on every lane (the vector warps and a batch of 2^13 nonces)
|
||||
//
|
||||
// Timing: each setting runs launches of 2^batch-log2 nonces until --seconds have passed (the first launch warms
|
||||
// up and is not counted; the worker's race times the same way), rate = hashes / wall seconds. A sampler thread runs
|
||||
// the --smi command every --smi-every seconds during each timed window and prints each line with the setting's
|
||||
// name (E17: nvidia-smi power.draw, clocks.sm, temperature.gpu, utilization.gpu). Every line that matters starts
|
||||
// with RESULT so a run job's report carries it.
|
||||
//
|
||||
// The kernel texts: the pack's kernel.cu (cache fill, dataset build), kernel_bound.cu (the honest hash) and
|
||||
// kernel_inline.cu + memhard_inline.h, which gen.py derives from the pack by text substitution. All four are handed
|
||||
// to NVRTC at run time with program.h and memhard.h, exactly as the shipped worker does; the exe prints the sha256
|
||||
// of every text it compiled.
|
||||
//
|
||||
// igneum-inline-bench --pack <dir> [--seconds 20] [--settings honest,inline256,inline64,inline32] [--batch-log2 22]
|
||||
// [--block-warps 1] [--check-mib 64] [--check] [--device 0] [--arch auto]
|
||||
// [--smi "nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu --format=csv,noheader"]
|
||||
// [--smi-every 2]
|
||||
//
|
||||
// The Mac check: built with -DIGNEUM_EMU against emu_backend.cpp (the driver API and NVRTC as host functions, the
|
||||
// kernel texts compiled by clang and run on host threads through proto-cuda/emu), `--check` runs the three checks
|
||||
// with no GPU. Rates from the emulation are not numbers.
|
||||
#include <cstdint>
|
||||
#include <cstdarg>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <chrono>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
#include <thread>
|
||||
#include <atomic>
|
||||
#include <mutex>
|
||||
#include <algorithm>
|
||||
|
||||
#include "../nvrtc/cuda_api.h"
|
||||
#include "../nvrtc/packfile.h"
|
||||
|
||||
#ifdef _WIN32
|
||||
#include <windows.h>
|
||||
#define POPEN _popen
|
||||
#define PCLOSE _pclose
|
||||
#else
|
||||
#include <dlfcn.h>
|
||||
#define POPEN popen
|
||||
#define PCLOSE pclose
|
||||
#endif
|
||||
|
||||
static const char* BENCH_VERSION = "1.0 (6 October 2026, ledger close round 2)";
|
||||
|
||||
static double wallMs() {
|
||||
using namespace std::chrono;
|
||||
return duration<double, std::milli>(steady_clock::now().time_since_epoch()).count();
|
||||
}
|
||||
static std::string fmt(const char* f, ...) {
|
||||
char b[4096];
|
||||
va_list ap; va_start(ap, f); std::vsnprintf(b, sizeof(b), f, ap); va_end(ap);
|
||||
return b;
|
||||
}
|
||||
static void line(const std::string& s) { std::fputs(s.c_str(), stdout); std::fputc('\n', stdout); std::fflush(stdout); }
|
||||
static void result(const std::string& s) { line("RESULT " + s); }
|
||||
static std::string readText(const std::string& path, bool& ok) {
|
||||
std::FILE* f = std::fopen(path.c_str(), "rb");
|
||||
if (!f) { ok = false; return ""; }
|
||||
std::string s; char buf[65536]; size_t n;
|
||||
while ((n = std::fread(buf, 1, sizeof(buf), f)) > 0) s.append(buf, n);
|
||||
std::fclose(f); ok = true; return s;
|
||||
}
|
||||
static std::string sha256Hex(const std::string& s) { char h[65]; pf_sha256_hex((const uint8_t*)s.data(), s.size(), h); return h; }
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// The two libraries (the same shape as the worker: run-time loading, no import library)
|
||||
|
||||
#ifndef IGNEUM_EMU
|
||||
static void* libOpen(const std::string& name) {
|
||||
#ifdef _WIN32
|
||||
return (void*)LoadLibraryA(name.c_str());
|
||||
#else
|
||||
return dlopen(name.c_str(), RTLD_NOW);
|
||||
#endif
|
||||
}
|
||||
static void* libSym(void* lib, const char* name) {
|
||||
#ifdef _WIN32
|
||||
return (void*)GetProcAddress((HMODULE)lib, name);
|
||||
#else
|
||||
return dlsym(lib, name);
|
||||
#endif
|
||||
}
|
||||
#endif
|
||||
#define LOAD_SYM(table, field, name) do { table.field = (decltype(table.field))libSym(lib, name); if (!table.field) { missing += std::string(missing.empty() ? "" : ", ") + name; } } while (0)
|
||||
|
||||
static bool loadDriver(Drv& d, std::string& err, std::string& libName) {
|
||||
#ifdef IGNEUM_EMU
|
||||
emu_fill_driver(d); libName = "emulation (host threads, no GPU)"; (void)err; return true;
|
||||
#else
|
||||
#ifdef _WIN32
|
||||
const char* names[] = { "nvcuda.dll" };
|
||||
#else
|
||||
const char* names[] = { "libcuda.so.1", "libcuda.so" };
|
||||
#endif
|
||||
void* lib = nullptr;
|
||||
for (const char* n : names) { lib = libOpen(n); if (lib) { libName = n; break; } }
|
||||
if (!lib) { err = "the CUDA driver library is not installed"; return false; }
|
||||
std::string missing;
|
||||
LOAD_SYM(d, init, "cuInit"); LOAD_SYM(d, driverGetVersion, "cuDriverGetVersion"); LOAD_SYM(d, deviceGetCount, "cuDeviceGetCount");
|
||||
LOAD_SYM(d, deviceGet, "cuDeviceGet"); LOAD_SYM(d, deviceGetName, "cuDeviceGetName"); LOAD_SYM(d, deviceGetAttribute, "cuDeviceGetAttribute");
|
||||
LOAD_SYM(d, deviceTotalMem, "cuDeviceTotalMem_v2"); LOAD_SYM(d, primaryCtxSetFlags, "cuDevicePrimaryCtxSetFlags_v2");
|
||||
LOAD_SYM(d, primaryCtxRetain, "cuDevicePrimaryCtxRetain"); LOAD_SYM(d, primaryCtxRelease, "cuDevicePrimaryCtxRelease_v2");
|
||||
LOAD_SYM(d, ctxSetCurrent, "cuCtxSetCurrent"); LOAD_SYM(d, ctxSynchronize, "cuCtxSynchronize"); LOAD_SYM(d, memGetInfo, "cuMemGetInfo_v2");
|
||||
LOAD_SYM(d, memAlloc, "cuMemAlloc_v2"); LOAD_SYM(d, memFree, "cuMemFree_v2"); LOAD_SYM(d, memcpyDtoH, "cuMemcpyDtoH_v2");
|
||||
LOAD_SYM(d, moduleLoadData, "cuModuleLoadData"); LOAD_SYM(d, moduleUnload, "cuModuleUnload"); LOAD_SYM(d, moduleGetFunction, "cuModuleGetFunction");
|
||||
LOAD_SYM(d, launchKernel, "cuLaunchKernel"); LOAD_SYM(d, streamCreate, "cuStreamCreate"); LOAD_SYM(d, streamSynchronize, "cuStreamSynchronize");
|
||||
LOAD_SYM(d, streamDestroy, "cuStreamDestroy_v2"); LOAD_SYM(d, funcGetAttribute, "cuFuncGetAttribute");
|
||||
LOAD_SYM(d, occupancy, "cuOccupancyMaxActiveBlocksPerMultiprocessor"); LOAD_SYM(d, getErrorString, "cuGetErrorString"); LOAD_SYM(d, getErrorName, "cuGetErrorName");
|
||||
if (!missing.empty()) { err = "the driver library lacks " + missing; return false; }
|
||||
return true;
|
||||
#endif
|
||||
}
|
||||
|
||||
#ifdef _WIN32
|
||||
static std::string exeDir() {
|
||||
char buf[MAX_PATH];
|
||||
DWORD n = GetModuleFileNameA(nullptr, buf, MAX_PATH);
|
||||
std::string p(buf, n);
|
||||
size_t i = p.find_last_of("\\/");
|
||||
return i == std::string::npos ? "." : p.substr(0, i);
|
||||
}
|
||||
#endif
|
||||
|
||||
static bool loadNvrtc(Rtc& r, std::string& err, std::string& libName) {
|
||||
#ifdef IGNEUM_EMU
|
||||
emu_fill_nvrtc(r); libName = "emulation (the texts are compiled by clang, see emu_backend.cpp)"; (void)err; return true;
|
||||
#else
|
||||
void* lib = nullptr;
|
||||
#ifdef _WIN32
|
||||
std::vector<std::string> names;
|
||||
if (const char* o = std::getenv("IGNEUM_NVRTC_DLL")) names.push_back(o);
|
||||
std::string dir = exeDir();
|
||||
WIN32_FIND_DATAA fd;
|
||||
HANDLE h = FindFirstFileA((dir + "\\nvrtc64_*_0.dll").c_str(), &fd);
|
||||
if (h != INVALID_HANDLE_VALUE) { do { std::string n = fd.cFileName; if (n.find(".alt.") == std::string::npos) names.push_back(dir + "\\" + n); } while (FindNextFileA(h, &fd)); FindClose(h); }
|
||||
names.push_back("nvrtc64_120_0.dll"); names.push_back("nvrtc64_130_0.dll");
|
||||
for (const std::string& n : names) { lib = libOpen(n); if (lib) { libName = n; break; } }
|
||||
if (!lib) { err = "nvrtc64_120_0.dll (and nvrtc-builtins64_128.dll) must sit next to the exe (copy them from the installed Igneum Miner)"; return false; }
|
||||
#else
|
||||
const char* names[] = { "libnvrtc.so.12", "libnvrtc.so" };
|
||||
for (const char* n : names) { lib = libOpen(n); if (lib) { libName = n; break; } }
|
||||
if (!lib) { err = "libnvrtc.so.12 not found"; return false; }
|
||||
#endif
|
||||
std::string missing;
|
||||
LOAD_SYM(r, version, "nvrtcVersion"); LOAD_SYM(r, createProgram, "nvrtcCreateProgram"); LOAD_SYM(r, destroyProgram, "nvrtcDestroyProgram");
|
||||
LOAD_SYM(r, compileProgram, "nvrtcCompileProgram"); LOAD_SYM(r, getProgramLogSize, "nvrtcGetProgramLogSize"); LOAD_SYM(r, getProgramLog, "nvrtcGetProgramLog");
|
||||
LOAD_SYM(r, getPTXSize, "nvrtcGetPTXSize"); LOAD_SYM(r, getPTX, "nvrtcGetPTX"); LOAD_SYM(r, getCUBINSize, "nvrtcGetCUBINSize"); LOAD_SYM(r, getCUBIN, "nvrtcGetCUBIN");
|
||||
LOAD_SYM(r, addNameExpression, "nvrtcAddNameExpression"); LOAD_SYM(r, getLoweredName, "nvrtcGetLoweredName"); LOAD_SYM(r, getErrorString, "nvrtcGetErrorString");
|
||||
if (!missing.empty()) { err = "the NVRTC library lacks " + missing; return false; }
|
||||
r.getNumSupportedArchs = (decltype(r.getNumSupportedArchs))libSym(lib, "nvrtcGetNumSupportedArchs");
|
||||
r.getSupportedArchs = (decltype(r.getSupportedArchs))libSym(lib, "nvrtcGetSupportedArchs");
|
||||
return true;
|
||||
#endif
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// Device and NVRTC (the worker's recipe: the device's own SASS when NVRTC knows it, else PTX for the driver)
|
||||
|
||||
struct Ctx {
|
||||
Drv drv; Rtc rtc;
|
||||
CUdevice dev = 0; CUcontext ctx = nullptr;
|
||||
std::string name; int major = 0, minor = 0, sms = 0, driverVersion = 0, rtcMajor = 0, rtcMinor = 0, l2Bytes = 0, smClockKhz = 0;
|
||||
std::string archOpt; bool ptx = false; std::string why;
|
||||
std::string err(CUresult r) { const char* s = nullptr; if (drv.getErrorString) drv.getErrorString(r, &s); return s ? s : "CUDA driver error"; }
|
||||
};
|
||||
#define DRV_CHECK(c, call, what) do { CUresult r_ = (call); if (r_ != CUDA_SUCCESS) { err = std::string(what) + ": " + (c).err(r_); return false; } } while (0)
|
||||
|
||||
static bool openDevice(Ctx& c, int device, const std::string& archArg, std::string& err) {
|
||||
DRV_CHECK(c, c.drv.init(0), "cuInit");
|
||||
int count = 0;
|
||||
DRV_CHECK(c, c.drv.deviceGetCount(&count), "cuDeviceGetCount");
|
||||
if (count == 0) { err = "no CUDA device"; return false; }
|
||||
if (device < 0 || device >= count) { err = fmt("device %d out of range (%d devices)", device, count); return false; }
|
||||
DRV_CHECK(c, c.drv.deviceGet(&c.dev, device), "cuDeviceGet");
|
||||
char name[256] = {0};
|
||||
DRV_CHECK(c, c.drv.deviceGetName(name, 255, c.dev), "cuDeviceGetName");
|
||||
c.name = name;
|
||||
for (char& ch : c.name) if (ch == ' ') ch = '_';
|
||||
DRV_CHECK(c, c.drv.deviceGetAttribute(&c.major, CU_DEVICE_ATTRIBUTE_COMPUTE_CAPABILITY_MAJOR, c.dev), "cc major");
|
||||
DRV_CHECK(c, c.drv.deviceGetAttribute(&c.minor, CU_DEVICE_ATTRIBUTE_COMPUTE_CAPABILITY_MINOR, c.dev), "cc minor");
|
||||
DRV_CHECK(c, c.drv.deviceGetAttribute(&c.sms, CU_DEVICE_ATTRIBUTE_MULTIPROCESSOR_COUNT, c.dev), "sm count");
|
||||
c.drv.deviceGetAttribute(&c.l2Bytes, CU_DEVICE_ATTRIBUTE_L2_CACHE_SIZE, c.dev);
|
||||
c.drv.deviceGetAttribute(&c.smClockKhz, CU_DEVICE_ATTRIBUTE_CLOCK_RATE, c.dev);
|
||||
c.drv.driverGetVersion(&c.driverVersion);
|
||||
c.drv.primaryCtxSetFlags(c.dev, CU_CTX_SCHED_BLOCKING_SYNC);
|
||||
DRV_CHECK(c, c.drv.primaryCtxRetain(&c.ctx, c.dev), "cuDevicePrimaryCtxRetain");
|
||||
DRV_CHECK(c, c.drv.ctxSetCurrent(c.ctx), "cuCtxSetCurrent");
|
||||
c.rtc.version(&c.rtcMajor, &c.rtcMinor);
|
||||
std::vector<int> archs;
|
||||
if (c.rtc.getNumSupportedArchs && c.rtc.getSupportedArchs) {
|
||||
int n = 0;
|
||||
if (c.rtc.getNumSupportedArchs(&n) == NVRTC_SUCCESS && n > 0 && n < 256) { archs.assign((size_t)n, 0); if (c.rtc.getSupportedArchs(archs.data()) != NVRTC_SUCCESS) archs.clear(); }
|
||||
}
|
||||
int cc = c.major * 10 + c.minor;
|
||||
if (archArg != "auto" && !archArg.empty()) { c.archOpt = archArg; c.ptx = archArg.rfind("compute_", 0) == 0; c.why = "--arch"; }
|
||||
else if (archs.empty()) { c.archOpt = fmt("sm_%d", cc); c.why = "the device's architecture (NVRTC did not list its targets)"; }
|
||||
else {
|
||||
bool known = false; int best = 0;
|
||||
for (int a : archs) { if (a == cc) known = true; if (a <= cc && a > best) best = a; }
|
||||
if (known) { c.archOpt = fmt("sm_%d", cc); c.why = "the device's architecture, listed by NVRTC"; }
|
||||
else if (best > 0) { c.archOpt = fmt("compute_%d", best); c.ptx = true; c.why = fmt("NVRTC does not know sm_%d; PTX for compute_%d", cc, best); }
|
||||
else { c.archOpt = fmt("compute_%d", archs.front()); c.ptx = true; c.why = "PTX for NVRTC's oldest target"; }
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
static const char* STUB_CUDA_RUNTIME =
|
||||
"#pragma once\n#ifndef __CUDACC_RTC__\n#error \"this stub is for NVRTC only\"\n#endif\n"
|
||||
"#ifdef __SIZE_TYPE__\ntypedef __SIZE_TYPE__ size_t;\n#elif defined(__LP64__) || defined(_LP64)\ntypedef unsigned long size_t;\n#else\ntypedef unsigned long long size_t;\n#endif\n";
|
||||
static const char* STUB_CSTDINT =
|
||||
"#pragma once\ntypedef signed char int8_t; typedef unsigned char uint8_t; typedef short int16_t; typedef unsigned short uint16_t;\n"
|
||||
"typedef int int32_t; typedef unsigned int uint32_t;\n"
|
||||
"#if defined(__LP64__) || defined(_LP64)\ntypedef long int64_t; typedef unsigned long uint64_t;\n#else\ntypedef long long int64_t; typedef unsigned long long uint64_t;\n#endif\n";
|
||||
|
||||
struct Compiled { std::vector<char> image; std::vector<std::string> lowered; double ms = 0; std::string log; };
|
||||
|
||||
static bool rtcCompile(Ctx& c, const std::string& src, const char* name, const std::vector<std::pair<std::string, std::string>>& hdrs,
|
||||
const std::vector<std::string>& nameExprs, Compiled& out, std::string& err) {
|
||||
double t0 = wallMs();
|
||||
std::vector<const char*> headers = { STUB_CUDA_RUNTIME, STUB_CSTDINT }, names = { "cuda_runtime.h", "cstdint" };
|
||||
for (const auto& h : hdrs) { names.push_back(h.first.c_str()); headers.push_back(h.second.c_str()); }
|
||||
nvrtcProgram prog = nullptr;
|
||||
nvrtcResult r = c.rtc.createProgram(&prog, src.c_str(), name, (int)headers.size(), headers.data(), names.data());
|
||||
if (r != NVRTC_SUCCESS) { err = std::string("nvrtcCreateProgram: ") + c.rtc.getErrorString(r); return false; }
|
||||
for (const std::string& e : nameExprs) {
|
||||
r = c.rtc.addNameExpression(prog, e.c_str());
|
||||
if (r != NVRTC_SUCCESS) { err = "nvrtcAddNameExpression " + e + ": " + c.rtc.getErrorString(r); c.rtc.destroyProgram(&prog); return false; }
|
||||
}
|
||||
std::string archOpt = "--gpu-architecture=" + c.archOpt;
|
||||
const char* opts[] = { archOpt.c_str(), "--std=c++17", "-default-device" };
|
||||
r = c.rtc.compileProgram(prog, 3, opts);
|
||||
{
|
||||
size_t logSize = 0;
|
||||
if (c.rtc.getProgramLogSize(prog, &logSize) == NVRTC_SUCCESS && logSize > 1) { std::vector<char> log(logSize); c.rtc.getProgramLog(prog, log.data()); out.log.assign(log.data(), logSize - 1); }
|
||||
}
|
||||
if (r != NVRTC_SUCCESS) {
|
||||
std::string one;
|
||||
for (char ch : out.log) { if (ch == '\n' || ch == '\r') { if (one.size() && one.back() != '|') one += " | "; } else one += ch; if (one.size() > 900) break; }
|
||||
err = std::string("nvrtcCompileProgram ") + name + " for " + c.archOpt + ": " + c.rtc.getErrorString(r) + ": " + one;
|
||||
c.rtc.destroyProgram(&prog); return false;
|
||||
}
|
||||
for (const std::string& e : nameExprs) {
|
||||
const char* lowered = nullptr;
|
||||
r = c.rtc.getLoweredName(prog, e.c_str(), &lowered);
|
||||
if (r != NVRTC_SUCCESS || !lowered) { err = "nvrtcGetLoweredName " + e + ": " + c.rtc.getErrorString(r); c.rtc.destroyProgram(&prog); return false; }
|
||||
out.lowered.push_back(lowered);
|
||||
}
|
||||
size_t n = 0;
|
||||
if (c.ptx) { r = c.rtc.getPTXSize(prog, &n); if (r == NVRTC_SUCCESS) { out.image.resize(n); r = c.rtc.getPTX(prog, out.image.data()); } }
|
||||
else { r = c.rtc.getCUBINSize(prog, &n); if (r == NVRTC_SUCCESS) { out.image.resize(n); r = c.rtc.getCUBIN(prog, out.image.data()); } }
|
||||
c.rtc.destroyProgram(&prog);
|
||||
if (r != NVRTC_SUCCESS || n == 0) { err = std::string(c.ptx ? "nvrtcGetPTX" : "nvrtcGetCUBIN") + ": " + c.rtc.getErrorString(r); return false; }
|
||||
out.ms = wallMs() - t0;
|
||||
return true;
|
||||
}
|
||||
|
||||
static bool deviceOnly(const std::string& text, std::string& out, std::string& err) {
|
||||
size_t cut = text.find("\n// Host-side launch wrappers");
|
||||
if (cut == std::string::npos) cut = text.find("\ncudaError_t ");
|
||||
if (cut == std::string::npos) { out = text; return true; }
|
||||
std::string tail = text.substr(cut + 1);
|
||||
if (tail.find("__global__") != std::string::npos || tail.find("__device__") != std::string::npos) { err = "device code after the host launch wrappers"; return false; }
|
||||
out = text.substr(0, cut + 1);
|
||||
return true;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// Options
|
||||
|
||||
struct Options {
|
||||
std::string pack, settings = "honest,inline256,inline64,inline32", arch = "auto";
|
||||
std::string smi = "nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu,utilization.gpu --format=csv,noheader";
|
||||
double seconds = 20; int batchLog2 = 22, blockWarps = 1, checkMib = 64, device = 0, smiEvery = 2; bool checkOnly = false;
|
||||
};
|
||||
static void usage() {
|
||||
line("igneum-inline-bench --pack <dir> [--seconds 20] [--settings honest,inline256,inline64,inline32] [--batch-log2 22] [--block-warps 1]");
|
||||
line(" [--check-mib 64] [--check] [--device 0] [--arch auto] [--smi \"<command>\"] [--smi-every 2]");
|
||||
}
|
||||
static Options parseArgs(int argc, char** argv) {
|
||||
Options o;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
std::string a = argv[i];
|
||||
auto next = [&]() -> std::string { if (i + 1 >= argc) { usage(); std::exit(2); } return argv[++i]; };
|
||||
if (a == "--pack") o.pack = next();
|
||||
else if (a == "--seconds") o.seconds = std::atof(next().c_str());
|
||||
else if (a == "--settings") o.settings = next();
|
||||
else if (a == "--batch-log2") o.batchLog2 = std::atoi(next().c_str());
|
||||
else if (a == "--block-warps") o.blockWarps = std::atoi(next().c_str());
|
||||
else if (a == "--check-mib") o.checkMib = std::atoi(next().c_str());
|
||||
else if (a == "--check") o.checkOnly = true;
|
||||
else if (a == "--device") o.device = std::atoi(next().c_str());
|
||||
else if (a == "--arch") o.arch = next();
|
||||
else if (a == "--smi") o.smi = next();
|
||||
else if (a == "--smi-every") o.smiEvery = std::atoi(next().c_str());
|
||||
else if (a == "-h" || a == "--help") { usage(); std::exit(0); }
|
||||
else { line("unknown argument " + a); usage(); std::exit(2); }
|
||||
}
|
||||
if (o.pack.empty()) { usage(); std::exit(2); }
|
||||
if (o.batchLog2 < 10 || o.batchLog2 > 26 || o.blockWarps < 1 || o.blockWarps > 32) { line("--batch-log2 10..26, --block-warps 1..32"); std::exit(2); }
|
||||
if (o.checkMib < 1 || o.checkMib > 1024 || (o.checkMib & (o.checkMib - 1))) { line("--check-mib must be a power of two between 1 and 1024"); std::exit(2); }
|
||||
return o;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
// The sampler (E17): one line of the --smi command every --smi-every seconds while a window is timed
|
||||
|
||||
struct Sampler {
|
||||
std::string cmd; int everyS = 2; std::atomic<bool> stop{false}; std::thread th; std::mutex mu; std::vector<std::string> lines;
|
||||
void start(const std::string& label) {
|
||||
if (cmd.empty()) return;
|
||||
stop = false;
|
||||
th = std::thread([this, label]() {
|
||||
while (!stop) {
|
||||
std::string out; std::FILE* p = POPEN(cmd.c_str(), "r");
|
||||
if (p) { char b[512]; while (std::fgets(b, sizeof(b), p)) out += b; PCLOSE(p); }
|
||||
while (!out.empty() && (out.back() == '\n' || out.back() == '\r')) out.pop_back();
|
||||
if (out.empty()) out = "(no output)";
|
||||
for (char& ch : out) if (ch == '\n' || ch == '\r') ch = ' ';
|
||||
{ std::lock_guard<std::mutex> g(mu); lines.push_back(out); }
|
||||
result("smi " + label + " " + out);
|
||||
for (int i = 0; i < everyS * 10 && !stop; ++i) std::this_thread::sleep_for(std::chrono::milliseconds(100));
|
||||
}
|
||||
});
|
||||
}
|
||||
std::vector<std::string> finish() { if (!th.joinable()) return {}; stop = true; th.join(); std::lock_guard<std::mutex> g(mu); std::vector<std::string> v = lines; lines.clear(); return v; }
|
||||
};
|
||||
// the median of the first comma field of each sample, read as a number ("123.45 W" -> 123.45); -1 when none parses
|
||||
static double medianFirstField(const std::vector<std::string>& v) {
|
||||
std::vector<double> xs;
|
||||
for (const std::string& s : v) { double x = 0; if (std::sscanf(s.c_str(), "%lf", &x) == 1) xs.push_back(x); }
|
||||
if (xs.empty()) return -1;
|
||||
std::sort(xs.begin(), xs.end());
|
||||
return xs[xs.size() / 2];
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------
|
||||
|
||||
struct IgneumInitWordsArg { uint32_t w[8]; };
|
||||
|
||||
struct Bench {
|
||||
Ctx& c; const PfPack& pk; CUfunction fCacheFill = nullptr, fBuild = nullptr, fHashBound = nullptr, fBuildInline = nullptr, fHashInline = nullptr;
|
||||
CUdeviceptr cache = 0, ds = 0, dOut = 0;
|
||||
uint32_t cacheWords = 0, dsWords = 0, lineMaskFull = 0, blockWarps = 1;
|
||||
Bench(Ctx& c_, const PfPack& pk_) : c(c_), pk(pk_) {}
|
||||
|
||||
bool launchHonest(CUdeviceptr dsBuf, uint32_t dsMask, CUdeviceptr out, uint32_t base, uint32_t nonces, std::string& err) {
|
||||
uint32_t block = 32u * blockWarps;
|
||||
IgneumInitWordsArg a; std::memcpy(a.w, pk.seedw, 32);
|
||||
void* args[5] = { &dsBuf, &out, &base, &dsMask, &a };
|
||||
DRV_CHECK(c, c.drv.launchKernel(fHashBound, nonces / block, 1, 1, block, 1, 1, 0, nullptr, args, nullptr), "cuLaunchKernel igneum_hash_bound");
|
||||
return true;
|
||||
}
|
||||
bool launchInline(uint32_t lineMask, uint32_t dsMask, CUdeviceptr out, uint32_t base, uint32_t nonces, std::string& err) {
|
||||
uint32_t block = 32u * blockWarps;
|
||||
IgneumInitWordsArg a; std::memcpy(a.w, pk.seedw, 32);
|
||||
void* args[6] = { &cache, &out, &base, &dsMask, &a, &lineMask };
|
||||
DRV_CHECK(c, c.drv.launchKernel(fHashInline, nonces / block, 1, 1, block, 1, 1, 0, nullptr, args, nullptr), "cuLaunchKernel igneum_hash_inline");
|
||||
return true;
|
||||
}
|
||||
bool sync(std::string& err) { DRV_CHECK(c, c.drv.streamSynchronize(nullptr), "cuStreamSynchronize"); return true; }
|
||||
bool readOut(CUdeviceptr out, std::vector<uint64_t>& v, size_t n, std::string& err) { v.resize(n); DRV_CHECK(c, c.drv.memcpyDtoH(v.data(), out, n * 8), "cuMemcpyDtoH"); return true; }
|
||||
};
|
||||
|
||||
static int compareLanes(const std::vector<uint64_t>& a, const std::vector<uint64_t>& b, size_t n, size_t& firstBad) {
|
||||
int bad = 0; firstBad = n;
|
||||
for (size_t i = 0; i < n; ++i) if (a[i] != b[i]) { if (bad == 0) firstBad = i; ++bad; }
|
||||
return bad;
|
||||
}
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
Options o = parseArgs(argc, argv);
|
||||
Ctx c;
|
||||
std::string err, drvLib, rtcLib;
|
||||
if (!loadDriver(c.drv, err, drvLib)) { result("error " + err); return 2; }
|
||||
if (!loadNvrtc(c.rtc, err, rtcLib)) { result("error " + err); return 2; }
|
||||
if (!openDevice(c, o.device, o.arch, err)) { result("error " + err); return 2; }
|
||||
result(fmt("bench igneum-inline-bench %s: device %d %s (sm_%d%d, %d SMs, L2 %d MiB, SM clock %d MHz), driver %d.%d from %s, NVRTC %d.%d from %s, target %s (%s)",
|
||||
BENCH_VERSION, o.device, c.name.c_str(), c.major, c.minor, c.sms, c.l2Bytes >> 20, c.smClockKhz / 1000, c.driverVersion / 1000, (c.driverVersion % 100) / 10,
|
||||
drvLib.c_str(), c.rtcMajor, c.rtcMinor, rtcLib.c_str(), c.archOpt.c_str(), c.why.c_str()));
|
||||
|
||||
PfPack pk; char perr[512];
|
||||
if (!pf_load(o.pack.c_str(), &pk, perr, sizeof(perr))) { result(std::string("error pack ") + o.pack + ": " + perr); return 2; }
|
||||
if (pk.datasetMode != 1) { result("error the pack is not memory-hard (IGNEUM_DATASET_MODE 1); the inline kernel has no meaning for a closed-form dataset"); return 2; }
|
||||
if (!pk.haveVectors) { result("error the pack carries no vectors.h; the bit-exact checks need it"); return 2; }
|
||||
bool ok1, ok2, ok3, ok4, ok5, ok6;
|
||||
std::string kernelCu = readText(o.pack + "/kernel.cu", ok1), boundCu = readText(o.pack + "/kernel_bound.cu", ok2);
|
||||
std::string programH = readText(o.pack + "/program.h", ok3), memhardH = readText(o.pack + "/memhard.h", ok4);
|
||||
std::string inlineCu = readText(o.pack + "/kernel_inline.cu", ok5), memhardInlineH = readText(o.pack + "/memhard_inline.h", ok6);
|
||||
if (!ok1 || !ok2 || !ok3 || !ok4) { result("error the pack lacks kernel.cu, kernel_bound.cu, program.h or memhard.h"); return 2; }
|
||||
if (!ok5 || !ok6) { result("error the pack lacks kernel_inline.cu or memhard_inline.h: run inline-bench/gen.py <pack> <pack> first"); return 2; }
|
||||
std::string kernelDev, boundDev;
|
||||
if (!deviceOnly(kernelCu, kernelDev, err) || !deviceOnly(boundCu, boundDev, err)) { result("error " + err); return 2; }
|
||||
{
|
||||
size_t n = 0, p = 0;
|
||||
while ((p = inlineCu.find("mhi_word(cache,", p)) != std::string::npos) { ++n; p += 10; }
|
||||
if (n == 0 || inlineCu.find("igneum_hash_inline") == std::string::npos || inlineCu.find("igneum_build_inline") == std::string::npos) { result("error kernel_inline.cu is not gen.py's output"); return 2; }
|
||||
result(fmt("texts pack %s seed %s: kernel.cu %s kernel_bound.cu %s program.h %s memhard.h %s kernel_inline.cu %s (%zu inline loads) memhard_inline.h %s",
|
||||
o.pack.c_str(), pk.seedString, sha256Hex(kernelCu).substr(0, 16).c_str(), sha256Hex(boundCu).substr(0, 16).c_str(), sha256Hex(programH).substr(0, 16).c_str(),
|
||||
sha256Hex(memhardH).substr(0, 16).c_str(), sha256Hex(inlineCu).substr(0, 16).c_str(), n, sha256Hex(memhardInlineH).substr(0, 16).c_str()));
|
||||
}
|
||||
|
||||
// Compile the three programs
|
||||
std::vector<std::pair<std::string, std::string>> hdrs = { { "program.h", programH }, { "memhard.h", memhardH }, { "memhard_inline.h", memhardInlineH } };
|
||||
Compiled ck, cb, ci;
|
||||
if (!rtcCompile(c, kernelDev, "kernel.cu", hdrs, { "igneum_cache_fill", "igneum_build" }, ck, err)) { result("error " + err); return 2; }
|
||||
if (!rtcCompile(c, boundDev, "kernel_bound.cu", hdrs, { "igneum_hash_bound" }, cb, err)) { result("error " + err); return 2; }
|
||||
if (!rtcCompile(c, inlineCu, "kernel_inline.cu", hdrs, { "igneum_build_inline", "igneum_hash_inline" }, ci, err)) { result("error " + err); return 2; }
|
||||
CUmodule mk = nullptr, mb = nullptr, mi = nullptr;
|
||||
Bench b(c, pk);
|
||||
b.blockWarps = (uint32_t)o.blockWarps;
|
||||
{
|
||||
CUresult r;
|
||||
if ((r = c.drv.moduleLoadData(&mk, ck.image.data())) != CUDA_SUCCESS) { result("error cuModuleLoadData kernel.cu: " + c.err(r)); return 2; }
|
||||
if ((r = c.drv.moduleLoadData(&mb, cb.image.data())) != CUDA_SUCCESS) { result("error cuModuleLoadData kernel_bound.cu: " + c.err(r)); return 2; }
|
||||
if ((r = c.drv.moduleLoadData(&mi, ci.image.data())) != CUDA_SUCCESS) { result("error cuModuleLoadData kernel_inline.cu: " + c.err(r)); return 2; }
|
||||
if (c.drv.moduleGetFunction(&b.fCacheFill, mk, ck.lowered[0].c_str()) != CUDA_SUCCESS || c.drv.moduleGetFunction(&b.fBuild, mk, ck.lowered[1].c_str()) != CUDA_SUCCESS) { result("error kernel.cu functions missing"); return 2; }
|
||||
if (c.drv.moduleGetFunction(&b.fHashBound, mb, cb.lowered[0].c_str()) != CUDA_SUCCESS) { result("error igneum_hash_bound missing"); return 2; }
|
||||
if (c.drv.moduleGetFunction(&b.fBuildInline, mi, ci.lowered[0].c_str()) != CUDA_SUCCESS || c.drv.moduleGetFunction(&b.fHashInline, mi, ci.lowered[1].c_str()) != CUDA_SUCCESS) { result("error kernel_inline.cu functions missing"); return 2; }
|
||||
}
|
||||
int regsH = 0, regsI = 0, occH = 0, occI = 0;
|
||||
c.drv.funcGetAttribute(®sH, CU_FUNC_ATTRIBUTE_NUM_REGS, b.fHashBound); c.drv.funcGetAttribute(®sI, CU_FUNC_ATTRIBUTE_NUM_REGS, b.fHashInline);
|
||||
c.drv.occupancy(&occH, b.fHashBound, 32 * o.blockWarps, 0); c.drv.occupancy(&occI, b.fHashInline, 32 * o.blockWarps, 0);
|
||||
result(fmt("compile %s: kernel.cu %.0f ms, kernel_bound.cu %.0f ms, kernel_inline.cu %.0f ms; registers honest %d inline %d; resident blocks/SM at %d warp(s)/block honest %d inline %d",
|
||||
c.archOpt.c_str(), ck.ms, cb.ms, ci.ms, regsH, regsI, o.blockWarps, occH, occI));
|
||||
|
||||
// Memory: the cache, the honest dataset, the check dataset, the outputs
|
||||
b.cacheWords = 1u << pk.cacheLog2Words; b.dsWords = 1u << pk.datasetLog2; b.lineMaskFull = (b.cacheWords / 16u) - 1u;
|
||||
uint32_t nonces = 1u << o.batchLog2, block = 32u * (uint32_t)o.blockWarps;
|
||||
if (nonces % block) { result("error the batch is not a multiple of the block"); return 2; }
|
||||
uint32_t checkWords = (uint32_t)(((uint64_t)o.checkMib << 20) / 4u);
|
||||
size_t cacheBytes = (size_t)b.cacheWords * 4u, dsBytes = (size_t)b.dsWords * 4u, checkBytes = (size_t)checkWords * 4u;
|
||||
{
|
||||
size_t freeB = 0, totalB = 0;
|
||||
if (c.drv.memGetInfo(&freeB, &totalB) == CUDA_SUCCESS) result(fmt("memory %llu MiB free of %llu; this run needs %llu MiB (cache %llu + dataset %llu + check dataset %llu + outputs)",
|
||||
(unsigned long long)(freeB >> 20), (unsigned long long)(totalB >> 20), (unsigned long long)((cacheBytes + dsBytes + checkBytes) >> 20) + 64, (unsigned long long)(cacheBytes >> 20), (unsigned long long)(dsBytes >> 20), (unsigned long long)(checkBytes >> 20)));
|
||||
CUresult r;
|
||||
if ((r = c.drv.memAlloc(&b.cache, cacheBytes)) != CUDA_SUCCESS) { result("error cuMemAlloc cache: " + c.err(r)); return 2; }
|
||||
if ((r = c.drv.memAlloc(&b.ds, dsBytes)) != CUDA_SUCCESS) { result("error cuMemAlloc dataset: " + c.err(r)); return 2; }
|
||||
if ((r = c.drv.memAlloc(&b.dOut, (size_t)nonces * 8u)) != CUDA_SUCCESS) { result("error cuMemAlloc out: " + c.err(r)); return 2; }
|
||||
}
|
||||
// Cache fill, dataset build
|
||||
double t0 = wallMs();
|
||||
{
|
||||
uint32_t nSeg = pk.cacheSegments, blk = 256u, grid = (nSeg + blk - 1u) / blk;
|
||||
void* args[2] = { &b.cache, &nSeg };
|
||||
CUresult r = c.drv.launchKernel(b.fCacheFill, grid, 1, 1, blk, 1, 1, 0, nullptr, args, nullptr);
|
||||
if (r == CUDA_SUCCESS) r = c.drv.streamSynchronize(nullptr);
|
||||
if (r != CUDA_SUCCESS) { result("error cache fill: " + c.err(r)); return 2; }
|
||||
}
|
||||
double cacheMs = wallMs() - t0; t0 = wallMs();
|
||||
{
|
||||
uint32_t nItems = b.dsWords / 16u, blk = 256u, grid = (nItems + blk - 1u) / blk;
|
||||
void* args[3] = { &b.ds, &b.cache, &nItems };
|
||||
CUresult r = c.drv.launchKernel(b.fBuild, grid, 1, 1, blk, 1, 1, 0, nullptr, args, nullptr);
|
||||
if (r == CUDA_SUCCESS) r = c.drv.streamSynchronize(nullptr);
|
||||
if (r != CUDA_SUCCESS) { result("error dataset build: " + c.err(r)); return 2; }
|
||||
}
|
||||
double dsMs = wallMs() - t0;
|
||||
result(fmt("fill cache %u MiB in %.1f ms (%u segments); dataset %u MiB built in %.1f ms (%.0f M items/s)", (unsigned)(cacheBytes >> 20), cacheMs, pk.cacheSegments, (unsigned)(dsBytes >> 20), dsMs, (double)(b.dsWords / 16u) / 1e3 / dsMs));
|
||||
|
||||
// Check 1: the pack's self-test through the honest kernel
|
||||
uint32_t dsMask = b.dsWords - 1u;
|
||||
std::vector<uint64_t> vecH((size_t)pk.vecWarps * 32u), vecI((size_t)pk.vecWarps * 32u), tmp;
|
||||
bool allPass = true;
|
||||
{
|
||||
std::vector<uint32_t> whole(b.cacheWords);
|
||||
if (c.drv.memcpyDtoH(whole.data(), b.cache, cacheBytes) != CUDA_SUCCESS) { result("error cuMemcpyDtoH cache"); return 2; }
|
||||
uint32_t cacheHead[16], cacheLast[16], dsHead[16], dsLast = 0;
|
||||
std::memcpy(cacheHead, whole.data(), 64); std::memcpy(cacheLast, whole.data() + b.cacheWords - 16u, 64);
|
||||
uint64_t fnv = pf_fnv1a64(whole.data(), cacheBytes);
|
||||
whole.clear(); whole.shrink_to_fit();
|
||||
c.drv.memcpyDtoH(dsHead, b.ds, 64);
|
||||
if (pk.dsLastIndex < b.dsWords) c.drv.memcpyDtoH(&dsLast, b.ds + (CUdeviceptr)pk.dsLastIndex * 4u, 4);
|
||||
std::vector<uint32_t> samples((size_t)(pk.nSamples > 0 ? pk.nSamples : 1), 0u);
|
||||
for (int i = 0; i < pk.nSamples; ++i) if (pk.sampleIdx[i] < b.dsWords) c.drv.memcpyDtoH(&samples[(size_t)i], b.ds + (CUdeviceptr)pk.sampleIdx[i] * 4u, 4);
|
||||
for (int w = 0; w < pk.vecWarps; ++w) {
|
||||
if (!b.launchHonest(b.ds, dsMask, b.dOut, pk.vecBase[w], block, err) || !b.sync(err) || !b.readOut(b.dOut, tmp, 32, err)) { result("error check 1: " + err); return 2; }
|
||||
std::memcpy(&vecH[(size_t)w * 32u], tmp.data(), 256);
|
||||
}
|
||||
char ln[1024];
|
||||
int pass = pf_selftest(&pk, cacheHead, cacheLast, fnv, dsHead, dsLast, samples.data(), vecH.data(), ln, sizeof(ln));
|
||||
result(std::string("check 1 honest kernel, the pack's self-test: ") + ln);
|
||||
allPass = allPass && pass != 0;
|
||||
}
|
||||
// Check 2: inline at the pack's mask == the pack's vectors
|
||||
{
|
||||
for (int w = 0; w < pk.vecWarps; ++w) {
|
||||
if (!b.launchInline(b.lineMaskFull, dsMask, b.dOut, pk.vecBase[w], block, err) || !b.sync(err) || !b.readOut(b.dOut, tmp, 32, err)) { result("error check 2: " + err); return 2; }
|
||||
std::memcpy(&vecI[(size_t)w * 32u], tmp.data(), 256);
|
||||
}
|
||||
std::vector<uint64_t> want((size_t)pk.vecWarps * 32u);
|
||||
for (int w = 0; w < pk.vecWarps; ++w) for (int l = 0; l < 32; ++l) want[(size_t)w * 32u + l] = pk.vecOut[w][l];
|
||||
size_t firstBad = 0; int bad = compareLanes(vecI, want, want.size(), firstBad);
|
||||
if (bad == 0) result(fmt("check 2 inline kernel at the %u MiB mask (0x%08x) on the %d vector lanes: PASS, every lane equals the pack's vectors (the dataset was never read)", (unsigned)(cacheBytes >> 20), b.lineMaskFull, pk.vecWarps * 32));
|
||||
else { result(fmt("check 2 inline kernel at the pack's mask: FAIL, %d of %d lanes differ, first at warp %zu lane %zu: device %016llx expected %016llx", bad, pk.vecWarps * 32, firstBad / 32, firstBad % 32, (unsigned long long)vecI[firstBad], (unsigned long long)want[firstBad])); allPass = false; }
|
||||
}
|
||||
// Check 3: at each smaller mask, the stored twin (igneum_build_inline at the mask, read by the honest kernel)
|
||||
// against the recomputed path (igneum_hash_inline at the mask), on a --check-mib dataset
|
||||
struct Setting { std::string name; uint32_t lineMask; bool honest; };
|
||||
std::vector<Setting> settings;
|
||||
{
|
||||
std::string s = o.settings; size_t p = 0;
|
||||
while (p <= s.size()) {
|
||||
size_t q = s.find(',', p); std::string t = s.substr(p, q == std::string::npos ? std::string::npos : q - p);
|
||||
if (t == "honest") settings.push_back({ t, 0, true });
|
||||
else if (t.rfind("inline", 0) == 0) { int mib = std::atoi(t.c_str() + 6); if (mib < 1 || mib > (int)(cacheBytes >> 20) || (mib & (mib - 1))) { result("error setting " + t + ": the MiB must be a power of two up to the cache size"); return 2; } settings.push_back({ t, (uint32_t)(((uint64_t)mib << 20) / 64u) - 1u, false }); }
|
||||
else if (!t.empty()) { result("error unknown setting " + t); return 2; }
|
||||
if (q == std::string::npos) break;
|
||||
p = q + 1;
|
||||
}
|
||||
}
|
||||
{
|
||||
CUdeviceptr dsCheck = 0; CUresult r;
|
||||
if ((r = c.drv.memAlloc(&dsCheck, checkBytes)) != CUDA_SUCCESS) { result("error cuMemAlloc check dataset: " + c.err(r)); return 2; }
|
||||
uint32_t checkMask = checkWords - 1u, nCheck = 1u << 13;
|
||||
if (nCheck < block) nCheck = block;
|
||||
for (const Setting& st : settings) {
|
||||
if (st.honest || st.lineMask == b.lineMaskFull) continue;
|
||||
uint32_t nItems = checkWords / 16u, blk = 256u, grid = (nItems + blk - 1u) / blk, lm = st.lineMask;
|
||||
void* args[4] = { &dsCheck, &b.cache, &nItems, &lm };
|
||||
r = c.drv.launchKernel(b.fBuildInline, grid, 1, 1, blk, 1, 1, 0, nullptr, args, nullptr);
|
||||
if (r == CUDA_SUCCESS) r = c.drv.streamSynchronize(nullptr);
|
||||
if (r != CUDA_SUCCESS) { result("error check 3 build at " + st.name + ": " + c.err(r)); return 2; }
|
||||
std::vector<uint64_t> hs, is; int bad = 0; size_t firstBad = 0, lanes = 0;
|
||||
for (int w = 0; w < pk.vecWarps; ++w) {
|
||||
if (!b.launchHonest(dsCheck, checkMask, b.dOut, pk.vecBase[w], block, err) || !b.sync(err) || !b.readOut(b.dOut, hs, block, err)) { result("error check 3: " + err); return 2; }
|
||||
if (!b.launchInline(st.lineMask, checkMask, b.dOut, pk.vecBase[w], block, err) || !b.sync(err) || !b.readOut(b.dOut, is, block, err)) { result("error check 3: " + err); return 2; }
|
||||
size_t fb; int bd = compareLanes(hs, is, block, fb); if (bd && bad == 0) firstBad = lanes + fb; bad += bd; lanes += block;
|
||||
}
|
||||
if (!b.launchHonest(dsCheck, checkMask, b.dOut, 0x20000000u, nCheck, err) || !b.sync(err) || !b.readOut(b.dOut, hs, nCheck, err)) { result("error check 3: " + err); return 2; }
|
||||
if (!b.launchInline(st.lineMask, checkMask, b.dOut, 0x20000000u, nCheck, err) || !b.sync(err) || !b.readOut(b.dOut, is, nCheck, err)) { result("error check 3: " + err); return 2; }
|
||||
{ size_t fb; int bd = compareLanes(hs, is, nCheck, fb); if (bd && bad == 0) firstBad = lanes + fb; bad += bd; lanes += nCheck; }
|
||||
// the two paths must also differ from the pack's vectors (a smaller cache is a different construction)
|
||||
size_t fbv; int diffFromPack = compareLanes(hs, vecH, (size_t)std::min<uint32_t>(32u, block), fbv);
|
||||
if (bad == 0) result(fmt("check 3 %s (mask 0x%08x, %d MiB check dataset): PASS, stored twin == recomputed on all %zu lanes%s", st.name.c_str(), st.lineMask, o.checkMib, lanes, diffFromPack ? "; differs from the pack's vectors, as a smaller cache must" : "; WARNING equals the pack's vectors"));
|
||||
else { result(fmt("check 3 %s: FAIL, %d of %zu lanes differ, first lane %zu: stored %016llx recomputed %016llx", st.name.c_str(), bad, lanes, firstBad, (unsigned long long)(firstBad < hs.size() ? hs[firstBad] : 0), (unsigned long long)(firstBad < is.size() ? is[firstBad] : 0))); allPass = false; }
|
||||
}
|
||||
c.drv.memFree(dsCheck);
|
||||
}
|
||||
result(std::string("checks ") + (allPass ? "PASS" : "FAIL") + ": every number below " + (allPass ? "stands" : "is VOID"));
|
||||
if (o.checkOnly || !allPass) { c.drv.memFree(b.dOut); c.drv.memFree(b.ds); c.drv.memFree(b.cache); return allPass ? 0 : 1; }
|
||||
|
||||
// Timing
|
||||
#ifdef IGNEUM_EMU
|
||||
result("note: emulation build, the rates below are host-thread rates and are not numbers");
|
||||
#endif
|
||||
Sampler smi; smi.cmd = o.smi; smi.everyS = o.smiEvery;
|
||||
double honestMhs = 0;
|
||||
std::vector<std::string> table;
|
||||
for (const Setting& st : settings) {
|
||||
smi.start(st.name);
|
||||
double tStart = 0; uint64_t hashes = 0; int launches = 0; bool ok = true;
|
||||
while (true) {
|
||||
uint32_t base = 0x40000000u + (uint32_t)launches * nonces;
|
||||
bool l = st.honest ? b.launchHonest(b.ds, dsMask, b.dOut, base, nonces, err) : b.launchInline(st.lineMask, dsMask, b.dOut, base, nonces, err);
|
||||
if (!l || !b.sync(err)) { result("error timing " + st.name + ": " + err); ok = false; break; }
|
||||
double now = wallMs();
|
||||
if (launches == 0) tStart = now; else hashes += nonces;
|
||||
++launches;
|
||||
if (launches >= 2 && now - tStart >= o.seconds * 1000.0) break;
|
||||
}
|
||||
std::vector<std::string> samples = smi.finish();
|
||||
if (!ok) break;
|
||||
double secs = (wallMs() - tStart) / 1000.0, mhs = (double)hashes / secs / 1e6;
|
||||
if (st.honest) honestMhs = mhs;
|
||||
double power = medianFirstField(samples);
|
||||
std::string ratio = honestMhs > 0 ? fmt("%.3f", mhs / honestMhs) : "n/a";
|
||||
std::string row = fmt("setting=%s lineMask=0x%08x cacheTouched=%u MiB blockWarps=%d batch=2^%d launches=%d hashes=%llu seconds=%.2f mhs=%.3f ratio_vs_honest=%s smi_samples=%zu smi_first_field_median=%s",
|
||||
st.name.c_str(), st.lineMask, st.honest ? (unsigned)(dsBytes >> 20) : (unsigned)(((uint64_t)(st.lineMask + 1u) * 64u) >> 20), o.blockWarps, o.batchLog2, launches - 1, (unsigned long long)hashes, secs, mhs, ratio.c_str(), samples.size(), power < 0 ? "none" : fmt("%.1f", power).c_str());
|
||||
result(row); table.push_back(row);
|
||||
}
|
||||
result("summary " + std::to_string(table.size()) + " settings timed; honest " + fmt("%.3f", honestMhs) + " Mhash/s; the inline ratios are the attacker's rate over the honest rate on this card");
|
||||
c.drv.memFree(b.dOut); c.drv.memFree(b.ds); c.drv.memFree(b.cache);
|
||||
c.drv.moduleUnload(mi); c.drv.moduleUnload(mb); c.drv.moduleUnload(mk);
|
||||
c.drv.primaryCtxRelease(c.dev);
|
||||
return 0;
|
||||
}
|
||||
88
proto-cuda/inline-bench/job-pc2.ps1
Normal file
88
proto-cuda/inline-bench/job-pc2.ps1
Normal file
|
|
@ -0,0 +1,88 @@
|
|||
# Igneum run job: ledger M16 (the inline-cache kernel, the on-die SRAM emulation) and E17 (the draw line per setting)
|
||||
# on PC 2's RTX 5090 (machine 1ccfe586), 6 October 2026, ledger close round 2. A signed `run` job, shell powershell,
|
||||
# not elevated. The card is taken whole for about 12 minutes: the app's miners are paused and the live prover switched
|
||||
# off for the run, both restored in `finally` (the prover-socket rule: every server and socket touched is put back).
|
||||
# The kit (igneum-inline-bench.exe + the pack with the inline texts) is fetched by https and refused unless its sha256
|
||||
# is the one below (the kit-path rule, C32: the presence check comes first and fails loudly). Every line that matters
|
||||
# starts with RESULT so `node tools/jobs.mjs <job id>` shows it.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
$KitUrl = 'https://dl.igneum.network/dl/__DL_TOKEN__/igneum-inline-bench-kit.zip'
|
||||
$KitSha = '__KIT_SHA256__'
|
||||
$Seconds = 15
|
||||
$Passes = @(1, 8) # warps per block: the worker's base (1) and a wide block (8), as the hot-table job ran
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
function Smi($q) { try { ((& nvidia-smi --query-gpu=$q --format=csv,noheader 2>$null) -join ' | ').Trim() } catch { 'nvidia-smi failed' } }
|
||||
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
$base = $null
|
||||
if (Test-Path $urlFile) { $base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/') }
|
||||
function State { if (-not $base) { return $null }; try { Invoke-RestMethod -Uri "$base/api/state" -TimeoutSec 20 } catch { $null } }
|
||||
function Post($path, $body) { if (-not $base) { return 'no app url' }; try { (Invoke-RestMethod -Method Post -Uri "$base$path" -ContentType 'application/json' -Body ($body | ConvertTo-Json -Compress) -TimeoutSec 15) | ConvertTo-Json -Compress } catch { "error: $_" } }
|
||||
function Card($st) { if ($null -eq $st) { return $null }; $st.mining.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 }
|
||||
|
||||
$job = $env:IGNEUM_JOB_DIR
|
||||
if (-not $job) { $job = Join-Path $env:TEMP 'igneum-inline-bench-job' }
|
||||
New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
"RESULT start $(Stamp) machine=$env:IGNEUM_MACHINE_ID job=$env:IGNEUM_JOB_ID app=$env:IGNEUM_APP_VERSION"
|
||||
|
||||
# 1. the kit: fetched, sha256-checked, extracted; fail loudly on anything missing (the kit-path rule)
|
||||
$zip = Join-Path $job 'kit.zip'
|
||||
& curl.exe -fsSL --retry 3 -m 300 -o $zip $KitUrl 2>&1 | ForEach-Object { "curl: $_" }
|
||||
if (-not (Test-Path $zip)) { "RESULT error kit missing: $KitUrl did not download"; exit 2 }
|
||||
$got = (Get-FileHash -Algorithm SHA256 $zip).Hash.ToLower()
|
||||
if ($got -ne $KitSha) { "RESULT error kit sha256 $got is not $KitSha; refused"; exit 2 }
|
||||
$kitDir = Join-Path $job 'kit'
|
||||
if (Test-Path $kitDir) { Remove-Item -Recurse -Force $kitDir }
|
||||
Expand-Archive -Path $zip -DestinationPath $kitDir -Force
|
||||
$exe = Join-Path $kitDir 'igneum-inline-bench\igneum-inline-bench.exe'
|
||||
$pack = Join-Path $kitDir 'igneum-inline-bench\pack'
|
||||
foreach ($f in @($exe, (Join-Path $pack 'kernel.cu'), (Join-Path $pack 'kernel_bound.cu'), (Join-Path $pack 'kernel_inline.cu'), (Join-Path $pack 'memhard_inline.h'), (Join-Path $pack 'program.json'), (Join-Path $pack 'vectors.h'))) {
|
||||
if (-not (Test-Path $f)) { "RESULT error kit file missing: $f"; exit 2 }
|
||||
}
|
||||
"RESULT kit $(Stamp) $KitUrl sha256 $got ok; exe $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower())"
|
||||
# the NVRTC DLLs come from the installed Igneum Miner (the same files the shipped worker loads)
|
||||
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
|
||||
if (-not $inst) { "RESULT error no installed igneum-worker-cuda.exe (the NVRTC DLLs come from there)"; exit 2 }
|
||||
Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination (Split-Path $exe) -Force
|
||||
$dlls = (Get-ChildItem (Split-Path $exe) -Filter 'nvrtc*.dll').Count
|
||||
if ($dlls -lt 2) { "RESULT error only $dlls NVRTC DLL(s) found in $inst"; exit 2 }
|
||||
"RESULT nvrtc $(Stamp) $dlls DLL(s) from $inst"
|
||||
|
||||
# 2. the card: the app's miners paused and the live prover off for the run, both restored in finally
|
||||
$st0 = State
|
||||
$proveWas = $null; $paused = $false; $rc = 1
|
||||
if ($st0) { $proveWas = [bool]$st0.settings.prove; "RESULT app $(Stamp) version=$($st0.version) machine=$($st0.machine_id) prove_setting=$proveWas proving=$($st0.proving.status) paused=$($st0.mining.paused)" } else { "RESULT app none: the app is not answering; measuring with whatever else runs on the card (labelled)" }
|
||||
try {
|
||||
if ($st0) {
|
||||
if ($proveWas) { "RESULT prove_off $(Stamp) $(Post '/api/prove' @{on=$false})" }
|
||||
"RESULT pause $(Stamp) $(Post '/api/pause' @{})"; $paused = $true
|
||||
$t = 0
|
||||
while ($t -lt 120) {
|
||||
Start-Sleep -Seconds 5; $t += 5
|
||||
$c = Card (State)
|
||||
$apps = (& nvidia-smi --query-compute-apps=pid,process_name --format=csv,noheader 2>$null) -join ' | '
|
||||
if ((-not $c -or $c.state -eq 'off' -or $c.pid -eq 0 -or $c.hash_now -le 0) -and -not $apps) { break }
|
||||
}
|
||||
$apps = (& nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv,noheader 2>$null) -join ' | '
|
||||
"RESULT card-quiet after $t s: compute apps on the card = $(if ($apps) { $apps } else { 'none' })$(if ($apps) { ' (SHARED CARD: the numbers below are labelled shared)' })"
|
||||
Start-Sleep -Seconds 5
|
||||
}
|
||||
"RESULT gpu-before $(Stamp) $(Smi 'name,driver_version,power.draw,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu,utilization.gpu')"
|
||||
|
||||
# 3. the bench: checks, then the four settings per pass, the E17 line sampled every 2 s inside each timed window
|
||||
$rc = 0
|
||||
foreach ($w in $Passes) {
|
||||
"RESULT pass block-warps=$w start $(Stamp)"
|
||||
& $exe --pack $pack --seconds $Seconds --settings honest,inline256,inline64,inline32 --batch-log2 22 --block-warps $w --check-mib 64 --smi-every 2 2>&1 | ForEach-Object { if ("$_" -match '^RESULT') { "$_" } else { "RESULT out $_" } }
|
||||
if ($LASTEXITCODE -ne 0) { "RESULT pass block-warps=$w exit $LASTEXITCODE"; $rc = $LASTEXITCODE }
|
||||
}
|
||||
"RESULT gpu-after $(Stamp) $(Smi 'power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu,utilization.gpu')"
|
||||
} finally {
|
||||
if ($paused) { "RESULT resume $(Stamp) $(Post '/api/resume' @{})" }
|
||||
if ($proveWas) { "RESULT prove_on $(Stamp) $(Post '/api/prove' @{on=$true})" }
|
||||
Start-Sleep -Seconds 10
|
||||
$st1 = State
|
||||
if ($st1) { "RESULT restored $(Stamp) prove_setting=$([bool]$st1.settings.prove) paused=$($st1.mining.paused) nvidia_state=$((Card $st1).state)" }
|
||||
}
|
||||
"RESULT end $(Stamp) exit $rc"
|
||||
exit $rc
|
||||
46
proto-cuda/inline-bench/make-kit.sh
Executable file
46
proto-cuda/inline-bench/make-kit.sh
Executable file
|
|
@ -0,0 +1,46 @@
|
|||
#!/usr/bin/env bash
|
||||
# The PC kit for the M16 run job: igneum-inline-bench.exe (build-windows.sh) and a copy of the pack with the inline
|
||||
# texts (gen.py) and the M28 stamp (stamp-pack.py), zipped as igneum-inline-bench-kit.zip with its sha256, copied to the
|
||||
# downloads folder (dl/<token>/, as push-build-inputs.sh does) and deployed. The run job (job-pc2.ps1) downloads the
|
||||
# zip by https, checks the sha256 and refuses anything else. Nothing secret goes in.
|
||||
# Usage: make-kit.sh [pack dir] [--no-deploy]
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
CUDA="$(cd "$HERE/.." && pwd)"
|
||||
PACK_SRC="$CUDA/packs/igneum-devnet-v4-epoch0"; DEPLOY=1
|
||||
for a in "$@"; do case "$a" in --no-deploy) DEPLOY=0 ;; *) PACK_SRC="$a" ;; esac; done
|
||||
EXE="$HERE/build-win/igneum-inline-bench.exe"
|
||||
[ -f "$EXE" ] || { echo "no $EXE: run build-windows.sh first" >&2; exit 1; }
|
||||
TOKEN="$(tr -d '[:space:]' < "$HOME/.config/igneum/dl-token")"
|
||||
DLSITE="${IGNEUM_DLSITE:-$(tr -d '[:space:]' < "$HOME/.config/igneum/dlsite-dir")}"
|
||||
[ -d "$DLSITE/dl/$TOKEN" ] || { echo "no downloads folder $DLSITE/dl/<token>" >&2; exit 1; }
|
||||
TMP="$(mktemp -d)"; KIT="$TMP/igneum-inline-bench"
|
||||
mkdir -p "$KIT/pack"
|
||||
cp "$EXE" "$KIT/"
|
||||
cp "$PACK_SRC"/* "$KIT/pack/"
|
||||
python3 "$HERE/gen.py" "$KIT/pack" "$KIT/pack"
|
||||
python3 "$HERE/stamp-pack.py" "$KIT/pack"
|
||||
cp "$HERE/job-pc2.ps1" "$KIT/" 2>/dev/null || true
|
||||
( cd "$KIT" && find . -type f | sort | xargs shasum -a 256 ) > "$KIT/MANIFEST.sha256"
|
||||
OUT="$DLSITE/dl/$TOKEN/igneum-inline-bench-kit.zip"
|
||||
rm -f "$OUT"
|
||||
( cd "$TMP" && zip -qr "$OUT" igneum-inline-bench -x '*.DS_Store' )
|
||||
SUM="$(shasum -a 256 "$OUT" | cut -d' ' -f1)"; SIZE="$(stat -f %z "$OUT")"
|
||||
printf '%s\n' "$SUM" > "${OUT%.zip}.sha256"
|
||||
echo "kit: $OUT $SIZE bytes sha256 $SUM"
|
||||
cat "$KIT/MANIFEST.sha256"
|
||||
rm -rf "$TMP"
|
||||
if [ "$DEPLOY" = 1 ]; then
|
||||
echo "deploying $DLSITE"
|
||||
(cd "$DLSITE" && npx --yes vercel@latest --global-config "$HOME/.config/igneum/vercel" deploy --prod --yes 2>&1 | grep -v "$TOKEN" || true)
|
||||
LIVE="$(mktemp)"; ok=0
|
||||
for try in 1 2 3 4 5 6; do
|
||||
code="$(curl -s -o "$LIVE" -w '%{http_code}' "https://dl.igneum.network/dl/$TOKEN/igneum-inline-bench-kit.sha256")"
|
||||
echo "https://dl.igneum.network/dl/<token>/igneum-inline-bench-kit.sha256 -> HTTP $code (try $try)"
|
||||
if [ "$code" = 200 ] && [ "$(tr -d '[:space:]' < "$LIVE")" = "$SUM" ]; then ok=1; break; fi
|
||||
sleep 10
|
||||
done
|
||||
rm -f "$LIVE"
|
||||
[ "$ok" = 1 ] || { echo "the live sha256 is not this zip's after 6 tries" >&2; exit 1; }
|
||||
echo "live: igneum-inline-bench-kit.zip verified by sha256 ($SUM, $SIZE bytes)"
|
||||
fi
|
||||
16
proto-cuda/inline-bench/stamp-pack.py
Executable file
16
proto-cuda/inline-bench/stamp-pack.py
Executable file
|
|
@ -0,0 +1,16 @@
|
|||
#!/usr/bin/env python3
|
||||
# Ledger M28: the workers refuse a pack whose program.json carries no kernel_sha256 stamp over the kernel files the
|
||||
# miner wrote from the seed. igneum-miner stamps the packs it writes; a pack copied from proto-cuda/packs (exported by
|
||||
# igneum-pow) carries none, so this stamps a COPY the way the miner does (the same snippet as nvrtc/emu/test.sh).
|
||||
# Usage: stamp-pack.py <pack dir>
|
||||
import hashlib, sys, os
|
||||
d = sys.argv[1]
|
||||
names = ["kernel.cu", "kernel_bound.cu", "kernel.cl", "kernel_bound.cl", "program.h", "memhard.h"]
|
||||
entries = ['"%s": "%s"' % (n, hashlib.sha256(open(os.path.join(d, n), 'rb').read()).hexdigest()) for n in names if os.path.exists(os.path.join(d, n))]
|
||||
p = os.path.join(d, "program.json")
|
||||
lines = [l for l in open(p).read().split("\n") if not l.strip().startswith('"kernel_sha256')]
|
||||
i = next(k for k, l in enumerate(lines) if l.strip() == "{")
|
||||
lines[i + 1:i + 1] = [' "kernel_sha256": {%s},' % ", ".join(entries),
|
||||
' "kernel_sha256_rule": "SHA-256 of each file as the miner wrote it from the seed; the chain commits to the seed, not to this text",']
|
||||
open(p, "w").write("\n".join(lines))
|
||||
print("stamped %s: %d files" % (p, len(entries)))
|
||||
68
tools/p17-conformance/run-2026-10-06.log
Normal file
68
tools/p17-conformance/run-2026-10-06.log
Normal file
|
|
@ -0,0 +1,68 @@
|
|||
23:23:02.293 t=0.3s fork binaries: /Users/joshm/Projects/igneum/vendor/igneum-node-ledger/target/release/igneumd /Users/joshm/Projects/igneum/vendor/igneum-node-ledger/target/release/igneum-miner
|
||||
23:23:03.119 t=1.1s n1 up pid 61205 json 29962 evm 29963 p2p 29961
|
||||
23:23:03.926 t=1.9s n0 up pid 61208 json 29952 evm 29953 p2p 29951
|
||||
23:23:04.730 t=2.7s n2 up pid 61222 json 29972 evm 29973 p2p 29971
|
||||
23:23:14.742 t=12.7s STEP fallback_before_first_lock: {"tip":8,"view":{"finalityActive":false,"finalityDepthDaa":"0x2d0","finalityReason":"window filling, 7 of 120","finalizedHeight":"0x0","finalizedSource":"genesis: the chain is younger than the finality depth","latestLockedHash":null,"latestLockedHeight":null,"latestLockedIndex":null,"readAtTip":"0x8","tip":"0x8"},"finalizedTagBlock":0,"latestEqualsFinalized":false}
|
||||
23:25:50.939 t=168.9s STEP first_lock: {"secs":168.9,"view":{"finalityActive":true,"finalityDepthDaa":"0x2d0","finalityReason":"active","finalizedHeight":"0x81","finalizedSource":"locked checkpoint","latestLockedHash":"0xda99631b6476541614a941a66c08f774e62da8b0597f4a8d3353d6fe38ae770b","latestLockedHeight":"0x81","latestLockedIndex":"0x5","readAtTip":"0x92","tip":"0x92"}}
|
||||
23:25:50.940 t=168.9s funded 0x70997970C51812dc3A010C7d01b50e0d17dc79C8 balance 147075638380000000000 wei
|
||||
23:25:50.947 t=168.9s tx1: sent to n0 nonce 0 hash 0x0ea89ea8842d8c03c793d23aa209c590e134dc1e6c0a8ebff9fb578dcaeb9f46
|
||||
23:25:50.949 t=168.9s tx1 n0: pending (block -, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:25:50.949 t=168.9s tx1 n1: unknown (block -, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:25:50.949 t=168.9s tx1 n2: unknown (block -, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:25:54.207 t=172.2s tx1 n0: executed (block 0x93, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:25:54.411 t=172.4s tx1 n0: unknown / reorged out (block -, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:25:56.453 t=174.4s tx1 n1: executed (block 0x95, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:25:56.657 t=174.6s tx1 n2: executed (block 0x95, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:19.494 t=197.5s tx1 n1: executed / finality paused (block 0x95, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:26:19.697 t=197.7s tx1 n0: executed / finality paused (block 0x95, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:26:19.697 t=197.7s tx1 n2: executed / finality paused (block 0x95, finalized tag 0x81 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:26:20.103 t=198.1s tx1 n1: finalised (block 0x95, finalized tag 0x99 by locked checkpoint, lockedCovered true, finality active)
|
||||
23:26:20.103 t=198.1s tx1 n2: finalised (block 0x95, finalized tag 0x99 by locked checkpoint, lockedCovered true, finality active)
|
||||
23:26:20.308 t=198.3s tx1 n0: finalised (block 0x95, finalized tag 0x99 by locked checkpoint, lockedCovered true, finality active)
|
||||
23:26:20.312 t=198.3s STEP happy_path: {"reachedFinalised":true,"chainBlock":"0x95","finalizedTag":153,"finalizedSource":"locked checkpoint","latest":173,"tagBelowTip":true}
|
||||
23:26:20.319 t=198.3s skipped pair: n0 0xbd03d37f0eee9fc5d5a4b969195f806bae957e1305e109f4542f45dbce01b3db; n2 0x317dc380fd5916ac7bd09e4f94c7cbfeed006298e9e6e22f0508d9d2793c2e75
|
||||
23:26:20.320 t=198.3s pairA n0: pending (block -, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:20.320 t=198.3s pairA n1: unknown (block -, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:20.320 t=198.3s pairA n2: unknown (block -, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:22.152 t=200.1s pairA n0: executed (block 0xb0, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:22.356 t=200.3s pairA n0: unknown / reorged out (block -, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:23.169 t=201.1s pairA n0: included / skipped (block -, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:23.169 t=201.1s pairA n1: included / skipped (block -, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:23.170 t=201.1s pairA n2: included / skipped (block -, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:23.171 t=201.1s pairB n0: executed (block 0xb0, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:23.171 t=201.1s pairB n1: executed (block 0xb0, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:23.171 t=201.1s pairB n2: executed (block 0xb0, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:26.174 t=204.2s STEP skipped: {"a":{"state":"included","failure":"skipped","includedIn":1},"b":{"state":"executed","failure":null,"includedIn":1,"skipReason":null}}
|
||||
23:26:26.175 t=204.2s link n2 -> n1 cut; n2 mines alone at 1/3
|
||||
23:26:32.179 t=210.2s tx3: sent to n2 nonce 2 hash 0x41e2f5a849db9865d4746565e3976d7e34a3d3bd6eba9bf97464d908f46e12cf
|
||||
23:26:32.180 t=210.2s tx3-isolated n2: pending (block -, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:26:35.847 t=213.8s tx3-isolated n2: executed (block 0xb7, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:27:00.854 t=238.8s STEP isolated: {"n2state":"executed","n2block":"0xb7","finalityOnN2AtCut":{"finalityActive":true,"finalityDepthDaa":"0x2d0","finalityReason":"active","finalizedHeight":"0x99","finalizedSource":"locked checkpoint","latestLockedHash":"0xd8d81ebc5b17d0ef960f16fbabe88b618023f9f95e7ab2eabca2528fa3f30d6a","latestLockedHeight":"0x99","latestLockedIndex":"0x6","readAtTip":"0xb7","tip":"0xb7"},"finalityOnN2After25s":{"finalityActive":true,"finalityDepthDaa":"0x2d0","finalityReason":"active","finalizedHeight":"0x99","finalizedSource":"locked checkpoint","latestLockedHash":"0xd8d81ebc5b17d0ef960f16fbabe88b618023f9f95e7ab2eabca2528fa3f30d6a","latestLockedHeight":"0x99","latestLockedIndex":"0x6","readAtTip":"0xbf","tip":"0xbf"},"tx3OnN2After25s":{"state":"executed","failure":null}}
|
||||
23:27:00.854 t=238.8s link n2 -> n1 healed; the heavier 2/3 chain should win on n2
|
||||
23:27:00.856 t=238.8s tx3-heal n2: executed (block 0xb7, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:27:20.521 t=258.5s tx3-heal n2: executed / finality paused (block 0xb7, finalized tag 0x99 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:28:05.761 t=303.7s tx3-heal n2: unknown / reorged out (block -, finalized tag 0xd3 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:28:05.764 t=303.7s tx3-reincluded n2: unknown / reorged out (block -, finalized tag 0xd3 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:28:05.764 t=303.7s tx3-reincluded n0: unknown (block -, finalized tag 0xd3 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:29:05.801 t=363.8s STEP reorged_out: {"observedReorgedOut":true,"n2ReorgLogLines":11,"afterHeal":{"state":"unknown","failure":"reorged out","reorgedFrom":"0xb7"},"later":{"state":"unknown","block":null},"healAt":238.8}
|
||||
23:29:20.969 t=378.9s v1 and v2 restarted with --no-vote (2/3 of the weight silent): finality should pause
|
||||
23:29:40.997 t=399.0s tx4: sent to n0 nonce 2 hash 0x4989eeb0735f024169e0f377e2e4ffba3d1b951db4fe34f92c07e70c0b7dda5d
|
||||
23:29:40.999 t=399.0s tx4-paused n0: pending (block -, finalized tag 0x126 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:29:40.999 t=399.0s tx4-paused n1: unknown (block -, finalized tag 0x126 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:29:40.999 t=399.0s tx4-paused n2: unknown (block -, finalized tag 0x126 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:29:44.463 t=402.4s tx4-paused n0: executed / finality paused (block 0x156, finalized tag 0x126 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:29:44.667 t=402.6s tx4-paused n1: executed / finality paused (block 0x156, finalized tag 0x126 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:29:44.870 t=402.8s tx4-paused n2: executed / finality paused (block 0x156, finalized tag 0x126 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:29:44.872 t=402.9s STEP paused: {"pausedAfter":{"secs":399,"view":{"finalityActive":false,"finalityDepthDaa":"0x2d0","finalityReason":"paused","finalizedHeight":"0x126","finalizedSource":"locked checkpoint","latestLockedHash":"0x380067b980c6c510029e9201b41d148dfa5705203ca72ee91289206492d21197","latestLockedHeight":"0x126","latestLockedIndex":"0xb","readAtTip":"0x153","tip":"0x153"}},"tx1":{"state":"finality not active","lockedCovered":true,"finality":"not active: paused"},"tx4":{"state":"executed","failure":"finality paused"},"finalizedTag":294,"finalizedSource":"locked checkpoint"}
|
||||
23:29:55.038 t=413.0s v1 and v2 sign again: finality should resume and tx4 reach finalised
|
||||
23:29:55.039 t=413.0s tx4-resumed n0: executed / finality paused (block 0x156, finalized tag 0x126 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:29:55.039 t=413.0s tx4-resumed n1: executed / finality paused (block 0x156, finalized tag 0x126 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:29:55.039 t=413.0s tx4-resumed n2: executed / finality paused (block 0x156, finalized tag 0x126 by locked checkpoint, lockedCovered false, finality not active: paused)
|
||||
23:29:55.243 t=413.2s tx4-resumed n0: executed (block 0x156, finalized tag 0x140 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:29:55.243 t=413.2s tx4-resumed n1: executed (block 0x156, finalized tag 0x140 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:29:55.244 t=413.2s tx4-resumed n2: executed (block 0x156, finalized tag 0x140 by locked checkpoint, lockedCovered false, finality active)
|
||||
23:30:19.278 t=437.3s tx4-resumed n0: finalised (block 0x156, finalized tag 0x15c by locked checkpoint, lockedCovered true, finality active)
|
||||
23:30:19.278 t=437.3s tx4-resumed n1: finalised (block 0x156, finalized tag 0x15c by locked checkpoint, lockedCovered true, finality active)
|
||||
23:30:19.481 t=437.5s tx4-resumed n2: finalised (block 0x156, finalized tag 0x15c by locked checkpoint, lockedCovered true, finality active)
|
||||
23:30:19.483 t=437.5s STEP resumed: {"tx4Finalised":true,"tx1":"finalised","view":{"finalityActive":true,"finalityDepthDaa":"0x2d0","finalityReason":"active","finalizedHeight":"0x15c","finalizedSource":"locked checkpoint","latestLockedHash":"0x8e36a46dcea6b8d41e0f19cc69aa4e6ebb77cf3dbcd3d4c0e611cbcbc593aa95","latestLockedHeight":"0x15c","latestLockedIndex":"0xd","readAtTip":"0x16c","tip":"0x16c"}}
|
||||
23:30:19.483 t=437.5s RESULT P17 conformance: PASSED in 437.5 s
|
||||
299
tools/p17-conformance/run.mjs
Normal file
299
tools/p17-conformance/run.mjs
Normal file
|
|
@ -0,0 +1,299 @@
|
|||
#!/usr/bin/env node
|
||||
// Ledger P17, O-7.2 (6 October 2026, ledger close round 2): the four-state word of design 2.4 observed on a fast-time
|
||||
// 3-node network built from the fork branch `ledger-fixes-2` (igneum_getTransactionStatus returns one `state` word,
|
||||
// `failure` names the failure path, `igneum_getFinalityView` says what the `finalized` tag resolves to and why).
|
||||
//
|
||||
// What is observed, each with the time it took and on which node:
|
||||
// happy path one transfer from the funded account: pending -> included -> executed -> finalised, the
|
||||
// `finalized` block tag following the latest locked checkpoint (never the tip)
|
||||
// fallback before the first lock: `finalized` resolves to genesis, source "the chain is younger than the
|
||||
// finality depth" (spec 3.9, the proof-of-work guidance), while `latest` is the tip
|
||||
// skipped two transfers with one nonce sent to two nodes in the same instant: one executes, the other copy is
|
||||
// skipped by rule and reports `included` with failure `skipped` (NonceTooLow)
|
||||
// reorged out node 2 is cut off behind its proxy and mines alone at 1/3 of the blocks; a transfer sent only to it
|
||||
// executes on its own chain; when the link heals the heavier 2/3 chain wins and node 2 unwinds the
|
||||
// segment: the transaction reports failure `reorged out` with the height it came from, then
|
||||
// `executed` again once node 2's mempool puts it in the new chain
|
||||
// paused the two voters holding 2/3 of the weight stop signing (restarted with --no-vote): finality pauses on
|
||||
// every node (spec 3.7 item 2), the finalised transaction reports `finality not active` with
|
||||
// lockedCovered true (the certificate still binds), a new transfer reports `executed` with failure
|
||||
// `finality paused`; when the voters sign again, both reach `finalised`
|
||||
// proven NOT exercised here: no shard is proved on this network tonight (the SP1 CPU prove takes minutes on
|
||||
// a loaded Mac and the proving loop is not part of the harness); the `proven` word is covered by the
|
||||
// unit test rpc::p17_tests on the paid-shard rule. Said in the bench-log.
|
||||
//
|
||||
// Ports 29950 to 29973 (node i: gRPC 29950+10i, p2p +1, json +2, EVM +3), proxies 29955 and 29965, network
|
||||
// igneum-devnet-1950, data under /tmp/igneum-p17. The live devnet and every other agent's range are never touched.
|
||||
// Under the run lock: tools/lock/with-lock.sh run node tools/p17-conformance/run.mjs
|
||||
// IGNEUMD, IGNEUM_MINER the fork build (default vendor/igneum-node-ledger/target/release, the ledger-fixes-2 build)
|
||||
// WARM (s, default 200) mining before the first transfer (the first lock needs min_daa 120 DAA s plus a checkpoint)
|
||||
import { spawn } from 'node:child_process';
|
||||
import { mkdirSync, rmSync, writeFileSync, readFileSync, openSync, existsSync, appendFileSync } from 'node:fs';
|
||||
import { createServer, connect as tcpConnect } from 'node:net';
|
||||
|
||||
const ROOT = new URL('../../', import.meta.url).pathname;
|
||||
const MAIN = process.env.IGNEUM_MAIN_ROOT || '/Users/joshm/Projects/igneum/';
|
||||
process.env.IGNEUM_FIN_TMP ||= '/tmp/igneum-p17';
|
||||
process.env.IGNEUM_FAST_TIME ||= '1';
|
||||
const { connectRpc } = await import(`${MAIN}tools/finality-attacks/lib/rpc.mjs`);
|
||||
const { privateKeyToAccount } = await import(`${MAIN}tools/prove-fixtures/node_modules/viem/_esm/accounts/index.js`);
|
||||
|
||||
const TMP = process.env.IGNEUM_FIN_TMP;
|
||||
const REL = process.env.IGNEUM_FORK_BIN || `${MAIN}vendor/igneum-node-ledger/target/release`;
|
||||
const IGNEUMD = process.env.IGNEUMD || `${REL}/igneumd`;
|
||||
const MINER = process.env.IGNEUM_MINER || `${REL}/igneum-miner`;
|
||||
const BASE = 29950, SUFFIX = 1950, CHAIN_ID = 4463;
|
||||
const WARM = +(process.env.WARM || 200);
|
||||
const DELAY_MS = +(process.env.DELAY_MS || 50);
|
||||
for (const b of [IGNEUMD, MINER]) if (!existsSync(b)) { console.error(`missing ${b}`); process.exit(2); }
|
||||
|
||||
const t0 = Date.now();
|
||||
const since = () => ((Date.now() - t0) / 1000).toFixed(1);
|
||||
const out = (...a) => { const l = `${new Date().toISOString().slice(11, 23)} t=${since()}s ${a.join(' ')}`; console.log(l); appendFileSync(`${TMP}/results.log`, l + '\n'); };
|
||||
const sleep = (ms) => new Promise(r => setTimeout(r, ms));
|
||||
const started = [];
|
||||
const report = { steps: {}, transitions: {} };
|
||||
const step = (k, v) => { report.steps[k] = v; out(`STEP ${k}: ${JSON.stringify(v)}`); };
|
||||
|
||||
// the funded account: the voters' payout address (a throwaway key, the one tools/proving-v0 uses)
|
||||
const funded = privateKeyToAccount('0x59c6995e998f97a5a0044966f0945389dc9e86dae88c7a8412f4603b6b78690d');
|
||||
const TO = '0x1111111111111111111111111111111111111111';
|
||||
|
||||
// the port gate: a previous run's nodes take seconds to exit after SIGINT; wait for the range to be free
|
||||
{
|
||||
const { execSync } = await import('node:child_process');
|
||||
for (let k = 0; k < 60; k++) {
|
||||
let busy = '';
|
||||
try { busy = execSync(`lsof -nP -iTCP:${BASE}-${BASE + 25} -sTCP:LISTEN 2>/dev/null | tail -n +2`, { encoding: 'utf8' }).trim(); } catch { busy = ''; }
|
||||
if (!busy) break;
|
||||
if (k === 59) { console.error(`ports ${BASE}..${BASE + 25} still held:\n${busy}`); process.exit(2); }
|
||||
await sleep(1000);
|
||||
}
|
||||
}
|
||||
rmSync(TMP, { recursive: true, force: true }); mkdirSync(TMP, { recursive: true });
|
||||
process.env.IGNEUM_FIN_BASE_PORT = String(BASE); process.env.IGNEUM_FIN_SUFFIX = String(SUFFIX);
|
||||
// the 60x fast-time file of THIS tree (the main checkout's copy still carries the dead `timestamp_deviation_tolerance`
|
||||
// field the node refuses, ledger close round 1) with skip_proof_of_work true; text edits, since the file holds u64::MAX
|
||||
// values JavaScript numbers cannot carry. Finality: min_daa 120, checkpoints every 30 DAA s, presence window 1,
|
||||
// finality depth 720 blocks.
|
||||
const override = `${TMP}/override.json`;
|
||||
{
|
||||
const text = readFileSync(`${ROOT}infra/fast-time/override-60x.json`, 'utf8').replace(/"skip_proof_of_work":\s*(true|false)/, '"skip_proof_of_work": true');
|
||||
if (!/"skip_proof_of_work": true/.test(text) || /timestamp_deviation_tolerance/.test(text)) throw new Error('override edit failed');
|
||||
writeFileSync(override, text);
|
||||
}
|
||||
|
||||
class Proxy { // one directed p2p link (the finality harness's shape); cut() drops it, heal() opens it again
|
||||
constructor(port, targetPort, delayMs = 0) { this.port = port; this.targetPort = targetPort; this.openGate = true; this.socks = new Set(); this.delayMs = delayMs; }
|
||||
get addr() { return `127.0.0.1:${this.port}`; }
|
||||
start() {
|
||||
return new Promise((resolve) => {
|
||||
this.server = createServer((client) => {
|
||||
if (!this.openGate) { client.destroy(); return; }
|
||||
const up = tcpConnect(this.targetPort, '127.0.0.1');
|
||||
this.socks.add(client); this.socks.add(up);
|
||||
const relay = (from, to) => from.on('data', (chunk) => setTimeout(() => { if (!to.destroyed) to.write(chunk); }, this.delayMs));
|
||||
relay(client, up); relay(up, client);
|
||||
const bye = () => { client.destroy(); up.destroy(); this.socks.delete(client); this.socks.delete(up); };
|
||||
client.on('error', bye); up.on('error', bye); client.on('close', bye); up.on('close', bye);
|
||||
});
|
||||
this.server.listen(this.port, '127.0.0.1', () => { started.push(this); resolve(this); });
|
||||
});
|
||||
}
|
||||
cut() { this.openGate = false; for (const s of this.socks) s.destroy(); this.socks.clear(); }
|
||||
heal() { this.openGate = true; }
|
||||
async stop() { this.cut(); await new Promise(r => this.server ? this.server.close(() => r()) : r()); }
|
||||
}
|
||||
|
||||
class Node {
|
||||
constructor(i, connect = []) {
|
||||
this.i = i; this.name = `n${i}`; this.grpcPort = BASE + i * 10; this.p2pPort = BASE + i * 10 + 1; this.jsonPort = BASE + i * 10 + 2; this.evmPort = BASE + i * 10 + 3;
|
||||
this.connect = connect; this.dir = `${TMP}/n${i}`; this.logFile = `${this.dir}/node.log`;
|
||||
}
|
||||
get grpc() { return `grpc://127.0.0.1:${this.grpcPort}`; }
|
||||
get evm() { return `http://127.0.0.1:${this.evmPort}`; }
|
||||
async start() {
|
||||
mkdirSync(this.dir, { recursive: true });
|
||||
const a = ['--devnet', `--devnet-suffix=${SUFFIX}`, '--nodnsseed', '--disable-upnp', '--nologfiles', '--enable-unsynced-mining', '--utxoindex', '--unsaferpc',
|
||||
`--appdir=${this.dir}`, `--rpclisten=127.0.0.1:${this.grpcPort}`, `--rpclisten-json=127.0.0.1:${this.jsonPort}`, `--evm-rpclisten=127.0.0.1:${this.evmPort}`,
|
||||
`--listen=127.0.0.1:${this.p2pPort}`, `--override-params-file=${override}`, '--loglevel=info', '--yes'];
|
||||
if (this.connect.length) a.push(...this.connect.map(c => `--connect=${c}`)); else a.push('--outpeers=0');
|
||||
const o = openSync(this.logFile, 'a');
|
||||
this.proc = spawn(IGNEUMD, a, { stdio: ['ignore', o, o], env: { ...process.env, IGNEUM_PROOF_VERIFY: 'trust' } });
|
||||
started.push(this);
|
||||
await sleep(800);
|
||||
this.rpc = await connectRpc(`ws://127.0.0.1:${this.jsonPort}`);
|
||||
out(`${this.name} up pid ${this.proc.pid} json ${this.jsonPort} evm ${this.evmPort} p2p ${this.p2pPort}`);
|
||||
return this;
|
||||
}
|
||||
async eth(method, params = []) {
|
||||
const r = await fetch(this.evm, { method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify({ jsonrpc: '2.0', id: 1, method, params }) });
|
||||
const j = await r.json();
|
||||
if (j.error) throw new Error(`${method}: ${j.error.message}`);
|
||||
return j.result;
|
||||
}
|
||||
async status(hash) { return this.eth('igneum_getTransactionStatus', [hash]).catch(e => ({ state: `rpc error: ${e.message}` })); }
|
||||
async view() { return this.eth('igneum_getFinalityView').catch(() => null); }
|
||||
async checkpoints(last = 400) { return this.rpc.call('getFinalityCheckpoints', { last }).catch(() => null); }
|
||||
grepLog(re) { try { return readFileSync(this.logFile, 'utf8').split('\n').filter(l => re.test(l)); } catch { return []; } }
|
||||
async stop() { if (this.rpc) { this.rpc.close(); this.rpc = null; } if (this.proc) { this.proc.kill('SIGINT'); await sleep(1500); try { this.proc.kill('SIGKILL'); } catch { } } }
|
||||
}
|
||||
class Miner {
|
||||
constructor(node, label, share, extra = []) { this.node = node; this.label = label; this.share = share; this.extra = extra; this.logFile = `${TMP}/vmine-${label}.log`; }
|
||||
start(secs = 3600) {
|
||||
const o = openSync(this.logFile, 'a');
|
||||
this.proc = spawn(MINER, ['vmine', this.node.grpc, String(secs), '--label', this.label, '--share', String(this.share), '--bps', '1', ...this.extra], { stdio: ['ignore', o, o] });
|
||||
started.push(this);
|
||||
return this;
|
||||
}
|
||||
async stop() { if (this.proc) { this.proc.kill('SIGINT'); for (let i = 0; i < 50 && this.proc.exitCode === null; i++) await sleep(100); try { this.proc.kill('SIGKILL'); } catch { } } }
|
||||
}
|
||||
async function stopAll() { for (const s of started.splice(0).reverse()) { try { await s.stop(); } catch { } } }
|
||||
process.on('SIGINT', async () => { await stopAll(); process.exit(130); });
|
||||
|
||||
/// Polls the status of `hash` on `nodes` every `everyMs` until `until(state)` holds on every node or the timeout,
|
||||
/// recording every distinct (state, failure) a node reports with the time it was first seen.
|
||||
async function watch(label, hash, nodes, until, timeoutMs, everyMs = 200) {
|
||||
const seen = {}; const first = {};
|
||||
const deadline = Date.now() + timeoutMs;
|
||||
let done = false;
|
||||
while (Date.now() < deadline && !done) {
|
||||
const states = await Promise.all(nodes.map(n => n.status(hash)));
|
||||
states.forEach((s, k) => {
|
||||
const key = `${s.state}${s.failure ? ' / ' + s.failure : ''}`;
|
||||
const nm = nodes[k].name;
|
||||
seen[nm] ||= []; first[nm] ||= {};
|
||||
if (!(key in first[nm])) { first[nm][key] = +since(); seen[nm].push(key); out(`${label} ${nm}: ${key} (block ${s.chainBlockNumber ?? '-'}, finalized tag ${s.finalizedHeight} by ${s.finalizedSource}, lockedCovered ${s.lockedCovered}, finality ${s.finality})`); }
|
||||
});
|
||||
done = states.every(until);
|
||||
if (!done) await sleep(everyMs);
|
||||
}
|
||||
report.transitions[label] = { seen, first, reached: done };
|
||||
return done;
|
||||
}
|
||||
// the fee: four times the node's folded gas price (the execution base fee on this line is 100 gwei), read per send
|
||||
async function fee(node) { const g = BigInt(await node.eth('eth_gasPrice')); return g * 4n + 1_000_000_000n; }
|
||||
async function signed(node, nonce, value) { return funded.signTransaction({ type: 'eip1559', chainId: CHAIN_ID, nonce, to: TO, value, gas: 21000n, maxFeePerGas: await fee(node), maxPriorityFeePerGas: 1_000_000_000n }); }
|
||||
async function sendTransfer(node, nonce, value, label) {
|
||||
const raw = await signed(node, nonce, value);
|
||||
const h = await node.eth('eth_sendRawTransaction', [raw]);
|
||||
out(`${label}: sent to ${node.name} nonce ${nonce} hash ${h}`);
|
||||
return h;
|
||||
}
|
||||
const hex = (v) => parseInt(v, 16);
|
||||
|
||||
try {
|
||||
out(`fork binaries: ${IGNEUMD} ${MINER}`);
|
||||
// topology: n1 listens; n0 and n2 each dial n1 through a proxy, so n2 can be cut off alone
|
||||
const n1 = await new Node(1).start();
|
||||
const p0 = await new Proxy(BASE + 5, n1.p2pPort, DELAY_MS).start();
|
||||
const p2 = await new Proxy(BASE + 15, n1.p2pPort, DELAY_MS).start();
|
||||
const n0 = await new Node(0, [p0.addr]).start();
|
||||
const n2 = await new Node(2, [p2.addr]).start();
|
||||
const nodes = [n0, n1, n2];
|
||||
await sleep(2000);
|
||||
// three voters at 1/3 each; v0 pays the funded account
|
||||
const v0 = new Miner(n0, 'v0', 1 / 3, ['--evm-address', funded.address]).start();
|
||||
let v1 = new Miner(n1, 'v1', 1 / 3).start();
|
||||
let v2 = new Miner(n2, 'v2', 1 / 3).start();
|
||||
await sleep(8000);
|
||||
const view0 = await n0.view();
|
||||
const tip0 = hex(await n0.eth('eth_blockNumber'));
|
||||
const finBlock = await n0.eth('eth_getBlockByNumber', ['finalized', false]);
|
||||
step('fallback_before_first_lock', { tip: tip0, view: view0, finalizedTagBlock: hex(finBlock.number), latestEqualsFinalized: hex(finBlock.number) === tip0 });
|
||||
|
||||
// warm until the first lock, watching the view flip
|
||||
let active = false;
|
||||
while (+since() < WARM) {
|
||||
await sleep(3000);
|
||||
const v = await n0.view();
|
||||
if (v && v.finalityActive && !active) { active = true; step('first_lock', { secs: +since(), view: v }); }
|
||||
if (active && +since() > 40 && v.latestLockedHeight != null) break;
|
||||
}
|
||||
if (!active) { const v = await n0.view(); step('first_lock', { secs: null, note: 'no lock within WARM', view: v }); }
|
||||
|
||||
// 1. the happy path
|
||||
const bal = BigInt(await n0.eth('eth_getBalance', [funded.address, 'latest']));
|
||||
out(`funded ${funded.address} balance ${bal} wei`);
|
||||
if (bal === 0n) throw new Error('the funded account has no rewards');
|
||||
let nonce = hex(await n0.eth('eth_getTransactionCount', [funded.address, 'latest']));
|
||||
const h1 = await sendTransfer(n0, nonce++, 1_000_000_000_000_000n, 'tx1');
|
||||
const ok1 = await watch('tx1', h1, nodes, s => s.state === 'finalised', 120_000);
|
||||
const s1 = await n0.status(h1);
|
||||
const finAfter = await n0.eth('eth_getBlockByNumber', ['finalized', false]);
|
||||
const latestAfter = await n0.eth('eth_getBlockByNumber', ['latest', false]);
|
||||
step('happy_path', { reachedFinalised: ok1, chainBlock: s1.chainBlockNumber, finalizedTag: hex(finAfter.number), finalizedSource: s1.finalizedSource, latest: hex(latestAfter.number), tagBelowTip: hex(finAfter.number) < hex(latestAfter.number) });
|
||||
|
||||
// 2. skipped: one nonce, two transfers, two nodes, the same instant
|
||||
const nSkip = nonce++;
|
||||
const rawA = await signed(n0, nSkip, 1n);
|
||||
const rawB = await signed(n2, nSkip, 2n);
|
||||
const [ra, rb] = await Promise.allSettled([n0.eth('eth_sendRawTransaction', [rawA]), n2.eth('eth_sendRawTransaction', [rawB])]);
|
||||
out(`skipped pair: n0 ${ra.status === 'fulfilled' ? ra.value : 'refused: ' + ra.reason.message}; n2 ${rb.status === 'fulfilled' ? rb.value : 'refused: ' + rb.reason.message}`);
|
||||
const pair = [ra, rb].filter(r => r.status === 'fulfilled').map(r => r.value);
|
||||
let skipped = null;
|
||||
if (pair.length === 2) {
|
||||
await watch('pairA', pair[0], nodes, s => s.state !== 'pending' && s.state !== 'unknown', 60_000);
|
||||
await watch('pairB', pair[1], nodes, s => s.state !== 'pending' && s.state !== 'unknown', 60_000);
|
||||
await sleep(3000);
|
||||
const sa = await n0.status(pair[0]), sb = await n0.status(pair[1]);
|
||||
skipped = { a: { state: sa.state, failure: sa.failure, includedIn: sa.includedIn.length }, b: { state: sb.state, failure: sb.failure, includedIn: sb.includedIn.length, skipReason: (sb.includedIn[0] || {}).skipReason } };
|
||||
} else {
|
||||
skipped = { note: 'the second copy was refused by the mempool before it reached a block (the relay carried the first copy over first), so no block holds a skipped copy; the skipped path is covered by the unit test', refused: [ra, rb].filter(r => r.status === 'rejected').map(r => r.reason.message) };
|
||||
}
|
||||
step('skipped', skipped);
|
||||
|
||||
// 3. reorged out: n2 alone behind a cut link
|
||||
p2.cut();
|
||||
out('link n2 -> n1 cut; n2 mines alone at 1/3');
|
||||
await sleep(6000);
|
||||
const h3 = await sendTransfer(n2, nonce, 3n, 'tx3'); // the nonce is reused on the main side below if n2 loses it
|
||||
await watch('tx3-isolated', h3, [n2], s => s.state === 'executed' || s.state === 'proven', 60_000);
|
||||
const isolated = await n2.status(h3);
|
||||
const viewIso = await n2.view();
|
||||
await sleep(25_000);
|
||||
const viewIso2 = await n2.view();
|
||||
const s3before = await n2.status(h3);
|
||||
step('isolated', { n2state: isolated.state, n2block: isolated.chainBlockNumber, finalityOnN2AtCut: viewIso, finalityOnN2After25s: viewIso2, tx3OnN2After25s: { state: s3before.state, failure: s3before.failure } });
|
||||
p2.heal();
|
||||
out('link n2 -> n1 healed; the heavier 2/3 chain should win on n2');
|
||||
const healT = +since();
|
||||
const reorg = await watch('tx3-heal', h3, [n2], s => s.failure === 'reorged out', 90_000, 100);
|
||||
const reorgLines = n2.grepLog(/selected-chain reorg/).length;
|
||||
const s3 = await n2.status(h3);
|
||||
await watch('tx3-reincluded', h3, [n2, n0], s => s.state === 'executed' || s.state === 'proven' || s.state === 'finalised', 60_000);
|
||||
const s3b = await n2.status(h3);
|
||||
step('reorged_out', { observedReorgedOut: reorg, n2ReorgLogLines: reorgLines, afterHeal: { state: s3.state, failure: s3.failure, reorgedFrom: s3.reorgedFrom }, later: { state: s3b.state, block: s3b.chainBlockNumber }, healAt: healT });
|
||||
nonce = hex(await n0.eth('eth_getTransactionCount', [funded.address, 'latest']));
|
||||
|
||||
// 4. finality paused: 2/3 of the weight stops signing
|
||||
await sleep(5000);
|
||||
await v1.stop(); await v2.stop();
|
||||
v1 = new Miner(n1, 'v1', 1 / 3, ['--no-vote']).start();
|
||||
v2 = new Miner(n2, 'v2', 1 / 3, ['--no-vote']).start();
|
||||
out('v1 and v2 restarted with --no-vote (2/3 of the weight silent): finality should pause');
|
||||
let paused = null;
|
||||
for (let k = 0; k < 60 && !paused; k++) { await sleep(2000); const v = await n0.view(); if (v && !v.finalityActive) paused = { secs: +since(), view: v }; }
|
||||
const s1p = await n0.status(h1);
|
||||
const h4 = await sendTransfer(n0, nonce++, 4n, 'tx4');
|
||||
await watch('tx4-paused', h4, nodes, s => s.state === 'executed' && s.failure === 'finality paused', 60_000);
|
||||
const s4p = await n0.status(h4);
|
||||
const finPaused = await n0.eth('eth_getBlockByNumber', ['finalized', false]);
|
||||
step('paused', { pausedAfter: paused, tx1: { state: s1p.state, lockedCovered: s1p.lockedCovered, finality: s1p.finality }, tx4: { state: s4p.state, failure: s4p.failure }, finalizedTag: hex(finPaused.number), finalizedSource: s4p.finalizedSource });
|
||||
await v1.stop(); await v2.stop();
|
||||
v1 = new Miner(n1, 'v1', 1 / 3).start();
|
||||
v2 = new Miner(n2, 'v2', 1 / 3).start();
|
||||
out('v1 and v2 sign again: finality should resume and tx4 reach finalised');
|
||||
const ok4 = await watch('tx4-resumed', h4, nodes, s => s.state === 'finalised', 150_000);
|
||||
const s1r = await n0.status(h1);
|
||||
step('resumed', { tx4Finalised: ok4, tx1: s1r.state, view: await n0.view() });
|
||||
report.ok = ok1 && ok4;
|
||||
out(`RESULT P17 conformance: ${report.ok ? 'PASSED' : 'FAILED'} in ${since()} s`);
|
||||
} catch (e) {
|
||||
report.error = e.message;
|
||||
out(`FAILED: ${e.message}`);
|
||||
} finally {
|
||||
await stopAll();
|
||||
writeFileSync(`${TMP}/report.json`, JSON.stringify(report, null, 2));
|
||||
console.log(JSON.stringify(report.steps, null, 2));
|
||||
}
|
||||
Loading…
Reference in a new issue