diff --git a/docs/bench-log.md b/docs/bench-log.md index f4c1097bc..85997dca6 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1728,6 +1728,7 @@ The 3-node run (`tools/txgen/relay-net.mjs`, new; this Mac, load 7 to 8 at the e Reading. Every transaction given to A was mined by B or C within 6 s, two thirds of them within 3 s, with no skipped copy: the hold on block-added kept B's and C's parallel blocks from carrying the same transfer. The afternoon run on the live devnet, through one node with the cooldown, had p50 40.7 s and p90 110.8 s with 50-s quiet stretches; here the 1.5 s p50 is one fast-time block plus the relay and the executor's lag. The pool depth matching on all three nodes at every sample is the convergence. Not measured here: a transaction flood above the per-peer rate (the bucket is unit-tested only), a 13 peer in the fleet (the digest check and the version gate are the evidence), and the hold's 30-s expiry on a block that never reaches the chain (not seen in 149 chain blocks). Commands: `IGNEUMD=vendor/igneum-node/target-txgossip/release/igneumd IGNEUM_MINER=vendor/igneum-node/target-txgossip/release/igneum-miner tools/lock/with-lock.sh run node tools/txgen/relay-net.mjs --rate 2 --duration 120 --wallets 16 --fund 2`; the Mac binaries from the fork worktree with `CARGO_TARGET_DIR=vendor/igneum-node/target-txgossip cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow` under the build lock (an APFS clone of `target-036`, 2 min 15 s to clone, 5 min 06 s to build); the suites with `node tools/build-job.mjs run --target 1ccfe586 --node vendor/igneum-node-txgossip --targets linux --node-tests "igneum-exec kaspa-p2p-flows" --no-app`. + ## 5 October 2026 (night), the SP1 CPU prover on PC 1 beside the miners, and the backend survey: no zkVM proves on AMD (amd-prove agent) the project lead, 22:50 BST: "test proving on the amd card?" and "can we test proving on mac?". The analysis with the backend table and the tier consequences: `docs/analysis/amd-proving.md`. The survey (SP1 v6.8.1 and `dev` 318dd530 of 28 Sep 2026, RISC Zero, Jolt, OpenVM, ICICLE, sppark; every claim cites a file or page there): on 5 October 2026 no zkVM proves on an AMD GPU; Apple silicon has RISC Zero's shipped Metal prover and ICICLE's Metal backend; SP1, the prover here, is CPU-only off NVIDIA. @@ -1884,6 +1885,7 @@ Reading, with the miner-on pairs above (empty shard 15,670 MiB, full prototype s | Unit tests | `cargo test --release -p kaspa-consensus-core -p igneum-exec --lib -- proving config::params::tests::override_params_carry_the_proving_v1 config::params::tests::consensus_digest` on this Mac (target `vendor/igneum-node/target-pv1`, 19:09Z): consensus core 13 passed (the segment record round trip, signature and the three nested sections; the credit split; the params switch and the digest that moves only once the switch is set), exec 8 passed (the segment grid and the split; the record checks: alignment, block, chain length, the veto naming the field, the deadline, the window; the chain rule both ways; the unproven restart; the shard side at 90%; the pool offering the segment section). The six full node suites go to PC 2 as a build job when the fleet is back | | The fast-time 3-node harness (`tools/proving-v1/net.mjs`, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac) | run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (`tools/proving-v1/report-2026-10-05.json`). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; `segmentsInWindow` proven 3, unproven 1. The shard side: a v1 shard's `shardWei` = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed. Run 3 on the FINAL fork tree (ece42979 on the 0.3.10 commit 21d4c73c, protocol 15, N = 8 both in the params default and `--segment 8`, the fast-time file's four fields), 20:52:41Z to 20:56:45Z: PASSED, 21 checks in 244.4 s (segments of 8: 119..126 paid in 1.0 s after submission, 127..134 refused fresh and paid continuing with chain_len 16, 135..142 left unproven and skipped, 143..150 restarted the chain) | + ## 5 October 2026 (night), dp4a-class throughput on the M5 Max: the dot4 emulation against the ALU chain (Counter ASIC 2.0 layer 7) Apple M5 Max, macOS 26, branch `ca2-analysis` (base `readwidth` 4badcee). The probes are standalone (no pack, no lottery kernel): `proto-metal/dot4-probe.swift` (built `swiftc -O -o dot4-probe dot4-probe.swift -framework Metal` under `with-lock.sh build`), `proto-opencl/dot4-probe.c` (built `cc -std=c99 -O2 -o dot4-probe-cl dot4-probe.c -framework OpenCL`), both run under `with-lock.sh measure` (exclusive; nothing else built or measured on the Mac during the runs). Shape: a dependent chain of one dot4 per step per lane, `acc = dot4(x, y, acc); x = x * 0x9E3779B1 + acc; y = rotl(y, 7) ^ (acc + s)`, 1,048,576 lanes x 4,096 steps, work-group 256, best of 3 with a fresh seed per repetition, device time (Metal: command buffer GPU start to end; OpenCL: event profiling). Beside it the ALU chain of the 9070 XT entry (`x = x * K + rotl(y, 7); y = (y ^ x) + s`, 5 ops per step counted). Every kernel is checked bit for bit against a CPU reference on lanes 0 and 1,048,575 in every repetition ("ok"). Design context: `docs/analysis/int8-matrix-family.md`. @@ -1911,6 +1913,29 @@ Reading: on this GPU a signed-byte dot4 costs 4.7 ALU-chain steps and an unsigne Reading: one `dp4a` on the 5090 costs about one ALU-chain step (7.45 T dot4/s, 0.85 of the chain's 8.75 T steps/s); one `v_dot4_i32_iu8` on the 9070 XT the same (0.66 T, 0.95 of its chain). Emulating the signed dot4 costs 6.0x the instruction on NVIDIA (the OpenCL compiler does not fold the four sign-extended products into `dp4a`) and 1.38x on AMD. Vendor ratios: the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on hardware dot4; against the M5 Max's best (unsigned emulation, 0.55 T) it is 10x on the chain and 13.6x on dot4. The hash itself is bound by DRAM reads, so these per-op numbers bound a family's cost and are not hash rates (`docs/analysis/int8-matrix-family.md` section 4). Adrenalin's OpenCL C accepts the clang builtin and emits the instruction on RDNA 4 (the third-party RDNA 3 report of the same route is now confirmed on this card); no PC platform lists the Khronos integer-dot extension. The 5090 SM clock read 2,505 MHz before and after (nvidia-smi; 2,850 MHz while mining in the telemetry entry), so the card was idle for the probe. +## 5 October 2026, layer 3 scratch soundness (Counter ASIC 2.0 step 4; branch ca2-soundness on readwidth b970dda; cryptographer) + +Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0, other agents' builds and the readwidth measurements running beside (the Mac measure lock was free during the GPU runs; nothing here is a hash-rate figure). Write-up `docs/analysis/scratch-soundness.md`; tests `igneum-pow/tests/scratch.rs`; harness `proto-metal/packbench` built from this branch (`--batch-base` added) into the session scratchpad with `swiftc -O -target arm64-apple-macos11 -framework Metal`. + +CPU, `with-lock.sh build nice -n 19 ~/.cargo/bin/cargo test -j4 --test scratch -- --nocapture` (3.6 s): 7 of 7 pass. Stats, 6 classes x 3 seeds x 2^11 units, every read-modify-write traced (3.1 to 12.6 million per class): written-word bias within 6 sigma (worst 3.63); re-hit rate measured against the uniform birthday rate 12.58 vs 10.91 percent (scr2k32), 21.47 vs 20.83 (scr4k32), 36.99 vs 36.50 (scr8k32), 3.84 vs 2.88 (scr2k128), 6.25 vs 5.82 (scr4k128), 11.89 vs 11.37 (scr8k128); slot histogram non-uniform (hottest slot 1.39x to 5.10x the mean: the slot is a register's low bits); deepest chain 5 to 9. Edge: 7 hand-built programs x 2 geometries x 4 bases against an independent hand model, 56 of 56, and 56 of 56 mismatches with the hand model's rewrite words swapped. Static scratch check: 42 of 42 emitted kernels of the 7 scr packs (regenerated byte for byte from program.json first), 6 deliberate breaks caught. Fuzz: 200 generated scratch programs, contract and acceptance on every instruction, 800 units; `IGNEUM_SCRATCH_PACKS_OUT` wrote 214 packs (57 s, three memory-hard caches). The crate's other tests: 33 of 34 lib tests pass; `verify::tests::fold_and_wide_fetch` fails on the readwidth tip itself (`verify.rs:508`, `k as u32 * 0x9E37_79B1` overflows under the test profile; not touched here). + +Metal, `with-lock.sh run