diff --git a/site/bench.html b/site/bench.html index 1ce0ea162..80fdb27a5 100644 --- a/site/bench.html +++ b/site/bench.html @@ -169,12 +169,12 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:
-
58 entries, newest at the bottom
+
59 entries, newest at the bottom

Engineering log

Every measurement the project has made, newest at the bottom, written by the people and agents who ran it, with the commands and hardware. Prototype numbers are not mining numbers and say so.

- +

Igneum bench log

Append-only. Every number here was measured on the machine named, on the date given.

2026-10-03 proto-metal / igneum-bench, first run

@@ -421,7 +421,19 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:
Measurev2 (rule as specified, view-local weights)v3 (plus the frozen table)
50/50, 60/40, 55/45 honest partitions, 12 days: first lock alone per sideday 10.2 / 10.1; 5.2 / never; 7.9 / 12.0never / never in every split; 0 conflicts; every pre-heal lock kept; first lock 0 min after the heal
50/50 and 60/40 for 31 days(as above)both sides at day 30.00, when the frozen table expires; first conflict day 30.00 to 30.06
70/30 for 150 and 360 minthe 70 side from minute 0 to 4, the 30 side never, 0 conflictsthe same
67/33 (the 4/2 split at exactly two thirds), 360 minthe 67 side locked in 1 of 2 seeds after 239 min (the 2.2% outage knife edge)never in 360 min
35% and 50% stop mining and signing at oncefirst lock day 1.7 and 10.1day 30.00 for both (the frozen table holds the departed keys until it expires)
equivocator across a 50/50 split, 30% / 33% / 34% of total(H: 0 / 0 / conflicts)0 / 0 / 21 to 69 conflicts from minute 14 to 78: the one-third bound of 3.11.2 is unchanged

Fast-time 3-node network (node tools/finality-attacks/v3.mjs, the target-finality build, infra/fast-time/override-60x.json merged with skip_proof_of_work and finality_v3_activation_daa 0, or the file's own "never" for the v2 control; W = 120 DAA; n1 listens, n0 and n2 dial it through TCP proxies that hold every byte 300 ms each way, the stand-in for tc/netem which macOS lacks, so n0 to n2 is 600 ms plus n1's relay; six vmine voters at 1/6 of 1 block/s in all; machine shared with other agents' builds and the live devnet). Raw tables and the v3 split's finality log lines in docs/benchmarks/finality-v3-2026-10-04/.

RunCriterionMeasuredVerdict
fold, v2 control, 480 s(baseline)12 locked indices per node; the certificate each node held at the end carried 4 or 5 of 6 votes (means 4.67 / 4.50 / 4.25 on n0 / n1 / n2), never 6; median lock latency 1,008 msthe F22 state reproduced with 300-ms links
fold, v3, 480 sthe held certificate carries at least 95% of connected keys' votes (6 of 6)11 locked indices per node; first-built certificates 4 or 5 of 6 (means 4.75 / 4.36 / 4.45, as under v2); held certificates 6 of 6 at 11 of 11 indices on every node (100%); 9 to 10 fold lines per node ("every voter signed", 0.3 to 1.4 s after the first build) and 1 to 4 replacements by a heavier gossiped certificate; 0 conflicting certificates; median lock latency 1,008 ms, unchangedPASS
split50, v2 control: 3/3 split 150 s (old bound W / (3R) = 80 s at R = 0.5 blocks/s per side), heal window 200 s(the fork of 3.7 item 9)side B (n1, n2) locked alone from 126 s after the cut, 3 locks; side A none; n0 redialled 72 s after the gate reopened; at the end 7 CONFLICTING certificate lines (4 on n0, 3 on n1) and 4 locked indices disagreeing across the three nodesthe fork, as on the morning's three-node and cloud runs
split50, v3: the same cutneither side locks during the split; heal; locking resumes on one chain; 0 conflicting certificates0 / 0 / 0 new locks during the 150 s (the v2 control locked at 126 s, so the frozen table held side B for the checkpoints of the last 24 s; the frozen table would have expired at 240 s); n0 redialled 72 s after the gate reopened; all three nodes resumed at index 7 and reached 13 inside the heal window; 0 conflicting certificates; 0 disagreeing locked indices; every post-heal LOCKED line names the frozen lock and its fraction (81 to 94% of the frozen table)PASS
split70, v3: 4/2 keys, the 4 side at 70% of weight (shares 0.175 x 4 against 0.15 x 2), 150 sthe 4 side locks during the split, the 2 side does not; 0 conflicts4 side: 4 new locks, the first 30 s after the cut; 2 side: 0; heal: all three at 17; 0 conflicting certificates; 0 disagreeing indicesPASS (6B at 70/30; exactly 4/6 is a knife edge under both rules, simulator row above)
-

What remains uncertain. (1) The v3 split's hold was observed over the last 24 s of a 150-s split (the control crossed at 126 s); a longer split under the 240-s expiry (say 200 s) would show more held checkpoints, and the "held by the frozen table" line is logged at debug, which the runs did not enable. (2) No cloud rehearsal: the 12-node the cloud provider network was destroyed at 15:30 UTC, so the 95% target is shown on three nodes with emulated 300-ms links and on the cloud logs' arithmetic, not on the cloud topology itself; the rollout plan names the re-creation and the partition experiment to run first. (3) The frozen reference is the node's own highest lock on the chain, not the certificate carried in C_i's past, so two honest nodes can test one checkpoint against tables 30 s apart; in a connected network those tables differ by a minute of blocks, under a partition both are pre-split, and no run showed a disagreement, but it is a property argued, not proved. (4) The price: a sudden departure of a third or more now pauses finality for a full window (30 days on mainnet) instead of 1.4 to 10 days; a gradual one costs nothing. the maintainers asked for the pause over the fork; the number is stated in spec 3.7 item 2. (5) The fold clock is in memory: a restarted node folds from daa(C_i) + depth, a few seconds late at worst. (6) Binaries, all from finality-fixes 6aa69a45, hashes and checks in the rollout plan's section 2: Mac native fe982a1d... (verified running), Linux 7c100fc2... (cargo-zigbuild, 34 min, not run on a Linux host), Windows cc1d1001... (mingw, 12 min 28 s, the v2 exe's DLL set, cannot run here); the Windows payload inputs were staged with push-inputs.sh --no-deploy into a scratch folder and NOT deployed (plan 7a).

+

What remains uncertain. (1) The v3 split's hold was observed over the last 24 s of a 150-s split (the control crossed at 126 s); a longer split under the 240-s expiry (say 200 s) would show more held checkpoints, and the "held by the frozen table" line is logged at debug, which the runs did not enable. (2) No cloud rehearsal: the 12-node the cloud provider network was destroyed at 15:30 UTC, so the 95% target is shown on three nodes with emulated 300-ms links and on the cloud logs' arithmetic, not on the cloud topology itself; the rollout plan names the re-creation and the partition experiment to run first. (3) The frozen reference is the node's own highest lock on the chain, not the certificate carried in C_i's past, so two honest nodes can test one checkpoint against tables 30 s apart; in a connected network those tables differ by a minute of blocks, under a partition both are pre-split, and no run showed a disagreement, but it is a property argued, not proved. (4) The price: a sudden departure of a third or more now pauses finality for a full window (30 days on mainnet) instead of 1.4 to 10 days; a gradual one costs nothing. the maintainers asked for the pause over the fork; the number is stated in spec 3.7 item 2. (5) The fold clock is in memory: a restarted node folds from daa(C_i) + depth, a few seconds late at worst. (6) Binaries, all from finality-fixes 6aa69a45, hashes and checks in the rollout plan's section 2: Mac native fe982a1d... (verified running), Linux 7c100fc2... (cargo-zigbuild, 34 min, not run on a Linux host), Windows cc1d1001... (mingw, 12 min 28 s, the v2 exe's DLL set, cannot run here); the Windows payload inputs were staged with push-inputs.sh --no-deploy into a scratch folder and NOT deployed (plan 7a).

+

4 October 2026, miner performance: variant racing (Metal worker on the M5 Max; the RTX 5090 job is ready, not run)

+

Method (docs/design/miner-tuning.md): at every hourly prepare the worker compiles the bound kernel in several variants (unroll, load path, register budget, threads per group, combinations), checks each bit for bit against the base kernel, times each for 2 s with the job loop paused, and keeps the fastest for the hour. Base is the kernel as it has always shipped. Code: proto-metal/main.swift (raceProgram, --race-test), proto-cuda/nvrtc/worker.cpp (racePair, --race), branch miner-perf, commit 460a99a.

+

Machine: Apple M5 Max, Darwin 25.6.0 (macOS 26.6.2), 64 GiB. CONDITIONS: the live Igneum Miner app's own Metal worker (igneum-bench --serve, pid 14687) was mining on the same GPU throughout, and the load average was 130 at the build and 14 to 67 during the races (other agents' cargo builds). The absolute MH/s below are therefore about half of the card's (the app reported 26.7 MH/s on 4 October with the GPU to itself) and each window was contended; the numbers to read are the ratios, taken as the best of three interleaved rounds per variant so the contention hits every variant alike. A re-run with the Apple M5 Max card paused is listed under "next".

+

Command (under the measure lock, which holds the build lock too):

+

tools/lock/with-lock.sh measure bash scratchpad/metal/measure.sh = swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal (47 s under load 130) igneum-bench --race-test --seed igneum-genesis --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000 igneum-bench --race-test --seed igneum-hourly --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000

+

Dataset 2^28 words (1 GiB, memory-hard, built in 285 and 295 ms), batch 2^22 nonces per launch, programs by the version-2 generator (128 loads per hash, no wide loads). 14 variants, every one bit-exact with base over 2^16 nonces (no variant discarded). MH/s = best of 3 rounds, 2 s windows, first launch of each window not counted.

+
variantthreads/groupmax threads/groupseed igneum-genesis MH/svs baseseed igneum-hourly MH/svs base
base32102411.180010.8000
g6464102411.582+3.6%11.801+9.3%
g128128102412.736+13.9%12.306+14.0%
g256256102413.114+17.3%13.089+21.2%
u232102410.683-4.4%10.596-1.9%
u832102410.823-3.2%10.280-4.8%
mt2563225610.953-2.0%10.664-1.3%
mt5123251211.022-1.4%10.636-1.5%
mt102432102410.666-4.6%11.150+3.2%
osize32102410.833-3.1%10.437-3.4%
u2-g128128102412.793+14.4%12.928+19.7%
u8-g128128102412.496+11.8%11.923+10.4%
mt256-g12812825612.628+13.0%12.675+17.4%
mt512-g25625651213.101+17.2%12.817+18.7%
+

Race cost: compile 1,798 ms (first seed; the Metal compiler cold) and 267 ms, timing 108 s for 14 variants x 3 rounds (2 s windows plus the 2^16-nonce check); in --serve the race runs one round, about 40 s, inside a 600-DAA lead, with mining paused only inside the windows.

+

Reading. On Apple silicon the win is threads per threadgroup: the Metal worker has dispatched one 32-thread group per 32 nonces since 3 October, and 256-thread groups are 17 to 21% faster on both programs under these conditions, with 128 close behind; the unroll, register-budget and size-optimisation knobs are within noise or worse on their own. The winner agrees across the two programs, so a tuning entry Apple_M5_Max: g256 would be the first fleet default; the race itself finds it in one round. These two programs are two points, under contention; the figure for the Apple M5 Max's own card with the GPU to itself is still to take. Nothing here says anything about NVIDIA: w8 (8 warps per block) is the CUDA cousin of g256, and whether the 5090 moves at all is what the RTX 5090 machine job (docs/plans/miner-perf.md) measures. Range to measure there: from no gain to what the block-size and load-path variants give on a 1 GiB random-read kernel; no claim.

+

Serve-protocol check (the same binary, --serve --race-rounds 1, scripted stdin: two inline jobs on pair A, the deferred race on A, prepare of pair B with its race, jobs on A meanwhile, the swap to B, a job across the 32-bit nonce boundary): 44 jobs done, 0 errors, no found line missed; the inline compile of pair A 220 ms, the deferred race on A winner g256 14.207 base 11.404 gain +24.58% (compile 563 ms, 36 s of windows); prepared for B after 35,554 ms = program 58 ms, dataset 524 ms, race 34,972 ms (winner mt512-g256 12.807 base 10.346 gain +23.79%), the swap to B in 0.01 ms, the 64-nonce job across the 32-bit boundary 14.8 ms. Found by this check: the job queued during a race waited for the whole race (job 2 done after 35,946 ms; the mutex is not fair), so the race now pauses 150 ms after every window (commit 32d1c01). Re-check with the pause (--serve, a job every 3 s through both races, load average 134 to 183): every job during the deferred race on A and the prepare race on B finished in 0.17 to 2.8 s (33 jobs, none over 2,831 ms, versus 35,946 ms before), the race on B 39.8 s inside a 40.2 s prepare, winner g256 both times, swap 0.01 ms, 0 errors; the race's own windows were 2 to 3 s longer in total than without the pause, as expected.

+

NVIDIA side, what the Apple M5 Max could check: proto-cuda/nvrtc/emu/test.sh PASS on the race build (the race off under emulation, "variants 1 base only, no race (emulation)" logged per pair; 9 source checks PASS, the --serve protocol with prepare, swap and self-heal unchanged, 17 sampled hashes equal to igneum-pow hash-bound); mingw cross-compile of igneum-worker-cuda.exe with the race (build-windows.sh, mingw, static): 1,509,376 bytes, the same imports as the shipped worker (KERNEL32 and the Universal CRT), icon and version block verified; zipped as igneum-worker-cuda-race.zip (429,387 bytes, sha256 321a086e...c4c049) for the RTX 5090 machine 1 job. The race has not run on a GPU.

+

Next: the RTX 5090 machine 1 job (ready in docs/plans/miner-perf.md); the Apple M5 Max card paused for a clean absolute table; the Mac app's own worker on this build (its hourly prepare then races by itself and logs the TUNING record).

Generated from the repository at build time. Times are UTC. Machine names are model names.

diff --git a/site/index.html b/site/index.html index 9e01d74cf..5b6d2f2a2 100644 --- a/site/index.html +++ b/site/index.html @@ -560,7 +560,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var( - +