igneum/docs/benchmarks/repro.md

32 KiB

The reproducible benchmark package

6 October 2026 (built on the evening of 5 October). One command per platform reproduces the numbers on the bench table on an outsider's own machine on day one: bench/repro.sh (Linux, macOS) and bench/repro.ps1 (Windows), shipped as igneum-repro-<tag>.tar.gz and .zip next to the public downloads (packaging/ota/publish-public.sh --repro, aliases /public/igneum-repro.tar.gz and /public/igneum-repro.zip; not deployed tonight). Package v0.1.0-repro was built from commit 39141f5 plus the package sources of branch repro-bench, commit 66ccd25 (the shipped binaries were built from that tree before the commit; the script fixes after the runs change no binary), and run end to end on the three project machines. The deltas against the bench log are below.

Why it exists: docs/evidence.md has 30 claims and none is reproduced externally. The ladder from "tested by the team" to "reproduced externally" needs "the command published and a third party's run with the same result". This is the command. The reward for running it is section 8.1 of docs/benchmarks/proving-e2e.md, quoted unchanged below.

1. What one run does

Step What runs Binary What it reports
1 The machine the workers' --list OS, CPU, memory; every GPU each worker sees (name, driver, memory)
2 The lottery hash vectors of the published genesis pack (proto-cuda/packs/igneum-genesis-mh: seed igneum-genesis, day 2026-10-03, generator 2, program id bcc1248b10cc90f2, memory-hard 1 GiB dataset) igneum-pow check-pack on the CPU; igneum-worker-cuda --bench, igneum-worker-opencl --bench --pack, igneum-bench --pack on every GPU bit-exact or not: 3 warps, 96 lanes, the 256 MiB cache FNV, on the CPU through the Rust interpreter and on every GPU through its own compiler (NVRTC, the vendor's OpenCL compiler, Metal)
3 The hash benchmark, 120 s per card at 1 GiB, --batch-log2 24 the same workers, --seconds 120 MH/s over the summed dispatch time, the hashes done, the fingerprint of the first 2^24 outputs at base nonce 0 (FNV-1a 64 over 16.7 million hashes: equal on two machines means every one of them agreed)
4 The random-read probe at 4, 64, 256 and 1024 MiB --memprobe dependent random 4-byte reads per second (the hash's access pattern) and their latency at 256 lanes, eight independent chains, 16 and 64-byte lines, the coalesced stream, an integer chain
5 The chip-resistance sweep: the same program at 4, 64, 256 and 1024 MiB, 5 batches each --bench --dataset-mib N (--sweep on Metal) MH/s per size; in-cache rate over the 1 GiB rate; the hash's share of the card's random-read ceiling at 1 GiB (MH/s x 128 loads against the chase)
6 One fixture shard proven with the pinned guest and verified (optional) igneum-prove-host --mode shard --shard 0, then --mode verify execute, core and compressed proof times, proof bytes, VERIFIED or not; skipped with the reason when there is no 12 GB NVIDIA card or no prover host
7 The result the script results/igneum-repro-<os>-<time>.json (format igneum-repro-1: machine, every binary's sha256, the pack files' sha256, every command run with its exit status, the rows in the bench table's own shape, the tolerances) and the same as a table in .md, plus the full log

The workers were extended for this (the flags are in every worker now, where the read-width experiment of 5 October had put --bench and --memprobe on the CUDA worker and --memprobe on the OpenCL worker, on branches): --list, --bench --pack <dir> --seconds N --dataset-mib N, --memprobe at 4, 64, 256 and 1024 MiB with one RESULT line per size and a DEVICE line per card; the Metal worker got --list, --pack (the published vectors against the GPU's warps), --seconds, --sweep and --memprobe; igneum-pow got check-pack. The pack reader accepts a string-seed pack (the genesis pack's 14-byte seed; the chain's are 32 bytes), the same change the read-width branch made.

What a run needs: no secrets, no node, no wallet, no network. The CUDA worker needs the NVIDIA driver only on Windows (NVIDIA's runtime compiler is in the package, THIRD-PARTY.md) and the driver plus libnvrtc.so.12 on Linux. The OpenCL worker needs the vendor's OpenCL. Nothing here earns anything.

2. The three runs of 5 October 2026 against the bench log

Package v0.1.0-repro, three machines asked for, two run tonight (the Mac in full, PC 2 with its 5090 shared with the live miner): PC 1 (ae432dc7: RTX 5090, RX 9070 XT, gfx1036) went down at 22:31:06 UTC on 5 October before its second slot (its first run at 21:04 UTC ended in a PowerShell parse error within a second per card, section 7) and nothing reaches it until the morning; its run is owed. The Mac ran under the measure lock (every build and simulation slot held; load average 6.19 at the start, 4.15 at the end). PC 2's run is in the coordinator's queue (section 2.2 is filled from it; "pending" until then).

2.1 The Mac (Apple M5 Max, 40 GPU cores, 64 GB, macOS 26.6.2), run 2fdd7365db4a5deb, 21:51 UTC, 5 October 2026

Result files: docs/benchmarks/repro-2026-10-06/mac/. 7 min 52 s for the two backends (Metal 120 s, Apple OpenCL 30 s as the cross-check, each with the sweep and the probe table).

Number The package The bench log Delta Reading
Vectors, CPU (Rust interpreter) 96/96 lanes, cache FNV 48c4f5bf24166b2e 96/96 (4 October 2026, generator version 2 adopted) exact
Vectors, Metal 96/96, program id bcc1248b10cc90f2 matches 96/96 (4 October 2026) exact
Vectors, Apple OpenCL 96/96 96/96 (proto-opencl README, 4 October 2026) exact
Fingerprint of 2^24 outputs at base 0 25f96e7dce90bd4e on Metal and on Apple OpenCL none at 2^24 for this pack (the README's f2a95d5bb84d961e is at 2^13) new: two compilers agree over 16.7 million nonces the number a second machine must match
Hash MH/s at 1 GiB, Metal, 120 s, GPU time 27.674 (194 dispatches of 2^24, 3.25 G hashes) 27.9 through Apple OpenCL on this pack (README, 5 x 2^24, wall); 26.7 mining through Metal on the live devnet (4 October) -0.8% against the bench, +3.6% against mining the bench row
Hash MH/s at 1 GiB, Apple OpenCL, 30 s, wall 26.875 27.9 (the same README figure) -3.7% outside the 3% tolerance against a 5-batch wall-time figure taken on 4 October; the cross-check path is 2.9% under Metal here, where the 3 October runs had the two Apple paths equal (45.2 against 45.0 on the version 1 program). Not the card's row
CPU verify, ms per 32-lane warp (avg of 50, one performance core) 0.611 (cold 0.625) 0.631 (4 October, version 2 units) -3% 16x under the 10 ms gate
Sweep, MH/s at 4 / 64 / 256 / 1024 MiB, Metal 275.2 / 103.2 / 56.6 / 27.8 569 / 183 / 94 / 44 (3 October, version 1 program with 104 loads, closed-form dataset) not comparable: the generator changed in-cache over 1 GiB 9.9x against 12.9x on version 1
Random reads at 1 GiB, G loads/s (chase) 3.40 Metal, 3.40 Apple OpenCL none for Apple before this (4.6 G was derived from the version 1 hash rate on 3 October) new the hash does 27.67 x 128 = 3.54 G loads/s, 1.04 of the single-chain chase: on this chip the hash is at the random-read ceiling the probe sees
Dependent-load latency, ns at 256 lanes 522 Metal (GPU time), 1,273 Apple OpenCL (wall, launch included) none new
Stream, GB/s at 1 GiB 546 Metal, 515 Apple OpenCL 427 GB/s dataset fill (3 October, a write) reads 28% above the write figure
Proving skipped: the GPU prover needs a 12 GB NVIDIA card as designed

2.2 PC 2 (RTX 5090, gfx1036, Windows 11), 1ccfe586

Run job run-repro-pc2-20261006 (fetch-repro-pc2-20261006 first), 00:44 to 01:09 UTC, 6 October 2026, on Igneum Miner 0.3.11, one card at a time through bench/jobs/pc-repro.ps1. What came back: every RESULT line in the log intake (the three per-card transcript logs are in docs/benchmarks/repro-2026-10-06/pc2/; the JSON and the markdown were never written, section 7). The 5090 block ran BESIDE THE APP'S LIVE CUDA WORKER: the card-off POST was accepted, but api/state answered {} (a 0.3.11 defect on a proving machine) so nothing confirmed a stopped worker, and nvidia-smi showed 10,176 MiB in use before the bench and the card at 348 W after it. Every 5090 rate below is therefore a shared-card figure, about half of what the card does alone (131 to 137 MH/s at this batch in every other run of the night), and is not the card's row. The checks (vectors, fingerprints, the proof) stand regardless of the sharing. The gfx1036 block is the integrated chip's own figure (3.312 MH/s, where it mines at 3.3 in the app).

Number The package (PC 2) The bench log Delta Reading
Vectors, CPU (Ryzen 7 9800X3D) 96/96, cache FNV 48c4f5bf24166b2e 96/96 on the M5 Max exact first CPU of the second architecture
Vectors, RTX 5090 through CUDA (NVRTC, sm_120) 96/96, program id matches the version 2 pack had not run on real NVIDIA silicon (evidence row 4) exact: closes that gap
Vectors, RTX 5090 through NVIDIA OpenCL 96/96 96/96 on the version 1 pack (3 October) exact
Vectors, gfx1036 through AMD OpenCL 96/96 96/96 on the version 1 pack (3 October) exact: the version 2 pack on real AMD silicon
Fingerprint of 2^24 outputs at base 0 25f96e7dce90bd4e on CUDA, NVIDIA OpenCL and AMD OpenCL 25f96e7dce90bd4e on Metal and Apple OpenCL (the Mac run) exact across five compilers and three vendors the cross-vendor claim over 16.7 million nonces
Hash MH/s at 1 GiB, 5090, CUDA, 120 s 62.412 (447 dispatches, 7.50 G hashes), beside the live miner 139.7 with the card to itself on a version 2 pack (4 October, variant racing); 124.2 mining in the app not comparable: shared card re-run owed with the worker confirmed stopped
Hash MH/s at 1 GiB, 5090, NVIDIA OpenCL, 30 s 62.283, beside the live miner 219.6 against 229.0 CUDA on the version 1 pack (3 October) not comparable the two paths agree with each other to 0.2% even shared
Hash MH/s at 1 GiB, gfx1036, AMD OpenCL, 120 s 3.312 (24 dispatches) 4.38 on the version 1 pack (104 loads, 3 October); 3.3 mining on the live devnet (4 October) +0.4% against mining; the version 1 figure is another program the card's row
Sweep 4 / 64 / 256 / 1024 MiB, 5090 CUDA (shared) 374 / 367 / 74 / 62 1,340 / 1,353 / 270 / 229 (3 October, version 1, card alone) shape only: in-cache over 1 GiB 6.0x against 5.8x the L2 edge is the same
Sweep, gfx1036 3.72 / 3.38 / 3.32 / 3.31 none new flat: the integrated chip is not cache-bound at any size, its memory is the system's
Random reads at 1 GiB, 5090 (chase, best lanes) 17.5 G loads/s CUDA (wall), 18.2 NVIDIA OpenCL (event); latency 415 ns (OpenCL event) 23.7 G derived from the version 1 hash rate (3 October); about 16 to 18 G implied by the 9070 XT entry's 6.6x (5 October) inside the implied range, shared card
Random reads at 1 GiB, gfx1036 0.458 G loads/s, 566 ns none new the hash does 3.312 x 128 = 0.424 G loads/s: 0.93 of the chase; at the ceiling like the 9070 XT (0.92)
Stream at 1 GiB, 5090 1,662 GB/s (OpenCL event), 339 (CUDA wall, the launch dominates) 1,638 GB/s dataset fill (3 October) +1.5% the memory clock was in its full state
Integer chain, 5090 13.4 T int ops/s (CUDA wall), 45.2 (OpenCL event) none the wall figure is launch-bound; the event figure is the card's
Proving, block-338-shard1 shard 0, the live host /opt/igneum/igneum-prove-host (pinned shard id 0x2b1a81cb...) execute 1.48 s (60.4 M cycles), core 20.2 s (18.1 MB, verify 0.60 s), compressed 33.9 s (1.27 MB, verify 0.038 s), VERIFIED; two tampered witnesses REJECTED the proving agent the same night: 10.8 to 11.4 s compressed alone, 33 s with the miner running the "with the miner" figure consistent with the shared card
Job size, 5090 CUDA (shared): 2^21 against 2^24 genesis pack 60.0 against 63.2 MH/s; live devnet pack 60.9 against 63.4 the app mines at 2^21: 114.0 wall against 116.0 inside jobs at 22:40 UTC (PC 2's STATUS lines, the consequences reviewer) 5% lower at 2^21 here, 1.8% wall-to-inside in the app the per-job cost at 2^21 is about 5% on a shared card; the 15% gap the reviewer found between bench and app is not this alone. The rows are PC 2's; PC 1's are owed

2.3 PC 1 (RTX 5090, RX 9070 XT, gfx1036, Windows 11), ae432dc7

Owed. The first job (run-repro-pc1-20261005, 21:04 to 21:09 UTC) switched each card off and on for nothing: repro.ps1 had three PowerShell parse errors (two missing parentheses on the wsl.exe lines, one (if ...) used as an expression), found only by the Windows parser on the PC because the Mac has no PowerShell. The fix, the parse gate (the PC job now parses the package script before any card is touched, and the bench/ folder is in the Windows PowerShell 5.1 parse job of .github/workflows/windows.yml) and the second package were ready at 21:45 UTC; PC 1 went down at 22:31:06 UTC before its second slot. One fact from the first job: the RX 9070 XT (gfx1201) was on the bus and switched off and on by the app at 21:06 UTC, where the coordinator's queue had it absent since 20:40 UTC.

3. The deltas, and what they mean for each user tier

What each number means for each user tier, and what the package does about it (the standing rule of 5 October 2026):

Number Home miner, one card (8, 12, 16, 24 or 32 GB) Rig Pool user What the package does
The bench needs 1.3 GB of device memory (1 GiB dataset, 256 MiB cache, a 128 MiB output buffer at 2^24) every tier runs it, 8 GB included; a 4 GB card or an integrated chip with 4 GB shared runs it too one card at a time, --only, so the rig keeps mining on the rest the same reports the free memory it found when a dataset does not fit
120 s per card, plus 3 short sweep sizes and the probe (about 4 min per card) a few minutes with the card idle a 6-card rig is 25 min, and --only splits it --seconds and the skip flags
The hash share of the random-read ceiling (1.04 on the M5 Max, PC 2 pending) the number that says a card is at its memory's limit, not the kernel's: a card under about 0.9 has a worker problem, not a card problem the same per card the probe table is in every result, so a low share names the step
Apple OpenCL 2.9% under Metal on the same chip an Apple user on the app gets the Metal rate (the app's worker is Metal); the OpenCL figure is a compiler cross-check, never the app's the OpenCL row is marked cross-check and makes no bench-table row
CPU verify 0.611 ms per warp on an M5 Max performance core the chain's gate is 10 ms on any core; a 2019-class laptop core is still unmeasured (evidence row 5, O-1.14) a pool verifying shares has 16x margin on this core check-pack prints the figure on any machine that runs the package; the first slow core to run it answers O-1.14
The proving step: skipped on macOS by design; on NVIDIA it needs a built prover host and a 12 GB card by the gate, and the 5090 run of the proving agent tonight peaked at 28.3 GB (bench-log, 5 October, "the 5090 alone proves block-338-shard1 shard 0 compressed in 10.8 to 11.4 s at a 28.3 GB peak") a 12 GB, 16 GB or 24 GB owner cannot run step 6 today: the gate lets them start it and the prover would fail on memory. The litepaper's "a 12 GB card proves a shard" stays "designed" (evidence row 16) until the prover's peak is under 12 GB step 6 says why it skipped; the next package raises the gate to the measured peak (28 GB) until the prover fits 12 GB, so nobody's run fails late. Filed below as owed work
One PC run per night at most on the project's own PCs the package itself takes minutes; the project's PCs are shared by many agents and PC 1 was down tonight, so the project's three-machine table is not complete on the first night. The outsider's run does not have that constraint

PC 2 (6 October 2026): the 5090's package numbers are void as the card's figures (beside the live miner) and stand only as checks; the morning re-runs the 5090 block alone with the worker confirmed stopped by nvidia-smi's compute-apps list (bench/jobs/pc-repro.ps1 does that now, and labels a run "UNCONFIRMED" when it cannot). For a 5090 owner the package's own figure is still owed; for a gfx1036 owner the figure is 3.31 MH/s, flat across dataset sizes, at 0.93 of the chip's random-read ceiling. For a proving owner: the live host proved and verified the fixture through the package's step 6 on a 32 GB card even with a miner on it (33.9 s compressed); the 12 GB gate stands as written in the table above.

4. Tolerances and how a result becomes a row

Number Tolerance Why this width
Vectors (3 warps, 96 lanes, cache FNV) exact a hash either matches or it does not
Fingerprint of 2^24 outputs at base 0 exact across machines at the same --batch-log2 the same
Hash MH/s at 1 GiB (120 s) 3% two 120-s runs of the same card on this evening's machines agree within about 1% with the card to itself; 3% leaves room for a driver version and a warmer card
Sweep MH/s at 4, 64, 256 MiB 10% 5 batches each, not 120 s; the in-cache sizes are the noisiest (a few hundred ms per batch)
Random reads at 1 GiB (G loads/s), stream GB/s 10% best of 3 at 8 lane counts; the memory clock's power state moves these
Dependent-load latency (ns at 256 lanes) 15% one warp, a few hundred ns, timer resolution
CPU verify per warp under 10 ms (the chain's gate), no tolerance on the figure itself any core must verify a warp inside the gate; the published cores are under 1 ms
Proof times (execute, core, compressed) 25% SP1's GPU prover varies run to run with the card's state and the server's warm-up

The path from a submitted file to the bench table (site/miner-bench.json, rendered at /miners by site/build.mjs):

  1. node bench/ingest.mjs check <result.json> prints the agreement table: every number against the team's reference for that card model (bench/reference.json, which names the bench-log entry or this document's run behind every reference), with the tolerance and the delta.
  2. node bench/ingest.mjs add <result.json> --by "reproduced externally" appends one row per card with that label when every number agrees, and refuses the label (recording "submitted" with the deltas in the note) when one does not. A failing run is as public as a passing one (proving-e2e.md 8.2 step 4). The site build then fails on any private string (site/forbidden-strings.txt), so a hostname in a note never reaches the page.
  3. The row's source is repro:<run id>, the random id the script drew; the result file is kept under docs/benchmarks/repro-<date>/, as the three of tonight are.

The relay's console (relay/api/console.mjs, results) shows bench entries synced from docs/bench-log.md by tools/console.mjs sync-bench; a submitted result enters there through the bench-log entry that records it, after the table row. The relay's bench role (relay/lib/relay.mjs) is for a machine of ours running benches, not for outsiders. Until the repository is public (the public testnet), results arrive by email to the address on igneum.network; after it, as an issue with the JSON attached.

5. Operator instructions

See bench/README.md, shipped in the package. In short: unpack, stop mining on the card, run the one command, send results/igneum-repro-<os>-<time>.json back, say which card was idle. About 4 minutes per card plus the optional proving step. The file carries no hostname, user name or address; it carries the GPU and CPU models, the OS and the driver version. --only backend:index runs one card; --no-prove, --no-probe, --no-sweep skip steps.

Our own machines run it the same way, with one difference the script does not know about: the card under test is switched off in the Igneum Miner app first (bench/jobs/pc-repro.ps1 does it through the app's api/cards, one card at a time, and restores it), and on the Mac the run takes the measure lock.

6. The reward terms (proving-e2e.md section 8.1, unchanged)

8.1 What counts as unrelated

Three operators are unrelated when every row holds for every pair:

Test Requirement
Person Different natural or legal persons; none is the project, an agent of it, or paid by it for the run (a published fixed reproduction reward, equal for everyone and announced before the run, is allowed and disclosed)
Hardware Bought separately; no shared host, card, rack or power meter
Network Different autonomous systems, verified by the IP in the published logs; not the same residential ISP account
Location Different physical sites
Software The same published release by hash; nobody receives a private build
Money No payment, loan or equipment between them or from the project, beyond the disclosed reproduction reward

An operator declares each row in their report and signs it with the vote key that mined on the devnet under the same fingerprint, so a report is tied to a key with a history.

The amount is the challenge reward row of docs/plans/funding.md (USD 1,000 per operator per workload set, three operators, about USD 3,000 per campaign, approximate; not funded as of 6 October 2026). Reproduction is asked for without a reward, which the standard allows. The result file of this package is signed by nothing; the signature with the vote key of 8.1 is the operator's own step, on the file.

7. What is unverified

Item Why What closes it
PC 1's run (RTX 5090, RX 9070 XT, gfx1036 on Windows) PC 1 went down at 22:31:06 UTC before its slot; the first job hit the parse errors the morning's slot: bench/jobs/pc-repro.ps1 as a run job after the package fetch, one card at a time
The Linux binaries (bin/linux-x86_64) cross-compiled with zig on the Mac, loaded nowhere tonight (no Linux GPU host; HiveOS is the first) a run on a Linux box with an NVIDIA or AMD driver; repro.sh is the same script the Mac ran
The CUDA worker's --bench, --memprobe and --list on a real card tonight's only CUDA runs were the Mac's emulation (--list, --bench on the genesis pack: self-test PASS, fingerprint e7d68ec2a49d0671 at 2^14, equal to Apple OpenCL's at 2^14) and PC 1's run that never reached the worker PC 2's slot (pending) and PC 1's morning run
repro.ps1 end to end parsed only by the Windows parser on PC 1 after the fix; never run to completion on a PC PC 2's slot
The proving step (step 6) never run through the script; on PC 2 it will use the live host /opt/igneum/igneum-prove-host as the app's WSL user with the live prover off PC 2's slot; then the memory gate against the measured 28.3 GB peak
A second machine of the same model the tolerance table has one Apple machine; "reproduced externally" needs an unrelated operator's run, and the repository is private until the public testnet the publish step (publish-public.sh --repro, not run tonight) and the first outside result
Apple OpenCL 3.7% under the README's 27.9 MH/s one 30-s wall-time run against one 5-batch run of 4 October; either could be the odd one a second 120-s run of each path on the Mac with the card to itself
The random-read probe on Apple as a ceiling the hash exceeds the single-chain chase by 4% on the M5 Max, so on Apple the chase at 4 M lanes is not the ceiling the 9070 XT entry took it for more lanes in flight, or the eight-chain probe (3.50 G here) as the Apple ceiling; a probe question, not a hash one
PC 2's 5090 rows the card-off step could not confirm a stopped worker (api/state answered {}), and the numbers are half the card's the morning's 5090-only run on PC 2 with the compute-apps confirmation; until then the rows are "beside the live miner"
PC 2's result files repro.ps1's close threw System.OutOfMemoryException in ConvertTo-Json and gave Set-Content an empty path: PowerShell variable names are case-insensitive, the result table $Cpu held $cpu (its own name string) and so contained itself, and the markdown lines $md wiped the markdown path $Md. Fixed (distinct names; bench/jobs/ps-case-check.sh fails CI on any case-only pair, shown to fire on a bad file); the RESULT lines in the intake and the three transcript logs are the record the morning's run writes the files
The prover on PC 2 the first script read the prover's state from api/state too, saw nothing, and left the live prover off from 00:49 to the restore job (run-prover-on-pc2-20261006); the script now reads settings.json and restores unconditionally done tonight; the rule in the job
The reward the amount and payer are docs/plans/funding.md, not funded; the terms are quoted unchanged Josh's decision

8. The vendor-share metric (Counter ASIC 3.0 item 7)

6 October 2026, branch ca3-reserve. The history audit's rank 7 (docs/analysis/asic-resistance-history.md section 4.3): a one-vendor fleet is a softer form of chip capture (lesson 8: Equihash's NVIDIA tilt, Ethash's balance), so the share of hash rate by vendor is published beside the benchmark and the AMD gap gets a 3.0 target. The observer computes it every 60 s into live_state.vendor_share (tools/observer/vendor-share.mjs; the one hook in observer.mjs; node --test tools/observer/vendor-share.test.mjs fires on a known-finished and a known-failed case; node tools/observer/vendor-share.mjs --dry prints the live reading and writes nothing).

Definition. Hash rate by vendor (nvidia, amd, apple, intel, unknown), two readings, each named for what it is:

Reading What it is What it covers What it misses
fleet-reported the sum of now=X MH/s wall from the newest STATUS line of every miner worker that uploaded to the log intake in the last 10 minutes, by the vendor in the worker's label (miner-nvidia-, miner-amd-, miner-mac-; miner-other- by the app's cards: line) only machines that upload to the intake: the project's own fleet and any app install carrying the intake key the rest of the network entirely; a worker that stopped uploading
chain-attributed each miner id's (vote_key_hash) share of the BLUE blocks of the last 10 minutes times the node's network hash-rate estimate, attributed to the vendor of the fleet worker that logged that vote key (identity N '...' vote_key_hash= in its first uploads), else "unknown" the whole network's blocks the vendor of any miner that is not a reporting worker (that is the "unknown" row, the honest hole); pending and red blocks are left out

Today's devnet, read-only from the live tables (dry mode, 07:36 UTC, 6 October 2026; nothing written, the live observer not restarted; the hook goes live with the merge):

Vendor fleet-reported MH/s (workers) chain-attributed share of blue blocks (10 min) chain-attributed MH/s at the 135.8 MH/s estimate Note
NVIDIA 120.61 (1: PC 2's RTX 5090, with the prover on the card) 0.986 (552 of 560) 133.86 PC 1's RTX 5090 was off the network at the reading (Josh's desk)
AMD 0 (0) 0 0 the RX 9070 XT is on PC 1, off at the reading
Apple 0 (0) 0 0 the M5 Max is paused for measurements
Intel 1.69 (1: the Windows laptop's UHD Graphics) 0.014 (8 of 560) 1.94 the first outside machine, which carries the intake key
unknown 0 0 0 coverage 1.00: every blue block of the window came from a reporting worker

At 07:21 UTC, before PC 1 went off, the same tables read 243 MH/s and 18 miner ids; the window then was NVIDIA about 0.9 (PC 1's and PC 2's 5090s) and AMD about 0.01 (the 9070 XT's first 3 blocks after its restart). The devnet today is the project's own machines, so the two readings agree; on a public testnet the chain-attributed "unknown" row is the number that matters, and the fleet-reported row is only the project's own share.

Rate per watt and per pound, from the bench rows that exist (every price approximate, UK list prices from memory, 6 October 2026):

Card MH/s (class v3) Power MH/s per W Price (approximate) MH/s per £ Source
RTX 5090 136.1 about 326 W (328.6 W peak with the prover, 5 October; the app's stability line today: PC 1 p95 307 W mining alone, PC 2 p95 338 W with the prover) 0.42 (at 326 W); 0.44 at 307 W about £2,000 0.068 bench-log "Counter ASIC 2.0, the numbers"; miner_logs stability lines
RX 9070 XT 18.6 to 19.2 OWED: no measured figure (the app logs no stability line for the Radeon; PC 1 not available today); the board's 304 W TBP is a list figure, approximate about 0.06 at the list TBP (approximate) about £600 0.031 bench-log; price approximate
Apple M5 Max 27.9 OWED: no powermetrics figure in the bench-log not computed the laptop, about £3,500 and up; not a GPU purchase 0.008 (approximate, against the laptop price) bench-log; price approximate
Intel UHD (laptop iGPU) 1.85 (the US laptop's STATUS line) not measured miner_logs

The 2.0 record put AMD at 2.2x worse per pound and 4.9x worse per watt than the 5090 (approximate); the rows above read 2.2x per pound and about 7x per watt at the list TBP, so the per-watt figure is the one the owed measurement must settle.

Consequence per tier:

Tier What a 90% NVIDIA share means What the project does
An AMD owner (RX 9070 XT, 16 GB) at equal network difficulty the card earns 19 / 136 = 0.14 of a 5090's income at about 0.3 of its price, so about half the income per pound; the share itself does not change that, but a 90% NVIDIA network sets the difficulty by NVIDIA's rate, and a chip built against one vendor's memory system would hit AMD owners first (lesson 8) the line-width question stays open per the plan (docs/plans/counter-asic-3.md: w16 closes nothing, w64 makes the 5090 bandwidth-bound); the 3.0 target for the AMD gap is the per-read gap (2.4 G against 17.5 G dependent reads per second) measured per driver release, with the first target a 2x closing by the public testnet or a stated reason it cannot close; the vendor share is published so the tilt is visible
An Apple owner (M-series) 27.9 MH/s is 0.2 of a 5090 on a machine bought for other reasons; the share means the Mac is a minority that no family in the reserve may cost more than 5% (1.13.2) the reserve order of item 6 puts the two families Apple pays for (perm, shfla) at R1 and R3 with the costs measured; the 5% rule is checked live before each unlock
An NVIDIA owner the majority sets the difficulty; a chip that beats NVIDIA is the only chip that matters the chip model and the bounty (bench-log "Counter ASIC 2.0, the numbers")
A rig or a pool a pool's share by vendor is what the chain-attributed reading cannot see (one vote key per pool) pools publish their own vendor mix or appear as "unknown"; the detector (item 4) reads the per-program spread per card model
The project a fleet-reported reading above the chain-attributed one means a reporting worker is not finding blocks (a stuck worker, the 4 October class); below it means miners outside the intake the daily report carries both readings from the merge on

9. Files

File What
bench/repro.sh, bench/repro.ps1 the operator commands
bench/README.md the package's README (operator instructions, tolerances, how to send a result back)
bench/make-package.sh builds the package from a tagged commit with the existing cross-build scripts (zig for Linux, mingw for Windows, swiftc and cc for macOS, cargo for the CPU tool)
bench/ingest.mjs, bench/reference.json the agreement check and the bench-table ingestion; the references
bench/jobs/pc-repro.ps1, bench/jobs/collect-results.mjs the PC job (one card at a time through the app) and the Mac-side extraction of its result files
docs/benchmarks/repro-2026-10-06/ the three result files of tonight, their tables and logs
packaging/ota/publish-public.sh --repro the public downloads entry (tar.gz, zip, their sha256, the two aliases, the downloads index)