32 KiB
The reproducible benchmark package
6 October 2026 (built on the evening of 5 October). One command per platform reproduces the numbers on the bench table
on an outsider's own machine on day one: bench/repro.sh (Linux, macOS) and bench/repro.ps1 (Windows), shipped as
igneum-repro-<tag>.tar.gz and .zip next to the public downloads (packaging/ota/publish-public.sh --repro, aliases
/public/igneum-repro.tar.gz and /public/igneum-repro.zip; not deployed tonight). Package v0.1.0-repro was built from
commit 39141f5 plus the package sources of branch repro-bench, commit 66ccd25 (the shipped binaries were built from that
tree before the commit; the script fixes after the runs change no binary),
and run end to end on the three project machines. The deltas against the bench log are below.
Why it exists: docs/evidence.md has 30 claims and none is reproduced externally. The ladder from "tested by the team"
to "reproduced externally" needs "the command published and a third party's run with the same result". This is the
command. The reward for running it is section 8.1 of docs/benchmarks/proving-e2e.md, quoted unchanged below.
1. What one run does
| Step | What runs | Binary | What it reports |
|---|---|---|---|
| 1 | The machine | the workers' --list |
OS, CPU, memory; every GPU each worker sees (name, driver, memory) |
| 2 | The lottery hash vectors of the published genesis pack (proto-cuda/packs/igneum-genesis-mh: seed igneum-genesis, day 2026-10-03, generator 2, program id bcc1248b10cc90f2, memory-hard 1 GiB dataset) |
igneum-pow check-pack on the CPU; igneum-worker-cuda --bench, igneum-worker-opencl --bench --pack, igneum-bench --pack on every GPU |
bit-exact or not: 3 warps, 96 lanes, the 256 MiB cache FNV, on the CPU through the Rust interpreter and on every GPU through its own compiler (NVRTC, the vendor's OpenCL compiler, Metal) |
| 3 | The hash benchmark, 120 s per card at 1 GiB, --batch-log2 24 |
the same workers, --seconds 120 |
MH/s over the summed dispatch time, the hashes done, the fingerprint of the first 2^24 outputs at base nonce 0 (FNV-1a 64 over 16.7 million hashes: equal on two machines means every one of them agreed) |
| 4 | The random-read probe at 4, 64, 256 and 1024 MiB | --memprobe |
dependent random 4-byte reads per second (the hash's access pattern) and their latency at 256 lanes, eight independent chains, 16 and 64-byte lines, the coalesced stream, an integer chain |
| 5 | The chip-resistance sweep: the same program at 4, 64, 256 and 1024 MiB, 5 batches each | --bench --dataset-mib N (--sweep on Metal) |
MH/s per size; in-cache rate over the 1 GiB rate; the hash's share of the card's random-read ceiling at 1 GiB (MH/s x 128 loads against the chase) |
| 6 | One fixture shard proven with the pinned guest and verified (optional) | igneum-prove-host --mode shard --shard 0, then --mode verify |
execute, core and compressed proof times, proof bytes, VERIFIED or not; skipped with the reason when there is no 12 GB NVIDIA card or no prover host |
| 7 | The result | the script | results/igneum-repro-<os>-<time>.json (format igneum-repro-1: machine, every binary's sha256, the pack files' sha256, every command run with its exit status, the rows in the bench table's own shape, the tolerances) and the same as a table in .md, plus the full log |
The workers were extended for this (the flags are in every worker now, where the read-width experiment of 5 October had
put --bench and --memprobe on the CUDA worker and --memprobe on the OpenCL worker, on branches): --list, --bench --pack <dir> --seconds N --dataset-mib N, --memprobe at 4, 64, 256 and 1024 MiB with one RESULT line per size and
a DEVICE line per card; the Metal worker got --list, --pack (the published vectors against the GPU's warps),
--seconds, --sweep and --memprobe; igneum-pow got check-pack. The pack reader accepts a string-seed pack
(the genesis pack's 14-byte seed; the chain's are 32 bytes), the same change the read-width branch made.
What a run needs: no secrets, no node, no wallet, no network. The CUDA worker needs the NVIDIA driver only on Windows
(NVIDIA's runtime compiler is in the package, THIRD-PARTY.md) and the driver plus libnvrtc.so.12 on Linux. The
OpenCL worker needs the vendor's OpenCL. Nothing here earns anything.
2. The three runs of 5 October 2026 against the bench log
Package v0.1.0-repro, three machines asked for, two run tonight (the Mac in full, PC 2 with its 5090 shared with the live miner): PC 1 (ae432dc7: RTX 5090, RX 9070 XT, gfx1036) went down at 22:31:06 UTC on 5 October before its second slot (its first run at 21:04 UTC ended in a PowerShell parse error within a second per card, section 7) and nothing reaches it until the morning; its run is owed. The Mac ran under the measure lock (every build and simulation slot held; load average 6.19 at the start, 4.15 at the end). PC 2's run is in the coordinator's queue (section 2.2 is filled from it; "pending" until then).
2.1 The Mac (Apple M5 Max, 40 GPU cores, 64 GB, macOS 26.6.2), run 2fdd7365db4a5deb, 21:51 UTC, 5 October 2026
Result files: docs/benchmarks/repro-2026-10-06/mac/. 7 min 52 s for the two backends (Metal 120 s, Apple OpenCL 30 s as
the cross-check, each with the sweep and the probe table).
| Number | The package | The bench log | Delta | Reading |
|---|---|---|---|---|
| Vectors, CPU (Rust interpreter) | 96/96 lanes, cache FNV 48c4f5bf24166b2e |
96/96 (4 October 2026, generator version 2 adopted) | exact | |
| Vectors, Metal | 96/96, program id bcc1248b10cc90f2 matches |
96/96 (4 October 2026) | exact | |
| Vectors, Apple OpenCL | 96/96 | 96/96 (proto-opencl README, 4 October 2026) | exact | |
| Fingerprint of 2^24 outputs at base 0 | 25f96e7dce90bd4e on Metal and on Apple OpenCL |
none at 2^24 for this pack (the README's f2a95d5bb84d961e is at 2^13) |
new: two compilers agree over 16.7 million nonces | the number a second machine must match |
| Hash MH/s at 1 GiB, Metal, 120 s, GPU time | 27.674 (194 dispatches of 2^24, 3.25 G hashes) | 27.9 through Apple OpenCL on this pack (README, 5 x 2^24, wall); 26.7 mining through Metal on the live devnet (4 October) | -0.8% against the bench, +3.6% against mining | the bench row |
| Hash MH/s at 1 GiB, Apple OpenCL, 30 s, wall | 26.875 | 27.9 (the same README figure) | -3.7% | outside the 3% tolerance against a 5-batch wall-time figure taken on 4 October; the cross-check path is 2.9% under Metal here, where the 3 October runs had the two Apple paths equal (45.2 against 45.0 on the version 1 program). Not the card's row |
| CPU verify, ms per 32-lane warp (avg of 50, one performance core) | 0.611 (cold 0.625) | 0.631 (4 October, version 2 units) | -3% | 16x under the 10 ms gate |
| Sweep, MH/s at 4 / 64 / 256 / 1024 MiB, Metal | 275.2 / 103.2 / 56.6 / 27.8 | 569 / 183 / 94 / 44 (3 October, version 1 program with 104 loads, closed-form dataset) | not comparable: the generator changed | in-cache over 1 GiB 9.9x against 12.9x on version 1 |
| Random reads at 1 GiB, G loads/s (chase) | 3.40 Metal, 3.40 Apple OpenCL | none for Apple before this (4.6 G was derived from the version 1 hash rate on 3 October) | new | the hash does 27.67 x 128 = 3.54 G loads/s, 1.04 of the single-chain chase: on this chip the hash is at the random-read ceiling the probe sees |
| Dependent-load latency, ns at 256 lanes | 522 Metal (GPU time), 1,273 Apple OpenCL (wall, launch included) | none | new | |
| Stream, GB/s at 1 GiB | 546 Metal, 515 Apple OpenCL | 427 GB/s dataset fill (3 October, a write) | reads 28% above the write figure | |
| Proving | skipped: the GPU prover needs a 12 GB NVIDIA card | as designed |
2.2 PC 2 (RTX 5090, gfx1036, Windows 11), 1ccfe586
Run job run-repro-pc2-20261006 (fetch-repro-pc2-20261006 first), 00:44 to 01:09 UTC, 6 October 2026, on Igneum Miner
0.3.11, one card at a time through bench/jobs/pc-repro.ps1. What came back: every RESULT line in the log intake (the
three per-card transcript logs are in docs/benchmarks/repro-2026-10-06/pc2/; the JSON and the markdown were never
written, section 7). The 5090 block ran BESIDE THE APP'S LIVE CUDA WORKER: the card-off POST was accepted, but
api/state answered {} (a 0.3.11 defect on a proving machine) so nothing confirmed a stopped worker, and nvidia-smi
showed 10,176 MiB in use before the bench and the card at 348 W after it. Every 5090 rate below is therefore a shared-card
figure, about half of what the card does alone (131 to 137 MH/s at this batch in every other run of the night), and is
not the card's row. The checks (vectors, fingerprints, the proof) stand regardless of the sharing. The gfx1036 block is the
integrated chip's own figure (3.312 MH/s, where it mines at 3.3 in the app).
| Number | The package (PC 2) | The bench log | Delta | Reading |
|---|---|---|---|---|
| Vectors, CPU (Ryzen 7 9800X3D) | 96/96, cache FNV 48c4f5bf24166b2e |
96/96 on the M5 Max | exact | first CPU of the second architecture |
| Vectors, RTX 5090 through CUDA (NVRTC, sm_120) | 96/96, program id matches | the version 2 pack had not run on real NVIDIA silicon (evidence row 4) | exact: closes that gap | |
| Vectors, RTX 5090 through NVIDIA OpenCL | 96/96 | 96/96 on the version 1 pack (3 October) | exact | |
| Vectors, gfx1036 through AMD OpenCL | 96/96 | 96/96 on the version 1 pack (3 October) | exact: the version 2 pack on real AMD silicon | |
| Fingerprint of 2^24 outputs at base 0 | 25f96e7dce90bd4e on CUDA, NVIDIA OpenCL and AMD OpenCL |
25f96e7dce90bd4e on Metal and Apple OpenCL (the Mac run) |
exact across five compilers and three vendors | the cross-vendor claim over 16.7 million nonces |
| Hash MH/s at 1 GiB, 5090, CUDA, 120 s | 62.412 (447 dispatches, 7.50 G hashes), beside the live miner | 139.7 with the card to itself on a version 2 pack (4 October, variant racing); 124.2 mining in the app | not comparable: shared card | re-run owed with the worker confirmed stopped |
| Hash MH/s at 1 GiB, 5090, NVIDIA OpenCL, 30 s | 62.283, beside the live miner | 219.6 against 229.0 CUDA on the version 1 pack (3 October) | not comparable | the two paths agree with each other to 0.2% even shared |
| Hash MH/s at 1 GiB, gfx1036, AMD OpenCL, 120 s | 3.312 (24 dispatches) | 4.38 on the version 1 pack (104 loads, 3 October); 3.3 mining on the live devnet (4 October) | +0.4% against mining; the version 1 figure is another program | the card's row |
| Sweep 4 / 64 / 256 / 1024 MiB, 5090 CUDA (shared) | 374 / 367 / 74 / 62 | 1,340 / 1,353 / 270 / 229 (3 October, version 1, card alone) | shape only: in-cache over 1 GiB 6.0x against 5.8x | the L2 edge is the same |
| Sweep, gfx1036 | 3.72 / 3.38 / 3.32 / 3.31 | none | new | flat: the integrated chip is not cache-bound at any size, its memory is the system's |
| Random reads at 1 GiB, 5090 (chase, best lanes) | 17.5 G loads/s CUDA (wall), 18.2 NVIDIA OpenCL (event); latency 415 ns (OpenCL event) | 23.7 G derived from the version 1 hash rate (3 October); about 16 to 18 G implied by the 9070 XT entry's 6.6x (5 October) | inside the implied range, shared card | |
| Random reads at 1 GiB, gfx1036 | 0.458 G loads/s, 566 ns | none | new | the hash does 3.312 x 128 = 0.424 G loads/s: 0.93 of the chase; at the ceiling like the 9070 XT (0.92) |
| Stream at 1 GiB, 5090 | 1,662 GB/s (OpenCL event), 339 (CUDA wall, the launch dominates) | 1,638 GB/s dataset fill (3 October) | +1.5% | the memory clock was in its full state |
| Integer chain, 5090 | 13.4 T int ops/s (CUDA wall), 45.2 (OpenCL event) | none | the wall figure is launch-bound; the event figure is the card's | |
Proving, block-338-shard1 shard 0, the live host /opt/igneum/igneum-prove-host (pinned shard id 0x2b1a81cb...) |
execute 1.48 s (60.4 M cycles), core 20.2 s (18.1 MB, verify 0.60 s), compressed 33.9 s (1.27 MB, verify 0.038 s), VERIFIED; two tampered witnesses REJECTED | the proving agent the same night: 10.8 to 11.4 s compressed alone, 33 s with the miner running | the "with the miner" figure | consistent with the shared card |
| Job size, 5090 CUDA (shared): 2^21 against 2^24 | genesis pack 60.0 against 63.2 MH/s; live devnet pack 60.9 against 63.4 | the app mines at 2^21: 114.0 wall against 116.0 inside jobs at 22:40 UTC (PC 2's STATUS lines, the consequences reviewer) | 5% lower at 2^21 here, 1.8% wall-to-inside in the app | the per-job cost at 2^21 is about 5% on a shared card; the 15% gap the reviewer found between bench and app is not this alone. The rows are PC 2's; PC 1's are owed |
2.3 PC 1 (RTX 5090, RX 9070 XT, gfx1036, Windows 11), ae432dc7
Owed. The first job (run-repro-pc1-20261005, 21:04 to 21:09 UTC) switched each card off and on for nothing: repro.ps1
had three PowerShell parse errors (two missing parentheses on the wsl.exe lines, one (if ...) used as an expression),
found only by the Windows parser on the PC because the Mac has no PowerShell. The fix, the parse gate (the PC job now
parses the package script before any card is touched, and the bench/ folder is in the Windows PowerShell 5.1 parse job
of .github/workflows/windows.yml) and the second package were ready at 21:45 UTC; PC 1 went down at 22:31:06 UTC before
its second slot. One fact from the first job: the RX 9070 XT (gfx1201) was on the bus and switched off and on by the
app at 21:06 UTC, where the coordinator's queue had it absent since 20:40 UTC.
3. The deltas, and what they mean for each user tier
What each number means for each user tier, and what the package does about it (the standing rule of 5 October 2026):
| Number | Home miner, one card (8, 12, 16, 24 or 32 GB) | Rig | Pool user | What the package does |
|---|---|---|---|---|
| The bench needs 1.3 GB of device memory (1 GiB dataset, 256 MiB cache, a 128 MiB output buffer at 2^24) | every tier runs it, 8 GB included; a 4 GB card or an integrated chip with 4 GB shared runs it too | one card at a time, --only, so the rig keeps mining on the rest |
the same | reports the free memory it found when a dataset does not fit |
| 120 s per card, plus 3 short sweep sizes and the probe (about 4 min per card) | a few minutes with the card idle | a 6-card rig is 25 min, and --only splits it |
--seconds and the skip flags |
|
| The hash share of the random-read ceiling (1.04 on the M5 Max, PC 2 pending) | the number that says a card is at its memory's limit, not the kernel's: a card under about 0.9 has a worker problem, not a card problem | the same per card | the probe table is in every result, so a low share names the step | |
| Apple OpenCL 2.9% under Metal on the same chip | an Apple user on the app gets the Metal rate (the app's worker is Metal); the OpenCL figure is a compiler cross-check, never the app's | the OpenCL row is marked cross-check and makes no bench-table row | ||
| CPU verify 0.611 ms per warp on an M5 Max performance core | the chain's gate is 10 ms on any core; a 2019-class laptop core is still unmeasured (evidence row 5, O-1.14) | a pool verifying shares has 16x margin on this core | check-pack prints the figure on any machine that runs the package; the first slow core to run it answers O-1.14 |
|
| The proving step: skipped on macOS by design; on NVIDIA it needs a built prover host and a 12 GB card by the gate, and the 5090 run of the proving agent tonight peaked at 28.3 GB (bench-log, 5 October, "the 5090 alone proves block-338-shard1 shard 0 compressed in 10.8 to 11.4 s at a 28.3 GB peak") | a 12 GB, 16 GB or 24 GB owner cannot run step 6 today: the gate lets them start it and the prover would fail on memory. The litepaper's "a 12 GB card proves a shard" stays "designed" (evidence row 16) until the prover's peak is under 12 GB | step 6 says why it skipped; the next package raises the gate to the measured peak (28 GB) until the prover fits 12 GB, so nobody's run fails late. Filed below as owed work | ||
| One PC run per night at most on the project's own PCs | the package itself takes minutes; the project's PCs are shared by many agents and PC 1 was down tonight, so the project's three-machine table is not complete on the first night. The outsider's run does not have that constraint |
PC 2 (6 October 2026): the 5090's package numbers are void as the card's figures (beside the live miner) and stand only
as checks; the morning re-runs the 5090 block alone with the worker confirmed stopped by nvidia-smi's compute-apps list
(bench/jobs/pc-repro.ps1 does that now, and labels a run "UNCONFIRMED" when it cannot). For a 5090 owner the package's
own figure is still owed; for a gfx1036 owner the figure is 3.31 MH/s, flat across dataset sizes, at 0.93 of the chip's
random-read ceiling. For a proving owner: the live host proved and verified the fixture through the package's step 6 on a
32 GB card even with a miner on it (33.9 s compressed); the 12 GB gate stands as written in the table above.
4. Tolerances and how a result becomes a row
| Number | Tolerance | Why this width |
|---|---|---|
| Vectors (3 warps, 96 lanes, cache FNV) | exact | a hash either matches or it does not |
| Fingerprint of 2^24 outputs at base 0 | exact across machines at the same --batch-log2 |
the same |
| Hash MH/s at 1 GiB (120 s) | 3% | two 120-s runs of the same card on this evening's machines agree within about 1% with the card to itself; 3% leaves room for a driver version and a warmer card |
| Sweep MH/s at 4, 64, 256 MiB | 10% | 5 batches each, not 120 s; the in-cache sizes are the noisiest (a few hundred ms per batch) |
| Random reads at 1 GiB (G loads/s), stream GB/s | 10% | best of 3 at 8 lane counts; the memory clock's power state moves these |
| Dependent-load latency (ns at 256 lanes) | 15% | one warp, a few hundred ns, timer resolution |
| CPU verify per warp | under 10 ms (the chain's gate), no tolerance on the figure itself | any core must verify a warp inside the gate; the published cores are under 1 ms |
| Proof times (execute, core, compressed) | 25% | SP1's GPU prover varies run to run with the card's state and the server's warm-up |
The path from a submitted file to the bench table (site/miner-bench.json, rendered at /miners by site/build.mjs):
node bench/ingest.mjs check <result.json>prints the agreement table: every number against the team's reference for that card model (bench/reference.json, which names the bench-log entry or this document's run behind every reference), with the tolerance and the delta.node bench/ingest.mjs add <result.json> --by "reproduced externally"appends one row per card with that label when every number agrees, and refuses the label (recording "submitted" with the deltas in the note) when one does not. A failing run is as public as a passing one (proving-e2e.md 8.2 step 4). The site build then fails on any private string (site/forbidden-strings.txt), so a hostname in a note never reaches the page.- The row's
sourceisrepro:<run id>, the random id the script drew; the result file is kept underdocs/benchmarks/repro-<date>/, as the three of tonight are.
The relay's console (relay/api/console.mjs, results) shows bench entries synced from docs/bench-log.md by
tools/console.mjs sync-bench; a submitted result enters there through the bench-log entry that records it, after the
table row. The relay's bench role (relay/lib/relay.mjs) is for a machine of ours running benches, not for outsiders.
Until the repository is public (the public testnet), results arrive by email to the address on igneum.network; after it,
as an issue with the JSON attached.
5. Operator instructions
See bench/README.md, shipped in the package. In short: unpack, stop mining on the card, run the one command, send
results/igneum-repro-<os>-<time>.json back, say which card was idle. About 4 minutes per card plus the optional
proving step. The file carries no hostname, user name or address; it carries the GPU and CPU models, the OS and the
driver version. --only backend:index runs one card; --no-prove, --no-probe, --no-sweep skip steps.
Our own machines run it the same way, with one difference the script does not know about: the card under test is
switched off in the Igneum Miner app first (bench/jobs/pc-repro.ps1 does it through the app's api/cards, one card at
a time, and restores it), and on the Mac the run takes the measure lock.
6. The reward terms (proving-e2e.md section 8.1, unchanged)
8.1 What counts as unrelated
Three operators are unrelated when every row holds for every pair:
| Test | Requirement |
|---|---|
| Person | Different natural or legal persons; none is the project, an agent of it, or paid by it for the run (a published fixed reproduction reward, equal for everyone and announced before the run, is allowed and disclosed) |
| Hardware | Bought separately; no shared host, card, rack or power meter |
| Network | Different autonomous systems, verified by the IP in the published logs; not the same residential ISP account |
| Location | Different physical sites |
| Software | The same published release by hash; nobody receives a private build |
| Money | No payment, loan or equipment between them or from the project, beyond the disclosed reproduction reward |
An operator declares each row in their report and signs it with the vote key that mined on the devnet under the same fingerprint, so a report is tied to a key with a history.
The amount is the challenge reward row of docs/plans/funding.md (USD 1,000 per operator per workload set, three
operators, about USD 3,000 per campaign, approximate; not funded as of 6 October 2026). Reproduction is asked for without
a reward, which the standard allows. The result file of this package is signed by nothing; the signature with the vote
key of 8.1 is the operator's own step, on the file.
7. What is unverified
| Item | Why | What closes it |
|---|---|---|
| PC 1's run (RTX 5090, RX 9070 XT, gfx1036 on Windows) | PC 1 went down at 22:31:06 UTC before its slot; the first job hit the parse errors | the morning's slot: bench/jobs/pc-repro.ps1 as a run job after the package fetch, one card at a time |
The Linux binaries (bin/linux-x86_64) |
cross-compiled with zig on the Mac, loaded nowhere tonight (no Linux GPU host; HiveOS is the first) | a run on a Linux box with an NVIDIA or AMD driver; repro.sh is the same script the Mac ran |
The CUDA worker's --bench, --memprobe and --list on a real card |
tonight's only CUDA runs were the Mac's emulation (--list, --bench on the genesis pack: self-test PASS, fingerprint e7d68ec2a49d0671 at 2^14, equal to Apple OpenCL's at 2^14) and PC 1's run that never reached the worker |
PC 2's slot (pending) and PC 1's morning run |
repro.ps1 end to end |
parsed only by the Windows parser on PC 1 after the fix; never run to completion on a PC | PC 2's slot |
| The proving step (step 6) | never run through the script; on PC 2 it will use the live host /opt/igneum/igneum-prove-host as the app's WSL user with the live prover off |
PC 2's slot; then the memory gate against the measured 28.3 GB peak |
| A second machine of the same model | the tolerance table has one Apple machine; "reproduced externally" needs an unrelated operator's run, and the repository is private until the public testnet | the publish step (publish-public.sh --repro, not run tonight) and the first outside result |
| Apple OpenCL 3.7% under the README's 27.9 MH/s | one 30-s wall-time run against one 5-batch run of 4 October; either could be the odd one | a second 120-s run of each path on the Mac with the card to itself |
| The random-read probe on Apple as a ceiling | the hash exceeds the single-chain chase by 4% on the M5 Max, so on Apple the chase at 4 M lanes is not the ceiling the 9070 XT entry took it for | more lanes in flight, or the eight-chain probe (3.50 G here) as the Apple ceiling; a probe question, not a hash one |
| PC 2's 5090 rows | the card-off step could not confirm a stopped worker (api/state answered {}), and the numbers are half the card's |
the morning's 5090-only run on PC 2 with the compute-apps confirmation; until then the rows are "beside the live miner" |
| PC 2's result files | repro.ps1's close threw System.OutOfMemoryException in ConvertTo-Json and gave Set-Content an empty path: PowerShell variable names are case-insensitive, the result table $Cpu held $cpu (its own name string) and so contained itself, and the markdown lines $md wiped the markdown path $Md. Fixed (distinct names; bench/jobs/ps-case-check.sh fails CI on any case-only pair, shown to fire on a bad file); the RESULT lines in the intake and the three transcript logs are the record |
the morning's run writes the files |
| The prover on PC 2 | the first script read the prover's state from api/state too, saw nothing, and left the live prover off from 00:49 to the restore job (run-prover-on-pc2-20261006); the script now reads settings.json and restores unconditionally |
done tonight; the rule in the job |
| The reward | the amount and payer are docs/plans/funding.md, not funded; the terms are quoted unchanged |
Josh's decision |
8. The vendor-share metric (Counter ASIC 3.0 item 7)
6 October 2026, branch ca3-reserve. The history audit's rank 7 (docs/analysis/asic-resistance-history.md section 4.3): a one-vendor fleet is a softer form of chip capture (lesson 8: Equihash's NVIDIA tilt, Ethash's balance), so the share of hash rate by vendor is published beside the benchmark and the AMD gap gets a 3.0 target. The observer computes it every 60 s into live_state.vendor_share (tools/observer/vendor-share.mjs; the one hook in observer.mjs; node --test tools/observer/vendor-share.test.mjs fires on a known-finished and a known-failed case; node tools/observer/vendor-share.mjs --dry prints the live reading and writes nothing).
Definition. Hash rate by vendor (nvidia, amd, apple, intel, unknown), two readings, each named for what it is:
| Reading | What it is | What it covers | What it misses |
|---|---|---|---|
| fleet-reported | the sum of now=X MH/s wall from the newest STATUS line of every miner worker that uploaded to the log intake in the last 10 minutes, by the vendor in the worker's label (miner-nvidia-, miner-amd-, miner-mac-; miner-other- by the app's cards: line) |
only machines that upload to the intake: the project's own fleet and any app install carrying the intake key | the rest of the network entirely; a worker that stopped uploading |
| chain-attributed | each miner id's (vote_key_hash) share of the BLUE blocks of the last 10 minutes times the node's network hash-rate estimate, attributed to the vendor of the fleet worker that logged that vote key (identity N '...' vote_key_hash= in its first uploads), else "unknown" |
the whole network's blocks | the vendor of any miner that is not a reporting worker (that is the "unknown" row, the honest hole); pending and red blocks are left out |
Today's devnet, read-only from the live tables (dry mode, 07:36 UTC, 6 October 2026; nothing written, the live observer not restarted; the hook goes live with the merge):
| Vendor | fleet-reported MH/s (workers) | chain-attributed share of blue blocks (10 min) | chain-attributed MH/s at the 135.8 MH/s estimate | Note |
|---|---|---|---|---|
| NVIDIA | 120.61 (1: PC 2's RTX 5090, with the prover on the card) | 0.986 (552 of 560) | 133.86 | PC 1's RTX 5090 was off the network at the reading (Josh's desk) |
| AMD | 0 (0) | 0 | 0 | the RX 9070 XT is on PC 1, off at the reading |
| Apple | 0 (0) | 0 | 0 | the M5 Max is paused for measurements |
| Intel | 1.69 (1: the Windows laptop's UHD Graphics) | 0.014 (8 of 560) | 1.94 | the first outside machine, which carries the intake key |
| unknown | 0 | 0 | 0 | coverage 1.00: every blue block of the window came from a reporting worker |
At 07:21 UTC, before PC 1 went off, the same tables read 243 MH/s and 18 miner ids; the window then was NVIDIA about 0.9 (PC 1's and PC 2's 5090s) and AMD about 0.01 (the 9070 XT's first 3 blocks after its restart). The devnet today is the project's own machines, so the two readings agree; on a public testnet the chain-attributed "unknown" row is the number that matters, and the fleet-reported row is only the project's own share.
Rate per watt and per pound, from the bench rows that exist (every price approximate, UK list prices from memory, 6 October 2026):
| Card | MH/s (class v3) | Power | MH/s per W | Price (approximate) | MH/s per £ | Source |
|---|---|---|---|---|---|---|
| RTX 5090 | 136.1 | about 326 W (328.6 W peak with the prover, 5 October; the app's stability line today: PC 1 p95 307 W mining alone, PC 2 p95 338 W with the prover) | 0.42 (at 326 W); 0.44 at 307 W | about £2,000 | 0.068 | bench-log "Counter ASIC 2.0, the numbers"; miner_logs stability lines |
| RX 9070 XT | 18.6 to 19.2 | OWED: no measured figure (the app logs no stability line for the Radeon; PC 1 not available today); the board's 304 W TBP is a list figure, approximate | about 0.06 at the list TBP (approximate) | about £600 | 0.031 | bench-log; price approximate |
| Apple M5 Max | 27.9 | OWED: no powermetrics figure in the bench-log | not computed | the laptop, about £3,500 and up; not a GPU purchase | 0.008 (approximate, against the laptop price) | bench-log; price approximate |
| Intel UHD (laptop iGPU) | 1.85 (the US laptop's STATUS line) | not measured | miner_logs |
The 2.0 record put AMD at 2.2x worse per pound and 4.9x worse per watt than the 5090 (approximate); the rows above read 2.2x per pound and about 7x per watt at the list TBP, so the per-watt figure is the one the owed measurement must settle.
Consequence per tier:
| Tier | What a 90% NVIDIA share means | What the project does |
|---|---|---|
| An AMD owner (RX 9070 XT, 16 GB) | at equal network difficulty the card earns 19 / 136 = 0.14 of a 5090's income at about 0.3 of its price, so about half the income per pound; the share itself does not change that, but a 90% NVIDIA network sets the difficulty by NVIDIA's rate, and a chip built against one vendor's memory system would hit AMD owners first (lesson 8) | the line-width question stays open per the plan (docs/plans/counter-asic-3.md: w16 closes nothing, w64 makes the 5090 bandwidth-bound); the 3.0 target for the AMD gap is the per-read gap (2.4 G against 17.5 G dependent reads per second) measured per driver release, with the first target a 2x closing by the public testnet or a stated reason it cannot close; the vendor share is published so the tilt is visible |
| An Apple owner (M-series) | 27.9 MH/s is 0.2 of a 5090 on a machine bought for other reasons; the share means the Mac is a minority that no family in the reserve may cost more than 5% (1.13.2) | the reserve order of item 6 puts the two families Apple pays for (perm, shfla) at R1 and R3 with the costs measured; the 5% rule is checked live before each unlock |
| An NVIDIA owner | the majority sets the difficulty; a chip that beats NVIDIA is the only chip that matters | the chip model and the bounty (bench-log "Counter ASIC 2.0, the numbers") |
| A rig or a pool | a pool's share by vendor is what the chain-attributed reading cannot see (one vote key per pool) | pools publish their own vendor mix or appear as "unknown"; the detector (item 4) reads the per-program spread per card model |
| The project | a fleet-reported reading above the chain-attributed one means a reporting worker is not finding blocks (a stuck worker, the 4 October class); below it means miners outside the intake | the daily report carries both readings from the merge on |
9. Files
| File | What |
|---|---|
bench/repro.sh, bench/repro.ps1 |
the operator commands |
bench/README.md |
the package's README (operator instructions, tolerances, how to send a result back) |
bench/make-package.sh |
builds the package from a tagged commit with the existing cross-build scripts (zig for Linux, mingw for Windows, swiftc and cc for macOS, cargo for the CPU tool) |
bench/ingest.mjs, bench/reference.json |
the agreement check and the bench-table ingestion; the references |
bench/jobs/pc-repro.ps1, bench/jobs/collect-results.mjs |
the PC job (one card at a time through the app) and the Mac-side extraction of its result files |
docs/benchmarks/repro-2026-10-06/ |
the three result files of tonight, their tables and logs |
packaging/ota/publish-public.sh --repro |
the public downloads entry (tar.gz, zip, their sha256, the two aliases, the downloads index) |