Reproducible benchmark package: one command per platform (bench/repro.sh, repro.ps1), run on the Mac and PC 2; PC 1 owed

bench/: repro.sh (Linux, macOS) and repro.ps1 (Windows) print the machine, check the genesis pack's lottery hash vectors on
the CPU and every GPU, hash for 120 s per card at 1 GiB, probe random reads at 4, 64, 256 and 1024 MiB, sweep the same
sizes, prove and verify one fixture shard when a 12 GB NVIDIA card and a prover host exist, and write one JSON result (the
bench table's row shape, every command, every binary's sha256, the tolerances) plus a table. make-package.sh builds
igneum-repro-<tag>.tar.gz and .zip from the existing cross-build scripts; ingest.mjs checks a result against
bench/reference.json and adds "reproduced externally" rows to site/miner-bench.json only when every number agrees;
publish-public.sh --repro publishes next to the downloads (not deployed). The workers gained --list, --bench --pack
--seconds --dataset-mib and --memprobe (CUDA from the readwidth branch, OpenCL memprobe from opencl-rdna4, Metal new) with
RESULT and DEVICE lines; igneum-pow gained check-pack; the pack reader accepts a string-seed pack.

Runs (docs/benchmarks/repro.md, the bench-log entry, evidence rows 4, 5, 15): the Mac under the measure lock (Metal
27.674 MH/s, Apple OpenCL 26.875, 96/96 on CPU, Metal and OpenCL, 2^24 fingerprint 25f96e7dce90bd4e, 3.40 G random
reads/s, CPU verify 0.611 ms per warp); PC 2 (96/96 through NVRTC, NVIDIA OpenCL and AMD OpenCL with the same
fingerprint; gfx1036 3.312 MH/s; the 5090 rows beside the app's live miner, 62.4 MH/s, not the card's; the live host
proved and verified the shard through step 6, compressed 33.9 s). PC 1 went down before its slot; owed.

Classes closed: the PC job parses the package script before touching a card and bench/ is in the Windows 5.1 parse job;
the card-off is confirmed by nvidia-smi's compute-apps list and the prover's state read from settings.json, never
api/state; PowerShell case-only variable pairs fail CI (bench/jobs/ps-case-check.sh); the package builder keeps the Mac
workers outside the stage and checks every binary with file -b.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 01:16:24 +00:00
parent 1ffb7454c1
commit eca4111c34
30 changed files with 3805 additions and 45 deletions

View file

@ -47,6 +47,7 @@ on:
- 'proving/windows-wsl2/**'
- 'relay/clients/**'
- 'relay/playbooks/**'
- 'bench/**'
- 'brand/icons/**'
- 'tools/ci/windows/**'
- '.github/workflows/windows.yml'
@ -85,7 +86,7 @@ jobs:
Install-Module -Name PSScriptAnalyzer -Force -Scope CurrentUser -AllowClobber
}
Import-Module PSScriptAnalyzer
$folders = @('proto-cuda/windows-app', 'proto-cuda/windows-miner', 'proto-cuda/windows-node', 'proving/windows-wsl2', 'relay/clients', 'relay/playbooks', 'packaging/windows', 'tools/ci/windows')
$folders = @('proto-cuda/windows-app', 'proto-cuda/windows-miner', 'proto-cuda/windows-node', 'proving/windows-wsl2', 'relay/clients', 'relay/playbooks', 'packaging/windows', 'tools/ci/windows', 'bench')
$total = 0
foreach ($f in $folders) {
$results = Invoke-ScriptAnalyzer -Path $f -Recurse -Severity Warning, Error -ExcludeRule PSAvoidUsingWriteHost, PSUseShouldProcessForStateChangingFunctions, PSUseSingularNouns, PSAvoidUsingPositionalParameters

3
.gitignore vendored
View file

@ -31,3 +31,6 @@ vendor/igneum-node-ship/
# Trademark instruction packs name the director and the applicant company; never in the repository
brand/trademark/pbip-pack/
brand/trademark/*.zip
# the reproducible benchmark package (bench/make-package.sh output)
bench/build/

80
bench/README.md Normal file
View file

@ -0,0 +1,80 @@
# Igneum reproducible benchmark
One command reproduces the numbers on the Igneum bench table on your own machine, on day one, with no secrets, no node, no
wallet and no network connection. The result is one JSON file you can send back, signed by nothing, that names every
command run and the sha256 of every binary used.
## Run it
| Platform | Command |
|---|---|
| Linux (x86_64, NVIDIA driver or any OpenCL driver) | `tar xzf igneum-repro-<tag>.tar.gz && cd igneum-repro-<tag> && ./repro.sh` |
| macOS (Apple silicon) | the same tar.gz; `./repro.sh` |
| Windows 10 or 11 (x64) | unzip `igneum-repro-<tag>.zip`, then in PowerShell: `powershell -ExecutionPolicy Bypass -File repro.ps1` |
Options: `--seconds 120` (per card, the default), `--only cuda:0,opencl:1` (one card; the indexes are the ones `--list`
prints at the start), `--no-probe`, `--no-sweep`, `--no-prove`, `--out <dir>`. On Windows the same as `-Seconds`,
`-Only`, `-NoProbe`, `-NoSweep`, `-NoProve`, `-Out`.
Stop mining on the card first. A number taken while a miner, a game or another benchmark holds the card is not a number;
the result file does not know, so you do.
Time: about 4 minutes per card (120 s hash benchmark, three short sweep sizes, the probe table), plus the CPU check (one
second). The optional proving step takes about a minute on an RTX 5090 class card and needs a prover host you built
(below).
## What runs, in order
| Step | What | What it needs | What it reports |
|---|---|---|---|
| 1 | The machine | nothing | OS, CPU, memory, every GPU each worker sees (name, driver, memory) |
| 2 | The lottery hash vectors of the published genesis pack (`packs/igneum-genesis-mh`: seed `igneum-genesis`, day 2026-10-03, generator 2, memory-hard 1 GiB dataset, program id `bcc1248b10cc90f2`) | a CPU; every GPU present | bit-exact or not: 3 warps, 96 lanes, the cache FNV, on the CPU through the Rust interpreter and on every GPU through its own compiler (NVRTC, the OpenCL driver, Metal) |
| 3 | The hash benchmark, 120 s per card at 1 GiB | a GPU with 1.3 GB free | MH/s, the hashes done, and the fingerprint of the first 2^24 outputs (equal fingerprints on two machines mean every one of those 16.7 million hashes agreed) |
| 4 | The random-read probe at 4, 64, 256 and 1024 MiB | a GPU | dependent random 4-byte reads per second (the hash's access pattern) and their latency, eight independent chains, 64-byte lines, the coalesced stream bandwidth, an integer chain |
| 5 | The chip-resistance sweep: the same program at 4, 64, 256 and 1024 MiB | a GPU | MH/s per size; the ratio in-cache to 1 GiB; the hash's share of the card's random-read ceiling (MH/s x 128 loads against step 4's chase at 1 GiB) |
| 6 | One fixture shard proven with the pinned guest and verified (optional) | an NVIDIA card with 12 GB or more, Linux or WSL2, a built `igneum-prove-host` (`IGNEUM_PROVE_HOST`) | execute, core and compressed proof times, the verify time, VERIFIED or not; otherwise the line says why it was skipped |
| 7 | The result | nothing | `results/igneum-repro-<os>-<time>.json` and `.md`, plus the full log |
The CUDA worker needs only the NVIDIA driver on Windows (the runtime compiler ships in the package) and the driver plus
`libnvrtc.so.12` on Linux (from the CUDA toolkit or NVIDIA's nvrtc redistributable). The OpenCL worker needs the GPU
driver's OpenCL (AMD Adrenalin or ROCm, Intel, NVIDIA, Apple). An NVIDIA card is run through CUDA for the full time and
through OpenCL for a quarter of it as a cross-check; an Apple card through Metal and through Apple's OpenCL the same way.
## The pass marks and the tolerances
| Number | Pass mark |
|---|---|
| Vectors | exact: every lane of every warp, on the CPU and on every GPU |
| Fingerprint | exact across machines at the same `--batch-log2` |
| Hash MH/s at 1 GiB | within 3% of the published figure for the same card model |
| Sweep MH/s | within 10% |
| Random reads at 1 GiB (G loads/s), stream GB/s | within 10% |
| CPU verify per warp | under 10 ms on any core (the chain's verification gate); the published cores are under 1 ms |
| Proof times | within 25% |
A row on the bench table (`/miners`) becomes "reproduced externally" when an unrelated operator's result agrees with the
team's reference within these tolerances. Disagreeing results are published too, as "submitted", with the deltas.
## Send it back
Attach `results/igneum-repro-<os>-<time>.json` (and the `.md` if you like) to an issue on the Igneum repository once it
is public, or email it to the address on igneum.network. Say which card was idle (no miner, no game) during the run.
The file carries no hostname, no user name and no address; it carries your GPU and CPU models, your OS version and the
driver version.
## The reproduction reward
The reward terms are the ones of `docs/benchmarks/proving-e2e.md` section 8.1, unchanged: a published fixed reproduction
reward, equal for everyone and announced before the run, is allowed and disclosed; three operators are unrelated when
they are different persons not paid by the project beyond that reward, on hardware bought separately with no shared
host, card, rack or power meter, on different autonomous systems and at different sites, running the same published
release by hash, with no payment, loan or equipment between them or from the project beyond the disclosed reward. The
amount and who pays it are in `docs/plans/funding.md` (challenge reward row); nothing is funded as of 6 October 2026,
and reproduction is asked for without a reward, which the standard allows.
## Build the package yourself
From the Igneum repository at the tagged commit: `bench/make-package.sh` (macOS with swiftc, cargo, zig and mingw-w64;
every binary comes from the repository's own cross-build scripts). The sha256 of every file is in `SHA256SUMS`, and the
`VERSION` file names the tag and commit. `bin/windows-x86_64/THIRD-PARTY.md` lists the one third-party component (NVIDIA's
NVRTC runtime compiler) with its licence.

104
bench/ingest.mjs Normal file
View file

@ -0,0 +1,104 @@
#!/usr/bin/env node
// Ingests a reproducible-benchmark result (bench/repro.sh or repro.ps1, format igneum-repro-1) into the public bench
// table site/miner-bench.json, the one /miners renders (site/build.mjs), and says per number whether it agrees with the
// team's reference within the stated tolerance (bench/reference.json, every reference a bench-log entry).
//
// node bench/ingest.mjs check <result.json> [...] agreement table only, writes nothing
// node bench/ingest.mjs add <result.json> [...] [--by "reproduced externally"|"measured by the team"|"submitted"]
// append one row per card to site/miner-bench.json
// node bench/ingest.mjs merge <out.json> <result.json> [...] one file from per-card runs of one machine (--only runs)
//
// The rule (docs/evidence.md, "What would move a row"): a card's row becomes "reproduced externally" when an unrelated
// operator's result agrees with the team's reference within tolerance (hash rate 3%, probe loads/s 10%, vectors and
// fingerprint exact); `add` refuses that label when the agreement table says no. "Measured by the team" is our own
// hardware (the three runs in docs/benchmarks/repro.md). A result that disagrees is still recorded as "submitted" with
// the deltas in its note: a failing run is as public as a passing one (proving-e2e.md 8.2 step 4).
// Scrubbing: the site build fails on any private string (site/forbidden-strings.txt); a card name or note that carries
// one is refused here first. No dependencies.
import { readFileSync, writeFileSync, existsSync } from 'node:fs';
import { dirname, join, resolve } from 'node:path';
import { fileURLToPath } from 'node:url';
const ROOT = resolve(dirname(fileURLToPath(import.meta.url)), '..');
const argv = process.argv.slice(2);
const cmd = argv[0];
const flags = {}; const files = [];
for (let i = 1; i < argv.length; i++) { if (argv[i].startsWith('--')) { flags[argv[i].slice(2)] = argv[i + 1]; i++; } else files.push(argv[i]); }
if (!cmd || !files.length) { console.log(readFileSync(fileURLToPath(import.meta.url), 'utf8').split('\n').slice(1, 20).map(l => l.replace(/^\/\/ ?/, '')).join('\n')); process.exit(2); }
const reference = JSON.parse(readFileSync(join(ROOT, 'bench', 'reference.json'), 'utf8'));
const forbidden = readFileSync(join(ROOT, 'site', 'forbidden-strings.txt'), 'utf8').split('\n').map(l => l.trim()).filter(l => l && !l.startsWith('#')).map(p => new RegExp(p));
const load = f => { const j = JSON.parse(readFileSync(f, 'utf8')); if (j.format !== 'igneum-repro-1') throw new Error(`${f}: format ${j.format} is not igneum-repro-1`); return j; };
const within = (got, ref, tol) => got != null && ref != null && ref !== 0 && Math.abs(got - ref) / Math.abs(ref) <= tol;
const cmp = (got, ref, tol) => ref == null ? 'no reference yet' : within(got, ref, tol) ? 'agree' : 'DISAGREE';
const pct = (got, ref) => (got == null || ref == null || !ref) ? 'n/a' : ((got - ref) / ref * 100).toFixed(1) + '%';
const scrubOk = s => !forbidden.some(re => re.test(s));
const refFor = (card) => reference.cards.find(r => card.name.toLowerCase().includes(r.match.toLowerCase()));
// one card against its reference: every number with its tolerance, and the verdict
function agreement(res, card) {
const tol = res.tolerance || {};
const ref = refFor(card);
const rows = [];
rows.push(['vectors', card.vectors, 'PASS', card.vectors === 'PASS' ? 'agree' : 'DISAGREE', 'exact']);
if (ref) {
const fp = ref.fingerprint && ref.fingerprint[String(card.hash.batch_log2)];
if (fp) rows.push([`fingerprint 2^${card.hash.batch_log2}`, card.fingerprint, fp, card.fingerprint === fp ? 'agree' : 'DISAGREE', 'exact']);
rows.push(['hash MH/s at 1 GiB', card.hash.mhs, ref.mhs, cmp(card.hash.mhs, ref.mhs, tol.hash_mhs ?? 0.03), `${(tol.hash_mhs ?? 0.03) * 100}% (${pct(card.hash.mhs, ref.mhs)})`, ref.source]);
for (const mib of ['4', '64', '256']) if (ref.sweep_mhs && ref.sweep_mhs[mib] != null) rows.push([`sweep ${mib} MiB MH/s`, card.sweep_mhs[mib], ref.sweep_mhs[mib], cmp(card.sweep_mhs[mib], ref.sweep_mhs[mib], tol.sweep_mhs ?? 0.10), `${(tol.sweep_mhs ?? 0.10) * 100}% (${pct(card.sweep_mhs[mib], ref.sweep_mhs[mib])})`]);
const p = card.memprobe && card.memprobe['1024'];
if (ref.chase_gloads_1024 != null) rows.push(['random reads at 1 GiB, G loads/s', p && p.chase_gloads, ref.chase_gloads_1024, cmp(p && p.chase_gloads, ref.chase_gloads_1024, tol.chase_gloads ?? 0.10), `${(tol.chase_gloads ?? 0.10) * 100}% (${pct(p && p.chase_gloads, ref.chase_gloads_1024)})`]);
if (ref.stream_gbps_1024 != null) rows.push(['stream GB/s at 1 GiB', p && p.stream_gbps, ref.stream_gbps_1024, cmp(p && p.stream_gbps, ref.stream_gbps_1024, tol.stream_gbps ?? 0.10), `${(tol.stream_gbps ?? 0.10) * 100}% (${pct(p && p.stream_gbps, ref.stream_gbps_1024)})`]);
} else rows.push(['reference', 'none for this card', '', 'no reference (first run of this model: it becomes the reference once a second machine agrees)', '']);
const verdict = rows.some(r => r[3] === 'DISAGREE') ? 'outside tolerance' : rows.some(r => r[3].startsWith('no reference')) ? 'no reference yet for some numbers' : 'agrees within tolerance';
return { ref, rows, verdict };
}
function printAgreement(res, f) {
console.log(`\n${f}: ${res.run.os}, ${res.machine.os}; ${res.machine.cpu}; package ${res.package.version} (${res.package.commit}); run ${res.run.id}`);
if (res.cpu) console.log(` CPU vectors ${res.cpu.vectors} (${res.cpu.lanes_pass} of ${res.cpu.lanes} lanes), verify ${res.cpu.verify_ms_per_warp} ms per warp (reference ${reference.cpu_verify_ms_per_warp_max} ms gate: ${res.cpu.verify_ms_per_warp <= reference.cpu_verify_ms_per_warp_max ? 'under' : 'OVER'})`);
const verdicts = [];
for (const card of res.cards) {
const a = agreement(res, card);
console.log(` ${card.backend}:${card.index} ${card.name}${card.cross_check ? ' (cross-check)' : ''}: ${a.verdict}`);
for (const r of a.rows) console.log(` ${r[0].padEnd(34)} got ${String(r[1]).padEnd(20)} ref ${String(r[2]).padEnd(20)} ${r[3]}${r[4] ? ' tol ' + r[4] : ''}${r[5] ? ' (' + r[5] + ')' : ''}`);
verdicts.push({ card, verdict: a.verdict, ref: a.ref });
}
if (res.proving) console.log(` proving: ${res.proving.status}${res.proving.verify ? `, compressed ${res.proving.compressed_prove_seconds} s, ${res.proving.verify} in ${res.proving.verify_seconds} s` : ''}`);
return verdicts;
}
if (cmd === 'check') {
for (const f of files) printAgreement(load(f), f);
} else if (cmd === 'merge') {
const [out, ...ins] = files;
const all = ins.map(load);
const base = JSON.parse(JSON.stringify(all[0]));
base.cards = []; base.rows = []; base.commands = []; base.run.only = ''; base.run.merged_from = all.map(r => r.run.id);
for (const r of all) { base.cards.push(...r.cards); base.rows.push(...r.rows); base.commands.push(...r.commands); if (r.proving && r.proving.status === 'run') base.proving = r.proving; }
writeFileSync(out, JSON.stringify(base, null, 1) + '\n');
console.log(`wrote ${out}: ${base.cards.length} cards from ${all.length} runs`);
} else if (cmd === 'add') {
const by = flags.by || 'submitted';
const tablePath = join(ROOT, 'site', 'miner-bench.json');
const table = JSON.parse(readFileSync(tablePath, 'utf8'));
let added = 0;
for (const f of files) {
const res = load(f);
const verdicts = printAgreement(res, f);
for (const { card, verdict } of verdicts) {
if (card.cross_check || card.hash.mhs == null) continue;
const row = res.rows.find(r => r.source === `repro:${res.run.id}` && r.card === card.name && r.miner.includes(`(${card.backend} `)) || res.rows.find(r => r.card === card.name);
if (!row) continue;
let label = by;
if (by === 'reproduced externally' && verdict !== 'agrees within tolerance') { console.log(` refusing "reproduced externally" for ${card.name}: ${verdict}; recorded as submitted`); label = 'submitted'; }
const note = `${row.note}; ${verdict}`;
if (!scrubOk(card.name) || !scrubOk(note)) { console.log(` refusing ${card.name}: a forbidden string (site/forbidden-strings.txt) in the name or note`); continue; }
if (table.rows.some(r => r.source === row.source && r.card === row.card && r.miner === row.miner)) { console.log(` already in the table: ${row.card} from ${row.source}`); continue; }
table.rows.push({ ...row, by: label, note });
added++;
console.log(` added ${row.card} ${row.mh_s} MH/s as "${label}" (source ${row.source})`);
}
}
if (added) { writeFileSync(tablePath, JSON.stringify(table, null, 2) + '\n'); console.log(`wrote ${tablePath} (+${added} rows; rebuild the site with node site/build.mjs)`); }
} else { console.error(`unknown command ${cmd}`); process.exit(2); }

View file

@ -0,0 +1,33 @@
#!/usr/bin/env node
// Cuts the result files out of a PC repro job's upload (bench/jobs/pc-repro.ps1 prints each file between RESULT-FILE-BEGIN
// and RESULT-FILE-END with every line prefixed "J "), checks the sha256 the PC printed, and writes them to --out.
// node tools/jobs.mjs <job id> | node bench/jobs/collect-results.mjs --out docs/benchmarks/repro-2026-10-06/pc1
// Lines may arrive prefixed by the job runner ("RESULT ...") and by the machine; only the "J " payload matters.
import { writeFileSync, mkdirSync } from 'node:fs';
import { join } from 'node:path';
import { createHash } from 'node:crypto';
const out = process.argv[process.argv.indexOf('--out') + 1] || '.';
mkdirSync(out, { recursive: true });
let text = '';
process.stdin.setEncoding('utf8');
process.stdin.on('data', d => text += d);
process.stdin.on('end', () => {
const lines = text.split('\n').map(l => l.replace(/\r$/, ''));
let cur = null, body = [], want = null;
let n = 0;
for (const raw of lines) {
const l = raw.replace(/^.*?(RESULT-FILE-BEGIN|RESULT-FILE-END|J )/, '$1');
if (l.startsWith('RESULT-FILE-BEGIN ')) { const p = l.split(' '); cur = p[1]; want = p[3] || null; body = []; continue; }
if (l.startsWith('RESULT-FILE-END ') && cur) {
const content = body.join('\n') + '\n';
const sha = createHash('sha256').update(content.replace(/\n$/, '')).digest('hex');
// the PC hashed the file bytes (CRLF, maybe a BOM); the lines here are what the upload kept, so the hash is advisory
writeFileSync(join(out, cur), content);
console.log(`${cur}: ${body.length} lines written to ${out}${want ? ` (PC sha256 ${want.slice(0, 12)}..., ours over LF text ${sha.slice(0, 12)}...)` : ''}`);
n++; cur = null; body = []; want = null; continue;
}
if (cur && l.startsWith('J ')) body.push(l.slice(2));
}
if (!n) { console.error('no RESULT-FILE blocks in the input'); process.exit(1); }
});

168
bench/jobs/pc-repro.ps1 Normal file
View file

@ -0,0 +1,168 @@
# Igneum run job: the reproducible benchmark package on a PC, one card at a time (bench/README.md, docs/benchmarks/repro.md).
# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; for each card the
# package's workers list, this script switches only that card off in the app through POST <app.url>api/cards, waits for its
# worker to stop, runs repro.ps1 -Only <backend:index> on it, and switches the card back on with the settings it had.
# Every result line starts with RESULT so `node tools/jobs.mjs <job id>` shows them; each result file is printed whole
# between RESULT-FILE-BEGIN and RESULT-FILE-END lines (tools: bench/jobs/collect-results.mjs cuts them out).
# Needs the fetch job before it: the package zip extracted under the app's jobs folder (any depth; the newest repro.ps1 wins).
# Environment: IGNEUM_REPRO_PROVE=1 runs the proving step on the NVIDIA card (PC 2: the WSL2 prover host),
# IGNEUM_PROVE_HOST and IGNEUM_WSL_DISTRO as repro.ps1 reads them, IGNEUM_REPRO_SECONDS (default 120).
$ErrorActionPreference = 'Continue'
function Say([string] $m) { Write-Output ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
$jobs = Split-Path $env:IGNEUM_JOB_DIR
$pkgScript = Get-ChildItem -Path $jobs -Recurse -Depth 4 -Filter 'repro.ps1' -ErrorAction SilentlyContinue | Sort-Object LastWriteTime -Descending | Select-Object -First 1
if (-not $pkgScript) { Write-Output "RESULT error no repro.ps1 under $jobs (the fetch job runs first)"; exit 2 }
$pkg = $pkgScript.DirectoryName
# parse the package script before any card is touched (5 October 2026: a parse error cost three cards 90 s each for nothing)
$tokens = $null; $perrs = $null
[void] [System.Management.Automation.Language.Parser]::ParseFile($pkgScript.FullName, [ref] $tokens, [ref] $perrs)
if ($perrs -and $perrs.Count -gt 0) { $perrs | ForEach-Object { Write-Output ("RESULT error repro.ps1 does not parse: line " + $_.Extent.StartLineNumber + ": " + $_.Message) }; exit 2 }
Write-Output "RESULT repro.ps1 parses"
$seconds = if ($env:IGNEUM_REPRO_SECONDS) { [int] $env:IGNEUM_REPRO_SECONDS } else { 120 }
$prove = ($env:IGNEUM_REPRO_PROVE -eq '1')
Write-Output "RESULT package $pkg version $((Get-Content (Join-Path $pkg 'VERSION') | Select-Object -First 2) -join ' ') sha256 repro.ps1 $((Get-FileHash -Algorithm SHA256 $pkgScript.FullName).Hash.ToLower())"
$results = Join-Path $env:IGNEUM_JOB_DIR 'results'
New-Item -ItemType Directory -Force -Path $results | Out-Null
# the app: its URL, its cards
$appDir = $env:IGNEUM_APP_DIR
if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
$urlFile = Join-Path $appDir 'app.url'
$url = $null
if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
# the app's cards from its settings.json (key = vendor:device:name, with enabled, identities, power_pct), never from
# api/state: on a proving machine 0.3.11 answers "{}" after about 15 paid shards (the coordinator's rule, 6 October 2026)
function AppCards() {
$sp = Join-Path $appDir 'settings.json'
if (-not (Test-Path $sp)) { Say "no settings.json at $sp"; return @() }
try {
$st = Get-Content -LiteralPath $sp -Raw | ConvertFrom-Json
$list = @()
foreach ($prop in $st.cards.PSObject.Properties) {
$v = $prop.Value
$pp = [int] $v.power_pct; if ($pp -lt 50) { $pp = 80 }
$list += @{ key = $prop.Name; vendor = ($prop.Name -split ':')[0]; enabled = [bool] $v.enabled; identities = [int] $v.identities; power_pct = $pp }
}
return $list
} catch { Say ("settings.json: " + $_.Exception.Message); return @() }
}
function SetCard($card, [bool] $enabled) {
$body = @{ cards = @(@{ key = $card.key; enabled = $enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; return $true } catch { Say ("api/cards: " + $_.Exception.Message); return $false }
}
# waits for the card's worker to stop and CONFIRMS it: on NVIDIA by nvidia-smi's compute-apps list (no igneum worker
# process on the card), as the M16 job does; api/state is not trusted (6 October 2026: run-repro-pc2-20261006 read "{}",
# waited 30 s, and benched the 5090 beside the app's live worker at half the card's rate). On other vendors the app's
# own process list (api/state when it answers) and a fixed wait; the result says "unconfirmed" when nothing confirmed it.
function CudaWorkers() {
$apps = @(& nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv,noheader 2>$null | ForEach-Object { "$_" } | Where-Object { $_ -match 'igneum' })
return $apps
}
function WaitOff($card) {
$t = 0; $confirmed = $false
while ($t -lt 120) {
Start-Sleep -Seconds 5; $t += 5
if ($card.vendor -eq 'nvidia') {
$w = CudaWorkers
if ($w.Count -eq 0) { $confirmed = $true; break }
continue
}
$c2 = $null
try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $cs = $st.mining.cards; if (-not $cs) { $cs = $st.cards }; $c2 = $cs | Where-Object { $_.key -eq $card.key } | Select-Object -First 1 } catch { }
if ($c2) { if ($c2.state -eq 'off' -and $c2.pid -eq 0) { $confirmed = $true; break } }
elseif ($t -ge 30) { break }
}
return @{ seconds = $t; confirmed = $confirmed }
}
$cards = AppCards
$cards | ForEach-Object { Write-Output ("RESULT app-card " + $_.key + " enabled=" + $_.enabled + " identities=" + $_.identities + " power_pct=" + $_.power_pct) }
# the devices the package's workers see
$bin = Join-Path $pkg 'bin\windows-x86_64'
$devices = @()
foreach ($w in @(@{ exe = 'igneum-worker-cuda.exe'; backend = 'cuda' }, @{ exe = 'igneum-worker-opencl.exe'; backend = 'opencl' })) {
$exe = Join-Path $bin $w.exe
if (-not (Test-Path $exe)) { continue }
$out = & $exe --list 2>&1 | ForEach-Object { "$_" }
foreach ($l in $out) {
if ($l -like 'DEVICE *') {
if ($w.backend -eq 'opencl' -and -not ($l -match ' type=GPU')) { continue }
$i = [regex]::Match($l, ' index=(\d+)').Groups[1].Value
$n = [regex]::Match($l, ' name="([^"]*)"').Groups[1].Value
$devices += @{ backend = $w.backend; index = $i; name = $n }
}
}
}
$devices | ForEach-Object { Write-Output ("RESULT device " + $_.backend + ":" + $_.index + " " + $_.name) }
if ($devices.Count -eq 0) { Write-Output 'RESULT error no device listed by the package workers'; exit 2 }
foreach ($d in $devices) {
# the app's card for this device: NVIDIA by vendor; an OpenCL device by the index in the key (vendor:index:name) or the name
$card = $null
if ($d.backend -eq 'cuda') { $card = $cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 }
else { $card = $cards | Where-Object { $_.vendor -ne 'nvidia' -and ($_.key -like ("*:" + $d.index + ":*") -or $_.key -like ("*" + $d.name + "*")) } | Select-Object -First 1 }
# an NVIDIA card through OpenCL: the app's NVIDIA card again (the cross-check runs with that card off too)
if (-not $card -and $d.backend -eq 'opencl' -and $d.name -match 'NVIDIA') { $card = $cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 }
$was = $null
if ($card) {
$was = $card
if (SetCard $card $false) {
$w = WaitOff $card
if ($w.confirmed) { Write-Output ("RESULT card-off " + $card.key + " confirmed after " + $w.seconds + " s") }
else { Write-Output ("RESULT card-off-UNCONFIRMED " + $card.key + " after " + $w.seconds + " s: the numbers below are beside whatever still runs on the card" + $(if ($card.vendor -eq 'nvidia') { " (compute apps: " + ((CudaWorkers) -join '; ') + ")" })) }
} else { Write-Output ("RESULT card-off-failed " + $card.key + " (measuring with the app's worker still on it)") }
Start-Sleep -Seconds 5
} else { Write-Output ("RESULT card none-found for " + $d.backend + ":" + $d.index + " " + $d.name + " (not in the app; measuring as is)") }
if ($d.backend -eq 'cuda') { & nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" } }
$argv = @('-ExecutionPolicy', 'Bypass', '-File', $pkgScript.FullName, '-Seconds', "$seconds", '-Only', ($d.backend + ':' + $d.index), '-Out', $results)
$proveHere = ($prove -and $d.backend -eq 'cuda')
if (-not $proveHere) { $argv += '-NoProve' }
$proverWasOn = $false
if ($proveHere) {
# the live prover off (its GPU server is the one a client would connect to) and its sockets cleared, before and after
# (5 October 2026: a root-owned /tmp/sp1-cuda-*.sock blinded the live prover; the proving agent's rule)
# the prover's state from settings.json, never api/state (it answered "{}" and the first run left the prover off)
$proverWasOn = $true
try { $sj = Get-Content -LiteralPath (Join-Path $appDir 'settings.json') -Raw | ConvertFrom-Json; $proverWasOn = [bool] $sj.prove } catch { }
try { Invoke-RestMethod -Uri ($url + 'api/prove') -Method POST -Body '{"on":false}' -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT prover-off (was on: " + $proverWasOn + ")") } catch { Write-Output ("RESULT prover-off failed: " + $_.Exception.Message) }
$distro = if ($env:IGNEUM_WSL_DISTRO) { $env:IGNEUM_WSL_DISTRO } else { 'Ubuntu-24.04' }
$wu = @(); if ($env:IGNEUM_WSL_USER) { $wu = @('-u', $env:IGNEUM_WSL_USER) }
Start-Sleep -Seconds 10
& wsl.exe -d $distro @wu -- bash -lc 'pkill -f sp1-gpu-server; sleep 2; rm -f /tmp/sp1-cuda-*.sock; echo "sockets cleared: $(ls /tmp/sp1-cuda-*.sock 2>/dev/null | wc -l) left"' 2>&1 | ForEach-Object { "RESULT prover-clean-before $_" }
}
Write-Output ("RESULT run " + $d.backend + ":" + $d.index + " start " + (Get-Date -Format HH:mm:ss) + " " + ($argv -join ' '))
& powershell.exe @argv 2>&1 | Where-Object { $_ -match '^RESULT|^DEVICE|^pack |^vector|^cache:|^warm-up|^rate:|^wrote|^done|error|FAIL|^\+ ' } | ForEach-Object { "RESULT $_" }
Write-Output ("RESULT run " + $d.backend + ":" + $d.index + " end " + (Get-Date -Format HH:mm:ss) + " exit " + $LASTEXITCODE)
if ($d.backend -eq 'cuda') {
# the job-size question (consequences reviewer, 5 October 2026): the same card at the app's 2^21 job size against 2^24,
# on the package's genesis pack and, when the app has exported one, on the live devnet pack (packs\devnet)
$cudaExe = Join-Path $bin 'igneum-worker-cuda.exe'
$packsToTry = @(@{ name = 'genesis'; dir = (Join-Path $pkg 'packs\igneum-genesis-mh') })
$live = Join-Path $appDir 'packs\devnet'
if (Test-Path (Join-Path $live 'program.h')) { $packsToTry += @{ name = 'live-devnet'; dir = $live } }
foreach ($pk in $packsToTry) {
foreach ($bl in @(21, 24)) {
$batches = if ($bl -eq 21) { 64 } else { 8 }
Write-Output ("RESULT jobsize " + $pk.name + " batch_log2=" + $bl + " start " + (Get-Date -Format HH:mm:ss))
& $cudaExe --bench --pack $pk.dir --device $d.index --batches $batches --batch-log2 $bl 2>&1 | Where-Object { $_ -match '^RESULT|^warm-up|^rate:|error|FAIL' } | ForEach-Object { "RESULT jobsize " + $pk.name + " " + $_ }
}
}
}
if ($d.backend -eq 'cuda') { & nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" } }
if ($proveHere) {
& wsl.exe -d $distro @wu -- bash -lc 'pkill -f sp1-gpu-server; sleep 2; rm -f /tmp/sp1-cuda-*.sock; echo "sockets cleared: $(ls /tmp/sp1-cuda-*.sock 2>/dev/null | wc -l) left"' 2>&1 | ForEach-Object { "RESULT prover-clean-after $_" }
if ($proverWasOn) { try { Invoke-RestMethod -Uri ($url + 'api/prove') -Method POST -Body '{"on":true}' -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output "RESULT prover-restored on" } catch { Write-Output ("RESULT error prover restore: " + $_.Exception.Message) } }
}
if ($was) {
if (SetCard $was ([bool]$was.enabled)) { Write-Output ("RESULT card-restored " + $was.key + " enabled=" + $was.enabled) } else { Write-Output ("RESULT error card restore " + $was.key) }
}
}
# every result file, whole, so the Mac side can rebuild them from the job's upload
Get-ChildItem -Path $results -Filter '*.json' | Sort-Object Name | ForEach-Object {
Write-Output ("RESULT-FILE-BEGIN " + $_.Name + " " + $_.Length + " " + (Get-FileHash -Algorithm SHA256 $_.FullName).Hash.ToLower())
Get-Content -LiteralPath $_.FullName | ForEach-Object { "J $_" }
Write-Output ("RESULT-FILE-END " + $_.Name)
}
Get-ChildItem -Path $results -Filter '*.md' | Sort-Object Name | ForEach-Object { Write-Output ("RESULT-FILE-BEGIN " + $_.Name); Get-Content -LiteralPath $_.FullName | ForEach-Object { "J $_" }; Write-Output ("RESULT-FILE-END " + $_.Name) }
Write-Output "RESULT all cards restored; done"
exit 0

13
bench/jobs/ps-case-check.sh Executable file
View file

@ -0,0 +1,13 @@
#!/usr/bin/env bash
# Fails when a PowerShell file under bench/ uses two variable names that differ only by case (PowerShell variables are
# case-insensitive: on 6 October 2026 repro.ps1's $Cpu result table held $cpu, its own name string, so the table contained
# itself and ConvertTo-Json ran out of memory; $md, the markdown lines, wiped $Md, the markdown path). Run by CI and by hand.
set -uo pipefail
cd "$(dirname "$0")/../.."
rc=0
for f in $(git ls-files 'bench/*.ps1' 'bench/**/*.ps1' 'relay/playbooks/*.ps1' 2>/dev/null); do
dups="$(grep -oE '\$[A-Za-z_][A-Za-z0-9_]*' "$f" | sed 's/^\$//' | grep -vE '^(_|PSItem|null|true|false|env|script|Matches|LASTEXITCODE|PSVersionTable|MyInvocation|PSScriptRoot|ErrorActionPreference|args|input|PSBoundParameters)$' | sort -u | awk '{ k = tolower($0); seen[k] = seen[k] ? seen[k] " " $0 : $0; n[k]++ } END { for (k in n) if (n[k] > 1) print seen[k] }')"
if [ -n "$dups" ]; then echo "$f: variable names that differ only by case: $dups" >&2; rc=1; fi
done
[ "$rc" = 0 ] && echo "ps-case-check: no case-only variable pairs in the PowerShell files"
exit $rc

93
bench/make-package.sh Executable file
View file

@ -0,0 +1,93 @@
#!/usr/bin/env bash
# Builds the Igneum reproducible benchmark package from the commit this tree is on (bench/README.md; the operator
# side is repro.sh and repro.ps1). Every binary comes from the existing cross-build scripts; nothing is downloaded.
#
# bench/make-package.sh [--tag v0.1.0-repro] [--out bench/build] [--skip-build]
#
# Contents of igneum-repro-<tag>/ (tar.gz for Linux and macOS, zip for Windows, the same files in both):
# repro.sh, repro.ps1, README.md, VERSION (tag, commit, built), SHA256SUMS
# bin/macos-arm64/ igneum-pow, igneum-bench (Metal), igneum-bench-cl (Apple OpenCL) built here with swiftc, cc and cargo
# bin/linux-x86_64/ igneum-pow, igneum-worker-cuda, igneum-worker-opencl infra/cross/build-workers-linux.sh (zig), cargo zigbuild
# bin/windows-x86_64/ igneum-pow.exe, igneum-worker-cuda.exe, igneum-worker-opencl.exe, proto-cuda/nvrtc/build-windows.sh (mingw), cargo
# nvrtc64_120_0.dll, nvrtc-builtins64_128.dll (NVIDIA's redistributable runtime compiler, THIRD-PARTY.md)
# packs/igneum-genesis-mh/ the published genesis vectors (seed "igneum-genesis", day 2026-10-03, generator 2, memory-hard)
# proving/ block-338-shard1.json (the shard fixture) and manifest.json (the pinned guest ids and hashes)
# The tag is recorded, not created: tag the commit first (git tag -a), or pass --tag for a dry build from an untagged tree
# (VERSION then says "untagged"). Builds run under the Mac build lock, one at a time.
set -euo pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
ROOT="$(cd "$HERE/.." && pwd)"
LOCK="${IGNEUM_LOCK:-/Users/joshm/Projects/igneum/tools/lock/with-lock.sh}"; [ -x "$LOCK" ] || LOCK=""
TAG=""; OUT="$HERE/build"; BUILD=1
while [ $# -gt 0 ]; do
case "$1" in
--tag) TAG="$2"; shift 2 ;;
--out) OUT="$2"; shift 2 ;;
--skip-build) BUILD=0; shift ;;
*) echo "unknown argument: $1" >&2; exit 2 ;;
esac
done
COMMIT="$(git -C "$ROOT" rev-parse --short HEAD)"
EXACT="$(git -C "$ROOT" describe --tags --exact-match 2>/dev/null || true)"
[ -n "$TAG" ] || TAG="${EXACT:-untagged-$COMMIT}"
[ -n "$EXACT" ] || echo "note: HEAD $COMMIT carries no tag; VERSION will say $TAG (tag the commit before publishing)"
if [ -n "$(git -C "$ROOT" status --porcelain -- bench igneum-pow proto-cuda proto-opencl proto-metal infra/cross proving/fixtures proving/igneum-prove/elf)" ]; then
echo "note: the tree has uncommitted changes in the package's sources; the package names commit $COMMIT but is not that commit"
fi
NAME="igneum-repro-$TAG"
STAGE="$OUT/$NAME"
rm -rf "$STAGE"; mkdir -p "$STAGE/bin/macos-arm64" "$STAGE/bin/linux-x86_64" "$STAGE/bin/windows-x86_64" "$STAGE/packs" "$STAGE/proving"
log() { printf '%s %s\n' "$(date -u +%H:%M:%S)" "$*"; }
locked() { if [ -n "$LOCK" ]; then "$LOCK" build "$@"; else "$@"; fi; }
export PATH="$HOME/.cargo/bin:$PATH"
RED="${RED:-$ROOT/proto-cuda/nvrtc/redist}"
[ -d "$RED/include" ] || RED="/Users/joshm/Projects/igneum/proto-cuda/nvrtc/redist"
if [ "$BUILD" = 1 ]; then
log "igneum-pow: macOS, Linux (cargo zigbuild) and Windows (mingw)"
(cd "$ROOT/igneum-pow" && locked nice -n 19 cargo build --release -j 4 --quiet \
&& locked nice -n 19 cargo zigbuild --release -j 4 --quiet --target x86_64-unknown-linux-gnu.2.36 \
&& locked nice -n 19 cargo build --release -j 4 --quiet --target x86_64-pc-windows-gnu)
mkdir -p "$OUT/macos-arm64"
log "Metal worker (swiftc)"
(cd "$ROOT/proto-metal" && locked nice -n 19 swiftc -O -target arm64-apple-macos11 -o "$OUT/macos-arm64/igneum-bench" main.swift -framework Metal)
log "Apple OpenCL worker (cc, generic --pack mode against the placeholder pack)"
(cd "$ROOT/proto-opencl" && locked nice -n 19 cc -arch arm64 -std=c99 -O2 -Wall -Wextra -Wno-deprecated-declarations -I "$ROOT/proto-cuda/packs/igneum-devnet-v4-epoch0" \
-DIGNEUM_KERNEL_PATH='"kernel_bound.cl"' -o "$OUT/macos-arm64/igneum-bench-cl" host.c -framework OpenCL)
log "Linux workers (infra/cross/build-workers-linux.sh)"
OUT_DIR="$OUT/linux-workers" RED="$RED" locked bash "$ROOT/infra/cross/build-workers-linux.sh" >/dev/null
log "Windows workers (proto-cuda/nvrtc/build-windows.sh)"
[ -e "$ROOT/proto-cuda/nvrtc/redist" ] || ln -s "$RED" "$ROOT/proto-cuda/nvrtc/redist"
locked bash "$ROOT/proto-cuda/nvrtc/build-windows.sh" >/dev/null
fi
# --skip-build reuses what the last build left under $OUT (5 October 2026: a --skip-build package shipped without the two
# Mac workers because they had been built straight into the stage, which every build deletes first)
for b in igneum-bench igneum-bench-cl; do [ -x "$OUT/macos-arm64/$b" ] || { echo "no $OUT/macos-arm64/$b (run without --skip-build)" >&2; exit 1; }; done
cp "$OUT/macos-arm64/igneum-bench" "$OUT/macos-arm64/igneum-bench-cl" "$STAGE/bin/macos-arm64/"
cp "$ROOT/igneum-pow/target/release/igneum-pow" "$STAGE/bin/macos-arm64/"
cp "$ROOT/igneum-pow/target/x86_64-unknown-linux-gnu/release/igneum-pow" "$STAGE/bin/linux-x86_64/"
cp "$ROOT/igneum-pow/target/x86_64-pc-windows-gnu/release/igneum-pow.exe" "$STAGE/bin/windows-x86_64/"
cp "$OUT/linux-workers/igneum-worker-cuda" "$OUT/linux-workers/igneum-worker-opencl" "$STAGE/bin/linux-x86_64/"
cp "$ROOT/proto-cuda/nvrtc/igneum-worker-cuda.exe" "$ROOT/proto-opencl/igneum-worker-opencl.exe" "$STAGE/bin/windows-x86_64/"
for dll in "$RED"/bin/nvrtc*.dll; do [ -f "$dll" ] && cp "$dll" "$STAGE/bin/windows-x86_64/"; done
cp "$ROOT/proto-cuda/nvrtc/THIRD-PARTY.md" "$STAGE/bin/windows-x86_64/"
cp -R "$ROOT/proto-cuda/packs/igneum-genesis-mh" "$STAGE/packs/"
cp "$ROOT/proving/fixtures/block-338-shard1.json" "$ROOT/proving/igneum-prove/elf/manifest.json" "$STAGE/proving/"
cp "$HERE/repro.sh" "$HERE/repro.ps1" "$HERE/README.md" "$STAGE/"
chmod +x "$STAGE/repro.sh" "$STAGE"/bin/macos-arm64/* "$STAGE"/bin/linux-x86_64/*
printf '%s\ncommit %s\nbuilt %s\nsource https://github.com/igneum-network/igneum (bench/)\n' "$TAG" "$COMMIT" "$(date -u +%Y-%m-%dT%H:%M:%SZ)" > "$STAGE/VERSION"
(cd "$STAGE" && find . -type f ! -name SHA256SUMS | sort | while read -r f; do shasum -a 256 "$f" | sed 's#\./##'; done > SHA256SUMS)
# every binary is present and for its platform (file -b: the path itself carries "arm64" and "x86_64", so a missing file
# must not pass on its name; the same class as above)
for b in bin/macos-arm64/igneum-pow bin/macos-arm64/igneum-bench bin/macos-arm64/igneum-bench-cl bin/linux-x86_64/igneum-pow bin/linux-x86_64/igneum-worker-cuda bin/linux-x86_64/igneum-worker-opencl bin/windows-x86_64/igneum-pow.exe bin/windows-x86_64/igneum-worker-cuda.exe bin/windows-x86_64/igneum-worker-opencl.exe bin/windows-x86_64/nvrtc64_120_0.dll; do
[ -f "$STAGE/$b" ] || { echo "missing $b" >&2; exit 1; }
done
for b in "$STAGE"/bin/linux-x86_64/*; do file -b "$b" | grep -q 'ELF 64-bit' || { echo "$b is not a Linux ELF" >&2; exit 1; }; done
for b in "$STAGE"/bin/windows-x86_64/*.exe; do file -b "$b" | grep -q 'PE32+' || { echo "$b is not a Windows PE" >&2; exit 1; }; done
for b in "$STAGE"/bin/macos-arm64/*; do file -b "$b" | grep -q 'arm64' || { echo "$b is not arm64" >&2; exit 1; }; done
# the scripts here are the ones in the package (the class of "edited after copying")
cmp -s "$HERE/repro.sh" "$STAGE/repro.sh" && cmp -s "$HERE/repro.ps1" "$STAGE/repro.ps1" || { echo "script copy differs" >&2; exit 1; }
(cd "$OUT" && rm -f "$NAME.tar.gz" "$NAME.zip" && tar -czf "$NAME.tar.gz" "$NAME" && zip -qr "$NAME.zip" "$NAME")
log "wrote $OUT/$NAME.tar.gz ($(du -h "$OUT/$NAME.tar.gz" | cut -f1), sha256 $(shasum -a 256 "$OUT/$NAME.tar.gz" | cut -d' ' -f1))"
log "wrote $OUT/$NAME.zip ($(du -h "$OUT/$NAME.zip" | cut -f1), sha256 $(shasum -a 256 "$OUT/$NAME.zip" | cut -d' ' -f1))"
log "contents: $(find "$STAGE" -type f | wc -l | tr -d ' ') files; VERSION: $(head -1 "$STAGE/VERSION") $COMMIT"

64
bench/reference.json Normal file
View file

@ -0,0 +1,64 @@
{
"_about": "The team's reference numbers per card model for bench/ingest.mjs: the bench-log entry each comes from (docs/bench-log.md) and the three package runs of docs/benchmarks/repro.md. A submitted result agrees when every number is inside the tolerance the result file states (hash 3%, probe 10%, vectors and fingerprint exact). 'match' is a case-insensitive substring of the card name the workers report. The cpu gate is the chain's rule (under 10 ms per 32-lane warp on any core).",
"cpu_verify_ms_per_warp_max": 10,
"cards": [
{
"match": "RTX 5090",
"mhs": null,
"sweep_mhs": {
"4": null,
"64": null,
"256": null
},
"chase_gloads_1024": null,
"stream_gbps_1024": null,
"fingerprint": {
"24": "25f96e7dce90bd4e"
},
"source": "pending: the 5090-only re-run on PC 2 (the 6 October run was beside the live miner: 62.4 MH/s, not the card's figure; its fingerprint stands); the bench log's 5090 numbers before the package are 139.7 MH/s with the card to itself on a version 2 pack (4 October 2026, miner performance: variant racing) and 124.2 MH/s mining in the app"
},
{
"match": "Apple M5 Max",
"mhs": 27.674,
"sweep_mhs": {
"4": 275.247,
"64": 103.227,
"256": 56.625
},
"chase_gloads_1024": 3.401,
"stream_gbps_1024": 546.3,
"fingerprint": {
"24": "25f96e7dce90bd4e"
},
"source": "docs/benchmarks/repro-2026-10-06/mac (package v0.1.0-repro, Metal, 5 October 2026, run 2fdd7365db4a5deb, load average 6.19 at start); the bench log before the package: 27.9 MH/s through Apple OpenCL on the genesis pack (4 October 2026, proto-opencl README) and 26.7 MH/s mining through Metal"
},
{
"match": "gfx1201",
"mhs": null,
"sweep_mhs": {
"4": null,
"64": null,
"256": null
},
"chase_gloads_1024": null,
"stream_gbps_1024": null,
"fingerprint": {},
"source": "pending: the PC 1 package run; the bench log's 9070 XT numbers before the package are 18.0 to 18.1 MH/s in the bench and 2.42 to 2.68 G random reads/s at 1 GiB (5 October 2026, the 9070 XT on the eGPU)"
},
{
"match": "gfx1036",
"mhs": 3.312,
"sweep_mhs": {
"4": 3.722,
"64": 3.383,
"256": 3.32
},
"chase_gloads_1024": 0.458,
"stream_gbps_1024": 63.5,
"fingerprint": {
"24": "25f96e7dce90bd4e"
},
"source": "run-repro-pc2-20261006 (package v0.1.0-repro, AMD OpenCL, 6 October 2026, docs/benchmarks/repro-2026-10-06/pc2); the bench log before the package: 3.3 MH/s mining on the live devnet (4 October)"
}
]
}

242
bench/repro.ps1 Normal file
View file

@ -0,0 +1,242 @@
# Igneum reproducible benchmark, Windows (Linux and macOS: repro.sh). One command, no secrets, no node, no network.
#
# powershell -ExecutionPolicy Bypass -File repro.ps1 [-Seconds 120] [-Only "cuda:0,opencl:1"] [-Out results] [-NoProve] [-NoProbe] [-NoSweep]
#
# The same seven steps as repro.sh (docs/benchmarks/repro.md): the machine, the vectors on the CPU and every GPU, the hash
# benchmark for a fixed time per card on the published genesis pack, the random-read probe at 4, 64, 256 and 1024 MiB,
# the chip-resistance sweep, one fixture shard proven and verified when a 12 GB NVIDIA card and a prover host exist
# (IGNEUM_PROVE_HOST names a Linux path inside WSL, IGNEUM_WSL_DISTRO the distribution; otherwise the step says so),
# and one JSON result file plus a human table. Signed by nothing. No hostnames, no user names in the result.
#
# Nothing here earns anything: no pool, no wallet, no node, no network connection is made.
param(
[int] $Seconds = 120,
[string] $Only = "",
[string] $Out = "",
[switch] $NoProve,
[switch] $NoProbe,
[switch] $NoSweep,
[int] $BatchLog2 = 24
)
$ErrorActionPreference = 'Continue'
$Here = Split-Path -Parent $MyInvocation.MyCommand.Path
$Bin = Join-Path $Here 'bin\windows-x86_64'
$Pack = Join-Path $Here 'packs\igneum-genesis-mh'
if (-not $Out) { $Out = Join-Path $Here 'results' }
if (-not (Test-Path $Bin)) { Write-Error "no binaries at $Bin"; exit 2 }
if (-not (Test-Path (Join-Path $Pack 'vectors.json'))) { Write-Error "no pack at $Pack"; exit 2 }
New-Item -ItemType Directory -Force -Path $Out | Out-Null
$Stamp = (Get-Date).ToUniversalTime().ToString('yyyyMMddTHHmmssZ')
$RunId = -join ((1..16) | ForEach-Object { '{0:x}' -f (Get-Random -Maximum 16) })
$Json = Join-Path $Out "igneum-repro-windows-$Stamp.json"
$MdPath = Join-Path $Out "igneum-repro-windows-$Stamp.md"
$Log = Join-Path $Out "igneum-repro-windows-$Stamp.log"
Start-Transcript -Path $Log | Out-Null
Write-Output "igneum repro windows, run $RunId, $Stamp, $Seconds s per card, results in $Out"
$Version = 'unknown'; $Commit = ''
if (Test-Path (Join-Path $Here 'VERSION')) { $vl = Get-Content (Join-Path $Here 'VERSION'); $Version = $vl[0]; $Commit = ($vl | Where-Object { $_ -like 'commit *' } | Select-Object -First 1) -replace '^commit ', '' }
# PowerShell variable names are case-insensitive: no two variables here differ by case only (6 October 2026: the CPU result
# table held the CPU name variable of the same name, so it contained itself and ConvertTo-Json ran out of memory; the
# markdown lines wiped the markdown path). bench/jobs/ps-case-check.sh fails CI when such a pair comes back.
# ---- helpers ---------------------------------------------------------------------------------------------------------
function Sha256([string] $p) { (Get-FileHash -Algorithm SHA256 -LiteralPath $p).Hash.ToLower() }
function KV([string] $line, [string] $key) {
if ($line -match " $key=`"([^`"]*)`"") { return $Matches[1] }
if ($line -match " $key=([^ ]+)") { return $Matches[1] }
return ''
}
function Num([string] $v) { $d = 0.0; if ([double]::TryParse($v, [System.Globalization.NumberStyles]::Float, [System.Globalization.CultureInfo]::InvariantCulture, [ref] $d)) { return $d } else { return $null } }
$Commands = New-Object System.Collections.ArrayList
$script:Last = @()
function Run([string] $step, [string] $exe, [string[]] $argv) {
$t0 = Get-Date
Write-Host ("+ " + $exe + " " + ($argv -join ' '))
$outLines = @(& $exe @argv 2>&1 | ForEach-Object { "$_" })
$rc = $LASTEXITCODE
$script:Last = $outLines
$outLines | ForEach-Object { Write-Host $_ }
[void] $Commands.Add(@{ step = $step; cmd = ($exe + ' ' + ($argv -join ' ')); exit = $rc; seconds = [int] ((Get-Date) - $t0).TotalSeconds })
return $rc
}
function OnlyWants([string] $backend, [string] $index) { if (-not $Only) { return $true }; return (",$Only," -like "*,$backend`:$index,*") }
# ---- 1. the machine ------------------------------------------------------------------------------------------------
Write-Output "`n=== 1. machine ==="
$os = Get-CimInstance Win32_OperatingSystem
$cpuName = (Get-CimInstance Win32_Processor | Select-Object -First 1).Name.Trim()
$OsDesc = "$($os.Caption) $($os.Version) ($env:PROCESSOR_ARCHITECTURE)"
$RamMib = [int] ($os.TotalVisibleMemorySize / 1024)
Write-Output "os: $OsDesc"; Write-Output "cpu: $cpuName"; Write-Output "memory: $RamMib MiB"
$videoCtl = @(Get-CimInstance Win32_VideoController | ForEach-Object { @{ name = $_.Name; driver = $_.DriverVersion } })
$videoCtl | ForEach-Object { Write-Output ("video controller: " + $_.name + ", driver " + $_.driver) }
$NvSmi = @()
if (Get-Command nvidia-smi -ErrorAction SilentlyContinue) { $NvSmi = @(& nvidia-smi --query-gpu=index,name,driver_version,memory.total --format=csv,noheader 2>$null); $NvSmi | ForEach-Object { Write-Output "nvidia-smi: $_" } }
$Gpus = New-Object System.Collections.ArrayList
$Devices = New-Object System.Collections.ArrayList # @{backend, index, name}
$CudaExe = Join-Path $Bin 'igneum-worker-cuda.exe'
$ClExe = Join-Path $Bin 'igneum-worker-opencl.exe'
if (Test-Path $CudaExe) {
if ((Run 'list cuda' $CudaExe @('--list')) -eq 0) {
foreach ($l in $script:Last) { if ($l -like 'DEVICE *') {
$i = KV $l 'index'; $drv = ''
$smi = $NvSmi | Where-Object { $_ -like "$i, *" } | Select-Object -First 1
if ($smi) { $drv = ($smi -split ', ')[2] }
[void] $Gpus.Add(@{ backend = 'cuda'; index = [int] $i; name = (KV $l 'name'); arch = (KV $l 'arch'); driver = $drv; memory_mib = (Num (KV $l 'memory_mib')) })
[void] $Devices.Add(@{ backend = 'cuda'; index = $i; name = (KV $l 'name') })
} }
}
}
if (Test-Path $ClExe) {
if ((Run 'list opencl' $ClExe @('--list')) -eq 0) {
foreach ($l in $script:Last) { if ($l -like 'DEVICE *' -and (KV $l 'type') -eq 'GPU') {
[void] $Gpus.Add(@{ backend = 'opencl'; index = [int] (KV $l 'index'); name = (KV $l 'name'); vendor = (KV $l 'vendor'); platform = (KV $l 'platform'); driver = (KV $l 'driver'); memory_mib = (Num (KV $l 'memory_mib')) })
[void] $Devices.Add(@{ backend = 'opencl'; index = (KV $l 'index'); name = (KV $l 'name') })
} }
}
}
if ($Devices.Count -eq 0) { Write-Output "no GPU found by any worker (CPU checks still run)" }
# ---- binaries and pack files: sha256 of everything used ----------------------------------------------------------------
$Binaries = @(Get-ChildItem -File $Bin | ForEach-Object { @{ file = "bin/windows-x86_64/$($_.Name)"; sha256 = (Sha256 $_.FullName); bytes = $_.Length } })
$PackFiles = @{}
Get-ChildItem -File $Pack | ForEach-Object { $PackFiles[$_.Name] = Sha256 $_.FullName }
$ph = Get-Content (Join-Path $Pack 'program.h')
$PackId = (($ph | Where-Object { $_ -like '#define IGNEUM_PROGRAM_ID *' }) -replace '.*0x([0-9a-f]+)ull.*', '$1')
$PackSeed = (($ph | Where-Object { $_ -like '#define IGNEUM_SEED_STRING *' }) -replace '.*"(.*)".*', '$1')
# ---- 2a. the CPU --------------------------------------------------------------------------------------------------------
Write-Output "`n=== 2. vectors on the CPU ==="
$CpuResult = @{ vectors = 'not run' }
[void] (Run 'cpu check-pack' (Join-Path $Bin 'igneum-pow.exe') @('check-pack', '--pack', $Pack, '--warps', '50'))
$r = $script:Last | Where-Object { $_ -like 'RESULT check-pack*' } | Select-Object -Last 1
if ($r) { $CpuResult = @{ vectors = (KV $r 'verdict'); program_id = (KV $r 'program_id'); id_match = ((KV $r 'id_match') -eq 'true'); cache_match = ((KV $r 'cache_match') -eq 'true'); warps = (Num (KV $r 'warps')); warps_pass = (Num (KV $r 'warps_pass')); lanes = (Num (KV $r 'lanes')); lanes_pass = (Num (KV $r 'lanes_pass')); verify_ms_per_warp = (Num (KV $r 'cpu_verify_ms')); verify_cold_ms = (Num (KV $r 'cpu_verify_cold_ms')); cpu = $cpuName } }
# ---- 2 to 5, per card ----------------------------------------------------------------------------------------------------
$Cards = New-Object System.Collections.ArrayList
$Rows = New-Object System.Collections.ArrayList
$CardsMd = New-Object System.Collections.ArrayList
$Today = (Get-Date).ToUniversalTime().ToString('yyyy-MM-dd')
foreach ($d in $Devices) {
$backend = $d.backend; $index = "$($d.index)"; $name = $d.name
if (-not (OnlyWants $backend $index)) { Write-Output "skipping $backend`:$index $name (--only)"; continue }
Write-Output "`n=== card $backend`:$index $name ==="
$exe = if ($backend -eq 'cuda') { $CudaExe } else { $ClExe }
$secs = $Seconds; $cross = $false
# a card another backend already runs (NVIDIA through CUDA) is a cross-check of that path through OpenCL, not a second card
if ($backend -eq 'opencl' -and ($Devices | Where-Object { $_.backend -ne 'opencl' -and $_.name -eq $name })) { $cross = $true; $secs = [Math]::Max(10, [int] ($Seconds / 4)) }
$vectors = 'not run'; $fp = ''; $mhs = $null; $hashes = $null; $disp = $null; $bsecs = $null
$sweep = @{ '4' = $null; '64' = $null; '256' = $null; '1024' = $null }
[void] (Run "bench $backend`:$index" $exe @('--bench', '--pack', $Pack, '--device', $index, '--seconds', "$secs", '--batch-log2', "$BatchLog2"))
$rl = $script:Last | Where-Object { $_ -like 'RESULT bench*' } | Select-Object -Last 1
if ($rl) { $vectors = KV $rl 'check'; $fp = KV $rl 'fingerprint'; $mhs = Num (KV $rl 'mhs'); $hashes = Num (KV $rl 'hashes'); $disp = Num (KV $rl 'dispatches'); $bsecs = Num (KV $rl 'seconds'); $sweep['1024'] = $mhs }
else { $e = ($script:Last | Where-Object { $_ -match 'error|FAIL' } | Select-Object -First 1); $vectors = "FAIL (no result: $e)" }
if (-not $NoSweep -and $mhs -ne $null) {
foreach ($mib in @(4, 64, 256)) {
[void] (Run "sweep $backend`:$index $mib MiB" $exe @('--bench', '--pack', $Pack, '--device', $index, '--batches', '5', '--batch-log2', "$BatchLog2", '--dataset-mib', "$mib"))
$rl = $script:Last | Where-Object { $_ -like 'RESULT bench*' } | Select-Object -Last 1
if ($rl) { $sweep["$mib"] = Num (KV $rl 'mhs') }
}
}
$probe = @{}
$chase1024 = $null
if (-not $NoProbe) {
[void] (Run "memprobe $backend`:$index" $exe @('--memprobe', '--device', $index))
foreach ($rl in ($script:Last | Where-Object { $_ -like 'RESULT memprobe*' })) {
$mib = KV $rl 'mib'
if ($mib -eq '0') { $probe['alu_gops'] = Num (KV $rl 'alu_gops') }
else {
$probe[$mib] = @{ chase_gloads = (Num (KV $rl 'chase_gloads')); chase_latency_ns = (Num (KV $rl 'chase_ns')); chase_lanes = (Num (KV $rl 'chase_lanes')); indep_gloads = (Num (KV $rl 'indep_gloads')); line64_glines = (Num (KV $rl 'line64_glines')); stream_gbps = (Num (KV $rl 'stream_gbps')) }
if ($mib -eq '1024') { $chase1024 = Num (KV $rl 'chase_gloads') }
}
}
}
$chip = @{ in_cache_over_1gib = $null; share_of_random_read_ceiling = $null }
if ($sweep['4'] -ne $null -and $sweep['1024'] -ne $null -and $sweep['1024'] -gt 0) { $chip.in_cache_over_1gib = [Math]::Round($sweep['4'] / $sweep['1024'], 3) }
if ($mhs -ne $null -and $chase1024 -ne $null -and $chase1024 -gt 0) { $chip.share_of_random_read_ceiling = [Math]::Round($mhs * 128 / ($chase1024 * 1000), 3) }
[void] $Cards.Add(@{ backend = $backend; index = [int] $index; name = $name; cross_check = $cross; vectors = $vectors; fingerprint = $fp; hash = @{ mhs = $mhs; seconds = $bsecs; hashes = $hashes; dispatches = $disp; batch_log2 = $BatchLog2; dataset_mib = 1024; loads_per_hash = 128 }; sweep_mhs = $sweep; memprobe = $probe; chip = $chip })
$p1024 = if ($probe.ContainsKey('1024')) { $probe['1024'] } else { @{} }
[void] $CardsMd.Add("| $backend`:$index $name$(if ($cross) { ' (cross-check)' }) | $vectors | $fp | $mhs | $($sweep['4']) / $($sweep['64']) / $($sweep['256']) / $($sweep['1024']) | $($p1024.chase_gloads), $($p1024.chase_latency_ns) | $($p1024.stream_gbps) | $($chip.in_cache_over_1gib) | $($chip.share_of_random_read_ceiling) |")
if ($mhs -ne $null -and -not $cross) {
[void] $Rows.Add(@{ card = $name; generator = 'v2'; mh_s = $mhs; mh_per_w = $null; miner = "igneum-repro $Version ($backend bench, not mining)"; date = $Today; source = "repro:$RunId"; by = 'submitted'; note = "genesis pack, 128 loads per hash, 1 GiB dataset, $bsecs s, vectors $vectors" })
}
}
# ---- 6. proving (optional): a 12 GB NVIDIA card and a prover host inside WSL ------------------------------------------------
Write-Output "`n=== 6. proving (optional) ==="
$Prove = @{ status = 'skipped: -NoProve' }
if (-not $NoProve) {
$nvmem = 0
if ($NvSmi.Count -gt 0) { $nvmem = [int] ((($NvSmi[0] -split ', ')[3]) -replace '[^0-9]', '') }
$proveHost = $env:IGNEUM_PROVE_HOST
$distro = if ($env:IGNEUM_WSL_DISTRO) { $env:IGNEUM_WSL_DISTRO } else { 'Ubuntu-24.04' }
$wslUser = @(); if ($env:IGNEUM_WSL_USER) { $wslUser = @('-u', $env:IGNEUM_WSL_USER) } # the user the live prover runs as (its GPU server socket is per user)
$fix = Join-Path $Here 'proving\block-338-shard1.json'
if ($nvmem -lt 11000) { $Prove = @{ status = "skipped: no NVIDIA card with 12 GB or more (nvidia-smi reports $nvmem MiB)" } }
elseif (-not $proveHost) { $Prove = @{ status = 'skipped: no prover host (set IGNEUM_PROVE_HOST to the igneum-prove-host path inside WSL, built with --features igneum-prove-host/cuda, and IGNEUM_WSL_DISTRO)' } }
elseif (-not (Get-Command wsl.exe -ErrorAction SilentlyContinue)) { $Prove = @{ status = 'skipped: no WSL (the GPU prover runs on Linux)' } }
elseif (-not (Test-Path $fix)) { $Prove = @{ status = 'skipped: no fixture proving/block-338-shard1.json in the package' } }
else {
$fixL = '/mnt/' + $fix.Substring(0, 1).ToLower() + ($fix.Substring(2) -replace '\\', '/')
$outL = '/tmp/igneum-repro-prove-' + $Stamp + '.json'
[void] (Run 'prove id' 'wsl.exe' (@('-d', $distro) + $wslUser + @('--', 'bash', '-lc', "`"$proveHost`" --mode id")))
$idLine = $script:Last | Where-Object { $_ -like 'RESULT id*' } | Select-Object -First 1
$t0 = Get-Date
$prc = Run 'prove shard' 'wsl.exe' (@('-d', $distro) + $wslUser + @('--', 'bash', '-lc', "SP1_PROVER=cuda `"$proveHost`" '$fixL' --mode shard --shard 0 --out '$outL'"))
$wall = [int] ((Get-Date) - $t0).TotalSeconds
$resText = (& wsl.exe -d $distro @wslUser -- bash -lc "cat '$outL' 2>/dev/null") -join "`n"
$res = $null; try { $res = $resText | ConvertFrom-Json } catch { }
$verify = 'not run'; $vsecs = $null
if ($prc -eq 0 -and $res -and $res.proof_file) {
$vrc = Run 'prove verify' 'wsl.exe' (@('-d', $distro) + $wslUser + @('--', 'bash', '-lc', "`"$proveHost`" --mode verify --proof '$($res.proof_file)' --statement '$($res.statement)'"))
$verify = if ($vrc -eq 0) { 'VERIFIED' } else { 'NOT VERIFIED' }
$vl = $script:Last | Where-Object { $_ -like 'RESULT verify:*' } | Select-Object -First 1
if ($vl -match ' in ([0-9.]+) s') { $vsecs = Num $Matches[1] }
}
$hostSha = (& wsl.exe -d $distro @wslUser -- bash -lc "sha256sum `"$proveHost`" | cut -d' ' -f1") -join ''
$pstatus = if ($prc -eq 0) { 'run' } else { "failed (exit $prc)" }
$Prove = @{ status = $pstatus; host_sha256 = $hostSha.Trim(); pinned = "$idLine"; fixture = 'proving/block-338-shard1.json'; fixture_sha256 = (Sha256 $fix); shard = 0; wall_seconds = $wall; verify = $verify; verify_seconds = $vsecs }
if ($res) { foreach ($k in 'cycles', 'execute_seconds', 'core_prove_seconds', 'compressed_prove_seconds', 'compressed_verify_seconds', 'compressed_proof_bytes') { if ($res.PSObject.Properties[$k]) { $Prove[$k] = $res.$k } } }
}
}
Write-Output ("proving: " + ($Prove | ConvertTo-Json -Compress))
# ---- 7. the result file and the table ----------------------------------------------------------------------------------------
$Finished = (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ')
$Result = [ordered] @{
format = 'igneum-repro-1'
package = @{ version = $Version; commit = $Commit }
run = @{ id = $RunId; started = $Stamp; finished = $Finished; os = 'windows'; seconds_per_card = $Seconds; batch_log2 = $BatchLog2; only = $Only; script = 'repro.ps1'; script_sha256 = (Sha256 $MyInvocation.MyCommand.Path) }
machine = @{ os = $OsDesc; cpu = $cpuName; memory_mib = $RamMib; gpus = @($Gpus); video_controllers = @($videoCtl) }
binaries = @($Binaries)
pack = @{ dir = 'packs/igneum-genesis-mh'; seed = $PackSeed; program_id = $PackId; generator = 2; files_sha256 = $PackFiles }
cpu = $CpuResult
cards = @($Cards)
proving = $Prove
rows = @($Rows)
tolerance = @{ vectors = 'exact'; fingerprint = 'exact'; hash_mhs = 0.03; sweep_mhs = 0.10; chase_gloads = 0.10; chase_latency_ns = 0.15; stream_gbps = 0.10; cpu_verify_ms = 0.50; prove_seconds = 0.25 }
commands = @($Commands)
}
$Result | ConvertTo-Json -Depth 8 | Set-Content -Encoding UTF8 -LiteralPath $Json
$mdLines = @()
$mdLines += "# Igneum reproducible benchmark, windows, $Stamp"
$mdLines += ""
$mdLines += "Package $Version ($Commit), run $RunId, $Seconds s per card. Signed by nothing; the JSON beside this file is the record."
$mdLines += ""
$mdLines += "| Machine | |"; $mdLines += "|---|---|"; $mdLines += "| OS | $OsDesc |"; $mdLines += "| CPU | $cpuName |"; $mdLines += "| Memory | $RamMib MiB |"
foreach ($d in $Devices) { $mdLines += "| GPU ($($d.backend) $($d.index)) | $($d.name) |" }
$mdLines += ""
$mdLines += "| CPU vectors | verify ms per warp |"; $mdLines += "|---|---|"; $mdLines += "| $($CpuResult.vectors) ($($CpuResult.lanes_pass) of $($CpuResult.lanes) lanes) | $($CpuResult.verify_ms_per_warp) |"
$mdLines += ""
$mdLines += "| Card | Vectors | Fingerprint (2^$BatchLog2 at base 0) | MH/s at 1 GiB | 4 / 64 / 256 / 1024 MiB | Random reads at 1 GiB (G/s, ns) | Stream GB/s | In-cache / 1 GiB | Hash share of read ceiling |"
$mdLines += "|---|---|---|---|---|---|---|---|---|"
$mdLines += @($CardsMd)
$mdLines += ""
$mdLines += "Proving: $($Prove.status)$(if ($Prove.ContainsKey('verify')) { ", compressed proof $($Prove.compressed_prove_seconds) s, $($Prove.verify) in $($Prove.verify_seconds) s" })"
$mdLines += ""
$mdLines += "Every command run and the sha256 of every binary are in the JSON. Vectors: bit-exact means every lane of the pack's three published warps matched. The fingerprint is the FNV-1a 64 of all 2^$BatchLog2 outputs at base nonce 0: equal fingerprints on two machines mean every one of those hashes agreed."
$mdLines -join "`n" | Set-Content -Encoding UTF8 -LiteralPath $MdPath
Write-Output ""; Write-Output "wrote $Json"; Write-Output "wrote $MdPath"; Write-Output "log $Log"
Write-Output "done: $($Cards.Count) card run(s), CPU vectors $($CpuResult.vectors)"
Stop-Transcript | Out-Null
exit 0

290
bench/repro.sh Executable file
View file

@ -0,0 +1,290 @@
#!/usr/bin/env bash
# Igneum reproducible benchmark, Linux and macOS (Windows: repro.ps1). One command, no secrets, no node, no network.
#
# ./repro.sh [--seconds 120] [--only cuda:0,opencl:1,metal:0] [--out results] [--no-prove] [--no-probe] [--no-sweep]
#
# What it does, in order (docs/benchmarks/repro.md in the Igneum repository):
# 1. prints the machine: OS, CPU, memory, every GPU each worker can see (name, driver, memory)
# 2. checks the lottery hash vectors of the published genesis pack (packs/igneum-genesis-mh, generator 2, seed
# "igneum-genesis", day 2026-10-03, memory-hard 1 GiB dataset) on the CPU and on every GPU present: bit-exact or not
# 3. runs the hash benchmark for a fixed --seconds (120 by default) per card on that pack at 1 GiB
# 4. runs the random-read probe at 4, 64, 256 and 1024 MiB on every card (the dependent 4-byte chase is the hash's access pattern)
# 5. runs the chip-resistance sweep: the same program at 4, 64, 256 and 1024 MiB (inside a card's cache against beyond it)
# 6. proves one published fixture shard with the pinned guest and verifies it, when a 12 GB NVIDIA card and a built
# igneum-prove-host are present (IGNEUM_PROVE_HOST, or igneum-prove-host on the PATH); otherwise says so
# 7. writes results/igneum-repro-<os>-<time>.json (one object, the rows the Igneum bench table ingests, every command run,
# the sha256 of every binary used) and the same as a human table in .md. Signed by nothing. No hostnames, no user names.
#
# Nothing here earns anything: no pool, no wallet, no node, no network connection is made.
set -uo pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
SECONDS_PER_CARD=120; ONLY=""; OUT="$HERE/results"; PROVE=1; PROBE=1; SWEEP=1; BATCH_LOG2=24
while [ $# -gt 0 ]; do
case "$1" in
--seconds) SECONDS_PER_CARD="$2"; shift 2 ;;
--only) ONLY="$2"; shift 2 ;;
--out) OUT="$2"; shift 2 ;;
--no-prove) PROVE=0; shift ;;
--no-probe) PROBE=0; shift ;;
--no-sweep) SWEEP=0; shift ;;
--batch-log2) BATCH_LOG2="$2"; shift 2 ;;
-h|--help) sed -n '2,20p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
*) echo "unknown argument: $1" >&2; exit 2 ;;
esac
done
case "$(uname -s)" in
Darwin) OS=macos; BIN="$HERE/bin/macos-arm64" ;;
Linux) OS=linux; BIN="$HERE/bin/linux-x86_64" ;;
*) echo "unsupported OS $(uname -s) (Windows: repro.ps1)" >&2; exit 2 ;;
esac
PACK="$HERE/packs/igneum-genesis-mh"
[ -d "$BIN" ] || { echo "no binaries at $BIN (build the package with bench/make-package.sh, or download it)" >&2; exit 2; }
[ -f "$PACK/vectors.json" ] || { echo "no pack at $PACK" >&2; exit 2; }
mkdir -p "$OUT"
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
RUN_ID="$(head -c 8 /dev/urandom | od -An -tx1 | tr -d ' \n')"
JSON="$OUT/igneum-repro-$OS-$STAMP.json"
MD="$OUT/igneum-repro-$OS-$STAMP.md"
LOG="$OUT/igneum-repro-$OS-$STAMP.log"
exec > >(tee -a "$LOG") 2>&1
echo "igneum repro $OS, run $RUN_ID, $STAMP, $SECONDS_PER_CARD s per card, results in $OUT"
VERSION="$(cat "$HERE/VERSION" 2>/dev/null | head -1 || echo unknown)"
COMMIT="$(sed -n 's/^commit //p' "$HERE/VERSION" 2>/dev/null | head -1)"
# ---- helpers -------------------------------------------------------------------------------------------------------
sha256() { if command -v sha256sum >/dev/null 2>&1; then sha256sum "$1" | cut -d' ' -f1; else shasum -a 256 "$1" | cut -d' ' -f1; fi; }
jesc() { printf '%s' "$1" | sed 's/\\/\\\\/g; s/"/\\"/g' | tr -d '\r' | awk 'BEGIN{ORS="\\n"} {print}' | sed 's/\\n$//'; }
kv() { # kv "<line>" key: the value of key=... (quoted or bare), empty if absent
local v; v="$(printf '%s\n' "$1" | sed -n "s/.*[ ]$2=\"\([^\"]*\)\".*/\1/p")"
[ -n "$v" ] || v="$(printf '%s\n' "$1" | sed -n "s/.*[ ]$2=\([^ ]*\).*/\1/p")"
printf '%s' "$v"
}
num() { local v="$1"; case "$v" in ''|*[!0-9.eE+-]*) printf 'null' ;; *) printf '%s' "$v" ;; esac; }
COMMANDS="" # JSON array body of every command run
run_logged() { # run_logged <step> <cmd...>: runs, tees to the log, records the exact command and exit status; output in $LAST
local step="$1"; shift
local t0 t1 rc
t0=$(date +%s)
echo "+ $*"
LAST="$("$@" 2>&1)"; rc=$?
t1=$(date +%s)
printf '%s\n' "$LAST"
COMMANDS="$COMMANDS{\"step\":\"$(jesc "$step")\",\"cmd\":\"$(jesc "$*")\",\"exit\":$rc,\"seconds\":$((t1 - t0))},"
return $rc
}
only_wants() { # only_wants backend index: 0 when --only names other cards
[ -z "$ONLY" ] && return 0
case ",$ONLY," in *",$1:$2,"*) return 0 ;; esac
return 1
}
# ---- 1. the machine ------------------------------------------------------------------------------------------------
echo; echo "=== 1. machine ==="
if [ "$OS" = macos ]; then
OS_DESC="$(sw_vers -productName) $(sw_vers -productVersion) ($([ "$(sysctl -n hw.optional.arm64 2>/dev/null)" = 1 ] && echo arm64 || uname -m))"
CPU="$(sysctl -n machdep.cpu.brand_string 2>/dev/null || echo unknown)"
RAM_MIB=$(( $(sysctl -n hw.memsize) / 1048576 ))
else
OS_DESC="$( (. /etc/os-release 2>/dev/null && echo "$PRETTY_NAME") || uname -sr) ($(uname -m), kernel $(uname -r))"
CPU="$(sed -n 's/^model name[ \t]*: //p' /proc/cpuinfo | head -1)"; [ -n "$CPU" ] || CPU="$(uname -m)"
RAM_MIB=$(( $(sed -n 's/^MemTotal:[ \t]*\([0-9]*\) kB/\1/p' /proc/meminfo) / 1024 ))
fi
loadavg() { if [ "$OS" = macos ]; then sysctl -n vm.loadavg | tr -d '{}' | awk '{print $1}'; else cut -d' ' -f1 /proc/loadavg; fi; }
LOAD_BEFORE="$(loadavg)"
echo "os: $OS_DESC"; echo "cpu: $CPU"; echo "memory: $RAM_MIB MiB"; echo "load average (1 min) at start: $LOAD_BEFORE"
NVSMI=""
if command -v nvidia-smi >/dev/null 2>&1; then
NVSMI="$(nvidia-smi --query-gpu=index,name,driver_version,memory.total --format=csv,noheader 2>/dev/null || true)"
[ -n "$NVSMI" ] && { echo "nvidia-smi:"; printf '%s\n' "$NVSMI" | sed 's/^/ /'; }
fi
GPUS_JSON=""
DEVICES="" # lines "backend index name"
if [ "$OS" = macos ]; then
if run_logged "list metal" "$BIN/igneum-bench" --list; then
while IFS= read -r l; do
case "$l" in DEVICE*) GPUS_JSON="$GPUS_JSON{\"backend\":\"metal\",\"index\":$(kv "$l" index),\"name\":\"$(jesc "$(kv "$l" name)")\",\"driver\":\"Metal (macOS $(sw_vers -productVersion))\",\"memory_mib\":$(num "$(kv "$l" memory_mib)")},"; DEVICES="$DEVICES"$'\n'"metal $(kv "$l" index) $(kv "$l" name)" ;; esac
done <<< "$LAST"
fi
CL="$BIN/igneum-bench-cl"
else
CL="$BIN/igneum-worker-opencl"
if [ -x "$BIN/igneum-worker-cuda" ] && run_logged "list cuda" "$BIN/igneum-worker-cuda" --list; then
while IFS= read -r l; do
case "$l" in DEVICE*) GPUS_JSON="$GPUS_JSON{\"backend\":\"cuda\",\"index\":$(kv "$l" index),\"name\":\"$(jesc "$(kv "$l" name)")\",\"arch\":\"$(kv "$l" arch)\",\"driver\":\"$(printf '%s\n' "$NVSMI" | sed -n "$(( $(kv "$l" index) + 1 ))p" | awk -F', ' '{print $3}')\",\"memory_mib\":$(num "$(kv "$l" memory_mib)")},"; DEVICES="$DEVICES"$'\n'"cuda $(kv "$l" index) $(kv "$l" name)" ;; esac
done <<< "$LAST"
fi
fi
if [ -x "$CL" ] && run_logged "list opencl" "$CL" --list; then
while IFS= read -r l; do
case "$l" in DEVICE*)
[ "$(kv "$l" type)" = GPU ] || continue
GPUS_JSON="$GPUS_JSON{\"backend\":\"opencl\",\"index\":$(kv "$l" index),\"name\":\"$(jesc "$(kv "$l" name)")\",\"vendor\":\"$(jesc "$(kv "$l" vendor)")\",\"platform\":\"$(jesc "$(kv "$l" platform)")\",\"driver\":\"$(jesc "$(kv "$l" driver)")\",\"memory_mib\":$(num "$(kv "$l" memory_mib)")},"
DEVICES="$DEVICES"$'\n'"opencl $(kv "$l" index) $(kv "$l" name)" ;;
esac
done <<< "$LAST"
fi
[ -n "$DEVICES" ] || echo "no GPU found by any worker (CPU checks still run)"
# ---- binaries and pack files: sha256 of everything used --------------------------------------------------------------
BIN_JSON=""
for f in "$BIN"/*; do [ -f "$f" ] || continue; BIN_JSON="$BIN_JSON{\"file\":\"$(jesc "bin/$(basename "$BIN")/$(basename "$f")")\",\"sha256\":\"$(sha256 "$f")\",\"bytes\":$(wc -c < "$f" | tr -d ' ')},"; done
PACK_JSON=""
for f in "$PACK"/*; do [ -f "$f" ] || continue; PACK_JSON="$PACK_JSON\"$(basename "$f")\":\"$(sha256 "$f")\","; done
PACK_ID="$(sed -n 's/^#define IGNEUM_PROGRAM_ID 0x\([0-9a-f]*\)ull/\1/p' "$PACK/program.h")"
PACK_SEED="$(sed -n 's/^#define IGNEUM_SEED_STRING "\(.*\)"/\1/p' "$PACK/program.h")"
# ---- 2a. the CPU: the vectors through the Rust interpreter, and the verify time per warp --------------------------------
echo; echo "=== 2. vectors on the CPU ==="
CPU_JSON='{"vectors":"not run"}'
if run_logged "cpu check-pack" "$BIN/igneum-pow" check-pack --pack "$PACK" --warps 50; then :; fi
R="$(printf '%s\n' "$LAST" | grep '^RESULT check-pack' | tail -1)"
if [ -n "$R" ]; then
CPU_JSON="{\"vectors\":\"$(kv "$R" verdict)\",\"program_id\":\"$(kv "$R" program_id)\",\"id_match\":$(kv "$R" id_match),\"cache_match\":$(kv "$R" cache_match),\"warps\":$(num "$(kv "$R" warps)"),\"warps_pass\":$(num "$(kv "$R" warps_pass)"),\"lanes\":$(num "$(kv "$R" lanes)"),\"lanes_pass\":$(num "$(kv "$R" lanes_pass)"),\"verify_ms_per_warp\":$(num "$(kv "$R" cpu_verify_ms)"),\"verify_cold_ms\":$(num "$(kv "$R" cpu_verify_cold_ms)"),\"cpu\":\"$(jesc "$CPU")\"}"
fi
# ---- 2 to 5, per card ----------------------------------------------------------------------------------------------
CARDS_JSON=""
CARDS_MD=""
ROWS_JSON=""
TODAY="$(date -u +%Y-%m-%d)"
bench_card() { # bench_card backend index name
local backend="$1" index="$2" name="$3" exe secs="$SECONDS_PER_CARD" cross=false
local vectors="not run" fp="" mhs="" hashes="" disp="" bsecs="" check="" sweep4="" sweep64="" sweep256="" sweep1024="" probe_json="" sweep_json="" chip_json="" rl
echo; echo "=== card $backend:$index $name ==="
case "$backend" in
cuda) exe="$BIN/igneum-worker-cuda" ;;
opencl) exe="$CL" ;;
metal) exe="$BIN/igneum-bench" ;;
esac
# a card that another backend already runs (NVIDIA through CUDA, Apple through Metal) is a cross-check of that path through
# OpenCL (the bench log's fourth compiler path), not a second card: a quarter of the time, and no bench-table row
if [ "$backend" = opencl ] && printf '%s\n' "$DEVICES" | grep -v '^opencl ' | grep -q " $name\$"; then cross=true; secs=$(( SECONDS_PER_CARD / 4 )); [ "$secs" -ge 10 ] || secs=10; fi
if [ "$backend" = metal ]; then
# one run: the vectors (RESULT vectors), the timed bench (RESULT bench) and the sweep (RESULT sweep)
local args=(--pack "$PACK" --seconds "$secs" --batch-log2 "$BATCH_LOG2")
[ "$SWEEP" = 1 ] && args+=(--sweep)
run_logged "bench metal:$index" "$exe" "${args[@]}"
rl="$(printf '%s\n' "$LAST" | grep '^RESULT vectors' | tail -1)"; [ -n "$rl" ] && vectors="$(kv "$rl" verdict)"
rl="$(printf '%s\n' "$LAST" | grep '^RESULT bench' | head -1)"
if [ -n "$rl" ]; then fp="$(kv "$rl" fingerprint)"; mhs="$(kv "$rl" mhs)"; hashes="$(kv "$rl" hashes)"; disp="$(kv "$rl" dispatches)"; bsecs="$(kv "$rl" seconds)"; check="$(kv "$rl" check)"; fi
rl="$(printf '%s\n' "$LAST" | grep '^RESULT sweep' | tail -1)"
if [ -n "$rl" ]; then sweep4="$(kv "$rl" mhs_4)"; sweep64="$(kv "$rl" mhs_64)"; sweep256="$(kv "$rl" mhs_256)"; sweep1024="$(kv "$rl" mhs_1024)"; fi
else
run_logged "bench $backend:$index" "$exe" --bench --pack "$PACK" --device "$index" --seconds "$secs" --batch-log2 "$BATCH_LOG2"
rl="$(printf '%s\n' "$LAST" | grep '^RESULT bench' | tail -1)"
if [ -n "$rl" ]; then
check="$(kv "$rl" check)"; vectors="$check"; fp="$(kv "$rl" fingerprint)"; mhs="$(kv "$rl" mhs)"; hashes="$(kv "$rl" hashes)"; disp="$(kv "$rl" dispatches)"; bsecs="$(kv "$rl" seconds)"; sweep1024="$mhs"
else
vectors="FAIL (no result: $(printf '%s\n' "$LAST" | grep -i 'error\|FAIL' | head -1 | cut -c1-200))"
fi
if [ "$SWEEP" = 1 ] && [ -n "$mhs" ]; then
for mib in 4 64 256; do
run_logged "sweep $backend:$index $mib MiB" "$exe" --bench --pack "$PACK" --device "$index" --batches 5 --batch-log2 "$BATCH_LOG2" --dataset-mib "$mib"
rl="$(printf '%s\n' "$LAST" | grep '^RESULT bench' | tail -1)"
case "$mib" in 4) sweep4="$(kv "$rl" mhs)" ;; 64) sweep64="$(kv "$rl" mhs)" ;; 256) sweep256="$(kv "$rl" mhs)" ;; esac
done
fi
fi
if [ "$PROBE" = 1 ]; then
run_logged "memprobe $backend:$index" "$exe" --memprobe --device "$index"
while IFS= read -r rl; do
case "$rl" in "RESULT memprobe"*)
local mib; mib="$(kv "$rl" mib)"
if [ "$mib" = 0 ]; then probe_json="$probe_json\"alu_gops\":$(num "$(kv "$rl" alu_gops)"),"
else probe_json="$probe_json\"$mib\":{\"chase_gloads\":$(num "$(kv "$rl" chase_gloads)"),\"chase_latency_ns\":$(num "$(kv "$rl" chase_ns)"),\"chase_lanes\":$(num "$(kv "$rl" chase_lanes)"),\"indep_gloads\":$(num "$(kv "$rl" indep_gloads)"),\"line64_glines\":$(num "$(kv "$rl" line64_glines)"),\"stream_gbps\":$(num "$(kv "$rl" stream_gbps)")},"; fi ;;
esac
done <<< "$LAST"
fi
sweep_json="{\"4\":$(num "$sweep4"),\"64\":$(num "$sweep64"),\"256\":$(num "$sweep256"),\"1024\":$(num "$sweep1024")}"
# the chip-resistance numbers: in-cache against beyond-cache rate, and the hash's share of the card's dependent random-read ceiling at 1 GiB
local chase1024; chase1024="$(printf '%s\n' "$LAST" | grep '^RESULT memprobe' | grep ' mib=1024 ' | tail -1)"; chase1024="$(kv "$chase1024" chase_gloads)"
chip_json="$(awk -v a="$sweep4" -v b="$sweep1024" -v m="$mhs" -v c="$chase1024" 'BEGIN{
r1 = (a != "" && b != "" && b > 0) ? sprintf("%.3f", a / b) : "null";
r2 = (m != "" && c != "" && c > 0) ? sprintf("%.3f", m * 128 / (c * 1000)) : "null";
printf("{\"in_cache_over_1gib\":%s,\"share_of_random_read_ceiling\":%s}", r1, r2)}')"
local p1024; p1024="$(printf '%s\n' "$LAST" | grep '^RESULT memprobe' | grep ' mib=1024 ' | tail -1)"
CARDS_MD="$CARDS_MD| $backend:$index $name$([ "$cross" = true ] && echo ' (cross-check)') | $vectors | $fp | ${mhs:-null} | ${sweep4:-null} / ${sweep64:-null} / ${sweep256:-null} / ${sweep1024:-null} | $(kv "$p1024" chase_gloads), $(kv "$p1024" chase_ns) | $(kv "$p1024" stream_gbps) | $(printf '%s' "$chip_json" | sed -n 's/.*"in_cache_over_1gib":\([0-9.nul]*\).*/\1/p') | $(printf '%s' "$chip_json" | sed -n 's/.*"share_of_random_read_ceiling":\([0-9.nul]*\).*/\1/p') |
"
CARDS_JSON="$CARDS_JSON{\"backend\":\"$backend\",\"index\":$index,\"name\":\"$(jesc "$name")\",\"cross_check\":$cross,\"vectors\":\"$(jesc "$vectors")\",\"fingerprint\":\"$fp\",\"hash\":{\"mhs\":$(num "$mhs"),\"seconds\":$(num "$bsecs"),\"hashes\":$(num "$hashes"),\"dispatches\":$(num "$disp"),\"batch_log2\":$BATCH_LOG2,\"dataset_mib\":1024,\"loads_per_hash\":128},\"sweep_mhs\":$sweep_json,\"memprobe\":{${probe_json%,}},\"chip\":$chip_json},"
if [ -n "$mhs" ] && [ "$cross" = false ]; then
ROWS_JSON="$ROWS_JSON{\"card\":\"$(jesc "$name")\",\"generator\":\"v2\",\"mh_s\":$(num "$mhs"),\"mh_per_w\":null,\"miner\":\"igneum-repro $VERSION ($backend bench, not mining)\",\"date\":\"$TODAY\",\"source\":\"repro:$RUN_ID\",\"by\":\"submitted\",\"note\":\"genesis pack, 128 loads per hash, 1 GiB dataset, $bsecs s, vectors $vectors\"},"
fi
}
while read -r backend index name; do
[ -n "$backend" ] || continue
only_wants "$backend" "$index" || { echo "skipping $backend:$index $name (--only)"; continue; }
bench_card "$backend" "$index" "$name"
done <<< "$(printf '%s\n' "$DEVICES" | sed '/^$/d')"
[ -n "$CARDS_JSON" ] || echo "no card was benchmarked"
# ---- 6. one fixture shard proven and verified with the pinned guest, when the hardware and the host are present ---------
echo; echo "=== 6. proving (optional) ==="
PROVE_JSON='{"status":"skipped: --no-prove"}'
if [ "$PROVE" = 1 ]; then
HOST="${IGNEUM_PROVE_HOST:-$(command -v igneum-prove-host 2>/dev/null || true)}"
NVMEM="$(printf '%s\n' "$NVSMI" | head -1 | awk -F', ' '{print $4}' | sed 's/[^0-9]//g')"
FIX="$HERE/proving/block-338-shard1.json"
if [ "$OS" = macos ]; then PROVE_JSON='{"status":"skipped: the GPU prover needs a 12 GB NVIDIA card (SP1 cuda); this is macOS"}'
elif [ -z "$NVMEM" ] || [ "$NVMEM" -lt 11000 ]; then PROVE_JSON="{\"status\":\"skipped: no NVIDIA card with 12 GB or more (nvidia-smi reports ${NVMEM:-none} MiB)\"}"
elif [ -z "$HOST" ] || [ ! -x "$HOST" ]; then PROVE_JSON='{"status":"skipped: no igneum-prove-host (build proving/igneum-prove with --features igneum-prove-host/cuda and set IGNEUM_PROVE_HOST)"}'
elif [ ! -f "$FIX" ]; then PROVE_JSON='{"status":"skipped: no fixture proving/block-338-shard1.json in the package"}'
else
HOST_SHA="$(sha256 "$HOST")"
run_logged "prove id" "$HOST" --mode id
PID_LINE="$(printf '%s\n' "$LAST" | grep '^RESULT id' | head -1)"
WANT_ID="$(sed -n 's/.*"shard_program_id": *"\([^"]*\)".*/\1/p' "$HERE/proving/manifest.json" 2>/dev/null | head -1)"
OUTJ="$OUT/prove-$STAMP.json"
T0=$(date +%s)
SP1_PROVER=cuda run_logged "prove shard" "$HOST" "$FIX" --mode shard --shard 0 --out "$OUTJ"
PRC=$?; T1=$(date +%s)
jnum() { sed -n "s/.*\"$1\": *\([0-9.eE+-]*\).*/\1/p" "$OUTJ" 2>/dev/null | head -1; }
jstr() { sed -n "s/.*\"$1\": *\"\([^\"]*\)\".*/\1/p" "$OUTJ" 2>/dev/null | head -1; }
VERIFY="not run"; VSECS=""
if [ "$PRC" = 0 ] && [ -f "$OUTJ" ] && [ -n "$(jstr proof_file)" ]; then
run_logged "prove verify" "$HOST" --mode verify --proof "$(jstr proof_file)" --statement "$(jstr statement)"
VRC=$?; VERIFY="$([ "$VRC" = 0 ] && echo VERIFIED || echo "NOT VERIFIED")"
VSECS="$(printf '%s\n' "$LAST" | sed -n 's/^RESULT verify: [A-Z ]* in \([0-9.]*\) s.*/\1/p' | head -1)"
fi
PROVE_JSON="{\"status\":\"$([ "$PRC" = 0 ] && echo run || echo "failed (exit $PRC)")\",\"host_sha256\":\"$HOST_SHA\",\"pinned\":\"$(jesc "$PID_LINE")\",\"pinned_manifest_shard_id\":\"$WANT_ID\",\"fixture\":\"proving/block-338-shard1.json\",\"fixture_sha256\":\"$(sha256 "$FIX")\",\"shard\":0,\"cycles\":$(num "$(jnum cycles)"),\"execute_seconds\":$(num "$(jnum execute_seconds)"),\"core_prove_seconds\":$(num "$(jnum core_prove_seconds)"),\"compressed_prove_seconds\":$(num "$(jnum compressed_prove_seconds)"),\"compressed_verify_seconds\":$(num "$(jnum compressed_verify_seconds)"),\"compressed_proof_bytes\":$(num "$(jnum compressed_proof_bytes)"),\"wall_seconds\":$((T1 - T0)),\"verify\":\"$VERIFY\",\"verify_seconds\":$(num "$VSECS")}"
fi
fi
echo "proving: $PROVE_JSON"
# ---- 7. the result file and the table ----------------------------------------------------------------------------------
FINISHED="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
cat > "$JSON" <<EOF
{"format":"igneum-repro-1",
"package":{"version":"$(jesc "$VERSION")","commit":"$(jesc "$COMMIT")"},
"run":{"id":"$RUN_ID","started":"$STAMP","finished":"$FINISHED","os":"$OS","seconds_per_card":$SECONDS_PER_CARD,"batch_log2":$BATCH_LOG2,"only":"$(jesc "$ONLY")","script":"repro.sh","script_sha256":"$(sha256 "$0")"},
"machine":{"os":"$(jesc "$OS_DESC")","cpu":"$(jesc "$CPU")","memory_mib":$RAM_MIB,"load_1min_start":$(num "$LOAD_BEFORE"),"load_1min_end":$(num "$(loadavg)"),"gpus":[${GPUS_JSON%,}]},
"binaries":[${BIN_JSON%,}],
"pack":{"dir":"packs/igneum-genesis-mh","seed":"$(jesc "$PACK_SEED")","program_id":"$PACK_ID","generator":2,"files_sha256":{${PACK_JSON%,}}},
"cpu":$CPU_JSON,
"cards":[${CARDS_JSON%,}],
"proving":$PROVE_JSON,
"rows":[${ROWS_JSON%,}],
"tolerance":{"vectors":"exact","fingerprint":"exact","hash_mhs":0.03,"sweep_mhs":0.10,"chase_gloads":0.10,"chase_latency_ns":0.15,"stream_gbps":0.10,"cpu_verify_ms":0.50,"prove_seconds":0.25},
"commands":[${COMMANDS%,}]
}
EOF
{
echo "# Igneum reproducible benchmark, $OS, $STAMP"
echo
echo "Package $VERSION ($COMMIT), run $RUN_ID, $SECONDS_PER_CARD s per card. Signed by nothing; the JSON beside this file is the record."
echo
echo "| Machine | |"; echo "|---|---|"; echo "| OS | $OS_DESC |"; echo "| CPU | $CPU |"; echo "| Memory | $RAM_MIB MiB |"
printf '%s\n' "$DEVICES" | sed '/^$/d' | while read -r b i n; do echo "| GPU ($b $i) | $n |"; done
echo
echo "| CPU vectors | verify ms per warp |"; echo "|---|---|"
echo "| $(printf '%s' "$CPU_JSON" | sed -n 's/.*"vectors":"\([^"]*\)".*/\1/p') ($(printf '%s' "$CPU_JSON" | sed -n 's/.*"lanes_pass":\([0-9]*\).*/\1/p') of $(printf '%s' "$CPU_JSON" | sed -n 's/.*"lanes":\([0-9]*\).*/\1/p') lanes) | $(printf '%s' "$CPU_JSON" | sed -n 's/.*"verify_ms_per_warp":\([0-9.]*\).*/\1/p') |"
echo
echo "| Card | Vectors | Fingerprint (2^$BATCH_LOG2 at base 0) | MH/s at 1 GiB | 4 / 64 / 256 / 1024 MiB | Random reads at 1 GiB (G/s, ns) | Stream GB/s | In-cache / 1 GiB | Hash share of read ceiling |"
echo "|---|---|---|---|---|---|---|---|---|"
printf '%s' "$CARDS_MD"
echo
echo "Proving: $(printf '%s' "$PROVE_JSON" | sed -n 's/.*"status":"\([^"]*\)".*/\1/p')$(printf '%s' "$PROVE_JSON" | grep -q '"verify":"' && printf ', compressed proof %s s, %s in %s s' "$(printf '%s' "$PROVE_JSON" | sed -n 's/.*"compressed_prove_seconds":\([0-9.nul]*\).*/\1/p')" "$(printf '%s' "$PROVE_JSON" | sed -n 's/.*"verify":"\([^"]*\)".*/\1/p')" "$(printf '%s' "$PROVE_JSON" | sed -n 's/.*"verify_seconds":\([0-9.nul]*\).*/\1/p')")"
echo
echo "Every command run and the sha256 of every binary are in the JSON. Vectors: bit-exact means every lane of the pack's three published warps matched. The fingerprint is the FNV-1a 64 of all 2^$BATCH_LOG2 outputs at base nonce 0: equal fingerprints on two machines mean every one of those hashes agreed."
} > "$MD"
echo; echo "wrote $JSON"; echo "wrote $MD"; echo "log $LOG"
echo "done: $(printf '%s' "$CARDS_JSON" | grep -o '"backend"' | wc -l | tr -d ' ') card run(s), CPU vectors $(printf '%s' "$CPU_JSON" | sed -n 's/.*"vectors":"\([^"]*\)".*/\1/p')"

View file

@ -1597,3 +1597,113 @@ Reading: the kernel is the same 116.0 ms on both paths (18.08 MH/s pure kernel,
**A second defect found on the way: the pack export race.** PC 1's app log since its 19:02 UTC restart (`node tools/logs.mjs win-ae432dc7-20261005-190232`): `worker error: error 0 pack packs\devnet: the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT` at 19:07:03, 19:07:19 and 19:08:07, so the 9070 XT was not mining at all in the app while this entry was written (my job `rdna4-serve-1` at 18:43 hit the same folder in the same state). Cause, from `app/igneum-app/src/engine.rs` `prepare_worker`: one thread per card, each running `igneum-miner export-pack` into the one folder `packs\devnet`; across an epoch change the two exports interleave and the folder keeps one epoch's `program.h` with the other's `seeds.txt` until the next export. Fix on this branch: a process-wide mutex around both export sites (`EXPORT_LOCK`); the second export rewrites the same pack. Not measured in the app yet: it ships with the branch.
**Answer to the project lead.** The 9070 XT does 2.5 G random 4-byte reads per second from its memory for this access pattern, and the hash needs 128 of them, so about 19 MH/s is this card's ceiling for the current program class, on any slot; it was running at 92% of that. The eGPU link cost 6% per job through the read-back, now removed (17.87 against 16.88 MH/s inside jobs standalone). The duplicate platform that halved it to 8.9 + 9.4 is folded away. The pack race that stopped it is serialised. Nothing else in the worker's control moves the number: the next step for this card is the program class itself (fewer, wider loads per hash would favour AMD's 64-byte lines), which is a consensus question, not a worker one.
## 5 October 2026 (night), the reproducible benchmark package: one command per platform, run on the Mac and on PC 2; PC 1 owed
The package (`bench/`, `docs/benchmarks/repro.md`): `repro.sh` (Linux, macOS) and `repro.ps1` (Windows), built by
`bench/make-package.sh` from the existing cross-build scripts as `igneum-repro-v0.1.0-repro.tar.gz` and `.zip` (31 files;
the Metal, Apple OpenCL and CPU tools for macOS arm64, the NVRTC and OpenCL workers for Linux and Windows, `igneum-pow`
for all three, the genesis pack, the shard fixture and the pinned-guest manifest). No secrets, no node, no network. Seven
steps: the machine, the vectors on the CPU and every GPU, 120 s of hashing per card at 1 GiB on the genesis pack, the
random-read probe at 4, 64, 256 and 1024 MiB, the chip-resistance sweep at the same sizes, one shard proven and verified
when a 12 GB NVIDIA card and a prover host exist, one JSON result with every command and every binary's sha256. The workers
grew `--list`, `--bench --pack --seconds --dataset-mib` and `--memprobe` for it (the CUDA ones carried from the readwidth
branch, the OpenCL memprobe from opencl-rdna4, the Metal ones new), `igneum-pow` grew `check-pack`, and the pack reader
accepts the genesis pack's string seed.
Commands: `bench/make-package.sh --tag v0.1.0-repro` (under the build lock, 2 min 11 s); on the Mac
`tools/lock/with-lock.sh measure bash ./repro.sh --seconds 120` from the extracted archive (load average 6.19 at the start,
4.15 at the end; the Igneum Miner app's Metal worker was off, "paused"); result files in
`docs/benchmarks/repro-2026-10-06/mac/`.
Apple M5 Max (40 GPU cores, 64 GB, macOS 26.6.2), package v0.1.0-repro, run 2fdd7365db4a5deb, 21:51 UTC:
| Step | Metal | Apple OpenCL (cross-check, 30 s) | CPU |
|---|---|---|---|
| Vectors (3 warps, 96 lanes, cache FNV) | 96/96 PASS, program id bcc1248b10cc90f2 | 96/96 PASS | 96/96 PASS, FNV 48c4f5bf24166b2e |
| Fingerprint of 2^24 outputs at base 0 | 25f96e7dce90bd4e | 25f96e7dce90bd4e | |
| Hash at 1 GiB | 27.674 MH/s, 120.5 s GPU time, 194 dispatches, 3.25 G hashes | 26.875 MH/s (wall) | |
| Sweep 4 / 64 / 256 / 1024 MiB | 275.2 / 103.2 / 56.6 / 27.8 MH/s | 335.2 / 104.7 / 47.2 / 26.9 | |
| Random reads at 1 GiB (chase, best lanes) | 3.401 G loads/s; 522 ns per dependent load at 256 lanes | 3.404 G loads/s; 1,273 ns (wall, launch included) | |
| Random reads at 4 / 64 / 256 MiB | 43.8 / 12.9 / 7.3 G loads/s | 37.0 / 12.9 / 6.6 | |
| Stream at 1 GiB | 546 GB/s | 515 GB/s | |
| Integer chain | 4,263 G int ops/s (5 ops per step counted, approximate) | 3,952 | |
| CPU verify per 32-lane warp | | | 0.611 ms (avg of 50), cold 0.625 |
| Proving | skipped: needs a 12 GB NVIDIA card | | |
Against the bench log: Metal 27.674 MH/s is 0.8% under the 27.9 the README measured through Apple OpenCL on this pack
(4 October) and 3.6% over the 26.7 the app mined at on the live devnet; the CPU verify 0.611 ms against 0.631 ms
(4 October). The Apple OpenCL path is 2.9% under Metal here where the 3 October version 1 runs had them equal, and 3.7%
under the README's figure: outside the package's 3% tolerance, one 30-s wall-time run against one 5-batch run; the card's
row is the Metal one. The 2^24 fingerprint is the same through Metal and through Apple's OpenCL compiler (16.7 million
nonces), and the same value through the CUDA worker under CPU emulation at 2^14 (e7d68ec2a49d0671, equal to Apple
OpenCL's at 2^14). The sweep is the first on a version 2 program on Apple: in-cache over 1 GiB 9.9x against 12.9x on the
version 1 program of 3 October. The hash's share of the random-read ceiling is 1.04: 27.674 x 128 = 3.54 G loads/s against
a single-chain chase of 3.40, so on this chip the hash sits at the limit the probe sees, and the chase at 4 M lanes is
not a ceiling above the hash (the eight-chain probe gives 3.50).
Consequences (the rule of 5 October): the bench needs 1.3 GB of device memory, so every tier down to an 8 GB card and an
integrated chip runs it; about 4 min per card, `--only` for a rig. The proving step's gate (a 12 GB card) is the design
floor, and the only card that has proven the fixture peaked at 28.3 GB (the proving agent, PC 2, tonight), so a 12, 16 or
24 GB owner who runs step 6 today fails late on memory: the next package gates on the measured peak until the prover fits
12 GB, and evidence row 16 stays "designed".
PC 1 (ae432dc7): the first job (run-repro-pc1-20261005, 21:04 to 21:09 UTC) switched the 5090, the 9070 XT and the gfx1036
off and on for nothing: `repro.ps1` had three PowerShell parse errors (two missing parentheses on the wsl.exe lines, one
`(if ...)` as an expression in a hashtable), found by the Windows parser on the PC because the Mac has no PowerShell;
the second PowerShell playbook tonight to fail that way. The class is closed: the PC job parses the package script with
`[System.Management.Automation.Language.Parser]::ParseFile` before any card is touched and exits 2 with the line, and
`bench/` is in the Windows PowerShell 5.1 parse job of `.github/workflows/windows.yml` (every `.ps1` under it is parsed on
each push, like the relay playbooks). Also found there: the RX 9070 XT (gfx1201) was on PC 1's bus at 21:06 UTC and the
app switched it off and on, where the Counter ASIC 2.0 queue had it absent since 20:40 UTC. PC 1 went down at 22:31:06
UTC before the second slot; its run is owed to the morning. Two package defects found by the Mac re-runs and fixed the
same hour: a `--skip-build` package shipped without the two Mac workers (built straight into the stage, which every
build deletes), and the Apple OpenCL worker was built x86_64 under the Rosetta shell and passed the arm64 check on its
path name (`file -b` now, every binary listed by name).
PC 2 (1ccfe586, RTX 5090, gfx1036, Windows 11): run job run-repro-pc2-20261006 (00:44 to 01:09 UTC, 6 October, on Igneum Miner 0.3.11; the fetch job first). The 5090
block ran BESIDE THE APP'S LIVE CUDA WORKER: the card-off POST was accepted but `api/state` answered `{}` (a 0.3.11
defect on a proving machine) and nothing confirmed a stopped worker; nvidia-smi showed 10,176 MiB in use before the bench
and 348 W after it. The 5090 rates are shared-card figures and not the card's row; the checks stand.
| Step | RTX 5090, CUDA (NVRTC sm_120, wall time), shared card | RTX 5090, NVIDIA OpenCL (event time), shared | gfx1036, AMD OpenCL (event time) | CPU (Ryzen 7 9800X3D) |
|---|---|---|---|---|
| Vectors | 96/96 PASS, program id bcc1248b10cc90f2 | 96/96 PASS | 96/96 PASS | 96/96 PASS, FNV 48c4f5bf24166b2e |
| Fingerprint of 2^24 outputs at base 0 | 25f96e7dce90bd4e | 25f96e7dce90bd4e | 25f96e7dce90bd4e | |
| Hash at 1 GiB | 62.412 MH/s, 120.2 s, 447 dispatches (beside the miner) | 62.283 MH/s, 29.9 s (beside the miner) | 3.312 MH/s, 121.6 s, 24 dispatches | |
| Sweep 4 / 64 / 256 / 1024 MiB | 374 / 367 / 74 / 62 | 375 / 386 / 72 / 62 | 3.72 / 3.38 / 3.32 / 3.31 | |
| Random reads at 1 GiB (chase, best lanes) | 17.5 G loads/s; 1,766 ns (wall, launch included) | 18.2 G loads/s; 415 ns | 0.458 G loads/s; 566 ns | |
| Stream at 1 GiB | 339 GB/s (wall, launch-bound) | 1,662 GB/s | 63.5 GB/s | |
| Integer chain | 13.4 T int ops/s (wall) | 45.2 T | 0.20 T | |
| CPU verify per warp | | | | the line is in the intake (the check-pack RESULT line of the job) |
| Proving (live host /opt/igneum/igneum-prove-host, pinned shard id 0x2b1a81cb..., SP1 cuda, the live prover switched off first) | execute 1.48 s (60.4 M cycles), core 20.2 s (18.1 MB, verify 0.60 s), compressed 33.9 s (1.27 MB, verify 0.038 s), VERIFIED; two tampered witnesses REJECTED | | | |
| Job size, 2^21 against 2^24 (the consequences reviewer's question) | genesis pack 60.0 against 63.2 MH/s; live devnet pack (program id 5c7e497fda6ef223) 60.9 against 63.4 | | | |
Readings. The version 2 genesis pack is now bit-exact on real NVIDIA silicon through NVRTC and through NVIDIA's OpenCL,
and on real AMD silicon (the integrated gfx1036), with the same 2^24 fingerprint as Metal and Apple OpenCL on the Mac:
five compilers, three vendors, 16.7 million nonces each, one value (evidence row 4's owed run). The gfx1036 is the only
PC 2 card with a clean figure: 3.312 MH/s at 1 GiB, flat across 4 MiB to 1 GiB (its memory is the system's), at 0.93 of
its random-read ceiling (3.312 x 128 = 0.424 G loads/s against a 0.458 G chase), the same reading as the 9070 XT's 0.92
(5 October). The 5090's chase of 17.5 to 18.2 G loads/s at 1 GiB and stream of 1,662 GB/s (against the 1,638 GB/s fill
of 3 October) are the card's, probe bursts being short; its hash, sweep, proof and job-size rows are shared-card
figures: 62.4 MH/s against 131 to 137 in every other run of the night, 33.9 s compressed against the proving agent's 10.8
to 11.4 s alone and 33 s "with the miner running" the same night. The job-size answer on the shared card is 5% (60.0
against 63.2 at 2^21 against 2^24), against the 1.8% between the app's wall and inside-jobs rates (114.0 against 116.0,
PC 2's STATUS lines at 22:40 UTC); the reviewer's 15% gap between the bench and the app is not the job size alone, and
the clean 2^21 row waits for the 5090-only re-run. Consequence either way: a 2^22 or 2^23 job for cards over 100 MH/s
costs nothing and removes whatever part of the gap is per-job work; it goes in the next cut's list as a measurement,
not a change, until the clean pair exists.
Three defects of the night, each with its class closed: (1) the card-off step trusted `api/state` (the `{}` defect) and
a fixed wait, so the 5090 ran shared; `pc-repro.ps1` now confirms by nvidia-smi's compute-apps list (no igneum worker on
the card) and labels a run "card-off-UNCONFIRMED" when it cannot, as the M16 job does. (2) The same reading left PC 2's
live prover OFF from 00:49 UTC until run-prover-on-pc2-20261006 restored it (the script read "was on" from `api/state`,
got nothing, and restored nothing); the state now comes from settings.json and the restore is unconditional. (3) The
result writer on Windows threw `System.OutOfMemoryException` in ConvertTo-Json and gave Set-Content an empty path, so PC 2
wrote no JSON and no markdown: PowerShell variable names are case-insensitive, and `repro.ps1` used `$Cpu` for the result
table and `$cpu` for the CPU name it held (the table contained itself) and `$md` for the markdown lines over `$Md` for its
path. Distinct names now, and `bench/jobs/ps-case-check.sh` fails CI on any pair of variable names that differ only by
case (shown to fire on a bad file and pass on the real ones). The three per-card transcript logs are in
`docs/benchmarks/repro-2026-10-06/pc2/`; the RESULT lines of the job are the record. The class for the other playbook
owners: a PowerShell playbook written on the Mac meets the 5.1 parser, the case-insensitive variable table and
`api/state`'s `{}` only on the PC; parse first, confirm by the device's own tools, name variables distinctly.

View file

@ -0,0 +1,13 @@
{"format":"igneum-repro-1",
"package":{"version":"v0.1.0-repro","commit":"39141f5"},
"run":{"id":"2fdd7365db4a5deb","started":"20261005T215136Z","finished":"2026-10-05T21:55:24Z","os":"macos","seconds_per_card":120,"batch_log2":24,"only":"","script":"repro.sh","script_sha256":"2d8d3e4ecba1fa01326ef9daf4ca8f7f776f606674ce826dd1edbb059946fb1e"},
"machine":{"os":"macOS 26.6.2 (arm64)","cpu":"Apple M5 Max","memory_mib":65536,"load_1min_start":6.19,"load_1min_end":4.15,"gpus":[{"backend":"metal","index":0,"name":"Apple M5 Max","driver":"Metal (macOS 26.6.2)","memory_mib":53084},{"backend":"opencl","index":0,"name":"Apple M5 Max","vendor":"Apple","platform":"Apple","driver":"1.2 1.0","memory_mib":53084}]},
"binaries":[{"file":"bin/macos-arm64/igneum-bench","sha256":"19927950c40f9d001d0ec9ed6a90e6218ec066eedd0b3ec132107fa6c61f8ba9","bytes":689961},{"file":"bin/macos-arm64/igneum-bench-cl","sha256":"49ed1b558cccd631dc70c7f96e5f501e44bc613775b08d9aa9dbf24a051f8308","bytes":122220},{"file":"bin/macos-arm64/igneum-pow","sha256":"2198a20203b3e8dc8ae06ea5ddea94359dfa06e7de6ef80f7782a4d34b6e9786","bytes":659392}],
"pack":{"dir":"packs/igneum-genesis-mh","seed":"igneum-genesis","program_id":"bcc1248b10cc90f2","generator":2,"files_sha256":{"kernel.cl":"ad45c0d447e45c354c64bd27275df6a123dcc0a77bccbf11540191aeec5e5378","kernel.cu":"f40fa27a78d2be3560ac6f42d6cbc8d858620e344c403bcc5e5d9bfd217d7f79","kernel_bound.cl":"6d2ae7909d842cdce48f27a07c05fc522774691c4decc1c3d9b43d9cb576f9e8","kernel_bound.cu":"d3ac56050c014856f2a5d31150a829662cc2956477656b6a36cefb8ebd2104c1","memhard.h":"9040c6415fac0d9792c9486ddd02e1d8d429d48af3065f0707055700c9771ce1","memhard.metal":"1ad3381e444b3583edcb66bed839e7c072139d8728b828d5a99b6dcfc236729e","program.h":"8e0ff3469934f882d00eccc039a5ef612c95dd51f76da0485e65ea52884097cd","program.json":"2bab5c397ecdcdbc4104380ad468f41ee69f638a162ce07e81b50aa1d0ce0cd1","program.metal":"f990bd9791f108d7b0cff6b3c54ab54ef05c8613b62da8002a715fb9c9f0cc79","program_bound.metal":"15dde5c74d30745d66149917b5ead855a6d02a0106f6e4da5f76e04eab6a1b4a","vectors.h":"660ff7a57824550fa17e52993a1f1a4c695c1feaa49458f39ecf59decd1433be","vectors.json":"fdd286c014d7f4776c6b400da0ff80e1c77db858d263697464f9cb496e5b3569"}},
"cpu":{"vectors":"PASS","program_id":"bcc1248b10cc90f2","id_match":true,"cache_match":true,"warps":3,"warps_pass":3,"lanes":96,"lanes_pass":96,"verify_ms_per_warp":0.611,"verify_cold_ms":0.686,"cpu":"Apple M5 Max"},
"cards":[{"backend":"metal","index":0,"name":"Apple M5 Max","cross_check":false,"vectors":"PASS","fingerprint":"25f96e7dce90bd4e","hash":{"mhs":27.674,"seconds":120.035,"hashes":3321888768,"dispatches":198,"batch_log2":24,"dataset_mib":1024,"loads_per_hash":128},"sweep_mhs":{"4":275.247,"64":103.227,"256":56.625,"1024":27.762},"memprobe":{"4":{"chase_gloads":43.752,"chase_latency_ns":240.1,"chase_lanes":4194304,"indep_gloads":42.194,"line64_glines":41.616,"stream_gbps":662.3},"64":{"chase_gloads":12.859,"chase_latency_ns":393.1,"chase_lanes":16384,"indep_gloads":11.512,"line64_glines":15.350,"stream_gbps":697.5},"256":{"chase_gloads":7.262,"chase_latency_ns":476.2,"chase_lanes":16384,"indep_gloads":7.170,"line64_glines":7.520,"stream_gbps":568.9},"1024":{"chase_gloads":3.401,"chase_latency_ns":521.5,"chase_lanes":4194304,"indep_gloads":3.505,"line64_glines":3.477,"stream_gbps":546.3},"alu_gops":4262.9},"chip":{"in_cache_over_1gib":9.915,"share_of_random_read_ceiling":1.042}},{"backend":"opencl","index":0,"name":"Apple M5 Max","cross_check":true,"vectors":"PASS","fingerprint":"25f96e7dce90bd4e","hash":{"mhs":26.875,"seconds":30.589,"hashes":822083584,"dispatches":49,"batch_log2":24,"dataset_mib":1024,"loads_per_hash":128},"sweep_mhs":{"4":335.165,"64":104.658,"256":47.209,"1024":26.875},"memprobe":{"4":{"chase_gloads":34.990,"chase_latency_ns":867.2,"chase_lanes":4194304,"indep_gloads":36.010,"line64_glines":40.001,"stream_gbps":28.5},"64":{"chase_gloads":12.843,"chase_latency_ns":1003.9,"chase_lanes":4194304,"indep_gloads":12.898,"line64_glines":9.383,"stream_gbps":171.2},"256":{"chase_gloads":6.743,"chase_latency_ns":1363.3,"chase_lanes":4194304,"indep_gloads":6.918,"line64_glines":6.580,"stream_gbps":406.1},"1024":{"chase_gloads":3.404,"chase_latency_ns":1273.4,"chase_lanes":4194304,"indep_gloads":3.449,"line64_glines":3.380,"stream_gbps":515.0},"alu_gops":4100.6},"chip":{"in_cache_over_1gib":12.471,"share_of_random_read_ceiling":1.011}}],
"proving":{"status":"skipped: the GPU prover needs a 12 GB NVIDIA card (SP1 cuda); this is macOS"},
"rows":[{"card":"Apple M5 Max","generator":"v2","mh_s":27.674,"mh_per_w":null,"miner":"igneum-repro v0.1.0-repro (metal bench, not mining)","date":"2026-10-05","source":"repro:2fdd7365db4a5deb","by":"submitted","note":"genesis pack, 128 loads per hash, 1 GiB dataset, 120.035 s, vectors PASS"}],
"tolerance":{"vectors":"exact","fingerprint":"exact","hash_mhs":0.03,"sweep_mhs":0.10,"chase_gloads":0.10,"chase_latency_ns":0.15,"stream_gbps":0.10,"cpu_verify_ms":0.50,"prove_seconds":0.25},
"commands":[{"step":"list metal","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench --list","exit":0,"seconds":0},{"step":"list opencl","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --list","exit":0,"seconds":0},{"step":"cpu check-pack","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-pow check-pack --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --warps 50","exit":0,"seconds":1},{"step":"bench metal:0","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --seconds 120 --batch-log2 24 --sweep","exit":0,"seconds":129},{"step":"memprobe metal:0","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench --memprobe --device 0","exit":0,"seconds":28},{"step":"bench opencl:0","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --device 0 --seconds 30 --batch-log2 24","exit":0,"seconds":32},{"step":"sweep opencl:0 4 MiB","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 4","exit":0,"seconds":0},{"step":"sweep opencl:0 64 MiB","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 64","exit":0,"seconds":2},{"step":"sweep opencl:0 256 MiB","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 256","exit":0,"seconds":2},{"step":"memprobe opencl:0","cmd":"/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --memprobe --device 0","exit":0,"seconds":28}]
}

View file

@ -0,0 +1,467 @@
igneum repro macos, run 2fdd7365db4a5deb, 20261005T215136Z, 120 s per card, results in /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/results
=== 1. machine ===
os: macOS 26.6.2 (arm64)
cpu: Apple M5 Max
memory: 65536 MiB
load average (1 min) at start: 6.19
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench --list
metal devices (1):
[0] Apple M5 Max maxBufferLength 39813 MiB unified true recommendedMaxWorkingSetSize 53084 MiB
DEVICE index=0 backend=metal name="Apple M5 Max" memory_mib=53084 unified=true low_power=false
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --list
igneum-bench-cl pack "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000" (test harness: no pool, no network, no wallet)
OpenCL devices (1):
[0] Apple M5 Max | Apple (OpenCL 1.2 (Jul 31 2026 20:36:30))
GPU, vendor Apple, driver 1.2 1.0, OpenCL C 1.2 , 40 compute units, 1000 MHz
global 53084 MiB, max alloc 9953 MiB, local 32 KiB, max work-group 256, sub-group extension: none
DEVICE index=0 backend=opencl type=GPU name="Apple M5 Max" vendor="Apple" platform="Apple" driver="1.2 1.0" compute_units=40 memory_mib=53084
=== 2. vectors on the CPU ===
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-pow check-pack --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --warps 50
check-pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh: seed "igneum-genesis" generator v2 attempt 0 program id bcc1248b10cc90f2 (matches program.json), dataset 2^28 words (memory-hard), 128 loads/hash; built in 175 ms
cache: FNV-1a 64 48c4f5bf24166b2e == vectors.json (48c4f5bf24166b2e)
vector warp base 0 (nonces 0..31): PASS (32 of 32 lanes)
vector warp base 4096 (nonces 4096..4127): PASS (32 of 32 lanes)
vector warp base 1000000 (nonces 1000000..1000031): PASS (32 of 32 lanes)
CPU verify: 0.611 ms per 32-lane warp, avg of 50 (cold 0.686 ms max over the vector warps; checksum 19297e99c7b9a55e)
RESULT check-pack pack=/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh seed=igneum-genesis program_id=bcc1248b10cc90f2 id_match=true dataset_log2=28 mode=memory-hard cache_fnv=48c4f5bf24166b2e cache_match=true warps=3 warps_pass=3 lanes=96 lanes_pass=96 cpu_verify_ms=0.611 cpu_verify_cold_ms=0.686 verdict=PASS
=== card metal:0 Apple M5 Max ===
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --seconds 120 --batch-log2 24 --sweep
pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh: seed "igneum-genesis", day "2026-10-03", dataset 2^28 words (memory-hard), program id bcc1248b10cc90f2, 3 vector warps
igneum-bench
GPU: Apple M5 Max (maxBufferLength 39813 MiB, unified memory true)
dataset: 2^28 uint32 = 1024 MiB, day "2026-10-03", construction memory-hard (default; MEMHARD.md)
dataset kernels: compile 1.09 ms
cache fill (GPU): 2.03 ms GPU time first (11.26 ms wall), 2.01 ms GPU time second; 65536 chains x 64 ChaCha12 blocks, 256 MiB written
dataset build (memory-hard): 29.96 ms GPU time first (56.49 ms wall), 21.12 ms second; 16777216 items, 794.3 M items/s, 6.35 G cache-line reads/s (second)
cache: CPU fill 191.8 ms on one core (2^26 words, 65536 chains of 64 ChaCha12 blocks)
cache check: GPU cache == CPU cache, all 67108864 words compared, FNV-1a 64 48c4f5bf24166b2e
dataset sample check (memory-hard, 2^24 words): 1024 GPU words vs CPU derivation, 0 mismatches
vector warp base 0 (nonces 0..31): PASS (32 of 32 lanes)
vector warp base 4096 (nonces 4096..4127): PASS (32 of 32 lanes)
vector warp base 1000000 (nonces 1000000..1000031): PASS (32 of 32 lanes)
cache: FNV-1a 64 48c4f5bf24166b2e == vectors.json (48c4f5bf24166b2e)
RESULT vectors backend=metal device="Apple M5 Max" pack=/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 id_match=true cache_match=true warps=3 warps_pass=3 lanes=96 lanes_pass=96 verdict=PASS
=== epoch seed "igneum-genesis" ===
program: 64 instructions x 8 iterations, loads/hash = 128 (wide 0), distinct items per warp = 4096
op mix: load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1
compile: library 0.04 ms, pipeline 0.10 ms, total 0.14 ms
threadExecutionWidth = 32, maxTotalThreadsPerThreadgroup = 1024
warm-up batch: 16777216 hashes in 618.95 ms wall, 615.94 ms GPU
batch fingerprint (FNV-1a 64 of 2^24 outputs at base nonce 0): 25f96e7dce90bd4e
timed: 198 batches x 16777216 hashes = 3321888768 hashes in 120.1 s wall (--seconds 120)
wall 120083.91 ms -> 27.663 Mhash/s, 14.16 GB/s useful (loads x 4 B)
GPU 120034.97 ms -> 27.674 Mhash/s, 14.17 GB/s useful (loads x 4 B)
verify warp 0 (nonces 0..31): PASS cpu 2.150 ms single, 1.068 ms avg of 20, 4096 items derived
verify warp 262145 (nonces 8388640..8388671): PASS cpu 1.813 ms single, 1.021 ms avg of 20, 4096 items derived
verify warp 524287 (nonces 16777184..16777215): PASS cpu 1.829 ms single, 1.048 ms avg of 20, 4096 items derived
RESULT bench backend=metal device="Apple M5 Max" seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=1024 loads=128 check=PASS cpu_warps=3 fingerprint=25f96e7dce90bd4e batch_log2=24 dispatches=198 hashes=3321888768 seconds=120.035 wall_seconds=120.084 mhs=27.674 mhs_wall=27.663 time=gpu
=== sweep (the same program, dataset 4, 64, 256 and 1024 MiB) ===
=== epoch seed "igneum-genesis" ===
program: 64 instructions x 8 iterations, loads/hash = 128 (wide 0), distinct items per warp = 4096
op mix: load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1
compile: library 0.05 ms, pipeline 0.09 ms, total 0.14 ms
threadExecutionWidth = 32, maxTotalThreadsPerThreadgroup = 1024
warm-up batch: 16777216 hashes in 53.07 ms wall, 49.13 ms GPU
batch fingerprint (FNV-1a 64 of 2^24 outputs at base nonce 0): 1468be5e1a85e771
timed: 5 batches x 16777216 hashes = 83886080 hashes
wall 305.35 ms -> 274.718 Mhash/s, 140.66 GB/s useful (loads x 4 B)
GPU 304.77 ms -> 275.247 Mhash/s, 140.93 GB/s useful (loads x 4 B)
verify warp 0 (nonces 0..31): PASS cpu 1.928 ms single, 1.064 ms avg of 20, 4094 items derived
RESULT bench backend=metal device="Apple M5 Max" seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=4 loads=128 check=PASS cpu_warps=1 fingerprint=1468be5e1a85e771 batch_log2=24 dispatches=5 hashes=83886080 seconds=0.305 wall_seconds=0.305 mhs=275.247 mhs_wall=274.718 time=gpu
=== epoch seed "igneum-genesis" ===
program: 64 instructions x 8 iterations, loads/hash = 128 (wide 0), distinct items per warp = 4096
op mix: load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1
compile: library 0.05 ms, pipeline 0.11 ms, total 0.16 ms
threadExecutionWidth = 32, maxTotalThreadsPerThreadgroup = 1024
warm-up batch: 16777216 hashes in 167.12 ms wall, 163.65 ms GPU
batch fingerprint (FNV-1a 64 of 2^24 outputs at base nonce 0): 48a2e75ca0dbbd7c
timed: 5 batches x 16777216 hashes = 83886080 hashes
wall 813.28 ms -> 103.145 Mhash/s, 52.81 GB/s useful (loads x 4 B)
GPU 812.64 ms -> 103.227 Mhash/s, 52.85 GB/s useful (loads x 4 B)
verify warp 0 (nonces 0..31): PASS cpu 2.048 ms single, 0.982 ms avg of 20, 4095 items derived
RESULT bench backend=metal device="Apple M5 Max" seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=64 loads=128 check=PASS cpu_warps=1 fingerprint=48a2e75ca0dbbd7c batch_log2=24 dispatches=5 hashes=83886080 seconds=0.813 wall_seconds=0.813 mhs=103.227 mhs_wall=103.145 time=gpu
=== epoch seed "igneum-genesis" ===
program: 64 instructions x 8 iterations, loads/hash = 128 (wide 0), distinct items per warp = 4096
op mix: load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1
compile: library 0.05 ms, pipeline 0.10 ms, total 0.15 ms
threadExecutionWidth = 32, maxTotalThreadsPerThreadgroup = 1024
warm-up batch: 16777216 hashes in 299.62 ms wall, 296.41 ms GPU
batch fingerprint (FNV-1a 64 of 2^24 outputs at base nonce 0): 3d1523681daf7c2b
timed: 5 batches x 16777216 hashes = 83886080 hashes
wall 1482.04 ms -> 56.602 Mhash/s, 28.98 GB/s useful (loads x 4 B)
GPU 1481.42 ms -> 56.625 Mhash/s, 28.99 GB/s useful (loads x 4 B)
verify warp 0 (nonces 0..31): PASS cpu 2.238 ms single, 1.193 ms avg of 20, 4095 items derived
RESULT bench backend=metal device="Apple M5 Max" seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=256 loads=128 check=PASS cpu_warps=1 fingerprint=3d1523681daf7c2b batch_log2=24 dispatches=5 hashes=83886080 seconds=1.481 wall_seconds=1.482 mhs=56.625 mhs_wall=56.602 time=gpu
=== epoch seed "igneum-genesis" ===
program: 64 instructions x 8 iterations, loads/hash = 128 (wide 0), distinct items per warp = 4096
op mix: load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1
compile: library 0.04 ms, pipeline 0.10 ms, total 0.15 ms
threadExecutionWidth = 32, maxTotalThreadsPerThreadgroup = 1024
warm-up batch: 16777216 hashes in 652.38 ms wall, 634.38 ms GPU
batch fingerprint (FNV-1a 64 of 2^24 outputs at base nonce 0): 25f96e7dce90bd4e
timed: 5 batches x 16777216 hashes = 83886080 hashes
wall 3022.41 ms -> 27.755 Mhash/s, 14.21 GB/s useful (loads x 4 B)
GPU 3021.64 ms -> 27.762 Mhash/s, 14.21 GB/s useful (loads x 4 B)
verify warp 0 (nonces 0..31): PASS cpu 2.428 ms single, 1.043 ms avg of 20, 4096 items derived
RESULT bench backend=metal device="Apple M5 Max" seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=1024 loads=128 check=PASS cpu_warps=1 fingerprint=25f96e7dce90bd4e batch_log2=24 dispatches=5 hashes=83886080 seconds=3.022 wall_seconds=3.022 mhs=27.762 mhs_wall=27.755 time=gpu
| dataset MiB | Mhash/s (GPU time) |
|---|---|
| 4 | 275.247 |
| 64 | 103.227 |
| 256 | 56.625 |
| 1024 | 27.762 |
RESULT sweep backend=metal device="Apple M5 Max" seed="igneum-genesis" mhs_4=275.247 mhs_64=103.227 mhs_256=56.625 mhs_1024=27.762 time=gpu
=== summary (Apple M5 Max, dataset 2^28 words memory-hard, batch 2^24 x 4) ===
| seed | compile ms (lib+pipe) | Mhash/s (wall) | Mhash/s (GPU) | GB/s useful (wall) | loads/hash | items/warp | CPU verify ms/warp (avg of 20) | verify |
|---|---|---|---|---|---|---|---|---|
| igneum-genesis | 0.1 | 27.663 | 27.674 | 14.16 | 128 | 4096 | 1.046 | PASS (3 warps) |
cache fill: 2.01 ms GPU, 191.8 ms one CPU core; dataset build: 21.12 ms GPU for 1024 MiB (memory-hard)
OVERALL: PASS
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench --memprobe --device 0
memprobe on Apple M5 Max (Metal, unified memory, maxBufferLength 39813 MiB), GPU time per command buffer, best of 3
| probe | MiB | threadgroup | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |
|---|---|---|---|---|---|---|---|
| chase | 4 | 32 | 256 | 256 | 0.061 | 1.066 | 240 |
| chase | 4 | 32 | 1024 | 256 | 0.057 | 4.569 | 224 |
| chase | 4 | 32 | 4096 | 256 | 0.047 | 22.134 | 185 |
| chase | 4 | 32 | 16384 | 256 | 0.100 | 41.891 | 391 |
| chase | 4 | 32 | 65536 | 256 | 0.440 | 38.130 | 1719 |
| chase | 4 | 32 | 262144 | 256 | 1.746 | 38.441 | 6819 |
| chase | 4 | 32 | 1048576 | 256 | 6.248 | 42.961 | 24408 |
| chase | 4 | 32 | 4194304 | 256 | 24.541 | 43.752 | 95865 |
| chase | 4 | 256 | 256 | 256 | 0.066 | 0.990 | 259 |
| chase | 4 | 256 | 1024 | 256 | 0.059 | 4.465 | 229 |
| chase | 4 | 256 | 4096 | 256 | 0.056 | 18.739 | 219 |
| chase | 4 | 256 | 16384 | 256 | 0.106 | 39.383 | 416 |
| chase | 4 | 256 | 65536 | 256 | 0.416 | 40.302 | 1626 |
| chase | 4 | 256 | 262144 | 256 | 1.614 | 41.590 | 6303 |
| chase | 4 | 256 | 1048576 | 256 | 6.186 | 43.393 | 24165 |
| chase | 4 | 256 | 4194304 | 256 | 25.642 | 41.875 | 100162 |
| indep x8 | 4 | 256 | 65536 | 256 | 3.360 | 39.947 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 262144 | 256 | 12.724 | 42.194 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 1048576 | 256 | 56.694 | 37.878 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 4194304 | 256 | 215.655 | 39.832 | (8 loads in flight per lane) |
| line 16 B | 4 | 256 | 16384 | 256 | 0.112 | 37.575 G reads/s | 601.2 GB/s in 16 B reads |
| line 16 B | 4 | 256 | 65536 | 256 | 0.435 | 38.583 G reads/s | 617.3 GB/s in 16 B reads |
| line 16 B | 4 | 256 | 262144 | 256 | 1.716 | 39.105 G reads/s | 625.7 GB/s in 16 B reads |
| line 16 B | 4 | 256 | 1048576 | 256 | 6.864 | 39.110 G reads/s | 625.8 GB/s in 16 B reads |
| line 16 B | 4 | 256 | 4194304 | 256 | 27.584 | 38.926 G reads/s | 622.8 GB/s in 16 B reads |
| line 64 B | 4 | 256 | 16384 | 256 | 0.104 | 40.378 G lines/s | 2584.2 GB/s in lines |
| line 64 B | 4 | 256 | 65536 | 256 | 0.439 | 38.239 G lines/s | 2447.3 GB/s in lines |
| line 64 B | 4 | 256 | 262144 | 256 | 1.624 | 41.331 G lines/s | 2645.2 GB/s in lines |
| line 64 B | 4 | 256 | 1048576 | 256 | 6.467 | 41.510 G lines/s | 2656.6 GB/s in lines |
| line 64 B | 4 | 256 | 4194304 | 256 | 25.801 | 41.616 G lines/s | 2663.4 GB/s in lines |
| stream | 4 | 256 | 262144 | 1 | 0.006 | 662.3 GB/s coalesced | (4 MiB read once) |
RESULT memprobe backend=metal device="Apple M5 Max" mib=4 chase_gloads=43.752 chase_ns=240.1 chase_lanes=4194304 chase_best_ms=24.541 indep_gloads=42.194 line16_greads=39.110 line64_glines=41.616 stream_gbps=662.3 time=gpu
| chase | 64 | 32 | 256 | 256 | 0.101 | 0.651 | 393 |
| chase | 64 | 32 | 1024 | 256 | 0.106 | 2.463 | 416 |
| chase | 64 | 32 | 4096 | 256 | 0.132 | 7.941 | 516 |
| chase | 64 | 32 | 16384 | 256 | 0.326 | 12.859 | 1274 |
| chase | 64 | 32 | 65536 | 256 | 1.330 | 12.612 | 5196 |
| chase | 64 | 32 | 262144 | 256 | 5.539 | 12.115 | 21639 |
| chase | 64 | 32 | 1048576 | 256 | 22.056 | 12.171 | 86154 |
| chase | 64 | 32 | 4194304 | 256 | 90.968 | 11.803 | 355345 |
| chase | 64 | 256 | 256 | 256 | 0.109 | 0.601 | 426 |
| chase | 64 | 256 | 1024 | 256 | 0.114 | 2.292 | 447 |
| chase | 64 | 256 | 4096 | 256 | 0.139 | 7.557 | 542 |
| chase | 64 | 256 | 16384 | 256 | 0.358 | 11.709 | 1399 |
| chase | 64 | 256 | 65536 | 256 | 1.465 | 11.452 | 5723 |
| chase | 64 | 256 | 262144 | 256 | 5.774 | 11.623 | 22554 |
| chase | 64 | 256 | 1048576 | 256 | 23.208 | 11.567 | 90656 |
| chase | 64 | 256 | 4194304 | 256 | 96.464 | 11.131 | 376812 |
| indep x8 | 64 | 256 | 65536 | 256 | 11.813 | 11.361 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 262144 | 256 | 47.058 | 11.409 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 1048576 | 256 | 186.550 | 11.512 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 4194304 | 256 | 757.335 | 11.342 | (8 loads in flight per lane) |
| line 16 B | 64 | 256 | 16384 | 256 | 0.326 | 12.856 G reads/s | 205.7 GB/s in 16 B reads |
| line 16 B | 64 | 256 | 65536 | 256 | 1.667 | 10.066 G reads/s | 161.1 GB/s in 16 B reads |
| line 16 B | 64 | 256 | 262144 | 256 | 6.210 | 10.807 G reads/s | 172.9 GB/s in 16 B reads |
| line 16 B | 64 | 256 | 1048576 | 256 | 23.438 | 11.453 G reads/s | 183.3 GB/s in 16 B reads |
| line 16 B | 64 | 256 | 4194304 | 256 | 95.622 | 11.229 G reads/s | 179.7 GB/s in 16 B reads |
| line 64 B | 64 | 256 | 16384 | 256 | 0.273 | 15.350 G lines/s | 982.4 GB/s in lines |
| line 64 B | 64 | 256 | 65536 | 256 | 1.351 | 12.418 G lines/s | 794.8 GB/s in lines |
| line 64 B | 64 | 256 | 262144 | 256 | 5.348 | 12.549 G lines/s | 803.1 GB/s in lines |
| line 64 B | 64 | 256 | 1048576 | 256 | 22.405 | 11.981 G lines/s | 766.8 GB/s in lines |
| line 64 B | 64 | 256 | 4194304 | 256 | 91.801 | 11.696 G lines/s | 748.6 GB/s in lines |
| stream | 64 | 256 | 1048576 | 4 | 0.096 | 697.5 GB/s coalesced | (64 MiB read once) |
RESULT memprobe backend=metal device="Apple M5 Max" mib=64 chase_gloads=12.859 chase_ns=393.1 chase_lanes=16384 chase_best_ms=0.326 indep_gloads=11.512 line16_greads=12.856 line64_glines=15.350 stream_gbps=697.5 time=gpu
| chase | 256 | 32 | 256 | 256 | 0.122 | 0.538 | 476 |
| chase | 256 | 32 | 1024 | 256 | 0.127 | 2.063 | 496 |
| chase | 256 | 32 | 4096 | 256 | 0.198 | 5.296 | 773 |
| chase | 256 | 32 | 16384 | 256 | 0.758 | 5.534 | 2961 |
| chase | 256 | 32 | 65536 | 256 | 2.807 | 5.977 | 10965 |
| chase | 256 | 32 | 262144 | 256 | 11.155 | 6.016 | 43574 |
| chase | 256 | 32 | 1048576 | 256 | 40.901 | 6.563 | 159769 |
| chase | 256 | 32 | 4194304 | 256 | 154.132 | 6.966 | 602080 |
| chase | 256 | 256 | 256 | 256 | 0.105 | 0.626 | 409 |
| chase | 256 | 256 | 1024 | 256 | 0.113 | 2.320 | 441 |
| chase | 256 | 256 | 4096 | 256 | 0.183 | 5.726 | 715 |
| chase | 256 | 256 | 16384 | 256 | 0.578 | 7.262 | 2256 |
| chase | 256 | 256 | 65536 | 256 | 2.343 | 7.160 | 9153 |
| chase | 256 | 256 | 262144 | 256 | 9.698 | 6.920 | 37885 |
| chase | 256 | 256 | 1048576 | 256 | 38.777 | 6.922 | 151474 |
| chase | 256 | 256 | 4194304 | 256 | 152.711 | 7.031 | 596529 |
| indep x8 | 256 | 256 | 65536 | 256 | 19.066 | 7.040 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 262144 | 256 | 74.873 | 7.170 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 1048576 | 256 | 305.193 | 7.036 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 4194304 | 256 | 1217.784 | 7.054 | (8 loads in flight per lane) |
| line 16 B | 256 | 256 | 16384 | 256 | 0.592 | 7.083 G reads/s | 113.3 GB/s in 16 B reads |
| line 16 B | 256 | 256 | 65536 | 256 | 2.418 | 6.938 G reads/s | 111.0 GB/s in 16 B reads |
| line 16 B | 256 | 256 | 262144 | 256 | 9.717 | 6.907 G reads/s | 110.5 GB/s in 16 B reads |
| line 16 B | 256 | 256 | 1048576 | 256 | 38.751 | 6.927 G reads/s | 110.8 GB/s in 16 B reads |
| line 16 B | 256 | 256 | 4194304 | 256 | 155.833 | 6.890 G reads/s | 110.2 GB/s in 16 B reads |
| line 64 B | 256 | 256 | 16384 | 256 | 0.558 | 7.520 G lines/s | 481.3 GB/s in lines |
| line 64 B | 256 | 256 | 65536 | 256 | 2.340 | 7.170 G lines/s | 458.9 GB/s in lines |
| line 64 B | 256 | 256 | 262144 | 256 | 9.633 | 6.967 G lines/s | 445.9 GB/s in lines |
| line 64 B | 256 | 256 | 1048576 | 256 | 39.033 | 6.877 G lines/s | 440.1 GB/s in lines |
| line 64 B | 256 | 256 | 4194304 | 256 | 157.705 | 6.809 G lines/s | 435.7 GB/s in lines |
| stream | 256 | 256 | 1048576 | 16 | 0.472 | 568.9 GB/s coalesced | (256 MiB read once) |
RESULT memprobe backend=metal device="Apple M5 Max" mib=256 chase_gloads=7.262 chase_ns=476.2 chase_lanes=16384 chase_best_ms=0.578 indep_gloads=7.170 line16_greads=7.083 line64_glines=7.520 stream_gbps=568.9 time=gpu
| chase | 1024 | 32 | 256 | 256 | 0.133 | 0.491 | 521 |
| chase | 1024 | 32 | 1024 | 256 | 0.142 | 1.852 | 553 |
| chase | 1024 | 32 | 4096 | 256 | 0.353 | 2.969 | 1380 |
| chase | 1024 | 32 | 16384 | 256 | 1.395 | 3.007 | 5449 |
| chase | 1024 | 32 | 65536 | 256 | 5.640 | 2.975 | 22031 |
| chase | 1024 | 32 | 262144 | 256 | 22.671 | 2.960 | 88558 |
| chase | 1024 | 32 | 1048576 | 256 | 91.088 | 2.947 | 355814 |
| chase | 1024 | 32 | 4194304 | 256 | 340.927 | 3.149 | 1331745 |
| chase | 1024 | 256 | 256 | 256 | 0.125 | 0.524 | 488 |
| chase | 1024 | 256 | 1024 | 256 | 0.136 | 1.922 | 533 |
| chase | 1024 | 256 | 4096 | 256 | 0.329 | 3.185 | 1286 |
| chase | 1024 | 256 | 16384 | 256 | 1.240 | 3.383 | 4843 |
| chase | 1024 | 256 | 65536 | 256 | 5.108 | 3.285 | 19952 |
| chase | 1024 | 256 | 262144 | 256 | 20.659 | 3.248 | 80699 |
| chase | 1024 | 256 | 1048576 | 256 | 81.485 | 3.294 | 318301 |
| chase | 1024 | 256 | 4194304 | 256 | 315.703 | 3.401 | 1233215 |
| indep x8 | 1024 | 256 | 65536 | 256 | 39.237 | 3.421 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 262144 | 256 | 154.794 | 3.468 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 1048576 | 256 | 612.776 | 3.505 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 4194304 | 256 | 2461.838 | 3.489 | (8 loads in flight per lane) |
| line 16 B | 1024 | 256 | 16384 | 256 | 1.214 | 3.456 G reads/s | 55.3 GB/s in 16 B reads |
| line 16 B | 1024 | 256 | 65536 | 256 | 4.886 | 3.433 G reads/s | 54.9 GB/s in 16 B reads |
| line 16 B | 1024 | 256 | 262144 | 256 | 19.892 | 3.374 G reads/s | 54.0 GB/s in 16 B reads |
| line 16 B | 1024 | 256 | 1048576 | 256 | 79.130 | 3.392 G reads/s | 54.3 GB/s in 16 B reads |
| line 16 B | 1024 | 256 | 4194304 | 256 | 318.974 | 3.366 G reads/s | 53.9 GB/s in 16 B reads |
| line 64 B | 1024 | 256 | 16384 | 256 | 1.206 | 3.477 G lines/s | 222.5 GB/s in lines |
| line 64 B | 1024 | 256 | 65536 | 256 | 4.840 | 3.466 G lines/s | 221.8 GB/s in lines |
| line 64 B | 1024 | 256 | 262144 | 256 | 19.807 | 3.388 G lines/s | 216.8 GB/s in lines |
| line 64 B | 1024 | 256 | 1048576 | 256 | 79.537 | 3.375 G lines/s | 216.0 GB/s in lines |
| line 64 B | 1024 | 256 | 4194304 | 256 | 317.574 | 3.381 G lines/s | 216.4 GB/s in lines |
| stream | 1024 | 256 | 1048576 | 64 | 1.965 | 546.3 GB/s coalesced | (1024 MiB read once) |
RESULT memprobe backend=metal device="Apple M5 Max" mib=1024 chase_gloads=3.401 chase_ns=521.5 chase_lanes=4194304 chase_best_ms=315.703 indep_gloads=3.505 line16_greads=3.456 line64_glines=3.477 stream_gbps=546.3 time=gpu
| alu | 0 | 256 | 1048576 | 4096 | 5.038 | 4262.9 G int ops/s | (approximate: 5 ops per step counted) |
RESULT memprobe backend=metal device="Apple M5 Max" mib=0 alu_gops=4262.9 time=gpu
memprobe: done
=== card opencl:0 Apple M5 Max ===
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --device 0 --seconds 30 --batch-log2 24
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (1):
*[0] Apple M5 Max | Apple (OpenCL 1.2 (Jul 31 2026 20:36:30))
GPU, vendor Apple, driver 1.2 1.0, OpenCL C 1.2 , 40 compute units, 1000 MHz
global 53084 MiB, max alloc 9953 MiB, local 32 KiB, max work-group 256, sub-group extension: none
DEVICE index=0 backend=opencl type=GPU name="Apple M5 Max" vendor="Apple" platform="Apple" driver="1.2 1.0" compute_units=40 memory_mib=53084
using device [0] Apple M5 Max
timing: host wall time (Apple's OpenCL event timestamps are not usable; the rate is still a device rate, see README.md)
kernel source: /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh on Apple M5 Max: cache 12 dataset 49 check 298 ms (359 ms in all); self-test PASS (cache head, last line and FNV-1a 64 48c4f5bf24166b2e; dataset head, word [268435455] and 64 samples; 96 of 96 vector lanes)
kernel: igneum_hash_bound max work-group 32, preferred multiple 32, local memory 256 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size could not be queried (neither clGetKernelSubGroupInfoKHR nor clGetKernelSubGroupInfo is available))
warm-up dispatch (base 0): 629.71 ms, fingerprint 25f96e7dce90bd4e; 49 timed dispatches of 16777216 nonces in 30.6 s: mean 624.26 ms
rate: 26.875 Mhash/s, 13.76 GB/s useful (loads x 4 B), 3.440 G random loads/s
RESULT bench backend=opencl device="Apple M5 Max" platform="Apple" pack=/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=1024 loads=128 check=PASS fingerprint=25f96e7dce90bd4e batch_log2=24 dispatches=49 hashes=822083584 seconds=30.589 mhs=26.875 group_warps=1 exchange=0 time=wall
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 4
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (1):
*[0] Apple M5 Max | Apple (OpenCL 1.2 (Jul 31 2026 20:36:30))
GPU, vendor Apple, driver 1.2 1.0, OpenCL C 1.2 , 40 compute units, 1000 MHz
global 53084 MiB, max alloc 9953 MiB, local 32 KiB, max work-group 256, sub-group extension: none
DEVICE index=0 backend=opencl type=GPU name="Apple M5 Max" vendor="Apple" platform="Apple" driver="1.2 1.0" compute_units=40 memory_mib=53084
using device [0] Apple M5 Max
timing: host wall time (Apple's OpenCL event timestamps are not usable; the rate is still a device rate, see README.md)
kernel source: /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh on Apple M5 Max: cache 11 dataset 1 check 0 ms (12 ms in all); self-test skipped (dataset 4 MiB is not the pack's 1024 MiB: the vectors are for the pack size)
kernel: igneum_hash_bound max work-group 32, preferred multiple 32, local memory 256 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size could not be queried (neither clGetKernelSubGroupInfoKHR nor clGetKernelSubGroupInfo is available))
warm-up dispatch (base 0): 51.54 ms, fingerprint 1468be5e1a85e771; 5 timed dispatches of 16777216 nonces in 0.3 s: mean 50.06 ms
rate: 335.165 Mhash/s, 171.60 GB/s useful (loads x 4 B), 42.901 G random loads/s
RESULT bench backend=opencl device="Apple M5 Max" platform="Apple" pack=/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=4 loads=128 check=skipped fingerprint=1468be5e1a85e771 batch_log2=24 dispatches=5 hashes=83886080 seconds=0.250 mhs=335.165 group_warps=1 exchange=0 time=wall
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 64
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (1):
*[0] Apple M5 Max | Apple (OpenCL 1.2 (Jul 31 2026 20:36:30))
GPU, vendor Apple, driver 1.2 1.0, OpenCL C 1.2 , 40 compute units, 1000 MHz
global 53084 MiB, max alloc 9953 MiB, local 32 KiB, max work-group 256, sub-group extension: none
DEVICE index=0 backend=opencl type=GPU name="Apple M5 Max" vendor="Apple" platform="Apple" driver="1.2 1.0" compute_units=40 memory_mib=53084
using device [0] Apple M5 Max
timing: host wall time (Apple's OpenCL event timestamps are not usable; the rate is still a device rate, see README.md)
kernel source: /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh on Apple M5 Max: cache 10 dataset 3 check 0 ms (13 ms in all); self-test skipped (dataset 64 MiB is not the pack's 1024 MiB: the vectors are for the pack size)
kernel: igneum_hash_bound max work-group 32, preferred multiple 32, local memory 256 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size could not be queried (neither clGetKernelSubGroupInfoKHR nor clGetKernelSubGroupInfo is available))
warm-up dispatch (base 0): 161.48 ms, fingerprint 48a2e75ca0dbbd7c; 5 timed dispatches of 16777216 nonces in 0.8 s: mean 160.31 ms
rate: 104.658 Mhash/s, 53.58 GB/s useful (loads x 4 B), 13.396 G random loads/s
RESULT bench backend=opencl device="Apple M5 Max" platform="Apple" pack=/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=64 loads=128 check=skipped fingerprint=48a2e75ca0dbbd7c batch_log2=24 dispatches=5 hashes=83886080 seconds=0.802 mhs=104.658 group_warps=1 exchange=0 time=wall
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --bench --pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 256
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (1):
*[0] Apple M5 Max | Apple (OpenCL 1.2 (Jul 31 2026 20:36:30))
GPU, vendor Apple, driver 1.2 1.0, OpenCL C 1.2 , 40 compute units, 1000 MHz
global 53084 MiB, max alloc 9953 MiB, local 32 KiB, max work-group 256, sub-group extension: none
DEVICE index=0 backend=opencl type=GPU name="Apple M5 Max" vendor="Apple" platform="Apple" driver="1.2 1.0" compute_units=40 memory_mib=53084
using device [0] Apple M5 Max
timing: host wall time (Apple's OpenCL event timestamps are not usable; the rate is still a device rate, see README.md)
kernel source: /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh on Apple M5 Max: cache 12 dataset 14 check 0 ms (26 ms in all); self-test skipped (dataset 256 MiB is not the pack's 1024 MiB: the vectors are for the pack size)
kernel: igneum_hash_bound max work-group 32, preferred multiple 32, local memory 256 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size could not be queried (neither clGetKernelSubGroupInfoKHR nor clGetKernelSubGroupInfo is available))
warm-up dispatch (base 0): 319.60 ms, fingerprint 3d1523681daf7c2b; 5 timed dispatches of 16777216 nonces in 1.8 s: mean 355.38 ms
rate: 47.209 Mhash/s, 24.17 GB/s useful (loads x 4 B), 6.043 G random loads/s
RESULT bench backend=opencl device="Apple M5 Max" platform="Apple" pack=/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/packs/igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=256 loads=128 check=skipped fingerprint=3d1523681daf7c2b batch_log2=24 dispatches=5 hashes=83886080 seconds=1.777 mhs=47.209 group_warps=1 exchange=0 time=wall
+ /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/bin/macos-arm64/igneum-bench-cl --memprobe --device 0
igneum-bench-cl pack "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000" (test harness: no pool, no network, no wallet)
OpenCL devices (1):
*[0] Apple M5 Max | Apple (OpenCL 1.2 (Jul 31 2026 20:36:30))
GPU, vendor Apple, driver 1.2 1.0, OpenCL C 1.2 , 40 compute units, 1000 MHz
global 53084 MiB, max alloc 9953 MiB, local 32 KiB, max work-group 256, sub-group extension: none
DEVICE index=0 backend=opencl type=GPU name="Apple M5 Max" vendor="Apple" platform="Apple" driver="1.2 1.0" compute_units=40 memory_mib=53084
using device [0] Apple M5 Max
timing: host wall time (Apple's OpenCL event timestamps are not usable; the rate is still a device rate, see README.md)
memprobe on [Apple] Apple M5 Max, driver 1.2 1.0, 40 compute units, 1000 MHz, wall time
memprobe kernel: probe_chase max work-group 256, preferred multiple 32, local memory 0 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size could not be queried (neither clGetKernelSubGroupInfoKHR nor clGetKernelSubGroupInfo is available))
memprobe kernel: probe_chase max work-group 256, preferred multiple 32, local memory 0 bytes, private memory 0 bytes, work-group 256, sub-group size 0 (the sub-group size could not be queried (neither clGetKernelSubGroupInfoKHR nor clGetKernelSubGroupInfo is available))
| probe | MiB | work-group | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |
|---|---|---|---|---|---|---|---|
| chase | 4 | 32 | 256 | 256 | 0.222 | 0.295 | 867 |
| chase | 4 | 32 | 1024 | 256 | 0.214 | 1.225 | 836 |
| chase | 4 | 32 | 4096 | 256 | 0.204 | 5.140 | 797 |
| chase | 4 | 32 | 16384 | 256 | 0.257 | 16.320 | 1004 |
| chase | 4 | 32 | 65536 | 256 | 0.636 | 26.379 | 2484 |
| chase | 4 | 32 | 262144 | 256 | 2.150 | 31.213 | 8398 |
| chase | 4 | 32 | 1048576 | 256 | 7.835 | 34.261 | 30605 |
| chase | 4 | 32 | 4194304 | 256 | 30.687 | 34.990 | 119871 |
| chase | 4 | 256 | 256 | 256 | 0.264 | 0.248 | 1031 |
| chase | 4 | 256 | 1024 | 256 | 0.239 | 1.097 | 934 |
| chase | 4 | 256 | 4096 | 256 | 0.228 | 4.599 | 891 |
| chase | 4 | 256 | 16384 | 256 | 0.288 | 14.564 | 1125 |
| chase | 4 | 256 | 65536 | 256 | 0.648 | 25.891 | 2531 |
| chase | 4 | 256 | 262144 | 256 | 2.144 | 31.301 | 8375 |
| chase | 4 | 256 | 1048576 | 256 | 7.868 | 34.117 | 30734 |
| chase | 4 | 256 | 4194304 | 256 | 31.226 | 34.386 | 121977 |
| indep x8 | 4 | 256 | 65536 | 256 | 4.378 | 30.657 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 262144 | 256 | 15.971 | 33.615 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 1048576 | 256 | 62.304 | 34.468 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 4194304 | 256 | 238.541 | 36.010 | (8 loads in flight per lane) |
| line 64 B | 4 | 256 | 16384 | 256 | 0.282 | 14.873 G lines/s | 951.9 GB/s in lines |
| line 64 B | 4 | 256 | 65536 | 256 | 0.608 | 27.594 G lines/s | 1766.0 GB/s in lines |
| line 64 B | 4 | 256 | 262144 | 256 | 1.788 | 37.533 G lines/s | 2402.1 GB/s in lines |
| line 64 B | 4 | 256 | 1048576 | 256 | 6.954 | 38.602 G lines/s | 2470.5 GB/s in lines |
| line 64 B | 4 | 256 | 4194304 | 256 | 26.843 | 40.001 G lines/s | 2560.1 GB/s in lines |
| stream | 4 | 256 | 262144 | 1 | 0.147 | 28.5 GB/s coalesced | (4 MiB read once) |
RESULT memprobe backend=opencl device="Apple M5 Max" mib=4 chase_gloads=34.990 chase_ns=867.2 chase_lanes=4194304 chase_best_ms=30.687 indep_gloads=36.010 line64_glines=40.001 stream_gbps=28.5 time=wall
| chase | 64 | 32 | 256 | 256 | 0.257 | 0.255 | 1004 |
| chase | 64 | 32 | 1024 | 256 | 0.287 | 0.913 | 1121 |
| chase | 64 | 32 | 4096 | 256 | 0.345 | 3.039 | 1348 |
| chase | 64 | 32 | 16384 | 256 | 0.489 | 8.577 | 1910 |
| chase | 64 | 32 | 65536 | 256 | 1.586 | 10.578 | 6195 |
| chase | 64 | 32 | 262144 | 256 | 5.671 | 11.834 | 22152 |
| chase | 64 | 32 | 1048576 | 256 | 22.290 | 12.043 | 87070 |
| chase | 64 | 32 | 4194304 | 256 | 85.947 | 12.493 | 335730 |
| chase | 64 | 256 | 256 | 256 | 0.245 | 0.267 | 957 |
| chase | 64 | 256 | 1024 | 256 | 0.248 | 1.057 | 969 |
| chase | 64 | 256 | 4096 | 256 | 0.277 | 3.785 | 1082 |
| chase | 64 | 256 | 16384 | 256 | 0.531 | 7.899 | 2074 |
| chase | 64 | 256 | 65536 | 256 | 1.459 | 11.499 | 5699 |
| chase | 64 | 256 | 262144 | 256 | 5.382 | 12.469 | 21023 |
| chase | 64 | 256 | 1048576 | 256 | 21.298 | 12.604 | 83195 |
| chase | 64 | 256 | 4194304 | 256 | 83.607 | 12.843 | 326590 |
| indep x8 | 64 | 256 | 65536 | 256 | 10.550 | 12.722 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 262144 | 256 | 41.661 | 12.887 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 1048576 | 256 | 166.502 | 12.898 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 4194304 | 256 | 721.529 | 11.905 | (8 loads in flight per lane) |
| line 64 B | 64 | 256 | 16384 | 256 | 0.649 | 6.463 G lines/s | 413.6 GB/s in lines |
| line 64 B | 64 | 256 | 65536 | 256 | 1.906 | 8.802 G lines/s | 563.3 GB/s in lines |
| line 64 B | 64 | 256 | 262144 | 256 | 7.386 | 9.086 G lines/s | 581.5 GB/s in lines |
| line 64 B | 64 | 256 | 1048576 | 256 | 28.953 | 9.271 G lines/s | 593.4 GB/s in lines |
| line 64 B | 64 | 256 | 4194304 | 256 | 114.431 | 9.383 G lines/s | 600.5 GB/s in lines |
| stream | 64 | 256 | 1048576 | 4 | 0.392 | 171.2 GB/s coalesced | (64 MiB read once) |
RESULT memprobe backend=opencl device="Apple M5 Max" mib=64 chase_gloads=12.843 chase_ns=1003.9 chase_lanes=4194304 chase_best_ms=83.607 indep_gloads=12.898 line64_glines=9.383 stream_gbps=171.2 time=wall
| chase | 256 | 32 | 256 | 256 | 0.349 | 0.188 | 1363 |
| chase | 256 | 32 | 1024 | 256 | 0.390 | 0.672 | 1523 |
| chase | 256 | 32 | 4096 | 256 | 0.525 | 1.997 | 2051 |
| chase | 256 | 32 | 16384 | 256 | 1.309 | 3.204 | 5113 |
| chase | 256 | 32 | 65536 | 256 | 4.766 | 3.520 | 18617 |
| chase | 256 | 32 | 262144 | 256 | 18.515 | 3.625 | 72324 |
| chase | 256 | 32 | 1048576 | 256 | 61.710 | 4.350 | 241055 |
| chase | 256 | 32 | 4194304 | 256 | 185.356 | 5.793 | 724047 |
| chase | 256 | 256 | 256 | 256 | 0.248 | 0.264 | 969 |
| chase | 256 | 256 | 1024 | 256 | 0.269 | 0.975 | 1051 |
| chase | 256 | 256 | 4096 | 256 | 0.386 | 2.717 | 1508 |
| chase | 256 | 256 | 16384 | 256 | 0.867 | 4.838 | 3387 |
| chase | 256 | 256 | 65536 | 256 | 2.910 | 5.765 | 11367 |
| chase | 256 | 256 | 262144 | 256 | 11.224 | 5.979 | 43844 |
| chase | 256 | 256 | 1048576 | 256 | 42.996 | 6.243 | 167953 |
| chase | 256 | 256 | 4194304 | 256 | 159.241 | 6.743 | 622035 |
| indep x8 | 256 | 256 | 65536 | 256 | 19.899 | 6.745 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 262144 | 256 | 77.610 | 6.918 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 1048576 | 256 | 310.603 | 6.914 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 4194304 | 256 | 1380.417 | 6.223 | (8 loads in flight per lane) |
| line 64 B | 256 | 256 | 16384 | 256 | 0.739 | 5.676 G lines/s | 363.2 GB/s in lines |
| line 64 B | 256 | 256 | 65536 | 256 | 2.622 | 6.399 G lines/s | 409.5 GB/s in lines |
| line 64 B | 256 | 256 | 262144 | 256 | 10.436 | 6.431 G lines/s | 411.6 GB/s in lines |
| line 64 B | 256 | 256 | 1048576 | 256 | 40.795 | 6.580 G lines/s | 421.1 GB/s in lines |
| line 64 B | 256 | 256 | 4194304 | 256 | 164.633 | 6.522 G lines/s | 417.4 GB/s in lines |
| stream | 256 | 256 | 1048576 | 16 | 0.661 | 406.1 GB/s coalesced | (256 MiB read once) |
RESULT memprobe backend=opencl device="Apple M5 Max" mib=256 chase_gloads=6.743 chase_ns=1363.3 chase_lanes=4194304 chase_best_ms=159.241 indep_gloads=6.918 line64_glines=6.580 stream_gbps=406.1 time=wall
| chase | 1024 | 32 | 256 | 256 | 0.326 | 0.201 | 1273 |
| chase | 1024 | 32 | 1024 | 256 | 0.316 | 0.830 | 1234 |
| chase | 1024 | 32 | 4096 | 256 | 0.519 | 2.020 | 2027 |
| chase | 1024 | 32 | 16384 | 256 | 1.608 | 2.608 | 6281 |
| chase | 1024 | 32 | 65536 | 256 | 5.920 | 2.834 | 23125 |
| chase | 1024 | 32 | 262144 | 256 | 23.237 | 2.888 | 90770 |
| chase | 1024 | 32 | 1048576 | 256 | 91.116 | 2.946 | 355922 |
| chase | 1024 | 32 | 4194304 | 256 | 342.235 | 3.137 | 1336855 |
| chase | 1024 | 256 | 256 | 256 | 0.349 | 0.188 | 1363 |
| chase | 1024 | 256 | 1024 | 256 | 0.359 | 0.730 | 1402 |
| chase | 1024 | 256 | 4096 | 256 | 0.501 | 2.093 | 1957 |
| chase | 1024 | 256 | 16384 | 256 | 1.478 | 2.838 | 5773 |
| chase | 1024 | 256 | 65536 | 256 | 5.483 | 3.060 | 21418 |
| chase | 1024 | 256 | 262144 | 256 | 21.568 | 3.112 | 84250 |
| chase | 1024 | 256 | 1048576 | 256 | 82.553 | 3.252 | 322473 |
| chase | 1024 | 256 | 4194304 | 256 | 315.402 | 3.404 | 1232039 |
| indep x8 | 1024 | 256 | 65536 | 256 | 39.256 | 3.419 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 262144 | 256 | 157.467 | 3.409 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 1048576 | 256 | 622.614 | 3.449 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 4194304 | 256 | 2516.818 | 3.413 | (8 loads in flight per lane) |
| line 64 B | 1024 | 256 | 16384 | 256 | 1.424 | 2.945 G lines/s | 188.5 GB/s in lines |
| line 64 B | 1024 | 256 | 65536 | 256 | 5.082 | 3.301 G lines/s | 211.3 GB/s in lines |
| line 64 B | 1024 | 256 | 262144 | 256 | 20.003 | 3.355 G lines/s | 214.7 GB/s in lines |
| line 64 B | 1024 | 256 | 1048576 | 256 | 79.412 | 3.380 G lines/s | 216.3 GB/s in lines |
| line 64 B | 1024 | 256 | 4194304 | 256 | 318.413 | 3.372 G lines/s | 215.8 GB/s in lines |
| stream | 1024 | 256 | 1048576 | 64 | 2.085 | 515.0 GB/s coalesced | (1024 MiB read once) |
RESULT memprobe backend=opencl device="Apple M5 Max" mib=1024 chase_gloads=3.404 chase_ns=1273.4 chase_lanes=4194304 chase_best_ms=315.402 indep_gloads=3.449 line64_glines=3.380 stream_gbps=515.0 time=wall
| alu | 0 | 256 | 1048576 | 4096 | 5.237 | 4100.6 G int ops/s | 20.503 G steps/s per compute unit (approximate: 5 ops per step counted) |
RESULT memprobe backend=opencl device="Apple M5 Max" mib=0 alu_gops=4100.6 time=wall
memprobe: done
=== 6. proving (optional) ===
proving: {"status":"skipped: the GPU prover needs a 12 GB NVIDIA card (SP1 cuda); this is macOS"}
wrote /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/results/igneum-repro-macos-20261005T215136Z.json
wrote /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/results/igneum-repro-macos-20261005T215136Z.md
log /private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/mac-pkg/igneum-repro-v0.1.0-repro/results/igneum-repro-macos-20261005T215136Z.log
done: 2 card run(s), CPU vectors PASS

View file

@ -0,0 +1,24 @@
# Igneum reproducible benchmark, macos, 20261005T215136Z
Package v0.1.0-repro (39141f5), run 2fdd7365db4a5deb, 120 s per card. Signed by nothing; the JSON beside this file is the record.
| Machine | |
|---|---|
| OS | macOS 26.6.2 (arm64) |
| CPU | Apple M5 Max |
| Memory | 65536 MiB |
| GPU (metal 0) | Apple M5 Max |
| GPU (opencl 0) | Apple M5 Max |
| CPU vectors | verify ms per warp |
|---|---|
| PASS (96 of 96 lanes) | 0.611 |
| Card | Vectors | Fingerprint (2^24 at base 0) | MH/s at 1 GiB | 4 / 64 / 256 / 1024 MiB | Random reads at 1 GiB (G/s, ns) | Stream GB/s | In-cache / 1 GiB | Hash share of read ceiling |
|---|---|---|---|---|---|---|---|---|
| metal:0 Apple M5 Max | PASS | 25f96e7dce90bd4e | 27.674 | 275.247 / 103.227 / 56.625 / 27.762 | 3.401, 521.5 | 546.3 | 9.915 | 1.042 |
| opencl:0 Apple M5 Max (cross-check) | PASS | 25f96e7dce90bd4e | 26.875 | 335.165 / 104.658 / 47.209 / 26.875 | 3.404, 1273.4 | 515.0 | 12.471 | 1.011 |
Proving: skipped: the GPU prover needs a 12 GB NVIDIA card (SP1 cuda); this is macOS
Every command run and the sha256 of every binary are in the JSON. Vectors: bit-exact means every lane of the pack's three published warps matched. The fingerprint is the FNV-1a 64 of all 2^24 outputs at base nonce 0: equal fingerprints on two machines mean every one of those hashes agreed.

View file

@ -0,0 +1,288 @@
IGNEUM-APP version=0.3.11 machine=1ccfe586 platform=windows node=igneumd_2.1.0
**********************
Windows PowerShell transcript start
Start time: 20261006014515
Username: <pc-hostname>\Admin
RunAs User: <pc-hostname>\Admin
Configuration Name:
Machine: <pc-hostname> (Microsoft Windows NT 10.0.26200.0)
Host Application: C:\WINDOWS\System32\WindowsPowerShell\v1.0\powershell.exe -ExecutionPolicy Bypass -File <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1 -Seconds 120 -Only cuda:0 -Out <app data>\app\jobs\run-repro-pc2-20261006\results
Process ID: 29272
PSVersion: 5.1.26100.9444
PSEdition: Desktop
PSCompatibleVersions: 1.0, 2.0, 3.0, 4.0, 5.0, 5.1.26100.9444
BuildVersion: 10.0.26100.9444
CLRVersion: 4.0.30319.42000
WSManStackVersion: 3.0
PSRemotingProtocolVersion: 2.3
SerializationVersion: 1.1.0.1
**********************
igneum repro windows, run 31d75f1d674c2a77, 20261006T004515Z, 120 s per card, results in <app data>\app\jobs\run-repro-pc2-20261006\results
=== 1. machine ===
os: Microsoft Windows 11 Pro 10.0.26200 (AMD64)
cpu: AMD Ryzen 7 9800X3D 8-Core Processor
memory: 63132 MiB
video controller: AMD Radeon(TM) Graphics, driver 32.0.21042.62
video controller: NVIDIA GeForce RTX 5090, driver 32.0.16.1047
nvidia-smi: 0, NVIDIA GeForce RTX 5090, 610.47, 32607 MiB
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-cuda.exe --list
cuda devices (1), driver 13.3:
[0] NVIDIA GeForce RTX 5090 sm_120 170 SMs 32606 MiB
DEVICE index=0 backend=cuda name="NVIDIA GeForce RTX 5090" arch=sm_120 sms=170 memory_mib=32606
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --list
igneum-bench-cl pack "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000" (test harness: no pool, no network, no wallet)
OpenCL devices (2):
[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
=== 2. vectors on the CPU ===
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-pow.exe check-pack --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --warps 50
check-pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh: seed "igneum-genesis" generator v2 attempt 0 program id bcc1248b10cc90f2 (matches program.json), dataset 2^28 words (memory-hard), 128 loads/hash; built in 153 ms
cache: FNV-1a 64 48c4f5bf24166b2e == vectors.json (48c4f5bf24166b2e)
vector warp base 0 (nonces 0..31): PASS (32 of 32 lanes)
vector warp base 4096 (nonces 4096..4127): PASS (32 of 32 lanes)
vector warp base 1000000 (nonces 1000000..1000031): PASS (32 of 32 lanes)
CPU verify: 0.602 ms per 32-lane warp, avg of 50 (cold 1.310 ms max over the vector warps; checksum 19297e99c7b9a55e)
RESULT check-pack pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed=igneum-genesis program_id=bcc1248b10cc90f2 id_match=true dataset_log2=28 mode=memory-hard cache_fnv=48c4f5bf24166b2e cache_match=true warps=3 warps_pass=3 lanes=96 lanes_pass=96 cpu_verify_ms=0.602 cpu_verify_cold_ms=1.310 verdict=PASS
=== card cuda:0 NVIDIA GeForce RTX 5090 ===
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-cuda.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 0 --seconds 120 --batch-log2 24
info igneum-worker-cuda 1.0 (4 October 2026): device 0 NVIDIA_GeForce_RTX_5090 (sm_120, 170 SMs), driver 13.3 from nvcuda.dll, NVRTC 12.8 from <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\nvrtc64_120_0.dll, target sm_120 (the device's architecture, listed by NVRTC)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on NVIDIA_GeForce_RTX_5090: nvrtc 157 cache 5 dataset 42 check 265 race 0 ms variant base; self-test PASS (cache head, last line and FNV-1a 64 48c4f5bf24166b2e; dataset head, word [268435455] and 64 samples; 96 of 96 vector lanes)
warm-up dispatch (base 0): 261.63 ms, fingerprint 25f96e7dce90bd4e; 447 timed dispatches of 16777216 nonces in 120.2 s: mean 268.81 ms
rate: 62.412 Mhash/s, 31.95 GB/s useful (loads x 4 B), 7.989 G random loads/s
RESULT bench backend=cuda device="NVIDIA_GeForce_RTX_5090" arch=sm_120 pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=1024 loads=128 check=PASS fingerprint=25f96e7dce90bd4e batch_log2=24 dispatches=447 hashes=7499415552 seconds=120.160 mhs=62.412 regs=31 blocks_per_sm=24 block_warps=1 time=wall
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-cuda.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 4
info igneum-worker-cuda 1.0 (4 October 2026): device 0 NVIDIA_GeForce_RTX_5090 (sm_120, 170 SMs), driver 13.3 from nvcuda.dll, NVRTC 12.8 from <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\nvrtc64_120_0.dll, target sm_120 (the device's architecture, listed by NVRTC)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on NVIDIA_GeForce_RTX_5090: nvrtc 150 cache 5 dataset 2 check 0 race 0 ms variant base; self-test skipped (dataset 2^20 words is not the pack's 2^28: the vectors are for the pack size)
warm-up dispatch (base 0): 44.94 ms, fingerprint 1468be5e1a85e771; 5 timed dispatches of 16777216 nonces in 0.2 s: mean 44.83 ms
rate: 374.217 Mhash/s, 191.60 GB/s useful (loads x 4 B), 47.900 G random loads/s
RESULT bench backend=cuda device="NVIDIA_GeForce_RTX_5090" arch=sm_120 pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=4 loads=128 check=skipped fingerprint=1468be5e1a85e771 batch_log2=24 dispatches=5 hashes=83886080 seconds=0.224 mhs=374.217 regs=31 blocks_per_sm=24 block_warps=1 time=wall
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-cuda.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 64
info igneum-worker-cuda 1.0 (4 October 2026): device 0 NVIDIA_GeForce_RTX_5090 (sm_120, 170 SMs), driver 13.3 from nvcuda.dll, NVRTC 12.8 from <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\nvrtc64_120_0.dll, target sm_120 (the device's architecture, listed by NVRTC)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on NVIDIA_GeForce_RTX_5090: nvrtc 157 cache 5 dataset 3 check 0 race 0 ms variant base; self-test skipped (dataset 2^24 words is not the pack's 2^28: the vectors are for the pack size)
warm-up dispatch (base 0): 47.12 ms, fingerprint 48a2e75ca0dbbd7c; 5 timed dispatches of 16777216 nonces in 0.2 s: mean 45.77 ms
rate: 366.518 Mhash/s, 187.66 GB/s useful (loads x 4 B), 46.914 G random loads/s
RESULT bench backend=cuda device="NVIDIA_GeForce_RTX_5090" arch=sm_120 pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=64 loads=128 check=skipped fingerprint=48a2e75ca0dbbd7c batch_log2=24 dispatches=5 hashes=83886080 seconds=0.229 mhs=366.518 regs=31 blocks_per_sm=24 block_warps=1 time=wall
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-cuda.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 256
info igneum-worker-cuda 1.0 (4 October 2026): device 0 NVIDIA_GeForce_RTX_5090 (sm_120, 170 SMs), driver 13.3 from nvcuda.dll, NVRTC 12.8 from <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\nvrtc64_120_0.dll, target sm_120 (the device's architecture, listed by NVRTC)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on NVIDIA_GeForce_RTX_5090: nvrtc 151 cache 5 dataset 13 check 0 race 0 ms variant base; self-test skipped (dataset 2^26 words is not the pack's 2^28: the vectors are for the pack size)
warm-up dispatch (base 0): 223.70 ms, fingerprint 3d1523681daf7c2b; 5 timed dispatches of 16777216 nonces in 1.1 s: mean 226.12 ms
rate: 74.196 Mhash/s, 37.99 GB/s useful (loads x 4 B), 9.497 G random loads/s
RESULT bench backend=cuda device="NVIDIA_GeForce_RTX_5090" arch=sm_120 pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=256 loads=128 check=skipped fingerprint=3d1523681daf7c2b batch_log2=24 dispatches=5 hashes=83886080 seconds=1.131 mhs=74.196 regs=31 blocks_per_sm=24 block_warps=1 time=wall
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-cuda.exe --memprobe --device 0
info igneum-worker-cuda 1.0 (4 October 2026): device 0 NVIDIA_GeForce_RTX_5090 (sm_120, 170 SMs), driver 13.3 from nvcuda.dll, NVRTC 12.8 from <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\nvrtc64_120_0.dll, target sm_120 (the device's architecture, listed by NVRTC)
memprobe on NVIDIA_GeForce_RTX_5090 (sm_120, 170 SMs, driver 13.3, NVRTC 12.8), wall time around cuStreamSynchronize, best of 3
| probe | MiB | block | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |
|---|---|---|---|---|---|---|---|
| chase | 4 | 32 | 256 | 256 | 1.053 | 0.062 | 4114 |
| chase | 4 | 32 | 1024 | 256 | 1.116 | 0.235 | 4359 |
| chase | 4 | 32 | 4096 | 256 | 1.422 | 0.737 | 5555 |
| chase | 4 | 32 | 16384 | 256 | 0.793 | 5.288 | 3098 |
| chase | 4 | 32 | 65536 | 256 | 1.327 | 12.644 | 5183 |
| chase | 4 | 32 | 262144 | 256 | 0.640 | 104.835 | 2501 |
| chase | 4 | 32 | 1048576 | 256 | 3.522 | 76.212 | 13759 |
| chase | 4 | 32 | 4194304 | 256 | 20.289 | 52.922 | 79254 |
| chase | 4 | 256 | 256 | 256 | 2.111 | 0.031 | 8246 |
| chase | 4 | 256 | 1024 | 256 | 0.924 | 0.284 | 3609 |
| chase | 4 | 256 | 4096 | 256 | 0.619 | 1.694 | 2418 |
| chase | 4 | 256 | 16384 | 256 | 1.813 | 2.313 | 7082 |
| chase | 4 | 256 | 65536 | 256 | 1.093 | 15.353 | 4269 |
| chase | 4 | 256 | 262144 | 256 | 0.657 | 102.109 | 2567 |
| chase | 4 | 256 | 1048576 | 256 | 3.380 | 79.422 | 13203 |
| chase | 4 | 256 | 4194304 | 256 | 20.901 | 51.373 | 81644 |
| indep x8 | 4 | 256 | 65536 | 256 | 2.619 | 51.245 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 262144 | 256 | 8.906 | 60.282 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 1048576 | 256 | 40.979 | 52.404 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 4194304 | 256 | 164.808 | 52.121 | (8 loads in flight per lane) |
| line 16 B | 4 | 256 | 16384 | 256 | 1.076 | 3.898 G reads/s | 62.4 GB/s in 16 B reads |
| line 16 B | 4 | 256 | 65536 | 256 | 0.226 | 74.211 G reads/s | 1187.4 GB/s in 16 B reads |
| line 16 B | 4 | 256 | 262144 | 256 | 0.620 | 108.305 G reads/s | 1732.9 GB/s in 16 B reads |
| line 16 B | 4 | 256 | 1048576 | 256 | 4.714 | 56.943 G reads/s | 911.1 GB/s in 16 B reads |
| line 16 B | 4 | 256 | 4194304 | 256 | 17.395 | 61.727 G reads/s | 987.6 GB/s in 16 B reads |
| line 64 B | 4 | 256 | 16384 | 256 | 0.680 | 6.169 G lines/s | 394.8 GB/s in lines |
| line 64 B | 4 | 256 | 65536 | 256 | 1.261 | 13.305 G lines/s | 851.5 GB/s in lines |
| line 64 B | 4 | 256 | 262144 | 256 | 1.632 | 41.125 G lines/s | 2632.0 GB/s in lines |
| line 64 B | 4 | 256 | 1048576 | 256 | 11.347 | 23.657 G lines/s | 1514.1 GB/s in lines |
| line 64 B | 4 | 256 | 4194304 | 256 | 43.025 | 24.956 G lines/s | 1597.2 GB/s in lines |
| stream | 4 | 256 | 262144 | 1 | 1.226 | 3.4 GB/s coalesced | (4 MiB read once) |
RESULT memprobe backend=cuda device="NVIDIA_GeForce_RTX_5090" mib=4 chase_gloads=104.835 chase_ns=4114.2 chase_lanes=262144 chase_best_ms=0.640 indep_gloads=60.282 line16_greads=108.305 line64_glines=41.125 stream_gbps=3.4 time=wall
| chase | 64 | 32 | 256 | 256 | 0.398 | 0.165 | 1554 |
| chase | 64 | 32 | 1024 | 256 | 1.776 | 0.148 | 6937 |
| chase | 64 | 32 | 4096 | 256 | 1.378 | 0.761 | 5383 |
| chase | 64 | 32 | 16384 | 256 | 0.157 | 26.677 | 614 |
| chase | 64 | 32 | 65536 | 256 | 0.210 | 79.999 | 819 |
| chase | 64 | 32 | 262144 | 256 | 0.639 | 105.036 | 2496 |
| chase | 64 | 32 | 1048576 | 256 | 4.710 | 56.993 | 18398 |
| chase | 64 | 32 | 4194304 | 256 | 22.656 | 47.393 | 88500 |
| chase | 64 | 256 | 256 | 256 | 1.304 | 0.050 | 5095 |
| chase | 64 | 256 | 1024 | 256 | 0.607 | 0.432 | 2371 |
| chase | 64 | 256 | 4096 | 256 | 0.770 | 1.362 | 3008 |
| chase | 64 | 256 | 16384 | 256 | 1.104 | 3.799 | 4313 |
| chase | 64 | 256 | 65536 | 256 | 0.208 | 80.562 | 813 |
| chase | 64 | 256 | 262144 | 256 | 0.659 | 101.844 | 2574 |
| chase | 64 | 256 | 1048576 | 256 | 3.889 | 69.026 | 15191 |
| chase | 64 | 256 | 4194304 | 256 | 24.084 | 44.583 | 94078 |
| indep x8 | 64 | 256 | 65536 | 256 | 3.127 | 42.923 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 262144 | 256 | 9.972 | 53.838 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 1048576 | 256 | 44.887 | 47.842 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 4194304 | 256 | 175.309 | 48.999 | (8 loads in flight per lane) |
| line 16 B | 64 | 256 | 16384 | 256 | 1.756 | 2.389 G reads/s | 38.2 GB/s in 16 B reads |
| line 16 B | 64 | 256 | 65536 | 256 | 0.964 | 17.402 G reads/s | 278.4 GB/s in 16 B reads |
| line 16 B | 64 | 256 | 262144 | 256 | 2.710 | 24.762 G reads/s | 396.2 GB/s in 16 B reads |
| line 16 B | 64 | 256 | 1048576 | 256 | 4.704 | 57.064 G reads/s | 913.0 GB/s in 16 B reads |
| line 16 B | 64 | 256 | 4194304 | 256 | 19.486 | 55.103 G reads/s | 881.6 GB/s in 16 B reads |
| line 64 B | 64 | 256 | 16384 | 256 | 1.315 | 3.190 G lines/s | 204.2 GB/s in lines |
| line 64 B | 64 | 256 | 65536 | 256 | 2.450 | 6.848 G lines/s | 438.3 GB/s in lines |
| line 64 B | 64 | 256 | 262144 | 256 | 3.456 | 19.419 G lines/s | 1242.8 GB/s in lines |
| line 64 B | 64 | 256 | 1048576 | 256 | 10.065 | 26.670 G lines/s | 1706.9 GB/s in lines |
| line 64 B | 64 | 256 | 4194304 | 256 | 46.140 | 23.271 G lines/s | 1489.4 GB/s in lines |
| stream | 64 | 256 | 1048576 | 4 | 0.467 | 143.7 GB/s coalesced | (64 MiB read once) |
RESULT memprobe backend=cuda device="NVIDIA_GeForce_RTX_5090" mib=64 chase_gloads=105.036 chase_ns=1554.5 chase_lanes=262144 chase_best_ms=0.639 indep_gloads=53.838 line16_greads=57.064 line64_glines=26.670 stream_gbps=143.7 time=wall
| chase | 256 | 32 | 256 | 256 | 0.983 | 0.067 | 3839 |
| chase | 256 | 32 | 1024 | 256 | 1.647 | 0.159 | 6434 |
| chase | 256 | 32 | 4096 | 256 | 0.735 | 1.427 | 2871 |
| chase | 256 | 32 | 16384 | 256 | 1.985 | 2.113 | 7754 |
| chase | 256 | 32 | 65536 | 256 | 0.829 | 20.235 | 3239 |
| chase | 256 | 32 | 262144 | 256 | 6.382 | 10.515 | 24930 |
| chase | 256 | 32 | 1048576 | 256 | 25.417 | 10.561 | 99286 |
| chase | 256 | 32 | 4194304 | 256 | 109.449 | 9.810 | 427535 |
| chase | 256 | 256 | 256 | 256 | 2.249 | 0.029 | 8785 |
| chase | 256 | 256 | 1024 | 256 | 1.289 | 0.203 | 5035 |
| chase | 256 | 256 | 4096 | 256 | 1.075 | 0.975 | 4199 |
| chase | 256 | 256 | 16384 | 256 | 2.376 | 1.765 | 9281 |
| chase | 256 | 256 | 65536 | 256 | 2.035 | 8.245 | 7949 |
| chase | 256 | 256 | 262144 | 256 | 4.478 | 14.987 | 17491 |
| chase | 256 | 256 | 1048576 | 256 | 25.814 | 10.399 | 100836 |
| chase | 256 | 256 | 4194304 | 256 | 110.207 | 9.743 | 430495 |
| indep x8 | 256 | 256 | 65536 | 256 | 13.970 | 9.608 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 262144 | 256 | 54.400 | 9.869 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 1048576 | 256 | 228.235 | 9.409 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 4194304 | 256 | 934.530 | 9.192 | (8 loads in flight per lane) |
| line 16 B | 256 | 256 | 16384 | 256 | 0.975 | 4.301 G reads/s | 68.8 GB/s in 16 B reads |
| line 16 B | 256 | 256 | 65536 | 256 | 0.703 | 23.869 G reads/s | 381.9 GB/s in 16 B reads |
| line 16 B | 256 | 256 | 262144 | 256 | 4.816 | 13.934 G reads/s | 222.9 GB/s in 16 B reads |
| line 16 B | 256 | 256 | 1048576 | 256 | 24.106 | 11.136 G reads/s | 178.2 GB/s in 16 B reads |
| line 16 B | 256 | 256 | 4194304 | 256 | 100.201 | 10.716 G reads/s | 171.5 GB/s in 16 B reads |
| line 64 B | 256 | 256 | 16384 | 256 | 0.233 | 18.008 G lines/s | 1152.5 GB/s in lines |
| line 64 B | 256 | 256 | 65536 | 256 | 0.511 | 32.833 G lines/s | 2101.3 GB/s in lines |
| line 64 B | 256 | 256 | 262144 | 256 | 8.712 | 7.703 G lines/s | 493.0 GB/s in lines |
| line 64 B | 256 | 256 | 1048576 | 256 | 36.245 | 7.406 G lines/s | 474.0 GB/s in lines |
| line 64 B | 256 | 256 | 4194304 | 256 | 176.297 | 6.091 G lines/s | 389.8 GB/s in lines |
| stream | 256 | 256 | 1048576 | 16 | 0.916 | 293.1 GB/s coalesced | (256 MiB read once) |
RESULT memprobe backend=cuda device="NVIDIA_GeForce_RTX_5090" mib=256 chase_gloads=20.235 chase_ns=3838.5 chase_lanes=65536 chase_best_ms=0.829 indep_gloads=9.869 line16_greads=23.869 line64_glines=32.833 stream_gbps=293.1 time=wall
| chase | 1024 | 32 | 256 | 256 | 0.452 | 0.145 | 1766 |
| chase | 1024 | 32 | 1024 | 256 | 1.017 | 0.258 | 3973 |
| chase | 1024 | 32 | 4096 | 256 | 0.781 | 1.343 | 3051 |
| chase | 1024 | 32 | 16384 | 256 | 0.588 | 7.134 | 2296 |
| chase | 1024 | 32 | 65536 | 256 | 1.896 | 8.848 | 7407 |
| chase | 1024 | 32 | 262144 | 256 | 8.805 | 7.622 | 34395 |
| chase | 1024 | 32 | 1048576 | 256 | 32.211 | 8.334 | 125824 |
| chase | 1024 | 32 | 4194304 | 256 | 133.279 | 8.056 | 520621 |
| chase | 1024 | 256 | 256 | 256 | 0.826 | 0.079 | 3227 |
| chase | 1024 | 256 | 1024 | 256 | 2.167 | 0.121 | 8465 |
| chase | 1024 | 256 | 4096 | 256 | 0.855 | 1.226 | 3341 |
| chase | 1024 | 256 | 16384 | 256 | 0.282 | 14.887 | 1101 |
| chase | 1024 | 256 | 65536 | 256 | 0.959 | 17.495 | 3746 |
| chase | 1024 | 256 | 262144 | 256 | 6.607 | 10.157 | 25809 |
| chase | 1024 | 256 | 1048576 | 256 | 32.176 | 8.343 | 125688 |
| chase | 1024 | 256 | 4194304 | 256 | 134.493 | 7.984 | 525364 |
| indep x8 | 1024 | 256 | 65536 | 256 | 13.180 | 10.183 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 262144 | 256 | 67.818 | 7.916 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 1048576 | 256 | 270.518 | 7.938 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 4194304 | 256 | 1100.428 | 7.806 | (8 loads in flight per lane) |
| line 16 B | 1024 | 256 | 16384 | 256 | 1.958 | 2.142 G reads/s | 34.3 GB/s in 16 B reads |
| line 16 B | 1024 | 256 | 65536 | 256 | 0.910 | 18.443 G reads/s | 295.1 GB/s in 16 B reads |
| line 16 B | 1024 | 256 | 262144 | 256 | 8.613 | 7.792 G reads/s | 124.7 GB/s in 16 B reads |
| line 16 B | 1024 | 256 | 1048576 | 256 | 31.067 | 8.640 G reads/s | 138.2 GB/s in 16 B reads |
| line 16 B | 1024 | 256 | 4194304 | 256 | 135.396 | 7.930 G reads/s | 126.9 GB/s in 16 B reads |
| line 64 B | 1024 | 256 | 16384 | 256 | 2.288 | 1.833 G lines/s | 117.3 GB/s in lines |
| line 64 B | 1024 | 256 | 65536 | 256 | 1.064 | 15.769 G lines/s | 1009.2 GB/s in lines |
| line 64 B | 1024 | 256 | 262144 | 256 | 14.335 | 4.681 G lines/s | 299.6 GB/s in lines |
| line 64 B | 1024 | 256 | 1048576 | 256 | 64.783 | 4.144 G lines/s | 265.2 GB/s in lines |
| line 64 B | 1024 | 256 | 4194304 | 256 | 249.461 | 4.304 G lines/s | 275.5 GB/s in lines |
| stream | 1024 | 256 | 1048576 | 64 | 3.165 | 339.3 GB/s coalesced | (1024 MiB read once) |
RESULT memprobe backend=cuda device="NVIDIA_GeForce_RTX_5090" mib=1024 chase_gloads=17.495 chase_ns=1766.2 chase_lanes=65536 chase_best_ms=0.959 indep_gloads=10.183 line16_greads=18.443 line64_glines=15.769 stream_gbps=339.3 time=wall
| alu | 0 | 256 | 1048576 | 4096 | 1.598 | 13439.4 G int ops/s | 15.811 G steps/s per SM (approximate: 5 ops per step counted) |
RESULT memprobe backend=cuda device="NVIDIA_GeForce_RTX_5090" mib=0 alu_gops=13439.4 time=wall
memprobe: done
skipping opencl:0 NVIDIA GeForce RTX 5090 (--only)
skipping opencl:1 gfx1036 (--only)
=== 6. proving (optional) ===
+ wsl.exe -d Ubuntu-24.04 -- bash -lc "/opt/igneum/igneum-prove-host" --mode id
RESULT id: pinned guests: shard program id 0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a (2832504 bytes, sha256 0x150f4c05a2951fc5) aggregator id 0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896 (319744 bytes), pinned 2026-10-05T16:20:38Z on Darwin MacBook-Pro.local 25.6.0 Darwin Kernel Version 25.6.0: Fri Jul 31 19:19:08 PDT 2026; root:xnu-12377.161.14~5/RELEASE_ARM64_T6050 arm64, SP1 6.8.1 circuit v6.1.0
+ wsl.exe -d Ubuntu-24.04 -- bash -lc SP1_PROVER=cuda "/opt/igneum/igneum-prove-host" '/mnt/c/%USERPROFILE%/AppData/Local/igneum/app/jobs/fetch-repro-pc2-20261006/repro/igneum-repro-v0.1.0-repro/proving/block-338-shard1.json' --mode shard --shard 0 --out '/tmp/igneum-repro-prove-20261006T004515Z.json'
igneum-prove-host sources 08a603690c518d75: fixture /mnt/c/%USERPROFILE%/AppData/Local/igneum/app/jobs/fetch-repro-pc2-20261006/repro/igneum-repro-v0.1.0-repro/proving/block-338-shard1.json: chain 4463 block 338 (0x0a7b66a57fde7b012999f54d4e2a1c50b86171485cf7d8ac23b22f5c2b79df47) at DAA score 0, 11 transactions in 1 including blocks, 32 accounts in the pre-state, plan 1 shard(s) at S_p = 7500000 pgas; fee schedule prototype, calibrated v1 from DAA score never: this block meters with prototype (intrinsic 200 pgas, B_p 30000000); SP1_PROVER=cuda; prover payout 0x1919191919191919191919191919191919191919; 2026-10-06T00:47:41Z
STAGE native start 2026-10-06T00:47:41Z
RESULT native: 0.0089 s, pre 0xd363fab9a828ebcd5b2788455a37ab778baaf82fba2a7f243e0f2d86e5a815a4 post 0xa7dc035ced7f1d1cf258ac851c66f7d6b3ab63d400c745100d4c35702c0278fe receipts 0x5908eb7bd934c28990289a3e1764c9a935133f346c0a83c372eb0cb191b6607f gas 1390773 pgas 6751568 executed 11 skipped 0: MATCHES the fixture's expected values at 2026-10-06T00:47:41Z
RESULT shard 0 native: txs 0..11 (11 executed, 0 skipped), gas 1390773, pgas 6751568, witness 19 accounts 3 slots 24 leaves 13 hashes, input 21611 bytes, pre 0xd363fab9a828ebcd5b2788455a37ab778baaf82fba2a7f243e0f2d86e5a815a4 post 0xa7dc035ced7f1d1cf258ac851c66f7d6b3ab63d400c745100d4c35702c0278fe
RESULT plan: 1 shard(s) chain from 0xd363fab9a828ebcd5b2788455a37ab778baaf82fba2a7f243e0f2d86e5a815a4 to 0xa7dc035ced7f1d1cf258ac851c66f7d6b3ab63d400c745100d4c35702c0278fe and sum to gas 1390773 pgas 6751568: the cut equals the block; native aggregation receipts 0x86aa587dff84019338e587bbeaecc0d74ca64f99d8f1b5be535b0e5d74e2fb2a provers 0xe8a6c07ee13e9afa662a1bf0bf69bae726aa8fdb541308bdd4b63bbc60175f27
RESULT tamper account balance: REJECTED (pre-root 0xa8ffbad9fece604586ce4a4efa5815a491ec79fff5a3a4460f316c0bc332e951 is not the node's 0xd363fab9a828ebcd5b2788455a37ab778baaf82fba2a7f243e0f2d86e5a815a4)
RESULT tamper storage or code: REJECTED (the statement cannot be proven)
RESULT tamper dropped account: REJECTED (the statement cannot be proven)
RESULT stub: ProofSystem v0 program 0x09a477342fcc874f0fbd45d307bfee4306efa7d12eb6c587213e0300df5fce65 verify_segment true
STAGE setup start 2026-10-06T00:47:41Z
Running sp1-gpu-server 6.8.1 with device 0
RESULT setup: 15.80 s (prover client 0.72 s, shard keys 15.05 s, aggregator keys 0.03 s), ProofSystem v1 shard program id 0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a aggregator id 0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896 at 2026-10-06T00:47:57Z
RESULT pinned: setup matches the manifest (pinned guests: shard program id 0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a (2832504 bytes, sha256 0x150f4c05a2951fc5) aggregator id 0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896 (319744 bytes), pinned 2026-10-05T16:20:38Z on Darwin MacBook-Pro.local 25.6.0 Darwin Kernel Version 25.6.0: Fri Jul 31 19:19:08 PDT 2026; root:xnu-12377.161.14~5/RELEASE_ARM64_T6050 arm64, SP1 6.8.1 circuit v6.1.0)
STAGE execute shard 0 start 2026-10-06T00:47:57Z
RESULT execute shard 0: 60415376 cycles, prover gas 50304365, 1.48 s, 43 cycles per EVM gas, 9 cycles per pgas, syscalls 8711 at 2026-10-06T00:47:58Z
STAGE core shard 0 start 2026-10-06T00:47:58Z
RESULT core shard 0: prove 20.2 s, proof 18112711 bytes, verify 0.597 s, VERIFIED; post-root 0xa7dc035ced7f1d1cf258ac851c66f7d6b3ab63d400c745100d4c35702c0278fe at 2026-10-06T00:48:19Z
saved /tmp/block-338-shard-0-core.bin in 0.0 s
STAGE compressed shard 0 start 2026-10-06T00:48:19Z
RESULT compressed shard 0: prove 33.9 s, proof 1272897 bytes, verify 0.038 s, VERIFIED; post-root 0xa7dc035ced7f1d1cf258ac851c66f7d6b3ab63d400c745100d4c35702c0278fe prover 0x1919191919191919191919191919191919191919 at 2026-10-06T00:48:53Z
saved /tmp/block-338-shard-0-compressed.bin in 0.0 s
results written to /tmp/igneum-repro-prove-20261006T004515Z.json
proving: {"fixture":"proving/block-338-shard1.json","core_prove_seconds":20.187076321,"cycles":60415376,"wall_seconds":72,"shard":0,"compressed_verify_seconds":0.037670188,"fixture_sha256":"93bbb9b49471a928482146b674a74b24f7071c6747118b545fe3b2020d7984c6","compressed_proof_bytes":1272897,"execute_seconds":1.4813646249999999,"verify":"not run","verify_seconds":null,"host_sha256":"29cc476898c3ab2666110ef65f5b432ac243d1e87e565b57b28e4d140fcbfeee","status":"run","compressed_prove_seconds":33.911097916,"pinned":"RESULT id: pinned guests: shard program id 0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a (2832504 bytes, sha256 0x150f4c05a2951fc5) aggregator id 0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896 (319744 bytes), pinned 2026-10-05T16:20:38Z on Darwin MacBook-Pro.local 25.6.0 Darwin Kernel Version 25.6.0: Fri Jul 31 19:19:08 PDT 2026; root:xnu-12377.161.14~5/RELEASE_ARM64_T6050 arm64, SP1 6.8.1 circuit v6.1.0"}
PS>TerminatingError(ConvertTo-Json): "Exception of type 'System.OutOfMemoryException' was thrown."
ConvertTo-Json : Exception of type 'System.OutOfMemoryException' was thrown.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:217
char:11
+ $Result | ConvertTo-Json -Depth 8 | Set-Content -Encoding UTF8 -Liter ...
+ ~~~~~~~~~~~~~~~~~~~~~~~
+ CategoryInfo : NotSpecified: (:) [ConvertTo-Json], OutOfMemoryException
+ FullyQualifiedErrorId : System.OutOfMemoryException,Microsoft.PowerShell.Commands.ConvertToJsonCommand
ConvertTo-Json : Exception of type 'System.OutOfMemoryException' was thrown.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:217
char:11
+ $Result | ConvertTo-Json -Depth 8 | Set-Content -Encoding UTF8 -Liter ...
+ ~~~~~~~~~~~~~~~~~~~~~~~
+ CategoryInfo : NotSpecified: (:) [ConvertTo-Json], OutOfMemoryException
+ FullyQualifiedErrorId : System.OutOfMemoryException,Microsoft.PowerShell.Commands.ConvertToJsonCommand
>> TerminatingError(Set-Content): "Cannot bind argument to parameter 'LiteralPath' because it is an empty string."
Set-Content : Cannot bind argument to parameter 'LiteralPath' because it is an empty string.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:235
char:58
+ $md -join "`n" | Set-Content -Encoding UTF8 -LiteralPath $Md
+ ~~~
+ CategoryInfo : InvalidData: (:) [Set-Content], ParameterBindingValidationException
+ FullyQualifiedErrorId :
ParameterArgumentValidationErrorEmptyStringNotAllowed,Microsoft.PowerShell.Commands.SetContentCommand
Set-Content : Cannot bind argument to parameter 'LiteralPath' because it is an empty string.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:235
char:58
+ $md -join "`n" | Set-Content -Encoding UTF8 -LiteralPath $Md
+ ~~~
+ CategoryInfo : InvalidData: (:) [Set-Content], ParameterBindingValidationException
+ FullyQualifiedErrorId : ParameterArgumentValidationErrorEmptyStringNotAllowed,Microsoft.PowerShell.Commands.SetC
ontentCommand
wrote <app data>\app\jobs\run-repro-pc2-20261006\results\igneum-repro-windows-20261006T004515Z.json
wrote # Igneum reproducible benchmark, windows, 20261006T004515Z Package v0.1.0-repro (39141f5), run 31d75f1d674c2a77, 120 s per card. Signed by nothing; the JSON beside this file is the record. | Machine | | |---|---| | OS | Microsoft Windows 11 Pro 10.0.26200 (AMD64) | | CPU | System.Collections.Hashtable | | Memory | 63132 MiB | | GPU (cuda 0) | NVIDIA GeForce RTX 5090 | | GPU (opencl 0) | NVIDIA GeForce RTX 5090 | | GPU (opencl 1) | gfx1036 | | CPU vectors | verify ms per warp | |---|---| | PASS (96 of 96 lanes) | 0.602 | | Card | Vectors | Fingerprint (2^24 at base 0) | MH/s at 1 GiB | 4 / 64 / 256 / 1024 MiB | Random reads at 1 GiB (G/s, ns) | Stream GB/s | In-cache / 1 GiB | Hash share of read ceiling | |---|---|---|---|---|---|---|---|---| | cuda:0 NVIDIA GeForce RTX 5090 | PASS | 25f96e7dce90bd4e | 62.412 | 374.217 / 366.518 / 74.196 / 62.412 | 17.495, 1766.2 | 339.3 | 5.996 | 0.457 | Proving: run, compressed proof 33.911097916 s, not run in s Every command run and the sha256 of every binary are in the JSON. Vectors: bit-exact means every lane of the pack's three published warps matched. The fingerprint is the FNV-1a 64 of all 2^24 outputs at base nonce 0: equal fingerprints on two machines mean every one of those hashes agreed.
log <app data>\app\jobs\run-repro-pc2-20261006\results\igneum-repro-windows-20261006T004515Z.log
done: 1 card run(s), CPU vectors PASS
**********************
Windows PowerShell transcript end
End time: 20261006015127
**********************

View file

@ -0,0 +1,316 @@
IGNEUM-APP version=0.3.11 machine=1ccfe586 platform=windows node=igneumd_2.1.0
**********************
Windows PowerShell transcript start
Start time: 20261006015218
Username: <pc-hostname>\Admin
RunAs User: <pc-hostname>\Admin
Configuration Name:
Machine: <pc-hostname> (Microsoft Windows NT 10.0.26200.0)
Host Application: C:\WINDOWS\System32\WindowsPowerShell\v1.0\powershell.exe -ExecutionPolicy Bypass -File <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1 -Seconds 120 -Only opencl:0 -Out <app data>\app\jobs\run-repro-pc2-20261006\results -NoProve
Process ID: 23408
PSVersion: 5.1.26100.9444
PSEdition: Desktop
PSCompatibleVersions: 1.0, 2.0, 3.0, 4.0, 5.0, 5.1.26100.9444
BuildVersion: 10.0.26100.9444
CLRVersion: 4.0.30319.42000
WSManStackVersion: 3.0
PSRemotingProtocolVersion: 2.3
SerializationVersion: 1.1.0.1
**********************
igneum repro windows, run 208b02bccab33633, 20261006T005218Z, 120 s per card, results in <app data>\app\jobs\run-repro-pc2-20261006\results
=== 1. machine ===
os: Microsoft Windows 11 Pro 10.0.26200 (AMD64)
cpu: AMD Ryzen 7 9800X3D 8-Core Processor
memory: 63132 MiB
video controller: AMD Radeon(TM) Graphics, driver 32.0.21042.62
video controller: NVIDIA GeForce RTX 5090, driver 32.0.16.1047
nvidia-smi: 0, NVIDIA GeForce RTX 5090, 610.47, 32607 MiB
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-cuda.exe --list
cuda devices (1), driver 13.3:
[0] NVIDIA GeForce RTX 5090 sm_120 170 SMs 32606 MiB
DEVICE index=0 backend=cuda name="NVIDIA GeForce RTX 5090" arch=sm_120 sms=170 memory_mib=32606
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --list
igneum-bench-cl pack "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000" (test harness: no pool, no network, no wallet)
OpenCL devices (2):
[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
=== 2. vectors on the CPU ===
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-pow.exe check-pack --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --warps 50
check-pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh: seed "igneum-genesis" generator v2 attempt 0 program id bcc1248b10cc90f2 (matches program.json), dataset 2^28 words (memory-hard), 128 loads/hash; built in 156 ms
cache: FNV-1a 64 48c4f5bf24166b2e == vectors.json (48c4f5bf24166b2e)
vector warp base 0 (nonces 0..31): PASS (32 of 32 lanes)
vector warp base 4096 (nonces 4096..4127): PASS (32 of 32 lanes)
vector warp base 1000000 (nonces 1000000..1000031): PASS (32 of 32 lanes)
CPU verify: 0.640 ms per 32-lane warp, avg of 50 (cold 0.731 ms max over the vector warps; checksum 19297e99c7b9a55e)
RESULT check-pack pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed=igneum-genesis program_id=bcc1248b10cc90f2 id_match=true dataset_log2=28 mode=memory-hard cache_fnv=48c4f5bf24166b2e cache_match=true warps=3 warps_pass=3 lanes=96 lanes_pass=96 cpu_verify_ms=0.640 cpu_verify_cold_ms=0.731 verdict=PASS
skipping cuda:0 NVIDIA GeForce RTX 5090 (--only)
=== card opencl:0 NVIDIA GeForce RTX 5090 ===
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 0 --seconds 30 --batch-log2 24
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (2):
*[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [0] NVIDIA GeForce RTX 5090
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
kernel source: <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on NVIDIA GeForce RTX 5090: cache 4 dataset 38 check 239 ms (282 ms in all); self-test PASS (cache head, last line and FNV-1a 64 48c4f5bf24166b2e; dataset head, word [268435455] and 64 samples; 96 of 96 vector lanes)
kernel: igneum_hash_bound max work-group 256, preferred multiple 32, local memory 260 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size query failed: clGetKernelSubGroupInfoKHR returned CL_INVALID_OPERATION (-59))
warm-up dispatch (base 0): 274.19 ms, fingerprint 25f96e7dce90bd4e; 111 timed dispatches of 16777216 nonces in 29.9 s: mean 269.37 ms
rate: 62.283 Mhash/s, 31.89 GB/s useful (loads x 4 B), 7.972 G random loads/s
RESULT bench backend=opencl device="NVIDIA GeForce RTX 5090" platform="NVIDIA CUDA" pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=1024 loads=128 check=PASS fingerprint=25f96e7dce90bd4e batch_log2=24 dispatches=111 hashes=1862270976 seconds=29.900 mhs=62.283 group_warps=1 exchange=0 time=event
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 4
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (2):
*[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [0] NVIDIA GeForce RTX 5090
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
kernel source: <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on NVIDIA GeForce RTX 5090: cache 5 dataset 4 check 0 ms (9 ms in all); self-test skipped (dataset 4 MiB is not the pack's 1024 MiB: the vectors are for the pack size)
kernel: igneum_hash_bound max work-group 256, preferred multiple 32, local memory 260 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size query failed: clGetKernelSubGroupInfoKHR returned CL_INVALID_OPERATION (-59))
warm-up dispatch (base 0): 41.47 ms, fingerprint 1468be5e1a85e771; 5 timed dispatches of 16777216 nonces in 0.2 s: mean 44.71 ms
rate: 375.252 Mhash/s, 192.13 GB/s useful (loads x 4 B), 48.032 G random loads/s
RESULT bench backend=opencl device="NVIDIA GeForce RTX 5090" platform="NVIDIA CUDA" pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=4 loads=128 check=skipped fingerprint=1468be5e1a85e771 batch_log2=24 dispatches=5 hashes=83886080 seconds=0.224 mhs=375.252 group_warps=1 exchange=0 time=event
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 64
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (2):
*[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [0] NVIDIA GeForce RTX 5090
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
kernel source: <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on NVIDIA GeForce RTX 5090: cache 4 dataset 3 check 0 ms (7 ms in all); self-test skipped (dataset 64 MiB is not the pack's 1024 MiB: the vectors are for the pack size)
kernel: igneum_hash_bound max work-group 256, preferred multiple 32, local memory 260 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size query failed: clGetKernelSubGroupInfoKHR returned CL_INVALID_OPERATION (-59))
warm-up dispatch (base 0): 44.99 ms, fingerprint 48a2e75ca0dbbd7c; 5 timed dispatches of 16777216 nonces in 0.2 s: mean 43.50 ms
rate: 385.662 Mhash/s, 197.46 GB/s useful (loads x 4 B), 49.365 G random loads/s
RESULT bench backend=opencl device="NVIDIA GeForce RTX 5090" platform="NVIDIA CUDA" pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=64 loads=128 check=skipped fingerprint=48a2e75ca0dbbd7c batch_log2=24 dispatches=5 hashes=83886080 seconds=0.218 mhs=385.662 group_warps=1 exchange=0 time=event
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 0 --batches 5 --batch-log2 24 --dataset-mib 256
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (2):
*[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [0] NVIDIA GeForce RTX 5090
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
kernel source: <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on NVIDIA GeForce RTX 5090: cache 4 dataset 12 check 0 ms (16 ms in all); self-test skipped (dataset 256 MiB is not the pack's 1024 MiB: the vectors are for the pack size)
kernel: igneum_hash_bound max work-group 256, preferred multiple 32, local memory 260 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size query failed: clGetKernelSubGroupInfoKHR returned CL_INVALID_OPERATION (-59))
warm-up dispatch (base 0): 238.51 ms, fingerprint 3d1523681daf7c2b; 5 timed dispatches of 16777216 nonces in 1.2 s: mean 231.49 ms
rate: 72.476 Mhash/s, 37.11 GB/s useful (loads x 4 B), 9.277 G random loads/s
RESULT bench backend=opencl device="NVIDIA GeForce RTX 5090" platform="NVIDIA CUDA" pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=256 loads=128 check=skipped fingerprint=3d1523681daf7c2b batch_log2=24 dispatches=5 hashes=83886080 seconds=1.157 mhs=72.476 group_warps=1 exchange=0 time=event
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --memprobe --device 0
igneum-bench-cl pack "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000" (test harness: no pool, no network, no wallet)
OpenCL devices (2):
*[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [0] NVIDIA GeForce RTX 5090
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
memprobe on [NVIDIA CUDA] NVIDIA GeForce RTX 5090, driver 610.47, 170 compute units, 2407 MHz, device event time
memprobe kernel: probe_chase max work-group 256, preferred multiple 32, local memory 1 bytes, private memory 0 bytes, work-group 32, sub-group size 0 (the sub-group size query failed: clGetKernelSubGroupInfoKHR returned CL_INVALID_OPERATION (-59))
memprobe kernel: probe_chase max work-group 256, preferred multiple 32, local memory 1 bytes, private memory 0 bytes, work-group 256, sub-group size 0 (the sub-group size query failed: clGetKernelSubGroupInfoKHR returned CL_INVALID_OPERATION (-59))
| probe | MiB | work-group | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |
|---|---|---|---|---|---|---|---|
| chase | 4 | 32 | 256 | 256 | 0.106 | 0.621 | 412 |
| chase | 4 | 32 | 1024 | 256 | 0.072 | 3.641 | 281 |
| chase | 4 | 32 | 4096 | 256 | 0.047 | 22.156 | 185 |
| chase | 4 | 32 | 16384 | 256 | 0.048 | 87.556 | 187 |
| chase | 4 | 32 | 65536 | 256 | 0.170 | 98.587 | 665 |
| chase | 4 | 32 | 262144 | 256 | 0.602 | 111.432 | 2352 |
| chase | 4 | 32 | 1048576 | 256 | 4.158 | 64.553 | 16244 |
| chase | 4 | 32 | 4194304 | 256 | 20.395 | 52.647 | 79668 |
| chase | 4 | 256 | 256 | 256 | 0.102 | 0.640 | 400 |
| chase | 4 | 256 | 1024 | 256 | 0.100 | 2.609 | 392 |
| chase | 4 | 256 | 4096 | 256 | 0.047 | 22.521 | 182 |
| chase | 4 | 256 | 16384 | 256 | 0.055 | 76.028 | 216 |
| chase | 4 | 256 | 65536 | 256 | 0.173 | 96.768 | 677 |
| chase | 4 | 256 | 262144 | 256 | 0.606 | 110.744 | 2367 |
| chase | 4 | 256 | 1048576 | 256 | 4.641 | 57.835 | 18131 |
| chase | 4 | 256 | 4194304 | 256 | 20.554 | 52.240 | 80289 |
| indep x8 | 4 | 256 | 65536 | 256 | 1.188 | 112.984 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 262144 | 256 | 6.520 | 82.343 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 1048576 | 256 | 43.273 | 49.627 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 4194304 | 256 | 171.618 | 50.053 | (8 loads in flight per lane) |
| line 64 B | 4 | 256 | 16384 | 256 | 0.080 | 52.597 G lines/s | 3366.2 GB/s in lines |
| line 64 B | 4 | 256 | 65536 | 256 | 0.280 | 59.850 G lines/s | 3830.4 GB/s in lines |
| line 64 B | 4 | 256 | 262144 | 256 | 1.216 | 55.182 G lines/s | 3531.7 GB/s in lines |
| line 64 B | 4 | 256 | 1048576 | 256 | 9.484 | 28.303 G lines/s | 1811.4 GB/s in lines |
| line 64 B | 4 | 256 | 4194304 | 256 | 43.509 | 24.678 G lines/s | 1579.4 GB/s in lines |
| stream | 4 | 256 | 262144 | 1 | 0.006 | 740.5 GB/s coalesced | (4 MiB read once) |
RESULT memprobe backend=opencl device="NVIDIA GeForce RTX 5090" mib=4 chase_gloads=111.432 chase_ns=412.2 chase_lanes=262144 chase_best_ms=0.602 indep_gloads=112.984 line64_glines=59.850 stream_gbps=740.5 time=event
| chase | 64 | 32 | 256 | 256 | 0.104 | 0.628 | 408 |
| chase | 64 | 32 | 1024 | 256 | 0.106 | 2.471 | 414 |
| chase | 64 | 32 | 4096 | 256 | 0.117 | 8.941 | 458 |
| chase | 64 | 32 | 16384 | 256 | 0.133 | 31.500 | 520 |
| chase | 64 | 32 | 65536 | 256 | 0.160 | 105.026 | 624 |
| chase | 64 | 32 | 262144 | 256 | 0.604 | 111.137 | 2359 |
| chase | 64 | 32 | 1048576 | 256 | 4.754 | 56.470 | 18569 |
| chase | 64 | 32 | 4194304 | 256 | 17.907 | 59.962 | 69949 |
| chase | 64 | 256 | 256 | 256 | 0.104 | 0.630 | 407 |
| chase | 64 | 256 | 1024 | 256 | 0.105 | 2.508 | 408 |
| chase | 64 | 256 | 4096 | 256 | 0.119 | 8.804 | 465 |
| chase | 64 | 256 | 16384 | 256 | 0.148 | 28.432 | 576 |
| chase | 64 | 256 | 65536 | 256 | 0.163 | 103.145 | 635 |
| chase | 64 | 256 | 262144 | 256 | 0.618 | 108.537 | 2415 |
| chase | 64 | 256 | 1048576 | 256 | 3.737 | 71.829 | 14598 |
| chase | 64 | 256 | 4194304 | 256 | 18.488 | 58.079 | 72217 |
| indep x8 | 64 | 256 | 65536 | 256 | 1.272 | 105.520 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 262144 | 256 | 8.506 | 63.114 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 1048576 | 256 | 44.155 | 48.636 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 4194304 | 256 | 183.862 | 46.719 | (8 loads in flight per lane) |
| line 64 B | 64 | 256 | 16384 | 256 | 0.135 | 31.133 G lines/s | 1992.5 GB/s in lines |
| line 64 B | 64 | 256 | 65536 | 256 | 0.320 | 52.408 G lines/s | 3354.1 GB/s in lines |
| line 64 B | 64 | 256 | 262144 | 256 | 1.314 | 51.058 G lines/s | 3267.7 GB/s in lines |
| line 64 B | 64 | 256 | 1048576 | 256 | 9.957 | 26.959 G lines/s | 1725.4 GB/s in lines |
| line 64 B | 64 | 256 | 4194304 | 256 | 44.425 | 24.170 G lines/s | 1546.9 GB/s in lines |
| stream | 64 | 256 | 1048576 | 4 | 0.045 | 1502.3 GB/s coalesced | (64 MiB read once) |
RESULT memprobe backend=opencl device="NVIDIA GeForce RTX 5090" mib=64 chase_gloads=111.137 chase_ns=407.6 chase_lanes=262144 chase_best_ms=0.604 indep_gloads=105.520 line64_glines=52.408 stream_gbps=1502.3 time=event
| chase | 256 | 32 | 256 | 256 | 0.110 | 0.595 | 430 |
| chase | 256 | 32 | 1024 | 256 | 0.106 | 2.462 | 416 |
| chase | 256 | 32 | 4096 | 256 | 0.139 | 7.550 | 542 |
| chase | 256 | 32 | 16384 | 256 | 0.256 | 16.382 | 1000 |
| chase | 256 | 32 | 65536 | 256 | 0.807 | 20.794 | 3152 |
| chase | 256 | 32 | 262144 | 256 | 5.221 | 12.853 | 20395 |
| chase | 256 | 32 | 1048576 | 256 | 25.799 | 10.405 | 100776 |
| chase | 256 | 32 | 4194304 | 256 | 114.536 | 9.375 | 447406 |
| chase | 256 | 256 | 256 | 256 | 0.103 | 0.639 | 401 |
| chase | 256 | 256 | 1024 | 256 | 0.106 | 2.472 | 414 |
| chase | 256 | 256 | 4096 | 256 | 0.127 | 8.237 | 497 |
| chase | 256 | 256 | 16384 | 256 | 0.256 | 16.415 | 998 |
| chase | 256 | 256 | 65536 | 256 | 0.805 | 20.852 | 3143 |
| chase | 256 | 256 | 262144 | 256 | 3.939 | 17.039 | 15385 |
| chase | 256 | 256 | 1048576 | 256 | 27.211 | 9.865 | 106292 |
| chase | 256 | 256 | 4194304 | 256 | 111.792 | 9.605 | 436689 |
| indep x8 | 256 | 256 | 65536 | 256 | 8.720 | 15.392 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 262144 | 256 | 54.796 | 9.798 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 1048576 | 256 | 228.993 | 9.378 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 4194304 | 256 | 931.504 | 9.222 | (8 loads in flight per lane) |
| line 64 B | 256 | 256 | 16384 | 256 | 0.191 | 21.999 G lines/s | 1408.0 GB/s in lines |
| line 64 B | 256 | 256 | 65536 | 256 | 0.508 | 33.045 G lines/s | 2114.9 GB/s in lines |
| line 64 B | 256 | 256 | 262144 | 256 | 7.167 | 9.364 G lines/s | 599.3 GB/s in lines |
| line 64 B | 256 | 256 | 1048576 | 256 | 38.511 | 6.970 G lines/s | 446.1 GB/s in lines |
| line 64 B | 256 | 256 | 4194304 | 256 | 167.877 | 6.396 G lines/s | 409.3 GB/s in lines |
| stream | 256 | 256 | 1048576 | 16 | 0.184 | 1458.4 GB/s coalesced | (256 MiB read once) |
RESULT memprobe backend=opencl device="NVIDIA GeForce RTX 5090" mib=256 chase_gloads=20.852 chase_ns=430.5 chase_lanes=65536 chase_best_ms=0.805 indep_gloads=15.392 line64_glines=33.045 stream_gbps=1458.4 time=event
| chase | 1024 | 32 | 256 | 256 | 0.106 | 0.617 | 415 |
| chase | 1024 | 32 | 1024 | 256 | 0.113 | 2.329 | 440 |
| chase | 1024 | 32 | 4096 | 256 | 0.122 | 8.607 | 476 |
| chase | 1024 | 32 | 16384 | 256 | 0.251 | 16.680 | 982 |
| chase | 1024 | 32 | 65536 | 256 | 0.927 | 18.093 | 3622 |
| chase | 1024 | 32 | 262144 | 256 | 5.227 | 12.839 | 20418 |
| chase | 1024 | 32 | 1048576 | 256 | 31.285 | 8.580 | 122208 |
| chase | 1024 | 32 | 4194304 | 256 | 127.392 | 8.429 | 497624 |
| chase | 1024 | 256 | 256 | 256 | 0.109 | 0.601 | 426 |
| chase | 1024 | 256 | 1024 | 256 | 0.119 | 2.196 | 466 |
| chase | 1024 | 256 | 4096 | 256 | 0.122 | 8.605 | 476 |
| chase | 1024 | 256 | 16384 | 256 | 0.251 | 16.706 | 981 |
| chase | 1024 | 256 | 65536 | 256 | 0.923 | 18.168 | 3607 |
| chase | 1024 | 256 | 262144 | 256 | 7.510 | 8.936 | 29334 |
| chase | 1024 | 256 | 1048576 | 256 | 31.888 | 8.418 | 124561 |
| chase | 1024 | 256 | 4194304 | 256 | 134.392 | 7.990 | 524969 |
| indep x8 | 1024 | 256 | 65536 | 256 | 11.017 | 12.182 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 262144 | 256 | 65.525 | 8.193 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 1048576 | 256 | 264.802 | 8.110 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 4194304 | 256 | 1083.377 | 7.929 | (8 loads in flight per lane) |
| line 64 B | 1024 | 256 | 16384 | 256 | 0.272 | 15.395 G lines/s | 985.3 GB/s in lines |
| line 64 B | 1024 | 256 | 65536 | 256 | 1.056 | 15.884 G lines/s | 1016.6 GB/s in lines |
| line 64 B | 1024 | 256 | 262144 | 256 | 11.554 | 5.808 G lines/s | 371.7 GB/s in lines |
| line 64 B | 1024 | 256 | 1048576 | 256 | 57.203 | 4.693 G lines/s | 300.3 GB/s in lines |
| line 64 B | 1024 | 256 | 4194304 | 256 | 254.929 | 4.212 G lines/s | 269.6 GB/s in lines |
| stream | 1024 | 256 | 1048576 | 64 | 0.646 | 1661.9 GB/s coalesced | (1024 MiB read once) |
RESULT memprobe backend=opencl device="NVIDIA GeForce RTX 5090" mib=1024 chase_gloads=18.168 chase_ns=414.8 chase_lanes=65536 chase_best_ms=0.923 indep_gloads=12.182 line64_glines=15.884 stream_gbps=1661.9 time=event
| alu | 0 | 256 | 1048576 | 4096 | 0.476 | 45157.7 G int ops/s | 53.127 G steps/s per compute unit (approximate: 5 ops per step counted) |
RESULT memprobe backend=opencl device="NVIDIA GeForce RTX 5090" mib=0 alu_gops=45157.7 time=event
memprobe: done
skipping opencl:1 gfx1036 (--only)
=== 6. proving (optional) ===
proving: {"status":"skipped: -NoProve"}
PS>TerminatingError(ConvertTo-Json): "Exception of type 'System.OutOfMemoryException' was thrown."
ConvertTo-Json : Exception of type 'System.OutOfMemoryException' was thrown.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:217
char:11
+ $Result | ConvertTo-Json -Depth 8 | Set-Content -Encoding UTF8 -Liter ...
+ ~~~~~~~~~~~~~~~~~~~~~~~
+ CategoryInfo : NotSpecified: (:) [ConvertTo-Json], OutOfMemoryException
+ FullyQualifiedErrorId : System.OutOfMemoryException,Microsoft.PowerShell.Commands.ConvertToJsonCommand
ConvertTo-Json : Exception of type 'System.OutOfMemoryException' was thrown.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:217
char:11
+ $Result | ConvertTo-Json -Depth 8 | Set-Content -Encoding UTF8 -Liter ...
+ ~~~~~~~~~~~~~~~~~~~~~~~
+ CategoryInfo : NotSpecified: (:) [ConvertTo-Json], OutOfMemoryException
+ FullyQualifiedErrorId : System.OutOfMemoryException,Microsoft.PowerShell.Commands.ConvertToJsonCommand
>> TerminatingError(Set-Content): "Cannot bind argument to parameter 'LiteralPath' because it is an empty string."
Set-Content : Cannot bind argument to parameter 'LiteralPath' because it is an empty string.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:235
char:58
+ $md -join "`n" | Set-Content -Encoding UTF8 -LiteralPath $Md
+ ~~~
+ CategoryInfo : InvalidData: (:) [Set-Content], ParameterBindingValidationException
+ FullyQualifiedErrorId :
ParameterArgumentValidationErrorEmptyStringNotAllowed,Microsoft.PowerShell.Commands.SetContentCommand
Set-Content : Cannot bind argument to parameter 'LiteralPath' because it is an empty string.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:235
char:58
+ $md -join "`n" | Set-Content -Encoding UTF8 -LiteralPath $Md
+ ~~~
+ CategoryInfo : InvalidData: (:) [Set-Content], ParameterBindingValidationException
+ FullyQualifiedErrorId : ParameterArgumentValidationErrorEmptyStringNotAllowed,Microsoft.PowerShell.Commands.SetC
ontentCommand
wrote <app data>\app\jobs\run-repro-pc2-20261006\results\igneum-repro-windows-20261006T005218Z.json
wrote # Igneum reproducible benchmark, windows, 20261006T005218Z Package v0.1.0-repro (39141f5), run 208b02bccab33633, 120 s per card. Signed by nothing; the JSON beside this file is the record. | Machine | | |---|---| | OS | Microsoft Windows 11 Pro 10.0.26200 (AMD64) | | CPU | System.Collections.Hashtable | | Memory | 63132 MiB | | GPU (cuda 0) | NVIDIA GeForce RTX 5090 | | GPU (opencl 0) | NVIDIA GeForce RTX 5090 | | GPU (opencl 1) | gfx1036 | | CPU vectors | verify ms per warp | |---|---| | PASS (96 of 96 lanes) | 0.64 | | Card | Vectors | Fingerprint (2^24 at base 0) | MH/s at 1 GiB | 4 / 64 / 256 / 1024 MiB | Random reads at 1 GiB (G/s, ns) | Stream GB/s | In-cache / 1 GiB | Hash share of read ceiling | |---|---|---|---|---|---|---|---|---| | opencl:0 NVIDIA GeForce RTX 5090 (cross-check) | PASS | 25f96e7dce90bd4e | 62.283 | 375.252 / 385.662 / 72.476 / 62.283 | 18.168, 414.8 | 1661.9 | 6.025 | 0.439 | Proving: skipped: -NoProve Every command run and the sha256 of every binary are in the JSON. Vectors: bit-exact means every lane of the pack's three published warps matched. The fingerprint is the FNV-1a 64 of all 2^24 outputs at base nonce 0: equal fingerprints on two machines mean every one of those hashes agreed.
log <app data>\app\jobs\run-repro-pc2-20261006\results\igneum-repro-windows-20261006T005218Z.log
done: 1 card run(s), CPU vectors PASS
**********************
Windows PowerShell transcript end
End time: 20261006015546
**********************

View file

@ -0,0 +1,316 @@
IGNEUM-APP version=0.3.11 machine=1ccfe586 platform=windows node=igneumd_2.1.0
**********************
Windows PowerShell transcript start
Start time: 20261006015622
Username: <pc-hostname>\Admin
RunAs User: <pc-hostname>\Admin
Configuration Name:
Machine: <pc-hostname> (Microsoft Windows NT 10.0.26200.0)
Host Application: C:\WINDOWS\System32\WindowsPowerShell\v1.0\powershell.exe -ExecutionPolicy Bypass -File <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1 -Seconds 120 -Only opencl:1 -Out <app data>\app\jobs\run-repro-pc2-20261006\results -NoProve
Process ID: 24240
PSVersion: 5.1.26100.9444
PSEdition: Desktop
PSCompatibleVersions: 1.0, 2.0, 3.0, 4.0, 5.0, 5.1.26100.9444
BuildVersion: 10.0.26100.9444
CLRVersion: 4.0.30319.42000
WSManStackVersion: 3.0
PSRemotingProtocolVersion: 2.3
SerializationVersion: 1.1.0.1
**********************
igneum repro windows, run 4134a181cd309190, 20261006T005622Z, 120 s per card, results in <app data>\app\jobs\run-repro-pc2-20261006\results
=== 1. machine ===
os: Microsoft Windows 11 Pro 10.0.26200 (AMD64)
cpu: AMD Ryzen 7 9800X3D 8-Core Processor
memory: 63132 MiB
video controller: AMD Radeon(TM) Graphics, driver 32.0.21042.62
video controller: NVIDIA GeForce RTX 5090, driver 32.0.16.1047
nvidia-smi: 0, NVIDIA GeForce RTX 5090, 610.47, 32607 MiB
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-cuda.exe --list
cuda devices (1), driver 13.3:
[0] NVIDIA GeForce RTX 5090 sm_120 170 SMs 32606 MiB
DEVICE index=0 backend=cuda name="NVIDIA GeForce RTX 5090" arch=sm_120 sms=170 memory_mib=32606
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --list
igneum-bench-cl pack "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000" (test harness: no pool, no network, no wallet)
OpenCL devices (2):
[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
=== 2. vectors on the CPU ===
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-pow.exe check-pack --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --warps 50
check-pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh: seed "igneum-genesis" generator v2 attempt 0 program id bcc1248b10cc90f2 (matches program.json), dataset 2^28 words (memory-hard), 128 loads/hash; built in 155 ms
cache: FNV-1a 64 48c4f5bf24166b2e == vectors.json (48c4f5bf24166b2e)
vector warp base 0 (nonces 0..31): PASS (32 of 32 lanes)
vector warp base 4096 (nonces 4096..4127): PASS (32 of 32 lanes)
vector warp base 1000000 (nonces 1000000..1000031): PASS (32 of 32 lanes)
CPU verify: 0.633 ms per 32-lane warp, avg of 50 (cold 0.729 ms max over the vector warps; checksum 19297e99c7b9a55e)
RESULT check-pack pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed=igneum-genesis program_id=bcc1248b10cc90f2 id_match=true dataset_log2=28 mode=memory-hard cache_fnv=48c4f5bf24166b2e cache_match=true warps=3 warps_pass=3 lanes=96 lanes_pass=96 cpu_verify_ms=0.633 cpu_verify_cold_ms=0.729 verdict=PASS
skipping cuda:0 NVIDIA GeForce RTX 5090 (--only)
skipping opencl:0 NVIDIA GeForce RTX 5090 (--only)
=== card opencl:1 gfx1036 ===
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 1 --seconds 120 --batch-log2 24
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (2):
[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
*[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [1] gfx1036
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
kernel source: <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on gfx1036: cache 29 dataset 382 check 238 ms (649 ms in all); self-test PASS (cache head, last line and FNV-1a 64 48c4f5bf24166b2e; dataset head, word [268435455] and 64 samples; 96 of 96 vector lanes)
kernel: igneum_hash_bound max work-group 32, preferred multiple 32, local memory 256 bytes, private memory 0 bytes, work-group 32, sub-group size 32 (queried through clGetKernelSubGroupInfoKHR)
warm-up dispatch (base 0): 5080.35 ms, fingerprint 25f96e7dce90bd4e; 24 timed dispatches of 16777216 nonces in 121.6 s: mean 5065.43 ms
rate: 3.312 Mhash/s, 1.70 GB/s useful (loads x 4 B), 0.424 G random loads/s
RESULT bench backend=opencl device="gfx1036" platform="AMD Accelerated Parallel Processing" pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=1024 loads=128 check=PASS fingerprint=25f96e7dce90bd4e batch_log2=24 dispatches=24 hashes=402653184 seconds=121.570 mhs=3.312 group_warps=1 exchange=0 time=event
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 1 --batches 5 --batch-log2 24 --dataset-mib 4
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (2):
[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
*[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [1] gfx1036
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
kernel source: <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on gfx1036: cache 28 dataset 3 check 0 ms (31 ms in all); self-test skipped (dataset 4 MiB is not the pack's 1024 MiB: the vectors are for the pack size)
kernel: igneum_hash_bound max work-group 32, preferred multiple 32, local memory 256 bytes, private memory 0 bytes, work-group 32, sub-group size 32 (queried through clGetKernelSubGroupInfoKHR)
warm-up dispatch (base 0): 4513.78 ms, fingerprint 1468be5e1a85e771; 5 timed dispatches of 16777216 nonces in 22.5 s: mean 4508.01 ms
rate: 3.722 Mhash/s, 1.91 GB/s useful (loads x 4 B), 0.476 G random loads/s
RESULT bench backend=opencl device="gfx1036" platform="AMD Accelerated Parallel Processing" pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=4 loads=128 check=skipped fingerprint=1468be5e1a85e771 batch_log2=24 dispatches=5 hashes=83886080 seconds=22.540 mhs=3.722 group_warps=1 exchange=0 time=event
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 1 --batches 5 --batch-log2 24 --dataset-mib 64
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (2):
[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
*[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [1] gfx1036
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
kernel source: <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on gfx1036: cache 28 dataset 26 check 0 ms (54 ms in all); self-test skipped (dataset 64 MiB is not the pack's 1024 MiB: the vectors are for the pack size)
kernel: igneum_hash_bound max work-group 32, preferred multiple 32, local memory 256 bytes, private memory 0 bytes, work-group 32, sub-group size 32 (queried through clGetKernelSubGroupInfoKHR)
warm-up dispatch (base 0): 4959.47 ms, fingerprint 48a2e75ca0dbbd7c; 5 timed dispatches of 16777216 nonces in 24.8 s: mean 4958.92 ms
rate: 3.383 Mhash/s, 1.73 GB/s useful (loads x 4 B), 0.433 G random loads/s
RESULT bench backend=opencl device="gfx1036" platform="AMD Accelerated Parallel Processing" pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=64 loads=128 check=skipped fingerprint=48a2e75ca0dbbd7c batch_log2=24 dispatches=5 hashes=83886080 seconds=24.795 mhs=3.383 group_warps=1 exchange=0 time=event
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --bench --pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh --device 1 --batches 5 --batch-log2 24 --dataset-mib 256
igneum-bench-cl pack "igneum-genesis" (test harness: no pool, no network, no wallet) [--pack: generic serve mode]
OpenCL devices (2):
[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
*[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [1] gfx1036
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
kernel source: <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh/kernel_bound.cl (19923 bytes)
build options: -cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0
exchange: local-memory exchange with barrier (device lists no sub-group shuffle extension)
pack <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh on gfx1036: cache 27 dataset 97 check 0 ms (124 ms in all); self-test skipped (dataset 256 MiB is not the pack's 1024 MiB: the vectors are for the pack size)
kernel: igneum_hash_bound max work-group 32, preferred multiple 32, local memory 256 bytes, private memory 0 bytes, work-group 32, sub-group size 32 (queried through clGetKernelSubGroupInfoKHR)
warm-up dispatch (base 0): 5057.03 ms, fingerprint 3d1523681daf7c2b; 5 timed dispatches of 16777216 nonces in 25.3 s: mean 5053.86 ms
rate: 3.320 Mhash/s, 1.70 GB/s useful (loads x 4 B), 0.425 G random loads/s
RESULT bench backend=opencl device="gfx1036" platform="AMD Accelerated Parallel Processing" pack=<app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\packs\igneum-genesis-mh seed="igneum-genesis" program_id=bcc1248b10cc90f2 dataset_mib=256 loads=128 check=skipped fingerprint=3d1523681daf7c2b batch_log2=24 dispatches=5 hashes=83886080 seconds=25.269 mhs=3.320 group_warps=1 exchange=0 time=event
+ <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\bin\windows-x86_64\igneum-worker-opencl.exe --memprobe --device 1
igneum-bench-cl pack "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000" (test harness: no pool, no network, no wallet)
OpenCL devices (2):
[0] NVIDIA GeForce RTX 5090 | NVIDIA CUDA (OpenCL 3.0 CUDA 13.3.44)
GPU, vendor NVIDIA Corporation, driver 610.47, OpenCL C 1.2 , 170 compute units, 2407 MHz
global 32606 MiB, max alloc 8151 MiB, local 48 KiB, max work-group 1024, sub-group extension: none, NVIDIA warp size 32
DEVICE index=0 backend=opencl type=GPU name="NVIDIA GeForce RTX 5090" vendor="NVIDIA Corporation" platform="NVIDIA CUDA" driver="610.47" compute_units=170 memory_mib=32606
*[1] gfx1036 | AMD Accelerated Parallel Processing (OpenCL 2.1 AMD-APP (3652.0))
GPU, vendor Advanced Micro Devices, Inc., driver 3652.0 (PAL,LC), OpenCL C 2.0 , 1 compute units, 2200 MHz
global 37109 MiB, max alloc 29802 MiB, local 64 KiB, max work-group 256, sub-group extension: cl_khr_subgroups (no shuffle extension), AMD wavefront width 32
DEVICE index=1 backend=opencl type=GPU name="gfx1036" vendor="Advanced Micro Devices, Inc." platform="AMD Accelerated Parallel Processing" driver="3652.0 (PAL,LC)" compute_units=1 memory_mib=37109
using device [1] gfx1036
timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time
memprobe on [AMD Accelerated Parallel Processing] gfx1036, driver 3652.0 (PAL,LC), 1 compute units, 2200 MHz, device event time
memprobe kernel: probe_chase max work-group 256, preferred multiple 32, local memory 0 bytes, private memory 0 bytes, work-group 32, sub-group size 32 (queried through clGetKernelSubGroupInfoKHR)
memprobe kernel: probe_chase max work-group 256, preferred multiple 32, local memory 0 bytes, private memory 0 bytes, work-group 256, sub-group size 32 (queried through clGetKernelSubGroupInfoKHR)
| probe | MiB | work-group | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |
|---|---|---|---|---|---|---|---|
| chase | 4 | 32 | 256 | 256 | 0.129 | 0.508 | 504 |
| chase | 4 | 32 | 1024 | 256 | 0.485 | 0.541 | 1893 |
| chase | 4 | 32 | 4096 | 256 | 1.928 | 0.544 | 7532 |
| chase | 4 | 32 | 16384 | 256 | 7.843 | 0.535 | 30637 |
| chase | 4 | 32 | 65536 | 256 | 33.485 | 0.501 | 130799 |
| chase | 4 | 32 | 262144 | 256 | 135.807 | 0.494 | 530498 |
| chase | 4 | 32 | 1048576 | 256 | 549.573 | 0.488 | 2146769 |
| chase | 4 | 32 | 4194304 | 256 | 2200.918 | 0.488 | 8597336 |
| chase | 4 | 256 | 256 | 256 | 0.128 | 0.512 | 500 |
| chase | 4 | 256 | 1024 | 256 | 0.487 | 0.539 | 1901 |
| chase | 4 | 256 | 4096 | 256 | 1.900 | 0.552 | 7420 |
| chase | 4 | 256 | 16384 | 256 | 7.721 | 0.543 | 30161 |
| chase | 4 | 256 | 65536 | 256 | 32.841 | 0.511 | 128287 |
| chase | 4 | 256 | 262144 | 256 | 133.267 | 0.504 | 520574 |
| chase | 4 | 256 | 1048576 | 256 | 548.374 | 0.490 | 2142085 |
| chase | 4 | 256 | 4194304 | 256 | 2178.067 | 0.493 | 8508076 |
| indep x8 | 4 | 256 | 65536 | 256 | 274.357 | 0.489 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 262144 | 256 | 1101.550 | 0.487 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 1048576 | 256 | 4351.166 | 0.494 | (8 loads in flight per lane) |
| indep x8 | 4 | 256 | 4194304 | 256 | 17318.075 | 0.496 | (8 loads in flight per lane) |
| line 64 B | 4 | 256 | 16384 | 256 | 7.940 | 0.528 G lines/s | 33.8 GB/s in lines |
| line 64 B | 4 | 256 | 65536 | 256 | 33.904 | 0.495 G lines/s | 31.7 GB/s in lines |
| line 64 B | 4 | 256 | 262144 | 256 | 130.045 | 0.516 G lines/s | 33.0 GB/s in lines |
| line 64 B | 4 | 256 | 1048576 | 256 | 525.637 | 0.511 G lines/s | 32.7 GB/s in lines |
| line 64 B | 4 | 256 | 4194304 | 256 | 2103.688 | 0.510 G lines/s | 32.7 GB/s in lines |
| stream | 4 | 256 | 262144 | 1 | 0.092 | 45.6 GB/s coalesced | (4 MiB read once) |
RESULT memprobe backend=opencl device="gfx1036" mib=4 chase_gloads=0.552 chase_ns=504.4 chase_lanes=4096 chase_best_ms=1.900 indep_gloads=0.496 line64_glines=0.528 stream_gbps=45.6 time=event
| chase | 64 | 32 | 256 | 256 | 0.142 | 0.462 | 554 |
| chase | 64 | 32 | 1024 | 256 | 0.566 | 0.463 | 2211 |
| chase | 64 | 32 | 4096 | 256 | 2.260 | 0.464 | 8830 |
| chase | 64 | 32 | 16384 | 256 | 9.090 | 0.461 | 35508 |
| chase | 64 | 32 | 65536 | 256 | 38.049 | 0.441 | 148630 |
| chase | 64 | 32 | 262144 | 256 | 152.015 | 0.441 | 593808 |
| chase | 64 | 32 | 1048576 | 256 | 610.369 | 0.440 | 2384255 |
| chase | 64 | 32 | 4194304 | 256 | 2445.044 | 0.439 | 9550953 |
| chase | 64 | 256 | 256 | 256 | 0.146 | 0.448 | 572 |
| chase | 64 | 256 | 1024 | 256 | 0.566 | 0.463 | 2211 |
| chase | 64 | 256 | 4096 | 256 | 2.259 | 0.464 | 8825 |
| chase | 64 | 256 | 16384 | 256 | 9.065 | 0.463 | 35410 |
| chase | 64 | 256 | 65536 | 256 | 37.898 | 0.443 | 148039 |
| chase | 64 | 256 | 262144 | 256 | 151.855 | 0.442 | 593184 |
| chase | 64 | 256 | 1048576 | 256 | 610.919 | 0.439 | 2386403 |
| chase | 64 | 256 | 4194304 | 256 | 2442.272 | 0.440 | 9540125 |
| indep x8 | 64 | 256 | 65536 | 256 | 305.276 | 0.440 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 262144 | 256 | 1220.763 | 0.440 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 1048576 | 256 | 4882.786 | 0.440 | (8 loads in flight per lane) |
| indep x8 | 64 | 256 | 4194304 | 256 | 19540.168 | 0.440 | (8 loads in flight per lane) |
| line 64 B | 64 | 256 | 16384 | 256 | 8.807 | 0.476 G lines/s | 30.5 GB/s in lines |
| line 64 B | 64 | 256 | 65536 | 256 | 37.720 | 0.445 G lines/s | 28.5 GB/s in lines |
| line 64 B | 64 | 256 | 262144 | 256 | 152.015 | 0.441 G lines/s | 28.3 GB/s in lines |
| line 64 B | 64 | 256 | 1048576 | 256 | 611.038 | 0.439 G lines/s | 28.1 GB/s in lines |
| line 64 B | 64 | 256 | 4194304 | 256 | 2443.542 | 0.439 G lines/s | 28.1 GB/s in lines |
| stream | 64 | 256 | 1048576 | 4 | 1.072 | 62.6 GB/s coalesced | (64 MiB read once) |
RESULT memprobe backend=opencl device="gfx1036" mib=64 chase_gloads=0.464 chase_ns=553.9 chase_lanes=4096 chase_best_ms=2.259 indep_gloads=0.440 line64_glines=0.476 stream_gbps=62.6 time=event
| chase | 256 | 32 | 256 | 256 | 0.156 | 0.419 | 611 |
| chase | 256 | 32 | 1024 | 256 | 0.589 | 0.445 | 2300 |
| chase | 256 | 32 | 4096 | 256 | 2.295 | 0.457 | 8966 |
| chase | 256 | 32 | 16384 | 256 | 9.192 | 0.456 | 35905 |
| chase | 256 | 32 | 65536 | 256 | 38.481 | 0.436 | 150318 |
| chase | 256 | 32 | 262144 | 256 | 154.592 | 0.434 | 603874 |
| chase | 256 | 32 | 1048576 | 256 | 616.953 | 0.435 | 2409971 |
| chase | 256 | 32 | 4194304 | 256 | 2474.658 | 0.434 | 9666634 |
| chase | 256 | 256 | 256 | 256 | 0.144 | 0.454 | 564 |
| chase | 256 | 256 | 1024 | 256 | 0.575 | 0.456 | 2244 |
| chase | 256 | 256 | 4096 | 256 | 2.294 | 0.457 | 8962 |
| chase | 256 | 256 | 16384 | 256 | 9.199 | 0.456 | 35932 |
| chase | 256 | 256 | 65536 | 256 | 38.452 | 0.436 | 150202 |
| chase | 256 | 256 | 262144 | 256 | 153.788 | 0.436 | 600734 |
| chase | 256 | 256 | 1048576 | 256 | 619.350 | 0.433 | 2419337 |
| chase | 256 | 256 | 4194304 | 256 | 2473.977 | 0.434 | 9663974 |
| indep x8 | 256 | 256 | 65536 | 256 | 309.074 | 0.434 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 262144 | 256 | 1234.677 | 0.435 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 1048576 | 256 | 4946.973 | 0.434 | (8 loads in flight per lane) |
| indep x8 | 256 | 256 | 4194304 | 256 | 19788.057 | 0.434 | (8 loads in flight per lane) |
| line 64 B | 256 | 256 | 16384 | 256 | 9.381 | 0.447 G lines/s | 28.6 GB/s in lines |
| line 64 B | 256 | 256 | 65536 | 256 | 39.275 | 0.427 G lines/s | 27.3 GB/s in lines |
| line 64 B | 256 | 256 | 262144 | 256 | 156.017 | 0.430 G lines/s | 27.5 GB/s in lines |
| line 64 B | 256 | 256 | 1048576 | 256 | 625.957 | 0.429 G lines/s | 27.4 GB/s in lines |
| line 64 B | 256 | 256 | 4194304 | 256 | 2509.253 | 0.428 G lines/s | 27.4 GB/s in lines |
| stream | 256 | 256 | 1048576 | 16 | 4.239 | 63.3 GB/s coalesced | (256 MiB read once) |
RESULT memprobe backend=opencl device="gfx1036" mib=256 chase_gloads=0.457 chase_ns=611.1 chase_lanes=4096 chase_best_ms=2.294 indep_gloads=0.435 line64_glines=0.447 stream_gbps=63.3 time=event
| chase | 1024 | 32 | 256 | 256 | 0.145 | 0.452 | 566 |
| chase | 1024 | 32 | 1024 | 256 | 0.575 | 0.456 | 2246 |
| chase | 1024 | 32 | 4096 | 256 | 2.303 | 0.455 | 8995 |
| chase | 1024 | 32 | 16384 | 256 | 9.220 | 0.455 | 36016 |
| chase | 1024 | 32 | 65536 | 256 | 38.454 | 0.436 | 150211 |
| chase | 1024 | 32 | 262144 | 256 | 154.089 | 0.436 | 601910 |
| chase | 1024 | 32 | 1048576 | 256 | 619.025 | 0.434 | 2418066 |
| chase | 1024 | 32 | 4194304 | 256 | 2475.697 | 0.434 | 9670691 |
| chase | 1024 | 256 | 256 | 256 | 0.224 | 0.293 | 874 |
| chase | 1024 | 256 | 1024 | 256 | 0.578 | 0.453 | 2260 |
| chase | 1024 | 256 | 4096 | 256 | 2.291 | 0.458 | 8949 |
| chase | 1024 | 256 | 16384 | 256 | 9.328 | 0.450 | 36438 |
| chase | 1024 | 256 | 65536 | 256 | 38.466 | 0.436 | 150260 |
| chase | 1024 | 256 | 262144 | 256 | 154.251 | 0.435 | 602543 |
| chase | 1024 | 256 | 1048576 | 256 | 618.853 | 0.434 | 2417395 |
| chase | 1024 | 256 | 4194304 | 256 | 2477.883 | 0.433 | 9679230 |
| indep x8 | 1024 | 256 | 65536 | 256 | 309.134 | 0.434 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 262144 | 256 | 1239.159 | 0.433 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 1048576 | 256 | 4956.469 | 0.433 | (8 loads in flight per lane) |
| indep x8 | 1024 | 256 | 4194304 | 256 | 19831.795 | 0.433 | (8 loads in flight per lane) |
| line 64 B | 1024 | 256 | 16384 | 256 | 9.435 | 0.445 G lines/s | 28.5 GB/s in lines |
| line 64 B | 1024 | 256 | 65536 | 256 | 39.092 | 0.429 G lines/s | 27.5 GB/s in lines |
| line 64 B | 1024 | 256 | 262144 | 256 | 157.228 | 0.427 G lines/s | 27.3 GB/s in lines |
| line 64 B | 1024 | 256 | 1048576 | 256 | 629.086 | 0.427 G lines/s | 27.3 GB/s in lines |
| line 64 B | 1024 | 256 | 4194304 | 256 | 2520.001 | 0.426 G lines/s | 27.3 GB/s in lines |
| stream | 1024 | 256 | 1048576 | 64 | 16.922 | 63.5 GB/s coalesced | (1024 MiB read once) |
RESULT memprobe backend=opencl device="gfx1036" mib=1024 chase_gloads=0.458 chase_ns=565.9 chase_lanes=4096 chase_best_ms=2.291 indep_gloads=0.434 line64_glines=0.445 stream_gbps=63.5 time=event
| alu | 0 | 256 | 1048576 | 4096 | 105.925 | 202.7 G int ops/s | 40.547 G steps/s per compute unit (approximate: 5 ops per step counted) |
RESULT memprobe backend=opencl device="gfx1036" mib=0 alu_gops=202.7 time=event
memprobe: done
=== 6. proving (optional) ===
proving: {"status":"skipped: -NoProve"}
PS>TerminatingError(ConvertTo-Json): "Exception of type 'System.OutOfMemoryException' was thrown."
ConvertTo-Json : Exception of type 'System.OutOfMemoryException' was thrown.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:217
char:11
+ $Result | ConvertTo-Json -Depth 8 | Set-Content -Encoding UTF8 -Liter ...
+ ~~~~~~~~~~~~~~~~~~~~~~~
+ CategoryInfo : NotSpecified: (:) [ConvertTo-Json], OutOfMemoryException
+ FullyQualifiedErrorId : System.OutOfMemoryException,Microsoft.PowerShell.Commands.ConvertToJsonCommand
ConvertTo-Json : Exception of type 'System.OutOfMemoryException' was thrown.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:217
char:11
+ $Result | ConvertTo-Json -Depth 8 | Set-Content -Encoding UTF8 -Liter ...
+ ~~~~~~~~~~~~~~~~~~~~~~~
+ CategoryInfo : NotSpecified: (:) [ConvertTo-Json], OutOfMemoryException
+ FullyQualifiedErrorId : System.OutOfMemoryException,Microsoft.PowerShell.Commands.ConvertToJsonCommand
>> TerminatingError(Set-Content): "Cannot bind argument to parameter 'LiteralPath' because it is an empty string."
Set-Content : Cannot bind argument to parameter 'LiteralPath' because it is an empty string.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:235
char:58
+ $md -join "`n" | Set-Content -Encoding UTF8 -LiteralPath $Md
+ ~~~
+ CategoryInfo : InvalidData: (:) [Set-Content], ParameterBindingValidationException
+ FullyQualifiedErrorId :
ParameterArgumentValidationErrorEmptyStringNotAllowed,Microsoft.PowerShell.Commands.SetContentCommand
Set-Content : Cannot bind argument to parameter 'LiteralPath' because it is an empty string.
At <app data>\app\jobs\fetch-repro-pc2-20261006\repro\igneum-repro-v0.1.0-repro\repro.ps1:235
char:58
+ $md -join "`n" | Set-Content -Encoding UTF8 -LiteralPath $Md
+ ~~~
+ CategoryInfo : InvalidData: (:) [Set-Content], ParameterBindingValidationException
+ FullyQualifiedErrorId : ParameterArgumentValidationErrorEmptyStringNotAllowed,Microsoft.PowerShell.Commands.SetC
ontentCommand
wrote <app data>\app\jobs\run-repro-pc2-20261006\results\igneum-repro-windows-20261006T005622Z.json
wrote # Igneum reproducible benchmark, windows, 20261006T005622Z Package v0.1.0-repro (39141f5), run 4134a181cd309190, 120 s per card. Signed by nothing; the JSON beside this file is the record. | Machine | | |---|---| | OS | Microsoft Windows 11 Pro 10.0.26200 (AMD64) | | CPU | System.Collections.Hashtable | | Memory | 63132 MiB | | GPU (cuda 0) | NVIDIA GeForce RTX 5090 | | GPU (opencl 0) | NVIDIA GeForce RTX 5090 | | GPU (opencl 1) | gfx1036 | | CPU vectors | verify ms per warp | |---|---| | PASS (96 of 96 lanes) | 0.633 | | Card | Vectors | Fingerprint (2^24 at base 0) | MH/s at 1 GiB | 4 / 64 / 256 / 1024 MiB | Random reads at 1 GiB (G/s, ns) | Stream GB/s | In-cache / 1 GiB | Hash share of read ceiling | |---|---|---|---|---|---|---|---|---| | opencl:1 gfx1036 | PASS | 25f96e7dce90bd4e | 3.312 | 3.722 / 3.383 / 3.32 / 3.312 | 0.458, 565.9 | 63.5 | 1.124 | 0.926 | Proving: skipped: -NoProve Every command run and the sha256 of every binary are in the JSON. Vectors: bit-exact means every lane of the pack's three published warps matched. The fingerprint is the FNV-1a 64 of all 2^24 outputs at base nonce 0: equal fingerprints on two machines mean every one of those hashes agreed.
log <app data>\app\jobs\run-repro-pc2-20261006\results\igneum-repro-windows-20261006T005622Z.log
done: 1 card run(s), CPU vectors PASS
**********************
Windows PowerShell transcript end
End time: 20261006020937
**********************

218
docs/benchmarks/repro.md Normal file
View file

@ -0,0 +1,218 @@
# The reproducible benchmark package
6 October 2026 (built on the evening of 5 October). One command per platform reproduces the numbers on the bench table
on an outsider's own machine on day one: `bench/repro.sh` (Linux, macOS) and `bench/repro.ps1` (Windows), shipped as
`igneum-repro-<tag>.tar.gz` and `.zip` next to the public downloads (`packaging/ota/publish-public.sh --repro`, aliases
`/public/igneum-repro.tar.gz` and `/public/igneum-repro.zip`; not deployed tonight). Package v0.1.0-repro was built from
commit `39141f5` plus the uncommitted package sources of this branch (the commit that carries them is named at the end),
and run end to end on the three project machines. The deltas against the bench log are below.
Why it exists: `docs/evidence.md` has 30 claims and none is reproduced externally. The ladder from "tested by the team"
to "reproduced externally" needs "the command published and a third party's run with the same result". This is the
command. The reward for running it is section 8.1 of `docs/benchmarks/proving-e2e.md`, quoted unchanged below.
## 1. What one run does
| Step | What runs | Binary | What it reports |
|---|---|---|---|
| 1 | The machine | the workers' `--list` | OS, CPU, memory; every GPU each worker sees (name, driver, memory) |
| 2 | The lottery hash vectors of the published genesis pack (`proto-cuda/packs/igneum-genesis-mh`: seed `igneum-genesis`, day 2026-10-03, generator 2, program id `bcc1248b10cc90f2`, memory-hard 1 GiB dataset) | `igneum-pow check-pack` on the CPU; `igneum-worker-cuda --bench`, `igneum-worker-opencl --bench --pack`, `igneum-bench --pack` on every GPU | bit-exact or not: 3 warps, 96 lanes, the 256 MiB cache FNV, on the CPU through the Rust interpreter and on every GPU through its own compiler (NVRTC, the vendor's OpenCL compiler, Metal) |
| 3 | The hash benchmark, 120 s per card at 1 GiB, `--batch-log2 24` | the same workers, `--seconds 120` | MH/s over the summed dispatch time, the hashes done, the fingerprint of the first 2^24 outputs at base nonce 0 (FNV-1a 64 over 16.7 million hashes: equal on two machines means every one of them agreed) |
| 4 | The random-read probe at 4, 64, 256 and 1024 MiB | `--memprobe` | dependent random 4-byte reads per second (the hash's access pattern) and their latency at 256 lanes, eight independent chains, 16 and 64-byte lines, the coalesced stream, an integer chain |
| 5 | The chip-resistance sweep: the same program at 4, 64, 256 and 1024 MiB, 5 batches each | `--bench --dataset-mib N` (`--sweep` on Metal) | MH/s per size; in-cache rate over the 1 GiB rate; the hash's share of the card's random-read ceiling at 1 GiB (MH/s x 128 loads against the chase) |
| 6 | One fixture shard proven with the pinned guest and verified (optional) | `igneum-prove-host --mode shard --shard 0`, then `--mode verify` | execute, core and compressed proof times, proof bytes, VERIFIED or not; skipped with the reason when there is no 12 GB NVIDIA card or no prover host |
| 7 | The result | the script | `results/igneum-repro-<os>-<time>.json` (format `igneum-repro-1`: machine, every binary's sha256, the pack files' sha256, every command run with its exit status, the rows in the bench table's own shape, the tolerances) and the same as a table in `.md`, plus the full log |
The workers were extended for this (the flags are in every worker now, where the read-width experiment of 5 October had
put `--bench` and `--memprobe` on the CUDA worker and `--memprobe` on the OpenCL worker, on branches): `--list`, `--bench
--pack <dir> --seconds N --dataset-mib N`, `--memprobe` at 4, 64, 256 and 1024 MiB with one `RESULT` line per size and
a `DEVICE` line per card; the Metal worker got `--list`, `--pack` (the published vectors against the GPU's warps),
`--seconds`, `--sweep` and `--memprobe`; `igneum-pow` got `check-pack`. The pack reader accepts a string-seed pack
(the genesis pack's 14-byte seed; the chain's are 32 bytes), the same change the read-width branch made.
What a run needs: no secrets, no node, no wallet, no network. The CUDA worker needs the NVIDIA driver only on Windows
(NVIDIA's runtime compiler is in the package, `THIRD-PARTY.md`) and the driver plus `libnvrtc.so.12` on Linux. The
OpenCL worker needs the vendor's OpenCL. Nothing here earns anything.
## 2. The three runs of 5 October 2026 against the bench log
Package v0.1.0-repro, three machines asked for, two run tonight (the Mac in full, PC 2 with its 5090 shared with the live miner): PC 1 (ae432dc7: RTX 5090, RX 9070 XT, gfx1036) went
down at 22:31:06 UTC on 5 October before its second slot (its first run at 21:04 UTC ended in a PowerShell parse error
within a second per card, section 7) and nothing reaches it until the morning; its run is owed. The Mac ran under the
measure lock (every build and simulation slot held; load average 6.19 at the start, 4.15 at the end). PC 2's run is in the
coordinator's queue (section 2.2 is filled from it; "pending" until then).
### 2.1 The Mac (Apple M5 Max, 40 GPU cores, 64 GB, macOS 26.6.2), run 2fdd7365db4a5deb, 21:51 UTC, 5 October 2026
Result files: `docs/benchmarks/repro-2026-10-06/mac/`. 7 min 52 s for the two backends (Metal 120 s, Apple OpenCL 30 s as
the cross-check, each with the sweep and the probe table).
| Number | The package | The bench log | Delta | Reading |
|---|---|---|---|---|
| Vectors, CPU (Rust interpreter) | 96/96 lanes, cache FNV `48c4f5bf24166b2e` | 96/96 (4 October 2026, generator version 2 adopted) | exact | |
| Vectors, Metal | 96/96, program id `bcc1248b10cc90f2` matches | 96/96 (4 October 2026) | exact | |
| Vectors, Apple OpenCL | 96/96 | 96/96 (proto-opencl README, 4 October 2026) | exact | |
| Fingerprint of 2^24 outputs at base 0 | `25f96e7dce90bd4e` on Metal and on Apple OpenCL | none at 2^24 for this pack (the README's `f2a95d5bb84d961e` is at 2^13) | new: two compilers agree over 16.7 million nonces | the number a second machine must match |
| Hash MH/s at 1 GiB, Metal, 120 s, GPU time | 27.674 (194 dispatches of 2^24, 3.25 G hashes) | 27.9 through Apple OpenCL on this pack (README, 5 x 2^24, wall); 26.7 mining through Metal on the live devnet (4 October) | -0.8% against the bench, +3.6% against mining | the bench row |
| Hash MH/s at 1 GiB, Apple OpenCL, 30 s, wall | 26.875 | 27.9 (the same README figure) | -3.7% | outside the 3% tolerance against a 5-batch wall-time figure taken on 4 October; the cross-check path is 2.9% under Metal here, where the 3 October runs had the two Apple paths equal (45.2 against 45.0 on the version 1 program). Not the card's row |
| CPU verify, ms per 32-lane warp (avg of 50, one performance core) | 0.611 (cold 0.625) | 0.631 (4 October, version 2 units) | -3% | 16x under the 10 ms gate |
| Sweep, MH/s at 4 / 64 / 256 / 1024 MiB, Metal | 275.2 / 103.2 / 56.6 / 27.8 | 569 / 183 / 94 / 44 (3 October, version 1 program with 104 loads, closed-form dataset) | not comparable: the generator changed | in-cache over 1 GiB 9.9x against 12.9x on version 1 |
| Random reads at 1 GiB, G loads/s (chase) | 3.40 Metal, 3.40 Apple OpenCL | none for Apple before this (4.6 G was derived from the version 1 hash rate on 3 October) | new | the hash does 27.67 x 128 = 3.54 G loads/s, 1.04 of the single-chain chase: on this chip the hash is at the random-read ceiling the probe sees |
| Dependent-load latency, ns at 256 lanes | 522 Metal (GPU time), 1,273 Apple OpenCL (wall, launch included) | none | new | |
| Stream, GB/s at 1 GiB | 546 Metal, 515 Apple OpenCL | 427 GB/s dataset fill (3 October, a write) | reads 28% above the write figure | |
| Proving | skipped: the GPU prover needs a 12 GB NVIDIA card | | | as designed |
### 2.2 PC 2 (RTX 5090, gfx1036, Windows 11), 1ccfe586
Run job run-repro-pc2-20261006 (fetch-repro-pc2-20261006 first), 00:44 to 01:09 UTC, 6 October 2026, on Igneum Miner
0.3.11, one card at a time through `bench/jobs/pc-repro.ps1`. What came back: every RESULT line in the log intake (the
three per-card transcript logs are in `docs/benchmarks/repro-2026-10-06/pc2/`; the JSON and the markdown were never
written, section 7). The 5090 block ran BESIDE THE APP'S LIVE CUDA WORKER: the card-off POST was accepted, but
`api/state` answered `{}` (a 0.3.11 defect on a proving machine) so nothing confirmed a stopped worker, and nvidia-smi
showed 10,176 MiB in use before the bench and the card at 348 W after it. Every 5090 rate below is therefore a shared-card
figure, about half of what the card does alone (131 to 137 MH/s at this batch in every other run of the night), and is
not the card's row. The checks (vectors, fingerprints, the proof) stand regardless of the sharing. The gfx1036 block is the
integrated chip's own figure (3.312 MH/s, where it mines at 3.3 in the app).
| Number | The package (PC 2) | The bench log | Delta | Reading |
|---|---|---|---|---|
| Vectors, CPU (Ryzen 7 9800X3D) | 96/96, cache FNV `48c4f5bf24166b2e` | 96/96 on the M5 Max | exact | first CPU of the second architecture |
| Vectors, RTX 5090 through CUDA (NVRTC, sm_120) | 96/96, program id matches | the version 2 pack had not run on real NVIDIA silicon (evidence row 4) | exact: closes that gap | |
| Vectors, RTX 5090 through NVIDIA OpenCL | 96/96 | 96/96 on the version 1 pack (3 October) | exact | |
| Vectors, gfx1036 through AMD OpenCL | 96/96 | 96/96 on the version 1 pack (3 October) | exact: the version 2 pack on real AMD silicon | |
| Fingerprint of 2^24 outputs at base 0 | `25f96e7dce90bd4e` on CUDA, NVIDIA OpenCL and AMD OpenCL | `25f96e7dce90bd4e` on Metal and Apple OpenCL (the Mac run) | exact across five compilers and three vendors | the cross-vendor claim over 16.7 million nonces |
| Hash MH/s at 1 GiB, 5090, CUDA, 120 s | 62.412 (447 dispatches, 7.50 G hashes), beside the live miner | 139.7 with the card to itself on a version 2 pack (4 October, variant racing); 124.2 mining in the app | not comparable: shared card | re-run owed with the worker confirmed stopped |
| Hash MH/s at 1 GiB, 5090, NVIDIA OpenCL, 30 s | 62.283, beside the live miner | 219.6 against 229.0 CUDA on the version 1 pack (3 October) | not comparable | the two paths agree with each other to 0.2% even shared |
| Hash MH/s at 1 GiB, gfx1036, AMD OpenCL, 120 s | 3.312 (24 dispatches) | 4.38 on the version 1 pack (104 loads, 3 October); 3.3 mining on the live devnet (4 October) | +0.4% against mining; the version 1 figure is another program | the card's row |
| Sweep 4 / 64 / 256 / 1024 MiB, 5090 CUDA (shared) | 374 / 367 / 74 / 62 | 1,340 / 1,353 / 270 / 229 (3 October, version 1, card alone) | shape only: in-cache over 1 GiB 6.0x against 5.8x | the L2 edge is the same |
| Sweep, gfx1036 | 3.72 / 3.38 / 3.32 / 3.31 | none | new | flat: the integrated chip is not cache-bound at any size, its memory is the system's |
| Random reads at 1 GiB, 5090 (chase, best lanes) | 17.5 G loads/s CUDA (wall), 18.2 NVIDIA OpenCL (event); latency 415 ns (OpenCL event) | 23.7 G derived from the version 1 hash rate (3 October); about 16 to 18 G implied by the 9070 XT entry's 6.6x (5 October) | inside the implied range, shared card | |
| Random reads at 1 GiB, gfx1036 | 0.458 G loads/s, 566 ns | none | new | the hash does 3.312 x 128 = 0.424 G loads/s: 0.93 of the chase; at the ceiling like the 9070 XT (0.92) |
| Stream at 1 GiB, 5090 | 1,662 GB/s (OpenCL event), 339 (CUDA wall, the launch dominates) | 1,638 GB/s dataset fill (3 October) | +1.5% | the memory clock was in its full state |
| Integer chain, 5090 | 13.4 T int ops/s (CUDA wall), 45.2 (OpenCL event) | none | | the wall figure is launch-bound; the event figure is the card's |
| Proving, block-338-shard1 shard 0, the live host `/opt/igneum/igneum-prove-host` (pinned shard id `0x2b1a81cb...`) | execute 1.48 s (60.4 M cycles), core 20.2 s (18.1 MB, verify 0.60 s), compressed 33.9 s (1.27 MB, verify 0.038 s), VERIFIED; two tampered witnesses REJECTED | the proving agent the same night: 10.8 to 11.4 s compressed alone, 33 s with the miner running | the "with the miner" figure | consistent with the shared card |
| Job size, 5090 CUDA (shared): 2^21 against 2^24 | genesis pack 60.0 against 63.2 MH/s; live devnet pack 60.9 against 63.4 | the app mines at 2^21: 114.0 wall against 116.0 inside jobs at 22:40 UTC (PC 2's STATUS lines, the consequences reviewer) | 5% lower at 2^21 here, 1.8% wall-to-inside in the app | the per-job cost at 2^21 is about 5% on a shared card; the 15% gap the reviewer found between bench and app is not this alone. The rows are PC 2's; PC 1's are owed |
### 2.3 PC 1 (RTX 5090, RX 9070 XT, gfx1036, Windows 11), ae432dc7
Owed. The first job (run-repro-pc1-20261005, 21:04 to 21:09 UTC) switched each card off and on for nothing: repro.ps1
had three PowerShell parse errors (two missing parentheses on the wsl.exe lines, one `(if ...)` used as an expression),
found only by the Windows parser on the PC because the Mac has no PowerShell. The fix, the parse gate (the PC job now
parses the package script before any card is touched, and the bench/ folder is in the Windows PowerShell 5.1 parse job
of `.github/workflows/windows.yml`) and the second package were ready at 21:45 UTC; PC 1 went down at 22:31:06 UTC before
its second slot. One fact from the first job: the RX 9070 XT (gfx1201) was on the bus and switched off and on by the
app at 21:06 UTC, where the coordinator's queue had it absent since 20:40 UTC.
## 3. The deltas, and what they mean for each user tier
What each number means for each user tier, and what the package does about it (the standing rule of 5 October 2026):
| Number | Home miner, one card (8, 12, 16, 24 or 32 GB) | Rig | Pool user | What the package does |
|---|---|---|---|---|
| The bench needs 1.3 GB of device memory (1 GiB dataset, 256 MiB cache, a 128 MiB output buffer at 2^24) | every tier runs it, 8 GB included; a 4 GB card or an integrated chip with 4 GB shared runs it too | one card at a time, `--only`, so the rig keeps mining on the rest | the same | reports the free memory it found when a dataset does not fit |
| 120 s per card, plus 3 short sweep sizes and the probe (about 4 min per card) | a few minutes with the card idle | a 6-card rig is 25 min, and `--only` splits it | | `--seconds` and the skip flags |
| The hash share of the random-read ceiling (1.04 on the M5 Max, PC 2 pending) | the number that says a card is at its memory's limit, not the kernel's: a card under about 0.9 has a worker problem, not a card problem | the same per card | | the probe table is in every result, so a low share names the step |
| Apple OpenCL 2.9% under Metal on the same chip | an Apple user on the app gets the Metal rate (the app's worker is Metal); the OpenCL figure is a compiler cross-check, never the app's | | | the OpenCL row is marked cross-check and makes no bench-table row |
| CPU verify 0.611 ms per warp on an M5 Max performance core | the chain's gate is 10 ms on any core; a 2019-class laptop core is still unmeasured (evidence row 5, O-1.14) | | a pool verifying shares has 16x margin on this core | `check-pack` prints the figure on any machine that runs the package; the first slow core to run it answers O-1.14 |
| The proving step: skipped on macOS by design; on NVIDIA it needs a built prover host and a 12 GB card by the gate, and the 5090 run of the proving agent tonight peaked at 28.3 GB (bench-log, 5 October, "the 5090 alone proves block-338-shard1 shard 0 compressed in 10.8 to 11.4 s at a 28.3 GB peak") | a 12 GB, 16 GB or 24 GB owner cannot run step 6 today: the gate lets them start it and the prover would fail on memory. The litepaper's "a 12 GB card proves a shard" stays "designed" (evidence row 16) until the prover's peak is under 12 GB | | | step 6 says why it skipped; the next package raises the gate to the measured peak (28 GB) until the prover fits 12 GB, so nobody's run fails late. Filed below as owed work |
| One PC run per night at most on the project's own PCs | | | | the package itself takes minutes; the project's PCs are shared by many agents and PC 1 was down tonight, so the project's three-machine table is not complete on the first night. The outsider's run does not have that constraint |
PC 2 (6 October 2026): the 5090's package numbers are void as the card's figures (beside the live miner) and stand only
as checks; the morning re-runs the 5090 block alone with the worker confirmed stopped by nvidia-smi's compute-apps list
(`bench/jobs/pc-repro.ps1` does that now, and labels a run "UNCONFIRMED" when it cannot). For a 5090 owner the package's
own figure is still owed; for a gfx1036 owner the figure is 3.31 MH/s, flat across dataset sizes, at 0.93 of the chip's
random-read ceiling. For a proving owner: the live host proved and verified the fixture through the package's step 6 on a
32 GB card even with a miner on it (33.9 s compressed); the 12 GB gate stands as written in the table above.
## 4. Tolerances and how a result becomes a row
| Number | Tolerance | Why this width |
|---|---|---|
| Vectors (3 warps, 96 lanes, cache FNV) | exact | a hash either matches or it does not |
| Fingerprint of 2^24 outputs at base 0 | exact across machines at the same `--batch-log2` | the same |
| Hash MH/s at 1 GiB (120 s) | 3% | two 120-s runs of the same card on this evening's machines agree within about 1% with the card to itself; 3% leaves room for a driver version and a warmer card |
| Sweep MH/s at 4, 64, 256 MiB | 10% | 5 batches each, not 120 s; the in-cache sizes are the noisiest (a few hundred ms per batch) |
| Random reads at 1 GiB (G loads/s), stream GB/s | 10% | best of 3 at 8 lane counts; the memory clock's power state moves these |
| Dependent-load latency (ns at 256 lanes) | 15% | one warp, a few hundred ns, timer resolution |
| CPU verify per warp | under 10 ms (the chain's gate), no tolerance on the figure itself | any core must verify a warp inside the gate; the published cores are under 1 ms |
| Proof times (execute, core, compressed) | 25% | SP1's GPU prover varies run to run with the card's state and the server's warm-up |
The path from a submitted file to the bench table (`site/miner-bench.json`, rendered at `/miners` by `site/build.mjs`):
1. `node bench/ingest.mjs check <result.json>` prints the agreement table: every number against the team's reference
for that card model (`bench/reference.json`, which names the bench-log entry or this document's run behind every
reference), with the tolerance and the delta.
2. `node bench/ingest.mjs add <result.json> --by "reproduced externally"` appends one row per card with that label when
every number agrees, and refuses the label (recording "submitted" with the deltas in the note) when one does not. A
failing run is as public as a passing one (proving-e2e.md 8.2 step 4). The site build then fails on any private
string (`site/forbidden-strings.txt`), so a hostname in a note never reaches the page.
3. The row's `source` is `repro:<run id>`, the random id the script drew; the result file is kept under
`docs/benchmarks/repro-<date>/`, as the three of tonight are.
The relay's console (`relay/api/console.mjs`, `results`) shows bench entries synced from `docs/bench-log.md` by
`tools/console.mjs sync-bench`; a submitted result enters there through the bench-log entry that records it, after the
table row. The relay's `bench` role (`relay/lib/relay.mjs`) is for a machine of ours running benches, not for outsiders.
Until the repository is public (the public testnet), results arrive by email to the address on igneum.network; after it,
as an issue with the JSON attached.
## 5. Operator instructions
See `bench/README.md`, shipped in the package. In short: unpack, stop mining on the card, run the one command, send
`results/igneum-repro-<os>-<time>.json` back, say which card was idle. About 4 minutes per card plus the optional
proving step. The file carries no hostname, user name or address; it carries the GPU and CPU models, the OS and the
driver version. `--only backend:index` runs one card; `--no-prove`, `--no-probe`, `--no-sweep` skip steps.
Our own machines run it the same way, with one difference the script does not know about: the card under test is
switched off in the Igneum Miner app first (`bench/jobs/pc-repro.ps1` does it through the app's `api/cards`, one card at
a time, and restores it), and on the Mac the run takes the measure lock.
## 6. The reward terms (proving-e2e.md section 8.1, unchanged)
### 8.1 What counts as unrelated
Three operators are unrelated when every row holds for every pair:
| Test | Requirement |
|---|---|
| Person | Different natural or legal persons; none is the project, an agent of it, or paid by it for the run (a published fixed reproduction reward, equal for everyone and announced before the run, is allowed and disclosed) |
| Hardware | Bought separately; no shared host, card, rack or power meter |
| Network | Different autonomous systems, verified by the IP in the published logs; not the same residential ISP account |
| Location | Different physical sites |
| Software | The same published release by hash; nobody receives a private build |
| Money | No payment, loan or equipment between them or from the project, beyond the disclosed reproduction reward |
An operator declares each row in their report and signs it with the vote key that mined on the devnet under the same fingerprint, so a report is tied to a key with a history.
The amount is the challenge reward row of `docs/plans/funding.md` (USD 1,000 per operator per workload set, three
operators, about USD 3,000 per campaign, approximate; not funded as of 6 October 2026). Reproduction is asked for without
a reward, which the standard allows. The result file of this package is signed by nothing; the signature with the vote
key of 8.1 is the operator's own step, on the file.
## 7. What is unverified
| Item | Why | What closes it |
|---|---|---|
| PC 1's run (RTX 5090, RX 9070 XT, gfx1036 on Windows) | PC 1 went down at 22:31:06 UTC before its slot; the first job hit the parse errors | the morning's slot: `bench/jobs/pc-repro.ps1` as a `run` job after the package fetch, one card at a time |
| The Linux binaries (`bin/linux-x86_64`) | cross-compiled with zig on the Mac, loaded nowhere tonight (no Linux GPU host; HiveOS is the first) | a run on a Linux box with an NVIDIA or AMD driver; `repro.sh` is the same script the Mac ran |
| The CUDA worker's `--bench`, `--memprobe` and `--list` on a real card | tonight's only CUDA runs were the Mac's emulation (`--list`, `--bench` on the genesis pack: self-test PASS, fingerprint `e7d68ec2a49d0671` at 2^14, equal to Apple OpenCL's at 2^14) and PC 1's run that never reached the worker | PC 2's slot (pending) and PC 1's morning run |
| `repro.ps1` end to end | parsed only by the Windows parser on PC 1 after the fix; never run to completion on a PC | PC 2's slot |
| The proving step (step 6) | never run through the script; on PC 2 it will use the live host `/opt/igneum/igneum-prove-host` as the app's WSL user with the live prover off | PC 2's slot; then the memory gate against the measured 28.3 GB peak |
| A second machine of the same model | the tolerance table has one Apple machine; "reproduced externally" needs an unrelated operator's run, and the repository is private until the public testnet | the publish step (`publish-public.sh --repro`, not run tonight) and the first outside result |
| Apple OpenCL 3.7% under the README's 27.9 MH/s | one 30-s wall-time run against one 5-batch run of 4 October; either could be the odd one | a second 120-s run of each path on the Mac with the card to itself |
| The random-read probe on Apple as a ceiling | the hash exceeds the single-chain chase by 4% on the M5 Max, so on Apple the chase at 4 M lanes is not the ceiling the 9070 XT entry took it for | more lanes in flight, or the eight-chain probe (3.50 G here) as the Apple ceiling; a probe question, not a hash one |
| PC 2's 5090 rows | the card-off step could not confirm a stopped worker (`api/state` answered `{}`), and the numbers are half the card's | the morning's 5090-only run on PC 2 with the compute-apps confirmation; until then the rows are "beside the live miner" |
| PC 2's result files | `repro.ps1`'s close threw `System.OutOfMemoryException` in ConvertTo-Json and gave Set-Content an empty path: PowerShell variable names are case-insensitive, the result table `$Cpu` held `$cpu` (its own name string) and so contained itself, and the markdown lines `$md` wiped the markdown path `$Md`. Fixed (distinct names; `bench/jobs/ps-case-check.sh` fails CI on any case-only pair, shown to fire on a bad file); the RESULT lines in the intake and the three transcript logs are the record | the morning's run writes the files |
| The prover on PC 2 | the first script read the prover's state from `api/state` too, saw nothing, and left the live prover off from 00:49 to the restore job (run-prover-on-pc2-20261006); the script now reads settings.json and restores unconditionally | done tonight; the rule in the job |
| The reward | the amount and payer are `docs/plans/funding.md`, not funded; the terms are quoted unchanged | the project lead's decision |
## 8. Files
| File | What |
|---|---|
| `bench/repro.sh`, `bench/repro.ps1` | the operator commands |
| `bench/README.md` | the package's README (operator instructions, tolerances, how to send a result back) |
| `bench/make-package.sh` | builds the package from a tagged commit with the existing cross-build scripts (zig for Linux, mingw for Windows, swiftc and cc for macOS, cargo for the CPU tool) |
| `bench/ingest.mjs`, `bench/reference.json` | the agreement check and the bench-table ingestion; the references |
| `bench/jobs/pc-repro.ps1`, `bench/jobs/collect-results.mjs` | the PC job (one card at a time through the app) and the Mac-side extraction of its result files |
| `docs/benchmarks/repro-2026-10-06/` | the three result files of tonight, their tables and logs |
| `packaging/ota/publish-public.sh --repro` | the public downloads entry (tar.gz, zip, their sha256, the two aliases, the downloads index) |

View file

@ -28,8 +28,8 @@ Versions in the table: `igneum-pow` is the Rust crate at `igneum-pow/Cargo.toml`
| 1 | A new mining program every hour, compiled by the miner, with no human in the loop and no pause in mining | Homepage hero and "This hour's program"; litepaper Mining | tested by the team | igneum-pow 0.2.0; repo `b27da39`, `1292110`, `100c5d7`; fork `devnet-v4` `6457ca95` | The live devnet v4: the node announces `next_epoch_seed` 150 DAA past the seed score, `igneum-miner` sends `prepare` to its worker, the worker builds the next program while the current one mines; Metal (`proto-metal/igneum-bench`), CUDA and OpenCL workers; bench-log "first hourly program swap on the live devnet" | Epoch boundary at DAA 3,600 (11:05:07 BST, 4 October 2026) crossed live on three vendors: the Mac M5 Max (Metal) compiled the next program in 82 ms, 449 DAA before the boundary, swapped in 0.01 ms, 26.7 MH/s before and after; the RTX 5090 compiled in 1,285 ms, swapped in 0.00 ms, 121.8 before and 123.4 MH/s after; the integrated AMD chip 2.74 MH/s before and after. 0 restarts, 0 rejected blocks, 0 rebuilds. The epoch seed is the epoch block hash; the delay of row 2 is not wired in. At a later boundary (DAA 18,000) one OpenCL worker on PC 2 stayed on the previous epoch after an app reinstall and answered 514 jobs with a seed mismatch; fixed in the miner (`3bfe346f`, workers emit a `need` line), the swap time of that forced prepare not measured | none yet |
| 2 | The program seed passes through a 10-minute verifiable delay from a certified checkpoint, so nobody can grind the seed | Litepaper Mining, vs RandomX ("Closed by a verifiable delay") | implemented | repo `792776e`; `proto-vdf/` | `proto-vdf` full 10-minute runs and the tamper cases in `proto-vdf/README.md`; bench-log "proto-vdf" | Class group 1024-bit: 163,000 squarings per second, 10-min eval 585.4 s, prove 9.1 s on 12 threads, verify 4.47 ms, 516-byte proof; wrong checkpoint, flipped seed bit and T+1 all rejected; grinding model gains 0 blocks per epoch with the delay against +3.62 at a 30% advantage without it. 3 October 2026, Apple M5 Max, one core. Prototype only: not in the node on 4 October either, not reviewed against chiavdf (O-4.1) | none yet |
| 3 | The dataset is memory-hard: computing an item costs more than loading it, and every hash does 128 distinct dataset reads | Litepaper Mining and vs RandomX; homepage vs RandomX ("Memory 2 GB, growing") | tested by the team | repo `58a5a63` (memory-hard), `b27da39` (generator 2); igneum-pow 0.2.0 (`memhard.rs`, `generator.rs`, `accept.rs`); spec 01 sections 1.4.2 to 1.4.6; `proto-metal/MEMHARD.md` | `proto-metal/igneum-bench --inline-dataset` against the honest run at 1 GiB and 256 MiB; `igneum-census --gen v2 --warps 64` over 20,000 programs; bench-log "memory-hard dataset" and "generator version 2 adopted" | Honest 45.2 Mhash/s, inline (never reads the dataset) 9.49 Mhash/s, ratio 0.21 at 1 GiB, 0.10 at 256 MiB. 3 October 2026, Apple M5 Max. Generator 2, 4 October 2026: 20,000 programs, 128 static loads on every program, distinct addresses per hash mean 127.9, minimum 120.1; 5.2% of candidates rejected by the acceptance rule. The price of the 128 fresh reads is the hash rate: Apple OpenCL 45.0 MH/s on a version 1 program with 80 distinct loads against 27.5 to 27.9 on version 2; the RTX 5090 229 MH/s on a 104-load version 1 program at 1 GiB (3 October) against 121.8 to 124.2 MH/s mining version 2 on the live devnet (4 October). Apple only for the shortcut ratio (O-1.5); the on-die cache question of ledger M16 is unchanged | none yet |
| 4 | The same program produces identical hashes on three GPU vendors, cache and dataset included | Litepaper vs RandomX ("Bit-exact on Apple, NVIDIA and AMD, measured"), For miners; homepage | tested by the team | repo `f2e903e`, `0f1fdaf` (version 1 packs), `b27da39` (version 2 packs `igneum-genesis-mh`, `igneum-devnet-v4-epoch0`); igneum-pow 0.2.0 | The 96 test vectors of a pack through `proto-metal/igneum-bench`, `proto-cuda/host.cu`, `proto-opencl/host.c`; batch fingerprint at `--batch-log2 24`; the miner's CPU re-check of every share a GPU worker finds on the devnet; bench-log entries "RTX 5090, memory-hard dataset", "AMD gfx1036", "RTX 5090 through NVIDIA OpenCL", "generator version 2 adopted", "the gfx1036 worker fault" | Version 1: 96/96 on Apple Metal (M5 Max), NVIDIA CUDA and NVIDIA OpenCL (RTX 5090, Windows), AMD OpenCL (Ryzen 7 9800X3D integrated gfx1036, 1 compute unit), Apple OpenCL, pocl and two CPU references; batch fingerprint `98af644e993239e2` over 16.7 million nonces identical on the AMD chip and the 5090, 3 October 2026. Version 2: 96/96 on Apple Metal, Apple OpenCL and the CUDA and OpenCL emulators with identical fingerprints; on real NVIDIA and AMD silicon the version 2 vectors have not run as a pack, but both mined accepted blocks on the live devnet with the CPU re-check clean on every share (RTX 5090 at 124.2 MH/s, gfx1036 at 3.3 MH/s), 4 October 2026. The AMD device is an integrated chip; no discrete AMD card and no Intel card has run anything (O-1.15) | none yet |
| 5 | A CPU verifies one hash in under 10 ms by simulating one warp | Litepaper Mining ("about ten milliseconds"), vs RandomX; roadmap gate 2 | tested by the team | repo `75cac18`, `b27da39`; igneum-pow 0.2.0 (`verify.rs`) | `cargo test` and the crate bench in `igneum-pow/`; bench-log "igneum-pow: Rust crate bit-exact with proto-metal" and "generator version 2 adopted" | 0.411 to 0.579 ms per 32-lane warp steady, 0.41 to 0.87 ms cold, average of 20, 1 GiB dataset, cache held, one M5 Max performance core, 3 October 2026; version 2 units 0.631 ms (average of 20), cold 0.67 to 0.81 ms, 4 October 2026. Gate margin about 16x on this core. Not measured on a 2019-class laptop core (O-1.14) | none yet |
| 4 | The same program produces identical hashes on three GPU vendors, cache and dataset included | Litepaper vs RandomX ("Bit-exact on Apple, NVIDIA and AMD, measured"), For miners; homepage | tested by the team | repo `f2e903e`, `0f1fdaf` (version 1 packs), `b27da39` (version 2 packs `igneum-genesis-mh`, `igneum-devnet-v4-epoch0`); igneum-pow 0.2.0 | The 96 test vectors of a pack through `proto-metal/igneum-bench`, `proto-cuda/host.cu`, `proto-opencl/host.c`; batch fingerprint at `--batch-log2 24`; the miner's CPU re-check of every share a GPU worker finds on the devnet; bench-log entries "RTX 5090, memory-hard dataset", "AMD gfx1036", "RTX 5090 through NVIDIA OpenCL", "generator version 2 adopted", "the gfx1036 worker fault" | Version 1: 96/96 on Apple Metal (M5 Max), NVIDIA CUDA and NVIDIA OpenCL (RTX 5090, Windows), AMD OpenCL (Ryzen 7 9800X3D integrated gfx1036, 1 compute unit), Apple OpenCL, pocl and two CPU references; batch fingerprint `98af644e993239e2` over 16.7 million nonces identical on the AMD chip and the 5090, 3 October 2026. Version 2: 96/96 on Apple Metal, Apple OpenCL and the CUDA and OpenCL emulators with identical fingerprints; on real NVIDIA and AMD silicon the version 2 vectors have not run as a pack, but both mined accepted blocks on the live devnet with the CPU re-check clean on every share (RTX 5090 at 124.2 MH/s, gfx1036 at 3.3 MH/s), 4 October 2026. Through the reproducible benchmark package (5 October 2026 night, `docs/benchmarks/repro.md`): the version 2 genesis pack 96/96 on the CPU, on Metal and on Apple OpenCL, with one fingerprint `25f96e7dce90bd4e` over 2^24 nonces on both Apple compilers; the same package's CUDA worker under CPU emulation agrees at 2^14. PC 2's run (6 October 2026, run-repro-pc2-20261006): the version 2 pack 96/96 on the RTX 5090 through NVRTC and through NVIDIA's OpenCL, and on the gfx1036 through AMD's OpenCL, with the same 2^24 fingerprint `25f96e7dce90bd4e` on all three: five compilers, three vendors, one value. PC 1's run (the discrete RX 9070 XT) is owed to the morning. The AMD device is an integrated chip; no discrete AMD card and no Intel card has run anything (O-1.15) | none yet |
| 5 | A CPU verifies one hash in under 10 ms by simulating one warp | Litepaper Mining ("about ten milliseconds"), vs RandomX; roadmap gate 2 | tested by the team | repo `75cac18`, `b27da39`; igneum-pow 0.2.0 (`verify.rs`) | `cargo test` and the crate bench in `igneum-pow/`; bench-log "igneum-pow: Rust crate bit-exact with proto-metal" and "generator version 2 adopted" | 0.411 to 0.579 ms per 32-lane warp steady, 0.41 to 0.87 ms cold, average of 20, 1 GiB dataset, cache held, one M5 Max performance core, 3 October 2026; version 2 units 0.631 ms (average of 20), cold 0.67 to 0.81 ms, 4 October 2026; through the reproducible benchmark package (`bench/repro.sh`, `igneum-pow check-pack`, 5 October 2026 night) 0.611 ms (average of 50), cold 0.625 ms, on the same core. Gate margin about 16x on this core. Not measured on a 2019-class laptop core (O-1.14): the package prints the figure on any machine that runs it | none yet |
| 6 | The hash is bound to the header: one nonce serves one header, and a wrong nonce is rejected | Spec 1.6; litepaper Mining (implied by "checks a hash") | tested by the team | repo `33f7b33`, `9812466`, `b27da39`; igneum-pow 0.2.0 (`bind.rs`, bound vectors re-cut for version 2, 39 crate tests) | `igneum-miner bad-nonce` against a devnet node; `igneum-pow hash-bound` for the 96-nonce job across the 2^32 lane boundary; bench-log "first devnet blocks on the real lottery hash" and "generator version 2 adopted" | 833 blocks accepted by `igneum-lottery-v1-bound` on 3 nodes, 0 rejections; `bad-nonce` gave Reject(BlockInvalid); Metal, OpenCL and CUDA (emulated) workers bit-exact with the crate on the lane-boundary job, 3 October 2026, Apple M5 Max. Version 2: the node's engine reports `igneum-lottery-v2-bound`, 39 of 39 crate tests, and the live devnet v4 accepts its blocks under it, 4 October 2026 | none yet |
| 7 | The devnet runs at one block a second | Homepage stats ("1 / s"); litepaper Speed; roadmap phase 3 | tested by the team | repo `9812466`, `e9328c6`, `8dae48b`; fork `devnet-v4` `dc749905` | The merged node's 3-node test network (`igneum-devnet-880`, 960 s); the live devnet v4 record `sim/difficulty/records/live-2026-10-04.csv`; the 12-node cloud network's arrival logs; bench-log "devnet-v4 integration", "difficulty rule v2", "first devnet blocks" | Merged node, 4 October 2026, Apple M5 Max: 1,055 blocks in 960 s, 1.03 blocks/s, sink identical on 3 nodes at 31 of 31 samples, 0 rejected. Live devnet v4 the same day: 49 to 81 blocks a minute while two RTX 5090s joined and left (row 12), 1.1 to 1.2 blocks/s in the oscillating window, then within 1.3% per minute with one PC and the Mac. The 12-node cloud network at one block a second: 644 blocks in a 10-minute window. The 3 October CPU devnet: 1.29 blocks/s over 641 s, 1.03 after the first retarget. The phase 3 gate also asks for proofs under 60 s behind the tip; no proof is on the chain (row 15) | none yet |
| 8 | Blocks are mined by GPUs on Apple and NVIDIA | Homepage live strip; journey phase 3 ("GPU miners on three vendors") | tested by the team | repo `9812466`, `e9328c6`, `d7e1f89`, `2309c8d`; fork `devnet-v4` | Metal worker `proto-metal/igneum-bench --serve` driven by `igneum-miner --worker`; the live devnet v4 hash-rate record `sim/difficulty/records/live-2026-10-04-hashrate.csv` (587 worker STATUS lines by run id); bench-log "first devnet blocks", "devnet v4 cut-over", "difficulty rule v2", "first machine on the Igneum Miner app" | Metal: 506 jobs, 5,636 blocks found and accepted, 0 rejected, 0 CPU/GPU mismatches, 28.2 MH/s wall, 3 October 2026. Live devnet v4, 4 October 2026: PC 1's RTX 5090 at 122 MH/s with 8 identities, PC 2's at 124 MH/s with 8 identities (117 to 119 MH/s inside the one-click app, 34 accepted blocks in its first minute, CPU re-check OK on every share), the Mac's Metal worker at 26.7 MH/s; 17 vote keys signed the first finality lock (row 10); from the afternoon an Apple silicon laptop outside the project at 21.0 MH/s through the app (row 30). Two RTX 5090s and two Apple chips; no other NVIDIA model has mined | none yet |
@ -39,7 +39,7 @@ Versions in the table: `igneum-pow` is the Rust crate at `igneum-pow/Cargo.toml`
| 12 | The difficulty rule recovers from a hashrate step within minutes, where Kaspa's sampled rule never settles. A step inside an epoch set the rule oscillating on the live devnet on 4 October 2026; rule v2 removes it in the simulator and on a test network and is built but not yet rolled out | Spec 2.3; litepaper Speed (implied); bench page | tested by the team | repo `e9328c6`, `abb5a5d` (attacks), `67bf226` (rule v2); fork `difficulty` branch (timestamp fix) and `devnet-v4` `a21ff239` (`difficulty_v2_activation_daa`, `REF_WINDOW_V2 = 600`); `sim/difficulty/sim.py --live` | The live record `sim/difficulty/records/live-2026-10-04.csv` (8,090 headers, `pull_live.py`) and the hash-rate record beside it; `sim/difficulty/sim.py` on the synthetic set and the DAG replay; `sim/difficulty/attacks/attacks.py`; `sim/difficulty/testnet_v2.py` (3 nodes, activation at DAA 900); `cargo test --release -p kaspa-consensus --lib difficulty` (15 pass); bench-log "difficulty controller", "difficulty rule under attack", "timestamp attack fixed", "difficulty rule v2" | Live devnet v4, 4 October 2026 (UTC): a second RTX 5090 joining 7 minutes into an epoch (about 152 to 280 MH/s) hardened the difficulty 70M to 144M in 90 s and then swung by about a third for 40 minutes around the true level of 139M while the epoch-long reference lane carried the join; that card leaving for 4 minutes eased 116M to 67M and back to 106M; the epoch boundary with both PCs restarting took 152M to 77M in 3 minutes, after which the rule held within 1.3% per minute with no flips. Cause: the reference lane covered the whole epoch, so a mid-epoch step polluted it for the hour and the 25% trigger flipped on the short lane's noise. The DAG replay reproduces the record (std of log difficulty 0.115 against 0.134, 4.3 peaks against 4). Rule v2 (reference window 600 DAA) on the replay: std 0.026, 0 flips, mean 142.6M against 139M true; on a 3-node test network the v2 nodes eased a leave with no peak and held a rejoin within 3% after 60 s, and a node without the activation height forked off at it as designed. Rule v2 rolled onto the 12-node cloud network on 4 October (all nodes crossed the height on one chain; a hash-rate step then settled in 160 to 270 s with no swing) and activates on the devnet at DAA 33,000 the same evening. Timestamp forging (ledger M23) fixed the same day: a 50% forger drifts the rate under 1.1% where the 3 October rule gave it a 9.9x difficulty. Simulator, settled seconds: x50 step 62 to 66 (Kaspa 1,542), /50 step 657 to 753 (Kaspa 12,296). Apple M5 Max under load 7 to 442; the DAG model is fitted on one scale; the pool hopper's 0.7-point excess over Kaspa's rule stays open | none yet |
| 13 | Every node executes the ordered transactions natively and reaches the same state root | Litepaper Proving ("Every node executes ... natively"), Building ("runs on Igneum unchanged") | tested by the team | repo `f5f8c80`, `8dae48b`; fork `devnet-v4` `dc749905`; revm 43.0.3 | `node tools/evm-smoke/smoke.mjs` against a 3-node `igneumd`; `igneum-exec-diff seq.json`; bench-log "execution layer devnet v3" and "devnet-v4 integration" | Simnet, 3 October 2026: 87 of 87 viem checks, state roots identical on 3 nodes at four heights, 57 executed and 19 skipped transactions agree with plain revm, 0 mismatches. Merged node on real proof of work, 4 October 2026: 84 of 85 checks (the miss needs parallel blocks the network did not produce in 36 s), 59 transfers in 10 chain blocks, state roots identical on 3 nodes, `igneum-exec-diff` 0 mismatches over 59 transactions; the live devnet v4 runs this execution layer. Apple M5 Max. The prover is a stub; state is rebuilt from genesis at start; no EVM transaction relay between nodes | none yet |
| 14 | Ethereum bytecode runs unchanged, with the documented differences of spec 7.1 | Homepage Build card; litepaper Building | tested by the team | as row 13; fixes `F-exec-A`, `F-exec-B` (spec 7.5) | `tools/evm-smoke/smoke.mjs`: deploy via viem, `increment`, `hashLoop`, `eth_estimateGas`, `eth_getLogs`; `tools/exec-attacks` scenarios 1 and 3; bench-log "execution layer attack fixes" | Deployment, calls, reverts, logs and gas estimates behave as viem expects; chain id 4463; the prototype pgas table gives 0.0095 to 0.028 pgas per gas, below the design's band before calibration, 3 October 2026. 4 October 2026: a transaction that would cross the block's proving budget is refused by the mempool and, if forced in, aborted and charged with its nonce advanced (25 of 25 checks; 30 of 30 malformed cases). Apple M5 Max. The `Prover` precompile, proof records and the shard planner are not in the node | none yet |
| 15 | Every block is proven, with the proof landing within about a minute at launch | Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate | implemented | repo `d7e1f89` (GPU proof), `e01a3cc`, `292e800`, `eedd136` (`proving/igneum-prove`: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6 | `proving/windows-wsl2` (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; `igneum-prove-host --mode block` on `proving/fixtures/`; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards" | First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture `block-78-increment` (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in `docs/benchmarks/proving-e2e.md`. Second RTX 5090 run, 4 October 2026 evening (job run-20261004-173115): a full shard at the provisional S_p (6.75 M pgas, 60.8 M cycles) executed in 1.63 s, core proof 8.3 s (18.1 MB), compressed proof 10.9 s (1.27 MB, verify 0.040 s); a two-shard block (13.5 M pgas) proved shard by shard (11.7 s and 10.0 s) and aggregated in 2.2 s, 24 s of GPU stages end to end, every proof verified, six tampered witnesses rejected. The two host defects (an abort after the upload, an idle wait that turned out to be an unbuffered 18 MB proof save through the WSL2 file bridge, 24 minutes) are fixed (ledger P20) 5 October 2026, live devnet with real transactions (bench-log "real transactions, the first non-empty shard proven and paid"): block 72704 shard 0, 29 transfers, 5,800 pgas, proven on PC 2 in 34 s, verified on the Mac in 0.297 s and paid 1.7623 IGN, 53 s after the chain block executed; of about 1,400 blocks in the 20-minute window 36 were proven (the one prover takes the newest shard assigned to it), so "every block" is not yet true; a second content shard (72803, all copies skipped) failed the native-execution veto on the exporter's block structure, fixed with fixtures the same day, the node side pending the 0.3.9 rollout | none yet |
| 15 | Every block is proven, with the proof landing within about a minute at launch | Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate | implemented | repo `d7e1f89` (GPU proof), `e01a3cc`, `292e800`, `eedd136` (`proving/igneum-prove`: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6 | `proving/windows-wsl2` (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; `igneum-prove-host --mode block` on `proving/fixtures/`; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards" | First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture `block-78-increment` (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in `docs/benchmarks/proving-e2e.md`. Second RTX 5090 run, 4 October 2026 evening (job run-20261004-173115): a full shard at the provisional S_p (6.75 M pgas, 60.8 M cycles) executed in 1.63 s, core proof 8.3 s (18.1 MB), compressed proof 10.9 s (1.27 MB, verify 0.040 s); a two-shard block (13.5 M pgas) proved shard by shard (11.7 s and 10.0 s) and aggregated in 2.2 s, 24 s of GPU stages end to end, every proof verified, six tampered witnesses rejected. The two host defects (an abort after the upload, an idle wait that turned out to be an unbuffered 18 MB proof save through the WSL2 file bridge, 24 minutes) are fixed (ledger P20). Through the reproducible benchmark package's optional step 6 (6 October 2026, run-repro-pc2-20261006, the live host on PC 2's RTX 5090 beside the app's live miner): shard 0 of `block-338-shard1` executed in 1.48 s, core proof 20.2 s, compressed 33.9 s, verified in 0.038 s with the pinned key, two tampered witnesses rejected; the same command an outsider with a prover host runs 5 October 2026, live devnet with real transactions (bench-log "real transactions, the first non-empty shard proven and paid"): block 72704 shard 0, 29 transfers, 5,800 pgas, proven on PC 2 in 34 s, verified on the Mac in 0.297 s and paid 1.7623 IGN, 53 s after the chain block executed; of about 1,400 blocks in the 20-minute window 36 were proven (the one prover takes the newest shard assigned to it), so "every block" is not yet true; a second content shard (72803, all copies skipped) failed the native-execution veto on the exporter's block structure, fixed with fixtures the same day, the node side pending the 0.3.9 rollout | none yet |
| 16 | A 12 GB card proves one shard in about 20 s | Litepaper Proving ("The proving budget"); roadmap gate 2 | designed | spec 5.1 (Target), 7.6 (`S_p` provisional, 7,500,000 pgas = `B_p` / 4) | `PROVE-SHARD.bat` on the RTX 5090 (pending); the end-to-end standard in `docs/benchmarks/proving-e2e.md`; bench-log "proving: devnet v4 shards" | Measured on a 32 GB card, not yet on a 12 GB card. A shard at the provisional `S_p` is 60.8 M SP1 cycles on the prototype pgas table (9 cycles per pgas, 44 per EVM gas; the modexp entry about 100x its SP1 cost); on an RTX 5090 (4 October 2026 evening, job run-20261004-173115) it executed in 1.63 s and its compressed proof took 10.9 s, verified in 0.040 s, so the 32 GB card is inside the 20 s target with margin. Whether a 12 GB card proves it at all, and in what time, is the next measurement (an RTX 3060 and an RTX 5060 Ti 16 GB are on order). A per-shard time can be met by shrinking the shard, so the project does not use it as a pass mark | none yet |
| 17 | The chip resistance target: a chip gains under 2x over a GPU | Litepaper Mining, "What Igneum does not claim"; homepage "no chip can be built for it" | designed | spec 0.2 (Target); O-1.17 | Public benchmark with a leaderboard by card model and a standing bounty, January 2027 (O-1.17); the on-die-SRAM test on the RTX 5090 (R3.5) | A target, not a measurement. Review round 3 priced a recompute chip with the 256 MiB cache on die at about 2.4x, approximate, before the usual chip-versus-GPU integer gain; the design answer (cache larger than any die) is open (spec 1.16) | none yet |
| 18 | The chip resistance measurements: the program is random-access bound, not bandwidth bound, and sits beyond a card's on-chip cache | Litepaper Mining ("bound by memory bandwidth", to be corrected), vs RandomX "Measured so far" | tested by the team | repo `aba248d`, `f2a1a64`, `4b95c5e` | RTX 5090 dataset sweep 4 MiB to 1 GiB with `proto-cuda/host.cu`; bench-log "RTX 5090 first run" and "dataset sweep" | At 1 GiB: 228.1 Mhash/s, 23.7 G random loads/s, 94.9 GB/s useful against a 1,638 GB/s dataset fill; inside the 96 MiB L2 (4 and 64 MiB) 1,340 to 1,353 Mhash/s, about 5.8x faster; 104 against 128 loads per hash gives 228 against 185 Mhash/s, proportional. 3 October 2026, RTX 5090, Windows, CUDA 12.8, version 1 programs. Prototype dataset 1 GiB against 2 GB at genesis; a pure random-read microbenchmark (R3 chip designer, attack 2) has not run; the sweep has not been repeated on version 2 | none yet |
@ -94,6 +94,6 @@ Versions in the table: `igneum-pow` is the Rust crate at `igneum-pow/Cargo.toml`
|---|---|---|
| designed | implemented | Code in this repository with test vectors that pass |
| implemented | tested by the team | A bench-log entry with the machine, the date, the command and the number |
| tested by the team | reproduced externally | The repository public (at the public testnet), the command published, and a third party's run with the same result, linked from the row |
| tested by the team | reproduced externally | The repository public (at the public testnet), the command published, and a third party's run with the same result, linked from the row. The command exists since 5 October 2026: the reproducible benchmark package (`bench/`, `docs/benchmarks/repro.md`; `igneum-repro-<tag>.tar.gz` and `.zip` next to the public downloads once published), one command per platform, with the tolerances per number and `bench/ingest.mjs` to turn an agreeing result into a bench-table row labelled "reproduced externally". Rows 3, 4, 5, 18 and 19 are the ones it covers today; row 15 and 16 through its optional proving step |
| reproduced externally | reviewed independently | A named reviewer's published finding on that version. Funding for review is `docs/plans/funding.md` |
| any | the row's status falls back | A new version of the code or rule the row names |

View file

@ -7,6 +7,10 @@
//! [--epoch-hex <64 hex> --day-hex <hex>] byte seeds instead of strings (Epoch::from_seed_bytes)
//! igneum-pow accept --seed <s> [--epoch-hex <64 hex>] every candidate of the seed with its verdict (spec 01 section 1.4.6)
//! igneum-pow show --seed <s> [--epoch-hex <64 hex>] the accepted program, one instruction per line
//! igneum-pow check-pack --pack <dir> [--warps 20] the CPU side of the reproducible benchmark (bench/repro.sh, 6 October
//! 2026): rebuild the pack's program and dataset from program.json, recompute every vector warp of
//! vectors.json and the cache FNV, print PASS or FAIL per warp, the program id and the CPU verify
//! time per warp, and one RESULT line. Exit 0 only when every vector is bit-exact.
use igneum_pow::emit::export_pack;
use igneum_pow::memhard::Cache;
@ -26,6 +30,7 @@ struct Args {
prehash: String,
epoch_hex: Option<String>,
day_hex: Option<String>,
pack: Option<String>,
}
fn usage() -> ! {
@ -36,7 +41,8 @@ fn usage() -> ! {
\x20 hash --nonce <n> print the 64-bit hash of one nonce (pack form, init words = seed words)\n\
\x20 hash-bound --prehash <64 hex> --nonce <u64> print the header-bound hash (bind.rs) of one 64-bit nonce\n\
\x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\
\x20 show the accepted program, one instruction per line"
\x20 show the accepted program, one instruction per line\n\
\x20 check-pack --pack <dir> [--warps 20] recompute the pack's vectors.json on the CPU, PASS or FAIL per warp, one RESULT line"
);
std::process::exit(2)
}
@ -54,6 +60,7 @@ fn parse() -> Args {
prehash: "00".repeat(32),
epoch_hex: None,
day_hex: None,
pack: None,
};
let mut it = std::env::args().skip(1);
a.cmd = it.next().unwrap_or_else(|| usage());
@ -70,6 +77,7 @@ fn parse() -> Args {
"--prehash" => a.prehash = val(),
"--epoch-hex" => a.epoch_hex = Some(val()),
"--day-hex" => a.day_hex = Some(val()),
"--pack" => a.pack = Some(val()),
_ => usage(),
}
}
@ -84,6 +92,7 @@ fn main() {
"export" => export(&a, mode),
"accept" => accept(&a),
"show" => show(&a),
"check-pack" => check_pack(&a),
"hash" => {
let e = Epoch::new(&a.seed, &a.day, mode, a.dataset_log2);
println!("{:016x}", e.hash(a.nonce as u32));
@ -258,3 +267,119 @@ fn show(a: &Args) {
);
}
}
// ---- check-pack: the CPU side of the reproducible benchmark (bench/, 6 October 2026) ---------------------------------
//
// No JSON crate (the crate has no dependency outside std): the few fields needed are read with a scanner that finds
// `"name":` and takes the string or number after it. program.json: seed_bytes, dataset.day_bytes (the first
// "day_bytes"), dataset_mode, the first "log2_words" (the dataset's; the cache's comes later), program_id.
// vectors.json: every "base_nonce" with the 32 "expected" strings after it, and cache_fnv1a64.
fn json_after<'a>(text: &'a str, key: &str, from: usize) -> Option<(usize, &'a str)> {
let pat = format!("\"{key}\":");
let at = text[from..].find(&pat)? + from + pat.len();
let rest = text[at..].trim_start();
let off = text.len() - rest.len();
if let Some(r) = rest.strip_prefix('"') {
let end = r.find('"')?;
Some((off + 1 + end + 1, &r[..end]))
} else {
let end = rest.find(|c: char| c == ',' || c == '}' || c == ']' || c.is_whitespace()).unwrap_or(rest.len());
Some((off + end, &rest[..end]))
}
}
fn json_string_array<'a>(text: &'a str, key: &str, from: usize) -> Option<(usize, Vec<&'a str>)> {
let pat = format!("\"{key}\":");
let at = text[from..].find(&pat)? + from + pat.len();
let open = text[at..].find('[')? + at;
let close = text[open..].find(']')? + open;
let items = text[open + 1..close].split(',').map(|s| s.trim().trim_matches('"')).filter(|s| !s.is_empty()).collect();
Some((close + 1, items))
}
fn check_pack(a: &Args) {
let dir = a.pack.clone().unwrap_or_else(|| usage());
let read = |name: &str| -> String {
std::fs::read_to_string(format!("{dir}/{name}")).unwrap_or_else(|e| {
eprintln!("FAIL: cannot read {dir}/{name}: {e}");
std::process::exit(2)
})
};
let pj = read("program.json");
let vj = read("vectors.json");
let field = |key: &str| json_after(&pj, key, 0).map(|(_, v)| v.to_string()).unwrap_or_else(|| { eprintln!("FAIL: program.json has no {key}"); std::process::exit(2) });
let seed_label = field("seed");
let seed_bytes = igneum_pow::bind::unhex(&field("seed_bytes")).unwrap_or_else(|| usage());
let day_bytes = igneum_pow::bind::unhex(&field("day_bytes")).unwrap_or_else(|| usage());
let mode = if field("dataset_mode") == "memory-hard" { DatasetMode::MemoryHard } else { DatasetMode::ClosedForm };
let log2: u32 = field("log2_words").parse().unwrap_or_else(|_| usage());
let want_id = u64::from_str_radix(field("program_id").trim_start_matches("0x"), 16).unwrap_or(0);
let t0 = Instant::now();
let program = igneum_pow::generator::generate_from_seed_bytes(&seed_label, &seed_bytes);
let key = igneum_pow::seed::seed_words_from_bytes(&day_bytes);
let mut dataset = igneum_pow::verify::DatasetSource::from_key(key, mode, log2);
dataset.key_bytes = day_bytes.clone();
let e = Epoch { program, dataset };
let build_ms = t0.elapsed().as_secs_f64() * 1e3;
let id_ok = e.program.program_id() == want_id;
println!(
"check-pack {dir}: seed \"{}\" generator v{} attempt {} program id {:016x} ({}), dataset 2^{log2} words ({}), {} loads/hash; built in {build_ms:.0} ms",
e.program.seed_string, e.program.generator, e.program.attempt, e.program.program_id(),
if id_ok { "matches program.json" } else { "DOES NOT MATCH program.json" }, mode.name(), e.program.loads_per_hash()
);
let mut cache_ok = true;
let mut cache_fnv = 0u64;
if mode == DatasetMode::MemoryHard {
if let Some((_, want)) = json_after(&vj, "cache_fnv1a64", 0) {
let want = u64::from_str_radix(want.trim_start_matches("0x"), 16).unwrap_or(0);
cache_fnv = e.dataset.memhard().map(|m| m.cache.fnv1a64()).unwrap_or(0);
cache_ok = cache_fnv == want;
println!("cache: FNV-1a 64 {cache_fnv:016x} {} vectors.json ({want:016x})", if cache_ok { "==" } else { "!=" });
}
}
// every vector warp of vectors.json
let mut pos = 0usize;
let mut warps = 0usize;
let mut lanes_ok = 0usize;
let mut lanes = 0usize;
let mut warps_ok = 0usize;
let mut single_ms = Vec::new();
while let Some((p1, base)) = json_after(&vj, "base_nonce", pos) {
let base: u32 = base.parse().unwrap_or_else(|_| usage());
let (p2, expected) = json_string_array(&vj, "expected", p1).unwrap_or_else(|| usage());
pos = p2;
let t = Instant::now();
let got = e.hash_warp(base);
single_ms.push(t.elapsed().as_secs_f64() * 1e3);
let mut bad = Vec::new();
for (l, want) in expected.iter().enumerate().take(32) {
let w = u64::from_str_radix(want.trim_start_matches("0x"), 16).unwrap_or(1);
lanes += 1;
if got[l] == w { lanes_ok += 1; } else { bad.push(l); }
}
warps += 1;
if bad.is_empty() { warps_ok += 1; }
println!(
"vector warp base {base} (nonces {base}..{}): {} ({} of 32 lanes){}",
base + 31, if bad.is_empty() { "PASS" } else { "FAIL" }, 32 - bad.len(),
if bad.is_empty() { String::new() } else { format!(" mismatched lanes {bad:?}, lane {} cpu {:016x} vectors {}", bad[0], got[bad[0]], expected[bad[0]]) }
);
}
// the CPU verify time per warp, steady state (the gate: under 10 ms on one core)
let n = a.warps.max(1);
let t = Instant::now();
let mut sink = 0u64;
for i in 0..n {
sink ^= e.hash_warp((i as u32) * 32 + 65536)[0];
}
let avg = t.elapsed().as_secs_f64() * 1e3 / n as f64;
let cold = single_ms.iter().cloned().fold(0.0, f64::max);
println!("CPU verify: {avg:.3} ms per 32-lane warp, avg of {n} (cold {cold:.3} ms max over the vector warps; checksum {sink:016x})");
let pass = id_ok && cache_ok && warps > 0 && warps_ok == warps;
println!(
"RESULT check-pack pack={dir} seed={} program_id={:016x} id_match={} dataset_log2={log2} mode={} cache_fnv={cache_fnv:016x} cache_match={} warps={warps} warps_pass={warps_ok} lanes={lanes} lanes_pass={lanes_ok} cpu_verify_ms={avg:.3} cpu_verify_cold_ms={cold:.3} verdict={}",
e.program.seed_string, e.program.program_id(), id_ok, mode.name(), cache_ok, if pass { "PASS" } else { "FAIL" }
);
std::process::exit(if pass { 0 } else { 1 });
}

View file

@ -7,6 +7,8 @@
# https://dl.igneum.network/public/igneum-miner-mac.dmg -> dl/public/Igneum-Miner-<v>.dmg
# https://dl.igneum.network/public/igneum-miner-hive.tar.gz -> dl/public/igneum-hive-<v>.tar.gz
# https://dl.igneum.network/public/igneum-wallet-mac.dmg -> dl/public/Igneum-Wallet-<v>.dmg
# https://dl.igneum.network/public/igneum-repro.tar.gz -> dl/public/igneum-repro-<tag>.tar.gz (the reproducible benchmark, bench/)
# https://dl.igneum.network/public/igneum-repro.zip -> dl/public/igneum-repro-<tag>.zip (the same files for Windows)
#
# The aliases are rewrites in <dlsite>/vercel.json, written by this script from what dl/public/ holds, so they can
# never point at a file that is not there and they follow every release by themselves (the bytes exist once).
@ -16,6 +18,8 @@
# dl/public/igneum-app-latest.json(.sig) with public URLs, signed
# packaging/ota/publish-public.sh --wallet the same for igneum-wallet-latest.json
# packaging/ota/publish-public.sh --hive <igneum-hive-<v>.tar.gz> copy the HiveOS package (plus its .sha256)
# packaging/ota/publish-public.sh --repro <igneum-repro-<tag>.tar.gz> copy the reproducible benchmark package: the tar.gz, the
# zip next to it (bench/make-package.sh writes both) and their .sha256
# packaging/ota/publish-public.sh --aliases only rewrite vercel.json from the folder (runs after every step above)
# --dry-run read everything, print the plan, write nothing --deploy deploy the downloads folder afterwards
# --verify HEAD every public file and alias, GET the manifests (also after --deploy) --no-prune keep old files
@ -32,12 +36,13 @@ KEY="$CFG/ota-signing-key"; PUB_KEY="$CFG/ota-signing-key.pub"
SIGNER="${IGNEUM_OTA_SIGN:-$ROOT/app/igneum-app/target/release/igneum-ota-sign}" # a built signer elsewhere (another worktree)
HOST="https://dl.igneum.network"
DO_APP=0 DO_WALLET=0 HIVE="" DO_ALIASES=0 DRY=0 DEPLOY=0 VERIFY=0 PRUNE=1 DEST="" BASE="" TRIES=12
DO_APP=0 DO_WALLET=0 HIVE="" REPRO="" DO_ALIASES=0 DRY=0 DEPLOY=0 VERIFY=0 PRUNE=1 DEST="" BASE="" TRIES=12
while [ $# -gt 0 ]; do
case "$1" in
--app) DO_APP=1; shift ;;
--wallet) DO_WALLET=1; shift ;;
--hive) HIVE="$2"; shift 2 ;;
--repro) REPRO="$2"; shift 2 ;;
--aliases) DO_ALIASES=1; shift ;;
--dry-run) DRY=1; shift ;;
--deploy) DEPLOY=1; shift ;;
@ -49,7 +54,7 @@ while [ $# -gt 0 ]; do
*) echo "unknown argument: $1" >&2; exit 2 ;;
esac
done
[ "$DO_APP$DO_WALLET$DO_ALIASES$VERIFY" != 0000 ] || [ -n "$HIVE" ] || { echo "nothing to do: --app, --wallet, --hive <file>, --aliases or --verify" >&2; exit 2; }
[ "$DO_APP$DO_WALLET$DO_ALIASES$VERIFY" != 0000 ] || [ -n "$HIVE" ] || [ -n "$REPRO" ] || { echo "nothing to do: --app, --wallet, --hive <file>, --repro <file>, --aliases or --verify" >&2; exit 2; }
command -v python3 >/dev/null || { echo "python3 is needed" >&2; exit 1; }
[ -x "$SIGNER" ] || { echo "no $SIGNER (cargo build --release --bin igneum-ota-sign in app/igneum-app)" >&2; exit 1; }
@ -119,7 +124,22 @@ if [ -n "$HIVE" ]; then
else cp "$HIVE" "$PUB/$hn.part" && mv "$PUB/$hn.part" "$PUB/$hn"; "$SIGNER" sha256 "$PUB/$hn" | awk -v n="$hn" '{print $1 " " n}' > "$PUB/$hn.sha256"; log " $hn: copied into dl/public/ ($(stat -f %z "$PUB/$hn") B, sha256 $(sha "$PUB/$hn"))"; fi
fi
# the aliases: what the two manifests and the newest HiveOS package in dl/public/ name, nothing else
if [ -n "$REPRO" ]; then
[ -f "$REPRO" ] || { echo "missing: $REPRO" >&2; exit 1; }
rn="$(basename "$REPRO")"
case "$rn" in igneum-repro-*.tar.gz) ;; *) echo "$REPRO: the package is igneum-repro-<tag>.tar.gz (bench/make-package.sh)" >&2; exit 1 ;; esac
rz="${REPRO%.tar.gz}.zip"
[ -f "$rz" ] || { echo "no $rz next to the tar.gz (make-package.sh writes both)" >&2; exit 1; }
log "reproducible benchmark package -> dl/public/"
for f in "$REPRO" "$rz"; do
fn="$(basename "$f")"
if [ -f "$PUB/$fn" ] && [ "$(sha "$PUB/$fn")" = "$(sha "$f")" ]; then log " $fn: already in dl/public/ (same sha256)"
elif [ "$DRY" = 1 ]; then log " $fn: would copy into dl/public/ ($(stat -f %z "$f") B, sha256 $(sha "$f"))"
else cp "$f" "$PUB/$fn.part" && mv "$PUB/$fn.part" "$PUB/$fn"; "$SIGNER" sha256 "$PUB/$fn" | awk -v n="$fn" '{print $1 " " n}' > "$PUB/$fn.sha256"; log " $fn: copied into dl/public/ ($(stat -f %z "$PUB/$fn") B, sha256 $(sha "$PUB/$fn"))"; fi
done
fi
# the aliases: what the two manifests, the newest HiveOS package and the newest repro package in dl/public/ name, nothing else
ALIASES_JSON="$(python3 - "$PUB" "$DRY" <<'PY'
import json, os, re, sys
pub, dry = sys.argv[1], sys.argv[2] == "1"
@ -132,11 +152,16 @@ def entry(name, plat):
return f if os.path.exists(os.path.join(pub, f)) or dry else None
hive = sorted([f for f in os.listdir(pub) if re.fullmatch(r"igneum-hive-\d+\.\d+\.\d+\.tar\.gz", f)],
key=lambda f: tuple(int(x) for x in re.findall(r"\d+", f)[:3])) if os.path.isdir(pub) else []
# the repro package: the newest by the version numbers in its tag (igneum-repro-v0.1.0-repro.tar.gz), the zip beside it
repro = sorted([f for f in os.listdir(pub) if re.fullmatch(r"igneum-repro-.+\.tar\.gz", f)],
key=lambda f: tuple(int(x) for x in re.findall(r"\d+", f)[:3])) if os.path.isdir(pub) else []
aliases = {
"igneum-miner-windows.exe": entry("igneum-app-latest.json", "windows"),
"igneum-miner-mac.dmg": entry("igneum-app-latest.json", "mac"),
"igneum-miner-hive.tar.gz": hive[-1] if hive else None,
"igneum-wallet-mac.dmg": entry("igneum-wallet-latest.json", "mac"),
"igneum-repro.tar.gz": repro[-1] if repro else None,
"igneum-repro.zip": (repro[-1][:-7] + ".zip") if repro and os.path.exists(os.path.join(pub, repro[-1][:-7] + ".zip")) else None,
}
print(json.dumps(aliases))
PY
@ -177,9 +202,10 @@ def sha(p):
return h.hexdigest()
def version_of(name, f):
if name.startswith("igneum-miner-hive"): return ".".join(re.findall(r"\d+", f)[:3])
if name.startswith("igneum-repro"): return re.sub(r"^igneum-repro-|\.tar\.gz$|\.zip$", "", f)
m = json.load(open(os.path.join(pub, "igneum-wallet-latest.json" if name.startswith("igneum-wallet") else "igneum-app-latest.json")))
return m.get("version")
keys = {"igneum-miner-windows.exe": "miner-windows", "igneum-miner-mac.dmg": "miner-mac", "igneum-miner-hive.tar.gz": "miner-hive", "igneum-wallet-mac.dmg": "wallet-mac"}
keys = {"igneum-miner-windows.exe": "miner-windows", "igneum-miner-mac.dmg": "miner-mac", "igneum-miner-hive.tar.gz": "miner-hive", "igneum-wallet-mac.dmg": "wallet-mac", "igneum-repro.tar.gz": "repro-tar", "igneum-repro.zip": "repro-zip"}
files = {}
for alias, f in aliases.items():
if not f: continue

View file

@ -274,13 +274,17 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) {
strcpy(ehex, e2); strcpy(dhex, d2);
}
if (!ehex[0] || !dhex[0]) return pf_fail(err, cap, "no seeds: neither seeds.txt nor IGNEUM_SEED_BYTES_HEX / IGNEUM_DAY_BYTES_HEX in program.h (a pack from igneum-pow export --seed <name> has no byte seeds)");
if (strlen(ehex) != 64) return pf_fail(err, cap, "epoch seed is not 64 hex characters");
// The chain's epoch seed is 32 bytes (64 hex characters). A pack exported from a seed STRING (igneum-pow export
// --seed <name>, the read-width experiment's packs of 5 October 2026) carries the string's bytes instead; the
// seed-word re-derivation below checks either form, so any even-length hex seed is accepted here. The serve
// protocol still carries 64-hex seeds; a string-seed pack can only be benched (--bench, --bench-pack, --check).
if (strlen(ehex) < 2 || strlen(ehex) % 2 != 0) return pf_fail(err, cap, "epoch seed is not an even-length hex string");
strcpy(pk->epochHex, ehex); strcpy(pk->dayHex, dhex);
// The seed words derived from the bytes must be the pack's own words: otherwise the pack and the seeds disagree
{
uint32_t w[8];
if (!pf_unhex(ehex, bytes, 32, &blen) || blen != 32) return pf_fail(err, cap, "epoch seed hex is malformed");
pf_seed_words_from_bytes(bytes, 32, w);
if (!pf_unhex(ehex, bytes, sizeof(bytes), &blen) || blen == 0) return pf_fail(err, cap, "epoch seed hex is malformed");
pf_seed_words_from_bytes(bytes, blen, w);
if (memcmp(w, pk->seedw, 32) != 0) return pf_fail(err, cap, "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT (wrong seeds.txt for this pack?)");
if (!pf_unhex(dhex, bytes, sizeof(bytes), &blen)) return pf_fail(err, cap, "day seed hex is malformed");
pf_seed_words_from_bytes(bytes, blen, w);

View file

@ -53,6 +53,7 @@
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <cctype>
#include <chrono>
#include <string>
#include <vector>
@ -255,6 +256,10 @@ struct Ctx {
int raceBenchMs = 2000, raceBudgetS = 120, raceRounds = 1, batchLog2 = 22;
std::string pinned; // --variant: use this variant, no race
std::string tuning; // the tuning file's text ("" = none)
// the reproducible benchmark (bench/, 6 October 2026): --bench times dispatches for --seconds (or --batches),
// --dataset-mib builds the pack's program over another dataset size (the self-test is then skipped: the vectors
// are for the pack size), as proto-cuda/host.cu --dataset-mib and --sweep do for the ahead-of-time harness
int batches = 5, seconds = 0, datasetLog2Override = 0;
std::string err(CUresult r) { const char* s = nullptr; if (drv.getErrorString) drv.getErrorString(r, &s); return s ? s : "CUDA driver error"; }
};
@ -751,6 +756,8 @@ static Pair* buildPair(Ctx& c, const std::string& dir, CUstream s, std::string&
p->dir = dir; p->epochHex = pk.epochHex; p->dayHex = pk.dayHex; p->seedString = pk.seedString;
std::memcpy(p->sw, pk.seedw, 32); std::memcpy(p->kw, pk.keyw, 32);
p->datasetLog2 = pk.datasetLog2; p->words = 1u << pk.datasetLog2; p->cacheWords = 1u << pk.cacheLog2Words; p->cacheSegments = pk.cacheSegments;
bool sizeOverride = c.datasetLog2Override > 0 && (uint32_t)c.datasetLog2Override != pk.datasetLog2;
if (sizeOverride) { p->datasetLog2 = (uint32_t)c.datasetLog2Override; p->words = 1u << p->datasetLog2; }
// Compile
Compiled ck, cb;
if (!rtcCompile(c, kernelDev, "kernel.cu", programH, memhardH, { "igneum_cache_fill", "igneum_build" }, ck, err)) { releasePair(c, p); return nullptr; }
@ -802,7 +809,10 @@ static Pair* buildPair(Ctx& c, const std::string& dir, CUstream s, std::string&
p->dsMs = wallMs() - t0;
// Self-test against vectors.h
t0 = wallMs();
if (!pk.haveVectors) {
if (sizeOverride) {
p->checked = false; p->checkPass = true;
p->check = fmt("self-test skipped (dataset 2^%u words is not the pack's 2^%u: the vectors are for the pack size)", p->datasetLog2, pk.datasetLog2);
} else if (!pk.haveVectors) {
p->checked = false; p->checkPass = true;
p->check = "self-test skipped (no vectors.h in the pack); the miner's CPU re-check covers every found nonce";
} else {
@ -876,6 +886,8 @@ static void prepareRun(Ctx* c, PrepareTask* t) {
struct Options {
bool serve = false, check = false, raceOnly = false;
bool bench = false, memprobe = false, list = false; // the reproducible benchmark (bench/, 6 October 2026)
int batches = 5, seconds = 0, probeMib = 0, datasetMib = 0;
int device = 0, batchLog2 = 22, blockWarps = 1;
std::string pack, arch = "auto";
std::string race = "on", pinned, tuningPath;
@ -890,6 +902,16 @@ static void usage() {
" --batch-log2 B nonces per dispatch = 2^B (default 22)\n"
" --block-warps W warps per thread block (default 1)\n"
" --arch sm_XY|compute_XY|auto NVRTC target (default auto: the device's architecture)\n"
" --list every CUDA device (index, name, architecture, SMs, memory), then exit; no pack\n"
" --bench --pack <dir> the reproducible benchmark: build and self-test the pack (its vectors through the bound kernel with\n"
" the pack's seed words), then time dispatches of 2^B nonces for --seconds N (default: --batches N\n"
" dispatches); one RESULT line with the rate, the hashes and the 2^B fingerprint at base nonce 0\n"
" --seconds N --bench: time dispatches until N seconds of wall time have passed (0 = --batches dispatches)\n"
" --batches N --bench with --seconds 0: timed dispatches (default 5)\n"
" --dataset-mib N --bench: build the dataset at N MiB (power of two) instead of the pack's size; vectors skipped\n"
" --memprobe [--probe-mib N] no pack: dependent random 4-byte reads (the hash's pattern), eight independent chains,\n"
" 16 and 64-byte reads, a coalesced stream and an integer chain at 4, 64, 256 and 1024 MiB\n"
" (the same table as igneum-worker-opencl --memprobe); --probe-mib N for one size\n"
" --race --pack <dir> the variant race alone (3 rounds): one line per variant, the race line, exit 0 or 1\n"
" --race on|off|a,b,c in --serve: race every variant (default), none, or these names\n"
" --race-bench-ms N timed window per variant (default 2000)\n"
@ -906,6 +928,13 @@ static Options parseArgs(int argc, char** argv) {
auto next = [&]() -> std::string { if (i + 1 >= argc) { usage(); std::exit(2); } return argv[++i]; };
if (a == "--serve") o.serve = true;
else if (a == "--check") o.check = true;
else if (a == "--bench") o.bench = true;
else if (a == "--memprobe") o.memprobe = true;
else if (a == "--list") o.list = true;
else if (a == "--batches") o.batches = std::atoi(next().c_str());
else if (a == "--seconds") o.seconds = std::atoi(next().c_str());
else if (a == "--probe-mib") o.probeMib = std::atoi(next().c_str());
else if (a == "--dataset-mib") o.datasetMib = std::atoi(next().c_str());
else if (a == "--race" && (i + 1 >= argc || std::string(argv[i + 1]).rfind("--", 0) == 0)) o.raceOnly = true;
else if (a == "--race") o.race = next();
else if (a == "--race-bench-ms") o.raceBenchMs = std::atoi(next().c_str());
@ -924,12 +953,15 @@ static Options parseArgs(int argc, char** argv) {
}
if (o.batchLog2 < 10 || o.batchLog2 > 28) { std::printf("--batch-log2 must be between 10 and 28\n"); std::exit(2); }
if (o.blockWarps < 1 || o.blockWarps > 32) { std::printf("--block-warps must be between 1 and 32\n"); std::exit(2); }
if (!o.serve && !o.check && !o.raceOnly) { usage(); std::exit(2); }
if (!o.serve && !o.check && !o.raceOnly && !o.bench && !o.memprobe && !o.list) { usage(); std::exit(2); }
if (o.batches < 1) { std::printf("--batches must be at least 1\n"); std::exit(2); }
if (o.seconds < 0 || o.seconds > 3600) { std::printf("--seconds must be between 0 and 3600\n"); std::exit(2); }
if (o.datasetMib < 0 || (o.datasetMib > 0 && ((o.datasetMib & (o.datasetMib - 1)) != 0 || o.datasetMib < 1 || o.datasetMib > 16384))) { std::printf("--dataset-mib must be a power of two between 1 and 16384\n"); std::exit(2); }
if (o.raceBenchMs < 200 || o.raceBenchMs > 20000) { std::printf("--race-bench-ms must be between 200 and 20000\n"); std::exit(2); }
if (o.raceBudgetS < 5 || o.raceBudgetS > 540) { std::printf("--race-budget-s must be between 5 and 540 (the prepare lead is 600 DAA)\n"); std::exit(2); }
if (o.raceRounds == 0) o.raceRounds = o.raceOnly ? 3 : 1;
if (o.tuningPath.empty()) if (const char* t = std::getenv("IGNEUM_TUNING_FILE")) o.tuningPath = t;
if (o.pack.empty()) { std::printf("--pack <dir> is required (igneum-miner export-pack <node> <dir> writes one)\n"); std::exit(2); }
if (o.pack.empty() && !o.memprobe && !o.list) { std::printf("--pack <dir> is required (igneum-miner export-pack <node> <dir> writes one)\n"); std::exit(2); }
while (o.pack.size() > 1 && (o.pack.back() == '/' || o.pack.back() == '\\')) o.pack.pop_back();
return o;
}
@ -1127,22 +1159,290 @@ static int runServe(Ctx& c, const Options& o, Pair* cur) {
// ---------------------------------------------------------------------------------------------
// Main
// ---------------------------------------------------------------------------------------------
// The reproducible benchmark (bench/repro.sh and repro.ps1, 6 October 2026): --list, --bench and --memprobe.
// --bench and --memprobe were first written for the read-width experiment of 5 October 2026 (branch readwidth) and
// are carried here without that experiment's pack classes; the probe table is the OpenCL worker's (host.c).
static uint64_t fnv1a64Bytes(const void* p, size_t n) {
const uint8_t* b = (const uint8_t*)p;
uint64_t h = 0xcbf29ce484222325ull;
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
return h;
}
// --list: every CUDA device the driver sees, one line each, no context created.
static int runList(Ctx& c) {
std::string err;
CUresult r = c.drv.init(0);
if (r != CUDA_SUCCESS) { std::printf("cuda: cuInit failed: %s\n", c.err(r).c_str()); return 2; }
int count = 0;
if (c.drv.deviceGetCount(&count) != CUDA_SUCCESS) { std::printf("cuda: cuDeviceGetCount failed\n"); return 2; }
int drv = 0; c.drv.driverGetVersion(&drv);
std::printf("cuda devices (%d), driver %d.%d:\n", count, drv / 1000, (drv % 100) / 10);
for (int i = 0; i < count; ++i) {
CUdevice d = 0; char name[256] = {0}; int major = 0, minor = 0, sms = 0; size_t mem = 0;
if (c.drv.deviceGet(&d, i) != CUDA_SUCCESS) continue;
c.drv.deviceGetName(name, 255, d);
c.drv.deviceGetAttribute(&major, CU_DEVICE_ATTRIBUTE_COMPUTE_CAPABILITY_MAJOR, d);
c.drv.deviceGetAttribute(&minor, CU_DEVICE_ATTRIBUTE_COMPUTE_CAPABILITY_MINOR, d);
c.drv.deviceGetAttribute(&sms, CU_DEVICE_ATTRIBUTE_MULTIPROCESSOR_COUNT, d);
if (c.drv.deviceTotalMem) c.drv.deviceTotalMem(&mem, d);
std::printf(" [%d] %s sm_%d%d %d SMs %llu MiB\n", i, name, major, minor, sms, (unsigned long long)(mem >> 20));
std::printf("DEVICE index=%d backend=cuda name=\"%s\" arch=sm_%d%d sms=%d memory_mib=%llu\n", i, name, major, minor, sms, (unsigned long long)(mem >> 20));
}
return 0;
}
// The pack's program id, from program.h (packfile.h does not carry it).
static std::string packProgramId(const std::string& dir) {
bool ok = false;
std::string h = readText(dir + "/program.h", ok);
if (!ok) return "";
size_t at = h.find("#define IGNEUM_PROGRAM_ID ");
if (at == std::string::npos) return "";
at += 26;
size_t end = at;
while (end < h.size() && (std::isalnum((unsigned char)h[end]) || h[end] == 'x')) ++end;
std::string v = h.substr(at, end - at);
if (v.rfind("0x", 0) == 0) v = v.substr(2);
while (!v.empty() && (v.back() == 'u' || v.back() == 'l' || v.back() == 'U' || v.back() == 'L')) v.pop_back();
return v;
}
// --bench: the pair is built and self-tested (the vectors through the bound kernel with the pack's seed words, which is
// igneum_hash of kernel.cu); then a warm-up dispatch at base nonce 0 (fingerprinted) and timed dispatches of 2^B nonces,
// wall time around cuStreamSynchronize (the driver API path loads no event symbols; a 2^24 dispatch is 60 to 900 ms on
// the cards measured so far, so the launch overhead is under 1 percent). With --seconds N the dispatches continue until
// N seconds of wall time have passed; the rate is hashes over the summed dispatch time.
static int runBench(Ctx& c, const Options& o, Pair* p) {
uint32_t nonces = 1u << o.batchLog2, block = 32u * (uint32_t)c.blockWarps;
CUdeviceptr dOut = 0;
std::string err;
if (c.drv.memAlloc(&dOut, (size_t)nonces * 8u) != CUDA_SUCCESS) { std::printf("FAIL: cuMemAlloc out\n"); return 2; }
std::vector<uint64_t> hOut(nonces);
double sum = 0, warm = 0, wall0 = 0;
uint64_t fp = 0;
int done = 0;
for (int b = -1;; ++b) {
double t0 = wallMs();
if (!launchHash(c, p, dOut, (uint32_t)((uint64_t)(b + 1) * nonces), p->sw, nonces, block, nullptr, err)) { std::printf("FAIL: %s\n", err.c_str()); return 2; }
CUresult r = c.drv.streamSynchronize(nullptr);
if (r != CUDA_SUCCESS) { std::printf("FAIL: dispatch %d: %s\n", b, c.err(r).c_str()); return 2; }
double ms = wallMs() - t0;
if (b < 0) {
warm = ms;
if (c.drv.memcpyDtoH(hOut.data(), dOut, (size_t)nonces * 8u) != CUDA_SUCCESS) { std::printf("FAIL: read-back\n"); return 2; }
fp = fnv1a64Bytes(hOut.data(), (size_t)nonces * 8u);
wall0 = wallMs();
continue;
}
sum += ms; ++done;
if (o.seconds > 0) { if (wallMs() - wall0 >= o.seconds * 1000.0) break; }
else if (done >= o.batches) break;
}
c.drv.memFree(dOut);
std::string dev = c.name;
double hashes = (double)nonces * (double)done;
double mhs = hashes / (sum / 1000.0) / 1e6;
std::printf("warm-up dispatch (base 0): %.2f ms, fingerprint %016llx; %d timed dispatches of %u nonces in %.1f s: mean %.2f ms\n", warm, (unsigned long long)fp, done, nonces, sum / 1000.0, sum / done);
std::printf("rate: %.3f Mhash/s, %.2f GB/s useful (loads x 4 B), %.3f G random loads/s\n", mhs, mhs * 1e6 * 128.0 * 4.0 / 1e9, mhs * 1e6 * 128.0 / 1e9);
std::printf("RESULT bench backend=cuda device=\"%s\" arch=%s pack=%s seed=\"%s\" program_id=%s dataset_mib=%llu loads=128 check=%s fingerprint=%016llx batch_log2=%d dispatches=%d hashes=%.0f seconds=%.3f mhs=%.3f regs=%d blocks_per_sm=%d block_warps=%d time=wall\n",
dev.c_str(), c.archOpt.c_str(), p->dir.c_str(), p->seedString.c_str(), packProgramId(p->dir).c_str(), (unsigned long long)(((uint64_t)p->words * 4u) >> 20),
p->checked ? (p->checkPass ? "PASS" : "FAIL") : "skipped", (unsigned long long)fp, o.batchLog2, done, hashes, sum / 1000.0, mhs, p->regs, p->blocksPerSM, c.blockWarps);
return 0;
}
// --memprobe: the OpenCL worker's table (proto-opencl/host.c, 5 October 2026) in CUDA C through NVRTC, so the two
// vendors are probed with the same access patterns: a dependent chain of random 4-byte reads (the hash's pattern),
// eight independent chains per lane, dependent random 16-byte and 64-byte reads, a coalesced stream and an integer
// chain, at 4, 64, 256 and 1024 MiB. Wall time around cuStreamSynchronize, best of 3, a fresh seed per repetition.
static const char* PROBE_CUDA =
"#include <cstdint>\n"
"__device__ __forceinline__ uint32_t pm_mix(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }\n"
"extern \"C\" __global__ void probe_fill(uint32_t* ds, uint32_t n) { uint32_t i = blockIdx.x * blockDim.x + threadIdx.x; if (i < n) ds[i] = pm_mix(i ^ 0x9E3779B9u); }\n"
"extern \"C\" __global__ void probe_chase(const uint32_t* ds, uint32_t mask, uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed);\n"
" for (uint32_t s = 0u; s < steps; ++s) x = ds[x & mask] ^ (x * 0x9E3779B1u + s);\n"
" out[g] = x;\n"
"}\n"
"extern \"C\" __global__ void probe_indep(const uint32_t* ds, uint32_t mask, uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x;\n"
" uint32_t x0 = pm_mix(g * 8u ^ seed), x1 = pm_mix((g * 8u + 1u) ^ seed), x2 = pm_mix((g * 8u + 2u) ^ seed), x3 = pm_mix((g * 8u + 3u) ^ seed);\n"
" uint32_t x4 = pm_mix((g * 8u + 4u) ^ seed), x5 = pm_mix((g * 8u + 5u) ^ seed), x6 = pm_mix((g * 8u + 6u) ^ seed), x7 = pm_mix((g * 8u + 7u) ^ seed);\n"
" for (uint32_t s = 0u; s < steps; ++s) {\n"
" x0 = ds[x0 & mask] ^ (x0 * 0x9E3779B1u + s); x1 = ds[x1 & mask] ^ (x1 * 0x9E3779B1u + s);\n"
" x2 = ds[x2 & mask] ^ (x2 * 0x9E3779B1u + s); x3 = ds[x3 & mask] ^ (x3 * 0x9E3779B1u + s);\n"
" x4 = ds[x4 & mask] ^ (x4 * 0x9E3779B1u + s); x5 = ds[x5 & mask] ^ (x5 * 0x9E3779B1u + s);\n"
" x6 = ds[x6 & mask] ^ (x6 * 0x9E3779B1u + s); x7 = ds[x7 & mask] ^ (x7 * 0x9E3779B1u + s);\n"
" }\n"
" out[g] = x0 ^ x1 ^ x2 ^ x3 ^ x4 ^ x5 ^ x6 ^ x7;\n"
"}\n"
"extern \"C\" __global__ void probe_line16(const uint4* ds, uint32_t vecMask, uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed);\n"
" for (uint32_t s = 0u; s < steps; ++s) { uint4 a = ds[x & vecMask]; x = (a.x ^ a.y ^ a.z ^ a.w) ^ (x * 0x9E3779B1u + s); }\n"
" out[g] = x;\n"
"}\n"
"extern \"C\" __global__ void probe_line(const uint4* ds, uint32_t lineMask, uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed);\n"
" for (uint32_t s = 0u; s < steps; ++s) { uint32_t l = (x & lineMask) * 4u; uint4 a = ds[l], b = ds[l + 1u], c = ds[l + 2u], d = ds[l + 3u]; x = (a.x ^ b.y ^ c.z ^ d.w) ^ (x * 0x9E3779B1u + s); }\n"
" out[g] = x;\n"
"}\n"
"extern \"C\" __global__ void probe_stream(const uint4* ds, uint32_t perLane, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x, n = gridDim.x * blockDim.x; uint4 acc = make_uint4(0u, 0u, 0u, 0u);\n"
" for (uint32_t s = 0u; s < perLane; ++s) { uint4 v = ds[s * n + g]; acc.x ^= v.x; acc.y ^= v.y; acc.z ^= v.z; acc.w ^= v.w; }\n"
" out[g] = acc.x ^ acc.y ^ acc.z ^ acc.w;\n"
"}\n"
"extern \"C\" __global__ void probe_alu(uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u;\n"
" for (uint32_t s = 0u; s < steps; ++s) { x = x * 0x9E3779B1u + ((y << 7u) | (y >> 25u)); y = (y ^ x) + s; }\n"
" out[g] = x ^ y;\n"
"}\n";
static double probeLaunch(Ctx& c, CUfunction f, size_t lanes, size_t local, int reps, int seedArg, uint32_t seed, void** args) {
double best = -1;
for (int r = 0; r < reps; ++r) {
uint32_t s = seed + (uint32_t)r * 0x9E3779B9u;
if (seedArg >= 0) args[seedArg] = &s;
double t0 = wallMs();
if (c.drv.launchKernel(f, (unsigned)(lanes / local), 1, 1, (unsigned)local, 1, 1, 0, nullptr, args, nullptr) != CUDA_SUCCESS) return -1;
if (c.drv.streamSynchronize(nullptr) != CUDA_SUCCESS) return -1;
double ms = wallMs() - t0;
if (best < 0 || ms < best) best = ms;
}
return best;
}
static int runMemprobe(Ctx& c, const Options& o) {
Compiled cp;
std::string err;
if (!rtcCompile(c, PROBE_CUDA, "probe.cu", "", "", {}, cp, err)) { std::printf("memprobe: build FAILED: %s\n", err.c_str()); return 2; }
CUmodule mod = nullptr;
if (c.drv.moduleLoadData(&mod, cp.image.data()) != CUDA_SUCCESS) { std::printf("memprobe: cuModuleLoadData failed\n"); return 2; }
CUfunction kFill, kChase, kIndep, kLine16, kLine, kStream, kAlu;
const char* names[7] = { "probe_fill", "probe_chase", "probe_indep", "probe_line16", "probe_line", "probe_stream", "probe_alu" };
CUfunction* fns[7] = { &kFill, &kChase, &kIndep, &kLine16, &kLine, &kStream, &kAlu };
for (int i = 0; i < 7; ++i) if (c.drv.moduleGetFunction(fns[i], mod, names[i]) != CUDA_SUCCESS) { std::printf("memprobe: %s not in the module\n", names[i]); return 2; }
int sizes[4] = { 4, 64, 256, 1024 }, nSizes = 4;
if (o.probeMib > 0) { sizes[0] = o.probeMib; nSizes = 1; }
const size_t lanesList[8] = { 256, 1024, 1u << 12, 1u << 14, 1u << 16, 1u << 18, 1u << 20, 1u << 22 };
const size_t groups[2] = { 32, 256 };
const uint32_t STEPS = 256u, ALU_STEPS = 4096u;
const size_t maxLanes = 1u << 22;
CUdeviceptr dOut = 0;
if (c.drv.memAlloc(&dOut, maxLanes * 4u) != CUDA_SUCCESS) { std::printf("memprobe: cuMemAlloc out\n"); return 2; }
std::printf("memprobe on %s (sm_%d%d, %d SMs, driver %d.%d, NVRTC %d.%d), wall time around cuStreamSynchronize, best of 3\n", c.name.c_str(), c.major, c.minor, c.sms, c.driverVersion / 1000, (c.driverVersion % 100) / 10, c.rtcMajor, c.rtcMinor);
std::printf("| probe | MiB | block | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |\n|---|---|---|---|---|---|---|---|\n");
for (int si = 0; si < nSizes; ++si) {
int mib = sizes[si];
uint64_t bytes = (uint64_t)mib << 20;
uint32_t words = (uint32_t)(bytes / 4ull), mask = words - 1u, n = words;
CUdeviceptr dDs = 0;
double bestChase = 0, bestChaseMs = 0, latNs = 0; size_t bestLanes = 0;
if (c.drv.memAlloc(&dDs, (size_t)bytes) != CUDA_SUCCESS) { std::printf("| chase | %d | skipped: cuMemAlloc failed | | | | | |\n", mib); continue; }
{ void* a[2] = { &dDs, &n }; probeLaunch(c, kFill, ((size_t)words + 255) / 256 * 256, 256, 1, -1, 0, a); }
for (int gi = 0; gi < 2; ++gi) {
size_t local = groups[gi];
for (int li = 0; li < 8; ++li) {
size_t lanes = lanesList[li];
if (lanes < local) continue;
uint32_t seed = 0x1234567u + (uint32_t)li * 977u, steps = STEPS;
void* a[5] = { &dDs, &mask, &steps, &seed, &dOut };
double ms = probeLaunch(c, kChase, lanes, local, 3, 3, seed, a);
double gl = (double)lanes * STEPS / (ms / 1000.0) / 1e9;
if (gl > bestChase) { bestChase = gl; bestChaseMs = ms; bestLanes = lanes; }
if (local == 32 && lanes == 256) latNs = ms * 1e6 / STEPS; // one warp of 8 chains: the dependent-load latency
std::printf("| chase | %d | %zu | %zu | %u | %.3f | %.3f | %.0f |\n", mib, local, lanes, STEPS, ms, gl, ms * 1e6 / STEPS);
std::fflush(stdout);
}
}
double bestIndep = 0;
for (size_t lanes = 1u << 16; lanes <= maxLanes; lanes <<= 2) {
uint32_t seed = 0x7654321u, steps = STEPS;
void* a[5] = { &dDs, &mask, &steps, &seed, &dOut };
double ms = probeLaunch(c, kIndep, lanes, 256, 3, 3, seed, a);
double gl = (double)lanes * 8.0 * STEPS / (ms / 1000.0) / 1e9;
if (gl > bestIndep) bestIndep = gl;
std::printf("| indep x8 | %d | 256 | %zu | %u | %.3f | %.3f | (8 loads in flight per lane) |\n", mib, lanes, STEPS, ms, gl);
}
double bestLine16 = 0;
for (size_t lanes = 1u << 14; lanes <= maxLanes; lanes <<= 2) {
uint32_t vecMask = (words / 4u) - 1u, seed = 0x2718281u, steps = STEPS;
void* a[5] = { &dDs, &vecMask, &steps, &seed, &dOut };
double ms = probeLaunch(c, kLine16, lanes, 256, 3, 3, seed, a);
double gl = (double)lanes * STEPS / (ms / 1000.0) / 1e9;
if (gl > bestLine16) bestLine16 = gl;
std::printf("| line 16 B | %d | 256 | %zu | %u | %.3f | %.3f G reads/s | %.1f GB/s in 16 B reads |\n", mib, lanes, STEPS, ms, gl, (double)lanes * STEPS * 16.0 / (ms / 1000.0) / 1e9);
}
double bestLine = 0;
for (size_t lanes = 1u << 14; lanes <= maxLanes; lanes <<= 2) {
uint32_t lineMask = (words / 16u) - 1u, seed = 0x3141592u, steps = STEPS;
void* a[5] = { &dDs, &lineMask, &steps, &seed, &dOut };
double ms = probeLaunch(c, kLine, lanes, 256, 3, 3, seed, a);
double gl = (double)lanes * STEPS / (ms / 1000.0) / 1e9;
if (gl > bestLine) bestLine = gl;
std::printf("| line 64 B | %d | 256 | %zu | %u | %.3f | %.3f G lines/s | %.1f GB/s in lines |\n", mib, lanes, STEPS, ms, gl, (double)lanes * STEPS * 64.0 / (ms / 1000.0) / 1e9);
}
double streamGbps = 0;
{
size_t lanes = 1u << 20;
uint32_t perLane = (uint32_t)((uint64_t)words / 4ull / (uint64_t)lanes);
if (perLane == 0) { perLane = 1; lanes = (size_t)words / 4u; }
double bytesRead = (double)perLane * (double)lanes * 16.0;
void* a[3] = { &dDs, &perLane, &dOut };
double ms = probeLaunch(c, kStream, lanes, 256, 3, -1, 0, a);
streamGbps = bytesRead / (ms / 1000.0) / 1e9;
std::printf("| stream | %d | 256 | %zu | %u | %.3f | %.1f GB/s coalesced | (%.0f MiB read once) |\n", mib, lanes, perLane, ms, streamGbps, bytesRead / 1048576.0);
}
c.drv.memFree(dDs);
std::printf("RESULT memprobe backend=cuda device=\"%s\" mib=%d chase_gloads=%.3f chase_ns=%.1f chase_lanes=%zu chase_best_ms=%.3f indep_gloads=%.3f line16_greads=%.3f line64_glines=%.3f stream_gbps=%.1f time=wall\n",
c.name.c_str(), mib, bestChase, latNs, bestLanes, bestChaseMs, bestIndep, bestLine16, bestLine, streamGbps);
std::fflush(stdout);
}
{
size_t lanes = 1u << 20;
uint32_t seed = 0x2468aceu, steps = ALU_STEPS;
void* a[3] = { &steps, &seed, &dOut };
double ms = probeLaunch(c, kAlu, lanes, 256, 3, 1, seed, a);
double ops = (double)lanes * ALU_STEPS * 5.0;
std::printf("| alu | 0 | 256 | %zu | %u | %.3f | %.1f G int ops/s | %.3f G steps/s per SM (approximate: 5 ops per step counted) |\n", lanes, ALU_STEPS, ms, ops / (ms / 1000.0) / 1e9, (double)lanes * ALU_STEPS / (ms / 1000.0) / 1e9 / (c.sms ? c.sms : 1));
std::printf("RESULT memprobe backend=cuda device=\"%s\" mib=0 alu_gops=%.1f time=wall\n", c.name.c_str(), ops / (ms / 1000.0) / 1e9);
}
c.drv.memFree(dOut);
c.drv.moduleUnload(mod);
std::printf("memprobe: done\n");
return 0;
}
int main(int argc, char** argv) {
Options o = parseArgs(argc, argv);
Ctx c;
c.blockWarps = o.blockWarps;
c.race = o.race; c.raceBenchMs = o.raceBenchMs; c.raceBudgetS = o.raceBudgetS; c.raceRounds = o.raceRounds; c.batchLog2 = o.batchLog2; c.pinned = o.pinned;
c.batches = o.batches; c.seconds = o.seconds;
if (o.datasetMib > 0) { int l = 0; while ((1 << l) < o.datasetMib) ++l; c.datasetLog2Override = l + 18; } // MiB -> log2 words (4-byte words)
if (o.bench || o.memprobe || o.list) c.race = "off";
if (!o.tuningPath.empty()) { bool ok = false; c.tuning = readText(o.tuningPath, ok); if (!ok) c.tuning.clear(); }
std::string err, drvLib, rtcLib;
if (!loadDriver(c.drv, err, drvLib)) { emit("error 0 " + err); return 2; }
if (o.list) return runList(c);
if (!loadNvrtc(c.rtc, err, rtcLib)) { emit("error 0 " + err); return 2; }
if (!openDevice(c, o.device, o.arch, err)) { emit("error 0 " + err); return 2; }
info(fmt("igneum-worker-cuda %s: device %d %s (sm_%d%d, %d SMs), driver %d.%d from %s, NVRTC %d.%d from %s, target %s (%s)",
WORKER_VERSION, o.device, c.name.c_str(), c.major, c.minor, c.sms, c.driverVersion / 1000, (c.driverVersion % 100) / 10, drvLib.c_str(), c.rtcMajor, c.rtcMinor, rtcLib.c_str(), c.archOpt.c_str(), c.why.c_str()));
if (!c.tuning.empty()) info(fmt("tuning file %s (%zu bytes): %s", o.tuningPath.c_str(), c.tuning.size(), readTuning(c.tuning, c.name).found ? "has an entry for this card" : "no entry for this card"));
if (o.memprobe) { int rc = runMemprobe(c, o); c.drv.primaryCtxRelease(c.dev); return rc; }
double t0 = wallMs();
Pair* cur = buildPair(c, o.pack, nullptr, err, !o.check);
Pair* cur = buildPair(c, o.pack, nullptr, err, !o.check && !o.bench);
if (!cur) { emit("error 0 " + err); return 1; }
if (o.bench) {
std::printf("pack %s on %s: %s\n", o.pack.c_str(), c.name.c_str(), pairSummary(cur).c_str());
int rc = runBench(c, o, cur);
releasePair(c, cur);
c.drv.primaryCtxRelease(c.dev);
return rc;
}
if (o.raceOnly) {
std::printf("race %s on %s (%s, %d SMs, driver %d.%d, NVRTC %d.%d, %s): %s\n", o.pack.c_str(), c.name.c_str(), c.archOpt.c_str(), c.sms, c.driverVersion / 1000, (c.driverVersion % 100) / 10, c.rtcMajor, c.rtcMinor, c.why.c_str(), pairSummary(cur).c_str());
std::printf("%s\n", cur->raceLine.c_str());

View file

@ -45,6 +45,14 @@ struct Options {
var pinnedVariant: String? = nil // --variant: this one, no race
var tuningPath: String? = nil // --tuning <file> (default IGNEUM_TUNING_FILE)
var raceTest = false // --race-test: the race for --seed on --day, rounds, table, exit
// The reproducible benchmark (bench/repro.sh, 6 October 2026): the same flags as the CUDA and OpenCL workers.
var list = false // --list: every Metal device, one DEVICE line each, exit
var pack: String? = nil // --pack <dir>: seed, day, size and mode from program.json; the GPU's vector warps against vectors.json
var seconds = 0 // --seconds N: time batches until N seconds have passed (0 = --batches batches)
var sweep = false // --sweep: the same program at 4, 64, 256 and 1024 MiB (the chip-resistance sweep)
var memprobe = false // --memprobe: the random-read probe table (chase, indep x8, 16 B, 64 B, stream, alu)
var probeMib = 0 // --probe-mib N: --memprobe at one size
var device = 0 // --device N: the Nth device of MTLCopyAllDevices (0 = the system default); the same flag as the other workers
var anyTest: Bool { fuzz != nil || edge || stats || determinism || memcheck }
}
@ -83,6 +91,13 @@ func parseArgs() -> Options {
case "--variant": o.pinnedVariant = take()
case "--tuning": o.tuningPath = take()
case "--race-test": o.raceTest = true
case "--list": o.list = true
case "--pack": o.pack = take()
case "--seconds": o.seconds = Int(take()) ?? 0
case "--sweep": o.sweep = true
case "--memprobe": o.memprobe = true
case "--probe-mib": o.probeMib = Int(take()) ?? 0
case "--device": o.device = Int(take()) ?? 0
case "-h", "--help":
print("""
igneum-bench [--seed <string>] [--hours N] [--batch-log2 22] [--batches 4]
@ -105,6 +120,13 @@ func parseArgs() -> Options {
[--race on|off|a,b,c] [--race-bench-ms 2000] [--race-budget-s 120] [--race-rounds N]
[--variant <name>] use this variant, no race [--tuning <file>] per-card tuning (IGNEUM_TUNING_FILE)
[--race-test] the race alone for --seed on --day (3 rounds): a table per variant, exit 0/1
the reproducible benchmark (bench/repro.sh, 6 October 2026; the same flags as the CUDA and OpenCL workers):
[--list] every Metal device, one DEVICE line each, then exit
[--pack <dir>] seed, day, dataset size and mode from the pack's program.json; the GPU's vector
warps are checked against vectors.json (bit-exact with the published vectors)
[--seconds N] time batches until N seconds have passed (default 0 = --batches batches)
[--sweep] the same program at 4, 64, 256 and 1024 MiB after the main run (5 batches each)
[--memprobe [--probe-mib N]] the random-read probe table at 4, 64, 256 and 1024 MiB, then exit
""")
exit(0)
default:
@ -1725,12 +1747,15 @@ func fmt(_ v: Double, _ digits: Int = 2) -> String { String(format: "%.\(digits)
// MARK: - GPU context
var selectedDeviceIndex = 0 // --device N (set before the first GPU() is made)
final class GPU {
let device: MTLDevice
let queue: MTLCommandQueue
init() {
guard let d = MTLCreateSystemDefaultDevice(), let q = d.makeCommandQueue() else {
print("FAIL: no Metal device"); exit(1)
let all = MTLCopyAllDevices()
let pick: MTLDevice? = selectedDeviceIndex > 0 ? (selectedDeviceIndex < all.count ? all[selectedDeviceIndex] : nil) : MTLCreateSystemDefaultDevice()
guard let d = pick, let q = d.makeCommandQueue() else {
print("FAIL: no Metal device\(selectedDeviceIndex > 0 ? " at index \(selectedDeviceIndex) (\(all.count) devices)" : "")"); exit(1)
}
device = d; queue = q
}
@ -1982,27 +2007,50 @@ func runEpoch(gpu: GPU, opts: Options, seedString: String, dataset: MTLBuffer, c
var gpuOutputs = [[UInt64]]()
for w in warps { gpuOutputs.append((0..<32).map { outPtr[w * 32 + $0] }) }
// Timed batches
var cbs = [MTLCommandBuffer]()
for b in 0..<opts.batches {
let cb = gpu.queue.makeCommandBuffer()!
encodeBatch(cb, base: UInt32(truncatingIfNeeded: (b + 1) * n))
cbs.append(cb)
// The fingerprint of batch 0 (FNV-1a 64 over the 2^B outputs at base nonce 0): the same value as the CUDA and
// OpenCL workers print for this pack at the same --batch-log2, so three vendors can be compared over every nonce.
let batchFingerprint = fnv64(outBuf.contents(), n * 8)
print("batch fingerprint (FNV-1a 64 of 2^\(opts.batchLog2) outputs at base nonce 0): \(h64(batchFingerprint))")
// Timed batches: --batches of them in one go, or (--seconds N) one at a time until N seconds have passed
var gpuSeconds = 0.0, wallSeconds = 0.0, batchesDone = 0
if opts.seconds > 0 {
let s0 = nowNs()
var b = 0
while true {
let cb = gpu.queue.makeCommandBuffer()!
encodeBatch(cb, base: UInt32(truncatingIfNeeded: (b + 1) * n))
cb.commit(); cb.waitUntilCompleted()
if let e = cb.error { print("FAIL: batch error \(e)"); exit(1) }
gpuSeconds += cb.gpuEndTime - cb.gpuStartTime
b += 1
if Double(nowNs() - s0) / 1e9 >= Double(opts.seconds) { break }
}
wallSeconds = Double(nowNs() - s0) / 1e9
batchesDone = b
} else {
var cbs = [MTLCommandBuffer]()
for b in 0..<opts.batches {
let cb = gpu.queue.makeCommandBuffer()!
encodeBatch(cb, base: UInt32(truncatingIfNeeded: (b + 1) * n))
cbs.append(cb)
}
let s0 = nowNs()
for cb in cbs { cb.commit() }
cbs.last!.waitUntilCompleted()
let s1 = nowNs()
for cb in cbs { if let e = cb.error { print("FAIL: batch error \(e)"); exit(1) } }
gpuSeconds = cbs.reduce(0.0) { $0 + ($1.gpuEndTime - $1.gpuStartTime) }
wallSeconds = Double(s1 - s0) / 1e9
batchesDone = opts.batches
}
let s0 = nowNs()
for cb in cbs { cb.commit() }
cbs.last!.waitUntilCompleted()
let s1 = nowNs()
for cb in cbs { if let e = cb.error { print("FAIL: batch error \(e)"); exit(1) } }
let gpuSeconds = cbs.reduce(0.0) { $0 + ($1.gpuEndTime - $1.gpuStartTime) }
let wallSeconds = Double(s1 - s0) / 1e9
let totalHashes = Double(n * opts.batches)
let totalHashes = Double(n * batchesDone)
let hpsWall = totalHashes / wallSeconds
let hpsGPU = totalHashes / gpuSeconds
let bytesPerHash = Double(program.loadsPerHash * 4)
let gbpsWall = hpsWall * bytesPerHash / 1e9
let gbpsGPU = hpsGPU * bytesPerHash / 1e9
print("timed: \(opts.batches) batches x \(n) hashes = \(Int(totalHashes)) hashes")
print("timed: \(batchesDone) batches x \(n) hashes = \(Int(totalHashes)) hashes\(opts.seconds > 0 ? " in \(fmt(wallSeconds, 1)) s wall (--seconds \(opts.seconds))" : "")")
print(" wall \(fmt(wallSeconds * 1000)) ms -> \(fmt(hpsWall / 1e6, 3)) Mhash/s, \(fmt(gbpsWall)) GB/s useful (loads x 4 B)")
print(" GPU \(fmt(gpuSeconds * 1000)) ms -> \(fmt(hpsGPU / 1e6, 3)) Mhash/s, \(fmt(gbpsGPU)) GB/s useful (loads x 4 B)")
@ -2035,6 +2083,8 @@ func runEpoch(gpu: GPU, opts: Options, seedString: String, dataset: MTLBuffer, c
print("verify warp \(w) (nonces \(base)..\(base + 31)): \(pass ? "PASS" : "FAIL") cpu \(fmt(single, 3)) ms single, \(fmt(rep, 3)) ms avg of \(reps)\(items)\(detail)")
}
let allOk = verify.allSatisfy { $0.pass }
print("RESULT bench backend=metal device=\"\(gpu.device.name)\" seed=\"\(seedString)\" program_id=\(h64(programId(program))) dataset_mib=\((1 << opts.datasetLog2) * 4 / (1 << 20)) loads=\(program.loadsPerHash) check=\(allOk ? "PASS" : "FAIL") cpu_warps=\(verify.count) fingerprint=\(h64(batchFingerprint)) batch_log2=\(opts.batchLog2) dispatches=\(batchesDone) hashes=\(Int(totalHashes)) seconds=\(fmt(gpuSeconds, 3)) wall_seconds=\(fmt(wallSeconds, 3)) mhs=\(fmt(hpsGPU / 1e6, 3)) mhs_wall=\(fmt(hpsWall / 1e6, 3)) time=gpu")
return EpochResult(seed: seedString, libraryMs: libMs, pipelineMs: pipeMs,
hashesPerSecWall: hpsWall, hashesPerSecGPU: hpsGPU, gbpsWall: gbpsWall, gbpsGPU: gbpsGPU,
loadsPerHash: program.loadsPerHash, itemsPerWarp: program.itemsPerWarp, verify: verify)
@ -3161,7 +3211,215 @@ func runRaceTest(_ opts: Options) -> Never {
exit(line.contains("winner ") ? 0 : 1)
}
// MARK: - The reproducible benchmark (bench/repro.sh, 6 October 2026): --list, --pack, --memprobe
// --list: every Metal device, as the CUDA and OpenCL workers list theirs (one DEVICE line each for the scripts).
func runList() -> Never {
let devs = MTLCopyAllDevices()
print("metal devices (\(devs.count)):")
for (i, d) in devs.enumerated() {
print(" [\(i)] \(d.name) maxBufferLength \(d.maxBufferLength / (1 << 20)) MiB unified \(d.hasUnifiedMemory) recommendedMaxWorkingSetSize \(d.recommendedMaxWorkingSetSize / (1 << 20)) MiB")
print("DEVICE index=\(i) backend=metal name=\"\(d.name)\" memory_mib=\(d.recommendedMaxWorkingSetSize / (1 << 20)) unified=\(d.hasUnifiedMemory) low_power=\(d.isLowPower)")
}
exit(0)
}
// A pack directory: what --pack takes from program.json and vectors.json (Foundation's JSON reader; the Rust crate and
// the C workers read the same files).
struct PackInfo {
var seed = "", day = "", seedBytesHex = "", programId: UInt64 = 0
var datasetLog2 = 28, memoryHard = true
var bases = [UInt32](), expected = [[UInt64]]()
var cacheFnv: UInt64 = 0
}
func readPack(_ dir: String) -> PackInfo {
func json(_ name: String) -> [String: Any] {
guard let d = FileManager.default.contents(atPath: "\(dir)/\(name)"),
let j = try? JSONSerialization.jsonObject(with: d) as? [String: Any] else { print("FAIL: cannot read \(dir)/\(name)"); exit(2) }
return j
}
func hex64(_ v: Any?) -> UInt64 { UInt64((v as? String ?? "0").replacingOccurrences(of: "0x", with: ""), radix: 16) ?? 0 }
let pj = json("program.json"), vj = json("vectors.json")
var p = PackInfo()
p.seed = pj["seed"] as? String ?? ""
p.seedBytesHex = pj["seed_bytes"] as? String ?? ""
p.programId = hex64(pj["program_id"])
p.memoryHard = (pj["dataset_mode"] as? String) == "memory-hard"
let ds = pj["dataset"] as? [String: Any] ?? [:]
p.day = ds["day"] as? String ?? ""
p.datasetLog2 = ds["log2_words"] as? Int ?? 28
for w in vj["warps"] as? [[String: Any]] ?? [] {
p.bases.append(UInt32(w["base_nonce"] as? Int ?? 0))
p.expected.append((w["expected"] as? [String] ?? []).map { UInt64($0.replacingOccurrences(of: "0x", with: ""), radix: 16) ?? 1 })
}
p.cacheFnv = hex64(vj["cache_fnv1a64"])
return p
}
// The random-read probe table of the CUDA and OpenCL workers (proto-opencl/host.c --memprobe, 5 October 2026) in
// Metal: a dependent chain of random 4-byte reads (the hash's pattern), eight independent chains per lane, dependent
// random 16-byte and 64-byte reads, a coalesced stream and an integer chain, at 4, 64, 256 and 1024 MiB. GPU time per
// command buffer, best of 3, a fresh seed per repetition (a replay of the same addresses is served from cache).
let probeMSL = """
#include <metal_stdlib>
using namespace metal;
inline uint pm_mix(uint x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }
kernel void probe_fill(device uint* ds [[buffer(0)]], constant uint& n [[buffer(1)]], uint g [[thread_position_in_grid]]) { if (g < n) ds[g] = pm_mix(g ^ 0x9E3779B9u); }
kernel void probe_chase(device const uint* ds [[buffer(0)]], constant uint& mask [[buffer(1)]], constant uint& steps [[buffer(2)]], constant uint& seed [[buffer(3)]], device uint* out [[buffer(4)]], uint g [[thread_position_in_grid]]) {
uint x = pm_mix(g ^ seed);
for (uint s = 0u; s < steps; ++s) x = ds[x & mask] ^ (x * 0x9E3779B1u + s);
out[g] = x;
}
kernel void probe_indep(device const uint* ds [[buffer(0)]], constant uint& mask [[buffer(1)]], constant uint& steps [[buffer(2)]], constant uint& seed [[buffer(3)]], device uint* out [[buffer(4)]], uint g [[thread_position_in_grid]]) {
uint x0 = pm_mix(g * 8u ^ seed), x1 = pm_mix((g * 8u + 1u) ^ seed), x2 = pm_mix((g * 8u + 2u) ^ seed), x3 = pm_mix((g * 8u + 3u) ^ seed);
uint x4 = pm_mix((g * 8u + 4u) ^ seed), x5 = pm_mix((g * 8u + 5u) ^ seed), x6 = pm_mix((g * 8u + 6u) ^ seed), x7 = pm_mix((g * 8u + 7u) ^ seed);
for (uint s = 0u; s < steps; ++s) {
x0 = ds[x0 & mask] ^ (x0 * 0x9E3779B1u + s); x1 = ds[x1 & mask] ^ (x1 * 0x9E3779B1u + s);
x2 = ds[x2 & mask] ^ (x2 * 0x9E3779B1u + s); x3 = ds[x3 & mask] ^ (x3 * 0x9E3779B1u + s);
x4 = ds[x4 & mask] ^ (x4 * 0x9E3779B1u + s); x5 = ds[x5 & mask] ^ (x5 * 0x9E3779B1u + s);
x6 = ds[x6 & mask] ^ (x6 * 0x9E3779B1u + s); x7 = ds[x7 & mask] ^ (x7 * 0x9E3779B1u + s);
}
out[g] = x0 ^ x1 ^ x2 ^ x3 ^ x4 ^ x5 ^ x6 ^ x7;
}
kernel void probe_line16(device const uint4* ds [[buffer(0)]], constant uint& vecMask [[buffer(1)]], constant uint& steps [[buffer(2)]], constant uint& seed [[buffer(3)]], device uint* out [[buffer(4)]], uint g [[thread_position_in_grid]]) {
uint x = pm_mix(g ^ seed);
for (uint s = 0u; s < steps; ++s) { uint4 a = ds[x & vecMask]; x = (a.x ^ a.y ^ a.z ^ a.w) ^ (x * 0x9E3779B1u + s); }
out[g] = x;
}
kernel void probe_line(device const uint4* ds [[buffer(0)]], constant uint& lineMask [[buffer(1)]], constant uint& steps [[buffer(2)]], constant uint& seed [[buffer(3)]], device uint* out [[buffer(4)]], uint g [[thread_position_in_grid]]) {
uint x = pm_mix(g ^ seed);
for (uint s = 0u; s < steps; ++s) { uint l = (x & lineMask) * 4u; uint4 a = ds[l], b = ds[l + 1u], c = ds[l + 2u], d = ds[l + 3u]; x = (a.x ^ b.y ^ c.z ^ d.w) ^ (x * 0x9E3779B1u + s); }
out[g] = x;
}
kernel void probe_stream(device const uint4* ds [[buffer(0)]], constant uint& perLane [[buffer(1)]], constant uint& lanes [[buffer(2)]], device uint* out [[buffer(3)]], uint g [[thread_position_in_grid]]) {
uint4 acc = uint4(0u);
for (uint s = 0u; s < perLane; ++s) { uint4 v = ds[s * lanes + g]; acc ^= v; }
out[g] = acc.x ^ acc.y ^ acc.z ^ acc.w;
}
kernel void probe_alu(constant uint& steps [[buffer(0)]], constant uint& seed [[buffer(1)]], device uint* out [[buffer(2)]], uint g [[thread_position_in_grid]]) {
uint x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u;
for (uint s = 0u; s < steps; ++s) { x = x * 0x9E3779B1u + ((y << 7u) | (y >> 25u)); y = (y ^ x) + s; }
out[g] = x ^ y;
}
"""
func runMemprobe(_ opts: Options) -> Never {
let gpu = GPU()
let lib: MTLLibrary
do { lib = try gpu.device.makeLibrary(source: probeMSL, options: MTLCompileOptions()) } catch { print("memprobe: build FAILED: \(error)"); exit(2) }
func pipe(_ n: String) -> MTLComputePipelineState {
guard let f = lib.makeFunction(name: n), let p = try? gpu.device.makeComputePipelineState(function: f) else { print("memprobe: \(n) not in the library"); exit(2) }
return p
}
let kFill = pipe("probe_fill"), kChase = pipe("probe_chase"), kIndep = pipe("probe_indep"), kLine16 = pipe("probe_line16"), kLine = pipe("probe_line"), kStream = pipe("probe_stream"), kAlu = pipe("probe_alu")
var sizes = [4, 64, 256, 1024]
if opts.probeMib > 0 { sizes = [opts.probeMib] }
let lanesList = [256, 1024, 1 << 12, 1 << 14, 1 << 16, 1 << 18, 1 << 20, 1 << 22]
let groups = [32, 256]
let STEPS: UInt32 = 256, ALU_STEPS: UInt32 = 4096
let maxLanes = 1 << 22
guard let out = gpu.device.makeBuffer(length: maxLanes * 4, options: .storageModePrivate) else { print("memprobe: no out buffer"); exit(2) }
// one launch: scalars by setBytes in the order the kernel declares; returns the best GPU ms of `reps`, a fresh seed per repetition
func launch(_ k: MTLComputePipelineState, lanes: Int, group: Int, reps: Int, buffers: [(MTLBuffer, Int)], scalars: [(UInt32, Int)], seedIndex: Int?, seed: UInt32) -> Double {
var best = -1.0
for r in 0..<reps {
let cb = gpu.queue.makeCommandBuffer()!
let enc = cb.makeComputeCommandEncoder()!
enc.setComputePipelineState(k)
for (b, i) in buffers { enc.setBuffer(b, offset: 0, index: i) }
for (v, i) in scalars { var x = v; enc.setBytes(&x, length: 4, index: i) }
if let si = seedIndex { var sv = seed &+ UInt32(r) &* 0x9E3779B9; enc.setBytes(&sv, length: 4, index: si) }
enc.dispatchThreadgroups(MTLSize(width: lanes / group, height: 1, depth: 1), threadsPerThreadgroup: MTLSize(width: group, height: 1, depth: 1))
enc.endEncoding()
cb.commit(); cb.waitUntilCompleted()
if cb.error != nil { return -1 }
let ms = (cb.gpuEndTime - cb.gpuStartTime) * 1000
if best < 0 || ms < best { best = ms }
}
return best
}
print("memprobe on \(gpu.device.name) (Metal, \(gpu.device.hasUnifiedMemory ? "unified memory" : "discrete"), maxBufferLength \(gpu.device.maxBufferLength / (1 << 20)) MiB), GPU time per command buffer, best of 3")
print("| probe | MiB | threadgroup | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |\n|---|---|---|---|---|---|---|---|")
for mib in sizes {
let bytes = mib << 20
let words = UInt32(bytes / 4), mask = words - 1
guard let ds = gpu.device.makeBuffer(length: bytes, options: .storageModePrivate) else { print("| chase | \(mib) | skipped: no buffer | | | | | |"); continue }
_ = launch(kFill, lanes: (bytes / 4 + 255) / 256 * 256, group: 256, reps: 1, buffers: [(ds, 0)], scalars: [(words, 1)], seedIndex: nil, seed: 0)
var bestChase = 0.0, bestChaseMs = 0.0, bestLanes = 0, latNs = 0.0
for group in groups {
for (li, lanes) in lanesList.enumerated() where lanes >= group {
let seed = UInt32(0x1234567) &+ UInt32(li) &* 977
let ms = launch(kChase, lanes: lanes, group: group, reps: 3, buffers: [(ds, 0), (out, 4)], scalars: [(mask, 1), (STEPS, 2)], seedIndex: 3, seed: seed)
let gl = Double(lanes) * Double(STEPS) / (ms / 1000) / 1e9
if gl > bestChase { bestChase = gl; bestChaseMs = ms; bestLanes = lanes }
if group == 32 && lanes == 256 { latNs = ms * 1e6 / Double(STEPS) }
print("| chase | \(mib) | \(group) | \(lanes) | \(STEPS) | \(fmt(ms, 3)) | \(fmt(gl, 3)) | \(fmt(ms * 1e6 / Double(STEPS), 0)) |")
}
}
var bestIndep = 0.0
var lanes = 1 << 16
while lanes <= maxLanes {
let ms = launch(kIndep, lanes: lanes, group: 256, reps: 3, buffers: [(ds, 0), (out, 4)], scalars: [(mask, 1), (STEPS, 2)], seedIndex: 3, seed: 0x7654321)
let gl = Double(lanes) * 8 * Double(STEPS) / (ms / 1000) / 1e9
if gl > bestIndep { bestIndep = gl }
print("| indep x8 | \(mib) | 256 | \(lanes) | \(STEPS) | \(fmt(ms, 3)) | \(fmt(gl, 3)) | (8 loads in flight per lane) |")
lanes <<= 2
}
var bestLine16 = 0.0
lanes = 1 << 14
while lanes <= maxLanes {
let ms = launch(kLine16, lanes: lanes, group: 256, reps: 3, buffers: [(ds, 0), (out, 4)], scalars: [(words / 4 - 1, 1), (STEPS, 2)], seedIndex: 3, seed: 0x2718281)
let gl = Double(lanes) * Double(STEPS) / (ms / 1000) / 1e9
if gl > bestLine16 { bestLine16 = gl }
print("| line 16 B | \(mib) | 256 | \(lanes) | \(STEPS) | \(fmt(ms, 3)) | \(fmt(gl, 3)) G reads/s | \(fmt(Double(lanes) * Double(STEPS) * 16 / (ms / 1000) / 1e9, 1)) GB/s in 16 B reads |")
lanes <<= 2
}
var bestLine = 0.0
lanes = 1 << 14
while lanes <= maxLanes {
let ms = launch(kLine, lanes: lanes, group: 256, reps: 3, buffers: [(ds, 0), (out, 4)], scalars: [(words / 16 - 1, 1), (STEPS, 2)], seedIndex: 3, seed: 0x3141592)
let gl = Double(lanes) * Double(STEPS) / (ms / 1000) / 1e9
if gl > bestLine { bestLine = gl }
print("| line 64 B | \(mib) | 256 | \(lanes) | \(STEPS) | \(fmt(ms, 3)) | \(fmt(gl, 3)) G lines/s | \(fmt(Double(lanes) * Double(STEPS) * 64 / (ms / 1000) / 1e9, 1)) GB/s in lines |")
lanes <<= 2
}
var streamGbps = 0.0
do {
var sl = 1 << 20
var perLane = UInt32(bytes / 16 / sl)
if perLane == 0 { perLane = 1; sl = bytes / 16 }
let read = Double(perLane) * Double(sl) * 16
let ms = launch(kStream, lanes: sl, group: 256, reps: 3, buffers: [(ds, 0), (out, 3)], scalars: [(perLane, 1), (UInt32(sl), 2)], seedIndex: nil, seed: 0)
streamGbps = read / (ms / 1000) / 1e9
print("| stream | \(mib) | 256 | \(sl) | \(perLane) | \(fmt(ms, 3)) | \(fmt(streamGbps, 1)) GB/s coalesced | (\(fmt(read / 1048576, 0)) MiB read once) |")
}
print("RESULT memprobe backend=metal device=\"\(gpu.device.name)\" mib=\(mib) chase_gloads=\(fmt(bestChase, 3)) chase_ns=\(fmt(latNs, 1)) chase_lanes=\(bestLanes) chase_best_ms=\(fmt(bestChaseMs, 3)) indep_gloads=\(fmt(bestIndep, 3)) line16_greads=\(fmt(bestLine16, 3)) line64_glines=\(fmt(bestLine, 3)) stream_gbps=\(fmt(streamGbps, 1)) time=gpu")
}
do {
let lanes = 1 << 20
let ms = launch(kAlu, lanes: lanes, group: 256, reps: 3, buffers: [(out, 2)], scalars: [(ALU_STEPS, 0)], seedIndex: 1, seed: 0x2468ace)
let ops = Double(lanes) * Double(ALU_STEPS) * 5
print("| alu | 0 | 256 | \(lanes) | \(ALU_STEPS) | \(fmt(ms, 3)) | \(fmt(ops / (ms / 1000) / 1e9, 1)) G int ops/s | (approximate: 5 ops per step counted) |")
print("RESULT memprobe backend=metal device=\"\(gpu.device.name)\" mib=0 alu_gops=\(fmt(ops / (ms / 1000) / 1e9, 1)) time=gpu")
}
print("memprobe: done")
exit(0)
}
var opts = parseArgs()
selectedDeviceIndex = opts.device
if opts.list { runList() }
if opts.memprobe { runMemprobe(opts) }
var packInfo: PackInfo? = nil
if let dir = opts.pack {
// --pack: the run is defined by the pack (seed string, day, size, construction), and its vectors are the check
let pk = readPack(dir)
packInfo = pk
opts.seed = pk.seed; opts.day = pk.day; opts.datasetLog2 = pk.datasetLog2; opts.closedForm = !pk.memoryHard; opts.hours = 1
if pk.seedBytesHex != Array(pk.seed.utf8).map({ String(format: "%02x", $0) }).joined() { print("FAIL: the pack's seed_bytes are not the seed string's bytes (a byte-seed pack; use the seed string form)"); exit(2) }
print("pack \(dir): seed \"\(pk.seed)\", day \"\(pk.day)\", dataset 2^\(pk.datasetLog2) words (\(pk.memoryHard ? "memory-hard" : "closed-form")), program id \(h64(pk.programId)), \(pk.bases.count) vector warps")
}
if opts.raceRounds == 0 { opts.raceRounds = opts.raceTest ? 3 : 1 }
generatorConfig = GeneratorConfig(loadWeight: opts.loadWeight, wideFrac: opts.wideFrac)
if opts.exportPack != nil { exportPack(opts) }
@ -3209,12 +3467,61 @@ if ctx.closed {
if !sc.ok { print("FAIL: GPU dataset words differ from the CPU derivation"); exit(1) }
}
// --pack: the published vectors (vectors.json) against this GPU, warp by warp, before the timed run
var vectorsVerdict = "none"
if let pk = packInfo {
let program = generateProgram(seedString: pk.seed)
let idOk = programId(program) == pk.programId
var lanesOk = 0, lanes = 0, warpsOk = 0
do {
let k = try compileHash(gpu, msl: generateMSL(program, datasetLog2: opts.datasetLog2))
if let got = gpuWarps(gpu, k, dataset: dataset, bases: pk.bases) {
for (i, base) in pk.bases.enumerated() {
let want = pk.expected[i]
let bad = (0..<32).filter { $0 >= want.count || got[i][$0] != want[$0] }
lanes += 32; lanesOk += 32 - bad.count
if bad.isEmpty { warpsOk += 1 }
print("vector warp base \(base) (nonces \(base)..\(base + 31)): \(bad.isEmpty ? "PASS" : "FAIL") (\(32 - bad.count) of 32 lanes)" + (bad.isEmpty ? "" : " mismatched lanes \(bad), lane \(bad[0]) gpu \(h64(got[i][bad[0]])) vectors \(h64(want[bad[0]]))"))
}
} else { print("FAIL: vector warps did not run") }
} catch { print("FAIL: \(error)"); exit(1) }
var cacheOk = true
if !ctx.closed, let c = ctx.cpuSide() {
let fp = fnv64(c.cache, cacheWords * 4)
cacheOk = fp == pk.cacheFnv
print("cache: FNV-1a 64 \(h64(fp)) \(cacheOk ? "==" : "!=") vectors.json (\(h64(pk.cacheFnv)))")
}
let pass = idOk && cacheOk && warpsOk == pk.bases.count && !pk.bases.isEmpty
vectorsVerdict = pass ? "PASS" : "FAIL"
print("RESULT vectors backend=metal device=\"\(gpu.device.name)\" pack=\(opts.pack!) seed=\"\(pk.seed)\" program_id=\(h64(programId(program))) id_match=\(idOk) cache_match=\(cacheOk) warps=\(pk.bases.count) warps_pass=\(warpsOk) lanes=\(lanes) lanes_pass=\(lanesOk) verdict=\(vectorsVerdict)")
if !pass { print("OVERALL: FAIL (vectors)"); exit(1) }
}
var results = [EpochResult]()
for epoch in 0..<max(opts.hours, 1) {
let seedString = epoch == 0 ? opts.seed : "\(opts.seed)/epoch\(epoch)"
results.append(runEpoch(gpu: gpu, opts: opts, seedString: seedString, dataset: dataset, ctx: ctx))
}
// --sweep: the same program over 4, 64, 256 and 1024 MiB datasets, 5 batches each (the chip-resistance sweep of the
// bench log: inside a card's cache against beyond it)
if opts.sweep {
var sweepOpts = opts
sweepOpts.seconds = 0; sweepOpts.batches = 5; sweepOpts.verifyWarps = 1; sweepOpts.hours = 1
print("\n=== sweep (the same program, dataset 4, 64, 256 and 1024 MiB) ===")
var rows = [(Int, Double)]()
for log2 in [20, 24, 26, 28] {
sweepOpts.datasetLog2 = log2
let ds = log2 == opts.datasetLog2 ? dataset : ctx.makeDataset(log2: log2)
let r = runEpoch(gpu: gpu, opts: sweepOpts, seedString: opts.seed, dataset: ds, ctx: ctx)
rows.append(((1 << log2) * 4 / (1 << 20), r.hashesPerSecGPU / 1e6))
if !r.allPass { print("OVERALL: FAIL (sweep verify at 2^\(log2))"); exit(1) }
}
print("| dataset MiB | Mhash/s (GPU time) |\n|---|---|")
for (mib, mhs) in rows { print("| \(mib) | \(fmt(mhs, 3)) |") }
print("RESULT sweep backend=metal device=\"\(gpu.device.name)\" seed=\"\(opts.seed)\" " + rows.map { "mhs_\($0.0)=\(fmt($0.1, 3))" }.joined(separator: " ") + " time=gpu")
}
// Summary table
print("\n=== summary (\(gpu.device.name), dataset 2^\(opts.datasetLog2) words \(ctx.modeName)\(opts.inlineDataset ? " INLINE shortcut kernel" : ""), batch 2^\(opts.batchLog2) x \(opts.batches)) ===")
print("| seed | compile ms (lib+pipe) | Mhash/s (wall) | Mhash/s (GPU) | GB/s useful (wall) | loads/hash | items/warp | CPU verify ms/warp (avg of 20) | verify |")

View file

@ -206,7 +206,10 @@ typedef struct {
// (every output word comes back, 8 bytes per nonce, the path before 5 October 2026)
int memprobe; // --memprobe: dependent-load latency and throughput, independent-load throughput and an ALU
// chain on the chosen device, no pack needed (5 October 2026, the 9070 XT on the eGPU)
int probeMib; // --probe-mib N: --memprobe at that one buffer size only (default 0 = 4, 64 and 1024 MiB)
int probeMib; // --probe-mib N: --memprobe at that one buffer size only (default 0 = 4, 64, 256 and 1024 MiB)
int bench; // --bench: with --pack D, the reproducible benchmark (bench/, 6 October 2026): the pack read at run
// time, built and self-tested as --serve does, then its bound kernel timed for --seconds
int seconds; // --seconds N: --bench times dispatches until N seconds have passed (0 = --batches dispatches)
} Options;
static int packMib(void) { return (int)(((1ull << IGNEUM_DATASET_LOG2) * 4ull) >> 20); }
@ -237,7 +240,11 @@ static void usage(void) {
" --readback M with --serve: select (default) reads back only the hits and 34 sentinel words of each dispatch through a\n"
" GPU-side pass; full reads back every output (8 bytes per nonce). IGNEUM_READBACK=full does the same.\n"
" --memprobe no pack: dependent random loads (latency and throughput against lanes in flight), independent random\n"
" loads and an ALU chain on the chosen device, at 4, 64 and 1024 MiB (--probe-mib N for one size)\n", packMib(), IGNEUM_KERNEL_PATH);
" loads and an ALU chain on the chosen device, at 4, 64, 256 and 1024 MiB (--probe-mib N for one size)\n"
" --bench --pack D the reproducible benchmark: read the pack at run time, build and self-test it (its vectors through the\n"
" bound kernel with the pack's seed words), time dispatches of 2^B nonces for --seconds N (else --batches),\n"
" one RESULT line; --dataset-mib N builds the program over another size (vectors then skipped)\n"
" --seconds N --bench: time dispatches until N seconds have passed (default 0 = --batches dispatches)\n", packMib(), IGNEUM_KERNEL_PATH);
}
static int isPow2(long long v) { return v > 0 && (v & (v - 1)) == 0; }
@ -248,7 +255,7 @@ static Options parseArgs(int argc, char** argv) {
int i;
o.datasetMib = 1024; o.batchLog2 = 24; o.batches = 5; o.groupWarps = 1; o.sweep = 0; o.device = -1;
o.exchange = 0; o.list = 0; o.timeWall = -1; o.kernelPath = IGNEUM_KERNEL_PATH; o.extraOpts = ""; o.serve = 0; o.noPrepare = 0; o.kernelGiven = 0; o.vendor = NULL; o.packDir = NULL;
o.readback = (getenv("IGNEUM_READBACK") && strcmp(getenv("IGNEUM_READBACK"), "full") == 0) ? 1 : 0; o.memprobe = 0; o.probeMib = 0;
o.readback = (getenv("IGNEUM_READBACK") && strcmp(getenv("IGNEUM_READBACK"), "full") == 0) ? 1 : 0; o.memprobe = 0; o.probeMib = 0; o.bench = 0; o.seconds = 0;
for (i = 1; i < argc; ++i) {
const char* a = argv[i];
int needs = (strcmp(a, "--dataset-mib") == 0 || strcmp(a, "--batch-log2") == 0 || strcmp(a, "--batches") == 0 ||
@ -274,6 +281,8 @@ static Options parseArgs(int argc, char** argv) {
else { printf("--readback must be select or full\n"); exit(2); }
}
else if (strcmp(a, "--memprobe") == 0) o.memprobe = 1;
else if (strcmp(a, "--bench") == 0) o.bench = 1;
else if (strcmp(a, "--seconds") == 0) { if (i + 1 >= argc) { usage(); exit(2); } o.seconds = atoi(argv[++i]); }
else if (strcmp(a, "--probe-mib") == 0) { if (i + 1 >= argc) { usage(); exit(2); } o.probeMib = atoi(argv[++i]); }
else if (strcmp(a, "--build-opts") == 0) o.extraOpts = argv[++i];
else if (strcmp(a, "--time") == 0) {
@ -298,6 +307,7 @@ static Options parseArgs(int argc, char** argv) {
if (o.batchLog2 < 10 || o.batchLog2 > 28) { printf("--batch-log2 must be between 10 and 28\n"); exit(2); }
if (o.batches < 1) { printf("--batches must be at least 1\n"); exit(2); }
if (o.groupWarps < 1 || o.groupWarps > 8) { printf("--group-warps must be between 1 and 8\n"); exit(2); }
if (o.seconds < 0 || o.seconds > 3600) { printf("--seconds must be between 0 and 3600\n"); exit(2); }
return o;
}
@ -442,6 +452,9 @@ static void printDevice(int idx, const DeviceInfo* d, int chosen) {
if (d->amdWavefront) printf(", AMD wavefront width %u", d->amdWavefront);
if (d->nvWarp) printf(", NVIDIA warp size %u", d->nvWarp);
printf("\n");
/* one machine-readable line per listed device for bench/repro.sh and repro.ps1 (6 October 2026) */
printf("DEVICE index=%d backend=opencl type=%s name=\"%s\" vendor=\"%s\" platform=\"%s\" driver=\"%s\" compute_units=%u memory_mib=%llu\n",
idx, typeName(d->type), d->name, d->vendor, d->platformName, d->driver, d->computeUnits, (unsigned long long)(d->globalMem >> 20));
}
// ---------------------------------------------------------------------------------------------
@ -1658,7 +1671,7 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) {
cl_program prog;
cl_kernel kFill, kChase, kIndep, kAlu, kLine, kStream;
size_t srcLen = strlen(PROBE_SRC);
int sizes[3] = { 4, 64, 1024 }, nSizes = 3, si;
int sizes[4] = { 4, 64, 256, 1024 }, nSizes = 4, si;
size_t lanesList[8] = { 256, 1024, 1u << 12, 1u << 14, 1u << 16, 1u << 18, 1u << 20, 1u << 22 };
const int nLanes = 8;
size_t groups[2] = { 32, 256 };
@ -1696,6 +1709,8 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) {
size_t gi, li;
size_t fillLocal = di->maxWorkGroup < 256 ? di->maxWorkGroup : 256;
size_t fillGlobal = ((size_t)words + fillLocal - 1) / fillLocal * fillLocal;
double bestChase = 0, bestChaseMs = 0, bestIndep = 0, bestLine = 0, streamGbps = 0, latNs = 0;
size_t bestLanes = 0;
if ((uint64_t)di->maxAlloc < bytes) { printf("| chase | %d | skipped: max alloc %llu MiB | | | | | |\n", mib, (unsigned long long)(di->maxAlloc >> 20)); continue; }
dDs = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, (size_t)bytes, NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer probe dataset");
CL_CHECK(clSetKernelArg(kFill, 0, sizeof(cl_mem), &dDs));
@ -1715,6 +1730,8 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) {
CL_CHECK(clSetKernelArg(kChase, 3, sizeof(cl_uint), &seed));
CL_CHECK(clSetKernelArg(kChase, 4, sizeof(cl_mem), &dOut));
ms = probeLaunch(dv, o, kChase, lanes, local, 3, 3, seed);
if ((double)lanes * (double)STEPS / (ms / 1000.0) / 1e9 > bestChase) { bestChase = (double)lanes * (double)STEPS / (ms / 1000.0) / 1e9; bestChaseMs = ms; bestLanes = lanes; }
if (local == 32 && lanes == 256) latNs = ms * 1e6 / (double)STEPS; /* one work-group of 8 chains: the dependent-load latency */
printf("| chase | %d | %llu | %llu | %u | %.3f | %.3f | %.0f |\n", mib, (unsigned long long)local, (unsigned long long)lanes, STEPS, ms,
(double)lanes * (double)STEPS / (ms / 1000.0) / 1e9, ms * 1e6 / (double)STEPS);
fflush(stdout);
@ -1732,6 +1749,7 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) {
CL_CHECK(clSetKernelArg(kIndep, 3, sizeof(cl_uint), &seed));
CL_CHECK(clSetKernelArg(kIndep, 4, sizeof(cl_mem), &dOut));
ms = probeLaunch(dv, o, kIndep, lanes, local, 3, 3, seed);
if ((double)lanes * 8.0 * (double)STEPS / (ms / 1000.0) / 1e9 > bestIndep) bestIndep = (double)lanes * 8.0 * (double)STEPS / (ms / 1000.0) / 1e9;
printf("| indep x8 | %d | %llu | %llu | %u | %.3f | %.3f | (8 loads in flight per lane) |\n", mib, (unsigned long long)local, (unsigned long long)lanes, STEPS, ms,
(double)lanes * 8.0 * (double)STEPS / (ms / 1000.0) / 1e9);
fflush(stdout);
@ -1754,6 +1772,7 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) {
CL_CHECK(clSetKernelArg(kLine, 3, sizeof(cl_uint), &seed));
CL_CHECK(clSetKernelArg(kLine, 4, sizeof(cl_mem), &dOut));
ms = probeLaunch(dv, o, kLine, lanes, local, 3, 3, seed);
if ((double)lanes * (double)STEPS / (ms / 1000.0) / 1e9 > bestLine) bestLine = (double)lanes * (double)STEPS / (ms / 1000.0) / 1e9;
printf("| line 64 B | %d | %llu | %llu | %u | %.3f | %.3f G lines/s | %.1f GB/s in lines |\n", mib, (unsigned long long)local, (unsigned long long)lanes, STEPS, ms,
(double)lanes * (double)STEPS / (ms / 1000.0) / 1e9, (double)lanes * (double)STEPS * 64.0 / (ms / 1000.0) / 1e9);
fflush(stdout);
@ -1772,10 +1791,14 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) {
CL_CHECK(clSetKernelArg(kStream, 1, sizeof(cl_uint), &perLane));
CL_CHECK(clSetKernelArg(kStream, 2, sizeof(cl_mem), &dOut));
ms = probeLaunch(dv, o, kStream, lanes, local, 3, -1, 0u);
streamGbps = bytes / (ms / 1000.0) / 1e9;
printf("| stream | %d | %llu | %llu | %u | %.3f | %.1f GB/s coalesced | (%.0f MiB read once) |\n", mib, (unsigned long long)local, (unsigned long long)lanes, perLane, ms, bytes / (ms / 1000.0) / 1e9, bytes / 1048576.0);
fflush(stdout);
}
clReleaseMemObject(dDs);
printf("RESULT memprobe backend=opencl device=\"%s\" mib=%d chase_gloads=%.3f chase_ns=%.1f chase_lanes=%llu chase_best_ms=%.3f indep_gloads=%.3f line64_glines=%.3f stream_gbps=%.1f time=%s\n",
di->name, mib, bestChase, latNs, (unsigned long long)bestLanes, bestChaseMs, bestIndep, bestLine, streamGbps, o->timeWall ? "wall" : "event");
fflush(stdout);
}
{
size_t local = di->maxWorkGroup < 256 ? di->maxWorkGroup : 256;
@ -1790,6 +1813,7 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) {
printf("| alu | 0 | %llu | %llu | %u | %.3f | %.1f G int ops/s | %.3f G steps/s per compute unit (approximate: 5 ops per step counted) |\n",
(unsigned long long)local, (unsigned long long)lanes, ALU_STEPS, ms, ops / (ms / 1000.0) / 1e9,
(double)lanes * (double)ALU_STEPS / (ms / 1000.0) / 1e9 / (double)(di->computeUnits ? di->computeUnits : 1));
printf("RESULT memprobe backend=opencl device=\"%s\" mib=0 alu_gops=%.1f time=%s\n", di->name, ops / (ms / 1000.0) / 1e9, o->timeWall ? "wall" : "event");
}
clReleaseMemObject(dOut);
clReleaseKernel(kFill); clReleaseKernel(kChase); clReleaseKernel(kIndep); clReleaseKernel(kAlu); clReleaseKernel(kLine); clReleaseKernel(kStream);
@ -1798,6 +1822,100 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) {
return 0;
}
/* ---------------------------------------------------------------------------------------------
* --bench --pack D (the reproducible benchmark, bench/repro.sh and repro.ps1, 6 October 2026): the pack in D is read
* at run time and built and self-tested exactly as the first pair of --serve is (cache head, last line and FNV-1a 64,
* dataset head, last word and samples, the vector warps through igneum_hash_bound with the pack's own seed words,
* which is igneum_hash of kernel.cl), then the bound kernel is timed: a warm-up dispatch at base nonce 0 whose 2^B
* outputs are fingerprinted, and dispatches until --seconds have passed (or --batches dispatches). One exe, any pack,
* no toolkit: the same path the one-click worker mines on. --dataset-mib N builds the dataset at another size for the
* chip-resistance sweep; the vectors are for the pack size, so the self-test is then skipped and the line says so. */
static int runBenchPack(Device* dv, const DeviceInfo* di, const Options* o) {
#if IGNEUM_DATASET_MODE != 1
(void)dv; (void)di; (void)o;
printf("FAIL: --bench --pack needs a memory-hard placeholder pack (IGNEUM_DATASET_MODE 1)\n");
return 2;
#else
uint32_t words = gServeWords;
int sizeOverride = 0;
const uint32_t nonces = 1u << o->batchLog2;
size_t groupSize = 32 * (size_t)o->groupWarps;
size_t g = (size_t)nonces;
cl_int err = 0;
cl_mem dOut, dInit;
ServePair* cur;
char perr[512];
double t0 = wallMs(), wall0, sum = 0, warm = 0;
uint64_t fp = 0xcbf29ce484222325ull;
uint64_t* all;
int done = 0, b;
uint32_t mask;
const char* programIdText = "";
static char pid[40];
if (!dv->kHashBound) { printf("FAIL: the kernel source has no igneum_hash_bound (is %s a pack directory?)\n", o->packDir); return 2; }
if (o->datasetMib > 0 && (uint32_t)o->datasetMib != (words >> 18)) { words = (uint32_t)o->datasetMib << 18; sizeOverride = 1; }
mask = words - 1u;
cur = (ServePair*)calloc(1, sizeof(ServePair));
cur->kHashBound = dv->kHashBound; cur->kCacheFill = dv->kCacheFill; cur->kBuild = dv->kBuild; cur->prog = dv->prog;
dv->kHashBound = dv->kCacheFill = dv->kBuild = NULL; dv->prog = NULL;
memcpy(cur->sw, gPack.seedw, 32); memcpy(cur->kw, gPack.keyw, 32);
if ((uint64_t)di->maxAlloc < (uint64_t)words * 4u) { printf("FAIL: CL_DEVICE_MAX_MEM_ALLOC_SIZE is %llu MiB, the dataset needs %u MiB in one buffer\n", (unsigned long long)(di->maxAlloc >> 20), words >> 18); return 2; }
if (!pairBuffers(dv, di, dv->q, cur, words, gServeCacheWords, gServeSegments, sizeOverride ? NULL : o->packDir, perr, sizeof(perr))) { printf("FAIL: pack %s: %s\n", o->packDir, perr); return 1; }
if (sizeOverride) snprintf(cur->check, sizeof(cur->check), "self-test skipped (dataset %u MiB is not the pack's %u MiB: the vectors are for the pack size)", words >> 18, gServeWords >> 18);
printf("pack %s on %s: cache %.0f dataset %.0f check %.0f ms (%.0f ms in all); %s\n", o->packDir, di->name, cur->cacheMs, cur->datasetMs, cur->checkMs, wallMs() - t0, cur->check);
printKernelInfo(di, cur->kHashBound, "igneum_hash_bound", (int)groupSize, "");
{
/* the pack's program id, from program.h (packfile.h does not carry it) */
size_t len = 0;
char* h = pf_read_pack_file(o->packDir, "program.h", &len);
const char* at = h ? strstr(h, "#define IGNEUM_PROGRAM_ID 0x") : NULL;
if (at) { size_t k = 0; at += 28; while (k < 16 && at[k] && isxdigit((unsigned char)at[k])) { pid[k] = at[k]; ++k; } pid[k] = 0; programIdText = pid; }
free(h);
}
dOut = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, (size_t)nonces * sizeof(uint64_t), NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer out");
dInit = clCreateBuffer(dv->ctx, CL_MEM_READ_ONLY | CL_MEM_COPY_HOST_PTR, 32, cur->sw, &err); CL_CHECK_ERR(err, "clCreateBuffer init words");
CL_CHECK(clSetKernelArg(cur->kHashBound, 0, sizeof(cl_mem), &cur->ds));
CL_CHECK(clSetKernelArg(cur->kHashBound, 1, sizeof(cl_mem), &dOut));
CL_CHECK(clSetKernelArg(cur->kHashBound, 3, sizeof(cl_uint), &mask));
CL_CHECK(clSetKernelArg(cur->kHashBound, 4, sizeof(cl_mem), &dInit));
all = (uint64_t*)malloc((size_t)nonces * sizeof(uint64_t));
for (b = -1;; ++b) {
cl_uint base = (cl_uint)((uint64_t)(b + 1) * (uint64_t)nonces);
cl_event ev = NULL;
double w0 = wallMs(), ms;
CL_CHECK(clSetKernelArg(cur->kHashBound, 2, sizeof(cl_uint), &base));
CL_CHECK(clEnqueueNDRangeKernel(dv->q, cur->kHashBound, 1, NULL, &g, &groupSize, 0, NULL, &ev));
CL_CHECK(clWaitForEvents(1, &ev));
ms = o->timeWall ? wallMs() - w0 : eventMs(ev);
if (ms < 0) ms = wallMs() - w0;
clReleaseEvent(ev);
if (b < 0) {
size_t i;
warm = ms;
CL_CHECK(clEnqueueReadBuffer(dv->q, dOut, CL_TRUE, 0, (size_t)nonces * sizeof(uint64_t), all, 0, NULL, NULL));
for (i = 0; i < (size_t)nonces * 8u; ++i) { fp ^= ((const uint8_t*)all)[i]; fp *= 0x100000001b3ull; }
wall0 = wallMs();
continue;
}
sum += ms; ++done;
if (o->seconds > 0) { if (wallMs() - wall0 >= o->seconds * 1000.0) break; }
else if (done >= o->batches) break;
}
free(all);
{
double hashes = (double)nonces * (double)done, mhs = hashes / (sum / 1000.0) / 1e6;
printf("warm-up dispatch (base 0): %.2f ms, fingerprint %016llx; %d timed dispatches of %u nonces in %.1f s: mean %.2f ms\n", warm, (unsigned long long)fp, done, nonces, sum / 1000.0, sum / done);
printf("rate: %.3f Mhash/s, %.2f GB/s useful (loads x 4 B), %.3f G random loads/s\n", mhs, mhs * 1e6 * 128.0 * 4.0 / 1e9, mhs * 1e6 * 128.0 / 1e9);
printf("RESULT bench backend=opencl device=\"%s\" platform=\"%s\" pack=%s seed=\"%s\" program_id=%s dataset_mib=%u loads=128 check=%s fingerprint=%016llx batch_log2=%d dispatches=%d hashes=%.0f seconds=%.3f mhs=%.3f group_warps=%d exchange=%d time=%s\n",
di->name, di->platformName, o->packDir, gPack.seedString, programIdText, words >> 18, sizeOverride ? "skipped" : (cur->checked ? "PASS" : "skipped"), (unsigned long long)fp, o->batchLog2, done, hashes, sum / 1000.0, mhs, o->groupWarps, dv->exchange, o->timeWall ? "wall" : "event");
}
clReleaseMemObject(dOut); clReleaseMemObject(dInit);
releasePair(cur);
return 0;
#endif
}
int main(int argc, char** argv) {
Options o = parseArgs(argc, argv);
DeviceInfo* devs = NULL;
@ -1821,7 +1939,7 @@ int main(int argc, char** argv) {
static char boundPath[1200];
char perr[512];
size_t n = strlen(o.packDir);
if (!o.serve) { printf("FAIL: --pack goes with --serve (the bench runs the compiled-in pack)\n"); return 2; }
if (!o.serve && !o.bench) { printf("FAIL: --pack goes with --serve or --bench (the plain bench runs the compiled-in pack)\n"); return 2; }
if (n > 1 && (o.packDir[n - 1] == '/' || o.packDir[n - 1] == '\\')) ((char*)o.packDir)[n - 1] = 0;
if (!pf_load(o.packDir, &gPack, perr, sizeof(perr))) { printf("error 0 pack %s: %s\n", o.packDir, perr); fflush(stdout); return 2; }
gGeneric = 1;
@ -1894,6 +2012,10 @@ int main(int argc, char** argv) {
printf("build options: %s\n", dv.buildOptions);
printf("exchange: %s\n", dv.exchangeNote);
if (o.serve) return runServe(&dv, di, &o);
if (o.bench) {
if (!gGeneric) { printf("FAIL: --bench needs --pack <dir> (the plain bench of the compiled-in pack runs without --bench)\n"); return 2; }
return runBenchPack(&dv, di, &o);
}
printKernelInfo(di, dv.kHash, "igneum_hash", dv.groupSize, "");
printf("program: %d instructions x %d iterations, loads/hash %d, op mix %s\n",
IGNEUM_INSTR_COUNT, IGNEUM_ITERATIONS, IGNEUM_LOADS_PER_HASH, IGNEUM_OP_MIX);

File diff suppressed because one or more lines are too long

View file

@ -345,7 +345,7 @@ for (const [file, active] of PAGES) {
'<p>One row per card, generator version and miner version. The rate is the best one measured. Integrated GPUs are not listed. Prototype rows are bench numbers from before the devnet and say so in the miner column.</p>',
table,
'<h2 id="how">How a row gets here</h2>',
'<p>Every row names the engineering log entry or the job it came from. "Measured by the team" means our own hardware and our own log. "Reported by the fleet" means a machine we do not own, read from the status lines its miner uploads.</p>',
'<p>Every row names the engineering log entry or the job it came from. "Measured by the team" means our own hardware and our own log. "Reported by the fleet" means a machine we do not own, read from the status lines its miner uploads. "Reproduced externally" means somebody outside the project ran the published reproducible benchmark package (<a href="/public/igneum-repro.tar.gz">tar.gz</a>, <a href="/public/igneum-repro.zip">zip</a>; bench/ in the repository) on their own card and the result agreed with ours within the stated tolerance (hash rate 3%, random reads 10%, vectors and fingerprint exact); "submitted" is such a result that did not, published all the same with its deltas.</p>',
'<p>MH per watt needs the card\'s power draw during the run. The app reads it on NVIDIA cards through the driver. Rows get the figure when a run records it.</p>',
'<p>There is no other Igneum miner to compare with yet, so this table compares cards, not miners. The app that produces these rows: <a href="/miner">the miner page</a>.</p>',
`<p>Rows: ${rows.length}. Source file: <code>site/miner-bench.json</code> in the repository.</p>`,

View file

@ -182,7 +182,7 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:
<p>One row per card, generator version and miner version. The rate is the best one measured. Integrated GPUs are not listed. Prototype rows are bench numbers from before the devnet and say so in the miner column.</p>
<div class="tbl"><table><thead><tr><th>Card</th><th>Generator</th><th>Best MH/s</th><th>MH per watt</th><th>Miner</th><th>Date</th><th>Source</th><th>Who measured it</th></tr></thead><tbody><tr><td>Apple M5 Max (40 GPU cores, Metal)</td><td>v1</td><td>45.2</td><td>not measured</td><td>proto-metal bench (prototype, not mining)</td><td>2026-10-03</td><td>bench log: 3 October 2026, RTX 5090 first run (the Apple row of the same table)</td><td>measured by the team. genesis program, 1 GiB dataset</td></tr><tr><td>Apple M5 Max (40 GPU cores, Metal)</td><td>v2</td><td>26.7</td><td>not measured</td><td>igneum-miner devnet v4, Metal worker with prepare</td><td>2026-10-04</td><td>bench log: 4 October 2026, first hourly program swap on the live devnet: compile-ahead, no pause, two cards</td><td>measured by the team. live devnet v4, unbroken through the hour boundary</td></tr><tr><td>Apple silicon laptop (model not reported)</td><td>v2</td><td>24.3</td><td>not measured</td><td>Igneum Miner 0.3.1 (DMG)</td><td>2026-10-04</td><td>bench log: 4 October 2026, first outside machine on the devnet: an Apple silicon laptop through the Igneum Miner app</td><td>reported by the fleet. 21.0 MH/s average over 7 minutes, 24.3 MH/s at the moment of the report, 33 accepted blocks</td></tr><tr><td>NVIDIA RTX 5090 (32 GB)</td><td>v1</td><td>229</td><td>not measured</td><td>proto-cuda bench (prototype, not mining)</td><td>2026-10-03</td><td>bench log: 3 October 2026, RTX 5090, memory-hard dataset (pack igneum-genesis-mh)</td><td>measured by the team. genesis program, 104 loads per hash, 1 GiB dataset</td></tr><tr><td>NVIDIA RTX 5090 (32 GB)</td><td>v1</td><td>185.3</td><td>not measured</td><td>proto-cuda bench (prototype, not mining)</td><td>2026-10-03</td><td>bench log: 3 October 2026, RTX 5090 first run, dataset sweep and second program</td><td>measured by the team. hourly program, 128 loads per hash, 1 GiB dataset</td></tr><tr><td>NVIDIA RTX 5090 (32 GB)</td><td>v2</td><td>124.2</td><td>not measured</td><td>Igneum Miner 0.3.0 package, prebuilt NVRTC worker</td><td>2026-10-04</td><td>bench log: 4 October 2026, the gfx1036 worker fault and what the Apple M5 Max could and could not reproduce</td><td>measured by the team. live devnet v4, 128 loads per hash, CPU re-check clean, 0 rejected</td></tr></tbody></table></div>
<h2 id="how">How a row gets here</h2>
<p>Every row names the engineering log entry or the job it came from. "Measured by the team" means our own hardware and our own log. "Reported by the fleet" means a machine we do not own, read from the status lines its miner uploads.</p>
<p>Every row names the engineering log entry or the job it came from. "Measured by the team" means our own hardware and our own log. "Reported by the fleet" means a machine we do not own, read from the status lines its miner uploads. "Reproduced externally" means somebody outside the project ran the published reproducible benchmark package (<a href="/public/igneum-repro.tar.gz">tar.gz</a>, <a href="/public/igneum-repro.zip">zip</a>; bench/ in the repository) on their own card and the result agreed with ours within the stated tolerance (hash rate 3%, random reads 10%, vectors and fingerprint exact); "submitted" is such a result that did not, published all the same with its deltas.</p>
<p>MH per watt needs the card's power draw during the run. The app reads it on NVIDIA cards through the driver. Rows get the figure when a run records it.</p>
<p>There is no other Igneum miner to compare with yet, so this table compares cards, not miners. The app that produces these rows: <a href="/miner">the miner page</a>.</p>
<p>Rows: 6. Source file: <code>site/miner-bench.json</code> in the repository.</p></article>

View file

@ -10,7 +10,7 @@
param(
[string]$Root = '',
[string[]]$Folders = @('proto-cuda/windows-app', 'proto-cuda/windows-miner', 'proto-cuda/windows-node',
'proving/windows-wsl2', 'relay/clients', 'relay/playbooks', 'packaging/windows', 'app/windows', 'tools/ci/windows'),
'proving/windows-wsl2', 'relay/clients', 'relay/playbooks', 'packaging/windows', 'app/windows', 'tools/ci/windows', 'bench'),
[switch]$AnyVersion
)
$ErrorActionPreference = 'Stop'