docs/benchmarks/repro.md: carried from repro-bench 208c9d8

This commit is contained in:
igneum-labs 2026-10-06 07:34:58 +00:00
parent ac2bca8a0a
commit d758d0a135

219
docs/benchmarks/repro.md Normal file
View file

@ -0,0 +1,219 @@
# The reproducible benchmark package
6 October 2026 (built on the evening of 5 October). One command per platform reproduces the numbers on the bench table
on an outsider's own machine on day one: `bench/repro.sh` (Linux, macOS) and `bench/repro.ps1` (Windows), shipped as
`igneum-repro-<tag>.tar.gz` and `.zip` next to the public downloads (`packaging/ota/publish-public.sh --repro`, aliases
`/public/igneum-repro.tar.gz` and `/public/igneum-repro.zip`; not deployed tonight). Package v0.1.0-repro was built from
commit `39141f5` plus the package sources of branch `repro-bench`, commit `66ccd25` (the shipped binaries were built from that
tree before the commit; the script fixes after the runs change no binary),
and run end to end on the three project machines. The deltas against the bench log are below.
Why it exists: `docs/evidence.md` has 30 claims and none is reproduced externally. The ladder from "tested by the team"
to "reproduced externally" needs "the command published and a third party's run with the same result". This is the
command. The reward for running it is section 8.1 of `docs/benchmarks/proving-e2e.md`, quoted unchanged below.
## 1. What one run does
| Step | What runs | Binary | What it reports |
|---|---|---|---|
| 1 | The machine | the workers' `--list` | OS, CPU, memory; every GPU each worker sees (name, driver, memory) |
| 2 | The lottery hash vectors of the published genesis pack (`proto-cuda/packs/igneum-genesis-mh`: seed `igneum-genesis`, day 2026-10-03, generator 2, program id `bcc1248b10cc90f2`, memory-hard 1 GiB dataset) | `igneum-pow check-pack` on the CPU; `igneum-worker-cuda --bench`, `igneum-worker-opencl --bench --pack`, `igneum-bench --pack` on every GPU | bit-exact or not: 3 warps, 96 lanes, the 256 MiB cache FNV, on the CPU through the Rust interpreter and on every GPU through its own compiler (NVRTC, the vendor's OpenCL compiler, Metal) |
| 3 | The hash benchmark, 120 s per card at 1 GiB, `--batch-log2 24` | the same workers, `--seconds 120` | MH/s over the summed dispatch time, the hashes done, the fingerprint of the first 2^24 outputs at base nonce 0 (FNV-1a 64 over 16.7 million hashes: equal on two machines means every one of them agreed) |
| 4 | The random-read probe at 4, 64, 256 and 1024 MiB | `--memprobe` | dependent random 4-byte reads per second (the hash's access pattern) and their latency at 256 lanes, eight independent chains, 16 and 64-byte lines, the coalesced stream, an integer chain |
| 5 | The chip-resistance sweep: the same program at 4, 64, 256 and 1024 MiB, 5 batches each | `--bench --dataset-mib N` (`--sweep` on Metal) | MH/s per size; in-cache rate over the 1 GiB rate; the hash's share of the card's random-read ceiling at 1 GiB (MH/s x 128 loads against the chase) |
| 6 | One fixture shard proven with the pinned guest and verified (optional) | `igneum-prove-host --mode shard --shard 0`, then `--mode verify` | execute, core and compressed proof times, proof bytes, VERIFIED or not; skipped with the reason when there is no 12 GB NVIDIA card or no prover host |
| 7 | The result | the script | `results/igneum-repro-<os>-<time>.json` (format `igneum-repro-1`: machine, every binary's sha256, the pack files' sha256, every command run with its exit status, the rows in the bench table's own shape, the tolerances) and the same as a table in `.md`, plus the full log |
The workers were extended for this (the flags are in every worker now, where the read-width experiment of 5 October had
put `--bench` and `--memprobe` on the CUDA worker and `--memprobe` on the OpenCL worker, on branches): `--list`, `--bench
--pack <dir> --seconds N --dataset-mib N`, `--memprobe` at 4, 64, 256 and 1024 MiB with one `RESULT` line per size and
a `DEVICE` line per card; the Metal worker got `--list`, `--pack` (the published vectors against the GPU's warps),
`--seconds`, `--sweep` and `--memprobe`; `igneum-pow` got `check-pack`. The pack reader accepts a string-seed pack
(the genesis pack's 14-byte seed; the chain's are 32 bytes), the same change the read-width branch made.
What a run needs: no secrets, no node, no wallet, no network. The CUDA worker needs the NVIDIA driver only on Windows
(NVIDIA's runtime compiler is in the package, `THIRD-PARTY.md`) and the driver plus `libnvrtc.so.12` on Linux. The
OpenCL worker needs the vendor's OpenCL. Nothing here earns anything.
## 2. The three runs of 5 October 2026 against the bench log
Package v0.1.0-repro, three machines asked for, two run tonight (the Mac in full, PC 2 with its 5090 shared with the live miner): PC 1 (ae432dc7: RTX 5090, RX 9070 XT, gfx1036) went
down at 22:31:06 UTC on 5 October before its second slot (its first run at 21:04 UTC ended in a PowerShell parse error
within a second per card, section 7) and nothing reaches it until the morning; its run is owed. The Mac ran under the
measure lock (every build and simulation slot held; load average 6.19 at the start, 4.15 at the end). PC 2's run is in the
coordinator's queue (section 2.2 is filled from it; "pending" until then).
### 2.1 The Mac (Apple M5 Max, 40 GPU cores, 64 GB, macOS 26.6.2), run 2fdd7365db4a5deb, 21:51 UTC, 5 October 2026
Result files: `docs/benchmarks/repro-2026-10-06/mac/`. 7 min 52 s for the two backends (Metal 120 s, Apple OpenCL 30 s as
the cross-check, each with the sweep and the probe table).
| Number | The package | The bench log | Delta | Reading |
|---|---|---|---|---|
| Vectors, CPU (Rust interpreter) | 96/96 lanes, cache FNV `48c4f5bf24166b2e` | 96/96 (4 October 2026, generator version 2 adopted) | exact | |
| Vectors, Metal | 96/96, program id `bcc1248b10cc90f2` matches | 96/96 (4 October 2026) | exact | |
| Vectors, Apple OpenCL | 96/96 | 96/96 (proto-opencl README, 4 October 2026) | exact | |
| Fingerprint of 2^24 outputs at base 0 | `25f96e7dce90bd4e` on Metal and on Apple OpenCL | none at 2^24 for this pack (the README's `f2a95d5bb84d961e` is at 2^13) | new: two compilers agree over 16.7 million nonces | the number a second machine must match |
| Hash MH/s at 1 GiB, Metal, 120 s, GPU time | 27.674 (194 dispatches of 2^24, 3.25 G hashes) | 27.9 through Apple OpenCL on this pack (README, 5 x 2^24, wall); 26.7 mining through Metal on the live devnet (4 October) | -0.8% against the bench, +3.6% against mining | the bench row |
| Hash MH/s at 1 GiB, Apple OpenCL, 30 s, wall | 26.875 | 27.9 (the same README figure) | -3.7% | outside the 3% tolerance against a 5-batch wall-time figure taken on 4 October; the cross-check path is 2.9% under Metal here, where the 3 October runs had the two Apple paths equal (45.2 against 45.0 on the version 1 program). Not the card's row |
| CPU verify, ms per 32-lane warp (avg of 50, one performance core) | 0.611 (cold 0.625) | 0.631 (4 October, version 2 units) | -3% | 16x under the 10 ms gate |
| Sweep, MH/s at 4 / 64 / 256 / 1024 MiB, Metal | 275.2 / 103.2 / 56.6 / 27.8 | 569 / 183 / 94 / 44 (3 October, version 1 program with 104 loads, closed-form dataset) | not comparable: the generator changed | in-cache over 1 GiB 9.9x against 12.9x on version 1 |
| Random reads at 1 GiB, G loads/s (chase) | 3.40 Metal, 3.40 Apple OpenCL | none for Apple before this (4.6 G was derived from the version 1 hash rate on 3 October) | new | the hash does 27.67 x 128 = 3.54 G loads/s, 1.04 of the single-chain chase: on this chip the hash is at the random-read ceiling the probe sees |
| Dependent-load latency, ns at 256 lanes | 522 Metal (GPU time), 1,273 Apple OpenCL (wall, launch included) | none | new | |
| Stream, GB/s at 1 GiB | 546 Metal, 515 Apple OpenCL | 427 GB/s dataset fill (3 October, a write) | reads 28% above the write figure | |
| Proving | skipped: the GPU prover needs a 12 GB NVIDIA card | | | as designed |
### 2.2 PC 2 (RTX 5090, gfx1036, Windows 11), 1ccfe586
Run job run-repro-pc2-20261006 (fetch-repro-pc2-20261006 first), 00:44 to 01:09 UTC, 6 October 2026, on Igneum Miner
0.3.11, one card at a time through `bench/jobs/pc-repro.ps1`. What came back: every RESULT line in the log intake (the
three per-card transcript logs are in `docs/benchmarks/repro-2026-10-06/pc2/`; the JSON and the markdown were never
written, section 7). The 5090 block ran BESIDE THE APP'S LIVE CUDA WORKER: the card-off POST was accepted, but
`api/state` answered `{}` (a 0.3.11 defect on a proving machine) so nothing confirmed a stopped worker, and nvidia-smi
showed 10,176 MiB in use before the bench and the card at 348 W after it. Every 5090 rate below is therefore a shared-card
figure, about half of what the card does alone (131 to 137 MH/s at this batch in every other run of the night), and is
not the card's row. The checks (vectors, fingerprints, the proof) stand regardless of the sharing. The gfx1036 block is the
integrated chip's own figure (3.312 MH/s, where it mines at 3.3 in the app).
| Number | The package (PC 2) | The bench log | Delta | Reading |
|---|---|---|---|---|
| Vectors, CPU (Ryzen 7 9800X3D) | 96/96, cache FNV `48c4f5bf24166b2e` | 96/96 on the M5 Max | exact | first CPU of the second architecture |
| Vectors, RTX 5090 through CUDA (NVRTC, sm_120) | 96/96, program id matches | the version 2 pack had not run on real NVIDIA silicon (evidence row 4) | exact: closes that gap | |
| Vectors, RTX 5090 through NVIDIA OpenCL | 96/96 | 96/96 on the version 1 pack (3 October) | exact | |
| Vectors, gfx1036 through AMD OpenCL | 96/96 | 96/96 on the version 1 pack (3 October) | exact: the version 2 pack on real AMD silicon | |
| Fingerprint of 2^24 outputs at base 0 | `25f96e7dce90bd4e` on CUDA, NVIDIA OpenCL and AMD OpenCL | `25f96e7dce90bd4e` on Metal and Apple OpenCL (the Mac run) | exact across five compilers and three vendors | the cross-vendor claim over 16.7 million nonces |
| Hash MH/s at 1 GiB, 5090, CUDA, 120 s | 62.412 (447 dispatches, 7.50 G hashes), beside the live miner | 139.7 with the card to itself on a version 2 pack (4 October, variant racing); 124.2 mining in the app | not comparable: shared card | re-run owed with the worker confirmed stopped |
| Hash MH/s at 1 GiB, 5090, NVIDIA OpenCL, 30 s | 62.283, beside the live miner | 219.6 against 229.0 CUDA on the version 1 pack (3 October) | not comparable | the two paths agree with each other to 0.2% even shared |
| Hash MH/s at 1 GiB, gfx1036, AMD OpenCL, 120 s | 3.312 (24 dispatches) | 4.38 on the version 1 pack (104 loads, 3 October); 3.3 mining on the live devnet (4 October) | +0.4% against mining; the version 1 figure is another program | the card's row |
| Sweep 4 / 64 / 256 / 1024 MiB, 5090 CUDA (shared) | 374 / 367 / 74 / 62 | 1,340 / 1,353 / 270 / 229 (3 October, version 1, card alone) | shape only: in-cache over 1 GiB 6.0x against 5.8x | the L2 edge is the same |
| Sweep, gfx1036 | 3.72 / 3.38 / 3.32 / 3.31 | none | new | flat: the integrated chip is not cache-bound at any size, its memory is the system's |
| Random reads at 1 GiB, 5090 (chase, best lanes) | 17.5 G loads/s CUDA (wall), 18.2 NVIDIA OpenCL (event); latency 415 ns (OpenCL event) | 23.7 G derived from the version 1 hash rate (3 October); about 16 to 18 G implied by the 9070 XT entry's 6.6x (5 October) | inside the implied range, shared card | |
| Random reads at 1 GiB, gfx1036 | 0.458 G loads/s, 566 ns | none | new | the hash does 3.312 x 128 = 0.424 G loads/s: 0.93 of the chase; at the ceiling like the 9070 XT (0.92) |
| Stream at 1 GiB, 5090 | 1,662 GB/s (OpenCL event), 339 (CUDA wall, the launch dominates) | 1,638 GB/s dataset fill (3 October) | +1.5% | the memory clock was in its full state |
| Integer chain, 5090 | 13.4 T int ops/s (CUDA wall), 45.2 (OpenCL event) | none | | the wall figure is launch-bound; the event figure is the card's |
| Proving, block-338-shard1 shard 0, the live host `/opt/igneum/igneum-prove-host` (pinned shard id `0x2b1a81cb...`) | execute 1.48 s (60.4 M cycles), core 20.2 s (18.1 MB, verify 0.60 s), compressed 33.9 s (1.27 MB, verify 0.038 s), VERIFIED; two tampered witnesses REJECTED | the proving agent the same night: 10.8 to 11.4 s compressed alone, 33 s with the miner running | the "with the miner" figure | consistent with the shared card |
| Job size, 5090 CUDA (shared): 2^21 against 2^24 | genesis pack 60.0 against 63.2 MH/s; live devnet pack 60.9 against 63.4 | the app mines at 2^21: 114.0 wall against 116.0 inside jobs at 22:40 UTC (PC 2's STATUS lines, the consequences reviewer) | 5% lower at 2^21 here, 1.8% wall-to-inside in the app | the per-job cost at 2^21 is about 5% on a shared card; the 15% gap the reviewer found between bench and app is not this alone. The rows are PC 2's; PC 1's are owed |
### 2.3 PC 1 (RTX 5090, RX 9070 XT, gfx1036, Windows 11), ae432dc7
Owed. The first job (run-repro-pc1-20261005, 21:04 to 21:09 UTC) switched each card off and on for nothing: repro.ps1
had three PowerShell parse errors (two missing parentheses on the wsl.exe lines, one `(if ...)` used as an expression),
found only by the Windows parser on the PC because the Mac has no PowerShell. The fix, the parse gate (the PC job now
parses the package script before any card is touched, and the bench/ folder is in the Windows PowerShell 5.1 parse job
of `.github/workflows/windows.yml`) and the second package were ready at 21:45 UTC; PC 1 went down at 22:31:06 UTC before
its second slot. One fact from the first job: the RX 9070 XT (gfx1201) was on the bus and switched off and on by the
app at 21:06 UTC, where the coordinator's queue had it absent since 20:40 UTC.
## 3. The deltas, and what they mean for each user tier
What each number means for each user tier, and what the package does about it (the standing rule of 5 October 2026):
| Number | Home miner, one card (8, 12, 16, 24 or 32 GB) | Rig | Pool user | What the package does |
|---|---|---|---|---|
| The bench needs 1.3 GB of device memory (1 GiB dataset, 256 MiB cache, a 128 MiB output buffer at 2^24) | every tier runs it, 8 GB included; a 4 GB card or an integrated chip with 4 GB shared runs it too | one card at a time, `--only`, so the rig keeps mining on the rest | the same | reports the free memory it found when a dataset does not fit |
| 120 s per card, plus 3 short sweep sizes and the probe (about 4 min per card) | a few minutes with the card idle | a 6-card rig is 25 min, and `--only` splits it | | `--seconds` and the skip flags |
| The hash share of the random-read ceiling (1.04 on the M5 Max, PC 2 pending) | the number that says a card is at its memory's limit, not the kernel's: a card under about 0.9 has a worker problem, not a card problem | the same per card | | the probe table is in every result, so a low share names the step |
| Apple OpenCL 2.9% under Metal on the same chip | an Apple user on the app gets the Metal rate (the app's worker is Metal); the OpenCL figure is a compiler cross-check, never the app's | | | the OpenCL row is marked cross-check and makes no bench-table row |
| CPU verify 0.611 ms per warp on an M5 Max performance core | the chain's gate is 10 ms on any core; a 2019-class laptop core is still unmeasured (evidence row 5, O-1.14) | | a pool verifying shares has 16x margin on this core | `check-pack` prints the figure on any machine that runs the package; the first slow core to run it answers O-1.14 |
| The proving step: skipped on macOS by design; on NVIDIA it needs a built prover host and a 12 GB card by the gate, and the 5090 run of the proving agent tonight peaked at 28.3 GB (bench-log, 5 October, "the 5090 alone proves block-338-shard1 shard 0 compressed in 10.8 to 11.4 s at a 28.3 GB peak") | a 12 GB, 16 GB or 24 GB owner cannot run step 6 today: the gate lets them start it and the prover would fail on memory. The litepaper's "a 12 GB card proves a shard" stays "designed" (evidence row 16) until the prover's peak is under 12 GB | | | step 6 says why it skipped; the next package raises the gate to the measured peak (28 GB) until the prover fits 12 GB, so nobody's run fails late. Filed below as owed work |
| One PC run per night at most on the project's own PCs | | | | the package itself takes minutes; the project's PCs are shared by many agents and PC 1 was down tonight, so the project's three-machine table is not complete on the first night. The outsider's run does not have that constraint |
PC 2 (6 October 2026): the 5090's package numbers are void as the card's figures (beside the live miner) and stand only
as checks; the morning re-runs the 5090 block alone with the worker confirmed stopped by nvidia-smi's compute-apps list
(`bench/jobs/pc-repro.ps1` does that now, and labels a run "UNCONFIRMED" when it cannot). For a 5090 owner the package's
own figure is still owed; for a gfx1036 owner the figure is 3.31 MH/s, flat across dataset sizes, at 0.93 of the chip's
random-read ceiling. For a proving owner: the live host proved and verified the fixture through the package's step 6 on a
32 GB card even with a miner on it (33.9 s compressed); the 12 GB gate stands as written in the table above.
## 4. Tolerances and how a result becomes a row
| Number | Tolerance | Why this width |
|---|---|---|
| Vectors (3 warps, 96 lanes, cache FNV) | exact | a hash either matches or it does not |
| Fingerprint of 2^24 outputs at base 0 | exact across machines at the same `--batch-log2` | the same |
| Hash MH/s at 1 GiB (120 s) | 3% | two 120-s runs of the same card on this evening's machines agree within about 1% with the card to itself; 3% leaves room for a driver version and a warmer card |
| Sweep MH/s at 4, 64, 256 MiB | 10% | 5 batches each, not 120 s; the in-cache sizes are the noisiest (a few hundred ms per batch) |
| Random reads at 1 GiB (G loads/s), stream GB/s | 10% | best of 3 at 8 lane counts; the memory clock's power state moves these |
| Dependent-load latency (ns at 256 lanes) | 15% | one warp, a few hundred ns, timer resolution |
| CPU verify per warp | under 10 ms (the chain's gate), no tolerance on the figure itself | any core must verify a warp inside the gate; the published cores are under 1 ms |
| Proof times (execute, core, compressed) | 25% | SP1's GPU prover varies run to run with the card's state and the server's warm-up |
The path from a submitted file to the bench table (`site/miner-bench.json`, rendered at `/miners` by `site/build.mjs`):
1. `node bench/ingest.mjs check <result.json>` prints the agreement table: every number against the team's reference
for that card model (`bench/reference.json`, which names the bench-log entry or this document's run behind every
reference), with the tolerance and the delta.
2. `node bench/ingest.mjs add <result.json> --by "reproduced externally"` appends one row per card with that label when
every number agrees, and refuses the label (recording "submitted" with the deltas in the note) when one does not. A
failing run is as public as a passing one (proving-e2e.md 8.2 step 4). The site build then fails on any private
string (`site/forbidden-strings.txt`), so a hostname in a note never reaches the page.
3. The row's `source` is `repro:<run id>`, the random id the script drew; the result file is kept under
`docs/benchmarks/repro-<date>/`, as the three of tonight are.
The relay's console (`relay/api/console.mjs`, `results`) shows bench entries synced from `docs/bench-log.md` by
`tools/console.mjs sync-bench`; a submitted result enters there through the bench-log entry that records it, after the
table row. The relay's `bench` role (`relay/lib/relay.mjs`) is for a machine of ours running benches, not for outsiders.
Until the repository is public (the public testnet), results arrive by email to the address on igneum.network; after it,
as an issue with the JSON attached.
## 5. Operator instructions
See `bench/README.md`, shipped in the package. In short: unpack, stop mining on the card, run the one command, send
`results/igneum-repro-<os>-<time>.json` back, say which card was idle. About 4 minutes per card plus the optional
proving step. The file carries no hostname, user name or address; it carries the GPU and CPU models, the OS and the
driver version. `--only backend:index` runs one card; `--no-prove`, `--no-probe`, `--no-sweep` skip steps.
Our own machines run it the same way, with one difference the script does not know about: the card under test is
switched off in the Igneum Miner app first (`bench/jobs/pc-repro.ps1` does it through the app's `api/cards`, one card at
a time, and restores it), and on the Mac the run takes the measure lock.
## 6. The reward terms (proving-e2e.md section 8.1, unchanged)
### 8.1 What counts as unrelated
Three operators are unrelated when every row holds for every pair:
| Test | Requirement |
|---|---|
| Person | Different natural or legal persons; none is the project, an agent of it, or paid by it for the run (a published fixed reproduction reward, equal for everyone and announced before the run, is allowed and disclosed) |
| Hardware | Bought separately; no shared host, card, rack or power meter |
| Network | Different autonomous systems, verified by the IP in the published logs; not the same residential ISP account |
| Location | Different physical sites |
| Software | The same published release by hash; nobody receives a private build |
| Money | No payment, loan or equipment between them or from the project, beyond the disclosed reproduction reward |
An operator declares each row in their report and signs it with the vote key that mined on the devnet under the same fingerprint, so a report is tied to a key with a history.
The amount is the challenge reward row of `docs/plans/funding.md` (USD 1,000 per operator per workload set, three
operators, about USD 3,000 per campaign, approximate; not funded as of 6 October 2026). Reproduction is asked for without
a reward, which the standard allows. The result file of this package is signed by nothing; the signature with the vote
key of 8.1 is the operator's own step, on the file.
## 7. What is unverified
| Item | Why | What closes it |
|---|---|---|
| PC 1's run (RTX 5090, RX 9070 XT, gfx1036 on Windows) | PC 1 went down at 22:31:06 UTC before its slot; the first job hit the parse errors | the morning's slot: `bench/jobs/pc-repro.ps1` as a `run` job after the package fetch, one card at a time |
| The Linux binaries (`bin/linux-x86_64`) | cross-compiled with zig on the Mac, loaded nowhere tonight (no Linux GPU host; HiveOS is the first) | a run on a Linux box with an NVIDIA or AMD driver; `repro.sh` is the same script the Mac ran |
| The CUDA worker's `--bench`, `--memprobe` and `--list` on a real card | tonight's only CUDA runs were the Mac's emulation (`--list`, `--bench` on the genesis pack: self-test PASS, fingerprint `e7d68ec2a49d0671` at 2^14, equal to Apple OpenCL's at 2^14) and PC 1's run that never reached the worker | PC 2's slot (pending) and PC 1's morning run |
| `repro.ps1` end to end | parsed only by the Windows parser on PC 1 after the fix; never run to completion on a PC | PC 2's slot |
| The proving step (step 6) | never run through the script; on PC 2 it will use the live host `/opt/igneum/igneum-prove-host` as the app's WSL user with the live prover off | PC 2's slot; then the memory gate against the measured 28.3 GB peak |
| A second machine of the same model | the tolerance table has one Apple machine; "reproduced externally" needs an unrelated operator's run, and the repository is private until the public testnet | the publish step (`publish-public.sh --repro`, not run tonight) and the first outside result |
| Apple OpenCL 3.7% under the README's 27.9 MH/s | one 30-s wall-time run against one 5-batch run of 4 October; either could be the odd one | a second 120-s run of each path on the Mac with the card to itself |
| The random-read probe on Apple as a ceiling | the hash exceeds the single-chain chase by 4% on the M5 Max, so on Apple the chase at 4 M lanes is not the ceiling the 9070 XT entry took it for | more lanes in flight, or the eight-chain probe (3.50 G here) as the Apple ceiling; a probe question, not a hash one |
| PC 2's 5090 rows | the card-off step could not confirm a stopped worker (`api/state` answered `{}`), and the numbers are half the card's | the morning's 5090-only run on PC 2 with the compute-apps confirmation; until then the rows are "beside the live miner" |
| PC 2's result files | `repro.ps1`'s close threw `System.OutOfMemoryException` in ConvertTo-Json and gave Set-Content an empty path: PowerShell variable names are case-insensitive, the result table `$Cpu` held `$cpu` (its own name string) and so contained itself, and the markdown lines `$md` wiped the markdown path `$Md`. Fixed (distinct names; `bench/jobs/ps-case-check.sh` fails CI on any case-only pair, shown to fire on a bad file); the RESULT lines in the intake and the three transcript logs are the record | the morning's run writes the files |
| The prover on PC 2 | the first script read the prover's state from `api/state` too, saw nothing, and left the live prover off from 00:49 to the restore job (run-prover-on-pc2-20261006); the script now reads settings.json and restores unconditionally | done tonight; the rule in the job |
| The reward | the amount and payer are `docs/plans/funding.md`, not funded; the terms are quoted unchanged | the project lead's decision |
## 8. Files
| File | What |
|---|---|
| `bench/repro.sh`, `bench/repro.ps1` | the operator commands |
| `bench/README.md` | the package's README (operator instructions, tolerances, how to send a result back) |
| `bench/make-package.sh` | builds the package from a tagged commit with the existing cross-build scripts (zig for Linux, mingw for Windows, swiftc and cc for macOS, cargo for the CPU tool) |
| `bench/ingest.mjs`, `bench/reference.json` | the agreement check and the bench-table ingestion; the references |
| `bench/jobs/pc-repro.ps1`, `bench/jobs/collect-results.mjs` | the PC job (one card at a time through the app) and the Mac-side extraction of its result files |
| `docs/benchmarks/repro-2026-10-06/` | the three result files of tonight, their tables and logs |
| `packaging/ota/publish-public.sh --repro` | the public downloads entry (tar.gz, zip, their sha256, the two aliases, the downloads index) |