diff --git a/docs/bench-log.md b/docs/bench-log.md index 5b83ef4fb..59391e7d3 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1962,7 +1962,6 @@ Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan. -<<<<<<< HEAD ## 6 October 2026, Counter ASIC 3.0 item 2: the per-day derivation Branch `ca3-derive` (worker "derive", from ca3-coord 50df751; commits acb96ee and after), design, spec text and @@ -2043,7 +2042,6 @@ laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare **RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2 lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED. -======= ## 6 October 2026, Counter ASIC 3.0 item 8: program work in the latency shadow Branch `ca3-shadow`, worker "shadow" (`docs/analysis/latency-shadow-2026-10-06.md` carries the design, the chip side and the consequences; this entry carries the measurements). The knob: `LoadClass::shadow`, class name `+shx`, a block of `S` ALU instructions drawn from the program stream after the 64 base instructions and run `R` times at the end of every iteration (no load; the base program, its attempt and the acceptance verdict are the class's without the shadow; v2 and v3 byte-identical, `cargo test` in igneum-pow 54 + 4 + 19 + 7 green). Packs `proto-cuda/packs-ca3-shadow/*` over `mx8` for seed igneum-genesis; the control is the pinned class v3 pack `packs-ca2-mixer/mx8-genesis`. Ops per hash = shadow instructions x 1.83 (counted from the emitted statements: add 5, rotr 2, shfl 2, the rest 1, weighted over the non-load weights) + 930 (the base program's 384 ALU instructions and 128 loads). @@ -2070,4 +2068,42 @@ Reading: the M5 Max stays latency-bound to about 100,000 ops per hash and its 5 **RX 9070 XT (PC 1, ae432dc7)**: OWED (PC 1 is the project lead's desk and not released today); the OpenCL kernels are in every pack. Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (`sh256x27`) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 by the model holds its rate at about 378 W (per watt 0.77x; the job measures it), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the `f = 1` chip's edge per joule falls from 1.7x to 1.4x against the M5 Max and from 5.0x to 2.7x against the 5090 (model watts) at `k = 1`, where `k` is the chip core's energy per op over the GPU's marginal 5.5 pJ: the number that decides the item. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000. ->>>>>>> ca3-shadow +## 6 October 2026, Counter ASIC 3.0 item 6: the reserve families' step costs + +Branch `ca3-reserve`, worker "reserve" (`docs/plans/counter-asic-3-reserve.md` carries the proposed order and spec text; this entry carries the measurements). Method: the dot4 probe's dependent chain (`docs/analysis/int8-matrix-family.md` section 4), one op of the family per step per lane, 1,048,576 lanes x 4,096 steps, best of 3 dispatches per run, three runs, bit-exact against a CPU reference on two whole 32-lane warps (the shuffle rows need the whole warp). Every chain has the same glue (`acc = OP(acc, x, y); x = x * K + acc; y = rotl(y, 7) ^ (acc + s)`), so the "step cost" is the family's one op plus four glue ops against the add-xor-rotate chain of the 9070 XT bench-log entry (`alu`: `x = x * K + rotl(y, 7); y = (y ^ x) + s`, 5 ops per step counted, no `acc`). Reference rows are live families (`alu`; `rotr` = the live `rotr_var` text; `shflx` = the live `shfl`, lane XOR 8). Candidate rows are the seven families of spec 1.13.2 (`shl`, `shr`, `bfe` with the vendor's extract function and `bfec` the C form `(y >> 7) & 0x1fff`, `andn`, `perm` = bytes (b1, b3, b0, b2), `popc` and `clz` folded by add, `sel` on bit 5, `shfla` = lane + 3 mod 32). Comparison rows: `dot4u` and `dot4s` (Apple, emulated), `dot4i` (`__dp4a`) and `mm8` (one `mma.sync.m8n8k16` u8 per step per warp, inline PTX) on CUDA. Sources: `proto-metal/family-probe.swift`, `proto-cuda/family-probe.cu`, the PC 2 job `tools/ca3-reserve/pc2-family-probe.ps1` (made by `make-pc2-playbook.sh`). G steps/s is the whole card's dependent-step throughput; ops per step counted = the family's op plus 4 glue (`alu` 5, `shfla` and `shflx` 6: shuffle plus xor, `mm8` 1 mma plus 4). + +**Apple M5 Max, Metal** (`swiftc -O -o family-probe family-probe.swift -framework Metal` under `with-lock.sh build`; three runs of `with-lock.sh measure ./family-probe --reps 3`, 07:29:05 to 07:29:08 UTC, load average 7.64 / 7.59 / 7.14 before and after every run (the Mac was loaded by other agents' builds the whole morning; the measure lock held, the GPU idle: the Mac mines nothing), GPU start-to-end time): + +| kernel | best ms, runs 1 / 2 / 3 | best of the three, ms | G steps/s (best) | ns per step (best) | ops per step counted | step cost (ratio to `alu`, best) | bit-exact, 3 runs | +|---|---|---|---|---|---|---|---| +| alu | 5.039 / 4.918 / 4.873 | 4.873 | 881 | 1,190 | 5 | 1.00 | yes | +| rotr (live) | 5.610 / 5.533 / 5.492 | 5.492 | 782 | 1,341 | 5 | 1.13 | yes | +| shflx (live) | 4.196 / 4.259 / 4.223 | 4.196 | 1,024 | 1,024 | 6 | 0.86 | yes | +| shl | 4.120 / 4.202 / 4.119 | 4.119 | 1,043 | 1,006 | 5 | 0.85 | yes | +| shr | 4.296 / 4.194 / 4.319 | 4.194 | 1,024 | 1,024 | 5 | 0.86 | yes | +| bfe (`extract_bits`) | 3.731 / 3.805 / 3.816 | 3.731 | 1,151 | 911 | 5 | 0.77 | yes | +| bfec (C form) | 3.762 / 3.818 / 3.646 | 3.646 | 1,178 | 890 | 5 | 0.75 | yes | +| andn | 3.680 / 3.732 / 3.676 | 3.676 | 1,168 | 898 | 5 | 0.75 | yes | +| perm | 5.500 / 5.548 / 5.524 | 5.500 | 781 | 1,343 | 5 | 1.13 | yes | +| popc | 4.261 / 4.223 / 4.262 | 4.223 | 1,017 | 1,031 | 5 | 0.87 | yes | +| clz | 4.918 / 4.916 / 4.914 | 4.914 | 874 | 1,200 | 5 | 1.01 | yes | +| sel | 3.718 / 3.831 / 3.718 | 3.718 | 1,155 | 908 | 5 | 0.76 | yes | +| shfla (lane + 3) | 9.337 / 9.323 / 9.287 | 9.287 | 462 | 2,267 | 6 | 1.91 | yes | +| dot4u (emulated) | 7.818 / 8.044 / 7.925 | 7.818 | 549 | 1,909 | 5 | 1.60 | yes | +| dot4s (emulated) | 23.264 / 23.254 / 23.045 | 23.045 | 186 | 5,626 | 5 | 4.73 | yes | + +Reading of the Mac rows. The run-to-run spread is under 4% on every row. The dot4 rows reproduce the 5 October figures (1.6x unsigned, 4.7x signed), which is the check on the method. A step cost under 1.00 means the family's op plus the glue is cheaper than the five-op reference chain: the reference's two registers are a tighter dependency than the three-register candidate chains, and Apple's shifts, extract, andn and select each cost about what an xor costs. Three rows cost more than the reference: `perm` (1.13: no byte-permute function in MSL; the `uchar4` swizzle compiles to shifts and masks, so a byte permute is emulated on Apple at about the price of the live `rotr`), `clz` (1.01) and `shfla` (1.91: a shuffle by a computed lane index costs 2.2x the live xor shuffle on Apple, `simd_shuffle` against `simd_shuffle_xor`; the second shuffle form is the one candidate Apple pays for). `mm8` as a chain on Apple is owed (Metal 4 `matmul2d`; this toolchain is Swift 5.8 without the tensor API). + +**RTX 5090 (PC 2, 1ccfe586), CUDA**: PENDING the PC 2 job (`run-ca3-family-pc2-20261006`, published only after `/tmp/igneum-devnet/pc2-ca3.clear` and under the mkdir lock; the card to itself: `--stop-miners`, prover off for the run). The rows land in this entry when the closing report is read. + +**RX 9070 XT (PC 1, ae432dc7), OpenCL**: OWED. PC 1 is the project lead's desk and not released today (the brief's rule); the OpenCL twin of the probe (`__builtin_amdgcn_*` paths for `v_bfe_u32`, `v_perm_b32`, `v_bcnt_u32_b32`, `v_cndmask_b32`, `ds_bpermute_b32`) is the next job on that card. + +Consequences per tier, Mac rows (the hash is latency-bound by 128 dependent DRAM reads; a family at `W_new` = 4 points is about 4% of the 64 instructions, so these per-op costs bound a family's hash-rate cost and are not hash rates; the 5% rule of 1.13.2 is argued from them, not measured, until a family is live): + +| Tier | What the rows mean | What is being done | +|---|---|---| +| Apple user (M-series laptop or desktop, the app's Metal worker) | six of the seven candidates (`shl`, `shr`, `bfe`, `andn`, `popc`, `sel`) cost at most the live `rotr` step; `clz` the same as the reference; `perm` 1.13x (emulated); `shfla` 1.91x, the only candidate over the live `shfl`'s cost by more than 2x on this card. At 4 points of 64 a 1.91x op costs under 1% of the program's ALU time, itself a small share of a latency-bound hash (approximate: argued, measured when live) | the proposed order puts `shfla` after the plain datapath families (R6), so Apple pays it last; `mm8` stays last | +| NVIDIA user (8 to 32 GB card) | pending the 5090 rows above | the PC 2 job | +| AMD user (RX 9070 XT, 16 GB) | owed: no row today | the PC 1 job when the desk is free | +| A rig or a pool user | the same per-card figures; no family changes the dependent-read bound | nothing until a family is live | +| A chip | every candidate but `mm8` is a 32-bit datapath structure (barrel shifter, byte crossbar, popcount tree, 32-lane shuffle crossbar: `docs/plans/counter-asic-3-reserve.md` section 3 names them with approximate areas); none is licensable as a block the way an int8 matrix unit is | the reserve order of that document | diff --git a/docs/benchmarks/repro.md b/docs/benchmarks/repro.md new file mode 100644 index 000000000..2527fd7aa --- /dev/null +++ b/docs/benchmarks/repro.md @@ -0,0 +1,263 @@ +# The reproducible benchmark package + +6 October 2026 (built on the evening of 5 October). One command per platform reproduces the numbers on the bench table +on an outsider's own machine on day one: `bench/repro.sh` (Linux, macOS) and `bench/repro.ps1` (Windows), shipped as +`igneum-repro-.tar.gz` and `.zip` next to the public downloads (`packaging/ota/publish-public.sh --repro`, aliases +`/public/igneum-repro.tar.gz` and `/public/igneum-repro.zip`; not deployed tonight). Package v0.1.0-repro was built from +commit `39141f5` plus the package sources of branch `repro-bench`, commit `66ccd25` (the shipped binaries were built from that +tree before the commit; the script fixes after the runs change no binary), +and run end to end on the three project machines. The deltas against the bench log are below. + +Why it exists: `docs/evidence.md` has 30 claims and none is reproduced externally. The ladder from "tested by the team" +to "reproduced externally" needs "the command published and a third party's run with the same result". This is the +command. The reward for running it is section 8.1 of `docs/benchmarks/proving-e2e.md`, quoted unchanged below. + +## 1. What one run does + +| Step | What runs | Binary | What it reports | +|---|---|---|---| +| 1 | The machine | the workers' `--list` | OS, CPU, memory; every GPU each worker sees (name, driver, memory) | +| 2 | The lottery hash vectors of the published genesis pack (`proto-cuda/packs/igneum-genesis-mh`: seed `igneum-genesis`, day 2026-10-03, generator 2, program id `bcc1248b10cc90f2`, memory-hard 1 GiB dataset) | `igneum-pow check-pack` on the CPU; `igneum-worker-cuda --bench`, `igneum-worker-opencl --bench --pack`, `igneum-bench --pack` on every GPU | bit-exact or not: 3 warps, 96 lanes, the 256 MiB cache FNV, on the CPU through the Rust interpreter and on every GPU through its own compiler (NVRTC, the vendor's OpenCL compiler, Metal) | +| 3 | The hash benchmark, 120 s per card at 1 GiB, `--batch-log2 24` | the same workers, `--seconds 120` | MH/s over the summed dispatch time, the hashes done, the fingerprint of the first 2^24 outputs at base nonce 0 (FNV-1a 64 over 16.7 million hashes: equal on two machines means every one of them agreed) | +| 4 | The random-read probe at 4, 64, 256 and 1024 MiB | `--memprobe` | dependent random 4-byte reads per second (the hash's access pattern) and their latency at 256 lanes, eight independent chains, 16 and 64-byte lines, the coalesced stream, an integer chain | +| 5 | The chip-resistance sweep: the same program at 4, 64, 256 and 1024 MiB, 5 batches each | `--bench --dataset-mib N` (`--sweep` on Metal) | MH/s per size; in-cache rate over the 1 GiB rate; the hash's share of the card's random-read ceiling at 1 GiB (MH/s x 128 loads against the chase) | +| 6 | One fixture shard proven with the pinned guest and verified (optional) | `igneum-prove-host --mode shard --shard 0`, then `--mode verify` | execute, core and compressed proof times, proof bytes, VERIFIED or not; skipped with the reason when there is no 12 GB NVIDIA card or no prover host | +| 7 | The result | the script | `results/igneum-repro--