Merge branch 'ca3-reserve' into ca3-coord
# Conflicts: # docs/bench-log.md # tools/observer/observer.mjs
This commit is contained in:
commit
e3ee3049bd
11 changed files with 1218 additions and 3 deletions
|
|
@ -1962,7 +1962,6 @@ Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB.
|
|||
|
||||
Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan.
|
||||
|
||||
<<<<<<< HEAD
|
||||
## 6 October 2026, Counter ASIC 3.0 item 2: the per-day derivation
|
||||
|
||||
Branch `ca3-derive` (worker "derive", from ca3-coord 50df751; commits acb96ee and after), design, spec text and
|
||||
|
|
@ -2043,7 +2042,6 @@ laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare
|
|||
|
||||
**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2
|
||||
lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED.
|
||||
=======
|
||||
## 6 October 2026, Counter ASIC 3.0 item 8: program work in the latency shadow
|
||||
|
||||
Branch `ca3-shadow`, worker "shadow" (`docs/analysis/latency-shadow-2026-10-06.md` carries the design, the chip side and the consequences; this entry carries the measurements). The knob: `LoadClass::shadow`, class name `<class>+sh<S>x<R>`, a block of `S` ALU instructions drawn from the program stream after the 64 base instructions and run `R` times at the end of every iteration (no load; the base program, its attempt and the acceptance verdict are the class's without the shadow; v2 and v3 byte-identical, `cargo test` in igneum-pow 54 + 4 + 19 + 7 green). Packs `proto-cuda/packs-ca3-shadow/*` over `mx8` for seed igneum-genesis; the control is the pinned class v3 pack `packs-ca2-mixer/mx8-genesis`. Ops per hash = shadow instructions x 1.83 (counted from the emitted statements: add 5, rotr 2, shfl 2, the rest 1, weighted over the non-load weights) + 930 (the base program's 384 ALU instructions and 128 loads).
|
||||
|
|
@ -2070,4 +2068,42 @@ Reading: the M5 Max stays latency-bound to about 100,000 ops per hash and its 5
|
|||
**RX 9070 XT (PC 1, ae432dc7)**: OWED (PC 1 is the project lead's desk and not released today); the OpenCL kernels are in every pack.
|
||||
|
||||
Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (`sh256x27`) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 by the model holds its rate at about 378 W (per watt 0.77x; the job measures it), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the `f = 1` chip's edge per joule falls from 1.7x to 1.4x against the M5 Max and from 5.0x to 2.7x against the 5090 (model watts) at `k = 1`, where `k` is the chip core's energy per op over the GPU's marginal 5.5 pJ: the number that decides the item. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000.
|
||||
>>>>>>> ca3-shadow
|
||||
## 6 October 2026, Counter ASIC 3.0 item 6: the reserve families' step costs
|
||||
|
||||
Branch `ca3-reserve`, worker "reserve" (`docs/plans/counter-asic-3-reserve.md` carries the proposed order and spec text; this entry carries the measurements). Method: the dot4 probe's dependent chain (`docs/analysis/int8-matrix-family.md` section 4), one op of the family per step per lane, 1,048,576 lanes x 4,096 steps, best of 3 dispatches per run, three runs, bit-exact against a CPU reference on two whole 32-lane warps (the shuffle rows need the whole warp). Every chain has the same glue (`acc = OP(acc, x, y); x = x * K + acc; y = rotl(y, 7) ^ (acc + s)`), so the "step cost" is the family's one op plus four glue ops against the add-xor-rotate chain of the 9070 XT bench-log entry (`alu`: `x = x * K + rotl(y, 7); y = (y ^ x) + s`, 5 ops per step counted, no `acc`). Reference rows are live families (`alu`; `rotr` = the live `rotr_var` text; `shflx` = the live `shfl`, lane XOR 8). Candidate rows are the seven families of spec 1.13.2 (`shl`, `shr`, `bfe` with the vendor's extract function and `bfec` the C form `(y >> 7) & 0x1fff`, `andn`, `perm` = bytes (b1, b3, b0, b2), `popc` and `clz` folded by add, `sel` on bit 5, `shfla` = lane + 3 mod 32). Comparison rows: `dot4u` and `dot4s` (Apple, emulated), `dot4i` (`__dp4a`) and `mm8` (one `mma.sync.m8n8k16` u8 per step per warp, inline PTX) on CUDA. Sources: `proto-metal/family-probe.swift`, `proto-cuda/family-probe.cu`, the PC 2 job `tools/ca3-reserve/pc2-family-probe.ps1` (made by `make-pc2-playbook.sh`). G steps/s is the whole card's dependent-step throughput; ops per step counted = the family's op plus 4 glue (`alu` 5, `shfla` and `shflx` 6: shuffle plus xor, `mm8` 1 mma plus 4).
|
||||
|
||||
**Apple M5 Max, Metal** (`swiftc -O -o family-probe family-probe.swift -framework Metal` under `with-lock.sh build`; three runs of `with-lock.sh measure ./family-probe --reps 3`, 07:29:05 to 07:29:08 UTC, load average 7.64 / 7.59 / 7.14 before and after every run (the Mac was loaded by other agents' builds the whole morning; the measure lock held, the GPU idle: the Mac mines nothing), GPU start-to-end time):
|
||||
|
||||
| kernel | best ms, runs 1 / 2 / 3 | best of the three, ms | G steps/s (best) | ns per step (best) | ops per step counted | step cost (ratio to `alu`, best) | bit-exact, 3 runs |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| alu | 5.039 / 4.918 / 4.873 | 4.873 | 881 | 1,190 | 5 | 1.00 | yes |
|
||||
| rotr (live) | 5.610 / 5.533 / 5.492 | 5.492 | 782 | 1,341 | 5 | 1.13 | yes |
|
||||
| shflx (live) | 4.196 / 4.259 / 4.223 | 4.196 | 1,024 | 1,024 | 6 | 0.86 | yes |
|
||||
| shl | 4.120 / 4.202 / 4.119 | 4.119 | 1,043 | 1,006 | 5 | 0.85 | yes |
|
||||
| shr | 4.296 / 4.194 / 4.319 | 4.194 | 1,024 | 1,024 | 5 | 0.86 | yes |
|
||||
| bfe (`extract_bits`) | 3.731 / 3.805 / 3.816 | 3.731 | 1,151 | 911 | 5 | 0.77 | yes |
|
||||
| bfec (C form) | 3.762 / 3.818 / 3.646 | 3.646 | 1,178 | 890 | 5 | 0.75 | yes |
|
||||
| andn | 3.680 / 3.732 / 3.676 | 3.676 | 1,168 | 898 | 5 | 0.75 | yes |
|
||||
| perm | 5.500 / 5.548 / 5.524 | 5.500 | 781 | 1,343 | 5 | 1.13 | yes |
|
||||
| popc | 4.261 / 4.223 / 4.262 | 4.223 | 1,017 | 1,031 | 5 | 0.87 | yes |
|
||||
| clz | 4.918 / 4.916 / 4.914 | 4.914 | 874 | 1,200 | 5 | 1.01 | yes |
|
||||
| sel | 3.718 / 3.831 / 3.718 | 3.718 | 1,155 | 908 | 5 | 0.76 | yes |
|
||||
| shfla (lane + 3) | 9.337 / 9.323 / 9.287 | 9.287 | 462 | 2,267 | 6 | 1.91 | yes |
|
||||
| dot4u (emulated) | 7.818 / 8.044 / 7.925 | 7.818 | 549 | 1,909 | 5 | 1.60 | yes |
|
||||
| dot4s (emulated) | 23.264 / 23.254 / 23.045 | 23.045 | 186 | 5,626 | 5 | 4.73 | yes |
|
||||
|
||||
Reading of the Mac rows. The run-to-run spread is under 4% on every row. The dot4 rows reproduce the 5 October figures (1.6x unsigned, 4.7x signed), which is the check on the method. A step cost under 1.00 means the family's op plus the glue is cheaper than the five-op reference chain: the reference's two registers are a tighter dependency than the three-register candidate chains, and Apple's shifts, extract, andn and select each cost about what an xor costs. Three rows cost more than the reference: `perm` (1.13: no byte-permute function in MSL; the `uchar4` swizzle compiles to shifts and masks, so a byte permute is emulated on Apple at about the price of the live `rotr`), `clz` (1.01) and `shfla` (1.91: a shuffle by a computed lane index costs 2.2x the live xor shuffle on Apple, `simd_shuffle` against `simd_shuffle_xor`; the second shuffle form is the one candidate Apple pays for). `mm8` as a chain on Apple is owed (Metal 4 `matmul2d`; this toolchain is Swift 5.8 without the tensor API).
|
||||
|
||||
**RTX 5090 (PC 2, 1ccfe586), CUDA**: PENDING the PC 2 job (`run-ca3-family-pc2-20261006`, published only after `/tmp/igneum-devnet/pc2-ca3.clear` and under the mkdir lock; the card to itself: `--stop-miners`, prover off for the run). The rows land in this entry when the closing report is read.
|
||||
|
||||
**RX 9070 XT (PC 1, ae432dc7), OpenCL**: OWED. PC 1 is the project lead's desk and not released today (the brief's rule); the OpenCL twin of the probe (`__builtin_amdgcn_*` paths for `v_bfe_u32`, `v_perm_b32`, `v_bcnt_u32_b32`, `v_cndmask_b32`, `ds_bpermute_b32`) is the next job on that card.
|
||||
|
||||
Consequences per tier, Mac rows (the hash is latency-bound by 128 dependent DRAM reads; a family at `W_new` = 4 points is about 4% of the 64 instructions, so these per-op costs bound a family's hash-rate cost and are not hash rates; the 5% rule of 1.13.2 is argued from them, not measured, until a family is live):
|
||||
|
||||
| Tier | What the rows mean | What is being done |
|
||||
|---|---|---|
|
||||
| Apple user (M-series laptop or desktop, the app's Metal worker) | six of the seven candidates (`shl`, `shr`, `bfe`, `andn`, `popc`, `sel`) cost at most the live `rotr` step; `clz` the same as the reference; `perm` 1.13x (emulated); `shfla` 1.91x, the only candidate over the live `shfl`'s cost by more than 2x on this card. At 4 points of 64 a 1.91x op costs under 1% of the program's ALU time, itself a small share of a latency-bound hash (approximate: argued, measured when live) | the proposed order puts `shfla` after the plain datapath families (R6), so Apple pays it last; `mm8` stays last |
|
||||
| NVIDIA user (8 to 32 GB card) | pending the 5090 rows above | the PC 2 job |
|
||||
| AMD user (RX 9070 XT, 16 GB) | owed: no row today | the PC 1 job when the desk is free |
|
||||
| A rig or a pool user | the same per-card figures; no family changes the dependent-read bound | nothing until a family is live |
|
||||
| A chip | every candidate but `mm8` is a 32-bit datapath structure (barrel shifter, byte crossbar, popcount tree, 32-lane shuffle crossbar: `docs/plans/counter-asic-3-reserve.md` section 3 names them with approximate areas); none is licensable as a block the way an int8 matrix unit is | the reserve order of that document |
|
||||
|
|
|
|||
263
docs/benchmarks/repro.md
Normal file
263
docs/benchmarks/repro.md
Normal file
|
|
@ -0,0 +1,263 @@
|
|||
# The reproducible benchmark package
|
||||
|
||||
6 October 2026 (built on the evening of 5 October). One command per platform reproduces the numbers on the bench table
|
||||
on an outsider's own machine on day one: `bench/repro.sh` (Linux, macOS) and `bench/repro.ps1` (Windows), shipped as
|
||||
`igneum-repro-<tag>.tar.gz` and `.zip` next to the public downloads (`packaging/ota/publish-public.sh --repro`, aliases
|
||||
`/public/igneum-repro.tar.gz` and `/public/igneum-repro.zip`; not deployed tonight). Package v0.1.0-repro was built from
|
||||
commit `39141f5` plus the package sources of branch `repro-bench`, commit `66ccd25` (the shipped binaries were built from that
|
||||
tree before the commit; the script fixes after the runs change no binary),
|
||||
and run end to end on the three project machines. The deltas against the bench log are below.
|
||||
|
||||
Why it exists: `docs/evidence.md` has 30 claims and none is reproduced externally. The ladder from "tested by the team"
|
||||
to "reproduced externally" needs "the command published and a third party's run with the same result". This is the
|
||||
command. The reward for running it is section 8.1 of `docs/benchmarks/proving-e2e.md`, quoted unchanged below.
|
||||
|
||||
## 1. What one run does
|
||||
|
||||
| Step | What runs | Binary | What it reports |
|
||||
|---|---|---|---|
|
||||
| 1 | The machine | the workers' `--list` | OS, CPU, memory; every GPU each worker sees (name, driver, memory) |
|
||||
| 2 | The lottery hash vectors of the published genesis pack (`proto-cuda/packs/igneum-genesis-mh`: seed `igneum-genesis`, day 2026-10-03, generator 2, program id `bcc1248b10cc90f2`, memory-hard 1 GiB dataset) | `igneum-pow check-pack` on the CPU; `igneum-worker-cuda --bench`, `igneum-worker-opencl --bench --pack`, `igneum-bench --pack` on every GPU | bit-exact or not: 3 warps, 96 lanes, the 256 MiB cache FNV, on the CPU through the Rust interpreter and on every GPU through its own compiler (NVRTC, the vendor's OpenCL compiler, Metal) |
|
||||
| 3 | The hash benchmark, 120 s per card at 1 GiB, `--batch-log2 24` | the same workers, `--seconds 120` | MH/s over the summed dispatch time, the hashes done, the fingerprint of the first 2^24 outputs at base nonce 0 (FNV-1a 64 over 16.7 million hashes: equal on two machines means every one of them agreed) |
|
||||
| 4 | The random-read probe at 4, 64, 256 and 1024 MiB | `--memprobe` | dependent random 4-byte reads per second (the hash's access pattern) and their latency at 256 lanes, eight independent chains, 16 and 64-byte lines, the coalesced stream, an integer chain |
|
||||
| 5 | The chip-resistance sweep: the same program at 4, 64, 256 and 1024 MiB, 5 batches each | `--bench --dataset-mib N` (`--sweep` on Metal) | MH/s per size; in-cache rate over the 1 GiB rate; the hash's share of the card's random-read ceiling at 1 GiB (MH/s x 128 loads against the chase) |
|
||||
| 6 | One fixture shard proven with the pinned guest and verified (optional) | `igneum-prove-host --mode shard --shard 0`, then `--mode verify` | execute, core and compressed proof times, proof bytes, VERIFIED or not; skipped with the reason when there is no 12 GB NVIDIA card or no prover host |
|
||||
| 7 | The result | the script | `results/igneum-repro-<os>-<time>.json` (format `igneum-repro-1`: machine, every binary's sha256, the pack files' sha256, every command run with its exit status, the rows in the bench table's own shape, the tolerances) and the same as a table in `.md`, plus the full log |
|
||||
|
||||
The workers were extended for this (the flags are in every worker now, where the read-width experiment of 5 October had
|
||||
put `--bench` and `--memprobe` on the CUDA worker and `--memprobe` on the OpenCL worker, on branches): `--list`, `--bench
|
||||
--pack <dir> --seconds N --dataset-mib N`, `--memprobe` at 4, 64, 256 and 1024 MiB with one `RESULT` line per size and
|
||||
a `DEVICE` line per card; the Metal worker got `--list`, `--pack` (the published vectors against the GPU's warps),
|
||||
`--seconds`, `--sweep` and `--memprobe`; `igneum-pow` got `check-pack`. The pack reader accepts a string-seed pack
|
||||
(the genesis pack's 14-byte seed; the chain's are 32 bytes), the same change the read-width branch made.
|
||||
|
||||
What a run needs: no secrets, no node, no wallet, no network. The CUDA worker needs the NVIDIA driver only on Windows
|
||||
(NVIDIA's runtime compiler is in the package, `THIRD-PARTY.md`) and the driver plus `libnvrtc.so.12` on Linux. The
|
||||
OpenCL worker needs the vendor's OpenCL. Nothing here earns anything.
|
||||
|
||||
## 2. The three runs of 5 October 2026 against the bench log
|
||||
|
||||
Package v0.1.0-repro, three machines asked for, two run tonight (the Mac in full, PC 2 with its 5090 shared with the live miner): PC 1 (ae432dc7: RTX 5090, RX 9070 XT, gfx1036) went
|
||||
down at 22:31:06 UTC on 5 October before its second slot (its first run at 21:04 UTC ended in a PowerShell parse error
|
||||
within a second per card, section 7) and nothing reaches it until the morning; its run is owed. The Mac ran under the
|
||||
measure lock (every build and simulation slot held; load average 6.19 at the start, 4.15 at the end). PC 2's run is in the
|
||||
coordinator's queue (section 2.2 is filled from it; "pending" until then).
|
||||
|
||||
### 2.1 The Mac (Apple M5 Max, 40 GPU cores, 64 GB, macOS 26.6.2), run 2fdd7365db4a5deb, 21:51 UTC, 5 October 2026
|
||||
|
||||
Result files: `docs/benchmarks/repro-2026-10-06/mac/`. 7 min 52 s for the two backends (Metal 120 s, Apple OpenCL 30 s as
|
||||
the cross-check, each with the sweep and the probe table).
|
||||
|
||||
| Number | The package | The bench log | Delta | Reading |
|
||||
|---|---|---|---|---|
|
||||
| Vectors, CPU (Rust interpreter) | 96/96 lanes, cache FNV `48c4f5bf24166b2e` | 96/96 (4 October 2026, generator version 2 adopted) | exact | |
|
||||
| Vectors, Metal | 96/96, program id `bcc1248b10cc90f2` matches | 96/96 (4 October 2026) | exact | |
|
||||
| Vectors, Apple OpenCL | 96/96 | 96/96 (proto-opencl README, 4 October 2026) | exact | |
|
||||
| Fingerprint of 2^24 outputs at base 0 | `25f96e7dce90bd4e` on Metal and on Apple OpenCL | none at 2^24 for this pack (the README's `f2a95d5bb84d961e` is at 2^13) | new: two compilers agree over 16.7 million nonces | the number a second machine must match |
|
||||
| Hash MH/s at 1 GiB, Metal, 120 s, GPU time | 27.674 (194 dispatches of 2^24, 3.25 G hashes) | 27.9 through Apple OpenCL on this pack (README, 5 x 2^24, wall); 26.7 mining through Metal on the live devnet (4 October) | -0.8% against the bench, +3.6% against mining | the bench row |
|
||||
| Hash MH/s at 1 GiB, Apple OpenCL, 30 s, wall | 26.875 | 27.9 (the same README figure) | -3.7% | outside the 3% tolerance against a 5-batch wall-time figure taken on 4 October; the cross-check path is 2.9% under Metal here, where the 3 October runs had the two Apple paths equal (45.2 against 45.0 on the version 1 program). Not the card's row |
|
||||
| CPU verify, ms per 32-lane warp (avg of 50, one performance core) | 0.611 (cold 0.625) | 0.631 (4 October, version 2 units) | -3% | 16x under the 10 ms gate |
|
||||
| Sweep, MH/s at 4 / 64 / 256 / 1024 MiB, Metal | 275.2 / 103.2 / 56.6 / 27.8 | 569 / 183 / 94 / 44 (3 October, version 1 program with 104 loads, closed-form dataset) | not comparable: the generator changed | in-cache over 1 GiB 9.9x against 12.9x on version 1 |
|
||||
| Random reads at 1 GiB, G loads/s (chase) | 3.40 Metal, 3.40 Apple OpenCL | none for Apple before this (4.6 G was derived from the version 1 hash rate on 3 October) | new | the hash does 27.67 x 128 = 3.54 G loads/s, 1.04 of the single-chain chase: on this chip the hash is at the random-read ceiling the probe sees |
|
||||
| Dependent-load latency, ns at 256 lanes | 522 Metal (GPU time), 1,273 Apple OpenCL (wall, launch included) | none | new | |
|
||||
| Stream, GB/s at 1 GiB | 546 Metal, 515 Apple OpenCL | 427 GB/s dataset fill (3 October, a write) | reads 28% above the write figure | |
|
||||
| Proving | skipped: the GPU prover needs a 12 GB NVIDIA card | | | as designed |
|
||||
|
||||
### 2.2 PC 2 (RTX 5090, gfx1036, Windows 11), 1ccfe586
|
||||
|
||||
Run job run-repro-pc2-20261006 (fetch-repro-pc2-20261006 first), 00:44 to 01:09 UTC, 6 October 2026, on Igneum Miner
|
||||
0.3.11, one card at a time through `bench/jobs/pc-repro.ps1`. What came back: every RESULT line in the log intake (the
|
||||
three per-card transcript logs are in `docs/benchmarks/repro-2026-10-06/pc2/`; the JSON and the markdown were never
|
||||
written, section 7). The 5090 block ran BESIDE THE APP'S LIVE CUDA WORKER: the card-off POST was accepted, but
|
||||
`api/state` answered `{}` (a 0.3.11 defect on a proving machine) so nothing confirmed a stopped worker, and nvidia-smi
|
||||
showed 10,176 MiB in use before the bench and the card at 348 W after it. Every 5090 rate below is therefore a shared-card
|
||||
figure, about half of what the card does alone (131 to 137 MH/s at this batch in every other run of the night), and is
|
||||
not the card's row. The checks (vectors, fingerprints, the proof) stand regardless of the sharing. The gfx1036 block is the
|
||||
integrated chip's own figure (3.312 MH/s, where it mines at 3.3 in the app).
|
||||
|
||||
| Number | The package (PC 2) | The bench log | Delta | Reading |
|
||||
|---|---|---|---|---|
|
||||
| Vectors, CPU (Ryzen 7 9800X3D) | 96/96, cache FNV `48c4f5bf24166b2e` | 96/96 on the M5 Max | exact | first CPU of the second architecture |
|
||||
| Vectors, RTX 5090 through CUDA (NVRTC, sm_120) | 96/96, program id matches | the version 2 pack had not run on real NVIDIA silicon (evidence row 4) | exact: closes that gap | |
|
||||
| Vectors, RTX 5090 through NVIDIA OpenCL | 96/96 | 96/96 on the version 1 pack (3 October) | exact | |
|
||||
| Vectors, gfx1036 through AMD OpenCL | 96/96 | 96/96 on the version 1 pack (3 October) | exact: the version 2 pack on real AMD silicon | |
|
||||
| Fingerprint of 2^24 outputs at base 0 | `25f96e7dce90bd4e` on CUDA, NVIDIA OpenCL and AMD OpenCL | `25f96e7dce90bd4e` on Metal and Apple OpenCL (the Mac run) | exact across five compilers and three vendors | the cross-vendor claim over 16.7 million nonces |
|
||||
| Hash MH/s at 1 GiB, 5090, CUDA, 120 s | 62.412 (447 dispatches, 7.50 G hashes), beside the live miner | 139.7 with the card to itself on a version 2 pack (4 October, variant racing); 124.2 mining in the app | not comparable: shared card | re-run owed with the worker confirmed stopped |
|
||||
| Hash MH/s at 1 GiB, 5090, NVIDIA OpenCL, 30 s | 62.283, beside the live miner | 219.6 against 229.0 CUDA on the version 1 pack (3 October) | not comparable | the two paths agree with each other to 0.2% even shared |
|
||||
| Hash MH/s at 1 GiB, gfx1036, AMD OpenCL, 120 s | 3.312 (24 dispatches) | 4.38 on the version 1 pack (104 loads, 3 October); 3.3 mining on the live devnet (4 October) | +0.4% against mining; the version 1 figure is another program | the card's row |
|
||||
| Sweep 4 / 64 / 256 / 1024 MiB, 5090 CUDA (shared) | 374 / 367 / 74 / 62 | 1,340 / 1,353 / 270 / 229 (3 October, version 1, card alone) | shape only: in-cache over 1 GiB 6.0x against 5.8x | the L2 edge is the same |
|
||||
| Sweep, gfx1036 | 3.72 / 3.38 / 3.32 / 3.31 | none | new | flat: the integrated chip is not cache-bound at any size, its memory is the system's |
|
||||
| Random reads at 1 GiB, 5090 (chase, best lanes) | 17.5 G loads/s CUDA (wall), 18.2 NVIDIA OpenCL (event); latency 415 ns (OpenCL event) | 23.7 G derived from the version 1 hash rate (3 October); about 16 to 18 G implied by the 9070 XT entry's 6.6x (5 October) | inside the implied range, shared card | |
|
||||
| Random reads at 1 GiB, gfx1036 | 0.458 G loads/s, 566 ns | none | new | the hash does 3.312 x 128 = 0.424 G loads/s: 0.93 of the chase; at the ceiling like the 9070 XT (0.92) |
|
||||
| Stream at 1 GiB, 5090 | 1,662 GB/s (OpenCL event), 339 (CUDA wall, the launch dominates) | 1,638 GB/s dataset fill (3 October) | +1.5% | the memory clock was in its full state |
|
||||
| Integer chain, 5090 | 13.4 T int ops/s (CUDA wall), 45.2 (OpenCL event) | none | | the wall figure is launch-bound; the event figure is the card's |
|
||||
| Proving, block-338-shard1 shard 0, the live host `/opt/igneum/igneum-prove-host` (pinned shard id `0x2b1a81cb...`) | execute 1.48 s (60.4 M cycles), core 20.2 s (18.1 MB, verify 0.60 s), compressed 33.9 s (1.27 MB, verify 0.038 s), VERIFIED; two tampered witnesses REJECTED | the proving agent the same night: 10.8 to 11.4 s compressed alone, 33 s with the miner running | the "with the miner" figure | consistent with the shared card |
|
||||
| Job size, 5090 CUDA (shared): 2^21 against 2^24 | genesis pack 60.0 against 63.2 MH/s; live devnet pack 60.9 against 63.4 | the app mines at 2^21: 114.0 wall against 116.0 inside jobs at 22:40 UTC (PC 2's STATUS lines, the consequences reviewer) | 5% lower at 2^21 here, 1.8% wall-to-inside in the app | the per-job cost at 2^21 is about 5% on a shared card; the 15% gap the reviewer found between bench and app is not this alone. The rows are PC 2's; PC 1's are owed |
|
||||
|
||||
### 2.3 PC 1 (RTX 5090, RX 9070 XT, gfx1036, Windows 11), ae432dc7
|
||||
|
||||
Owed. The first job (run-repro-pc1-20261005, 21:04 to 21:09 UTC) switched each card off and on for nothing: repro.ps1
|
||||
had three PowerShell parse errors (two missing parentheses on the wsl.exe lines, one `(if ...)` used as an expression),
|
||||
found only by the Windows parser on the PC because the Mac has no PowerShell. The fix, the parse gate (the PC job now
|
||||
parses the package script before any card is touched, and the bench/ folder is in the Windows PowerShell 5.1 parse job
|
||||
of `.github/workflows/windows.yml`) and the second package were ready at 21:45 UTC; PC 1 went down at 22:31:06 UTC before
|
||||
its second slot. One fact from the first job: the RX 9070 XT (gfx1201) was on the bus and switched off and on by the
|
||||
app at 21:06 UTC, where the coordinator's queue had it absent since 20:40 UTC.
|
||||
|
||||
## 3. The deltas, and what they mean for each user tier
|
||||
|
||||
What each number means for each user tier, and what the package does about it (the standing rule of 5 October 2026):
|
||||
|
||||
| Number | Home miner, one card (8, 12, 16, 24 or 32 GB) | Rig | Pool user | What the package does |
|
||||
|---|---|---|---|---|
|
||||
| The bench needs 1.3 GB of device memory (1 GiB dataset, 256 MiB cache, a 128 MiB output buffer at 2^24) | every tier runs it, 8 GB included; a 4 GB card or an integrated chip with 4 GB shared runs it too | one card at a time, `--only`, so the rig keeps mining on the rest | the same | reports the free memory it found when a dataset does not fit |
|
||||
| 120 s per card, plus 3 short sweep sizes and the probe (about 4 min per card) | a few minutes with the card idle | a 6-card rig is 25 min, and `--only` splits it | | `--seconds` and the skip flags |
|
||||
| The hash share of the random-read ceiling (1.04 on the M5 Max, PC 2 pending) | the number that says a card is at its memory's limit, not the kernel's: a card under about 0.9 has a worker problem, not a card problem | the same per card | | the probe table is in every result, so a low share names the step |
|
||||
| Apple OpenCL 2.9% under Metal on the same chip | an Apple user on the app gets the Metal rate (the app's worker is Metal); the OpenCL figure is a compiler cross-check, never the app's | | | the OpenCL row is marked cross-check and makes no bench-table row |
|
||||
| CPU verify 0.611 ms per warp on an M5 Max performance core | the chain's gate is 10 ms on any core; a 2019-class laptop core is still unmeasured (evidence row 5, O-1.14) | | a pool verifying shares has 16x margin on this core | `check-pack` prints the figure on any machine that runs the package; the first slow core to run it answers O-1.14 |
|
||||
| The proving step: skipped on macOS by design; on NVIDIA it needs a built prover host and a 12 GB card by the gate, and the 5090 run of the proving agent tonight peaked at 28.3 GB (bench-log, 5 October, "the 5090 alone proves block-338-shard1 shard 0 compressed in 10.8 to 11.4 s at a 28.3 GB peak") | a 12 GB, 16 GB or 24 GB owner cannot run step 6 today: the gate lets them start it and the prover would fail on memory. The litepaper's "a 12 GB card proves a shard" stays "designed" (evidence row 16) until the prover's peak is under 12 GB | | | step 6 says why it skipped; the next package raises the gate to the measured peak (28 GB) until the prover fits 12 GB, so nobody's run fails late. Filed below as owed work |
|
||||
| One PC run per night at most on the project's own PCs | | | | the package itself takes minutes; the project's PCs are shared by many agents and PC 1 was down tonight, so the project's three-machine table is not complete on the first night. The outsider's run does not have that constraint |
|
||||
|
||||
PC 2 (6 October 2026): the 5090's package numbers are void as the card's figures (beside the live miner) and stand only
|
||||
as checks; the morning re-runs the 5090 block alone with the worker confirmed stopped by nvidia-smi's compute-apps list
|
||||
(`bench/jobs/pc-repro.ps1` does that now, and labels a run "UNCONFIRMED" when it cannot). For a 5090 owner the package's
|
||||
own figure is still owed; for a gfx1036 owner the figure is 3.31 MH/s, flat across dataset sizes, at 0.93 of the chip's
|
||||
random-read ceiling. For a proving owner: the live host proved and verified the fixture through the package's step 6 on a
|
||||
32 GB card even with a miner on it (33.9 s compressed); the 12 GB gate stands as written in the table above.
|
||||
|
||||
## 4. Tolerances and how a result becomes a row
|
||||
|
||||
| Number | Tolerance | Why this width |
|
||||
|---|---|---|
|
||||
| Vectors (3 warps, 96 lanes, cache FNV) | exact | a hash either matches or it does not |
|
||||
| Fingerprint of 2^24 outputs at base 0 | exact across machines at the same `--batch-log2` | the same |
|
||||
| Hash MH/s at 1 GiB (120 s) | 3% | two 120-s runs of the same card on this evening's machines agree within about 1% with the card to itself; 3% leaves room for a driver version and a warmer card |
|
||||
| Sweep MH/s at 4, 64, 256 MiB | 10% | 5 batches each, not 120 s; the in-cache sizes are the noisiest (a few hundred ms per batch) |
|
||||
| Random reads at 1 GiB (G loads/s), stream GB/s | 10% | best of 3 at 8 lane counts; the memory clock's power state moves these |
|
||||
| Dependent-load latency (ns at 256 lanes) | 15% | one warp, a few hundred ns, timer resolution |
|
||||
| CPU verify per warp | under 10 ms (the chain's gate), no tolerance on the figure itself | any core must verify a warp inside the gate; the published cores are under 1 ms |
|
||||
| Proof times (execute, core, compressed) | 25% | SP1's GPU prover varies run to run with the card's state and the server's warm-up |
|
||||
|
||||
The path from a submitted file to the bench table (`site/miner-bench.json`, rendered at `/miners` by `site/build.mjs`):
|
||||
|
||||
1. `node bench/ingest.mjs check <result.json>` prints the agreement table: every number against the team's reference
|
||||
for that card model (`bench/reference.json`, which names the bench-log entry or this document's run behind every
|
||||
reference), with the tolerance and the delta.
|
||||
2. `node bench/ingest.mjs add <result.json> --by "reproduced externally"` appends one row per card with that label when
|
||||
every number agrees, and refuses the label (recording "submitted" with the deltas in the note) when one does not. A
|
||||
failing run is as public as a passing one (proving-e2e.md 8.2 step 4). The site build then fails on any private
|
||||
string (`site/forbidden-strings.txt`), so a hostname in a note never reaches the page.
|
||||
3. The row's `source` is `repro:<run id>`, the random id the script drew; the result file is kept under
|
||||
`docs/benchmarks/repro-<date>/`, as the three of tonight are.
|
||||
|
||||
The relay's console (`relay/api/console.mjs`, `results`) shows bench entries synced from `docs/bench-log.md` by
|
||||
`tools/console.mjs sync-bench`; a submitted result enters there through the bench-log entry that records it, after the
|
||||
table row. The relay's `bench` role (`relay/lib/relay.mjs`) is for a machine of ours running benches, not for outsiders.
|
||||
Until the repository is public (the public testnet), results arrive by email to the address on igneum.network; after it,
|
||||
as an issue with the JSON attached.
|
||||
|
||||
## 5. Operator instructions
|
||||
|
||||
See `bench/README.md`, shipped in the package. In short: unpack, stop mining on the card, run the one command, send
|
||||
`results/igneum-repro-<os>-<time>.json` back, say which card was idle. About 4 minutes per card plus the optional
|
||||
proving step. The file carries no hostname, user name or address; it carries the GPU and CPU models, the OS and the
|
||||
driver version. `--only backend:index` runs one card; `--no-prove`, `--no-probe`, `--no-sweep` skip steps.
|
||||
|
||||
Our own machines run it the same way, with one difference the script does not know about: the card under test is
|
||||
switched off in the Igneum Miner app first (`bench/jobs/pc-repro.ps1` does it through the app's `api/cards`, one card at
|
||||
a time, and restores it), and on the Mac the run takes the measure lock.
|
||||
|
||||
## 6. The reward terms (proving-e2e.md section 8.1, unchanged)
|
||||
|
||||
### 8.1 What counts as unrelated
|
||||
|
||||
Three operators are unrelated when every row holds for every pair:
|
||||
|
||||
| Test | Requirement |
|
||||
|---|---|
|
||||
| Person | Different natural or legal persons; none is the project, an agent of it, or paid by it for the run (a published fixed reproduction reward, equal for everyone and announced before the run, is allowed and disclosed) |
|
||||
| Hardware | Bought separately; no shared host, card, rack or power meter |
|
||||
| Network | Different autonomous systems, verified by the IP in the published logs; not the same residential ISP account |
|
||||
| Location | Different physical sites |
|
||||
| Software | The same published release by hash; nobody receives a private build |
|
||||
| Money | No payment, loan or equipment between them or from the project, beyond the disclosed reproduction reward |
|
||||
|
||||
An operator declares each row in their report and signs it with the vote key that mined on the devnet under the same fingerprint, so a report is tied to a key with a history.
|
||||
|
||||
The amount is the challenge reward row of `docs/plans/funding.md` (USD 1,000 per operator per workload set, three
|
||||
operators, about USD 3,000 per campaign, approximate; not funded as of 6 October 2026). Reproduction is asked for without
|
||||
a reward, which the standard allows. The result file of this package is signed by nothing; the signature with the vote
|
||||
key of 8.1 is the operator's own step, on the file.
|
||||
|
||||
## 7. What is unverified
|
||||
|
||||
| Item | Why | What closes it |
|
||||
|---|---|---|
|
||||
| PC 1's run (RTX 5090, RX 9070 XT, gfx1036 on Windows) | PC 1 went down at 22:31:06 UTC before its slot; the first job hit the parse errors | the morning's slot: `bench/jobs/pc-repro.ps1` as a `run` job after the package fetch, one card at a time |
|
||||
| The Linux binaries (`bin/linux-x86_64`) | cross-compiled with zig on the Mac, loaded nowhere tonight (no Linux GPU host; HiveOS is the first) | a run on a Linux box with an NVIDIA or AMD driver; `repro.sh` is the same script the Mac ran |
|
||||
| The CUDA worker's `--bench`, `--memprobe` and `--list` on a real card | tonight's only CUDA runs were the Mac's emulation (`--list`, `--bench` on the genesis pack: self-test PASS, fingerprint `e7d68ec2a49d0671` at 2^14, equal to Apple OpenCL's at 2^14) and PC 1's run that never reached the worker | PC 2's slot (pending) and PC 1's morning run |
|
||||
| `repro.ps1` end to end | parsed only by the Windows parser on PC 1 after the fix; never run to completion on a PC | PC 2's slot |
|
||||
| The proving step (step 6) | never run through the script; on PC 2 it will use the live host `/opt/igneum/igneum-prove-host` as the app's WSL user with the live prover off | PC 2's slot; then the memory gate against the measured 28.3 GB peak |
|
||||
| A second machine of the same model | the tolerance table has one Apple machine; "reproduced externally" needs an unrelated operator's run, and the repository is private until the public testnet | the publish step (`publish-public.sh --repro`, not run tonight) and the first outside result |
|
||||
| Apple OpenCL 3.7% under the README's 27.9 MH/s | one 30-s wall-time run against one 5-batch run of 4 October; either could be the odd one | a second 120-s run of each path on the Mac with the card to itself |
|
||||
| The random-read probe on Apple as a ceiling | the hash exceeds the single-chain chase by 4% on the M5 Max, so on Apple the chase at 4 M lanes is not the ceiling the 9070 XT entry took it for | more lanes in flight, or the eight-chain probe (3.50 G here) as the Apple ceiling; a probe question, not a hash one |
|
||||
| PC 2's 5090 rows | the card-off step could not confirm a stopped worker (`api/state` answered `{}`), and the numbers are half the card's | the morning's 5090-only run on PC 2 with the compute-apps confirmation; until then the rows are "beside the live miner" |
|
||||
| PC 2's result files | `repro.ps1`'s close threw `System.OutOfMemoryException` in ConvertTo-Json and gave Set-Content an empty path: PowerShell variable names are case-insensitive, the result table `$Cpu` held `$cpu` (its own name string) and so contained itself, and the markdown lines `$md` wiped the markdown path `$Md`. Fixed (distinct names; `bench/jobs/ps-case-check.sh` fails CI on any case-only pair, shown to fire on a bad file); the RESULT lines in the intake and the three transcript logs are the record | the morning's run writes the files |
|
||||
| The prover on PC 2 | the first script read the prover's state from `api/state` too, saw nothing, and left the live prover off from 00:49 to the restore job (run-prover-on-pc2-20261006); the script now reads settings.json and restores unconditionally | done tonight; the rule in the job |
|
||||
| The reward | the amount and payer are `docs/plans/funding.md`, not funded; the terms are quoted unchanged | the project lead's decision |
|
||||
|
||||
## 8. The vendor-share metric (Counter ASIC 3.0 item 7)
|
||||
|
||||
6 October 2026, branch `ca3-reserve`. The history audit's rank 7 (`docs/analysis/asic-resistance-history.md` section 4.3): a one-vendor fleet is a softer form of chip capture (lesson 8: Equihash's NVIDIA tilt, Ethash's balance), so the share of hash rate by vendor is published beside the benchmark and the AMD gap gets a 3.0 target. The observer computes it every 60 s into `live_state.vendor_share` (`tools/observer/vendor-share.mjs`; the one hook in `observer.mjs`; `node --test tools/observer/vendor-share.test.mjs` fires on a known-finished and a known-failed case; `node tools/observer/vendor-share.mjs --dry` prints the live reading and writes nothing).
|
||||
|
||||
Definition. Hash rate by vendor (nvidia, amd, apple, intel, unknown), two readings, each named for what it is:
|
||||
|
||||
| Reading | What it is | What it covers | What it misses |
|
||||
|---|---|---|---|
|
||||
| fleet-reported | the sum of `now=X MH/s wall` from the newest STATUS line of every miner worker that uploaded to the log intake in the last 10 minutes, by the vendor in the worker's label (`miner-nvidia-`, `miner-amd-`, `miner-mac-`; `miner-other-` by the app's `cards:` line) | only machines that upload to the intake: the project's own fleet and any app install carrying the intake key | the rest of the network entirely; a worker that stopped uploading |
|
||||
| chain-attributed | each miner id's (`vote_key_hash`) share of the BLUE blocks of the last 10 minutes times the node's network hash-rate estimate, attributed to the vendor of the fleet worker that logged that vote key (`identity N '...' vote_key_hash=` in its first uploads), else "unknown" | the whole network's blocks | the vendor of any miner that is not a reporting worker (that is the "unknown" row, the honest hole); pending and red blocks are left out |
|
||||
|
||||
Today's devnet, read-only from the live tables (dry mode, 07:36 UTC, 6 October 2026; nothing written, the live observer not restarted; the hook goes live with the merge):
|
||||
|
||||
| Vendor | fleet-reported MH/s (workers) | chain-attributed share of blue blocks (10 min) | chain-attributed MH/s at the 135.8 MH/s estimate | Note |
|
||||
|---|---|---|---|---|
|
||||
| NVIDIA | 120.61 (1: PC 2's RTX 5090, with the prover on the card) | 0.986 (552 of 560) | 133.86 | PC 1's RTX 5090 was off the network at the reading (the project lead's desk) |
|
||||
| AMD | 0 (0) | 0 | 0 | the RX 9070 XT is on PC 1, off at the reading |
|
||||
| Apple | 0 (0) | 0 | 0 | the M5 Max is paused for measurements |
|
||||
| Intel | 1.69 (1: the Windows laptop's UHD Graphics) | 0.014 (8 of 560) | 1.94 | the first outside machine, which carries the intake key |
|
||||
| unknown | 0 | 0 | 0 | coverage 1.00: every blue block of the window came from a reporting worker |
|
||||
|
||||
At 07:21 UTC, before PC 1 went off, the same tables read 243 MH/s and 18 miner ids; the window then was NVIDIA about 0.9 (PC 1's and PC 2's 5090s) and AMD about 0.01 (the 9070 XT's first 3 blocks after its restart). The devnet today is the project's own machines, so the two readings agree; on a public testnet the chain-attributed "unknown" row is the number that matters, and the fleet-reported row is only the project's own share.
|
||||
|
||||
Rate per watt and per pound, from the bench rows that exist (every price approximate, UK list prices from memory, 6 October 2026):
|
||||
|
||||
| Card | MH/s (class v3) | Power | MH/s per W | Price (approximate) | MH/s per £ | Source |
|
||||
|---|---|---|---|---|---|---|
|
||||
| RTX 5090 | 136.1 | about 326 W (328.6 W peak with the prover, 5 October; the app's stability line today: PC 1 p95 307 W mining alone, PC 2 p95 338 W with the prover) | 0.42 (at 326 W); 0.44 at 307 W | about £2,000 | 0.068 | bench-log "Counter ASIC 2.0, the numbers"; miner_logs stability lines |
|
||||
| RX 9070 XT | 18.6 to 19.2 | OWED: no measured figure (the app logs no stability line for the Radeon; PC 1 not available today); the board's 304 W TBP is a list figure, approximate | about 0.06 at the list TBP (approximate) | about £600 | 0.031 | bench-log; price approximate |
|
||||
| Apple M5 Max | 27.9 | OWED: no powermetrics figure in the bench-log | not computed | the laptop, about £3,500 and up; not a GPU purchase | 0.008 (approximate, against the laptop price) | bench-log; price approximate |
|
||||
| Intel UHD (laptop iGPU) | 1.85 (the US laptop's STATUS line) | not measured | | | | miner_logs |
|
||||
|
||||
The 2.0 record put AMD at 2.2x worse per pound and 4.9x worse per watt than the 5090 (approximate); the rows above read 2.2x per pound and about 7x per watt at the list TBP, so the per-watt figure is the one the owed measurement must settle.
|
||||
|
||||
Consequence per tier:
|
||||
|
||||
| Tier | What a 90% NVIDIA share means | What the project does |
|
||||
|---|---|---|
|
||||
| An AMD owner (RX 9070 XT, 16 GB) | at equal network difficulty the card earns 19 / 136 = 0.14 of a 5090's income at about 0.3 of its price, so about half the income per pound; the share itself does not change that, but a 90% NVIDIA network sets the difficulty by NVIDIA's rate, and a chip built against one vendor's memory system would hit AMD owners first (lesson 8) | the line-width question stays open per the plan (`docs/plans/counter-asic-3.md`: w16 closes nothing, w64 makes the 5090 bandwidth-bound); the 3.0 target for the AMD gap is the per-read gap (2.4 G against 17.5 G dependent reads per second) measured per driver release, with the first target a 2x closing by the public testnet or a stated reason it cannot close; the vendor share is published so the tilt is visible |
|
||||
| An Apple owner (M-series) | 27.9 MH/s is 0.2 of a 5090 on a machine bought for other reasons; the share means the Mac is a minority that no family in the reserve may cost more than 5% (1.13.2) | the reserve order of item 6 puts the two families Apple pays for (perm, shfla) at R1 and R3 with the costs measured; the 5% rule is checked live before each unlock |
|
||||
| An NVIDIA owner | the majority sets the difficulty; a chip that beats NVIDIA is the only chip that matters | the chip model and the bounty (bench-log "Counter ASIC 2.0, the numbers") |
|
||||
| A rig or a pool | a pool's share by vendor is what the chain-attributed reading cannot see (one vote key per pool) | pools publish their own vendor mix or appear as "unknown"; the detector (item 4) reads the per-program spread per card model |
|
||||
| The project | a fleet-reported reading above the chain-attributed one means a reporting worker is not finding blocks (a stuck worker, the 4 October class); below it means miners outside the intake | the daily report carries both readings from the merge on |
|
||||
|
||||
## 9. Files
|
||||
|
||||
| File | What |
|
||||
|---|---|
|
||||
| `bench/repro.sh`, `bench/repro.ps1` | the operator commands |
|
||||
| `bench/README.md` | the package's README (operator instructions, tolerances, how to send a result back) |
|
||||
| `bench/make-package.sh` | builds the package from a tagged commit with the existing cross-build scripts (zig for Linux, mingw for Windows, swiftc and cc for macOS, cargo for the CPU tool) |
|
||||
| `bench/ingest.mjs`, `bench/reference.json` | the agreement check and the bench-table ingestion; the references |
|
||||
| `bench/jobs/pc-repro.ps1`, `bench/jobs/collect-results.mjs` | the PC job (one card at a time through the app) and the Mac-side extraction of its result files |
|
||||
| `docs/benchmarks/repro-2026-10-06/` | the three result files of tonight, their tables and logs |
|
||||
| `packaging/ota/publish-public.sh --repro` | the public downloads entry (tar.gz, zip, their sha256, the two aliases, the downloads index) |
|
||||
137
docs/plans/counter-asic-3-reserve.md
Normal file
137
docs/plans/counter-asic-3-reserve.md
Normal file
|
|
@ -0,0 +1,137 @@
|
|||
# Counter ASIC 3.0 item 6: the reserve ordered by chip-unfriendliness (PROPOSED)
|
||||
|
||||
6 October 2026, branch `ca3-reserve`. The history audit's rank 6 (`docs/analysis/asic-resistance-history.md` section 4.3: families that force a full 32-bit datapath per lane first, `mm8` last, because int8 matrix blocks are licensable IP at every node and Least Authority's ProgPoW suggestion 5 was "watch ML hardware"). Everything here is PROPOSED text for spec 1.13.2: the order and the weights are a decision for the project lead, and nothing here touches the generator, a vector, the manifest or a consensus parameter. The measurements are in `docs/bench-log.md`, entry "6 October 2026, Counter ASIC 3.0 item 6"; this document carries the readings, the ISA facts, the chip structures and the proposed spec text.
|
||||
|
||||
## 1. The step cost per family per card
|
||||
|
||||
Method (the int8 analysis's, `docs/analysis/int8-matrix-family.md` section 4): a dependent chain of one op per step per lane, 1,048,576 lanes x 4,096 steps, best of 3 dispatches, three runs, bit-exact against a CPU reference on two whole 32-lane warps. The step cost is the chain's best time over the add-xor-rotate chain's (`alu`, the live reference, 5 ops per step). Every candidate chain is the family's one op plus the same four glue ops. Sources `proto-metal/family-probe.swift`, `proto-cuda/family-probe.cu`, job `tools/ca3-reserve/pc2-family-probe.ps1`.
|
||||
|
||||
| Family (1.13.2) | Probe row | M5 Max, Metal: step cost (best ms; G steps/s), load 7.64 | RTX 5090, CUDA (PC 2) | RX 9070 XT, OpenCL (PC 1) |
|
||||
|---|---|---|---|---|
|
||||
| reference: add-xor-rotate | alu | 1.00 (4.873; 881) | PENDING the PC 2 job | OWED (PC 1 not released today) |
|
||||
| reference: live `rotr` | rotr | 1.13 (5.492; 782) | PENDING | OWED |
|
||||
| reference: live `shfl` (lane XOR mask) | shflx | 0.86 (4.196; 1,024) | PENDING | OWED |
|
||||
| variable left shift | shl | 0.85 (4.119; 1,043) | PENDING | OWED |
|
||||
| variable logical right shift | shr | 0.86 (4.194; 1,024) | PENDING | OWED |
|
||||
| bit-field extract, immediate offset and width | bfe (`extract_bits`) 0.77; bfec (C form) 0.75 | 0.77 (3.731; 1,151) | PENDING (`bfe.u32` and the C form) | OWED |
|
||||
| andn | andn | 0.75 (3.676; 1,168) | PENDING | OWED |
|
||||
| byte permute, immediate selector | perm | 1.13 (5.500; 781), EMULATED (shifts and masks) | PENDING (`prmt`) | OWED (`v_perm_b32`) |
|
||||
| popcount folded by add | popc | 0.87 (4.223; 1,017) | PENDING | OWED |
|
||||
| clz folded by add | clz | 1.01 (4.914; 874) | PENDING | OWED |
|
||||
| three-register select | sel | 0.76 (3.718; 1,155) | PENDING | OWED |
|
||||
| second shuffle form, lane + delta mod 32 | shfla | 1.91 (9.287; 462) | PENDING (`shfl.sync.idx`) | OWED (`ds_bpermute_b32`) |
|
||||
| comparison: dp4a | dot4u 1.60 (emulated), dot4s 4.73 (emulated) | as the 5 October rows (1.6x, 4.7x) | PENDING (`__dp4a`, 1.17x on 5 October) | 1.06x on 5 October (`v_dot4_i32_iu8`) |
|
||||
| comparison: mm8 as a chain | mm8 | OWED (Metal 4 `matmul2d`; Swift 5.8 toolchain has no tensor API) | PENDING (`mma.sync.m8n8k16.u8`, inline PTX) | OWED (WMMA) |
|
||||
|
||||
The three Mac runs agree within 4% on every row; the 5090 column is filled from the job's closing report the moment it is read (bench-log entry). The AMD column is the owed row: PC 1 is the project lead's desk today.
|
||||
|
||||
What the Mac rows say. Six of the seven candidates cost at most the live `rotr` on Apple. The two Apple pays for: `perm` at 1.13x (no byte-permute function in MSL, the `uchar4` swizzle compiles to shifts and masks) and `shfla` at 1.91x (`simd_shuffle` by a computed lane index against `simd_shuffle_xor`, 2.2x the live shuffle). Both are inside the 8x per-op bound of 1.13.2 by a wide margin, and at `W_new` = 4 points of the 64-instruction program a 1.91x op is under 1% of the program's ALU time on a hash that spends its time on 128 dependent DRAM reads (approximate: argued from the step cost and the weight, not measured; the 5% rule is checked on the vendor's card with the family live, as 1.13.2 says).
|
||||
|
||||
## 2. Native or emulated, per vendor, cited
|
||||
|
||||
Sources read 6 October 2026: PTX ISA 9.4 (https://docs.nvidia.com/cuda/parallel-thread-execution/index.html, section numbers below); the LLVM AMDGPU backend's instruction tables (https://github.com/llvm/llvm-project/tree/main/llvm/lib/Target/AMDGPU: `VOP1Instructions.td`, `VOP2Instructions.td`, `VOP3Instructions.td`, `DSInstructions.td`, `VOP3PInstructions.td`, main, line numbers below; AMD's CDN refused the RDNA 3 and RDNA 4 ISA PDFs again today, as on 5 October, so the mnemonics are the compiler's, not quoted from the ISA guide); Metal Shading Language Specification version 4.1 (https://developer.apple.com/metal/Metal-Shading-Language-Specification.pdf, section 6.4 Integer Functions and 6.10.2 SIMD-Group Functions). Apple's GPU instruction set is not published, so "native" on Apple means "an MSL function exists"; the step cost is the only measure of what the compiler emits.
|
||||
|
||||
| Family | NVIDIA (PTX) | AMD (RDNA, by the LLVM mnemonic) | Apple (MSL) | Emulated where |
|
||||
|---|---|---|---|---|
|
||||
| variable shifts | `shl.b32`, `shr.b32` (9.7.9.8, 9.7.9.9; "supported on all target architectures") | `v_lshlrev_b32`, `v_lshrrev_b32` (VOP2Instructions.td 899, 901) | `<<`, `>>` operators (6.4) | nowhere |
|
||||
| bit-field extract | `bfe.u32` (9.7.1.20, sm_20+); whether SASS keeps it as one instruction on Blackwell is what the `bfe` against `bfec` rows measure | `v_bfe_u32` (VOP3Instructions.td 276) | `extract_bits` (6.4) | nowhere by the ISA; Apple's lowering unknown (0.77x measured, so no penalty) |
|
||||
| andn | `and.b32` with `not.b32`, folded into `lop3.b32` on sm_50+ (9.7.9.1, 9.7.9.4, 9.7.9.6) | `v_and_b32` with `v_not_b32`, or one `v_bfi_b32` (VOP2 902, VOP1 378, VOP3 278) | `&`, `~` (6.4) | nowhere |
|
||||
| byte permute | `prmt.b32` (9.7.10.7, sm_20+) | `v_perm_b32` (VOP3Instructions.td 426) | NO function; `uchar4` swizzle, 1.13x measured | Apple |
|
||||
| popcount, clz | `popc.b32`, `clz.b32` (9.7.1.15, 9.7.1.16, sm_20+) | `v_bcnt_u32_b32` (VOP2 980: popcount plus an add in one instruction, the folded form exactly), `v_ffbh_u32`, renamed `v_clz_i32_u32` on gfx11+ (VOP1 380, 1296) | `popcount`, `clz` (6.4; `clz(0)` = 32 on all three) | nowhere |
|
||||
| three-register select | `selp.b32` (9.7.7.3) | `v_cndmask_b32` (VOP2Instructions.td 878) | `select` (6.4) | nowhere |
|
||||
| second shuffle form (lane + delta) | `shfl.sync.idx.b32` (9.7.10.6, sm_30+) | `ds_bpermute_b32` (DSInstructions.td 839: a backward permute through the LDS crossbar, not a VALU op; the live xor form has the same route or DPP) | `simd_shuffle` (6.10.2) | nowhere by the ISA; Apple 1.91x measured, AMD's LDS route owed |
|
||||
| dp4a (comparison) | `dp4a` (9.7.1.24, sm_61+) | `v_dot4_i32_iu8` (VOP3PInstructions.td 859) | none; emulated 1.6x / 4.7x | Apple |
|
||||
| mm8 (comparison) | `mma.sync.m8n8k16` `.u8` (9.7.16.5.3, sm_75+) | `v_wmma_i32_16x16x16_iu8` (VOP3PInstructions.td 1725, 2307) | Metal 4 `matmul2d` (a library path, cost owed) | nowhere by the ISA |
|
||||
|
||||
## 3. What a chip pays per family
|
||||
|
||||
A 12-op chip (the ledger's M1: a sequencer over the live families and 8 registers) carries per lane: a 32-bit adder (also the subtractor), a 32x32 multiplier with the high half (`mul`, `mulhi`, `mad`), xor, or, a 32-bit rotator (`rotl` by immediate, `rotr` by a register), the xor-shuffle butterfly per warp (5 stages, masks 1, 2, 4, 8, 16) and the load path. Every structure below is named with its approximate area relative to that lane's 32-bit adder (approximate, from memory of standard-cell estimates: a ripple adder is about 32 full-adder cells; a 2:1 mux is about a third of a full adder; these are order-of-magnitude figures to rank the families, not a layout).
|
||||
|
||||
| Family | Structure the chip must add | Approximate area, in 32-bit adders per lane | Note |
|
||||
|---|---|---|---|
|
||||
| variable shifts | a 32-bit barrel shifter: 5 stages x 32 2:1 muxes, plus the zero fill | 1 to 2; but the live `rotr` already needs a 32-bit barrel rotator, and a shifter is the rotator with a fill mask: about 0.2 more | the smallest addition to a chip that already runs class v3 |
|
||||
| bit-field extract | the shifter plus a 32-bit mask generator (a decoder over the width) | 0.3 beyond the shifter | small |
|
||||
| andn | 32 inverters on one input of the AND | under 0.1 | nothing a chip lacks |
|
||||
| byte permute | a 4x4 byte crossbar: 4 output bytes x 4:1 mux x 8 bits = 128 mux bits, plus the selector decode | 1 to 2 | a structure no live family needs; a chip that rotates by multiples of 8 does half of it, the other half is new |
|
||||
| popcount, clz | a 32-bit popcount tree (31 small adders, 5 levels) and a 32-bit priority encoder (leading-zero count) | 1.5 to 3 together | the tree is a new structure; neither is in the live set |
|
||||
| three-register select | 32 2:1 muxes and a bit select | about 0.3 | nothing a chip lacks; every sequencer has operand muxes |
|
||||
| second shuffle form | a full 32-lane x 32-bit crossbar per warp (32 outputs x 32:1 mux x 32 bits = 32,768 mux bits) in place of the 5-stage butterfly (5 x 32 x 32 = 5,120 mux bits) | about 6x the butterfly per warp, so about 3 to 6 adders per lane (approximate) | the largest 32-bit structure of the seven; on a GPU it is the existing shuffle network, so free for the honest card |
|
||||
| dp4a (comparison) | four 8x8 multipliers and a 4-input adder tree | 4 to 6 | the chip's multiplier can be split into four 8x8 blocks; cheap for a chip |
|
||||
| mm8 | a 8x8x16 u8 MAC tile per warp (1,024 MACs; 32 per lane) | about 100 per lane (approximate) | licensable IP at every node (rank 6); the structure the history says to order last |
|
||||
|
||||
## 4. The bounds of 1.13.2 per family
|
||||
|
||||
| Family | 8x per-op emulation bound | 5% hash-rate bound at `W_new` = 4 | Reading |
|
||||
|---|---|---|---|
|
||||
| variable shifts | holds everywhere: no emulation (Apple 0.85 / 0.86x, native PTX and RDNA) | holds, argued: a 0.86x op at 4 points of 64 moves the ALU time by under 1%, and the hash is read-bound (latency-bound share 0.95 to 1.06 on the three cards, bench-log "Counter ASIC 2.0, the numbers") | in |
|
||||
| bit-field extract | holds: Apple 0.77x; native on PTX and RDNA | holds, argued | in |
|
||||
| andn | holds: 0.75x; native everywhere | holds, argued | in |
|
||||
| byte permute | holds: Apple emulated at 1.13x (inside 8x by 7x); native PTX `prmt`, RDNA `v_perm_b32` | holds, argued: a 1.13x op at 4 points is under 0.5% of the ALU time | in; Apple is the vendor that can only emulate, inside both bounds |
|
||||
| popcount, clz | holds: 0.87x / 1.01x; native everywhere | holds, argued | in |
|
||||
| three-register select | holds: 0.76x; native everywhere | holds, argued | in |
|
||||
| second shuffle form | holds: Apple 1.91x (inside 8x by 4x); native PTX; AMD's `ds_bpermute_b32` is an LDS op whose step cost is owed | holds, argued for Apple (under 1% at 4 points); AMD argued from the live `shfl` sharing the route, measured when the PC 1 row lands | in, pending the AMD row |
|
||||
| mm8 | Metal 4 `matmul2d` is a path, cost owed; the per-lane dot4 is 1.6x emulated | holds, argued (`int8-matrix-family.md` section 4) | in, R8 |
|
||||
|
||||
Every bound above is argued from the step cost and the weight, not measured on a live program: no reserve family is in a live program today. The measurement is the family-live run of 1.13.2's own text, owed per family at its unlock rehearsal.
|
||||
|
||||
## 5. The proposed order
|
||||
|
||||
Ranked by the structure a chip must add beyond the class v3 datapath (section 3), then by what the honest vendors pay (section 1): the largest new 32-bit structure first, the trivial ones last, `mm8` last of all.
|
||||
|
||||
| Reserve slot | Family | Why here | `W_new` | Unlock by the rule of 1.13.2 (family n at era n; era = 180 days) |
|
||||
|---|---|---|---|---|
|
||||
| R1 | byte permute (`perm`) | a byte crossbar per lane that no live family needs; native on NVIDIA and AMD; Apple pays 1.13x, inside both bounds | 4 | era 1 (day 180) |
|
||||
| R2 | popcount and clz folded by add (`popc`, `clz`) | a popcount tree and a priority encoder, both new; native on all three; costs no vendor anything | 4 | era 2 |
|
||||
| R3 | second shuffle form (`shfla`) | the 32-lane crossbar is the largest 32-bit structure of the seven; native on all three; Apple pays 1.91x; the AMD LDS cost is the owed number that could move this row to R1 (cheap on AMD) or to R5 (dear) | 4 | era 3 |
|
||||
| R4 | bit-field extract (`bfe`) | a mask generator on the shifter; native everywhere; 0.77x on Apple | 4 | era 4 |
|
||||
| R5 | variable shifts (`shl`, `shr` by `src AND 31`, the direction an immediate bit) | the smallest addition to a chip that already rotates; native everywhere | 4 | era 5 |
|
||||
| R6 | three-register select (`sel`) | operand muxes a sequencer has anyway; its value is the data-dependent path, not the silicon | 4 | era 6 |
|
||||
| R7 | andn | an inverter; last of the datapath families | 4 | era 7 |
|
||||
| R8 | `mm8` (integer matrix) | licensable IP at every node; the vendor-emulation rule written for it | 4 (as decided 5 October) | era 8 (day 1,440), or earlier by the 90% signal |
|
||||
|
||||
The decision this moves for the project lead: `mm8` was reserved on 5 October as R1 with an unlock at era 4. Under the rule "family n at era n" an eighth slot unlocks at era 8 (four years). Either the rule stays and `mm8` waits, or the entry keeps its era-4 unlock as a named exception (the text below keeps the era-4 date as an exception and says so).
|
||||
|
||||
## 6. PROPOSED spec text for 1.13.2 (replaces the paragraph from "The order and `W_new` are Open" to the end of the `mm8` entry)
|
||||
|
||||
> The reserve, in order. Reserve family `n` becomes live at the start of era `n` (1.13.1), or earlier by the 90% signalling path of section 5.7; never by a release. Every entry takes `W_new` = 4 points proportionally from the live non-load families (the load weight and count are untouched). Every entry's edge vectors are hand-built units run on every vendor of 1.15 before genesis. The ordering rule: the family whose silicon a class v3 chip lacks most comes first, the integer matrix family last (`docs/plans/counter-asic-3-reserve.md`).
|
||||
>
|
||||
> R1, `perm` (byte permute). Semantics: `dst = permute(dst, sel4)`, result byte `i` = byte `(sel4 >> 2i) AND 3` of `dst`, `sel4` an 8-bit immediate drawn per instruction. Native: PTX `prmt.b32`, RDNA `v_perm_b32`; Apple by `uchar4` swizzle (emulation, 1.13x per op measured 6 October 2026). Edge vectors: `dst` = 0x00FF00FF with every `sel4` of the form (3, 2, 1, 0), (0, 1, 2, 3), (0, 0, 0, 0), (3, 3, 3, 3); `dst` = 0x80808080 (the sign bits travel as bytes, no sign extension); `dst` = 0x01020304 with `sel4` = (1, 3, 0, 2) (the probe's selector: result 0x03010402); alternating 0x00 and 0xFF by lane.
|
||||
>
|
||||
> R2, `popc` (population count and count-leading-zeros). Semantics: `dst = dst + popcount(src)` when the immediate bit is 0, `dst = dst + clz(src)` when it is 1, `clz(0)` = 32, add modulo 2^32. Native: PTX `popc.b32` and `clz.b32`, RDNA `v_bcnt_u32_b32` and `v_clz_i32_u32`, MSL `popcount` and `clz`. Edge vectors: `src` = 0 (popcount 0, clz 32); `src` = 0xFFFFFFFF (32, 0); `src` = 1 (1, 31); `src` = 0x80000000 (1, 0); `dst` = 0xFFFFFFFF with `src` = 0xFFFFFFFF (the wrap to 31); both bits on the same operands.
|
||||
>
|
||||
> R3, `shfla` (shuffle by lane plus delta). Semantics: `dst = dst XOR src_of_lane((lane + delta) mod 32)`, `delta` in 1..31 drawn per instruction, within the lane's own 32-lane warp. Native: PTX `shfl.sync.idx.b32`, RDNA `ds_bpermute_b32`, MSL `simd_shuffle`. Edge vectors: `delta` = 1 and 31 on a warp whose `src` is the lane index (the wrap at lane 31 and lane 0); `delta` = 16 against the live `shfl` mask 16 (equal results); `src` all equal (dst unchanged except by the xor with itself: 0 on every lane); one lane's `src` = 0xFFFFFFFF, the rest 0, at `delta` = 7.
|
||||
>
|
||||
> R4, `bfe` (bit-field extract). Semantics: `dst = (src >> off) AND (2^w - 1)`, `off` in 0..31 and `w` in 1..32 drawn per instruction, `off + w` clamped to 32 (bits past 31 read as 0). Native: PTX `bfe.u32`, RDNA `v_bfe_u32`, MSL `extract_bits`. Edge vectors: `off` = 0, `w` = 32 (identity); `off` = 31, `w` = 1; `off` = 7, `w` = 13 (the probe); `off` = 24, `w` = 16 (the clamp: 8 bits come back); `src` = 0xFFFFFFFF at every pair above; `src` = 0.
|
||||
>
|
||||
> R5, `shl` (variable shifts). Semantics: `dst = dst << (src AND 31)` when the immediate bit is 0, `dst = dst >> (src AND 31)` (logical) when it is 1. Native everywhere (PTX `shl.b32`, `shr.b32`; RDNA `v_lshlrev_b32`, `v_lshrrev_b32`; MSL operators). Edge vectors: `src AND 31` = 0 (identity) and 31; `dst` = 0xFFFFFFFF at both; `dst` = 0x80000000 right by 31 (1) and left by 1 (0); `src` = 0xFFFFFFE0 (the mask reads 0, not 32); both bits on the same operands.
|
||||
>
|
||||
> R6, `sel` (three-register select). Semantics: `dst = (bit `bit` of src2) ? src : dst`, `bit` in 0..31 drawn per instruction. Native everywhere (PTX `selp.b32`, RDNA `v_cndmask_b32`, MSL `select`). Edge vectors: `src2` = 0 and 0xFFFFFFFF at `bit` = 0 and 31; `src` = `dst`; `src2` = `dst` (the selected bit reads the register being written: the value before the instruction); alternating by lane.
|
||||
>
|
||||
> R7, `andn`. Semantics: `dst = dst AND NOT src`. Native everywhere. Edge vectors: `src` = 0 (identity), 0xFFFFFFFF (0), `src` = `dst` (0); `dst` = 0xAAAAAAAA with `src` = 0x55555555 (unchanged).
|
||||
>
|
||||
> R8, `mm8` (integer matrix): the entry decided 5 October 2026, text unchanged (semantics, `W_new` = 4, the six edge vectors, native paths, the vendor-that-can-only-emulate rule), with one change: it is the eighth reserve family. Its unlock stays the start of era 4 (DAA 62,208,000) as a named exception to "family n at era n", or moves to era 8 if the project lead keeps the rule; one of the two is decided before the public testnet genesis.
|
||||
|
||||
The `andn` and `sel` entries carry a note for the generator: both are injecting in the sense of 1.4.1 only when the acceptance rule counts them so (`sel` writes `src` into `dst` on half the lanes; `andn` is not bijective in `dst`), so neither satisfies rule (b) of 1.4.6 alone; the weak-program census is re-run with each family at its unlock rehearsal.
|
||||
|
||||
## 7. Consequences per tier
|
||||
|
||||
| Tier | Today (the reserve is not live) | At each unlock, from the rows above |
|
||||
|---|---|---|
|
||||
| Apple user (M-series; the app's Metal worker) | nothing changes | R1 `perm` costs about 1.13x on 4 points of 64 (under 0.5% of ALU time, argued); R3 `shfla` 1.91x on 4 points (under 1%); every other family costs less than the live `rotr`. The measurement with the family live is owed at each unlock rehearsal; if a live run shows over 5% on Apple the family does not unlock (1.13.2) |
|
||||
| NVIDIA user (8 to 32 GB) | nothing changes | every family native; the 5090 rows (pending the PC 2 job) give the per-op cost; `bfe` is the one to watch (SASS has no single bfe on recent architectures by the measured `bfe` against `bfec` rows) |
|
||||
| AMD user (RX 9070 XT) | nothing changes | every family native by the LLVM tables; the `shfla` cost through `ds_bpermute_b32` (an LDS op) is the owed number, and the one that could move R3 |
|
||||
| Intel iGPU user | nothing changes | not measured anywhere; OpenCL C exposes every function (`popcount`, `clz`, `sub_group_shuffle`); owed with the first Intel measurement |
|
||||
| A rig, a pool user | nothing changes | the same per-card figures; a pool verifies with the CPU reference, which gains one op per family and stays under the 10 ms gate (the verifier runs 2.08 ms per warp at x8 against a 10 ms gate) |
|
||||
| A chip built against class v3 | nothing changes | R1 to R3 each add a structure it lacks (byte crossbar, popcount tree, 32-lane crossbar) within 18 months of genesis; R4 to R7 add little silicon but take weight from the families it was built for; `mm8` at R8 adds the licensable block last |
|
||||
|
||||
## 8. What is owed
|
||||
|
||||
| Item | State |
|
||||
|---|---|
|
||||
| RTX 5090 step costs (every row) | PENDING: the PC 2 job, after `/tmp/igneum-devnet/pc2-ca3.clear` and under the mkdir lock; one job carries every family three times |
|
||||
| RX 9070 XT step costs (every row) | OWED: PC 1 is the project lead's desk and not used today; the OpenCL twin of the probe is the next job on that card |
|
||||
| `mm8` as a chain on Apple (Metal 4 `matmul2d`) | OWED (toolchain) |
|
||||
| AMD `ds_bpermute_b32` cost for `shfla` | OWED (the PC 1 row); decides whether R3 moves |
|
||||
| The 5% hash-rate rule per family with the family live | argued, not measured, for every family (no reserve family is in a live program); measured at each unlock rehearsal |
|
||||
| RDNA 3 and RDNA 4 ISA guides (the instruction text) | AMD's CDN refused the PDFs again; the mnemonics are from the LLVM tables |
|
||||
| the project lead's decisions | the order R1 to R8; `W_new` = 4 per family; `mm8` at era 4 by exception or at era 8 by the rule |
|
||||
179
proto-cuda/family-probe.cu
Normal file
179
proto-cuda/family-probe.cu
Normal file
|
|
@ -0,0 +1,179 @@
|
|||
// family-probe (CUDA): the step cost of every reserve candidate family of spec 1.13.2 on NVIDIA, standalone (no pack).
|
||||
// Counter ASIC 3.0 item 6 (docs/plans/counter-asic-3-reserve.md), 6 October 2026. PC job; not run on the Mac. Same
|
||||
// method as dot4-probe.cu (docs/analysis/int8-matrix-family.md section 4): a dependent chain of one op per step per
|
||||
// lane, 1,048,576 lanes x 4,096 steps, best of N, event time, bit-exact against a CPU reference on two whole warps.
|
||||
//
|
||||
// Every chain has the dot4 probe's glue: acc = OP(acc, x, y); x = x * K + acc; y = rotl(y, 7) ^ (acc + s). Reference
|
||||
// rows (live families): alu (add-xor-rotate, 5 ops per step counted, no acc), rotr (the live rotr_var text), shflx
|
||||
// (the live shfl: dst ^= src of lane (lane XOR 8)). Candidate rows, the seven families of 1.13.2:
|
||||
// shl acc = y << (x & 31) shr acc = y >> (x & 31)
|
||||
// bfe acc = bfe.u32(y, 7, 13) (inline PTX) bfec acc = (y >> 7) & 0x1fff (the C form, what nvcc emits)
|
||||
// andn acc = y & ~x perm acc = __byte_perm(y, y, 0x2031): bytes (y.b1, y.b3, y.b0, y.b2)
|
||||
// popc acc = acc + __popc(x) clz acc = acc + __clz(x) (clz(0) = 32)
|
||||
// sel acc = bit 5 of y ? x : acc shfla acc ^= __shfl_sync(x, (lane + 3) & 31)
|
||||
// Comparison rows: dot4i (__dp4a, PTX dp4a.s32.s32, the dot4 probe's intrinsic row) and mm8 (one mma.sync
|
||||
// m8n8k16 u8 x u8 -> s32 per step per warp, inline PTX, sm_75+; PTX ISA 9.4 section 9.7.16.5.3 fragment layout:
|
||||
// lane i holds A[i/4][4(i%4)..+3] as the bytes of x, B[4(i%4)..+3][i/4] as the bytes of y, C and D [i/4][2(i%4)+j]
|
||||
// as acc (j = 0) and acc2 (j = 1); the chain carries d0 forward as acc, d1 as acc2). The shifted and extracted forms
|
||||
// take y (fresh every step) as the shifted register so the chain never drains to 0.
|
||||
//
|
||||
// Build (Linux or WSL2, CUDA Toolkit): nvcc -O2 -arch=sm_120 -o family-probe family-probe.cu
|
||||
// (older toolkits: nvcc -O2 -gencode arch=compute_89,code=compute_89 ...)
|
||||
// Run: family-probe [--lanes N] [--steps N] [--reps N] [--device N]
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <cstring>
|
||||
#include <cstdint>
|
||||
|
||||
__host__ __device__ inline uint32_t pm_mix(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }
|
||||
__host__ __device__ inline uint32_t rotl32(uint32_t v, uint32_t n) { return (v << n) | (v >> (32u - n)); }
|
||||
__host__ __device__ inline uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
#define K32 0x9E3779B1u
|
||||
#define CHAIN(NAME, OP) \
|
||||
__global__ void NAME(uint32_t steps, uint32_t seed, uint32_t* out) { \
|
||||
uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t lane = threadIdx.x & 31u; (void)lane; \
|
||||
uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; uint32_t acc = pm_mix(x); \
|
||||
for (uint32_t s = 0; s < steps; ++s) { OP; x = x * K32 + acc; y = rotl32(y, 7u) ^ (acc + s); } \
|
||||
out[g] = acc ^ x ^ y; \
|
||||
}
|
||||
|
||||
__global__ void probe_alu(uint32_t steps, uint32_t seed, uint32_t* out) {
|
||||
uint32_t g = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u;
|
||||
for (uint32_t s = 0; s < steps; ++s) { x = x * K32 + rotl32(y, 7u); y = (y ^ x) + s; }
|
||||
out[g] = x ^ y;
|
||||
}
|
||||
__device__ inline uint32_t bfe_ptx(uint32_t v) { uint32_t r; asm("bfe.u32 %0, %1, 7, 13;" : "=r"(r) : "r"(v)); return r; }
|
||||
CHAIN(probe_rotr, acc = rotr_var(y, x))
|
||||
CHAIN(probe_shflx, acc = acc ^ __shfl_xor_sync(0xffffffffu, x, 8))
|
||||
CHAIN(probe_shl, acc = y << (x & 31u))
|
||||
CHAIN(probe_shr, acc = y >> (x & 31u))
|
||||
CHAIN(probe_bfe, acc = bfe_ptx(y))
|
||||
CHAIN(probe_bfec, acc = (y >> 7u) & 0x1fffu)
|
||||
CHAIN(probe_andn, acc = y & ~x)
|
||||
CHAIN(probe_perm, acc = __byte_perm(y, y, 0x2031))
|
||||
CHAIN(probe_popc, acc = acc + (uint32_t)__popc((int)x))
|
||||
CHAIN(probe_clz, acc = acc + (uint32_t)__clz((int)x))
|
||||
CHAIN(probe_sel, acc = ((y >> 5u) & 1u) ? x : acc)
|
||||
CHAIN(probe_shfla, acc = acc ^ __shfl_sync(0xffffffffu, x, (int)((lane + 3u) & 31u)))
|
||||
CHAIN(probe_dot4i, acc = (uint32_t)__dp4a((int)x, (int)y, (int)acc))
|
||||
|
||||
// one mma.sync m8n8k16 per step per warp; the chain carries d0 as acc, d1 as acc2
|
||||
__global__ void probe_mm8(uint32_t steps, uint32_t seed, uint32_t* out) {
|
||||
uint32_t g = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; uint32_t acc = pm_mix(x), acc2 = pm_mix(acc);
|
||||
for (uint32_t s = 0; s < steps; ++s) {
|
||||
uint32_t d0, d1;
|
||||
asm volatile("mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32 {%0,%1}, {%2}, {%3}, {%4,%5};"
|
||||
: "=r"(d0), "=r"(d1) : "r"(x), "r"(y), "r"(acc), "r"(acc2));
|
||||
acc = d0; acc2 = d1;
|
||||
x = x * K32 + acc; y = rotl32(y, 7u) ^ (acc + s);
|
||||
}
|
||||
out[g] = acc ^ acc2 ^ x ^ y;
|
||||
}
|
||||
|
||||
// ---- CPU reference: one whole warp (lanes g0 .. g0+31, lane = g AND 31) ----
|
||||
static uint32_t clz32(uint32_t x) { if (!x) return 32; uint32_t n = 0; while (!(x & 0x80000000u)) { x <<= 1; ++n; } return n; }
|
||||
static uint32_t popc32(uint32_t x) { uint32_t n = 0; while (x) { n += x & 1u; x >>= 1; } return n; }
|
||||
static uint32_t perm_ref(uint32_t y) { return ((y >> 8) & 0xffu) | (((y >> 24) & 0xffu) << 8) | ((y & 0xffu) << 16) | (((y >> 16) & 0xffu) << 24); }
|
||||
static uint32_t dot4s_ref(uint32_t a, uint32_t b, uint32_t acc) {
|
||||
int32_t r = (int32_t)acc;
|
||||
for (int i = 0; i < 4; ++i) { int32_t ba = (int32_t)(int8_t)((a >> (8 * i)) & 0xffu); int32_t bb = (int32_t)(int8_t)((b >> (8 * i)) & 0xffu); r = (int32_t)((uint32_t)r + (uint32_t)(ba * bb)); }
|
||||
return (uint32_t)r;
|
||||
}
|
||||
static void warp_ref(const char* name, uint32_t g0, uint32_t seed, uint32_t steps, uint32_t* res) {
|
||||
uint32_t x[32], y[32], acc[32], acc2[32], xs[32];
|
||||
for (int l = 0; l < 32; ++l) { x[l] = pm_mix((g0 + l) ^ seed); y[l] = x[l] ^ 0x5bd1e995u; acc[l] = pm_mix(x[l]); acc2[l] = pm_mix(acc[l]); }
|
||||
if (!strcmp(name, "alu")) {
|
||||
for (uint32_t s = 0; s < steps; ++s) for (int l = 0; l < 32; ++l) { x[l] = x[l] * K32 + rotl32(y[l], 7u); y[l] = (y[l] ^ x[l]) + s; }
|
||||
for (int l = 0; l < 32; ++l) res[l] = x[l] ^ y[l];
|
||||
return;
|
||||
}
|
||||
for (uint32_t s = 0; s < steps; ++s) {
|
||||
memcpy(xs, x, sizeof xs);
|
||||
if (!strcmp(name, "mm8")) {
|
||||
uint32_t D[8][8];
|
||||
for (int r = 0; r < 8; ++r) for (int n = 0; n < 8; ++n) {
|
||||
uint32_t c = (n & 1) ? acc2[r * 4 + n / 2] : acc[r * 4 + n / 2];
|
||||
for (int k = 0; k < 16; ++k) {
|
||||
uint32_t a = (xs[r * 4 + k / 4] >> (8 * (k % 4))) & 0xffu;
|
||||
uint32_t b = (y[n * 4 + k / 4] >> (8 * (k % 4))) & 0xffu;
|
||||
c += a * b;
|
||||
}
|
||||
D[r][n] = c;
|
||||
}
|
||||
for (int l = 0; l < 32; ++l) { acc[l] = D[l / 4][2 * (l % 4)]; acc2[l] = D[l / 4][2 * (l % 4) + 1]; }
|
||||
for (int l = 0; l < 32; ++l) { x[l] = xs[l] * K32 + acc[l]; y[l] = rotl32(y[l], 7u) ^ (acc[l] + s); }
|
||||
continue;
|
||||
}
|
||||
for (int l = 0; l < 32; ++l) {
|
||||
uint32_t xv = xs[l], yv = y[l];
|
||||
if (!strcmp(name, "rotr")) acc[l] = rotr_var(yv, xv);
|
||||
else if (!strcmp(name, "shflx")) acc[l] = acc[l] ^ xs[l ^ 8];
|
||||
else if (!strcmp(name, "shl")) acc[l] = yv << (xv & 31u);
|
||||
else if (!strcmp(name, "shr")) acc[l] = yv >> (xv & 31u);
|
||||
else if (!strcmp(name, "bfe") || !strcmp(name, "bfec")) acc[l] = (yv >> 7u) & 0x1fffu;
|
||||
else if (!strcmp(name, "andn")) acc[l] = yv & ~xv;
|
||||
else if (!strcmp(name, "perm")) acc[l] = perm_ref(yv);
|
||||
else if (!strcmp(name, "popc")) acc[l] = acc[l] + popc32(xv);
|
||||
else if (!strcmp(name, "clz")) acc[l] = acc[l] + clz32(xv);
|
||||
else if (!strcmp(name, "sel")) acc[l] = ((yv >> 5u) & 1u) ? xv : acc[l];
|
||||
else if (!strcmp(name, "shfla")) acc[l] = acc[l] ^ xs[(l + 3) & 31];
|
||||
else if (!strcmp(name, "dot4i")) acc[l] = dot4s_ref(xv, yv, acc[l]);
|
||||
else { printf("no reference for %s\n", name); exit(3); }
|
||||
x[l] = xv * K32 + acc[l];
|
||||
y[l] = rotl32(yv, 7u) ^ (acc[l] + s);
|
||||
}
|
||||
}
|
||||
for (int l = 0; l < 32; ++l) res[l] = acc[l] ^ x[l] ^ y[l] ^ (!strcmp(name, "mm8") ? acc2[l] : 0u);
|
||||
}
|
||||
|
||||
#define CK(x) do { cudaError_t e = (x); if (e != cudaSuccess) { printf("CUDA error %s at %s:%d\n", cudaGetErrorString(e), __FILE__, __LINE__); return 1; } } while (0)
|
||||
typedef void (*kernel_t)(uint32_t, uint32_t, uint32_t*);
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
uint32_t lanes = 1u << 20, steps = 4096u; int reps = 3, device = 0;
|
||||
for (int i = 1; i < argc; ++i) {
|
||||
if (!strcmp(argv[i], "--lanes") && i + 1 < argc) lanes = (uint32_t)strtoul(argv[++i], 0, 10);
|
||||
else if (!strcmp(argv[i], "--steps") && i + 1 < argc) steps = (uint32_t)strtoul(argv[++i], 0, 10);
|
||||
else if (!strcmp(argv[i], "--reps") && i + 1 < argc) reps = atoi(argv[++i]);
|
||||
else if (!strcmp(argv[i], "--device") && i + 1 < argc) device = atoi(argv[++i]);
|
||||
else { printf("unknown argument %s\n", argv[i]); return 2; }
|
||||
}
|
||||
CK(cudaSetDevice(device));
|
||||
cudaDeviceProp p; CK(cudaGetDeviceProperties(&p, device));
|
||||
printf("family-probe (CUDA) on %s, sm_%d%d, %d SMs, %d MHz, lanes %u, steps %u, best of %d, event time\n", p.name, p.major, p.minor, p.multiProcessorCount, p.clockRate / 1000, lanes, steps, reps);
|
||||
uint32_t* d_out; CK(cudaMalloc(&d_out, (size_t)lanes * 4));
|
||||
uint32_t* h_out = (uint32_t*)malloc((size_t)lanes * 4);
|
||||
cudaEvent_t e0, e1; CK(cudaEventCreate(&e0)); CK(cudaEventCreate(&e1));
|
||||
const char* names[] = { "alu", "rotr", "shflx", "shl", "shr", "bfe", "bfec", "andn", "perm", "popc", "clz", "sel", "shfla", "dot4i", "mm8" };
|
||||
kernel_t kernels[] = { probe_alu, probe_rotr, probe_shflx, probe_shl, probe_shr, probe_bfe, probe_bfec, probe_andn, probe_perm, probe_popc, probe_clz, probe_sel, probe_shfla, probe_dot4i, probe_mm8 };
|
||||
const int nk = (int)(sizeof(names) / sizeof(names[0]));
|
||||
printf("| kernel | lanes | steps | best ms | G steps/s | ns per step | ratio to alu | warps 0 and last ok |\n|---|---|---|---|---|---|---|---|\n");
|
||||
float alu_best = 0;
|
||||
for (int k = 0; k < nk; ++k) {
|
||||
float best = 1e30f; int ok = 1;
|
||||
for (int r = 0; r < reps; ++r) {
|
||||
uint32_t seed = 0x2468aceu + (uint32_t)r * 0x9E3779B9u;
|
||||
CK(cudaEventRecord(e0));
|
||||
kernels[k]<<<lanes / 256, 256>>>(steps, seed, d_out);
|
||||
CK(cudaEventRecord(e1)); CK(cudaEventSynchronize(e1)); CK(cudaGetLastError());
|
||||
float ms = 0; CK(cudaEventElapsedTime(&ms, e0, e1)); if (ms < best) best = ms;
|
||||
CK(cudaMemcpy(h_out, d_out, (size_t)lanes * 4, cudaMemcpyDeviceToHost));
|
||||
uint32_t g0s[2] = { 0u, lanes - 32u };
|
||||
for (int j = 0; j < 2; ++j) {
|
||||
uint32_t want[32]; warp_ref(names[k], g0s[j], seed, steps, want);
|
||||
for (int l = 0; l < 32; ++l) if (h_out[g0s[j] + l] != want[l]) { ok = 0; printf("MISMATCH %s lane %u: gpu %08x cpu %08x\n", names[k], g0s[j] + l, h_out[g0s[j] + l], want[l]); break; }
|
||||
}
|
||||
}
|
||||
if (k == 0) alu_best = best;
|
||||
double sps = (double)lanes * (double)steps / (best / 1000.0);
|
||||
double ratio = alu_best > 0 ? best / alu_best : 0;
|
||||
printf("| %s | %u | %u | %.3f | %.2f | %.3f | %.2f | %s |\n", names[k], lanes, steps, best, sps / 1e9, best * 1e6 / (double)steps, ratio, ok ? "yes" : "NO");
|
||||
printf("RESULT FAMILY vendor=nvidia device=\"%s\" kernel=%s lanes=%u steps=%u best_ms=%.3f gsteps_per_s=%.2f ns_per_step=%.3f ratio_alu=%.3f ok=%d\n", p.name, names[k], lanes, steps, best, sps / 1e9, best * 1e6 / (double)steps, ratio, ok);
|
||||
}
|
||||
printf("family-probe: done\n");
|
||||
return 0;
|
||||
}
|
||||
199
proto-metal/family-probe.swift
Normal file
199
proto-metal/family-probe.swift
Normal file
|
|
@ -0,0 +1,199 @@
|
|||
// family-probe: the step cost of every reserve candidate family of spec 1.13.2 on Apple silicon, standalone (no pack).
|
||||
// Counter ASIC 3.0 item 6 (docs/plans/counter-asic-3-reserve.md), 6 October 2026. Same method as dot4-probe.swift
|
||||
// (docs/analysis/int8-matrix-family.md section 4): a dependent chain of one op per step per lane, 1,048,576 lanes x
|
||||
// 4,096 steps, best of N, GPU start-to-end time, bit-exact against a CPU reference on two whole 32-lane SIMD groups.
|
||||
//
|
||||
// Every chain has the dot4 probe's glue: acc = OP(acc, x, y); x = x * K + acc; y = rotate(y, 7) ^ (acc + s). The
|
||||
// reference rows are the live families: alu (the add-xor-rotate chain of the 9070 XT bench-log entry, 5 ops per step
|
||||
// counted, no acc), rotr (dst = rotr(src, src2 AND 31), the live rotr_var text), shflx (dst = dst XOR src of lane
|
||||
// (lane XOR 8), the live shfl). The candidate rows are the seven families of 1.13.2:
|
||||
// shl acc = y << (x AND 31) variable left shift
|
||||
// shr acc = y >> (x AND 31) variable logical right shift
|
||||
// bfe acc = extract_bits(y, 7, 13) bit-field extract, immediate offset and width (bfec = the C form)
|
||||
// andn acc = y AND NOT x
|
||||
// perm acc = bytes (y.b1, y.b3, y.b0, y.b2) byte permute by an immediate selector
|
||||
// popc acc = acc + popcount(x) popcount folded by add
|
||||
// clz acc = acc + clz(x) count-leading-zeros folded by add (clz(0) = 32)
|
||||
// sel acc = bit 5 of y ? x : acc three-register select
|
||||
// shfla acc = acc XOR x of lane ((lane + 3) mod 32) the second shuffle form
|
||||
// and for comparison the dp4a-class rows of the dot4 probe (dot4u unsigned, dot4s signed, both emulated on Apple).
|
||||
// The shifted and extracted forms take y (fresh every step) as the shifted register so the chain never drains to 0;
|
||||
// the op count per step is the family's one op plus the same glue in every row. mm8 as a chain (Metal 4 matmul2d) is
|
||||
// owed: this toolchain (Swift 5.8) has no Metal 4 tensor API.
|
||||
//
|
||||
// Build: swiftc -O -o family-probe family-probe.swift -framework Metal
|
||||
// Run: ./family-probe [--lanes N] [--steps N] [--reps N]
|
||||
import Foundation
|
||||
import Metal
|
||||
|
||||
let source = """
|
||||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
inline uint pm_mix(uint x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline int dot4_s(uint a, uint b, int acc) {
|
||||
int4 va = int4(as_type<char4>(a)); int4 vb = int4(as_type<char4>(b));
|
||||
return acc + va.x * vb.x + va.y * vb.y + va.z * vb.z + va.w * vb.w;
|
||||
}
|
||||
inline uint dot4_u(uint a, uint b, uint acc) {
|
||||
uint4 va = uint4(as_type<uchar4>(a)); uint4 vb = uint4(as_type<uchar4>(b));
|
||||
return acc + va.x * vb.x + va.y * vb.y + va.z * vb.z + va.w * vb.w;
|
||||
}
|
||||
|
||||
#define CHAIN(NAME, OP) \\
|
||||
kernel void NAME(constant uint& steps [[buffer(0)]], constant uint& seed [[buffer(1)]], device uint* out [[buffer(2)]], \\
|
||||
uint g [[thread_position_in_grid]], ushort lane [[thread_index_in_simdgroup]]) { \\
|
||||
uint x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; uint acc = pm_mix(x); \\
|
||||
for (uint s = 0u; s < steps; ++s) { OP; x = x * 0x9E3779B1u + acc; y = rotate(y, 7u) ^ (acc + s); } \\
|
||||
out[g] = acc ^ x ^ y; \\
|
||||
}
|
||||
|
||||
kernel void probe_alu(constant uint& steps [[buffer(0)]], constant uint& seed [[buffer(1)]], device uint* out [[buffer(2)]],
|
||||
uint g [[thread_position_in_grid]]) {
|
||||
uint x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u;
|
||||
for (uint s = 0u; s < steps; ++s) { x = x * 0x9E3779B1u + rotate(y, 7u); y = (y ^ x) + s; }
|
||||
out[g] = x ^ y;
|
||||
}
|
||||
CHAIN(probe_rotr, acc = rotr_var(y, x))
|
||||
CHAIN(probe_shflx, acc = acc ^ simd_shuffle_xor(x, (ushort)8))
|
||||
CHAIN(probe_shl, acc = y << (x & 31u))
|
||||
CHAIN(probe_shr, acc = y >> (x & 31u))
|
||||
CHAIN(probe_bfe, acc = extract_bits(y, 7u, 13u))
|
||||
CHAIN(probe_bfec, acc = (y >> 7u) & 0x1fffu)
|
||||
CHAIN(probe_andn, acc = y & ~x)
|
||||
CHAIN(probe_perm, uchar4 b_ = as_type<uchar4>(y); acc = as_type<uint>(uchar4(b_.y, b_.w, b_.x, b_.z)))
|
||||
CHAIN(probe_popc, acc = acc + popcount(x))
|
||||
CHAIN(probe_clz, acc = acc + clz(x))
|
||||
CHAIN(probe_sel, acc = select(acc, x, ((y >> 5u) & 1u) != 0u))
|
||||
CHAIN(probe_shfla, acc = acc ^ simd_shuffle(x, (ushort)((lane + 3u) & 31u)))
|
||||
CHAIN(probe_dot4u, acc = dot4_u(x, y, acc))
|
||||
CHAIN(probe_dot4s, acc = uint(dot4_s(x, y, int(acc))))
|
||||
"""
|
||||
|
||||
// ---- CPU reference: one whole 32-lane SIMD group (lanes g0 .. g0+31, lane = g AND 31), bit-exact ----
|
||||
func pmMix(_ v: UInt32) -> UInt32 {
|
||||
var x = v
|
||||
x ^= x >> 16; x = x &* 0x7feb352d; x ^= x >> 15; x = x &* 0x846ca68b; x ^= x >> 16
|
||||
return x
|
||||
}
|
||||
func rotl(_ v: UInt32, _ n: UInt32) -> UInt32 { (v << n) | (v >> (32 - n)) }
|
||||
func rotrVar(_ x: UInt32, _ nIn: UInt32) -> UInt32 { let n = nIn & 31; return (x >> n) | (x << ((32 - n) & 31)) }
|
||||
func clz32(_ x: UInt32) -> UInt32 { x == 0 ? 32 : UInt32(x.leadingZeroBitCount) }
|
||||
func perm(_ y: UInt32) -> UInt32 {
|
||||
((y >> 8) & 0xff) | (((y >> 24) & 0xff) << 8) | ((y & 0xff) << 16) | (((y >> 16) & 0xff) << 24)
|
||||
}
|
||||
func dot4sRef(_ a: UInt32, _ b: UInt32, _ acc: UInt32) -> UInt32 {
|
||||
var r = Int32(bitPattern: acc)
|
||||
for i in 0..<4 {
|
||||
let ba = Int32(Int8(truncatingIfNeeded: a >> (8 * UInt32(i))))
|
||||
let bb = Int32(Int8(truncatingIfNeeded: b >> (8 * UInt32(i))))
|
||||
r = r &+ ba &* bb
|
||||
}
|
||||
return UInt32(bitPattern: r)
|
||||
}
|
||||
func dot4uRef(_ a: UInt32, _ b: UInt32, _ acc: UInt32) -> UInt32 {
|
||||
var r = acc
|
||||
for i in 0..<4 {
|
||||
let ba = UInt32(UInt8(truncatingIfNeeded: a >> (8 * UInt32(i))))
|
||||
let bb = UInt32(UInt8(truncatingIfNeeded: b >> (8 * UInt32(i))))
|
||||
r = r &+ ba &* bb
|
||||
}
|
||||
return r
|
||||
}
|
||||
// the 32 outputs of the SIMD group whose first lane is g0
|
||||
func warpRef(kernel: String, g0: UInt32, seed: UInt32, steps: UInt32) -> [UInt32] {
|
||||
var x = [UInt32](repeating: 0, count: 32), y = [UInt32](repeating: 0, count: 32), acc = [UInt32](repeating: 0, count: 32)
|
||||
for l in 0..<32 { x[l] = pmMix((g0 + UInt32(l)) ^ seed); y[l] = x[l] ^ 0x5bd1e995; acc[l] = pmMix(x[l]) }
|
||||
if kernel == "probe_alu" {
|
||||
for s in 0..<steps { for l in 0..<32 { x[l] = x[l] &* 0x9E3779B1 &+ rotl(y[l], 7); y[l] = (y[l] ^ x[l]) &+ s } }
|
||||
return (0..<32).map { x[$0] ^ y[$0] }
|
||||
}
|
||||
for s in 0..<steps {
|
||||
let xs = x // shuffles read the other lanes' x of this step
|
||||
for l in 0..<32 {
|
||||
let xv = x[l], yv = y[l]
|
||||
switch kernel {
|
||||
case "probe_rotr": acc[l] = rotrVar(yv, xv)
|
||||
case "probe_shflx": acc[l] = acc[l] ^ xs[l ^ 8]
|
||||
case "probe_shl": acc[l] = yv << (xv & 31)
|
||||
case "probe_shr": acc[l] = yv >> (xv & 31)
|
||||
case "probe_bfe", "probe_bfec": acc[l] = (yv >> 7) & 0x1fff
|
||||
case "probe_andn": acc[l] = yv & ~xv
|
||||
case "probe_perm": acc[l] = perm(yv)
|
||||
case "probe_popc": acc[l] = acc[l] &+ UInt32(xv.nonzeroBitCount)
|
||||
case "probe_clz": acc[l] = acc[l] &+ clz32(xv)
|
||||
case "probe_sel": acc[l] = ((yv >> 5) & 1) != 0 ? xv : acc[l]
|
||||
case "probe_shfla": acc[l] = acc[l] ^ xs[(l + 3) & 31]
|
||||
case "probe_dot4u": acc[l] = dot4uRef(xv, yv, acc[l])
|
||||
case "probe_dot4s": acc[l] = dot4sRef(xv, yv, acc[l])
|
||||
default: fatalError("no reference for \(kernel)")
|
||||
}
|
||||
x[l] = xv &* 0x9E3779B1 &+ acc[l]
|
||||
y[l] = rotl(yv, 7) ^ (acc[l] &+ s)
|
||||
}
|
||||
}
|
||||
return (0..<32).map { acc[$0] ^ x[$0] ^ y[$0] }
|
||||
}
|
||||
|
||||
var lanes = 1 << 20, steps: UInt32 = 4096, reps = 3
|
||||
var args = Array(CommandLine.arguments.dropFirst())
|
||||
while !args.isEmpty {
|
||||
let a = args.removeFirst()
|
||||
switch a {
|
||||
case "--lanes": lanes = Int(args.removeFirst())!
|
||||
case "--steps": steps = UInt32(args.removeFirst())!
|
||||
case "--reps": reps = Int(args.removeFirst())!
|
||||
default: print("unknown argument \(a)"); exit(2)
|
||||
}
|
||||
}
|
||||
|
||||
guard let dev = MTLCreateSystemDefaultDevice() else { print("no Metal device"); exit(1) }
|
||||
let lib: MTLLibrary
|
||||
do { lib = try dev.makeLibrary(source: source, options: nil) } catch { print("compile failed: \(error)"); exit(1) }
|
||||
let queue = dev.makeCommandQueue()!
|
||||
let outBuf = dev.makeBuffer(length: lanes * 4, options: .storageModeShared)!
|
||||
let names = ["probe_alu", "probe_rotr", "probe_shflx", "probe_shl", "probe_shr", "probe_bfe", "probe_bfec", "probe_andn", "probe_perm",
|
||||
"probe_popc", "probe_clz", "probe_sel", "probe_shfla", "probe_dot4u", "probe_dot4s"]
|
||||
print("family-probe on \(dev.name), lanes \(lanes), steps \(steps), best of \(reps), GPU start-to-end time")
|
||||
print("| kernel | lanes | steps | best ms | G steps/s | ns per step | ratio to alu | warps 0 and last ok |")
|
||||
print("|---|---|---|---|---|---|---|---|")
|
||||
var aluBest = 0.0
|
||||
for name in names {
|
||||
let fn = lib.makeFunction(name: name)!
|
||||
let pso = try! dev.makeComputePipelineState(function: fn)
|
||||
let tg = min(256, pso.maxTotalThreadsPerThreadgroup)
|
||||
var best = Double.infinity
|
||||
var okAll = true
|
||||
for r in 0..<reps {
|
||||
var st = steps
|
||||
var seed = UInt32(0x2468ace) &+ UInt32(r) &* 0x9E3779B9
|
||||
let cb = queue.makeCommandBuffer()!
|
||||
let enc = cb.makeComputeCommandEncoder()!
|
||||
enc.setComputePipelineState(pso)
|
||||
enc.setBytes(&st, length: 4, index: 0)
|
||||
enc.setBytes(&seed, length: 4, index: 1)
|
||||
enc.setBuffer(outBuf, offset: 0, index: 2)
|
||||
enc.dispatchThreads(MTLSize(width: lanes, height: 1, depth: 1), threadsPerThreadgroup: MTLSize(width: tg, height: 1, depth: 1))
|
||||
enc.endEncoding()
|
||||
cb.commit()
|
||||
cb.waitUntilCompleted()
|
||||
let ms = (cb.gpuEndTime - cb.gpuStartTime) * 1000.0
|
||||
if ms < best { best = ms }
|
||||
let p = outBuf.contents().bindMemory(to: UInt32.self, capacity: lanes)
|
||||
for g0 in [UInt32(0), UInt32(lanes - 32)] {
|
||||
let want = warpRef(kernel: name, g0: g0, seed: seed, steps: steps)
|
||||
for l in 0..<32 where p[Int(g0) + l] != want[l] {
|
||||
okAll = false
|
||||
print("MISMATCH \(name) lane \(g0 + UInt32(l)): gpu \(String(p[Int(g0) + l], radix: 16)) cpu \(String(want[l], radix: 16))")
|
||||
break
|
||||
}
|
||||
}
|
||||
}
|
||||
if name == "probe_alu" { aluBest = best }
|
||||
let stepsPerS = Double(lanes) * Double(steps) / (best / 1000.0)
|
||||
let ratio = aluBest > 0 ? best / aluBest : 0
|
||||
print(String(format: "| %@ | %d | %u | %.3f | %.2f | %.3f | %.2f | %@ |", name, lanes, steps, best, stepsPerS / 1e9, best * 1e6 / Double(steps), ratio, okAll ? "yes" : "NO"))
|
||||
print(String(format: "RESULT FAMILY vendor=apple device=\"%@\" kernel=%@ lanes=%d steps=%u best_ms=%.3f gsteps_per_s=%.2f ns_per_step=%.3f ratio_alu=%.3f ok=%d", dev.name, name, lanes, steps, best, stepsPerS / 1e9, best * 1e6 / Double(steps), ratio, okAll ? 1 : 0))
|
||||
}
|
||||
print("family-probe: done")
|
||||
20
tools/ca3-reserve/make-pc2-playbook.sh
Executable file
20
tools/ca3-reserve/make-pc2-playbook.sh
Executable file
|
|
@ -0,0 +1,20 @@
|
|||
#!/usr/bin/env bash
|
||||
# Writes the PC 2 family-probe playbook: pc2-family-probe.ps1 with proto-cuda/family-probe.cu inlined at
|
||||
# CU_PLACEHOLDER, into $1 (default: $TMPDIR/pc2-family-probe.ps1). Prints the sha256 of the .cu so the job's
|
||||
# "RESULT source ... sha256" line can be checked against it. The shape follows tools/prover-floor/make-measure-playbook.sh.
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
ROOT="$(cd "$HERE/../.." && pwd)"
|
||||
OUT="${1:-${TMPDIR:-/tmp}/pc2-family-probe.ps1}"
|
||||
CU="$ROOT/proto-cuda/family-probe.cu"
|
||||
grep -q "^'@" "$CU" && { echo "the .cu has a line starting with '@ (would end the here-string)" >&2; exit 1; }
|
||||
python3 - "$HERE/pc2-family-probe.ps1" "$CU" "$OUT" <<'PY'
|
||||
import sys
|
||||
tpl, cu, out = sys.argv[1], sys.argv[2], sys.argv[3]
|
||||
s = open(tpl).read()
|
||||
c = open(cu).read().rstrip('\n')
|
||||
assert 'CU_PLACEHOLDER' in s
|
||||
open(out, 'w').write(s.replace('CU_PLACEHOLDER', c))
|
||||
print("wrote", out)
|
||||
PY
|
||||
echo "family-probe.cu sha256 $(shasum -a 256 "$CU" | cut -c1-64) bytes $(wc -c < "$CU" | tr -d ' ')"
|
||||
63
tools/ca3-reserve/pc2-family-probe.ps1
Normal file
63
tools/ca3-reserve/pc2-family-probe.ps1
Normal file
|
|
@ -0,0 +1,63 @@
|
|||
# Counter ASIC 3.0 item 6 (6 October 2026): the step cost of every reserve candidate family of spec 1.13.2 on PC 2's
|
||||
# RTX 5090 (docs/plans/counter-asic-3-reserve.md). A signed `run` job published with --stop-miners: the app stops its
|
||||
# miners before this script starts and restarts them when it ends (app/igneum-app/src/jobrun.rs, stop_miners_first);
|
||||
# the script never quits, restarts or updates the app. It switches the prover off for the run and back on at the end
|
||||
# (as tools/prover-floor/pc2-floor-measure.ps1 does), confirms the card is empty by nvidia-smi's compute-apps list
|
||||
# (never api/state, which answers {} on this machine), writes proto-cuda/family-probe.cu (inlined below by
|
||||
# make-pc2-playbook.sh, sha256 printed on both sides) into the job dir, compiles it with nvcc inside WSL2
|
||||
# Ubuntu-24.04 (/usr/local/cuda-12.*) and runs it three times. Every number is a RESULT line; the probe's own
|
||||
# "RESULT FAMILY ..." lines carry kernel, best ms, G steps/s, ns per step, the ratio to the alu chain and ok=1 for
|
||||
# bit-exact against the CPU reference on two whole warps. Read back with `node tools/jobs.mjs <job id>`.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
|
||||
function Smi($q) { try { (& nvidia-smi --query-gpu=$q --format=csv,noheader,nounits 2>$null) -join ' | ' } catch { 'nvidia-smi failed' } }
|
||||
function CudaApps { @(& nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv,noheader 2>$null | ForEach-Object { "$_" } | Where-Object { $_ -match 'igneum|sp1|prove' }) }
|
||||
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
|
||||
Start-Sleep -Seconds 30
|
||||
$t = 0; $apps = @(CudaApps)
|
||||
while ($t -lt 120 -and $apps.Count -gt 0) { Start-Sleep -Seconds 10; $t += 10; $apps = @(CudaApps) }
|
||||
if ($apps.Count -eq 0) { "RESULT card $(Stamp) empty after $t s (no igneum, sp1 or prove compute app on the card)" } else { "RESULT card $(Stamp) UNCONFIRMED after $t s: the numbers below are beside " + ($apps -join '; ') }
|
||||
"RESULT gpus $(Stamp) $(Smi 'index,name,driver_version,memory.used,clocks.sm,clocks.mem,power.draw,temperature.gpu')"
|
||||
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-ca3-family' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
$cu = @'
|
||||
CU_PLACEHOLDER
|
||||
'@
|
||||
$cuFile = Join-Path $job 'family-probe.cu'
|
||||
[IO.File]::WriteAllText($cuFile, (($cu -replace "`r`n", "`n") + "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
"RESULT source family-probe.cu sha256 $((Get-FileHash -Algorithm SHA256 $cuFile).Hash.ToLower()) bytes $((Get-Item $cuFile).Length)"
|
||||
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
|
||||
$bash = @'
|
||||
set -uo pipefail
|
||||
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"
|
||||
[ -n "$CUDA_DIR" ] || { echo "RESULT build_failed no /usr/local/cuda-12.*"; exit 2; }
|
||||
export PATH="$CUDA_DIR/bin:$PATH" LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
|
||||
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
SRC='SRC_PLACEHOLDER'
|
||||
W=/tmp/igneum-ca3-family; rm -rf "$W"; mkdir -p "$W"; cp "$SRC" "$W/family-probe.cu"; cd "$W"
|
||||
echo "RESULT nvcc $(nvcc --version | grep -o 'release [0-9.]*' | head -1) cuda_dir=$CUDA_DIR"
|
||||
echo "RESULT source_wsl sha256=$(sha256sum family-probe.cu | cut -c1-64) bytes=$(stat -c %s family-probe.cu)"
|
||||
if nvcc -O2 -arch=sm_120 -o family-probe family-probe.cu > build.log 2>&1; then echo "RESULT build ok arch=sm_120"
|
||||
else
|
||||
echo "RESULT build sm_120 failed: $(grep -m1 -i 'error' build.log | cut -c1-200)"
|
||||
if nvcc -O2 -gencode arch=compute_89,code=compute_89 -o family-probe family-probe.cu > build2.log 2>&1; then echo "RESULT build ok arch=compute_89 (PTX, driver JIT)"
|
||||
else echo "RESULT build_failed $(grep -m1 -i 'error' build2.log | cut -c1-200)"; exit 2; fi
|
||||
fi
|
||||
echo "RESULT binary sha256=$(sha256sum family-probe | cut -c1-16)"
|
||||
for i in 1 2 3; do
|
||||
echo "RESULT run $i start $(stamp) gpu_before $(nvidia-smi --query-gpu=clocks.sm,clocks.mem,power.draw,temperature.gpu --format=csv,noheader,nounits 2>/dev/null | head -1)"
|
||||
./family-probe --reps 3 2>&1 | sed "s/^RESULT /RESULT run=$i /; t; s/^/RESULT run=$i text /"
|
||||
echo "RESULT run $i end $(stamp) exit=${PIPESTATUS[0]} gpu_after $(nvidia-smi --query-gpu=clocks.sm,clocks.mem,power.draw,temperature.gpu --format=csv,noheader,nounits 2>/dev/null | head -1)"
|
||||
done
|
||||
echo "RESULT measure_end $(stamp)"
|
||||
'@
|
||||
$bash = $bash.Replace('SRC_PLACEHOLDER', (WslPath $cuFile))
|
||||
$bashFile = Join-Path $job 'measure.sh'
|
||||
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
|
||||
"RESULT gpus_after $(Stamp) $(Smi 'clocks.sm,clocks.mem,power.draw,temperature.gpu')"
|
||||
"RESULT end $(Stamp) prover back on: $(Prove $true)"
|
||||
exit 0
|
||||
|
|
@ -49,6 +49,7 @@ Created on start if missing.
|
|||
| `live_checkpoints` | one per checkpoint index, kept 7 days | `index`, `hash`, `blue_score`, `daa_score`, `state` (proposed, certified, locked), `signed_weight`, `active_weight`, `total_weight`, `fraction_active`, `fraction_total`, `votes_seen`, `voters`, `aggregators` (key hashes whose sortition proof made them aggregators), `locked_at`, `first_seen_at`, `updated_at`. |
|
||||
| `live_proofs` | one per planned shard of a chain block, kept `LIVE_RETAIN_HOURS` | `block_hash`, `shard`, `shards` (in the plan), `block_number`, `block_daa`, `block_ts`, `pgas`, `state` (planned, proving, verified, paid), `prover` (id8), `verified`, `carried_by`, `carrier_number`, `carrier_daa`, `lag_daa`, `payout_wei`, `received_at`, `updated_at`. Primary key `(block_hash, shard)`, index on `received_at`. |
|
||||
| `live_state.proving` | jsonb, updated every 2 s | `supported` (false with `reason` when the node has no proving RPCs or the endpoint is unreachable), `active`, `activation_daa`, `tip_daa`, `verifier`, `pool` (entries, pending, verified, failed), `paid_shards_total`, `shard_budget_pgas`, `blocks_10m`, `blocks_fully_proven_10m`, `shards_proven_10m`, `shards_paid_10m`, `median_proof_lag_s` (median `lag_daa` of the last 10 minutes; the devnet targets one DAA step per second), `provers_10m`, `open_blocks`, `pending_plans`, `evm_rpc`. |
|
||||
| `live_state.vendor_share` | jsonb, updated every 60 s (`tools/observer/vendor-share.mjs`, Counter ASIC 3.0 item 7) | `at`, `window_s` (600), `network_hps`, `fleet_reported` (`by_vendor` {nvidia, amd, apple, intel, unknown} with `mhs`, `workers`, `share`; `total_mhs`, `workers`, `machines`, `rows` per worker with label, machine id8, card, mhs, ages), `chain_attributed` (`by_vendor` with `blue_blocks`, `miners`, `share`, `mhs`; `blue_blocks_total`, `mapped_keys`, `unknown_miner_ids`), `by_vendor` (the two readings side by side), `coverage` (1 minus the unknown share). Fleet-reported = the sum of `now=MH/s wall` from each intake-reporting worker's newest STATUS line in the window (our fleet only); chain-attributed = each vote key's share of the window's blue blocks times the network estimate, vendor from the worker that logged the key, else unknown. `docs/benchmarks/repro.md` section 8. |
|
||||
| `live_state.finality` | jsonb, updated every 2 s | `params`, `chain_id`, `next_index`, `finality_active`, `latest_locked_index`, `latest_locked_hash`, `latest_locked_blue_score`, `weights` (`total_weight`, `active_weight`, `voters`, `keys[]` with `id`, `blocks`, `voter`, `participation`, `stripped_until_daa`, `revealed`). |
|
||||
|
||||
## Keeping it current (autosync)
|
||||
|
|
|
|||
|
|
@ -47,6 +47,7 @@ import { homedir } from 'node:os';
|
|||
import { blockSubsidy } from '../../site/lib/emission.mjs';
|
||||
import { keccak256 } from '../../site/lib/eth.mjs';
|
||||
import { makeRunner as makeDetector, optsFromEnv as detectorOpts } from './detector.mjs'; // Counter ASIC 3.0 item 4a: the share-pattern detector
|
||||
import { vendorShareTick } from './vendor-share.mjs'; // Counter ASIC 3.0 item 7: hash rate by vendor
|
||||
|
||||
const RPC = process.env.IGNEUM_RPC || 'ws://127.0.0.1:28610';
|
||||
const RETAIN_HOURS = Number(process.env.LIVE_RETAIN_HOURS || 24);
|
||||
|
|
@ -130,6 +131,8 @@ async function setupSchema() {
|
|||
`ALTER TABLE ${TS} ADD COLUMN IF NOT EXISTS rpc_load jsonb`,
|
||||
`ALTER TABLE ${TS} ADD COLUMN IF NOT EXISTS supply_check jsonb`,
|
||||
`ALTER TABLE ${TS} ADD COLUMN IF NOT EXISTS detector jsonb`,
|
||||
// Vendor share (Counter ASIC 3.0 item 7, tools/observer/vendor-share.mjs): hash rate by vendor, fleet-reported and chain-attributed
|
||||
`ALTER TABLE ${TS} ADD COLUMN IF NOT EXISTS vendor_share jsonb`,
|
||||
// One row per planned shard of a chain block (spec 7.7 item 8): the plan as the block joins the chain, then the
|
||||
// record's progress. prover is the first 8 hex characters of the record's vote key hash (never the full key, R4.6.2).
|
||||
`CREATE TABLE IF NOT EXISTS ${TP} (
|
||||
|
|
@ -1010,6 +1013,7 @@ async function provingState() {
|
|||
// of every block it merges (utxo_validation.rs:176 pays a merged block what its own payload declares; blues and reds
|
||||
// inside the DAA window alike). Written to live_state.supply_check; /api/supply shows it beside the rule's total.
|
||||
const SUPPLY_CHECK_EVERY_MS = 60 * 60_000, SUPPLY_SAMPLE = 500;
|
||||
const VENDOR_SHARE_EVERY_MS = 60_000; // live_state.vendor_share, written by tools/observer/vendor-share.mjs
|
||||
let lastSupplyCheck = null;
|
||||
async function supplyCheck() {
|
||||
try {
|
||||
|
|
@ -1143,6 +1147,7 @@ async function main() {
|
|||
// Detector (6 Oct 2026): once a minute, live_state.detector and live_events kind `detector`; tools/observer/detector.mjs
|
||||
const detector = makeDetector({ sql, TB, TS, recordEvent, log }, detectorOpts());
|
||||
setInterval(() => detector().catch(e => log('detector failed', e.message)), 60_000); setTimeout(() => detector().catch(e => log('detector failed', e.message)), 40_000);
|
||||
setInterval(() => vendorShareTick(sql, T, log), VENDOR_SHARE_EVERY_MS); setTimeout(() => vendorShareTick(sql, T, log), 30_000);
|
||||
tick(rpc); prune();
|
||||
}
|
||||
|
||||
|
|
|
|||
233
tools/observer/vendor-share.mjs
Normal file
233
tools/observer/vendor-share.mjs
Normal file
|
|
@ -0,0 +1,233 @@
|
|||
// The vendor-share metric (Counter ASIC 3.0 item 7, docs/analysis/asic-resistance-history.md section 4.3 rank 7;
|
||||
// docs/benchmarks/repro.md section 8). Hash rate by vendor (nvidia, amd, apple, intel, unknown), two ways, each
|
||||
// named for what it is:
|
||||
//
|
||||
// (a) fleet_reported: the sum of "now=X MH/s wall" from the newest STATUS line of every miner worker that reported
|
||||
// to the log intake (miner_logs, labels miner-<vendor>-<machine id8>-<card index + 1>) in the last 10 minutes.
|
||||
// It covers ONLY the machines that upload to the intake: the project's own fleet and any app install that
|
||||
// carries the intake key. It says nothing about the rest of the network.
|
||||
// (b) chain_attributed: every miner id's (vote_key_hash) share of the BLUE blocks of the last 10 minutes
|
||||
// (live_blocks.color = 'blue'; pending and red blocks are left out) times the node's network hash-rate
|
||||
// estimate (live_state.hashes_per_second_estimate), attributed to a vendor where the vote key is one a fleet
|
||||
// worker logged ("identity N '<label>-<k>' vote_key_hash=<hex>" in its upload; the vendor from the label, and
|
||||
// for miner-other-* workers from the app's "cards:" line), and to "unknown" otherwise. A key that mined and
|
||||
// belongs to no reporting machine is the honest hole in the metric.
|
||||
//
|
||||
// Exported pure functions (tested in vendor-share.test.mjs on fabricated rows); `vendorShareTick(sql, prefix, log)`
|
||||
// is the observer's one hook: it reads the three queries below (no table is written but live_state.vendor_share,
|
||||
// one UPDATE of that column every 60 s, the way live_state.proving is a jsonb column of the same row) and returns
|
||||
// the object it wrote. `node tools/observer/vendor-share.mjs --dry` prints the object from the live tables and
|
||||
// writes nothing (DATABASE_URL from ~/.config/igneum/env).
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { homedir } from 'node:os';
|
||||
|
||||
export const VENDORS = ['nvidia', 'amd', 'apple', 'intel', 'unknown'];
|
||||
export const WINDOW_S = 600;
|
||||
|
||||
/** The vendor of a card name as the app prints it in its "cards:" line. */
|
||||
export function vendorOfCard(name) {
|
||||
const n = String(name || '').toLowerCase();
|
||||
if (/nvidia|geforce|rtx|gtx|quadro/.test(n)) return 'nvidia';
|
||||
if (/amd|radeon|rx \d|gfx\d/.test(n)) return 'amd';
|
||||
if (/apple m\d|apple gpu/.test(n)) return 'apple';
|
||||
if (/intel|uhd|iris|arc/.test(n)) return 'intel';
|
||||
return 'unknown';
|
||||
}
|
||||
|
||||
/** The miner label's parts: miner-<vendor word>-<machine id8>-<n>; vendor word nvidia, amd, mac or other. */
|
||||
export function parseMinerLabel(label) {
|
||||
const m = /^miner-(nvidia|amd|mac|other)-([0-9a-f]{8})-(\d+)$/.exec(String(label || ''));
|
||||
if (!m) return null;
|
||||
return { word: m[1], machine: m[2], index: Number(m[3]) - 1 };
|
||||
}
|
||||
|
||||
/** The vendor of a worker: from the label's word, or for "other" from the machine's card list (cards: lines). */
|
||||
export function vendorOfWorker(label, cards) {
|
||||
const p = parseMinerLabel(label);
|
||||
if (!p) return 'unknown';
|
||||
if (p.word === 'nvidia' || p.word === 'amd') return p.word;
|
||||
if (p.word === 'mac') return 'apple';
|
||||
const card = cardOfWorker(label, cards);
|
||||
return card ? vendorOfCard(card) : 'unknown';
|
||||
}
|
||||
|
||||
/** The card name of a worker from the machine's "cards:" line: the entry at the label's index when it exists. */
|
||||
export function cardOfWorker(label, cards) {
|
||||
const p = parseMinerLabel(label);
|
||||
if (!p || !Array.isArray(cards) || !cards.length) return null;
|
||||
if (p.index >= 0 && p.index < cards.length) return cards[p.index];
|
||||
return cards[0];
|
||||
}
|
||||
|
||||
/** "1791271300 cards: NVIDIA GeForce RTX 5090 [discrete, mining] | AMD Radeon(TM) Graphics [integrated, off]" -> names. */
|
||||
export function parseCardsLine(line) {
|
||||
const m = /cards: (.*)$/.exec(String(line || ''));
|
||||
if (!m) return [];
|
||||
return m[1].split('|').map(s => s.replace(/\s*\[[^\]]*\]\s*$/, '').trim()).filter(Boolean);
|
||||
}
|
||||
|
||||
/** One STATUS line -> { t (unix s), label, nowMhs } ; null for any other line. */
|
||||
export function parseStatus(line) {
|
||||
const s = String(line || '');
|
||||
const m = /^(\d+(?:\.\d+)?) STATUS '([^']+)' .*? now=(\d+(?:\.\d+)?) MH\/s wall/.exec(s);
|
||||
if (!m) return null;
|
||||
return { t: Number(m[1]), label: m[2], nowMhs: Number(m[3]) };
|
||||
}
|
||||
|
||||
/** Every vote key hash a worker's upload carries ("identity N '...' vote_key_hash=<64 hex>"). */
|
||||
export function parseVoteKeys(lines) {
|
||||
const out = new Set();
|
||||
for (const m of String(lines || '').matchAll(/vote_key_hash=([0-9a-f]{64})/g)) out.add(m[1]);
|
||||
return [...out];
|
||||
}
|
||||
|
||||
function emptyByVendor(fields) {
|
||||
const o = {};
|
||||
for (const v of VENDORS) { o[v] = {}; for (const f of fields) o[v][f] = 0; }
|
||||
return o;
|
||||
}
|
||||
|
||||
/**
|
||||
* (a) fleet-reported. workers: [{ label, machine, received_at (ms or ISO), status (the newest STATUS line), keys [] }],
|
||||
* machineCards: { id8 -> [card names] }. Only uploads received within the window count; a worker whose newest STATUS
|
||||
* line is older than the window inside the upload is still counted (the upload is the freshness signal) but its
|
||||
* status age is reported.
|
||||
*/
|
||||
export function fleetReported(workers, machineCards, nowMs = Date.now(), windowS = WINDOW_S) {
|
||||
const by = emptyByVendor(['mhs', 'workers']);
|
||||
const rows = [];
|
||||
const machines = new Set();
|
||||
for (const w of workers || []) {
|
||||
const recv = typeof w.received_at === 'number' ? w.received_at : new Date(w.received_at).getTime();
|
||||
if (!(recv > nowMs - windowS * 1000)) continue;
|
||||
const p = parseMinerLabel(w.label);
|
||||
if (!p) continue;
|
||||
const cards = machineCards && machineCards[p.machine];
|
||||
const vendor = vendorOfWorker(w.label, cards);
|
||||
const st = parseStatus(w.status);
|
||||
const mhs = st ? st.nowMhs : 0;
|
||||
by[vendor].mhs += mhs; by[vendor].workers += 1;
|
||||
machines.add(p.machine);
|
||||
rows.push({ label: w.label, machine: p.machine, vendor, card: cardOfWorker(w.label, cards), mhs: Math.round(mhs * 100) / 100,
|
||||
upload_age_s: Math.round((nowMs - recv) / 1000), status_age_s: st ? Math.max(0, Math.round(nowMs / 1000 - st.t)) : null, keys: (w.keys || []).length });
|
||||
}
|
||||
let total = 0; for (const v of VENDORS) { by[v].mhs = Math.round(by[v].mhs * 100) / 100; total += by[v].mhs; }
|
||||
total = Math.round(total * 100) / 100;
|
||||
for (const v of VENDORS) by[v].share = total > 0 ? Math.round(by[v].mhs / total * 10000) / 10000 : 0;
|
||||
return { by_vendor: by, total_mhs: total, workers: rows.length, machines: machines.size, rows };
|
||||
}
|
||||
|
||||
/** vote key -> vendor from the workers' uploads (the newest upload per worker; a key seen by two workers keeps the first). */
|
||||
export function keyMapOf(workers, machineCards) {
|
||||
const map = {};
|
||||
for (const w of workers || []) {
|
||||
const p = parseMinerLabel(w.label);
|
||||
if (!p) continue;
|
||||
const vendor = vendorOfWorker(w.label, machineCards && machineCards[p.machine]);
|
||||
for (const k of w.keys || []) if (!(k in map)) map[k] = vendor;
|
||||
}
|
||||
return map;
|
||||
}
|
||||
|
||||
/**
|
||||
* (b) chain-attributed. blocks: [{ vote_key_hash, blue (count of blue blocks in the window) }], keyMap: vote key -> vendor,
|
||||
* networkHps: the node's estimate in H/s (null when unknown: shares are still computed, MH/s is null).
|
||||
*/
|
||||
export function chainAttributed(blocks, keyMap, networkHps) {
|
||||
const by = emptyByVendor(['blue_blocks', 'miners']);
|
||||
let total = 0;
|
||||
const unknownKeys = [];
|
||||
for (const b of blocks || []) {
|
||||
const n = Number(b.blue) || 0;
|
||||
if (!b.vote_key_hash || n <= 0) continue;
|
||||
const vendor = keyMap[b.vote_key_hash] || 'unknown';
|
||||
if (vendor === 'unknown') unknownKeys.push(String(b.vote_key_hash).slice(0, 8));
|
||||
by[vendor].blue_blocks += n; by[vendor].miners += 1; total += n;
|
||||
}
|
||||
const hps = networkHps === null || networkHps === undefined || !(Number(networkHps) > 0) ? null : Number(networkHps);
|
||||
for (const v of VENDORS) {
|
||||
by[v].share = total > 0 ? Math.round(by[v].blue_blocks / total * 10000) / 10000 : 0;
|
||||
by[v].mhs = hps === null ? null : Math.round(by[v].share * hps / 1e6 * 100) / 100;
|
||||
}
|
||||
return { by_vendor: by, blue_blocks_total: total, miners: (blocks || []).filter(b => b.vote_key_hash && Number(b.blue) > 0).length,
|
||||
mapped_keys: Object.keys(keyMap).length, unknown_miner_ids: unknownKeys, network_mhs: hps === null ? null : Math.round(hps / 1e6 * 100) / 100 };
|
||||
}
|
||||
|
||||
/** The whole object from rows already read. */
|
||||
export function compute({ workers, machineCards, blocks, networkHps, nowMs = Date.now(), hpsSource = null }) {
|
||||
const keyMap = keyMapOf(workers, machineCards);
|
||||
const fleet = fleetReported(workers, machineCards, nowMs);
|
||||
const chain = chainAttributed(blocks, keyMap, networkHps);
|
||||
const summary = {};
|
||||
for (const v of VENDORS) summary[v] = { fleet_reported_mhs: fleet.by_vendor[v].mhs, fleet_share: fleet.by_vendor[v].share, chain_share: chain.by_vendor[v].share, chain_attributed_mhs: chain.by_vendor[v].mhs };
|
||||
return {
|
||||
at: new Date(nowMs).toISOString(), window_s: WINDOW_S,
|
||||
network_hps: networkHps === null || networkHps === undefined ? null : Number(networkHps), network_hps_source: hpsSource,
|
||||
fleet_reported: { what: 'sum of now=MH/s wall from the newest STATUS line of each worker that uploaded to the intake in the window; intake-reporting machines only', ...fleet },
|
||||
chain_attributed: { what: 'blue blocks per miner id in the window as a share of all blue blocks, times the network hash-rate estimate; vendor from the fleet worker that logged the vote key, else unknown', ...chain },
|
||||
by_vendor: summary,
|
||||
// coverage: how much of the chain's attributed rate the fleet's own reports explain (1.0 = every blue block came from a reporting worker)
|
||||
coverage: chain.blue_blocks_total > 0 ? Math.round((1 - chain.by_vendor.unknown.share) * 10000) / 10000 : null,
|
||||
};
|
||||
}
|
||||
|
||||
// ---- the queries (read-only; the only write is the one column in vendorShareTick) ----
|
||||
// The newest upload per worker (its last STATUS line and run id) and the vote keys that upload still carries. A worker
|
||||
// logs its identity lines once at start, and an upload is the last 256 KB of its log, so after a few hours they have
|
||||
// scrolled out: the keys of a run come from its first uploads (KEYS_SQL), read once per (label, run_id) and cached.
|
||||
const WORKERS_SQL = `
|
||||
SELECT label, machine, run_id, received_at,
|
||||
reverse(substring(reverse(lines) from '[^\n]* SUTATS [^\n]*')) AS status,
|
||||
array(SELECT DISTINCT m[1] FROM regexp_matches(lines, 'vote_key_hash=([0-9a-f]{64})', 'g') m) AS keys
|
||||
FROM (SELECT DISTINCT ON (label) label, machine, run_id, received_at, lines FROM miner_logs
|
||||
WHERE label LIKE 'miner-%' AND received_at > now() - interval '1 day' ORDER BY label, received_at DESC) s`;
|
||||
const KEYS_SQL = `
|
||||
SELECT array(SELECT DISTINCT m[1] FROM regexp_matches(string_agg(lines, E'\n'), 'vote_key_hash=([0-9a-f]{64})', 'g') m) AS keys
|
||||
FROM (SELECT lines FROM miner_logs WHERE label = $1 AND run_id = $2 ORDER BY received_at ASC LIMIT 3) s`;
|
||||
const runKeys = new Map(); // `${label}|${run_id}` -> [vote key hashes], for this process's lifetime
|
||||
const CARDS_SQL = `
|
||||
SELECT label, reverse(substring(reverse(lines) from '[^\n]* :sdrac [^\n]*')) AS cards
|
||||
FROM (SELECT DISTINCT ON (label) label, lines FROM miner_logs
|
||||
WHERE (label LIKE 'win-%' OR label LIKE 'mac-%' OR label LIKE 'linux-%') AND received_at > now() - interval '1 day' ORDER BY label, received_at DESC) s`;
|
||||
const BLOCKS_SQL = (T) => `
|
||||
SELECT vote_key_hash, count(*) FILTER (WHERE color = 'blue')::int AS blue
|
||||
FROM ${T}live_blocks WHERE received_at > now() - interval '${WINDOW_S} seconds' AND vote_key_hash IS NOT NULL GROUP BY 1`;
|
||||
const STATE_SQL = (T) => `SELECT hashes_per_second_estimate FROM ${T}live_state WHERE id = 1`;
|
||||
|
||||
/** Reads the live tables (read-only) and returns the object; `sql(query, params)` returns rows. T = table prefix. */
|
||||
export async function readVendorShare(sql, T = '') {
|
||||
const [workers, cardRows, blocks, state] = await Promise.all([sql(WORKERS_SQL), sql(CARDS_SQL), sql(BLOCKS_SQL(T)), sql(STATE_SQL(T))]);
|
||||
for (const w of workers) {
|
||||
const k = `${w.label}|${w.run_id}`;
|
||||
if (!runKeys.has(k)) { const r = await sql(KEYS_SQL, [w.label, w.run_id]); runKeys.set(k, (r[0] && r[0].keys) || []); }
|
||||
w.keys = [...new Set([...(w.keys || []), ...runKeys.get(k)])];
|
||||
}
|
||||
const machineCards = {};
|
||||
for (const r of cardRows) { const m = /^(?:win|mac|linux)-([0-9a-f]{8})$/.exec(r.label); if (m && r.cards) machineCards[m[1]] = parseCardsLine(r.cards); }
|
||||
const hps = state.length && state[0].hashes_per_second_estimate !== null ? Number(state[0].hashes_per_second_estimate) : null;
|
||||
return compute({ workers, machineCards, blocks, networkHps: hps, hpsSource: 'live_state.hashes_per_second_estimate (the node\'s estimateNetworkHashesPerSecond, else blue work per second over 10 min)' });
|
||||
}
|
||||
|
||||
/** The observer's hook: read, write live_state.vendor_share, return the object. Errors are logged, never thrown. */
|
||||
export async function vendorShareTick(sql, T = '', log = () => {}) {
|
||||
try {
|
||||
const v = await readVendorShare(sql, T);
|
||||
await sql(`UPDATE ${T}live_state SET vendor_share = $1::jsonb WHERE id = 1`, [JSON.stringify(v)]);
|
||||
return v;
|
||||
} catch (e) { log('vendor share failed', e.message); return null; }
|
||||
}
|
||||
|
||||
// ---- CLI: --dry prints the live reading and writes nothing ----
|
||||
if (process.argv[1] && process.argv[1].endsWith('vendor-share.mjs') && process.argv.includes('--dry')) {
|
||||
const env = readFileSync(`${homedir()}/.config/igneum/env`, 'utf8');
|
||||
const m = /^DATABASE_URL=(.*)$/m.exec(env);
|
||||
if (!m) { console.error('DATABASE_URL not found in ~/.config/igneum/env'); process.exit(1); }
|
||||
const url = m[1].trim().replace(/^['"]|['"]$/g, '');
|
||||
const host = new URL(url).hostname.replace('-pooler', '');
|
||||
const sql = async (query, params = []) => {
|
||||
const r = await fetch(`https://${host}/sql`, { method: 'POST', headers: { 'Neon-Connection-String': url, 'Content-Type': 'application/json' }, body: JSON.stringify({ query, params }) });
|
||||
const j = await r.json(); if (!r.ok) throw new Error(j.message || JSON.stringify(j)); return j.rows || [];
|
||||
};
|
||||
const T = (process.env.LIVE_TABLE_PREFIX || '').replace(/[^a-z0-9_]/gi, '');
|
||||
console.log(JSON.stringify(await readVendorShare(sql, T), null, 1));
|
||||
}
|
||||
79
tools/observer/vendor-share.test.mjs
Normal file
79
tools/observer/vendor-share.test.mjs
Normal file
|
|
@ -0,0 +1,79 @@
|
|||
// node --test tools/observer/vendor-share.test.mjs
|
||||
// Fabricated rows: one known-finished case (two reporting workers, every blue block mapped) and one known-failed
|
||||
// case (a vote key no reporting machine owns lands in "unknown"; a stale upload is left out; a 0 MH/s worker counts 0).
|
||||
import { test } from 'node:test';
|
||||
import assert from 'node:assert/strict';
|
||||
import { compute, parseStatus, parseVoteKeys, parseCardsLine, vendorOfWorker, cardOfWorker, VENDORS } from './vendor-share.mjs';
|
||||
|
||||
const NOW = Date.parse('2026-10-06T08:00:00Z');
|
||||
const status = (label, mhs, ageS = 20) => `${(NOW / 1000 - ageS).toFixed(3)} STATUS '${label}' [worker]: 20212s jobs=137109 accepted=16053 rejected=0 fee=152 mismatched=0 extra=1059 rate=0.79 blocks/s hash=113.81 MH/s wall (114.08 MH/s inside jobs) now=${mhs.toFixed(2)} MH/s wall (113.92 MH/s inside jobs, 204 jobs, seed walk 0 calls) template_age=0.80s synced=true idle=0.2% (last 30s: 0.0%) queued=1 restarts=0 faults=0 identities=2 accepted_by_identity=8044/8009`;
|
||||
const K1 = 'a'.repeat(64), K2 = 'b'.repeat(64), K3 = 'c'.repeat(64), K4 = 'd'.repeat(64), KX = 'e'.repeat(64);
|
||||
const cards = {
|
||||
'1ccfe586': parseCardsLine('1791271300 cards: NVIDIA GeForce RTX 5090 [discrete, mining] | AMD Radeon(TM) Graphics [integrated, off]'),
|
||||
'ae432dc7': parseCardsLine('1791271259 cards: NVIDIA GeForce RTX 5090 [discrete, mining] | AMD Radeon(TM) Graphics [integrated, off] | AMD Radeon RX 9070 XT [discrete, mining]'),
|
||||
'37ba0461': parseCardsLine('1791271315 cards: Intel(R) UHD Graphics [integrated, mining]'),
|
||||
'd937c69d': parseCardsLine('1791271752 cards: Apple M5 Max [apple, off]'),
|
||||
};
|
||||
|
||||
test('the parsers read the app\'s own line shapes', () => {
|
||||
const s = parseStatus(status('nvidia-1ccfe586-1', 113.87));
|
||||
assert.equal(s.label, 'nvidia-1ccfe586-1'); assert.equal(s.nowMhs, 113.87);
|
||||
assert.equal(parseStatus('1791270154.425 identity 8 \'nvidia-ae432dc7-1-8\' vote_key_hash=' + K1), null);
|
||||
assert.deepEqual(parseVoteKeys(`x\n1 identity 1 'a' vote_key_hash=${K1} pubkey=ff\n2 identity 2 'a' vote_key_hash=${K2}\n3 identity 1 'a' vote_key_hash=${K1}`), [K1, K2]);
|
||||
assert.deepEqual(cards['ae432dc7'], ['NVIDIA GeForce RTX 5090', 'AMD Radeon(TM) Graphics', 'AMD Radeon RX 9070 XT']);
|
||||
assert.equal(vendorOfWorker('miner-amd-ae432dc7-3', cards['ae432dc7']), 'amd');
|
||||
assert.equal(cardOfWorker('miner-amd-ae432dc7-3', cards['ae432dc7']), 'AMD Radeon RX 9070 XT');
|
||||
assert.equal(vendorOfWorker('miner-other-37ba0461-1', cards['37ba0461']), 'intel');
|
||||
assert.equal(vendorOfWorker('miner-other-37ba0461-1', undefined), 'unknown');
|
||||
assert.equal(vendorOfWorker('miner-mac-d937c69d-1', cards['d937c69d']), 'apple');
|
||||
});
|
||||
|
||||
test('known-finished: two workers, every blue block mapped, both readings agree on the vendor order', () => {
|
||||
const workers = [
|
||||
{ label: 'miner-nvidia-1ccfe586-1', machine: 'DESKTOP-1ccfe586', received_at: NOW - 30_000, status: status('nvidia-1ccfe586-1', 113.87), keys: [K1, K2] },
|
||||
{ label: 'miner-amd-ae432dc7-3', machine: 'DESKTOP-ae432dc7', received_at: NOW - 45_000, status: status('amd-ae432dc7-3', 19.09), keys: [K3] },
|
||||
];
|
||||
const blocks = [{ vote_key_hash: K1, blue: 45 }, { vote_key_hash: K2, blue: 40 }, { vote_key_hash: K3, blue: 15 }];
|
||||
const v = compute({ workers, machineCards: cards, blocks, networkHps: 133e6, nowMs: NOW, hpsSource: 'test' });
|
||||
assert.equal(v.fleet_reported.total_mhs, 132.96);
|
||||
assert.equal(v.fleet_reported.by_vendor.nvidia.mhs, 113.87);
|
||||
assert.equal(v.fleet_reported.by_vendor.amd.mhs, 19.09);
|
||||
assert.equal(v.fleet_reported.workers, 2); assert.equal(v.fleet_reported.machines, 2);
|
||||
assert.equal(v.chain_attributed.blue_blocks_total, 100);
|
||||
assert.equal(v.chain_attributed.by_vendor.nvidia.share, 0.85);
|
||||
assert.equal(v.chain_attributed.by_vendor.nvidia.mhs, 113.05);
|
||||
assert.equal(v.chain_attributed.by_vendor.amd.share, 0.15);
|
||||
assert.equal(v.chain_attributed.by_vendor.unknown.blue_blocks, 0);
|
||||
assert.equal(v.coverage, 1);
|
||||
assert.deepEqual(Object.keys(v.by_vendor), VENDORS);
|
||||
assert.equal(v.fleet_reported.rows[1].card, 'AMD Radeon RX 9070 XT');
|
||||
});
|
||||
|
||||
test('known-failed: a vote key no reporting machine owns lands in unknown; a stale upload is left out; 0 MH/s counts 0', () => {
|
||||
const workers = [
|
||||
{ label: 'miner-nvidia-1ccfe586-1', machine: 'DESKTOP-1ccfe586', received_at: NOW - 30_000, status: status('nvidia-1ccfe586-1', 113.87), keys: [K1, K2] },
|
||||
// the Mac paused 40 minutes ago: its upload is outside the window and must not count, but its keys still map
|
||||
{ label: 'miner-mac-d937c69d-1', machine: 'MacBook-Pro-d937c69d', received_at: NOW - 40 * 60_000, status: status('mac-d937c69d-1', 27.9, 2400), keys: [K4] },
|
||||
// a worker whose newest STATUS line says 0 (the watchdog's zero case)
|
||||
{ label: 'miner-other-37ba0461-1', machine: 'DESKTOP-37ba0461', received_at: NOW - 10_000, status: status('other-37ba0461-1', 0), keys: [] },
|
||||
];
|
||||
const blocks = [{ vote_key_hash: K1, blue: 50 }, { vote_key_hash: K2, blue: 30 }, { vote_key_hash: KX, blue: 20 }, { vote_key_hash: K4, blue: 0 }];
|
||||
const v = compute({ workers, machineCards: cards, blocks, networkHps: 200e6, nowMs: NOW });
|
||||
assert.equal(v.fleet_reported.workers, 2, 'the stale Mac upload is out');
|
||||
assert.equal(v.fleet_reported.by_vendor.apple.mhs, 0);
|
||||
assert.equal(v.fleet_reported.by_vendor.intel.mhs, 0);
|
||||
assert.equal(v.fleet_reported.by_vendor.intel.workers, 1);
|
||||
assert.equal(v.fleet_reported.total_mhs, 113.87);
|
||||
assert.equal(v.chain_attributed.by_vendor.unknown.blue_blocks, 20);
|
||||
assert.equal(v.chain_attributed.by_vendor.unknown.share, 0.2);
|
||||
assert.equal(v.chain_attributed.by_vendor.unknown.mhs, 40);
|
||||
assert.deepEqual(v.chain_attributed.unknown_miner_ids, [KX.slice(0, 8)]);
|
||||
assert.equal(v.chain_attributed.by_vendor.apple.blue_blocks, 0, 'a key with no blue blocks adds nothing');
|
||||
assert.equal(v.chain_attributed.mapped_keys, 3, 'the paused Mac\'s key is still in the map');
|
||||
assert.equal(v.coverage, 0.8);
|
||||
// no network estimate: shares stand, MH/s is null
|
||||
const w = compute({ workers, machineCards: cards, blocks, networkHps: null, nowMs: NOW });
|
||||
assert.equal(w.chain_attributed.by_vendor.nvidia.share, 0.8);
|
||||
assert.equal(w.chain_attributed.by_vendor.nvidia.mhs, null);
|
||||
assert.equal(w.network_hps, null);
|
||||
});
|
||||
Loading…
Reference in a new issue