diff --git a/docs/analysis/chip-model-v3.md b/docs/analysis/chip-model-v3.md
index 0c005d97..0510d7ce 100644
--- a/docs/analysis/chip-model-v3.md
+++ b/docs/analysis/chip-model-v3.md
@@ -23,6 +23,7 @@ is a measurement of a chip; every GPU figure says where it was measured. "Approx
| Cache mirror plus a 96 MB hot table, N5 headline | 175 mm^2, $68 | same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate) |
| 512 MiB and 1 GiB mirrors, N5 headline | 255 mm^2 and 510 mm^2; $111 to $306 | same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node) |
| GPU-class die | 750 mm^2 (the equal-silicon comparison) | M16 section 3 |
+| CPU verifier, one M5 Max core (loaded, load average 5.6; ratios are the measurement) | v2 1.31 to 1.36 ms per unit, x4 1.92 to 1.96 (1.45x), x8 2.79 (2.1x); worst cold 1.58 / 2.04 / 2.94 ms | `docs/plans/mixer-x4.md` section 6.4, 5 October 2026 21:40 UTC |
## 2. The rows
diff --git a/docs/bench-log.md b/docs/bench-log.md
index 8a82a73b..8550d873 100644
--- a/docs/bench-log.md
+++ b/docs/bench-log.md
@@ -1645,3 +1645,22 @@ Rates (5 dispatches of 2^24 after a warm-up; 5090 `--bench --block-warps 1`, 907
| hot96k4a | 114.4 | 0.84 | 14.56 | 0.80 | 1 |
Reading: the probe promises a full hit rate on the 5090 (every S inside the 96 MiB L2 at one ceiling, 6.4x DRAM) and the hash gets 2 to 8% at k = 4 and 20% at k = 8; the 9070 XT the same shape. The dataset's random lines evict the table from the shared cache on every card. The added form costs 13 to 20% of the rate. Recommendation in `docs/plans/hot-table.md` section 6.4: do not adopt layer 5 in either form on these measurements.
+## 5 October 2026 (night), mixer x4 and the cache growth rule: the class v3 dataset construction, with the x8 candidate (Counter ASIC 2.0; branch ca2-mixer on ca2-v3 6c75dad; cryptographer's lane)
+
+Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0. Write-up `docs/plans/mixer-x4.md`; chip model `docs/analysis/chip-model-v3.md`; code `igneum-pow` (LoadClass mixer_mult and growth, memhard::Shape, the schedule, the three emitters), packs `proto-cuda/packs-ca2-mixer/`, tests `igneum-pow/tests/mixer.rs` and `tests/packs.rs`. Commits 0fc0ad1, 66eeba3, e4c04a7, 7ce8d1e, 504cae4, fe4e193 and this entry's.
+
+What changed. Under program class v3 (`V3_CLASS = LoadClass::MX4`) every mixer application of the item derivation is `m = 4` applications with round keys `(r m + j + 1) x 0x9E3779B9`, the 8 dependent cache reads per item unchanged; the cache doubles when the dataset doubles (`growth_doublings(d) = floor(log2(1 + d / 1460))`: 2^26 words to day 1,459, 2^27 from day 1,460, 2^28 from day 4,380). Version 2 is byte-identical: fresh exports of igneum-genesis-mh and igneum-devnet-v4-epoch0 `diff -r` IDENTICAL against the checked-in packs, and the crate tests regenerate every pinned file. A v3 program of a seed is the v2 program of that seed instruction for instruction (v2 loads take no width roll); only the dataset words and the hashes change.
+
+Bit-exactness, `with-lock.sh run`, 22:05 and 21:45 UTC: the two pinned v3 packs (mx4-genesis, mx4-devnet-epoch0: dataset words 0..15 `61ff2180 0d4c7e6c ...` and `afe80d67 b9fbd029 ...`, word MASK `5020180e` and `e6a99c7a`, unit at base 0 lane 0 `63acd2d273f475ba` and `212c6442b51e87ae`) and the two x8 candidate packs on Metal (`packbench`, built from this branch) and Apple OpenCL (`igneum-bench-cl --bench-pack`): 3/3 standalone and 3/3 in batch, 96 of 96 lanes, cache FNV-1a 64 unchanged from v2 (48c4f5bf24166b2e, 448274a57f508cbc), dataset head, word MASK and 64 samples PASS, one 2^24 fingerprint per pack across both harnesses (mx4 6f48d5a2aa0dbe5f and 73caaebb28e808fe; mx8 7c28cfb06c5c65a9 and bbb183f72692f840); hash rate the v2 rate (27.5 to 27.7 MH/s GPU time, the hash kernel is unchanged). Fuzz: 200 class v3 programs (4 units each across the 32-bit range, one in the top 256 nonces) interpreted twice on the CPU, 800 of 800; the same 200 packs on Metal 200 of 200 (`--batch-log2 9 --batch-base 4294967040`, the wrapping unit inside the window), every tenth on Apple OpenCL 20 of 20; x8: 50 of 50 on Metal, 5 of 5 on OpenCL. Stats (8,192 outputs per seed, two seeds): v3 avalanche 49.97 to 49.99 percent, worst bit z 1.92 to 3.09, 0 duplicates (v2 beside it 49.87 to 49.98, z 2.25 to 2.30). Edges: items 0, 1, 2^28 - 1, 2^32 - 1 by hand at m = 1, 2, 4, 8; words 0, 15, 16, 17, MASK - 1, MASK through the fetch path. Determinism: two epochs, every vector and file equal and equal to the pinned pack. The scratch soundness tests of ca2-soundness (cherry-pick 0d8f745) 7 of 7 on this tree. Crate: 44 lib + 12 packs + 4 mixer + 7 scratch tests pass. A first Metal fuzz run reported 200 of 200 FAIL on an empty RESULT line (a packbench built before the `--batch-base` cherry-pick); it was read as a failure, the harness rebuilt, the run repeated.
+
+Timings, `with-lock.sh measure`, one session 21:40:12 to 21:40:23 UTC, one core, two rounds; the box carried a load average of 5.6 (one minute) and 26 (fifteen minutes) from unlocked processes, so the absolute figures are about 2.2x the quiet readwidth night's 0.604 ms v2 row and the ratios are the measurement:
+
+| Construction | Verifier ms per 32-lane unit, avg of 50 (two rounds) | Worst cold unit | Against v2 | 256 MiB fill, one core | Metal 1 GiB build, GPU ms |
+|---|---|---|---|---|---|
+| v2 | 1.361 / 1.310 | 1.579 | 1 | 172 to 173 ms | 29.7 (first touch) / 21.0 |
+| x4 (class v3) | 1.956 / 1.923 | 2.043 | 1.45x | 172 to 175 ms | 20.9 / 21.0 |
+| x8 (candidate) | 2.785 / 2.790 | 2.942 | 2.09x | 172 ms | 21.9 / 21.9 |
+
+Reading: the mixer multiplies the verifier's ALU part only (the 8 dependent misses per item are unchanged), hence 1.45x and 2.1x and not 4x and 8x; the Mac's GPU build is latency-bound and does not move with the mixer, so the "under 1 s on every discrete card" half of the x8 rule is the PC job (five packs, `relay/playbooks/mixer-x4-pc1-bench.ps1`, waiting for the go). Verification throughput (C19): a quiet 2026 core serves about 1,100 shares per second at x4 and 800 at x8 (1,660 at v2, re-cutting spec 09's 2,270), a 22,000-member pool at one share per 10 s needs 2 cores at x4 and 3 at x8, IBD over 108,000 headers is 1.6 min at x4 and 2.3 at x8 on that core; the 10 ms gate keeps 8.0 ms (x4) and 7.1 ms (x8) of margin on the loaded core, 6 to 7 ms on a 2019-class laptop core (approximate, unmeasured, O-1.14).
+
+Chip model (`docs/analysis/chip-model-v3.md`): the on-die-cache recompute chip at 50 T op/s against the 5090's measured 136.1 MH/s: v2 334 MH/s, 2.45x bare, 7.4x with the 3x fixed-function factor; x4 83.5 MH/s, 0.61x bare, 1.84x with the factor, 1.53x with the 128 mm^2 N5 mirror deducted at equal silicon; x8 41.7 MH/s, 0.31x, 0.92x, 0.76x. The claim at x4 is "under 2x" with the margin thin on the equal-budget convention (a 3.3x factor or a 10 percent larger budget reads 2.0x); the hot table in the added form would have raised it to 2.1x to 2.2x at the 5090's g (kept as measured, not adopted). Nothing here is a measurement of a chip.
diff --git a/docs/plans/mixer-x4.md b/docs/plans/mixer-x4.md
index f4c92f58..ce91b34d 100644
--- a/docs/plans/mixer-x4.md
+++ b/docs/plans/mixer-x4.md
@@ -236,10 +236,61 @@ x4 build on the GPU is the row the measure lock will settle.
| Metal and Apple OpenCL on the two pinned v3 packs | section 6.2 | 3/3 standalone, 3/3 in batch, 96 of 96 lanes, dataset head, MASK word and 64 samples, one fingerprint per pack across both harnesses |
| Metal fuzz: the 200 packs, 4 units each standalone and the top-256 unit inside a 512-nonce batch at base 4,294,967,040; every tenth pack on Apple OpenCL as well | `packbench --pack
--batches 1 --batch-log2 9 --batch-base 4294967040` (Metal), `igneum-bench-cl-igneum-genesis-mh --bench-pack --pack --batches 1 --batch-log2 10` (Apple OpenCL), `with-lock.sh run`, 21:22 to 21:24 UTC | Metal 200 of 200 packs PASS (800 of 800 standalone units, 200 of 200 inside the wrapping window, cache and dataset self-tests on every pack); Apple OpenCL 20 of 20 packs PASS (the three vectors.h units, the self-tests); a first run with a packbench built before the `--batch-base` cherry-pick reported 200 of 200 FAIL on an empty RESULT line and was read as such (the watcher rule), the harness rebuilt and the run repeated |
-### 6.4 Timings (`with-lock.sh measure`)
+### 6.4 Timings (`with-lock.sh measure`, one session, 21:40:12 to 21:40:23 UTC, commit 504cae4)
-Owed until the measure lock frees (the session script takes v2, x4 and x8 together: verifier avg of 50 and the
-three cold units, the 256 MiB fill on one core, the Metal 1 GiB build, two rounds each).
+Script `measure-v3.sh` (session scratchpad): `igneum-pow bench --seed igneum-genesis --day 2026-10-03 --warps 50`
+(v2), `... --program-class v3` (x4), `... --class mx8` (x8), two rounds each, then the devnet seeds, then
+`packbench --pack --batches 2 --batch-log2 22 --group 256` on igneum-genesis-mh, mx4-genesis and mx8-genesis,
+two rounds. The lock was exclusive among the agents' builds and measurements, but the box was not quiet: load
+average 5.6 (one minute) and 26 (fifteen minutes) at the start, from unlocked processes (the devnet node, other
+agents' editors); the v2 row reads 1.31 to 1.36 ms where the quiet readwidth night read 0.604 to 0.626. So the
+absolute numbers below are a loaded-core figure, about 2.2x the quiet one, and the ratios between the rows are the
+measurement (two rounds within 4 percent). A quiet-box re-run is owed (section 9).
+
+| Construction | Verifier, ms per 32-lane unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | 256 MiB cache fill, one core | Metal 1 GiB dataset build, GPU ms (round 1 / round 2) |
+|---|---|---|---|---|---|
+| v2 (igneum-genesis) | 1.361 / 1.310 | 1.579 | 1 | 172.1 / 172.6 ms | 29.7 / 21.0 |
+| x4 (mx4, class v3) | 1.956 / 1.923 | 2.043 | 1.45x | 175.3 / 172.3 ms | 20.9 / 21.0 |
+| x8 (mx8) | 2.785 / 2.790 | 2.942 | 2.09x | 172.3 / 172.3 ms | 21.9 / 21.9 |
+| x4, the devnet seeds (mx4-devnet-epoch0) | 1.923 | 2.012 | | 173.9 ms | |
+| x8, the devnet seeds | 2.972 | 2.885 | | 173.6 ms | |
+
+Reading. The verifier's ALU part is what grows: x4 adds 0.6 ms per unit for 27 more mixer applications on each of
+4,096 items (110,592 applications, about 5.5 ns each on this core, the lanes' chains interleaved), x8 another
+0.85 ms for 36 more; the latency part (8 dependent misses per item) is the same in every row, which is why the
+measured ratios are 1.45x and 2.1x and not the 4x and 8x of the M16 table's scaling. Shape B (32 rounds of one
+read, section 1) would have multiplied the latency part too; x4 is under the 4.8 ms bar even on the loaded core, so
+B stays unimplemented. The cache fill does not depend on the mixer (it is the ChaCha chain): 172 to 175 ms, the
+spec's 175 to 181 ms of 1.8.3. The Metal 1 GiB build does not move with the mixer at all (21 ms at v2, x4 and x8
+once warm; the 29.7 ms first v2 run is the first-touch cost the hosts fill twice for): on this card the build is
+bound by the 8 dependent cache-line reads per item, not by the arithmetic, so the Mac says nothing about whether
+the 5090's or the 9070 XT's build is arithmetic-bound; that is the PC job (section 8).
+
+Against the x4 / x8 rule (section 6.5): the verifier half passes for x8 with 7.1 ms of the 10 ms gate to spare on
+this loaded core (worst cold 2.94 ms; the quiet-core figure would be about 1.3 ms, scaled by the 2.2x of the v2
+row, approximate); x4 leaves 8.0 ms. The build half waits on the PC rows.
+
+### 6.5 Verification throughput per tier (consequences row C19), from the loaded-core figures above
+
+Warps verified per second on one core = 1,000 / (ms per warp); a pool core verifying members' shares handles that
+many shares per second; a node verifies a block with one unit (plus the header path, under 0.1 ms, not measured
+here); IBD over the 108,000-header pruning window (spec 02) on one core = 108,000 x ms per warp.
+
+| Figure | v2 | x4 | x8 | Note |
+|---|---|---|---|---|
+| ms per warp, steady (this session, loaded core) | 1.33 | 1.94 | 2.79 | avg of the two rounds |
+| ms per warp, quiet M5 Max core (scaled by 0.604 / 1.33 = 0.45, approximate) | 0.60 | 0.88 | 1.26 | the readwidth night's v2 figure is measured; x4 and x8 scaled |
+| ms per warp, 2019-class laptop core (approximate: 2.5x the quiet M5 Max figure, the ratio the design document assumes for the gate; unmeasured, O-1.14) | 1.5 | 2.2 | 3.2 | the figure that fixes the gate is a measurement, not this row |
+| Shares per second per core (loaded / quiet, approximate) | 750 / 1,660 | 515 / 1,140 | 358 / 790 | spec 09 section 9.8 item 5 carried 2,270 at v2; re-cut from the quiet row: 1,660 |
+| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second), loaded / quiet | 2.9 / 1.3 | 4.3 / 1.9 | 6.1 / 2.8 | |
+| Node: worst cold single unit (loaded core) | 1.58 ms | 2.04 ms | 2.94 ms | per block |
+| IBD over 108,000 headers on one core, loaded / quiet, minutes | 2.4 / 1.1 | 3.5 / 1.6 | 5.0 / 2.3 | laptop (approximate): 2.7 / 4.0 / 5.8 min; a seed VM core (unmeasured) sits between the laptop and the quiet M5 Max |
+| Margin left under the 10 ms gate for Counter ASIC 3.0 (worst cold, loaded core) | 8.4 ms | 8.0 ms | 7.1 ms | on the 2019-class laptop row (approximate) 7.5 / 6.8 / 5.9 ms steady |
+
+Reading: at x4 a pool core serves about 1,100 shares per second on a quiet 2026 core (a 22,000-member pool needs
+two cores); at x8 about 800 (three cores). A node's block verification stays a few milliseconds. The gate's
+remaining margin is what Counter ASIC 3.0 has to spend, and on the unmeasured laptop core it is 6 to 7 ms at x4 and
+about 6 at x8, which is the number the 2019-class measurement (O-1.14) must confirm before x8 is final.
### 6.5 The daily build per tier, and the x4 / x8 rule
@@ -271,8 +322,30 @@ prepare), or to restart per epoch. The node agent is asked whether the per-day r
## 8. What is unverified
-Filled at the end.
+1. The 5090's and the 9070 XT's dataset build at x4 and x8 (the "under 1 s on every discrete card" half of the x8
+ rule): the PC 1 job (`relay/playbooks/mixer-x4-pc1-bench.ps1`, package `mixer-x4-pcjob.zip` sha256
+ ec3be97e...bdf87, five packs, the worker's `cache ... dataset ... ms` line per pack) waits for the coordinator's
+ go after the era job; by the M16 arithmetic the 5090 is 54 ms at x4 and 107 ms at x8 if its build is
+ arithmetic-bound and 13.4 ms if it is latency-bound like the Mac, both far under 1 s; the gfx1036 is the tier
+ that fails (section 6.5 of the build table), and its consequence (per-day dataset reuse in the workers) is with
+ the node agent.
+2. The absolute verifier figures were taken on a loaded core (load average 5.6); the ratios are the measurement
+ and the quiet-core figures are scaled. A 2019-class laptop core has not run any construction (O-1.14).
+3. The mixer has had no cryptanalysis (spec 1.8.4); `m` applications with distinct round keys is `m` times the
+ work only if no shortcut composes them, which is the same open question as for one application.
+4. The x8 packs are generator 2 with the class in the id (`--class mx8`); if x8 is chosen, the pinned v3 packs are
+ re-cut through the seam (`V3_CLASS = MX8`, generator 3) and the tests re-pinned, one commit.
+5. The 5090's rate for the v3 program is the v2 rate by construction (the hash kernel is unchanged, the Mac shows
+ 27.7 MH/s at v2, x4 and x8); the chip row's denominator stays the readwidth table's 136.1 MH/s until a v3 pack
+ runs on the card, which the PC job also gives.
## 9. Owed
-Filled at the end.
+| Item | Owner | When |
+|---|---|---|
+| PC 1 run of the five packs (the build-time rows, the v3 fingerprints on NVIDIA and AMD) | ca2-mixer, on the coordinator's go | after the era job, about 22:05 UTC |
+| Quiet-box re-run of the verifier session for absolute numbers | ca2-mixer | when the Mac is quiet |
+| The x4 / x8 choice recorded from the rule, then the vectors re-cut once through the seam | coordinator, then ca2-mixer | after the PC rows |
+| `Epoch::chain_dataset_day` wired to the genesis day index in the node (`days_since_genesis(day_index(header), day_index(genesis))`) and `pow_genesis_dataset_log2` in the override | ca2-node | the integration |
+| The spec text of section 2 into `docs/spec/01-lottery-hash.md` 1.8.5 and 1.13.3 (with the v3 vectors into 1.17) | the integration | after the choice |
+| The 2019-class laptop core measurement that fixes the gate (O-1.14) | cryptographer | gate 1 |
diff --git a/relay/playbooks/mixer-x4-pc1-bench.ps1 b/relay/playbooks/mixer-x4-pc1-bench.ps1
new file mode 100644
index 00000000..6a92de28
--- /dev/null
+++ b/relay/playbooks/mixer-x4-pc1-bench.ps1
@@ -0,0 +1,95 @@
+# Igneum run job: the class v3 dataset construction, mixer x4 and the x8 candidate (docs/plans/mixer-x4.md), on PC 1
+# (machine ae432dc7): the RTX 5090 through igneum-worker-cuda.exe --bench, then the AMD card through
+# igneum-worker-opencl.exe --bench-pack, each card switched off in the app ONLY while it is under test and restored
+# after with the settings it had (the ca2-era-pc1.ps1 shape). 5 October 2026. The AMD card: the RX 9070 XT (gfx1201)
+# when its eGPU box is on the bus, else the integrated gfx1036. The number this job is for: the worker's own
+# `cache ... dataset ... ms` line per pack, the daily 1 GiB dataset build time at x1, x4 and x8 (the coordinator's rule:
+# x8 enters v3 only if that build stays under 1 s on every discrete card we own), plus the vectors and the 2^24
+# fingerprint per pack (bit-exactness against the Mac: mx4-genesis 6f48d5a2aa0dbe5f, mx4-devnet-epoch0 73caaebb28e808fe,
+# mx8-genesis 7c28cfb06c5c65a9, mx8-devnet-epoch0 bbb183f72692f840, v2-genesis-mh 25f96e7dce90bd4e).
+# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining. Packs:
+# proto-cuda/packs-ca2-mixer/{mx4-genesis, mx4-devnet-epoch0, mx8-genesis, mx8-devnet-epoch0} from branch ca2-mixer and
+# proto-cuda/packs/igneum-genesis-mh as v2-genesis-mh (the control). Every result line starts with RESULT.
+$ErrorActionPreference = 'Continue'
+function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
+$jobs = Split-Path $env:IGNEUM_JOB_DIR
+$fetched = Join-Path $jobs 'fetch-mixer-x4-20261005'
+$cuda = Join-Path $fetched 'igneum-worker-cuda.exe'
+$ocl = Join-Path $fetched 'igneum-worker-opencl.exe'
+$packs = Join-Path $fetched 'packs-ca2-mixer'
+$packList = @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0', 'mx8-genesis', 'mx8-devnet-epoch0')
+if (-not (Test-Path $cuda)) { Write-Output "RESULT error cuda worker missing at $cuda (the fetch job runs first)"; exit 2 }
+if (-not (Test-Path $ocl)) { Write-Output "RESULT error opencl worker missing at $ocl"; exit 2 }
+if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 }
+foreach ($pk in $packList) { if (-not (Test-Path (Join-Path $packs $pk))) { Write-Output "RESULT error pack $pk missing"; exit 2 } }
+$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
+if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe (the NVRTC DLLs come from there)'; exit 2 }
+Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination $fetched -Force
+Write-Output "RESULT worker-cuda $cuda sha256 $((Get-FileHash -Algorithm SHA256 $cuda).Hash.ToLower()) with $((Get-ChildItem $fetched -Filter 'nvrtc*.dll').Count) NVRTC DLL(s) from $inst"
+Write-Output "RESULT worker-opencl $ocl sha256 $((Get-FileHash -Algorithm SHA256 $ocl).Hash.ToLower())"
+
+# the app
+$appDir = $env:IGNEUM_APP_DIR
+if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
+$urlFile = Join-Path $appDir 'app.url'
+$url = $null
+if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
+function Find-Card([string] $vendor, [string] $keyMatch) {
+ if (-not $url) { return $null }
+ try {
+ $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10
+ $c = $st.mining.cards | Where-Object { $_.vendor -eq $vendor -and $_.key -match $keyMatch } | Select-Object -First 1
+ if (-not $c) { $c = $st.cards | Where-Object { $_.vendor -eq $vendor -and $_.key -match $keyMatch } | Select-Object -First 1 }
+ return $c
+ } catch { Say ("api/state: " + $_.Exception.Message); return $null }
+}
+function Card-Off($card) {
+ Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state)
+ $body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
+ try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) }
+ $t = 0
+ while ($t -lt 90) {
+ Start-Sleep -Seconds 5; $t += 5
+ try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { }
+ }
+ Write-Output ("RESULT card-off " + $card.key + " after " + $t + " s")
+ Start-Sleep -Seconds 5
+}
+function Card-Restore($card) {
+ $body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
+ try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored " + $card.key + " enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore " + $card.key + ": " + $_.Exception.Message) }
+}
+
+# ---- the RTX 5090 (CUDA, NVRTC) ----
+$nv = Find-Card 'nvidia' '.'
+if ($nv) { Card-Off $nv } else { Write-Output 'RESULT card none-found nvidia (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' }
+& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" }
+foreach ($pk in $packList) {
+ $d = Join-Path $packs $pk
+ Write-Output "RESULT bench-5090 $pk start $(Get-Date -Format HH:mm:ss)"
+ & $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" }
+ & $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT $_" }
+}
+& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" }
+if ($nv) { Card-Restore $nv }
+
+# ---- the AMD card (OpenCL): the gfx1201 on the eGPU when present, else the integrated gfx1036 ----
+$list = & $ocl --list 2>&1
+$list | ForEach-Object { "RESULT list $_" }
+$dev = $null; $gfx = $null
+foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1201' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; $gfx = 'gfx1201'; break } }
+if ($null -eq $dev) { foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1036' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; $gfx = 'gfx1036'; break } } }
+if ($null -eq $dev) { Write-Output 'RESULT error no gfx1201 and no gfx1036 device in --list'; exit 2 }
+Write-Output "RESULT device $dev $gfx"
+$amd = Find-Card 'amd' $gfx
+if ($amd) { Card-Off $amd } else { Write-Output "RESULT card none-found $gfx (the app is not running or has no such card); measuring with whatever else runs on the GPU" }
+# the 9070 XT: the full shape; the gfx1036 (about 3 MH/s): the same 2^24 fingerprint range, 2 timed dispatches, one shape
+$batches = 5; if ($gfx -eq 'gfx1036') { $batches = 2 }
+foreach ($pk in $packList) {
+ $d = Join-Path $packs $pk
+ Write-Output "RESULT bench-$gfx $pk start $(Get-Date -Format HH:mm:ss)"
+ & $ocl --bench-pack --pack $d --batches $batches --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" }
+ if ($gfx -eq 'gfx1201') { & $ocl --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev --group-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT $_" } }
+}
+if ($amd) { Card-Restore $amd }
+exit 0