From 4605314315ae77334ffeb30e3526fe40b4d7b49e Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 08:08:26 +0000 Subject: [PATCH] Counter ASIC 3.0 item 8: the Mac rows The M5 Max ladder (Metal, packbench, IOReport GPU and DRAM watts without root): latency-bound to about 100,000 ops per hash, the 5 percent point about 130,000, 11 to 27 W GPU at 100,000 ops, 0.78 to 1.40 microjoules per hash; the verifier's law 2.06 ms + 3.2 us per 1,000 shadow instructions per warp on one core; every pack bit-exact. The analysis file with the knob, the chip side (k = 1, 1.5, 0.3), the gates and the consequences; the bench-log entry; the 5090 rows pending the PC 2 job (the playbook now carries the core-clock rows and the sh256x40 rung). Co-Authored-By: Claude Fable 5.1 --- docs/analysis/latency-shadow-2026-10-06.md | 152 +++++++++++++++++++++ docs/bench-log.md | 27 ++++ tools/ca3-shadow/pc2-shadow-bench.ps1 | 41 +++++- 3 files changed, 214 insertions(+), 6 deletions(-) create mode 100644 docs/analysis/latency-shadow-2026-10-06.md diff --git a/docs/analysis/latency-shadow-2026-10-06.md b/docs/analysis/latency-shadow-2026-10-06.md new file mode 100644 index 000000000..36c2ea5ea --- /dev/null +++ b/docs/analysis/latency-shadow-2026-10-06.md @@ -0,0 +1,152 @@ +# Program work in the latency shadow (Counter ASIC 3.0 item 8, 6 October 2026) + +Branch `ca3-shadow`, worker "shadow", on `ca3-coord` 870363c. The item was opened by item 1's finding (`docs/analysis/chip-model-v3.md` section 5.7): the `f = 1` chip stores the dataset in DRAM and recomputes nothing, so the only lever that moves its per-joule row is program work that the honest card hides behind its 128 dependent reads and a chip must pay for with an ALU core. This file holds the knob, the measurements on the Apple M5 Max and the RTX 5090, the verifier's cost law, the chip side re-evaluated on measured watts, the gates a class v4 candidate would need, and the consequences per tier. Every GPU figure says where it was measured and the load average it was taken at; every chip figure is arithmetic on the cited figures of chip-model-v3.md section 5 and is approximate. Nothing here is live: the knob sits behind a `LoadClass` field that version 2 and class v3 never set, no default changed, and nothing was published to the devnet. + +## 1. What the hash does today, counted from the code + +| Quantity | Value | Source | +|---|---|---| +| Instructions per program | 64, of which 16 are `load` | `igneum-pow/src/generator.rs` `INSTR_COUNT`, `LOAD_SLOTS`; spec 01 section 1.4 | +| Iterations per hash | 8 | `ITERATIONS` | +| Instructions executed per hash | 512: 128 loads, 384 ALU | 64 x 8 | +| ALU op mix of the 48 non-load slots | add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4 (weights, sum 75) | spec 01 section 1.4.2 | +| Integer ops per instruction, counted from the emitted statements | add 5 (shift, and, select, two adds), rotr 2 (and, funnel shift), shfl 2 (shuffle, xor), the seven others 1; load 2 (mask, xor) | `igneum-pow/src/emit.rs` instruction lines; weighted mean over the non-load weights 137 / 75 = 1.83 ops per ALU instruction | +| Integer ops per hash, this count | about 930: 384 x 1.83 = 703 ALU plus 128 x 2 = 256 for the loads' address mask and fold | arithmetic; the chip model's "512 ops per hash" (section 5.1) counted instructions, not ops | +| Per-hash loop | `for it in 0..8 { sel = r0; 64 instructions in order }`, then the fold | spec 01 section 1.7 | +| RTX 5090 integer budget | 45.2 T op/s (OpenCL event, integer chain) | `docs/benchmarks/repro.md` 2.2 (branch repro-bench) | +| Ops the 5090 could hide before compute binds at 136.1 MH/s | about 332,000 per hash | 45.2 T / 136.1 M | + +Vendor cost per op differs from this count. Item 6's step costs on the M5 Max (`docs/bench-log.md` "6 October 2026, Counter ASIC 3.0 item 6", branch ca3-reserve 192a683): the live `rotr` step 1.13x and the live `shfl` step 0.86x of the add-xor-rotate chain. The "ops" column below is the count above; the measured watts and rates are what the cards do with it. + +## 2. The knob + +`LoadClass::shadow: Option`, class name `+shx` (`mx8+sh256x27`), `igneum-pow/src/generator.rs`. The shadow block is `S` ALU instructions drawn from the program stream AFTER the 64 base instructions (the ten non-load families at the weights of 1.4.2, the same nine draws per instruction as the program's, the source drawn as on an ALU slot, the width roll and the era windows drawn and ignored when the class takes them), executed `R` times at the end of every iteration, after instruction 63 and before the next iteration samples `sel`, with the iteration's `sel`. Per hash it adds `8 x S x R` ALU instructions and no load. + +| Property | How the knob keeps it | +|---|---| +| 16 loads per program, 128 per hash, 4,096 items per unit | the base program is untouched; the block holds no load (`shadow_class_leaves_the_base_program_and_class_v3_untouched`, `generator.rs` tests) | +| The acceptance rule of 1.4.6 | unchanged: it interprets the base program (`accept.rs` reads `p.instrs`), and the block is drawn after the base draws, so the attempt and the verdict are those of the class without the shadow, draw for draw. A class v4 that adopts the block would decide whether rule (c) also runs the block (its cost scales with N: 2,048 evaluations x N ops, about 0.7 s at N = 330,000 on one core, approximate) | +| v2 and v3 byte-identical | the field is `None` on `V2`, `MX4`, `MX8`, `V3_CLASS`; the interpreter, the six kernel bodies, program.h and program.json emit nothing when it is `None`; `cargo test` in igneum-pow: 54 + 4 + 19 + 7 green, the pinned v2 and v3 packs byte for byte (`tests/packs.rs`) | +| Program id | the read-width id with `shadow/ || instrs_le16 || reps_le16` appended, so two block sizes of one seed never share an id | +| The CPU verifier | `verify.rs` runs the block `R` times per iteration after the base instructions with the same `step`; bit-exact against Metal on every pack below | +| Kernel text | one `for (sh = 0; sh < R; ++sh) { S lines }` block inside the iteration loop, the same statements the base instructions use, in Metal, CUDA and OpenCL, both kernels each; program.h carries `IGNEUM_SHADOW_INSTRS`, `IGNEUM_SHADOW_REPS`, `IGNEUM_SHADOW_INSTRS_PER_HASH`, `IGNEUM_SHADOW_OP_MIX` | + +Why a block with a repeat count and not a longer straight line: the block is one site in the kernel text, so a 256-instruction block at 88 passes is 256 lines of code, not 22,528, and the compile-ahead per epoch stays at the v3 figure (Metal compile 85 to 97 ms for a 256-instruction block, 231 ms for 1,024, against 1 ms cached for the control; the 1,024-instruction block also costs the M5 Max 17 percent of its hash rate at the same N, section 3). A repeated random block is RandomX's shape (256 instructions x 2,048 iterations per program), and a chip pays it with a general ALU core either way. + +## 3. The M5 Max rows (Metal, `with-lock.sh measure`, two sessions, 07:51 to 08:03 UTC) + +Harness: `proto-metal/packbench --pack --batches 60 --batch-log2 24 --group 256` (built from this branch; vectors are the Rust interpreter's, the fingerprint FNV-1a 64 over the 2^24 outputs at base 0). Power: the GPU and DRAM energy channels of the IOReport "Energy Model" group, sampled at 2 Hz by a 60-line Swift tool in the scratchpad (`gpupower`: `dlopen("/usr/lib/libIOReport.dylib")`, `IOReportCopyChannelsInGroup("Energy Model")`, the `GPU` and `DRAM` channels in mJ; the `GPU Energy` channel in nJ agrees with `GPU` to 2 percent), no root, the mean over the samples after the first 6 s of each run (compile, fills, warm-up) and before its last second. These are the GPU and the memory, not the package: Ember Tune's Apple row reports 38 W for the whole package at 26.7 MH/s (`docs/plans/ember-tune.md`, branch ember-tune, approximate), so the SoC fabric, the CPU and the rest add about 17 W over the 21 W here. The 5090's watts are whole-card. Idle GPU 0.44 W, idle DRAM 0.64 W. Load averages 3.3 to 6.8 (other agents' builds queued behind the measure lock; the GPU was idle: the Mac mines nothing). The control is the pinned class v3 pack `proto-cuda/packs-ca2-mixer/mx8-genesis`; the shadow packs are `proto-cuda/packs-ca3-shadow/*` over the same seed and class. + +| Pack (class over mx8) | Shadow instrs per hash | Ops per hash (x1.83 + 930) | MH/s (GPU time) | Against the control | GPU W | DRAM W | Microjoules per hash (GPU + DRAM) | Verifier ms per warp, one core, avg of 20 (worst cold) | Bit-exact (3 vectors, 3 in batch, fingerprint) | Load average at start | Compile ms | +|---|---|---|---|---|---|---|---|---|---|---|---| +| mx8-genesis (control, runs 1, 2, 3) | 0 | 930 | 27.07, 27.07, 27.10 | | 11.2, 9.2, 12.3 | 10.2, 10.0, 10.2 | 0.78 (mean 21.0 W) | 2.062 (2.188) | yes, 7c28cfb06c5c65a9 (the 5 October fingerprint) | 4.6, 6.8, 4.2 | 1 (cached) | +| sh256x2 | 4,096 | 8,400 | 26.85 | -0.8% | 15.1 | 10.4 | 0.95 | 2.080 (2.249) | yes, 33e8bbe4c35b2e54 | 4.6 | 85 | +| sh256x7 | 14,336 | 27,200 | 26.74 | -1.3% | 18.0 | 10.6 | 1.07 | 2.112 (2.224) | yes, 6cfb70911007520a | 4.2 | 93 | +| sh256x13 | 26,624 | 49,700 | 26.75 | -1.2% | 20.6 | 10.5 | 1.16 | 2.212 (2.237) | yes, 59ac286fe2a5a9ef | 3.7 | 93 | +| sh64x52 (the same N as sh256x13, a 64-instruction block; runs 1, 2) | 26,624 | 49,700 | 27.72, 27.79 | +2.5% | 19.8, 20.3 | 10.2, 10.3 | 1.09 | 2.137 (2.237) | yes, 9dd010f79d8ca9f4 | 4.4, 4.1 | 59 | +| sh1024x3 (a 1,024-instruction block) | 24,576 | 45,900 | 22.53 | -16.8% | 20.2 | 9.7 | 1.33 | 2.150 (2.274) | yes, a05399c819b79aad | 6.0 | 231 | +| sh256x27 (runs 1, 2) | 55,296 | 102,100 | 26.48, 26.86 | -1.5% (mean 26.67) | 26.9, 26.9 | 10.4, 10.2 | 1.40 | 2.229 (2.276) | yes, 3d2e8245cc084d07 | 3.3, 5.6 | 96 | +| sh256x40 | 81,920 | 150,800 | 25.10 | -7.3% | 26.7 | 10.1 | 1.46 | 2.329 (2.396) | yes, 0844b706302f1c9c | 5.0 | 97 | +| sh256x53 (runs 1, 2) | 108,544 | 199,600 | 24.40, 24.09 | -10.4% (mean 24.25) | 28.8, 28.5 | 9.5, 9.7 | 1.58 | 2.427 (2.461) | yes, 4f824b15cf2b124a | 3.9, 4.6 | 85 | +| sh256x88 | 180,224 | 330,700 | 21.39 | -21.0% | 31.7 | 8.4 | 1.88 | 2.619 (2.771) | yes, 0572522e39a94d8a | 3.8 | 87 | + +Commands: `with-lock.sh measure /mac-shadow-session.sh` and `mac-shadow-session2.sh` (each: `uptime`, `gpupower --interval-ms 500` in the background, `packbench` as above, then `igneum-pow bench --seed igneum-genesis --class --warps 20` for the verifier; logs `mac-shadow-session.log`, `mac-shadow-session2.log` in the scratchpad). The control reads 2.5 percent under the 5 October figure for this pack (27.7 MH/s, `docs/plans/mixer-x4.md` 6.2) on three runs over 12 minutes, so the day's denominator is 27.08 and every row is read against it. + +What the rows say: + +1. The M5 Max stays latency-bound to about 100,000 ops per hash (sh256x27: 1.5 percent under the control, inside the run-to-run spread of 1.4 percent) and stops there: 7.3 percent down at 151,000, 10.4 at 200,000, 21 at 331,000. By the 5 percent rule (`docs/plans/counter-asic-2-public.md`, the 2.0 rule) its ceiling is about 130,000 ops per hash (interpolated between the 102,000 and 151,000 rungs), 70,000 shadow instructions, `sh256x35`. Its counted throughput at the compute-bound end is 21.39 M x 330,700 = 7.1 T ops/s, against the 4.4 T op/s the add-xor-rotate chain of item 6 gives (881 G steps/s x 5): the shadow's op mix issues faster than a dependent chain, as it should. The chip model's "about 290,000" for the M5 Max (section 5.7, from memory) was 2.2x too high; the measured bind point is the row. +2. The block size matters on Apple at the same N: a 64-instruction block (sh64x52) runs 2.5 percent ABOVE the control on both runs, 256 holds, 1,024 costs 17 percent (the kernel's instruction footprint; the 1,024-line block at 3 passes is 3,072 instructions of straight-line code per iteration if the compiler unrolls it). A class v4 would fix the block at 64 to 256 instructions. Why a block can be faster than nothing: the shadow spaces the loads of a warp in time, so fewer reads queue at once (the probe's 522 ns at 256 lanes is a queued latency); it is a 2.5 percent effect on this card and is not claimed for the others. +3. Watts: the GPU rises from 11 W to 27 W at 100,000 ops and to 32 W at 331,000, the DRAM stays at 10 W while the rate holds and falls with it. The marginal energy per counted op on the latency-bound rungs: 22 pJ at 8,400 ops (the first rung wakes the ALUs), 11 at 27,000, 7.8 at 50,000, 6.9 at 100,000; at the compute-bound end 2.9 pJ (31.7 W over 7.1 T op/s). The M5 Max's marginal ALU energy at the useful N is therefore 6 to 8 pJ per counted op, the same class as the 5.5 pJ the model assumed for the 5090 (section 5.7, from (575 - 326) W / 45.2 T op/s). +4. Energy per hash, GPU plus DRAM: 0.78 microjoules at the control, 1.40 at 100,000 ops, 1.88 at 331,000. The M5 Max at the hash is 3.0x better per joule than the 5090's 2.34 microjoules (290 W at 124 MH/s in the app, the 5 October sweep record: `docs/plans/miner-eff.md`, branch miner-eff; 326 W was a peak with the prover on) on these channels, 1.7x on the package figure. + +## 4. The verifier's cost law + +One M5 Max performance core, the Rust interpreter, 20 warps, loads 3.3 to 6.8: ms per warp = 2.06 + 3.2 x 10^-3 x (shadow instructions per hash) / 1,000, the slope fitted on the three largest rungs (3.09, 3.36, 3.26 microseconds per 1,000 shadow instructions per warp). That is 0.1 ns per lane-instruction (the register-major loops vectorise over the 32 lanes) or 1.75 microseconds per 1,000 counted ops per warp. The worst cold unit sits 0.05 to 0.17 ms above the average on every rung. On a 2019-class laptop core the rule of chip-model-v3.md section 3 item 1 (2.5x slower) gives 8.0 microseconds per 1,000 shadow instructions per warp. + +The headroom under the 10 ms gate that the shadow may spend, per pairing (item 2's figures, `docs/plans/counter-asic-3-derivation.md` section 0, bdc07d3, steady / worst cold; the derivation and the shadow are paid by the same verifier on the same unit): + +| Pairing | Headroom, M5 Max core (steady / worst cold) | Shadow instrs per hash that fit on the M5 Max core | Ops per hash | Headroom, 2019-class core (approximate) | Shadow instrs that fit on the 2019-class core | Ops per hash | +|---|---|---|---|---|---|---| +| x8 (class v3) + shadow | 7.9 / 7.8 ms | 2.4 million | 4.5 million | 4.8 / 4.6 ms | 575,000 | 1.05 million | +| dr368 + shadow | 7.3 / 7.1 | 2.2 million | 4.1 million | 3.3 / 2.8 | 350,000 | 640,000 | +| dr736 + shadow | 5.1 / 4.8 | 1.5 million | 2.7 million | none (dr736 alone is 12 ms on that core) | 0 | 0 | + +So the verifier does not bind the shadow at any N on the ladder in any pairing where the derivation itself fits the gate: 330,700 ops per hash costs 0.56 ms on the M5 Max core and about 1.4 ms on the 2019-class core, which is 12 to 30 percent of the headroom. The binding constraint on N is the honest cards' compute (section 3 for the M5 Max, section 5 for the 5090), not the node. The measured rows: x8 + sh256x88 2.62 ms steady, 2.77 worst cold on the M5 Max core, about 6.5 and 6.9 ms on a 2019-class core by the 2.5x rule; dr736 + sh256x88 about 5.5 ms on the M5 Max core (4.9 + 0.56). The 2019-class core itself is unmeasured (O-1.14); the 2.5x rule stands in for it. + +## 5. The RTX 5090 rows (PC 2, 1ccfe586) + +PENDING the PC 2 job `run-ca3-shadow-pc2-20261006` (`tools/ca3-shadow/pc2-shadow-bench.ps1`, published after `/tmp/igneum-devnet/pc2-ca3.clear` under the mkdir lock; the fetch job `fetch-ca3-shadow-20261006` carries the ten packs). The job: `--stop-miners`, the prover off for the run and back on at the end, the card confirmed empty by nvidia-smi's compute-apps list, the installed `igneum-worker-cuda.exe` (NVRTC, the packs' own text) with `--bench --batches 250 --batch-log2 24 --block-warps 1` on every pack, the control first and last, `--block-warps 8` on the control and two rungs, `nvidia-smi -l 1` sampling power.draw, utilization.gpu, clocks.sm, clocks.mem and temperature every second with the per-pack mean over the bench's own window; then the core-clock rows. The rows land here with the closing report. + +Power caps. The brief asked for `nvidia-smi -pl 200, 250` and the default. The 5 October sweep on this card (`docs/plans/counter-asic-2-status.md` "22:16 the 5090 power-limit sweep", relay #224, 22:09 to 22:15Z on PC 2, corrected in `counter-asic-2-rollout.md` 7b) already carries the cap rows at the current program length: + +| Cap | Limit W | Draw W | MH/s (the app's API) | MH/W | SM MHz | +|---|---|---|---|---|---| +| 100% | 575 | 316.2 | 115.42 | 0.365 | 3,051 | +| 80% | 460 | 316.1 | 115.60 | 0.366 | 3,050 | +| 65% | 400 (the floor) | 310.6 | 114.46 | 0.369 | 3,051 | +| 50% | 400 (clamped) | 302.4 | 109.20 | 0.361 | 3,050 | + +Reading: the card draws 302 to 316 W under this program whatever the cap, `power.min_limit` is 400 W, so `-pl 200` and `-pl 250` cannot be set on a 5090 and the cap never binds at the hash. Those rows are re-used, not re-run. The lever that can move the watts is the core clock, and the job carries `nvidia-smi -lgc 0,` at the Ember plan's four caps (2,781 / 2,472 / 2,163 / 1,854 MHz, `docs/plans/ember-tune.md`), about 75 s each, at the control and at sh256x88, with `-rgc` in a finally block; if nvidia-smi refuses the lock the rows are OWED and the job does not try to gain rights. Ember Tune's own 5090 steps did not run (its second engine never mined; re-run pending the project lead); its idle readbacks on PC 1: 90.6 W idle, 2,505 MHz core, 14,001 MHz memory, limit 450 of 575 W. + +The item 1 denominator moves: the 5090 at the hash alone is about 290 W (p95 290, max 295 W at 124 MH/s in the app, the 5 October record on branch miner-eff), not 326 W, so 2.34 microjoules per hash, and the `f = 1` rows of section 5.4 read 5.0x on GDDR7 (was 5.1x), 7.3x on one HBM3 stack (was 7.5x), 8.9x on eight (was 9.2x): a 2 to 3 percent move. The job's own watts at the hash replace this when they land. + +## 6. The chip side, re-evaluated (approximate) + +The `f = 1` chip of chip-model-v3.md section 5.4, at the memory's activate ceiling: GDDR7 (16 devices) 166.4 MH/s at 0.466 microjoules per hash, one HBM3 stack 83.6 MH/s at 0.321, eight stacks 0.262; those are memory, static and controller only. With the shadow the chip adds a core that runs the per-epoch random block at `k` times the GPU's marginal energy per op; the model's 5.5 pJ (section 5.7) is kept as the unit, and the measured marginal figures of section 3 (6 to 8 pJ on the M5 Max at the useful N) say it is the right order. Three columns: `k = 1` (a chip core as good as the GPU's ALU on a random program: RandomX's argument and the project lead's test), `k = 1.5` (1.5x worse), and `k = 0.3` (a wide-SIMD fixed-datapath array at N5: the int32 datapath energy of section 5.1, 0.06 pJ per add and 0.52 per multiply at the op mix's 29 percent multiplies is 0.19 pJ, with a 2x pipeline overhead and about 4x for the register file, operand wires and the shared instruction fetch over a wide SIMD row, approximate; the floor a chip maker would claim). Section 5.7's "k = 1.5" column was computed as the chip being 1.5x BETTER (0.83 at N = 100,000 is 0.466 + 0.55 / 1.5); the column below computes it as written, 1.5x worse. Energy per hash = memory + N x 5.5 pJ x k. + +ALU silicon for the core, at the 5090's 45.2 T op/s (what N = 330,700 at 136.1 MH/s needs): 22,600 32-bit lanes at 2.0 GHz. At N5 an int32 multiply-add lane with its register-file slice is about 0.002 mm^2 (approximate: 3,000 to 6,000 gates at about 0.3 square microns per NAND2-equivalent plus registers), so about 45 mm^2 of datapath, 60 to 100 mm^2 with SIMD control and operand networks, $25 to $40 of silicon at the sram-mirror.md yield model ($20,000 wafer, about $0.36 per mm^2 on a 128 mm^2 die; section 5 of that file); at 28 nm the same lanes are about 8x the area, 360 mm^2 of datapath and 500 mm^2 or more with control, a reticle-class die that a $5M to $30M controller project does not carry. Watts for the core: at `k = 1` 5.5 pJ x 45.2 T = 249 W (the GPU's own marginal power), at `k = 0.3` 75 W, at `k = 1.5` 373 W. So at N = 330,700 the chip is a memory controller plus a GPU-class ALU array at 250 W on an advanced node, which is the brief's test in numbers; at N = 100,000 the array is a third of that (14,000 lanes, about 30 mm^2 at N5, 75 W at `k = 1`). + +Gain per joule against the honest cards, chip energy = memory + N x 5.5 pJ x k, N in counted ops: + +| N ops per hash | Card microjoules: M5 Max (GPU + DRAM) / 5090 | Chip, GDDR7, k = 1 / 1.5 / 0.3 | Gain per joule, GDDR7, against the M5 Max, k = 1 / 1.5 / 0.3 | Against the 5090 (section 5.7's linear watts until the job's rows land) | Chip, one HBM3 stack, k = 1 / 1.5 / 0.3 | Gain, HBM3, against the M5 Max | Against the 5090 (model) | +|---|---|---|---|---|---|---|---| +| 930 (today) | 0.78 / 2.34 (measured, 290 W) | 0.47 | 1.7x | 5.0x | 0.33 | 2.4x | 7.3x | +| 49,700 (sh256x13) | 1.16 / about 2.45 (333 W, model) | 0.74 / 0.88 / 0.55 | 1.6x / 1.3x / 2.1x | 3.3x / 2.8x / 4.5x | 0.59 / 0.73 / 0.40 | 2.0x / 1.6x / 2.9x | 4.2x / 3.4x / 6.1x | +| 102,100 (sh256x27) | 1.40 / about 2.78 (378 W, model) | 1.03 / 1.31 / 0.63 | 1.4x / 1.1x / 2.2x | 2.7x / 2.1x / 4.4x | 0.88 / 1.16 / 0.49 | 1.6x / 1.2x / 2.9x | 3.2x / 2.4x / 5.7x | +| 199,600 (sh256x53; the M5 Max is 10 percent down here) | 1.58 / about 3.40 (462 W, model) | 1.56 / 2.11 / 0.80 | 1.0x / 0.75x / 2.0x | 2.2x / 1.6x / 4.3x | 1.42 / 1.97 / 0.65 | 1.1x / 0.80x / 2.4x | 2.4x / 1.7x / 5.2x | +| 330,700 (sh256x88; the M5 Max is 21 percent down) | 1.88 / about 4.22 (575 W, model) | 2.29 / 3.19 / 1.01 | 0.82x / 0.59x / 1.9x | 1.8x / 1.3x / 4.2x | 2.14 / 3.05 / 0.87 | 0.88x / 0.61x / 2.2x | 2.0x / 1.4x / 4.9x | + +Reading: the lever works exactly as far as `k` is near 1. At N = 100,000 the chip's edge over the M5 Max is 1.4x at `k = 1` and 1.1x at `k = 1.5`, and over the 5090 (model watts) 2.7x and 2.1x; at `k = 0.3` the edge barely moves at any N (2.2x and 4.4x), because a fixed-datapath array runs the block for a third of the card's marginal energy and the memory row stays the same. The number that decides the verdict is therefore `k`: the energy per op of a chip core that executes a per-epoch random integer program with a 32-lane shuffle, divided by a GPU's marginal 5.5 to 7 pJ. Measured on the cards: the M5 Max pays 6.9 pJ per counted op at 100,000 and 2.9 at its compute-bound end; the 5090's figure is the job's. RandomX's argument (and the precedent: no RandomX chip has beaten a CPU per joule) is that a general core cannot reach `k` under about 0.5 on a random program; nothing in this file measures a chip, so the `k = 0.3` column is the attacker's claim and the `k = 1` column ours. + +## 7. The gates a class v4 candidate would need + +The six gates of `docs/plans/counter-asic-2-rollout.md` section 7 plus the verifier on a 2019-class core (O-1.14), with the state today: + +| Gate | State on the shadow class | What closes it | +|---|---|---| +| G1 bit-exact on all three vendors against the Mac reference | Metal GREEN on ten packs (3 of 3 vectors standalone and in batch, cache and dataset PASS, fingerprints above); CUDA PENDING the PC 2 job (the worker's self-test runs the vectors through the bound kernel and prints the 2^24 fingerprint); AMD OWED (PC 1) | the PC 2 closing report; a PC 1 job on the 9070 XT (the OpenCL kernels are emitted) | +| G2 the CPU verifier exact on 1,000 random hashes per card | not run: the evidence is the 96 vectors and the 2^24 fingerprint per pack on Metal | a serve-mode job per card as in 2.0 | +| G3 the generator soundness suite green | `cargo test` 54 + 4 + 19 + 7 green with the new test; the Metal fuzz, edge, stats and determinism runs were NOT run on the shadow class | those four runs at the chosen S and R | +| G4 the fast-time 3-node network across an activation | not run (no node change exists: the class is a `LoadClass` field, and the chain's `V3_CLASS` does not set it) | the node seam for v4 (`program_class_v4_activation_daa`) and the run | +| G5 the PC-built Windows workers and the Mac workers from the same commit | not applicable yet (no worker change: the workers run the packs' own text) | the release step | +| G6 the node change on a fork branch with suites green on PC 2 | not applicable yet | the v4 seam | +| O-1.14 the verifier on a 2019-class core | by the 2.5x rule every rung fits with x8 (6.9 ms worst cold at 330,700 ops); the core itself is unmeasured | one run on such a core | +| Spec text | none written (no PROPOSED entry until the rows are complete); the block would be a class v4 field beside the class v3 ones in 1.4 and 1.7, the acceptance rule's treatment of the block decided (section 2) | the v4 proposal after the 5090 and AMD rows | + +## 8. Consequences per tier + +A longer program costs the card watts and nothing else while the card stays latency-bound; past its bind point it costs hash rate. Income per pound of card is unchanged while the rate holds; income per watt falls by the watts ratio. The numbers are this file's; the 5090's are the model's until the job lands, and the AMD row is the budget's. + +| Tier | At N = 100,000 ops per hash (sh256x27, the candidate) | At N = 200,000 | At N = 331,000 (the 5090's full budget) | What is being done | +|---|---|---|---|---| +| Apple user (M-series laptop or desktop; M5 Max measured) | rate -1.5%, GPU + DRAM 21 to 37 W (+16 W; a laptop on battery feels it), income per watt 0.56x, per pound 0.99x | rate -10%, 38 W; per watt 0.49x, per pound 0.90x | rate -21%, 40 W; per watt 0.41x, per pound 0.79x | the project's N is capped by this card: no more than 130,000 ops (the 5 percent rule), 100,000 recommended | +| NVIDIA, one 24 or 32 GB card (5090 class) | model: rate holds, about 290 to 378 W (+30%); per watt 0.77x, per pound 1.0x | about 462 W; per watt 0.63x | compute-bound at 136 MH/s, 575 W; per watt 0.50x | the PC 2 job measures the rate, the watts and the clock caps; the per-watt loss is the price of the lever and is reported, not hidden | +| NVIDIA, one 8, 12 or 16 GB card | the same shape per card: the ALU budget scales with the SM count, so a 4070-class card (about 30 T op/s, approximate) binds near 200,000 ops at its rate; at 100,000 it holds | a 4060-class card (about 15 T op/s) binds near 100,000 | | no card we own is in this tier; the ladder runs on any CUDA card through the same job | +| AMD, one 16 GB card (RX 9070 XT) | OWED: the budget (about 650,000 ops, approximate, section 5.7) says it holds at 100,000 and at 200,000; its watts at the hash are unmeasured | | | the PC 1 job when the desk is free; the OpenCL kernels are in every pack | +| A rig | per card as above; a rig's bill is watts, so at 100,000 ops a 5090 rig pays about 30 percent more electricity for the same hash (model) | | | the recommended N is the smallest that moves the chip row; the rig pays it only if class v4 is adopted | +| A pool user | nothing changes in shares or payout: the hash rate holds at the recommended N on every card measured | | | | +| A chip | must add a 14,000-lane ALU array at N = 100,000 (about 30 mm^2 at N5, 75 W at `k = 1`) and a 22,600-lane array at 331,000 (60 to 100 mm^2, 250 W); its per-joule edge over the 5090 falls from 5.0x to 2.7x at 100,000 and `k = 1` (model watts), over the M5 Max from 1.7x to 1.4x | | | the `k` question goes to the external cryptanalysis and chip review (item 3): can a core run a random 32-lane program under half a GPU's marginal energy per op | +| The public claim | the honest M5 Max is 1.7x from the `f = 1` GDDR7 chip per joule at N = 0 by these channels, the 5090 5.0x; at N = 100,000 and `k = 1` they read 1.4x and 2.7x (model) | | | nothing on the site changes from this file | + +The N the project should pick, by the 2.0 rule (no card we own loses more than 5 percent): 100,000 ops per hash, `S = 256`, `R = 27` (55,296 shadow instructions per hash), subject to the 5090 row holding within 5 percent in the job and the 9070 XT row when PC 1 is free; the M5 Max's ceiling is 130,000. The block size is 64 to 256 instructions, never 1,024. + +## 9. Unverified and owed + +- The RTX 5090 rows (hash rate, watts, clock caps, the CUDA fingerprints): the PC 2 job, section 5. +- The RX 9070 XT rows: OWED (PC 1 not released today). +- The Mac watts are the IOReport GPU and DRAM channels (no root, the private IOReport library through dlopen), not powermetrics and not the package; Ember Tune's 38 W package figure is approximate and from another session. +- The ops count (1.83 per ALU instruction) is a convention counted from the emitted statements; the vendors' real cost per op differs (item 6's step ratios), which is why the rows carry instructions and watts, not ops alone. +- The chip side is arithmetic: the 5.5 pJ unit, the `k` columns, the lane area and the 28 nm scaling are approximate; no chip was measured, and the `k = 0.3` floor is an estimate of what a fixed-datapath array could claim. +- The acceptance rule does not run the block; whether a class v4 should make it is a design decision with the cost stated in section 2. +- The 2019-class core is unmeasured (O-1.14); the 2.5x rule stands in. +- The Metal fuzz, edge, stats and determinism runs were not made on the shadow class (gate G3 is the crate suite only). diff --git a/docs/bench-log.md b/docs/bench-log.md index 85997dca6..27dddbb38 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1961,3 +1961,30 @@ Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/ep Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU. Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan. + +## 6 October 2026, Counter ASIC 3.0 item 8: program work in the latency shadow + +Branch `ca3-shadow`, worker "shadow" (`docs/analysis/latency-shadow-2026-10-06.md` carries the design, the chip side and the consequences; this entry carries the measurements). The knob: `LoadClass::shadow`, class name `+shx`, a block of `S` ALU instructions drawn from the program stream after the 64 base instructions and run `R` times at the end of every iteration (no load; the base program, its attempt and the acceptance verdict are the class's without the shadow; v2 and v3 byte-identical, `cargo test` in igneum-pow 54 + 4 + 19 + 7 green). Packs `proto-cuda/packs-ca3-shadow/*` over `mx8` for seed igneum-genesis; the control is the pinned class v3 pack `packs-ca2-mixer/mx8-genesis`. Ops per hash = shadow instructions x 1.83 (counted from the emitted statements: add 5, rotr 2, shfl 2, the rest 1, weighted over the non-load weights) + 930 (the base program's 384 ALU instructions and 128 loads). + +**Apple M5 Max, Metal** (`proto-metal/packbench --pack --batches 60 --batch-log2 24 --group 256` under `with-lock.sh measure`, two sessions 07:51 to 08:03 UTC, GPU time; power = the IOReport "Energy Model" `GPU` and `DRAM` channels at 2 Hz through a dlopen of libIOReport, no root, the mean over each run after its first 6 s; idle GPU 0.44 W, DRAM 0.64 W; the package is about 17 W more, Ember Tune's 38 W row, approximate; load averages 3.3 to 6.8, the GPU idle: the Mac mines nothing): + +| Pack | Shadow instrs per hash | Ops per hash | MH/s | Against the control | GPU W | DRAM W | Microjoules per hash (GPU + DRAM) | Verifier ms per warp, one core, avg of 20 (worst cold) | Bit-exact, fingerprint 2^24 | Load at start | +|---|---|---|---|---|---|---|---|---|---|---| +| mx8-genesis (control; runs 1, 2, 3) | 0 | 930 | 27.07, 27.07, 27.10 | | 11.2, 9.2, 12.3 | 10.2, 10.0, 10.2 | 0.78 | 2.062 (2.188) | yes, 7c28cfb06c5c65a9 | 4.6, 6.8, 4.2 | +| sh256x2 | 4,096 | 8,400 | 26.85 | -0.8% | 15.1 | 10.4 | 0.95 | 2.080 (2.249) | yes, 33e8bbe4c35b2e54 | 4.6 | +| sh256x7 | 14,336 | 27,200 | 26.74 | -1.3% | 18.0 | 10.6 | 1.07 | 2.112 (2.224) | yes, 6cfb70911007520a | 4.2 | +| sh256x13 | 26,624 | 49,700 | 26.75 | -1.2% | 20.6 | 10.5 | 1.16 | 2.212 (2.237) | yes, 59ac286fe2a5a9ef | 3.7 | +| sh64x52 (64-instruction block; runs 1, 2) | 26,624 | 49,700 | 27.72, 27.79 | +2.5% | 19.8, 20.3 | 10.2, 10.3 | 1.09 | 2.137 (2.237) | yes, 9dd010f79d8ca9f4 | 4.4, 4.1 | +| sh1024x3 (1,024-instruction block) | 24,576 | 45,900 | 22.53 | -16.8% | 20.2 | 9.7 | 1.33 | 2.150 (2.274) | yes, a05399c819b79aad | 6.0 | +| sh256x27 (runs 1, 2) | 55,296 | 102,100 | 26.48, 26.86 | -1.5% | 26.9, 26.9 | 10.4, 10.2 | 1.40 | 2.229 (2.276) | yes, 3d2e8245cc084d07 | 3.3, 5.6 | +| sh256x40 | 81,920 | 150,800 | 25.10 | -7.3% | 26.7 | 10.1 | 1.46 | 2.329 (2.396) | yes, 0844b706302f1c9c | 5.0 | +| sh256x53 (runs 1, 2) | 108,544 | 199,600 | 24.40, 24.09 | -10.4% | 28.8, 28.5 | 9.5, 9.7 | 1.58 | 2.427 (2.461) | yes, 4f824b15cf2b124a | 3.9, 4.6 | +| sh256x88 | 180,224 | 330,700 | 21.39 | -21.0% | 31.7 | 8.4 | 1.88 | 2.619 (2.771) | yes, 0572522e39a94d8a | 3.8 | + +Reading: the M5 Max stays latency-bound to about 100,000 ops per hash and its 5 percent point is about 130,000 (between the 102,100 and 150,800 rungs), 2.2x under the chip model's 290,000 (from memory); the block size matters on Apple (64 instructions +2.5 percent, 256 holds, 1,024 costs 17 percent at the same N: the instruction footprint); the GPU rises from 11 to 27 W at 100,000 ops (marginal 6.9 pJ per counted op; 2.9 pJ at the compute-bound end) and the energy per hash from 0.78 to 1.40 microjoules, GPU plus DRAM. The verifier's law on this core: 2.06 ms + 3.2 microseconds per 1,000 shadow instructions per warp (0.1 ns per lane-instruction), so 330,700 ops cost 0.56 ms here and about 1.4 ms on a 2019-class core by the 2.5x rule: inside every pairing's headroom (x8 7.8 / 4.6 ms, dr368 7.1 / 2.8, dr736 5.1 / none, M5 Max / 2019-class, item 2's figures), so the node never binds the shadow before the cards do. The control is 2.5 percent under the 5 October figure for this pack (27.7, mixer-x4.md 6.2) on three runs; every row is read against today's 27.08. + +**RTX 5090 (PC 2, 1ccfe586), CUDA**: PENDING the PC 2 job `run-ca3-shadow-pc2-20261006` (`tools/ca3-shadow/pc2-shadow-bench.ps1`, after `/tmp/igneum-devnet/pc2-ca3.clear` under the mkdir lock): the same ladder through the installed worker's `--bench` with `nvidia-smi -l 1` sampling, and the core-clock rows at 2,781 / 2,472 / 2,163 / 1,854 MHz (`-lgc`, `-rgc` in a finally; OWED if the lock needs rights the job does not have). The power-cap rows are the 5 October sweep's (counter-asic-2-status.md 22:16: 302 to 316 W at every cap, floor 400 W, so `-pl 200` and `250` cannot be set and the cap never binds at the hash); the card at the hash alone is about 290 W (miner-eff record), 2.34 microjoules per hash at 124 MH/s in the app. + +**RX 9070 XT (PC 1, ae432dc7)**: OWED (PC 1 is the project lead's desk and not released today); the OpenCL kernels are in every pack. + +Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (`sh256x27`) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 by the model holds its rate at about 378 W (per watt 0.77x; the job measures it), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the `f = 1` chip's edge per joule falls from 1.7x to 1.4x against the M5 Max and from 5.0x to 2.7x against the 5090 (model watts) at `k = 1`, where `k` is the chip core's energy per op over the GPU's marginal 5.5 pJ: the number that decides the item. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000. diff --git a/tools/ca3-shadow/pc2-shadow-bench.ps1 b/tools/ca3-shadow/pc2-shadow-bench.ps1 index fcc2f27c7..6fca93bfc 100644 --- a/tools/ca3-shadow/pc2-shadow-bench.ps1 +++ b/tools/ca3-shadow/pc2-shadow-bench.ps1 @@ -5,12 +5,13 @@ # tools/prover-floor/pc2-floor-measure.ps1 does), confirms the card is empty by nvidia-smi's compute-apps list (never # api/state, which answers {} on this machine), then runs the INSTALLED igneum-worker-cuda.exe (NVRTC, the pack's own # kernel text) with --bench on every pack of the fetched kit (fetch job fetch-ca3-shadow-20261006: the class v3 control -# mx8-genesis and the shadow ladder sh256x2 .. sh256x88, sh64x52, sh1024x3) while `nvidia-smi -l 1` samples +# mx8-genesis, run first and last, and the shadow ladder sh256x2 .. sh256x88, sh64x52, sh1024x3) while `nvidia-smi -l 1` samples # power.draw, utilization.gpu, clocks.sm, clocks.mem and temperature every second into a file. Every number is a RESULT # line: the worker's own lines (vectors, fingerprint, MH/s), and per pack the mean, min and max power, the mean # utilisation and clocks over the bench's own window (the timed dispatches), plus every raw sample. No power cap is set # (the 5 October sweep on this card, counter-asic-2-status.md 22:16, showed the cap never binds above the 400 W floor and -# -pl 200 / 250 are below power.min_limit); nothing elevated. Read back with `node tools/jobs.mjs `. +# -pl 200 / 250 are below power.min_limit); the core clock is locked for the clock rows only if nvidia-smi allows it +# unprompted (else OWED); nothing elevated. Read back with `node tools/jobs.mjs `. $ErrorActionPreference = 'Continue' function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') } $urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' } @@ -45,7 +46,7 @@ if (-not $smiExe) { $smiExe = Join-Path $env:SystemRoot 'System32\nvidia-smi.exe $sampler = Start-Process -FilePath $smiExe -ArgumentList @('--query-gpu=timestamp,power.draw,utilization.gpu,clocks.sm,clocks.mem,temperature.gpu,memory.used', '--format=csv,noheader,nounits', '-l', '1') -RedirectStandardOutput $samples -NoNewWindow -PassThru Start-Sleep -Seconds 5 $windows = @() -$order = @('mx8-genesis', 'sh256x2', 'sh256x7', 'sh256x13', 'sh256x27', 'sh256x53', 'sh256x88', 'sh64x52', 'sh1024x3') +$order = @('mx8-genesis', 'sh256x2', 'sh256x7', 'sh256x13', 'sh256x27', 'sh256x40', 'sh256x53', 'sh256x88', 'sh64x52', 'sh1024x3', 'mx8-genesis') try { "RESULT idle $(Stamp) 10 s before the first bench: $(Smi 'power.draw,utilization.gpu,clocks.sm,clocks.mem')" Start-Sleep -Seconds 10 @@ -60,11 +61,39 @@ try { & $exe --bench --pack $d --batches $batches --batch-log2 24 --block-warps $bw 2>&1 | ForEach-Object { "RESULT bench $pk bw$bw $_" } $t1 = Get-Date "RESULT bench $pk bw$bw end $(Stamp) exit=$LASTEXITCODE wall_s=$([int]($t1 - $t0).TotalSeconds)" - $windows += [pscustomobject]@{ pack = $pk; bw = $bw; t0 = $t0; t1 = $t1 } + $windows += [pscustomobject]@{ pack = $pk; bw = "bw$bw"; t0 = $t0; t1 = $t1 } + if ($pk -eq 'mx8-genesis' -and $bw -eq 1 -and $windows.Count -gt 2) { break } + Start-Sleep -Seconds 3 + } + } + # the core-clock rows (coordinator, 6 October 2026 08:0x UTC): the 5 October sweep showed `-pl` cannot bind on this kernel + # (the card draws about 290 to 316 W under a 400 W floor), so the lever that can move the watts is the core clock. Four + # caps of the Ember plan, at the control and at the largest N that fits the verifier gate, about 75 s each. If + # nvidia-smi refuses the lock (it needs administrator rights this job does not have and must not ask for), the rows + # are OWED and nothing else is tried; -rgc runs in the finally block whatever happens. + $clockOk = $true + foreach ($pk in @('mx8-genesis', 'sh256x88')) { + if (-not $clockOk) { break } + $d = Join-Path $packs $pk + if (-not (Test-Path $d)) { continue } + foreach ($mhz in @(2781, 2472, 2163, 1854)) { + $r = (& nvidia-smi -i 0 -lgc 0,$mhz 2>&1 | Out-String).Trim() + "RESULT clock $pk $mhz lgc: $($r -replace "`r?`n", ' | ')" + if ($r -match 'Insufficient|[Pp]ermission|denied|administrator|not supported|Unknown Error|No such') { "RESULT clock OWED: nvidia-smi -lgc refused ($mhz MHz); the rows need rights this job does not have and does not ask for"; $clockOk = $false; break } + Start-Sleep -Seconds 3 + "RESULT clock $pk $mhz readback: $(Smi 'clocks.sm,clocks.max.sm,clocks.mem,power.draw,power.limit')" + $t0 = Get-Date + "RESULT bench $pk clk$mhz start $(Stamp)" + & $exe --bench --pack $d --batches 600 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT bench $pk clk$mhz $_" } + $t1 = Get-Date + "RESULT bench $pk clk$mhz end $(Stamp) exit=$LASTEXITCODE wall_s=$([int]($t1 - $t0).TotalSeconds)" + $windows += [pscustomobject]@{ pack = $pk; bw = "clk$mhz"; t0 = $t0; t1 = $t1 } Start-Sleep -Seconds 3 } } } finally { + $r = (& nvidia-smi -i 0 -rgc 2>&1 | Out-String).Trim() + "RESULT clock reset rgc: $($r -replace "`r?`n", ' | ') readback: $(Smi 'clocks.sm,clocks.max.sm,power.limit')" Start-Sleep -Seconds 2 try { Stop-Process -Id $sampler.Id -Force -ErrorAction SilentlyContinue } catch { } Start-Sleep -Seconds 1 @@ -88,8 +117,8 @@ foreach ($win in $windows) { $s = ($in | Measure-Object -Property sm -Average).Average $m = ($in | Measure-Object -Property mem -Average).Average $tp = ($in | Measure-Object -Property temp -Maximum).Maximum - "RESULT power $($win.pack) bw$($win.bw) samples=$($in.Count) watts_mean=$([math]::Round($w.Average, 1)) watts_min=$($w.Minimum) watts_max=$($w.Maximum) util_mean=$([math]::Round($u, 1)) sm_mhz_mean=$([math]::Round($s)) mem_mhz_mean=$([math]::Round($m)) temp_max=$tp window_s=$([int]($win.t1 - $win.t0).TotalSeconds)" - } else { "RESULT power $($win.pack) bw$($win.bw) no samples in window" } + "RESULT power $($win.pack) $($win.bw) samples=$($in.Count) watts_mean=$([math]::Round($w.Average, 1)) watts_min=$($w.Minimum) watts_max=$($w.Maximum) util_mean=$([math]::Round($u, 1)) sm_mhz_mean=$([math]::Round($s)) mem_mhz_mean=$([math]::Round($m)) temp_max=$tp window_s=$([int]($win.t1 - $win.t0).TotalSeconds)" + } else { "RESULT power $($win.pack) $($win.bw) no samples in window" } } foreach ($r in $rows) { "RESULT sample $($r.ts.ToString('HH:mm:ss')) $($r.w) $($r.util) $($r.sm) $($r.mem) $($r.temp) $($r.used)" } "RESULT gpus_after $(Stamp) $(Smi 'power.draw,power.limit,clocks.sm,clocks.mem,temperature.gpu,memory.used')"