# Program work in the latency shadow (Counter ASIC 3.0 item 8, 6 October 2026) Branch `ca3-shadow`, worker "shadow", on `ca3-coord` 870363c. The item was opened by item 1's finding (`docs/analysis/chip-model-v3.md` section 5.7): the `f = 1` chip stores the dataset in DRAM and recomputes nothing, so the only lever that moves its per-joule row is program work that the honest card hides behind its 128 dependent reads and a chip must pay for with an ALU core. This file holds the knob, the measurements on the Apple M5 Max and the RTX 5090, the verifier's cost law, the chip side re-evaluated on measured watts, the gates a class v4 candidate would need, and the consequences per tier. Every GPU figure says where it was measured and the load average it was taken at; every chip figure is arithmetic on the cited figures of chip-model-v3.md section 5 and is approximate. Nothing here is live: the knob sits behind a `LoadClass` field that version 2 and class v3 never set, no default changed, and nothing was published to the devnet. ## 1. What the hash does today, counted from the code | Quantity | Value | Source | |---|---|---| | Instructions per program | 64, of which 16 are `load` | `igneum-pow/src/generator.rs` `INSTR_COUNT`, `LOAD_SLOTS`; spec 01 section 1.4 | | Iterations per hash | 8 | `ITERATIONS` | | Instructions executed per hash | 512: 128 loads, 384 ALU | 64 x 8 | | ALU op mix of the 48 non-load slots | add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4 (weights, sum 75) | spec 01 section 1.4.2 | | Integer ops per instruction, counted from the emitted statements | add 5 (shift, and, select, two adds), rotr 2 (and, funnel shift), shfl 2 (shuffle, xor), the seven others 1; load 2 (mask, xor) | `igneum-pow/src/emit.rs` instruction lines; weighted mean over the non-load weights 137 / 75 = 1.83 ops per ALU instruction | | Integer ops per hash, this count | about 930: 384 x 1.83 = 703 ALU plus 128 x 2 = 256 for the loads' address mask and fold | arithmetic; the chip model's "512 ops per hash" (section 5.1) counted instructions, not ops | | Per-hash loop | `for it in 0..8 { sel = r0; 64 instructions in order }`, then the fold | spec 01 section 1.7 | | RTX 5090 integer budget | 45.2 T op/s (OpenCL event, integer chain) | `docs/benchmarks/repro.md` 2.2 (branch repro-bench) | | Ops the 5090 could hide before compute binds at 136.1 MH/s | about 332,000 per hash | 45.2 T / 136.1 M | Vendor cost per op differs from this count. Item 6's step costs on the M5 Max (`docs/bench-log.md` "6 October 2026, Counter ASIC 3.0 item 6", branch ca3-reserve 192a683): the live `rotr` step 1.13x and the live `shfl` step 0.86x of the add-xor-rotate chain. The "ops" column below is the count above; the measured watts and rates are what the cards do with it. ## 2. The knob `LoadClass::shadow: Option`, class name `+shx` (`mx8+sh256x27`), `igneum-pow/src/generator.rs`. The shadow block is `S` ALU instructions drawn from the program stream AFTER the 64 base instructions (the ten non-load families at the weights of 1.4.2, the same nine draws per instruction as the program's, the source drawn as on an ALU slot, the width roll and the era windows drawn and ignored when the class takes them), executed `R` times at the end of every iteration, after instruction 63 and before the next iteration samples `sel`, with the iteration's `sel`. Per hash it adds `8 x S x R` ALU instructions and no load. | Property | How the knob keeps it | |---|---| | 16 loads per program, 128 per hash, 4,096 items per unit | the base program is untouched; the block holds no load (`shadow_class_leaves_the_base_program_and_class_v3_untouched`, `generator.rs` tests) | | The acceptance rule of 1.4.6 | unchanged: it interprets the base program (`accept.rs` reads `p.instrs`), and the block is drawn after the base draws, so the attempt and the verdict are those of the class without the shadow, draw for draw. A class v4 that adopts the block would decide whether rule (c) also runs the block (its cost scales with N: 2,048 evaluations x N ops, about 0.7 s at N = 330,000 on one core, approximate) | | v2 and v3 byte-identical | the field is `None` on `V2`, `MX4`, `MX8`, `V3_CLASS`; the interpreter, the six kernel bodies, program.h and program.json emit nothing when it is `None`; `cargo test` in igneum-pow: 54 + 4 + 19 + 7 green, the pinned v2 and v3 packs byte for byte (`tests/packs.rs`) | | Program id | the read-width id with `shadow/ || instrs_le16 || reps_le16` appended, so two block sizes of one seed never share an id | | The CPU verifier | `verify.rs` runs the block `R` times per iteration after the base instructions with the same `step`; bit-exact against Metal on every pack below | | Kernel text | one `for (sh = 0; sh < R; ++sh) { S lines }` block inside the iteration loop, the same statements the base instructions use, in Metal, CUDA and OpenCL, both kernels each; program.h carries `IGNEUM_SHADOW_INSTRS`, `IGNEUM_SHADOW_REPS`, `IGNEUM_SHADOW_INSTRS_PER_HASH`, `IGNEUM_SHADOW_OP_MIX` | Why a block with a repeat count and not a longer straight line: the block is one site in the kernel text, so a 256-instruction block at 88 passes is 256 lines of code, not 22,528, and the compile-ahead per epoch stays at the v3 figure (Metal compile 85 to 97 ms for a 256-instruction block, 231 ms for 1,024, against 1 ms cached for the control; the 1,024-instruction block also costs the M5 Max 17 percent of its hash rate at the same N, section 3). A repeated random block is RandomX's shape (256 instructions x 2,048 iterations per program), and a chip pays it with a general ALU core either way. ## 3. The M5 Max rows (Metal, `with-lock.sh measure`, two sessions, 07:51 to 08:03 UTC) Harness: `proto-metal/packbench --pack --batches 60 --batch-log2 24 --group 256` (built from this branch; vectors are the Rust interpreter's, the fingerprint FNV-1a 64 over the 2^24 outputs at base 0). Power: the GPU and DRAM energy channels of the IOReport "Energy Model" group, sampled at 2 Hz by a 60-line Swift tool in the scratchpad (`gpupower`: `dlopen("/usr/lib/libIOReport.dylib")`, `IOReportCopyChannelsInGroup("Energy Model")`, the `GPU` and `DRAM` channels in mJ; the `GPU Energy` channel in nJ agrees with `GPU` to 2 percent), no root, the mean over the samples after the first 6 s of each run (compile, fills, warm-up) and before its last second. These are the GPU and the memory, not the package: Ember Tune's Apple row reports 38 W for the whole package at 26.7 MH/s (`docs/plans/ember-tune.md`, branch ember-tune, approximate), so the SoC fabric, the CPU and the rest add about 17 W over the 21 W here. The 5090's watts are whole-card. Idle GPU 0.44 W, idle DRAM 0.64 W. Load averages 3.3 to 6.8 (other agents' builds queued behind the measure lock; the GPU was idle: the Mac mines nothing). The control is the pinned class v3 pack `proto-cuda/packs-ca2-mixer/mx8-genesis`; the shadow packs are `proto-cuda/packs-ca3-shadow/*` over the same seed and class. | Pack (class over mx8) | Shadow instrs per hash | Ops per hash (x1.83 + 930) | MH/s (GPU time) | Against the control | GPU W | DRAM W | Microjoules per hash (GPU + DRAM) | Verifier ms per warp, one core, avg of 20 (worst cold) | Bit-exact (3 vectors, 3 in batch, fingerprint) | Load average at start | Compile ms | |---|---|---|---|---|---|---|---|---|---|---|---| | mx8-genesis (control, runs 1, 2, 3) | 0 | 930 | 27.07, 27.07, 27.10 | | 11.2, 9.2, 12.3 | 10.2, 10.0, 10.2 | 0.78 (mean 21.0 W) | 2.062 (2.188) | yes, 7c28cfb06c5c65a9 (the 5 October fingerprint) | 4.6, 6.8, 4.2 | 1 (cached) | | sh256x2 | 4,096 | 8,400 | 26.85 | -0.8% | 15.1 | 10.4 | 0.95 | 2.080 (2.249) | yes, 33e8bbe4c35b2e54 | 4.6 | 85 | | sh256x7 | 14,336 | 27,200 | 26.74 | -1.3% | 18.0 | 10.6 | 1.07 | 2.112 (2.224) | yes, 6cfb70911007520a | 4.2 | 93 | | sh256x13 | 26,624 | 49,700 | 26.75 | -1.2% | 20.6 | 10.5 | 1.16 | 2.212 (2.237) | yes, 59ac286fe2a5a9ef | 3.7 | 93 | | sh64x52 (the same N as sh256x13, a 64-instruction block; runs 1, 2) | 26,624 | 49,700 | 27.72, 27.79 | +2.5% | 19.8, 20.3 | 10.2, 10.3 | 1.09 | 2.137 (2.237) | yes, 9dd010f79d8ca9f4 | 4.4, 4.1 | 59 | | sh1024x3 (a 1,024-instruction block) | 24,576 | 45,900 | 22.53 | -16.8% | 20.2 | 9.7 | 1.33 | 2.150 (2.274) | yes, a05399c819b79aad | 6.0 | 231 | | sh256x27 (runs 1, 2) | 55,296 | 102,100 | 26.48, 26.86 | -1.5% (mean 26.67) | 26.9, 26.9 | 10.4, 10.2 | 1.40 | 2.229 (2.276) | yes, 3d2e8245cc084d07 | 3.3, 5.6 | 96 | | sh256x40 | 81,920 | 150,800 | 25.10 | -7.3% | 26.7 | 10.1 | 1.46 | 2.329 (2.396) | yes, 0844b706302f1c9c | 5.0 | 97 | | sh256x53 (runs 1, 2) | 108,544 | 199,600 | 24.40, 24.09 | -10.4% (mean 24.25) | 28.8, 28.5 | 9.5, 9.7 | 1.58 | 2.427 (2.461) | yes, 4f824b15cf2b124a | 3.9, 4.6 | 85 | | sh256x88 | 180,224 | 330,700 | 21.39 | -21.0% | 31.7 | 8.4 | 1.88 | 2.619 (2.771) | yes, 0572522e39a94d8a | 3.8 | 87 | Commands: `with-lock.sh measure /mac-shadow-session.sh` and `mac-shadow-session2.sh` (each: `uptime`, `gpupower --interval-ms 500` in the background, `packbench` as above, then `igneum-pow bench --seed igneum-genesis --class --warps 20` for the verifier; logs `mac-shadow-session.log`, `mac-shadow-session2.log` in the scratchpad). The control reads 2.5 percent under the 5 October figure for this pack (27.7 MH/s, `docs/plans/mixer-x4.md` 6.2) on three runs over 12 minutes, so the day's denominator is 27.08 and every row is read against it. What the rows say: 1. The M5 Max stays latency-bound to about 100,000 ops per hash (sh256x27: 1.5 percent under the control, inside the run-to-run spread of 1.4 percent) and stops there: 7.3 percent down at 151,000, 10.4 at 200,000, 21 at 331,000. By the 5 percent rule (`docs/plans/counter-asic-2-public.md`, the 2.0 rule) its ceiling is about 130,000 ops per hash (interpolated between the 102,000 and 151,000 rungs), 70,000 shadow instructions, `sh256x35`. Its counted throughput at the compute-bound end is 21.39 M x 330,700 = 7.1 T ops/s, against the 4.4 T op/s the add-xor-rotate chain of item 6 gives (881 G steps/s x 5): the shadow's op mix issues faster than a dependent chain, as it should. The chip model's "about 290,000" for the M5 Max (section 5.7, from memory) was 2.2x too high; the measured bind point is the row. 2. The block size matters on Apple at the same N: a 64-instruction block (sh64x52) runs 2.5 percent ABOVE the control on both runs, 256 holds, 1,024 costs 17 percent (the kernel's instruction footprint; the 1,024-line block at 3 passes is 3,072 instructions of straight-line code per iteration if the compiler unrolls it). A class v4 would fix the block at 64 to 256 instructions. Why a block can be faster than nothing: the shadow spaces the loads of a warp in time, so fewer reads queue at once (the probe's 522 ns at 256 lanes is a queued latency); it is a 2.5 percent effect on this card and is not claimed for the others. 3. Watts: the GPU rises from 11 W to 27 W at 100,000 ops and to 32 W at 331,000, the DRAM stays at 10 W while the rate holds and falls with it. The marginal energy per counted op on the latency-bound rungs: 22 pJ at 8,400 ops (the first rung wakes the ALUs), 11 at 27,000, 7.8 at 50,000, 6.9 at 100,000; at the compute-bound end 2.9 pJ (31.7 W over 7.1 T op/s). The M5 Max's marginal ALU energy at the useful N is therefore 6 to 8 pJ per counted op, the same class as the 5.5 pJ the model assumed for the 5090 (section 5.7, from (575 - 326) W / 45.2 T op/s). 4. Energy per hash, GPU plus DRAM: 0.78 microjoules at the control, 1.40 at 100,000 ops, 1.88 at 331,000. The M5 Max at the hash is 3.0x better per joule than the 5090's 2.34 microjoules (290 W at 124 MH/s in the app, the 5 October sweep record: `docs/plans/miner-eff.md`, branch miner-eff; 326 W was a peak with the prover on) on these channels, 1.7x on the package figure. ## 4. The verifier's cost law One M5 Max performance core, the Rust interpreter, 20 warps, loads 3.3 to 6.8: ms per warp = 2.06 + 3.2 x 10^-3 x (shadow instructions per hash) / 1,000, the slope fitted on the three largest rungs (3.09, 3.36, 3.26 microseconds per 1,000 shadow instructions per warp). That is 0.1 ns per lane-instruction (the register-major loops vectorise over the 32 lanes) or 1.75 microseconds per 1,000 counted ops per warp. The worst cold unit sits 0.05 to 0.17 ms above the average on every rung. On a 2019-class laptop core the rule of chip-model-v3.md section 3 item 1 (2.5x slower) gives 8.0 microseconds per 1,000 shadow instructions per warp. The headroom under the 10 ms gate that the shadow may spend, per pairing (item 2's figures, `docs/plans/counter-asic-3-derivation.md` section 0, bdc07d3, steady / worst cold; the derivation and the shadow are paid by the same verifier on the same unit): | Pairing | Headroom, M5 Max core (steady / worst cold) | Shadow instrs per hash that fit on the M5 Max core | Ops per hash | Headroom, 2019-class core (approximate) | Shadow instrs that fit on the 2019-class core | Ops per hash | |---|---|---|---|---|---|---| | x8 (class v3) + shadow | 7.9 / 7.8 ms | 2.4 million | 4.5 million | 4.8 / 4.6 ms | 575,000 | 1.05 million | | dr368 + shadow | 7.3 / 7.1 | 2.2 million | 4.1 million | 3.3 / 2.8 | 350,000 | 640,000 | | dr736 + shadow | 5.1 / 4.8 | 1.5 million | 2.7 million | none (dr736 alone is 12 ms on that core) | 0 | 0 | So the verifier does not bind the shadow at any N on the ladder in any pairing where the derivation itself fits the gate: 330,700 ops per hash costs 0.56 ms on the M5 Max core and about 1.4 ms on the 2019-class core, which is 12 to 30 percent of the headroom. The binding constraint on N is the honest cards' compute (section 3 for the M5 Max, section 5 for the 5090), not the node. The measured rows: x8 + sh256x88 2.62 ms steady, 2.77 worst cold on the M5 Max core, about 6.5 and 6.9 ms on a 2019-class core by the 2.5x rule; dr736 + sh256x88 about 5.5 ms on the M5 Max core (4.9 + 0.56). The 2019-class core itself is unmeasured (O-1.14); the 2.5x rule stands in for it. ## 5. The RTX 5090 rows (PC 2, 1ccfe586, CUDA through NVRTC) Job `run-ca3-shadow-pc2-20261006` (`tools/ca3-shadow/pc2-shadow-bench.ps1`; fetch job `fetch-ca3-shadow-20261006`, zip sha256 65a65f92..., extracted with sha256 ok at 08:30:26Z), published at 08:29:54Z after `/tmp/igneum-devnet/pc2-ca3.clear` (08:24:27Z) under the mkdir lock (taken 08:29:22Z, released 08:41:03Z after the closing report was read); the job ran 08:30:26Z to 08:38:57Z, 511 s, exit 0, `--stop-miners`, the prover off for the run and back on at the end (`{"ok":true}` both times). The card was EMPTY before the ladder: `nvidia-smi --query-compute-apps` listed no igneum, sp1 or prove process after 0 s (the app had stopped its miners), so these are the card's own figures, not loaded-card ratios (item 2's job found the app's `POST api/cards` with the state's key switched nothing; this job did not use it). Harness: the installed `igneum-worker-cuda.exe` (0.3.11, sha256 2b3b8c92...) `--bench --pack --batches 250 --batch-log2 24 --block-warps 1` (`--block-warps 8`, 120 batches, on three packs), NVRTC 12.8, sm_120, driver 13.3, the packs' own kernel text, the vectors through the bound kernel and the 2^24 fingerprint at base 0; power = `nvidia-smi -l 1` (power.draw, utilization.gpu, clocks.sm, clocks.mem, temperature), the mean over each bench's window after its first 12 s and before its last 2 s (17 to 34 samples per row). Card state before: 431 W limit (the app's 75 percent setting; min 400, max 600), 73.8 W idle at 862 MHz, 56 C; after: 85 W, 65 C. | Pack (class over mx8) | Shadow instrs per hash | Ops per hash | MH/s (wall, 250 x 2^24) | Against the control | Watts, mean (min to max) | SM MHz, mean | Microjoules per hash | Bit-exact (self-test, 96 of 96 lanes; fingerprint 2^24 = the Mac's) | Registers, blocks per SM | NVRTC ms | |---|---|---|---|---|---|---|---|---|---|---| | mx8-genesis (control, first and last) | 0 | 930 | 131.94, 132.47 | | 342.4 (341 to 343), 357.4 (356 to 359) | 3,052, 3,037 | 2.65 (mean 350 W) | yes, 7c28cfb06c5c65a9 | 31, 24 | 158, 160 | | mx8-genesis, 8 warps per block | 0 | 930 | 131.87 | -0.3% | 344.5 | 3,052 | 2.61 | yes | 31, 6 | 158 | | sh256x2 | 4,096 | 8,400 | 132.34 | +0.1% | 354.4 (353 to 356) | 3,037 | 2.68 | yes, 33e8bbe4c35b2e54 | 48, 24 | 248 | | sh256x7 | 14,336 | 27,200 | 132.32 | +0.1% | 385.3 (383 to 388) | 3,037 | 2.91 | yes, 6cfb70911007520a | 48, 24 | 250 | | sh256x13 | 26,624 | 49,700 | 132.28 | +0.1% | 424.6 (422 to 426) | 3,034 | 3.21 | yes, 59ac286fe2a5a9ef | 48, 24 | 240 | | sh256x13, 8 warps per block | 26,624 | 49,700 | 131.86 | -0.3% | 426.1 | 3,030 | 3.23 | yes | 48, 5 | 239 | | sh64x52 (64-instruction block) | 26,624 | 49,700 | 136.77 | +3.5% | 431.5 (the cap; 431 to 433) | 3,024 | 3.15 | yes, 9dd010f79d8ca9f4 | 30, 24 | 181 | | sh1024x3 (1,024-instruction block) | 24,576 | 45,900 | 132.02 | -0.1% | 428.3 (427 to 429) | 3,030 | 3.24 | yes, a05399c819b79aad | 80, 24 | 510 | | sh256x27 | 55,296 | 102,100 | 131.95 | -0.2% | 431.0 (the cap) | 2,824 | 3.27 | yes, 3d2e8245cc084d07 | 48, 24 | 242 | | sh256x40 | 81,920 | 150,800 | 131.75 | -0.3% | 431.0 (the cap) | 2,427 | 3.27 | yes, 0844b706302f1c9c | 48, 24 | 242 | | sh256x53 | 108,544 | 199,600 | 128.67 | -2.7% | 431.0 (the cap) | 1,753 | 3.35 | yes, 4f824b15cf2b124a | 48, 24 | 241 | | sh256x88 | 180,224 | 330,700 | 86.39 | -34.7% | 431.0 (the cap) | 1,834 | 4.99 | yes, 0572522e39a94d8a | 48, 24 | 241 | | sh256x88, 8 warps per block | 180,224 | 330,700 | 85.85 | -35.1% | 431.0 | 1,846 | 5.02 | yes | 48, 5 | 244 | What the rows say: 1. The 5090 holds its rate to 150,800 ops per hash (-0.3 percent) and loses 2.7 percent at 199,600, under a 431 W power limit that the control never reaches (342 to 357 W, the card warming from 62 to 71 C over the job) and that binds from 102,100 ops upward: the card sits at 431.0 W from sh256x27 on and the SM clock falls (3,037 MHz at the control, 2,824 at 102,100 ops, 2,427 at 150,800, 1,753 at 199,600, 1,834 at 330,700). At 330,700 ops it is compute-bound at the capped clock: 86.39 MH/s x 330,700 = 28.6 T counted op/s at 1,834 MHz, which is the 45.2 T op/s budget scaled by 1,834 / 3,050 (27.2 T). So the brief's cap rows are in effect measured: at a 431 W cap the card does the hash plus 150,000 ops at full rate, and the model's "332,000 before compute binds" holds only with the cap at 575 W (where the card would draw about 575 W for 136 MH/s; the 5 October sweep's 400 W floor rows and this 431 W cap bracket what a capped 5090 pays). The 5 percent point at 431 W: about 210,000 ops per hash (interpolated between 199,600 at -2.7 percent and 330,700 at -34.7 percent). 2. The control reads 132.2 MH/s against 135.9 to 137.7 for this pack on 5 October (`docs/bench-log.md`, the mixer job): 3 percent lower, the same harness and card, the card warmer; every row is read against today's control. 3. The block size matters on the 5090 as on the M5 Max: the 64-instruction block (sh64x52) runs 3.5 percent ABOVE the control at 49,700 ops (30 registers against 48, the same occupancy; the spaced loads queue less), 256 holds, and 1,024 holds the rate too (80 registers, still 24 blocks per SM) but compiles in 510 ms against 240. A class v4 block is 64 to 256 instructions. 4. Watts and the marginal energy per counted op, read against the 350 W control: 4.5 pJ at 8,400 ops (noise: a 4 W step), 10.2 pJ at 27,200, 11.6 at 49,700 (12.2 on sh64x52, 13.2 on sh1024x3); above that the cap holds the watts and the clock gives, so the marginal cannot be read. The 5090's marginal ALU energy on a random 32-lane program at its shipping clock is therefore 10 to 13 pJ per counted op, about twice the 5.5 pJ the model assumed (section 5.7: TGP minus 326 W over the whole 45.2 T budget, a figure that mixes the clock-down the cap forces with the ALU energy). The M5 Max's 6.9 pJ at 100,000 ops (section 3) is 0.6 of the 5090's, so on this program the Apple GPU is the better ALU per joule as well as the better memory system per joule. 5. Energy per hash, whole card: 2.65 microjoules at the control (350 W; the app's 290 W at 124 MH/s on 5 October gave 2.34: the bench pushes the card 20 percent harder than the app's job loop does), 3.21 at 49,700 ops, 3.27 at 102,100 and 150,800 (the cap), 3.35 at 199,600, 4.99 at 330,700. 6. Clock rows: OWED. `nvidia-smi -i 0 -lgc 0,2781` answered "The current user does not have permission to change clocks for GPU 00000000:01:00.0" and the job did not try to gain rights (the brief's rule); `-rgc` in the finally block answered the same and the readback showed the clock unlocked (3,037 MHz, max 3,090). The rows need an elevated job (the sweep-5090.ps1 shape, `--elevated`), which is a decision for the coordinator since an elevated job is the administrator-prompt class on PC 2. What the ladder already shows about clocks: at the 431 W cap the card's own governor took the SM clock to 1,753 to 1,834 MHz at 200,000 to 331,000 ops and the rate fell with it, so a locked clock under 2,400 MHz costs hash rate once the program carries over 150,000 ops, and nothing at the control (the control is latency-bound: 136 MH/s at 3,050 MHz and 115 MH/s in the app at the same clock in the 5 October sweep). The brief's `-pl` rows: the 5 October sweep on this card (`docs/plans/counter-asic-2-status.md` "22:16 the 5090 power-limit sweep", relay #224, 22:09 to 22:15Z on PC 2, corrected in `counter-asic-2-rollout.md` 7b): | Cap | Limit W | Draw W | MH/s (the app's API) | MH/W | SM MHz | |---|---|---|---|---|---| | 100% | 575 | 316.2 | 115.42 | 0.365 | 3,051 | | 80% | 460 | 316.1 | 115.60 | 0.366 | 3,050 | | 65% | 400 (the floor) | 310.6 | 114.46 | 0.369 | 3,051 | | 50% | 400 (clamped) | 302.4 | 109.20 | 0.361 | 3,050 | `power.min_limit` is 400 W (read again by this job: min 400, max 600), so `-pl 200` and `-pl 250` cannot be set on a 5090; those rows are re-used, not re-run, and the ladder above adds what the cap does once the program carries work. Ember Tune's own 5090 steps did not run (its second engine never mined; re-run pending the founder); its idle readbacks on PC 1: 90.6 W idle, 2,505 MHz core, 14,001 MHz memory, limit 450 of 575 W. The item 1 denominator moves: at the hash alone the 5090 is 290 W in the app (p95 290, max 295 W at 124 MH/s, the 5 October record on branch miner-eff: 2.34 microjoules) and 342 to 357 W in this bench (132 MH/s: 2.65 microjoules), not 326 W at 136.1 (2.40). The `f = 1` rows of chip-model-v3.md section 5.4 therefore read 5.0x (app) to 5.7x (bench) on GDDR7 against 5.1x, 7.3x to 8.3x on one HBM3 stack against 7.5x, 8.9x to 10.1x on eight against 9.2x: a move of 2 to 11 percent, inside the model's own margin. ## 6. The chip side, re-evaluated (approximate) The `f = 1` chip of chip-model-v3.md section 5.4, at the memory's activate ceiling: GDDR7 (16 devices) 166.4 MH/s at 0.466 microjoules per hash, one HBM3 stack 83.6 MH/s at 0.321, eight stacks 0.262; those are memory, static and controller only. With the shadow the chip adds a core that runs the per-epoch random block at `k` times a GPU's marginal energy per op. The unit is now measured: 11 pJ per counted op, the 5090's marginal at its shipping clock (10.2 to 13.2 pJ on the three rungs under the cap, section 5); the model's 5.5 pJ (section 5.7, from TGP minus 326 W over the whole budget) is half that and is kept as the `k = 0.5` column, which is also about the M5 Max's 6.9 pJ. Four columns: `k = 1` (a chip core as good as the 5090's ALU on a random program: RandomX's argument and the founder's test), `k = 1.5` (1.5x worse), `k = 0.5` (as good as the M5 Max's ALU, or the model's old unit), and `k = 0.3` (a wide-SIMD fixed-datapath array at N5: the int32 datapath energy of section 5.1, 0.06 pJ per add and 0.52 per multiply at the op mix's 29 percent multiplies is 0.19 pJ, with a 2x pipeline overhead and about 8x for the register file, operand wires and the shared instruction fetch over a wide SIMD row, approximate; the floor a chip maker would claim). Section 5.7's "k = 1.5" column was computed as the chip being 1.5x BETTER (0.83 at N = 100,000 is 0.466 + 0.55 / 1.5); the columns below compute `k` as written. Energy per hash = memory + N x 11 pJ x k; N in counted ops. ALU silicon for the core, at the 5090's 45.2 T op/s (what N = 330,700 at 136.1 MH/s needs): 22,600 32-bit lanes at 2.0 GHz. At N5 an int32 multiply-add lane with its register-file slice is about 0.002 mm^2 (approximate: 3,000 to 6,000 gates at about 0.3 square microns per NAND2-equivalent plus registers), so about 45 mm^2 of datapath, 60 to 100 mm^2 with SIMD control and operand networks, $25 to $40 of silicon at the sram-mirror.md yield model ($20,000 wafer, about $0.36 per mm^2 on a 128 mm^2 die; section 5 of that file); at 28 nm the same lanes are about 8x the area, 360 mm^2 of datapath and 500 mm^2 or more with control, a reticle-class die that a $5M to $30M controller project does not carry. Watts for the core at N = 330,700 and 136 MH/s: at `k = 1` 11 pJ x 45.2 T = 497 W (more than the whole 5090 draws for the same work, which is what `k = 1` means), at `k = 0.5` 249 W, at `k = 0.3` 149 W. At N = 100,000 the array is a third of that: 14,000 lanes, about 30 mm^2 at N5, 150 W at `k = 1`, 45 W at `k = 0.3`. Gain per joule against the honest cards, measured watts on both (the M5 Max GPU plus DRAM, the 5090 whole card under its 431 W cap): | N ops per hash | Card microjoules: M5 Max / 5090 | Chip, GDDR7, k = 1 / 1.5 / 0.5 / 0.3 | Gain, GDDR7, against the M5 Max | Against the 5090 | Chip, one HBM3 stack, k = 1 / 1.5 / 0.5 / 0.3 | Gain, HBM3, against the M5 Max | Against the 5090 | |---|---|---|---|---|---|---|---| | 930 (today) | 0.78 / 2.65 (2.34 in the app) | 0.48 / 0.48 / 0.47 / 0.47 | 1.6x | 5.6x (5.0x on the app's watts) | 0.33 | 2.3x | 8.0x (7.3x) | | 49,700 (sh256x13) | 1.16 / 3.21 | 1.01 / 1.29 / 0.74 / 0.63 | 1.15x / 0.90x / 1.6x / 1.9x | 3.2x / 2.5x / 4.3x / 5.1x | 0.87 / 1.14 / 0.59 / 0.48 | 1.3x / 1.0x / 2.0x / 2.4x | 3.7x / 2.8x / 5.4x / 6.6x | | 102,100 (sh256x27, the candidate) | 1.39 / 3.27 | 1.59 / 2.15 / 1.03 / 0.80 | 0.88x / 0.65x / 1.4x / 1.7x | 2.1x / 1.5x / 3.2x / 4.1x | 1.44 / 2.01 / 0.88 / 0.66 | 0.97x / 0.70x / 1.6x / 2.1x | 2.3x / 1.6x / 3.7x / 5.0x | | 150,800 (sh256x40; the M5 Max is 7 percent down) | 1.46 / 3.27 | 2.13 / 2.95 / 1.30 / 0.96 | 0.69x / 0.49x / 1.1x / 1.5x | 1.5x / 1.1x / 2.5x / 3.4x | 1.98 / 2.81 / 1.15 / 0.82 | 0.74x / 0.52x / 1.3x / 1.8x | 1.7x / 1.2x / 2.8x / 4.0x | | 199,600 (sh256x53; the M5 Max 10 percent down, the 5090 2.7) | 1.58 / 3.35 | 2.66 / 3.76 / 1.56 / 1.12 | 0.59x / 0.42x / 1.0x / 1.4x | 1.3x / 0.89x / 2.1x / 3.0x | 2.52 / 3.61 / 1.42 / 0.98 | 0.63x / 0.44x / 1.1x / 1.6x | 1.3x / 0.93x / 2.4x / 3.4x | | 330,700 (sh256x88; both compute-bound) | 1.87 / 4.99 | 4.10 / 5.92 / 2.29 / 1.56 | 0.46x / 0.32x / 0.82x / 1.2x | 1.2x / 0.84x / 2.2x / 3.2x | 3.96 / 5.78 / 2.14 / 1.41 | 0.47x / 0.32x / 0.88x / 1.3x | 1.3x / 0.86x / 2.3x / 3.5x | Reading: at N = 100,000 ops per hash the `f = 1` chip's edge over the 5090 falls from 5.6x to 2.1x on GDDR7 (2.3x on one HBM3 stack) when its core costs what the 5090's ALU costs per op, and to 1.5x (1.6x) when it costs 1.5x more; over the M5 Max it falls from 1.6x to under 1x at `k = 1`. The 2x line against the 5090 is crossed only if the chip's core beats the 5090's ALU by 2x per op (`k = 0.5`: 3.2x, the model's old unit) or more (`k = 0.3`: 4.1x), and against the M5 Max only at `k` under about 0.4. The number that decides the verdict is therefore `k`: the energy per counted op of a chip core that executes a per-epoch random integer program with a 32-lane shuffle, divided by the 5090's measured 11 pJ. Measured on the cards: the 5090 pays 10 to 13 pJ, the M5 Max 6.9 pJ at 100,000 ops and 2.7 to 2.9 pJ at their compute-bound ends (where the clock has dropped and every op is useful). RandomX's argument (and the precedent: no RandomX chip has beaten a CPU per joule) is that a general core cannot reach `k` under about 0.5 on a random program; nothing in this file measures a chip, so the `k = 0.3` column is the attacker's claim and the `k = 1` column ours. The N5 datapath floor (0.19 pJ per op before overheads, section 5.1) says a chip CAN in principle reach `k = 0.3` on a fixed-datapath array, so the shadow lowers the chip's edge, it does not remove it; what removes it is the honest card's own watts (the M5 Max at 0.78 microjoules is already inside 2x of the GDDR7 chip with no shadow at all). ## 7. The gates a class v4 candidate would need The six gates of `docs/plans/counter-asic-2-rollout.md` section 7 plus the verifier on a 2019-class core (O-1.14), with the state today: | Gate | State on the shadow class | What closes it | |---|---|---| | G1 bit-exact on all three vendors against the Mac reference | Metal GREEN on ten packs (3 of 3 vectors standalone and in batch, cache and dataset PASS, fingerprints above); the CUDA text GREEN in the clang emulation on the Mac (`proto-cuda/emu/emu.sh` on sh256x2: cache, dataset and the 3 vector warps standalone and in batch PASS, `with-lock.sh run`, 08:10 UTC) and GREEN on the card (the PC 2 job: self-test PASS on all ten packs, 96 of 96 vector lanes through the bound kernel, every 2^24 fingerprint equal to the Mac's, section 5); the OpenCL text GREEN on Apple OpenCL (`proto-opencl/igneum-bench-cl --bench-pack` on sh256x2 and sh256x88: 96 of 96 vector lanes, self-test PASS); AMD silicon OWED (PC 1) | the PC 2 closing report; a PC 1 job on the 9070 XT (the OpenCL kernels are emitted) | | G2 the CPU verifier exact on 1,000 random hashes per card | not run: the evidence is the 96 vectors and the 2^24 fingerprint per pack on Metal | a serve-mode job per card as in 2.0 | | G3 the generator soundness suite green | `cargo test` 54 + 4 + 19 + 7 green with the new test; the Metal fuzz, edge, stats and determinism runs were NOT run on the shadow class | those four runs at the chosen S and R | | G4 the fast-time 3-node network across an activation | not run (no node change exists: the class is a `LoadClass` field, and the chain's `V3_CLASS` does not set it) | the node seam for v4 (`program_class_v4_activation_daa`) and the run | | G5 the PC-built Windows workers and the Mac workers from the same commit | not applicable yet (no worker change: the workers run the packs' own text) | the release step | | G6 the node change on a fork branch with suites green on PC 2 | not applicable yet | the v4 seam | | O-1.14 the verifier on a 2019-class core | by the 2.5x rule every rung fits with x8 (6.9 ms worst cold at 330,700 ops); the core itself is unmeasured | one run on such a core | | Spec text | none written (no PROPOSED entry until the rows are complete); the block would be a class v4 field beside the class v3 ones in 1.4 and 1.7, the acceptance rule's treatment of the block decided (section 2) | the v4 proposal after the 5090 and AMD rows | ## 8. Consequences per tier A longer program costs the card watts and nothing else while the card stays latency-bound; past its bind point, or once its power cap binds, it costs hash rate. Income per pound of card is unchanged while the rate holds; income per watt falls by the watts ratio. The numbers are this file's (sections 3 and 5); the AMD row is the budget's. | Tier | At N = 100,000 ops per hash (sh256x27, the candidate) | At N = 200,000 | At N = 331,000 (the 5090's full budget at 575 W) | What is being done | |---|---|---|---|---| | Apple user (M-series laptop or desktop; M5 Max measured) | rate -1.5%, GPU + DRAM 21 to 37 W (+16 W; a laptop on battery feels it), income per watt 0.56x, per pound 0.99x | rate -10%, 38 W; per watt 0.49x, per pound 0.90x | rate -21%, 40 W; per watt 0.41x, per pound 0.79x | the project's N is capped by this card: no more than 130,000 ops (the 5 percent rule), 100,000 recommended; the block is 64 to 256 instructions | | NVIDIA, one 24 or 32 GB card (5090 measured, 431 W cap) | rate -0.2%, 350 to 431 W (the cap; +23%), per watt 0.81x, per pound 1.0x | rate -2.7% at the cap, 431 W; per watt 0.79x | rate -35% at the cap (86 MH/s at 1,834 MHz); at a 575 W cap the model says 136 MH/s at about 575 W, per watt 0.61x | measured; the clock rows are OWED (an elevated job); a 5090 owner on the app's 75 percent cap loses nothing at 100,000 and 2.7 percent at 200,000 | | NVIDIA, one 8, 12 or 16 GB card | the same shape per card: the ALU budget scales with the SM count and the cap with the board, so a 4070-class card (about 30 T op/s, approximate) binds near 200,000 ops at its rate and holds at 100,000 | a 4060-class card (about 15 T op/s) binds near 100,000 and is the first NVIDIA tier to pay in rate | | no card we own is in this tier; the ladder runs on any CUDA card through the same job, which is the next NVIDIA measurement | | AMD, one 16 GB card (RX 9070 XT) | OWED: the budget (about 650,000 ops, approximate, section 5.7) says it holds at 100,000 and at 200,000; its watts at the hash are unmeasured | | | the PC 1 job when the desk is free; the OpenCL kernels are in every pack and pass on Apple OpenCL | | A rig | per card as above; a rig's bill is watts, so at 100,000 ops a 5090 rig pays about 23 percent more electricity for the same hash (its cards at their cap), an Apple rig 76 percent more on the GPU and memory channels | | | the recommended N is the smallest that halves the chip row; the rig pays it only if class v4 is adopted | | A pool user | nothing changes in shares or payout: the hash rate holds at the recommended N on every card measured | | | | | A chip | must add a 14,000-lane ALU array at N = 100,000 (about 30 mm^2 at N5, 150 W at `k = 1`) and a 22,600-lane array at 331,000 (60 to 100 mm^2, 500 W); its per-joule edge over the 5090 falls from 5.6x to 2.1x at 100,000 and `k = 1`, over the M5 Max from 1.6x to 0.9x | | | the `k` question goes to the external cryptanalysis and chip review (item 3): can a core run a random 32-lane program under half the 5090's 11 pJ per op | | The public claim | at N = 0 the honest M5 Max is 1.6x from the `f = 1` GDDR7 chip per joule by these channels and the 5090 5.0x to 5.6x; at N = 100,000 and `k = 1` they read 0.9x and 2.1x | | | nothing on the site changes from this file | The N the project should pick, by the 2.0 rule (no card we own loses more than 5 percent): 100,000 ops per hash, `S = 256`, `R = 27` (55,296 shadow instructions per hash): the M5 Max loses 1.5 percent, the 5090 0.2 percent under its 431 W cap, the 9070 XT holds by its budget (owed). The M5 Max's ceiling is about 130,000, the 5090's at 431 W about 210,000. The block size is 64 to 256 instructions, never 1,024. ## 9. Unverified and owed - The RTX 5090 clock rows (`-lgc` at the four Ember caps): OWED, nvidia-smi refused the lock without administrator rights and the job did not ask for them; an elevated job (the sweep-5090.ps1 shape) is the coordinator's decision. The 5090 rows were taken under the app's 431 W cap, which binds from 102,100 ops; the rows at 575 W are the model's. - The RX 9070 XT rows: OWED (PC 1 not released today). - The Mac watts are the IOReport GPU and DRAM channels (no root, the private IOReport library through dlopen), not powermetrics and not the package; Ember Tune's 38 W package figure is approximate and from another session. - The ops count (1.83 per ALU instruction) is a convention counted from the emitted statements; the vendors' real cost per op differs (item 6's step ratios), which is why the rows carry instructions and watts, not ops alone. - The chip side is arithmetic: the 11 pJ unit is the 5090's measured marginal on three rungs (10.2 to 13.2), the `k` columns, the lane area and the 28 nm scaling are approximate; no chip was measured, and the `k = 0.3` floor is an estimate of what a fixed-datapath array could claim. - The acceptance rule does not run the block; whether a class v4 should make it is a design decision with the cost stated in section 2. - The 2019-class core is unmeasured (O-1.14); the 2.5x rule stands in. - The Metal fuzz, edge, stats and determinism runs were not made on the shadow class (gate G3 is the crate suite only). ## 10. Verdict GO as a class v4 candidate at N = 100,000 ops per hash (`mx8+sh256x27`, a 256-instruction block at 27 passes; 64 to 256 instructions per block), subject to the 9070 XT row and the gates of section 7. The number that decides it: at N = 100,000 the `f = 1` chip's per-joule edge over the 5090 falls from 5.6x to 2.1x (GDDR7) and 2.3x (one HBM3 stack) at `k = 1`, and over the M5 Max from 1.6x to 0.9x, for a cost of 1.5 percent of the Mac's rate and none of the 5090's, 16 W on the Mac and 81 W on the 5090. It is a GO for the lever, not a closed verdict on the chip: the edge stays over 2x against the 5090 at every `k` under about 0.9, and the honest Apple card already sits under 2x with no shadow at all, so the project's public line should name the honest card's joules (0.78 microjoules on the M5 Max's GPU and memory, 2.34 to 2.65 on the 5090) and the chip core's `k` as the two numbers, not "under 2x". NO-GO above 130,000 ops (the M5 Max loses more than 5 percent) and NO-GO for a block over 256 instructions (17 percent on the M5 Max at 1,024).