igneum/docs/analysis/latency-shadow-2026-10-06.md
igneum-labs 9417a14ac7 Counter ASIC 3.0 item 8: the RTX 5090 rows, the chip side on the measured 11 pJ per op, the verdict
PC 2 job run-ca3-shadow-pc2-20261006 (card empty, every pack bit-exact against the Mac): the 5090 holds its rate to
150,800 ops per hash and loses 2.7 percent at 199,600 under the app's 431 W cap, which binds from 102,100 ops up and
takes the clock from 3,037 to 1,834 MHz (86 MH/s at 330,700 ops); 350 W at the control, 2.65 to 3.27 microjoules per
hash; marginal ALU energy 10 to 13 pJ per counted op. Clock rows OWED (nvidia-smi refused -lgc without rights). Chip
side at N = 100,000 and k = 1: 2.1x over the 5090 on GDDR7, 0.9x over the M5 Max. GO at mx8+sh256x27.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:44:57 +00:00

36 KiB

Program work in the latency shadow (Counter ASIC 3.0 item 8, 6 October 2026)

Branch ca3-shadow, worker "shadow", on ca3-coord 870363c. The item was opened by item 1's finding (docs/analysis/chip-model-v3.md section 5.7): the f = 1 chip stores the dataset in DRAM and recomputes nothing, so the only lever that moves its per-joule row is program work that the honest card hides behind its 128 dependent reads and a chip must pay for with an ALU core. This file holds the knob, the measurements on the Apple M5 Max and the RTX 5090, the verifier's cost law, the chip side re-evaluated on measured watts, the gates a class v4 candidate would need, and the consequences per tier. Every GPU figure says where it was measured and the load average it was taken at; every chip figure is arithmetic on the cited figures of chip-model-v3.md section 5 and is approximate. Nothing here is live: the knob sits behind a LoadClass field that version 2 and class v3 never set, no default changed, and nothing was published to the devnet.

1. What the hash does today, counted from the code

Quantity Value Source
Instructions per program 64, of which 16 are load igneum-pow/src/generator.rs INSTR_COUNT, LOAD_SLOTS; spec 01 section 1.4
Iterations per hash 8 ITERATIONS
Instructions executed per hash 512: 128 loads, 384 ALU 64 x 8
ALU op mix of the 48 non-load slots add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4 (weights, sum 75) spec 01 section 1.4.2
Integer ops per instruction, counted from the emitted statements add 5 (shift, and, select, two adds), rotr 2 (and, funnel shift), shfl 2 (shuffle, xor), the seven others 1; load 2 (mask, xor) igneum-pow/src/emit.rs instruction lines; weighted mean over the non-load weights 137 / 75 = 1.83 ops per ALU instruction
Integer ops per hash, this count about 930: 384 x 1.83 = 703 ALU plus 128 x 2 = 256 for the loads' address mask and fold arithmetic; the chip model's "512 ops per hash" (section 5.1) counted instructions, not ops
Per-hash loop for it in 0..8 { sel = r0; 64 instructions in order }, then the fold spec 01 section 1.7
RTX 5090 integer budget 45.2 T op/s (OpenCL event, integer chain) docs/benchmarks/repro.md 2.2 (branch repro-bench)
Ops the 5090 could hide before compute binds at 136.1 MH/s about 332,000 per hash 45.2 T / 136.1 M

Vendor cost per op differs from this count. Item 6's step costs on the M5 Max (docs/bench-log.md "6 October 2026, Counter ASIC 3.0 item 6", branch ca3-reserve 192a683): the live rotr step 1.13x and the live shfl step 0.86x of the add-xor-rotate chain. The "ops" column below is the count above; the measured watts and rates are what the cards do with it.

2. The knob

LoadClass::shadow: Option<ShadowClass { instrs, reps }>, class name <class>+sh<S>x<R> (mx8+sh256x27), igneum-pow/src/generator.rs. The shadow block is S ALU instructions drawn from the program stream AFTER the 64 base instructions (the ten non-load families at the weights of 1.4.2, the same nine draws per instruction as the program's, the source drawn as on an ALU slot, the width roll and the era windows drawn and ignored when the class takes them), executed R times at the end of every iteration, after instruction 63 and before the next iteration samples sel, with the iteration's sel. Per hash it adds 8 x S x R ALU instructions and no load.

Property How the knob keeps it
16 loads per program, 128 per hash, 4,096 items per unit the base program is untouched; the block holds no load (shadow_class_leaves_the_base_program_and_class_v3_untouched, generator.rs tests)
The acceptance rule of 1.4.6 unchanged: it interprets the base program (accept.rs reads p.instrs), and the block is drawn after the base draws, so the attempt and the verdict are those of the class without the shadow, draw for draw. A class v4 that adopts the block would decide whether rule (c) also runs the block (its cost scales with N: 2,048 evaluations x N ops, about 0.7 s at N = 330,000 on one core, approximate)
v2 and v3 byte-identical the field is None on V2, MX4, MX8, V3_CLASS; the interpreter, the six kernel bodies, program.h and program.json emit nothing when it is None; cargo test in igneum-pow: 54 + 4 + 19 + 7 green, the pinned v2 and v3 packs byte for byte (tests/packs.rs)
Program id the read-width id with `shadow/
The CPU verifier verify.rs runs the block R times per iteration after the base instructions with the same step; bit-exact against Metal on every pack below
Kernel text one for (sh = 0; sh < R; ++sh) { S lines } block inside the iteration loop, the same statements the base instructions use, in Metal, CUDA and OpenCL, both kernels each; program.h carries IGNEUM_SHADOW_INSTRS, IGNEUM_SHADOW_REPS, IGNEUM_SHADOW_INSTRS_PER_HASH, IGNEUM_SHADOW_OP_MIX

Why a block with a repeat count and not a longer straight line: the block is one site in the kernel text, so a 256-instruction block at 88 passes is 256 lines of code, not 22,528, and the compile-ahead per epoch stays at the v3 figure (Metal compile 85 to 97 ms for a 256-instruction block, 231 ms for 1,024, against 1 ms cached for the control; the 1,024-instruction block also costs the M5 Max 17 percent of its hash rate at the same N, section 3). A repeated random block is RandomX's shape (256 instructions x 2,048 iterations per program), and a chip pays it with a general ALU core either way.

3. The M5 Max rows (Metal, with-lock.sh measure, two sessions, 07:51 to 08:03 UTC)

Harness: proto-metal/packbench --pack <dir> --batches 60 --batch-log2 24 --group 256 (built from this branch; vectors are the Rust interpreter's, the fingerprint FNV-1a 64 over the 2^24 outputs at base 0). Power: the GPU and DRAM energy channels of the IOReport "Energy Model" group, sampled at 2 Hz by a 60-line Swift tool in the scratchpad (gpupower: dlopen("/usr/lib/libIOReport.dylib"), IOReportCopyChannelsInGroup("Energy Model"), the GPU and DRAM channels in mJ; the GPU Energy channel in nJ agrees with GPU to 2 percent), no root, the mean over the samples after the first 6 s of each run (compile, fills, warm-up) and before its last second. These are the GPU and the memory, not the package: Ember Tune's Apple row reports 38 W for the whole package at 26.7 MH/s (docs/plans/ember-tune.md, branch ember-tune, approximate), so the SoC fabric, the CPU and the rest add about 17 W over the 21 W here. The 5090's watts are whole-card. Idle GPU 0.44 W, idle DRAM 0.64 W. Load averages 3.3 to 6.8 (other agents' builds queued behind the measure lock; the GPU was idle: the Mac mines nothing). The control is the pinned class v3 pack proto-cuda/packs-ca2-mixer/mx8-genesis; the shadow packs are proto-cuda/packs-ca3-shadow/* over the same seed and class.

Pack (class over mx8) Shadow instrs per hash Ops per hash (x1.83 + 930) MH/s (GPU time) Against the control GPU W DRAM W Microjoules per hash (GPU + DRAM) Verifier ms per warp, one core, avg of 20 (worst cold) Bit-exact (3 vectors, 3 in batch, fingerprint) Load average at start Compile ms
mx8-genesis (control, runs 1, 2, 3) 0 930 27.07, 27.07, 27.10 11.2, 9.2, 12.3 10.2, 10.0, 10.2 0.78 (mean 21.0 W) 2.062 (2.188) yes, 7c28cfb06c5c65a9 (the 5 October fingerprint) 4.6, 6.8, 4.2 1 (cached)
sh256x2 4,096 8,400 26.85 -0.8% 15.1 10.4 0.95 2.080 (2.249) yes, 33e8bbe4c35b2e54 4.6 85
sh256x7 14,336 27,200 26.74 -1.3% 18.0 10.6 1.07 2.112 (2.224) yes, 6cfb70911007520a 4.2 93
sh256x13 26,624 49,700 26.75 -1.2% 20.6 10.5 1.16 2.212 (2.237) yes, 59ac286fe2a5a9ef 3.7 93
sh64x52 (the same N as sh256x13, a 64-instruction block; runs 1, 2) 26,624 49,700 27.72, 27.79 +2.5% 19.8, 20.3 10.2, 10.3 1.09 2.137 (2.237) yes, 9dd010f79d8ca9f4 4.4, 4.1 59
sh1024x3 (a 1,024-instruction block) 24,576 45,900 22.53 -16.8% 20.2 9.7 1.33 2.150 (2.274) yes, a05399c819b79aad 6.0 231
sh256x27 (runs 1, 2) 55,296 102,100 26.48, 26.86 -1.5% (mean 26.67) 26.9, 26.9 10.4, 10.2 1.40 2.229 (2.276) yes, 3d2e8245cc084d07 3.3, 5.6 96
sh256x40 81,920 150,800 25.10 -7.3% 26.7 10.1 1.46 2.329 (2.396) yes, 0844b706302f1c9c 5.0 97
sh256x53 (runs 1, 2) 108,544 199,600 24.40, 24.09 -10.4% (mean 24.25) 28.8, 28.5 9.5, 9.7 1.58 2.427 (2.461) yes, 4f824b15cf2b124a 3.9, 4.6 85
sh256x88 180,224 330,700 21.39 -21.0% 31.7 8.4 1.88 2.619 (2.771) yes, 0572522e39a94d8a 3.8 87

Commands: with-lock.sh measure <scratchpad>/mac-shadow-session.sh and mac-shadow-session2.sh (each: uptime, gpupower --interval-ms 500 in the background, packbench as above, then igneum-pow bench --seed igneum-genesis --class <class> --warps 20 for the verifier; logs mac-shadow-session.log, mac-shadow-session2.log in the scratchpad). The control reads 2.5 percent under the 5 October figure for this pack (27.7 MH/s, docs/plans/mixer-x4.md 6.2) on three runs over 12 minutes, so the day's denominator is 27.08 and every row is read against it.

What the rows say:

  1. The M5 Max stays latency-bound to about 100,000 ops per hash (sh256x27: 1.5 percent under the control, inside the run-to-run spread of 1.4 percent) and stops there: 7.3 percent down at 151,000, 10.4 at 200,000, 21 at 331,000. By the 5 percent rule (docs/plans/counter-asic-2-public.md, the 2.0 rule) its ceiling is about 130,000 ops per hash (interpolated between the 102,000 and 151,000 rungs), 70,000 shadow instructions, sh256x35. Its counted throughput at the compute-bound end is 21.39 M x 330,700 = 7.1 T ops/s, against the 4.4 T op/s the add-xor-rotate chain of item 6 gives (881 G steps/s x 5): the shadow's op mix issues faster than a dependent chain, as it should. The chip model's "about 290,000" for the M5 Max (section 5.7, from memory) was 2.2x too high; the measured bind point is the row.
  2. The block size matters on Apple at the same N: a 64-instruction block (sh64x52) runs 2.5 percent ABOVE the control on both runs, 256 holds, 1,024 costs 17 percent (the kernel's instruction footprint; the 1,024-line block at 3 passes is 3,072 instructions of straight-line code per iteration if the compiler unrolls it). A class v4 would fix the block at 64 to 256 instructions. Why a block can be faster than nothing: the shadow spaces the loads of a warp in time, so fewer reads queue at once (the probe's 522 ns at 256 lanes is a queued latency); it is a 2.5 percent effect on this card and is not claimed for the others.
  3. Watts: the GPU rises from 11 W to 27 W at 100,000 ops and to 32 W at 331,000, the DRAM stays at 10 W while the rate holds and falls with it. The marginal energy per counted op on the latency-bound rungs: 22 pJ at 8,400 ops (the first rung wakes the ALUs), 11 at 27,000, 7.8 at 50,000, 6.9 at 100,000; at the compute-bound end 2.9 pJ (31.7 W over 7.1 T op/s). The M5 Max's marginal ALU energy at the useful N is therefore 6 to 8 pJ per counted op, the same class as the 5.5 pJ the model assumed for the 5090 (section 5.7, from (575 - 326) W / 45.2 T op/s).
  4. Energy per hash, GPU plus DRAM: 0.78 microjoules at the control, 1.40 at 100,000 ops, 1.88 at 331,000. The M5 Max at the hash is 3.0x better per joule than the 5090's 2.34 microjoules (290 W at 124 MH/s in the app, the 5 October sweep record: docs/plans/miner-eff.md, branch miner-eff; 326 W was a peak with the prover on) on these channels, 1.7x on the package figure.

4. The verifier's cost law

One M5 Max performance core, the Rust interpreter, 20 warps, loads 3.3 to 6.8: ms per warp = 2.06 + 3.2 x 10^-3 x (shadow instructions per hash) / 1,000, the slope fitted on the three largest rungs (3.09, 3.36, 3.26 microseconds per 1,000 shadow instructions per warp). That is 0.1 ns per lane-instruction (the register-major loops vectorise over the 32 lanes) or 1.75 microseconds per 1,000 counted ops per warp. The worst cold unit sits 0.05 to 0.17 ms above the average on every rung. On a 2019-class laptop core the rule of chip-model-v3.md section 3 item 1 (2.5x slower) gives 8.0 microseconds per 1,000 shadow instructions per warp.

The headroom under the 10 ms gate that the shadow may spend, per pairing (item 2's figures, docs/plans/counter-asic-3-derivation.md section 0, bdc07d3, steady / worst cold; the derivation and the shadow are paid by the same verifier on the same unit):

Pairing Headroom, M5 Max core (steady / worst cold) Shadow instrs per hash that fit on the M5 Max core Ops per hash Headroom, 2019-class core (approximate) Shadow instrs that fit on the 2019-class core Ops per hash
x8 (class v3) + shadow 7.9 / 7.8 ms 2.4 million 4.5 million 4.8 / 4.6 ms 575,000 1.05 million
dr368 + shadow 7.3 / 7.1 2.2 million 4.1 million 3.3 / 2.8 350,000 640,000
dr736 + shadow 5.1 / 4.8 1.5 million 2.7 million none (dr736 alone is 12 ms on that core) 0 0

So the verifier does not bind the shadow at any N on the ladder in any pairing where the derivation itself fits the gate: 330,700 ops per hash costs 0.56 ms on the M5 Max core and about 1.4 ms on the 2019-class core, which is 12 to 30 percent of the headroom. The binding constraint on N is the honest cards' compute (section 3 for the M5 Max, section 5 for the 5090), not the node. The measured rows: x8 + sh256x88 2.62 ms steady, 2.77 worst cold on the M5 Max core, about 6.5 and 6.9 ms on a 2019-class core by the 2.5x rule; dr736 + sh256x88 about 5.5 ms on the M5 Max core (4.9 + 0.56). The 2019-class core itself is unmeasured (O-1.14); the 2.5x rule stands in for it.

5. The RTX 5090 rows (PC 2, 1ccfe586, CUDA through NVRTC)

Job run-ca3-shadow-pc2-20261006 (tools/ca3-shadow/pc2-shadow-bench.ps1; fetch job fetch-ca3-shadow-20261006, zip sha256 65a65f92..., extracted with sha256 ok at 08:30:26Z), published at 08:29:54Z after /tmp/igneum-devnet/pc2-ca3.clear (08:24:27Z) under the mkdir lock (taken 08:29:22Z, released 08:41:03Z after the closing report was read); the job ran 08:30:26Z to 08:38:57Z, 511 s, exit 0, --stop-miners, the prover off for the run and back on at the end ({"ok":true} both times). The card was EMPTY before the ladder: nvidia-smi --query-compute-apps listed no igneum, sp1 or prove process after 0 s (the app had stopped its miners), so these are the card's own figures, not loaded-card ratios (item 2's job found the app's POST api/cards with the state's key switched nothing; this job did not use it). Harness: the installed igneum-worker-cuda.exe (0.3.11, sha256 2b3b8c92...) --bench --pack <dir> --batches 250 --batch-log2 24 --block-warps 1 (--block-warps 8, 120 batches, on three packs), NVRTC 12.8, sm_120, driver 13.3, the packs' own kernel text, the vectors through the bound kernel and the 2^24 fingerprint at base 0; power = nvidia-smi -l 1 (power.draw, utilization.gpu, clocks.sm, clocks.mem, temperature), the mean over each bench's window after its first 12 s and before its last 2 s (17 to 34 samples per row). Card state before: 431 W limit (the app's 75 percent setting; min 400, max 600), 73.8 W idle at 862 MHz, 56 C; after: 85 W, 65 C.

Pack (class over mx8) Shadow instrs per hash Ops per hash MH/s (wall, 250 x 2^24) Against the control Watts, mean (min to max) SM MHz, mean Microjoules per hash Bit-exact (self-test, 96 of 96 lanes; fingerprint 2^24 = the Mac's) Registers, blocks per SM NVRTC ms
mx8-genesis (control, first and last) 0 930 131.94, 132.47 342.4 (341 to 343), 357.4 (356 to 359) 3,052, 3,037 2.65 (mean 350 W) yes, 7c28cfb06c5c65a9 31, 24 158, 160
mx8-genesis, 8 warps per block 0 930 131.87 -0.3% 344.5 3,052 2.61 yes 31, 6 158
sh256x2 4,096 8,400 132.34 +0.1% 354.4 (353 to 356) 3,037 2.68 yes, 33e8bbe4c35b2e54 48, 24 248
sh256x7 14,336 27,200 132.32 +0.1% 385.3 (383 to 388) 3,037 2.91 yes, 6cfb70911007520a 48, 24 250
sh256x13 26,624 49,700 132.28 +0.1% 424.6 (422 to 426) 3,034 3.21 yes, 59ac286fe2a5a9ef 48, 24 240
sh256x13, 8 warps per block 26,624 49,700 131.86 -0.3% 426.1 3,030 3.23 yes 48, 5 239
sh64x52 (64-instruction block) 26,624 49,700 136.77 +3.5% 431.5 (the cap; 431 to 433) 3,024 3.15 yes, 9dd010f79d8ca9f4 30, 24 181
sh1024x3 (1,024-instruction block) 24,576 45,900 132.02 -0.1% 428.3 (427 to 429) 3,030 3.24 yes, a05399c819b79aad 80, 24 510
sh256x27 55,296 102,100 131.95 -0.2% 431.0 (the cap) 2,824 3.27 yes, 3d2e8245cc084d07 48, 24 242
sh256x40 81,920 150,800 131.75 -0.3% 431.0 (the cap) 2,427 3.27 yes, 0844b706302f1c9c 48, 24 242
sh256x53 108,544 199,600 128.67 -2.7% 431.0 (the cap) 1,753 3.35 yes, 4f824b15cf2b124a 48, 24 241
sh256x88 180,224 330,700 86.39 -34.7% 431.0 (the cap) 1,834 4.99 yes, 0572522e39a94d8a 48, 24 241
sh256x88, 8 warps per block 180,224 330,700 85.85 -35.1% 431.0 1,846 5.02 yes 48, 5 244

What the rows say:

  1. The 5090 holds its rate to 150,800 ops per hash (-0.3 percent) and loses 2.7 percent at 199,600, under a 431 W power limit that the control never reaches (342 to 357 W, the card warming from 62 to 71 C over the job) and that binds from 102,100 ops upward: the card sits at 431.0 W from sh256x27 on and the SM clock falls (3,037 MHz at the control, 2,824 at 102,100 ops, 2,427 at 150,800, 1,753 at 199,600, 1,834 at 330,700). At 330,700 ops it is compute-bound at the capped clock: 86.39 MH/s x 330,700 = 28.6 T counted op/s at 1,834 MHz, which is the 45.2 T op/s budget scaled by 1,834 / 3,050 (27.2 T). So the brief's cap rows are in effect measured: at a 431 W cap the card does the hash plus 150,000 ops at full rate, and the model's "332,000 before compute binds" holds only with the cap at 575 W (where the card would draw about 575 W for 136 MH/s; the 5 October sweep's 400 W floor rows and this 431 W cap bracket what a capped 5090 pays). The 5 percent point at 431 W: about 210,000 ops per hash (interpolated between 199,600 at -2.7 percent and 330,700 at -34.7 percent).
  2. The control reads 132.2 MH/s against 135.9 to 137.7 for this pack on 5 October (docs/bench-log.md, the mixer job): 3 percent lower, the same harness and card, the card warmer; every row is read against today's control.
  3. The block size matters on the 5090 as on the M5 Max: the 64-instruction block (sh64x52) runs 3.5 percent ABOVE the control at 49,700 ops (30 registers against 48, the same occupancy; the spaced loads queue less), 256 holds, and 1,024 holds the rate too (80 registers, still 24 blocks per SM) but compiles in 510 ms against 240. A class v4 block is 64 to 256 instructions.
  4. Watts and the marginal energy per counted op, read against the 350 W control: 4.5 pJ at 8,400 ops (noise: a 4 W step), 10.2 pJ at 27,200, 11.6 at 49,700 (12.2 on sh64x52, 13.2 on sh1024x3); above that the cap holds the watts and the clock gives, so the marginal cannot be read. The 5090's marginal ALU energy on a random 32-lane program at its shipping clock is therefore 10 to 13 pJ per counted op, about twice the 5.5 pJ the model assumed (section 5.7: TGP minus 326 W over the whole 45.2 T budget, a figure that mixes the clock-down the cap forces with the ALU energy). The M5 Max's 6.9 pJ at 100,000 ops (section 3) is 0.6 of the 5090's, so on this program the Apple GPU is the better ALU per joule as well as the better memory system per joule.
  5. Energy per hash, whole card: 2.65 microjoules at the control (350 W; the app's 290 W at 124 MH/s on 5 October gave 2.34: the bench pushes the card 20 percent harder than the app's job loop does), 3.21 at 49,700 ops, 3.27 at 102,100 and 150,800 (the cap), 3.35 at 199,600, 4.99 at 330,700.
  6. Clock rows: OWED. nvidia-smi -i 0 -lgc 0,2781 answered "The current user does not have permission to change clocks for GPU 00000000:01:00.0" and the job did not try to gain rights (the brief's rule); -rgc in the finally block answered the same and the readback showed the clock unlocked (3,037 MHz, max 3,090). The rows need an elevated job (the sweep-5090.ps1 shape, --elevated), which is a decision for the coordinator since an elevated job is the administrator-prompt class on PC 2. What the ladder already shows about clocks: at the 431 W cap the card's own governor took the SM clock to 1,753 to 1,834 MHz at 200,000 to 331,000 ops and the rate fell with it, so a locked clock under 2,400 MHz costs hash rate once the program carries over 150,000 ops, and nothing at the control (the control is latency-bound: 136 MH/s at 3,050 MHz and 115 MH/s in the app at the same clock in the 5 October sweep).

The brief's -pl rows: the 5 October sweep on this card (docs/plans/counter-asic-2-status.md "22:16 the 5090 power-limit sweep", relay #224, 22:09 to 22:15Z on PC 2, corrected in counter-asic-2-rollout.md 7b):

Cap Limit W Draw W MH/s (the app's API) MH/W SM MHz
100% 575 316.2 115.42 0.365 3,051
80% 460 316.1 115.60 0.366 3,050
65% 400 (the floor) 310.6 114.46 0.369 3,051
50% 400 (clamped) 302.4 109.20 0.361 3,050

power.min_limit is 400 W (read again by this job: min 400, max 600), so -pl 200 and -pl 250 cannot be set on a 5090; those rows are re-used, not re-run, and the ladder above adds what the cap does once the program carries work. Ember Tune's own 5090 steps did not run (its second engine never mined; re-run pending the project lead); its idle readbacks on PC 1: 90.6 W idle, 2,505 MHz core, 14,001 MHz memory, limit 450 of 575 W.

The item 1 denominator moves: at the hash alone the 5090 is 290 W in the app (p95 290, max 295 W at 124 MH/s, the 5 October record on branch miner-eff: 2.34 microjoules) and 342 to 357 W in this bench (132 MH/s: 2.65 microjoules), not 326 W at 136.1 (2.40). The f = 1 rows of chip-model-v3.md section 5.4 therefore read 5.0x (app) to 5.7x (bench) on GDDR7 against 5.1x, 7.3x to 8.3x on one HBM3 stack against 7.5x, 8.9x to 10.1x on eight against 9.2x: a move of 2 to 11 percent, inside the model's own margin.

6. The chip side, re-evaluated (approximate)

The f = 1 chip of chip-model-v3.md section 5.4, at the memory's activate ceiling: GDDR7 (16 devices) 166.4 MH/s at 0.466 microjoules per hash, one HBM3 stack 83.6 MH/s at 0.321, eight stacks 0.262; those are memory, static and controller only. With the shadow the chip adds a core that runs the per-epoch random block at k times a GPU's marginal energy per op. The unit is now measured: 11 pJ per counted op, the 5090's marginal at its shipping clock (10.2 to 13.2 pJ on the three rungs under the cap, section 5); the model's 5.5 pJ (section 5.7, from TGP minus 326 W over the whole budget) is half that and is kept as the k = 0.5 column, which is also about the M5 Max's 6.9 pJ. Four columns: k = 1 (a chip core as good as the 5090's ALU on a random program: RandomX's argument and the project lead's test), k = 1.5 (1.5x worse), k = 0.5 (as good as the M5 Max's ALU, or the model's old unit), and k = 0.3 (a wide-SIMD fixed-datapath array at N5: the int32 datapath energy of section 5.1, 0.06 pJ per add and 0.52 per multiply at the op mix's 29 percent multiplies is 0.19 pJ, with a 2x pipeline overhead and about 8x for the register file, operand wires and the shared instruction fetch over a wide SIMD row, approximate; the floor a chip maker would claim). Section 5.7's "k = 1.5" column was computed as the chip being 1.5x BETTER (0.83 at N = 100,000 is 0.466 + 0.55 / 1.5); the columns below compute k as written. Energy per hash = memory + N x 11 pJ x k; N in counted ops.

ALU silicon for the core, at the 5090's 45.2 T op/s (what N = 330,700 at 136.1 MH/s needs): 22,600 32-bit lanes at 2.0 GHz. At N5 an int32 multiply-add lane with its register-file slice is about 0.002 mm^2 (approximate: 3,000 to 6,000 gates at about 0.3 square microns per NAND2-equivalent plus registers), so about 45 mm^2 of datapath, 60 to 100 mm^2 with SIMD control and operand networks, $25 to $40 of silicon at the sram-mirror.md yield model ($20,000 wafer, about $0.36 per mm^2 on a 128 mm^2 die; section 5 of that file); at 28 nm the same lanes are about 8x the area, 360 mm^2 of datapath and 500 mm^2 or more with control, a reticle-class die that a $5M to $30M controller project does not carry. Watts for the core at N = 330,700 and 136 MH/s: at k = 1 11 pJ x 45.2 T = 497 W (more than the whole 5090 draws for the same work, which is what k = 1 means), at k = 0.5 249 W, at k = 0.3 149 W. At N = 100,000 the array is a third of that: 14,000 lanes, about 30 mm^2 at N5, 150 W at k = 1, 45 W at k = 0.3.

Gain per joule against the honest cards, measured watts on both (the M5 Max GPU plus DRAM, the 5090 whole card under its 431 W cap):

N ops per hash Card microjoules: M5 Max / 5090 Chip, GDDR7, k = 1 / 1.5 / 0.5 / 0.3 Gain, GDDR7, against the M5 Max Against the 5090 Chip, one HBM3 stack, k = 1 / 1.5 / 0.5 / 0.3 Gain, HBM3, against the M5 Max Against the 5090
930 (today) 0.78 / 2.65 (2.34 in the app) 0.48 / 0.48 / 0.47 / 0.47 1.6x 5.6x (5.0x on the app's watts) 0.33 2.3x 8.0x (7.3x)
49,700 (sh256x13) 1.16 / 3.21 1.01 / 1.29 / 0.74 / 0.63 1.15x / 0.90x / 1.6x / 1.9x 3.2x / 2.5x / 4.3x / 5.1x 0.87 / 1.14 / 0.59 / 0.48 1.3x / 1.0x / 2.0x / 2.4x 3.7x / 2.8x / 5.4x / 6.6x
102,100 (sh256x27, the candidate) 1.39 / 3.27 1.59 / 2.15 / 1.03 / 0.80 0.88x / 0.65x / 1.4x / 1.7x 2.1x / 1.5x / 3.2x / 4.1x 1.44 / 2.01 / 0.88 / 0.66 0.97x / 0.70x / 1.6x / 2.1x 2.3x / 1.6x / 3.7x / 5.0x
150,800 (sh256x40; the M5 Max is 7 percent down) 1.46 / 3.27 2.13 / 2.95 / 1.30 / 0.96 0.69x / 0.49x / 1.1x / 1.5x 1.5x / 1.1x / 2.5x / 3.4x 1.98 / 2.81 / 1.15 / 0.82 0.74x / 0.52x / 1.3x / 1.8x 1.7x / 1.2x / 2.8x / 4.0x
199,600 (sh256x53; the M5 Max 10 percent down, the 5090 2.7) 1.58 / 3.35 2.66 / 3.76 / 1.56 / 1.12 0.59x / 0.42x / 1.0x / 1.4x 1.3x / 0.89x / 2.1x / 3.0x 2.52 / 3.61 / 1.42 / 0.98 0.63x / 0.44x / 1.1x / 1.6x 1.3x / 0.93x / 2.4x / 3.4x
330,700 (sh256x88; both compute-bound) 1.87 / 4.99 4.10 / 5.92 / 2.29 / 1.56 0.46x / 0.32x / 0.82x / 1.2x 1.2x / 0.84x / 2.2x / 3.2x 3.96 / 5.78 / 2.14 / 1.41 0.47x / 0.32x / 0.88x / 1.3x 1.3x / 0.86x / 2.3x / 3.5x

Reading: at N = 100,000 ops per hash the f = 1 chip's edge over the 5090 falls from 5.6x to 2.1x on GDDR7 (2.3x on one HBM3 stack) when its core costs what the 5090's ALU costs per op, and to 1.5x (1.6x) when it costs 1.5x more; over the M5 Max it falls from 1.6x to under 1x at k = 1. The 2x line against the 5090 is crossed only if the chip's core beats the 5090's ALU by 2x per op (k = 0.5: 3.2x, the model's old unit) or more (k = 0.3: 4.1x), and against the M5 Max only at k under about 0.4. The number that decides the verdict is therefore k: the energy per counted op of a chip core that executes a per-epoch random integer program with a 32-lane shuffle, divided by the 5090's measured 11 pJ. Measured on the cards: the 5090 pays 10 to 13 pJ, the M5 Max 6.9 pJ at 100,000 ops and 2.7 to 2.9 pJ at their compute-bound ends (where the clock has dropped and every op is useful). RandomX's argument (and the precedent: no RandomX chip has beaten a CPU per joule) is that a general core cannot reach k under about 0.5 on a random program; nothing in this file measures a chip, so the k = 0.3 column is the attacker's claim and the k = 1 column ours. The N5 datapath floor (0.19 pJ per op before overheads, section 5.1) says a chip CAN in principle reach k = 0.3 on a fixed-datapath array, so the shadow lowers the chip's edge, it does not remove it; what removes it is the honest card's own watts (the M5 Max at 0.78 microjoules is already inside 2x of the GDDR7 chip with no shadow at all).

7. The gates a class v4 candidate would need

The six gates of docs/plans/counter-asic-2-rollout.md section 7 plus the verifier on a 2019-class core (O-1.14), with the state today:

Gate State on the shadow class What closes it
G1 bit-exact on all three vendors against the Mac reference Metal GREEN on ten packs (3 of 3 vectors standalone and in batch, cache and dataset PASS, fingerprints above); the CUDA text GREEN in the clang emulation on the Mac (proto-cuda/emu/emu.sh on sh256x2: cache, dataset and the 3 vector warps standalone and in batch PASS, with-lock.sh run, 08:10 UTC) and GREEN on the card (the PC 2 job: self-test PASS on all ten packs, 96 of 96 vector lanes through the bound kernel, every 2^24 fingerprint equal to the Mac's, section 5); the OpenCL text GREEN on Apple OpenCL (proto-opencl/igneum-bench-cl --bench-pack on sh256x2 and sh256x88: 96 of 96 vector lanes, self-test PASS); AMD silicon OWED (PC 1) the PC 2 closing report; a PC 1 job on the 9070 XT (the OpenCL kernels are emitted)
G2 the CPU verifier exact on 1,000 random hashes per card not run: the evidence is the 96 vectors and the 2^24 fingerprint per pack on Metal a serve-mode job per card as in 2.0
G3 the generator soundness suite green cargo test 54 + 4 + 19 + 7 green with the new test; the Metal fuzz, edge, stats and determinism runs were NOT run on the shadow class those four runs at the chosen S and R
G4 the fast-time 3-node network across an activation not run (no node change exists: the class is a LoadClass field, and the chain's V3_CLASS does not set it) the node seam for v4 (program_class_v4_activation_daa) and the run
G5 the PC-built Windows workers and the Mac workers from the same commit not applicable yet (no worker change: the workers run the packs' own text) the release step
G6 the node change on a fork branch with suites green on PC 2 not applicable yet the v4 seam
O-1.14 the verifier on a 2019-class core by the 2.5x rule every rung fits with x8 (6.9 ms worst cold at 330,700 ops); the core itself is unmeasured one run on such a core
Spec text none written (no PROPOSED entry until the rows are complete); the block would be a class v4 field beside the class v3 ones in 1.4 and 1.7, the acceptance rule's treatment of the block decided (section 2) the v4 proposal after the 5090 and AMD rows

8. Consequences per tier

A longer program costs the card watts and nothing else while the card stays latency-bound; past its bind point, or once its power cap binds, it costs hash rate. Income per pound of card is unchanged while the rate holds; income per watt falls by the watts ratio. The numbers are this file's (sections 3 and 5); the AMD row is the budget's.

Tier At N = 100,000 ops per hash (sh256x27, the candidate) At N = 200,000 At N = 331,000 (the 5090's full budget at 575 W) What is being done
Apple user (M-series laptop or desktop; M5 Max measured) rate -1.5%, GPU + DRAM 21 to 37 W (+16 W; a laptop on battery feels it), income per watt 0.56x, per pound 0.99x rate -10%, 38 W; per watt 0.49x, per pound 0.90x rate -21%, 40 W; per watt 0.41x, per pound 0.79x the project's N is capped by this card: no more than 130,000 ops (the 5 percent rule), 100,000 recommended; the block is 64 to 256 instructions
NVIDIA, one 24 or 32 GB card (5090 measured, 431 W cap) rate -0.2%, 350 to 431 W (the cap; +23%), per watt 0.81x, per pound 1.0x rate -2.7% at the cap, 431 W; per watt 0.79x rate -35% at the cap (86 MH/s at 1,834 MHz); at a 575 W cap the model says 136 MH/s at about 575 W, per watt 0.61x measured; the clock rows are OWED (an elevated job); a 5090 owner on the app's 75 percent cap loses nothing at 100,000 and 2.7 percent at 200,000
NVIDIA, one 8, 12 or 16 GB card the same shape per card: the ALU budget scales with the SM count and the cap with the board, so a 4070-class card (about 30 T op/s, approximate) binds near 200,000 ops at its rate and holds at 100,000 a 4060-class card (about 15 T op/s) binds near 100,000 and is the first NVIDIA tier to pay in rate no card we own is in this tier; the ladder runs on any CUDA card through the same job, which is the next NVIDIA measurement
AMD, one 16 GB card (RX 9070 XT) OWED: the budget (about 650,000 ops, approximate, section 5.7) says it holds at 100,000 and at 200,000; its watts at the hash are unmeasured the PC 1 job when the desk is free; the OpenCL kernels are in every pack and pass on Apple OpenCL
A rig per card as above; a rig's bill is watts, so at 100,000 ops a 5090 rig pays about 23 percent more electricity for the same hash (its cards at their cap), an Apple rig 76 percent more on the GPU and memory channels the recommended N is the smallest that halves the chip row; the rig pays it only if class v4 is adopted
A pool user nothing changes in shares or payout: the hash rate holds at the recommended N on every card measured
A chip must add a 14,000-lane ALU array at N = 100,000 (about 30 mm^2 at N5, 150 W at k = 1) and a 22,600-lane array at 331,000 (60 to 100 mm^2, 500 W); its per-joule edge over the 5090 falls from 5.6x to 2.1x at 100,000 and k = 1, over the M5 Max from 1.6x to 0.9x the k question goes to the external cryptanalysis and chip review (item 3): can a core run a random 32-lane program under half the 5090's 11 pJ per op
The public claim at N = 0 the honest M5 Max is 1.6x from the f = 1 GDDR7 chip per joule by these channels and the 5090 5.0x to 5.6x; at N = 100,000 and k = 1 they read 0.9x and 2.1x nothing on the site changes from this file

The N the project should pick, by the 2.0 rule (no card we own loses more than 5 percent): 100,000 ops per hash, S = 256, R = 27 (55,296 shadow instructions per hash): the M5 Max loses 1.5 percent, the 5090 0.2 percent under its 431 W cap, the 9070 XT holds by its budget (owed). The M5 Max's ceiling is about 130,000, the 5090's at 431 W about 210,000. The block size is 64 to 256 instructions, never 1,024.

9. Unverified and owed

  • The RTX 5090 clock rows (-lgc at the four Ember caps): OWED, nvidia-smi refused the lock without administrator rights and the job did not ask for them; an elevated job (the sweep-5090.ps1 shape) is the coordinator's decision. The 5090 rows were taken under the app's 431 W cap, which binds from 102,100 ops; the rows at 575 W are the model's.
  • The RX 9070 XT rows: OWED (PC 1 not released today).
  • The Mac watts are the IOReport GPU and DRAM channels (no root, the private IOReport library through dlopen), not powermetrics and not the package; Ember Tune's 38 W package figure is approximate and from another session.
  • The ops count (1.83 per ALU instruction) is a convention counted from the emitted statements; the vendors' real cost per op differs (item 6's step ratios), which is why the rows carry instructions and watts, not ops alone.
  • The chip side is arithmetic: the 11 pJ unit is the 5090's measured marginal on three rungs (10.2 to 13.2), the k columns, the lane area and the 28 nm scaling are approximate; no chip was measured, and the k = 0.3 floor is an estimate of what a fixed-datapath array could claim.
  • The acceptance rule does not run the block; whether a class v4 should make it is a design decision with the cost stated in section 2.
  • The 2019-class core is unmeasured (O-1.14); the 2.5x rule stands in.
  • The Metal fuzz, edge, stats and determinism runs were not made on the shadow class (gate G3 is the crate suite only).

10. Verdict

GO as a class v4 candidate at N = 100,000 ops per hash (mx8+sh256x27, a 256-instruction block at 27 passes; 64 to 256 instructions per block), subject to the 9070 XT row and the gates of section 7. The number that decides it: at N = 100,000 the f = 1 chip's per-joule edge over the 5090 falls from 5.6x to 2.1x (GDDR7) and 2.3x (one HBM3 stack) at k = 1, and over the M5 Max from 1.6x to 0.9x, for a cost of 1.5 percent of the Mac's rate and none of the 5090's, 16 W on the Mac and 81 W on the 5090. It is a GO for the lever, not a closed verdict on the chip: the edge stays over 2x against the 5090 at every k under about 0.9, and the honest Apple card already sits under 2x with no shadow at all, so the project's public line should name the honest card's joules (0.78 microjoules on the M5 Max's GPU and memory, 2.34 to 2.65 on the 5090) and the chip core's k as the two numbers, not "under 2x". NO-GO above 130,000 ops (the M5 Max loses more than 5 percent) and NO-GO for a block over 256 instructions (17 percent on the M5 Max at 1,024).