Counter ASIC 3.0 item 8: the RTX 5090 rows, the chip side on the measured 11 pJ per op, the verdict

PC 2 job run-ca3-shadow-pc2-20261006 (card empty, every pack bit-exact against the Mac): the 5090 holds its rate to
150,800 ops per hash and loses 2.7 percent at 199,600 under the app's 431 W cap, which binds from 102,100 ops up and
takes the clock from 3,037 to 1,834 MHz (86 MH/s at 330,700 ops); 350 W at the control, 2.65 to 3.27 microjoules per
hash; marginal ALU energy 10 to 13 pJ per counted op. Clock rows OWED (nvidia-smi refused -lgc without rights). Chip
side at N = 100,000 and k = 1: 2.1x over the 5090 on GDDR7, 0.9x over the M5 Max. GO at mx8+sh256x27.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 08:44:57 +00:00
parent 53382993b2
commit 9417a14ac7
2 changed files with 75 additions and 30 deletions

View file

@ -73,11 +73,36 @@ The headroom under the 10 ms gate that the shadow may spend, per pairing (item 2
So the verifier does not bind the shadow at any N on the ladder in any pairing where the derivation itself fits the gate: 330,700 ops per hash costs 0.56 ms on the M5 Max core and about 1.4 ms on the 2019-class core, which is 12 to 30 percent of the headroom. The binding constraint on N is the honest cards' compute (section 3 for the M5 Max, section 5 for the 5090), not the node. The measured rows: x8 + sh256x88 2.62 ms steady, 2.77 worst cold on the M5 Max core, about 6.5 and 6.9 ms on a 2019-class core by the 2.5x rule; dr736 + sh256x88 about 5.5 ms on the M5 Max core (4.9 + 0.56). The 2019-class core itself is unmeasured (O-1.14); the 2.5x rule stands in for it.
## 5. The RTX 5090 rows (PC 2, 1ccfe586)
## 5. The RTX 5090 rows (PC 2, 1ccfe586, CUDA through NVRTC)
PENDING the PC 2 job `run-ca3-shadow-pc2-20261006` (`tools/ca3-shadow/pc2-shadow-bench.ps1`, published after `/tmp/igneum-devnet/pc2-ca3.clear` under the mkdir lock; the fetch job `fetch-ca3-shadow-20261006` carries the ten packs). The job: `--stop-miners`, the prover off for the run and back on at the end, the card confirmed empty by nvidia-smi's compute-apps list, the installed `igneum-worker-cuda.exe` (NVRTC, the packs' own text) with `--bench --batches 250 --batch-log2 24 --block-warps 1` on every pack, the control first and last, `--block-warps 8` on the control and two rungs, `nvidia-smi -l 1` sampling power.draw, utilization.gpu, clocks.sm, clocks.mem and temperature every second with the per-pack mean over the bench's own window; then the core-clock rows. The rows land here with the closing report.
Job `run-ca3-shadow-pc2-20261006` (`tools/ca3-shadow/pc2-shadow-bench.ps1`; fetch job `fetch-ca3-shadow-20261006`, zip sha256 65a65f92..., extracted with sha256 ok at 08:30:26Z), published at 08:29:54Z after `/tmp/igneum-devnet/pc2-ca3.clear` (08:24:27Z) under the mkdir lock (taken 08:29:22Z, released 08:41:03Z after the closing report was read); the job ran 08:30:26Z to 08:38:57Z, 511 s, exit 0, `--stop-miners`, the prover off for the run and back on at the end (`{"ok":true}` both times). The card was EMPTY before the ladder: `nvidia-smi --query-compute-apps` listed no igneum, sp1 or prove process after 0 s (the app had stopped its miners), so these are the card's own figures, not loaded-card ratios (item 2's job found the app's `POST api/cards` with the state's key switched nothing; this job did not use it). Harness: the installed `igneum-worker-cuda.exe` (0.3.11, sha256 2b3b8c92...) `--bench --pack <dir> --batches 250 --batch-log2 24 --block-warps 1` (`--block-warps 8`, 120 batches, on three packs), NVRTC 12.8, sm_120, driver 13.3, the packs' own kernel text, the vectors through the bound kernel and the 2^24 fingerprint at base 0; power = `nvidia-smi -l 1` (power.draw, utilization.gpu, clocks.sm, clocks.mem, temperature), the mean over each bench's window after its first 12 s and before its last 2 s (17 to 34 samples per row). Card state before: 431 W limit (the app's 75 percent setting; min 400, max 600), 73.8 W idle at 862 MHz, 56 C; after: 85 W, 65 C.
Power caps. The brief asked for `nvidia-smi -pl 200, 250` and the default. The 5 October sweep on this card (`docs/plans/counter-asic-2-status.md` "22:16 the 5090 power-limit sweep", relay #224, 22:09 to 22:15Z on PC 2, corrected in `counter-asic-2-rollout.md` 7b) already carries the cap rows at the current program length:
| Pack (class over mx8) | Shadow instrs per hash | Ops per hash | MH/s (wall, 250 x 2^24) | Against the control | Watts, mean (min to max) | SM MHz, mean | Microjoules per hash | Bit-exact (self-test, 96 of 96 lanes; fingerprint 2^24 = the Mac's) | Registers, blocks per SM | NVRTC ms |
|---|---|---|---|---|---|---|---|---|---|---|
| mx8-genesis (control, first and last) | 0 | 930 | 131.94, 132.47 | | 342.4 (341 to 343), 357.4 (356 to 359) | 3,052, 3,037 | 2.65 (mean 350 W) | yes, 7c28cfb06c5c65a9 | 31, 24 | 158, 160 |
| mx8-genesis, 8 warps per block | 0 | 930 | 131.87 | -0.3% | 344.5 | 3,052 | 2.61 | yes | 31, 6 | 158 |
| sh256x2 | 4,096 | 8,400 | 132.34 | +0.1% | 354.4 (353 to 356) | 3,037 | 2.68 | yes, 33e8bbe4c35b2e54 | 48, 24 | 248 |
| sh256x7 | 14,336 | 27,200 | 132.32 | +0.1% | 385.3 (383 to 388) | 3,037 | 2.91 | yes, 6cfb70911007520a | 48, 24 | 250 |
| sh256x13 | 26,624 | 49,700 | 132.28 | +0.1% | 424.6 (422 to 426) | 3,034 | 3.21 | yes, 59ac286fe2a5a9ef | 48, 24 | 240 |
| sh256x13, 8 warps per block | 26,624 | 49,700 | 131.86 | -0.3% | 426.1 | 3,030 | 3.23 | yes | 48, 5 | 239 |
| sh64x52 (64-instruction block) | 26,624 | 49,700 | 136.77 | +3.5% | 431.5 (the cap; 431 to 433) | 3,024 | 3.15 | yes, 9dd010f79d8ca9f4 | 30, 24 | 181 |
| sh1024x3 (1,024-instruction block) | 24,576 | 45,900 | 132.02 | -0.1% | 428.3 (427 to 429) | 3,030 | 3.24 | yes, a05399c819b79aad | 80, 24 | 510 |
| sh256x27 | 55,296 | 102,100 | 131.95 | -0.2% | 431.0 (the cap) | 2,824 | 3.27 | yes, 3d2e8245cc084d07 | 48, 24 | 242 |
| sh256x40 | 81,920 | 150,800 | 131.75 | -0.3% | 431.0 (the cap) | 2,427 | 3.27 | yes, 0844b706302f1c9c | 48, 24 | 242 |
| sh256x53 | 108,544 | 199,600 | 128.67 | -2.7% | 431.0 (the cap) | 1,753 | 3.35 | yes, 4f824b15cf2b124a | 48, 24 | 241 |
| sh256x88 | 180,224 | 330,700 | 86.39 | -34.7% | 431.0 (the cap) | 1,834 | 4.99 | yes, 0572522e39a94d8a | 48, 24 | 241 |
| sh256x88, 8 warps per block | 180,224 | 330,700 | 85.85 | -35.1% | 431.0 | 1,846 | 5.02 | yes | 48, 5 | 244 |
What the rows say:
1. The 5090 holds its rate to 150,800 ops per hash (-0.3 percent) and loses 2.7 percent at 199,600, under a 431 W power limit that the control never reaches (342 to 357 W, the card warming from 62 to 71 C over the job) and that binds from 102,100 ops upward: the card sits at 431.0 W from sh256x27 on and the SM clock falls (3,037 MHz at the control, 2,824 at 102,100 ops, 2,427 at 150,800, 1,753 at 199,600, 1,834 at 330,700). At 330,700 ops it is compute-bound at the capped clock: 86.39 MH/s x 330,700 = 28.6 T counted op/s at 1,834 MHz, which is the 45.2 T op/s budget scaled by 1,834 / 3,050 (27.2 T). So the brief's cap rows are in effect measured: at a 431 W cap the card does the hash plus 150,000 ops at full rate, and the model's "332,000 before compute binds" holds only with the cap at 575 W (where the card would draw about 575 W for 136 MH/s; the 5 October sweep's 400 W floor rows and this 431 W cap bracket what a capped 5090 pays). The 5 percent point at 431 W: about 210,000 ops per hash (interpolated between 199,600 at -2.7 percent and 330,700 at -34.7 percent).
2. The control reads 132.2 MH/s against 135.9 to 137.7 for this pack on 5 October (`docs/bench-log.md`, the mixer job): 3 percent lower, the same harness and card, the card warmer; every row is read against today's control.
3. The block size matters on the 5090 as on the M5 Max: the 64-instruction block (sh64x52) runs 3.5 percent ABOVE the control at 49,700 ops (30 registers against 48, the same occupancy; the spaced loads queue less), 256 holds, and 1,024 holds the rate too (80 registers, still 24 blocks per SM) but compiles in 510 ms against 240. A class v4 block is 64 to 256 instructions.
4. Watts and the marginal energy per counted op, read against the 350 W control: 4.5 pJ at 8,400 ops (noise: a 4 W step), 10.2 pJ at 27,200, 11.6 at 49,700 (12.2 on sh64x52, 13.2 on sh1024x3); above that the cap holds the watts and the clock gives, so the marginal cannot be read. The 5090's marginal ALU energy on a random 32-lane program at its shipping clock is therefore 10 to 13 pJ per counted op, about twice the 5.5 pJ the model assumed (section 5.7: TGP minus 326 W over the whole 45.2 T budget, a figure that mixes the clock-down the cap forces with the ALU energy). The M5 Max's 6.9 pJ at 100,000 ops (section 3) is 0.6 of the 5090's, so on this program the Apple GPU is the better ALU per joule as well as the better memory system per joule.
5. Energy per hash, whole card: 2.65 microjoules at the control (350 W; the app's 290 W at 124 MH/s on 5 October gave 2.34: the bench pushes the card 20 percent harder than the app's job loop does), 3.21 at 49,700 ops, 3.27 at 102,100 and 150,800 (the cap), 3.35 at 199,600, 4.99 at 330,700.
6. Clock rows: OWED. `nvidia-smi -i 0 -lgc 0,2781` answered "The current user does not have permission to change clocks for GPU 00000000:01:00.0" and the job did not try to gain rights (the brief's rule); `-rgc` in the finally block answered the same and the readback showed the clock unlocked (3,037 MHz, max 3,090). The rows need an elevated job (the sweep-5090.ps1 shape, `--elevated`), which is a decision for the coordinator since an elevated job is the administrator-prompt class on PC 2. What the ladder already shows about clocks: at the 431 W cap the card's own governor took the SM clock to 1,753 to 1,834 MHz at 200,000 to 331,000 ops and the rate fell with it, so a locked clock under 2,400 MHz costs hash rate once the program carries over 150,000 ops, and nothing at the control (the control is latency-bound: 136 MH/s at 3,050 MHz and 115 MH/s in the app at the same clock in the 5 October sweep).
The brief's `-pl` rows: the 5 October sweep on this card (`docs/plans/counter-asic-2-status.md` "22:16 the 5090 power-limit sweep", relay #224, 22:09 to 22:15Z on PC 2, corrected in `counter-asic-2-rollout.md` 7b):
| Cap | Limit W | Draw W | MH/s (the app's API) | MH/W | SM MHz |
|---|---|---|---|---|---|
@ -86,27 +111,28 @@ Power caps. The brief asked for `nvidia-smi -pl 200, 250` and the default. The 5
| 65% | 400 (the floor) | 310.6 | 114.46 | 0.369 | 3,051 |
| 50% | 400 (clamped) | 302.4 | 109.20 | 0.361 | 3,050 |
Reading: the card draws 302 to 316 W under this program whatever the cap, `power.min_limit` is 400 W, so `-pl 200` and `-pl 250` cannot be set on a 5090 and the cap never binds at the hash. Those rows are re-used, not re-run. The lever that can move the watts is the core clock, and the job carries `nvidia-smi -lgc 0,<MHz>` at the Ember plan's four caps (2,781 / 2,472 / 2,163 / 1,854 MHz, `docs/plans/ember-tune.md`), about 75 s each, at the control and at sh256x88, with `-rgc` in a finally block; if nvidia-smi refuses the lock the rows are OWED and the job does not try to gain rights. Ember Tune's own 5090 steps did not run (its second engine never mined; re-run pending the project lead); its idle readbacks on PC 1: 90.6 W idle, 2,505 MHz core, 14,001 MHz memory, limit 450 of 575 W.
`power.min_limit` is 400 W (read again by this job: min 400, max 600), so `-pl 200` and `-pl 250` cannot be set on a 5090; those rows are re-used, not re-run, and the ladder above adds what the cap does once the program carries work. Ember Tune's own 5090 steps did not run (its second engine never mined; re-run pending the project lead); its idle readbacks on PC 1: 90.6 W idle, 2,505 MHz core, 14,001 MHz memory, limit 450 of 575 W.
The item 1 denominator moves: the 5090 at the hash alone is about 290 W (p95 290, max 295 W at 124 MH/s in the app, the 5 October record on branch miner-eff), not 326 W, so 2.34 microjoules per hash, and the `f = 1` rows of section 5.4 read 5.0x on GDDR7 (was 5.1x), 7.3x on one HBM3 stack (was 7.5x), 8.9x on eight (was 9.2x): a 2 to 3 percent move. The job's own watts at the hash replace this when they land.
The item 1 denominator moves: at the hash alone the 5090 is 290 W in the app (p95 290, max 295 W at 124 MH/s, the 5 October record on branch miner-eff: 2.34 microjoules) and 342 to 357 W in this bench (132 MH/s: 2.65 microjoules), not 326 W at 136.1 (2.40). The `f = 1` rows of chip-model-v3.md section 5.4 therefore read 5.0x (app) to 5.7x (bench) on GDDR7 against 5.1x, 7.3x to 8.3x on one HBM3 stack against 7.5x, 8.9x to 10.1x on eight against 9.2x: a move of 2 to 11 percent, inside the model's own margin.
## 6. The chip side, re-evaluated (approximate)
The `f = 1` chip of chip-model-v3.md section 5.4, at the memory's activate ceiling: GDDR7 (16 devices) 166.4 MH/s at 0.466 microjoules per hash, one HBM3 stack 83.6 MH/s at 0.321, eight stacks 0.262; those are memory, static and controller only. With the shadow the chip adds a core that runs the per-epoch random block at `k` times the GPU's marginal energy per op; the model's 5.5 pJ (section 5.7) is kept as the unit, and the measured marginal figures of section 3 (6 to 8 pJ on the M5 Max at the useful N) say it is the right order. Three columns: `k = 1` (a chip core as good as the GPU's ALU on a random program: RandomX's argument and the project lead's test), `k = 1.5` (1.5x worse), and `k = 0.3` (a wide-SIMD fixed-datapath array at N5: the int32 datapath energy of section 5.1, 0.06 pJ per add and 0.52 per multiply at the op mix's 29 percent multiplies is 0.19 pJ, with a 2x pipeline overhead and about 4x for the register file, operand wires and the shared instruction fetch over a wide SIMD row, approximate; the floor a chip maker would claim). Section 5.7's "k = 1.5" column was computed as the chip being 1.5x BETTER (0.83 at N = 100,000 is 0.466 + 0.55 / 1.5); the column below computes it as written, 1.5x worse. Energy per hash = memory + N x 5.5 pJ x k.
The `f = 1` chip of chip-model-v3.md section 5.4, at the memory's activate ceiling: GDDR7 (16 devices) 166.4 MH/s at 0.466 microjoules per hash, one HBM3 stack 83.6 MH/s at 0.321, eight stacks 0.262; those are memory, static and controller only. With the shadow the chip adds a core that runs the per-epoch random block at `k` times a GPU's marginal energy per op. The unit is now measured: 11 pJ per counted op, the 5090's marginal at its shipping clock (10.2 to 13.2 pJ on the three rungs under the cap, section 5); the model's 5.5 pJ (section 5.7, from TGP minus 326 W over the whole budget) is half that and is kept as the `k = 0.5` column, which is also about the M5 Max's 6.9 pJ. Four columns: `k = 1` (a chip core as good as the 5090's ALU on a random program: RandomX's argument and the project lead's test), `k = 1.5` (1.5x worse), `k = 0.5` (as good as the M5 Max's ALU, or the model's old unit), and `k = 0.3` (a wide-SIMD fixed-datapath array at N5: the int32 datapath energy of section 5.1, 0.06 pJ per add and 0.52 per multiply at the op mix's 29 percent multiplies is 0.19 pJ, with a 2x pipeline overhead and about 8x for the register file, operand wires and the shared instruction fetch over a wide SIMD row, approximate; the floor a chip maker would claim). Section 5.7's "k = 1.5" column was computed as the chip being 1.5x BETTER (0.83 at N = 100,000 is 0.466 + 0.55 / 1.5); the columns below compute `k` as written. Energy per hash = memory + N x 11 pJ x k; N in counted ops.
ALU silicon for the core, at the 5090's 45.2 T op/s (what N = 330,700 at 136.1 MH/s needs): 22,600 32-bit lanes at 2.0 GHz. At N5 an int32 multiply-add lane with its register-file slice is about 0.002 mm^2 (approximate: 3,000 to 6,000 gates at about 0.3 square microns per NAND2-equivalent plus registers), so about 45 mm^2 of datapath, 60 to 100 mm^2 with SIMD control and operand networks, $25 to $40 of silicon at the sram-mirror.md yield model ($20,000 wafer, about $0.36 per mm^2 on a 128 mm^2 die; section 5 of that file); at 28 nm the same lanes are about 8x the area, 360 mm^2 of datapath and 500 mm^2 or more with control, a reticle-class die that a $5M to $30M controller project does not carry. Watts for the core: at `k = 1` 5.5 pJ x 45.2 T = 249 W (the GPU's own marginal power), at `k = 0.3` 75 W, at `k = 1.5` 373 W. So at N = 330,700 the chip is a memory controller plus a GPU-class ALU array at 250 W on an advanced node, which is the brief's test in numbers; at N = 100,000 the array is a third of that (14,000 lanes, about 30 mm^2 at N5, 75 W at `k = 1`).
ALU silicon for the core, at the 5090's 45.2 T op/s (what N = 330,700 at 136.1 MH/s needs): 22,600 32-bit lanes at 2.0 GHz. At N5 an int32 multiply-add lane with its register-file slice is about 0.002 mm^2 (approximate: 3,000 to 6,000 gates at about 0.3 square microns per NAND2-equivalent plus registers), so about 45 mm^2 of datapath, 60 to 100 mm^2 with SIMD control and operand networks, $25 to $40 of silicon at the sram-mirror.md yield model ($20,000 wafer, about $0.36 per mm^2 on a 128 mm^2 die; section 5 of that file); at 28 nm the same lanes are about 8x the area, 360 mm^2 of datapath and 500 mm^2 or more with control, a reticle-class die that a $5M to $30M controller project does not carry. Watts for the core at N = 330,700 and 136 MH/s: at `k = 1` 11 pJ x 45.2 T = 497 W (more than the whole 5090 draws for the same work, which is what `k = 1` means), at `k = 0.5` 249 W, at `k = 0.3` 149 W. At N = 100,000 the array is a third of that: 14,000 lanes, about 30 mm^2 at N5, 150 W at `k = 1`, 45 W at `k = 0.3`.
Gain per joule against the honest cards, chip energy = memory + N x 5.5 pJ x k, N in counted ops:
Gain per joule against the honest cards, measured watts on both (the M5 Max GPU plus DRAM, the 5090 whole card under its 431 W cap):
| N ops per hash | Card microjoules: M5 Max (GPU + DRAM) / 5090 | Chip, GDDR7, k = 1 / 1.5 / 0.3 | Gain per joule, GDDR7, against the M5 Max, k = 1 / 1.5 / 0.3 | Against the 5090 (section 5.7's linear watts until the job's rows land) | Chip, one HBM3 stack, k = 1 / 1.5 / 0.3 | Gain, HBM3, against the M5 Max | Against the 5090 (model) |
| N ops per hash | Card microjoules: M5 Max / 5090 | Chip, GDDR7, k = 1 / 1.5 / 0.5 / 0.3 | Gain, GDDR7, against the M5 Max | Against the 5090 | Chip, one HBM3 stack, k = 1 / 1.5 / 0.5 / 0.3 | Gain, HBM3, against the M5 Max | Against the 5090 |
|---|---|---|---|---|---|---|---|
| 930 (today) | 0.78 / 2.34 (measured, 290 W) | 0.47 | 1.7x | 5.0x | 0.33 | 2.4x | 7.3x |
| 49,700 (sh256x13) | 1.16 / about 2.45 (333 W, model) | 0.74 / 0.88 / 0.55 | 1.6x / 1.3x / 2.1x | 3.3x / 2.8x / 4.5x | 0.59 / 0.73 / 0.40 | 2.0x / 1.6x / 2.9x | 4.2x / 3.4x / 6.1x |
| 102,100 (sh256x27) | 1.40 / about 2.78 (378 W, model) | 1.03 / 1.31 / 0.63 | 1.4x / 1.1x / 2.2x | 2.7x / 2.1x / 4.4x | 0.88 / 1.16 / 0.49 | 1.6x / 1.2x / 2.9x | 3.2x / 2.4x / 5.7x |
| 199,600 (sh256x53; the M5 Max is 10 percent down here) | 1.58 / about 3.40 (462 W, model) | 1.56 / 2.11 / 0.80 | 1.0x / 0.75x / 2.0x | 2.2x / 1.6x / 4.3x | 1.42 / 1.97 / 0.65 | 1.1x / 0.80x / 2.4x | 2.4x / 1.7x / 5.2x |
| 330,700 (sh256x88; the M5 Max is 21 percent down) | 1.88 / about 4.22 (575 W, model) | 2.29 / 3.19 / 1.01 | 0.82x / 0.59x / 1.9x | 1.8x / 1.3x / 4.2x | 2.14 / 3.05 / 0.87 | 0.88x / 0.61x / 2.2x | 2.0x / 1.4x / 4.9x |
| 930 (today) | 0.78 / 2.65 (2.34 in the app) | 0.48 / 0.48 / 0.47 / 0.47 | 1.6x | 5.6x (5.0x on the app's watts) | 0.33 | 2.3x | 8.0x (7.3x) |
| 49,700 (sh256x13) | 1.16 / 3.21 | 1.01 / 1.29 / 0.74 / 0.63 | 1.15x / 0.90x / 1.6x / 1.9x | 3.2x / 2.5x / 4.3x / 5.1x | 0.87 / 1.14 / 0.59 / 0.48 | 1.3x / 1.0x / 2.0x / 2.4x | 3.7x / 2.8x / 5.4x / 6.6x |
| 102,100 (sh256x27, the candidate) | 1.39 / 3.27 | 1.59 / 2.15 / 1.03 / 0.80 | 0.88x / 0.65x / 1.4x / 1.7x | 2.1x / 1.5x / 3.2x / 4.1x | 1.44 / 2.01 / 0.88 / 0.66 | 0.97x / 0.70x / 1.6x / 2.1x | 2.3x / 1.6x / 3.7x / 5.0x |
| 150,800 (sh256x40; the M5 Max is 7 percent down) | 1.46 / 3.27 | 2.13 / 2.95 / 1.30 / 0.96 | 0.69x / 0.49x / 1.1x / 1.5x | 1.5x / 1.1x / 2.5x / 3.4x | 1.98 / 2.81 / 1.15 / 0.82 | 0.74x / 0.52x / 1.3x / 1.8x | 1.7x / 1.2x / 2.8x / 4.0x |
| 199,600 (sh256x53; the M5 Max 10 percent down, the 5090 2.7) | 1.58 / 3.35 | 2.66 / 3.76 / 1.56 / 1.12 | 0.59x / 0.42x / 1.0x / 1.4x | 1.3x / 0.89x / 2.1x / 3.0x | 2.52 / 3.61 / 1.42 / 0.98 | 0.63x / 0.44x / 1.1x / 1.6x | 1.3x / 0.93x / 2.4x / 3.4x |
| 330,700 (sh256x88; both compute-bound) | 1.87 / 4.99 | 4.10 / 5.92 / 2.29 / 1.56 | 0.46x / 0.32x / 0.82x / 1.2x | 1.2x / 0.84x / 2.2x / 3.2x | 3.96 / 5.78 / 2.14 / 1.41 | 0.47x / 0.32x / 0.88x / 1.3x | 1.3x / 0.86x / 2.3x / 3.5x |
Reading: the lever works exactly as far as `k` is near 1. At N = 100,000 the chip's edge over the M5 Max is 1.4x at `k = 1` and 1.1x at `k = 1.5`, and over the 5090 (model watts) 2.7x and 2.1x; at `k = 0.3` the edge barely moves at any N (2.2x and 4.4x), because a fixed-datapath array runs the block for a third of the card's marginal energy and the memory row stays the same. The number that decides the verdict is therefore `k`: the energy per op of a chip core that executes a per-epoch random integer program with a 32-lane shuffle, divided by a GPU's marginal 5.5 to 7 pJ. Measured on the cards: the M5 Max pays 6.9 pJ per counted op at 100,000 and 2.9 at its compute-bound end; the 5090's figure is the job's. RandomX's argument (and the precedent: no RandomX chip has beaten a CPU per joule) is that a general core cannot reach `k` under about 0.5 on a random program; nothing in this file measures a chip, so the `k = 0.3` column is the attacker's claim and the `k = 1` column ours.
Reading: at N = 100,000 ops per hash the `f = 1` chip's edge over the 5090 falls from 5.6x to 2.1x on GDDR7 (2.3x on one HBM3 stack) when its core costs what the 5090's ALU costs per op, and to 1.5x (1.6x) when it costs 1.5x more; over the M5 Max it falls from 1.6x to under 1x at `k = 1`. The 2x line against the 5090 is crossed only if the chip's core beats the 5090's ALU by 2x per op (`k = 0.5`: 3.2x, the model's old unit) or more (`k = 0.3`: 4.1x), and against the M5 Max only at `k` under about 0.4. The number that decides the verdict is therefore `k`: the energy per counted op of a chip core that executes a per-epoch random integer program with a 32-lane shuffle, divided by the 5090's measured 11 pJ. Measured on the cards: the 5090 pays 10 to 13 pJ, the M5 Max 6.9 pJ at 100,000 ops and 2.7 to 2.9 pJ at their compute-bound ends (where the clock has dropped and every op is useful). RandomX's argument (and the precedent: no RandomX chip has beaten a CPU per joule) is that a general core cannot reach `k` under about 0.5 on a random program; nothing in this file measures a chip, so the `k = 0.3` column is the attacker's claim and the `k = 1` column ours. The N5 datapath floor (0.19 pJ per op before overheads, section 5.1) says a chip CAN in principle reach `k = 0.3` on a fixed-datapath array, so the shadow lowers the chip's edge, it does not remove it; what removes it is the honest card's own watts (the M5 Max at 0.78 microjoules is already inside 2x of the GDDR7 chip with no shadow at all).
## 7. The gates a class v4 candidate would need
@ -114,7 +140,7 @@ The six gates of `docs/plans/counter-asic-2-rollout.md` section 7 plus the verif
| Gate | State on the shadow class | What closes it |
|---|---|---|
| G1 bit-exact on all three vendors against the Mac reference | Metal GREEN on ten packs (3 of 3 vectors standalone and in batch, cache and dataset PASS, fingerprints above); the CUDA text GREEN in the clang emulation on the Mac (`proto-cuda/emu/emu.sh` on sh256x2: cache, dataset and the 3 vector warps standalone and in batch PASS, `with-lock.sh run`, 08:10 UTC) and PENDING on the card (the PC 2 job's self-test runs the vectors through the bound kernel and prints the 2^24 fingerprint); the OpenCL text GREEN on Apple OpenCL (`proto-opencl/igneum-bench-cl --bench-pack` on sh256x2 and sh256x88: 96 of 96 vector lanes, self-test PASS); AMD silicon OWED (PC 1) | the PC 2 closing report; a PC 1 job on the 9070 XT (the OpenCL kernels are emitted) |
| G1 bit-exact on all three vendors against the Mac reference | Metal GREEN on ten packs (3 of 3 vectors standalone and in batch, cache and dataset PASS, fingerprints above); the CUDA text GREEN in the clang emulation on the Mac (`proto-cuda/emu/emu.sh` on sh256x2: cache, dataset and the 3 vector warps standalone and in batch PASS, `with-lock.sh run`, 08:10 UTC) and GREEN on the card (the PC 2 job: self-test PASS on all ten packs, 96 of 96 vector lanes through the bound kernel, every 2^24 fingerprint equal to the Mac's, section 5); the OpenCL text GREEN on Apple OpenCL (`proto-opencl/igneum-bench-cl --bench-pack` on sh256x2 and sh256x88: 96 of 96 vector lanes, self-test PASS); AMD silicon OWED (PC 1) | the PC 2 closing report; a PC 1 job on the 9070 XT (the OpenCL kernels are emitted) |
| G2 the CPU verifier exact on 1,000 random hashes per card | not run: the evidence is the 96 vectors and the 2^24 fingerprint per pack on Metal | a serve-mode job per card as in 2.0 |
| G3 the generator soundness suite green | `cargo test` 54 + 4 + 19 + 7 green with the new test; the Metal fuzz, edge, stats and determinism runs were NOT run on the shadow class | those four runs at the chosen S and R |
| G4 the fast-time 3-node network across an activation | not run (no node change exists: the class is a `LoadClass` field, and the chain's `V3_CLASS` does not set it) | the node seam for v4 (`program_class_v4_activation_daa`) and the run |
@ -125,28 +151,32 @@ The six gates of `docs/plans/counter-asic-2-rollout.md` section 7 plus the verif
## 8. Consequences per tier
A longer program costs the card watts and nothing else while the card stays latency-bound; past its bind point it costs hash rate. Income per pound of card is unchanged while the rate holds; income per watt falls by the watts ratio. The numbers are this file's; the 5090's are the model's until the job lands, and the AMD row is the budget's.
A longer program costs the card watts and nothing else while the card stays latency-bound; past its bind point, or once its power cap binds, it costs hash rate. Income per pound of card is unchanged while the rate holds; income per watt falls by the watts ratio. The numbers are this file's (sections 3 and 5); the AMD row is the budget's.
| Tier | At N = 100,000 ops per hash (sh256x27, the candidate) | At N = 200,000 | At N = 331,000 (the 5090's full budget) | What is being done |
| Tier | At N = 100,000 ops per hash (sh256x27, the candidate) | At N = 200,000 | At N = 331,000 (the 5090's full budget at 575 W) | What is being done |
|---|---|---|---|---|
| Apple user (M-series laptop or desktop; M5 Max measured) | rate -1.5%, GPU + DRAM 21 to 37 W (+16 W; a laptop on battery feels it), income per watt 0.56x, per pound 0.99x | rate -10%, 38 W; per watt 0.49x, per pound 0.90x | rate -21%, 40 W; per watt 0.41x, per pound 0.79x | the project's N is capped by this card: no more than 130,000 ops (the 5 percent rule), 100,000 recommended |
| NVIDIA, one 24 or 32 GB card (5090 class) | model: rate holds, about 290 to 378 W (+30%); per watt 0.77x, per pound 1.0x | about 462 W; per watt 0.63x | compute-bound at 136 MH/s, 575 W; per watt 0.50x | the PC 2 job measures the rate, the watts and the clock caps; the per-watt loss is the price of the lever and is reported, not hidden |
| NVIDIA, one 8, 12 or 16 GB card | the same shape per card: the ALU budget scales with the SM count, so a 4070-class card (about 30 T op/s, approximate) binds near 200,000 ops at its rate; at 100,000 it holds | a 4060-class card (about 15 T op/s) binds near 100,000 | | no card we own is in this tier; the ladder runs on any CUDA card through the same job |
| AMD, one 16 GB card (RX 9070 XT) | OWED: the budget (about 650,000 ops, approximate, section 5.7) says it holds at 100,000 and at 200,000; its watts at the hash are unmeasured | | | the PC 1 job when the desk is free; the OpenCL kernels are in every pack |
| A rig | per card as above; a rig's bill is watts, so at 100,000 ops a 5090 rig pays about 30 percent more electricity for the same hash (model) | | | the recommended N is the smallest that moves the chip row; the rig pays it only if class v4 is adopted |
| Apple user (M-series laptop or desktop; M5 Max measured) | rate -1.5%, GPU + DRAM 21 to 37 W (+16 W; a laptop on battery feels it), income per watt 0.56x, per pound 0.99x | rate -10%, 38 W; per watt 0.49x, per pound 0.90x | rate -21%, 40 W; per watt 0.41x, per pound 0.79x | the project's N is capped by this card: no more than 130,000 ops (the 5 percent rule), 100,000 recommended; the block is 64 to 256 instructions |
| NVIDIA, one 24 or 32 GB card (5090 measured, 431 W cap) | rate -0.2%, 350 to 431 W (the cap; +23%), per watt 0.81x, per pound 1.0x | rate -2.7% at the cap, 431 W; per watt 0.79x | rate -35% at the cap (86 MH/s at 1,834 MHz); at a 575 W cap the model says 136 MH/s at about 575 W, per watt 0.61x | measured; the clock rows are OWED (an elevated job); a 5090 owner on the app's 75 percent cap loses nothing at 100,000 and 2.7 percent at 200,000 |
| NVIDIA, one 8, 12 or 16 GB card | the same shape per card: the ALU budget scales with the SM count and the cap with the board, so a 4070-class card (about 30 T op/s, approximate) binds near 200,000 ops at its rate and holds at 100,000 | a 4060-class card (about 15 T op/s) binds near 100,000 and is the first NVIDIA tier to pay in rate | | no card we own is in this tier; the ladder runs on any CUDA card through the same job, which is the next NVIDIA measurement |
| AMD, one 16 GB card (RX 9070 XT) | OWED: the budget (about 650,000 ops, approximate, section 5.7) says it holds at 100,000 and at 200,000; its watts at the hash are unmeasured | | | the PC 1 job when the desk is free; the OpenCL kernels are in every pack and pass on Apple OpenCL |
| A rig | per card as above; a rig's bill is watts, so at 100,000 ops a 5090 rig pays about 23 percent more electricity for the same hash (its cards at their cap), an Apple rig 76 percent more on the GPU and memory channels | | | the recommended N is the smallest that halves the chip row; the rig pays it only if class v4 is adopted |
| A pool user | nothing changes in shares or payout: the hash rate holds at the recommended N on every card measured | | | |
| A chip | must add a 14,000-lane ALU array at N = 100,000 (about 30 mm^2 at N5, 75 W at `k = 1`) and a 22,600-lane array at 331,000 (60 to 100 mm^2, 250 W); its per-joule edge over the 5090 falls from 5.0x to 2.7x at 100,000 and `k = 1` (model watts), over the M5 Max from 1.7x to 1.4x | | | the `k` question goes to the external cryptanalysis and chip review (item 3): can a core run a random 32-lane program under half a GPU's marginal energy per op |
| The public claim | the honest M5 Max is 1.7x from the `f = 1` GDDR7 chip per joule at N = 0 by these channels, the 5090 5.0x; at N = 100,000 and `k = 1` they read 1.4x and 2.7x (model) | | | nothing on the site changes from this file |
| A chip | must add a 14,000-lane ALU array at N = 100,000 (about 30 mm^2 at N5, 150 W at `k = 1`) and a 22,600-lane array at 331,000 (60 to 100 mm^2, 500 W); its per-joule edge over the 5090 falls from 5.6x to 2.1x at 100,000 and `k = 1`, over the M5 Max from 1.6x to 0.9x | | | the `k` question goes to the external cryptanalysis and chip review (item 3): can a core run a random 32-lane program under half the 5090's 11 pJ per op |
| The public claim | at N = 0 the honest M5 Max is 1.6x from the `f = 1` GDDR7 chip per joule by these channels and the 5090 5.0x to 5.6x; at N = 100,000 and `k = 1` they read 0.9x and 2.1x | | | nothing on the site changes from this file |
The N the project should pick, by the 2.0 rule (no card we own loses more than 5 percent): 100,000 ops per hash, `S = 256`, `R = 27` (55,296 shadow instructions per hash), subject to the 5090 row holding within 5 percent in the job and the 9070 XT row when PC 1 is free; the M5 Max's ceiling is 130,000. The block size is 64 to 256 instructions, never 1,024.
The N the project should pick, by the 2.0 rule (no card we own loses more than 5 percent): 100,000 ops per hash, `S = 256`, `R = 27` (55,296 shadow instructions per hash): the M5 Max loses 1.5 percent, the 5090 0.2 percent under its 431 W cap, the 9070 XT holds by its budget (owed). The M5 Max's ceiling is about 130,000, the 5090's at 431 W about 210,000. The block size is 64 to 256 instructions, never 1,024.
## 9. Unverified and owed
- The RTX 5090 rows (hash rate, watts, clock caps, the CUDA fingerprints): the PC 2 job, section 5.
- The RTX 5090 clock rows (`-lgc` at the four Ember caps): OWED, nvidia-smi refused the lock without administrator rights and the job did not ask for them; an elevated job (the sweep-5090.ps1 shape) is the coordinator's decision. The 5090 rows were taken under the app's 431 W cap, which binds from 102,100 ops; the rows at 575 W are the model's.
- The RX 9070 XT rows: OWED (PC 1 not released today).
- The Mac watts are the IOReport GPU and DRAM channels (no root, the private IOReport library through dlopen), not powermetrics and not the package; Ember Tune's 38 W package figure is approximate and from another session.
- The ops count (1.83 per ALU instruction) is a convention counted from the emitted statements; the vendors' real cost per op differs (item 6's step ratios), which is why the rows carry instructions and watts, not ops alone.
- The chip side is arithmetic: the 5.5 pJ unit, the `k` columns, the lane area and the 28 nm scaling are approximate; no chip was measured, and the `k = 0.3` floor is an estimate of what a fixed-datapath array could claim.
- The chip side is arithmetic: the 11 pJ unit is the 5090's measured marginal on three rungs (10.2 to 13.2), the `k` columns, the lane area and the 28 nm scaling are approximate; no chip was measured, and the `k = 0.3` floor is an estimate of what a fixed-datapath array could claim.
- The acceptance rule does not run the block; whether a class v4 should make it is a design decision with the cost stated in section 2.
- The 2019-class core is unmeasured (O-1.14); the 2.5x rule stands in.
- The Metal fuzz, edge, stats and determinism runs were not made on the shadow class (gate G3 is the crate suite only).
## 10. Verdict
GO as a class v4 candidate at N = 100,000 ops per hash (`mx8+sh256x27`, a 256-instruction block at 27 passes; 64 to 256 instructions per block), subject to the 9070 XT row and the gates of section 7. The number that decides it: at N = 100,000 the `f = 1` chip's per-joule edge over the 5090 falls from 5.6x to 2.1x (GDDR7) and 2.3x (one HBM3 stack) at `k = 1`, and over the M5 Max from 1.6x to 0.9x, for a cost of 1.5 percent of the Mac's rate and none of the 5090's, 16 W on the Mac and 81 W on the 5090. It is a GO for the lever, not a closed verdict on the chip: the edge stays over 2x against the 5090 at every `k` under about 0.9, and the honest Apple card already sits under 2x with no shadow at all, so the project's public line should name the honest card's joules (0.78 microjoules on the M5 Max's GPU and memory, 2.34 to 2.65 on the 5090) and the chip core's `k` as the two numbers, not "under 2x". NO-GO above 130,000 ops (the M5 Max loses more than 5 percent) and NO-GO for a block over 256 instructions (17 percent on the M5 Max at 1,024).

View file

@ -1983,8 +1983,23 @@ Branch `ca3-shadow`, worker "shadow" (`docs/analysis/latency-shadow-2026-10-06.m
Reading: the M5 Max stays latency-bound to about 100,000 ops per hash and its 5 percent point is about 130,000 (between the 102,100 and 150,800 rungs), 2.2x under the chip model's 290,000 (from memory); the block size matters on Apple (64 instructions +2.5 percent, 256 holds, 1,024 costs 17 percent at the same N: the instruction footprint); the GPU rises from 11 to 27 W at 100,000 ops (marginal 6.9 pJ per counted op; 2.9 pJ at the compute-bound end) and the energy per hash from 0.78 to 1.40 microjoules, GPU plus DRAM. The verifier's law on this core: 2.06 ms + 3.2 microseconds per 1,000 shadow instructions per warp (0.1 ns per lane-instruction), so 330,700 ops cost 0.56 ms here and about 1.4 ms on a 2019-class core by the 2.5x rule: inside every pairing's headroom (x8 7.8 / 4.6 ms, dr368 7.1 / 2.8, dr736 5.1 / none, M5 Max / 2019-class, item 2's figures), so the node never binds the shadow before the cards do. The control is 2.5 percent under the 5 October figure for this pack (27.7, mixer-x4.md 6.2) on three runs; every row is read against today's 27.08.
**RTX 5090 (PC 2, 1ccfe586), CUDA**: PENDING the PC 2 job `run-ca3-shadow-pc2-20261006` (`tools/ca3-shadow/pc2-shadow-bench.ps1`, after `/tmp/igneum-devnet/pc2-ca3.clear` under the mkdir lock): the same ladder through the installed worker's `--bench` with `nvidia-smi -l 1` sampling, and the core-clock rows at 2,781 / 2,472 / 2,163 / 1,854 MHz (`-lgc`, `-rgc` in a finally; OWED if the lock needs rights the job does not have). The power-cap rows are the 5 October sweep's (counter-asic-2-status.md 22:16: 302 to 316 W at every cap, floor 400 W, so `-pl 200` and `250` cannot be set and the cap never binds at the hash); the card at the hash alone is about 290 W (miner-eff record), 2.34 microjoules per hash at 124 MH/s in the app.
**RTX 5090 (PC 2, 1ccfe586), CUDA through NVRTC** (job `run-ca3-shadow-pc2-20261006`, 08:30:26Z to 08:38:57Z, 511 s, exit 0, `--stop-miners`, the prover off for the run and back on, the card EMPTY before the ladder by nvidia-smi's compute-apps list; the installed worker 0.3.11 `--bench --batches 250 --batch-log2 24 --block-warps 1`, wall time; power = `nvidia-smi -l 1` means over each bench's window; the card under the app's 431 W limit, 73.8 W idle, 56 to 71 C):
| Pack | Shadow instrs per hash | Ops per hash | MH/s | Against the control | Watts, mean | SM MHz | Microjoules per hash | Bit-exact, fingerprint 2^24 = the Mac's | NVRTC ms |
|---|---|---|---|---|---|---|---|---|---|
| mx8-genesis (control, first and last) | 0 | 930 | 131.94, 132.47 | | 342.4, 357.4 | 3,052, 3,037 | 2.65 | yes, 7c28cfb06c5c65a9 | 158, 160 |
| sh256x2 | 4,096 | 8,400 | 132.34 | +0.1% | 354.4 | 3,037 | 2.68 | yes | 248 |
| sh256x7 | 14,336 | 27,200 | 132.32 | +0.1% | 385.3 | 3,037 | 2.91 | yes | 250 |
| sh256x13 | 26,624 | 49,700 | 132.28 | +0.1% | 424.6 | 3,034 | 3.21 | yes | 240 |
| sh64x52 (64-instruction block) | 26,624 | 49,700 | 136.77 | +3.5% | 431.5 (the cap) | 3,024 | 3.15 | yes | 181 |
| sh1024x3 (1,024-instruction block) | 24,576 | 45,900 | 132.02 | -0.1% | 428.3 | 3,030 | 3.24 | yes | 510 |
| sh256x27 | 55,296 | 102,100 | 131.95 | -0.2% | 431.0 (the cap) | 2,824 | 3.27 | yes | 242 |
| sh256x40 | 81,920 | 150,800 | 131.75 | -0.3% | 431.0 | 2,427 | 3.27 | yes | 242 |
| sh256x53 | 108,544 | 199,600 | 128.67 | -2.7% | 431.0 | 1,753 | 3.35 | yes | 241 |
| sh256x88 | 180,224 | 330,700 | 86.39 | -34.7% | 431.0 | 1,834 | 4.99 | yes | 241 |
Reading: the 5090 holds to 150,800 ops (-0.3 percent) and loses 2.7 percent at 199,600, under a 431 W cap that the control never reaches (342 to 357 W) and that binds from 102,100 ops up: the SM clock falls from 3,037 to 1,834 MHz and at 330,700 ops the card is compute-bound at the capped clock (28.6 T counted op/s, the 45.2 T budget scaled by the clock). The 5 percent point at 431 W is about 210,000 ops. The marginal ALU energy at the shipping clock, read on the three rungs under the cap: 10.2 to 13.2 pJ per counted op, twice the 5.5 pJ the chip model assumed. The 64-instruction block runs 3.5 percent above the control here too. Clock rows (`-lgc`): OWED, nvidia-smi refused the lock without administrator rights and the job did not ask for them. Power-cap rows: the 5 October sweep's (floor 400 W, so `-pl 200` and `250` cannot be set; the cap never binds at the control).
**RX 9070 XT (PC 1, ae432dc7)**: OWED (PC 1 is the project lead's desk and not released today); the OpenCL kernels are in every pack.
Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (`sh256x27`) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 by the model holds its rate at about 378 W (per watt 0.77x; the job measures it), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the `f = 1` chip's edge per joule falls from 1.7x to 1.4x against the M5 Max and from 5.0x to 2.7x against the 5090 (model watts) at `k = 1`, where `k` is the chip core's energy per op over the GPU's marginal 5.5 pJ: the number that decides the item. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000.
Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (`sh256x27`) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 holds its rate at its 431 W cap (350 W at the control: per watt 0.81x, measured), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the `f = 1` chip's edge per joule falls from 1.6x to 0.9x against the M5 Max and from 5.6x to 2.1x against the 5090 at `k = 1`, where `k` is the chip core's energy per op over the 5090's measured 11 pJ: the number that decides the item. Verdict: GO as a class v4 candidate at N = 100,000 (`mx8+sh256x27`), subject to the 9070 XT row and the gates; NO-GO above 130,000 or with a block over 256 instructions. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000.