Counter ASIC 4.0 research: the identity that forbids a zero-premium chip lever, the ranked designs, and the SM-sparse worker variant

docs/analysis/counter-asic-4-research.md (excluded from the public export): edge = (E_card + F) / (E_mem + k F); at zero premium the stored-dataset chip keeps E_card / E_mem, 3.6x on a 5090 locked at 1,400 MHz (measured card, modelled chip), about 2x at that card's idle-plus-memory bound, 1.7x on the M5 Max; class v5 at zero shadow leaves it at 5.1x to 9.1x; the shadow buys 2.1x at k = 1 for the measured 88 W (81.8 W at the knee) and turns against us below 1.8 pJ per op; per-card classes, the refresh as a cost, proof of useful work, memory shaping and the tensor block priced out with their arithmetic; ten designs ranked with the chip edge, the 5090 and 4070 premiums, the verifier cost, agent hours and what breaks each; the recommendation in five sentences; consequences per tier.

proto-cuda/nvrtc/worker.cpp: opt-in SM-sparse variants sp<N> and sp<N>-w<W> (a persistent grid of N blocks of W warps over the dispatch's nonces; the bound kernel rewritten at compile time with exact anchors into a unit function plus a wrapper of the kernel's name; made on demand by name, never in the default race; refused on variant-5 packs and with minBlocks). Gate: mingw cross-compile on igneum-build-1 under lease pool (class measure, label "ca4 research"), exit 0 with -Wall -Wextra, 22:48 UTC; exe sha256 ee8d0e70dd101f125f42c0c7cf07481a794ee18a1317acf68561b37ea18d72be; unrun on a card (the hash lane's PC 1 job run-ca4-pc1-ca4sparse-5090-20261007 is its first run).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 20:55:27 +00:00
parent adfc46c555
commit 5abb666aea
3 changed files with 304 additions and 2 deletions

View file

@ -0,0 +1,221 @@
# Counter ASIC 4.0 research: can the chip's per-joule edge be held near 2x without the GPU paying an energy premium?
7 October 2026, 20:27 to 23:2x UTC (21:27 to 00:2x UK), branch `counter-asic-4`, the Counter ASIC 4.0 research lane. Written for the founder's question of this evening: class v4 (the latency shadow, about 100,000 counted integer operations per hash hidden in the memory wait) brings the strongest chip's per-joule edge over an RTX 5090 from 5x to 9x down to 2.1x to 3.9x, and costs the 5090 145 W at its stock clock and 88 W at a 1,400 MHz core lock for the same hash rate. The brief: find the design that keeps the chip's edge at or under about 2x per joule WITHOUT the GPU paying an energy premium, or prove that no such design exists and say what the floor is.
Every number carries a label: **measured** (a card on a named night, the file named), **modelled** (arithmetic on the chip model's cited memory figures, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling estimate). Nothing in this file changes a consensus object, a pack or a parameter; the one code change (section 12) is a miner-side kernel variant that is opt-in by name and never raced by default.
## 0. One page
**The answer: no hash-side design reaches a chip edge of about 2x at zero premium, and the reason is an identity, not a missing idea.** Against the chip that matters (the stored-dataset chip, "a GPU's memory system without the GPU", chip-model-v3 section 5.5) the per-joule edge is
edge = (E_card + F) / (E_mem + k F)
where `E_card` is the honest card's whole-card energy per hash on class v3, `F` the premium per hash the card pays for forced work, `E_mem` the chip's memory-system energy per hash (0.466 microjoules on GDDR7, 0.321 on one HBM3 stack, modelled), and `k` the chip core's energy per forced operation over the card's. A lever with `F = 0` is a lever with `k F = 0`: every joule the chip must spend on forced work is a joule the card spends first. So at zero premium the edge is `E_card / E_mem`, which no hash change touches, and the ONLY things that move it are the card's own watts and the chip's memory choice.
The floor, on tonight's measured rows (the hash lane's efficiency pass, PC 2, 7 October 2026, `docs/plans/counter-asic-3-status.md` on branch ca3-coord, section "the efficiency pass"):
| Honest card, class v3 (no shadow) | Energy per hash | Chip edge at zero premium, GDDR7 / one HBM3 stack / eight stacks | Label |
|---|---|---|---|
| RTX 5090, core unlocked (136.59 MH/s at 330.2 W) | 2.42 microjoules | 5.2x / 7.5x / 9.2x | measured card, modelled chip |
| RTX 5090, app stock point, 5 October (124 MH/s at 290 W) | 2.34 | 5.0x / 7.3x / 8.9x | measured, modelled |
| **RTX 5090, 1,400 MHz core lock (134.68 MH/s at 228.0 W)** | **1.69** | **3.6x / 5.3x / 6.5x** | measured, modelled |
| RTX 5090, the knee (the hash lane's second pass, 22:1x UTC): class v3 best at 1,300 MHz, 134.62 MH/s at 223.3 W; class v4 best at 1,200 MHz, 133.80 MH/s at 305.1 W; the premium at the best points 81.8 W; the floor below 1,100 MHz unmeasured (a third pass runs) | 1.66 (v3); 2.28 (v4) | 3.6x / 5.2x / 6.3x on v3; class v4 at k = 1 (6.1 pJ per counted op) 2.1x | measured, modelled |
| RTX 5090, the bound any operating point can reach: idle 74 W (PC 2) plus the memory system's 55 W at 17.5 G reads/s, 135 MH/s | 0.96 | 2.1x / 3.0x / 3.7x | approximate (idle measured; the memory share modelled) |
| RTX 4070 at its Ember tune point (30.95 MH/s at 79.5 W) | 2.57 | 5.5x / 8.0x / 9.8x | measured, modelled |
| RX 9070 XT (about 18.6 MH/s at about 304 W) | 10.6 | 23x / 33x / 40x | approximate card figure |
| Apple M5 Max, GPU plus DRAM channels (27.08 MH/s at 21.0 W) | 0.78 | 1.7x / 2.4x / 3.0x | measured, modelled |
So the premium-free floor against the GDDR7 chip is 3.6x on a 5090 locked at 1,400 MHz, about 2x at the bound set by that card's own idle and memory power, and 1.7x on an Apple GPU today with no shadow at all. The 2x line at zero premium is a property of the honest card's idle, not of the hash. The shadow buys the rest only with a premium, and only if the chip's core is no better than about half a GPU lane per op: at the 1,400 lock the measured premium (88.3 W, 0.654 microjoules, 6.5 pJ per counted op) gives 2.1x at `k = 1`, 2.9x at a core of 3.3 pJ per op, 4.1x at 1 pJ, and below about 1.8 pJ per op (`k` under 0.28) the shadow turns against us, because it then dilutes the card more than the chip.
**The ranked designs (section 9 has every column):**
1. **The operating point as the shipped default** (miner-only; core lock at the knee plus undervolt where the driver allows; AMD through ADLX or nothing; Apple has no lever). Premium: none by construction, it lowers `E_card`. Measured: v3 2.42 to 1.69 microjoules, the v4 premium 145 to 88 W, the class v4 edge at `k = 1` 2.1x. 2 to 4 agent hours (already ordered into Ember Tune for 0.3.24). Breaks: a locked card's free shadow shrinks (about 150,000 ops at 1,400 MHz), so ladder rung 2 is not free on it.
2. **The SM-sparse miner kernel** (miner-only; the hash on a fraction of the SMs with 32 warps per block and the rest clock-gated). Reads the unmeasured 100 W between the 5090's idle-plus-memory (about 130 W) and its 228 W at the lock; if half is SM activity the premium-free edge falls toward 2.5x. Built tonight as worker variants `sp<N>-w<W>` (section 12), one PC 1 job owed. Breaks if the 100 W is clock tree and leakage.
3. **The shadow kept at rung 0, its op mix re-weighted toward the families where a chip gains least over a GPU lane** (shuffle and multiply heavy). Same premium, the pessimistic-core edge 3.4x to about 2.9x, `k = 1` unchanged at 2.1x. 4 to 6 agent hours plus the six gates (a class change under the 95 percent rule). Held until the knee and the SM-sparse rows are read (the coordinator, 22:0x UK).
Dead by arithmetic (sections 3 to 7): dropping the shadow after class v5 (v5 leaves the chip at 5.1x to 9.1x), the per-window refresh as a cost (a 1 GiB rebuild is about 0.1 J on a chip core), extra reads or longer chains (both sides scale together), memory-system shaping (the same DRAM on both sides), per-card adaptive classes (a chip picks the class that maximises its own reward per joule, so a menu never beats the card's best single class), proof of useful work (the useful fraction is bounded at 8 percent of the hash budget at 1 GH/s and falls with scale; no ZK ASIC has beaten a GPU per joule on real hardware in the public record; every deployed useful-work scheme was gamed), and the tensor block (`k` near 1 but 0.056 pJ per multiply-add, so there are no joules in it to force without 26x the verifier cost).
## 1. The measured facts this rests on
| Fact | Value | Label | Source |
|---|---|---|---|
| RTX 5090 class v3 and v4 across the core-clock grid, the same hash rate | unlocked: v3 136.59 MH/s at 330.2 W, v4 136.84 at 475.5 (premium 145.3 W); 1,400 MHz lock: v3 134.68 at 228.0 W, v4 134.98 at 316.3 (premium 88.3 W); the rate memory-bound and flat from 2,850 to 1,400 MHz (136.8 to 135.0); the knee below 1,400 being read now | measured, PC 2, 7 October 2026 | `docs/plans/counter-asic-3-status.md` (ca3-coord), the efficiency pass table |
| The shadow's marginal energy per counted op on the 5090 | 10.4 pJ unlocked (145.3 W over 136.84 M hashes x 102,100 ops); 6.5 pJ at the 1,400 lock (88.3 W over 134.98 M x 102,100); 10.2 to 13.2 pJ on the 6 October ladder under the 431 W cap | measured (arithmetic on measured rows) | the same; `docs/analysis/latency-shadow-2026-10-06.md` section 5 |
| The 5090's ALU budget and where the shadow binds | 45.2 T op/s at 3,050 MHz; about 332,000 ops per hash before compute binds at 136 MH/s; at a 1,400 MHz lock the budget scales to about 20.7 T op/s and the bind point to about 150,000 ops | measured budget; the scaling approximate | `docs/benchmarks/repro.md` 2.2; latency-shadow section 5 item 1 |
| RTX 5090 idle | 73.8 W at 862 MHz (PC 2, 6 October); 90.6 W (PC 1, 5 October) | measured | latency-shadow section 5; the Ember readbacks in the same file |
| RTX 4070 at its tune point (1,860 MHz lock, 160 W cap) | class v3 30.95 MH/s at 79.5 W (2.57 microjoules); class v4 3.51 microjoules, 109 W: the owner pays 30 W more; the rate holds to about 200,000 ops | measured, PC 1, 6 October | status file item 8, the 4070 rows; `docs/analysis/horizon/algorithm.md` 5.1 |
| RX 9070 XT | about 18.6 MH/s at about 304 W (10.6 microjoules); the shadow costs it no rate at any rung (+3.6 percent at rung 3); its watts under class v4 OWED | approximate (watts from memory); rate measured | algorithm.md 5.1; `docs/design/latency-ladder.md` section 8 |
| Apple M5 Max, GPU plus DRAM channels | class v3 27.08 MH/s at 21.0 W (0.78 microjoules); class v4 at 100,000 ops 1.40 microjoules (+16 W, -1.5 percent of rate); marginal 6.9 pJ per counted op | measured, 6 October | latency-shadow section 3 |
| The stored-dataset chip (f = 1) | GDDR7, the 5090's own 16 devices without the GPU: 166.4 MH/s at 77.6 W, 0.466 microjoules; one HBM3 stack 83.6 MH/s at 26.8 W, 0.321; eight stacks 666 MH/s at 174 W, 0.262; a node and executor beside it about 85 W (approximate) | modelled (activate-bound ceilings, 2.0 and 1.2 nJ per random 32-byte read, static and controller allowances) | chip-model-v3 sections 5.3 to 5.5 and 5.10 (the Counter lane's class v5 rows, relayed 20:5x UTC) |
| Class v5 at zero shadow | GDDR7 5.1x per joule against the 5090 at 2.40 microjoules (2.5x with a node per chip); one HBM3 stack 7.2x (1.8x with a node per chip); eight stacks 9.1x (6.2x); the leaves ship at 16.5 KB/s to 10,000 members today, so a farm shares one node | modelled | chip-model-v3 section 5.10; `docs/design/class-v5-stored-state.md` 2a.2 |
| The tensor-shaped block (mm8) on an RTX 4090 | free in rate to 4,096 tiles per hash; 2.9 to 14.7 W; 0.70 falling to 0.056 pJ per multiply-add; verifier +1.1 microseconds per step per unit scalar | measured, 6 October | `docs/analysis/horizon/new-pow.md` 5.1 |
| Proving on the 5090 | a shard at the adopted v1 budget (4.7 M SP1 cycles) 4.3 s alone, 13.2 s beside the miner; the prover alone peaks at 190.6 W on empty shards; a GPU-proved block's useful fraction bound 8 percent of network hashes at 1 GH/s | measured; the bound is arithmetic | `docs/bench-log.md` "proving v1" and "the S_p curve"; new-pow section 6 |
| The verifier gate | class v4 rung 0 8.77 ms cold with the sibling loaded on the reference core (5.14 alone); rungs 0 to 2 admissible, rung 3 out by 0.08 ms; class v5 adds 0.20 to 0.28 ms | measured, igneum-build-1, 6 and 7 October | latency-ladder section 5; class-v5 section 7 |
## 2. The identity, and what it forbids
Write the card's energy per hash on class v3 as `E_card = E_idle + E_memsys + E_wait`: the idle floor spread over the hashes (74 W over 135 MH/s is 0.55 microjoules on the 5090), the memory system's own reads and standby (about 55 W, 0.41 microjoules, modelled), and the SMs, fabric and clock tree while they wait on loads (the rest: 228 - 74 - 55 = 99 W, 0.73 microjoules at the lock, unmeasured as a breakdown). The chip pays `E_mem` (its memory, controller and static: the whole 0.466) and nothing else. Add forced work `W` ops per hash at `e_gpu` per op on the card and `e_chip` on the chip, `k = e_chip / e_gpu`, `F = W e_gpu`:
| Quantity | Formula | 5090 at the 1,400 lock, GDDR7 chip | 5090 unlocked, GDDR7 | 5090 lock, one HBM3 stack |
|---|---|---|---|---|
| Edge at zero premium | `E_card / E_mem` | 3.63x | 5.19x | 5.27x |
| Edge with the premium `F` | `(E_card + F) / (E_mem + k F)` | at F = 0.654 (88 W): 2.07x at k = 1 (e_chip 6.5 pJ); 2.92x at 3.3 pJ; 3.50x at 2.0 pJ; 4.12x at 1.0 pJ | at F = 1.062 (145 W): 2.27x at k = 1 (10.4 pJ); 4.33x at 3.3 pJ; 6.12x at 1.0 pJ | 2.38x at k = 1; 3.56x at 3.3 pJ; 5.54x at 1.0 pJ |
| The asymptote at infinite premium | `1 / k` | 1x at k = 1, 2x at k = 0.5 | the same | the same |
| The premium that reaches 2x at k = 1 | `E_card - 2 E_mem` | 0.761 microjoules, 103 W | 1.486, 203 W | 1.051, 142 W |
| The premium that reaches 2x at k at or under 0.5 | none | | | |
| The slope at F = 0 | `(E_mem - k E_card) / E_mem^2` per microjoule | k = 1: -5.65x per microjoule; k = 0.5: -1.75; k = 0.3: -0.19; k = 0.275: 0 | k = 1: -9.0; k = 0.19: 0 | k = 1: -13.3; k = 0.19: 0 |
Three consequences, each labelled arithmetic on the rows above:
1. **Zero premium is zero forcing.** `F = 0` gives `k F = 0` for every `k`. A design that costs the GPU nothing extra costs the chip nothing extra. The only exceptions would be levers with an asymmetry the identity does not see: work the card already pays for in `E_idle + E_wait` that the chip does not pay for. Sections 4 and 5 look for one and find none that a consensus rule can check (the GPU's L2 is the nearest, section 4 row 5, and it costs rate).
2. **The shadow's worth is a bet on `k`, and the bet gets narrower at the knee.** At the 1,400 lock the shadow buys 1.5x of edge at `k = 1`, 0.7x at a 3.3 pJ core, nothing at 1.8 pJ, and below that it costs us. The reference the project has used for a pessimistic core (Bitmain's withdrawn X9 at about a third of a desktop CPU's energy per RandomX hash, claimed; latency-ladder section 5a) is a ratio against a CPU core, and a CPU core spends about 100 pJ or more per instruction against a GPU lane's 6.5 to 10.4 pJ measured here (Horowitz 2014 gives 70 pJ per in-order instruction at 45 nm, Dally 2023 gives 250 pJ per out-of-order ARM instruction; both claimed). A chip three times better than a CPU core on a random program is a chip WORSE than a GPU lane per op. So the X9 implies nothing about `k` against a GPU; the honest band for a chip maker who builds a SIMD array like a GPU's is `k` 0.5 to 1 (approximate), with the N5 fixed-datapath floor (0.19 pJ of datapath before register file, operand wires and instruction fetch; chip-model-v3 5.1) as the theoretical bound that nobody has shown on a random 32-lane program with shuffles. The honest headline stays 2.1x at `k = 1`, and the project should stop quoting the X9 as a `k` of 0.33 against a GPU.
3. **What moves the premium-free floor is `E_card`, and the hash cannot touch it.** The 5090's own idle (74 to 91 W measured) plus its memory system (55 W modelled) bound the floor near 2.0x on GDDR7 at any operating point; the Apple GPU at 21 W shows a GPU designed for low idle sitting at 1.7x with no shadow. The measurements that read how much of the 99 W of `E_wait` is reachable are the knee sweep (running) and the SM-sparse kernel (section 12).
## 3. Thread 1: class v5 with the shadow at zero
The question: once class v5 is live (the dataset built from the chain's execution state, refreshed per window), does the shadow still buy anything? Re-run of the chip model with v5 on and the shadow at zero (ladder rung below 0, which is class v3's energy, since no rung below 0 exists and dropping the shadow is the class object moving back to v3 under the 95 percent rule):
| Chip under class v5, shadow at zero | Rate, watts | Energy per hash | Edge over the 5090 at 2.40 microjoules (the model's denominator) | Edge at the 1,400 lock (1.69) | With a node and executor per chip (85 W, approximate) | Label |
|---|---|---|---|---|---|---|
| GDDR7, 16 devices | 166 MH/s at 78 W | 0.466 | 5.1x | 3.6x | 0.977 microjoules: 2.5x (1.7x at the lock) | modelled (the Counter lane, chip-model-v3 5.10) |
| HBM3, one stack | 84 MH/s at 28 W | 0.321 | 7.2x | 5.3x | 1.34: 1.8x (1.3x) | modelled |
| HBM3, eight stacks | 666 MH/s at 175 W | 0.262 | 9.1x | 6.5x | 0.389: 6.2x (4.3x) | modelled |
What v5 adds to the chip's bill is the window's leaves, 64 bytes per state record, 5,952 bytes at today's devnet state and at most the dataset's size at the sample cap (class-v5 2a.2), delivered over a link: 16.5 KB/s to 10,000 members today; the WAN line is about 700,000 state records (45 MB per member per hour at 1 Gbit/s), a LAN farm is never bounded. So the "node per chip" column is not reachable by construction: one node serves a farm, as the v5 design states itself. The arithmetic the founder asked for: under v5 alone the strongest chip keeps 5.1x (GDDR7) to 9.1x (eight HBM3 stacks) at the stock point and 3.6x to 6.5x at the lock, all over 2x. **The answer is not "drop the shadow after v5."** What v5 does remove is the f = 0 recompute chip as a category and the pool miner who holds nothing but the day key (class-v5 section 1); its merit is there, not against the memory-system chip.
A construction that WOULD force a node per card does not exist below the state size: everything the lottery derives is derived from a seed and the state, so the only bytes a central node cannot compress away are the state's, and today's state is 6 KB. The refresh cadence does not help either (section 4 row 7).
## 4. Thread 2: energy-shaped work instead of op-shaped work
The brief's hypothesis: some operations cost a GPU almost nothing extra in energy because the hardware already pays for them during the memory wait. The identity says what to look for: an operation class with `k` as high as possible (a chip can improve on the GPU least) at a cost `e_gpu` that is REAL joules (so the forcing exists). A class with tiny `e_gpu` forces nothing; a class with tiny `k` forces the card and spares the chip. Per class, the figures:
| Operation class | Card's marginal energy per op | A chip's cost to match it | k band | Premium per 0.1x of edge bought at the 1,400 lock, GDDR7 (from the slope at F = 0) | Verifier cost | Reading | Labels |
|---|---|---|---|---|---|---|---|
| Random 32-bit integer ALU program (the class v4 shadow: add, xor, mul, mad, rotates, sub, mulhi, or) | 5090: 10.4 pJ unlocked, 6.5 pJ at the 1,400 lock; M5 Max 6.9 pJ | a SIMD array at N5: 0.19 pJ of datapath plus register file, operand network and instruction fetch; the model's columns 3.3 pJ (k 0.3 of 11) to 11 pJ | 0.5 to 1 honest band; 0.15 to 0.3 the attacker's claim (unshown) | k = 1: 0.018 microjoules (2.4 W) per 0.1x; k = 0.5: 0.057 (7.7 W); k = 0.3: 0.53 (71 W) | 3.2 microseconds per 1,000 shadow instructions per warp on the M5 Max core; rung 0 fits with 1.2 ms to spare on the loaded reference core | the right lever IF k is near 1; its value collapses below k = 0.5 and reverses at 0.28 | card measured; chip approximate; slope arithmetic |
| Shuffle-heavy mix (shfl, shfla through the warp crossbar) | the shfl step 0.86x of the add-xor-rotate chain on the M5 Max, 1.13x for rotr; the 5090's per-family step costs in the item 6 record; AMD shfla 0.75 to 0.84 | a crossbar is the same wire physics on any die; the algorithm lane's k floor rises from 0.32 to 0.46 with a shuffle-heavy weight table (approximate) | 0.46 to 1 | at k = 0.46: 0.075 microjoules (10 W) per 0.1x | the same law; shfl is one op in the interpreter | the content to prefer inside the shadow: the same premium, less of it reaches a chip | measured steps; k floor approximate (algorithm.md 5.3) |
| 32-bit multiply heavy | a 32 x 32 multiplier is 0.52 pJ at N5 datapath on both sides (claimed, mlsysbook citing Horowitz and Dally); the GPU's overhead around it is the operand delivery | the same multiplier | about 0.5 to 1 | as the shuffle row | one op | second choice behind shuffles; the saturation and lossy-source rules of sub-version 3 already bound what a multiply can do to the read map | claimed figures |
| Tensor-shaped int8 tiles (mm8) | 0.056 pJ per multiply-add at 4,096 tiles per hash on the 4090 (57 pJ per 1,024-MAC tile); 14.7 W in all | a licensable int8 MMA block: the same circuit | 0.5 to 1 | to reach the shadow's 0.654 microjoules the block needs about 11,500 tiles per hash (R about 1,450); the 4090 binds near R 3,000 | scalar 1.1 microseconds per step per unit: about 10.6 ms at R = 1,450 on the M5 Max core, FAILS the gate alone; a byte-dot verifier (AVX-VNNI, NEON udot, AVX2 pmaddubsw) 4x to 16x cheaper, approximate, unmeasured | k is good and the joules are too cheap: no forcing without the verifier paying 26x per joule forced (new-pow 5.1 item 3). Kept as reserve R8 for datapath diversity | measured block; verifier scaling approximate |
| L2-resident hot reads (a second table sized to GPU L2, read beside the dataset) | an L2 hit on a 96 MB L2: about 0.1 to 0.3 nJ per 32-byte sector (approximate: Horowitz's 1 MB cache at 100 pJ per 64 bits at 45 nm, scaled to N5 and a 96 MB array with its crossbar); measured in the added form the table cost the 5090 13 to 16 percent of RATE because the dataset stream evicts it (5 October, no cache-policy hints) | a 64 MiB SRAM read on a 128 mm^2-class array: 0.2 to 0.5 nJ (chip-model-v3 5.1, approximate) | 0.7 to 3: the one class where a chip may be WORSE than the GPU per op, because a big SRAM array is not cheaper than a GPU's L2 slice | at k = 1.5: 0.008 microjoules (1 W) per 0.1x, if the hot reads were free in rate | the hot table's items are derived lazily on the CPU (unmeasured per unit) | worth ONE measurement, not a claim: the 5 October rows used no cache hints; CUDA `createpolicy.fractional.L2::evict_last` for the hot table and `ld.global.L2::evict_first` for the dataset stream, AMD's slc bits; nothing on Metal, so Apple pays rate. If the rate holds with hints, this is the only lever with k possibly over 1 | approximate throughout |
| Dependent pointer chasing: more reads per hash, a longer chain | the card's incremental energy per read about 3.1 nJ (55 W over 17.5 G reads/s, modelled); its overhead power is per second, so a longer chain lowers both rates together | 2.0 nJ per read | the edge is unchanged: both sides' energy per hash scale by the same factor (E_card and E_mem are both per-second powers over a rate that moves together) | none bought at any premium | +0.13 ms per 128 reads per unit, approximate | not a lever (chip-model-v3 5.7 row "latency") | modelled |
| Extra random reads at the same bandwidth (wider loads, w16, w32) | the 5090 fetches a 32-byte sector per 4-byte read already; w16 moved the 5090 2.7 percent and the chip not at all; w64 made the 5090 bandwidth-bound (71.9 MH/s) | the same 32-byte atom per channel access | 1 | nothing to force: the chip already pays the sector | none | not a lever; the 4-byte decision stands | measured (read-width) |
| Per-lane live state (a scratch, register-resident permutations across the hash) | a chip keeps 1,172 lanes at 55 ns of controller latency against the 5090's 7,262 at 415 ns; lane state is 64 B to 320 B | SRAM at picojoules per access | not applicable: lane state is not an energy term on either side | none | | `docs/analysis/scratch-soundness.md`: the live state is bounded by the read-modify-write count and sits in a chip's SRAM at under 5 percent of its mirror | measured bound |
| The dataset refresh as the cost (class v5's rebuild per window, per epoch or per block) | a 1 GiB rebuild is 2^24 items at about 9,360 ops each, 157 G ops: 32 ms on a 4090, 13.4 ms on the 5090 | 157 G ops at 1 to 3 pJ per op is 0.16 to 0.5 J per rebuild; a rebuild per block is 0.16 to 0.5 W against 78 W of hashing; the write is 1 GiB per second against 1.8 TB/s | the refresh is 0.1 percent of the chip's item traffic | nothing | a rebuild per second costs the honest 4090 3 percent of its hash time (rate, not joules) | dead by arithmetic; the refresh cadence is a liveness tool (proof of following), not an energy lever | arithmetic on measured build times and approximate chip energies |
What the table says: the founder's instinct is right that the shadow spends energy because it uses ALUs, and the public energy figures agree on why a GPU lane is expensive per op (Dally, Hot Chips 2023, claimed: fetch, decode and operand delivery about 30 pJ per lane instruction at 45 nm against 0.1 pJ for the add itself; a tensor instruction amortises that overhead over 1,024 multiply-adds, which is exactly why the mm8 block is cheap for the card and useless as a forcing lever). The classes where the card's energy is "already paid" during the wait (lane state, in-flight reads, the memory burst) are classes where the chip's cost is zero or equal, so they force nothing. The one class where a chip might pay more than the GPU per op is the GPU's own L2, and the one measurement of it lost rate; a cache-hinted re-run is the open item. Everything else reduces to choosing the shadow's content for the highest `k` at a fixed premium, which is design 3.
## 5. Thread 3: memory-system binding beyond random reads
Row-buffer-miss patterns, bank-conflict shaping, refresh-aware access and burst-sized items were each checked against the identity. The hash already opens a row per read (random 4-byte reads over 1 GiB: a row hit rate near zero on both sides); a 32-byte item is already the GDDR7 channel burst (the 5090 measures a 32-byte sector per read; the 9070 XT a 64-byte line, which is AMD's cost, not the chip's); bank parallelism is the same 1,024 banks on the 5090's board and on a chip built from the same 16 devices; refresh is the DRAM's. **There is no asymmetry to shape: the chip and the card have the same DRAM, the same row cycle (tRC about 45 ns across DDR4, GDDR and HBM; Li, Reddy and Jacob, MEMSYS 2018, claimed) and the same activate window.** The two levers a memory system offers are both the chip's: a lower energy per read (HBM's 1.2 nJ against GDDR7's 2.0 through 0.8 pJ per bit of interposer I/O against a PCB, modelled; fine-grained DRAM would cut the 909 pJ activation further, O'Connor et al. MICRO 2017, https://www.cs.utexas.edu/~skeckler/pubs/MICRO_2017_Fine_Grained_DRAM.pdf, claimed) and more activates per second per watt (AMD's "Folded Banks", ISCA 2025, 6.7x the irregular bandwidth from 8x the activate parallelism, https://dl.acm.org/doi/10.1145/3695053.3731111, claimed). Both move `E_mem` down, so the premium-free edge can only rise with memory generations, which is the HBM4 point the Horizon lane made (algorithm.md proposal 5).
What the Ethash and RandomX record measured (the literature sub-agent's pass tonight, 7 October 2026; every figure claimed by its vendor unless marked):
| Algorithm, what it relied on | The chip, its date and figures | The GPU at the time | Chip edge per joule | Where the resistance failed | Sources |
|---|---|---|---|---|---|
| Ethash: 4 GB+ DAG, 64 random 128-byte reads per hash, bandwidth-bound | Antminer E3 (Jul 2018) 190 MH/s at 760 W; Innosilicon A10 Pro (2020) 500 at 950; Linzhi Phoenix (Dec 2020) 2,733 MH/s at about 3,000 W (F2Pool test, measured); Jasminer X4 (Nov 2021) 2,500 at 1,200; Antminer E9 Pro (2023) 3,680 at 2,200; Jasminer X16-P (2024) 5,800 at 1,900 | RTX 3080 about 100 MH/s at 220 W (0.45 MH/W, measured) | E3 0.55x; A10 Pro 1.2x; Phoenix 2.0x; X4 4.6x; E9 Pro 3.7x; X16-P 6.8x | the chips won by moving the same bytes at lower energy per bit (custom controllers, on-package and on-die memory: Linzhi's 72 mixers and 2.8 Tb/s of internal memory); nothing in the hash's logic | https://www.theblock.co/post/88622/questions-new-ethash-asic-ethereum ; https://f2pool.io/mining/pow-round-up/20201228-pow-round-up/ ; https://www.asicminervalue.com/miners/jasminer/x4 ; https://pool.kryptex.com/ms/device/asic/jasminer/x16-p ; https://ethereum-magicians.org/t/progpow-audit-delay-issue/3309?page=4 |
| RandomX: 2 GiB dataset, one 64-byte read per iteration, 8 chained random programs, CPU-shaped | Antminer X5 (Sep 2023) 212 kH/s at 1,350 W; Antminer X9 announced 26 Dec 2025 at 1,000 kH/s and 2,472 W, withdrawn May 2026 before any unit shipped; Pinecone R1X 1,200 kH/s at 2,055 W (unverified) | Ryzen 9 3950X 12.9 kH/s at 105 W (measured); GPUs 18x worse per joule than the CPU | X5 1.3x; X9 3.3x (claimed, never shipped); R1X 4.7x (unverified) | against its honest device the latency-bound random program held a chip to 1.3x shipped and 3x to 5x claimed after seven years; RandomX v2 testnet forked 5 October 2026 | https://github.com/tevador/RandomX/blob/master/doc/design.md ; https://pool.kryptex.com/articles/antminer-x5-en ; https://github.com/monero-project/monero/issues/10270 ; https://oneminers.com/blogs/news/whatever-happened-to-the-antminer-x9-bitmain-monero-miner |
| ProgPoW and KawPoW: random math per period on the GPU's datapath, 256-byte DAG loads | none shipped as of 2026 | RTX 4090 65 MH/s at 330 W | the authors' claim 1.1x to 1.2x; Rao's audit (Sep 2019): conventional chips gain little, memory-intensive algorithms face an "imminent threat" in general | the resistance held; the energy per hash on GPUs is the price (KawPoW 0.2 MH/W against Ethash 0.45 on the same cards) | https://github.com/ifdefelse/ProgPOW/blob/master/README.md ; https://leastauthority.com/?p=759 |
| Equihash 200,9: a 144 MB sort | Antminer Z9 mini (May 2018) 10 kSol/s at 247 to 300 W; Z15 (Jun 2020) 420 kSol/s at 1,510 W; Z15 Pro (2023) 840 kSol/s at 2,780 W | Vega 64 506 Sol/s at about 250 W | 17x to 151x | the working set fitted on-die and the sort pipelines; Zcash voted 45 to 19 against resistance (Jun 2018) | https://overclock3d.net/news/misc/bitmain-launches-their-antminer-z9-mini-zcash-mining-asic/ ; https://support.bitmain.com/hc/en-us/articles/19259431024281 |
| Cuckatoo32: a 512 MB SRAM random walk | iPollo G1 (Dec 2020) 36 graphs/s at 2,800 W | RTX 3090 1.0 to 1.43 graphs/s at 285 W | 2.6x to 3.7x | the random walk stayed latency-bound even in SRAM: the smallest edge of any compute-shaped chip | https://pool.kryptex.com/device/asic/ipollo/g1 ; https://forum.grin.mw/t/mining-hardware-comparision/9011 |
| kHeavyHash, Blake3, SHA (compute only) | IceRiver KS0 to Antminer KS5 Pro (2023 to 2024); Antminer AL1 (2024) | RTX 4090 | 160x to 700x | nothing memory-bound to hold them | https://asicminervalue.com/miners/bitmain/antminer-ks5-pro-21th ; https://pool.kryptex.com/device/asic/bitmain/antminer-al1 |
The reading for Igneum: Igneum's hash is in the Ethash and RandomX class (DRAM-bound), and the Ethash chips' 2x to 7x is exactly `E_card / E_mem` with a better memory system; the Cuckatoo and RandomX rows show what latency-bound work does to a chip's logic edge (3x against a CPU, under 1.3x shipped). The identity predicts both.
## 6. Thread 4: per-card adaptive work
The question: could the protocol admit per-class programs where the shadow is scaled to the card's own memory-to-compute ratio, without letting a chip pick the cheapest? **No, and the proof is one line.** Let the menu be a set of classes `c`, each with a difficulty weight `w(c)` (reward per hash under that class). An honest card's best reward per joule is `max_c w(c) / E_card(c)`, a chip's is `max_c w(c) / E_chip(c)`. The chip's edge under free choice is the ratio of the two maxima. Since the chip may pick the card's own best class `c*`, `max_c w(c) / E_chip(c) >= w(c*) / E_chip(c*)`, so
edge(menu) >= w(c*) / E_chip(c*) over w(c*) / E_card(c*) = edge(c*),
the edge at the card's best single class. A menu can never beat the best single class for the honest card, and whenever any class on the menu is cheaper for the chip than `c*` is, the menu is worse. The chain cannot verify hardware, only values, so no rule can bind a class to a card type (new-pow 3.2 makes the same point about timings and capacities). The ladder (`docs/design/latency-ladder.md`) is the construction that survives this: one network-wide `N`, stepped by miner signal within a verifier-bounded genesis list, which the attack-pass lane's F10 row gated (monotone, 89 percent does not move it, no two-step jump). What a per-card ratio CAN do is sit in the miner, not the protocol: a card picks its own clock, grid and kernel shape (designs 1 and 2) for the one network class.
## 7. Thread 5: proof of useful work
Igneum's miners are its provers. Could the proving replace or fund the shadow?
| Quantity | Value | Label | Source |
|---|---|---|---|
| Energy per proven shard on the 5090 tonight (the adopted v1 budget, 4.7 M SP1 cycles) | 4.3 s alone at up to 190.6 W: at most about 820 J per shard, at least about 5,700 cycles per joule (174 microjoules per cycle); beside the miner 13.2 s | measured time and the prover's peak watts (the watts for the v1 shard itself unmeasured, the empty-shard peak is the bound) | bench-log "the S_p curve" and "proving v1 step 1" |
| The hash for comparison | 590,000 hashes per joule at the 1,400 lock | measured | section 1 |
| What the chain needs proven | 1 block per second, 2 shards per block at the v1 budget: about 1.6 kW of 5090 proving network-wide, independent of the hash rate | arithmetic | the same |
| The useful fraction of a proving-as-lottery scheme | proving work per segment over network hashes per segment: 8 percent at 1 GH/s, 0.08 percent at 100 GH/s; gas sets the numerator, the security budget the denominator; a verifier that recomputes the piece or checks a 32 to 40 ms proof cannot sit under the 10 ms gate | arithmetic on measured figures | new-pow sections 3.1 and 6 (scheme A, NEVER) |
| Public GPU zkVM proving rates | Airbender 21.8 MHz of RISC-V cycles on an H100 and 9.7 MHz on a 4090 (about 16 to 46 J per million cycles at TDP); SP1 Turbo 3.45 MHz on an H100; RISC Zero 1.1 MHz; SP1 Hypercube 16 RTX 5090s for 99.7 percent of Ethereum blocks under 12 s (Nov 2025), "12.5x GPU efficiency in 6 months" | claimed | https://zksync.io/airbender ; https://blog.succinct.xyz/real-time-proving-16-gpus/ |
| Proving chips | Cysic ZK Air "comparable to 10 RTX 4090", ZK Pro "50", no watts published, shipping slipped to 2026; Ingonyama's ZPU paper model "about 13x the efficiency of an A40" with no silicon; Fabric Cryptography's VPU "orders of magnitude", no number, no watts, still pre-order; academic ASICs (SZKP, NoCap, UniZK, zkSpeed) 46x to 800x against CPUs or a 2017 V100, all simulation, bandwidth-hungry (2 to 3 TB/s of HBM) | claimed | https://docs.cysic.xyz/hardware-products/zk-asic-products/ ; https://www.ingonyama.com/oldblogs/zpu-the-zero-knowledge-processing-unit ; https://www.fabriccryptography.com/ ; https://arxiv.org/html/2408.05890v1 ; https://arxiv.org/html/2504.06211v1 |
| The one measured non-GPU ZK data point | Jane Street's Hardcaml MSM, ZPrize FPGA winner: 2^26 MSM in 5.08 s at 52 W, about 264 J; the GPU winner on an A40 about 0.56 s per MSM, at most about 170 J at TDP: the FPGA lost on energy | measured times; energy approximate | https://blog.janestreet.com/zero-knowledge-fpgas-hardcaml/ ; https://github.com/matter-labs/z-prize-msm-gpu-combined |
| What happened to useful-work schemes | Aleo dropped PoSW consensus after "a small number of provers developed specialized hardware" and rewrote the puzzle when miners optimised MSM/NTT (2022 to 2024); Qubic's rented CPUs took 27 to 51 percent of Monero's hash and reorged it (Aug 2025); Boundless PoVW's token fell 96 percent from its high and RISC Zero closed its hosted prover (Dec 2025); Pearl cuPOW's 112 MW did "zero useful AI computation" | measured events | https://equilibrium.co/writing/aleo-mainnet-launch-reflecting-on-the-journey ; https://aleo.org/post/prover-incentives-2-retrospective/ ; https://www.theblock.co/post/364496/qubics-monero-hashrate-controversy ; https://blockeden.xyz/blog/2026/01/14/boundless-risc-zero-decentralized-proof-market-zk/ ; https://arxiv.org/abs/2606.04819 |
Does it change the question? **No, in two ways.** (1) The useful fraction is bounded by gas, not by the hash budget: at any scale past 1 GH/s the chain cannot spend its security budget on proving because there is not enough proving to spend it on, and the per-hash verifier cannot check a proof's piece under 10 ms. (2) Proving is compute-shaped (NTT, hashing over 31-bit fields, MSM), the shape in which chips reach 46x to 800x in simulation and the shape that Aleo's provers specialised against; making the lottery out of it hands the chip the lottery. Proving chips do not exist as measured products today (no vendor publishes watts; the one measured FPGA lost to the GPU), and the public GPU proving software moved 12.5x in six months, which a fixed pipeline cannot follow; so for 2026 to 2028 a proving chip's credible edge over a 5090 is 5x to 30x per joule (approximate, from the simulation claims against obsolete GPUs discounted by the software trend), which is worse than the hash's 2.1x to 3.6x. The 80/20 split stands; the energy spent on proving is defensible only because paying users demand the proofs.
## 8. Thread 6: the literature since 2023
| Item | What it did | Cost to GPUs in energy | What chips did | Date, source |
|---|---|---|---|---|
| Kaspa's chip turn | kept kHeavyHash; "ASIC resistance only delays the unavoidable" | none (the GPUs left) | IceRiver KS0 162x per joule over a 4090, Antminer KS5 Pro 700x; GPU share negligible within months | 2023 to 2024; https://kaspa.org/?p=45054 ; https://www.asicminervalue.com/miners/iceriver/ks0 |
| Karlsen, Iron Fish to FishHash (Ethash-derived, 4.6 to 5 GB DAG) | the Kaspa forks that wanted GPUs back took an Ethash-class DAG | 4090 82.5 MH/s at 260 W (0.32 MH/W, measured) | none has followed; the Ethash chips' 2x to 7x is the precedent waiting | Apr and Sep 2024; https://pool.kryptex.com/articles/ironfish-hardfork-en |
| Alephium Blake3 | compute | n/a | Antminer AL1 about 200x per joule | 2024; https://pool.kryptex.com/device/asic/bitmain/antminer-al1 |
| Autolykos v2 (Ergo) | a 2 GB table growing 5 percent per 51,200 blocks to about 66 GB | 4090 265 MH/s at 240 W (1.1 MH/W) | none through 2026 | https://docs.ergoplatform.com/mining/asic/ |
| NexaPoW (Schnorr per attempt) | "useful ASICs" | 4090 0.84 MH/W | Dragonball A21 1.89 MH/W: 2.2x | Dec 2024; https://pool.kryptex.com/articles/dragonball-a21-en |
| XelisHash v2 and v3 | ChaCha8 plus Blake3 scratchpad, sequential | GPU-mined | none; forked v2 for FPGA resistance (Jul 2024), v3 Dec 2025 | https://github.com/xelis-project/xelis-hash |
| RandomX v2.0 | released 25 to 30 Mar 2026, testnet fork 5 Oct 2026 | CPU-mined | the response to the X5 and the announced X9 | https://www.coinpro.ch/en/?p=42081 |
| PoSME (Condrey, arXiv 2604.15751, 17 Apr 2026; IETF draft May 2026) | latency-bound pointer chasing, 32 MiB to 2 GiB arena; claims an HBM3 bound of 1.6x to 2.0x, wafer-scale SRAM 5x to 28x, verifier about 6 ms | GPUs 14x to 19x slower than a CPU (claimed) | no deployment | https://arxiv.org/abs/2604.15751 |
| Bandwidth-hard functions from random permutations | formal bandwidth hardness | theory | | https://arxiv.org/pdf/2207.11519 (2022) |
| Qubic uPoW on Monero | rented CPU fleets at 27 to 51 percent of hash | economic, not silicon | a six-block reorg, 12 Aug 2025 | https://coinmetrics.substack.com/p/state-of-the-network-issue-326 |
Nothing since 2023 found a lever outside the identity. The latency-bound family (RandomX, PoSME, Cuckatoo, Igneum) holds a chip's LOGIC edge to about 1.3x to 3x against the honest device; the memory-side edge is `E_card / E_mem` everywhere, and the one project that has lived with it longest (Monero) answers with re-tunes, not with a trick.
## 9. The ranked list of designs
Each row: (a) the chip's per-joule edge over the 5090 with its labels, (b) the 5090's and the 4070's energy premium over class v3, (c) the verifier cost against the 10 ms gate, (d) build cost in agent hours, (e) what breaks it. The chip is the f = 1 GDDR7 chip unless named; HBM3 one stack in brackets.
| Rank | Design | (a) Chip edge per joule over the 5090 | (b) Premium over class v3: 5090 / 4070 | (c) Verifier | (d) Agent hours | (e) What breaks it |
|---|---|---|---|---|---|---|
| 1 | **The operating point as the shipped default**: core lock at the knee plus undervolt where the driver allows, per card model, in Ember Tune (miner-only, no consensus) | class v3 at the 1,400 lock 3.6x (5.3x), measured card, modelled chip; at the knee below 1,400 lower (being read); class v4 at the lock 2.1x at k = 1, 2.9x at a 3.3 pJ core (measured premium, modelled chip) | lowers the base: 330 to 228 W on v3; the v4 premium 145 to 88 W (both measured); the 4070 already sits at its tune point (79.5 W; its premium 30 W measured) | none | 2 to 4 (the Ember core-clock knob is ordered for 0.3.24; add the undervolt line, the AMD ADLX line, the Apple "no lever" line) | AMD has no `-lgc` (ADLX or nothing); Apple has no lever; a locked 5090's free shadow is about 150,000 ops, so ladder rung 2 costs it rate; the knee may sit where the driver refuses the lock |
| 2 | **The SM-sparse miner kernel**: the hash on about N of the SMs at 32 warps per block, the rest clock-gated (miner-only; worker variants `sp<N>-w<W>`, section 12) | unmeasured: reads the 99 W of `E_wait` at the lock; if half is SM activity, class v3 at the lock about 2.8x (4.0x) and class v4 at k = 1 about 1.7x; if none, unchanged | lowers the base by what it finds; no premium | none | 1 (built tonight) + 1 (one PC 1 job) + 3 to 4 to ship as the default kernel through the tuning file and the Devnet 2 gate | the overhead is clock tree and leakage, not SMs (the knee sweep shows this first); fewer SMs cannot hold 17.5 G reads in flight (MSHR limits), so the rate falls with the watts; AMD and Apple need their own variants |
| 3 | **The shadow at rung 0 with a high-k op mix** (shuffle and multiply heavy weight table; a class change under the 95 percent rule) | at the 1,400 lock and the same premium: 2.1x at k = 1 (unchanged), 2.9x at the pessimistic core instead of 3.4x (k floor 0.32 to 0.46, approximate) | unchanged: 88 W locked, 145 W unlocked / 30 W | unchanged (shfl is one op; about 8.8 ms loaded at rung 0) | 4 to 6 (the weight-table knob, packs, three-vendor fingerprints) plus the six gates across lanes | Apple pays shfl at 1.91x per op (under 1 percent of its ALU time at rung 0, argued); held by the coordinator until designs 1 and 2 are read |
| 4 | **The L2-resident hot table with cache-policy hints** (a 32 to 64 MiB table beside the dataset; dataset loads evict-first, hot loads evict-last) | if the rate holds: the one lever with k possibly over 1 (a 64 MiB SRAM read on a chip 0.2 to 0.5 nJ against an L2 hit 0.1 to 0.3 nJ, both approximate): 512 hot reads per hash is about 0.1 microjoules on both sides, class v3 at the lock about 3.2x (4.4x) | about 0.1 microjoules (13 W) if the hints hold the table; the 5 October rows without hints cost 13 to 16 percent of RATE on the 5090 and 16 to 20 on the 9070 XT | the hot items derived lazily per unit: unmeasured, likely under 0.5 ms | 4 for the measurement (a hinted variant of the 5 October hot-table packs on PC 2); 10 or more if adopted (a class change) | Metal has no cache hints (Apple pays rate); AMD's slc bits unverified; the chip's SRAM is $12 of die at 64 MiB, so the lever is energy only; it is a measurement, not a result |
| 5 | **The tensor block with a byte-dot verifier** (mm8 at R about 1,450, the shadow's premium in tiles) | k 0.5 to 1 (the same circuit): at the shadow's premium 2.1x at k = 1, 2.9x at k = 0.5 (modelled on the 4090's measured block energy) | the same 0.65 microjoules by construction / the 4070 untested | scalar FAILS (about 10.6 ms at R = 1,450 on the M5 Max core); VNNI, udot or pmaddubsw 4x to 16x cheaper (approximate): 1 to 3 ms, inside the gate on a 2019 core only with AVX2 | 8 to 12 (the two-output block exists in `proto-newpow/mma-shadow`; the SIMD verifier, the AMD WMMA layout gate, Apple's emulation measurement) | Pascal, RDNA 2 and every Apple card emulate at 10 to 16 steps per tile; the AMD fragment layout is unverified; buys nothing over design 3 at equal premium except a better-bounded k |
| 6 | **Class v5 with the shadow dropped** | 5.1x (7.2x), modelled; 3.6x (5.3x) at the lock | none / none | -0.65 ms (the shadow gone), +0.2 ms (v5) | 0 | fails the brief: the memory-system chip is untouched; v5 is shipped for its own reasons (the recompute chip and the stateless pool miner gone) |
| 7 | **The refresh as the cost** (per epoch or per block rebuild) | unchanged (the rebuild is 0.1 to 0.5 J on a chip core) | the 4090 loses 3 percent of hash time at one rebuild per second / the same | none | 0 | dead by arithmetic |
| 8 | **Per-card adaptive classes** | at least the edge of the card's best single class (the proof in section 6) | | | 0 | dead by the maximum argument; the ladder is what survives |
| 9 | **Proof of useful work** (proving in the lottery) | a proving chip's credible 5x to 30x (approximate) | | a proof piece cannot be checked under 10 ms | 0 | dead by the gas bound and the Aleo precedent |
| 10 | **Memory-system shaping** (row, bank, refresh, burst; more reads; wider loads) | unchanged | none | none | 0 | no asymmetry: the same DRAM both sides |
The three strategic answers that are not designs, for the record: (i) accept chips (the Kaspa turn: the chip edge is then the point, and the honest number to publish is the chip's USD 0.00021 per MH/s-hour against an owned 5090's 0.00113, algorithm.md 5.1); (ii) hold the line with the shadow, the ladder and the share-pattern detector, paying the premium at the knee and saying so; (iii) retire the premium by card-side efficiency (designs 1 and 2) and keep the shadow as the floor-setter. This file recommends (iii) with (ii)'s honesty.
## 10. The recommendation, in five sentences
The premium cannot reach zero against a chip that is a GPU's memory system without the GPU, because every joule we force the chip to spend is a joule the card spends first, at a ratio set by silicon; what zero premium buys is the card's own whole-card energy over the chip's memory energy, which is 3.6x on a 5090 locked at 1,400 MHz, about 2x at that card's idle-plus-memory bound, and 1.7x on an Apple GPU today. The shadow is therefore kept at rung 0 as the floor-setter, with its measured price stated at the knee (88 W at 1,400 MHz, lower if the knee is lower) and its measured worth (2.1x at a chip core equal to a GPU lane, 2.9x at a core three times better, nothing below 1.8 pJ per op), and its content re-weighted toward shuffles and multiplies once designs 1 and 2 are read, because that is the only way to raise the chip's `k` at the same premium. The lever that moves every row with no consensus change is the honest card's operating point, so the knee lock and the undervolt ship as the default in Ember Tune for every NVIDIA card model, with the AMD and Apple rows stated honestly as "ADLX or nothing" and "no lever". The SM-sparse kernel built tonight is the one measurement that can lower the premium-free floor further on the same card, and it decides itself in one PC 1 job: if a quarter of the SMs holds the rate at much lower watts it ships as the miner's default kernel, if the watts track the rate it is dead. Class v5 ships for its own reasons and does not retire the shadow; the public text should say the floor and the premium as numbers, and should stop quoting the withdrawn X9 as a chip core a third as costly as a GPU lane, since that ratio was against a CPU.
## 11. Consequences per tier
| Tier | What this file means tonight | What is being done |
|---|---|---|
| Home miner, one 8 GB card (NVIDIA; 4060 Ti class, 3.81 microjoules rented and untuned) | No chip exists; against the modelled GDDR7 chip the untuned card sits at 8.2x per joule on class v3 and about 3.1x on class v4 at k = 1 (algorithm.md 5.1). The two levers that move its row are in the miner, not the hash: the Ember tune (10 to 30 percent per joule measured on the 4070) and, if it reads well, the SM-sparse kernel. Class v4's premium on a small card is about 30 W at the tune point (the 4070's row) | design 1 in 0.3.24; design 2's PC 1 job |
| One 12 GB card (4070, 5070) | the 4070 at its tune point is 5.5x (v3) and about 2.2x (v4, k = 1) against the GDDR7 chip, measured card, modelled chip; its v4 premium is 30 W (79.5 to 109 W) for no rate | the same |
| One 16 GB card (RX 9070 XT) | 23x per joule behind the GDDR7 chip on class v3 (approximate watts); the shadow costs it no rate (+3.6 percent at rung 3) and its watts under v4 are OWED; AMD has no clock-lock path in the app (ADLX line or nothing) | the AMD watts row; the ADLX line in Ember Tune's text |
| One 24 or 32 GB card (5090 class) | the measured rows of this file: 5.2x unlocked, 3.6x at the 1,400 lock on class v3; class v4 2.1x at k = 1 for 88 W at the lock; a 5090 owner who locks the core pays 316 W instead of 476 W for 1.4 percent of rate | the knee sweep (running), the Ember knob, the SM-sparse job |
| Apple (M-series) | the honest best per joule the project owns: 1.7x from the GDDR7 chip with no shadow, 0.9x at class v4 and k = 1; no clock lever; a shuffle-heavy mix costs it 1.91x per shfl op (under 1 percent of its ALU time at rung 0, argued) | nothing to do for Apple in designs 1 and 2; design 3 is checked against its 5 percent rule |
| A rig | a rig's bill is watts: the lock takes a 5090 rig from 476 to 316 W per card on class v4 (-34 percent of electricity) for 1.4 percent of rate; the shadow's premium at the knee is the rig's price for the 2.1x floor | design 1 as the default |
| A pool user | nothing changes in shares or payout from any design here; designs 3 to 5 would be class changes announced by the 95 percent signal | |
| A node operator (the verifier) | designs 1 and 2 cost nothing; design 3 nothing; design 4 under 0.5 ms (unmeasured); design 5 needs a SIMD byte-dot path to fit the gate | |
| A chip | its edge is `E_card / E_mem` plus whatever the shadow forces at its own `k`: 3.6x to 5.2x on GDDR7 with no shadow depending on the card's operating point, 2.1x with the shadow at k = 1; a core cheaper than 1.8 pJ per counted op would turn the shadow in its favour at the knee, which no public figure shows for a random 32-lane program | the `k` question stays with the external brief (algorithm.md proposal 8) |
| The public claim | the honest sentence is the floor and the premium as numbers: "a chip that stores the dataset keeps the card's whole-card energy over its memory energy, 3.6x on a locked 5090 and 1.7x on an Apple GPU with no extra work; the shadow holds it to 2.1x against a core as good as a GPU lane for 88 W on a 5090 at the knee" | to the Counter lane for the chip texts once the knee is a number (the coordinator, 22:0x UK) |
## 12. What was built tonight: the SM-sparse worker variant
`proto-cuda/nvrtc/worker.cpp` (this branch) gains opt-in variants `sp<N>` and `sp<N>-w<W>` (N persistent blocks of W warps, W default 32): the pack's bound kernel text is rewritten at compile time with exact anchors (the race's existing mechanism, `variantSource`): `__global__ void igneum_hash_bound(...)` becomes `__device__ __forceinline__ void igneum_hash_bound_unit(..., uint32_t gid)` with its `gid` line removed, and a persistent wrapper of the kernel's own name loops the unit over the dispatch's nonces with stride `gridDim.x * blockDim.x`, taking a trailing `uint32_t nonces` argument. `launchHash` launches it as a grid of N blocks; `raceTime` carries the entry's shape while timing it; the winner installs it. The variants are made on demand by name (`parseSparseVariant`) and are not in the catalogue, so a default race never runs them and no shipped worker changes behaviour. Refused on a persistent (variant-5) pack and when combined with a `__launch_bounds__` minBlocks variant. The rewrite was mirrored in Python on the class v4 devnet pack and reads as intended; the exe was cross-compiled on igneum-build-1 under `lease pool 4 --class measure --label "ca4 research: worker cross-compile (SM-sparse variant)"` (22:48 UTC; `/srv/builds/ca4-research/nvrtc/igneum-worker-cuda-ca4sparse.exe`, sha256 ee8d0e70dd101f125f42c0c7cf07481a794ee18a1317acf68561b37ea18d72be, no icon or version block: a scratch exe for one job, not a shipping one). Not yet run on a card: the first run's self-test (the vector warps through the bound kernel and the 2^24 fingerprint equal to the base kernel's) is its gate, and an NVRTC compile error shows as the variant dropped with "compile:" in the race line. The job is with the hash lane for PC 1 after the knee sweep (rungs sp170, sp85, sp43, sp21, sp11 at the 1,400 lock and unlocked, class v3 and v4 packs; MH/s, watts, SM MHz, fingerprint per row).
## 13. Unverified and owed
- The knee: read at 22:1x UTC as 1,300 MHz on class v3 (223.3 W, 1.66 microjoules) and 1,200 MHz on class v4 (305.1 W, 2.28), the premium 81.8 W at the best points; every "at the lock" row in this file is the 1,400 row and moves by under 3 percent to the knee (3.63x to 3.56x on GDDR7; the class v4 edge at k = 1 stays 2.1x). The floor below 1,100 MHz is unmeasured (a third pass runs to it).
- The breakdown of the 5090's 228 W at the lock into idle, memory and SM activity is unmeasured; the SM-sparse job reads it.
- The RX 9070 XT's watts under class v4; the 2019-class verifier core (O-1.14); the M5 Max package watts.
- The chip side is the model (chip-model-v3 section 5): activate-bound ceilings, 2.0 and 1.2 nJ per random read, the static and controller allowances, the 85 W node; no chip has been measured. The `k` band 0.5 to 1 for a SIMD-array core is approximate; the N5 datapath floor is a bound, not a design.
- The L2-hit and 64 MiB SRAM energies of design 4 are scalings from Horowitz 2014 and chip-model-v3 5.1 (approximate); the cache-hinted hot table is unmeasured.
- The tensor block's SIMD verifier cost is a 4x to 16x scaling, unmeasured.
- Every external figure is the vendor's or the author's claim at its URL, read 7 October 2026 by the literature sub-agents of this lane; where a page refused the fetch the figure came through a search summary and is marked claimed.
## 14. Sources
Internal: `docs/analysis/chip-model-v3.md` (sections 5 and 6; section 5.10 on ca3-coord), `docs/analysis/latency-shadow-2026-10-06.md`, `docs/plans/counter-asic-3-status.md` (ca3-coord: the efficiency pass, the 4070 and 9070 XT rows), `docs/plans/counter-asic-3.md`, `docs/plans/counter-asic-3-derivation.md`, `docs/plans/counter-asic-2.md`, `docs/design/class-v5-stored-state.md` (class-v5), `docs/design/latency-ladder.md` (ladder and attack-ladder-5a), `docs/analysis/attack-pass/f1-shadow.md` and `f10-ladder.md` (attack-pass), `docs/plans/cryptanalysis/in-house-pass.md` (crypto-engage), `docs/analysis/horizon/new-pow.md` and `algorithm.md`, `docs/analysis/asic-resistance-history.md`, `docs/analysis/scratch-soundness.md`, `docs/bench-log.md`, `docs/evidence.md` row 15.
External (all read 7 October 2026): O'Connor et al., Fine-Grained DRAM, MICRO 2017, https://www.cs.utexas.edu/users/skeckler/pubs/MICRO_2017_Fine_Grained_DRAM.pdf ; Chatterjee et al., subchannels, HPCA 2017, https://www.cs.utexas.edu/users/skeckler/pubs/HPCA_2017_Subchannels.pdf (GUPS about 9 pJ per bit against STREAM about 3, claimed); Horowitz, ISSCC 2014, https://pages.cs.wisc.edu/~markhill/restricted/isscc2014_horowitz_power_scaling.pdf ; Dally, Hot Chips 2023 keynote, https://www.hc2023.hotchips.org/assets/program/conference/day2/Keynote%202/Keynote-NVIDIA_Hardware-for-Deep-Learning.pdf ; Micron GDDR7 4.5 pJ per bit via https://www.club386.com/micron-gddr7-improves-nvidia-gpus-by-up-to-30-in-gaming/ ; Samsung HBM3E 3.9 pJ per bit via https://www.igorslab.de/en/samsung-ends-afterburner-hbm4e-with-325-tb-s-bandwidth-to-exceed-nvidias-requirements/ ; AccelWattch, MICRO 2021, https://paragon.cs.northwestern.edu/papers/2021-MICRO-AccelWattch-Kandiah.pdf ; tensor cores on memory-bound kernels, https://arxiv.org/html/2502.16851v2 ; Shuhai, FCCM 2020, https://arxiv.org/pdf/2005.04324 ; Folded Banks, ISCA 2025, https://dl.acm.org/doi/10.1145/3695053.3731111 ; Li, Reddy, Jacob, MEMSYS 2018, https://terpconnect.umd.edu/~blj/papers/memsys2018-dramsim.pdf ; the Ethash, RandomX, ProgPoW, Equihash, Cuckatoo, Kaspa, Alephium, FishHash, Autolykos, NexaPoW, Xelis and PoSME URLs in sections 5 and 8; the proving URLs in section 7.

View file

@ -423,6 +423,7 @@ struct Pair {
int regs = 0, blocksPerSM = 0;
int blockWarps = 1; // threads per block = 32 x this (the winning variant's, else the worker's default)
std::string variant = "base"; // the bound kernel in service: a variant name (see allVariants)
int sparseBlocks = 0; // the SM-sparse variants (sp<N>): the persistent grid in blocks, 0 = the plain grid
std::string raceLine; // the race's one-line report, emitted by the main thread with "prepared"
double raceMs = 0;
// read-width experiment (5 October 2026): the pack's load class and, for variant 5, the persistent-warp scratch
@ -498,6 +499,15 @@ static bool launchHash(Ctx& c, Pair* p, CUdeviceptr out, uint32_t baseNonce, con
DRV_CHECK(c, c.drv.launchKernel(p->fHashBound, warps, 1, 1, 32, 1, 1, 0, s, args, nullptr), "cuLaunchKernel igneum_hash_bound (persistent)");
return true;
}
if (p->sparseBlocks > 0) {
// The SM-sparse wrapper: a grid of sparseBlocks blocks, the trailing argument the dispatch's nonce count (after
// the hot table when the pack has one); a dispatch smaller than the grid leaves the surplus blocks idle.
if (nonces % block != 0u) { err = "nonces not a multiple of the block"; return false; }
void* args[7] = { &p->ds, &out, &baseNonce, &mask, &a, &nonces, nullptr };
if (p->hot) { args[5] = &p->hot; args[6] = &nonces; }
DRV_CHECK(c, c.drv.launchKernel(p->fHashBound, (uint32_t)p->sparseBlocks, 1, 1, block, 1, 1, 0, s, args, nullptr), "cuLaunchKernel igneum_hash_bound (SM-sparse)");
return true;
}
void* args[6] = { &p->ds, &out, &baseNonce, &mask, &a, &p->hot };
DRV_CHECK(c, c.drv.launchKernel(p->fHashBound, nonces / block, 1, 1, block, 1, 1, 0, s, args, nullptr), "cuLaunchKernel igneum_hash_bound");
return true;
@ -514,6 +524,10 @@ struct Variant {
int maxrreg = 0; // 0: none; N: --maxrregcount=N (registers per thread, occupancy against spills)
int blockWarps = 0; // 0: the worker's --block-warps; N: 32 x N threads per block
int minBlocks = 0; // N > 0: __launch_bounds__(32 x blockWarps, N) (the compiler fits N blocks per SM)
int sparseBlocks = 0; // N > 0: the SM-sparse shape (Counter ASIC 4.0 research, 7 October 2026): a persistent grid of N
// blocks, every thread looping over the dispatch's nonces with stride gridDim.x * blockDim.x, so the
// hash runs on about N SMs at 32 warps per block and the other SMs idle; the kernel gains a
// trailing `uint32_t nonces` argument. Opt-in by name only (sp<N> or sp<N>-w<W>), never in a race by default
};
// The catalogue. Names are stable: the tuning file and the fleet records use them. "base" is the pack's text as
@ -541,8 +555,32 @@ static std::vector<Variant> allVariants() {
return v;
}
// sp<N> and sp<N>-w<W>: the SM-sparse variants, made on demand (not in the catalogue, so a default race never runs them):
// N persistent blocks (1 to 4096) of W warps (default 32, the largest block, one block per SM at the hash's register count).
static bool parseSparseVariant(const std::string& name, Variant& out) {
if (name.rfind("sp", 0) != 0) return false;
size_t i = 2, n = 0; int blocks = 0, warps = 32;
while (i < name.size() && name[i] >= '0' && name[i] <= '9') { blocks = blocks * 10 + (name[i] - '0'); ++i; ++n; }
if (n == 0 || blocks < 1 || blocks > 4096) return false;
if (i < name.size()) {
if (name.compare(i, 2, "-w") != 0) return false;
i += 2; n = 0; warps = 0;
while (i < name.size() && name[i] >= '0' && name[i] <= '9') { warps = warps * 10 + (name[i] - '0'); ++i; ++n; }
if (n == 0 || i != name.size() || warps < 1 || warps > 32) return false;
}
out = Variant(); out.name = name; out.blockWarps = warps; out.sparseBlocks = blocks;
return true;
}
static const Variant* findVariant(const std::vector<Variant>& all, const std::string& name) {
for (const Variant& v : all) if (v.name == name) return &v;
static std::vector<Variant> made; // the on-demand sparse variants, kept so the pointer stays valid
Variant sp;
if (parseSparseVariant(name, sp)) {
for (const Variant& v : made) if (v.name == name) return &v;
made.reserve(64);
if (made.size() < 64) { made.push_back(sp); return &made.back(); }
}
return nullptr;
}
@ -576,6 +614,42 @@ static bool variantSource(const std::string& base, const Variant& v, int blockWa
if (p == std::string::npos) { why = "no kernel declaration anchor"; return false; }
out.replace(p, std::strlen(a), fmt("__global__ void __launch_bounds__(%d, %d) igneum_hash_bound(", 32 * blockWarps, v.minBlocks));
}
if (v.sparseBlocks > 0) {
// The SM-sparse shape: the emitted kernel becomes a per-thread unit function taking its gid, and a persistent
// wrapper of the kernel's name loops the unit over the dispatch's nonces. Anchors: the kernel declaration as
// igneum-pow emits it (no __launch_bounds__ on it: the two rewrites are not combined) and its first line.
if (v.minBlocks > 0) { why = "sp and __launch_bounds__ minBlocks are not combined"; return false; }
const char* a = "__global__ void igneum_hash_bound(";
size_t p = out.find(a);
if (p == std::string::npos) { why = "no kernel declaration anchor"; return false; }
size_t close = out.find(") {", p);
if (close == std::string::npos) { why = "no parameter list close"; return false; }
std::string params = out.substr(p + std::strlen(a), close - (p + std::strlen(a)));
if (params.find("scratch") != std::string::npos || params.find("units") != std::string::npos) { why = "a persistent (variant-5) pack has its own loop"; return false; }
// the argument names: the last token of each parameter
std::string args; size_t from = 0;
while (from <= params.size()) {
size_t comma = params.find(',', from);
std::string one = params.substr(from, comma == std::string::npos ? std::string::npos : comma - from);
size_t e = one.find_last_not_of(" \t");
if (e == std::string::npos) { why = "an empty parameter"; return false; }
size_t b = one.find_last_of(" \t*&", e);
std::string nm = one.substr(b == std::string::npos ? 0 : b + 1, e - (b == std::string::npos ? 0 : b + 1));
if (nm.empty()) { why = "a parameter without a name"; return false; }
args += (args.empty() ? "" : ", ") + nm;
if (comma == std::string::npos) break;
from = comma + 1;
}
const char* gidLine = "\n uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;\n";
size_t g = out.find(gidLine, close);
if (g == std::string::npos) { why = "no gid line after the declaration"; return false; }
out.replace(g, std::strlen(gidLine), "\n");
out.replace(p, close - p, "__device__ __forceinline__ void igneum_hash_bound_unit(" + params + ", uint32_t gid");
out += fmt("\n// SM-sparse wrapper (variant %s): %d persistent blocks of %d threads over the dispatch's nonces\n", v.name.c_str(), v.sparseBlocks, 32 * blockWarps);
out += fmt("__global__ void __launch_bounds__(%d) igneum_hash_bound(", 32 * blockWarps) + params + ", uint32_t nonces) {\n";
out += " for (uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; gid < nonces; gid += gridDim.x * blockDim.x)\n";
out += " igneum_hash_bound_unit(" + args + ", gid);\n}\n";
}
return true;
}
@ -643,7 +717,9 @@ struct RaceEntry {
static bool raceTime(Ctx& c, Pair* p, const PfPack& pk, RaceEntry& e, CUdeviceptr dOut, uint32_t batch, int benchMs, CUstream s, bool selfTest) {
std::lock_guard<std::mutex> hold(gpuMutex);
CUfunction keep = p->fHashBound;
int keepSparse = p->sparseBlocks;
p->fHashBound = e.fn;
p->sparseBlocks = e.v.sparseBlocks;
std::string err;
uint32_t block = 32u * (uint32_t)e.blockWarps;
bool ok = true;
@ -673,6 +749,7 @@ static bool raceTime(Ctx& c, Pair* p, const PfPack& pk, RaceEntry& e, CUdevicept
}
}
p->fHashBound = keep;
p->sparseBlocks = keepSparse;
return ok;
}
@ -781,9 +858,10 @@ static void racePair(Ctx& c, Pair* p, const PfPack& pk, const std::string& bound
RaceEntry& w = entries[win];
c.drv.moduleUnload(p->modBound);
p->modBound = w.mod; p->fHashBound = w.fn; p->regs = w.regs; p->blocksPerSM = w.blocksPerSM; p->blockWarps = w.blockWarps; p->variant = w.v.name;
p->sparseBlocks = w.v.sparseBlocks;
w.mod = nullptr;
} else {
p->blockWarps = c.blockWarps; p->variant = "base";
p->blockWarps = c.blockWarps; p->variant = "base"; p->sparseBlocks = 0;
}
for (size_t i = 1; i < entries.size(); ++i) if (entries[i].mod) c.drv.moduleUnload(entries[i].mod);
p->raceMs = wallMs() - t0;
@ -1026,7 +1104,8 @@ static void usage() {
" --race-bench-ms N timed window per variant (default 2000)\n"
" --race-budget-s N a race stops compiling and timing after this (default 120; base is kept)\n"
" --race-rounds N interleaved rounds, best per variant (default 1 in --serve, 3 in --race)\n"
" --variant <name> use this variant without a race (also from the tuning file)\n"
" --variant <name> use this variant without a race (also from the tuning file); sp<N> or sp<N>-w<W> is the\n"
" SM-sparse shape (N persistent blocks of W warps, default 32; Counter ASIC 4.0 research)\n"
" --tuning <file> the per-card tuning file (default: IGNEUM_TUNING_FILE from the environment)\n", WORKER_VERSION);
}

View file

@ -15,3 +15,5 @@ docs/analysis/ci-failures-2026-10-06.md
# 7 October 2026: the last research round (the mission lanes and the closed list): internal research written for the
# owner, quoting his words and the operations record; the public spec mirror carries none of it
docs/analysis/mission
# 7 October 2026 (night): the Counter ASIC 4.0 research record (an internal research document: the founder's words, box names, lane records)
docs/analysis/counter-asic-4-research.md