diff --git a/docs/analysis/counter-asic-4-research.md b/docs/analysis/counter-asic-4-research.md index 34d8b9823..a6a13ddb4 100644 --- a/docs/analysis/counter-asic-4-research.md +++ b/docs/analysis/counter-asic-4-research.md @@ -1,6 +1,6 @@ # Counter ASIC 4.0 research: can the chip's per-joule edge be held near 2x without the GPU paying an energy premium? -7 October 2026, 20:27 to 23:2x UTC (21:27 to 00:2x UK), branch `counter-asic-4`, the Counter ASIC 4.0 research lane. Written for the founder's question of this evening: class v4 (the latency shadow, about 100,000 counted integer operations per hash hidden in the memory wait) brings the strongest chip's per-joule edge over an RTX 5090 from 5x to 9x down to 2.1x to 3.9x, and costs the 5090 145 W at its stock clock and 88 W at a 1,400 MHz core lock for the same hash rate. The brief: find the design that keeps the chip's edge at or under about 2x per joule WITHOUT the GPU paying an energy premium, or prove that no such design exists and say what the floor is. +7 October 2026, 20:27 to 21:0x UTC (21:27 to 22:0x UK), branch `counter-asic-4`, the Counter ASIC 4.0 research lane. Written for the founder's question of this evening: class v4 (the latency shadow, about 100,000 counted integer operations per hash hidden in the memory wait) brings the strongest chip's per-joule edge over an RTX 5090 from 5x to 9x down to 2.1x to 3.9x, and costs the 5090 145 W at its stock clock and 88 W at a 1,400 MHz core lock for the same hash rate. The brief: find the design that keeps the chip's edge at or under about 2x per joule WITHOUT the GPU paying an energy premium, or prove that no such design exists and say what the floor is. Every number carries a label: **measured** (a card on a named night, the file named), **modelled** (arithmetic on the chip model's cited memory figures, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling estimate). Nothing in this file changes a consensus object, a pack or a parameter; the one code change (section 12) is a miner-side kernel variant that is opt-in by name and never raced by default. @@ -19,7 +19,7 @@ The floor, on tonight's measured rows (the hash lane's efficiency pass, PC 2, 7 | RTX 5090, core unlocked (136.59 MH/s at 330.2 W) | 2.42 microjoules | 5.2x / 7.5x / 9.2x | measured card, modelled chip | | RTX 5090, app stock point, 5 October (124 MH/s at 290 W) | 2.34 | 5.0x / 7.3x / 8.9x | measured, modelled | | **RTX 5090, 1,400 MHz core lock (134.68 MH/s at 228.0 W)** | **1.69** | **3.6x / 5.3x / 6.5x** | measured, modelled | -| RTX 5090, the knee (the hash lane's second pass, 22:1x UTC): class v3 best at 1,300 MHz, 134.62 MH/s at 223.3 W; class v4 best at 1,200 MHz, 133.80 MH/s at 305.1 W; the premium at the best points 81.8 W; the floor below 1,100 MHz unmeasured (a third pass runs) | 1.66 (v3); 2.28 (v4) | 3.6x / 5.2x / 6.3x on v3; class v4 at k = 1 (6.1 pJ per counted op) 2.1x | measured, modelled | +| RTX 5090, the knee (the hash lane's second pass, 20:5x UTC): class v3 best at 1,300 MHz, 134.62 MH/s at 223.3 W; class v4 best at 1,200 MHz, 133.80 MH/s at 305.1 W; the premium at the best points 81.8 W; the floor below 1,100 MHz unmeasured (a third pass runs) | 1.66 (v3); 2.28 (v4) | 3.6x / 5.2x / 6.3x on v3; class v4 at k = 1 (6.1 pJ per counted op) 2.1x | measured, modelled | | RTX 5090, the bound any operating point can reach: idle 74 W (PC 2) plus the memory system's 55 W at 17.5 G reads/s, 135 MH/s | 0.96 | 2.1x / 3.0x / 3.7x | approximate (idle measured; the memory share modelled) | | RTX 4070 at its Ember tune point (30.95 MH/s at 79.5 W) | 2.57 | 5.5x / 8.0x / 9.8x | measured, modelled | | RX 9070 XT (about 18.6 MH/s at about 304 W) | 10.6 | 23x / 33x / 40x | approximate card figure | @@ -31,7 +31,7 @@ So the premium-free floor against the GDDR7 chip is 3.6x on a 5090 locked at 1,4 1. **The operating point as the shipped default** (miner-only; core lock at the knee plus undervolt where the driver allows; AMD through ADLX or nothing; Apple has no lever). Premium: none by construction, it lowers `E_card`. Measured: v3 2.42 to 1.69 microjoules, the v4 premium 145 to 88 W, the class v4 edge at `k = 1` 2.1x. 2 to 4 agent hours (already ordered into Ember Tune for 0.3.24). Breaks: a locked card's free shadow shrinks (about 150,000 ops at 1,400 MHz), so ladder rung 2 is not free on it. 2. **The SM-sparse miner kernel** (miner-only; the hash on a fraction of the SMs with 32 warps per block and the rest clock-gated). Reads the unmeasured 100 W between the 5090's idle-plus-memory (about 130 W) and its 228 W at the lock; if half is SM activity the premium-free edge falls toward 2.5x. Built tonight as worker variants `sp-w` (section 12), one PC 1 job owed. Breaks if the 100 W is clock tree and leakage. -3. **The shadow kept at rung 0, its op mix re-weighted toward the families where a chip gains least over a GPU lane** (shuffle and multiply heavy). Same premium, the pessimistic-core edge 3.4x to about 2.9x, `k = 1` unchanged at 2.1x. 4 to 6 agent hours plus the six gates (a class change under the 95 percent rule). Held until the knee and the SM-sparse rows are read (the coordinator, 22:0x UK). +3. **The shadow kept at rung 0, its op mix re-weighted toward the families where a chip gains least over a GPU lane** (shuffle and multiply heavy). Same premium, the pessimistic-core edge 3.4x to about 2.9x, `k = 1` unchanged at 2.1x. 4 to 6 agent hours plus the six gates (a class change under the 95 percent rule). Held until the knee and the SM-sparse rows are read (the coordinator, 21:5x UK). Dead by arithmetic (sections 3 to 7): dropping the shadow after class v5 (v5 leaves the chip at 5.1x to 9.1x), the per-window refresh as a cost (a 1 GiB rebuild is about 0.1 J on a chip core), extra reads or longer chains (both sides scale together), memory-system shaping (the same DRAM on both sides), per-card adaptive classes (a chip picks the class that maximises its own reward per joule, so a menu never beats the card's best single class), proof of useful work (the useful fraction is bounded at 8 percent of the hash budget at 1 GH/s and falls with scale; no ZK ASIC has beaten a GPU per joule on real hardware in the public record; every deployed useful-work scheme was gamed), and the tensor block (`k` near 1 but 0.056 pJ per multiply-add, so there are no joules in it to force without 26x the verifier cost). @@ -198,15 +198,15 @@ The premium cannot reach zero against a chip that is a GPU's memory system witho | A pool user | nothing changes in shares or payout from any design here; designs 3 to 5 would be class changes announced by the 95 percent signal | | | A node operator (the verifier) | designs 1 and 2 cost nothing; design 3 nothing; design 4 under 0.5 ms (unmeasured); design 5 needs a SIMD byte-dot path to fit the gate | | | A chip | its edge is `E_card / E_mem` plus whatever the shadow forces at its own `k`: 3.6x to 5.2x on GDDR7 with no shadow depending on the card's operating point, 2.1x with the shadow at k = 1; a core cheaper than 1.8 pJ per counted op would turn the shadow in its favour at the knee, which no public figure shows for a random 32-lane program | the `k` question stays with the external brief (algorithm.md proposal 8) | -| The public claim | the honest sentence is the floor and the premium as numbers: "a chip that stores the dataset keeps the card's whole-card energy over its memory energy, 3.6x on a locked 5090 and 1.7x on an Apple GPU with no extra work; the shadow holds it to 2.1x against a core as good as a GPU lane for 88 W on a 5090 at the knee" | to the Counter lane for the chip texts once the knee is a number (the coordinator, 22:0x UK) | +| The public claim | the honest sentence is the floor and the premium as numbers: "a chip that stores the dataset keeps the card's whole-card energy over its memory energy, 3.6x on a locked 5090 and 1.7x on an Apple GPU with no extra work; the shadow holds it to 2.1x against a core as good as a GPU lane for 88 W on a 5090 at the knee" | to the Counter lane for the chip texts once the knee is a number (the coordinator, 21:5x UK) | ## 12. What was built tonight: the SM-sparse worker variant -`proto-cuda/nvrtc/worker.cpp` (this branch) gains opt-in variants `sp` and `sp-w` (N persistent blocks of W warps, W default 32): the pack's bound kernel text is rewritten at compile time with exact anchors (the race's existing mechanism, `variantSource`): `__global__ void igneum_hash_bound(...)` becomes `__device__ __forceinline__ void igneum_hash_bound_unit(..., uint32_t gid)` with its `gid` line removed, and a persistent wrapper of the kernel's own name loops the unit over the dispatch's nonces with stride `gridDim.x * blockDim.x`, taking a trailing `uint32_t nonces` argument. `launchHash` launches it as a grid of N blocks; `raceTime` carries the entry's shape while timing it; the winner installs it. The variants are made on demand by name (`parseSparseVariant`) and are not in the catalogue, so a default race never runs them and no shipped worker changes behaviour. Refused on a persistent (variant-5) pack and when combined with a `__launch_bounds__` minBlocks variant. The rewrite was mirrored in Python on the class v4 devnet pack and reads as intended; the exe was cross-compiled on igneum-build-1 under `lease pool 4 --class measure --label "ca4 research: worker cross-compile (SM-sparse variant)"` (22:48 UTC; `/srv/builds/ca4-research/nvrtc/igneum-worker-cuda-ca4sparse.exe`, sha256 ee8d0e70dd101f125f42c0c7cf07481a794ee18a1317acf68561b37ea18d72be, no icon or version block: a scratch exe for one job, not a shipping one). Not yet run on a card: the first run's self-test (the vector warps through the bound kernel and the 2^24 fingerprint equal to the base kernel's) is its gate, and an NVRTC compile error shows as the variant dropped with "compile:" in the race line. The job is with the hash lane for PC 1 after the knee sweep (rungs sp170, sp85, sp43, sp21, sp11 at the 1,400 lock and unlocked, class v3 and v4 packs; MH/s, watts, SM MHz, fingerprint per row). +`proto-cuda/nvrtc/worker.cpp` (this branch) gains opt-in variants `sp` and `sp-w` (N persistent blocks of W warps, W default 32): the pack's bound kernel text is rewritten at compile time with exact anchors (the race's existing mechanism, `variantSource`): `__global__ void igneum_hash_bound(...)` becomes `__device__ __forceinline__ void igneum_hash_bound_unit(..., uint32_t gid)` with its `gid` line removed, and a persistent wrapper of the kernel's own name loops the unit over the dispatch's nonces with stride `gridDim.x * blockDim.x`, taking a trailing `uint32_t nonces` argument. `launchHash` launches it as a grid of N blocks; `raceTime` carries the entry's shape while timing it; the winner installs it. The variants are made on demand by name (`parseSparseVariant`) and are not in the catalogue, so a default race never runs them and no shipped worker changes behaviour. Refused on a persistent (variant-5) pack and when combined with a `__launch_bounds__` minBlocks variant. The rewrite was mirrored in Python on the class v4 devnet pack and reads as intended; the exe was cross-compiled on igneum-build-1 under `lease pool 4 --class measure --label "ca4 research: worker cross-compile (SM-sparse variant)"` (20:48 UTC; `/srv/builds/ca4-research/nvrtc/igneum-worker-cuda-ca4sparse.exe`, sha256 ee8d0e70dd101f125f42c0c7cf07481a794ee18a1317acf68561b37ea18d72be, no icon or version block: a scratch exe for one job, not a shipping one). Not yet run on a card: the first run's self-test (the vector warps through the bound kernel and the 2^24 fingerprint equal to the base kernel's) is its gate, and an NVRTC compile error shows as the variant dropped with "compile:" in the race line. The job is with the hash lane for PC 1 after the knee sweep (rungs sp170, sp85, sp43, sp21, sp11 at the 1,400 lock and unlocked, class v3 and v4 packs; MH/s, watts, SM MHz, fingerprint per row). ## 13. Unverified and owed -- The knee: read at 22:1x UTC as 1,300 MHz on class v3 (223.3 W, 1.66 microjoules) and 1,200 MHz on class v4 (305.1 W, 2.28), the premium 81.8 W at the best points; every "at the lock" row in this file is the 1,400 row and moves by under 3 percent to the knee (3.63x to 3.56x on GDDR7; the class v4 edge at k = 1 stays 2.1x). The floor below 1,100 MHz is unmeasured (a third pass runs to it). +- The knee: read at 20:5x UTC as 1,300 MHz on class v3 (223.3 W, 1.66 microjoules) and 1,200 MHz on class v4 (305.1 W, 2.28), the premium 81.8 W at the best points; every "at the lock" row in this file is the 1,400 row and moves by under 3 percent to the knee (3.63x to 3.56x on GDDR7; the class v4 edge at k = 1 stays 2.1x). The floor below 1,100 MHz is unmeasured (a third pass runs to it). - The breakdown of the 5090's 228 W at the lock into idle, memory and SM activity is unmeasured; the SM-sparse job reads it. - The RX 9070 XT's watts under class v4; the 2019-class verifier core (O-1.14); the M5 Max package watts. - The chip side is the model (chip-model-v3 section 5): activate-bound ceilings, 2.0 and 1.2 nJ per random read, the static and controller allowances, the 85 W node; no chip has been measured. The `k` band 0.5 to 1 for a SIMD-array core is approximate; the N5 datapath floor is a bound, not a design.