From 03871c83594c0550ca6dfb1360b89d6e02b53c80 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 05:16:12 +0000 Subject: [PATCH] Counter ASIC 4.0 research: 20.3b the SM-sparse reading (a quarter of the SMs holds 98 percent of the rate at the same draw; the draw follows the work, not the SM count): rank 2 dead on measured rows Co-Authored-By: Claude Fable 5.1 --- docs/analysis/counter-asic-4-research.md | 21 +++++++++++++++++++-- 1 file changed, 19 insertions(+), 2 deletions(-) diff --git a/docs/analysis/counter-asic-4-research.md b/docs/analysis/counter-asic-4-research.md index bd5a892e4..b80247e2a 100644 --- a/docs/analysis/counter-asic-4-research.md +++ b/docs/analysis/counter-asic-4-research.md @@ -301,7 +301,7 @@ So the inverse lever's best design is the shadow per load: it costs the honest c | Rank | Design | Chip edge per joule (GDDR7; 1,400 lock) | Premium: 5090 / 4070 | Verifier | Agent hours | Breaks on | Change | |---|---|---|---|---|---|---|---| | 1 | The operating point as the shipped default | class v3 3.6x; class v4 2.1x at k = 1 | lowers the base (330 to 223 W); the v4 premium 145 to 82 W | none | 2 to 4 | AMD and Apple have no lock | unchanged | -| 2 | The SM-sparse miner kernel | unmeasured; toward 2.8x (v3) and 1.7x (v4) if half the SM-side 99 W is reachable | none | none | 1 + 1 + 4 | the 99 W is clock tree and leakage | unchanged; the PC 1 job is queued | +| 2 | The SM-sparse miner kernel | **DEAD on measured rows (20.3b, 05:12 UTC): a quarter of the SMs holds 98 percent of the rate at the same draw; the draw follows the work, not the SM count** | none | none | 1 + 1 + 4 | measured: the energy per hash never falls below base | closed | | 3 | **The tensor shadow with a SIMD byte-dot verifier** (int8 tiles carrying the premium; a class change) | at the ALU shadow's premium: 2.1x at k = 1, 1.6x at k = 1.5, 1.3x at k = 2; the chip's k 0.3 downside removed | the same as class v4 by construction / the 4070 tensor-bound above about R 600 (approximate) | scalar FAILS; AVX2 and VNNI 2.9 to 5.1 ms on the M5 Max core, 7 to 13 ms on a 2019 core (approximate, unwritten) | 8 to 12 (the SIMD verifier, the AMD layout gate, Apple's emulation row) plus the six gates | the 2019-class core; the AMD layout; Apple loses 5 to 10 percent of rate on emulation; k near 1 is claimed, not measured | NEW this pass: the k-above-1 candidate | | 4 | The shadow per load (capex) | unchanged in energy | none by construction (block-size effect unmeasured at 16) | unchanged | 4 to 6 plus the gates | compile-ahead at 16 sites; a class change; **the 16 x 27 construction is DEAD (20.2a-close, 22:53 UTC): the acceptance rule in execution order accepts 1.4 percent of its candidates; the sound form (one pass of 432 per load) is undrawn** | NEW: doubles the break-even cap to about USD 200 M, on a construction not yet shown to exist | | 5 | The ALU shadow re-weighted toward shuffles and multiplies | 2.1x at k = 1; 2.9x at the pessimistic k 0.46 | unchanged | unchanged | 4 to 6 plus the gates | Apple pays shfl 1.91x | down from 3: the tensor tile bounds k better | @@ -402,6 +402,23 @@ The hash lane's run `run-ca4-pc1-ca4sparse-5090-20261007` (the ca4sparse exe, th What the run did read, and keeps (the 5090 at the 1,300 MHz lock, base kernel, measured): class v4 134.26 MH/s at 309.9 W (0.433 MH/W, SM 1,290 MHz), class v3 134.03 at 219.4 W (0.611 MH/W); the class v4 premium 90.5 W at the lock against 133.9 W unlocked on this run (145.3 on the efficiency pass; the unlocked v4 rows ran hot, 465 to 485 W). The premium per counted op at the lock: 90.5 W over 134.26 M x 102,100 = 6.6 pJ (6.5 on the efficiency pass). +### 20.3b The SM-sparse reading (PC 1, the 5090 alone, 04:35:49 to 05:12:48 UTC, 8 October 2026 (05:35 to 06:12 UK); the hash lane's job `run-ca4-pc1-ca4sparse-5090-20261008-c` on the fourth exe, sha256 a4550202...): the candidate is dead + +The run as a gate: the card-free check on the card's own exe read `variants=2 names=base,sp43-w32`, `rewrite=applied`, `call="igneum_hash_bound_unit(ds, out, baseNonce, mask, iw, gid)" names_match=1`, `nvrtc64_120_0.dll sm_120 compiled=1 image_bytes=31384`; every sparse row served its variant (`served=sp-w32 sparse_blocks=N`, no "compile:" text) and every fingerprint on every row equals the Mac's (class v4 e370fb2080b7dbb1, class v3 90f794dd556f7a3b), so the rewritten persistent kernel is bit-exact. 32-second rows, the three power fields agreeing, idle 74 W; the drift check at the end repeats the unlocked rows within 1 W and 0.2 MH/s. + +| Variant (blocks of 32 warps, about one per SM) | Class v4 unlocked: MH/s / W / MH/W | Class v3 unlocked | Class v4 at the 1,300 lock | Class v3 at the 1,300 lock | Label | +|---|---|---|---|---|---| +| base (the plain grid, one warp per block) | 137.07 / 450.8 / 0.304 | 136.77 / 313.9 / 0.436 | 134.31 / 301.0 / 0.446 | 134.04 / 211.4 / 0.634 | measured | +| sp170-w32 (the persistent shape on every SM) | 136.09 / 461.2 / 0.295 | 136.02 / 313.3 / 0.434 | 128.95 / 297.6 / 0.433 | 129.20 / 208.2 / 0.621 | measured | +| sp85-w32 (half the SMs) | 136.29 / 461.7 / 0.295 | 136.87 / 310.7 / 0.441 | 125.37 / 290.1 / 0.432 | 129.80 / 207.8 / 0.625 | measured | +| **sp43-w32 (a quarter)** | **134.58 / 460.1 / 0.293** | **136.55 / 309.8 / 0.441** | 64.33 / 208.2 / 0.309 | 126.25 / 205.4 / 0.615 | measured | +| sp21-w32 (an eighth) | 70.75 / 327.4 / 0.216 | 132.82 / 303.4 / 0.438 | 31.48 / 155.2 / 0.203 | 84.78 / 174.9 / 0.485 | measured | +| sp11-w32 (a sixteenth) | 37.34 / 250.9 / 0.149 | 100.14 / 274.5 / 0.365 | 16.55 / 116.7 / 0.142 | 44.57 / 147.5 / 0.302 | measured | + +The two readings the design asked for. (1) The shape itself: sp170-w32 against base costs 0.7 percent of rate on class v4 and 0.5 on v3, +10 W on v4 and 0 W on v3: within noise. (2) Watts minus idle per MH/s against base: class v4 unlocked base 2.75 W per MH/s, sp43 2.87, sp21 3.58, sp11 4.74; class v3 unlocked base 1.75, sp43 1.73, sp21 1.73, sp11 2.00. A quarter of the SMs holds 98.2 percent of the class v4 rate at the SAME draw (460 W against 451) and 99.8 percent of the class v3 rate at 4 W less; the draw falls only when the rate falls, and the energy per hash never falls below base on either class. **So the SM-side 99 W of section 2 is not activity that idle SMs would save: the card's draw follows the work, not the SM count; the ALU work per hash costs the same on 43 SMs as on 170, and the 43 SMs run it at the same energy.** The candidate is dead by its own rule (watts track the work one for one), and the lock rows say the same from the other side (class v4 sp43 at 1,300 MHz is compute-bound at 64 MH/s: the shadow's ops are real throughput work). The one number kept for the chip model: the class v4 premium at sp43 unlocked, 150.3 W over class v3 at a held rate, equal to the full-card premium (136.9 W base, 151.2 on the packs job), so the shadow's energy does not depend on how many SMs carry it; it is the energy of the ops. + +What this closes in the ranking: rank 2 (the SM-sparse miner kernel) is dead; the premium-free floor of section 0 stays the card's idle plus its memory system plus whatever `E_wait` is at the knee, and the only miner-side lever on `E_card` is the operating point (rank 1). The 5090's four runs of this night (three failed, one read) also gave four repeats of the knee pass: the class v4 premium 133.9 to 151.2 W unlocked and 82.8 to 90.5 W at 1,300 MHz, with every one of the three power fields agreeing within 0.2 W. + ### 20.4 Into the chip model (the rows measured; the chip side modelled: f = 1 GDDR7 0.466 microjoules per hash, one HBM3 stack 0.321) Energy per hash = the card's measured; the chip's = `E_mem + k x F`, `F` the measured premium per hash, `k` the chip's cost per forced operation over the 5090's at the same state. @@ -418,4 +435,4 @@ Energy per hash = the card's measured; the chip's = `E_mem + k x F`, `F` the mea The k column was priced at this point: section 15.2 set the tile count from the 4090's marginal 0.056 pJ per MAC so that the block would carry the ALU shadow's 0.654 microjoules (11,400 tiles per hash, R about 1,430), and the exported pack is 1,430 tiles per iteration, 11,440 per hash; the 5090 reads 0.091 pJ per MAC unlocked and 0.048 at the lock at that point, so the GPU-side cost of the chip-model-v3 5.11 column (2.1x at k = 1, 1.6x at k = 1.5) is now measured at the premium it was priced for (1.067 microjoules unlocked, 0.564 at the lock, against the ALU shadow's 1.10 and 0.652). The open side is Apple: the Metal emulation costs 35 percent of the M5 Max's rate at 1,024 tiles per hash and 78 percent at 4,096, so the 11,440-tile point is out of the Apple tier's reach without an integer matrix path in Metal; nothing served moves (the chip texts rest on the ALU shadow). Reading: at `k = 1` the tile block and the ALU shadow give the same 2.1x to 2.2x, and the tile block's premium at the knee is 14 percent lower for it; the difference is the chip's reachable `k`. On the ALU shadow a fixed SIMD array can plausibly reach `k` 0.3 to 0.5 (the served pessimistic 3.4x to 3.5x); on the int8 tile the GPU's own tensor core is the dense array, measured at 0.048 to 0.091 pJ per MAC, and a chip at the same node cannot be 2x to 3x cheaper per MAC at the same voltage, so the chip's downside column collapses: the tile block's edge at the measured premium reads 2.2x at `k = 1` and under 1.8x at any `k` over 1.3, against the ALU shadow's 3.5x at `k` 0.3. That is the finding of the k hunt made measured on the card: **the tile block does not beat class v4's premium (it matches it at the same rate, 14 percent cheaper at the knee) and it beats class v4's chip edge at the pessimistic end (about 2.2x against 3.5x) and not at `k = 1`.** What stands against it as a class: the verifier (AVX2 0.047 microseconds per tile per unit, `mm1430` 5.8 ms alone and 10.14 ms with the sibling loaded on the box's core, a 0.14 ms miss at load 84 to 95; scalar 13x worse; NEON unwritten), the Apple tier (the Metal emulation at 1,024 tiles costs 35 percent of the M5 Max's rate and 4,096 costs 78 percent, so 11,440 is out of reach without an integer matrix path in Metal), the AMD WMMA layout (unverified; the OpenCL reference path emitted but unmeasured on the 9070 XT), and the compile-ahead (1.2 to 2.1 s per pack on the 5090, inside the budget). -**The first sentence, on measured rows (04:22 UTC, 8 October 2026): neither prototype beats class v4's premium; the tile block matches it at the same hash rate (146.7 W against 151.2 W unlocked, 71.5 against 82.8 W at the 1,300 lock) and beats class v4's chip edge only at the pessimistic end (about 2.2x against 3.5x, because a chip's int8 MAC cannot undercut the 5090's measured 0.048 to 0.091 pJ per MAC the way a fixed datapath undercuts its 6.4 to 10.8 pJ per ALU op), not at k = 1 (2.2x either way); the per-load placement is dead as a construction, and its energy rows read 13 to 14 W under the whole block for the same instructions.** The SM-sparse reading is still owed (three failed runs: the race switch, the race list, the rewrite's capture; the fourth exe checked through NVRTC card-free, queued). +**The first sentence, on measured rows (04:22 UTC, 8 October 2026): neither prototype beats class v4's premium; the tile block matches it at the same hash rate (146.7 W against 151.2 W unlocked, 71.5 against 82.8 W at the 1,300 lock) and beats class v4's chip edge only at the pessimistic end (about 2.2x against 3.5x, because a chip's int8 MAC cannot undercut the 5090's measured 0.048 to 0.091 pJ per MAC the way a fixed datapath undercuts its 6.4 to 10.8 pJ per ALU op), not at k = 1 (2.2x either way); the per-load placement is dead as a construction, and its energy rows read 13 to 14 W under the whole block for the same instructions.** The SM-sparse reading landed at 05:12 UTC after three failed runs (the race switch, the race list, the rewrite's capture) and the candidate is dead: a quarter of the SMs holds 98 percent of the rate at the same draw (20.3b).