From 4072cd7f077881aaf2e0bf53a4aee245cf81a9ad Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Wed, 7 Oct 2026 22:53:36 +0000 Subject: [PATCH] Counter ASIC 4.0 research: 20.2a-close, the per-load 16 x 27 construction dead as a chain class (1.4 percent acceptance, 42 of 64 seeds without a program), the sound form named, rank 4 marked Co-Authored-By: Claude Fable 5.1 --- docs/analysis/counter-asic-4-research.md | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/analysis/counter-asic-4-research.md b/docs/analysis/counter-asic-4-research.md index 8f1c6746c..d87f6dfd2 100644 --- a/docs/analysis/counter-asic-4-research.md +++ b/docs/analysis/counter-asic-4-research.md @@ -303,7 +303,7 @@ So the inverse lever's best design is the shadow per load: it costs the honest c | 1 | The operating point as the shipped default | class v3 3.6x; class v4 2.1x at k = 1 | lowers the base (330 to 223 W); the v4 premium 145 to 82 W | none | 2 to 4 | AMD and Apple have no lock | unchanged | | 2 | The SM-sparse miner kernel | unmeasured; toward 2.8x (v3) and 1.7x (v4) if half the SM-side 99 W is reachable | none | none | 1 + 1 + 4 | the 99 W is clock tree and leakage | unchanged; the PC 1 job is queued | | 3 | **The tensor shadow with a SIMD byte-dot verifier** (int8 tiles carrying the premium; a class change) | at the ALU shadow's premium: 2.1x at k = 1, 1.6x at k = 1.5, 1.3x at k = 2; the chip's k 0.3 downside removed | the same as class v4 by construction / the 4070 tensor-bound above about R 600 (approximate) | scalar FAILS; AVX2 and VNNI 2.9 to 5.1 ms on the M5 Max core, 7 to 13 ms on a 2019 core (approximate, unwritten) | 8 to 12 (the SIMD verifier, the AMD layout gate, Apple's emulation row) plus the six gates | the 2019-class core; the AMD layout; Apple loses 5 to 10 percent of rate on emulation; k near 1 is claimed, not measured | NEW this pass: the k-above-1 candidate | -| 4 | The shadow per load (capex) | unchanged in energy | none by construction (block-size effect unmeasured at 16) | unchanged | 4 to 6 plus the gates | compile-ahead at 16 sites; a class change | NEW: doubles the break-even cap to about USD 200 M | +| 4 | The shadow per load (capex) | unchanged in energy | none by construction (block-size effect unmeasured at 16) | unchanged | 4 to 6 plus the gates | compile-ahead at 16 sites; a class change; **the 16 x 27 construction is DEAD (20.2a-close, 22:53 UTC): the acceptance rule in execution order accepts 1.4 percent of its candidates; the sound form (one pass of 432 per load) is undrawn** | NEW: doubles the break-even cap to about USD 200 M, on a construction not yet shown to exist | | 5 | The ALU shadow re-weighted toward shuffles and multiplies | 2.1x at k = 1; 2.9x at the pessimistic k 0.46 | unchanged | unchanged | 4 to 6 plus the gates | Apple pays shfl 1.91x | down from 3: the tensor tile bounds k better | | 6 | The L2-resident hot table with cache hints | energy 0.1 microjoules at k 0.7 to 3; capex +USD 15 | 13 W if the hint holds the rate | under 0.5 ms (unmeasured) | 4 for the ldcs measurement (queued) | Metal has no hint; the rate without it | unchanged; a measurement | | 7 to 11 | class v5 without the shadow; the refresh as a cost; per-card classes; proof of useful work; memory shaping | as section 9 | | | 0 | dead by arithmetic | unchanged | @@ -365,6 +365,10 @@ The tile cost law from these rows: AVX2 (5.82 - 5.33) ms over (1,430 - 128) x 8 Two readings from the Metal row that move section 15.2's tile verdict. First, the reference tile (`mm8_ref`: 12 index shuffles and 32 byte products per lane) is bit-exact against the Rust verifier on all three vector warps of both tile packs, so the tile semantics are pinned on two implementations before any card runs the PTX form. Second, the Apple cost is far above the "10 ALU steps per tile" estimate: 1,024 tiles per hash cost the M5 Max 35 percent of its rate and 4,096 tiles 78 percent, so a tile shadow at the ALU shadow's premium (11,440 tiles per hash) would take the Apple tier out entirely unless Metal gains an integer matrix path reachable from the toolchain (`mpp::tensor_ops::matmul2d` with `uchar` operands, unverified). The tile shadow therefore stands as a class only with an Apple exemption nobody has designed, which moves it from rank 3 to beside rank 5 until that path is measured; the k question it answers is unchanged. +### 20.2a-close The per-load construction, closed (22:44 to 22:53 UTC): dead as a chain class + +The fix of 20.2a held for distinctness and then met the value-level requirement of 20.2b, and the construction did not survive it. With the acceptance rule stepping the per-load sub-blocks in the order the class executes and judging both the duplicate lanes at a load row and the one-count of every index bit per site over the 64 units (6-sigma band), the 16 x 27 per-load class accepts 22 of 1,621 candidates over 64 seeds (1.4 percent); 42 of 64 seeds exhaust the chain's 32 attempts, which on the chain is an epoch without a program. The first failing test per candidate: a biased index bit 775, duplicate lanes 643, the base rule's lane-constant site 110, (b) 43, (a) 28. The genesis seed accepts none of 32; candidate 0 of the class carries index bit 0 set in 40 of 1,024 addresses at site 0 (z 29.5; the sub-block writer a `rotl` of a `mad` result). The structural reason, read from the record: 27 passes of a 16-instruction map immediately before a load is a tight iteration of a small function, and whatever lossy or product arithmetic it carries (or inherits from the base writer before it) collapses or biases the load's address register before any base instruction can re-randomise it; the class v4 shape places the same 6,912 instructions after instruction 63, where the next iteration's 64 base instructions and 16 loads stand between the block and every load. Both exports (854050a4293f0615, bd64b207a30413fb) were accepted only because the rule did not model the placement; their PC 1 rows stay as the energy reading of the placement, labelled "unsound construction, energy reading only". Design 4 (the shadow per load) and the USD 200 M capex row of 16.2 therefore rest on a construction that does not exist yet; the sound form to try is one pass of a 432-instruction sub-block per load (or 64 x 7), the same N, where the sub-block is a program segment rather than an iterated map; it is a new class to draw, accept and measure (kernel text 6,912 lines per iteration against the 1,024-line block's measured 17 percent on the M5 Max, so the footprint is its own gate), not tonight's. The class v4 shape across the same 16 drawn eras reads 0 duplicate pairs and carries the adv-cache-2 product bit at address bit R exactly in 14 of 17 eras on this pre-amendment generator (one-count 250 or 780 of 1,024, z 15 to 19; sites with `mul`, `rotl` and `or` writers), which is that lane's finding replicated by a second instrument. + ### 20.2b A named requirement for every CA4 prototype: no biased product bits in an address (the crypto lane's adv-cache-2 reading, 7 October 2026, 22:1x UTC) The crypto lane's finding: a product's low bits are biased (P(bit 0) = 1/4, measured exactly), the bias survives the odd stride multiplier, and the stride rotation places the biased bits at address bits R and up, inside the 28-bit item index unless R is 28 or more. The devnet era draws R = 29, which cuts them off, so 31 of 32 devnet-era programs read clean while 6 of 16 drawn-era programs (R from 3 to 22) show a site over 1.04x (2 over 1.2x, the worst 1.51x); under the 2 GiB genesis dataset (D = 29) R = 29 would show it too. The devnet's cleanliness is an accident of its era draw; the chain prevalence is the drawn-era figure; the price to a partial-store chip stays under 0.1 percent of a hash's reads per site, so no chip number moves. The requirement for any class this file proposes (the per-load shadow, the tile block, a re-weighted shadow): (1) a load whose source register's last writer is a product (`mul`, `mulhi`, `mad`) carries biased low bits into the address, and the acceptance must judge it at the VALUE level (the bit bias of the index at the site over the units), not by the index-distinctness ratio alone, which the duplicate-lane test above is; (2) the census of any candidate reads across the drawn eras split by R (3 to 22 against 28 to 31), as 20.2a now does for the per-load class, never the devnet era alone. The per-load class's dynamic rule covers distinctness, not bias; the value-level test is owed and is the same item for class v5's acceptance. Nothing in class v4 or v5 moves on this without main's word.