From 0e23ca3f89295b057a95419aaa3bd7eb5fab1bf3 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 12:28:21 +0000 Subject: [PATCH] counter-asic-4 documents: the branch's text at 086411627 taken whole for the documents-only landing (what the replayed commits did not carry: the branch's merge resolutions) --- docs/analysis/chip-model-v3.md | 13 ++++++++++++- docs/analysis/counter-asic-4-research.md | 8 ++++++-- tools/ci/export-exclude.txt | 4 ++++ 3 files changed, 22 insertions(+), 3 deletions(-) diff --git a/docs/analysis/chip-model-v3.md b/docs/analysis/chip-model-v3.md index 3dae344e7..82ea952f3 100644 --- a/docs/analysis/chip-model-v3.md +++ b/docs/analysis/chip-model-v3.md @@ -365,7 +365,18 @@ The k column. k is the chip core's energy per op over the GPU's at the same oper Correction, 8 October 2026 (the research lane's microbench on the 5090, counter-asic-4-research.md 15.1a and the corrected 20.3 and 20.4 at 71fd465b): the per-MAC figures above are wrong by a factor of 32. A `mma.m8n8k16` tile is 1,024 multiply-adds per warp, 32 per lane, so a hash does 32 MACs per tile, not 1,024; the 4090's "0.056 pJ per MAC" is 1.8 pJ, and the 5090 at the ALU shadow's premium reads 2.9 pJ per MAC unlocked and 1.5 pJ at the 1,300 MHz lock (the packs job, 366,080 MACs per hash; the microbench's dependent u8 tile 4.1 and 2.2, the wide s8 m16n8k32 tile 1.36 and 0.83). Against the same 5 nm MAC array figures (0.04 to 0.4 pJ per INT8-class MAC, claimed) a chip's k on tile work is therefore 0.03 to 0.3, below the ALU shadow's 0.3 to 0.8, not near 1: at the same premium a tensor-shaped shadow leaves the chip 3.5x to 6.7x where the ALU shadow leaves it 2.1x to 3.5x. The tensor-tile column (2.1x at k = 1, 1.6x at k = 1.5) is withdrawn as a candidate; its premise, that a chip's MAC is no cheaper than the GPU's, is false by 4x to 30x on the public figures. The shipped row is unchanged: 2.1x at k = 1 and 3.4x at k about 0.33, the ALU shadow at the operating point's knee, measured four times at 82 to 90 W. -The capex column. The `f = 1` GDDR7 chip of 5.5 is USD 2.8 per MH/s of silicon and memory, which is USD 0.00016 per MH/s-hour of capex over two years against USD 0.000023 of electricity: capex-dominated 7x, as the 5090 is (USD 14.7 per MH/s at MSRP, 10x). A 64 MiB hot table adds about USD 15 of N5 die, the shadow core USD 25 to 40, an interposer USD 200, so the chip's capex reaches at most about USD 4.3 per MH/s: the per-unit capex wall is unreachable by 3x to 7x, and the break-even market cap moves only through the project cost (the mission lane's model: about USD 100 M with the N5 shadow core, about 200 M if the shadow runs per load and forces one die or an interposer; the per-load form behind that figure, the 16 x 27 placement, was closed on 7 October 2026 at night when it failed the value-level acceptance test across drawn eras, so the 200 M row rests on no construction shown to exist until a sound per-load class, one pass of a 432-instruction sub-block per load, is drawn, accepted and measured). Every figure here is modelled on cited or claimed parts; the research lane's microbench (20 probes, the mma_u8 and l2 rows the ones this model would take) is on PC 1's queue after the hot-table job. +The capex column. The `f = 1` GDDR7 chip of 5.5 is USD 2.8 per MH/s of silicon and memory, which is USD 0.00016 per MH/s-hour of capex over two years against USD 0.000023 of electricity: capex-dominated 7x, as the 5090 is (USD 14.7 per MH/s at MSRP, 10x). A 64 MiB hot table adds about USD 15 of N5 die, the shadow core USD 25 to 40, an interposer USD 200, so the chip's capex reaches at most about USD 4.3 per MH/s: the per-unit capex wall is unreachable by 3x to 7x, and the break-even market cap moves only through the project cost (the mission lane's model: about USD 100 M with the N5 shadow core, about 200 M if the shadow runs per load and forces one die or an interposer; the per-load form behind that figure is REOPENED (8 October 2026, 11:0x UTC): the 16 x 27 iterated placement is dead on the no-era census (0.986 to 0.990 rejection per candidate on both instruments), the 7 October night's drawn-era death was an instrument artefact (the per-load acceptance test counted the era window's fixed top index bits as biased, fixed on counter-asic-4 at 10:5x UTC), and the sound form, one pass of a 256-instruction sub-block per load (`mx8+shl4096x1`), accepts 234 of 256 seeds on the no-era census with its drawn-era verdict the invention lane's 15:00 BST read; the 200 M row rests on that form, class v6's layer 5). Every figure here is modelled on cited or claimed parts; the research lane's microbench (20 probes, the mma_u8 and l2 rows the ones this model would take) is on PC 1's queue after the hot-table job. + +### 5.12 Two chips from the hardware-future lane, and the honest denominator (8 October 2026, 11:3x UK, main's order; the figures lane B's, `docs/analysis/class-v6/hardware-future.md` at master 34f63b3c, modelled on this file's method unless marked claimed; the research lane carries them here as the model's own rows) + +| Chip | When | Energy per random read | Rate per die | Energy per hash at zero shadow | Edge per joule against the 5090 at 2.40 microjoules | Against the M5 Max at 0.78 | With the class v4 shadow (F = 1.10 microjoules) at k = 0.5 / k = 1 | Silicon and project | What the four layers of class v6 do to it | Labels | +|---|---|---|---|---|---|---|---|---|---|---| +| **A 2 GiB SRAM full store on one N2 reticle** (TSMC N2 38 Mb/mm^2 HD SRAM, +11 percent over N3E, claimed, IEEE Spectrum, 12 December 2024; 2 GiB is about 452 mm^2 of macro under the 858 mm^2 reticle) | 2027 to 2028 (an N2 project; wafers about USD 30,000 and booked to 2028, claimed) | 1.0 nJ for a 64-byte read across the array (0.5 to 2.0: the wire term dominates, this file's 5.1 convention) | power-bound: about 2,100 MH/s at 300 W (0.14 microjoules per hash) | 0.14 | **17x (8x to 30x)** | 5.6x | **4.8x / 2.7x**; at the 5090's whole latency shadow (F about 2.0) 3.7x / 2.0x | about USD 400 to 600 of silicon per die, USD 0.25 to 0.4 per MH/s; an N2 project USD 100 M to 500 M and 18 to 24 months (claimed, the 2.5 economics; the mission lane's break-even model then reads a cap of about USD 330 M to 1.7 B in years 1 to 2 at s = 0.30) | layers 1, 3 and 4: nothing (the dataset is the SRAM's content, every draw firmware; the core an N2 core, k 0.3 to 0.5); layer 2 moves its CAPEX, not its joules: two dies at 4 GiB 15x and about USD 1,000, four at 8 GiB 13x and USD 2,000 to 2,500 | modelled; the N2 density and price claimed | +| **A custom HBM4E base die** (the controller and PHY in the stack on N3P at 0.75 V, "2x the power efficiency", claimed, TrendForce, 1 December 2025; Micron 2027, SK hynix 2026; a non-hyperscaler customer from about 2028) | 2028 or later | 0.9 to 1.0 nJ | 166 MH/s per stack at the model's activate ceiling (36 at the JEDEC tFAW: the ceiling is unmeasured, 5.3) | 0.18 (0.37) | **14x (6.5x)** | 4.4x (2.1x) | 4.4x / 2.6x | about USD 5 per MH/s (approximate); the stack at a premium: Samsung asking about USD 4 to 5 per Gbit for HBM4 against 1.5 for HBM3E (claimed, 2 October 2026) | untouched by all four layers; the brake until about 2028 is HBM allocation and price, not the hash | modelled; the die and price claimed | +| Per-bank processing in memory (Samsung HBM-PIM, SK hynix GDDR6-AiM and AiMX, Samsung LPDDR5X-PIM at Hot Chips 25 August 2026, LPDDR6-PIM in JEDEC work) | now to 2027 | the host memory's own | | | **under 1x: structurally blind to the hash's dependent random reads** | | | | a per-bank unit sees its own bank only: 1.6 percent of a hash's reads land in-bank at a 2 GiB dataset on a 32 MB bank, 0.4 percent at 8 GiB; every other read crosses the channel as it does for a controller chip; layer 2's growth lowers the share further | modelled (arithmetic on the bank size) | +| The honest denominator | | | | the M5 Max at 0.78 microjoules per hash (the GPU and DRAM channels, measured 6 October), 3.1x the 5090 per joule; the 5090 at its knee 1.67 (measured 7 October) | | | | | against the best honest joule every chip edge is 2x to 3x smaller than against the 5090's stock point: the SRAM die 5.6x and the base die 3.6x to 4.4x at zero shadow, about 3x with the shadow at k = 0.5 | measured cards | + +Reading, as the research lane reads it for the class v6 synthesis: the number this file's 5.6 verdict carried ("over 2x", the GDDR7 board at 5.1x) is no longer the strongest chip in the five-year window. The SRAM full store on merchant N2 reaches 17x at zero shadow and 2.7x to 4.8x with the class v4 shadow, and none of the rotating layers reaches it; what holds it is money and time (an N2 project at USD 100 M to 500 M, 18 to 24 months, the break-even cap in the hundreds of millions to about USD 1.7 B) and the one lever the chain has on it is capex through the dataset's size (layer 2's floor and its schedule, priced in the class v6 synthesis). The served chip line is reviewed against these rows at 20:00 UK today in the synthesis's last section; nothing served moves on this section. ## 6. The per-day derivation (item 2) diff --git a/docs/analysis/counter-asic-4-research.md b/docs/analysis/counter-asic-4-research.md index 6c26aded7..e88bbc7ca 100644 --- a/docs/analysis/counter-asic-4-research.md +++ b/docs/analysis/counter-asic-4-research.md @@ -350,7 +350,7 @@ So the inverse lever's best design is the shadow per load: it costs the honest c | 1 | The operating point as the shipped default | class v3 3.6x; class v4 2.1x at k = 1 | lowers the base (330 to 223 W); the v4 premium 145 to 82 W | none | 2 to 4 | AMD and Apple have no lock | unchanged | | 2 | The SM-sparse miner kernel | **DEAD on measured rows (20.3b, 05:12 UTC): a quarter of the SMs holds 98 percent of the rate at the same draw; the draw follows the work, not the SM count** | none | none | 1 + 1 + 4 | measured: the energy per hash never falls below base | closed | | 3 | **The tensor shadow with a SIMD byte-dot verifier** (int8 tiles carrying the premium; a class change) | **DEAD on measured rows (15.1a, 20.4): the 5090's int8 MAC is 1.5 to 4 pJ measured, a 5 nm array 0.04 to 0.4 claimed, k 0.03 to 0.3, under the ALU shadow's band; at the measured premium the chip keeps 3.5x to 6.7x** | the same as class v4 by construction / the 4070 tensor-bound above about R 600 (approximate) | scalar FAILS; AVX2 and VNNI 2.9 to 5.1 ms on the M5 Max core, 7 to 13 ms on a 2019 core (approximate, unwritten) | 8 to 12 (the SIMD verifier, the AMD layout gate, Apple's emulation row) plus the six gates | the 2019-class core; the AMD layout; Apple loses 5 to 10 percent of rate on emulation; k near 1 is claimed, not measured | NEW this pass: the k-above-1 candidate | -| 4 | The shadow per load (capex) | unchanged in energy | none by construction (block-size effect unmeasured at 16) | unchanged | 4 to 6 plus the gates | compile-ahead at 16 sites; a class change; **the 16 x 27 construction is DEAD (20.2a-close, 22:53 UTC): the acceptance rule in execution order accepts 1.4 percent of its candidates; the sound form (one pass of 432 per load) is undrawn** | NEW: doubles the break-even cap to about USD 200 M, on a construction not yet shown to exist | +| 4 | The shadow per load (capex) | unchanged in energy | none by construction (the sound form's 5090 rows in the hash lane's v6 job) | the sound form 8.28 to 8.82 ms loaded on the reference core (measured, the invention lane) | 4 to 6 plus the gates | **REOPENED (8 October, 11:0x UTC): the 16 x 27 iterated form is dead on the no-era census (0.989 rejection); a drawn-era death before the 10:5x fix was an instrument artefact (the window's fixed bits); the sound form `mx8+shl4096x1` accepts 234 of 256 seeds no-era and its drawn-era verdict is the invention lane's 15:00 BST read** | doubles the break-even cap to about USD 200 M on the sound form (class v6 layer 5) | | 5 | The ALU shadow re-weighted toward shuffles and multiplies | 2.1x at k = 1; 2.9x at the pessimistic k 0.46 (modelled) | **raised: the 5090 pays 55.8 pJ per shuffle against 11.3 per add (15.1a), so a shuffle-heavy mix at the same instruction count costs the card up to 5x the premium per instruction** | unchanged | 4 to 6 plus the gates | Apple pays shfl 1.91x; the GPU side measured against it | HELD; the measured GPU side is against it | | 6 | The L2-resident hot table with cache hints | **DEAD on the GPU side (15.1a): an L2 hit costs the 5090 2.4 nJ unlocked and 1.4 at the lock against a chip's SRAM read at 0.2 to 0.5 nJ, k 0.1 to 0.3**; capex +USD 15 | 13 W if the hint holds the rate | under 0.5 ms (unmeasured) | 4 for the ldcs measurement (queued) | Metal has no hint; the rate without it | unchanged; a measurement | | 7 to 11 | class v5 without the shadow; the refresh as a cost; per-card classes; proof of useful work; memory shaping | as section 9 | | | 0 | dead by arithmetic | unchanged | @@ -412,10 +412,14 @@ The tile cost law from these rows: AVX2 (5.82 - 5.33) ms over (1,430 - 128) x 8 Two readings from the Metal row that move section 15.2's tile verdict. First, the reference tile (`mm8_ref`: 12 index shuffles and 32 byte products per lane) is bit-exact against the Rust verifier on all three vector warps of both tile packs, so the tile semantics are pinned on two implementations before any card runs the PTX form. Second, the Apple cost is far above the "10 ALU steps per tile" estimate: 1,024 tiles per hash cost the M5 Max 35 percent of its rate and 4,096 tiles 78 percent, so a tile shadow at the ALU shadow's premium (11,440 tiles per hash) would take the Apple tier out entirely unless Metal gains an integer matrix path reachable from the toolchain (`mpp::tensor_ops::matmul2d` with `uchar` operands, unverified). The tile shadow therefore stands as a class only with an Apple exemption nobody has designed, which moves it from rank 3 to beside rank 5 until that path is measured; the k question it answers is unchanged. -### 20.2a-close The per-load construction, closed (22:44 to 22:53 UTC): dead as a chain class +### 20.2a-close The per-load construction: REOPENED (8 October 2026, 11:0x UTC, the coordinator's wording): the 22:5x UTC death of the 16 x 27 form stands on the no-era census (both instruments: 0.986 to 0.990 rejection per candidate); a drawn-era read of any per-load form before the 10:5x UTC fix measured the instrument (the window's fixed top bits counted as biased); the sound form (one pass of a 256-instruction sub-block per load, `mx8+shl4096x1`) accepts 234 of 256 seeds on the no-era census and its verdict on drawn eras is the invention lane's 15:00 BST read The fix of 20.2a held for distinctness and then met the value-level requirement of 20.2b, and the construction did not survive it. With the acceptance rule stepping the per-load sub-blocks in the order the class executes and judging both the duplicate lanes at a load row and the one-count of every index bit per site over the 64 units (6-sigma band), the 16 x 27 per-load class accepts 22 of 1,621 candidates over 64 seeds (1.4 percent); 42 of 64 seeds exhaust the chain's 32 attempts, which on the chain is an epoch without a program. The first failing test per candidate: a biased index bit 775, duplicate lanes 643, the base rule's lane-constant site 110, (b) 43, (a) 28. The genesis seed accepts none of 32; candidate 0 of the class carries index bit 0 set in 40 of 1,024 addresses at site 0 (z 29.5; the sub-block writer a `rotl` of a `mad` result). The structural reason, read from the record: 27 passes of a 16-instruction map immediately before a load is a tight iteration of a small function, and whatever lossy or product arithmetic it carries (or inherits from the base writer before it) collapses or biases the load's address register before any base instruction can re-randomise it; the class v4 shape places the same 6,912 instructions after instruction 63, where the next iteration's 64 base instructions and 16 loads stand between the block and every load. Both exports (854050a4293f0615, bd64b207a30413fb) were accepted only because the rule did not model the placement; their PC 1 rows stay as the energy reading of the placement, labelled "unsound construction, energy reading only". Design 4 (the shadow per load) and the USD 200 M capex row of 16.2 therefore rest on a construction that does not exist yet; the sound form to try is one pass of a 432-instruction sub-block per load (or 64 x 7), the same N, where the sub-block is a program segment rather than an iterated map; it is a new class to draw, accept and measure (kernel text 6,912 lines per iteration against the 1,024-line block's measured 17 percent on the M5 Max, so the footprint is its own gate), not tonight's. The class v4 shape across the same 16 drawn eras reads 0 duplicate pairs and carries the adv-cache-2 product bit at address bit R exactly in 14 of 17 eras on this pre-amendment generator (one-count 250 or 780 of 1,024, z 15 to 19; sites with `mul`, `rotl` and `or` writers), which is that lane's finding replicated by a second instrument. +### 20.2a-correction (8 October 2026, 10:5x UTC, the invention lane's reading of the per-load tests) + +The 20.2a-close figures (22 of 1,621 candidates accepted, 42 of 64 seeds exhausting 32 attempts) are the NO-ERA figures: the census ran `mx8+shl256x27` with no era draw, so the index was the plain `x AND mask` and every bit was fair to judge. The instrument as first written was wrong under an era: `BiasedIndexBit` judged every index bit, and under an era a narrow-window site's top bits are fixed by design (`verify::window`), so under any drawn era every per-load form read 0 of 256 for the instrument's reason, not the construction's; the invention lane's census (build-1, this crate at 5984ffab) found it. Fixed in this branch at 10:5x UTC: the test judges only the bits inside each site's window mask. The invention lane's own no-era census confirms the construction's figure for the iterated forms (0.989 to 0.990 rejection per candidate for the 16-instruction iterated forms, the 16 x 27 form dead) and finds the sound form: `mx8+shl4096x1` (one pass of a 256-instruction sub-block after every load) accepts 234 of 256 seeds within 32 attempts at 0.927 rejection per candidate (P(exhaust at 256) about 4e-9), the constant-work ladder flat at 0.927 to 0.949 from a 36-instruction sub-block up; the remaining 0.93 is the product's low-bit law at the load's source (bit 0 in 1,563 of 2,329 bias rejections), which class v4's (a') dataflow rule removes at the draw and the per-load class never applied: the named fix. Its verifier on core 40 with core 88 loaded 8.28 to 8.82 ms against 8.33 to 8.63 for class v4's shape in the same minutes. The class v6 document carries the sound form as layer 5 (`docs/design/class-v6-rotating-family.md` section 7c) with its 5090 rows owed from the hash lane's v6 job. + ### 20.2b A named requirement for every CA4 prototype: no biased product bits in an address (the crypto lane's adv-cache-2 reading, 7 October 2026, 22:1x UTC) The crypto lane's finding: a product's low bits are biased (P(bit 0) = 1/4, measured exactly), the bias survives the odd stride multiplier, and the stride rotation places the biased bits at address bits R and up, inside the 28-bit item index unless R is 28 or more. The devnet era draws R = 29, which cuts them off, so 31 of 32 devnet-era programs read clean while 6 of 16 drawn-era programs (R from 3 to 22) show a site over 1.04x (2 over 1.2x, the worst 1.51x); under the 2 GiB genesis dataset (D = 29) R = 29 would show it too. The devnet's cleanliness is an accident of its era draw; the chain prevalence is the drawn-era figure; the price to a partial-store chip stays under 0.1 percent of a hash's reads per site, so no chip number moves. The requirement for any class this file proposes (the per-load shadow, the tile block, a re-weighted shadow): (1) a load whose source register's last writer is a product (`mul`, `mulhi`, `mad`) carries biased low bits into the address, and the acceptance must judge it at the VALUE level (the bit bias of the index at the site over the units), not by the index-distinctness ratio alone, which the duplicate-lane test above is; (2) the census of any candidate reads across the drawn eras split by R (3 to 22 against 28 to 31), as 20.2a now does for the per-load class, never the devnet era alone. The per-load class's dynamic rule covers distinctness, not bias; the value-level test is owed and is the same item for class v5's acceptance. Nothing in class v4 or v5 moves on this without main's word. diff --git a/tools/ci/export-exclude.txt b/tools/ci/export-exclude.txt index eeeb9b8f6..c2b651bfc 100644 --- a/tools/ci/export-exclude.txt +++ b/tools/ci/export-exclude.txt @@ -15,6 +15,10 @@ docs/analysis/ci-failures-2026-10-06.md # 7 October 2026: the last research round (the mission lanes and the closed list): internal research written for the # owner, quoting his words and the operations record; the public spec mirror carries none of it docs/analysis/mission +# 7 October 2026 (night): the Counter ASIC 4.0 research record (an internal research document: the founder's words, box names, lane records) +docs/analysis/counter-asic-4-research.md +# 8 October 2026: the class v6 design record (an internal research document: the founder's words, lane records) +docs/design/class-v6-rotating-family.md # 7 October 2026: the in-house adversarial cryptanalysis pass (crypto-engage and the adv-* lanes): rule set, plans, reports and # copied logs are research and operations documents, not public export; the public text is the served sentence main landed. docs/analysis/cryptanalysis