Counter ASIC 2.0: layers 1, 2, 3 decided from the readwidth table, layer 5 Mac rows and the added-form redesign, user tiers, public level 2 edit

This commit is contained in:
igneum-labs 2026-10-05 20:27:01 +00:00
parent 6a9b7c1283
commit fff1994c8a
3 changed files with 33 additions and 8 deletions

View file

@ -10,7 +10,7 @@ Built for graphics cards. A custom chip gains under 2x, and the model and the bo
Three ideas, no layer names, no widths, no SRAM.
**The hash rewrites itself.** A new program every hour, drawn from the chain. Its memory pattern and its read widths change with it. The rules change on a schedule fixed at launch. No release, no vote.
**The hash rewrites itself.** A new program every hour, drawn from the chain. Its memory pattern changes with it. The rules change on a schedule fixed at launch. No release, no vote.
**It waits on memory, not maths.** Every hash is a chain of random reads into a table too big for a chip to carry. The wait is the same physics for everyone.
@ -30,9 +30,11 @@ Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per h
| Card | v2 MH/s | v3 MH/s | v3 bytes per hash | Latency-bound share v3 | Verifier ms per warp v3 |
|---|---|---|---|---|---|
| Apple M5 Max (Metal) | 27.7 | [owed: readwidth table] | | | |
| RTX 5090 (CUDA) | [owed] | [owed] | | | |
| RX 9070 XT (OpenCL, eGPU) | 18.0 | [owed] | | | |
| Apple M5 Max (Metal) | 27.74 | [owed: the v3 pack] | 512 + hot | [owed] | [owed: ca2-mixer] |
| RTX 5090 (CUDA) | 136.1 | [owed] | 512 + hot | 0.96 at v2 | [owed] |
| RX 9070 XT (OpenCL, eGPU) | 18.15 | [owed] | 512 + hot | 0.87 at v2 | [owed] |
AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), near parity per pound and about 4.5x worse per watt; the card's memory system, not a tuning gap.
Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant].

View file

@ -47,8 +47,8 @@ the project lead's rules, applied by the coordinator and recorded here with the
| Decision | the project lead's rule | Choice | The number |
|---|---|---|---|
| Width (layers 1 and 2) | the widest read that keeps every card we own latency-bound (achieved loads within 90% of the probe ceiling) with margin on the 5090 (its bytes per hash under a third of its bandwidth at the measured rate) | `<pending the readwidth table>` | |
| Per-load mix (layer 2) | in, if the min-to-max spread across six programs is under 5% per card | `<pending>` | |
| Width (layer 1) | the widest read that keeps every card we own latency-bound (achieved loads within 90% of the probe ceiling) with margin on the 5090 (its bytes per hash under a third of its bandwidth at the measured rate) | DECIDED (5 October 2026, delegated): keep v2, 128 x 4 B. w16 passes the rule (shares 0.90 / 0.84 / 1.03, 18% of the 5090's stream) but closes nothing and does not move the chip row; w64 and w64x4 make the 5090 bandwidth-bound (share 0.58 / 0.56, 37% of stream) | `docs/plans/read-width.md` (readwidth e752fc7): v2 5090 136.1 MH/s, 9070 XT 18.15, M5 Max 27.74 (gap 7.5x); w16 139.8 / 17.90 / 28.26 (gap 7.8x); w64 71.9 / 17.59 / 28.27 (gap 4.1x); the 9070 XT does 2.4 G dependent reads/s at every width |
| Per-load mix (layer 2) | in, if the min-to-max spread across six programs is under 5% per card | DECIDED (5 October 2026, delegated): out | spreads of the median over six programs: mix 50/35/15 5090 18.8%, 9070 XT 7.4%, M5 Max 11.3%; mix 25/50/25 22.3% / 5.5% / 8.1% |
| Scratch share (layer 3) | the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap | DECIDED (5 October 2026, delegated): 0. No share under the cap moves the on-die-cache recompute chip, so layer 3 is not adopted into v3 | `docs/analysis/scratch-soundness.md` (ca2-soundness a465881): chip 333 MH/s against the 5090's measured 139.7 = 2.4x at 0% RMW; 2.4x at 12.5 / 25 / 50% replaced (chip 381 / 443 / 661 against 160 / 186 / 279 projected) and 2.4x or more added, at 32 and 128 KB; the chip keeps the scratch implicitly in 80 to 320 B per lane because the verifier resets it per unit |
| Activation height N4 | devnet tip + 14,400 at publish, checked >= 10,800 at publish, rounded up to the epoch boundary | `<at publish>` | |
@ -56,6 +56,10 @@ the project lead's rules, applied by the coordinator and recorded here with the
The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2).
### 6b. The user tiers (the consequences rule)
AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second at 1 GiB, measured tonight), near parity per pound and about 4.5x worse per watt (0.089 against 0.398 MH/W, measured 5 October 2026). This is the card's memory system, not a tuning gap: no read width closes it without making the 5090 bandwidth-bound. The level 3 numbers page states it.
## 7. Gates before any publish (all of them, no exceptions)
| # | Gate | Evidence required | State |
@ -89,8 +93,8 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa
|---|---|---|---|
| The width rule (layers 1 and 2) | 4 B fixed; 16 B; 64 B; the per-program mix | `<the readwidth table's decision rule: the widest read that keeps every card latency-bound with margin on the 5090>` | `docs/plans/read-width.md` |
| The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `<from the readwidth table and docs/analysis/scratch-soundness.md>` | the same |
| The hot table size (layer 5) | 32, 64, 96 MB | `<from docs/plans/hot-table.md>` | |
| The era draw bounds (layers 4 and 8) | the draw as written in `docs/plans/era-layout.md` | | |
| The hot table (layer 5) | 32, 64, 96 MB; replaced or added | In the ADDED form only (16 dataset loads plus k hot loads): the replaced form lets the on-die-cache chip skip item derivations (k = 4: chip x1.33 against the GPU's measured x1.05 to x1.22). Size = the largest table resident on every card we own, pending the PC rows; Mac rows (replaced form, M5 Max): hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k8 x1.71; Apple OpenCL dependent-read probe 32 / 64 / 96 / 1024 MiB: 21.7 / 12.8 / 12.3 / 3.50 G loads/s | `docs/plans/hot-table.md` (ca2-cache 65bc7a7) |
| The era draws (layers 4 and 8) | in, if the min-to-max spread across six drawn eras is under 5% per card; item size fixed at 4 B, draws of stride, interleave and the working-set window at or above 256 MiB | pending the six-era table | `docs/plans/era-layout.md` |
| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate |
| Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) |
| The activation height N4 | the rule of section 3 | | |

View file

@ -114,3 +114,22 @@ Gate G3 note: the scratch tests (igneum-pow/tests/scratch.rs) join the v3 suite
Coordinator's decision under the project lead's delegation (recorded in the rollout plan section 6a): the mixer multiplier x4 and the cache growth rule (option C) enter class v3 behind the same activation; layer 3 stays out at scratch share 0, its soundness document and pack-contract tests kept. Agent af345b1e2c541ffbb (branch ca2-mixer) implements `mixer_mult` as a class parameter (m mixer applications per round, the 8 dependent reads unchanged), the `cache_log2_words(day)` schedule (doublings at years 4 and 12 with the dataset stepping to the next power of two), re-cuts the v3 dataset vectors, re-runs the soundness suite, measures the verifier (v2 0.604 ms per warp; v3 expected 1.6 to 4.8 ms) and the 1 GiB build on the Mac, prepares the 5090 and 9070 XT build-time job, and writes docs/analysis/chip-model-v3.md with the combined headline row (fixed-function factor included). The claim on the site reads "under 2x" only if that row does; else qualified, with the mixer x8 and the hot table named as the next levers.
Agents now: ca2-era (a452664c512c73b9b), ca2-cache (a5271cf269757b118), ca2-node (a3f505a9d981300cd), ca2-mixer (af345b1e2c541ffbb). Done: ca2-analysis, ca2-soundness. Waiting: the readwidth PC table; PC 2 (proving memsweep runs) and PC 1 (readwidth 9070 round).
## 22:15 the readwidth table landed; layers 1, 2, 3 decided; layer 5 measured on the Mac and redesigned
Readwidth e752fc7 (`docs/plans/read-width.md`), bit-exact on Metal, Apple OpenCL, the 5090 (NVRTC) and the 9070 XT, both PCs released. MH/s (latency-bound share):
| Class | RTX 5090 | RX 9070 XT | M5 Max | Gap |
|---|---|---|---|---|
| v2 (128 x 4 B) | 136.1 (0.96) | 18.15 (0.87) | 27.74 (1.01) | 7.5x |
| w16 | 139.8 (0.90) | 17.90 (0.84) | 28.26 (1.03) | 7.8x |
| w64 | 71.9 (0.58, 37% of stream) | 17.59 (0.78) | 28.27 (1.03) | 4.1x |
| w64x4 (32 loads) | 275.3 (0.56) | 75.19 (0.84) | 109.7 (1.00) | 3.7x |
| mix 50/35/15, six programs, spread of median | 18.8% | 7.4% | 11.3% | |
| mix 25/50/25 | 22.3% | 5.5% | 8.1% | |
| scratch 32 KB at 12.5 / 25 / 50% | -18 / -21 / -12% | -18 / -22 / -21% | -7 / +12 / +74% | |
| scratch 128 KB | -21 / -30 / -48% | -21 / -27 / -33% | -7 / -1 / +25% | |
Decisions (the project lead's rules, delegated): layer 1 keep v2 (w16 passes the rule but closes nothing and does not move the chip row; the vector re-cut is not worth it); layer 2 out (spread over 5% on every card); layer 3 out (scratch share 0). The AMD gap is the card's dependent-read rate (2.4 G/s at every width), stated for the user tiers in the rollout plan 6b.
Layer 5 (ca2-cache 53ef59f, 011cc0a, 86726cd, 65bc7a7; `docs/plans/hot-table.md`): five packs bit-exact on Metal and Apple OpenCL (96/96 each). M5 Max rates against v2 27.68: hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k2 x1.00, hot64k8 x1.71; verifier 0.344 to 0.560 ms against 0.626; hot fill per epoch 24 / 46 / 73 ms on one core, 0.07 / 0.15 / 0.22 ms on the GPU; Apple OpenCL probe 32 / 64 / 96 / 1024 MiB 21.7 / 12.8 / 12.3 / 3.50 G loads/s. Redesign ordered: hot loads ADDED beside the 16 dataset loads (the replaced form lets the on-die-cache chip skip item derivations and worsens the gain); the agent re-measures the added form and rebuilds the PC job. PC 1 is given to the dot4 probe (under 15 min), then to the era agent, then the hot-table job; PC 2 stays the proving agent's.