Counter ASIC 2.0: the shipped-density correction to the SRAM mirror (130 to 165 mm^2), carried into the rollout and public docs

This commit is contained in:
igneum-labs 2026-10-05 20:18:39 +00:00
parent b258094890
commit cbbc81f002
3 changed files with 7 additions and 3 deletions

View file

@ -32,7 +32,7 @@ Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per h
| RTX 5090 (CUDA) | [owed] | [owed] | | | |
| RX 9070 XT (OpenCL, eGPU) | 18.0 | [owed] | | | |
Chip model, before and after: [owed: from docs/analysis/sram-mirror.md and the hot-table and scratch analyses].
Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant].
## Level 4: the analysis documents

View file

@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the
### 6a. The chip model's headline, and how the scratch share is chosen
The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (54 to 83 mm^2, $13 to $26 of silicon, approximate, `docs/analysis/sram-mirror.md`) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. If no scratch share under the 6 GB cap gets that chip under 2x, the status file and the public copy's level 3 say so plainly, and the site's "under 2x" claim is qualified until the mixer multiplier or the cache rule closes it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2).
The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. If no scratch share under the 6 GB cap gets that chip under 2x, the status file and the public copy's level 3 say so plainly, and the site's "under 2x" claim is qualified until the mixer multiplier or the cache rule closes it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2).
## 7. Gates before any publish (all of them, no exceptions)
@ -91,6 +91,6 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa
| The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `<from the readwidth table and docs/analysis/scratch-soundness.md>` | the same |
| The hot table size (layer 5) | 32, 64, 96 MB | `<from docs/plans/hot-table.md>` | |
| The era draw bounds (layers 4 and 8) | the draw as written in `docs/plans/era-layout.md` | | |
| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate |
| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate |
| Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) |
| The activation height N4 | the rule of section 3 | | |

View file

@ -77,3 +77,7 @@ Node-fork agent a3f505a9d981300cd started 20:50: ca2-v3 (igneum-pow seam: Progra
## 21:05 decisions recorded: layer 6 option C, layer 7 R1 = mm8; the chip headline
Layer 6 DECIDED (delegated; the project lead confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles (256 MiB genesis, 512 MiB year 4, 1 GiB year 12); one-core fill 0.2 / 0.4 / 0.8 s, under 1 s at every step. Layer 7 DECIDED: reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal. The chip model's headline now names the on-die-cache recompute chip (54 to 83 mm^2) as a row per variant; the scratch share is chosen as the smallest share at which that chip's gain falls under 1.5x, else said plainly and the public "under 2x" claim qualified. The soundness agent carries that table (its question 2) with M16's mixer multiplier beside it.
## 21:15 correction to the SRAM mirror figures (chip-economics research, cluster D)
Bit cell x 0.70 understates real die area. Shipped cache dies: AMD 3D V-Cache 64 MB on 41 mm^2 at 7 nm (1.56 MB/mm^2, Tom's Hardware, Hot Chips August 2021); Graphcore GC200 900 MB on 823 mm^2 with compute (1.09 MB/mm^2); Groq TSP 220 MB on 725 mm^2 at 14 nm (0.30 MB/mm^2); TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30% assist overhead (SemiAnalysis, December 2022). A 256 MiB mirror is about 165 mm^2 at 7 nm on the densest shipped cache-only die and about 130 mm^2 at N5/N3E, not 54 to 83 mm^2; cost per die 2 to 3x the earlier figure; the conclusion (affordable for a funded chip) stands. The analysis agent is redoing the table with both columns; the soundness agent carries the corrected density into the chip row. Latency citations behind the latency-bound rule, to be added: DRAM row cycle 40 to 48 ns across DDR4, GDDR5, HBM2 (Li, Reddy, Jacob, MEMSYS 2018); latency 1.3x in two decades against bandwidth 20x (Chang 2017); no shipped mining chip used HBM or stacked memory.