diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index ce0615898..1eaeb0d07 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -24,6 +24,8 @@ Site card placement: the Mine section of `site/index.html` beside "no chip can b ## Level 3: the numbers page (`site/bench.html`, section "Counter ASIC") +Headline of the chip model (5 October 2026, night): the strongest chip holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) and computes dataset items on the fly; its gain over the RTX 5090 is 2.4x as the parameters stand, and no write-scratch share within an 8 GB card's budget changes that. The lever that does is the dataset item's mixer cost (x4: 1.8x with a 3x fixed-function factor, verifier 1.6 to 4.8 ms per warp). Until that or the cache rule is in the class, the claim reads: a chip gains about 2.4x by the published model, the target is under 2x, the model and the bounty are public. + Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per hash, the latency-bound share (rate over the card's random-read ceiling per load), the CPU verifier per warp, with machine, date and command. The chip model before and after Counter ASIC 2.0 (the m16 model's gain arithmetic at the v2 class and at the v3 class, with the SRAM a mirror needs, cited or approximate as the analysis says). The bounty terms (spec O-1.17: the leaderboard by card model, the standing bounty for any chip design beating a GPU by more than 2x, January 2027). Here the layers are named next to their numbers: read width, per-program mix, scratch, era layout, working set, hot table, cache schedule, the reserved integer-matrix family. | Card | v2 MH/s | v3 MH/s | v3 bytes per hash | Latency-bound share v3 | Verifier ms per warp v3 | diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 198d38e93..a49e22ee6 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -49,12 +49,12 @@ the project lead's rules, applied by the coordinator and recorded here with the |---|---|---|---| | Width (layers 1 and 2) | the widest read that keeps every card we own latency-bound (achieved loads within 90% of the probe ceiling) with margin on the 5090 (its bytes per hash under a third of its bandwidth at the measured rate) | `` | | | Per-load mix (layer 2) | in, if the min-to-max spread across six programs is under 5% per card | `` | | -| Scratch share (layer 3) | the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap | `` | | +| Scratch share (layer 3) | the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap | DECIDED (5 October 2026, delegated): 0. No share under the cap moves the on-die-cache recompute chip, so layer 3 is not adopted into v3 | `docs/analysis/scratch-soundness.md` (ca2-soundness a465881): chip 333 MH/s against the 5090's measured 139.7 = 2.4x at 0% RMW; 2.4x at 12.5 / 25 / 50% replaced (chip 381 / 443 / 661 against 160 / 186 / 279 projected) and 2.4x or more added, at 32 and 128 KB; the chip keeps the scratch implicitly in 80 to 320 B per lane because the verifier resets it per unit | | Activation height N4 | devnet tip + 14,400 at publish, checked >= 10,800 at publish, rounded up to the epoch boundary | `` | | ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. If no scratch share under the 6 GB cap gets that chip under 2x, the status file and the public copy's level 3 say so plainly, and the site's "under 2x" claim is qualified until the mixer multiplier or the cache rule closes it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. Whether x4 goes into v3 tonight or into Counter ASIC 3.0 is with the coordinator; the plan proceeds on 3.0 unless told otherwise. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ## 7. Gates before any publish (all of them, no exceptions) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 00716e78c..27ab3f693 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -100,3 +100,11 @@ Two columns, headline = shipped-product density (AMD V-Cache 41 mm^2 per 64 MiB | N2, $30,000 | 106 / 54, $56 / $26 | 146 / 74, $81 / $37 | 425 / 215, $343 / $131 | Year 10 at the 6% per year trend: 59 mm^2 for the flat cache (82 with the hot table), 8% of a 750 mm^2 die. One reticle holds 1.3 GiB (N7) to 1.9 GiB (N2); mirror share of a 750 mm^2 die at year 0: 14% (22% with the hot table), inside M16's 13 to 40% band. Recommendation unchanged: C. Latency section added (MEMSYS 2018, Chang 2017, the mining-chip memory-type note), marked as research the agent did not re-read tonight apart from the V-Cache figure. + +## 21:50 layer 3 soundness landed: the scratch does not move the chip; scratch share decided 0 + +ca2-soundness (0d8f745 tests and trace hook, a465881 doc and bench-log). The on-die-cache recompute chip (N5 headline 128 mm^2, $46) at 50 T op/s: 333 MH/s against the 5090's measured 139.7, 2.4x; at 12.5 / 25 / 50% RMW replaced, chip 381 / 443 / 661 against 5090 projected 160 / 186 / 279, 2.4x each; added, 2.4x or more; 32 or 128 KB alike. The chip keeps the scratch implicitly in 80 to 320 B per lane (the verifier resets it per unit), needs about 530 units in flight, dense scratch 6.2 / 3.1 mm^2 at N5. Under the project lead's rule the scratch share is 0: layer 3 is NOT adopted into v3; the public "under 2x" claim is qualified (public copy level 3 rewritten). The measured lever is M16's mixer multiplier (x2 1.2x at 0.8 to 2.4 ms verify; x4 0.6x at 1.6 to 4.8 ms; 3.6x and 1.8x with a 3x fixed-function factor); whether x4 enters v3 tonight is asked of the coordinator; default: Counter ASIC 3.0. + +Soundness results (Metal, M5 Max): 28/28 edge launches, 200/200 fuzz packs (91 s), 56/56 hand-model edge checks, 42/42 kernels pass the static scratch-mask check with 6 deliberate breaks caught, broken tag and broken lazy fill caught, fingerprint 8c07620f4d9adefd warp-count-independent; re-hit rates 2 to 33% above the birthday bound (slot addresses are register low bits); written words unbiased (worst 3.63 of 6 sigma). Verifier exactness needs a host contract (zero the arena at allocation and at the 32-bit tag wrap, tags from 1), which neither host gives today. Pre-existing on readwidth b970dda: verify::tests::fold_and_wide_fetch overflows under the test profile (wrapping_mul fixes it); passed to the readwidth agent with the class sweep. + +Gate G3 note: the scratch tests (igneum-pow/tests/scratch.rs) join the v3 suite even though the class carries no scratch, parametric over the class; they guard the v2 path's scratch-free invariant at zero cost.