Counter ASIC 2.0: the mixer x4 decision into v3, ca2-mixer agent, status 22:00

This commit is contained in:
igneum-josh 2026-10-05 20:24:55 +00:00
parent ff65d41dfc
commit 4b550fa6d1
3 changed files with 9 additions and 3 deletions

View file

@ -24,7 +24,7 @@ Site card placement: the Mine section of `site/index.html` beside "no chip can b
## Level 3: the numbers page (`site/bench.html`, section "Counter ASIC")
Headline of the chip model (5 October 2026, night): the strongest chip holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) and computes dataset items on the fly; its gain over the RTX 5090 is 2.4x as the parameters stand, and no write-scratch share within an 8 GB card's budget changes that. The lever that does is the dataset item's mixer cost (x4: 1.8x with a 3x fixed-function factor, verifier 1.6 to 4.8 ms per warp). Until that or the cache rule is in the class, the claim reads: a chip gains about 2.4x by the published model, the target is under 2x, the model and the bounty are public.
Headline of the chip model (5 October 2026, night): the strongest chip holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) and computes dataset items on the fly; its gain over the RTX 5090 is 2.4x as the parameters stand, and no write-scratch share within an 8 GB card's budget changes that. The lever that does is the dataset item's mixer cost (x4: 1.8x with a 3x fixed-function factor, verifier 1.6 to 4.8 ms per warp). Decided 5 October 2026 (delegated): the mixer x4 and the cache growth rule enter class v3, so the headline row is the on-die-cache chip against v3 with everything combined. [owed: the combined row from docs/analysis/chip-model-v3.md; if it reads 1.8x, the claim is "under 2x" with the margin stated as thin, and the next levers are named: the mixer x8 and the hot table.]
Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per hash, the latency-bound share (rate over the card's random-read ceiling per load), the CPU verifier per warp, with machine, date and command. The chip model before and after Counter ASIC 2.0 (the m16 model's gain arithmetic at the v2 class and at the v3 class, with the SRAM a mirror needs, cited or approximate as the analysis says). The bounty terms (spec O-1.17: the leaderboard by card model, the standing bounty for any chip design beating a GPU by more than 2x, January 2027). Here the layers are named next to their numbers: read width, per-program mix, scratch, era layout, working set, hot table, cache schedule, the reserved integer-matrix family.

View file

@ -8,7 +8,7 @@ Shape and rules follow `docs/plans/finality-v3-rollout-devnet.md` and the publis
Only the lottery hash's program class changes, and only from the first epoch at or above the height. One switch, `program_class_v3_activation_daa`, in `Params` and `OverrideParams` like `difficulty_v2_activation_daa`; default `u64::MAX` (never) on every network. Because one epoch has one program (spec 01 section 1.12), the switch keys on the EPOCH: epoch `e` is class v3 when `3,600 e >= N4`, so the activation is rounded up to an epoch boundary and a block's class is a function of its DAA score alone, as today.
Class v3 = generator version 3: the width rule `<W>` (layer 1 or 2, decision 1), the per-warp scratch at `<share>` RMW and `<KB>` per warp (layer 3, under the 6 GB working-set cap), the era draw of the table layout and the working set (layers 4 and 8), the hot table of `<S>` MB from the epoch seed (layer 5). New program id (`generator = 3` in the id's preimage, spec 1.4.6), new packs and vectors, new `IGNEUM_GENERATOR` in every pack, a pack of the other version refused by every implementation (spec 1.4.5 already says so).
Class v3 = generator version 3: the width rule `<W>` (layer 1 or 2, decision 1), no scratch (layer 3 decided out: scratch share 0), the era draw of the table layout and the working set (layers 4 and 8), the hot table of `<S>` MB from the epoch seed (layer 5), the cache growth rule of layer 6 (option C) and the M16 mixer x4 in the dataset item construction (decided 5 October 2026, delegated). New program id (`generator = 3` in the id's preimage, spec 1.4.6), new packs and vectors, new `IGNEUM_GENERATOR` in every pack, a pack of the other version refused by every implementation (spec 1.4.5 already says so).
What does not change: the chain, the genesis, the databases, the day key and the 256 MiB cache fill, the dataset items (spec 1.8.5, if the era interleave keeps the item values; the era-layout document says what it costs otherwise), finality, fees, proving. The SP1 guest does not read the lottery hash (`proving/igneum-prove` has no dependency on `igneum-pow`; the pinned guest of DAA 210,000 is a fee-table switch), so no new guest is pinned. The node's `EpochSeeds` gains the class and the era bytes; `IgneumEngine::epoch_for` builds the v3 `Epoch` from them; the miner's `export-pack` and the serve protocol's job line carry the class so a GPU worker regenerates the right pack from the seed bytes.
@ -54,7 +54,7 @@ Josh's rules, applied by the coordinator and recorded here with the number that
### 6a. The chip model's headline, and how the scratch share is chosen
The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. Whether x4 goes into v3 tonight or into Counter ASIC 3.0 is with the coordinator; the plan proceeds on 3.0 unless told otherwise. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2).
The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; Josh confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2).
## 7. Gates before any publish (all of them, no exceptions)

View file

@ -108,3 +108,9 @@ ca2-soundness (0d8f745 tests and trace hook, a465881 doc and bench-log). The on-
Soundness results (Metal, M5 Max): 28/28 edge launches, 200/200 fuzz packs (91 s), 56/56 hand-model edge checks, 42/42 kernels pass the static scratch-mask check with 6 deliberate breaks caught, broken tag and broken lazy fill caught, fingerprint 8c07620f4d9adefd warp-count-independent; re-hit rates 2 to 33% above the birthday bound (slot addresses are register low bits); written words unbiased (worst 3.63 of 6 sigma). Verifier exactness needs a host contract (zero the arena at allocation and at the 32-bit tag wrap, tags from 1), which neither host gives today. Pre-existing on readwidth b970dda: verify::tests::fold_and_wide_fetch overflows under the test profile (wrapping_mul fixes it); passed to the readwidth agent with the class sweep.
Gate G3 note: the scratch tests (igneum-pow/tests/scratch.rs) join the v3 suite even though the class carries no scratch, parametric over the class; they guard the v2 path's scratch-free invariant at zero cost.
## 22:00 decided: M16 mixer x4 into v3; agent ca2-mixer started
Coordinator's decision under Josh's delegation (recorded in the rollout plan section 6a): the mixer multiplier x4 and the cache growth rule (option C) enter class v3 behind the same activation; layer 3 stays out at scratch share 0, its soundness document and pack-contract tests kept. Agent af345b1e2c541ffbb (branch ca2-mixer) implements `mixer_mult` as a class parameter (m mixer applications per round, the 8 dependent reads unchanged), the `cache_log2_words(day)` schedule (doublings at years 4 and 12 with the dataset stepping to the next power of two), re-cuts the v3 dataset vectors, re-runs the soundness suite, measures the verifier (v2 0.604 ms per warp; v3 expected 1.6 to 4.8 ms) and the 1 GiB build on the Mac, prepares the 5090 and 9070 XT build-time job, and writes docs/analysis/chip-model-v3.md with the combined headline row (fixed-function factor included). The claim on the site reads "under 2x" only if that row does; else qualified, with the mixer x8 and the hot table named as the next levers.
Agents now: ca2-era (a452664c512c73b9b), ca2-cache (a5271cf269757b118), ca2-node (a3f505a9d981300cd), ca2-mixer (af345b1e2c541ffbb). Done: ca2-analysis, ca2-soundness. Waiting: the readwidth PC table; PC 2 (proving memsweep runs) and PC 1 (readwidth 9070 round).