Counter ASIC 2.0: layer 6 option C and layer 7 R1 decided (delegated), the chip headline rule for the scratch share
This commit is contained in:
parent
38c113e957
commit
bdfb7a3467
2 changed files with 10 additions and 2 deletions
|
|
@ -52,6 +52,10 @@ Josh's rules, applied by the coordinator and recorded here with the number that
|
|||
| Scratch share (layer 3) | the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap | `<pending the readwidth table and docs/analysis/scratch-soundness.md>` | |
|
||||
| Activation height N4 | devnet tip + 14,400 at publish, checked >= 10,800 at publish, rounded up to the epoch boundary | `<at publish>` | |
|
||||
|
||||
### 6a. The chip model's headline, and how the scratch share is chosen
|
||||
|
||||
The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (54 to 83 mm^2, $13 to $26 of silicon, approximate, `docs/analysis/sram-mirror.md`) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. If no scratch share under the 6 GB cap gets that chip under 2x, the status file and the public copy's level 3 say so plainly, and the site's "under 2x" claim is qualified until the mixer multiplier or the cache rule closes it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2).
|
||||
|
||||
## 7. Gates before any publish (all of them, no exceptions)
|
||||
|
||||
| # | Gate | Evidence required | State |
|
||||
|
|
@ -87,6 +91,6 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa
|
|||
| The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `<from the readwidth table and docs/analysis/scratch-soundness.md>` | the same |
|
||||
| The hot table size (layer 5) | 32, 64, 96 MB | `<from docs/plans/hot-table.md>` | |
|
||||
| The era draw bounds (layers 4 and 8) | the draw as written in `docs/plans/era-layout.md` | | |
|
||||
| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | Recommended C: the cache doubles when the dataset doubles (512 MiB at year 4); not delegated, a gate 1 decision for Josh | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate |
|
||||
| Layer 7 | reserved family, unlock by era height or 90% signal | Reserve R1 = mm8 (uint8 8x16 by 16x8 tile per unit), W_new 4, unlock era 4 or 90% signal; switched off, no consensus effect tonight | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) |
|
||||
| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; Josh confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate |
|
||||
| Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) |
|
||||
| The activation height N4 | the rule of section 3 | | |
|
||||
|
|
|
|||
|
|
@ -73,3 +73,7 @@ Branch ca2-analysis (5d5ba15, f59708d). Layer 6: no cache growth rule exists in
|
|||
Node-fork agent a3f505a9d981300cd started 20:50: ca2-v3 (igneum-pow seam: ProgramClass, Epoch::from_chain_seeds, generator 3 in the program id, class in the pack) and ca2-v3-node (vendor/igneum-node-ca2 under the ca2-v3 worktree, from 21d4c73c, with pack-loop 05ef0fa3 merged): the field in Params, OverrideParams and the digest, the epoch-boundary rounding, the era stand-in, the job line, the fast-time gate script.
|
||||
|
||||
0.3.11 scope (coordinator, 20:52): carries program_class_v3 and proving v1 together (one override object, one digest, one publish; each half under the same gates); the ready half ships as 0.3.11 and the other as 0.3.12 if one lags. Next-cut list recorded in the rollout plan section 8a.
|
||||
|
||||
## 21:05 decisions recorded: layer 6 option C, layer 7 R1 = mm8; the chip headline
|
||||
|
||||
Layer 6 DECIDED (delegated; Josh confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles (256 MiB genesis, 512 MiB year 4, 1 GiB year 12); one-core fill 0.2 / 0.4 / 0.8 s, under 1 s at every step. Layer 7 DECIDED: reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal. The chip model's headline now names the on-die-cache recompute chip (54 to 83 mm^2) as a row per variant; the scratch share is chosen as the smallest share at which that chip's gain falls under 1.5x, else said plainly and the public "under 2x" claim qualified. The soundness agent carries that table (its question 2) with M16's mixer multiplier beside it.
|
||||
|
|
|
|||
Loading…
Reference in a new issue