Counter ASIC 2.0: layer 5 decided OUT on the PC rows; status 21:36

This commit is contained in:
igneum-labs 2026-10-05 21:36:30 +00:00
parent 030858063c
commit 9eb6b5e654
3 changed files with 19 additions and 2 deletions

View file

@ -115,7 +115,7 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa
|---|---|---|---|
| The width rule (layers 1 and 2) | 4 B fixed; 16 B; 64 B; the per-program mix | `<the readwidth table's decision rule: the widest read that keeps every card latency-bound with margin on the 5090>` | `docs/plans/read-width.md` |
| The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `<from the readwidth table and docs/analysis/scratch-soundness.md>` | the same |
| The hot table (layer 5) | 32, 64, 96 MB; replaced or added | In the ADDED form only (16 dataset loads plus k hot loads): the replaced form lets the on-die-cache chip skip item derivations (k = 4: chip x1.33 against the GPU's measured x1.05 to x1.22). Size = the largest table resident on every card we own, pending the PC rows; Mac rows (replaced form, M5 Max): hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k8 x1.71; Apple OpenCL dependent-read probe 32 / 64 / 96 / 1024 MiB: 21.7 / 12.8 / 12.3 / 3.50 G loads/s | `docs/plans/hot-table.md` (ca2-cache 65bc7a7) |
| The hot table (layer 5) | 32, 64, 96 MB; replaced or added | DECIDED (5 October 2026, 21:36 UTC, delegated): OUT of v3. The rule was g at or above 0.97 on both PC cards in the added form; measured g is 0.87 / 0.85 / 0.84 on the 5090 and 0.84 / 0.81 / 0.80 on the 9070 XT at 32 / 64 / 96 MiB: neither card keeps even the 32 MiB table resident while the dataset streams, and the replaced form helps the on-die-cache chip. Layer 5 stays a measured option for 3.0 | `docs/plans/hot-table.md` (ca2-cache): 5090 MH/s v2 136.1; replaced hot32k4 146.6, hot64k4 140.8, hot96k4 138.5, hot64k2 137.5, hot64k8 163.6; added hot32k4a 118.7, hot64k4a 115.4, hot96k4a 114.4. 9070 XT v2 18.15; replaced 19.79 / 18.73 / 18.33 / 18.17 / 22.32; added 15.27 / 14.62 / 14.56. M5 Max added 0.93 / 0.87 / 0.83. All eight packs bit-exact on the 5090 and the 9070 XT with the Mac's fingerprints (21:29 to 21:35 UTC, no restart straddled) |
| The era draws (layers 4 and 8) | in, if the min-to-max spread across six drawn eras is under 5% per card; item size fixed at 4 B, draws of stride, interleave and the working-set window at or above 256 MiB | pending the six-era table | `docs/plans/era-layout.md` |
| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate |
| Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) |

View file

@ -367,3 +367,20 @@ Node agent: the Metal worker's ServeStore already keys datasets by day and progr
Gate G4: the first fast-time run failed at node start on the script, not the node: JSON.parse turned a never height (18446744073709551615) into 1.8446744073709552e+19 and the node refused the override; fixed as text merging in class-v3.mjs and simnet.mjs (d5ff532; no other script in tools/, infra/ or sim/ has the shape). The gate runs again on the mixer-x4 class (binaries from 79bd8e10 + ca2-v3 66eeba3). Main-repo tip for the suite job title: d5ff532 (era and cache not yet merged).
21:34. PC 2: agg-cost-pc2-1 closed 21:25:11Z (done, miners and prover back on); agg-cost-pc2-2 went out at 21:33Z (a missed close), self-limited to 16.5 min, closes about 21:51Z; the 0.3.11 suites publish after it AND after PC 2's app shows the 0.3.10 STATUS line (the update-now goes to PC 2 now). PC 1: the hot-table job runs (release about 21:40); PC 1's update-now follows the release; the era job starts after PC 1's 0.3.10 STATUS line.
## 21:36 layer 5 measured on both PC cards: OUT of v3
Hot-table jobs on PC 1 (fetch 21:29:01Z; run-ca2-hot-5090 126 s, closed 21:31:38Z; run-ca2-hot-9070 247 s, closed 21:35:39Z; both cards restored; the 9070 XT WAS on the bus, full path; app 0.3.9 throughout, no straddle). All eight packs bit-exact on the 5090 and the 9070 XT with the Mac's fingerprints.
| Pack | RTX 5090 MH/s (v2 136.1) | RX 9070 XT MH/s (v2 18.15) | M5 Max g |
|---|---|---|---|
| hot32k4 (replaced) | 146.6 (x1.08) | 19.79 (x1.09) | x1.22 |
| hot64k4 | 140.8 (x1.03) | 18.73 (x1.03) | x1.12 |
| hot96k4 | 138.5 (x1.02) | 18.33 (x1.01) | x1.05 |
| hot64k2 | 137.5 (x1.01) | 18.17 (x1.00) | x1.00 |
| hot64k8 | 163.6 (x1.20) | 22.32 (x1.23) | x1.71 |
| hot32k4a (added) | 118.7 (0.87) | 15.27 (0.84) | 0.93 |
| hot64k4a | 115.4 (0.85) | 14.62 (0.81) | 0.87 |
| hot96k4a | 114.4 (0.84) | 14.56 (0.80) | 0.83 |
Decision (the 0.97 rule): layer 5 is OUT of v3. Neither card keeps even the 32 MiB table resident while the 1 GiB dataset streams (the replaced form gains 1.02 to 1.08x at k = 4 against an ideal 1.33x), and the added form costs 13 to 20%; the chip row moves the wrong way with it. Layer 5 stays a measured option for 3.0 (a table small enough to stay resident, or a different access pattern). The probe rows and the writeup follow on ca2-cache. PC 1 is released to the shipper for PC 1's update-now; the era job starts after PC 1's 0.3.10 STATUS line.

View file

@ -26,7 +26,7 @@ Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGP
| 2 | out | spread across six programs 5.5 to 22.3% per card, over the 5% rule |
| 3 | out (scratch share 0); the construct is sound and its tests stay | the on-die-cache recompute chip stays at 2.4x at every share under the 6 GB cap |
| 4 and 8 | the era draw of stride, interleave and the working-set window, in if the six-era spread is under 5% per card | pending the PC rows |
| 5 | in, in the ADDED form only (16 dataset loads plus k hot loads), size = the largest table resident on every card | the replaced form lets the chip skip item derivations (x1.33 at k = 4); Mac rows hot32k4 x1.22, hot64k4 x1.12 |
| 5 | OUT of v3 (a measured option for 3.0) | the added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% (g 0.87 / 0.85 / 0.84 and 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB against the 0.97 rule): no card keeps even 32 MiB resident while the dataset streams; the replaced form helps the on-die-cache chip (x1.33 at k = 4) |
| 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 |
| 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple |
| M16 mixer x4 | into v3 | the only measured lever that moves the named chip: 0.6x bare, 1.8x with a 3x fixed-function factor; verifier 1.6 to 4.8 ms per warp |