From 9eb6b5e654e290ea31f167dd3245e99ae1add5a7 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:36:30 +0000 Subject: [PATCH] Counter ASIC 2.0: layer 5 decided OUT on the PC rows; status 21:36 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 17 +++++++++++++++++ docs/plans/counter-asic-2.md | 2 +- 3 files changed, 19 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 1ed5a42d4..4f95f216c 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -115,7 +115,7 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa |---|---|---|---| | The width rule (layers 1 and 2) | 4 B fixed; 16 B; 64 B; the per-program mix | `` | `docs/plans/read-width.md` | | The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `` | the same | -| The hot table (layer 5) | 32, 64, 96 MB; replaced or added | In the ADDED form only (16 dataset loads plus k hot loads): the replaced form lets the on-die-cache chip skip item derivations (k = 4: chip x1.33 against the GPU's measured x1.05 to x1.22). Size = the largest table resident on every card we own, pending the PC rows; Mac rows (replaced form, M5 Max): hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k8 x1.71; Apple OpenCL dependent-read probe 32 / 64 / 96 / 1024 MiB: 21.7 / 12.8 / 12.3 / 3.50 G loads/s | `docs/plans/hot-table.md` (ca2-cache 65bc7a7) | +| The hot table (layer 5) | 32, 64, 96 MB; replaced or added | DECIDED (5 October 2026, 21:36 UTC, delegated): OUT of v3. The rule was g at or above 0.97 on both PC cards in the added form; measured g is 0.87 / 0.85 / 0.84 on the 5090 and 0.84 / 0.81 / 0.80 on the 9070 XT at 32 / 64 / 96 MiB: neither card keeps even the 32 MiB table resident while the dataset streams, and the replaced form helps the on-die-cache chip. Layer 5 stays a measured option for 3.0 | `docs/plans/hot-table.md` (ca2-cache): 5090 MH/s v2 136.1; replaced hot32k4 146.6, hot64k4 140.8, hot96k4 138.5, hot64k2 137.5, hot64k8 163.6; added hot32k4a 118.7, hot64k4a 115.4, hot96k4a 114.4. 9070 XT v2 18.15; replaced 19.79 / 18.73 / 18.33 / 18.17 / 22.32; added 15.27 / 14.62 / 14.56. M5 Max added 0.93 / 0.87 / 0.83. All eight packs bit-exact on the 5090 and the 9070 XT with the Mac's fingerprints (21:29 to 21:35 UTC, no restart straddled) | | The era draws (layers 4 and 8) | in, if the min-to-max spread across six drawn eras is under 5% per card; item size fixed at 4 B, draws of stride, interleave and the working-set window at or above 256 MiB | pending the six-era table | `docs/plans/era-layout.md` | | The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate | | Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index b2b88c697..5cd2bc6d5 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -367,3 +367,20 @@ Node agent: the Metal worker's ServeStore already keys datasets by day and progr Gate G4: the first fast-time run failed at node start on the script, not the node: JSON.parse turned a never height (18446744073709551615) into 1.8446744073709552e+19 and the node refused the override; fixed as text merging in class-v3.mjs and simnet.mjs (d5ff532; no other script in tools/, infra/ or sim/ has the shape). The gate runs again on the mixer-x4 class (binaries from 79bd8e10 + ca2-v3 66eeba3). Main-repo tip for the suite job title: d5ff532 (era and cache not yet merged). 21:34. PC 2: agg-cost-pc2-1 closed 21:25:11Z (done, miners and prover back on); agg-cost-pc2-2 went out at 21:33Z (a missed close), self-limited to 16.5 min, closes about 21:51Z; the 0.3.11 suites publish after it AND after PC 2's app shows the 0.3.10 STATUS line (the update-now goes to PC 2 now). PC 1: the hot-table job runs (release about 21:40); PC 1's update-now follows the release; the era job starts after PC 1's 0.3.10 STATUS line. + +## 21:36 layer 5 measured on both PC cards: OUT of v3 + +Hot-table jobs on PC 1 (fetch 21:29:01Z; run-ca2-hot-5090 126 s, closed 21:31:38Z; run-ca2-hot-9070 247 s, closed 21:35:39Z; both cards restored; the 9070 XT WAS on the bus, full path; app 0.3.9 throughout, no straddle). All eight packs bit-exact on the 5090 and the 9070 XT with the Mac's fingerprints. + +| Pack | RTX 5090 MH/s (v2 136.1) | RX 9070 XT MH/s (v2 18.15) | M5 Max g | +|---|---|---|---| +| hot32k4 (replaced) | 146.6 (x1.08) | 19.79 (x1.09) | x1.22 | +| hot64k4 | 140.8 (x1.03) | 18.73 (x1.03) | x1.12 | +| hot96k4 | 138.5 (x1.02) | 18.33 (x1.01) | x1.05 | +| hot64k2 | 137.5 (x1.01) | 18.17 (x1.00) | x1.00 | +| hot64k8 | 163.6 (x1.20) | 22.32 (x1.23) | x1.71 | +| hot32k4a (added) | 118.7 (0.87) | 15.27 (0.84) | 0.93 | +| hot64k4a | 115.4 (0.85) | 14.62 (0.81) | 0.87 | +| hot96k4a | 114.4 (0.84) | 14.56 (0.80) | 0.83 | + +Decision (the 0.97 rule): layer 5 is OUT of v3. Neither card keeps even the 32 MiB table resident while the 1 GiB dataset streams (the replaced form gains 1.02 to 1.08x at k = 4 against an ideal 1.33x), and the added form costs 13 to 20%; the chip row moves the wrong way with it. Layer 5 stays a measured option for 3.0 (a table small enough to stay resident, or a different access pattern). The probe rows and the writeup follow on ca2-cache. PC 1 is released to the shipper for PC 1's update-now; the era job starts after PC 1's 0.3.10 STATUS line. diff --git a/docs/plans/counter-asic-2.md b/docs/plans/counter-asic-2.md index 8e9f58d86..542ef627c 100644 --- a/docs/plans/counter-asic-2.md +++ b/docs/plans/counter-asic-2.md @@ -26,7 +26,7 @@ Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGP | 2 | out | spread across six programs 5.5 to 22.3% per card, over the 5% rule | | 3 | out (scratch share 0); the construct is sound and its tests stay | the on-die-cache recompute chip stays at 2.4x at every share under the 6 GB cap | | 4 and 8 | the era draw of stride, interleave and the working-set window, in if the six-era spread is under 5% per card | pending the PC rows | -| 5 | in, in the ADDED form only (16 dataset loads plus k hot loads), size = the largest table resident on every card | the replaced form lets the chip skip item derivations (x1.33 at k = 4); Mac rows hot32k4 x1.22, hot64k4 x1.12 | +| 5 | OUT of v3 (a measured option for 3.0) | the added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% (g 0.87 / 0.85 / 0.84 and 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB against the 0.97 rule): no card keeps even 32 MiB resident while the dataset streams; the replaced form helps the on-die-cache chip (x1.33 at k = 4) | | 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 | | 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple | | M16 mixer x4 | into v3 | the only measured lever that moves the named chip: 0.6x bare, 1.8x with a 3x fixed-function factor; verifier 1.6 to 4.8 ms per warp |