From 14303e9bf883071fcf8e65f210789cf4cb0e3e63 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 19:56:18 +0000 Subject: [PATCH] Counter ASIC 2.0 status: agents respawned, first readwidth Mac rows --- docs/plans/counter-asic-2-status.md | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 74a6bc4c4..6ccc74b14 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -6,12 +6,12 @@ Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rew | Layer | Branch | Agent | State | Numbers so far | Blockers | |---|---|---|---|---|---| -| 1, 2, 3 (widths, mix, scratch) | readwidth 019b014 | a451c9935bfb1bc19 (not ours) | packs and verifier done; Mac Metal and Apple OpenCL runs next, then one job per PC | CPU verifier, M5 Max one core, avg of 50, ms per warp: v2 0.604, w16 0.610, w64 0.630, w64x4 0.160, mixA 0.620, mixB 0.614, scr0 0.600, scr2 0.534, scr4 0.457, scr8 0.314 | holds the Mac measure lock and both PCs first | -| 4 + 8 (era layout, working set) | ca2-era | spawning | design | none | PC time behind readwidth | -| 5 (cache-sized second table) | ca2-cache | spawning | design | none | PC time behind readwidth | -| 6 (SRAM schedule) | ca2-analysis | spawning | cited analysis | none | none | -| 7 (integer matrix family) | ca2-analysis | spawning | design; dp4a throughput owed unless PC time frees | none | Apple int8 path to check against Metal docs | -| 3 soundness | ca2-soundness | spawning | analysis + tests | none | none | +| 1, 2, 3 (widths, mix, scratch) | readwidth 019b014 | a451c9935bfb1bc19 (not ours) | Mac Metal rows in; Apple OpenCL next, then one job per PC; scratch re-run at 32 and 128 KB per warp | CPU verifier, M5 Max one core, avg of 50, ms per warp: v2 0.604, w16 0.610, w64 0.630, w64x4 0.160, mixA 0.620, mixB 0.614, scr0 0.600, scr2 0.534, scr4 0.457, scr8 0.314. Metal M5 Max hash rate (5 x 2^24, GPU time, 3 of 3 vectors on every pack): v2 27.7 MH/s, w16 28.3, w64 28.2, w64x4 109.7 (32 loads of 64 B), mix 50/35/15 over 6 programs 25.4 to 28.4 (median 26.7), mix 25/50/25 over 6 programs 22.0 to 25.2 (median 24.4). Scratch at 1 MiB per warp, 1,024 warps: 23.1 / 18.2 / 16.3 / 14.9 MH/s at 0 / 12.5 / 25 / 50% RMW, superseded by the cap | holds the Mac measure lock and both PCs first | +| 4 + 8 (era layout, working set) | ca2-era | a452664c512c73b9b | design, respawned 19:56 | none | PC time behind readwidth | +| 5 (cache-sized second table) | ca2-cache | a5271cf269757b118 | design, respawned 19:57 | none | PC time behind readwidth | +| 6 (SRAM schedule) | ca2-analysis | a5c6bc2dfcc4613ef (respawned 19:58) | cited analysis | none | none | +| 7 (integer matrix family) | ca2-analysis | a5c6bc2dfcc4613ef | design; dp4a throughput owed unless PC time frees | none | Apple int8 path to check against Metal docs | +| 3 soundness | ca2-soundness | a548aadeefd1ab3b2 (spawned 19:59; sizes 32 and 128 KB per warp) | analysis + tests | none | none | | Integration v3 | ca2-v3 | after the above | waiting | none | the readwidth table and the four branches | Budget rule received from the coordinator at the restart: the per-warp scratch is capped so the whole working set (1 GiB table + hot table + scratch for every resident warp + buffers) stays under 6 GB on an 8 GB card, which puts the scratch in the tens of KB per warp; the layer 5 hot table shares that budget.