Counter ASIC 2.0 status: agents respawned, first readwidth Mac rows

This commit is contained in:
igneum-labs 2026-10-05 19:56:18 +00:00
parent 68e9c9fd29
commit 14303e9bf8

View file

@ -6,12 +6,12 @@ Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rew
| Layer | Branch | Agent | State | Numbers so far | Blockers |
|---|---|---|---|---|---|
| 1, 2, 3 (widths, mix, scratch) | readwidth 019b014 | a451c9935bfb1bc19 (not ours) | packs and verifier done; Mac Metal and Apple OpenCL runs next, then one job per PC | CPU verifier, M5 Max one core, avg of 50, ms per warp: v2 0.604, w16 0.610, w64 0.630, w64x4 0.160, mixA 0.620, mixB 0.614, scr0 0.600, scr2 0.534, scr4 0.457, scr8 0.314 | holds the Mac measure lock and both PCs first |
| 4 + 8 (era layout, working set) | ca2-era | spawning | design | none | PC time behind readwidth |
| 5 (cache-sized second table) | ca2-cache | spawning | design | none | PC time behind readwidth |
| 6 (SRAM schedule) | ca2-analysis | spawning | cited analysis | none | none |
| 7 (integer matrix family) | ca2-analysis | spawning | design; dp4a throughput owed unless PC time frees | none | Apple int8 path to check against Metal docs |
| 3 soundness | ca2-soundness | spawning | analysis + tests | none | none |
| 1, 2, 3 (widths, mix, scratch) | readwidth 019b014 | a451c9935bfb1bc19 (not ours) | Mac Metal rows in; Apple OpenCL next, then one job per PC; scratch re-run at 32 and 128 KB per warp | CPU verifier, M5 Max one core, avg of 50, ms per warp: v2 0.604, w16 0.610, w64 0.630, w64x4 0.160, mixA 0.620, mixB 0.614, scr0 0.600, scr2 0.534, scr4 0.457, scr8 0.314. Metal M5 Max hash rate (5 x 2^24, GPU time, 3 of 3 vectors on every pack): v2 27.7 MH/s, w16 28.3, w64 28.2, w64x4 109.7 (32 loads of 64 B), mix 50/35/15 over 6 programs 25.4 to 28.4 (median 26.7), mix 25/50/25 over 6 programs 22.0 to 25.2 (median 24.4). Scratch at 1 MiB per warp, 1,024 warps: 23.1 / 18.2 / 16.3 / 14.9 MH/s at 0 / 12.5 / 25 / 50% RMW, superseded by the cap | holds the Mac measure lock and both PCs first |
| 4 + 8 (era layout, working set) | ca2-era | a452664c512c73b9b | design, respawned 19:56 | none | PC time behind readwidth |
| 5 (cache-sized second table) | ca2-cache | a5271cf269757b118 | design, respawned 19:57 | none | PC time behind readwidth |
| 6 (SRAM schedule) | ca2-analysis | a5c6bc2dfcc4613ef (respawned 19:58) | cited analysis | none | none |
| 7 (integer matrix family) | ca2-analysis | a5c6bc2dfcc4613ef | design; dp4a throughput owed unless PC time frees | none | Apple int8 path to check against Metal docs |
| 3 soundness | ca2-soundness | a548aadeefd1ab3b2 (spawned 19:59; sizes 32 and 128 KB per warp) | analysis + tests | none | none |
| Integration v3 | ca2-v3 | after the above | waiting | none | the readwidth table and the four branches |
Budget rule received from the coordinator at the restart: the per-warp scratch is capped so the whole working set (1 GiB table + hot table + scratch for every resident warp + buffers) stays under 6 GB on an 8 GB card, which puts the scratch in the tens of KB per warp; the layer 5 hot table shares that budget.