From 41afd097cd357a1043d696f666095542ec89de10 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:20:53 +0000 Subject: [PATCH] Counter ASIC 2.0 status 21:20: the added-form hot table on the Mac, the decision rule for layer 5 --- docs/plans/counter-asic-2-status.md | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index d9e8e7510..8eb71e997 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -305,3 +305,15 @@ Design (reserve-only): epoch_len base 3,600, ladder 600 to 7,200, SIGNAL ONLY (t | Radeon iGPU under load (M11) | 124 s with the dataset | 20.7% (needs per-day dataset reuse before any signal below the base) | FPGA citations: PRflow (FPT 2019) 42 min typical, 160 min worst for a monolithic Vivado compile; Aldec hours on Virtex UltraScale; partial reconfiguration milliseconds per region (ICAP 400 MB/s) shortens the load, not the compile; at 600 s a per-program bitstream mines 0% of each epoch, 47% at 3,600. Riders: the race defaults off (M11: base wins on both cards; 6.3% at the floor); per-day dataset reuse in the workers for the iGPU tier. Owed: the 9070 XT clBuildProgram time, the difficulty settle at a 15% step and 600-s epochs in sim.py, the 3-bit signal encoding against Kaspa's version bits (spec 5.8). + +## 21:20 the hot table in the added form, measured on the Mac: the big-die chip comes out ahead + +ca2-cache (hot-table.md 6.2 to 6.4): M5 Max, 21:03 to 21:19 UTC, load average 7 to 14, Metal packbench 2^24 x 5 (GPU time) and Apple OpenCL --bench-pack; v2 in the same session 27.63 / 27.59 MH/s. + +| Pack | Metal / OpenCL MH/s | g against v2 (probe predicted) | Fingerprint (2^20, base 0) | Verifier ms per warp (v2 0.602) | Fill, one core | +|---|---|---|---|---|---| +| hot32k4a | 25.76 / 25.72 | 0.93 (0.96) | 8a3414735db4523c | 0.631 | 21.7 ms | +| hot64k4a | 23.92 / 23.87 | 0.87 (0.94) | 45668f34105f6307 | 0.609 | 43.3 ms | +| hot96k4a | 22.92 / 22.88 | 0.83 (0.93) | af763997dfee4c82 | 0.614 | 64.9 ms | + +Reading: on Apple the added form costs 7 to 17% of the rate for four extra loads per iteration, more than the probe predicts as the table grows; only the 32 MiB table is near free. Chip arithmetic with these g: a chip serving H from DRAM keeps 0.86 / 0.92 / 0.96 of its gain; a 100 mm^2 die with the SRAM keeps 0.96 / 0.92 / 0.88; a 750 mm^2 die with the SRAM comes out 6 to 15% AHEAD, because the GPU pays the hits in rate and a big die pays them in 1.6 to 4.8% of area. So on the Mac's numbers layer 5 does not pass its own test; the decision waits for the 5090 (96 MiB L2) and 9070 XT (64 MB Infinity Cache) rows, where the hits may be near free (g close to 1). Rule for the decision: layer 5 goes into v3 only if, on every card we own, g is at or above 0.97 at the chosen size AND the on-die-cache chip row (chip-model-v3.md) moves down with it; otherwise layer 5 is out of v3 and stays a measured option for 3.0.