Counter ASIC 3.0 status: the lanes' own 00:00 lines, the spec-text re-derivation ordered, the class v5 packs test's derivation-text diff and its disposition, the one 0.3.24 kit named (e6c088bb); chip model section 5: the second partial-store correction from adv-cache-2's window-layer reading

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 00:11:21 +00:00
parent 28bde0f17b
commit fbbd1b7c16
2 changed files with 3 additions and 1 deletions

View file

@ -205,6 +205,8 @@ Per row: reads per hash = 128 f; items recomputed = 128 (1 - f); ops per hash =
Correction, 7 October 2026 (the in-house adversarial pass, lane adv-cache-3, report 091edc34): the partial-store rows above and adv-cache's Q1b table price a chip that holds every k-th line of the 64-line chain and recomputes a read at offset o in o evaluations ((k - 1) / 2 on average). The exact pebbling optimum for the chain (dynamic programming, checked against exhaustive search at 10 to 16 lines) sits under that curve: blocks per read 16.0 against 31.5 at f = 1/64 (the one held line belongs at line 32, not line 0), 10.5 against 15.5 at 2/64, 6.09 against 7.5 at 4/64, 3.17 against 3.5 at 8/64, 1.45 against 1.5 at 16/64, equal from f = 1/2. So a chip holding 1/64 of the cache pays 9.3x the item's ops, not 17.4x; at f = 1/2 and above nothing moves, and the SRAM column and the full-store verdict stand (no point on the curve beats the full store under the op budget or under energy).
Memory-bound rate = the ceiling / (128 f). Compute-bound rate = 50 T op/s / ops per hash (the section 1 budget).
Second correction, 8 October 2026 (the in-house adversarial pass, lane adv-cache-2, report section 2.3 and its Q3(3) window-layer reading, tip 3f50d6c4): the partial-store rows model a chip that holds a fraction f of the ITEMS chosen uniformly, so it serves f of the reads and recomputes 1 - f. The item-read distribution is not uniform: the exact window-layer distribution (matching the 4,096-program census to four digits; top quarter mean 0.3382, top half 0.5811 of reads) lets a chip holding the hottest f of items serve 0.4219, 0.7188 and 0.8907 of reads at f = 0.25, 0.5 and 0.75 on the measured programs, so the f = 0.25 and f = 0.5 rows overstate the recompute share by up to 2.3x (1.8x on the first shard's read) and the f = 0.75 row by about 2.3x on the miss side. The f = 1 row, the SRAM column and the full-store verdict do not move (a chip that holds everything recomputes nothing either way), and no served number rests on f under 1; the partial-store rows stay as the uniform-store bound with this note until the hottest-f rows are drawn from the window distribution, which is the next pass of this section.
The rate is the smaller; "binding" names it. Power = rate x (128 f x E_read + 128 (1 - f) x 6.3 nJ) + static (memory,
controller, and 20 W for the recompute die's clocks and leakage when `f < 1`). Energy per hash = power / rate. "Gain,
rate" = rate / 136.1 MH/s (per chip, the section 2 metric); "with the 3x factor" multiplies the compute-bound rate by

File diff suppressed because one or more lines are too long