From 3e237db050105f49800dd0374c62ba03b62978cd Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:53:45 +0000 Subject: [PATCH] counter-asic-2.md: layer 9 row and the decisions table of 5 October 2026 (night) --- docs/plans/counter-asic-2.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/docs/plans/counter-asic-2.md b/docs/plans/counter-asic-2.md index f6598d19b..8e9f58d86 100644 --- a/docs/plans/counter-asic-2.md +++ b/docs/plans/counter-asic-2.md @@ -16,6 +16,21 @@ Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGP | 6 | Cache growth on the genesis schedule | the SRAM mirror stays unaffordable | none | already in the design; confirm the schedule against SRAM density | | 7 | Integer matrix ops in the program (INT8 x INT8 into INT32, exact) | matrix hardware at GPU scale | none on NVIDIA and AMD; Apple to check | reserved family, not at launch | | 8 | Working-set size drawn per program | one memory design cannot fit every hour | none | folds into 4 and 5 | +| 9 | Epoch length as a signalled era parameter: base 3,600 DAA s, ladder 600 to 7,200, set by 90% signal at a day boundary, lead and `T_epoch` fixed | a per-program bitstream (an FPGA with a hard datapath): at 600 s nothing it compiles ever runs (42 to 160 min per compile, PRflow FPT 2019; hours on large parts, Aldec) | compile-ahead 1 s per epoch on the 5090, 0.5 s on the M5 Max (38 s with the race on); one CPU core `600 / epoch_len` busy on the VDF | reserve-only tonight: `docs/plans/epoch-length.md` | + +## Decided 5 October 2026 (night), under the project lead's delegation for the devnet (`docs/plans/counter-asic-2-rollout.md` section 6) + +| # | Decision | The number that decided it | +|---|---|---| +| 1 | keep v2's 128 x 4 B | w16 passes the rule but closes nothing (5090 139.8 against 136.1 MH/s, 9070 XT 17.90 against 18.15); w64 makes the 5090 bandwidth-bound (share 0.58, 37% of stream) | +| 2 | out | spread across six programs 5.5 to 22.3% per card, over the 5% rule | +| 3 | out (scratch share 0); the construct is sound and its tests stay | the on-die-cache recompute chip stays at 2.4x at every share under the 6 GB cap | +| 4 and 8 | the era draw of stride, interleave and the working-set window, in if the six-era spread is under 5% per card | pending the PC rows | +| 5 | in, in the ADDED form only (16 dataset loads plus k hot loads), size = the largest table resident on every card | the replaced form lets the chip skip item derivations (x1.33 at k = 4); Mac rows hot32k4 x1.22, hot64k4 x1.12 | +| 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 | +| 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple | +| M16 mixer x4 | into v3 | the only measured lever that moves the named chip: 0.6x bare, 1.8x with a 3x fixed-function factor; verifier 1.6 to 4.8 ms per warp | +| 9 | reserve-only, no change to the devnet's hour | the floor 600 s from the slowest compile-ahead (38 s) | Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors).