From d8e6ad3341b8f1e32a0d270e243493140aa7ca18 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:28:07 +0000 Subject: [PATCH] Counter ASIC 2.0: the x8 refinement in the rollout plan; status 21:28: hot-table go, installer attempt 3 failed, PC 1 order --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 755a8da48..ab0a9a8c4 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. Refinement (coordinator, 21:28 UTC, delegated under "as strong as the measurements allow"): x8 is built and measured beside x4 on the same packs (verifier per warp on one Mac core, the 1 GiB daily build on the 5090, the M5 Max and the 9070 XT or gfx1036, the chip row at equal silicon with the 3x factor); x8 goes into v3 if the per-warp verify stays under 10 ms on one core and the daily build stays under 1 s on every card we own, otherwise x4 with the thin margin stated in level 3 and x8 named as the next lever. The vectors are re-cut once, after this choice. Measured so far: x4 chip row 0.61x bare, 1.84x with the factor, 1.53x with the 128 mm^2 mirror deducted; x8 from the m16 table 0.31x bare, 0.92x with the factor (approximate until measured). The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ### 6b. The user tiers (the consequences rule) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 2ab098092..c5a14cb2d 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -351,3 +351,5 @@ Consequence for layer 5: unless the 5090 and 9070 XT rows show g at or above 0.9 ## 21:27 decision rule for the mixer: x8 beside x4 Delegated (coordinator, the project lead's "as strong as the measurements allow"): the mixer agent builds LoadClass::MX8 beside MX4, exports mx8-genesis and mx8-devnet-epoch0, runs the Mac bit-exactness, and measures v2, x4 and x8 in one measure-lock session (verifier ms per warp on one core, avg of 50 and worst cold; the 256 MiB fill; the Metal 1 GiB build); the 5090 and AMD daily-build times come from a 2-minute prepare job on PC 1 after the era job. x8 goes into v3 if the per-warp verify stays under 10 ms on one core AND the daily build stays under 1 s on every card we own; else x4 with the thin margin stated (1.84x with the factor) and x8 named as the next lever (0.92x with the factor, from the m16 table). The vectors are re-cut once after the choice. The hot-table rule stands (into v3 only at g >= 0.97 on both PC cards). + +21:28. Installer attempt 3 failed at 21:27:20Z (exit 2 in 4 s: an inline bash -c string lost a quote through PowerShell; the fix is the 0.3.6 cut's: the WSL part as a file run with bash , plus a read-only path probe before attempt 4). By the rule, "go PC 1" went to the hot-table job at 21:28 (10 min). PC 1 order from here: the hot-table job, the shipper's path probe and attempt 4 (about 16 min), the era job, the 5090 power sweep, the mixer daily-build job (2 min), Ember's build, the repro re-run, the AMD sweep (conditional), Ember's run.