Counter ASIC 2.0: mixer x8 decided into v3 on the PC build rows; status 22:06
This commit is contained in:
parent
08d562a8fb
commit
d76aad40bc
3 changed files with 8 additions and 2 deletions
|
|
@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the
|
|||
|
||||
### 6a. The chip model's headline, and how the scratch share is chosen
|
||||
|
||||
The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. Refinement (coordinator, 21:28 UTC, delegated under "as strong as the measurements allow"): x8 is built and measured beside x4 on the same packs (verifier per warp on one Mac core, the 1 GiB daily build on the 5090, the M5 Max and the 9070 XT or gfx1036, the chip row at equal silicon with the 3x factor); x8 goes into v3 if the per-warp verify stays under 10 ms on one core and the daily build stays under 1 s on every card we own, otherwise x4 with the thin margin stated in level 3 and x8 named as the next lever. The vectors are re-cut once, after this choice. Measured so far: x4 chip row 0.61x bare, 1.84x with the factor, 1.53x with the 128 mm^2 mirror deducted; x8 from the m16 table 0.31x bare, 0.92x with the factor (approximate until measured). The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2).
|
||||
The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. Refinement (coordinator, 21:28 UTC, delegated under "as strong as the measurements allow"): x8 is built and measured beside x4 on the same packs (verifier per warp on one Mac core, the 1 GiB daily build on the 5090, the M5 Max and the 9070 XT or gfx1036, the chip row at equal silicon with the 3x factor); x8 goes into v3 if the per-warp verify stays under 10 ms on one core and the daily build stays under 1 s on every card we own, otherwise x4 with the thin margin stated in level 3 and x8 named as the next lever. DECIDED (5 October 2026, 22:06 UTC, delegated): x8. Both halves pass: the per-warp verify at x8 is 2.79 ms on a loaded M5 Max core (2.1x v2; about 1.3 ms quiet, approximate) against the 10 ms gate; the daily 1 GiB build does not move with the mixer on any discrete card (RTX 5090 23 to 25 ms, RX 9070 XT 72 to 77 ms, M5 Max 21 ms at x1, x4 and x8: latency-bound), 13x to 40x under the 1 s bar (PC 1 jobs fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005, 22:00:08 to 22:04:39Z, both cards restored, every pack's fingerprint equal to the Mac's). V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }. Chip row at x8: 1,198,080 ops per hash, 41.7 MH/s at 50 T op/s, 0.31x bare, 0.92x with the 3x fixed-function factor, 0.76x at equal silicon: the claim reads under 1x with the factor, margin stated. The vectors are re-cut once on this class. Cost: pool shares per core and IBD time scale with the verifier (x8: 2.1x v2); the integrated tier mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2).
|
||||
|
||||
### 6b. The user tiers (the consequences rule)
|
||||
|
||||
|
|
|
|||
|
|
@ -480,3 +480,9 @@ PC 1: fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005 published 21:59:32Z
|
|||
## 22:03 PC 2's prover is back; suite job 3 of 3 published
|
||||
|
||||
socketfix-pc2-pv1 (22:01:14 to 22:02:12Z): /tmp/sp1-cuda-0.sock owned by root removed (the aggregation-cost job's run), the prover switched off and on; the next shard (block 89011 shard 0) "proven and submitted in 34 s" at 22:02:13Z and paid 0.93116546 IGN at 22:02:24Z. PC 2's prover had been dark from 21:25:24Z to 22:02 (the root-socket class; the CI check tools/ci/prover-socket-check.sh now fails any playbook without the two restore lines). Suite job 3 of 3 (consensus-core, igneum-exec, kaspa-pow, igneum-miner, kaspa-p2p-flows and the app tests; main 8ea6740, fork 79bd8e10) is packing and publishing from the ca2-v3 worktree now. After it on PC 2: the prover-floor agent's toolchain check and its 60-minute build, then the aggregation-cost re-run.
|
||||
|
||||
## 22:06 decided: mixer x8 into v3; the verifier fix found; ca2-v3 at 795472e
|
||||
|
||||
The mixer PC 1 job (22:00:08 to 22:04:39Z, both cards restored, the 9070 XT present): the daily 1 GiB build per pack, two dispatches, wall ms: RTX 5090 v2 25 / 23, mx4 24 / 24 and 23 / 25, mx8 23 / 23 and 23 / 23 (cache 4); RX 9070 XT v2 74 / 74, mx4 77 / 75 and 73 / 74, mx8 72 / 76 and 75 / 74 (cache 8 to 9). The build is latency-bound on every card; the x8 rule's build half passes with 13x to 40x margin; its verifier half passed at 2.79 ms per warp on the loaded core. DECIDED (delegated): x8 into v3; V3_CLASS = { era: None, hot: None, ..MX8 }; the chip row at x8 reads 0.92x with the 3x factor (the claim "under 1x with the factor", margin stated). Every mixer pack's fingerprint equals the Mac's on both cards (v2 25f96e7dce90bd4e; mx4 6f48d5a2aa0dbe5f, 73caaebb28e808fe; mx8 7c28cfb06c5c65a9, bbb183f72692f840); hash rates at the v2 rate on both (5090 136.5 to 137.4, 9070 XT 18.0 to 18.2 at every class).
|
||||
|
||||
The verifier regression is found: not the mask but inlining; the item loop inlined into MemhardCpu::fetch runs at 1.33 ms per unit, the same loop out of line (#[inline(never)], one instance per cache size, the line mask a constant) at 0.60 to 0.62 against readwidth's 0.60 to 0.64 in the same minute. The fix, the MX8 V3_CLASS, the pinned pack re-export and the final v2 / x4 / x8 session land as one commit on ca2-v3 795472e (the node agent merged the mixer's 16dfd1e as 4e733bb with two one-line field fixes; 53 + 4 + 19 + 7 tests, packfile 0 failures, the fork check clean). Then: the node agent merges, rebuilds, gate run 3 on the final class; the six era packs and the pinned v3 pack re-exported on it; the final bit-exactness and G2 (1,000 random hashes per card re-hashed by the CPU) job on PC 1.
|
||||
|
|
|
|||
|
|
@ -29,7 +29,7 @@ Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGP
|
|||
| 5 | OUT of v3 (a measured option for 3.0) | the added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% (g 0.87 / 0.85 / 0.84 and 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB against the 0.97 rule): no card keeps even 32 MiB resident while the dataset streams; the replaced form helps the on-die-cache chip (x1.33 at k = 4) |
|
||||
| 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 |
|
||||
| 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple |
|
||||
| M16 mixer x4 | into v3 | the only measured lever that moves the named chip: 0.6x bare, 1.8x with a 3x fixed-function factor; verifier 1.6 to 4.8 ms per warp |
|
||||
| M16 mixer x8 (x4 measured beside it) | into v3 | the only measured lever that moves the named chip: x8 0.31x bare, 0.92x with a 3x fixed-function factor (x4: 0.61x, 1.84x); verifier 2.1x v2 per warp (2.79 ms on a loaded core, about 1.3 ms quiet); the daily build unmoved on every card (latency-bound) |
|
||||
| 9 | reserve-only, no change to the devnet's hour | the floor 600 s from the slowest compile-ahead (38 s) |
|
||||
|
||||
Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors).
|
||||
|
|
|
|||
Loading…
Reference in a new issue