diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index a03414bf2..4e166ee87 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -34,7 +34,7 @@ Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per h | RTX 5090 (CUDA) | 136.1 | [owed] | 512 + hot | 0.96 at v2 | [owed] | | RX 9070 XT (OpenCL, eGPU) | 18.15 | [owed] | 512 + hot | 0.87 at v2 | [owed] | -AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), near parity per pound and about 4.5x worse per watt; the card's memory system, not a tuning gap. +AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), 2.2x worse per pound at list prices and 4.9x worse per watt (read-width.md section 4.1, approximate); the card's memory system, not a tuning gap. Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant]. diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 7dccb83e0..0ace07a27 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -58,7 +58,7 @@ The layer 6 finding changes the headline: the strongest chip holds the whole cac ### 6b. The user tiers (the consequences rule) -AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second at 1 GiB, measured tonight), near parity per pound and about 4.5x worse per watt (0.089 against 0.398 MH/W, measured 5 October 2026). This is the card's memory system, not a tuning gap: no read width closes it without making the 5090 bandwidth-bound. The level 3 numbers page states it. +AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second at 1 GiB, measured tonight), 2.2x worse per pound at list prices (0.032 against 0.072 MH/s per pound, approximate) and 4.9x worse per watt (read-width.md section 4.1). This is the card's memory system, not a tuning gap: no read width closes it without making the 5090 bandwidth-bound. The level 3 numbers page states it. ## 7. Gates before any publish (all of them, no exceptions) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index eb781fad4..7c0e34b20 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -140,7 +140,7 @@ Layer 9 (Josh: faster program changes): the epoch length becomes an era paramete ca2-era mid-way (a452664c512c73b9b): era draw behind LoadClass::era (EraParams beside mix, load_slots, scratch, scratch_kb; verify::load_index; memhard::Layout; one load form in the three emitters; --era, --era-widths), 54 crate tests green, pinned packs byte-identical; six era packs bit-exact on Metal (packbench 6/6), Apple OpenCL (6/6, same fingerprints) and the CUDA emulation (6/6); re-exporting with the width pinned at 4 B and rebasing onto e752fc7; PC job in about 20 minutes. Two class bugs found and fixed on its branch, both outside its layer: (1) proto-cuda/nvrtc/packfile.h re-derived the seed words as attempt 0 only, so ANY pack with IGNEUM_PROGRAM_ATTEMPT >= 1 (5.14% of chain epochs under v2) is refused by the one-click workers with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", the error PC 1 logged on 5 October and attributed to the export race; fixed attempt-aware, with a tampered-attempt refusal test. This fix must ship in the 0.3.11 workers whatever else does. (2) proto-cuda/host.cu and proto-opencl/host.c derived dataset words on the host as mh_item(w >> 4)[w AND 15] instead of the pack's mh_word; fixed. -Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (20:23, 20:24 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.5x the electricity per hash (0.089 against 0.398 MH/W), near parity per pound; the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. +Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (20:23, 20:24 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.9x the electricity per hash and 2.2x worse per pound at list prices (0.032 against 0.072 MH/s per pound, read-width.md 4.1, approximate); the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. ## 20:29 C1 decided for the morning; one packfile.h fix for 0.3.11 @@ -197,3 +197,5 @@ ca2-node (a3f505a9d981300cd): ca2-v3 commits d2cd6e1 (pack-loop af983a7 merged: Gate item found: main.swift refuses v3 lines "until Swift has generator 3"; the Mac worker never regenerates a program (igneum-miner export-pack writes the pack), so the fix ordered is to accept the pack's class and era against the line, as the CUDA and OpenCL workers do. Without it Mac node 1 and the Mac app cannot mine v3 (gates G1 and G4). PC 1: the AMD-proving small fixture is running; Ember Tune's 8-minute CPU-only build takes the next CPU gap, its 30-minute both-cards run is job 5. + +20:50. Consequences round 2 (C18 to C20). C18: "near parity per pound" was wrong and is struck everywhere; read-width.md section 4.1 gives the 5090 at 2.2x the 9070 XT per pound at list (0.072 against 0.032 MH/s per pound, approximate), 4.9x per watt, 7.5x in rate; the level 3 page carries those. C19 (to ca2-mixer, already sent by the reviewer): the x4 verifier cost in IBD minutes over the 108,000-header pruning window per tier (8.6 min against 1.1 on one M5 Max core at the top of the range), pool shares per core per second, a scaled 2019-class figure (approximate), and the 10 ms gate margin left for 3.0 go into mixer-x4.md; the seeds' header-verify load goes into the testnet go checklist. C20 (to ca2-epoch, already sent): both rig miners run --exit-on-seed-change and re-export on exit 42, so a 10-minute epoch restarts every card's miner six times an hour and the Mac fleet's prepare pause goes from 35 s to 3.5 min an hour; the epoch-length document gets a per-tier restart-cost row and the compile-ahead margin against the VDF at 600 DAA s, and the rig installer drops the exit-42 path for prepare-ahead before any short epoch can be drawn (next-cut list).