diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index d36a93004..a944248d8 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ### 6b. The user tiers (the consequences rule) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 65f20f2f7..2fd7f81b4 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -1,6 +1,6 @@ # Counter ASIC 2.0: status -Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rewritten every 45 minutes while the work runs. Times UTC, 5 October 2026 (night). Base for every ca2 branch: `readwidth` at 019b014 (the LoadClass flag, fold_words, the scratch op, the three emitters, 20 packs). +Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rewritten every 45 minutes while the work runs. Times UTC, 5 October 2026 (night). Heading times before 20:40 were corrected at 20:42 from the commit clock (the coordinator had written them from a guessed clock, up to 2 h 40 min ahead); every entry's true time is its commit's author time in UTC. Base for every ca2 branch: `readwidth` at 019b014 (the LoadClass flag, fold_words, the scratch op, the three emitters, 20 packs). ## 19:55 first status (the 19:50 start was cut off by exhausted credits at about 19:58 before any sub-agent work landed; respawned at 19:55 on the restart) @@ -35,7 +35,7 @@ Finding that bears on layer 3 (readwidth, Metal, M5 Max, capped sizes): the scra Reading (readwidth agent): 4,096 warps x 32 KB = 128 MB sits in the chip's caches. Consequence for the decision: a scratch that fits a GPU's cache fits a chip's SRAM at the same size, so at the capped size the writes cost everyone a cache-bound op in place of a latency-bound load. Passed to the soundness agent: what size would make the writes cost DRAM latency, whether that fits the 6 GB cap, and whether the RMW share should be added to the 16 dataset loads rather than taken from them. -## 20:15 mandate: v3 on the devnet tonight, by the project lead's rules +## 20:08 mandate: v3 on the devnet tonight, by the project lead's rules the project lead has gone to bed and delegated the three decisions for the DEVNET only (not the public testnet): width, per-load mix and scratch share by the rules now written in `docs/plans/counter-asic-2-rollout.md` section 6, the activation height = devnet tip + 14,400 at publish (checked >= 10,800), published the way finality v3 was. Six gates before any publish (rollout section 7): bit-exact v3 on the three cards; the CPU verifier exact on 1,000 random GPU hashes per card; the soundness suite green with the new scratch tests; the fast-time 3-node network mining across a v3 activation with 0 rejected blocks and 0 forks; Windows and Mac workers from one commit; the node change on a fork from the 0.3.10 tip (21d4c73c) with suites green on PC 2. Release 0.3.11 through the shipper's pipeline; the 0.3.10 shipper (ae892a8b0f78fe31c) has been asked for its state and the handoff. If a gate fails: stop, write why here, do not publish. @@ -43,7 +43,7 @@ Node build note for the integration: the node links `igneum-pow` by path (`../.. After the publish: Counter ASIC 3.0 from the ASIC-history agent's ranked additions (a202a09dcd24ba1d3), as class v4 behind its own activation, same gates, `docs/plans/counter-asic-3.md`. -## 20:25 the shipper's answer, the node fork convention +## 20:08 the shipper's answer, the node fork convention 0.3.10 (shipper ae892a8b0f78fe31c): staged and blocked on GitHub's Actions incident (run 37365130137 queued since 19:42:43Z under a re-dispatching watcher); nothing on the network has moved, the live manifest is still 0.3.9. Once CI is green: ship (5 min), update-now (apps restart 1 to 10 min later), hand nodes and seed (5 min), digest sweep; 0.3.10 finished about 30 min after green. The app version per machine on the console's cards is the restart signal for re-running any straddling measurement. HiveOS is published by the ship's --public step, nothing separate. @@ -51,7 +51,7 @@ Node fork for v3: base on COMMIT 21d4c73c (release-0.3.10 in vendor/igneum-node 0.3.11 shipper: the coordinator assigns it (the 0.3.10 shipper stops at its report). Inputs the ship needs: the fork commit with its PC 2 suite results recorded, the main tip, the override object with every switch (the four live fields plus program_class_v3_activation_daa), the expected digest read on a 22-s scratch node, the deadline note ("program class v3"), the activation height, the one-line changelog, and whether the pinned proving guest changes (it does not: the prover has no igneum-pow dependency; confirmed by grep of proving/igneum-prove Cargo files). -## 20:35 readwidth round 2 on the PCs; the width arithmetic under the project lead's rule +## 20:10 readwidth round 2 on the PCs; the width arithmetic under the project lead's rule Round 1 of the readwidth PC jobs refused every pack (the workers demand a 32-byte chain seed; the experiment packs carried string seeds); fixed at readwidth 1ea7a52 (packfile.h), republished 20:09:19Z as `run-readwidth-5090-20261005c` (PC 2, about 6 min) and `run-readwidth-9070-20261005c` (PC 1, about 10 min). The era and cache agents were told to rebase onto 1ea7a52 and to prove their packs load on the Mac OpenCL host before any PC job. Every ca2 branch now bases on 1ea7a52. @@ -64,9 +64,9 @@ Probe ceilings from round 1 (dependent reads per second at 1024 MiB, device time the project lead's width rule applied to the ceilings alone (the measured v3 rates will replace this when the table lands): a 128-load hash at 64 B reads 8,192 B; at the 5090's 64 B ceiling that is 71 MH/s and 584 GB/s, 37% of its stream bandwidth, over the one-third margin the rule sets; at 16 B it is 2,048 B per hash, 141 MH/s and 288 GB/s, 18%, inside the margin; 4 B is 9%. On the 9070 XT every width costs the same line fetch (2.4 G/s), so 16 B is where AMD gains 4x the bytes per hash at no cost and the 5090 stays latency-bound with margin. Provisional width under the rule: 16 B (w16), pending the measured rates and the latency-bound share per card. -Added deliverable (20:40): the public description in four levels, `docs/plans/counter-asic-2-public.md` (aec53bb): levels 1 and 2 are written as copy; level 3 carries the bench table with `[owed]` markers for every number not yet measured; level 4 lists the documents. The integration branch applies levels 1 to 3 to `site/index.html`, `site/litepaper.html` and `site/bench.html` with the final numbers. +Added deliverable (20:11): the public description in four levels, `docs/plans/counter-asic-2-public.md` (aec53bb): levels 1 and 2 are written as copy; level 3 carries the bench table with `[owed]` markers for every number not yet measured; level 4 lists the documents. The integration branch applies levels 1 to 3 to `site/index.html`, `site/litepaper.html` and `site/bench.html` with the final numbers. -## 20:55 layers 6 and 7 landed; the node-fork agent started; 0.3.11 scope +## 20:15 layers 6 and 7 landed; the node-fork agent started; 0.3.11 scope Branch ca2-analysis (5d5ba15, f59708d). Layer 6: no cache growth rule exists in the spec; cited bit cells N7 0.027, N5 / N3E / Intel 18A 0.021, N3B 0.0199, N2 0.0175 um^2, array factor 0.70; the 256 MiB mirror is 83 / 64 / 54 mm^2 at N7 / N5 / N2 (74 at N2 with a 96 MB hot table), $13 to $26 of silicon per good die (approximate). The mirror was never unaffordable; the cache's job is to stay above GPU L2 (5090 96 MB, GB202 128 MB). Recommendation C for the project lead (gate 1): cache doubles when the dataset doubles. Layer 7: Metal 4 matmul2d has int8 x int8 -> int32, so a unit-level mm8 tile is native on all three vendors; per-lane dot4 is emulation on Apple (M5 Max: 548 G unsigned dot4/s against an 880 G ALU chain, 1.6x; signed 4.7x). Reserve R1 = mm8, W_new 4, unlock era 4 or 90% signal. Owed: the 5090 and 9070 XT dot4 probe (job prepared: relay/playbooks/dot4-probe.ps1, exe sha256 5adaeb1a...6416f4; publishes when a PC frees). @@ -74,21 +74,21 @@ Node-fork agent a3f505a9d981300cd started 20:50: ca2-v3 (igneum-pow seam: Progra 0.3.11 scope (coordinator, 20:52): carries program_class_v3 and proving v1 together (one override object, one digest, one publish; each half under the same gates); the ready half ships as 0.3.11 and the other as 0.3.12 if one lags. Next-cut list recorded in the rollout plan section 8a. -## 21:05 decisions recorded: layer 6 option C, layer 7 R1 = mm8; the chip headline +## 20:16 decisions recorded: layer 6 option C, layer 7 R1 = mm8; the chip headline Layer 6 DECIDED (delegated; the project lead confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles (256 MiB genesis, 512 MiB year 4, 1 GiB year 12); one-core fill 0.2 / 0.4 / 0.8 s, under 1 s at every step. Layer 7 DECIDED: reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal. The chip model's headline now names the on-die-cache recompute chip (54 to 83 mm^2) as a row per variant; the scratch share is chosen as the smallest share at which that chip's gain falls under 1.5x, else said plainly and the public "under 2x" claim qualified. The soundness agent carries that table (its question 2) with M16's mixer multiplier beside it. -## 21:15 correction to the SRAM mirror figures (chip-economics research, cluster D) +## 20:18 correction to the SRAM mirror figures (chip-economics research, cluster D) Bit cell x 0.70 understates real die area. Shipped cache dies: AMD 3D V-Cache 64 MB on 41 mm^2 at 7 nm (1.56 MB/mm^2, Tom's Hardware, Hot Chips August 2021); Graphcore GC200 900 MB on 823 mm^2 with compute (1.09 MB/mm^2); Groq TSP 220 MB on 725 mm^2 at 14 nm (0.30 MB/mm^2); TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30% assist overhead (SemiAnalysis, December 2022). A 256 MiB mirror is about 165 mm^2 at 7 nm on the densest shipped cache-only die and about 130 mm^2 at N5/N3E, not 54 to 83 mm^2; cost per die 2 to 3x the earlier figure; the conclusion (affordable for a funded chip) stands. The analysis agent is redoing the table with both columns; the soundness agent carries the corrected density into the chip row. Latency citations behind the latency-bound rule, to be added: DRAM row cycle 40 to 48 ns across DDR4, GDDR5, HBM2 (Li, Reddy, Jacob, MEMSYS 2018); latency 1.3x in two decades against bandwidth 20x (Chang 2017); no shipped mining chip used HBM or stacked memory. -## 21:25 proving v1 state for 0.3.11; PC 2 occupancy +## 20:19 proving v1 state for 0.3.11; PC 2 occupancy Proving v1 (acd4f36bc2c07a4e2): fork proving-v1 b177718e on a24ab01a (told to rebase onto commit 21d4c73c now), app proving-v1 79bc820 on a93199a. Override fields proving_v1_activation_daa (tip + 14,400 at publish), proving_v1_segment_blocks 4, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Harness: `tools/proving-v1/net.mjs --secs 1500` PASSED (21 checks) in 197.3 s on b177718e; rerun owed on the final tree. The pinned guests do not change. Shared files with ca2-v3-node: params.rs, daemon.rs, igneum/miner/src/main.rs, override-60x.json; both agents keep separable hunks. PC 2 is held by the proving agent's memsweep-pc2-pv1 (about 20 min, miners stopped) and a second run (about 10 min). Queue after it: the ca2 node suites, then the readwidth, era, cache and dot4 measurement jobs. PC 1 is held by readwidth's run-readwidth-9070-20261005c until it reports. -## 21:35 sram-mirror.md revision 2 (ca2-analysis) +## 20:21 sram-mirror.md revision 2 (ca2-analysis) Two columns, headline = shipped-product density (AMD V-Cache 41 mm^2 per 64 MiB at N7, scaled by the bit-cell ratio), lower bound = bit cell x 0.70. mm^2 and $ per good die (D0 0.1 per cm^2, wafer prices approximate), headline / lower bound: @@ -101,7 +101,7 @@ Two columns, headline = shipped-product density (AMD V-Cache 41 mm^2 per 64 MiB Year 10 at the 6% per year trend: 59 mm^2 for the flat cache (82 with the hot table), 8% of a 750 mm^2 die. One reticle holds 1.3 GiB (N7) to 1.9 GiB (N2); mirror share of a 750 mm^2 die at year 0: 14% (22% with the hot table), inside M16's 13 to 40% band. Recommendation unchanged: C. Latency section added (MEMSYS 2018, Chang 2017, the mining-chip memory-type note), marked as research the agent did not re-read tonight apart from the V-Cache figure. -## 21:50 layer 3 soundness landed: the scratch does not move the chip; scratch share decided 0 +## 20:23 layer 3 soundness landed: the scratch does not move the chip; scratch share decided 0 ca2-soundness (0d8f745 tests and trace hook, a465881 doc and bench-log). The on-die-cache recompute chip (N5 headline 128 mm^2, $46) at 50 T op/s: 333 MH/s against the 5090's measured 139.7, 2.4x; at 12.5 / 25 / 50% RMW replaced, chip 381 / 443 / 661 against 5090 projected 160 / 186 / 279, 2.4x each; added, 2.4x or more; 32 or 128 KB alike. The chip keeps the scratch implicitly in 80 to 320 B per lane (the verifier resets it per unit), needs about 530 units in flight, dense scratch 6.2 / 3.1 mm^2 at N5. Under the project lead's rule the scratch share is 0: layer 3 is NOT adopted into v3; the public "under 2x" claim is qualified (public copy level 3 rewritten). The measured lever is M16's mixer multiplier (x2 1.2x at 0.8 to 2.4 ms verify; x4 0.6x at 1.6 to 4.8 ms; 3.6x and 1.8x with a 3x fixed-function factor); whether x4 enters v3 tonight is asked of the coordinator; default: Counter ASIC 3.0. @@ -109,13 +109,13 @@ Soundness results (Metal, M5 Max): 28/28 edge launches, 200/200 fuzz packs (91 s Gate G3 note: the scratch tests (igneum-pow/tests/scratch.rs) join the v3 suite even though the class carries no scratch, parametric over the class; they guard the v2 path's scratch-free invariant at zero cost. -## 22:00 decided: M16 mixer x4 into v3; agent ca2-mixer started +## 20:24 decided: M16 mixer x4 into v3; agent ca2-mixer started Coordinator's decision under the project lead's delegation (recorded in the rollout plan section 6a): the mixer multiplier x4 and the cache growth rule (option C) enter class v3 behind the same activation; layer 3 stays out at scratch share 0, its soundness document and pack-contract tests kept. Agent af345b1e2c541ffbb (branch ca2-mixer) implements `mixer_mult` as a class parameter (m mixer applications per round, the 8 dependent reads unchanged), the `cache_log2_words(day)` schedule (doublings at years 4 and 12 with the dataset stepping to the next power of two), re-cuts the v3 dataset vectors, re-runs the soundness suite, measures the verifier (v2 0.604 ms per warp; v3 expected 1.6 to 4.8 ms) and the 1 GiB build on the Mac, prepares the 5090 and 9070 XT build-time job, and writes docs/analysis/chip-model-v3.md with the combined headline row (fixed-function factor included). The claim on the site reads "under 2x" only if that row does; else qualified, with the mixer x8 and the hot table named as the next levers. Agents now: ca2-era (a452664c512c73b9b), ca2-cache (a5271cf269757b118), ca2-node (a3f505a9d981300cd), ca2-mixer (af345b1e2c541ffbb). Done: ca2-analysis, ca2-soundness. Waiting: the readwidth PC table; PC 2 (proving memsweep runs) and PC 1 (readwidth 9070 round). -## 22:15 the readwidth table landed; layers 1, 2, 3 decided; layer 5 measured on the Mac and redesigned +## 20:27 the readwidth table landed; layers 1, 2, 3 decided; layer 5 measured on the Mac and redesigned Readwidth e752fc7 (`docs/plans/read-width.md`), bit-exact on Metal, Apple OpenCL, the 5090 (NVRTC) and the 9070 XT, both PCs released. MH/s (latency-bound share): @@ -134,31 +134,31 @@ Decisions (the project lead's rules, delegated): layer 1 keep v2 (w16 passes the Layer 5 (ca2-cache 53ef59f, 011cc0a, 86726cd, 65bc7a7; `docs/plans/hot-table.md`): five packs bit-exact on Metal and Apple OpenCL (96/96 each). M5 Max rates against v2 27.68: hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k2 x1.00, hot64k8 x1.71; verifier 0.344 to 0.560 ms against 0.626; hot fill per epoch 24 / 46 / 73 ms on one core, 0.07 / 0.15 / 0.22 ms on the GPU; Apple OpenCL probe 32 / 64 / 96 / 1024 MiB 21.7 / 12.8 / 12.3 / 3.50 G loads/s. Redesign ordered: hot loads ADDED beside the 16 dataset loads (the replaced form lets the on-die-cache chip skip item derivations and worsens the gain); the agent re-measures the added form and rebuilds the PC job. PC 1 is given to the dot4 probe (under 15 min), then to the era agent, then the hot-table job; PC 2 stays the proving agent's. -## 22:30 layer 9 added; the era draw passes Mac bit-exactness; two class bugs in the harnesses; C1 +## 20:27 layer 9 added; the era draw passes Mac bit-exactness; two class bugs in the harnesses; C1 Layer 9 (the project lead: faster program changes): the epoch length becomes an era parameter in the genesis reserve, 1 hour at launch, 10 minutes to 2 hours by draw or 90% signal, reserve-only tonight; an agent (ca2-epoch) designs it beside layers 4 and 8 and measures the compile-ahead cost per card at a 10-minute epoch, the seed-path consequence and the FPGA threat it answers; one row in the rollout plan section 6, one in the level 3 numbers. Spawns when the dot4 probe frees its slot. ca2-era mid-way (a452664c512c73b9b): era draw behind LoadClass::era (EraParams beside mix, load_slots, scratch, scratch_kb; verify::load_index; memhard::Layout; one load form in the three emitters; --era, --era-widths), 54 crate tests green, pinned packs byte-identical; six era packs bit-exact on Metal (packbench 6/6), Apple OpenCL (6/6, same fingerprints) and the CUDA emulation (6/6); re-exporting with the width pinned at 4 B and rebasing onto e752fc7; PC job in about 20 minutes. Two class bugs found and fixed on its branch, both outside its layer: (1) proto-cuda/nvrtc/packfile.h re-derived the seed words as attempt 0 only, so ANY pack with IGNEUM_PROGRAM_ATTEMPT >= 1 (5.14% of chain epochs under v2) is refused by the one-click workers with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", the error PC 1 logged on 5 October and attributed to the export race; fixed attempt-aware, with a tampered-attempt refusal test. This fix must ship in the 0.3.11 workers whatever else does. (2) proto-cuda/host.cu and proto-opencl/host.c derived dataset words on the host as mh_item(w >> 4)[w AND 15] instead of the pack's mh_word; fixed. -Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (21:50, 22:00 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.5x the electricity per hash (0.089 against 0.398 MH/W), near parity per pound; the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. +Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (20:23, 20:24 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.5x the electricity per hash (0.089 against 0.398 MH/W), near parity per pound; the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. -## 22:40 C1 decided for the morning; one packfile.h fix for 0.3.11 +## 20:29 C1 decided for the morning; one packfile.h fix for 0.3.11 C1 (the fee switch H = 210,000 at about 19:50Z on 6 October; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before 19:50Z. The proving agent recommends (b) unless (a) is certain; the check decides it. The attempt-0 packfile.h bug (the era agent's find) is the one that took the fleet down at 18:23Z (epoch 34, attempt 1); it is already fixed attempt-aware on pack-loop af983a7 and merged into the 0.3.10 tree. Rule for 0.3.11: one derivation, one test set: the era and cache agents build their workers on that packfile.h and keep only tests that add a case; the host.cu / host.c mh_word derivation fix (new, the era agent's) stays with its own test. -## 22:50 card lifetime merged; cache freed after the build; growth mapping (b) recommended +## 20:31 card lifetime merged; cache freed after the build; growth mapping (b) recommended `docs/analysis/card-lifetime-2026-10-05.md` (branch card-lifetime 1fecfe2) merged into ca2-coord. Decided (delegated): the GPU frees the 256 / 512 / 1,024 MiB cache after the daily dataset build (the hash never reads it; the rebuild costs 0.67 ms fill + 13.4 ms build on the 5090 at x1, about 54 ms at x4, owed); hot-table.md's resident reading is corrected, era-layout.md's freed reading stands. Recommended for the project lead: growth mapping (b), power-of-two steps at years 4, 12, 28, 60 with AND MASK, the only mapping under which the litepaper's "4 GB about four years, 8 GB more than a decade" holds (4 GB: year 4 under (b), 1.0 to 1.5 years under (a); 8 GB: year 12 or 6.3 to 7.5; 12 GB: year 28 with the cache freed; 24 GB: year 60). Public lines and the evidence row go on the integration branch (rollout plan 6c); hot-table.md line 73 (the 8 GB row counted a 5090's warps) goes to the cache agent. -## 23:00 layer 9 agent started; the dot4 probe is on PC 1 +## 20:32 layer 9 agent started; the dot4 probe is on PC 1 ca2-epoch (a32a3ece66c02417a): the epoch length as an era parameter (600 to 7,200 DAA s, base 3,600; draw or 90% signal; the VDF rule; the difficulty-window constraint; the FPGA threat with citations), Mac compile-ahead measured now, the 5090 and 9070 XT compile times cited from the bench log, docs/plans/epoch-length.md. The dot4 probe job is running on PC 1 (ca2-analysis tip ee42d7c carries the playbook; the 5090 confirmation is the agent's watch). PC 1 queue after it: the era six-pack job, then the hot-table added-form job. PC 2: the proving agent's, then the ca2 node suites. Agents running: ca2-era, ca2-cache, ca2-node, ca2-mixer, ca2-epoch; ca2-analysis watching its PC job. Done: ca2-soundness. -## 23:10 layer 7 complete on all three cards (ca2-analysis ee42d7c); PC 1 free +## 20:38 layer 7 complete on all three cards (ca2-analysis ee42d7c); PC 1 free dot4 probe on PC 1 (jobs fetch-dot4-20261005, run-dot4-20261005, exit 0, 101 s; both cards restored and mining; 5090 SM clock 2,505 MHz before and after): @@ -173,7 +173,7 @@ All bit-exact against the CPU reference. One dp4a costs about one ALU step on NV PC 1 is free: the era six-pack job goes next when its package arrives, then the hot-table added-form job. -## 23:20 PC 1 scheduler (the coordinator's role from now): the queue +## 20:39 PC 1 scheduler (the coordinator's role from now): the queue Rule: one PC 1 job at a time; an agent asks by message before publishing, gets a "go PC 1" from this coordinator, and reports when its RESULT lines are in and both cards are restored; a hash-rate or power number taken while another job holds a card is not a number. The CPU-only job runs only in a slot where no measurement overlaps it. @@ -188,4 +188,4 @@ Rule: one PC 1 job at a time; an agent asks by message before publishing, gets a If job 1's package is more than 15 minutes away when job 2 is ready, job 2 goes first; the short job 3 fills any gap of under 10 minutes between packages. -23:30. No more agents are spawned tonight (the project lead: no unnecessary credits); the running ones finish. PC 1 queue change: the AMD-proving CPU fallback's small fixture (block-56-transfers-3shards, minutes) runs NOW in the gap before the era package; its S_p shard (block-338-shard1, up to 30 min of every core) stays job 6, last. Next-cut list gains the rig installer's two follow-ups (the Linux manifest entry, the Linux prover build), tied to whichever release carries the proving half (rollout plan 8a). +20:39. No more agents are spawned tonight (the project lead: no unnecessary credits); the running ones finish. PC 1 queue change: the AMD-proving CPU fallback's small fixture (block-56-transfers-3shards, minutes) runs NOW in the gap before the era package; its S_p shard (block-338-shard1, up to 30 min of every core) stays job 6, last. Next-cut list gains the rig installer's two follow-ups (the Linux manifest entry, the Linux prover build), tied to whichever release carries the proving half (rollout plan 8a).