igneum/docs/plans/counter-asic-2-rollout.md

42 KiB

Generator class v3 on the live devnet: rollout plan (5 October 2026, night)

Scope: the DEVNET only. the project lead delegated the three decisions for the devnet before going to bed (5 October 2026, about 20:10 UTC, through the coordinator): "Counter ASIC 2.0 fully deployed" tonight. The public testnet is not open; its genesis takes v3 from day one. The devnet is ours and a reset is acceptable.

Shape and rules follow docs/plans/finality-v3-rollout-devnet.md and the publish record docs/plans/finality-v3-devnet-publish.md: one height switch read from the override file, every node carries the same object before the height, the PCs get igneumd only through an OTA app version, the activation height leaves at least three hours from the manifest publish. Nothing in this file has run on the devnet. The numbers marked <...> are filled by the integration branch ca2-v3 and the decisions of section 6; the plan is published with them, not before.

1. What changes and what does not

Only the lottery hash's program class changes, and only from the first epoch at or above the height. One switch, program_class_v3_activation_daa, in Params and OverrideParams like difficulty_v2_activation_daa; default u64::MAX (never) on every network. Because one epoch has one program (spec 01 section 1.12), the switch keys on the EPOCH: epoch e is class v3 when 3,600 e >= N4, so the activation is rounded up to an epoch boundary and a block's class is a function of its DAA score alone, as today.

Class v3 = generator version 3: the width rule <W> (layer 1 or 2, decision 1), no scratch (layer 3 decided out: scratch share 0), the era draw of the table layout and the working set (layers 4 and 8), the hot table of <S> MB from the epoch seed (layer 5), the cache growth rule of layer 6 (option C) and the M16 mixer x4 in the dataset item construction (decided 5 October 2026, delegated). New program id (generator = 3 in the id's preimage, spec 1.4.6), new packs and vectors, new IGNEUM_GENERATOR in every pack, a pack of the other version refused by every implementation (spec 1.4.5 already says so).

What does not change: the chain, the genesis, the databases, the day key and the 256 MiB cache fill, the dataset items (spec 1.8.5, if the era interleave keeps the item values; the era-layout document says what it costs otherwise), finality, fees, proving. The SP1 guest does not read the lottery hash (proving/igneum-prove has no dependency on igneum-pow; the pinned guest of DAA 210,000 is a fee-table switch), so no new guest is pinned. The node's EpochSeeds gains the class and the era bytes; IgneumEngine::epoch_for builds the v3 Epoch from them; the miner's export-pack and the serve protocol's job line carry the class so a GPU worker regenerates the right pack from the seed bytes.

The consensus digest (Params::consensus_digest, ledger X18) covers every activation height, so the new field enters the digest and every node must carry the same object before any node reaches the height; a node without the field is refused at the handshake once the others carry it, which is the protection the digest exists for. Binary rollout first (the digest flips when the binary carries the field at never), the height second.

2. The binaries

Built from <node branch> at <commit> on the PCs through tools/build-job.mjs (standing rule 5 October 2026), the Mac binary under the build lock. The table is filled at build time: platform, path, sha256, how it was verified (--version, the switch's first line on a private suffix, strings carries the field name).

3. The activation height N4, and how every node learns it

N4 = DAA at the manifest publish + 10,800 at least, chosen as DAA now + 14,400 rounded up to the next epoch boundary (a multiple of 3,600), checked at publish (N4 - DAA >= 10,800). The packaged line carries every switch:

NODE_OVERRIDE_PARAMS='{"difficulty_v2_activation_daa":33000,"proving_v0_activation_daa":84100,"fees_v1_activation_daa":210000,"finality_v3_activation_daa":135200,"program_class_v3_activation_daa":N4,"proving_v1_activation_daa":N5,"proving_v1_segment_blocks":8,"proving_v1_unproven_daa":600,"proving_v1_aggregator_share_bps":1000}'

The same nine-field object goes verbatim into the override files of Mac node 1, the observer node and the seed, and into the manifest's consensus.override (publish-manifest.sh --override). N5 = DAA at publish + 14,400 (the proving v1 switch, no rounding). Two digests, read on the 0.3.11 Mac node (igneumd bd7f043c..., fork 89dfcb95, 22:5x UTC): with no override file c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (the rolling-upgrade digest, equal to the node agent's pinned test); with the nine-field object at N4 = N5 = 154,800 the ACTIVATION digest 0139ab9dc2992d449ec787d8f021974933631eb55740ab4b6ce9d5c226e72888, the value every node must print after the publish; the node logs "Program class v3 ... active from epoch 43 (DAA 154800 ... epochs of 3600)" and "Proving v1 ... paid from DAA 154800, 8 blocks a segment, unproven after 600 DAA, aggregator share 1000 bps". Binaries from release-0.3.11 23bc2b2: Igneum-Miner-0.3.11.dmg b7e81d4f6f3af9f9179e29faa72c844b795df757dfb2e4cd1d56cd454e78e1e7 (41,592,041 bytes; the packaged json carries the nine fields, read back from the image); the seed's Linux igneumd 63cf490d... (glibc 2.34, 89dfcb95 inside). The era seed for the devnet: the stand-in of docs/plans/era-layout.md (the hash of the last selected-chain block below 15,552,000 n - 7,200; era 0 on the devnet uses the genesis hash), until the 1-hour VDF of spec 4.4 is in the node.

4. The order

  1. The digest flip: a node build that carries program_class_v3_activation_daa at never on every node (hand nodes and the seed first: infra/devnet/restart-hand-nodes.sh '<object without the new field>', then the app version through the manifest; every node prints the same Consensus params digest).
  2. Fix N4, cut the app version (packaging/mac/packaged-config.sh, the three version files), commit as igneum-labs.
  3. The Windows payload inputs (packaging/windows/push-inputs.sh) and the Mac DMG (packaging/mac/build-dmg.sh), the manifest (publish-manifest.sh --activation-height N4 --deadline-note "program class v3" --override '<object>' --deploy), fetch-ci-artifacts.sh --deploy.
  4. The observer, the seed, Mac node 1 with the object; each node's first lines show every switch and Program class v3 from the override file: active from epoch <N4 / 3,600>.
  5. HiveOS: packaging/hive/make-hive-package.sh republished with the v3 igneum-miner and workers, same version string as the apps.
  6. The watch: before N4 - 1,800 both PCs on the new app version (STATUS lines); at the boundary every miner's prepare of the v3 pack (the hot-swap entry's shape) and the first v3 block's program id on the observer; the hash rate per card against the measured v3 numbers of docs/plans/counter-asic-2.md's table; zero pack refused lines; the CPU verifier time per block in the node log against the measured ms per warp.

4a. Two rules for the publish and every PC job (C32)

  • The 0.3.11 update-now goes to PC 2 only after the prover-floor agent's server build (floor-build-3, under /opt/igneum-floor in WSL2, published 22:27 UTC, 25 to 90 minutes) has closed: an app restart ends the running job. The update-now takes a machine list (as 0.3.10's did): the Mac, the laptop and PC 1 first, PC 2 last.

  • Re-fetch after every app update: an app update clears the jobs folder (the 0.3.10 install at 21:49 UTC took PC 1's AMD kit with it), so every fetch-then-run pair re-publishes its fetch after an update, and every run playbook opens with a presence check of its kit that fails with "kit missing: republish the fetch after the app update".

  • No PC job raises an elevation prompt on either PC for the rest of the night (C35): both unexplained app quits tonight came 20 to 41 s after an administrator prompt beside the running installed app (PC 2 20:00:49 to 20:01:09Z, PC 1 22:30:25 to 22:31:06Z), the engine has no self-relaunch after a quit, and nobody is at a keyboard; the two-minute test of that class (raise one prompt from a job while the app mines, cancel it, read the quit line, which ember-tune b671c8b stamps with its source) runs in the morning with the project lead present, or tonight only after the relay relaunch path is proven and the rollout is done. The sweeps and Ember's run stay held; the floor, aggregation-cost and M16 jobs raise no prompt.

5. Rollback

Before N4: remove the field on every node and restart; nothing has happened (the digest flips back, so every node at once). After N4: there is no rollback by restart, because blocks mined under v3 verify only under v3. A rollback is a second height switch back to v2 at a later epoch, carried the same way. This is why the measurements of the plan come first.

6. The decisions, by the project lead's rules (devnet)

the project lead's rules, applied by the coordinator and recorded here with the number that decided each:

Decision the project lead's rule Choice The number
Width (layer 1) the widest read that keeps every card we own latency-bound (achieved loads within 90% of the probe ceiling) with margin on the 5090 (its bytes per hash under a third of its bandwidth at the measured rate) DECIDED (5 October 2026, delegated): keep v2, 128 x 4 B. w16 passes the rule (shares 0.90 / 0.84 / 1.03, 18% of the 5090's stream) but closes nothing and does not move the chip row; w64 and w64x4 make the 5090 bandwidth-bound (share 0.58 / 0.56, 37% of stream) docs/plans/read-width.md (readwidth e752fc7): v2 5090 136.1 MH/s, 9070 XT 18.15, M5 Max 27.74 (gap 7.5x); w16 139.8 / 17.90 / 28.26 (gap 7.8x); w64 71.9 / 17.59 / 28.27 (gap 4.1x); the 9070 XT does 2.4 G dependent reads/s at every width
Per-load mix (layer 2) in, if the min-to-max spread across six programs is under 5% per card DECIDED (5 October 2026, delegated): out spreads of the median over six programs: mix 50/35/15 5090 18.8%, 9070 XT 7.4%, M5 Max 11.3%; mix 25/50/25 22.3% / 5.5% / 8.1%
Scratch share (layer 3) the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap DECIDED (5 October 2026, delegated): 0. No share under the cap moves the on-die-cache recompute chip, so layer 3 is not adopted into v3 docs/analysis/scratch-soundness.md (ca2-soundness a465881): chip 333 MH/s against the 5090's measured 139.7 = 2.4x at 0% RMW; 2.4x at 12.5 / 25 / 50% replaced (chip 381 / 443 / 661 against 160 / 186 / 279 projected) and 2.4x or more added, at 32 and 128 KB; the chip keeps the scratch implicitly in 80 to 320 B per lane because the verifier resets it per unit
Activation height N4 devnet tip + 14,400 at publish, checked >= 10,800 at publish, rounded up to the epoch boundary <at publish>

6a. The chip model's headline, and how the scratch share is chosen

The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; docs/analysis/sram-mirror.md after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. Refinement (coordinator, 21:28 UTC, delegated under "as strong as the measurements allow"): x8 is built and measured beside x4 on the same packs (verifier per warp on one Mac core, the 1 GiB daily build on the 5090, the M5 Max and the 9070 XT or gfx1036, the chip row at equal silicon with the 3x factor); x8 goes into v3 if the per-warp verify stays under 10 ms on one core and the daily build stays under 1 s on every card we own, otherwise x4 with the thin margin stated in level 3 and x8 named as the next lever. DECIDED (5 October 2026, 22:06 UTC, delegated): x8. Both halves pass: the per-warp verify at x8 is 2.79 ms on a loaded M5 Max core (2.1x v2; about 1.3 ms quiet, approximate) against the 10 ms gate; the daily 1 GiB build does not move with the mixer on any discrete card (RTX 5090 23 to 25 ms, RX 9070 XT 72 to 77 ms, M5 Max 21 ms at x1, x4 and x8: latency-bound), 13x to 40x under the 1 s bar (PC 1 jobs fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005, 22:00:08 to 22:04:39Z, both cards restored, every pack's fingerprint equal to the Mac's). V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }. Chip row at x8: 1,198,080 ops per hash, 41.7 MH/s at 50 T op/s, 0.31x bare, 0.92x with the 3x fixed-function factor, 0.76x at equal silicon: the claim reads under 1x with the factor, margin 8% on the factor and 9% on the budget (chip-model-v3.md). Verifier on the fixed crate (ca2-mixer 1ab8b21, 22:07 UTC, same input beside readwidth's binary, load 5.5): v2 0.609 / 0.611 ms per unit (readwidth 0.607 / 0.610), x4 1.238 / 1.237 (2.0x, worst cold 1.40), x8 2.077 / 2.058 (3.4x, worst cold 2.15): 4.8x inside the 10 ms gate on this loaded core. The vectors are re-cut once on this class. Cost: pool shares per core and IBD time scale with the verifier (x8: 2.1x v2); the integrated tier mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). The scratch-soundness analysis carries this table (docs/analysis/scratch-soundness.md, question 2).

6b. The user tiers (the consequences rule)

AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second at 1 GiB, measured tonight), 2.2x worse per pound at list prices (0.032 against 0.072 MH/s per pound, approximate) and 4.9x worse per watt (read-width.md section 4.1). This is the card's memory system, not a tuning gap: no read width closes it without making the 5090 bandwidth-bound. The level 3 numbers page states it.

7. Gates before any publish (all of them, no exceptions)

# Gate Evidence required State
G1 bit-exact v3 on all three vendors against the Mac reference GREEN on the final class (job run-ca2-era-pc1b-20261005, 22:16 to 22:21Z, exit 0 in 304 s, both cards restored, app 0.3.10, the 9070 XT present as gfx1201): the seven final-class packs' 2^24 fingerprints equal on the RTX 5090 (CUDA/NVRTC), the RX 9070 XT (OpenCL) and the M5 Max (Metal): mx8-devnet-epoch0 90f794dd556f7a3b, era-0 8e8e070db4eea52d, era-1 891c01b8563bb47e, era-2 e54279fed2831b5d, era-3 77e0ba8abbd0ae62, era-4 d898d8f4f2e7684b, era-5 a6927db380f7efb2; self-test PASS on every pack on both cards. RULING (coordinator, 5 October 2026, about 21:10 UTC): the integrated gfx1036 (RDNA 2, AMD OpenCL 3683.0) satisfies the AMD vendor tonight, because G1 is a compiler-and-ISA property and gfx1036 carried the v1 and v2 conformance; the 9070 XT's hash-rate and power rows are owed and taken when its link is back GREEN on the final class (run-ca2-era-pc1b-20261005, 22:16 to 22:21 UTC)
G2 the CPU verifier exact on 1,000 random hashes per card GREEN: one serve-mode job of 1,024 nonces at target ff..ff per card and pack (every nonce a found line), re-hashed on the Mac with igneum-pow hash-bound --prehash 00..01 --count 1024 on the same pack: RTX 5090 era-0 1,024 of 1,024 and mx8-devnet-epoch0 1,024 of 1,024; RX 9070 XT era-0 1,024 of 1,024 and mx8-devnet-epoch0 1,024 of 1,024 (the same job) GREEN
G3 the generator soundness suite green, the new scratch tests included cargo test in igneum-pow, tests/packs.rs, the Metal fuzz, edge, stats, determinism runs on the v3 class GREEN (the crate suite 53 + 4 + 19 + 7 on ca2-mixer 1ab8b21 and the release tree; the Metal fuzz 200 of 200 and 50 of 50 on x8, the edge, stats and determinism runs; the scratch tests 7 of 7; 22:12 UTC)
G4 the fast-time 3-node network mining across a v3 activation 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. docs/plans/counter-asic-2-node.md section 5; summary docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms Run 3 PASS on the FINAL class (22:15:04 to 22:19:44Z, binaries from ca2-v3 d233fa1 = x8 + era + the verifier fix, fork 89dfcb95): 182 / 122 blocks around DAA 180, 304 in all; v3 ids e3 a6523b90cff501e3, e4 bc811b3c4b8b1ced, e5 e784541f19cdebe5 on all three miners; epoch 0's v2 id 8f8806638d59850f the same in all three runs; rejected 0/0/0; one sink a9ce45df8beeaf13 at 303/303/303; 3 of 3 switch lines; cache ready v2 179 ms, first v3 191 ms. Summary docs/plans/counter-asic-2-gate/class-v3-20261005-2215Z-era-mx8.json GREEN (runs 1, 2 and 3; run 3 on the final class)
G5 the PC-built Windows workers and the Mac workers from the same commit From release-0.3.11 23bc2b2: igneum-worker-cuda.exe 2b3b8c92885442179f6bf2907c6f3eb453dc4a19908d90fd05981a09b7c2674c (1,536,512 bytes), igneum-worker-opencl.exe edc4a75da3b93d814caa69fd635010780d63d5b622ec24c3741d433c584f91e3 (478,208), both with the resource block, different from 0.3.10's pair; the Mac worker and the DMG from the same tree (the shipper's step report) GREEN at the workers; the DMG and the node builds in flight
G4b the Mac mines v3: a real Metal miner across a v3 boundary through the miner's --prepare-packs flow, and the app passes that flag to the Metal worker Found 22:21 UTC: app/igneum-app/src/engine.rs miner_args pushes --prepare-packs only when card.worker != "Metal" (line 1468 on a223ca9), so the Mac app never hands its Metal worker a prepare pack, and under class v3 the Metal worker compiles v3 only from a prepared pack (servePackProgram): at the first v3 epoch every Mac would answer need lines and stop, the 18:23Z outage class. Two closes before the ship, both in hand: (a) the app fix on 0.3.11's app branch (push the flag for every worker with the platform's path separator, a unit test on miner_args for a Metal card; assigned to the proving agent on top of a223ca9); (b) gate run 4: the fast-time network with one real Metal miner on the Mac across the activation (the prepare lines on both sides, the worker's prepared line for the v3 pack, found or accepted blocks on v3, no need or mismatch line; assigned to the node agent). If (a) is not in the app tree at the cut, the ship does not go: a Mac that cannot mine v3 at activation is a fleet outage, and the activation height (tip + 14,400) is not far enough to carry the fix in 0.3.12 safely GREEN. (b) gate 4 (22:25 to 22:31Z, a real Metal miner on node 0, igneum-bench from ca2-v3 00c55aa): three v3 PREPARE lines with the pack dir and class=v3 era=; the worker's v3 prepared lines (252.5 ms the first: program 55.5, dataset 196.9, cache fill 0.9, build 41.1; then 51 and 46 ms with the day resident); every swap "with no pause"; 124 blocks accepted on v3 (301 in the run), cpu re-check mismatched 0, need 0, no mismatch or refusal, no exit 42 or 44; chain 182 / 123 across DAA 180, 0 rejected, one sink. A second outage found and fixed before the run: the Metal worker's serveDataset built every day with the Swift version 2 construction keyed by day only, so a v3 program would have hashed over an x1 dataset; now a v3 prepare builds the day from the pack's memhard.metal and the store keys datasets by (day, class, era), commit 00c55aa
G6 the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 build-20261005-221237 (published 22:12:37Z: main d233fa1 = era + cache + the mixer fix with V3_CLASS = MX8, fork 89dfcb95) ran the five crates and the app on the final tree: every stage ok (22:12:37 to 22:15:44Z): kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8, 0 failed. With job 2 (kaspa-consensus alone 97) G6 is GREEN on the final tree

If any gate fails: stop at that gate, write why in docs/plans/counter-asic-2-status.md, do not publish.

8. The release

0.3.11 through the shipper's pipeline (tools/ship-app.mjs, the plan shape of docs/plans/release-0.3.10.md). The 0.3.10 shipper finishes 0.3.10 first, then takes the 0.3.11 tree, or the coordinator ships 0.3.11 with the same runbook if the shipper has stopped. The publish order is that of docs/plans/finality-v3-devnet-publish.md: the override field, the digest handshake, hand nodes and the seed first, then the manifest with --activation-height and the deadline note, then the apps, then the digest sweep, then the HiveOS package republish.

6c. Card lifetime: cache residency and the growth mapping (from docs/analysis/card-lifetime-2026-10-05.md, merged)

Decision Choice The number
Cache residency on the GPU DECIDED (5 October 2026, delegated): the cache is FREED after the daily dataset build; the hash reads the dataset and the hot table only, never the cache. hot-table.md's resident reading is corrected to this The daily rebuild is the only cost: cache fill 0.67 ms and dataset build 13.4 ms on the RTX 5090 at x1 (bench-log, 3 October 2026), 2 ms and 13 to 30 ms on the M5 Max; at mixer x4 the build is about 54 ms (measurement owed on ca2-mixer). Freeing it moves the 12 GB tier from year 12 to year 28 under mapping (b)
Dataset growth mapping (a: continuous 2 + 0.5 GiB a year with a multiply-shift index; b: power-of-two steps at years 4, 12, 28, 60 with AND MASK) RECOMMENDED for the project lead: (b). It keeps AND MASK and every vector's size, it is what option C's "doubles when the dataset doubles" already assumes, and it is the only mapping under which the litepaper's "4 GB about four years, 8 GB more than a decade" is true (under (a) a 4 GB card is out within 1 to 1.5 years, an 8 GB card at 6 to 7.5 years) card-lifetime table 2: 4 GB out at year 4 (b) or 1.0 to 1.5 (a); 8 GB year 12 or 6.3 to 7.5; 12 GB year 28 freed or 12 resident; 24 GB year 60 freed

Public lines to fix on the integration branch (card-lifetime table 3): site/index.html "2 GB, growing" gains the rate ("2 GB at genesis, doubling at years 4, 12 and 28"); "Any 4 GB card" becomes "any 4 GB card at launch, 8 GB from year 4"; the litepaper's "4 GB about four years, 8 GB more than a decade" stays with mapping (b) and gains "under the step schedule"; docs/evidence.md gains a row for the card-lifetime claim labelled designed. hot-table.md line 73's 8 GB row is corrected (it counted a 5090's 8,160 warps; a real 8 GB card has 20 to 24 SMs).

7b. Hardware events (for the morning summary)

When (UTC) Event What the app did For the project lead
5 October, at install (earlier today) the RX 9070 XT in the Sonnet Breakaway Box 850T5 over USB4 went Code 43 came back after a driver reinstall and a reboot
5 October, about 20:40 the 9070 XT dropped off PC 1's bus: Get-PnpDevice -Class Display lists only the integrated AMD Radeon Graphics (gfx1036) and the RTX 5090; after pnputil /scan-devices at 20:45:34Z the card is still absent and the USB4 list shows only the host and root routers: the "USB4 Router (2.0), Sonnet Technologies Breakaway Box 850T5" present at 17:18Z is gone, so the box itself is off the link the AMD worker (igneum-worker-opencl --device 1) mines the gfx1036 at 3.12 MH/s; the 5090 keeps mining; nobody was woken, PC 1's app was not restarted; a 10-second rescan probe (pnputil /scan-devices, the USB4 router status) was granted the second eGPU link fault today: reseat the USB4 cable and the eGPU's power; the 0.3.10 hot-plug code shows the card as "removed" and picks it up again without a restart
5 October, 21:22:59 to 22:21 the card dropped again at 21:22:59 (the third drop), was back and used by the hot-table (21:31 to 21:35), the era (21:46 to 21:51 and 22:16 to 22:21) and the mixer (22:00 to 22:04) jobs as gfx1201, and was gone again by 22:24:09 (the fourth drop; no Sonnet or USB4 router device) the OpenCL worker falls to the gfx1036 (--device 1) when the card is gone; every measurement that names gfx1201 ran while it was present the link flaps on a scale of tens of minutes: reseat the USB4 cable and the eGPU's power, try another port or cable; the AMD clock and power sweep is owed on this
5 October, by 21:09 the 9070 XT is back on PC 1's bus: the reproducible-benchmark package's OpenCL worker listed it as opencl:1 and the app switched it off and on through api/cards; no restart, nobody touched the box the app's AMD worker returns to it at the next prepare the link drops and returns by itself; the reseat is still worth doing in the morning
5 October, by 21:22:59 the third drop: no Sonnet or USB4 Router (2.0) device present, the display list shows only the gfx1036 and the 5090 the app lists nvidia:0 and amd:1:gfx1036; the 5090 keeps mining at 124.5 MH/s (app log STATUS lines through 21:23:26Z) the link is flapping: reseat the USB4 cable and the eGPU's power in the morning, and consider a different USB4 port or cable; every AMD measurement tonight runs on the gfx1036 fallback unless the card is present at the job's own probe

7a. A dated constraint from the consequences review (C1)

The fee switch H = 210,000 arrives about 18:45Z on 6 October (DAA 137,041 at 22:32:54Z on 5 October; 1.002 DAA/s averaged since 15:40Z; the 19:50Z in fee-switch-devnet.md is an hour late). The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H.

8a. Proving v1 rides with it

the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL e0de2ab on 5b0d54f (a223ca9 the resume fix, e0de2ab the Metal --prepare-packs fix; docs-only commits between) (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, cargo test -p igneum-app resume 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; cargo test -p igneum-app provedefault 6 of 6 on the Mac). Harness on the final fork tree: IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8 gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which.

Merge rule for the ship (C31): the app branch proving-v1 (e0de2ab) rewrote docs/evidence.md row 16 (WITHDRAWN, the 24 GB measurement) and the litepaper's proving sentences; ca2-coord carries the older row 16 and its own litepaper edits, so at the merge take proving-v1's row 16 and its proving sentences, and ca2-coord's everything else; the stale row must not win by accident. fud-close (the ledger closer's main-repo branch: 45 public-text fixes on the site and litepaper, spec 8.3 and 8.8, two CI checks, the relay fixes; a merge-tree onto ca2-coord shows 0 conflicts) is NOT in 0.3.11: the tree closed at 23bc2b2 (its workers, DMG and PC build job carry it) before the branch reached the ship order, and the shipper takes no late branch (the 0.3.10 rule); fud-close heads the next cut's list, with a coupling the next cut must respect: fud-close's worker change for ledger M28 (packfile.h's kernel_sha256 check, host.c refusing a pack that fails it) pairs with the fork-side miner change on ledger-fixes 3d4ec451 that stamps the hashes into program.json; the workers without that miner commit refuse every pack, the miner without the workers is harmless, so both go in one cut or the worker half of M28 is held back. fud-close tip 647b08c (its checks green); ledger-fixes is not yet rebased onto 89dfcb95 (two conflicting files: igneum/miner/src/main.rs, protocol/flows/src/ibd/proof.rs). The fork-side ledger-fixes branch (from release-0.3.6, not 21d4c73c) is NOT in 0.3.11: it rebases onto the 0.3.11 fork for the next cut. Branches merged into the 0.3.11 main tree beside ca2-v3 and ca2-coord (tooling, no consensus): consequences (the reviewer's rows), bash-body-check (7adb1ca, 6805125: tools/ci/bash-body-check.sh with fixtures and a self-test in ci.yml, the convention paragraph in packaging/README-ship.md, publish-jobs.sh add --kind run running the bash-body and prover-socket checks before signing; add/add with proving-v1 c2544be on tools/ci/prover-socket-check.sh: take bash-body-check's superset and proving-v1's ci.yml step; tools/amd-prove/pc1-cpu-prove.ps1 goes on the allow list, it starts no GPU server), amd-prove (f1d7a7d), card-lifetime (1fecfe2). Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), ember-tune 8ab9068 (the second-engine rule: no pipe into a second engine, its process tree killed at the end and on the budget, the installed app's miners restarted after; applied to ember-tune-pc1.ps1 and sweep-5090.ps1; tools/ci/second-engine-check.sh in ci.yml, shown to fire on a known-bad playbook and pass the fixed pair; every Cmd::Quit stamped with its source; the AMD gmax offset fix); rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = publish-public.sh --hive adds a platforms.linux entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' --exit-on-seed-change path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths.

9. After the publish: Counter ASIC 3.0

The ASIC-history agent sends its ranked additions; they are measured the same way, folded into the class as v4 behind its own activation, same gates, same rollout; docs/plans/counter-asic-3.md.

10. Decisions table (superseded by section 6; kept for the options)

Decision Options Recommendation Source
The width rule (layers 1 and 2) 4 B fixed; 16 B; 64 B; the per-program mix <the readwidth table's decision rule: the widest read that keeps every card latency-bound with margin on the 5090> docs/plans/read-width.md
The scratch share and size (layer 3) 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp <from the readwidth table and docs/analysis/scratch-soundness.md> the same
The hot table (layer 5) 32, 64, 96 MB; replaced or added DECIDED (5 October 2026, 21:36 UTC, delegated): OUT of v3. The rule was g at or above 0.97 on both PC cards in the added form; measured g is 0.87 / 0.85 / 0.84 on the 5090 and 0.84 / 0.81 / 0.80 on the 9070 XT at 32 / 64 / 96 MiB: neither card keeps even the 32 MiB table resident while the dataset streams, and the replaced form helps the on-die-cache chip. Layer 5 stays a measured option for 3.0 docs/plans/hot-table.md (ca2-cache): 5090 MH/s v2 136.1; replaced hot32k4 146.6, hot64k4 140.8, hot96k4 138.5, hot64k2 137.5, hot64k8 163.6; added hot32k4a 118.7, hot64k4a 115.4, hot96k4a 114.4. 9070 XT v2 18.15; replaced 19.79 / 18.73 / 18.33 / 18.17 / 22.32; added 15.27 / 14.62 / 14.56. M5 Max added 0.93 / 0.87 / 0.83. All eight packs bit-exact on the 5090 and the 9070 XT with the Mac's fingerprints (21:29 to 21:35 UTC, no restart straddled)
The era draws (layers 4 and 8) in, if the min-to-max spread across six drawn eras is under 5% per card; item size fixed at 4 B, draws of stride, interleave and the working-set window at or above 256 MiB DECIDED (5 October 2026, 21:58 UTC, delegated): IN. Spread over six eras: RTX 5090 1.3% (136.18 / 136.44 / 138.01 MH/s against v2 137.2), RX 9070 XT 3.2% (18.61 / 18.93 / 19.21 against 18.09), M5 Max 0.8% (28.35 / 28.48 / 28.58 against 27.68); all under 5%; latency-bound share 1.01 / 0.95 / 1.06; every pack's 2^24 fingerprint equal on all three vendors; the CPU verifier 1.00 to 1.02 of v2 within one binary docs/plans/era-layout.md (ca2-era 78c0ee4; PC 1 jobs fetch-ca2-era-20261005 and run-ca2-era-pc1-20261005, 301 s, both cards restored). Chip line: 512 B read per hash, 120 to 128 distinct lines, the mirror is the whole dataset every hour; the interleave's value against a chip with a programmable address decoder is nil (stated), the stride is a bijection with no cryptanalysis yet
The cache schedule (layer 6) flat 256 MiB or a growth schedule DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB docs/analysis/sram-mirror.md: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate
Layer 7 reserved family, unlock by era height or 90% signal DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) docs/analysis/int8-matrix-family.md: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned)
The activation height N4 the rule of section 3 N4 = N5 = 154,800 (the shipper, 22:45 UTC: DAA 136,967 at 0.965 blocks/s puts the publish near 140,200, tip + 14,400 near 154,600, the first multiple of 3,600 at or above is 154,800; one number for both switches); the 10,800 floor holds until DAA 144,000, about 00:45 UTC, re-pinned and rebuilt past that release-0.3.11 23bc2b2 (packaging/mac/packaged-config.sh, the nine-field line)