igneum/docs/plans/counter-asic-2-rollout.md

22 KiB

Generator class v3 on the live devnet: rollout plan (5 October 2026, night)

Scope: the DEVNET only. the project lead delegated the three decisions for the devnet before going to bed (5 October 2026, about 20:10 UTC, through the coordinator): "Counter ASIC 2.0 fully deployed" tonight. The public testnet is not open; its genesis takes v3 from day one. The devnet is ours and a reset is acceptable.

Shape and rules follow docs/plans/finality-v3-rollout-devnet.md and the publish record docs/plans/finality-v3-devnet-publish.md: one height switch read from the override file, every node carries the same object before the height, the PCs get igneumd only through an OTA app version, the activation height leaves at least three hours from the manifest publish. Nothing in this file has run on the devnet. The numbers marked <...> are filled by the integration branch ca2-v3 and the decisions of section 6; the plan is published with them, not before.

1. What changes and what does not

Only the lottery hash's program class changes, and only from the first epoch at or above the height. One switch, program_class_v3_activation_daa, in Params and OverrideParams like difficulty_v2_activation_daa; default u64::MAX (never) on every network. Because one epoch has one program (spec 01 section 1.12), the switch keys on the EPOCH: epoch e is class v3 when 3,600 e >= N4, so the activation is rounded up to an epoch boundary and a block's class is a function of its DAA score alone, as today.

Class v3 = generator version 3: the width rule <W> (layer 1 or 2, decision 1), no scratch (layer 3 decided out: scratch share 0), the era draw of the table layout and the working set (layers 4 and 8), the hot table of <S> MB from the epoch seed (layer 5), the cache growth rule of layer 6 (option C) and the M16 mixer x4 in the dataset item construction (decided 5 October 2026, delegated). New program id (generator = 3 in the id's preimage, spec 1.4.6), new packs and vectors, new IGNEUM_GENERATOR in every pack, a pack of the other version refused by every implementation (spec 1.4.5 already says so).

What does not change: the chain, the genesis, the databases, the day key and the 256 MiB cache fill, the dataset items (spec 1.8.5, if the era interleave keeps the item values; the era-layout document says what it costs otherwise), finality, fees, proving. The SP1 guest does not read the lottery hash (proving/igneum-prove has no dependency on igneum-pow; the pinned guest of DAA 210,000 is a fee-table switch), so no new guest is pinned. The node's EpochSeeds gains the class and the era bytes; IgneumEngine::epoch_for builds the v3 Epoch from them; the miner's export-pack and the serve protocol's job line carry the class so a GPU worker regenerates the right pack from the seed bytes.

The consensus digest (Params::consensus_digest, ledger X18) covers every activation height, so the new field enters the digest and every node must carry the same object before any node reaches the height; a node without the field is refused at the handshake once the others carry it, which is the protection the digest exists for. Binary rollout first (the digest flips when the binary carries the field at never), the height second.

2. The binaries

Built from <node branch> at <commit> on the PCs through tools/build-job.mjs (standing rule 5 October 2026), the Mac binary under the build lock. The table is filled at build time: platform, path, sha256, how it was verified (--version, the switch's first line on a private suffix, strings carries the field name).

3. The activation height N4, and how every node learns it

N4 = DAA at the manifest publish + 10,800 at least, chosen as DAA now + 14,400 rounded up to the next epoch boundary (a multiple of 3,600), checked at publish (N4 - DAA >= 10,800). The packaged line carries every switch:

NODE_OVERRIDE_PARAMS='{"difficulty_v2_activation_daa":33000,"proving_v0_activation_daa":84100,"fees_v1_activation_daa":210000,"finality_v3_activation_daa":135200,"program_class_v3_activation_daa":N4}'

The same object goes verbatim into the override files of Mac node 1, the observer node and the seed, and into the manifest's consensus.override (publish-manifest.sh --override). The era seed for the devnet: the stand-in of docs/plans/era-layout.md (the hash of the last selected-chain block below 15,552,000 n - 7,200; era 0 on the devnet uses the genesis hash), until the 1-hour VDF of spec 4.4 is in the node.

4. The order

  1. The digest flip: a node build that carries program_class_v3_activation_daa at never on every node (hand nodes and the seed first: infra/devnet/restart-hand-nodes.sh '<object without the new field>', then the app version through the manifest; every node prints the same Consensus params digest).
  2. Fix N4, cut the app version (packaging/mac/packaged-config.sh, the three version files), commit as igneum-labs.
  3. The Windows payload inputs (packaging/windows/push-inputs.sh) and the Mac DMG (packaging/mac/build-dmg.sh), the manifest (publish-manifest.sh --activation-height N4 --deadline-note "program class v3" --override '<object>' --deploy), fetch-ci-artifacts.sh --deploy.
  4. The observer, the seed, Mac node 1 with the object; each node's first lines show every switch and Program class v3 from the override file: active from epoch <N4 / 3,600>.
  5. HiveOS: packaging/hive/make-hive-package.sh republished with the v3 igneum-miner and workers, same version string as the apps.
  6. The watch: before N4 - 1,800 both PCs on the new app version (STATUS lines); at the boundary every miner's prepare of the v3 pack (the hot-swap entry's shape) and the first v3 block's program id on the observer; the hash rate per card against the measured v3 numbers of docs/plans/counter-asic-2.md's table; zero pack refused lines; the CPU verifier time per block in the node log against the measured ms per warp.

5. Rollback

Before N4: remove the field on every node and restart; nothing has happened (the digest flips back, so every node at once). After N4: there is no rollback by restart, because blocks mined under v3 verify only under v3. A rollback is a second height switch back to v2 at a later epoch, carried the same way. This is why the measurements of the plan come first.

6. The decisions, by the project lead's rules (devnet)

the project lead's rules, applied by the coordinator and recorded here with the number that decided each:

Decision the project lead's rule Choice The number
Width (layer 1) the widest read that keeps every card we own latency-bound (achieved loads within 90% of the probe ceiling) with margin on the 5090 (its bytes per hash under a third of its bandwidth at the measured rate) DECIDED (5 October 2026, delegated): keep v2, 128 x 4 B. w16 passes the rule (shares 0.90 / 0.84 / 1.03, 18% of the 5090's stream) but closes nothing and does not move the chip row; w64 and w64x4 make the 5090 bandwidth-bound (share 0.58 / 0.56, 37% of stream) docs/plans/read-width.md (readwidth e752fc7): v2 5090 136.1 MH/s, 9070 XT 18.15, M5 Max 27.74 (gap 7.5x); w16 139.8 / 17.90 / 28.26 (gap 7.8x); w64 71.9 / 17.59 / 28.27 (gap 4.1x); the 9070 XT does 2.4 G dependent reads/s at every width
Per-load mix (layer 2) in, if the min-to-max spread across six programs is under 5% per card DECIDED (5 October 2026, delegated): out spreads of the median over six programs: mix 50/35/15 5090 18.8%, 9070 XT 7.4%, M5 Max 11.3%; mix 25/50/25 22.3% / 5.5% / 8.1%
Scratch share (layer 3) the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap DECIDED (5 October 2026, delegated): 0. No share under the cap moves the on-die-cache recompute chip, so layer 3 is not adopted into v3 docs/analysis/scratch-soundness.md (ca2-soundness a465881): chip 333 MH/s against the 5090's measured 139.7 = 2.4x at 0% RMW; 2.4x at 12.5 / 25 / 50% replaced (chip 381 / 443 / 661 against 160 / 186 / 279 projected) and 2.4x or more added, at 32 and 128 KB; the chip keeps the scratch implicitly in 80 to 320 B per lane because the verifier resets it per unit
Activation height N4 devnet tip + 14,400 at publish, checked >= 10,800 at publish, rounded up to the epoch boundary <at publish>

6a. The chip model's headline, and how the scratch share is chosen

The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; docs/analysis/sram-mirror.md after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (docs/analysis/scratch-soundness.md, question 2).

6b. The user tiers (the consequences rule)

AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second at 1 GiB, measured tonight), 2.2x worse per pound at list prices (0.032 against 0.072 MH/s per pound, approximate) and 4.9x worse per watt (read-width.md section 4.1). This is the card's memory system, not a tuning gap: no read width closes it without making the 5090 bandwidth-bound. The level 3 numbers page states it.

7. Gates before any publish (all of them, no exceptions)

# Gate Evidence required State
G1 bit-exact v3 on all three vendors against the Mac reference vectors PASS on Metal, CUDA (5090), AMD OpenCL for the v3 packs; batch fingerprints equal. RULING (coordinator, 5 October 2026, about 21:10 UTC): the integrated gfx1036 (RDNA 2, AMD OpenCL 3683.0) satisfies the AMD vendor tonight, because G1 is a compiler-and-ISA property and gfx1036 carried the v1 and v2 conformance; the 9070 XT's hash-rate and power rows are owed and taken when its link is back open
G2 the CPU verifier exact on 1,000 random hashes per card 1,000 GPU hashes per card re-hashed by igneum-pow on the Mac, 0 mismatches open
G3 the generator soundness suite green, the new scratch tests included cargo test in igneum-pow, tests/packs.rs, the Metal fuzz, edge, stats, determinism runs on the v3 class open
G4 the fast-time 3-node network mining across a v3 activation 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary open
G5 the PC-built Windows workers and the Mac workers from the same commit sha256 of each worker and the commit in the bench log open
G6 the node change on a fork branch from the 0.3.10 tip 21d4c73c (release-0.3.6 once the shipper lands it) with suites green on PC 2 the build job id and its SUMMARY line open

If any gate fails: stop at that gate, write why in docs/plans/counter-asic-2-status.md, do not publish.

8. The release

0.3.11 through the shipper's pipeline (tools/ship-app.mjs, the plan shape of docs/plans/release-0.3.10.md). The 0.3.10 shipper finishes 0.3.10 first, then takes the 0.3.11 tree, or the coordinator ships 0.3.11 with the same runbook if the shipper has stopped. The publish order is that of docs/plans/finality-v3-devnet-publish.md: the override field, the digest handshake, hand nodes and the seed first, then the manifest with --activation-height and the deadline note, then the apps, then the digest sweep, then the HiveOS package republish.

6c. Card lifetime: cache residency and the growth mapping (from docs/analysis/card-lifetime-2026-10-05.md, merged)

Decision Choice The number
Cache residency on the GPU DECIDED (5 October 2026, delegated): the cache is FREED after the daily dataset build; the hash reads the dataset and the hot table only, never the cache. hot-table.md's resident reading is corrected to this The daily rebuild is the only cost: cache fill 0.67 ms and dataset build 13.4 ms on the RTX 5090 at x1 (bench-log, 3 October 2026), 2 ms and 13 to 30 ms on the M5 Max; at mixer x4 the build is about 54 ms (measurement owed on ca2-mixer). Freeing it moves the 12 GB tier from year 12 to year 28 under mapping (b)
Dataset growth mapping (a: continuous 2 + 0.5 GiB a year with a multiply-shift index; b: power-of-two steps at years 4, 12, 28, 60 with AND MASK) RECOMMENDED for the project lead: (b). It keeps AND MASK and every vector's size, it is what option C's "doubles when the dataset doubles" already assumes, and it is the only mapping under which the litepaper's "4 GB about four years, 8 GB more than a decade" is true (under (a) a 4 GB card is out within 1 to 1.5 years, an 8 GB card at 6 to 7.5 years) card-lifetime table 2: 4 GB out at year 4 (b) or 1.0 to 1.5 (a); 8 GB year 12 or 6.3 to 7.5; 12 GB year 28 freed or 12 resident; 24 GB year 60 freed

Public lines to fix on the integration branch (card-lifetime table 3): site/index.html "2 GB, growing" gains the rate ("2 GB at genesis, doubling at years 4, 12 and 28"); "Any 4 GB card" becomes "any 4 GB card at launch, 8 GB from year 4"; the litepaper's "4 GB about four years, 8 GB more than a decade" stays with mapping (b) and gains "under the step schedule"; docs/evidence.md gains a row for the card-lifetime claim labelled designed. hot-table.md line 73's 8 GB row is corrected (it counted a 5090's 8,160 warps; a real 8 GB card has 20 to 24 SMs).

7b. Hardware events (for the morning summary)

When (UTC) Event What the app did For the project lead
5 October, at install (earlier today) the RX 9070 XT in the Sonnet Breakaway Box 850T5 over USB4 went Code 43 came back after a driver reinstall and a reboot
5 October, about 20:40 the 9070 XT dropped off PC 1's bus: Get-PnpDevice -Class Display lists only the integrated AMD Radeon Graphics (gfx1036) and the RTX 5090; after pnputil /scan-devices at 20:45:34Z the card is still absent and the USB4 list shows only the host and root routers: the "USB4 Router (2.0), Sonnet Technologies Breakaway Box 850T5" present at 17:18Z is gone, so the box itself is off the link the AMD worker (igneum-worker-opencl --device 1) mines the gfx1036 at 3.12 MH/s; the 5090 keeps mining; nobody was woken, PC 1's app was not restarted; a 10-second rescan probe (pnputil /scan-devices, the USB4 router status) was granted the second eGPU link fault today: reseat the USB4 cable and the eGPU's power; the 0.3.10 hot-plug code shows the card as "removed" and picks it up again without a restart

7a. A dated constraint from the consequences review (C1)

The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H.

8a. Proving v1 rides with it

the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2): fork branch proving-v1 (from a24ab01a: segment records, the pool split, five RPCs, p2p message 72, parameters proving_v1_activation_daa, segment_blocks, unproven_daa, aggregator_share_bps, in the digest only once set) and its app branch, rebased onto the 0.3.10 tips. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which.

Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = publish-public.sh --hive adds a platforms.linux entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; the rig miners' --exit-on-seed-change path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths.

9. After the publish: Counter ASIC 3.0

The ASIC-history agent sends its ranked additions; they are measured the same way, folded into the class as v4 behind its own activation, same gates, same rollout; docs/plans/counter-asic-3.md.

10. Decisions table (superseded by section 6; kept for the options)

Decision Options Recommendation Source
The width rule (layers 1 and 2) 4 B fixed; 16 B; 64 B; the per-program mix <the readwidth table's decision rule: the widest read that keeps every card latency-bound with margin on the 5090> docs/plans/read-width.md
The scratch share and size (layer 3) 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp <from the readwidth table and docs/analysis/scratch-soundness.md> the same
The hot table (layer 5) 32, 64, 96 MB; replaced or added In the ADDED form only (16 dataset loads plus k hot loads): the replaced form lets the on-die-cache chip skip item derivations (k = 4: chip x1.33 against the GPU's measured x1.05 to x1.22). Size = the largest table resident on every card we own, pending the PC rows; Mac rows (replaced form, M5 Max): hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k8 x1.71; Apple OpenCL dependent-read probe 32 / 64 / 96 / 1024 MiB: 21.7 / 12.8 / 12.3 / 3.50 G loads/s docs/plans/hot-table.md (ca2-cache 65bc7a7)
The era draws (layers 4 and 8) in, if the min-to-max spread across six drawn eras is under 5% per card; item size fixed at 4 B, draws of stride, interleave and the working-set window at or above 256 MiB pending the six-era table docs/plans/era-layout.md
The cache schedule (layer 6) flat 256 MiB or a growth schedule DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB docs/analysis/sram-mirror.md: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate
Layer 7 reserved family, unlock by era height or 90% signal DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) docs/analysis/int8-matrix-family.md: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned)
The activation height N4 the rule of section 3