From 578a903b36a6574e5ba3ab0f85af02bbec3e48c3 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 19:50:36 +0000 Subject: [PATCH 001/131] Counter ASIC 2.0: coordinator status file, first entry --- docs/plans/counter-asic-2-status.md | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) create mode 100644 docs/plans/counter-asic-2-status.md diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md new file mode 100644 index 000000000..00b43e9ea --- /dev/null +++ b/docs/plans/counter-asic-2-status.md @@ -0,0 +1,17 @@ +# Counter ASIC 2.0: status + +Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rewritten every 45 minutes while the work runs. Times UTC, 5 October 2026 (night). Base for every ca2 branch: `readwidth` at 019b014 (the LoadClass flag, fold_words, the scratch op, the three emitters, 20 packs). + +## 20:05 first status + +| Layer | Branch | Agent | State | Numbers so far | Blockers | +|---|---|---|---|---|---| +| 1, 2, 3 (widths, mix, scratch) | readwidth 019b014 | a451c9935bfb1bc19 (not ours) | packs and verifier done; Mac Metal and Apple OpenCL runs next, then one job per PC | CPU verifier, M5 Max one core, avg of 50, ms per warp: v2 0.604, w16 0.610, w64 0.630, w64x4 0.160, mixA 0.620, mixB 0.614, scr0 0.600, scr2 0.534, scr4 0.457, scr8 0.314 | holds the Mac measure lock and both PCs first | +| 4 + 8 (era layout, working set) | ca2-era | spawning | design | none | PC time behind readwidth | +| 5 (cache-sized second table) | ca2-cache | spawning | design | none | PC time behind readwidth | +| 6 (SRAM schedule) | ca2-analysis | spawning | cited analysis | none | none | +| 7 (integer matrix family) | ca2-analysis | spawning | design; dp4a throughput owed unless PC time frees | none | Apple int8 path to check against Metal docs | +| 3 soundness | ca2-soundness | spawning | analysis + tests | none | none | +| Integration v3 | ca2-v3 | after the above | waiting | none | the readwidth table and the four branches | + +Decisions for the project lead so far: none. Nothing here touches consensus or any live node. From 68e9c9fd2970a09eee57233774d82e319d53611e Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 19:53:57 +0000 Subject: [PATCH 002/131] Counter ASIC 2.0 status: restart after the credit cut-off, the 6 GB working-set budget --- docs/plans/counter-asic-2-status.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 00b43e9ea..74a6bc4c4 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -2,7 +2,7 @@ Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rewritten every 45 minutes while the work runs. Times UTC, 5 October 2026 (night). Base for every ca2 branch: `readwidth` at 019b014 (the LoadClass flag, fold_words, the scratch op, the three emitters, 20 packs). -## 20:05 first status +## 19:55 first status (the 19:50 start was cut off by exhausted credits at about 19:58 before any sub-agent work landed; respawned at 19:55 on the restart) | Layer | Branch | Agent | State | Numbers so far | Blockers | |---|---|---|---|---|---| @@ -14,4 +14,6 @@ Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rew | 3 soundness | ca2-soundness | spawning | analysis + tests | none | none | | Integration v3 | ca2-v3 | after the above | waiting | none | the readwidth table and the four branches | +Budget rule received from the coordinator at the restart: the per-warp scratch is capped so the whole working set (1 GiB table + hot table + scratch for every resident warp + buffers) stays under 6 GB on an 8 GB card, which puts the scratch in the tens of KB per warp; the layer 5 hot table shares that budget. + Decisions for the project lead so far: none. Nothing here touches consensus or any live node. From 14303e9bf883071fcf8e65f210789cf4cb0e3e63 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 19:56:18 +0000 Subject: [PATCH 003/131] Counter ASIC 2.0 status: agents respawned, first readwidth Mac rows --- docs/plans/counter-asic-2-status.md | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 74a6bc4c4..6ccc74b14 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -6,12 +6,12 @@ Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rew | Layer | Branch | Agent | State | Numbers so far | Blockers | |---|---|---|---|---|---| -| 1, 2, 3 (widths, mix, scratch) | readwidth 019b014 | a451c9935bfb1bc19 (not ours) | packs and verifier done; Mac Metal and Apple OpenCL runs next, then one job per PC | CPU verifier, M5 Max one core, avg of 50, ms per warp: v2 0.604, w16 0.610, w64 0.630, w64x4 0.160, mixA 0.620, mixB 0.614, scr0 0.600, scr2 0.534, scr4 0.457, scr8 0.314 | holds the Mac measure lock and both PCs first | -| 4 + 8 (era layout, working set) | ca2-era | spawning | design | none | PC time behind readwidth | -| 5 (cache-sized second table) | ca2-cache | spawning | design | none | PC time behind readwidth | -| 6 (SRAM schedule) | ca2-analysis | spawning | cited analysis | none | none | -| 7 (integer matrix family) | ca2-analysis | spawning | design; dp4a throughput owed unless PC time frees | none | Apple int8 path to check against Metal docs | -| 3 soundness | ca2-soundness | spawning | analysis + tests | none | none | +| 1, 2, 3 (widths, mix, scratch) | readwidth 019b014 | a451c9935bfb1bc19 (not ours) | Mac Metal rows in; Apple OpenCL next, then one job per PC; scratch re-run at 32 and 128 KB per warp | CPU verifier, M5 Max one core, avg of 50, ms per warp: v2 0.604, w16 0.610, w64 0.630, w64x4 0.160, mixA 0.620, mixB 0.614, scr0 0.600, scr2 0.534, scr4 0.457, scr8 0.314. Metal M5 Max hash rate (5 x 2^24, GPU time, 3 of 3 vectors on every pack): v2 27.7 MH/s, w16 28.3, w64 28.2, w64x4 109.7 (32 loads of 64 B), mix 50/35/15 over 6 programs 25.4 to 28.4 (median 26.7), mix 25/50/25 over 6 programs 22.0 to 25.2 (median 24.4). Scratch at 1 MiB per warp, 1,024 warps: 23.1 / 18.2 / 16.3 / 14.9 MH/s at 0 / 12.5 / 25 / 50% RMW, superseded by the cap | holds the Mac measure lock and both PCs first | +| 4 + 8 (era layout, working set) | ca2-era | a452664c512c73b9b | design, respawned 19:56 | none | PC time behind readwidth | +| 5 (cache-sized second table) | ca2-cache | a5271cf269757b118 | design, respawned 19:57 | none | PC time behind readwidth | +| 6 (SRAM schedule) | ca2-analysis | a5c6bc2dfcc4613ef (respawned 19:58) | cited analysis | none | none | +| 7 (integer matrix family) | ca2-analysis | a5c6bc2dfcc4613ef | design; dp4a throughput owed unless PC time frees | none | Apple int8 path to check against Metal docs | +| 3 soundness | ca2-soundness | a548aadeefd1ab3b2 (spawned 19:59; sizes 32 and 128 KB per warp) | analysis + tests | none | none | | Integration v3 | ca2-v3 | after the above | waiting | none | the readwidth table and the four branches | Budget rule received from the coordinator at the restart: the per-warp scratch is capped so the whole working set (1 GiB table + hot table + scratch for every resident warp + buffers) stays under 6 GB on an 8 GB card, which puts the scratch in the tens of KB per warp; the layer 5 hot table shares that budget. From 8988f9e03c7b98d7d2d885532eeb47c1c895d08c Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 19:57:30 +0000 Subject: [PATCH 004/131] Counter ASIC 2.0: rollout plan draft (program_class_v3 height switch, digest first, decisions table) --- docs/plans/counter-asic-2-rollout.md | 52 ++++++++++++++++++++++++++++ 1 file changed, 52 insertions(+) create mode 100644 docs/plans/counter-asic-2-rollout.md diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md new file mode 100644 index 000000000..827aedab8 --- /dev/null +++ b/docs/plans/counter-asic-2-rollout.md @@ -0,0 +1,52 @@ +# Generator class v3 on the live devnet: rollout plan (draft, 5 October 2026, night) + +Shape and rules follow `docs/plans/finality-v3-rollout-devnet.md` and the publish record `docs/plans/finality-v3-devnet-publish.md`: one height switch read from the override file, every node carries the same object before the height, the PCs get igneumd only through an OTA app version, the activation height leaves at least three hours from the manifest publish. Nothing in this file has run on the devnet. The numbers marked `<...>` are filled by the integration branch `ca2-v3` and the decisions of section 6; the plan is published with them, not before. + +## 1. What changes and what does not + +Only the lottery hash's program class changes, and only from the first epoch at or above the height. One switch, `program_class_v3_activation_daa`, in `Params` and `OverrideParams` like `difficulty_v2_activation_daa`; default `u64::MAX` (never) on every network. Because one epoch has one program (spec 01 section 1.12), the switch keys on the EPOCH: epoch `e` is class v3 when `3,600 e >= N4`, so the activation is rounded up to an epoch boundary and a block's class is a function of its DAA score alone, as today. + +Class v3 = generator version 3: the width rule `` (layer 1 or 2, decision 1), the per-warp scratch at `` RMW and `` per warp (layer 3, under the 6 GB working-set cap), the era draw of the table layout and the working set (layers 4 and 8), the hot table of `` MB from the epoch seed (layer 5). New program id (`generator = 3` in the id's preimage, spec 1.4.6), new packs and vectors, new `IGNEUM_GENERATOR` in every pack, a pack of the other version refused by every implementation (spec 1.4.5 already says so). + +What does not change: the chain, the genesis, the databases, the day key and the 256 MiB cache fill, the dataset items (spec 1.8.5, if the era interleave keeps the item values; the era-layout document says what it costs otherwise), finality, fees, proving. The SP1 guest does not read the lottery hash (`proving/igneum-prove` has no dependency on `igneum-pow`; the pinned guest of DAA 210,000 is a fee-table switch), so no new guest is pinned. The node's `EpochSeeds` gains the class and the era bytes; `IgneumEngine::epoch_for` builds the v3 `Epoch` from them; the miner's `export-pack` and the serve protocol's job line carry the class so a GPU worker regenerates the right pack from the seed bytes. + +The consensus digest (`Params::consensus_digest`, ledger X18) covers every activation height, so the new field enters the digest and every node must carry the same object before any node reaches the height; a node without the field is refused at the handshake once the others carry it, which is the protection the digest exists for. Binary rollout first (the digest flips when the binary carries the field at `never`), the height second. + +## 2. The binaries + +Built from `` at `` on the PCs through `tools/build-job.mjs` (standing rule 5 October 2026), the Mac binary under the build lock. The table is filled at build time: platform, path, sha256, how it was verified (`--version`, the switch's first line on a private suffix, `strings` carries the field name). + +## 3. The activation height N4, and how every node learns it + +`N4 = DAA at the manifest publish + 10,800` at least, chosen as `DAA now + 14,400` rounded up to the next epoch boundary (a multiple of 3,600), checked at publish (`N4 - DAA >= 10,800`). The packaged line carries every switch: + +``` +NODE_OVERRIDE_PARAMS='{"difficulty_v2_activation_daa":33000,"proving_v0_activation_daa":84100,"fees_v1_activation_daa":210000,"finality_v3_activation_daa":135200,"program_class_v3_activation_daa":N4}' +``` + +The same object goes verbatim into the override files of Mac node 1, the observer node and the seed, and into the manifest's `consensus.override` (`publish-manifest.sh --override`). The era seed for the devnet: the stand-in of `docs/plans/era-layout.md` (the hash of the last selected-chain block below `15,552,000 n - 7,200`; era 0 on the devnet uses the genesis hash), until the 1-hour VDF of spec 4.4 is in the node. + +## 4. The order + +1. The digest flip: a node build that carries `program_class_v3_activation_daa` at never on every node (hand nodes and the seed first: `infra/devnet/restart-hand-nodes.sh ''`, then the app version through the manifest; every node prints the same `Consensus params digest`). +2. Fix N4, cut the app version (`packaging/mac/packaged-config.sh`, the three version files), commit as igneum-labs. +3. The Windows payload inputs (`packaging/windows/push-inputs.sh`) and the Mac DMG (`packaging/mac/build-dmg.sh`), the manifest (`publish-manifest.sh --activation-height N4 --deadline-note "program class v3" --override '' --deploy`), `fetch-ci-artifacts.sh --deploy`. +4. The observer, the seed, Mac node 1 with the object; each node's first lines show every switch and `Program class v3 from the override file: active from epoch `. +5. HiveOS: `packaging/hive/make-hive-package.sh` republished with the v3 `igneum-miner` and workers, same version string as the apps. +6. The watch: before `N4 - 1,800` both PCs on the new app version (STATUS lines); at the boundary every miner's `prepare` of the v3 pack (the hot-swap entry's shape) and the first v3 block's program id on the observer; the hash rate per card against the measured v3 numbers of `docs/plans/counter-asic-2.md`'s table; zero `pack refused` lines; the CPU verifier time per block in the node log against the measured ms per warp. + +## 5. Rollback + +Before N4: remove the field on every node and restart; nothing has happened (the digest flips back, so every node at once). After N4: there is no rollback by restart, because blocks mined under v3 verify only under v3. A rollback is a second height switch back to v2 at a later epoch, carried the same way. This is why the measurements of the plan come first. + +## 6. What the project lead decides + +| Decision | Options | Recommendation | Source | +|---|---|---|---| +| The width rule (layers 1 and 2) | 4 B fixed; 16 B; 64 B; the per-program mix | `` | `docs/plans/read-width.md` | +| The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `` | the same | +| The hot table size (layer 5) | 32, 64, 96 MB | `` | | +| The era draw bounds (layers 4 and 8) | the draw as written in `docs/plans/era-layout.md` | | | +| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | `` | | +| Layer 7 | reserved family, unlock by era height or 90% signal | `` | | +| The activation height N4 | the rule of section 3 | | | From b3b935f69f3c032ac5f24223273661f4ad18d9f5 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:04:04 +0000 Subject: [PATCH 005/131] Counter ASIC 2.0 status 20:05: readwidth PC jobs, the capped-scratch finding --- docs/plans/counter-asic-2-status.md | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 6ccc74b14..fdf4bce49 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -17,3 +17,20 @@ Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rew Budget rule received from the coordinator at the restart: the per-warp scratch is capped so the whole working set (1 GiB table + hot table + scratch for every resident warp + buffers) stays under 6 GB on an 8 GB card, which puts the scratch in the tens of KB per warp; the layer 5 hot table shares that budget. Decisions for the project lead so far: none. Nothing here touches consensus or any live node. + +## 20:05 readwidth PC jobs out; the scratch finding + +Base moved: every ca2 branch rebases onto readwidth b970dda (scratch per warp is a class parameter, 32 or 128 KiB; the SCRATCH_* constants are gone; Metal pack harness `proto-metal/packbench.swift`; OpenCL `--bench-pack`). The coordinator branch is rebased; the four agents were told. + +Readwidth PC jobs published 20:02:51Z: `fetch-readwidth-20261005` (both PCs), `run-readwidth-5090-20261005` on PC 2 (6 to 10 min once it starts; a build job from another session is queued ahead of it), `run-readwidth-9070-20261005` on PC 1 (10 to 15 min; only the gfx1201 card is switched off). My layer 4, 5 and 7 PC jobs queue behind these. + +Finding that bears on layer 3 (readwidth, Metal, M5 Max, capped sizes): the scratch read-modify-writes are cheaper than the dataset loads they replace, so the rate RISES with the RMW share. + +| Class | MH/s | +|---|---| +| v2 | 27.7 | +| scr0k32 (control) | 28.1 | +| 32 KB per warp, 12.5 / 25 / 50% RMW | 25.4-26.1 / 29.4-31.7 / 44.4-49.1 | +| 128 KB per warp, 12.5 / 25 / 50% RMW | 24.3-26.2 / 26.4-28.1 / 34.1-35.4 | + +Reading (readwidth agent): 4,096 warps x 32 KB = 128 MB sits in the chip's caches. Consequence for the decision: a scratch that fits a GPU's cache fits a chip's SRAM at the same size, so at the capped size the writes cost everyone a cache-bound op in place of a latency-bound load. Passed to the soundness agent: what size would make the writes cost DRAM latency, whether that fits the 6 GB cap, and whether the RMW share should be added to the 16 dataset loads rather than taken from them. From 34ba9cd92e7d91f22d6e98605eb3868afe89d57b Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:07:26 +0000 Subject: [PATCH 006/131] Counter ASIC 2.0 rollout: the project lead's devnet decision rules, the six gates, the 0.3.11 release path --- docs/plans/counter-asic-2-rollout.md | 38 ++++++++++++++++++++++++++-- 1 file changed, 36 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 827aedab8..0e37bf481 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -1,4 +1,6 @@ -# Generator class v3 on the live devnet: rollout plan (draft, 5 October 2026, night) +# Generator class v3 on the live devnet: rollout plan (5 October 2026, night) + +Scope: the DEVNET only. the project lead delegated the three decisions for the devnet before going to bed (5 October 2026, about 20:10 UTC, through the coordinator): "Counter ASIC 2.0 fully deployed" tonight. The public testnet is not open; its genesis takes v3 from day one. The devnet is ours and a reset is acceptable. Shape and rules follow `docs/plans/finality-v3-rollout-devnet.md` and the publish record `docs/plans/finality-v3-devnet-publish.md`: one height switch read from the override file, every node carries the same object before the height, the PCs get igneumd only through an OTA app version, the activation height leaves at least three hours from the manifest publish. Nothing in this file has run on the devnet. The numbers marked `<...>` are filled by the integration branch `ca2-v3` and the decisions of section 6; the plan is published with them, not before. @@ -39,7 +41,39 @@ The same object goes verbatim into the override files of Mac node 1, the observe Before N4: remove the field on every node and restart; nothing has happened (the digest flips back, so every node at once). After N4: there is no rollback by restart, because blocks mined under v3 verify only under v3. A rollback is a second height switch back to v2 at a later epoch, carried the same way. This is why the measurements of the plan come first. -## 6. What the project lead decides +## 6. The decisions, by the project lead's rules (devnet) + +the project lead's rules, applied by the coordinator and recorded here with the number that decided each: + +| Decision | the project lead's rule | Choice | The number | +|---|---|---|---| +| Width (layers 1 and 2) | the widest read that keeps every card we own latency-bound (achieved loads within 90% of the probe ceiling) with margin on the 5090 (its bytes per hash under a third of its bandwidth at the measured rate) | `` | | +| Per-load mix (layer 2) | in, if the min-to-max spread across six programs is under 5% per card | `` | | +| Scratch share (layer 3) | the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap | `` | | +| Activation height N4 | devnet tip + 14,400 at publish, checked >= 10,800 at publish, rounded up to the epoch boundary | `` | | + +## 7. Gates before any publish (all of them, no exceptions) + +| # | Gate | Evidence required | State | +|---|---|---|---| +| G1 | bit-exact v3 on all three cards against the Mac reference | vectors PASS on Metal, CUDA (5090), OpenCL (9070 XT) for the v3 packs; batch fingerprints equal | open | +| G2 | the CPU verifier exact on 1,000 random hashes per card | 1,000 GPU hashes per card re-hashed by `igneum-pow` on the Mac, 0 mismatches | open | +| G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | +| G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary | open | +| G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | +| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c (release-0.3.6 once the shipper lands it) with suites green on PC 2 | the build job id and its SUMMARY line | open | + +If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. + +## 8. The release + +0.3.11 through the shipper's pipeline (`tools/ship-app.mjs`, the plan shape of `docs/plans/release-0.3.10.md`). The 0.3.10 shipper finishes 0.3.10 first, then takes the 0.3.11 tree, or the coordinator ships 0.3.11 with the same runbook if the shipper has stopped. The publish order is that of `docs/plans/finality-v3-devnet-publish.md`: the override field, the digest handshake, hand nodes and the seed first, then the manifest with `--activation-height` and the deadline note, then the apps, then the digest sweep, then the HiveOS package republish. + +## 9. After the publish: Counter ASIC 3.0 + +The ASIC-history agent sends its ranked additions; they are measured the same way, folded into the class as v4 behind its own activation, same gates, same rollout; `docs/plans/counter-asic-3.md`. + +## 10. Decisions table (superseded by section 6; kept for the options) | Decision | Options | Recommendation | Source | |---|---|---|---| From 4b11f3a0cba48eab735e96afc8ba734f7f8389e9 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:08:02 +0000 Subject: [PATCH 007/131] Counter ASIC 2.0 status 20:15: the devnet mandate, gates, node path note --- docs/plans/counter-asic-2-status.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index fdf4bce49..2c137c91f 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -34,3 +34,11 @@ Finding that bears on layer 3 (readwidth, Metal, M5 Max, capped sizes): the scra | 128 KB per warp, 12.5 / 25 / 50% RMW | 24.3-26.2 / 26.4-28.1 / 34.1-35.4 | Reading (readwidth agent): 4,096 warps x 32 KB = 128 MB sits in the chip's caches. Consequence for the decision: a scratch that fits a GPU's cache fits a chip's SRAM at the same size, so at the capped size the writes cost everyone a cache-bound op in place of a latency-bound load. Passed to the soundness agent: what size would make the writes cost DRAM latency, whether that fits the 6 GB cap, and whether the RMW share should be added to the 16 dataset loads rather than taken from them. + +## 20:15 mandate: v3 on the devnet tonight, by the project lead's rules + +the project lead has gone to bed and delegated the three decisions for the DEVNET only (not the public testnet): width, per-load mix and scratch share by the rules now written in `docs/plans/counter-asic-2-rollout.md` section 6, the activation height = devnet tip + 14,400 at publish (checked >= 10,800), published the way finality v3 was. Six gates before any publish (rollout section 7): bit-exact v3 on the three cards; the CPU verifier exact on 1,000 random GPU hashes per card; the soundness suite green with the new scratch tests; the fast-time 3-node network mining across a v3 activation with 0 rejected blocks and 0 forks; Windows and Mac workers from one commit; the node change on a fork from the 0.3.10 tip (21d4c73c) with suites green on PC 2. Release 0.3.11 through the shipper's pipeline; the 0.3.10 shipper (ae892a8b0f78fe31c) has been asked for its state and the handoff. If a gate fails: stop, write why here, do not publish. + +Node build note for the integration: the node links `igneum-pow` by path (`../../../../igneum-pow` from `consensus/pow` and `igneum/miner`), so the node fork worktree for v3 must live under the v3 worktree's `vendor/` so that the path resolves to the v3 crate, not master's. + +After the publish: Counter ASIC 3.0 from the ASIC-history agent's ranked additions (a202a09dcd24ba1d3), as class v4 behind its own activation, same gates, `docs/plans/counter-asic-3.md`. From 045e8e5708f513e4e2575d50815dae44e025548a Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:08:32 +0000 Subject: [PATCH 008/131] Counter ASIC 2.0 status 20:25: 0.3.10 blocked on GitHub, node fork convention, 0.3.11 inputs --- docs/plans/counter-asic-2-status.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 2c137c91f..e456edc6e 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -42,3 +42,11 @@ the project lead has gone to bed and delegated the three decisions for the DEVNE Node build note for the integration: the node links `igneum-pow` by path (`../../../../igneum-pow` from `consensus/pow` and `igneum/miner`), so the node fork worktree for v3 must live under the v3 worktree's `vendor/` so that the path resolves to the v3 crate, not master's. After the publish: Counter ASIC 3.0 from the ASIC-history agent's ranked additions (a202a09dcd24ba1d3), as class v4 behind its own activation, same gates, `docs/plans/counter-asic-3.md`. + +## 20:25 the shipper's answer, the node fork convention + +0.3.10 (shipper ae892a8b0f78fe31c): staged and blocked on GitHub's Actions incident (run 37365130137 queued since 19:42:43Z under a re-dispatching watcher); nothing on the network has moved, the live manifest is still 0.3.9. Once CI is green: ship (5 min), update-now (apps restart 1 to 10 min later), hand nodes and seed (5 min), digest sweep; 0.3.10 finished about 30 min after green. The app version per machine on the console's cards is the restart signal for re-running any straddling measurement. HiveOS is published by the ship's --public step, nothing separate. + +Node fork for v3: base on COMMIT 21d4c73c (release-0.3.10 in vendor/igneum-node = release-0.3.6 a24ab01a + housekeeping 4fb32865 + tx-gossip e242acd0 + c4-fix e18f1e0e). Fork-side pack-loop 05ef0fa3 (the miner's pack check, exit 44) is not in it and touches the same miner paths as v3 (export-pack, the job line): merge it. Because the node links igneum-pow by relative path, the v3 node worktree goes under the v3 worktree: `git -C /Users/joshm/Projects/igneum/vendor/igneum-node worktree add /Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2 -b ca2-v3-node 21d4c73c`. + +0.3.11 shipper: the coordinator assigns it (the 0.3.10 shipper stops at its report). Inputs the ship needs: the fork commit with its PC 2 suite results recorded, the main tip, the override object with every switch (the four live fields plus program_class_v3_activation_daa), the expected digest read on a 22-s scratch node, the deadline note ("program class v3"), the activation height, the one-line changelog, and whether the pinned proving guest changes (it does not: the prover has no igneum-pow dependency; confirmed by grep of proving/igneum-prove Cargo files). From 1509a2c477313ba039a1c3999bfd972a961d7055 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:10:32 +0000 Subject: [PATCH 009/131] Counter ASIC 2.0 status 20:35: readwidth round 2, probe ceilings, width arithmetic --- docs/plans/counter-asic-2-status.md | 13 +++++++++++++ 1 file changed, 13 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index e456edc6e..2867c71e7 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -50,3 +50,16 @@ After the publish: Counter ASIC 3.0 from the ASIC-history agent's ranked additio Node fork for v3: base on COMMIT 21d4c73c (release-0.3.10 in vendor/igneum-node = release-0.3.6 a24ab01a + housekeeping 4fb32865 + tx-gossip e242acd0 + c4-fix e18f1e0e). Fork-side pack-loop 05ef0fa3 (the miner's pack check, exit 44) is not in it and touches the same miner paths as v3 (export-pack, the job line): merge it. Because the node links igneum-pow by relative path, the v3 node worktree goes under the v3 worktree: `git -C /Users/joshm/Projects/igneum/vendor/igneum-node worktree add /Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2 -b ca2-v3-node 21d4c73c`. 0.3.11 shipper: the coordinator assigns it (the 0.3.10 shipper stops at its report). Inputs the ship needs: the fork commit with its PC 2 suite results recorded, the main tip, the override object with every switch (the four live fields plus program_class_v3_activation_daa), the expected digest read on a 22-s scratch node, the deadline note ("program class v3"), the activation height, the one-line changelog, and whether the pinned proving guest changes (it does not: the prover has no igneum-pow dependency; confirmed by grep of proving/igneum-prove Cargo files). + +## 20:35 readwidth round 2 on the PCs; the width arithmetic under the project lead's rule + +Round 1 of the readwidth PC jobs refused every pack (the workers demand a 32-byte chain seed; the experiment packs carried string seeds); fixed at readwidth 1ea7a52 (packfile.h), republished 20:09:19Z as `run-readwidth-5090-20261005c` (PC 2, about 6 min) and `run-readwidth-9070-20261005c` (PC 1, about 10 min). The era and cache agents were told to rebase onto 1ea7a52 and to prove their packs load on the Mac OpenCL host before any PC job. Every ca2 branch now bases on 1ea7a52. + +Probe ceilings from round 1 (dependent reads per second at 1024 MiB, device time): + +| Card | 4 B | 16 B | 64 B | Stream | +|---|---|---|---|---| +| RTX 5090 | 17.5 G | 18.0 G | 9.1 G (584 GB/s) | 1,579 GB/s | +| RX 9070 XT | 2.42 G | 2.43 G | 2.47 G (158 GB/s) | 636 GB/s | + +the project lead's width rule applied to the ceilings alone (the measured v3 rates will replace this when the table lands): a 128-load hash at 64 B reads 8,192 B; at the 5090's 64 B ceiling that is 71 MH/s and 584 GB/s, 37% of its stream bandwidth, over the one-third margin the rule sets; at 16 B it is 2,048 B per hash, 141 MH/s and 288 GB/s, 18%, inside the margin; 4 B is 9%. On the 9070 XT every width costs the same line fetch (2.4 G/s), so 16 B is where AMD gains 4x the bytes per hash at no cost and the 5090 stays latency-bound with margin. Provisional width under the rule: 16 B (w16), pending the measured rates and the latency-bound share per card. From 7262df1964e5fee88c34a774396ae1646d364c94 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:11:13 +0000 Subject: [PATCH 010/131] Counter ASIC: the public description in four levels (levels 1 and 2 as copy, 3 and 4 with owed markers) --- docs/plans/counter-asic-2-public.md | 39 +++++++++++++++++++++++++++++ 1 file changed, 39 insertions(+) create mode 100644 docs/plans/counter-asic-2-public.md diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md new file mode 100644 index 000000000..63be2fbca --- /dev/null +++ b/docs/plans/counter-asic-2-public.md @@ -0,0 +1,39 @@ +# Counter ASIC: the public description in four levels + +the project lead, 5 October 2026 (night): "not an information overload". Four levels; the layer names appear only from level 3 down, next to their numbers. Numbers come from the final table of `docs/plans/counter-asic-2-status.md`; a number still owed is marked `[owed: ...]`, never guessed. the project lead's copy law throughout. + +## Level 1: one sentence (site hero, litepaper abstract) + +Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public. + +## Level 2: one site card, one short litepaper section + +Three ideas, no layer names, no widths, no SRAM. + +**The hash rewrites itself.** A new program every hour, drawn from the chain. Its memory pattern and its read widths change with it. The rules change on a schedule fixed at launch. No release, no vote. + +**It waits on memory, not maths.** Every hash is a chain of random reads into a table too big for a chip to carry. The wait is the same physics for everyone. + +**Miners hold the switch.** Spare defences are written into the rules, switched off. A 90% miner signal turns one on. No fork. + +A custom chip gains under 2x. Model published, bounty standing. [link: the numbers page] + +Litepaper only, a fourth paragraph: No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured. + +Site card placement: the Mine section of `site/index.html` beside "no chip can be built for it" (which this card replaces: the claim is a bounded gain, not impossibility). Litepaper placement: `site/litepaper.html` section `mining`, replacing the paragraph that begins "Everything above is automatic" and the "What Igneum does not claim" line on chips; the vs RandomX table keeps its rows, with the "Changes over time" row's Igneum cell reading "A new program every hour, its memory pattern and read widths with it; era draws and reserved families on a schedule fixed at genesis". + +## Level 3: the numbers page (`site/bench.html`, section "Counter ASIC") + +Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per hash, the latency-bound share (rate over the card's random-read ceiling per load), the CPU verifier per warp, with machine, date and command. The chip model before and after Counter ASIC 2.0 (the m16 model's gain arithmetic at the v2 class and at the v3 class, with the SRAM a mirror needs, cited or approximate as the analysis says). The bounty terms (spec O-1.17: the leaderboard by card model, the standing bounty for any chip design beating a GPU by more than 2x, January 2027). Here the layers are named next to their numbers: read width, per-program mix, scratch, era layout, working set, hot table, cache schedule, the reserved integer-matrix family. + +| Card | v2 MH/s | v3 MH/s | v3 bytes per hash | Latency-bound share v3 | Verifier ms per warp v3 | +|---|---|---|---|---|---| +| Apple M5 Max (Metal) | 27.7 | [owed: readwidth table] | | | | +| RTX 5090 (CUDA) | [owed] | [owed] | | | | +| RX 9070 XT (OpenCL, eGPU) | 18.0 | [owed] | | | | + +Chip model, before and after: [owed: from docs/analysis/sram-mirror.md and the hot-table and scratch analyses]. + +## Level 4: the analysis documents + +`docs/plans/counter-asic-2.md` (the plan and the layer table), `docs/plans/read-width.md`, `docs/plans/era-layout.md`, `docs/plans/hot-table.md`, `docs/analysis/scratch-soundness.md`, `docs/analysis/sram-mirror.md`, `docs/analysis/int8-matrix-family.md`, `docs/analysis/m16-recompute-attacker-2026-10-05.md`, the bench log entries of 5 October 2026 (night), the specification sections 1.4 to 1.13. From 002c455eaaa19c16d07f07eceb72e4bbae9a8a7d Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:11:20 +0000 Subject: [PATCH 011/131] Counter ASIC 2.0 status: the four-level public description deliverable --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 2867c71e7..333c2c006 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -63,3 +63,5 @@ Probe ceilings from round 1 (dependent reads per second at 1024 MiB, device time | RX 9070 XT | 2.42 G | 2.43 G | 2.47 G (158 GB/s) | 636 GB/s | the project lead's width rule applied to the ceilings alone (the measured v3 rates will replace this when the table lands): a 128-load hash at 64 B reads 8,192 B; at the 5090's 64 B ceiling that is 71 MH/s and 584 GB/s, 37% of its stream bandwidth, over the one-third margin the rule sets; at 16 B it is 2,048 B per hash, 141 MH/s and 288 GB/s, 18%, inside the margin; 4 B is 9%. On the 9070 XT every width costs the same line fetch (2.4 G/s), so 16 B is where AMD gains 4x the bytes per hash at no cost and the 5090 stays latency-bound with margin. Provisional width under the rule: 16 B (w16), pending the measured rates and the latency-bound share per card. + +Added deliverable (20:40): the public description in four levels, `docs/plans/counter-asic-2-public.md` (aec53bb): levels 1 and 2 are written as copy; level 3 carries the bench table with `[owed]` markers for every number not yet measured; level 4 lists the documents. The integration branch applies levels 1 to 3 to `site/index.html`, `site/litepaper.html` and `site/bench.html` with the final numbers. From ece48696f32f62749eddef509a81d1d94acb9bc1 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:15:00 +0000 Subject: [PATCH 012/131] Counter ASIC 2.0 status 20:55: layers 6 and 7, node agent, 0.3.11 scope with proving v1 --- docs/plans/counter-asic-2-rollout.md | 10 ++++++++-- docs/plans/counter-asic-2-status.md | 8 ++++++++ 2 files changed, 16 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 0e37bf481..81f355d32 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,6 +69,12 @@ If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-st 0.3.11 through the shipper's pipeline (`tools/ship-app.mjs`, the plan shape of `docs/plans/release-0.3.10.md`). The 0.3.10 shipper finishes 0.3.10 first, then takes the 0.3.11 tree, or the coordinator ships 0.3.11 with the same runbook if the shipper has stopped. The publish order is that of `docs/plans/finality-v3-devnet-publish.md`: the override field, the digest handshake, hand nodes and the seed first, then the manifest with `--activation-height` and the deadline note, then the apps, then the digest sweep, then the HiveOS package republish. +## 8a. Proving v1 rides with it + +the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2): fork branch proving-v1 (from a24ab01a: segment records, the pool split, five RPCs, p2p message 72, parameters proving_v1_activation_daa, segment_blocks, unproven_daa, aggregator_share_bps, in the digest only once set) and its app branch, rebased onto the 0.3.10 tips. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. + +Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (second update key with revocation), rig-install, pool-v0 (its own service, no app change), repro-bench; fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. + ## 9. After the publish: Counter ASIC 3.0 The ASIC-history agent sends its ranked additions; they are measured the same way, folded into the class as v4 behind its own activation, same gates, same rollout; `docs/plans/counter-asic-3.md`. @@ -81,6 +87,6 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa | The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `` | the same | | The hot table size (layer 5) | 32, 64, 96 MB | `` | | | The era draw bounds (layers 4 and 8) | the draw as written in `docs/plans/era-layout.md` | | | -| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | `` | | -| Layer 7 | reserved family, unlock by era height or 90% signal | `` | | +| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | Recommended C: the cache doubles when the dataset doubles (512 MiB at year 4); not delegated, a gate 1 decision for the project lead | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate | +| Layer 7 | reserved family, unlock by era height or 90% signal | Reserve R1 = mm8 (uint8 8x16 by 16x8 tile per unit), W_new 4, unlock era 4 or 90% signal; switched off, no consensus effect tonight | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) | | The activation height N4 | the rule of section 3 | | | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 333c2c006..3671826d0 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -65,3 +65,11 @@ Probe ceilings from round 1 (dependent reads per second at 1024 MiB, device time the project lead's width rule applied to the ceilings alone (the measured v3 rates will replace this when the table lands): a 128-load hash at 64 B reads 8,192 B; at the 5090's 64 B ceiling that is 71 MH/s and 584 GB/s, 37% of its stream bandwidth, over the one-third margin the rule sets; at 16 B it is 2,048 B per hash, 141 MH/s and 288 GB/s, 18%, inside the margin; 4 B is 9%. On the 9070 XT every width costs the same line fetch (2.4 G/s), so 16 B is where AMD gains 4x the bytes per hash at no cost and the 5090 stays latency-bound with margin. Provisional width under the rule: 16 B (w16), pending the measured rates and the latency-bound share per card. Added deliverable (20:40): the public description in four levels, `docs/plans/counter-asic-2-public.md` (aec53bb): levels 1 and 2 are written as copy; level 3 carries the bench table with `[owed]` markers for every number not yet measured; level 4 lists the documents. The integration branch applies levels 1 to 3 to `site/index.html`, `site/litepaper.html` and `site/bench.html` with the final numbers. + +## 20:55 layers 6 and 7 landed; the node-fork agent started; 0.3.11 scope + +Branch ca2-analysis (5d5ba15, f59708d). Layer 6: no cache growth rule exists in the spec; cited bit cells N7 0.027, N5 / N3E / Intel 18A 0.021, N3B 0.0199, N2 0.0175 um^2, array factor 0.70; the 256 MiB mirror is 83 / 64 / 54 mm^2 at N7 / N5 / N2 (74 at N2 with a 96 MB hot table), $13 to $26 of silicon per good die (approximate). The mirror was never unaffordable; the cache's job is to stay above GPU L2 (5090 96 MB, GB202 128 MB). Recommendation C for the project lead (gate 1): cache doubles when the dataset doubles. Layer 7: Metal 4 matmul2d has int8 x int8 -> int32, so a unit-level mm8 tile is native on all three vendors; per-lane dot4 is emulation on Apple (M5 Max: 548 G unsigned dot4/s against an 880 G ALU chain, 1.6x; signed 4.7x). Reserve R1 = mm8, W_new 4, unlock era 4 or 90% signal. Owed: the 5090 and 9070 XT dot4 probe (job prepared: relay/playbooks/dot4-probe.ps1, exe sha256 5adaeb1a...6416f4; publishes when a PC frees). + +Node-fork agent a3f505a9d981300cd started 20:50: ca2-v3 (igneum-pow seam: ProgramClass, Epoch::from_chain_seeds, generator 3 in the program id, class in the pack) and ca2-v3-node (vendor/igneum-node-ca2 under the ca2-v3 worktree, from 21d4c73c, with pack-loop 05ef0fa3 merged): the field in Params, OverrideParams and the digest, the epoch-boundary rounding, the era stand-in, the job line, the fast-time gate script. + +0.3.11 scope (coordinator, 20:52): carries program_class_v3 and proving v1 together (one override object, one digest, one publish; each half under the same gates); the ready half ships as 0.3.11 and the other as 0.3.12 if one lags. Next-cut list recorded in the rollout plan section 8a. From b258094890bffb10116885f0e43139b4b7128fd1 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:16:01 +0000 Subject: [PATCH 013/131] Counter ASIC 2.0: layer 6 option C and layer 7 R1 decided (delegated), the chip headline rule for the scratch share --- docs/plans/counter-asic-2-rollout.md | 8 ++++++-- docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 10 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 81f355d32..68e45d7bd 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -52,6 +52,10 @@ the project lead's rules, applied by the coordinator and recorded here with the | Scratch share (layer 3) | the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap | `` | | | Activation height N4 | devnet tip + 14,400 at publish, checked >= 10,800 at publish, rounded up to the epoch boundary | `` | | +### 6a. The chip model's headline, and how the scratch share is chosen + +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (54 to 83 mm^2, $13 to $26 of silicon, approximate, `docs/analysis/sram-mirror.md`) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. If no scratch share under the 6 GB cap gets that chip under 2x, the status file and the public copy's level 3 say so plainly, and the site's "under 2x" claim is qualified until the mixer multiplier or the cache rule closes it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). + ## 7. Gates before any publish (all of them, no exceptions) | # | Gate | Evidence required | State | @@ -87,6 +91,6 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa | The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `` | the same | | The hot table size (layer 5) | 32, 64, 96 MB | `` | | | The era draw bounds (layers 4 and 8) | the draw as written in `docs/plans/era-layout.md` | | | -| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | Recommended C: the cache doubles when the dataset doubles (512 MiB at year 4); not delegated, a gate 1 decision for the project lead | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate | -| Layer 7 | reserved family, unlock by era height or 90% signal | Reserve R1 = mm8 (uint8 8x16 by 16x8 tile per unit), W_new 4, unlock era 4 or 90% signal; switched off, no consensus effect tonight | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) | +| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate | +| Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) | | The activation height N4 | the rule of section 3 | | | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 3671826d0..b03f2cb9c 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -73,3 +73,7 @@ Branch ca2-analysis (5d5ba15, f59708d). Layer 6: no cache growth rule exists in Node-fork agent a3f505a9d981300cd started 20:50: ca2-v3 (igneum-pow seam: ProgramClass, Epoch::from_chain_seeds, generator 3 in the program id, class in the pack) and ca2-v3-node (vendor/igneum-node-ca2 under the ca2-v3 worktree, from 21d4c73c, with pack-loop 05ef0fa3 merged): the field in Params, OverrideParams and the digest, the epoch-boundary rounding, the era stand-in, the job line, the fast-time gate script. 0.3.11 scope (coordinator, 20:52): carries program_class_v3 and proving v1 together (one override object, one digest, one publish; each half under the same gates); the ready half ships as 0.3.11 and the other as 0.3.12 if one lags. Next-cut list recorded in the rollout plan section 8a. + +## 21:05 decisions recorded: layer 6 option C, layer 7 R1 = mm8; the chip headline + +Layer 6 DECIDED (delegated; the project lead confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles (256 MiB genesis, 512 MiB year 4, 1 GiB year 12); one-core fill 0.2 / 0.4 / 0.8 s, under 1 s at every step. Layer 7 DECIDED: reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal. The chip model's headline now names the on-die-cache recompute chip (54 to 83 mm^2) as a row per variant; the scratch share is chosen as the smallest share at which that chip's gain falls under 1.5x, else said plainly and the public "under 2x" claim qualified. The soundness agent carries that table (its question 2) with M16's mixer multiplier beside it. From cbbc81f00270e4fec775e92685ecfc1100a71ed7 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:18:39 +0000 Subject: [PATCH 014/131] Counter ASIC 2.0: the shipped-density correction to the SRAM mirror (130 to 165 mm^2), carried into the rollout and public docs --- docs/plans/counter-asic-2-public.md | 2 +- docs/plans/counter-asic-2-rollout.md | 4 ++-- docs/plans/counter-asic-2-status.md | 4 ++++ 3 files changed, 7 insertions(+), 3 deletions(-) diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index 63be2fbca..ce0615898 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -32,7 +32,7 @@ Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per h | RTX 5090 (CUDA) | [owed] | [owed] | | | | | RX 9070 XT (OpenCL, eGPU) | 18.0 | [owed] | | | | -Chip model, before and after: [owed: from docs/analysis/sram-mirror.md and the hot-table and scratch analyses]. +Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant]. ## Level 4: the analysis documents diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 68e45d7bd..198d38e93 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (54 to 83 mm^2, $13 to $26 of silicon, approximate, `docs/analysis/sram-mirror.md`) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. If no scratch share under the 6 GB cap gets that chip under 2x, the status file and the public copy's level 3 say so plainly, and the site's "under 2x" claim is qualified until the mixer multiplier or the cache rule closes it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. If no scratch share under the 6 GB cap gets that chip under 2x, the status file and the public copy's level 3 say so plainly, and the site's "under 2x" claim is qualified until the mixer multiplier or the cache rule closes it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ## 7. Gates before any publish (all of them, no exceptions) @@ -91,6 +91,6 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa | The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `` | the same | | The hot table size (layer 5) | 32, 64, 96 MB | `` | | | The era draw bounds (layers 4 and 8) | the draw as written in `docs/plans/era-layout.md` | | | -| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate | +| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate | | Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) | | The activation height N4 | the rule of section 3 | | | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index b03f2cb9c..423ecb1f0 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -77,3 +77,7 @@ Node-fork agent a3f505a9d981300cd started 20:50: ca2-v3 (igneum-pow seam: Progra ## 21:05 decisions recorded: layer 6 option C, layer 7 R1 = mm8; the chip headline Layer 6 DECIDED (delegated; the project lead confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles (256 MiB genesis, 512 MiB year 4, 1 GiB year 12); one-core fill 0.2 / 0.4 / 0.8 s, under 1 s at every step. Layer 7 DECIDED: reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal. The chip model's headline now names the on-die-cache recompute chip (54 to 83 mm^2) as a row per variant; the scratch share is chosen as the smallest share at which that chip's gain falls under 1.5x, else said plainly and the public "under 2x" claim qualified. The soundness agent carries that table (its question 2) with M16's mixer multiplier beside it. + +## 21:15 correction to the SRAM mirror figures (chip-economics research, cluster D) + +Bit cell x 0.70 understates real die area. Shipped cache dies: AMD 3D V-Cache 64 MB on 41 mm^2 at 7 nm (1.56 MB/mm^2, Tom's Hardware, Hot Chips August 2021); Graphcore GC200 900 MB on 823 mm^2 with compute (1.09 MB/mm^2); Groq TSP 220 MB on 725 mm^2 at 14 nm (0.30 MB/mm^2); TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30% assist overhead (SemiAnalysis, December 2022). A 256 MiB mirror is about 165 mm^2 at 7 nm on the densest shipped cache-only die and about 130 mm^2 at N5/N3E, not 54 to 83 mm^2; cost per die 2 to 3x the earlier figure; the conclusion (affordable for a funded chip) stands. The analysis agent is redoing the table with both columns; the soundness agent carries the corrected density into the chip row. Latency citations behind the latency-bound rule, to be added: DRAM row cycle 40 to 48 ns across DDR4, GDDR5, HBM2 (Li, Reddy, Jacob, MEMSYS 2018); latency 1.3x in two decades against bandwidth 20x (Chang 2017); no shipped mining chip used HBM or stacked memory. From bcb3f20bd9b4091123c89f7d309abcdc68390696 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:19:32 +0000 Subject: [PATCH 015/131] Counter ASIC 2.0 status 21:25: proving v1 state, PC 2 queue --- docs/plans/counter-asic-2-status.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 423ecb1f0..d670e1b34 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -81,3 +81,9 @@ Layer 6 DECIDED (delegated; the project lead confirms for the public testnet gen ## 21:15 correction to the SRAM mirror figures (chip-economics research, cluster D) Bit cell x 0.70 understates real die area. Shipped cache dies: AMD 3D V-Cache 64 MB on 41 mm^2 at 7 nm (1.56 MB/mm^2, Tom's Hardware, Hot Chips August 2021); Graphcore GC200 900 MB on 823 mm^2 with compute (1.09 MB/mm^2); Groq TSP 220 MB on 725 mm^2 at 14 nm (0.30 MB/mm^2); TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30% assist overhead (SemiAnalysis, December 2022). A 256 MiB mirror is about 165 mm^2 at 7 nm on the densest shipped cache-only die and about 130 mm^2 at N5/N3E, not 54 to 83 mm^2; cost per die 2 to 3x the earlier figure; the conclusion (affordable for a funded chip) stands. The analysis agent is redoing the table with both columns; the soundness agent carries the corrected density into the chip row. Latency citations behind the latency-bound rule, to be added: DRAM row cycle 40 to 48 ns across DDR4, GDDR5, HBM2 (Li, Reddy, Jacob, MEMSYS 2018); latency 1.3x in two decades against bandwidth 20x (Chang 2017); no shipped mining chip used HBM or stacked memory. + +## 21:25 proving v1 state for 0.3.11; PC 2 occupancy + +Proving v1 (acd4f36bc2c07a4e2): fork proving-v1 b177718e on a24ab01a (told to rebase onto commit 21d4c73c now), app proving-v1 79bc820 on a93199a. Override fields proving_v1_activation_daa (tip + 14,400 at publish), proving_v1_segment_blocks 4, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Harness: `tools/proving-v1/net.mjs --secs 1500` PASSED (21 checks) in 197.3 s on b177718e; rerun owed on the final tree. The pinned guests do not change. Shared files with ca2-v3-node: params.rs, daemon.rs, igneum/miner/src/main.rs, override-60x.json; both agents keep separable hunks. + +PC 2 is held by the proving agent's memsweep-pc2-pv1 (about 20 min, miners stopped) and a second run (about 10 min). Queue after it: the ca2 node suites, then the readwidth, era, cache and dot4 measurement jobs. PC 1 is held by readwidth's run-readwidth-9070-20261005c until it reports. From 5c3a35c6deda18755183a837be59f4971fda30f6 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:21:05 +0000 Subject: [PATCH 016/131] Counter ASIC 2.0 status 21:35: the corrected SRAM mirror table --- docs/plans/counter-asic-2-status.md | 13 +++++++++++++ 1 file changed, 13 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index d670e1b34..00716e78c 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -87,3 +87,16 @@ Bit cell x 0.70 understates real die area. Shipped cache dies: AMD 3D V-Cache 64 Proving v1 (acd4f36bc2c07a4e2): fork proving-v1 b177718e on a24ab01a (told to rebase onto commit 21d4c73c now), app proving-v1 79bc820 on a93199a. Override fields proving_v1_activation_daa (tip + 14,400 at publish), proving_v1_segment_blocks 4, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Harness: `tools/proving-v1/net.mjs --secs 1500` PASSED (21 checks) in 197.3 s on b177718e; rerun owed on the final tree. The pinned guests do not change. Shared files with ca2-v3-node: params.rs, daemon.rs, igneum/miner/src/main.rs, override-60x.json; both agents keep separable hunks. PC 2 is held by the proving agent's memsweep-pc2-pv1 (about 20 min, miners stopped) and a second run (about 10 min). Queue after it: the ca2 node suites, then the readwidth, era, cache and dot4 measurement jobs. PC 1 is held by readwidth's run-readwidth-9070-20261005c until it reports. + +## 21:35 sram-mirror.md revision 2 (ca2-analysis) + +Two columns, headline = shipped-product density (AMD V-Cache 41 mm^2 per 64 MiB at N7, scaled by the bit-cell ratio), lower bound = bit cell x 0.70. mm^2 and $ per good die (D0 0.1 per cm^2, wafer prices approximate), headline / lower bound: + +| Node, wafer $ | 256 MiB | 256 + 96 MB hot table | 1 GiB | +|---|---|---|---| +| N7, $9,500 | 164 / 83 mm^2, $30 / $13 | 226 / 114, $44 / $19 | 656 / 331, $224 / $75 | +| N5, N3E, 18A, $20,000 | 128 / 64, $46 / $21 | 175 / 89, $68 / $30 | 510 / 258, $306 / $111 | +| N3B, $20,000 | 121 / 61, $43 / $20 | 166 / 84, $63 / $28 | 483 / 244, $280 / $103 | +| N2, $30,000 | 106 / 54, $56 / $26 | 146 / 74, $81 / $37 | 425 / 215, $343 / $131 | + +Year 10 at the 6% per year trend: 59 mm^2 for the flat cache (82 with the hot table), 8% of a 750 mm^2 die. One reticle holds 1.3 GiB (N7) to 1.9 GiB (N2); mirror share of a 750 mm^2 die at year 0: 14% (22% with the hot table), inside M16's 13 to 40% band. Recommendation unchanged: C. Latency section added (MEMSYS 2018, Chang 2017, the mining-chip memory-type note), marked as research the agent did not re-read tonight apart from the V-Cache figure. From f4fc6515128cceab380652e98d35e4188ca26292 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:23:13 +0000 Subject: [PATCH 017/131] Counter ASIC 2.0: layer 3 decided (scratch share 0, chip 2.4x at every share), the claim qualified, public level 3 headline --- docs/plans/counter-asic-2-public.md | 2 ++ docs/plans/counter-asic-2-rollout.md | 4 ++-- docs/plans/counter-asic-2-status.md | 8 ++++++++ 3 files changed, 12 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index ce0615898..1eaeb0d07 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -24,6 +24,8 @@ Site card placement: the Mine section of `site/index.html` beside "no chip can b ## Level 3: the numbers page (`site/bench.html`, section "Counter ASIC") +Headline of the chip model (5 October 2026, night): the strongest chip holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) and computes dataset items on the fly; its gain over the RTX 5090 is 2.4x as the parameters stand, and no write-scratch share within an 8 GB card's budget changes that. The lever that does is the dataset item's mixer cost (x4: 1.8x with a 3x fixed-function factor, verifier 1.6 to 4.8 ms per warp). Until that or the cache rule is in the class, the claim reads: a chip gains about 2.4x by the published model, the target is under 2x, the model and the bounty are public. + Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per hash, the latency-bound share (rate over the card's random-read ceiling per load), the CPU verifier per warp, with machine, date and command. The chip model before and after Counter ASIC 2.0 (the m16 model's gain arithmetic at the v2 class and at the v3 class, with the SRAM a mirror needs, cited or approximate as the analysis says). The bounty terms (spec O-1.17: the leaderboard by card model, the standing bounty for any chip design beating a GPU by more than 2x, January 2027). Here the layers are named next to their numbers: read width, per-program mix, scratch, era layout, working set, hot table, cache schedule, the reserved integer-matrix family. | Card | v2 MH/s | v3 MH/s | v3 bytes per hash | Latency-bound share v3 | Verifier ms per warp v3 | diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 198d38e93..a49e22ee6 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -49,12 +49,12 @@ the project lead's rules, applied by the coordinator and recorded here with the |---|---|---|---| | Width (layers 1 and 2) | the widest read that keeps every card we own latency-bound (achieved loads within 90% of the probe ceiling) with margin on the 5090 (its bytes per hash under a third of its bandwidth at the measured rate) | `` | | | Per-load mix (layer 2) | in, if the min-to-max spread across six programs is under 5% per card | `` | | -| Scratch share (layer 3) | the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap | `` | | +| Scratch share (layer 3) | the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap | DECIDED (5 October 2026, delegated): 0. No share under the cap moves the on-die-cache recompute chip, so layer 3 is not adopted into v3 | `docs/analysis/scratch-soundness.md` (ca2-soundness a465881): chip 333 MH/s against the 5090's measured 139.7 = 2.4x at 0% RMW; 2.4x at 12.5 / 25 / 50% replaced (chip 381 / 443 / 661 against 160 / 186 / 279 projected) and 2.4x or more added, at 32 and 128 KB; the chip keeps the scratch implicitly in 80 to 320 B per lane because the verifier resets it per unit | | Activation height N4 | devnet tip + 14,400 at publish, checked >= 10,800 at publish, rounded up to the epoch boundary | `` | | ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. If no scratch share under the 6 GB cap gets that chip under 2x, the status file and the public copy's level 3 say so plainly, and the site's "under 2x" claim is qualified until the mixer multiplier or the cache rule closes it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. Whether x4 goes into v3 tonight or into Counter ASIC 3.0 is with the coordinator; the plan proceeds on 3.0 unless told otherwise. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ## 7. Gates before any publish (all of them, no exceptions) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 00716e78c..27ab3f693 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -100,3 +100,11 @@ Two columns, headline = shipped-product density (AMD V-Cache 41 mm^2 per 64 MiB | N2, $30,000 | 106 / 54, $56 / $26 | 146 / 74, $81 / $37 | 425 / 215, $343 / $131 | Year 10 at the 6% per year trend: 59 mm^2 for the flat cache (82 with the hot table), 8% of a 750 mm^2 die. One reticle holds 1.3 GiB (N7) to 1.9 GiB (N2); mirror share of a 750 mm^2 die at year 0: 14% (22% with the hot table), inside M16's 13 to 40% band. Recommendation unchanged: C. Latency section added (MEMSYS 2018, Chang 2017, the mining-chip memory-type note), marked as research the agent did not re-read tonight apart from the V-Cache figure. + +## 21:50 layer 3 soundness landed: the scratch does not move the chip; scratch share decided 0 + +ca2-soundness (0d8f745 tests and trace hook, a465881 doc and bench-log). The on-die-cache recompute chip (N5 headline 128 mm^2, $46) at 50 T op/s: 333 MH/s against the 5090's measured 139.7, 2.4x; at 12.5 / 25 / 50% RMW replaced, chip 381 / 443 / 661 against 5090 projected 160 / 186 / 279, 2.4x each; added, 2.4x or more; 32 or 128 KB alike. The chip keeps the scratch implicitly in 80 to 320 B per lane (the verifier resets it per unit), needs about 530 units in flight, dense scratch 6.2 / 3.1 mm^2 at N5. Under the project lead's rule the scratch share is 0: layer 3 is NOT adopted into v3; the public "under 2x" claim is qualified (public copy level 3 rewritten). The measured lever is M16's mixer multiplier (x2 1.2x at 0.8 to 2.4 ms verify; x4 0.6x at 1.6 to 4.8 ms; 3.6x and 1.8x with a 3x fixed-function factor); whether x4 enters v3 tonight is asked of the coordinator; default: Counter ASIC 3.0. + +Soundness results (Metal, M5 Max): 28/28 edge launches, 200/200 fuzz packs (91 s), 56/56 hand-model edge checks, 42/42 kernels pass the static scratch-mask check with 6 deliberate breaks caught, broken tag and broken lazy fill caught, fingerprint 8c07620f4d9adefd warp-count-independent; re-hit rates 2 to 33% above the birthday bound (slot addresses are register low bits); written words unbiased (worst 3.63 of 6 sigma). Verifier exactness needs a host contract (zero the arena at allocation and at the 32-bit tag wrap, tags from 1), which neither host gives today. Pre-existing on readwidth b970dda: verify::tests::fold_and_wide_fetch overflows under the test profile (wrapping_mul fixes it); passed to the readwidth agent with the class sweep. + +Gate G3 note: the scratch tests (igneum-pow/tests/scratch.rs) join the v3 suite even though the class carries no scratch, parametric over the class; they guard the v2 path's scratch-free invariant at zero cost. From 6a9b7c1283a923deb833818bddbc754ad9762b51 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:24:55 +0000 Subject: [PATCH 018/131] Counter ASIC 2.0: the mixer x4 decision into v3, ca2-mixer agent, status 22:00 --- docs/plans/counter-asic-2-public.md | 2 +- docs/plans/counter-asic-2-rollout.md | 4 ++-- docs/plans/counter-asic-2-status.md | 6 ++++++ 3 files changed, 9 insertions(+), 3 deletions(-) diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index 1eaeb0d07..6294f496c 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -24,7 +24,7 @@ Site card placement: the Mine section of `site/index.html` beside "no chip can b ## Level 3: the numbers page (`site/bench.html`, section "Counter ASIC") -Headline of the chip model (5 October 2026, night): the strongest chip holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) and computes dataset items on the fly; its gain over the RTX 5090 is 2.4x as the parameters stand, and no write-scratch share within an 8 GB card's budget changes that. The lever that does is the dataset item's mixer cost (x4: 1.8x with a 3x fixed-function factor, verifier 1.6 to 4.8 ms per warp). Until that or the cache rule is in the class, the claim reads: a chip gains about 2.4x by the published model, the target is under 2x, the model and the bounty are public. +Headline of the chip model (5 October 2026, night): the strongest chip holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) and computes dataset items on the fly; its gain over the RTX 5090 is 2.4x as the parameters stand, and no write-scratch share within an 8 GB card's budget changes that. The lever that does is the dataset item's mixer cost (x4: 1.8x with a 3x fixed-function factor, verifier 1.6 to 4.8 ms per warp). Decided 5 October 2026 (delegated): the mixer x4 and the cache growth rule enter class v3, so the headline row is the on-die-cache chip against v3 with everything combined. [owed: the combined row from docs/analysis/chip-model-v3.md; if it reads 1.8x, the claim is "under 2x" with the margin stated as thin, and the next levers are named: the mixer x8 and the hot table.] Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per hash, the latency-bound share (rate over the card's random-read ceiling per load), the CPU verifier per warp, with machine, date and command. The chip model before and after Counter ASIC 2.0 (the m16 model's gain arithmetic at the v2 class and at the v3 class, with the SRAM a mirror needs, cited or approximate as the analysis says). The bounty terms (spec O-1.17: the leaderboard by card model, the standing bounty for any chip design beating a GPU by more than 2x, January 2027). Here the layers are named next to their numbers: read width, per-program mix, scratch, era layout, working set, hot table, cache schedule, the reserved integer-matrix family. diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index a49e22ee6..ae3a61583 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -8,7 +8,7 @@ Shape and rules follow `docs/plans/finality-v3-rollout-devnet.md` and the publis Only the lottery hash's program class changes, and only from the first epoch at or above the height. One switch, `program_class_v3_activation_daa`, in `Params` and `OverrideParams` like `difficulty_v2_activation_daa`; default `u64::MAX` (never) on every network. Because one epoch has one program (spec 01 section 1.12), the switch keys on the EPOCH: epoch `e` is class v3 when `3,600 e >= N4`, so the activation is rounded up to an epoch boundary and a block's class is a function of its DAA score alone, as today. -Class v3 = generator version 3: the width rule `` (layer 1 or 2, decision 1), the per-warp scratch at `` RMW and `` per warp (layer 3, under the 6 GB working-set cap), the era draw of the table layout and the working set (layers 4 and 8), the hot table of `` MB from the epoch seed (layer 5). New program id (`generator = 3` in the id's preimage, spec 1.4.6), new packs and vectors, new `IGNEUM_GENERATOR` in every pack, a pack of the other version refused by every implementation (spec 1.4.5 already says so). +Class v3 = generator version 3: the width rule `` (layer 1 or 2, decision 1), no scratch (layer 3 decided out: scratch share 0), the era draw of the table layout and the working set (layers 4 and 8), the hot table of `` MB from the epoch seed (layer 5), the cache growth rule of layer 6 (option C) and the M16 mixer x4 in the dataset item construction (decided 5 October 2026, delegated). New program id (`generator = 3` in the id's preimage, spec 1.4.6), new packs and vectors, new `IGNEUM_GENERATOR` in every pack, a pack of the other version refused by every implementation (spec 1.4.5 already says so). What does not change: the chain, the genesis, the databases, the day key and the 256 MiB cache fill, the dataset items (spec 1.8.5, if the era interleave keeps the item values; the era-layout document says what it costs otherwise), finality, fees, proving. The SP1 guest does not read the lottery hash (`proving/igneum-prove` has no dependency on `igneum-pow`; the pinned guest of DAA 210,000 is a fee-table switch), so no new guest is pinned. The node's `EpochSeeds` gains the class and the era bytes; `IgneumEngine::epoch_for` builds the v3 `Epoch` from them; the miner's `export-pack` and the serve protocol's job line carry the class so a GPU worker regenerates the right pack from the seed bytes. @@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. Whether x4 goes into v3 tonight or into Counter ASIC 3.0 is with the coordinator; the plan proceeds on 3.0 unless told otherwise. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ## 7. Gates before any publish (all of them, no exceptions) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 27ab3f693..016c34900 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -108,3 +108,9 @@ ca2-soundness (0d8f745 tests and trace hook, a465881 doc and bench-log). The on- Soundness results (Metal, M5 Max): 28/28 edge launches, 200/200 fuzz packs (91 s), 56/56 hand-model edge checks, 42/42 kernels pass the static scratch-mask check with 6 deliberate breaks caught, broken tag and broken lazy fill caught, fingerprint 8c07620f4d9adefd warp-count-independent; re-hit rates 2 to 33% above the birthday bound (slot addresses are register low bits); written words unbiased (worst 3.63 of 6 sigma). Verifier exactness needs a host contract (zero the arena at allocation and at the 32-bit tag wrap, tags from 1), which neither host gives today. Pre-existing on readwidth b970dda: verify::tests::fold_and_wide_fetch overflows under the test profile (wrapping_mul fixes it); passed to the readwidth agent with the class sweep. Gate G3 note: the scratch tests (igneum-pow/tests/scratch.rs) join the v3 suite even though the class carries no scratch, parametric over the class; they guard the v2 path's scratch-free invariant at zero cost. + +## 22:00 decided: M16 mixer x4 into v3; agent ca2-mixer started + +Coordinator's decision under the project lead's delegation (recorded in the rollout plan section 6a): the mixer multiplier x4 and the cache growth rule (option C) enter class v3 behind the same activation; layer 3 stays out at scratch share 0, its soundness document and pack-contract tests kept. Agent af345b1e2c541ffbb (branch ca2-mixer) implements `mixer_mult` as a class parameter (m mixer applications per round, the 8 dependent reads unchanged), the `cache_log2_words(day)` schedule (doublings at years 4 and 12 with the dataset stepping to the next power of two), re-cuts the v3 dataset vectors, re-runs the soundness suite, measures the verifier (v2 0.604 ms per warp; v3 expected 1.6 to 4.8 ms) and the 1 GiB build on the Mac, prepares the 5090 and 9070 XT build-time job, and writes docs/analysis/chip-model-v3.md with the combined headline row (fixed-function factor included). The claim on the site reads "under 2x" only if that row does; else qualified, with the mixer x8 and the hot table named as the next levers. + +Agents now: ca2-era (a452664c512c73b9b), ca2-cache (a5271cf269757b118), ca2-node (a3f505a9d981300cd), ca2-mixer (af345b1e2c541ffbb). Done: ca2-analysis, ca2-soundness. Waiting: the readwidth PC table; PC 2 (proving memsweep runs) and PC 1 (readwidth 9070 round). From fff1994c8a6a45e8f925897ae8c8d15899f79e77 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:27:01 +0000 Subject: [PATCH 019/131] Counter ASIC 2.0: layers 1, 2, 3 decided from the readwidth table, layer 5 Mac rows and the added-form redesign, user tiers, public level 2 edit --- docs/plans/counter-asic-2-public.md | 10 ++++++---- docs/plans/counter-asic-2-rollout.md | 12 ++++++++---- docs/plans/counter-asic-2-status.md | 19 +++++++++++++++++++ 3 files changed, 33 insertions(+), 8 deletions(-) diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index 6294f496c..4ccc5e2da 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -10,7 +10,7 @@ Built for graphics cards. A custom chip gains under 2x, and the model and the bo Three ideas, no layer names, no widths, no SRAM. -**The hash rewrites itself.** A new program every hour, drawn from the chain. Its memory pattern and its read widths change with it. The rules change on a schedule fixed at launch. No release, no vote. +**The hash rewrites itself.** A new program every hour, drawn from the chain. Its memory pattern changes with it. The rules change on a schedule fixed at launch. No release, no vote. **It waits on memory, not maths.** Every hash is a chain of random reads into a table too big for a chip to carry. The wait is the same physics for everyone. @@ -30,9 +30,11 @@ Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per h | Card | v2 MH/s | v3 MH/s | v3 bytes per hash | Latency-bound share v3 | Verifier ms per warp v3 | |---|---|---|---|---|---| -| Apple M5 Max (Metal) | 27.7 | [owed: readwidth table] | | | | -| RTX 5090 (CUDA) | [owed] | [owed] | | | | -| RX 9070 XT (OpenCL, eGPU) | 18.0 | [owed] | | | | +| Apple M5 Max (Metal) | 27.74 | [owed: the v3 pack] | 512 + hot | [owed] | [owed: ca2-mixer] | +| RTX 5090 (CUDA) | 136.1 | [owed] | 512 + hot | 0.96 at v2 | [owed] | +| RX 9070 XT (OpenCL, eGPU) | 18.15 | [owed] | 512 + hot | 0.87 at v2 | [owed] | + +AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), near parity per pound and about 4.5x worse per watt; the card's memory system, not a tuning gap. Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant]. diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index ae3a61583..17843bdee 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -47,8 +47,8 @@ the project lead's rules, applied by the coordinator and recorded here with the | Decision | the project lead's rule | Choice | The number | |---|---|---|---| -| Width (layers 1 and 2) | the widest read that keeps every card we own latency-bound (achieved loads within 90% of the probe ceiling) with margin on the 5090 (its bytes per hash under a third of its bandwidth at the measured rate) | `` | | -| Per-load mix (layer 2) | in, if the min-to-max spread across six programs is under 5% per card | `` | | +| Width (layer 1) | the widest read that keeps every card we own latency-bound (achieved loads within 90% of the probe ceiling) with margin on the 5090 (its bytes per hash under a third of its bandwidth at the measured rate) | DECIDED (5 October 2026, delegated): keep v2, 128 x 4 B. w16 passes the rule (shares 0.90 / 0.84 / 1.03, 18% of the 5090's stream) but closes nothing and does not move the chip row; w64 and w64x4 make the 5090 bandwidth-bound (share 0.58 / 0.56, 37% of stream) | `docs/plans/read-width.md` (readwidth e752fc7): v2 5090 136.1 MH/s, 9070 XT 18.15, M5 Max 27.74 (gap 7.5x); w16 139.8 / 17.90 / 28.26 (gap 7.8x); w64 71.9 / 17.59 / 28.27 (gap 4.1x); the 9070 XT does 2.4 G dependent reads/s at every width | +| Per-load mix (layer 2) | in, if the min-to-max spread across six programs is under 5% per card | DECIDED (5 October 2026, delegated): out | spreads of the median over six programs: mix 50/35/15 5090 18.8%, 9070 XT 7.4%, M5 Max 11.3%; mix 25/50/25 22.3% / 5.5% / 8.1% | | Scratch share (layer 3) | the smallest share at which the chip model's gain falls under 1.5x at the lowest GPU cost, within the 6 GB working-set cap | DECIDED (5 October 2026, delegated): 0. No share under the cap moves the on-die-cache recompute chip, so layer 3 is not adopted into v3 | `docs/analysis/scratch-soundness.md` (ca2-soundness a465881): chip 333 MH/s against the 5090's measured 139.7 = 2.4x at 0% RMW; 2.4x at 12.5 / 25 / 50% replaced (chip 381 / 443 / 661 against 160 / 186 / 279 projected) and 2.4x or more added, at 32 and 128 KB; the chip keeps the scratch implicitly in 80 to 320 B per lane because the verifier resets it per unit | | Activation height N4 | devnet tip + 14,400 at publish, checked >= 10,800 at publish, rounded up to the epoch boundary | `` | | @@ -56,6 +56,10 @@ the project lead's rules, applied by the coordinator and recorded here with the The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +### 6b. The user tiers (the consequences rule) + +AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second at 1 GiB, measured tonight), near parity per pound and about 4.5x worse per watt (0.089 against 0.398 MH/W, measured 5 October 2026). This is the card's memory system, not a tuning gap: no read width closes it without making the 5090 bandwidth-bound. The level 3 numbers page states it. + ## 7. Gates before any publish (all of them, no exceptions) | # | Gate | Evidence required | State | @@ -89,8 +93,8 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa |---|---|---|---| | The width rule (layers 1 and 2) | 4 B fixed; 16 B; 64 B; the per-program mix | `` | `docs/plans/read-width.md` | | The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `` | the same | -| The hot table size (layer 5) | 32, 64, 96 MB | `` | | -| The era draw bounds (layers 4 and 8) | the draw as written in `docs/plans/era-layout.md` | | | +| The hot table (layer 5) | 32, 64, 96 MB; replaced or added | In the ADDED form only (16 dataset loads plus k hot loads): the replaced form lets the on-die-cache chip skip item derivations (k = 4: chip x1.33 against the GPU's measured x1.05 to x1.22). Size = the largest table resident on every card we own, pending the PC rows; Mac rows (replaced form, M5 Max): hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k8 x1.71; Apple OpenCL dependent-read probe 32 / 64 / 96 / 1024 MiB: 21.7 / 12.8 / 12.3 / 3.50 G loads/s | `docs/plans/hot-table.md` (ca2-cache 65bc7a7) | +| The era draws (layers 4 and 8) | in, if the min-to-max spread across six drawn eras is under 5% per card; item size fixed at 4 B, draws of stride, interleave and the working-set window at or above 256 MiB | pending the six-era table | `docs/plans/era-layout.md` | | The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate | | Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) | | The activation height N4 | the rule of section 3 | | | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 016c34900..d3e41daa3 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -114,3 +114,22 @@ Gate G3 note: the scratch tests (igneum-pow/tests/scratch.rs) join the v3 suite Coordinator's decision under the project lead's delegation (recorded in the rollout plan section 6a): the mixer multiplier x4 and the cache growth rule (option C) enter class v3 behind the same activation; layer 3 stays out at scratch share 0, its soundness document and pack-contract tests kept. Agent af345b1e2c541ffbb (branch ca2-mixer) implements `mixer_mult` as a class parameter (m mixer applications per round, the 8 dependent reads unchanged), the `cache_log2_words(day)` schedule (doublings at years 4 and 12 with the dataset stepping to the next power of two), re-cuts the v3 dataset vectors, re-runs the soundness suite, measures the verifier (v2 0.604 ms per warp; v3 expected 1.6 to 4.8 ms) and the 1 GiB build on the Mac, prepares the 5090 and 9070 XT build-time job, and writes docs/analysis/chip-model-v3.md with the combined headline row (fixed-function factor included). The claim on the site reads "under 2x" only if that row does; else qualified, with the mixer x8 and the hot table named as the next levers. Agents now: ca2-era (a452664c512c73b9b), ca2-cache (a5271cf269757b118), ca2-node (a3f505a9d981300cd), ca2-mixer (af345b1e2c541ffbb). Done: ca2-analysis, ca2-soundness. Waiting: the readwidth PC table; PC 2 (proving memsweep runs) and PC 1 (readwidth 9070 round). + +## 22:15 the readwidth table landed; layers 1, 2, 3 decided; layer 5 measured on the Mac and redesigned + +Readwidth e752fc7 (`docs/plans/read-width.md`), bit-exact on Metal, Apple OpenCL, the 5090 (NVRTC) and the 9070 XT, both PCs released. MH/s (latency-bound share): + +| Class | RTX 5090 | RX 9070 XT | M5 Max | Gap | +|---|---|---|---|---| +| v2 (128 x 4 B) | 136.1 (0.96) | 18.15 (0.87) | 27.74 (1.01) | 7.5x | +| w16 | 139.8 (0.90) | 17.90 (0.84) | 28.26 (1.03) | 7.8x | +| w64 | 71.9 (0.58, 37% of stream) | 17.59 (0.78) | 28.27 (1.03) | 4.1x | +| w64x4 (32 loads) | 275.3 (0.56) | 75.19 (0.84) | 109.7 (1.00) | 3.7x | +| mix 50/35/15, six programs, spread of median | 18.8% | 7.4% | 11.3% | | +| mix 25/50/25 | 22.3% | 5.5% | 8.1% | | +| scratch 32 KB at 12.5 / 25 / 50% | -18 / -21 / -12% | -18 / -22 / -21% | -7 / +12 / +74% | | +| scratch 128 KB | -21 / -30 / -48% | -21 / -27 / -33% | -7 / -1 / +25% | | + +Decisions (the project lead's rules, delegated): layer 1 keep v2 (w16 passes the rule but closes nothing and does not move the chip row; the vector re-cut is not worth it); layer 2 out (spread over 5% on every card); layer 3 out (scratch share 0). The AMD gap is the card's dependent-read rate (2.4 G/s at every width), stated for the user tiers in the rollout plan 6b. + +Layer 5 (ca2-cache 53ef59f, 011cc0a, 86726cd, 65bc7a7; `docs/plans/hot-table.md`): five packs bit-exact on Metal and Apple OpenCL (96/96 each). M5 Max rates against v2 27.68: hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k2 x1.00, hot64k8 x1.71; verifier 0.344 to 0.560 ms against 0.626; hot fill per epoch 24 / 46 / 73 ms on one core, 0.07 / 0.15 / 0.22 ms on the GPU; Apple OpenCL probe 32 / 64 / 96 / 1024 MiB 21.7 / 12.8 / 12.3 / 3.50 G loads/s. Redesign ordered: hot loads ADDED beside the 16 dataset loads (the replaced form lets the on-die-cache chip skip item derivations and worsens the gain); the agent re-measures the added form and rebuilds the PC job. PC 1 is given to the dot4 probe (under 15 min), then to the era agent, then the hot-table job; PC 2 stays the proving agent's. From 724c5eee3755035d94afbb97efdb21a7cf26317a Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:27:58 +0000 Subject: [PATCH 020/131] Counter ASIC 2.0: layer 9 row, C1 constraint, era mid-way and the attempt-aware packfile bug, C10 and C11 answers --- docs/plans/counter-asic-2-rollout.md | 4 ++++ docs/plans/counter-asic-2-status.md | 8 ++++++++ 2 files changed, 12 insertions(+) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 17843bdee..9393ea34c 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -77,6 +77,10 @@ If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-st 0.3.11 through the shipper's pipeline (`tools/ship-app.mjs`, the plan shape of `docs/plans/release-0.3.10.md`). The 0.3.10 shipper finishes 0.3.10 first, then takes the 0.3.11 tree, or the coordinator ships 0.3.11 with the same runbook if the shipper has stopped. The publish order is that of `docs/plans/finality-v3-devnet-publish.md`: the override field, the digest handshake, hand nodes and the seed first, then the manifest with `--activation-height` and the deadline note, then the apps, then the digest sweep, then the HiveOS package republish. +## 7a. A dated constraint from the consequences review (C1) + +The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H. + ## 8a. Proving v1 rides with it the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2): fork branch proving-v1 (from a24ab01a: segment records, the pool split, five RPCs, p2p message 72, parameters proving_v1_activation_daa, segment_blocks, unproven_daa, aggregator_share_bps, in the digest only once set) and its app branch, rebased onto the 0.3.10 tips. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index d3e41daa3..1670df9dc 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -133,3 +133,11 @@ Readwidth e752fc7 (`docs/plans/read-width.md`), bit-exact on Metal, Apple OpenCL Decisions (the project lead's rules, delegated): layer 1 keep v2 (w16 passes the rule but closes nothing and does not move the chip row; the vector re-cut is not worth it); layer 2 out (spread over 5% on every card); layer 3 out (scratch share 0). The AMD gap is the card's dependent-read rate (2.4 G/s at every width), stated for the user tiers in the rollout plan 6b. Layer 5 (ca2-cache 53ef59f, 011cc0a, 86726cd, 65bc7a7; `docs/plans/hot-table.md`): five packs bit-exact on Metal and Apple OpenCL (96/96 each). M5 Max rates against v2 27.68: hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k2 x1.00, hot64k8 x1.71; verifier 0.344 to 0.560 ms against 0.626; hot fill per epoch 24 / 46 / 73 ms on one core, 0.07 / 0.15 / 0.22 ms on the GPU; Apple OpenCL probe 32 / 64 / 96 / 1024 MiB 21.7 / 12.8 / 12.3 / 3.50 G loads/s. Redesign ordered: hot loads ADDED beside the 16 dataset loads (the replaced form lets the on-die-cache chip skip item derivations and worsens the gain); the agent re-measures the added form and rebuilds the PC job. PC 1 is given to the dot4 probe (under 15 min), then to the era agent, then the hot-table job; PC 2 stays the proving agent's. + +## 22:30 layer 9 added; the era draw passes Mac bit-exactness; two class bugs in the harnesses; C1 + +Layer 9 (the project lead: faster program changes): the epoch length becomes an era parameter in the genesis reserve, 1 hour at launch, 10 minutes to 2 hours by draw or 90% signal, reserve-only tonight; an agent (ca2-epoch) designs it beside layers 4 and 8 and measures the compile-ahead cost per card at a 10-minute epoch, the seed-path consequence and the FPGA threat it answers; one row in the rollout plan section 6, one in the level 3 numbers. Spawns when the dot4 probe frees its slot. + +ca2-era mid-way (a452664c512c73b9b): era draw behind LoadClass::era (EraParams beside mix, load_slots, scratch, scratch_kb; verify::load_index; memhard::Layout; one load form in the three emitters; --era, --era-widths), 54 crate tests green, pinned packs byte-identical; six era packs bit-exact on Metal (packbench 6/6), Apple OpenCL (6/6, same fingerprints) and the CUDA emulation (6/6); re-exporting with the width pinned at 4 B and rebasing onto e752fc7; PC job in about 20 minutes. Two class bugs found and fixed on its branch, both outside its layer: (1) proto-cuda/nvrtc/packfile.h re-derived the seed words as attempt 0 only, so ANY pack with IGNEUM_PROGRAM_ATTEMPT >= 1 (5.14% of chain epochs under v2) is refused by the one-click workers with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", the error PC 1 logged on 5 October and attributed to the export race; fixed attempt-aware, with a tampered-attempt refusal test. This fix must ship in the 0.3.11 workers whatever else does. (2) proto-cuda/host.cu and proto-opencl/host.c derived dataset words on the host as mh_item(w >> 4)[w AND 15] instead of the pack's mh_word; fixed. + +Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (21:50, 22:00 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.5x the electricity per hash (0.089 against 0.398 MH/W), near parity per pound; the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. From 35b3e38d2a79eb15bdbcd156302fabe5f5029234 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:29:09 +0000 Subject: [PATCH 021/131] Counter ASIC 2.0 status 22:40: C1 decision and the 16:00Z check, one packfile.h fix --- docs/plans/counter-asic-2-status.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 1670df9dc..1cef880d8 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -141,3 +141,9 @@ Layer 9 (the project lead: faster program changes): the epoch length becomes an ca2-era mid-way (a452664c512c73b9b): era draw behind LoadClass::era (EraParams beside mix, load_slots, scratch, scratch_kb; verify::load_index; memhard::Layout; one load form in the three emitters; --era, --era-widths), 54 crate tests green, pinned packs byte-identical; six era packs bit-exact on Metal (packbench 6/6), Apple OpenCL (6/6, same fingerprints) and the CUDA emulation (6/6); re-exporting with the width pinned at 4 B and rebasing onto e752fc7; PC job in about 20 minutes. Two class bugs found and fixed on its branch, both outside its layer: (1) proto-cuda/nvrtc/packfile.h re-derived the seed words as attempt 0 only, so ANY pack with IGNEUM_PROGRAM_ATTEMPT >= 1 (5.14% of chain epochs under v2) is refused by the one-click workers with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", the error PC 1 logged on 5 October and attributed to the export race; fixed attempt-aware, with a tampered-attempt refusal test. This fix must ship in the 0.3.11 workers whatever else does. (2) proto-cuda/host.cu and proto-opencl/host.c derived dataset words on the host as mh_item(w >> 4)[w AND 15] instead of the pack's mh_word; fixed. Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (21:50, 22:00 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.5x the electricity per hash (0.089 against 0.398 MH/W), near parity per pound; the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. + +## 22:40 C1 decided for the morning; one packfile.h fix for 0.3.11 + +C1 (the fee switch H = 210,000 at about 19:50Z on 6 October; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before 19:50Z. The proving agent recommends (b) unless (a) is certain; the check decides it. + +The attempt-0 packfile.h bug (the era agent's find) is the one that took the fleet down at 18:23Z (epoch 34, attempt 1); it is already fixed attempt-aware on pack-loop af983a7 and merged into the 0.3.10 tree. Rule for 0.3.11: one derivation, one test set: the era and cache agents build their workers on that packfile.h and keep only tests that add a case; the host.cu / host.c mh_word derivation fix (new, the era agent's) stays with its own test. From efeab4b23db82f8a55682403b5544d04180c4d92 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:29:37 +0000 Subject: [PATCH 022/131] Counter ASIC 2.0 rollout: ota-k2 commit on the next-cut list --- docs/plans/counter-asic-2-rollout.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 9393ea34c..b51f04861 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -85,7 +85,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2): fork branch proving-v1 (from a24ab01a: segment records, the pool split, five RPCs, p2p message 72, parameters proving_v1_activation_daa, segment_blocks, unproven_daa, aggregator_share_bps, in the digest only once set) and its app branch, rebased onto the 0.3.10 tips. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. -Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (second update key with revocation), rig-install, pool-v0 (its own service, no app change), repro-bench; fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. +Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install, pool-v0 (its own service, no app change), repro-bench; fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. ## 9. After the publish: Counter ASIC 3.0 From 99ab478f1e1b5ee9d3f68b39c547620041ea743f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:30:46 +0000 Subject: [PATCH 023/131] docs/analysis/card-lifetime-2026-10-05.md: years each card tier keeps mining under the dataset schedule and the working-set rule, against the site and litepaper claims Co-Authored-By: Claude Fable 5.1 --- docs/analysis/card-lifetime-2026-10-05.md | 85 +++++++++++++++++++++++ 1 file changed, 85 insertions(+) create mode 100644 docs/analysis/card-lifetime-2026-10-05.md diff --git a/docs/analysis/card-lifetime-2026-10-05.md b/docs/analysis/card-lifetime-2026-10-05.md new file mode 100644 index 000000000..4a6d565a2 --- /dev/null +++ b/docs/analysis/card-lifetime-2026-10-05.md @@ -0,0 +1,85 @@ +# Card lifetime per tier: how many years a card keeps mining + +5 October 2026. Consequences review, sub-agent of the consequences reviewer. Desk arithmetic only; nothing was run. + +## 1. Inputs + +| Input | Source | Value used | +|---|---|---| +| Dataset schedule | `docs/spec/01-lottery-hash.md` 432 to 437 | 2 GiB at genesis plus 0.5 GiB a year (2,048 + 512 x years MiB) | +| Index mapping (a) | same file, line 442 | multiply-shift: the dataset grows every day, continuous | +| Index mapping (b) | same file, line 442 | power-of-two steps 2, 4, 8 GiB on the schedule's average: 4 GiB at year 4, 8 GiB at year 12; my extrapolation: 16 GiB at year 28, 32 GiB at year 60 | +| Scratch per resident warp | `igneum-wt-ca2-cache/docs/plans/hot-table.md` 66 to 73 | 32 or 128 KiB per warp; 5090 = 170 SMs x 48 warps = 8,160 (approximate, from memory) | +| Hot table, buffers | same file, 70 | hot table 32, 64 or 96 MiB (96 used here); buffers 128 MiB | +| Cache | hot-table.md 70 (resident, 256 MiB in every total) against `igneum-wt-ca2-era/docs/plans/era-layout.md` 93 ("resident only while the day's dataset is built, then free") | both readings carried: resident = worst case, freed = best case. The two plans disagree and gate 1 should say which | +| Cache growth | `igneum-wt-ca2-coord/docs/plans/counter-asic-2-status.md` 79 (layer 6 option C) | 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12; by the same rule 2 GiB at year 28, 4 GiB at year 60 | +| Budget rule | same file, 17: the whole working set stays under 6 GB on an 8 GB card | my reading: 75% of card memory at every tier. Apple: 50% of unified memory, because macOS, the display and the node share it; that share is my assumption | +| Public claims | `site/index.html` 443, 461; `site/litepaper.html` 560; `docs/evidence.md` | quoted in Table 3. evidence.md has no row on card lifetime | + +Card memory is binary (8 GB = 8,192 MiB). The hot-table row "An 8 GB card at 5090 occupancy" (line 73) counts 8,160 warps; a real 8 GB card has 20 to 24 SMs, so its scratch is about a tenth of that row. Resident warps below are SMs x 48 (NVIDIA Ampere and later), SM counts from memory, approximate; Apple uses the 2,048 warps the Metal harness launches (hot-table.md 66). + +## 2. Table 1: non-dataset working set per tier (MiB) + +Worst = scratch 128 KiB, cache resident. Columns g / y4 / y12 = genesis, year 4, year 12 (the cache doublings). Freed = era-layout's reading, constant over the years. + +| Tier | Card assumed (SMs, approximate) | Warps | Scratch 128 KiB | Scratch 32 KiB | Cache resident, 128 KiB: g / y4 / y12 | Cache resident, 32 KiB: g / y4 / y12 | Cache freed: 128 / 32 KiB | +|---|---|---|---|---|---|---|---| +| 4 GB | GTX 1650 (14 SMs x 32 warps, Turing) | 448 | 56 | 14 | 536 / 792 / 1,304 | 494 / 750 / 1,262 | 280 / 238 | +| 8 GB | RTX 3050 (20) | 960 | 120 | 30 | 600 / 856 / 1,368 | 510 / 766 / 1,278 | 344 / 254 | +| 12 GB | RTX 3060 (28) | 1,344 | 168 | 42 | 648 / 904 / 1,416 | 522 / 778 / 1,290 | 392 / 266 | +| 16 GB | RTX 5060 Ti (36) | 1,728 | 216 | 54 | 696 / 952 / 1,464 | 534 / 790 / 1,302 | 440 / 278 | +| 24 GB | RTX 4090 (128) | 6,144 | 768 | 192 | 1,248 / 1,504 / 2,016 | 672 / 928 / 1,440 | 992 / 416 | +| 32 GB | RTX 5090 (170) | 8,160 | 1,020 | 255 | 1,500 / 1,756 / 2,268 | 735 / 991 / 1,503 | 1,244 / 479 | +| Apple 8 to 64 GB | M-series, harness launch count | 2,048 | 256 | 64 | 736 / 992 / 1,504 | 544 / 800 / 1,312 | 480 / 288 | + +Every row = scratch + 96 (hot table) + 128 (buffers) + cache (256 / 512 / 1,024 when resident). The freed reading still peaks at dataset + cache during the daily build, but that peak is smaller than the resident total whenever hashing pauses for the build, so the freed column is the steady-state set. + +## 3. Table 2: dataset room and the year the dataset outgrows it + +Room = usable memory (75%, Apple 50%) minus Table 1. Worst = 128 KiB scratch, cache resident (room shrinks at years 4, 12, 28, 60). Best = 32 KiB scratch, cache freed. Option (a): the year 2,048 + 512 x y exceeds the room. Option (b): the first step the room cannot hold; the card mines up to that day. + +| Tier | Usable MiB (share) | Room at genesis, worst / best | (a) ends, years, worst / best | (b) ends, year, worst / best | +|---|---|---|---|---| +| 4 GB | 3,072 (75%) | 2,536 / 2,834 | 1.0 / 1.5 | 4 / 4 | +| 8 GB | 6,144 (75%) | 5,544 / 5,890 | 6.3 / 7.5 | 12 / 12 | +| 12 GB | 9,216 (75%) | 8,568 / 8,950 | 12.0 / 13.5 | 12 / 28 | +| 16 GB | 12,288 (75%) | 11,592 / 12,010 | 17.1 / 19.5 | 28 / 28 | +| 24 GB | 18,432 (75%) | 17,184 / 18,016 | 28.0 / 31.2 | 28 / 60 | +| 32 GB | 24,576 (75%) | 23,076 / 24,097 | 37.6 / 43.1 | 60 / 60 | +| Apple 8 GB | 4,096 (50%) | 3,360 / 3,808 | 2.6 / 3.4 | 4 / 4 | +| Apple 16 GB | 8,192 (50%) | 7,456 / 7,904 | 10.1 / 11.4 | 12 / 12 | +| Apple 32 GB | 16,384 (50%) | 15,648 / 16,096 | 25.1 / 27.4 | 28 / 28 | +| Apple 64 GB | 32,768 (50%) | 32,032 / 32,480 | 55.1 / 59.4 | 60 / 60 | + +What the table says per tier: + +| Tier | Reading | +|---|---| +| 4 GB | Mines at genesis with 488 to 786 MiB spare. Under (a) it is out within 1 to 1.5 years. Under (b) it lasts to the year-4 step, as the spec's own remark says (line 442) | +| 8 GB | 6 to 7.5 years under (a). 12 years under (b): "more than a decade" is true only under (b), and only just | +| 12 GB | The year-12 cache doubling (1 GiB resident) is what ends it, under both options, if the cache stays resident. With the cache freed it reaches year 28 under (b). This tier's lifetime is decided by the cache residency question, not by the dataset | +| 16 GB | 17 to 19.5 years under (a), year 28 under (b) | +| 24 GB | Under the resident reading the year-28 cache doubling (2 GiB) ends it the same day under both options. Freed: 31 years or year 60 | +| 32 GB | 38 to 43 years under (a), year 60 under (b). Not a constraint for any plan | +| Apple 8 GB | 2.6 to 3.4 years under (a), year 4 under (b). The base MacBook Air is a short-lived miner | +| Apple 16 GB | 10 to 11.4 years under (a), year 12 under (b): the same shape as an 8 GB card | +| Apple 32 / 64 GB | 25 years and 55 years or more. No constraint | + +Proving is a separate budget (the 15.6 GB peak the 12 GB mine-and-prove question came from); this file covers mining only. + +## 4. Table 3: the public sentences against the numbers + +| Where | Sentence now | What the tables give | Proposed sentence (the project lead decides the wording) | +|---|---|---|---| +| `site/index.html` 443 | Memory: "2 GB, fixed" (RandomX) / "2 GB, growing" (Igneum) | 2 GiB at genesis, plus 0.5 GiB a year on average under either option | "2 GB, growing 0.5 GB a year". The row is right; the rate is the useful addition | +| `site/index.html` 461 | "Any 4 GB card, approximate." | True at genesis (2,584 to 2,834 MiB of a 3,072 MiB budget). Ends at 1 to 1.5 years under (a), year 4 under (b) | "Any 4 GB card at launch, 8 GB for the long run, approximate." | +| `site/litepaper.html` 560 | "a 4 GB card mines for about four years and an 8 GB card for more than a decade, approximate." | 4 GB: 1 to 1.5 years (a) or 4 years (b). 8 GB: 6.3 to 7.5 years (a) or 12 years (b). Both numbers hold only under option (b) | If gate 1 picks (b): "a 4 GB card mines until the first dataset step at year 4, an 8 GB card until the second at year 12 and a 16 GB card until year 28, approximate." If (a): "a 4 GB card mines for about a year, an 8 GB card for about seven and a 16 GB card for about seventeen, approximate." | +| `site/litepaper.html` 560 | "12 GB or more proves full shards." | Not a lifetime claim; left as is. For mining, 12 GB lasts 12 years with the cache resident, year 28 with it freed under (b) | No change from this file | +| `docs/evidence.md` | No row on card lifetime | The litepaper sentence is a public claim with no row | Add a row, label "designed", sources: spec 1.13.3 and this file; status moves to "tested" once a 4 GB and an 8 GB card run the genesis working set under the cap | + +## 5. Reading + +- The two index-mapping options end on the same day where a cache doubling takes the last of the room. With the cache resident that is the 12 GB tier at year 12 and the 24 GB tier at year 28 (Table 2, worst column). Under option (b) every tier ends on a step day by construction, so a tier ends on the same day under both options exactly when option (a) also ends it on a doubling day. +- Everywhere else option (b) is kinder: 4 GB gains about 2.5 years, 8 GB about 5, 16 GB about 10. The site and litepaper numbers are option (b) numbers. If gate 1 picks (a), both public sentences are wrong today by 2.5 to 5 years. +- The cache residency disagreement (hot-table.md 70 against era-layout.md 93) decides the 12 GB tier's lifetime (12 against 28 years) and nothing else. It should be settled at gate 1 beside the mapping choice. +- The 75% rule is my generalisation of "under 6 GB on an 8 GB card"; at 4 GB it leaves 1 GB for the driver and the display, which a headless rig would not need. A 4 GB card on a bare Linux rig might hold out to year 2 under (a). Not measured. From 770145f7becf1afc370fde4763762194423d693b Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:31:48 +0000 Subject: [PATCH 024/131] Counter ASIC 2.0: card lifetime merged, cache freed after the build, growth mapping (b) recommended, public lines listed --- docs/plans/counter-asic-2-rollout.md | 9 +++++++++ docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 13 insertions(+) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index b51f04861..a25fad592 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -77,6 +77,15 @@ If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-st 0.3.11 through the shipper's pipeline (`tools/ship-app.mjs`, the plan shape of `docs/plans/release-0.3.10.md`). The 0.3.10 shipper finishes 0.3.10 first, then takes the 0.3.11 tree, or the coordinator ships 0.3.11 with the same runbook if the shipper has stopped. The publish order is that of `docs/plans/finality-v3-devnet-publish.md`: the override field, the digest handshake, hand nodes and the seed first, then the manifest with `--activation-height` and the deadline note, then the apps, then the digest sweep, then the HiveOS package republish. +### 6c. Card lifetime: cache residency and the growth mapping (from `docs/analysis/card-lifetime-2026-10-05.md`, merged) + +| Decision | Choice | The number | +|---|---|---| +| Cache residency on the GPU | DECIDED (5 October 2026, delegated): the cache is FREED after the daily dataset build; the hash reads the dataset and the hot table only, never the cache. hot-table.md's resident reading is corrected to this | The daily rebuild is the only cost: cache fill 0.67 ms and dataset build 13.4 ms on the RTX 5090 at x1 (bench-log, 3 October 2026), 2 ms and 13 to 30 ms on the M5 Max; at mixer x4 the build is about 54 ms (measurement owed on ca2-mixer). Freeing it moves the 12 GB tier from year 12 to year 28 under mapping (b) | +| Dataset growth mapping (a: continuous 2 + 0.5 GiB a year with a multiply-shift index; b: power-of-two steps at years 4, 12, 28, 60 with `AND MASK`) | RECOMMENDED for the project lead: (b). It keeps `AND MASK` and every vector's size, it is what option C's "doubles when the dataset doubles" already assumes, and it is the only mapping under which the litepaper's "4 GB about four years, 8 GB more than a decade" is true (under (a) a 4 GB card is out within 1 to 1.5 years, an 8 GB card at 6 to 7.5 years) | card-lifetime table 2: 4 GB out at year 4 (b) or 1.0 to 1.5 (a); 8 GB year 12 or 6.3 to 7.5; 12 GB year 28 freed or 12 resident; 24 GB year 60 freed | + +Public lines to fix on the integration branch (card-lifetime table 3): `site/index.html` "2 GB, growing" gains the rate ("2 GB at genesis, doubling at years 4, 12 and 28"); "Any 4 GB card" becomes "any 4 GB card at launch, 8 GB from year 4"; the litepaper's "4 GB about four years, 8 GB more than a decade" stays with mapping (b) and gains "under the step schedule"; `docs/evidence.md` gains a row for the card-lifetime claim labelled designed. hot-table.md line 73's 8 GB row is corrected (it counted a 5090's 8,160 warps; a real 8 GB card has 20 to 24 SMs). + ## 7a. A dated constraint from the consequences review (C1) The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 1cef880d8..6434067a0 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -147,3 +147,7 @@ Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout C1 (the fee switch H = 210,000 at about 19:50Z on 6 October; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before 19:50Z. The proving agent recommends (b) unless (a) is certain; the check decides it. The attempt-0 packfile.h bug (the era agent's find) is the one that took the fleet down at 18:23Z (epoch 34, attempt 1); it is already fixed attempt-aware on pack-loop af983a7 and merged into the 0.3.10 tree. Rule for 0.3.11: one derivation, one test set: the era and cache agents build their workers on that packfile.h and keep only tests that add a case; the host.cu / host.c mh_word derivation fix (new, the era agent's) stays with its own test. + +## 22:50 card lifetime merged; cache freed after the build; growth mapping (b) recommended + +`docs/analysis/card-lifetime-2026-10-05.md` (branch card-lifetime 1fecfe2) merged into ca2-coord. Decided (delegated): the GPU frees the 256 / 512 / 1,024 MiB cache after the daily dataset build (the hash never reads it; the rebuild costs 0.67 ms fill + 13.4 ms build on the 5090 at x1, about 54 ms at x4, owed); hot-table.md's resident reading is corrected, era-layout.md's freed reading stands. Recommended for the project lead: growth mapping (b), power-of-two steps at years 4, 12, 28, 60 with AND MASK, the only mapping under which the litepaper's "4 GB about four years, 8 GB more than a decade" holds (4 GB: year 4 under (b), 1.0 to 1.5 years under (a); 8 GB: year 12 or 6.3 to 7.5; 12 GB: year 28 with the cache freed; 24 GB: year 60). Public lines and the evidence row go on the integration branch (rollout plan 6c); hot-table.md line 73 (the 8 GB row counted a 5090's warps) goes to the cache agent. From 3f33c6fd009d4dfde8c0fefe1537ffb000745ae6 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:32:53 +0000 Subject: [PATCH 025/131] Counter ASIC 2.0 status 23:00: layer 9 agent, PC queues --- docs/plans/counter-asic-2-status.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 6434067a0..f89836921 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -151,3 +151,9 @@ The attempt-0 packfile.h bug (the era agent's find) is the one that took the fle ## 22:50 card lifetime merged; cache freed after the build; growth mapping (b) recommended `docs/analysis/card-lifetime-2026-10-05.md` (branch card-lifetime 1fecfe2) merged into ca2-coord. Decided (delegated): the GPU frees the 256 / 512 / 1,024 MiB cache after the daily dataset build (the hash never reads it; the rebuild costs 0.67 ms fill + 13.4 ms build on the 5090 at x1, about 54 ms at x4, owed); hot-table.md's resident reading is corrected, era-layout.md's freed reading stands. Recommended for the project lead: growth mapping (b), power-of-two steps at years 4, 12, 28, 60 with AND MASK, the only mapping under which the litepaper's "4 GB about four years, 8 GB more than a decade" holds (4 GB: year 4 under (b), 1.0 to 1.5 years under (a); 8 GB: year 12 or 6.3 to 7.5; 12 GB: year 28 with the cache freed; 24 GB: year 60). Public lines and the evidence row go on the integration branch (rollout plan 6c); hot-table.md line 73 (the 8 GB row counted a 5090's warps) goes to the cache agent. + +## 23:00 layer 9 agent started; the dot4 probe is on PC 1 + +ca2-epoch (a32a3ece66c02417a): the epoch length as an era parameter (600 to 7,200 DAA s, base 3,600; draw or 90% signal; the VDF rule; the difficulty-window constraint; the FPGA threat with citations), Mac compile-ahead measured now, the 5090 and 9070 XT compile times cited from the bench log, docs/plans/epoch-length.md. The dot4 probe job is running on PC 1 (ca2-analysis tip ee42d7c carries the playbook; the 5090 confirmation is the agent's watch). PC 1 queue after it: the era six-pack job, then the hot-table added-form job. PC 2: the proving agent's, then the ca2 node suites. + +Agents running: ca2-era, ca2-cache, ca2-node, ca2-mixer, ca2-epoch; ca2-analysis watching its PC job. Done: ca2-soundness. From 4ce31bbbe021a46426018df575e805c0266aa461 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:38:08 +0000 Subject: [PATCH 026/131] Counter ASIC 2.0 status 23:10: layer 7 PC rows, PC 1 free --- docs/plans/counter-asic-2-status.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index f89836921..545de7048 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -157,3 +157,18 @@ The attempt-0 packfile.h bug (the era agent's find) is the one that took the fle ca2-epoch (a32a3ece66c02417a): the epoch length as an era parameter (600 to 7,200 DAA s, base 3,600; draw or 90% signal; the VDF rule; the difficulty-window constraint; the FPGA threat with citations), Mac compile-ahead measured now, the 5090 and 9070 XT compile times cited from the bench log, docs/plans/epoch-length.md. The dot4 probe job is running on PC 1 (ca2-analysis tip ee42d7c carries the playbook; the 5090 confirmation is the agent's watch). PC 1 queue after it: the era six-pack job, then the hot-table added-form job. PC 2: the proving agent's, then the ca2 node suites. Agents running: ca2-era, ca2-cache, ca2-node, ca2-mixer, ca2-epoch; ca2-analysis watching its PC job. Done: ca2-soundness. + +## 23:10 layer 7 complete on all three cards (ca2-analysis ee42d7c); PC 1 free + +dot4 probe on PC 1 (jobs fetch-dot4-20261005, run-dot4-20261005, exit 0, 101 s; both cards restored and mining; 5090 SM clock 2,505 MHz before and after): + +| Device | ALU chain, G steps/s | signed dot4 emulation, G dot4/s | dot4 instruction, G dot4/s | emulation vs instruction | +|---|---|---|---|---| +| RTX 5090 (NVIDIA OpenCL 3.0, driver 617.14) | 8,753.5 | 1,239.1 (7.1x an ALU step) | 7,453.6 (inline PTX dp4a.s32.s32, 1.17x) | 6.0x | +| RX 9070 XT gfx1201 (AMD-APP 3683.0) | 701.4 | 480.8 (1.46x) | 664.3 (__builtin_amdgcn_sudot4, 1.06x) | 1.38x | +| gfx1036 (RDNA 2 iGPU) | 40.6 | 15.8 (2.6x) | sudot4 does not build (needs dot8-insts) | | +| M5 Max (Metal) | 879.8 | 188.2 (4.7x); unsigned 548.2 (1.6x) | none exists | | + +All bit-exact against the CPU reference. One dp4a costs about one ALU step on NVIDIA and AMD; the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on dot4 (the family does not widen the AMD gap), 10x the M5 Max on the chain and 13.6x on dot4 (Apple's emulation widens its gap 1.4x). At W_new 4 the family adds about 21 ops per hash per lane; hash-rate losses expected under 5% on every card (to be measured with the family live). Owed: the CUDA __dp4a cross-check (needs nvcc), Metal 4 matmul2d int8 on the M5, sdot4 on RDNA 2. cl_khr_integer_dot_product is listed by no driver we own. + +PC 1 is free: the era six-pack job goes next when its package arrives, then the hot-table added-form job. From 9dcb9697cfa9ae1f20a40700af333d916908e3e9 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:39:03 +0000 Subject: [PATCH 027/131] Counter ASIC 2.0 status 23:20: the PC 1 queue --- docs/plans/counter-asic-2-status.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 545de7048..151349239 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -172,3 +172,18 @@ dot4 probe on PC 1 (jobs fetch-dot4-20261005, run-dot4-20261005, exit 0, 101 s; All bit-exact against the CPU reference. One dp4a costs about one ALU step on NVIDIA and AMD; the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on dot4 (the family does not widen the AMD gap), 10x the M5 Max on the chain and 13.6x on dot4 (Apple's emulation widens its gap 1.4x). At W_new 4 the family adds about 21 ops per hash per lane; hash-rate losses expected under 5% on every card (to be measured with the family live). Owed: the CUDA __dp4a cross-check (needs nvcc), Metal 4 matmul2d int8 on the M5, sdot4 on RDNA 2. cl_khr_integer_dot_product is listed by no driver we own. PC 1 is free: the era six-pack job goes next when its package arrives, then the hot-table added-form job. + +## 23:20 PC 1 scheduler (the coordinator's role from now): the queue + +Rule: one PC 1 job at a time; an agent asks by message before publishing, gets a "go PC 1" from this coordinator, and reports when its RESULT lines are in and both cards are restored; a hash-rate or power number taken while another job holds a card is not a number. The CPU-only job runs only in a slot where no measurement overlaps it. + +| # | Job | Agent | Cards | Length | State | +|---|---|---|---|---|---| +| 1 | Era six-pack (layers 4 and 8), 5090 and 9070 XT | ca2-era a452664c512c73b9b | one card at a time | about 15 to 20 min | waiting for the package (packs on the pack-loop packfile.h) | +| 2 | Hot table, added form, probe 32/64/96 MiB plus packs | ca2-cache a5271cf269757b118 | one card at a time | about 15 min | waiting for the re-measured Mac rows and the rebuilt zip | +| 3 | Reproducible benchmark run | a0b9f574775ef1693 | one card at a time, 120 s per card | about 5 min | queued | +| 4 | AMD sweep on the 9070 XT (core clock and power steps) | a01dcb34ae16d867c | 9070 XT only; the 5090 keeps mining | about 20 min | queued | +| 5 | Ember Tune end to end, both cards | a855dcc4bd05e0615 | both | to be stated | queued after the AMD sweep | +| 6 | AMD-proving CPU fallback (CPU-only SP1 run, both cards mining) | a39db54d4de4af51e | none; loads the CPU | to be stated | last, or in a gap where no measurement runs for its whole length | + +If job 1's package is more than 15 minutes away when job 2 is ready, job 2 goes first; the short job 3 fills any gap of under 10 minutes between packages. From 7f3201f53358168ef4c400e490bc4dbf55554187 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:39:44 +0000 Subject: [PATCH 028/131] Counter ASIC 2.0: rig-install follow-ups on the next-cut list, the CPU fixture in the PC 1 gap, no more agents --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index a25fad592..d36a93004 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -94,7 +94,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2): fork branch proving-v1 (from a24ab01a: segment records, the pool split, five RPCs, p2p message 72, parameters proving_v1_activation_daa, segment_blocks, unproven_daa, aggregator_share_bps, in the digest only once set) and its app branch, rebased onto the 0.3.10 tips. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. -Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install, pool-v0 (its own service, no app change), repro-bench; fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. +Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. ## 9. After the publish: Counter ASIC 3.0 diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 151349239..65f20f2f7 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -187,3 +187,5 @@ Rule: one PC 1 job at a time; an agent asks by message before publishing, gets a | 6 | AMD-proving CPU fallback (CPU-only SP1 run, both cards mining) | a39db54d4de4af51e | none; loads the CPU | to be stated | last, or in a gap where no measurement runs for its whole length | If job 1's package is more than 15 minutes away when job 2 is ready, job 2 goes first; the short job 3 fills any gap of under 10 minutes between packages. + +23:30. No more agents are spawned tonight (the project lead: no unnecessary credits); the running ones finish. PC 1 queue change: the AMD-proving CPU fallback's small fixture (block-56-transfers-3shards, minutes) runs NOW in the gap before the era package; its S_p shard (block-338-shard1, up to 30 min of every core) stays job 6, last. Next-cut list gains the rig installer's two follow-ups (the Linux manifest entry, the Linux prover build), tied to whichever release carries the proving half (rollout plan 8a). From 33713368b68c8743bb92405a857f8dcab9595257 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:41:03 +0000 Subject: [PATCH 029/131] Counter ASIC 2.0 status: heading times corrected to the commit clock (UTC) --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 42 ++++++++++++++-------------- 2 files changed, 22 insertions(+), 22 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index d36a93004..a944248d8 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 21:15 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (21:50): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ### 6b. The user tiers (the consequences rule) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 65f20f2f7..2fd7f81b4 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -1,6 +1,6 @@ # Counter ASIC 2.0: status -Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rewritten every 45 minutes while the work runs. Times UTC, 5 October 2026 (night). Base for every ca2 branch: `readwidth` at 019b014 (the LoadClass flag, fold_words, the scratch op, the three emitters, 20 packs). +Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rewritten every 45 minutes while the work runs. Times UTC, 5 October 2026 (night). Heading times before 20:40 were corrected at 20:42 from the commit clock (the coordinator had written them from a guessed clock, up to 2 h 40 min ahead); every entry's true time is its commit's author time in UTC. Base for every ca2 branch: `readwidth` at 019b014 (the LoadClass flag, fold_words, the scratch op, the three emitters, 20 packs). ## 19:55 first status (the 19:50 start was cut off by exhausted credits at about 19:58 before any sub-agent work landed; respawned at 19:55 on the restart) @@ -35,7 +35,7 @@ Finding that bears on layer 3 (readwidth, Metal, M5 Max, capped sizes): the scra Reading (readwidth agent): 4,096 warps x 32 KB = 128 MB sits in the chip's caches. Consequence for the decision: a scratch that fits a GPU's cache fits a chip's SRAM at the same size, so at the capped size the writes cost everyone a cache-bound op in place of a latency-bound load. Passed to the soundness agent: what size would make the writes cost DRAM latency, whether that fits the 6 GB cap, and whether the RMW share should be added to the 16 dataset loads rather than taken from them. -## 20:15 mandate: v3 on the devnet tonight, by the project lead's rules +## 20:08 mandate: v3 on the devnet tonight, by the project lead's rules the project lead has gone to bed and delegated the three decisions for the DEVNET only (not the public testnet): width, per-load mix and scratch share by the rules now written in `docs/plans/counter-asic-2-rollout.md` section 6, the activation height = devnet tip + 14,400 at publish (checked >= 10,800), published the way finality v3 was. Six gates before any publish (rollout section 7): bit-exact v3 on the three cards; the CPU verifier exact on 1,000 random GPU hashes per card; the soundness suite green with the new scratch tests; the fast-time 3-node network mining across a v3 activation with 0 rejected blocks and 0 forks; Windows and Mac workers from one commit; the node change on a fork from the 0.3.10 tip (21d4c73c) with suites green on PC 2. Release 0.3.11 through the shipper's pipeline; the 0.3.10 shipper (ae892a8b0f78fe31c) has been asked for its state and the handoff. If a gate fails: stop, write why here, do not publish. @@ -43,7 +43,7 @@ Node build note for the integration: the node links `igneum-pow` by path (`../.. After the publish: Counter ASIC 3.0 from the ASIC-history agent's ranked additions (a202a09dcd24ba1d3), as class v4 behind its own activation, same gates, `docs/plans/counter-asic-3.md`. -## 20:25 the shipper's answer, the node fork convention +## 20:08 the shipper's answer, the node fork convention 0.3.10 (shipper ae892a8b0f78fe31c): staged and blocked on GitHub's Actions incident (run 37365130137 queued since 19:42:43Z under a re-dispatching watcher); nothing on the network has moved, the live manifest is still 0.3.9. Once CI is green: ship (5 min), update-now (apps restart 1 to 10 min later), hand nodes and seed (5 min), digest sweep; 0.3.10 finished about 30 min after green. The app version per machine on the console's cards is the restart signal for re-running any straddling measurement. HiveOS is published by the ship's --public step, nothing separate. @@ -51,7 +51,7 @@ Node fork for v3: base on COMMIT 21d4c73c (release-0.3.10 in vendor/igneum-node 0.3.11 shipper: the coordinator assigns it (the 0.3.10 shipper stops at its report). Inputs the ship needs: the fork commit with its PC 2 suite results recorded, the main tip, the override object with every switch (the four live fields plus program_class_v3_activation_daa), the expected digest read on a 22-s scratch node, the deadline note ("program class v3"), the activation height, the one-line changelog, and whether the pinned proving guest changes (it does not: the prover has no igneum-pow dependency; confirmed by grep of proving/igneum-prove Cargo files). -## 20:35 readwidth round 2 on the PCs; the width arithmetic under the project lead's rule +## 20:10 readwidth round 2 on the PCs; the width arithmetic under the project lead's rule Round 1 of the readwidth PC jobs refused every pack (the workers demand a 32-byte chain seed; the experiment packs carried string seeds); fixed at readwidth 1ea7a52 (packfile.h), republished 20:09:19Z as `run-readwidth-5090-20261005c` (PC 2, about 6 min) and `run-readwidth-9070-20261005c` (PC 1, about 10 min). The era and cache agents were told to rebase onto 1ea7a52 and to prove their packs load on the Mac OpenCL host before any PC job. Every ca2 branch now bases on 1ea7a52. @@ -64,9 +64,9 @@ Probe ceilings from round 1 (dependent reads per second at 1024 MiB, device time the project lead's width rule applied to the ceilings alone (the measured v3 rates will replace this when the table lands): a 128-load hash at 64 B reads 8,192 B; at the 5090's 64 B ceiling that is 71 MH/s and 584 GB/s, 37% of its stream bandwidth, over the one-third margin the rule sets; at 16 B it is 2,048 B per hash, 141 MH/s and 288 GB/s, 18%, inside the margin; 4 B is 9%. On the 9070 XT every width costs the same line fetch (2.4 G/s), so 16 B is where AMD gains 4x the bytes per hash at no cost and the 5090 stays latency-bound with margin. Provisional width under the rule: 16 B (w16), pending the measured rates and the latency-bound share per card. -Added deliverable (20:40): the public description in four levels, `docs/plans/counter-asic-2-public.md` (aec53bb): levels 1 and 2 are written as copy; level 3 carries the bench table with `[owed]` markers for every number not yet measured; level 4 lists the documents. The integration branch applies levels 1 to 3 to `site/index.html`, `site/litepaper.html` and `site/bench.html` with the final numbers. +Added deliverable (20:11): the public description in four levels, `docs/plans/counter-asic-2-public.md` (aec53bb): levels 1 and 2 are written as copy; level 3 carries the bench table with `[owed]` markers for every number not yet measured; level 4 lists the documents. The integration branch applies levels 1 to 3 to `site/index.html`, `site/litepaper.html` and `site/bench.html` with the final numbers. -## 20:55 layers 6 and 7 landed; the node-fork agent started; 0.3.11 scope +## 20:15 layers 6 and 7 landed; the node-fork agent started; 0.3.11 scope Branch ca2-analysis (5d5ba15, f59708d). Layer 6: no cache growth rule exists in the spec; cited bit cells N7 0.027, N5 / N3E / Intel 18A 0.021, N3B 0.0199, N2 0.0175 um^2, array factor 0.70; the 256 MiB mirror is 83 / 64 / 54 mm^2 at N7 / N5 / N2 (74 at N2 with a 96 MB hot table), $13 to $26 of silicon per good die (approximate). The mirror was never unaffordable; the cache's job is to stay above GPU L2 (5090 96 MB, GB202 128 MB). Recommendation C for the project lead (gate 1): cache doubles when the dataset doubles. Layer 7: Metal 4 matmul2d has int8 x int8 -> int32, so a unit-level mm8 tile is native on all three vendors; per-lane dot4 is emulation on Apple (M5 Max: 548 G unsigned dot4/s against an 880 G ALU chain, 1.6x; signed 4.7x). Reserve R1 = mm8, W_new 4, unlock era 4 or 90% signal. Owed: the 5090 and 9070 XT dot4 probe (job prepared: relay/playbooks/dot4-probe.ps1, exe sha256 5adaeb1a...6416f4; publishes when a PC frees). @@ -74,21 +74,21 @@ Node-fork agent a3f505a9d981300cd started 20:50: ca2-v3 (igneum-pow seam: Progra 0.3.11 scope (coordinator, 20:52): carries program_class_v3 and proving v1 together (one override object, one digest, one publish; each half under the same gates); the ready half ships as 0.3.11 and the other as 0.3.12 if one lags. Next-cut list recorded in the rollout plan section 8a. -## 21:05 decisions recorded: layer 6 option C, layer 7 R1 = mm8; the chip headline +## 20:16 decisions recorded: layer 6 option C, layer 7 R1 = mm8; the chip headline Layer 6 DECIDED (delegated; the project lead confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles (256 MiB genesis, 512 MiB year 4, 1 GiB year 12); one-core fill 0.2 / 0.4 / 0.8 s, under 1 s at every step. Layer 7 DECIDED: reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal. The chip model's headline now names the on-die-cache recompute chip (54 to 83 mm^2) as a row per variant; the scratch share is chosen as the smallest share at which that chip's gain falls under 1.5x, else said plainly and the public "under 2x" claim qualified. The soundness agent carries that table (its question 2) with M16's mixer multiplier beside it. -## 21:15 correction to the SRAM mirror figures (chip-economics research, cluster D) +## 20:18 correction to the SRAM mirror figures (chip-economics research, cluster D) Bit cell x 0.70 understates real die area. Shipped cache dies: AMD 3D V-Cache 64 MB on 41 mm^2 at 7 nm (1.56 MB/mm^2, Tom's Hardware, Hot Chips August 2021); Graphcore GC200 900 MB on 823 mm^2 with compute (1.09 MB/mm^2); Groq TSP 220 MB on 725 mm^2 at 14 nm (0.30 MB/mm^2); TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30% assist overhead (SemiAnalysis, December 2022). A 256 MiB mirror is about 165 mm^2 at 7 nm on the densest shipped cache-only die and about 130 mm^2 at N5/N3E, not 54 to 83 mm^2; cost per die 2 to 3x the earlier figure; the conclusion (affordable for a funded chip) stands. The analysis agent is redoing the table with both columns; the soundness agent carries the corrected density into the chip row. Latency citations behind the latency-bound rule, to be added: DRAM row cycle 40 to 48 ns across DDR4, GDDR5, HBM2 (Li, Reddy, Jacob, MEMSYS 2018); latency 1.3x in two decades against bandwidth 20x (Chang 2017); no shipped mining chip used HBM or stacked memory. -## 21:25 proving v1 state for 0.3.11; PC 2 occupancy +## 20:19 proving v1 state for 0.3.11; PC 2 occupancy Proving v1 (acd4f36bc2c07a4e2): fork proving-v1 b177718e on a24ab01a (told to rebase onto commit 21d4c73c now), app proving-v1 79bc820 on a93199a. Override fields proving_v1_activation_daa (tip + 14,400 at publish), proving_v1_segment_blocks 4, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Harness: `tools/proving-v1/net.mjs --secs 1500` PASSED (21 checks) in 197.3 s on b177718e; rerun owed on the final tree. The pinned guests do not change. Shared files with ca2-v3-node: params.rs, daemon.rs, igneum/miner/src/main.rs, override-60x.json; both agents keep separable hunks. PC 2 is held by the proving agent's memsweep-pc2-pv1 (about 20 min, miners stopped) and a second run (about 10 min). Queue after it: the ca2 node suites, then the readwidth, era, cache and dot4 measurement jobs. PC 1 is held by readwidth's run-readwidth-9070-20261005c until it reports. -## 21:35 sram-mirror.md revision 2 (ca2-analysis) +## 20:21 sram-mirror.md revision 2 (ca2-analysis) Two columns, headline = shipped-product density (AMD V-Cache 41 mm^2 per 64 MiB at N7, scaled by the bit-cell ratio), lower bound = bit cell x 0.70. mm^2 and $ per good die (D0 0.1 per cm^2, wafer prices approximate), headline / lower bound: @@ -101,7 +101,7 @@ Two columns, headline = shipped-product density (AMD V-Cache 41 mm^2 per 64 MiB Year 10 at the 6% per year trend: 59 mm^2 for the flat cache (82 with the hot table), 8% of a 750 mm^2 die. One reticle holds 1.3 GiB (N7) to 1.9 GiB (N2); mirror share of a 750 mm^2 die at year 0: 14% (22% with the hot table), inside M16's 13 to 40% band. Recommendation unchanged: C. Latency section added (MEMSYS 2018, Chang 2017, the mining-chip memory-type note), marked as research the agent did not re-read tonight apart from the V-Cache figure. -## 21:50 layer 3 soundness landed: the scratch does not move the chip; scratch share decided 0 +## 20:23 layer 3 soundness landed: the scratch does not move the chip; scratch share decided 0 ca2-soundness (0d8f745 tests and trace hook, a465881 doc and bench-log). The on-die-cache recompute chip (N5 headline 128 mm^2, $46) at 50 T op/s: 333 MH/s against the 5090's measured 139.7, 2.4x; at 12.5 / 25 / 50% RMW replaced, chip 381 / 443 / 661 against 5090 projected 160 / 186 / 279, 2.4x each; added, 2.4x or more; 32 or 128 KB alike. The chip keeps the scratch implicitly in 80 to 320 B per lane (the verifier resets it per unit), needs about 530 units in flight, dense scratch 6.2 / 3.1 mm^2 at N5. Under the project lead's rule the scratch share is 0: layer 3 is NOT adopted into v3; the public "under 2x" claim is qualified (public copy level 3 rewritten). The measured lever is M16's mixer multiplier (x2 1.2x at 0.8 to 2.4 ms verify; x4 0.6x at 1.6 to 4.8 ms; 3.6x and 1.8x with a 3x fixed-function factor); whether x4 enters v3 tonight is asked of the coordinator; default: Counter ASIC 3.0. @@ -109,13 +109,13 @@ Soundness results (Metal, M5 Max): 28/28 edge launches, 200/200 fuzz packs (91 s Gate G3 note: the scratch tests (igneum-pow/tests/scratch.rs) join the v3 suite even though the class carries no scratch, parametric over the class; they guard the v2 path's scratch-free invariant at zero cost. -## 22:00 decided: M16 mixer x4 into v3; agent ca2-mixer started +## 20:24 decided: M16 mixer x4 into v3; agent ca2-mixer started Coordinator's decision under the project lead's delegation (recorded in the rollout plan section 6a): the mixer multiplier x4 and the cache growth rule (option C) enter class v3 behind the same activation; layer 3 stays out at scratch share 0, its soundness document and pack-contract tests kept. Agent af345b1e2c541ffbb (branch ca2-mixer) implements `mixer_mult` as a class parameter (m mixer applications per round, the 8 dependent reads unchanged), the `cache_log2_words(day)` schedule (doublings at years 4 and 12 with the dataset stepping to the next power of two), re-cuts the v3 dataset vectors, re-runs the soundness suite, measures the verifier (v2 0.604 ms per warp; v3 expected 1.6 to 4.8 ms) and the 1 GiB build on the Mac, prepares the 5090 and 9070 XT build-time job, and writes docs/analysis/chip-model-v3.md with the combined headline row (fixed-function factor included). The claim on the site reads "under 2x" only if that row does; else qualified, with the mixer x8 and the hot table named as the next levers. Agents now: ca2-era (a452664c512c73b9b), ca2-cache (a5271cf269757b118), ca2-node (a3f505a9d981300cd), ca2-mixer (af345b1e2c541ffbb). Done: ca2-analysis, ca2-soundness. Waiting: the readwidth PC table; PC 2 (proving memsweep runs) and PC 1 (readwidth 9070 round). -## 22:15 the readwidth table landed; layers 1, 2, 3 decided; layer 5 measured on the Mac and redesigned +## 20:27 the readwidth table landed; layers 1, 2, 3 decided; layer 5 measured on the Mac and redesigned Readwidth e752fc7 (`docs/plans/read-width.md`), bit-exact on Metal, Apple OpenCL, the 5090 (NVRTC) and the 9070 XT, both PCs released. MH/s (latency-bound share): @@ -134,31 +134,31 @@ Decisions (the project lead's rules, delegated): layer 1 keep v2 (w16 passes the Layer 5 (ca2-cache 53ef59f, 011cc0a, 86726cd, 65bc7a7; `docs/plans/hot-table.md`): five packs bit-exact on Metal and Apple OpenCL (96/96 each). M5 Max rates against v2 27.68: hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k2 x1.00, hot64k8 x1.71; verifier 0.344 to 0.560 ms against 0.626; hot fill per epoch 24 / 46 / 73 ms on one core, 0.07 / 0.15 / 0.22 ms on the GPU; Apple OpenCL probe 32 / 64 / 96 / 1024 MiB 21.7 / 12.8 / 12.3 / 3.50 G loads/s. Redesign ordered: hot loads ADDED beside the 16 dataset loads (the replaced form lets the on-die-cache chip skip item derivations and worsens the gain); the agent re-measures the added form and rebuilds the PC job. PC 1 is given to the dot4 probe (under 15 min), then to the era agent, then the hot-table job; PC 2 stays the proving agent's. -## 22:30 layer 9 added; the era draw passes Mac bit-exactness; two class bugs in the harnesses; C1 +## 20:27 layer 9 added; the era draw passes Mac bit-exactness; two class bugs in the harnesses; C1 Layer 9 (the project lead: faster program changes): the epoch length becomes an era parameter in the genesis reserve, 1 hour at launch, 10 minutes to 2 hours by draw or 90% signal, reserve-only tonight; an agent (ca2-epoch) designs it beside layers 4 and 8 and measures the compile-ahead cost per card at a 10-minute epoch, the seed-path consequence and the FPGA threat it answers; one row in the rollout plan section 6, one in the level 3 numbers. Spawns when the dot4 probe frees its slot. ca2-era mid-way (a452664c512c73b9b): era draw behind LoadClass::era (EraParams beside mix, load_slots, scratch, scratch_kb; verify::load_index; memhard::Layout; one load form in the three emitters; --era, --era-widths), 54 crate tests green, pinned packs byte-identical; six era packs bit-exact on Metal (packbench 6/6), Apple OpenCL (6/6, same fingerprints) and the CUDA emulation (6/6); re-exporting with the width pinned at 4 B and rebasing onto e752fc7; PC job in about 20 minutes. Two class bugs found and fixed on its branch, both outside its layer: (1) proto-cuda/nvrtc/packfile.h re-derived the seed words as attempt 0 only, so ANY pack with IGNEUM_PROGRAM_ATTEMPT >= 1 (5.14% of chain epochs under v2) is refused by the one-click workers with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", the error PC 1 logged on 5 October and attributed to the export race; fixed attempt-aware, with a tampered-attempt refusal test. This fix must ship in the 0.3.11 workers whatever else does. (2) proto-cuda/host.cu and proto-opencl/host.c derived dataset words on the host as mh_item(w >> 4)[w AND 15] instead of the pack's mh_word; fixed. -Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (21:50, 22:00 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.5x the electricity per hash (0.089 against 0.398 MH/W), near parity per pound; the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. +Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (20:23, 20:24 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.5x the electricity per hash (0.089 against 0.398 MH/W), near parity per pound; the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. -## 22:40 C1 decided for the morning; one packfile.h fix for 0.3.11 +## 20:29 C1 decided for the morning; one packfile.h fix for 0.3.11 C1 (the fee switch H = 210,000 at about 19:50Z on 6 October; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before 19:50Z. The proving agent recommends (b) unless (a) is certain; the check decides it. The attempt-0 packfile.h bug (the era agent's find) is the one that took the fleet down at 18:23Z (epoch 34, attempt 1); it is already fixed attempt-aware on pack-loop af983a7 and merged into the 0.3.10 tree. Rule for 0.3.11: one derivation, one test set: the era and cache agents build their workers on that packfile.h and keep only tests that add a case; the host.cu / host.c mh_word derivation fix (new, the era agent's) stays with its own test. -## 22:50 card lifetime merged; cache freed after the build; growth mapping (b) recommended +## 20:31 card lifetime merged; cache freed after the build; growth mapping (b) recommended `docs/analysis/card-lifetime-2026-10-05.md` (branch card-lifetime 1fecfe2) merged into ca2-coord. Decided (delegated): the GPU frees the 256 / 512 / 1,024 MiB cache after the daily dataset build (the hash never reads it; the rebuild costs 0.67 ms fill + 13.4 ms build on the 5090 at x1, about 54 ms at x4, owed); hot-table.md's resident reading is corrected, era-layout.md's freed reading stands. Recommended for the project lead: growth mapping (b), power-of-two steps at years 4, 12, 28, 60 with AND MASK, the only mapping under which the litepaper's "4 GB about four years, 8 GB more than a decade" holds (4 GB: year 4 under (b), 1.0 to 1.5 years under (a); 8 GB: year 12 or 6.3 to 7.5; 12 GB: year 28 with the cache freed; 24 GB: year 60). Public lines and the evidence row go on the integration branch (rollout plan 6c); hot-table.md line 73 (the 8 GB row counted a 5090's warps) goes to the cache agent. -## 23:00 layer 9 agent started; the dot4 probe is on PC 1 +## 20:32 layer 9 agent started; the dot4 probe is on PC 1 ca2-epoch (a32a3ece66c02417a): the epoch length as an era parameter (600 to 7,200 DAA s, base 3,600; draw or 90% signal; the VDF rule; the difficulty-window constraint; the FPGA threat with citations), Mac compile-ahead measured now, the 5090 and 9070 XT compile times cited from the bench log, docs/plans/epoch-length.md. The dot4 probe job is running on PC 1 (ca2-analysis tip ee42d7c carries the playbook; the 5090 confirmation is the agent's watch). PC 1 queue after it: the era six-pack job, then the hot-table added-form job. PC 2: the proving agent's, then the ca2 node suites. Agents running: ca2-era, ca2-cache, ca2-node, ca2-mixer, ca2-epoch; ca2-analysis watching its PC job. Done: ca2-soundness. -## 23:10 layer 7 complete on all three cards (ca2-analysis ee42d7c); PC 1 free +## 20:38 layer 7 complete on all three cards (ca2-analysis ee42d7c); PC 1 free dot4 probe on PC 1 (jobs fetch-dot4-20261005, run-dot4-20261005, exit 0, 101 s; both cards restored and mining; 5090 SM clock 2,505 MHz before and after): @@ -173,7 +173,7 @@ All bit-exact against the CPU reference. One dp4a costs about one ALU step on NV PC 1 is free: the era six-pack job goes next when its package arrives, then the hot-table added-form job. -## 23:20 PC 1 scheduler (the coordinator's role from now): the queue +## 20:39 PC 1 scheduler (the coordinator's role from now): the queue Rule: one PC 1 job at a time; an agent asks by message before publishing, gets a "go PC 1" from this coordinator, and reports when its RESULT lines are in and both cards are restored; a hash-rate or power number taken while another job holds a card is not a number. The CPU-only job runs only in a slot where no measurement overlaps it. @@ -188,4 +188,4 @@ Rule: one PC 1 job at a time; an agent asks by message before publishing, gets a If job 1's package is more than 15 minutes away when job 2 is ready, job 2 goes first; the short job 3 fills any gap of under 10 minutes between packages. -23:30. No more agents are spawned tonight (the project lead: no unnecessary credits); the running ones finish. PC 1 queue change: the AMD-proving CPU fallback's small fixture (block-56-transfers-3shards, minutes) runs NOW in the gap before the era package; its S_p shard (block-338-shard1, up to 30 min of every core) stays job 6, last. Next-cut list gains the rig installer's two follow-ups (the Linux manifest entry, the Linux prover build), tied to whichever release carries the proving half (rollout plan 8a). +20:39. No more agents are spawned tonight (the project lead: no unnecessary credits); the running ones finish. PC 1 queue change: the AMD-proving CPU fallback's small fixture (block-56-transfers-3shards, minutes) runs NOW in the gap before the era package; its S_p shard (block-338-shard1, up to 30 min of every core) stays job 6, last. Next-cut list gains the rig installer's two follow-ups (the Linux manifest entry, the Linux prover build), tied to whichever release carries the proving half (rollout plan 8a). From 3f9bf6057672f5cc1eae711d1bc644d8aeb0476e Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:41:37 +0000 Subject: [PATCH 030/131] Counter ASIC 2.0 status 20:45: the node switch written, the Metal worker gate item, PC 1 gaps --- docs/plans/counter-asic-2-status.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 2fd7f81b4..c14ddf8cd 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -189,3 +189,11 @@ Rule: one PC 1 job at a time; an agent asks by message before publishing, gets a If job 1's package is more than 15 minutes away when job 2 is ready, job 2 goes first; the short job 3 fills any gap of under 10 minutes between packages. 20:39. No more agents are spawned tonight (the project lead: no unnecessary credits); the running ones finish. PC 1 queue change: the AMD-proving CPU fallback's small fixture (block-56-transfers-3shards, minutes) runs NOW in the gap before the era package; its S_p shard (block-338-shard1, up to 30 min of every core) stays job 6, last. Next-cut list gains the rig installer's two follow-ups (the Linux manifest entry, the Linux prover build), tied to whichever release carries the proving half (rollout plan 8a). + +## 20:45 the node switch is written; the Metal worker is a gate item + +ca2-node (a3f505a9d981300cd): ca2-v3 commits d2cd6e1 (pack-loop af983a7 merged: packcheck.rs and the attempt rule; one packfile.h conflict resolved), 50d5c86 (the seam: ProgramClass { V2, V3 }, V3_CLASS placeholder w16, generator 3 in the program id, Epoch::from_chain_seeds, Epoch::chain_program, IGNEUM_PROGRAM_CLASS and IGNEUM_ERA_SEED_HEX in the packs, packcheck refuses wrong class / era / generator; v2 packs byte-identical; 53 tests), 9ed787e (workers: packfile.h reads class and era, CUDA and OpenCL workers take `class=v3 era=` tokens on job and prepare lines, pair identity includes them, mismatch answers `need`; 13 packfile checks). Fast-time gate script infra/fast-time/class-v3.mjs written, not run (needs the fork binaries). Node (ca2-v3-node, uncommitted until cargo check passes, queued behind the measure lock): the field in Params, OverrideParams, override_params and the digest (unconditional, its own statement; 11-entry digest test), the daemon line, program_class_for_epoch_at (v3 iff 3600 e >= N4, first epoch = ceil), POW_ERA_BLOCKS 15,552,000 and POW_ERA_LEAD 7,200 with the era stand-in (era 0 = genesis), EpochSeeds { epoch, day, class, era }, template and RPC fields 12 to 16, the miner's job line and seeds.txt. PC 2 command ready (run from the ca2-v3 worktree so push-build-inputs.sh packs the v3 igneum-pow); held until "ready for PC 2". + +Gate item found: main.swift refuses v3 lines "until Swift has generator 3"; the Mac worker never regenerates a program (igneum-miner export-pack writes the pack), so the fix ordered is to accept the pack's class and era against the line, as the CUDA and OpenCL workers do. Without it Mac node 1 and the Mac app cannot mine v3 (gates G1 and G4). + +PC 1: the AMD-proving small fixture is running; Ember Tune's 8-minute CPU-only build takes the next CPU gap, its 30-minute both-cards run is job 5. From 3e3b12b5050dd78a4a1e124794616721e32b716a Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:41:50 +0000 Subject: [PATCH 031/131] Counter ASIC 2.0: C18 per-pound correction struck everywhere, C19 and C20 routed --- docs/plans/counter-asic-2-public.md | 2 +- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 4 +++- 3 files changed, 5 insertions(+), 3 deletions(-) diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index 4ccc5e2da..eebe3a4c9 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -34,7 +34,7 @@ Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per h | RTX 5090 (CUDA) | 136.1 | [owed] | 512 + hot | 0.96 at v2 | [owed] | | RX 9070 XT (OpenCL, eGPU) | 18.15 | [owed] | 512 + hot | 0.87 at v2 | [owed] | -AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), near parity per pound and about 4.5x worse per watt; the card's memory system, not a tuning gap. +AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), 2.2x worse per pound at list prices and 4.9x worse per watt (read-width.md section 4.1, approximate); the card's memory system, not a tuning gap. Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant]. diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index a944248d8..be09320fa 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -58,7 +58,7 @@ The layer 6 finding changes the headline: the strongest chip holds the whole cac ### 6b. The user tiers (the consequences rule) -AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second at 1 GiB, measured tonight), near parity per pound and about 4.5x worse per watt (0.089 against 0.398 MH/W, measured 5 October 2026). This is the card's memory system, not a tuning gap: no read width closes it without making the 5090 bandwidth-bound. The level 3 numbers page states it. +AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second at 1 GiB, measured tonight), 2.2x worse per pound at list prices (0.032 against 0.072 MH/s per pound, approximate) and 4.9x worse per watt (read-width.md section 4.1). This is the card's memory system, not a tuning gap: no read width closes it without making the 5090 bandwidth-bound. The level 3 numbers page states it. ## 7. Gates before any publish (all of them, no exceptions) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index c14ddf8cd..f68c96bdb 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -140,7 +140,7 @@ Layer 9 (the project lead: faster program changes): the epoch length becomes an ca2-era mid-way (a452664c512c73b9b): era draw behind LoadClass::era (EraParams beside mix, load_slots, scratch, scratch_kb; verify::load_index; memhard::Layout; one load form in the three emitters; --era, --era-widths), 54 crate tests green, pinned packs byte-identical; six era packs bit-exact on Metal (packbench 6/6), Apple OpenCL (6/6, same fingerprints) and the CUDA emulation (6/6); re-exporting with the width pinned at 4 B and rebasing onto e752fc7; PC job in about 20 minutes. Two class bugs found and fixed on its branch, both outside its layer: (1) proto-cuda/nvrtc/packfile.h re-derived the seed words as attempt 0 only, so ANY pack with IGNEUM_PROGRAM_ATTEMPT >= 1 (5.14% of chain epochs under v2) is refused by the one-click workers with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", the error PC 1 logged on 5 October and attributed to the export race; fixed attempt-aware, with a tampered-attempt refusal test. This fix must ship in the 0.3.11 workers whatever else does. (2) proto-cuda/host.cu and proto-opencl/host.c derived dataset words on the host as mh_item(w >> 4)[w AND 15] instead of the pack's mh_word; fixed. -Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (20:23, 20:24 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.5x the electricity per hash (0.089 against 0.398 MH/W), near parity per pound; the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. +Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (20:23, 20:24 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.9x the electricity per hash and 2.2x worse per pound at list prices (0.032 against 0.072 MH/s per pound, read-width.md 4.1, approximate); the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound. ## 20:29 C1 decided for the morning; one packfile.h fix for 0.3.11 @@ -197,3 +197,5 @@ ca2-node (a3f505a9d981300cd): ca2-v3 commits d2cd6e1 (pack-loop af983a7 merged: Gate item found: main.swift refuses v3 lines "until Swift has generator 3"; the Mac worker never regenerates a program (igneum-miner export-pack writes the pack), so the fix ordered is to accept the pack's class and era against the line, as the CUDA and OpenCL workers do. Without it Mac node 1 and the Mac app cannot mine v3 (gates G1 and G4). PC 1: the AMD-proving small fixture is running; Ember Tune's 8-minute CPU-only build takes the next CPU gap, its 30-minute both-cards run is job 5. + +20:50. Consequences round 2 (C18 to C20). C18: "near parity per pound" was wrong and is struck everywhere; read-width.md section 4.1 gives the 5090 at 2.2x the 9070 XT per pound at list (0.072 against 0.032 MH/s per pound, approximate), 4.9x per watt, 7.5x in rate; the level 3 page carries those. C19 (to ca2-mixer, already sent by the reviewer): the x4 verifier cost in IBD minutes over the 108,000-header pruning window per tier (8.6 min against 1.1 on one M5 Max core at the top of the range), pool shares per core per second, a scaled 2019-class figure (approximate), and the 10 ms gate margin left for 3.0 go into mixer-x4.md; the seeds' header-verify load goes into the testnet go checklist. C20 (to ca2-epoch, already sent): both rig miners run --exit-on-seed-change and re-export on exit 42, so a 10-minute epoch restarts every card's miner six times an hour and the Mac fleet's prepare pause goes from 35 s to 3.5 min an hour; the epoch-length document gets a per-tier restart-cost row and the compile-ahead margin against the VDF at 600 DAA s, and the rig installer drops the exit-42 path for prepare-ahead before any short epoch can be drawn (next-cut list). From 0e337349d027717207e65a1d3e37958197f39bd5 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:42:03 +0000 Subject: [PATCH 032/131] Counter ASIC 2.0 rollout: the rig miners' exit-42 path on the next-cut list as a layer 9 precondition --- docs/plans/counter-asic-2-rollout.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index be09320fa..0190f57f0 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -94,7 +94,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2): fork branch proving-v1 (from a24ab01a: segment records, the pool split, five RPCs, p2p message 72, parameters proving_v1_activation_daa, segment_blocks, unproven_daa, aggregator_share_bps, in the digest only once set) and its app branch, rebased onto the 0.3.10 tips. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. -Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. +Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. ## 9. After the publish: Counter ASIC 3.0 From 878c0581ef99bd1e1f111b78b83b265c90c53ab8 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:44:16 +0000 Subject: [PATCH 033/131] Counter ASIC 2.0 status 21:00: the Metal pack path, the --prepare-packs app item, the seam signatures, the build queue --- docs/plans/counter-asic-2-status.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index f68c96bdb..ef1d065e2 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -199,3 +199,11 @@ Gate item found: main.swift refuses v3 lines "until Swift has generator 3"; the PC 1: the AMD-proving small fixture is running; Ember Tune's 8-minute CPU-only build takes the next CPU gap, its 30-minute both-cards run is job 5. 20:50. Consequences round 2 (C18 to C20). C18: "near parity per pound" was wrong and is struck everywhere; read-width.md section 4.1 gives the 5090 at 2.2x the 9070 XT per pound at list (0.072 against 0.032 MH/s per pound, approximate), 4.9x per watt, 7.5x in rate; the level 3 page carries those. C19 (to ca2-mixer, already sent by the reviewer): the x4 verifier cost in IBD minutes over the 108,000-header pruning window per tier (8.6 min against 1.1 on one M5 Max core at the top of the range), pool shares per core per second, a scaled 2019-class figure (approximate), and the 10 ms gate margin left for 3.0 go into mixer-x4.md; the seeds' header-verify load goes into the testnet go checklist. C20 (to ca2-epoch, already sent): both rig miners run --exit-on-seed-change and re-export on exit 42, so a 10-minute epoch restarts every card's miner six times an hour and the Mac fleet's prepare pause goes from 35 s to 3.5 min an hour; the epoch-length document gets a per-tier restart-cost row and the compile-ahead margin against the VDF at 600 DAA s, and the rig installer drops the exit-42 path for prepare-ahead before any short epoch can be drawn (next-cut list). + +## 21:00 the Metal worker's v3 path; the app flag; the seam + +Correction to the 20:45 entry: igneum-bench --serve DOES regenerate every program in Swift (serveProgram calls generateProgramV2), so a v3 line could not be trusted blind. ca2-node added servePackProgram in main.swift: a `prepare class=v3 era=` line compiles program_bound.metal from the pack the miner wrote (--prepare-packs) after the packfile.h checks in Swift (generator 2 or 3, class against generator, seed bytes against the line, IGNEUM_SEEDW_INIT against attempt_words, class and era against the line); the program store keys on (seed, class, era); a v3 job with no resident pack answers `need` + `error ... program class mismatch`; v2 lines unchanged. A pack program never races variants (the Mac loses the variant race on v3 epochs; its cost is the race's gain, from the miner-perf entry, to be quoted). Integration item: the Mac app and Mac node 1's miner command must pass --prepare-packs, else the first v3 epoch on the Mac worker ends in `need` lines; check app/igneum-app/src/engine.rs. The gate network's Mac miner is the CPU miner (igneum-miner --engine igneum-pow), unaffected. + +The seam the node relies on (kept by every branch): Epoch::from_chain_seeds(epoch, day, era, class, label), Epoch::chain_program(epoch, era, class, label), Epoch::chain_dataset(day, class), generate_from_seed_bytes_program_class(label, seed, class, era), ProgramClass::{load_class, generator_version, from_generator, name, parse}, V3_CLASS, Program::era_bytes, packcheck::verify_pack_dir_chain. + +Mac build queue: the measure lock has been held by a packbench run since 20:31Z with three build slots held and five builds waiting; the node's cargo check and the swiftc recompile wait behind it. This is the lock working as designed; it sets the pace of the gates tonight. From fbeb3bb94d32dcb3825361ca87bda0b6ac9cd461 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:45:35 +0000 Subject: [PATCH 034/131] Counter ASIC 2.0 status 21:05: the 9070 XT absent on PC 1, app restart facts, era package ETA --- docs/plans/counter-asic-2-status.md | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index ef1d065e2..b236063be 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -207,3 +207,13 @@ Correction to the 20:45 entry: igneum-bench --serve DOES regenerate every progra The seam the node relies on (kept by every branch): Epoch::from_chain_seeds(epoch, day, era, class, label), Epoch::chain_program(epoch, era, class, label), Epoch::chain_dataset(day, class), generate_from_seed_bytes_program_class(label, seed, class, era), ProgramClass::{load_class, generator_version, from_generator, name, parse}, V3_CLASS, Program::era_bytes, packcheck::verify_pack_dir_chain. Mac build queue: the measure lock has been held by a packbench run since 20:31Z with three build slots held and five builds waiting; the node's cargo check and the swiftc recompile wait behind it. This is the lock working as designed; it sets the pace of the gates tonight. + +## 21:05 the 9070 XT has dropped off PC 1's bus; app restart facts; the era package ETA + +The AMD sweep agent (a01dcb34ae16d867c) reports from a read-only probe at about 20:40 UTC: Get-PnpDevice lists only the integrated "AMD Radeon(TM) Graphics" (gfx1036) and the RTX 5090; the app's AMD worker now mines the gfx1036 at 3.12 MH/s; the 9070 XT is absent from PnP (the eGPU link: the Sonnet box or the USB4 router; earlier today it went Code 43 and came back after a driver reinstall and reboot). A 10-second rescan probe is granted (pnputil /scan-devices, the USB4 router status). the project lead is asleep and is not woken. Consequence if the card stays absent: gate G1 (bit-exact v3 on all three cards) and the 9070 XT rows of the era and hot-table tables cannot be taken tonight; the AMD-vendor stand-in available is the gfx1036 (RDNA 2, AMD OpenCL 3683.0, 3 MH/s), which ran the version 1 and version 2 conformance; whether it satisfies G1 for the devnet publish is asked of the coordinator. Every 9070 XT row taken before 20:40 (readwidth, dot4) stands. + +App restart facts from the log intake (`node tools/logs.mjs`, 20:45): PC 1's app run is win-ae432dc7-20261005-190232 (started 19:02:32, no restart since), so no PC 1 measurement tonight straddled an app restart; PC 2's run is win-1ccfe586-20261005-200114 (started 20:01:14, before the readwidth 5090 round at 20:09). The 0.3.10 manifest is still unpublished (the shipper's CI is queued); the "jobs folder cleared by the 0.3.10 update" reading was wrong: the folder is cleared by fetch jobs. + +The --prepare-packs item of 21:00 is resolved: app/igneum-app/src/engine.rs line 1225 passes it on master and on the 0.3.10 tree. + +Era package: 10 to 15 minutes away (the pack-loop packfile merged over readwidth's; the OpenCL verification and the mingw rebuild queued behind three held build slots); zip ~/Desktop/igneum-ca2-era-pc1.zip, fetch id fetch-ca2-era-20261005, one playbook relay/playbooks/ca2-era-pc1.ps1 doing both cards (about 6 to 10 min). Widths pinned at 4 B in every era pack (512 B per hash), windows identical across the six packs, so the six-era spread isolates stride plus interleave. The CPU-only proving fixture holds PC 1 until about 21:15 to 21:30; the hot-table job is not ready either, so the order stays era, then hot table. From c8f0fe7472d7de11d05ae7bbd76dacb21be86ea5 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:46:20 +0000 Subject: [PATCH 035/131] Counter ASIC 2.0: the G1 ruling (gfx1036 as the AMD vendor tonight), hardware events section for the morning --- docs/plans/counter-asic-2-rollout.md | 9 ++++++++- docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 12 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 0190f57f0..d01772f6c 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -64,7 +64,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | # | Gate | Evidence required | State | |---|---|---|---| -| G1 | bit-exact v3 on all three cards against the Mac reference | vectors PASS on Metal, CUDA (5090), OpenCL (9070 XT) for the v3 packs; batch fingerprints equal | open | +| G1 | bit-exact v3 on all three vendors against the Mac reference | vectors PASS on Metal, CUDA (5090), AMD OpenCL for the v3 packs; batch fingerprints equal. RULING (coordinator, 5 October 2026, about 21:10 UTC): the integrated gfx1036 (RDNA 2, AMD OpenCL 3683.0) satisfies the AMD vendor tonight, because G1 is a compiler-and-ISA property and gfx1036 carried the v1 and v2 conformance; the 9070 XT's hash-rate and power rows are owed and taken when its link is back | open | | G2 | the CPU verifier exact on 1,000 random hashes per card | 1,000 GPU hashes per card re-hashed by `igneum-pow` on the Mac, 0 mismatches | open | | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary | open | @@ -86,6 +86,13 @@ If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-st Public lines to fix on the integration branch (card-lifetime table 3): `site/index.html` "2 GB, growing" gains the rate ("2 GB at genesis, doubling at years 4, 12 and 28"); "Any 4 GB card" becomes "any 4 GB card at launch, 8 GB from year 4"; the litepaper's "4 GB about four years, 8 GB more than a decade" stays with mapping (b) and gains "under the step schedule"; `docs/evidence.md` gains a row for the card-lifetime claim labelled designed. hot-table.md line 73's 8 GB row is corrected (it counted a 5090's 8,160 warps; a real 8 GB card has 20 to 24 SMs). +## 7b. Hardware events (for the morning summary) + +| When (UTC) | Event | What the app did | For the project lead | +|---|---|---|---| +| 5 October, at install (earlier today) | the RX 9070 XT in the Sonnet Breakaway Box 850T5 over USB4 went Code 43 | came back after a driver reinstall and a reboot | | +| 5 October, about 20:40 | the 9070 XT dropped off PC 1's bus: Get-PnpDevice -Class Display lists only the integrated AMD Radeon Graphics (gfx1036) and the RTX 5090 | the AMD worker (igneum-worker-opencl --device 1) mines the gfx1036 at 3.12 MH/s; the 5090 keeps mining; nobody was woken, PC 1's app was not restarted; a 10-second rescan probe (pnputil /scan-devices, the USB4 router status) was granted | the second eGPU link fault today: reseat the USB4 cable and the eGPU's power; the 0.3.10 hot-plug code shows the card as "removed" and picks it up again without a restart | + ## 7a. A dated constraint from the consequences review (C1) The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index b236063be..4c8ee79ee 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -217,3 +217,7 @@ App restart facts from the log intake (`node tools/logs.mjs`, 20:45): PC 1's app The --prepare-packs item of 21:00 is resolved: app/igneum-app/src/engine.rs line 1225 passes it on master and on the 0.3.10 tree. Era package: 10 to 15 minutes away (the pack-loop packfile merged over readwidth's; the OpenCL verification and the mingw rebuild queued behind three held build slots); zip ~/Desktop/igneum-ca2-era-pc1.zip, fetch id fetch-ca2-era-20261005, one playbook relay/playbooks/ca2-era-pc1.ps1 doing both cards (about 6 to 10 min). Widths pinned at 4 B in every era pack (512 B per hash), windows identical across the six packs, so the six-era spread isolates stride plus interleave. The CPU-only proving fixture holds PC 1 until about 21:15 to 21:30; the hot-table job is not ready either, so the order stays era, then hot table. + +## 21:10 ruling on G1; hardware event recorded + +Ruling (coordinator): gfx1036 satisfies the AMD vendor for gate G1 tonight (a compiler-and-ISA property; it carried the v1 and v2 conformance); the 9070 XT's hash-rate and power rows are owed and taken when the link is back; every 9070 XT row before 20:40 UTC stands. Nobody is woken, PC 1's app is not restarted. The event is in the rollout plan section 7b (hardware events) for the morning summary: the second eGPU link fault today (Code 43 at install, a bus drop at about 20:40); the project lead reseats the USB4 cable and the eGPU power; the 0.3.10 hot-plug code shows "removed" and picks the card up without a restart. The publish proceeds when every other gate is green. From 9343d96a98e411efb0804994e54d117d75549844 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:47:57 +0000 Subject: [PATCH 036/131] Counter ASIC 2.0 status 21:15: the Sonnet box off the link, PC 1 facts corrected, the queue --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 8 ++++++++ 2 files changed, 9 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index d01772f6c..e8ad72d27 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -91,7 +91,7 @@ Public lines to fix on the integration branch (card-lifetime table 3): `site/ind | When (UTC) | Event | What the app did | For the project lead | |---|---|---|---| | 5 October, at install (earlier today) | the RX 9070 XT in the Sonnet Breakaway Box 850T5 over USB4 went Code 43 | came back after a driver reinstall and a reboot | | -| 5 October, about 20:40 | the 9070 XT dropped off PC 1's bus: Get-PnpDevice -Class Display lists only the integrated AMD Radeon Graphics (gfx1036) and the RTX 5090 | the AMD worker (igneum-worker-opencl --device 1) mines the gfx1036 at 3.12 MH/s; the 5090 keeps mining; nobody was woken, PC 1's app was not restarted; a 10-second rescan probe (pnputil /scan-devices, the USB4 router status) was granted | the second eGPU link fault today: reseat the USB4 cable and the eGPU's power; the 0.3.10 hot-plug code shows the card as "removed" and picks it up again without a restart | +| 5 October, about 20:40 | the 9070 XT dropped off PC 1's bus: Get-PnpDevice -Class Display lists only the integrated AMD Radeon Graphics (gfx1036) and the RTX 5090; after `pnputil /scan-devices` at 20:45:34Z the card is still absent and the USB4 list shows only the host and root routers: the "USB4 Router (2.0), Sonnet Technologies Breakaway Box 850T5" present at 17:18Z is gone, so the box itself is off the link | the AMD worker (igneum-worker-opencl --device 1) mines the gfx1036 at 3.12 MH/s; the 5090 keeps mining; nobody was woken, PC 1's app was not restarted; a 10-second rescan probe (pnputil /scan-devices, the USB4 router status) was granted | the second eGPU link fault today: reseat the USB4 cable and the eGPU's power; the 0.3.10 hot-plug code shows the card as "removed" and picks it up again without a restart | ## 7a. A dated constraint from the consequences review (C1) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 4c8ee79ee..0a6d684b4 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -221,3 +221,11 @@ Era package: 10 to 15 minutes away (the pack-loop packfile merged over readwidth ## 21:10 ruling on G1; hardware event recorded Ruling (coordinator): gfx1036 satisfies the AMD vendor for gate G1 tonight (a compiler-and-ISA property; it carried the v1 and v2 conformance); the 9070 XT's hash-rate and power rows are owed and taken when the link is back; every 9070 XT row before 20:40 UTC stands. Nobody is woken, PC 1's app is not restarted. The event is in the rollout plan section 7b (hardware events) for the morning summary: the second eGPU link fault today (Code 43 at install, a bus drop at about 20:40); the project lead reseats the USB4 cable and the eGPU power; the 0.3.10 hot-plug code shows "removed" and picks the card up without a restart. The publish proceeds when every other gate is green. + +## 21:15 the Sonnet box is off the link; PC 1 facts corrected; the queue after the era job + +Rescan at 20:45:34Z (relay probe #203, 10 s): the 9070 XT stays absent after pnputil /scan-devices; the USB4 list shows only the host and root routers, the Sonnet Breakaway Box 850T5 router present at 17:18Z is gone: the box is off the link, not just the card. Job 4 (the 9070 XT sweep) is dropped, its rows owed with this reason and time. The era and hot-table PC jobs run their AMD half on the gfx1036 for bit-exactness only (the G1 ruling); their 9070 XT hash-rate and probe rows are owed. + +Correction to the 21:05 entry: PC 1's app is 0.3.9 (file 15:47:20Z) and its process started at 20:01:14Z (pid 12340), a restart, not a 0.3.10 install; the log intake's run id dates the log file, not the process. Both PCs restarted at about 20:01Z, before every readwidth PC job (from 20:02:51Z) and the dot4 probe (20:27Z), so no measurement tonight straddled a restart. 0.3.10 is still unpublished. + +PC 1 queue now: (1) the AMD-proving small fixture (running, release expected 21:15 to 21:30), (2) the era job (both halves, about 6 to 10 min), (3) the 5090 power-limit sweep (575 / 460 / 400 / 400 W, 90 s each, cap restored to 431 W, about 8 min; SM and memory clocks in the RESULT lines), (4) the hot-table job, (5) the reproducible benchmark (5 min), (6) Ember Tune's 8-minute build in a CPU gap then its 30-minute both-cards run, (7) the AMD-proving S_p shard (up to 90 min, CPU only). From 27471633c7054bd5438a829130b5eb952965bf9e Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:48:20 +0000 Subject: [PATCH 037/131] Counter ASIC 2.0 status: heading times 20:42 to 20:47 corrected to the commit clock --- docs/plans/counter-asic-2-status.md | 20 ++++++++++---------- 1 file changed, 10 insertions(+), 10 deletions(-) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 0a6d684b4..f6102f5a3 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -1,6 +1,6 @@ # Counter ASIC 2.0: status -Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rewritten every 45 minutes while the work runs. Times UTC, 5 October 2026 (night). Heading times before 20:40 were corrected at 20:42 from the commit clock (the coordinator had written them from a guessed clock, up to 2 h 40 min ahead); every entry's true time is its commit's author time in UTC. Base for every ca2 branch: `readwidth` at 019b014 (the LoadClass flag, fold_words, the scratch op, the three emitters, 20 packs). +Coordinator's running status for the plan in `docs/plans/counter-asic-2.md`. Rewritten every 45 minutes while the work runs. Times UTC, 5 October 2026 (night). Heading times before 20:40 were corrected at 20:42 from the commit clock (the coordinator had written them from a guessed clock, up to 2 h 40 min ahead; corrected again at 20:49 for the 20:42 to 20:47 entries); every entry's true time is its commit's author time in UTC, and from 20:49 every heading is stamped from `date -u`. Base for every ca2 branch: `readwidth` at 019b014 (the LoadClass flag, fold_words, the scratch op, the three emitters, 20 packs). ## 19:55 first status (the 19:50 start was cut off by exhausted credits at about 19:58 before any sub-agent work landed; respawned at 19:55 on the restart) @@ -190,7 +190,7 @@ If job 1's package is more than 15 minutes away when job 2 is ready, job 2 goes 20:39. No more agents are spawned tonight (the project lead: no unnecessary credits); the running ones finish. PC 1 queue change: the AMD-proving CPU fallback's small fixture (block-56-transfers-3shards, minutes) runs NOW in the gap before the era package; its S_p shard (block-338-shard1, up to 30 min of every core) stays job 6, last. Next-cut list gains the rig installer's two follow-ups (the Linux manifest entry, the Linux prover build), tied to whichever release carries the proving half (rollout plan 8a). -## 20:45 the node switch is written; the Metal worker is a gate item +## 20:42 the node switch is written; the Metal worker is a gate item ca2-node (a3f505a9d981300cd): ca2-v3 commits d2cd6e1 (pack-loop af983a7 merged: packcheck.rs and the attempt rule; one packfile.h conflict resolved), 50d5c86 (the seam: ProgramClass { V2, V3 }, V3_CLASS placeholder w16, generator 3 in the program id, Epoch::from_chain_seeds, Epoch::chain_program, IGNEUM_PROGRAM_CLASS and IGNEUM_ERA_SEED_HEX in the packs, packcheck refuses wrong class / era / generator; v2 packs byte-identical; 53 tests), 9ed787e (workers: packfile.h reads class and era, CUDA and OpenCL workers take `class=v3 era=` tokens on job and prepare lines, pair identity includes them, mismatch answers `need`; 13 packfile checks). Fast-time gate script infra/fast-time/class-v3.mjs written, not run (needs the fork binaries). Node (ca2-v3-node, uncommitted until cargo check passes, queued behind the measure lock): the field in Params, OverrideParams, override_params and the digest (unconditional, its own statement; 11-entry digest test), the daemon line, program_class_for_epoch_at (v3 iff 3600 e >= N4, first epoch = ceil), POW_ERA_BLOCKS 15,552,000 and POW_ERA_LEAD 7,200 with the era stand-in (era 0 = genesis), EpochSeeds { epoch, day, class, era }, template and RPC fields 12 to 16, the miner's job line and seeds.txt. PC 2 command ready (run from the ca2-v3 worktree so push-build-inputs.sh packs the v3 igneum-pow); held until "ready for PC 2". @@ -198,34 +198,34 @@ Gate item found: main.swift refuses v3 lines "until Swift has generator 3"; the PC 1: the AMD-proving small fixture is running; Ember Tune's 8-minute CPU-only build takes the next CPU gap, its 30-minute both-cards run is job 5. -20:50. Consequences round 2 (C18 to C20). C18: "near parity per pound" was wrong and is struck everywhere; read-width.md section 4.1 gives the 5090 at 2.2x the 9070 XT per pound at list (0.072 against 0.032 MH/s per pound, approximate), 4.9x per watt, 7.5x in rate; the level 3 page carries those. C19 (to ca2-mixer, already sent by the reviewer): the x4 verifier cost in IBD minutes over the 108,000-header pruning window per tier (8.6 min against 1.1 on one M5 Max core at the top of the range), pool shares per core per second, a scaled 2019-class figure (approximate), and the 10 ms gate margin left for 3.0 go into mixer-x4.md; the seeds' header-verify load goes into the testnet go checklist. C20 (to ca2-epoch, already sent): both rig miners run --exit-on-seed-change and re-export on exit 42, so a 10-minute epoch restarts every card's miner six times an hour and the Mac fleet's prepare pause goes from 35 s to 3.5 min an hour; the epoch-length document gets a per-tier restart-cost row and the compile-ahead margin against the VDF at 600 DAA s, and the rig installer drops the exit-42 path for prepare-ahead before any short epoch can be drawn (next-cut list). +20:43. Consequences round 2 (C18 to C20). C18: "near parity per pound" was wrong and is struck everywhere; read-width.md section 4.1 gives the 5090 at 2.2x the 9070 XT per pound at list (0.072 against 0.032 MH/s per pound, approximate), 4.9x per watt, 7.5x in rate; the level 3 page carries those. C19 (to ca2-mixer, already sent by the reviewer): the x4 verifier cost in IBD minutes over the 108,000-header pruning window per tier (8.6 min against 1.1 on one M5 Max core at the top of the range), pool shares per core per second, a scaled 2019-class figure (approximate), and the 10 ms gate margin left for 3.0 go into mixer-x4.md; the seeds' header-verify load goes into the testnet go checklist. C20 (to ca2-epoch, already sent): both rig miners run --exit-on-seed-change and re-export on exit 42, so a 10-minute epoch restarts every card's miner six times an hour and the Mac fleet's prepare pause goes from 35 s to 3.5 min an hour; the epoch-length document gets a per-tier restart-cost row and the compile-ahead margin against the VDF at 600 DAA s, and the rig installer drops the exit-42 path for prepare-ahead before any short epoch can be drawn (next-cut list). -## 21:00 the Metal worker's v3 path; the app flag; the seam +## 20:43 the Metal worker's v3 path; the app flag; the seam -Correction to the 20:45 entry: igneum-bench --serve DOES regenerate every program in Swift (serveProgram calls generateProgramV2), so a v3 line could not be trusted blind. ca2-node added servePackProgram in main.swift: a `prepare class=v3 era=` line compiles program_bound.metal from the pack the miner wrote (--prepare-packs) after the packfile.h checks in Swift (generator 2 or 3, class against generator, seed bytes against the line, IGNEUM_SEEDW_INIT against attempt_words, class and era against the line); the program store keys on (seed, class, era); a v3 job with no resident pack answers `need` + `error ... program class mismatch`; v2 lines unchanged. A pack program never races variants (the Mac loses the variant race on v3 epochs; its cost is the race's gain, from the miner-perf entry, to be quoted). Integration item: the Mac app and Mac node 1's miner command must pass --prepare-packs, else the first v3 epoch on the Mac worker ends in `need` lines; check app/igneum-app/src/engine.rs. The gate network's Mac miner is the CPU miner (igneum-miner --engine igneum-pow), unaffected. +Correction to the 20:42 entry: igneum-bench --serve DOES regenerate every program in Swift (serveProgram calls generateProgramV2), so a v3 line could not be trusted blind. ca2-node added servePackProgram in main.swift: a `prepare class=v3 era=` line compiles program_bound.metal from the pack the miner wrote (--prepare-packs) after the packfile.h checks in Swift (generator 2 or 3, class against generator, seed bytes against the line, IGNEUM_SEEDW_INIT against attempt_words, class and era against the line); the program store keys on (seed, class, era); a v3 job with no resident pack answers `need` + `error ... program class mismatch`; v2 lines unchanged. A pack program never races variants (the Mac loses the variant race on v3 epochs; its cost is the race's gain, from the miner-perf entry, to be quoted). Integration item: the Mac app and Mac node 1's miner command must pass --prepare-packs, else the first v3 epoch on the Mac worker ends in `need` lines; check app/igneum-app/src/engine.rs. The gate network's Mac miner is the CPU miner (igneum-miner --engine igneum-pow), unaffected. The seam the node relies on (kept by every branch): Epoch::from_chain_seeds(epoch, day, era, class, label), Epoch::chain_program(epoch, era, class, label), Epoch::chain_dataset(day, class), generate_from_seed_bytes_program_class(label, seed, class, era), ProgramClass::{load_class, generator_version, from_generator, name, parse}, V3_CLASS, Program::era_bytes, packcheck::verify_pack_dir_chain. Mac build queue: the measure lock has been held by a packbench run since 20:31Z with three build slots held and five builds waiting; the node's cargo check and the swiftc recompile wait behind it. This is the lock working as designed; it sets the pace of the gates tonight. -## 21:05 the 9070 XT has dropped off PC 1's bus; app restart facts; the era package ETA +## 20:45 the 9070 XT has dropped off PC 1's bus; app restart facts; the era package ETA The AMD sweep agent (a01dcb34ae16d867c) reports from a read-only probe at about 20:40 UTC: Get-PnpDevice lists only the integrated "AMD Radeon(TM) Graphics" (gfx1036) and the RTX 5090; the app's AMD worker now mines the gfx1036 at 3.12 MH/s; the 9070 XT is absent from PnP (the eGPU link: the Sonnet box or the USB4 router; earlier today it went Code 43 and came back after a driver reinstall and reboot). A 10-second rescan probe is granted (pnputil /scan-devices, the USB4 router status). the project lead is asleep and is not woken. Consequence if the card stays absent: gate G1 (bit-exact v3 on all three cards) and the 9070 XT rows of the era and hot-table tables cannot be taken tonight; the AMD-vendor stand-in available is the gfx1036 (RDNA 2, AMD OpenCL 3683.0, 3 MH/s), which ran the version 1 and version 2 conformance; whether it satisfies G1 for the devnet publish is asked of the coordinator. Every 9070 XT row taken before 20:40 (readwidth, dot4) stands. App restart facts from the log intake (`node tools/logs.mjs`, 20:45): PC 1's app run is win-ae432dc7-20261005-190232 (started 19:02:32, no restart since), so no PC 1 measurement tonight straddled an app restart; PC 2's run is win-1ccfe586-20261005-200114 (started 20:01:14, before the readwidth 5090 round at 20:09). The 0.3.10 manifest is still unpublished (the shipper's CI is queued); the "jobs folder cleared by the 0.3.10 update" reading was wrong: the folder is cleared by fetch jobs. -The --prepare-packs item of 21:00 is resolved: app/igneum-app/src/engine.rs line 1225 passes it on master and on the 0.3.10 tree. +The --prepare-packs item of 20:43 is resolved: app/igneum-app/src/engine.rs line 1225 passes it on master and on the 0.3.10 tree. Era package: 10 to 15 minutes away (the pack-loop packfile merged over readwidth's; the OpenCL verification and the mingw rebuild queued behind three held build slots); zip ~/Desktop/igneum-ca2-era-pc1.zip, fetch id fetch-ca2-era-20261005, one playbook relay/playbooks/ca2-era-pc1.ps1 doing both cards (about 6 to 10 min). Widths pinned at 4 B in every era pack (512 B per hash), windows identical across the six packs, so the six-era spread isolates stride plus interleave. The CPU-only proving fixture holds PC 1 until about 21:15 to 21:30; the hot-table job is not ready either, so the order stays era, then hot table. -## 21:10 ruling on G1; hardware event recorded +## 20:46 ruling on G1; hardware event recorded Ruling (coordinator): gfx1036 satisfies the AMD vendor for gate G1 tonight (a compiler-and-ISA property; it carried the v1 and v2 conformance); the 9070 XT's hash-rate and power rows are owed and taken when the link is back; every 9070 XT row before 20:40 UTC stands. Nobody is woken, PC 1's app is not restarted. The event is in the rollout plan section 7b (hardware events) for the morning summary: the second eGPU link fault today (Code 43 at install, a bus drop at about 20:40); the project lead reseats the USB4 cable and the eGPU power; the 0.3.10 hot-plug code shows "removed" and picks the card up without a restart. The publish proceeds when every other gate is green. -## 21:15 the Sonnet box is off the link; PC 1 facts corrected; the queue after the era job +## 20:47 the Sonnet box is off the link; PC 1 facts corrected; the queue after the era job Rescan at 20:45:34Z (relay probe #203, 10 s): the 9070 XT stays absent after pnputil /scan-devices; the USB4 list shows only the host and root routers, the Sonnet Breakaway Box 850T5 router present at 17:18Z is gone: the box is off the link, not just the card. Job 4 (the 9070 XT sweep) is dropped, its rows owed with this reason and time. The era and hot-table PC jobs run their AMD half on the gfx1036 for bit-exactness only (the G1 ruling); their 9070 XT hash-rate and probe rows are owed. -Correction to the 21:05 entry: PC 1's app is 0.3.9 (file 15:47:20Z) and its process started at 20:01:14Z (pid 12340), a restart, not a 0.3.10 install; the log intake's run id dates the log file, not the process. Both PCs restarted at about 20:01Z, before every readwidth PC job (from 20:02:51Z) and the dot4 probe (20:27Z), so no measurement tonight straddled a restart. 0.3.10 is still unpublished. +Correction to the 20:45 entry: PC 1's app is 0.3.9 (file 15:47:20Z) and its process started at 20:01:14Z (pid 12340), a restart, not a 0.3.10 install; the log intake's run id dates the log file, not the process. Both PCs restarted at about 20:01Z, before every readwidth PC job (from 20:02:51Z) and the dot4 probe (20:27Z), so no measurement tonight straddled a restart. 0.3.10 is still unpublished. PC 1 queue now: (1) the AMD-proving small fixture (running, release expected 21:15 to 21:30), (2) the era job (both halves, about 6 to 10 min), (3) the 5090 power-limit sweep (575 / 460 / 400 / 400 W, 90 s each, cap restored to 431 W, about 8 min; SM and memory clocks in the RESULT lines), (4) the hot-table job, (5) the reproducible benchmark (5 min), (6) Ember Tune's 8-minute build in a CPU gap then its 30-minute both-cards run, (7) the AMD-proving S_p shard (up to 90 min, CPU only). From 4a1f872cb62af93a98c71d5158db594cd3d6f496 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:49:39 +0000 Subject: [PATCH 038/131] Counter ASIC 2.0 status 20:49: mixer construction, the lock cleared, PC 2 state, one merged-tree suite job --- docs/plans/counter-asic-2-status.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index f6102f5a3..60eb9b3d5 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -229,3 +229,11 @@ Rescan at 20:45:34Z (relay probe #203, 10 s): the 9070 XT stays absent after pnp Correction to the 20:45 entry: PC 1's app is 0.3.9 (file 15:47:20Z) and its process started at 20:01:14Z (pid 12340), a restart, not a 0.3.10 install; the log intake's run id dates the log file, not the process. Both PCs restarted at about 20:01Z, before every readwidth PC job (from 20:02:51Z) and the dot4 probe (20:27Z), so no measurement tonight straddled a restart. 0.3.10 is still unpublished. PC 1 queue now: (1) the AMD-proving small fixture (running, release expected 21:15 to 21:30), (2) the era job (both halves, about 6 to 10 min), (3) the 5090 power-limit sweep (575 / 460 / 400 / 400 W, 90 s each, cap restored to 431 W, about 8 min; SM and memory clocks in the RESULT lines), (4) the hot-table job, (5) the reproducible benchmark (5 min), (6) Ember Tune's 8-minute build in a CPU gap then its 30-minute both-cards run, (7) the AMD-proving S_p shard (up to 90 min, CPU only). + +## 20:49 mixer construction written; the measure lock cleared; PC 2 and the merged-tree suites + +ca2-mixer (af345b1e2c541ffbb), no commit yet (lands when the crate tests and the v2 pack diff are green): LoadClass gains mixer_mult (1 or 4) and growth; LoadClass::MX4 = v2 loads, mixer x4, growth on, no width roll, so its program stream is version 2's draw for draw; memhard::Shape { mixer_mult, cache_log2_words } in MixParams; derive_items applies the mixer with keys round_key(r x m + j), j in 0..m, before each of the 8 reads and round_key(8m + j) after; Cache::fill_log2; growth_doublings(d) = ilog2(1 + d / 1460) (doublings at years 4, 12, 28, 60), cache_log2_words(d) = 26 + doublings, dataset_log2_words capped at 32; d = 1 on the devnet pack keeps 2^26 and 2^28. Emitters emit the m-loop only when m > 1 (v2 text byte for byte otherwise); program.h carries IGNEUM_MIXER_MULT and IGNEUM_CACHE_GROWTH. Seam addition (additive): Epoch::chain_dataset_day(day_bytes, class, days_since_genesis, genesis_dataset_log2); the node agent was told to wire the genesis day index from Params.genesis.timestamp and a pow_genesis_dataset_log2 field (28 on the devnet, in the digest). V3_CLASS becomes LoadClass::MX4 composed with the era and hot fields at integration. Numbers follow the lock. + +The Mac measure lock: the packbench loop was the readwidth agent's (Metal currentAllocatedSize at 32 and 128 KiB scratch for the consequences reviewer's C12, not a decided row); it stopped the loop at about 20:52, so the queued builds (the node's cargo check, the mixer's tests, the epoch and mixer measurements) proceed. + +PC 2: spcurve-stopped-pc2-pv1b closed 20:47:46Z (the card alone: an empty shard 13,875 MiB in 2.1 s; a v1 shard 20,435 MiB, 4.2 s; 2.25 M pgas 28,371 MiB, 6.6 s; 4.5 M 28,307 MiB, 8.5 s; the prototype 6.75 M 28,275 MiB, 11.2 s: the peak plateaus at 28.3 GB from 20 M cycles up); spcurve-miner-pc2-pv1 (the same with the miner on, about 5 min) runs now; PC 2 is released after it. The proving fork tip is 3203c8d0 (eb32c645 plus the pool test's field and N = 8), app 440fd59; its unit tests ran on the Mac (consensus-core 13, exec 8, flows, 0 failed); its harness runs on the Mac. Decision: ONE PC 2 suite job on the merged tree (ca2-v3-node plus 3203c8d0) once the ca2 node branch is committed, covering both halves of 0.3.11. From 4f042f37dc0cc90c8c429faef570130767f0a7d9 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:50:05 +0000 Subject: [PATCH 039/131] Counter ASIC 2.0 status 20:50: the measure lock holder, the with-lock status defect --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 60eb9b3d5..bc2fbea90 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -237,3 +237,7 @@ ca2-mixer (af345b1e2c541ffbb), no commit yet (lands when the crate tests and the The Mac measure lock: the packbench loop was the readwidth agent's (Metal currentAllocatedSize at 32 and 128 KiB scratch for the consequences reviewer's C12, not a decided row); it stopped the loop at about 20:52, so the queued builds (the node's cargo check, the mixer's tests, the epoch and mixer measurements) proceed. PC 2: spcurve-stopped-pc2-pv1b closed 20:47:46Z (the card alone: an empty shard 13,875 MiB in 2.1 s; a v1 shard 20,435 MiB, 4.2 s; 2.25 M pgas 28,371 MiB, 6.6 s; 4.5 M 28,307 MiB, 8.5 s; the prototype 6.75 M 28,275 MiB, 11.2 s: the peak plateaus at 28.3 GB from 20 M cycles up); spcurve-miner-pc2-pv1 (the same with the miner on, about 5 min) runs now; PC 2 is released after it. The proving fork tip is 3203c8d0 (eb32c645 plus the pool test's field and N = 8), app 440fd59; its unit tests ran on the Mac (consensus-core 13, exec 8, flows, 0 failed); its harness runs on the Mac. Decision: ONE PC 2 suite job on the merged tree (ca2-v3-node plus 3203c8d0) once the ca2 node branch is committed, covering both halves of 0.3.11. + +## 20:50 the measure lock holder is a prover measurement; a lock-status defect + +Correction to the 20:50 entry above: the measure lock has been held since 20:31Z by pid 45000, a `with-lock.sh measure` of igneum-wt-agg-cost's igneum-prove-host (the aggregation-cost agent's prover measurement, 18 minutes so far), not by the readwidth packbench; `with-lock.sh status` prints the LAST WRITER's command text, not the holder's, which is why it named the w4 run. Defect for the next cut (tools/lock/with-lock.sh: the status line must read the holder's pid and command, not the last writer's; the class of CLAUDE.md's watcher rule). Builds proceed in the three build slots (build-0 taken at 20:49:51 after an 885-s wait); GPU measurements (the mixer's verifier and build timings, the epoch compile-ahead, the Mac bit-exactness runs under `run` are not blocked) queue behind the prover measurement. Readwidth is at 46dcc3e (per-watt rows for the consequences reviewer). From 949c4582e1ae0fc1fff6ae97e2c0a8a7f0c0f66c Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:50:30 +0000 Subject: [PATCH 040/131] Counter ASIC 2.0 status: readwidth head hash corrected to 06dcb31 --- docs/plans/counter-asic-2-status.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index bc2fbea90..dadbc3b71 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -240,4 +240,4 @@ PC 2: spcurve-stopped-pc2-pv1b closed 20:47:46Z (the card alone: an empty shard ## 20:50 the measure lock holder is a prover measurement; a lock-status defect -Correction to the 20:50 entry above: the measure lock has been held since 20:31Z by pid 45000, a `with-lock.sh measure` of igneum-wt-agg-cost's igneum-prove-host (the aggregation-cost agent's prover measurement, 18 minutes so far), not by the readwidth packbench; `with-lock.sh status` prints the LAST WRITER's command text, not the holder's, which is why it named the w4 run. Defect for the next cut (tools/lock/with-lock.sh: the status line must read the holder's pid and command, not the last writer's; the class of CLAUDE.md's watcher rule). Builds proceed in the three build slots (build-0 taken at 20:49:51 after an 885-s wait); GPU measurements (the mixer's verifier and build timings, the epoch compile-ahead, the Mac bit-exactness runs under `run` are not blocked) queue behind the prover measurement. Readwidth is at 46dcc3e (per-watt rows for the consequences reviewer). +Correction to the 20:50 entry above: the measure lock has been held since 20:31Z by pid 45000, a `with-lock.sh measure` of igneum-wt-agg-cost's igneum-prove-host (the aggregation-cost agent's prover measurement, 18 minutes so far), not by the readwidth packbench; `with-lock.sh status` prints the LAST WRITER's command text, not the holder's, which is why it named the w4 run. Defect for the next cut (tools/lock/with-lock.sh: the status line must read the holder's pid and command, not the last writer's; the class of CLAUDE.md's watcher rule). Builds proceed in the three build slots (build-0 taken at 20:49:51 after an 885-s wait); GPU measurements (the mixer's verifier and build timings, the epoch compile-ahead, the Mac bit-exactness runs under `run` are not blocked) queue behind the prover measurement. Readwidth head is 30ff674 (per-watt rows for the consequences reviewer; e752fc7 stays the table commit). From 1e34c3a4e487671c7adee38b9ee74405851e0d6e Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:51:13 +0000 Subject: [PATCH 041/131] Site and litepaper: the level 2 chip-resistance copy (three ideas, the fourth paragraph), the latency correction, the card-lifetime lines under the step schedule --- site/index.html | 4 ++-- site/litepaper.html | 9 +++++---- 2 files changed, 7 insertions(+), 6 deletions(-) diff --git a/site/index.html b/site/index.html index 2ccd3a077..f067ce671 100644 --- a/site/index.html +++ b/site/index.html @@ -440,7 +440,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var( Who minesCPUsGPUs Program changesEvery hashEvery hour - Memory2 GB, fixed2 GB, growing + Memory2 GB, fixed2 GB at genesis, doubling at years 4, 12 and 28 Verified byAny CPUAny CPU @@ -458,7 +458,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var(

Mine

-

Any 4 GB card, approximate. 80% of each block to its finder.

+

Any 4 GB card at launch, 8 GB from year 4, approximate. 80% of each block to its finder.

About the miner
diff --git a/site/litepaper.html b/site/litepaper.html index 44f156721..533f17c63 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -408,7 +408,7 @@ body.all .pager{display:none}

Mining: a program that never holds still

Every GPU chain that promised ASIC resistance shipped a fixed algorithm, and a fixed algorithm gets a chip the moment the prize pays for one. Igneum does not have a fixed algorithm.

-

Each hour the chain derives a seed from a locked checkpoint one epoch back, passes it through a ten-minute verifiable delay so no miner can see which program a seed implies before choosing whether to publish a block, and feeds it to a deterministic generator. The generator emits a random integer program built from what graphics cards are uniquely good at: wide parallel integer maths, shuffles between the 32 lanes of a warp, and random reads over a multi-gigabyte dataset that changes daily, so the program is bound by memory bandwidth. The memory footprint and instruction count are fixed and only the maths sequence is random, so no hour favours one vendor's cards and nobody gains by grinding the seed. Miners compile the program once per hour. Anyone running a node, a wallet or an exchange checks a hash on an ordinary CPU in under ten milliseconds by simulating one warp, so nobody needs a GPU except to mine. Measured: 0.41 to 0.58 ms per warp on one Apple M5 Max core with the 256 MB cache, about 17x inside the 10 ms gate; a 2019-class laptop core is not yet measured.

+

Each hour the chain derives a seed from a locked checkpoint one epoch back, passes it through a ten-minute verifiable delay so no miner can see which program a seed implies before choosing whether to publish a block, and feeds it to a deterministic generator. The generator emits a random integer program built from what graphics cards are uniquely good at: wide parallel integer maths, shuffles between the 32 lanes of a warp, and random reads over a multi-gigabyte dataset that changes daily, so the program waits on memory latency, not on maths or bandwidth. The memory footprint and instruction count are fixed and only the maths sequence is random, so no hour favours one vendor's cards and nobody gains by grinding the seed. Miners compile the program once per hour. Anyone running a node, a wallet or an exchange checks a hash on an ordinary CPU in under ten milliseconds by simulating one warp, so nobody needs a GPU except to mine. Measured: 0.41 to 0.58 ms per warp on one Apple M5 Max core with the 256 MB cache, about 17x inside the 10 ms gate; a 2019-class laptop core is not yet measured.

The hash is a lottery, not a general-purpose cryptographic hash. It has to be unpredictable per nonce, free of any shortcut cheaper than honest evaluation, and free of bias a miner can exploit. It does not need preimage or collision resistance. Open: no analysis of the lottery properties exists yet. It is the first job of the external review in phase 1, and until then the hash is a design claim backed by the measurements below.

@@ -420,7 +420,8 @@ body.all .pager{display:none}
ClockWhat changesMiner update needed?
ContinuouslyThe dataset grows on a schedule fixed at genesis, slowly enough that consumer cards keep up for years. A chip is built with fixed memory, so it is on a countdown from the day it ships. Ethereum's growing dataset ran Bitmain's E3 out of memory in 2020 this way, approximate, with nobody doing anythingNo
-

Everything above is automatic. Nobody writes a new program, nobody schedules a fork, and Igneum runs on the generator fixed at genesis, as Monero has run on RandomX since 2019 with no chip publicly shipped, approximate. Monero is precedent, not proof: absence of a public chip does not show that none can exist, which is why a bounty exists. The generator is designed so that the best hardware for any program it can emit is a graphics card. Target: a chip that dropped the graphics parts and kept the parallel cores and the memory gains under 2x, below what pays for a tapeout. That is a design target, not a measurement. Ethash chips reached roughly 1.5 to 2x, approximate, and the standing bounty exists to test the target. On top of that, the widening program space and the growing dataset mean a chip designed for this year's Igneum meets a harder Igneum next year without anyone lifting a finger.

+

Three ideas carry the chip resistance. The hash rewrites itself. A new program every hour, drawn from the chain. Its memory pattern changes with it. The rules change on a schedule fixed at launch. No release, no vote. It waits on memory, not maths. Every hash is a chain of random reads into a table too big for a chip to carry. The wait is the same physics for everyone. Miners hold the switch. Spare defences are written into the rules, switched off. A 90% miner signal turns one on. No fork.

+

No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured. The model and the bounty are public: the numbers. Monero has run on RandomX since 2019 with no chip publicly shipped, approximate; that is precedent, not proof.

One thing takes a person, here and on every chain that exists: writing new code. A chain cannot safely write its own generator, and it cannot safely tell a chip from a wave of honest new cards by hashrate alone. If the design above ever failed, anyone could publish a new generator and miners would switch it on by signalling, as Monero's community can fork. Igneum is built to make that day unlikely, and does not depend on avoiding it.

@@ -432,9 +433,9 @@ body.all .pager{display:none} Hardware it is built forCPUs. GPUs run it badly on purposeGPUs. Any card, any vendor. Bit-exact on Apple, NVIDIA and AMD, measured Random programPer hash, interpreted in a virtual machinePer hour, compiled to native GPU code. Per hash, the 128 dataset addresses change with the nonce - DatasetAbout 2 GB, the same size since 2019, approximate2 GB at genesis, growing every year past any chip's memory + DatasetAbout 2 GB, the same size since 2019, approximate2 GB at genesis, doubling at years 4, 12 and 28 under the step schedule; a 4 GB card mines about four years, an 8 GB card past a decade Light verification256 MB cache on a CPU, milliseconds256 MB cache on a CPU, one warp under 10 ms, the gate. Measured 0.41 to 0.58 ms on one Apple M5 Max core; a 2019-class core not yet - Changes over timeNone. A fixed design, unchanged for seven yearsAutomatic era draws and a reserve of instruction families that unlock by height. Nobody touches it + Changes over timeNone. A fixed design, unchanged for seven yearsA new program every hour, its memory pattern with it; era draws and reserved families on a schedule fixed at genesis. Nobody touches it Seed grindingNot applicable, the program comes from the hash inputClosed by a verifiable delay between seed and program Useful workNone. Hashing onlyThe same card proves every block and sells proofs to other chains Track recordNo chip publicly shipped in seven years, approximateZero years. Every number above is measured and logged with the commands that produced it. The specification, reference hash, test vectors and simulators are public now (github.com/igneum-network/spec). The node, the miner and the wallet are in a private repository until the public testnet From 4a7b72da9b56513083a6337bb37b31499185b696 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:51:26 +0000 Subject: [PATCH 042/131] Litepaper hardware line on the step schedule; status 20:51: public copy applied, level 1 and the limits bullet wait for the chip row --- docs/plans/counter-asic-2-status.md | 4 ++++ site/litepaper.html | 2 +- 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index dadbc3b71..06ad51d57 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -241,3 +241,7 @@ PC 2: spcurve-stopped-pc2-pv1b closed 20:47:46Z (the card alone: an empty shard ## 20:50 the measure lock holder is a prover measurement; a lock-status defect Correction to the 20:50 entry above: the measure lock has been held since 20:31Z by pid 45000, a `with-lock.sh measure` of igneum-wt-agg-cost's igneum-prove-host (the aggregation-cost agent's prover measurement, 18 minutes so far), not by the readwidth packbench; `with-lock.sh status` prints the LAST WRITER's command text, not the holder's, which is why it named the w4 run. Defect for the next cut (tools/lock/with-lock.sh: the status line must read the holder's pid and command, not the last writer's; the class of CLAUDE.md's watcher rule). Builds proceed in the three build slots (build-0 taken at 20:49:51 after an 885-s wait); GPU measurements (the mixer's verifier and build timings, the epoch compile-ahead, the Mac bit-exactness runs under `run` are not blocked) queue behind the prover measurement. Readwidth head is 30ff674 (per-watt rows for the consequences reviewer; e752fc7 stays the table commit). + +## 20:51 the number-free public copy is applied on ca2-coord (0ad70ba and the next commit) + +site/litepaper.html: the Mining section's "bound by memory bandwidth" is corrected to "waits on memory latency, not on maths or bandwidth"; the "Everything above is automatic" paragraph is replaced by the level 2 three ideas and the fourth paragraph with a link to the numbers page; the vs RandomX rows "Changes over time" and "Dataset" and the Hardware paragraph carry the step schedule (years 4, 12, 28; 4 GB about four years, 8 GB about twelve). site/index.html: the Memory row and the Mine card carry the step schedule ("2 GB at genesis, doubling at years 4, 12 and 28"; "any 4 GB card at launch, 8 GB from year 4"). NOT yet applied, because they carry the chip number: the level 1 sentence in the hero and the abstract ("a custom chip gains under 2x") and the limits bullet "A chip is impossible"; they wait for the combined chip row from docs/analysis/chip-model-v3.md (if 1.8x: "under 2x" with the margin stated as thin; else qualified). The bench page's Counter ASIC section (level 3) waits for the final table. docs/evidence.md's card-lifetime row (designed) is still to add. diff --git a/site/litepaper.html b/site/litepaper.html index 533f17c63..ce7b0620b 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -558,7 +558,7 @@ body.all .pager{display:none}

The honest bear-market case rests on cost. A miner's card is already running and the power is often domestic, so Igneum miners' marginal cost in the proving market is close to power, which is an edge over data-centre provers and nothing more.

Hardware

-

The dataset starts at 2 GB and grows by half a gigabyte a year, so a 4 GB card mines for about four years and an 8 GB card for more than a decade, approximate. 12 GB or more proves full shards. NVIDIA and AMD both work, because the mining program is generated for the architecture both share and the proof system is hash-based. Apple's chips are GPUs with unified memory, so Macs mine too, at about a fifth of a flagship card: Measured, 26.7 against 123 million hashes a second, an Apple M5 Max beside an RTX 5090 on the live devnet, 4 October 2026. A Mac is a poor miner per dollar. There is no CPU mining lane, on purpose, because CPU mining is what botnets farm. Nodes, wallets and exchanges need no GPU at all.

+

The dataset starts at 2 GB and doubles on a step schedule fixed at genesis (years 4, 12 and 28, the average of half a gigabyte a year), so a 4 GB card mines for about four years and an 8 GB card for about twelve, approximate. 12 GB or more proves full shards. NVIDIA and AMD both work, because the mining program is generated for the architecture both share and the proof system is hash-based. Apple's chips are GPUs with unified memory, so Macs mine too, at about a fifth of a flagship card: Measured, 26.7 against 123 million hashes a second, an Apple M5 Max beside an RTX 5090 on the live devnet, 4 October 2026. A Mac is a poor miner per dollar. There is no CPU mining lane, on purpose, because CPU mining is what botnets farm. Nodes, wallets and exchanges need no GPU at all.

What a miner's hour looks like

The card hashes the lottery continuously. When the client sees a shard or an external job it can win, it switches the card to proving for a few seconds, posts the proof, and goes back to hashing. The client does the switching and the miner sees one balance.

The protocol carries no fee: no dev fund, no cut to any team. Ember, the miner software, takes an optional 1% dev fee, the way other GPU miners do. One block template in 100 is requested with the dev address instead of yours, by a counter, not a random draw, so it is exactly 1 in 100 and anyone can check it from the source or from the chain. One flag turns it off (--dev-fee 0, a switch in the app, a line in the HiveOS config). The miner prints the fee and the address when it starts. Any other client is welcome.

From b8f8105af4fa791dafec0b91ca04ccdb16b03434 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:51:38 +0000 Subject: [PATCH 043/131] evidence.md: the card-lifetime claim as a designed row under the step schedule --- docs/evidence.md | 1 + 1 file changed, 1 insertion(+) diff --git a/docs/evidence.md b/docs/evidence.md index 89e002fe8..a00304a06 100644 --- a/docs/evidence.md +++ b/docs/evidence.md @@ -87,6 +87,7 @@ Versions in the table: `igneum-pow` is the Rust crate at `igneum-pow/Cargo.toml` |---|---|---|---| | 15 | implemented | implemented, with a live result | the first non-empty shard (block 72704, 29 transfers) proven, verified and paid on the devnet; not every block is proven yet | | 21 | designed | tested by the team | 388 shards paid from the pool on the live devnet, the rule in `proving.rs`, the numbers in the bench log | +| 22 | Card lifetime: a 4 GB card mines about four years and an 8 GB card about twelve, under the dataset's step schedule (2 GB at genesis, doubling at years 4, 12, 28, 60) with the cache freed after the daily build | Litepaper Hardware and vs RandomX ("Dataset" row); homepage Mine card and "Memory" row | designed | `docs/analysis/card-lifetime-2026-10-05.md` (branch card-lifetime 1fecfe2); spec 1.13.3 option (b) recommended to the project lead 5 October 2026 (`docs/plans/counter-asic-2-rollout.md` 6c) | The per-tier working-set arithmetic of that document (GTX 1650, RTX 3050, RTX 3060, RTX 4090 tiers) against the step schedule | A design claim: under the continuous mapping (a) a 4 GB card is out within 1 to 1.5 years and an 8 GB card at 6 to 7.5 years, so the sentence is true only under the step schedule (b), which the spec has not yet fixed (O-1.13) | none yet | ## What would move a row From cf0839a231a82bd7cc3f0da50d692ea3704a4b4f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:51:52 +0000 Subject: [PATCH 044/131] Counter ASIC 2.0 status 20:51: chain_dataset_day wired in the node, evidence row 22 --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 06ad51d57..f6344a276 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -245,3 +245,7 @@ Correction to the 20:50 entry above: the measure lock has been held since 20:31Z ## 20:51 the number-free public copy is applied on ca2-coord (0ad70ba and the next commit) site/litepaper.html: the Mining section's "bound by memory bandwidth" is corrected to "waits on memory latency, not on maths or bandwidth"; the "Everything above is automatic" paragraph is replaced by the level 2 three ideas and the fourth paragraph with a link to the numbers page; the vs RandomX rows "Changes over time" and "Dataset" and the Hardware paragraph carry the step schedule (years 4, 12, 28; 4 GB about four years, 8 GB about twelve). site/index.html: the Memory row and the Mine card carry the step schedule ("2 GB at genesis, doubling at years 4, 12 and 28"; "any 4 GB card at launch, 8 GB from year 4"). NOT yet applied, because they carry the chip number: the level 1 sentence in the hero and the abstract ("a custom chip gains under 2x") and the limits bullet "A chip is impossible"; they wait for the combined chip row from docs/analysis/chip-model-v3.md (if 1.8x: "under 2x" with the margin stated as thin; else qualified). The bench page's Counter ASIC section (level 3) waits for the final table. docs/evidence.md's card-lifetime row (designed) is still to add. + +## 20:51 the node builds every day cache through chain_dataset_day + +ca2-v3 6c75dad (node agent): Epoch::chain_dataset_day(day_bytes, class, days_since_genesis, genesis_dataset_log2) with a placeholder body (the mixer branch fills growth_doublings under that signature), verify::days_since_genesis; the Metal worker takes v3 from a pack (compiled). Fork (uncommitted, in the cargo check holding build-1 since 20:50Z): Params::pow_genesis_dataset_log2 (28 on every network, in the digest in its own statement), Params::genesis_day_index(), install_pow_genesis in the daemon after the class switch, the engine's build_day through chain_dataset_day, the pack export through the same build_day, PowEpochInfo / RPC / proto fields 17 and 18 (genesis_day_index, genesis_dataset_log2; an old node's 0 reads as 28), override-60x.json and the redteam override carry pow_genesis_dataset_log2 28. Next: the check result, then "ready for PC 2" with the fork commit. evidence.md row 22 (card lifetime, designed) added on ca2-coord (3dc29fe). From ac0d7c1592c73791170850542f710d5db589d883 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:52:58 +0000 Subject: [PATCH 045/131] Spec 01 and 04: layer 9 epoch_len (1.12, 1.13.1, 4.3), reserve family R1 and the emulation rule (1.13.2), the cache growth rule option C and the step mapping (1.13.3), the mixer_mult row --- docs/spec/01-lottery-hash.md | 20 ++++++++++++++++---- docs/spec/04-seeds-and-vdf.md | 2 +- 2 files changed, 17 insertions(+), 5 deletions(-) diff --git a/docs/spec/01-lottery-hash.md b/docs/spec/01-lottery-hash.md index e9680184c..34743c41e 100644 --- a/docs/spec/01-lottery-hash.md +++ b/docs/spec/01-lottery-hash.md @@ -396,11 +396,11 @@ All times are DAA seconds since genesis (section 0.6). At 1 block per second one | Clock | Length | What changes | Label | |---|---|---|---| -| Epoch | 3,600 DAA s | The program: new seed words from the VDF of section 4, new kernel | Designed (design document, "Always evolving, on three clocks"); the epoch length is a prototype value, to be fixed at gate 2 by the difficulty-tracking measurement (fork map c1: the hash rate steps by program, 35 to 48 Mhash/s across seeds on the M5 Max, so the DAA window must track within an epoch) | +| Epoch | `epoch_len(d)` DAA s, base 3,600; the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200; set by 90% miner signal at a day boundary (sections 1.13.1 and 5.7) | The program: new seed words from the VDF of section 4, new kernel | Designed (Counter ASIC 2.0 layer 9, 5 October 2026, `docs/plans/epoch-length.md`); 3,600 stays the value on every network until a signal moves it, and stays the prototype value to be fixed at gate 2 by the difficulty-tracking measurement (fork map c1: the hash rate steps by program, 35 to 48 Mhash/s across seeds on the M5 Max, so the DAA window must track within an epoch) | | Day | 86,400 DAA s | The day key, hence the cache and the dataset | Designed | | Era | 15,552,000 DAA s (180 days) | Era parameters and one instruction-family unlock, section 1.13 | Designed; the length is a prototype value (the design says "every 6 months") | -Epoch `e` covers DAA scores `[3,600 e, 3,600 (e + 1))`. The epoch of a block is the epoch of its own DAA score, so "which program was this block mined under" is a function of the header alone once the seed is known. The program for epoch `e` is `generate_from_seed_bytes(program_seed_e)`: attempt 0 is drawn from `S_e = seed_words_from_bytes(program_seed_e)`, and a rejected attempt is replaced as 1.4.6 says; `program_seed_e` is the 32-byte VDF output of section 4.3. +Epoch `(d, e)` covers DAA scores `[86,400 d + L e, 86,400 d + L (e + 1))` with `L = epoch_len(d)` and `e` in `0 .. 86,400 / L`; every ladder step divides 86,400, so day boundaries are epoch boundaries, and the epoch is identified by its start score `s = 86,400 d + L e`. At the base `L = 3,600` this is `[3,600 e, 3,600 (e + 1))` and nothing below differs from the earlier text. The epoch of a block is the epoch of its own DAA score, and `epoch_len(d)` is a function of the blue blocks of the signalling window that closed at least 2 days before day `d` (section 1.13.1), which are in the header's past, so "which program was this block mined under" is a function of the header alone once the seed is known. `T_epoch` and the 1,200-s lead of section 4.3 are genesis constants and do not follow `epoch_len`: the program of every epoch is known 600 s before it starts on the reference core at every length. The program for epoch `e` is `generate_from_seed_bytes(program_seed_e)`: attempt 0 is drawn from `S_e = seed_words_from_bytes(program_seed_e)`, and a rejected attempt is replaced as 1.4.6 says; `program_seed_e` is the 32-byte VDF output of section 4.3. Implementation note (devnet, 3 October 2026, `docs/fork-divergence.md` "Epoch seed"): until the VDF of section 4 is in the node, `program_seed_e` is the hash of the last selected-chain block whose DAA score is below `3,600 e - 600`. The 600-DAA-score lead stands in for section 4.3's 20-minute lead: the program of epoch `e` is knowable about 10 minutes before it starts, every block template reports it (`pow_epoch.next_epoch_seed`), and a GPU worker compiles it in the background and swaps at the boundary with no pause (serve protocol `prepare`, `proto-metal/main.swift`, `proto-cuda/host.cu`, `proto-opencl/host.c`). Measured across boundaries on a short-epoch test network in `docs/bench-log.md` (hot-swap entry). The program schedule is a protocol constant; a miner that cannot compile ahead sees the same seed at the same time as everyone else, only later. @@ -422,12 +422,24 @@ The era seed `E_n` is the 32-byte output of the 1-hour VDF of section 4.4. One S | Load count | 16 of 64 | not drawn | fixed, so every era is equally memory-bound | | Output fold rotations | (7, 14, 21), (9, 18, 27) | each `1 + below(31)` | 1..31 | | Mixer round count | 8 | not drawn | fixed, so the verify budget holds | +| Mixer applications per round `mixer_mult` | 4 (class v3; 1 under class v2) | not drawn | fixed at genesis (Counter ASIC 2.0, 5 October 2026: the M16 recompute chip's only measured lever; section 1.8.5 carries the form; `docs/plans/mixer-x4.md`) | +| Epoch length `epoch_len` | 3,600 DAA s | one draw of the era stream consumed and not used (the value is set by signal) | the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 (layer 9, `docs/plans/epoch-length.md`) | + +`epoch_len` is the one era-table parameter set by miners rather than by the draw: 90% of blue blocks over a 7-day window carrying the same ladder index (3 bits of the header version, encoding Open in section 5.8) sets that length from the first day boundary at least 2 days after the window closes (section 5.7). It is not a code upgrade: the rule, the ladder and the window are genesis constants, and the chain carries no release. The era stream consumes its draw so that a future draw of this parameter changes no other parameter's value. The threat it answers is a per-program hard datapath (an FPGA fleet: 42 to 160 minutes per compile on a mid-size part, PRflow, FPT 2019, hours on large parts; at 600 s nothing it compiles ever runs); it does not answer a programmable chip, which the other layers answer. The floor 600 is set by the slowest compile-ahead measured (the Metal variant race, 38 s on the M5 Max, 6.3% of a 600-s epoch and inside the 600-s seed window; `docs/plans/epoch-length.md` section 6). + +The table layout and the working-set window (Counter ASIC 2.0 layers 4 and 8) are drawn by the same stream; their rows and the draw order are in `docs/plans/era-layout.md` and enter this table with class v3's vectors (pending its six-era measurement, 5 October 2026). "Memory pattern" in the design document is read here as the item-address pattern (the cache line index word, `s[0]` in 1.8.5, and the XOR-all-sixteen rule); the proposal is to leave it fixed at era 0 and let the unlocked families change the kernel instead, because every change to the item derivation changes the verify time and must be re-measured. ### 1.13.2 Instruction-family reserve -At genesis the generator carries the eleven families of 1.4.1 live and a reserve list of further families in a fixed order. At the start of era `n >= 1`, reserve family `n` becomes live with weight `W_new` taken proportionally from the live non-load families. A family may enter the reserve only if it is integer-exact and has passed the cross-vendor conformance of section 1.15 on every vendor in the benchmark (Metal, CUDA, OpenCL on NVIDIA and AMD), with its own edge-case vectors, before genesis. Candidate families, all integer ALU operations present on Apple, NVIDIA and AMD: variable left shift and logical right shift by `src AND 31`; bit-field extract with an immediate offset and width; `andn` (`dst = dst AND NOT src`); byte permute of `dst` by an immediate selector; population count and count-leading-zeros folded into `dst` by add; a three-register select (`dst = bit of src2 ? src : dst`); a second shuffle form (`lane + delta mod 32`). The order and `W_new` are Open. A family that is not in the genesis reserve can only be added by the upgrade path of section 5.7. +At genesis the generator carries the eleven families of 1.4.1 live and a reserve list of further families in a fixed order. At the start of era `n >= 1`, reserve family `n` becomes live with weight `W_new` taken proportionally from the live non-load families. A family may enter the reserve only if it is integer-exact and has passed the cross-vendor conformance of section 1.15 on every vendor in the benchmark (Metal, CUDA, OpenCL on NVIDIA and AMD), with its own edge-case vectors, before genesis. Candidate families, all integer ALU operations present on Apple, NVIDIA and AMD: variable left shift and logical right shift by `src AND 31`; bit-field extract with an immediate offset and width; `andn` (`dst = dst AND NOT src`); byte permute of `dst` by an immediate selector; population count and count-leading-zeros folded into `dst` by add; a three-register select (`dst = bit of src2 ? src : dst`); a second shuffle form (`lane + delta mod 32`). The order and `W_new` are Open, except the first entry, decided 5 October 2026 (Counter ASIC 2.0 layer 7, delegated; the project lead confirms for the public testnet genesis; `docs/analysis/int8-matrix-family.md`): + +> Reserve family R1, `mm8` (integer matrix). Semantics: section 2.2 of `docs/analysis/int8-matrix-family.md`, uint8 operands from `src` and `src2` in the m8n8k16 fragment layout, one int32 element of C per lane selected by the immediate `bit`, added into `dst` modulo 2^32. Weight at unlock `W_new = 4` points, taken proportionally from the ten live non-load families (the load weight and count are untouched). Edge vectors, each a hand-built unit run on every vendor: all bytes 0xFF in A and B (C = 1,040,400 everywhere); all bytes 0x80 (C = 262,144); A all zero (C = 0); `dst` = 0xFFFFFFFF with a nonzero C (the wrap); alternating 0x00 and 0xFF by lane; `bit` = 0 and 1 on the same fragments. Unlock: at the start of era n = 4 (DAA 62,208,000), or earlier by the 90% signalling path of section 5.7; never by a release. Native paths: PTX `mma.sync` `.u8` (sm_75+), AMD WMMA `i32_16x16x16_iu8` (RDNA 3 and 4), Metal 4 `mpp::tensor_ops::matmul2d` (`uchar x uchar -> int`); the per-lane `dot4` form is emulation on Apple (1.6x per op unsigned, measured 5 October 2026) and is not the reserved form. + +A vendor that can only emulate. A family enters the reserve when it is bit-exact on every vendor of 1.15. A vendor that reaches the result only by emulation (no instruction or library path) does not block entry if the measured penalty of the emulation on that vendor, on the family's own probe (a dependent chain of the op against the same vendor's integer ALU chain), is at most 8x per op, AND the family's weight at unlock keeps the emulating vendor's hash-rate loss under 5% on the memory-hard hash, checked on the vendor's card with the family live. A family whose emulation exceeds either bound stays out of the reserve until the vendor ships a path. + +A family that is not in the genesis reserve can only be added by the upgrade path of section 5.7. ### 1.13.3 Dataset growth @@ -437,7 +449,7 @@ Designed: 2 GiB at genesis plus 0.5 GiB per year (design document, "Which cards N_d = floor((2 GiB + 0.5 GiB * (86,400 d / 31,536,000)) / 64 bytes) ``` -evaluated in integers (bytes), with one year = 31,536,000 DAA seconds. The dataset grows by about 23 KiB per day and is recomputed with the day key. Two consequences are Open: +evaluated in integers (bytes), with one year = 31,536,000 DAA seconds. The dataset grows by about 23 KiB per day on this average and is recomputed with the day key. Decided 5 October 2026 (Counter ASIC 2.0 layer 6, delegated; the project lead confirms for the public testnet genesis): the cache grows with the dataset, doubling when the dataset doubles: `cache_log2_words(d) = 26 + growth_doublings(d)`, `growth_doublings(d) = floor(log2(1 + d / 1,460))` for day `d` since genesis (doublings at years 4, 12, 28 and 60), so the cache is 256 MiB at genesis, 512 MiB from year 4, 1 GiB from year 12; the verifier's one-core fill is 0.2, 0.4 and 0.8 s at those steps (0.2 s per 256 MiB, section 1.12), under 1 s at every step of the schedule. The dataset steps to the next power of two on the same doublings (option (b) below, recommended to the project lead with the card-lifetime consequences in `docs/analysis/card-lifetime-2026-10-05.md`: a 4 GB card mines to year 4, an 8 GB card to year 12, a 12 GB card to year 28 with the cache freed after the daily build). Why the cache grows at all: an SRAM mirror of a flat 256 MiB cache is about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density (AMD V-Cache, 64 MB on 41 mm^2 at 7 nm; `docs/analysis/sram-mirror.md`), so the cache size never prices a chip out; its job is to stay above any GPU's on-die cache (96 MB on the RTX 5090, 128 MB on GB202), which a flat 256 MiB loses within the decade. The GPU frees the cache after the daily dataset build; the hash never reads it. Two consequences remain Open: - Index mapping. `src AND MASK` requires a power-of-two size. For a non-power-of-two `N_d` the proposed mapping is `idx = (src * N_words) >> 32` computed in 64 bits (a multiply-shift range reduction; uniform to within 2^-32, branch-free, integer only). At `N_words = 2^28` this gives `src >> 4`, not `src AND MASK`, so adopting it changes the 1 GiB vectors; gate 1 chooses between (a) the multiply-shift mapping with new vectors, or (b) power-of-two sizes only, growing in steps (2 GiB, 4 GiB) on the same schedule's average, which keeps `AND MASK` and means a 4 GiB card lasts until the 4 GiB step instead of fading. - The item index `t` is 32 bits, so the construction as written tops out at 2^32 items = 256 GiB, which the schedule reaches after 508 years. No action needed. diff --git a/docs/spec/04-seeds-and-vdf.md b/docs/spec/04-seeds-and-vdf.md index 31204de18..b62d9c92c 100644 --- a/docs/spec/04-seeds-and-vdf.md +++ b/docs/spec/04-seeds-and-vdf.md @@ -49,7 +49,7 @@ Designed. Epoch e is the DAA-score interval `[3,600 e, 3,600 (e + 1))` (section 1. **Seed checkpoint.** `C(e)` is the highest-index checkpoint (section 3, C1) whose checkpoint block has DAA score at most `3,600 e - 1,200`: the latest checkpoint at least 20 minutes of DAA time before the epoch starts. The checkpoint block hash is Kaspa's full header hash, which covers the nonce, as the grinding defence requires (`proto-vdf/README.md`: a hash that covers only the body would let a miner start the VDF while still searching nonces). 2. **Evaluation.** `input = hash(C(e))`, `T = T_epoch`, run 4.2. `program_seed_e` is the 32-byte output; `proof_e` is (T, y, pi). -3. **Program.** `S_e = seed_words_from_bytes(program_seed_e)` (section 1.3.1), program = `generate_from_words(S_e)` (section 1.4). The 1,200-s lead is 2x the reference evaluation time, so a core half as fast as the reference still finishes before the epoch (4.6). +3. **Program.** `S_e = seed_words_from_bytes(program_seed_e)` (section 1.3.1), program = `generate_from_words(S_e)` (section 1.4). The 1,200-s lead is 2x the reference evaluation time, so a core half as fast as the reference still finishes before the epoch (4.6). The lead and `T_epoch` are fixed whatever the epoch length of section 1.12: at the floor of 600 DAA s the checkpoint is two epochs back and the program is known one full epoch ahead; at the base it is known for the last sixth of the previous epoch (`docs/plans/epoch-length.md`, section 3). 4. **Header.** Every header carries `seed_source = hash(C(e))` for its own epoch (section 2.4). A header is valid under the lottery only if its `seed_source` is a block on its own selected chain at the blue score of checkpoint index `i(C(e))`, and its PoW verifies under the program derived from that block. The proof is not in the header: a node verifies `proof_e` once per epoch (4.5 ms) and caches `S_e`. Determinism of step 1 is the point of the rule: which block is "the checkpoint at blue score 30 i" is a function of the header's own past, so two nodes validating the same header derive the same program, and a header mined under a reorged-away checkpoint names a block that is not on its chain and is invalid. Whether `C(e)` must be certified (section 3) or merely be the selected-chain block at that blue score in the header's past is Open (O-4.3): requiring certification couples mining to finality liveness (a stall longer than the lead would stop the program from being derivable), which the design document accepts ("an epoch cannot start without a valid proof") and this specification argues against, because the chain is meant to keep running on plain GHOSTDAG through a finality pause (section 3.7 item 2). The proposal: the selected-chain block at that blue score, certified or not, deep enough (1,200 DAA s plus d) that a reorg across it is a merge-depth-scale event. From 9eee20f97c948ce9ade8ec511bea01448afcc8c8 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:53:12 +0000 Subject: [PATCH 046/131] Counter ASIC 2.0 status 20:53: spec text applied for layers 6, 7, 9; what still lands from the branches --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index f6344a276..0f7a1429e 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -249,3 +249,7 @@ site/litepaper.html: the Mining section's "bound by memory bandwidth" is correct ## 20:51 the node builds every day cache through chain_dataset_day ca2-v3 6c75dad (node agent): Epoch::chain_dataset_day(day_bytes, class, days_since_genesis, genesis_dataset_log2) with a placeholder body (the mixer branch fills growth_doublings under that signature), verify::days_since_genesis; the Metal worker takes v3 from a pack (compiled). Fork (uncommitted, in the cargo check holding build-1 since 20:50Z): Params::pow_genesis_dataset_log2 (28 on every network, in the digest in its own statement), Params::genesis_day_index(), install_pow_genesis in the daemon after the class switch, the engine's build_day through chain_dataset_day, the pack export through the same build_day, PowEpochInfo / RPC / proto fields 17 and 18 (genesis_day_index, genesis_dataset_log2; an old node's 0 reads as 28), override-60x.json and the redteam override carry pow_genesis_dataset_log2 28. Next: the check result, then "ready for PC 2" with the fork commit. evidence.md row 22 (card lifetime, designed) added on ca2-coord (3dc29fe). + +## 20:53 spec text applied for the decided layers (ca2-coord 9b1f849) + +docs/spec/01-lottery-hash.md: 1.12 carries epoch_len (the ladder 600 to 7,200, 90% signal at a day boundary, T_epoch and the lead fixed) with 3,600 unchanged on every network; 1.13.1 gains the mixer_mult row (4 under class v3) and the epoch_len row with the signal rule, the FPGA threat and the 600-s floor, and a pointer to era-layout.md for the layer 4 and 8 rows; 1.13.2 carries reserve family R1 = mm8 (uint8, W_new 4, unlock era 4 or 90% signal, the edge vectors, the native paths) and the emulation rule (8x per op, 5% hash-rate cap); 1.13.3 carries the cache growth rule (option C, growth_doublings(d) = floor(log2(1 + d / 1,460)), 256 / 512 / 1,024 MiB at genesis / year 4 / year 12, fill 0.2 / 0.4 / 0.8 s), the step mapping (b) as recommended, the shipped-density reason, and the cache freed after the daily build. docs/spec/04-seeds-and-vdf.md 4.3: the lead and T_epoch fixed at every epoch length. Still to land in the spec from the branches: 1.8.5 (the mixer x4 form, from mixer-x4.md), the 1.13.1 rows for stride, interleave and the window (era-layout.md), 1.5 and 1.8 for the hot table in the added form (hot-table.md), 1.17 and 1.15 for the v3 vectors and the conformance runs, 1.4.5 and 1.4.6 for generator 3 and the class in the pack. From f2c99cff9576c0538e57bf241046932338e7256c Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:53:26 +0000 Subject: [PATCH 047/131] Spec 01 1.4.5 and 1.4.6: generator 3, the class and era seed in the pack and on the job line, the refusal rule --- docs/spec/01-lottery-hash.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/spec/01-lottery-hash.md b/docs/spec/01-lottery-hash.md index 34743c41e..cb87e61ff 100644 --- a/docs/spec/01-lottery-hash.md +++ b/docs/spec/01-lottery-hash.md @@ -164,7 +164,7 @@ Every emitted instruction satisfies: `rot` in 1..31, `mask` in {1, 2, 4, 8, 16}, ### 1.4.5 Encoding -A program is transmitted as the seed bytes, never as instructions. A node hands a miner the pack it emits itself (`igneum-pow/src/emit.rs`: `kernel.cu`, `kernel.cl`, `program.metal`, `program.h`, `program.json`, `memhard.h`, `vectors.*`, the three `*_bound` kernels), and a miner MAY regenerate everything from the seed bytes by the procedure of 1.4.6. `program.json` (format `igneum-program-pack-3`) is the interchange form; its field names are those of `Instr` in `generator.rs`, and it carries `generator` (2), `attempt`, `program_id` and `seed_bytes`. `program.h` carries the same as `IGNEUM_GENERATOR`, `IGNEUM_PROGRAM_ATTEMPT`, `IGNEUM_PROGRAM_ID` and `IGNEUM_SEED_BYTES_HEX`. An implementation MUST refuse a pack whose generator version is not its own. +A program is transmitted as the seed bytes, never as instructions. A node hands a miner the pack it emits itself (`igneum-pow/src/emit.rs`: `kernel.cu`, `kernel.cl`, `program.metal`, `program.h`, `program.json`, `memhard.h`, `vectors.*`, the three `*_bound` kernels), and a miner MAY regenerate everything from the seed bytes by the procedure of 1.4.6. `program.json` (format `igneum-program-pack-3`) is the interchange form; its field names are those of `Instr` in `generator.rs`, and it carries `generator` (2), `attempt`, `program_id` and `seed_bytes`. `program.h` carries the same as `IGNEUM_GENERATOR`, `IGNEUM_PROGRAM_ATTEMPT`, `IGNEUM_PROGRAM_ID` and `IGNEUM_SEED_BYTES_HEX`. An implementation MUST refuse a pack whose generator version is not its own. Program class v3 (Counter ASIC 2.0, 5 October 2026, activated by the height switch `program_class_v3_activation_daa` from the first epoch whose start score is at or above it, `docs/plans/counter-asic-2-rollout.md`) writes `generator` 3, and every pack of it carries `IGNEUM_PROGRAM_CLASS` (`v3`) and `IGNEUM_ERA_SEED_HEX` (the 32-byte era seed of section 1.13.1, or its devnet stand-in) beside `IGNEUM_GENERATOR`; the serve protocol's `prepare` and `job` lines carry `class=v3 era=` for v3 epochs and nothing for v2 ones. A worker MUST refuse a pack whose class or era seed does not match the line it was prepared for (`igneum-pow/src/packcheck.rs`, `verify_pack_dir_chain`; `proto-cuda/nvrtc/packfile.h`), and a pack of a generator other than 2 or 3. ### 1.4.6 Program acceptance @@ -178,7 +178,7 @@ Implemented (`igneum-pow/src/accept.rs`, `proto-metal/main.swift`; ledger M6 Fix Attempts. Attempt 0 of a program seed `b` (the 32-byte epoch seed, or the UTF-8 of a seed string) is the candidate drawn from `seed_words_from_bytes(b)`. If it fails, attempt `k = 1, 2, ...` is drawn from `seed_words_from_bytes(b || k_le32)`; the first accepted candidate is the program of the epoch. Measured rejection rate under this generator: 5.14 percent over 100,000 seeds (census section 7) and the 20,000-seed confirmation of `docs/bench-log.md` (4 October 2026), so the probability that 32 consecutive candidates fail is below 2^-136, and an implementation MAY treat 32 consecutive failures as a consensus fault (`MAX_ATTEMPTS`). -Program id. `FNV-1a-64("igneum-program/" || generator_le32 || seed words as little-endian bytes || attempt_le32)` with `generator = 2`, written into every pack. Two implementations that agree on the id agree on the generator version, the seed words and the attempt. +Program id. `FNV-1a-64("igneum-program/" || generator_le32 || seed words as little-endian bytes || attempt_le32)` with `generator = 2` under class v2 and `generator = 3` under class v3, written into every pack. Two implementations that agree on the id agree on the generator version, the seed words and the attempt. Why the closed form: the test is then a pure function of the program (no cache, no day), costs 1.3 to 3.4 ms on one core, and the census checked on 100,000 programs that its verdict agrees with the memory-hard dataset's on all but 39 threshold-edge cases (section 7.3). What the three parts catch: (a) the empty-list fallback of 1.4.3; (b) registers that saturate to all ones (2.4 percent of candidates); (c) zero-absorbing register sets, lane-constant load sites, output bias and value-level address repeats (2.1 percent). Not in the rule, and why: a contraction as the last write (80 percent of programs) and the `or` count are too common and (c) already catches the cases that matter; the load critical path is a hash-rate question, not a weakness. From 3e237db050105f49800dd0374c62ba03b62978cd Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:53:45 +0000 Subject: [PATCH 048/131] counter-asic-2.md: layer 9 row and the decisions table of 5 October 2026 (night) --- docs/plans/counter-asic-2.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/docs/plans/counter-asic-2.md b/docs/plans/counter-asic-2.md index f6598d19b..8e9f58d86 100644 --- a/docs/plans/counter-asic-2.md +++ b/docs/plans/counter-asic-2.md @@ -16,6 +16,21 @@ Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGP | 6 | Cache growth on the genesis schedule | the SRAM mirror stays unaffordable | none | already in the design; confirm the schedule against SRAM density | | 7 | Integer matrix ops in the program (INT8 x INT8 into INT32, exact) | matrix hardware at GPU scale | none on NVIDIA and AMD; Apple to check | reserved family, not at launch | | 8 | Working-set size drawn per program | one memory design cannot fit every hour | none | folds into 4 and 5 | +| 9 | Epoch length as a signalled era parameter: base 3,600 DAA s, ladder 600 to 7,200, set by 90% signal at a day boundary, lead and `T_epoch` fixed | a per-program bitstream (an FPGA with a hard datapath): at 600 s nothing it compiles ever runs (42 to 160 min per compile, PRflow FPT 2019; hours on large parts, Aldec) | compile-ahead 1 s per epoch on the 5090, 0.5 s on the M5 Max (38 s with the race on); one CPU core `600 / epoch_len` busy on the VDF | reserve-only tonight: `docs/plans/epoch-length.md` | + +## Decided 5 October 2026 (night), under the project lead's delegation for the devnet (`docs/plans/counter-asic-2-rollout.md` section 6) + +| # | Decision | The number that decided it | +|---|---|---| +| 1 | keep v2's 128 x 4 B | w16 passes the rule but closes nothing (5090 139.8 against 136.1 MH/s, 9070 XT 17.90 against 18.15); w64 makes the 5090 bandwidth-bound (share 0.58, 37% of stream) | +| 2 | out | spread across six programs 5.5 to 22.3% per card, over the 5% rule | +| 3 | out (scratch share 0); the construct is sound and its tests stay | the on-die-cache recompute chip stays at 2.4x at every share under the 6 GB cap | +| 4 and 8 | the era draw of stride, interleave and the working-set window, in if the six-era spread is under 5% per card | pending the PC rows | +| 5 | in, in the ADDED form only (16 dataset loads plus k hot loads), size = the largest table resident on every card | the replaced form lets the chip skip item derivations (x1.33 at k = 4); Mac rows hot32k4 x1.22, hot64k4 x1.12 | +| 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 | +| 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple | +| M16 mixer x4 | into v3 | the only measured lever that moves the named chip: 0.6x bare, 1.8x with a 3x fixed-function factor; verifier 1.6 to 4.8 ms per warp | +| 9 | reserve-only, no change to the devnet's hour | the floor 600 s from the slowest compile-ahead (38 s) | Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors). From 88740ce75e8bded15d7286969b01093190ac89de Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:56:23 +0000 Subject: [PATCH 049/131] Counter ASIC 2.0 status 20:56: PC 2 released, the proving fork tip, the S_p curve with the miner on --- docs/plans/counter-asic-2-status.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 0f7a1429e..9f3a14000 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -253,3 +253,9 @@ ca2-v3 6c75dad (node agent): Epoch::chain_dataset_day(day_bytes, class, days_sin ## 20:53 spec text applied for the decided layers (ca2-coord 9b1f849) docs/spec/01-lottery-hash.md: 1.12 carries epoch_len (the ladder 600 to 7,200, 90% signal at a day boundary, T_epoch and the lead fixed) with 3,600 unchanged on every network; 1.13.1 gains the mixer_mult row (4 under class v3) and the epoch_len row with the signal rule, the FPGA threat and the 600-s floor, and a pointer to era-layout.md for the layer 4 and 8 rows; 1.13.2 carries reserve family R1 = mm8 (uint8, W_new 4, unlock era 4 or 90% signal, the edge vectors, the native paths) and the emulation rule (8x per op, 5% hash-rate cap); 1.13.3 carries the cache growth rule (option C, growth_doublings(d) = floor(log2(1 + d / 1,460)), 256 / 512 / 1,024 MiB at genesis / year 4 / year 12, fill 0.2 / 0.4 / 0.8 s), the step mapping (b) as recommended, the shipped-density reason, and the cache freed after the daily build. docs/spec/04-seeds-and-vdf.md 4.3: the lead and T_epoch fixed at every epoch length. Still to land in the spec from the branches: 1.8.5 (the mixer x4 form, from mixer-x4.md), the 1.13.1 rows for stride, interleave and the window (era-layout.md), 1.5 and 1.8 for the hot table in the added form (hot-table.md), 1.17 and 1.15 for the v3 vectors and the conformance runs, 1.4.5 and 1.4.6 for generator 3 and the class in the pack. + +## 20:56 PC 2 released; the proving fork tip for the merged tree + +PC 2 is free (the proving agent's last job closed; the live prover is back on). Proving fork tip ece42979 on 21d4c73c (N = 8, the digest test edit), app 440fd59 or later; its suites ride with the ca2 node suites on the merged tree; its harness on the final tree is running on the Mac. The S_p curve with the miner on the card (PC 2's 5090): empty shard 15,585 MiB 7.5 s; the adopted v1 shard (30,000 pgas, 4.7 M cycles) 22,210 MiB 13.2 s (20,435 MiB, 4.2 s alone); 2.25 M pgas 30,049 MiB 17.9 s; 4.5 M 29,954 MiB 26.3 s; the prototype shard 30,083 MiB 33.3 s. Tiers as the proving agent published them: 32 GB mines and proves today, 24 GB from the fee switch (2.3 GB spare on the adopted shard), 16 GB empty shards only, 12 GB nothing on this build (D2 to the project lead). + +PC 2 queue: the ca2 node suites on the merged tree (ca2-v3-node + ece42979) as soon as the node agent sends "ready for PC 2"; nothing else is queued on PC 2. From b4e94a482d65946d1ed7cc80265a3244f489c18c Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:57:12 +0000 Subject: [PATCH 050/131] Counter ASIC 2.0 status 20:57: the mixer construction committed, MX4 is V3_CLASS --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 9f3a14000..b0133a07f 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -259,3 +259,7 @@ docs/spec/01-lottery-hash.md: 1.12 carries epoch_len (the ladder 600 to 7,200, 9 PC 2 is free (the proving agent's last job closed; the live prover is back on). Proving fork tip ece42979 on 21d4c73c (N = 8, the digest test edit), app 440fd59 or later; its suites ride with the ca2 node suites on the merged tree; its harness on the final tree is running on the Mac. The S_p curve with the miner on the card (PC 2's 5090): empty shard 15,585 MiB 7.5 s; the adopted v1 shard (30,000 pgas, 4.7 M cycles) 22,210 MiB 13.2 s (20,435 MiB, 4.2 s alone); 2.25 M pgas 30,049 MiB 17.9 s; 4.5 M 29,954 MiB 26.3 s; the prototype shard 30,083 MiB 33.3 s. Tiers as the proving agent published them: 32 GB mines and proves today, 24 GB from the fee switch (2.3 GB spare on the adopted shard), 16 GB empty shards only, 12 GB nothing on this build (D2 to the project lead). PC 2 queue: the ca2 node suites on the merged tree (ca2-v3-node + ece42979) as soon as the node agent sends "ready for PC 2"; nothing else is queued on PC 2. + +## 20:57 the mixer construction is committed on ca2-v3's base + +ca2-mixer 0fc0ad1 (rebased onto ca2-v3 6c75dad): LoadClass::MX4 = V3_CLASS (v2 loads, mixer x4, the growth rule; a class v3 program is the v2 program of its seed instruction for instruction, generator 3 in its id); chain_dataset_day has its real body (Shape::for_class_day: cache 2^cache_log2_words(d), dataset 2^dataset_log2_words(D_0, d)); 44 lib + 11 pack tests green, the two pinned v2 packs byte for byte. Next from it: --program-class v3 / --era-hex on the CLI, the pinned v3 packs mx4-genesis and mx4-devnet-epoch0 (era = the genesis-hash stand-in), the design and 1.8.5 spec text, the chip-model row; bit-exactness runs under the run lock now; the verifier and build timings wait for the measure lock (held by a live prover measurement from another worktree, pid 45000, 25 min in at 20:56). The node agent merges ca2-mixer before its fast-time gate, so the gate runs the real construction minus the era and hot fields. From 2abdb561946abc479d4b1e8b88a78b304e9fc562 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:58:57 +0000 Subject: [PATCH 051/131] Counter ASIC 2.0 status 20:58: PC 1 and PC 2 occupancy --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index b0133a07f..fc77168d6 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -263,3 +263,5 @@ PC 2 queue: the ca2 node suites on the merged tree (ca2-v3-node + ece42979) as s ## 20:57 the mixer construction is committed on ca2-v3's base ca2-mixer 0fc0ad1 (rebased onto ca2-v3 6c75dad): LoadClass::MX4 = V3_CLASS (v2 loads, mixer x4, the growth rule; a class v3 program is the v2 program of its seed instruction for instruction, generator 3 in its id); chain_dataset_day has its real body (Shape::for_class_day: cache 2^cache_log2_words(d), dataset 2^dataset_log2_words(D_0, d)); 44 lib + 11 pack tests green, the two pinned v2 packs byte for byte. Next from it: --program-class v3 / --era-hex on the CLI, the pinned v3 packs mx4-genesis and mx4-devnet-epoch0 (era = the genesis-hash stand-in), the design and 1.8.5 spec text, the chip-model row; bit-exactness runs under the run lock now; the verifier and build timings wait for the measure lock (held by a live prover measurement from another worktree, pid 45000, 25 min in at 20:56). The node agent merges ca2-mixer before its fast-time gate, so the gate runs the real construction minus the era and hot fields. + +20:58. PC 1: cpu-prove-pc1-small2 running since 20:54:02 (CPU only). PC 2: a fetch job from the aggregation-cost agent (job-fetch-prove-aggcost, 20:55:39) landed after the proving agent's release, so PC 2 is NOT idle for the ca2 suites until that agent's run closes; the suite publish checks `node tools/jobs.mjs status` for an idle PC 2 first. Readiness in hand: the repro benchmark (8 min, 5090 then gfx1036) waits for a gap; the 5090 power sweep (8 min) follows the era job; Ember Tune's build (8 min, CPU) and run (30 min, both cards) follow; the S_p CPU shard (up to 90 min) is last. From b93d1897ebd60094128ad242804ccd081b3ca402 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:00:58 +0000 Subject: [PATCH 052/131] Counter ASIC 2.0 status 21:00: PC 1 to the repro run, PC 2 to agg-cost, the CPU fallback numbers, the S_p shard dropped --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index fc77168d6..9cc162575 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -265,3 +265,7 @@ PC 2 queue: the ca2 node suites on the merged tree (ca2-v3-node + ece42979) as s ca2-mixer 0fc0ad1 (rebased onto ca2-v3 6c75dad): LoadClass::MX4 = V3_CLASS (v2 loads, mixer x4, the growth rule; a class v3 program is the v2 program of its seed instruction for instruction, generator 3 in its id); chain_dataset_day has its real body (Shape::for_class_day: cache 2^cache_log2_words(d), dataset 2^dataset_log2_words(D_0, d)); 44 lib + 11 pack tests green, the two pinned v2 packs byte for byte. Next from it: --program-class v3 / --era-hex on the CLI, the pinned v3 packs mx4-genesis and mx4-devnet-epoch0 (era = the genesis-hash stand-in), the design and 1.8.5 spec text, the chip-model row; bit-exactness runs under the run lock now; the verifier and build timings wait for the measure lock (held by a live prover measurement from another worktree, pid 45000, 25 min in at 20:56). The node agent merges ca2-mixer before its fast-time gate, so the gate runs the real construction minus the era and hot fields. 20:58. PC 1: cpu-prove-pc1-small2 running since 20:54:02 (CPU only). PC 2: a fetch job from the aggregation-cost agent (job-fetch-prove-aggcost, 20:55:39) landed after the proving agent's release, so PC 2 is NOT idle for the ca2 suites until that agent's run closes; the suite publish checks `node tools/jobs.mjs status` for an idle PC 2 first. Readiness in hand: the repro benchmark (8 min, 5090 then gfx1036) waits for a gap; the 5090 power sweep (8 min) follows the era job; Ember Tune's build (8 min, CPU) and run (30 min, both cards) follow; the S_p CPU shard (up to 90 min) is last. + +## 21:00 PC 1 released by the CPU fixture; the repro run has it; PC 2 to the aggregation-cost agent + +cpu-prove-pc1-small2 finished 20:59:49Z: the SP1 CPU prover on PC 1 with the miners running: block-56-transfers-3shards shard 0 (200 pgas) 312 s wall, peak RSS 29.5 GB, 978% CPU; block-78-increment 322 s, 30.5 GB; the 5090 untouched (89% mean). The S_p shard job is DROPPED tonight: 312 s for 315 k cycles extrapolates the 60.8 M-cycle shard to many hours of every core and over 30 GB (approximate), which answers the CPU-fallback question (not viable for S_p shards; viable for empty or tiny shards only). "go PC 1" given to the reproducible benchmark (the 5090 then the gfx1036, about 8 min); the era job follows it, then the 5090 power sweep, then the hot table, then Ember Tune's build and run. "go PC 2" given to the aggregation-cost agent (20 min, GPU proving with the miner on then paused, prover restored); then the ca2 node suites, then the repro run's PC 2 slot (10 min). From 64ab9ce8d03d1bc0675ff5bbde9cf831aed4a120 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:01:20 +0000 Subject: [PATCH 053/131] Counter ASIC 2.0 status 21:01: the 0.3.10 installer build takes PC 1 after the repro run --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 9cc162575..c00a63d7a 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -269,3 +269,5 @@ ca2-mixer 0fc0ad1 (rebased onto ca2-v3 6c75dad): LoadClass::MX4 = V3_CLASS (v2 l ## 21:00 PC 1 released by the CPU fixture; the repro run has it; PC 2 to the aggregation-cost agent cpu-prove-pc1-small2 finished 20:59:49Z: the SP1 CPU prover on PC 1 with the miners running: block-56-transfers-3shards shard 0 (200 pgas) 312 s wall, peak RSS 29.5 GB, 978% CPU; block-78-increment 322 s, 30.5 GB; the 5090 untouched (89% mean). The S_p shard job is DROPPED tonight: 312 s for 315 k cycles extrapolates the 60.8 M-cycle shard to many hours of every core and over 30 GB (approximate), which answers the CPU-fallback question (not viable for S_p shards; viable for empty or tiny shards only). "go PC 1" given to the reproducible benchmark (the 5090 then the gfx1036, about 8 min); the era job follows it, then the 5090 power sweep, then the hot table, then Ember Tune's build and run. "go PC 2" given to the aggregation-cost agent (20 min, GPU proving with the miner on then paused, prover restored); then the ca2 node suites, then the repro run's PC 2 slot (10 min). + +21:01. GitHub Actions is in a major outage (six queued runs since 19:26Z, none acquired); the coordinator gave the 0.3.10 shipper the fallback at 21:00Z: build the Windows installer on PC 1 (MSVC window host, the payload under Git Bash, Inno Setup; CPU only, about 15 min). PC 1 order now: the repro run (until about 21:09), then the 0.3.10 installer build (the fleet's release, ahead of every measurement), then the era job, the 5090 power sweep, the hot table, Ember Tune. Any measurement that straddles the build window is re-run. From 4366c6ad908233d0a1940af9ba4f940565aeee3b Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:02:13 +0000 Subject: [PATCH 054/131] Counter ASIC 2.0 status 21:02: the measure lock convoy, found by lsof --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index c00a63d7a..03a6b203d 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -271,3 +271,5 @@ ca2-mixer 0fc0ad1 (rebased onto ca2-v3 6c75dad): LoadClass::MX4 = V3_CLASS (v2 l cpu-prove-pc1-small2 finished 20:59:49Z: the SP1 CPU prover on PC 1 with the miners running: block-56-transfers-3shards shard 0 (200 pgas) 312 s wall, peak RSS 29.5 GB, 978% CPU; block-78-increment 322 s, 30.5 GB; the 5090 untouched (89% mean). The S_p shard job is DROPPED tonight: 312 s for 315 k cycles extrapolates the 60.8 M-cycle shard to many hours of every core and over 30 GB (approximate), which answers the CPU-fallback question (not viable for S_p shards; viable for empty or tiny shards only). "go PC 1" given to the reproducible benchmark (the 5090 then the gfx1036, about 8 min); the era job follows it, then the 5090 power sweep, then the hot table, then Ember Tune's build and run. "go PC 2" given to the aggregation-cost agent (20 min, GPU proving with the miner on then paused, prover restored); then the ca2 node suites, then the repro run's PC 2 slot (10 min). 21:01. GitHub Actions is in a major outage (six queued runs since 19:26Z, none acquired); the coordinator gave the 0.3.10 shipper the fallback at 21:00Z: build the Windows installer on PC 1 (MSVC window host, the payload under Git Bash, Inno Setup; CPU only, about 15 min). PC 1 order now: the repro run (until about 21:09), then the 0.3.10 installer build (the fleet's release, ahead of every measurement), then the era job, the 5090 power sweep, the hot table, Ember Tune. Any measurement that straddles the build window is re-run. + +21:02. The Mac measure lock, found by `lsof`: the recorded holder pid 43916 is dead; the files are held open by two WAITERS, the readwidth agent's re-queued footprint loop (pid 78893, holding the measure and build files 11 min, waiting for the three build slots, which cargo tests keep re-acquiring: a convoy) and the epoch agent's compile-ahead measurement (pid 78476, waiting behind it). The readwidth agent is asked to kill 78893; the epoch measurement then runs when the build slots drain. Two defects for the next cut (one task filed): the status line shows the last writer, not the holder; a measure waiter can hold the master lock while build slots keep being granted to new builds, so a measurement can wait indefinitely under a steady stream of cargo tests. From ffa2633845366cc05c2b5f6f48f969a227059ab7 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:02:59 +0000 Subject: [PATCH 055/131] docs/analysis/amd-proving.md: no zkVM proves on AMD (SP1, RISC Zero, Jolt, OpenVM, ICICLE cited), the SP1 CPU prover measured on PC 1 beside the miners (282 s a shard at any size, 30 GB RSS: no CPU tier), the tier consequences and the public line; PC 1 job scripts with the bash -n gate; bench-log entry Co-Authored-By: Claude Fable 5.1 --- docs/analysis/amd-proving.md | 131 +++++++++++++++++++++++++++ docs/bench-log.md | 16 ++++ tools/amd-prove/check-job-bash.sh | 26 ++++++ tools/amd-prove/pc1-cpu-prove-sp.ps1 | 81 +++++++++++++++++ tools/amd-prove/pc1-cpu-prove.ps1 | 81 +++++++++++++++++ 5 files changed, 335 insertions(+) create mode 100644 docs/analysis/amd-proving.md create mode 100755 tools/amd-prove/check-job-bash.sh create mode 100644 tools/amd-prove/pc1-cpu-prove-sp.ps1 create mode 100644 tools/amd-prove/pc1-cpu-prove.ps1 diff --git a/docs/analysis/amd-proving.md b/docs/analysis/amd-proving.md new file mode 100644 index 000000000..b64c94a1d --- /dev/null +++ b/docs/analysis/amd-proving.md @@ -0,0 +1,131 @@ +# Proving on AMD and Apple cards: what exists, what the CPU can do, what to tell the public + +5 October 2026, from the project lead's two questions that evening: "test proving on the amd card?" and "can we test proving on +mac?". PC 1 holds an RTX 5090 and an RX 9070 XT (gfx1201, 16 GB) in an eGPU; this Mac is an M5 Max. The prover is +SP1 (`proving/igneum-prove`, `docs/plans/proving-v0.md`, `proving-v1.md`), run on the GPU only through SP1's CUDA +server. Every figure below is measured (with its bench-log entry or job id) or cited (with its file or page); the +rest is labelled approximate. Status words follow `docs/spec/00-overview.md` 0.2. + +**The answer in three lines.** No zkVM proves on an AMD GPU on 5 October 2026: not SP1, not RISC Zero, not Jolt, not +OpenVM, and the ICICLE library underneath them has no AMD backend either. Apple silicon has a shipped Metal prover in +RISC Zero and a Metal backend in ICICLE, but SP1, the prover Igneum runs, is CPU-only on a Mac. So an AMD-only or +Apple-only machine mines and does not prove on its card; it can prove on its CPU, at the times measured in section 2. + +## 1. The backends (read 5 October 2026, 20:30 to 20:50 UTC) + +| Prover | Version read | CPU | NVIDIA (CUDA) | AMD (ROCm or HIP) | Apple (Metal) | Vulkan or WebGPU | Where it says so | +|---|---|---|---|---|---|---|---| +| SP1 (ours) | v6.8.1, 24 Sep 2026 (pinned); `dev` head 318dd530, 28 Sep 2026 | yes; AVX2 and AVX-512 on x86 through Plonky3 | yes: `sp1-gpu-server`, "Compute Capability 8.0 or higher", "24GB or more VRAM", "the CUDA 12 runtime and a compatible NVIDIA driver", Linux x86_64 | **no** | **no** | **no** | docs.succinct.xyz, SP1 docs "Hardware acceleration" page; `crates/sdk/src/lib.rs` (`pub mod cpu`, `mock`, `light`, `#[cfg(feature = "cuda")] pub mod cuda`, `#[cfg(feature = "network")] pub mod network`: no other backend module); `sp1-gpu/README.md` (`CUDA_ARCHS` 89, 90, 100, 120; NTT by NVIDIA cuPQC or sppark); release notes v6.2.3 to v6.8.1 (the only backend line: "add optional cuPQC NTT backend", v6.8.0); a code search of the repository on 5 October: "rocm" 0 files, "metal" 0, "vulkan" 0, "webgpu" 0; `cuobjdump` of sp1-gpu-server 6.8.1: sm_80, 86, 89, 90, 100, 120 and compute_120 PTX, nothing else (`docs/bench-log.md`, "proving v1", 5 October 2026) | +| sppark (SP1's NTT fallback, vendored at `sp1-gpu/crates/sys/sppark`) | `main` README, read 5 October 2026 | | yes: "x86_64 with Nvidia's Volta+ GPU hardware platforms on Linux and Windows" | "A limited support for AMD's RDNA and CDNA GPUs is provided" (upstream README). SP1's tree carries no HIP build: the 0 "rocm" files above, and `sp1-gpu/crates/sys/sppark/util/gpu_t.cuh` is CUDA only | no | no | github.com/supranational/sppark README; the SP1 files named | +| RISC Zero | latest release v3.0.6, 17 Jul 2026 (a v5.0.0-rc.1 of 15 Jan 2026 is also on the releases page) | yes, "nearly any modern CPU (x86 or ARM)" | yes, "RISC Zero targets NVIDIA GPUs using the CUDA framework" | **no** ("rocm", "vulkan": 0 files in the repository) | **yes**: `metal = ["prove"]` in `risc0/zkvm/Cargo.toml`; kernels in `risc0/sys/kernels/zkp/metal/*.metal` (zk, fri, mix, sha); docs: "RISC Zero will use the integrated Metal compute cores" on Apple silicon. The Groth16 wrapper "only works on x86 architecture, and so Apple Silicon is currently unsupported (even via Docker)" | no | dev.risczero.com "Local proving"; `risc0/zkvm/Cargo.toml` features `cuda = [... risc0-zkp/cuda ...]`, `metal = ["prove"]` | +| Jolt (a16z) | v0.3.0-alpha, 1 Oct 2025; "Jolt is in alpha and is not suitable for production use" | yes, "state-of-the-art performance on CPU" | no | **no** | a **draft** PR #1733 (opened 3 Aug 2026, not merged): titled "feat: Metal GPU backend (Apple Silicon)" and marked experimental: 92 Metal kernels, "2.12x speedup at 2^20 scale" on an M4 mini and "3.20x vs same-binary CPU" on an M5 Max, "Apple Silicon + macOS only. No CI coverage" | no | github.com/a16z/jolt README and book (jolt.a16zcrypto.com); PR #1733 | +| OpenVM | v2.0.2, 14 Aug 2026 | yes | yes: `cuda-backend` (v1.4.2 notes), "Improves the Halo2 GPU prover" (v2.0.2) | **no** | **no** | no | github.com/openvm-org/openvm releases | +| ICICLE (Ingonyama; the GPU library behind several provers, not SP1) | v4.0.0, 11 Jul 2025 | yes (MIT) | yes, "CUDA (for NVIDIA GPUs)" | **no** backend listed | yes, "Metal (for Apple Silicon GPUs)"; both under a special licence with a free research licence | Vulkan in the build system (PR #735, merged Jan 2025) and a draft "Vulkan NTT" PR #1019 (Jul 2025, "still wip"); nothing installable | dev.ingonyama.com "Install GPU backend"; the releases page; the README ("backends ... are distributed under a special license") | + +Said plainly: **on 5 October 2026 no zkVM proves on an AMD GPU.** The only AMD code in the whole chain is sppark's +limited HIP path, which SP1 does not build. For Apple silicon the answer is split: RISC Zero ships a Metal prover +and ICICLE a Metal backend; SP1, Jolt and OpenVM do not. Nothing read tonight names an AMD plan with a date. + +## 2. The CPU fallback, measured + +SP1's CPU prover is the path an AMD-only or Apple-only machine has today. Three machines, the same pinned guests +(shard program id `0x2b1a81cb...`, aggregator `0x474678f3...`, pinned 2026-10-05T16:20:38Z), `SP1_PROVER=cpu`, +`--mode shard --shard 0` (execute, core proof, compressed proof, each verified). The RTX 5090 rows are the reference. + +| Fixture (SP1 cycles) | Stage | PC 1 CPU, miner running on both cards (job `cpu-prove-pc1-small2`) | Apple M5 Max CPU (4 October, loaded; bench-log) | RTX 5090 (bench-log) | +|---|---|---|---|---| +| block-56-transfers-3shards shard 0, 200 pgas (315 k) | core | 82.5 s, 7,310,257 B, verify 0.210 s | 83.1 s, 7,310,257 B | not run on the 5090; the nearest rows are block-78 below and an empty live shard: 7.0 to 7.7 s compressed with the miner on the card (5 Oct, `chain-pc2-pv1b`, `pv1c`) | +| | compressed | 199.2 s, 1,272,897 B, verify 0.035 s; 312 s wall for setup 22.8 s, execute 0.14 s, core, compressed | 272.3 s, 1,272,897 B | | +| | peak RSS, CPU | 29.5 GB peak RSS; 978% CPU (9.8 of 16 cores), user 2,516 s, system 537 s | not recorded | | +| block-78-increment, 2 transactions (626 k) | core | 87.0 s, 7,317,857 B, verify 0.209 s | 22.0 s, 7.3 MB (3 October, v0 guest) | 1.4 s (4 October, mining paused) | +| | compressed | 202.3 s, 1,272,897 B, verify 0.034 s; 322 s wall (setup 21.8 s) | 55.7 s, 1.27 MB | 2.7 s | +| | peak RSS, CPU | 30.5 GB peak RSS; 979% CPU, user 2,616 s, system 541 s | not recorded | | +| block-338-shard1, one shard at `S_p` (60.8 M) | core | **not run**, by the PC 1 scheduler's decision at 21:05Z (PC 1's time tonight belongs to the Counter ASIC 2.0 gates; the job `cpu-prove-pc1-sp`, script `tools/amd-prove/pc1-cpu-prove-sp.ps1`, is written and unpublished). Extrapolation, approximate: 60.8 M cycles is about 29 SP1 shards of 2^21 cycles where the small fixtures are one, so the core proof alone is about 29 x 80 s, 40 min, and the compressed recursion over 29 shard proofs adds hours; the floor from the 5090's own ratios (6x on core, 4x on compressed between block-78 and `S_p`) is 9 min core and 13 min compressed. Either way far outside every deadline | not run on the CPU (execute alone 6.9 s) | 8.3 s | +| | compressed | not run (see the core cell) | not run | 10.9 s with the card to itself (4 Oct); 33.0 s with the miner running (5 Oct, `memminer-pc2-pv1`); 7.3 to 7.7 s per EMPTY shard with the miner running (`chain-pc2-pv1c`) | +| | peak RSS, CPU | not run; at least the 30 GB of the small rows | | GPU peak 28,295 MiB alone, 30,039 MiB beside the miner | + +PC 1: Windows 11, WSL2 Ubuntu 24.04 as root, 16 cores and 46,994 MB visible to the VM, the Igneum Miner app 0.3.9 mining on the RTX 5090 (89% mean utilisation through both runs, 59 to 70% minimum: the miner, untouched) and on the RX 9070 XT (not visible to nvidia-smi, mining through the app's OpenCL worker). Job `cpu-prove-pc1-small2`, 20:49:00Z to 20:59:49Z, 649 s wall including a 6-s warm build; the host built without the `cuda` feature from the hosted package `igneum-prove-wsl2-pv1b.zip`, `--mode id` the pinned pair. The first job, `cpu-prove-pc1-small` (20:44 to 20:46Z), built the host cold in 126 s and proved nothing: an apostrophe inside a single-quoted awk program ended the quote, bash refused the whole loop and the job reported exit 0. The class fix: `tools/amd-prove/check-job-bash.sh` runs `bash -n` on the bash body of a PowerShell job before it is published, and the job itself runs `bash -n` inside the distro before the run; both were shown to fire on the bad body and pass the fixed one. Host RAM in the VM: 968 MB used before, 2,351 MB after; the prover's own peak 29.5 to 30.5 GB. + +Mac, fresh run tonight: not taken. The Mac measure lock was held from 20:31Z (a read-width `packbench` under `measure`, three build slots, then a 1,500-s proving-v1 network under `run`) and did not free inside the 10-minute window the coordinator set, so the Apple column is the 4 October rows (M5 Max, 18 cores, 64 GB, load 38 to 47, `nice -n 19`): the same host modes on the same fixture, under heavier load than PC 1 tonight. The Mac's RAM peak was not recorded on 4 October; PC 1's 30 GB says a Mac needs more than 32 GB for the CPU prover, which a 64 GB M5 Max has and a 16 or 24 GB Mac does not. + +The deadlines a CPU proof has to fit (all in the spec and the v1 plan): the exclusive window of an assigned shard is +10 s of DAA time (spec 7.2 item 3; after it anyone may prove and be paid first); the litepaper promises the block's +proof "within about a minute"; the launch target is 20 to 60 s behind the tip; from proving v1 a segment nobody has +proven in `T` = 600 DAA s (10 min) pays nothing (`docs/plans/proving-v1.md`, decisions). So a CPU shard proof is +useful only if it lands inside 10 min and competitive only if it lands inside about a minute. + +Reading. On PC 1 the CPU proof of the smallest shard (315 k cycles) and of the two-transaction block (631 k cycles) cost the same: 82.5 and 87.0 s core, 199.2 and 202.3 s compressed. Doubling the cycles added 4.5 s to the core proof and 3.1 s to the compressed one, so about 280 s of every CPU proof is fixed cost (the recursion that turns the core proof into the 1.27 MB compressed proof the chain carries), and no shard size removes it. Against the deadlines: 282 s a shard (core plus compressed, the client already set up, as the app's loop runs it) is 28x the 10-s assignment window, 4.7x the minute the litepaper promises, and inside the 600-s unproven deadline of v1 with 5 min to spare; but an NVIDIA card proves the same shard in 2.7 to 7.7 s, so a CPU prover only ever wins a shard that no card has taken in 10 minutes. The Mac's 83.1 and 272.3 s of 4 October have the same shape. The RAM peak of 29.5 to 30.5 GB is the second finding: the SP1 CPU prover does not fit a 16 GB machine at all, and WSL2 gives a Windows VM half the host's RAM by default, so the CPU path needs a 64 GB Windows PC or a 32 GB Linux or Mac machine. The `S_p` shard on the CPU can only be slower (the 5090 takes 6x longer at `S_p` than on block-78: 8.3 s against 1.4 s core); it was not run tonight (the scheduler kept PC 1 for the Counter ASIC 2.0 gates) and could not change the conclusion. + +## 3. What this means for each tier (the every-number rule, CLAUDE.md 5 October 2026) + +| Tier | Mines | Proves on the card | The 20% proving-pool share (spec 2.5) | What the software does today | +|---|---|---|---|---| +| AMD-only home miner, one card of 8, 12 or 16 GB (an RX 9070 XT is 16 GB), Windows or Linux | yes (OpenCL worker, `proto-opencl`; PC 1's 9070 XT mines on the devnet) | **no**: no prover exists for the card | **lost**, unless CPU proving at a small shard size becomes a tier (section 4a) | the rig installer: `prover_decision` in `packaging/linux/bin/igneum-rig-lib.sh` (branch `rig-install`) skips every non-NVIDIA card (`[[ "$vendor" == nvidia ]] \|\| continue`) and prints "proving off by default: no NVIDIA card (no CUDA prover for AMD or Intel yet)"; the app: `provedefault.rs` (branch `proving-v1`) considers NVIDIA cards only. Both already right; neither offers the CPU path | +| Apple silicon (M-series, unified memory) | yes: the M5 Max at 26.7 MH/s (bench-log 4 October, "first hourly program swap", Metal `prepare 1` row) | **no** with SP1; RISC Zero and ICICLE have Metal, SP1 does not | **lost** today; a Metal prover behind the swappable interface would restore it (section 4b) | `provedefault.rs`: "proving stays off on Apple silicon: the M5 Max CPU took 41 to 55 s for an empty shard and minutes for a full one; Settings switches it on (CPU, slow)". Right | +| Mixed rig (NVIDIA and AMD cards in one box) | every card | the NVIDIA cards prove for the box; the AMD cards mine | kept, earned by the NVIDIA cards | the rig installer picks the biggest NVIDIA card (`prover_decision`, `PROVER_CARD` overrides), pauses its miner under 20 GB, keeps it mining at 20 GB or more; the AMD cards get a miner unit each. **The prover unit must never select an AMD card**: it does not (the vendor filter above), and that filter is now a stated requirement, not an accident | +| NVIDIA home miner, 8 or 12 GB | yes | no on this SP1 build (13.9 GB floor on an empty shard, `memsweep-pc2-pv1`) | lost unless the shard size moves | unchanged from `proving-v1.md` | +| NVIDIA 16 GB | yes | prove-only, miner paused per shard | kept | unchanged | +| NVIDIA 24 or 32 GB | yes | mines and proves (peak 16.8 GB on empty shards, 30.0 GB on a full prototype shard beside the miner) | kept | unchanged | +| Pool user | through the pool | the pool's own NVIDIA cards prove the shards assigned to the pool's keys (approximate: the pool protocol, spec 09, does not yet say who proves) | by the pool's rules | open, spec 09 | + +## 4. The options + +### 4a. CPU proving at a small shard size, as a tier + +What it is: an AMD-only or Apple machine proves shards cut at a smaller budget than `S_p` on its CPU, through the +same host (`SP1_PROVER=cpu`; the host's `--budget` re-plan from branch `proving-v1`, commit c2544be, cuts a fixture at +any budget). The miner keeps the card; the prover takes the CPU. + +What the numbers say: the fixed cost kills it. 282 s a shard on a 16-core PC and 355 s on the loaded M5 Max, with 30 GB of RAM, at the smallest shard there is; the time sits in the compressed-proof recursion, not in the cycles, so cutting shards smaller does not help, and the launch deadline (20 to 60 s behind the tip) is missed by 5x. It fits only the v1 unproven deadline (600 s), which pays a CPU prover only when no card has proven the shard in 10 minutes: on a chain with one NVIDIA prover that never happens. Recommendation: **no CPU tier**. Settings may still switch the CPU prover on (it does on macOS today), and the Proving tile must then say the proof takes about five minutes and is paid only when no card proves first. + +What it costs the chain: a block cut into more, smaller shards costs more aggregation work (the aggregator guest +verifies one deferred proof per shard; 1.66 M cycles for four shards on the executor, bench-log 4 October; the +chained aggregation is 9.6 to 9.7 s per block on a mining 5090, `chain-pc2-pv1c`) and more records; the assignment +rule (8 assignees, 10 s window, spec 7.2) would need a CPU class with a longer window or the CPU provers only ever +win the open phase. None of that is measured. Status: Designed, nothing implemented. + +### 4b. A second prover backend behind the swappable interface + +The seam exists: `proving/igneum-prove/host/src/proof_system.rs` (`ProofSystem` trait, `Sp1ProofSystem`, +`StubProofSystem`), versioned per the design. The candidates: + +| Target | Most likely backend | What exists | What adopting it costs | +|---|---|---|---| +| Apple silicon | RISC Zero's Metal prover (`metal` feature, shipped) | a shipped feature with kernels in the tree; ICICLE's Metal backend as the other library | a second guest program (the shard statement, `core/` is plain Rust and ports; the precompile patches for keccak and secp256k1 differ), a second pinned program id and verifying key in `elf/manifest.json`, the node's verifier for both proof formats (RISC Zero receipt and SP1 compressed proof) in `--mode verify` and the proof pool, and an aggregation problem: SP1's aggregator folds SP1 proofs by deferred verification; it cannot fold a RISC Zero receipt, so a block with shards from both families needs two aggregations or a wrapper. Approximate: weeks of a person's time, no measurement of a Metal shard time exists; RISC Zero's Groth16 wrapper for light clients does not run on Apple silicon at all | +| AMD | nothing | sppark's limited HIP path (not in SP1's tree); ICICLE's and Jolt's Vulkan and Metal work are not AMD | no backend to adopt. The honest statement is that it lands when a zkVM ships one | + +### 4c. The public line + +The site says today (read 5 October 2026 from `site/litepaper.html`, `site/miner.html`, `site/index.html`): "The same +card proves every block", "The card mines and proves", "Ember finds your GPU, makes a wallet for you and runs the +node, the miner and the prover as one app", "Target: shard size will be set so a 12 GB card proves one shard in about +20 seconds". Every one of those is true of an NVIDIA card with enough memory and false of an AMD or Apple card, and the +litepaper's own rule is "If consumer GPUs cannot prove shards fast enough, Igneum says so and does not launch on promises". + +The recommended line, for the litepaper's proving section, the miner page and the app's Proving tile (copy law): + +> Proving needs an NVIDIA card with 16 GB or more today (20 GB to mine and prove on the same card). AMD and Apple +> cards mine. A prover for them lands when a zkVM ships one. A CPU can prove a small shard in about five minutes with 32 GB of RAM free; the chain pays the first proof, which a card delivers in seconds, so CPU proving is for testing, not income. + +Where the numbers come from: 16 GB and 20 GB are the measured gates of `proving-v1.md` (13.8 GB prover-alone peak, +16.8 GB mine-and-prove peak); "when a zkVM ships one" is section 1. The line changes when the memory sweep moves the +gates or a backend ships; it is reviewed with every prover release. + +## 5. What this analysis does about it (the consequences, before anyone asks) + +| Consequence | Action | Owner | +|---|---|---| +| An AMD-only miner loses the proving share | the CPU tier of 4a is measured here (section 2); whether it becomes a tier is a decision for the project lead on those numbers | this analysis; the project lead | +| The rig's prover unit must select NVIDIA cards only | already true in `prover_decision`; told the rig-installer agent to keep it as a stated rule and to print the CPU-fallback line for AMD-only rigs | rig-installer agent | +| The app's Proving tile on an AMD-only or Apple machine should say why it is off and name the CPU path | the `provedefault.rs` lines already say so for Apple; AMD-only Windows machines get "no NVIDIA card ..." | proving agent (told) | +| The site and litepaper over-promise for AMD and Apple | the line of 4c, to land with the next site pass (copy law; `node site/build.mjs`; link-check) | site-pages owner; not changed here | +| A Metal prover is the only non-NVIDIA path with a shipped backend | 4b names RISC Zero's Metal path and its cost; no work started | proving agent (told) | + +## 6. Commands, jobs and sources + +| What | Where | +|---|---| +| The PC 1 jobs (signed `run` jobs, PowerShell, not elevated, miners untouched, SP1_PROVER=cpu, host built without the `cuda` feature) | `tools/amd-prove/pc1-cpu-prove.ps1` (small fixtures), `pc1-cpu-prove-sp.ps1` (the `S_p` shard); published as `cpu-prove-pc1-small` (built, proved nothing: the quote bug) and `cpu-prove-pc1-small2` (the numbers) by `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --shell powershell --timeout-minutes 60`; `cpu-prove-pc1-sp` written, not published; `check-job-bash.sh` gates the bash body of every job script here; the package `igneum-prove-wsl2-pv1b.zip` (sha256 df50dee5...), the URL and hash filled at publish time, never committed | +| The Mac run | `tools/lock/with-lock.sh measure /usr/bin/time -l igneum-prove-host block-56-transfers-3shards.json --mode shard --shard 0` with the `proving-v1` worktree's host (pinned ids checked with `--mode id`) | +| Results | `node tools/jobs.mjs cpu-prove-pc1-small2 --all`, `docs/bench-log.md` entry "5 October 2026, the CPU prover on PC 1 and the backend survey" | +| Pages read | SP1: docs.succinct.xyz hardware-acceleration page, github.com/succinctlabs/sp1 (releases, `crates/sdk/src/lib.rs`, `sp1-gpu/README.md`, code search); RISC Zero: dev.risczero.com local-proving, `risc0/zkvm/Cargo.toml`; Jolt: README, book, PR #1733; OpenVM: releases; ICICLE: install_gpu_backend page, releases, README, PRs #735 and #1019; sppark README | diff --git a/docs/bench-log.md b/docs/bench-log.md index 9c0171fc2..a4e913320 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1524,3 +1524,19 @@ What is measured: one BLS12-381 aggregate signature over 16 summed G1 keys plus | on, split 90 s | v3 | 0 / 2 | none / 3 | 278 / 265 | apart | none | 3 on n0 | 2 (n0 reconnected 6 s after the heal, A's chain at about 58 DAA, inside the table) | Reading (the NEW finding, ledger C4). With the module off GHOSTDAG alone converges on the heavier chain and the losing side's records re-determine (F24 works when the chain moves). With the module on the overlay holds during the split (A, with 30% of the frozen table, locks nothing; B locks 7 and 8) and then fails at the heal in the shipped node: B's certificates for blocks off n0's chain are "kept pending until the chain decides (no lock at this index)", n0's chain never decides because GHOSTDAG keeps its heavier tip and nothing turns the certificate into a fork-choice constraint, and once n0's last lock (index 7, DAA 209) is one window old (DAA 329) the frozen table stops applying on A's chain ("no frozen table (no lock on this chain inside the window)"), A's two keys are 100% of A's own window (B's post-cut blocks are red there) and n0 locks 10, 11, 12 alone; B's certificates for 10 and 11 then log CONFLICTING on n0 (n0 log, 17:27:04 to 17:29:54 BST). A finality fork from a 96-s honest partition, no attacker, table intact at the heal; the 150-s run and the v2 control end the same way. The spec's fork choice ("GHOSTDAG among tips through all certified checkpoints", 3.5) is therefore implemented only for certificates over blocks already on the node's chain. Fix named in the ledger entry: verify an off-chain certificate against the table at its own block and let it constrain fork choice (a certificate-driven reorg), then re-determine. Raw: `scratchpad fud-a/c4-results-*.md`, node logs `c4-on90-tmp/`, `c4-v2-control-tmp/`. + +## 5 October 2026 (night), the SP1 CPU prover on PC 1 beside the miners, and the backend survey: no zkVM proves on AMD (amd-prove agent) + +the project lead, 22:50 BST: "test proving on the amd card?" and "can we test proving on mac?". The analysis with the backend table and the tier consequences: `docs/analysis/amd-proving.md`. The survey (SP1 v6.8.1 and `dev` 318dd530 of 28 Sep 2026, RISC Zero, Jolt, OpenVM, ICICLE, sppark; every claim cites a file or page there): on 5 October 2026 no zkVM proves on an AMD GPU; Apple silicon has RISC Zero's shipped Metal prover and ICICLE's Metal backend; SP1, the prover here, is CPU-only off NVIDIA. + +Machine: PC 1 (machine ae432dc7), Windows 11, WSL2 Ubuntu 24.04 as root, 16 cores and 46,994 MB visible to the VM, the Igneum Miner app 0.3.9 mining on the RTX 5090 and the RX 9070 XT throughout (the 5090 at 89% mean utilisation, 59 to 70% minimum, from a 1-s `nvidia-smi` sampler under the run: the job never touched a card). Signed `run` job `cpu-prove-pc1-small2` (`tools/amd-prove/pc1-cpu-prove.ps1`), 20:49:00Z to 20:59:49Z: the hosted package `igneum-prove-wsl2-pv1b.zip` (sha256 df50dee5...) built WITHOUT the `cuda` feature (6 s warm; the first job `cpu-prove-pc1-small` built it cold in 126 s), `--mode id` the pinned pair (shard `0x2b1a81cb...`, aggregator `0x474678f3...`, pinned 2026-10-05T16:20:38Z), `SP1_PROVER=cpu`, `--mode shard --shard 0` under `/usr/bin/time -v`. Log: `node tools/jobs.mjs cpu-prove-pc1-small2 --all`. + +| Fixture | SP1 cycles | Setup s | Core prove s (bytes, verify s) | Compressed prove s (bytes, verify s) | Wall s | Peak RSS | CPU | +|---|---|---|---|---|---|---|---| +| block-56-transfers-3shards shard 0 (200 pgas, one transfer) | 315,479 | 22.75 (client 19.46, shard keys 1.85, aggregator keys 1.44) | 82.5 (7,310,257, 0.210) VERIFIED | 199.2 (1,272,897, 0.035) VERIFIED | 312.1 | 29,503,652 kB (29.5 GB) | 978% (9.8 of 16 cores), user 2,516 s, system 537 s, load max 11.3 | +| block-78-increment (2 transactions, 1 executed 1 skipped) | 631,127 | 21.75 | 87.0 (7,317,857, 0.209) VERIFIED | 202.3 (1,272,897, 0.034) VERIFIED | 322.3 | 30,517,916 kB (30.5 GB) | 979%, user 2,616 s, system 541 s, load max 13.1 | +| block-338-shard1 (one shard at `S_p`, 60.8 M cycles) | | not run: the PC 1 scheduler kept the machine for the Counter ASIC 2.0 gates (21:05Z). Approximate extrapolation: about 29 SP1 shards of 2^21 cycles at about 80 s each, 40 min of core proof, then hours of recursion; floor from the 5090's ratios (6x core, 4x compressed, block-78 to `S_p`): 9 min core, 13 min compressed | | | | | | + +For comparison (this log): the Apple M5 Max CPU on 4 October, loaded, block-56 shard 0: core 83.1 s, compressed 272.3 s; on 3 October the v0 guest on block-78: core 22.0 s, compressed 55.7 s. The RTX 5090: block-78 core 1.4 s, compressed 2.7 s (4 October, mining paused); a full shard at `S_p` compressed 10.9 s alone and 33.0 s beside the miner; an empty live shard 7.0 to 7.7 s beside the miner (5 October). No fresh Mac run tonight: the measure lock was held from 20:31Z (a read-width `packbench`, three builds, a 1,500-s proving-v1 network under `run`) and did not free inside the 10-minute window set for it. + +Reading, and the consequences (CLAUDE.md, every number). Doubling the cycles added 4.5 s to the core proof and 3.1 s to the compressed proof: about 280 s of a CPU proof is fixed cost in the compressed-proof recursion, so no shard size brings a CPU proof under the launch deadline (20 to 60 s behind the tip) or near the 10-s assignment window; it fits only the v1 unproven deadline (600 s), which pays a CPU prover only when no card has proven the shard in 10 minutes. The 29.5 to 30.5 GB peak RSS means the CPU prover needs 32 GB free: a 64 GB Windows PC (WSL2 takes half the host's RAM by default), a 32 GB Linux machine, a 64 GB Mac; a 16 GB machine cannot run it at all. Per tier: an AMD-only home miner (8, 12 or 16 GB, Windows or Linux) mines and does not prove, and loses the 20% proving-pool share; Apple silicon the same (the M5 Max mines at 26.7 MH/s, this log, 4 October); a mixed rig proves on its NVIDIA cards and the rig installer's `prover_decision` already skips every non-NVIDIA card (`packaging/linux/bin/igneum-rig-lib.sh`, branch `rig-install`), now a stated requirement; the app's `provedefault.rs` already keeps proving off on Apple silicon and off without an NVIDIA card. Decision asked of nobody: no CPU tier (the analysis, section 4a); the public line for the site, litepaper and Proving tile is in section 4c ("Proving needs an NVIDIA card with 16 GB or more today ... AMD and Apple cards mine. A prover for them lands when a zkVM ships one"). The first job proved nothing because an apostrophe inside a single-quoted awk program ended the quote and bash refused the loop while the job reported exit 0; the class fix is `tools/amd-prove/check-job-bash.sh` (`bash -n` on the embedded bash body before publishing) and the same `bash -n` inside the job before the run, both shown to refuse the bad body and pass the fixed one. diff --git a/tools/amd-prove/check-job-bash.sh b/tools/amd-prove/check-job-bash.sh new file mode 100755 index 000000000..198365ffd --- /dev/null +++ b/tools/amd-prove/check-job-bash.sh @@ -0,0 +1,26 @@ +#!/usr/bin/env bash +# Checks the bash body embedded in a PowerShell job script (the here-string assigned to $bash) with `bash -n` +# before it is published. The class of 5 October 2026 (job cpu-prove-pc1-small): an apostrophe inside a single-quoted +# awk program ended the quote, bash refused the whole for-loop, and the job reported exit 0 with no measurement. +# Usage: tools/amd-prove/check-job-bash.sh ... (exit 1 when a body does not parse) +set -uo pipefail +rc=0 +tmp="$(mktemp)" +for f in "$@"; do + python3 - "$f" > "$tmp" <<'PY' +import re,sys +t=open(sys.argv[1],encoding='utf-8').read() +m=re.search(r'^\$bash = @"\n(.*?)\n"@', t, re.S|re.M) +if not m: print("NO-HERE-STRING"); sys.exit(0) +b=m.group(1) +bt=chr(96) +# PowerShell expands $NAME and $(...) in a double-quoted here-string; a backtick-dollar is a literal dollar. +for v in ('pkgW','jobW','FIXTURES'): b=b.replace('$'+v, 'PLACEHOLDER_'+v) +b=b.replace(bt+'$','$').replace(bt+bt,bt) +print(b) +PY + if grep -q '^NO-HERE-STRING$' "$tmp"; then echo "$f: no bash here-string, skipped"; continue; fi + if bash -n "$tmp" 2>"$tmp.err"; then echo "$f: bash body parses"; else echo "$f: BASH BODY DOES NOT PARSE"; cat "$tmp.err"; rc=1; fi +done +rm -f "$tmp" "$tmp.err" +exit $rc diff --git a/tools/amd-prove/pc1-cpu-prove-sp.ps1 b/tools/amd-prove/pc1-cpu-prove-sp.ps1 new file mode 100644 index 000000000..e725ff808 --- /dev/null +++ b/tools/amd-prove/pc1-cpu-prove-sp.ps1 @@ -0,0 +1,81 @@ +# AMD proving question (5 October 2026, the project lead: "test proving on the amd card?"): the SP1 CPU prover on PC 1 (machine +# ae432dc7, RTX 5090 + RX 9070 XT on the eGPU), as a signed `run` job (shell powershell, not elevated, the miners keep +# mining on BOTH cards; the job never touches a card: SP1_PROVER=cpu, the host built WITHOUT the cuda feature). +# No zkVM proves on AMD today (docs/analysis/amd-proving.md), so the CPU path is the only prover an AMD-only machine has. +# 1. downloads the hosted prover package (igneum-prove-wsl2-pv1b.zip, sha256 checked) into the job folder +# 2. inside WSL2 (root, Ubuntu-24.04): package -> ~/igneum-prove-cpu, every file re-stamped, the host built +# with `cargo build --release -p igneum-prove-host` (no cuda feature; the guests are pinned, no Succinct toolchain) +# 3. for each fixture in $FIXTURES: `--mode shard --shard 0` (execute, core, compressed, each verified) under +# /usr/bin/time -v for the wall time, the peak resident set and the CPU percentage; a 1-s sampler of +# /proc/loadavg and nvidia-smi utilisation underneath (the 5090's utilisation stays at the miner's level: the +# job did not touch it; the 9070 XT is not visible to nvidia-smi) +# Every number is a RESULT line. The results JSON per fixture is printed at the end (RESULTS-JSON ... END). +$ErrorActionPreference = 'Continue' +$FIXTURES = 'block-338-shard1' +$ZIP_URL = '__ZIP_URL__' # filled at publish time from the downloads folder (the dl token never enters the repository) +$ZIP_SHA = '__ZIP_SHA__' +function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') } +$job = $env:IGNEUM_JOB_DIR +if (-not $job) { $job = Join-Path $env:TEMP 'igneum-cpu-prove' } +New-Item -ItemType Directory -Force -Path $job | Out-Null +"RESULT start $(Stamp) machine=$env:IGNEUM_MACHINE_ID fixtures=$FIXTURES" +# 1. the package +$zip = Join-Path $job 'igneum-prove-wsl2-pv1b.zip' +& curl.exe -s -S -L -m 600 -o $zip $ZIP_URL 2>&1 | ForEach-Object { "curl: $_" } +if (-not (Test-Path $zip)) { "RESULT package FAILED: no download"; exit 1 } +$sha = (Get-FileHash -Algorithm SHA256 $zip).Hash.ToLower() +if ($sha -ne $ZIP_SHA) { "RESULT package FAILED: sha256 $sha is not $ZIP_SHA"; exit 1 } +$pkg = Join-Path $job 'pkg' +if (Test-Path $pkg) { Remove-Item -Recurse -Force $pkg } +Expand-Archive -Path $zip -DestinationPath $pkg -Force +$pkgRoot = Get-ChildItem -Path $pkg -Directory | Select-Object -First 1 +"RESULT package $(Stamp) $((Get-Item $zip).Length) bytes, sha256 ok, extracted to $($pkgRoot.FullName)" +function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } } +$pkgW = WslPath $pkgRoot.FullName; $jobW = WslPath $job +$bash = @" +set -uo pipefail +export PATH="`$HOME/.cargo/bin:`$PATH"; [ -f "`$HOME/.cargo/env" ] && . "`$HOME/.cargo/env" +stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; } +PKG='$pkgW'; JOB='$jobW'; DEST="`$HOME/igneum-prove-cpu" +echo "RESULT wsl `$(stamp) host `$(hostname) user `$(id -un) cores `$(nproc) ram_total_mb `$(free -m | awk '/Mem:/ {print `$2}') ram_used_mb `$(free -m | awk '/Mem:/ {print `$3}') loadavg `$(cut -d' ' -f1-3 /proc/loadavg) cargo `$(command -v cargo || echo MISSING)" +nvidia-smi --query-gpu=name,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null | sed 's/^/RESULT gpu-before /' || echo "RESULT gpu-before nvidia-smi not available" +# tools the build and the measurement need (the node build on this distro had the compilers; protoc and GNU time may be missing) +command -v protoc >/dev/null && command -v /usr/bin/time >/dev/null || { apt-get update -qq >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -qq protobuf-compiler time pkg-config libssl-dev >/dev/null 2>&1 || echo "RESULT apt FAILED (continuing)"; } +echo "RESULT tools protoc `$(command -v protoc || echo MISSING) time `$(command -v /usr/bin/time || echo MISSING)" +mkdir -p "`$DEST" +rsync -a --delete --exclude target "`$PKG/package/" "`$DEST/" 2>/dev/null || { rm -rf "`$DEST"; mkdir -p "`$DEST"; cp -r "`$PKG/package/." "`$DEST/"; } +# re-stamp every copied file: cargo rebuilds by mtime and this side keeps its target dir (the stale-build class, 4 and 5 October 2026) +find "`$DEST" -name target -prune -o -type f -exec touch {} + 2>/dev/null +grep -o '"program_id": "0x[0-9a-f]*"' "`$DEST/proving/igneum-prove/elf/manifest.json" | sed 's/^/RESULT manifest /' +cd "`$DEST/proving/igneum-prove" +echo "RESULT build start `$(stamp) (no cuda feature; cold target dir unless this job ran before)" +t0=`$(date +%s) +if ! cargo build --release -p igneum-prove-host 2>&1 | tail -3; then echo "RESULT build FAILED"; exit 1; fi +echo "RESULT build `$(stamp) exit 0 in `$(( `$(date +%s) - t0 )) s" +H="`$DEST/proving/igneum-prove/target/release/igneum-prove-host" +`$H --mode id | sed 's/^/RESULT cpu-host /' +mkdir -p "`$JOB/results" +for FX in $FIXTURES; do + F="`$DEST/proving/fixtures/`$FX.json" + [ -f "`$F" ] || { echo "RESULT `$FX FAILED: no fixture `$F"; continue; } + `$H "`$F" --mode native 2>&1 | grep -E "^RESULT (native|plan)" | sed "s/^/`$FX /" + # the samplers: load average and the 5090's utilisation, one line a second + ( while true; do echo "`$(date -u +%H:%M:%S) `$(cut -d' ' -f1 /proc/loadavg) `$(nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader,nounits 2>/dev/null | head -1 | tr -d ' ')"; sleep 1; done ) > "`$JOB/results/`$FX-sampler.txt" 2>/dev/null & + S=`$! + echo "RESULT `$FX prove start `$(stamp) SP1_PROVER=cpu mode shard --shard 0" + SP1_PROVER=cpu RUST_LOG=off /usr/bin/time -v -o "`$JOB/results/`$FX-time.txt" `$H "`$F" --mode shard --shard 0 --out "`$JOB/results/`$FX-cpu.json" 2>&1 | grep -E "^(RESULT|STAGE|igneum-prove-host sources)" | sed "s/^/`$FX: /" + echo "RESULT `$FX prove end `$(stamp) exit `${PIPESTATUS[0]}" + kill `$S 2>/dev/null; sleep 1 + awk -v fx="`$FX" '/Elapsed .wall clock./ {w=`$NF} /Maximum resident set size/ {r=`$NF} /Percent of CPU this job got/ {c=`$NF} /User time/ {u=`$NF} /System time/ {s=`$NF} END {print "RESULT " fx " time wall=" w " max_rss_kb=" r " cpu_percent=" c " user_s=" u " sys_s=" s}' "`$JOB/results/`$FX-time.txt" + awk -v fx="`$FX" -F'[ ,]' 'NF>=2 { n++; if (`$2+0 > lmax) lmax=`$2+0; if (`$3 != "") { gs += `$3+0; gn++; if (`$3+0 < gmin || gn==1) gmin=`$3+0 } } END { printf "RESULT %s sampler samples=%d loadavg_max=%.1f gpu0_util_mean=%.0f gpu0_util_min=%d (the level of the miner: the job did not use the card)\n", fx, n, lmax, (gn? gs/gn : -1), gmin }' "`$JOB/results/`$FX-sampler.txt" +done +free -m | awk '/Mem:/ {print "RESULT wsl_ram_after total_mb=" `$2 " used_mb=" `$3}' +nvidia-smi --query-gpu=name,utilization.gpu --format=csv,noheader 2>/dev/null | sed 's/^/RESULT gpu-after /' +for FX in $FIXTURES; do [ -f "`$JOB/results/`$FX-cpu.json" ] && { echo "RESULTS-JSON `$FX"; cat "`$JOB/results/`$FX-cpu.json"; echo; echo "END"; }; done +echo "RESULT done `$(stamp)" +"@ +$bashFile = Join-Path $job 'cpu-prove.sh' +[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false)) +$chk = (& wsl.exe -d Ubuntu-24.04 -u root -- bash -n (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", "") }); if ($LASTEXITCODE -ne 0) { "RESULT syntax FAILED: $chk"; exit 1 } else { "RESULT syntax ok (bash -n)" } +& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') } +"RESULT end $(Stamp)" diff --git a/tools/amd-prove/pc1-cpu-prove.ps1 b/tools/amd-prove/pc1-cpu-prove.ps1 new file mode 100644 index 000000000..52e8e26b4 --- /dev/null +++ b/tools/amd-prove/pc1-cpu-prove.ps1 @@ -0,0 +1,81 @@ +# AMD proving question (5 October 2026, the project lead: "test proving on the amd card?"): the SP1 CPU prover on PC 1 (machine +# ae432dc7, RTX 5090 + RX 9070 XT on the eGPU), as a signed `run` job (shell powershell, not elevated, the miners keep +# mining on BOTH cards; the job never touches a card: SP1_PROVER=cpu, the host built WITHOUT the cuda feature). +# No zkVM proves on AMD today (docs/analysis/amd-proving.md), so the CPU path is the only prover an AMD-only machine has. +# 1. downloads the hosted prover package (igneum-prove-wsl2-pv1b.zip, sha256 checked) into the job folder +# 2. inside WSL2 (root, Ubuntu-24.04): package -> ~/igneum-prove-cpu, every file re-stamped, the host built +# with `cargo build --release -p igneum-prove-host` (no cuda feature; the guests are pinned, no Succinct toolchain) +# 3. for each fixture in $FIXTURES: `--mode shard --shard 0` (execute, core, compressed, each verified) under +# /usr/bin/time -v for the wall time, the peak resident set and the CPU percentage; a 1-s sampler of +# /proc/loadavg and nvidia-smi utilisation underneath (the 5090's utilisation stays at the miner's level: the +# job did not touch it; the 9070 XT is not visible to nvidia-smi) +# Every number is a RESULT line. The results JSON per fixture is printed at the end (RESULTS-JSON ... END). +$ErrorActionPreference = 'Continue' +$FIXTURES = 'block-56-transfers-3shards block-78-increment' +$ZIP_URL = '__ZIP_URL__' # filled at publish time from the downloads folder (the dl token never enters the repository) +$ZIP_SHA = '__ZIP_SHA__' +function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') } +$job = $env:IGNEUM_JOB_DIR +if (-not $job) { $job = Join-Path $env:TEMP 'igneum-cpu-prove' } +New-Item -ItemType Directory -Force -Path $job | Out-Null +"RESULT start $(Stamp) machine=$env:IGNEUM_MACHINE_ID fixtures=$FIXTURES" +# 1. the package +$zip = Join-Path $job 'igneum-prove-wsl2-pv1b.zip' +& curl.exe -s -S -L -m 600 -o $zip $ZIP_URL 2>&1 | ForEach-Object { "curl: $_" } +if (-not (Test-Path $zip)) { "RESULT package FAILED: no download"; exit 1 } +$sha = (Get-FileHash -Algorithm SHA256 $zip).Hash.ToLower() +if ($sha -ne $ZIP_SHA) { "RESULT package FAILED: sha256 $sha is not $ZIP_SHA"; exit 1 } +$pkg = Join-Path $job 'pkg' +if (Test-Path $pkg) { Remove-Item -Recurse -Force $pkg } +Expand-Archive -Path $zip -DestinationPath $pkg -Force +$pkgRoot = Get-ChildItem -Path $pkg -Directory | Select-Object -First 1 +"RESULT package $(Stamp) $((Get-Item $zip).Length) bytes, sha256 ok, extracted to $($pkgRoot.FullName)" +function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } } +$pkgW = WslPath $pkgRoot.FullName; $jobW = WslPath $job +$bash = @" +set -uo pipefail +export PATH="`$HOME/.cargo/bin:`$PATH"; [ -f "`$HOME/.cargo/env" ] && . "`$HOME/.cargo/env" +stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; } +PKG='$pkgW'; JOB='$jobW'; DEST="`$HOME/igneum-prove-cpu" +echo "RESULT wsl `$(stamp) host `$(hostname) user `$(id -un) cores `$(nproc) ram_total_mb `$(free -m | awk '/Mem:/ {print `$2}') ram_used_mb `$(free -m | awk '/Mem:/ {print `$3}') loadavg `$(cut -d' ' -f1-3 /proc/loadavg) cargo `$(command -v cargo || echo MISSING)" +nvidia-smi --query-gpu=name,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null | sed 's/^/RESULT gpu-before /' || echo "RESULT gpu-before nvidia-smi not available" +# tools the build and the measurement need (the node build on this distro had the compilers; protoc and GNU time may be missing) +command -v protoc >/dev/null && command -v /usr/bin/time >/dev/null || { apt-get update -qq >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -qq protobuf-compiler time pkg-config libssl-dev >/dev/null 2>&1 || echo "RESULT apt FAILED (continuing)"; } +echo "RESULT tools protoc `$(command -v protoc || echo MISSING) time `$(command -v /usr/bin/time || echo MISSING)" +mkdir -p "`$DEST" +rsync -a --delete --exclude target "`$PKG/package/" "`$DEST/" 2>/dev/null || { rm -rf "`$DEST"; mkdir -p "`$DEST"; cp -r "`$PKG/package/." "`$DEST/"; } +# re-stamp every copied file: cargo rebuilds by mtime and this side keeps its target dir (the stale-build class, 4 and 5 October 2026) +find "`$DEST" -name target -prune -o -type f -exec touch {} + 2>/dev/null +grep -o '"program_id": "0x[0-9a-f]*"' "`$DEST/proving/igneum-prove/elf/manifest.json" | sed 's/^/RESULT manifest /' +cd "`$DEST/proving/igneum-prove" +echo "RESULT build start `$(stamp) (no cuda feature; cold target dir unless this job ran before)" +t0=`$(date +%s) +if ! cargo build --release -p igneum-prove-host 2>&1 | tail -3; then echo "RESULT build FAILED"; exit 1; fi +echo "RESULT build `$(stamp) exit 0 in `$(( `$(date +%s) - t0 )) s" +H="`$DEST/proving/igneum-prove/target/release/igneum-prove-host" +`$H --mode id | sed 's/^/RESULT cpu-host /' +mkdir -p "`$JOB/results" +for FX in $FIXTURES; do + F="`$DEST/proving/fixtures/`$FX.json" + [ -f "`$F" ] || { echo "RESULT `$FX FAILED: no fixture `$F"; continue; } + `$H "`$F" --mode native 2>&1 | grep -E "^RESULT (native|plan)" | sed "s/^/`$FX /" + # the samplers: load average and the 5090's utilisation, one line a second + ( while true; do echo "`$(date -u +%H:%M:%S) `$(cut -d' ' -f1 /proc/loadavg) `$(nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader,nounits 2>/dev/null | head -1 | tr -d ' ')"; sleep 1; done ) > "`$JOB/results/`$FX-sampler.txt" 2>/dev/null & + S=`$! + echo "RESULT `$FX prove start `$(stamp) SP1_PROVER=cpu mode shard --shard 0" + SP1_PROVER=cpu RUST_LOG=off /usr/bin/time -v -o "`$JOB/results/`$FX-time.txt" `$H "`$F" --mode shard --shard 0 --out "`$JOB/results/`$FX-cpu.json" 2>&1 | grep -E "^(RESULT|STAGE|igneum-prove-host sources)" | sed "s/^/`$FX: /" + echo "RESULT `$FX prove end `$(stamp) exit `${PIPESTATUS[0]}" + kill `$S 2>/dev/null; sleep 1 + awk -v fx="`$FX" '/Elapsed .wall clock./ {w=`$NF} /Maximum resident set size/ {r=`$NF} /Percent of CPU this job got/ {c=`$NF} /User time/ {u=`$NF} /System time/ {s=`$NF} END {print "RESULT " fx " time wall=" w " max_rss_kb=" r " cpu_percent=" c " user_s=" u " sys_s=" s}' "`$JOB/results/`$FX-time.txt" + awk -v fx="`$FX" -F'[ ,]' 'NF>=2 { n++; if (`$2+0 > lmax) lmax=`$2+0; if (`$3 != "") { gs += `$3+0; gn++; if (`$3+0 < gmin || gn==1) gmin=`$3+0 } } END { printf "RESULT %s sampler samples=%d loadavg_max=%.1f gpu0_util_mean=%.0f gpu0_util_min=%d (the level of the miner: the job did not use the card)\n", fx, n, lmax, (gn? gs/gn : -1), gmin }' "`$JOB/results/`$FX-sampler.txt" +done +free -m | awk '/Mem:/ {print "RESULT wsl_ram_after total_mb=" `$2 " used_mb=" `$3}' +nvidia-smi --query-gpu=name,utilization.gpu --format=csv,noheader 2>/dev/null | sed 's/^/RESULT gpu-after /' +for FX in $FIXTURES; do [ -f "`$JOB/results/`$FX-cpu.json" ] && { echo "RESULTS-JSON `$FX"; cat "`$JOB/results/`$FX-cpu.json"; echo; echo "END"; }; done +echo "RESULT done `$(stamp)" +"@ +$bashFile = Join-Path $job 'cpu-prove.sh' +[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false)) +$chk = (& wsl.exe -d Ubuntu-24.04 -u root -- bash -n (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", "") }); if ($LASTEXITCODE -ne 0) { "RESULT syntax FAILED: $chk"; exit 1 } else { "RESULT syntax ok (bash -n)" } +& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') } +"RESULT end $(Stamp)" From 3c129fb7a9bb24b09278a8dd07dd96991d028e60 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:04:11 +0000 Subject: [PATCH 056/131] Site, litepaper, miner page: the proving line (NVIDIA 24 GB from the fee switch, 32 GB today; AMD and Apple cards mine) replaces "the card mines and proves" --- site/index.html | 4 ++-- site/litepaper.html | 3 ++- site/miner.html | 10 +++++----- 3 files changed, 9 insertions(+), 8 deletions(-) diff --git a/site/index.html b/site/index.html index f067ce671..4f11aa90f 100644 --- a/site/index.html +++ b/site/index.html @@ -340,7 +340,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var(
GPUs are back · for good

Mined by GPUs.
Proven by fire.

-

A chain built so a chip gains too little to take your place. A mining program that rewrites itself every hour, so a chip built for one hour is useless the next. The same card proves every block and gets paid for it. No premine, no stake, no foundation, no merge to proof of stake.

+

A chain built so a chip gains too little to take your place. A mining program that rewrites itself every hour, so a chip built for one hour is useless the next. An NVIDIA card with 24 GB or more proves every block and gets paid for it; AMD and Apple cards mine. No premine, no stake, no foundation, no merge to proof of stake.

See the miner Read the litepaper @@ -481,7 +481,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var(
-

One click. The card mines and proves.

+

One click. The card mines; an NVIDIA card with 24 GB proves too.

Igneum Ember finds your GPU, makes a wallet for you and runs the node, the miner and the prover as one app.

diff --git a/site/litepaper.html b/site/litepaper.html index ce7b0620b..37ed6bb39 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -437,7 +437,7 @@ body.all .pager{display:none} Light verification256 MB cache on a CPU, milliseconds256 MB cache on a CPU, one warp under 10 ms, the gate. Measured 0.41 to 0.58 ms on one Apple M5 Max core; a 2019-class core not yet Changes over timeNone. A fixed design, unchanged for seven yearsA new program every hour, its memory pattern with it; era draws and reserved families on a schedule fixed at genesis. Nobody touches it Seed grindingNot applicable, the program comes from the hash inputClosed by a verifiable delay between seed and program - Useful workNone. Hashing onlyThe same card proves every block and sells proofs to other chains + Useful workNone. Hashing onlyNVIDIA cards with 24 GB or more prove every block and sell proofs to other chains; AMD and Apple cards mine, and a prover for them lands when a zkVM ships one Track recordNo chip publicly shipped in seven years, approximateZero years. Every number above is measured and logged with the commands that produced it. The specification, reference hash, test vectors and simulators are public now (github.com/igneum-network/spec). The node, the miner and the wallet are in a private repository until the public testnet
@@ -449,6 +449,7 @@ body.all .pager{display:none}

Every Igneum block is proven with a zero-knowledge proof, and the miners produce it. Proving is a useful GPU workload that is cheaply verifiable by construction. A proof is right or it is not, and a phone can check it in milliseconds.

How a block gets proven

Blocks carry transactions only and make no claim about state. Every node executes the ordered transactions natively at once, so users see their transaction land in about a second. The execution is then split into shards of a fixed proving cost. Shards are assigned by lot to eight provers for ten seconds, then open to anyone; there is no bond. Provers run them on consumer cards, and the shard proofs are folded by recursive aggregation into one proof for the block. That proof lands on-chain within about a minute at launch. Because the proof computes the state from the ordered sequence, no node accepts a block with a wrong state root. Full nodes also execute every block natively and reject a proof record whose result differs from their own execution, so a forged proof is a light-client problem and never a chain split. Implemented: the native-execution check on every carried proof record, proving v0 on the devnet (specification section 7). The emergency path for a soundness bug in the proof system is a human one: a new proof-system version is written by people and activates only on miner signalling. Invalid transactions are skipped by rule, the way Kaspa skips conflicting spends.

+

Proving needs an NVIDIA card with 24 GB or more (32 GB until the fee switch of 6 October 2026; from it a 24 GB card mines and proves on the same card: 22.2 GB peak measured with the miner on, 5 October 2026). AMD and Apple cards mine. A prover for them lands when a zkVM ships one.

The proving budget

Gas prices execution. Proving cost is a different number, so Igneum meters it separately: every transaction pays in both dimensions, and each block has a proving-cost budget set in consensus from measured prover throughput. A transaction that is cheap to run and expensive to prove pays for what it costs the provers. Target: shard size will be set so a 12 GB card proves one shard in about 20 seconds. That number is the phase 2 gate on the roadmap and is not measured yet on the card the gate names. The first proofs exist: on 4 October 2026 an RTX 5090 proved a small two-transaction block in 1.4 seconds (2.7 seconds compressed), verified in 0.22 and 0.038 seconds, and a laptop CPU proved a three-shard block end to end in 19 minutes. Later that day the same card proved a full shard at the provisional size, 6.75 million prover gas, which executed in 60.8 million cycles: core proof 8.3 seconds, compressed proof 10.9 seconds, verified in 0.040 seconds; a four-shard block took 44.5 seconds of GPU stages end to end. Since 5 October 2026 shards are assigned and proven on the live devnet. The gate asks for a mid-range card, and an RTX 5090 is not one, so the gate stands open. Once the gate is measured, the budget rises by schedule as hardware improves. The proof system is hash-based, which is what runs on consumer cards, and sits behind a versioned interface, so Igneum can adopt a better proof system when one exists by a miner-signalled release, and runs for ever on the current one if none is adopted.

Proving for everyone else

diff --git a/site/miner.html b/site/miner.html index 20687283a..f170bf4db 100644 --- a/site/miner.html +++ b/site/miner.html @@ -4,13 +4,13 @@ The Igneum miner app - + - + @@ -18,7 +18,7 @@ - + @@ -249,7 +249,7 @@ pre b{color:var(--molten);font-weight:500}
Igneum Ember · one click
-

Install. Start. The card mines and proves.

+

Install. Start. The card mines. An NVIDIA card with 24 GB proves too.

Ember finds your GPU, makes a wallet for you and runs the node, the miner and the prover as one app. Every hour it compiles the next program while the current one mines, so the card never stops.

Downloads: public testnet @@ -305,7 +305,7 @@ pre b{color:var(--molten);font-weight:500}
01 · One click

Install, press Start

Five steps from download to the first block on a friend's laptop: drag to Applications, open, Get started, Make me an address, Start mining.

-
The card mines and proves

One button starts the node, waits for sync, starts one miner per card and the prover. Quit stops them in order.

+
The card mines; an NVIDIA card proves

One button starts the node, waits for sync, starts one miner per card and, on an NVIDIA card with 24 GB or more, the prover. AMD and Apple cards mine; a prover for them lands when a zkVM ships one. Quit stops them in order.

Finds your GPU by itself

NVIDIA through the driver, AMD and Intel through OpenCL, Apple silicon through Metal. Real card names, one worker per card.

A wallet made for you

A payout key made on your machine and locked to your user, with a backup step before mining starts. Or paste an address you already hold.

Earnings in IGN

Every block this machine finds pays one address in IGN. The dashboard counts them per run, per hour and for life. The wallet shows the balance.

From 006b2fe5850cbeba52e3539676a2a5d3abe03745 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:04:35 +0000 Subject: [PATCH 057/131] Counter ASIC 2.0: proving v1 handoff recorded in the rollout plan; status 21:04 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index e8ad72d27..826559f0e 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -99,7 +99,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's ## 8a. Proving v1 rides with it -the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2): fork branch proving-v1 (from a24ab01a: segment records, the pool split, five RPCs, p2p message 72, parameters proving_v1_activation_daa, segment_blocks, unproven_daa, aggregator_share_bps, in the digest only once set) and its app branch, rebased onto the 0.3.10 tips. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. +the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 on 5b0d54f (90d3299, the final hash follows its gate tests). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 03a6b203d..0028ba93e 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -273,3 +273,7 @@ cpu-prove-pc1-small2 finished 20:59:49Z: the SP1 CPU prover on PC 1 with the min 21:01. GitHub Actions is in a major outage (six queued runs since 19:26Z, none acquired); the coordinator gave the 0.3.10 shipper the fallback at 21:00Z: build the Windows installer on PC 1 (MSVC window host, the payload under Git Bash, Inno Setup; CPU only, about 15 min). PC 1 order now: the repro run (until about 21:09), then the 0.3.10 installer build (the fleet's release, ahead of every measurement), then the era job, the 5090 power sweep, the hot table, Ember Tune. Any measurement that straddles the build window is re-run. 21:02. The Mac measure lock, found by `lsof`: the recorded holder pid 43916 is dead; the files are held open by two WAITERS, the readwidth agent's re-queued footprint loop (pid 78893, holding the measure and build files 11 min, waiting for the three build slots, which cargo tests keep re-acquiring: a convoy) and the epoch agent's compile-ahead measurement (pid 78476, waiting behind it). The readwidth agent is asked to kill 78893; the epoch measurement then runs when the build slots drain. Two defects for the next cut (one task filed): the status line shows the last writer, not the holder; a measure waiter can hold the master lock while build slots keep being granted to new builds, so a measurement can wait indefinitely under a steady stream of cargo tests. + +## 21:04 proving v1 handoff received; the AMD-proving line on the site; amd-prove merged + +Proving v1 for 0.3.11 (rollout plan 8a): fork ece42979 on 21d4c73c, harness PASSED (21 checks) in 244.4 s at 20:56:45Z on the final fork tree, override fields and the mixed-fleet rule recorded; the app's final hash follows its gate tests (90d3299 before it). Branch amd-prove (f1d7a7d) merged into ca2-coord (the append-only bench-log conflict kept both entries); its finding: no zkVM proves on an AMD GPU as of 5 October 2026 (SP1 CPU and CUDA; RISC Zero and ICICLE add Metal; nothing for AMD), the CPU fallback is about 5 minutes per small shard at 30 GB RSS, not a tier. The public line is applied on ca2-coord (1c8439f) with the proving agent's measured tiers in place of the doc's 16 and 20 GB: "Proving needs an NVIDIA card with 24 GB or more (32 GB until the fee switch of 6 October 2026; from it a 24 GB card mines and proves on the same card: 22.2 GB peak with the miner on). AMD and Apple cards mine. A prover for them lands when a zkVM ships one." It replaces "the card mines and proves" on the index, the litepaper's vs RandomX row and proving section, and the miner page (title, meta, hero, feature). Left as it was: the app's Proving tile text (the proving agent's). From 1319adf588243148368e741d83b199858a0f04d6 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:05:20 +0000 Subject: [PATCH 058/131] Counter ASIC 2.0 status 21:05: the node merge ba43cf0f, the hot-table package, unproven_daa at fast time --- docs/plans/counter-asic-2-status.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 0028ba93e..ec66d3535 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -277,3 +277,9 @@ cpu-prove-pc1-small2 finished 20:59:49Z: the SP1 CPU prover on PC 1 with the min ## 21:04 proving v1 handoff received; the AMD-proving line on the site; amd-prove merged Proving v1 for 0.3.11 (rollout plan 8a): fork ece42979 on 21d4c73c, harness PASSED (21 checks) in 244.4 s at 20:56:45Z on the final fork tree, override fields and the mixed-fleet rule recorded; the app's final hash follows its gate tests (90d3299 before it). Branch amd-prove (f1d7a7d) merged into ca2-coord (the append-only bench-log conflict kept both entries); its finding: no zkVM proves on an AMD GPU as of 5 October 2026 (SP1 CPU and CUDA; RISC Zero and ICICLE add Metal; nothing for AMD), the CPU fallback is about 5 minutes per small shard at 30 GB RSS, not a tier. The public line is applied on ca2-coord (1c8439f) with the proving agent's measured tiers in place of the doc's 16 and 20 GB: "Proving needs an NVIDIA card with 24 GB or more (32 GB until the fee switch of 6 October 2026; from it a 24 GB card mines and proves on the same card: 22.2 GB peak with the miner on). AMD and Apple cards mine. A prover for them lands when a zkVM ships one." It replaces "the card mines and proves" on the index, the litepaper's vs RandomX row and proving section, and the miner page (title, meta, hero, feature). Left as it was: the app's Proving tile text (the proving agent's). + +## 21:05 the node merge for 0.3.11 is done; the hot-table package is ready; unproven_daa 10 at fast time + +ca2-v3-node: 2e464e81 (the class switch, the era stand-in, the template, the miner) and the merge of proving-v1 ece42979 = ba43cf0f. One conflict, params.rs's digest-test edits array, resolved by keeping both sides (14 entries); every other shared hunk auto-merged as separate blocks. The merged tree is in cargo check (with igneum-exec and kaspa-p2p-flows); "ready for PC 2" follows with ba43cf0f once it and the Mac pow and consensus-core tests are green. Main repo ca2-v3: ca2-mixer fast-forwarded (66eeba3), then 43ca289 (the fast-time and redteam overrides carry the proving v1 fields and pow_genesis_dataset_log2 28; class-v3.mjs prints the dataset build ms per epoch). proving_v1_unproven_daa is a DAA clock (exec/src/proving.rs segment_status), so 10 at fast time is right (the proving agent confirms; its harness passes --unproven itself). The PC 2 suite command is the node agent's final shape (75-minute budget, six node crates plus the app tests), published from the ca2-v3 worktree when PC 2 is idle (the aggregation-cost job closes at about 21:21). + +ca2-cache 196db96 (rebased on ca2-v3 464d6e1, the added form): 47 + 13 tests; the three added packs bit-exact on Metal and Apple OpenCL (fingerprints at 2^20, base 0: hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5; 96/96; hot table PASS); first-pass rates under load (not numbers): 25.7 / 23.9 / 22.9 MH/s against v2 27.7 (the probe predicts 0.96 / 0.94 / 0.94); the measured rows wait for the measure lock behind the epoch measurement. PC package: ~/Desktop/igneum-ca2-hot.zip sha256 bd49faa1c9d48024f49c615481faff5c68a4c09f0889dbaf009c208674d67b3f (workers 956c4ab3... and 32d3d343... from 196db96, eight packs), fetch-ca2-hot-20261005, playbooks ca2-hot-5090-bench.ps1 and ca2-hot-9070-bench.ps1 (gfx1036 fallback). Its go follows the 0.3.10 build on PC 1 unless the era package is there first. From 451b02ea42dcb043b46620d6f3fafce4d4d6fa49 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:11:19 +0000 Subject: [PATCH 059/131] Counter ASIC 2.0 status 21:11: the 9070 XT back on the bus, PC 1 to the 0.3.10 build, the playbook parse-gate rule --- docs/plans/counter-asic-2-rollout.md | 1 + docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 826559f0e..626e41e11 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -92,6 +92,7 @@ Public lines to fix on the integration branch (card-lifetime table 3): `site/ind |---|---|---|---| | 5 October, at install (earlier today) | the RX 9070 XT in the Sonnet Breakaway Box 850T5 over USB4 went Code 43 | came back after a driver reinstall and a reboot | | | 5 October, about 20:40 | the 9070 XT dropped off PC 1's bus: Get-PnpDevice -Class Display lists only the integrated AMD Radeon Graphics (gfx1036) and the RTX 5090; after `pnputil /scan-devices` at 20:45:34Z the card is still absent and the USB4 list shows only the host and root routers: the "USB4 Router (2.0), Sonnet Technologies Breakaway Box 850T5" present at 17:18Z is gone, so the box itself is off the link | the AMD worker (igneum-worker-opencl --device 1) mines the gfx1036 at 3.12 MH/s; the 5090 keeps mining; nobody was woken, PC 1's app was not restarted; a 10-second rescan probe (pnputil /scan-devices, the USB4 router status) was granted | the second eGPU link fault today: reseat the USB4 cable and the eGPU's power; the 0.3.10 hot-plug code shows the card as "removed" and picks it up again without a restart | +| 5 October, by 21:09 | the 9070 XT is back on PC 1's bus: the reproducible-benchmark package's OpenCL worker listed it as opencl:1 and the app switched it off and on through api/cards; no restart, nobody touched the box | the app's AMD worker returns to it at the next prepare | the link drops and returns by itself; the reseat is still worth doing in the morning | ## 7a. A dated constraint from the consequences review (C1) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index ec66d3535..fe2754889 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -283,3 +283,7 @@ Proving v1 for 0.3.11 (rollout plan 8a): fork ece42979 on 21d4c73c, harness PASS ca2-v3-node: 2e464e81 (the class switch, the era stand-in, the template, the miner) and the merge of proving-v1 ece42979 = ba43cf0f. One conflict, params.rs's digest-test edits array, resolved by keeping both sides (14 entries); every other shared hunk auto-merged as separate blocks. The merged tree is in cargo check (with igneum-exec and kaspa-p2p-flows); "ready for PC 2" follows with ba43cf0f once it and the Mac pow and consensus-core tests are green. Main repo ca2-v3: ca2-mixer fast-forwarded (66eeba3), then 43ca289 (the fast-time and redteam overrides carry the proving v1 fields and pow_genesis_dataset_log2 28; class-v3.mjs prints the dataset build ms per epoch). proving_v1_unproven_daa is a DAA clock (exec/src/proving.rs segment_status), so 10 at fast time is right (the proving agent confirms; its harness passes --unproven itself). The PC 2 suite command is the node agent's final shape (75-minute budget, six node crates plus the app tests), published from the ca2-v3 worktree when PC 2 is idle (the aggregation-cost job closes at about 21:21). ca2-cache 196db96 (rebased on ca2-v3 464d6e1, the added form): 47 + 13 tests; the three added packs bit-exact on Metal and Apple OpenCL (fingerprints at 2^20, base 0: hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5; 96/96; hot table PASS); first-pass rates under load (not numbers): 25.7 / 23.9 / 22.9 MH/s against v2 27.7 (the probe predicts 0.96 / 0.94 / 0.94); the measured rows wait for the measure lock behind the epoch measurement. PC package: ~/Desktop/igneum-ca2-hot.zip sha256 bd49faa1c9d48024f49c615481faff5c68a4c09f0889dbaf009c208674d67b3f (workers 956c4ab3... and 32d3d343... from 196db96, eight packs), fetch-ca2-hot-20261005, playbooks ca2-hot-5090-bench.ps1 and ca2-hot-9070-bench.ps1 (gfx1036 fallback). Its go follows the 0.3.10 build on PC 1 unless the era package is there first. + +## 21:11 the 9070 XT is back; PC 1 to the 0.3.10 build; the repro run failed at parse time + +The repro run (run-repro-pc1-20261005, 21:04 to 21:09:24Z) switched every card off and on and restored them, and found the 9070 XT (gfx1201) ON the bus again (opencl:1; the app switched it off and on), so the era and hot-table jobs run their gfx1201 halves as planned and the G1 ruling's fallback is not needed unless the link drops again; recorded in the rollout plan's hardware events. The run produced no numbers: repro.ps1 failed with a PowerShell parse error on each card (MissingEndParenthesisInExpression), the second playbook tonight that passed no local parse (the Mac has no pwsh); the class fix ordered: every PowerShell playbook parses itself on the PC as its first step (System.Management.Automation.Language.Parser::ParseFile, errors printed, non-zero exit), the way the dot4 playbook gates its bash body with bash -n. "PC 1 is yours" given to the 0.3.10 shipper at 21:10 for the installer build (CPU only, about 15 min); then the hot-table and era jobs, the 5090 power sweep, the repro re-run (about 22:00), Ember Tune. From 5a6c30560abcff85c14fd385912424bff9fe8169 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:20:13 +0000 Subject: [PATCH 060/131] Counter ASIC 2.0: the proving v1 app branch final 6dc686a; the 0.3.11 tree composition; status 21:20 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 626e41e11..925b4d128 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -100,7 +100,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's ## 8a. Proving v1 rides with it -the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 on 5b0d54f (90d3299, the final hash follows its gate tests). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. +the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL 6dc686a on 5b0d54f (the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index fe2754889..de855b439 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -287,3 +287,5 @@ ca2-cache 196db96 (rebased on ca2-v3 464d6e1, the added form): 47 + 13 tests; th ## 21:11 the 9070 XT is back; PC 1 to the 0.3.10 build; the repro run failed at parse time The repro run (run-repro-pc1-20261005, 21:04 to 21:09:24Z) switched every card off and on and restored them, and found the 9070 XT (gfx1201) ON the bus again (opencl:1; the app switched it off and on), so the era and hot-table jobs run their gfx1201 halves as planned and the G1 ruling's fallback is not needed unless the link drops again; recorded in the rollout plan's hardware events. The run produced no numbers: repro.ps1 failed with a PowerShell parse error on each card (MissingEndParenthesisInExpression), the second playbook tonight that passed no local parse (the Mac has no pwsh); the class fix ordered: every PowerShell playbook parses itself on the PC as its first step (System.Management.Automation.Language.Parser::ParseFile, errors printed, non-zero exit), the way the dot4 playbook gates its bash body with bash -n. "PC 1 is yours" given to the 0.3.10 shipper at 21:10 for the installer build (CPU only, about 15 min); then the hot-table and era jobs, the 5090 power sweep, the repro re-run (about 22:00), Ember Tune. + +21:20. Proving v1 app branch final: 6dc686a on 5b0d54f (provedefault 6 of 6; the app's 0.3.11 inputs are complete on that side); fork stays ece42979 (merged into ca2-v3-node ba43cf0f). The 0.3.11 app tree = the 0.3.10 release tree 5b0d54f + proving-v1 6dc686a + whatever the app needs for v3 (the --prepare-packs flag is already there; the Metal worker change is in the worker, not the app); the 0.3.11 main tree = master + ca2-v3 (igneum-pow, workers, fast-time, docs) + ca2-coord (the plans, the spec, the site copy). From 8420f074d5bec621acaab2ea49637174d67e07ea Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:20:25 +0000 Subject: [PATCH 061/131] Counter ASIC 2.0 status 21:20: layer 9 complete with the compile-ahead table and the floor --- docs/plans/counter-asic-2-status.md | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index de855b439..d9e8e7510 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -289,3 +289,19 @@ ca2-cache 196db96 (rebased on ca2-v3 464d6e1, the added form): 47 + 13 tests; th The repro run (run-repro-pc1-20261005, 21:04 to 21:09:24Z) switched every card off and on and restored them, and found the 9070 XT (gfx1201) ON the bus again (opencl:1; the app switched it off and on), so the era and hot-table jobs run their gfx1201 halves as planned and the G1 ruling's fallback is not needed unless the link drops again; recorded in the rollout plan's hardware events. The run produced no numbers: repro.ps1 failed with a PowerShell parse error on each card (MissingEndParenthesisInExpression), the second playbook tonight that passed no local parse (the Mac has no pwsh); the class fix ordered: every PowerShell playbook parses itself on the PC as its first step (System.Management.Automation.Language.Parser::ParseFile, errors printed, non-zero exit), the way the dot4 playbook gates its bash body with bash -n. "PC 1 is yours" given to the 0.3.10 shipper at 21:10 for the installer build (CPU only, about 15 min); then the hot-table and era jobs, the 5090 power sweep, the repro re-run (about 22:00), Ember Tune. 21:20. Proving v1 app branch final: 6dc686a on 5b0d54f (provedefault 6 of 6; the app's 0.3.11 inputs are complete on that side); fork stays ece42979 (merged into ca2-v3-node ba43cf0f). The 0.3.11 app tree = the 0.3.10 release tree 5b0d54f + proving-v1 6dc686a + whatever the app needs for v3 (the --prepare-packs flag is already there; the Metal worker change is in the worker, not the app); the 0.3.11 main tree = master + ca2-v3 (igneum-pow, workers, fast-time, docs) + ca2-coord (the plans, the spec, the site copy). + +## 21:20 layer 9 complete (ca2-epoch 4300608, e95e8b5) + +Design (reserve-only): epoch_len base 3,600, ladder 600 to 7,200, SIGNAL ONLY (the era stream consumes draw 8 and ignores it); day-anchored epochs so a change lands in days; VDF option A (T_epoch 600 s and the 1,200-s lead stay genesis constants: the program is known 600 s ahead at every length, grinding margin 300x; option B rejected at 50x and a 200-s-deep checkpoint); REF_WINDOW_V2 = min(600, L); the 144-s settle per step is 24% of a 600-s epoch. Floor 600 DAA s: the slowest compile-ahead is the variant race at 38 s on the Mac and the 5090, 6.3% of the epoch and inside the window. + +| Card | Compile-ahead | Share of a 600-s epoch | +|---|---|---| +| M5 Max Metal, race off (measured 21:18 UTC: 15.9 / 17.7 / 20.4 ms over 10 fresh programs, pack 79 ms then 1 ms) | 0.5 s (1.8 s cold) | 0.1% | +| M5 Max, race on (M11) | 38 s | 6.3% | +| RTX 5090 NVRTC, race off (M11, hot-swap) | 1.0 s | 0.2% | +| RTX 5090, race on | 38 s | 6.3% | +| RX 9070 XT OpenCL | 0.31 s + compile owed | | +| Intel UHD OpenCL (M11) | 6.4 s | 1.1% | +| Radeon iGPU under load (M11) | 124 s with the dataset | 20.7% (needs per-day dataset reuse before any signal below the base) | + +FPGA citations: PRflow (FPT 2019) 42 min typical, 160 min worst for a monolithic Vivado compile; Aldec hours on Virtex UltraScale; partial reconfiguration milliseconds per region (ICAP 400 MB/s) shortens the load, not the compile; at 600 s a per-program bitstream mines 0% of each epoch, 47% at 3,600. Riders: the race defaults off (M11: base wins on both cards; 6.3% at the floor); per-day dataset reuse in the workers for the iGPU tier. Owed: the 9070 XT clBuildProgram time, the difficulty settle at a 15% step and 600-s epochs in sim.py, the 3-bit signal encoding against Kaspa's version bits (spec 5.8). From 41afd097cd357a1043d696f666095542ec89de10 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:20:53 +0000 Subject: [PATCH 062/131] Counter ASIC 2.0 status 21:20: the added-form hot table on the Mac, the decision rule for layer 5 --- docs/plans/counter-asic-2-status.md | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index d9e8e7510..8eb71e997 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -305,3 +305,15 @@ Design (reserve-only): epoch_len base 3,600, ladder 600 to 7,200, SIGNAL ONLY (t | Radeon iGPU under load (M11) | 124 s with the dataset | 20.7% (needs per-day dataset reuse before any signal below the base) | FPGA citations: PRflow (FPT 2019) 42 min typical, 160 min worst for a monolithic Vivado compile; Aldec hours on Virtex UltraScale; partial reconfiguration milliseconds per region (ICAP 400 MB/s) shortens the load, not the compile; at 600 s a per-program bitstream mines 0% of each epoch, 47% at 3,600. Riders: the race defaults off (M11: base wins on both cards; 6.3% at the floor); per-day dataset reuse in the workers for the iGPU tier. Owed: the 9070 XT clBuildProgram time, the difficulty settle at a 15% step and 600-s epochs in sim.py, the 3-bit signal encoding against Kaspa's version bits (spec 5.8). + +## 21:20 the hot table in the added form, measured on the Mac: the big-die chip comes out ahead + +ca2-cache (hot-table.md 6.2 to 6.4): M5 Max, 21:03 to 21:19 UTC, load average 7 to 14, Metal packbench 2^24 x 5 (GPU time) and Apple OpenCL --bench-pack; v2 in the same session 27.63 / 27.59 MH/s. + +| Pack | Metal / OpenCL MH/s | g against v2 (probe predicted) | Fingerprint (2^20, base 0) | Verifier ms per warp (v2 0.602) | Fill, one core | +|---|---|---|---|---|---| +| hot32k4a | 25.76 / 25.72 | 0.93 (0.96) | 8a3414735db4523c | 0.631 | 21.7 ms | +| hot64k4a | 23.92 / 23.87 | 0.87 (0.94) | 45668f34105f6307 | 0.609 | 43.3 ms | +| hot96k4a | 22.92 / 22.88 | 0.83 (0.93) | af763997dfee4c82 | 0.614 | 64.9 ms | + +Reading: on Apple the added form costs 7 to 17% of the rate for four extra loads per iteration, more than the probe predicts as the table grows; only the 32 MiB table is near free. Chip arithmetic with these g: a chip serving H from DRAM keeps 0.86 / 0.92 / 0.96 of its gain; a 100 mm^2 die with the SRAM keeps 0.96 / 0.92 / 0.88; a 750 mm^2 die with the SRAM comes out 6 to 15% AHEAD, because the GPU pays the hits in rate and a big die pays them in 1.6 to 4.8% of area. So on the Mac's numbers layer 5 does not pass its own test; the decision waits for the 5090 (96 MiB L2) and 9070 XT (64 MB Infinity Cache) rows, where the hits may be near free (g close to 1). Rule for the decision: layer 5 goes into v3 only if, on every card we own, g is at or above 0.97 at the chosen size AND the on-die-cache chip row (chip-model-v3.md) moves down with it; otherwise layer 5 is out of v3 and stays a measured option for 3.0. From e8547b1bcb5831f0e6aa0c6d3a7b744ea1529662 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:21:27 +0000 Subject: [PATCH 063/131] Counter ASIC 2.0 status 21:21: PC 2 to agg-cost-pc2-2 until 21:45, the PC 1 order --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 8eb71e997..16b533179 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -317,3 +317,5 @@ ca2-cache (hot-table.md 6.2 to 6.4): M5 Max, 21:03 to 21:19 UTC, load average 7 | hot96k4a | 22.92 / 22.88 | 0.83 (0.93) | af763997dfee4c82 | 0.614 | 64.9 ms | Reading: on Apple the added form costs 7 to 17% of the rate for four extra loads per iteration, more than the probe predicts as the table grows; only the 32 MiB table is near free. Chip arithmetic with these g: a chip serving H from DRAM keeps 0.86 / 0.92 / 0.96 of its gain; a 100 mm^2 die with the SRAM keeps 0.96 / 0.92 / 0.88; a 750 mm^2 die with the SRAM comes out 6 to 15% AHEAD, because the GPU pays the hits in rate and a big die pays them in 1.6 to 4.8% of area. So on the Mac's numbers layer 5 does not pass its own test; the decision waits for the 5090 (96 MiB L2) and 9070 XT (64 MB Infinity Cache) rows, where the hits may be near free (g close to 1). Rule for the decision: layer 5 goes into v3 only if, on every card we own, g is at or above 0.97 at the chosen size AND the on-die-cache chip row (chip-model-v3.md) moves down with it; otherwise layer 5 is out of v3 and stays a measured option for 3.0. + +21:21. PC 2: agg-cost-pc2-1 closes at about 21:25Z (its own-miner phases ran the iGPU miner by a script fault; the curve is unmeasured); "go PC 2" given for agg-cost-pc2-2 (about 17 min, to 21:44Z), release due by 21:45Z; the 0.3.11 suites on ba43cf0f take PC 2 next (the node agent's Mac suites, release build and fast-time gate are in flight). PC 1: the 0.3.10 installer build (from 21:10); the hot-table job (package in hand) goes the moment the shipper reports the build closed; the era package is still being built. From ced3b703d4b02d44e5a1026846f63e3606e299ec Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:22:20 +0000 Subject: [PATCH 064/131] Counter ASIC 2.0 status 21:22: Ember Tune queued --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 16b533179..4e9a84531 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -319,3 +319,5 @@ ca2-cache (hot-table.md 6.2 to 6.4): M5 Max, 21:03 to 21:19 UTC, load average 7 Reading: on Apple the added form costs 7 to 17% of the rate for four extra loads per iteration, more than the probe predicts as the table grows; only the 32 MiB table is near free. Chip arithmetic with these g: a chip serving H from DRAM keeps 0.86 / 0.92 / 0.96 of its gain; a 100 mm^2 die with the SRAM keeps 0.96 / 0.92 / 0.88; a 750 mm^2 die with the SRAM comes out 6 to 15% AHEAD, because the GPU pays the hits in rate and a big die pays them in 1.6 to 4.8% of area. So on the Mac's numbers layer 5 does not pass its own test; the decision waits for the 5090 (96 MiB L2) and 9070 XT (64 MB Infinity Cache) rows, where the hits may be near free (g close to 1). Rule for the decision: layer 5 goes into v3 only if, on every card we own, g is at or above 0.97 at the chosen size AND the on-die-cache chip row (chip-model-v3.md) moves down with it; otherwise layer 5 is out of v3 and stays a measured option for 3.0. 21:21. PC 2: agg-cost-pc2-1 closes at about 21:25Z (its own-miner phases ran the iGPU miner by a script fault; the curve is unmeasured); "go PC 2" given for agg-cost-pc2-2 (about 17 min, to 21:44Z), release due by 21:45Z; the 0.3.11 suites on ba43cf0f take PC 2 next (the node agent's Mac suites, release build and fast-time gate are in flight). PC 1: the 0.3.10 installer build (from 21:10); the hot-table job (package in hand) goes the moment the shipper reports the build closed; the era package is still being built. + +21:22. Ember Tune (f9bf552 on ember-tune) is queued: its Windows app build (8 min, CPU) in the first gap after the installer build, its baseline run on the 5090 (8 min, no prompt; the 9070 XT if it can be taken without one) after the repro re-run. PC 1 order: the 0.3.10 installer build (running), the hot-table job, the era job, the 5090 power sweep, Ember's build in the first CPU gap, the repro re-run (about 22:00), Ember's run. From 877b45a1168e98b4818e8e2b2f8914fc35e97c13 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:22:37 +0000 Subject: [PATCH 065/131] Counter ASIC 2.0 status 21:22: the installer build job failed in 2 s, PC 1 decision pending --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 4e9a84531..8a23fe346 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -321,3 +321,5 @@ Reading: on Apple the added form costs 7 to 17% of the rate for four extra loads 21:21. PC 2: agg-cost-pc2-1 closes at about 21:25Z (its own-miner phases ran the iGPU miner by a script fault; the curve is unmeasured); "go PC 2" given for agg-cost-pc2-2 (about 17 min, to 21:44Z), release due by 21:45Z; the 0.3.11 suites on ba43cf0f take PC 2 next (the node agent's Mac suites, release build and fast-time gate are in flight). PC 1: the 0.3.10 installer build (from 21:10); the hot-table job (package in hand) goes the moment the shipper reports the build closed; the era package is still being built. 21:22. Ember Tune (f9bf552 on ember-tune) is queued: its Windows app build (8 min, CPU) in the first gap after the installer build, its baseline run on the 5090 (8 min, no prompt; the 9070 XT if it can be taken without one) after the repro re-run. PC 1 order: the 0.3.10 installer build (running), the hot-table job, the era job, the 5090 power sweep, Ember's build in the first CPU gap, the repro re-run (about 22:00), Ember's run. + +21:22. PC 1: the 0.3.10 installer build job (job-fb-installer-pc1) FAILED at 21:20:05Z, exit 2 after 2 s (the script, not the build); the shipper is asked to republish within 5 minutes or yield PC 1 to the hot-table measurement (10 min) and follow it. PC 2: agg-cost-pc2-1 still running at 21:21. From 123af2d107475e9cf134ff7711eacbe44c9cdec0 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:23:10 +0000 Subject: [PATCH 066/131] Counter ASIC 2.0 status 21:23: the installer build republished, PC 1 order --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 8a23fe346..00120d2d5 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -323,3 +323,5 @@ Reading: on Apple the added form costs 7 to 17% of the rate for four extra loads 21:22. Ember Tune (f9bf552 on ember-tune) is queued: its Windows app build (8 min, CPU) in the first gap after the installer build, its baseline run on the 5090 (8 min, no prompt; the 9070 XT if it can be taken without one) after the repro re-run. PC 1 order: the 0.3.10 installer build (running), the hot-table job, the era job, the 5090 power sweep, Ember's build in the first CPU gap, the repro re-run (about 22:00), Ember's run. 21:22. PC 1: the 0.3.10 installer build job (job-fb-installer-pc1) FAILED at 21:20:05Z, exit 2 after 2 s (the script, not the build); the shipper is asked to republish within 5 minutes or yield PC 1 to the hot-table measurement (10 min) and follow it. PC 2: agg-cost-pc2-1 still running at 21:21. + +21:23. The shipper republished the installer build as fb-installer-pc1-2 (the first exit 2 was Test-Path on a \\wsl$ root path refused to the non-elevated session; the engine is now copied out with wsl -u root), cap 25 min, so PC 1 is the build's until about 21:50; the AMD presence probe (10 s, nothing held) runs beside it. Then: the hot-table job, the era job, the 5090 power sweep, Ember's build and fetches, the repro re-run, the AMD sweep (26 min), Ember's run (25 min). From ec0c36f611ab343bdfa264a1e599b18fbe42e994 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:23:38 +0000 Subject: [PATCH 067/131] Counter ASIC 2.0 status 21:23: the era package, the V3_CLASS composition rule --- docs/plans/counter-asic-2-status.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 00120d2d5..60083c48d 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -325,3 +325,9 @@ Reading: on Apple the added form costs 7 to 17% of the rate for four extra loads 21:22. PC 1: the 0.3.10 installer build job (job-fb-installer-pc1) FAILED at 21:20:05Z, exit 2 after 2 s (the script, not the build); the shipper is asked to republish within 5 minutes or yield PC 1 to the hot-table measurement (10 min) and follow it. PC 2: agg-cost-pc2-1 still running at 21:21. 21:23. The shipper republished the installer build as fb-installer-pc1-2 (the first exit 2 was Test-Path on a \\wsl$ root path refused to the non-elevated session; the engine is now copied out with wsl -u root), cap 25 min, so PC 1 is the build's until about 21:50; the AMD presence probe (10 s, nothing held) runs beside it. Then: the hot-table job, the era job, the 5090 power sweep, Ember's build and fetches, the repro re-run, the AMD sweep (26 min), Ember's run (25 min). + +## 21:23 the era package is ready; the V3_CLASS composition rule + +ca2-era PC 1 package: ~/Desktop/igneum-ca2-era-pc1.zip sha256 f26f997602d94b0a408a6974ee484a7b3249ce18dc91dc00783aa9754e5ff040 (workers 5dd3bc16... and 9615efb3... from ca2-era on ca2-v3 464d6e1 with the pack-loop packfile.h merge and the host.c mh_word fix; packs v2 and era-0 to era-5, the six as class v3 chain packs, generator 3, program id 6b02c7c49eb126bd shared, width 4 B, the same devnet epoch seed and day); fetch-ca2-era-20261005; relay/playbooks/ca2-era-pc1.ps1 (the 5090 then gfx1201, gfx1036 fallback; 6 to 10 min). Mac bit-exactness on these packs: Metal 6/6, Apple OpenCL 6/6 (same fingerprints), CUDA CPU emulation 6/6. Its go follows the hot-table job, about 22:00. + +Composition rule for the integration (three branches define V3_CLASS): V3_CLASS = LoadClass::MX4's fields (v2 loads, mixer_mult 4, growth on) + era: None (drawn per program inside generate_from_seed_bytes_program_class from the era bytes, LoadClass::era(V3_CLASS, era, &V3_ALLOWED)) + hot: None until the PC rows decide layer 5; the layout rides with the program (program.class.layout() in the interpreter and Epoch::dataset_word), so one day cache (chain_dataset_day) serves every era. The era agent rebases onto ca2-v3 HEAD with that literal; the node agent takes it into the integration tree. The era stream is seeded from "igneum-era/" || E_n (the index dropped: E_n commits to n through the VDF input); the spec text in era-layout.md says so. From c0c44cab3936dfffe47076ac27e6ca98500c6f2e Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:24:45 +0000 Subject: [PATCH 068/131] Counter ASIC 2.0 status 21:24: the third eGPU drop, the miners are mining, the installer build running --- docs/plans/counter-asic-2-rollout.md | 1 + docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 925b4d128..57e1922ea 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -93,6 +93,7 @@ Public lines to fix on the integration branch (card-lifetime table 3): `site/ind | 5 October, at install (earlier today) | the RX 9070 XT in the Sonnet Breakaway Box 850T5 over USB4 went Code 43 | came back after a driver reinstall and a reboot | | | 5 October, about 20:40 | the 9070 XT dropped off PC 1's bus: Get-PnpDevice -Class Display lists only the integrated AMD Radeon Graphics (gfx1036) and the RTX 5090; after `pnputil /scan-devices` at 20:45:34Z the card is still absent and the USB4 list shows only the host and root routers: the "USB4 Router (2.0), Sonnet Technologies Breakaway Box 850T5" present at 17:18Z is gone, so the box itself is off the link | the AMD worker (igneum-worker-opencl --device 1) mines the gfx1036 at 3.12 MH/s; the 5090 keeps mining; nobody was woken, PC 1's app was not restarted; a 10-second rescan probe (pnputil /scan-devices, the USB4 router status) was granted | the second eGPU link fault today: reseat the USB4 cable and the eGPU's power; the 0.3.10 hot-plug code shows the card as "removed" and picks it up again without a restart | | 5 October, by 21:09 | the 9070 XT is back on PC 1's bus: the reproducible-benchmark package's OpenCL worker listed it as opencl:1 and the app switched it off and on through api/cards; no restart, nobody touched the box | the app's AMD worker returns to it at the next prepare | the link drops and returns by itself; the reseat is still worth doing in the morning | +| 5 October, by 21:22:59 | the third drop: no Sonnet or USB4 Router (2.0) device present, the display list shows only the gfx1036 and the 5090 | the app lists nvidia:0 and amd:1:gfx1036; the 5090 keeps mining at 124.5 MH/s (app log STATUS lines through 21:23:26Z) | the link is flapping: reseat the USB4 cable and the eGPU's power in the morning, and consider a different USB4 port or cable; every AMD measurement tonight runs on the gfx1036 fallback unless the card is present at the job's own probe | ## 7a. A dated constraint from the consequences review (C1) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 60083c48d..1134c9b39 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -331,3 +331,5 @@ Reading: on Apple the added form costs 7 to 17% of the rate for four extra loads ca2-era PC 1 package: ~/Desktop/igneum-ca2-era-pc1.zip sha256 f26f997602d94b0a408a6974ee484a7b3249ce18dc91dc00783aa9754e5ff040 (workers 5dd3bc16... and 9615efb3... from ca2-era on ca2-v3 464d6e1 with the pack-loop packfile.h merge and the host.c mh_word fix; packs v2 and era-0 to era-5, the six as class v3 chain packs, generator 3, program id 6b02c7c49eb126bd shared, width 4 B, the same devnet epoch seed and day); fetch-ca2-era-20261005; relay/playbooks/ca2-era-pc1.ps1 (the 5090 then gfx1201, gfx1036 fallback; 6 to 10 min). Mac bit-exactness on these packs: Metal 6/6, Apple OpenCL 6/6 (same fingerprints), CUDA CPU emulation 6/6. Its go follows the hot-table job, about 22:00. Composition rule for the integration (three branches define V3_CLASS): V3_CLASS = LoadClass::MX4's fields (v2 loads, mixer_mult 4, growth on) + era: None (drawn per program inside generate_from_seed_bytes_program_class from the era bytes, LoadClass::era(V3_CLASS, era, &V3_ALLOWED)) + hot: None until the PC rows decide layer 5; the layout rides with the program (program.class.layout() in the interpreter and Epoch::dataset_word), so one day cache (chain_dataset_day) serves every era. The era agent rebases onto ca2-v3 HEAD with that literal; the node agent takes it into the integration tree. The era stream is seeded from "igneum-era/" || E_n (the index dropped: E_n commits to n through the VDF input); the spec text in era-layout.md says so. + +21:24. The 9070 XT dropped off PC 1's bus again (probe #205 at 21:22:59Z: no Sonnet or USB4 router device; the third drop today; rollout plan 7b updated: the link is flapping). The era and hot-table jobs run their gfx1036 fallback unless the card is present at run time; the AMD sweep slot is conditional on a presence probe at 22:00. A probe reading of "miners off at 0.0 MH/s" was wrong: the app log shows the 5090 at 124.5 MH/s through 21:23:26Z; the AMD agent fixes its probe's precondition. The 0.3.10 installer build fb-installer-pc1-2 runs on PC 1 (cap 25 min from 21:23). From 48baa91575139c9d82a0ed2b8363e09280300ef9 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:25:03 +0000 Subject: [PATCH 069/131] Counter ASIC 2.0 status 21:25: the mixer x4 bit-exact on the Mac, the chip row 1.84x, the hot table's effect on the row --- docs/plans/counter-asic-2-status.md | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 1134c9b39..0b075be94 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -333,3 +333,13 @@ ca2-era PC 1 package: ~/Desktop/igneum-ca2-era-pc1.zip sha256 f26f997602d94b0a40 Composition rule for the integration (three branches define V3_CLASS): V3_CLASS = LoadClass::MX4's fields (v2 loads, mixer_mult 4, growth on) + era: None (drawn per program inside generate_from_seed_bytes_program_class from the era bytes, LoadClass::era(V3_CLASS, era, &V3_ALLOWED)) + hot: None until the PC rows decide layer 5; the layout rides with the program (program.class.layout() in the interpreter and Epoch::dataset_word), so one day cache (chain_dataset_day) serves every era. The era agent rebases onto ca2-v3 HEAD with that literal; the node agent takes it into the integration tree. The era stream is seeded from "igneum-era/" || E_n (the index dropped: E_n commits to n through the VDF input); the spec text in era-layout.md says so. 21:24. The 9070 XT dropped off PC 1's bus again (probe #205 at 21:22:59Z: no Sonnet or USB4 router device; the third drop today; rollout plan 7b updated: the link is flapping). The era and hot-table jobs run their gfx1036 fallback unless the card is present at run time; the AMD sweep slot is conditional on a presence probe at 22:00. A probe reading of "miners off at 0.0 MH/s" was wrong: the app log shows the 5090 at 124.5 MH/s through 21:23:26Z; the AMD agent fixes its probe's precondition. The 0.3.10 installer build fb-installer-pc1-2 runs on PC 1 (cap 25 min from 21:23). + +## 21:25 the mixer x4 construction is bit-exact on the Mac; the chip row reads 1.84x (thin) + +ca2-mixer commits: 0fc0ad1 (construction, V3_CLASS = mx4, chain_dataset_day body), 66eeba3 (the pinned v3 packs mx4-genesis and mx4-devnet-epoch0, --program-class v3 / --era-hex), e4c04a7 (tests/mixer.rs: 200-program v3 fuzz, stats, edges, determinism; scratch.rs cherry-picked, 7 of 7), 7ce8d1e (docs/analysis/chip-model-v3.md). Spec text for 1.8.5 and 1.13.3 in docs/plans/mixer-x4.md section 2: the multiplied mixer with keys (r m + j + 1) x 0x9E3779B9; option C as doublings(d) = floor(log2(1 + d / 1460)); the day table with the verifier fill per step. + +Bit-exactness (run lock): both v3 packs on Metal and Apple OpenCL, 3/3 standalone and 3/3 in batch, 96 of 96 lanes, dataset head / MASK / 64 samples PASS, one fingerprint per pack across both harnesses (6f48d5a2aa0dbe5f, 73caaebb28e808fe); the 200-pack v3 fuzz on Metal 200 of 200, every tenth on Apple OpenCL 20 of 20; CPU 44 lib, 12 packs, 7 scratch, 4 mixer tests; v2 exports IDENTICAL. Indicative (run lock, not a number): the Metal 1 GiB build at x4 30.2 ms (mx4-genesis) and 21.7 ms (mx4-devnet); the measured verifier and build timings wait on the measure lock. + +The chip row (chip-model-v3.md), v3 at x4: 599,040 ops per hash; the on-die-cache recompute chip at 50 T op/s does 83.5 MH/s, 0.61x bare against the 5090's 136.1, 1.84x with the 3x fixed-function factor, 1.53x with the 128 mm^2 mirror deducted at equal silicon. With the hot table in the ADDED form at the Mac's g: 1.98x (32 MiB) and 2.12x (64 MiB) at equal budget, 1.60x / 1.67x with the SRAM deducted: the added hot table costs the card and not this chip, so it moves the row the WRONG way on the Mac's numbers. The claim holds "under 2x" on the equal-silicon convention, and on the equal-budget one only without the hot table; thin everywhere (a 3.3x factor or 10% on the budget reads 2.0x). Next lever: x8 (0.31x bare, 0.92x with the factor). + +Consequence for layer 5: unless the 5090 and 9070 XT rows show g at or above 0.97 (the hits near free), the hot table stays OUT of v3 tonight (the rule in the 21:20 entry) and the public level 3 names it as a measured option, not a lever. From e2be67fc3fc77fff5787e668662568b68a96ee2c Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:25:28 +0000 Subject: [PATCH 070/131] Counter ASIC 2.0: G6 state and the expected 0.3.11 digest; status 21:25 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 57e1922ea..755a8da48 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,7 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary | open | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c (release-0.3.6 once the shipper lands it) with suites green on PC 2 | the build job id and its SUMMARY line | open | +| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green; PC 2 pending | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 0b075be94..8ac5f9424 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -343,3 +343,5 @@ Bit-exactness (run lock): both v3 packs on Metal and Apple OpenCL, 3/3 standalon The chip row (chip-model-v3.md), v3 at x4: 599,040 ops per hash; the on-die-cache recompute chip at 50 T op/s does 83.5 MH/s, 0.61x bare against the 5090's 136.1, 1.84x with the 3x fixed-function factor, 1.53x with the 128 mm^2 mirror deducted at equal silicon. With the hot table in the ADDED form at the Mac's g: 1.98x (32 MiB) and 2.12x (64 MiB) at equal budget, 1.60x / 1.67x with the SRAM deducted: the added hot table costs the card and not this chip, so it moves the row the WRONG way on the Mac's numbers. The claim holds "under 2x" on the equal-silicon convention, and on the equal-budget one only without the hot table; thin everywhere (a 3.3x factor or 10% on the budget reads 2.0x). Next lever: x8 (0.31x bare, 0.92x with the factor). Consequence for layer 5: unless the 5090 and 9070 XT rows show g at or above 0.97 (the hits near free), the hot table stays OUT of v3 tonight (the rule in the 21:20 entry) and the public level 3 names it as a measured option, not a lever. + +21:25. "ready for PC 2" from the node agent: fork ca2-v3-node 79bd8e10, Mac checks and tests green (rollout plan G6 row); the expected 0.3.11 digest with the two new fields at never is c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c. The coordinator publishes the suite job from the ca2-v3 worktree at 21:45 when the aggregation-cost agent releases PC 2, on the worktree's tip at that moment (the era and cache merges go in first if the era commit arrives in time). From 8981d15f593f4b64b5ac1ea7c1aaa252dad9fc85 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:27:00 +0000 Subject: [PATCH 071/131] Counter ASIC 2.0 status 21:27: installer attempt 3, the fast-failure rule --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 8ac5f9424..55fbd559e 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -345,3 +345,5 @@ The chip row (chip-model-v3.md), v3 at x4: 599,040 ops per hash; the on-die-cach Consequence for layer 5: unless the 5090 and 9070 XT rows show g at or above 0.97 (the hits near free), the hot table stays OUT of v3 tonight (the rule in the 21:20 entry) and the public level 3 names it as a measured option, not a lever. 21:25. "ready for PC 2" from the node agent: fork ca2-v3-node 79bd8e10, Mac checks and tests green (rollout plan G6 row); the expected 0.3.11 digest with the two new fields at never is c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c. The coordinator publishes the suite job from the ca2-v3 worktree at 21:45 when the aggregation-cost agent releases PC 2, on the worktree's tip at that moment (the era and cache merges go in first if the era commit arrives in time). + +21:27. The installer build's second attempt failed in 4 s on a path (the engine is not under /root/igneum-build/app/...); attempt 3 (fb-installer-pc1-3) is publishing with a find-based path; rule: if it fails inside 5 minutes, the hot-table measurement takes PC 1 before attempt 4. The shipper's jobs started no second app instance (the failed attempts exited before any exe ran); its post-build listing of igneum-app.exe processes settles the "off" reading. From 52e41787ecff3f5d14385aa47c7816734c8cad46 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:27:46 +0000 Subject: [PATCH 072/131] Counter ASIC 2.0: the x4/x8 decision rule recorded (delegated); status 21:27 --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 55fbd559e..2ab098092 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -347,3 +347,7 @@ Consequence for layer 5: unless the 5090 and 9070 XT rows show g at or above 0.9 21:25. "ready for PC 2" from the node agent: fork ca2-v3-node 79bd8e10, Mac checks and tests green (rollout plan G6 row); the expected 0.3.11 digest with the two new fields at never is c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c. The coordinator publishes the suite job from the ca2-v3 worktree at 21:45 when the aggregation-cost agent releases PC 2, on the worktree's tip at that moment (the era and cache merges go in first if the era commit arrives in time). 21:27. The installer build's second attempt failed in 4 s on a path (the engine is not under /root/igneum-build/app/...); attempt 3 (fb-installer-pc1-3) is publishing with a find-based path; rule: if it fails inside 5 minutes, the hot-table measurement takes PC 1 before attempt 4. The shipper's jobs started no second app instance (the failed attempts exited before any exe ran); its post-build listing of igneum-app.exe processes settles the "off" reading. + +## 21:27 decision rule for the mixer: x8 beside x4 + +Delegated (coordinator, the project lead's "as strong as the measurements allow"): the mixer agent builds LoadClass::MX8 beside MX4, exports mx8-genesis and mx8-devnet-epoch0, runs the Mac bit-exactness, and measures v2, x4 and x8 in one measure-lock session (verifier ms per warp on one core, avg of 50 and worst cold; the 256 MiB fill; the Metal 1 GiB build); the 5090 and AMD daily-build times come from a 2-minute prepare job on PC 1 after the era job. x8 goes into v3 if the per-warp verify stays under 10 ms on one core AND the daily build stays under 1 s on every card we own; else x4 with the thin margin stated (1.84x with the factor) and x8 named as the next lever (0.92x with the factor, from the m16 table). The vectors are re-cut once after the choice. The hot-table rule stands (into v3 only at g >= 0.97 on both PC cards). From d8e6ad3341b8f1e32a0d270e243493140aa7ca18 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:28:07 +0000 Subject: [PATCH 073/131] Counter ASIC 2.0: the x8 refinement in the rollout plan; status 21:28: hot-table go, installer attempt 3 failed, PC 1 order --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 755a8da48..ab0a9a8c4 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. Refinement (coordinator, 21:28 UTC, delegated under "as strong as the measurements allow"): x8 is built and measured beside x4 on the same packs (verifier per warp on one Mac core, the 1 GiB daily build on the 5090, the M5 Max and the 9070 XT or gfx1036, the chip row at equal silicon with the 3x factor); x8 goes into v3 if the per-warp verify stays under 10 ms on one core and the daily build stays under 1 s on every card we own, otherwise x4 with the thin margin stated in level 3 and x8 named as the next lever. The vectors are re-cut once, after this choice. Measured so far: x4 chip row 0.61x bare, 1.84x with the factor, 1.53x with the 128 mm^2 mirror deducted; x8 from the m16 table 0.31x bare, 0.92x with the factor (approximate until measured). The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ### 6b. The user tiers (the consequences rule) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 2ab098092..c5a14cb2d 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -351,3 +351,5 @@ Consequence for layer 5: unless the 5090 and 9070 XT rows show g at or above 0.9 ## 21:27 decision rule for the mixer: x8 beside x4 Delegated (coordinator, the project lead's "as strong as the measurements allow"): the mixer agent builds LoadClass::MX8 beside MX4, exports mx8-genesis and mx8-devnet-epoch0, runs the Mac bit-exactness, and measures v2, x4 and x8 in one measure-lock session (verifier ms per warp on one core, avg of 50 and worst cold; the 256 MiB fill; the Metal 1 GiB build); the 5090 and AMD daily-build times come from a 2-minute prepare job on PC 1 after the era job. x8 goes into v3 if the per-warp verify stays under 10 ms on one core AND the daily build stays under 1 s on every card we own; else x4 with the thin margin stated (1.84x with the factor) and x8 named as the next lever (0.92x with the factor, from the m16 table). The vectors are re-cut once after the choice. The hot-table rule stands (into v3 only at g >= 0.97 on both PC cards). + +21:28. Installer attempt 3 failed at 21:27:20Z (exit 2 in 4 s: an inline bash -c string lost a quote through PowerShell; the fix is the 0.3.6 cut's: the WSL part as a file run with bash , plus a read-only path probe before attempt 4). By the rule, "go PC 1" went to the hot-table job at 21:28 (10 min). PC 1 order from here: the hot-table job, the shipper's path probe and attempt 4 (about 16 min), the era job, the 5090 power sweep, the mixer daily-build job (2 min), Ember's build, the repro re-run, the AMD sweep (conditional), Ember's run. From 3791a29ebbfde90b35d7ec368d076e0e10b19831 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:31:06 +0000 Subject: [PATCH 074/131] Counter ASIC 2.0 status 21:31: C23 and C24 --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index c5a14cb2d..15ea5fe1c 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -353,3 +353,5 @@ Consequence for layer 5: unless the 5090 and 9070 XT rows show g at or above 0.9 Delegated (coordinator, the project lead's "as strong as the measurements allow"): the mixer agent builds LoadClass::MX8 beside MX4, exports mx8-genesis and mx8-devnet-epoch0, runs the Mac bit-exactness, and measures v2, x4 and x8 in one measure-lock session (verifier ms per warp on one core, avg of 50 and worst cold; the 256 MiB fill; the Metal 1 GiB build); the 5090 and AMD daily-build times come from a 2-minute prepare job on PC 1 after the era job. x8 goes into v3 if the per-warp verify stays under 10 ms on one core AND the daily build stays under 1 s on every card we own; else x4 with the thin margin stated (1.84x with the factor) and x8 named as the next lever (0.92x with the factor, from the m16 table). The vectors are re-cut once after the choice. The hot-table rule stands (into v3 only at g >= 0.97 on both PC cards). 21:28. Installer attempt 3 failed at 21:27:20Z (exit 2 in 4 s: an inline bash -c string lost a quote through PowerShell; the fix is the 0.3.6 cut's: the WSL part as a file run with bash , plus a read-only path probe before attempt 4). By the rule, "go PC 1" went to the hot-table job at 21:28 (10 min). PC 1 order from here: the hot-table job, the shipper's path probe and attempt 4 (about 16 min), the era job, the 5090 power sweep, the mixer daily-build job (2 min), Ember's build, the repro re-run, the AMD sweep (conditional), Ember's run. + +21:31. Consequences round 3. C23: the x8 table gains a gfx1036 row (the integrated tier builds the dataset per PREPARE: 7 to 12 s at x1, 55 to 124 s under load, so 28 to 48 s an hour at x4 and 56 to 96 s at x8, every boundary missed) and a scaled 8 GB-class row; the "under 1 s" rule applies to the discrete cards' daily build; for the integrated tier either per-day dataset reuse in the three workers lands with v3 (the node agent is asked whether it is bounded tonight) or the level 3 page says the iGPU tier mines v3 with a restart per epoch; recorded with the x4/x8 choice. C24: two inline bash bodies lost a quote through PowerShell tonight (amd-prove's awk at 20:44, the installer's bash -c at 21:27); the reviewer's sub-agent adds a repo-wide CI check (every inline bash body through bash -n, an unextractable one fails CI) and the convention line in packaging/README-ship.md; no collision with the PC 1 queue. From 2432a01f6c83646bacf67ce8264f8465074e8da5 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:31:32 +0000 Subject: [PATCH 075/131] Counter ASIC 2.0 status 21:31: the runner recovered, the 0.3.10 rollout, the restart rule --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 15ea5fe1c..9600ce468 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -355,3 +355,7 @@ Delegated (coordinator, the project lead's "as strong as the measurements allow" 21:28. Installer attempt 3 failed at 21:27:20Z (exit 2 in 4 s: an inline bash -c string lost a quote through PowerShell; the fix is the 0.3.6 cut's: the WSL part as a file run with bash , plus a read-only path probe before attempt 4). By the rule, "go PC 1" went to the hot-table job at 21:28 (10 min). PC 1 order from here: the hot-table job, the shipper's path probe and attempt 4 (about 16 min), the era job, the 5090 power sweep, the mixer daily-build job (2 min), Ember's build, the repro re-run, the AMD sweep (conditional), Ember's run. 21:31. Consequences round 3. C23: the x8 table gains a gfx1036 row (the integrated tier builds the dataset per PREPARE: 7 to 12 s at x1, 55 to 124 s under load, so 28 to 48 s an hour at x4 and 56 to 96 s at x8, every boundary missed) and a scaled 8 GB-class row; the "under 1 s" rule applies to the discrete cards' daily build; for the integrated tier either per-day dataset reuse in the three workers lands with v3 (the node agent is asked whether it is bounded tonight) or the level 3 page says the iGPU tier mines v3 with a restart per epoch; recorded with the x4/x8 choice. C24: two inline bash bodies lost a quote through PowerShell tonight (amd-prove's awk at 20:44, the installer's bash -c at 21:27); the reviewer's sub-agent adds a repo-wide CI check (every inline bash body through bash -n, an unextractable one fails CI) and the convention line in packaging/README-ship.md; no collision with the PC 1 queue. + +## 21:31 GitHub's runner recovered: the 0.3.10 rollout starts; the restart rule for every PC job + +The 0.3.10 Windows run 37374158235 went green at 21:30:29Z on the release tree; the three fb-installer jobs are done and expired; nothing more of the shipper's touches PC 1. The rollout runs now (manifest, update-now, hand nodes, seed, digest sweep): every app restarts once within the next quarter hour. Rules: the shipper is asked to hold PC 1's update-now until the hot-table job releases (about 21:40) and the era job starts only after PC 1's new STATUS line; the 0.3.11 suites publish on PC 2 only after PC 2's app shows 0.3.10 (a build job dies with the app); any measurement that straddles a restart is re-run; every PC job's RESULT lines carry the app version before and after. From d6a503ca93829e7561dfe949d80012044c85e56b Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:32:35 +0000 Subject: [PATCH 076/131] Counter ASIC 2.0: the iGPU restart-per-epoch statement, dataset reuse on the next-cut list, the gate's script fix; status 21:32 --- docs/plans/counter-asic-2-public.md | 2 +- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 6 ++++++ 3 files changed, 8 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index eebe3a4c9..f215b3ff5 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -34,7 +34,7 @@ Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per h | RTX 5090 (CUDA) | 136.1 | [owed] | 512 + hot | 0.96 at v2 | [owed] | | RX 9070 XT (OpenCL, eGPU) | 18.15 | [owed] | 512 + hot | 0.87 at v2 | [owed] | -AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), 2.2x worse per pound at list prices and 4.9x worse per watt (read-width.md section 4.1, approximate); the card's memory system, not a tuning gap. +The integrated tier (Radeon iGPU, Intel UHD) on the CUDA and OpenCL one-click workers mines v3 with a restart per hourly epoch, because those workers rebuild the dataset on every prepare (7 to 12 s at x1, about 4x at the x4 mixer; the Metal worker keeps its day's dataset); per-day dataset reuse in those two workers is the first item after the publish (0.3.12). AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), 2.2x worse per pound at list prices and 4.9x worse per watt (read-width.md section 4.1, approximate); the card's memory system, not a tuning gap. Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant]. diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index ab0a9a8c4..1ed5a42d4 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -103,7 +103,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL 6dc686a on 5b0d54f (the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. -Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. +Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. ## 9. After the publish: Counter ASIC 3.0 diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 9600ce468..4a40e2fc7 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -359,3 +359,9 @@ Delegated (coordinator, the project lead's "as strong as the measurements allow" ## 21:31 GitHub's runner recovered: the 0.3.10 rollout starts; the restart rule for every PC job The 0.3.10 Windows run 37374158235 went green at 21:30:29Z on the release tree; the three fb-installer jobs are done and expired; nothing more of the shipper's touches PC 1. The rollout runs now (manifest, update-now, hand nodes, seed, digest sweep): every app restarts once within the next quarter hour. Rules: the shipper is asked to hold PC 1's update-now until the hot-table job releases (about 21:40) and the era job starts only after PC 1's new STATUS line; the 0.3.11 suites publish on PC 2 only after PC 2's app shows 0.3.10 (a build job dies with the app); any measurement that straddles a restart is re-run; every PC job's RESULT lines carry the app version before and after. + +## 21:32 per-day dataset reuse: Metal has it, CUDA and OpenCL do not (0.3.12); the gate re-runs after a script fix + +Node agent: the Metal worker's ServeStore already keys datasets by day and programs by epoch (a prepare on a resident day builds the program only); the CUDA worker.cpp and OpenCL host.c bundle program, cache, dataset and the self-test in one Pair, and splitting a Day object out touches buffer ownership, releasePair, the prepare thread and the self-test in both: over an hour, not shipped untested tonight; first item after the publish (0.3.12; next-cut list). Decision recorded (C23): the level 3 page states that the integrated tier on the one-click workers mines v3 with a restart per epoch (public copy and rollout plan updated). + +Gate G4: the first fast-time run failed at node start on the script, not the node: JSON.parse turned a never height (18446744073709551615) into 1.8446744073709552e+19 and the node refused the override; fixed as text merging in class-v3.mjs and simnet.mjs (d5ff532; no other script in tools/, infra/ or sim/ has the shape). The gate runs again on the mixer-x4 class (binaries from 79bd8e10 + ca2-v3 66eeba3). Main-repo tip for the suite job title: d5ff532 (era and cache not yet merged). From 030858063cf0f79bdbd31340a3db2a3653fd37fa Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:34:16 +0000 Subject: [PATCH 077/131] Counter ASIC 2.0 status 21:34: PC 2 timing, the suites after the restart --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 4a40e2fc7..b2b88c697 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -365,3 +365,5 @@ The 0.3.10 Windows run 37374158235 went green at 21:30:29Z on the release tree; Node agent: the Metal worker's ServeStore already keys datasets by day and programs by epoch (a prepare on a resident day builds the program only); the CUDA worker.cpp and OpenCL host.c bundle program, cache, dataset and the self-test in one Pair, and splitting a Day object out touches buffer ownership, releasePair, the prepare thread and the self-test in both: over an hour, not shipped untested tonight; first item after the publish (0.3.12; next-cut list). Decision recorded (C23): the level 3 page states that the integrated tier on the one-click workers mines v3 with a restart per epoch (public copy and rollout plan updated). Gate G4: the first fast-time run failed at node start on the script, not the node: JSON.parse turned a never height (18446744073709551615) into 1.8446744073709552e+19 and the node refused the override; fixed as text merging in class-v3.mjs and simnet.mjs (d5ff532; no other script in tools/, infra/ or sim/ has the shape). The gate runs again on the mixer-x4 class (binaries from 79bd8e10 + ca2-v3 66eeba3). Main-repo tip for the suite job title: d5ff532 (era and cache not yet merged). + +21:34. PC 2: agg-cost-pc2-1 closed 21:25:11Z (done, miners and prover back on); agg-cost-pc2-2 went out at 21:33Z (a missed close), self-limited to 16.5 min, closes about 21:51Z; the 0.3.11 suites publish after it AND after PC 2's app shows the 0.3.10 STATUS line (the update-now goes to PC 2 now). PC 1: the hot-table job runs (release about 21:40); PC 1's update-now follows the release; the era job starts after PC 1's 0.3.10 STATUS line. From 9eb6b5e654e290ea31f167dd3245e99ae1add5a7 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:36:30 +0000 Subject: [PATCH 078/131] Counter ASIC 2.0: layer 5 decided OUT on the PC rows; status 21:36 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 17 +++++++++++++++++ docs/plans/counter-asic-2.md | 2 +- 3 files changed, 19 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 1ed5a42d4..4f95f216c 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -115,7 +115,7 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa |---|---|---|---| | The width rule (layers 1 and 2) | 4 B fixed; 16 B; 64 B; the per-program mix | `` | `docs/plans/read-width.md` | | The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `` | the same | -| The hot table (layer 5) | 32, 64, 96 MB; replaced or added | In the ADDED form only (16 dataset loads plus k hot loads): the replaced form lets the on-die-cache chip skip item derivations (k = 4: chip x1.33 against the GPU's measured x1.05 to x1.22). Size = the largest table resident on every card we own, pending the PC rows; Mac rows (replaced form, M5 Max): hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k8 x1.71; Apple OpenCL dependent-read probe 32 / 64 / 96 / 1024 MiB: 21.7 / 12.8 / 12.3 / 3.50 G loads/s | `docs/plans/hot-table.md` (ca2-cache 65bc7a7) | +| The hot table (layer 5) | 32, 64, 96 MB; replaced or added | DECIDED (5 October 2026, 21:36 UTC, delegated): OUT of v3. The rule was g at or above 0.97 on both PC cards in the added form; measured g is 0.87 / 0.85 / 0.84 on the 5090 and 0.84 / 0.81 / 0.80 on the 9070 XT at 32 / 64 / 96 MiB: neither card keeps even the 32 MiB table resident while the dataset streams, and the replaced form helps the on-die-cache chip. Layer 5 stays a measured option for 3.0 | `docs/plans/hot-table.md` (ca2-cache): 5090 MH/s v2 136.1; replaced hot32k4 146.6, hot64k4 140.8, hot96k4 138.5, hot64k2 137.5, hot64k8 163.6; added hot32k4a 118.7, hot64k4a 115.4, hot96k4a 114.4. 9070 XT v2 18.15; replaced 19.79 / 18.73 / 18.33 / 18.17 / 22.32; added 15.27 / 14.62 / 14.56. M5 Max added 0.93 / 0.87 / 0.83. All eight packs bit-exact on the 5090 and the 9070 XT with the Mac's fingerprints (21:29 to 21:35 UTC, no restart straddled) | | The era draws (layers 4 and 8) | in, if the min-to-max spread across six drawn eras is under 5% per card; item size fixed at 4 B, draws of stride, interleave and the working-set window at or above 256 MiB | pending the six-era table | `docs/plans/era-layout.md` | | The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate | | Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index b2b88c697..5cd2bc6d5 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -367,3 +367,20 @@ Node agent: the Metal worker's ServeStore already keys datasets by day and progr Gate G4: the first fast-time run failed at node start on the script, not the node: JSON.parse turned a never height (18446744073709551615) into 1.8446744073709552e+19 and the node refused the override; fixed as text merging in class-v3.mjs and simnet.mjs (d5ff532; no other script in tools/, infra/ or sim/ has the shape). The gate runs again on the mixer-x4 class (binaries from 79bd8e10 + ca2-v3 66eeba3). Main-repo tip for the suite job title: d5ff532 (era and cache not yet merged). 21:34. PC 2: agg-cost-pc2-1 closed 21:25:11Z (done, miners and prover back on); agg-cost-pc2-2 went out at 21:33Z (a missed close), self-limited to 16.5 min, closes about 21:51Z; the 0.3.11 suites publish after it AND after PC 2's app shows the 0.3.10 STATUS line (the update-now goes to PC 2 now). PC 1: the hot-table job runs (release about 21:40); PC 1's update-now follows the release; the era job starts after PC 1's 0.3.10 STATUS line. + +## 21:36 layer 5 measured on both PC cards: OUT of v3 + +Hot-table jobs on PC 1 (fetch 21:29:01Z; run-ca2-hot-5090 126 s, closed 21:31:38Z; run-ca2-hot-9070 247 s, closed 21:35:39Z; both cards restored; the 9070 XT WAS on the bus, full path; app 0.3.9 throughout, no straddle). All eight packs bit-exact on the 5090 and the 9070 XT with the Mac's fingerprints. + +| Pack | RTX 5090 MH/s (v2 136.1) | RX 9070 XT MH/s (v2 18.15) | M5 Max g | +|---|---|---|---| +| hot32k4 (replaced) | 146.6 (x1.08) | 19.79 (x1.09) | x1.22 | +| hot64k4 | 140.8 (x1.03) | 18.73 (x1.03) | x1.12 | +| hot96k4 | 138.5 (x1.02) | 18.33 (x1.01) | x1.05 | +| hot64k2 | 137.5 (x1.01) | 18.17 (x1.00) | x1.00 | +| hot64k8 | 163.6 (x1.20) | 22.32 (x1.23) | x1.71 | +| hot32k4a (added) | 118.7 (0.87) | 15.27 (0.84) | 0.93 | +| hot64k4a | 115.4 (0.85) | 14.62 (0.81) | 0.87 | +| hot96k4a | 114.4 (0.84) | 14.56 (0.80) | 0.83 | + +Decision (the 0.97 rule): layer 5 is OUT of v3. Neither card keeps even the 32 MiB table resident while the 1 GiB dataset streams (the replaced form gains 1.02 to 1.08x at k = 4 against an ideal 1.33x), and the added form costs 13 to 20%; the chip row moves the wrong way with it. Layer 5 stays a measured option for 3.0 (a table small enough to stay resident, or a different access pattern). The probe rows and the writeup follow on ca2-cache. PC 1 is released to the shipper for PC 1's update-now; the era job starts after PC 1's 0.3.10 STATUS line. diff --git a/docs/plans/counter-asic-2.md b/docs/plans/counter-asic-2.md index 8e9f58d86..542ef627c 100644 --- a/docs/plans/counter-asic-2.md +++ b/docs/plans/counter-asic-2.md @@ -26,7 +26,7 @@ Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGP | 2 | out | spread across six programs 5.5 to 22.3% per card, over the 5% rule | | 3 | out (scratch share 0); the construct is sound and its tests stay | the on-die-cache recompute chip stays at 2.4x at every share under the 6 GB cap | | 4 and 8 | the era draw of stride, interleave and the working-set window, in if the six-era spread is under 5% per card | pending the PC rows | -| 5 | in, in the ADDED form only (16 dataset loads plus k hot loads), size = the largest table resident on every card | the replaced form lets the chip skip item derivations (x1.33 at k = 4); Mac rows hot32k4 x1.22, hot64k4 x1.12 | +| 5 | OUT of v3 (a measured option for 3.0) | the added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% (g 0.87 / 0.85 / 0.84 and 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB against the 0.97 rule): no card keeps even 32 MiB resident while the dataset streams; the replaced form helps the on-die-cache chip (x1.33 at k = 4) | | 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 | | 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple | | M16 mixer x4 | into v3 | the only measured lever that moves the named chip: 0.6x bare, 1.8x with a 3x fixed-function factor; verifier 1.6 to 4.8 ms per warp | From 29e1b28db4250261980feedae3373135aabb4485 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:38:34 +0000 Subject: [PATCH 079/131] Counter ASIC 2.0 status 21:38: ca2-cache final 2de19e5 --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 5cd2bc6d5..91ba964ba 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -384,3 +384,5 @@ Hot-table jobs on PC 1 (fetch 21:29:01Z; run-ca2-hot-5090 126 s, closed 21:31:38 | hot96k4a | 114.4 (0.84) | 14.56 (0.80) | 0.83 | Decision (the 0.97 rule): layer 5 is OUT of v3. Neither card keeps even the 32 MiB table resident while the 1 GiB dataset streams (the replaced form gains 1.02 to 1.08x at k = 4 against an ideal 1.33x), and the added form costs 13 to 20%; the chip row moves the wrong way with it. Layer 5 stays a measured option for 3.0 (a table small enough to stay resident, or a different access pattern). The probe rows and the writeup follow on ca2-cache. PC 1 is released to the shipper for PC 1's update-now; the era job starts after PC 1's 0.3.10 STATUS line. + +21:38. ca2-cache final: 2de19e5 (nine commits from 55e285c, on ca2-v3 464d6e1); hot-table.md carries the probe rows for all three cards (5090 112.6 G loads/s at 32 / 64 / 96 MiB inside its L2 against 17.6 at 1 GiB; 9070 XT 9.88 / 9.47 / 8.18 / 2.43; M5 Max 21.7 / 12.8 / 12.3 / 3.50), the PC tables with g, the chip arithmetic at the measured g, the decision, the 3.0 note ("what would make it pay": a resident size found by a hash sweep below 32 MiB, k only with residency, a line-unit or streamed access shape) and the unverified list; the bench-log entry and two addenda carry the job ids and the worker sha256s. The probe promises full hits inside the 5090's L2 but the hash gets 2 to 8% at k = 4 because the streaming dataset evicts the table. From a61bbc45246241ea94e6066d5211383490743495 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:39:01 +0000 Subject: [PATCH 080/131] Counter ASIC 2.0: gate G4 run 1 PASS recorded; status 21:39 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 4f95f216c..73a2f2aee 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -67,7 +67,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G1 | bit-exact v3 on all three vendors against the Mac reference | vectors PASS on Metal, CUDA (5090), AMD OpenCL for the v3 packs; batch fingerprints equal. RULING (coordinator, 5 October 2026, about 21:10 UTC): the integrated gfx1036 (RDNA 2, AMD OpenCL 3683.0) satisfies the AMD vendor tonight, because G1 is a compiler-and-ISA property and gfx1036 carried the v1 and v2 conformance; the 9070 XT's hash-rate and power rows are owed and taken when its link is back | open | | G2 | the CPU verifier exact on 1,000 random hashes per card | 1,000 GPU hashes per card re-hashed by `igneum-pow` on the Mac, 0 mismatches | open | | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | -| G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary | open | +| G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 on the composed class (era merged) follows | run 1 green; run 2 pending the era commit | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | | G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green; PC 2 pending | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 91ba964ba..0a612b5e6 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -386,3 +386,7 @@ Hot-table jobs on PC 1 (fetch 21:29:01Z; run-ca2-hot-5090 126 s, closed 21:31:38 Decision (the 0.97 rule): layer 5 is OUT of v3. Neither card keeps even the 32 MiB table resident while the 1 GiB dataset streams (the replaced form gains 1.02 to 1.08x at k = 4 against an ideal 1.33x), and the added form costs 13 to 20%; the chip row moves the wrong way with it. Layer 5 stays a measured option for 3.0 (a table small enough to stay resident, or a different access pattern). The probe rows and the writeup follow on ca2-cache. PC 1 is released to the shipper for PC 1's update-now; the era job starts after PC 1's 0.3.10 STATUS line. 21:38. ca2-cache final: 2de19e5 (nine commits from 55e285c, on ca2-v3 464d6e1); hot-table.md carries the probe rows for all three cards (5090 112.6 G loads/s at 32 / 64 / 96 MiB inside its L2 against 17.6 at 1 GiB; 9070 XT 9.88 / 9.47 / 8.18 / 2.43; M5 Max 21.7 / 12.8 / 12.3 / 3.50), the PC tables with g, the chip arithmetic at the measured g, the decision, the 3.0 note ("what would make it pay": a resident size found by a hash sweep below 32 MiB, k only with residency, a line-unit or streamed access shape) and the unverified list; the bench-log entry and two addenda carry the job ids and the worker sha256s. The probe promises full hits inside the 5090's L2 but the hash gets 2 to 8% at k = 4 because the streaming dataset evicts the table. + +## 21:39 gate G4 run 1 PASS on the mixer-x4 class + +Fast-time 3-node network, 21:33 to 21:38 UTC (fork 79bd8e10 + igneum-pow 66eeba3): the switch line on 3 of 3 nodes (active from epoch 3, DAA 150 rounded up to 180), templates class 2 then 3 from epoch 3, 181 blocks before and 124 after the boundary, program ids agree on all three miners (v2 e0 to e2, v3 e3 to e5), 0 rejected on miners and nodes, one sink on all three (082fd39ba65df2ff, 304/304/304), a new (day, class) cache 177 to 235 ms on one core. Main-repo tip bf04c56 (the doc, the summary JSON, the script's --connect fix). Run 2 on the composed class follows the era commit and the cache merge (rebuild about 10 min, gate 5 min). The x4 dataset build time comes from the mixer's PC prepare job (the CPU miner derives words from the cache). From 208487c30042ecfb8e570b1b2003fa63f8163040 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:40:50 +0000 Subject: [PATCH 081/131] Counter ASIC 2.0 status 21:40: x8 built and bit-exact, the mixer PC job --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 0a612b5e6..315de4a7f 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -390,3 +390,7 @@ Decision (the 0.97 rule): layer 5 is OUT of v3. Neither card keeps even the 32 M ## 21:39 gate G4 run 1 PASS on the mixer-x4 class Fast-time 3-node network, 21:33 to 21:38 UTC (fork 79bd8e10 + igneum-pow 66eeba3): the switch line on 3 of 3 nodes (active from epoch 3, DAA 150 rounded up to 180), templates class 2 then 3 from epoch 3, 181 blocks before and 124 after the boundary, program ids agree on all three miners (v2 e0 to e2, v3 e3 to e5), 0 rejected on miners and nodes, one sink on all three (082fd39ba65df2ff, 304/304/304), a new (day, class) cache 177 to 235 ms on one core. Main-repo tip bf04c56 (the doc, the summary JSON, the script's --connect fix). Run 2 on the composed class follows the era commit and the cache merge (rebuild about 10 min, gate 5 min). The x4 dataset build time comes from the mixer's PC prepare job (the CPU miner derives words from the cache). + +## 21:40 x8 built and bit-exact beside x4; the mixer PC job retargeted to PC 1 + +ca2-mixer 504cae4 (LoadClass::MX8 "mx8", packs mx8-genesis and mx8-devnet-epoch0, the fuzz takes IGNEUM_MIXER_CLASS, playbooks with the x8 packs) and fe4e193 (mixer-x4.md per-tier build table and the x4/x8 rule; chip-model-v3.md with the mixer row as the headline, the layer 5 rows kept as measured not adopted with the PC g beside the Mac's; x8 rows 0.31x bare, 0.92x with the factor, 0.76x at equal silicon at year 0). x8 bit-exactness (run lock): both packs on Metal and Apple OpenCL 3/3 + 3/3, 96 of 96 lanes, one fingerprint per pack across both harnesses (7c28cfb06c5c65a9, bbb183f72692f840); 50-program x8 fuzz on Metal 50 of 50, every tenth on OpenCL 5 of 5. Indicative Mac builds (run lock): 30.0 ms at x8 against 30.2 at x4 (genesis pack), 22.0 against 21.7 (devnet pack): the Mac's build is latency-bound. The timing session (verifier v2 / x4 / x8, the fill, the build) is queued behind the measure lock. PC job: mixer-x4-pcjob.zip sha256 55a2913cb8790cd3b106dc3d0d29b6e2952b5a08378cd815935915ab82898a8e (v2 control plus the mx4 and mx8 genesis and devnet packs; a prepare per pack printing the worker's cache and dataset build ms); retargeted so both halves run on PC 1 (its 5090 and its AMD card), after the era job, about 22:05. From f394016bc38c5c3e5f84f6a3949098d1ff6e7607 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:41:15 +0000 Subject: [PATCH 082/131] Counter ASIC 2.0 status 21:41: the mixer timing session, x8 provisional --- docs/plans/counter-asic-2-status.md | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 315de4a7f..92cc11d38 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -394,3 +394,15 @@ Fast-time 3-node network, 21:33 to 21:38 UTC (fork 79bd8e10 + igneum-pow 66eeba3 ## 21:40 x8 built and bit-exact beside x4; the mixer PC job retargeted to PC 1 ca2-mixer 504cae4 (LoadClass::MX8 "mx8", packs mx8-genesis and mx8-devnet-epoch0, the fuzz takes IGNEUM_MIXER_CLASS, playbooks with the x8 packs) and fe4e193 (mixer-x4.md per-tier build table and the x4/x8 rule; chip-model-v3.md with the mixer row as the headline, the layer 5 rows kept as measured not adopted with the PC g beside the Mac's; x8 rows 0.31x bare, 0.92x with the factor, 0.76x at equal silicon at year 0). x8 bit-exactness (run lock): both packs on Metal and Apple OpenCL 3/3 + 3/3, 96 of 96 lanes, one fingerprint per pack across both harnesses (7c28cfb06c5c65a9, bbb183f72692f840); 50-program x8 fuzz on Metal 50 of 50, every tenth on OpenCL 5 of 5. Indicative Mac builds (run lock): 30.0 ms at x8 against 30.2 at x4 (genesis pack), 22.0 against 21.7 (devnet pack): the Mac's build is latency-bound. The timing session (verifier v2 / x4 / x8, the fill, the build) is queued behind the measure lock. PC job: mixer-x4-pcjob.zip sha256 55a2913cb8790cd3b106dc3d0d29b6e2952b5a08378cd815935915ab82898a8e (v2 control plus the mx4 and mx8 genesis and devnet packs; a prepare per pack printing the worker's cache and dataset build ms); retargeted so both halves run on PC 1 (its 5090 and its AMD card), after the era job, about 22:05. + +## 21:41 the mixer timing session: x8 passes the verifier half of the rule + +M5 Max, one core, measure lock, 21:40:12 to 21:40:23 UTC, on a loaded box (load average 5.6 one-minute, 26 fifteen-minute: other agents' unlocked processes), so the absolute figures are about 2x the quiet 0.604 ms v2 baseline and the RATIOS are the measurement (two rounds, within 4%); a quiet-box re-run is owed for absolute numbers. + +| Class | Verifier ms per warp, avg of 50 (round 1 / 2) | Worst cold unit | Ratio to v2 | +|---|---|---|---| +| v2 | 1.361 / 1.310 | 1.579 | 1 | +| x4 (genesis; devnet pack 1.923) | 1.956 / 1.923 | 2.043 | 1.45x | +| x8 (genesis; devnet pack 2.972) | 2.785 / 2.790 | 2.942 | 2.1x | + +256 MiB cache fill on one core 172 to 175 ms. Metal 1 GiB build, GPU time: v2 21.0 ms (29.7 cold), x4 20.9 / 21.0, x8 21.9 / 21.9: the Mac's build is bound by the 8 dependent cache-line reads per item, not the arithmetic, so the "under 1 s on every discrete card" half of the rule is decided by the 5090 and 9070 XT rows of the mixer PC job (by the M16 arithmetic the 5090 is 54 ms at x4 and 107 ms at x8 if arithmetic-bound, 13.4 ms if latency-bound: far under 1 s either way). Verifier half: x8 passes with 7.1 ms of the 10 ms gate to spare on the loaded core (about 1.3 ms on a quiet core, approximate); x4 leaves 8.0 ms. Chip row at x8: 1,198,080 ops per hash, 41.7 MH/s, 0.31x bare, 0.92x with the 3x factor, 0.76x at equal silicon; x4 1.84x / 1.53x. C19 at these loaded figures: shares per core per second 735 / 511 / 358 (v2 / x4 / x8), a 22,000-member pool at one share per 10 s needs 3.0 / 4.3 / 6.1 cores; IBD over 108,000 headers on one core 2.4 / 3.5 / 5.0 min. Provisional choice under the rule: x8, confirmed when the PC build rows land (about 22:10). From 5f125f5aeeee635add87abdedc233d9c72b4c3d8 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:41:42 +0000 Subject: [PATCH 083/131] Counter ASIC 2.0 status 21:41: the era draw's Mac spread 0.8%, the final package, the merges --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 92cc11d38..20c1919c7 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -406,3 +406,7 @@ M5 Max, one core, measure lock, 21:40:12 to 21:40:23 UTC, on a loaded box (load | x8 (genesis; devnet pack 2.972) | 2.785 / 2.790 | 2.942 | 2.1x | 256 MiB cache fill on one core 172 to 175 ms. Metal 1 GiB build, GPU time: v2 21.0 ms (29.7 cold), x4 20.9 / 21.0, x8 21.9 / 21.9: the Mac's build is bound by the 8 dependent cache-line reads per item, not the arithmetic, so the "under 1 s on every discrete card" half of the rule is decided by the 5090 and 9070 XT rows of the mixer PC job (by the M16 arithmetic the 5090 is 54 ms at x4 and 107 ms at x8 if arithmetic-bound, 13.4 ms if latency-bound: far under 1 s either way). Verifier half: x8 passes with 7.1 ms of the 10 ms gate to spare on the loaded core (about 1.3 ms on a quiet core, approximate); x4 leaves 8.0 ms. Chip row at x8: 1,198,080 ops per hash, 41.7 MH/s, 0.31x bare, 0.92x with the 3x factor, 0.76x at equal silicon; x4 1.84x / 1.53x. C19 at these loaded figures: shares per core per second 735 / 511 / 358 (v2 / x4 / x8), a 22,000-member pool at one share per 10 s needs 3.0 / 4.3 / 6.1 cores; IBD over 108,000 headers on one core 2.4 / 3.5 / 5.0 min. Provisional choice under the rule: x8, confirmed when the PC build rows land (about 22:10). + +## 21:41 the era draw passes the 5% rule on the Mac; the final era package; the merges for G4 run 2 + +ca2-era 9f98af2 (one commit on ca2-v3 HEAD; 48 lib + 17 integration tests; the pinned v2 packs byte-identical; mx4-genesis unchanged; mx4-devnet-epoch0 re-exported with the era inside the class). V3_CLASS = LoadClass { era: None, ..LoadClass::MX4 }; the era class is drawn inside generate_from_seed_bytes_program_class from the era bytes; chain_dataset_day and Epoch::dataset_word compose unchanged; an era program takes 11 draws per instruction. Mac (M5 Max, 21:38Z, Metal packbench 5 x 2^24, a loaded box): v2 27.68 MH/s; era-0 to era-5 28.58, 28.48, 28.35, 28.38, 28.49, 28.48: min 28.35, median 28.48, max 28.58, SPREAD 0.8% (under the 5% rule); 3/3 vectors and the in-batch vectors PASS on every pack; CPU verify 1.319 to 1.345 ms per warp against v2 1.334 in the same loaded run (quiet re-run owed). FINAL PC 1 package: ~/Desktop/igneum-ca2-era-pc1.zip sha256 f79c0607bb4187e3cf16fce3f533e7d525673d766d7edb799f27fd81af5dcee1 (workers 0fbfd50a... and 8c8caff7... from the merged tree; packs re-exported, attempt 0, program id 73bcbfe8ccf988f1 in all six with the era seed beside it); fetch-ca2-era-20261005; ca2-era-pc1.ps1. The 0.3.10 update-now reached every machine at 21:39:59Z (manifest live 21:33Z); the era job starts on PC 1's 0.3.10 STATUS line. The node agent merges 9f98af2 then ca2-cache 2de19e5 into ca2-v3 for G4 run 2 and the suites' tip. From 978c430a8aca4160f892a9b1df1fcfa273177462 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:42:41 +0000 Subject: [PATCH 084/131] Counter ASIC 2.0 status 21:42: the era fast-forward 1558971, the cache rebase, the suites' tip --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 20c1919c7..a2f247af1 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -410,3 +410,5 @@ M5 Max, one core, measure lock, 21:40:12 to 21:40:23 UTC, on a loaded box (load ## 21:41 the era draw passes the 5% rule on the Mac; the final era package; the merges for G4 run 2 ca2-era 9f98af2 (one commit on ca2-v3 HEAD; 48 lib + 17 integration tests; the pinned v2 packs byte-identical; mx4-genesis unchanged; mx4-devnet-epoch0 re-exported with the era inside the class). V3_CLASS = LoadClass { era: None, ..LoadClass::MX4 }; the era class is drawn inside generate_from_seed_bytes_program_class from the era bytes; chain_dataset_day and Epoch::dataset_word compose unchanged; an era program takes 11 draws per instruction. Mac (M5 Max, 21:38Z, Metal packbench 5 x 2^24, a loaded box): v2 27.68 MH/s; era-0 to era-5 28.58, 28.48, 28.35, 28.38, 28.49, 28.48: min 28.35, median 28.48, max 28.58, SPREAD 0.8% (under the 5% rule); 3/3 vectors and the in-batch vectors PASS on every pack; CPU verify 1.319 to 1.345 ms per warp against v2 1.334 in the same loaded run (quiet re-run owed). FINAL PC 1 package: ~/Desktop/igneum-ca2-era-pc1.zip sha256 f79c0607bb4187e3cf16fce3f533e7d525673d766d7edb799f27fd81af5dcee1 (workers 0fbfd50a... and 8c8caff7... from the merged tree; packs re-exported, attempt 0, program id 73bcbfe8ccf988f1 in all six with the era seed beside it); fetch-ca2-era-20261005; ca2-era-pc1.ps1. The 0.3.10 update-now reached every machine at 21:39:59Z (manifest live 21:33Z); the era job starts on PC 1's 0.3.10 STATUS line. The node agent merges 9f98af2 then ca2-cache 2de19e5 into ca2-v3 for G4 run 2 and the suites' tip. + +21:42. ca2-v3 now carries the era draw: ca2-era's tip b105a55 (9f98af2 rebased onto the node agent's 88dafbc) fast-forwarded, no conflict; the suite job's main-repo tip is b105a55. ca2-cache 2de19e5 does NOT merge (it bases on 464d6e1, before the mixer and era commits rewrote the class literal, the load emitters, the pack fields and the pinned-pack tests: 8 files, 35 hunks); the node agent aborted cleanly and the cache agent is rebasing onto b105a55 as a squashed commit with hot: None kept; the composed class under test is unchanged by the cache code, so the igneum-pow and packfile checks, the igneumd and igneum-miner rebuild and gate run 2 proceed on b105a55 now. From 17da4045b59aecb7edcfcf63d0213aeb24631daf Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:45:43 +0000 Subject: [PATCH 085/131] Counter ASIC 2.0 status 21:45: the PC 2 resume no-op defect, the miners off since 21:25 --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index a2f247af1..c0d7b84bb 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -412,3 +412,5 @@ M5 Max, one core, measure lock, 21:40:12 to 21:40:23 UTC, on a loaded box (load ca2-era 9f98af2 (one commit on ca2-v3 HEAD; 48 lib + 17 integration tests; the pinned v2 packs byte-identical; mx4-genesis unchanged; mx4-devnet-epoch0 re-exported with the era inside the class). V3_CLASS = LoadClass { era: None, ..LoadClass::MX4 }; the era class is drawn inside generate_from_seed_bytes_program_class from the era bytes; chain_dataset_day and Epoch::dataset_word compose unchanged; an era program takes 11 draws per instruction. Mac (M5 Max, 21:38Z, Metal packbench 5 x 2^24, a loaded box): v2 27.68 MH/s; era-0 to era-5 28.58, 28.48, 28.35, 28.38, 28.49, 28.48: min 28.35, median 28.48, max 28.58, SPREAD 0.8% (under the 5% rule); 3/3 vectors and the in-batch vectors PASS on every pack; CPU verify 1.319 to 1.345 ms per warp against v2 1.334 in the same loaded run (quiet re-run owed). FINAL PC 1 package: ~/Desktop/igneum-ca2-era-pc1.zip sha256 f79c0607bb4187e3cf16fce3f533e7d525673d766d7edb799f27fd81af5dcee1 (workers 0fbfd50a... and 8c8caff7... from the merged tree; packs re-exported, attempt 0, program id 73bcbfe8ccf988f1 in all six with the era seed beside it); fetch-ca2-era-20261005; ca2-era-pc1.ps1. The 0.3.10 update-now reached every machine at 21:39:59Z (manifest live 21:33Z); the era job starts on PC 1's 0.3.10 STATUS line. The node agent merges 9f98af2 then ca2-cache 2de19e5 into ca2-v3 for G4 run 2 and the suites' tip. 21:42. ca2-v3 now carries the era draw: ca2-era's tip b105a55 (9f98af2 rebased onto the node agent's 88dafbc) fast-forwarded, no conflict; the suite job's main-repo tip is b105a55. ca2-cache 2de19e5 does NOT merge (it bases on 464d6e1, before the mixer and era commits rewrote the class literal, the load emitters, the pack fields and the pinned-pack tests: 8 files, 35 hunks); the node agent aborted cleanly and the cache agent is rebasing onto b105a55 as a squashed commit with hot: None kept; the composed class under test is unchanged by the cache code, so the igneum-pow and packfile checks, the igneumd and igneum-miner rebuild and gate run 2 proceed on b105a55 now. + +21:45. PC 2 defect: since a job's /api/resume at 21:25:11Z the 0.3.9 app answered ok and never restarted the NVIDIA miner (nor the iGPU one): the 5090 worker "off" at hash 0 holding 1.7 GB, so agg-cost-pc2-2's mining phases are void (its idle phases run; closes about 21:52Z) and the devnet has been short PC 2's rate since 21:25. The 0.3.10 restart should bring the miners back; the shipper confirms PC 2's 5090 STATUS rate after the 0.3.10 line, else the aggregation-cost agent's restore script (tools/proving-v1/pc2-agg-cost-restore.ps1, 30 s) runs. Defect for the next cut: a resume that answers ok without a miner restart; the app must re-check the miner processes after a resume and report a failure. The aggregation-cost agent gets a 20-minute re-run slot on PC 2 after the 0.3.11 suites. From 705e1294461bff633c11f15cce2c9586bde35f87 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:46:13 +0000 Subject: [PATCH 086/131] Counter ASIC 2.0 status 21:46: PC 1 on 0.3.10, the era go --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index c0d7b84bb..059e68fa0 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -414,3 +414,7 @@ ca2-era 9f98af2 (one commit on ca2-v3 HEAD; 48 lib + 17 integration tests; the p 21:42. ca2-v3 now carries the era draw: ca2-era's tip b105a55 (9f98af2 rebased onto the node agent's 88dafbc) fast-forwarded, no conflict; the suite job's main-repo tip is b105a55. ca2-cache 2de19e5 does NOT merge (it bases on 464d6e1, before the mixer and era commits rewrote the class literal, the load emitters, the pack fields and the pinned-pack tests: 8 files, 35 hunks); the node agent aborted cleanly and the cache agent is rebasing onto b105a55 as a squashed commit with hot: None kept; the composed class under test is unchanged by the cache code, so the igneum-pow and packfile checks, the igneumd and igneum-miner rebuild and gate run 2 proceed on b105a55 now. 21:45. PC 2 defect: since a job's /api/resume at 21:25:11Z the 0.3.9 app answered ok and never restarted the NVIDIA miner (nor the iGPU one): the 5090 worker "off" at hash 0 holding 1.7 GB, so agg-cost-pc2-2's mining phases are void (its idle phases run; closes about 21:52Z) and the devnet has been short PC 2's rate since 21:25. The 0.3.10 restart should bring the miners back; the shipper confirms PC 2's 5090 STATUS rate after the 0.3.10 line, else the aggregation-cost agent's restore script (tools/proving-v1/pc2-agg-cost-restore.ps1, 30 s) runs. Defect for the next cut: a resume that answers ok without a miner restart; the app must re-check the miner processes after a resume and report a failure. The aggregation-cost agent gets a 20-minute re-run slot on PC 2 after the 0.3.11 suites. + +## 21:46 PC 1 is on 0.3.10; the era job has the go + +PC 1 restarted on 0.3.10 at 21:40:41Z (engine run win-ae432dc7-20261005-214041), the 5090 at 141.4 MH/s by 21:45:17Z; "go PC 1" to the era job at 21:46 (fetch-ca2-era-20261005, zip f79c0607...; the 5090 then the AMD card; 6 to 10 min). The mixer daily-build job follows it, then the 5090 power sweep, Ember's build and fetches, the repro re-run, the AMD sweep (conditional), Ember's run. PC 2 is still on 0.3.9 at 21:45 (its restart pending); the 0.3.11 suites publish after its 0.3.10 line and a confirmed 5090 rate. Every measurement before the restart on PC 1 (hot table 21:29 to 21:35) stands: it did not straddle. From 110a9b1af1fac5960358dc150484e81e25cd4dbe Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:50:54 +0000 Subject: [PATCH 087/131] Counter ASIC 2.0: the resume fix assigned to 0.3.11's app half; status 21:50 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 73a2f2aee..9e351b2dd 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -101,7 +101,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's ## 8a. Proving v1 rides with it -the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL 6dc686a on 5b0d54f (the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. +the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL 6dc686a on 5b0d54f (the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 059e68fa0..eb0060185 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -418,3 +418,5 @@ ca2-era 9f98af2 (one commit on ca2-v3 HEAD; 48 lib + 17 integration tests; the p ## 21:46 PC 1 is on 0.3.10; the era job has the go PC 1 restarted on 0.3.10 at 21:40:41Z (engine run win-ae432dc7-20261005-214041), the 5090 at 141.4 MH/s by 21:45:17Z; "go PC 1" to the era job at 21:46 (fetch-ca2-era-20261005, zip f79c0607...; the 5090 then the AMD card; 6 to 10 min). The mixer daily-build job follows it, then the 5090 power sweep, Ember's build and fetches, the repro re-run, the AMD sweep (conditional), Ember's run. PC 2 is still on 0.3.9 at 21:45 (its restart pending); the 0.3.11 suites publish after its 0.3.10 line and a confirmed 5090 rate. Every measurement before the restart on PC 1 (hot table 21:29 to 21:35) stands: it did not straddle. + +21:50. The resume fix (the miners not restarted after POST /api/resume: PC 2 since 21:25Z tonight, the Mac this afternoon) is assigned to the proving agent on 0.3.11's app branch (engine.rs resume: restart every enabled card's worker and a stale pack export, re-check within one tick, a state-machine unit test plus the known-failed case from PC 2's log); rollout plan 8a. If its commit is not in hand at the app cut, it heads 0.3.12's list. From 4127c80108a0227cce82a50ab9f994753fda6a32 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:52:11 +0000 Subject: [PATCH 088/131] Counter ASIC 2.0: gate G4 green (runs 1 and 2); status 21:52: the suites packing --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 9e351b2dd..0525bf98e 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -67,7 +67,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G1 | bit-exact v3 on all three vendors against the Mac reference | vectors PASS on Metal, CUDA (5090), AMD OpenCL for the v3 packs; batch fingerprints equal. RULING (coordinator, 5 October 2026, about 21:10 UTC): the integrated gfx1036 (RDNA 2, AMD OpenCL 3683.0) satisfies the AMD vendor tonight, because G1 is a compiler-and-ISA property and gfx1036 carried the v1 and v2 conformance; the 9070 XT's hash-rate and power rows are owed and taken when its link is back | open | | G2 | the CPU verifier exact on 1,000 random hashes per card | 1,000 GPU hashes per card re-hashed by `igneum-pow` on the Mac, 0 mismatches | open | | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | -| G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 on the composed class (era merged) follows | run 1 green; run 2 pending the era commit | +| G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms | GREEN (runs 1 and 2) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | | G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green; PC 2 pending | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index eb0060185..bd031d25b 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -420,3 +420,7 @@ ca2-era 9f98af2 (one commit on ca2-v3 HEAD; 48 lib + 17 integration tests; the p PC 1 restarted on 0.3.10 at 21:40:41Z (engine run win-ae432dc7-20261005-214041), the 5090 at 141.4 MH/s by 21:45:17Z; "go PC 1" to the era job at 21:46 (fetch-ca2-era-20261005, zip f79c0607...; the 5090 then the AMD card; 6 to 10 min). The mixer daily-build job follows it, then the 5090 power sweep, Ember's build and fetches, the repro re-run, the AMD sweep (conditional), Ember's run. PC 2 is still on 0.3.9 at 21:45 (its restart pending); the 0.3.11 suites publish after its 0.3.10 line and a confirmed 5090 rate. Every measurement before the restart on PC 1 (hot table 21:29 to 21:35) stands: it did not straddle. 21:50. The resume fix (the miners not restarted after POST /api/resume: PC 2 since 21:25Z tonight, the Mac this afternoon) is assigned to the proving agent on 0.3.11's app branch (engine.rs resume: restart every enabled card's worker and a stale pack export, re-check within one tick, a state-machine unit test plus the known-failed case from PC 2's log); rollout plan 8a. If its commit is not in hand at the app cut, it heads 0.3.12's list. + +## 21:52 gate G4 run 2 PASS on the composed class; the 0.3.11 suites are packing for PC 2 + +G4 run 2 (21:46:36 to 21:51:31 UTC, b105a55 = era + mixer, hot None): every check true; 181 / 124 blocks around DAA 180; v3 ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners; the v2 epoch-0 id 8f8806638d59850f unchanged from run 1; 0 rejected; one sink 712c1b212091dcdc at 303/303/303; the switch line on 3 of 3; cache ready 178 / 181 ms. G4 is GREEN. The 0.3.11 suite job is packing from the ca2-v3 worktree (fork 79bd8e10, main b105a55; build-inputs.zip 9,563,672 bytes sha256 bf89ab4c...) for PC 2, which is on 0.3.10 since 21:49:41Z with its 5090 worker back on the first try. From fc5721a51670785be463307c9a59449664a55307 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:53:24 +0000 Subject: [PATCH 089/131] Counter ASIC 2.0 status 21:53: the suite job published --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 0525bf98e..b08e1e62e 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,7 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms | GREEN (runs 1 and 2) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green; PC 2 pending | +| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green; PC 2 job build-20261005-215219 published 21:52:19Z (75-minute budget), result pending | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index bd031d25b..ac6354587 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -424,3 +424,5 @@ PC 1 restarted on 0.3.10 at 21:40:41Z (engine run win-ae432dc7-20261005-214041), ## 21:52 gate G4 run 2 PASS on the composed class; the 0.3.11 suites are packing for PC 2 G4 run 2 (21:46:36 to 21:51:31 UTC, b105a55 = era + mixer, hot None): every check true; 181 / 124 blocks around DAA 180; v3 ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners; the v2 epoch-0 id 8f8806638d59850f unchanged from run 1; 0 rejected; one sink 712c1b212091dcdc at 303/303/303; the switch line on 3 of 3; cache ready 178 / 181 ms. G4 is GREEN. The 0.3.11 suite job is packing from the ca2-v3 worktree (fork 79bd8e10, main b105a55; build-inputs.zip 9,563,672 bytes sha256 bf89ab4c...) for PC 2, which is on 0.3.10 since 21:49:41Z with its 5090 worker back on the first try. + +21:53. The 0.3.11 suite job is published: build-20261005-215219 to PC 2 (fork 79bd8e10, main b105a55; linux build 30 min, tests 35 min: kaspa-consensus-core, igneum-exec, kaspa-pow, kaspa-consensus, igneum-miner, kaspa-p2p-flows, igneum-app); PC 2 is on 0.3.10 with its 5090 back. The worktree freeze is lifted for the node agent (the run-2 summary commit, then the cache merge). The proving agent's resume fix waits on a test run behind the Mac's held lock slots. From 1610a0744b5639dfc4f9b43d940037eb882f255f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:55:39 +0000 Subject: [PATCH 090/131] Counter ASIC 2.0 status 21:55: both PCs on 0.3.10, the PC 2 prover socket item, ca2-v3 495c552 --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index ac6354587..a30df065b 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -426,3 +426,5 @@ PC 1 restarted on 0.3.10 at 21:40:41Z (engine run win-ae432dc7-20261005-214041), G4 run 2 (21:46:36 to 21:51:31 UTC, b105a55 = era + mixer, hot None): every check true; 181 / 124 blocks around DAA 180; v3 ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners; the v2 epoch-0 id 8f8806638d59850f unchanged from run 1; 0 rejected; one sink 712c1b212091dcdc at 303/303/303; the switch line on 3 of 3; cache ready 178 / 181 ms. G4 is GREEN. The 0.3.11 suite job is packing from the ca2-v3 worktree (fork 79bd8e10, main b105a55; build-inputs.zip 9,563,672 bytes sha256 bf89ab4c...) for PC 2, which is on 0.3.10 since 21:49:41Z with its 5090 worker back on the first try. 21:53. The 0.3.11 suite job is published: build-20261005-215219 to PC 2 (fork 79bd8e10, main b105a55; linux build 30 min, tests 35 min: kaspa-consensus-core, igneum-exec, kaspa-pow, kaspa-consensus, igneum-miner, kaspa-p2p-flows, igneum-app); PC 2 is on 0.3.10 with its 5090 back. The worktree freeze is lifted for the node agent (the run-2 summary commit, then the cache merge). The proving agent's resume fix waits on a test run behind the Mac's held lock slots. + +21:55. The 0.3.10 rollout: both PCs mine on 0.3.10 with the rebuilt workers (PC 2 120.6 MH/s at 21:54:33Z, PC 1 141.3); the hand nodes (21:49:38Z, 21:49:50Z) and the seed (21:50:15Z) on 21d4c73c, digest 1f4b4425 everywhere; not yet on 0.3.10: the US laptop 37ba0461 (installer downloaded 21:40:52Z, app not back after 13 min; nothing to drive from here) and Sam's Mac (quit since 20:47Z). Open on PC 2: the prover fails with "CudaClientError: Connect(PermissionDenied)" since the restart (three shards 21:49:56 to 21:50:32Z); the likely cause is the sp1-gpu-server socket handling of the aggregation-cost jobs; the proving agent owns it and publishes a fix after the suite job (PC 2 is the suite job's until it closes). The ca2-v3 tip is 63dabb2 (run-2 summary and doc, the G6 job id recorded). From 5fc5e100ad3a5568ffa4ecd9a241ae38621bcbb5 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:56:54 +0000 Subject: [PATCH 091/131] Counter ASIC 2.0: G6 first job failed on the known kaspa-consensus flake, split re-run; the app branch final a223ca9; status 21:56 --- docs/plans/counter-asic-2-rollout.md | 4 ++-- docs/plans/counter-asic-2-status.md | 6 ++++++ 2 files changed, 8 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index b08e1e62e..5a0a69d1d 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,7 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms | GREEN (runs 1 and 2) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green; PC 2 job build-20261005-215219 published 21:52:19Z (75-minute budget), result pending | +| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 runs kaspa-consensus alone, job 3 of 3 the other five crates; G6 is green only when both pass | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. @@ -101,7 +101,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's ## 8a. Proving v1 rides with it -the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL 6dc686a on 5b0d54f (the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. +the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL a223ca9 on 5b0d54f (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index a30df065b..73351e2f6 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -428,3 +428,9 @@ G4 run 2 (21:46:36 to 21:51:31 UTC, b105a55 = era + mixer, hot None): every chec 21:53. The 0.3.11 suite job is published: build-20261005-215219 to PC 2 (fork 79bd8e10, main b105a55; linux build 30 min, tests 35 min: kaspa-consensus-core, igneum-exec, kaspa-pow, kaspa-consensus, igneum-miner, kaspa-p2p-flows, igneum-app); PC 2 is on 0.3.10 with its 5090 back. The worktree freeze is lifted for the node agent (the run-2 summary commit, then the cache merge). The proving agent's resume fix waits on a test run behind the Mac's held lock slots. 21:55. The 0.3.10 rollout: both PCs mine on 0.3.10 with the rebuilt workers (PC 2 120.6 MH/s at 21:54:33Z, PC 1 141.3); the hand nodes (21:49:38Z, 21:49:50Z) and the seed (21:50:15Z) on 21d4c73c, digest 1f4b4425 everywhere; not yet on 0.3.10: the US laptop 37ba0461 (installer downloaded 21:40:52Z, app not back after 13 min; nothing to drive from here) and Sam's Mac (quit since 20:47Z). Open on PC 2: the prover fails with "CudaClientError: Connect(PermissionDenied)" since the restart (three shards 21:49:56 to 21:50:32Z); the likely cause is the sp1-gpu-server socket handling of the aggregation-cost jobs; the proving agent owns it and publishes a fix after the suite job (PC 2 is the suite job's until it closes). The ca2-v3 tip is 63dabb2 (run-2 summary and doc, the G6 job id recorded). + +## 21:56 gate G6: the first PC 2 job failed on the known kaspa-consensus flake; split re-run + +build-20261005-215219 (21:53:01 to 21:55:48Z): Linux build ok (igneumd 49,164,264 bytes sha256 11979b49..., igneum-miner d25a8270..., igneum-app 68007173...), igneum-app tests 78 + 26 + 8 passed; the node stage exit 101: kaspa-consensus 96 passed, 1 failed, processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(..., 487112384, 487129578) in mine_on_all. This is the flake the 0.3.10 cut met on the same base 21d4c73c under the six-package parallel run (release-0.3.10.md: it passed alone twice, job build-20261005-182804, 97 passed), a timing-dependent difficulty in the test helper, not a v3 change (v3 touches no finality code). Per the gate rule the publish stops here until the suites are green: job 2 of 3 (kaspa-consensus alone) is published now; job 3 of 3 (the other five crates) follows; the flake itself goes on the next-cut list (make mine_on_all deterministic under parallel load). + +The proving v1 app branch is final at a223ca9 (6dc686a plus the resume fix with the PC 2 case as a unit test); rollout plan 8a updated. PC 2's prover fault is the root-socket class (agg-cost-pc2-1 ran the host as root in WSL2 and left /tmp/sp1-cuda-0.sock owned by root; the app's prover has failed every shard since 21:25:24Z); the proving agent's pc2-socket-fix.ps1 (60 s, miners untouched) runs between my two suite jobs. From 3a6beaa0b8aba8dfd6c30d8707ecea8fd785777f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:57:48 +0000 Subject: [PATCH 092/131] Counter ASIC 2.0 status 21:57: the era PC job, the verifier regression gate item, the PC 2 order --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 8 ++++++++ 2 files changed, 9 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 5a0a69d1d..97e698f04 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -101,7 +101,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's ## 8a. Proving v1 rides with it -the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL a223ca9 on 5b0d54f (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. +the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL a223ca9 on 5b0d54f, docs-only 8b47073 after it (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 73351e2f6..8fec3c089 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -434,3 +434,11 @@ G4 run 2 (21:46:36 to 21:51:31 UTC, b105a55 = era + mixer, hot None): every chec build-20261005-215219 (21:53:01 to 21:55:48Z): Linux build ok (igneumd 49,164,264 bytes sha256 11979b49..., igneum-miner d25a8270..., igneum-app 68007173...), igneum-app tests 78 + 26 + 8 passed; the node stage exit 101: kaspa-consensus 96 passed, 1 failed, processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(..., 487112384, 487129578) in mine_on_all. This is the flake the 0.3.10 cut met on the same base 21d4c73c under the six-package parallel run (release-0.3.10.md: it passed alone twice, job build-20261005-182804, 97 passed), a timing-dependent difficulty in the test helper, not a v3 change (v3 touches no finality code). Per the gate rule the publish stops here until the suites are green: job 2 of 3 (kaspa-consensus alone) is published now; job 3 of 3 (the other five crates) follows; the flake itself goes on the next-cut list (make mine_on_all deterministic under parallel load). The proving v1 app branch is final at a223ca9 (6dc686a plus the resume fix with the PC 2 case as a unit test); rollout plan 8a updated. PC 2's prover fault is the root-socket class (agg-cost-pc2-1 ran the host as root in WSL2 and left /tmp/sp1-cuda-0.sock owned by root; the app's prover has failed every shard since 21:25:24Z); the proving agent's pc2-socket-fix.ps1 (60 s, miners untouched) runs between my two suite jobs. + +## 21:57 the era PC job is done; a 2.2x CPU-verifier regression on ca2-v3 HEAD (gate item); the PC 2 order + +Era job run-ca2-era-pc1-20261005 (exit 0, 301 s), both cards restored, the 9070 XT ON the bus at run time (gfx1201, 32 CUs): first rows 18.96 to 19.21 MH/s on the era packs on the 9070 XT, self-test PASS, the 2^24 fingerprints equal to the Mac's (era-4 3ace11ad84c053ae, era-5 a8897d82adceb4a1); the full 5090 and 9070 XT tables with the spread per card follow. "go PC 1" to the mixer daily-build job at 21:57. + +GATE ITEM (found by the era agent, two binaries on the same v2 input, checksum 19297e99c7b9a55e, same minute, load 4 to 5): the CPU verifier on ca2-v3 HEAD (88dafbc) takes 1.332 ms per warp and on the era branch 1.310, against readwidth's binary at 0.604 and 0.606: the mixer branch's derive_items / Shape path costs 2.2x at m = 1, on the v2 path the live devnet verifies with. The mixer agent's "loaded box" reading of its 1.31 to 1.36 ms v2 figure was the code, not the load. Ordered: find and fix on ca2-mixer, restore v2 to within 5% of 0.604 ms measured the same way, re-measure v2 / x4 / x8 on the fixed binary (the C19 figures too), then the node agent merges it; the publish waits on it (G3's suite does not catch a slowdown; the 10 ms gate and the pool and IBD figures depend on it). Next-cut rule: a verifier benchmark with a pinned bound in the crate's CI. + +PC 2 order: suite job 2 of 3 (kaspa-consensus alone, running), the proving agent's socket fix (60 s; the aggregation-cost restore is skipped), suite job 3 of 3 (the five other crates), the aggregation-cost re-run (20 min), the prover-floor agent's windows (through the proving agent). The proving app branch: a223ca9 is the last code change (8b47073 docs only after it). From e83084a1c1489d98f7f2c695b2aa1170e6977ec7 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:58:03 +0000 Subject: [PATCH 093/131] Counter ASIC 2.0 status 21:58: ca2-cache rebased 39ecd7c --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 8fec3c089..a83af6e80 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -442,3 +442,5 @@ Era job run-ca2-era-pc1-20261005 (exit 0, 301 s), both cards restored, the 9070 GATE ITEM (found by the era agent, two binaries on the same v2 input, checksum 19297e99c7b9a55e, same minute, load 4 to 5): the CPU verifier on ca2-v3 HEAD (88dafbc) takes 1.332 ms per warp and on the era branch 1.310, against readwidth's binary at 0.604 and 0.606: the mixer branch's derive_items / Shape path costs 2.2x at m = 1, on the v2 path the live devnet verifies with. The mixer agent's "loaded box" reading of its 1.31 to 1.36 ms v2 figure was the code, not the load. Ordered: find and fix on ca2-mixer, restore v2 to within 5% of 0.604 ms measured the same way, re-measure v2 / x4 / x8 on the fixed binary (the C19 figures too), then the node agent merges it; the publish waits on it (G3's suite does not catch a slowdown; the 10 ms gate and the pool and IBD figures depend on it). Next-cut rule: a verifier benchmark with a pinned bound in the crate's CI. PC 2 order: suite job 2 of 3 (kaspa-consensus alone, running), the proving agent's socket fix (60 s; the aggregation-cost restore is skipped), suite job 3 of 3 (the five other crates), the aggregation-cost re-run (20 min), the prover-floor agent's windows (through the proving agent). The proving app branch: a223ca9 is the last code change (8b47073 docs only after it). + +21:58. ca2-cache is rebased: one squashed commit 1950661 on ca2-v3 63dabb2 (fast-forwardable; history under tag ca2-cache-history-2026-10-05); V3_CLASS = { era: None, hot: None, ..MX4 }; every hunk kept the ca2-v3 side and appended the hot code; 53 + 19 tests; the pinned v2, mx4, era and readwidth packs untouched; the eight hot packs re-exported with unchanged vectors, 96/96 and the pre-rebase fingerprints on Metal and Apple OpenCL. The node agent merges it after the mixer's verifier fix. From 4c4a947baac99a4274efce85e60e74886b1a25b6 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:58:33 +0000 Subject: [PATCH 094/131] Counter ASIC 2.0: layers 4 and 8 decided IN on the PC rows; status 21:58 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 10 ++++++++++ docs/plans/counter-asic-2.md | 2 +- 3 files changed, 12 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 97e698f04..46013392b 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -116,7 +116,7 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa | The width rule (layers 1 and 2) | 4 B fixed; 16 B; 64 B; the per-program mix | `` | `docs/plans/read-width.md` | | The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `` | the same | | The hot table (layer 5) | 32, 64, 96 MB; replaced or added | DECIDED (5 October 2026, 21:36 UTC, delegated): OUT of v3. The rule was g at or above 0.97 on both PC cards in the added form; measured g is 0.87 / 0.85 / 0.84 on the 5090 and 0.84 / 0.81 / 0.80 on the 9070 XT at 32 / 64 / 96 MiB: neither card keeps even the 32 MiB table resident while the dataset streams, and the replaced form helps the on-die-cache chip. Layer 5 stays a measured option for 3.0 | `docs/plans/hot-table.md` (ca2-cache): 5090 MH/s v2 136.1; replaced hot32k4 146.6, hot64k4 140.8, hot96k4 138.5, hot64k2 137.5, hot64k8 163.6; added hot32k4a 118.7, hot64k4a 115.4, hot96k4a 114.4. 9070 XT v2 18.15; replaced 19.79 / 18.73 / 18.33 / 18.17 / 22.32; added 15.27 / 14.62 / 14.56. M5 Max added 0.93 / 0.87 / 0.83. All eight packs bit-exact on the 5090 and the 9070 XT with the Mac's fingerprints (21:29 to 21:35 UTC, no restart straddled) | -| The era draws (layers 4 and 8) | in, if the min-to-max spread across six drawn eras is under 5% per card; item size fixed at 4 B, draws of stride, interleave and the working-set window at or above 256 MiB | pending the six-era table | `docs/plans/era-layout.md` | +| The era draws (layers 4 and 8) | in, if the min-to-max spread across six drawn eras is under 5% per card; item size fixed at 4 B, draws of stride, interleave and the working-set window at or above 256 MiB | DECIDED (5 October 2026, 21:58 UTC, delegated): IN. Spread over six eras: RTX 5090 1.3% (136.18 / 136.44 / 138.01 MH/s against v2 137.2), RX 9070 XT 3.2% (18.61 / 18.93 / 19.21 against 18.09), M5 Max 0.8% (28.35 / 28.48 / 28.58 against 27.68); all under 5%; latency-bound share 1.01 / 0.95 / 1.06; every pack's 2^24 fingerprint equal on all three vendors; the CPU verifier 1.00 to 1.02 of v2 within one binary | `docs/plans/era-layout.md` (ca2-era 78c0ee4; PC 1 jobs fetch-ca2-era-20261005 and run-ca2-era-pc1-20261005, 301 s, both cards restored). Chip line: 512 B read per hash, 120 to 128 distinct lines, the mirror is the whole dataset every hour; the interleave's value against a chip with a programmable address decoder is nil (stated), the stride is a bijection with no cryptanalysis yet | | The cache schedule (layer 6) | flat 256 MiB or a growth schedule | DECIDED (5 October 2026, delegated under "execute the full 2.0 plan, tonight 1-8"; the project lead confirms for the public testnet genesis; the mirror's mm^2 are being corrected to shipped-product density, 2 to 3x the bit-cell figures, conclusion unchanged): option C, the cache doubles when the dataset doubles: 256 MiB at genesis, 512 MiB at year 4, 1 GiB at year 12. Verifier fill on one M5 Max core at 0.2 s per 256 MiB (spec 1.12): 0.2 s, 0.4 s, 0.8 s at each step, under 1 s at every step of the schedule; verifier memory 256 MiB, 512 MiB, 1 GiB | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate | | Layer 7 | reserved family, unlock by era height or 90% signal | DECIDED (5 October 2026, delegated): reserve family R1 = mm8 (uint8 8x16 by 16x8 tile per unit, unsigned bytes), W_new 4, unlock at era 4 or 90% signal, the emulation rule in spec 1.13.2; switched off, no consensus effect tonight; the 5090 and 9070 XT dp4a numbers when the PCs free (PC 2 first) | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) | | The activation height N4 | the rule of section 3 | | | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index a83af6e80..88e89ea9c 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -444,3 +444,13 @@ GATE ITEM (found by the era agent, two binaries on the same v2 input, checksum 1 PC 2 order: suite job 2 of 3 (kaspa-consensus alone, running), the proving agent's socket fix (60 s; the aggregation-cost restore is skipped), suite job 3 of 3 (the five other crates), the aggregation-cost re-run (20 min), the prover-floor agent's windows (through the proving agent). The proving app branch: a223ca9 is the last code change (8b47073 docs only after it). 21:58. ca2-cache is rebased: one squashed commit 1950661 on ca2-v3 63dabb2 (fast-forwardable; history under tag ca2-cache-history-2026-10-05); V3_CLASS = { era: None, hot: None, ..MX4 }; every hunk kept the ca2-v3 side and appended the hot code; 53 + 19 tests; the pinned v2, mx4, era and readwidth packs untouched; the eight hot packs re-exported with unchanged vectors, 96/96 and the pre-rebase fingerprints on Metal and Apple OpenCL. The node agent merges it after the mixer's verifier fix. + +## 21:58 layers 4 and 8 decided IN on the PC rows (ca2-era 78c0ee4) + +| Card | v2 MH/s | six eras min / median / max | Spread | Fingerprints = Mac | Latency-bound share | +|---|---|---|---|---|---| +| RTX 5090 (PC 1, CUDA, 1 warp per block) | 137.2 | 136.18 / 136.44 / 138.01 | 1.3% | 7/7 | 1.01 | +| RX 9070 XT (PC 1, OpenCL, group 256) | 18.09 | 18.61 / 18.93 / 19.21 | 3.2% | 7/7 | 0.95 | +| M5 Max (Metal, 5 x 2^24) | 27.68 | 28.35 / 28.48 / 28.58 | 0.8% | 7/7 | 1.06 | + +All under the 5% rule; the CPU verifier 1.00 to 1.02 of v2 within one binary; the dataset build with the scatter store 20.7 to 21.8 ms against 21.1 linear on the M5 Max. Chip line: 512 B per hash, 120 to 128 distinct 64-B lines, the SRAM mirror is the whole dataset every hour (a windows-union census over 300 programs); the interleave buys nothing against a chip with a programmable address decoder (stated in the doc); the stride is a bijection with no cryptanalysis yet. Commits: b105a55 (implementation, history under tag ca2-era-pre-squash), c570da3, 669a27a, 78c0ee4 (the PC rows). The suite job 2 of 3 is build-20261005-215712 on PC 2 (kaspa-consensus alone). diff --git a/docs/plans/counter-asic-2.md b/docs/plans/counter-asic-2.md index 542ef627c..229a5a5a5 100644 --- a/docs/plans/counter-asic-2.md +++ b/docs/plans/counter-asic-2.md @@ -25,7 +25,7 @@ Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGP | 1 | keep v2's 128 x 4 B | w16 passes the rule but closes nothing (5090 139.8 against 136.1 MH/s, 9070 XT 17.90 against 18.15); w64 makes the 5090 bandwidth-bound (share 0.58, 37% of stream) | | 2 | out | spread across six programs 5.5 to 22.3% per card, over the 5% rule | | 3 | out (scratch share 0); the construct is sound and its tests stay | the on-die-cache recompute chip stays at 2.4x at every share under the 6 GB cap | -| 4 and 8 | the era draw of stride, interleave and the working-set window, in if the six-era spread is under 5% per card | pending the PC rows | +| 4 and 8 | IN: the era draw of stride, interleave and the working-set window (width pinned at 4 B) | six-era spread 1.3% on the 5090, 3.2% on the 9070 XT, 0.8% on the M5 Max; bit-exact on all three vendors | | 5 | OUT of v3 (a measured option for 3.0) | the added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% (g 0.87 / 0.85 / 0.84 and 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB against the 0.97 rule): no card keeps even 32 MiB resident while the dataset streams; the replaced form helps the on-die-cache chip (x1.33 at k = 4) | | 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 | | 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple | From d60caeb60832c94ace04de8c14fab15e1e31f219 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:58:53 +0000 Subject: [PATCH 095/131] Counter ASIC 2.0 status 21:58: ca2-v3 def920b with the cache merged --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 88e89ea9c..696150ffe 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -454,3 +454,5 @@ PC 2 order: suite job 2 of 3 (kaspa-consensus alone, running), the proving agent | M5 Max (Metal, 5 x 2^24) | 27.68 | 28.35 / 28.48 / 28.58 | 0.8% | 7/7 | 1.06 | All under the 5% rule; the CPU verifier 1.00 to 1.02 of v2 within one binary; the dataset build with the scatter store 20.7 to 21.8 ms against 21.1 linear on the M5 Max. Chip line: 512 B per hash, 120 to 128 distinct 64-B lines, the SRAM mirror is the whole dataset every hour (a windows-union census over 300 programs); the interleave buys nothing against a chip with a programmable address decoder (stated in the doc); the stride is a bijection with no cryptanalysis yet. Commits: b105a55 (implementation, history under tag ca2-era-pre-squash), c570da3, 669a27a, 78c0ee4 (the PC rows). The suite job 2 of 3 is build-20261005-215712 on PC 2 (kaspa-consensus alone). + +21:58. ca2-v3 tip fbf958e: ca2-cache 1950661 fast-forwarded (no conflict), V3_CLASS = { era: None, hot: None, ..MX4 }; igneum-pow 53 + 19 tests, packfile-test 0 failures; the fork's check against the new crate running; the fork stays 79bd8e10. The node doc carries the CPU hash rate across the switch in run 2 (about 30% under v2 on the CPU interpreter, approximate, a shared Mac) and the verifier before/after slot for the mixer fix. Gate run 3 on the fixed tree follows the mixer fix. From be67d769141948a19bc07e21f834f9ceda94f1a6 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:59:55 +0000 Subject: [PATCH 096/131] Counter ASIC 2.0 status 21:59: 0.3.10 shipped and merged, the 0.3.11 base is master 2a9dbd8 --- docs/plans/counter-asic-2-status.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 696150ffe..d9e12e8bb 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -456,3 +456,9 @@ PC 2 order: suite job 2 of 3 (kaspa-consensus alone, running), the proving agent All under the 5% rule; the CPU verifier 1.00 to 1.02 of v2 within one binary; the dataset build with the scatter store 20.7 to 21.8 ms against 21.1 linear on the M5 Max. Chip line: 512 B per hash, 120 to 128 distinct 64-B lines, the SRAM mirror is the whole dataset every hour (a windows-union census over 300 programs); the interleave buys nothing against a chip with a programmable address decoder (stated in the doc); the stride is a bijection with no cryptanalysis yet. Commits: b105a55 (implementation, history under tag ca2-era-pre-squash), c570da3, 669a27a, 78c0ee4 (the PC rows). The suite job 2 of 3 is build-20261005-215712 on PC 2 (kaspa-consensus alone). 21:58. ca2-v3 tip fbf958e: ca2-cache 1950661 fast-forwarded (no conflict), V3_CLASS = { era: None, hot: None, ..MX4 }; igneum-pow 53 + 19 tests, packfile-test 0 failures; the fork's check against the new crate running; the fork stays 79bd8e10. The node doc carries the CPU hash rate across the switch in run 2 (about 30% under v2 on the CPU interpreter, approximate, a shared Mac) and the verifier before/after slot for the mixer fix. Gate run 3 on the fixed tree follows the mixer fix. + +## 21:59 0.3.10 is shipped and merged to master (cde561c, pushed 21:58:17Z) + +The shipper's report: master cde561c (the 0.3.10 merge) + 1f0d62c (the plan); fork release-0.3.10 21d4c73c; digest 1f4b4425... on every node; six suites green on 21d4c73c on PC 2 with the same ban_is_decided flake under the parallel run (passes alone twice), recorded for the c4 agent. Open from it: PC 37ba0461 (the US laptop) stuck in its install since 21:41:16Z; Sam's Mac quit since 20:47Z; PC 2's prover dark (the root-socket cause is now named, the fix queued); C1 at 16:00Z; the /api/resume no-op (fixed on the 0.3.11 app branch). + +Consequence for 0.3.11: the main tree base is now master cde561c, so the integration merge is ca2-v3 (8ea6740 plus the mixer fix) and ca2-coord into master, with the app branch a223ca9 (on 5b0d54f, which master contains). The fork base stays 21d4c73c (= release-0.3.10's tip), so the fork merge is clean by construction. From 1852593e58471fd7c6fc217fb4edb925eb97ebff Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:00:30 +0000 Subject: [PATCH 097/131] Counter ASIC 2.0: G6 job 2 green; the master merge dry-run; status 22:00 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 6 ++++++ 2 files changed, 7 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 46013392b..e761bdd49 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,7 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms | GREEN (runs 1 and 2) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 runs kaspa-consensus alone, job 3 of 3 the other five crates; G6 is green only when both pass | +| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 (the other five crates and the app tests) follows the PC 2 prover socket fix; G6 is green only when job 3 passes | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index d9e12e8bb..bf439a669 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -462,3 +462,9 @@ All under the 5% rule; the CPU verifier 1.00 to 1.02 of v2 within one binary; th The shipper's report: master cde561c (the 0.3.10 merge) + 1f0d62c (the plan); fork release-0.3.10 21d4c73c; digest 1f4b4425... on every node; six suites green on 21d4c73c on PC 2 with the same ban_is_decided flake under the parallel run (passes alone twice), recorded for the c4 agent. Open from it: PC 37ba0461 (the US laptop) stuck in its install since 21:41:16Z; Sam's Mac quit since 20:47Z; PC 2's prover dark (the root-socket cause is now named, the fix queued); C1 at 16:00Z; the /api/resume no-op (fixed on the 0.3.11 app branch). Consequence for 0.3.11: the main tree base is now master cde561c, so the integration merge is ca2-v3 (8ea6740 plus the mixer fix) and ca2-coord into master, with the app branch a223ca9 (on 5b0d54f, which master contains). The fork base stays 21d4c73c (= release-0.3.10's tip), so the fork merge is clean by construction. + +## 22:00 G6 job 2 of 3 green; the merge into master dry-run + +build-20261005-215712 (21:57:12 to 21:59:45Z): kaspa-consensus alone 97 passed, 0 failed, 3 ignored (ban_is_decided ... ok), every stage ok. The PC 2 prover socket fix runs now (the proving agent, 60 s), then job 3 of 3 (consensus-core, igneum-exec, kaspa-pow, igneum-miner, kaspa-p2p-flows and the app tests). + +Merge dry-run into master cde561c (a scratch worktree, aborted): ca2-v3 (8ea6740) conflicts in docs/bench-log.md and proto-opencl/host.c; ca2-coord conflicts in docs/bench-log.md, proto-cuda/nvrtc/packfile.h and proto-opencl/host.c (master's 0.3.10 merge brought pack-loop's packfile.h and opencl-rdna4's host.c). The bench log is append-only (keep both); packfile.h and host.c take the ca2-v3 side (it carries the pack-loop rule plus the class, era, mixer and hot fields) re-checked against master's hunks. The integration merge is the ship's first step. From 65600a5caacced75d9285ecb20c30ca9020bff57 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:01:10 +0000 Subject: [PATCH 098/131] Counter ASIC 2.0 status 22:01: the regression bisected, the mixer PC job, the PC 2 queue, the ship template --- docs/plans/counter-asic-2-status.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index bf439a669..c22846e10 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -468,3 +468,11 @@ Consequence for 0.3.11: the main tree base is now master cde561c, so the integra build-20261005-215712 (21:57:12 to 21:59:45Z): kaspa-consensus alone 97 passed, 0 failed, 3 ignored (ban_is_decided ... ok), every stage ok. The PC 2 prover socket fix runs now (the proving agent, 60 s), then job 3 of 3 (consensus-core, igneum-exec, kaspa-pow, igneum-miner, kaspa-p2p-flows and the app tests). Merge dry-run into master cde561c (a scratch worktree, aborted): ca2-v3 (8ea6740) conflicts in docs/bench-log.md and proto-opencl/host.c; ca2-coord conflicts in docs/bench-log.md, proto-cuda/nvrtc/packfile.h and proto-opencl/host.c (master's 0.3.10 merge brought pack-loop's packfile.h and opencl-rdna4's host.c). The bench log is append-only (keep both); packfile.h and host.c take the ca2-v3 side (it carries the pack-loop rule plus the class, era, mixer and hot fields) re-checked against master's hunks. The integration merge is the ship's first step. + +## 22:01 the verifier regression bisected to the mixer's 0fc0ad1; the mixer PC 1 job is running; PC 2 queue + +Bisection (mixer agent, measure lock, 21:59 UTC, same input 19297e99c7b9a55e, avg of 50, two rounds): readwidth e752fc7 0.607 / 0.609 ms per warp; the ca2-v3 seam 6c75dad 0.610 / 0.609; the mixer's 0fc0ad1 1.332 / 1.316; ca2-v3 HEAD 88dafbc 1.325 / 1.347. The 2.2x is in 0fc0ad1's derive_items / Cache path at m = 1 (not the era layout, not the load). Three candidate fixes building (the constant line mask back in Cache::line; an m == 1 fast path that is readwidth's loop verbatim; both); the one that restores 0.61 goes on top of ca2-v3 8ea6740 with the six mixer commits rebased, measured the same way; then the node agent merges, rebuilds and runs gate 3. + +PC 1: fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005 published 21:59:32Z (zip 152fcf93...; workers e6007918... and e4334aaf... from ca2-mixer; the five packs; 25-minute timeout). PC 2: the proving agent's socket fix (60 s) now, then suite job 3 of 3, then the prover-floor agent's toolchain check (3 min) and its 60-minute niced sp1-gpu-server build (CPU only; the shipped server panics on any card under 24 GB, sp1-gpu builder.rs 35 to 39, and allocates every prover at Setup: the 13.9 GB floor's cause), then the aggregation-cost re-run (20 min), then the prover-floor measurements. + +0.3.11 ship template (release-0.3.10.md section 7): push the release branch, `gh workflow run windows.yml --ref `, `node tools/ship-app.mjs 0.3.11 --node --branch --public --activation-height N4 --deadline-note "program class v3 + proving v1" --notes "..." [--from ci]` with the override object carrying every switch; gh must be on igneum-labs; the pre-push hook flips two site files (restore with `git checkout -- site/`). From 27a7ca6550d232c40fc5674a46f21d90daac7c65 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:03:35 +0000 Subject: [PATCH 099/131] Counter ASIC 2.0 status 22:03: the PC 2 prover back, suite job 3 published --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index c22846e10..bfc37ffe6 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -476,3 +476,7 @@ Bisection (mixer agent, measure lock, 21:59 UTC, same input 19297e99c7b9a55e, av PC 1: fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005 published 21:59:32Z (zip 152fcf93...; workers e6007918... and e4334aaf... from ca2-mixer; the five packs; 25-minute timeout). PC 2: the proving agent's socket fix (60 s) now, then suite job 3 of 3, then the prover-floor agent's toolchain check (3 min) and its 60-minute niced sp1-gpu-server build (CPU only; the shipped server panics on any card under 24 GB, sp1-gpu builder.rs 35 to 39, and allocates every prover at Setup: the 13.9 GB floor's cause), then the aggregation-cost re-run (20 min), then the prover-floor measurements. 0.3.11 ship template (release-0.3.10.md section 7): push the release branch, `gh workflow run windows.yml --ref `, `node tools/ship-app.mjs 0.3.11 --node --branch --public --activation-height N4 --deadline-note "program class v3 + proving v1" --notes "..." [--from ci]` with the override object carrying every switch; gh must be on igneum-labs; the pre-push hook flips two site files (restore with `git checkout -- site/`). + +## 22:03 PC 2's prover is back; suite job 3 of 3 published + +socketfix-pc2-pv1 (22:01:14 to 22:02:12Z): /tmp/sp1-cuda-0.sock owned by root removed (the aggregation-cost job's run), the prover switched off and on; the next shard (block 89011 shard 0) "proven and submitted in 34 s" at 22:02:13Z and paid 0.93116546 IGN at 22:02:24Z. PC 2's prover had been dark from 21:25:24Z to 22:02 (the root-socket class; the CI check tools/ci/prover-socket-check.sh now fails any playbook without the two restore lines). Suite job 3 of 3 (consensus-core, igneum-exec, kaspa-pow, igneum-miner, kaspa-p2p-flows and the app tests; main 8ea6740, fork 79bd8e10) is packing and publishing from the ca2-v3 worktree now. After it on PC 2: the prover-floor agent's toolchain check and its 60-minute build, then the aggregation-cost re-run. From a941b8b467fdc600241fc0c0ed35786614687d7a Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:06:59 +0000 Subject: [PATCH 100/131] Counter ASIC 2.0: mixer x8 decided into v3 on the PC build rows; status 22:06 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 6 ++++++ docs/plans/counter-asic-2.md | 2 +- 3 files changed, 8 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index e761bdd49..c365c3366 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. Refinement (coordinator, 21:28 UTC, delegated under "as strong as the measurements allow"): x8 is built and measured beside x4 on the same packs (verifier per warp on one Mac core, the 1 GiB daily build on the 5090, the M5 Max and the 9070 XT or gfx1036, the chip row at equal silicon with the 3x factor); x8 goes into v3 if the per-warp verify stays under 10 ms on one core and the daily build stays under 1 s on every card we own, otherwise x4 with the thin margin stated in level 3 and x8 named as the next lever. The vectors are re-cut once, after this choice. Measured so far: x4 chip row 0.61x bare, 1.84x with the factor, 1.53x with the 128 mm^2 mirror deducted; x8 from the m16 table 0.31x bare, 0.92x with the factor (approximate until measured). The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. Refinement (coordinator, 21:28 UTC, delegated under "as strong as the measurements allow"): x8 is built and measured beside x4 on the same packs (verifier per warp on one Mac core, the 1 GiB daily build on the 5090, the M5 Max and the 9070 XT or gfx1036, the chip row at equal silicon with the 3x factor); x8 goes into v3 if the per-warp verify stays under 10 ms on one core and the daily build stays under 1 s on every card we own, otherwise x4 with the thin margin stated in level 3 and x8 named as the next lever. DECIDED (5 October 2026, 22:06 UTC, delegated): x8. Both halves pass: the per-warp verify at x8 is 2.79 ms on a loaded M5 Max core (2.1x v2; about 1.3 ms quiet, approximate) against the 10 ms gate; the daily 1 GiB build does not move with the mixer on any discrete card (RTX 5090 23 to 25 ms, RX 9070 XT 72 to 77 ms, M5 Max 21 ms at x1, x4 and x8: latency-bound), 13x to 40x under the 1 s bar (PC 1 jobs fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005, 22:00:08 to 22:04:39Z, both cards restored, every pack's fingerprint equal to the Mac's). V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }. Chip row at x8: 1,198,080 ops per hash, 41.7 MH/s at 50 T op/s, 0.31x bare, 0.92x with the 3x fixed-function factor, 0.76x at equal silicon: the claim reads under 1x with the factor, margin stated. The vectors are re-cut once on this class. Cost: pool shares per core and IBD time scale with the verifier (x8: 2.1x v2); the integrated tier mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ### 6b. The user tiers (the consequences rule) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index bfc37ffe6..1dcf5933c 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -480,3 +480,9 @@ PC 1: fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005 published 21:59:32Z ## 22:03 PC 2's prover is back; suite job 3 of 3 published socketfix-pc2-pv1 (22:01:14 to 22:02:12Z): /tmp/sp1-cuda-0.sock owned by root removed (the aggregation-cost job's run), the prover switched off and on; the next shard (block 89011 shard 0) "proven and submitted in 34 s" at 22:02:13Z and paid 0.93116546 IGN at 22:02:24Z. PC 2's prover had been dark from 21:25:24Z to 22:02 (the root-socket class; the CI check tools/ci/prover-socket-check.sh now fails any playbook without the two restore lines). Suite job 3 of 3 (consensus-core, igneum-exec, kaspa-pow, igneum-miner, kaspa-p2p-flows and the app tests; main 8ea6740, fork 79bd8e10) is packing and publishing from the ca2-v3 worktree now. After it on PC 2: the prover-floor agent's toolchain check and its 60-minute build, then the aggregation-cost re-run. + +## 22:06 decided: mixer x8 into v3; the verifier fix found; ca2-v3 at 795472e + +The mixer PC 1 job (22:00:08 to 22:04:39Z, both cards restored, the 9070 XT present): the daily 1 GiB build per pack, two dispatches, wall ms: RTX 5090 v2 25 / 23, mx4 24 / 24 and 23 / 25, mx8 23 / 23 and 23 / 23 (cache 4); RX 9070 XT v2 74 / 74, mx4 77 / 75 and 73 / 74, mx8 72 / 76 and 75 / 74 (cache 8 to 9). The build is latency-bound on every card; the x8 rule's build half passes with 13x to 40x margin; its verifier half passed at 2.79 ms per warp on the loaded core. DECIDED (delegated): x8 into v3; V3_CLASS = { era: None, hot: None, ..MX8 }; the chip row at x8 reads 0.92x with the 3x factor (the claim "under 1x with the factor", margin stated). Every mixer pack's fingerprint equals the Mac's on both cards (v2 25f96e7dce90bd4e; mx4 6f48d5a2aa0dbe5f, 73caaebb28e808fe; mx8 7c28cfb06c5c65a9, bbb183f72692f840); hash rates at the v2 rate on both (5090 136.5 to 137.4, 9070 XT 18.0 to 18.2 at every class). + +The verifier regression is found: not the mask but inlining; the item loop inlined into MemhardCpu::fetch runs at 1.33 ms per unit, the same loop out of line (#[inline(never)], one instance per cache size, the line mask a constant) at 0.60 to 0.62 against readwidth's 0.60 to 0.64 in the same minute. The fix, the MX8 V3_CLASS, the pinned pack re-export and the final v2 / x4 / x8 session land as one commit on ca2-v3 795472e (the node agent merged the mixer's 16dfd1e as 4e733bb with two one-line field fixes; 53 + 4 + 19 + 7 tests, packfile 0 failures, the fork check clean). Then: the node agent merges, rebuilds, gate run 3 on the final class; the six era packs and the pinned v3 pack re-exported on it; the final bit-exactness and G2 (1,000 random hashes per card re-hashed by the CPU) job on PC 1. diff --git a/docs/plans/counter-asic-2.md b/docs/plans/counter-asic-2.md index 229a5a5a5..7482f9934 100644 --- a/docs/plans/counter-asic-2.md +++ b/docs/plans/counter-asic-2.md @@ -29,7 +29,7 @@ Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGP | 5 | OUT of v3 (a measured option for 3.0) | the added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% (g 0.87 / 0.85 / 0.84 and 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB against the 0.97 rule): no card keeps even 32 MiB resident while the dataset streams; the replaced form helps the on-die-cache chip (x1.33 at k = 4) | | 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 | | 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple | -| M16 mixer x4 | into v3 | the only measured lever that moves the named chip: 0.6x bare, 1.8x with a 3x fixed-function factor; verifier 1.6 to 4.8 ms per warp | +| M16 mixer x8 (x4 measured beside it) | into v3 | the only measured lever that moves the named chip: x8 0.31x bare, 0.92x with a 3x fixed-function factor (x4: 0.61x, 1.84x); verifier 2.1x v2 per warp (2.79 ms on a loaded core, about 1.3 ms quiet); the daily build unmoved on every card (latency-bound) | | 9 | reserve-only, no change to the devnet's hour | the floor 600 s from the slowest compile-ahead (38 s) | Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors). From 90589ff60cbbccf35e2987de7fae267a1b3cd88b Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:08:07 +0000 Subject: [PATCH 101/131] Counter ASIC 2.0: G6 job 3 failed on a stale fork test, job 4 on the final tree; status 22:08 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index c365c3366..94f09708e 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,7 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms | GREEN (runs 1 and 2) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 (the other five crates and the app tests) follows the PC 2 prover socket fix; G6 is green only when job 3 passes | +| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 re-runs the five crates and the app on the final tree (the fork fix and the mixer's verifier fix with MX8 merged); G6 is green only when job 4 passes | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 1dcf5933c..de8dab039 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -486,3 +486,7 @@ socketfix-pc2-pv1 (22:01:14 to 22:02:12Z): /tmp/sp1-cuda-0.sock owned by root re The mixer PC 1 job (22:00:08 to 22:04:39Z, both cards restored, the 9070 XT present): the daily 1 GiB build per pack, two dispatches, wall ms: RTX 5090 v2 25 / 23, mx4 24 / 24 and 23 / 25, mx8 23 / 23 and 23 / 23 (cache 4); RX 9070 XT v2 74 / 74, mx4 77 / 75 and 73 / 74, mx8 72 / 76 and 75 / 74 (cache 8 to 9). The build is latency-bound on every card; the x8 rule's build half passes with 13x to 40x margin; its verifier half passed at 2.79 ms per warp on the loaded core. DECIDED (delegated): x8 into v3; V3_CLASS = { era: None, hot: None, ..MX8 }; the chip row at x8 reads 0.92x with the 3x factor (the claim "under 1x with the factor", margin stated). Every mixer pack's fingerprint equals the Mac's on both cards (v2 25f96e7dce90bd4e; mx4 6f48d5a2aa0dbe5f, 73caaebb28e808fe; mx8 7c28cfb06c5c65a9, bbb183f72692f840); hash rates at the v2 rate on both (5090 136.5 to 137.4, 9070 XT 18.0 to 18.2 at every class). The verifier regression is found: not the mask but inlining; the item loop inlined into MemhardCpu::fetch runs at 1.33 ms per unit, the same loop out of line (#[inline(never)], one instance per cache size, the line mask a constant) at 0.60 to 0.62 against readwidth's 0.60 to 0.64 in the same minute. The fix, the MX8 V3_CLASS, the pinned pack re-export and the final v2 / x4 / x8 session land as one commit on ca2-v3 795472e (the node agent merged the mixer's 16dfd1e as 4e733bb with two one-line field fixes; 53 + 4 + 19 + 7 tests, packfile 0 failures, the fork check clean). Then: the node agent merges, rebuilds, gate run 3 on the final class; the six era packs and the pinned v3 pack re-exported on it; the final bit-exactness and G2 (1,000 random hashes per card re-hashed by the CPU) job on PC 1. + +## 22:08 G6 job 3 failed on a stale fork test (era inside the class); job 4 on the final tree + +build-20261005-220351 (22:04 to 22:06:55Z): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8; kaspa-pow 13 passed, 1 failed: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache asserts the program's class equals V3_CLASS with era: None, but the merged crate carries the drawn era inside the class (left: era Some(EraParams { ... stride_mul 3969900165 ... }); right: era None). A stale fork test written before the era merge, not a behaviour fault; the node agent fixes it on the fork. Correction to the G6 note: the PC's test stage DOES run the v3 engine test, so the PC job is the evidence. Job 4 (the five crates and the app) runs on the final tree once the fork fix and the mixer's verifier fix (with V3_CLASS = MX8) are merged; the publish waits on it. PC 1: the 5090 power sweep has the go (8 min), the AMD sweep may follow on its own presence probe; the final v3 vectors and G2 job (the era agent, 1,000 hashes per card re-hashed on the Mac) is being prepared for about 22:40. From c597a7dcdb015db24426012799eebf6f473d2f74 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:09:59 +0000 Subject: [PATCH 102/131] Counter ASIC 2.0 status 22:09: the fork test fixed at 89dfcb95 --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index de8dab039..f96a11bdd 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -490,3 +490,5 @@ The verifier regression is found: not the mask but inlining; the item loop inlin ## 22:08 G6 job 3 failed on a stale fork test (era inside the class); job 4 on the final tree build-20261005-220351 (22:04 to 22:06:55Z): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8; kaspa-pow 13 passed, 1 failed: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache asserts the program's class equals V3_CLASS with era: None, but the merged crate carries the drawn era inside the class (left: era Some(EraParams { ... stride_mul 3969900165 ... }); right: era None). A stale fork test written before the era merge, not a behaviour fault; the node agent fixes it on the fork. Correction to the G6 note: the PC's test stage DOES run the v3 engine test, so the PC job is the evidence. Job 4 (the five crates and the app) runs on the final tree once the fork fix and the mixer's verifier fix (with V3_CLASS = MX8) are merged; the publish waits on it. PC 1: the 5090 power sweep has the go (8 min), the AMD sweep may follow on its own presence probe; the final v3 vectors and G2 job (the era agent, 1,000 hashes per card re-hashed on the Mac) is being prepared for about 22:40. + +22:09. Fork ca2-v3-node 89dfcb95: the v3 engine test asserts what it meant on the era crate (LoadClass { era: None, ..class } == V3_CLASS and class.era.is_some(); the era bytes equal the seeds'; another era seed keeps the program id, draws another era class, hashes another pow, adds no day cache); Mac `cargo test -p kaspa-pow --features igneum-pow` 14 passed. Main-repo tip 6a705a2. Waiting on the mixer's fix commit for the merge, the rebuild, gate run 3 and suite job 4. PC 2: the prover-floor agent's 3-minute toolchain check has the go; its build waits for job 4. From 615837dbd0c2e09a5af1bd16448f436df1066526 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:11:10 +0000 Subject: [PATCH 103/131] Counter ASIC 2.0 status 22:11: the toolchain check, the build held behind job 4, C26 and C27 --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index f96a11bdd..308207df9 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -492,3 +492,5 @@ The verifier regression is found: not the mask but inlining; the item loop inlin build-20261005-220351 (22:04 to 22:06:55Z): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8; kaspa-pow 13 passed, 1 failed: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache asserts the program's class equals V3_CLASS with era: None, but the merged crate carries the drawn era inside the class (left: era Some(EraParams { ... stride_mul 3969900165 ... }); right: era None). A stale fork test written before the era merge, not a behaviour fault; the node agent fixes it on the fork. Correction to the G6 note: the PC's test stage DOES run the v3 engine test, so the PC job is the evidence. Job 4 (the five crates and the app) runs on the final tree once the fork fix and the mixer's verifier fix (with V3_CLASS = MX8) are merged; the publish waits on it. PC 1: the 5090 power sweep has the go (8 min), the AMD sweep may follow on its own presence probe; the final v3 vectors and G2 job (the era agent, 1,000 hashes per card re-hashed on the Mac) is being prepared for about 22:40. 22:09. Fork ca2-v3-node 89dfcb95: the v3 engine test asserts what it meant on the era crate (LoadClass { era: None, ..class } == V3_CLASS and class.era.is_some(); the era bytes equal the seeds'; another era seed keeps the program id, draws another era class, hashes another pow, adds no day cache); Mac `cargo test -p kaspa-pow --features igneum-pow` 14 passed. Main-repo tip 6a705a2. Waiting on the mixer's fix commit for the merge, the rebuild, gate run 3 and suite job 4. PC 2: the prover-floor agent's 3-minute toolchain check has the go; its build waits for job 4. + +22:11. PC 2: the prover-floor toolchain check done (floor-toolchain-1, 22:10:03 to 22:10:06Z: nvcc 12.8, cmake 3.28.3, gcc 13.3, clang 18, cargo 1.99.0; no go, so the rebuilt server drops native-gnark, the Groth16 wrap that compressed proofs never use; 16 cores, 30 GB WSL RAM; the live server untouched); its 90-minute build (sm_86, sm_89, sm_120 after C26) is HELD behind suite job 4 (the gate), with a 22:40 fallback: if job 4 is not published by then, the build goes first. Consequences C26 (the arch list, the card, the packaging row, the verify-segment run, "24 GB" kept until the 3060 proves) is with the prover-floor agent; C27 (publish-jobs.sh add runs the prover-socket and bash-body checks and refuses on failure) is being wired by the reviewer's sub-agent, no collision. From 199c974f3d8d88b15682e34c11da81b842da85d0 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:11:49 +0000 Subject: [PATCH 104/131] Counter ASIC 2.0: the verifier fix and the x8 class committed; status 22:11 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 94f09708e..8d24204a7 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -54,7 +54,7 @@ the project lead's rules, applied by the coordinator and recorded here with the ### 6a. The chip model's headline, and how the scratch share is chosen -The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. Refinement (coordinator, 21:28 UTC, delegated under "as strong as the measurements allow"): x8 is built and measured beside x4 on the same packs (verifier per warp on one Mac core, the 1 GiB daily build on the 5090, the M5 Max and the 9070 XT or gfx1036, the chip row at equal silicon with the 3x factor); x8 goes into v3 if the per-warp verify stays under 10 ms on one core and the daily build stays under 1 s on every card we own, otherwise x4 with the thin margin stated in level 3 and x8 named as the next lever. DECIDED (5 October 2026, 22:06 UTC, delegated): x8. Both halves pass: the per-warp verify at x8 is 2.79 ms on a loaded M5 Max core (2.1x v2; about 1.3 ms quiet, approximate) against the 10 ms gate; the daily 1 GiB build does not move with the mixer on any discrete card (RTX 5090 23 to 25 ms, RX 9070 XT 72 to 77 ms, M5 Max 21 ms at x1, x4 and x8: latency-bound), 13x to 40x under the 1 s bar (PC 1 jobs fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005, 22:00:08 to 22:04:39Z, both cards restored, every pack's fingerprint equal to the Mac's). V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }. Chip row at x8: 1,198,080 ops per hash, 41.7 MH/s at 50 T op/s, 0.31x bare, 0.92x with the 3x fixed-function factor, 0.76x at equal silicon: the claim reads under 1x with the factor, margin stated. The vectors are re-cut once on this class. Cost: pool shares per core and IBD time scale with the verifier (x8: 2.1x v2); the integrated tier mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). +The layer 6 finding changes the headline: the strongest chip holds the whole cache on-die (about 130 mm^2 at N5/N3E and 165 mm^2 at 7 nm by shipped-product SRAM density, AMD 3D V-Cache 1.56 MB/mm^2 and TSMC N5 HD macros; 54 to 83 mm^2 is the bit-cell-only lower bound; `docs/analysis/sram-mirror.md` after the 20:18 correction) and computes dataset items on the fly through M16's mixer. The before-and-after table must carry that chip as a named row against the GPU for each variant, and the scratch share is chosen by that row: the smallest share at which the on-die-cache chip's gain falls under 1.5x. Result (20:23 UTC): no scratch share under the 6 GB cap gets that chip under 2x; the gain is 2.4x at every share. The site's "under 2x" claim is therefore qualified until the mixer multiplier or the cache rule closes it: M16's mixer multiplier x2 gives 1.2x (3.6x with a 3x fixed-function factor) at 0.8 to 2.4 ms verify per warp, x4 gives 0.6x (1.8x with the factor) at 1.6 to 4.8 ms, inside the 10 ms gate, with the 5090's daily dataset build at 27 and 54 ms. DECIDED (5 October 2026, delegated under "execute the full 2.0 plan" and "deploy what is absolute best"; the project lead confirms for the public testnet genesis): M16 mixer x4 goes into v3 behind the same activation. Numbers: attacker 0.083 Ghash/s at 50 T op/s (0.36x bare against 229 MH/s; 1.8x against the 5090's measured 139.7 MH/s with a 3x fixed-function factor, approximate); verifier 1.6 to 4.8 ms per warp (inside the 10 ms gate); the 5090's daily dataset build 54 ms (13.4 x 4, measurement owed). Layer 3 stays out (scratch share 0); its soundness document and pack-contract tests are kept because the construct is sound and may return. The chip model's headline row becomes the on-die-cache recompute chip against v3 with everything combined (x4 mixer, the hot table, the width rule, the era draws, the cache growth), fixed-function factor included; the public level 3 shows that row: if it reads 1.8x the claim is "under 2x" with the margin stated as thin and the mixer x8 and the hot table named as the next levers. The dataset and cache vectors are re-cut once for v3 (cache growth rule and x4 together), the soundness suite re-run on the new construction, bit-exact on the three cards, the verifier per-warp time measured on the Mac; the daily dataset build time quoted for the 5090, the Mac and the 9070 XT. Branch ca2-mixer carries it. Refinement (coordinator, 21:28 UTC, delegated under "as strong as the measurements allow"): x8 is built and measured beside x4 on the same packs (verifier per warp on one Mac core, the 1 GiB daily build on the 5090, the M5 Max and the 9070 XT or gfx1036, the chip row at equal silicon with the 3x factor); x8 goes into v3 if the per-warp verify stays under 10 ms on one core and the daily build stays under 1 s on every card we own, otherwise x4 with the thin margin stated in level 3 and x8 named as the next lever. DECIDED (5 October 2026, 22:06 UTC, delegated): x8. Both halves pass: the per-warp verify at x8 is 2.79 ms on a loaded M5 Max core (2.1x v2; about 1.3 ms quiet, approximate) against the 10 ms gate; the daily 1 GiB build does not move with the mixer on any discrete card (RTX 5090 23 to 25 ms, RX 9070 XT 72 to 77 ms, M5 Max 21 ms at x1, x4 and x8: latency-bound), 13x to 40x under the 1 s bar (PC 1 jobs fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005, 22:00:08 to 22:04:39Z, both cards restored, every pack's fingerprint equal to the Mac's). V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }. Chip row at x8: 1,198,080 ops per hash, 41.7 MH/s at 50 T op/s, 0.31x bare, 0.92x with the 3x fixed-function factor, 0.76x at equal silicon: the claim reads under 1x with the factor, margin 8% on the factor and 9% on the budget (chip-model-v3.md). Verifier on the fixed crate (ca2-mixer 1ab8b21, 22:07 UTC, same input beside readwidth's binary, load 5.5): v2 0.609 / 0.611 ms per unit (readwidth 0.607 / 0.610), x4 1.238 / 1.237 (2.0x, worst cold 1.40), x8 2.077 / 2.058 (3.4x, worst cold 2.15): 4.8x inside the 10 ms gate on this loaded core. The vectors are re-cut once on this class. Cost: pool shares per core and IBD time scale with the verifier (x8: 2.1x v2); the integrated tier mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). The scratch-soundness analysis carries this table (`docs/analysis/scratch-soundness.md`, question 2). ### 6b. The user tiers (the consequences rule) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 308207df9..812ad2afb 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -494,3 +494,7 @@ build-20261005-220351 (22:04 to 22:06:55Z): igneum-exec 17, igneum-miner 18, con 22:09. Fork ca2-v3-node 89dfcb95: the v3 engine test asserts what it meant on the era crate (LoadClass { era: None, ..class } == V3_CLASS and class.era.is_some(); the era bytes equal the seeds'; another era seed keeps the program id, draws another era class, hashes another pow, adds no day cache); Mac `cargo test -p kaspa-pow --features igneum-pow` 14 passed. Main-repo tip 6a705a2. Waiting on the mixer's fix commit for the merge, the rebuild, gate run 3 and suite job 4. PC 2: the prover-floor agent's 3-minute toolchain check has the go; its build waits for job 4. 22:11. PC 2: the prover-floor toolchain check done (floor-toolchain-1, 22:10:03 to 22:10:06Z: nvcc 12.8, cmake 3.28.3, gcc 13.3, clang 18, cargo 1.99.0; no go, so the rebuilt server drops native-gnark, the Groth16 wrap that compressed proofs never use; 16 cores, 30 GB WSL RAM; the live server untouched); its 90-minute build (sm_86, sm_89, sm_120 after C26) is HELD behind suite job 4 (the gate), with a 22:40 fallback: if job 4 is not published by then, the build goes first. Consequences C26 (the arch list, the card, the packaging row, the verify-segment run, "24 GB" kept until the 3060 proves) is with the prover-floor agent; C27 (publish-jobs.sh add runs the prover-socket and bash-body checks and refuses on failure) is being wired by the reviewer's sub-agent, no collision. + +## 22:11 the verifier fix and the x8 class are committed (ca2-mixer 1ab8b21); the final code is in + +Before / after, the era agent's way (one measure session, 22:07 UTC, readwidth e752fc7's binary beside the fixed one, the same v2 input, load 5.5): readwidth 0.607 / 0.610 ms per unit; the fixed binary 0.609 / 0.611 (was 1.332 / 1.316 on 0fc0ad1). On the fixed binary: x4 1.238 / 1.237 (2.0x v2, worst cold 1.40), x8 2.077 / 2.058 (3.4x, worst cold 2.15), x8 on the devnet seeds 2.058; so the x8 class verifies at 2.1 ms per warp on a loaded core, 4.8x inside the 10 ms gate. Cause and fix: the item loop inlined into MemhardCpu::fetch ran at 2.2x whatever the mask; derive_items_mask #[inline(never)], one instance per cache size (2^26 to 2^30) with the line mask a constant, restores 0.61. V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }; mx8-genesis (program id e323b9dcaf283a6f, fingerprint 7c28cfb06c5c65a9) and mx8-devnet-epoch0 through the chain path with the era inside (class mx8-erad810f22d, program id 73bcbfe8ccf988f1, unit 0 lane 0 d424577fce4a7a60, fingerprint 90f794dd556f7a3b on Metal and Apple OpenCL, 3/3 + 3/3, 96 of 96); the suite 53 + 4 + 19 + 7. chip-model-v3.md headline: x8, 0.92x with the factor, margin 8% on the factor and 9% on the budget; x4 kept as the measured candidate. Owed: the composed mx8-devnet-epoch0's fingerprint on the PCs (in the final-vectors job). Next: the node agent fast-forwards ca2-v3, re-checks, rebuilds, gate run 3; suite job 4 publishes when the tree is settled; the era agent builds the final-vectors and G2 package on 1ab8b21. From 271db8e5ac21fa204e05527ee6a90e0b71e81cd3 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:13:25 +0000 Subject: [PATCH 105/131] Counter ASIC 2.0 status 22:13: suite job 4 published on the final tree --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 8d24204a7..fe6713752 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,7 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms | GREEN (runs 1 and 2) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 re-runs the five crates and the app on the final tree (the fork fix and the mixer's verifier fix with MX8 merged); G6 is green only when job 4 passes | +| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 build-20261005-221237 (published 22:12:37Z: main d233fa1 = era + cache + the mixer fix with V3_CLASS = MX8, fork 89dfcb95) runs the five crates and the app on the final tree; G6 is green only when job 4 passes | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 812ad2afb..e826725f6 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -498,3 +498,5 @@ build-20261005-220351 (22:04 to 22:06:55Z): igneum-exec 17, igneum-miner 18, con ## 22:11 the verifier fix and the x8 class are committed (ca2-mixer 1ab8b21); the final code is in Before / after, the era agent's way (one measure session, 22:07 UTC, readwidth e752fc7's binary beside the fixed one, the same v2 input, load 5.5): readwidth 0.607 / 0.610 ms per unit; the fixed binary 0.609 / 0.611 (was 1.332 / 1.316 on 0fc0ad1). On the fixed binary: x4 1.238 / 1.237 (2.0x v2, worst cold 1.40), x8 2.077 / 2.058 (3.4x, worst cold 2.15), x8 on the devnet seeds 2.058; so the x8 class verifies at 2.1 ms per warp on a loaded core, 4.8x inside the 10 ms gate. Cause and fix: the item loop inlined into MemhardCpu::fetch ran at 2.2x whatever the mask; derive_items_mask #[inline(never)], one instance per cache size (2^26 to 2^30) with the line mask a constant, restores 0.61. V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }; mx8-genesis (program id e323b9dcaf283a6f, fingerprint 7c28cfb06c5c65a9) and mx8-devnet-epoch0 through the chain path with the era inside (class mx8-erad810f22d, program id 73bcbfe8ccf988f1, unit 0 lane 0 d424577fce4a7a60, fingerprint 90f794dd556f7a3b on Metal and Apple OpenCL, 3/3 + 3/3, 96 of 96); the suite 53 + 4 + 19 + 7. chip-model-v3.md headline: x8, 0.92x with the factor, margin 8% on the factor and 9% on the budget; x4 kept as the measured candidate. Owed: the composed mx8-devnet-epoch0's fingerprint on the PCs (in the final-vectors job). Next: the node agent fast-forwards ca2-v3, re-checks, rebuilds, gate run 3; suite job 4 publishes when the tree is settled; the era agent builds the final-vectors and G2 package on 1ab8b21. + +22:13. Suite job 4 published: build-20261005-221237 (main d233fa1 = ca2-v3 with the mixer fix merged, V3_CLASS = { era: None, hot: None, ..MX8 }; fork 89dfcb95 with the engine test fixed; the five crates and the app). The node agent's checks, rebuild and gate run 3 on d233fa1 follow; the era agent's final-vectors and G2 package is being built on the final class; PC 1 is the AMD sweep's until about 22:25. From 0f82258d7d065584b1c38f81390c1eb9f45bab1b Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:15:07 +0000 Subject: [PATCH 106/131] Counter ASIC 2.0: the final-vectors package, the bash-body-check branch in the 0.3.11 tree; status 22:15 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index fe6713752..b0403c541 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -103,7 +103,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL a223ca9 on 5b0d54f, docs-only 8b47073 after it (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. -Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. +Branches merged into the 0.3.11 main tree beside ca2-v3 and ca2-coord (tooling, no consensus): consequences (the reviewer's rows), bash-body-check (7adb1ca, 6805125: tools/ci/bash-body-check.sh with fixtures and a self-test in ci.yml, the convention paragraph in packaging/README-ship.md, `publish-jobs.sh add --kind run` running the bash-body and prover-socket checks before signing; add/add with proving-v1 c2544be on tools/ci/prover-socket-check.sh: take bash-body-check's superset and proving-v1's ci.yml step; tools/amd-prove/pc1-cpu-prove.ps1 goes on the allow list, it starts no GPU server), amd-prove (f1d7a7d), card-lifetime (1fecfe2). Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. ## 9. After the publish: Counter ASIC 3.0 diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index e826725f6..7b3e3fb14 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -500,3 +500,5 @@ build-20261005-220351 (22:04 to 22:06:55Z): igneum-exec 17, igneum-miner 18, con Before / after, the era agent's way (one measure session, 22:07 UTC, readwidth e752fc7's binary beside the fixed one, the same v2 input, load 5.5): readwidth 0.607 / 0.610 ms per unit; the fixed binary 0.609 / 0.611 (was 1.332 / 1.316 on 0fc0ad1). On the fixed binary: x4 1.238 / 1.237 (2.0x v2, worst cold 1.40), x8 2.077 / 2.058 (3.4x, worst cold 2.15), x8 on the devnet seeds 2.058; so the x8 class verifies at 2.1 ms per warp on a loaded core, 4.8x inside the 10 ms gate. Cause and fix: the item loop inlined into MemhardCpu::fetch ran at 2.2x whatever the mask; derive_items_mask #[inline(never)], one instance per cache size (2^26 to 2^30) with the line mask a constant, restores 0.61. V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }; mx8-genesis (program id e323b9dcaf283a6f, fingerprint 7c28cfb06c5c65a9) and mx8-devnet-epoch0 through the chain path with the era inside (class mx8-erad810f22d, program id 73bcbfe8ccf988f1, unit 0 lane 0 d424577fce4a7a60, fingerprint 90f794dd556f7a3b on Metal and Apple OpenCL, 3/3 + 3/3, 96 of 96); the suite 53 + 4 + 19 + 7. chip-model-v3.md headline: x8, 0.92x with the factor, margin 8% on the factor and 9% on the budget; x4 kept as the measured candidate. Owed: the composed mx8-devnet-epoch0's fingerprint on the PCs (in the final-vectors job). Next: the node agent fast-forwards ca2-v3, re-checks, rebuilds, gate run 3; suite job 4 publishes when the tree is settled; the era agent builds the final-vectors and G2 package on 1ab8b21. 22:13. Suite job 4 published: build-20261005-221237 (main d233fa1 = ca2-v3 with the mixer fix merged, V3_CLASS = { era: None, hot: None, ..MX8 }; fork 89dfcb95 with the engine test fixed; the five crates and the app). The node agent's checks, rebuild and gate run 3 on d233fa1 follow; the era agent's final-vectors and G2 package is being built on the final class; PC 1 is the AMD sweep's until about 22:25. + +22:15. The era final-class package is ready (zip igneum-ca2-era-pc1b.zip sha256 cb0e9db07e304b11fd4c0591351af46090442ea4f51d60eb86945b96bd28aba3; workers f8d19f0a... and 87647c15... from 1ab8b21 plus the era branch; seven packs on the final class: mx8-devnet-epoch0 and era-0 to era-5, program id 73bcbfe8ccf988f1; Mac 7/7 on Metal and Apple OpenCL, fingerprints equal: mx8-devnet-epoch0 a6752e037514c91a, era-0 64c0ee90bac42624, ..., era-5 43673acc89954d5e). G2 method: one serve-mode job of 1,024 nonces at target ff..ff per card for era-0 and the pinned pack, every nonce a "found g2" line, re-hashed on the Mac with igneum-pow hash-bound --count 1024 (dry run through Apple OpenCL: 1,024 of 1,024 on both packs). Its PC 1 go follows the 5090 power sweep (ahead of the AMD sweep). C24 and C27 closed on branch bash-body-check (7adb1ca, 6805125); the integration merge takes its prover-socket-check.sh over proving-v1's and puts pc1-cpu-prove.ps1 on the allow list. From a6be364c87ba06faef6af9c5109c6fb0c5e99c82 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:16:15 +0000 Subject: [PATCH 107/131] Counter ASIC 2.0 status 22:16: the 5090 power sweep, the G2 job go --- docs/plans/counter-asic-2-status.md | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 7b3e3fb14..24912998c 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -502,3 +502,14 @@ Before / after, the era agent's way (one measure session, 22:07 UTC, readwidth e 22:13. Suite job 4 published: build-20261005-221237 (main d233fa1 = ca2-v3 with the mixer fix merged, V3_CLASS = { era: None, hot: None, ..MX8 }; fork 89dfcb95 with the engine test fixed; the five crates and the app). The node agent's checks, rebuild and gate run 3 on d233fa1 follow; the era agent's final-vectors and G2 package is being built on the final class; PC 1 is the AMD sweep's until about 22:25. 22:15. The era final-class package is ready (zip igneum-ca2-era-pc1b.zip sha256 cb0e9db07e304b11fd4c0591351af46090442ea4f51d60eb86945b96bd28aba3; workers f8d19f0a... and 87647c15... from 1ab8b21 plus the era branch; seven packs on the final class: mx8-devnet-epoch0 and era-0 to era-5, program id 73bcbfe8ccf988f1; Mac 7/7 on Metal and Apple OpenCL, fingerprints equal: mx8-devnet-epoch0 a6752e037514c91a, era-0 64c0ee90bac42624, ..., era-5 43673acc89954d5e). G2 method: one serve-mode job of 1,024 nonces at target ff..ff per card for era-0 and the pinned pack, every nonce a "found g2" line, re-hashed on the Mac with igneum-pow hash-bound --count 1024 (dry run through Apple OpenCL: 1,024 of 1,024 on both packs). Its PC 1 go follows the 5090 power sweep (ahead of the AMD sweep). C24 and C27 closed on branch bash-body-check (7adb1ca, 6805125); the integration merge takes its prover-socket-check.sh over proving-v1's and puts pc1-cpu-prove.ps1 on the allow list. + +## 22:16 the 5090 power-limit sweep (PC 1, relay #224, 22:09 to 22:15:37Z); the G2 job has the go + +| Cap | Limit W | Draw W | MH/s | MH/W | SM MHz | Memory MHz | Busy | +|---|---|---|---|---|---|---|---| +| 100% | 575 | 316.2 | 115.42 | 0.365 | 3,051 | 13,801 | 92.9% | +| 80% | 460 | 316.1 | 115.60 | 0.366 | 3,050 | 13,801 | 92.5% | +| 65% | 400 (the floor) | 310.6 | 114.46 | 0.369 | 3,051 | 13,801 | 90.1% | +| 50% | 400 (clamped) | 302.4 | 109.20 | 0.361 | 3,050 | 13,801 | 87.5% | + +The card draws 302 to 316 W under this program whatever the cap, so a cap above 400 W never binds; the readwidth, era, hot-table and mixer numbers taken at 431 W sit on the flat part of the curve (within 1% of stock); best per watt 65% (400 W) at 0.369 MH/W, a 0.8% hash cost. The cap was restored to 431 W and read back. (The app's 0.3.9 rate of 124 to 141 MH/s in the STATUS lines against 115 here: the API's hash_now sampled every 5 s under the sweep's own load; the bench rows of 136 to 137 MH/s are device time.) "go PC 1" given to the era agent's final-vectors and G2 job at 22:16; the AMD sweep follows it on a fresh probe. From 4fa1320a0aaba02bd996f5de847e89453d18ea2c Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:16:50 +0000 Subject: [PATCH 108/131] Counter ASIC 2.0: gate G6 green; status 22:16 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index b0403c541..96d0da579 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,7 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms | GREEN (runs 1 and 2) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 build-20261005-221237 (published 22:12:37Z: main d233fa1 = era + cache + the mixer fix with V3_CLASS = MX8, fork 89dfcb95) runs the five crates and the app on the final tree; G6 is green only when job 4 passes | +| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 build-20261005-221237 (published 22:12:37Z: main d233fa1 = era + cache + the mixer fix with V3_CLASS = MX8, fork 89dfcb95) ran the five crates and the app on the final tree: every stage ok (22:12:37 to 22:15:44Z): kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8, 0 failed. With job 2 (kaspa-consensus alone 97) G6 is GREEN on the final tree | GREEN | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 24912998c..62e53508b 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -513,3 +513,7 @@ Before / after, the era agent's way (one measure session, 22:07 UTC, readwidth e | 50% | 400 (clamped) | 302.4 | 109.20 | 0.361 | 3,050 | 13,801 | 87.5% | The card draws 302 to 316 W under this program whatever the cap, so a cap above 400 W never binds; the readwidth, era, hot-table and mixer numbers taken at 431 W sit on the flat part of the curve (within 1% of stock); best per watt 65% (400 W) at 0.369 MH/W, a 0.8% hash cost. The cap was restored to 431 W and read back. (The app's 0.3.9 rate of 124 to 141 MH/s in the STATUS lines against 115 here: the API's hash_now sampled every 5 s under the sweep's own load; the bench rows of 136 to 137 MH/s are device time.) "go PC 1" given to the era agent's final-vectors and G2 job at 22:16; the AMD sweep follows it on a fresh probe. + +## 22:16 gate G6 GREEN on the final tree + +build-20261005-221237 (main d233fa1, fork 89dfcb95, 185 s): every stage ok; kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test, with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8; 0 failed; with job 2's kaspa-consensus 97 alone, G6 is green. PC 2 goes to the prover-floor agent's 90-minute build (its go at 22:17), then the aggregation-cost re-run, then the prover-floor sweep. Gates: G3 green (the Mac suites on the final class: 53 + 4 + 19 + 7 crate tests, the Metal fuzz, edge, stats and determinism runs, the scratch tests), G4 green (runs 1 and 2; run 3 on the final x8 + era class pending), G6 green; G1 and G2 pending the era agent's PC 1 job (running from 22:16); G5 (the Windows and Mac workers from the same commit) is the ship's build step on the merged tree. From 31814153eec6e968376cefc1c32903ae9133d79e Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:17:33 +0000 Subject: [PATCH 109/131] Spec 01: class v3's mixer x8 form in 1.8.5 with the measured costs, the class v3 vectors in 1.17, the cache note in 1.5 --- docs/spec/01-lottery-hash.md | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/docs/spec/01-lottery-hash.md b/docs/spec/01-lottery-hash.md index cb87e61ff..e63876400 100644 --- a/docs/spec/01-lottery-hash.md +++ b/docs/spec/01-lottery-hash.md @@ -207,7 +207,7 @@ Implemented for the prototype size; Designed for genesis. A load reads one 4-byte word at `src AND MASK`. Every load in every emitted kernel has exactly this form; the static check in `TESTS.md` section 5 is part of conformance (section 1.15). Because item values do not depend on the dataset size (section 1.8.5), the 1 GiB vectors remain valid for words below 2^28 at any larger size. -Growth beyond genesis is in section 1.13. +Growth beyond genesis is in section 1.13; under program class v3 the cache follows the dataset's doublings (1.13.3) and the verifier holds 256 MiB, then 512 MiB from year 4 and 1 GiB from year 12. ## 1.6 Register initialisation @@ -319,7 +319,21 @@ s = M_8(s) item(t) = s ``` -Eight dependent cache reads (`ITEM_ROUNDS = 8`, prototype value): the address of read `r` depends on every earlier read. Nine mixer applications. `dataset[w] = item(w >> 4)[w AND 15]`. A dataset of 2^D words is the prefix of items `0 .. 2^(D-4) - 1`, so an item has the same value at every dataset size. +Eight dependent cache reads (`ITEM_ROUNDS = 8`, prototype value): the address of read `r` depends on every earlier read. Nine mixer applications under program class v2. + +Program class v3 (Counter ASIC 2.0, decided 5 October 2026, delegated; the project lead confirms for the public testnet genesis) applies the mixer `m = 8` times per round with distinct round keys (`LoadClass::mixer_mult`; `docs/plans/mixer-x4.md` section 2), the eight dependent reads unchanged: + +``` +for r in 0..7: + for j in 0..m-1: + s = M(s, rk = (r * m + j + 1) * 0x9E3779B9) + a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache (C from 1.13.3) + s[i] = s[i] XOR cache[line a][i] for i in 0..15 +for j in 0..m-1: + s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9) +``` + +Under `m = 1` the keys are `(r + 1) * 0x9E3779B9` and `9 * 0x9E3779B9`, the class v2 text exactly; the `9 m` keys are the first `9 m` values of the sequence `k * 0x9E3779B9`, all distinct. Why `m = 8`: the recompute attacker's cost is operations per item (`docs/analysis/m16-recompute-attacker-2026-10-05.md`); the honest miner pays the mixer once a day in the dataset build, which stays latency-bound (RTX 5090 23 to 25 ms, RX 9070 XT 72 to 77 ms, M5 Max 21 ms at `m` = 1, 4 and 8, measured 5 October 2026); the verifier pays `m` per item it derives: 0.61 ms per warp at `m = 1`, 1.24 at 4, 2.08 at 8 on one loaded M5 Max core (measured 5 October 2026, `docs/plans/mixer-x4.md` 6.4a), inside the 10 ms gate. The on-die-cache recompute chip's gain against the RTX 5090 falls from 2.4x (`m = 1`) to 0.92x with a 3x fixed-function factor at `m = 8` (`docs/analysis/chip-model-v3.md`, approximate factor). Under class v3 an item's value also depends on `C` through the line mask, so the items change on the day the cache doubles (1.13.3); the emitted `mh_item` carries the `m` loop only for `m > 1`, so every class v2 pack keeps its text. Class v3 also draws the dataset layout and the load windows per era (`docs/plans/era-layout.md`; the strided windowed load address and the interleaved mapping `mh_addr`, one text form in the three dialects). `dataset[w] = item(w >> 4)[w AND 15]`. A dataset of 2^D words is the prefix of items `0 .. 2^(D-4) - 1`, so an item has the same value at every dataset size. What the construction buys (Measured, `MEMHARD.md` section 2.2, M5 Max, seed igneum-genesis, 1 GiB): @@ -521,4 +535,6 @@ Pack `proto-cuda/packs/igneum-devnet-v4-epoch0/`: epoch seed bytes `edc4fa844da9 Header-bound vectors (section 1.6 rule, seed `igneum-genesis`, day `2026-10-03`): `igneum-pow/README.md`, eight values, for example H = 32 zero bytes and nonce 0 give `746c567b090acf6a`. +Program class v3 vectors (5 October 2026, `proto-cuda/packs-ca2-mixer/` and `proto-cuda/packs-ca2-era/`; generator 3, mixer x8, the cache growth rule, the era draw inside the class): pack `mx8-genesis` (seed `igneum-genesis`, no era, program id `e323b9dcaf283a6f`, batch fingerprint `7c28cfb06c5c65a9`) and pack `mx8-devnet-epoch0` (the devnet epoch seed and day of 1.17 with the era stand-in E_0 = the devnet genesis hash inside the class, program id `73bcbfe8ccf988f1`, unit 0 lane 0 `d424577fce4a7a60`, fingerprint `90f794dd556f7a3b` over the harness's 2^20 outputs and `a6752e037514c91a` over 2^16), reproduced by the Rust interpreter, Metal and Apple OpenCL on 5 October 2026 (3/3 standalone, 3/3 in batch, 96 of 96 lanes each); the six era packs `era-0` to `era-5` (the same epoch seed and day, era test seeds 0 to 5, program id `73bcbfe8ccf988f1`, fingerprints over 2^16 `64c0ee90bac42624`, `fb276b04bab43db2`, `51a15e86ce7afdad`, `2d4f94bae8356fd8`, `194ce26f9ebd508e`, `43673acc89954d5e`), 7/7 on Metal and Apple OpenCL; the RTX 5090 and RX 9070 XT runs of these exact packs and the 1,024-hash CPU re-check per card are the job `run-ca2-era-pc1b-20261005` of `docs/bench-log.md` (5 October 2026, night). The class v2 vectors above stand unchanged (the v2 exports are byte-identical on the class v3 crate, `igneum-pow/tests/packs.rs`). + Cache, dataset and mixer vectors: section 1.8.4 and 1.8.5. Seed words: section 1.3.1. Generator: section 1.4.3. Acceptance: section 1.4.6. Batch fingerprints: section 1.15 item 5. From 5959651452ac6905d593d55f22c5524074da0732 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:17:59 +0000 Subject: [PATCH 110/131] Spec 01 1.13.1: the class v3 era draw (stride, interleave, windows) as decided, with the devnet stand-in and the measured spread --- docs/spec/01-lottery-hash.md | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/docs/spec/01-lottery-hash.md b/docs/spec/01-lottery-hash.md index e63876400..6ee2a5414 100644 --- a/docs/spec/01-lottery-hash.md +++ b/docs/spec/01-lottery-hash.md @@ -441,7 +441,14 @@ The era seed `E_n` is the 32-byte output of the 1-hour VDF of section 4.4. One S `epoch_len` is the one era-table parameter set by miners rather than by the draw: 90% of blue blocks over a 7-day window carrying the same ladder index (3 bits of the header version, encoding Open in section 5.8) sets that length from the first day boundary at least 2 days after the window closes (section 5.7). It is not a code upgrade: the rule, the ladder and the window are genesis constants, and the chain carries no release. The era stream consumes its draw so that a future draw of this parameter changes no other parameter's value. The threat it answers is a per-program hard datapath (an FPGA fleet: 42 to 160 minutes per compile on a mid-size part, PRflow, FPT 2019, hours on large parts; at 600 s nothing it compiles ever runs); it does not answer a programmable chip, which the other layers answer. The floor 600 is set by the slowest compile-ahead measured (the Metal variant race, 38 s on the M5 Max, 6.3% of a 600-s epoch and inside the 600-s seed window; `docs/plans/epoch-length.md` section 6). -The table layout and the working-set window (Counter ASIC 2.0 layers 4 and 8) are drawn by the same stream; their rows and the draw order are in `docs/plans/era-layout.md` and enter this table with class v3's vectors (pending its six-era measurement, 5 October 2026). +The table layout and the working-set window (Counter ASIC 2.0 layers 4 and 8, decided IN on 5 October 2026, delegated: the six-era hash-rate spread is 1.3% on the RTX 5090, 3.2% on the RX 9070 XT and 0.8% on the M5 Max, under the 5% rule; `docs/plans/era-layout.md`) are drawn under program class v3 by a second stream `S` seeded with words 0 and 1 of `seed_words_from_bytes("igneum-era/" || E_n)` (the index is not in the preimage: `E_n` commits to `n` through the VDF input), seven draws in this order whether or not a value is used: + +1. `W = allowed[below(|allowed|)]`: the width in words of every dataset load of the era, from the genesis-fixed set `allowed`; the set is `{1}` (4 bytes, the read-width decision of 5 October 2026), so the draw is consumed and the width pinned. +2. `M = low32(next()) OR 1`: the stride multiplier, odd, so `x -> x * M` is a bijection. +3. `R = 1 + below(31)`: the stride rotation. +4. to 7. `r_i = next()` for `i` in 0..3: the interleave draws. With `b = log2(W)` and `free = 4 - b`, `c = [b, ..., 15]`; for `i` in `0..free`: `j = i + (r_i mod (16 - b - i))`, swap `c[i]` and `c[j]`; the interleave is `pos = [0, ..., b - 1] ++ sort(c[0..free])`, four ascending bit positions below 16. + +The era parameters are `(W, M, R, pos)`. Dataset mapping under class v3: word `w` holds word `j(w)` of item `t(w)`, where bit `i` of `j(w)` is bit `pos[i]` of `w` and `t(w)` is `w` with bits `pos[0..3]` removed; with `pos = [0, 1, 2, 3]` this is `dataset[w] = item(w >> 4)[w AND 15]` byte for byte; an item keeps its value at every dataset size of at least 2^16 words, and the `W` words of one aligned load lie in one item, so the 4,096-item verifier bound of 1.11 holds. Load address under class v3, for a load site with window draws `(k_off, o)` and a dataset of `2^D` words: `k = min(k_off, D - 26)`, `y = rotl(x * M, R)`, `idx = ((y AND (MASK >> k)) OR ((o AND (2^k - 1)) << (D - k))) AND MASK` (uniform on the window, branch-free, three operations before the mask), one text form in Metal, CUDA and OpenCL. The window draws per instruction (layer 8), after the nine draws of 1.4.3: `k_off = below(3)` (the dataset, a half or a quarter) and `o = low32(next()) AND (2^k_off - 1)`, used only on a load slot, so a class v3 program takes 720 draws; the window never goes below 2^26 words (256 MiB, above the largest on-chip cache in the benchmark) nor above the dataset, and sixteen sites with drawn offsets cover the dataset with high probability (a windows-union census over 300 programs: the SRAM mirror a chip would need is the whole dataset in every hour). The acceptance rule of 1.4.6 is unchanged in its tests and mirrors this address at its constant `D = 28`. Devnet stand-in for `E_n` until the VDF of 4.4 is in the node: era 0 the genesis block hash; era `n >= 1` the hash of the last selected-chain block whose DAA score is below `15,552,000 n - 7,200`. What the interleave buys and does not: a chip that hard-wires one layout reads the wrong 15 words with every word once the era draws another; a chip whose address decoder can permute its address lines pays nothing (stated in the plan). The stride is a bijection with no cryptanalysis yet (Open). "Memory pattern" in the design document is read here as the item-address pattern (the cache line index word, `s[0]` in 1.8.5, and the XOR-all-sixteen rule); the proposal is to leave it fixed at era 0 and let the unlocked families change the kernel instead, because every change to the item derivation changes the verify time and must be re-measured. From 9488399fcc4bb6030e9f4857cbce0dd42dafd107 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:18:39 +0000 Subject: [PATCH 111/131] Site and litepaper: level 1 on the hero and abstract, the chip bullet on the measured row; the level 3 table filled; status 22:18 --- docs/plans/counter-asic-2-public.md | 18 ++++++++++++------ docs/plans/counter-asic-2-status.md | 4 ++++ site/index.html | 2 +- site/litepaper.html | 4 ++-- 4 files changed, 19 insertions(+), 9 deletions(-) diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index f215b3ff5..1c76354d9 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -28,13 +28,19 @@ Headline of the chip model (5 October 2026, night): the strongest chip holds the Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per hash, the latency-bound share (rate over the card's random-read ceiling per load), the CPU verifier per warp, with machine, date and command. The chip model before and after Counter ASIC 2.0 (the m16 model's gain arithmetic at the v2 class and at the v3 class, with the SRAM a mirror needs, cited or approximate as the analysis says). The bounty terms (spec O-1.17: the leaderboard by card model, the standing bounty for any chip design beating a GPU by more than 2x, January 2027). Here the layers are named next to their numbers: read width, per-program mix, scratch, era layout, working set, hot table, cache schedule, the reserved integer-matrix family. -| Card | v2 MH/s | v3 MH/s | v3 bytes per hash | Latency-bound share v3 | Verifier ms per warp v3 | -|---|---|---|---|---|---| -| Apple M5 Max (Metal) | 27.74 | [owed: the v3 pack] | 512 + hot | [owed] | [owed: ca2-mixer] | -| RTX 5090 (CUDA) | 136.1 | [owed] | 512 + hot | 0.96 at v2 | [owed] | -| RX 9070 XT (OpenCL, eGPU) | 18.15 | [owed] | 512 + hot | 0.87 at v2 | [owed] | +| Card | v2 MH/s | v3 MH/s (era packs, six eras) | Bytes per hash | Latency-bound share | Verifier ms per warp (v2 / v3, one loaded M5 Max core) | Daily 1 GiB build (v2 / v3) | +|---|---|---|---|---|---|---| +| Apple M5 Max (Metal) | 27.68 | 28.35 to 28.58 (spread 0.8%) | 512 | 1.06 | 0.61 / 2.08 | 21 / 21 ms | +| RTX 5090 (CUDA) | 137.2 | 136.18 to 138.01 (spread 1.3%) | 512 | 1.01 | the same verifier | 25 / 23 ms | +| RX 9070 XT (OpenCL) | 18.09 | 18.61 to 19.21 (spread 3.2%) | 512 | 0.95 | the same verifier | 74 / 75 ms | -The integrated tier (Radeon iGPU, Intel UHD) on the CUDA and OpenCL one-click workers mines v3 with a restart per hourly epoch, because those workers rebuild the dataset on every prepare (7 to 12 s at x1, about 4x at the x4 mixer; the Metal worker keeps its day's dataset); per-day dataset reuse in those two workers is the first item after the publish (0.3.12). AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), 2.2x worse per pound at list prices and 4.9x worse per watt (read-width.md section 4.1, approximate); the card's memory system, not a tuning gap. +Every number measured 5 October 2026 (`docs/plans/era-layout.md`, `docs/plans/mixer-x4.md`, `docs/bench-log.md`); the v3 verifier figure is the x8 mixer on a loaded core (about 1.3 ms quiet, approximate). Bit-exact: every v3 pack's fingerprint equal on the three vendors. + +Chip model, before and after (`docs/analysis/chip-model-v3.md`): the on-die-cache recompute chip (the whole 256 MiB cache in SRAM, about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) against the RTX 5090's measured 136.1 MH/s at 50 T integer op/s: class v2 333 MH/s, 2.4x; class v3 (mixer x8) 41.7 MH/s, 0.31x bare, 0.92x with a 3x fixed-function allowance (approximate), 0.76x at equal silicon. The claim "under 2x" holds with the margin stated: 8% on the allowance (a 3.3x allowance reads 1.0x), 9% on the budget. Next levers, named: the mixer at x16 (the verifier at about 4 ms per warp, inside the 10 ms gate; a 2019-class core unmeasured), a hot table small enough to stay resident beside the streaming dataset (measured and not adopted tonight: 32 to 96 MiB tables cost the GPU 7 to 20% and help the chip). + +Levers measured and not adopted (5 October 2026): wider reads (16 and 64 B: no card gains, the 5090 goes bandwidth-bound at 64 B), the per-load width mix (5.5 to 22.3% spread), a per-warp write scratch (the chip keeps it implicitly: 2.4x at every share), the hot table (above). Reserved, switched off: the integer matrix family R1 (mm8, native on all three vendors as a tile; dp4a 1.17x a step on the 5090, 1.06x on the 9070 XT, emulation 1.6x on Apple) and the epoch length (600 s to 2 hours by 90% signal; a per-program FPGA bitstream mines 0% of a 600-s epoch). + +AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), 2.2x worse per pound at list prices and 4.9x worse per watt (read-width.md section 4.1, approximate); the card's memory system, not a tuning gap. Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant]. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 62e53508b..a29201fd0 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -517,3 +517,7 @@ The card draws 302 to 316 W under this program whatever the cap, so a cap above ## 22:16 gate G6 GREEN on the final tree build-20261005-221237 (main d233fa1, fork 89dfcb95, 185 s): every stage ok; kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test, with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8; 0 failed; with job 2's kaspa-consensus 97 alone, G6 is green. PC 2 goes to the prover-floor agent's 90-minute build (its go at 22:17), then the aggregation-cost re-run, then the prover-floor sweep. Gates: G3 green (the Mac suites on the final class: 53 + 4 + 19 + 7 crate tests, the Metal fuzz, edge, stats and determinism runs, the scratch tests), G4 green (runs 1 and 2; run 3 on the final x8 + era class pending), G6 green; G1 and G2 pending the era agent's PC 1 job (running from 22:16); G5 (the Windows and Mac workers from the same commit) is the ship's build step on the merged tree. + +## 22:18 the spec and the public copy carry the final class + +docs/spec/01-lottery-hash.md on ca2-coord: 1.8.5 (the mixer x8 form with the measured costs), 1.13.1 (the era draw: stride, interleave, windows, the devnet stand-in, the measured spread), 1.17 (the class v3 vectors), 1.5 (the cache note); earlier tonight 1.12 and 1.13.1 (epoch_len), 1.13.2 (R1 and the emulation rule), 1.13.3 (option C and the step mapping), 1.4.5 and 1.4.6 (generator 3, the class and era in the pack), and 4.3. Public copy: level 1 on the hero and the abstract ("Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public."), the limits bullet rewritten on the chip row (0.92x with the allowance, approximate; the margin on the numbers page; the next lever named), the level 3 table filled from the final numbers in counter-asic-2-public.md (the bench page section is written from it at the ship). Waiting: the era agent's PC 1 job (G1 on the final class and G2), the node agent's gate run 3; then the integration merge and the ship. diff --git a/site/index.html b/site/index.html index 4f11aa90f..60976e26f 100644 --- a/site/index.html +++ b/site/index.html @@ -340,7 +340,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var(
GPUs are back · for good

Mined by GPUs.
Proven by fire.

-

A chain built so a chip gains too little to take your place. A mining program that rewrites itself every hour, so a chip built for one hour is useless the next. An NVIDIA card with 24 GB or more proves every block and gets paid for it; AMD and Apple cards mine. No premine, no stake, no foundation, no merge to proof of stake.

+

Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public. A mining program that rewrites itself every hour, so a chip built for one hour is useless the next. An NVIDIA card with 24 GB or more proves every block and gets paid for it; AMD and Apple cards mine. No premine, no stake, no foundation, no merge to proof of stake.

See the miner Read the litepaper diff --git a/site/litepaper.html b/site/litepaper.html index 37ed6bb39..bc7e21beb 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -298,7 +298,7 @@ body.all .pager{display:none}

Abstract

-

Igneum is a proof-of-work blockchain mined on graphics cards, where the same cards prove every block with zero-knowledge proofs and sell proving to other chains.

+

Igneum is a proof-of-work blockchain built for graphics cards, where NVIDIA cards also prove every block with zero-knowledge proofs and sell proving to other chains. A custom chip gains under 2x, and the model and the bounty are public.

It runs the Ethereum virtual machine, so anything built for Ethereum runs on Igneum unchanged. Transactions are included in about one second, proven within about a minute at launch, and locked by miners within about two. There is no premine, no pre-sale, no treasury taken from emission, no stake anywhere in consensus, and no dependence on any other chain. Mining stays open to anyone with a GPU because the mining program changes every hour, so a chip built for one program is useless for the next, and a chip for the whole program space is a GPU without the graphics parts. No scheduled human release is needed to keep it that way. Writing new code, including an emergency fix to the proof system, is the one thing that takes a person, and it activates only on miner signalling.

1 / s
blocks, rising to 10
@@ -730,7 +730,7 @@ body.all .pager{display:none}

Here are the limits, stated before anyone else states them.

  • A proof in seconds. Not at launch. Proving a full block today needs a cluster of 100 to 200 consumer GPUs, approximate, so Igneum launches with proofs within about a minute and tightens as hardware improves. Users still see their transaction land in one second.
  • -
  • A chip is impossible. No. A chip is a bad bet, because the target moves before it ships. The efficiency ceiling for a fixed chip on a memory-bound program is a target of under 2x, not a measurement, and Igneum's generator changes under it every hour. Monero's seven years without a public chip are precedent, not proof.
  • +
  • A chip is impossible. No. A chip is a bad bet, because the target moves before it ships. The published model (5 October 2026) prices the strongest chip we can name, one with the whole cache on-die computing dataset items on the fly, at 0.92x the hash rate of an RTX 5090 per unit of silicon with a 3x fixed-function allowance, approximate; the claim is under 2x, the margin is stated on the numbers page, and the next lever is named there. No hash has stayed free of chips forever; Igneum does not claim to. Monero's seven years without a public chip are precedent, not proof.
  • A guaranteed income floor. No. External proving is a small market today. Igneum's miners' marginal cost in it is close to power, which is an edge and nothing more.
  • A memory-hard prototype on every vendor. Not yet. The 256 MB cache closed the shortcut on Apple silicon (computing items runs 4.8x slower than loading them, measured 3 October 2026). The same ratio on NVIDIA and on a discrete AMD card is Open.
  • Finality in the first month. No. No checkpoint locks until the 30-day window has 30 days of history. The first month of mainnet is proof of work with a 12-hour depth, and the text above says so wherever a day count appears.
  • From 48cb0a191dc53046e30eafa599d0e612e7065452 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:19:12 +0000 Subject: [PATCH 112/131] evidence.md rows 17, 18, 19 on the measured class v3: the chip model, the latency bound, the soundness and bit-exactness --- docs/evidence.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/evidence.md b/docs/evidence.md index a00304a06..93a3c2d5b 100644 --- a/docs/evidence.md +++ b/docs/evidence.md @@ -41,9 +41,9 @@ Versions in the table: `igneum-pow` is the Rust crate at `igneum-pow/Cargo.toml` | 14 | Ethereum bytecode runs unchanged, with the documented differences of spec 7.1 | Homepage Build card; litepaper Building | tested by the team | as row 13; fixes `F-exec-A`, `F-exec-B` (spec 7.5) | `tools/evm-smoke/smoke.mjs`: deploy via viem, `increment`, `hashLoop`, `eth_estimateGas`, `eth_getLogs`; `tools/exec-attacks` scenarios 1 and 3; bench-log "execution layer attack fixes" | Deployment, calls, reverts, logs and gas estimates behave as viem expects; chain id 4463; the prototype pgas table gives 0.0095 to 0.028 pgas per gas, below the design's band before calibration, 3 October 2026. 4 October 2026: a transaction that would cross the block's proving budget is refused by the mempool and, if forced in, aborted and charged with its nonce advanced (25 of 25 checks; 30 of 30 malformed cases). Apple M5 Max. The `Prover` precompile, proof records and the shard planner are not in the node | none yet | | 15 | Every block is proven, with the proof landing within about a minute at launch | Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate | implemented | repo `d7e1f89` (GPU proof), `e01a3cc`, `292e800`, `eedd136` (`proving/igneum-prove`: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6 | `proving/windows-wsl2` (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; `igneum-prove-host --mode block` on `proving/fixtures/`; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards" | First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture `block-78-increment` (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in `docs/benchmarks/proving-e2e.md`. Second RTX 5090 run, 4 October 2026 evening (job run-20261004-173115): a full shard at the provisional S_p (6.75 M pgas, 60.8 M cycles) executed in 1.63 s, core proof 8.3 s (18.1 MB), compressed proof 10.9 s (1.27 MB, verify 0.040 s); a two-shard block (13.5 M pgas) proved shard by shard (11.7 s and 10.0 s) and aggregated in 2.2 s, 24 s of GPU stages end to end, every proof verified, six tampered witnesses rejected. The two host defects (an abort after the upload, an idle wait that turned out to be an unbuffered 18 MB proof save through the WSL2 file bridge, 24 minutes) are fixed (ledger P20) 5 October 2026, live devnet with real transactions (bench-log "real transactions, the first non-empty shard proven and paid"): block 72704 shard 0, 29 transfers, 5,800 pgas, proven on PC 2 in 34 s, verified on the Mac in 0.297 s and paid 1.7623 IGN, 53 s after the chain block executed; of about 1,400 blocks in the 20-minute window 36 were proven (the one prover takes the newest shard assigned to it), so "every block" is not yet true; a second content shard (72803, all copies skipped) failed the native-execution veto on the exporter's block structure, fixed with fixtures the same day, the node side pending the 0.3.9 rollout | none yet | | 16 | A 12 GB card proves one shard in about 20 s | Litepaper Proving ("The proving budget"); roadmap gate 2 | designed | spec 5.1 (Target), 7.6 (`S_p` provisional, 7,500,000 pgas = `B_p` / 4) | `PROVE-SHARD.bat` on the RTX 5090 (pending); the end-to-end standard in `docs/benchmarks/proving-e2e.md`; bench-log "proving: devnet v4 shards" | Measured on a 32 GB card, not yet on a 12 GB card. A shard at the provisional `S_p` is 60.8 M SP1 cycles on the prototype pgas table (9 cycles per pgas, 44 per EVM gas; the modexp entry about 100x its SP1 cost); on an RTX 5090 (4 October 2026 evening, job run-20261004-173115) it executed in 1.63 s and its compressed proof took 10.9 s, verified in 0.040 s, so the 32 GB card is inside the 20 s target with margin. Whether a 12 GB card proves it at all, and in what time, is the next measurement (an RTX 3060 and an RTX 5060 Ti 16 GB are on order). A per-shard time can be met by shrinking the shard, so the project does not use it as a pass mark | none yet | -| 17 | The chip resistance target: a chip gains under 2x over a GPU | Litepaper Mining, "What Igneum does not claim"; homepage "no chip can be built for it" | designed | spec 0.2 (Target); O-1.17 | Public benchmark with a leaderboard by card model and a standing bounty, January 2027 (O-1.17); the on-die-SRAM test on the RTX 5090 (R3.5) | A target, not a measurement. Review round 3 priced a recompute chip with the 256 MiB cache on die at about 2.4x, approximate, before the usual chip-versus-GPU integer gain; the design answer (cache larger than any die) is open (spec 1.16) | none yet | -| 18 | The chip resistance measurements: the program is random-access bound, not bandwidth bound, and sits beyond a card's on-chip cache | Litepaper Mining ("bound by memory bandwidth", to be corrected), vs RandomX "Measured so far" | tested by the team | repo `aba248d`, `f2a1a64`, `4b95c5e` | RTX 5090 dataset sweep 4 MiB to 1 GiB with `proto-cuda/host.cu`; bench-log "RTX 5090 first run" and "dataset sweep" | At 1 GiB: 228.1 Mhash/s, 23.7 G random loads/s, 94.9 GB/s useful against a 1,638 GB/s dataset fill; inside the 96 MiB L2 (4 and 64 MiB) 1,340 to 1,353 Mhash/s, about 5.8x faster; 104 against 128 loads per hash gives 228 against 185 Mhash/s, proportional. 3 October 2026, RTX 5090, Windows, CUDA 12.8, version 1 programs. Prototype dataset 1 GiB against 2 GB at genesis; a pure random-read microbenchmark (R3 chip designer, attack 2) has not run; the sweep has not been repeated on version 2 | none yet | -| 19 | The lottery hash is sound as a hash: uniform output, deterministic, no out-of-bounds read, fuzzed | Litepaper vs RandomX ("Every number above is measured and logged") | tested by the team | repo `c52307e`, `58a5a63`, `b27da39`; `proto-metal/TESTS.md` | `proto-metal/igneum-bench --fuzz --edge --stats --determinism --memcheck`; `--fuzz 2000` on the version 2 generator; `igneum-census`; bench-log "hardening tests", the re-run on the memory-hard dataset, "generator version 2 adopted" | Version 1: 10,200 random programs, 1,305,600 hashes, 0 mismatches; 14 of 14 edge cases; bit frequency within 2.90 sigma, avalanche mean 31.99 to 32.04 of 32; deterministic fingerprint across 5 runs; every dataset read masked, 3 October 2026. Version 2, 4 October 2026: 2,000 random programs through the Metal cross-check, 8,000 warps, 0 mismatches, 128 loads per hash on every program; 20,000-program census, 5.2% rejected (4.1% static, 1.1% dynamic). Apple M5 Max. Statistics are not a security proof; the edge, stats and memcheck sections were not re-run on version 2 (they do not depend on the generator); the seed derivation review (O-1.4) is open; the fuzz set has run on Metal and the CPU only | none yet | +| 17 | The chip resistance target: a chip gains under 2x over a GPU | Homepage hero and litepaper abstract ("a custom chip gains under 2x, and the model and the bounty are public"), litepaper "What Igneum does not claim" | tested by the team (the model), designed (the target) | program class v3 (Counter ASIC 2.0, 5 October 2026): branches ca2-v3 d233fa1 and after, ca2-mixer 1ab8b21, ca2-era 78c0ee4; `docs/analysis/chip-model-v3.md`, `docs/analysis/sram-mirror.md`, `docs/analysis/scratch-soundness.md` | The m16 recompute model re-run on the measured v3 rates and verifier times; the on-die-cache chip row | The on-die-cache recompute chip against the RTX 5090's measured 136.1 MH/s: class v2 2.4x; class v3 (mixer x8) 0.31x bare, 0.92x with a 3x fixed-function allowance (approximate), 0.76x at equal silicon; margin 8% on the allowance, 9% on the budget. 5 October 2026, M5 Max, RTX 5090, RX 9070 XT. The 2x target is a target: no chip has been built; the bounty stands (O-1.17) | none yet | +| 18 | The chip resistance measurements: the program is latency-bound (random reads), not bandwidth-bound, on every card we own, and sits beyond a card's on-chip cache | Litepaper Mining ("waits on memory latency, not on maths or bandwidth"), vs RandomX; the numbers page | tested by the team | readwidth e752fc7 (`docs/plans/read-width.md`), ca2-era 78c0ee4, ca2-cache 2de19e5 (`docs/plans/hot-table.md`) | The dependent-read probes at 32 to 1,024 MiB and the hash rate per class on the three cards; the latency-bound share = rate over the probe ceiling per load | Latency-bound share at the 1 GiB dataset: RTX 5090 0.96 (v2) and 1.01 (v3), RX 9070 XT 0.87 and 0.95, M5 Max 1.01 and 1.06; wider reads do not close the AMD gap (the 9070 XT does 2.4 G dependent reads per second at every width; the 5090 goes bandwidth-bound at 64 B, share 0.58); a 32 to 96 MiB hot table is not kept resident by any card while the dataset streams (g 0.80 to 0.87 in the added form). 5 October 2026 | none yet | +| 19 | The lottery hash is sound as a hash: uniform output, deterministic, no out-of-bounds read, fuzzed; class v3 bit-exact on the three vendors | Litepaper vs RandomX ("Every number above is measured and logged"), the numbers page | tested by the team | ca2-mixer 1ab8b21 (`tests/mixer.rs`, `tests/scratch.rs`), ca2-era 78c0ee4, ca2-soundness a465881 (`docs/analysis/scratch-soundness.md`), `igneum-pow/tests/packs.rs` | The crate suite (53 + 4 + 19 + 7), the Metal fuzz, edge, stats and determinism runs on the v3 construction, the pack vectors and 2^24 fingerprints on Metal, Apple OpenCL, the RTX 5090 and the RX 9070 XT, the 1,024-hash CPU re-check per card | Class v3 (mixer x8 + era): 200-program fuzz 200 of 200 on Metal, every tenth on Apple OpenCL; the pinned v3 packs 3/3 + 3/3 and 96 of 96 lanes on Metal and Apple OpenCL; the six era packs' fingerprints equal on the three vendors (PC 1 job run-ca2-era-pc1-20261005, 5 October 2026); the v2 exports byte-identical on the v3 crate; the final-class PC rows and the G2 re-check: job run-ca2-era-pc1b-20261005 (pending at the time of writing) | none yet | | 20 | No premine, no pre-sale, no allocation: every coin is minted by the schedule and every coin goes to the block producer (80%) and the proving pool (20%) | Homepage stats and Economics tiles; litepaper Supply, Economics | implemented | repo `6ac80a3`; fork "igneum-node devnet v0"; `consensus/core/src/igneum.rs`, `coinbase.rs` | `cargo test -p kaspa-consensus-core igneum` (8 pass: subsidy table, ramp, split, cap) and `cargo test -p kaspa-consensus coinbase` (8 pass); `igneum-miner inspect 40`; bench-log "igneum-node devnet v0" | Coinbases on the devnet: 80/20 exact on 39 of 39 single-payee blocks, the 20% to the `igneum-proving-pool-v0` output; the per-second schedule sums to under the 4,000,000,000 cap by less than 100 coins; 3,168,808,781 units per DAA second in years 0 to 2, halving at 63,115,200 DAA s. 3 October 2026, Apple M5 Max. The devnet genesis carries no allocation; the mainnet genesis does not exist yet, so the claim is about the code and the stated rule, not a launch that has happened | none yet | | 21 | The proving pool's 20% reaches shard provers and aggregators | Litepaper Economics; homepage "20% provers" | tested by the team | spec 5.3; `proving/igneum-prove` carries the prover's payout address in every shard proof (ledger P12) | None. The pool output exists (row 20); the payout from it against proof records is unwritten. Since 5 October 2026: the payout rule is live on the devnet (`proving.rs shard_payouts`, the carrying segment pays the first valid record per shard its part of the segment's pool credit) | The escrow accumulated on the simnet (92.55 IGN at the end of the v3 run) and nothing can draw it. Rule decided: per block, divided among shards by consensus proving cost, sortition to 8 provers for 10 s then open (spec 7.2). The economy model of 4 October 2026 (`sim/economy`, 1,000 operators, 30 days) kept every block proven within 60 s under six stress scenarios; a model, not hardware Live devnet, 5 October 2026: 388 shards paid by 16:02 UTC, 446.13 IGN from the pool to PC 2's payout address, 0.8813 IGN per mergeset block of the proven segment (bench-log entries of 5 October: "the first shards proven, verified and paid" and "real transactions, the first non-empty shard proven and paid") | none yet | | 22 | The base fee is burned in full and the priority fee splits 80% to the miner and provers, 20% to the apps whose code ran | Homepage Economics caption and Build card; litepaper "Where fees go" | tested by the team | repo `f5f8c80`; fork worktree `vendor/igneum-node-exec` | `tools/evm-smoke/smoke.mjs` receipt checks; bench-log "execution layer devnet v3" | Transfer receipt: `burnedProvingFee` 200 gwei, `minerTip` 16,800 gwei (80%), unregistered developer share 4,200 gwei burned; contract call: 80% to the miner, 20% credited to the payee the constructor registered, balance delta equal. 3 October 2026, Apple M5 Max simnet. The provers' part of the 80% is not split out (no provers exist); the base fee stayed at the 1 gwei floor throughout | none yet | From f8df1b7caa95497ade15fab3dd8edbe40706d4f9 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:19:18 +0000 Subject: [PATCH 113/131] Counter ASIC 2.0 status 22:19: proving-methods recorded --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index a29201fd0..f08f00c63 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -521,3 +521,5 @@ build-20261005-221237 (main d233fa1, fork 89dfcb95, 185 s): every stage ok; kasp ## 22:18 the spec and the public copy carry the final class docs/spec/01-lottery-hash.md on ca2-coord: 1.8.5 (the mixer x8 form with the measured costs), 1.13.1 (the era draw: stride, interleave, windows, the devnet stand-in, the measured spread), 1.17 (the class v3 vectors), 1.5 (the cache note); earlier tonight 1.12 and 1.13.1 (epoch_len), 1.13.2 (R1 and the emulation rule), 1.13.3 (option C and the step mapping), 1.4.5 and 1.4.6 (generator 3, the class and era in the pack), and 4.3. Public copy: level 1 on the hero and the abstract ("Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public."), the limits bullet rewritten on the chip row (0.92x with the allowance, approximate; the margin on the numbers page; the next lever named), the level 3 table filled from the final numbers in counter-asic-2-public.md (the bench page section is written from it at the ship). Waiting: the era agent's PC 1 job (G1 on the final class and G2), the node agent's gate run 3; then the integration merge and the ship. + +22:19. docs/analysis/proving-methods.md (branch proving-methods e7e0db7, not consensus): top recommendation re-size SP1's own GPU server (the floor is its code: the under-20 GB panic at sp1-gpu builder.rs:37, trace buffers at the maximum shard, a CUDA mempool that never releases), S_p as the dial, one server per card on rigs; the pinned ids stay; fallback and the Apple route: RISC Zero as proof-system version 2 behind the ProofSystem seam (8 GB at po2 19, 16 GB at po2 20, shipped Metal; 3 to 4 agent days); no 12 GB card has run a prover here, so a 4070 or 3060 in the loop is the first action. That branch merges into the 0.3.11 main tree as documentation (no code). evidence.md rows 17, 18 and 19 are rewritten on the measured class v3 (ba8379d). From a59e782159d4fe1a5d8dd52fa3cc94d242d0f55d Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:20:05 +0000 Subject: [PATCH 114/131] Bench log: Counter ASIC 2.0, the numbers (the level 3 page section); litepaper anchor --- docs/bench-log.md | 32 ++++++++++++++++++++++++++++++++ site/litepaper.html | 2 +- 2 files changed, 33 insertions(+), 1 deletion(-) diff --git a/docs/bench-log.md b/docs/bench-log.md index 5b674a242..bc0cfe101 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1586,3 +1586,35 @@ Machine: PC 1 (machine ae432dc7), Windows 11, WSL2 Ubuntu 24.04 as root, 16 core For comparison (this log): the Apple M5 Max CPU on 4 October, loaded, block-56 shard 0: core 83.1 s, compressed 272.3 s; on 3 October the v0 guest on block-78: core 22.0 s, compressed 55.7 s. The RTX 5090: block-78 core 1.4 s, compressed 2.7 s (4 October, mining paused); a full shard at `S_p` compressed 10.9 s alone and 33.0 s beside the miner; an empty live shard 7.0 to 7.7 s beside the miner (5 October). No fresh Mac run tonight: the measure lock was held from 20:31Z (a read-width `packbench`, three builds, a 1,500-s proving-v1 network under `run`) and did not free inside the 10-minute window set for it. Reading, and the consequences (CLAUDE.md, every number). Doubling the cycles added 4.5 s to the core proof and 3.1 s to the compressed proof: about 280 s of a CPU proof is fixed cost in the compressed-proof recursion, so no shard size brings a CPU proof under the launch deadline (20 to 60 s behind the tip) or near the 10-s assignment window; it fits only the v1 unproven deadline (600 s), which pays a CPU prover only when no card has proven the shard in 10 minutes. The 29.5 to 30.5 GB peak RSS means the CPU prover needs 32 GB free: a 64 GB Windows PC (WSL2 takes half the host's RAM by default), a 32 GB Linux machine, a 64 GB Mac; a 16 GB machine cannot run it at all. Per tier: an AMD-only home miner (8, 12 or 16 GB, Windows or Linux) mines and does not prove, and loses the 20% proving-pool share; Apple silicon the same (the M5 Max mines at 26.7 MH/s, this log, 4 October); a mixed rig proves on its NVIDIA cards and the rig installer's `prover_decision` already skips every non-NVIDIA card (`packaging/linux/bin/igneum-rig-lib.sh`, branch `rig-install`), now a stated requirement; the app's `provedefault.rs` already keeps proving off on Apple silicon and off without an NVIDIA card. Decision asked of nobody: no CPU tier (the analysis, section 4a); the public line for the site, litepaper and Proving tile is in section 4c ("Proving needs an NVIDIA card with 16 GB or more today ... AMD and Apple cards mine. A prover for them lands when a zkVM ships one"). The first job proved nothing because an apostrophe inside a single-quoted awk program ended the quote and bash refused the loop while the job reported exit 0; the class fix is `tools/amd-prove/check-job-bash.sh` (`bash -n` on the embedded bash body before publishing) and the same `bash -n` inside the job before the run, both shown to refuse the bad body and pass the fixed one. + +## Counter ASIC 2.0, the numbers + +5 October 2026 (night). The chip-resistance layers measured on the three cards we own (Apple M5 Max, RTX 5090 on PC 1 and PC 2, RX 9070 XT on PC 1's eGPU), the decisions taken under the project lead's delegated rules for the devnet, and the chip model before and after. Every number is from an entry above or from the plan documents named; approximate is marked. Levels: `docs/plans/counter-asic-2-public.md`. + +**Program class v3 (the devnet, activation by height switch `program_class_v3_activation_daa`)** = class v2's 128 x 4-byte loads, the era draw of the table layout and the working-set windows (layers 4 and 8), the cache growth rule (layer 6, option C: the cache doubles when the dataset doubles), the mixer at x8 (M16's multiplier), reserve family R1 (integer matrix, switched off) and the epoch length as a signalled reserve parameter (layer 9, 3,600 DAA s until a 90% signal). Not adopted on the measurements: wider reads (layer 1), the per-load width mix (layer 2), the per-warp write scratch (layer 3), the hot table (layer 5). + +| Card | v2 MH/s | v3 MH/s, six eras (spread) | Bytes per hash | Latency-bound share | Daily 1 GiB build, v2 / v3 | +|---|---|---|---|---|---| +| Apple M5 Max, Metal | 27.68 | 28.35 to 28.58 (0.8%) | 512 | 1.06 | 21 / 21 ms | +| RTX 5090, CUDA | 137.2 | 136.18 to 138.01 (1.3%) | 512 | 1.01 | 25 / 23 ms | +| RX 9070 XT, OpenCL | 18.09 | 18.61 to 19.21 (3.2%) | 512 | 0.95 | 74 / 75 ms | + +CPU verifier, one M5 Max core (loaded box; ratios are the measurement): v2 0.61 ms per warp, v3 (x8) 2.08 ms, worst cold 2.15; the 10 ms gate holds 4.8x. Bit-exact: every v3 pack's fingerprint equal on Metal, Apple OpenCL, CUDA and AMD OpenCL. + +| Layer | Measured | Decision | The number | +|---|---|---|---| +| 1 wider reads | w16 139.8 / 17.90 / 28.26 MH/s (5090 / 9070 XT / M5 Max) against v2 136.1 / 18.15 / 27.74; w64 71.9 on the 5090 (share 0.58, 37% of its stream) | out: keep 4 B | the 9070 XT does 2.4 G dependent reads/s at every width; wider reads make the 5090 bandwidth-bound | +| 2 width mix per load | spread over six programs 18.8 / 7.4 / 11.3% and 22.3 / 5.5 / 8.1% | out | the 5% rule | +| 3 write scratch | GPU cost 12 to 48% at 32 and 128 KB per warp; the on-die-cache chip 2.4x at every share | out (the construct is sound; its tests stay) | the verifier resets the scratch per unit, so a chip keeps it in 80 to 320 B per lane | +| 4 and 8 era layout and windows | six-era spread 1.3 / 3.2 / 0.8% | in | under the 5% rule; the SRAM mirror a chip needs is the whole dataset every hour | +| 5 hot table | added form g 0.87 / 0.85 / 0.84 (5090), 0.84 / 0.81 / 0.80 (9070 XT) at 32 / 64 / 96 MiB | out (a 3.0 option) | no card keeps 32 MiB resident while the dataset streams; the replaced form helps the chip | +| 6 cache schedule | the 256 MiB mirror is 128 mm^2 and $46 at N5 by shipped cache-die density, approximate | option C, in | the cache's job is to stay above GPU L2 (96 MB on the 5090, 128 MB on GB202) | +| 7 integer matrix | dp4a 1.17x a step on the 5090, 1.06x on the 9070 XT, 1.6x emulated on Apple; mm8 native on all three as a tile | reserved R1, off | unlock at era 4 or 90% signal | +| M16 mixer | x4: verifier 1.24 ms, chip 1.84x with the allowance; x8: 2.08 ms, 0.92x; the daily build unmoved on every card | x8 in | the only lever that moves the named chip | +| 9 epoch length | compile-ahead 0.5 s (M5 Max, race off), 1.0 s (5090), 38 s with the race; FPGA compiles 42 to 160 min (PRflow, FPT 2019) | reserved, 600 s to 2 h by signal | at 600 s a per-program bitstream mines 0% of each epoch | + +**The chip model, before and after** (`docs/analysis/chip-model-v3.md`, `docs/analysis/sram-mirror.md`): the strongest chip we can name holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5, approximate) and computes dataset items on the fly at 50 T integer op/s. Against the RTX 5090's measured 136.1 MH/s: class v2 333 MH/s, 2.4x; class v3 41.7 MH/s, 0.31x bare, 0.92x with a 3x fixed-function allowance (approximate), 0.76x at equal silicon. The claim is "under 2x"; the margin is thin on the allowance (3.3x reads 1.0x) and 9% on the budget. Next levers, named: the mixer at x16 (the verifier at about 4 ms per warp; a 2019-class core unmeasured), a hot table small enough to stay resident beside the streaming dataset. + +**The user tiers.** AMD RDNA 4 sits at about a seventh of a 5090 on this hash (its dependent-read rate: 2.4 G against 17.5 G per second), 2.2x worse per pound at list prices and 4.9x worse per watt (approximate); the card's memory system, not a tuning gap. The integrated tier on the CUDA and OpenCL one-click workers mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). Card lifetime under the step schedule: a 4 GB card to year 4, 8 GB to year 12, 12 GB to year 28 with the cache freed after the daily build. + +**The bounty.** A standing bounty for any chip design beating a GPU by more than 2x on the published model, with a leaderboard by card model, January 2027 (spec O-1.17). diff --git a/site/litepaper.html b/site/litepaper.html index bc7e21beb..73aaa2049 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -421,7 +421,7 @@ body.all .pager{display:none}

Three ideas carry the chip resistance. The hash rewrites itself. A new program every hour, drawn from the chain. Its memory pattern changes with it. The rules change on a schedule fixed at launch. No release, no vote. It waits on memory, not maths. Every hash is a chain of random reads into a table too big for a chip to carry. The wait is the same physics for everyone. Miners hold the switch. Spare defences are written into the rules, switched off. A 90% miner signal turns one on. No fork.

-

No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured. The model and the bounty are public: the numbers. Monero has run on RandomX since 2019 with no chip publicly shipped, approximate; that is precedent, not proof.

+

No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured. The model and the bounty are public: the numbers. Monero has run on RandomX since 2019 with no chip publicly shipped, approximate; that is precedent, not proof.

One thing takes a person, here and on every chain that exists: writing new code. A chain cannot safely write its own generator, and it cannot safely tell a chip from a wave of honest new cards by hashrate alone. If the design above ever failed, anyone could publish a new generator and miners would switch it on by signalling, as Monero's community can fork. Igneum is built to make that day unlikely, and does not depend on avoiding it.

From 84aac1baed9d4f8e24be334bd95379613743cd32 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:20:18 +0000 Subject: [PATCH 115/131] Litepaper: the verifier figures on class v3; status 22:20 --- docs/plans/counter-asic-2-status.md | 2 ++ site/litepaper.html | 4 ++-- 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index f08f00c63..dfc6fb525 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -523,3 +523,5 @@ build-20261005-221237 (main d233fa1, fork 89dfcb95, 185 s): every stage ok; kasp docs/spec/01-lottery-hash.md on ca2-coord: 1.8.5 (the mixer x8 form with the measured costs), 1.13.1 (the era draw: stride, interleave, windows, the devnet stand-in, the measured spread), 1.17 (the class v3 vectors), 1.5 (the cache note); earlier tonight 1.12 and 1.13.1 (epoch_len), 1.13.2 (R1 and the emulation rule), 1.13.3 (option C and the step mapping), 1.4.5 and 1.4.6 (generator 3, the class and era in the pack), and 4.3. Public copy: level 1 on the hero and the abstract ("Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public."), the limits bullet rewritten on the chip row (0.92x with the allowance, approximate; the margin on the numbers page; the next lever named), the level 3 table filled from the final numbers in counter-asic-2-public.md (the bench page section is written from it at the ship). Waiting: the era agent's PC 1 job (G1 on the final class and G2), the node agent's gate run 3; then the integration merge and the ship. 22:19. docs/analysis/proving-methods.md (branch proving-methods e7e0db7, not consensus): top recommendation re-size SP1's own GPU server (the floor is its code: the under-20 GB panic at sp1-gpu builder.rs:37, trace buffers at the maximum shard, a CUDA mempool that never releases), S_p as the dial, one server per card on rigs; the pinned ids stay; fallback and the Apple route: RISC Zero as proof-system version 2 behind the ProofSystem seam (8 GB at po2 19, 16 GB at po2 20, shipped Metal; 3 to 4 agent days); no 12 GB card has run a prover here, so a 4070 or 3060 in the loop is the first action. That branch merges into the 0.3.11 main tree as documentation (no code). evidence.md rows 17, 18 and 19 are rewritten on the measured class v3 (ba8379d). + +22:20. The numbers page's Counter ASIC section is written as the bench-log entry "Counter ASIC 2.0, the numbers" (the page is built from the bench log; the litepaper links /bench#counter-asic-2-0-the-numbers); the litepaper's verifier figures moved to the v3 class (2.1 ms per warp on a loaded core, 4.8x inside the gate; the cache 512 MB from year 4). Everything public now carries the final class except the PC rows of the final-class packs and the G2 counts, which the running PC 1 job supplies. diff --git a/site/litepaper.html b/site/litepaper.html index 73aaa2049..3c6bce685 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -408,7 +408,7 @@ body.all .pager{display:none}

Mining: a program that never holds still

Every GPU chain that promised ASIC resistance shipped a fixed algorithm, and a fixed algorithm gets a chip the moment the prize pays for one. Igneum does not have a fixed algorithm.

-

Each hour the chain derives a seed from a locked checkpoint one epoch back, passes it through a ten-minute verifiable delay so no miner can see which program a seed implies before choosing whether to publish a block, and feeds it to a deterministic generator. The generator emits a random integer program built from what graphics cards are uniquely good at: wide parallel integer maths, shuffles between the 32 lanes of a warp, and random reads over a multi-gigabyte dataset that changes daily, so the program waits on memory latency, not on maths or bandwidth. The memory footprint and instruction count are fixed and only the maths sequence is random, so no hour favours one vendor's cards and nobody gains by grinding the seed. Miners compile the program once per hour. Anyone running a node, a wallet or an exchange checks a hash on an ordinary CPU in under ten milliseconds by simulating one warp, so nobody needs a GPU except to mine. Measured: 0.41 to 0.58 ms per warp on one Apple M5 Max core with the 256 MB cache, about 17x inside the 10 ms gate; a 2019-class laptop core is not yet measured.

+

Each hour the chain derives a seed from a locked checkpoint one epoch back, passes it through a ten-minute verifiable delay so no miner can see which program a seed implies before choosing whether to publish a block, and feeds it to a deterministic generator. The generator emits a random integer program built from what graphics cards are uniquely good at: wide parallel integer maths, shuffles between the 32 lanes of a warp, and random reads over a multi-gigabyte dataset that changes daily, so the program waits on memory latency, not on maths or bandwidth. The memory footprint and instruction count are fixed and only the maths sequence is random, so no hour favours one vendor's cards and nobody gains by grinding the seed. Miners compile the program once per hour. Anyone running a node, a wallet or an exchange checks a hash on an ordinary CPU in under ten milliseconds by simulating one warp, so nobody needs a GPU except to mine. Measured: 0.61 ms per warp on one Apple M5 Max core for class v2 and 2.1 ms for class v3 (the mixer at x8, 5 October 2026, a loaded core; about 1.3 ms quiet, approximate), 4.8x inside the 10 ms gate; a 2019-class laptop core is not yet measured.

The hash is a lottery, not a general-purpose cryptographic hash. It has to be unpredictable per nonce, free of any shortcut cheaper than honest evaluation, and free of bias a miner can exploit. It does not need preimage or collision resistance. Open: no analysis of the lottery properties exists yet. It is the first job of the external review in phase 1, and until then the hash is a design claim backed by the measurements below.

@@ -434,7 +434,7 @@ body.all .pager{display:none} - + From ce01cb75e6ca1cd5b0eb24bf94701b693e648aa4 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:20:42 +0000 Subject: [PATCH 116/131] Counter ASIC 2.0: gate G4 run 3 green on the final class; status 22:20 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 96d0da579..83cbfa6ee 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -67,7 +67,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G1 | bit-exact v3 on all three vendors against the Mac reference | vectors PASS on Metal, CUDA (5090), AMD OpenCL for the v3 packs; batch fingerprints equal. RULING (coordinator, 5 October 2026, about 21:10 UTC): the integrated gfx1036 (RDNA 2, AMD OpenCL 3683.0) satisfies the AMD vendor tonight, because G1 is a compiler-and-ISA property and gfx1036 carried the v1 and v2 conformance; the 9070 XT's hash-rate and power rows are owed and taken when its link is back | open | | G2 | the CPU verifier exact on 1,000 random hashes per card | 1,000 GPU hashes per card re-hashed by `igneum-pow` on the Mac, 0 mismatches | open | | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | -| G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms | GREEN (runs 1 and 2) | +| G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms Run 3 PASS on the FINAL class (22:15:04 to 22:19:44Z, binaries from ca2-v3 d233fa1 = x8 + era + the verifier fix, fork 89dfcb95): 182 / 122 blocks around DAA 180, 304 in all; v3 ids e3 a6523b90cff501e3, e4 bc811b3c4b8b1ced, e5 e784541f19cdebe5 on all three miners; epoch 0's v2 id 8f8806638d59850f the same in all three runs; rejected 0/0/0; one sink a9ce45df8beeaf13 at 303/303/303; 3 of 3 switch lines; cache ready v2 179 ms, first v3 191 ms. Summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2215Z-era-mx8.json` | GREEN (runs 1, 2 and 3; run 3 on the final class) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | | G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 build-20261005-221237 (published 22:12:37Z: main d233fa1 = era + cache + the mixer fix with V3_CLASS = MX8, fork 89dfcb95) ran the five crates and the app on the final tree: every stage ok (22:12:37 to 22:15:44Z): kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8, 0 failed. With job 2 (kaspa-consensus alone 97) G6 is GREEN on the final tree | GREEN | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index dfc6fb525..ce0f1571b 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -525,3 +525,7 @@ docs/spec/01-lottery-hash.md on ca2-coord: 1.8.5 (the mixer x8 form with the mea 22:19. docs/analysis/proving-methods.md (branch proving-methods e7e0db7, not consensus): top recommendation re-size SP1's own GPU server (the floor is its code: the under-20 GB panic at sp1-gpu builder.rs:37, trace buffers at the maximum shard, a CUDA mempool that never releases), S_p as the dial, one server per card on rigs; the pinned ids stay; fallback and the Apple route: RISC Zero as proof-system version 2 behind the ProofSystem seam (8 GB at po2 19, 16 GB at po2 20, shipped Metal; 3 to 4 agent days); no 12 GB card has run a prover here, so a 4070 or 3060 in the loop is the first action. That branch merges into the 0.3.11 main tree as documentation (no code). evidence.md rows 17, 18 and 19 are rewritten on the measured class v3 (ba8379d). 22:20. The numbers page's Counter ASIC section is written as the bench-log entry "Counter ASIC 2.0, the numbers" (the page is built from the bench log; the litepaper links /bench#counter-asic-2-0-the-numbers); the litepaper's verifier figures moved to the v3 class (2.1 ms per warp on a loaded core, 4.8x inside the gate; the cache 512 MB from year 4). Everything public now carries the final class except the PC rows of the final-class packs and the G2 counts, which the running PC 1 job supplies. + +## 22:20 gate G4 run 3 PASS on the final class; the node side is final + +Fast-time 3-node network on ca2-v3 d233fa1 (x8 + era + the verifier fix) and fork 89dfcb95, 22:15:04 to 22:19:44Z: every check true; 182 / 122 blocks around the boundary; the v3 ids agree on all three miners; 0 rejected; one sink at 303/303/303; the switch line on 3 of 3; cache ready 179 / 191 ms. Final tips: ca2-v3 5eb2331 (docs only since d233fa1), ca2-v3-node 89dfcb95, both clean; the gate network stopped. Gates: G3 green, G4 green, G6 green; G1 and G2 on the final-class PC 1 job (running since 22:16); G5 at the ship's build step. From dd2ede567d5c9dd24e6c069721aaf351a2ecc925 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:21:57 +0000 Subject: [PATCH 117/131] Counter ASIC 2.0: gate G4b (the Mac mines v3: the app's --prepare-packs gap and the Metal gate run); status 22:21 --- docs/plans/counter-asic-2-rollout.md | 1 + docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 83cbfa6ee..2f4ce2626 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,6 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms Run 3 PASS on the FINAL class (22:15:04 to 22:19:44Z, binaries from ca2-v3 d233fa1 = x8 + era + the verifier fix, fork 89dfcb95): 182 / 122 blocks around DAA 180, 304 in all; v3 ids e3 a6523b90cff501e3, e4 bc811b3c4b8b1ced, e5 e784541f19cdebe5 on all three miners; epoch 0's v2 id 8f8806638d59850f the same in all three runs; rejected 0/0/0; one sink a9ce45df8beeaf13 at 303/303/303; 3 of 3 switch lines; cache ready v2 179 ms, first v3 191 ms. Summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2215Z-era-mx8.json` | GREEN (runs 1, 2 and 3; run 3 on the final class) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | +| G4b | the Mac mines v3: a real Metal miner across a v3 boundary through the miner's `--prepare-packs` flow, and the app passes that flag to the Metal worker | Found 22:21 UTC: `app/igneum-app/src/engine.rs` `miner_args` pushes `--prepare-packs` only when `card.worker != "Metal"` (line 1468 on a223ca9), so the Mac app never hands its Metal worker a prepare pack, and under class v3 the Metal worker compiles v3 only from a prepared pack (servePackProgram): at the first v3 epoch every Mac would answer `need` lines and stop, the 18:23Z outage class. Two closes before the ship, both in hand: (a) the app fix on 0.3.11's app branch (push the flag for every worker with the platform's path separator, a unit test on `miner_args` for a Metal card; assigned to the proving agent on top of a223ca9); (b) gate run 4: the fast-time network with one real Metal miner on the Mac across the activation (the prepare lines on both sides, the worker's `prepared` line for the v3 pack, found or accepted blocks on v3, no `need` or mismatch line; assigned to the node agent). If (a) is not in the app tree at the cut, the ship does not go: a Mac that cannot mine v3 at activation is a fleet outage, and the activation height (tip + 14,400) is not far enough to carry the fix in 0.3.12 safely | open: (a) and (b) in progress | | G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 build-20261005-221237 (published 22:12:37Z: main d233fa1 = era + cache + the mixer fix with V3_CLASS = MX8, fork 89dfcb95) ran the five crates and the app on the final tree: every stage ok (22:12:37 to 22:15:44Z): kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8, 0 failed. With job 2 (kaspa-consensus alone 97) G6 is GREEN on the final tree | GREEN | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index ce0f1571b..76dac616d 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -529,3 +529,7 @@ docs/spec/01-lottery-hash.md on ca2-coord: 1.8.5 (the mixer x8 form with the mea ## 22:20 gate G4 run 3 PASS on the final class; the node side is final Fast-time 3-node network on ca2-v3 d233fa1 (x8 + era + the verifier fix) and fork 89dfcb95, 22:15:04 to 22:19:44Z: every check true; 182 / 122 blocks around the boundary; the v3 ids agree on all three miners; 0 rejected; one sink at 303/303/303; the switch line on 3 of 3; cache ready 179 / 191 ms. Final tips: ca2-v3 5eb2331 (docs only since d233fa1), ca2-v3-node 89dfcb95, both clean; the gate network stopped. Gates: G3 green, G4 green, G6 green; G1 and G2 on the final-class PC 1 job (running since 22:16); G5 at the ship's build step. + +## 22:21 gate G4b added: the Mac must mine v3 (the app's --prepare-packs gap) + +The node agent's owed item is an app gap, confirmed in the 0.3.11 app branch: engine.rs miner_args pushes --prepare-packs only for non-Metal workers (line 1468 on a223ca9), so the Mac app's Metal worker would get no prepare pack and under class v3 would answer `need` at the first v3 epoch (every Mac stops: the 18:23Z class). Closes in hand: (a) the proving agent adds the flag for every worker (packs/prepare on macOS) with a unit test on miner_args for a Metal card, on the app branch above a223ca9; (b) the node agent runs gate 4: a real Metal miner on this Mac across a v3 boundary on the fast-time network through the miner's --prepare-packs flow. The ship does not go without (a) in the app tree and (b) green; recorded as G4b in the rollout plan's gate list. From 787c8e39bd091200f0f6327e514c501b1687a15f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:23:07 +0000 Subject: [PATCH 118/131] Counter ASIC 2.0: G4b (a) done, app e0de2ab; status 22:23 --- docs/plans/counter-asic-2-rollout.md | 4 ++-- docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 2f4ce2626..d53aa9347 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,7 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms Run 3 PASS on the FINAL class (22:15:04 to 22:19:44Z, binaries from ca2-v3 d233fa1 = x8 + era + the verifier fix, fork 89dfcb95): 182 / 122 blocks around DAA 180, 304 in all; v3 ids e3 a6523b90cff501e3, e4 bc811b3c4b8b1ced, e5 e784541f19cdebe5 on all three miners; epoch 0's v2 id 8f8806638d59850f the same in all three runs; rejected 0/0/0; one sink a9ce45df8beeaf13 at 303/303/303; 3 of 3 switch lines; cache ready v2 179 ms, first v3 191 ms. Summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2215Z-era-mx8.json` | GREEN (runs 1, 2 and 3; run 3 on the final class) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G4b | the Mac mines v3: a real Metal miner across a v3 boundary through the miner's `--prepare-packs` flow, and the app passes that flag to the Metal worker | Found 22:21 UTC: `app/igneum-app/src/engine.rs` `miner_args` pushes `--prepare-packs` only when `card.worker != "Metal"` (line 1468 on a223ca9), so the Mac app never hands its Metal worker a prepare pack, and under class v3 the Metal worker compiles v3 only from a prepared pack (servePackProgram): at the first v3 epoch every Mac would answer `need` lines and stop, the 18:23Z outage class. Two closes before the ship, both in hand: (a) the app fix on 0.3.11's app branch (push the flag for every worker with the platform's path separator, a unit test on `miner_args` for a Metal card; assigned to the proving agent on top of a223ca9); (b) gate run 4: the fast-time network with one real Metal miner on the Mac across the activation (the prepare lines on both sides, the worker's `prepared` line for the v3 pack, found or accepted blocks on v3, no `need` or mismatch line; assigned to the node agent). If (a) is not in the app tree at the cut, the ship does not go: a Mac that cannot mine v3 at activation is a fleet outage, and the activation height (tip + 14,400) is not far enough to carry the fix in 0.3.12 safely | open: (a) and (b) in progress | +| G4b | the Mac mines v3: a real Metal miner across a v3 boundary through the miner's `--prepare-packs` flow, and the app passes that flag to the Metal worker | Found 22:21 UTC: `app/igneum-app/src/engine.rs` `miner_args` pushes `--prepare-packs` only when `card.worker != "Metal"` (line 1468 on a223ca9), so the Mac app never hands its Metal worker a prepare pack, and under class v3 the Metal worker compiles v3 only from a prepared pack (servePackProgram): at the first v3 epoch every Mac would answer `need` lines and stop, the 18:23Z outage class. Two closes before the ship, both in hand: (a) the app fix on 0.3.11's app branch (push the flag for every worker with the platform's path separator, a unit test on `miner_args` for a Metal card; assigned to the proving agent on top of a223ca9); (b) gate run 4: the fast-time network with one real Metal miner on the Mac across the activation (the prepare lines on both sides, the worker's `prepared` line for the v3 pack, found or accepted blocks on v3, no `need` or mismatch line; assigned to the node agent). If (a) is not in the app tree at the cut, the ship does not go: a Mac that cannot mine v3 at activation is a fleet outage, and the activation height (tip + 14,400) is not far enough to carry the fix in 0.3.12 safely | (a) DONE: app proving-v1 e0de2ab (prepare_packs_arg() for every worker, `packs/prepare` on macOS; start_miner gives the Metal worker the app data folder as cwd, which was None; unit test every_worker_gets_the_prepare_directory_in_the_platform_form; 113 + 27 + 8 passed). (b) gate run 4 pending | | G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 build-20261005-221237 (published 22:12:37Z: main d233fa1 = era + cache + the mixer fix with V3_CLASS = MX8, fork 89dfcb95) ran the five crates and the app on the final tree: every stage ok (22:12:37 to 22:15:44Z): kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8, 0 failed. With job 2 (kaspa-consensus alone 97) G6 is GREEN on the final tree | GREEN | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. @@ -102,7 +102,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's ## 8a. Proving v1 rides with it -the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL a223ca9 on 5b0d54f, docs-only 8b47073 after it (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. +the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL e0de2ab on 5b0d54f (a223ca9 the resume fix, e0de2ab the Metal --prepare-packs fix; docs-only commits between) (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. Branches merged into the 0.3.11 main tree beside ca2-v3 and ca2-coord (tooling, no consensus): consequences (the reviewer's rows), bash-body-check (7adb1ca, 6805125: tools/ci/bash-body-check.sh with fixtures and a self-test in ci.yml, the convention paragraph in packaging/README-ship.md, `publish-jobs.sh add --kind run` running the bash-body and prover-socket checks before signing; add/add with proving-v1 c2544be on tools/ci/prover-socket-check.sh: take bash-body-check's superset and proving-v1's ci.yml step; tools/amd-prove/pc1-cpu-prove.ps1 goes on the allow list, it starts no GPU server), amd-prove (f1d7a7d), card-lifetime (1fecfe2). Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 76dac616d..926613eb9 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -533,3 +533,5 @@ Fast-time 3-node network on ca2-v3 d233fa1 (x8 + era + the verifier fix) and for ## 22:21 gate G4b added: the Mac must mine v3 (the app's --prepare-packs gap) The node agent's owed item is an app gap, confirmed in the 0.3.11 app branch: engine.rs miner_args pushes --prepare-packs only for non-Metal workers (line 1468 on a223ca9), so the Mac app's Metal worker would get no prepare pack and under class v3 would answer `need` at the first v3 epoch (every Mac stops: the 18:23Z class). Closes in hand: (a) the proving agent adds the flag for every worker (packs/prepare on macOS) with a unit test on miner_args for a Metal card, on the app branch above a223ca9; (b) the node agent runs gate 4: a real Metal miner on this Mac across a v3 boundary on the fast-time network through the miner's --prepare-packs flow. The ship does not go without (a) in the app tree and (b) green; recorded as G4b in the rollout plan's gate list. + +22:23. G4b (a) done: app proving-v1 e0de2ab on 5b0d54f: miner_args pushes --prepare-packs for every worker through prepare_packs_arg() (packs\prepare on Windows, packs/prepare elsewhere), the OpenCL-only --job-nonces kept, and start_miner now gives the Metal worker the app data folder as cwd (it had None, which would have broken the relative path: a second latent Mac fault closed); unit test every_worker_gets_the_prepare_directory_in_the_platform_form; 113 + 27 + 8 passed. The 0.3.11 app tree's final code commit is e0de2ab. (b), the Metal gate run, is with the node agent. From 02b165a4144cec4ca7483a014a396d20473380ae Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:23:46 +0000 Subject: [PATCH 119/131] Counter ASIC 2.0 status 22:23: floor-build-1 failed on its log path, floor-build-2 running --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 926613eb9..f55f9cd16 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -535,3 +535,5 @@ Fast-time 3-node network on ca2-v3 d233fa1 (x8 + era + the verifier fix) and for The node agent's owed item is an app gap, confirmed in the 0.3.11 app branch: engine.rs miner_args pushes --prepare-packs only for non-Metal workers (line 1468 on a223ca9), so the Mac app's Metal worker would get no prepare pack and under class v3 would answer `need` at the first v3 epoch (every Mac stops: the 18:23Z class). Closes in hand: (a) the proving agent adds the flag for every worker (packs/prepare on macOS) with a unit test on miner_args for a Metal card, on the app branch above a223ca9; (b) the node agent runs gate 4: a real Metal miner on this Mac across a v3 boundary on the fast-time network through the miner's --prepare-packs flow. The ship does not go without (a) in the app tree and (b) green; recorded as G4b in the rollout plan's gate list. 22:23. G4b (a) done: app proving-v1 e0de2ab on 5b0d54f: miner_args pushes --prepare-packs for every worker through prepare_packs_arg() (packs\prepare on Windows, packs/prepare elsewhere), the OpenCL-only --job-nonces kept, and start_miner now gives the Metal worker the app data folder as cwd (it had None, which would have broken the relative path: a second latent Mac fault closed); unit test every_worker_gets_the_prepare_directory_in_the_platform_form; 113 + 27 + 8 passed. The 0.3.11 app tree's final code commit is e0de2ab. (b), the Metal gate run, is with the node agent. + +22:23. PC 2: floor-build-1 (22:17 to 22:21:12Z) FAILED at cargo exit 101 after 240 s with the error not uploaded (the playbook named the cargo log with a timestamp and could not find it; fixed: a fixed path, the error lines printed on failure); floor-build-2 (the same 90-minute shape) has the go at 22:24; the aggregation-cost re-run follows it. PC 1: the era agent's final-class job runs (from 22:16); the AMD sweep follows on a fresh probe. From 0839e259ccde1299e19e52238c5aa746f9e9307f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:24:19 +0000 Subject: [PATCH 120/131] Counter ASIC 2.0: gates G1 and G2 green on the final class; spec 1.17 fingerprints; status 22:24 --- docs/plans/counter-asic-2-rollout.md | 4 ++-- docs/plans/counter-asic-2-status.md | 4 ++++ docs/spec/01-lottery-hash.md | 2 +- 3 files changed, 7 insertions(+), 3 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index d53aa9347..1b18a36eb 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -64,8 +64,8 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | # | Gate | Evidence required | State | |---|---|---|---| -| G1 | bit-exact v3 on all three vendors against the Mac reference | vectors PASS on Metal, CUDA (5090), AMD OpenCL for the v3 packs; batch fingerprints equal. RULING (coordinator, 5 October 2026, about 21:10 UTC): the integrated gfx1036 (RDNA 2, AMD OpenCL 3683.0) satisfies the AMD vendor tonight, because G1 is a compiler-and-ISA property and gfx1036 carried the v1 and v2 conformance; the 9070 XT's hash-rate and power rows are owed and taken when its link is back | open | -| G2 | the CPU verifier exact on 1,000 random hashes per card | 1,000 GPU hashes per card re-hashed by `igneum-pow` on the Mac, 0 mismatches | open | +| G1 | bit-exact v3 on all three vendors against the Mac reference | GREEN on the final class (job run-ca2-era-pc1b-20261005, 22:16 to 22:21Z, exit 0 in 304 s, both cards restored, app 0.3.10, the 9070 XT present as gfx1201): the seven final-class packs' 2^24 fingerprints equal on the RTX 5090 (CUDA/NVRTC), the RX 9070 XT (OpenCL) and the M5 Max (Metal): mx8-devnet-epoch0 90f794dd556f7a3b, era-0 8e8e070db4eea52d, era-1 891c01b8563bb47e, era-2 e54279fed2831b5d, era-3 77e0ba8abbd0ae62, era-4 d898d8f4f2e7684b, era-5 a6927db380f7efb2; self-test PASS on every pack on both cards. RULING (coordinator, 5 October 2026, about 21:10 UTC): the integrated gfx1036 (RDNA 2, AMD OpenCL 3683.0) satisfies the AMD vendor tonight, because G1 is a compiler-and-ISA property and gfx1036 carried the v1 and v2 conformance; the 9070 XT's hash-rate and power rows are owed and taken when its link is back | open | +| G2 | the CPU verifier exact on 1,000 random hashes per card | GREEN: one serve-mode job of 1,024 nonces at target ff..ff per card and pack (every nonce a found line), re-hashed on the Mac with `igneum-pow hash-bound --prehash 00..01 --count 1024` on the same pack: RTX 5090 era-0 1,024 of 1,024 and mx8-devnet-epoch0 1,024 of 1,024; RX 9070 XT era-0 1,024 of 1,024 and mx8-devnet-epoch0 1,024 of 1,024 (the same job) | GREEN | | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms Run 3 PASS on the FINAL class (22:15:04 to 22:19:44Z, binaries from ca2-v3 d233fa1 = x8 + era + the verifier fix, fork 89dfcb95): 182 / 122 blocks around DAA 180, 304 in all; v3 ids e3 a6523b90cff501e3, e4 bc811b3c4b8b1ced, e5 e784541f19cdebe5 on all three miners; epoch 0's v2 id 8f8806638d59850f the same in all three runs; rejected 0/0/0; one sink a9ce45df8beeaf13 at 303/303/303; 3 of 3 switch lines; cache ready v2 179 ms, first v3 191 ms. Summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2215Z-era-mx8.json` | GREEN (runs 1, 2 and 3; run 3 on the final class) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index f55f9cd16..494a2ae89 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -537,3 +537,7 @@ The node agent's owed item is an app gap, confirmed in the 0.3.11 app branch: en 22:23. G4b (a) done: app proving-v1 e0de2ab on 5b0d54f: miner_args pushes --prepare-packs for every worker through prepare_packs_arg() (packs\prepare on Windows, packs/prepare elsewhere), the OpenCL-only --job-nonces kept, and start_miner now gives the Metal worker the app data folder as cwd (it had None, which would have broken the relative path: a second latent Mac fault closed); unit test every_worker_gets_the_prepare_directory_in_the_platform_form; 113 + 27 + 8 passed. The 0.3.11 app tree's final code commit is e0de2ab. (b), the Metal gate run, is with the node agent. 22:23. PC 2: floor-build-1 (22:17 to 22:21:12Z) FAILED at cargo exit 101 after 240 s with the error not uploaded (the playbook named the cargo log with a timestamp and could not find it; fixed: a fixed path, the error lines printed on failure); floor-build-2 (the same 90-minute shape) has the go at 22:24; the aggregation-cost re-run follows it. PC 1: the era agent's final-class job runs (from 22:16); the AMD sweep follows on a fresh probe. + +## 22:24 gates G1 and G2 GREEN on the final class + +Job run-ca2-era-pc1b-20261005 (22:16 to 22:21Z, 304 s, both cards restored, app 0.3.10, the 9070 XT present): the seven final-class packs' 2^24 fingerprints equal on the 5090, the 9070 XT and the M5 Max (mx8-devnet-epoch0 90f794dd556f7a3b; era-0 8e8e070db4eea52d, era-1 891c01b8563bb47e, era-2 e54279fed2831b5d, era-3 77e0ba8abbd0ae62, era-4 d898d8f4f2e7684b, era-5 a6927db380f7efb2); G2: 1,024 of 1,024 per card on era-0 and the pinned pack re-hashed by the Rust verifier. Rates on the final class (MH/s): 5090 135.90 to 137.70 (spread 1.3%), 9070 XT 18.59 to 19.18 (3.1%), M5 Max 27.85 to 27.98 (0.5%); the pinned pack 136.10 / 18.90 / 27.80. Spec 1.17 carries these fingerprints. Gates now: G1, G2, G3, G4, G6 green; G4b (a) done, (b) the Metal gate run pending; G5 at the ship's build. PC 1 goes to the AMD sweep on its presence probe. diff --git a/docs/spec/01-lottery-hash.md b/docs/spec/01-lottery-hash.md index 6ee2a5414..74297072f 100644 --- a/docs/spec/01-lottery-hash.md +++ b/docs/spec/01-lottery-hash.md @@ -542,6 +542,6 @@ Pack `proto-cuda/packs/igneum-devnet-v4-epoch0/`: epoch seed bytes `edc4fa844da9 Header-bound vectors (section 1.6 rule, seed `igneum-genesis`, day `2026-10-03`): `igneum-pow/README.md`, eight values, for example H = 32 zero bytes and nonce 0 give `746c567b090acf6a`. -Program class v3 vectors (5 October 2026, `proto-cuda/packs-ca2-mixer/` and `proto-cuda/packs-ca2-era/`; generator 3, mixer x8, the cache growth rule, the era draw inside the class): pack `mx8-genesis` (seed `igneum-genesis`, no era, program id `e323b9dcaf283a6f`, batch fingerprint `7c28cfb06c5c65a9`) and pack `mx8-devnet-epoch0` (the devnet epoch seed and day of 1.17 with the era stand-in E_0 = the devnet genesis hash inside the class, program id `73bcbfe8ccf988f1`, unit 0 lane 0 `d424577fce4a7a60`, fingerprint `90f794dd556f7a3b` over the harness's 2^20 outputs and `a6752e037514c91a` over 2^16), reproduced by the Rust interpreter, Metal and Apple OpenCL on 5 October 2026 (3/3 standalone, 3/3 in batch, 96 of 96 lanes each); the six era packs `era-0` to `era-5` (the same epoch seed and day, era test seeds 0 to 5, program id `73bcbfe8ccf988f1`, fingerprints over 2^16 `64c0ee90bac42624`, `fb276b04bab43db2`, `51a15e86ce7afdad`, `2d4f94bae8356fd8`, `194ce26f9ebd508e`, `43673acc89954d5e`), 7/7 on Metal and Apple OpenCL; the RTX 5090 and RX 9070 XT runs of these exact packs and the 1,024-hash CPU re-check per card are the job `run-ca2-era-pc1b-20261005` of `docs/bench-log.md` (5 October 2026, night). The class v2 vectors above stand unchanged (the v2 exports are byte-identical on the class v3 crate, `igneum-pow/tests/packs.rs`). +Program class v3 vectors (5 October 2026, `proto-cuda/packs-ca2-mixer/` and `proto-cuda/packs-ca2-era/`; generator 3, mixer x8, the cache growth rule, the era draw inside the class): pack `mx8-genesis` (seed `igneum-genesis`, no era, program id `e323b9dcaf283a6f`, batch fingerprint `7c28cfb06c5c65a9`) and pack `mx8-devnet-epoch0` (the devnet epoch seed and day of 1.17 with the era stand-in E_0 = the devnet genesis hash inside the class, program id `73bcbfe8ccf988f1`, unit 0 lane 0 `d424577fce4a7a60`, batch fingerprint `90f794dd556f7a3b` over 2^24 outputs at base nonce 0), reproduced by the Rust interpreter, Metal and Apple OpenCL on 5 October 2026 (3/3 standalone, 3/3 in batch, 96 of 96 lanes each); the six era packs `era-0` to `era-5` (the same epoch seed and day, era test seeds 0 to 5, program id `73bcbfe8ccf988f1`; 2^24 fingerprints `8e8e070db4eea52d`, `891c01b8563bb47e`, `e54279fed2831b5d`, `77e0ba8abbd0ae62`, `d898d8f4f2e7684b`, `a6927db380f7efb2`). All seven fingerprints are equal on the RTX 5090 (CUDA/NVRTC), the RX 9070 XT (AMD OpenCL) and the M5 Max (Metal), self-test PASS on every pack, and 1,024 random nonces per card on `era-0` and `mx8-devnet-epoch0` re-hash to the same value on the Rust verifier (1,024 of 1,024 each): job `run-ca2-era-pc1b-20261005`, 5 October 2026, `docs/bench-log.md`. The class v2 vectors above stand unchanged (the v2 exports are byte-identical on the class v3 crate, `igneum-pow/tests/packs.rs`). Cache, dataset and mixer vectors: section 1.8.4 and 1.8.5. Seed words: section 1.3.1. Generator: section 1.4.3. Acceptance: section 1.4.6. Batch fingerprints: section 1.15 item 5. From c877e11ff5d1c648fb3305ee9f688bd684ea72fd Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:25:02 +0000 Subject: [PATCH 121/131] C29 and D11: the verifier line on the fixed crate, the bounty struck until escrowed, the final-class rates in the public tables; status 22:25 --- docs/bench-log.md | 10 +++++----- docs/plans/counter-asic-2-public.md | 12 ++++++------ docs/plans/counter-asic-2-status.md | 2 ++ site/index.html | 2 +- site/litepaper.html | 8 ++++---- 5 files changed, 18 insertions(+), 16 deletions(-) diff --git a/docs/bench-log.md b/docs/bench-log.md index bc0cfe101..cd648250d 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1595,11 +1595,11 @@ Reading, and the consequences (CLAUDE.md, every number). Doubling the cycles add | Card | v2 MH/s | v3 MH/s, six eras (spread) | Bytes per hash | Latency-bound share | Daily 1 GiB build, v2 / v3 | |---|---|---|---|---|---| -| Apple M5 Max, Metal | 27.68 | 28.35 to 28.58 (0.8%) | 512 | 1.06 | 21 / 21 ms | -| RTX 5090, CUDA | 137.2 | 136.18 to 138.01 (1.3%) | 512 | 1.01 | 25 / 23 ms | -| RX 9070 XT, OpenCL | 18.09 | 18.61 to 19.21 (3.2%) | 512 | 0.95 | 74 / 75 ms | +| Apple M5 Max, Metal | 27.68 | 27.85 to 27.98 (0.5%) | 512 | 1.06 | 21 / 21 ms | +| RTX 5090, CUDA | 137.2 | 135.90 to 137.70 (1.3%) | 512 | 1.01 | 25 / 23 ms | +| RX 9070 XT, OpenCL | 18.09 | 18.59 to 19.18 (3.1%) | 512 | 0.95 | 74 / 75 ms | -CPU verifier, one M5 Max core (loaded box; ratios are the measurement): v2 0.61 ms per warp, v3 (x8) 2.08 ms, worst cold 2.15; the 10 ms gate holds 4.8x. Bit-exact: every v3 pack's fingerprint equal on Metal, Apple OpenCL, CUDA and AMD OpenCL. +CPU verifier, one M5 Max core at load average 5.5 (the fixed crate, ca2-mixer 1ab8b21): v2 0.61 ms per warp, v3 (x8) 2.08 ms (3.4x), worst cold 2.15; the 10 ms gate holds 4.8x (4.6x on the worst cold unit). Bit-exact: every v3 pack's fingerprint equal on Metal, Apple OpenCL, CUDA and AMD OpenCL. | Layer | Measured | Decision | The number | |---|---|---|---| @@ -1617,4 +1617,4 @@ CPU verifier, one M5 Max core (loaded box; ratios are the measurement): v2 0.61 **The user tiers.** AMD RDNA 4 sits at about a seventh of a 5090 on this hash (its dependent-read rate: 2.4 G against 17.5 G per second), 2.2x worse per pound at list prices and 4.9x worse per watt (approximate); the card's memory system, not a tuning gap. The integrated tier on the CUDA and OpenCL one-click workers mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). Card lifetime under the step schedule: a 4 GB card to year 4, 8 GB to year 12, 12 GB to year 28 with the cache freed after the daily build. -**The bounty.** A standing bounty for any chip design beating a GPU by more than 2x on the published model, with a leaderboard by card model, January 2027 (spec O-1.17). +**The bounty.** A bounty for any chip design beating a GPU by more than 2x on the published model, with a leaderboard by card model, follows the external review (spec O-1.17, January 2027); it is named publicly only once escrowed (`docs/plans/funding.md`, rule 3), which it is not yet. diff --git a/docs/plans/counter-asic-2-public.md b/docs/plans/counter-asic-2-public.md index 1c76354d9..05c79ff4d 100644 --- a/docs/plans/counter-asic-2-public.md +++ b/docs/plans/counter-asic-2-public.md @@ -4,7 +4,7 @@ the project lead, 5 October 2026 (night): "not an information overload". Four le ## Level 1: one sentence (site hero, litepaper abstract) -Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public. +Built for graphics cards. A custom chip gains under 2x, and the model is public. (The bounty is named only once it is escrowed: docs/plans/funding.md rule 3; D11 for the project lead.) ## Level 2: one site card, one short litepaper section @@ -16,7 +16,7 @@ Three ideas, no layer names, no widths, no SRAM. **Miners hold the switch.** Spare defences are written into the rules, switched off. A 90% miner signal turns one on. No fork. -A custom chip gains under 2x. Model published, bounty standing. [link: the numbers page] +A custom chip gains under 2x. Model published; a bounty follows the external review. [link: the numbers page] Litepaper only, a fourth paragraph: No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured. @@ -30,11 +30,11 @@ Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per h | Card | v2 MH/s | v3 MH/s (era packs, six eras) | Bytes per hash | Latency-bound share | Verifier ms per warp (v2 / v3, one loaded M5 Max core) | Daily 1 GiB build (v2 / v3) | |---|---|---|---|---|---|---| -| Apple M5 Max (Metal) | 27.68 | 28.35 to 28.58 (spread 0.8%) | 512 | 1.06 | 0.61 / 2.08 | 21 / 21 ms | -| RTX 5090 (CUDA) | 137.2 | 136.18 to 138.01 (spread 1.3%) | 512 | 1.01 | the same verifier | 25 / 23 ms | -| RX 9070 XT (OpenCL) | 18.09 | 18.61 to 19.21 (spread 3.2%) | 512 | 0.95 | the same verifier | 74 / 75 ms | +| Apple M5 Max (Metal) | 27.68 | 27.85 to 27.98 (spread 0.5%, the final class) | 512 | 1.06 | 0.61 / 2.08 (3.4x; worst cold 2.15) | 21 / 21 ms | +| RTX 5090 (CUDA) | 137.2 | 135.90 to 137.70 (spread 1.3%, the final class) | 512 | 1.01 | the same verifier | 25 / 23 ms | +| RX 9070 XT (OpenCL) | 18.09 | 18.59 to 19.18 (spread 3.1%, the final class) | 512 | 0.95 | the same verifier | 74 / 75 ms | -Every number measured 5 October 2026 (`docs/plans/era-layout.md`, `docs/plans/mixer-x4.md`, `docs/bench-log.md`); the v3 verifier figure is the x8 mixer on a loaded core (about 1.3 ms quiet, approximate). Bit-exact: every v3 pack's fingerprint equal on the three vendors. +Every number measured 5 October 2026 (`docs/plans/era-layout.md`, `docs/plans/mixer-x4.md`, `docs/bench-log.md`); the v3 verifier figure is the x8 mixer on one core at load average 5.5 (the fixed crate; the same session matched readwidth's quiet v2 figure within 1%). Bit-exact: every v3 pack's fingerprint equal on the three vendors. Chip model, before and after (`docs/analysis/chip-model-v3.md`): the on-die-cache recompute chip (the whole 256 MiB cache in SRAM, about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) against the RTX 5090's measured 136.1 MH/s at 50 T integer op/s: class v2 333 MH/s, 2.4x; class v3 (mixer x8) 41.7 MH/s, 0.31x bare, 0.92x with a 3x fixed-function allowance (approximate), 0.76x at equal silicon. The claim "under 2x" holds with the margin stated: 8% on the allowance (a 3.3x allowance reads 1.0x), 9% on the budget. Next levers, named: the mixer at x16 (the verifier at about 4 ms per warp, inside the 10 ms gate; a 2019-class core unmeasured), a hot table small enough to stay resident beside the streaming dataset (measured and not adopted tonight: 32 to 96 MiB tables cost the GPU 7 to 20% and help the chip). diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 494a2ae89..dbe668240 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -541,3 +541,5 @@ The node agent's owed item is an app gap, confirmed in the 0.3.11 app branch: en ## 22:24 gates G1 and G2 GREEN on the final class Job run-ca2-era-pc1b-20261005 (22:16 to 22:21Z, 304 s, both cards restored, app 0.3.10, the 9070 XT present): the seven final-class packs' 2^24 fingerprints equal on the 5090, the 9070 XT and the M5 Max (mx8-devnet-epoch0 90f794dd556f7a3b; era-0 8e8e070db4eea52d, era-1 891c01b8563bb47e, era-2 e54279fed2831b5d, era-3 77e0ba8abbd0ae62, era-4 d898d8f4f2e7684b, era-5 a6927db380f7efb2); G2: 1,024 of 1,024 per card on era-0 and the pinned pack re-hashed by the Rust verifier. Rates on the final class (MH/s): 5090 135.90 to 137.70 (spread 1.3%), 9070 XT 18.59 to 19.18 (3.1%), M5 Max 27.85 to 27.98 (0.5%); the pinned pack 136.10 / 18.90 / 27.80. Spec 1.17 carries these fingerprints. Gates now: G1, G2, G3, G4, G6 green; G4b (a) done, (b) the Metal gate run pending; G5 at the ship's build. PC 1 goes to the AMD sweep on its presence probe. + +22:25. Consequences C29 and D11 applied. C29: the litepaper's verifier line now reads 2.1 ms per warp for class v3 on one M5 Max core at load average 5.5 (the fixed crate; worst cold 2.15), 3.4x the v2 verifier's 0.61 ms, the gate leaving 4.8x (4.6x on the worst cold unit); the "about 1.3 ms quiet" scaling is struck: it came from the slow binary's 2.79 ms session (21:40), and the fixed binary's session (22:07) matched readwidth's quiet v2 figure within 1%, so the fixed numbers are near-quiet. D11: "and the bounty" struck from the hero and the abstract, "standing" from level 1; the copy says "a bounty follows the external review"; the bounty is named only once escrowed (funding.md rule 3; USD 50,000 not funded); the project lead decides (D11). The level 3 table and the numbers-page entry now carry the final-class rates (5090 135.90 to 137.70, 9070 XT 18.59 to 19.18, M5 Max 27.85 to 27.98). diff --git a/site/index.html b/site/index.html index 60976e26f..5cdbb14c4 100644 --- a/site/index.html +++ b/site/index.html @@ -340,7 +340,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var(
GPUs are back · for good

Mined by GPUs.
Proven by fire.

-

Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public. A mining program that rewrites itself every hour, so a chip built for one hour is useless the next. An NVIDIA card with 24 GB or more proves every block and gets paid for it; AMD and Apple cards mine. No premine, no stake, no foundation, no merge to proof of stake.

+

Built for graphics cards. A custom chip gains under 2x, and the model is public. A mining program that rewrites itself every hour, so a chip built for one hour is useless the next. An NVIDIA card with 24 GB or more proves every block and gets paid for it; AMD and Apple cards mine. No premine, no stake, no foundation, no merge to proof of stake.

See the miner Read the litepaper diff --git a/site/litepaper.html b/site/litepaper.html index 3c6bce685..9952d8b6f 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -298,7 +298,7 @@ body.all .pager{display:none}

Abstract

-

Igneum is a proof-of-work blockchain built for graphics cards, where NVIDIA cards also prove every block with zero-knowledge proofs and sell proving to other chains. A custom chip gains under 2x, and the model and the bounty are public.

+

Igneum is a proof-of-work blockchain built for graphics cards, where NVIDIA cards also prove every block with zero-knowledge proofs and sell proving to other chains. A custom chip gains under 2x, and the model is public; a bounty follows the external review.

It runs the Ethereum virtual machine, so anything built for Ethereum runs on Igneum unchanged. Transactions are included in about one second, proven within about a minute at launch, and locked by miners within about two. There is no premine, no pre-sale, no treasury taken from emission, no stake anywhere in consensus, and no dependence on any other chain. Mining stays open to anyone with a GPU because the mining program changes every hour, so a chip built for one program is useless for the next, and a chip for the whole program space is a GPU without the graphics parts. No scheduled human release is needed to keep it that way. Writing new code, including an emergency fix to the proof system, is the one thing that takes a person, and it activates only on miner signalling.

1 / s
blocks, rising to 10
@@ -408,7 +408,7 @@ body.all .pager{display:none}

Mining: a program that never holds still

Every GPU chain that promised ASIC resistance shipped a fixed algorithm, and a fixed algorithm gets a chip the moment the prize pays for one. Igneum does not have a fixed algorithm.

-

Each hour the chain derives a seed from a locked checkpoint one epoch back, passes it through a ten-minute verifiable delay so no miner can see which program a seed implies before choosing whether to publish a block, and feeds it to a deterministic generator. The generator emits a random integer program built from what graphics cards are uniquely good at: wide parallel integer maths, shuffles between the 32 lanes of a warp, and random reads over a multi-gigabyte dataset that changes daily, so the program waits on memory latency, not on maths or bandwidth. The memory footprint and instruction count are fixed and only the maths sequence is random, so no hour favours one vendor's cards and nobody gains by grinding the seed. Miners compile the program once per hour. Anyone running a node, a wallet or an exchange checks a hash on an ordinary CPU in under ten milliseconds by simulating one warp, so nobody needs a GPU except to mine. Measured: 0.61 ms per warp on one Apple M5 Max core for class v2 and 2.1 ms for class v3 (the mixer at x8, 5 October 2026, a loaded core; about 1.3 ms quiet, approximate), 4.8x inside the 10 ms gate; a 2019-class laptop core is not yet measured.

+

Each hour the chain derives a seed from a locked checkpoint one epoch back, passes it through a ten-minute verifiable delay so no miner can see which program a seed implies before choosing whether to publish a block, and feeds it to a deterministic generator. The generator emits a random integer program built from what graphics cards are uniquely good at: wide parallel integer maths, shuffles between the 32 lanes of a warp, and random reads over a multi-gigabyte dataset that changes daily, so the program waits on memory latency, not on maths or bandwidth. The memory footprint and instruction count are fixed and only the maths sequence is random, so no hour favours one vendor's cards and nobody gains by grinding the seed. Miners compile the program once per hour. Anyone running a node, a wallet or an exchange checks a hash on an ordinary CPU in under ten milliseconds by simulating one warp, so nobody needs a GPU except to mine. Measured: 0.61 ms per warp on one Apple M5 Max core for class v2 and 2.1 ms for class v3 (the mixer at x8, 5 October 2026, one core at load average 5.5, worst cold unit 2.15 ms), 3.4x the class v2 verifier; the 10 ms gate leaves 4.8x (4.6x on the worst cold unit); a 2019-class laptop core is not yet measured.

The hash is a lottery, not a general-purpose cryptographic hash. It has to be unpredictable per nonce, free of any shortcut cheaper than honest evaluation, and free of bias a miner can exploit. It does not need preimage or collision resistance. Open: no analysis of the lottery properties exists yet. It is the first job of the external review in phase 1, and until then the hash is a design claim backed by the measurements below.

ClockWhat changesMiner update needed?
Hardware it is built forCPUs. GPUs run it badly on purposeGPUs. Any card, any vendor. Bit-exact on Apple, NVIDIA and AMD, measured
Random programPer hash, interpreted in a virtual machinePer hour, compiled to native GPU code. Per hash, the 128 dataset addresses change with the nonce
DatasetAbout 2 GB, the same size since 2019, approximate2 GB at genesis, doubling at years 4, 12 and 28 under the step schedule; a 4 GB card mines about four years, an 8 GB card past a decade
Light verification256 MB cache on a CPU, milliseconds256 MB cache on a CPU, one warp under 10 ms, the gate. Measured 0.41 to 0.58 ms on one Apple M5 Max core; a 2019-class core not yet
Light verification256 MB cache on a CPU, milliseconds256 MB cache on a CPU (512 MB from year 4), one warp under 10 ms, the gate. Measured 2.1 ms on one loaded Apple M5 Max core for class v3; a 2019-class core not yet
Changes over timeNone. A fixed design, unchanged for seven yearsA new program every hour, its memory pattern with it; era draws and reserved families on a schedule fixed at genesis. Nobody touches it
Seed grindingNot applicable, the program comes from the hash inputClosed by a verifiable delay between seed and program
Useful workNone. Hashing onlyNVIDIA cards with 24 GB or more prove every block and sell proofs to other chains; AMD and Apple cards mine, and a prover for them lands when a zkVM ships one
@@ -421,7 +421,7 @@ body.all .pager{display:none}
ClockWhat changesMiner update needed?

Three ideas carry the chip resistance. The hash rewrites itself. A new program every hour, drawn from the chain. Its memory pattern changes with it. The rules change on a schedule fixed at launch. No release, no vote. It waits on memory, not maths. Every hash is a chain of random reads into a table too big for a chip to carry. The wait is the same physics for everyone. Miners hold the switch. Spare defences are written into the rules, switched off. A 90% miner signal turns one on. No fork.

-

No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured. The model and the bounty are public: the numbers. Monero has run on RandomX since 2019 with no chip publicly shipped, approximate; that is precedent, not proof.

+

No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured. The model is public: the numbers; a bounty follows the external review. Monero has run on RandomX since 2019 with no chip publicly shipped, approximate; that is precedent, not proof.

One thing takes a person, here and on every chain that exists: writing new code. A chain cannot safely write its own generator, and it cannot safely tell a chip from a wave of honest new cards by hashrate alone. If the design above ever failed, anyone could publish a new generator and miners would switch it on by signalling, as Monero's community can fork. Igneum is built to make that day unlikely, and does not depend on avoiding it.

@@ -434,7 +434,7 @@ body.all .pager{display:none} Hardware it is built forCPUs. GPUs run it badly on purposeGPUs. Any card, any vendor. Bit-exact on Apple, NVIDIA and AMD, measured Random programPer hash, interpreted in a virtual machinePer hour, compiled to native GPU code. Per hash, the 128 dataset addresses change with the nonce DatasetAbout 2 GB, the same size since 2019, approximate2 GB at genesis, doubling at years 4, 12 and 28 under the step schedule; a 4 GB card mines about four years, an 8 GB card past a decade - Light verification256 MB cache on a CPU, milliseconds256 MB cache on a CPU (512 MB from year 4), one warp under 10 ms, the gate. Measured 2.1 ms on one loaded Apple M5 Max core for class v3; a 2019-class core not yet + Light verification256 MB cache on a CPU, milliseconds256 MB cache on a CPU (512 MB from year 4), one warp under 10 ms, the gate. Measured 2.1 ms on one Apple M5 Max core for class v3 (3.4x class v2's 0.61 ms); a 2019-class core not yet Changes over timeNone. A fixed design, unchanged for seven yearsA new program every hour, its memory pattern with it; era draws and reserved families on a schedule fixed at genesis. Nobody touches it Seed grindingNot applicable, the program comes from the hash inputClosed by a verifiable delay between seed and program Useful workNone. Hashing onlyNVIDIA cards with 24 GB or more prove every block and sell proofs to other chains; AMD and Apple cards mine, and a prover for them lands when a zkVM ships one From 78c040095d9702462b8ac4fd8ed908cf514a6f3a Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:25:24 +0000 Subject: [PATCH 122/131] Counter ASIC 2.0: the fourth eGPU drop in the hardware events, the era branch final, PC 1 to Ember; status 22:25 --- docs/plans/counter-asic-2-rollout.md | 1 + docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 5 insertions(+) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 1b18a36eb..82d2a611d 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -93,6 +93,7 @@ Public lines to fix on the integration branch (card-lifetime table 3): `site/ind |---|---|---|---| | 5 October, at install (earlier today) | the RX 9070 XT in the Sonnet Breakaway Box 850T5 over USB4 went Code 43 | came back after a driver reinstall and a reboot | | | 5 October, about 20:40 | the 9070 XT dropped off PC 1's bus: Get-PnpDevice -Class Display lists only the integrated AMD Radeon Graphics (gfx1036) and the RTX 5090; after `pnputil /scan-devices` at 20:45:34Z the card is still absent and the USB4 list shows only the host and root routers: the "USB4 Router (2.0), Sonnet Technologies Breakaway Box 850T5" present at 17:18Z is gone, so the box itself is off the link | the AMD worker (igneum-worker-opencl --device 1) mines the gfx1036 at 3.12 MH/s; the 5090 keeps mining; nobody was woken, PC 1's app was not restarted; a 10-second rescan probe (pnputil /scan-devices, the USB4 router status) was granted | the second eGPU link fault today: reseat the USB4 cable and the eGPU's power; the 0.3.10 hot-plug code shows the card as "removed" and picks it up again without a restart | +| 5 October, 21:22:59 to 22:21 | the card dropped again at 21:22:59 (the third drop), was back and used by the hot-table (21:31 to 21:35), the era (21:46 to 21:51 and 22:16 to 22:21) and the mixer (22:00 to 22:04) jobs as gfx1201, and was gone again by 22:24:09 (the fourth drop; no Sonnet or USB4 router device) | the OpenCL worker falls to the gfx1036 (--device 1) when the card is gone; every measurement that names gfx1201 ran while it was present | the link flaps on a scale of tens of minutes: reseat the USB4 cable and the eGPU's power, try another port or cable; the AMD clock and power sweep is owed on this | | 5 October, by 21:09 | the 9070 XT is back on PC 1's bus: the reproducible-benchmark package's OpenCL worker listed it as opencl:1 and the app switched it off and on through api/cards; no restart, nobody touched the box | the app's AMD worker returns to it at the next prepare | the link drops and returns by itself; the reseat is still worth doing in the morning | | 5 October, by 21:22:59 | the third drop: no Sonnet or USB4 Router (2.0) device present, the display list shows only the gfx1036 and the 5090 | the app lists nvidia:0 and amd:1:gfx1036; the 5090 keeps mining at 124.5 MH/s (app log STATUS lines through 21:23:26Z) | the link is flapping: reseat the USB4 cable and the eGPU's power in the morning, and consider a different USB4 port or cable; every AMD measurement tonight runs on the gfx1036 fallback unless the card is present at the job's own probe | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index dbe668240..3b86457a1 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -543,3 +543,7 @@ The node agent's owed item is an app gap, confirmed in the 0.3.11 app branch: en Job run-ca2-era-pc1b-20261005 (22:16 to 22:21Z, 304 s, both cards restored, app 0.3.10, the 9070 XT present): the seven final-class packs' 2^24 fingerprints equal on the 5090, the 9070 XT and the M5 Max (mx8-devnet-epoch0 90f794dd556f7a3b; era-0 8e8e070db4eea52d, era-1 891c01b8563bb47e, era-2 e54279fed2831b5d, era-3 77e0ba8abbd0ae62, era-4 d898d8f4f2e7684b, era-5 a6927db380f7efb2); G2: 1,024 of 1,024 per card on era-0 and the pinned pack re-hashed by the Rust verifier. Rates on the final class (MH/s): 5090 135.90 to 137.70 (spread 1.3%), 9070 XT 18.59 to 19.18 (3.1%), M5 Max 27.85 to 27.98 (0.5%); the pinned pack 136.10 / 18.90 / 27.80. Spec 1.17 carries these fingerprints. Gates now: G1, G2, G3, G4, G6 green; G4b (a) done, (b) the Metal gate run pending; G5 at the ship's build. PC 1 goes to the AMD sweep on its presence probe. 22:25. Consequences C29 and D11 applied. C29: the litepaper's verifier line now reads 2.1 ms per warp for class v3 on one M5 Max core at load average 5.5 (the fixed crate; worst cold 2.15), 3.4x the v2 verifier's 0.61 ms, the gate leaving 4.8x (4.6x on the worst cold unit); the "about 1.3 ms quiet" scaling is struck: it came from the slow binary's 2.79 ms session (21:40), and the fixed binary's session (22:07) matched readwidth's quiet v2 figure within 1%, so the fixed numbers are near-quiet. D11: "and the bounty" struck from the hero and the abstract, "standing" from level 1; the copy says "a bounty follows the external review"; the bounty is named only once escrowed (funding.md rule 3; USD 50,000 not funded); the project lead decides (D11). The level 3 table and the numbers-page entry now carry the final-class rates (5090 135.90 to 137.70, 9070 XT 18.59 to 19.18, M5 Max 27.85 to 27.98). + +## 22:25 the fourth eGPU drop; the era branch final; PC 1 to Ember + +The 9070 XT is absent again at 22:24:09Z (relay probe #230: no Sonnet or USB4 router device), the fourth drop today, after serving the hot-table, era (twice) and mixer jobs between 21:31 and 22:21; the AMD clock and power sweep did not start and its rows are owed with this time and reason; the hardware-events table in the rollout plan carries the flapping for the morning. ca2-era is final at 95955c3 on ca2-v3 5eb2331 (six commits: the implementation, the Mac rows, the PC round 1 rows, the G2 hash-bound --count flag, the final-class re-export, the round 2 rows; 53 + 30 tests). "go PC 1 build" given to Ember Tune (its two fetches and the 8-minute Windows app build), then its 8-minute run on the 5090 baseline. PC 2: floor-build-2 running (to about 23:50), then the aggregation-cost re-run, then the prover-floor sweep. From 79880291b994ce442948d4b007ee115c5917ae93 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:27:04 +0000 Subject: [PATCH 123/131] Counter ASIC 2.0 status 22:27: the integration merge plan and the ship order --- docs/plans/counter-asic-2-status.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 3b86457a1..fdfd0d4ac 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -547,3 +547,7 @@ Job run-ca2-era-pc1b-20261005 (22:16 to 22:21Z, 304 s, both cards restored, app ## 22:25 the fourth eGPU drop; the era branch final; PC 1 to Ember The 9070 XT is absent again at 22:24:09Z (relay probe #230: no Sonnet or USB4 router device), the fourth drop today, after serving the hot-table, era (twice) and mixer jobs between 21:31 and 22:21; the AMD clock and power sweep did not start and its rows are owed with this time and reason; the hardware-events table in the rollout plan carries the flapping for the morning. ca2-era is final at 95955c3 on ca2-v3 5eb2331 (six commits: the implementation, the Mac rows, the PC round 1 rows, the G2 hash-bound --count flag, the final-class re-export, the round 2 rows; 53 + 30 tests). "go PC 1 build" given to Ember Tune (its two fetches and the 8-minute Windows app build), then its 8-minute run on the 5090 baseline. PC 2: floor-build-2 running (to about 23:50), then the aggregation-cost re-run, then the prover-floor sweep. + +## 22:27 the integration merge, in the node agent's hands; the ship order + +A dry run of ca2-v3 into master in a scratch worktree conflicts in docs/bench-log.md (append-only, both kept) and proto-opencl/host.c, where taking ca2-v3's side would drop master's 0.3.10 hunks (the PCI topology and duplicate-platform detection, the select read-back): it must be resolved by hand keeping both. ca2-v3 also lacks readwidth's last three commits (d0018cf the OpenCL scratch local-buffer rule, e752fc7 read-width.md and the overflow fix, 30ff674 the per-watt rows). So the node agent, as ca2-v3's owner, merges readwidth 30ff674 and origin/master (cde561c, 1f0d62c) into ca2-v3 beside gate run 4, re-runs the igneum-pow suite, the packfile test, test-generic.sh on Apple OpenCL and the fork check, and sends the tip. ca2-coord (docs, spec, site, evidence, the plans) stays on its a9e002c base: its code files are untouched readwidth copies, so its merge onto that tip takes ca2-v3's code and brings only the documents (a rebase attempt replayed readwidth's own commits and was aborted). Ship order: master <- ca2-v3' <- ca2-coord <- the tooling and analysis branches (consequences, bash-body-check, amd-prove, card-lifetime, proving-methods, asic-history when it lands), the app branch proving-v1 e0de2ab (on 5b0d54f, in master), the fork ca2-v3-node 89dfcb95 (on 21d4c73c, release-0.3.10's tip). The ASIC-history agent (a202a09dcd24ba1d3) resumed at 22:52 (its clock) in ../igneum-wt-asic-history; its ranked additions are Counter ASIC 3.0's input after the publish. From 50e80cb11e8c9a5b3a78fd35aa8c193e3ac6fda7 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:27:36 +0000 Subject: [PATCH 124/131] C1: H = 210,000 lands about 19:00Z on 6 October (18:50 to 19:35), not 19:50 --- docs/plans/counter-asic-2-rollout.md | 2 +- docs/plans/counter-asic-2-status.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 82d2a611d..086ee5bd2 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -99,7 +99,7 @@ Public lines to fix on the integration branch (card-lifetime table 3): `site/ind ## 7a. A dated constraint from the consequences review (C1) -The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H. +The fee switch H = 210,000 arrives about 19:00Z on 6 October (18:50 to 19:35Z: the devnet read DAA 136,578 at 22:24:31Z on 5 October, 0.909 blocks/s over the stats window and 1.005 DAA/s since 15:40Z; the 19:50Z in fee-switch-devnet.md is up to an hour late). The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H. ## 8a. Proving v1 rides with it diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index fdfd0d4ac..21d0bc55a 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -144,7 +144,7 @@ Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout ## 20:29 C1 decided for the morning; one packfile.h fix for 0.3.11 -C1 (the fee switch H = 210,000 at about 19:50Z on 6 October; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before 19:50Z. The proving agent recommends (b) unless (a) is certain; the check decides it. +C1 (the fee switch H = 210,000 at about 19:00Z on 6 October, 18:50 to 19:35Z by the 22:24Z DAA read; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before about 19:00Z (18:50 to 19:35Z). The proving agent recommends (b) unless (a) is certain; the check decides it. The attempt-0 packfile.h bug (the era agent's find) is the one that took the fleet down at 18:23Z (epoch 34, attempt 1); it is already fixed attempt-aware on pack-loop af983a7 and merged into the 0.3.10 tree. Rule for 0.3.11: one derivation, one test set: the era and cache agents build their workers on that packfile.h and keep only tests that add a case; the host.cu / host.c mh_word derivation fix (new, the era agent's) stays with its own test. From de77871a848c9d0f504daab063fb9799742fe4a4 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:30:41 +0000 Subject: [PATCH 125/131] C31: the litepaper's 12 GB clause struck, the dataset schedule written as the gate 1 proposal on the site and litepaper, the merge rule for evidence row 16 and the proving sentences --- docs/plans/counter-asic-2-rollout.md | 2 +- site/index.html | 4 ++-- site/litepaper.html | 4 ++-- 3 files changed, 5 insertions(+), 5 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 086ee5bd2..965204f26 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -105,7 +105,7 @@ The fee switch H = 210,000 arrives about 19:00Z on 6 October (18:50 to 19:35Z: t the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL e0de2ab on 5b0d54f (a223ca9 the resume fix, e0de2ab the Metal --prepare-packs fix; docs-only commits between) (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. -Branches merged into the 0.3.11 main tree beside ca2-v3 and ca2-coord (tooling, no consensus): consequences (the reviewer's rows), bash-body-check (7adb1ca, 6805125: tools/ci/bash-body-check.sh with fixtures and a self-test in ci.yml, the convention paragraph in packaging/README-ship.md, `publish-jobs.sh add --kind run` running the bash-body and prover-socket checks before signing; add/add with proving-v1 c2544be on tools/ci/prover-socket-check.sh: take bash-body-check's superset and proving-v1's ci.yml step; tools/amd-prove/pc1-cpu-prove.ps1 goes on the allow list, it starts no GPU server), amd-prove (f1d7a7d), card-lifetime (1fecfe2). Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. +Merge rule for the ship (C31): the app branch proving-v1 (e0de2ab) rewrote docs/evidence.md row 16 (WITHDRAWN, the 24 GB measurement) and the litepaper's proving sentences; ca2-coord carries the older row 16 and its own litepaper edits, so at the merge take proving-v1's row 16 and its proving sentences, and ca2-coord's everything else; the stale row must not win by accident. Branches merged into the 0.3.11 main tree beside ca2-v3 and ca2-coord (tooling, no consensus): consequences (the reviewer's rows), bash-body-check (7adb1ca, 6805125: tools/ci/bash-body-check.sh with fixtures and a self-test in ci.yml, the convention paragraph in packaging/README-ship.md, `publish-jobs.sh add --kind run` running the bash-body and prover-socket checks before signing; add/add with proving-v1 c2544be on tools/ci/prover-socket-check.sh: take bash-body-check's superset and proving-v1's ci.yml step; tools/amd-prove/pc1-cpu-prove.ps1 goes on the allow list, it starts no GPU server), amd-prove (f1d7a7d), card-lifetime (1fecfe2). Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. ## 9. After the publish: Counter ASIC 3.0 diff --git a/site/index.html b/site/index.html index 5cdbb14c4..94ac604a1 100644 --- a/site/index.html +++ b/site/index.html @@ -440,7 +440,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var( Who minesCPUsGPUs Program changesEvery hashEvery hour - Memory2 GB, fixed2 GB at genesis, doubling at years 4, 12 and 28 + Memory2 GB, fixed2 GB at genesis, growing on a genesis schedule (proposed steps at years 4, 12 and 28) Verified byAny CPUAny CPU @@ -458,7 +458,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var(

Mine

-

Any 4 GB card at launch, 8 GB from year 4, approximate. 80% of each block to its finder.

+

Any 4 GB card at launch, 8 GB from about year 4, approximate. 80% of each block to its finder.

About the miner
diff --git a/site/litepaper.html b/site/litepaper.html index 9952d8b6f..6522c9114 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -433,7 +433,7 @@ body.all .pager{display:none} Hardware it is built forCPUs. GPUs run it badly on purposeGPUs. Any card, any vendor. Bit-exact on Apple, NVIDIA and AMD, measured Random programPer hash, interpreted in a virtual machinePer hour, compiled to native GPU code. Per hash, the 128 dataset addresses change with the nonce - DatasetAbout 2 GB, the same size since 2019, approximate2 GB at genesis, doubling at years 4, 12 and 28 under the step schedule; a 4 GB card mines about four years, an 8 GB card past a decade + DatasetAbout 2 GB, the same size since 2019, approximate2 GB at genesis, growing on a schedule fixed at genesis (the proposed steps: years 4, 12 and 28, decided at gate 1); a 4 GB card mines about four years, an 8 GB card about twelve Light verification256 MB cache on a CPU, milliseconds256 MB cache on a CPU (512 MB from year 4), one warp under 10 ms, the gate. Measured 2.1 ms on one Apple M5 Max core for class v3 (3.4x class v2's 0.61 ms); a 2019-class core not yet Changes over timeNone. A fixed design, unchanged for seven yearsA new program every hour, its memory pattern with it; era draws and reserved families on a schedule fixed at genesis. Nobody touches it Seed grindingNot applicable, the program comes from the hash inputClosed by a verifiable delay between seed and program @@ -559,7 +559,7 @@ body.all .pager{display:none}

The honest bear-market case rests on cost. A miner's card is already running and the power is often domestic, so Igneum miners' marginal cost in the proving market is close to power, which is an edge over data-centre provers and nothing more.

Hardware

-

The dataset starts at 2 GB and doubles on a step schedule fixed at genesis (years 4, 12 and 28, the average of half a gigabyte a year), so a 4 GB card mines for about four years and an 8 GB card for about twelve, approximate. 12 GB or more proves full shards. NVIDIA and AMD both work, because the mining program is generated for the architecture both share and the proof system is hash-based. Apple's chips are GPUs with unified memory, so Macs mine too, at about a fifth of a flagship card: Measured, 26.7 against 123 million hashes a second, an Apple M5 Max beside an RTX 5090 on the live devnet, 4 October 2026. A Mac is a poor miner per dollar. There is no CPU mining lane, on purpose, because CPU mining is what botnets farm. Nodes, wallets and exchanges need no GPU at all.

+

The dataset starts at 2 GB and grows on a schedule fixed at genesis; the proposed schedule, decided at gate 1, doubles it at years 4, 12 and 28 (the average of half a gigabyte a year), so a 4 GB card mines for about four years and an 8 GB card for about twelve, approximate. NVIDIA and AMD both work, because the mining program is generated for the architecture both share and the proof system is hash-based. Apple's chips are GPUs with unified memory, so Macs mine too, at about a fifth of a flagship card: Measured, 26.7 against 123 million hashes a second, an Apple M5 Max beside an RTX 5090 on the live devnet, 4 October 2026. A Mac is a poor miner per dollar. There is no CPU mining lane, on purpose, because CPU mining is what botnets farm. Nodes, wallets and exchanges need no GPU at all.

What a miner's hour looks like

The card hashes the lottery continuously. When the client sees a shard or an external job it can win, it switches the card to proving for a few seconds, posts the proof, and goes back to hashing. The client does the switching and the miner sees one balance.

The protocol carries no fee: no dev fund, no cut to any team. Ember, the miner software, takes an optional 1% dev fee, the way other GPU miners do. One block template in 100 is requested with the dev address instead of yours, by a counter, not a random draw, so it is exactly 1 in 100 and anyone can check it from the source or from the chain. One flag turns it off (--dev-fee 0, a switch in the app, a line in the HiveOS config). The miner prints the fee and the address when it starts. Any other client is welcome.

From 1bec7efddf00846a11f8df62dfe36daa65faafff Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:30:56 +0000 Subject: [PATCH 126/131] Counter ASIC 2.0 status 22:30: C31, Ember's run, floor-build-3 --- docs/plans/counter-asic-2-status.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 21d0bc55a..3b79dbd7d 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -551,3 +551,5 @@ The 9070 XT is absent again at 22:24:09Z (relay probe #230: no Sonnet or USB4 ro ## 22:27 the integration merge, in the node agent's hands; the ship order A dry run of ca2-v3 into master in a scratch worktree conflicts in docs/bench-log.md (append-only, both kept) and proto-opencl/host.c, where taking ca2-v3's side would drop master's 0.3.10 hunks (the PCI topology and duplicate-platform detection, the select read-back): it must be resolved by hand keeping both. ca2-v3 also lacks readwidth's last three commits (d0018cf the OpenCL scratch local-buffer rule, e752fc7 read-width.md and the overflow fix, 30ff674 the per-watt rows). So the node agent, as ca2-v3's owner, merges readwidth 30ff674 and origin/master (cde561c, 1f0d62c) into ca2-v3 beside gate run 4, re-runs the igneum-pow suite, the packfile test, test-generic.sh on Apple OpenCL and the fork check, and sends the tip. ca2-coord (docs, spec, site, evidence, the plans) stays on its a9e002c base: its code files are untouched readwidth copies, so its merge onto that tip takes ca2-v3's code and brings only the documents (a rebase attempt replayed readwidth's own commits and was aborted). Ship order: master <- ca2-v3' <- ca2-coord <- the tooling and analysis branches (consequences, bash-body-check, amd-prove, card-lifetime, proving-methods, asic-history when it lands), the app branch proving-v1 e0de2ab (on 5b0d54f, in master), the fork ca2-v3-node 89dfcb95 (on 21d4c73c, release-0.3.10's tip). The ASIC-history agent (a202a09dcd24ba1d3) resumed at 22:52 (its clock) in ../igneum-wt-asic-history; its ranked additions are Counter ASIC 3.0's input after the publish. + +22:30. C31 applied (a7be43f): the litepaper's "12 GB or more proves full shards" struck (one page, one number: 24 GB); the dataset's step schedule written as the gate 1 proposal on the site and litepaper (D4 is the project lead's); the merge rule "take proving-v1's evidence row 16 and its proving sentences" in the rollout plan. PC 1: Ember's build done (build-20261005-222558, 97 s, 6 outputs verified), its tune run on the 5090 has the go. PC 2: floor-build-3 (with the pinned Go toolchain; builds 1 and 2 failed on a missing go, the first unreported by a log-path bug) runs, 25 to 45 minutes expected. From de847c6cae06b9c77ea3e69d0eb7d409f858cdd8 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:31:11 +0000 Subject: [PATCH 127/131] Site: one wording for the dataset schedule in all five places (the proposed schedule, fixed at the testnet genesis), miner.html brought to it --- site/index.html | 4 ++-- site/litepaper.html | 4 ++-- site/miner.html | 2 +- 3 files changed, 5 insertions(+), 5 deletions(-) diff --git a/site/index.html b/site/index.html index 94ac604a1..35e43ae9a 100644 --- a/site/index.html +++ b/site/index.html @@ -440,7 +440,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var( Who minesCPUsGPUs Program changesEvery hashEvery hour - Memory2 GB, fixed2 GB at genesis, growing on a genesis schedule (proposed steps at years 4, 12 and 28) + Memory2 GB, fixed2 GB, growing (the proposed schedule, fixed at the testnet genesis: 2 GB, doubling at years 4, 12 and 28) Verified byAny CPUAny CPU @@ -458,7 +458,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var(

Mine

-

Any 4 GB card at launch, 8 GB from about year 4, approximate. 80% of each block to its finder.

+

Any 4 GB card at launch, 8 GB from about year 4 under the proposed schedule, fixed at the testnet genesis: 2 GB, doubling at years 4, 12 and 28, approximate. 80% of each block to its finder.

About the miner
diff --git a/site/litepaper.html b/site/litepaper.html index 6522c9114..886652dba 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -433,7 +433,7 @@ body.all .pager{display:none} Hardware it is built forCPUs. GPUs run it badly on purposeGPUs. Any card, any vendor. Bit-exact on Apple, NVIDIA and AMD, measured Random programPer hash, interpreted in a virtual machinePer hour, compiled to native GPU code. Per hash, the 128 dataset addresses change with the nonce - DatasetAbout 2 GB, the same size since 2019, approximate2 GB at genesis, growing on a schedule fixed at genesis (the proposed steps: years 4, 12 and 28, decided at gate 1); a 4 GB card mines about four years, an 8 GB card about twelve + DatasetAbout 2 GB, the same size since 2019, approximate2 GB, growing (the proposed schedule, fixed at the testnet genesis: 2 GB, doubling at years 4, 12 and 28); a 4 GB card mines about four years, an 8 GB card about twelve, approximate Light verification256 MB cache on a CPU, milliseconds256 MB cache on a CPU (512 MB from year 4), one warp under 10 ms, the gate. Measured 2.1 ms on one Apple M5 Max core for class v3 (3.4x class v2's 0.61 ms); a 2019-class core not yet Changes over timeNone. A fixed design, unchanged for seven yearsA new program every hour, its memory pattern with it; era draws and reserved families on a schedule fixed at genesis. Nobody touches it Seed grindingNot applicable, the program comes from the hash inputClosed by a verifiable delay between seed and program @@ -559,7 +559,7 @@ body.all .pager{display:none}

The honest bear-market case rests on cost. A miner's card is already running and the power is often domestic, so Igneum miners' marginal cost in the proving market is close to power, which is an edge over data-centre provers and nothing more.

Hardware

-

The dataset starts at 2 GB and grows on a schedule fixed at genesis; the proposed schedule, decided at gate 1, doubles it at years 4, 12 and 28 (the average of half a gigabyte a year), so a 4 GB card mines for about four years and an 8 GB card for about twelve, approximate. NVIDIA and AMD both work, because the mining program is generated for the architecture both share and the proof system is hash-based. Apple's chips are GPUs with unified memory, so Macs mine too, at about a fifth of a flagship card: Measured, 26.7 against 123 million hashes a second, an Apple M5 Max beside an RTX 5090 on the live devnet, 4 October 2026. A Mac is a poor miner per dollar. There is no CPU mining lane, on purpose, because CPU mining is what botnets farm. Nodes, wallets and exchanges need no GPU at all.

+

The dataset starts at 2 GB and grows (the proposed schedule, fixed at the testnet genesis: 2 GB, doubling at years 4, 12 and 28, the average of half a gigabyte a year), so a 4 GB card mines for about four years and an 8 GB card for about twelve, approximate. NVIDIA and AMD both work, because the mining program is generated for the architecture both share and the proof system is hash-based. Apple's chips are GPUs with unified memory, so Macs mine too, at about a fifth of a flagship card: Measured, 26.7 against 123 million hashes a second, an Apple M5 Max beside an RTX 5090 on the live devnet, 4 October 2026. A Mac is a poor miner per dollar. There is no CPU mining lane, on purpose, because CPU mining is what botnets farm. Nodes, wallets and exchanges need no GPU at all.

What a miner's hour looks like

The card hashes the lottery continuously. When the client sees a shard or an external job it can win, it switches the card to proving for a few seconds, posts the proof, and goes back to hashing. The client does the switching and the miner sees one balance.

The protocol carries no fee: no dev fund, no cut to any team. Ember, the miner software, takes an optional 1% dev fee, the way other GPU miners do. One block template in 100 is requested with the dev address instead of yours, by a counter, not a random draw, so it is exactly 1 in 100 and anyone can check it from the source or from the chain. One flag turns it off (--dev-fee 0, a switch in the app, a line in the HiveOS config). The miner prints the fee and the address when it starts. Any other client is welcome.

diff --git a/site/miner.html b/site/miner.html index f170bf4db..ed3887367 100644 --- a/site/miner.html +++ b/site/miner.html @@ -469,7 +469,7 @@ pre b{color:var(--molten);font-weight:500}

Miner: Custom. Installation URL: https://dl.igneum.network/dl/public/igneum-hive-0.3.9.tar.gz. Miner name igneum, wallet and worker 0x<your 40-hex payout address>.%WORKER_NAME%, pool URL grpc://<your node>:26610 or local for the bundled node, extra config DEV_FEE=1 IDENTITIES=8 WORKER=auto VOTE=1. The Linux button above is the same archive: the Linux miner, node and both GPU workers. Hive itself is untested so far. Report what breaks.

sha256 7a58a30fd47c9ecb3d4aeaa0c0a464550f7e4b2b33eacf87512f1b72164c829e

-

Testnet coin for testing: the faucet sends 10 IGN per address per day. Any 4 GB card. The dataset starts at 2 GB and grows by half a gigabyte a year, approximate. Updates arrive signed; the app installs them itself. The wallet reads this machine's node when the miner is installed.

+

Testnet coin for testing: the faucet sends 10 IGN per address per day. Any 4 GB card at launch. The dataset starts at 2 GB and grows (the proposed schedule, fixed at the testnet genesis: 2 GB, doubling at years 4, 12 and 28), so a 4 GB card mines for about four years, approximate. Updates arrive signed; the app installs them itself. The wallet reads this machine's node when the miner is installed.

Read more in the litepaper
From cf6e6f82b400b4834be3511c5c212ab9137950ee Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:32:44 +0000 Subject: [PATCH 128/131] Counter ASIC 2.0: C32 rules (PC 2's update-now after the floor build; re-fetch after every app update); status 22:32 --- docs/plans/counter-asic-2-rollout.md | 5 +++++ docs/plans/counter-asic-2-status.md | 2 ++ 2 files changed, 7 insertions(+) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 965204f26..f881f4d8d 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -37,6 +37,11 @@ The same object goes verbatim into the override files of Mac node 1, the observe 5. HiveOS: `packaging/hive/make-hive-package.sh` republished with the v3 `igneum-miner` and workers, same version string as the apps. 6. The watch: before `N4 - 1,800` both PCs on the new app version (STATUS lines); at the boundary every miner's `prepare` of the v3 pack (the hot-swap entry's shape) and the first v3 block's program id on the observer; the hash rate per card against the measured v3 numbers of `docs/plans/counter-asic-2.md`'s table; zero `pack refused` lines; the CPU verifier time per block in the node log against the measured ms per warp. +### 4a. Two rules for the publish and every PC job (C32) + +- The 0.3.11 update-now goes to PC 2 only after the prover-floor agent's server build (floor-build-3, under /opt/igneum-floor in WSL2, published 22:27 UTC, 25 to 90 minutes) has closed: an app restart ends the running job. The update-now takes a machine list (as 0.3.10's did): the Mac, the laptop and PC 1 first, PC 2 last. +- Re-fetch after every app update: an app update clears the jobs folder (the 0.3.10 install at 21:49 UTC took PC 1's AMD kit with it), so every fetch-then-run pair re-publishes its fetch after an update, and every run playbook opens with a presence check of its kit that fails with "kit missing: republish the fetch after the app update". + ## 5. Rollback Before N4: remove the field on every node and restart; nothing has happened (the digest flips back, so every node at once). After N4: there is no rollback by restart, because blocks mined under v3 verify only under v3. A rollback is a second height switch back to v2 at a later epoch, carried the same way. This is why the measurements of the plan come first. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 3b79dbd7d..71b6c1873 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -553,3 +553,5 @@ The 9070 XT is absent again at 22:24:09Z (relay probe #230: no Sonnet or USB4 ro A dry run of ca2-v3 into master in a scratch worktree conflicts in docs/bench-log.md (append-only, both kept) and proto-opencl/host.c, where taking ca2-v3's side would drop master's 0.3.10 hunks (the PCI topology and duplicate-platform detection, the select read-back): it must be resolved by hand keeping both. ca2-v3 also lacks readwidth's last three commits (d0018cf the OpenCL scratch local-buffer rule, e752fc7 read-width.md and the overflow fix, 30ff674 the per-watt rows). So the node agent, as ca2-v3's owner, merges readwidth 30ff674 and origin/master (cde561c, 1f0d62c) into ca2-v3 beside gate run 4, re-runs the igneum-pow suite, the packfile test, test-generic.sh on Apple OpenCL and the fork check, and sends the tip. ca2-coord (docs, spec, site, evidence, the plans) stays on its a9e002c base: its code files are untouched readwidth copies, so its merge onto that tip takes ca2-v3's code and brings only the documents (a rebase attempt replayed readwidth's own commits and was aborted). Ship order: master <- ca2-v3' <- ca2-coord <- the tooling and analysis branches (consequences, bash-body-check, amd-prove, card-lifetime, proving-methods, asic-history when it lands), the app branch proving-v1 e0de2ab (on 5b0d54f, in master), the fork ca2-v3-node 89dfcb95 (on 21d4c73c, release-0.3.10's tip). The ASIC-history agent (a202a09dcd24ba1d3) resumed at 22:52 (its clock) in ../igneum-wt-asic-history; its ranked additions are Counter ASIC 3.0's input after the publish. 22:30. C31 applied (a7be43f): the litepaper's "12 GB or more proves full shards" struck (one page, one number: 24 GB); the dataset's step schedule written as the gate 1 proposal on the site and litepaper (D4 is the project lead's); the merge rule "take proving-v1's evidence row 16 and its proving sentences" in the rollout plan. PC 1: Ember's build done (build-20261005-222558, 97 s, 6 outputs verified), its tune run on the 5090 has the go. PC 2: floor-build-3 (with the pinned Go toolchain; builds 1 and 2 failed on a missing go, the first unreported by a log-path bug) runs, 25 to 45 minutes expected. + +22:32. C32: an app update clears the jobs folder (the 0.3.10 install took PC 1's AMD kit, fetched at 21:23:59Z; the 21:05 reading "cleared by fetch jobs" was wrong), so two rules enter the rollout plan (4a): the 0.3.11 update-now reaches PC 2 only after floor-build-3 closes (PC 2 last on the machine list), and every fetch-then-run pair re-fetches after an update, with a kit presence check at the top of every run playbook. From 91713112daa5f20126dc2ee4993b4dcef4f3ef09 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:33:15 +0000 Subject: [PATCH 129/131] Counter ASIC 2.0: gate G4b green (the Metal miner across a v3 boundary; the serveDataset outage fixed); status 22:33 --- docs/plans/counter-asic-2-status.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 71b6c1873..5b22995e6 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -555,3 +555,9 @@ A dry run of ca2-v3 into master in a scratch worktree conflicts in docs/bench-lo 22:30. C31 applied (a7be43f): the litepaper's "12 GB or more proves full shards" struck (one page, one number: 24 GB); the dataset's step schedule written as the gate 1 proposal on the site and litepaper (D4 is the project lead's); the merge rule "take proving-v1's evidence row 16 and its proving sentences" in the rollout plan. PC 1: Ember's build done (build-20261005-222558, 97 s, 6 outputs verified), its tune run on the 5090 has the go. PC 2: floor-build-3 (with the pinned Go toolchain; builds 1 and 2 failed on a missing go, the first unreported by a log-path bug) runs, 25 to 45 minutes expected. 22:32. C32: an app update clears the jobs folder (the 0.3.10 install took PC 1's AMD kit, fetched at 21:23:59Z; the 21:05 reading "cleared by fetch jobs" was wrong), so two rules enter the rollout plan (4a): the 0.3.11 update-now reaches PC 2 only after floor-build-3 closes (PC 2 last on the machine list), and every fetch-then-run pair re-fetches after an update, with a kit presence check at the top of every run playbook. + +## 22:33 gate G4b GREEN: the Mac mines v3 end to end; a second Metal outage found and fixed + +Gate 4 (22:25 to 22:31Z, a real Metal miner through the miner's --prepare-packs flow on the fast-time network, igneum-bench from ca2-v3 00c55aa): three v3 prepares with the pack dir and the class and era tokens; the worker's v3 prepared lines (252.5 ms the first: program 55.5, dataset 196.9, cache fill 0.9, build 41.1; then 51 and 46 ms with the day resident); every swap "with no pause"; 124 blocks accepted on v3, cpu re-check mismatched 0, need 0, no mismatch or refusal; chain 182 / 123 across the boundary, 0 rejected, one sink. Found and fixed before the run: the Metal worker's serveDataset built every day with the Swift version 2 construction keyed by day only, so a v3 program would have hashed over an x1 dataset and every found would have been refused by the CPU re-check, the same fleet-outage class as the app's missing flag; now a v3 prepare builds the day from the pack's memhard.metal and the store keys datasets by (day, class, era) (00c55aa). Two Mac outages caught by one gate run that the CPU-miner gates could not see; the rule for the next cut: every worker path (Metal, CUDA, OpenCL) mines across a boundary in the gate network before a class change ships. The integration merges (readwidth 30ff674, then origin/master) are in progress on ca2-v3 with the conflicts resolved by hand keeping both sides. + +Gates: G1, G2, G3, G4, G4b, G6 GREEN; G5 at the ship's build step. The ship waits only on the merged tip and its checks. From 607587dc2a4627b242b46da173a7339d6b5c01c2 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:33:42 +0000 Subject: [PATCH 130/131] Counter ASIC 2.0: H at about 18:45Z, the G4b row's state text, the floor build green and the sweep go; status 22:33 --- docs/plans/counter-asic-2-rollout.md | 4 ++-- docs/plans/counter-asic-2-status.md | 4 +++- 2 files changed, 5 insertions(+), 3 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index f881f4d8d..96767d920 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -74,7 +74,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms Run 3 PASS on the FINAL class (22:15:04 to 22:19:44Z, binaries from ca2-v3 d233fa1 = x8 + era + the verifier fix, fork 89dfcb95): 182 / 122 blocks around DAA 180, 304 in all; v3 ids e3 a6523b90cff501e3, e4 bc811b3c4b8b1ced, e5 e784541f19cdebe5 on all three miners; epoch 0's v2 id 8f8806638d59850f the same in all three runs; rejected 0/0/0; one sink a9ce45df8beeaf13 at 303/303/303; 3 of 3 switch lines; cache ready v2 179 ms, first v3 191 ms. Summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2215Z-era-mx8.json` | GREEN (runs 1, 2 and 3; run 3 on the final class) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G4b | the Mac mines v3: a real Metal miner across a v3 boundary through the miner's `--prepare-packs` flow, and the app passes that flag to the Metal worker | Found 22:21 UTC: `app/igneum-app/src/engine.rs` `miner_args` pushes `--prepare-packs` only when `card.worker != "Metal"` (line 1468 on a223ca9), so the Mac app never hands its Metal worker a prepare pack, and under class v3 the Metal worker compiles v3 only from a prepared pack (servePackProgram): at the first v3 epoch every Mac would answer `need` lines and stop, the 18:23Z outage class. Two closes before the ship, both in hand: (a) the app fix on 0.3.11's app branch (push the flag for every worker with the platform's path separator, a unit test on `miner_args` for a Metal card; assigned to the proving agent on top of a223ca9); (b) gate run 4: the fast-time network with one real Metal miner on the Mac across the activation (the prepare lines on both sides, the worker's `prepared` line for the v3 pack, found or accepted blocks on v3, no `need` or mismatch line; assigned to the node agent). If (a) is not in the app tree at the cut, the ship does not go: a Mac that cannot mine v3 at activation is a fleet outage, and the activation height (tip + 14,400) is not far enough to carry the fix in 0.3.12 safely | (a) DONE: app proving-v1 e0de2ab (prepare_packs_arg() for every worker, `packs/prepare` on macOS; start_miner gives the Metal worker the app data folder as cwd, which was None; unit test every_worker_gets_the_prepare_directory_in_the_platform_form; 113 + 27 + 8 passed). (b) gate run 4 pending | +| G4b | the Mac mines v3: a real Metal miner across a v3 boundary through the miner's `--prepare-packs` flow, and the app passes that flag to the Metal worker | Found 22:21 UTC: `app/igneum-app/src/engine.rs` `miner_args` pushes `--prepare-packs` only when `card.worker != "Metal"` (line 1468 on a223ca9), so the Mac app never hands its Metal worker a prepare pack, and under class v3 the Metal worker compiles v3 only from a prepared pack (servePackProgram): at the first v3 epoch every Mac would answer `need` lines and stop, the 18:23Z outage class. Two closes before the ship, both in hand: (a) the app fix on 0.3.11's app branch (push the flag for every worker with the platform's path separator, a unit test on `miner_args` for a Metal card; assigned to the proving agent on top of a223ca9); (b) gate run 4: the fast-time network with one real Metal miner on the Mac across the activation (the prepare lines on both sides, the worker's `prepared` line for the v3 pack, found or accepted blocks on v3, no `need` or mismatch line; assigned to the node agent). If (a) is not in the app tree at the cut, the ship does not go: a Mac that cannot mine v3 at activation is a fleet outage, and the activation height (tip + 14,400) is not far enough to carry the fix in 0.3.12 safely | GREEN. (b) gate 4 (22:25 to 22:31Z, a real Metal miner on node 0, igneum-bench from ca2-v3 00c55aa): three v3 PREPARE lines with the pack dir and class=v3 era=; the worker's v3 `prepared` lines (252.5 ms the first: program 55.5, dataset 196.9, cache fill 0.9, build 41.1; then 51 and 46 ms with the day resident); every swap "with no pause"; 124 blocks accepted on v3 (301 in the run), cpu re-check mismatched 0, need 0, no mismatch or refusal, no exit 42 or 44; chain 182 / 123 across DAA 180, 0 rejected, one sink. A second outage found and fixed before the run: the Metal worker's serveDataset built every day with the Swift version 2 construction keyed by day only, so a v3 program would have hashed over an x1 dataset; now a v3 prepare builds the day from the pack's memhard.metal and the store keys datasets by (day, class, era), commit 00c55aa | | G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 build-20261005-221237 (published 22:12:37Z: main d233fa1 = era + cache + the mixer fix with V3_CLASS = MX8, fork 89dfcb95) ran the five crates and the app on the final tree: every stage ok (22:12:37 to 22:15:44Z): kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8, 0 failed. With job 2 (kaspa-consensus alone 97) G6 is GREEN on the final tree | GREEN | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. @@ -104,7 +104,7 @@ Public lines to fix on the integration branch (card-lifetime table 3): `site/ind ## 7a. A dated constraint from the consequences review (C1) -The fee switch H = 210,000 arrives about 19:00Z on 6 October (18:50 to 19:35Z: the devnet read DAA 136,578 at 22:24:31Z on 5 October, 0.909 blocks/s over the stats window and 1.005 DAA/s since 15:40Z; the 19:50Z in fee-switch-devnet.md is up to an hour late). The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H. +The fee switch H = 210,000 arrives about 18:45Z on 6 October (DAA 137,041 at 22:32:54Z on 5 October; 1.002 DAA/s averaged since 15:40Z; the 19:50Z in fee-switch-devnet.md is an hour late). The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H. ## 8a. Proving v1 rides with it diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 5b22995e6..54a22d5fc 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -144,7 +144,7 @@ Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout ## 20:29 C1 decided for the morning; one packfile.h fix for 0.3.11 -C1 (the fee switch H = 210,000 at about 19:00Z on 6 October, 18:50 to 19:35Z by the 22:24Z DAA read; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before about 19:00Z (18:50 to 19:35Z). The proving agent recommends (b) unless (a) is certain; the check decides it. +C1 (the fee switch H = 210,000 at about 18:45Z on 6 October by the 22:32:54Z DAA read, 1.002 DAA/s since 15:40Z; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before about 18:45Z. The proving agent recommends (b) unless (a) is certain; the check decides it. The attempt-0 packfile.h bug (the era agent's find) is the one that took the fleet down at 18:23Z (epoch 34, attempt 1); it is already fixed attempt-aware on pack-loop af983a7 and merged into the 0.3.10 tree. Rule for 0.3.11: one derivation, one test set: the era and cache agents build their workers on that packfile.h and keep only tests that add a case; the host.cu / host.c mh_word derivation fix (new, the era agent's) stays with its own test. @@ -561,3 +561,5 @@ A dry run of ca2-v3 into master in a scratch worktree conflicts in docs/bench-lo Gate 4 (22:25 to 22:31Z, a real Metal miner through the miner's --prepare-packs flow on the fast-time network, igneum-bench from ca2-v3 00c55aa): three v3 prepares with the pack dir and the class and era tokens; the worker's v3 prepared lines (252.5 ms the first: program 55.5, dataset 196.9, cache fill 0.9, build 41.1; then 51 and 46 ms with the day resident); every swap "with no pause"; 124 blocks accepted on v3, cpu re-check mismatched 0, need 0, no mismatch or refusal; chain 182 / 123 across the boundary, 0 rejected, one sink. Found and fixed before the run: the Metal worker's serveDataset built every day with the Swift version 2 construction keyed by day only, so a v3 program would have hashed over an x1 dataset and every found would have been refused by the CPU re-check, the same fleet-outage class as the app's missing flag; now a v3 prepare builds the day from the pack's memhard.metal and the store keys datasets by (day, class, era) (00c55aa). Two Mac outages caught by one gate run that the CPU-miner gates could not see; the rule for the next cut: every worker path (Metal, CUDA, OpenCL) mines across a boundary in the gate network before a class change ships. The integration merges (readwidth 30ff674, then origin/master) are in progress on ca2-v3 with the conflicts resolved by hand keeping both sides. Gates: G1, G2, G3, G4, G4b, G6 GREEN; G5 at the ship's build step. The ship waits only on the merged tip and its checks. + +22:33. H = 210,000 lands about 18:45Z on 6 October (DAA 137,041 at 22:32:54Z, 1.002 DAA/s since 15:40Z); the 16:00Z check keeps 2.75 hours. PC 2: floor-build-3 green at 22:32:29Z (240 s on the warm target; the patched sp1-gpu-server 166,768,224 bytes sha256 5568108b..., v6.8.1 c84ada1e with patch 700173fe, sm_86 sm_89 sm_120, the Go tarball's sha256 matched; the live server untouched, miners never stopped); "go PC 2 sweep" given for floor-sweep-1 (9 points, 8 to 10 min) ahead of the aggregation-cost re-run. From 8f03662db2e8e4237e2d2a99f57430cafa1f0124 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:38:33 +0000 Subject: [PATCH 131/131] Counter ASIC 2.0 status 22:38: the merged tip 2112721, the 0.3.11 ship assigned with its inputs --- docs/plans/counter-asic-2-status.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 54a22d5fc..3ed9a6d24 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -563,3 +563,9 @@ Gate 4 (22:25 to 22:31Z, a real Metal miner through the miner's --prepare-packs Gates: G1, G2, G3, G4, G4b, G6 GREEN; G5 at the ship's build step. The ship waits only on the merged tip and its checks. 22:33. H = 210,000 lands about 18:45Z on 6 October (DAA 137,041 at 22:32:54Z, 1.002 DAA/s since 15:40Z); the 16:00Z check keeps 2.75 hours. PC 2: floor-build-3 green at 22:32:29Z (240 s on the warm target; the patched sp1-gpu-server 166,768,224 bytes sha256 5568108b..., v6.8.1 c84ada1e with patch 700173fe, sm_86 sm_89 sm_120, the Go tarball's sha256 matched; the live server untouched, miners never stopped); "go PC 2 sweep" given for floor-sweep-1 (9 points, 8 to 10 min) ahead of the aggregation-cost re-run. + +## 22:38 the merged tip is in; the 0.3.11 ship is assigned + +ca2-v3 fa3c932 (code 49c7e78): readwidth 30ff674 merged at 3566afd (emit.rs two hunks, HEAD's superset; packbench.swift's footprint lines added to the hot-aware RESULT line; bench-log both; read-width.md readwidth's version), origin/master 1f0d62c merged at 49c7e78 (host.c one struct hunk, both fields kept; master's topology, duplicate-platform and select read-back code and ca2-v3's class, era, mh_word, hot and mixer code all auto-merged; packfile.h no conflict); checks all green on 49c7e78: igneum-pow 53 + 4 + 19 + 7, packfile-test 0 failures, the NVRTC emulator test PASS (17 sampled hashes equal to hash-bound), test-generic.sh on Apple OpenCL PASS, the fork's check clean. Fork ca2-v3-node 89dfcb95. + +The ship is assigned to the 0.3.10 shipper (ae892a8b0f78fe31c) with every input (rollout plan sections 3, 4, 4a, 7, 8, 8a): release-0.3.11 from origin/master; merges in order ca2-v3 fa3c932, ca2-coord, the app branch proving-v1 (e0de2ab plus the prover log-line commit), bash-body-check e3bd761, consequences, proving-methods e7e0db7, asic-history if it lands; the fork 89dfcb95 under the release tree's vendor/; the override object with every switch; N4 = (DAA at publish + 14,400) rounded up to a multiple of 3,600, N5 = DAA + 14,400; the expected digest c562d70e...; the deadline note "program class v3 + proving v1"; update-now machine by machine with PC 2 last after the prover-floor sweep; the hand nodes, the seed, the digest sweep, the HiveOS package, the merge to master and the push. The node and proving agents stand by for fixes. PC 1: Ember's tune run. PC 2: floor-sweep-1 (from 22:34:11Z), then the aggregation-cost re-run.