igneum/docs/plans/counter-asic-2-status.md
2026-10-06 07:02:44 +00:00

183 KiB

Counter ASIC 2.0: status

Coordinator's running status for the plan in docs/plans/counter-asic-2.md. Rewritten every 45 minutes while the work runs. Times UTC, 5 October 2026 (night). Heading times before 20:40 were corrected at 20:42 from the commit clock (the coordinator had written them from a guessed clock, up to 2 h 40 min ahead; corrected again at 20:49 for the 20:42 to 20:47 entries); every entry's true time is its commit's author time in UTC, and from 20:49 every heading is stamped from date -u. Base for every ca2 branch: readwidth at 019b014 (the LoadClass flag, fold_words, the scratch op, the three emitters, 20 packs).

19:55 first status (the 19:50 start was cut off by exhausted credits at about 19:58 before any sub-agent work landed; respawned at 19:55 on the restart)

Layer Branch Agent State Numbers so far Blockers
1, 2, 3 (widths, mix, scratch) readwidth 019b014 a451c9935bfb1bc19 (not ours) Mac Metal rows in; Apple OpenCL next, then one job per PC; scratch re-run at 32 and 128 KB per warp CPU verifier, M5 Max one core, avg of 50, ms per warp: v2 0.604, w16 0.610, w64 0.630, w64x4 0.160, mixA 0.620, mixB 0.614, scr0 0.600, scr2 0.534, scr4 0.457, scr8 0.314. Metal M5 Max hash rate (5 x 2^24, GPU time, 3 of 3 vectors on every pack): v2 27.7 MH/s, w16 28.3, w64 28.2, w64x4 109.7 (32 loads of 64 B), mix 50/35/15 over 6 programs 25.4 to 28.4 (median 26.7), mix 25/50/25 over 6 programs 22.0 to 25.2 (median 24.4). Scratch at 1 MiB per warp, 1,024 warps: 23.1 / 18.2 / 16.3 / 14.9 MH/s at 0 / 12.5 / 25 / 50% RMW, superseded by the cap holds the Mac measure lock and both PCs first
4 + 8 (era layout, working set) ca2-era a452664c512c73b9b design, respawned 19:56 none PC time behind readwidth
5 (cache-sized second table) ca2-cache a5271cf269757b118 design, respawned 19:57 none PC time behind readwidth
6 (SRAM schedule) ca2-analysis a5c6bc2dfcc4613ef (respawned 19:58) cited analysis none none
7 (integer matrix family) ca2-analysis a5c6bc2dfcc4613ef design; dp4a throughput owed unless PC time frees none Apple int8 path to check against Metal docs
3 soundness ca2-soundness a548aadeefd1ab3b2 (spawned 19:59; sizes 32 and 128 KB per warp) analysis + tests none none
Integration v3 ca2-v3 after the above waiting none the readwidth table and the four branches

Budget rule received from the coordinator at the restart: the per-warp scratch is capped so the whole working set (1 GiB table + hot table + scratch for every resident warp + buffers) stays under 6 GB on an 8 GB card, which puts the scratch in the tens of KB per warp; the layer 5 hot table shares that budget.

Decisions for the project lead so far: none. Nothing here touches consensus or any live node.

20:05 readwidth PC jobs out; the scratch finding

Base moved: every ca2 branch rebases onto readwidth b970dda (scratch per warp is a class parameter, 32 or 128 KiB; the SCRATCH_* constants are gone; Metal pack harness proto-metal/packbench.swift; OpenCL --bench-pack). The coordinator branch is rebased; the four agents were told.

Readwidth PC jobs published 20:02:51Z: fetch-readwidth-20261005 (both PCs), run-readwidth-5090-20261005 on PC 2 (6 to 10 min once it starts; a build job from another session is queued ahead of it), run-readwidth-9070-20261005 on PC 1 (10 to 15 min; only the gfx1201 card is switched off). My layer 4, 5 and 7 PC jobs queue behind these.

Finding that bears on layer 3 (readwidth, Metal, M5 Max, capped sizes): the scratch read-modify-writes are cheaper than the dataset loads they replace, so the rate RISES with the RMW share.

Class MH/s
v2 27.7
scr0k32 (control) 28.1
32 KB per warp, 12.5 / 25 / 50% RMW 25.4-26.1 / 29.4-31.7 / 44.4-49.1
128 KB per warp, 12.5 / 25 / 50% RMW 24.3-26.2 / 26.4-28.1 / 34.1-35.4

Reading (readwidth agent): 4,096 warps x 32 KB = 128 MB sits in the chip's caches. Consequence for the decision: a scratch that fits a GPU's cache fits a chip's SRAM at the same size, so at the capped size the writes cost everyone a cache-bound op in place of a latency-bound load. Passed to the soundness agent: what size would make the writes cost DRAM latency, whether that fits the 6 GB cap, and whether the RMW share should be added to the 16 dataset loads rather than taken from them.

20:08 mandate: v3 on the devnet tonight, by the project lead's rules

the project lead has gone to bed and delegated the three decisions for the DEVNET only (not the public testnet): width, per-load mix and scratch share by the rules now written in docs/plans/counter-asic-2-rollout.md section 6, the activation height = devnet tip + 14,400 at publish (checked >= 10,800), published the way finality v3 was. Six gates before any publish (rollout section 7): bit-exact v3 on the three cards; the CPU verifier exact on 1,000 random GPU hashes per card; the soundness suite green with the new scratch tests; the fast-time 3-node network mining across a v3 activation with 0 rejected blocks and 0 forks; Windows and Mac workers from one commit; the node change on a fork from the 0.3.10 tip (21d4c73c) with suites green on PC 2. Release 0.3.11 through the shipper's pipeline; the 0.3.10 shipper (ae892a8b0f78fe31c) has been asked for its state and the handoff. If a gate fails: stop, write why here, do not publish.

Node build note for the integration: the node links igneum-pow by path (../../../../igneum-pow from consensus/pow and igneum/miner), so the node fork worktree for v3 must live under the v3 worktree's vendor/ so that the path resolves to the v3 crate, not master's.

After the publish: Counter ASIC 3.0 from the ASIC-history agent's ranked additions (a202a09dcd24ba1d3), as class v4 behind its own activation, same gates, docs/plans/counter-asic-3.md.

20:08 the shipper's answer, the node fork convention

0.3.10 (shipper ae892a8b0f78fe31c): staged and blocked on GitHub's Actions incident (run 37365130137 queued since 19:42:43Z under a re-dispatching watcher); nothing on the network has moved, the live manifest is still 0.3.9. Once CI is green: ship (5 min), update-now (apps restart 1 to 10 min later), hand nodes and seed (5 min), digest sweep; 0.3.10 finished about 30 min after green. The app version per machine on the console's cards is the restart signal for re-running any straddling measurement. HiveOS is published by the ship's --public step, nothing separate.

Node fork for v3: base on COMMIT 21d4c73c (release-0.3.10 in vendor/igneum-node = release-0.3.6 a24ab01a + housekeeping 4fb32865 + tx-gossip e242acd0 + c4-fix e18f1e0e). Fork-side pack-loop 05ef0fa3 (the miner's pack check, exit 44) is not in it and touches the same miner paths as v3 (export-pack, the job line): merge it. Because the node links igneum-pow by relative path, the v3 node worktree goes under the v3 worktree: git -C /Users/joshm/Projects/igneum/vendor/igneum-node worktree add /Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2 -b ca2-v3-node 21d4c73c.

0.3.11 shipper: the coordinator assigns it (the 0.3.10 shipper stops at its report). Inputs the ship needs: the fork commit with its PC 2 suite results recorded, the main tip, the override object with every switch (the four live fields plus program_class_v3_activation_daa), the expected digest read on a 22-s scratch node, the deadline note ("program class v3"), the activation height, the one-line changelog, and whether the pinned proving guest changes (it does not: the prover has no igneum-pow dependency; confirmed by grep of proving/igneum-prove Cargo files).

20:10 readwidth round 2 on the PCs; the width arithmetic under the project lead's rule

Round 1 of the readwidth PC jobs refused every pack (the workers demand a 32-byte chain seed; the experiment packs carried string seeds); fixed at readwidth 1ea7a52 (packfile.h), republished 20:09:19Z as run-readwidth-5090-20261005c (PC 2, about 6 min) and run-readwidth-9070-20261005c (PC 1, about 10 min). The era and cache agents were told to rebase onto 1ea7a52 and to prove their packs load on the Mac OpenCL host before any PC job. Every ca2 branch now bases on 1ea7a52.

Probe ceilings from round 1 (dependent reads per second at 1024 MiB, device time):

Card 4 B 16 B 64 B Stream
RTX 5090 17.5 G 18.0 G 9.1 G (584 GB/s) 1,579 GB/s
RX 9070 XT 2.42 G 2.43 G 2.47 G (158 GB/s) 636 GB/s

the project lead's width rule applied to the ceilings alone (the measured v3 rates will replace this when the table lands): a 128-load hash at 64 B reads 8,192 B; at the 5090's 64 B ceiling that is 71 MH/s and 584 GB/s, 37% of its stream bandwidth, over the one-third margin the rule sets; at 16 B it is 2,048 B per hash, 141 MH/s and 288 GB/s, 18%, inside the margin; 4 B is 9%. On the 9070 XT every width costs the same line fetch (2.4 G/s), so 16 B is where AMD gains 4x the bytes per hash at no cost and the 5090 stays latency-bound with margin. Provisional width under the rule: 16 B (w16), pending the measured rates and the latency-bound share per card.

Added deliverable (20:11): the public description in four levels, docs/plans/counter-asic-2-public.md (aec53bb): levels 1 and 2 are written as copy; level 3 carries the bench table with [owed] markers for every number not yet measured; level 4 lists the documents. The integration branch applies levels 1 to 3 to site/index.html, site/litepaper.html and site/bench.html with the final numbers.

20:15 layers 6 and 7 landed; the node-fork agent started; 0.3.11 scope

Branch ca2-analysis (5d5ba15, f59708d). Layer 6: no cache growth rule exists in the spec; cited bit cells N7 0.027, N5 / N3E / Intel 18A 0.021, N3B 0.0199, N2 0.0175 um^2, array factor 0.70; the 256 MiB mirror is 83 / 64 / 54 mm^2 at N7 / N5 / N2 (74 at N2 with a 96 MB hot table), $13 to $26 of silicon per good die (approximate). The mirror was never unaffordable; the cache's job is to stay above GPU L2 (5090 96 MB, GB202 128 MB). Recommendation C for the project lead (gate 1): cache doubles when the dataset doubles. Layer 7: Metal 4 matmul2d has int8 x int8 -> int32, so a unit-level mm8 tile is native on all three vendors; per-lane dot4 is emulation on Apple (M5 Max: 548 G unsigned dot4/s against an 880 G ALU chain, 1.6x; signed 4.7x). Reserve R1 = mm8, W_new 4, unlock era 4 or 90% signal. Owed: the 5090 and 9070 XT dot4 probe (job prepared: relay/playbooks/dot4-probe.ps1, exe sha256 5adaeb1a...6416f4; publishes when a PC frees).

Node-fork agent a3f505a9d981300cd started 20:50: ca2-v3 (igneum-pow seam: ProgramClass, Epoch::from_chain_seeds, generator 3 in the program id, class in the pack) and ca2-v3-node (vendor/igneum-node-ca2 under the ca2-v3 worktree, from 21d4c73c, with pack-loop 05ef0fa3 merged): the field in Params, OverrideParams and the digest, the epoch-boundary rounding, the era stand-in, the job line, the fast-time gate script.

0.3.11 scope (coordinator, 20:52): carries program_class_v3 and proving v1 together (one override object, one digest, one publish; each half under the same gates); the ready half ships as 0.3.11 and the other as 0.3.12 if one lags. Next-cut list recorded in the rollout plan section 8a.

20:16 decisions recorded: layer 6 option C, layer 7 R1 = mm8; the chip headline

Layer 6 DECIDED (delegated; the project lead confirms for the public testnet genesis): option C, the cache doubles when the dataset doubles (256 MiB genesis, 512 MiB year 4, 1 GiB year 12); one-core fill 0.2 / 0.4 / 0.8 s, under 1 s at every step. Layer 7 DECIDED: reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal. The chip model's headline now names the on-die-cache recompute chip (54 to 83 mm^2) as a row per variant; the scratch share is chosen as the smallest share at which that chip's gain falls under 1.5x, else said plainly and the public "under 2x" claim qualified. The soundness agent carries that table (its question 2) with M16's mixer multiplier beside it.

20:18 correction to the SRAM mirror figures (chip-economics research, cluster D)

Bit cell x 0.70 understates real die area. Shipped cache dies: AMD 3D V-Cache 64 MB on 41 mm^2 at 7 nm (1.56 MB/mm^2, Tom's Hardware, Hot Chips August 2021); Graphcore GC200 900 MB on 823 mm^2 with compute (1.09 MB/mm^2); Groq TSP 220 MB on 725 mm^2 at 14 nm (0.30 MB/mm^2); TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30% assist overhead (SemiAnalysis, December 2022). A 256 MiB mirror is about 165 mm^2 at 7 nm on the densest shipped cache-only die and about 130 mm^2 at N5/N3E, not 54 to 83 mm^2; cost per die 2 to 3x the earlier figure; the conclusion (affordable for a funded chip) stands. The analysis agent is redoing the table with both columns; the soundness agent carries the corrected density into the chip row. Latency citations behind the latency-bound rule, to be added: DRAM row cycle 40 to 48 ns across DDR4, GDDR5, HBM2 (Li, Reddy, Jacob, MEMSYS 2018); latency 1.3x in two decades against bandwidth 20x (Chang 2017); no shipped mining chip used HBM or stacked memory.

20:19 proving v1 state for 0.3.11; PC 2 occupancy

Proving v1 (acd4f36bc2c07a4e2): fork proving-v1 b177718e on a24ab01a (told to rebase onto commit 21d4c73c now), app proving-v1 79bc820 on a93199a. Override fields proving_v1_activation_daa (tip + 14,400 at publish), proving_v1_segment_blocks 4, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Harness: tools/proving-v1/net.mjs --secs 1500 PASSED (21 checks) in 197.3 s on b177718e; rerun owed on the final tree. The pinned guests do not change. Shared files with ca2-v3-node: params.rs, daemon.rs, igneum/miner/src/main.rs, override-60x.json; both agents keep separable hunks.

PC 2 is held by the proving agent's memsweep-pc2-pv1 (about 20 min, miners stopped) and a second run (about 10 min). Queue after it: the ca2 node suites, then the readwidth, era, cache and dot4 measurement jobs. PC 1 is held by readwidth's run-readwidth-9070-20261005c until it reports.

20:21 sram-mirror.md revision 2 (ca2-analysis)

Two columns, headline = shipped-product density (AMD V-Cache 41 mm^2 per 64 MiB at N7, scaled by the bit-cell ratio), lower bound = bit cell x 0.70. mm^2 and $ per good die (D0 0.1 per cm^2, wafer prices approximate), headline / lower bound:

Node, wafer $ 256 MiB 256 + 96 MB hot table 1 GiB
N7, $9,500 164 / 83 mm^2, $30 / $13 226 / 114, $44 / $19 656 / 331, $224 / $75
N5, N3E, 18A, $20,000 128 / 64, $46 / $21 175 / 89, $68 / $30 510 / 258, $306 / $111
N3B, $20,000 121 / 61, $43 / $20 166 / 84, $63 / $28 483 / 244, $280 / $103
N2, $30,000 106 / 54, $56 / $26 146 / 74, $81 / $37 425 / 215, $343 / $131

Year 10 at the 6% per year trend: 59 mm^2 for the flat cache (82 with the hot table), 8% of a 750 mm^2 die. One reticle holds 1.3 GiB (N7) to 1.9 GiB (N2); mirror share of a 750 mm^2 die at year 0: 14% (22% with the hot table), inside M16's 13 to 40% band. Recommendation unchanged: C. Latency section added (MEMSYS 2018, Chang 2017, the mining-chip memory-type note), marked as research the agent did not re-read tonight apart from the V-Cache figure.

20:23 layer 3 soundness landed: the scratch does not move the chip; scratch share decided 0

ca2-soundness (0d8f745 tests and trace hook, a465881 doc and bench-log). The on-die-cache recompute chip (N5 headline 128 mm^2, $46) at 50 T op/s: 333 MH/s against the 5090's measured 139.7, 2.4x; at 12.5 / 25 / 50% RMW replaced, chip 381 / 443 / 661 against 5090 projected 160 / 186 / 279, 2.4x each; added, 2.4x or more; 32 or 128 KB alike. The chip keeps the scratch implicitly in 80 to 320 B per lane (the verifier resets it per unit), needs about 530 units in flight, dense scratch 6.2 / 3.1 mm^2 at N5. Under the project lead's rule the scratch share is 0: layer 3 is NOT adopted into v3; the public "under 2x" claim is qualified (public copy level 3 rewritten). The measured lever is M16's mixer multiplier (x2 1.2x at 0.8 to 2.4 ms verify; x4 0.6x at 1.6 to 4.8 ms; 3.6x and 1.8x with a 3x fixed-function factor); whether x4 enters v3 tonight is asked of the coordinator; default: Counter ASIC 3.0.

Soundness results (Metal, M5 Max): 28/28 edge launches, 200/200 fuzz packs (91 s), 56/56 hand-model edge checks, 42/42 kernels pass the static scratch-mask check with 6 deliberate breaks caught, broken tag and broken lazy fill caught, fingerprint 8c07620f4d9adefd warp-count-independent; re-hit rates 2 to 33% above the birthday bound (slot addresses are register low bits); written words unbiased (worst 3.63 of 6 sigma). Verifier exactness needs a host contract (zero the arena at allocation and at the 32-bit tag wrap, tags from 1), which neither host gives today. Pre-existing on readwidth b970dda: verify::tests::fold_and_wide_fetch overflows under the test profile (wrapping_mul fixes it); passed to the readwidth agent with the class sweep.

Gate G3 note: the scratch tests (igneum-pow/tests/scratch.rs) join the v3 suite even though the class carries no scratch, parametric over the class; they guard the v2 path's scratch-free invariant at zero cost.

20:24 decided: M16 mixer x4 into v3; agent ca2-mixer started

Coordinator's decision under the project lead's delegation (recorded in the rollout plan section 6a): the mixer multiplier x4 and the cache growth rule (option C) enter class v3 behind the same activation; layer 3 stays out at scratch share 0, its soundness document and pack-contract tests kept. Agent af345b1e2c541ffbb (branch ca2-mixer) implements mixer_mult as a class parameter (m mixer applications per round, the 8 dependent reads unchanged), the cache_log2_words(day) schedule (doublings at years 4 and 12 with the dataset stepping to the next power of two), re-cuts the v3 dataset vectors, re-runs the soundness suite, measures the verifier (v2 0.604 ms per warp; v3 expected 1.6 to 4.8 ms) and the 1 GiB build on the Mac, prepares the 5090 and 9070 XT build-time job, and writes docs/analysis/chip-model-v3.md with the combined headline row (fixed-function factor included). The claim on the site reads "under 2x" only if that row does; else qualified, with the mixer x8 and the hot table named as the next levers.

Agents now: ca2-era (a452664c512c73b9b), ca2-cache (a5271cf269757b118), ca2-node (a3f505a9d981300cd), ca2-mixer (af345b1e2c541ffbb). Done: ca2-analysis, ca2-soundness. Waiting: the readwidth PC table; PC 2 (proving memsweep runs) and PC 1 (readwidth 9070 round).

20:27 the readwidth table landed; layers 1, 2, 3 decided; layer 5 measured on the Mac and redesigned

Readwidth e752fc7 (docs/plans/read-width.md), bit-exact on Metal, Apple OpenCL, the 5090 (NVRTC) and the 9070 XT, both PCs released. MH/s (latency-bound share):

Class RTX 5090 RX 9070 XT M5 Max Gap
v2 (128 x 4 B) 136.1 (0.96) 18.15 (0.87) 27.74 (1.01) 7.5x
w16 139.8 (0.90) 17.90 (0.84) 28.26 (1.03) 7.8x
w64 71.9 (0.58, 37% of stream) 17.59 (0.78) 28.27 (1.03) 4.1x
w64x4 (32 loads) 275.3 (0.56) 75.19 (0.84) 109.7 (1.00) 3.7x
mix 50/35/15, six programs, spread of median 18.8% 7.4% 11.3%
mix 25/50/25 22.3% 5.5% 8.1%
scratch 32 KB at 12.5 / 25 / 50% -18 / -21 / -12% -18 / -22 / -21% -7 / +12 / +74%
scratch 128 KB -21 / -30 / -48% -21 / -27 / -33% -7 / -1 / +25%

Decisions (the project lead's rules, delegated): layer 1 keep v2 (w16 passes the rule but closes nothing and does not move the chip row; the vector re-cut is not worth it); layer 2 out (spread over 5% on every card); layer 3 out (scratch share 0). The AMD gap is the card's dependent-read rate (2.4 G/s at every width), stated for the user tiers in the rollout plan 6b.

Layer 5 (ca2-cache 53ef59f, 011cc0a, 86726cd, 65bc7a7; docs/plans/hot-table.md): five packs bit-exact on Metal and Apple OpenCL (96/96 each). M5 Max rates against v2 27.68: hot32k4 x1.22, hot64k4 x1.12, hot96k4 x1.05, hot64k2 x1.00, hot64k8 x1.71; verifier 0.344 to 0.560 ms against 0.626; hot fill per epoch 24 / 46 / 73 ms on one core, 0.07 / 0.15 / 0.22 ms on the GPU; Apple OpenCL probe 32 / 64 / 96 / 1024 MiB 21.7 / 12.8 / 12.3 / 3.50 G loads/s. Redesign ordered: hot loads ADDED beside the 16 dataset loads (the replaced form lets the on-die-cache chip skip item derivations and worsens the gain); the agent re-measures the added form and rebuilds the PC job. PC 1 is given to the dot4 probe (under 15 min), then to the era agent, then the hot-table job; PC 2 stays the proving agent's.

20:27 layer 9 added; the era draw passes Mac bit-exactness; two class bugs in the harnesses; C1

Layer 9 (the project lead: faster program changes): the epoch length becomes an era parameter in the genesis reserve, 1 hour at launch, 10 minutes to 2 hours by draw or 90% signal, reserve-only tonight; an agent (ca2-epoch) designs it beside layers 4 and 8 and measures the compile-ahead cost per card at a 10-minute epoch, the seed-path consequence and the FPGA threat it answers; one row in the rollout plan section 6, one in the level 3 numbers. Spawns when the dot4 probe frees its slot.

ca2-era mid-way (a452664c512c73b9b): era draw behind LoadClass::era (EraParams beside mix, load_slots, scratch, scratch_kb; verify::load_index; memhard::Layout; one load form in the three emitters; --era, --era-widths), 54 crate tests green, pinned packs byte-identical; six era packs bit-exact on Metal (packbench 6/6), Apple OpenCL (6/6, same fingerprints) and the CUDA emulation (6/6); re-exporting with the width pinned at 4 B and rebasing onto e752fc7; PC job in about 20 minutes. Two class bugs found and fixed on its branch, both outside its layer: (1) proto-cuda/nvrtc/packfile.h re-derived the seed words as attempt 0 only, so ANY pack with IGNEUM_PROGRAM_ATTEMPT >= 1 (5.14% of chain epochs under v2) is refused by the one-click workers with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", the error PC 1 logged on 5 October and attributed to the export race; fixed attempt-aware, with a tampered-attempt refusal test. This fix must ship in the 0.3.11 workers whatever else does. (2) proto-cuda/host.cu and proto-opencl/host.c derived dataset words on the host as mh_item(w >> 4)[w AND 15] instead of the pack's mh_word; fixed.

Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout plan 7a (0.3.11 on every prover before 16:00Z on 6 October, else the fee switch is republished at tip + 86,400); C10 was resolved by the mixer x4 decision (20:23, 20:24 entries), the pool-core and node verification numbers are being measured on ca2-mixer; C11: AMD RDNA 4 is about a seventh of a 5090 on this hash by dependent-read rate, 4.9x the electricity per hash and 2.2x worse per pound at list prices (0.032 against 0.072 MH/s per pound, read-width.md 4.1, approximate); the card's memory system, not a tuning gap; MH/W and MH per pound columns (list prices, approximate) go into the final table and the level 3 page; the line-width question is a 3.0 question since no width closes the gap without making the 5090 bandwidth-bound.

20:29 C1 decided for the morning; one packfile.h fix for 0.3.11

C1 (the fee switch H = 210,000 at about 18:45Z on 6 October by the 22:32:54Z DAA read, 1.002 DAA/s since 15:40Z; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before about 18:45Z. The proving agent recommends (b) unless (a) is certain; the check decides it.

The attempt-0 packfile.h bug (the era agent's find) is the one that took the fleet down at 18:23Z (epoch 34, attempt 1); it is already fixed attempt-aware on pack-loop af983a7 and merged into the 0.3.10 tree. Rule for 0.3.11: one derivation, one test set: the era and cache agents build their workers on that packfile.h and keep only tests that add a case; the host.cu / host.c mh_word derivation fix (new, the era agent's) stays with its own test.

docs/analysis/card-lifetime-2026-10-05.md (branch card-lifetime 1fecfe2) merged into ca2-coord. Decided (delegated): the GPU frees the 256 / 512 / 1,024 MiB cache after the daily dataset build (the hash never reads it; the rebuild costs 0.67 ms fill + 13.4 ms build on the 5090 at x1, about 54 ms at x4, owed); hot-table.md's resident reading is corrected, era-layout.md's freed reading stands. Recommended for the project lead: growth mapping (b), power-of-two steps at years 4, 12, 28, 60 with AND MASK, the only mapping under which the litepaper's "4 GB about four years, 8 GB more than a decade" holds (4 GB: year 4 under (b), 1.0 to 1.5 years under (a); 8 GB: year 12 or 6.3 to 7.5; 12 GB: year 28 with the cache freed; 24 GB: year 60). Public lines and the evidence row go on the integration branch (rollout plan 6c); hot-table.md line 73 (the 8 GB row counted a 5090's warps) goes to the cache agent.

20:32 layer 9 agent started; the dot4 probe is on PC 1

ca2-epoch (a32a3ece66c02417a): the epoch length as an era parameter (600 to 7,200 DAA s, base 3,600; draw or 90% signal; the VDF rule; the difficulty-window constraint; the FPGA threat with citations), Mac compile-ahead measured now, the 5090 and 9070 XT compile times cited from the bench log, docs/plans/epoch-length.md. The dot4 probe job is running on PC 1 (ca2-analysis tip ee42d7c carries the playbook; the 5090 confirmation is the agent's watch). PC 1 queue after it: the era six-pack job, then the hot-table added-form job. PC 2: the proving agent's, then the ca2 node suites.

Agents running: ca2-era, ca2-cache, ca2-node, ca2-mixer, ca2-epoch; ca2-analysis watching its PC job. Done: ca2-soundness.

20:38 layer 7 complete on all three cards (ca2-analysis ee42d7c); PC 1 free

dot4 probe on PC 1 (jobs fetch-dot4-20261005, run-dot4-20261005, exit 0, 101 s; both cards restored and mining; 5090 SM clock 2,505 MHz before and after):

Device ALU chain, G steps/s signed dot4 emulation, G dot4/s dot4 instruction, G dot4/s emulation vs instruction
RTX 5090 (NVIDIA OpenCL 3.0, driver 617.14) 8,753.5 1,239.1 (7.1x an ALU step) 7,453.6 (inline PTX dp4a.s32.s32, 1.17x) 6.0x
RX 9070 XT gfx1201 (AMD-APP 3683.0) 701.4 480.8 (1.46x) 664.3 (__builtin_amdgcn_sudot4, 1.06x) 1.38x
gfx1036 (RDNA 2 iGPU) 40.6 15.8 (2.6x) sudot4 does not build (needs dot8-insts)
M5 Max (Metal) 879.8 188.2 (4.7x); unsigned 548.2 (1.6x) none exists

All bit-exact against the CPU reference. One dp4a costs about one ALU step on NVIDIA and AMD; the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on dot4 (the family does not widen the AMD gap), 10x the M5 Max on the chain and 13.6x on dot4 (Apple's emulation widens its gap 1.4x). At W_new 4 the family adds about 21 ops per hash per lane; hash-rate losses expected under 5% on every card (to be measured with the family live). Owed: the CUDA __dp4a cross-check (needs nvcc), Metal 4 matmul2d int8 on the M5, sdot4 on RDNA 2. cl_khr_integer_dot_product is listed by no driver we own.

PC 1 is free: the era six-pack job goes next when its package arrives, then the hot-table added-form job.

20:39 PC 1 scheduler (the coordinator's role from now): the queue

Rule: one PC 1 job at a time; an agent asks by message before publishing, gets a "go PC 1" from this coordinator, and reports when its RESULT lines are in and both cards are restored; a hash-rate or power number taken while another job holds a card is not a number. The CPU-only job runs only in a slot where no measurement overlaps it.

# Job Agent Cards Length State
1 Era six-pack (layers 4 and 8), 5090 and 9070 XT ca2-era a452664c512c73b9b one card at a time about 15 to 20 min waiting for the package (packs on the pack-loop packfile.h)
2 Hot table, added form, probe 32/64/96 MiB plus packs ca2-cache a5271cf269757b118 one card at a time about 15 min waiting for the re-measured Mac rows and the rebuilt zip
3 Reproducible benchmark run a0b9f574775ef1693 one card at a time, 120 s per card about 5 min queued
4 AMD sweep on the 9070 XT (core clock and power steps) a01dcb34ae16d867c 9070 XT only; the 5090 keeps mining about 20 min queued
5 Ember Tune end to end, both cards a855dcc4bd05e0615 both to be stated queued after the AMD sweep
6 AMD-proving CPU fallback (CPU-only SP1 run, both cards mining) a39db54d4de4af51e none; loads the CPU to be stated last, or in a gap where no measurement runs for its whole length

If job 1's package is more than 15 minutes away when job 2 is ready, job 2 goes first; the short job 3 fills any gap of under 10 minutes between packages.

20:39. No more agents are spawned tonight (the project lead: no unnecessary credits); the running ones finish. PC 1 queue change: the AMD-proving CPU fallback's small fixture (block-56-transfers-3shards, minutes) runs NOW in the gap before the era package; its S_p shard (block-338-shard1, up to 30 min of every core) stays job 6, last. Next-cut list gains the rig installer's two follow-ups (the Linux manifest entry, the Linux prover build), tied to whichever release carries the proving half (rollout plan 8a).

20:42 the node switch is written; the Metal worker is a gate item

ca2-node (a3f505a9d981300cd): ca2-v3 commits d2cd6e1 (pack-loop af983a7 merged: packcheck.rs and the attempt rule; one packfile.h conflict resolved), 50d5c86 (the seam: ProgramClass { V2, V3 }, V3_CLASS placeholder w16, generator 3 in the program id, Epoch::from_chain_seeds, Epoch::chain_program, IGNEUM_PROGRAM_CLASS and IGNEUM_ERA_SEED_HEX in the packs, packcheck refuses wrong class / era / generator; v2 packs byte-identical; 53 tests), 9ed787e (workers: packfile.h reads class and era, CUDA and OpenCL workers take class=v3 era=<hex> tokens on job and prepare lines, pair identity includes them, mismatch answers need; 13 packfile checks). Fast-time gate script infra/fast-time/class-v3.mjs written, not run (needs the fork binaries). Node (ca2-v3-node, uncommitted until cargo check passes, queued behind the measure lock): the field in Params, OverrideParams, override_params and the digest (unconditional, its own statement; 11-entry digest test), the daemon line, program_class_for_epoch_at (v3 iff 3600 e >= N4, first epoch = ceil), POW_ERA_BLOCKS 15,552,000 and POW_ERA_LEAD 7,200 with the era stand-in (era 0 = genesis), EpochSeeds { epoch, day, class, era }, template and RPC fields 12 to 16, the miner's job line and seeds.txt. PC 2 command ready (run from the ca2-v3 worktree so push-build-inputs.sh packs the v3 igneum-pow); held until "ready for PC 2".

Gate item found: main.swift refuses v3 lines "until Swift has generator 3"; the Mac worker never regenerates a program (igneum-miner export-pack writes the pack), so the fix ordered is to accept the pack's class and era against the line, as the CUDA and OpenCL workers do. Without it Mac node 1 and the Mac app cannot mine v3 (gates G1 and G4).

PC 1: the AMD-proving small fixture is running; Ember Tune's 8-minute CPU-only build takes the next CPU gap, its 30-minute both-cards run is job 5.

20:43. Consequences round 2 (C18 to C20). C18: "near parity per pound" was wrong and is struck everywhere; read-width.md section 4.1 gives the 5090 at 2.2x the 9070 XT per pound at list (0.072 against 0.032 MH/s per pound, approximate), 4.9x per watt, 7.5x in rate; the level 3 page carries those. C19 (to ca2-mixer, already sent by the reviewer): the x4 verifier cost in IBD minutes over the 108,000-header pruning window per tier (8.6 min against 1.1 on one M5 Max core at the top of the range), pool shares per core per second, a scaled 2019-class figure (approximate), and the 10 ms gate margin left for 3.0 go into mixer-x4.md; the seeds' header-verify load goes into the testnet go checklist. C20 (to ca2-epoch, already sent): both rig miners run --exit-on-seed-change and re-export on exit 42, so a 10-minute epoch restarts every card's miner six times an hour and the Mac fleet's prepare pause goes from 35 s to 3.5 min an hour; the epoch-length document gets a per-tier restart-cost row and the compile-ahead margin against the VDF at 600 DAA s, and the rig installer drops the exit-42 path for prepare-ahead before any short epoch can be drawn (next-cut list).

20:43 the Metal worker's v3 path; the app flag; the seam

Correction to the 20:42 entry: igneum-bench --serve DOES regenerate every program in Swift (serveProgram calls generateProgramV2), so a v3 line could not be trusted blind. ca2-node added servePackProgram in main.swift: a prepare <e> <d> <dir> class=v3 era=<hex> line compiles program_bound.metal from the pack the miner wrote (--prepare-packs) after the packfile.h checks in Swift (generator 2 or 3, class against generator, seed bytes against the line, IGNEUM_SEEDW_INIT against attempt_words, class and era against the line); the program store keys on (seed, class, era); a v3 job with no resident pack answers need + error ... program class mismatch; v2 lines unchanged. A pack program never races variants (the Mac loses the variant race on v3 epochs; its cost is the race's gain, from the miner-perf entry, to be quoted). Integration item: the Mac app and Mac node 1's miner command must pass --prepare-packs, else the first v3 epoch on the Mac worker ends in need lines; check app/igneum-app/src/engine.rs. The gate network's Mac miner is the CPU miner (igneum-miner --engine igneum-pow), unaffected.

The seam the node relies on (kept by every branch): Epoch::from_chain_seeds(epoch, day, era, class, label), Epoch::chain_program(epoch, era, class, label), Epoch::chain_dataset(day, class), generate_from_seed_bytes_program_class(label, seed, class, era), ProgramClass::{load_class, generator_version, from_generator, name, parse}, V3_CLASS, Program::era_bytes, packcheck::verify_pack_dir_chain.

Mac build queue: the measure lock has been held by a packbench run since 20:31Z with three build slots held and five builds waiting; the node's cargo check and the swiftc recompile wait behind it. This is the lock working as designed; it sets the pace of the gates tonight.

20:45 the 9070 XT has dropped off PC 1's bus; app restart facts; the era package ETA

The AMD sweep agent (a01dcb34ae16d867c) reports from a read-only probe at about 20:40 UTC: Get-PnpDevice lists only the integrated "AMD Radeon(TM) Graphics" (gfx1036) and the RTX 5090; the app's AMD worker now mines the gfx1036 at 3.12 MH/s; the 9070 XT is absent from PnP (the eGPU link: the Sonnet box or the USB4 router; earlier today it went Code 43 and came back after a driver reinstall and reboot). A 10-second rescan probe is granted (pnputil /scan-devices, the USB4 router status). the project lead is asleep and is not woken. Consequence if the card stays absent: gate G1 (bit-exact v3 on all three cards) and the 9070 XT rows of the era and hot-table tables cannot be taken tonight; the AMD-vendor stand-in available is the gfx1036 (RDNA 2, AMD OpenCL 3683.0, 3 MH/s), which ran the version 1 and version 2 conformance; whether it satisfies G1 for the devnet publish is asked of the coordinator. Every 9070 XT row taken before 20:40 (readwidth, dot4) stands.

App restart facts from the log intake (node tools/logs.mjs, 20:45): PC 1's app run is win-ae432dc7-20261005-190232 (started 19:02:32, no restart since), so no PC 1 measurement tonight straddled an app restart; PC 2's run is win-1ccfe586-20261005-200114 (started 20:01:14, before the readwidth 5090 round at 20:09). The 0.3.10 manifest is still unpublished (the shipper's CI is queued); the "jobs folder cleared by the 0.3.10 update" reading was wrong: the folder is cleared by fetch jobs.

The --prepare-packs item of 20:43 is resolved: app/igneum-app/src/engine.rs line 1225 passes it on master and on the 0.3.10 tree.

Era package: 10 to 15 minutes away (the pack-loop packfile merged over readwidth's; the OpenCL verification and the mingw rebuild queued behind three held build slots); zip ~/Desktop/igneum-ca2-era-pc1.zip, fetch id fetch-ca2-era-20261005, one playbook relay/playbooks/ca2-era-pc1.ps1 doing both cards (about 6 to 10 min). Widths pinned at 4 B in every era pack (512 B per hash), windows identical across the six packs, so the six-era spread isolates stride plus interleave. The CPU-only proving fixture holds PC 1 until about 21:15 to 21:30; the hot-table job is not ready either, so the order stays era, then hot table.

20:46 ruling on G1; hardware event recorded

Ruling (coordinator): gfx1036 satisfies the AMD vendor for gate G1 tonight (a compiler-and-ISA property; it carried the v1 and v2 conformance); the 9070 XT's hash-rate and power rows are owed and taken when the link is back; every 9070 XT row before 20:40 UTC stands. Nobody is woken, PC 1's app is not restarted. The event is in the rollout plan section 7b (hardware events) for the morning summary: the second eGPU link fault today (Code 43 at install, a bus drop at about 20:40); the project lead reseats the USB4 cable and the eGPU power; the 0.3.10 hot-plug code shows "removed" and picks the card up without a restart. The publish proceeds when every other gate is green.

Rescan at 20:45:34Z (relay probe #203, 10 s): the 9070 XT stays absent after pnputil /scan-devices; the USB4 list shows only the host and root routers, the Sonnet Breakaway Box 850T5 router present at 17:18Z is gone: the box is off the link, not just the card. Job 4 (the 9070 XT sweep) is dropped, its rows owed with this reason and time. The era and hot-table PC jobs run their AMD half on the gfx1036 for bit-exactness only (the G1 ruling); their 9070 XT hash-rate and probe rows are owed.

Correction to the 20:45 entry: PC 1's app is 0.3.9 (file 15:47:20Z) and its process started at 20:01:14Z (pid 12340), a restart, not a 0.3.10 install; the log intake's run id dates the log file, not the process. Both PCs restarted at about 20:01Z, before every readwidth PC job (from 20:02:51Z) and the dot4 probe (20:27Z), so no measurement tonight straddled a restart. 0.3.10 is still unpublished.

PC 1 queue now: (1) the AMD-proving small fixture (running, release expected 21:15 to 21:30), (2) the era job (both halves, about 6 to 10 min), (3) the 5090 power-limit sweep (575 / 460 / 400 / 400 W, 90 s each, cap restored to 431 W, about 8 min; SM and memory clocks in the RESULT lines), (4) the hot-table job, (5) the reproducible benchmark (5 min), (6) Ember Tune's 8-minute build in a CPU gap then its 30-minute both-cards run, (7) the AMD-proving S_p shard (up to 90 min, CPU only).

20:49 mixer construction written; the measure lock cleared; PC 2 and the merged-tree suites

ca2-mixer (af345b1e2c541ffbb), no commit yet (lands when the crate tests and the v2 pack diff are green): LoadClass gains mixer_mult (1 or 4) and growth; LoadClass::MX4 = v2 loads, mixer x4, growth on, no width roll, so its program stream is version 2's draw for draw; memhard::Shape { mixer_mult, cache_log2_words } in MixParams; derive_items applies the mixer with keys round_key(r x m + j), j in 0..m, before each of the 8 reads and round_key(8m + j) after; Cache::fill_log2; growth_doublings(d) = ilog2(1 + d / 1460) (doublings at years 4, 12, 28, 60), cache_log2_words(d) = 26 + doublings, dataset_log2_words capped at 32; d = 1 on the devnet pack keeps 2^26 and 2^28. Emitters emit the m-loop only when m > 1 (v2 text byte for byte otherwise); program.h carries IGNEUM_MIXER_MULT and IGNEUM_CACHE_GROWTH. Seam addition (additive): Epoch::chain_dataset_day(day_bytes, class, days_since_genesis, genesis_dataset_log2); the node agent was told to wire the genesis day index from Params.genesis.timestamp and a pow_genesis_dataset_log2 field (28 on the devnet, in the digest). V3_CLASS becomes LoadClass::MX4 composed with the era and hot fields at integration. Numbers follow the lock.

The Mac measure lock: the packbench loop was the readwidth agent's (Metal currentAllocatedSize at 32 and 128 KiB scratch for the consequences reviewer's C12, not a decided row); it stopped the loop at about 20:52, so the queued builds (the node's cargo check, the mixer's tests, the epoch and mixer measurements) proceed.

PC 2: spcurve-stopped-pc2-pv1b closed 20:47:46Z (the card alone: an empty shard 13,875 MiB in 2.1 s; a v1 shard 20,435 MiB, 4.2 s; 2.25 M pgas 28,371 MiB, 6.6 s; 4.5 M 28,307 MiB, 8.5 s; the prototype 6.75 M 28,275 MiB, 11.2 s: the peak plateaus at 28.3 GB from 20 M cycles up); spcurve-miner-pc2-pv1 (the same with the miner on, about 5 min) runs now; PC 2 is released after it. The proving fork tip is 3203c8d0 (eb32c645 plus the pool test's field and N = 8), app 440fd59; its unit tests ran on the Mac (consensus-core 13, exec 8, flows, 0 failed); its harness runs on the Mac. Decision: ONE PC 2 suite job on the merged tree (ca2-v3-node plus 3203c8d0) once the ca2 node branch is committed, covering both halves of 0.3.11.

20:50 the measure lock holder is a prover measurement; a lock-status defect

Correction to the 20:50 entry above: the measure lock has been held since 20:31Z by pid 45000, a with-lock.sh measure of igneum-wt-agg-cost's igneum-prove-host (the aggregation-cost agent's prover measurement, 18 minutes so far), not by the readwidth packbench; with-lock.sh status prints the LAST WRITER's command text, not the holder's, which is why it named the w4 run. Defect for the next cut (tools/lock/with-lock.sh: the status line must read the holder's pid and command, not the last writer's; the class of CLAUDE.md's watcher rule). Builds proceed in the three build slots (build-0 taken at 20:49:51 after an 885-s wait); GPU measurements (the mixer's verifier and build timings, the epoch compile-ahead, the Mac bit-exactness runs under run are not blocked) queue behind the prover measurement. Readwidth head is 30ff674 (per-watt rows for the consequences reviewer; e752fc7 stays the table commit).

20:51 the number-free public copy is applied on ca2-coord (0ad70ba and the next commit)

site/litepaper.html: the Mining section's "bound by memory bandwidth" is corrected to "waits on memory latency, not on maths or bandwidth"; the "Everything above is automatic" paragraph is replaced by the level 2 three ideas and the fourth paragraph with a link to the numbers page; the vs RandomX rows "Changes over time" and "Dataset" and the Hardware paragraph carry the step schedule (years 4, 12, 28; 4 GB about four years, 8 GB about twelve). site/index.html: the Memory row and the Mine card carry the step schedule ("2 GB at genesis, doubling at years 4, 12 and 28"; "any 4 GB card at launch, 8 GB from year 4"). NOT yet applied, because they carry the chip number: the level 1 sentence in the hero and the abstract ("a custom chip gains under 2x") and the limits bullet "A chip is impossible"; they wait for the combined chip row from docs/analysis/chip-model-v3.md (if 1.8x: "under 2x" with the margin stated as thin; else qualified). The bench page's Counter ASIC section (level 3) waits for the final table. docs/evidence.md's card-lifetime row (designed) is still to add.

20:51 the node builds every day cache through chain_dataset_day

ca2-v3 6c75dad (node agent): Epoch::chain_dataset_day(day_bytes, class, days_since_genesis, genesis_dataset_log2) with a placeholder body (the mixer branch fills growth_doublings under that signature), verify::days_since_genesis; the Metal worker takes v3 from a pack (compiled). Fork (uncommitted, in the cargo check holding build-1 since 20:50Z): Params::pow_genesis_dataset_log2 (28 on every network, in the digest in its own statement), Params::genesis_day_index(), install_pow_genesis in the daemon after the class switch, the engine's build_day through chain_dataset_day, the pack export through the same build_day, PowEpochInfo / RPC / proto fields 17 and 18 (genesis_day_index, genesis_dataset_log2; an old node's 0 reads as 28), override-60x.json and the redteam override carry pow_genesis_dataset_log2 28. Next: the check result, then "ready for PC 2" with the fork commit. evidence.md row 22 (card lifetime, designed) added on ca2-coord (3dc29fe).

20:53 spec text applied for the decided layers (ca2-coord 9b1f849)

docs/spec/01-lottery-hash.md: 1.12 carries epoch_len (the ladder 600 to 7,200, 90% signal at a day boundary, T_epoch and the lead fixed) with 3,600 unchanged on every network; 1.13.1 gains the mixer_mult row (4 under class v3) and the epoch_len row with the signal rule, the FPGA threat and the 600-s floor, and a pointer to era-layout.md for the layer 4 and 8 rows; 1.13.2 carries reserve family R1 = mm8 (uint8, W_new 4, unlock era 4 or 90% signal, the edge vectors, the native paths) and the emulation rule (8x per op, 5% hash-rate cap); 1.13.3 carries the cache growth rule (option C, growth_doublings(d) = floor(log2(1 + d / 1,460)), 256 / 512 / 1,024 MiB at genesis / year 4 / year 12, fill 0.2 / 0.4 / 0.8 s), the step mapping (b) as recommended, the shipped-density reason, and the cache freed after the daily build. docs/spec/04-seeds-and-vdf.md 4.3: the lead and T_epoch fixed at every epoch length. Still to land in the spec from the branches: 1.8.5 (the mixer x4 form, from mixer-x4.md), the 1.13.1 rows for stride, interleave and the window (era-layout.md), 1.5 and 1.8 for the hot table in the added form (hot-table.md), 1.17 and 1.15 for the v3 vectors and the conformance runs, 1.4.5 and 1.4.6 for generator 3 and the class in the pack.

20:56 PC 2 released; the proving fork tip for the merged tree

PC 2 is free (the proving agent's last job closed; the live prover is back on). Proving fork tip ece42979 on 21d4c73c (N = 8, the digest test edit), app 440fd59 or later; its suites ride with the ca2 node suites on the merged tree; its harness on the final tree is running on the Mac. The S_p curve with the miner on the card (PC 2's 5090): empty shard 15,585 MiB 7.5 s; the adopted v1 shard (30,000 pgas, 4.7 M cycles) 22,210 MiB 13.2 s (20,435 MiB, 4.2 s alone); 2.25 M pgas 30,049 MiB 17.9 s; 4.5 M 29,954 MiB 26.3 s; the prototype shard 30,083 MiB 33.3 s. Tiers as the proving agent published them: 32 GB mines and proves today, 24 GB from the fee switch (2.3 GB spare on the adopted shard), 16 GB empty shards only, 12 GB nothing on this build (D2 to the project lead).

PC 2 queue: the ca2 node suites on the merged tree (ca2-v3-node + ece42979) as soon as the node agent sends "ready for PC 2"; nothing else is queued on PC 2.

20:57 the mixer construction is committed on ca2-v3's base

ca2-mixer 0fc0ad1 (rebased onto ca2-v3 6c75dad): LoadClass::MX4 = V3_CLASS (v2 loads, mixer x4, the growth rule; a class v3 program is the v2 program of its seed instruction for instruction, generator 3 in its id); chain_dataset_day has its real body (Shape::for_class_day: cache 2^cache_log2_words(d), dataset 2^dataset_log2_words(D_0, d)); 44 lib + 11 pack tests green, the two pinned v2 packs byte for byte. Next from it: --program-class v3 / --era-hex on the CLI, the pinned v3 packs mx4-genesis and mx4-devnet-epoch0 (era = the genesis-hash stand-in), the design and 1.8.5 spec text, the chip-model row; bit-exactness runs under the run lock now; the verifier and build timings wait for the measure lock (held by a live prover measurement from another worktree, pid 45000, 25 min in at 20:56). The node agent merges ca2-mixer before its fast-time gate, so the gate runs the real construction minus the era and hot fields.

20:58. PC 1: cpu-prove-pc1-small2 running since 20:54:02 (CPU only). PC 2: a fetch job from the aggregation-cost agent (job-fetch-prove-aggcost, 20:55:39) landed after the proving agent's release, so PC 2 is NOT idle for the ca2 suites until that agent's run closes; the suite publish checks node tools/jobs.mjs status for an idle PC 2 first. Readiness in hand: the repro benchmark (8 min, 5090 then gfx1036) waits for a gap; the 5090 power sweep (8 min) follows the era job; Ember Tune's build (8 min, CPU) and run (30 min, both cards) follow; the S_p CPU shard (up to 90 min) is last.

21:00 PC 1 released by the CPU fixture; the repro run has it; PC 2 to the aggregation-cost agent

cpu-prove-pc1-small2 finished 20:59:49Z: the SP1 CPU prover on PC 1 with the miners running: block-56-transfers-3shards shard 0 (200 pgas) 312 s wall, peak RSS 29.5 GB, 978% CPU; block-78-increment 322 s, 30.5 GB; the 5090 untouched (89% mean). The S_p shard job is DROPPED tonight: 312 s for 315 k cycles extrapolates the 60.8 M-cycle shard to many hours of every core and over 30 GB (approximate), which answers the CPU-fallback question (not viable for S_p shards; viable for empty or tiny shards only). "go PC 1" given to the reproducible benchmark (the 5090 then the gfx1036, about 8 min); the era job follows it, then the 5090 power sweep, then the hot table, then Ember Tune's build and run. "go PC 2" given to the aggregation-cost agent (20 min, GPU proving with the miner on then paused, prover restored); then the ca2 node suites, then the repro run's PC 2 slot (10 min).

21:01. GitHub Actions is in a major outage (six queued runs since 19:26Z, none acquired); the coordinator gave the 0.3.10 shipper the fallback at 21:00Z: build the Windows installer on PC 1 (MSVC window host, the payload under Git Bash, Inno Setup; CPU only, about 15 min). PC 1 order now: the repro run (until about 21:09), then the 0.3.10 installer build (the fleet's release, ahead of every measurement), then the era job, the 5090 power sweep, the hot table, Ember Tune. Any measurement that straddles the build window is re-run.

21:02. The Mac measure lock, found by lsof: the recorded holder pid 43916 is dead; the files are held open by two WAITERS, the readwidth agent's re-queued footprint loop (pid 78893, holding the measure and build files 11 min, waiting for the three build slots, which cargo tests keep re-acquiring: a convoy) and the epoch agent's compile-ahead measurement (pid 78476, waiting behind it). The readwidth agent is asked to kill 78893; the epoch measurement then runs when the build slots drain. Two defects for the next cut (one task filed): the status line shows the last writer, not the holder; a measure waiter can hold the master lock while build slots keep being granted to new builds, so a measurement can wait indefinitely under a steady stream of cargo tests.

21:04 proving v1 handoff received; the AMD-proving line on the site; amd-prove merged

Proving v1 for 0.3.11 (rollout plan 8a): fork ece42979 on 21d4c73c, harness PASSED (21 checks) in 244.4 s at 20:56:45Z on the final fork tree, override fields and the mixed-fleet rule recorded; the app's final hash follows its gate tests (90d3299 before it). Branch amd-prove (f1d7a7d) merged into ca2-coord (the append-only bench-log conflict kept both entries); its finding: no zkVM proves on an AMD GPU as of 5 October 2026 (SP1 CPU and CUDA; RISC Zero and ICICLE add Metal; nothing for AMD), the CPU fallback is about 5 minutes per small shard at 30 GB RSS, not a tier. The public line is applied on ca2-coord (1c8439f) with the proving agent's measured tiers in place of the doc's 16 and 20 GB: "Proving needs an NVIDIA card with 24 GB or more (32 GB until the fee switch of 6 October 2026; from it a 24 GB card mines and proves on the same card: 22.2 GB peak with the miner on). AMD and Apple cards mine. A prover for them lands when a zkVM ships one." It replaces "the card mines and proves" on the index, the litepaper's vs RandomX row and proving section, and the miner page (title, meta, hero, feature). Left as it was: the app's Proving tile text (the proving agent's).

21:05 the node merge for 0.3.11 is done; the hot-table package is ready; unproven_daa 10 at fast time

ca2-v3-node: 2e464e81 (the class switch, the era stand-in, the template, the miner) and the merge of proving-v1 ece42979 = ba43cf0f. One conflict, params.rs's digest-test edits array, resolved by keeping both sides (14 entries); every other shared hunk auto-merged as separate blocks. The merged tree is in cargo check (with igneum-exec and kaspa-p2p-flows); "ready for PC 2" follows with ba43cf0f once it and the Mac pow and consensus-core tests are green. Main repo ca2-v3: ca2-mixer fast-forwarded (66eeba3), then 43ca289 (the fast-time and redteam overrides carry the proving v1 fields and pow_genesis_dataset_log2 28; class-v3.mjs prints the dataset build ms per epoch). proving_v1_unproven_daa is a DAA clock (exec/src/proving.rs segment_status), so 10 at fast time is right (the proving agent confirms; its harness passes --unproven itself). The PC 2 suite command is the node agent's final shape (75-minute budget, six node crates plus the app tests), published from the ca2-v3 worktree when PC 2 is idle (the aggregation-cost job closes at about 21:21).

ca2-cache 196db96 (rebased on ca2-v3 464d6e1, the added form): 47 + 13 tests; the three added packs bit-exact on Metal and Apple OpenCL (fingerprints at 2^20, base 0: hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5; 96/96; hot table PASS); first-pass rates under load (not numbers): 25.7 / 23.9 / 22.9 MH/s against v2 27.7 (the probe predicts 0.96 / 0.94 / 0.94); the measured rows wait for the measure lock behind the epoch measurement. PC package: ~/Desktop/igneum-ca2-hot.zip sha256 bd49faa1c9d48024f49c615481faff5c68a4c09f0889dbaf009c208674d67b3f (workers 956c4ab3... and 32d3d343... from 196db96, eight packs), fetch-ca2-hot-20261005, playbooks ca2-hot-5090-bench.ps1 and ca2-hot-9070-bench.ps1 (gfx1036 fallback). Its go follows the 0.3.10 build on PC 1 unless the era package is there first.

21:11 the 9070 XT is back; PC 1 to the 0.3.10 build; the repro run failed at parse time

The repro run (run-repro-pc1-20261005, 21:04 to 21:09:24Z) switched every card off and on and restored them, and found the 9070 XT (gfx1201) ON the bus again (opencl:1; the app switched it off and on), so the era and hot-table jobs run their gfx1201 halves as planned and the G1 ruling's fallback is not needed unless the link drops again; recorded in the rollout plan's hardware events. The run produced no numbers: repro.ps1 failed with a PowerShell parse error on each card (MissingEndParenthesisInExpression), the second playbook tonight that passed no local parse (the Mac has no pwsh); the class fix ordered: every PowerShell playbook parses itself on the PC as its first step (System.Management.Automation.Language.Parser::ParseFile, errors printed, non-zero exit), the way the dot4 playbook gates its bash body with bash -n. "PC 1 is yours" given to the 0.3.10 shipper at 21:10 for the installer build (CPU only, about 15 min); then the hot-table and era jobs, the 5090 power sweep, the repro re-run (about 22:00), Ember Tune.

21:20. Proving v1 app branch final: 6dc686a on 5b0d54f (provedefault 6 of 6; the app's 0.3.11 inputs are complete on that side); fork stays ece42979 (merged into ca2-v3-node ba43cf0f). The 0.3.11 app tree = the 0.3.10 release tree 5b0d54f + proving-v1 6dc686a + whatever the app needs for v3 (the --prepare-packs flag is already there; the Metal worker change is in the worker, not the app); the 0.3.11 main tree = master + ca2-v3 (igneum-pow, workers, fast-time, docs) + ca2-coord (the plans, the spec, the site copy).

21:20 layer 9 complete (ca2-epoch 4300608, e95e8b5)

Design (reserve-only): epoch_len base 3,600, ladder 600 to 7,200, SIGNAL ONLY (the era stream consumes draw 8 and ignores it); day-anchored epochs so a change lands in days; VDF option A (T_epoch 600 s and the 1,200-s lead stay genesis constants: the program is known 600 s ahead at every length, grinding margin 300x; option B rejected at 50x and a 200-s-deep checkpoint); REF_WINDOW_V2 = min(600, L); the 144-s settle per step is 24% of a 600-s epoch. Floor 600 DAA s: the slowest compile-ahead is the variant race at 38 s on the Mac and the 5090, 6.3% of the epoch and inside the window.

Card Compile-ahead Share of a 600-s epoch
M5 Max Metal, race off (measured 21:18 UTC: 15.9 / 17.7 / 20.4 ms over 10 fresh programs, pack 79 ms then 1 ms) 0.5 s (1.8 s cold) 0.1%
M5 Max, race on (M11) 38 s 6.3%
RTX 5090 NVRTC, race off (M11, hot-swap) 1.0 s 0.2%
RTX 5090, race on 38 s 6.3%
RX 9070 XT OpenCL 0.31 s + compile owed
Intel UHD OpenCL (M11) 6.4 s 1.1%
Radeon iGPU under load (M11) 124 s with the dataset 20.7% (needs per-day dataset reuse before any signal below the base)

FPGA citations: PRflow (FPT 2019) 42 min typical, 160 min worst for a monolithic Vivado compile; Aldec hours on Virtex UltraScale; partial reconfiguration milliseconds per region (ICAP 400 MB/s) shortens the load, not the compile; at 600 s a per-program bitstream mines 0% of each epoch, 47% at 3,600. Riders: the race defaults off (M11: base wins on both cards; 6.3% at the floor); per-day dataset reuse in the workers for the iGPU tier. Owed: the 9070 XT clBuildProgram time, the difficulty settle at a 15% step and 600-s epochs in sim.py, the 3-bit signal encoding against Kaspa's version bits (spec 5.8).

21:20 the hot table in the added form, measured on the Mac: the big-die chip comes out ahead

ca2-cache (hot-table.md 6.2 to 6.4): M5 Max, 21:03 to 21:19 UTC, load average 7 to 14, Metal packbench 2^24 x 5 (GPU time) and Apple OpenCL --bench-pack; v2 in the same session 27.63 / 27.59 MH/s.

Pack Metal / OpenCL MH/s g against v2 (probe predicted) Fingerprint (2^20, base 0) Verifier ms per warp (v2 0.602) Fill, one core
hot32k4a 25.76 / 25.72 0.93 (0.96) 8a3414735db4523c 0.631 21.7 ms
hot64k4a 23.92 / 23.87 0.87 (0.94) 45668f34105f6307 0.609 43.3 ms
hot96k4a 22.92 / 22.88 0.83 (0.93) af763997dfee4c82 0.614 64.9 ms

Reading: on Apple the added form costs 7 to 17% of the rate for four extra loads per iteration, more than the probe predicts as the table grows; only the 32 MiB table is near free. Chip arithmetic with these g: a chip serving H from DRAM keeps 0.86 / 0.92 / 0.96 of its gain; a 100 mm^2 die with the SRAM keeps 0.96 / 0.92 / 0.88; a 750 mm^2 die with the SRAM comes out 6 to 15% AHEAD, because the GPU pays the hits in rate and a big die pays them in 1.6 to 4.8% of area. So on the Mac's numbers layer 5 does not pass its own test; the decision waits for the 5090 (96 MiB L2) and 9070 XT (64 MB Infinity Cache) rows, where the hits may be near free (g close to 1). Rule for the decision: layer 5 goes into v3 only if, on every card we own, g is at or above 0.97 at the chosen size AND the on-die-cache chip row (chip-model-v3.md) moves down with it; otherwise layer 5 is out of v3 and stays a measured option for 3.0.

21:21. PC 2: agg-cost-pc2-1 closes at about 21:25Z (its own-miner phases ran the iGPU miner by a script fault; the curve is unmeasured); "go PC 2" given for agg-cost-pc2-2 (about 17 min, to 21:44Z), release due by 21:45Z; the 0.3.11 suites on ba43cf0f take PC 2 next (the node agent's Mac suites, release build and fast-time gate are in flight). PC 1: the 0.3.10 installer build (from 21:10); the hot-table job (package in hand) goes the moment the shipper reports the build closed; the era package is still being built.

21:22. Ember Tune (f9bf552 on ember-tune) is queued: its Windows app build (8 min, CPU) in the first gap after the installer build, its baseline run on the 5090 (8 min, no prompt; the 9070 XT if it can be taken without one) after the repro re-run. PC 1 order: the 0.3.10 installer build (running), the hot-table job, the era job, the 5090 power sweep, Ember's build in the first CPU gap, the repro re-run (about 22:00), Ember's run.

21:22. PC 1: the 0.3.10 installer build job (job-fb-installer-pc1) FAILED at 21:20:05Z, exit 2 after 2 s (the script, not the build); the shipper is asked to republish within 5 minutes or yield PC 1 to the hot-table measurement (10 min) and follow it. PC 2: agg-cost-pc2-1 still running at 21:21.

21:23. The shipper republished the installer build as fb-installer-pc1-2 (the first exit 2 was Test-Path on a \wsl$ root path refused to the non-elevated session; the engine is now copied out with wsl -u root), cap 25 min, so PC 1 is the build's until about 21:50; the AMD presence probe (10 s, nothing held) runs beside it. Then: the hot-table job, the era job, the 5090 power sweep, Ember's build and fetches, the repro re-run, the AMD sweep (26 min), Ember's run (25 min).

21:23 the era package is ready; the V3_CLASS composition rule

ca2-era PC 1 package: ~/Desktop/igneum-ca2-era-pc1.zip sha256 f26f997602d94b0a408a6974ee484a7b3249ce18dc91dc00783aa9754e5ff040 (workers 5dd3bc16... and 9615efb3... from ca2-era on ca2-v3 464d6e1 with the pack-loop packfile.h merge and the host.c mh_word fix; packs v2 and era-0 to era-5, the six as class v3 chain packs, generator 3, program id 6b02c7c49eb126bd shared, width 4 B, the same devnet epoch seed and day); fetch-ca2-era-20261005; relay/playbooks/ca2-era-pc1.ps1 (the 5090 then gfx1201, gfx1036 fallback; 6 to 10 min). Mac bit-exactness on these packs: Metal 6/6, Apple OpenCL 6/6 (same fingerprints), CUDA CPU emulation 6/6. Its go follows the hot-table job, about 22:00.

Composition rule for the integration (three branches define V3_CLASS): V3_CLASS = LoadClass::MX4's fields (v2 loads, mixer_mult 4, growth on) + era: None (drawn per program inside generate_from_seed_bytes_program_class from the era bytes, LoadClass::era(V3_CLASS, era, &V3_ALLOWED)) + hot: None until the PC rows decide layer 5; the layout rides with the program (program.class.layout() in the interpreter and Epoch::dataset_word), so one day cache (chain_dataset_day) serves every era. The era agent rebases onto ca2-v3 HEAD with that literal; the node agent takes it into the integration tree. The era stream is seeded from "igneum-era/" || E_n (the index dropped: E_n commits to n through the VDF input); the spec text in era-layout.md says so.

21:24. The 9070 XT dropped off PC 1's bus again (probe #205 at 21:22:59Z: no Sonnet or USB4 router device; the third drop today; rollout plan 7b updated: the link is flapping). The era and hot-table jobs run their gfx1036 fallback unless the card is present at run time; the AMD sweep slot is conditional on a presence probe at 22:00. A probe reading of "miners off at 0.0 MH/s" was wrong: the app log shows the 5090 at 124.5 MH/s through 21:23:26Z; the AMD agent fixes its probe's precondition. The 0.3.10 installer build fb-installer-pc1-2 runs on PC 1 (cap 25 min from 21:23).

21:25 the mixer x4 construction is bit-exact on the Mac; the chip row reads 1.84x (thin)

ca2-mixer commits: 0fc0ad1 (construction, V3_CLASS = mx4, chain_dataset_day body), 66eeba3 (the pinned v3 packs mx4-genesis and mx4-devnet-epoch0, --program-class v3 / --era-hex), e4c04a7 (tests/mixer.rs: 200-program v3 fuzz, stats, edges, determinism; scratch.rs cherry-picked, 7 of 7), 7ce8d1e (docs/analysis/chip-model-v3.md). Spec text for 1.8.5 and 1.13.3 in docs/plans/mixer-x4.md section 2: the multiplied mixer with keys (r m + j + 1) x 0x9E3779B9; option C as doublings(d) = floor(log2(1 + d / 1460)); the day table with the verifier fill per step.

Bit-exactness (run lock): both v3 packs on Metal and Apple OpenCL, 3/3 standalone and 3/3 in batch, 96 of 96 lanes, dataset head / MASK / 64 samples PASS, one fingerprint per pack across both harnesses (6f48d5a2aa0dbe5f, 73caaebb28e808fe); the 200-pack v3 fuzz on Metal 200 of 200, every tenth on Apple OpenCL 20 of 20; CPU 44 lib, 12 packs, 7 scratch, 4 mixer tests; v2 exports IDENTICAL. Indicative (run lock, not a number): the Metal 1 GiB build at x4 30.2 ms (mx4-genesis) and 21.7 ms (mx4-devnet); the measured verifier and build timings wait on the measure lock.

The chip row (chip-model-v3.md), v3 at x4: 599,040 ops per hash; the on-die-cache recompute chip at 50 T op/s does 83.5 MH/s, 0.61x bare against the 5090's 136.1, 1.84x with the 3x fixed-function factor, 1.53x with the 128 mm^2 mirror deducted at equal silicon. With the hot table in the ADDED form at the Mac's g: 1.98x (32 MiB) and 2.12x (64 MiB) at equal budget, 1.60x / 1.67x with the SRAM deducted: the added hot table costs the card and not this chip, so it moves the row the WRONG way on the Mac's numbers. The claim holds "under 2x" on the equal-silicon convention, and on the equal-budget one only without the hot table; thin everywhere (a 3.3x factor or 10% on the budget reads 2.0x). Next lever: x8 (0.31x bare, 0.92x with the factor).

Consequence for layer 5: unless the 5090 and 9070 XT rows show g at or above 0.97 (the hits near free), the hot table stays OUT of v3 tonight (the rule in the 21:20 entry) and the public level 3 names it as a measured option, not a lever.

21:25. "ready for PC 2" from the node agent: fork ca2-v3-node 79bd8e10, Mac checks and tests green (rollout plan G6 row); the expected 0.3.11 digest with the two new fields at never is c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c. The coordinator publishes the suite job from the ca2-v3 worktree at 21:45 when the aggregation-cost agent releases PC 2, on the worktree's tip at that moment (the era and cache merges go in first if the era commit arrives in time).

21:27. The installer build's second attempt failed in 4 s on a path (the engine is not under /root/igneum-build/app/...); attempt 3 (fb-installer-pc1-3) is publishing with a find-based path; rule: if it fails inside 5 minutes, the hot-table measurement takes PC 1 before attempt 4. The shipper's jobs started no second app instance (the failed attempts exited before any exe ran); its post-build listing of igneum-app.exe processes settles the "off" reading.

21:27 decision rule for the mixer: x8 beside x4

Delegated (coordinator, the project lead's "as strong as the measurements allow"): the mixer agent builds LoadClass::MX8 beside MX4, exports mx8-genesis and mx8-devnet-epoch0, runs the Mac bit-exactness, and measures v2, x4 and x8 in one measure-lock session (verifier ms per warp on one core, avg of 50 and worst cold; the 256 MiB fill; the Metal 1 GiB build); the 5090 and AMD daily-build times come from a 2-minute prepare job on PC 1 after the era job. x8 goes into v3 if the per-warp verify stays under 10 ms on one core AND the daily build stays under 1 s on every card we own; else x4 with the thin margin stated (1.84x with the factor) and x8 named as the next lever (0.92x with the factor, from the m16 table). The vectors are re-cut once after the choice. The hot-table rule stands (into v3 only at g >= 0.97 on both PC cards).

21:28. Installer attempt 3 failed at 21:27:20Z (exit 2 in 4 s: an inline bash -c string lost a quote through PowerShell; the fix is the 0.3.6 cut's: the WSL part as a file run with bash , plus a read-only path probe before attempt 4). By the rule, "go PC 1" went to the hot-table job at 21:28 (10 min). PC 1 order from here: the hot-table job, the shipper's path probe and attempt 4 (about 16 min), the era job, the 5090 power sweep, the mixer daily-build job (2 min), Ember's build, the repro re-run, the AMD sweep (conditional), Ember's run.

21:31. Consequences round 3. C23: the x8 table gains a gfx1036 row (the integrated tier builds the dataset per PREPARE: 7 to 12 s at x1, 55 to 124 s under load, so 28 to 48 s an hour at x4 and 56 to 96 s at x8, every boundary missed) and a scaled 8 GB-class row; the "under 1 s" rule applies to the discrete cards' daily build; for the integrated tier either per-day dataset reuse in the three workers lands with v3 (the node agent is asked whether it is bounded tonight) or the level 3 page says the iGPU tier mines v3 with a restart per epoch; recorded with the x4/x8 choice. C24: two inline bash bodies lost a quote through PowerShell tonight (amd-prove's awk at 20:44, the installer's bash -c at 21:27); the reviewer's sub-agent adds a repo-wide CI check (every inline bash body through bash -n, an unextractable one fails CI) and the convention line in packaging/README-ship.md; no collision with the PC 1 queue.

21:31 GitHub's runner recovered: the 0.3.10 rollout starts; the restart rule for every PC job

The 0.3.10 Windows run 37374158235 went green at 21:30:29Z on the release tree; the three fb-installer jobs are done and expired; nothing more of the shipper's touches PC 1. The rollout runs now (manifest, update-now, hand nodes, seed, digest sweep): every app restarts once within the next quarter hour. Rules: the shipper is asked to hold PC 1's update-now until the hot-table job releases (about 21:40) and the era job starts only after PC 1's new STATUS line; the 0.3.11 suites publish on PC 2 only after PC 2's app shows 0.3.10 (a build job dies with the app); any measurement that straddles a restart is re-run; every PC job's RESULT lines carry the app version before and after.

21:32 per-day dataset reuse: Metal has it, CUDA and OpenCL do not (0.3.12); the gate re-runs after a script fix

Node agent: the Metal worker's ServeStore already keys datasets by day and programs by epoch (a prepare on a resident day builds the program only); the CUDA worker.cpp and OpenCL host.c bundle program, cache, dataset and the self-test in one Pair, and splitting a Day object out touches buffer ownership, releasePair, the prepare thread and the self-test in both: over an hour, not shipped untested tonight; first item after the publish (0.3.12; next-cut list). Decision recorded (C23): the level 3 page states that the integrated tier on the one-click workers mines v3 with a restart per epoch (public copy and rollout plan updated).

Gate G4: the first fast-time run failed at node start on the script, not the node: JSON.parse turned a never height (18446744073709551615) into 1.8446744073709552e+19 and the node refused the override; fixed as text merging in class-v3.mjs and simnet.mjs (d5ff532; no other script in tools/, infra/ or sim/ has the shape). The gate runs again on the mixer-x4 class (binaries from 79bd8e10 + ca2-v3 66eeba3). Main-repo tip for the suite job title: d5ff532 (era and cache not yet merged).

21:34. PC 2: agg-cost-pc2-1 closed 21:25:11Z (done, miners and prover back on); agg-cost-pc2-2 went out at 21:33Z (a missed close), self-limited to 16.5 min, closes about 21:51Z; the 0.3.11 suites publish after it AND after PC 2's app shows the 0.3.10 STATUS line (the update-now goes to PC 2 now). PC 1: the hot-table job runs (release about 21:40); PC 1's update-now follows the release; the era job starts after PC 1's 0.3.10 STATUS line.

21:36 layer 5 measured on both PC cards: OUT of v3

Hot-table jobs on PC 1 (fetch 21:29:01Z; run-ca2-hot-5090 126 s, closed 21:31:38Z; run-ca2-hot-9070 247 s, closed 21:35:39Z; both cards restored; the 9070 XT WAS on the bus, full path; app 0.3.9 throughout, no straddle). All eight packs bit-exact on the 5090 and the 9070 XT with the Mac's fingerprints.

Pack RTX 5090 MH/s (v2 136.1) RX 9070 XT MH/s (v2 18.15) M5 Max g
hot32k4 (replaced) 146.6 (x1.08) 19.79 (x1.09) x1.22
hot64k4 140.8 (x1.03) 18.73 (x1.03) x1.12
hot96k4 138.5 (x1.02) 18.33 (x1.01) x1.05
hot64k2 137.5 (x1.01) 18.17 (x1.00) x1.00
hot64k8 163.6 (x1.20) 22.32 (x1.23) x1.71
hot32k4a (added) 118.7 (0.87) 15.27 (0.84) 0.93
hot64k4a 115.4 (0.85) 14.62 (0.81) 0.87
hot96k4a 114.4 (0.84) 14.56 (0.80) 0.83

Decision (the 0.97 rule): layer 5 is OUT of v3. Neither card keeps even the 32 MiB table resident while the 1 GiB dataset streams (the replaced form gains 1.02 to 1.08x at k = 4 against an ideal 1.33x), and the added form costs 13 to 20%; the chip row moves the wrong way with it. Layer 5 stays a measured option for 3.0 (a table small enough to stay resident, or a different access pattern). The probe rows and the writeup follow on ca2-cache. PC 1 is released to the shipper for PC 1's update-now; the era job starts after PC 1's 0.3.10 STATUS line.

21:38. ca2-cache final: 2de19e5 (nine commits from 55e285c, on ca2-v3 464d6e1); hot-table.md carries the probe rows for all three cards (5090 112.6 G loads/s at 32 / 64 / 96 MiB inside its L2 against 17.6 at 1 GiB; 9070 XT 9.88 / 9.47 / 8.18 / 2.43; M5 Max 21.7 / 12.8 / 12.3 / 3.50), the PC tables with g, the chip arithmetic at the measured g, the decision, the 3.0 note ("what would make it pay": a resident size found by a hash sweep below 32 MiB, k only with residency, a line-unit or streamed access shape) and the unverified list; the bench-log entry and two addenda carry the job ids and the worker sha256s. The probe promises full hits inside the 5090's L2 but the hash gets 2 to 8% at k = 4 because the streaming dataset evicts the table.

21:39 gate G4 run 1 PASS on the mixer-x4 class

Fast-time 3-node network, 21:33 to 21:38 UTC (fork 79bd8e10 + igneum-pow 66eeba3): the switch line on 3 of 3 nodes (active from epoch 3, DAA 150 rounded up to 180), templates class 2 then 3 from epoch 3, 181 blocks before and 124 after the boundary, program ids agree on all three miners (v2 e0 to e2, v3 e3 to e5), 0 rejected on miners and nodes, one sink on all three (082fd39ba65df2ff, 304/304/304), a new (day, class) cache 177 to 235 ms on one core. Main-repo tip bf04c56 (the doc, the summary JSON, the script's --connect fix). Run 2 on the composed class follows the era commit and the cache merge (rebuild about 10 min, gate 5 min). The x4 dataset build time comes from the mixer's PC prepare job (the CPU miner derives words from the cache).

21:40 x8 built and bit-exact beside x4; the mixer PC job retargeted to PC 1

ca2-mixer 504cae4 (LoadClass::MX8 "mx8", packs mx8-genesis and mx8-devnet-epoch0, the fuzz takes IGNEUM_MIXER_CLASS, playbooks with the x8 packs) and fe4e193 (mixer-x4.md per-tier build table and the x4/x8 rule; chip-model-v3.md with the mixer row as the headline, the layer 5 rows kept as measured not adopted with the PC g beside the Mac's; x8 rows 0.31x bare, 0.92x with the factor, 0.76x at equal silicon at year 0). x8 bit-exactness (run lock): both packs on Metal and Apple OpenCL 3/3 + 3/3, 96 of 96 lanes, one fingerprint per pack across both harnesses (7c28cfb06c5c65a9, bbb183f72692f840); 50-program x8 fuzz on Metal 50 of 50, every tenth on OpenCL 5 of 5. Indicative Mac builds (run lock): 30.0 ms at x8 against 30.2 at x4 (genesis pack), 22.0 against 21.7 (devnet pack): the Mac's build is latency-bound. The timing session (verifier v2 / x4 / x8, the fill, the build) is queued behind the measure lock. PC job: mixer-x4-pcjob.zip sha256 55a2913cb8790cd3b106dc3d0d29b6e2952b5a08378cd815935915ab82898a8e (v2 control plus the mx4 and mx8 genesis and devnet packs; a prepare per pack printing the worker's cache and dataset build ms); retargeted so both halves run on PC 1 (its 5090 and its AMD card), after the era job, about 22:05.

21:41 the mixer timing session: x8 passes the verifier half of the rule

M5 Max, one core, measure lock, 21:40:12 to 21:40:23 UTC, on a loaded box (load average 5.6 one-minute, 26 fifteen-minute: other agents' unlocked processes), so the absolute figures are about 2x the quiet 0.604 ms v2 baseline and the RATIOS are the measurement (two rounds, within 4%); a quiet-box re-run is owed for absolute numbers.

Class Verifier ms per warp, avg of 50 (round 1 / 2) Worst cold unit Ratio to v2
v2 1.361 / 1.310 1.579 1
x4 (genesis; devnet pack 1.923) 1.956 / 1.923 2.043 1.45x
x8 (genesis; devnet pack 2.972) 2.785 / 2.790 2.942 2.1x

256 MiB cache fill on one core 172 to 175 ms. Metal 1 GiB build, GPU time: v2 21.0 ms (29.7 cold), x4 20.9 / 21.0, x8 21.9 / 21.9: the Mac's build is bound by the 8 dependent cache-line reads per item, not the arithmetic, so the "under 1 s on every discrete card" half of the rule is decided by the 5090 and 9070 XT rows of the mixer PC job (by the M16 arithmetic the 5090 is 54 ms at x4 and 107 ms at x8 if arithmetic-bound, 13.4 ms if latency-bound: far under 1 s either way). Verifier half: x8 passes with 7.1 ms of the 10 ms gate to spare on the loaded core (about 1.3 ms on a quiet core, approximate); x4 leaves 8.0 ms. Chip row at x8: 1,198,080 ops per hash, 41.7 MH/s, 0.31x bare, 0.92x with the 3x factor, 0.76x at equal silicon; x4 1.84x / 1.53x. C19 at these loaded figures: shares per core per second 735 / 511 / 358 (v2 / x4 / x8), a 22,000-member pool at one share per 10 s needs 3.0 / 4.3 / 6.1 cores; IBD over 108,000 headers on one core 2.4 / 3.5 / 5.0 min. Provisional choice under the rule: x8, confirmed when the PC build rows land (about 22:10).

21:41 the era draw passes the 5% rule on the Mac; the final era package; the merges for G4 run 2

ca2-era 9f98af2 (one commit on ca2-v3 HEAD; 48 lib + 17 integration tests; the pinned v2 packs byte-identical; mx4-genesis unchanged; mx4-devnet-epoch0 re-exported with the era inside the class). V3_CLASS = LoadClass { era: None, ..LoadClass::MX4 }; the era class is drawn inside generate_from_seed_bytes_program_class from the era bytes; chain_dataset_day and Epoch::dataset_word compose unchanged; an era program takes 11 draws per instruction. Mac (M5 Max, 21:38Z, Metal packbench 5 x 2^24, a loaded box): v2 27.68 MH/s; era-0 to era-5 28.58, 28.48, 28.35, 28.38, 28.49, 28.48: min 28.35, median 28.48, max 28.58, SPREAD 0.8% (under the 5% rule); 3/3 vectors and the in-batch vectors PASS on every pack; CPU verify 1.319 to 1.345 ms per warp against v2 1.334 in the same loaded run (quiet re-run owed). FINAL PC 1 package: ~/Desktop/igneum-ca2-era-pc1.zip sha256 f79c0607bb4187e3cf16fce3f533e7d525673d766d7edb799f27fd81af5dcee1 (workers 0fbfd50a... and 8c8caff7... from the merged tree; packs re-exported, attempt 0, program id 73bcbfe8ccf988f1 in all six with the era seed beside it); fetch-ca2-era-20261005; ca2-era-pc1.ps1. The 0.3.10 update-now reached every machine at 21:39:59Z (manifest live 21:33Z); the era job starts on PC 1's 0.3.10 STATUS line. The node agent merges 9f98af2 then ca2-cache 2de19e5 into ca2-v3 for G4 run 2 and the suites' tip.

21:42. ca2-v3 now carries the era draw: ca2-era's tip b105a55 (9f98af2 rebased onto the node agent's 88dafbc) fast-forwarded, no conflict; the suite job's main-repo tip is b105a55. ca2-cache 2de19e5 does NOT merge (it bases on 464d6e1, before the mixer and era commits rewrote the class literal, the load emitters, the pack fields and the pinned-pack tests: 8 files, 35 hunks); the node agent aborted cleanly and the cache agent is rebasing onto b105a55 as a squashed commit with hot: None kept; the composed class under test is unchanged by the cache code, so the igneum-pow and packfile checks, the igneumd and igneum-miner rebuild and gate run 2 proceed on b105a55 now.

21:45. PC 2 defect: since a job's /api/resume at 21:25:11Z the 0.3.9 app answered ok and never restarted the NVIDIA miner (nor the iGPU one): the 5090 worker "off" at hash 0 holding 1.7 GB, so agg-cost-pc2-2's mining phases are void (its idle phases run; closes about 21:52Z) and the devnet has been short PC 2's rate since 21:25. The 0.3.10 restart should bring the miners back; the shipper confirms PC 2's 5090 STATUS rate after the 0.3.10 line, else the aggregation-cost agent's restore script (tools/proving-v1/pc2-agg-cost-restore.ps1, 30 s) runs. Defect for the next cut: a resume that answers ok without a miner restart; the app must re-check the miner processes after a resume and report a failure. The aggregation-cost agent gets a 20-minute re-run slot on PC 2 after the 0.3.11 suites.

21:46 PC 1 is on 0.3.10; the era job has the go

PC 1 restarted on 0.3.10 at 21:40:41Z (engine run win-ae432dc7-20261005-214041), the 5090 at 141.4 MH/s by 21:45:17Z; "go PC 1" to the era job at 21:46 (fetch-ca2-era-20261005, zip f79c0607...; the 5090 then the AMD card; 6 to 10 min). The mixer daily-build job follows it, then the 5090 power sweep, Ember's build and fetches, the repro re-run, the AMD sweep (conditional), Ember's run. PC 2 is still on 0.3.9 at 21:45 (its restart pending); the 0.3.11 suites publish after its 0.3.10 line and a confirmed 5090 rate. Every measurement before the restart on PC 1 (hot table 21:29 to 21:35) stands: it did not straddle.

21:50. The resume fix (the miners not restarted after POST /api/resume: PC 2 since 21:25Z tonight, the Mac this afternoon) is assigned to the proving agent on 0.3.11's app branch (engine.rs resume: restart every enabled card's worker and a stale pack export, re-check within one tick, a state-machine unit test plus the known-failed case from PC 2's log); rollout plan 8a. If its commit is not in hand at the app cut, it heads 0.3.12's list.

21:52 gate G4 run 2 PASS on the composed class; the 0.3.11 suites are packing for PC 2

G4 run 2 (21:46:36 to 21:51:31 UTC, b105a55 = era + mixer, hot None): every check true; 181 / 124 blocks around DAA 180; v3 ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners; the v2 epoch-0 id 8f8806638d59850f unchanged from run 1; 0 rejected; one sink 712c1b212091dcdc at 303/303/303; the switch line on 3 of 3; cache ready 178 / 181 ms. G4 is GREEN. The 0.3.11 suite job is packing from the ca2-v3 worktree (fork 79bd8e10, main b105a55; build-inputs.zip 9,563,672 bytes sha256 bf89ab4c...) for PC 2, which is on 0.3.10 since 21:49:41Z with its 5090 worker back on the first try.

21:53. The 0.3.11 suite job is published: build-20261005-215219 to PC 2 (fork 79bd8e10, main b105a55; linux build 30 min, tests 35 min: kaspa-consensus-core, igneum-exec, kaspa-pow, kaspa-consensus, igneum-miner, kaspa-p2p-flows, igneum-app); PC 2 is on 0.3.10 with its 5090 back. The worktree freeze is lifted for the node agent (the run-2 summary commit, then the cache merge). The proving agent's resume fix waits on a test run behind the Mac's held lock slots.

21:55. The 0.3.10 rollout: both PCs mine on 0.3.10 with the rebuilt workers (PC 2 120.6 MH/s at 21:54:33Z, PC 1 141.3); the hand nodes (21:49:38Z, 21:49:50Z) and the seed (21:50:15Z) on 21d4c73c, digest 1f4b4425 everywhere; not yet on 0.3.10: the US laptop 37ba0461 (installer downloaded 21:40:52Z, app not back after 13 min; nothing to drive from here) and Sam's Mac (quit since 20:47Z). Open on PC 2: the prover fails with "CudaClientError: Connect(PermissionDenied)" since the restart (three shards 21:49:56 to 21:50:32Z); the likely cause is the sp1-gpu-server socket handling of the aggregation-cost jobs; the proving agent owns it and publishes a fix after the suite job (PC 2 is the suite job's until it closes). The ca2-v3 tip is 63dabb2 (run-2 summary and doc, the G6 job id recorded).

21:56 gate G6: the first PC 2 job failed on the known kaspa-consensus flake; split re-run

build-20261005-215219 (21:53:01 to 21:55:48Z): Linux build ok (igneumd 49,164,264 bytes sha256 11979b49..., igneum-miner d25a8270..., igneum-app 68007173...), igneum-app tests 78 + 26 + 8 passed; the node stage exit 101: kaspa-consensus 96 passed, 1 failed, processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(..., 487112384, 487129578) in mine_on_all. This is the flake the 0.3.10 cut met on the same base 21d4c73c under the six-package parallel run (release-0.3.10.md: it passed alone twice, job build-20261005-182804, 97 passed), a timing-dependent difficulty in the test helper, not a v3 change (v3 touches no finality code). Per the gate rule the publish stops here until the suites are green: job 2 of 3 (kaspa-consensus alone) is published now; job 3 of 3 (the other five crates) follows; the flake itself goes on the next-cut list (make mine_on_all deterministic under parallel load).

The proving v1 app branch is final at a223ca9 (6dc686a plus the resume fix with the PC 2 case as a unit test); rollout plan 8a updated. PC 2's prover fault is the root-socket class (agg-cost-pc2-1 ran the host as root in WSL2 and left /tmp/sp1-cuda-0.sock owned by root; the app's prover has failed every shard since 21:25:24Z); the proving agent's pc2-socket-fix.ps1 (60 s, miners untouched) runs between my two suite jobs.

21:57 the era PC job is done; a 2.2x CPU-verifier regression on ca2-v3 HEAD (gate item); the PC 2 order

Era job run-ca2-era-pc1-20261005 (exit 0, 301 s), both cards restored, the 9070 XT ON the bus at run time (gfx1201, 32 CUs): first rows 18.96 to 19.21 MH/s on the era packs on the 9070 XT, self-test PASS, the 2^24 fingerprints equal to the Mac's (era-4 3ace11ad84c053ae, era-5 a8897d82adceb4a1); the full 5090 and 9070 XT tables with the spread per card follow. "go PC 1" to the mixer daily-build job at 21:57.

GATE ITEM (found by the era agent, two binaries on the same v2 input, checksum 19297e99c7b9a55e, same minute, load 4 to 5): the CPU verifier on ca2-v3 HEAD (88dafbc) takes 1.332 ms per warp and on the era branch 1.310, against readwidth's binary at 0.604 and 0.606: the mixer branch's derive_items / Shape path costs 2.2x at m = 1, on the v2 path the live devnet verifies with. The mixer agent's "loaded box" reading of its 1.31 to 1.36 ms v2 figure was the code, not the load. Ordered: find and fix on ca2-mixer, restore v2 to within 5% of 0.604 ms measured the same way, re-measure v2 / x4 / x8 on the fixed binary (the C19 figures too), then the node agent merges it; the publish waits on it (G3's suite does not catch a slowdown; the 10 ms gate and the pool and IBD figures depend on it). Next-cut rule: a verifier benchmark with a pinned bound in the crate's CI.

PC 2 order: suite job 2 of 3 (kaspa-consensus alone, running), the proving agent's socket fix (60 s; the aggregation-cost restore is skipped), suite job 3 of 3 (the five other crates), the aggregation-cost re-run (20 min), the prover-floor agent's windows (through the proving agent). The proving app branch: a223ca9 is the last code change (8b47073 docs only after it).

21:58. ca2-cache is rebased: one squashed commit 1950661 on ca2-v3 63dabb2 (fast-forwardable; history under tag ca2-cache-history-2026-10-05); V3_CLASS = { era: None, hot: None, ..MX4 }; every hunk kept the ca2-v3 side and appended the hot code; 53 + 19 tests; the pinned v2, mx4, era and readwidth packs untouched; the eight hot packs re-exported with unchanged vectors, 96/96 and the pre-rebase fingerprints on Metal and Apple OpenCL. The node agent merges it after the mixer's verifier fix.

21:58 layers 4 and 8 decided IN on the PC rows (ca2-era 78c0ee4)

Card v2 MH/s six eras min / median / max Spread Fingerprints = Mac Latency-bound share
RTX 5090 (PC 1, CUDA, 1 warp per block) 137.2 136.18 / 136.44 / 138.01 1.3% 7/7 1.01
RX 9070 XT (PC 1, OpenCL, group 256) 18.09 18.61 / 18.93 / 19.21 3.2% 7/7 0.95
M5 Max (Metal, 5 x 2^24) 27.68 28.35 / 28.48 / 28.58 0.8% 7/7 1.06

All under the 5% rule; the CPU verifier 1.00 to 1.02 of v2 within one binary; the dataset build with the scatter store 20.7 to 21.8 ms against 21.1 linear on the M5 Max. Chip line: 512 B per hash, 120 to 128 distinct 64-B lines, the SRAM mirror is the whole dataset every hour (a windows-union census over 300 programs); the interleave buys nothing against a chip with a programmable address decoder (stated in the doc); the stride is a bijection with no cryptanalysis yet. Commits: b105a55 (implementation, history under tag ca2-era-pre-squash), c570da3, 669a27a, 78c0ee4 (the PC rows). The suite job 2 of 3 is build-20261005-215712 on PC 2 (kaspa-consensus alone).

21:58. ca2-v3 tip fbf958e: ca2-cache 1950661 fast-forwarded (no conflict), V3_CLASS = { era: None, hot: None, ..MX4 }; igneum-pow 53 + 19 tests, packfile-test 0 failures; the fork's check against the new crate running; the fork stays 79bd8e10. The node doc carries the CPU hash rate across the switch in run 2 (about 30% under v2 on the CPU interpreter, approximate, a shared Mac) and the verifier before/after slot for the mixer fix. Gate run 3 on the fixed tree follows the mixer fix.

21:59 0.3.10 is shipped and merged to master (cde561c, pushed 21:58:17Z)

The shipper's report: master cde561c (the 0.3.10 merge) + 1f0d62c (the plan); fork release-0.3.10 21d4c73c; digest 1f4b4425... on every node; six suites green on 21d4c73c on PC 2 with the same ban_is_decided flake under the parallel run (passes alone twice), recorded for the c4 agent. Open from it: PC 37ba0461 (the US laptop) stuck in its install since 21:41:16Z; Sam's Mac quit since 20:47Z; PC 2's prover dark (the root-socket cause is now named, the fix queued); C1 at 16:00Z; the /api/resume no-op (fixed on the 0.3.11 app branch).

Consequence for 0.3.11: the main tree base is now master cde561c, so the integration merge is ca2-v3 (8ea6740 plus the mixer fix) and ca2-coord into master, with the app branch a223ca9 (on 5b0d54f, which master contains). The fork base stays 21d4c73c (= release-0.3.10's tip), so the fork merge is clean by construction.

22:00 G6 job 2 of 3 green; the merge into master dry-run

build-20261005-215712 (21:57:12 to 21:59:45Z): kaspa-consensus alone 97 passed, 0 failed, 3 ignored (ban_is_decided ... ok), every stage ok. The PC 2 prover socket fix runs now (the proving agent, 60 s), then job 3 of 3 (consensus-core, igneum-exec, kaspa-pow, igneum-miner, kaspa-p2p-flows and the app tests).

Merge dry-run into master cde561c (a scratch worktree, aborted): ca2-v3 (8ea6740) conflicts in docs/bench-log.md and proto-opencl/host.c; ca2-coord conflicts in docs/bench-log.md, proto-cuda/nvrtc/packfile.h and proto-opencl/host.c (master's 0.3.10 merge brought pack-loop's packfile.h and opencl-rdna4's host.c). The bench log is append-only (keep both); packfile.h and host.c take the ca2-v3 side (it carries the pack-loop rule plus the class, era, mixer and hot fields) re-checked against master's hunks. The integration merge is the ship's first step.

22:01 the verifier regression bisected to the mixer's 0fc0ad1; the mixer PC 1 job is running; PC 2 queue

Bisection (mixer agent, measure lock, 21:59 UTC, same input 19297e99c7b9a55e, avg of 50, two rounds): readwidth e752fc7 0.607 / 0.609 ms per warp; the ca2-v3 seam 6c75dad 0.610 / 0.609; the mixer's 0fc0ad1 1.332 / 1.316; ca2-v3 HEAD 88dafbc 1.325 / 1.347. The 2.2x is in 0fc0ad1's derive_items / Cache path at m = 1 (not the era layout, not the load). Three candidate fixes building (the constant line mask back in Cache::line; an m == 1 fast path that is readwidth's loop verbatim; both); the one that restores 0.61 goes on top of ca2-v3 8ea6740 with the six mixer commits rebased, measured the same way; then the node agent merges, rebuilds and runs gate 3.

PC 1: fetch-mixer-x4-20261005 and run-mixer-x4-pc1-20261005 published 21:59:32Z (zip 152fcf93...; workers e6007918... and e4334aaf... from ca2-mixer; the five packs; 25-minute timeout). PC 2: the proving agent's socket fix (60 s) now, then suite job 3 of 3, then the prover-floor agent's toolchain check (3 min) and its 60-minute niced sp1-gpu-server build (CPU only; the shipped server panics on any card under 24 GB, sp1-gpu builder.rs 35 to 39, and allocates every prover at Setup: the 13.9 GB floor's cause), then the aggregation-cost re-run (20 min), then the prover-floor measurements.

0.3.11 ship template (release-0.3.10.md section 7): push the release branch, gh workflow run windows.yml --ref <branch>, node tools/ship-app.mjs 0.3.11 --node <fork worktree> --branch <branch> --public --activation-height N4 --deadline-note "program class v3 + proving v1" --notes "..." [--from ci] with the override object carrying every switch; gh must be on igneum-labs; the pre-push hook flips two site files (restore with git checkout -- site/).

22:03 PC 2's prover is back; suite job 3 of 3 published

socketfix-pc2-pv1 (22:01:14 to 22:02:12Z): /tmp/sp1-cuda-0.sock owned by root removed (the aggregation-cost job's run), the prover switched off and on; the next shard (block 89011 shard 0) "proven and submitted in 34 s" at 22:02:13Z and paid 0.93116546 IGN at 22:02:24Z. PC 2's prover had been dark from 21:25:24Z to 22:02 (the root-socket class; the CI check tools/ci/prover-socket-check.sh now fails any playbook without the two restore lines). Suite job 3 of 3 (consensus-core, igneum-exec, kaspa-pow, igneum-miner, kaspa-p2p-flows and the app tests; main 8ea6740, fork 79bd8e10) is packing and publishing from the ca2-v3 worktree now. After it on PC 2: the prover-floor agent's toolchain check and its 60-minute build, then the aggregation-cost re-run.

22:06 decided: mixer x8 into v3; the verifier fix found; ca2-v3 at 795472e

The mixer PC 1 job (22:00:08 to 22:04:39Z, both cards restored, the 9070 XT present): the daily 1 GiB build per pack, two dispatches, wall ms: RTX 5090 v2 25 / 23, mx4 24 / 24 and 23 / 25, mx8 23 / 23 and 23 / 23 (cache 4); RX 9070 XT v2 74 / 74, mx4 77 / 75 and 73 / 74, mx8 72 / 76 and 75 / 74 (cache 8 to 9). The build is latency-bound on every card; the x8 rule's build half passes with 13x to 40x margin; its verifier half passed at 2.79 ms per warp on the loaded core. DECIDED (delegated): x8 into v3; V3_CLASS = { era: None, hot: None, ..MX8 }; the chip row at x8 reads 0.92x with the 3x factor (the claim "under 1x with the factor", margin stated). Every mixer pack's fingerprint equals the Mac's on both cards (v2 25f96e7dce90bd4e; mx4 6f48d5a2aa0dbe5f, 73caaebb28e808fe; mx8 7c28cfb06c5c65a9, bbb183f72692f840); hash rates at the v2 rate on both (5090 136.5 to 137.4, 9070 XT 18.0 to 18.2 at every class).

The verifier regression is found: not the mask but inlining; the item loop inlined into MemhardCpu::fetch runs at 1.33 ms per unit, the same loop out of line (#[inline(never)], one instance per cache size, the line mask a constant) at 0.60 to 0.62 against readwidth's 0.60 to 0.64 in the same minute. The fix, the MX8 V3_CLASS, the pinned pack re-export and the final v2 / x4 / x8 session land as one commit on ca2-v3 795472e (the node agent merged the mixer's 16dfd1e as 4e733bb with two one-line field fixes; 53 + 4 + 19 + 7 tests, packfile 0 failures, the fork check clean). Then: the node agent merges, rebuilds, gate run 3 on the final class; the six era packs and the pinned v3 pack re-exported on it; the final bit-exactness and G2 (1,000 random hashes per card re-hashed by the CPU) job on PC 1.

22:08 G6 job 3 failed on a stale fork test (era inside the class); job 4 on the final tree

build-20261005-220351 (22:04 to 22:06:55Z): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8; kaspa-pow 13 passed, 1 failed: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache asserts the program's class equals V3_CLASS with era: None, but the merged crate carries the drawn era inside the class (left: era Some(EraParams { ... stride_mul 3969900165 ... }); right: era None). A stale fork test written before the era merge, not a behaviour fault; the node agent fixes it on the fork. Correction to the G6 note: the PC's test stage DOES run the v3 engine test, so the PC job is the evidence. Job 4 (the five crates and the app) runs on the final tree once the fork fix and the mixer's verifier fix (with V3_CLASS = MX8) are merged; the publish waits on it. PC 1: the 5090 power sweep has the go (8 min), the AMD sweep may follow on its own presence probe; the final v3 vectors and G2 job (the era agent, 1,000 hashes per card re-hashed on the Mac) is being prepared for about 22:40.

22:09. Fork ca2-v3-node 89dfcb95: the v3 engine test asserts what it meant on the era crate (LoadClass { era: None, ..class } == V3_CLASS and class.era.is_some(); the era bytes equal the seeds'; another era seed keeps the program id, draws another era class, hashes another pow, adds no day cache); Mac cargo test -p kaspa-pow --features igneum-pow 14 passed. Main-repo tip 6a705a2. Waiting on the mixer's fix commit for the merge, the rebuild, gate run 3 and suite job 4. PC 2: the prover-floor agent's 3-minute toolchain check has the go; its build waits for job 4.

22:11. PC 2: the prover-floor toolchain check done (floor-toolchain-1, 22:10:03 to 22:10:06Z: nvcc 12.8, cmake 3.28.3, gcc 13.3, clang 18, cargo 1.99.0; no go, so the rebuilt server drops native-gnark, the Groth16 wrap that compressed proofs never use; 16 cores, 30 GB WSL RAM; the live server untouched); its 90-minute build (sm_86, sm_89, sm_120 after C26) is HELD behind suite job 4 (the gate), with a 22:40 fallback: if job 4 is not published by then, the build goes first. Consequences C26 (the arch list, the card, the packaging row, the verify-segment run, "24 GB" kept until the 3060 proves) is with the prover-floor agent; C27 (publish-jobs.sh add runs the prover-socket and bash-body checks and refuses on failure) is being wired by the reviewer's sub-agent, no collision.

22:11 the verifier fix and the x8 class are committed (ca2-mixer 1ab8b21); the final code is in

Before / after, the era agent's way (one measure session, 22:07 UTC, readwidth e752fc7's binary beside the fixed one, the same v2 input, load 5.5): readwidth 0.607 / 0.610 ms per unit; the fixed binary 0.609 / 0.611 (was 1.332 / 1.316 on 0fc0ad1). On the fixed binary: x4 1.238 / 1.237 (2.0x v2, worst cold 1.40), x8 2.077 / 2.058 (3.4x, worst cold 2.15), x8 on the devnet seeds 2.058; so the x8 class verifies at 2.1 ms per warp on a loaded core, 4.8x inside the 10 ms gate. Cause and fix: the item loop inlined into MemhardCpu::fetch ran at 2.2x whatever the mask; derive_items_mask #[inline(never)], one instance per cache size (2^26 to 2^30) with the line mask a constant, restores 0.61. V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }; mx8-genesis (program id e323b9dcaf283a6f, fingerprint 7c28cfb06c5c65a9) and mx8-devnet-epoch0 through the chain path with the era inside (class mx8-erad810f22d, program id 73bcbfe8ccf988f1, unit 0 lane 0 d424577fce4a7a60, fingerprint 90f794dd556f7a3b on Metal and Apple OpenCL, 3/3 + 3/3, 96 of 96); the suite 53 + 4 + 19 + 7. chip-model-v3.md headline: x8, 0.92x with the factor, margin 8% on the factor and 9% on the budget; x4 kept as the measured candidate. Owed: the composed mx8-devnet-epoch0's fingerprint on the PCs (in the final-vectors job). Next: the node agent fast-forwards ca2-v3, re-checks, rebuilds, gate run 3; suite job 4 publishes when the tree is settled; the era agent builds the final-vectors and G2 package on 1ab8b21.

22:13. Suite job 4 published: build-20261005-221237 (main d233fa1 = ca2-v3 with the mixer fix merged, V3_CLASS = { era: None, hot: None, ..MX8 }; fork 89dfcb95 with the engine test fixed; the five crates and the app). The node agent's checks, rebuild and gate run 3 on d233fa1 follow; the era agent's final-vectors and G2 package is being built on the final class; PC 1 is the AMD sweep's until about 22:25.

22:15. The era final-class package is ready (zip igneum-ca2-era-pc1b.zip sha256 cb0e9db07e304b11fd4c0591351af46090442ea4f51d60eb86945b96bd28aba3; workers f8d19f0a... and 87647c15... from 1ab8b21 plus the era branch; seven packs on the final class: mx8-devnet-epoch0 and era-0 to era-5, program id 73bcbfe8ccf988f1; Mac 7/7 on Metal and Apple OpenCL, fingerprints equal: mx8-devnet-epoch0 a6752e037514c91a, era-0 64c0ee90bac42624, ..., era-5 43673acc89954d5e). G2 method: one serve-mode job of 1,024 nonces at target ff..ff per card for era-0 and the pinned pack, every nonce a "found g2" line, re-hashed on the Mac with igneum-pow hash-bound --count 1024 (dry run through Apple OpenCL: 1,024 of 1,024 on both packs). Its PC 1 go follows the 5090 power sweep (ahead of the AMD sweep). C24 and C27 closed on branch bash-body-check (7adb1ca, 6805125); the integration merge takes its prover-socket-check.sh over proving-v1's and puts pc1-cpu-prove.ps1 on the allow list.

22:16 the 5090 power-limit sweep (PC 1, relay #224, 22:09 to 22:15:37Z); the G2 job has the go

Cap Limit W Draw W MH/s MH/W SM MHz Memory MHz Busy
100% 575 316.2 115.42 0.365 3,051 13,801 92.9%
80% 460 316.1 115.60 0.366 3,050 13,801 92.5%
65% 400 (the floor) 310.6 114.46 0.369 3,051 13,801 90.1%
50% 400 (clamped) 302.4 109.20 0.361 3,050 13,801 87.5%

The card draws 302 to 316 W under this program whatever the cap, so a cap above 400 W never binds; the readwidth, era, hot-table and mixer numbers taken at 431 W sit on the flat part of the curve (within 1% of stock); best per watt 65% (400 W) at 0.369 MH/W, a 0.8% hash cost. The cap was restored to 431 W and read back. (The app's 0.3.9 rate of 124 to 141 MH/s in the STATUS lines against 115 here: the API's hash_now sampled every 5 s under the sweep's own load; the bench rows of 136 to 137 MH/s are device time.) "go PC 1" given to the era agent's final-vectors and G2 job at 22:16; the AMD sweep follows it on a fresh probe.

22:16 gate G6 GREEN on the final tree

build-20261005-221237 (main d233fa1, fork 89dfcb95, 185 s): every stage ok; kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test, with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8; 0 failed; with job 2's kaspa-consensus 97 alone, G6 is green. PC 2 goes to the prover-floor agent's 90-minute build (its go at 22:17), then the aggregation-cost re-run, then the prover-floor sweep. Gates: G3 green (the Mac suites on the final class: 53 + 4 + 19 + 7 crate tests, the Metal fuzz, edge, stats and determinism runs, the scratch tests), G4 green (runs 1 and 2; run 3 on the final x8 + era class pending), G6 green; G1 and G2 pending the era agent's PC 1 job (running from 22:16); G5 (the Windows and Mac workers from the same commit) is the ship's build step on the merged tree.

22:18 the spec and the public copy carry the final class

docs/spec/01-lottery-hash.md on ca2-coord: 1.8.5 (the mixer x8 form with the measured costs), 1.13.1 (the era draw: stride, interleave, windows, the devnet stand-in, the measured spread), 1.17 (the class v3 vectors), 1.5 (the cache note); earlier tonight 1.12 and 1.13.1 (epoch_len), 1.13.2 (R1 and the emulation rule), 1.13.3 (option C and the step mapping), 1.4.5 and 1.4.6 (generator 3, the class and era in the pack), and 4.3. Public copy: level 1 on the hero and the abstract ("Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public."), the limits bullet rewritten on the chip row (0.92x with the allowance, approximate; the margin on the numbers page; the next lever named), the level 3 table filled from the final numbers in counter-asic-2-public.md (the bench page section is written from it at the ship). Waiting: the era agent's PC 1 job (G1 on the final class and G2), the node agent's gate run 3; then the integration merge and the ship.

22:19. docs/analysis/proving-methods.md (branch proving-methods e7e0db7, not consensus): top recommendation re-size SP1's own GPU server (the floor is its code: the under-20 GB panic at sp1-gpu builder.rs:37, trace buffers at the maximum shard, a CUDA mempool that never releases), S_p as the dial, one server per card on rigs; the pinned ids stay; fallback and the Apple route: RISC Zero as proof-system version 2 behind the ProofSystem seam (8 GB at po2 19, 16 GB at po2 20, shipped Metal; 3 to 4 agent days); no 12 GB card has run a prover here, so a 4070 or 3060 in the loop is the first action. That branch merges into the 0.3.11 main tree as documentation (no code). evidence.md rows 17, 18 and 19 are rewritten on the measured class v3 (ba8379d).

22:20. The numbers page's Counter ASIC section is written as the bench-log entry "Counter ASIC 2.0, the numbers" (the page is built from the bench log; the litepaper links /bench#counter-asic-2-0-the-numbers); the litepaper's verifier figures moved to the v3 class (2.1 ms per warp on a loaded core, 4.8x inside the gate; the cache 512 MB from year 4). Everything public now carries the final class except the PC rows of the final-class packs and the G2 counts, which the running PC 1 job supplies.

22:20 gate G4 run 3 PASS on the final class; the node side is final

Fast-time 3-node network on ca2-v3 d233fa1 (x8 + era + the verifier fix) and fork 89dfcb95, 22:15:04 to 22:19:44Z: every check true; 182 / 122 blocks around the boundary; the v3 ids agree on all three miners; 0 rejected; one sink at 303/303/303; the switch line on 3 of 3; cache ready 179 / 191 ms. Final tips: ca2-v3 5eb2331 (docs only since d233fa1), ca2-v3-node 89dfcb95, both clean; the gate network stopped. Gates: G3 green, G4 green, G6 green; G1 and G2 on the final-class PC 1 job (running since 22:16); G5 at the ship's build step.

22:21 gate G4b added: the Mac must mine v3 (the app's --prepare-packs gap)

The node agent's owed item is an app gap, confirmed in the 0.3.11 app branch: engine.rs miner_args pushes --prepare-packs only for non-Metal workers (line 1468 on a223ca9), so the Mac app's Metal worker would get no prepare pack and under class v3 would answer need at the first v3 epoch (every Mac stops: the 18:23Z class). Closes in hand: (a) the proving agent adds the flag for every worker (packs/prepare on macOS) with a unit test on miner_args for a Metal card, on the app branch above a223ca9; (b) the node agent runs gate 4: a real Metal miner on this Mac across a v3 boundary on the fast-time network through the miner's --prepare-packs flow. The ship does not go without (a) in the app tree and (b) green; recorded as G4b in the rollout plan's gate list.

22:23. G4b (a) done: app proving-v1 e0de2ab on 5b0d54f: miner_args pushes --prepare-packs for every worker through prepare_packs_arg() (packs\prepare on Windows, packs/prepare elsewhere), the OpenCL-only --job-nonces kept, and start_miner now gives the Metal worker the app data folder as cwd (it had None, which would have broken the relative path: a second latent Mac fault closed); unit test every_worker_gets_the_prepare_directory_in_the_platform_form; 113 + 27 + 8 passed. The 0.3.11 app tree's final code commit is e0de2ab. (b), the Metal gate run, is with the node agent.

22:23. PC 2: floor-build-1 (22:17 to 22:21:12Z) FAILED at cargo exit 101 after 240 s with the error not uploaded (the playbook named the cargo log with a timestamp and could not find it; fixed: a fixed path, the error lines printed on failure); floor-build-2 (the same 90-minute shape) has the go at 22:24; the aggregation-cost re-run follows it. PC 1: the era agent's final-class job runs (from 22:16); the AMD sweep follows on a fresh probe.

22:24 gates G1 and G2 GREEN on the final class

Job run-ca2-era-pc1b-20261005 (22:16 to 22:21Z, 304 s, both cards restored, app 0.3.10, the 9070 XT present): the seven final-class packs' 2^24 fingerprints equal on the 5090, the 9070 XT and the M5 Max (mx8-devnet-epoch0 90f794dd556f7a3b; era-0 8e8e070db4eea52d, era-1 891c01b8563bb47e, era-2 e54279fed2831b5d, era-3 77e0ba8abbd0ae62, era-4 d898d8f4f2e7684b, era-5 a6927db380f7efb2); G2: 1,024 of 1,024 per card on era-0 and the pinned pack re-hashed by the Rust verifier. Rates on the final class (MH/s): 5090 135.90 to 137.70 (spread 1.3%), 9070 XT 18.59 to 19.18 (3.1%), M5 Max 27.85 to 27.98 (0.5%); the pinned pack 136.10 / 18.90 / 27.80. Spec 1.17 carries these fingerprints. Gates now: G1, G2, G3, G4, G6 green; G4b (a) done, (b) the Metal gate run pending; G5 at the ship's build. PC 1 goes to the AMD sweep on its presence probe.

22:25. Consequences C29 and D11 applied. C29: the litepaper's verifier line now reads 2.1 ms per warp for class v3 on one M5 Max core at load average 5.5 (the fixed crate; worst cold 2.15), 3.4x the v2 verifier's 0.61 ms, the gate leaving 4.8x (4.6x on the worst cold unit); the "about 1.3 ms quiet" scaling is struck: it came from the slow binary's 2.79 ms session (21:40), and the fixed binary's session (22:07) matched readwidth's quiet v2 figure within 1%, so the fixed numbers are near-quiet. D11: "and the bounty" struck from the hero and the abstract, "standing" from level 1; the copy says "a bounty follows the external review"; the bounty is named only once escrowed (funding.md rule 3; USD 50,000 not funded); the project lead decides (D11). The level 3 table and the numbers-page entry now carry the final-class rates (5090 135.90 to 137.70, 9070 XT 18.59 to 19.18, M5 Max 27.85 to 27.98).

22:25 the fourth eGPU drop; the era branch final; PC 1 to Ember

The 9070 XT is absent again at 22:24:09Z (relay probe #230: no Sonnet or USB4 router device), the fourth drop today, after serving the hot-table, era (twice) and mixer jobs between 21:31 and 22:21; the AMD clock and power sweep did not start and its rows are owed with this time and reason; the hardware-events table in the rollout plan carries the flapping for the morning. ca2-era is final at 95955c3 on ca2-v3 5eb2331 (six commits: the implementation, the Mac rows, the PC round 1 rows, the G2 hash-bound --count flag, the final-class re-export, the round 2 rows; 53 + 30 tests). "go PC 1 build" given to Ember Tune (its two fetches and the 8-minute Windows app build), then its 8-minute run on the 5090 baseline. PC 2: floor-build-2 running (to about 23:50), then the aggregation-cost re-run, then the prover-floor sweep.

22:27 the integration merge, in the node agent's hands; the ship order

A dry run of ca2-v3 into master in a scratch worktree conflicts in docs/bench-log.md (append-only, both kept) and proto-opencl/host.c, where taking ca2-v3's side would drop master's 0.3.10 hunks (the PCI topology and duplicate-platform detection, the select read-back): it must be resolved by hand keeping both. ca2-v3 also lacks readwidth's last three commits (d0018cf the OpenCL scratch local-buffer rule, e752fc7 read-width.md and the overflow fix, 30ff674 the per-watt rows). So the node agent, as ca2-v3's owner, merges readwidth 30ff674 and origin/master (cde561c, 1f0d62c) into ca2-v3 beside gate run 4, re-runs the igneum-pow suite, the packfile test, test-generic.sh on Apple OpenCL and the fork check, and sends the tip. ca2-coord (docs, spec, site, evidence, the plans) stays on its a9e002c base: its code files are untouched readwidth copies, so its merge onto that tip takes ca2-v3's code and brings only the documents (a rebase attempt replayed readwidth's own commits and was aborted). Ship order: master <- ca2-v3' <- ca2-coord <- the tooling and analysis branches (consequences, bash-body-check, amd-prove, card-lifetime, proving-methods, asic-history when it lands), the app branch proving-v1 e0de2ab (on 5b0d54f, in master), the fork ca2-v3-node 89dfcb95 (on 21d4c73c, release-0.3.10's tip). The ASIC-history agent (a202a09dcd24ba1d3) resumed at 22:52 (its clock) in ../igneum-wt-asic-history; its ranked additions are Counter ASIC 3.0's input after the publish.

22:30. C31 applied (a7be43f): the litepaper's "12 GB or more proves full shards" struck (one page, one number: 24 GB); the dataset's step schedule written as the gate 1 proposal on the site and litepaper (D4 is the project lead's); the merge rule "take proving-v1's evidence row 16 and its proving sentences" in the rollout plan. PC 1: Ember's build done (build-20261005-222558, 97 s, 6 outputs verified), its tune run on the 5090 has the go. PC 2: floor-build-3 (with the pinned Go toolchain; builds 1 and 2 failed on a missing go, the first unreported by a log-path bug) runs, 25 to 45 minutes expected.

22:32. C32: an app update clears the jobs folder (the 0.3.10 install took PC 1's AMD kit, fetched at 21:23:59Z; the 21:05 reading "cleared by fetch jobs" was wrong), so two rules enter the rollout plan (4a): the 0.3.11 update-now reaches PC 2 only after floor-build-3 closes (PC 2 last on the machine list), and every fetch-then-run pair re-fetches after an update, with a kit presence check at the top of every run playbook.

22:33 gate G4b GREEN: the Mac mines v3 end to end; a second Metal outage found and fixed

Gate 4 (22:25 to 22:31Z, a real Metal miner through the miner's --prepare-packs flow on the fast-time network, igneum-bench from ca2-v3 00c55aa): three v3 prepares with the pack dir and the class and era tokens; the worker's v3 prepared lines (252.5 ms the first: program 55.5, dataset 196.9, cache fill 0.9, build 41.1; then 51 and 46 ms with the day resident); every swap "with no pause"; 124 blocks accepted on v3, cpu re-check mismatched 0, need 0, no mismatch or refusal; chain 182 / 123 across the boundary, 0 rejected, one sink. Found and fixed before the run: the Metal worker's serveDataset built every day with the Swift version 2 construction keyed by day only, so a v3 program would have hashed over an x1 dataset and every found would have been refused by the CPU re-check, the same fleet-outage class as the app's missing flag; now a v3 prepare builds the day from the pack's memhard.metal and the store keys datasets by (day, class, era) (00c55aa). Two Mac outages caught by one gate run that the CPU-miner gates could not see; the rule for the next cut: every worker path (Metal, CUDA, OpenCL) mines across a boundary in the gate network before a class change ships. The integration merges (readwidth 30ff674, then origin/master) are in progress on ca2-v3 with the conflicts resolved by hand keeping both sides.

Gates: G1, G2, G3, G4, G4b, G6 GREEN; G5 at the ship's build step. The ship waits only on the merged tip and its checks.

22:33. H = 210,000 lands about 18:45Z on 6 October (DAA 137,041 at 22:32:54Z, 1.002 DAA/s since 15:40Z); the 16:00Z check keeps 2.75 hours. PC 2: floor-build-3 green at 22:32:29Z (240 s on the warm target; the patched sp1-gpu-server 166,768,224 bytes sha256 5568108b..., v6.8.1 c84ada1e with patch 700173fe, sm_86 sm_89 sm_120, the Go tarball's sha256 matched; the live server untouched, miners never stopped); "go PC 2 sweep" given for floor-sweep-1 (9 points, 8 to 10 min) ahead of the aggregation-cost re-run.

22:38 the merged tip is in; the 0.3.11 ship is assigned

ca2-v3 fa3c932 (code 49c7e78): readwidth 30ff674 merged at 3566afd (emit.rs two hunks, HEAD's superset; packbench.swift's footprint lines added to the hot-aware RESULT line; bench-log both; read-width.md readwidth's version), origin/master 1f0d62c merged at 49c7e78 (host.c one struct hunk, both fields kept; master's topology, duplicate-platform and select read-back code and ca2-v3's class, era, mh_word, hot and mixer code all auto-merged; packfile.h no conflict); checks all green on 49c7e78: igneum-pow 53 + 4 + 19 + 7, packfile-test 0 failures, the NVRTC emulator test PASS (17 sampled hashes equal to hash-bound), test-generic.sh on Apple OpenCL PASS, the fork's check clean. Fork ca2-v3-node 89dfcb95.

The ship is assigned to the 0.3.10 shipper (ae892a8b0f78fe31c) with every input (rollout plan sections 3, 4, 4a, 7, 8, 8a): release-0.3.11 from origin/master; merges in order ca2-v3 fa3c932, ca2-coord, the app branch proving-v1 (e0de2ab plus the prover log-line commit), bash-body-check e3bd761, consequences, proving-methods e7e0db7, asic-history if it lands; the fork 89dfcb95 under the release tree's vendor/; the override object with every switch; N4 = (DAA at publish + 14,400) rounded up to a multiple of 3,600, N5 = DAA + 14,400; the expected digest c562d70e...; the deadline note "program class v3 + proving v1"; update-now machine by machine with PC 2 last after the prover-floor sweep; the hand nodes, the seed, the digest sweep, the HiveOS package, the merge to master and the push. The node and proving agents stand by for fixes. PC 1: Ember's tune run. PC 2: floor-sweep-1 (from 22:34:11Z), then the aggregation-cost re-run.

22:41 the release tree is assembled; counter-asic-3.md written; PC 2 to the aggregation-cost re-run

Release worktree /Users/joshm/Projects/igneum-wt-ship0311, branch release-0.3.11 from origin/master b38f3de: merges in order ca2-v3 fa3c932 (clean), ca2-coord (bench-log both kept), the app branch 22c2363 (bench-log both; docs/evidence.md rows 15 and 16 from proving-v1, 17 to 22 from ca2-coord; the litepaper's schedule sentence from ca2-coord with proving-v1's 24 GB sentence appended), ca2-coord 57844e9 (the nine-field override line), bash-body-check e3bd761 (ci.yml both steps kept, prover-socket-check.sh from bash-body-check), consequences 99fd988, proving-methods e7e0db7, asic-history 9e4af7f; tip b968ee0; the fork worktree vendor/igneum-node-0311 at 89dfcb95 under the release tree. The checks (igneum-pow suite, the app suite, test-generic.sh on Apple OpenCL) run now; then the handoff to the shipper (the coordinator has told it to ship). docs/plans/counter-asic-3.md is written from the ASIC-history agent's seven ranked additions (the partial-store chip and the time-memory curve first, the random daily derivation as a reserve, the mixer cryptanalysis as a genesis gate, the detector and the issuance-triggered bounty, the FPGA lane, the reserve order, the vendor-share metric) with the decisions for the project lead.

PC 2: floor-sweep-1 (22:34 to 22:37:55Z, every proof verified by the unpatched host): the shard term is gone but a second floor binds at 12.7 GB measured (10.95 GB the server's own): at Setup the server pre-builds five recursion keys at the fixed 2^27 capacity (3.75 GB) plus the shrink and core keys, 9.7 GB before the first shard; a second patch (every trace buffer sized to its program or shard) follows in about 30 minutes; nothing is under 11.0 GB yet; the public line stays 24 GB. "go PC 2" to the aggregation-cost re-run (20 min); the floor's second build after it; the 0.3.11 update-now reaches PC 2 after both.

22:42 the release tree is handed to the shipper; the ship runs

Checks on the merge tip b968ee0 on the Mac: igneum-pow 53 + 4 + 19 + 7, the app 113 + 27 + 8, 0 failed (the OpenCL generic test wants the NVRTC emulator run first; both ran green on the same code at 49c7e78). The coordinator told the shipper directly to ship; it has the tree (its version bump 21173c4 sits on b968ee0 in /Users/joshm/Projects/igneum-wt-ship0311) and I touch nothing in it from here. Handoff message sent with the tip, the fork worktree (vendor/igneum-node-0311 at 89dfcb95), the nine-field override line, the two digests and the machine order (PC 2 last, after the aggregation-cost re-run closing about 23:02 and the prover-floor agent's build 4 and sweep 2 on the v2 patch eda49ab: every trace buffer sized to its padded need, the recursion keys 0.75 to about 0.21 GB each). The shipper reports to the coordinator and me at each step; the node and proving agents stand by.

22:43. PC 1: Ember's tune run (ember-tune-pc1-1) failed at 22:31:06Z after 47 s (the second engine exited at once; no knob touched: the 5090 at its 450 W limit, 2,850 MHz; the 9070 XT present on bus 98 at factory); a collect job reads the exit reason; Ember's re-run comes after PC 1's 0.3.11 update with its fetches republished (C32). Finding from its probe: the AMD helper's gmax range is an offset (-500 to 1000), not MHz; the probe is fixed to keep the clock knob closed on offset ranges and bound the power ladder by plimit_range. The window before PC 1's update goes to the AMD sweep (a fresh kit fetch, the two-read probe, 26 minutes if the card is present); the shipper holds PC 1's update-now until "PC 1 clear" (and PC 2's until "PC 2 clear"); the Mac and the laptop update first. PC 2: the aggregation-cost re-run (to about 23:02), then floor-build-4 and floor-sweep-2.

22:44. The coordinator stopped the AMD telemetry agent at about 23:00Z (its clock; the 9070 XT will not be back on the bus tonight); its AMD sweep rows are owed to the morning (relay/playbooks/amd-card-test.ps1, kit amd-kit.zip); it holds no slot. PC 1 is clear for the 0.3.11 update-now once Ember's collect job closes (the shipper reads the intake for no running job on ae432dc7); PC 2's update still waits on the aggregation-cost re-run and the prover-floor build 4 and sweep 2.

22:45 the shipper has the tree; the PC queues for the rollout

The shipper took over release-0.3.11 at cc72f4a (b968ee0 + the bump 21173c4 + one CI commit): three of the tree's own checks had failed and are fixed there (the identity check's "MacBook" pattern matched prose in card-lifetime and proving-methods, reworded; the C32 kit-path check flagged the three tools/proving-v1/pc2-*.ps1 scripts for a bare jobs\ literal, now Test-Path'd; the socket check flagged tools/amd-prove/pc1-cpu-prove.ps1 and its -sp sibling, now allow-listed with the reason); every check green (identity 0 of 220, bash-body 15 of 28, kit-path 14 of 14, socket, markers, workflow shell, relay 17, UI 23). The reviewer's C34 follow-up is with the shipper: packaging/mac/packaged-config.sh line 31 must carry the nine-field object before the DMG and the Windows inputs (a fresh install would otherwise start on the four-field digest). Its order: the Mac node build and the two digest readings, the PC 1 build job (node and app, Linux and Windows), the PC 2 suites from the release worktree, the seed's Linux cross-build, the workers, the app, the DMG, the inputs, the push and CI; then the manifest and the machine order.

PC 1: yours to the shipper now (no running job; Ember's run was "aborted (the app is quitting)" at 22:31:06Z, cause unexplained: no 0.3.11 action existed then; the shipper's PC 1 job output may say). PC 2: agg-cost-pc2-3 runs 22:41:15Z to about 22:59Z (PC 2's app log; the intake lags), then the shipper's suite job, then the prover-floor build 4 and sweep 2, then "PC 2 clear" for the update-now. The AMD telemetry agent is stopped (its rows owed to the morning); Ember's re-run after PC 1 shows 0.3.11 (its engine 054e041, the fetches republished).

22:48. C35: tonight's two unattributed app quits share a shape (PC 2 at 20:01:09Z, 20 s after the efficiency sweep's elevated helper was cancelled; PC 1 at 22:31:06Z, 45 s after Ember's job started a second engine beside the installed app), both killing a measurement in flight, both "quit:" lines naming no source. Ember's re-run is HELD until the cause is named: the Ember owner reads the engine's quit path for every caller (the single-instance lock, the API port bind, the helper protocol, the jobs runner's quit command, the updater) and the line before each quit in both PC logs, gives the quit line its source, and makes the second engine unable to make the installed one quit; one line goes to the 0.3.11 tree through the shipper, else the next cut. The AMD agent's last probe (22:45:01Z): the 9070 XT absent by the PnP list 14 minutes after Ember's ADLX read saw it on bus 98; the link flaps on a scale of minutes; its kit amd-kit-2 is on PC 1 for the morning's window.

22:48 the ship's first step report: N4 = N5 = 154,800; the workers built (G5)

Tree release-0.3.11 23bc2b2 (cc72f4a plus the nine-field packaged line). N4 = N5 = 154,800 (DAA 136,967 at 22:45Z at 0.965 blocks/s: the publish near 140,200, tip + 14,400 near 154,600, the first multiple of 3,600 at or above it; the 10,800 floor holds until DAA 144,000, about 00:45Z, else re-pinned). The two Windows workers from this tree: igneum-worker-cuda.exe 2b3b8c92... (1,536,512), igneum-worker-opencl.exe edc4a75d... (478,208), both with the resource block, different from 0.3.10's pair. PC 1's build job build-20261005-224654 (node 89dfcb95, app 23bc2b2, Linux and Windows). The app suite 113 + 27 + 8 and igneum-pow 53 + 4 + 19 + 7 green on the tree. In flight: the fork's Mac node and its two digest readings, the seed's Linux cross-build, the prover build and fixtures; then the inputs, the DMG, the push and CI; PC 2's suite job on "PC 2 suites go" (after the aggregation-cost close at about 22:59).

22:50. Two corrections from the reviewer's log reading, for the morning. (1) PC 2's app quit at 20:01:09Z was an update: its log shows "job update-now-0310 (update-now) starts: 0.3.10 is published: re-read the manifest and install now" at 20:01:04Z, the 49 MB download and the installer start; so a 0.3.10 manifest and an update-now job reached PC 2 at 20:01Z, ninety minutes before the fleet publish at 21:39:59Z; the shipper is asked which publish and jobs file that was and whether it was the same build (an unexplained early publish is a release-process question for the morning). (2) C35's order: "node stopped (exit Some(1))" is written by the engine's stop_node inside the quit path, so on both PCs the node's exit is a consequence of the quit, not its cause; the installed app's Quit comes only from the host's close or /api/quit; for PC 1 at 22:31:06Z the open question is whether Ember's playbook sends /api/quit to the installed app before starting its second engine (the reviewer reads the playbook; Ember's re-run stays held).

22:50. WITHDRAWN, the 22:50 entry's point (1): the update-now-0310 lines on PC 2 are at 21:49:24Z (the fleet update), not 20:01Z; the shipper's "nothing published before 21:33Z" stands, and PC 2's 20:01:09Z quit stays unexplained (a cancelled administrator prompt at 20:00:49Z, a PermissionDenied at 20:00:56Z, then a quit with no source; a next-cut item). C35 narrows to Ember's playbook: relay/playbooks/ember-tune-pc1.ps1 lines 125 to 127 POST api/quit to the URL file in $u when its budget is spent; if $u resolved to the installed app's URL file and the budget check fired at once, the playbook quit the installed app 46 s in. The Ember owner confirms before any re-run; the fix would be that the playbook never addresses the installed app's URL file (its own scratch app dir only) and never calls /api/quit on a URL it did not create.

22:52 the digests and the DMG; PC 1's app down since 22:31 (the fleet's hash rate and the ship's PC 1 step)

The shipper's second report (tree 23bc2b2): the two digests on the 0.3.11 Mac node: c562d70e... with no override file (= the node agent's pinned test), 0139ab9dc2992d449ec787d8f021974933631eb55740ab4b6ce9d5c226e72888 with the nine-field object at N4 = N5 = 154,800, the value every node must print after the publish (the node logs the class switch "active from epoch 43" and the proving v1 line); the DMG b7e81d4f... (41,592,041 bytes, the nine fields read back from the image); the seed's Linux node 63cf490d.... Waiting on PC 1's exes (build-20261005-224654) for the inputs, the pin, the push and CI.

PC 1 (Ember's finding, confirmed on the intake at 22:51): the installed app has not come back since its quit at 22:31:06Z (the last upload 22:31:08Z, no new run id, no job since), so PC 1 is not mining (the devnet short its 141 MH/s for 20 minutes), the shipper's build job cannot start there, and no update-now can reach it. The relay agent on PC 1 is alive (read-only probes ran through it at 22:45); the shipper is asked to relaunch the installed app through a relay task in the interactive session and to verify a new run id; if the relay cannot reach the user's session, PC 1 waits for the project lead in the morning and the 0.3.11 rollout goes without it (its update lands at its relaunch). the project lead is not woken. The quit's cause (C35): the senders are the tray Quit, stdin EOF in wrapper mode and POST /api/quit with the token; Ember's playbook POSTs api/quit to the URL file in $u when its budget is spent (lines 125 to 127); the Ember owner is checking what $u resolved to at 22:31; Ember's re-run stays held, and the playbook rule becomes: never read the installed app's URL file, never POST quit to a URL it did not create.

22:53. PC 1: the relaunch of the installed app went out through the relay at 22:52:31Z (item #241: igneum-app.exe started through explorer.exe so the app gets the user's own token, a one-shot scheduled task at limited run level as the fallback); the shipper reports the new run id and the first STATUS line; its build job starts when the app fetches the jobs file. PC 2 queue after the aggregation-cost job (closing about 22:59): the shipper's 0.3.11 suites, the prover-floor build 4 and sweep 2, the aggregation-cost re-run (20 min, with a plain-text parser: under 0.3.10 the app's /api/state comes back empty to PowerShell 5.1's JSON reader while the body is there, a class hit by three playbooks tonight; the app owner fixes the response shape or every playbook parses the text), then "PC 2 clear" for the update-now, then the ledger-pc2 agent's M16 / E17 / P17 job (the inline-cache kernel at 64 and 256 MiB on the 5090 against the honest kernel, the nvidia-smi line per setting, the igneum-exec suite in WSL2; about 12 minutes, after its kit lands on the dl folder in 1 to 2 hours).

22:54. C35 resolution (Ember owner): $u resolved to %LOCALAPPDATA%\igneum-tune\app\app.url, the second engine's own file (lines 34 to 36 and 125 of the playbook; line 74 deletes it before the engine starts; platform::data_root() honours IGNEUM_APP_DATA, set to the scratch root at line 93), and the budget branch never ran (its "RESULT TUNE error=budget_exceeded" line is absent); so the playbook did not quit the installed app. What the reading did find: the second engine counted --sweep as Power control and raised a UAC prompt at about 22:30:25Z (apply_power_limits at start), 40 s before the installed app's quit; closed on ember-tune b671c8b (Power control alone decides, no cap at start under --sweep, every Cmd::Quit names its source, the budget floored at 5 minutes, the playbook refuses a quit to any URL under the installed igneum\app). The remaining question is whether an unanswered elevation prompt can take the installed app's window host down (stdin EOF): the same shape as PC 2's quit at 20:01:09Z, 20 s after a cancelled administrator prompt (C16). "go PC 1 collect" given: the Application event log at 22:31Z and the installed app's log tail, once PC 1's app is back; the re-run stays held until the source is named.

22:55. C35 narrows to one common factor: both unexplained quits came 20 to 41 s after an administrator prompt beside the running installed app (PC 2 at 20:00:49Z then 20:01:09Z; PC 1 at about 22:30:25Z then 22:31:06Z); Ember's playbook is ruled out. Rule for the night (rollout plan 4a): no PC job raises an elevation prompt on either PC; the two-minute class test (one prompt raised and cancelled beside the mining app, the quit line read) waits for the morning with the project lead present, or tonight only after the relay relaunch path is proven and the rollout is done; Ember's quit-source stamping (b671c8b) goes to the next cut. The devnet has been short PC 1's 141 MH/s since 22:31Z (96 MH/s at 22:51); the relay relaunch #241 is the recovery, else PC 1 is on the morning's hands list beside the eGPU reseat and the 16:00Z check.

22:55. PC 1's engine did not exit: relay task #241 found igneum-app.exe ALIVE (pid 26696, started 21:49:40Z, session 1) answering nothing on /api/state in 60 s and uploading nothing since 22:31:08Z: it logged its quit at 22:31:06Z, stopped the miners and the node, and hung instead of exiting (a quit that never ends: a C35 fact and a next-cut defect: the quit path must end the process or the watchdog must end it after a bound). Task #243 (22:55:10Z) ends that engine by pid, as the app's own updater does, then starts the per-user install through explorer.exe (the user's token), the limited-run-level scheduled task as the fallback; the new run id, the first STATUS line and the build job's first STAGE line follow in the intake.

22:56. fud-close (the ledger closer's 45 public-text fixes, two CI checks, the relay fixes; 0 conflicts with ca2-coord) is added to the ship order after ca2-coord if its ready tip reaches the shipper before the inputs are pushed (the workers and the DMG rebuilt from the merged tip, G5); else it heads the next cut. The fork-side ledger-fixes (from release-0.3.6) is not in 0.3.11.

22:57. Correction: fud-close is NOT in 0.3.11 (the tree closed at 23bc2b2 before it reached the ship order; no late branch, the 0.3.10 rule); it heads the next cut with the fork-side ledger-fixes rebased onto the 0.3.11 fork. The push and CI go the moment PC 1's exes land.

22:57. Next-cut coupling recorded (the ledger closer): fud-close 647b08c's worker half of M28 (the kernel_sha256 check in packfile.h and host.c) must ship with the fork-side ledger-fixes miner commit 3d4ec451 that stamps the hashes, or every worker refuses every pack; ledger-fixes is not yet rebased onto 89dfcb95 (two conflicting files); round 2's ledger branches stay behind tonight.

22:59. Next-cut note from the reviewer's merge-tree against 23bc2b2 (0 conflicts for explorer d7e797c, ember-tune b671c8b, rig-install 086008a, pool-v0 425b875, ota-k2 89a76b1, hive-words 2d056e8, fud-close 647b08c and the others already in): explorer d7e797c makes tools/ci/public-api-check.mjs fail when the live /api/stats lacks proving, and that check runs on master pushes against the live site, which Vercel redeploys only after the push, so the first master CI after a ship carrying it goes red through no fault; the next cut holds d7e797c or gives the check a retry loop. Waiting now on two watchers: the aggregation-cost close on PC 2 (then the shipper's suites) and PC 1's app relaunch (then the exes, the push and CI).

22:59 PC 1: the hang explained; orphaned miners from the second engine hold both GPUs

The shipper's relay task #244 (22:55:24Z) counted 1 igneumd, 2 igneum-miner, 2 igneum-worker-cuda and 1 igneum-worker-opencl running although the installed app had stopped its miners at 22:30:20Z and its node at 22:31:06Z: Ember's second engine's children, orphaned when its job was aborted, mining on both GPUs; they would fight the relaunched app's miners and void every number. The relay lane kills the tree (taskkill /F /T on every miner, worker and non-app igneumd) and relaunches the app. The hang (Ember's reading): the installed engine's quit got stuck in the jobs runner's abort, whose reader waits for EOF on the script's stdout pipe; the pipe's write end was inherited by the second engine and its miners (PowerShell's Process.Start with redirection inherits every inheritable handle), so EOF never came and the engine sat "responding" until #243 ended it. Class rule for every playbook that starts a second engine (ember-tune-pc1.ps1, relay/playbooks/sweep-5090.ps1 and any job script of the shipper's): no inherited pipe into the second engine, its whole process tree killed at the end and on abort, the installed app's miners restarted only after; a CI check that fails a playbook starting an engine without those lines (Ember's branch). The quit's sender is still open (the tray excluded by the missing 45-s host timer kill: stdin EOF or POST /api/quit; the event-log collect decides, when the app is back).

23:00. The second-engine rule is closed on ember-tune 8ab9068 (docs/plans/ember-tune.md section 5; the playbook pair fixed; tools/ci/second-engine-check.sh in ci.yml, shown to fire on a known-bad playbook and pass the fixed pair); next cut. The shipper's relay kill-and-relaunch task on PC 1 goes out at about 23:03:30Z (or the kill alone at once if the new engine reports first); the build job follows on a clean PC 1.

23:03. C38 (the reviewer): the documents of ca2-analysis ee42d7c (sram-mirror.md, int8-matrix-family.md, the dot4 probe sources), ca2-soundness a465881 (scratch-soundness.md; its scratch tests reached ca2-v3 through the mixer branch), ca2-epoch e95e8b5 (epoch-length.md) and prover-floor (prover-floor.md) exist on their branches only: my ship list omitted them, and evidence rows 17 and 19, the litepaper's chip bullet, chip-model-v3.md and the rollout plan's layer 9 row cite them. The shipper decides before the push: a docs-plus-standalone-sources merge of the four (0 conflicts for docs/, nothing the gates ran on, the shipped code paths' diff verified empty), or, under the 0.3.10 no-late-branch rule, the citations changed to "on branch " on ca2-coord as one docs commit and the four heading the next cut beside fud-close.

23:05. C38 closed on ca2-coord: ca2-analysis ee42d7c, ca2-soundness a465881, ca2-epoch e95e8b5 and prover-floor cfe3d80 merged (8c00f84, 271cd63, f940101, ead67e0; the bench log both sides each time), the soundness branch's two code files set to the release tree's versions (0b505f9; an empty diff against 23bc2b2), so the five cited documents are on the branch the ship takes and its merge is one docs commit; the shipper decides whether at 0.3.11's master merge or the morning's cut.

23:07 CORRECTION: the relay's "PC1" is PC 2; the 9070 XT never dropped during the app jobs; PC 2 was restarted twice; PC 1 is unreachable tonight

The shipper found it from the intake: PC 2 has two new engine runs (win-1ccfe586-20261005-225528 and -230330) at the exact times of its relay tasks #243 and #245, and PC 1 none since 21:40:41Z; the relay machine named "PC1" is the 1ccfe586 box and PC 1 (ae432dc7) has no relay agent. Consequences, corrected in the rollout plan 7b: every relay probe that reported the 9070 XT "absent" read PC 2 (no 9070 XT there), so the card was present on PC 1 at every app job (21:04 to 22:31) and the eGPU "drops" are withdrawn (the reseat is off the morning list; the only real fault was the install-time Code 43); the 5090 power-limit sweep ran on PC 2's 5090 (its numbers stand, the machine corrected); PC 2 was force-restarted at 22:55:28Z and 23:03:30Z (its agg-cost-pc2-3 killed; its app back at 23:04Z, pid 30484, miners up); PC 1's app is down since 22:31:06Z and unreachable tonight: the project lead relaunches it in the morning (its 0.3.11 lands then through the manifest); the fleet runs short its 141 MH/s until then; the Windows exes for 0.3.11 come from PC 2 instead. PC 2 order: the shipper's combined suites-and-exes job (now), the prover-floor build 4 and sweep 2, the aggregation-cost re-run, "PC 2 clear" for the update-now, the ledger-pc2 M16 job. No relay task to "PC1" without my word. Next-cut item: relay clients named by machine id, and a refusal of a name two boxes could answer.

23:08. PC 2 after the restarts: agg-cost-pc2-3 died in its own-miner phases (last upload 22:51:16Z) and its finally block never ran, leaving the live prover OFF (its job switches it off at start and the app persisted it: PC 2 has not proved since 22:44Z), a stray miner beside the app's restarted 5090 miner, and a root-owned socket; "go PC 2 restore" given for tools/proving-v1/pc2-agg-cost-restore.ps1 (20 s: stray miners and workers stopped, the root server killed and the socket unlinked, the card re-enabled, the prover on); it queues behind the shipper's combined job if that is already in the file. Measured in job 3 before the restart: phase A (the app's 5090 miner at 117 MH/s mean) 7.6 to 8.0 s shards, 8.0 s unchained and 10.0 s chained aggregations, 93.8% GPU, the same as job 1. The shipper: C38's documents are in the release tree (ca2-analysis, ca2-epoch and prover-floor as docs under the built-input gate, scratch-soundness.md alone, then ca2-coord 8bef299's four docs files; tree 5debb36; identity 0 of 224, kit-path 19 of 19; G5 untouched); the Windows node exes are being cross-built on the Mac (the 0.3.9 way) as the fallback while PC 2's combined job is the preferred source (the PC's exes ship if its job lands first; the plan records which).

23:09. The ledger-pc2 M16 kit is on the dl folder (igneum-inline-bench-kit.zip, 202,138 bytes, sha256 80ce0290a456b5c221ffde02bb35f74a25e48394e4cbccf53335a83e8aeb08ce; the Mac's three bit-exact checks PASS on the real kernel texts): one run job, two passes (block-warps 1 and 8) over the honest 1 GiB kernel and the inline kernel at 256, 64 and 32 MiB, 15 s a setting, the miners paused and the prover off by the script and restored in its finally; its go after PC 2's 0.3.11 update. The morning's hands list is in the rollout plan 7c (start PC 1's app first, then the 16:00Z check).

23:09. PC 2 order, final for the night (one job at a time, each on the previous release): the shipper's combined suites-and-exes job; the aggregation-cost restore (20 s, queued behind it: the prover back on); the prover-floor build 4 (5 min) and sweep 2 (4 min); the aggregation-cost re-run (20 min); "PC 2 clear" and the 0.3.11 update-now; the ledger-pc2 M16 job (4 min); the repro benchmark (14 min: the 5090 CUDA and OpenCL, the gfx1036, the proving step, the 2^21 against 2^24 job-size bench). PC 1: nothing (its app down; the morning's first item).

23:10. C32's premise withdrawn (the reviewer, from the code and the relay finding): an app update does not clear the jobs folder (the "AMD kit gone" read was PC 2, where it was never fetched); what stands is that an app restart aborts every running job, so the update-now is sequenced after the jobs that must not die, and the kit presence check is hygiene; rollout plan 4a reworded; the bench-log line 39f02ff on the stopped AMD agent's branch is corrected by this note (its owner is stopped).

23:11. For the morning: N4 = N5 = 154,800 (epoch 43) lands at about 03:30 UTC on 6 October (DAA 139,1xx at 23:1x at 1.002 DAA/s): class v3 and proving v1 switch on while everyone sleeps, the Mac, the laptop and PC 2 cross it on 0.3.11; PC 1 (down, 0.3.10, no class v3 field in its digest) shows "0 peers" at its relaunch until it takes 0.3.11, the digest rule working as written; 7c's first line says relaunch, update, reconnect.

23:13. C39, confirmed to the shipper as the order: two manifests, never one. The first 0.3.11 manifest carries the live four-field override unchanged (a 0.3.10 node dies on a file with the new fields: deny_unknown_fields, and the engine restarts the node with a taken manifest's override whatever binary is installed); the nine-field object goes in a second manifest only when every reporting app and node is on 0.3.11 (N4 = N5 re-pinned then if the 10,800 floor has passed); the hand nodes and the seed switch files at that step; the digest sweep to 0139ab9d... after it. Rollout plan section 4.

23:14. C40, confirmed to the shipper as the rule and written into section 4: publish 1 moves the hand nodes and the seed to the 0.3.11 binaries with the four-field file FIRST (each printing c562d70e), then the manifest and the apps; otherwise the apps on c562d70e and the hand nodes on 1f4b4425 would partition for the window between the publishes. Publish 2 flips everything to 0139ab9d in one sweep.

23:14. The shipper confirms the order with one measured correction: the 0.3.11 node with the fleet's live four-field file prints 4d8f8bb668828a3dcf7b783b995f3d3ebfde32a092dd1dbd5bf4373c5c65a62c (a 22-s scratch node, 23:13:54Z); c562d70e is the no-file case (the pinned test); 0139ab9d the nine-field case. So publish 1 flips the fleet from 1f4b4425 to 4d8f8bb6 (the hand nodes and the seed first, then the manifest and the apps; the 0.3.10 apps refused for the minutes until each updates, PC 1 until the morning), publish 2 to 0139ab9d in one sweep. Step 1 starts when PC 2's job gives the exes and CI is green on the pushed tree. Rollout plan section 3 carries the three digests.

23:17 the combined PC 2 job is green: the suites on the final release tree and the Windows exes (G5 and G6 complete)

build-20261005-230745 (23:07:45 to 23:16:49Z, 469 s, every stage ok): the six node suites exit 0 in 54 s (kaspa-consensus among them, no flake this run), the app suite exit 0; igneumd.exe be8e83c07aeae5eb6842768071735289f3ef149ed9a7088592d8c54a4c252c08 (51,321,856), igneum-miner.exe 1ba1a249..., igneum-app.exe ab104cc0..., nine outputs verified. The shipper pushes the inputs and the tree and starts CI now; publish 1 follows CI (the hand nodes and the seed at 4d8f8bb6 first, then the carried-over manifest and update-now to the Mac and the laptop; PC 2 on my "PC 2 clear"). PC 2 now: the aggregation-cost restore (queued), floor-build-4 (the go at 23:17), floor-sweep-2, the aggregation-cost re-run, then "PC 2 clear".

23:18. PC 2 restored (agg-cost-restore-1, 23:16:50 to 23:17:13Z): the prover ON and proving at once (block 92435 shard 0 assigned at 23:17:03Z), the 5090 miner restarted by the app at 23:17:14Z, no stray miner left (the 23:03 restart had taken them). Side effect: the card switch re-sent the identities as the state reported them, so the 5090 runs 2 identities (vote keys) instead of 8 until set back; the aggregation-cost re-run's first step restores 8 through /api/cards and the restore script carries the identities it read (the class closed). Closed tonight: the nvidia-smi time-slice lever ("Not Supported" without elevation on driver 13.3). PC 2 now: floor-build-4, then floor-sweep-2, then the aggregation-cost re-run, then "PC 2 clear".

23:19 release-0.3.11 is on origin (3b0262f); the Windows inputs are live; CI acquired

The shipper: release-0.3.11 pushed at 3b0262f (the tree 5debb36 plus the plan, the pin and the CI fixes); the signed Windows inputs live (payload-inputs.zip ba62e866818eb30ca173b8767468988f7daaa207c2b1e4ebeae326b6e5a365d9, 72,180,607 bytes, 12 files: igneumd.exe be8e83c0..., igneum-miner.exe 1ba1a249..., the class-aware workers 2b3b8c92... and edc4a75d..., the NVRTC DLLs; node 89dfcb95); the Windows CI run 37387737179 dispatched 23:18:40Z and acquired at once (GitHub operational again). On its green: the fetch and the DMG copy; step 1a (the observer, node 1 and the seed on the 0.3.11 binaries with the four-field file, 4d8f8bb6); 1b (the manifest with the override carried over, --public with the HiveOS package c606a170...); 1c (update-now to the Mac and the laptop; PC 2 on "PC 2 clear"); then publish 2 with the DAA read sent to me first.

23:20. floor-build-4 green (23:20:14Z, 120 s warm, 6 crates): the v2 server (patch 08ce0555: every trace buffer sized to its padded need) /opt/igneum-floor/bin/sp1-gpu-server 166,748,808 bytes sha256 fb3165d809d2031cc05312c79a545065b2d22ced0a2cd134fb40462262449edf, v6.8.1, sm_86 sm_89 sm_120, the live server untouched, miners never stopped. "go PC 2 sweep" given at 23:20 for floor-sweep-2 (9 points, 4 to 6 minutes); then the aggregation-cost re-run, then "PC 2 clear".

23:32 publish 1 is live: the hand nodes and the seed on 0.3.11 at 4d8f8bb6, the manifest carried over, the Mac updated

The shipper's step 1 report. 1a: the observer restarted 23:23:54Z (pid 61754), node 1 at 23:24:06Z (pid 61864), the seed at 23:24:25Z, all three igneumd/2.1.0-89dfcb95 with the four-field file, each printing digest 4d8f8bb6 (the fleet's live digest, so no split). 1b: the 0.3.11 manifest live at about 23:25Z with the consensus override carried over unchanged (mac b7e81d4f..., windows 84a21443..., HiveOS package c606a170...). 1c: update-now-0311-d937c69d-37ba0461 published 23:28:06Z; the Mac on 0.3.11 since 23:29:05Z (attached to node 1 at 4d8f8bb6), the laptop's update in flight (job done exit 0 at 23:28:48Z). CI 37388453875 green. Publish 2 waits on the DAA read, which the shipper sends to me first.

The consequence for PC 2 (the shipper's fact, 23:31): since the hand nodes moved to 89dfcb95 at 23:24Z, PC 2's 0.3.10 node is refused by every peer (0 peers, 0.0 MH/s on its card from 23:29:27Z) until its update-now, and PC 1 is already down, so the devnet runs on the Mac and the laptop alone until "PC 2 clear". Decision: PC 2 updates the minute floor-sweep-2 closes (it stops the miners anyway, so no hash rate is lost by letting it finish, and an update-now restart would abort it mid-sweep and need another restore); the aggregation-cost re-run moves to after PC 2 is on 0.3.11 (its parser already handles 0.3.11's CardState, the identities step first), then the M16 job, then the repro job. floor-sweep-2 at 23:31:18Z: started 23:21:17Z, the prover off, the v2 server fb3165d8 confirmed, idle 2,060 MiB, no point rows in the 23:31 progress upload yet (cap 30 min).

23:40. "PC 2 clear" given to the shipper. floor-sweep-2 never wrote a point row (three identical 1,139-byte progress uploads at 23:26, 23:31 and 23:36Z: start, gpus, the v2 server fb3165d8, the live server untouched, idle 2,060 MiB, nothing after); the prover-floor agent's reading: rows are written per point and sweep 1's points took 11 to 18 s, so 18 minutes without one means the first proof never returned (hung, not slow). So nothing was in flight to protect and PC 2 updates now; the update's restart kills the job tree, so its finally block (prover on, the patched server's process ended) never runs: the prover-floor agent's restore-and-diagnose job is PC 2's first job after the update, then the aggregation-cost re-run (its parser takes 0.3.11's CardState; the identities step first), then M16, then the repro job. The 12 GB prover-floor rows stay unmeasured until that diagnosis says why the v2 server hangs on point 1; sweep 3 only if it is under 10 minutes tonight. Also the consequences reviewer's C42: the new side had no miner from 23:24Z (the Mac app came back paused from a persisted Pause; the observer read 0 miners), so the new side's DAA clock stood still until the laptop's and the Mac's miners joined; N4 = N5 and H slip by the stall's length, and PC 2's old-side chain wins the reorg when it joins (allowed by the C4 rule). The shipper carries the fix (resume the Mac's miner; "the new side has at least one miner" as a step-1 check next time).

23:47. PC 2's update-now (update-now-0311-1ccfe586, published 23:39:44Z) queues behind the hung sweep: the app fetched it at 23:40:13Z ("1 new for this machine, 1 queued") and runs jobs one after another, so it waits for floor-sweep-2's 30-minute cap (about 23:51:17Z) rather than killing it; PC 2 reads "0.00 MH/s, waiting | 0 peers, syncing" meanwhile (refused by the new side, its miners stopped by the sweep). Decision: let the cap expire, no relay task (the two relay tasks tonight that touched that box both force-restarted it). The new side mines on the Mac at 21 to 23 MH/s; the laptop's 0.3.11 install is 15 minutes silent (its 0.3.10 install took 17). Then: PC 2's 0.3.11 STATUS line, "go PC 2" to floor-restore-1 (40 s, the diagnosis of the hang), the aggregation-cost re-run, M16, the repro job. Publish 2's DAA reading after PC 2 and the laptop show 0.3.11; the floor holds until DAA 144,000 (about 01:00Z after the stall). Next-cut item: an update-now pre-empts a running job, or the STATUS line reports the queue wait. Rollout plan updated (7e63cf5): 7d the step-1 stall (11 minutes, N4 about 03:40Z, H about 18:56Z, the step-1 miner check), the prover memory tiers (16 GB open, 12 GB waiting on patch v2), C41 HiveOS on the next-cut list and the tier line.

23:57 PC 2 on 0.3.11; publish 2 has the go; N4 = N5 = 154,800 stand

PC 2: the queued update-now ran 23:51:19Z (two seconds after the sweep's cap), the installer verified 23:51:22Z, the quit 23:51:24Z, the new engine 23:51:30Z (run win-1ccfe586-20261005-235130), igneumd 23:51:32Z, the 5090 miner 23:51:43Z with a clean first pack (self-test PASS); its node joined the new side at once (1 peer, DAA tracking the observer's). The hourly boundary at DAA 140,401 (23:52:21Z) fell 40 s after the worker's first pack: the worker refused the new epoch's jobs for 80 s until the app's attempt-aware path prepared it (attempt 0, the worker switched 23:53:42Z), then 113.4 MH/s wall at 23:54:14Z; the Mac's worker crossed the same boundary in 241 ms. The laptop is silent since its update-now at 23:28:28Z (its 0.3.10 install took 55 minutes of silence); publish 2 does not wait on it.

Step 2a reading (the shipper, 23:55:42Z): tip DAA 140,706, floor 14,094 against the 10,800 minimum, so N4 = N5 = 154,800 stand, the nine-field object is unchanged and the target digest stays 0139ab9d...; a re-pin would only be needed past DAA 144,000 (about 00:55Z). Publish 2 given the go at 23:57: the hand nodes' and the seed's files switched and restarted, the nine-field manifest, update-now to the Mac and the laptop; PC 2's second restart on "PC 2 clear" after floor-restore-1 (its go at 23:56, about 40 s: the hang's diagnosis, the sweep's leftovers killed inside WSL2, the prover back on). Then on PC 2 once on the final state: the aggregation-cost re-run, M16, the repro job, the ledger-fixes-0311 suites (window about 00:50 to 01:10Z).

23:58. floor-restore-1 closed (23:57:06 to 23:57:29Z, exit 0): the hang diagnosed from the point's own log. The v2 server (patch 08ce0555, every trace buffer sized to its padded need) panicked inside the proof, not at Setup: thread 'tokio-rt-worker' panicked at sp1-gpu/crates/jagged_tracegen/src/lib.rs:240:70: range end index 37,428,736 out of range for slice of length 36,700,160 while laying out a recursion key, after four tracegen allocations had run (device 6,703 to 9,033 MiB); the client then waited on the socket forever (the sampler ran 1,740 s), which is the hang. Peak device memory on the hung point 9,966 MiB. The arithmetic names the bug (the prover-floor agent): 36,700,160 is the key's buffer as v2 sized it (the padded preprocessed traces, 35,651,584, plus one stacking height of slack) and 37,428,736 is that end plus one main trace of 1,777,152 elements, so a prove path appends the shard's main traces into the key's buffer, which upstream sized for a whole shard and v2 sized for the preprocessed phase alone. Before the panic v2 was doing what it should: 6,567 MiB after Setup against 9,703 on v1 (a 3.1 GB cut), the core shard buffers 0.02 to 0.42 GB against 0.76. Fix (v3): keys that feed that path keep the main allowance; the build is 2 minutes warm and the sweep 4; it runs on PC 2 right after the second restart, before the aggregation-cost re-run (two gos: build, then sweep). PC 2 after the restore: no leftover processes, the live server c2642ad1 untouched, the prover on (23:57:08Z), the 5090 at 96 percent on the miner. "PC 2 clear" given for the second restart at 23:58; the aggregation-cost re-run on PC 2's STATUS line at 0139ab9d.

00:09 publish 2 is on the nodes: the hand nodes, the seed and PC 2 at 0139ab9d; the nine-field manifest live

The shipper: the observer, node 1 and the seed restarted with the nine-field override at 23:56:59Z, 23:57:12Z and 23:57:30Z, each at digest 0139ab9d; the live manifest carries the nine-field object since 23:57:49Z. PC 2: the switch job update-now-0311-switch-1ccfe586 ran 00:01:10Z, its node restarted with the nine-field file 00:01:42Z, STATUS at 00:06:31Z: 106.48 MH/s, node 141,351 blocks, 1 peer, synced, 830 accepted this run, 0 faults. The expected gap of the two-publish order: PC 2 (still on 4d8f8bb6) refused the seed between 23:57:42Z and 23:58:12Z, closed by the switch. One finding for the next cut (the miner): the Mac's miner lost node 1 at its 23:57:12Z restart and never reconnected (5,300 "Not connected to server" submit errors, templates frozen at 2,350, hash burned at 26 MH/s); the shipper published restart-miners-0311-d937c69d at 00:08:01Z for the Mac only; the gRPC subscription does not resubscribe after the node restarts. Still open on the ship: the laptop's return (silent since 23:28:28Z), the digest sweep, the merge to master and the push, the shipper's report. PC 2 now: floor-build-5 (the go at 00:09; patch v3 3604c25: the key's buffer stays preprocessed-sized at Setup and main_tracegen grows it on first use; a short bound aborts loudly instead of leaving the client on the socket), then floor-sweep-3 (9 points, 4 to 6 min), then the aggregation-cost re-run, M16, the repro job, the ledger-fixes-0311 suites (fork tip fbb0082a, published with the ledger-rebase worktree's tools/build-job.mjs; the Mac run of the same suites is in flight under the build lock).

00:18 0.3.11 SHIPPED: master 30cd292 then 630da6b pushed; every live node at 0139ab9d

The shipper's close: release-0.3.11 (7ab72f3) merged onto master 063baf6 without conflict as 30cd292, then 630da6b (the plan's merge section), pushed 00:13:04Z and 00:15Z; CI on master in progress (37392831312 ci, 37392831184 windows-ci). The live site deployed from 30cd292 at about 00:13:45Z with every corrected sentence (24 GB proves, the schedule as the gate 1 proposal, no bounty words, no 12 GB clause). The digest sweep closed for every live node: the observer 23:56:59Z, node 1 23:57:12Z, the seed 23:57:30Z, PC 2 00:01:42Z, the Mac through node 1; all 0139ab9dc2992d449ec787d8f021974933631eb55740ab4b6ce9d5c226e72888; tip DAA 141,715 at 00:12:36Z; PC 2 105 to 116 MH/s; the Mac 25.7 MH/s with 55 accepted since its miner restart at 00:08:33Z. Pending their relaunch by hand: PC 1 (down since 22:31:06Z, on 0.3.10), the laptop 37ba0461 (silent since 23:28:28Z, off the console), Sam's Mac (0.3.9, quit 20:47Z); each takes 0.3.11 and the nine-field object through the manifest, with one possible old-node death inside the update (C39's shape) to read in the morning intake. No release tag was cut (0.3.9 and 0.3.10 have none; a convention needs the project lead's word). The shipper's plan: docs/plans/release-0.3.11.md sections 8 (step 2 times), 9 (the sweep), 11 (open items: the miner's dead gRPC channel, the HiveOS override, the prepared pack near the hour, PC 2's first-shard socket error at 23:52:14Z, the relay naming, the job queue, the edge index, C1 at 16:00Z), 12 (the merge). The activation: program_class_v3 and proving v1 at N4 = N5 = 154,800 (epoch 43), about 03:40Z on 6 October at the new side's rate; H = 210,000 about 18:56Z. The shipper is done with PC 2 for the night. Still running on PC 2 under my schedule: floor-sweep-3 (the 12 GB tier under patch v3), then the aggregation-cost re-run, M16, the repro job, the ledger-fixes-0311 suites.

00:20 floor-sweep-3 green: nine points proved and verified on the v3 server; the aggregation-cost re-run has its go

floor-sweep-3 (00:13:01 to 00:16:57Z, 237 s, exit 0; the v3 server b37defef, the live server c2642ad1 untouched, the prover off then back on, idle 2,089 MiB with the miners stopped). Every point VERIFIED by the unpatched host (proof 1,272,897 bytes, verify 37 to 40 ms). Peak device MiB (idle included), prove seconds:

cfg fixture cycles peak MiB prove s
budget 12 GB block-83616 (empty) 280,706 9,939 2.6
budget 12 GB block-56-transfers 556,369 10,003 3.5
budget 12 GB fees-v1-shards2 (v1, the adopted shard) 4,717,439 12,915 4.3
budget 12 GB block-338-shard1 (prototype) 60,415,376 13,459 16.8
budget 12 GB, recursion alloc 2^26.6 fees-v1-shards2 4,717,439 12,883 4.3
budget 16 GB fees-v1-shards2 4,717,439 16,115 4.0
budget 32 GB fees-v1-shards2 4,717,439 16,851 4.0
element threshold 2^26 block-83616 280,706 9,971 3.3
element threshold 2^26 fees-v1-shards2 4,717,439 10,291 5.7

The key-buffer growth fired as designed ("grow key buffer 36,700,160 -> 91,226,112 elements"), the panic class of sweep 2 is closed. Reading, pending the prover-floor agent's write-up: under a 12 GB budget the prototype shard peaks at 13,459 MiB and the adopted v1 shard at 12,915 MiB with 2,089 MiB of idle inside, against about 12,288 MiB reported by a 12 GB card; the 2^26 element threshold brings the v1 shard to 10,291 MiB at 5.7 s against 4.3 (a third slower), which is the first row under a 12 GB card's size; the tier call is the agent's, in docs/analysis/proving-methods.md and the bench log. PC 2 released at 00:17Z; "go PC 2" to agg-cost-pc2-4 at 00:19 (about 20 min; the identities set back to 8 first), then M16, the repro job, the ledger-fixes-0311 suites.

00:24 agg-cost-pc2-4 measured nothing (the parser against 0.3.11's state); sweep 4 running; a fork test finding

agg-cost-pc2-4 (00:18:26 to 00:22:05Z, exit 0): card_off and card_on both read "no nvidia card in the state" against 0.3.11's /api/state, so the app's CUDA worker kept running, the pause left 3 miner/worker processes after 120 s, and every phase (D, E20, E18, E16, E0, H) was voided by the job's own double-mining guard; the miner resumed and the prover switched on at 00:22:05Z (the 5090 at 95 percent). Fix assigned to its agent: the parser for 0.3.11's state shape and one refusal before the pause; job 5 takes the slot after sweep 4, M16 and the repro job (about 00:50Z). floor-sweep-4 (mine-and-prove: 2^26 on the v1 and empty shards, 2^25 and 2^27 on the v1 shard, the 5090 mining at full rate, 3 minutes) has its go at 00:23. The prover-floor write-ups are in (prover-floor 251154d: the bench-log entry with the nine rows and the idle-subtracted figures, docs/analysis/prover-floor.md with the model, the hang diagnosis, the patch's three steps and the tier table; the arithmetic for mine-and-prove on 12 GB, 8.2 + 1.8 = 10.0 GB of 12.3 before the display, so the "under 9.0 GB" row is not met alone and sweep 4 decides 2^25); proving-methods 0a00c08 (RISC Zero stays the Apple route only); proving-v1 docs 0a1ad9d (the tier table).

A finding on the shipped fork from the ledger closer (00:21): kaspa-consensus test processes::pruning_proof::igneum_m20_tests::witnesses_are_checked_in_epoch_order_under_their_own_seeds fails on its Mac run of the ledger-fixes-0311 suites (fork fbb0082a on 89dfcb95; 99 passed, 1 failed, 4 ignored): the pruning-proof stand-in (igneum_pow.rs 114) fills era 0 with the genesis hash and the test expects EpochSeeds::v2's era ZERO_HASH (consensus/pow/src/igneum.rs 88). The code read: the node's RPC always reports Some(genesis) in era 0 (consensus/src/consensus/mod.rs 838 to 869) and the miner's legacy walk gives genesis (main.rs 504 to 536), so the proof code matches the node and the test's expectation is the odd one out; seeds_from_info's unwrap_or(ZERO_HASH) differs only against a pre-field node (a mixed case that ends at the three relaunches), and a v2 program reads neither field, so no hash changes. PC 2's combined suite job on the release tree reported kaspa-consensus exit 0, which disagrees with the Mac run; the single test is running on 89dfcb95 on the Mac under the build lock to settle it. Either way the next cut carries the test fix (the expectation, or v2() taking the genesis as its era) and the miner's unwrap_or aligned to genesis.

00:28 the 0.3.12 list, in order; the /api/state defect (decided by main: next cut, no re-ship tonight)

  1. App proving-v1 6714a45: /api/state answers {} on every proving machine about 15 paid shards after an app start (paid_wei is a u128; serde_json refuses it above u64::MAX = 18.45 IGN; a paid shard averages 1.23 IGN; engine.rs 180 swallowed the error into an empty object). Consequence per tier: the dashboard on any proving 24 or 32 GB machine goes blank within about 12 minutes of its first payout and stays blank until a restart; mining, proving and payouts unaffected; every PC playbook that read /api/state failed the same way (agg-cost-pc2-4 tonight). Fix: paid_wei as a decimal string, the error logged once, a reply carrying error and version; a unit test; 114 app tests green; 2 files, 37 lines, no UI change. Workaround until 0.3.12: restart the app (12 more minutes of dashboard). Rule for every PC 2 playbook tonight (main, 00:27): card keys from the app's settings.json, never from /api/state (agg-cost-pc2-5 is built that way).
  2. The miner resubscribes after a node restart, and the watchdog's "accepted or a template change in 120 s" rule (C43).
  3. Per-day dataset reuse in the CUDA and OpenCL workers (the integrated tier's restart per epoch).
  4. The M20 test expectation and the miner's era unwrap_or aligned to the genesis stand-in (pending the feature-gated run on 89dfcb95).
  5. HiveOS: the override file in local mode (C41), rigs-mine-only README, IDENTITIES=auto, with the M28 worker-miner coupling.
  6. The update-now job pre-empts a running job, or the STATUS line reports the queue wait; the relay clients named by machine id.
  7. fud-close 647b08c (with ledger-fixes 3d4ec451 rebased onto the 0.3.11 fork), ota-k2, ember-tune, rig-install's two follow-ups, the explorer split (3e01212 safe, d7e797c waits), the step-1 "every new-side node has a miner" check in the publish runbook. No further rollout tonight beyond the planned object sweep (done: every live node at 0139ab9d).

00:29. The M20 test finding is CONFIRMED on the shipped fork: on 89dfcb95 itself, cargo test -p kaspa-consensus --features igneum-pow --lib witnesses_are_checked_in_epoch_order_under_their_own_seeds on the Mac under the build lock fails at igneum_m20_tests.rs:122 (left era = the test's genesis 0x5e51 from the stand-in, right era = ZERO_HASH from EpochSeeds::v2; 20.2 s). Without the feature the test is not compiled ("102 filtered out"): the whole igneum_m20_tests module is behind #[cfg(feature = "igneum-pow")], and tools/build-job.mjs passes no feature flag, so PC 2's combined suite job (G6, build-20261005-230745, "six node suites exit 0") never ran the lottery-hash consensus tests. Two items, both 0.3.12 (item 4 of the list above, now split): (a) the test's expectation (the era-0 stand-in is the genesis hash on every path that has it: the node's RPC, the miner's legacy walk, the pruning proof; the class fix is one era0_seed(genesis) helper used by all three and by seeds_from_info's fallback instead of ZERO_HASH, with v2 seeds compared on (epoch, day, class) where the era is unread); (b) the gate: build-job.mjs runs the node suites with --features igneum-pow so the gated tests count, and G6's evidence names the feature set. What it does NOT mean: no consensus or hash change (a v2 program reads neither class nor era; the proof-path stand-in and the node agree), nothing on the devnet is affected; it is a test and a gate gap.

00:36 the crossing watcher is armed (C45); floor-sweep-4 live after a refused publish

The crossing into class v3 at N4 = 154,800 (epoch 43) lands at about 03:50Z (DAA 143,104 at 00:35:28Z, the observer, about 1.0 DAA per second) with nobody awake. Armed at 00:35:28Z on the Mac, detached (nohup, pid 49195): infra/devnet/crossing-watch.mjs (ca2-coord), read-only against the observer's JSON-RPC (ws://127.0.0.1:28640, getBlockDagInfo and getBlockTemplate every 15 s) and the fleet's job lines every 5 minutes, writing to /tmp/igneum-devnet/crossing-154800.out: a RESULT line per epoch change (class, generator, era, seed, activation, boundary, DAA, blocks), the crossing line at the first epoch at or above N4, the chain's blocks per minute on each side, and a verdict (FAIL on a non-v3 first epoch, no block for 10 minutes after the crossing, the observer not answering, or the after-rate under a quarter of the before-rate; PASS otherwise 90 minutes after the crossing; deadline 08:00Z). Dry run at 00:34:32Z read epoch 39, class v2, generator 2, era 0, activation 154,800, boundary 144,000, DAA 143,065. It touches no node, miner or job. The morning reads that file first (grep RESULT /tmp/igneum-devnet/crossing-154800.out), PC 1's relaunch second. What it cannot see: a miner's own "pack refused" or "need" lines (those are in each machine's intake; the chain's rate after the crossing is the proxy), and a fork between the hand nodes (the observer's sink only).

floor-sweep-4: the agent's first publish at 00:22:49Z was refused by publish-jobs.sh's kit-path check (added 21:49Z; the measurement template reached the chain job's kit without a Test-Path) and its command showed only the last line, so it waited on a job that did not exist until my repeated go; fixed in the template (the kit tested first; an update that wiped it gives measure_failed rows instead of a killed run), republished 00:33:32Z, 3 to 4 minutes. Next on PC 2 after its close: the ledger M16 job, the repro job, agg-cost-pc2-5, the ledger-fixes-0311 suites.

00:40 floor-sweep-4 green: mine-and-prove beside the 5090 miner; the M16 job has its go

floor-sweep-4 (00:34:19 to 00:37:49Z, 210 s, exit 0): the v3 server b37defef beside the 0.3.11 miner at full rate (the card at 95 percent, 338 W, idle-with-miner 3,833 MiB inside every peak), the prover off then back on, the live server untouched, every proof VERIFIED by the unpatched host (1,272,897 bytes, verify 37 to 40 ms):

cfg fixture cycles peak MiB (miner inside) prove s alone (sweep 3)
element threshold 2^26 fees-v1-shards2 (the adopted v1 shard) 4,717,439 12,066 24.4 10,291 MiB, 5.7 s
element threshold 2^25 fees-v1-shards2 4,717,439 12,066 43.5
element threshold 2^26 block-83616 (empty) 280,706 11,586 13.0 9,971 MiB, 3.3 s
budget 12 GB (2^27) fees-v1-shards2 4,717,439 14,786 17.4 12,915 MiB, 4.3 s

Reading (the tier call is the prover-floor agent's write-up): the prover's own share beside the miner is 12,066 minus 3,833 = about 8.2 GB on the v1 shard at 2^26, the same as alone, so a 12 GB card fits mine-and-prove on paper (8.2 GB plus a miner's 1 GiB dataset and working set, about 10.0 GB of 12.3 before the display); 2^25 buys no memory (12,066 either way) and costs 1.8x the time. The cost is time, not memory: 24.4 s per shard beside the miner against 5.7 s alone (4.3x; the miner holds the card), 13.0 s on the empty shard against 3.3 s; that is inside proving v1's 600 DAA-second deadline by 25x, so a 12 GB card that mines and proves at once meets the deadline and simply takes fewer shards per hour (the per-prover throughput share, not a tier exclusion). The prover-floor agent's tier reading, on the 5090's allocation: a 12 GB card (12,288 MiB) mining and proving holds the server's 8.2 GB plus the miner's 1.7 GB plus its display (0.5 to 1 GB, not measured) = 10.4 to 10.9 GB, so it fits with 1.4 to 1.9 GB spare at 2^26 and 24 s a shard; the floor is now the Setup keys and the recursion stage (4.4 GB own after Setup plus about 3.8 GB at the recursion peak), so 2^25 is not a lever; a 16 GB card mines and proves at 2^27 (10.95 GB own plus 1.7) with 2.9 GB spare at 17 s. My earlier sentence here ("proves slower than the chain makes segments") was wrong against the deadline rule and is withdrawn. Not measured: any real 12 GB card (the on-order card runs the same two points). PC 2 released 00:37:49Z; "go PC 2" to m16-inline-pc2-1 at 00:39 (about 8 min), then the repro pair (fetch-repro-pc2-20261006, run-repro-pc2-20261006, about 14 min), then agg-cost-pc2-5 (about 12 min), then the ledger-fixes-0311 suites.

00:45 M16 inline bench on the 5090 (ledger-pc2 564acab); the repro run is on PC 2

m16-inline-pc2-1 (00:38:54 to 00:43:17Z, 263 s, exit 0; app 0.3.11; the miners resumed and the prover back on at 00:43:06Z). The recompute attacker on the same silicon, version-2 program, 15 s windows, power the median of 7 nvidia-smi samples:

Setting (1 warp per block) MH/s of honest W
honest, 1 GiB dataset 132.20 1.00 326.6 (3,060 MHz)
inline, 256 MiB cache in VRAM 11.26 0.085 415.7
inline, 64 MiB (inside the L2: the SRAM-class emulation) 33.88 0.256 431.0 (the power limit, 2,835 MHz)
inline, 32 MiB 33.87 0.256 431.0

8 warps per block: 131.15 / 10.89 / 29.32 / 29.46. All three bit-exact checks passed on the card. Reading: a recompute attacker with the cache in SRAM-class memory reaches a quarter of the honest rate on the same silicon and is 5.1x worse per joule; the 1,024 dependent cache-line reads per hash bound it, not the integer budget (about 6 T op/s reached of the card's 50 T, approximate). This is the GPU-side measurement behind the chip model's recompute row (the on-die-cache chip in docs/analysis/chip-model-v3.md is the ASIC-side projection; both say the mixer's dependent reads, not the arithmetic, set the attacker's rate). Bench-log and ledger M16 and E17 updated on ledger-pc2 564acab. PC 2: the repro kit fetched 00:44:29Z (40,887,373 bytes, sha256 ok), run-repro-pc2-20261006 running (about 14 min), then agg-cost-pc2-5, then the ledger-fixes-0311 suites.

00:59 the chain past DAA 144,000: N4 = N5 = 154,800 is final; the repro run in its tenth minute

The crossing watcher (pid 49195) at 00:55:29Z: DAA 144,209, epoch 40, class v2, 57.9 blocks per minute on the observer (PC 2 at 105 to 116 MH/s and the Mac at about 26 MH/s on the new side). DAA 144,000 is behind the chain, so the 10,800 floor can no longer be re-pinned and N4 = N5 = 154,800 stands as published; at 57.9 blocks per minute the crossing lands at about 03:58Z (10,591 DAA to go at 00:55Z). Machines: PC 2 and the Mac on 0.3.11 at 0139ab9d; the laptop silent since its 23:28:48Z update-now (its install), PC 1 down since 22:31:06Z, Sam's Mac quit at 20:47Z; each takes 0.3.11 and the nine fields at relaunch. PC 2 queue: run-repro-pc2-20261006 running since 00:49:30Z (about 14 min), then agg-cost-pc2-5 (12 min), then the ledger-fixes-0311 suites (about 8 min; I publish it from the ledger-rebase worktree). The final report to main follows the suites' close, about 01:40Z.

01:10 the repro run closed; its 5090 rows read as beside the live miner; agg-cost-pc2-5 has its go

run-repro-pc2-20261006 (00:44:29Z fetch, the run 00:49:30 to 01:09:38Z, 1,509 s, exit 0, "all cards restored; done"). The package's bit-exact checks PASS on every backend (fingerprint 25f96e7dce90bd4e at 1 GiB on CUDA, NVIDIA OpenCL and gfx1036 OpenCL; program bcc1248b10cc90f2, 128 loads). The rates: RTX 5090 CUDA 62.412 MH/s (120 s, batch 2^24), 5090 OpenCL 62.283 (30 s), gfx1036 3.312 (the integrated RDNA 2, 1 CU); the memprobes 17.5 G dependent reads per second at 1 GiB on CUDA (chase 1,766 ns), 339 GB/s stream, 13.4 T ALU op/s. The 5090 rows are HALF this card's rate in every other run tonight (M16's honest 132.20 at 00:40Z; the app 105 to 116 wall) and the package's proving step took 33.9 s compressed and 20.2 s core on block-338-shard1 against 16.8 s in sweep 3, so the run most likely shared the card with the app's live CUDA worker: its card-off step reads nothing from /api/state (the "{}" defect) and waits 30 s without confirming a stopped worker. Reading passed to the repro agent to confirm from the job's own worker-count lines; if so, the PC 2 5090 rows are "beside the live miner" rows and the morning re-runs the 5090 block alone (gfx1036 and the memprobes stand). The script's close failed twice (ConvertTo-Json out of memory on the sample set; Set-Content with an empty path for the markdown), so no result file was written on PC 2; every row is in the intake. Package fixes, the class: the card-off step confirms by nvidia-smi's compute-apps list, the writer drops the samples from the JSON. "go PC 2" to agg-cost-pc2-5 at 01:10 (about 12 min), then the ledger-fixes-0311 suites.

01:13. Confirmed by the repro agent: the run's 5090 rows (62.4 CUDA, 62.3 OpenCL, 20.2 s core, 33.9 s compressed, 60.0 at 2^21 against 63.2 at 2^24) are beside-the-live-miner rows (the card-off POST was accepted but nothing confirmed a stopped worker; nvidia-smi showed 10,176 MiB used and 347 W after the run); the morning re-runs the 5090 block alone with the card-off confirmed by nvidia-smi's compute-apps list. And a withdrawal: the script read the prover's state from /api/state, got "{}", and did NOT switch the prover back on, so PC 2 has not proved since 00:49Z; the restore job run-prover-on-pc2-20261006 (POST api/prove on, about 10 s, no card) is published and runs on the app's queue; the time without the prover is the repro run plus the queue, about 25 to 35 minutes of PC 2's shard payouts. The close failures were a PowerShell case-insensitivity collision ($Cpu the result table over $cpu its own name string, so the table contained itself and ConvertTo-Json ran out of memory; $md the lines over $Md the path, so Set-Content got an empty path): both renamed; the class goes on the hygiene item for every playbook owner (no two variables that differ by case). agg-cost-pc2-5 failed at parse (01:10:44Z, 1 s, $RestoreIdentities: in a string at line 113); the fixed agg-cost-pc2-6 has the go (01:13), then the ledger suites.

01:34 the PC 2 chain is closed: job 6's curve, the ledger suites on PC 2, the job-runner finding; the coordinator stops here

agg-cost-pc2-6 (01:12:09 to 01:24:14Z, 725 s, exit 0; the parse fix held; the job's own miner alone on the 5090, the finally block put the card and the prover back at 01:24:09Z). The curve, the same four live blocks 96556 to 96559, one empty shard each: batch-log2 22: shards 7.8 to 8.1 s, chained aggregation 10.0 to 10.4 s, 18.1 s a block, 103.9 MH/s; 20: the same (18.0 s, 103.7); 18: 6.7 to 7.0 and 8.8 to 8.9 s, 15.6 s a block, 99.3 MH/s (minus 4.4%); 16: 4.9 to 5.1 and 6.1 to 6.2 s, 11.1 s a block, 83.8 MH/s (minus 19%), reproduced (11.1 s, 84.0). Against the card alone (4.1 s a block, 2.1 s an aggregation): the shortest kernel buys the prover 1.6x for a fifth of the hash rate and the slowdown stays 2.7x, so the 3-second aggregation on a mining card is not reachable by the kernel length; the defaults stay; the 2^16 trade and the batch fold (a new pinned guest, about 0.7 s a block alone by the step costs, estimated) are the project lead's decision in docs/plans/proving-v1.md "Aggregation cost (5 October, night)". Branch agg-cost ea38ece (worktree igneum-wt-agg-cost, on proving-v1's docs tip 9be5817, not pushed).

The ledger-fixes-0311 suites on PC 2 (build-20261006-012543, published by me from the ledger-rebase worktree at 01:25:43Z, 01:26:15 to 01:32:33Z, 378 s): the Linux node build 137 s, the Windows node build 163 s, the seven suites exit 0 in 46 s, 7 files uploaded. The 46 s says the suites ran WITHOUT the igneum-pow feature (the G6 caveat of 00:29 applies: the lottery-hash consensus tests were not compiled; the Mac run with the feature is the suite evidence, 99/1/4 on kaspa-consensus with the M20 era test the one failure, 108/0/2, 52/0, 36/0, 14/0, 23/0, 20/0 on the rest). Fork tip fbb0082a, docs tip 9872299 (ledger-rebase).

The job-runner finding (the aggregation-cost agent, 01:30): the app ran three jobs on PC 2 at once from 01:24Z (run-prover-on-pc2 at 01:24:21Z and the suites build at 01:26:15Z landed while nothing else of the queue was expected to run), whereas at 23:40Z PC 2's update-now queued behind the hung sweep ("1 new for this machine, 1 queued"). So the app serialises the jobs it fetches in one poll and runs jobs from different polls side by side; "one job on PC 2 at a time" held tonight by coordination only. Next-cut item (with the update-now item): the runner takes one job at a time per machine whatever the poll, and a collect or update-now may pre-empt.

PC 2's prover: off from 00:49Z (the repro job) to 01:24:09Z (job 6's finally block), then confirmed on by run-prover-on-pc2 at 01:24:21Z; about 35 minutes of PC 2's shard payouts lost. Every PC 2 job of the night is closed; nothing is queued. The crossing watcher (pid 49195) runs to 08:00Z; the morning reads /tmp/igneum-devnet/crossing-154800.out first. The coordinator's final report to main follows this entry; the status file ends here unless the morning adds to it.

03:56 THE CROSSING HELD: epoch 43 is class v3 (generator 3) from DAA 154,814 at 03:51:42Z; the chain continued

The watcher (pid 49195, /tmp/igneum-devnet/crossing-154800.out): epoch 42 read "next v3" from 02:51:53Z (DAA 151,212); the first template of epoch 43 at 03:51:42Z: class v3, generator 3, era 0, seed edc4fa84 (the genesis stand-in), boundary 158,400, DAA 154,814. The consequences reviewer's own watcher read DAA 154,808 at 03:51:38Z. The chain after the crossing: DAA 154,943 at 03:54:06Z (block time 1.22 s, 4 miners, 137 MH/s unchanged, the reviewer's read); the observer at 03:55Z: DAA 155,004, blocks 109,467, sink moving; 61 to 62 blocks per minute before, the after-rate and the verdict (PASS if the after-rate holds a quarter of the before-rate with no 10-minute gap) land at about 05:21Z in the same file. The block count fell from 153,857 to 108,621 between 03:35:41Z and 03:40:41Z: the observer's pruning at 03:39:45Z ("SMT root was rebuilt successfully following pruning"), not a reorg (the DAA rose through it); it sits inside the before-window only, so the verdict's before-rate (the last 10 minutes before the crossing) is clean. Still to read in the morning: each machine's STATUS rate after the swap and the intake's "pack refused", "need" and "mismatch" counts (the reviewer takes them at 04:02Z into the ledger), PC 1's relaunch (it crosses at its relaunch through the manifest), and the laptop's return. C46 (the reviewer, 03:54Z): /api/stats still prints "generator v2" from a literal in site/api/stats.mjs line 39, so the public API misdescribes the chain from the first v3 epoch; next cut: the handler names the class and generator the observer records from the epoch line.

04:17 proving v1 active with nothing proven (C47): the morning summary says so

The consequences reviewer at 04:16Z: node 1's igneum_getProvingStatus.v1 reads active, segmentsInWindow pending 54, proven 0, unproven 21, paidSegments 0, no segment record carried; PC 2's upload shows no proving line since 03:44Z while v0 shards still paid at 13 per 10 minutes. So the honest line for the morning is "class v3 crossed and held; proving v1 active since 03:51:42Z with 0 segments proven in its first 25 minutes (cause pending from the proving agent)", not "0.3.11 activated". Consequence per tier while it lasts: no segment pays and no aggregator share (10%) is earned, v0 shards keep paying; the fee switch at H = 210,000 (about 18:56Z) is unaffected but C1 (every prover on 0.3.11 by 16:00Z) now has a second condition, that a prover proves v1 segments at all. The proving agent has the ask (PC 2's prover state after the crossing, the app's v1 segment path, aggregate_once); any PC 2 job it needs is one job on my go; a 0.3.11 app cause joins the 0.3.12 list at the top beside 6714a45.

04:17. C47's cause, read from the shipped app (app/igneum-app/src/prover.rs, the aggregate_once doc and line 657): one aggregation attempt needs one shard proof per shard of EVERY block of the segment in this node's pool before it runs igneum-prove-host --mode aggregate, else it answers "segment a..b: waiting for shard proofs ... in this node's pool". With one prover (PC 2) at about 2.7 percent block coverage, eight consecutive proven blocks never occur, so proving v1 yields zero segment records on tonight's devnet by arithmetic, not by a fault; the aggregator share accumulates in escrow until the fleet reaches the proving plan's coverage rows (47 mining 5090-class cards with the shipped shard loop, 18 through the chain mode, or about 6 proving-only cards at one block per second; the proving agent's fleet table, corrected 04:30Z). The morning line: "class v3 crossed and held; proving v1 active, 0 segments proven and none expected at one prover". No PC 2 job; the proving agent confirms from PC 2's log and writes it into docs/plans/proving-v1.md. What it means for the public page: the proving line stays "every block proven" as a design, and the devnet shows the per-block shards (v0) paying while segments wait on coverage; the C1 check at 16:00Z is about the binaries, not about segments.

04:28. Confirmed by the proving agent from PC 2's app log (run win-1ccfe586-20261005-235130): the prover is on and the aggregator loop runs the v1 path every 42 s ("aggregator: segment N..N+7: waiting for shard proofs N/0 ... N+7/0 in this node's pool", all eight missing on every pass); one 5090 proves 13 shards per 10 minutes of about 600 blocks (2.2 percent), so each segment passes its 600-DAA deadline unproven; node 1 at 04:2xZ: pending 55, proven 0, unproven 20, paid 0. No fault, no 0.3.12 item. DECISION FOR [user] (7, the morning): the fix's shape is cards (47 mining 5090-class cards with the shard loop as shipped in 0.3.11, 18 through the chain mode on mining cards, or about 6 proving-only cards, at one block per second on empty blocks; the proving agent's fleet table, corrected 04:30Z) or a smaller proving_v1_segment_blocks for a small devnet (1 or 2 instead of 8), which is a consensus parameter and so a new override object, a new digest and a two-manifest publish at a new height (tip + 14,400, the same rules as tonight); not tonight (main's rule: no further rollout). Until then the aggregator share sits in escrow and v0 shard payouts continue; the public testnet's genesis carries v1 from day one with whatever segment length the coverage rows justify.

05:59 VERDICT PASS; H re-cut to about 19:12 to 19:20Z

The watcher's verdict at 05:21:48Z: PASS, class v3 after the crossing, 58.7 blocks per minute in the ten minutes before against 59.2 in the ninety after, no gap of ten minutes, DAA 160,148 and 114,611 blocks since the observer's pruning; epoch 44 (DAA 158,400) crossed clean. The reviewer's measured rate from the crossing to 05:57:38Z (DAA 162,295): 0.99 DAA/s, so H = 210,000 lands about 19:12 to 19:20Z, later than the 18:45 to 18:56 quoted overnight (the stalls pulled the earlier average forward); the 16:00Z C1 check keeps about 3.2 hours of margin and nothing in the order changes. The watcher exits on its own; the status file ends here.

07:03 PC 1 is back; the source of its 22:31Z quit is named (C35 closed); the Ember re-run waits on the project lead

PC 1's app came back at about 06:59Z (job pc1-morning-intake-1 ran 07:01:45Z); it takes 0.3.11 and the nine fields from the manifest at that start. The C35 source, from the second engine's own log (collect ember-c35-collect-1, the Ember agent, 06:59Z): the second engine the Ember playbook started reported version 0.3.9 (the branch's Cargo version), under the manifest's min_supported_version, so its updater treated 0.3.10 as urgent (the urgent rule beats the auto_update = false the playbook wrote into the copied settings), downloaded it at 22:31:02Z and ran the per-user installer at 22:31:05Z, whose PrepareToInstall sent POST /api/quit to the INSTALLED app (quit logged 22:31:06Z); the second engine then hung on the inherited pipe until the relay lane ended it. So PC 1's quit was the playbook's second engine through the installer, not an administrator prompt and not the host. Fix e600e63 on ember-tune: a second engine never runs the updater (Engine.no_ota from IGNEUM_APP_NO_OTA=1, implied by --sweep; the OTA tick skipped and Check now refused; logged at start), both playbooks set it, tools/ci/second-engine-check.sh demands it beside the file-not-pipe and tree-kill lines; 93 app tests pass. PC 2's 20:01Z quit is a different case (it came back as 0.3.9, so not an installer) and keeps the morning's prompt test. The Ember re-run (5090 baseline, the 9070 XT power ladder, no prompt) is built and held: PC 1 is the project lead's desk and he is at the machine, so it runs on his word through main, not on the night scheduler's; it needs the build of e600e63 and the two fetches republished after PC 1's update.