Merge remote-tracking branch 'origin/master' into peer-directory

This commit is contained in:
igneum-labs 2026-10-07 14:07:42 +00:00
commit b6a0a78e12
155 changed files with 6848 additions and 1278 deletions

45
.github/workflows/ci-red.yml vendored Normal file
View file

@ -0,0 +1,45 @@
# The red watcher as its own workflow, on workflow_run, so the copy on master watches EVERY branch's ci run whatever
# ci.yml that branch carries: GitHub runs a workflow_run workflow from the default branch only, and the branch's own
# ci.yml never enters it (7 October 2026: the inline `red` job of ci.yml was conditioned on master and release-*, and
# a feature branch would have waited for a merge of master before its reds were posted at all).
#
# One line per failed run (tools/ci/red-watch.mjs record, idempotent per run attempt) to /srv/ci-red/red.jsonl on the
# box; the box's igneum-ci-red.timer posts each new line once to the hidden updates channel, naming the branch, the
# commit, the red check and the pushing author. Runs on the box's own runner (not a GitHub-hosted machine: the billing
# block of 6 October 2026, 18:37Z to 20:10Z, failed every hosted job at start and nobody was told). Never blocks a
# release: it reads the run, writes one line, and ends.
name: ci-red
on:
workflow_run:
workflows: [ci]
types: [completed]
jobs:
red:
name: red watcher (every branch; one line per failed run, with the branch, commit, red check and pushing author, to the updates channel and the box file)
if: ${{ github.event.workflow_run.conclusion == 'failure' }}
# the label ci-red is on igneum-build-1 only (added through the runners API on 7 October 2026; the default of
# RUNNER_LABELS in provision.sh carries it): the record file and the poster (igneum-ci-red.timer, the webhook file)
# live on that box, and the pool label igneum-build-1 is shared with igneum-build-2 since the same day
runs-on: [self-hosted, linux, x64, ci-red]
timeout-minutes: 5
permissions:
actions: read # the failed run's jobs API (the first real red run, 21:19Z on 6 October: the default token answered 403 and the line carried no step)
contents: read
steps:
- uses: actions/checkout@v4
with:
sparse-checkout: tools/ci
- name: record the failed run (one line, the branch, the commit, the failed jobs and their first failed step from the run's own API, the pushing author)
env:
GITHUB_TOKEN: ${{ github.token }}
RED_WATCH_RUN_ID: ${{ github.event.workflow_run.id }}
RED_WATCH_ATTEMPT: ${{ github.event.workflow_run.run_attempt }}
RED_WATCH_WORKFLOW: ${{ github.event.workflow_run.name }}
RED_WATCH_BRANCH: ${{ github.event.workflow_run.head_branch }}
RED_WATCH_SHA: ${{ github.event.workflow_run.head_sha }}
RED_WATCH_EVENT: ${{ github.event.workflow_run.event }}
RED_WATCH_URL: ${{ github.event.workflow_run.html_url }}
RED_WATCH_ACTOR: ${{ github.event.workflow_run.actor.login }}
RED_WATCH_TITLE: ${{ github.event.workflow_run.head_commit.message }}
RED_WATCH_AUTHOR: ${{ github.event.workflow_run.head_commit.author.name }}
run: node tools/ci/red-watch.mjs record --file /srv/ci-red/red.jsonl

View file

@ -8,13 +8,16 @@
# local gate and CI cannot drift (6 October 2026: 131 red `ci` runs in three days, 92 of them on master, every one a
# tree check that would have failed on the pushing machine in under 25 s; docs/analysis/ci-failures-2026-10-06.md).
#
# Where it runs: `pow` and `sims` go to the box's runner (igneum-build-1, rustc pinned, sccache read-only, 48 jobs)
# Where it runs: `pow` and `sims` go to the self-hosted pool (label igneum-build-1: the runners on igneum-build-1 and, since
# 7 October 2026, igneum-build-2, which carries that label too; rustc pinned, sccache read-only) and only when the push
# touched code (the `changes` job; a docs-only push skips them)
# when the repository variable IGNEUM_CI_RUNNER is `box`, else to ubuntu-latest (docs/plans/ci-self-hosted.md; GitHub
# has no fallback in runs-on, the variable is the switch). The `site` job stays on GitHub's machines. The `red` job
# runs on the box after any failed run on ANY branch and records the failure for the watcher
# (tools/ci/red-watch.mjs; infra/build-server/ci-red): one line per run, naming the branch, the commit, the red check
# and the pushing author, to the hidden updates channel and to /srv/ci-red/red.jsonl, so nobody opens the Actions page
# to learn a branch is red (master and release-* only until 7 October 2026, when eight red runs on ca3-v4-node went unseen).
# has no fallback in runs-on, the variable is the switch). The `site` job stays on GitHub's machines. The red watcher
# is its own workflow, .github/workflows/ci-red.yml (workflow_run, so the copy on master watches every branch's run
# whatever ci.yml that branch carries): one line per failed run, naming the branch, the commit, the red check and the
# pushing author, to the hidden updates channel and to /srv/ci-red/red.jsonl (tools/ci/red-watch.mjs;
# infra/build-server/ci-red), so nobody opens the Actions page to learn a branch is red (the inline `red` job here
# watched master and release-* only until 7 October 2026, when eight red runs on ca3-v4-node went unseen).
#
# What does not run, on purpose: the node fork (vendor/igneum-node*, a rusty-kaspa fork of about 500 crates with
# rocksdb, blst and the execution layer) is gitignored here and too big for the free runners today (a cold build is
@ -25,8 +28,39 @@ on:
push:
pull_request:
jobs:
changes:
# What the push touched (tools/ci/docs-only-check.sh): a push of documents only (docs/, site/, *.md) skips the two
# compile-or-compute jobs below, which read none of those paths, so the self-hosted queue carries only runs that can
# change their result (7 October 2026: 31 runs queued on one runner, most of them status-document pushes). The tree
# gate (the `site` job) runs on ubuntu-latest for every push. A pull request, a new branch or a force push answers
# code=true (no `before` to compare from), as does any error reading the compare API: when in doubt, run.
name: what the push touched (docs-only runs skip the Rust and simulator jobs)
runs-on: ubuntu-latest
outputs:
code: ${{ steps.classify.outputs.code }}
steps:
- uses: actions/checkout@v4
with:
sparse-checkout: tools/ci
- id: classify
env:
GH_TOKEN: ${{ github.token }}
BEFORE: ${{ github.event.before }}
AFTER: ${{ github.sha }}
REPO: ${{ github.repository }}
EVENT: ${{ github.event_name }}
run: |
if [ "$EVENT" != push ] || [ -z "$BEFORE" ] || [ "$BEFORE" = 0000000000000000000000000000000000000000 ]; then
echo "code=true" >> "$GITHUB_OUTPUT"; echo "no base to compare from ($EVENT): the compile jobs run"; exit 0
fi
files="$(gh api "repos/$REPO/compare/$BEFORE...$AFTER" --paginate --jq '.files[].filename' 2>/dev/null || true)"
line="$(printf '%s\n' "$files" | bash tools/ci/docs-only-check.sh)"
echo "$line" >> "$GITHUB_OUTPUT"
echo "$line: $(printf '%s\n' "$files" | grep -c .) changed path(s) between ${BEFORE:0:8} and ${AFTER:0:8}"
pow:
name: igneum-pow tests, igneum-census build
needs: changes
if: ${{ needs.changes.outputs.code == 'true' }}
runs-on: ${{ vars.IGNEUM_CI_RUNNER == 'box' && fromJSON('["self-hosted", "linux", "x64", "igneum-build-1"]') || 'ubuntu-latest' }}
steps:
- uses: actions/checkout@v4
@ -42,6 +76,10 @@ jobs:
run: cargo build --release
sims:
name: simulators, quick modes
needs: changes
# master and release-* pushes, and pull requests into them, only (main, 7 October 2026: every code push cost two box jobs and the
# queue read 22); a feature-branch code push runs the igneum-pow tests alone. tools/ci/sims-branch-check.sh holds this rule.
if: ${{ needs.changes.outputs.code == 'true' && ((github.event_name == 'push' && (github.ref == 'refs/heads/master' || startsWith(github.ref, 'refs/heads/release-'))) || (github.event_name == 'pull_request' && (github.base_ref == 'master' || startsWith(github.base_ref, 'release-')))) }}
runs-on: ${{ vars.IGNEUM_CI_RUNNER == 'box' && fromJSON('["self-hosted", "linux", "x64", "igneum-build-1"]') || 'ubuntu-latest' }}
steps:
- uses: actions/checkout@v4
@ -71,33 +109,13 @@ jobs:
- uses: actions/setup-node@v4
with:
node-version: '22'
- name: a headless Chromium for the text-overlap sweep (Playwright outside the tree; the gate finds it through IGNEUM_PLAYWRIGHT_DIR)
run: |
mkdir -p /tmp/pw && cd /tmp/pw && npm init -y >/dev/null && npm i --no-audit --no-fund playwright@1.56 | tail -1
npx playwright install --with-deps chromium | tail -1
echo "IGNEUM_PLAYWRIGHT_DIR=/tmp/pw" >> "$GITHUB_ENV"
- name: the tree gate, tools/ci/pre-push.sh --ci (the same script the pre-push hook runs; one line per check, a red check prints its output)
run: bash tools/ci/pre-push.sh --ci
- name: public stats API answers with the documented fields (the live site; master only, the endpoints exist there after the merge)
if: github.ref == 'refs/heads/master'
run: node tools/ci/public-api-check.mjs https://igneum.network
red:
# Runs when a run on any branch has a failed job, on the box's own runner (not a GitHub-hosted machine:
# the billing block of 6 October 2026, 18:37Z to 20:10Z, failed every hosted job at start and nobody was told).
# tools/ci/red-watch.mjs record appends ONE line for this run to /srv/ci-red/red.jsonl (idempotent per run attempt);
# the box's igneum-ci-red.timer posts each new line once to the hidden updates channel. Never blocks a release:
# it reads the run, writes one line, and ends.
name: red watcher (every branch; one line per failed run, with the branch, commit, red check and pushing author, to the updates channel and the box file)
needs: [pow, sims, site]
if: ${{ failure() }}
runs-on: [self-hosted, linux, x64, igneum-build-1]
timeout-minutes: 5
permissions:
actions: read # the run's jobs API (the first real red run, 21:19Z: the default token answered 403 and the line carried no step)
contents: read
steps:
- uses: actions/checkout@v4
with:
sparse-checkout: tools/ci
- name: record this run (one line, the branch, the commit, the failed jobs and their first failed step from the run's own API, the pushing author)
env:
GITHUB_TOKEN: ${{ github.token }}
RED_WATCH_TITLE: ${{ github.event.head_commit.message }}
RED_WATCH_AUTHOR: ${{ github.event.head_commit.author.name }}
run: node tools/ci/red-watch.mjs record --file /srv/ci-red/red.jsonl

View file

@ -0,0 +1,65 @@
// Igneum vendor and OS marks (gpu-logos, 7 October 2026): the strings the miner app ships in app/igneum-app/ui/app.js
// (View.MARKS and View.VENDORS), exported verbatim for the site and anything else that names the hardware.
// view.test.mjs fails when this file and app.js drift apart, so edit app.js first and regenerate this file with
// `node brand/marks/regen.mjs` (or copy the strings by hand; the test says which one moved).
//
// Each glyph: a hand-drawn simplified monochrome mark of the vendor's public geometry (never a copied logo file,
// never a raster), 24 x 24 viewBox at 22 px, under 460 bytes, fill or stroke through currentColor so the element's
// colour tints it. Nominative use that names the hardware; the ember accent is for state and never tints a brand.
//
// The treatment in the app (app.css, the block at the end): a 44 x 44 well, radius 12, background the vendor colour
// at .14 alpha (dark) or .10 (light), a 1 px ring in the vendor colour at .45 alpha, the glyph in the full colour;
// hover and focus-within add a 3 px halo of the well colour; nothing animates. The light hex of every vendor reads at
// 3:1 or better on its well over white (nvidia 3.87, amd 5.01, intel 4.29, apple 7.52, gpu 4.65).
//
// Class names in the app: .badge.<vendor> (nvidia | amd | intel | apple | gpu), .badge.mini for a 26 px inline mark,
// .gen for the series line under the name. Tokens: --mark-<vendor>, --mark-<vendor>-well, --mark-<vendor>-ring.
export const VENDORS = {
"nvidia": {
"label": "NVIDIA",
"dark": "#8BE37A",
"light": "#2F8A22"
},
"amd": {
"label": "AMD Radeon",
"dark": "#FF5A5A",
"light": "#C41E2A"
},
"intel": {
"label": "Intel",
"dark": "#7CC4FF",
"light": "#1C6FD6"
},
"apple": {
"label": "Apple",
"dark": "#E6E3DD",
"light": "#4A4A50"
},
"gpu": {
"label": "GPU",
"dark": "#9A9A9E",
"light": "#6B6B70"
}
};
export const WELL_ALPHA = { dark: 0.14, light: 0.1 };
export const MARKS = {
"nvidia": "<svg viewBox=\"0 0 24 24\" width=\"22\" height=\"22\" aria-hidden=\"true\" focusable=\"false\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.8\" stroke-linecap=\"round\" stroke-linejoin=\"round\"><path d=\"M2.6 12C5 7.9 8.3 5.8 12 5.8s7 2.1 9.4 6.2c-2.4 4.1-5.7 6.2-9.4 6.2S5 16.1 2.6 12z\"/><path d=\"M15.6 12A3.6 3.6 0 1 0 12 15.6\"/><circle cx=\"12\" cy=\"12\" r=\"1.2\" fill=\"currentColor\" stroke=\"none\"/></svg>",
"amd": "<svg viewBox=\"0 0 24 24\" width=\"22\" height=\"22\" aria-hidden=\"true\" focusable=\"false\" fill=\"currentColor\"><path d=\"M8 3h13v13l-4-4V7h-5z\"/><path d=\"M3 8l4 4v5h5l4 4H3z\"/></svg>",
"intel": "<svg viewBox=\"0 0 24 24\" width=\"22\" height=\"22\" aria-hidden=\"true\" focusable=\"false\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.8\" stroke-linecap=\"round\"><path d=\"M7.4 7.6C4.9 8.8 3.4 10.5 3.4 12.3c0 3.4 4.3 5.9 9.8 5.9 4.5 0 7.6-1.6 7.6-3.9 0-1.4-1.3-2.6-3.5-3.3\"/><path d=\"M11.2 9.6v6.2\"/><circle cx=\"11.2\" cy=\"6.6\" r=\"1.1\" fill=\"currentColor\" stroke=\"none\"/></svg>",
"apple": "<svg viewBox=\"0 0 24 24\" width=\"22\" height=\"22\" aria-hidden=\"true\" focusable=\"false\" fill=\"currentColor\"><path d=\"M16.4 12.6c0-2.5 2-3.6 2.1-3.7-1.2-1.7-3-1.9-3.6-2-1.5-.2-3 .9-3.8.9-.8 0-2-.9-3.3-.8-1.7 0-3.2 1-4.1 2.5-1.8 3-.5 7.6 1.3 10.1.9 1.2 1.9 2.6 3.2 2.5 1.3 0 1.8-.8 3.3-.8 1.6 0 2 .8 3.3.8 1.4 0 2.3-1.2 3.1-2.5 1-1.4 1.4-2.8 1.4-2.9 0 0-2.7-1-2.9-4.1zM13.9 5.3c.7-.8 1.2-2 1-3.2-1 0-2.2.7-2.9 1.5-.6.7-1.2 1.9-1 3 1.1.1 2.2-.5 2.9-1.3z\"/></svg>",
"gpu": "<svg viewBox=\"0 0 24 24\" width=\"22\" height=\"22\" aria-hidden=\"true\" focusable=\"false\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.7\" stroke-linecap=\"round\"><rect x=\"5.5\" y=\"5.5\" width=\"13\" height=\"13\" rx=\"2.2\"/><rect x=\"9.5\" y=\"9.5\" width=\"5\" height=\"5\" rx=\"1\"/><path d=\"M9 2.5v3M12 2.5v3M15 2.5v3M9 18.5v3M12 18.5v3M15 18.5v3M2.5 9h3M2.5 12h3M2.5 15h3M18.5 9h3M18.5 12h3M18.5 15h3\"/></svg>"
};
// OS marks in the same treatment: Apple is the vendor glyph; Windows is the four slanted panes
export const OS_MARKS = {
macos: MARKS.apple,
windows: "<svg viewBox=\"0 0 24 24\" width=\"22\" height=\"22\" aria-hidden=\"true\" focusable=\"false\" fill=\"currentColor\"><path d=\"M3 5.6l7.3-1v7.1H3zM11.4 4.4L21 3v8.7h-9.6zM3 12.3h7.3v7.1L3 18.4zM11.4 12.3H21V21l-9.6-1.4z\"/></svg>"
};
export const OS_COLOURS = { macos: VENDORS.apple, windows: { label: 'Windows', dark: '#7CC4FF', light: '#1C6FD6' } };
// the CSS tokens for both themes, as the app declares them
export function tokensCss() {
const line = (theme) => Object.entries(VENDORS).map(([v, c]) => { const h = c[theme], r = parseInt(h.slice(1, 3), 16), g = parseInt(h.slice(3, 5), 16), b = parseInt(h.slice(5, 7), 16), a = String(WELL_ALPHA[theme]).replace(/^0/, ''); return `--mark-${v}:${h};--mark-${v}-well:rgba(${r},${g},${b},${a});--mark-${v}-ring:rgba(${r},${g},${b},.45)`; }).join(';');
return `:root{${line('dark')}}\n@media (prefers-color-scheme:light){:root:not([data-theme="dark"]){${line('light')}}}\n:root[data-theme="light"]{${line('light')}}`;
}
// the well: <div class="badge nvidia"><svg…></svg></div>
export function markHtml(vendor, size) { const v = MARKS[vendor] ? vendor : 'gpu'; return '<div class="badge ' + (size ? size + ' ' : '') + v + '" data-mark="' + v + '" title="' + VENDORS[v].label + '">' + MARKS[v] + '</div>'; }

View file

@ -377,3 +377,28 @@ spill-over. Now (lib.sh `bs_route_spill`, master from this commit):
them keeps its number; since this commit a bounded run takes the band its slot owns (slot 0 the last 32 cores, slot 1 the 32
below, slot 2 the 32 below that), so three bounded runs never share a core. A gate still takes its slot ahead of queued suites.
- `--box N` still pins. Self-test: tools/ci/route-spill-check.sh (thirteen cases through `BS_ROUTE_STATE_<n>`, no ssh), in the gate.
## 8. The measure file is retired: per-core leases and the quiet class (7 October 2026, 15:07 UK)
Main's reading at 15:07 UK: on build-1 a sync-fuzz probe (the capacity lane's, SIGSTOPped since 09:45Z, no owner) held the global
measure flock for five and a half hours; beside it attack-f6's phase2b waited exclusive on the same file behind seven F2 solvers
holding it shared on cores 6-11,54-59, and every new shared taker (every build) queued behind the exclusive waiter: load 120, slots
free, nine waiting. On build-2 the era VDF bench held the file exclusive while pinned to one core and five jobs waited. One global
exclusive lock across unrelated measurements was the wrong design. Now:
- `infra/build-server/lease.sh`, installed on every box at `/srv/builds/_bin/lease` (provision.sh; by hand on build-1 and build-2
at 14:2x BST): `lease cores <set> --label "..." [--owner <agent>] -- <cmd>` takes a lease on THOSE CORES ONLY (one flock per
core, `_locks/core-<n>`, a `wait-<pid>` file with the label while it waits), runs the command under nice 10 and taskset, and
releases. Nothing else is excluded. `lease quiet --label "..." --owner <agent> -- <cmd>` is the whole-box class: refused (exit
73) while any slot or core lease is held, capped at 20 minutes, holder line with the owner in `_locks/quiet`. `lease status`,
`lease reap`.
- remote-run.sh: an unbounded run (nice 0, the full set) takes `quiet` shared and waits for it; a bounded run (suites, benches,
everything on box 2) never takes it. Every run keeps off leased cores (its set minus the `core-<n>` flocks, said once). A run's
keeper refreshes its holder file's mtime every 20 s and calls `lease reap`: a holder of a lease, the quiet file or a slot whose
process has been STOPPED for 5 minutes is killed and its file cleared, one line each in `_log/reaped.log`. The keeper closes the
lock descriptors it inherits (an orphaned `sleep 20` held a slot and a worktree lock 20 s past the release). BR_MEASURE=1 is the
quiet class with the same refusals; it needs a named owner (IGNEUM_AGENT).
- The lanes' own `flock -s /srv/builds/_locks/measure -c "nice -n 10 taskset -c <set> ..."` lines no longer hold anything a build
waits for; they become `/srv/builds/_bin/lease cores <set> --label "..." -- <cmd>`, and `flock -x .../measure` becomes
`lease quiet`. Self-tests: lease.sh --self-test and remote-run.sh --self-test-slots, run on build-1 by tools/ci/box-locks-check.sh
in the gate.

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,503 @@
# Igneum Miner 0.3.17: the 0.3.17 tree as cut on 7 October 2026
The hourly cut after 0.3.16, under the coordinator's rule: every new switch at never on the devnet so the digest does not move; cut what is green and list what waited. Worktree `igneum-wt-ship0317` (release-0.3.17 from release-0.3.15 a9eb58f1), node `release-0.3.17-node` in `vendor/igneum-node-0317` (from f1ea7a38).
> Renumbered 7 October 2026, 03:3xZ (main): the tree this plan describes is **0.3.18**. Its canary on 12153428 failed on two further faults (the idle-peer guard closes every peer mid headers-proof IBD; a window-below-epoch log flood at 1,900 lines a minute), and tonight's **0.3.17** became the node-only hotfix on the 0.3.16 app tree (f1ea7a38 plus 90aaf38e = b3c228fa; see docs/plans/release-0.3.17.md on release-0.3.17). The version scheme is three-part everywhere, so "0.3.16.1" was not available. Rows below keep "0.3.17" where they were written; read it as this tree. Branches: release-0.3.18 (app), release-0.3.18-node (fork, 12153428). Decimals is 0.3.19.
> Renumbered again 7 October 2026, 08:2xZ (main): this feature tree is **0.3.19**; **0.3.18** is the app-only cut on the 0.3.17 tree (the engine's clock sample gated on igneum_getExecStatus, ledger N7; docs/plans/release-0.3.18.md on release-0.3.18-app); decimals is 0.3.20. Branches now: release-0.3.19 (app), release-0.3.19-node (fork, the mirror). Rows below keep the "0.3.18" they were written with; read it as this tree.
> Renumbered a third time, 7 October 2026 10:4x UK (main): this feature tree is **0.3.20** (the dc141409 line plus the class v4 amendment, option A on AP-F8-1, and the proof archive; its canary gate as ordered). **0.3.19** is an app-only cut on the 0.3.18 app tree plus miner-ui-5 on the node pin 5899f603 (docs/plans/release-0.3.19.md on release-0.3.19). Branches: release-0.3.20 (app), release-0.3.20-node (fork, the mirror). Rows below keep the numbers they were written with.
## 1. The node tree
| Branch | Tip | In 0.3.17 | Why |
|---|---|---|---|
| ca3-v4-0316 (node lane) | 9a44fcb8 (df787727 the last functional commit: igneum_getNodeInfo; the miner stall guard e6e1fbe2 with exit 45, the idle-peer drop 96619b7d) | yes | the unwrap class, the seven-window rule, the finality cause fields, igneum_getRecentBlocks, difficulty rule v3, the fork gate, vote-or-burn, the signing bonus, the miner's stall exit (every switch absent from the live object = never) |
| ladder-node (ladder lane) | 1591ee1d | yes | the latency ladder behind latency_ladder_activation_daa (never); the receive gate reads header_signals_active (the lane's merge note, applied as c6e12005's parent commit) |
| exec-sync-0313 (proving lane) | 40fd8d8c | yes (merged last, 02d15a87; params.rs both sides kept) | consensus proof verification behind proving_consensus_verify_daa (never), the override's program ids as the statement's ids, sp1 as cfg(not(windows)) dependencies with the host-backed verify on Windows (the first merge of a344fea3 broke the Windows cross-build in sp1-jit; 40fd8d8c fixed it) |
| finality-pause-node-0317 (finality lane) | fdc71bb5, rebased | yes (cherry-picked as d8d1de1a onto the exec-free base) | W7 the departure announcement behind finality_leave_activation_daa (never on the devnet; 0 and leave_delay 3600 in the testnet genesis per the project lead); protocol 17; the devnet digest unchanged by test |
| tail-emission-node-0317 (economy lane) | 5647533f, rebased | yes (cherry-picked as 996b592f, its params hunks without exec-sync's proving fields) | its coinbase table replaces the subsidy table the silence rules build on; the rebase resolves it; no digest move while CURRENT |
| kaspad igneum-pow default feature | 75467244 | yes | the node lane's N5 finding: a bare `cargo build -p kaspad` validated PoW with the kHeavyHash stub; every box reads igneum_getNodeInfo's powEngine "igneum-pow" in the table, "stub" is a FAIL |
| kaspa-testing-integration | cc78a21f | fixed here | compiles again (the block store's evm argument, Header's vote_key_hash, the Params const as a let, the four finality ops in the RPC sanity match); the box suite is to compile every test crate from here |
Final node tree 02d15a87 (75467244 + exec-sync-0313 40fd8d8c). Digests on the exec-free tree c6e12005 (Mac binary, direct run): the live sixteen-field object eada4bda (unchanged), the thirteen-field file b18ed271, no file c562d70e. A harness note: scratchpad digest.sh printed the no-file digest for a file that reads eada4bda while other nodes held its ports; it now says UNTRUSTED when the override lines are missing.
## 2. The app tree
| Branch | Tip | In 0.3.17 |
|---|---|---|
| ladder (repo side) | 7003f9f | yes (igneum-pow's chain_program_shadow, which the fork's kaspa-pow needs) |
| exec-app-0314 | 032f03c | yes (clean; docs as the union) |
| relay | 28c028b | yes (ci.yml on the release side) |
| ember-tune-0317 (rebased on a9eb58f1) | bb37c993 | yes: the release lines kept (job_active, install_asked, the stale-exporter line), live.rs one merged module, publish-manifest.sh Ember's guard |
| miner-ui-4 (rebased on a9eb58f1) | afc331bd | yes: ui/ whole with Ember's finality words, extnode.rs on igneum_getNodeInfo, N4 stall exits, the port-collision check, POST /api/shot, an external node that goes away hands the ports to the app after 60 s; src/live.rs stays Ember's |
| proving-v2 | d8055a7 | NO (its exec content is in c00e608; its extra is the proof-system v2 slot, its own cut, per its owner) |
App tests on the merged tree: 175 + 28 + 8; UI 41. A whole-file take of a lane's branch on an older base dropped the release tree's own fixes (seen and reverted): the lanes rebase onto the release tree and resolve there; the shipper merges, never substitutes files.
## 2a. Linux binaries by glibc (main's rule, 7 October 2026, 01:3x UK)
HiveOS images are Ubuntu 20.04 (glibc 2.31); the canary pods on 22.04 (2.35) refused the box's binaries. From now: the HiveOS package carries the 2.31 zig build (the build-server lane produces and container-tests it; the 0.3.17 HiveOS row stays at 0.3.16 in the index until that package lands, then joins with its own read-back); seeds and generic Linux binaries are the 2.35 build; the fleet's own boxes (Ubuntu 24.04) take the box's native build. The Mac and Windows rows publish on their gates.
## 2b. Checklist line for every cut (the node lane, 7 October 2026)
`cargo check -p kaspa-testing-integration --tests` runs in the box suite of every cut and in the fork's CI (the crate had not compiled since the finality fields landed; no integration test ran on any cut until this one). The 0.3.17 box chain carries it as its last step (rebuild-on-fix.sh step_box / the node chain's "integration check"). 0.3.18 node notes: relay pipelining bb397f99 and the sync-request fuzz gate 6cf33d36 on ca3-v4-0318; aa0182aa and 812c3ac2 from ca3-v4-0316.
## 2c. A kill matches a pid, never a name (main's row, 7 October 2026, 01:xx UK)
My `pkill -f build-remote.sh` at 00:17 UK (to stop my own chained box builds before a rerun) killed the decimals lane's box run on the same Mac: the pattern matched every build-remote.sh, not mine. Rule: every kill in the ship tooling matches a pid file or the exact command line (the run's log path, the worktree path), never a tool's name; the ship scripts run through tools/ci/pre-push.sh's kill-by-name check before the next cut. Tonight's scratch runbooks hold no pkill by tool name any more (the chains are started with their log path in the command and stopped by that path).
## 2d. The canary's FAIL (01:37Z): the IBD guard refuses signal-version relay blocks
c17-1 (02d15a87, the live sixteen-field object) could not join: "IBD with headers proof from <peer> was unsuccessful (peer relayed block ... header version mismatch: got 1026, expected 2 at DAA score 237583)", 483 lines in 15 minutes against all four peers, peers 0, blocks 0. Cause: protocol/flows/src/ibd/flow.rs:390 (upstream's Toccata guard, f94053a0) compares the syncer's relay-block version with the plain block version before the pruning proof; under the open window the sink carries legal 1026 signal blocks. The line is on f1ea7a38 too, so the LIVE devnet has refused every fresh join (the headers-proof IBD path) since publish 2 opened the window at 22:43Z; tonight's moves passed because every node had a datadir (the relay path) or began its IBD before the hub moved. Fix (node lane): the comparison gated like the pre-ghostdag rule (block_version_of under header_signals_active), known-failed test first; main decides between 0.3.17 carrying it and a 0.3.16.1 node hotfix. Checklist line from here: the canary's target is a FRESH node joining the LIVE object's chain, which carries signal headers once a window is open; the pre-cut harness (node-compat.mjs) gets a chain with signal-version headers under an open window.
## 2e. The rebuild on the fix (7 October 2026, 02:3x to 02:4x Z)
The node lane's IBD-guard fix (ca3-v4-0317-fix 90aaf38e) merged into release-0.3.17-node as 12153428. Every binary rebuilt on it:
| piece | commit string | sha256 (first 8) | bytes |
|---|---|---|---|
| Linux igneumd, box native (glibc 2.39, the fleet's canary sha) | 12153428 | 5a4a0d68 | 57,347,232 |
| Windows igneumd.exe (cross, box) | 12153428 | 1e0b49ef | 52,367,360 |
| Windows igneum-miner.exe | 12153428 | 7ca9ca01 | |
| Mac igneumd (Apple Silicon) | 12153428 | 7416d4a5 | 47,837,248 |
| Igneum-Miner-0.3.17.dmg (packaged object = the live sixteen-field object, prover pair aboard) | | 916248b8 | 44,244,489 |
| HiveOS igneumd (zig, glibc 2.31) | 12153428 | 378340e4 | 56,022,096 |
| igneum-hive-0.3.17.tar.gz (2.31 node pair + the box's CUDA/OpenCL workers) | | 04fb7d45 | 27,193,392 |
Suite on 12153428: green (two load flakes pass alone), `cargo check -p kaspa-testing-integration --tests` rc 0. Digests read directly from the rebuilt node: sixteen-field eada4bda (igneum_getNodeInfo powEngine igneum-pow), thirteen-field b18ed271, no file c562d70e: unchanged from 0.3.16.
Windows inputs pushed 02:37Z (igneumd.exe 1e0b49ef, workers d7a413c7, 6f4bb57f, 0d68d06e, the Linux prover pair), pin b5a16f5b on release-0.3.17, windows.yml run 37562948420 dispatched 02:38Z.
HiveOS 2.31 smoke in an ubuntu:20.04 container on igneum-build-1 (ldd 2.31): igneumd, igneum-miner, igneum-worker-cuda and igneum-worker-opencl all load and answer. Found on the way: GNU tar materialises the Mac's extended headers as `._` files beside every file in the package (the live 0.3.16 tar has twelve of them; harmless, HiveOS ran it). Fixed in make-hive-package.sh (COPYFILE_DISABLE, no Mac metadata, c5187a70); the 0.3.17 tar extracts clean (10 files, 0 `._`).
The scratch staging copy (`r0317/dlsite-stage`) had the superseded DMG d25ab372 and installer d1065ac0 removed; the manifest is re-written there once the Windows run's installer is fetched. The live downloads folder holds no 0.3.17 file until the deploy step, which waits on the fleet's second canary line.
## 2f. Ready at the line (7 October 2026, 02:5x Z)
- Windows: run 37562948420 green (engine, window host, payload, installer, smoke run), Igneum-Miner-Setup-0.3.17.exe 93e29580, 62,814,273 bytes, fetched into the scratch copy. The scratch manifests (token folder and public) read: version 0.3.17, mac 916248b8, windows 93e29580, consensus.override = the live sixteen-field object (floor 831,600, window 86,400), HiveOS alias held at 0.3.16 until its own row.
- The deploy runbook is `scratchpad/r0317/deploy.sh` (step_preflight, step_manifest = the live folder written and deployed in one step with the same inputs, step_update_now for d937c69d, ae432dc7 and 1ccfe586, step_hive, step_readback). It runs only on the fleet's second canary line.
- Hands and seed: the build-server lane builds the seed's 2.35 pair and the hands' native pair on 12153428 and restarts them LAST on my line (observer, node 1, then the seed), each read back by commit string, digest and powEngine.
- The Discord card is dry-run at `tools/community/out/release_0.3.17.json` (four reader-facing change lines, no em dash); it posts only once the HiveOS row is live, since the card carries every platform.
- The 0.3.16.1 fallback (main's rule: if the retry fails off the IBD-guard fix) is prepared and not pushed: scratch worktree `r0317/fallback-03161-node`, branch release-0.3.16.1-node at b3c228fa = f1ea7a38 plus a hand-port of 90aaf38e (the helper gates on class_signal_active() alone, since f1ea7a38 has no ladder; the ladder assert dropped from the test). `cargo check` on consensus-core, consensus and p2p-flows clean, the unit test green. The node lane confirms it is the same adaptation it made for its 0318 tree (f2b25fdf).
- The node lane's two-daemon window test (10226194 on the fix branch, separable, touches only testing/integration) is NOT in 0.3.17's pin; run on the box against 12153428 plus the commit (worktree vendor/igneum-node-twodaemon, so the path igneum-pow resolves): `a_fresh_node_joins_a_chain_of_signalling_headers_and_no_window_refuses_them ... ok`, 11.46 s, 03:01Z. It rides the next tree.
- Pairing rule (the build-server lane, 02:5xZ): a fork build takes igneum-pow by path from the igneum worktree it sits in; a fork worktree under a master checkout fails in kaspa-pow (`chain_program_shadow` missing). 0.3.17's artefacts paired with release-0.3.17 at a73e400c to 6f8d7a7e, whose igneum-pow is unchanged since 6b30e855. The tool logs the pairing from the next master.
## 3. Owed
- exec-sync-0313's consensus verification behind a feature off for Windows (the proving lane), then its two commits.
- finality-pause-node rebased onto the 0.3.17 node (the finality lane).
- decimals (B10 gate), vote-weigh and tail-emission docs: the next cut.
- The fork gate's fast-time line is GREEN on aa0182aa (gate on: a 27.5 percent key's private chain 183 DAA deep refused, 268 refusal lines, the honest node kept its chain; gate off, the known-failed case reorgs), one commit past the 9a44fcb8 this tree carries; the switch is at never on every network, so 9a44fcb8 stands for 0.3.17 (the chains were building) and aa0182aa rides 0.3.18. vote-or-burn and the signing bonus: the replay gate FAILED on 9a44fcb8 (this tree) and PASSES on 812c3ac2 (the silence read as a function of the block's own past), which rides 0.3.18 with aa0182aa; both switches are at never here, so a devnet node behaves the same; the signing bonus is settable at the testnet genesis on the project lead's word, the devnet not a candidate until the longer-span replay passes.
- The canary in the full form before publish (a new node mining beside an old one with a poisoned peer, ten minutes, a mid-window re-sync), the staged-in-scratch rule, every new switch at never.
## 4. For the 0.3.18 tree (not in 0.3.17)
- miner-ui-4 past afc331bd: 74c12665 (the merge check: a node whose blocks never merge reads "behind" and the miner is held; src/merge.rs, observer poll, one UI line; 172 box + 42 UI tests green). The UI lane's word, 03:2xZ.
- ca3-v4-0317-fix 10226194: the two-daemon window test (green on 12153428 plus the commit, 03:01Z).
- The node lane's ca3-v4-0318 tree (98c78d9a, f2b25fdf).
- The build-server lane's pairing log line in build-remote.sh (which igneum-pow a fork build paired with).
## 5. The 0.3.18 cut after the hotfix (main's order, 7 October 2026, 03:5xZ)
Once "0.3.17 live" (the hotfix) is read back: rebuild release-0.3.18-node on 12153428 plus the node lane's 4f4bb9c9 (ca3-v4-0317-fix: the idle-peer drop counts headers, proof, trusted data and UTXO chunks as deliveries and never fires during a sync; the class-signal missing-history warning once per epoch per process; kaspa-p2p-lib 20, kaspa-p2p-flows 35, class_signal 1 green on the box; the same patch onto ca3-v4-0318 by 05:15 UK), every platform, digests unchanged, suites, the integration check, the window test again on the merged tip; merge the UI lane's list (afc331bd in, then 74c12665, 89501ce8, b7d33e8e, the doc-only d29a8663), ember-tune-0317 c9b31edc; re-stage 0.3.18 in scratch; the fleet's full canary on its pods with the headers-proof fresh join decisive; publish 0.3.18 on its line, whenever that falls, morning included. Ember's PC 1 window ("PC 1 on 0.3.18") and the UI lane's Mac check follow its update-nows.
Tips (node lane, 03:5xZ): release tree 4f4bb9c9 on ca3-v4-0317-fix (cherry-pick alone onto 12153428); feature tree a3fe9f67 on ca3-v4-0318 (on e1197983, a lockfile-only commit). A follow-up on both is due (the ping flow's `syncing` flag IBD-only; "sink at genesis" broke the peer-drop gate's `--case still` and is redundant for the canary): take it WITH 4f4bb9c9. The hotfix canary (0.3.17, node d712b498) started 03:58Z on c17-1; its clock in UTC: headers through about 04:17, synced 04:40, reads 04:50, cases 05:25.
**Final node inputs (main, 04:0xZ): the 0.3.18 node tree is ca3-v4-0318 at 6e4ace3f** (a3fe9f67 the fix, b21659e4 the IBD-only follow-up, 6e4ace3f igneum_getNodeInfo's `blockrate` object for the merge check; exec-only, digest untouched; all suites green). The release-branch form (4f4bb9c9 + ddda7d52 on 12153428) is the equivalent and not used. After "0.3.17 live": release-0.3.18-node = 6e4ace3f, every platform rebuilt, then the canary.
**Dated fault for this tree (the node lane, 05:05Z, found by the Mac's headers-proof join gate, not by a canary):** when the devnet's pruning point first leaves genesis (chain DAA 185,799, about eight hours from 05:00Z at 1 bps), every fresh join by headers proof fails for the next 2,644 DAA (about 45 minutes) with "IBD with headers proof ... unsuccessful (DAA window data has only N entries)": upstream's sampled-window walk (661 samples at rate 4) runs out at a gap in the trusted blocks instead of reaching genesis. Nodes already on the chain are untouched. Fix: joiner-side on the trusted path only, digest-neutral (the walk accepts a window that runs out at origin when the block is younger than the span), coming on ca3-v4-0318 (hash to follow); 0.3.18 carries it. Any fresh-join canary timed inside that 45-minute window FAILS from this, not from the idle-peer fix: time 0.3.18's canary outside it (before about 12:50Z or after about 13:40Z, to be firmed from the live DAA).
**Main's rule for the 0.3.18 cut (05:1xZ):** it carries the node lane's joiner-side fix for the pruning-point fault (on the 0.3.18 tree the N6 guard would otherwise ban the syncer for ten minutes per retry), and **0.3.18 is live on every node before 13:30 UK (12:30Z)**, ahead of the devnet's first pruning (DAA 185,799, about 14:00 UK). If the full canary cannot close by then, publish 1 goes on the clean form plus the node lane's harness for that case, and the cases follow as confirmation, as with the hotfix. Every new chain, the testnet genesis included, meets the same fault once at its first pruning: the testnet go checklist (docs/plans/testnet-go.md) gets the line.
**Correction (the node lane, 05:1xZ): the devnet does NOT meet the young-window fault, and the 185,799 clock was wrong** (the canary's 155,700 headers were taken for the chain's DAA; live DAA was 250,809 at 05:00Z, the hub's 39th pruning move was at 03:46Z and the 03:50Z fresh join synced clean). Pruning samples sit at multiples of the finality depth (pruning.rs is_pruning_sample); the devnet's finality depth is 43,200 blocks, so its first non-genesis pruning point was at blue score 43,200, sixteen window spans past genesis, and the walk from any devnet pruning point fills its window. The fault needs a pruning point within 2,644 DAA of genesis, which only a chain with finality depth under 2,644 can have: the fast-time profile (120) where the gate found it. The testnet at the compiled numbers never meets it either. So: no fresh-join window to avoid, the 45-minute rule is dropped, main's "before 13:30 UK" rule for 0.3.18 no longer rests on this fault (main to confirm), and the testnet go line is softened to a note. The fix still rides 0.3.18 (joiner-side, digest-neutral, relaxes the rule only when the walk ended within one sample of genesis, impossible on the devnet; unit test green on the box; hash to follow).
**Main, 05:2xZ: the 13:30 UK deadline is withdrawn.** 0.3.18 ships on its normal gates with the full canary, no shortcut; the young-window fix stays in the tree as a harness-profile correctness fix. ca3-v4-0318 tip 06:04 UK: e3798a17 (69647bd6 per-path finality timers plus the join bench, test-only plus accounting; e3798a17 the young-window fix, unit test green). The fresh-join cost fix is next on the branch with its own hash and before/after numbers; its "before" figure comes from the 0.3.18 canary's IBD-end line (the new "finality time: bodies ..., virtual ..., weight tables ..., signatures ..., persists ..." line), which I relay verbatim (route (b); a read-only joiner from the box, route (a), is main's word).
ca3-v4-0318 tip 06:06 UK: **aa49613f** (on e3798a17; carries the IBD-end "finality time" line; p2p-flows and kaspad check clean on the box). The cut takes aa49613f or the later tip the node lane names; the canary's fresh-join IBD-end line goes to the node lane verbatim.
## 6. The 0.3.18 inputs as of 05:2xZ (main's list)
- Node: the ca3-v4-0318 tip the node lane names at the cut; now **8220c944** (06:10 UK; on aa49613f): a block's votes BLS-checked across cores before the finality lock, exact by the two-path test. Box, 500-checkpoint chain of 22,221 blocks: join 218 s to 104 s, finality 130 s to 30 s, signatures 97 s to 9 s. Suites at 8220c944 on igneum-build-1: kaspa-consensus lib 110 passed, 0 failed, 3 ignored (the two new finality tests inside); kaspa-p2p-flows 37 passed; kaspad check clean at aa49613f with nothing in kaspad changed since. The real split of the canary's 8 s per checkpoint comes from the IBD-end timers line of the 0.3.18 canary's fresh join, relayed verbatim to the node lane.
- App: miner-ui-4 afc331bd (in), 74c12665, 89501ce8, b7d33e8e, **bac43ac4** (main's row from PC 1's 04:51Z relaunch: the card watchdog judges a miner only while the node is synced; a silence fault from the sync is released when the node syncs and the card starts again with an Activity line; a worker silent from its start is faulted at 60 s; one Activity line per card at its first start; known-failed test from the PC 1 row in release-0.3.17.md; 173 box + 42 UI tests green) plus the doc commits (d29a8663, 0e2e8a31); ember-tune bb37c993 (c9b31edc is job-script only); exec-app 032f03c; relay 28c028b.
- Gates: the same as every cut; the full canary with the fresh join decisive; publish on its line.
## 7. The node tree is a merge (05:5xZ)
ca3-v4-0318 (8220c944) does not contain release-0.3.18-node 12153428: their merge-base is 9a44fcb8 (ca3-v4-0316), so the branch lacks the ladder 1591ee1d, exec-sync 40fd8d8c, finality d8d1de1a, tail-emission 996b592f, the kaspad igneum-pow default 75467244, cc78a21f and 90aaf38e (it carries f2b25fdf, its own adaptation of the canary fix). The 0.3.18 node is the merge of 8220c944 (or later) into 12153428. A no-commit test merge conflicts in six files (consensus/core/src/igneum.rs, pre_ghostdag_validation.rs, ibd/flow.rs, v10/blockrelay/flow.rs, the two integration test files): the node lane resolves it, pushes release-0.3.18-node to the mirror, runs the suites and the integration check, and names the tip. The branch stays at 12153428 until then.
PC 1's 0.3.17 relaunch (section 7 of release-0.3.17.md) adds a node row for this tree: the template RPC blocks under the finality catch-up after a restart on a synced node (every getBlockTemplate timed out at 5 s for 90 s); the app row (bac43ac4) needs its rule changed to "template timeouts while the node reports synced are not a card fault".
## 8. The app tree assembled (05:4xZ)
release-0.3.18 at da5c40ba: miner-ui-4 afc331bd (in), 74c12665, 89501ce8, b7d33e8e, d29a8663, 0e2e8a31, bac43ac4, 5d9c55a2 (the UI lane's rule from PC 1's read-back: "template fetch timed out" is the miner's heartbeat, the card reads "waiting for the node to answer block templates", judged again from its last line; 173 box + 42 UI tests on its branch), ember-tune c9b31edc (job script; bb37c993 in), exec-app 032f03c (in), relay 28c028b (in). One merge fix by the shipper: 74c12665's merge view called the three-argument `live::fetch`; this tree's is the four-argument form (e5336dc7 gave it the engine's Shared), so engine.rs:3200 passes `&shared` as server.rs does (da5c40ba). App tests on the Mac: 179 + 28 + 8 passed. The node pin, binaries, inputs and staging wait on the node lane's merged release-0.3.18-node tip.
UI lane on the merge fix (05:5xZ): correct; the four-argument form wins. Note: Ember's fetch clamps the window to 300 s, so the merge view sees 300 s rather than 600; the check needs a block of ours inside 120 s, so it holds. The merge view, TemplateTimeout and the synced gate are present on origin/release-0.3.18.
**The node merge, the node lane, 06:0xZ (not yet pushed):** 8220c944 into 12153428 resolved: five conflicts taken as the release side (the ladder's header_signals_active reading, the test fixes), the sixth the relay flow (the pipelined handle_block carries the 0.3.16 IgneumProofMissing retry arm ahead of the N6 arm); one copy of header_version_acceptable kept. Box at the merged tree: kaspad check clean, integration check clean, consensus-core 123, p2p-flows 37, p2p-lib 20. A real finding: kaspa-consensus failed 1 of 112 then 4 of 112 under a box load of 178, every panic WrongBlockVersion(1026 or 32770, 2) in a plain-block test: a test-order race both parents carry (the signal and ladder tables are process-wide statics; two tests install and restore them; a block-building test inside that window reads the installed version; the two load flakes of the 12153428 suite were this). Fix: an install lock the two installers hold; the suite runs twice on the quiet box; if green the merge is pushed as release-0.3.18-node with the lock commit on top (tip about 06:20Z); if a reader still races, the two installing tests run serially as a second pass, said before the push.
**Suite rule change with the merged node (the node lane, 06:1xZ):** the install lock did not close the row (2 of 112 still WrongBlockVersion(32770, 2) on the quiet box), so the two installing tests (block_template_uses_current_block_version, cheap_checks_run_before_the_pow_engine) move out of the lib unit-test binary into their own target consensus/tests/igneum_installed_signals. From the 0.3.18 node on, the suite runs `cargo test -p kaspa-consensus` without `--lib` (all targets) so both run; a `--lib` run silently skips them. Push with the merged tree once the lib suite is green twice and the new target once, about 06:35Z.
## 9. The node tree: release-0.3.18-node = 1cf43254 (the node lane, 06:09Z)
The merge of ca3-v4-0318 (8220c944) into 12153428 with the six resolutions, plus the race fix in the same commit (the two installing tests moved to consensus/tests/igneum_installed_signals.rs and igneum_order_tests.rs; INSTALL_TEST_LOCK for installers sharing a binary). Box at 1cf43254: kaspad check clean, integration check clean, kaspa-consensus lib 110 twice on the quiet box, the two new targets 1 each, consensus-core 123, p2p-flows 37, p2p-lib 20. A second merge may follow within the half hour (PC 1's two node rules: template answers under catch-up, synced reads behind), else the cut is at 1cf43254. The shipper's box and Mac chains started on 1cf43254 at 06:1xZ (scratch only; the pin, inputs and staging wait for the node lane's cut word).
**Main, 06:1xZ:** the chain runs on 1cf43254 now. The node lane's two PC 1 rules (getBlockTemplate answers inside its timeout during the finality catch-up; isSynced false while the catch-up is behind the sink) land as a second merge with a node rebuild only if they reach the release branch before the inputs push; otherwise they are 0.3.19, said here. Full canary with the fresh join decisive, the IBD-end timers line to the node lane, publish on its line.
## 10. Builds on 1cf43254 (06:10Z on)
| piece | commit string | sha256 (first 8) | bytes | note |
|---|---|---|---|---|
| Linux igneumd, box native (the fleet's canary sha) | 1cf43254 | 96858a88 | 57,408,800 | 06:11Z; the fleet has it with GO for the full form |
| Linux igneum-miner, box native | | fac45489 | 10,213,112 | |
| Windows igneumd.exe (cross) | 1cf43254 | c29ee9e5 | 52,340,224 | 06:12Z |
| Windows igneum-miner.exe | | 7ca9ca01 | 11,219,456 | unchanged from the 12153428 build |
| Mac igneumd | 1cf43254 | 25b464b4 | 47,888,704 | |
| Igneum-Miner-0.3.18.dmg | | a3822c4d | 44,248,348 | packaged object = the live sixteen; prover pair aboard; version 0.3.18 |
Digests on the Mac binary: sixteen eada4bda (igneum_getNodeInfo: powEngine igneum-pow, blockrate {bps 1, finalityDepth 43200, ghostdagK 18, mergeDepth 3600, pruningDepth 108000}), thirteen b18ed271, no file c562d70e (re-read on free ports, 06:15Z). All three unchanged from 0.3.16/0.3.17.
(suite, integration check, seed, hive, inputs, Windows run, staging, canary: rows as they land)
## 11. The cut: release-0.3.18-node = ae17ad00 (the node lane's word, 06:14Z)
1cf43254 plus the second merge of ca3-v4-0318 (e3a00bf0: 4f7c56a0 PC 1's two node rules, getBlockTemplate answers inside its timeout during the finality catch-up and isSynced reads false while the catch-up is behind; e3a00bf0 the installing-tests move on that branch too). Box at ae17ad00: `cargo test -p kaspa-consensus` 111 in the lib plus the two moved targets 1 each, p2p-flows 37, `cargo check -p kaspad -p kaspa-rpc-service -p kaspa-testing-integration --tests` clean. One hand fix in the merge: the W7 leaves block in template_section rides the template snapshot like the other items (newest 32, emitted last under the leave_active gate). The 1cf43254 chain was stopped at its seed step (its suite: see r0318/ae17/box-1cf43254.log) and both chains restarted on ae17ad00 at 06:2xZ. The canary form gains two reads for PC 1's rules: on a restart with a kept datadir, every getBlockTemplate inside 5 s through the catch-up (no "template fetch timed out" lines), and isSynced false while the catch-up holds the finality state, true after.
**Canary on ae17ad00 started 06:16:47Z** (c18-1, RTX 3070 pod, wiped datadir, node b06c1a97, 57,447,712 bytes, miner fac45489; the 1cf43254 early run stopped by pid; it had read vline 1cf43254, digest eada4bda, igneum_getNodeInfo with blockrate, IBD from 4 peers). Clock: headers through about 06:52Z, synced 07:12Z, mining reads to 07:22Z, the restart reads (isSynced false then true every 10 s; max template_ms, timeout count, template count) to about 07:40Z, then the relay and poison cases with c18-1 as the target.
## 12. Builds on ae17ad00 (06:15Z on)
| piece | commit string | sha256 (first 8) | bytes | note |
|---|---|---|---|---|
| Linux igneumd, box native (the canary sha) | ae17ad00 | b06c1a97 | 57,447,712 | 06:16Z; the canary runs on it from 06:16:47Z |
| Linux igneum-miner | | fac45489 | 10,213,112 | unchanged from 1cf43254 |
| Windows igneumd.exe (cross) | ae17ad00 | e4979672 | 52,391,424 | pow link 9 |
| Mac igneumd | ae17ad00 | a56469d7 | 47,905,856 | |
| Igneum-Miner-0.3.18.dmg | | f617b63b | 44,241,681 | packaged object = the live sixteen; prover pair aboard |
| Windows inputs | ae17ad00 | | | pushed 06:19Z (workers d7a413c7 etc, the Linux prover pair), pin 1c8fb76a on release-0.3.18, windows.yml run 37580969266 dispatched 06:20Z |
Digests on the Mac binary: sixteen eada4bda (igneum_getNodeInfo powEngine igneum-pow, blockrate as before), thirteen b18ed271, no file c562d70e (free ports). All three unchanged. Box suite on ae17ad00 rc 0 (kaspa-consensus without --lib, so the two moved targets ran). Identity grep on the payload UI clean. The 1cf43254 builds (96858a88, c29ee9e5, 25b464b4, DMG a3822c4d) are superseded and not staged.
| generic Linux igneumd (class seed, glibc 2.35) | ae17ad00 | 99809615 | 56,141,136 | pow link 6; the seed's own build comes from the build-server lane |
|---|---|---|---|---|
| HiveOS igneumd (class hive, glibc 2.31) | ae17ad00 | a2e734d1 | 56,141,776 | pow link 6 |
| igneum-hive-0.3.18.tar.gz | | 21ea06eb | 27,234,062 | 2.31 node pair + the box's workers 193ec36f/7a35ff2a; no `._` entries; ubuntu:20.04 container smoke on the box: all four load (06:23Z) |
| hands igneumd (box native, the build-server lane, pairs with igneum 1c8fb76a) | ae17ad00 | 17209647 | 57,448,096 | held for the line; OUT_DIR apart from b06c1a97 |
| seed igneumd (class seed, the build-server lane) | ae17ad00 | fa9c2f12 | 56,139,920 | GLIBC_2.34; held for the line |
| hands / seed igneum-miner | | 9adcb707 / 2ebaee58 | 10,213,112 / 10,234,128 | |
Integration check on ae17ad00 rc 0 (06:18Z). The hands and seed go last on the shipper's line after the miners (observer, node 1, seed), each read back by commit string, digest and igneum_getNodeInfo powEngine with its blockrate object.
## 13. Staged at the line (06:29Z)
Windows run 37580969266 green (06:27Z): Igneum-Miner-Setup-0.3.18.exe 2fcaf093, 62,822,510 bytes. Scratch copy r0318/dlsite-stage (fresh from the live folder): manifests at 0.3.18, mac f617b63b, windows 2fcaf093, consensus.override = the live sixteen-field object, HiveOS alias held at 0.3.17 until its own row; the live folder holds no 0.3.18 file. Runbook r0318/deploy.sh (preflight passes; step_manifest, step_update_now, step_hive with the 0.3.18 tar, step_readback with igneum_getNodeInfo). Discord card dry-run (four reader-facing lines, no em dash). Everything waits on the canary's line.
## 14. The canary (c18-1, ae17ad00 / b06c1a97)
- 06:34:54Z: headers 47 percent (59,425) in one IBD session since 06:17:32Z, 17 minutes in and 5 past the old guard mark; SendPingsFlow 0, idle-drop lines 0, "completed with error" 0; the class-signal warning is ONE line (49,844 on the hotfix). Headers through about 06:55Z on the 3070, synced about 07:15Z.
## 15. HOLD on ae17ad00: the exec RPC panic (the node lane, 06:4xZ)
The Mac's headers-proof join gate killed its joining node mid-IBD on the canary-fix binaries: a panic at igneum/exec/src/rpc.rs ("index out of bounds: the len is 0 but the index is 0") on a tokio worker, and the node's panic hook exits the process. The site: eth_getBlockByNumber reading state.records[n] after resolve_block mapped "latest" to tip_number() = 0 on an exec state with no records (the follower still "waiting for consensus to sync"); a local client polled 127.0.0.1:26790 during the IBD and the node died at 66 percent of the headers stage. Every records index in that file is unchecked on both trees (0.3.17's 5899f603: rpc.rs lines 474, 550, 591 to 629; ae17ad00: 575, 651, 692 to 730), so any fresh or restarting node with the exec RPC bound and a client asking eth_getBlockByNumber("latest") before the follower has a record dies. The shipped app itself calls only eth_blockNumber and igneum_* methods (none index records by block number), so the live 0.3.17 risk is a wallet or third-party client on a fresh node mid-IBD; the fix (every index bounds-checked; "latest" on an empty state answers null; a test on an empty state) goes onto ca3-v4-0318 and into release-0.3.18-node as a third merge. The pin stays at ae17ad00 on origin and nothing is staged further until the third tip; the ae17ad00 canary runs on (nothing attached to its eth_ port) as an early read of the other rows.
**The 0.3.17.1 question settled (07:0xZ, main's rule: a hotfix only if the shipped app sends the panic class):** no. The dead harness node's site is eth_getBlockByNumber; the shipped app sends eth_blockNumber and eleven igneum_* methods only (grep of the whole 0.3.17 app tree); its two records-indexing polls (igneum_getAssignedShards, igneum_getProofRecords) run only after the node reads synced, when a devnet follower already holds records from the packaged restart block; the two it sends while unsynced index nothing. The wallet app sends eth_getBlockByNumber to its own node. Ledger N7 carries the method map. The public testnet filter now blocks ten eth_ and six igneum_ records-indexing methods (seed1, verified; tree copy committed).
## 16. The cut moves to release-0.3.18-node = e69e8a39 (the node lane, 06:53Z; the pin released)
ae17ad00 plus the third merge of ca3-v4-0318 (500ddd66): c6a62e00 (every chain-block read in the exec JSON-RPC bounds-checked; an exec state with no record answers an error instead of indexing; test on the empty state) and 500ddd66 (GetBlockTemplate stage timers: a debug line per request, an info line once per 10 s past 500 ms; the epoch-seed walk and the live sink tally memoised). One hand resolution: the payout parameter on igneum_getProofRecords beside the bounds-checked read. Box at e69e8a39: igneum-exec 26, kaspa-consensus 111 plus the two moved targets, p2p-flows 37, kaspad / rpc-service / testing-integration checks clean. Both chains restarted on it 06:55Z (the ae17ad00 logs kept in r0318/ae17/). The canary form gains two reads: an eth_ client polling the joining node's exec port through the IBD (the node survives), and the first slow "GetBlockTemplate N ms" line on the pool's loaded node after the publish.
## 17. Builds on e69e8a39, the pin (06:55Z on)
| piece | commit string | sha256 (first 8) | bytes | note |
|---|---|---|---|---|
| Linux igneumd, box native (the canary sha) | e69e8a39 | 252c8dad | 57,461,600 | the form started from the wipe 07:02:54Z |
| Linux igneum-miner | | fac45489 | 10,213,112 | unchanged |
| Windows igneumd.exe (cross) | e69e8a39 | c1a42970 | 52,379,136 | pow link 9; inputs pushed 07:03Z, pin 0911b00f, windows.yml run 37585128393 dispatched 07:04Z |
| Mac igneumd | e69e8a39 | 066807ae | 47,921,280 | |
| Igneum-Miner-0.3.18.dmg | | cb30080e | 44,247,374 | packaged object = the live sixteen; prover pair aboard; version 0.3.18 |
| generic Linux igneumd (class seed, 2.35) | e69e8a39 | 7d67cb51 | 56,152,400 | |
| HiveOS igneumd (class hive, 2.31) | e69e8a39 | 6293b443 | 56,153,104 | |
| igneum-hive-0.3.18.tar.gz | | 72699158 | 27,236,213 | 20.04 container smoke clean, no `._` entries (07:08Z) |
| hands igneumd / seed igneumd (the build-server lane, pairs with igneum e9ccb3a1) | e69e8a39 | 37610a77 / 63cd49be | 57,461,984 / 56,152,720 | held for the line |
Digests on the Mac binary (explicit arguments; a zsh loop had not split them on the first pass): sixteen eada4bda, thirteen b18ed271, no file c562d70e. igneum_getNodeInfo: powEngine igneum-pow, blockrate {bps 1, finalityDepth 43200, ghostdagK 18, mergeDepth 3600, pruningDepth 108000}. The N7 class on the Mac binary: eth_getBlockByNumber ["latest", false] on an empty exec state answers {"result": null} and the node lives (the node lane: null is Ethereum's shape for a block that is not there; every other read on an empty or short state answers an error object). The branch tip moved to 2a014bf1 (test-only: all 46 exec RPC methods through the real dispatcher on empty and one-record states, no panic); the pin stays at e69e8a39, whose shipped bytes it equals.
The chain's own log was lost mid-run (the build-server lane removed scratchpad/r0317 and r0318 whole while dropping its superseded pairs; rule now: no lane removes a scratch directory it did not create; this lane's scratch is r0318-ship); the suite and the integration check are re-run on the box for the record (r0318-ship/suite-e69e8a39.log).
**Suite on e69e8a39 (the box, 07:08Z, re-run for the record):** igneum-exec 26, kaspa-pow 19, kaspa-consensus 111 (lib) + the two moved targets 1 each, consensus-core 123 + 7, kaspa-p2p-flows 37, igneum-miner 16; 0 failed; rc 0. `cargo check -p kaspa-testing-integration --tests` rc 0. Together with the node lane's own runs at e69e8a39 and 2a014bf1 this is the gate's suite line.
## 18. Staged at the line (07:1xZ)
Windows run 37585128393 green (07:10Z): Igneum-Miner-Setup-0.3.18.exe 669d5676, 62832534 bytes. Scratch copy r0318-ship/dlsite-stage (fresh from the live folder): manifests at 0.3.18, mac cb30080e, windows 669d5676, consensus.override = the live sixteen-field object, HiveOS alias held at 0.3.17 until its own row; the live folder holds no 0.3.18 file. Runbook r0318-ship/deploy.sh (preflight passes). Discord card dry-run (five reader-facing lines, no em dash). Everything waits on the canary's cases line.
**For 0.3.19 (not the pin), the node lane 07:1xZ:** ca3-v4-0318 = fd7de1b4 (the live class-signal tally refreshed off the request path; kaspad check clean, consensus 109). e69e8a39 memoises the tally for 10 s per sink, so the pool's canary should show the "pow epoch" stage near zero on nine requests in ten and the full walk (1 to 1.6 s) on the tenth; if the stage line shows it high on most requests the reading is wrong. fd7de1b4 is the first 0.3.19 node commit (main, 07:1xZ: no four-part numbers, the parser refuses them); 0.3.18 ships as pinned at e69e8a39 on its canary's line.
## 19. The canary on the pin (c18-1, e69e8a39 / 252c8dad, from the wipe 07:02:54Z)
- 07:21Z: 17 minutes into one IBD session, headers 37 percent (48,160 at 07:16:45Z), past the old guard mark; SendPingsFlow 0, idle-drop 0, IBD errors 0, the class-signal warning 1 line, node alive; the eth_getBlockByNumber poller at 208 polls, 0 error answers, all null. Headers through about 07:45Z, synced about 08:05Z.
## 20. PC 1 and PC 2 after the project lead reopened the apps (07:28Z reads, read-only)
- PC 2 (1ccfe586): back online; its app came back on 0.3.16 and applied 0.3.17 at reopen (update: version 0.3.17, current, updated_from 0.3.16); node 2.1.0 synced at DAA 262,124 with 5 peers; the RTX 5090 mining at 100 MH/s (pid 19620, 11 accepted in its first minutes), no fault lines. The update-now takes it to 0.3.18 directly.
- PC 1 (ae432dc7): app 0.3.17 (reopened, uptime 531 s at the read), node binary 5899f603 (3 marks, pow link 9), digest eada4bda; but the node is in a restart loop: state "restarting", starts 16, igneumd not running at the read; every card "waiting for the node to sync" (mining "waiting", not faulted this time). Cause being read from the engine's node lines and the node logs (read-only job). Rule for the publish: PC 1 and PC 2 take their update-nows first and their workers are read back at their rates as the per-box table's first line.
**Latency ladder rung 3 re-measured (the node lane, 07:46Z, in the shipper's box window under the measure hold):** reps 88, cold alone 9.04 ms, cold with the SMT sibling loaded 10.85 ms, averages of 50 at 5.95 and 10.11 ms; over the 10 ms gate by 0.85, so rung 3 stays inadmissible and the testnet genesis freezes with 88 false. Rung 2 as the control on the same run: 9.25 ms cold loaded, admissible.
**PC 1's restart loop IS the N7 class on the shipped kit (07:5xZ, the engine log app-20733-072141.log):** every node start ends 2 to 5 s later with "panicked at igneum/exec/src/rpc.rs:591:41: index out of bounds: the len is 0 but the index is 0" (eth_getBlockByNumber's records[n] in 0.3.17's tree), the app restarts it every 40 s (starts 16 at 07:28Z), the miners are stopped each time. CORRECTED 08:0xZ: the client is the shipped engine itself: `latest_block_time` (update.rs:46) POSTs eth_getBlockByNumber ["latest", false] through curl to the exec port every 9 s as the app's clock sample once the node reports blocks > 0 and peers > 0 (engine.rs:3752); export-pack's lines were its own failure after the node died (igneum-miner's export-pack speaks gRPC only). So every 0.3.17 node with the app attached dies within 9 s of blocks arriving whenever its exec follower has no record: PC 1's loop, and any fresh install's IBD (the hotfix canary had no app attached). Proving-off would change nothing and was not run. c6a62e00 (0.3.18) ends it; the 0.3.18 canary's eth_ poller is the app's own call and the node lives on it. PC 1 takes the 0.3.18 update-now first on the publish line.
**Main's publish rule (08:0xZ):** 0.3.18 publishes on the canary's synced line plus the poller summary (the decisive reads for the N7 fault, the app's own 9-second call surviving the whole IBD), not the cases line. Order: PC 1's update-now first (it is crash-looping), then PC 2, the Mac, the fleet one box at a time under the lock rule, the hands and the seed by their owner. The restart reads and the cases follow on the same pod as confirmation; a FAIL there is a rebuild, not a rollback of this fix. Then "0.3.18 live" with the per-box table and the card.
## 21. HOLD on e69e8a39 (the fleet, 08:2xZ): the block stage cannot complete against any peer
The canary passed its headers proof at 08:00Z (one session, zero guard lines, the app's eth_ poll alive at 918 polls) and froze at 2,728 blocks: every IBD attempt (110 by 08:2xZ, every peer) ends "block f52e64f4... carries 1 proof records whose proofs this peer did not deliver in 20 s". The node lane's reading: exec-sync's daemon installs the proof oracle whatever proving_consensus_verify_daa says (daemon.rs, after the keys branch), so the IBD flow (flow.rs, the 0.3.16 "carried proofs come from the syncer" rule) demands the proof bytes of every record-carrying block before queuing it while the serve flow answers only from a peer's bounded recent pool; block 2,728 is months old, so a fresh node of this tree can join NO network, and a synced one would verify every relayed record natively and refuse blocks its 0.3.17 peers accept (a split in waiting). Fix, joiner-side: the oracle carries its switch (active_from); the body rule, the IBD fetch and the relay retry apply to blocks at or above it only, so at never the node is 0.3.17 on this path. A direct commit on release-0.3.19-node; the pin moves to it; the canary re-runs from the wipe.
## 22. For the 0.3.19 node: ledger N8 (the project lead's word, 08:5x UK)
The execution layer credits merged blocks at the chain block's DAA (the UTXO coinbase pays each its own): d840537b on ca3-v4-0318, field `subsidy_per_block_activation_daa` (u64::MAX = never as compiled on every network; in the digest once set). The devnet takes it at an upgrade height carried by 0.3.19 (a DAA score past the rollout's last box under the lock rule, set in DEVNET_PARAMS at the cut; every node must carry it before that height or its EVM state diverges); the testnet from genesis. The node lane names the value when the 0.3.19 cut line is named (its canary's pass).
**App tree (08:3xZ):** miner-ui-4 taken through 3661dc9e (a2670ace the two site captures; 3661dc9e carries 0e6c1248's exec-record clock gate without its version bumps), on top of the earlier chain; app tests 180 + 28 + 8 green. Ember: c870c383 (job-script only, equals bb37c993 on the app) at the rebuild.
**The exec lane agrees (08:4xZ):** exec-sync-0313 HEAD carries the same shape (ProofOracle::active_from() = proving_consensus_verify_daa, set by the daemon before the oracle is handed out; the IBD pre-fetch gated on it; the body rule and the relay retry already gated); igneum-exec 23 green. The node lane's direct commit on release-0.3.19-node is the one the pin takes; the shapes reconcile at the next merge; the switch must live on the oracle (one source). Exec lane's 0.3.19 tips: node exec-sync-0313 HEAD, app exec-app-0314 731268a1. Its "0.3.19.1" (the finality-horizon skip) is a four-part number and so 0.3.20 material.
## 23. The pin: release-0.3.19-node = dc141409 (the node lane, 08:36Z)
e69e8a39 plus the oracle switch: `ProofOracle::active_from` carries `proving_consensus_verify_daa`; the body rule, the IBD proof fetch and the relay retry apply to blocks at or above it only (one source, on the oracle: igneum/exec/src/proving.rs:1499, read at flow.rs:1024 and body_validation_in_context.rs:47), so at never a node demands no proof of any block; the sink takes the switch from Params through proof_sink (daemon.rs:1018). Box at dc141409: igneum-exec 27, consensus-core 123, kaspa-consensus 111 plus the two moved targets, p2p-flows 37, kaspad and testing-integration checks clean. The exec lane's own commit ac6e32cd on exec-sync-0313 has the same shape and stays as the record (reconciled at the next merge). Both chains restarted on dc141409 at 08:4xZ (scratch r0319-ship); the pairs rebuild by the build-server lane. Canary reads for this tip: the block stage passes 2,728 and every record-carrying block with no "Proofs: asking" line, the IBD-end line with its finality time, the synced line, the app's eth_ poll alive throughout, the restart reads, the cases, and on pool-1 after the publish the first slow "GetBlockTemplate N ms" line. The 0.3.19 cut line is that canary's pass; the node lane then names the devnet DAA for the per-block subsidy height (N8).
**Hands and seed pairs on dc141409 (the build-server lane, 08:4xZ, pairs with igneum 5010536a):** hands igneumd 8be8a5a5 (57,460,512 B), igneum-miner 92740aea; seed-class igneumd 4910352b (56,154,320 B, GLIBC_2.34), igneum-miner 9d7c4b8a; held for the line after the miners.
## 24. Builds on dc141409 (08:43Z on)
| piece | commit string | sha256 (first 8) | bytes | note |
|---|---|---|---|---|
| Linux igneumd, box native (the canary sha) | dc141409 | 3f9aca09 | 57,461,792 | GO sent to the fleet 08:5xZ; the form from the wipe on c18-1 |
| Linux igneum-miner | | fac45489 | 10,213,112 | unchanged |
(cross, suite, integration check, seed, hive, Mac, DMG, inputs, Windows run, staging: rows as they land; the first chain start on this tip at 08:37Z had picked up the hotfix tree's script copy and was stopped at 08:39Z; its one native build, d712b498, is the 0.3.17 binary again and is not used)
## 25. The canary on dc141409 (c18-1, from the wipe 08:51:50Z)
Node 3f9aca09, POLL_EXEC=1 and the exec-status fields; the e69e8a39 node stopped by its kill file (frozen at 2,728 blocks since 08:00Z). The form prints the count of "Proofs: asking" lines and of "carries N proof record" IBD errors beside the IBD-end line. Clock: headers through about 09:27Z, the block stage past 2,728 about 09:35Z, synced about 09:50Z, mining reads to 10:00Z, the restart reads to about 10:15Z, then the relay and poison cases; pool-1's "GetBlockTemplate N ms" line after the publish.
## 26. Owed to the 0.4.0 cut (the launch-pack lane, 08:5xZ; not this tree)
- The signing step (docs/plans/launch-pack.md section 2.4 on master 9b98b837, gate LG-3 in docs/plans/testnet-go.md): implemented only once the certificates exist (Windows EV Authenticode and an Apple Developer organisation account, both under Igneum Labs LTD, enrolled the day the entity exists; the owner approved both 7 October 2026). Until then 0.4.0 ships unsigned and the download page says so. The step keeps ship-app.mjs's list and adds: windows.yml signtool through the CA's cloud KSP (two secrets; unsigned and marked so without them), Inno SignTool=; ship-app.mjs fetch osslsigncode verify (CN Igneum Labs LTD, SHA-256, timestamp; unsigned fails); build-dmg.sh codesign Developer ID with runtime and timestamp, notarytool --wait Accepted, stapler on the app and the DMG, no quarantine strip; publish-manifest.sh signed_by and notarized per platform, --public refused without; ship-app.mjs verify osslsigncode and spctl --assess on the live files; the download page and the Discord card say "Signed by Igneum Labs LTD" only when the manifest does; tools/ci/signed-release-check.sh in pre-push. Keys never on igneum-build-1.
- docs/analysis/income-tiers.md regenerated from tools/launch/income-tiers.json at the 0.4.0 cut, the rows re-measured on class v4 (node tools/launch/income-tiers.mjs; --check in pre-push).
Names to build against (the launch-pack lane): GitHub secrets IGNEUM_CODESIGN_ACCOUNT and IGNEUM_CODESIGN_KEY (TOTP seed or API key by the CA; absent on forks and PRs, so unsigned and marked); notary keychain profile igneum-notary (xcrun notarytool store-credentials, once on the Mac); the codesign identity "Developer ID Application: Igneum Labs LTD (TEAMID)" read from one place (an env or a packaged-config field), the team id dropping in without a code change. Values reach the shipper by relay, never the repository (launch-pack.md section 5, items 3 and 4).
**Branch tip past the pin (the node lane, 09:04Z):** release-0.3.19-node = aea0ca5c on the mirror, on dc141409: the proof archive (ledger N9's second half; every carried record's proof kept under the exec db dir's proofs/ for the pruning window, served to a joiner's IgneumRequestProofRecords past the pool's 600-block window, dropped below the pruning point; no digest field, no behaviour on the devnet at never). Box: igneum-exec 28, p2p-flows 37, kaspad and testing-integration checks clean. **The pin stays at dc141409** (the canary running on it since 08:51Z; a re-pin only on a FAIL); aea0ca5c is 0.3.20's first node commit otherwise.
**Suite coverage note (the node lane, 09:1xZ):** `processes::pruning_proof::igneum_m20_tests::witnesses_are_checked_in_epoch_order_under_their_own_seeds` fails deterministically (MissingEpochSeed(1, _, 109) not matched) on an untouched e69e8a39 worktree with `--features igneum-pow` at any box load; it sits behind `#[cfg(feature = "igneum-pow")]`, so no lane's suite line (`cargo test -p kaspa-consensus` without the feature) compiles it, which is why 111 reads green. Owner: the M20 pruning-proof witness work (the epoch seed table of a proof-synced node). Not a publish blocker by itself (the gated path is what the live binaries run; the red is a test or seed-table reading), but the suite line must either carry the feature or the red must be fixed before "all green" covers it: a row for the 0.3.20 suite rule.
## 27. miner-ui-5 is the 0.3.19 app (the project lead's word at 13:1x UK: "deploy the miner look")
Taken onto release-0.3.19 as five cherry-picks on the miner-ui-4 chain: 2b4775b4 (the ladder's engine half: ladder.json, api/ladder chain facts, the block card), 449fb207 (the ladder on every page, the count-up, the block card, Earnings in IGN), b49db57d (the plan, the captures, the first-share runbook, the Discord rules), 3973ff8b (live-dag.js 2.0.1, the site lane's real-data options), e9fe1106 (u04 and u14 reshot); tip a199160d. Gate: app 197 + 28 + 8, UI 47 (notices 8, tune-line 5, update-card 4, view 30), all green. Next: push, windows.yml, the DMG (after the dc141409 Mac chain), staging; the read-back line adds "app miner-ui-5" and the pack's two shas. If the canary slips past 15:00 UK, main decides whether the app ships alone as 0.3.19 on the 5899f603 pin and the feature node becomes 0.3.20.
## 28. The prover pair (main's order, 09:0xZ): handed over, verified, the CI check in
The fleet's 14 standing provers had proved nothing since the hotfix: their harness (tools/fleet/box-prover.py) exported igneum_exportSegments from the restart block 27276 to the tip on every claim and replayed 133,000 blocks through 72142, where every port (the 6 October floor exporter, master's 6c3cc8b9 core, the kit's 263bf4ce) computes state root 0x8e1bef4b against the nodes' 0x9851d7e2 (hub-1, p1-5090 and PC 2's 0.3.17 node agree: 27276 0xed27bb2d, 72141 0x103190f4, 72142 0x9851d7e2, 72143 0x212e7703). The shipped prover never replays that far: the app exports [first-1, last] with the account dump (prover.rs, the 0.3.14 rule). On p1-5090 the kit's pair (host 71bc2438, export 263bf4ce; the same bytes in the 0.3.17 and 0.3.19 kits; the proving crate and the exec types identical across both trees) fed the app's form reads "account dump: 83 accounts after chain block 160829, state root equals the node's; replayed 8 segments, every state root equals the node's", so the pair is right; the fleet rolls it one box at a time with box-prover.py exporting the app's way, the two pairing lines read after one segment. The from-27276 divergence at 72142 is with the node and exec lanes (not on any proof path). "e809e396" is not a prover on the box (the hands run igneumd only). CI: tools/ci/prover-pair-check.sh (the prover's evm-types tree equals the pinned node's; --self-test; in pre-push, 32 checks) reads ok on the pin. One oddity for the node lane: igneum_exportSegments ["0x2747d","0x27485"] answered with segments 160829 to 160837, 64 blocks below the ask.
## 29. The 0.3.20 tree as of 10:2xZ (after the split and the Mac's reboot)
- Numbers: 0.3.19 = the app-only miner-ui-5 cut on pin 5899f603 (release-0.3.19); this tree = 0.3.20 (release-0.3.20 app, release-0.3.20-node = dc141409 on the mirror, the canary on it running as 0.3.20's gate); decimals 0.3.21 unless main says otherwise.
- Node side to come on release-0.3.20-node (the node lane): aea0ca5c the proof archive; the amended class v4 (CLASS_SIGNAL_V4 = 5; byte-4 signals never count; the kaspa-pow vector test; the daemon's window line naming object and sub-version; igneum-pow at the hash lane's a0aaca92 with the seven packs under the sub-version-1 stamp); the exec RPC listener watchdog (a dead listener rebound once with a log line, a second death within a minute exits; gated on the every-method test and a kill-the-listener unit test); N8's subsidy height (DEVNET_PARAMS at the cut). The flip arithmetic (the Counter ASIC lane, plan 6.6 on ca3-v4-node fa5bc9e6): the two v4 fields publish only after the one-sweep rollout and every worker on the 0.3.20 tree; the floor moves to the publish DAA + 604,800 rounded up to the epoch boundary (882,000 for a 12:00 UK publish on 7 October; recomputed from the live DAA at the publish), the window 86,400 unchanged; the earliest flip about 6 days 10 hours after the publish, never before every node has had the sweep plus a week. The rollout clock for that: 32 minutes for the 14 standing boxes under the lock rule, the hands and the seed about 3 minutes after, the Mac and PCs within minutes.
- App side: ember-heat (worktree igneum-wt-ember-heat, 7c779035, cd1034a2, 0d5fc6d2 on e9fe1106: heat mode in the engine and Ember, the region and price step at first run, the cost-against-rent row, two gate checks; app 198, UI 56, gate 32 green), main's word; its PC 1 4-hour hold is a runbook, not a gate. The ui-ota channel (the same lane): a "ui" object inside the signed manifest body, one signature; the bundle under dl/ as the installers; ui.version three-part; publish-manifest.sh gains --ui <tar> (the one writer); the kill switch is the object's absence plus a fallback on any mismatch; a CI check that the manifest's ui.sha256 equals the public tar. The 0.3.19 gate module (execrpc) and version carry over at the merge.
- Ledger: N10 (eadae138): the 72142 port-versus-node root is the thin-record export after a snapshot cut at FULL_RECORDS 1,200, not a rule; the exec lane owes the one-time re-execution that fattens thin records and an exporter that refuses a thin segment; "accounts" in an export is the tip state.
- The Mac rule (main, after the 10:5x UK crash): the Mac builds only the macOS binaries and the DMG, one at a time under the build lock; every other build, suite and the Windows cross-build on the box or a PC; the app gate runs on the box.
**Node lane, 10:3xZ:** (1) the exec RPC listener watchdog is written for the node line (`rpc::serve_watched`: the listener task polled every 10 s, rebound once on a death with "exec RPC listener died: {reason}; rebound on {listen}", a second death within a minute exits 3 "so the app sees it"; gated by the every-method test and a tokio kill-the-task test). (2) **The igneum-pow of the 0.3.20 cut is the hash lane's 8c728ca3** (ca3-v4-amend), not a0aaca92: a0aaca92 keyed the load-source rule on the whole class with the shadow's pass count inside, so the base program moved with the ladder rung (caught by the fork's ladder test on the box 10:06Z) and did not compile against dc141409 (no chain_program_shadow); 8c728ca3 keys it on the class with the pass count set aside and carries the ladder igneum-pow underneath; the pinned ids do not move. (3) On the node line for the observer: igneum_claimSegment, igneum_getProofClaims and `claims` on igneum_getProofRecords (the claim posted before a prove), plus the app lane's four finality methods. Tips as each lands.
## 30. The dc141409 canary: synced 10:29:19Z, every decisive read clean (the fleet)
c18-1 (RTX 3070 pod, wiped datadir, from 08:51:50Z; the Mac's reboot cut the form's shell at 10:5x UK and it re-attached): synced at 141,357 blocks (headers 141,617, 3 peers). IBD-end line: "10:29:13.927+00:00 [INFO ] IBD with peer 213.173.107.74:16516 completed successfully; finality time: bodies 16592238 ms over 141699 blocks, virtual 1061858 ms over 36183 changes, weight tables 968243 ms over 359709 tables (3832285840 blocks walked), signatures 711416 ms over 439307, persists 40537 ms over 3324"; the relay catch-ups after it a second each. Counts: "Proofs: asking" 0, "carries N proof record" 0 (the oracle switch holds), SendPingsFlow 0, idle-drop 0, the class-signal warning 1 line, 6 IBD sessions, 2 "completed with error" before the re-attach (lines owed with the restart reads). The app's poller: 4,217 eth_getBlockByNumber calls, 0 errors, 0 non-null, the node alive. At the synced line the exec layer read "exec not synced: this node's consensus starts at pruning point 36a7ba0d" (executedTipHash null). Next: ten minutes of mining, the hub read about 10:40Z, the restart reads to about 10:55Z, the cases on the fresh pods. Prover roll 7 of 13 paid; hub-1's prover moved to the pair and the [first-1, last] export.
**App side taken for 0.3.20:** ui-ota 0247b065, c337f768, 10c881dc (on miner-ui-5 b322e9fa: the "ui" object inside the signed body plus the entry's own signature, kept; "size"; src/uiota.rs; Settings > Interface; publish-manifest.sh --ui/--no-ui; tools/ui-ota/publish.mjs with --verify as the post-deploy step; ui/VERSION 1.0.0 embedded, 1.0.1 the first bundle, min_engine 0.3.20 so a 0.3.19 engine ignores it) and ember-heat 7c779035, cd1034a2, 0d5fc6d2; both green on the box. The miner-reliability branch (workers start only when READY = synced and an executed tip; the node watchdog never counts the catch-up; the restart ladder with no permanent fault; fault lines to the intake) is main's call: a 0.3.19 follow-up on the same pin, or this tree's app.
**Fleet kit rule (main, 10:3xZ):** one prover identity per box; the kit never ships an identity file; box-prover generates its own at first start from the box's label and keeps it in the registry row; a shipped or duplicated identity is refused at start with a line. Migration on the shared boxes (9e4ba6b0 on three, faa34a1a on one) one at a time, prover only; prover identity keys are not vote keys (no weight, no signal), so the lock rule does not apply.
**Main (10:4xZ): miner-reliability rides 0.3.20's app** with ember-heat and ui-ota (ee09ae8b, docs only, goes with its three); no app-only follow-up and no renumber. One exception: if the reliability lane's watchdog-clock and export-lock changes land small and green before the node line is ready, main decides an app-only 0.3.20 then (the feature node would become 0.3.21). PC 1 today: the Arc held off and the packs-ahead restart on 0.3.19.
## 31. Rows from PC 1 on 0.3.19 (11:3x to 11:4x UK)
- The template path on the 0.3.17 node: on PC 1 (24 identities, 3 cards x 8) every getBlockTemplate times out at 5 s on every identity ("template fetch timed out (5 s) for identity N", 59 + 43 + 28 lines in two minutes), no STATUS line; the 0.3.19 watchdog treats it as the miner's heartbeat ("waiting for the node to answer block templates") so the cards wait rather than fault. The node logs no per-request time on this tree (the timers are 500ddd66, 0.3.20). Mitigation today: identities down to 2 per card by the project lead's tap (no signed job can set identities: a script may not POST /api/cards, and the runner's --cards-off only toggles enabled; a runner option for identities is a small jobrun.rs change, main's word). The fix is the 0.3.20 node on PC 1 first (the memoised tally 500ddd66; fd7de1b4 the off-path refresh is 0.3.21 material unless main moves it).
- Two app faults for the reliability lane's rules, both 0.3.20's app (main): (1) engine.rs:3032, the runner's --stop-miners hold is not released while a following job runs (the resume fires only when job_hold is set and no job holds the miners), so a read-only watch job kept both cards off for three minutes; (2) orphan igneum-miner.exe processes the app no longer tracks (two alive under --stop-miners with their rows at pid 0, one after) hammer the node's template RPC beside the tracked miners, part of why PC 1's template calls ran past 5 s. Sizing of the orphan-kill for an app-only 0.3.20: below.
- The fleet's prover-identity finding was withdrawn (no box shares a key; the "shared" hashes were a carrier block's record list); no migration. The kit rule stands on another ground: `igneum-miner key-hash <label>` is a pure function of the label, so the 0.3.20 kit seeds the identity from the box, keeps the hash on the registry row, refuses a duplicate.
- The node lane's IBD-end read (ca3-v4-node d49c71d0): the weight-table cache (64 tables, cleared past that) made every 2,000-body IBD batch walk the window again (359,709 tables for 5,335 checkpoints, 968 s); fix on release-0.3.20-node: WEIGHT_TABLE_CACHE = 1,024, oldest-first eviction, never a clear, a unit test; bodies 16,592 s is thread time behind the finality state lock, not CPU; the next fresh join on the pod should read 60 to 70 minutes.
**Main (11:5x UK):** a one-off exception to the 6 October rule, the project lead's word ("you will have to fix it"): exactly one POST to /api/cards on PC 1 from the Intel lane's scratch script (not in the tree), the 5090 and the 9070 XT to 2 identities and the Arc off; the engine restarts the workers on the change. The permanent path is the reliability lane's signed `cards` job kind (per-card enabled, identities, power_pct, through the app's own card path, read back), which replaces any runner option: 0.3.20.
## 29. The app side assembled on release-0.3.20 (11:5x UK)
Picked by hash onto ca1742cd, in order: miner-ui-5 810bf5a1 and b322e9fa; ui-ota 0247b065, c337f768, 10c881dc, ee09ae8b; ember-heat 7c779035, cd1034a2, 0d5fc6d2. A merge of ui-ota was tried first and aborted (14 files, its base is miner-ui-5's lineage, which this tree carries as picks). Five conflicts, each resolved by union: publish-manifest.sh keeps Ember's `activation_passed` form and takes ui-ota's `--ui`/`--no-ui` arms, UI block and the `"${UI:-}"` argument (one python call, `sys.argv[1:13]`); the Settings literal in config.rs and the SettingsState literal in engine.rs carry both heat and `ui_builtin`; app.js's export carries both lanes' names; pre-push.sh runs the ui-ota self-test and the heat gate reader. The Mac's `igneum-ota-sign` was rebuilt under the build lock (the stale binary lacked `sign-ui`, the pre-push self-test read RED). UI tests on the box: heat-region 8, notices 8, tune-line 5, ui-ota 4, update-card 4, view 30, all pass. The app gate (`cargo test --release`) runs on the box; the push follows its green.
**Held in the 0.3.20 verdict (the fleet, 11:4x UK):** isSynced reads false at the tip on dc141409 (7 of 12 reads false with blocks equal to headers and no IBD session; a 17-read sample 13 false; the 0.3.17 control reads true 12 of 12). Mining, acceptance and the exec follower unaffected. Relayed to the node lane: a fix on the node line or a ruling that the kit reads the tip. The publish waits on that word.
**Main (12:0x UK), the Arc B580 for 0.3.20's worker kit:** Intel's OpenCL compiler folds the kernel's `rotr_var` helper `rotate(x, (0u - n) & 31u)` into a left rotate by n, so every variable right-rotate is wrong on Intel. The fix is worker-side only: on an Intel platform the worker rewrites that one helper line to the shift form before clBuildProgram, behind the vendor check; no consensus or pack change; the vectors self-test already refuses the wrong hash. When the Intel lane reports bench d green on the Arc (vectors pass, a rate row), its host.c patch goes into the shipped igneum-worker-opencl for 0.3.20 (Windows and Linux) with a unit test that the Intel path builds the shift form and the AMD path is unchanged; the Intel vendor badge and card row from intel-arc ride the same cut. Asked of the Intel lane: the patch as one commit, the test and its command, the badge and card-row hashes. The 0.3.20 worker build holds on that; nothing else in the cut waits on it.
**The fleet (12:0x UK):** the 0.3.20 cases (relay and poison against c18-1 on dc141409) start at about 11:20Z after a re-rent (relay on a 3090, poison on an A5000); the ten-member pool window keys on CASES END; prover roll 10 of 13 paired.
**The Arc fix is in (12:2x UK):** the Intel lane's bench d green on the Arc B580 (run-ia-arc-bench-20261007-d, 11:20:41Z to 11:22:26Z, the Arc alone through --cards-off; self-test PASS on class v4, the v3 control and the live pack, 96 of 96 vector lanes each; fingerprints f410c731b6bc2d31 v4 and 90f794dd556f7a3b v3 equal to the Mac's and the 5090's; 11.011 MH/s on v4, 11.019 v3, 11.002 live pack, 10.882 with --exchange local, memprobe ceiling 11.02; the lane-0 register trace identical to the CPU interpreter in all 55,809 snapshots). Picked: 26e135a3 → 9088293a (proto-opencl/intel_rotr.h, host.c +17, test_intel_rotr.c, the pre-push run line; the C test passes on the box); 18453bb1's six app files as their own diff → adf79cec (detect.rs, sweep.rs, app.css, app.js, view.test.mjs, ui-mock server.mjs; the view test's prove sentence in this tree's wording). UI tests on the box after: heat-region 8, notices 8, tune-line 5, ui-ota 4, update-card 4, view 31, all pass. Still to come from the Intel lane: the cards_leave_off runner param (jobs.rs, jobrun.rs, publish-jobs.sh --cards-leave-off), hash when its box tests are green.
**App gate green on the box (12:28 UK):** `cargo test --release` on 5b3bd023 (211 + 28 + 8, rc=0, 11:26:51Z) and again on 357b88b3 after the Intel changes (213 + 28 + 8, rc=0, 11:28:34Z; the two new tests are detect.rs's B580 rows). The pre-push gate's 34 checks pass on the tree, the ui-ota self-test among them. release-0.3.20 pushed on this row.
**Main (12:3x UK): hold for the isSynced fix.** The 0.3.20 app's execrpc gate and the worker-start rule both key on the node reporting synced; a flapping flag would hold workers back on the machines the cut is meant to make plug-and-play (PC 1's morning in a new form). The node lane carries the fix as a blocker on the node line with a test; the pin waits for it. If it slips past 16:00 UK, main hears the node lane's estimate and decides again. Everything else as set.
**cards_leave_off in (12:33 UK):** the Intel lane's 9e794503 → afb0cfe9 (jobrun.rs: a run job's --cards-off cards restored as OFF with identities and cap kept when the job's bool param cards_leave_off is set, the report line "cards LEFT OFF ... persisted by the app"; jobs.rs kinds doc; publish-jobs.sh --cards-leave-off emitting "cards_leave_off": true beside cards_off). App gate on the box, third pass: 214 + 28 + 8, rc=0, 11:32Z. Nothing else from the Intel lane for the cut. The reliability lane's hashes are app 38a30397 and fork f067f7c1, held until its pod injector finishes (about 45 minutes) and it sends the lines.
**The fleet (12:4x UK), prover roll: a 12 GB prover on the open devnet is paired and unpaid.** p2-4070-1 with the 0.3.17 pair: 7 segments "pair ok", 56 shards accepted, 0 refused, and 0 of 8 claims paid in 25 minutes (7 segment_refused "segment already paid; end to end 203 s", one "does not chain to ... which is proven"; 4 held, 3 held_expired). Its claims are fresh at claim time (margin 432 to 518 DAA, 5 to 7 candidates, rank_by fnv), so a 24 GB box claims the same segment and pays it inside the 4070's 203 s. p1-4070 paid once this morning (2.6 IGN after 176 s). Per tier: a home prover on a 12 GB card earns nothing while a 3090 or 4090 is awake on the same segments; the fix is the claim rule (a settled-depth claim with a per-key reservation, or the candidate hash spread over keys), which the node lane carries on the 0.3.20 line; until then a 4070 prover is a verifier that never pays, not a box fault. The roll continues with p1-3080 and hub-1; p2-4070-1 recorded "pair ok, unpaid (race)". Proving feed 11:35Z: provers_10m 4, 62 shards, lag 248 s. Asked of the node lane: confirm the claim-rule change is on the tip I pin.
## 30. The node line: release-0.3.20-node = 8097d600 (the node lane, 12:5x UK), the cut rule
8097d600 = dc141409, the archive aea0ca5c, then one commit: isSynced is the consensus rule alone on GetInfo and the template (headers and blocks at the tip), `behind` answers from the hook's stamp and never from a lock try, the finality catch-up reported apart as `finalityBehind` on igneum_getNodeInfo (the fix main holds for); the amended class v4 as object 5 (byte-4 blocks never count; devnet epoch-0 vectors pinned; the window line names the object and sub-version 1); the weight-table cache bounded at 1,024, oldest-first; the template snapshot refreshed only while templates are wanted (join bench 220 to 92 ms a checkpoint); the submit reply ahead of the virtual state; template wait 100 ms; igneum_getFinalityWeights and igneum_getFinalityCheckpoints {last} with signers. igneum-pow pair = the hash lane's 8c728ca3 (not a0aaca92). Suites green on build-2 12:28 to 12:37 UK (consensus-core 123, kaspa-pow 17, exec RPC, four finality tests, flows, rpc-service). Binaries building on build-1 for the two gates (mixed-version Devnet 2 beside the 5899f603 pair, the digest test).
A second commit follows behind the gates: igneum_getFinalityKey, igneum_getProofRecordsByKey, the observer's claims (igneum_claimSegment, igneum_getProofClaims, claims on getProofRecords), the settled claim floor for the provers (the fleet's 4070 race), and the listener watchdog (N7's macOS shape); its suites on build-2 now.
**The cut rule (shipper, 12:5x UK):** the pin is the second commit if its suites and gates are green by 15:30 UK (the watchdog is in 0.3.20's scope, the claim floor answers a live finding); past 15:30 UK the pin is 8097d600, the rest goes to 0.3.21, and main hears at the 16:00 checkpoint with the node lane's estimate. Asked of the node lane: its estimate, the commit string and gate lines on green, and what the kit's wait steps read on this line (isSynced alone, or finalityBehind too). Note: two notes meant for the node lane went to ae892a8b0f78fe31c by mistake earlier; the node lane is a283f5f0d364ceef0.
**Main (13:0x UK): the cut rule accepted, two conditions on the second commit.** The claim floor ships with its own test line (a 12 GB prover paid at least once in the ten-member window, or a harness equivalent); the watchdog with the N7 shape reproduced then clean; both named in the tip. Publish order unchanged: canary from the wipe on c18-1, PC 1 first, then PC 2 and the Mac, one box at a time with lock lines. The pin is reported the moment it is named. Sent to the node lane and the fleet (the fleet may be asked for the 12 GB line on a pod inside its pool window, the p2-4070-1 read).
**The node lane's estimate (13:1x UK), UK clock:** second-commit suites green about 13:35; the commit on release-0.3.20-node right after (one line on top of 8097d600); its binaries on build-1 about 13:50; the two gates about 14:15; the fleet's 12 GB line about 15:00 if started by 14:20. Any slip past 15:30 pins 8097d600. The conditions as the node lane will meet them: the watchdog's line is its unit test (the listener task killed, nothing answers, the watchdog rebinds once, the exec RPC answers again, a second kill reaches the exit hook); a live reproduction from outside the process is not possible (a task death, not a signal), so the test is the line, named in the tip. The claim floor's line comes from the fleet on a 12 GB pod: igneum_getProvingStatus settledNumber/settledDaa/settledBy and `settled` on every igneum_getAssignedShards row, box-prover claiming only settled rows. The kit's wait steps on this line: synced = GetInfo's isSynced alone (headers and blocks at the tip, as 0.3.17 read it); finalityBehind for display and the app's "catching up finality" line only; box-prover does not wait on finalityBehind (settled rows answer it).
**The floor at the publish (the node lane's note, the coordinator has it):** the live override file (6 October 22:49Z) carries program_class_v4_activation_daa 831,600 and window 86,400, so every 0.3.17 node signals byte 4 today and flips to the OLD v4 stream at epoch 231, about 13 October 09:00 UK, whatever is signalled. The 0.3.20 file must carry the moved floor (publish DAA + 604,800 rounded up to the epoch boundary, about 882,000 for a publish today), and the sweep must replace every 0.3.17 node before 13 October 09:00 UK, or the straggler forks alone then.
**Main's rulings (13:2x UK).** (1) The watchdog's unit test (listener task killed in process, rebind once, exec RPC answers again, second kill reaches the exit hook) meets the condition; named in the tip, no live line. The claim floor's live 12 GB line stands. (2) The binaries publish stays on the 15:30 UK pin rule. (3) The floor-moved live file (program_class_v4_activation_daa = publish DAA + 604,800 rounded up to the epoch boundary, window 86,400) is a live manifest change and goes out only on the project lead's explicit word: staged beside the release with its digest and a one-line diff against eada4bda, named in the 16:00 UK report, NOT published. The binaries publish carries the live file as it is (eada4bda). (4) The sweep of every 0.3.17 node (fleet, hands, seed, PC 1, PC 2, the Mac) runs with the binaries publish and must finish before 13 October 09:00 UK; that date goes in every rollout lock line. (5) The 16:00 report confirms what a 0.3.20 node on the OLD file does at epoch 231 (the Counter lane: it flips to the amended stream and the devnet stays whole if every node is 0.3.20), so the project lead chooses between the file move and the sweep alone.
**Epoch 231 on the old file, confirmed by the node lane with the reference (13:4x UK):** `program_class_for_epoch_signalled` (consensus/core/src/igneum.rs line 452 on 8097d600) answers the floor first through `program_class_for_epoch` / `program_class_for_epoch_at` (line 320, the floor rounded up to the epoch boundary): V4 at every epoch at or past 231 for 831,600; the signal branch is consulted only where the floor says V3. The class is one enum value on both binaries; the stream is each binary's own igneum-pow (`pow_class_of`, `Epoch::chain_program_shadow`): 8c728ca3 with sub-version 1 on 0.3.20, the 6 October generator on 0.3.17, and `program_id` (the "sub/" suffix on 8c728ca3) refuses the other's blocks. So at epoch 231 on the old file every 0.3.20 node flips to the amended stream; if every node is 0.3.20 by then the devnet stays whole with no file change; a 0.3.17 node flips to the old stream and forks alone. A 0.3.20 node does nothing differently between the old file and the moved file: the signalling byte is 5 either way (`template_signal_byte` line 499 stamps it whenever both fields are set), the window line prints whichever floor it reads, the rounding is the same on a different number, the tally and the seven-window rule untouched. For the project lead: the sweep alone works if every node is 0.3.20 before 13 October 09:00 UK and leaves no second file; the file move buys a week's margin for stragglers at the price of a digest change on every node (the sweep anyway). Digest: 8097d600 with the live file reads eada4bda, as 5899f603 does.
**Staging recipe (scratch r0319-app, 13:4x UK):** `digest-box.sh <igneumd on the box> <file|''> <port>` starts the binary on the box alone (no peers, own appdir) and prints "Consensus params digest: <hex>", killed by pid (the lost r0312/digest.sh replaced; the Mac builds and runs nothing). `stage-floor-file.sh <live daa> <igneumd on the box> <port>` writes ov16-floor-<floor>.json from the live sixteen-field object (floor = ceil((daa + 604,800) / 3,600) * 3,600, window untouched), prints the one-line diff and both digests. Dry run on the 0.3.17 binary at DAA 275,300: live eada4bda, moved floor 882,000 reads 844ebde1 (the moved digest is read again on the pinned binary at the publish). Not published.
**8097d600 alone is not a pin (the node lane's honesty line, 13:5x UK):** the whole finality test module on build-2 shows one red unit test on 8097d600, `the_template_answers_and_synced_reads_behind_while_the_catch_up_holds_the_state` (500ddd66's PC 1 rule test, line 37 asserting `behind` TRUE while the state is held, the expectation the isSynced ruling inverted); the code is right, the test is stale; on 8097d600 the lane had run the four new finality tests, consensus-core, kaspa-pow and the exec RPC suite, not the whole module. Shipper's word: the fallback pin is a one-line test-only commit directly on 8097d600 (behind false with the state held, the snapshot still answers), pushed as release-0.3.20-node, the second commit rebased on it without the test change; the whole module run on the fallback; the gate lines on the 8097d600 build carry to the fallback (byte-identical code, stated in the tip) but the pinned binary is rebuilt from the fallback commit for the commit-string read-back. No known-red pins.
**The fallback pin is on the mirror (14:0x UK): release-0.3.20-node = 6b94c823**, the test-only commit directly on 8097d600 (no code line touched; binaries byte-identical to 8097d600's; the two gate lines on the 8097d600 build carry to it, stated in the tip). The whole finality module runs on 6b94c823 from a clean worktree on build-2. The second commit (key methods, claims, settled floor, watchdog) lands as 6b94c823's child when its suites are green; both binaries rebuild on build-1 then (the fallback's for the read-back, the second commit's for the fleet's 12 GB line). The vendor worktree has the mirror branch fetched; the checkout happens at the pin.
## 31. The candidate pin 6a3432a3, the builds started (14:5x UK)
release-0.3.20-node = 6a3432a3 (the child of the fallback 6b94c823, itself the test-only child of 8097d600): the key methods, the observer's claims, the settled claim floor, the listener watchdog. Suites on build-2 13:46 to 13:48 UK: the whole finality module 25 passed, the exec suite 29 passed with the watchdog's test `rpc::watchdog_tests::the_watchdog_rebinds_a_dead_listener_once_and_exits_on_the_second_death`, kaspad, flows, rpc-service green; consensus-core 123 and kaspa-pow 17 on the same code. The fallback's own line: the whole finality module on 6b94c823 from a clean worktree, 25 passed. Gate rule (shipper): lines on one binary do not carry to a binary whose code changed; 6a3432a3 runs the digest gate and the ten-minute mixed-version gate on its own binary (lines about 14:20 UK with the commit string read back), the fallback on its own (about 14:40 UK); the first mixed-version run on the 8097d600 build is void (the path was replaced mid-run). The digest harness reads the thirteen-field b18ed271 and the sixteen-field moved digest by design; eada4bda on the pinned binary with the live file is the deploy gate's read.
igneum-pow: release-0.3.20 carries the hash lane's 8c728ca3 tree at 00249643 (the pair for both candidates; byte-equal). Asked of the build-server lane: the seed and hands pairs for both candidates, held until the pin and the canary. Asked of the Counter lane: that nothing past 8c728ca3 touches igneum-pow, and the hash lane's pairing line.
Speculative builds on the candidate from the ship worktree, box only, sequential (scratch r0320/build-candidate.sh, started 12:52Z): igneum-pow tests, fleet-native node (target-0320), hive class (target-0320-hive), Windows cross of the node and the app, the app's Linux binary; shas and commit strings at the end. If the pin falls back, the same script runs on 6b94c823. The Mac builds only its own binaries and the DMG, under the lock, after the pin.
**The amended class v4 packs (the hash lane, 15:0x UK), what the fleet reads on the canary.** igneum-pow 8c728ca3, byte-equal to 00249643's tree; program id 1a4230699a6b9c60 (generator 4, class v4, sub-version 1) on every v4 pack; the pre-amendment id c120d7963abdcd96 is the must-differ vector (tests/recheck.rs and the node's kaspa-pow test); the v3 control mx8-devnet-epoch0 id 73bcbfe8ccf988f1 fingerprint 90f794dd556f7a3b unchanged. Fingerprints 2^24 at base 0, equal on Metal, Apple OpenCL and the RTX 5090 (NVRTC 12.8 sm_120): v4-devnet-epoch0 867dbc45cfb36b4d (vectors 756301bf1739a7ee); v4-era-0 2146ecacc8c75a8e; era-1 fe52602393f6d3d4; era-2 3b206471a13912b4; era-3 c3f03c4a5d7333aa; era-4 f1dfd7209f15bb97; era-5 8c194da64fadf31d. Packs at proto-cuda/packs-ca3-v4 on ca3-v4-amend; the eight-pack kit zip sha256 889ec99976d2728b4b5035bfa476032e5b6a13b928968fc45236d5f25084aa39. G1 (d8859522): PC 2 run-ca3-v4-amend-g1-pc2-20261007 09:41Z, the 0.3.17 worker, self-test PASS on all eight packs, every fingerprint equal. The pairing line (igneum-pow 8c728ca3 against the fork's kaspa-pow, `test --release -p kaspa-pow --features igneum-pow`) is in flight on the box; the node lane's own run of the source-rule line on the 0.3.20 fork passed (1a4230699a6b9c60 equal, c120d7963abdcd96 differs, v3 unchanged, object byte 5).
**Interop fact from the void 8097d600 run (the node lane, 15:1x UK):** on the live sixteen-field file plus genesis_bits (one digest on all five nodes, f92a0675), the 5899f603 hub accepted every block the 8097d600 node mined, 235 accepted and 0 rejected, the old node's headers at version 1026 (block version 2, object byte 4) and the new node's at 1282 (byte 5), counts equal on all five at 472 before the restart step. So a 0.3.17 node takes byte-5 headers from a 0.3.20 node and relays them: the amended class's one interop question, answered. The run's four FAILED checks are the binary swap (n1 restarted two seconds before 6a3432a3's build replaced the path) and one harness expectation written for a file without v4 fields (`every_header_version_is_the_block_version`); the clean runs use the thirteen-field object. 6a3432a3's own gates started 13:53 UK, lines about 14:08 UK; the fallback's follow.
**Clock correction (12:55 BST, read from `date`):** the UK stamps in the rows from "Main's rulings" to the interop fact above ran ahead of the clock by one to two hours (the lanes' quoted "13:4x", "13:53", "14:08", "15:0x UK" included). The true times, from the commit times of the rows: main's conditions 12:40 BST; main's rulings 12:41; the epoch-231 confirmation and the staging recipe 12:43; the stale test and the fallback shape 12:47; the fallback 6b94c823 on the mirror 12:48; igneum-pow 8c728ca3 taken 12:51; section 31 and the speculative builds 12:52 (the script started 11:52:06Z = 12:52 BST); the pack table 12:53; the interop fact 12:54. The 15:30 BST pin rule and the 16:00 BST report stand on the true clock; the node lane's estimates (gates about 14:20, the 12 GB line about 15:00) are re-read against `date` when its lines land. From here every stamp in this plan is `TZ=Europe/London date`.
**The build-server lane (12:53 BST):** the four 0.3.20 pairs build serially on build-1 in its own worktree igneum-wt-bs0320 (igneum 00249643 detached, igneum-pow byte-identical to 8c728ca3; vendor/igneum-node-0320a = 6a3432a3, 0320b = 6b94c823): a hands (native 2.39), a seed (--ship seed, 2.35), b hands, b seed; artefacts in its scratch bs0320/<a|b>-<hands|seed>/; pairing line "pairs with igneum 00249643 (detached): igneum-pow 0.2.0", rustc 1.99.0 both sides. Lines (sha256, commit string, glibc need) as each lands; nothing to the hands or the seed before the pin and the canary. Found: neither 6a3432a3 nor 6b94c823 carries rust-toolchain.toml (the fork's copy is on fork master 37f1206b, not an ancestor); inside an app worktree rustup walks up to the app tree's pin, so every build on the line is pinned; a standalone checkout of release-0.3.20-node is not. Owed for 0.3.21 (not now: a file commit would move the pin and re-run the gates for no code change): rust-toolchain.toml on the node line.
## 32. Two defects on the line, the pin moves to b7cc37e7, no fallback commit (13:0x BST, `date`)
**N12, the watchdog and a held port (6a3432a3's own digest gate):** the gate's four nodes share one exec JSON-RPC port; the watchdog counted "cannot bind: address in use" as a listener death (rebind at 10 s, "died twice within a minute" at 20 s, exit 3), three of four nodes gone before the harness read their peers, where 0.3.17 and 8097d600 warn and live without the exec RPC. Digest facts came out right before the exits: thirteen fields a89be8a7 on both binaries, the sixteen-field object db9a85f9 refused with the mismatch line. Fix 09124180: a bind failure is a retry every poll with one line a minute and no death counted; a death after a successful bind keeps rebind-once-then-exit. Tests: `the_watchdog_rebinds_a_dead_listener_once_and_exits_on_the_second_death` and `a_held_port_is_retried_and_never_counted_as_a_death`; exec suite 30 passed.
**N13, the kept-datadir death (the fleet, starting 6a3432a3 on pool-1's kept 0.3.17 copy):** `called Result::unwrap() on an Err value: DeserializationError(Io(Kind(UnexpectedEof)))` at consensus/src/model/stores/virtual_state.rs:250. Cause: 10db4b61 (0.3.16 feature line, vote-or-burn and the signing bonus) added `silent: bool` to `BlockRewardData` under `#[serde(default)]`; bincode is not self-describing and ignores serde defaults, so the virtual-state row a 0.3.17 node wrote (three-field rewards in `mergeset_rewards`) reads short on every build from 10db4b61 on: dc141409, 8097d600, 6b94c823, 6a3432a3, 09124180 all die at start on any kept 0.3.17 datadir; no canary saw it because every canary wiped. Fix b7cc37e7: the store reads the live row in the current layout first; on a deserialization error it decodes the row as a mirror of the v1 layout, converts with `silent` false and rewrites it under the same key in the current layout; version suffix unchanged. Test green on build-2 (a v1 row in a temp DB: the current layout reads it short, the store reads and rewrites it, a second open reads first-try) plus the kaspad check.
**The pin rule now:** candidate b7cc37e7 (8097d600 → 6b94c823 → 6a3432a3 → 09124180 → b7cc37e7), igneum-pow 8c728ca3. No fallback commit on the line (6b94c823 dies on a kept datadir); if b7cc37e7's gates are not green by 15:30 BST, 5899f603 stays live and 0.3.20 ships later on green. New gate before the canary, whatever the pin: the kept-datadir start, the pinned binary on a copy of a standing 0.3.17 box's datadir on a scratch pod, the rewrite line as the pass (the fleet). The node lane's clock: b7cc37e7's build about 14:45 BST, the kept-datadir read about 14:50, digest and ten-minute gates about 15:05, the 12 GB claim line about 14:50 to 15:00 on the 6a3432a3 pod (claim code unchanged). The build-server lane's a/b pairs are void and rebuild on b7cc37e7; the shipper's speculative builds on 6a3432a3 stopped by pid (script 30520 and its child) and restart on b7cc37e7; the vendor worktree now at b7cc37e7. Also: the amended v4 packs were stale in the app tree (igneum-pow's recheck tests read the old program id c120d7963abdcd96 from program.json); proto-cuda/packs-ca3-v4 taken from 8c728ca3 as its own commit; tests rerunning on the box.
**Main (13:1x BST):** the fallback reading stands (b7cc37e7 by 15:30 BST or 5899f603 stays live and 0.3.20 ships later today on green). Three additions: (1) standing rule, every canary runs a wiped and a kept datadir, the kept-datadir start a named gate in every release, in docs/plans/release-rules.md (rule 4) and the miner-reliability register as its own fault class (asked of the reliability lane); (2) the Windows shape: the pinned Windows node once against a copy of PC 2's datadir (PC 2 only) before PC 1 gets the build (asked of the build-server lane, with the pairs moved to b7cc37e7); (3) the 16:00 report stands, but on green earlier the publish goes out on green with the clock time. The fleet has the rule and runs the kept start on the pod copy, then the wipe canary on c18-1, then the kept read on c18-1 before its restart step.
**6b94c823's gates on its own binary (the node lane; the lane's "13:56 to 14:08 UK" = about 12:56 to 13:08 BST):** sha b1b7d47b, string 6b94c823. Digest gate: thirteen fields a89be8a7 on both binaries (compat), the sixteen-field object db9a85f9 refused with the mismatch line; the one FAILED check `n3_has_no_peer` is the refused connection's reconnect in flight at the read (harness fix 36d3efdc: the minimum of five reads). Mixed-version gate, ten minutes on the thirteen-field file: digest b0afb2ee on all five, the 5899f603 hub accepted every block the 6b94c823 node mined (146 new, 246 old, 0 rejected), header versions plain 2, counts equal at 536, 695 and 785 through the clean join via the old hub, the join served by the new node, and the new node's restart; the one FAILED check `no_panic_in_any_node_log` is six "Address already in use" panics in the two old nodes' server threads at start (the gates overlapped on ports 29830/29831; now a 20 s gap). So 8097d600's code is clean on both gates by the checks that bear on the binary; the fallback is still not a pin (N13). From here only b7cc37e7 gets gates, on its own binary when its build lands (about 13:48 BST by the clock), lines about 14:05 BST with sha and string.
**Main (13:2x BST):** the driver check (the Intel lane's b9487dcf on driver-check: the table, the install thread with one elevated prompt, the manifest's drivers object, the row strip; box 219 + 32 + 8, cross green, pre-push 35, UI 32) rides 0.3.21 with the drivers manifest once the NVIDIA and AMD hashes are real; 0.3.20's app tree is closed. Clock: the node lane's "UK" stamps read an hour fast (UTC+2); every estimate is re-read against `TZ=Europe/London date` before it is quoted; the node lane and the fleet told to quote that clock. The 15:30 BST checkpoint is the real clock.
**The fleet's prover roll, p1-3080 (13:1x BST):** the pair right (28 "pair ok"), 0 paid in its history (2,167 claims), every proof since the pair dies at the compressed step at the memory wall (device_used 9,859 of 9,885 MiB), 29 of 29, miner off the proving step. Per tier: a 10 GB card does not prove on this pair; the failing prover took the 3080's GPU from its miner for 16 hours for nothing. On the shipper's word (fleet config, no key moves): PROVER=0 on p1-3080 with its lock line and the miner's rate before and after; the same reading for standing boxes under 12 GB; one rented 3080 for an hour to measure the lower-memory SP1 threshold. For main: the kit and the app refusing to start proving under 12 GB with the reason shown, until the measured threshold lands. hub-1 is the roll's last box.
## 33. The pin moves to c4459193; the 12 GB rule in the app; the roll complete (13:3x BST, `date`)
**b7cc37e7 struck:** its digest gate passed (12:14 to 12:16Z: thirteen fields a89be8a7 on both binaries, the sixteen-field object db9a85f9 refused, the live file's own digest eada4bda as 5899f603 reads it) and the fleet's kept-datadir read passed on p12-vast (12:17Z: "Virtual state: the row was written by a build before the silence field (a kept datadir); read as the v1 layout and rewritten in the current one (1 mergeset rewards)", the finality blob converted, 1,747 locks, no panic; a second start first-try; 6a3432a3 the known-failed shape), but its ten-minute mixed-version gate FAILED at the restart step (12:24:19Z): the node died on its own datadir LOCK (conn_builder.rs:167) because the 6a3432a3 watchdog sleeps its whole 10 s poll before checking shutdown, so every node from 6a3432a3 on takes up to 10 s longer to stop than 5899f603 (the fleet's "a 12-second timeout does not stop the node" was this) and a restart inside that window meets the lock. Fix c4459193 (b7cc37e7's child, rpc.rs alone): the poll in 250 ms steps returning the moment shutdown is set; test `a_shutdown_returns_within_a_second_whatever_the_poll`; exec suite 31 passed with the three watchdog tests. **The candidate pin is c4459193**, igneum-pow 8c728ca3; its build on build-1 from 12:30Z, digest and ten-minute gates on its own binary, lines about 12:50Z (13:50 BST). Carry ruling (shipper): the b7cc37e7 kept read is evidence, not the gate line; the gate line is the kept read on c18-1 on c4459193's binary before the canary's restart step. The b7cc37e7 pairs (hands bc6b3b3b / 240eb0a4, seed ffaa441d / 64207cc0) are evidence only; the build-server lane rebuilds on c4459193 and runs the PC 2 Windows kept-datadir job on c4459193's exe.
**Main (13:2x BST):** the attack-pass lane's F8 re-gate on object 5 heads to FAIL at 1.2x on nine of the first thirty seeds (worst p31 at 29.3x), far better than the old stream; 0.3.20 ships object 5 as it stands (the live floor flips every node to the OLD stream on 13 October otherwise); sub-version 2 on the Counter lane is 0.3.21's; the 16:00 report tells the project lead the floor move is recommended, not optional. Build slots: suites to build-2, builds and gates on build-1 (tools worktree fast-forwarded to master 1cf850c9; build-remote routes test and bench to box 2 by class).
**The 12 GB rule in the app (4c89372a):** MIN_VRAM_MB_PROVE_ANY 11,800 and PROVE_UNDER_12GB_LINE "proving needs a 12 GB card; mining continues" in provedefault.rs with `the_prover_refuses_every_nvidia_card_under_12gb_and_says_why`; the prover loop refuses before the sync wait when every present NVIDIA card is under it (status off, the sentence); the tile sentence in app.js (PROVE_MIN_GB 12) with its view test. App gate on the box: 215 + 28 + 8, rc=0 (13:13 BST); UI 31 of 31; the Windows cross of the app running. In the fleet's kit: box-prover refuses the same way (gpu-fleet f7d40af5, on 18 boxes at 12:14Z; p1-3080's line "RESULT refused 2026-10-07T12:14:30Z proving needs a 12 GB card; mining continues (memory.total 10240 MiB)"). PROVER=0 on p1-3080 (12:13:57Z, node and miner untouched, no key moved; lock line 98.6 percent at checkpoint 8931): the miner 38.5 to 39.8 MH/s before, 44.7 to 45.9 after, plus 17 percent.
**The prover roll complete (12:19:40Z):** 11 of 13 standing provers paid on the 71bc2438/263bf4ce pair (hub-1 last, 4.6410 IGN after 143 s); p2-4070-1 pair ok and unpaid (the claim race), p1-3080 prover off. Feed 12:25Z: provers_10m 5, 67 shards, lag 351 s.
**Prover kit facts from the 3080 hour (the fleet, corrected 12:3xZ):** every standing box's sp1-gpu-server is a binary built for its own card (p1-3080 sm_86, p2-3090-3 sm_86, p1-4070 sm_89, p1-5090 sm_120); a server for the wrong architecture fails every proof in 12 s with "CudaRustError: named symbol not found" and the miner never notices. Kit consequence (0.3.21 design, routed to main): one server per architecture picked by compute capability at install, or a fat binary. The 610-driver question is open again (the first host had both faults at once); one more rented 610 host with an sm_86 server separates them. The threshold series restarted on thr-3080b with the sm_86 server.
**For 0.3.21 (the Intel lane):** driver-check tip 1982dbc7 on release-0.3.20's 4c89372a (the stray #[test] at detect.rs:990 removed; the table measured: NVIDIA 617.42 f115c927…, AMD 26.9.2 593c1d73… with the https://www.amd.com/ referer, Intel 32.0.101.9034; box 219 + 32 + 8, gate 35, UI 32).
**The 3080 hour's number (the fleet, 12:33Z):** on thr-3080b (RTX 3080 10 GB, driver 570.211.01, the sm_86 server, p1-3080's own failed segment 163366..163373), at the DEFAULT SP1_GPU_ELEMENT_THRESHOLD the chain proof completes, rc 0, 66 s wall for the eight-block segment, device_used 8,642 MiB of 9,883 with nothing else on the card. p1-3080 fails because its miner holds about 1,547 MiB and 8,642 plus 1,547 passes 9,885 (its "free_mib=25" on every one of 29 failures). Per tier: a 10 GB card is a mine-only card OR a prove-only card, never both; the sentence "proving needs a 12 GB card; mining continues" is right for the default mine-and-prove install and stays; a 10 GB owner who wants to prove instead can, at the cost of the miner (box-prover's MINER=pause on the fleet; an app choice to route for 0.3.21). Lower thresholds reading now for whether the compressed step fits beside the miner's 1.5 GB; the number goes in the bench log with this sentence. The 610-driver host approved (one rent, the sm_86 server). Build-server lane: c4459193 as vendor/igneum-node-0320d, the sequence from 13:33 BST (hands, seed, the Windows cross of c4459193, then 6a3432a3's as the known-failed shape); the PC 2 kept-datadir job on those two exes.
**The wipe canary's clock (13:4x BST):** c18-1 is held by the 0.3.20 cases until about 14:00 BST, so a wipe canary on it would read synced about 15:40 BST, past the checkpoint. The shipper's call: the fleet rents a second one-shot pod of c18-1's class now and starts the wipe canary on c4459193's binary the moment its build lands (synced about 15:15 BST by the 98-minute class), with the kept read on the pool-1 0.3.17 copy and the restart step on that pod; c18-1 keeps the cases and the pool window. The wipe is the decisive read by rule (release-rules.md 3), so the pin never cuts without it and b7cc37e7's lineage does not stand in; a slip past 15:30 BST holds the pin to the wipe line and main hears the clock.
**Main (13:4x BST):** the clock change accepted as set: the wipe canary on the second rented pod is the decisive read, the pin follows its synced line, and a slip past 15:30 BST is reported as a clock, not cut on the other lines.
**The wipe canary pod (the fleet, 12:37Z):** c19-1, RunPod pod wpuke4tfu0vr49, RTX 3070 community, USD 0.13/h, c18-1's class; c4459193's igneumd sha256 45be9b02d1b002f5 (string c4459193 read back, build-1 12:34Z) and igneum-miner c7cfc40b; the dc141409 form (wipe, IBD from the pruning-point proof, synced, ten minutes mining with the exec poller, the hub read, the restart read on the kept datadir) with the kept read on pool-1's 0.3.17 copy armed behind the synced line. Synced about 14:20Z (15:20 BST) by the 98-minute class.
**The 10 GB tier, closed (the fleet, 12:3xZ; docs/bench-log.md "The 10 GB prover tier"):** below the default threshold the server dies before the compressed step ("Failed to read the response: early eof" at 12 to 13 s, about 5,000 MiB used) at 524288, 262144 and 131072 alike; the default is the only working value and completes alone at 8,642 MiB. Per tier: a 10 GB card is mine-only or prove-only, never both; the rule "proving needs a 12 GB card; mining continues" stands for the default install; a 10 GB owner who chooses to prove gives up the miner (MINER=pause on the fleet; the app's switch is a 0.3.21 routing for main). A 12 GB card fits both with about 2 GB spare, a 16 GB card with 5.8 GB. The 610 read runs on thr-610 (driver 610.57.04, the sm_86 server), one proof.
**CASES END, PASS on dc141409 (the fleet, 12:37:18Z; run from 11:18:10Z):** relay c17-relay (the 0.3.17 relay build, pool-1's kept 1026 datadir, 22 peers, 1,038 blocks relayed) and poison c17-poison (the poisoned copy, mining from 12:24:44Z) against c18-1 (dc141409): over 13 polls, poison wbv 0 and got 1 (one block at first contact, nothing after, the hub refused it), relay R_wbv 0, c18-1 T_wbv 0 throughout, T_from_relay 0, T_1026 headers seen and none accepted, blocks moving with the tip, powEngine igneum-pow; the hub holds 900 of c18-1's blocks in its last 900 and names it in 1 reject (the right outcome). The isSynced flap visible on this line (4 of 6 polls false at the tip), fixed from 8097d600. The three pods destroyed. The ten-member pool window rented at 12:37:36Z. The carry to c4459193 goes to main on the node lane's diff argument (every file from dc141409 to c4459193 against block and header acceptance, the version check, relay and peer handling); if any touches them the cases rerun on c4459193 beside c19-1 whatever the clock.
**The wipe canary started (c19-1, 12:38:55Z):** the node on a wiped datadir on c4459193 (sha256 45be9b02d1b002f5486d0f0108571c3b6042094113ad9da6f3d3d9ffc0072bba asserted before the put; the node's own line igneumd/2.1.0-c4459193, digest eada4bda8aa8368c, finalityBehind false, pruningDepth 108000, 4 peers at 12:39:26Z, IBD from the pruning-point proof), the exec poller every 5 s, the miner c7cfc40b staged. Clock by the 98-minute class: synced about 14:17Z (15:17 BST), the ten-minute mining read and the hub read to 14:30Z, the restart read to 14:40Z, the kept read on pool-1's copy right after. Polls every four minutes.
**The cases rerun on c4459193 (13:4x BST):** the node lane's diff dc141409 → c4459193 (13 files, +1,155 −48) touches one of the cases' four categories: consensus/core/src/igneum.rs sets CLASS_SIGNAL_V4 4 → 5 (the node stamps byte 5, the tally counts byte 5 and above); header_version_acceptable, signalled_version, class_signal_of and the IBD guard untouched; flow_context.rs keeps the relay broadcast's place and payload (submit_rpc_block returns after the block task, on_new_block spawned); nothing on inbound relay or peer paths; the rest is finality, stores, exec RPC and tests. By rule the cases rerun on c4459193; the expected shape is dc141409's with the header byte read 5 (the void 8097d600 run showed the old hub taking 235 byte-5 headers, 0 rejected). To land inside the checkpoint the fleet runs it now beside c19-1: the target a fresh pod with c4459193 on pool-1's kept 0.3.17 copy (synced in minutes by the N13 path, itself a kept-datadir read on the pinned binary), the relay and poison pods against it, CASES END about 15:05 BST if the pods are up by 12:48Z.
**The 610 read (the fleet, 12:40:14Z):** thr-610 (RTX 3080 10 GB, driver 610.57.04, the sm_86 server, the same segment) completes the chain proof, rc 0, 87 s, device_used 8,729 MiB. The 610-series driver proves; this morning's "named symbol not found" was the architecture alone (an Ada server on Ampere cards). The kit's driver rule keeps its floor (570 or newer, CUDA 12.8) with no upper bound; the rule that matters is one GPU server per compute capability, picked by nvidia-smi compute_cap at install (a mismatch fails every proof in 12 s and the miner never notices): 0.3.21's kit design for main. The 10 GB tier unchanged (proves alone at 8.6 to 8.7 GB, never beside its miner; 12 GB fits both with 2 GB spare).
**Main (13:4x BST):** all three accepted: the cases rerun on the pin's binary; the 610 question closed (driver floor 570, no upper bound; the per-architecture GPU server is the prover rule for 0.3.21); 1982dbc7 is the driver-check tip for 0.3.21. The cut set and the clock stand; nothing further from main until the wipe line or a slip.
**The cases rerun started (the fleet, 12:42Z):** target c20-1 with c4459193 (in/igneumd-45be9b02d1b002f5, the string read back) on pool-1's kept 1026 datadir (its start recorded as a kept-datadir read on the pinned binary: the rewrite line, then the catch-up from 129,398 blocks, 15 to 20 minutes), a fresh c17-relay and c17-poison, the dc141409 form on the target's synced line with the byte-5 header shape expected; CASES END about 14:15Z (15:15 BST). The pool window's daemon on pool-1 (03457d96) since 12:40:39Z with its ten members.
**c4459193 hands pair held (the build-server lane, 13:41 BST):** native glibc 2.39, 381 s, rc 0, pairs with igneum 00249643: igneum-pow 0.2.0, rustc 1.99.0 by the tree's own pin. igneumd 57,628,576 B sha256 a80ed39caf885d314f97ce88863afcb307cbb4acc45a10828b2443a41e5d27d2 (string c4459193 twice, 6 igneum-pow/src/ paths, the N13 line present, needs GLIBC_2.39); igneum-miner 10,214,648 B sha256 70a5180f30fab73fde0b3f33cfd37d2d68899eab70290b86c924e02980b09afd (8 paths). Held at scratch bs0320/d-hands/release/. The seed pair from 13:41 BST, then the two Windows cross builds.
**c4459193's node gates, both PASS on its own binary (the node lane; built 12:33Z, sha256 45be9b02d1b002f5, string read back):** digest gate 12:33:27 to 12:35:05Z: thirteen fields a89be8a7 on both binaries, sixteen fields db9a85f9 refused with the line and no peer, the live file's digest eada4bda unmoved. Mixed-version gate 12:35:26 to 12:45:39Z: digest b0afb2ee on all five, 215 new and 314 old blocks accepted, 0 rejected, plain header version 2, counts equal at 312, 441 and 529 through both clean joins and the restart step (the new node restarted 12:43:08Z on its own datadir and resynced, where b7cc37e7 met the lock), no panic in any log. The node side's lines are complete; outstanding for the pin: the wipe canary on c19-1, the cases rerun on c20-1, the 12 GB settled-claim line.
**The Mac's binaries start now (13:4x BST):** on the node gates' pass the Mac builds its own igneumd and igneum-miner from c4459193 under the build lock, one at a time (the Mac rule), so the DMG follows the wipe line by minutes; a canary fail voids them.
**The cases rerun's target up (the fleet, 12:45:46Z):** c20-1 on c4459193 (sha 45be9b02d1b002f5 read back on the pod) started on pool-1's 0.3.17 copy with the N13 rewrite line and "state blob of layout 2 read and converted" once, panicked 0, catching up from 129,398 blocks: the pinned binary's kept-datadir read on a second pod. Relay j84mqzlai3yax9 and poison vnm1dlx8hkocq2 against it on its synced line; CASES END about 14:20Z (15:20 BST). c19-1's wipe IBD at 15 percent of the headers at 12:43Z, on the class. The Mac's node build started under the lock (scratch r0320/mac-build.sh: igneumd and igneum-miner, then the DMG).
**c4459193 seed pair held (the build-server lane, 13:48 BST):** zigbuild 2.35, 374 s, rc 0; igneumd 56,304,592 B sha256 4a2d8a8db0911a9236264ba7ef9ca8891e8a9df53f0892a695e176323cfaf415 (string c4459193, the N13 line, needs GLIBC_2.34 against the seed ceiling 2.35); igneum-miner 10,235,664 B sha256 18099e976b844653d2ccc7a2d37cef785d7537cbc3df042b073c7fcd8dd31fd0. Held at scratch bs0320/d-seed/x86_64-unknown-linux-gnu/release/. The c4459193 pair set is complete: hands a80ed39c / 70a5180f, seed 4a2d8a8d / 18099e97. The Windows cross of c4459193 from 13:48 BST, then 6a3432a3's, then the PC 2 gate job.
**Slip (the fleet, 12:5xZ; 13:5x BST):** the cases rerun's target starts from pool-1's kept copy 22,000 blocks behind the tip and this line walks the headers first (the morning's poison pod: 50 minutes on the same copy), so c20-1 syncs about 14:50 BST and CASES END lands about 16:10 BST, forty minutes past the checkpoint; no faster path. The rest holds: c19-1's wipe on the class (synced about 15:17 BST), p12-vast synced about 14:45 with the 12 GB line about 15:00 to 15:10. Put to main: (a) hold the pin to CASES END and publish about 16:15 BST, or (b) cut about 15:35 on the wipe line, the kept reads, the restart and the 12 GB line with the cases as a confirmation before the sweep's first box (the diff argument: the class byte is the only touch; the old hub took 235 byte-5 headers in the void run). The shipper's read: (a). The pods run on unchanged.
## 34. The 0.3.20 artefacts on the pin c4459193 (14:0x BST), staged for a one-step publish
**Main (13:5x BST):** (a) for miner-reliability: 0.3.21 takes both halves; 0.3.20's trees stay closed; its app half rebases onto the published 0.3.20 tree right after the publish. The report states what a 0.3.20 user still meets from MF-1 to MF-10 and what the pin removes (N13, N12 and the shutdown poll, the template latency memo).
**The shipper's box builds (scratch r0320, the vendor worktree at c4459193, app 4c89372a, igneum-pow 8c728ca3's tree):** fleet-native node (2.39) igneumd a8d08da5 57,628,128 B / igneum-miner 474273ce 10,214,648 B; hive class (2.31) igneumd b1c4e841 56,305,232 B / igneum-miner 4050c255 10,236,448 B; every igneumd carries c4459193 twice, no stub string, the N13 line; Windows cross igneumd.exe 49502cc7 52,569,088 B / igneum-miner.exe b4871b5d 11,219,968 B; the app's Windows exes igneum-app.exe b73120f1 3,956,224 B, igneum-ota-sign.exe 5355a5b6, igneum-prove-verify.exe 15998eab; the app's Linux binary igneum-app dd22a4ae 3,389,776 B. igneum-pow tests on the box with the amended packs: 61 + 7 + 4 + 19 + 2 + 7, 0 failed. The Mac, under the lock: igneumd b306baba 48,072,304 B and igneum-miner deb4d263 (c4459193 twice, the N13 line), the DMG Igneum-Miner-0.3.20.dmg 44,467,804 B sha256 73796c5febc20506 (build 202610071249, hdiutil VALID, the prover pair in).
**The workers with the Arc fix (built on the box from this tree's proto-opencl at 9088293a's content):** Linux 2.31 for HiveOS: igneum-worker-cuda c08e4694 6,761,464 B, igneum-worker-opencl 56cbe32f 312,304 B; Linux 2.35: cuda ed3dfe58, opencl 8ef81d8d; Windows igneum-worker-opencl.exe 618a2b10 530,432 B (mingw on the box, static, imports KERNEL32 and msvcrt only, the rewrite text present), in place of the 01:33 build (479,232 B, kept as igneum-worker-opencl-0133-old.exe in scratch); igneum-worker-cuda.exe unchanged since 0.3.19 (its sources untouched). Asked of the Intel lane: how bench d's exe (3c62470f) was built and whether it is the same source; if its exe differs the installer is rebuilt on it.
**HiveOS package:** igneum-hive-0.3.20.tar.gz 27 MB sha256 d9dd12dfe900bf9f (the 2.31 node pair, the 2.31 workers, version.txt naming c4459193 and 8c728ca3; no AppleDouble entries). Smoke in ubuntu:20.04 on the box (glibc 2.31): igneumd answers "igneumd 2.1.0" with c4459193 in the binary, igneum-miner its usage, both workers load and answer.
**Windows inputs and installer:** payload-inputs.zip published 13:59 BST (igneumd.exe 49502cc7, igneum-miner.exe b4871b5d, the three mingw runtime DLLs, igneum-worker-cuda.exe with nvrtc64_120_0 and nvrtc-builtins64_128, igneum-worker-opencl.exe 618a2b10, the 0.3.17 Linux prover pair 71bc2438 / 263bf4ce, which stands because the pre-push prover-pair-check read c4459193's evm-types equal to the exec pin's; IGNEUM_NODE_SRC = the vendor worktree at c4459193); windows.yml run 37625030022 dispatched 12:59:29Z on a255b095, the installer fetched with fetch-ci-artifacts.sh (no --deploy) when it lands.
**Staging plan at the publish (one step):** a scratch copy of the downloads folder, `IGNEUM_DLSITE=<copy> publish-manifest.sh --no-deploy --version 0.3.20 --mac <dmg> --win <installer> --override <the live sixteen-field file, digest eada4bda> --notes ...`, the hive tar through publish-public.sh, the floor-moved file beside the release with its digest and the one-line diff (not published); then the deploy step on CASES END, PC 1 first, the lock lines carrying "sweep complete before 13 October 09:00 UK", hands and seed last by the build-server lane.
## 35. the project lead's orders (15:1x UK by main's stamp; 14:0x BST on `date`): the floor file publishes, the sweep in waves, 0.3.21 tonight
(1) The floor-moved file publishes WITH 0.3.20: floor = live DAA at the publish + 604,800 rounded up to the epoch (3,600), window unchanged, the digest read back on c4459193's binary, the one-line diff against eada4bda in the publish record. Dry run at 14:07 BST, DAA 286,227: floor 892,800, digest 43725627 on c4459193 (the live file reads eada4bda on the same binary). Consequence: a node on the old file refuses a node on the new one as a peer, so an unswept box is isolated, not forked, until its turn. (2) Fast: the cut word on CASES END, everything staged so the publish is one step; the sweep in parallel waves as the lock lines allow (PC 1 first by the app's poller, then PC 2, the Mac, the seed, the hands and the fleet), each box read back by commit string; with the moved file the seed, the hands and the fleet move in the first wave. (3) Build-2 for anything that still builds. (4) The Windows OpenCL worker: the Intel lane's exe af53f194 (built on build-1 by build-windows.sh's line, static, KERNEL32 and msvcrt only; the Arc self-test PASS 96 of 96 on v4-devnet-epoch0 867dbc45cfb36b4d, the mx8 control and the live pack; 11.008 MH/s on v4) ships in place of the shipper's own box build 618a2b10 (same source, not Arc-tested); the inputs re-pushed with it and windows.yml re-dispatched: run 37625870876 (37625030022 and 37625590021 cancelled). (5) Standing rules from 0.3.21 (release-rules.md 4a, 4b, 5 at a36298c4): every gate starts on every candidate as it builds; warm pods per gate class; the sweep in waves. (6) 0.3.21 starts the moment the sweep ends: sub-version 2 (07a809a7, byte 7, pending the F8 census), miner-reliability both halves, driver-check 1982dbc7, the fork gate from horizon-node, the under-12 GB prove-instead routing, the per-architecture prover server; the tree staged tonight, the first gate about 19:30 BST (app, build-2), the node candidate's parallel gates from its first binary.
**The sweep's wave plan (the fleet, 14:1x BST), accepted:** from a 16:15 BST publish, each box's move = override.json replaced, c4459193 put as in/igneumd-45be9b02d1b002f5 with the sha asserted, the node killed by its kill file and started on its kept datadir (the rewrite line once, about 80 s to the tip), read back by the commit string in its log line and the digest from its handshake line, synced, the miner back; the prover pair under bin-0320 and the prover restarted by file after synced. Wave 1, 16:15 to 16:21, read back by 16:23: hub-1 first, pool-1 (its daemon stopped first), the eight heaviest voters (p1-5090, p1-4090, p2-4090-3, p2-4090-1b, p2-3090-1 to 4), with the seed and the hands in the same minutes from the build-server lane: more than two thirds of the live table's weight on the new digest at once; the first lock on the new side (the folded certificate on hub-1, about 16:27) is the line between waves. Wave 2, 16:27 to 16:32: p1-a5000, p1-4070, p2-4070-1, p1-3080; the lock line by 16:35. Wave 3, 16:35 to 16:40: the Devnet 2 six (dn2-seed first). The prover roll behind each wave's synced reads, 16:25 to 16:50. PC 1, PC 2 and the Mac on their pollers. Sweep end about 16:42 BST; provers paid on the pair by about 16:50. Risk named: a refusing datadir would cost a 98-minute resync (today's two kept reads say none will). The go = the override file with its sha256 and digest at the cut word. Warm pods for 0.3.21 (rule 4b): w-target and w-relay syncing on the kept copy, w-poison when a 3070 frees, p12-vast kept; USD 11.52 a day.
**0.3.21's clocks (14:1x BST):** the F8 census on byte 7 (07a809a7) lands about 15:00 BST; the pass line is all 64 seeds under 1.2x of the window model plus the hash lane's suite and G1; if it fails or slips past 20:00 BST, 0.3.21's node ships byte 5 again with the re-pin dropped. The node lane's dry merges into c4459193 are clean (miner-reliability-20 f067f7c1 one file, horizon-node 6eb21fc9 six commits on finality.rs); at the sweep-end word the branch commits, suites on build-2 about 16:50 to 17:05 BST, the first candidate binary on build-1 about 17:10 BST with every gate started from the build. App side: release-0.3.21 on origin at 0b75cf52 (release-0.3.20's closed tree, driver-check 1982dbc7's three commits, the 0.3.21 version strings, master 819d536b merged with its 51-check gate; the hands script's pgrep in the bracket form); the reliability lane rebases its app half by cherry-pick onto it as miner-reliability-21 (tip fb3a4a40, gate on build-2 running). Still to write on the app side: the under-12 GB prove-instead routing and the per-architecture prover server.
**The 12 GB settled-claim line, first half (the fleet, 13:23Z):** p12-vast (RTX 3060 12 GB) on c4459193 (45be9b02d1b002f5, the string read back) on pool-1's kept copy from 12:36:35Z (first-try start, the copy already rewritten by b7cc37e7's read), synced 13:22:47Z; box-prover (cf335220, the settled filter) started 13:23:01Z with the sm_86 server; first claim "segment 164198..164205 (8 shards, fresh) margin=414 tip=287564 settledNumber=0x28185 settledDaa=0x46317 settledBy=finality candidates=5 rank_by=fnv", pair ok. The floor read: settled number 164,229 at settled DAA 287,511 by finality, 31 blocks above the claimed segment, the claim behind the floor as the rule wants. The paid record by the pod's key is the second half, about 3 to 10 minutes.
**Main (14:3x BST):** the project lead withdrew the Harmony overlay; apps-harmony is parked. 0.3.21's UI branches in order: gpu-logos (a414b6bdc81d348d8, takes the Prove switch fix), ui-overlap-fixes (a7aa33bf8ae230f46; its wallet half is the wallet's cut), scene-parity (a75edb2ef8015f21c; live-dag.js and proof-core.js, the blank-canvas fix first); each instructed: rebase onto release-0.3.21, push as <name>-21, UI tests and the app gate on build-2, tip and lines to the shipper; 19:30 BST stands. A site rebuild lane ships through the site gate on master, outside 0.3.21. The Windows installer run 37625870876 failed its inputs check on the stale node-source pin (e69e8a39 at a255b095 against the payload's c4459193); the pin pushed (4ae4a54e) and run 37627661560 dispatched 14:20 BST.
## 36. CUT BLOCKER (the fleet's 12 GB line, 13:25Z; 14:3x BST): the 0.3.20 node reads the proving ids as unknown
On p12-vast, c4459193 (45be9b02, string read back) over pool-1's kept copy with the hub's override-16.json: the node's start lines read "[igneum-exec] proving v1: ... shard program id unknown, aggregator id unknown", then "verifying keys embedded: shard program id 0x2b1a81cb... aggregator id 0x474678f3..."; hub-1 on 5899f603 prints the ids in the proving v1 line. The node builds its segment statement with zeros for the ids and the pair's proof is refused: "RESULT seg 164198 FAILED: statement differs from the node's at hex offset 472 (lengths 616 vs 616); ours ...2b1a81cb413236cf... node ...0000000000" (the claim behind the settled floor was right: settledNumber 0x28185 by finality, margin 414; the 3060 proved 8 of 8 shards in 135.9 s at 8,487 MiB). Per tier: swept as it stands, every standing prover stops paying from the sweep minute and the feed goes to zero while the miners run on; a 0.3.20 home prover earns nothing. The pin HOLDS. Asked of the node lane: the cause (where 5899f603 takes the ids from, which of the five commits since dc141409 loses them), a fix with a test (a node on the live override file reports the ids; a zero-id statement the known-failed shape) as c4459193's child, its time; the fixed binary takes the full gate set from its build (rule 4a; the warm pods are up, about 80 minutes). The sweep plan, the floor file, the installer and the app tree are unaffected. Earliest publish by the shipper's read: about 17:30 BST if the cause is small. Main told.
**Main (14:3x BST):** the hold agreed (a prover cannot be split from its node); the pin waits for the ids fix as c4459193's child and takes the full set from its build; new rule 4c (release-rules.md 706b46d3): on every candidate a node on the live override file reports both proving ids on its start line and a prover's first statement is accepted, a zero-id statement the known-failed shape, run beside the kept-datadir read on the warm pod; the fleet has it. Publish on green with the clock; the 16:00 report carries the slip and the cause.
**The Windows installer (14:26 BST):** run 37627661560 on 4ae4a54e success; Igneum-Miner-Setup-0.3.20.exe 63,025,372 B sha256 45b2f3fb54f40f83 and igneum-windows-app.zip 91,038,211 B; inside the zip igneumd.exe 49502cc7 and igneum-worker-opencl.exe af53f194 (the Arc-proven worker), as pushed. fetch-ci-artifacts.sh copies into the live downloads folder by design; both files moved out to scratch r0320/win at once (rule 2: the live folder changes only in the deploy step). The installer carries c4459193's node: if the ids fix lands as a child commit, the node's Windows exe, the payload inputs and this installer are rebuilt on it (about 35 minutes: the cross on the box, push-inputs, one windows.yml run), as are the native, hive, seed, hands and Mac binaries and the DMG.
**The blocker is the start environment (the node lane's read of the code and hub-1's process, the fleet's restart on the pod, 14:3x to 14:4x BST):** the statement's ids come from the override object's two fields, then IGNEUM_PROOF_PROGRAM_IDS, then the verifier host's `--mode id` under IGNEUM_PROOF_VERIFIER; the live file carries no id fields; every standing box's supervisor sets the host (hub-1 included), and the app sets it whenever it finds igneum-prove-host. The pod's node was started bare for the kept read; 5899f603 bare reads the same. Restarted on p12-vast with the env at 13:30:28Z, the same c4459193 prints both ids and the 3060's first statement is accepted at 13:33:56Z (8 of 8 shards); the paid half then lost the claim race ("segment already paid; end to end 113.1 s", p2-4070-1's shape) and runs on. The swept fleet keeps the env through box-standing.sh. The node lane's child commit (`resolve_program_ids`: the override first, then the env or host, then the embedded keys; test known-failed first) makes a bare node self-sufficient; it is the fix for a hand-started node without the host, which has no prover to submit to it. Put to main: pin c4459193 as it stands (its gates landing: c19-1 synced about 14:40 BST, c20-1 synced 13:33:08Z and the cases form running, CASES END about 15:50 BST), publish about 15:55 BST, the ids commit as 0.3.21's first node commit with rule 4c's gate on every 0.3.21 candidate (ids-gate.py, the fleet's README item 19); or the child now with the full set again, publish about 16:35. The shipper's read: pin c4459193. The pool window ended 13:30:23Z with 0 shares (the daemon's template RPC timed out from 12:54Z; the members' prepare stall at the epoch boundary): the pool lane's, not 0.3.20's.
**Main (14:3x BST): THE PIN IS c4459193 as it stands.** The ids commit (55768f88, `resolve_program_ids`, the exec suite 32 with `a_bare_node_resolves_its_program_ids_from_the_embedded_keys`) is 0.3.21's first node commit, with the ids gate on every candidate from now. Publish on CASES END about 15:55 BST. The sweep: the project lead is taking PC 1 offline for cable work (a Mac mini comes online), so PC 1 is not first and not in the waves; its app updates on its poller when it returns, its lock line "PC 1 offline at the sweep, updates on return". First wave: this Mac, the fleet, the seed, the hands, and PC 2 if it polls. The Mac mini is a new machine and takes the 0.3.20 DMG as a fresh install when the project lead has it up. The shipper's parallel builds on 55768f88 (box and Mac, started 14:36) stopped by pid; the vendor worktree back on c4459193; the pinned Mac binaries and the DMG 73796c5f kept.
## 37. The pin's canary lines on c4459193 (c19-1; the fleet, 14:4x BST)
sha256 45be9b02d1b002f5 asserted at the put, igneumd/2.1.0-c4459193 in the node's log, digest eada4bda8aa8368c. **The wipe (decisive):** IBD from the pruning-point proof 12:38:55Z, synced 13:35:50Z (156,046 blocks, 3 peers), 57 minutes; the IBD-end line "finality time: bodies 3864115 ms over 156238 blocks, virtual 373190 ms over 16695 changes, weight tables 48145 ms over 5595 tables (53495046 blocks walked), signatures 607815 ms over 473073"; proof-asking 0, carried records 0, pings 0, idle drops 0; the exec poller 682 calls, 0 errors, the first non-null answer at the synced line (exec synced true). **Mining** from 13:36:00Z: at +166 s 34.2 MH/s, mined 16, accepted 12, rejected 0, got_reject 0, wrong_version 0, isSynced true at the tip on every read (the dc141409 flap is gone: 8097d600's fix). The ten-minute and hub reads about 13:46Z, the restart read about 13:50Z. **The kept read (rule 4):** pool-1's 0.3.17 copy on the same pod, first start 13:38:14Z with the N13 rewrite line and the finality blob converted (1,747 locks), no panic; second start 13:38:33Z first-try, no panic: PASS, c20-1's shape. **The cut word:** c20-1's CASES END about 14:50Z (15:50 BST; the poison pod's IBD from the relay started 13:38Z).
**The seeds (main's ruling, 14:4x BST):** the three testnet seeds run `--testnet --netsuffix=1` with no override at height 0 on 1c19441d; c4459193's seed-class build prints testnet digest 9537868d against 1c19441d's b7d8c915, so a swap would move the pre-go testnet's digest. Ruling (b): the seeds stay on 1c19441d and the RPC blocklist stays; the testnet genesis re-cut replaces the seeds' line at the go on the project lead's word. The devnet floor file has nothing to do with the seeds (the shipper's earlier order corrected). Wave 1 is the Mac, the hands and the fleet, PC 2 on its poller.
**PC 2 is back and the Windows kept-datadir gate PASSED (13:38 to 13:39Z):** PC 2's 10:46Z silence was the whole PC losing power (Kernel-Power 41, 6008, no bugcheck; two such events today), not the 0.3.19 update-now, which had returned cleanly at 10:34:38Z; it polls again since 13:38:32Z. run-20261007-125433 on it: c4459193's igneumd.exe (df19315d) on a robocopy of the app's datadir (1,147 MB), "verdict c4459193 (pass): the N13 line seen; consensus built and the servers started", the copy removed, 0 leftover nodes; 6a3432a3 on its own copy died at virtual_state.rs:250 with InvalidBoolEncoding(20), the Windows shape of the known-failed read. A datadir cut off mid-write twice by power loss started clean on the pin. The reliability register's MF-11 is renamed to the machine going silent with nothing reporting it.
**The testnet digest move, for the record (the build-server lane, 14:4x BST; not a 0.3.20 item):** between 1c19441d and c4459193 the testnet's genesis, base params, finality table, fee table and pow schedule are unchanged; five additions enter the digest on c4459193 and are the whole b7d8c915 → 9537868d move: finality_leave_activation_daa 0 with finality.leave_delay 3,600; program_class_v3_activation_daa 0 (unconditional); program_class_v4_activation_daa 0 (the signal window 0 stays out); pow_genesis_dataset_log2 28 (unconditional); emission EmissionSchedule::TESTNET_1 (launch_rate 100 UNIT, ramp 90 days from 10 percent, monthly steps with step_decay_q32 4,172,697,914, tail 100 bps a year). Everything else new sits at never and stays out; dns_seeders is not a digest field. No override pins the old value (1c19441d has none of the five fields), so every testnet node swaps together or not at all: the testnet-go checklist's digest line reads 9537868d on this line, to be re-read on the genesis re-cut's binary at the go.
**PC 2 out of the waves too (main via the Counter lane, 14:5x BST):** the project lead is taking PC 2 down for cable work (PC 1 is back, but it is his desk and no job goes to it); both PCs update on their pollers on return, their lock lines "offline at the sweep, updates on return". The sub-version-2 Windows G1 completed on PC 2 before it went down (13:46:36 to 13:46:50Z, exit 0, eight of eight fingerprints equal to the Mac's). Wave 1 is the Mac, the hands and the fleet.
**c19-1's last pin lines (the fleet; FORM END rc 0 at 13:50:53Z):** mining at +666 s 34.3 MH/s, mined 66, accepted 66, rejected 0, got_reject 0, wrong_version 0, isSynced true at the tip on every read, the exec follower moving; the hub holds 41 blocks by c19-1's key 2c7cc291d38579e0 in its last 700, rejects naming the pod 0; the restart on the kept datadir 13:47:15Z: the stop and the start inside seven seconds (09124180's poll proving itself against b7cc37e7's LOCK death), synced again 13:48:39Z (157,048 blocks, 4 peers), 84 seconds after the kill, isSynced true on the first read and never false after; 109 templates in the read, max template_ms 3,432, "template fetch timed out" 0; the kept read passed at 13:38Z. Every line the rule reads is in from c19-1; CASES END from c20-1 is the one left (about 14:50Z). c19-1 stays up until the shipper's word. p1-5090 ran the sub-version-2 G1 meanwhile (eight of eight equal to the Mac's) and is back under its supervisor.

View file

@ -0,0 +1,48 @@
# Release 0.3.21 (the cut after 0.3.20; staged 7 October 2026, 14:2x BST)
Tree: release-0.3.21 on origin, from release-0.3.20's closed tree (the 0.3.20 cut: pin c4459193, igneum-pow 8c728ca3 at byte 5, the floor-moved file). Rules: docs/plans/release-rules.md (4a every gate on every candidate from its build; 4b warm pods per gate class; 5 the sweep in waves).
## 1. What rides it (the project lead's list, 7 October 2026)
| Item | Owner | State |
|---|---|---|
| driver-check 1982dbc7 (the driver table measured: NVIDIA 617.42, AMD 26.9.2 with the amd.com referer, Intel 32.0.101.9034; the install thread with one elevated prompt; `--drivers` on the manifest) | the Intel lane | in: 39187097, afe4e2eb, 416d1f5d |
| miner-reliability, app half (miner-reliability-21 aa2bb07e, 18 commits: the readiness gate and the ladder, the load clock, the hold owner, the orphan sweep, the FAULT lines to the intake, the signed `cards` job kind, execrpc with `probe`; box 226 + 32 + 8, UI 49) | the reliability lane | merged into release-0.3.21 |
| miner-reliability, miner half (miner-reliability-20 f067f7c1: STATUS while waiting for a template, identities capped from the node's template time, the fetch back-off) | the node lane | staged on release-0.3.21-node at the sweep-end word |
| sub-version 2 (the hash lane's 07a809a7, byte 7, id a788661687db4bb3), pending the F8 census (about 15:00 BST; pass = all 64 seeds under 1.2x of the window model plus the suite and G1 lines); if it fails or slips past 20:00 BST the node ships byte 5 again with the re-pin dropped | the Counter lane, the node lane | held |
| the fork gate from horizon-node (6eb21fc9, six commits on finality.rs) | the node lane | dry-merged clean into c4459193 |
| the node's own rust-toolchain.toml | the node lane | in place on its worktree |
| the under-12 GB prove-instead switch (section 2) | the shipper | to write after the 0.3.20 publish |
| the per-architecture prover server (section 3) | the shipper, the fleet | to write after the 0.3.20 publish |
| pool-finish 434e8b9b (the pool crate: the TLS 1.3 member port with the authorize binding, the open pool, pool.md 10; fork pool-finish-node b0444f51 off 8097d600: the pool-mode miner, the IGNS/IGNP codec, pool_split_activation_daa at never on every network, so the digest does not move) and pool-mf-row 343dd83b (MF-11 in the register) | the pool lane | main's word 14:2x BST: rides 0.3.21 with the split switch at never; if it reds a 0.3.21 gate or costs more than one rebase round it moves to 0.3.22 |
| apps-harmony (the builder's redesign, Overview as the landing tab, the Prove switch fix), then gpu-logos on top of it, then ui-overlap-fixes, then scene-parity, in that order as they land, each with its own gate line | the UI lanes | to come |
| MF-11 (the app not returning after update-now; PC 2 silent since the 0.3.19 update-now at 10:34Z, its relay agent dead with it, so no job or task reaches it): the register row, the UI lane owning the cause, the relay agent as a service surviving the app as the restart path | the reliability lane, the UI lane, the relay lane | the row asked |
| the ids commit 55768f88 (`resolve_program_ids`: the override's fields, then the env or the host, then the embedded verifying keys; the exec suite 32 with the bare-node test known-failed first; built 13:35Z sha 279b1b690e854fc9): 0.3.21's FIRST node commit by main's word, with the ids gate (rule 4c) on every candidate | the node lane | on the mirror; its gates run under rule 4a |
| the 0.3.21 node order (main, 14:4x BST), each with its own gate line, one rebase round or 0.3.22: on c4459193, 55768f88; miner-reliability-20 f067f7c1 and the late-join commit 70e4601e; pool-finish-node b0444f51; horizon-node 6eb21fc9 (the deep fork-choice fix, fork_gate never set); peer-directory-node db28d331 (the peer directory prototype, peer_directory_activation_daa at never, the IGNP announce only under --announce); the sub-version-2 re-pin held for the census. The horizon lane rebases its two onto c4459193 at the sweep-end word | the node lane, the horizon lane | staged tonight |
| pool-finish-21 at 5e557171 on origin (release-0.3.21 5d153892 plus the ten pool commits, the tenth d5265fb3 the `--template-parallel` fix, default 2: the ten-member window's daemon queued ten template calls on one gRPC connection past its own timeout, 1,891 lines, no job for 36 minutes; site/miner.html carries the pool-fee sentence); igneum-pool 28 at 5e557171 on build-2, the app gate 226 + 32 + 8 at 54dd4c06 (the last commit touches pool/ and docs only); merged after the three UI branches with the app gate rerun at the merge | the pool lane | ready to merge |
| 0.3.20's version strings moved to 0.3.21 | the shipper | 0d7b9bd0 |
| master 819d536b merged (the 51-check gate; build-remote routes by load; the hands script's pgrep in the bracket form) | the shipper | c83ca904, 0b75cf52 |
## 2. The under-12 GB prove-instead switch (design, main's routing)
The fact (the fleet's 3080 hour, 7 October 2026): a 10 GB card completes the compressed step alone at the default threshold (8,642 to 8,729 MiB of 9,885) and never beside its miner's 1,547 MiB; no lower threshold fits (the server dies before the compressed step at 524288 and below). 0.3.20 ships the rule "proving needs a 12 GB card; mining continues" (provedefault::prove_refused_under_12gb). 0.3.21 adds the choice: a Settings switch `prove_instead` (off by default) that, on a machine whose only NVIDIA card is under 12 GB, stops that card's miner while the prover runs and restarts it when the prover is off; the tile sentence becomes "proving instead of mining on <card> (10 GB holds one, not both)"; the refusal line stays when the switch is off. Engine: the prover loop's refusal rung reads the switch; when on, it takes the miners hold for that card (the hold owner from miner-reliability) for the prover's lifetime. Test: the refusal lifts with the switch on and the card's miner reads held; a 12 GB card beside it never triggers the hold. UI: the switch under Settings > Proving with the sentence; the view test.
## 3. The per-architecture prover server (design, main's rule)
The fact (the fleet, 7 October 2026): sp1-gpu-server is built for one card's compute capability (sm_86 Ampere, sm_89 Ada, sm_120 Blackwell); a server for the wrong architecture fails every proof in 12 s with "CudaRustError: named symbol not found" and the miner never notices; the 610-series driver proves (the driver floor stays 570 or newer, no upper bound). The app's WSL2 setup (setup-wsl.sh) builds the server on the machine, so it matches by construction; the risk is a copied or stale server. 0.3.21: (1) the prover's setup records the compute capability the server was built for (nvidia-smi --query-gpu=compute_cap) beside the binary; (2) at every prover start the engine reads the card's compute_cap and the record, and refuses with "the prover was built for sm_NN; this card is sm_MM; Set up rebuilds it" when they differ (MF-10's capability check in prover.rs, owed by the reliability lane's register); (3) the fleet's kit ships one server per architecture under bin-<arch> picked by compute_cap at install (the fleet lane). Test: a record and a card that differ refuse with the sentence; equal ones pass; no record passes with a warning line.
## 4. Gate clock (rules 4a and 4b)
App: the gate on build-2 about 19:30 BST (`build-remote.sh --box 2 -- test --release` from app/igneum-app), the UI tests, the Windows cross on build-2, the reliability injector's eight steps on a one-shot pod (the fleet, about 45 minutes, after the sweep). Node: the branch commits at the sweep-end word (about 16:45 BST), suites on build-2 about 16:50 to 17:05, the first candidate binary on build-1 about 17:10 with the digest gate, the mixed-version gate, the kept start, the cases and the wipe all started from the build on the warm pods; the pin about 80 minutes after the last candidate builds. The Mac's own binaries and the DMG under the lock after the pin; the staging in a scratch copy; the publish on green with the clock time.
**gpu-logos-21 in (14:4x BST):** ba5d9cea, four commits on 8ed08dcf (1054616c the vendor marks, bca70dcd the captures, c98f2666 the Prove switch fix as its own commit: the DAG legend's bare .lg and the inspector's bare .track scoped to .legend .lg and #i-track, the switch's size class its own, the native input hidden the accessible way under a positioned label; ba5d9cea the shared marks module for the site and the re-taken shots). Lines: build-2 app gate 226 + 32 + 8 at c98f2666 (ba5d9cea touches nothing in the crate); UI 72 pass (view 41: 5 marks, 3 switch known-failed first, 1 module drift); pre-push 52 green. Captures docs/plans/gpu-logos-shots/ 01 to 12. Merged into release-0.3.21.
**earnings-tidy-21 in (14:4x BST):** a436c6f0 on 5d153892 (c7ec4573 the Earnings card tidied by the project lead's order "remove all the type a price stuff": the £ per IGN input, Use it, the remembered price, the price, balance-in-pounds, share, IGN per kWh and cost-of-hash lines gone; the card is the day rate with its reason, blocks and IGN this run, a row of three (weight, electricity, lifetime), the dev-fee switch; View.earningsLines(s, chain, pence, now, minor), FIELDS.price gone; a436c6f0 the captures docs/plans/earnings-tidy-shots/01 to 03). Lines: build-2 app gate 226 + 32 + 8; UI 75 (three new, known-failed first); pre-push 52. Merged into release-0.3.21.
**miner-reliability-21 at cbd6f3a4 in (14:5x BST):** four more commits merged: 4a28eb59 MF-10 (provedefault::server_mismatch and is_arch_failure, the test sm_86 on 8.9 and 12.0 refused, 8.6 and a fat binary pass; prover.rs server_arch_check reads the card's compute_cap and the server's sm_ words at the first probe, refuses the GPU path with the reason on the tile and a FAULT class=prover-arch line; "named symbol not found" logged as the same class) and MF-8's injector step (app-run.mjs --node-old: the previous release's node writes 150 blocks and stops, the new node opens the datadir and reads synced in one start); a751d391, 6a51c6c1, cbd6f3a4 the register (MF-11 the machine going silent with nothing reporting it, MF-12 the pool member stalled at an epoch boundary carried over from the pool lane, the rows moved from owed to code). Lines: build-2 app tests 227 + 32 + 8 on 4a28eb59; pre-push 52; UI 49 (the lane's run; the tree's UI tests re-run at the next merge). The 0.3.21 injector run: the payload staged (the app 786c3d37 from 8ed08dcf, node c4459193, miner 65d0837a, the old 0.3.17-line igneumd for --node-old), nine steps about 50 minutes on the fleet's pod after the sweep. This closes section 3's app half (the capability check); the per-architecture kit is the fleet's.
## 5. The node line's first candidate: 55768f88 (the node lane, 14:5x BST)
sha256 279b1b690e854fc9, the string read back, pairing 8c728ca3 at byte 5. The digest gate 13:35:41 to 13:37:19Z PASS (thirteen fields a89be8a7 on both binaries, sixteen fields db9a85f9 refused with no peer, the live file's digest eada4bda unmoved); the ten-minute mixed-version gate beside the 5899f603 pair 13:37:40 to 13:47:52Z PASS (digest b0afb2ee on all five, 223 new and 381 old blocks accepted, 0 rejected, plain header version 2, counts equal at 319, 486 and 604 through both joins and the restart step, no panic). The fleet's set on the same binary (the kept start with the ids gate, the cases on the warm set, the wipe) runs under rule 4a. Staging at the sweep-end word in main's order: f067f7c1 and 70e4601e, then b0444f51, then the horizon lane's rebased 6eb21fc9 and db28d331, then the re-pin on the Counter lane's word; suites on build-2 and the digest read after every merge; the next candidate's binary with every gate from its build.
**scene-parity-21 in (14:5x BST), ahead of ui-overlap-fixes-21 in the order because it was green on the exact tip (50ffa562) while ui-overlap-fixes still rebases:** 20c9153b (0d75bc1f the shared scene/ folder and live-dag.js 2.0.3, c1334faf the app side, 20c9153b the parity harness and its gate line). Lines: build-2 app gate 228 + 32 + 8; UI 72; pre-push 56 with four new chain-scene checks (sync byte-equal, the paint-on-push known-failed test, the feed contract, the parity render on build-2: home fold = /live = app Inspect at T+0, +2, +4 s). live-dag.js 2.0.2 → 2.0.3 (paint on every push whatever the visibility, the blank /live fix; the phone rule on the viewport width); proof-core.js unchanged 2.0.0; the site's copies moved on master f7743534 (igneum.network/live-dag.js reads 2.0.3); the source is scene/live-dag.js, both copies written by `node tools/scene/sync.mjs`, the gate refuses drift, so the 0.3.21 pack carries the 2.0.3 bytes. Users see: the light theme's ember at the brand package's #D0420D, the chain feed asking 300 s, Inspect 420 px tall (seven lanes), the phone layout no longer frozen at launch width, api/live keeping the key id in `miner`. Captures: build-2 /srv/builds/scene-parity/igneum-wt-scene-parity/_out and scratchpad/scene/parity-out.

View file

@ -0,0 +1,20 @@
# Release rules (the shipper's standing rules, as main set them)
Every cut of the Igneum Miner app and its node runs under these. The dated plan for each cut (docs/plans/release-<version>.md) records how each rule was met.
1. **Versions are three-part.** A hotfix takes the next number; the feature tree moves up. The app's parser returns None on a fourth part.
2. **Nothing is staged in the live downloads folder.** Stage in a scratch copy with `IGNEUM_DLSITE=<copy> publish-manifest.sh --no-deploy`; the live folder changes only in the deploy step. `publish-jobs.sh --deploy` is jobs-only. No lane removes a scratch directory it did not create.
3. **The deploy gate is a full canary on the fleet's pods:** the fresh join through the headers proof on a WIPED datadir (decisive), synced, ten minutes mining, the hub holding a block, the relay and poison cases. Main may call the deploy on the decisive read plus a diff argument.
4. **The kept-datadir start is a named gate in every release (main, 7 October 2026, after ledger N13).** Every canary runs both a wiped datadir and a KEPT one: a copy of a standing box's datadir from the live release, the pinned binary started on the copy on a scratch pod, "synced" or the store's rewrite line as the pass. The Windows shape too: the pinned Windows node once against a copy of PC 2's datadir (PC 2 only, never PC 1) before PC 1 gets the build. The miner-reliability register carries it as its own fault class. Why: every node build from 10db4b61 died at start on a kept 0.3.17 datadir (bincode ignores serde defaults) and no canary saw it because every canary wiped.
4c. **The proving ids gate (main, 7 October 2026, after the 0.3.20 blocker).** On every candidate, a node started on the LIVE override file reports both proving ids (the shard program id and the aggregator id) on its proving v1 start line, and a prover's first statement against it is accepted; a zero-id statement is the known-failed shape. It runs beside the kept-datadir read, since both share the warm pod. Why: c4459193 read the ids as unknown, built statements with zeros and refused every proof; a sweep would have stopped every prover's pay.
4a. **Every gate starts on every candidate the moment its binary builds, never after the pin (the project lead, 7 October 2026).** The digest and mixed-version gates, the kept-datadir start, the relay and poison cases and the wipe canary all begin on each candidate binary as it lands; a struck candidate's runs are stopped and its successor's begin. The post-pin wait is then the longest single form (about 80 minutes, the wipe), not the sum.
4b. **Warm pods per gate class (the project lead, 7 October 2026, ordered to the fleet lane).** The fleet keeps synced pods warm for each gate class so a case form's target starts at the tip (a kept copy of the live line, caught up), never from a kept copy far behind it; the wipe canary is the only full IBD in the set.
5. **Rollout in waves, each box read back (the project lead, 7 October 2026, replacing one-box-at-a-time):** PC 1 first, then PC 2, the Mac, the seed, the hands and the fleet in parallel waves as the lock lines allow; a lock line from the hub between waves; hold if the frozen table's signed share reads under 75; every box read back by its commit string. When the publish moves the consensus floor (a new digest), every 0.3.x node on the old file refuses the new ones as peers until it is swept, so the seed, the hands and the fleet move in the first wave with the apps' pollers, not last. Every lock line of a sweep that replaces nodes carrying a consensus floor names the date the sweep must finish (0.3.20: before 13 October 2026 09:00 UK).
6. **Read-back is by commit string plus digest plus engine:** on 0.3.18+ nodes igneum_getNodeInfo powEngine must read "igneum-pow" ("stub" = FAIL); on earlier trees `strings igneumd | grep -c igneum-pow/src/` above zero. The miner embeds no commit string; its pairing is the build line and the sha.
7. **igneum-pow pairing:** a fork build takes igneum-pow by path from the igneum worktree it sits in; build each node tree inside its own app worktree whose igneum-pow is the pinned tree; the pairing log line names it. Master's build tools need rust-toolchain.toml in the tree (the app tree's pin applies to a vendor worktree under it; a standalone node checkout is unpinned until the node line carries its own file).
8. **glibc classes:** HiveOS 2.31 (`--ship hive`, smoke in ubuntu:20.04 on the box), seeds and generic 2.35 (`--ship seed`), fleet 24.04 boxes native 2.39.
9. **The Mac builds only the macOS binaries and the DMG,** one at a time under the build lock; every other build, suite and the Windows cross-build runs on the box or a PC; the app gate is `build-remote.sh -- test --release` from the crate dir.
10. **A pin is green on its own suites and gates.** Lines taken on one binary carry to another only when the code is byte-identical, stated in the tip. No known-red pins: a stale test takes a test-only commit on top.
11. **Kill by pid, never by name,** on the shared Mac; a merge worktree never checks out master.
12. **The Discord card only when every platform is live.** Live manifest changes beyond the binaries (a moved consensus floor) go out only on the project lead's explicit word, staged beside the release with their digest and a one-line diff.
13. **Ship on green:** no calendar waits; when the gates are green, publish and state the clock time (UK). Checkpoints are for slips, not for waiting.

Binary file not shown.

After

Width:  |  Height:  |  Size: 224 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 225 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 238 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 240 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.2 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 586 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 597 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 478 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 478 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 367 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 364 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 549 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 548 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 549 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 547 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 332 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 332 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 277 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 276 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 590 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 605 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 518 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 528 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 517 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 516 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 309 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 311 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 381 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 385 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 263 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 266 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 335 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 335 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 276 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 276 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.2 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.2 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 558 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 586 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 328 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 336 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 327 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 337 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 345 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 346 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 374 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 372 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 488 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 489 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 396 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 398 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 613 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 634 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 393 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 416 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 407 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 407 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 350 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 351 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 677 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 682 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 477 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 285 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 225 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 223 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 244 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 242 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 426 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 425 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 333 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 331 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 396 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 393 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 396 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 392 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 571 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 201 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 427 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 451 KiB

View file

@ -37,11 +37,12 @@ CAP_NODE_BRANCH="${CAP_NODE_BRANCH:-}"; CAP_REPO_SHA="${CAP_REPO_SHA:-}"; CAP_FO
# ---- the build-slot and measure hold the layer yields to --------------------------------------------------------------
# a slot or the measure file is "held" when flock -n cannot take it (another process holds the fd). The same probe the
# collector uses (tools/workers/collect.mjs flockHeld). Returns 0 (true) when ANY build slot or the measure hold is taken.
# collector uses (tools/workers/collect.mjs flockHeld). Returns 0 (true) when ANY build slot, the quiet hold or a core lease is taken.
cap_build_active() {
local slots k f
slots=$(cat "$LOCKS/slots" 2>/dev/null || echo 1); [ "$slots" -ge 1 ] 2>/dev/null || slots=1
if [ -e "$LOCKS/measure" ] && ! flock -n "$LOCKS/measure" true 2>/dev/null; then return 0; fi
if [ -e "$LOCKS/quiet" ] && ! flock -n "$LOCKS/quiet" true 2>/dev/null; then return 0; fi # 7 Oct 2026: the quiet class replaced the measure file
for f in "$LOCKS"/core-*; do [ -e "$f" ] && ! flock -n "$f" true 2>/dev/null && return 0; done
for k in $(seq 0 $((slots - 1))); do
f="$LOCKS/build-$k"
[ -e "$f" ] || continue
@ -51,7 +52,8 @@ cap_build_active() {
}
cap_hold_reason() {
local slots k
if [ -e "$LOCKS/measure" ] && ! flock -n "$LOCKS/measure" true 2>/dev/null; then echo "measure hold"; return; fi
if [ -e "$LOCKS/quiet" ] && ! flock -n "$LOCKS/quiet" true 2>/dev/null; then echo "quiet hold"; return; fi
for f in "$LOCKS"/core-*; do [ -e "$f" ] && ! flock -n "$f" true 2>/dev/null && { echo "core lease"; return; }; done
slots=$(cat "$LOCKS/slots" 2>/dev/null || echo 1); [ "$slots" -ge 1 ] 2>/dev/null || slots=1
for k in $(seq 0 $((slots - 1))); do
[ -e "$LOCKS/build-$k" ] || continue

155
infra/build-server/lease.sh Executable file
View file

@ -0,0 +1,155 @@
#!/usr/bin/env bash
# Measurement leases on a build box (main, 7 October 2026, 15:07 UK: one global exclusive "measure" flock across unrelated
# measurements stalled build-1 at load 120 with free slots, an exclusive waiter queueing every new shared taker behind it, and
# a stopped probe held the box for five and a half hours). The measure file is retired. In its place:
#
# lease cores <set> --label "<text>" [--owner <agent>] [--nice N] -- <command...>
# A measurement that pins cores takes a lease on THOSE CORES ONLY (one flock per core, _locks/core-<n>, taken in ascending
# order, waited for up to 2 h with a wait-<pid> file carrying the label), runs the command under nice N (default 10) and
# taskset on the set, and releases. Builds and suites keep off leased cores (remote-run.sh reads the core files before it
# pins its own set). Nothing else is excluded: the box stays open.
# lease quiet --label "<text>" --owner <agent> [--cap-s N] -- <command...>
# A WHOLE-BOX quiet measurement: its own class, refused (exit 73) while any build slot or any core lease is held, capped at
# 20 minutes (timeout; --cap-s at most 1200), holder line with the owner in _locks/quiet. Unbounded builds take quiet
# shared and wait for it; bounded suites (nice 10, a 32-core band) never take it.
# lease status every lease, the quiet holder and every waiter with its label
# lease reap a holder (lease, quiet or build slot) whose process has been STOPPED (state T) for 5 minutes or more is
# killed and its file cleared, one line each in _log/reaped.log; remote-run.sh's keeper calls this every
# 20 s while any run is on the box (LEASE_REAP_S overrides the 300 s for the self-test)
# lease --self-test the known cases against a scratch lock directory (in the gate: tools/ci/pre-push.sh)
#
# Installed on every box at /srv/builds/_bin/lease by provision.sh (and copied by hand on 7 October 2026); the lock directory is
# IGNEUM_BUILD_SLOTS_DIR (the profile sets /srv/builds/_locks), the log directory IGNEUM_BUILD_LOG_DIR. Files: core-<n> (flock),
# lease-<pid> (holder line "pid N since HH:MM:SSZ cores <set>: <label>; owner=<agent>"), quiet (holder line), wait-<pid>.
set -uo pipefail
LOCKS="${IGNEUM_BUILD_SLOTS_DIR:-/srv/builds/_locks}"; LOGS="${IGNEUM_BUILD_LOG_DIR:-/srv/builds/_log}"
REAP_S="${LEASE_REAP_S:-300}"
now() { date -u +%H:%M:%SZ; }
say() { echo "lease: $*" >&2; }
expand_set() { # "6-11,54-59" -> "6 7 8 9 10 11 54 ..."
local out="" part lo hi; IFS=',' read -ra parts <<<"$1"
for part in "${parts[@]}"; do case "$part" in *-*) lo=${part%-*}; hi=${part#*-} ;; *) lo=$part; hi=$part ;; esac
[ "$lo" -le "$hi" ] 2>/dev/null || { say "bad core set '$1'"; return 2; }
for ((c = lo; c <= hi; c++)); do out="$out $c"; done; done
echo "${out# }"
}
held_cores() { # the cores whose lease file is flocked right now, as a space list
local f c out=""
for f in "$LOCKS"/core-*; do [ -e "$f" ] || continue; c=${f##*/core-}
exec 8>>"$f"; if ! flock -n 8; then out="$out $c"; fi; exec 8>&-; done
echo "${out# }"
}
any_slot_held() { local n f k; n=$(cat "$LOCKS/slots" 2>/dev/null || echo 1); for ((k = 0; k < n; k++)); do f="$LOCKS/build-$k"; [ -e "$f" ] || continue; exec 8>>"$f"; if ! flock -n 8; then exec 8>&-; return 0; fi; exec 8>&-; done; return 1; }
pid_of() { sed -n 's/^pid \([0-9]*\) .*/\1/p' "$1" 2>/dev/null | head -1; }
reap() { # stopped holders older than REAP_S are killed, their files cleared, one line each in reaped.log
local f p st age line n=0
for f in "$LOCKS"/lease-* "$LOCKS"/quiet "$LOCKS"/build-[0-9]*; do
[ -s "$f" ] || continue; p=$(pid_of "$f"); [ -n "$p" ] || continue
st=$(ps -o stat= -p "$p" 2>/dev/null | tr -d ' '); case "$st" in T*) ;; *) continue ;; esac
age=$(( $(date +%s) - $(stat -c %Y "$f") ))
# the file's mtime is refreshed by the holder's keeper every 20 s while it runs; a stopped holder's file goes stale
[ "$age" -ge "$REAP_S" ] || continue
line="$(now) reaped pid $p (stopped $age s, state $st) holding $(basename "$f"): $(head -c 160 "$f" | tr '\n' ' ')"
kill -KILL "$p" 2>/dev/null; pkill -KILL -P "$p" 2>/dev/null; : > "$f"; case "$f" in */lease-*) rm -f "$f" ;; esac
mkdir -p "$LOGS"; echo "$line" >> "$LOGS/reaped.log"; say "$line"; n=$((n + 1))
done
echo "$n"
}
status() {
echo "leases:"; for f in "$LOCKS"/lease-*; do [ -s "$f" ] && echo " $(cat "$f")"; done 2>/dev/null
echo "quiet: $(cat "$LOCKS/quiet" 2>/dev/null)"
echo "held cores: $(held_cores)"
echo "waiting:"; for f in "$LOCKS"/wait-*; do [ -s "$f" ] && echo " $(cat "$f")"; done 2>/dev/null
}
run_cores() {
local set="$1"; shift; local label="" owner="${IGNEUM_AGENT:-unknown}" nice=10
while [ $# -gt 0 ]; do case "$1" in --label) label="$2"; shift 2 ;; --owner) owner="$2"; shift 2 ;; --nice) nice="$2"; shift 2 ;; --) shift; break ;; *) say "unknown option $1"; exit 2 ;; esac; done
[ -n "$label" ] || { say "--label is required (the dashboard and the next lane read it)"; exit 2; }
[ $# -gt 0 ] || { say "no command"; exit 2; }
local cores c fd t0 waited=0 line waitfile="$LOCKS/wait-$$"
cores=$(expand_set "$set") || exit 2
mkdir -p "$LOCKS"; t0=$(date +%s)
line="pid $$ since $(now) waited 0 s: lease cores $set: $label; owner=$owner"; printf '%s\n' "$line" > "$waitfile"
trap 'rm -f "$waitfile" "$LOCKS/lease-$$"' EXIT
local fds=()
for c in $cores; do
exec {fd}>>"$LOCKS/core-$c"
if ! flock -n "$fd"; then say "core $c is leased ($(for f in "$LOCKS"/lease-*; do grep -l " core-$c\b\| cores [^:]*\b$c\b" "$f" 2>/dev/null; done | head -1 | xargs -r cat | cut -c1-120)); waiting"; flock -w 7200 "$fd" || { say "gave up waiting for core $c after 2 h"; exit 75; }; fi
fds+=("$fd")
done
waited=$(( $(date +%s) - t0 ))
line="pid $$ since $(now) waited $waited s: lease cores $set: $label; owner=$owner"; printf '%s\n' "$line" > "$LOCKS/lease-$$"; rm -f "$waitfile"
say "holding cores $set (waited $waited s): $label"
# the keeper closes the lock descriptors first: an inherited flock would outlive the release in its orphaned sleep
( for fd in "${fds[@]}"; do exec {fd}>&-; done; while kill -0 $$ 2>/dev/null; do touch "$LOCKS/lease-$$" 2>/dev/null; [ -s "$LOCKS/lease-$$" ] || printf '%s\n' "$line" > "$LOCKS/lease-$$"; sleep 20; done ) & local keeper=$!
nice -n "$nice" taskset -c "$set" "$@"; local rc=$?
pkill -P "$keeper" 2>/dev/null; kill "$keeper" 2>/dev/null; wait "$keeper" 2>/dev/null
rm -f "$LOCKS/lease-$$"; say "released cores $set after $(( $(date +%s) - t0 )) s, exit $rc"
exit $rc
}
run_quiet() {
local label="" owner="${IGNEUM_AGENT:-}" cap=1200
while [ $# -gt 0 ]; do case "$1" in --label) label="$2"; shift 2 ;; --owner) owner="$2"; shift 2 ;; --cap-s) cap="$2"; shift 2 ;; --) shift; break ;; *) say "unknown option $1"; exit 2 ;; esac; done
[ -n "$label" ] && [ -n "$owner" ] || { say "quiet needs --label and --owner (a named owner, the 5.5-hour rule)"; exit 2; }
[ "$cap" -le 1200 ] 2>/dev/null || cap=1200
[ $# -gt 0 ] || { say "no command"; exit 2; }
mkdir -p "$LOCKS"
if any_slot_held; then say "REFUSED: a build slot is held; a whole-box quiet measurement waits for an idle box (lease status)"; exit 73; fi
if [ -n "$(held_cores)" ]; then say "REFUSED: cores $(held_cores | tr ' ' ',') are leased; a whole-box quiet measurement waits for an idle box"; exit 73; fi
exec 7>>"$LOCKS/quiet"
if ! flock -n 7; then say "REFUSED: another quiet measurement holds the box: $(head -c 160 "$LOCKS/quiet")"; exit 73; fi
local t0 line; t0=$(date +%s)
line="pid $$ since $(now) waited 0 s: quiet (cap $cap s): $label; owner=$owner"; printf '%s\n' "$line" > "$LOCKS/quiet"
trap ': > "$LOCKS/quiet"' EXIT
( exec 7>&-; while kill -0 $$ 2>/dev/null; do touch "$LOCKS/quiet" 2>/dev/null; [ -s "$LOCKS/quiet" ] || printf '%s\n' "$line" > "$LOCKS/quiet"; sleep 20; done ) & local keeper=$!
say "holding the box quiet (cap $cap s): $label"
timeout --signal TERM --kill-after 30 "$cap" "$@"; local rc=$?
pkill -P "$keeper" 2>/dev/null; kill "$keeper" 2>/dev/null; wait "$keeper" 2>/dev/null
: > "$LOCKS/quiet"; [ "$rc" = 124 ] && say "the quiet measurement hit its ${cap}s cap and was ended"
say "released the box after $(( $(date +%s) - t0 )) s, exit $rc"; exit $rc
}
self_test() {
t=$(mktemp -d); trap 'rm -rf "$t"' EXIT
export IGNEUM_BUILD_SLOTS_DIR="$t/locks" IGNEUM_BUILD_LOG_DIR="$t/log" LEASE_REAP_S=2; mkdir -p "$t/locks"; echo 2 > "$t/locks/slots"
local me="$0" fail=0
f() { echo "lease self-test: FAIL: $*"; fail=1; }
# 1. two leases on disjoint cores run together; the same core waits
bash "$me" cores 0-1 --label A -- bash -c "date +%s.%N > '$t/a.start'; sleep 2; date +%s.%N > '$t/a.end'" 2>/dev/null &
sleep 0.5; bash "$me" cores 2-3 --label B -- bash -c "date +%s.%N > '$t/b.start'" 2>/dev/null; sleep 0.2
python3 -c "import sys; sys.exit(0 if float(open('$t/b.start').read()) < float(open('$t/a.start').read()) + 1.5 else 1)" || f "a lease on other cores waited for an unrelated lease"
bash "$me" cores 1-2 --label C -- bash -c "date +%s.%N > '$t/c.start'" 2>/dev/null
python3 -c "import sys; sys.exit(0 if float(open('$t/c.start').read()) >= float(open('$t/a.end').read()) else 1)" || f "a lease sharing a core started before the holder released it"
wait
[ -z "$(ls "$t/locks" | grep -E '^(lease|wait)-')" ] || f "lease or wait files left behind: $(ls "$t/locks")"
# 2. quiet is refused while a slot is held, and runs on an idle box with a holder line naming the owner
( exec 9>>"$t/locks/build-0"; flock 9; echo "pid $BASHPID since x waited 0 s: fake build" > "$t/locks/build-0"; sleep 2 ) &
sleep 0.3; bash "$me" quiet --label Q --owner tester -- true 2>/dev/null; [ $? = 73 ] || f "quiet was not refused while a slot was held"
wait; : > "$t/locks/build-0"
bash "$me" quiet --label Q --owner tester -- bash -c "grep -q 'owner=tester' '$t/locks/quiet'" || f "quiet did not run on an idle box with the owner in its holder line"
bash "$me" quiet --label Q --owner tester --cap-s 1 -- sleep 5; [ $? = 124 ] || f "the quiet cap did not end the command"
# 3. a quiet is refused while a core is leased
bash "$me" cores 0 --label L -- sleep 2 2>/dev/null & sleep 0.5
bash "$me" quiet --label Q --owner tester -- true 2>/dev/null; [ $? = 73 ] || f "quiet was not refused while a core was leased"
wait
# 4. reap: a stopped holder older than the window is killed and its file cleared, with a line
bash "$me" cores 5 --label S -- sleep 60 2>/dev/null & local lp=$!; sleep 0.6
local hp; hp=$(sed -n 's/^pid \([0-9]*\) .*/\1/p' "$t"/locks/lease-* | head -1); kill -STOP "$hp"; sleep 2.5
touch -d '-10 seconds' "$t"/locks/lease-* 2>/dev/null
[ "$(bash "$me" reap 2>/dev/null)" = 1 ] || f "the stopped lease holder was not reaped"
grep -q "reaped pid $hp" "$t/log/reaped.log" 2>/dev/null || f "no reaped line was written"
[ -z "$(ls "$t/locks" | grep '^lease-')" ] || f "the reaped lease file remains"
kill -KILL "$lp" 2>/dev/null; wait 2>/dev/null
# 5. a running (not stopped) holder is left alone
bash "$me" cores 6 --label R -- sleep 3 2>/dev/null & sleep 0.6; touch -d '-10 seconds' "$t"/locks/lease-*
[ "$(bash "$me" reap 2>/dev/null)" = 0 ] || f "a running holder was reaped"; wait
[ "$fail" = 0 ] && echo "lease self-test: disjoint leases run together, a shared core waits, quiet is refused beside a slot or a lease and capped, a stopped holder is reaped after the window, a running one is kept"
return $fail
}
case "${1:-}" in
cores) shift; run_cores "$@" ;;
quiet) shift; run_quiet "$@" ;;
status) status ;;
reap) reap ;;
--self-test) self_test ;;
*) sed -n '2,24p' "$0" | sed 's/^# \{0,1\}//'; exit 2 ;;
esac

View file

@ -45,7 +45,9 @@ bs_box_state() { # <box> -> "free=<n> slots=<n> load1=<x>" | "absent" | "down"
v=$(eval "printf '%s' \"\${BS_ROUTE_STATE_$b:-}\""); if [ -n "$v" ]; then printf '%s' "$v"; return 0; fi
f=$(bs_box_file "$b"); [ -s "$f" ] || { printf 'absent'; return 0; }
h="$(head -1 "$f" | tr -d '[:space:]')"
ssh "${BS_SSH_OPTS[@]}" -o ConnectTimeout=8 "$h" 'd=/srv/builds/_locks; n=$(cat $d/slots 2>/dev/null || echo 1); free=0; k=0; while [ $k -lt $n ]; do exec 9>>$d/build-$k; if flock -n 9; then free=$((free+1)); fi; exec 9>&-; k=$((k+1)); done; printf "free=%s slots=%s load1=%s" $free $n "$(cut -d" " -f1 /proc/loadavg)"' 2>/dev/null || printf 'down'
# the probe runs BEFORE bs_host builds BS_SSH_OPTS (the pool lane, 7 Oct 2026 15:1x UK: bash 3.2 under set -u refuses an
# unset array), so it carries its own options: the ops key, batch mode, a short connect timeout, no control socket
ssh -i "$BS_KEY" -o BatchMode=yes -o StrictHostKeyChecking=accept-new -o ConnectTimeout=8 "$h" 'd=/srv/builds/_locks; n=$(cat $d/slots 2>/dev/null || echo 1); free=0; k=0; while [ $k -lt $n ]; do exec 9>>$d/build-$k; if flock -n 9; then free=$((free+1)); fi; exec 9>&-; k=$((k+1)); done; printf "free=%s slots=%s load1=%s" $free $n "$(cut -d" " -f1 /proc/loadavg)"' 2>/dev/null || printf 'down'
}
bs_state_ok() { # <state> -> 0 when the box can take a job now (a free slot, load1 at or under the line)
local st="$1" free load

View file

@ -0,0 +1,57 @@
#!/usr/bin/env bash
# A headless Chromium for the text-overlap sweep (tools/ci/overlap-check.mjs) on a build box, without root. The build user
# has no sudo, so Playwright's `install --with-deps` cannot add the X and GTK libraries Chromium links; this script downloads
# the same Ubuntu packages with apt-get download (no root needed), unpacks them into a sysroot and points LD_LIBRARY_PATH
# at it, plus two font packages behind a private fontconfig file (the box ships no fonts at all: fc-list is empty, and
# text with no font has no width). Idempotent; re-run after a box re-provision.
#
# infra/build-server/overlap-browser.sh on the box (run-from-mac.sh copies and runs it; or ssh and run by hand)
# . /srv/builds/_bin/overlap/env.sh what a caller sources before `node tools/ci/overlap-check.mjs`
#
# Installed 7 October 2026 on igneum-build-2 (Ubuntu 24.04.5, node 22): playwright 1.56 with chromium 1194 (headless shell
# and full), 36 packages in the sysroot, Liberation and DejaVu fonts. A page with 20 px sans-serif text measured 23 px high
# through it, so glyphs have real metrics there.
set -euo pipefail
ROOT="${OVERLAP_ROOT:-/srv/builds/_bin/overlap}"
PW_VERSION="${PW_VERSION:-1.56}"
mkdir -p "$ROOT/debs" "$ROOT/sysroot" "$ROOT/fonts" "$ROOT/fc-cache"
cd "$ROOT"
[ -f package.json ] || npm init -y >/dev/null
if ! node -e "require('playwright')" 2>/dev/null; then npm i --no-audit --no-fund "playwright@$PW_VERSION" | tail -1; fi
# the browser download (Playwright keeps it under ~/.cache/ms-playwright; a second run finds it and does nothing)
PLAYWRIGHT_SKIP_VALIDATE_HOST_REQUIREMENTS=1 npx playwright install chromium 2>&1 | tail -1 || true
PKGS="libatk1.0-0t64 libatk-bridge2.0-0t64 libatspi2.0-0t64 libx11-6 libxcomposite1 libxdamage1 libxext6 libxfixes3 libxrandr2 libgbm1 libxcb1
libasound2t64 libxrender1 libxau6 libxdmcp6 libwayland-server0 libwayland-client0 libcups2t64 libcairo2 libpango-1.0-0 libpangocairo-1.0-0
libpangoft2-1.0-0 libharfbuzz0b libfribidi0 libthai0 libdatrie1 libpixman-1-0 libxcb-render0 libxcb-shm0 libavahi-client3 libavahi-common3
libxi6 libfontconfig1 libgraphite2-3 fonts-liberation fonts-dejavu-core"
cd "$ROOT/debs"
# shellcheck disable=SC2086
apt-get download $PKGS 2>&1 | grep -v '^Get:' | tail -1 || true
for d in ./*.deb; do dpkg -x "$d" "$ROOT/sysroot"; done
find "$ROOT/sysroot" -name '*.ttf' -exec cp -n {} "$ROOT/fonts/" \;
cat >"$ROOT/fonts.conf" <<EOF
<?xml version="1.0"?><!DOCTYPE fontconfig SYSTEM "fonts.dtd">
<fontconfig><dir>$ROOT/fonts</dir><cachedir>$ROOT/fc-cache</cachedir>
<alias><family>sans-serif</family><prefer><family>Liberation Sans</family></prefer></alias>
<alias><family>serif</family><prefer><family>Liberation Serif</family></prefer></alias>
<alias><family>monospace</family><prefer><family>Liberation Mono</family></prefer></alias></fontconfig>
EOF
cat >"$ROOT/env.sh" <<EOF
export LD_LIBRARY_PATH=$ROOT/sysroot/usr/lib/x86_64-linux-gnu\${LD_LIBRARY_PATH:+:\$LD_LIBRARY_PATH}
export FONTCONFIG_FILE=$ROOT/fonts.conf
export NODE_PATH=$ROOT/node_modules\${NODE_PATH:+:\$NODE_PATH}
export IGNEUM_PLAYWRIGHT_DIR=$ROOT
EOF
# shellcheck disable=SC1091
. "$ROOT/env.sh"
SHELL_BIN="$(ls -d "$HOME"/.cache/ms-playwright/chromium_headless_shell-*/chrome-linux/headless_shell | tail -1)"
if ldd "$SHELL_BIN" | grep -q 'not found'; then echo "overlap-browser: libraries still missing:"; ldd "$SHELL_BIN" | grep 'not found'; exit 1; fi
cat >"$ROOT/smoke.cjs" <<'EOF2'
const { chromium } = require("playwright");
(async () => { const b = await chromium.launch({ args: ["--no-sandbox"] }); const p = await b.newPage();
await p.setContent('<p id=a style="font:20px sans-serif">Hello overlap world</p>');
const h = await p.evaluate(() => document.getElementById("a").getBoundingClientRect().height); await b.close();
if (!(h > 15 && h < 40)) { console.error("overlap-browser: text has no metrics (height " + h + ")"); process.exit(1); }
console.log("overlap-browser: ready at " + __dirname + " (20 px text measures " + h + " px; source " + __dirname + "/env.sh)"); })();
EOF2
cd "$ROOT" && node smoke.cjs

View file

@ -69,7 +69,12 @@ RUNNER_VERSION="${RUNNER_VERSION:-2.338.0}" # github.com/actions
RUNNER_SHA256="${RUNNER_SHA256:-af4b794c1bc41d73d40535e3fe092a39f9679cd8d965954c2aca25a05ca41d32}" # the release note's linux-x64 line
RUNNER_REPO_URL="${RUNNER_REPO_URL:-https://github.com/igneum-network/igneum}"
RUNNER_NAME="${RUNNER_NAME:-$BOX_HOSTNAME}"
RUNNER_LABELS="${RUNNER_LABELS:-igneum-build-1}" # added to the defaults self-hosted, linux, x64
RUNNER_LABELS="${RUNNER_LABELS:-igneum-build-1,ci-red}" # added to the defaults self-hosted, linux, x64. igneum-build-1 is the POOL label
# (every box that takes pow and sims carries it); ci-red marks the one box that
# holds the red watcher's record file and poster. A second box: BOX_HOSTNAME=igneum-build-2
# RUNNER_LABELS=igneum-build-1,igneum-build-2 RUNNER_CPUS=0-31 RUNNER_JOBS=32 (register.sh --host)
RUNNER_CPUS="${RUNNER_CPUS:-}" # AllowedCPUs for the runner's service when set (a second box is bounded like a suite: 32 cores, nice 10)
case "$BOX_HOSTNAME" in *-2|*-3) [ -n "$RUNNER_CPUS" ] || RUNNER_CPUS="0-31" ;; esac # build-2 and build-3 join the pool (label igneum-build-1) bounded to 32 cores at Nice 10 (main, 7 Oct 2026)
RUNNER_JOBS="${RUNNER_JOBS:-48}" # cargo jobs for a CI job: half the box, the agents' builds keep the rest
RUNNER_TOKEN="${RUNNER_TOKEN:-}" # a registration token (1 h), from infra/build-server/runner/register.sh over stdin; never logged
RUNNER_SCCACHE_PORT="${RUNNER_SCCACHE_PORT:-4227}" # the runner's own sccache server; 4226 is the build user's
@ -146,6 +151,10 @@ APT_PACKAGES=(
libicu74 python3-numpy
# innoextract: tools/repro reads the shipped igneumd.exe and igneum-miner.exe out of the public Inno Setup installer
innoextract
# headless Chromium (the dashboard lane's site captures): the 16 system libraries it dlopens, found missing on both boxes
# on 7 October 2026 (installed by hand at 14:5x UK; step_headless_check launches the shell once and prints its version)
libatk1.0-0t64 libatk-bridge2.0-0t64 libatspi2.0-0t64 libcairo2 libcups2t64 libgbm1 libpango-1.0-0 libx11-6 libxcb1
libxcomposite1 libxdamage1 libxext6 libxfixes3 libxrandr2 libasound2t64 libxkbcommon0 fonts-liberation
)
step_apt() {
local need=() p
@ -239,6 +248,27 @@ step_dirs() {
done
[ "$any" = 1 ] && changed dirs "/srv/builds ($(worktree_count) worktree dirs), /srv/sccache, slots=$SLOTS" || ok dirs "$(worktree_count) worktree dirs, slots=$SLOTS"
}
# the lease tool (infra/build-server/lease.sh: per-core measurement leases, the quiet class, the reaper; 7 October 2026) from the
# mirror's master at /srv/builds/_bin/lease, which remote-run.sh's keeper calls; the mirror is pushed by step_mirrors' caller
step_lease_tool() {
local src="/srv/igneum.git" want have=""
install -d -m 755 -o "$BUILD_USER" -g "$BUILD_USER" /srv/builds/_bin
want=$(git -C "$src" show master:infra/build-server/lease.sh 2>/dev/null) || { ok lease-tool "mirror has no master yet; run-from-mac.sh installs it on the next provision"; return; }
[ -f /srv/builds/_bin/lease ] && have=$(cat /srv/builds/_bin/lease)
if [ "$want" = "$have" ]; then ok lease-tool "/srv/builds/_bin/lease is the mirror's master copy"; return; fi
printf '%s\n' "$want" > /srv/builds/_bin/lease.new; chmod 755 /srv/builds/_bin/lease.new; chown "$BUILD_USER:$BUILD_USER" /srv/builds/_bin/lease.new
mv /srv/builds/_bin/lease.new /srv/builds/_bin/lease; changed lease-tool "/srv/builds/_bin/lease installed from the mirror's master"
}
# headless Chromium self-test: the shell the dashboard lane's node tooling downloads (playwright or puppeteer cache of the build
# user) launches once with --headless and prints its version; when no shell is downloaded yet the libraries are checked by ldd of
# nothing, so the step only says so (the apt list above carries them)
step_headless_check() {
local shell="" v
shell=$(find "/home/$BUILD_USER/.cache/ms-playwright" "/home/$BUILD_USER/.cache/puppeteer" /srv -maxdepth 6 -type f \( -name headless_shell -o -name chrome-headless-shell -o -name chrome \) 2>/dev/null | head -1)
[ -n "$shell" ] || { ok headless "no headless shell downloaded yet (the 16 libraries are installed; the lane's first capture downloads it)"; return; }
if v=$(sudo -u "$BUILD_USER" timeout 60 "$shell" --headless --no-sandbox --disable-gpu --version 2>&1 | head -1) && [ -n "$v" ]; then ok headless "$shell: $v"
else die "headless shell $shell does not launch: $v"; fi
}
step_mirrors() {
local r any=0
@ -465,7 +495,7 @@ step_ufw() {
# documentation as remembered on 6 October 2026, the docs host answered 404 to the fetch that evening).
runner_env_file() {
cat <<EOF
# igneum-build-1 (infra/build-server/provision.sh step_runner): the environment every CI job on this runner starts with
# $BOX_HOSTNAME (infra/build-server/provision.sh step_runner): the environment every CI job on this runner starts with
RUSTC_WRAPPER=/usr/local/bin/sccache
SCCACHE_CONF=$RUNNER_HOME/.config/sccache/config
SCCACHE_SERVER_PORT=$RUNNER_SCCACHE_PORT
@ -536,7 +566,7 @@ step_runner() {
RUNNER_TOKEN="$RUNNER_TOKEN" runuser -u "$RUNNER_USER" -- bash -c "cd '$RUNNER_DIR' && ./config.sh --unattended --replace --url '$RUNNER_REPO_URL' --token \"\$RUNNER_TOKEN\" --name '$RUNNER_NAME' --labels '$RUNNER_LABELS' --work _work" >/dev/null \
|| die "runner: config.sh failed (an expired token? register.sh fetches a fresh one)"
any=1
log "runner: registered as $RUNNER_NAME with labels self-hosted, linux, x64, $RUNNER_LABELS"
log "runner: registered as $RUNNER_NAME with labels self-hosted, linux, x64, $RUNNER_LABELS${RUNNER_CPUS:+, AllowedCPUs $RUNNER_CPUS}"
fi
# 7. the service: GitHub's unit (User=runner, KillMode=process) plus Nice and a restart on failure
svc="actions.runner.$(sed -n 's/.*"gitHubUrl": *"https:\/\/github.com\/\([^"]*\)".*/\1/p' "$RUNNER_DIR/.runner" | tr '/' '-').$RUNNER_NAME.service"
@ -552,7 +582,8 @@ step_runner() {
rm -f "$tmp"
dropin="/etc/systemd/system/$svc.d/igneum.conf"
tmp=$(mktemp)
printf '# igneum-build-1 (infra/build-server/provision.sh step_runner)\n[Service]\nNice=10\nIOSchedulingClass=best-effort\nIOSchedulingPriority=7\nRestart=on-failure\nRestartSec=30\n' > "$tmp"
printf '# %s (infra/build-server/provision.sh step_runner)\n[Service]\nNice=10\nIOSchedulingClass=best-effort\nIOSchedulingPriority=7\nRestart=on-failure\nRestartSec=30\n' "$BOX_HOSTNAME" > "$tmp"
[ -z "$RUNNER_CPUS" ] || printf 'AllowedCPUs=%s\n' "$RUNNER_CPUS" >> "$tmp"
install -d -m 755 "$(dirname "$dropin")"
if ! cmp -s "$tmp" "$dropin"; then install -m 644 "$tmp" "$dropin"; systemctl daemon-reload; any=1; fi
rm -f "$tmp"
@ -634,6 +665,7 @@ do_provision() {
step_docker
step_dirs
step_mirrors
step_lease_tool
step_rustup
step_sccache
step_cargo_config
@ -647,6 +679,7 @@ do_provision() {
step_cargo_tools
step_zig
step_night
step_headless_check
step_summary
log "done"
}

View file

@ -181,8 +181,8 @@ if [ "${1:-}" = --self-test-slots ]; then
mkdir -p "$t/locks" "$t/log" "$t/dir"; echo 2 > "$t/locks/slots"
fake() { # <name> <seconds> [BR_MEASURE=1]: a fake run that records its start and end epoch and the job count it was given
local name="$1" secs="$2" measure="${3:-0}"
IGNEUM_BUILD_SLOTS_DIR="$t/locks" IGNEUM_BUILD_LOG_DIR="$t/log" BR_MEASURE="$measure" BR_DIR="$t/dir" BR_CMD="date +%s.%N > '$t/$name.start'; echo JOBS=\${CARGO_BUILD_JOBS:-none} > '$t/$name.jobs'; sleep $secs; date +%s.%N > '$t/$name.end'" \
BR_LABEL="self-test $name" BR_TOOL=self-test BR_KIND=other BR_WT=t BR_CRATE=t BR_BRANCH=t BR_SHA=0 BR_AGENT=self-test BR_COMMAND="fake $name" \
IGNEUM_BUILD_SLOTS_DIR="$t/locks" IGNEUM_BUILD_LOG_DIR="$t/log" BR_MEASURE="$measure" BR_CORES="${BR_CORES:-0}" BR_NICE="${BR_NICE:-0}" BR_DIR="$t/dir" BR_CMD="date +%s.%N > '$t/$name.start'; echo JOBS=\${CARGO_BUILD_JOBS:-none} > '$t/$name.jobs'; sleep $secs; date +%s.%N > '$t/$name.end'" \
BR_LABEL="self-test $name" BR_TOOL=self-test BR_KIND=other BR_WT="wt-$name" BR_CRATE=t BR_BRANCH=t BR_SHA=0 BR_AGENT=self-test BR_COMMAND="fake $name" \
bash "$me" >"$t/$name.out" 2>&1
}
fail() { echo "self-test-slots: FAIL: $*"; exit 1; }
@ -193,33 +193,41 @@ if [ "${1:-}" = --self-test-slots ]; then
# 2. one build alone: 90
fake c 1
[ "$(cat "$t/c.jobs")" = JOBS=90 ] || fail "a lone build got $(cat "$t/c.jobs") (want JOBS=90)"
# 3. a measure blocks a build: the build starts only after the measure ended
fake m 3 1 & sleep 0.5; fake d 1 & wait
after "$t/d.start" "$t/m.end" || fail "a build started while a measurement held the box (build start $(cat "$t/d.start"), measure end $(cat "$t/m.end"))"
# 3. a quiet measurement blocks an unbounded build (it starts only after the quiet ended) and lets a bounded suite run beside it
fake m 3 1 & sleep 0.5; fake d 1 & BR_CORES=1 BR_NICE=10 fake s 1 & wait
after "$t/d.start" "$t/m.end" || fail "an unbounded build started while a quiet measurement held the box (build start $(cat "$t/d.start"), quiet end $(cat "$t/m.end"))"
python3 -c "import sys; sys.exit(0 if float(open('$t/s.start').read()) < float(open('$t/m.end').read()) else 1)" || fail "a bounded suite waited for the quiet measurement"
[ "$(cat "$t/m.jobs")" = JOBS=none ] || fail "a measurement was given a job count"
# 4. a build blocks a measure: the measure starts only after the build ended
fake e 3 & sleep 0.5; fake n 1 1 & wait
after "$t/n.start" "$t/e.end" || fail "a measurement started while a build ran (measure start $(cat "$t/n.start"), build end $(cat "$t/e.end"))"
grep -q 'owner=self-test' "$t/m.out" "$t/log/builds.jsonl" 2>/dev/null || fail "the quiet holder line lacks its owner"
# 4. a quiet measurement is REFUSED (exit 73) while a build holds a slot or a core is leased
fake e 3 & sleep 0.5; fake n 1 1; rc=$?; wait
[ "$rc" = 73 ] && [ ! -f "$t/n.start" ] || fail "a quiet measurement was not refused while a build ran (rc $rc)"
( exec 9>>"$t/locks/core-7"; flock 9; sleep 6 ) & sleep 0.5; fake o 1 1; rc=$?; wait
[ "$rc" = 73 ] || fail "a quiet measurement was not refused while a core was leased (rc $rc)"
# 4b. a run keeps off leased cores: with core 1 leased, a 2-core bounded run on a 2-core box says so (the exclusion line)
# the lease outlives the fake's slot take and settle (a loaded box took them past 2 s and the first version read no lease)
( exec 9>>"$t/locks/core-$(( $(nproc) - 1 ))"; flock 9; sleep 12 ) & sleep 0.5; BR_CORES=2 BR_NICE=10 fake p 1; wait
grep -q 'are leased to a measurement; this run keeps to' "$t/p.out" || fail "a run beside a leased core did not exclude it: $(cat "$t/p.out" | tail -3)"
# 5. a probing build leaves a busy slot's holder line intact
fake f 3 & sleep 1.2; fake g 1 & sleep 0.3
grep -q 'self-test f' "$t/locks/build-0" || fail "the holder line of the busy slot build-0 was lost when another build probed it: '$(cat "$t/locks/build-0")'"
wait
# 6. the log carries the job count and the measure flag
grep -q '"jobs":45' "$t/log/builds.jsonl" && grep -q '"measure":true' "$t/log/builds.jsonl" || fail "builds.jsonl lacks jobs or measure fields"
echo "self-test-slots: two concurrent builds 45 each, a lone build 90, a measure blocks a build, a build blocks a measure, a probe keeps the holder line, the log carries jobs and measure"; exit 0
echo "self-test-slots: two concurrent builds 45 each, a lone build 90, a quiet blocks an unbounded build and not a bounded suite, a quiet is refused beside a slot or a lease, a run keeps off leased cores, a probe keeps the holder line, the log carries jobs and measure"; exit 0
fi
# One run per worktree directory at a time (6 October 2026, 19:51:09 UK: two runs of one worktree started in the same second;
# one found no crate directory while the other's checkout was replacing the tree, exit 2). The checkout and the run each take
# the worktree's lock (append mode, held to exit) and wait up to 2 h for it instead of dying; the wait is said on stderr.
wt_lock() { # <worktree name>
local name="${1//\//_}" fd
local name="${1//\//_}"
[ -n "$name" ] || return 0
mkdir -p "$IGNEUM_BUILD_SLOTS_DIR" 2>/dev/null || return 0
exec {fd}>>"$IGNEUM_BUILD_SLOTS_DIR/wt-$name.lock" || return 0
if ! flock -n "$fd"; then
exec {WT_FD}>>"$IGNEUM_BUILD_SLOTS_DIR/wt-$name.lock" || return 0 # WT_FD is global: the keeper closes it (an orphaned sleep held it 20 s)
if ! flock -n "$WT_FD"; then
echo "build-remote: another run holds worktree $1 on this box, waiting for it (up to 2 h)" >&2
flock -w 7200 "$fd" || { echo "build-remote: gave up waiting for worktree $1 after 2 h" >&2; exit 75; }
flock -w 7200 "$WT_FD" || { echo "build-remote: gave up waiting for worktree $1 after 2 h" >&2; exit 75; }
fi
}
if [ "${BR_MODE:-run}" = checkout ]; then
@ -318,25 +326,30 @@ give_up() { # <what>
rm -f "$waitfile"; exit 75
}
waitfile="$SLOTS_DIR/wait-$BR_PID"
# The measure file is RETIRED (main, 7 October 2026, 15:07 UK: one global exclusive flock across unrelated measurements stalled
# build-1 at load 120 with free slots). A measurement that pins cores takes a lease on those cores only (lease.sh, installed at
# /srv/builds/_bin/lease), and this runner keeps its command off leased cores (see cores_str below). A WHOLE-BOX quiet measurement
# (BR_MEASURE=1) is its own class: refused with exit 73 while any slot or core lease is held, capped at 20 minutes, holder line with
# the owner in $SLOTS_DIR/quiet. An unbounded run (nice 0, the full core set) takes quiet shared and waits for it; a bounded run
# (BR_CORES > 0: suites, benches, everything on box 2) never takes it.
held_cores() { local f c out=""; for f in "$SLOTS_DIR"/core-*; do [ -e "$f" ] || continue; c=${f##*/core-}; exec {cfd}>>"$f"; if ! flock -n "$cfd"; then out="$out $c"; fi; exec {cfd}>&-; done; echo "${out# }"; }
any_slot_held() { local k f; for ((k = 0; k < slots; k++)); do f="$SLOTS_DIR/build-$k"; [ -e "$f" ] || continue; exec {sfd}>>"$f"; if ! flock -n "$sfd"; then exec {sfd}>&-; return 0; fi; exec {sfd}>&-; done; return 1; }
# append mode: opening a lock file must never truncate the holder line another run wrote into it
exec {mfd}>>"$SLOTS_DIR/measure"
exec {mfd}>>"$SLOTS_DIR/quiet"
if [ "${BR_MEASURE:-0}" = 1 ]; then
# a measurement: the measure file exclusively; every running build holds it shared, so this waits for them and blocks new ones
if ! flock -n "$mfd"; then
echo "build-remote: measure waits for the running build(s) (up to 2 h): $(for f in "$SLOTS_DIR"/build-*; do head -c 120 "$f" 2>/dev/null; done | tr '\n' ' ')" >&2
holder_line 0 > "$waitfile"; trap 'rm -f "$waitfile"' EXIT
flock -w 7200 "$mfd" || give_up "the measure hold"
rm -f "$waitfile"; trap - EXIT
fi
if any_slot_held; then echo "build-remote: QUIET REFUSED: a build slot is held ($(for f in "$SLOTS_DIR"/build-*; do head -c 100 "$f" 2>/dev/null; done | tr '\n' ' ')); a whole-box measurement needs an idle box" >&2; BR_CLASS=quiet-refused jsonlog 73 0 0 "" "$(date +%s)" "" "" "" "" "" ""; exit 73; fi
lc=$(held_cores); if [ -n "$lc" ]; then echo "build-remote: QUIET REFUSED: cores $(echo $lc | tr ' ' ',') are leased; a whole-box measurement needs an idle box" >&2; BR_CLASS=quiet-refused jsonlog 73 0 0 "" "$(date +%s)" "" "" "" "" "" ""; exit 73; fi
if ! flock -n "$mfd"; then echo "build-remote: QUIET REFUSED: another quiet measurement holds the box: $(head -c 160 "$SLOTS_DIR/quiet")" >&2; BR_CLASS=quiet-refused jsonlog 73 0 0 "" "$(date +%s)" "" "" "" "" "" ""; exit 73; fi
[ -n "${BR_AGENT:-}" ] && [ "$BR_AGENT" != unknown ] || { echo "build-remote: QUIET REFUSED: a whole-box measurement needs a named owner (IGNEUM_AGENT)" >&2; exit 73; }
waited=$(( $(date +%s) - BR_T0 )); got=measure
holder_line "$waited" > "$SLOTS_DIR/measure"
echo "build-remote: holding measure on $BR_HOST (waited $waited s; builds are excluded until this run ends)" >&2
BR_LABEL="quiet (cap 1200 s): $BR_LABEL; owner=$BR_AGENT"; holder_line "$waited" > "$SLOTS_DIR/quiet"
echo "build-remote: holding the box QUIET on $BR_HOST (waited $waited s; owner $BR_AGENT; cap 20 min; unbounded builds wait, bounded suites run beside it on their bands)" >&2
BR_CMD="timeout --signal TERM --kill-after 30 1200 bash -c $(printf '%q' "$BR_CMD")"
else
# a build: the measure file shared (a running measurement blocks us), then one exclusive slot
if ! flock -s -n "$mfd"; then
echo "build-remote: a measurement holds the box, waiting (up to 2 h): $(head -c 160 "$SLOTS_DIR/measure" 2>/dev/null)" >&2
if [ "${BR_CORES:-0}" = 0 ] && ! flock -s -n "$mfd"; then
echo "build-remote: a quiet measurement holds the box, waiting (up to 20 min): $(head -c 160 "$SLOTS_DIR/quiet" 2>/dev/null)" >&2
holder_line 0 > "$waitfile"; trap 'rm -f "$waitfile"' EXIT
flock -s -w 7200 "$mfd" || give_up "the measurement to end"
flock -s -w 1500 "$mfd" || give_up "the quiet measurement to end"
rm -f "$waitfile"; trap - EXIT
fi
# scheduling (main, 7 Oct 2026): a gate announces itself (gate-pending-<pid>) and takes the next free slot; a suite, bench or
@ -390,13 +403,16 @@ fi
# because a run from a worktree without last night's append-mode fix still opens a busy sibling's slot file with `>` on every
# probe). A keeper re-writes the holder line whenever it finds the file empty, every BR_KEEP_S seconds, until release_slot.
keeper_pid=""
keep_line() { # <file> <waited>
keep_line() { # <file> <waited>; the keeper also refreshes the file's mtime (a stopped holder's file goes stale) and reaps stopped
# holders of any lease, quiet or slot after 5 minutes through lease.sh (/srv/builds/_bin/lease reap, one line each in reaped.log)
local f="$1" w="$2"
( while kill -0 "$BR_PID" 2>/dev/null; do [ -s "$f" ] || holder_line "$w" > "$f" 2>/dev/null; sleep "${BR_KEEP_S:-20}"; done ) &
( exec {mfd}>&- {fd}>&- 2>/dev/null; [ -n "${WT_FD:-}" ] && exec {WT_FD}>&-; while kill -0 "$BR_PID" 2>/dev/null; do [ -s "$f" ] || holder_line "$w" > "$f" 2>/dev/null; touch "$f" 2>/dev/null
[ -x /srv/builds/_bin/lease ] && IGNEUM_BUILD_SLOTS_DIR="$SLOTS_DIR" IGNEUM_BUILD_LOG_DIR="$LOG_DIR" /srv/builds/_bin/lease reap >/dev/null 2>&1
sleep "${BR_KEEP_S:-20}"; done ) &
keeper_pid=$!
}
if [ "$got" = measure ]; then keep_line "$SLOTS_DIR/measure" "$waited"; else keep_line "$SLOTS_DIR/build-$got" "$waited"; fi
release_slot() { [ -n "$keeper_pid" ] && { kill "$keeper_pid" 2>/dev/null; wait "$keeper_pid" 2>/dev/null; keeper_pid=""; }; if [ "$got" = measure ]; then : > "$SLOTS_DIR/measure"; else : > "$SLOTS_DIR/build-$got"; fi; }
if [ "$got" = measure ]; then keep_line "$SLOTS_DIR/quiet" "$waited"; else keep_line "$SLOTS_DIR/build-$got" "$waited"; fi
release_slot() { [ -n "$keeper_pid" ] && { pkill -P "$keeper_pid" 2>/dev/null; kill "$keeper_pid" 2>/dev/null; wait "$keeper_pid" 2>/dev/null; keeper_pid=""; }; if [ "$got" = measure ]; then : > "$SLOTS_DIR/quiet"; else : > "$SLOTS_DIR/build-$got"; fi; }
cd "$BR_DIR" || { BR_CLASS=no-dir jsonlog 2 "$got" "$waited" "" "$(date +%s)" "" "" "" "" "" ""; redlog 2 0 no-dir; release_slot; exit 2; }
# PRE-FLIGHT for a cargo command (the instant-death class, 6 October 2026): the manifest parses and every -p package exists,
@ -457,6 +473,29 @@ if [ "${BR_CORES:-0}" -gt 0 ] && [ "${BR_CORES}" -lt "$ncpu" ]; then
lo=$((ncpu - BR_CORES * (band + 1))); [ "$lo" -ge 0 ] || lo=$((ncpu - BR_CORES))
cores_str="$lo-$((lo + BR_CORES - 1))"
fi
# leased cores (lease.sh: a pinned measurement's core-<n> flocks) are taken out of this run's set; a set that would be empty keeps
# its cores (the measurement is told by its own lease line); the exclusion is said once
if [ "$got" != measure ]; then
leased=$(held_cores)
if [ -n "$leased" ]; then
kept=$(python3 -c '
import sys
def expand(s):
out=set()
for part in s.split(","):
a,_,b=part.partition("-"); a=int(a); b=int(b) if b else a; out.update(range(a,b+1))
return out
mine=expand(sys.argv[1]); leased=set(int(x) for x in sys.argv[2].split()); keep=sorted(mine-leased)
if not keep: print(sys.argv[1]); sys.exit()
runs=[];
for c in keep:
if runs and runs[-1][1]==c-1: runs[-1][1]=c
else: runs.append([c,c])
print(",".join(f"{a}-{b}" if a!=b else f"{a}" for a,b in runs))' "$cores_str" "$leased")
[ "$kept" != "$cores_str" ] && echo "build-remote: cores $(echo $leased | tr ' ' ',') are leased to a measurement; this run keeps to $kept" >&2
cores_str="$kept"
fi
fi
( [ "${BR_NICE:-0}" -gt 0 ] && renice -n "$BR_NICE" -p $BASHPID >/dev/null 2>&1; [ "$cores_str" != "0-$((ncpu - 1))" ] && taskset -cp "$cores_str" $BASHPID >/dev/null 2>&1; eval "$BR_CMD" ) > >(tee -a "$BR_RUN_LOG") 2> >(tee -a "$BR_RUN_LOG" >&2)
rc=$?
# the keeper stops BEFORE the bare `wait` (which flushes the two tees): a bare wait also waits for the keeper, and the keeper waits

View file

@ -2,6 +2,12 @@
# Register (or re-register) the GitHub Actions self-hosted runner on igneum-build-1 from this Mac.
# infra/build-server/runner/register.sh fetch a registration token with gh, run provision.sh on the box with it
# infra/build-server/runner/register.sh --status list the repository's runners (name, status, labels) and the box's unit
# infra/build-server/runner/register.sh --host <ip> another box (7 October 2026, igneum-build-2): the same, against that ip;
# BOX_HOSTNAME, RUNNER_NAME, RUNNER_LABELS, RUNNER_CPUS and RUNNER_JOBS
# from this shell's environment travel with it (provision.sh would
# otherwise rename the box igneum-build-1), e.g.
# BOX_HOSTNAME=igneum-build-2 RUNNER_LABELS=igneum-build-1,igneum-build-2 RUNNER_CPUS=0-31 RUNNER_JOBS=32 \
# infra/build-server/runner/register.sh --host 142.132.249.238
#
# The token: `gh api -X POST repos/igneum-network/igneum/actions/runners/registration-token` as igneum-labs (the CLAUDE.md gh
# rule: that account must be ACTIVE; any other active account fails here before anything is fetched). It is a one-hour
@ -14,7 +20,20 @@ HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_SLUG="${IGNEUM_GH_REPO:-igneum-network/igneum}"
KEY="${IGNEUM_BUILD_KEY:-$HOME/.ssh/igneum_ed25519}"
HOST_LINE="$(head -1 "${IGNEUM_BUILD_HOST_FILE:-$HOME/.config/igneum/build-server}" | tr -d '[:space:]')"
IP="${HOST_LINE#*@}"; [ -n "$IP" ] || { echo "no build server in ~/.config/igneum/build-server (infra/build-server/run-from-mac.sh writes it)" >&2; exit 1; }
IP="${HOST_LINE#*@}"
if [ "${1:-}" = --host ]; then
IP="${2:-}"; [ -n "$IP" ] || { echo "--host needs an ip" >&2; exit 1; }; shift 2
[ -n "${BOX_HOSTNAME:-}" ] || { echo "--host: set BOX_HOSTNAME (provision.sh would otherwise rename the box igneum-build-1)" >&2; exit 1; }
fi
[ -n "$IP" ] || { echo "no build server in ~/.config/igneum/build-server (infra/build-server/run-from-mac.sh writes it)" >&2; exit 1; }
# the provisioning variables a second box needs, forwarded as assignments in front of the remote shell (values are plain
# words: a hostname, a label list, a cpu range, a number; anything else is refused)
FWD=""
for v in BOX_HOSTNAME RUNNER_NAME RUNNER_LABELS RUNNER_CPUS RUNNER_JOBS; do
val="${!v:-}"; [ -n "$val" ] || continue
printf '%s' "$val" | grep -qE '^[A-Za-z0-9,._-]+$' || { echo "$v='$val' is not a plain word" >&2; exit 1; }
FWD="$FWD $v=$val"
done
SSH=(ssh -i "$KEY" -o BatchMode=yes -o StrictHostKeyChecking=accept-new -o ConnectTimeout=15 "root@$IP")
gh_josh() {
@ -43,7 +62,7 @@ echo "token received (not shown); running provision.sh on root@$IP with it (firs
# dropped before it is shown (provision.sh never prints it; this is the belt)
OUT=$(mktemp); trap 'rm -f "$OUT"' EXIT
set +e
{ printf '%s\n' "$TOKEN"; cat "$HERE/../provision.sh"; } | "${SSH[@]}" 'IFS= read -r RUNNER_TOKEN; export RUNNER_TOKEN; MODE=provision bash -s' > "$OUT" 2>&1
{ printf '%s\n' "$TOKEN"; cat "$HERE/../provision.sh"; } | "${SSH[@]}" "IFS= read -r RUNNER_TOKEN; export RUNNER_TOKEN; MODE=provision$FWD bash -s" > "$OUT" 2>&1
RC=$?
set -e
grep -v -F "$TOKEN" "$OUT" || true

128
infra/build-server/wave1-0320.sh Executable file
View file

@ -0,0 +1,128 @@
#!/usr/bin/env bash
# The 0.3.20 FIRST wave (the shipper, 7 October 2026): the publish carries a floor-moved override file whose digest differs from
# the live one, and a node on the old file refuses a node on the new one as a peer, so the three testnet seeds and the hands move
# AT the publish minute, in parallel with the fleet's wave. Dry run by default (every line printed, no box touched); --go executes.
#
# infra/build-server/wave1-0320.sh seeds [--override <file>] --igneumd <seed-class igneumd> --sha256 <hex> [--miner <seed-class igneum-miner>] \
# --digest <hex> [--commit <sha8>] [--wipe-genesis --genesis <hex>] [--go]
# the three seeds in PARALLEL (seed1/2/3.testnet, root over the ops key, infra/seed-nodes/seeds-testnet.tsv): the binary put
# as /opt/igneum/bin/igneumd.new with its sha256 asserted on the box, the override written to /etc/igneum/override-params.json,
# EXTRA_ARGS in /etc/igneum/seed.env set to --override-params-file=<that>, the unit igneumd stopped, the binary swapped (the old
# one kept as igneumd.prev), the unit started on its kept datadir, then the read-back: commit string in the installed binary,
# the "Consensus params digest" line of the new run against --digest, eth_syncing over the loopback EVM RPC, peers; one line
# per seed, logs in the scratch directory. Without --override the seeds' env is left as it is (no override today).
# --wipe-genesis (the testnet go, docs/plans/testnet-go.md runbook step 2; the shipper, 7 Oct 2026): the genesis re-cut changes
# the genesis itself, so a binary put onto the kept datadir would refuse its own genesis; with the unit stopped the datadir
# /var/lib/igneum/igneum-testnet-1 (height 0, nothing but genesis) is moved aside as igneum-testnet-1.prev-<stamp> and the node
# starts fresh; the read-back adds the "[igneum-exec] genesis <hash> executed" line against --genesis (MATCH/MISMATCH)
# infra/build-server/wave1-0320.sh hands --node <fork worktree on the Mac> --override-json '<object>' --digest <hex> [--go]
# the hands through infra/build-server/hands/move-hand.sh: binary (build + install + override), restart observer-node, restart
# node1, each read back by that tool (first executing line, commit string, digest, powEngine)
# infra/build-server/wave1-0320.sh lift-rpc-filter [--go]
# on seed1 only, AFTER the three read-backs: BLOCKED_UNTIL_FIXED_NODE in /opt/igneum/bin/rpc-filter.py becomes the empty set
# (the N7 class is fixed from 8097d600; backup kept as rpc-filter.py.bak-<stamp>), python3 -m py_compile, the unit
# igneum-rpc-filter restarted, read back: unit active and eth_blockNumber answered through https://rpc.testnet.igneum.network
set -euo pipefail
HERE=$(cd "$(dirname "$0")" && pwd); ROOT=$(cd "$HERE/../.." && pwd)
SEEDS_TSV="$ROOT/infra/seed-nodes/seeds-testnet.tsv"; KEY="${IGNEUM_BUILD_KEY:-$HOME/.ssh/igneum_ed25519}"
S="${WAVE_LOG_DIR:-/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/wave1-0320}"; mkdir -p "$S"
SSH=(ssh -i "$KEY" -o BatchMode=yes -o StrictHostKeyChecking=accept-new -o ConnectTimeout=15)
now() { TZ=Europe/London date '+%H:%M:%S %Z'; }
say() { echo "$(now) wave1: $*" >&2; }
die() { say "ERROR: $*"; exit 1; }
seed_ip() { awk -F'\t' -v n="$1" '$1==n {print $4}' "$SEEDS_TSV"; }
mode="${1:-}"; [ -n "$mode" ] || { sed -n '2,24p' "$0" | sed 's/^# \{0,1\}//'; exit 2; }; shift
GO=0; OVERRIDE=""; BIN=""; SHA=""; MINER=""; DIGEST=""; COMMIT="c4459193"; NODE=""; OVJSON=""; WIPE=0; GENESIS=""
while [ $# -gt 0 ]; do case "$1" in
--go) GO=1; shift ;; --override) OVERRIDE="$2"; shift 2 ;; --igneumd) BIN="$2"; shift 2 ;; --sha256) SHA="$2"; shift 2 ;;
--miner) MINER="$2"; shift 2 ;; --digest) DIGEST="$2"; shift 2 ;; --commit) COMMIT="$2"; shift 2 ;; --node) NODE="$2"; shift 2 ;;
--override-json) OVJSON="$2"; shift 2 ;; --wipe-genesis) WIPE=1; shift ;; --genesis) GENESIS="$2"; shift 2 ;; *) die "unknown option $1" ;; esac; done
one_seed() { # <name>: runs in the background; prints one RESULT line at the end
local name="$1" ip log t0 line
ip=$(seed_ip "$name"); [ -n "$ip" ] || { echo "RESULT $name: no ip in $SEEDS_TSV"; return 1; }
log="$S/$name.log"; : > "$log"; t0=$(date +%s)
{
if [ "$GO" = 0 ]; then
echo "DRY: scp $BIN root@$ip:/opt/igneum/bin/igneumd.new; sha256 asserted = $SHA"
[ -n "$MINER" ] && echo "DRY: scp $MINER root@$ip:/opt/igneum/bin/igneum-miner.new (swapped with the node)"
[ -n "$OVERRIDE" ] && echo "DRY: scp $OVERRIDE root@$ip:/etc/igneum/override-params.json; seed.env EXTRA_ARGS=--override-params-file=/etc/igneum/override-params.json" || echo "DRY: no override: seed.env untouched"
[ "$WIPE" = 1 ] && echo "DRY: with the unit stopped: mv /var/lib/igneum/igneum-testnet-1 /var/lib/igneum/igneum-testnet-1.prev-<stamp> (kept); genesis read back against $GENESIS"
echo "DRY: systemctl stop igneumd; mv igneumd igneumd.prev; mv igneumd.new igneumd; systemctl start igneumd"
echo "DRY: read back: grep -c $COMMIT in the binary; journal 'Consensus params digest' == $DIGEST; eth_syncing on 127.0.0.1:26890; peers"
echo "DRY RUN, nothing touched; box $ip answers as $("${SSH[@]}" "root@$ip" 'hostname; systemctl is-active igneumd; sha256sum /opt/igneum/bin/igneumd | cut -c1-12' 2>/dev/null | tr '\n' ' ')"
else
scp -q -i "$KEY" -o BatchMode=yes "$BIN" "root@$ip:/opt/igneum/bin/igneumd.new"
[ -n "$MINER" ] && scp -q -i "$KEY" -o BatchMode=yes "$MINER" "root@$ip:/opt/igneum/bin/igneum-miner.new"
[ -n "$OVERRIDE" ] && scp -q -i "$KEY" -o BatchMode=yes "$OVERRIDE" "root@$ip:/etc/igneum/override-params.json"
"${SSH[@]}" "root@$ip" bash -s -- "$SHA" "$COMMIT" "$DIGEST" "$([ -n "$MINER" ] && echo 1 || echo 0)" "$([ -n "$OVERRIDE" ] && echo 1 || echo 0)" "$WIPE" "$GENESIS" <<'REMOTE'
set -euo pipefail
SHA="$1"; COMMIT="$2"; DIGEST="$3"; MINER="$4"; OV="$5"; WIPE="$6"; GENESIS="$7"; cd /opt/igneum/bin
got=$(sha256sum igneumd.new | cut -d' ' -f1); [ "$got" = "$SHA" ] || { echo "sha256 MISMATCH on the box: $got"; exit 1; }
if [ "$OV" = 1 ]; then
python3 -c 'import json,sys; json.load(open("/etc/igneum/override-params.json"))' || { echo "override file does not parse"; exit 1; }
grep -qE '^EXTRA_ARGS=' /etc/igneum/seed.env && sed -i -E 's#^EXTRA_ARGS=.*#EXTRA_ARGS="--override-params-file=/etc/igneum/override-params.json"#' /etc/igneum/seed.env || echo 'EXTRA_ARGS="--override-params-file=/etc/igneum/override-params.json"' >> /etc/igneum/seed.env
fi
chmod 755 igneumd.new; [ "$MINER" = 1 ] && chmod 755 igneum-miner.new
t0=$(date +%s); systemctl stop igneumd
wiped=""
if [ "$WIPE" = 1 ] && [ -d /var/lib/igneum/igneum-testnet-1 ]; then stamp=$(date -u +%Y%m%dT%H%M%SZ); mv /var/lib/igneum/igneum-testnet-1 "/var/lib/igneum/igneum-testnet-1.prev-$stamp"; wiped="datadir moved aside as igneum-testnet-1.prev-$stamp; "; fi
mv -f igneumd igneumd.prev; mv -f igneumd.new igneumd; [ "$MINER" = 1 ] && { mv -f igneum-miner igneum-miner.prev 2>/dev/null || true; mv -f igneum-miner.new igneum-miner; }
systemctl start igneumd; sleep 6
gline=""; if [ "$WIPE" = 1 ]; then for i in $(seq 1 30); do gline=$(journalctl -u igneumd --since "-90 s" --no-pager 2>/dev/null | grep -oE "genesis [0-9a-f]{8,} executed" | tail -1); [ -n "$gline" ] && break; sleep 2; done; fi
gmatch=""; if [ "$WIPE" = 1 ]; then gh=$(echo "$gline" | awk '{print $2}'); if [ -z "$gh" ]; then gmatch="genesis line not seen in 60 s; "; elif [ -n "$GENESIS" ] && [ "${gh#${GENESIS:0:8}}" = "$gh" ]; then gmatch="genesis MISMATCH($gh); "; else gmatch="genesis $gh MATCH; "; fi
# the re-cut's binary prints "Base unit: 10^18" at start and the 5 October one never does (the testnet lane, 7 Oct 2026)
if journalctl -u igneumd --since "-90 s" --no-pager 2>/dev/null | grep -q "Base unit: 10^18"; then gmatch="${gmatch}base unit 10^18 line seen; "; else gmatch="${gmatch}NO base unit line; "; fi; fi
down=$(( $(date +%s) - t0 ))
commits=$(grep -a -c "$COMMIT" igneumd || true)
first=$(journalctl -u igneumd --since "-40 s" --no-pager 2>/dev/null | grep -vE "Started|Stopped|Stopping|Deactivated|Consumed" | head -1 | cut -c1-120)
dig=$(journalctl -u igneumd --since "-40 s" --no-pager 2>/dev/null | grep -oE "Consensus params digest: [0-9a-f]+" | tail -1 | awk '{print $4}')
syncing=$(curl -s -m 5 -H 'content-type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"eth_syncing","params":[]}' http://127.0.0.1:26890 | cut -c1-80)
peers=$(curl -s -m 5 -H 'content-type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"net_peerCount","params":[]}' http://127.0.0.1:26890 | grep -oE '"result":"[^"]*"' | cut -d'"' -f4)
dmatch=no; [ -n "$dig" ] && [ "$dig" = "$DIGEST" ] && dmatch=MATCH; [ -n "$dig" ] && [ "$dig" != "$DIGEST" ] && dmatch="MISMATCH($dig)"
echo "unit $(systemctl is-active igneumd) down ${down}s; ${wiped}${gmatch}commit strings $commits; digest $dmatch; eth_syncing $syncing; peers $peers; first: $first"
REMOTE
fi
} >> "$log" 2>&1; rc=$?
line=$(tail -1 "$log" | cut -c1-300)
echo "RESULT $name ($ip) rc=$rc after $(( $(date +%s) - t0 )) s: $line"
}
case "$mode" in
seeds)
[ -n "$BIN" ] && [ -n "$SHA" ] && [ -n "$DIGEST" ] || die "seeds needs --igneumd --sha256 --digest (and --override when the seeds take one)"
[ -f "$BIN" ] || die "binary file missing"; [ -z "$OVERRIDE" ] || [ -f "$OVERRIDE" ] || die "override file missing"
[ "$WIPE" = 0 ] || [ -n "$GENESIS" ] || die "--wipe-genesis needs --genesis <hex> for the read-back"
local_sha=$(shasum -a 256 "$BIN" | cut -d' ' -f1); [ "$local_sha" = "$SHA" ] || die "the binary's sha256 is $local_sha, not $SHA"
[ -z "$OVERRIDE" ] || python3 -c 'import json,sys; json.load(open(sys.argv[1]))' "$OVERRIDE" || die "override file does not parse"
say "seeds: $( [ "$GO" = 1 ] && echo GO || echo DRY RUN ); igneumd $SHA; override ${OVERRIDE:-none}; digest $DIGEST; commit $COMMIT; wipe-genesis $WIPE${GENESIS:+ (genesis $GENESIS)}"
for n in seed1.testnet seed2.testnet seed3.testnet; do one_seed "$n" & done; wait
say "seeds done; logs in $S" ;;
hands)
[ -n "$NODE" ] && [ -n "$OVJSON" ] && [ -n "$DIGEST" ] || die "hands needs --node --override-json --digest"
g=""; [ "$GO" = 1 ] && g="--go"
say "hands: $( [ "$GO" = 1 ] && echo GO || echo DRY RUN ) through move-hand.sh"
"$HERE/hands/move-hand.sh" binary --node "$NODE" --override-json "$OVJSON" $g
"$HERE/hands/move-hand.sh" restart observer-node --digest "$DIGEST" $g
"$HERE/hands/move-hand.sh" restart node1 --digest "$DIGEST" $g ;;
lift-rpc-filter)
ip=$(seed_ip seed1.testnet)
if [ "$GO" = 0 ]; then
say "DRY RUN: on root@$ip: back up /opt/igneum/bin/rpc-filter.py, set BLOCKED_UNTIL_FIXED_NODE = set(), py_compile, systemctl restart igneum-rpc-filter, read back"
"${SSH[@]}" "root@$ip" 'echo "current: $(grep -c "^BLOCKED_UNTIL_FIXED_NODE = {" /opt/igneum/bin/rpc-filter.py) block set(s), unit $(systemctl is-active igneum-rpc-filter.service)"'; exit 0; fi
"${SSH[@]}" "root@$ip" bash -s <<'REMOTE'
set -euo pipefail
f=/opt/igneum/bin/rpc-filter.py; cp -a "$f" "$f.bak-$(date -u +%Y%m%dT%H%M%SZ)"
python3 - <<'PY'
import re
p='/opt/igneum/bin/rpc-filter.py'; s=open(p).read()
new, n = re.subn(r'BLOCKED_UNTIL_FIXED_NODE = \{.*?\n\}', 'BLOCKED_UNTIL_FIXED_NODE = set() # lifted 7 Oct 2026: every seed runs the fixed node (N7 class fixed from 8097d600)', s, count=1, flags=re.S)
assert n == 1, "block set not found"
open(p,'w').write(new)
PY
python3 -m py_compile "$f"; systemctl restart igneum-rpc-filter.service; sleep 2
echo "unit $(systemctl is-active igneum-rpc-filter.service); blocked set now: $(grep -E '^BLOCKED_UNTIL_FIXED_NODE' "$f" | cut -c1-60)"
REMOTE
say "public read-back: $(curl -s -m 8 -H 'content-type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"eth_blockNumber","params":[]}' https://rpc.testnet.igneum.network | cut -c1-120)" ;;
*) die "unknown mode $mode" ;;
esac

43
scene/README.md Normal file
View file

@ -0,0 +1,43 @@
# scene: the chain scene, once
The live devnet animation on the site's home fold and `/live`, and the miner app's "The chain, live" card and its Inspect view,
are ONE renderer fed ONE JSON shape. This folder is the source; everything else is a copy or a reader.
| File | What it is | Copies |
|---|---|---|
| `live-dag.js` | EMBER 02 `IgneumDag` 2.0.4: lanes per vote key, parent curves, the selected chain, checkpoint bands, proof glow; the box height follows the lanes (`onSize`, `autoHeight`) | `site/live-dag.js`, `app/igneum-app/ui/live-dag.js` |
| `proof-core.js` | `IgneumProof` 2.0.0: the selected block's shards | `site/proof-core.js`, `app/igneum-app/ui/proof-core.js` |
| `tokens.css` | The fourteen palette tokens the two scripts read from `:root`, dark and light, the brand package's values | the `scene-tokens` block in `site/site.css` and `app/igneum-app/ui/app.css` |
| `feed-contract.md`, `feed-contract.json` | The feed shape, in words and as key lists | read by `tools/scene/feed-contract.mjs` and the app's `src/live.rs` test |
| `fixtures/live-2026-10-07.json` | One recorded `/api/live?window=300` reply | the parity, paint and contract tests |
`node tools/scene/sync.mjs` writes the copies; `--check` is the gate line (byte-equal scripts, an equal token block, the token
names defined nowhere else, a print block excepted); `--self-test` proves the check on known-failed cases first. A side (site or
app) takes part once its `live-dag.js` copy exists; a `release-*` branch checks the app's copies only, because the release
branches carry the site tree as it was when they were cut and the site deploys from master alone.
## The one rule for the two surfaces
Same renderer, same tokens, same feed, same mount options: `window: 60`, `fps: 60`, `poll: false`, the page polls
`/api/live?window=300` every 2 s and pushes. What may differ is the box the scene sits in (the home hero is 500 px tall, `/live`
420 px, the app's card 220 px in `compact` mode and its Inspect view the `/live` height), and the app's viewpoint: its `mine`
option names this machine's vote keys, so its own blocks glow and its lane reads YOUR KEY. That is an overlay on the same
picture, never a second renderer, never a rewritten feed. The phone rule (30 s window, four lanes, bigger nodes) keys on the
viewport width under 720 px, not on the canvas width: a 640 px hero on a laptop is not a phone.
The box's height follows the lanes (2.0.4): wanted height = top padding + axis padding + lanes shown times `laneHeight`
(46 px, 40 on a phone, 26 compact; at most `maxLanes` 7, `narrowLanes` 4 on a phone, never under 2). The renderer calls
`onSize(heightPx, {lanes, laneHeight, narrow, compact})` whenever that changes and exposes `getWantedHeight()`; a page that
lays out its own box sets its height from the callback. With `autoHeight` (on by default when the host set no CSS height on
the canvas, forced either way by the option, never in compact mode) the renderer sets `canvas.style.height` itself and shows
every lane up to the cap. Every current host sets a height (/live 420 px, the hero 500 px, the app's card 220 px and its
Inspect view 420 px), so until a page opts in the frames are unchanged.
## The tests (all in `tools/ci/pre-push.sh`)
| Line | What it proves |
|---|---|
| `tools/scene/sync.mjs --self-test && --check` | the copies are the source, the tokens live once |
| `tools/scene/paint-test.cjs` | a push paints with the document hidden, the observer silent and no animation frame (the blank `/live` of 7 October 2026) |
| `node --test tools/scene/feed-contract.test.mjs` | the fixture validates; a rewritten miner, a float `now`, a stray key are refused |
| `tools/scene/parity-remote.sh` | on build-2: the fixture through the site's `/live` and the app's UI, three frames each, pixel-equal apart from the app's lane label; a changed token fails first |

20
scene/feed-contract.json Normal file
View file

@ -0,0 +1,20 @@
{
"version": 1,
"doc": "scene/feed-contract.md: the one JSON shape the chain scene reads, from the observer's /api/live and from the miner app's api/live alike",
"top": ["ok", "now", "state", "blocks", "miners", "events", "finality", "proving"],
"top_extensions": ["partial"],
"state": ["stale", "age_s", "network", "node_version", "height", "block_count", "header_count", "blue_score", "difficulty", "hashes_per_second_estimate", "peers", "mempool", "blocks_60s", "blocks_per_second_60s", "blocks_per_minute", "miners_10m", "observer_started_at", "observer_lag_s", "queue_depth", "updated_at"],
"state_extensions": ["source", "you_blocks", "node"],
"block": ["hash", "number", "blue_score", "daa", "ts", "rx", "parents", "chain", "miner", "color", "locked", "final", "shards", "proven"],
"miner": ["id", "blocks", "share", "last_seen", "engine"],
"event": ["ts", "kind", "text"],
"finality": ["supported", "active", "next_index", "latest_locked_index", "latest_locked_hash", "latest_locked_blue_score", "params", "weights", "checkpoints"],
"checkpoint": ["index", "hash", "blue_score", "daa", "state", "signed", "active", "total", "fraction_active", "fraction_total", "votes", "voters", "locked_at", "updated_at"],
"shard": ["i", "n", "state", "prover", "lag", "payout", "pgas"],
"proving_unsupported": ["supported", "reason"],
"proving_supported": ["supported", "active", "activation_daa", "tip_daa", "verifier", "pool", "paid_shards_total", "shard_budget_pgas", "blocks_10m", "blocks_fully_proven_10m", "shards_proven_10m", "shards_paid_10m", "median_proof_lag_s", "provers_10m"],
"color": ["blue", "red", "pending"],
"shard_state": ["planned", "proving", "verified", "paid"],
"checkpoint_state": ["pending", "locked"],
"miner_id": "^[0-9a-f]{8}$"
}

78
scene/feed-contract.md Normal file
View file

@ -0,0 +1,78 @@
# The live feed contract
One JSON shape. The chain scene (`scene/live-dag.js`, `scene/proof-core.js`) reads it, and two producers write it:
| Producer | Where | Source of the rows |
|---|---|---|
| The observer | `site/api/live.mjs` (GET `/api/live?window=N`, igneum.network) | Neon tables the observer (`tools/observer/observer.mjs`) fills from node 1 |
| The miner app | `app/igneum-app/src/live.rs` (GET `api/live?window=N` on the engine's port) | The observer's reply, passed through; when the observer is unreachable, the local node's `igneum_getRecentBlocks` shaped into the same rows (`partial: true`) |
The machine-readable key lists live in `scene/feed-contract.json`. `tools/scene/feed-contract.mjs` validates a reply against them
(the gate runs it over the recorded fixture and over known-bad shapes); the app's Rust test reads the same file, so neither side
can add a field or drop one without the other noticing.
## Top level
| Key | Type | Meaning |
|---|---|---|
| `ok` | true | A reply with `ok: false` carries `error` and nothing else |
| `now` | ISO 8601 string | The producer's clock when the reply was built |
| `state` | object | The network state, below |
| `blocks` | array | The window's blocks, oldest first, at most 1,800 |
| `miners` | array | Vote keys seen in the last 10 minutes |
| `events` | array | The last 30 observer events |
| `finality` | object | Checkpoints and weights |
| `proving` | object | The proving layer |
| `partial` | true, app only | The rows came from the local node: no parents, no proof state, no checkpoints |
## `state`
`stale`, `age_s`, `network`, `node_version`, `height` (the chain block number), `block_count`, `header_count`, `blue_score`,
`difficulty`, `hashes_per_second_estimate`, `peers`, `mempool`, `blocks_60s`, `blocks_per_second_60s`, `blocks_per_minute`
(10 numbers, oldest first), `miners_10m`, `observer_started_at`, `observer_lag_s`, `queue_depth`, `updated_at`. A value the
producer does not know is `null`, never missing.
App extensions, additive and ignored by the scene: `source` (`observer` or `node`), `you_blocks` (this machine's blocks in
the window), `node` (the local node's own numbers for the Cards tab).
## A block
| Key | Type | Meaning |
|---|---|---|
| `hash` | string, the first 16 hex | Both producers cut the hash the same way, so a hash reads the same everywhere |
| `number` | number or null | The chain block number once planned; null off the chain or before the plan |
| `blue_score`, `daa` | number or null | |
| `ts` | number | Header time in milliseconds |
| `rx` | number | When the producer first saw the block, milliseconds |
| `parents` | array of hash strings | Empty on a `partial` reply |
| `chain` | boolean | On the selected chain |
| `miner` | string, 8 hex | The vote key id. NEVER a word: the app marks its own blocks through the scene's `mine` option, not by rewriting this field |
| `color` | `blue`, `red` or `pending` | GHOSTDAG inclusion |
| `locked` | boolean | A locked checkpoint block |
| `final` | boolean | At or before the newest lock |
| `shards` | array of `{i, n, state, prover, lag, payout, pgas}` | `state` is `planned`, `proving`, `verified` or `paid` |
| `proven` | boolean | Every shard verified or paid |
## `miners[]`, `events[]`, `finality`, `proving`
A miner is `{id (8 hex), blocks, share (percent of the 10 minutes), last_seen (ISO or null), engine (string or null)}`.
An event is `{ts, kind, text}`.
`finality` is `{supported, active, next_index, latest_locked_index, latest_locked_hash, latest_locked_blue_score, params, weights,
checkpoints}`; a checkpoint is `{index, hash, blue_score, daa, state (pending or locked), signed, active, total, fraction_active,
fraction_total, votes, voters, locked_at, updated_at}`.
`proving` is `{supported: false, reason}` or the supported object (`active`, `activation_daa`, `tip_daa`, `verifier`, `pool`,
`paid_shards_total`, `shard_budget_pgas`, `blocks_10m`, `blocks_fully_proven_10m`, `shards_proven_10m`, `shards_paid_10m`,
`median_proof_lag_s`, `provers_10m`).
## What the scene does with it
`IgneumDag.mount(canvas, opts).push(reply)` normalises every block (unknown colours become `pending`, unknown shard states
`unknown`, a hash over 256 characters or a non-numeric `ts` drops the row), keeps the newest 340 s, and paints at once.
The window the scene shows (`opts.window`, 30 to 300 s) is independent of the window the page asks the producer for: both the
site and the app ask for 300 s so the viewer can pan and zoom out to the full range (`?window=300`).
## The recorded fixture
`scene/fixtures/live-2026-10-07.json` is one real `/api/live?window=300` reply (7 October 2026, 13:17 UTC, 905 blocks, 15
miners, 40 checkpoints, 14 proven blocks, 50 excluded). The contract test validates it; the parity test renders it through both
surfaces; the paint test pushes it into a hidden document.

File diff suppressed because one or more lines are too long

Some files were not shown because too many files have changed in this diff Show more