Thirteen failure emails between 15:26 and 16:53 UK. The classes and what closes them:
- the box-locks check on a hosted runner (10 runs): closed by bd6fcb88 and 165e8b35 earlier
- windows-ci's stale payload-inputs pin (3 runs): closed on master by 4b4e1bc1; update-return's dispatches still carry e69e8a39
- three hosted site jobs on master hung in the tree gate for over two hours (no timeout-minutes): ci.yml now carries
site 15, changes 10, pow 60, sims 45, the overlap sweep runs under a 10-minute wall clock where GNU timeout exists, and
tools/ci/workflow-timeouts-check.sh fails a job without a budget (self-test: a job without the key, a wrong budget)
- a branch merged with no ci run of its own (era-vdf 61421005, 16:31 UK): master's igneum-pow suite went red and five
docs-only merges landed green over it because their runs skip the compile job. tools/ci/ci-state.mjs reads the runs
API through gh (a commit's newest run, master's last COMPILED run, a branch's last red); merge-to-master.sh pushes an
unrun branch for a run, waits for a queued one printing the clock, refuses a red one and refuses any merge onto a red
master except the declared fix (--fixes-master); the pre-push hook refuses a push to master whose commit, or whose
merge's branch parent, has no green run on that exact sha; a feature-branch push prints the branch's previous red
first. Self-tests with a fake gh in all three.
- ci-red.yml fires on failure, cancelled and timed_out and hands the conclusion to red-watch.mjs, whose line names the
kind (CI red, CI cancelled, CI timed out); the self-test reads the workflow file for the three conclusions
- tools/ci/retry-once.sh: one retry before red for the box-locks check, the scene parity check and the live public API
check (each keeps its own skip line on a runner without the resource)
GitHub's branch protection cannot be applied: the organisation is on the free plan and the repository is private (the
API answers 403, "Upgrade to GitHub Pro or make this repository public"), so the two scripts are the enforcement; the
rule is one line in CLAUDE.md under the CI block.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
31 runs were queued on igneum-build-1's one runner at 13:15 UK on 7 October 2026 and nothing had concluded since 13:03Z, so no lane could read a conclusion.
- ci.yml: a `changes` job (ubuntu-latest) classifies the push with tools/ci/docs-only-check.sh (docs/, site/, *.md only = code=false; a new branch, a pull request, a force push or an API error = code=true); pow and sims need it and run only on code=true. The site job is unchanged on ubuntu-latest for every push. The classifier's self-test is in the gate.
- provision.sh and runner/register.sh: `--host <ip>` registers another box, forwarding BOX_HOSTNAME, RUNNER_NAME, RUNNER_LABELS, RUNNER_CPUS and RUNNER_JOBS (plain words only); RUNNER_CPUS writes AllowedCPUs into the service drop-in beside Nice=10, so igneum-build-2's runner is bounded like a suite (32 cores). The pool label is igneum-build-1 (both boxes carry it); ci-red marks the box with the record file and the poster, the default labels carry it, and igneum-build-1 got it through the runners API today.
- ci-red.yml runs on the ci-red label, so the red line always lands where the poster reads it.
- CLAUDE.md: the rule reads "read the conclusion when it lands, own a red before the next push"; pushes are never held.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
GitHub runs a workflow_run workflow from the default branch only, so a feature branch no longer waits for a merge of master before its reds are posted. record() reads the FAILED run from RED_WATCH_* (id, attempt, workflow, branch, sha, event, url, actor, title, author: the workflow_run payload) and asks the jobs API for that run, not the watcher's own; the inline GITHUB_* shape stays for a branch with the old `red` job (one line per run id either way). The inline job leaves ci.yml. Self-test covers the workflow_run shape and the full line.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>