The fleet's 22:09 UK incident (a Mac-side pkill -f <log file name> matched nothing, the roll-everything script lived on and wiped a held box) and the day's two pgrep self-matches are one class. The check flags pgrep -f / pkill -f with a plain literal (every one on a line), any pgrep/pkill on a file-name shape, and ps | grep with a literal; it allows the bracket form, -x, -F pidfile, kill $(cat pidfile), a variable and a full path; 11 banned and 16 allowed shapes in its self-test; 0.15 s over the tree. The 25 pkill -f sp1-gpu-server inside bash -c bodies (which matched the calling bash) are pkill -x; the other 11 literals take the bracket form; prover-socket-check accepts both. Row R in the record; the CLAUDE.md rule names the check and covers pkill and file names. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
23 KiB
CI failures, 4 to 6 October 2026: every non-green run classified, the fixes, the guards
Written 6 October 2026, 22:0x UK, from gh run list (430 runs, every run since the first workflow run at 09:57Z on
4 October) and gh run view --log-failed on one run per class. Times UTC (UK was UTC+1). This document is an operations
record and is listed in tools/ci/export-exclude.txt: it is not exported to the public mirror.
1. The totals
| Workflow | Runs | Green | Failed | Cancelled |
|---|---|---|---|---|
| ci | 336 | 205 | 131 | 0 |
| windows-ci | 94 | 57 | 20 | 17 |
| total | 430 | 262 | 151 | 17 |
126 of the 168 non-green runs were on master. Master was red without a break from 18:37Z to the end of the evening (20:23Z, run 37526027569) across three causes in a row: the billing block, the copied-sources check, the identity grep.
2. Every failure by root cause
One line per class. "Guard" is what now stops the class before it reaches master.
| Class | Runs | First | Last | Cause | Fix | Guard |
|---|---|---|---|---|---|---|
| A1 identity grep | 56 | 04 13:26Z | 05 01:46Z | A simulator's run log committed under sim/difficulty/records carried the Mac's home path in its first line; 43 master pushes in twelve hours each failed the same step |
The log was scrubbed; .log files joined the generic scrub (mirror e18256d) |
The pre-push gate runs the identity grep locally before any push to master or release-* (tools/ci/pre-push.sh --hook), so the hit lands on the pushing machine, not on master |
| A2 identity grep | 17 | 05 02:17Z | 05 10:36Z | vmmap dumps under docs/benchmarks/memory-floods-2026-10-04/vmmap/ carried a local time offset on their Date/Time lines |
The dumps were re-stamped to UTC | Same gate |
| A3 identity grep | 1 | 06 20:23Z | 06 20:23Z | docs/analysis/horizon/polish.md, a research review of internal tooling, quotes the overlay network's product name, the intake key's variable name and the identity guard's own regexes as text; docs/analysis is in the export list, so the grep is right to flag it |
The document is listed in tools/ci/export-exclude.txt; the identity check and the mirror's sync.sh both prune that list, so the guard's purpose (nothing the public reads carries these strings) is intact and the research text is untouched |
The exclusion list, plus the gate; identity-check.sh --self-test shows an excluded path may quote the patterns and an exported one may not |
| B1 hosted runner refused | 32 | 06 18:37Z | 06 20:05Z | "The job was not started because recent account payments have failed or your spending limit needs to be increased": every GitHub-hosted job of every run failed at start with zero steps, 26 master runs and 6 release-0.3.15 runs, and nothing told anyone | Cleared on the billing page by 20:08Z | The red job runs on the box's own runner after any failed master or release-* run and posts one line per run (section 4); the pow and sims jobs can move to the box with one repository variable (section 5) |
| C1 copied-sources | 1 | 06 18:00Z | 06 18:00Z | tools/workers/collect.mjs mentioned rsync and cargo in a comment; the check read comments |
The check skips comment lines | The gate |
| C2 copied-sources | 12 | 06 20:08Z | 06 20:19Z | infra/build-server/repro/rebuild-on-box.sh clones sources and runs cargo without touch; 10 master runs and 2 release-0.3.15 runs |
0f0abc6 (box-work 2bd3bec): the repro script re-stamps its clones. Release-0.3.15 still carries the old script at c25a3ca and will fail this step again until it takes master (or cherry-picks 2bd3bec) |
The gate |
| D no-foreign-tree-writes | 2 | 06 18:20Z | 06 18:24Z | The new check's warning pipeline (`grep | grep -v | sed |
| E Windows checkout | 3 | 04 13:53Z | 06 20:15Z | Nine screenshot files under docs/plans/site-ui-3-shots/ carried a colon from an address; git on windows-latest refuses the path, so actions/checkout died and with it every Windows build of the tree (the 0.3.15 installer waited on it) |
61f46cc renamed the nine files |
tools/ci/windows-paths-check.sh: colon and the other forbidden characters, trailing dot or space, reserved device names, over 240 characters; as the pre-commit hook on the staged paths and in the gate on every tracked path |
| F1 site build | 5 | 04 13:46Z | 04 13:53Z | site/scrub-bench.sh failed on docs/bench-log.md and build.mjs threw from the execSync |
Fixed in the bench log the same afternoon | The gate builds the site in a temporary copy before the push |
| F2 site build | 2 | 05 16:42Z | 05 16:43Z | Conflict markers left in site/journey.json; build.mjs parsed it as JSON |
Resolved by hand; no-conflict-markers.sh and the first pre-push hook (fb076de) followed |
The gate's first check, on every push |
| G link check | 1 | 05 09:28Z | 05 09:28Z | /#wallet linked from every page with no such id (release-0.3.6) |
Anchor added | The gate |
| H windows payload inputs | 4 | 05 16:30Z | 06 18:15Z | The signed inputs manifest on the downloads host pinned one node commit and packaging/windows/node-source.pin in the tree another: the Mac had pushed new inputs without committing the pin, or committed a pin without pushing inputs |
Each time, the pin and the inputs were brought level | Not a tree check and not in the gate: the shipper's push-inputs.sh writes the pin and the commit must carry it; the red job now reports the mismatch within a minute instead of the next person opening the Actions page |
| I windows installer | 1 | 04 10:32Z | 04 10:32Z | The runner image's Inno Setup was older than 6.3 | The workflow installs Inno Setup when the image's is too old | Resolved in the workflow |
| J windows engine | 2 | 04 13:53Z | 04 13:56Z | Rust that did not compile pushed to master (expected identifier, found keyword let) |
Fixed in the next push | A compile is not a 25-second check; the owner's merge rule (CLAUDE.md, "CI red is stop-the-line") covers it: the merger fixes or reverts inside 15 minutes |
| K igneum-census | 1 | 05 23:18Z | 05 23:18Z | igneum-pow's Instr, Program and a layout argument changed; igneum-census was not rebuilt (release-0.3.11) |
Updated with the crate | As J |
| L prover-socket | 1 | 06 08:49Z | 06 08:49Z | tools/proving-v1/pc2-agg-cost.ps1 ran the prover host as root without killing sp1-gpu-server (release-0.3.12) |
The playbook was fixed | The gate |
| M public API check | 1 | 06 15:12Z | 06 15:12Z | The live observer was 969 s stale when the master-only live check ran | The observer recovered; the hands moved to the box that evening | A live check stays in CI only, master only; it is not a tree fact and not in the gate |
| N no-secrets | 2 | 06 15:56Z | 06 16:23Z | A 64-hex test vector next to private_key in app/igneum-wallet/src/vault.rs (wallet-0.1.5) |
Allow-listed as a test value | The gate |
| P swallowed defaults line (no CI run: a silent class) | 0 | 06 19:5xZ | 06 21:xxZ | A comment appended to a line of shell assignments in tools/build-remote.sh and then tools/cross-remote.sh turned every assignment after the # into comment text; bash -n and shellcheck are silent on it; the default cross-build never ran and its chain kept the previous exes from about 20:40 to 22:00 UK |
36e4ee7 on master: the lines split; tools/ci/defaults-line-check.sh with its self-test |
In ci.yml at 36e4ee7 and in the gate from this change, so it runs on the pushing machine before the push |
| Q gate checks that read the machine, not the fact (found by the gate itself, 22:1x to 22:3x UK) | 0 | 06 21:1xZ | 06 21:3xZ | Two pushes of this change to master were refused by the new hook: (1) git hands a hook GIT_DIR, and remote-run.sh --self-test's nested git init, commits and reset then acted on THIS repository's worktree: it set core.bare, moved the local master to three fixture commits and broke the main checkout for ten minutes (restored from the reflog: master back to 36e4ee7, core.bare false; nothing was pushed, nothing lost); (2) the same self-test judged a stale index.lock by pgrep -x git over the whole machine, so the hook's own git push (or any other agent's git) made the fixture's checkout die with "index.lock: File exists" |
The gate unsets GIT_DIR and the other hook variables before any check; the staleness test is the lock's age (over 30 s), and the self-test backdates its fixture lock |
The gate itself: every self-test now runs inside a real hook before every master push, with a git push alive beside it |
| R kill by name (the fleet, no CI run) | 0 | 06 21:09Z | 06 21:09Z | A Mac-side pkill -f <log file name> matched nothing: the name was a shell redirect, not part of any command line; the roll-everything script lived on and wiped a box it had been told to hold. Earlier the same day, twice: a pgrep -f "<literal>" matched the calling shell's own command line. Six pkill -f sp1-gpu-server inside bash -c '...' bodies in tools/proving-v1 carried the same shape on master |
The fleet runs Mac-side jobs under tools/fleet/fleet-bg.sh (a pid file per job) and anchors every on-box kill on the binary's full path and first argument; the six prover lines use pkill -x on the binary name |
tools/ci/kill-by-name-check.sh in the gate: flags pgrep -f / pkill -f with a plain literal, any pgrep/pkill on a file-name shape (.log, .out, .pid, .json ...), and `ps |
| O1 windows-ci cancelled | 16 | 04 10:42Z | 06 18:12Z | concurrency: cancel-in-progress on windows.yml: a newer master push superseded the run. Not a failure |
None needed | None; they are listed because gh run list counts them as non-green |
| O2 hosted runner not acquired | 8 | 05 19:26Z | 05 20:54Z | "The job was not acquired by Runner of type hosted even after multiple attempts" on release-0.3.10 (7) and master (1): GitHub capacity, retried by hand | Re-run | The red job reports it; the box runner for pow and sims (section 5) takes those jobs off the hosted pool |
Sum: 168 runs, plus class P, which never reached CI because nothing checked for it. Classes A1 to A3, C1, C2, D, E, F1, F2, G, L and N are 102 runs (61 percent), every one a tree check that finishes in under 25 s on the pushing machine. B1 and O2 are 40 runs (24 percent) of GitHub-side refusals that nobody saw until the Actions page was opened. O1 is 16 runs (10 percent) of expected cancellations. H, I, J, K and M are the remaining 10.
3. The gate: one script, local and CI (tools/ci/pre-push.sh)
Every fast tree check CI runs is in one script. The site job of ci.yml calls tools/ci/pre-push.sh --ci; the pre-push
hook calls tools/ci/pre-push.sh --hook. The two cannot drift because there is one list. A check added to ci.yml alone
is the wrong place; it goes in the script.
| Mode | When | What |
|---|---|---|
--hook on a push to master or release-* |
installed by tools/ci/install-hooks.sh into the shared hooks directory (one set for every worktree) |
all 32 checks; a red check refuses the push and prints its output |
--hook on any other ref |
same | the two structural checks only (conflict markers, Windows paths) |
--ci |
the site job |
all 32 checks, with the site built in place |
| default | by hand in any worktree | all 32 checks |
--self-test |
in the gate itself | a known failure is RED and fails the gate; a known success is ok; master and release-* select the full gate, other refs the light one |
Measured 6 October 2026, 21:5x UK, on the Mac: 30 checks, GREEN, 25 s (no-secrets 11 s, identity grep 3 s, the rest
under 2 s each). The hook never writes into the worktree: the site is built in a temporary copy with
SITE_DOWNLOADS_OFFLINE=1 (063bbca: the earlier hook built in place and rewrote the downloads snapshot in five
worktrees); git status before and after the full gate is identical.
Checks that joined CI through the gate and were not in ci.yml before: the ledger sentence check
(ledger-text-check.mjs), the workflow shell parse (check-workflow-shell.mjs), the Windows paths check, and the three
self-tests (identity, Windows paths, the red watcher). The swallowed-defaults check (36e4ee7) and the kill-by-name check are in the gate too.
4. The red watcher (tools/ci/red-watch.mjs, infra/build-server/ci-red/)
A red job in ci.yml and windows.yml runs only when a master or release-* run has a failed job. It runs on the box's
own runner (igneum-build-1), not on a GitHub-hosted machine, because the hosted pool is the thing that was refused in
B1 and O2. It appends one JSON line for the run (id, workflow, branch, commit, title, the failed jobs and each one's first
failed step from the run's own API, the URL) to /srv/ci-red/red.jsonl, idempotent per run attempt. On the box,
igneum-ci-red.timer runs the poster every minute as build: each line not yet posted goes once to the hidden updates
channel through DISCORD_WEBHOOK_UPDATES in /srv/discord-hooks/env, then its run id is recorded in
/srv/discord-hooks/ci-red-posted.json. The orchestrator reads the file (ssh build@<box> cat /srv/ci-red/red.jsonl)
or the channel. No URL is ever printed; a missing key is logged by name.
Shown on 6 October 2026: the self-test (one line however often record runs; the dry run sends nothing; a missing key
is named, never a URL; one live send per run; a webhook error keeps the run pending). The poster is installed and
active on the box (22:55 CEST, "nothing to post (0 recorded)"). Open: the updates channel has no webhook yet, so the
first real red run will land in red.jsonl and the poster will log the missing key until DISCORD_WEBHOOK_UPDATES is
added to ~/.config/igneum/discord on the Mac and infra/build-server/discord-hooks/install.sh is re-run. The
Actions-side trigger has not fired on a real red run yet (master was made green in the same change); the first red
master or release-* run is its known-failed case.
5. Where CI runs, and why ci takes about three minutes
| Job | Where today | Time on ubuntu-latest | Time on the box (measured 6 October) | Note |
|---|---|---|---|---|
pow (igneum-pow cargo test --release, packfile test, igneum-census build) |
ubuntu-latest | 2 min 30 s to 3 min, cold every run (no cache action) | 42 s cold as the runner user (99 tests), sccache read-only hits after the first build | the long pole |
| sims (two Python simulators, --quick) | ubuntu-latest | about 1 min with setup-python and pip | python3 and numpy are on the box from provision.sh | |
| site (the gate) | ubuntu-latest | under 1 min | not moved: the live public API check belongs on a neutral egress | |
| red | the box's runner | seconds | only after a failed master or release-* run |
The workflow now reads the repository variable IGNEUM_CI_RUNNER: box sends pow and sims to
[self-hosted, linux, x64, igneum-build-1], anything else keeps ubuntu-latest (docs/plans/ci-self-hosted.md: GitHub
has no fallback in runs-on, so a variable is the switch; gh variable set IGNEUM_CI_RUNNER --body box as igneum-labs,
gh variable delete IGNEUM_CI_RUNNER to come back). Recommendation: flip it. The two compile-or-compute jobs are what
GitHub's minutes and the billing block were spent on, the box compiles the agents' own pinned rustc 1.99.0, and a ci
run drops from about three minutes to about one. The hosted runner then serves only the gate and the Windows
pipeline (MSVC, WebView2, Inno Setup, PowerShell 5.1, which a Linux box cannot provide). The flip is main's call after the
0.3.15 cut, per build-server.md section 7.1.
6. The build box's own red rows (the project lead, 22:3x UK: "also make sure we are fixing and learning from all the errors here")
Source: /srv/builds/_log/builds.jsonl on igneum-build-1, 139 rows from the first build at 18:38 UK to 22:11 UK on
6 October; 34 with a non-zero exit. Until this change the box kept no output of a run (it streamed to the agent's terminal),
so a red row could not be read afterwards; the classes below are from exit code, duration, compile count and what the next
row of the same worktree did. Times UK. "Iteration" = the same worktree and kind green inside ten minutes.
| Time | Worktree, kind | Exit, secs, compiles | What it was | Kind | Would a local pre-check have saved the round trip | Guard now |
|---|---|---|---|---|---|---|
| 18:49:12 | build-server, suite --lib jobbuild on app/igneum-app |
101, 0 s, 0 | Instant death: --lib on a crate whose tests live in the bin (the next row, --bin, was green the same minute) |
real class: instant | Yes, one second | pre-flight and the kept run log; class instant on the row and in the red-run file |
| 19:22:14, 19:51:09 | ship0314 then ship0315, prove --features igneum-prove-host/cuda |
101, 0 s, 0 (twice) | Instant death, the same command twice on two branches: the feature name (nothing on the box kept the message) | real class: failed the same way twice, instant | Yes | pre-flight refuses a feature the package does not have (preflight-feature, exit 3, before the slot) |
| 19:26:01 | master, check -p igneum-prove-core |
101, 0 s, 0 | Instant death: the package name (no such member) | real class: instant | Yes | pre-flight refuses a -p that names nothing (preflight-package) |
| 19:27:26, 19:27:54 | master, prove | 101, 0 s, 0; then 101, 55 s, 605 | An instant death, then a compile error; green 4 min later | iteration (with one instant) | The instant one, yes | as above; compile-error class with the run log kept |
| 19:35:56, 19:42:36, 19:48:37, 19:58:12, 21:02:56, 21:47:14 | ship0315, suite and node-linux | 101, 19 to 66 s, 44 to 724 compiles | The shipper's 0.3.15 node fixes: compile and test failures, each green within 1 to 8 min | iteration | No: a Mac check takes 12 to 18 min, the box IS the pre-check | run log kept; class compile-error or test-failure |
| 19:51:09 | ship0315, suite | 2, no duration | no-dir: the crate directory was not there, because a second run of the same worktree started in the same second and its checkout was replacing the tree |
real class: a shared resource (one worktree directory, two runs) | n/a | one run per worktree directory at a time: remote-run.sh takes a per-worktree lock in checkout and run mode and waits (shown: the second of two concurrent runs waited and both finished) |
| 19:50:57, 19:52:50 | ca3-v4-node, suite | 101, 12 s and 7 s | Test failures while fixing; green 2 min and 0 min later | iteration | No | run log kept |
| 20:45:55 | box-work, night battery dry run | 1, 234 s | The night battery's dry-run subset failed in the box-work agent's hands; its output was on that agent's side | one-off, unread | n/a | the run log is kept from now on; the night battery writes its own report |
| 21:01:12 | build-server, self-test-repro |
1, 0 s | The repro self-test's first version; green 1 min later | iteration | n/a | none needed |
| 21:17:57 | ca3-v4-node, cargo audit |
2, 0 s, 0 | Instant death: cargo-audit is not installed on the box | real class: a missing tool, instant | Yes | pre-flight refuses a cargo subcommand the box does not have (preflight-subcommand); the install belongs in provision.sh (open: add cargo-audit there, the night battery runs it) |
| 21:17:58 to 21:26:42 | finality-pause, two suites alternating | 101, 4 to 26 s | Six reds, the same two commands three times each while the finality pause tests were being fixed; green 2 to 11 min after each | iteration (the first pair outside ten minutes) | No | run log kept; the digest counts the class |
| 21:45:13, 21:46:06 | box-capacity, node-linux p2p-probe | 101, 39 s and 3 s | Compile errors; green 1 min and 0 min later | iteration | No | run log kept |
| 21:58:26 | ca3-v4-node, check | 101, 12 s | Compile error; green 1 min later | iteration | No | run log kept |
| 22:00:08, 22:06:11, 22:08:16 | ca3-v4-node, suite (finality, then clock_rule_v3 a_pause_carries ...) |
101, 24, 13, 12 s, 22 compiles | The clock rule v3 tests failing while being fixed; green at 22:10 | iteration | No | run log kept |
| 22:08:52, 22:08:55, 22:08:57 | ca3-v4-node, three suites in five seconds | 101, 0 s, 0 (three times) | Three instant deaths in a row with nothing compiled: cargo refused before building (a manifest or lock the overlay carried mid-edit, or a name); the three commands were green one minute later, so the tree was fixed under them | real class: instant, failed the same way three times | Yes, one second each | pre-flight (manifest, package, feature, subcommand) before the slot; class instant flagged; the run log kept so the next one is read, not guessed |
| 22:11:38 | ca3-v4-node, check --tests |
101, 49 s, 844 | A compile error in the test targets; still red when the log was read | iteration, open | No | run log kept |
Sum: 34 red rows. 22 are iteration (an agent taking a red to green inside ten minutes, the box doing the compile the Mac
cannot do in time). 12 are real classes, in three families: instant deaths (8 rows: a feature, a package, a target, a
tool, and three unread), a shared worktree directory (1 row), and the night battery's unread dry run (1 row, plus two
instant ones inside iterations). Every real class now has a guard in infra/build-server/remote-run.sh, shown on the box
in a sandbox (a pass, a failing test, an empty filter, a missing package, a missing feature, a missing subcommand, a compile
error, a broken manifest, two concurrent runs of one worktree):
| Guard | What it does | Class it stops |
|---|---|---|
| pre-flight | before the slot's time is spent: cargo <sub> exists, the manifest parses (cargo metadata --no-deps), every -p is a member or a dependency, every --features pkg/feat exists; a refusal is exit 3 in about a second with the reason |
instant (manifest, package, feature, tool) |
| kept output | the last 400 lines of every run in /srv/builds/_log/runs/<id>.log, named in the row (run_log) |
every unread class |
| class on the row | class in builds.jsonl: compile-error, link-error, test-failure, instant, slot-timeout, no-dir, no-test-matched, preflight-*, other |
the digest can count what kind of red it was |
| no-test-matched | a cargo test with a filter that ran 0 tests in every binary exits 3 instead of a green "0 tests" |
a wasted round trip that read as ok |
| one run per worktree | a per-worktree lock in checkout and run mode; the second waits up to 2 h and says so | no-dir, the half-replaced tree |
| the red-run file | every red row is appended to /srv/ci-red/red.jsonl (the file the CI watcher writes), "source":"box"; not posted alone |
|
| the daily digest | at the first pass at or after 09:00 London, one line to the updates channel: reds in 24 h, CI and box, per class, with each class's guard | the learning, read once a day |
7. What belongs on another branch
| Branch | One-line change |
|---|---|
| build-server (provision.sh) | install cargo-audit for the build user (the night battery and the ca3 lane call it; pre-flight now refuses it with a clear line instead of a 0-second exit 2) |
| release-0.3.15 | take master (or cherry-pick 2bd3bec): infra/build-server/repro/rebuild-on-box.sh re-stamps its clones, else the copied-sources step fails again at the next push. Its own tools/ci/windows-paths-check.sh (c25a3ca) is superseded by master's: on merge keep master's file and drop the extra ci.yml step, the gate runs it |