Merge box-work: the self-hosted CI runner, two build slots with jobs halved, the measure hold, the CPU prover trial

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 19:33:33 +00:00
commit bace6d52e2
8 changed files with 507 additions and 43 deletions

View file

@ -115,6 +115,71 @@ Also logged for context: the Mac's Linux cross-build with zig (`infra/cross/buil
|---|---|---|
| No zig / cargo-zigbuild | the devnet seed (Debian 12, glibc 2.36) takes the Mac's zig build; a native box build links glibc 2.39, which Debian 13 seeds accept and HiveOS (Ubuntu 18/20 base) does not | install zig 0.17 + cargo-zigbuild in provision.sh, add `--target x86_64-unknown-linux-gnu.2.36` mode to build-remote.sh |
| No macOS target | agents who run nodes on the Mac still build there | out of scope (needs the macOS SDK on Linux); the fleet or the box's own Devnet 2 seed takes the test-network runs instead |
| No CI runner | GitHub `ci.yml` and `windows.yml` run on GitHub's machines | install a self-hosted runner as user build once R1 is in |
| No CI runner | DONE 6 October 2026, 19:19Z (section 7): the runner `igneum-build-1` is online under user `runner`, never build; the workflow change is proposed in docs/plans/ci-self-hosted.md | main flips `IGNEUM_CI_RUNNER=box` after the shipper's cut |
| Byte identity with the Mac's exes | different C/C++ toolchain (Homebrew mingw vs Ubuntu GCC 13) and embedded source paths | not a goal; the box is identical with itself build to build, cross-remote.sh reports sha256 and the DLL list per exe |
| Robot API | `~/.config/igneum/hetzner-token` is the Cloud token (hcloud); the dedicated box lives in Robot, a separate credential | main sets the Robot server name in the UI; a webservice user goes to `~/.config/igneum/robot-credentials` when needed |
## 7. The box's second shift (6 October 2026, from 20:3x UK, the project lead: "what else can our building machine be working on?")
Four items, each its own commit on branch `box-work` with its section here. Times UTC.
### 7.1 The GitHub Actions runner (DONE 19:19Z)
| Fact | Value |
|---|---|
| Runner | `igneum-build-1`, actions/runner 2.338.0 (tarball sha256 af4b794c... checked against the release note), registered on igneum-network/igneum at 19:19:49Z, online, labels `self-hosted, Linux, X64, igneum-build-1` |
| User | `runner` (uid 1001, own group, no sudo, not in `build`'s group; /home/runner 750). Never `build`, never root. The runner's credential (`/opt/actions-runner/.credentials`, mode 600) is the only thing it holds; it signs nothing and reaches no hand |
| Service | `actions.runner.igneum-network-igneum.igneum-build-1.service` (GitHub's `svc.sh install runner`), drop-in `igneum.conf`: Nice 10, IO best-effort 7, Restart on-failure. Agents' builds (nice 0 through build-remote.sh) win the CPU over a CI job |
| Toolchains | rustup 1.99.0 pinned like the box (`RUST_TOOLCHAIN`), targets x86_64-unknown-linux-gnu and x86_64-pc-windows-gnu, clippy, rustfmt; mingw-w64 GCC 13 posix, Node 22, python3 + numpy from the system (numpy added to APT for the simulators) |
| sccache | `/usr/local/bin/sccache` (the build user's binary copied; /home/build is 750), config `/home/runner/.config/sccache/config` with `rw_mode = "READ_ONLY"` on /srv/sccache, own server port 4227. Shown: a job-shaped `cargo test --release` of igneum-pow as `runner` (99 tests pass, 42 s cold) made 6 compile requests, 0 hits, 6 cache WRITE ERRORS (the refusal, as wanted), and /srv/sccache stayed at 5,880,836 KB |
| Jobs | `.env`: RUSTC_WRAPPER, SCCACHE_CONF, SCCACHE_SERVER_PORT 4227, CARGO_INCREMENTAL 0, CARGO_BUILD_JOBS 48 (half the box); `.path`: the runner's cargo bin, /usr/local/bin, /usr/bin, /bin. A CI job takes NO build slot today (ci-self-hosted.md, open row) |
| Ephemeral | no. `--ephemeral` is for autoscaled fleets that register a fresh runner per job; one standing runner on a private repository keeps its registration and cleans `_work` per job (approximate: GitHub's docs host answered 404 to both fetches tonight, so this is the rule as remembered, labelled so) |
| Registration | `infra/build-server/runner/register.sh`: gh as igneum-labs (fails on any other active account, checks the login is igneum-labs), `POST repos/igneum-network/igneum/actions/runners/registration-token`, the token as the first stdin line to `provision.sh` on the box (never an argument, never a file, never logged; the output is filtered for it as a belt). `--status` lists the repository's runners and the unit |
| Idempotent | provision.sh runs 3 and 4 after the registration: `runner: ok`, no restart (`ActiveEnterTimestamp` unchanged). Run 2 had said `changed` because GitHub's `svc.sh install` runs `env.sh`, which rewrites `.env` and `.path`; the step now writes them AFTER the install |
| Workflows | NOT changed (the shipper owns them tonight). The proposed diff and the fallback (repository variable `IGNEUM_CI_RUNNER`; GitHub has no "else" in `runs-on`) are in `docs/plans/ci-self-hosted.md`. `windows.yml` cannot move to a Linux box (MSVC, WebView2, Inno Setup, PowerShell 5.1) |
Consequences: ci.yml's `pow` and `sims` jobs would run on a pinned 1.99.0 (GitHub's `ubuntu-latest` ships whatever stable it has), with 48 jobs; GitHub-hosted minutes on a private repository are the thing saved. A CI job on the box reads only what the checkout gives it; the mirrors and `/srv/builds` belong to `build` and are not readable by `runner` (git's safe.directory is set for the runner so a future job may clone a mirror read-only if main wants it).
### 7.0 The slots ruling (main, 20:3x UK; DONE 19:28Z, live on the box)
| Change | Where | Shown by |
|---|---|---|
| 2 build slots (`/srv/builds/_locks/slots` = 2) | provision.sh `SLOTS` default 2, applied 19:28:18Z | `dirs: changed (... slots=2)` |
| CARGO_BUILD_JOBS 90 when a build holds the only taken slot, 45 when it sees the other slot held (one second of settling after taking the slot, then a `flock -n` probe of the other file); a `-j` on the cargo line wins | remote-run.sh; tools/build-remote.sh and tools/cross-remote.sh pass `-j` only when `--jobs` is given | `remote-run.sh --self-test-slots` on the box: two concurrent fake builds get 45 each, a lone one 90 |
| A measurement (`BR_MEASURE=1`) takes the `measure` file exclusively; builds hold it shared for their whole run, so a measure waits for the running builds and blocks new ones, as with-lock.sh's `measure` on the Mac | remote-run.sh; `infra/build-server/prover/cpu-trial.sh` is its first user | the self-test: a measure blocks a build, a build blocks a measure; JSONL carries `"jobs"` and `"measure"` |
| Lock files opened in APPEND mode | remote-run.sh | the first version's `exec {fd}>build-k` truncated a BUSY slot's holder line each time another build probed it (the dashboard read empty lines for held slots); the self-test's case 5 keeps a holder line through a probe. The OLD script under the same cases: `JOBS=none JOBS=none`, FAIL (the known-failed run, 19:25Z) |
| An environment IGNEUM_BUILD_SLOTS_DIR or IGNEUM_BUILD_LOG_DIR wins over the profile | remote-run.sh | the first self-test run let the profile reset the scratch dir and took the box's REAL slot for 7 s (19:24Z, box idle) |
Open: a worktree whose remote-run.sh predates this keeps the old behaviour until it has master with it (the script is piped from each Mac worktree per build), so until every agent rebases, a build from an old worktree still asks `-j 90` beside a new one at 45. The build-server agent was told at 19:28Z.
### 7.4 The CPU prover trial (DONE 19:30Z; verdict: the box is NOT a prover)
`infra/build-server/prover/cpu-trial.sh` on the box under the measure hold (builds excluded), `igneum-prove-host` from master
7483fb37 (the 0.3.15 prover pair, sha256 71bc2438...; its `cuda` feature changes nothing under `SP1_PROVER=cpu`), fixture
`proving/fixtures/block-56-transfers.json` (the v0 block: 3 transfers, 600 pgas, one shard, 556,369 SP1 cycles), `--mode shard
--shard 0`, `RAYON_NUM_THREADS=96`, nice 19, box otherwise idle (load 2.6 at start). Log and results JSON:
`/srv/builds/_log/prover-trial/trial-20261006T192824Z.{log,json,txt}`; JSONL line kind `measure`.
| Stage | Box, 96 threads (EPYC 9454P) | Mac, same statement class (bench-log) |
|---|---|---|
| setup (prover client, shard and aggregator keys) | 14.4 s (client 12.6, keys 1.8) | 9.3 s per invocation in the v0 loop (4 Oct, "nearly all SP1 setup") |
| execute | 0.28 s, 556,369 cycles, 927 cycles per pgas | 0.19 s for block 78 (3 Oct) |
| core proof | **34.2 s**, 7,317,561 B, verify 0.34 s | 83 s for the 200-pgas shard on a Mac at load 40 (4 Oct); 71.7 s is main's Mac figure for this fixture (its stage not recorded here, labelled approximate) |
| compressed proof (what a record carries) | **85.9 s**, 1,272,897 B, verify 0.07 s | 272 s for the 200-pgas shard on the loaded Mac (4 Oct); 61 s for the smallest shard in the 3-node v0 loop |
| whole run, wall | 136.9 s | |
| peak RSS | 28.2 GB (VmHWM) | |
| CPU use | 64 of 96 threads busy on average (6,408 percent in `ps`) | |
What the numbers mean, per tier, and what follows:
| Number | Means | Done or proposed |
|---|---|---|
| core 34.2 s under 60 s, compressed 85.9 s over it | the proof a record carries is the compressed one, so the shard that matters takes 120 s of proving on 96 CPU threads for the SMALLEST shard the chain has (600 pgas, 0.56 M cycles); a shard at `S_p` is 60 M cycles (bench-log 4 Oct), about 100x, so hours per shard on this CPU against the 60 s proof lag the litepaper states | the 60 s test of the ask is NOT met on the stage that counts; NO standing CPU prover unit is written, nothing joins the devnet from the box (R6 holds: the box is never a node host for the live devnet) |
| 28.2 GB peak for the smallest shard | a CPU prover needs 32 GB of RAM for a toy shard; every home tier (8, 12, 16, 24 or 32 GB CARDS, 16 to 64 GB of RAM) is out of CPU proving, and the rented 4090 boxes' CPUs are not a fallback either | the app keeps "proving on the CPU (slow)" as a correctness lane only; the prover tiers are the real cards (docs/analysis, 6 Oct rented-card measurement) |
| 64 of 96 threads busy | SP1's CPU prover does not scale to the whole box; a second trial with `RAYON_NUM_THREADS=48` would show whether half the box proves as fast (then two shards side by side) | not run tonight (one slot of the box's evening); the script takes `--threads` |
| 14.4 s setup per invocation | the same per-process cost the Mac pays; a resident prover would pay it once | already the design of the app's prover loop |
The box stays a build and test machine. If main wants a CPU prover anyway for coverage (a prover that is always on, never fast),
the shape is a systemd unit as `build` with `SP1_PROVER=cpu`, a throwaway devnet key (never the OTA key, never a hand's key),
`--threads 48`, Nice 19 and the measure hold taken for the whole run, which would exclude builds for minutes at a time: that
is why it is not written.

View file

@ -0,0 +1,75 @@
# CI on the box: the self-hosted runner and the workflow change (proposal, 6 October 2026)
The runner `igneum-build-1` (labels `self-hosted, linux, x64, igneum-build-1`) is installed by `infra/build-server/provision.sh`
step_runner and registered by `infra/build-server/runner/register.sh` (docs/plans/build-server.md section 7). The workflows are
NOT changed here: the shipper owns `.github/workflows` tonight. This is the proposed diff for main.
## 1. The shape: one repository variable decides, GitHub-hosted is the fallback
GitHub has no "try this runner, else that one" in `runs-on`: a list of labels means ALL of them must match one runner, so
`[self-hosted, igneum-build-1, ubuntu-latest]` would never schedule. The fallback is therefore a repository variable read in
the expression. `IGNEUM_CI_RUNNER` = `box` sends the job to the box; unset or anything else keeps `ubuntu-latest`. Flipping it
back is one click in Settings > Secrets and variables > Actions > Variables (or `gh variable set IGNEUM_CI_RUNNER --body box`
and `gh variable delete IGNEUM_CI_RUNNER` as igneum-labs), with no commit and no queue lost: a job already queued for the box
stays queued; the next push goes to GitHub's machines.
## 2. ci.yml (the two jobs that compile or compute; the `site` job stays on GitHub's machines)
```diff
jobs:
pow:
name: igneum-pow tests, igneum-census build
- runs-on: ubuntu-latest
+ runs-on: ${{ vars.IGNEUM_CI_RUNNER == 'box' && fromJSON('["self-hosted", "linux", "x64", "igneum-build-1"]') || 'ubuntu-latest' }}
steps:
- uses: actions/checkout@v4
- name: toolchain
run: rustc --version && cargo --version
@@
sims:
name: simulators, quick modes
- runs-on: ubuntu-latest
+ runs-on: ${{ vars.IGNEUM_CI_RUNNER == 'box' && fromJSON('["self-hosted", "linux", "x64", "igneum-build-1"]') || 'ubuntu-latest' }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
+ if: vars.IGNEUM_CI_RUNNER != 'box' # the box has python3 and numpy from provision.sh; setup-python would download a second Python
with:
python-version: '3.12'
- - run: python3 -m pip install --quiet numpy
+ - run: python3 -m pip install --quiet numpy
+ if: vars.IGNEUM_CI_RUNNER != 'box'
```
What the box gives these two jobs: rustc 1.99.0 pinned (GitHub's `ubuntu-latest` carries whatever stable it ships; the box
is the Mac's version, so CI compiles what the agents compile), sccache hits from the agents' cache (read-only), 48 cargo
jobs. The `site` job is Node and shell checks and takes under a minute on GitHub's runners; moving it buys nothing and would
put `tools/ci/public-api-check.mjs` (a live HTTPS check) behind the box's egress for no reason.
Why the toolchain line still runs: on the box `rustc --version` must print 1.99.0; a mismatch means provision.sh and the
runner's rustup disagree (R5 in build-server.md) and the job should say so in its first step.
## 3. windows.yml: no change possible on this box
Every job of `windows.yml` runs on `windows-latest` for a reason the box cannot answer: the engine builds on the MSVC
target, the window host needs the Windows SDK and WebView2, the installer needs Inno Setup, the smoke run executes the exes
and the launcher under Windows PowerShell 5.1. A Linux runner has none of that. The only self-hosted option for this
workflow is a Windows runner on PC 1 or PC 2 (`actions/runner` for Windows under a service account), which conflicts with
the rule that the PCs keep only GPU and Windows-runtime JOBS through the signed job system, and is not proposed tonight.
What the box already does for Windows is upstream of this workflow: `tools/cross-remote.sh` builds `igneumd.exe` and
`igneum-miner.exe` (the payload inputs) in 1 min 44 s, and the night battery rebuilds them for the reproducibility record.
## 4. What to check after the flip (main, the first run on the box)
| Check | Where | Pass |
|---|---|---|
| the job landed on the box | the run's "Set up job" log says `Runner name: 'igneum-build-1'` | yes |
| the toolchain | the `toolchain` step prints `rustc 1.99.0` | yes |
| sccache hits | add `sccache --show-stats` as a step once, or read `/srv/sccache` size before and after: the runner's config is READ_ONLY, so the size must NOT change | size unchanged |
| the agents were not starved | `/srv/builds/_log/builds.jsonl` `secs` of the builds during the run against the same crate's earlier lines | within the usual spread |
| the fallback | `gh variable delete IGNEUM_CI_RUNNER`, push a no-op commit: the job runs on `ubuntu-latest` again | yes |
Open: a CI job on the box does not take a build slot (`/srv/builds/_locks/build-<k>`), it runs at Nice 10 with 48 jobs; if a
CI job ever delays a release build visibly, the fix is a step at the top of the job that takes a slot through
`infra/build-server/remote-run.sh`'s flock, the same file the agents use.

View file

@ -0,0 +1,66 @@
#!/usr/bin/env bash
# CPU prover trial on igneum-build-1 (6 October 2026): time the SP1 CPU prover on one real fixture shard with every thread
# of the box, at nice 19, under one build slot (a timing taken beside another build is not a number: CLAUDE.md).
# Runs ON the box as user build:
# infra/build-server/prover/cpu-trial.sh [--fixture <json>] [--host <igneum-prove-host>] [--threads 96] [--mode shard|execute] [--out <dir>]
# Defaults: the host built from the igneum repo's master on the box (/srv/builds/igneum/proving/igneum-prove/target/release,
# tools/build-remote.sh from proving/igneum-prove; the `cuda` feature in it changes nothing for SP1_PROVER=cpu), the fixture
# block-56-transfers.json (the v0 block, one shard), 96 threads, --mode shard --shard 0.
# Writes <out>/trial-<utc>.{log,json,txt}: the host's own results JSON (execute, core, compressed and verify seconds), the
# wall time, and the peak resident memory read from /proc/<pid>/status VmHWM every second (no `time` package on the box).
set -euo pipefail
[ -f /etc/profile.d/igneum-build.sh ] && . /etc/profile.d/igneum-build.sh
FIXTURE=/srv/builds/igneum/proving/fixtures/block-56-transfers.json
HOST=/srv/builds/igneum/proving/igneum-prove/target/release/igneum-prove-host
THREADS=$(nproc); MODE=shard; OUT=/srv/builds/_log/prover-trial
while [ $# -gt 0 ]; do
case "$1" in
--fixture) FIXTURE="$2"; shift 2 ;; --host) HOST="$2"; shift 2 ;; --threads) THREADS="$2"; shift 2 ;;
--mode) MODE="$2"; shift 2 ;; --out) OUT="$2"; shift 2 ;; *) echo "unknown argument $1" >&2; exit 2 ;;
esac
done
[ -x "$HOST" ] || { echo "no prover host at $HOST (cd proving/igneum-prove && tools/build-remote.sh -- build --release -p igneum-prove-host)" >&2; exit 2; }
[ -f "$FIXTURE" ] || { echo "no fixture at $FIXTURE" >&2; exit 2; }
mkdir -p "$OUT"
STAMP=$(date -u +%Y%m%dT%H%M%SZ); LOG="$OUT/trial-$STAMP.log"; RES="$OUT/trial-$STAMP.json"; SUM="$OUT/trial-$STAMP.txt"
# the measurement hold (main's ruling 6 October 2026): remote-run.sh BR_MEASURE=1 takes the `measure` file exclusively, waits
# for every running build and blocks new ones until this run ends, and writes the run's JSONL line (kind measure)
RR="${IGNEUM_REMOTE_RUN:-$(dirname "$0")/../remote-run.sh}"; [ -f "$RR" ] || RR=/srv/builds/_bin/remote-run.sh
[ -f "$RR" ] || { echo "no remote-run.sh beside this script or at /srv/builds/_bin" >&2; exit 2; }
echo "cpu-trial: host $HOST; fixture $FIXTURE; threads $THREADS; mode $MODE; load $(cut -d' ' -f1-3 /proc/loadavg)" | tee "$LOG"
"$HOST" --version 2>/dev/null | tee -a "$LOG" || true
args=("$FIXTURE" --mode "$MODE" --out "$RES"); [ "$MODE" = shard ] && args+=(--shard 0)
t0=$(date +%s.%N)
PEAKF=$(mktemp)
# the command remote-run.sh evals under the measure hold: the host at nice 19 with a VmHWM poller beside it
CMD="SP1_PROVER=cpu RAYON_NUM_THREADS=$THREADS RUST_LOG=${RUST_LOG:-info} nice -n 19 '$HOST' $(printf '%q ' "${args[@]}") >> '$LOG' 2>&1 & p=\$!; peak=0; while kill -0 \$p 2>/dev/null; do h=\$(awk '/VmHWM/ { print \$2 }' /proc/\$p/status 2>/dev/null || echo 0); [ \${h:-0} -gt \$peak ] && peak=\$h; sleep 1; done; wait \$p; rc=\$?; echo \$peak > '$PEAKF'; ( exit \$rc )"
BR_MEASURE=1 BR_DIR="$(dirname "$HOST")" BR_CMD="$CMD" BR_LABEL="prover cpu-trial $(basename "$FIXTURE") $THREADS threads; agent=${IGNEUM_AGENT:-box-work}" \
BR_TOOL=cpu-trial BR_KIND=measure BR_WT=igneum BR_CRATE=proving/igneum-prove BR_BRANCH="${BR_BRANCH:-master}" BR_SHA="${BR_SHA:-}" BR_AGENT="${IGNEUM_AGENT:-box-work}" \
BR_COMMAND="igneum-prove-host $(basename "$FIXTURE") --mode $MODE (SP1_PROVER=cpu, $THREADS threads, nice 19)" BR_TARGET=x86_64-unknown-linux-gnu \
bash "$RR" 2>&1 | tee -a "$LOG"; rc=${PIPESTATUS[0]}
peak=$(cat "$PEAKF" 2>/dev/null || echo 0); rm -f "$PEAKF"
t1=$(date +%s.%N)
wall=$(python3 -c "print(round($t1 - $t0, 1))")
{
echo "cpu-trial $STAMP: exit $rc, wall ${wall} s (hold included), peak RSS $(( peak / 1024 )) MB (VmHWM), threads $THREADS, nice 19, measure hold, load at end $(cut -d' ' -f1-3 /proc/loadavg)"
echo "host: $HOST ($(stat -c %s "$HOST") B, sha256 $(sha256sum "$HOST" | cut -c1-16)...)"
echo "fixture: $FIXTURE"
if [ -f "$RES" ]; then
python3 - "$RES" <<'PY'
import json, sys
d = json.load(open(sys.argv[1]))
def walk(o, p=""):
if isinstance(o, dict):
for k, v in o.items(): walk(v, f"{p}.{k}" if p else k)
elif isinstance(o, list):
for i, v in enumerate(o): walk(v, f"{p}[{i}]")
else:
s = str(o)
if any(t in p.lower() for t in ("sec", "time", "cycles", "bytes", "verified", "prover", "shard", "program", "stage")) and len(s) < 80:
print(f" {p} = {s}")
walk(d)
PY
fi
echo "log lines with times:"; grep -E -i 'prove|core|compress|verif|execute|cycles| s\b|seconds' "$LOG" | grep -v '^cpu-trial' | tail -20 | sed 's/^/ /'
} | tee "$SUM"
exit "$rc"

View file

@ -26,7 +26,8 @@
# NODE_MAJOR 22
# WORKTREES "" space-separated agent worktree names, one /srv/builds/<name> each (run-from-mac.sh fills it from
# `git worktree list` on the Mac; tools/build-remote.sh creates a missing one on first use)
# SLOTS 1 remote build slots (tools/build-remote.sh takes one; with 1 slot every build gets the box)
# SLOTS 2 remote build slots (tools/build-remote.sh takes one; remote-run.sh sets CARGO_BUILD_JOBS 90 when it holds
# the only taken slot and 45 when both are held; a measurement takes the `measure` file and excludes builds)
# P2P_PORTS "26611 26811" TCP ports ufw opens beside 22: the devnet seed's p2p (infra/seed-nodes/config.sh devnet
# P2P_PORT=26611) and the testnet seed's (26811). A suffixed devnet (Devnet 2, the fleet's staging
# chain) listens on the same 26611 (infra/cloud-devnet/config.sh P2P_PORT=26611). RPC ports
@ -52,13 +53,25 @@ SCCACHE_GB="${SCCACHE_GB:-100}"
SCCACHE_VERSION="${SCCACHE_VERSION:-}"
NODE_MAJOR="${NODE_MAJOR:-22}"
WORKTREES="${WORKTREES:-}"
SLOTS="${SLOTS:-1}"
SLOTS="${SLOTS:-2}" # 2 since main's ruling of 6 October 2026 (20:3x UK); remote-run.sh gives 90 jobs alone, 45 beside another
P2P_PORTS="${P2P_PORTS:-26611 26811}"
BOX_HOSTNAME="${BOX_HOSTNAME:-igneum-build-1}"
WORKERS_HOST="${WORKERS_HOST:-build.igneum.network}" # the dashboard feed's HTTPS name (A record in deSEC, 6 Oct 2026)
SSH_PUBKEY="${SSH_PUBKEY:-}"
BUILD_USER=build
BUILD_HOME=/home/$BUILD_USER
# the GitHub Actions runner (step_runner, 6 October 2026 evening): a dedicated user, never build and never root
RUNNER_USER=runner
RUNNER_HOME=/home/$RUNNER_USER
RUNNER_DIR=/opt/actions-runner
RUNNER_VERSION="${RUNNER_VERSION:-2.338.0}" # github.com/actions/runner releases, read 6 October 2026
RUNNER_SHA256="${RUNNER_SHA256:-af4b794c1bc41d73d40535e3fe092a39f9679cd8d965954c2aca25a05ca41d32}" # the release note's linux-x64 line
RUNNER_REPO_URL="${RUNNER_REPO_URL:-https://github.com/igneum-network/igneum}"
RUNNER_NAME="${RUNNER_NAME:-$BOX_HOSTNAME}"
RUNNER_LABELS="${RUNNER_LABELS:-igneum-build-1}" # added to the defaults self-hosted, linux, x64
RUNNER_JOBS="${RUNNER_JOBS:-48}" # cargo jobs for a CI job: half the box, the agents' builds keep the rest
RUNNER_TOKEN="${RUNNER_TOKEN:-}" # a registration token (1 h), from infra/build-server/runner/register.sh over stdin; never logged
RUNNER_SCCACHE_PORT="${RUNNER_SCCACHE_PORT:-4227}" # the runner's own sccache server; 4226 is the build user's
log() { printf '%s provision: %s\n' "$(date -u +%H:%M:%S)" "$*"; }
die() { log "ERROR: $*" >&2; exit 1; }
@ -127,6 +140,9 @@ APT_PACKAGES=(
build-essential clang lld llvm libclang-dev pkg-config libssl-dev cmake protobuf-compiler libprotobuf-dev
gcc-mingw-w64-x86-64 g++-mingw-w64-x86-64 binutils-mingw-w64-x86-64 mingw-w64-x86-64-dev mingw-w64-tools
git tmux curl ca-certificates xz-utils zstd unzip rsync jq python3 ufw htop file caddy
# the GitHub Actions runner's .NET runtime needs libicu (bin/installdependencies.sh would install it); the two Python
# simulators (sim/finality_v2.py, sim/difficulty/sim.py) need numpy, which ci.yml pip-installs on GitHub's runners
libicu74 python3-numpy
)
step_apt() {
local need=() p
@ -408,6 +424,110 @@ step_ufw() {
[ "$any" = 1 ] && changed ufw "22, 80, 443 and ${P2P_PORTS} open, everything else denied" || ok ufw "22, 80, 443 and ${P2P_PORTS}"
}
# The GitHub Actions self-hosted runner for igneum-network/igneum (asked for 6 October 2026, "what else can the box work on").
# A dedicated user `runner` (no sudo, not in the build group), the runner at /opt/actions-runner (sha256 of the tarball checked
# against the release note), its own rustup pinned to RUST_TOOLCHAIN with the two targets and clippy, Node 22 and mingw from
# the system, and sccache against the box's cache in READ-ONLY mode (`rw_mode = "READ_ONLY"`, accepted by sccache 0.18 on
# 6 October 2026): a CI job may take hits from what the agents built, never write a line into their cache, and talks to its
# own sccache server on RUNNER_SCCACHE_PORT so the build user's server never compiles as the wrong user. Registration needs
# RUNNER_TOKEN (infra/build-server/runner/register.sh fetches one with gh as igneum-labs and pipes it over stdin); without
# it the step installs everything and says what is missing. The service is GitHub's own `svc.sh install runner` unit plus a
# drop-in with Nice=10 (agents' builds through build-remote.sh win the CPU; a CI job is a check, not a release build).
# Not ephemeral: GitHub recommends `--ephemeral` for autoscaled fleets that register a fresh runner per job; one standing
# runner on a private repository keeps its registration and cleans `_work` per job (approximate: from GitHub's runner
# documentation as remembered on 6 October 2026, the docs host answered 404 to the fetch that evening).
runner_env_file() {
cat <<EOF
# igneum-build-1 (infra/build-server/provision.sh step_runner): the environment every CI job on this runner starts with
RUSTC_WRAPPER=/usr/local/bin/sccache
SCCACHE_CONF=$RUNNER_HOME/.config/sccache/config
SCCACHE_SERVER_PORT=$RUNNER_SCCACHE_PORT
CARGO_INCREMENTAL=0
CARGO_BUILD_JOBS=$RUNNER_JOBS
CARGO_NET_GIT_FETCH_WITH_CLI=true
IGNEUM_RUST_TOOLCHAIN=$RUST_TOOLCHAIN
EOF
}
step_runner() {
local any=0 tarball cargo="$RUNNER_HOME/.cargo/bin/cargo" rustup="$RUNNER_HOME/.cargo/bin/rustup" t tmp svc unit dropin
as_runner() { su - "$RUNNER_USER" -c "$*"; }
# 1. the user: own group, no sudo, no membership of build; /home/runner 750 like the other homes
if ! id -u "$RUNNER_USER" >/dev/null 2>&1; then useradd -m -s /bin/bash "$RUNNER_USER"; any=1; fi
id -nG "$RUNNER_USER" | tr ' ' '\n' | grep -qx sudo && die "the runner user must never be in sudo"
# 2. the runner itself, from the GitHub release with the published sha256
if [ ! -x "$RUNNER_DIR/run.sh" ] || ! grep -q "\"$RUNNER_VERSION\"" "$RUNNER_DIR/.runner_version" 2>/dev/null; then
[ ! -d "$RUNNER_DIR" ] || [ ! -f "$RUNNER_DIR/.runner" ] || log "runner: a configured runner is in place; the tarball is updated underneath it (the service restarts below)"
install -d -m 755 -o "$RUNNER_USER" -g "$RUNNER_USER" "$RUNNER_DIR"
tarball="actions-runner-linux-x64-$RUNNER_VERSION.tar.gz"
tmp=$(mktemp -d)
( cd "$tmp" && curl -fsSLO "https://github.com/actions/runner/releases/download/v$RUNNER_VERSION/$tarball" \
&& printf '%s %s\n' "$RUNNER_SHA256" "$tarball" | sha256sum -c --quiet - ) || { rm -rf "$tmp"; die "runner tarball download or sha256 check failed (version $RUNNER_VERSION)"; }
tar -xzf "$tmp/$tarball" -C "$RUNNER_DIR" # as root: the temp dir is root-only; ownership handed over below
chown -R "$RUNNER_USER:$RUNNER_USER" "$RUNNER_DIR"
rm -rf "$tmp"
printf '"%s"\n' "$RUNNER_VERSION" > "$RUNNER_DIR/.runner_version"; chown "$RUNNER_USER:$RUNNER_USER" "$RUNNER_DIR/.runner_version"
any=1
fi
# 3. toolchains for the runner user: rustup pinned like the box's, both targets, clippy and rustfmt
if [ ! -x "$rustup" ]; then
as_runner "curl -fsSL https://sh.rustup.rs | sh -s -- -y --profile minimal --no-modify-path --default-toolchain $RUST_TOOLCHAIN" >/dev/null; any=1
fi
as_runner "$rustup toolchain list" | grep -q "^$RUST_TOOLCHAIN-" || { as_runner "$rustup toolchain install $RUST_TOOLCHAIN --profile minimal" >/dev/null; any=1; }
[ "$(as_runner "$rustup default" | cut -d- -f1)" = "$RUST_TOOLCHAIN" ] || { as_runner "$rustup default $RUST_TOOLCHAIN" >/dev/null; any=1; }
for t in x86_64-pc-windows-gnu x86_64-unknown-linux-gnu; do
as_runner "$rustup target list --installed --toolchain $RUST_TOOLCHAIN" | grep -qx "$t" || { as_runner "$rustup target add $t --toolchain $RUST_TOOLCHAIN" >/dev/null; any=1; }
done
as_runner "$rustup component list --installed --toolchain $RUST_TOOLCHAIN" | grep -q '^clippy' || { as_runner "$rustup component add clippy rustfmt --toolchain $RUST_TOOLCHAIN" >/dev/null; any=1; }
as_runner "git config --global --get safe.directory >/dev/null 2>&1 || git config --global --add safe.directory '*'" # the mirrors are owned by build
# 4. sccache: the build user's binary copied system-wide (the runner cannot read /home/build), a read-only view of /srv/sccache
if [ ! -x /usr/local/bin/sccache ] || ! cmp -s "$BUILD_HOME/.cargo/bin/sccache" /usr/local/bin/sccache; then
install -m 755 "$BUILD_HOME/.cargo/bin/sccache" /usr/local/bin/sccache; any=1
fi
install -d -m 755 -o "$RUNNER_USER" -g "$RUNNER_USER" "$RUNNER_HOME/.config" "$RUNNER_HOME/.config/sccache"
tmp=$(mktemp)
printf '[cache.disk]\ndir = "/srv/sccache"\nsize = %s\nrw_mode = "READ_ONLY"\n' "$(( SCCACHE_GB * 1024 * 1024 * 1024 ))" > "$tmp"
if ! cmp -s "$tmp" "$RUNNER_HOME/.config/sccache/config"; then install -m 644 -o "$RUNNER_USER" -g "$RUNNER_USER" "$tmp" "$RUNNER_HOME/.config/sccache/config"; any=1; fi
rm -f "$tmp"
# 6. registration (once; --replace re-registers under the same name after a token is given again)
if [ ! -f "$RUNNER_DIR/.runner" ]; then
if [ -z "$RUNNER_TOKEN" ]; then
log "runner: NOT registered: no RUNNER_TOKEN. From the Mac: infra/build-server/runner/register.sh (gh as igneum-labs fetches a registration token and pipes it here)"
return
fi
# the token goes to config.sh as an argument of a process owned by runner for a second; it is a one-hour registration
# token (not the runner's credential, which config.sh writes to .credentials, mode 600, owner runner), never logged here
RUNNER_TOKEN="$RUNNER_TOKEN" runuser -u "$RUNNER_USER" -- bash -c "cd '$RUNNER_DIR' && ./config.sh --unattended --replace --url '$RUNNER_REPO_URL' --token \"\$RUNNER_TOKEN\" --name '$RUNNER_NAME' --labels '$RUNNER_LABELS' --work _work" >/dev/null \
|| die "runner: config.sh failed (an expired token? register.sh fetches a fresh one)"
any=1
log "runner: registered as $RUNNER_NAME with labels self-hosted, linux, x64, $RUNNER_LABELS"
fi
# 7. the service: GitHub's unit (User=runner, KillMode=process) plus Nice and a restart on failure
svc="actions.runner.$(sed -n 's/.*"gitHubUrl": *"https:\/\/github.com\/\([^"]*\)".*/\1/p' "$RUNNER_DIR/.runner" | tr '/' '-').$RUNNER_NAME.service"
unit="/etc/systemd/system/$svc"
if [ ! -f "$unit" ]; then ( cd "$RUNNER_DIR" && ./svc.sh install "$RUNNER_USER" >/dev/null ) || die "runner: svc.sh install failed"; any=1; fi
# 8. the job environment, AFTER svc.sh install: its env.sh rewrites .env and .path from the installing shell (6 October 2026:
# the second provision run found them changed and restarted the service for nothing); the runner reads both at start
tmp=$(mktemp); runner_env_file > "$tmp"
if ! cmp -s "$tmp" "$RUNNER_DIR/.env"; then install -m 644 -o "$RUNNER_USER" -g "$RUNNER_USER" "$tmp" "$RUNNER_DIR/.env"; any=1; fi
rm -f "$tmp"
tmp=$(mktemp); printf '%s\n' "$RUNNER_HOME/.cargo/bin:/usr/local/bin:/usr/bin:/bin" > "$tmp"
if ! cmp -s "$tmp" "$RUNNER_DIR/.path"; then install -m 644 -o "$RUNNER_USER" -g "$RUNNER_USER" "$tmp" "$RUNNER_DIR/.path"; any=1; fi
rm -f "$tmp"
dropin="/etc/systemd/system/$svc.d/igneum.conf"
tmp=$(mktemp)
printf '# igneum-build-1 (infra/build-server/provision.sh step_runner)\n[Service]\nNice=10\nIOSchedulingClass=best-effort\nIOSchedulingPriority=7\nRestart=on-failure\nRestartSec=30\n' > "$tmp"
install -d -m 755 "$(dirname "$dropin")"
if ! cmp -s "$tmp" "$dropin"; then install -m 644 "$tmp" "$dropin"; systemctl daemon-reload; any=1; fi
rm -f "$tmp"
systemctl enable --quiet "$svc" 2>/dev/null || true
if [ "$any" = 1 ]; then systemctl restart "$svc"; else systemctl is-active --quiet "$svc" || systemctl start "$svc"; fi
sleep 2
systemctl is-active --quiet "$svc" || die "runner: $svc is not active: journalctl -u '$svc' -n 30"
[ "$any" = 1 ] && changed runner "$svc active as $RUNNER_USER, runner $RUNNER_VERSION, $(as_runner "$cargo --version"), sccache read-only on /srv/sccache, jobs $RUNNER_JOBS" \
|| ok runner "$svc active, runner $RUNNER_VERSION, $(as_runner "$cargo --version")"
}
step_summary() {
log "summary:"
{
@ -424,6 +544,7 @@ step_summary() {
printf 'ufw: %s\n' "$(ufw status | grep -E 'ALLOW' | awk '{ print $1 }' | tr '\n' ' ')"
printf 'caddy: %s, %s\n' "$(caddy version 2>/dev/null | cut -d' ' -f1)" "$(systemctl is-active caddy 2>/dev/null) at https://$WORKERS_HOST/headline.json"
printf 'cuda headers: %s\n' "$(ls -d /usr/local/cuda-*/include 2>/dev/null | tr '\n' ' ')$( [ -f /usr/include/CL/cl.h ] && echo '+ CL/cl.h' )"
printf 'runner: %s\n' "$( [ -f "$RUNNER_DIR/.runner" ] && printf '%s, %s, user %s' "$(systemctl list-units --type=service --no-legend 'actions.runner.*' | awk '{ print $1 ": " $4 }' | head -1)" "v$(tr -d '"' < "$RUNNER_DIR/.runner_version" 2>/dev/null)" "$RUNNER_USER" || echo "installed, NOT registered (infra/build-server/runner/register.sh)" )"
printf 'ssh line: ssh -i ~/.ssh/igneum_ed25519 build@%s\n' "$(hostname -I 2>/dev/null | awk '{ print $1 }')"
} | sed 's/^/ /'
}
@ -448,6 +569,7 @@ do_provision() {
step_caddy
step_cuda
step_ufw
step_runner
step_summary
log "done"
}

View file

@ -10,8 +10,15 @@
# BR_ARTEFACTS space-separated paths (relative to BR_DIR) the Mac will fetch; empty for test, check, clippy
#
# 1. Takes a build slot: flock on $IGNEUM_BUILD_SLOTS_DIR/build-<k> for k below the count in .../slots (the box's own slot
# files, never the Mac's); when every slot is busy it waits up to 2 h on build-0 and exits 75 if it gives up. The holder
# line is `pid N since HH:MM:SSZ waited S s: <label>`, the format tools/lock/with-lock.sh writes on the Mac.
# files, never the Mac's; 2 since main's ruling of 6 October 2026, 20:3x UK); when every slot is busy it waits up to 2 h on
# build-0 and exits 75 if it gives up. The holder line is `pid N since HH:MM:SSZ waited S s: <label>`, the format
# tools/lock/with-lock.sh writes on the Mac. The cargo job count follows the slots: after one second of settling, a build
# that holds the only taken slot gets CARGO_BUILD_JOBS=90, one that sees the other slot held gets 45 (JOBS_ALONE and
# JOBS_SHARED), so two builds share the 96 threads without thrashing; a `-j` on the cargo line (--jobs on the Mac) wins.
# A MEASUREMENT (BR_MEASURE=1: the CPU prover timing, bench rows) takes the `measure` file exclusively and excludes every
# build, as with-lock.sh's `measure` does on the Mac: builds hold `measure` shared for their whole run, so a measure waits
# for the running builds and blocks new ones until it ends. Lock files are opened in APPEND mode: the first version opened
# them with `>` and truncated a busy slot's holder line every time another build probed it (found 6 October 2026, evening).
# 2. Runs BR_CMD in BR_DIR with sccache, prints one `build-remote: RESULT rc= secs= compiles= sccache_hits_total= ...` line.
# 3. Appends one JSON line to /srv/builds/_log/builds.jsonl (the worker dashboard reads it; asked for by main on 6 October
# 2026): on success, on failure and on the slot give-up. UTC ISO 8601 Z times, numbers unquoted, unknown fields omitted,
@ -21,9 +28,16 @@
# bs_push_and_checkout: BR_CO_DIR, BR_CO_MIRROR, BR_CO_BRANCH, BR_CO_SHA, BR_CO_WT); `--self-test` (first argument) builds a
# scratch mirror and clone, dirties the clone the way a build's overlay does, moves the mirror one commit on, and shows the
# checkout mode lands on the new commit with a clean tree (the 6 October 2026 case: a stale overlay made `git checkout -B`
# refuse with "local changes would be overwritten").
# refuse with "local changes would be overwritten"); `--self-test-slots` runs fake builds and a fake measurement against a
# scratch slots directory: two concurrent builds get 45 jobs each, one alone gets 90, a measure blocks a build and a build
# blocks a measure, and a probing build leaves a busy slot's holder line intact (IGNEUM_REMOTE_RUN_UNDER_TEST=<script> runs
# the cases against another copy, which is how the old script was shown to fail them).
set -uo pipefail
# the profile sets the box's paths; an IGNEUM_BUILD_SLOTS_DIR or IGNEUM_BUILD_LOG_DIR already in the environment wins (the slot
# self-test runs against a scratch directory; the first version let the profile reset it and the test took the REAL slot)
_slots_env="${IGNEUM_BUILD_SLOTS_DIR:-}"; _log_env="${IGNEUM_BUILD_LOG_DIR:-}"
[ -f /etc/profile.d/igneum-build.sh ] && . /etc/profile.d/igneum-build.sh
[ -n "$_slots_env" ] && IGNEUM_BUILD_SLOTS_DIR="$_slots_env"; [ -n "$_log_env" ] && IGNEUM_BUILD_LOG_DIR="$_log_env"
# discard the previous overlay (tracked edits and untracked files; target dirs, the sha stamps and anything ignored are kept),
# fetch, then the branch at the commit. Runs in BR_CO_DIR, clones it from BR_CO_MIRROR when it has no .git.
@ -85,6 +99,40 @@ if [ "${1:-}" = --self-test ]; then
echo "self-test: checkout mode lands on the new commit with a clean tree, target dirs and sha stamps kept at any depth, stale index.lock removed, and fires on a stray file"; exit 0
fi
if [ "${1:-}" = --self-test-slots ]; then
me="${IGNEUM_REMOTE_RUN_UNDER_TEST:-$0}"
t=$(mktemp -d); trap 'rm -rf "$t"' EXIT
mkdir -p "$t/locks" "$t/log" "$t/dir"; echo 2 > "$t/locks/slots"
fake() { # <name> <seconds> [BR_MEASURE=1]: a fake run that records its start and end epoch and the job count it was given
local name="$1" secs="$2" measure="${3:-0}"
IGNEUM_BUILD_SLOTS_DIR="$t/locks" IGNEUM_BUILD_LOG_DIR="$t/log" BR_MEASURE="$measure" BR_DIR="$t/dir" BR_CMD="date +%s.%N > '$t/$name.start'; echo JOBS=\${CARGO_BUILD_JOBS:-none} > '$t/$name.jobs'; sleep $secs; date +%s.%N > '$t/$name.end'" \
BR_LABEL="self-test $name" BR_TOOL=self-test BR_KIND=other BR_WT=t BR_CRATE=t BR_BRANCH=t BR_SHA=0 BR_AGENT=self-test BR_COMMAND="fake $name" \
bash "$me" >"$t/$name.out" 2>&1
}
fail() { echo "self-test-slots: FAIL: $*"; exit 1; }
after() { python3 -c "import sys; sys.exit(0 if float(open(sys.argv[1]).read()) >= float(open(sys.argv[2]).read()) else 1)" "$1" "$2"; }
# 1. two concurrent builds: 45 jobs each
fake a 3 & fake b 3 & wait
[ "$(cat "$t/a.jobs")" = JOBS=45 ] && [ "$(cat "$t/b.jobs")" = JOBS=45 ] || fail "two concurrent builds got $(cat "$t/a.jobs" "$t/b.jobs" | tr '\n' ' ') (want JOBS=45 JOBS=45)"
# 2. one build alone: 90
fake c 1
[ "$(cat "$t/c.jobs")" = JOBS=90 ] || fail "a lone build got $(cat "$t/c.jobs") (want JOBS=90)"
# 3. a measure blocks a build: the build starts only after the measure ended
fake m 3 1 & sleep 0.5; fake d 1 & wait
after "$t/d.start" "$t/m.end" || fail "a build started while a measurement held the box (build start $(cat "$t/d.start"), measure end $(cat "$t/m.end"))"
[ "$(cat "$t/m.jobs")" = JOBS=none ] || fail "a measurement was given a job count"
# 4. a build blocks a measure: the measure starts only after the build ended
fake e 3 & sleep 0.5; fake n 1 1 & wait
after "$t/n.start" "$t/e.end" || fail "a measurement started while a build ran (measure start $(cat "$t/n.start"), build end $(cat "$t/e.end"))"
# 5. a probing build leaves a busy slot's holder line intact
fake f 3 & sleep 1.2; fake g 1 & sleep 0.3
grep -q 'self-test f' "$t/locks/build-0" || fail "the holder line of the busy slot build-0 was lost when another build probed it: '$(cat "$t/locks/build-0")'"
wait
# 6. the log carries the job count and the measure flag
grep -q '"jobs":45' "$t/log/builds.jsonl" && grep -q '"measure":true' "$t/log/builds.jsonl" || fail "builds.jsonl lacks jobs or measure fields"
echo "self-test-slots: two concurrent builds 45 each, a lone build 90, a measure blocks a build, a build blocks a measure, a probe keeps the holder line, the log carries jobs and measure"; exit 0
fi
if [ "${BR_MODE:-run}" = checkout ]; then
: "${BR_CO_DIR:?}" "${BR_CO_MIRROR:?}" "${BR_CO_BRANCH:?}" "${BR_CO_SHA:?}"
checkout_tree "$BR_CO_DIR" "$BR_CO_MIRROR" "$BR_CO_BRANCH" "$BR_CO_SHA" "${BR_CO_WT:-}"; exit $?
@ -93,7 +141,7 @@ fi
: "${BR_DIR:?}" "${BR_CMD:?}" "${BR_LABEL:?}" "${BR_TOOL:?}" "${BR_KIND:?}"
BR_HOST=$(hostname); BR_PID=$$; BR_T0=$(date +%s)
export BR_HOST BR_PID BR_T0
LOG_DIR=/srv/builds/_log; mkdir -p "$LOG_DIR"
LOG_DIR="${IGNEUM_BUILD_LOG_DIR:-/srv/builds/_log}"; mkdir -p "$LOG_DIR"
# jsonlog <exit> <slot> <wait_s> <start> <end> <secs> <compiles> <hits> <misses> <hits_total> <misses_total>
jsonlog() {
@ -112,7 +160,7 @@ d = {
"target": e.get('BR_TARGET'), "branch": e.get('BR_BRANCH'), "sha": e.get('BR_SHA'), "label": e['BR_LABEL'],
"agent": e.get('BR_AGENT'), "slot": num(e['BR_SLOT']), "wait_s": num(e['BR_WAIT']), "queued_at": iso(e['BR_T0']),
"start": iso(e['BR_START']), "end": iso(e['BR_END']), "secs": num(e['BR_SECS']), "exit": num(e['BR_EXIT']),
"compiles": num(e['BR_COMPILES']),
"compiles": num(e['BR_COMPILES']), "jobs": num(e.get('BR_JOBS')), "measure": (e.get('BR_MEASURE') == '1') or None,
}
sc = {k: num(e[v]) for k, v in (("hits", "BR_HITS"), ("misses", "BR_MISSES"), ("hits_total", "BR_HITS_T"), ("misses_total", "BR_MISSES_T"))}
sc = {k: v for k, v in sc.items() if v is not None}
@ -143,33 +191,68 @@ with open(e['BR_LOG'], 'a') as f:
PY
}
slots=$(cat "$IGNEUM_BUILD_SLOTS_DIR/slots" 2>/dev/null || echo 1); [ "$slots" -ge 1 ] 2>/dev/null || slots=1
got=""
for k in $(seq 0 $((slots - 1))); do
exec {fd}>"$IGNEUM_BUILD_SLOTS_DIR/build-$k"
if flock -n "$fd"; then got=$k; break; fi
exec {fd}>&-
done
if [ -z "$got" ]; then
echo "build-remote: all $slots slot(s) busy, waiting (up to 2 h) for build-0: $(head -c 160 "$IGNEUM_BUILD_SLOTS_DIR/build-0" 2>/dev/null)" >&2
# the queue is visible while it waits (the worker dashboard reads wait-* files; asked for on 6 October 2026): the same line
# format as a slot file, removed the moment the slot is taken or the wait is given up
waitfile="$IGNEUM_BUILD_SLOTS_DIR/wait-$BR_PID"
printf 'pid %s since %sZ waited 0 s: %s\n' "$BR_PID" "$(date -u +%H:%M:%S)" "$BR_LABEL" > "$waitfile"
trap 'rm -f "$waitfile"' EXIT
exec {fd}>"$IGNEUM_BUILD_SLOTS_DIR/build-0"
if ! flock -w 7200 "$fd"; then
echo "build-remote: gave up waiting for a slot after 2 h" >&2
jsonlog 75 0 $(( $(date +%s) - BR_T0 )) "" "$(date +%s)" "" "" "" "" "" ""
rm -f "$waitfile"
exit 75
SLOTS_DIR="$IGNEUM_BUILD_SLOTS_DIR"
slots=$(cat "$SLOTS_DIR/slots" 2>/dev/null || echo 1); [ "$slots" -ge 1 ] 2>/dev/null || slots=1
JOBS_ALONE="${JOBS_ALONE:-90}"; JOBS_SHARED="${JOBS_SHARED:-45}"
holder_line() { printf 'pid %s since %sZ waited %s s: %s\n' "$BR_PID" "$(date -u +%H:%M:%S)" "$1" "$BR_LABEL"; }
give_up() { # <what>
echo "build-remote: gave up waiting for $1 after 2 h" >&2
jsonlog 75 0 $(( $(date +%s) - BR_T0 )) "" "$(date +%s)" "" "" "" "" "" ""
rm -f "$waitfile"; exit 75
}
waitfile="$SLOTS_DIR/wait-$BR_PID"
# append mode: opening a lock file must never truncate the holder line another run wrote into it
exec {mfd}>>"$SLOTS_DIR/measure"
if [ "${BR_MEASURE:-0}" = 1 ]; then
# a measurement: the measure file exclusively; every running build holds it shared, so this waits for them and blocks new ones
if ! flock -n "$mfd"; then
echo "build-remote: measure waits for the running build(s) (up to 2 h): $(for f in "$SLOTS_DIR"/build-*; do head -c 120 "$f" 2>/dev/null; done | tr '\n' ' ')" >&2
holder_line 0 > "$waitfile"; trap 'rm -f "$waitfile"' EXIT
flock -w 7200 "$mfd" || give_up "the measure hold"
rm -f "$waitfile"; trap - EXIT
fi
rm -f "$waitfile"; trap - EXIT
got=0
waited=$(( $(date +%s) - BR_T0 )); got=measure
holder_line "$waited" > "$SLOTS_DIR/measure"
echo "build-remote: holding measure on $BR_HOST (waited $waited s; builds are excluded until this run ends)" >&2
else
# a build: the measure file shared (a running measurement blocks us), then one exclusive slot
if ! flock -s -n "$mfd"; then
echo "build-remote: a measurement holds the box, waiting (up to 2 h): $(head -c 160 "$SLOTS_DIR/measure" 2>/dev/null)" >&2
holder_line 0 > "$waitfile"; trap 'rm -f "$waitfile"' EXIT
flock -s -w 7200 "$mfd" || give_up "the measurement to end"
rm -f "$waitfile"; trap - EXIT
fi
got=""
for k in $(seq 0 $((slots - 1))); do
exec {fd}>>"$SLOTS_DIR/build-$k"
if flock -n "$fd"; then got=$k; break; fi
exec {fd}>&-
done
if [ -z "$got" ]; then
echo "build-remote: all $slots slot(s) busy, waiting (up to 2 h) for build-0: $(head -c 160 "$SLOTS_DIR/build-0" 2>/dev/null)" >&2
# the queue is visible while it waits (the worker dashboard reads wait-* files; asked for on 6 October 2026): the same line
# format as a slot file, removed the moment the slot is taken or the wait is given up
holder_line 0 > "$waitfile"; trap 'rm -f "$waitfile"' EXIT
exec {fd}>>"$SLOTS_DIR/build-0"
flock -w 7200 "$fd" || give_up "a slot"
rm -f "$waitfile"; trap - EXIT
got=0
fi
waited=$(( $(date +%s) - BR_T0 ))
holder_line "$waited" > "$SLOTS_DIR/build-$got"
# the job count: let a build that started in the same second take its slot, then count the slots held (this one included)
sleep 1
held=0
for k in $(seq 0 $((slots - 1))); do
if [ "$k" = "$got" ]; then held=$((held + 1)); continue; fi
exec {tfd}>>"$SLOTS_DIR/build-$k"
if flock -n "$tfd"; then flock -u "$tfd"; else held=$((held + 1)); fi
exec {tfd}>&-
done
if [ "$held" -gt 1 ]; then BR_JOBS=$JOBS_SHARED; else BR_JOBS=$JOBS_ALONE; fi
export CARGO_BUILD_JOBS="$BR_JOBS" BR_JOBS
echo "build-remote: holding build-$got on $BR_HOST (waited $waited s; $held of $slots slots held, CARGO_BUILD_JOBS=$BR_JOBS)" >&2
fi
waited=$(( $(date +%s) - BR_T0 ))
printf 'pid %s since %sZ waited %s s: %s\n' "$BR_PID" "$(date -u +%H:%M:%S)" "$waited" "$BR_LABEL" > "$IGNEUM_BUILD_SLOTS_DIR/build-$got"
echo "build-remote: holding build-$got on $BR_HOST (waited $waited s)" >&2
cd "$BR_DIR" || { jsonlog 2 "$got" "$waited" "" "$(date +%s)" "" "" "" "" "" ""; exit 2; }
sccache --start-server >/dev/null 2>&1 || true
@ -183,8 +266,8 @@ rc=$?
t2=$(date +%s); secs=$(( t2 - t1 ))
exec_after=$(stat_field "Compile requests executed"); hits_after=$(stat_field "Cache hits "); misses_after=$(stat_field "Cache misses ")
compiles=$(( ${exec_after:-0} - ${exec_before:-0} )); hits=$(( ${hits_after:-0} - ${hits_before:-0} )); misses=$(( ${misses_after:-0} - ${misses_before:-0} ))
printf 'build-remote: RESULT rc=%s secs=%s compiles=%s sccache_hits=%s sccache_misses=%s sccache_hits_total=%s sccache_misses_total=%s load=%s\n' \
"$rc" "$secs" "$compiles" "$hits" "$misses" "${hits_after:-?}" "${misses_after:-?}" "$(cut -d' ' -f1-3 /proc/loadavg)"
jsonlog "$rc" "$got" "$waited" "$t1" "$t2" "$secs" "$compiles" "$hits" "$misses" "${hits_after:-}" "${misses_after:-}"
: > "$IGNEUM_BUILD_SLOTS_DIR/build-$got"
printf 'build-remote: RESULT rc=%s secs=%s compiles=%s sccache_hits=%s sccache_misses=%s sccache_hits_total=%s sccache_misses_total=%s jobs=%s load=%s\n' \
"$rc" "$secs" "$compiles" "$hits" "$misses" "${hits_after:-?}" "${misses_after:-?}" "${BR_JOBS:-measure}" "$(cut -d' ' -f1-3 /proc/loadavg)"
jsonlog "$rc" "$([ "$got" = measure ] && echo "" || echo "$got")" "$waited" "$t1" "$t2" "$secs" "$compiles" "$hits" "$misses" "${hits_after:-}" "${misses_after:-}"
if [ "$got" = measure ]; then : > "$SLOTS_DIR/measure"; else : > "$SLOTS_DIR/build-$got"; fi
exit "$rc"

View file

@ -0,0 +1,53 @@
#!/usr/bin/env bash
# Register (or re-register) the GitHub Actions self-hosted runner on igneum-build-1 from this Mac.
# infra/build-server/runner/register.sh fetch a registration token with gh, run provision.sh on the box with it
# infra/build-server/runner/register.sh --status list the repository's runners (name, status, labels) and the box's unit
#
# The token: `gh api -X POST repos/igneum-network/igneum/actions/runners/registration-token` as igneum-labs (the CLAUDE.md gh
# rule: that account must be ACTIVE; any other active account fails here before anything is fetched). It is a one-hour
# registration token, not a credential the runner keeps (config.sh writes its own into /opt/actions-runner/.credentials,
# owner runner, mode 600). It travels to the box on ssh stdin as the first line, followed by provision.sh itself; it is
# never an argument of ssh, never written to a file on the Mac, and provision.sh never logs it. The whole of provision.sh
# runs (idempotent, every other step says ok), so the box is also brought up to date.
set -euo pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_SLUG="${IGNEUM_GH_REPO:-igneum-network/igneum}"
KEY="${IGNEUM_BUILD_KEY:-$HOME/.ssh/igneum_ed25519}"
HOST_LINE="$(head -1 "${IGNEUM_BUILD_HOST_FILE:-$HOME/.config/igneum/build-server}" | tr -d '[:space:]')"
IP="${HOST_LINE#*@}"; [ -n "$IP" ] || { echo "no build server in ~/.config/igneum/build-server (infra/build-server/run-from-mac.sh writes it)" >&2; exit 1; }
SSH=(ssh -i "$KEY" -o BatchMode=yes -o StrictHostKeyChecking=accept-new -o ConnectTimeout=15 "root@$IP")
gh_josh() {
local active
active=$(gh auth status 2>/dev/null | awk '/Logged in to github.com account/ { acct=$7 } /Active account: true/ { print acct; exit }')
if [ "$active" != igneum-labs ]; then
gh auth switch --user igneum-labs >/dev/null 2>&1 || { echo "gh: cannot switch to igneum-labs (gh auth status: ${active:-no active account})" >&2; exit 1; }
fi
[ "$(gh api user --jq .login 2>/dev/null)" = igneum-labs ] || { echo "gh: the active token is not the igneum-labs login (stored as igneum-labs); refusing" >&2; exit 1; }
}
if [ "${1:-}" = --status ]; then
gh_josh
gh api "repos/$REPO_SLUG/actions/runners" --jq '.runners[] | "\(.name)\t\(.status)\tbusy=\(.busy)\t\(([.labels[].name]) | join(","))"' || echo "(no runners or no access)"
"${SSH[@]}" 'systemctl list-units --type=service --no-legend "actions.runner.*" ; ls -la /opt/actions-runner/.runner 2>/dev/null || echo "not registered on the box"'
exit 0
fi
gh_josh
echo "fetching a registration token for $REPO_SLUG as igneum-labs (igneum-labs) ..."
TOKEN=$(gh api -X POST "repos/$REPO_SLUG/actions/runners/registration-token" --jq .token 2>/dev/null) || { echo "gh refused the registration token: the account needs admin on $REPO_SLUG (gh api repos/$REPO_SLUG --jq .permissions)" >&2; exit 1; }
[ -n "$TOKEN" ] || { echo "empty token from gh" >&2; exit 1; }
echo "token received (not shown); running provision.sh on root@$IP with it (first line of stdin, then the script)"
# the first stdin line is the token, read by the remote shell before bash -s takes the rest as the script; the output is
# kept in a temp file so the ssh exit code is read (a filter in the pipe would hide it) and any line carrying the token is
# dropped before it is shown (provision.sh never prints it; this is the belt)
OUT=$(mktemp); trap 'rm -f "$OUT"' EXIT
set +e
{ printf '%s\n' "$TOKEN"; cat "$HERE/../provision.sh"; } | "${SSH[@]}" 'IFS= read -r RUNNER_TOKEN; export RUNNER_TOKEN; MODE=provision bash -s' > "$OUT" 2>&1
RC=$?
set -e
grep -v -F "$TOKEN" "$OUT" || true
unset TOKEN
[ "$RC" = 0 ] || { echo "provision.sh exited $RC on the box" >&2; exit "$RC"; }
echo "runners now registered on $REPO_SLUG:"
gh api "repos/$REPO_SLUG/actions/runners" --jq '.runners[] | " \(.name)\t\(.status)\t\(([.labels[].name]) | join(","))"'

View file

@ -36,7 +36,7 @@ BS_TOOL=build-remote
# shellcheck source=../infra/build-server/lib.sh
. "$HERE/../infra/build-server/lib.sh"
JOBS="${JOBS:-90}"; OUT=""; ARTEFACTS=""; TARGET_DIR="target"; FETCH=1; CARGO_ARGS=()
JOBS="${JOBS:-}"; # empty = the box decides: 90 alone, 45 beside another slot holder (remote-run.sh, main's ruling 6 Oct 2026) OUT=""; ARTEFACTS=""; TARGET_DIR="target"; FETCH=1; CARGO_ARGS=()
while [ $# -gt 0 ]; do
case "$1" in
--jobs) JOBS="$2"; shift 2 ;;
@ -72,7 +72,7 @@ esac
case "${CARGO_ARGS[0]}" in build) ;; *) [ -n "${ARTEFACTS_SET:-}" ] || { FETCH=0; ARTEFACTS=""; } ;; esac # test, check, clippy: nothing to fetch
[ -n "$OUT" ] || OUT="$BS_CRATE/target-remote"
bs_log "$BS_KIND crate $BS_WT/$BS_CRATE_REL at $BS_SHA ($BS_BRANCH) -> $BS_HOST:$BS_REMOTE_CRATE; cargo ${CARGO_ARGS[*]} -j $JOBS; target dir $TARGET_DIR"
bs_log "$BS_KIND crate $BS_WT/$BS_CRATE_REL at $BS_SHA ($BS_BRANCH) -> $BS_HOST:$BS_REMOTE_CRATE; cargo ${CARGO_ARGS[*]} -j ${JOBS:-auto}; target dir $TARGET_DIR"
bs_toolchain_check
t_sync0=$(date +%s)
bs_sync_sources
@ -85,7 +85,7 @@ pre=""
if [ "$BS_KIND" = node ] && [ "${CARGO_ARGS[0]}" = build ]; then
pre="[ \"\$(cat '.build-remote-sha-$TARGET_DIR' 2>/dev/null)\" = '$BS_SHA' ] || CARGO_TARGET_DIR='$TARGET_DIR' cargo clean -q --release -p kaspa-build-info 2>/dev/null; "
fi
cmd="${pre}CARGO_TARGET_DIR='$TARGET_DIR' cargo $(printf '%q ' "${CARGO_ARGS[@]}")-j $JOBS 2>&1 | tee -a '$BS_REMOTE_WT/.build-remote.log'; rc=\${PIPESTATUS[0]}; [ \$rc = 0 ] && echo '$BS_SHA' > '.build-remote-sha-$TARGET_DIR'; ( exit \$rc )" # a subshell exit: the runner reads \$? and still prints its RESULT line
cmd="${pre}CARGO_TARGET_DIR='$TARGET_DIR' cargo $(printf '%q ' "${CARGO_ARGS[@]}")${JOBS:+-j $JOBS} 2>&1 | tee -a '$BS_REMOTE_WT/.build-remote.log'; rc=\${PIPESTATUS[0]}; [ \$rc = 0 ] && echo '$BS_SHA' > '.build-remote-sha-$TARGET_DIR'; ( exit \$rc )" # a subshell exit: the runner reads \$? and still prints its RESULT line
label="$BS_WT/$BS_CRATE_REL cargo ${CARGO_ARGS[*]}"
BR_KIND=$(bs_kind build-remote "${CARGO_ARGS[0]}"); BR_COMMAND="cargo ${CARGO_ARGS[*]}"; BR_TARGET=x86_64-unknown-linux-gnu
for ((i = 0; i < ${#CARGO_ARGS[@]}; i++)); do [ "${CARGO_ARGS[$i]}" = --target ] && BR_TARGET="${CARGO_ARGS[$((i + 1))]:-}"; done

View file

@ -37,7 +37,7 @@ BS_TOOL=cross-remote
. "$HERE/../infra/build-server/lib.sh"
TARGET=x86_64-pc-windows-gnu
JOBS="${JOBS:-90}"; OUT=""; COMPARE=""; TARGET_DIR="target"; CARGO_ARGS=()
JOBS="${JOBS:-}"; # empty = the box decides: 90 alone, 45 beside another slot holder (remote-run.sh, main's ruling 6 Oct 2026) OUT=""; COMPARE=""; TARGET_DIR="target"; CARGO_ARGS=()
while [ $# -gt 0 ]; do
case "$1" in
--jobs) JOBS="$2"; shift 2 ;;
@ -65,7 +65,7 @@ esac
[ -n "$OUT" ] || OUT="$BS_CRATE/target-remote"
[ -z "$COMPARE" ] && [ -d "$BS_CRATE/target-integration/$TARGET/release" ] && COMPARE="$BS_CRATE/target-integration/$TARGET/release"
bs_log "$BS_KIND crate $BS_WT/$BS_CRATE_REL at $BS_SHA ($BS_BRANCH) -> $BS_HOST:$BS_REMOTE_CRATE; cargo ${CARGO_ARGS[*]} -j $JOBS; target dir $TARGET_DIR"
bs_log "$BS_KIND crate $BS_WT/$BS_CRATE_REL at $BS_SHA ($BS_BRANCH) -> $BS_HOST:$BS_REMOTE_CRATE; cargo ${CARGO_ARGS[*]} -j ${JOBS:-auto}; target dir $TARGET_DIR"
bs_toolchain_check
t_sync0=$(date +%s)
bs_sync_sources
@ -82,7 +82,7 @@ echo "cross-remote: $(x86_64-w64-mingw32-gcc-posix --version | head -1); libclan
pre="" # the same kaspa-build-info clean on a new commit as build-remote.sh (the Windows target has its own fingerprint)
[ "$BS_KIND" = node ] && pre="[ \"\$(cat '.cross-remote-sha-$TARGET_DIR' 2>/dev/null)\" = '$BS_SHA' ] || CARGO_TARGET_DIR='$TARGET_DIR' cargo clean -q --release -p kaspa-build-info --target $TARGET 2>/dev/null; "
cmd="$env_block
${pre}CARGO_TARGET_DIR='$TARGET_DIR' cargo $(printf '%q ' "${CARGO_ARGS[@]}")-j $JOBS 2>&1 | tee -a '$BS_REMOTE_WT/.cross-remote.log'; rc=\${PIPESTATUS[0]}; [ \$rc = 0 ] && echo '$BS_SHA' > '.cross-remote-sha-$TARGET_DIR'
${pre}CARGO_TARGET_DIR='$TARGET_DIR' cargo $(printf '%q ' "${CARGO_ARGS[@]}")${JOBS:+-j $JOBS} 2>&1 | tee -a '$BS_REMOTE_WT/.cross-remote.log'; rc=\${PIPESTATUS[0]}; [ \$rc = 0 ] && echo '$BS_SHA' > '.cross-remote-sha-$TARGET_DIR'
for exe in $EXES; do f='$TARGET_DIR/$TARGET/release/'\$exe; [ -f \"\$f\" ] && echo \"cross-remote: DLLS \$exe: \$(x86_64-w64-mingw32-objdump -p \"\$f\" | awk '/DLL Name/ { print \$3 }' | sort -u | tr '\n' ' ')\"; done
( exit \$rc )" # a subshell exit: the runner reads \$? and still prints its RESULT line
label="$BS_WT/$BS_CRATE_REL cross $TARGET"