diff --git a/docs/plans/build-server.md b/docs/plans/build-server.md index 4a4f1f32e..ba7a841fb 100644 --- a/docs/plans/build-server.md +++ b/docs/plans/build-server.md @@ -115,6 +115,71 @@ Also logged for context: the Mac's Linux cross-build with zig (`infra/cross/buil |---|---|---| | No zig / cargo-zigbuild | the devnet seed (Debian 12, glibc 2.36) takes the Mac's zig build; a native box build links glibc 2.39, which Debian 13 seeds accept and HiveOS (Ubuntu 18/20 base) does not | install zig 0.17 + cargo-zigbuild in provision.sh, add `--target x86_64-unknown-linux-gnu.2.36` mode to build-remote.sh | | No macOS target | agents who run nodes on the Mac still build there | out of scope (needs the macOS SDK on Linux); the fleet or the box's own Devnet 2 seed takes the test-network runs instead | -| No CI runner | GitHub `ci.yml` and `windows.yml` run on GitHub's machines | install a self-hosted runner as user build once R1 is in | +| No CI runner | DONE 6 October 2026, 19:19Z (section 7): the runner `igneum-build-1` is online under user `runner`, never build; the workflow change is proposed in docs/plans/ci-self-hosted.md | main flips `IGNEUM_CI_RUNNER=box` after the shipper's cut | | Byte identity with the Mac's exes | different C/C++ toolchain (Homebrew mingw vs Ubuntu GCC 13) and embedded source paths | not a goal; the box is identical with itself build to build, cross-remote.sh reports sha256 and the DLL list per exe | | Robot API | `~/.config/igneum/hetzner-token` is the Cloud token (hcloud); the dedicated box lives in Robot, a separate credential | main sets the Robot server name in the UI; a webservice user goes to `~/.config/igneum/robot-credentials` when needed | + +## 7. The box's second shift (6 October 2026, from 20:3x UK, the project lead: "what else can our building machine be working on?") + +Four items, each its own commit on branch `box-work` with its section here. Times UTC. + +### 7.1 The GitHub Actions runner (DONE 19:19Z) + +| Fact | Value | +|---|---| +| Runner | `igneum-build-1`, actions/runner 2.338.0 (tarball sha256 af4b794c... checked against the release note), registered on igneum-network/igneum at 19:19:49Z, online, labels `self-hosted, Linux, X64, igneum-build-1` | +| User | `runner` (uid 1001, own group, no sudo, not in `build`'s group; /home/runner 750). Never `build`, never root. The runner's credential (`/opt/actions-runner/.credentials`, mode 600) is the only thing it holds; it signs nothing and reaches no hand | +| Service | `actions.runner.igneum-network-igneum.igneum-build-1.service` (GitHub's `svc.sh install runner`), drop-in `igneum.conf`: Nice 10, IO best-effort 7, Restart on-failure. Agents' builds (nice 0 through build-remote.sh) win the CPU over a CI job | +| Toolchains | rustup 1.99.0 pinned like the box (`RUST_TOOLCHAIN`), targets x86_64-unknown-linux-gnu and x86_64-pc-windows-gnu, clippy, rustfmt; mingw-w64 GCC 13 posix, Node 22, python3 + numpy from the system (numpy added to APT for the simulators) | +| sccache | `/usr/local/bin/sccache` (the build user's binary copied; /home/build is 750), config `/home/runner/.config/sccache/config` with `rw_mode = "READ_ONLY"` on /srv/sccache, own server port 4227. Shown: a job-shaped `cargo test --release` of igneum-pow as `runner` (99 tests pass, 42 s cold) made 6 compile requests, 0 hits, 6 cache WRITE ERRORS (the refusal, as wanted), and /srv/sccache stayed at 5,880,836 KB | +| Jobs | `.env`: RUSTC_WRAPPER, SCCACHE_CONF, SCCACHE_SERVER_PORT 4227, CARGO_INCREMENTAL 0, CARGO_BUILD_JOBS 48 (half the box); `.path`: the runner's cargo bin, /usr/local/bin, /usr/bin, /bin. A CI job takes NO build slot today (ci-self-hosted.md, open row) | +| Ephemeral | no. `--ephemeral` is for autoscaled fleets that register a fresh runner per job; one standing runner on a private repository keeps its registration and cleans `_work` per job (approximate: GitHub's docs host answered 404 to both fetches tonight, so this is the rule as remembered, labelled so) | +| Registration | `infra/build-server/runner/register.sh`: gh as igneum-labs (fails on any other active account, checks the login is igneum-labs), `POST repos/igneum-network/igneum/actions/runners/registration-token`, the token as the first stdin line to `provision.sh` on the box (never an argument, never a file, never logged; the output is filtered for it as a belt). `--status` lists the repository's runners and the unit | +| Idempotent | provision.sh runs 3 and 4 after the registration: `runner: ok`, no restart (`ActiveEnterTimestamp` unchanged). Run 2 had said `changed` because GitHub's `svc.sh install` runs `env.sh`, which rewrites `.env` and `.path`; the step now writes them AFTER the install | +| Workflows | NOT changed (the shipper owns them tonight). The proposed diff and the fallback (repository variable `IGNEUM_CI_RUNNER`; GitHub has no "else" in `runs-on`) are in `docs/plans/ci-self-hosted.md`. `windows.yml` cannot move to a Linux box (MSVC, WebView2, Inno Setup, PowerShell 5.1) | + +Consequences: ci.yml's `pow` and `sims` jobs would run on a pinned 1.99.0 (GitHub's `ubuntu-latest` ships whatever stable it has), with 48 jobs; GitHub-hosted minutes on a private repository are the thing saved. A CI job on the box reads only what the checkout gives it; the mirrors and `/srv/builds` belong to `build` and are not readable by `runner` (git's safe.directory is set for the runner so a future job may clone a mirror read-only if main wants it). + +### 7.0 The slots ruling (main, 20:3x UK; DONE 19:28Z, live on the box) + +| Change | Where | Shown by | +|---|---|---| +| 2 build slots (`/srv/builds/_locks/slots` = 2) | provision.sh `SLOTS` default 2, applied 19:28:18Z | `dirs: changed (... slots=2)` | +| CARGO_BUILD_JOBS 90 when a build holds the only taken slot, 45 when it sees the other slot held (one second of settling after taking the slot, then a `flock -n` probe of the other file); a `-j` on the cargo line wins | remote-run.sh; tools/build-remote.sh and tools/cross-remote.sh pass `-j` only when `--jobs` is given | `remote-run.sh --self-test-slots` on the box: two concurrent fake builds get 45 each, a lone one 90 | +| A measurement (`BR_MEASURE=1`) takes the `measure` file exclusively; builds hold it shared for their whole run, so a measure waits for the running builds and blocks new ones, as with-lock.sh's `measure` on the Mac | remote-run.sh; `infra/build-server/prover/cpu-trial.sh` is its first user | the self-test: a measure blocks a build, a build blocks a measure; JSONL carries `"jobs"` and `"measure"` | +| Lock files opened in APPEND mode | remote-run.sh | the first version's `exec {fd}>build-k` truncated a BUSY slot's holder line each time another build probed it (the dashboard read empty lines for held slots); the self-test's case 5 keeps a holder line through a probe. The OLD script under the same cases: `JOBS=none JOBS=none`, FAIL (the known-failed run, 19:25Z) | +| An environment IGNEUM_BUILD_SLOTS_DIR or IGNEUM_BUILD_LOG_DIR wins over the profile | remote-run.sh | the first self-test run let the profile reset the scratch dir and took the box's REAL slot for 7 s (19:24Z, box idle) | + +Open: a worktree whose remote-run.sh predates this keeps the old behaviour until it has master with it (the script is piped from each Mac worktree per build), so until every agent rebases, a build from an old worktree still asks `-j 90` beside a new one at 45. The build-server agent was told at 19:28Z. + +### 7.4 The CPU prover trial (DONE 19:30Z; verdict: the box is NOT a prover) + +`infra/build-server/prover/cpu-trial.sh` on the box under the measure hold (builds excluded), `igneum-prove-host` from master +7483fb37 (the 0.3.15 prover pair, sha256 71bc2438...; its `cuda` feature changes nothing under `SP1_PROVER=cpu`), fixture +`proving/fixtures/block-56-transfers.json` (the v0 block: 3 transfers, 600 pgas, one shard, 556,369 SP1 cycles), `--mode shard +--shard 0`, `RAYON_NUM_THREADS=96`, nice 19, box otherwise idle (load 2.6 at start). Log and results JSON: +`/srv/builds/_log/prover-trial/trial-20261006T192824Z.{log,json,txt}`; JSONL line kind `measure`. + +| Stage | Box, 96 threads (EPYC 9454P) | Mac, same statement class (bench-log) | +|---|---|---| +| setup (prover client, shard and aggregator keys) | 14.4 s (client 12.6, keys 1.8) | 9.3 s per invocation in the v0 loop (4 Oct, "nearly all SP1 setup") | +| execute | 0.28 s, 556,369 cycles, 927 cycles per pgas | 0.19 s for block 78 (3 Oct) | +| core proof | **34.2 s**, 7,317,561 B, verify 0.34 s | 83 s for the 200-pgas shard on a Mac at load 40 (4 Oct); 71.7 s is main's Mac figure for this fixture (its stage not recorded here, labelled approximate) | +| compressed proof (what a record carries) | **85.9 s**, 1,272,897 B, verify 0.07 s | 272 s for the 200-pgas shard on the loaded Mac (4 Oct); 61 s for the smallest shard in the 3-node v0 loop | +| whole run, wall | 136.9 s | | +| peak RSS | 28.2 GB (VmHWM) | | +| CPU use | 64 of 96 threads busy on average (6,408 percent in `ps`) | | + +What the numbers mean, per tier, and what follows: + +| Number | Means | Done or proposed | +|---|---|---| +| core 34.2 s under 60 s, compressed 85.9 s over it | the proof a record carries is the compressed one, so the shard that matters takes 120 s of proving on 96 CPU threads for the SMALLEST shard the chain has (600 pgas, 0.56 M cycles); a shard at `S_p` is 60 M cycles (bench-log 4 Oct), about 100x, so hours per shard on this CPU against the 60 s proof lag the litepaper states | the 60 s test of the ask is NOT met on the stage that counts; NO standing CPU prover unit is written, nothing joins the devnet from the box (R6 holds: the box is never a node host for the live devnet) | +| 28.2 GB peak for the smallest shard | a CPU prover needs 32 GB of RAM for a toy shard; every home tier (8, 12, 16, 24 or 32 GB CARDS, 16 to 64 GB of RAM) is out of CPU proving, and the rented 4090 boxes' CPUs are not a fallback either | the app keeps "proving on the CPU (slow)" as a correctness lane only; the prover tiers are the real cards (docs/analysis, 6 Oct rented-card measurement) | +| 64 of 96 threads busy | SP1's CPU prover does not scale to the whole box; a second trial with `RAYON_NUM_THREADS=48` would show whether half the box proves as fast (then two shards side by side) | not run tonight (one slot of the box's evening); the script takes `--threads` | +| 14.4 s setup per invocation | the same per-process cost the Mac pays; a resident prover would pay it once | already the design of the app's prover loop | + +The box stays a build and test machine. If main wants a CPU prover anyway for coverage (a prover that is always on, never fast), +the shape is a systemd unit as `build` with `SP1_PROVER=cpu`, a throwaway devnet key (never the OTA key, never a hand's key), +`--threads 48`, Nice 19 and the measure hold taken for the whole run, which would exclude builds for minutes at a time: that +is why it is not written. diff --git a/docs/plans/ci-self-hosted.md b/docs/plans/ci-self-hosted.md new file mode 100644 index 000000000..10e0d1471 --- /dev/null +++ b/docs/plans/ci-self-hosted.md @@ -0,0 +1,75 @@ +# CI on the box: the self-hosted runner and the workflow change (proposal, 6 October 2026) + +The runner `igneum-build-1` (labels `self-hosted, linux, x64, igneum-build-1`) is installed by `infra/build-server/provision.sh` +step_runner and registered by `infra/build-server/runner/register.sh` (docs/plans/build-server.md section 7). The workflows are +NOT changed here: the shipper owns `.github/workflows` tonight. This is the proposed diff for main. + +## 1. The shape: one repository variable decides, GitHub-hosted is the fallback + +GitHub has no "try this runner, else that one" in `runs-on`: a list of labels means ALL of them must match one runner, so +`[self-hosted, igneum-build-1, ubuntu-latest]` would never schedule. The fallback is therefore a repository variable read in +the expression. `IGNEUM_CI_RUNNER` = `box` sends the job to the box; unset or anything else keeps `ubuntu-latest`. Flipping it +back is one click in Settings > Secrets and variables > Actions > Variables (or `gh variable set IGNEUM_CI_RUNNER --body box` +and `gh variable delete IGNEUM_CI_RUNNER` as igneum-labs), with no commit and no queue lost: a job already queued for the box +stays queued; the next push goes to GitHub's machines. + +## 2. ci.yml (the two jobs that compile or compute; the `site` job stays on GitHub's machines) + +```diff + jobs: + pow: + name: igneum-pow tests, igneum-census build +- runs-on: ubuntu-latest ++ runs-on: ${{ vars.IGNEUM_CI_RUNNER == 'box' && fromJSON('["self-hosted", "linux", "x64", "igneum-build-1"]') || 'ubuntu-latest' }} + steps: + - uses: actions/checkout@v4 + - name: toolchain + run: rustc --version && cargo --version +@@ + sims: + name: simulators, quick modes +- runs-on: ubuntu-latest ++ runs-on: ${{ vars.IGNEUM_CI_RUNNER == 'box' && fromJSON('["self-hosted", "linux", "x64", "igneum-build-1"]') || 'ubuntu-latest' }} + steps: + - uses: actions/checkout@v4 + - uses: actions/setup-python@v5 ++ if: vars.IGNEUM_CI_RUNNER != 'box' # the box has python3 and numpy from provision.sh; setup-python would download a second Python + with: + python-version: '3.12' +- - run: python3 -m pip install --quiet numpy ++ - run: python3 -m pip install --quiet numpy ++ if: vars.IGNEUM_CI_RUNNER != 'box' +``` + +What the box gives these two jobs: rustc 1.99.0 pinned (GitHub's `ubuntu-latest` carries whatever stable it ships; the box +is the Mac's version, so CI compiles what the agents compile), sccache hits from the agents' cache (read-only), 48 cargo +jobs. The `site` job is Node and shell checks and takes under a minute on GitHub's runners; moving it buys nothing and would +put `tools/ci/public-api-check.mjs` (a live HTTPS check) behind the box's egress for no reason. + +Why the toolchain line still runs: on the box `rustc --version` must print 1.99.0; a mismatch means provision.sh and the +runner's rustup disagree (R5 in build-server.md) and the job should say so in its first step. + +## 3. windows.yml: no change possible on this box + +Every job of `windows.yml` runs on `windows-latest` for a reason the box cannot answer: the engine builds on the MSVC +target, the window host needs the Windows SDK and WebView2, the installer needs Inno Setup, the smoke run executes the exes +and the launcher under Windows PowerShell 5.1. A Linux runner has none of that. The only self-hosted option for this +workflow is a Windows runner on PC 1 or PC 2 (`actions/runner` for Windows under a service account), which conflicts with +the rule that the PCs keep only GPU and Windows-runtime JOBS through the signed job system, and is not proposed tonight. + +What the box already does for Windows is upstream of this workflow: `tools/cross-remote.sh` builds `igneumd.exe` and +`igneum-miner.exe` (the payload inputs) in 1 min 44 s, and the night battery rebuilds them for the reproducibility record. + +## 4. What to check after the flip (main, the first run on the box) + +| Check | Where | Pass | +|---|---|---| +| the job landed on the box | the run's "Set up job" log says `Runner name: 'igneum-build-1'` | yes | +| the toolchain | the `toolchain` step prints `rustc 1.99.0` | yes | +| sccache hits | add `sccache --show-stats` as a step once, or read `/srv/sccache` size before and after: the runner's config is READ_ONLY, so the size must NOT change | size unchanged | +| the agents were not starved | `/srv/builds/_log/builds.jsonl` `secs` of the builds during the run against the same crate's earlier lines | within the usual spread | +| the fallback | `gh variable delete IGNEUM_CI_RUNNER`, push a no-op commit: the job runs on `ubuntu-latest` again | yes | + +Open: a CI job on the box does not take a build slot (`/srv/builds/_locks/build-`), it runs at Nice 10 with 48 jobs; if a +CI job ever delays a release build visibly, the fix is a step at the top of the job that takes a slot through +`infra/build-server/remote-run.sh`'s flock, the same file the agents use. diff --git a/infra/build-server/prover/cpu-trial.sh b/infra/build-server/prover/cpu-trial.sh new file mode 100644 index 000000000..256e0aa99 --- /dev/null +++ b/infra/build-server/prover/cpu-trial.sh @@ -0,0 +1,66 @@ +#!/usr/bin/env bash +# CPU prover trial on igneum-build-1 (6 October 2026): time the SP1 CPU prover on one real fixture shard with every thread +# of the box, at nice 19, under one build slot (a timing taken beside another build is not a number: CLAUDE.md). +# Runs ON the box as user build: +# infra/build-server/prover/cpu-trial.sh [--fixture ] [--host ] [--threads 96] [--mode shard|execute] [--out ] +# Defaults: the host built from the igneum repo's master on the box (/srv/builds/igneum/proving/igneum-prove/target/release, +# tools/build-remote.sh from proving/igneum-prove; the `cuda` feature in it changes nothing for SP1_PROVER=cpu), the fixture +# block-56-transfers.json (the v0 block, one shard), 96 threads, --mode shard --shard 0. +# Writes /trial-.{log,json,txt}: the host's own results JSON (execute, core, compressed and verify seconds), the +# wall time, and the peak resident memory read from /proc//status VmHWM every second (no `time` package on the box). +set -euo pipefail +[ -f /etc/profile.d/igneum-build.sh ] && . /etc/profile.d/igneum-build.sh +FIXTURE=/srv/builds/igneum/proving/fixtures/block-56-transfers.json +HOST=/srv/builds/igneum/proving/igneum-prove/target/release/igneum-prove-host +THREADS=$(nproc); MODE=shard; OUT=/srv/builds/_log/prover-trial +while [ $# -gt 0 ]; do + case "$1" in + --fixture) FIXTURE="$2"; shift 2 ;; --host) HOST="$2"; shift 2 ;; --threads) THREADS="$2"; shift 2 ;; + --mode) MODE="$2"; shift 2 ;; --out) OUT="$2"; shift 2 ;; *) echo "unknown argument $1" >&2; exit 2 ;; + esac +done +[ -x "$HOST" ] || { echo "no prover host at $HOST (cd proving/igneum-prove && tools/build-remote.sh -- build --release -p igneum-prove-host)" >&2; exit 2; } +[ -f "$FIXTURE" ] || { echo "no fixture at $FIXTURE" >&2; exit 2; } +mkdir -p "$OUT" +STAMP=$(date -u +%Y%m%dT%H%M%SZ); LOG="$OUT/trial-$STAMP.log"; RES="$OUT/trial-$STAMP.json"; SUM="$OUT/trial-$STAMP.txt" +# the measurement hold (main's ruling 6 October 2026): remote-run.sh BR_MEASURE=1 takes the `measure` file exclusively, waits +# for every running build and blocks new ones until this run ends, and writes the run's JSONL line (kind measure) +RR="${IGNEUM_REMOTE_RUN:-$(dirname "$0")/../remote-run.sh}"; [ -f "$RR" ] || RR=/srv/builds/_bin/remote-run.sh +[ -f "$RR" ] || { echo "no remote-run.sh beside this script or at /srv/builds/_bin" >&2; exit 2; } +echo "cpu-trial: host $HOST; fixture $FIXTURE; threads $THREADS; mode $MODE; load $(cut -d' ' -f1-3 /proc/loadavg)" | tee "$LOG" +"$HOST" --version 2>/dev/null | tee -a "$LOG" || true +args=("$FIXTURE" --mode "$MODE" --out "$RES"); [ "$MODE" = shard ] && args+=(--shard 0) +t0=$(date +%s.%N) +PEAKF=$(mktemp) +# the command remote-run.sh evals under the measure hold: the host at nice 19 with a VmHWM poller beside it +CMD="SP1_PROVER=cpu RAYON_NUM_THREADS=$THREADS RUST_LOG=${RUST_LOG:-info} nice -n 19 '$HOST' $(printf '%q ' "${args[@]}") >> '$LOG' 2>&1 & p=\$!; peak=0; while kill -0 \$p 2>/dev/null; do h=\$(awk '/VmHWM/ { print \$2 }' /proc/\$p/status 2>/dev/null || echo 0); [ \${h:-0} -gt \$peak ] && peak=\$h; sleep 1; done; wait \$p; rc=\$?; echo \$peak > '$PEAKF'; ( exit \$rc )" +BR_MEASURE=1 BR_DIR="$(dirname "$HOST")" BR_CMD="$CMD" BR_LABEL="prover cpu-trial $(basename "$FIXTURE") $THREADS threads; agent=${IGNEUM_AGENT:-box-work}" \ + BR_TOOL=cpu-trial BR_KIND=measure BR_WT=igneum BR_CRATE=proving/igneum-prove BR_BRANCH="${BR_BRANCH:-master}" BR_SHA="${BR_SHA:-}" BR_AGENT="${IGNEUM_AGENT:-box-work}" \ + BR_COMMAND="igneum-prove-host $(basename "$FIXTURE") --mode $MODE (SP1_PROVER=cpu, $THREADS threads, nice 19)" BR_TARGET=x86_64-unknown-linux-gnu \ + bash "$RR" 2>&1 | tee -a "$LOG"; rc=${PIPESTATUS[0]} +peak=$(cat "$PEAKF" 2>/dev/null || echo 0); rm -f "$PEAKF" +t1=$(date +%s.%N) +wall=$(python3 -c "print(round($t1 - $t0, 1))") +{ + echo "cpu-trial $STAMP: exit $rc, wall ${wall} s (hold included), peak RSS $(( peak / 1024 )) MB (VmHWM), threads $THREADS, nice 19, measure hold, load at end $(cut -d' ' -f1-3 /proc/loadavg)" + echo "host: $HOST ($(stat -c %s "$HOST") B, sha256 $(sha256sum "$HOST" | cut -c1-16)...)" + echo "fixture: $FIXTURE" + if [ -f "$RES" ]; then + python3 - "$RES" <<'PY' +import json, sys +d = json.load(open(sys.argv[1])) +def walk(o, p=""): + if isinstance(o, dict): + for k, v in o.items(): walk(v, f"{p}.{k}" if p else k) + elif isinstance(o, list): + for i, v in enumerate(o): walk(v, f"{p}[{i}]") + else: + s = str(o) + if any(t in p.lower() for t in ("sec", "time", "cycles", "bytes", "verified", "prover", "shard", "program", "stage")) and len(s) < 80: + print(f" {p} = {s}") +walk(d) +PY + fi + echo "log lines with times:"; grep -E -i 'prove|core|compress|verif|execute|cycles| s\b|seconds' "$LOG" | grep -v '^cpu-trial' | tail -20 | sed 's/^/ /' +} | tee "$SUM" +exit "$rc" diff --git a/infra/build-server/provision.sh b/infra/build-server/provision.sh index ee789e9fc..99bfdfc85 100755 --- a/infra/build-server/provision.sh +++ b/infra/build-server/provision.sh @@ -26,7 +26,8 @@ # NODE_MAJOR 22 # WORKTREES "" space-separated agent worktree names, one /srv/builds/ each (run-from-mac.sh fills it from # `git worktree list` on the Mac; tools/build-remote.sh creates a missing one on first use) -# SLOTS 1 remote build slots (tools/build-remote.sh takes one; with 1 slot every build gets the box) +# SLOTS 2 remote build slots (tools/build-remote.sh takes one; remote-run.sh sets CARGO_BUILD_JOBS 90 when it holds +# the only taken slot and 45 when both are held; a measurement takes the `measure` file and excludes builds) # P2P_PORTS "26611 26811" TCP ports ufw opens beside 22: the devnet seed's p2p (infra/seed-nodes/config.sh devnet # P2P_PORT=26611) and the testnet seed's (26811). A suffixed devnet (Devnet 2, the fleet's staging # chain) listens on the same 26611 (infra/cloud-devnet/config.sh P2P_PORT=26611). RPC ports @@ -52,13 +53,25 @@ SCCACHE_GB="${SCCACHE_GB:-100}" SCCACHE_VERSION="${SCCACHE_VERSION:-}" NODE_MAJOR="${NODE_MAJOR:-22}" WORKTREES="${WORKTREES:-}" -SLOTS="${SLOTS:-1}" +SLOTS="${SLOTS:-2}" # 2 since main's ruling of 6 October 2026 (20:3x UK); remote-run.sh gives 90 jobs alone, 45 beside another P2P_PORTS="${P2P_PORTS:-26611 26811}" BOX_HOSTNAME="${BOX_HOSTNAME:-igneum-build-1}" WORKERS_HOST="${WORKERS_HOST:-build.igneum.network}" # the dashboard feed's HTTPS name (A record in deSEC, 6 Oct 2026) SSH_PUBKEY="${SSH_PUBKEY:-}" BUILD_USER=build BUILD_HOME=/home/$BUILD_USER +# the GitHub Actions runner (step_runner, 6 October 2026 evening): a dedicated user, never build and never root +RUNNER_USER=runner +RUNNER_HOME=/home/$RUNNER_USER +RUNNER_DIR=/opt/actions-runner +RUNNER_VERSION="${RUNNER_VERSION:-2.338.0}" # github.com/actions/runner releases, read 6 October 2026 +RUNNER_SHA256="${RUNNER_SHA256:-af4b794c1bc41d73d40535e3fe092a39f9679cd8d965954c2aca25a05ca41d32}" # the release note's linux-x64 line +RUNNER_REPO_URL="${RUNNER_REPO_URL:-https://github.com/igneum-network/igneum}" +RUNNER_NAME="${RUNNER_NAME:-$BOX_HOSTNAME}" +RUNNER_LABELS="${RUNNER_LABELS:-igneum-build-1}" # added to the defaults self-hosted, linux, x64 +RUNNER_JOBS="${RUNNER_JOBS:-48}" # cargo jobs for a CI job: half the box, the agents' builds keep the rest +RUNNER_TOKEN="${RUNNER_TOKEN:-}" # a registration token (1 h), from infra/build-server/runner/register.sh over stdin; never logged +RUNNER_SCCACHE_PORT="${RUNNER_SCCACHE_PORT:-4227}" # the runner's own sccache server; 4226 is the build user's log() { printf '%s provision: %s\n' "$(date -u +%H:%M:%S)" "$*"; } die() { log "ERROR: $*" >&2; exit 1; } @@ -127,6 +140,9 @@ APT_PACKAGES=( build-essential clang lld llvm libclang-dev pkg-config libssl-dev cmake protobuf-compiler libprotobuf-dev gcc-mingw-w64-x86-64 g++-mingw-w64-x86-64 binutils-mingw-w64-x86-64 mingw-w64-x86-64-dev mingw-w64-tools git tmux curl ca-certificates xz-utils zstd unzip rsync jq python3 ufw htop file caddy + # the GitHub Actions runner's .NET runtime needs libicu (bin/installdependencies.sh would install it); the two Python + # simulators (sim/finality_v2.py, sim/difficulty/sim.py) need numpy, which ci.yml pip-installs on GitHub's runners + libicu74 python3-numpy ) step_apt() { local need=() p @@ -408,6 +424,110 @@ step_ufw() { [ "$any" = 1 ] && changed ufw "22, 80, 443 and ${P2P_PORTS} open, everything else denied" || ok ufw "22, 80, 443 and ${P2P_PORTS}" } +# The GitHub Actions self-hosted runner for igneum-network/igneum (asked for 6 October 2026, "what else can the box work on"). +# A dedicated user `runner` (no sudo, not in the build group), the runner at /opt/actions-runner (sha256 of the tarball checked +# against the release note), its own rustup pinned to RUST_TOOLCHAIN with the two targets and clippy, Node 22 and mingw from +# the system, and sccache against the box's cache in READ-ONLY mode (`rw_mode = "READ_ONLY"`, accepted by sccache 0.18 on +# 6 October 2026): a CI job may take hits from what the agents built, never write a line into their cache, and talks to its +# own sccache server on RUNNER_SCCACHE_PORT so the build user's server never compiles as the wrong user. Registration needs +# RUNNER_TOKEN (infra/build-server/runner/register.sh fetches one with gh as igneum-labs and pipes it over stdin); without +# it the step installs everything and says what is missing. The service is GitHub's own `svc.sh install runner` unit plus a +# drop-in with Nice=10 (agents' builds through build-remote.sh win the CPU; a CI job is a check, not a release build). +# Not ephemeral: GitHub recommends `--ephemeral` for autoscaled fleets that register a fresh runner per job; one standing +# runner on a private repository keeps its registration and cleans `_work` per job (approximate: from GitHub's runner +# documentation as remembered on 6 October 2026, the docs host answered 404 to the fetch that evening). +runner_env_file() { + cat </dev/null 2>&1; then useradd -m -s /bin/bash "$RUNNER_USER"; any=1; fi + id -nG "$RUNNER_USER" | tr ' ' '\n' | grep -qx sudo && die "the runner user must never be in sudo" + # 2. the runner itself, from the GitHub release with the published sha256 + if [ ! -x "$RUNNER_DIR/run.sh" ] || ! grep -q "\"$RUNNER_VERSION\"" "$RUNNER_DIR/.runner_version" 2>/dev/null; then + [ ! -d "$RUNNER_DIR" ] || [ ! -f "$RUNNER_DIR/.runner" ] || log "runner: a configured runner is in place; the tarball is updated underneath it (the service restarts below)" + install -d -m 755 -o "$RUNNER_USER" -g "$RUNNER_USER" "$RUNNER_DIR" + tarball="actions-runner-linux-x64-$RUNNER_VERSION.tar.gz" + tmp=$(mktemp -d) + ( cd "$tmp" && curl -fsSLO "https://github.com/actions/runner/releases/download/v$RUNNER_VERSION/$tarball" \ + && printf '%s %s\n' "$RUNNER_SHA256" "$tarball" | sha256sum -c --quiet - ) || { rm -rf "$tmp"; die "runner tarball download or sha256 check failed (version $RUNNER_VERSION)"; } + tar -xzf "$tmp/$tarball" -C "$RUNNER_DIR" # as root: the temp dir is root-only; ownership handed over below + chown -R "$RUNNER_USER:$RUNNER_USER" "$RUNNER_DIR" + rm -rf "$tmp" + printf '"%s"\n' "$RUNNER_VERSION" > "$RUNNER_DIR/.runner_version"; chown "$RUNNER_USER:$RUNNER_USER" "$RUNNER_DIR/.runner_version" + any=1 + fi + # 3. toolchains for the runner user: rustup pinned like the box's, both targets, clippy and rustfmt + if [ ! -x "$rustup" ]; then + as_runner "curl -fsSL https://sh.rustup.rs | sh -s -- -y --profile minimal --no-modify-path --default-toolchain $RUST_TOOLCHAIN" >/dev/null; any=1 + fi + as_runner "$rustup toolchain list" | grep -q "^$RUST_TOOLCHAIN-" || { as_runner "$rustup toolchain install $RUST_TOOLCHAIN --profile minimal" >/dev/null; any=1; } + [ "$(as_runner "$rustup default" | cut -d- -f1)" = "$RUST_TOOLCHAIN" ] || { as_runner "$rustup default $RUST_TOOLCHAIN" >/dev/null; any=1; } + for t in x86_64-pc-windows-gnu x86_64-unknown-linux-gnu; do + as_runner "$rustup target list --installed --toolchain $RUST_TOOLCHAIN" | grep -qx "$t" || { as_runner "$rustup target add $t --toolchain $RUST_TOOLCHAIN" >/dev/null; any=1; } + done + as_runner "$rustup component list --installed --toolchain $RUST_TOOLCHAIN" | grep -q '^clippy' || { as_runner "$rustup component add clippy rustfmt --toolchain $RUST_TOOLCHAIN" >/dev/null; any=1; } + as_runner "git config --global --get safe.directory >/dev/null 2>&1 || git config --global --add safe.directory '*'" # the mirrors are owned by build + # 4. sccache: the build user's binary copied system-wide (the runner cannot read /home/build), a read-only view of /srv/sccache + if [ ! -x /usr/local/bin/sccache ] || ! cmp -s "$BUILD_HOME/.cargo/bin/sccache" /usr/local/bin/sccache; then + install -m 755 "$BUILD_HOME/.cargo/bin/sccache" /usr/local/bin/sccache; any=1 + fi + install -d -m 755 -o "$RUNNER_USER" -g "$RUNNER_USER" "$RUNNER_HOME/.config" "$RUNNER_HOME/.config/sccache" + tmp=$(mktemp) + printf '[cache.disk]\ndir = "/srv/sccache"\nsize = %s\nrw_mode = "READ_ONLY"\n' "$(( SCCACHE_GB * 1024 * 1024 * 1024 ))" > "$tmp" + if ! cmp -s "$tmp" "$RUNNER_HOME/.config/sccache/config"; then install -m 644 -o "$RUNNER_USER" -g "$RUNNER_USER" "$tmp" "$RUNNER_HOME/.config/sccache/config"; any=1; fi + rm -f "$tmp" + # 6. registration (once; --replace re-registers under the same name after a token is given again) + if [ ! -f "$RUNNER_DIR/.runner" ]; then + if [ -z "$RUNNER_TOKEN" ]; then + log "runner: NOT registered: no RUNNER_TOKEN. From the Mac: infra/build-server/runner/register.sh (gh as igneum-labs fetches a registration token and pipes it here)" + return + fi + # the token goes to config.sh as an argument of a process owned by runner for a second; it is a one-hour registration + # token (not the runner's credential, which config.sh writes to .credentials, mode 600, owner runner), never logged here + RUNNER_TOKEN="$RUNNER_TOKEN" runuser -u "$RUNNER_USER" -- bash -c "cd '$RUNNER_DIR' && ./config.sh --unattended --replace --url '$RUNNER_REPO_URL' --token \"\$RUNNER_TOKEN\" --name '$RUNNER_NAME' --labels '$RUNNER_LABELS' --work _work" >/dev/null \ + || die "runner: config.sh failed (an expired token? register.sh fetches a fresh one)" + any=1 + log "runner: registered as $RUNNER_NAME with labels self-hosted, linux, x64, $RUNNER_LABELS" + fi + # 7. the service: GitHub's unit (User=runner, KillMode=process) plus Nice and a restart on failure + svc="actions.runner.$(sed -n 's/.*"gitHubUrl": *"https:\/\/github.com\/\([^"]*\)".*/\1/p' "$RUNNER_DIR/.runner" | tr '/' '-').$RUNNER_NAME.service" + unit="/etc/systemd/system/$svc" + if [ ! -f "$unit" ]; then ( cd "$RUNNER_DIR" && ./svc.sh install "$RUNNER_USER" >/dev/null ) || die "runner: svc.sh install failed"; any=1; fi + # 8. the job environment, AFTER svc.sh install: its env.sh rewrites .env and .path from the installing shell (6 October 2026: + # the second provision run found them changed and restarted the service for nothing); the runner reads both at start + tmp=$(mktemp); runner_env_file > "$tmp" + if ! cmp -s "$tmp" "$RUNNER_DIR/.env"; then install -m 644 -o "$RUNNER_USER" -g "$RUNNER_USER" "$tmp" "$RUNNER_DIR/.env"; any=1; fi + rm -f "$tmp" + tmp=$(mktemp); printf '%s\n' "$RUNNER_HOME/.cargo/bin:/usr/local/bin:/usr/bin:/bin" > "$tmp" + if ! cmp -s "$tmp" "$RUNNER_DIR/.path"; then install -m 644 -o "$RUNNER_USER" -g "$RUNNER_USER" "$tmp" "$RUNNER_DIR/.path"; any=1; fi + rm -f "$tmp" + dropin="/etc/systemd/system/$svc.d/igneum.conf" + tmp=$(mktemp) + printf '# igneum-build-1 (infra/build-server/provision.sh step_runner)\n[Service]\nNice=10\nIOSchedulingClass=best-effort\nIOSchedulingPriority=7\nRestart=on-failure\nRestartSec=30\n' > "$tmp" + install -d -m 755 "$(dirname "$dropin")" + if ! cmp -s "$tmp" "$dropin"; then install -m 644 "$tmp" "$dropin"; systemctl daemon-reload; any=1; fi + rm -f "$tmp" + systemctl enable --quiet "$svc" 2>/dev/null || true + if [ "$any" = 1 ]; then systemctl restart "$svc"; else systemctl is-active --quiet "$svc" || systemctl start "$svc"; fi + sleep 2 + systemctl is-active --quiet "$svc" || die "runner: $svc is not active: journalctl -u '$svc' -n 30" + [ "$any" = 1 ] && changed runner "$svc active as $RUNNER_USER, runner $RUNNER_VERSION, $(as_runner "$cargo --version"), sccache read-only on /srv/sccache, jobs $RUNNER_JOBS" \ + || ok runner "$svc active, runner $RUNNER_VERSION, $(as_runner "$cargo --version")" +} + step_summary() { log "summary:" { @@ -424,6 +544,7 @@ step_summary() { printf 'ufw: %s\n' "$(ufw status | grep -E 'ALLOW' | awk '{ print $1 }' | tr '\n' ' ')" printf 'caddy: %s, %s\n' "$(caddy version 2>/dev/null | cut -d' ' -f1)" "$(systemctl is-active caddy 2>/dev/null) at https://$WORKERS_HOST/headline.json" printf 'cuda headers: %s\n' "$(ls -d /usr/local/cuda-*/include 2>/dev/null | tr '\n' ' ')$( [ -f /usr/include/CL/cl.h ] && echo '+ CL/cl.h' )" + printf 'runner: %s\n' "$( [ -f "$RUNNER_DIR/.runner" ] && printf '%s, %s, user %s' "$(systemctl list-units --type=service --no-legend 'actions.runner.*' | awk '{ print $1 ": " $4 }' | head -1)" "v$(tr -d '"' < "$RUNNER_DIR/.runner_version" 2>/dev/null)" "$RUNNER_USER" || echo "installed, NOT registered (infra/build-server/runner/register.sh)" )" printf 'ssh line: ssh -i ~/.ssh/igneum_ed25519 build@%s\n' "$(hostname -I 2>/dev/null | awk '{ print $1 }')" } | sed 's/^/ /' } @@ -448,6 +569,7 @@ do_provision() { step_caddy step_cuda step_ufw + step_runner step_summary log "done" } diff --git a/infra/build-server/remote-run.sh b/infra/build-server/remote-run.sh index bb626a2f6..f04b87baa 100755 --- a/infra/build-server/remote-run.sh +++ b/infra/build-server/remote-run.sh @@ -10,8 +10,15 @@ # BR_ARTEFACTS space-separated paths (relative to BR_DIR) the Mac will fetch; empty for test, check, clippy # # 1. Takes a build slot: flock on $IGNEUM_BUILD_SLOTS_DIR/build- for k below the count in .../slots (the box's own slot -# files, never the Mac's); when every slot is busy it waits up to 2 h on build-0 and exits 75 if it gives up. The holder -# line is `pid N since HH:MM:SSZ waited S s: