- provision.sh step_runner: actions/runner 2.338.0 (sha256 checked) at /opt/actions-runner under a dedicated user `runner`
(no sudo, not in build's group), rustup 1.99.0 pinned with both targets, sccache against /srv/sccache in READ_ONLY mode
on its own server port, Node 22 and mingw from the system, GitHub's svc.sh unit with a Nice 10 drop-in; registered on
igneum-network/igneum as igneum-build-1 (labels self-hosted, linux, x64, igneum-build-1) through
infra/build-server/runner/register.sh (gh as igneum-labs, the token on ssh stdin, never logged). Idempotent after
the env files moved behind svc.sh install (its env.sh rewrites them). libicu74 and python3-numpy added to APT.
- main's slots ruling: SLOTS default 2; remote-run.sh sets CARGO_BUILD_JOBS 90 when it holds the only taken slot and 45
when both are held, BR_MEASURE=1 takes the `measure` file exclusively and excludes builds (builds hold it shared),
lock files open in append mode (the old `exec {fd}>` truncated a busy slot's holder line on every probe), env
IGNEUM_BUILD_SLOTS_DIR and IGNEUM_BUILD_LOG_DIR win over the profile, `--self-test-slots` with five cases (the old
script fails it with JOBS=none); build-remote.sh and cross-remote.sh pass -j only when --jobs is given.
- infra/build-server/prover/cpu-trial.sh: the SP1 CPU prover on one fixture shard under the measure hold with a VmHWM
poller; 6 Oct 2026 run: core 34.2 s, compressed 85.9 s, peak RSS 28.2 GB on 96 threads, so no standing CPU prover.
- docs/plans/ci-self-hosted.md: the proposed runs-on change for ci.yml behind the repository variable IGNEUM_CI_RUNNER
(GitHub-hosted is the fallback), and why windows.yml cannot move to a Linux box. Workflows untouched.
- docs/plans/build-server.md section 7.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
30 KiB
Build server: igneum-build-1 (plan, benchmarks, rules)
6 October 2026. the project lead ordered a Hetzner dedicated server at 18:2x UK: AX162-1-LTD, EPYC 9454P (48 cores, 96 threads), 128 GB, 2x 3.84 TB NVMe, Falkenstein (FSN1-DC24, Hetzner #3088308), 188.40.146.49, IPv6 2a01:4f8:2240:205a::/64. Purpose: this Mac is the compile queue for ten agents (8-minute rebuilds under contention) and the Windows cross-builds; the box takes over Rust builds, cross-builds and a shared compile cache, and later a CI runner and the Devnet 2 seed.
1. What is in place (all on branch build-server)
| Piece | File | State 6 Oct 2026 |
|---|---|---|
| Install + provision | infra/build-server/provision.sh |
ran: installimage 17:16 to 17:20 UTC (Ubuntu 24.04, RAID 1, no swap), provision 17:24 to 17:27 UTC, second run 2 s with zero changes (idempotent) |
| Mac side | infra/build-server/run-from-mac.sh <ip> |
ran: ~/.config/igneum/build-server = build@188.40.146.49, build remotes added, 124 igneum branches and 48 fork branches pushed to the bare mirrors |
| Shared library | infra/build-server/lib.sh |
host line, ssh options, mirror push, remote checkout, overlay, remote slot runner |
| Remote runner | infra/build-server/remote-run.sh |
runs on the box: slot, sccache, RESULT line, one JSON line per run in /srv/builds/_log/builds.jsonl (the worker dashboard's feed, fields as main asked on 6 Oct) |
| Remote cargo | tools/build-remote.sh |
any crate dir, any worktree: tools/build-remote.sh [-- cargo args]; IGNEUM_AGENT=<name> tags the slot label and the log |
| Remote Windows cross | tools/cross-remote.sh |
fork worktree or app/igneum-app: tools/cross-remote.sh [--compare <dir of the Mac's exes>] |
ssh line: ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 (root works with the same key; passwords are off).
Box facts (provision summary): rustc 1.99.0 (the Mac's version; NO rust-toolchain file exists in the repo or the fork, the
pin is RUST_TOOLCHAIN in provision.sh), targets x86_64-unknown-linux-gnu and x86_64-pc-windows-gnu, sccache 0.18.0 with a
100 GB disk cache at /srv/sccache, mingw-w64 GCC 13 posix threads (the PC job's recipe), clang and lld 18.1.3, Node v22.23.3,
git 2.43.0, tmux 3.4, ufw 22 + 26611 + 26811 (the devnet and testnet seed p2p ports; RPC stays on loopback as on every seed),
md0 /boot and md1 / as RAID 1, 3.3 TB free, swap none, 1 build slot in /srv/builds/_locks/slots.
2. How a build travels
bs_context(lib.sh) reads the crate: a fork worktree undervendor/(kind node, mirror/srv/igneum-node.git) or a crate of the igneum repo (kind repo, mirror/srv/igneum.git).cargo metadatalists every path dependency.- HEAD is pushed to the mirror; the box clones or fetches it at
/srv/builds/<worktree>/<same relative path>and checks out that commit. A real.gitis needed: the fork'skaspa-build-inforunsgit rev-parse HEADat build time and every release plan checks the commit in the binary's strings. - Uncommitted changes go by
rsync --checksumwithout-t(files that changed are written with the box's clock) and the written files are re-stamped withtouch(the copied-sources rule,tools/ci/copied-sources-check.shpasses). - The cargo command runs in the crate dir under one of the box's slot files (
flockon/srv/builds/_locks/build-<k>, never the Mac's~/.config/igneum/build-slots), with sccache and-j 90, logging wall time, compile count and cache hits. - Artefacts come back into
<crate>/target-remote/...with size and sha256. NEVER intotarget/: the box's native binaries are x86_64 Linux (glibc 2.39) and do not run on the Mac.
/srv/builds/<worktree> mirrors the Mac's igneum worktree ROOT because the fork's path dependency is
../../../../igneum-pow (vendor/igneum-node/consensus/pow/Cargo.toml).
3. Benchmarks
The Mac numbers are from the logs (docs/bench-log.md, docs/plans/release-0.3.10.md, release-0.3.11.md); no fresh Mac run
was made, because a Mac build would take a slot from the agents and the logs already hold several readings per case.
Mac = Apple M5 Max (18 cores), builds at nice -n 19 with 4 to 6 jobs under the build lock, usually loaded by other agents.
Box = igneum-build-1, -j 90, nothing else running.
| Case | Mac (logged) | Box | Box run |
|---|---|---|---|
Clean node build (cargo build --release -p kaspad -p igneum-miner --features kaspad/igneum-pow, every crate) |
12 min 36 s (0.3.10, new target dir, -j 4); 17 min 53 s at load 140 and 8 min 58 s second time (fud ledger, -j 4); rusty-kaspa kaspad alone 2 min 36 s on an idle Mac | 1 min 27 s cargo wall (1 min 34 s end to end from the Mac: push and sync 1 s, build, fetch of both binaries) | cold: 518 crates downloaded inside that time, sccache 0 hits / 993 misses, 563 crates compiled; igneumd 48,220,896 B, igneum-miner 9,107,232 B, ELF x86-64 PIE; igneumd --version runs on the box; 17:30:58 to 17:32:32 UTC |
| Incremental rebuild (one file changed, same target dir) | 2 min 08 s (0.3.10, 17:49Z); 3 min 19 s (0.3.11); 5 min 06 s (txgossip); 15 min 18 s at load 110 to 134 (M31); 35 s to 4 min (execution layer); "8 min under contention" (main, 6 Oct) | 7 s cargo wall (10 s end to end: sync 1 s, build 6.97 s, fetch) | one line appended to kaspad/src/main.rs in the fork worktree, uncommitted, carried by the overlay (1 file written and re-stamped); 1 crate compiled, igneumd relinked; box idle (load 12 from the clean build a minute earlier); 17:33:11 to 17:33:21 UTC |
Windows cross-build (--target x86_64-pc-windows-gnu, igneumd.exe + igneum-miner.exe) |
8 min 25 s clean (-j 6, 3 Oct); 4 min 49 s and 4 min 53 s with warm dependencies (4 Oct); 12 min 28 s (finality-fixes, 4 Oct) | 1 min 44 s cargo wall (1 min 48 s end to end) | cold: 995 compiles, sccache 0 hits; igneumd.exe 49,434,624 B, igneum-miner.exe 10,201,600 B; mingw GCC 13 posix, libclang 18; igneumd.exe imports libstdc++-6.dll (this fork head predates the housekeeping commit that made the C++ runtime static), so cross-remote.sh now fetches the three GCC 13 runtime DLLs beside the exes as the PC job does; 17:33:58 to 17:35:46 UTC. Second run on the same target dir, no source change: 8 s |
Also logged for context: the Mac's Linux cross-build with zig (infra/cross/build-linux.sh) took 30 min 56 s cold and
3 min 20 s incremental; PC 1 built Linux + Windows node and app and ran both test suites in 7 min 38 s cold, 5 min 09 s warm
(CLAUDE.md, 5 Oct).
What the numbers mean and what follows
| Number | Means | Done or proposed |
|---|---|---|
| Clean build 1 min 27 s against 12 to 18 min on the loaded Mac (8 to 12x) | a new worktree costs an agent a minute and a half, not a slot for a quarter of an hour; the first build of every one of the 123 worktree dirs on the box is this cold case, later ones are the 7 s case | R1 proposed; sccache was cold (0 hits of 1,982 requests) because every run so far was the first of its kind; the cache fills as agents build the same crate versions from different worktrees, so the second worktree's clean build will be mostly hits (measure when it happens, write the number here) |
| Incremental 7 s against 2 to 15 min on the Mac (the "8 min under contention") | an edit-build loop of seconds for Linux targets; end to end 10 s because sync is 1 s and the two binaries (57 MB) come back in 2 s | R1; the slot count stays 1 until two agents collide, then 2 with -j 48 each (SLOTS=2 run-from-mac.sh and --jobs 48) |
| Windows cross 1 min 44 s against 4 min 49 s to 12 min 28 s on the Mac (3 to 7x), 8 s incremental | a Windows exe per commit is cheap enough to build on every push; the PC build job (7 min 38 s cold, 5 min 09 s warm for Linux + Windows + tests) stays the second source |
R2 proposed |
| Mac arm64 binaries: not built here | agents who run nodes on the Mac (local devnets, the DMG) still take Mac slots; the box cannot remove that contention | R3; the real relief is to move test networks to the fleet or to a Devnet 2 seed on this box (its unit, ports and ufw are ready) |
| The commit hash is EMPTY in every Mac worktree build and was empty in the first box builds | kaspa-build-info (build-info/build.rs) needs .git to be a directory AND HEAD to be a symbolic ref to a loose branch file; a worktree's .git is a file and a detached HEAD is not a ref; and once it has found nothing it emits no rerun-if-changed, so cargo never runs it again in that target dir (release-0.3.11: cargo clean -p kaspa-build-info) |
fixed on the box: the remote checkout is git checkout -B <branch> <sha> and build-remote.sh runs cargo clean --release -p kaspa-build-info whenever the commit differs from the last one built in that target dir (.build-remote-sha-<target dir>); verified 17:40 UTC: 2 string hits for 3bfe346f, binary +1,024 B. The release plans' "commit in its strings" checks were passing against builds from the fork's MAIN checkout (a real .git directory on a branch), not from worktrees. DONE (main's decision): the PC job writes a minimal node/.git (HEAD -> refs/heads/build holding the manifest's new commit_full, written by push-build-inputs.sh) at extract and cleans kaspa-build-info on a new commit (jobbuild.rs, 4 unit tests pass on the box); the Mac's cross-build.sh refuses a worktree and cleans on a new commit; the gate tools/ci/commit-string-check.sh runs on every igneumd the three scripts produce (self-test in ci.yml; shown firing on the Mac's worktree-built igneumd and passing on the box's). Only kaspad depends on kaspa-build-info, so igneum-miner is out of the gate's scope |
| igneumd.exe hash changed build to build with no source change (c424aae0 then bfbff015) while the Linux igneumd stayed byte-identical across three builds | the mingw linker wrote a timestamp into the PE header; the Linux ELF has none | DONE (main's decision): -C link-arg=-Wl,--no-insert-timestamp in cross-remote.sh, the Mac's cross-build.sh and the PC job (jobbuild.rs); verified 17:48 and 17:49 UTC: two builds, igneumd.exe c38b7570... and igneum-miner.exe c3a0fc2b... identical both times |
| Second worktree's clean build (vendor/igneum-node-v4 from the main checkout, cold target dir, warm sccache): 1 min 18 s, 604 hits of 993 compiles (61 percent), 389 misses | the first build of each of the 123 worktree dirs costs 1 min 18 s to 1 min 27 s, not the Mac's 12 to 18 min; sccache saves 9 s of the 87 because the clean build is dominated by the fork's own 74 crates and the C++ (rocksdb) objects, which differ per tree or do not cache; the cache is 1.0 GB after four builds, capped at 100 GB | nothing; the hit rate is in every RESULT line and every JSONL line |
| Disk 3.3 TB free, RAM 125 GB, load peaked at 24 during the Windows build with 96 threads | room for 2 slots and the Devnet 2 seed without contention | nothing now |
4. Rules (ADOPTED by main on 6 October 2026, written into CLAUDE.md "Running agents on this Mac"; R2 narrowed: the PCs keep only GPU and Windows-runtime jobs from now, not two releases)
| Rule | Text |
|---|---|
| R1 | Every agent's cargo build, cargo test, cargo check and cargo clippy for Linux goes through tools/build-remote.sh from the crate directory. The box takes one remote slot per build; nobody runs cargo over ssh by hand. |
| R2 | Windows exes come from tools/cross-remote.sh (fork worktree: igneumd.exe, igneum-miner.exe; app/igneum-app: igneum-app.exe and the two tools). The PC build job stays as the second source until two releases have shipped from the box. |
| R3 | The Mac keeps what only it can do: aarch64-apple-darwin binaries (the DMG, nodes that agents run locally), tests that need Metal (proto-metal, the Metal worker), and measurements. Those still use tools/lock/with-lock.sh and the Mac's build slots. |
| R4 | A worktree builds once on the box per commit plus overlay; the next build is incremental in /srv/builds/<worktree>/.../target. Nobody deletes another worktree's target dir on the box. |
| R5 | RUST_TOOLCHAIN in provision.sh is bumped in the same commit as the Mac's rustup update; build-remote.sh refuses a mismatch. Add a rust-toolchain.toml to the fork and the repo (none exists today) so both sides pin from one file. |
| R6 | The box is never a node host for the live devnet and never holds a secret (no ~/.config/igneum there). The Devnet 2 seed on it runs under its own unit with --devnet --devnet-suffix=<n> on 26611 (ufw already open) when that work starts. |
| R6a | ONE recorded exception to R6 (main's ruling, 6 October 2026, 20:0x UK): the three Discord webhook URLs live at /srv/discord-hooks/env (mode 600, owner build), installed by infra/build-server/discord-hooks/install.sh with IGNEUM_SECRET_ON_BOX_OK=1, because the Mac sleeps and the bot's timer must not. The reasoning: a webhook URL signs no release, moves no funds and reaches none of the devnet's hands; whoever holds it can only post as the bot, and a leaked one is deleted in Discord in one click and the file replaced. No other secret joins it; check prints key names, never values. |
| R7 | CLAUDE.md line "nothing is built on a server" and "Windows builds go to the GitHub runner, Linux binaries come from infra/cross/build-linux.sh" are rewritten when R1 and R2 are adopted; until then the box is the measured option, not the rule. |
5. Gotchas met on the first day
| Case | What happened | Fix |
|---|---|---|
| Stale overlay blocks the next checkout (PC 1 worker, 6 Oct 2026) | build-remote.sh rsyncs uncommitted files over the box's checkout; on the next commit git checkout -B refused with "local changes would be overwritten" |
remote-run.sh checkout mode: git checkout -- . and git clean -fd (target dirs, sha stamps and ignored files kept) before the branch checkout, then the overlay; remote-run.sh --self-test reproduces the dirty tree and shows the mode landing on the new commit clean |
| An all-identical overlay listed only directories | the first run's touch pipeline got an empty file list and failed | files only are counted and re-stamped (lib.sh bs_overlay_dir) |
bash 3.2 on the Mac treats an empty array as unbound under set -u |
run-from-mac.sh died on PASS[*] |
a string instead of an array |
grep -q plus pipefail turned a strings hit into a miss |
the commit-string gate failed a stamped igneumd on its first use (SIGPIPE on strings) |
grep -c |
| A path dependency inside a vendor repo (the shipper's proving build, 6 Oct 2026) | proving/igneum-prove depends on vendor/igneum-node-exec/igneum/evm-types, a MEMBER of the fork's workspace (it inherits thiserror from the fork's root manifest); the first design synced that one directory, so cargo found no workspace root on the box ("failed to load manifest for workspace member"), and the shipper cross-built the Linux prove-host on the Mac with cargo-zigbuild meanwhile |
lib.sh groups path dependencies by git top level: one under vendor/ is a whole repository, pushed to its mirror (a fork worktree such as igneum-node-exec goes to /srv/igneum-node.git, which already held its branch; a repository of its own gets /srv/.git, created on first use) and checked out whole at /srv/builds//vendor/; run-from-mac.sh wires every vendor repo the Cargo.toml files reach (today only igneum-node-exec). The detector's first version tested "under BS_TOP" before "own repository" and missed it, since vendor/ lies under the igneum top level on disk. Then libprotobuf-dev was missing (sp1-prover-types's build script imports google/protobuf/empty.proto); added to provision.sh. Proof: cd proving/igneum-prove && tools/build-remote.sh -- build --release -p igneum-prove-host: igneum-prove-host 71,943,192 B, sha256 e9213e3a6c979512d7859f6d8e848105bab53f4355e99fb0d30fb4a72c2d5714, ELF x86-64, 1 min 05 s warm (the cold run compiled 605 crates in 55 s before protoc stopped it). A Mac worktree has no vendor/ of its own, so a worktree that builds proving needs git -C vendor/igneum-node worktree add <wt>/vendor/igneum-node-exec execution-layer first, as the fork worktrees do |
| One remote build lost its ssh session after 75 s (18:30:50 UTC, the first full proving build) | the remote bash died with it (slot line left behind, no JSONL line); no OOM, no reboot, the retry a minute later passed | the dashboard collector's one-pass unit finished within a second of the drop, so it was tested: a 90 s remote session through the same ControlMaster path survived two collector passes triggered by hand; the collector only reads (/proc, lock files, flock -n, sccache --show-stats, kill(pid, 0)). One event, no cause in the journal, the retry passed. A build that must survive a dropped connection would need the remote command under setsid with the Mac reconnecting to wait; not done, open if it happens again |
| A stamp file at a nested path failed every checkout of /srv/builds/igneum (18:31 to 18:5x UTC, the shipper's 18:52 run) | a repo-kind crate keeps its .build-remote-sha-<target> and target/ inside the crate dir; the checkout mode's clean-tree test only excused them at the tree root |
the test is depth-agnostic and uses --untracked-files=all (a wholly untracked directory is otherwise collapsed to ?? dir/); the self-test carries a nested stamp, a nested target dir and a stale .git/index.lock, which the mode now removes when no git runs there |
| Two runs on one worktree at once (the shipper, 18:48:56Z) | the second run's checkout replaced the first's sources mid-cargo; both died | lib.sh takes a per-worktree lock on the box (/srv/builds/_locks/wt-<worktree>, mkdir-atomic, holder line) across sync, build and fetch; a second run waits up to 2 h (a line every minute), a lock older than 3 h is taken over; released on EXIT. The build slot (build-<k>) is unchanged |
The 0.3.15 prover pair for the PCs needs --features igneum-prove-host/cuda (the PCs run SP1_PROVER=cuda) |
my first proving build named no feature | built on the box from master e1b5bc9: igneum-prove-host 73,161,528 B sha256 71bc2438856bb141f6cad3d18489f708568144fad5a06a002fbadefb9ce256f9, igneum-prove-export 3,609,360 B sha256 263bf4cef70af4a13a45b2e79b8dbab373282ea4c571f624d02ddb5791935361 (52 s warm, no CUDA needed at build time); handed to the shipper |
| Let's Encrypt saw NXDOMAIN for build.igneum.network | the deSEC record was minutes old; Ubuntu's Caddy then fell back to ZeroSSL and failed with HTTP 422 for ever | issuer pinned to Let's Encrypt; the retry got the certificate |
5a. The GPU workers (added 6 October 2026, 19:10 UTC, for the class v4 rehearsal)
| What | Fact |
|---|---|
| Headers on the box | provision.sh step_cuda: NVIDIA's ubuntu2404 apt repository, cuda-nvrtc-dev-12-8, cuda-cudart-dev-12-8, cuda-driver-dev-12-8 (the libcuda stub) and Ubuntu's opencl-c-headers; no nvcc (no build file calls it: both workers compile from C/C++ sources with the headers and dlopen libcuda, libnvrtc and libOpenCL at run time), no driver (no GPU here). 12.8 is the minor proto-cuda/nvrtc/fetch-redist.sh pins |
| Command | tools/workers-remote.sh [--out <dir>] from any igneum worktree: proto-cuda and proto-opencl through the mirror and overlay, clang++ 18 with -static-libstdc++ -static-libgcc, both workers, sha256 and the glibc ceiling printed, artefacts in <worktree>/infra/cross/out-workers-box/ |
| Proof (master worker.cpp at 3b2c840) | igneum-worker-cuda 1,613,992 B sha256 6db8a9ad295a9f7598a8b5500f1c271fbfdb9a8df90848b95ff340664007031e; igneum-worker-opencl 124,904 B sha256 0dea75bb8d2d54721ee44b55a3ba486241d3c6053e72a133f8f5e40d2355cbbd; 3 s on the box |
| Consequence: glibc ceiling 2.38 | runs on the fleet (Ubuntu 22.04 containers are glibc 2.35: NO, 2.38 > 2.35; Ubuntu 24.04 hosts yes). The Mac's zig build (infra/cross/build-workers-linux.sh, glibc 2.36) is the one for Debian 12 and HiveOS; the fleet agent must check its boxes' glibc before swapping the worker. Fix if needed: zig on the box (open row in section 6) or clang -target x86_64-linux-gnu.2.35 via zig; both a day's work, not done |
| Runner fix found on the way | a command string carrying set -e leaked into remote-run.sh through eval and killed the runner before its RESULT line (reported as rc 101); the runner now evaluates the command in a subshell |
6. What the box does not do yet
| Gap | Why it matters | Next step |
|---|---|---|
| No zig / cargo-zigbuild | the devnet seed (Debian 12, glibc 2.36) takes the Mac's zig build; a native box build links glibc 2.39, which Debian 13 seeds accept and HiveOS (Ubuntu 18/20 base) does not | install zig 0.17 + cargo-zigbuild in provision.sh, add --target x86_64-unknown-linux-gnu.2.36 mode to build-remote.sh |
| No macOS target | agents who run nodes on the Mac still build there | out of scope (needs the macOS SDK on Linux); the fleet or the box's own Devnet 2 seed takes the test-network runs instead |
| No CI runner | DONE 6 October 2026, 19:19Z (section 7): the runner igneum-build-1 is online under user runner, never build; the workflow change is proposed in docs/plans/ci-self-hosted.md |
main flips IGNEUM_CI_RUNNER=box after the shipper's cut |
| Byte identity with the Mac's exes | different C/C++ toolchain (Homebrew mingw vs Ubuntu GCC 13) and embedded source paths | not a goal; the box is identical with itself build to build, cross-remote.sh reports sha256 and the DLL list per exe |
| Robot API | ~/.config/igneum/hetzner-token is the Cloud token (hcloud); the dedicated box lives in Robot, a separate credential |
main sets the Robot server name in the UI; a webservice user goes to ~/.config/igneum/robot-credentials when needed |
7. The box's second shift (6 October 2026, from 20:3x UK, the project lead: "what else can our building machine be working on?")
Four items, each its own commit on branch box-work with its section here. Times UTC.
7.1 The GitHub Actions runner (DONE 19:19Z)
| Fact | Value |
|---|---|
| Runner | igneum-build-1, actions/runner 2.338.0 (tarball sha256 af4b794c... checked against the release note), registered on igneum-network/igneum at 19:19:49Z, online, labels self-hosted, Linux, X64, igneum-build-1 |
| User | runner (uid 1001, own group, no sudo, not in build's group; /home/runner 750). Never build, never root. The runner's credential (/opt/actions-runner/.credentials, mode 600) is the only thing it holds; it signs nothing and reaches no hand |
| Service | actions.runner.igneum-network-igneum.igneum-build-1.service (GitHub's svc.sh install runner), drop-in igneum.conf: Nice 10, IO best-effort 7, Restart on-failure. Agents' builds (nice 0 through build-remote.sh) win the CPU over a CI job |
| Toolchains | rustup 1.99.0 pinned like the box (RUST_TOOLCHAIN), targets x86_64-unknown-linux-gnu and x86_64-pc-windows-gnu, clippy, rustfmt; mingw-w64 GCC 13 posix, Node 22, python3 + numpy from the system (numpy added to APT for the simulators) |
| sccache | /usr/local/bin/sccache (the build user's binary copied; /home/build is 750), config /home/runner/.config/sccache/config with rw_mode = "READ_ONLY" on /srv/sccache, own server port 4227. Shown: a job-shaped cargo test --release of igneum-pow as runner (99 tests pass, 42 s cold) made 6 compile requests, 0 hits, 6 cache WRITE ERRORS (the refusal, as wanted), and /srv/sccache stayed at 5,880,836 KB |
| Jobs | .env: RUSTC_WRAPPER, SCCACHE_CONF, SCCACHE_SERVER_PORT 4227, CARGO_INCREMENTAL 0, CARGO_BUILD_JOBS 48 (half the box); .path: the runner's cargo bin, /usr/local/bin, /usr/bin, /bin. A CI job takes NO build slot today (ci-self-hosted.md, open row) |
| Ephemeral | no. --ephemeral is for autoscaled fleets that register a fresh runner per job; one standing runner on a private repository keeps its registration and cleans _work per job (approximate: GitHub's docs host answered 404 to both fetches tonight, so this is the rule as remembered, labelled so) |
| Registration | infra/build-server/runner/register.sh: gh as igneum-labs (fails on any other active account, checks the login is igneum-labs), POST repos/igneum-network/igneum/actions/runners/registration-token, the token as the first stdin line to provision.sh on the box (never an argument, never a file, never logged; the output is filtered for it as a belt). --status lists the repository's runners and the unit |
| Idempotent | provision.sh runs 3 and 4 after the registration: runner: ok, no restart (ActiveEnterTimestamp unchanged). Run 2 had said changed because GitHub's svc.sh install runs env.sh, which rewrites .env and .path; the step now writes them AFTER the install |
| Workflows | NOT changed (the shipper owns them tonight). The proposed diff and the fallback (repository variable IGNEUM_CI_RUNNER; GitHub has no "else" in runs-on) are in docs/plans/ci-self-hosted.md. windows.yml cannot move to a Linux box (MSVC, WebView2, Inno Setup, PowerShell 5.1) |
Consequences: ci.yml's pow and sims jobs would run on a pinned 1.99.0 (GitHub's ubuntu-latest ships whatever stable it has), with 48 jobs; GitHub-hosted minutes on a private repository are the thing saved. A CI job on the box reads only what the checkout gives it; the mirrors and /srv/builds belong to build and are not readable by runner (git's safe.directory is set for the runner so a future job may clone a mirror read-only if main wants it).
7.0 The slots ruling (main, 20:3x UK; DONE 19:28Z, live on the box)
| Change | Where | Shown by |
|---|---|---|
2 build slots (/srv/builds/_locks/slots = 2) |
provision.sh SLOTS default 2, applied 19:28:18Z |
dirs: changed (... slots=2) |
CARGO_BUILD_JOBS 90 when a build holds the only taken slot, 45 when it sees the other slot held (one second of settling after taking the slot, then a flock -n probe of the other file); a -j on the cargo line wins |
remote-run.sh; tools/build-remote.sh and tools/cross-remote.sh pass -j only when --jobs is given |
remote-run.sh --self-test-slots on the box: two concurrent fake builds get 45 each, a lone one 90 |
A measurement (BR_MEASURE=1) takes the measure file exclusively; builds hold it shared for their whole run, so a measure waits for the running builds and blocks new ones, as with-lock.sh's measure on the Mac |
remote-run.sh; infra/build-server/prover/cpu-trial.sh is its first user |
the self-test: a measure blocks a build, a build blocks a measure; JSONL carries "jobs" and "measure" |
| Lock files opened in APPEND mode | remote-run.sh | the first version's exec {fd}>build-k truncated a BUSY slot's holder line each time another build probed it (the dashboard read empty lines for held slots); the self-test's case 5 keeps a holder line through a probe. The OLD script under the same cases: JOBS=none JOBS=none, FAIL (the known-failed run, 19:25Z) |
| An environment IGNEUM_BUILD_SLOTS_DIR or IGNEUM_BUILD_LOG_DIR wins over the profile | remote-run.sh | the first self-test run let the profile reset the scratch dir and took the box's REAL slot for 7 s (19:24Z, box idle) |
Open: a worktree whose remote-run.sh predates this keeps the old behaviour until it has master with it (the script is piped from each Mac worktree per build), so until every agent rebases, a build from an old worktree still asks -j 90 beside a new one at 45. The build-server agent was told at 19:28Z.
7.4 The CPU prover trial (DONE 19:30Z; verdict: the box is NOT a prover)
infra/build-server/prover/cpu-trial.sh on the box under the measure hold (builds excluded), igneum-prove-host from master
7483fb37 (the 0.3.15 prover pair, sha256 71bc2438...; its cuda feature changes nothing under SP1_PROVER=cpu), fixture
proving/fixtures/block-56-transfers.json (the v0 block: 3 transfers, 600 pgas, one shard, 556,369 SP1 cycles), --mode shard --shard 0, RAYON_NUM_THREADS=96, nice 19, box otherwise idle (load 2.6 at start). Log and results JSON:
/srv/builds/_log/prover-trial/trial-20261006T192824Z.{log,json,txt}; JSONL line kind measure.
| Stage | Box, 96 threads (EPYC 9454P) | Mac, same statement class (bench-log) |
|---|---|---|
| setup (prover client, shard and aggregator keys) | 14.4 s (client 12.6, keys 1.8) | 9.3 s per invocation in the v0 loop (4 Oct, "nearly all SP1 setup") |
| execute | 0.28 s, 556,369 cycles, 927 cycles per pgas | 0.19 s for block 78 (3 Oct) |
| core proof | 34.2 s, 7,317,561 B, verify 0.34 s | 83 s for the 200-pgas shard on a Mac at load 40 (4 Oct); 71.7 s is main's Mac figure for this fixture (its stage not recorded here, labelled approximate) |
| compressed proof (what a record carries) | 85.9 s, 1,272,897 B, verify 0.07 s | 272 s for the 200-pgas shard on the loaded Mac (4 Oct); 61 s for the smallest shard in the 3-node v0 loop |
| whole run, wall | 136.9 s | |
| peak RSS | 28.2 GB (VmHWM) | |
| CPU use | 64 of 96 threads busy on average (6,408 percent in ps) |
What the numbers mean, per tier, and what follows:
| Number | Means | Done or proposed |
|---|---|---|
| core 34.2 s under 60 s, compressed 85.9 s over it | the proof a record carries is the compressed one, so the shard that matters takes 120 s of proving on 96 CPU threads for the SMALLEST shard the chain has (600 pgas, 0.56 M cycles); a shard at S_p is 60 M cycles (bench-log 4 Oct), about 100x, so hours per shard on this CPU against the 60 s proof lag the litepaper states |
the 60 s test of the ask is NOT met on the stage that counts; NO standing CPU prover unit is written, nothing joins the devnet from the box (R6 holds: the box is never a node host for the live devnet) |
| 28.2 GB peak for the smallest shard | a CPU prover needs 32 GB of RAM for a toy shard; every home tier (8, 12, 16, 24 or 32 GB CARDS, 16 to 64 GB of RAM) is out of CPU proving, and the rented 4090 boxes' CPUs are not a fallback either | the app keeps "proving on the CPU (slow)" as a correctness lane only; the prover tiers are the real cards (docs/analysis, 6 Oct rented-card measurement) |
| 64 of 96 threads busy | SP1's CPU prover does not scale to the whole box; a second trial with RAYON_NUM_THREADS=48 would show whether half the box proves as fast (then two shards side by side) |
not run tonight (one slot of the box's evening); the script takes --threads |
| 14.4 s setup per invocation | the same per-process cost the Mac pays; a resident prover would pay it once | already the design of the app's prover loop |
The box stays a build and test machine. If main wants a CPU prover anyway for coverage (a prover that is always on, never fast),
the shape is a systemd unit as build with SP1_PROVER=cpu, a throwaway devnet key (never the OTA key, never a hand's key),
--threads 48, Nice 19 and the measure hold taken for the whole run, which would exclude builds for minutes at a time: that
is why it is not written.