igneum/docs/plans/build-server.md
igneum-labs b04c4b3ac9 Build boxes: the class router is a preference with spill-over (a held or overloaded box hands the job to the other one); build-2 gets a third slot and every run there is bounded on its own 32-core band; the spill decision in the first route line, the slot label and the JSONL row
the project lead, 7 October 2026 15:02 UK: build-1 at load 139 / 114 / 90 with both slots held and a 1 h 40 min queue while build-2 read 4.5 with
free slots, because the class router pinned each class to its box. Now lib.sh bs_route_spill reads the preferred box (free slots,
1-minute load) with one ssh and hands the job to the other box when the preferred one has no free slot or sits above load 64 and the
other qualifies; neither qualifying queues on the class's own box. The decision travels as BR_ROUTE_* into the JSONL "route" object
for the dashboard. build-2's slots file reads 3; everything on box 2 runs at nice 10 / 32 cores / -j 32, and a bounded run takes
the band its slot owns so three never share a core. Self-test tools/ci/route-spill-check.sh (thirteen cases) in the gate.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 13:08:19 +00:00

56 KiB

Build server: igneum-build-1 (plan, benchmarks, rules)

6 October 2026. the project lead ordered a Hetzner dedicated server at 18:2x UK: AX162-1-LTD, EPYC 9454P (48 cores, 96 threads), 128 GB, 2x 3.84 TB NVMe, Falkenstein (FSN1-DC24, Hetzner #3088308), 188.40.146.49, IPv6 2a01:4f8:2240:205a::/64. Purpose: this Mac is the compile queue for ten agents (8-minute rebuilds under contention) and the Windows cross-builds; the box takes over Rust builds, cross-builds and a shared compile cache, and later a CI runner and the Devnet 2 seed.

1. What is in place (all on branch build-server)

Piece File State 6 Oct 2026
Install + provision infra/build-server/provision.sh ran: installimage 17:16 to 17:20 UTC (Ubuntu 24.04, RAID 1, no swap), provision 17:24 to 17:27 UTC, second run 2 s with zero changes (idempotent)
Mac side infra/build-server/run-from-mac.sh <ip> ran: ~/.config/igneum/build-server = build@188.40.146.49, build remotes added, 124 igneum branches and 48 fork branches pushed to the bare mirrors
Shared library infra/build-server/lib.sh host line, ssh options, mirror push, remote checkout, overlay, remote slot runner
Remote runner infra/build-server/remote-run.sh runs on the box: slot, sccache, RESULT line, one JSON line per run in /srv/builds/_log/builds.jsonl (the worker dashboard's feed, fields as main asked on 6 Oct)
Remote cargo tools/build-remote.sh any crate dir, any worktree: tools/build-remote.sh [-- cargo args]; IGNEUM_AGENT=<name> tags the slot label and the log
Remote Windows cross tools/cross-remote.sh fork worktree or app/igneum-app: tools/cross-remote.sh [--compare <dir of the Mac's exes>]

ssh line: ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 (root works with the same key; passwords are off).

Box facts (provision summary): rustc 1.99.0 (the Mac's version; NO rust-toolchain file exists in the repo or the fork, the pin is RUST_TOOLCHAIN in provision.sh), targets x86_64-unknown-linux-gnu and x86_64-pc-windows-gnu, sccache 0.18.0 with a 100 GB disk cache at /srv/sccache, mingw-w64 GCC 13 posix threads (the PC job's recipe), clang and lld 18.1.3, Node v22.23.3, git 2.43.0, tmux 3.4, ufw 22 + 26611 + 26811 (the devnet and testnet seed p2p ports; RPC stays on loopback as on every seed), md0 /boot and md1 / as RAID 1, 3.3 TB free, swap none, 1 build slot in /srv/builds/_locks/slots.

2. How a build travels

  1. bs_context (lib.sh) reads the crate: a fork worktree under vendor/ (kind node, mirror /srv/igneum-node.git) or a crate of the igneum repo (kind repo, mirror /srv/igneum.git). cargo metadata lists every path dependency.
  2. HEAD is pushed to the mirror; the box clones or fetches it at /srv/builds/<worktree>/<same relative path> and checks out that commit. A real .git is needed: the fork's kaspa-build-info runs git rev-parse HEAD at build time and every release plan checks the commit in the binary's strings.
  3. Uncommitted changes go by rsync --checksum without -t (files that changed are written with the box's clock) and the written files are re-stamped with touch (the copied-sources rule, tools/ci/copied-sources-check.sh passes).
  4. The cargo command runs in the crate dir under one of the box's slot files (flock on /srv/builds/_locks/build-<k>, never the Mac's ~/.config/igneum/build-slots), with sccache and -j 90, logging wall time, compile count and cache hits.
  5. Artefacts come back into <crate>/target-remote/... with size and sha256. NEVER into target/: the box's native binaries are x86_64 Linux (glibc 2.39) and do not run on the Mac.

/srv/builds/<worktree> mirrors the Mac's igneum worktree ROOT because the fork's path dependency is ../../../../igneum-pow (vendor/igneum-node/consensus/pow/Cargo.toml).

3. Benchmarks

The Mac numbers are from the logs (docs/bench-log.md, docs/plans/release-0.3.10.md, release-0.3.11.md); no fresh Mac run was made, because a Mac build would take a slot from the agents and the logs already hold several readings per case. Mac = Apple M5 Max (18 cores), builds at nice -n 19 with 4 to 6 jobs under the build lock, usually loaded by other agents. Box = igneum-build-1, -j 90, nothing else running.

Case Mac (logged) Box Box run
Clean node build (cargo build --release -p kaspad -p igneum-miner --features kaspad/igneum-pow, every crate) 12 min 36 s (0.3.10, new target dir, -j 4); 17 min 53 s at load 140 and 8 min 58 s second time (fud ledger, -j 4); rusty-kaspa kaspad alone 2 min 36 s on an idle Mac 1 min 27 s cargo wall (1 min 34 s end to end from the Mac: push and sync 1 s, build, fetch of both binaries) cold: 518 crates downloaded inside that time, sccache 0 hits / 993 misses, 563 crates compiled; igneumd 48,220,896 B, igneum-miner 9,107,232 B, ELF x86-64 PIE; igneumd --version runs on the box; 17:30:58 to 17:32:32 UTC
Incremental rebuild (one file changed, same target dir) 2 min 08 s (0.3.10, 17:49Z); 3 min 19 s (0.3.11); 5 min 06 s (txgossip); 15 min 18 s at load 110 to 134 (M31); 35 s to 4 min (execution layer); "8 min under contention" (main, 6 Oct) 7 s cargo wall (10 s end to end: sync 1 s, build 6.97 s, fetch) one line appended to kaspad/src/main.rs in the fork worktree, uncommitted, carried by the overlay (1 file written and re-stamped); 1 crate compiled, igneumd relinked; box idle (load 12 from the clean build a minute earlier); 17:33:11 to 17:33:21 UTC
Windows cross-build (--target x86_64-pc-windows-gnu, igneumd.exe + igneum-miner.exe) 8 min 25 s clean (-j 6, 3 Oct); 4 min 49 s and 4 min 53 s with warm dependencies (4 Oct); 12 min 28 s (finality-fixes, 4 Oct) 1 min 44 s cargo wall (1 min 48 s end to end) cold: 995 compiles, sccache 0 hits; igneumd.exe 49,434,624 B, igneum-miner.exe 10,201,600 B; mingw GCC 13 posix, libclang 18; igneumd.exe imports libstdc++-6.dll (this fork head predates the housekeeping commit that made the C++ runtime static), so cross-remote.sh now fetches the three GCC 13 runtime DLLs beside the exes as the PC job does; 17:33:58 to 17:35:46 UTC. Second run on the same target dir, no source change: 8 s

Also logged for context: the Mac's Linux cross-build with zig (infra/cross/build-linux.sh) took 30 min 56 s cold and 3 min 20 s incremental; PC 1 built Linux + Windows node and app and ran both test suites in 7 min 38 s cold, 5 min 09 s warm (CLAUDE.md, 5 Oct).

What the numbers mean and what follows

Number Means Done or proposed
Clean build 1 min 27 s against 12 to 18 min on the loaded Mac (8 to 12x) a new worktree costs an agent a minute and a half, not a slot for a quarter of an hour; the first build of every one of the 123 worktree dirs on the box is this cold case, later ones are the 7 s case R1 proposed; sccache was cold (0 hits of 1,982 requests) because every run so far was the first of its kind; the cache fills as agents build the same crate versions from different worktrees, so the second worktree's clean build will be mostly hits (measure when it happens, write the number here)
Incremental 7 s against 2 to 15 min on the Mac (the "8 min under contention") an edit-build loop of seconds for Linux targets; end to end 10 s because sync is 1 s and the two binaries (57 MB) come back in 2 s R1; the slot count stays 1 until two agents collide, then 2 with -j 48 each (SLOTS=2 run-from-mac.sh and --jobs 48)
Windows cross 1 min 44 s against 4 min 49 s to 12 min 28 s on the Mac (3 to 7x), 8 s incremental a Windows exe per commit is cheap enough to build on every push; the PC build job (7 min 38 s cold, 5 min 09 s warm for Linux + Windows + tests) stays the second source R2 proposed
Mac arm64 binaries: not built here agents who run nodes on the Mac (local devnets, the DMG) still take Mac slots; the box cannot remove that contention R3; the real relief is to move test networks to the fleet or to a Devnet 2 seed on this box (its unit, ports and ufw are ready)
The commit hash is EMPTY in every Mac worktree build and was empty in the first box builds kaspa-build-info (build-info/build.rs) needs .git to be a directory AND HEAD to be a symbolic ref to a loose branch file; a worktree's .git is a file and a detached HEAD is not a ref; and once it has found nothing it emits no rerun-if-changed, so cargo never runs it again in that target dir (release-0.3.11: cargo clean -p kaspa-build-info) fixed on the box: the remote checkout is git checkout -B <branch> <sha> and build-remote.sh runs cargo clean --release -p kaspa-build-info whenever the commit differs from the last one built in that target dir (.build-remote-sha-<target dir>); verified 17:40 UTC: 2 string hits for 3bfe346f, binary +1,024 B. The release plans' "commit in its strings" checks were passing against builds from the fork's MAIN checkout (a real .git directory on a branch), not from worktrees. DONE (main's decision): the PC job writes a minimal node/.git (HEAD -> refs/heads/build holding the manifest's new commit_full, written by push-build-inputs.sh) at extract and cleans kaspa-build-info on a new commit (jobbuild.rs, 4 unit tests pass on the box); the Mac's cross-build.sh refuses a worktree and cleans on a new commit; the gate tools/ci/commit-string-check.sh runs on every igneumd the three scripts produce (self-test in ci.yml; shown firing on the Mac's worktree-built igneumd and passing on the box's). Only kaspad depends on kaspa-build-info, so igneum-miner is out of the gate's scope
igneumd.exe hash changed build to build with no source change (c424aae0 then bfbff015) while the Linux igneumd stayed byte-identical across three builds the mingw linker wrote a timestamp into the PE header; the Linux ELF has none DONE (main's decision): -C link-arg=-Wl,--no-insert-timestamp in cross-remote.sh, the Mac's cross-build.sh and the PC job (jobbuild.rs); verified 17:48 and 17:49 UTC: two builds, igneumd.exe c38b7570... and igneum-miner.exe c3a0fc2b... identical both times
Second worktree's clean build (vendor/igneum-node-v4 from the main checkout, cold target dir, warm sccache): 1 min 18 s, 604 hits of 993 compiles (61 percent), 389 misses the first build of each of the 123 worktree dirs costs 1 min 18 s to 1 min 27 s, not the Mac's 12 to 18 min; sccache saves 9 s of the 87 because the clean build is dominated by the fork's own 74 crates and the C++ (rocksdb) objects, which differ per tree or do not cache; the cache is 1.0 GB after four builds, capped at 100 GB nothing; the hit rate is in every RESULT line and every JSONL line
Disk 3.3 TB free, RAM 125 GB, load peaked at 24 during the Windows build with 96 threads room for 2 slots and the Devnet 2 seed without contention nothing now

4. Rules (ADOPTED by main on 6 October 2026, written into CLAUDE.md "Running agents on this Mac"; R2 narrowed: the PCs keep only GPU and Windows-runtime jobs from now, not two releases)

Rule Text
R1 Every agent's cargo build, cargo test, cargo check and cargo clippy for Linux goes through tools/build-remote.sh from the crate directory. The box takes one remote slot per build; nobody runs cargo over ssh by hand.
R2 Windows exes come from tools/cross-remote.sh (fork worktree: igneumd.exe, igneum-miner.exe; app/igneum-app: igneum-app.exe and the two tools). The PC build job stays as the second source until two releases have shipped from the box.
R3 The Mac keeps what only it can do: aarch64-apple-darwin binaries (the DMG, nodes that agents run locally), tests that need Metal (proto-metal, the Metal worker), and measurements. Those still use tools/lock/with-lock.sh and the Mac's build slots.
R4 A worktree builds once on the box per commit plus overlay; the next build is incremental in /srv/builds/<worktree>/.../target. Nobody deletes another worktree's target dir on the box.
R4a A lane's scratch on the box mirror survives other lanes' builds (7 October 2026, after the attack rows lost attack-f3/, attack-f1-venv/ and tools/attack/*/target to each other's builds, hazard AP-H1). The checkout's git clean spares attack-*, scratch-*, target-attack-* and .build-remote.log at any depth, plus every glob in the mirror-local file /srv/builds/<worktree>/.igneum-scratch-spare (one glob per line, gitignore syntax, # comments; the file itself is spared). To declare a prefix: before the first build that must leave it alone, append one line named for the lane (bs-<name>, r03xx-ship, the scratchpad prefix rule of 7 October): ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 "printf 'bs-mylane-*\n' >> /srv/builds/<worktree>/.igneum-scratch-spare". Keep scratch outside the crate directories the overlay syncs (bs_overlay_dir rsyncs those with --delete, so an undeclared dir inside one goes the Mac's way regardless; a target-attack-* name is safe there too, rsync protects its excluded target-*/). remote-run.sh --self-test shows a declared and a fixed-prefix dir surviving and an undeclared one removed; tools/ci/scratch-spare-check.sh (pre-push gate) fails when the clean line loses the mechanism or gains -x.
R5 RUST_TOOLCHAIN in provision.sh is bumped in the same commit as the Mac's rustup update; build-remote.sh refuses a mismatch. Add a rust-toolchain.toml to the fork and the repo (none exists today) so both sides pin from one file.
R6 The box is never a node host for the live devnet and never holds a secret (no ~/.config/igneum there). The Devnet 2 seed on it runs under its own unit with --devnet --devnet-suffix=<n> on 26611 (ufw already open) when that work starts.
R6a ONE recorded exception to R6 (main's ruling, 6 October 2026, 20:0x UK): the three Discord webhook URLs live at /srv/discord-hooks/env (mode 600, owner build), installed by infra/build-server/discord-hooks/install.sh with IGNEUM_SECRET_ON_BOX_OK=1, because the Mac sleeps and the bot's timer must not. The reasoning: a webhook URL signs no release, moves no funds and reaches none of the devnet's hands; whoever holds it can only post as the bot, and a leaked one is deleted in Discord in one click and the file replaced. No other secret joins it; check prints key names, never values.
R7 CLAUDE.md line "nothing is built on a server" and "Windows builds go to the GitHub runner, Linux binaries come from infra/cross/build-linux.sh" are rewritten when R1 and R2 are adopted; until then the box is the measured option, not the rule.

5. Gotchas met on the first day

Case What happened Fix
Stale overlay blocks the next checkout (PC 1 worker, 6 Oct 2026) build-remote.sh rsyncs uncommitted files over the box's checkout; on the next commit git checkout -B refused with "local changes would be overwritten" remote-run.sh checkout mode: git checkout -- . and git clean -fd (target dirs, sha stamps and ignored files kept) before the branch checkout, then the overlay; remote-run.sh --self-test reproduces the dirty tree and shows the mode landing on the new commit clean
The same class met twice more by the UI lane on its mirror (worktrees whose tools predate 3e6a488 pipe the OLD remote-run.sh to the box per build) nothing builds until the tree is cleared by hand the fix is in master's remote-run.sh since 3e6a488; a worktree gets it by merging master; tools/ci/mirror-reset-check.sh (in the pre-push gate) reads checkout_tree() and fails when the reset and the clean do not both come before the branch checkout; its self-test shows the real script passing, a copy without the reset failing, and the same lines moved after the checkout failing. cargo-audit is owned by provision.sh step_cargo_tools (1eec354, the night battery's agent), confirmed ok (cargo-audit 0.22.2) on the box
An all-identical overlay listed only directories the first run's touch pipeline got an empty file list and failed files only are counted and re-stamped (lib.sh bs_overlay_dir)
bash 3.2 on the Mac treats an empty array as unbound under set -u run-from-mac.sh died on PASS[*] a string instead of an array
grep -q plus pipefail turned a strings hit into a miss the commit-string gate failed a stamped igneumd on its first use (SIGPIPE on strings) grep -c
A path dependency inside a vendor repo (the shipper's proving build, 6 Oct 2026) proving/igneum-prove depends on vendor/igneum-node-exec/igneum/evm-types, a MEMBER of the fork's workspace (it inherits thiserror from the fork's root manifest); the first design synced that one directory, so cargo found no workspace root on the box ("failed to load manifest for workspace member"), and the shipper cross-built the Linux prove-host on the Mac with cargo-zigbuild meanwhile lib.sh groups path dependencies by git top level: one under vendor/ is a whole repository, pushed to its mirror (a fork worktree such as igneum-node-exec goes to /srv/igneum-node.git, which already held its branch; a repository of its own gets /srv/.git, created on first use) and checked out whole at /srv/builds//vendor/; run-from-mac.sh wires every vendor repo the Cargo.toml files reach (today only igneum-node-exec). The detector's first version tested "under BS_TOP" before "own repository" and missed it, since vendor/ lies under the igneum top level on disk. Then libprotobuf-dev was missing (sp1-prover-types's build script imports google/protobuf/empty.proto); added to provision.sh. Proof: cd proving/igneum-prove && tools/build-remote.sh -- build --release -p igneum-prove-host: igneum-prove-host 71,943,192 B, sha256 e9213e3a6c979512d7859f6d8e848105bab53f4355e99fb0d30fb4a72c2d5714, ELF x86-64, 1 min 05 s warm (the cold run compiled 605 crates in 55 s before protoc stopped it). A Mac worktree has no vendor/ of its own, so a worktree that builds proving needs git -C vendor/igneum-node worktree add <wt>/vendor/igneum-node-exec execution-layer first, as the fork worktrees do
One remote build lost its ssh session after 75 s (18:30:50 UTC, the first full proving build) the remote bash died with it (slot line left behind, no JSONL line); no OOM, no reboot, the retry a minute later passed the dashboard collector's one-pass unit finished within a second of the drop, so it was tested: a 90 s remote session through the same ControlMaster path survived two collector passes triggered by hand; the collector only reads (/proc, lock files, flock -n, sccache --show-stats, kill(pid, 0)). One event, no cause in the journal, the retry passed. A build that must survive a dropped connection would need the remote command under setsid with the Mac reconnecting to wait; not done, open if it happens again
A stamp file at a nested path failed every checkout of /srv/builds/igneum (18:31 to 18:5x UTC, the shipper's 18:52 run) a repo-kind crate keeps its .build-remote-sha-<target> and target/ inside the crate dir; the checkout mode's clean-tree test only excused them at the tree root the test is depth-agnostic and uses --untracked-files=all (a wholly untracked directory is otherwise collapsed to ?? dir/); the self-test carries a nested stamp, a nested target dir and a stale .git/index.lock, which the mode now removes when no git runs there
The attack rows lost scratch dirs to each other's builds (7 Oct 2026, 09:2x UK, hazard AP-H1 from the attack-pass lane) checkout_tree's git clean -fd ran on the shared mirror before every build from any agent and took every untracked directory: attack-f3/, attack-f1-venv/, tools/attack/*/target the clean spares the fixed prefixes and the globs in .igneum-scratch-spare (R4a), still without -x so .git/info/exclude applies; the clean-tree test asks git clean -nd with the same excludes instead of filtering the status list, so a spared dir no longer reads as "not clean"; the self-test carries a fixed-prefix dir at the root and nested, a declared dir, the spare file and an undeclared dir; tools/ci/scratch-spare-check.sh in the gate
Two runs on one worktree at once (the shipper, 18:48:56Z) the second run's checkout replaced the first's sources mid-cargo; both died lib.sh takes a per-worktree lock on the box (/srv/builds/_locks/wt-<worktree>, mkdir-atomic, holder line) across sync, build and fetch; a second run waits up to 2 h (a line every minute), a lock older than 3 h is taken over; released on EXIT. The build slot (build-<k>) is unchanged
The 0.3.15 prover pair for the PCs needs --features igneum-prove-host/cuda (the PCs run SP1_PROVER=cuda) my first proving build named no feature built on the box from master e1b5bc9: igneum-prove-host 73,161,528 B sha256 71bc2438856bb141f6cad3d18489f708568144fad5a06a002fbadefb9ce256f9, igneum-prove-export 3,609,360 B sha256 263bf4cef70af4a13a45b2e79b8dbab373282ea4c571f624d02ddb5791935361 (52 s warm, no CUDA needed at build time); handed to the shipper
Let's Encrypt saw NXDOMAIN for build.igneum.network the deSEC record was minutes old; Ubuntu's Caddy then fell back to ZeroSSL and failed with HTTP 422 for ever issuer pinned to Let's Encrypt; the retry got the certificate

5b. Reproducible builds (main's rule, 6 October 2026, from the 0.3.14 repro docs/evidence/reproduced/0.3.14.md)

Class What differed Standard fix, in every build path
prost's generated protowire.rs embeds OUT_DIR two builds in differently named target dirs give different bytes ONE fixed target path per target: target (or the name --target-dir gives) on the box, $CARGO_TARGET_DIR on the Mac's cross-build.sh, the persistent dir of the PC job; never a per-run name
libmimalloc-sys compiles mimalloc's C with __DATE__/__TIME__ two builds a minute apart differ when mimalloc recompiles SOURCE_DATE_EPOCH = the node commit's author time and TZ=UTC, exported in lib.sh bs_repro_env (build-remote.sh, cross-remote.sh, workers-remote.sh), remote-run.sh (BR_SDE, logged in the JSONL line as source_date_epoch), proto-cuda/windows-node/cross-build.sh, and the PC job (push-build-inputs.sh writes node.commit_time, jobbuild.rs exports it before every cargo build of a stage; unit test asserts it)
sccache hid both a cache hit returns the first build's object the self-test runs with RUSTC_WRAPPER=/usr/bin/env (a pass-through; an EMPTY value is "unset" to cargo and would fall back to the configured sccache)

Self-test: tools/build-remote.sh --self-test-repro [--full] from a fork worktree. Run 6 Oct 20:02 UTC on the box (igneum-node-bs at 3bfe346f, epoch 1791120573): igneum-miner twice a minute apart, kaspa-grpc-core cleaned in between: MATCH 91e130f52438edf012466d3f1d3d634ab9263cd7858e67d2dd8d590b152f8a22; the same build into a per-run target path: 548671e7... (differs, the OUT_DIR class shown firing). --full (kaspad with libmimalloc-sys recompiled a minute later, with and without the epoch): run 6 Oct 20:04 to 20:08 UTC (247 s): kaspad with libmimalloc-sys recompiled a minute later, epoch set: MATCH 70219bc29cc98ac75da00702b9966f9f5d73cbf10efaca1c8841c00bb5aae7bf; without the epoch, a minute later: 45169e88... (differs, the DATE class shown firing); the miner line again MATCH 91e130f5..., the per-run path 310383f5... (differs). The box is deterministic with the rule and shown non-deterministic without it

5a. The GPU workers (added 6 October 2026, 19:10 UTC, for the class v4 rehearsal)

What Fact
Headers on the box provision.sh step_cuda: NVIDIA's ubuntu2404 apt repository, cuda-nvrtc-dev-12-8, cuda-cudart-dev-12-8, cuda-driver-dev-12-8 (the libcuda stub) and Ubuntu's opencl-c-headers; no nvcc (no build file calls it: both workers compile from C/C++ sources with the headers and dlopen libcuda, libnvrtc and libOpenCL at run time), no driver (no GPU here). 12.8 is the minor proto-cuda/nvrtc/fetch-redist.sh pins
Command tools/workers-remote.sh [--out <dir>] from any igneum worktree: proto-cuda and proto-opencl through the mirror and overlay, clang++ 18 with -static-libstdc++ -static-libgcc, both workers, sha256 and the glibc ceiling printed, artefacts in <worktree>/infra/cross/out-workers-box/
Proof (master worker.cpp at 3b2c840) igneum-worker-cuda 1,613,992 B sha256 6db8a9ad295a9f7598a8b5500f1c271fbfdb9a8df90848b95ff340664007031e; igneum-worker-opencl 124,904 B sha256 0dea75bb8d2d54721ee44b55a3ba486241d3c6053e72a133f8f5e40d2355cbbd; 3 s on the box
Consequence: glibc ceiling 2.38 runs on the fleet (Ubuntu 22.04 containers are glibc 2.35: NO, 2.38 > 2.35; Ubuntu 24.04 hosts yes). The Mac's zig build (infra/cross/build-workers-linux.sh, glibc 2.36) is the one for Debian 12 and HiveOS; the fleet agent must check its boxes' glibc before swapping the worker. Fix if needed: zig on the box (open row in section 6) or clang -target x86_64-linux-gnu.2.35 via zig; both a day's work, not done
Runner fix found on the way a command string carrying set -e leaked into remote-run.sh through eval and killed the runner before its RESULT line (reported as rc 101); the runner now evaluates the command in a subshell

5c. Deploy key for the observer clone (DONE: the project lead added the key on 7 October 2026, morning)

Done. the project lead added the public half as the read-only deploy key "igneum-build-1 observer (read-only)" (SHA256:51ice3W8...; the organisation's deploy-key policy had to be switched to Enabled first). The sync's own test then still said "mirror": ssh -T to GitHub exits 1 after its greeting and the script runs under pipefail, so su ... | grep -q reported failure although grep had matched; fixed by capturing the output first (install-hands.sh). First pass reading "source: github (deploy key accepted)": 7 Oct 2026 07:30:54 UTC, "observer files changed (8aab05ce); restarting igneum-observer": the clone went from the mirror's 8397781 to GitHub's master 8aab05ce in that pass, the observer restarted on it and is active; remote github = git@github-igneum-observer:igneum-network/igneum.git. The clone now follows origin/master every 5 minutes; the mirror stays the fallback whenever GitHub refuses.

The steps as they were, for the record:

The observer on the box runs tools/observer from a clone that follows the mirror /srv/igneum.git, which moves only when a Mac agent pushes. With a read-only deploy key it follows GitHub directly (every 5 minutes, igneum-observer-sync.timer). The key pair was made ON the box by infra/build-server/hands/install-hands.sh as user build; the private half is /srv/observer/.ssh/deploy_igneum (mode 600), never copied and never printed. The public half is what GitHub gets.

Step Where What
1 this Mac ssh -i ~/.ssh/igneum_ed25519 build@188.40.146.49 cat /srv/observer/.ssh/deploy_igneum.pub prints one line starting ssh-ed25519 and ending igneum-build-1 observer read-only (the public half; also at the end of this section)
2 github.com, signed in as the organisation owner (igneum-labs) https://github.com/igneum-network/igneum/settings/keys, "Add deploy key"
3 the form Title igneum-build-1 observer (read-only); Key: paste the line from step 1; leave "Allow write access" UNTICKED; "Add key"
4 this Mac ssh -i ~/.ssh/igneum_ed25519 root@188.40.146.49 'systemctl start igneum-observer-sync.service; journalctl -u igneum-observer-sync -o cat -n 4' must print source: github (deploy key accepted); until then it prints source: mirror (...) and nothing is broken

The sync tests the key with ssh -T git@github-igneum-observer on every pass and falls back to the mirror whenever GitHub refuses, so a revoked key never stops the observer; it only makes it follow the mirror again. Public half: ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIGS9oMQ4E9f0Zx+WTCemiUvun+6G45ecyIW8N1AGJfwi igneum-build-1 observer read-only

6. What the box does not do yet

Gap Why it matters Next step
zig / cargo-zigbuild: DONE 7 Oct 2026 (main's order after a seed took 14 restarts and three minutes down on a glibc 2.39 binary) provision.sh step_zig (zig 0.17.0 from ziglang.org, sha256 of the official index) and cargo-zigbuild 0.23.4 in step_cargo_tools; tools/build-remote.sh --ship [--glibc 2.36] runs cargo zigbuild --target x86_64-unknown-linux-gnu.2.36, fetches from the target-triple dir and runs tools/ci/glibc-ceiling-check.sh (need at most 2.36) on every artefact; tools/workers-remote.sh builds with zig cc/c++ -target x86_64-linux-gnu.2.36 by default (GLIBC=native for clang). Rule: anything that ships to a seed or a HiveOS rig is built with --ship; a plain build is glibc 2.39 for the box and Ubuntu 24.04 hosts only. Proof (fork 3bfe346f, cold through zig, 3 min 17 s): igneumd 47,023,120 B sha256 345dfb95... needs GLIBC_2.34; igneum-miner 9,248,808 B sha256 d09dc27b... needs GLIBC_2.34; the workers through zig at 2.36 (6 s): igneum-worker-cuda 6,759,272 B sha256 690c8e91... and igneum-worker-opencl 307,264 B sha256 329fb6a9..., each needing GLIBC_2.34 and libc only (the first zig attempt died on __isoc23_strtol because -I /usr/include for CL/cl.h put the host's glibc 2.39 headers before zig's; the OpenCL headers are reached through a CL-only symlink dir now); the check's self-test fires on 2.38 against 2.36 and passes 2.34 and 2.36 the Mac's infra/cross/build-linux.sh is the same recipe and can retire once two seed releases shipped from the box
Ceilings per class (main, 7 Oct 2026 00:2x UTC, after RunPod's Ubuntu 22.04 canaries, glibc 2.35, refused the GLIBC_2.38 box binaries, and HiveOS turned out Ubuntu 20.04 based, glibc 2.31) one table in tools/ci/glibc-ceiling-check.sh (--class, --ceiling-of): hive and rig 2.31, seed and linux 2.35, native unchecked; `tools/build-remote.sh --ship hive rig
No macOS target agents who run nodes on the Mac still build there out of scope (needs the macOS SDK on Linux); the fleet or the box's own Devnet 2 seed takes the test-network runs instead
CI runner: DONE by the box-work agent (actions.runner.igneum-network-igneum.igneum-build-1.service)
Byte identity with the Mac's exes different C/C++ toolchain (Homebrew mingw vs Ubuntu GCC 13) and embedded source paths not a goal; the box is identical with itself build to build, cross-remote.sh reports sha256 and the DLL list per exe
Robot API ~/.config/igneum/hetzner-token is the Cloud token (hcloud); the dedicated box lives in Robot, a separate credential main sets the Robot server name in the UI; a webservice user goes to ~/.config/igneum/robot-credentials when needed

7. The box's second shift (6 October 2026, from 20:3x UK, the project lead: "what else can our building machine be working on?")

Four items, each its own commit on branch box-work with its section here. Times UTC.

7.1 The GitHub Actions runner (DONE 19:19Z)

Fact Value
Runner igneum-build-1, actions/runner 2.338.0 (tarball sha256 af4b794c... checked against the release note), registered on igneum-network/igneum at 19:19:49Z, online, labels self-hosted, Linux, X64, igneum-build-1
User runner (uid 1001, own group, no sudo, not in build's group; /home/runner 750). Never build, never root. The runner's credential (/opt/actions-runner/.credentials, mode 600) is the only thing it holds; it signs nothing and reaches no hand
Service actions.runner.igneum-network-igneum.igneum-build-1.service (GitHub's svc.sh install runner), drop-in igneum.conf: Nice 10, IO best-effort 7, Restart on-failure. Agents' builds (nice 0 through build-remote.sh) win the CPU over a CI job
Toolchains rustup 1.99.0 pinned like the box (RUST_TOOLCHAIN), targets x86_64-unknown-linux-gnu and x86_64-pc-windows-gnu, clippy, rustfmt; mingw-w64 GCC 13 posix, Node 22, python3 + numpy from the system (numpy added to APT for the simulators)
sccache /usr/local/bin/sccache (the build user's binary copied; /home/build is 750), config /home/runner/.config/sccache/config with rw_mode = "READ_ONLY" on /srv/sccache, own server port 4227. Shown: a job-shaped cargo test --release of igneum-pow as runner (99 tests pass, 42 s cold) made 6 compile requests, 0 hits, 6 cache WRITE ERRORS (the refusal, as wanted), and /srv/sccache stayed at 5,880,836 KB
Jobs .env: RUSTC_WRAPPER, SCCACHE_CONF, SCCACHE_SERVER_PORT 4227, CARGO_INCREMENTAL 0, CARGO_BUILD_JOBS 48 (half the box); .path: the runner's cargo bin, /usr/local/bin, /usr/bin, /bin. A CI job takes NO build slot today (ci-self-hosted.md, open row)
Ephemeral no. --ephemeral is for autoscaled fleets that register a fresh runner per job; one standing runner on a private repository keeps its registration and cleans _work per job (approximate: GitHub's docs host answered 404 to both fetches tonight, so this is the rule as remembered, labelled so)
Registration infra/build-server/runner/register.sh: gh as igneum-labs (fails on any other active account, checks the login is igneum-labs), POST repos/igneum-network/igneum/actions/runners/registration-token, the token as the first stdin line to provision.sh on the box (never an argument, never a file, never logged; the output is filtered for it as a belt). --status lists the repository's runners and the unit
Idempotent provision.sh runs 3 and 4 after the registration: runner: ok, no restart (ActiveEnterTimestamp unchanged). Run 2 had said changed because GitHub's svc.sh install runs env.sh, which rewrites .env and .path; the step now writes them AFTER the install
Workflows NOT changed (the shipper owns them tonight). The proposed diff and the fallback (repository variable IGNEUM_CI_RUNNER; GitHub has no "else" in runs-on) are in docs/plans/ci-self-hosted.md. windows.yml cannot move to a Linux box (MSVC, WebView2, Inno Setup, PowerShell 5.1)

Consequences: ci.yml's pow and sims jobs would run on a pinned 1.99.0 (GitHub's ubuntu-latest ships whatever stable it has), with 48 jobs; GitHub-hosted minutes on a private repository are the thing saved. A CI job on the box reads only what the checkout gives it; the mirrors and /srv/builds belong to build and are not readable by runner (git's safe.directory is set for the runner so a future job may clone a mirror read-only if main wants it).

7.0 The slots ruling (main, 20:3x UK; DONE 19:28Z, live on the box)

Change Where Shown by
2 build slots (/srv/builds/_locks/slots = 2) provision.sh SLOTS default 2, applied 19:28:18Z dirs: changed (... slots=2)
CARGO_BUILD_JOBS 90 when a build holds the only taken slot, 45 when it sees the other slot held (one second of settling after taking the slot, then a flock -n probe of the other file); a -j on the cargo line wins remote-run.sh; tools/build-remote.sh and tools/cross-remote.sh pass -j only when --jobs is given remote-run.sh --self-test-slots on the box: two concurrent fake builds get 45 each, a lone one 90
A measurement (BR_MEASURE=1) takes the measure file exclusively; builds hold it shared for their whole run, so a measure waits for the running builds and blocks new ones, as with-lock.sh's measure on the Mac remote-run.sh; infra/build-server/prover/cpu-trial.sh is its first user the self-test: a measure blocks a build, a build blocks a measure; JSONL carries "jobs" and "measure"
Lock files opened in APPEND mode remote-run.sh the first version's exec {fd}>build-k truncated a BUSY slot's holder line each time another build probed it (the dashboard read empty lines for held slots); the self-test's case 5 keeps a holder line through a probe. The OLD script under the same cases: JOBS=none JOBS=none, FAIL (the known-failed run, 19:25Z)
An environment IGNEUM_BUILD_SLOTS_DIR or IGNEUM_BUILD_LOG_DIR wins over the profile remote-run.sh the first self-test run let the profile reset the scratch dir and took the box's REAL slot for 7 s (19:24Z, box idle)

Open: a worktree whose remote-run.sh predates this keeps the old behaviour until it has master with it (the script is piped from each Mac worktree per build), so until every agent rebases, a build from an old worktree still asks -j 90 beside a new one at 45. The build-server agent was told at 19:28Z.

7.2 The night battery (DONE 19:42Z installed, dry run 19:45 to 19:49Z)

infra/build-server/night/night-battery.sh, run by igneum-night-battery.timer at 02:00 Europe/London (the box's clock is Europe/Berlin, so the unit names the zone: next run Wed 2026-10-07 03:00 CEST = 02:00 BST; not Persistent, a missed night is not run by day) through igneum-night-battery.service (User build, Nice 19, idle IO, 8 h limit), whose ExecStart is remote-run.sh with the battery as BR_CMD, so ONE build slot spans the whole invocation and one JSONL line records it. Installed by provision.sh step_night from the mirror at NIGHT_REF (box-work tonight, master once merged); the battery re-execs itself from master's checkout at run time. Every row: cargo test --release --no-fail-fast per crate (the repo's five crates and the proving workspace's host, core and export; every fork workspace member, kaspad with igneum-pow), igneum-pow's two fuzz tests at 2,000 programs (10x), the three simulators in full, the fast-time harnesses (tools/finality-attacks/run.mjs --fast-time, tools/harness/run.mjs s3 s4 --fast-time --no-bench-log, tools/exec-sync/reorg.mjs) on igneumd, igneum-miner, igneum-harness-sim and igneum-p2p-probe built into the night checkout's target-integration, clippy per crate dir, cargo audit per Cargo.lock; then docs/benchmarks/night/<date>.md (pass/fail table, "new since last night" against the newest earlier report, commits moved), committed as igneum-labs on branch night-battery (rebuilt on master each night, earlier unmerged reports carried over) and force-pushed to /srv/igneum.git only. Main merges: git fetch build night-battery from the main checkout. The fork branch defaults to the newest release-*-node on the mirror (tonight release-0.3.15-node 713ef876).

Dry run (NIGHT_SUBSET=1, the second slot while the repro held the first, so CARGO_BUILD_JOBS 45): 3 min 54 s wall, 10 pass, 1 FAIL, 1 skip; report docs/benchmarks/night/2026-10-06-dryrun.md on the mirror's night-battery branch (fbb72e5).

Row Result Time Detail
suite igneum-pow pass 51 s 99 passed
suite fork/kaspa-pow pass 1 min 04 s 7 passed
suite fork/igneum-miner pass 49 s 18 passed
fuzz igneum-pow x200 pass 11 s 200 mx8 programs and 200 scratch programs, 800 units each
sim finality_sim.py, finality_v2.py --quick, difficulty/sim.py --quick pass 3 s, 41 s, 5 s 249, 117, 8 table lines
clippy igneum-pow pass 3 s 28 warnings
audit igneum-pow pass 3 s 0 vulnerabilities
audit vendor/igneum-node FAIL 2 s 22 advisories in the fork's lock file: h2 (RUSTSEC-2026-0258, unbounded empty DATA frames), quinn-proto (2026-0185, remote memory exhaustion), rustls (2026-0285, TLS 1.3 handshake across encryption levels), ruint (2026-0220), crossbeam-epoch, anyhow (2026-0190), event-listener, faster-hex (2026-0306, AVX2 read past src), lru (2026-0253), tracing-subscriber (2025-0055), chacha20, spin; 17 unmaintained-crate warnings (async-std discontinued, atty, bincode, derivative, instant, mach, paste, proc-macro-error, rustls-pemfile)
harness skip not in the subset; its first run is the 02:00 battery, so the first full report will show whether the Node harnesses run on Linux unchanged (c4, fud and v3.mjs default IGNEUM_NODE_ROOT to /Users/joshm/Projects/igneum/; the battery sets it)

What the FAIL means and what follows: every node binary shipped so far (and the 0.3.15 one tonight) links h2, quinn-proto and rustls at versions with published advisories; h2 and quinn-proto are in the gRPC and QUIC paths a peer can reach, so these are the remote ones. The fix is a dependency bump in the fork (cargo update -p h2 -p quinn-proto -p rustls -p ruint -p crossbeam-epoch -p anyhow -p event-listener -p faster-hex -p lru -p tracing-subscriber, then the suites), a consensus engineer's hour on a quiet branch, and the row goes green by itself the next night. Until then the row stays FAIL every night and "new since last night" stays quiet about it. The unmaintained-crate warnings are upstream rusty-kaspa's and do not fail the row.

Expected full-run time (not measured yet): the three suites above compile the fork's test targets once (about 1 min each for the first crates, seconds after), so 75 fork members plus the repo crates are estimated at 40 to 70 min; the full sims about 10 min (finality_v2.py is 5 min on the Mac); the harnesses 15 to 30 min; clippy and audit under 10 min. Under 2 h, inside the 8 h limit; the first report at 02:00 BST writes the real number.

7.3 Reproducible builds (DONE 19:52Z; A vs B MATCH on all four, DIFFER against the shipped bytes, both explained)

tools/repro/rebuild-release.sh <version> (the Mac) reads the pins from docs/plans/release-<v>.md (the heading "(node , app )" and the bold hashes of the Linux and Windows rows), makes sure both commits are on the mirrors, and runs infra/build-server/repro/rebuild-on-box.sh on the box: a clean clone of the fork at the node commit on a branch under a clean clone of the repo at the app commit, two clean passes per target in ONE target path each, no sccache, under build slots through remote-run.sh, SOURCE_DATE_EPOCH = the node commit's time, TZ=UTC; the shipped hashes come token-free from the public downloads (the HiveOS tarball for the Linux pair; the installer for the exes, when innoextract can open it) with the plan's hashes as the fallback; the evidence goes to docs/evidence/reproduced/<version>.md. The 0.3.14 run (node 4c6b129d, app a90f6a5): four passes of 70 to 78 s, whole run 5 min 06 s.

Artefact A vs shipped A vs B Why the shipped bytes differ (read off the binaries)
igneumd (box 03f35e05..., 49,600,096 B) DIFFER MATCH shipped 934f393c... was the Mac's zig build for glibc 2.36; the box's needs GLIBC_2.39 (native clang and lld). Commit string 4c6b129d in the box's: 1 hit
igneum-miner (box 900c1f0b..., 9,842,168 B) DIFFER MATCH the same toolchain difference
igneumd.exe (box 166e604e..., 51,758,592 B) DIFFER MATCH shipped 44fa74c0... came from the Mac's Homebrew mingw at 16:52Z, before the --no-insert-timestamp fix (17:48Z) and from a worktree (the empty-commit class); the box's is Ubuntu GCC 13 posix with a zero PE timestamp and the commit string (1 hit). innoextract 1.9 cannot open the Inno Setup 6 installer (setup loader revision 2), so the shipped exe's own header was not read
igneum-miner.exe (box fefd266c..., 10,994,688 B) DIFFER MATCH the same

Two non-determinisms found on the way, both in the SHIPPED builds too (first run 19:43Z, passes A and B differed on all four):

Class Fact Fix
OUT_DIR path in the binary prost's generated protowire.rs (kaspa-grpc-core, kaspa-p2p-lib) embeds its OUT_DIR path; a pass in a target dir of another NAME differs (igneum-miner matched byte for byte once the path was the same) one target path per target in the repro; for cross-machine identity a --remap-path-prefix of the target dir and the home (not done: the Mac and the box differ in every path anyway)
Build clock in the binary libmimalloc-sys compiles mimalloc's C with __DATE__ and __TIME__ ("Oct 6 2026", "21:38:15" sat in libmimalloc.a, next to the mimalloc option names); two builds a minute apart differ SOURCE_DATE_EPOCH exported for every pass (GCC and clang take the date and time from it); PROPOSED for build-remote.sh, cross-remote.sh, cross-build.sh and the PC job: export it from the commit time so two builds of one commit give one hash. The earlier "byte-identical across three builds" on the box was under sccache, which returns the first build's object and hides this class

0.3.15 as well (run 19:56 to 20:00Z, the moment it reached dl/public; node 713ef876, app 563485b; four clean passes of 63 to 76 s): A vs B MATCH on all four again (igneumd 1f1b6eee..., igneum-miner a34e0a56..., igneumd.exe 9b377455..., igneum-miner.exe 65b30edd...). Against the shipped Linux pair in igneum-hive-0.3.15.tar.gz (igneumd 1e51bfb6..., igneum-miner c5b48910...): DIFFER, and the binaries say why: the shipped pair needs GLIBC_2.34 and embeds /Users/joshm/.cargo/registry and the clock string 11:05:57, so the HiveOS package carries the Mac's zig build, not the box's 06211d55... of 19:00Z (the box build needs GLIBC_2.39, which HiveOS cannot run; the zig lane is right for that package). The Windows exes have no shipped hash yet (the public installer is still 0.3.14). docs/evidence/reproduced/0.3.15.md.

Re-run 20:13 to 20:22Z with sccache REALLY off (the build-server agent's finding: an empty RUSTC_WRAPPER= is read by cargo as unset and falls back to the box's config, so the first runs' passes could take hits; now RUSTC_WRAPPER=/usr/bin/env, a pass-through, and SOURCE_DATE_EPOCH = the author time of lib.sh bs_sde): both versions, all four artefacts, A vs B MATCH with the same hashes as above (0.3.14: 03f35e05, 900c1f0b, 166e604e, fefd266c; 0.3.15: 1f1b6eee, a34e0a56, 9b377455, 65b30edd); passes of 73 to 177 s with the two repros side by side on the two slots. The MATCH rows stand on their own.

What it means: the box is deterministic for a given commit and path, so a release built on it can be checked by anyone with the same toolchain by rebuilding and comparing; the shipped 0.3.14 and 0.3.15 Linux bytes came from the Mac's zig lane with a build clock inside and cannot be reproduced anywhere, and the Windows 0.3.14 exes carried a PE timestamp. The reason to ship every target from the box from 0.3.16 (R2) is this table, with SOURCE_DATE_EPOCH exported in every build script; for HiveOS (glibc 2.36 and under) the box needs zig + cargo-zigbuild first (section 6, row 1), or the package keeps the Mac's build and stays unreproducible until then. Open: a --reuse re-report of 0.3.15 once its Windows installer is public, and rebuild-release.sh 0.3.16 the moment it ships.

7.4 The CPU prover trial (DONE 19:30Z; verdict: the box is NOT a prover)

infra/build-server/prover/cpu-trial.sh on the box under the measure hold (builds excluded), igneum-prove-host from master 7483fb37 (the 0.3.15 prover pair, sha256 71bc2438...; its cuda feature changes nothing under SP1_PROVER=cpu), fixture proving/fixtures/block-56-transfers.json (the v0 block: 3 transfers, 600 pgas, one shard, 556,369 SP1 cycles), --mode shard --shard 0, RAYON_NUM_THREADS=96, nice 19, box otherwise idle (load 2.6 at start). Log and results JSON: /srv/builds/_log/prover-trial/trial-20261006T192824Z.{log,json,txt}; JSONL line kind measure.

Stage Box, 96 threads (EPYC 9454P) Mac, same statement class (bench-log)
setup (prover client, shard and aggregator keys) 14.4 s (client 12.6, keys 1.8) 9.3 s per invocation in the v0 loop (4 Oct, "nearly all SP1 setup")
execute 0.28 s, 556,369 cycles, 927 cycles per pgas 0.19 s for block 78 (3 Oct)
core proof 34.2 s, 7,317,561 B, verify 0.34 s 83 s for the 200-pgas shard on a Mac at load 40 (4 Oct); 71.7 s is main's Mac figure for this fixture (its stage not recorded here, labelled approximate)
compressed proof (what a record carries) 85.9 s, 1,272,897 B, verify 0.07 s 272 s for the 200-pgas shard on the loaded Mac (4 Oct); 61 s for the smallest shard in the 3-node v0 loop
whole run, wall 136.9 s
peak RSS 28.2 GB (VmHWM)
CPU use 64 of 96 threads busy on average (6,408 percent in ps)

What the numbers mean, per tier, and what follows:

Number Means Done or proposed
core 34.2 s under 60 s, compressed 85.9 s over it the proof a record carries is the compressed one, so the shard that matters takes 120 s of proving on 96 CPU threads for the SMALLEST shard the chain has (600 pgas, 0.56 M cycles); a shard at S_p is 60 M cycles (bench-log 4 Oct), about 100x, so hours per shard on this CPU against the 60 s proof lag the litepaper states the 60 s test of the ask is NOT met on the stage that counts; NO standing CPU prover unit is written, nothing joins the devnet from the box (R6 holds: the box is never a node host for the live devnet)
28.2 GB peak for the smallest shard a CPU prover needs 32 GB of RAM for a toy shard; every home tier (8, 12, 16, 24 or 32 GB CARDS, 16 to 64 GB of RAM) is out of CPU proving, and the rented 4090 boxes' CPUs are not a fallback either the app keeps "proving on the CPU (slow)" as a correctness lane only; the prover tiers are the real cards (docs/analysis, 6 Oct rented-card measurement)
64 of 96 threads busy SP1's CPU prover does not scale to the whole box; a second trial with RAYON_NUM_THREADS=48 would show whether half the box proves as fast (then two shards side by side) not run tonight (one slot of the box's evening); the script takes --threads
14.4 s setup per invocation the same per-process cost the Mac pays; a resident prover would pay it once already the design of the app's prover loop

The box stays a build and test machine. If main wants a CPU prover anyway for coverage (a prover that is always on, never fast), the shape is a systemd unit as build with SP1_PROVER=cpu, a throwaway devnet key (never the OTA key, never a hand's key), --threads 48, Nice 19 and the measure hold taken for the whole run, which would exclude builds for minutes at a time: that is why it is not written.

8. The capacity layer (6 October 2026, the project lead: "get the builder spun up to capacity")

The box sits idle most of the minute between builds. infra/build-server/capacity/ is a background workload layer that uses the idle cores and yields to builds. It runs as igneum-capacity.service (user build, Nice=19, chrt -i 0 SCHED_IDLE, IOSchedulingClass=idle, CPUQuota=8800% so 8 of the 96 threads are always free) and is a controller (run.sh) over a queue of jobs in jobs/. All work lives under /srv/capacity; the layer takes NO build slot and writes NOTHING under /srv/builds/<worktree>.

Rules (the layer obeys these; they are why a build never waits on it)

Rule How
Background work never takes a build slot the jobs run cargo directly in /srv/capacity and never call remote-run.sh; the controller's cap_build_active only READS /srv/builds/_locks with flock -n
It pauses whenever a build slot is taken or the measure hold exists run.sh polls /srv/builds/_locks every 5 s; it SIGSTOPs the running job's whole process group the instant any build-<k> or measure is held and SIGCONTs it when they clear; it does not even start a new slice while a build runs
It never touches /srv/builds/<worktree> trees its own checkout is /srv/capacity/src (repo) and /srv/capacity/src/vendor/igneum-node (fork), its targets are under /srv/capacity; the only /srv/builds access is a READ of an existing node binary and a READ of the lock files
It is killed by the night battery's start and restarted after igneum-night-battery.service has ExecStartPre=+-systemctl stop igneum-capacity.service and ExecStopPost=+-systemctl start igneum-capacity.service (the + runs as root; the - never fails the battery)

The jobs, in priority order (CAP_SEQUENCE gives the earlier ones more turns)

# Job What Out Dry-run / smoke
1 pow-fuzz continuous igneum-pow mixer and scratch fuzz; IGNEUM_FUZZ_SEED_BASE advances from a cursor so every slice walks fresh programs; counts programs and units, saves any mismatch with its seed base /srv/capacity/out/pow-fuzz/<date>/ --dry-run / --smoke (N 100, 10 min)
2 sync-fuzz the sync-request gate for the 28 unwrap sites (docs/analysis/horizon/consensus-security.md s5): a throwaway pruned igneumd on a simnet datadir on the box (net igneum-devnet-315, loopback only, NEVER the live devnet), fed random, boundary and below-retention locator/header/antipast/IBD/pruning-point requests by igneum-p2p-probe sync-fuzz; a gRPC alive check every 25 requests; any panic saved with the trace /srv/capacity/out/sync-fuzz/<date>/ --dry-run (builds one request of each kind) / --smoke (6 min of requests)
3 sim-sweeps GHOSTDAG (ghostdag_sim.py --seed-base), finality (finality_horizon.py, finality_v2.py --quick) and difficulty attacks (attacks.py --seed-base) across seeds 1 to 1,000; CSV appended; a daily note of any bound that moved /srv/capacity/out/sim-sweeps/<date>/sweeps.csv --dry-run / --smoke (one seed, quick)
4 model-sweeps the N-ladder and chip model (sim/horizon/algorithm/model.py) over every section, cached by the file's content hash /srv/capacity/out/model-sweeps/<date>/ --dry-run / --smoke
5 clippy-audit cargo clippy and cargo audit on every branch pushed to the box repo mirror in the last day, in its own checkout /srv/capacity/clippy, recorded per branch and commit so a slice only picks up new pushes /srv/capacity/out/clippy-audit/<date>/<branch>/ --dry-run / --smoke (one crate, newest branch)

Each job is a script in jobs/ with a --dry-run mode (no build, no node, no run) and a --smoke mode (a ~10 minute bounded run). Every job writes one line per slice into /srv/workers/capacity.json (summary.mjs, atomic), which the worker dashboard's collector (tools/workers/collect.mjs) reads into doc.background; the page (tools/workers/page) shows a "Background" lane from it.

Install

infra/build-server/capacity/install.sh (idempotent): copies the scripts and the unit, pushes the capacity-probe fork branch (the sync-fuzz subcommand) and the box-capacity repo branch to the box mirrors, and enables the service. The layer tracks the repo branch master; until this work merges, the mirror's master lacks the capacity tree, so install writes a drop-in pinning CAP_REPO_BRANCH=box-capacity and removes it once master carries infra/build-server/capacity/run.sh (self-healing after the merge). --no-start enables without starting; --smoke <job> installs then runs one job's 10-minute smoke and prints the summary.

7. Spill-over between the boxes (7 October 2026, 15:0x UK)

the project lead's reading at 15:02 UK: build-1 at load 139 / 114 / 90 with both slots held and a queue (the horizon lane waited 1 h 40 min) while build-2 read 4.5 / 27 / 46 with both slots free. The class router (section 6) pinned each class to its box with no spill-over. Now (lib.sh bs_route_spill, master from this commit):

  • The class is a PREFERENCE: a build, check or gate prefers box 1, a suite, bench or attack row box 2, a proving crate box 3.
  • Before a run, the preferred box is read with one ssh (free slots of its slot count, 1-minute load). It takes the job when it has a free slot and its load is at or under 64 (BS_SPILL_LOAD). Otherwise the other box (1 and 2 swap; 3 falls to 1) takes it when THAT one qualifies; when neither does, the job queues on its own box. A box without a host file is never chosen; an unreachable box reads as "down" and is skipped.
  • The decision is the first route line of the run ("route: class suite prefers box 2; box 2 (free=0 slots=3 load1=70) is full or over load 64: spilled to box 1 (free=1 slots=2 load1=20)"), the slot label carries "; spilled from box N", and the JSONL row carries "route": {"preferred", "box", "spilled", "reason"} for the dashboard's job card.
  • build-2 has THREE slots (its slots file reads 3 since 14:0x UK; provision.sh defaults a -2 hostname to 3). Everything that lands on box 2 runs at the bounded class, nice 10 on a 32-core band with -j 32, builds and gates included, so a suite beside them keeps its number; since this commit a bounded run takes the band its slot owns (slot 0 the last 32 cores, slot 1 the 32 below, slot 2 the 32 below that), so three bounded runs never share a core. A gate still takes its slot ahead of queued suites.
  • --box N still pins. Self-test: tools/ci/route-spill-check.sh (thirteen cases through BS_ROUTE_STATE_<n>, no ssh), in the gate.