docs/plans/miner-faults.md: MF-1 to MF-7, each with its rule, test and gate line.
- MF-1/MF-2: a worker starts and is judged only when the node is READY (synced and igneum_getExecStatus reports an
executed tip; execrpc::probe every 5 s off the engine thread); the node watchdog never counts the catch-up (settled
once read synced; 30 min cap before that; any RPC answer is a sign of life); the watchdog restarts on a ladder 10 s,
30 s, 2 min, 5 min, then every 5 min for ever (watchdog::RETRY_LADDER_S); the faulted state and the one-restart
budget are gone (tools/ci/permanent-fault-check.sh in the gate); a node-caused restart resets the ladder at sync.
- MF-3: the hot-plug pass starts a recovered or revived card's worker (unchanged rule, now in the register).
- MF-4: the status clock starts at ready (program loaded), loading bounded by 300 s; a self-test failure holds the
card 30 min with the reason on its row, released on a driver change; a crash loop climbs the ladder; the pack is
exported once a minute for every card (a refused pack forces one).
- MF-5: the app reads template_wait=, template_ms=, identities_active= from the 0.3.20 miner's STATUS; waiting on
the node is never the card's fault; the row says node slow; every node-wait label clears on the first rate.
- MF-6: a miners hold belongs to the job that took it and releases when that job is gone or at its own cap.
- MF-7: the engine owns every igneum-miner it started: an untracked one on this engine's node RPC is killed at start,
after every stop and every minute, one line and one fault report per kill; a restart kills the old process first.
- Every fault line posts one FAULT line to the log intake (label fault-<id8>, app and node version, 60/h cap).
- The signed cards job kind (per card enabled, identities, power_pct; refused for a card the machine lacks; applied
through the app's own card path, persisted, read back): packaging/ota/publish-jobs.sh add --kind cards.
- LG-4 as a job: relay/playbooks/first-share.ps1 and tools/fleet/first-share-gate.mjs (no Windows box yet).
- tools/reliability: the fault injector with one step per class (catch-up, card-appears, own-restart, zero-ladder,
no-status, node-silent, one-card-fails, orphan-miner); fake-worker.mjs lists devices and fails self-tests on command.
- master's build tooling (97255a4e) and release-0.3.20's igneum-pow taken into the worktree for the box routes.
Box: app 198 + 27 + 8 tests green on igneum-build-2; the tree gate green (33 checks).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
NVIDIA GeForce Game Ready 617.42 WHQL (6 October 2026): us.download.nvidia.com/Windows/617.42/..., 990,853,168
bytes, sha256 f115c927..., subject CN=NVIDIA Corporation (DigiCert G4). AMD Software Adrenalin 26.9.2 WHQL
(29 September 2026, the win11-b build): drivers.amd.com/drivers/whql-amd-software-adrenalin-edition-26.9.2-win11-b.exe
(an amd.com Referer required), 1,000,800,840 bytes, sha256 593c1d73..., subject CN=Advanced Micro Devices (Sectigo).
Read on the box: curl, sha256sum, the PKCS7 out of the PE security directory through openssl pkcs7 -print_certs.
The table's placeholders are gone: every row installs. The drivertable test sample, the mock and the view test name
the real NVIDIA release. detect.rs:990 carried a #[test] above the doc comment of the Intel test from the cherry-pick,
the test build's one warning ("duplicated attribute"): removed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 80787ea024)
the project lead: can the drivers be packaged with the miner, for all cards, with the system knowing which to install if not
present. Not bundled: detected and installed on one click. A per-vendor table rides the signed manifest (drivers.json:
min_version, the version on offer, the vendor's URL, size, sha256 and its source page, the silent arguments, the
restart exit codes, the Authenticode signer, the Linux and HiveOS package), validated by the signer and the app alike
(src/drivertable.rs, shared), written to <app data>/drivers.json by ota.rs. Each card's driver version is read at
every detection (nvidia-smi's driver_version; Windows' DriverVersion for AMD and Intel) and compared; a missing or
old driver puts the offer on the card's row, on the dashboard and on the first-run list. The click downloads with
curl (resume), checks size, sha256 and the Authenticode subject, runs the installer through one elevated prompt
(platform::elevated_command, the PC 1 driver job's shape), reports restart required with a Restart now button, and
never restarts by itself; the miners keep mining. macOS: no step; Linux and HiveOS: the package line. Dry run through
IGNEUM_DRIVER_DRY_RUN or the table. Tests: the table, the versions, the offers per tier, the exit codes, the
Authenticode verdicts, the download against a mocked vendor server on 127.0.0.1, the UI's strip per state; the mock's
drivers scenarios; captures light and dark in docs/plans/driver-check-shots. publish-manifest.sh --drivers carries the
table. Also the doubled #[test] in detect.rs from the cherry-pick.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 3f0ef65786)
The way to hold a card out of mining past a job without a script on api/cards (the 6 October rule): the runner's hold
is built with enabled=false (identities and cap kept), the report line says LEFT OFF, publish-jobs.sh carries
--cards-leave-off. For the Arc B580 on PC 1 while its worker fix rides to the shipped app. Test:
cards_leave_off_restores_the_card_as_off_with_its_settings_kept (igneum-app 157 of 157 on the box).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 9e794503d2e9689ae3a0996e703cb97ea6d95cef)
publish.mjs packs the fifteen served files with one fixed mtime (reproducible; the self-test checks it), hashes, signs
the entry through igneum-ota-sign sign-ui with the key in ~/.config/igneum (never read or printed here), copies the
bundle into the folder's ui/ and hands ui.json to publish-manifest.sh --ui, the one writer of the signed manifest,
which verifies the entry and the bundle's hash before signing; --dry-run writes nothing, --verify reads the live
manifest back against dl/<token>/ui and dl/public/ui; --no-ui withdraws the channel. Both self-tests sit on the one
gate. docs/plans/ui-ota.md: the shape, the engine, the security notes, the operator recipe, the tests, per tier.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit e02f5e14d0)
The app's install clears the jobs folder on a PC, so a run job whose kit was fetched by an earlier fetch job finds
nothing after an update and fails in seconds (5 October 2026, 21:49Z, the AMD kit; bench-log 4df339f). Rule: a run
playbook that reaches a path under the jobs folder other than its own tests the kit is there before its first use,
and the fetch is republished under a new id after any app update.
tools/ci/kit-path-check.sh reads every *.ps1 under relay/playbooks/ and tools/. A kit root is a path derived from
the jobs folder (`$jobs = Split-Path $env:IGNEUM_JOB_DIR` then `Join-Path $jobs '<fetch id>'`, the race-5090.ps1
shape) or one carrying a literal `jobs\` (the amd-card-test.ps1 shape); every path built from it belongs to that kit.
A presence check (Test-Path, [IO.File]::Exists, [IO.Directory]::Exists, Get-Item or Get-ChildItem with -ErrorAction)
on the root or anything under it covers the whole kit. A use before that line fails with "kit path used before a
presence check: republish the fetch after any app update", as does a literal jobs\ path in a command with no check.
The job's own folder ($env:IGNEUM_JOB_DIR) is not a kit path.
Fixtures: kit-path-ok.ps1 (both shapes, checked; a sibling pack file covered by the worker's check) and
kit-path-unchecked.ps1 (the worker run before its check, a literal never checked); --self-test asserts the lines.
Wired into ci.yml after the bash-body step, and into publish-jobs.sh add --kind run beside the other two checks;
test-publish-jobs.sh gains the refusal (34 passed, 0 failed). The current tree: race-5090.ps1 is the one playbook
with a kit, checked before use. README-ship.md: the rule.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A PC job is published from a worktree by packaging/ota/publish-jobs.sh and never passes CI before it runs; tonight
the root-socket fault came back from a job on a branch without the check. `add --kind run` now runs, on the script
being published and before anything is signed: tools/ci/bash-body-check.sh for a PowerShell script (every inline
bash body parses; a body it cannot read fails, never skips), `bash -n` for a .sh script, and
tools/ci/prover-socket-check.sh for both (a root prover run kills sp1-gpu-server and unlinks its socket). A failure
refuses the publish with the check's output; a missing check file refuses too. Kinds without a script (fetch,
collect, restart, update-now, shard-benchmark, build) are untouched.
tools/ci/prover-socket-check.sh is copied from proving-v1 (344cba8; master lacks it) with two additions: file
arguments check those files only (the publisher's call), and an allow list for packaging/ota/test-publish-jobs.sh,
which carries a known-bad root prover script on purpose. Its ci.yml step is left to proving-v1 to avoid a duplicate.
packaging/ota/test-publish-jobs.sh: four refusals (a lost quote in a PowerShell bash body, an unreadable body, a
.sh with a lost quote, a root prover script without the cleanup) and the envelope unchanged after a refusal.
32 passed, 0 failed on this Mac with the main checkout's signer. packaging/README-ship.md: the publish-time gate.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Relay (X23, X27): the intake key is its own tier (upload and file drops only, RELAY_INTAKE_COMPAT=0 closes it);
a run task needs an Ed25519 signature by the Mac run key over {to, nonce, body sha256, flags} (RELAY_RUN_PUB,
401 without) and an HMAC tag with the target's machine secret that the agent verifies before anything runs;
results and registration are bound to the machine the secret proves (403 on a forged from).
X24: every client and Mac tool sends x-relay-token as a header to /api/relay?fn=; the path token stays for the
phone page only. X25: the agent arms the logon task only for a restart a task asked for and disarms on start
and exit. X26: 30-day retention with blob deletion, feed capped at 100, the dl base as RELAY_DL_BASE held by the
agent, never in a body. X28: GET inbox never acks (POST inbox does), RELAY-REBOOT on its own line and only with a
reboot flag, 120/min and 10 failed auths/min per IP, no username or folder on register, WSL sudo scoped to
apt-get and dpkg with SETENV, no password on a command line. X29: the intake key reaches curl through -K in
upload.sh and both upload-log.bat; tools/ci/curl-header-check.sh fails the class. G14: TZ=UTC in ship-app.mjs
and publish-jobs.sh; tools/ci/commit-tz-check.sh fails the class; history-rewrite.md names the .old-2026-10-05
files as the values in the history. The handler moved to relay/lib/handler.mjs with injected sql and blobs
(relay/lib/blob.mjs holds @vercel/blob) so relay/test/handler.test.mjs drives it without a database:
47 tests across 6 suites, all green.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Downloads: packaging/ota/publish-public.sh publishes the current installers, the HiveOS package and the two signed
manifests into dl/public/ with no token in any URL, writes the four /public/ aliases as vercel.json rewrites and an
unsigned index for the site; publish-manifest.sh --public and ship-app.mjs --public run it on every release (dry run
and self-test cover it). Nothing removed from the token folders.
Site: the miner and wallet buttons link the public aliases and show the version and size from the index, read at
build time (site/downloads.json is the offline snapshot); TESTNET_OPEN in build.mjs drops the "Public testnet: not yet
open" line on the go; the HiveOS Flight Sheet install line on the miner page; /faucet page.
HiveOS: igneum-hive-0.3.8.tar.gz from the 0.3.8 node (2b6d23ef, PC build job) and the zig-built Linux workers.
Faucet: site/api/faucet.mjs (10 IGN per address and per IP per day, Neon table faucet_grants, EIP-1559 transfer signed
by site/lib/eth.mjs with no dependencies: keccak, RLP, secp256k1 with RFC 6979), FAUCET_KEY and FAUCET_RPC from the
Vercel env only; 15 unit tests with a fake database and node, run in CI.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The 13:19:41Z refusal on PC 2: fetch_jobs took igneum-jobs.json and .sig in two requests while the edge was
still serving the previous deployment for one of them. The signer wraps the verified pair into one object and
reads it back; the app fetches that object (the pair only when none is published); publish-jobs.sh writes and
mirrors all three files and verifies every folder after the deploy; tools/jobs.mjs reads the envelope.
Tests: jobs.rs signed_envelope_binds_file_and_signature, packaging/ota/test-publish-jobs.sh (24 checks).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
publish-jobs.sh --deploy POSTs the new stamp (published_at plus 8 hex of the file's sha256) and the added id to the
relay's /wake once the live file verifies. The relay token goes in a 600-mode header file, never on the command line
or the screen. Prints "woke the apps (stamp ...)" or a one-line warning; the apps' 2-minute poll still catches it.
tools/jobs.mjs status reads relay_wake (one row per publish with the ids it added) and prints "woken +N s after the
publish" for a machine's latest job that a publish added; nothing when the table does not exist yet.
docs/plans/release-0.3.6.md: "Instant jobs" section with the design, the expected latency and a TODO row per machine
for the measured number once 0.3.6 is live. packaging/ota/README.md: the 10-minute poll is history.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One notice at a time, the most important first (app/igneum-app/ui/app.js, Notices; pure, unit tested):
0 update installing, urgent or failed 1 job failed 2 clock 3 job running
4 update available, downloading, ready, waiting for permission, manual 5 job done 6 updated
Lower notices wait their turn. Every notice has a close control; closing hides that notice's key until the
state moves on (a new status, version or job id).
States and their rules:
update available "Igneum Miner X is available." Install now, Later. Key update:X:pending.
update downloading "Downloading Igneum Miner X: 43%." (no percent when unknown), progress bar; same key as
available and checking, so Later hides the whole download until it is ready.
update checking "Checking Igneum Miner X." (the engine's staging step).
update ready "Igneum Miner X is ready. It installs by itself at a quiet moment." (auto on) or just
"... is ready." Install now, Later. Key update:X:ready.
update waiting Windows, nobody answered the administrator prompt: "... is waiting for permission. It
installs the next time someone is at this PC. Mining continues."
update manual "... is downloaded. Open it and drag the app over the old one." Open the download.
update installing "Installing Igneum Miner X. The app restarts itself. Mining continues until then."
(on a Mac, where the engine quits at once: "The app restarts itself in a moment.")
Also while the engine says "installing now" after Install now.
update urgent the engine's consensus-deadline text, ember, downloading percent when it downloads.
update failed "The update to X failed." plus one line of cause and Try again; rolled back:
"Igneum Miner X did not stay up and was rolled back." A dev build with no manifest
configured shows nothing (Settings still says it).
updated "Updated to Igneum Miner X from Y." Gone 60 s after the new version started.
job running "Job: <title> running, N min. <Stage>." with the last RESULT line underneath.
job done "Job: <title> done after N min. Report uploaded." Gone after 5 minutes.
job failed "Job: <title> failed after N min, exit C. Report not uploaded." plus the first error
line (BUILD FAILED / error / failed / panic among the result lines, else the summary).
Stays until closed. Timeout and aborted are "hit its time cap" and "was stopped".
clock as before: the engine's words, Sync clock, the manual hint; on the setup screens only
(the node card carries it on the dashboard). Jobs show on the dashboard only.
Layout: the strip reserves no height while empty; when a notice appears or goes, main's top moves once with a
150 ms transition (none under prefers-reduced-motion). Existing tokens only, nothing newer than 2022 CSS.
Screenshots: ?update=<state> as before, ?job=running|done|failed added (packaging/ota/README.md).
Test: node --test app/igneum-app/ui/notices.test.mjs (ordering, dismissed keys, wording, the 5-minute and
60-second timers); added to the CI site job. Built once with cargo (include_str) and checked against the
ui-mock scenarios and a scratch engine instance (IGNEUM_APP_DATA in a temp dir, fake worker).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Conflict resolved: publish-jobs.sh keeps master's verify command and --tries (the retrying live check) alongside the
build kind's arguments; the usage range covers the merged header.
Conflicts resolved: state.rs keeps both the sweep fields (miner-eff) and the race fields (miner-perf); bench-log.md
keeps both appended entries; publish-manifest.sh keeps master's --override implementation (8082576, the "every
height switch" rule, --verify-only, --tries, the retrying live check) and adds miner-perf's --tuning / --no-tuning
with the carry-over of consensus.override and tuning from the current manifest. One --override case, one parser.
New job kind `build` (jobs.rs, jobbuild.rs, jobrun.rs run_build): free-space check on both sides (20 GB), the
build-inputs zip by sha256, setup inside the distro as root (mingw-w64 posix, clang for bindgen, protoc, zstd, the
Windows rust target; idempotent), sources extracted with the target dir persisting under /root/igneum-build,
cargo build --release native and for x86_64-pc-windows-gnu, cargo test for the manifest's packages, binaries
zstd-compressed and sent to the relay (fn=upload, Blob PUT, fn=drop; 50 MB each) with sha256 in RESULT lines,
STAGE lines with UTC times, a 40-minute default budget and per-stage caps, the Linux side killed on a cap. The
app's runner stays serial (one Active at a time), so a build never overlaps a shard job; nothing stops the miners.
From this version an unknown job kind is skipped by the app (parse_lenient) instead of rejecting the whole file;
the signer stays strict.
Mac side: packaging/windows/push-build-inputs.sh packs a fork worktree, app/igneum-app, brand/icons and
proto-cuda with a manifest (branch, commit, dirty, builds, tests) and the sha256; publish-jobs.sh add --kind build;
tools/build-job.mjs packs, publishes, watches, fetches, checks both sha256 per file and the PE header of every exe
(plus verify-exe.py on igneum-app.exe), and places the binaries where push-inputs.sh, make-payload.sh and the
cloud-devnet scripts look. relay.mjs drop <file> --body carries the body.
Tested on the Mac: 33 app tests (6 new) and the signer's 21; cargo check for x86_64-pc-windows-gnu; the packer
(7.9 MB zip, no target dirs); the publisher against a scratch folder with the rebuilt signer, the old signer
refusing the kind, a bad job refused at signing; the fetch path against the live relay with a real exe (sha256
and PE pass, a wrong sha256 refused; test items deleted). Not run on a PC: the job itself. docs/plans/build-job.md
has the first job for PC 1 and the rollout order (0.3.4 must be on the PCs before a build job is published).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The engine turns a worker's race line into one TUNING {json} line in the app log (card model as the worker names it,
driver, arch, program class loads and wide loads, every variant's MH/s, winner, gain, the card's power cap and
draw, MH per watt), which the existing intake receives; the card state carries the variant for the dashboard and
one event per race. tools/tuning.mjs aggregates the records from miner_logs per card model (median MH/s or MH per
watt, at least 3 samples, de-duplicated per race) and writes tuning.json; publish-manifest.sh --tuning puts it in
the signed manifest (and now takes --override for consensus.override; both are carried over from the current
manifest when not given, --no-tuning drops it); manifest.rs parses it; ota.rs writes <app data>/tuning.json and
removes it when the manifest drops it; procs::spawn takes an environment and every miner starts with
IGNEUM_TUNING_FILE, which its worker reads at every prepare. Dry run of the publisher against a scratch folder:
tuning and override written, carried over, dropped, signature verified.
docs/plans/miner-perf.md: the signed jobs for PC 1 (fetch the race build of the NVRTC worker, then
relay/playbooks/race-5090.ps1 with the miners stopped: 17 variants, 3 rounds, twice) with the exact publish
commands for the main session; not published by the agent.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
PC 2 on 0.3.3 logged 'needs wsl-prover' although the toolchain is there as user [user]. Every negative probe now
lands in the engine log with its exit code and the first 200 characters of output (it reaches the intake). The
probe sources ~/.cargo/env and sets ~/.cargo/bin and ~/.sp1/bin itself (a login shell from a console-less process
need not), runs as the job's wsl_user or IGNEUM_APP_WSL_USER first and the distro default user second, and a
payload with wsl2\bin\igneum-prove-host next to the app meets the requirement when the distro answers.
publish-jobs.sh: --requires none or "" publishes an empty list (none was a literal requirement, "" fell back
to the kind's default).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
relay/api/console.mjs reads the log intake (miner_logs) and console_items in Neon, the OTA manifest, the
CI json and the jobs file from the downloads host (DL_TOKEN in the project env, never in the client), and
igneum.network/api/live; 10 s cache per answer. tools/console.mjs: post --kind log|build|note, log, machines,
chain, jobs, builds, results, sync-bench, sync-dl, sync-hetzner, sync, url. The two Mac-side build scripts
post build events. Screenshots at 375 px and desktop in docs/design/console/.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>