Commit graph

15 commits

Author SHA1 Message Date
igneum-labs
2e05cf3b24 Merge commit '58cc162' into release-0.3.17
# Conflicts:
#	.github/workflows/ci.yml
2026-10-06 23:10:56 +00:00
igneum-labs
984727d828 Merge release-0.3.12 (095aa9e) into ember-tune: 0.3.11's six-section View and card order kept, Ember Tune's line and switches re-added on it; the tune fields move into hotplug::apply_pref; the power-cap plan keeps present(); both CI test lists; 132 app tests, 26 UI tests, every gate green
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:20:46 +00:00
igneum-labs
735887a309 Ember Tune: every card tuned for MH per watt out of the box, the fleet prior per card model in the signed manifest, the console and /miners priors table
the project lead, 5 October 2026, 22:45 BST: "make sure we have ember tuning every single card for efficiency out of the box, the
more data = the better the tune, make an awesome system." Built on lever 3 (docs/plans/miner-eff.md), lever 2's signed
tuning section (docs/design/miner-tuning.md), the AMD telemetry helper (35e3d26, its --tune/--set-gmax/--set-plimit/
--reset contract) and the Power control switch (652e848). Design, data flow, tiers and the privacy line:
docs/plans/ember-tune.md.

- src/ember.rs (new): two knobs per card (power limit %, core clock cap MHz; memory clock never touched), the full plan
  (power ladder 100..50%, then the clock ladder 90..60% at the chosen power), the confirm plan (the fleet prior and one
  neighbour), the baseline plan (measure only), the marks (faulted, hot, memory_clock_dropped, unapplied, no_readings),
  the choice (best MH/W within 1% of the top rate, then rate, then draw), the fleet record (a hash of the install id,
  no address), the prior lookup and the kill switch (tuning.ember), the state machine on a fake clock. 9 unit tests.
- engine.rs: tick_sweep schedules every NVIDIA, AMD and Apple card (120 s steady, 600 s to the boundary, no job hold,
  no pause, weekly, again after a driver major or program-class change, never under the manifest kill switch); the
  probe (nvidia-smi clocks.max.gr + driver_version and the direct/helper mode; igneum-gpu-telemetry --tune for AMD);
  tune_apply (nvidia-smi -pl / -lgc 0,<MHz> / -rgc directly or through the helper; the AMD helper per request);
  Cmd::TuneProbe, Cmd::TuneSet; faults from rejected and mismatched hashes mark the step; the TUNE lines and the TUNE
  {json} record, uploaded with the log; the Tuned line on the card state. The NVIDIA helper starts only with Power
  control on: the --sweep job never counts as permission (no prompt on a PC with nobody there).
- sweep.rs: the helper protocol gains lgc/rgc (clock cap and reset) and resets the clocks after 20 idle minutes.
- state.rs, config.rs: the tune fields (clock cap, driver, class, source, the Tuned line); the nvidia-smi telemetry
  query carries clocks.gr and clocks.mem; the AMD sample line's plimit_pct and gmax_mhz are parsed.
- ui: "Tuned: X MH/s at Y W (Z MH/W)" with the point, the source and when; measure-only cards say why; the Ember Tune
  switch; tune-line.test.mjs.
- relay/lib/ember.mjs + relay/test/ember.test.mjs: the aggregation per (card model | driver major | program class):
  median point, MH/W, spread, samples, machines; five samples converge, an outlier does not move the median, baselines
  make no prior, de-duplication, the manifest merge keeps lever 2's cards. api/console.mjs fn=tuning and
  tools/console.mjs tuning; tools/tuning.mjs --priors [--write tuning.json] [--site] [--tuning-off].
- site: the fleet priors table on /miners (site/miner-priors.json), the lever text.
- relay/playbooks/ember-tune-pc1.ps1: the PC 1 run (second engine with --sweep from a scratch copy of the install).

Measured tonight: see the bench log entry that follows the PC 1 run. The 9070 XT left PC 1's bus at 20:40 UTC and the
5090 needs the administrator prompt the project lead cannot answer asleep, so tonight's PC 1 run is the baseline plan on the 5090
through the whole pipeline; the two-knob tune on both cards is owed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:25:09 +00:00
igneum-labs
58cc162c4a relay: three auth tiers, signed run tasks, machine secrets, retention; clients on headers; TZ=UTC and curl -K checks (X23 X24 X25 X26 X27 X28 X29 G13 G14)
Relay (X23, X27): the intake key is its own tier (upload and file drops only, RELAY_INTAKE_COMPAT=0 closes it);
a run task needs an Ed25519 signature by the Mac run key over {to, nonce, body sha256, flags} (RELAY_RUN_PUB,
401 without) and an HMAC tag with the target's machine secret that the agent verifies before anything runs;
results and registration are bound to the machine the secret proves (403 on a forged from).
X24: every client and Mac tool sends x-relay-token as a header to /api/relay?fn=; the path token stays for the
phone page only. X25: the agent arms the logon task only for a restart a task asked for and disarms on start
and exit. X26: 30-day retention with blob deletion, feed capped at 100, the dl base as RELAY_DL_BASE held by the
agent, never in a body. X28: GET inbox never acks (POST inbox does), RELAY-REBOOT on its own line and only with a
reboot flag, 120/min and 10 failed auths/min per IP, no username or folder on register, WSL sudo scoped to
apt-get and dpkg with SETENV, no password on a command line. X29: the intake key reaches curl through -K in
upload.sh and both upload-log.bat; tools/ci/curl-header-check.sh fails the class. G14: TZ=UTC in ship-app.mjs
and publish-jobs.sh; tools/ci/commit-tz-check.sh fails the class; history-rewrite.md names the .old-2026-10-05
files as the values in the history. The handler moved to relay/lib/handler.mjs with injected sql and blobs
(relay/lib/blob.mjs holds @vercel/blob) so relay/test/handler.test.mjs drives it without a database:
47 tests across 6 suites, all green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 18:46:13 +00:00
igneum-labs
df9e0a6dcf app: GPU hot-plug (re-detection every minute, WM_DEVICECHANGE on Windows), faulty cards listed with the problem code, integrated GPUs off by default, the card list in the console
5 October 2026: an RX 9070 XT went into PC 1 through a Sonnet USB4 box while the app ran and nothing noticed; the
app detected cards once at start. Now src/hotplug.rs compares every enumeration with the list (key, else vendor +
name when unique; a card whose tool did not answer is never called removed): a new usable card starts a worker,
enabled by default like a card at start, with "New card: <name>, mining" on the strip and in the log; a card
Windows lists with a problem code (Win32_VideoController Status / ConfigManagerErrorCode) is shown as "<name>: not
usable (Code 43)" with the reboot-or-reinstall hint and no worker; a card that disappears has its worker stopped
(quit, 8 s) and its row says removed for five minutes, then hides; an unchanged list touches nothing. The Windows
host sends "detect" on WM_DEVICECHANGE; the engine polls every 60 s (300 s on macOS, no GPU hot-plug there).

detect.rs: the Ryzen iGPU is "gfx1036" to the OpenCL worker, so the APU gfx codes count as integrated, plus the
adapter row's Intel processor string and a dedicated memory under 1 GB; integrated defaults to off with "integrated
GPU, off by default (2 to 3 MH/s for 30 W)" on the row, and the user's choice is kept across re-detections and
restarts (settings, found by key or by vendor + name when the index moved).

Console: the engine logs "cards: <name> [<kind>, <state>] | ..." at start, on every change and every 10 minutes;
relay/lib/parse.mjs reads it and the hot-plug events, the machines API and tools/console.mjs machines show them.

Tests: hotplug.rs (added, removed, moved, errored, recovered, revived, unchanged, twins, user override kept,
the console line), detect.rs (PC 1's adapter lines, the Mac, kind classification, the unusable row),
notices.test.mjs (card notices), relay parse.test.mjs (cards line). cargo test -p igneum-app: 91 passed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 17:45:05 +00:00
igneum-labs
baa109d647 relay: /wake long-poll for the apps' remote jobs (public GET held 45 s, authenticated POST of the stamp)
GET /wake?since=<stamp> is public (the apps hold no token) and rate limited (30 a minute per IP). It holds up to
45 s, re-reading the stamp every 2 s, and answers {stamp, at, added, changed, held_ms} the moment the stored stamp
differs from since, else the unchanged stamp at the deadline. POST /r/<token>/wake {stamp, added} (the relay's
auth, also x-relay-token or x-igneum-key on /wake) records a stamp; one row per stamp in relay_wake, created by the
first POST. maxDuration 60 s for api/wake.mjs in vercel.json. api/relay.mjs is untouched.

The handler lives in lib/wake.mjs with its dependencies injected; relay/test/wake.test.mjs drives it with a fake
database, a fake clock and a fake sleep (the hold, the change, the deadline, the rate limit, the hold cap, auth, a
database error). CI's site job runs it with the other relay tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 08:22:02 +00:00
igneum-labs
df61a116ad Console: the last OTA state per machine and app run is remembered (console_ota_memo) and shown when the upload's 256 KiB tail no longer holds it
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 21:27:59 +00:00
igneum-labs
15e3fccebd Console: the app log's parsed tail is 400 KB so a stuck update's lines stay on the card (PC 1's 'installing' scrolled out of 60 KB in 30 min)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 21:26:05 +00:00
igneum-labs
c98b6a546d Console: a machine whose app logged a clean quit or an update, with no status line after it, shows 'stopped (quit|update) N ago' instead of 'silent' (parseAppTail moved to relay/lib/parse.mjs, test)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:04:18 +00:00
igneum-labs
235823b23d Bug hunt: console cards for other/intel workers and a stale mark on old STATUS lines (relay/lib/parse.mjs + test in CI); publish-jobs verifies the live file with retries and named reasons, a verify command, a failed deploy stops, a collect command without $_ is refused; the dl token masked in printed URLs; docs/bugs.md
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:15:05 +00:00
igneum-labs
df8c469723 Relay: run and task posts need the console token (round 4, X23); prove host saves proofs buffered (ledger P20, second gap)
The intake key sits in every miner package, so the relay now lets it report only (drop text and files, ack, done,
register, upload). Posting a run or task, or renaming and re-roling a machine, needs the console token.
The prove host wrote proofs through SP1's unbuffered save: on WSL2 under /mnt/c the 18 MB core proof of a shard
took longer to save than to prove. Proofs now go through a 4 MB buffer with a timed 'saved' line, and
prove-shard.sh keeps results on the Linux side and copies them per stage. Ledger P20 updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 18:14:10 +00:00
igneum-labs
8ab852ce21 Console: a job is done only on the app's closing report; prove-shard.sh takes every block fixture after the first argument
The console marked any job with a RESULT line as done, so a running shard job read as finished. Done now means
the SUMMARY line carries finished_at or the job's closing 'job <id>: <status> (exit N)' line is present.
prove-shard.sh dropped the third fixture argument (block-344-shards4) because it read only $2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 17:52:35 +00:00
igneum-labs
89dd615fef Console Machines tab: app version, OTA state, power cap, jobs and status lines from the app log (0.3.3 header)
The IGNEUM-APP header is read from the newest upload's first line, from any tail, or from the first upload of
the run; the app's status: line supplies peers, lifetime accepted and synced when the node log tail has none.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 16:50:21 +00:00
igneum-labs
8f08af5318 Igneum console at the relay URL: Machines, Jobs, Builds, Chain, Work log, Results and the relay as tabs
relay/api/console.mjs reads the log intake (miner_logs) and console_items in Neon, the OTA manifest, the
CI json and the jobs file from the downloads host (DL_TOKEN in the project env, never in the client), and
igneum.network/api/live; 10 s cache per answer. tools/console.mjs: post --kind log|build|note, log, machines,
chain, jobs, builds, results, sync-bench, sync-dl, sync-hetzner, sync, url. The two Mac-side build scripts
post build events. Screenshots at 375 px and desktop in docs/design/console/.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 13:53:33 +00:00
igneum-labs
12683c73cd Relay: text, files and runnable tasks between the Mac, the PCs and the phone (relay.igneum.network)
New Vercel project igneum-relay from relay/: one function (api/relay.mjs) over Neon tables relay_items and relay_machines,
files in Vercel Blob store igneum-relay (50 MB client uploads, 4 MB through the function), phone-first web page at /r/<token>/
with the site tokens. Mac CLI tools/relay.mjs (feed, read, drop, task, run, watch, inbox, machines, role, name).
Windows clients send.bat/send.ps1 and the igneum-agent (registers hostname, role, GPUs, WSL, nvcc; runs queued PowerShell
scripts, posts results, reboot-continue via scheduled task + RunOnce), bash twins send.sh and agent.sh (verified live),
playbooks for WSL setup, prover setup, prove-block, miner v4, one-click placeholder. make-clients.sh bakes the secrets
into a zip; the repo copies hold placeholders. Screenshots under docs/design/relay.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 10:04:30 +00:00