Commit graph

11 commits

Author SHA1 Message Date
igneum-labs
5c808b89e0 Ember Tune: every card tuned for MH per watt out of the box, the fleet prior per card model in the signed manifest, the console and /miners priors table
the project lead, 5 October 2026, 22:45 BST: "make sure we have ember tuning every single card for efficiency out of the box, the
more data = the better the tune, make an awesome system." Built on lever 3 (docs/plans/miner-eff.md), lever 2's signed
tuning section (docs/design/miner-tuning.md), the AMD telemetry helper (423936b, its --tune/--set-gmax/--set-plimit/
--reset contract) and the Power control switch (057f0ec). Design, data flow, tiers and the privacy line:
docs/plans/ember-tune.md.

- src/ember.rs (new): two knobs per card (power limit %, core clock cap MHz; memory clock never touched), the full plan
  (power ladder 100..50%, then the clock ladder 90..60% at the chosen power), the confirm plan (the fleet prior and one
  neighbour), the baseline plan (measure only), the marks (faulted, hot, memory_clock_dropped, unapplied, no_readings),
  the choice (best MH/W within 1% of the top rate, then rate, then draw), the fleet record (a hash of the install id,
  no address), the prior lookup and the kill switch (tuning.ember), the state machine on a fake clock. 9 unit tests.
- engine.rs: tick_sweep schedules every NVIDIA, AMD and Apple card (120 s steady, 600 s to the boundary, no job hold,
  no pause, weekly, again after a driver major or program-class change, never under the manifest kill switch); the
  probe (nvidia-smi clocks.max.gr + driver_version and the direct/helper mode; igneum-gpu-telemetry --tune for AMD);
  tune_apply (nvidia-smi -pl / -lgc 0,<MHz> / -rgc directly or through the helper; the AMD helper per request);
  Cmd::TuneProbe, Cmd::TuneSet; faults from rejected and mismatched hashes mark the step; the TUNE lines and the TUNE
  {json} record, uploaded with the log; the Tuned line on the card state. The NVIDIA helper starts only with Power
  control on: the --sweep job never counts as permission (no prompt on a PC with nobody there).
- sweep.rs: the helper protocol gains lgc/rgc (clock cap and reset) and resets the clocks after 20 idle minutes.
- state.rs, config.rs: the tune fields (clock cap, driver, class, source, the Tuned line); the nvidia-smi telemetry
  query carries clocks.gr and clocks.mem; the AMD sample line's plimit_pct and gmax_mhz are parsed.
- ui: "Tuned: X MH/s at Y W (Z MH/W)" with the point, the source and when; measure-only cards say why; the Ember Tune
  switch; tune-line.test.mjs.
- relay/lib/ember.mjs + relay/test/ember.test.mjs: the aggregation per (card model | driver major | program class):
  median point, MH/W, spread, samples, machines; five samples converge, an outlier does not move the median, baselines
  make no prior, de-duplication, the manifest merge keeps lever 2's cards. api/console.mjs fn=tuning and
  tools/console.mjs tuning; tools/tuning.mjs --priors [--write tuning.json] [--site] [--tuning-off].
- site: the fleet priors table on /miners (site/miner-priors.json), the lever text.
- relay/playbooks/ember-tune-pc1.ps1: the PC 1 run (second engine with --sweep from a scratch copy of the install).

Measured tonight: see the bench log entry that follows the PC 1 run. The 9070 XT left PC 1's bus at 20:40 UTC and the
5090 needs the administrator prompt the project lead cannot answer asleep, so tonight's PC 1 run is the baseline plan on the 5090
through the whole pipeline; the two-knob tune on both cards is owed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:25:09 +00:00
igneum-labs
0550f58857 relay: /wake long-poll for the apps' remote jobs (public GET held 45 s, authenticated POST of the stamp)
GET /wake?since=<stamp> is public (the apps hold no token) and rate limited (30 a minute per IP). It holds up to
45 s, re-reading the stamp every 2 s, and answers {stamp, at, added, changed, held_ms} the moment the stored stamp
differs from since, else the unchanged stamp at the deadline. POST /r/<token>/wake {stamp, added} (the relay's
auth, also x-relay-token or x-igneum-key on /wake) records a stamp; one row per stamp in relay_wake, created by the
first POST. maxDuration 60 s for api/wake.mjs in vercel.json. api/relay.mjs is untouched.

The handler lives in lib/wake.mjs with its dependencies injected; relay/test/wake.test.mjs drives it with a fake
database, a fake clock and a fake sleep (the hold, the change, the deadline, the rate limit, the hold cap, auth, a
database error). CI's site job runs it with the other relay tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 08:22:02 +00:00
igneum-labs
88017d69e1 Console: the last OTA state per machine and app run is remembered (console_ota_memo) and shown when the upload's 256 KiB tail no longer holds it
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 21:27:59 +00:00
igneum-labs
697b012954 Console: the app log's parsed tail is 400 KB so a stuck update's lines stay on the card (PC 1's 'installing' scrolled out of 60 KB in 30 min)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 21:26:05 +00:00
igneum-labs
21ccb65f40 Console: a machine whose app logged a clean quit or an update, with no status line after it, shows 'stopped (quit|update) N ago' instead of 'silent' (parseAppTail moved to relay/lib/parse.mjs, test)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:04:18 +00:00
igneum-labs
9915c883b1 Bug hunt: console cards for other/intel workers and a stale mark on old STATUS lines (relay/lib/parse.mjs + test in CI); publish-jobs verifies the live file with retries and named reasons, a verify command, a failed deploy stops, a collect command without $_ is refused; the dl token masked in printed URLs; docs/bugs.md
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:15:05 +00:00
igneum-labs
ec0ede62b9 Relay: run and task posts need the console token (round 4, X23); prove host saves proofs buffered (ledger P20, second gap)
The intake key sits in every miner package, so the relay now lets it report only (drop text and files, ack, done,
register, upload). Posting a run or task, or renaming and re-roling a machine, needs the console token.
The prove host wrote proofs through SP1's unbuffered save: on WSL2 under /mnt/c the 18 MB core proof of a shard
took longer to save than to prove. Proofs now go through a 4 MB buffer with a timed 'saved' line, and
prove-shard.sh keeps results on the Linux side and copies them per stage. Ledger P20 updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 18:14:10 +00:00
igneum-labs
1c1dbd93d6 Console: a job is done only on the app's closing report; prove-shard.sh takes every block fixture after the first argument
The console marked any job with a RESULT line as done, so a running shard job read as finished. Done now means
the SUMMARY line carries finished_at or the job's closing 'job <id>: <status> (exit N)' line is present.
prove-shard.sh dropped the third fixture argument (block-344-shards4) because it read only $2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 17:52:35 +00:00
igneum-labs
57993f79dc Console Machines tab: app version, OTA state, power cap, jobs and status lines from the app log (0.3.3 header)
The IGNEUM-APP header is read from the newest upload's first line, from any tail, or from the first upload of
the run; the app's status: line supplies peers, lifetime accepted and synced when the node log tail has none.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 16:50:21 +00:00
igneum-labs
552d5a30c2 Igneum console at the relay URL: Machines, Jobs, Builds, Chain, Work log, Results and the relay as tabs
relay/api/console.mjs reads the log intake (miner_logs) and console_items in Neon, the OTA manifest, the
CI json and the jobs file from the downloads host (DL_TOKEN in the project env, never in the client), and
igneum.network/api/live; 10 s cache per answer. tools/console.mjs: post --kind log|build|note, log, machines,
chain, jobs, builds, results, sync-bench, sync-dl, sync-hetzner, sync, url. The two Mac-side build scripts
post build events. Screenshots at 375 px and desktop in docs/design/console/.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 13:53:33 +00:00
igneum-labs
53f7f55796 Relay: text, files and runnable tasks between the Mac, the PCs and the phone (relay.igneum.network)
New Vercel project igneum-relay from relay/: one function (api/relay.mjs) over Neon tables relay_items and relay_machines,
files in Vercel Blob store igneum-relay (50 MB client uploads, 4 MB through the function), phone-first web page at /r/<token>/
with the site tokens. Mac CLI tools/relay.mjs (feed, read, drop, task, run, watch, inbox, machines, role, name).
Windows clients send.bat/send.ps1 and the igneum-agent (registers hostname, role, GPUs, WSL, nvcc; runs queued PowerShell
scripts, posts results, reboot-continue via scheduled task + RunOnce), bash twins send.sh and agent.sh (verified live),
playbooks for WSL setup, prover setup, prove-block, miner v4, one-click placeholder. make-clients.sh bakes the secrets
into a zip; the repo copies hold placeholders. Screenshots under docs/design/relay.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 10:04:30 +00:00