Commit graph

16 commits

Author SHA1 Message Date
igneum-labs
847211c7c2 Relay: the X23 tree on master (the signed and tagged run, lib/handler.mjs, guard, the per-machine secret), the start-app kind, the v3 agent as the per-user logon task IgneumRelayService with install-agent.ps1 and the hidden loop, the per-install hostname on registration, the signing tools/relay.mjs so no lane posts an unsigned run (the live relay refuses one since the 7 October 2026 deploy), tools/console.mjs and build-job.mjs on the header tier (X24), the playbooks on RELAY_DL_BASE (X26) and wsl-setup's sudo for apt-get and dpkg only (X28), the new playbooks (start-app, agent-install, pc2-crash-collect, boot-check, freeze-check), and the deploy record with the rollback line (docs/plans/relay-deploy-2026-10-07.md: production igneum-relay-iabqarnby, previous bgy767z40)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 14:47:59 +00:00
igneum-labs
53a4496dd9 Console: the job-channel ping (MF-11, 7 October 2026). Every 0.3.21 app names itself on its wake long-poll (machine, version, last job); /wake records one row per machine in relay_wake_seen before it holds; the Machines tab reads "job channel polled N ago, last job X" and, after 15 minutes without a poll, "job channel silent since <time>, last job X" in red, whatever the uploads say (PC 2 lost power at 10:46Z and nothing said so until a person read the intake). Tests in relay/test/wake.test.mjs.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 14:10:59 +00:00
igneum-labs
229ca3f091 Ember 2: the climb's tests (a synthetic memory-bound card converges in five probes; a refused probe backs the knob off for good; each goal picks its point; the £/day formula) and the fleet prior carrying the memory clock (climb records aggregate)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 15:50:34 +00:00
igneum-labs
563463b35f Merge release-0.3.12 (fda4684) into ember-tune: 0.3.11's six-section View and card order kept, Ember Tune's line and switches re-added on it; the tune fields move into hotplug::apply_pref; the power-cap plan keeps present(); both CI test lists; 132 app tests, 26 UI tests, every gate green
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:20:46 +00:00
igneum-labs
5c808b89e0 Ember Tune: every card tuned for MH per watt out of the box, the fleet prior per card model in the signed manifest, the console and /miners priors table
the project lead, 5 October 2026, 22:45 BST: "make sure we have ember tuning every single card for efficiency out of the box, the
more data = the better the tune, make an awesome system." Built on lever 3 (docs/plans/miner-eff.md), lever 2's signed
tuning section (docs/design/miner-tuning.md), the AMD telemetry helper (423936b, its --tune/--set-gmax/--set-plimit/
--reset contract) and the Power control switch (057f0ec). Design, data flow, tiers and the privacy line:
docs/plans/ember-tune.md.

- src/ember.rs (new): two knobs per card (power limit %, core clock cap MHz; memory clock never touched), the full plan
  (power ladder 100..50%, then the clock ladder 90..60% at the chosen power), the confirm plan (the fleet prior and one
  neighbour), the baseline plan (measure only), the marks (faulted, hot, memory_clock_dropped, unapplied, no_readings),
  the choice (best MH/W within 1% of the top rate, then rate, then draw), the fleet record (a hash of the install id,
  no address), the prior lookup and the kill switch (tuning.ember), the state machine on a fake clock. 9 unit tests.
- engine.rs: tick_sweep schedules every NVIDIA, AMD and Apple card (120 s steady, 600 s to the boundary, no job hold,
  no pause, weekly, again after a driver major or program-class change, never under the manifest kill switch); the
  probe (nvidia-smi clocks.max.gr + driver_version and the direct/helper mode; igneum-gpu-telemetry --tune for AMD);
  tune_apply (nvidia-smi -pl / -lgc 0,<MHz> / -rgc directly or through the helper; the AMD helper per request);
  Cmd::TuneProbe, Cmd::TuneSet; faults from rejected and mismatched hashes mark the step; the TUNE lines and the TUNE
  {json} record, uploaded with the log; the Tuned line on the card state. The NVIDIA helper starts only with Power
  control on: the --sweep job never counts as permission (no prompt on a PC with nobody there).
- sweep.rs: the helper protocol gains lgc/rgc (clock cap and reset) and resets the clocks after 20 idle minutes.
- state.rs, config.rs: the tune fields (clock cap, driver, class, source, the Tuned line); the nvidia-smi telemetry
  query carries clocks.gr and clocks.mem; the AMD sample line's plimit_pct and gmax_mhz are parsed.
- ui: "Tuned: X MH/s at Y W (Z MH/W)" with the point, the source and when; measure-only cards say why; the Ember Tune
  switch; tune-line.test.mjs.
- relay/lib/ember.mjs + relay/test/ember.test.mjs: the aggregation per (card model | driver major | program class):
  median point, MH/W, spread, samples, machines; five samples converge, an outlier does not move the median, baselines
  make no prior, de-duplication, the manifest merge keeps lever 2's cards. api/console.mjs fn=tuning and
  tools/console.mjs tuning; tools/tuning.mjs --priors [--write tuning.json] [--site] [--tuning-off].
- site: the fleet priors table on /miners (site/miner-priors.json), the lever text.
- relay/playbooks/ember-tune-pc1.ps1: the PC 1 run (second engine with --sweep from a scratch copy of the install).

Measured tonight: see the bench log entry that follows the PC 1 run. The 9070 XT left PC 1's bus at 20:40 UTC and the
5090 needs the administrator prompt the project lead cannot answer asleep, so tonight's PC 1 run is the baseline plan on the 5090
through the whole pipeline; the two-knob tune on both cards is owed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:25:09 +00:00
igneum-labs
d5da02a851 Merge pack-loop (4163143) into release-0.3.10: a pack's seed words are its program attempt's words, not the bare seed's (the epoch 34 outage)
# Conflicts:
#	relay/test/parse.test.mjs
2026-10-05 19:23:59 +00:00
igneum-labs
41631431c0 Workers: a pack's seed words are its program attempt's words, not the bare seed's (epoch 34 incident, 5 October 2026)
From 18:23Z both Windows workers (CUDA on PC 1 and PC 2, OpenCL on PC 1 after the 18:34Z node restart)
refused every pack for epoch 34 with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", and
the miner and the app restarted them every 5 to 60 s until 18:44Z and beyond. The packs were correct.
The generator retries a rejected candidate with seed || k_le32 (attempt_words); epoch 34's attempt 0 was
rejected (246 of 16384 final register values saturated, limit 163) and attempt 1 accepted, so the pack
carried attempt 1's words while pf_load (proto-cuda/nvrtc/packfile.h) derived the expected words from
the bare seed. Both workers also matched jobs to pairs by those bare-seed words, so even a loaded pack
of a retried program would have answered "epoch seed mismatch" on every job.

The rule, in one place per language:
- packfile.h: pf_program_words(bytes, attempt); pf_load reads IGNEUM_PROGRAM_ATTEMPT and checks the
  attempt's words; the refusal says "program pack and its seeds disagree: IGNEUM_SEEDW_INIT is not
  attempt N of the epoch seed ..." in plain words.
- worker.cpp and proto-opencl/host.c: a job belongs to a pair when the seed hex the node sent is the
  pair's (pairIs); the compiled-in placeholder pack keeps the word comparison.
- igneum-pow/src/packcheck.rs: verify_pack_texts / verify_pack_dir, the same rule in Rust; the miner
  checks every pack it writes with it before a worker sees it (vendor/igneum-node pack-loop branch).
  Tests pin the attempt vectors of epoch 34 on both sides (one vector, two implementations), that
  epoch 34 is attempt 1 and epoch 33 attempt 0, a known-good pack of a later attempt, a known-mismatched
  (out of date) pack, and self-contradicting packs.
- proto-cuda/nvrtc/emu/packfile-test.c (+ .sh, in CI): pf_load on a known-good attempt-1 pack, the
  checked-in attempt-0 pack, and the known-mismatched bare-words pack.

The app (app/igneum-app):
- watchdog.rs: PACK_OUT_OF_DATE_CODE 44, PackRebuilds (at most 3 pack exports per epoch, then the card
  shows the reason), pack_refusal (the worker's "error 0 pack" line and the miner's "PACK OUT OF DATE"
  line), pack_epoch_of; tests on the incident lines, known-good and known-mismatched.
- engine.rs: exit 44 exports the pack again before the restart instead of a blind restart, the strip
  says "program pack out of date, rebuilding", the card and the log name the condition; at the cap the
  card is marked failed with the reason and tries again in 10 minutes.

The relay (relay/lib/parse.mjs): PACK_MISMATCH; the card reads "pack mismatch, rebuilding (N refusals
in the tail, M restarts)" in `node tools/console.mjs machines` instead of a bare restart count; tests
on the PC 2 tail of 18:27Z and a healthy tail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 19:10:07 +00:00
igneum-labs
d7049fdfda app: GPU hot-plug (re-detection every minute, WM_DEVICECHANGE on Windows), faulty cards listed with the problem code, integrated GPUs off by default, the card list in the console
5 October 2026: an RX 9070 XT went into PC 1 through a Sonnet USB4 box while the app ran and nothing noticed; the
app detected cards once at start. Now src/hotplug.rs compares every enumeration with the list (key, else vendor +
name when unique; a card whose tool did not answer is never called removed): a new usable card starts a worker,
enabled by default like a card at start, with "New card: <name>, mining" on the strip and in the log; a card
Windows lists with a problem code (Win32_VideoController Status / ConfigManagerErrorCode) is shown as "<name>: not
usable (Code 43)" with the reboot-or-reinstall hint and no worker; a card that disappears has its worker stopped
(quit, 8 s) and its row says removed for five minutes, then hides; an unchanged list touches nothing. The Windows
host sends "detect" on WM_DEVICECHANGE; the engine polls every 60 s (300 s on macOS, no GPU hot-plug there).

detect.rs: the Ryzen iGPU is "gfx1036" to the OpenCL worker, so the APU gfx codes count as integrated, plus the
adapter row's Intel processor string and a dedicated memory under 1 GB; integrated defaults to off with "integrated
GPU, off by default (2 to 3 MH/s for 30 W)" on the row, and the user's choice is kept across re-detections and
restarts (settings, found by key or by vendor + name when the index moved).

Console: the engine logs "cards: <name> [<kind>, <state>] | ..." at start, on every change and every 10 minutes;
relay/lib/parse.mjs reads it and the hot-plug events, the machines API and tools/console.mjs machines show them.

Tests: hotplug.rs (added, removed, moved, errored, recovered, revived, unchanged, twins, user override kept,
the console line), detect.rs (PC 1's adapter lines, the Mac, kind classification, the unusable row),
notices.test.mjs (card notices), relay parse.test.mjs (cards line). cargo test -p igneum-app: 91 passed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 17:45:05 +00:00
igneum-labs
0550f58857 relay: /wake long-poll for the apps' remote jobs (public GET held 45 s, authenticated POST of the stamp)
GET /wake?since=<stamp> is public (the apps hold no token) and rate limited (30 a minute per IP). It holds up to
45 s, re-reading the stamp every 2 s, and answers {stamp, at, added, changed, held_ms} the moment the stored stamp
differs from since, else the unchanged stamp at the deadline. POST /r/<token>/wake {stamp, added} (the relay's
auth, also x-relay-token or x-igneum-key on /wake) records a stamp; one row per stamp in relay_wake, created by the
first POST. maxDuration 60 s for api/wake.mjs in vercel.json. api/relay.mjs is untouched.

The handler lives in lib/wake.mjs with its dependencies injected; relay/test/wake.test.mjs drives it with a fake
database, a fake clock and a fake sleep (the hold, the change, the deadline, the rate limit, the hold cap, auth, a
database error). CI's site job runs it with the other relay tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 08:22:02 +00:00
igneum-labs
67ef2fef96 Console: OTA version captured as major.minor.patch in every update line (the installing line carried a colon)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:47:28 +00:00
igneum-labs
1375c0a6b5 Console: every update line the app logs sets the OTA state, newest wins (the Mac's card said '0.3.3 is current' with 0.3.4 staged); test
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:46:17 +00:00
igneum-labs
88572621ed Console: a finished update shows as the OTA state ('update check: X is current'), not 'none' (PC 2 on 0.3.4); test
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:43:51 +00:00
igneum-labs
21ccb65f40 Console: a machine whose app logged a clean quit or an update, with no status line after it, shows 'stopped (quit|update) N ago' instead of 'silent' (parseAppTail moved to relay/lib/parse.mjs, test)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:04:18 +00:00
igneum-labs
4a32962417 Observer autosync restarts the observer whenever the checked-out tools/observer tree changes (marker + check mode); console stale mark at 180 s (one missed upload is not stale); bugs.md rows
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:48:45 +00:00
igneum-labs
df5f144642 Relay: secrets compared in constant time (relay/lib/auth.mjs, unit test in CI), HSTS header, tools/relay.mjs prints /r/<token> in list and watch (round 4, X28 and X24 part)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:31:22 +00:00
igneum-labs
9915c883b1 Bug hunt: console cards for other/intel workers and a stale mark on old STATUS lines (relay/lib/parse.mjs + test in CI); publish-jobs verifies the live file with retries and named reasons, a verify command, a failed deploy stops, a collect command without $_ is refused; the dl token masked in printed URLs; docs/bugs.md
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:15:05 +00:00