Commit graph

56 commits

Author SHA1 Message Date
igneum-labs
f225ccf842 run 4's cause: the playbook's PowerShell JSON round trip rewrote the copied settings (big integers as doubles), the engine read the file as defaults (no payout address, every card off, 96 old remote jobs run in the scratch root). Fix: the copies are verbatim and a --sweep engine applies Settings::for_measurement in memory (remote jobs, updates, proving and Power control off, not paused, tune on, every card due and unpinned); the CI check fails any playbook that rewrites settings.json through ConvertTo-Json; unit test
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 14:21:06 +00:00
igneum-labs
eb3d9e51f3 ember-tune-pc1.ps1: the watchdog's exits dump the engine's log tail and stderr first (run 4 left no engine line)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 14:18:29 +00:00
igneum-labs
280631fbc6 tools/ci/playbook-quit-check.sh: the standing rule of 5 October 2026 23:05 UTC as a gate (a playbook that reads the installed app's URL file and sends quit, pause or resume fails; the installer's own stop step is the one allowed sender; self-test on a bad and a good case); shard-test.ps1 loses its api/quit to the installed app; ember-tune-pc1.ps1's refusal guard reworded
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 11:52:08 +00:00
igneum-labs
32febb5f56 ember-tune-pc1.ps1: a fresh scratch root per run (run 3 at 11:45Z: the folder an earlier elevated engine had locked refused the settings copy, the engine exited in 6 s on stale files), the installed engine preferred from 0.3.12 on, the address guard hard
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 11:47:08 +00:00
igneum-labs
be0937bc41 ember-tune-pc1.ps1: budget 45 min (two NVIDIA cards, 19 steps)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 11:44:58 +00:00
igneum-labs
563463b35f Merge release-0.3.12 (fda4684) into ember-tune: 0.3.11's six-section View and card order kept, Ember Tune's line and switches re-added on it; the tune fields move into hotplug::apply_pref; the power-cap plan keeps present(); both CI test lists; 132 app tests, 26 UI tests, every gate green
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:20:46 +00:00
igneum-labs
6de4827ccb a job that cannot mine never burns its budget silently: the elevated job path follows its output file while the script runs (the 5-minute progress reports carry the lines; 0.3.12), and the tune playbook's watchdog fails a run that mines nothing within 120 s of its first status line (the engine's last log line in the RESULT, the tree ended, mining restored by the runner)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:03:13 +00:00
igneum-labs
d1705e0e1b playbooks: the scratch settings.json is written without a BOM (PowerShell 5.1's -Encoding utf8 adds one, the engine's JSON parser refuses it, the copy read as defaults with no payout address and nothing mined in runs 1 and 2); the address is read back and the job fails at once if it is empty; the CI check fails any playbook writing JSON with Set-Content -Encoding utf8
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 07:58:39 +00:00
igneum-labs
715890b0da ember-tune-pc1.ps1: find the AMD helper under any amd-kit job folder
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 07:20:36 +00:00
igneum-labs
6c38efc1c6 ember-tune-pc1.ps1: both cards stay on their best points (the project lead, 6 October 2026 07:25Z); no reset at the end
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 07:17:21 +00:00
igneum-labs
e941588c58 ember-tune-pc1.ps1: the copy runs with Power control off (the installed setting is read and reported, never a prompt: the project lead, 6 October 2026 07:20Z), the kit paths after the update, the AMD card reset to factory at the end
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 07:15:43 +00:00
igneum-labs
3d505766d2 C35 named: PC 1's 22:31 UTC quit was the per-user installer launched by the second engine's own updater (0.3.9 under min_supported_version = urgent, beating auto_update = false); a second engine never runs the updater (IGNEUM_APP_NO_OTA=1, implied by --sweep; the playbooks set it; the CI check demands it); bench log and plan carry the named source
Source: the scratch engine's own log in collect ember-c35-collect-1 (06:59Z): 22:31:02Z '0.3.10 is available: downloading',
22:31:05Z 'update: starting the installer first ... ota-apply.ps1', and the installed app's 'quit:' at 22:31:06Z.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 07:01:57 +00:00
igneum-labs
bfdc632a14 Merge ca2-analysis (c9ebe7b) into release-0.3.11: docs and standalone probe sources the evidence rows cite (C38); nothing a built artefact reads
# Conflicts:
#	docs/bench-log.md
2026-10-05 23:03:45 +00:00
igneum-labs
d19441c14d C35 class: a second engine gets no pipe (its output goes to a file the playbook tails) and its whole tree is ended at the end and on the budget; ember-tune-pc1.ps1 and sweep-5090.ps1 fixed; tools/ci/second-engine-check.sh fails any playbook without both; the rule in ember-tune.md
PC 1, 22:31 UTC: the installed engine's quit hung 24 minutes in the jobs runner's abort, waiting for EOF on the script's
stdout pipe whose write end the second engine and its miners had inherited (Process.Start with redirection inherits
every inheritable handle), while the orphaned miners mined on against the relaunched app.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 23:00:06 +00:00
igneum-labs
d69e1c9fae C35: every quit names its source (Cmd::Quit carries it: the window host's stdin, the host gone, POST /api/quit, the --sweep run's end); the --sweep job never counts as Power control and sets no cap at start (it raised a UAC prompt on PC 1 at 22:30 UTC); the tune playbook's budget quit goes only to its own scratch URL file and says so
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:53:32 +00:00
igneum-labs
19b30e31f0 Merge remote-tracking branch 'origin/master' into ca2-v3
# Conflicts:
#	docs/bench-log.md
#	proto-opencl/host.c
2026-10-05 22:34:07 +00:00
igneum-labs
829687b1ab mixer x4: the measure session (v2 / x4 / x8 verifier 1.33 / 1.94 / 2.79 ms per unit on a loaded core, 1.45x and 2.1x; the 256 MiB fill 172 to 175 ms; the Metal 1 GiB build flat at 21 ms, latency-bound), the verification-throughput consequences (C19), what is unverified and what is owed; the bench-log entry; the PC 1 two-card playbook; the measured verifier row in the chip model
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:02:28 +00:00
igneum-labs
147db8db44 mixer x8 candidate beside x4 (coordinator's rule, 5 October 2026 21:30 UTC): LoadClass::MX8 ("mx8", same stream, keys (r m + j + 1) x 0x9E3779B9), the two candidate packs mx8-genesis and mx8-devnet-epoch0 (generator 2 with the class in the id; the vectors are re-cut through the seam after the x4/x8 choice), tests/mixer.rs fuzz takes IGNEUM_MIXER_CLASS, the PC playbooks carry the x8 packs and the gfx1036 fallback; mixer-x4.md: the Metal fuzz row
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:02:23 +00:00
igneum-labs
39ecd7c52b hot table (Counter ASIC 2.0 layer 5), measured and not adopted, on the ca2-v3 composed class (squash of tag ca2-cache-history-2026-10-05)
LoadClass::hot (Option<HotClass { mb, k, added }>) beside mix, slots, scratch, mixer_mult, growth and era; V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 } (hot stays None: the hot code is behind the flag, measured and not adopted). Op::Hot, the hot slots drawn after the scratch slots, no width roll (v2_loads allows the added form's extra slots), the id suffix hot/<S><k>[added], the hot parsing inside parse_loads, the era branch first in name(). HotTable under seed_words("igneum-hot/" || epoch seed) with the cache chain and tag umHT, read at H[mulhi(src, HOT_WORDS)]; DatasetSource::hot attached by new_class_day and from_seed_bytes_class; the acceptance stand-in dataset_elem(idx, S[2], S[3]); the three emitters (hot argument after the init words, ht_segment and igneum_hot_fill beside the layout-aware cores); packfile.h hot fields beside class, era, attempt and mixerMult; OpenCL host, Metal packbench and NVRTC worker fill H on the device and self-test it. Eight packs under proto-cuda/packs-ca2-hot (replaced hot32k4 hot64k4 hot96k4 hot64k2 hot64k8, added hot32k4a hot64k4a hot96k4a) re-exported on the merged crate: vectors unchanged, program.h and program.json carry the mixer fields. docs/plans/hot-table.md (design, spec text, per-tier budget, chip model, Mac and PC measurements, the decision: layer 5 out of v3, the 3.0 note); bench-log entry and addenda with job ids and worker sha256s; the two PC playbooks.

Checks on this commit: cargo test --release 53 + 19 pass (the pinned v2, mx4, era, readwidth and hot packs); the pinned packs under proto-cuda/packs, packs-ca2-mixer, packs-ca2-era and packs-readwidth untouched; Metal packbench and Apple OpenCL --bench-pack on all eight hot packs (run lock, 2^20 at base 0): 96/96 lanes, hot table head, last line and FNV PASS, one fingerprint per pack on both harnesses, equal to the fingerprints before the rebase (hot32k4 679e5e83378d3790, hot64k4 d4c9e456b039fdef, hot96k4 7c98eceffee9fd73, hot64k2 f43b10a95879b8e5, hot64k8 c11309d743be9392, hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5). No era or mixer behaviour changed: every resolution kept the ca2-v3 side and appended the hot branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:57:39 +00:00
igneum-labs
1558971521 era layout (Counter ASIC 2.0 layers 4 and 8) behind the class flag, on the ca2-v3 seam: the 7-draw era stream from E_n (stride, interleave, width pinned at 4 B), per-site window draws (dataset, half, quarter at a 256 MiB floor), the strided windowed load address in the interpreter, the acceptance mirror and the three emitters, the interleaved dataset layout riding with the program (memhard::Layout, mh_t/mh_j/mh_addr, Epoch::dataset_word), V3_CLASS with the era drawn inside by generate_from_seed_bytes_program_class, --era / --era-widths on the CLI, six class v3 era packs (packs-ca2-era), tests, the design doc, host.cu/host.c deriving host words through the pack's mh_word, emu/test-layout.sh, the PC 1 playbook
Squashed from six commits (tag ca2-era-pre-squash) for one merge.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:41:33 +00:00
igneum-labs
5c808b89e0 Ember Tune: every card tuned for MH per watt out of the box, the fleet prior per card model in the signed manifest, the console and /miners priors table
the project lead, 5 October 2026, 22:45 BST: "make sure we have ember tuning every single card for efficiency out of the box, the
more data = the better the tune, make an awesome system." Built on lever 3 (docs/plans/miner-eff.md), lever 2's signed
tuning section (docs/design/miner-tuning.md), the AMD telemetry helper (423936b, its --tune/--set-gmax/--set-plimit/
--reset contract) and the Power control switch (057f0ec). Design, data flow, tiers and the privacy line:
docs/plans/ember-tune.md.

- src/ember.rs (new): two knobs per card (power limit %, core clock cap MHz; memory clock never touched), the full plan
  (power ladder 100..50%, then the clock ladder 90..60% at the chosen power), the confirm plan (the fleet prior and one
  neighbour), the baseline plan (measure only), the marks (faulted, hot, memory_clock_dropped, unapplied, no_readings),
  the choice (best MH/W within 1% of the top rate, then rate, then draw), the fleet record (a hash of the install id,
  no address), the prior lookup and the kill switch (tuning.ember), the state machine on a fake clock. 9 unit tests.
- engine.rs: tick_sweep schedules every NVIDIA, AMD and Apple card (120 s steady, 600 s to the boundary, no job hold,
  no pause, weekly, again after a driver major or program-class change, never under the manifest kill switch); the
  probe (nvidia-smi clocks.max.gr + driver_version and the direct/helper mode; igneum-gpu-telemetry --tune for AMD);
  tune_apply (nvidia-smi -pl / -lgc 0,<MHz> / -rgc directly or through the helper; the AMD helper per request);
  Cmd::TuneProbe, Cmd::TuneSet; faults from rejected and mismatched hashes mark the step; the TUNE lines and the TUNE
  {json} record, uploaded with the log; the Tuned line on the card state. The NVIDIA helper starts only with Power
  control on: the --sweep job never counts as permission (no prompt on a PC with nobody there).
- sweep.rs: the helper protocol gains lgc/rgc (clock cap and reset) and resets the clocks after 20 idle minutes.
- state.rs, config.rs: the tune fields (clock cap, driver, class, source, the Tuned line); the nvidia-smi telemetry
  query carries clocks.gr and clocks.mem; the AMD sample line's plimit_pct and gmax_mhz are parsed.
- ui: "Tuned: X MH/s at Y W (Z MH/W)" with the point, the source and when; measure-only cards say why; the Ember Tune
  switch; tune-line.test.mjs.
- relay/lib/ember.mjs + relay/test/ember.test.mjs: the aggregation per (card model | driver major | program class):
  median point, MH/W, spread, samples, machines; five samples converge, an outlier does not move the median, baselines
  make no prior, de-duplication, the manifest merge keeps lever 2's cards. api/console.mjs fn=tuning and
  tools/console.mjs tuning; tools/tuning.mjs --priors [--write tuning.json] [--site] [--tuning-off].
- site: the fleet priors table on /miners (site/miner-priors.json), the lever text.
- relay/playbooks/ember-tune-pc1.ps1: the PC 1 run (second engine with --sweep from a scratch copy of the install).

Measured tonight: see the bench log entry that follows the PC 1 run. The 9070 XT left PC 1's bus at 20:40 UTC and the
5090 needs the administrator prompt the project lead cannot answer asleep, so tonight's PC 1 run is the baseline plan on the 5090
through the whole pipeline; the two-knob tune on both cards is owed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:25:09 +00:00
igneum-labs
e08909f138 mixer x4 and the cache growth rule (Counter ASIC 2.0, class v3 construction): LoadClass mixer_mult and growth, LoadClass::MX4 (v2 loads, no width roll), memhard::Shape in MixParams, m mixer applications per round with keys round_key(r m + j), Cache::fill_log2, the option C schedule (growth_doublings, cache_log2_words, dataset_log2_words, days_since_genesis) with its test table, day-sized Epoch entries, the three emitters (m loop only for m > 1, v2 text unchanged), program.h and program.json fields, packfile.h mixerMult, packbench and OpenCL host prints, --class mx4 and --days on the CLI; docs/plans/mixer-x4.md design and spec text, docs/analysis/chip-model-v3.md, the 5090 and 9070 XT dataset-build playbooks (measurements and vectors to follow)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:54:10 +00:00
igneum-labs
d28a32e403 Merge commit '4163143' into ca2-v3
# Conflicts:
#	proto-cuda/nvrtc/packfile.h
2026-10-05 20:17:46 +00:00
igneum-labs
eeef9cd056 Counter ASIC 2.0 layer 7: integer matrix family design (vendor primitives cited: PTX dp4a and mma .u8/.s8, AMD v_dot4_i32_iu8 and WMMA iu8 on RDNA 3 and 4 via LLVM and GPUOpen, CDNA 3 MFMA i8, Metal 4 matmul2d char x char -> int found in MSL 4.1 table 7.3; dot4 and mm8 semantics, reserve entry R1, emulation rule); M5 Max dot4 probe numbers in the bench log; PC playbook dot4-probe.ps1 prepared, not published
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:12:53 +00:00
igneum-labs
a094972548 read-width: packfile accepts a string-seed pack (any even-length hex epoch seed; the seed-word re-derivation still checks it); bench-only playbooks for the second PC round
Round 1 (run-readwidth-{5090,9070}-20261005) delivered the probes and refused every pack: pf_load demanded the chain's 32-byte epoch seed and the experiment packs carry igneum-pow --seed strings.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:09:58 +00:00
igneum-labs
8c898d3f60 read-width: CUDA worker --bench and --memprobe, scratch arena from the occupancy capacity; PC playbooks (card under test off in the app, restored after); harness arena label
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:05:57 +00:00
igneum-labs
d5da02a851 Merge pack-loop (4163143) into release-0.3.10: a pack's seed words are its program attempt's words, not the bare seed's (the epoch 34 outage)
# Conflicts:
#	relay/test/parse.test.mjs
2026-10-05 19:23:59 +00:00
igneum-labs
41631431c0 Workers: a pack's seed words are its program attempt's words, not the bare seed's (epoch 34 incident, 5 October 2026)
From 18:23Z both Windows workers (CUDA on PC 1 and PC 2, OpenCL on PC 1 after the 18:34Z node restart)
refused every pack for epoch 34 with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", and
the miner and the app restarted them every 5 to 60 s until 18:44Z and beyond. The packs were correct.
The generator retries a rejected candidate with seed || k_le32 (attempt_words); epoch 34's attempt 0 was
rejected (246 of 16384 final register values saturated, limit 163) and attempt 1 accepted, so the pack
carried attempt 1's words while pf_load (proto-cuda/nvrtc/packfile.h) derived the expected words from
the bare seed. Both workers also matched jobs to pairs by those bare-seed words, so even a loaded pack
of a retried program would have answered "epoch seed mismatch" on every job.

The rule, in one place per language:
- packfile.h: pf_program_words(bytes, attempt); pf_load reads IGNEUM_PROGRAM_ATTEMPT and checks the
  attempt's words; the refusal says "program pack and its seeds disagree: IGNEUM_SEEDW_INIT is not
  attempt N of the epoch seed ..." in plain words.
- worker.cpp and proto-opencl/host.c: a job belongs to a pair when the seed hex the node sent is the
  pair's (pairIs); the compiled-in placeholder pack keeps the word comparison.
- igneum-pow/src/packcheck.rs: verify_pack_texts / verify_pack_dir, the same rule in Rust; the miner
  checks every pack it writes with it before a worker sees it (vendor/igneum-node pack-loop branch).
  Tests pin the attempt vectors of epoch 34 on both sides (one vector, two implementations), that
  epoch 34 is attempt 1 and epoch 33 attempt 0, a known-good pack of a later attempt, a known-mismatched
  (out of date) pack, and self-contradicting packs.
- proto-cuda/nvrtc/emu/packfile-test.c (+ .sh, in CI): pf_load on a known-good attempt-1 pack, the
  checked-in attempt-0 pack, and the known-mismatched bare-words pack.

The app (app/igneum-app):
- watchdog.rs: PACK_OUT_OF_DATE_CODE 44, PackRebuilds (at most 3 pack exports per epoch, then the card
  shows the reason), pack_refusal (the worker's "error 0 pack" line and the miner's "PACK OUT OF DATE"
  line), pack_epoch_of; tests on the incident lines, known-good and known-mismatched.
- engine.rs: exit 44 exports the pack again before the restart instead of a blind restart, the strip
  says "program pack out of date, rebuilding", the card and the log name the condition; at the cap the
  card is marked failed with the reason and tries again in 10 minutes.

The relay (relay/lib/parse.mjs): PACK_MISMATCH; the card reads "pack mismatch, rebuilding (N refusals
in the tail, M restarts)" in `node tools/console.mjs machines` instead of a bare restart count; tests
on the PC 2 tail of 18:27Z and a healthy tail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 19:10:07 +00:00
igneum-labs
d7049fdfda app: GPU hot-plug (re-detection every minute, WM_DEVICECHANGE on Windows), faulty cards listed with the problem code, integrated GPUs off by default, the card list in the console
5 October 2026: an RX 9070 XT went into PC 1 through a Sonnet USB4 box while the app ran and nothing noticed; the
app detected cards once at start. Now src/hotplug.rs compares every enumeration with the list (key, else vendor +
name when unique; a card whose tool did not answer is never called removed): a new usable card starts a worker,
enabled by default like a card at start, with "New card: <name>, mining" on the strip and in the log; a card
Windows lists with a problem code (Win32_VideoController Status / ConfigManagerErrorCode) is shown as "<name>: not
usable (Code 43)" with the reboot-or-reinstall hint and no worker; a card that disappears has its worker stopped
(quit, 8 s) and its row says removed for five minutes, then hides; an unchanged list touches nothing. The Windows
host sends "detect" on WM_DEVICECHANGE; the engine polls every 60 s (300 s on macOS, no GPU hot-plug there).

detect.rs: the Ryzen iGPU is "gfx1036" to the OpenCL worker, so the APU gfx codes count as integrated, plus the
adapter row's Intel processor string and a dedicated memory under 1 GB; integrated defaults to off with "integrated
GPU, off by default (2 to 3 MH/s for 30 W)" on the row, and the user's choice is kept across re-detections and
restarts (settings, found by key or by vendor + name when the index moved).

Console: the engine logs "cards: <name> [<kind>, <state>] | ..." at start, on every change and every 10 minutes;
relay/lib/parse.mjs reads it and the hot-plug events, the machines API and tools/console.mjs machines show them.

Tests: hotplug.rs (added, removed, moved, errored, recovered, revived, unchanged, twins, user override kept,
the console line), detect.rs (PC 1's adapter lines, the Mac, kind classification, the unusable row),
notices.test.mjs (card notices), relay parse.test.mjs (cards line). cargo test -p igneum-app: 91 passed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 17:45:05 +00:00
igneum-labs
2edc5f693c Merge job-wake: instant job wake-up (relay /wake long-poll, 2-minute fallback poll, publisher wake)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 08:26:07 +00:00
igneum-labs
2866737139 relay: the packaged intake key reports too (build job uploads got 'no token' after the key split)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 08:24:31 +00:00
igneum-labs
0550f58857 relay: /wake long-poll for the apps' remote jobs (public GET held 45 s, authenticated POST of the stamp)
GET /wake?since=<stamp> is public (the apps hold no token) and rate limited (30 a minute per IP). It holds up to
45 s, re-reading the stamp every 2 s, and answers {stamp, at, added, changed, held_ms} the moment the stored stamp
differs from since, else the unchanged stamp at the deadline. POST /r/<token>/wake {stamp, added} (the relay's
auth, also x-relay-token or x-igneum-key on /wake) records a stamp; one row per stamp in relay_wake, created by the
first POST. maxDuration 60 s for api/wake.mjs in vercel.json. api/relay.mjs is untouched.

The handler lives in lib/wake.mjs with its dependencies injected; relay/test/wake.test.mjs drives it with a fake
database, a fake clock and a fake sleep (the hold, the change, the deadline, the rate limit, the hold cap, auth, a
database error). CI's site job runs it with the other relay tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 08:22:02 +00:00
igneum-labs
7cd5f39b30 Merge remote-tracking branch 'origin/master' into release-0.3.5 2026-10-04 21:35:05 +00:00
igneum-labs
ca2ac6dfa8 Merge origin/miner-perf into release-0.3.5
Conflicts resolved: state.rs keeps both the sweep fields (miner-eff) and the race fields (miner-perf); bench-log.md
keeps both appended entries; publish-manifest.sh keeps master's --override implementation (8082576, the "every
height switch" rule, --verify-only, --tries, the retrying live check) and adds miner-perf's --tuning / --no-tuning
with the carry-over of consensus.override and tuning from the current manifest. One --override case, one parser.
2026-10-04 21:31:26 +00:00
igneum-labs
88017d69e1 Console: the last OTA state per machine and app run is remembered (console_ota_memo) and shown when the upload's 256 KiB tail no longer holds it
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 21:27:59 +00:00
igneum-labs
697b012954 Console: the app log's parsed tail is 400 KB so a stuck update's lines stay on the card (PC 1's 'installing' scrolled out of 60 KB in 30 min)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 21:26:05 +00:00
igneum-labs
68f3221530 Merge remote-tracking branch 'origin/miner-eff' into release-0.3.5 2026-10-04 21:25:44 +00:00
igneum-labs
cda444bbee READMEs: the console's Machines card as it is now (stale, stopped, OTA state, vendors), the parser tests, autosync's restart key and check mode, the observer's dependence on the app's node for proving
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 21:14:45 +00:00
igneum-labs
67ef2fef96 Console: OTA version captured as major.minor.patch in every update line (the installing line carried a colon)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:47:28 +00:00
igneum-labs
1375c0a6b5 Console: every update line the app logs sets the OTA state, newest wins (the Mac's card said '0.3.3 is current' with 0.3.4 staged); test
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:46:17 +00:00
igneum-labs
88572621ed Console: a finished update shows as the OTA state ('update check: X is current'), not 'none' (PC 2 on 0.3.4); test
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:43:51 +00:00
igneum-labs
7778dd7cde Miner efficiency sweep: hash per watt per NVIDIA card (lever 3)
src/sweep.rs (new): the cap steps 100% to 50% in 10% steps clamped to the card's limits, per-step rows (mean draw
from nvidia-smi power.draw, mean worker interval rate), the choice (best MH/W, ties to the higher rate then the lower
cap), the nvidia-smi power parser, a state machine on an explicit clock (15 s settle, 60 s hold, 30 s cap readback
limit), the elevated helper scripts (one administrator prompt per sweep: a command file polled by one elevated
process, self-restoring after 20 idle minutes), the unsupported reasons (Apple silicon, AMD). 9 unit tests with
PC 1's recorded RTX 5090 numbers (575 W default, 460 W cap, 290 W draw, memory temperature [N/A]).

Engine: scheduler (once after install, then weekly; one card at a time; only while the card mines, after 120 s
steady, never under a remote job hold, a pause, or inside 600 s of the hour boundary), the cap-mode probe (direct
when the engine runs elevated, else the helper), abort on any fault (card leaves mining, worker error, GPU 90 C,
job, pause, quit) with the cap restored, the chosen cap held and recorded, SWEEP table lines in the app log,
--sweep mode (sweep every supported card, print the table on stdout, leave the caps, quit). Cap floor 50% (was 60).
A readback that matches the asked cap now counts as applied (PC 1 showed "cap NOT applied" for hours at 460 W).

Dashboard: live eff MH/W on each tile, the sweep line (phase, last result, or why unsupported), Sweep now / Stop /
Unpin, "pinned" and "chosen by the sweep" on the cap line, the Settings toggle, the cards-page note, slider min 50.
A cap moved by hand pins the card: the sweep records but does not change it.

PC 1 measurement: relay/playbooks/sweep-5090.ps1 (a run job, elevated, miners stopped; a second engine with --sweep
in a scratch data folder, RESULT SWEEP lines) and docs/plans/miner-eff.md with the publish command. Not published.
Untested on a card.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:38:07 +00:00
igneum-labs
21ccb65f40 Console: a machine whose app logged a clean quit or an update, with no status line after it, shows 'stopped (quit|update) N ago' instead of 'silent' (parseAppTail moved to relay/lib/parse.mjs, test)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:04:18 +00:00
igneum-labs
4a32962417 Observer autosync restarts the observer whenever the checked-out tools/observer tree changes (marker + check mode); console stale mark at 180 s (one missed upload is not stale); bugs.md rows
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:48:45 +00:00
igneum-labs
df5f144642 Relay: secrets compared in constant time (relay/lib/auth.mjs, unit test in CI), HSTS header, tools/relay.mjs prints /r/<token> in list and watch (round 4, X28 and X24 part)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:31:22 +00:00
igneum-labs
9915c883b1 Bug hunt: console cards for other/intel workers and a stale mark on old STATUS lines (relay/lib/parse.mjs + test in CI); publish-jobs verifies the live file with retries and named reasons, a verify command, a failed deploy stops, a collect command without $_ is refused; the dl token masked in printed URLs; docs/bugs.md
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:15:05 +00:00
igneum-labs
fe3a3ab701 Fleet learning for the kernel race: the TUNING record, the aggregator, the manifest's tuning object, the PC 1 race job
The engine turns a worker's race line into one TUNING {json} line in the app log (card model as the worker names it,
driver, arch, program class loads and wide loads, every variant's MH/s, winner, gain, the card's power cap and
draw, MH per watt), which the existing intake receives; the card state carries the variant for the dashboard and
one event per race. tools/tuning.mjs aggregates the records from miner_logs per card model (median MH/s or MH per
watt, at least 3 samples, de-duplicated per race) and writes tuning.json; publish-manifest.sh --tuning puts it in
the signed manifest (and now takes --override for consensus.override; both are carried over from the current
manifest when not given, --no-tuning drops it); manifest.rs parses it; ota.rs writes <app data>/tuning.json and
removes it when the manifest drops it; procs::spawn takes an environment and every miner starts with
IGNEUM_TUNING_FILE, which its worker reads at every prepare. Dry run of the publisher against a scratch folder:
tuning and override written, carried over, dropped, signature verified.

docs/plans/miner-perf.md: the signed jobs for PC 1 (fetch the race build of the NVRTC worker, then
relay/playbooks/race-5090.ps1 with the miners stopped: 17 variants, 3 rounds, twice) with the exact publish
commands for the main session; not published by the agent.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:08:45 +00:00
igneum-labs
63174a5b6a Relay: its own key (relay-key) replaces the intake key for the Mac tools and clients; relay token rotated 4 Oct 2026 (round 4, X23); prove package excludes cross-build folders
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 18:25:19 +00:00
igneum-labs
ec0ede62b9 Relay: run and task posts need the console token (round 4, X23); prove host saves proofs buffered (ledger P20, second gap)
The intake key sits in every miner package, so the relay now lets it report only (drop text and files, ack, done,
register, upload). Posting a run or task, or renaming and re-roling a machine, needs the console token.
The prove host wrote proofs through SP1's unbuffered save: on WSL2 under /mnt/c the 18 MB core proof of a shard
took longer to save than to prove. Proofs now go through a 4 MB buffer with a timed 'saved' line, and
prove-shard.sh keeps results on the Linux side and copies them per stage. Ledger P20 updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 18:14:10 +00:00
igneum-labs
1c1dbd93d6 Console: a job is done only on the app's closing report; prove-shard.sh takes every block fixture after the first argument
The console marked any job with a RESULT line as done, so a running shard job read as finished. Done now means
the SUMMARY line carries finished_at or the job's closing 'job <id>: <status> (exit N)' line is present.
prove-shard.sh dropped the third fixture argument (block-344-shards4) because it read only $2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 17:52:35 +00:00