From 7e0fc7da657cc23739f8f555b4ef9e8da116153f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 08:22:02 +0000 Subject: [PATCH] publish-jobs: wake the apps after a verified deploy; jobs.mjs status shows the woken latency; 0.3.6 plan publish-jobs.sh --deploy POSTs the new stamp (published_at plus 8 hex of the file's sha256) and the added id to the relay's /wake once the live file verifies. The relay token goes in a 600-mode header file, never on the command line or the screen. Prints "woke the apps (stamp ...)" or a one-line warning; the apps' 2-minute poll still catches it. tools/jobs.mjs status reads relay_wake (one row per publish with the ids it added) and prints "woken +N s after the publish" for a machine's latest job that a publish added; nothing when the table does not exist yet. docs/plans/release-0.3.6.md: "Instant jobs" section with the design, the expected latency and a TODO row per machine for the measured number once 0.3.6 is live. packaging/ota/README.md: the 10-minute poll is history. Co-Authored-By: Claude Fable 5.1 --- docs/plans/release-0.3.6.md | 30 +++++++++++++++++++++++++++++ packaging/ota/README.md | 4 +++- packaging/ota/publish-jobs.sh | 36 ++++++++++++++++++++++++++++++----- tools/jobs.mjs | 17 +++++++++++++++-- 4 files changed, 79 insertions(+), 8 deletions(-) diff --git a/docs/plans/release-0.3.6.md b/docs/plans/release-0.3.6.md index 4c2a0a448..d8ab5a320 100644 --- a/docs/plans/release-0.3.6.md +++ b/docs/plans/release-0.3.6.md @@ -17,3 +17,33 @@ Written 5 October 2026, 08:45 BST, while proving v0 went live on the devnet at D - Consensus override changes must land on every node at once: a hand node restarted early with a different `proving_v0_activation_daa` was refused by every peer (digest handshake) and sat isolated at a lower height for 20 minutes. Order that works: publish the manifest override, `update-now` to every app, wait for every app node to log the new parameters, then restart the hand nodes and the seed with the same file. - `scratchpad/restart-hand-nodes.sh` died silently after `igneumd --version` (the 0.3.5 binary exits 1 after printing) under `set -e`; the restart it reported never happened. Every restart script ends by printing the new pids and their start times. - Switching proving on needs no app restart: `POST /api/prove {"on":true}` (the job `prove-on-pc2-84100` does this after installing the CUDA host into `/opt/igneum` for the app's WSL user). + +## Instant jobs (5 October 2026) + +the project lead: "why is it taking so long for pc2 and pc1s tasks to spin up? can we speed it up?". Before 0.3.6 every app polled +`igneum-jobs.json` every 10 minutes (`CHECK_EVERY_S = 600`), so a job published from the Mac waited up to 10 minutes +on each PC. The PCs have no inbound ports and one miner is off the LAN, so the fix is a wake signal the app pulls +over an outbound connection. Branch `job-wake`. + +| Piece | What it does | Where | +|---|---|---| +| Wake endpoint | `GET /wake?since=`: public (the apps hold no token), 30 a minute per IP, held up to 45 s, answers `{stamp, at, added, changed, held_ms}` the moment the stored stamp differs from `since`, else the unchanged stamp at the deadline. `POST /r//wake {stamp, added}` (the relay's auth) records a stamp. Rows in Neon `relay_wake`, created by the first POST. `maxDuration` 60 s. | `relay/api/wake.mjs`, `relay/lib/wake.mjs`, `relay/vercel.json`, test `relay/test/wake.test.mjs` (fake database and clock; in CI's site job) | +| App waker | One thread per app long-polls the endpoint with the last stamp; a changed stamp sends `Event::Wake` and the engine fetches the jobs file at once (signature check unchanged). First reply seeds the stamp. Backoff 5, 15, 60 s while the relay is unreachable, 10 s floor between requests, a full hold when the relay has no stamp yet. One log line per wake, one when it falls back, one when it recovers. `IGNEUM_APP_JOBS_WAKE_URL` overrides the URL (set and empty: no waker). | `app/igneum-app/src/jobrun.rs` (`WakeState`, `wake_loop`), unit tests for the stamp and backoff logic | +| Fallback poll | 2 minutes (was 10), the first poll still 40 to 60 s after start. An unchanged file is no longer logged on every poll. | `jobrun.rs` `CHECK_EVERY_S` | +| Publisher | After the live file verifies, `--deploy` POSTs the stamp (`published_at` plus 8 hex of the file's sha256) and the added id; the token goes in a header file, never on the command line; prints "woke the apps" or a one-line warning. | `packaging/ota/publish-jobs.sh` `wake_apps` | +| Status | `node tools/jobs.mjs status` shows "woken +N s after the publish" per machine for a job a publish added. | `tools/jobs.mjs` | + +Expected latency, publish to job start: the relay re-reads the stamp every 2 s inside the hold, the app's curl returns +at once, the engine fetches the file (two small downloads) and starts the job on its next tick. About 3 to 6 s when +the app is mid-hold, plus up to 10 s if the app was inside its floor after an earlier reply; the 2-minute poll is the +ceiling when the relay is down. Numbers below are measured, not estimated. + +| Machine | Publish to started_at | Date | Source | +|---|---|---|---| +| PC 1 (ae432dc7) | TODO (measure with `node tools/jobs.mjs status` after the first 0.3.6 job) | | | +| PC 2 (1ccfe586) | TODO | | | +| Mac | TODO | | | + +Not yet done on this branch: the relay deploy (`relay/`, by the owner; the `/wake` route and the 60 s `maxDuration` +go live with it, the `relay_wake` table appears on the first POST), the first live publish, and the Windows curl path +of the long-poll (curl.exe 8.x in System32; the 58 s `--max-time` was reviewed, not run). diff --git a/packaging/ota/README.md b/packaging/ota/README.md index 74b64be3a..428ec3eca 100644 --- a/packaging/ota/README.md +++ b/packaging/ota/README.md @@ -133,7 +133,9 @@ screenshots are `docs/design/app-screens/update-*.png`. Windows: reviewed only, the project lead's rule: one app on both PCs that the Mac can send commands and files to over the line, so everything is tested and built without a person at the PC. The channel is `igneum-jobs.json` plus `igneum-jobs.json.sig`, next to the -update manifest, signed with the same OTA key and verified by the same code; the apps poll it every 10 minutes. +update manifest, signed with the same OTA key and verified by the same code. Since 0.3.6 (5 October 2026) every app +holds a long-poll on the relay's public `/wake` and fetches the file within seconds of `--deploy` (the script records +the new stamp there); a poll every 2 minutes is the fallback (it was 10 minutes before 0.3.6). The relay (`relay/`) stays for the Mac and for humans; its PC agent is replaced by this. | Piece | Where | diff --git a/packaging/ota/publish-jobs.sh b/packaging/ota/publish-jobs.sh index b102c72f3..d5defb547 100755 --- a/packaging/ota/publish-jobs.sh +++ b/packaging/ota/publish-jobs.sh @@ -1,10 +1,11 @@ #!/usr/bin/env bash # Publishes signed remote jobs for the Igneum Miner apps: igneum-jobs.json (canonical JSON, sorted keys, no # whitespace) and its detached Ed25519 signature igneum-jobs.json.sig, next to the update manifest in the downloads -# folder (dl//), signed on this Mac with the OTA key ~/.config/igneum/ota-signing-key. Every app polls the -# file every 10 minutes (app/igneum-app/src/jobrun.rs), verifies it with the public key compiled into -# src/manifest.rs, runs each job that targets it ONCE per id, and reports to the log intake as -# run_id job-- (read back with tools/jobs.mjs). +# folder (dl//), signed on this Mac with the OTA key ~/.config/igneum/ota-signing-key. Every app holds a +# long-poll on the relay's public /wake (relay/api/wake.mjs) and fetches the file the moment --deploy records the new +# stamp there (0.3.6, app/igneum-app/src/jobrun.rs; a poll every 2 minutes is the fallback), verifies it with the +# public key compiled into src/manifest.rs, runs each job that targets it ONCE per id, and reports to the log intake +# as run_id job-- (read back with tools/jobs.mjs). # # packaging/ota/publish-jobs.sh add --kind run --target 1ccfe586 --script path.ps1 [--elevated] [--stop-miners] \ # [--timeout-minutes 60] [--shell powershell|bash] --title "..." [--expires-hours 48] [--deploy] @@ -153,6 +154,30 @@ verify_live() { # (5 s apart); 0 = verified, 1 = not, with the reason on return 1 } +# After a verified deploy: records the new jobs-file stamp on the relay (relay/api/wake.mjs), where every app holds a +# long-poll and fetches the file the moment the stamp moves (0.3.6). The stamp is published_at plus 8 hex of the +# file's sha256, so a re-signed file wakes the apps too. The relay token (~/.config/igneum/relay-token) travels in a +# header file, never on the command line or the screen. A failure here is a warning: the apps poll every 2 minutes. +wake_apps() { # + local tf="$HOME/.config/igneum/relay-token" rurl="https://relay.igneum.network" sum size pub stamp hdr out rc=0 added='[]' + [ -f "$tf" ] || { echo "warning: no $tf; the apps were not woken (they poll every 2 minutes)" >&2; return 0; } + [ -f "$HOME/.config/igneum/relay-url" ] && rurl="$(tr -d '[:space:]' < "$HOME/.config/igneum/relay-url")" + rurl="${rurl%/}" + read -r sum size < <("$SIGNER" sha256 "$JOBS") + pub="$(python3 -c 'import json,sys; print(json.load(open(sys.argv[1])).get("published_at", ""))' "$JOBS")" + stamp="$pub.${sum:0:8}" + if [ -n "${1:-}" ]; then added="[\"$1\"]"; fi + hdr="$(mktemp)"; chmod 600 "$hdr" + printf 'x-relay-token: %s\n' "$(tr -d '[:space:]' < "$tf")" > "$hdr" + out="$(curl -sS --max-time 20 -X POST "$rurl/wake" -H 'Content-Type: application/json' -H @"$hdr" --data-binary "{\"stamp\":\"$stamp\",\"added\":$added}" 2>&1)" || rc=$? + rm -f "$hdr" + if [ "$rc" = 0 ] && printf '%s' "$out" | grep -q '"ok":true'; then + echo "woke the apps (stamp $stamp)" + else + echo "warning: the wake call to $rurl/wake failed (${out:0:160}); the apps poll every 2 minutes and catch it" >&2 + fi +} + if [ "$CMD" = verify ]; then [ -f "$JOBS" ] || { echo "no local jobs file at $JOBS" >&2; exit 1; } verify_live "$TRIES"; exit $? @@ -315,7 +340,8 @@ if [ "$DEPLOY" = 1 ]; then (cd "$DLSITE" && npx --yes vercel@latest --global-config "$HOME/.config/igneum/vercel" deploy --prod --yes 2>&1 | sed "s#$TOKEN##g"; exit "${PIPESTATUS[0]}") \ || { echo "the deploy failed (the Vercel CLI's exit status above); nothing verified" >&2; exit 1; } verify_live "$TRIES" || exit 1 - echo "the apps pick it up within 10 minutes (Settings > remote jobs > Check now at once); results: node tools/jobs.mjs ${ID:-}" + wake_apps "$ID" + echo "the apps fetch it within seconds when woken, else within 2 minutes (Settings > remote jobs > Check now at once); results: node tools/jobs.mjs ${ID:-}" else if [ -n "$DLSITE" ]; then echo "not deployed: cd $DLSITE && npx --yes vercel@latest --global-config ~/.config/igneum/vercel deploy --prod --yes (or re-run with --deploy)" diff --git a/tools/jobs.mjs b/tools/jobs.mjs index d124169a9..5b1fadc55 100755 --- a/tools/jobs.mjs +++ b/tools/jobs.mjs @@ -1,7 +1,9 @@ #!/usr/bin/env node // Mac side of the remote jobs (app/igneum-app/src/jobs.rs, published by packaging/ota/publish-jobs.sh). // node tools/jobs.mjs the published jobs file: fetched from the downloads host, signature checked -// node tools/jobs.mjs status per machine, from the log intake: the latest job run and its SUMMARY line +// node tools/jobs.mjs status per machine, from the log intake: the latest job run and its SUMMARY line, +// plus the woken latency (job started_at minus the publish that added it, +// from relay_wake, written by publish-jobs.sh --deploy since 0.3.6) // node tools/jobs.mjs the result: the newest upload per label under run_id job-- // node tools/jobs.mjs --all every upload, oldest first (the 5-minute progress reports of a long job) // node tools/jobs.mjs watch poll the intake every 30 s until every reporting machine is final @@ -79,9 +81,20 @@ if (a === 'status') { SELECT DISTINCT ON (machine) machine, run_id, label, received_at, lines FROM miner_logs WHERE run_id LIKE 'job-%' AND label LIKE 'job-%' ORDER BY machine, received_at DESC`); if (!rows.length) { console.log('no job reports in the intake yet'); process.exit(0); } + // the woken latency: each publish is one relay_wake row (stamp, at, meta.added = the ids it added); a machine's + // latest job that a publish added shows its started_at minus that publish's at. No table yet: nothing shown. + let publishes = []; + try { publishes = await sql('SELECT stamp, at, meta FROM relay_wake ORDER BY id DESC LIMIT 200'); } catch { publishes = []; } + const tsOf = v => Date.parse(String(v || '').replace(' ', 'T').replace(/([+-]\d\d)$/, '$1:00')); + const publishedAt = id => { for (const p of publishes) { const a = p.meta && Array.isArray(p.meta.added) ? p.meta.added : []; if (a.includes(id)) return tsOf(p.at); } return NaN; }; + const woken = s => { + if (!s || !s.job || !s.started_at) return ''; + const d = Math.round((tsOf(s.started_at) - publishedAt(s.job)) / 1000); + return Number.isFinite(d) && d >= 0 && d < 86400 ? ` woken +${d} s after the publish` : ''; + }; for (const r of rows) { const s = summaryOf(r.lines); - console.log(`${r.machine.padEnd(28)} ${r.run_id.padEnd(44)} ${when(r.received_at)} ${s ? `${s.status} exit ${s.exit} after ${s.duration_s} s: ${s.summary}` : '(no SUMMARY line)'}`); + console.log(`${r.machine.padEnd(28)} ${r.run_id.padEnd(44)} ${when(r.received_at)} ${s ? `${s.status} exit ${s.exit} after ${s.duration_s} s: ${s.summary}` : '(no SUMMARY line)'}${woken(s)}`); if (s && s.errors && s.errors.length) for (const e of s.errors) console.log(`${''.padEnd(28)} ERROR ${e.slice(0, 160)}`); } process.exit(0);