igneum/docs/plans/release-0.3.6.md
igneum-labs 7e0fc7da65 publish-jobs: wake the apps after a verified deploy; jobs.mjs status shows the woken latency; 0.3.6 plan
publish-jobs.sh --deploy POSTs the new stamp (published_at plus 8 hex of the file's sha256) and the added id to the
relay's /wake once the live file verifies. The relay token goes in a 600-mode header file, never on the command line
or the screen. Prints "woke the apps (stamp ...)" or a one-line warning; the apps' 2-minute poll still catches it.

tools/jobs.mjs status reads relay_wake (one row per publish with the ids it added) and prints "woken +N s after the
publish" for a machine's latest job that a publish added; nothing when the table does not exist yet.

docs/plans/release-0.3.6.md: "Instant jobs" section with the design, the expected latency and a TODO row per machine
for the measured number once 0.3.6 is live. packaging/ota/README.md: the 10-minute poll is history.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 08:22:02 +00:00

5.9 KiB

Release 0.3.6: what the proving activation taught us

Written 5 October 2026, 08:45 BST, while proving v0 went live on the devnet at DAA 84,100.

Must ship in 0.3.6

Item Why Where
The app sets IGNEUM_PROOF_VERIFIER for its node On 5 October no node on the network ran a verifier (verifier: Off on every app node and the seed). Spec 7.7 item 4: a producer without a verifier never includes a record, so proofs were stored and never paid. The Mac app was relaunched by hand with the variable; Windows apps cannot be. app/igneum-app/src/engine.rs node spawn: Mac and Linux point at the shipped igneum-prove-host; Windows points at a small wrapper exe that runs the WSL host (wsl -d Ubuntu-24.04 -- /opt/igneum/igneum-prove-host --mode verify ... with wslpath for the proof file). Until the wrapper exists, devnet Windows nodes run IGNEUM_PROOF_VERIFY=trust (devnet only, never testnet).
Windows payload ships wsl2/bin/igneum-prove-host and igneum-prove-export The app's probe looks there first; today a PC needs the 20-minute WSL setup or a hand-installed /opt/igneum. The Linux binaries come from infra/cross/build-linux.sh (CUDA feature needs the CUDA toolchain in the cross image, else ship the CPU build and let setup-wsl.sh build the CUDA one). packaging/windows/make-payload.sh, Igneum-Miner.iss
The proving tile shows the verifier state A node that relays but never includes should say so on the tile and on /live. app/igneum-app/src/state.rs, site/api/live.mjs (verifier is already in the observer report)
Rotation phase 2 Branch rotation-2 (5317305): --dl-both, tools/logs.mjs --rotation, fresh-repo script. Plan: docs/plans/rotation-phase-2.md.
Testnet parameters behind fees_v1_activation_daa Branch testnet-prep and the fork's testnet-params (agent in progress).

Operational lessons from the activation (5 October 2026)

  • Consensus override changes must land on every node at once: a hand node restarted early with a different proving_v0_activation_daa was refused by every peer (digest handshake) and sat isolated at a lower height for 20 minutes. Order that works: publish the manifest override, update-now to every app, wait for every app node to log the new parameters, then restart the hand nodes and the seed with the same file.
  • scratchpad/restart-hand-nodes.sh died silently after igneumd --version (the 0.3.5 binary exits 1 after printing) under set -e; the restart it reported never happened. Every restart script ends by printing the new pids and their start times.
  • Switching proving on needs no app restart: POST <app.url>/api/prove {"on":true} (the job prove-on-pc2-84100 does this after installing the CUDA host into /opt/igneum for the app's WSL user).

Instant jobs (5 October 2026)

the project lead: "why is it taking so long for pc2 and pc1s tasks to spin up? can we speed it up?". Before 0.3.6 every app polled igneum-jobs.json every 10 minutes (CHECK_EVERY_S = 600), so a job published from the Mac waited up to 10 minutes on each PC. The PCs have no inbound ports and one miner is off the LAN, so the fix is a wake signal the app pulls over an outbound connection. Branch job-wake.

Piece What it does Where
Wake endpoint GET /wake?since=<stamp>: public (the apps hold no token), 30 a minute per IP, held up to 45 s, answers {stamp, at, added, changed, held_ms} the moment the stored stamp differs from since, else the unchanged stamp at the deadline. POST /r/<token>/wake {stamp, added} (the relay's auth) records a stamp. Rows in Neon relay_wake, created by the first POST. maxDuration 60 s. relay/api/wake.mjs, relay/lib/wake.mjs, relay/vercel.json, test relay/test/wake.test.mjs (fake database and clock; in CI's site job)
App waker One thread per app long-polls the endpoint with the last stamp; a changed stamp sends Event::Wake and the engine fetches the jobs file at once (signature check unchanged). First reply seeds the stamp. Backoff 5, 15, 60 s while the relay is unreachable, 10 s floor between requests, a full hold when the relay has no stamp yet. One log line per wake, one when it falls back, one when it recovers. IGNEUM_APP_JOBS_WAKE_URL overrides the URL (set and empty: no waker). app/igneum-app/src/jobrun.rs (WakeState, wake_loop), unit tests for the stamp and backoff logic
Fallback poll 2 minutes (was 10), the first poll still 40 to 60 s after start. An unchanged file is no longer logged on every poll. jobrun.rs CHECK_EVERY_S
Publisher After the live file verifies, --deploy POSTs the stamp (published_at plus 8 hex of the file's sha256) and the added id; the token goes in a header file, never on the command line; prints "woke the apps" or a one-line warning. packaging/ota/publish-jobs.sh wake_apps
Status node tools/jobs.mjs status shows "woken +N s after the publish" per machine for a job a publish added. tools/jobs.mjs

Expected latency, publish to job start: the relay re-reads the stamp every 2 s inside the hold, the app's curl returns at once, the engine fetches the file (two small downloads) and starts the job on its next tick. About 3 to 6 s when the app is mid-hold, plus up to 10 s if the app was inside its floor after an earlier reply; the 2-minute poll is the ceiling when the relay is down. Numbers below are measured, not estimated.

Machine Publish to started_at Date Source
PC 1 (ae432dc7) TODO (measure with node tools/jobs.mjs status after the first 0.3.6 job)
PC 2 (1ccfe586) TODO
Mac TODO

Not yet done on this branch: the relay deploy (relay/, by the owner; the /wake route and the 60 s maxDuration go live with it, the relay_wake table appears on the first POST), the first live publish, and the Windows curl path of the long-poll (curl.exe 8.x in System32; the 58 s --max-time was reviewed, not run).