diff --git a/CLAUDE.md b/CLAUDE.md index 40c780fe9..09f871b5b 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -84,4 +84,4 @@ Every Igneum task in Claude in Chrome runs in the Chrome profile "josh (igneum.n Josh: "this question and answer should have not needed to be asked ... so that suggestions like this get made without me" (the 12 GB mine-and-prove question, which followed from the 15.6 GB measurement and should have been raised by the agent that measured it). Rule: an agent that measures or reports a number also states, in the same report, what the number means for each user tier and what it will do about it, before anyone asks. The tiers: a home miner with one 8 GB card, one 12 GB card, one 16 GB card, one 24 or 32 GB card; a rig; a pool user; each on Windows, Linux and macOS; each on NVIDIA, AMD and Apple (Intel when it exists). A report that says "peak 15.6 GB" without "so 12 GB cards cannot do both; here is the profile that fits them, measuring now" is incomplete and goes back. The same for hash rates (per watt and per pound for the tiers), times (what deadline it fits), sizes (what disk or memory it needs), and prices. A standing reviewer agent reads every status file, bench-log entry and plan for missed consequences and opens the work; the coordinator does not wait for Josh to notice. ## A job never quits or restarts the installed app it did not start (standing rule, 5 October 2026, 23:05 UTC) -Ember Tune's PC 1 playbook started a second engine, that engine's own updater saw itself as 0.3.9 and launched the per-user installer, and the installer's stop step POSTed `/api/quit` to the installed app, taking PC 1 off the network from 22:31Z (141 MH/s gone, the 0.3.11 Windows build blocked) with nobody awake to relaunch (corrected 6 October 2026 from the collected log; the playbook's own quit never fired and pointed at its scratch engine; fixed at e600e63: a second engine never runs the updater). Rule: a test engine started by a job runs on its own port and data dir with its own URL file, and a job may quit, pause, resume or restart only an engine it started itself (the URL it created); the installed app is touched only through the signed `restart` and `update-now` job kinds. `tools/ci` carries a check (`playbook-quit-check`) that fails a playbook which reads `%LOCALAPPDATA%\igneum\app\app.url` or `~/Library/Application Support/Igneum/app/app.url` and sends quit, pause or resume to it; the reviewer's row C35 is the record. +Ember Tune's PC 1 playbook started a second engine, that engine's own updater saw itself as 0.3.9 and launched the per-user installer, and the installer's stop step POSTed `/api/quit` to the installed app, taking PC 1 off the network from 22:31Z (141 MH/s gone, the 0.3.11 Windows build blocked) with nobody awake to relaunch (corrected 6 October 2026 from the collected log; the playbook's own quit never fired and pointed at its scratch engine; fixed at e600e63: a second engine never runs the updater). Rule: a test engine started by a job runs on its own port and data dir with its own URL file, and a job may quit, pause, resume or restart only an engine it started itself (the URL it created); the installed app is touched only through the signed `restart` and `update-now` job kinds, and a job that needs the installed app's miners out of the way uses the runner's `--stop-miners` (the runner stops them before the script and restarts them on any exit), never `/api/pause` or `/api/resume` from the script, not even with a finally block (ruling 6 October 2026: a script that dies before its finally leaves the box paused unattended). `tools/ci` carries a check (`playbook-quit-check`) that fails a playbook which reads `%LOCALAPPDATA%\igneum\app\app.url` or `~/Library/Application Support/Igneum/app/app.url` and sends quit, pause or resume to it; the reviewer's row C35 is the record. diff --git a/docs/plans/morning-2026-10-06.md b/docs/plans/morning-2026-10-06.md index 220d6fb02..0ff5329d5 100644 --- a/docs/plans/morning-2026-10-06.md +++ b/docs/plans/morning-2026-10-06.md @@ -98,3 +98,14 @@ The full list with recommendations is in `consequences-decisions.md` (14) and th 1. **A fresh node never executes the EVM.** A node syncing from the seed through the pruning proof has no blocks below the pruning point, and the execution follower, which walks from genesis, sits silent: zero wallet, empty explorer, no proving work, mining fine. Every 0.3.12 joiner today is in that state; only genesis-era nodes execute. Fix: a loud "not synced" status, and an execution state snapshot at the pruning point verified against the header's state root, fetched from a peer. Stop-gap for tonight's fleet: a full sync from genesis, or a copy of the observer's execution data. 2. **The seed drops every fresh peer every 30 seconds.** The per-checkpoint certificate burst fills the finality route and the inherited rule closes a full route's connection; the baseline shows 11,700 already-known certificates resent in five minutes. Fix: a sized route, no replay during sync, drop instead of disconnect, and a guard counter. + +## Afternoon (16:00 UTC) + +| What | State | +|---|---| +| 0.3.13 | Approved by Josh at 15:20 UTC with the one-time devnet state reset: execution restarts empty at chain block 27,276 (11:40 UTC today) and re-derives forward; everything earned since is back, the first days' balances are gone; the chain, finality and the hash untouched. Carries the execution persistence and snapshot path, the fresh-joiner fix, the finality route fix (an echo of old certificates, 13,354 re-locks in 7 minutes at the seed) and Ember's helper. Two publishes, three new override fields, protocol 15 to 16, no prompt anywhere | +| Rented fleet, phase 1 | Done on 11 real cards for USD 9: every NVIDIA card from 8 GB proves; 12 GB and up mine and prove at once; 8 GB proves alone or mines beside a core-only prover (docs/analysis/prover-tiers-real-cards.md). The 8x 4090 rig is measuring on RunPod; an 8x 5090 is refused by both providers so far; no provider rents consumer AMD | +| Ember | Root cause of every failed run found and measured: the engine's own folder lock stripped the permissions off files a job copied in. Dry run with no prompt passed (5090 127 MH/s at 316 W, 4070 28.7 at 103 W). The table run goes on 0.3.13 with Josh's one click, which registers the helper and is the last prompt ever | +| PC 1 | 5090, 4070 and the 9070 XT all attached (two enclosures). Queue: 0.3.13, the Ember run, then the owed AMD rows for the class v4 candidate | +| Packaged prover server | Built and verified on PC 2 (prover-floor bf7b174): the patched server shipped signed in the app, fail-fast where it hung, a per-shard timeout with threshold step-down, tiers from the real-card rows. Ships in 0.3.14 with the public line | +| Apple | SP1 on the M5 Max CPU proves a shard in 72 s; RISC Zero's release has no Metal prover (575 s on the CPU). The Mac prover path uses the CPU; RISC Zero keeps the version slot for the 8 GB CUDA tier |