From d82de28e405cca1ac4f8be1786ac947ec0bdcfe1 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Wed, 7 Oct 2026 12:07:12 +0000 Subject: [PATCH] miner-faults MF-8: a node dying at start on a kept datadir (the serde(default) row, N13); the LG-4 job's tenth install keeps the datadir and a node restart loop fails the row Co-Authored-By: Claude Fable 5.1 --- docs/plans/miner-faults.md | 2 ++ relay/playbooks/first-share.ps1 | 11 ++++++++--- tools/fleet/first-share-gate.mjs | 7 +++++-- 3 files changed, 15 insertions(+), 5 deletions(-) diff --git a/docs/plans/miner-faults.md b/docs/plans/miner-faults.md index 5e7f0a211..c3bbbcc97 100644 --- a/docs/plans/miner-faults.md +++ b/docs/plans/miner-faults.md @@ -14,6 +14,7 @@ Standing rules behind every row (branch `miner-reliability`, off `release-0.3.19 | A card swap, a driver install or a restart needs no tap: the hot-plug enumeration every 60 s (`src/hotplug.rs`) starts the worker of a card that appears, recovers from a problem code or revives | `engine.rs` `merge_detection` → `plan_miners` | | Every fault line reports to the log intake the moment it happens, with the card, the class, the reason and the app version | `engine.rs` `fault_report` → `update::upload_text`; label `fault--`; read with `node tools/logs.mjs` | | Nothing on a user's machine is changed by a one-off script: card settings travel as the signed `cards` job kind (per card enabled, identities, power_pct), applied through the app's own card path, persisted, read back in the report, refused for a card the machine does not have | `src/jobs.rs` (`KINDS`, `validate_params`), `src/jobrun.rs` (`cards_job_choices`, `cards_applied`), `engine.rs` (`Action::ApplyCards`), `packaging/ota/publish-jobs.sh add --kind cards --cards "key=on:8"` | +| A cut never dies on a kept datadir: every stored-row schema change carries a versioned read path, and the canary and LG-4 start the new node on the previous version's datadir beside a wiped one (MF-8) | the node line's N13; `tools/fleet/first-share-gate.mjs` row "kept" (the installer over an existing install, the datadir kept); the canary form | | The fresh-install claim (LG-4) is a job, not a runbook: `tools/fleet/first-share-gate.mjs` on rented Windows boxes, on every cut, its line read by the shipper's publish | `relay/playbooks/first-share.ps1`, `tools/fleet/first-share-gate.mjs`, `site/evidence/first-share-.json` | ## The register @@ -26,6 +27,7 @@ Standing rules behind every row (branch `miner-reliability`, off `release-0.3.19 | MF-5 | 7 Oct 2026, PC 1, 0.3.17 node, 24 identities (the real cause of the 11:2x faults; MF-4 withdrawn as the cause) | Every card faulted "no status line from the miner for 90 s (restarted once already)" after the restart | Evidence (PC 1, 7 Oct 2026 11:4x UK): with 8 identities a card (24 template fetches a round) the 0.3.17 node answered no template in 5 s; with 2 a card (4 fetches) both cards mined at full rate (5090 122.4 MH/s, 9070 XT 18.9) within two minutes. The node's getBlockTemplate answered past 5 s with 24 identities fetching; the miners waited for a template inside their job-fill loop and printed no STATUS at all; the watchdog read the silence as the worker's; one restart, then faulted for good | Miner: STATUS every interval whatever the template state (`template_wait=` while it waits, `template_ms=` the node's last template time, `identities_active=`); the feed fetches only as many identities as fit one pass inside 8 s at the node's measured template time (`identities_for`: all of them when the node answers under 1 s; 24 at 8 s per template becomes 1), raised again when it answers faster; a pass whose fetches all fail backs off 2, 4, 8, 10 s and retries for ever with a `NODE SLOW` line once per 30 s. App: a STATUS with `template_wait>0` is the miner's heartbeat and the node's latency, never the card's fault (no zero-rate clock, no restart); the card reads "node slow: waiting for a block template for N s; the worker is kept" or "node slow: a template takes N s; k of n identities active"; the mitigation of the day (a one-off script POSTing /api/cards) is closed by the signed `cards` job kind | `watchdog` tests `a_slow_node_never_faults_the_card`, `parses_status_and_fault_lines` (the 0.3.20 line); `jobs` test for the `cards` kind; injector step `slow-node` (a template stub answering in 8 s while three cards run: open, needs the stub) | `app-run slow-node PASS` (0.3.20) | | MF-6 | 7 Oct 2026, PC 1, 0.3.19 | After a `--stop-miners` job, a following read-only job kept both cards "off, held for a remote job" for its whole three minutes | The engine released the hold only when no job held the miners; the next job's active state hid the release (engine.rs 3032 class) | A hold belongs to the job that took it (`job_hold_owner`) and releases the moment that job is no longer the running one, whatever runs next, or when its own cap passes (logged); a read-only job never holds (`jobrun::hold_release`) | `jobrun` test `a_hold_belongs_to_the_job_that_took_it` (owner running, another job, no job, cap passed, no owner) | `app tests: hold rule green` | | MF-7 | 7 Oct 2026, PC 1, 0.3.19 | Orphan `igneum-miner.exe` processes the app no longer tracked (two alive under `--stop-miners` with their rows at pid 0, one after) hammered the node's template RPC beside the tracked miners | The engine lost track of miners it had started (a stop that timed out, a restart over a live process) and never looked for them again | The engine owns every miner it started: at start, after every stop and every minute it kills any `igneum-miner` whose command line carries THIS engine's node RPC (the fence) and whose pid it does not track, one log line and one fault report per kill, never by name alone (`sweep_orphan_miners`, `platform::miner_processes`, `kill_pid`); a restart kills the slot's old process before the new one starts | injector step `orphan-miner` (a stray miner on the engine's node is killed inside the minute, the engine's own miner left alone) | `app-run orphan-miner PASS` | +| MF-8 | 7 Oct 2026, 0.3.19 on a kept datadir | The node binary died at start on a datadir the previous version had written; the app restarted it in a loop and every card waited | A field added to a stored row (`serde(default)` on `BlockRewardData` since 10db4b61) was read by bincode from the old row short: `DeserializationError UnexpectedEof` at `virtual_state.rs:250`; fixed on the 0.3.20 node line as N13 with a read-and-rewrite | Every schema change to a stored row ships with a versioned read path (the old shape read, rewritten in the new one); every canary and the fresh-install gate start the new node on a KEPT datadir of the previous version beside the wiped one, and a node that dies at start on a kept datadir is a FAIL of the cut, not a user's reset | The node line's N13 test (the old row read); the canary's kept-datadir start; `app-run.mjs` reads a node exit inside 10 s of its start as the engine's "igneumd exited at once" line and reports it (`node-exit` FAULT line to the intake) | `canary kept-datadir start PASS` and `LG-4` row "kept datadir" beside "wiped" | | MF-3 | 7 Oct 2026, PC 1, Intel Arc | The Intel driver's first install did not bind: the device sat in Code 12 at install time; the card never mined until a reboot | A driver installed while the device reports a problem code (12, 43, 31) does not bind; nothing re-scanned the device afterwards, and the app only re-enumerates | The app re-enumerates every 60 s and starts the worker the minute the OS drives the card (`hotplug::diff` recovered / revived, `settle_new`); the row says what to do while it does not ("reboot with the card attached; if it persists, reinstall the driver with the card attached"); a Windows host asks for a re-scan (`pnputil /scan-devices`) after a problem code is seen, every 5 minutes, at most 6 times (follow-up, host side) | `hotplug` test `a_driven_card_that_turns_faulty_is_errored_and_recovers_later`; `app-run.mjs` step `card-appears` (a card listed after 2 minutes starts without a tap) | `app-run card-appears PASS` | ## Commits (7 October 2026) diff --git a/relay/playbooks/first-share.ps1 b/relay/playbooks/first-share.ps1 index ac0ce4028..22930d1ef 100644 --- a/relay/playbooks/first-share.ps1 +++ b/relay/playbooks/first-share.ps1 @@ -13,19 +13,23 @@ param( [Parameter(Mandatory = $true)][string]$Url, [string]$Address = '0x4242424242424242424242424242424242424242', [int]$BudgetSeconds = 1800, - [switch]$AllowInstalled + [switch]$AllowInstalled, + # MF-8: install over the previous install and KEEP its datadir (the node must come up on the old version's rows) + [switch]$Keep ) $ErrorActionPreference = 'Stop' function Now { [int64]([DateTimeOffset]::UtcNow.ToUnixTimeMilliseconds()) } function Result([string]$line) { Write-Host ("RESULT " + $line) } $appData = Join-Path $env:LOCALAPPDATA 'igneum' $programs = Join-Path $env:LOCALAPPDATA 'Programs\Igneum Miner' +$kept = $false if ((Test-Path $appData) -or (Test-Path $programs)) { - if (-not $AllowInstalled) { Result "total_s=0 pass=false reason=installed_app_present (this is not a fresh box; -AllowInstalled overrides, with the owner's word)"; exit 2 } + if (-not $AllowInstalled -and -not $Keep) { Result "total_s=0 pass=false reason=installed_app_present (this is not a fresh box; -AllowInstalled overrides, with the owner's word)"; exit 2 } Get-Process -Name 'igneum-app', 'Igneum Miner', 'igneumd', 'igneum-miner' -ErrorAction SilentlyContinue | Stop-Process -Force -ErrorAction SilentlyContinue Start-Sleep -Seconds 3 - Remove-Item -Recurse -Force $appData -ErrorAction SilentlyContinue + if ($Keep) { $kept = Test-Path (Join-Path $appData 'devnet-v4') } else { Remove-Item -Recurse -Force $appData -ErrorAction SilentlyContinue } } +Result "kept=$($kept.ToString().ToLower())" # 1. download: the public installer through the dl host, timed from the first byte $setup = Join-Path $env:TEMP 'Igneum-Miner-Setup-first-share.exe' Remove-Item $setup -ErrorAction SilentlyContinue @@ -64,6 +68,7 @@ $deadline = $tDl0 + ($BudgetSeconds * 1000) while ((Now) -lt $deadline) { try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -TimeoutSec 5 } catch { Start-Sleep -Seconds 2; continue } if (-not $synced -and $st.node.synced) { $synced = Now; $stall = 'card'; Result "step=node end=$synced synced_at=$($st.ladder.first_synced_at)" } + if ($st.node.restarts -ge 2 -and -not $synced) { Result "step=node restarts=$($st.node.restarts) message=$($st.node.message)"; Result "total_s=$([int](((Now) - $tDl0) / 1000)) pass=false stalled_in=node reason=node_restart_loop kept=$($kept.ToString().ToLower())"; break } if ($st.ladder -and $st.ladder.first_mining_at -and -not $mining) { $mining = $st.ladder.first_mining_at } if ($st.ladder -and $st.ladder.first_block_at) { $first = Now; $hash = $st.ladder.first_block_hash; break } if ($st.ladder -and $st.ladder.first_share_at) { $first = Now; $hash = 'share'; break } diff --git a/tools/fleet/first-share-gate.mjs b/tools/fleet/first-share-gate.mjs index 57f16c8d4..abf502daa 100755 --- a/tools/fleet/first-share-gate.mjs +++ b/tools/fleet/first-share-gate.mjs @@ -11,6 +11,8 @@ // with no row the gate writes "LG-4 not run: no Windows box" and `check` fails. Each box runs installs in turn; boxes run in // parallel. Output: site/evidence/first-share-.json (the runbook's keys per row, the medians, the pass count, the // installer's sha256) and the gate line on stdout and in the file: "LG-4 PASS k/10" or "LG-4 FAIL k/10". +// The tenth install of a run is the KEPT-datadir case (MF-8): the installer runs over the ninth install with its datadir +// kept, and the node must come up on it; that row carries kept=true. The other nine are wiped. // Never touches PC 1 (the project lead's desk) and refuses a box whose installed app would be replaced unless --allow-installed is given // and the box's row says "job_box": true (PC 2's case needs the project lead's word, recorded in the row by the fleet lane). @@ -74,14 +76,15 @@ const started = Date.now(); // boxes in parallel (one ssh each, installs in turn inside it), the Mac waits const procs = boxes.map((b) => { const n = Math.min(perBox, installs - rows.length); - const remote = `$ErrorActionPreference='Continue'; $f=\"$env:TEMP\\first-share.ps1\"; [IO.File]::WriteAllText($f, [Text.Encoding]::UTF8.GetString([Convert]::FromBase64String('${Buffer.from(readFileSync(join(ROOT, 'relay', 'playbooks', 'first-share.ps1'))).toString('base64')}'))); 1..${n} | ForEach-Object { Write-Host \"INSTALL $_\"; powershell -ExecutionPolicy Bypass -File $f -Url '${url}' ${allowInstalled && b.job_box ? '-AllowInstalled' : ''} }`; + // the last install on the first box keeps the previous install's datadir (-Keep): the MF-8 row + const remote = `$ErrorActionPreference='Continue'; $f=\"$env:TEMP\\first-share.ps1\"; [IO.File]::WriteAllText($f, [Text.Encoding]::UTF8.GetString([Convert]::FromBase64String('${Buffer.from(readFileSync(join(ROOT, 'relay', 'playbooks', 'first-share.ps1'))).toString('base64')}'))); 1..${n} | ForEach-Object { Write-Host \"INSTALL $_\"; $keep = if ($_ -eq ${n} -and ${b === boxes[0] ? '$true' : '$false'}) { '-Keep' } else { '' }; powershell -ExecutionPolicy Bypass -File $f -Url '${url}' ${allowInstalled && b.job_box ? '-AllowInstalled' : ''} $keep }`; const key = (b.key || '~/.ssh/igneum-fleet').replace(/^~/, homedir()); const p = spawnSync('ssh', ['-o', 'ConnectTimeout=20', '-o', 'StrictHostKeyChecking=accept-new', '-i', key, '-p', String(b.port || 22), `${b.user}@${b.host}`, 'powershell', '-NoProfile', '-Command', remote], { encoding: 'utf8', timeout: (n * 1900 + 120) * 1000 }); return { box: b, out: (p.stdout || '') + (p.stderr || ''), status: p.status }; }); for (const { box, out } of procs) { for (const chunk of out.split(/^INSTALL \d+\r?$/m).slice(1)) { - const r = parse(chunk); r.box = box.label; r.gpu = box.gpu || ''; rows.push(r); + const r = parse(chunk); r.box = box.label; r.gpu = box.gpu || ''; r.kept = /kept=true/.test(chunk); rows.push(r); } if (!out.includes('RESULT')) rows.push({ box: box.label, gpu: box.gpu || '', pass: false, total_s: 0, error: out.trim().split('\n').slice(-3).join(' | ') }); }