bench log + plan: C35 corrected (the quit was not the 0.3.11 update; what is established, the hang, the orphans, the prompt, the fixes)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 23:00:52 +00:00
parent d19441c14d
commit f9d9805ae8
2 changed files with 4 additions and 4 deletions

View file

@ -1624,7 +1624,7 @@ Branch `ember-tune` (54ff1bc), docs/plans/ember-tune.md. Every card tuned for MH
**Tier consequences** (docs/plans/ember-tune.md section 7): a 9-step full tune costs about 12 minutes once and 3 minutes a week per card, under 1% of the hour, the worker never stops; a rig tunes one card at a time and every card of a known model after the first takes the 3-minute confirm; a pool user gives up the same 1% of shares at most; Apple silicon and AMD on Linux measure only and the row says so.
**The PC 1 run, 22:30 UTC (job ember-tune-pc1-1, engine aeea3228..., PC 1 on 0.3.10):** the job published at 22:29:40Z, the installed app stopped its miners and started the second engine at 22:30:21Z, and at 22:31:06Z the installed app quit for the 0.3.11 update-now (its log: `job ember-tune-pc1-1: aborted (the app is quitting)`), taking the second engine with it 47 s in, before any step. Nothing was set. What the run did record, the "before" snapshots with the miners stopped:
**The PC 1 run, 22:30 UTC (job ember-tune-pc1-1, engine aeea3228..., PC 1 on 0.3.10):** the job published at 22:29:40Z, the installed app stopped its miners and started the second engine at 22:30:21Z, and at 22:31:06Z the installed app quit (its log: `quit: stopping the miners, then the node`, then `job ember-tune-pc1-1: aborted (the app is quitting)`), 46 s in, before any step. Nothing was set. Corrected the same night (C35): the first reading, that the 0.3.11 update caused the quit, was wrong; no update, restart or relay task reached PC 1 then (its own jobs lines and the relay feed), and the quit's source is not in the log because the app did not name it (fixed at b671c8b: every `quit:` line now names its sender). What is established: the window host's tray quit is excluded (a host that sent the quit terminates the engine 45 s later, and the engine lived on until 22:55Z), leaving stdin EOF (the host process gone) or `POST /api/quit`; the engine's quit then HUNG for 24 minutes in the jobs runner's abort, waiting for EOF on the script's stdout pipe whose write end the second engine and its miners had inherited, and those miners (2 igneum-miner, 2 CUDA workers, 1 OpenCL worker) mined on, orphaned, until the relay lane killed them at about 23:00Z; the second engine also raised one administrator prompt at about 22:30:25Z (`apply_power_limits` at start counted `--sweep` as Power control), 41 s before the quit; PC 2's unexplained quit at 20:01:09Z came 20 s after a cancelled prompt of the same class, so the prompt is the common factor and the morning's test (one prompt raised beside the mining app on PC 2, the stamped quit line read). Fixed on the branch: b671c8b (quit sources, Power control alone decides, no cap at start under `--sweep`), 8ab9068 (no pipe into a second engine, its tree ended, the CI check). What the run did record, the "before" snapshots with the miners stopped:
| Card | Read back at 22:30:20Z | Meaning |
|---|---|---|
@ -1632,5 +1632,5 @@ Branch `ember-tune` (54ff1bc), docs/plans/ember-tune.md. Every card tuned for MH
| RX 9070 XT (bus 98, present again) | `tune 1 ... gmax 0 gmax_range -500 1000 plimit 0 plimit_range -30 10 factory 1 ok` | the helper's clock range is an OFFSET from stock in MHz, not a ceiling: a probe reading it as a 1,000 MHz maximum would have asked for `--set-gmax 900`, an overclock. Fixed at 054e041: an offset range closes the clock knob (until the stock clock is known) and the power ladder runs on the percent scale bounded by the range, so the 9070 XT's plan is 100, 90, 80, 70% (the -30 floor), 4 steps |
| Radeon(TM) Graphics (integrated) | `tune 0 ... gmax - ... factory 0 ok` | no manual tuning: measure only, and it is off by default anyway |
Consequence for the tiers: an AMD card is tuned on its power limit alone until its stock core clock is read (a 9070 XT at -30% is the floor the driver allows, 4 steps, 5 minutes); every NVIDIA card's two-knob plan waits on the user's one click on Power control; the re-run on PC 1 follows the 0.3.11 rollout (the update clears the jobs folder, so the engine and the helper are fetched again), with the scheduler's slot.
Consequence for the tiers: an AMD card is tuned on its power limit alone until its stock core clock is read (a 9070 XT at -30% is the floor the driver allows, 4 steps, 5 minutes); every NVIDIA card's two-knob plan waits on the user's one click on Power control; the re-run on PC 1 is held until the quit's source is named (the event-log collect) and follows the 0.3.11 rollout (the update clears the jobs folder, so the engine and the helper are fetched again), with the scheduler's slot.

View file

@ -153,8 +153,8 @@ measure, TUNE record, upload, aggregation, prior shape in a test manifest). The
9070 XT run are owed: the 5090 the moment Power control is switched on (one prompt, then the tune runs by itself
within 2 minutes of steady mining), the 9070 XT when the card is back on the bus.
Run 1 (ember-tune-pc1-1, 22:30 UTC): aborted 47 s in by the installed app quitting for the 0.3.11 update-now, before
any step; nothing set; the "before" snapshots are in the bench log (5090: 450 W of 575, 2,850 MHz core, 3,090 MHz
Run 1 (ember-tune-pc1-1, 22:30 UTC): aborted 46 s in by the installed app quitting (source unnamed by the 0.3.10 app;
not an update, not a job, not a relay task: C35 in the bench log), before any step; nothing set; the "before" snapshots are in the bench log (5090: 450 W of 575, 2,850 MHz core, 3,090 MHz
maximum, 14,001 MHz memory; 9070 XT present on bus 98 with OFFSET ranges `gmax_range -500 1000`, `plimit_range -30
10`). The offset finding changed the AMD mapping (054e041): an offset clock range closes the clock knob and the power
ladder runs on a percent scale bounded by `plimit_range`. The re-run follows the 0.3.11 rollout.