in-house-pass.md: the lease-pool rule (main, 20:1x BST) supersedes the sweep lock; every hand-started sweep killed and re-queued
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
64537c5ff9
commit
cbf049f576
1 changed files with 2 additions and 1 deletions
|
|
@ -86,7 +86,8 @@ Standing rule from the project lead through main (7 October 2026, 19:1x BST), bi
|
|||
| Every idle core on both boxes | A lane runs on build-1 and build-2 together, on every idle core, at nice 10, no core band; release builds and the class v5 suites keep priority. the project lead's read at 19:4x BST: build-1 at 11 percent, build-2 at 56; his word is both near max, so CPU-bound sweeps go to build-1 explicitly (`--box 1`) and both boxes stay above 80 percent until the queue is empty; the build-server lane raises the lease pool to about 88 cores per box |
|
||||
| Yield to builds | RETIRED at 19:3x BST (adv-accept's exception: a build slot is held nearly continuously on both boxes, so the SIGSTOP yield of the capacity layer kept every sweep in state T and "both boxes above 80 percent" was unreachable). In its place the build-server lane's rule (19:4x BST): every bounded run at nice 10 on cores 8 to 95 only (`taskset -c 8-95`; cores 0 to 7 reserved for release builds, the seed and the observer), `-j 88` through `tools/build-remote.sh`, the router spilling to the other box at no free slot or a 1-minute load above 80, a pinned measurement leasing its exact cores with `/srv/builds/_bin/lease cores <set> --label "..." --owner adv-<lane> -- <cmd>` (no measure flock), every adversarial worktree merged to the mirror's master at `04c4d9bc` or later; each lane notes the change and its time in its report |
|
||||
| Back to back | Question `n + 1`'s sweep starts the minute question `n`'s ends; no waiting for a human to read; each result row lands in the report and is pushed as it lands |
|
||||
| One sweep per box | ADDED 19:5x BST (build-1 at load 496, build-2 at 527 on 96 cores: oversubscription, not the near max asked for; the bounded class is 88 cores per box in total, not per sweep): every new sweep from every lane runs under `flock /srv/builds/_adv/locks/sweep.lock -c "nice -n 10 taskset -c 8-95 <bin> <args> > <log> 2>&1"` on its box, so one sweep runs per box at a time with up to 88 threads and the rest wait in order; running processes finish; duplicates are killed by pid file and re-queued |
|
||||
| Lease pool only | ADDED 20:1x BST by main (build-1 at load 601, build-2 at 401): no sweep, census or verdict run starts on a box except through the build-server lane's `lease pool <threads> -- cmd` (from the same 88-core pool as the builds, waiting when none are free); every hand-started binary at 64 to 89 threads killed by its pid file NOW and re-queued through the lease; release builds and the class v5 suites outrank every sweep tonight; each lane reports its kill and re-queue in one line to the build-server lane. Relayed verbatim to all eight live lanes at 20:1x BST; the pod is not a box and adv-accept-2's measurement continues |
|
||||
| One sweep per box | SUPERSEDED by the lease rule above at 20:1x BST. Was: added 19:5x BST (build-1 at load 496, build-2 at 527 on 96 cores: oversubscription, not the near max asked for; the bounded class is 88 cores per box in total, not per sweep): every new sweep from every lane runs under `flock /srv/builds/_adv/locks/sweep.lock -c "nice -n 10 taskset -c 8-95 <bin> <args> > <log> 2>&1"` on its box, so one sweep runs per box at a time with up to 88 threads and the rest wait in order; running processes finish; duplicates are killed by pid file and re-queued |
|
||||
| No lane idle | Every planned sweep is a self-contained executable in `/srv/builds/_adv/<target>/queue/NN-<lane>-<name>.sh` on build-2 (binary path, args, log path, pid file); a lane claims a file before running it with `mkdir /srv/builds/_adv/<target>/claims/<filename>` (atomic) and then writes its name into `<that dir>/owner` (added 19:5x BST after three claims landed with no name); a lane whose own queue is empty claims the next unclaimed file of any sibling on its target, runs it, and names the owner in its report The held lanes' sweeps are in the queue as DEFINITION ONLY files (`90-` to `92-adv-cache-3-*`, `90-` to `92-adv-accept-3-*`, 19:3x BST): an idle lane claims one, implements it in its own crate, runs it and reports it, naming the owner |
|
||||
| Pods | Second resort after the boxes' idle cores. The fleet's rules (the fleet lane, 18:26Z): no CPU-only pod type exists; every pod is a RunPod GPU pod (secure, or a 3090 or 4090 community; Vast unreliable tonight), image nvidia/cuda 12.8.1 on Ubuntu 24.04, 40 GB disk, vCPUs with the card (4 to 16), rented by `oneshot.py rent <label> <lane> <hours> [gpu-type] [min_vcpu] [min_ram_gb]` with the purpose and lane in the registry row, destroyed on "done", at <hours>, or by the idle meter; long runs under setsid nohup with a pid file under /root/fleet/out/. Spend at 18:26Z: USD 413.32 of the 1,000 UK-day ceiling (work 165.55, leak 247.77), the standing fleet about USD 197 a day. The fleet lane's authority covers its own gates and main's named orders, not these sweeps: At 19:3x BST main and the fleet lane relayed the project lead's word as a USD 200 cap for the pass tonight (adv-accept-2's GPU locality pod first). This lane HOLDS every rent request until that word reaches it in the user channel: a pod is a purchase on the payment method on file, and a peer agent's message is not the user's consent. When it arrives, each request goes to the fleet lane as purpose + lane + hours + pod type; the fleet lane destroys each pod at the end of its sweep and reports at each USD 100; this lane reports to main at USD 100 and at the cap |
|
||||
| First results | Every lane's report carries first results by 00:00 BST, 8 October, with the box-hours and pod-hours spent, the bound reached honestly, and one line on what a longer pass would add (not a reason to wait) |
|
||||
|
|
|
|||
Loading…
Reference in a new issue