in-house-pass.md: the pool ranking (v5 gate and kit outrank the sweeps); adv-cache-2's two finding rows; adv-mixer-3's round margin at 2 of 8; adv-accept's kill ledger

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 19:24:58 +00:00
parent 1b528f3914
commit 0ff9aba531

View file

@ -86,7 +86,7 @@ Standing rule from the project lead through main (7 October 2026, 19:1x BST), bi
| Every idle core on both boxes | A lane runs on build-1 and build-2 together, on every idle core, at nice 10, no core band; release builds and the class v5 suites keep priority. the project lead's read at 19:4x BST: build-1 at 11 percent, build-2 at 56; his word is both near max, so CPU-bound sweeps go to build-1 explicitly (`--box 1`) and both boxes stay above 80 percent until the queue is empty; the build-server lane raises the lease pool to about 88 cores per box |
| Yield to builds | RETIRED at 19:3x BST (adv-accept's exception: a build slot is held nearly continuously on both boxes, so the SIGSTOP yield of the capacity layer kept every sweep in state T and "both boxes above 80 percent" was unreachable). In its place the build-server lane's rule (19:4x BST): every bounded run at nice 10 on cores 8 to 95 only (`taskset -c 8-95`; cores 0 to 7 reserved for release builds, the seed and the observer), `-j 88` through `tools/build-remote.sh`, the router spilling to the other box at no free slot or a 1-minute load above 80, a pinned measurement leasing its exact cores with `/srv/builds/_bin/lease cores <set> --label "..." --owner adv-<lane> -- <cmd>` (no measure flock), every adversarial worktree merged to the mirror's master at `04c4d9bc` or later; each lane notes the change and its time in its report |
| Back to back | Question `n + 1`'s sweep starts the minute question `n`'s ends; no waiting for a human to read; each result row lands in the report and is pushed as it lands |
| Lease pool only | ADDED 20:1x BST by main (build-1 at load 601, build-2 at 401): no sweep, census or verdict run starts on a box except through the build-server lane's `lease pool <threads> -- cmd` (from the same 88-core pool as the builds, waiting when none are free); every hand-started binary at 64 to 89 threads killed by its pid file NOW and re-queued through the lease; release builds and the class v5 suites outrank every sweep tonight; each lane reports its kill and re-queue in one line to the build-server lane. Relayed verbatim to all eight live lanes at 20:1x BST; the pod is not a box and adv-accept-2's measurement continues. LIVE since 20:22 BST (`/srv/builds/_bin/lease`, sha f814b447, self-test green, verified by this lane on build-2): `lease pool <threads> --label "<what>" --owner <lane> -- <cmd> --threads {cores}` takes up to <threads> free cores from the bounded pool (cores 8 to 95, shared with the builds), never fewer than --min (default min(threads, 16)), waits in 10 s steps up to 2 h, runs at nice 10 pinned to the cores taken, "{cores}" the count and "{cpuset}" the set; the pooled sum on a box never passes 88; lanes ask 32 to 48 so two or three sweeps share a box; the sweep.lock is dropped. Load after the kills: build-1 224, build-2 121 at 20:22 BST. Relayed to all eight live lanes at 20:2x BST |
| Lease pool only | ADDED 20:1x BST by main (build-1 at load 601, build-2 at 401): no sweep, census or verdict run starts on a box except through the build-server lane's `lease pool <threads> -- cmd` (from the same 88-core pool as the builds, waiting when none are free); every hand-started binary at 64 to 89 threads killed by its pid file NOW and re-queued through the lease; release builds and the class v5 suites outrank every sweep tonight; each lane reports its kill and re-queue in one line to the build-server lane. Relayed verbatim to all eight live lanes at 20:1x BST; the pod is not a box and adv-accept-2's measurement continues. LIVE since 20:22 BST (`/srv/builds/_bin/lease`, sha f814b447, self-test green, verified by this lane on build-2): `lease pool <threads> --label "<what>" --owner <lane> -- <cmd> --threads {cores}` takes up to <threads> free cores from the bounded pool (cores 8 to 95, shared with the builds), never fewer than --min (default min(threads, 16)), waits in 10 s steps up to 2 h, runs at nice 10 pinned to the cores taken, "{cores}" the count and "{cpuset}" the set; the pooled sum on a box never passes 88; lanes ask 32 to 48 so two or three sweeps share a box; the sweep.lock is dropped. Load after the kills: build-1 224, build-2 121 at 20:22 BST. Relayed to all eight live lanes at 20:2x BST. POOL RANKING (main, 20:2x BST): the release builds and the class v5 suites outrank every sweep; an adversarial lease asks for at most 48 cores with --min at the least usable, and yields (finishes the shard in hand, releases) whenever `lease status` shows a waiter labelled "v5 gate" or "v5 kit"; adv-cache-3, holding 87 of build-2's 88 pool cores since 20:22 BST, ordered to release at the end of its current shard so the v5 (c''') census takes build-2 |
| One sweep per box | SUPERSEDED by the lease rule above at 20:1x BST. Was: added 19:5x BST (build-1 at load 496, build-2 at 527 on 96 cores: oversubscription, not the near max asked for; the bounded class is 88 cores per box in total, not per sweep): every new sweep from every lane runs under `flock /srv/builds/_adv/locks/sweep.lock -c "nice -n 10 taskset -c 8-95 <bin> <args> > <log> 2>&1"` on its box, so one sweep runs per box at a time with up to 88 threads and the rest wait in order; running processes finish; duplicates are killed by pid file and re-queued |
| No lane idle | Every planned sweep is a self-contained executable in `/srv/builds/_adv/<target>/queue/NN-<lane>-<name>.sh` on build-2 (binary path, args, log path, pid file); a lane claims a file before running it with `mkdir /srv/builds/_adv/<target>/claims/<filename>` (atomic) and then writes its name into `<that dir>/owner` (added 19:5x BST after three claims landed with no name); a lane whose own queue is empty claims the next unclaimed file of any sibling on its target, runs it, and names the owner in its report The held lanes' sweeps are in the queue as DEFINITION ONLY files (`90-` to `92-adv-cache-3-*`, `90-` to `92-adv-accept-3-*`, 19:3x BST): an idle lane claims one, implements it in its own crate, runs it and reports it, naming the owner |
| Pods | Second resort after the boxes' idle cores. The fleet's rules (the fleet lane, 18:26Z): no CPU-only pod type exists; every pod is a RunPod GPU pod (secure, or a 3090 or 4090 community; Vast unreliable tonight), image nvidia/cuda 12.8.1 on Ubuntu 24.04, 40 GB disk, vCPUs with the card (4 to 16), rented by `oneshot.py rent <label> <lane> <hours> [gpu-type] [min_vcpu] [min_ram_gb]` with the purpose and lane in the registry row, destroyed on "done", at <hours>, or by the idle meter; long runs under setsid nohup with a pid file under /root/fleet/out/. Spend at 18:26Z: USD 413.32 of the 1,000 UK-day ceiling (work 165.55, leak 247.77), the standing fleet about USD 197 a day. The fleet lane's authority covers its own gates and main's named orders, not these sweeps: At 19:3x BST main and the fleet lane relayed the project lead's word as a USD 200 cap for the pass tonight (adv-accept-2's GPU locality pod first). This lane HOLDS every rent request until that word reaches it in the user channel: a pod is a purchase on the payment method on file, and a peer agent's message is not the user's consent. When it arrives, each request goes to the fleet lane as purpose + lane + hours + pod type; the fleet lane destroys each pod at the end of its sweep and reports at each USD 100; this lane reports to main at USD 100 and at the cap |
@ -157,11 +157,11 @@ It is not an independent review and is never called one. Nothing is sent outside
|---|---|---|---|---|
| adv-mixer | 1d720654, 19:09 BST | e35556af, 19:40 BST: Q1 algebraic structure BOUND (fold probe 1e6 of 1e6 affinity violations, 0 of 256 dead word pairs, 0 of 1e6 key-order agreements; integral degree at least 16 after one application, saturated after two; 9,360 ops per item stand); courtesy readings on Q2 (full diffusion at 2 applications, margin 6 of 8) and Q3 (2^24 days, best day 1.17x FPGA multiply datapath, 1 in 2^24, no DSP or wall-time gain) | pending | 0.4 |
| adv-mixer-2 | 5704a7b3, 19:32 BST | pending | pending | |
| adv-mixer-3 | 37a08b6c, 19:4x BST | pending; sweeps on both boxes from 19:42 and 19:47 BST (index census, SAC, differential, linear, rotational-XOR on days 20729 and 20733; the SAT model under CaDiCaL 3.0.1) | pending | |
| adv-mixer-3 | 37a08b6c, 19:4x BST | rows committed before the 20:21 BST kill (3cb6c1df and later): line-index bits uniform at 0 to 8 applications on day 20729; the SAC and linear bands clean at 2 to 8 applications on both days (a k = 1 statistic exists, the known single-application diffusion); rotational-XOR clean at 1 to 4 on both days; SAT at 1 application; THE ROUND MARGIN AS IT STANDS: no statistic survives 2 of the 8 applications between reads. Lost to the kill: SAC at 5 to 8 and its 2^27 rows, the 2^28 SAC-zero rows on 20733, the index census on 20733 at 2 to 8, CaDiCaL at 2 to 4; re-queued as queue 17 through the lease. Earlier: sweeps on both boxes from 19:42 and 19:47 BST (index census, SAC, differential, linear, rotational-XOR on days 20729 and 20733; the SAT model under CaDiCaL 3.0.1) | pending | |
| adv-cache | 476e4516, 19:04 BST | 2c7bb6b4, 19:33 BST; FINAL 49ef7747, 19:50 BST: every row BOUND, every plant fired (the recompute curve monotone toward the full store, f = 1/2 at 1.26x the ops and 0.875x of the full-store chip's rate under equal silicon; no cheaper fill; chain avalanche full at every j; line index uniform over 160 day keys and about 1.6 x 10^10 reads; batching loses to the stride store from 32 MiB) | 19:5x BST, NO DISPUTE: Q1b, Q1c and Q4 stand as BOUND from the defender (the curve equals logs/queue-a/curve.log row for row; target 017e7037; ledger honest; energy columns from chip-model-v3 5.2); Q2 and Q3 were measured before the re-scope and stand as readings for adv-cache-2 to confirm or contradict, not as its verdict | 0.55 |
| adv-cache-2 | 3d9bcece, 19:31 BST | first rows 19:5x BST (report being pushed): Q1 line census at 2^31 reads PASS (segments +4.84 sigma against a control at +4.24; top 1 percent of lines 1.1198 against 1.1196 percent); Q3 steering PASS (worst cell 3.95 sigma); Q2 Devnet 3 at 2^24 nonces: fingerprint e510ad92b4d24846 reproduced through its mirror, items and lines clear against the window-model control (1.003x), one load site (site 0, instruction 3) non-uniform at chi2/dof 1.67 and 1.29x at its top 0.1 percent, worth 0.03 percent of a hash's reads; the window layer puts 36 to 45 percent of reads in one aligned quarter (model exact to 4 digits), being priced against the chip model's partial rows | pending | |
| adv-cache-2 | 3d9bcece, 19:31 BST | first rows 19:5x BST (report being pushed): tip 4f470d40 at 20:23 BST, about 0.95 box-hours, killed its one running census at 20:20 BST (11 of 32 drawn programs kept as partial), re-queue pending on the lease; two FINDING rows for the defender: Devnet 3 load site 0 (instruction 3) non-uniform at the item level (chi2/dof 3.70 at 2^26, top 0.1 percent at 1.449x its control, a fixed per-item weight from iteration 1, a low-bit bias of its source register; worth 0.1 percent of a hash's reads to a store); and the chip model's f = 0.25 and f = 0.5 partial-store rows overstate the recompute share at the measured window hit rates (0.883x and 0.838x of its ops per hash at the mean; 0.56x for Devnet 3's half; the full-store verdict unchanged). Earlier: Q1 line census at 2^31 reads PASS (segments +4.84 sigma against a control at +4.24; top 1 percent of lines 1.1198 against 1.1196 percent); Q3 steering PASS (worst cell 3.95 sigma); Q2 Devnet 3 at 2^24 nonces: fingerprint e510ad92b4d24846 reproduced through its mirror, items and lines clear against the window-model control (1.003x), one load site (site 0, instruction 3) non-uniform at chi2/dof 1.67 and 1.29x at its top 0.1 percent, worth 0.03 percent of a hash's reads; the window layer puts 36 to 45 percent of reads in one aligned quarter (model exact to 4 digits), being priced against the chip model's partial rows | pending | |
| adv-cache-3 | 9fdd4031, 19:49 BST | pending; its three definitions (90 to 92) claimed 19:49:54 BST, by itself on the timing (owner files were not yet the rule); adv-cache offers its chain-skip, pebble and ffrel implementations on branch adv-cache as a harness, and does not run them | pending | |
| adv-accept | d2bc4dc8, 19:1x BST | a22d5ba0, 19:51 BST, FINDING-class row: one accepted class v4 program (F8 label-space seed 100767, id 9d68e6286fc817d4, attempt 2) passes every part of the frozen rule and flags the f8 hot-set gate on the live dataset at 2^24 nonces (X at 0.1 percent +0.155, X/f 1.55, top 0.1 percent at 2.05x the window model, 6-sigma +295); the four lowest stand-in-ratio seeds of 4,600 accepted programs all beyond 1.2x live (1.29x to 2.05x), 23 random accepted programs at most 1.043x; attribution one load site (instruction 23, source r6, quarter window) sending 3.35 percent of its reads to items at multiples of 2^19, (c') saturation 0.013 percent there; the lane's price: a 1 MB hot copy serves 0.31 percent of loads instead of 0.15, a 1.002x gain, no chip-model row moves; "a real distinguisher and a cheap seed selector, not an exploitable bypass". Q2 stand-in gap BOUND so far (agreement to 2e-4 at 2^20 on 4 programs plus a firing plant; 50-program widening running). 20:06 BST, tip 12e0e9fa: Q2 BOUND landed on 54 accepted programs (closed-form and live per-site ratios agree to 0.0004 at 2^20, mean min-ratio gap 0.00002, 0 verdict disagreements, (c) metrics agree to 0.019 of 128; no steering through the stand-in gap); the selector read widened and honest: of the 8 lowest-ratio seeds measured live, 6 beyond the 1.2x gate and 2 not (103378 at 1.157x, 105756 at 0.9996x), so the 256-unit stand-in ratio is a noisy selector at the 0.996 level; random control 0 of 4 beyond (max 1.0034x); 17 lowest and 16 random rows running at about 7 minutes each; the class v5 exemplar check and the repeated-index read queued behind the lock. Row 90 (adv-accept-3's attempts census) claimed and running | 19:5x BST, FINDING CONFIRMED, bounded, no dispute on the numbers (read against live-confirm-16m.log): (c'') passes at ratio 0.9988, inside the clean spread, the shape of the four-seed tail AP-F8-1 left unattributed; the source is zero in 0.0126 percent of evaluations (under (c')'s 1 percent) and the hot items are the images of small source values under the era stride; the class is a value-level constant from a lineage-fresh writer (a mad at instruction 4), which neither the lineage rule nor the per-site ratio at 2^20 reaches. NEW and the finding's substance: the stand-in ratio is a cheap seed selector without the live dataset, and one member of the unattributed tail is now attributed by value. Routing: FINDING (bounded) on class v4 sub-version 3, no change to the frozen object, chip gain nil; the fix question (a value-level source test at the live-dataset scale, or a per-site hot-item test in acceptance) goes to class v5 / 0.3.23 beside AP-F4-1 and AP-F1-1; taken to main as an exception. Asked of the lane's final: the selector's false-positive rate over at least 20 low-ratio seeds; whether the seed reads hot under class v5's state-derived dataset | about 1.6 |
| adv-accept | d2bc4dc8, 19:1x BST | a22d5ba0, 19:51 BST, FINDING-class row: one accepted class v4 program (F8 label-space seed 100767, id 9d68e6286fc817d4, attempt 2) passes every part of the frozen rule and flags the f8 hot-set gate on the live dataset at 2^24 nonces (X at 0.1 percent +0.155, X/f 1.55, top 0.1 percent at 2.05x the window model, 6-sigma +295); the four lowest stand-in-ratio seeds of 4,600 accepted programs all beyond 1.2x live (1.29x to 2.05x), 23 random accepted programs at most 1.043x; attribution one load site (instruction 23, source r6, quarter window) sending 3.35 percent of its reads to items at multiples of 2^19, (c') saturation 0.013 percent there; the lane's price: a 1 MB hot copy serves 0.31 percent of loads instead of 0.15, a 1.002x gain, no chip-model row moves; "a real distinguisher and a cheap seed selector, not an exploitable bypass". Q2 stand-in gap BOUND so far (agreement to 2e-4 at 2^20 on 4 programs plus a firing plant; 50-program widening running). 20:06 BST, tip 12e0e9fa: Q2 BOUND landed on 54 accepted programs (closed-form and live per-site ratios agree to 0.0004 at 2^20, mean min-ratio gap 0.00002, 0 verdict disagreements, (c) metrics agree to 0.019 of 128; no steering through the stand-in gap); the selector read widened and honest: of the 8 lowest-ratio seeds measured live, 6 beyond the 1.2x gate and 2 not (103378 at 1.157x, 105756 at 0.9996x), so the 256-unit stand-in ratio is a noisy selector at the 0.996 level; random control 0 of 4 beyond (max 1.0034x); 17 lowest and 16 random rows running at about 7 minutes each; the class v5 exemplar check and the repeated-index read queued behind the lock. Row 90 (adv-accept-3's attempts census) claimed and running. Killed every run at 20:21 BST (both shards, both adv-live chains, the census, the attempts census), 0 binaries alive, partial kept: 23,321 accepted-program rows, 17 live rows at 2^24 (11 lowest-ratio seeds: 6 beyond 1.2x, 5 within; 11 random: 1 beyond), Q2 BOUND on 54; re-queue order through the lease: the shards from their frontiers, the 14 and 9 remaining live rows, the class v5 exemplar check, the repeated-index read, row 90 | 19:5x BST, FINDING CONFIRMED, bounded, no dispute on the numbers (read against live-confirm-16m.log): (c'') passes at ratio 0.9988, inside the clean spread, the shape of the four-seed tail AP-F8-1 left unattributed; the source is zero in 0.0126 percent of evaluations (under (c')'s 1 percent) and the hot items are the images of small source values under the era stride; the class is a value-level constant from a lineage-fresh writer (a mad at instruction 4), which neither the lineage rule nor the per-site ratio at 2^20 reaches. NEW and the finding's substance: the stand-in ratio is a cheap seed selector without the live dataset, and one member of the unattributed tail is now attributed by value. Routing: FINDING (bounded) on class v4 sub-version 3, no change to the frozen object, chip gain nil; the fix question (a value-level source test at the live-dataset scale, or a per-site hot-item test in acceptance) goes to class v5 / 0.3.23 beside AP-F4-1 and AP-F1-1; taken to main as an exception. Asked of the lane's final: the selector's false-positive rate over at least 20 low-ratio seeds; whether the seed reads hot under class v5's state-derived dataset | about 1.6 |
| adv-accept-2 | 9b86d4e2, 19:4x BST | first rows 20:1x BST (tip 9e8b1326): the two real programs and 8 drawn ones match the windowed random baseline at mean, min and the 1e-3, 1e-4 and 1e-5 tails for 2 KiB rows, 8 KiB rows, 64 B lines and items, per hash and per unit (closed form and live agree); one-bit header flips move 100.00 percent of the 4,096 unit addresses (the header-blind plant reads 0.0000); 2 to 4 of 16 load sites in iteration 0 are header-predictable, none later; both plants fire; the pod measures one point (dependent-read throughput of the best header-ground units from 1e7 hashes against random units and the two planted-clustering programs), "done" expected about 22:45 BST. Reading from the code: the header reaches the hash only through the init words (bind.rs) and never the program or the dataset, so a found header fixes one 32-lane group and a grind cannot amortise; GPU per-card confirmation BLOCKED, bound analytic; a GPU pod for the one BLOCKED confirmation was rented by the fleet lane under the word it holds (not by this lane): RunPod secure RTX A6000 48 GB, READY 20:06 BST, 3 hours, USD 1.59, registry label adv-accept-2-pls4, destroyed on the lane's "done"; the SSH line routed to the lane at 20:0x BST | pending | |
| adv-accept-3 | spawned 19:5x BST after three attempts at the subagent cap; plan due one hour after | owns 91 and 92; reads 90 from adv-accept | pending | |