Consequences ledger: C24 and C27 closed on branch bash-body-check (7adb1ca, 6805125) with the merge notes

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 22:14:37 +00:00
parent 2a8abfbe8a
commit 69c6dde37c

View file

@ -45,7 +45,7 @@ Sweep 5 (21:08 Mac clock) notes: the coordinator applied the public proving line
| # | Number | Source | Tier affected | Consequence | Action | Owner | State |
|---|---|---|---|---|---|---|---|
| C23 | Mixer x4 or x8: the 5090's daily dataset build 13.4 ms at x1 (54 ms at x4, about 107 ms at x8); the integrated gfx1036 builds the dataset per PREPARE, 7 to 12 s at x1 and 55 to 124 s under CPU load (epoch-length.md section 7); the x8 rule: "the daily build under 1 s on every card we own" | status 21:25 and 21:27; epoch-length.md section 7 | integrated GPUs (the iGPU tier), 8 GB cards (about a tenth of a 5090's rate), rigs with one weak card | The x8 rule names "every card we own": the gfx1036 is one, and at x8 its per-prepare build is 56 to 96 s per epoch (7 to 16 min under load), so it misses every epoch boundary and falls to the exit-42 path; at x4 it is 28 to 48 s per hour (0.8 to 1.3%) or 4 to 8 min under load. An 8 GB discrete card scales at about a tenth of the 5090: about 1 s at x8, on the rule's edge. The per-day dataset reuse in the worker (owed in epoch-length.md section 9) is what makes the mixer cheap for the iGPU tier; without it the mixer multiplies a per-epoch cost | The mixer decision names the gfx1036 and an 8 GB-class scaled row beside the 5090, M5 Max and 9070 XT in the "under 1 s" check, and the per-day dataset reuse in the worker lands before (or with) class v3, or the iGPU tier is stated as "mines v3 with a restart per epoch" on the level 3 page | ca2-mixer af345b1e2c541ffbb; coordinator | taken (coordinator and ca2-mixer, mixer-x4.md build-time table: the x8 table carries the gfx1036 row and a scaled 8 GB-class row; the under-1-s rule applies to the discrete cards' daily build; for the integrated tier either the per-day dataset reuse in the three workers lands with class v3, asked of the mixer and node agents as a bounded change tonight, or the level 3 page says the iGPU tier mines v3 with a restart per epoch; recorded with the x4/x8 choice) |
| C24 | Two PowerShell job scripts lost a quote inside an inline bash body tonight: amd-prove's awk program (cpu-prove-pc1-small, exit 0 with nothing proved) and the 0.3.10 installer's `bash -c` string (fb-installer-pc1-3, exit 2 in 4 s); amd-prove added `tools/amd-prove/check-job-bash.sh` (bash -n on its own scripts' bash bodies) | amd-proving.md section 2; status 21:28 | every PC job; the morning's rollouts | The class (CLAUDE.md, 5 October: fix the class the same day, add a check that fails when the shape comes back) is "a bash body inside a PowerShell job string"; the check exists for one agent's scripts and did not cover the shipper's, which failed the same way an hour later | A repo-wide CI check: every `.ps1` under relay/playbooks and tools that carries a bash body (`wsl ... bash -c`, `bash -lc`, here-strings fed to bash) has that body extracted and passed through `bash -n`; the shipper's rule (the WSL part as a file run with `bash <file>`) written in packaging/README-ship.md as the convention | sub-agent (bounded, no owner); coordinator told | in work (sub-agent bash-body-check; the coordinator: no collision with PC 1, the shipper moved its WSL part to a file for attempt 4, the repro agent added a PowerShell 5.1 parse gate for bench/; the check fails on an unextractable body rather than skipping it) |
| C24 | Two PowerShell job scripts lost a quote inside an inline bash body tonight: amd-prove's awk program (cpu-prove-pc1-small, exit 0 with nothing proved) and the 0.3.10 installer's `bash -c` string (fb-installer-pc1-3, exit 2 in 4 s); amd-prove added `tools/amd-prove/check-job-bash.sh` (bash -n on its own scripts' bash bodies) | amd-proving.md section 2; status 21:28 | every PC job; the morning's rollouts | The class (CLAUDE.md, 5 October: fix the class the same day, add a check that fails when the shape comes back) is "a bash body inside a PowerShell job string"; the check exists for one agent's scripts and did not cover the shipper's, which failed the same way an hour later | A repo-wide CI check: every `.ps1` under relay/playbooks and tools that carries a bash body (`wsl ... bash -c`, `bash -lc`, here-strings fed to bash) has that body extracted and passed through `bash -n`; the shipper's rule (the WSL part as a file run with `bash <file>`) written in packaging/README-ship.md as the convention | sub-agent (bounded, no owner); coordinator told | closed on branch bash-body-check 7adb1ca (tools/ci/bash-body-check.sh with fixtures and self-test, ci.yml, the convention in packaging/README-ship.md; the flagged existing playbooks are in the sub-agent's report for their owners) |
| C25 | Ember Tune: the signed manifest carries per-card-model tuning priors (power limit and core clock) that a new card applies and confirms in two steps; K1 signs it | ember-tune.md sections 4 and 5; bench-log Ember Tune entry | every NVIDIA and AMD card on the app; the keys | The manifest now sets clocks and power limits on every user's card, so the signing key's blast radius grew: a signed prior can underclock the fleet or push a card model to its power ceiling. The plan's clamps (inside power.min_limit / max_limit and clocks.max.gr, the confirm step, a faulted step reverted) are the bound; keys.md's "what K1 signs" table (ota-k2 00fcbb5) predates the priors and does not list them | ember-tune.md section 5 states the bound in one line (a prior can never set a value outside the card's own reported limits, and never a memory clock), with the test that proves it; keys.md's table gains the tuning priors under K1 with that bound | Ember Tune a855dcc4bd05e0615; OTA key agent (the table row) | closed (Ember Tune 5d7ced9: the rule as a row in section 5 with the two tests named, a prior of 9,000 MHz at 30% clamps to 3,090 MHz at 50%, a bad prior costs one confirm step per card; the OTA agent has the keys.md note: tuning priors and the kill switch under K1 with that bound) |
Sweep 7 (21:49 UTC) notes, no new row: the hot table is measured and NOT adopted (g 0.93 to 0.96 on the Mac against the 0.97 rule, bdab8df); the era draw passes the 5% rule on the Mac (spread 0.8%, c570da3) and its chip line says the union of a program's 16 windows covered the whole dataset in 300 of 300 programs, so a chip mirrors the whole dataset or nothing; the mixer x8 passes the verifier half of its rule (C19) and the PC build rows decide the other half about 22:10; PC 2's miners have been off since a job's /api/resume at 21:25 answered ok without restarting them (the devnet short PC 2's rate, the aggregation-cost mining phases void, a next-cut defect: a resume re-checks the miner processes), all stated by the coordinator at 21:45; PC 1 restarted on 0.3.10 at 21:40:41Z with no measurement straddling it.
@ -55,6 +55,6 @@ Sweep 7 (21:49 UTC) notes, no new row: the hot table is measured and NOT adopted
| # | Number | Source | Tier affected | Consequence | Action | Owner | State |
|---|---|---|---|---|---|---|---|
| C26 | The 13.9 GB floor's cause: the shipped sp1-gpu-server 6.8.1 panics on any card under 24 GB (builder.rs 35 to 39) and allocates its core, recursion, shrink and wrap provers at Setup; the prover-floor agent rebuilds it from source on PC 2 with those sizes cut, CUDA_ARCHS=120 | proving-v1.md 4c82e56; status 22:01 | 12 and 16 GB NVIDIA cards (RTX 3060, 4070, 5070, 5080, 4060 Ti 16 GB): the tier the project lead asked for; packaging and signing | A server built for CUDA_ARCHS=120 runs on the 5090 only; the 12 GB tier is sm_86 (3060) and sm_89 (4070), the 16 GB tier sm_89 and sm_120, so a cut-size server measured on the 5090 proves nothing about a 3060 until the on-order 3060 runs it, and the build must list sm_86, sm_89, sm_120 (sm_100 is datacentre) to serve the tier at all. Shipping our own 250 MB CUDA server means the project signs and distributes a build of someone else's prover: it enters the DMG and the WSL2 package, the K1-signed inputs, the SBOM-style notes in evidence.md, and every SP1 upgrade is re-done by hand. Cut buffer sizes do not change the verifying key (prover-side chunking), so no guest re-pin, but the recipe must say so with a verify-segment run on a proof from the rebuilt server | The prover-floor measurement states its arch list and the card it ran on; the 12 GB claim waits for the 3060; the rebuilt server's packaging path (who builds, who signs, where it lands) is a row in the proving plan before 0.3.12, and the public line keeps "24 GB" until the 3060 proves on it | prover-floor agent (through the coordinator); proving v1 | taken (the proving plan carries "A self-built CUDA server (the 12 GB path), before 0.3.12": arch list sm_86 / sm_89 / sm_120 with one measured row per family, the build on PC 1 from a pinned SP1 tag, the Mac signs, placement as wsl2/bin/sp1-gpu-server with its sha256 in payload-inputs.json and the DMG, the evidence.md note, the rebuild at each SP1 upgrade, the gate that verify-segment and verify show the pinned keys unchanged; the 12 GB claim waits for the 3060; the prover-floor agent abefda4c3872f866f has the measurement side) |
| C27 | The root-socket fault recurred at 21:25Z from another agent's job (agg-cost-pc2-1) after the class fix and CI check landed; PC 2's prover was dark 37 minutes; a resume at 21:25 answered ok without restarting the miners | bench-log 1d78979; status 21:45, 22:03 | every PC job; the devnet's proving and hash rate tonight | The class check lives in CI, but PC jobs are published from worktrees by `publish-jobs.sh` and never pass through CI before they run, so a job written on a branch without the check runs the old shape. The check must run where the job is published, not only where the repo is tested | `packaging/ota/publish-jobs.sh add` runs `tools/ci/prover-socket-check.sh` and the bash-body check on the script it publishes and refuses on a failure; the sub-agent on the bash-body check wires both; the resume defect is on the next-cut list (stated) | sub-agent bash-body-check (the wiring); coordinator (the rule) | in work (sub-agent wiring both checks into publish-jobs.sh add; the coordinator: no collision, jobs without a script pass trivially, an unextractable body fails) |
| C27 | The root-socket fault recurred at 21:25Z from another agent's job (agg-cost-pc2-1) after the class fix and CI check landed; PC 2's prover was dark 37 minutes; a resume at 21:25 answered ok without restarting the miners | bench-log 1d78979; status 21:45, 22:03 | every PC job; the devnet's proving and hash rate tonight | The class check lives in CI, but PC jobs are published from worktrees by `publish-jobs.sh` and never pass through CI before they run, so a job written on a branch without the check runs the old shape. The check must run where the job is published, not only where the repo is tested | `packaging/ota/publish-jobs.sh add` runs `tools/ci/prover-socket-check.sh` and the bash-body check on the script it publishes and refuses on a failure; the sub-agent on the bash-body check wires both; the resume defect is on the next-cut list (stated) | sub-agent bash-body-check (the wiring); coordinator (the rule) | closed on branch bash-body-check 6805125 (publish-jobs.sh add --kind run runs the bash-body check and prover-socket-check.sh before signing, refuses with the output, never skips; test-publish-jobs.sh 32 passed with four new refusals). Merge notes for the integrator: prover-socket-check.sh exists on both proving-v1 c2544be and this branch (add/add, take the superset here); the socket grep flags tools/amd-prove/pc1-cpu-prove.ps1 (CPU-only, -u root, no GPU server) so that job needs the cleanup lines or an allow-list entry before its next publish, told to the coordinator |
Sweep 8 (22:09 UTC) notes: mixer x8 DECIDED into v3 on the PC rows (the daily 1 GiB build latency-bound on every card: 5090 23 to 25 ms, 9070 XT 72 to 77 ms at every multiplier; the chip row 0.92x with the 3x factor), so the public claim holds with margin (D5 re-cut); the verifier regression (2.2x) bisected to inlining in the mixer's fetch loop and fixed, so the quiet-core figures of C19 return to about 0.6 / 0.9 / 1.3 ms; G6 job 3 failed on a stale fork test (era inside the class), job 4 on the final tree; the integration merge into master has three known conflicts (bench-log append-only, packfile.h and host.c take the ca2-v3 side).