From ac2a4f986209fefa3f325f54f60a226e36a9f3c4 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:11:05 +0000 Subject: [PATCH] Consequences ledger: C26 taken into the proving plan Co-Authored-By: Claude Fable 5.1 --- docs/plans/consequences-2026-10-05.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/plans/consequences-2026-10-05.md b/docs/plans/consequences-2026-10-05.md index b8ab4e8a6..6f03ee73b 100644 --- a/docs/plans/consequences-2026-10-05.md +++ b/docs/plans/consequences-2026-10-05.md @@ -54,7 +54,7 @@ Sweep 7 (21:49 UTC) notes, no new row: the hot table is measured and NOT adopted | # | Number | Source | Tier affected | Consequence | Action | Owner | State | |---|---|---|---|---|---|---|---| -| C26 | The 13.9 GB floor's cause: the shipped sp1-gpu-server 6.8.1 panics on any card under 24 GB (builder.rs 35 to 39) and allocates its core, recursion, shrink and wrap provers at Setup; the prover-floor agent rebuilds it from source on PC 2 with those sizes cut, CUDA_ARCHS=120 | proving-v1.md 4c82e56; status 22:01 | 12 and 16 GB NVIDIA cards (RTX 3060, 4070, 5070, 5080, 4060 Ti 16 GB): the tier the project lead asked for; packaging and signing | A server built for CUDA_ARCHS=120 runs on the 5090 only; the 12 GB tier is sm_86 (3060) and sm_89 (4070), the 16 GB tier sm_89 and sm_120, so a cut-size server measured on the 5090 proves nothing about a 3060 until the on-order 3060 runs it, and the build must list sm_86, sm_89, sm_120 (sm_100 is datacentre) to serve the tier at all. Shipping our own 250 MB CUDA server means the project signs and distributes a build of someone else's prover: it enters the DMG and the WSL2 package, the K1-signed inputs, the SBOM-style notes in evidence.md, and every SP1 upgrade is re-done by hand. Cut buffer sizes do not change the verifying key (prover-side chunking), so no guest re-pin, but the recipe must say so with a verify-segment run on a proof from the rebuilt server | The prover-floor measurement states its arch list and the card it ran on; the 12 GB claim waits for the 3060; the rebuilt server's packaging path (who builds, who signs, where it lands) is a row in the proving plan before 0.3.12, and the public line keeps "24 GB" until the 3060 proves on it | prover-floor agent (through the coordinator); proving v1 | sent in full to the prover-floor agent abefda4c3872f866f by the coordinator (arch list, the packaging row before 0.3.12, the verify-segment run, "24 GB" kept until the 3060) | +| C26 | The 13.9 GB floor's cause: the shipped sp1-gpu-server 6.8.1 panics on any card under 24 GB (builder.rs 35 to 39) and allocates its core, recursion, shrink and wrap provers at Setup; the prover-floor agent rebuilds it from source on PC 2 with those sizes cut, CUDA_ARCHS=120 | proving-v1.md 4c82e56; status 22:01 | 12 and 16 GB NVIDIA cards (RTX 3060, 4070, 5070, 5080, 4060 Ti 16 GB): the tier the project lead asked for; packaging and signing | A server built for CUDA_ARCHS=120 runs on the 5090 only; the 12 GB tier is sm_86 (3060) and sm_89 (4070), the 16 GB tier sm_89 and sm_120, so a cut-size server measured on the 5090 proves nothing about a 3060 until the on-order 3060 runs it, and the build must list sm_86, sm_89, sm_120 (sm_100 is datacentre) to serve the tier at all. Shipping our own 250 MB CUDA server means the project signs and distributes a build of someone else's prover: it enters the DMG and the WSL2 package, the K1-signed inputs, the SBOM-style notes in evidence.md, and every SP1 upgrade is re-done by hand. Cut buffer sizes do not change the verifying key (prover-side chunking), so no guest re-pin, but the recipe must say so with a verify-segment run on a proof from the rebuilt server | The prover-floor measurement states its arch list and the card it ran on; the 12 GB claim waits for the 3060; the rebuilt server's packaging path (who builds, who signs, where it lands) is a row in the proving plan before 0.3.12, and the public line keeps "24 GB" until the 3060 proves on it | prover-floor agent (through the coordinator); proving v1 | taken (the proving plan carries "A self-built CUDA server (the 12 GB path), before 0.3.12": arch list sm_86 / sm_89 / sm_120 with one measured row per family, the build on PC 1 from a pinned SP1 tag, the Mac signs, placement as wsl2/bin/sp1-gpu-server with its sha256 in payload-inputs.json and the DMG, the evidence.md note, the rebuild at each SP1 upgrade, the gate that verify-segment and verify show the pinned keys unchanged; the 12 GB claim waits for the 3060; the prover-floor agent abefda4c3872f866f has the measurement side) | | C27 | The root-socket fault recurred at 21:25Z from another agent's job (agg-cost-pc2-1) after the class fix and CI check landed; PC 2's prover was dark 37 minutes; a resume at 21:25 answered ok without restarting the miners | bench-log 1d78979; status 21:45, 22:03 | every PC job; the devnet's proving and hash rate tonight | The class check lives in CI, but PC jobs are published from worktrees by `publish-jobs.sh` and never pass through CI before they run, so a job written on a branch without the check runs the old shape. The check must run where the job is published, not only where the repo is tested | `packaging/ota/publish-jobs.sh add` runs `tools/ci/prover-socket-check.sh` and the bash-body check on the script it publishes and refuses on a failure; the sub-agent on the bash-body check wires both; the resume defect is on the next-cut list (stated) | sub-agent bash-body-check (the wiring); coordinator (the rule) | in work (sub-agent wiring both checks into publish-jobs.sh add; the coordinator: no collision, jobs without a script pass trivially, an unextractable body fails) | Sweep 8 (22:09 UTC) notes: mixer x8 DECIDED into v3 on the PC rows (the daily 1 GiB build latency-bound on every card: 5090 23 to 25 ms, 9070 XT 72 to 77 ms at every multiplier; the chip row 0.92x with the 3x factor), so the public claim holds with margin (D5 re-cut); the verifier regression (2.2x) bisected to inlining in the mixer's fetch loop and fixed, so the quiet-core figures of C19 return to about 0.6 / 0.9 / 1.3 ms; G6 job 3 failed on a stale fork test (era inside the class), job 4 on the final tree; the integration merge into master has three known conflicts (bench-log append-only, packfile.h and host.c take the ca2-v3 side).