Consequences: C28 host RAM per server on rigs; D9 wording reconciled; D10 the proving route (A now, RISC Zero as version 2 is the project lead's)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
f91dc13689
commit
d3b30f4940
2 changed files with 5 additions and 1 deletions
|
|
@ -58,3 +58,6 @@ Sweep 7 (21:49 UTC) notes, no new row: the hot table is measured and NOT adopted
|
|||
| C27 | The root-socket fault recurred at 21:25Z from another agent's job (agg-cost-pc2-1) after the class fix and CI check landed; PC 2's prover was dark 37 minutes; a resume at 21:25 answered ok without restarting the miners | bench-log 1d78979; status 21:45, 22:03 | every PC job; the devnet's proving and hash rate tonight | The class check lives in CI, but PC jobs are published from worktrees by `publish-jobs.sh` and never pass through CI before they run, so a job written on a branch without the check runs the old shape. The check must run where the job is published, not only where the repo is tested | `packaging/ota/publish-jobs.sh add` runs `tools/ci/prover-socket-check.sh` and the bash-body check on the script it publishes and refuses on a failure; the sub-agent on the bash-body check wires both; the resume defect is on the next-cut list (stated) | sub-agent bash-body-check (the wiring); coordinator (the rule) | closed on branch bash-body-check 6805125 (publish-jobs.sh add --kind run runs the bash-body check and prover-socket-check.sh before signing, refuses with the output, never skips; test-publish-jobs.sh 32 passed with four new refusals). Merge notes for the integrator: prover-socket-check.sh exists on both proving-v1 c2544be and this branch (add/add, take the superset here); the socket grep flags tools/amd-prove/pc1-cpu-prove.ps1 (CPU-only, -u root, no GPU server) so that job needs the cleanup lines or an allow-list entry before its next publish, told to the coordinator |
|
||||
|
||||
Sweep 8 (22:09 UTC) notes: mixer x8 DECIDED into v3 on the PC rows (the daily 1 GiB build latency-bound on every card: 5090 23 to 25 ms, 9070 XT 72 to 77 ms at every multiplier; the chip row 0.92x with the 3x factor), so the public claim holds with margin (D5 re-cut); the verifier regression (2.2x) bisected to inlining in the mixer's fetch loop and fixed, so the quiet-core figures of C19 return to about 0.6 / 0.9 / 1.3 ms; G6 job 3 failed on a stale fork test (era inside the class), job 4 on the final tree; the integration merge into master has three known conflicts (bench-log append-only, packfile.h and host.c take the ca2-v3 side).
|
||||
| C28 | One sp1-gpu-server per card on rigs pins about 6 GB of host RAM per server today (under 2 GB with route A's buffers, approximate) | `docs/analysis/proving-methods.md` section 5, rig row (proving-methods e7e0db7) | rigs with several 24 GB or 32 GB cards on the rig installer | A six-card rig that proves on every card pins about 36 GB of host RAM under the shipped server; the rig installer's preflight says 16 GB to prove (a warning, per card not per rig) and its prover unit runs one server, so the moment it moves to one server per card (the analysis's route C) the RAM check is wrong by the card count | The rig preflight scales its RAM warning by the number of proving cards (6 GB each today, the route A figure when measured) and the README's per-card table gains a host RAM column; the one-server-per-card unit lands only with that check | rig installer a3e7b2b03222f5cff | sent |
|
||||
|
||||
Sweep 8a (22:2x UTC) note: `docs/analysis/proving-methods.md` (branch proving-methods e7e0db7, 411 lines) carries its own per-tier table for today, route A, route D and route 3, and a public paragraph; it is the third draft of the proving public line (with proving-v1 440fd59 and ca2-coord 1c8439f), noted under D2 and D8; the route choice is D10.
|
||||
|
|
|
|||
|
|
@ -12,4 +12,5 @@ Sibling of `ledger-decisions.md`. Each is a consequence of a measured number tha
|
|||
| D6 | C11 | Whether class v3 should favour AMD (fewer, wider loads per hash) at a cost to the 5090's latency-bound share, or accept that AMD cards mine at about a seventh of a 5090 and 4.5x the electricity per hash | Accept it for v3 and say so on the miners page ("NVIDIA first; AMD mines at a lower rate per watt on this class"); open the AMD question as a Counter ASIC 3.0 item with its own measurement | Tonight's rule was the project lead's and the 5090 margin is the anti-chip argument; AMD's position is a public-copy question, not a gate |
|
||||
| D7 | C13 | Same as D3's supply wording | with D3 | |
|
||||
| D8 | C7 | The miner page says "One click: install, start, the card mines and proves" (site/miner.html 7, 13, 21; site/index.html 484). On HiveOS and the rig installer a rig mines only until a Linux prover build is published, and on the app a 12 GB card mines only, a 16 GB card proves with the miner paused, 20 GB and up does both (the sweep of 5 October, before the v1-shard row) | Qualify the sentence on the miner page and the homepage card: "the card mines; 24 GB cards prove too, 32 GB does both at once; rigs mine until the Linux prover ships". The same page's "a visible 1% software fee you can switch off" becomes "solo mining carries a 1% software fee you can switch off; in a pool the pool's own fee is the only one" once pool-v0 ships (the pool agent fixed the rule: no software dev fee in pool mode) | The sentence is the product's first promise and tonight's measurement bounds it by card |
|
||||
| D9 | C3, C15, C26 | No 12 GB NVIDIA card exists in the fleet, so every 12 GB number tonight is scaled from the 5090 and labelled approximate. The 13.9 GB floor is SP1's GPU server code (`docs/analysis/proving-methods.md`, branch proving-methods e7e0db7: builder.rs refuses cards under 20 GB, the trace is allocated at the maximum shard, the CUDA mempool never releases; the proving plan's read of the same file says 24 GB, 4c82e56: the prover-floor agent's rows settle which). The prover-floor patch and the per-card profiles cannot be measured without the hardware | Buy one 12 GB NVIDIA card this week for PC 1's spare slot (an RTX 3060 12 GB or 4070 12 GB, about £250 to £400, approximate; the coordinator's request). Check first whether the RTX 3060 and the RTX 5060 Ti 16 GB that evidence.md row 16 says are "on order" are real and arriving; if so, no purchase, only the delivery date. No public line says "12 GB proves" before a real 12 GB card runs the rebuilt server on the S_p-curve fixtures and recipe | Money, and the one measurement every 12 GB claim rests on; the prover-floor agent's rows tonight replace the approximate figures when they land |
|
||||
| D9 | C3, C15, C26 | No 12 GB NVIDIA card exists in the fleet, so every 12 GB number tonight is scaled from the 5090 and labelled approximate. The 13.9 GB floor is SP1's GPU server code (`docs/analysis/proving-methods.md`, branch proving-methods e7e0db7, section 1.3: builder.rs adds 4 GB to the card's physical memory and panics under 24, so a 16 GB card (16 + 4 = 20) and a 12 GB card never start and 20 GB is the smallest that does; the trace is allocated at the maximum shard; the CUDA mempool never releases; the proving plan's "under 24" (4c82e56) is the same test read before the addition). The prover-floor patch and the per-card profiles cannot be measured without the hardware | Buy one 12 GB NVIDIA card this week for PC 1's spare slot (an RTX 3060 12 GB or 4070 12 GB, about £250 to £400, approximate; the coordinator's request). Check first whether the RTX 3060 and the RTX 5060 Ti 16 GB that evidence.md row 16 says are "on order" are real and arriving; if so, no purchase, only the delivery date. No public line says "12 GB proves" before a real 12 GB card runs the rebuilt server on the S_p-curve fixtures and recipe | Money, and the one measurement every 12 GB claim rests on; the prover-floor agent's rows tonight replace the approximate figures when they land |
|
||||
| D10 | C3, C15, C26, C28 | The proving route for the 12 GB tier and for Apple: `docs/analysis/proving-methods.md` section 4 ranks (1) route A, a re-sized SP1 GPU server with `S_p` as the dial and one server per card on rigs, nothing in consensus moving; (2) route D, RISC Zero 3.0.x as proof-system version 2 (the only shipped prover with a documented sub-12 GB configuration and a Metal path), the fallback if A misses 11 GB and the Apple route either way, 3 to 4 days plus a 3-month two-verifier overlap and three spec 7.8 rules; (3) a sumcheck family without a codeword (Jolt-class) in years, not now. Route A is already running tonight; route D adds a second proof family to the node, which is a consensus and verifier change outside tonight's delegation | Route A on the prover-floor rows, gated as the document says (the adopted shard under 11 GB alone and under 60 s prove-only on a real 12 GB card, D9); route D's measurement (RISC Zero at po2 19 and 20 on PC 2 and the Mac's Metal row) may run as a measurement, but adopting a second proof family waits for the project lead and for route A's result; the public paragraph of section 5 ("Proving runs on NVIDIA cards with 24 GB or more today. A build for 12 GB and 16 GB cards is being measured ...") is the honest line meanwhile and is the one of the three drafts to prefer, because it names what is being measured instead of a tier | A second verifier in the node is the kind of change the testnet's genesis must carry from day one; measuring it costs nothing, adopting it is the project lead's |
|
||||
|
|
|
|||
Loading…
Reference in a new issue