Consequences: C28 closed; C29 the litepaper's verifier ms is the ratio; C30 the 5090's 15% app gap; D11 the unfunded chip bounty on the hero
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
20a2d19883
commit
a215abc809
2 changed files with 5 additions and 1 deletions
|
|
@ -58,6 +58,8 @@ Sweep 7 (21:49 UTC) notes, no new row: the hot table is measured and NOT adopted
|
|||
| C27 | The root-socket fault recurred at 21:25Z from another agent's job (agg-cost-pc2-1) after the class fix and CI check landed; PC 2's prover was dark 37 minutes; a resume at 21:25 answered ok without restarting the miners | bench-log 1d78979; status 21:45, 22:03 | every PC job; the devnet's proving and hash rate tonight | The class check lives in CI, but PC jobs are published from worktrees by `publish-jobs.sh` and never pass through CI before they run, so a job written on a branch without the check runs the old shape. The check must run where the job is published, not only where the repo is tested | `packaging/ota/publish-jobs.sh add` runs `tools/ci/prover-socket-check.sh` and the bash-body check on the script it publishes and refuses on a failure; the sub-agent on the bash-body check wires both; the resume defect is on the next-cut list (stated) | sub-agent bash-body-check (the wiring); coordinator (the rule) | closed on branch bash-body-check 6805125 (publish-jobs.sh add --kind run runs the bash-body check and prover-socket-check.sh before signing, refuses with the output, never skips; test-publish-jobs.sh 32 passed with four new refusals). Merge notes for the integrator: prover-socket-check.sh exists on both proving-v1 c2544be and this branch (add/add, take the superset here); the socket grep flags tools/amd-prove/pc1-cpu-prove.ps1 (CPU-only, -u root, no GPU server) so that job needs the cleanup lines or an allow-list entry before its next publish, told to the coordinator |
|
||||
|
||||
Sweep 8 (22:09 UTC) notes: mixer x8 DECIDED into v3 on the PC rows (the daily 1 GiB build latency-bound on every card: 5090 23 to 25 ms, 9070 XT 72 to 77 ms at every multiplier; the chip row 0.92x with the 3x factor), so the public claim holds with margin (D5 re-cut); the verifier regression (2.2x) bisected to inlining in the mixer's fetch loop and fixed, so the quiet-core figures of C19 return to about 0.6 / 0.9 / 1.3 ms; G6 job 3 failed on a stale fork test (era inside the class), job 4 on the final tree; the integration merge into master has three known conflicts (bench-log append-only, packfile.h and host.c take the ca2-v3 side).
|
||||
| C28 | One sp1-gpu-server per card on rigs pins about 6 GB of host RAM per server today (under 2 GB with route A's buffers, approximate) | `docs/analysis/proving-methods.md` section 5, rig row (proving-methods e7e0db7) | rigs with several 24 GB or 32 GB cards on the rig installer | A six-card rig that proves on every card pins about 36 GB of host RAM under the shipped server; the rig installer's preflight says 16 GB to prove (a warning, per card not per rig) and its prover unit runs one server, so the moment it moves to one server per card (the analysis's route C) the RAM check is wrong by the card count | The rig preflight scales its RAM warning by the number of proving cards (6 GB each today, the route A figure when measured) and the README's per-card table gains a host RAM column; the one-server-per-card unit lands only with that check | rig installer a3e7b2b03222f5cff | sent |
|
||||
| C28 | One sp1-gpu-server per card on rigs pins about 6 GB of host RAM per server today (under 2 GB with route A's buffers, approximate) | `docs/analysis/proving-methods.md` section 5, rig row (proving-methods e7e0db7) | rigs with several 24 GB or 32 GB cards on the rig installer | A six-card rig that proves on every card pins about 36 GB of host RAM under the shipped server; the rig installer's preflight says 16 GB to prove (a warning, per card not per rig) and its prover unit runs one server, so the moment it moves to one server per card (the analysis's route C) the RAM check is wrong by the card count | The rig preflight scales its RAM warning by the number of proving cards (6 GB each today, the route A figure when measured) and the README's per-card table gains a host RAM column; the one-server-per-card unit lands only with that check | rig installer a3e7b2b03222f5cff | closed (rig installer 086008a: the preflight checks host RAM against 8 GB plus about 6 GB per proving card at the 23,552 MB gate, PROVER_RAM_GB_PER_CARD default 6 marked approximate, the README host RAM row per tier; the one-server-per-card unit lands only with that check) |
|
||||
|
||||
Sweep 8a (22:2x UTC) note: `docs/analysis/proving-methods.md` (branch proving-methods e7e0db7, 411 lines) carries its own per-tier table for today, route A, route D and route 3, and a public paragraph; it is the third draft of the proving public line (with proving-v1 440fd59 and ca2-coord 1c8439f), noted under D2 and D8; the route choice is D10.
|
||||
| C29 | The litepaper's verifier line now reads "Measured 2.1 ms on one loaded Apple M5 Max core for class v3" and the status says "4.8x inside the gate"; the measurement (ca2-mixer 54bbfcc) is 2.79 ms per warp for x8 on the loaded core (2.1 was the RATIO to v2), worst cold unit 2.94 ms, quiet-core about 1.3 ms by scaling | site/litepaper.html on ca2-coord 8e65696; status 22:20 | the public page; every node operator who reads the gate margin | A ratio printed as milliseconds understates the verifier cost by a third and overstates the gate margin (10 / 2.79 = 3.6x, 3.4x on the worst cold unit, not 4.8x). The number is the one a reviewer will re-run first | The line reads "about 2.8 ms per warp on a loaded M5 Max core (about 1.3 ms quiet, approximate), 2.1x the v2 verifier; worst cold unit 2.9 ms; the 10 ms gate leaves 3.4x" until the quiet-core run lands | coordinator ada8afb62d752b1e2 | sent |
|
||||
| C30 | The 5090 mines in the app at 115.4 MH/s (the power sweep, 22:09 to 22:15Z, hash from the app's API) against 136 to 137 MH/s at device time in every bench row tonight, with the cap not binding (draw 316 W under a 431 W cap, SM at 3,051 MHz) | bench-log e304458; read-width and mixer PC rows | every 5090 owner on the app (and every big card: the gap is the app's job loop, not the kernel) | About 15% of a 5090's hash is lost between the kernel and the app, and the sweep entry explains it away as API sampling. M11 measured the 9070 XT at the app's 2^21 job size equal to its 2^24 rate, but no 5090 row exists at 2^21; the 5090 finishes a 2^21 job in about 15 ms, so per-job launch, read-back and template work can cost that much. The STATUS line prints "wall" and "inside jobs" rates and would show it | One measurement on PC 1: `igneum-worker-cuda --bench` on the live pack at --batch-log2 21 and 24 on the 5090, and the 5090's STATUS wall-against-inside gap over 10 minutes; if the job size is the cause, the app's job size for cards over 100 MH/s rises (2^22 or 2^23) in the next cut: a 15% gain for every 5090 owner | repro-bench agent a0b9f574775ef1693 (its PC 1 slot); coordinator | sent |
|
||||
|
|
|
|||
|
|
@ -14,6 +14,7 @@ Sibling of `ledger-decisions.md`. Each is a consequence of a measured number tha
|
|||
| D4 | gate 1 (before the testnet genesis) | Dataset growth mapping: (b) power-of-two steps (years 4, 12, 28, 60) keeps the public sentences true; (a) fades cards one year at a time | Choose; (b) recommended with the step calendar published |
|
||||
| D5 | before the next site push | The chip claim: mixer x8 measured at 0.92x with the 3x factor, "under 2x" holds with margin on the stated convention | Confirm the claim stays, with the convention named |
|
||||
| D6 | Counter ASIC 3.0 | AMD mines at a seventh of a 5090 and 4.9x the electricity per hash; no read width closes it | Accept for v3; the miners page says so |
|
||||
| D11 | before the next site push | The hero, the abstract and the level 1 copy now say "the model and the bounty are public" / "bounty standing"; funding.md prices the chip bounty at USD 50,000, marks it NOT FUNDED, and its rule 3 says a bounty is announced only when escrowed | Strike "and the bounty" from the hero and the abstract until the USD 50,000 is escrowed, or escrow it; the litepaper's older "a bounty is attached" (finality, line 686) is the same question |
|
||||
| D10 | after route A's rows and D9's card | Proving route: route A (re-sized SP1 server) now; RISC Zero as a second proof family (Apple, and the fallback) is a consensus and verifier change | Measure both; adopt a second family only on your say |
|
||||
|
||||
| # | Row | What needs deciding | Recommendation | Why |
|
||||
|
|
@ -28,3 +29,4 @@ Sibling of `ledger-decisions.md`. Each is a consequence of a measured number tha
|
|||
| D8 | C7 | The miner page says "One click: install, start, the card mines and proves" (site/miner.html 7, 13, 21; site/index.html 484). On HiveOS and the rig installer a rig mines only until a Linux prover build is published, and on the app a 12 GB card mines only, a 16 GB card proves with the miner paused, 20 GB and up does both (the sweep of 5 October, before the v1-shard row) | Qualify the sentence on the miner page and the homepage card: "the card mines; 24 GB cards prove too, 32 GB does both at once; rigs mine until the Linux prover ships". The same page's "a visible 1% software fee you can switch off" becomes "solo mining carries a 1% software fee you can switch off; in a pool the pool's own fee is the only one" once pool-v0 ships (the pool agent fixed the rule: no software dev fee in pool mode) | The sentence is the product's first promise and tonight's measurement bounds it by card |
|
||||
| D9 | C3, C15, C26 | No 12 GB NVIDIA card exists in the fleet, so every 12 GB number tonight is scaled from the 5090 and labelled approximate. The 13.9 GB floor is SP1's GPU server code (`docs/analysis/proving-methods.md`, branch proving-methods e7e0db7, section 1.3: builder.rs adds 4 GB to the card's physical memory and panics under 24, so a 16 GB card (16 + 4 = 20) and a 12 GB card never start and 20 GB is the smallest that does; the trace is allocated at the maximum shard; the CUDA mempool never releases; the proving plan's "under 24" (4c82e56) is the same test read before the addition). The prover-floor patch and the per-card profiles cannot be measured without the hardware | Buy one 12 GB NVIDIA card this week for PC 1's spare slot (an RTX 3060 12 GB or 4070 12 GB, about £250 to £400, approximate; the coordinator's request). Check first whether the RTX 3060 and the RTX 5060 Ti 16 GB that evidence.md row 16 says are "on order" are real and arriving; if so, no purchase, only the delivery date. No public line says "12 GB proves" before a real 12 GB card runs the rebuilt server on the S_p-curve fixtures and recipe | Money, and the one measurement every 12 GB claim rests on; the prover-floor agent's rows tonight replace the approximate figures when they land |
|
||||
| D10 | C3, C15, C26, C28 | The proving route for the 12 GB tier and for Apple: `docs/analysis/proving-methods.md` section 4 ranks (1) route A, a re-sized SP1 GPU server with `S_p` as the dial and one server per card on rigs, nothing in consensus moving; (2) route D, RISC Zero 3.0.x as proof-system version 2 (the only shipped prover with a documented sub-12 GB configuration and a Metal path), the fallback if A misses 11 GB and the Apple route either way, 3 to 4 days plus a 3-month two-verifier overlap and three spec 7.8 rules; (3) a sumcheck family without a codeword (Jolt-class) in years, not now. Route A is already running tonight; route D adds a second proof family to the node, which is a consensus and verifier change outside tonight's delegation | Route A on the prover-floor rows, gated as the document says (the adopted shard under 11 GB alone and under 60 s prove-only on a real 12 GB card, D9); route D's measurement (RISC Zero at po2 19 and 20 on PC 2 and the Mac's Metal row) may run as a measurement, but adopting a second proof family waits for the project lead and for route A's result; the public paragraph of section 5 ("Proving runs on NVIDIA cards with 24 GB or more today. A build for 12 GB and 16 GB cards is being measured ...") is the honest line meanwhile and is the one of the three drafts to prefer, because it names what is being measured instead of a tier | A second verifier in the node is the kind of change the testnet's genesis must carry from day one; measuring it costs nothing, adopting it is the project lead's |
|
||||
| D11 | C29 context; `docs/plans/funding.md` 36 and 63 | The chip bounty on the public pages: "the model and the bounty are public" (hero, abstract, ca2-coord cabec3b), "bounty standing" (level 1 copy); funding.md: USD 50,000 standing, "Not funded", "a bounty is announced only when it is escrowed"; the litepaper already says "a bounty is attached" to the finality review (line 686) | Strike "and the bounty" from the hero and the abstract and "standing" from level 1 until the money is escrowed, and say "a bounty follows the external review" if a sentence is wanted; or escrow USD 50,000 (and the USD 25,000 finality bounty) and keep the words. The coordinator is told to hold the words at the ship pending this | A public promise of money the project has not set aside is the FUD the ledger exists to prevent, by the project's own funding rule |
|
||||
|
|
|
|||
Loading…
Reference in a new issue