Horizon: new-pow section 6 carries the B and C verdicts in full and the ranking after measurement
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
ac4cc93772
commit
92db7c5ab1
1 changed files with 5 additions and 1 deletions
|
|
@ -278,7 +278,11 @@ Reading: B moves the f = 1 chip's edge by 1.4x to 2x at `k = 1` and by 1.1x at `
|
||||||
|
|
||||||
**A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.
|
**A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.
|
||||||
|
|
||||||
NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joulesC_PROSE
|
**B, in full.** The scheme asked whether a GPU-shaped puzzle could make a chip into a GPU. The answer the 4090 gives is that the tensor-shaped work is free for the honest card (rate unchanged to R = 512, 15 W at most for 4,096 tiles per hash) and therefore nearly free for the chip too: a lever against the f = 1 memory-controller chip works only through joules the honest card is forced to spend on work the chip cannot do more cheaply, and the ALU shadow of class v4 spends 11 pJ per op where the tensor path spends 0.056 to 0.70 pJ per multiply-add. The chip's edge moves from 6.9x to 4.9x at k = 1 (4090 denominator) where the ALU shadow moved the 5090's by 2.7x (5.6x to 2.1x), and the verifier pays 26x more per unit of forcing energy (4.39 ms per 0.23 microjoules against 0.17 ms per 0.6). There is also the vendor split (native on Turing and later and on RDNA 3 and later, emulated on Pascal, RDNA 2 and every Apple card) and the unverified AMD layout. So: never as class v5 content for this purpose. What the measurement does establish, and what should be kept: an mm8 family is exact on NVIDIA at every R tried (2^24 lanes PTX = reference, 1,024 lanes CPU = GPU at five rungs), cheap for the honest card, and bounded in verifier cost by a known law (1.1 microseconds per step per unit, scalar); that is the evidence the reserve entry R8 lacked, and the two-output form (both tile outputs consumed) is a correction R8 needs so a chip cannot halve the tile. The standing rule this leaves for the programme: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is LEAST efficient per op among the operations a chip cannot do much better (RandomX's argument), never where it is most efficient.
|
||||||
|
|
||||||
|
**C, in full.** The scheme does what it says at a cost that is measured and small: the hash kernel is byte for byte the shipped one, the 4090's rate and watts are equal within noise (63.083 against 63.088 MH/s, 205 to 208 W), the daily build grows by 1.4 ms resident or about 76 ms streamed over PCIe (about 150 ms at the designed 2 GiB, approximate), device memory during the build stays at dataset plus cache plus one 64 MiB chunk, the verifier grows by 0.11 to 0.21 ms per unit with the state in RAM (about 2.2 ms on the M5 Max core, about 5.7 ms on a 2019-class core by the 2.5x rule, inside the 10 ms gate), and every GPU word and lane agrees with the plain C derivation (1,024 of 1,024 items, 128 of 128 lanes). It is never weaker than today's dataset against any chip and it removes the f = 0 recompute chip as a category. What it adds has not been done on a shipped chain in this form: the dataset IS the keyed execution state, so every mining operation holds the chain, and every block names 128 random state leaves per lane that any full node can open against the day's state root. Its limits are stated: a pool can ship the dataset (a 2 GiB daily delivery per pool miner, where today a pool miner needs only the day key), a header verifier must hold the state or be shown 3.4 MiB of openings per unit (so light clients still rely on certificates, as spec 10 says), and the canonical serialisation is new consensus-critical code that must cross a day boundary and a finality pause on the fast-time harness before Devnet 2. Verdict: ship as the class v5 candidate, with the three spec items of 4.3 (the serialisation, the pruning-proof witness, the pause rule), the Devnet 2 gate, and the open decision on what a header verifier is asked to hold in front of it.
|
||||||
|
|
||||||
|
**The ranking after measurement.** C, then nothing else from this lane as class content; B retires into the reserve entry it improves; A is recorded as a bound. The brief's hope of "mining is proving" is answered with the reason it cannot be, which is worth more to the project than a design that pretended otherwise.
|
||||||
|
|
||||||
## 7. Ranked next steps
|
## 7. Ranked next steps
|
||||||
|
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue