Review B F09 and F05 on the model: (c) every cell generated from the inequality (the board's hand-written PASS at USD 18 M against 40 M corrected to FAIL; no chip meets (c) at any price; the board six of seven); (g) the hybrid operator with companion GPUs keeps 85 to 90 percent of the sunk owner's proving surplus, so (g) is not exclusive to the GPU owner; F05 the cached prefix tree priced beside the literal 63-fold at the adversary lane's placed figures (the chip's energy per hash -2.6 percent, the connected-state verdict unchanged)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
95a33769af
commit
5cb6b20888
2 changed files with 96 additions and 7 deletions
|
|
@ -71,9 +71,12 @@ ticket lands near; section 12c is the headline and 12b the superseded 300 W row.
|
|||
not under about a quarter of the GPU entrant's; (c) a fleet above a third of the chain costing more than a year's
|
||||
miner revenue; (d) GPUs keeping a resale market and a use outside mining; (e) the per-joule gap at the honest knee
|
||||
under about 3x; (f) the supplier's margin a normal return that does not rise with the halvings; (g) a second
|
||||
income (proving) for the commodity side that the specialised supplier cannot enter. The board passes all seven at
|
||||
a 1 to 3 year life; the die fails (a), (b), (c), (e) and (f) at every life, price path and electricity price once
|
||||
it exists. **Against the pass line**: the board's pass does not need a small network (it holds at 42 TH/s and IGN
|
||||
income (proving) for the commodity side that the specialised supplier's hash engine cannot earn (amended 19:5x UK
|
||||
on review B's F09(g): its owner can buy into it with companion GPUs at the GPU entrant's cost, so what remains to
|
||||
the commodity side is the sunk card's hardware term on the proving income, section 5a). The board passes six of
|
||||
seven at a 1 to 3 year life (it fails (c), as every chip does: corrected 19:5x UK on F09(c), the first cut's cell
|
||||
contradicted its own inequality); the die fails (a), (b), (c), (e) and (f) at every life, price path and
|
||||
electricity price once it exists at the silicon floor, and (a), (b), (c), (f) at the reconciled ticket (12c). **Against the pass line**: the board's pass does not need a small network (it holds at 42 TH/s and IGN
|
||||
1.00), token appreciation (it holds on the flat and shrinking paths) or scheduled ASIC death (its life axis is 1 to
|
||||
5 years and the 180-day rotation is not what holds it); the die's failure is not cured by any of the three either,
|
||||
and the only condition that holds the die is that nobody pays to build it (the surface: IGN 0.73 for a USD 150 M
|
||||
|
|
@ -214,6 +217,40 @@ day with a fault) the internal rows are 0.4 to 0.9 USD a day, still above mining
|
|||
classes. The condition this adds to section 12 is (g): a second income for the commodity side that the specialised
|
||||
supplier cannot enter; it holds while proving demand exists and the 16 GB and larger tiers can prove.
|
||||
|
||||
### 5a. Amendment, 19:5x UK (review B's F09(g)): the hybrid operator who owns companion GPUs or buys proofs
|
||||
|
||||
The first cut treated proving income as closed to the specialised supplier because its hash engine cannot prove. A
|
||||
company can own companion GPUs, or buy proofs; the single-device limitation does not exclude the operator. The test:
|
||||
for a hybrid operator (the chip for the hash, companion GPUs bought at MSRP for the proving), the proving surplus per
|
||||
card-day is the proving income less the companion card's entrant cost per day (hardware over two years with resale,
|
||||
plus power at 0.12 at the proving watts); for the sunk GPU owner it is the income less power alone. At IGN 0.10, the
|
||||
pool shared by a 5,000-card fleet, external demand USD 2,000 a day at launch and 20,000 at a spike (build-4, 19:54 UK;
|
||||
all modelled; the 12 GB tier's internal income the measured zero):
|
||||
|
||||
| Class | Proving income per card-day, launch / spike | Companion card's hardware per day at MSRP | Power per day at 0.12 | The hybrid operator's surplus, launch / spike | The sunk owner's surplus, launch / spike |
|
||||
|---|---|---|---|---|---|
|
||||
| RTX 5090 | 10.49 / 13.61 | 1.23 | 0.86 | 8.40 / 11.52 | 9.63 / 12.75 |
|
||||
| RTX 5080 | 8.26 / 10.72 | 0.68 | 0.63 | 6.95 / 9.40 | 7.63 / 10.09 |
|
||||
| RTX 5070 Ti | 9.44 / 12.25 | 0.51 | 0.58 | 8.36 / 11.16 | 8.87 / 11.68 |
|
||||
| RTX 5070 (12 GB) | 0.45 / 4.55 | 0.38 | 0.40 | -0.32 / 3.77 | 0.05 / 4.15 |
|
||||
| RTX 4090 | 10.49 / 13.61 | 1.20 | 0.81 | 8.48 / 11.60 | 9.69 / 12.81 |
|
||||
| RTX 4080 | 8.81 / 11.43 | 0.82 | 0.63 | 7.36 / 9.98 | 8.18 / 10.80 |
|
||||
| RTX 4070 (12 GB) | 0.18 / 1.80 | 0.45 | 0.32 | -0.59 / 1.04 | -0.14 / 1.49 |
|
||||
| RTX 3090 (used) | 4.44 / 5.76 | 1.44 | 0.66 | 2.34 / 3.66 | 3.77 / 5.09 |
|
||||
| H100 (hosted) | 22.04 / 28.59 | 17.11 | 1.01 | 3.92 / 10.47 | 21.03 / 27.58 |
|
||||
| A100 (used) | 13.22 / 17.15 | 8.90 | 0.72 | 3.60 / 7.53 | 12.50 / 16.43 |
|
||||
|
||||
Reading: at launch-shape demand the proving income on a 16 GB or larger consumer card is 4x to 20x the companion
|
||||
card's own daily cost, so a hybrid operator buys companion GPUs and captures the proving income with a surplus within
|
||||
10 to 15 percent of the sunk owner's (the hardware term, USD 0.5 to 1.4 a day on a consumer card); the datacentre
|
||||
parts are the exception (the hybrid keeps a fifth of the owner's surplus). Buying proofs instead of making them
|
||||
yields no income and is not a route to it. **Condition (g) is therefore not an exclusive advantage of the commodity
|
||||
side**: proving is a second market open to anyone who buys GPUs, which the chip's owner can and would; what remains to
|
||||
the commodity side is the sunk card's hardware term on that income (10 to 15 percent of the surplus on consumer
|
||||
cards), and the structural point that the proving fleet is GPUs whoever owns them. Condition (g) in section 12 is
|
||||
re-read accordingly: it holds for the GPU as a device (the proving fleet is GPUs) and not for the GPU owner as a
|
||||
business against a hybrid operator; it does not move any chip's verdict, since no chip passed (c) or (f) on it.
|
||||
|
||||
## 6. Private mining and hardware sales; cheaper derivative chips (D4)
|
||||
|
||||
From the surface (the floor file 4.4, the mission lane's shape), `p*` is the break-even price in USD per IGN:
|
||||
|
|
@ -388,14 +425,34 @@ operator simulation's D1), a GPU fleet of the same hash is tens of thousands of
|
|||
|---|---|---|---|---|
|
||||
| (a) the chip's all-in cost per accepted unit at its own electricity within about 1.5x of the best GPU owner's at the GPU's electricity | the owner at 0.12: 263 to 332 micro-USD (5070 Ti, 5080) | 307 at 3 y, 750 at 1 y: PASSES | 211 / 190 (node-for-node / a node ahead): 1.24x / 1.39x, PASSES | 72 to 134: FAILS at every life |
|
||||
| (b) the chip's annualised hardware per MH/s not below about a quarter of the GPU entrant's | the 5080 entrant's hardware 460 micro-USD | 239 to 677: PASSES | 0.27 of it: PASSES by a hair | 33 to 95 (0.07): FAILS |
|
||||
| (c) a fleet above a third of the chain's hash costs more than a year's miner revenue | at IGN 0.10 the chain is 9.6 TH/s, a third 3.2 | USD 18 M of boards against USD 40 to 80 M: PASSES above IGN 0.05 (at USD 4.84, 15 M: FAILS by 2.6x) | USD 9.4 M against 40 M: FAILS by 4.3x (2.2x to 9.7x across IGN 0.03 to 1.00) | USD 2.5 M of dies: FAILS by 16x; at every price under about 3 |
|
||||
| (c) a fleet above a third of the chain's hash costs more than a year's miner revenue (the inequality: a third of the equilibrium hash x the ticket in USD per MH/s > the year's miner revenue) | at IGN 0.03 / 0.10 / 0.30 / 1.00 the chain is 5.6 / 9.6 / 19.7 / 42.3 TH/s, a third 1.9 / 3.2 / 6.6 / 14.1; the year's revenue USD 12 / 40 / 120 / 400 M | USD 10 / 18 / 37 / 79 M of boards at USD 5.6 per MH/s: **FAILS at every price by 1.2x to 5.1x** (the first cut's "PASSES above IGN 0.05" was a hand-written cell contradicting the inequality, corrected 19:5x UK on review B's F09(c); at the final USD 4.84, 15 M against 40 M, 2.6x short) | USD 9.4 M against 40 M: FAILS by 4.3x (2.2x to 9.7x across IGN 0.03 to 1.00) | USD 1.5 / 2.5 / 5.3 / 11.3 M of dies at USD 0.8: FAILS by 8x to 35x |
|
||||
| (d) GPUs keep a resale market and a use outside mining | resale 25 to 55 percent after two years; rental 3x to 5x the mining cost | PASSES (the GPU side's property) | PASSES (the same) | PASSES (the same) |
|
||||
| (e) the per-joule gap at the honest knee stays under about 3x | Blackwell at the knee 1.70 to 2.06 microjoules | 2.2x to 2.6x: PASSES | 2.3x node-for-node: PASSES; 3.1x a node ahead: on the line | 3.7x to 4.5x at the chip-model energy: FAILS (the bare-lane floor 10x to 12x) |
|
||||
| (f) the supplier's gross margin is a normal return (under about 70 percent) and does not rise with the halvings | section 8 and 12b | 27 to 69 percent, falling with the halvings: PASSES | 72 to 79 percent on the growing path, 49 to 69 flat, a loss once it is the chain: FAILS | 85 to 93 percent, rising, or a loss once it is the chain: FAILS |
|
||||
| (g) a second income (proving) for the commodity side that the specialised supplier cannot enter | section 5: USD 3 to 7 a card-day at launch demand on the 16 GB and larger tiers; 0 for any hash engine | PASSES (the board cannot prove either, which is the GPU's advantage over it) | PASSES (the same) | PASSES (the same) |
|
||||
| (g) a second income (proving) that the specialised hardware cannot earn and its owner can buy into only at the GPU entrant's cost (amended 19:5x UK, F09(g), section 5a) | section 5: USD 3 to 7 a card-day at launch demand on the 16 GB and larger tiers; 0 for any hash engine; a hybrid operator with companion GPUs keeps 85 to 90 percent of the sunk owner's proving surplus on consumer cards | holds for the GPU as a device (the proving fleet is GPUs whoever owns them); NOT an exclusive advantage of the GPU owner as a business: the hybrid operator captures it | the same | the same |
|
||||
|
||||
**The result (as amended on the pinned ticket, 17:4x UK).** The success statement holds for the stored-dataset
|
||||
DRAM-board chip at a 1 to 3 year life on today's rows, in every price path, at every electricity price on the axis,
|
||||
**Every (c) cell, generated from the inequality and its inputs (review B's F09(c), 19:5x UK; `sim/economy/coexist/market.py`
|
||||
M3 and the equilibrium hashes of section 11):** a third of the chain's hash at the per-class equilibrium, times the
|
||||
chip's ticket, against the year's miner revenue.
|
||||
|
||||
| Ticket (USD per MH/s) | IGN 0.03 (USD 12 M a year; a third of the chain 1.9 TH/s) | 0.10 (40 M; 3.2 TH/s) | 0.30 (120 M; 6.6 TH/s) | 1.00 (400 M; 14.1 TH/s) | Verdict |
|
||||
|---|---|---|---|---|---|
|
||||
| GDDR7 board, chip model 5.6 | 10.4 M, 1.2x short | 17.8 M, 2.2x short | 36.8 M, 3.3x short | 78.9 M, 5.1x short | FAILS at every price |
|
||||
| GDDR7 machine, final 4.84 | 9.0 M, 1.3x | 15.4 M, 2.6x | 31.8 M, 3.8x | 68.2 M, 5.9x | FAILS |
|
||||
| Hybrid, chip model 2.44 | 4.5 M, 2.6x | 7.8 M, 5.1x | 16.0 M, 7.5x | 34.4 M, 11.6x | FAILS |
|
||||
| Hybrid, final 2.80 | 5.2 M, 2.3x | 8.9 M, 4.5x | 18.4 M, 6.5x | 39.5 M, 10.1x | FAILS |
|
||||
| SRAM die, reconciled 1.0 | 1.9 M, 6.5x | 3.2 M, 12.6x | 6.6 M, 18x | 14.1 M, 28x | FAILS |
|
||||
| SRAM die, silicon floor 0.8 with the board (0.6 bare) | 1.5 M, 8x | 2.5 M, 16x | 5.3 M, 23x | 11.3 M, 35x | FAILS |
|
||||
|
||||
Condition (c) as written is met by no chip at any price in the window: a chip fleet holding a third of the chain always
|
||||
costs less than a year of the chain's miner revenue, because every chip in the model is 3x to 15x cheaper per MH/s than
|
||||
the GPUs that set the equilibrium. The condition therefore does not discriminate between the chips; what does is the
|
||||
ratio by which each fails (1.2x to 5.9x for the DRAM board against 8x to 35x for the die), which is the dependence
|
||||
on a single supplier of section 11 read as money. The board's verdict is six of seven at both tickets (it fails (c)
|
||||
only), not seven; the first cut's seven was the hand-written cell.
|
||||
|
||||
**The result (as amended on the pinned ticket, 17:4x UK, and on F09(c) at 19:5x).** The success statement holds for the stored-dataset
|
||||
DRAM-board chip on six of the seven conditions at a 1 to 3 year life on today's rows, in every price path, at every electricity price on the axis,
|
||||
with the GPU side's own generation curve narrowing the gap further and proving as a second income the chip cannot
|
||||
enter; its pass needs no small network (it holds at 42 TH/s and IGN 1.00), no token appreciation (it holds on the flat
|
||||
and shrinking paths) and no scheduled ASIC death (the life axis, not the rotation, is what the board lives on). For
|
||||
|
|
|
|||
|
|
@ -541,6 +541,38 @@ half-profit split for hardware sales, the 20 percent residual, the 10 percent di
|
|||
re-entry at a higher price (which raises the hash and lowers the chip's revenue per MH/s only if GPU costs fall),
|
||||
the per-year `q` path instead of a constant, and the emission beyond year 8.
|
||||
|
||||
## 4a. Amendment, 20:0x UK (review B's F05): the address coupling's cost on the chip, the literal 63-fold against the cached prefix tree
|
||||
|
||||
Review B's F05 (findings.json on review-b-landing): the connected-state class's full-chain window computes a
|
||||
rotate-XOR fold over the 63 registers other than the source and XORs the source; it connects every register
|
||||
syntactically, but a physical implementation need not reread and fold all 63 on each load. With `a[k] = ROL32(r[k],
|
||||
63 - k)`, `S` the XOR of all `a[k]` and `P_s` the XOR of `a[k]` for `k < s`, the source expression is exactly `r[s] XOR
|
||||
ROR32(P_s, 1) XOR S XOR P_s XOR a[s]` (the review checked 33,024 comparisons including 4,096 state updates with no
|
||||
mismatch), and a prefix-XOR tree gives point updates and prefix queries in logarithmic time. The review's own
|
||||
caveat: cached prefix state, port bandwidth and update costs must all be priced; this is not evidence the hash is
|
||||
cheap. The chip-side cost of both designs at the adversary lane's PLACED per-lane-op figures (its mf core: 9.36 pJ per
|
||||
lane-op at ASAP7 placed and routed, 6.55 at N5, 4.72 at N3; the window's gated file charging per write, the k lane's
|
||||
+1.2 pJ per lane-op at N5 for the 64-register window; all modelled, no chip measured):
|
||||
|
||||
| Design | State per lane | Ports and traffic | Ops per hash for the fold (448-instruction text, 128 loads) | Energy per hash for the fold at N5 (6.55 pJ per lane-op) / N3 (4.72) | Against the class's shadow (55,296 lane-ops: 362 nJ at N5, 261 at N3) | Label |
|
||||
|---|---|---|---|---|---|---|
|
||||
| The literal 63-fold on every load | the 64 x 32-bit window only | 63 register reads per load (2,016 bits over the file's read ports, or a 63-cycle sequential read hidden by lane count) | 63 rotate-XORs x 128 loads = 8,064 | 52.8 nJ / 38.1 nJ | +14.6 percent of the chip's shadow energy; `E_chip` 0.466 + 0.362 + 0.053 = 0.881 microjoules (N5) | modelled |
|
||||
| The cached prefix tree (a Fenwick tree over the 64 `a[k]`, plus the running `S`) | +64 x 32 bits of prefix state and one 32-bit `S`: 260 bytes per lane, the window doubled (the k lane's +1.2 pJ per lane-op window term scales with it, +1.2 at N5, approximate) | per instruction: log2(64) = 6 read-modify-writes of prefix nodes plus the `S` update (7 ops, 7 x 32 bits each way); per load: 6 node reads for `P_s`, the rotation, 4 XORs (11 ops) | 7 x 448 updates + 11 x 128 queries = 4,544 | 29.8 nJ / 21.4 nJ, plus the doubled window's +1.2 pJ on 55,296 ops: +66 nJ at N5 if the whole file widens, +0 if the prefix state sits in a separate gated macro written 7 times per instruction (the designer's choice; the latter is the cheaper and is carried) | +8.2 percent of the shadow; `E_chip` 0.858 microjoules (N5) | modelled |
|
||||
| The difference | +260 B per lane | 7 ports' traffic per instruction against 63 reads per load | 1.8x fewer ops for the fold | 23 nJ per hash saved at N5 | **the chip's energy per hash falls 2.6 percent (0.881 to 0.858); the connected-state ratio of that lane's section 5 moves from 2.69x to 2.76x at the lock** (E_GPU 2.37 microjoules with the window at the lock, measured +1.6 percent) | modelled |
|
||||
|
||||
Reading: the cached-prefix design is the one the adversary builds (it halves the fold's ops for 260 bytes of state
|
||||
per lane and seven narrow writes per instruction), and it moves the chip's energy per hash by about 3 percent, which
|
||||
moves the connected-state verdict (KILL as a class, that lane's section 5: 1.08x node-for-node against a 1.25x gate)
|
||||
nowhere: the fold was never the chip's cost, the shadow's op count was. The GPU side: the measured window cost on the
|
||||
card (+1.6 percent of energy per hash at the 5090's lock, +5.9 unlocked, that lane's section 4) was taken with the
|
||||
compiler's own code for the fold, which may already hoist part of it; a byte-preserving incremental form on the GPU
|
||||
would be the matching experiment and is the review's "cheapest alternative first". The review's required regressions
|
||||
(native reference against incremental equivalence across every source register and real instruction updates; full
|
||||
kernel equivalence, registers, spills, wall power and accepted throughput on the target cards; the adversary's design
|
||||
carrying the prefix cost and never an assumed 63-read cost) are the connected-state lane's and the adversary lane's
|
||||
to run; this file carries the price only, and the price says the class's verdict is unchanged by F05. Nothing
|
||||
unmeasured and nonlinear is added in response, as the review asks.
|
||||
|
||||
## 5. Lever 4: everything else that reaches the die, checked
|
||||
|
||||
| Candidate | What it costs the SRAM die | What it costs the honest tiers | The fresh-join path | Verdict | Label |
|
||||
|
|
|
|||
Loading…
Reference in a new issue