Merge remote-tracking branch 'build/master' into tn-land
This commit is contained in:
commit
b42eb1413d
4 changed files with 884 additions and 1 deletions
299
docs/analysis/class-v6/coexistence-model.md
Normal file
299
docs/analysis/class-v6/coexistence-model.md
Normal file
|
|
@ -0,0 +1,299 @@
|
|||
# The coexistence model: can a specialised supplier earn a normal return while GPUs stay close enough to compete?
|
||||
|
||||
8 October 2026, 16:1x to 16:5x UK, branch `class-v6-floor-sram`, floor lane 3, on the research lane's word of 17:0x UK
|
||||
(the founder's accepted third external review: the capex wall is replaced by a coexistence model as the economic
|
||||
argument, with the profitability surface of `docs/analysis/class-v6/floor/sram-and-floor.md` section 4.4 as its base).
|
||||
First run on today's measured rows. **Every row is modelled**: the arithmetic is `scratchpad/coexist.py` run on
|
||||
build-3; the card rows are lane 4's class v5 table (`docs/analysis/class-v6/floor/denominator.md`, section 10.4 of
|
||||
the design document: measured at the floor where it says so, modelled knees elsewhere), the chip rows the chip model's
|
||||
(chip-model-v3 5.5 and 5.12, lane B, the k lane's 3.2 pJ per forced op), the prices street approximations. Nothing
|
||||
here is served; nothing is a measurement of a chip. The founder is not named.
|
||||
|
||||
The statement under test. **Success**: a specialised supplier earns a normal return and GPUs stay close enough in total
|
||||
cost, obtainable and useful outside mining, that entrants still compete. **Failure**: a supplier operating privately at
|
||||
much lower cost exhausts competitors' margins. The result is the set of conditions under which the success statement
|
||||
holds, never a level the chain stays below. The development cost is SUNK in the mandatory case (the opponent covers
|
||||
manufacturing, deployment and operation only); the paid-development cases sit beside it.
|
||||
|
||||
## 0. The result in one page
|
||||
|
||||
1. **On today's rows the N2 SRAM die fails the success statement in every scenario where its owner can buy a fleet
|
||||
worth a few percent of the chain's yearly miner revenue.** With development sunk, a 3-year life and power at USD
|
||||
0.06 per kWh, the die's all-in cost per accepted MH/s-hour is 72 micro-USD against the best GPU owner's 156 to 204
|
||||
at the same electricity (the 5070 Ti and 5080 at their knees, hardware sunk) and 415 to 588 for a new entrant: 2.2x
|
||||
to 2.8x on the existing owner, 5.8x to 8.2x on the entrant. At the GPU's more likely 0.12 per kWh it is 3.6x to
|
||||
4.6x and 7.2x to 9.9x; at 0.25, 6.8x to 8.5x and 10x to 14x. A USD 10 M fleet of such dies takes 37 to 73 percent
|
||||
of the chain's hash the year it lands at IGN 0.10 to 0.20 and 100 percent the year after on a flat price; a USD
|
||||
100 M fleet takes 100 percent on landing in every path and then runs at a loss because it is larger than the
|
||||
revenue (the self-limiting point: -1 to -300 percent margins).
|
||||
2. **The GDDR7 board chip passes the success statement on its 1-year life and fails it on its 3-year life.** At 1 year
|
||||
its cost per accepted unit (750 micro-USD at 0.06) is ABOVE every Blackwell owner's and above the Blackwell
|
||||
entrant's at 0.06 to 0.12 (0.8x to 1.0x the 5080 entrant, 0.6x to 0.7x the 5070 Ti), so it competes only against
|
||||
Ada and Ampere at dear electricity; at 3 years (307 micro-USD) it is 1.9x to 2.3x the Blackwell entrant and 0.9x
|
||||
to 1.1x the Blackwell owner: GPUs at their knee stay close enough. Its hardware is the term that holds it (USD 5.6
|
||||
per MH/s with the core and the system against the die's 0.8), not its joules.
|
||||
3. **The electricity axis is explicit and it is the GPU's.** The break-even electricity price above which an EXISTING
|
||||
GPU owner with sunk hardware cannot match the die at 0.06 is 0 to 5 cents per kWh for every card (a 3-year die) and
|
||||
1 to 5 cents (a 1-year die): no grid price in the world keeps a GPU owner level with a sunk SRAM die. Against the
|
||||
GDDR7 board at 3 years the owner's break-even is 9 to 15 cents on Blackwell, 4 to 6 on Ada and Ampere; at 1 year 27
|
||||
to 40 cents on Blackwell. Read the other way, the die breaks even against a 5080 owner at 0.12 only when the die
|
||||
pays 47 to 60 cents per kWh; the board at 3 years when it pays 1 to 9 cents.
|
||||
4. **Accessible supply is the structural fact.** At the GPU entry equilibrium (revenue per MH/s-hour equal to a new
|
||||
5070 Ti's cost at 0.12) the chain's hash is 2.6 TH/s at IGN 0.03, 8.8 at 0.10, 26 at 0.30, 88 at 1.00: that is
|
||||
34,000 to 1.1 M 5070 Ti-class cards, which exist in the world's installed base, against 480 to 16,000 SRAM dies, 8
|
||||
to 270 wafers of N2. One supplier holds the chain at every price in the window; the dependence on individual
|
||||
suppliers is total for the die and partial for the board (16,000 to 530,000 boards, a Bitmain-class run).
|
||||
5. **The conditions under which the success statement holds**, read off the tables: (a) the chip's all-in cost per
|
||||
accepted unit stays within about 1.5x of the best GPU owner's at the same electricity, which on today's rows is true
|
||||
of the GDDR7 board at a life of 1 to 2 years and false of the die at every life; (b) the chip's hardware per MH/s is
|
||||
not below about a quarter of the GPU's annualised hardware (the board's 0.7x to 1.9x the 5080 entrant passes, the
|
||||
die's 5x to 14x fails); (c) no single buyer can fund a fleet above about a third of the chain's hash for under a
|
||||
year's miner revenue (true for the board above IGN 0.3; false for the die at every price: a wafer is USD 30,000);
|
||||
(d) GPUs keep a resale market and a use outside mining (true for every card in the population; the chip has none,
|
||||
which is why its life is the axis that moves everything); (e) the per-joule gap at the honest knee stays under about
|
||||
3x (the board at 2.2x to 2.6x against Blackwell passes; the die at 3.7x to 4.5x does not).
|
||||
6. **What the chain controls, and what it does not.** It controls the honest cost per accepted unit (the operating
|
||||
point: the lock is worth 34 to 41 percent of a Blackwell card's draw), the chip's life against a fixed lane (the
|
||||
180-day rotation; nothing against a programmable one), and the visibility of a concentrated supplier (the share
|
||||
detector). It does not control the sunk development cost, electricity prices, the token price or a buyer's budget.
|
||||
On today's rows the SRAM die, once it exists, cannot be held to coexistence by anything in the hash; the board can.
|
||||
The condition that holds the die is that it does not get built, which is the surface of section 4.4 (a USD 150 M
|
||||
project at a third of the chain needs IGN 0.73 over three years), and that is a statement about who pays, not about
|
||||
the chain staying below a level.
|
||||
|
||||
## 1. The measure: cost per accepted unit of work
|
||||
|
||||
Cost per accepted MH/s-hour (micro-USD), both sides: annualised hardware plus power plus hosting, failures and fees,
|
||||
divided by accepted work. Accepted work is 97 percent of raw (rejects 0.5 percent, downtime 2.0, epoch preparation and
|
||||
propagation 0.5; approximate from the devnet's share rates), the pool fee 1 percent. The two GPU situations: the
|
||||
**existing owner** (hardware sunk; pays power, a wear allowance of 5 percent of the used price a year, fees; the
|
||||
alternative use is the card's rental yield, reported in section 2 and not deducted) and the **new entrant** (buys new
|
||||
or used, operates two years, resells at the table's fraction). Hosting: 0 at home, USD 0.02 per kWh-equivalent at a
|
||||
farm (the chip's case). Electricity axis: USD 0.06, 0.12, 0.25 per kWh.
|
||||
|
||||
| Input | Value | Label |
|
||||
|---|---|---|
|
||||
| Card joules per hash at the floor (the knee with the knobs) and at stock | lane 4's class v5 table: 5090 2.33 / 3.48, 5080 2.06 / 3.48, 5070 Ti 1.70 / 2.84, 5070 1.75 / 2.99, 5060 Ti 2.36 / 4.02, 5060 2.21 / 3.76, 4090 3.58 / 5.00, 4080 3.51 / 4.92, 4070 3.58 / 5.82, 4060 Ti 3.81 / 5.32, 3090 4.51 / 5.03, 3080 4.20 / 4.54, 3060 5.77 / 6.40, 9070 XT 7.90 / 10.7, H100 2.00 / 2.58, A100 2.89 / 2.99, M5 Max 1.40 microjoules | measured where lane 4 says so (5090, 5080, 4070, 9070 XT, M5 Max floors; most stock rows), modelled knees elsewhere |
|
||||
| Card rates at the floor | 5090 134.8, 5080 71.2, 5070 Ti 77, 5070 41, 5060 Ti 19, 5060 17, 4090 58, 4080 45, 4070 31.1, 4060 Ti 17.6, 3090 37.8, 3080 40.8, 3060 23.8, 9070 XT 18.9, H100 90, A100 60, M5 Max 27.1 MH/s | measured where the bench table has the row; approximate elsewhere |
|
||||
| Prices new / used, resale after two years | 5090 2,600 / 2,200 / 55 percent; 5080 1,100 / 900 / 50; 5070 Ti 800 / 650 / 50; 5070 560 / 450 / 50; 5060 Ti 450 / 360 / 45; 5060 310 / 250 / 45; 4090 1,700 / 1,300 / 45; 4080 1,000 / 700 / 40; 4070 550 / 400 / 40; 4060 Ti 420 / 290 / 35; 3090 900 / 650 / 30; 3080 450 / 330 / 25; 3060 260 / 190 / 25; 9070 XT 650 / 520 / 45; H100 25,000 / 18,000 / 50; A100 10,000 / 6,000 / 35; M5 Max 4,000 / 3,200 / 55 | approximate street, October 2026 |
|
||||
| Chip joules per hash (class v5, the shadow at 3.2 pJ per forced op, lane 4's chip columns) | GDDR7 board with an N5 core 0.79; HBM3 one stack 0.65; N2 SRAM die with the core 0.46 (W = 4); the die at the bare-lane floor 0.165 | modelled |
|
||||
| Chip hardware, USD per MH/s, silicon and board plus 30 percent system (PSU, chassis, cooling) | GDDR7 board 5.6 (4.3 with the core die, chip-model 5.5 and research 16.1); HBM3 8.6; SRAM die 0.8 (0.6 with the board, lane B) | modelled, approximate |
|
||||
| Chip failures, resale | 3 percent a year; no resale (single use) | approximate |
|
||||
| Chip lives | 0.5, 1, 2, 3, 5 years (0.5 is the fixed-lane chip under the 180-day rotation; 3 the programmable chip's default) | the review's axis |
|
||||
| Emission to miners | 0.77 / 0.80 / 0.40 / 0.40 / 0.20 B IGN in years 1 to 5 (the spec's constant, 80 percent to miners) | spec 05 |
|
||||
| GPU equilibrium | while any GPU mines, revenue per MH/s-hour settles at the cheapest entrant's cost (a new 5070 Ti at 0.12: 521 micro-USD) during growth and falls to the owners' costs during shrinkage; a GPU generation at year 3 cuts the entrant's cost 25 percent (1.5x per joule at the same price) with the old card resold at the table's fraction | modelled rule |
|
||||
|
||||
## 2. The GPU reference population: cost per accepted MH/s-hour (micro-USD)
|
||||
|
||||
| Card | Owner at 0.06 / 0.12 / 0.25 | Entrant, new, 2 years, at 0.06 / 0.12 / 0.25 | Entrant, used | Hardware share of the entrant's cost at 0.12 | Wh per MH/s-hour | Alternative use (rental yield, approximate) | Label |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| RTX 5090 | 243 / 388 / 704 | 661 / 807 / 1,122 | 800 / 945 / 1,261 | 64 percent | 2.33 | USD 0.3 to 0.5 an hour on a rental market: 2,200 to 3,700 micro-USD per MH/s-hour, 3x to 5x its mining cost | measured floor |
|
||||
| RTX 5080 | 204 / 332 / 611 | 588 / 716 / 995 | 650 / 779 / 1,058 | 64 | 2.06 | USD 0.15 to 0.25 an hour | measured floor |
|
||||
| RTX 5070 Ti | 156 / 263 / 493 | 415 / 521 / 751 | 453 / 560 / 790 | 59 | 1.70 | USD 0.1 to 0.2 an hour | modelled knee |
|
||||
| RTX 5070 | 175 / 284 / 521 | 515 / 624 / 861 | 558 / 668 / 905 | 65 | 1.75 | | modelled knee |
|
||||
| RTX 5060 Ti 16 GB | 260 / 407 / 727 | 921 / 1,069 / 1,388 | 956 / 1,104 / 1,423 | 72 | 2.36 | | modelled knee |
|
||||
| RTX 5060 | 225 / 364 / 663 | 734 / 872 / 1,171 | 768 / 906 / 1,205 | 68 | 2.21 | | modelled knee |
|
||||
| RTX 4090 | 357 / 580 / 1,065 | 1,181 / 1,405 / 1,890 | 1,163 / 1,387 / 1,872 | 68 | 3.58 | USD 0.3 to 0.4 an hour | modelled knee, stock measured |
|
||||
| RTX 4080 | 312 / 531 / 1,006 | 1,011 / 1,231 / 1,706 | 879 / 1,099 / 1,574 | 64 | 3.51 | | modelled knee |
|
||||
| RTX 4070 | 300 / 524 / 1,008 | 854 / 1,078 / 1,562 | 778 / 1,001 / 1,486 | 58 | 3.58 | | measured tune |
|
||||
| RTX 4060 Ti 16 GB | 336 / 574 / 1,090 | 1,159 / 1,397 / 1,913 | 969 / 1,207 / 1,723 | 66 | 3.81 | | modelled knee |
|
||||
| RTX 3090 (used) | 384 / 666 / 1,276 | 1,272 / 1,554 / 2,164 | 1,091 / 1,373 / 1,983 | 64 | 4.51 | | modelled cap |
|
||||
| RTX 3080 (used) | 310 / 573 / 1,141 | 754 / 1,016 / 1,585 | 661 / 923 / 1,492 | 48 | 4.20 | | modelled cap |
|
||||
| RTX 3060 (used) | 408 / 768 / 1,550 | 847 / 1,208 / 1,989 | 754 / 1,114 / 1,895 | 40 | 5.77 | | modelled cap |
|
||||
| RX 9070 XT | 657 / 1,151 / 2,220 | 1,617 / 2,111 / 3,180 | 1,668 / 2,162 / 3,231 | 53 | 7.90 | | measured |
|
||||
| H100 (hosted) | 1,313 / 1,438 / 1,709 | 8,374 / 8,499 / 8,770 | 7,880 / 8,004 / 8,275 | 97 | 2.00 | USD 2 to 3 an hour: never mines | modelled lock |
|
||||
| A100 (used) | 775 / 955 / 1,346 | 6,615 / 6,796 / 7,187 | 4,388 / 4,568 / 4,960 | 95 | 2.89 | USD 1 an hour: never mines | modelled |
|
||||
| Apple M5 Max (reported, not headlined) | 789 / 876 / 1,066 | 4,033 / 4,120 / 4,310 | 4,690 / 4,778 / 4,967 | 96 | 1.40 | a workstation: mines only as an owner | measured |
|
||||
|
||||
Reading: the owner's cost is 60 to 70 percent electricity on every consumer card, so the electricity price is the
|
||||
GPU's whole variable; the entrant's cost is 60 to 70 percent hardware, so the card price and its resale are the
|
||||
entrant's whole variable. The best honest owner on today's rows is a 5070 Ti at its knee (156 micro-USD at 0.06); the
|
||||
best entrant the same card (415). Datacentre parts and the Mac never enter as entrants (hardware 95 percent) and mine
|
||||
only as owners with nothing better to do, which their rental yields say they always have.
|
||||
|
||||
## 3. The chip rows
|
||||
|
||||
### 3.1 Development sunk (the mandatory case): cost per accepted MH/s-hour, micro-USD, farm hosting
|
||||
|
||||
| Chip | Life 0.5 y at 0.06 / 0.12 / 0.25 | 1 y | 2 y | 3 y | 5 y | Hardware / power split at 0.06, 3 y | Label |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| GDDR7 board with an N5 core | 1,414 / 1,463 / 1,570 | 750 / 799 / 906 | 418 / 467 / 574 | 307 / 356 / 463 | 219 / 268 / 375 | 239 / 65 | modelled |
|
||||
| HBM3 one stack with a core | 2,123 / 2,164 / 2,252 | 1,104 / 1,145 / 1,233 | 594 / 635 / 723 | 425 / 465 / 553 | 289 / 329 / 417 | 367 / 54 | modelled |
|
||||
| N2 SRAM die with the core | 226 / 255 / 317 | 134 / 163 / 225 | 87 / 116 / 178 | **72 / 101 / 163** | 60 / 88 / 151 | 33 / 38 | modelled |
|
||||
| N2 SRAM die at the bare-lane floor | 202 / 212 / 235 | 109 / 120 / 142 | 63 / 73 / 96 | 47 / 58 / 80 | 35 / 45 / 68 | 33 / 14 | modelled, the worst case |
|
||||
|
||||
### 3.2 Development paid: the same rows with `C_dev` spread over the fleet and the life
|
||||
|
||||
The fleet is sized to a share `q` of the network's hash at the GPU equilibrium (a new 5080 at 0.12 as the marginal
|
||||
entrant). All-in cost per accepted MH/s-hour of the SRAM die (the GDDR7 board in brackets), at 0.06:
|
||||
|
||||
| IGN price | Miner revenue (year 3) | Network hash | `C_dev` 20 M, 1 y, `q` 0.3 / 1.0 | 20 M, 3 y, 0.3 / 1.0 | 150 M, 1 y, 0.3 / 1.0 | 150 M, 3 y, 0.3 / 1.0 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 0.03 | USD 12 M | 1.9 TH/s | 4,277 / 1,377 (4,893 / 1,993) | 1,453 / 486 (1,688 / 721) | 31,211 / 9,457 | 10,431 / 3,180 |
|
||||
| 0.10 | 40 M | 6.4 TH/s | 1,377 / 507 (1,993 / 1,123) | 486 / 196 (721 / 431) | 9,457 / 2,931 | 3,180 / 1,004 |
|
||||
| 0.30 | 120 M | 19 TH/s | 548 / 258 (1,164 / 874) | 210 / 113 (445 / 349) | 3,242 / 1,066 | 1,108 / 383 |
|
||||
| 1.00 | 400 M | 64 TH/s | 258 / 171 (874 / 787) | 113 / 84 (349 / 320) | 1,066 / 414 | 383 / 165 |
|
||||
|
||||
Reading: a paid development cost puts every chip ABOVE the best GPU entrant (415 to 588) at IGN 0.10 and below, and
|
||||
the SRAM die below it only from IGN 0.30 on a 3-year life or IGN 1.00 on a 1-year life; the GDDR7 board with paid
|
||||
development is never below the Blackwell entrant inside the window. The sunk case is therefore the whole threat, and
|
||||
it is the case the review makes mandatory. The derivative design (a revision at 0.3 x `C_dev`, section 4.4 of the
|
||||
floor file) moves the paid rows a third of the way to the sunk rows.
|
||||
|
||||
## 4. The cost advantage, separated
|
||||
|
||||
The chip at 0.06 per kWh and farm hosting; the GPU at 0.06, 0.12 and 0.25. "Operating" is power only (joules times
|
||||
electricity); "hardware" the annualised hardware of the GPU entrant over the chip's.
|
||||
|
||||
| Chip, life | Against | Total: owner / entrant, at 0.06 | At 0.12 | At 0.25 | Operating advantage at 0.06 / 0.12 / 0.25 (joules x electricity) | Hardware advantage |
|
||||
|---|---|---|---|---|---|---|
|
||||
| GDDR7 board, 1 y | RTX 5080 | 0.3x / 0.8x | 0.4x / 1.0x | 0.8x / 1.3x | 2.6x / 5.2x / 10.9x (2.6 x 1, 2, 4.2) | 0.7x |
|
||||
| | RTX 5070 Ti | 0.2x / 0.6x | 0.4x / 0.7x | 0.7x / 1.0x | 2.2x / 4.3x / 9.0x | 0.5x |
|
||||
| | RTX 4090 | 0.5x / 1.6x | 0.8x / 1.9x | 1.4x / 2.5x | 4.5x / 9.1x / 18.9x | 1.4x |
|
||||
| | RTX 3080 (used) | 0.4x / 1.0x | 0.8x / 1.4x | 1.5x / 2.1x | 5.3x / 10.6x / 22.2x | 0.7x |
|
||||
| GDDR7 board, 3 y | RTX 5080 | 0.7x / 1.9x | 1.1x / 2.3x | 2.0x / 3.2x | the same | 1.9x |
|
||||
| | RTX 5070 Ti | 0.5x / 1.4x | 0.9x / 1.7x | 1.6x / 2.4x | | 1.3x |
|
||||
| | RTX 4090 | 1.2x / 3.8x | 1.9x / 4.6x | 3.5x / 6.2x | | 4.0x |
|
||||
| | RTX 3080 (used) | 1.0x / 2.5x | 1.9x / 3.3x | 3.7x / 5.2x | | 2.1x |
|
||||
| N2 SRAM die, 1 y | RTX 5080 | 1.5x / 4.4x | 2.5x / 5.4x | 4.6x / 7.4x | 4.5x / 9.0x / 18.7x | 4.9x |
|
||||
| | RTX 5070 Ti | 1.2x / 3.1x | 2.0x / 3.9x | 3.7x / 5.6x | 3.7x / 7.4x / 15.4x | 3.3x |
|
||||
| | RTX 4090 | 2.7x / 8.8x | 4.3x / 10.5x | 8.0x / 14.1x | 7.8x / 15.6x / 32.4x | 10.1x |
|
||||
| N2 SRAM die, 3 y | RTX 5080 | 2.8x / 8.2x | 4.6x / 9.9x | 8.5x / 13.8x | the same | 13.8x |
|
||||
| | RTX 5070 Ti | **2.2x / 5.8x** | **3.6x / 7.2x** | 6.8x / 10.4x | | 9.3x |
|
||||
| | RTX 4090 | 5.0x / 16.4x | 8.1x / 19.5x | 14.8x / 26.2x | | 28.7x |
|
||||
| | RTX 3080 (used) | 4.3x / 10.5x | 8.0x / 14.1x | 15.9x / 22.0x | | 14.7x |
|
||||
| | Apple M5 Max (reported) | 11.0x / 56x | 12.2x / 57x | 14.8x / 60x | 3.0x / 6.1x / 12.7x | 118x |
|
||||
|
||||
Reading: the electricity axis multiplies the operating advantage one for one (a 2.6x chip at 0.06 against a GPU at
|
||||
0.25 is 10.9x on power alone, the review's point), but on the consumer cards power is 60 to 70 percent of the owner's
|
||||
cost and 30 to 40 percent of the entrant's, so the total advantage is a third to a half of the operating one. The
|
||||
GDDR7 board's total advantage over a Blackwell card at its knee is under 1x (owner) to 2.3x (entrant) across the whole
|
||||
electricity axis at a 3-year life, which is the coexistence band; the die's is 2.2x to 14x, which is not.
|
||||
|
||||
## 5. The replacement economics
|
||||
|
||||
For an existing GPU owner, switching pays when the chip's all-in cost per accepted unit is below the owner's
|
||||
OPERATING cost (the hardware is sunk, the resale value is the only thing the switch recovers). For a new entrant, when
|
||||
the chip's all-in is below the GPU entrant's all-in. The chip must be purchasable for either (hardware sales; the
|
||||
manufacturer keeps about half the operator's profit through the price, floor file 4.4 table B, which roughly doubles
|
||||
the chip's hardware term for the buyer).
|
||||
|
||||
| Who | Against the GDDR7 board (bought, hardware term x2) | Against the SRAM die (bought, x2) | What it means |
|
||||
|---|---|---|---|
|
||||
| A Blackwell owner at 0.06 to 0.12 | never switches: the bought board at 3 years is 550 to 600 micro-USD against the owner's 156 to 332 | switches at 0.12 (the bought die 105 to 134 against 263 to 332) and is near indifferent at 0.06 (105 against 156 to 204) | the die replaces Blackwell owners at normal grid prices; the board never does |
|
||||
| An Ada or Ampere owner at 0.12 | near indifferent at 3 years (550 to 600 against 524 to 768); switches at 0.25 | switches at every electricity price | the board retires the oldest cards only at dear electricity, which the generation upgrade does anyway |
|
||||
| A new entrant choosing between a new 5070 Ti and a bought chip at 0.12 | the board at 3 years (about 600) is 1.15x the card's 521: the card wins; at 1 year the card wins 2x | the bought die (134) is 0.26x the card: the die wins 4x | an entrant market with a bought SRAM die has no GPU entrants; one with a bought board keeps them |
|
||||
| The GPU generation upgrade (year 3: 1.5x per joule at the same price, the old card resold) | cuts the entrant's cost about 25 percent and the owner's power 33 percent: the board at 3 years then reads 1.5x to 1.7x the new entrant, still in band | the die's advantage falls 25 to 33 percent, from 7x to 5x on the entrant: still out of band | the GPU side's own curve narrows the board's gap to nothing by the second generation and never closes the die's |
|
||||
|
||||
## 6. Break-even electricity prices
|
||||
|
||||
(a) The GPU electricity price above which an EXISTING owner (hardware sunk) cannot match the chip's all-in cost at
|
||||
0.06 per kWh:
|
||||
|
||||
| Chip, life | 5090 | 5080 | 5070 Ti | 5070 | 5060 Ti | 4090 | 4080 | 4070 | 3090 | 3080 | 3060 | 9070 XT | M5 Max |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| GDDR7 board, 1 y | 27 c | 32 c | 40 c | 38 c | 26 c | 17 c | 18 c | 18 c | 14 c | 16 c | 12 c | 7 c | 3 c |
|
||||
| GDDR7 board, 3 y | 9 c | 11 c | 15 c | 13 c | 8 c | 5 c | 6 c | 6 c | 4 c | 6 c | 4 c | 2 c | under 0 |
|
||||
| SRAM die, 1 y | 1.5 c | 2.7 c | 4.7 c | 3.8 c | 0.9 c | 0 c | 1.1 c | 1.5 c | 0.7 c | 2.0 c | 1.4 c | under 0 | under 0 |
|
||||
| SRAM die, 3 y | under 0 | under 0 | 1.2 c | 0.4 c | under 0 | under 0 | under 0 | under 0 | under 0 | 0.5 c | 0.4 c | under 0 | under 0 |
|
||||
|
||||
(b) The chip electricity price at which its all-in cost equals a GPU owner's at 0.12 (how much dearer the chip's
|
||||
hosting can be and still match): the GDDR7 board at 3 years 1 to 9 cents against Blackwell (it must be hosted cheaper
|
||||
than the GPU to match an owner) and 75 cents against the Mac; at 1 year it cannot match a Blackwell owner at any
|
||||
price. The SRAM die matches a 5070 Ti owner while paying up to 46 cents (3 years) or 33 cents (1 year), a 5080 owner
|
||||
up to 60 or 47, the Mac up to 174.
|
||||
|
||||
Reading: the board lives or dies on the GPU's grid price, which is the review's electricity axis doing what it should;
|
||||
the die does not see the axis at all.
|
||||
|
||||
## 7. Five years: growing, flat and shrinking networks, with the GPU side reacting
|
||||
|
||||
The sunk chip enters at the start of year 2 with a fleet bought for a budget `B` (USD 1 M, 10 M, 100 M at USD 0.8
|
||||
per MH/s for the die: 1.3, 13 and 128 TH/s); revenue per MH/s-hour `r` settles at the cheapest GPU entrant's cost
|
||||
(a new 5070 Ti at 0.12, 521 micro-USD; 391 after the year-3 generation) while entry continues, and when the chip's
|
||||
fleet alone exceeds the hash that revenue supports, GPU owners exit in cost order until the survivors' costs are
|
||||
covered or none are. Three price paths: growing (x2 a year from 0.10), flat (0.10), shrinking (x0.5 a year from
|
||||
0.30). Emission halves in year 3 and year 5.
|
||||
|
||||
| Path, budget | Year 2 | Year 3 | Year 5 | Reading |
|
||||
|---|---|---|---|---|
|
||||
| Growing, USD 1 M | chip 3.7 percent of 35 TH/s; `r` 521; margin 86 percent | 2.7 percent; owners all above water | 1.4 percent of 93 TH/s | coexistence: the supplier earns 82 to 86 percent margins on a tiny share, GPUs set the price |
|
||||
| Growing, USD 10 M | 37 percent of 35 TH/s | 27 percent | 14 percent | coexistence by dilution only: the share falls as the chain grows and the fleet is fixed; the supplier's margin stays 82 to 86 percent, far above normal |
|
||||
| Growing, USD 100 M | 100 percent; 0 of 17 owner classes above water; `r` falls to 142; margin 49 percent | the same | 100 percent; 2 of 17 owner classes above water at `r` 285 | failure: one buyer holds the chain for five years, GPUs exit in year 2 and only the two best classes could return in year 4 |
|
||||
| Flat, USD 1 M | 7 percent of 17.5 TH/s | 11 percent | 22 percent of 5.8 TH/s | coexistence, the share rising with each halving |
|
||||
| Flat, USD 10 M | 73 percent | 100 percent; 3 of 17 owner classes above water | 100 percent; 0 of 17 | failure by year 3: the halving does the rest |
|
||||
| Flat, USD 100 M | 100 percent; margin -1 percent | margin -102 percent | -305 percent | failure for both: the fleet is larger than the revenue; the buyer loses money and the GPUs are gone (the self-limiting point) |
|
||||
| Shrinking, USD 1 M | 5 percent | 15 percent | 100 percent; 3 of 17 above water | failure in year 5 at USD 4 M of revenue: even a USD 1 M fleet is the chain when the chain is small |
|
||||
| Shrinking, USD 10 M | 49 percent | 100 percent; 1 of 17 | 100 percent; margin -116 percent | failure from year 3 |
|
||||
| The GDDR7 board (sunk), flat, USD 10 M | 10 percent of 17.5 TH/s; margin 41 percent | 15 percent; margin 21 percent | 31 percent; margin 21 percent | coexistence: a normal return (21 to 41 percent gross) on a minority share with GPU entrants still setting the price |
|
||||
|
||||
Reading: the row that passes the success statement as the review states it is the GDDR7 board at a sunk USD 10 M (a
|
||||
21 to 41 percent gross margin, a 10 to 31 percent share, GPUs setting the price, entrants competing). The SRAM die
|
||||
passes only at a fleet under about 1 percent of the chain's yearly miner revenue in a growing network, and fails in
|
||||
every flat or shrinking path by year 3 to 5; above a tenth of a year's revenue it takes the chain in every path. The
|
||||
failure is not a margin the die extracts (its margins collapse once it is the chain); it is the exit of every GPU
|
||||
class, which is the review's definition.
|
||||
|
||||
## 8. Accessible supply and the dependence on individual suppliers
|
||||
|
||||
| IGN price | Network hash at the GPU entry equilibrium | In 5070 Ti-class cards | In 5090s | In N2 SRAM dies (5.5 GH/s) | In N2 wafers (about 60 good dies) | In GDDR7 boards (166 MH/s) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 0.03 | 2.6 TH/s | 34,000 | 19,500 | 480 | 8 | 15,800 |
|
||||
| 0.10 | 8.8 | 114,000 | 65,000 | 1,600 | 27 | 52,800 |
|
||||
| 0.30 | 26 | 341,000 | 195,000 | 4,800 | 80 | 158,000 |
|
||||
| 1.00 | 88 | 1.1 M | 650,000 | 16,000 | 270 | 528,000 |
|
||||
| 3.00 | 263 | 3.4 M | 1.9 M | 48,000 | 800 | 1.6 M |
|
||||
|
||||
The GPU side's accessible supply is the installed base of consumer cards (tens of millions of Ampere, Ada and
|
||||
Blackwell cards in the world, approximate) and the used market, with a use outside mining and a resale price that the
|
||||
population table carries; at every price in the window the hash the chain needs is under 4 percent of one
|
||||
generation's shipments. The die's supply is one supplier's wafer allocation (8 to 800 wafers of a node booked to 2028,
|
||||
claimed); the board's is a Bitmain-class production run (16,000 to 1.6 M units) with a commodity memory bill. The
|
||||
dependence on individual suppliers is total for the die at every price (one order holds the chain), partial for the
|
||||
board (a run of that size is visible and takes months), and nil for the GPU side.
|
||||
|
||||
## 9. The conditions under which the success statement holds (the result)
|
||||
|
||||
Success (a specialised supplier earns a normal return and GPUs stay close enough in total cost, obtainable and useful
|
||||
outside mining, that entrants still compete) holds on today's rows when ALL of the following do:
|
||||
|
||||
| Condition | The number on today's rows | GDDR7 board | SRAM die |
|
||||
|---|---|---|---|
|
||||
| (a) the chip's all-in cost per accepted unit at its own electricity is within about 1.5x of the best GPU owner's at the GPU's electricity | the owner at 0.12: 263 to 332 micro-USD (5070 Ti, 5080) | 307 at 3 years, 750 at 1 year: PASSES | 72 to 134: FAILS at every life |
|
||||
| (b) the chip's annualised hardware per MH/s is not below about a quarter of the GPU entrant's | the 5080 entrant's hardware 460 micro-USD | 239 to 677: PASSES | 33 to 95: FAILS |
|
||||
| (c) a fleet above a third of the chain's hash costs more than a year's miner revenue | at IGN 0.10 the chain is 8.8 TH/s; a third is 2.9 TH/s | USD 16 M of boards against USD 40 to 80 M of revenue: PASSES above IGN 0.05 | USD 2.3 M of dies: FAILS at every price under about 3 |
|
||||
| (d) GPUs keep a resale market and a use outside mining | every card in the population resells at 25 to 55 percent after two years and rents at 3x to 5x its mining cost | PASSES (the GPU side's property) | PASSES (the same) |
|
||||
| (e) the per-joule gap at the honest knee stays under about 3x | Blackwell at the knee 1.70 to 2.06 microjoules | 2.2x to 2.6x: PASSES | 3.7x to 4.5x: FAILS (the bare-lane floor 10x to 12x) |
|
||||
| (f) the supplier's gross margin on the share it holds is a normal return (under about 50 percent) and does not rise with each halving | | 21 to 41 percent in the five-year run: PASSES | 82 to 86 percent, or a loss once it is the chain: FAILS |
|
||||
|
||||
So: **the success statement holds for the stored-dataset DRAM-board chip at a 1 to 3 year life on today's rows, in
|
||||
every price path, at every electricity price on the axis, with the GPU side's own generation curve narrowing the gap
|
||||
further; it does not hold for the N2 SRAM die once that die exists with its development sunk, at any life, price path
|
||||
or electricity price, and the only condition that holds the die is that nobody pays to build it** (the surface of the
|
||||
floor file's section 4.4: a USD 150 M project at a third of the chain needs IGN 0.73 over three years, a maker who
|
||||
takes the chain 0.22, a revision 0.12 to 0.16). That last is a statement about an investor's decision, not about the
|
||||
chain staying below a level, and the review is right that it is the only honest form.
|
||||
|
||||
What the chain can do about the die, from the model: raise the honest side's efficiency (every cent of GPU electricity
|
||||
and every point of the knee moves condition (a) and (e); the lock already moves a Blackwell card 34 to 41 percent), keep
|
||||
the share detector as the instrument that makes condition (c) visible the week it fails, and keep the dataset floor as
|
||||
the ticket (section 3.2 of the floor file: USD 1,500 to 3,000 per die at the schedule, which moves condition (c) by 2x
|
||||
to 4x and nothing else). Nothing in the hash moves conditions (a), (b) or (e) for the die by the factor they need.
|
||||
|
||||
## 10. Unverified and owed
|
||||
|
||||
- Every chip-side figure is modelled; no chip has been measured. The chip's hardware per MH/s (USD 0.8 for the die
|
||||
with the system, 5.6 for the board) is the term conditions (b) and (c) rest on and is approximate within 2x.
|
||||
- The card prices are street approximations of October 2026; the 5090's street price has been 2x MSRP this year, which
|
||||
moves the entrant rows for that card by up to 2x and no other row.
|
||||
- The Ada and Ampere knees are modelled (no rented host allows the lock); the Blackwell floors are measured on the 5090,
|
||||
5080 and 4070, modelled on the 5070 Ti, 5070 and 5060 class.
|
||||
- The accepted-work factor (97 percent), the wear allowance (5 percent a year), the hosting (USD 0.02 per kWh) and the
|
||||
rental yields are approximate.
|
||||
- The five-year run's GPU reaction is a single-rule model (entry at the cheapest entrant's cost, exit in cost order);
|
||||
a second cut adds a per-class supply curve, the installed base as a cap on entry, re-entry on a price rise, and the
|
||||
halving's effect on the entrant cost through used-card prices.
|
||||
- The derivative-design rows (a revision at 0.3 x `C_dev`) are carried from the floor file's surface and not re-run
|
||||
here; the sunk case bounds them.
|
||||
- The emission beyond year 5 and the proving pool's 20 percent are outside the run.
|
||||
- Nothing was run on the Mac; the script ran on build-3.
|
||||
558
docs/analysis/class-v6/multi-family-adversary.md
Normal file
558
docs/analysis/class-v6/multi-family-adversary.md
Normal file
|
|
@ -0,0 +1,558 @@
|
|||
# The multi-family adversary: one programmable chip against the 180-day family bank (8 October 2026)
|
||||
|
||||
Branch `class-v6-adversary` from the box mirror's master (c86a7e23), the adversary lane of the second external review
|
||||
(the founder's 15:5x BST acceptance). The question the review put: does the 180-day family calendar add cost to a chip,
|
||||
or only change its firmware? The answer is built as the adversary would build it: ONE programmable design that executes
|
||||
every entry of the bank (`docs/analysis/rotation/layer3-family-bank.md`: 18 op families, 3 dataset atoms, the fold form
|
||||
with drawn constants, 3 block shapes, the 64-register window, the per-era parameter draws), with the designer free to
|
||||
choose lanes, clock, pipeline, banking, time-multiplexing, memory organisation and the memory arrangement itself, and
|
||||
priced on a complete machine. Every chip figure is synthesis and placement of RTL written for this lane (Yosys 0.68 and
|
||||
OpenROAD, the ORFS image, ASAP7, run on rented CPU hosts, never on the Mac, never on a Hetzner box); every GPU figure is
|
||||
the record's measurement (`docs/analysis/counter-asic-4-research.md` 15.1a, the floor programme's tier table in the
|
||||
design document 10.4). Labels as the record uses them: measured, modelled, claimed, approximate. The RTL, testbenches,
|
||||
flow and collector are under `tools/chip-model/mf/`; the k lane's rows (`floor/shadow-k.md`) are cited, never re-run.
|
||||
|
||||
The standing rules this file is written under (the review, 15:4x BST): the adversary's design is free; a synthesised
|
||||
k is a model, never a lower bound; a family transition is credited with an obsolescence benefit only where a loss of
|
||||
competitiveness is demonstrated; the stress life is 3 years, shown at 0.5, 1, 2 and 3; the cohort is the discrete-GPU
|
||||
population, Apple reported and not headlined; every defence is scored after the adversary re-optimises.
|
||||
|
||||
## 0. One page
|
||||
|
||||
(filled from the rows below at each cut; section 8 carries the three-part statement)
|
||||
|
||||
## 1. The design: what the adversary builds
|
||||
|
||||
### 1.1 What it must execute
|
||||
|
||||
| Bank entry | What the core does with it | Silicon or firmware |
|
||||
|---|---|---|
|
||||
| G1 to G7: add, sub, xor, or, rotl, rotr, and the lossy or | the ALU group, one unit per lane, operands isolated | silicon once; the mix weights are the program |
|
||||
| G8, G9, G6: mul, mulhi, mad | one 32 x 32 multiplier per lane; mulhi is the high word of the same product; mad adds the third read | silicon once |
|
||||
| G10 shfl (the 32-lane xor-mask shuffle) | a log2(LANES)-stage butterfly across the core's lanes, the mask from the instruction | silicon once per core |
|
||||
| R1 shfla (lane + delta) | a general LANES:1 crossbar per lane, the delta from a register | silicon once per core (the one network a butterfly cannot emulate in one op) |
|
||||
| R2 perm (prmt) | the 4-of-8 byte selector | silicon once |
|
||||
| R3 popc and clz | a popcount tree and a priority encoder | silicon once |
|
||||
| R4 bfe, R5 shl and shr, R6 sel, R7 andn | a shifter, a mask, a select, an and-not | silicon once (fractions of an adder each) |
|
||||
| R8 mm8 (the int8 tile) | a u8 dot4 accumulate per lane (4 MACs per op; the card's m8n8k16 tile is 32 MACs per lane, so 8 chip ops per card tile) | silicon once |
|
||||
| lop3 | the 8-bit truth table | silicon once |
|
||||
| the load and the fold form | the lane's own multiplier computes `x * M`, then the rotate and the three masks with the era's constants in registers; the returned word writes the destination | silicon once; M, R, WM, OFF, MASK are registers written at the era |
|
||||
| W = 4 wide reads | the fold step `x = rotl(x, r) * M ^ w` as an instruction (fwd), three per load | firmware |
|
||||
| the 64-register window | 64 x 32 bits per lane in an SRAM macro | silicon once (the macro) |
|
||||
| the 3 block shapes (64, 128, 256) | the program-length register; the imem holds 256 | firmware |
|
||||
| the drawn select tree | a 32-entry op permutation ahead of decode, written at the era | firmware (a 160-bit register) |
|
||||
| the op-mix band, the fold constants | the program and five registers | firmware |
|
||||
| the 3 dataset atoms (mixer x4, x8, dr368) | the per-window dataset build runs on the same core as a program (the atoms are straight-line ARX and multiply code over 16 registers); the per-hash path never executes an atom | firmware; the build's cost is section 7 |
|
||||
|
||||
### 1.2 The microarchitecture (the designer's choices)
|
||||
|
||||
- **The register state is an SRAM macro, not flops.** One FakeRAM2.0 `fakeram7_64x256` (64 words x 256 bits, single
|
||||
port) per 8 lanes: the lanes are SIMD, every lane reads the same register index, so one 256-bit access serves eight
|
||||
lanes' 32-bit reads. The macro's LEF and Liberty come with the ORFS ASAP7 platform (ABKGroup FakeRAM2.0, 7 nm,
|
||||
0.70 V; area 1,517 um^2 for the 64 x 256, 365 um^2 for the 256 x 34); the placement, the wiring and the clock tree
|
||||
see the macro as a real block. Its dynamic energy is NOT taken from the FakeRAM Liberty (a placeholder, 1.345 per
|
||||
clock edge identical for every size, "VALUES NOT REALISTIC" in the generator's own config); section 2.3 replaces it
|
||||
with a published macro energy band and the gate-level simulation supplies the exact access counts.
|
||||
- **The instruction memory is two `fakeram7_256x34` macros** shared by every lane of the core (40 bits used of 68).
|
||||
- **A single-port macro is time-multiplexed over a 5-phase slot:** read dst, read src, read src2 only when the op
|
||||
needs it (mad, lop3, sel, mm8, shfla), execute, write back. A core retires LANES lane-ops per 5 cycles; throughput is
|
||||
bought with lane count (die area), the cheapest resource the chip has, not with ports.
|
||||
- **Every unit's operands are isolated** (AND-gated by its own select), so an unused unit does not toggle: adding a
|
||||
family's unit costs leakage and a wider result mux, not switching on every op. The k lane's core evaluates every unit
|
||||
every cycle; this is one of the reasons its figure is not a lower bound.
|
||||
- **The era's draws are registers** (fold constants, program length, the select tree); a family epoch writes them.
|
||||
- Clock 1,500 ps (667 MHz) at the TC corner, as the k lane; nothing pipelined beyond the slot; two builds: 8 lanes
|
||||
(placed and routed) and 32 lanes (synthesised), each in two variants: `full` (every bank entry) and `base` (the 10
|
||||
genesis families with the load and the fold, the same microarchitecture): the difference between the two is what the
|
||||
bank adds to the chip.
|
||||
|
||||
### 1.3 What the card pays for the same instruction (the GPU side, measured)
|
||||
|
||||
The 5090's pJ per counted op at stock and at the 1,300 MHz lock (15.1a): add-class 11.3 / 6.2, mul and mad 13.9 / 8.3,
|
||||
mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, the u8 tile 4.1 / 2.2 per MAC. The reserve
|
||||
families by the measured NVIDIA step-cost ratio to the add step (design 4.2): shfla 1.53, popc 1.50, clz 1.63, bfe
|
||||
1.54, shl and shr 0.75, sel and andn about 1.0 (approximate), mm8 by the tile row (32 MACs per lane per instruction).
|
||||
|
||||
## 2. Method
|
||||
|
||||
### 2.1 The flow
|
||||
|
||||
ORFS on ASAP7 (7.5-track RVT, TC corner 0.70 V, NLDM), the default flow: Yosys with ABC, floorplan at 40 percent
|
||||
utilisation (30 for the 32-lane core) with the macros placed by the flow's macro placer under the BLOCKS power grid,
|
||||
global and detailed placement, CTS, global and detailed routing, OpenRCX parasitics. Power is OpenSTA `report_power`
|
||||
under the VCD of a random-input gate-level simulation of the netlist (iverilog; every instruction field drawn by
|
||||
`$random`, the window initialised with random words, a random returned word on every load), with a propagated 0.5
|
||||
activity as the cross-check. Synthesis-only rows (no wires, no clock tree) are marked; placed rows carry the SPEF. The routed runs clock at
|
||||
12,000 ps with the ABC target held at 1,500 ps: the unpipelined execute path (the multiplier, the fold and the
|
||||
lane reduction) is 10.6 ns at ASAP7 TC, and at 1,500 ps the flow's timing repair spent its time on a path the
|
||||
adversary would pipeline instead (two to three registers per lane, about 0.1 to 0.2 pJ per lane-op, inside the
|
||||
band). Energy per op does not depend on the period; the leakage term does, and the collector restates it at the
|
||||
1,500 ps equivalent for the routed rows (both are in `table.csv`).
|
||||
|
||||
### 2.2 The activity and the steady state
|
||||
|
||||
Each row is a tag: the op field fixed per family (`+fam=K`), or a drawn program: the class v4 draw over the 10 genesis
|
||||
families (`mix`), the same with one load in 16 (`mixld`), the draw with two reserve families live at 4 points each
|
||||
(`mix1`: shfla and mm8, the two dearest), every reserve family live at 4 points (`mix2`, the bank's bound, not a legal
|
||||
draw), the W = 4 form (`mixw4`), the 64-instruction shape (`mix64`). Two run lengths per tag (500 and 2,000 clocks, 100
|
||||
and 400 slots) bracket the 326-clock load phase (reset, the era's registers, the 256-word program, the 64-word window
|
||||
init) and the run-phase power is solved from the pair, as the k lane does.
|
||||
|
||||
### 2.3 The SRAM macro energy (the one modelled term on the chip side)
|
||||
|
||||
The FakeRAM Liberty's internal power is a placeholder, so the collector removes the macros' Liberty-attributed power
|
||||
(reported separately per VCD with `report_power -instances`) and adds a modelled access energy times the exact
|
||||
access count from the simulation (3 reads and 1 write per slot for a three-operand op, 2 reads and 1 write otherwise,
|
||||
per 8 lanes; 2 imem reads per slot per core). The band: a 64 x 256 single-port macro at a 7 nm class node 3.5 pJ per
|
||||
256-bit access (2.0 to 7.0); a 256 x 34 macro 1.5 pJ per access (0.8 to 3.0). Sources: Horowitz, ISSCC 2014 (45 nm:
|
||||
an 8 KB SRAM read of 64 bits 10 pJ, 32 KB 20 pJ), scaled by the bits moved and by the energy-per-bit reduction from
|
||||
45 nm to a 7 nm class node (about 0.1x to 0.2x, approximate, the same generation scaling the record applies to logic);
|
||||
the k lane's "2 to 4 pJ per 32-bit read, approximate" for a 4 KB imem; CACTI-class estimates for a 16 Kbit macro at
|
||||
7 nm (0.01 to 0.03 pJ per bit read, approximate). Every row carries the band; the low end is near a flop array with
|
||||
perfect clock gating, the high end a conservative compiler macro. The macros' switching on their output nets (256
|
||||
bits into the lanes' latches) stays in the logic figure, measured.
|
||||
|
||||
### 2.4 Node scaling (claimed) and what the method leaves out
|
||||
|
||||
ASAP7 is a predictive 7 nm-class PDK; the row is stated at ASAP7 and scaled by the foundry's headline per-node
|
||||
power reductions at the same speed (the k lane's factors, every one claimed): N5 = 0.70, N3 = 0.50, N2 = 0.36 of
|
||||
ASAP7. Node-for-node against the 5090 (TSMC 4N, N5 class) is the N5 column; a node ahead is N3. Left out on the chip
|
||||
side: the memory controller's queueing logic and the lane's address output (priced in the board model as the
|
||||
controller die), test and clock distribution beyond the block; on the card side the 15.1a figure is the whole card's
|
||||
marginal per counted op, which includes fetch, decode, operand collection and the register file, so the comparison
|
||||
is the chip's whole lane (fetch, decode, window, units, network) against the card's whole lane.
|
||||
|
||||
## 3. The rows: pJ per lane-op per family on the base core (synthesis only, 8 lanes)
|
||||
|
||||
Synthesis only (no wires, no clock tree), 8 lanes, 1,500 ps, ASAP7 TC; "pJ logic" is OpenSTA's figure for the
|
||||
standard cells under the VCD with the FakeRAM placeholder removed (its sequential and combinational parts beside it);
|
||||
"pJ SRAM" the modelled macro term at the simulated access count (low / nominal / high, section 2.3); the per-op
|
||||
figure is logic plus the nominal SRAM term, the band in brackets; the 5090 column is 15.1a (the reserve families
|
||||
by the measured step ratio; the mix rows against the card's 10.3 pJ per op on the class v4 draw at the lock, 18.8
|
||||
unlocked by the ARX ratio, approximate). The full core: 95,678 cells and 3 macros (the base core 66,973 and 3
|
||||
macros: the bank adds 43 percent of the standard cells, 0.13 mW of leakage per 8 lanes, and 11 percent to the
|
||||
energy of the class v4 draw on the same microarchitecture, the wider result mux and the leakage of the idle units).
|
||||
|
||||
The full core (every bank entry):
|
||||
|
||||
| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| mf8full | power_synth | mix | 95678 | 0.00514 | 0.000395 | 4.82 (0.88 / 3.6) | 0.99 / 1.76 / 3.52 | 6.58 (5.81 to 8.34) | 4.61 | 3.32 | 2.39 | 18.8 / 10.3 | 0.447 | 0.322 | 0.176 |
|
||||
| mf8full | power_synth | mixld | 95678 | 0.00513 | 0.000395 | 4.81 (0.88 / 3.6) | 0.986 / 1.75 / 3.5 | 6.56 (5.79 to 8.31) | 4.59 | 3.31 | 2.38 | 18.8 / 10.3 | 0.446 | 0.321 | 0.176 |
|
||||
| mf8full | power_synth | mix1 | 95678 | 0.00495 | 0.000395 | 4.64 (0.89 / 3.4) | 1.01 / 1.8 / 3.59 | 6.43 (5.65 to 8.23) | 4.5 | 3.24 | 2.33 | 18.8 / 10.3 | 0.437 | 0.315 | 0.172 |
|
||||
| mf8full | power_synth | mix2 | 95678 | 0.00436 | 0.000396 | 4.09 (0.88 / 2.8) | 1 / 1.78 / 3.56 | 5.87 (5.09 to 7.65) | 4.11 | 2.96 | 2.13 | 18.8 / 10.3 | 0.399 | 0.287 | 0.157 |
|
||||
| mf8full | power_synth | mixw4 | 95678 | 0.00508 | 0.000395 | 4.76 (0.88 / 3.5) | 0.977 / 1.73 / 3.47 | 6.49 (5.74 to 8.23) | 4.55 | 3.27 | 2.36 | 18.8 / 10.3 | 0.441 | 0.318 | 0.174 |
|
||||
| mf8full | power_synth | mix64 | 95678 | 0.00488 | 0.000395 | 4.58 (0.88 / 3.3) | 0.993 / 1.76 / 3.53 | 6.34 (5.57 to 8.1) | 4.44 | 3.19 | 2.3 | 18.8 / 10.3 | 0.431 | 0.31 | 0.17 |
|
||||
| mf8full | power_synth | add | 95678 | 0.003 | 0.000396 | 2.81 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.5 (3.76 to 6.18) | 3.15 | 2.27 | 1.63 | 11.3 / 6.2 | 0.508 | 0.366 | 0.201 |
|
||||
| mf8full | power_synth | sub | 95678 | 0.00299 | 0.000396 | 2.8 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.49 (3.75 to 6.18) | 3.14 | 2.26 | 1.63 | 11.3 / 6.2 | 0.507 | 0.365 | 0.2 |
|
||||
| mf8full | power_synth | xor | 95678 | 0.00308 | 0.000396 | 2.89 (0.86 / 1.7) | 0.95 / 1.69 / 3.38 | 4.57 (3.84 to 6.26) | 3.2 | 2.3 | 1.66 | 11.3 / 6.2 | 0.516 | 0.372 | 0.204 |
|
||||
| mf8full | power_synth | or | 95678 | 0.00183 | 0.000397 | 1.72 (0.84 / 0.5) | 0.95 / 1.69 / 3.38 | 3.41 (2.67 to 5.09) | 2.38 | 1.72 | 1.24 | 11.3 / 6.2 | 0.385 | 0.277 | 0.152 |
|
||||
| mf8full | power_synth | rotl | 95678 | 0.00318 | 0.000396 | 2.99 (0.86 / 1.8) | 0.95 / 1.69 / 3.38 | 4.67 (3.94 to 6.36) | 3.27 | 2.36 | 1.7 | 11.3 / 6.2 | 0.528 | 0.38 | 0.208 |
|
||||
| mf8full | power_synth | rotr | 95678 | 0.003 | 0.000396 | 2.81 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.5 (3.76 to 6.18) | 3.15 | 2.27 | 1.63 | 11.3 / 6.2 | 0.508 | 0.366 | 0.201 |
|
||||
| mf8full | power_synth | mul | 95678 | 0.00265 | 0.000396 | 2.48 (0.85 / 1.3) | 0.95 / 1.69 / 3.38 | 4.17 (3.43 to 5.86) | 2.92 | 2.1 | 1.51 | 13.9 / 8.3 | 0.352 | 0.253 | 0.151 |
|
||||
| mf8full | power_synth | mulhi | 95678 | 0.00249 | 0.000396 | 2.33 (0.85 / 1.1) | 0.95 / 1.69 / 3.38 | 4.02 (3.28 to 5.71) | 2.81 | 2.03 | 1.46 | 39.6 / 21 | 0.134 | 0.0965 | 0.0512 |
|
||||
| mf8full | power_synth | mad | 95678 | 0.00562 | 0.000395 | 5.26 (0.91 / 4) | 1.2 / 2.12 / 4.25 | 7.39 (6.46 to 9.51) | 5.17 | 3.72 | 2.68 | 13.9 / 8.3 | 0.623 | 0.449 | 0.268 |
|
||||
| mf8full | power_synth | shfl | 95678 | 0.00266 | 0.000397 | 2.49 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.18 (3.44 to 5.87) | 2.93 | 2.11 | 1.52 | 55.8 / 29.4 | 0.0995 | 0.0716 | 0.0377 |
|
||||
| mf8full | power_synth | load | 95678 | 0.00433 | 0.000395 | 4.06 (0.86 / 2.8) | 0.95 / 1.69 / 3.38 | 5.75 (5.01 to 7.44) | 4.03 | 2.9 | 2.09 | 13.9 / 8.3 | 0.485 | 0.349 | 0.209 |
|
||||
| mf8full | power_synth | fwd | 95678 | 0.004 | 0.000395 | 3.75 (0.86 / 2.5) | 0.95 / 1.69 / 3.38 | 5.44 (4.7 to 7.13) | 3.81 | 2.74 | 1.97 | 13.9 / 8.3 | 0.459 | 0.33 | 0.197 |
|
||||
| mf8full | power_synth | prmt | 95678 | 0.00255 | 0.000397 | 2.39 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4.07 (3.34 to 5.76) | 2.85 | 2.05 | 1.48 | 22.3 / 11.5 | 0.248 | 0.179 | 0.0921 |
|
||||
| mf8full | power_synth | lop3 | 95678 | 0.003 | 0.000397 | 2.81 (0.91 / 1.5) | 1.2 / 2.12 / 4.25 | 4.94 (4.01 to 7.06) | 3.46 | 2.49 | 1.79 | 24.1 / 13 | 0.266 | 0.191 | 0.103 |
|
||||
| mf8full | power_synth | shfla | 95678 | 0.00302 | 0.000397 | 2.83 (0.9 / 1.6) | 1.2 / 2.12 / 4.25 | 4.95 (4.03 to 7.08) | 3.47 | 2.5 | 1.8 | 17.3 / 9.49 | 0.365 | 0.263 | 0.144 |
|
||||
| mf8full | power_synth | popc | 95678 | 0.0017 | 0.000397 | 1.59 (0.85 / 0.38) | 0.95 / 1.69 / 3.38 | 3.28 (2.54 to 4.97) | 2.3 | 1.65 | 1.19 | 17 / 9.3 | 0.247 | 0.178 | 0.0976 |
|
||||
| mf8full | power_synth | clz | 95678 | 0.00173 | 0.000397 | 1.62 (0.85 / 0.41) | 0.95 / 1.69 / 3.38 | 3.31 (2.57 to 5) | 2.32 | 1.67 | 1.2 | 18.4 / 10.1 | 0.229 | 0.165 | 0.0906 |
|
||||
| mf8full | power_synth | bfe | 95678 | 0.00182 | 0.000397 | 1.71 (0.85 / 0.49) | 0.95 / 1.69 / 3.38 | 3.39 (2.66 to 5.08) | 2.38 | 1.71 | 1.23 | 17.4 / 9.55 | 0.249 | 0.179 | 0.0983 |
|
||||
| mf8full | power_synth | shl | 95678 | 0.00234 | 0.000397 | 2.19 (0.86 / 0.98) | 0.95 / 1.69 / 3.38 | 3.88 (3.14 to 5.57) | 2.71 | 1.95 | 1.41 | 8.48 / 4.65 | 0.584 | 0.42 | 0.231 |
|
||||
| mf8full | power_synth | shr | 95678 | 0.00221 | 0.000397 | 2.07 (0.85 / 0.84) | 0.95 / 1.69 / 3.38 | 3.76 (3.02 to 5.45) | 2.63 | 1.89 | 1.36 | 8.48 / 4.65 | 0.566 | 0.407 | 0.224 |
|
||||
| mf8full | power_synth | sel | 95678 | 0.0028 | 0.000397 | 2.63 (0.91 / 1.4) | 1.2 / 2.12 / 4.25 | 4.75 (3.83 to 6.88) | 3.33 | 2.4 | 1.72 | 11.3 / 6.2 | 0.537 | 0.386 | 0.212 |
|
||||
| mf8full | power_synth | andn | 95678 | 0.00234 | 0.000397 | 2.19 (0.86 / 0.98) | 0.95 / 1.69 / 3.38 | 3.88 (3.14 to 5.57) | 2.72 | 1.96 | 1.41 | 11.3 / 6.2 | 0.438 | 0.315 | 0.173 |
|
||||
| mf8full | power_synth | mm8 | 95678 | 0.0036 | 0.000396 | 3.37 (0.91 / 2.1) | 1.2 / 2.12 / 4.25 | 5.5 (4.57 to 7.62) | 3.85 | 2.77 | 2 | 16.4 / 8.8 | 0.437 | 0.315 | 0.169 |
|
||||
|
||||
The base core (the 10 genesis families on the same microarchitecture; the comparator):
|
||||
|
||||
| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| mf8base | power_synth | mix | 66973 | 0.00443 | 0.000263 | 4.15 (0.88 / 3) | 0.99 / 1.76 / 3.52 | 5.91 (5.14 to 7.66) | 4.13 | 2.98 | 2.14 | 18.8 / 10.3 | 0.401 | 0.289 | 0.158 |
|
||||
| mf8base | power_synth | mixld | 66973 | 0.00439 | 0.000263 | 4.12 (0.88 / 3) | 0.986 / 1.75 / 3.5 | 5.87 (5.1 to 7.62) | 4.11 | 2.96 | 2.13 | 18.8 / 10.3 | 0.399 | 0.287 | 0.157 |
|
||||
| mf8base | power_synth | add | 66973 | 0.00246 | 0.000264 | 2.31 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4 (3.26 to 5.68) | 2.8 | 2.01 | 1.45 | 11.3 / 6.2 | 0.451 | 0.325 | 0.178 |
|
||||
| mf8base | power_synth | sub | 66973 | 0.00245 | 0.000264 | 2.29 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 3.98 (3.24 to 5.67) | 2.79 | 2.01 | 1.44 | 11.3 / 6.2 | 0.449 | 0.324 | 0.178 |
|
||||
| mf8base | power_synth | xor | 66973 | 0.00247 | 0.000264 | 2.32 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4 (3.27 to 5.69) | 2.8 | 2.02 | 1.45 | 11.3 / 6.2 | 0.452 | 0.326 | 0.179 |
|
||||
| mf8base | power_synth | or | 66973 | 0.0016 | 0.000265 | 1.5 (0.84 / 0.42) | 0.95 / 1.69 / 3.38 | 3.19 (2.45 to 4.87) | 2.23 | 1.61 | 1.16 | 11.3 / 6.2 | 0.36 | 0.259 | 0.142 |
|
||||
| mf8base | power_synth | rotl | 66973 | 0.00259 | 0.000264 | 2.43 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.11 (3.38 to 5.8) | 2.88 | 2.07 | 1.49 | 11.3 / 6.2 | 0.464 | 0.334 | 0.183 |
|
||||
| mf8base | power_synth | rotr | 66973 | 0.00251 | 0.000264 | 2.35 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.04 (3.3 to 5.73) | 2.83 | 2.04 | 1.47 | 11.3 / 6.2 | 0.456 | 0.328 | 0.18 |
|
||||
| mf8base | power_synth | mul | 66973 | 0.00229 | 0.000264 | 2.14 (0.85 / 1.1) | 0.95 / 1.69 / 3.38 | 3.83 (3.09 to 5.52) | 2.68 | 1.93 | 1.39 | 13.9 / 8.3 | 0.323 | 0.233 | 0.139 |
|
||||
| mf8base | power_synth | mulhi | 66973 | 0.00217 | 0.000264 | 2.03 (0.84 / 0.96) | 0.95 / 1.69 / 3.38 | 3.72 (2.98 to 5.41) | 2.61 | 1.88 | 1.35 | 39.6 / 21 | 0.124 | 0.0893 | 0.0474 |
|
||||
| mf8base | power_synth | mad | 66973 | 0.00472 | 0.000263 | 4.43 (0.9 / 3.3) | 1.2 / 2.12 / 4.25 | 6.55 (5.62 to 8.67) | 4.58 | 3.3 | 2.38 | 13.9 / 8.3 | 0.552 | 0.398 | 0.237 |
|
||||
| mf8base | power_synth | shfl | 66973 | 0.00192 | 0.000265 | 1.8 (0.86 / 0.7) | 0.95 / 1.69 / 3.38 | 3.49 (2.75 to 5.17) | 2.44 | 1.76 | 1.27 | 55.8 / 29.4 | 0.083 | 0.0598 | 0.0315 |
|
||||
| mf8base | power_synth | load | 66973 | 0.00359 | 0.000263 | 3.36 (0.86 / 2.3) | 0.95 / 1.69 / 3.38 | 5.05 (4.31 to 6.74) | 3.53 | 2.54 | 1.83 | 13.9 / 8.3 | 0.426 | 0.307 | 0.183 |
|
||||
|
||||
Reading the rows. (1) The ALU-group families cost the chip 3.1 to 3.3 pJ per lane-op at N5 against the card's 6.2
|
||||
at the lock (k 0.51 to 0.53); or, popc, clz, bfe, prmt 2.3 to 2.9 (k 0.23 to 0.38); the multiply 2.9 (k 0.35);
|
||||
mulhi 2.8 against the card's 21 (k 0.13); the shuffle 2.9 against 29.4 (k 0.10); mad is the dearest op for the
|
||||
chip at 5.2 (three reads and two units, k 0.62) and the shifters the highest k (0.57 to 0.58) because the card
|
||||
does them cheapest. (2) The SRAM term is 1.7 pJ nominal per lane-op (0.95 to 3.4): a third of the row; the
|
||||
sequential term (the phase latches, the instruction register, the counters at five edges per op) 0.86 pJ; the rest
|
||||
is the units and the macro output nets. (3) The reserve families are cheaper for the chip than the genesis
|
||||
families: every live-family mix sits under the class v4 draw, and the bound with all eight live reads 11 percent
|
||||
under it, because the card's dearest instructions (mulhi, shfl, mad) are the genesis ones.
|
||||
|
||||
|
||||
## 4. The whole-hash energy and k per family, node-for-node and a node ahead
|
||||
|
||||
The whole-hash shadow is 102,612 chip ops (the 102,100 counted shadow ops of the record plus the 512-instruction
|
||||
base program; W = 4 adds 384 fold steps). The chip's cost is absolute (its own pJ at its node); the card's premium
|
||||
on the same draw is the measured 0.652 microjoules at the lock, moved by the live family's measured step ratio at 4
|
||||
points of 79 (modelled). Node-for-node is N5 (the 5090's own class), a node ahead N3; every factor claimed.
|
||||
|
||||
design mf8full stage power_synth; the class v4 draw on the core: 6.58 pJ per lane-op ASAP7, 4.61 N5, 3.32 N3 (band 4.07 to 5.84 at N5)
|
||||
|
||||
| Family live (4 points of 79) or draw | Chip pJ per op N5 / N3 (band) | 5090 pJ per op at the lock | k N5 / N3 | Chip shadow per hash, microjoules N5 / N3 (the mix with the family live) | Change vs the class v4 draw | The card's premium per hash at the lock (modelled from the step ratio) | Change |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| the class v4 draw (10 genesis families) | 4.61 / 3.32 (4.07 to 5.84) | 10.3 | 0.45 / 0.32 | 0.473 / 0.340 | +0.0 percent | 0.652 | +0.0 percent |
|
||||
| add (unit row, every instruction) | 3.15 / 2.27 (2.63 to 4.33) | 6.2 | 0.51 / 0.37 | 0.323 / 0.233 | -31.7 percent | 0.392 | -39.8 percent |
|
||||
| sub (unit row, every instruction) | 3.14 / 2.26 (2.63 to 4.32) | 6.2 | 0.51 / 0.36 | 0.322 / 0.232 | -31.8 percent | 0.392 | -39.8 percent |
|
||||
| xor (unit row, every instruction) | 3.20 / 2.30 (2.68 to 4.38) | 6.2 | 0.52 / 0.37 | 0.328 / 0.236 | -30.5 percent | 0.392 | -39.8 percent |
|
||||
| or (unit row, every instruction) | 2.38 / 1.72 (1.87 to 3.57) | 6.2 | 0.38 / 0.28 | 0.245 / 0.176 | -48.2 percent | 0.392 | -39.8 percent |
|
||||
| rotl (unit row, every instruction) | 3.27 / 2.36 (2.75 to 4.45) | 6.2 | 0.53 / 0.38 | 0.336 / 0.242 | -29.0 percent | 0.392 | -39.8 percent |
|
||||
| rotr (unit row, every instruction) | 3.15 / 2.27 (2.63 to 4.33) | 6.2 | 0.51 / 0.37 | 0.323 / 0.233 | -31.7 percent | 0.392 | -39.8 percent |
|
||||
| mul (unit row, every instruction) | 2.92 / 2.10 (2.40 to 4.10) | 8.3 | 0.35 / 0.25 | 0.299 / 0.216 | -36.7 percent | 0.525 | -19.4 percent |
|
||||
| mulhi (unit row, every instruction) | 2.81 / 2.03 (2.30 to 3.99) | 21.0 | 0.13 / 0.10 | 0.289 / 0.208 | -38.9 percent | 1.329 | +103.9 percent |
|
||||
| mad (unit row, every instruction) | 5.17 / 3.72 (4.52 to 6.66) | 8.3 | 0.62 / 0.45 | 0.531 / 0.382 | +12.3 percent | 0.525 | -19.4 percent |
|
||||
| shfl (unit row, every instruction) | 2.93 / 2.11 (2.41 to 4.11) | 29.4 | 0.10 / 0.07 | 0.300 / 0.216 | -36.5 percent | 1.861 | +185.4 percent |
|
||||
| load (unit row, every instruction) | 4.03 / 2.90 (3.51 to 5.21) | 8.3 | 0.48 / 0.35 | 0.413 / 0.297 | -12.6 percent | 0.525 | -19.4 percent |
|
||||
| prmt live | 2.85 / 2.05 (2.34 to 4.03) | 11.5 | 0.25 / 0.18 | 0.464 / 0.334 | -1.9 percent | 0.662 | +1.5 percent |
|
||||
| lop3 live | 3.46 / 2.49 (2.81 to 4.94) | 13.0 | 0.27 / 0.19 | 0.467 / 0.336 | -1.3 percent | 0.688 | +5.6 percent |
|
||||
| shfla live | 3.47 / 2.50 (2.82 to 4.95) | 9.5 | 0.37 / 0.26 | 0.467 / 0.336 | -1.3 percent | 0.669 | +2.7 percent |
|
||||
| popc live | 2.30 / 1.65 (1.78 to 3.48) | 9.3 | 0.25 / 0.18 | 0.461 / 0.332 | -2.5 percent | 0.669 | +2.5 percent |
|
||||
| clz live | 2.32 / 1.67 (1.80 to 3.50) | 10.1 | 0.23 / 0.17 | 0.461 / 0.332 | -2.5 percent | 0.673 | +3.2 percent |
|
||||
| bfe live | 2.38 / 1.71 (1.86 to 3.56) | 9.5 | 0.25 / 0.18 | 0.461 / 0.332 | -2.5 percent | 0.670 | +2.7 percent |
|
||||
| shl live | 2.71 / 1.95 (2.20 to 3.90) | 4.7 | 0.58 / 0.42 | 0.463 / 0.333 | -2.1 percent | 0.644 | -1.3 percent |
|
||||
| shr live | 2.63 / 1.89 (2.12 to 3.81) | 4.7 | 0.57 / 0.41 | 0.462 / 0.333 | -2.2 percent | 0.644 | -1.3 percent |
|
||||
| sel live | 3.33 / 2.40 (2.68 to 4.81) | 6.2 | 0.54 / 0.39 | 0.466 / 0.336 | -1.4 percent | 0.652 | +0.0 percent |
|
||||
| andn live | 2.72 / 1.96 (2.20 to 3.90) | 6.2 | 0.44 / 0.32 | 0.463 / 0.333 | -2.1 percent | 0.652 | +0.0 percent |
|
||||
| mm8 live | 3.85 / 2.77 (3.20 to 5.34) | 8.8 | 0.44 / 0.31 | 0.469 / 0.337 | -0.8 percent | 0.699 | +7.2 percent |
|
||||
| fwd (unit row, every instruction) | 3.81 / 2.74 (3.29 to 4.99) | 8.3 | 0.46 / 0.33 | 0.391 / 0.281 | -17.3 percent | 0.525 | -19.4 percent |
|
||||
| the draw with shfla and mm8 live (measured mix) | 4.50 / 3.24 (3.96 to 5.76) | 10.3 | 0.44 / 0.31 | 0.462 / 0.333 | -2.2 percent | 0.652 | +0.0 percent |
|
||||
| every reserve family live at 4 points (the bound, measured mix) | 4.11 / 2.96 (3.56 to 5.35) | 10.3 | 0.40 / 0.29 | 0.422 / 0.304 | -10.8 percent | 0.652 | +0.0 percent |
|
||||
| the draw with one load in 16 | 4.59 / 3.31 (4.06 to 5.82) | 10.3 | 0.45 / 0.32 | 0.471 / 0.339 | -0.3 percent | 0.652 | +0.0 percent |
|
||||
| W = 4: three fold steps per load | 4.55 / 3.27 (4.02 to 5.76) | 10.3 | 0.44 / 0.32 | 0.466 / 0.336 | -1.3 percent | 0.652 | +0.0 percent |
|
||||
| the 64-instruction shape | 4.44 / 3.19 (3.90 to 5.67) | 10.3 | 0.43 / 0.31 | 0.455 / 0.328 | -3.7 percent | 0.652 | +0.0 percent |
|
||||
|
||||
The whole-hash edge per joule with the class v4 shadow on the core (absolute: the chip pays its own pJ whatever the card does):
|
||||
| Memory (E_mem, microjoules) | Chip E_hash N5 / N3 | vs 5090 lock 2.33 | vs 5090 stock 3.36 | vs 5080 lock 2.06 | vs M5 Max 1.40 |
|
||||
|---|---|---|---|---|---|
|
||||
| GDDR7 board (0.466) | 0.939 / 0.806 | 2.5x / 2.9x | 3.6x / 4.2x | 2.2x / 2.6x | 1.5x / 1.7x |
|
||||
| HBM3 one stack (0.321) | 0.794 / 0.661 | 2.9x / 3.5x | 4.2x / 5.1x | 2.6x / 3.1x | 1.8x / 2.1x |
|
||||
| SRAM N2 die at W = 1 (0.036) | 0.509 / 0.376 | 4.6x / 6.2x | 6.6x / 8.9x | 4.1x / 5.5x | 2.8x / 3.7x |
|
||||
|
||||
The same edge on the base core (the genesis-only comparator): 2.6x / 3.0x microjoules per hash
|
||||
on the GDDR7 board at N5 / N3, 3.8x / 4.4x against the 5090 at its lock; the bank costs the
|
||||
adversary 5 percent of its edge (0.939 against 0.890 microjoules on the GDDR7 board, 2.5x against 2.6x), which is
|
||||
the whole answer to the review's question in one number: the 180-day calendar costs a chip that carries the bank
|
||||
about 11 percent of its shadow energy and 43 percent of its core cells, and no epoch costs it a part.
|
||||
|
||||
|
||||
## 5. The placed rows (8 lanes, routed, SPEF) and the 32-lane core
|
||||
|
||||
ROWS_PLACED
|
||||
|
||||
## 6. The board: joules per valid hash and USD per sustained MH/s for the complete machine
|
||||
|
||||
The complete machine: the memory devices at their modelled random-read energy and activate ceiling (chip-model-v3
|
||||
5.3, unmeasured; the 5090 reaches 82 percent of the GDDR7 figure, the sustained fraction here), the controller and
|
||||
PHY die, the core die sized to retire the hash's ops at the memory's sustained rate (lanes at 667 MHz over 5 phases;
|
||||
0.002 mm^2 per lane at N5 from the 8-lane core's floorplan at 40 percent utilisation scaled x0.55, approximate; USD
|
||||
0.36 per mm^2 of N5 from sram-mirror's yield model; 50 uW of leakage per lane, the synthesis figure), power delivery
|
||||
(PSU 92 percent, VRM 90 percent), cooling (3 percent), a board and assembly (USD 200), and one full node per 100
|
||||
machines (85 W and USD 1,500 shared: the state-derived dataset's host, priced as the review's rule 3 asks). The
|
||||
"high" case takes the memory's lower ceiling (the 5090's measured 17.5 G on GDDR7; the JEDEC tFAW floor on HBM3),
|
||||
the high read energy and the high SRAM term together. GPU rows: the 5090 at its lock 2.33 microjoules at 134.76
|
||||
MH/s, USD 1,999 plus USD 150 of rig share (USD 15.9 per MH/s; at the USD 3,000 street price 23.4); the 5080 at its
|
||||
lock 2.06 at 71.20 MH/s, USD 999 plus 150 (USD 16.1 per MH/s); the discrete-GPU cohort by count (design 3.4 and
|
||||
10.4: 8 GB 22 percent, 12 GB 22, 16 GB 28, 24 GB and up 16, the rest 10 and 11 GB) at its tuned points about 3.6
|
||||
microjoules (2.5 to 4.5, approximate: the 5070 and 5070 Ti 1.7 to 1.75 modelled, the 4070 3.58 measured, the 4090
|
||||
3.64, the 3090 6.4, the 9070 XT 8.1 measured) and about USD 18 per MH/s (14 to 28); the M5 Max 1.40 at 27.9 MH/s is
|
||||
reported in section 4, not headlined.
|
||||
|
||||
Node-for-node (N5), the full core at 4.61 pJ per lane-op:
|
||||
|
||||
core 4.61 pJ per lane-op (band 4.07 to 5.84), 102612 ops per hash, leakage 50 uW per lane, 0.002 mm^2 per lane at N5
|
||||
|
||||
| Memory arrangement | Case | Sustained MH/s | Machine W | microjoules per hash (whole machine) | Core lanes | Core mm^2 (N5) | Capex USD | USD per MH/s | vs 5090 lock 2.33 (J / USD) | vs 5080 lock 2.06 | vs cohort 3.6 / USD 18 |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | nominal | 136 | 175 | 1.280 | 104,960 | 210 | 661 | 4.84 | 1.8x / 3.3x | 1.6x / 3.3x | 2.8x / 3.7x |
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | low | 136 | 154 | 1.132 | 104,960 | 210 | 661 | 4.84 | 2.1x / 3.3x | 1.8x / 3.3x | 3.2x / 3.7x |
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | high | 112 | 180 | 1.603 | 86,235 | 172 | 647 | 5.77 | 1.5x / 2.8x | 1.3x / 2.8x | 2.2x / 3.1x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | nominal | 71 | 77 | 1.084 | 54,656 | 109 | 704 | 9.91 | 2.1x / 1.6x | 1.9x / 1.6x | 3.3x / 1.8x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | low | 71 | 70 | 0.984 | 54,656 | 109 | 704 | 9.91 | 2.4x / 1.6x | 2.1x / 1.6x | 3.7x / 1.8x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | high | 15 | 34 | 2.228 | 11,748 | 23 | 673 | 44.09 | 1.0x / 0.4x | 0.9x / 0.4x | 1.6x / 0.4x |
|
||||
| HBM3 eight stacks (an H100-class package) | nominal | 566 | 559 | 0.987 | 435,713 | 871 | 3,079 | 5.44 | 2.4x / 2.9x | 2.1x / 3.0x | 3.6x / 3.3x |
|
||||
| HBM3 eight stacks (an H100-class package) | low | 566 | 502 | 0.886 | 435,713 | 871 | 3,079 | 5.44 | 2.6x / 2.9x | 2.3x / 3.0x | 4.1x / 3.3x |
|
||||
| HBM3 eight stacks (an H100-class package) | high | 122 | 217 | 1.772 | 93,987 | 188 | 2,833 | 23.18 | 1.3x / 0.7x | 1.2x / 0.7x | 2.0x / 0.8x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | nominal | 535 | 400 | 0.747 | 411,225 | 822 | 1,211 | 2.27 | 3.1x / 7.0x | 2.8x / 7.1x | 4.8x / 7.9x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | low | 609 | 403 | 0.662 | 468,572 | 937 | 1,252 | 2.06 | 3.5x / 7.8x | 3.1x / 7.8x | 5.4x / 8.8x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | high | 419 | 394 | 0.940 | 322,466 | 645 | 1,147 | 2.74 | 2.5x / 5.8x | 2.2x / 5.9x | 3.8x / 6.6x |
|
||||
|
||||
Lifetime cost in USD per TH (10^12 hashes), capex spread over the life plus electricity at USD 0.08 per kWh:
|
||||
| Machine | 0.5 y | 1 y | 2 y | 3 y | of which electricity |
|
||||
|---|---|---|---|---|---|
|
||||
| 5090 at the lock | 1.063 | 0.557 | 0.305 | 0.220 | 0.0518 |
|
||||
| 5080 at the lock | 1.069 | 0.557 | 0.302 | 0.216 | 0.0458 |
|
||||
| cohort card | 1.222 | 0.651 | 0.365 | 0.270 | 0.0800 |
|
||||
| GDDR7 board, 16 devices, 64 channels chip | 0.335 | 0.182 | 0.105 | 0.080 | 0.0284 |
|
||||
| HBM3 one stack chip | 0.653 | 0.338 | 0.181 | 0.129 | 0.0241 |
|
||||
| HBM3 eight stacks chip | 0.367 | 0.194 | 0.108 | 0.079 | 0.0219 |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 chip | 0.160 | 0.088 | 0.053 | 0.041 | 0.0166 |
|
||||
|
||||
A node ahead (N3), the same machine:
|
||||
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | nominal | 136 | 150 | 1.101 | 104,960 | 147 | 638 | 4.67 | 2.1x / 3.4x | 1.9x / 3.5x | 3.3x / 3.9x |
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | low | 136 | 133 | 0.972 | 104,960 | 147 | 638 | 4.67 | 2.4x / 3.4x | 2.1x / 3.5x | 3.7x / 3.9x |
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | high | 112 | 155 | 1.380 | 86,235 | 121 | 628 | 5.61 | 1.7x / 2.8x | 1.5x / 2.9x | 2.6x / 3.2x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | nominal | 71 | 64 | 0.905 | 54,656 | 77 | 693 | 9.75 | 2.6x / 1.6x | 2.3x / 1.7x | 4.0x / 1.8x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | low | 71 | 59 | 0.824 | 54,656 | 77 | 693 | 9.75 | 2.8x / 1.6x | 2.5x / 1.7x | 4.4x / 1.8x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | high | 15 | 31 | 2.004 | 11,748 | 16 | 671 | 43.93 | 1.2x / 0.4x | 1.0x / 0.4x | 1.8x / 0.4x |
|
||||
| HBM3 eight stacks (an H100-class package) | nominal | 566 | 458 | 0.808 | 435,713 | 610 | 2,985 | 5.27 | 2.9x / 3.0x | 2.5x / 3.1x | 4.5x / 3.4x |
|
||||
| HBM3 eight stacks (an H100-class package) | low | 566 | 411 | 0.726 | 435,713 | 610 | 2,985 | 5.27 | 3.2x / 3.0x | 2.8x / 3.1x | 5.0x / 3.4x |
|
||||
| HBM3 eight stacks (an H100-class package) | high | 122 | 189 | 1.548 | 93,987 | 132 | 2,812 | 23.02 | 1.5x / 0.7x | 1.3x / 0.7x | 2.3x / 0.8x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | nominal | 724 | 398 | 0.550 | 557,288 | 780 | 1,196 | 1.65 | 4.2x / 9.7x | 3.7x / 9.8x | 6.5x / 10.9x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | low | 828 | 402 | 0.485 | 636,578 | 891 | 1,236 | 1.49 | 4.8x / 10.7x | 4.2x / 10.8x | 7.4x / 12.1x |
|
||||
|
||||
Reading. (1) The complete GDDR7 machine reads 1.8x the 5090 at its lock per joule node-for-node (1.5x to 2.1x) and
|
||||
2.1x a node ahead (1.7x to 2.4x); against the 5080 at its lock 1.6x (1.3x to 1.8x) and 1.9x; against the cohort
|
||||
card 2.8x and 3.3x. Per dollar of capex it is 3.3x the 5090 at MSRP (4.8x at the street price) and 3.7x the cohort,
|
||||
because the board carries the same USD 320 of memory and USD 76 of core silicon where the card carries a 750 mm^2
|
||||
GPU. (2) The HBM3 one-stack machine is no better per joule than GDDR7 once the controller, the host and the power
|
||||
train are in, and its dollars per MH/s are twice the card's at the modelled ceiling and 2.7x worse than the card at
|
||||
the JEDEC tFAW floor: the "same DRAM" assumption is tested and the adversary's cheapest memory IS the card's memory,
|
||||
so the GDDR7 board is the machine the economics must answer. (3) The N2 SRAM die at the hash's own width is the
|
||||
strongest machine at 3.1x per joule node-for-node (2.5x to 3.5x) and 4.2x a node ahead, and 7x per dollar of
|
||||
silicon, with the project cost (USD 100 M to 500 M, claimed) as its only hold. (4) Over a 3-year life the GDDR7
|
||||
machine's cost per TH is 0.080 USD against the 5090's 0.220 and the cohort's 0.270 (2.7x to 3.4x); at a 1-year life
|
||||
3.1x; electricity is a quarter of the machine's lifetime cost and a fifth of the card's, so the per-joule edge is
|
||||
the smaller half of the economic edge and the capex per MH/s the larger. The profitability surface the review asks
|
||||
for (development, fleet capital, share, margin, electricity, pre-production, discounting, residual, operator and
|
||||
manufacturer apart) takes these per-TH rows as its inputs; it is the economics lane's, not this file's.
|
||||
|
||||
|
||||
## 7. The lifetime answer: what each transition needs, and what is credited
|
||||
|
||||
The rule: a transition is credited only where the next entry needs a physical resource this design cannot supply
|
||||
economically; firmware changes (a program, a register, a weight) are credited zero. "Performance loss" is the change in
|
||||
the chip's energy per hash when the entry is live at its draw weight (4 points of 79 for a reserve family; the whole
|
||||
program for an atom or a shape), from the rows of sections 3 and 4; the GPU's own change on the same draw is beside it.
|
||||
|
||||
| Transition (the next epoch draws it) | Physical resource it needs | Does this design hold it? | Chip loss at the draw weight | The card's change on the same draw (measured ratio) | Credited obsolescence benefit |
|
||||
|---|---|---|---|---|---|
|
||||
| G1 to G7 (the ARX group) at any band weight | the ALU group | yes | 0 to +3 percent of the shadow energy across the band (the unit rows 3.1 to 3.3 pJ at N5 against the draw's 4.6) | 1.00 | 0 |
|
||||
| G8, G9, G6 (mul, mulhi, mad) at any band weight | the multiplier and the third read | yes | -1 to +2 percent at B = 4 (mul 2.9, mulhi 2.8, mad 5.2 pJ at N5) | mul 1.1, mulhi about 3.4 (the card's dearest ALU op) | 0 |
|
||||
| G10 shfl at its cap (8 points) | the lane butterfly | yes (per core) | 0 (2.9 pJ at N5, under the draw's mean) | 4.9x the add per op | 0 (the card pays 29.4 pJ for the move the chip pays about 1) |
|
||||
| R1 shfla live (4 points) | a general lane crossbar | yes (per core; the one network a butterfly cannot emulate in one op) | 0 (2.9 pJ at N5, under the draw's mean)A | 1.53 (NVIDIA), 1.91 (Apple) | 0 |
|
||||
| R2 perm live | a byte selector | yes | -1.9 percent | 1.30 | 0 |
|
||||
| R3 popc and clz live | a popcount tree, a priority encoder | yes | -2.5 percent | 1.50 and 1.63 | 0 |
|
||||
| R4 to R7 (bfe, shl and shr, sel, andn) live | a shifter, a mask, a select, an and-not | yes | -2.1 to -1.4 percent | 0.75 to 1.54 | 0 |
|
||||
| R8 mm8 live | a u8 dot4 per lane (8 chip ops per card tile) | yes | -0.8 percent (3.9 pJ per dp4a at N5, 1.0 pJ per MAC, against the card's 2.2 pJ per MAC at the lock) | the tile: 2.43 the add step per card instruction (32 MACs) | 0 (the chip's MAC is cheaper than the card's by 4x to 30x on the public figures; this is the card's loss, not the chip's) |
|
||||
| the op-mix band draw (B = 4 on injecting families) | nothing: the program | yes | within the family rows above | within 11 percent per instruction (shadow-k 6.2) | 0 |
|
||||
| the fold constants draw | five registers | yes | 0 | 0 | 0 |
|
||||
| the block shape draw (64, 128, 256) | the program-length register; 256 words of imem | yes | -3.7 percent at 64 (the imem term; 0 at 128 and 256) | 64 ran 2.5 to 3.5 percent faster than 256 on the 5090 and the M5 Max (measured) | 0 |
|
||||
| the read-width draw W = 4 (16 bytes) | three fwd instructions per load (firmware); the same DRAM sector | yes | +0.3 percent per hash (384 fold steps at 3.8 pJ) | within 2.7 percent on the 5090 and the 9070 XT, within 1 on the M5 Max (measured) | 0 |
|
||||
| the 64-register window | the 64 x 256 macro per 8 lanes | yes (built in) | 0 (it is the base) | modelled: occupancy to about half on a 5090 or 4090, the rate expected to hold (the hash lane) | 0 |
|
||||
| the dataset atom draw A1 / A2 / A3 (mixer x4, x8, dr368) | nothing per hash; the per-window build is a program on the same core (section 7.1) | yes | 0 per hash; the build 0.72 J per window per machine at 1 GiB (0.2 mW averaged), 1.45 J at 2 GiB | the card's rate unmoved by the atom (within 0.1 MH/s on the 5090, measured) | 0 |
|
||||
| a new op family outside the 18 (a bank refresh, a release) | a unit the die lacks | no: the chip emulates it from the 18 at the vendor penalty, as the cards do (1.5x to 2.4x per op, measured on the cards) or loses its 4 points | at 4 points of 79: at most 4 / 79 x (penalty - 1) of the shadow energy, about 2 to 6 percent (modelled) | the same emulation on every card that predates the release, 0 on a card with the native op | 0 unless the family is one the 18 cannot emulate; none proposed is |
|
||||
| a new read atom W = 8 (32 bytes, `admissible: false` today) | seven fwd instructions per load; the same GDDR7 sector; on an SRAM die +0.3 nJ of wire per read (floor lane 3) | yes | +0.7 percent per hash (896 fold steps) | free by the measured rows (one sector per load on NVIDIA) | 0 |
|
||||
| the dataset floor step (layer 2: 5.5 / 8.5 / 11.5 GiB) | device memory: 3 to 6 GDDR7 devices more, or 3 to 6 N2 reticles on the SRAM die | yes on DRAM (USD 60 to 120 more); the SRAM die's ticket rises USD 1,000 per step | 0 per joule on DRAM; the SRAM die's USD per MH/s unmoved (every die powered) | the tuned 5090 pays 4, 8 and 10 percent more energy per hash at 2, 4 and 8 GiB (measured) | 0 per joule; a capex ticket on the SRAM die only (floor lane 3) |
|
||||
|
||||
Reading, before the numbers: nothing in the bank asks for a resource the design lacks, because the bank is public at
|
||||
genesis and its whole op-family set is a few adders per lane and two networks per core. What the bank does to this
|
||||
chip is the per-op cost of carrying the unit set (sections 3 and 5: the full core against the base core on the same
|
||||
microarchitecture) and the leakage of the units that are not live, and that is the number the credit must come from.
|
||||
|
||||
### 7.1 The per-window dataset build on the chip (the atoms' only cost)
|
||||
|
||||
Under class v5 the dataset is rebuilt from the chain's state every window (3,600 s). A 1 GiB build is 157 G ops
|
||||
(measured as 13.4 ms on a 5090; the record's row 7). On this core at 4.61 pJ per lane-op (N5) that is 0.72 J per
|
||||
window per machine, 0.0002 W averaged over the window, against a machine of hundreds of watts: 0.0001 percent of
|
||||
the machine's energy, the same for every atom within the atom's op count (x4 half of x8, dr368 about x4's). The
|
||||
chip's node is the host the board model carries (one full node per 100 machines: 85 W and USD 1,500 shared); a
|
||||
specialised machine with a host keeping the dataset current is inside the adversary model by the review's rule (3),
|
||||
and this is its price: under 1 W and USD 15 per machine.
|
||||
|
||||
|
||||
## 8. Energy resistance, economic resistance and response capability, stated separately
|
||||
|
||||
Three separate things, each with its own number and its own holder.
|
||||
|
||||
**Energy resistance** is a property of the memory system and the shadow, not of the calendar. Against the
|
||||
adversary's re-optimised chip the complete GDDR7 machine reads 1.8x the 5090 at its lock per joule node-for-node
|
||||
(1.5x to 2.1x), 2.1x a node ahead; 1.6x and 1.9x against the 5080 at its lock; 2.8x and 3.3x against the cohort
|
||||
card; the N2 SRAM die 3.1x and 4.2x. The bank moves these by 5 percent in the honest side's favour (the chip's
|
||||
shadow energy +11 percent for carrying 18 families instead of 10) and no more; the register window moves them by
|
||||
0.03 to 0.13 of k and no more. What holds the per-joule number is the shadow's size on the card's own operating
|
||||
point and the memory system's activate ceiling; what would move it is a memory arrangement the chip cannot buy, and
|
||||
section 6 says there is none: the chip's cheapest memory is the card's own.
|
||||
|
||||
**Economic resistance** is capex per sustained MH/s and the project cost against the chain's revenue. The chip's
|
||||
machine costs USD 4.84 per MH/s against the card's 16 to 18, so over a 3-year life it mines at 0.080 USD per TH
|
||||
against 0.22 to 0.27: 2.7x to 3.4x, of which electricity is the smaller half. The calendar does not shorten that
|
||||
life: no transition in the bank retires the chip, so the 3-year stress life holds in full and the withdrawn headline
|
||||
("dies within an epoch, under 1x over its life") stays withdrawn. What holds the economics is the project (USD 20 M
|
||||
to 75 M for the GDDR7-board chip, floor lane 5; USD 100 M to 500 M for the SRAM die, claimed) against the miner
|
||||
revenue the chain pays, which is the profitability surface the review asks for and the economics lane owns; the
|
||||
dataset floor is a ticket on the SRAM die only (USD 1,000 per step) and nothing on the DRAM board (USD 60 per step).
|
||||
|
||||
**Response capability** is what the bank actually buys: the chain can change its object every 180 days without a
|
||||
release, and a chip that carries the bank follows by firmware at a per-transition loss of -3.7 to +0.7 percent of
|
||||
its shadow energy (zero credited). A family outside the bank costs that chip an emulation penalty of 2 to 6 percent
|
||||
of the shadow at 4 points, which every card that predates the release pays too; a structural change (a new read
|
||||
atom, a new derivation) costs the chip a host update and the cards a release. So the bank is a response channel
|
||||
whose value per event is the per-transition loss column, not a chip retirement; it is worth keeping for what it is
|
||||
(no fork for a family change, a fixed-function datapath dead on day one, section 7's table) and must not be served as
|
||||
energy or economic resistance.
|
||||
|
||||
|
||||
## 9. Consequences per tier (the standing rule)
|
||||
|
||||
| Tier | What the rows mean for it | What is being done |
|
||||
|---|---|---|
|
||||
| Home miner, one 8 GB card | Against the adversary's machine the cohort card sits 2.8x to 3.3x behind per joule and 3.7x per dollar; the bank does not change that, the operating-point lock does not reach this tier (no Blackwell lever on Ampere; Ada's lock is worth 1.2x); the first dataset step (5.5 GiB) retires this card from mining | the schedule's replacement cost (USD 350 to 450 to a 16 GB card) is stated beside the step; the 24 Gb device case shows the chip pays USD 700 once for the same horizon |
|
||||
| One 12 GB card | the same per joule; mines through the 8.5 GiB step, loses proving coexistence at the first step (7.3 + 8.6 GB over 12), leaves mining at 11.5 GiB | the mine-or-prove routing (0.3.21) and the replacement cost; the step offsets (section 11) |
|
||||
| One 16 GB card (the 5080 at its lock, the 9070 XT) | the 5080 at its lock is the honest NVIDIA floor: 1.6x behind the chip machine node-for-node, 1.9x a node ahead; the 9070 XT 4x to 5x behind; mines through every step, loses proving coexistence at 8.5 GiB | the per-dollar edge (3.3x) is the number the project-cost wall must answer; the lock stays the one lever |
|
||||
| One 24 or 32 GB card (the 5090 at its lock, the M5 Max) | 1.8x / 2.1x behind the chip machine; the M5 Max about 1.5x (reported); mines and proves through every step | nothing on the card side moves this; the rotation layers are response capability, not resistance |
|
||||
| A rig | per joule its cards; per dollar 3.3x behind the chip at MSRP, 4.8x at street; over 3 years 2.7x to 3.4x per TH | the issuance trigger and the share-pattern detector (the record's item 4) stay the instruments; the profitability surface is the economics lane's |
|
||||
| A pool user | a chip fleet is a few operators at 1.1 to 1.3 microjoules; the detector on the observer is what tells a pool a chip has arrived | unchanged |
|
||||
| The public claim | the chip line this file supports: "a chip that survives rotation is a GPU-shaped core on the card's own memory; it reads about 1.8x the best honest card per joule on the same node, 2.1x a node ahead, and about 3x per dollar; the rotation calendar changes its firmware, not its cost" | served only on the coordinator's word, after the placed rows |
|
||||
|
||||
## 10. Sources and what is owed
|
||||
|
||||
- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured);
|
||||
the floor programme's tier table (`docs/design/class-v6-rotating-family.md` 10.4: the 5090 at the 1,300 lock 2.33
|
||||
microjoules at 134.76 MH/s and 312.5 W, the 5080 at the 1,100 lock 2.06 at 71.20 MH/s and 146.6 W, stock 3.36 and
|
||||
3.48, the 4070, 4090, 3090 and 9070 XT rows, the M5 Max 1.40); the reserve families' step costs, design 4.2
|
||||
(measured on the cards); the card population by count, design 3.4 (approximate).
|
||||
- The chip side: this lane's RTL and flow (`tools/chip-model/mf/`), ORFS on ASAP7 (Clark et al., Microelectronics
|
||||
Journal 2016; the 7.5-track RVT library, TC corner), FakeRAM2.0 macros (ABKGroup, the ORFS platform's
|
||||
`fakeram7_64x256` and `fakeram7_256x34`, LEF and Liberty; the generator's own config marks its values "not
|
||||
realistic", so only area, pins and placement are taken from it); the k lane's rows (`floor/shadow-k.md`, the
|
||||
shuffle row 1.24 pJ per lane-op routed, relayed 16:0x BST) cited as published.
|
||||
- The SRAM access energy band: Horowitz, "Computing's energy problem (and what we can do about it)", ISSCC 2014
|
||||
(45 nm: 8 KB SRAM 10 pJ per 64-bit read, 32 KB 20 pJ); the scaling to a 7 nm class node approximate; the k lane's
|
||||
imem figure (2 to 4 pJ per 32-bit read, approximate); CACTI-class estimates (approximate).
|
||||
- Node scaling (claimed): TSMC's technology pages for N5, N3E and N2, read 8 October 2026 (shadow-k section 7).
|
||||
- The board: `docs/analysis/chip-model-v3.md` 5.3 and 5.5 (the GDDR7 and HBM3 random-read engines, the activate
|
||||
ceilings, the static and controller allowances, the prices, all modelled or claimed; the 5090 at 82 percent of the
|
||||
GDDR7 ceiling, measured); `floor/sram-and-floor.md` 2.1 (the SRAM die at W = 1, modelled); the power delivery,
|
||||
cooling and host allowances are this file's (approximate; the sensitivity is stated beside each).
|
||||
- The dataset build: the record's row 7 (a 1 GiB rebuild 157 G ops, 13.4 ms on a 5090, measured).
|
||||
|
||||
Owed: the k lane's crossbar, scratch and tile rows (in place and route at 16:0x BST); the placed 32-lane core (the
|
||||
8-lane core is placed; the 32-lane row is synthesis only); a real PDK memory compiler's figure for the two macros
|
||||
(FakeRAM gives area and pins only); the 64-register window's GPU cost measured (the hash lane's generator line);
|
||||
the HBM3 activate ceiling (unmeasured, the AWS F2 hour); the profitability surface of the review's rule (2) over the
|
||||
lifetime rows of section 6, which this file states as cost per TH and leaves the NPV to the economics lane.
|
||||
|
||||
|
||||
## 11. The data-local and memory-sharing adversary (the ProgPoW audit's threat; the third review, 17:5x BST)
|
||||
|
||||
The threat: split the dataset across processors and move the intermediate computation to the processor nearest
|
||||
the next item, share datasets, keep partial caches, recompute, re-lay the dataset, run several engines off one
|
||||
set-up; price the cheapest combination of moving state, moving data, recomputing and local resources against the
|
||||
live state of class v6, and say whether the necessary state makes data-local execution dear enough to erase any
|
||||
memory-system saving. The cost model, explicit, every term labelled:
|
||||
|
||||
| Term | Per dependent read | Source |
|
||||
|---|---|---|
|
||||
| The live state a hash carries across a read | 64 registers x 32 bits = 2,048 bits (61 of 64 necessary across the chain, the review's reading of the fold rule) plus pc, nonce and era pointer about 64 bits: **about 2,100 bits** | the class program (`verify::fold_words`: every register is consumed) |
|
||||
| The data a read returns, with its address and control | 32 bits of data, 32 of address, about 16 of control: **about 80 bits** at W = 1 (176 at W = 4, 304 at W = 8) | floor lane 3, section 2.1 (modelled) |
|
||||
| Moving bits on one die (global wire) | 1.3 pJ per bit across a 24 mm die (0.65 to 2.6) | floor lane 3 (approximate) |
|
||||
| Moving bits between dies on a package | UCIe 0.5 pJ per bit (0.25 to 0.5) | claimed (UCIe via SNIA), floor lane 3 |
|
||||
| Moving bits between packages on a board | 5 to 10 pJ per bit (a PCB SerDes link; approximate, from memory) | approximate |
|
||||
| A dependent random 32-byte read from GDDR7 / HBM3 | 2.0 / 1.2 nJ (1.5 to 2.6 / 1.0 to 1.5) | chip-model-v3 5.3 (modelled) |
|
||||
| An on-die SRAM read at W = 1 | 0.25 nJ (0.20 to 0.35) including the wire | floor lane 3 (modelled) |
|
||||
| Recomputing one item instead of reading it | 9,360 ops, about 43 nJ on this core at N5 (4.61 pJ per op) and 6.3 nJ on a wired mixer pipeline (chip-model-v3 5.2) | modelled |
|
||||
| The per-window set-up (the 1 GiB build and the host) | 0.7 J per machine per window, 0.85 W and USD 15 of host per machine | section 7.1 |
|
||||
|
||||
The four forms, priced per read at W = 1 (the hash's own width), nominal with the band:
|
||||
|
||||
| Form | What moves | Cost per read | Against the baseline (data to the lane, 80 bits) |
|
||||
|---|---|---|---|
|
||||
| Baseline: the state stays in its lane, the data travels to it | 80 bits of data, address and control | on a die 0.10 nJ (0.05 to 0.21); on a package 0.04; the DRAM read itself 2.0 nJ beside it | 1x |
|
||||
| Data-local on one die: the state travels to the macro holding the item | about 2,100 bits | 2.7 nJ (1.4 to 5.5) of wire per read: **27x the baseline's wire, and 11x the whole GDDR7 read** | 27x |
|
||||
| Data-local across dies on a package (the dataset split over chiplets) | about 2,100 bits over UCIe plus the local read | 1.05 nJ (0.5 to 1.05) of hop per read: half a GDDR7 read, four SRAM reads | 26x the baseline hop |
|
||||
| Data-local across packages on a board (the dataset split over boards) | about 2,100 bits over a SerDes link | 10 to 21 nJ per read: 5x to 10x a GDDR7 read | 260x |
|
||||
| Memory sharing: N engines over one dataset | nothing new: the engines are the lanes in flight the baseline already has (1,172 on the GDDR7 board at the activate ceiling); the set-up is shared at 0.7 J per window | 1x | 1x |
|
||||
| A partial store that recomputes the rest (the pebbling curve, chip-model-v3 5.4 with the two corrections) | a recomputed item costs 9,360 ops | 43 nJ on this core, 6.3 nJ wired, against 2.0 nJ read: **monotone, the full store is the cheapest point at every f** | 3x to 21x per recomputed item |
|
||||
| A hot-item SRAM beside the DRAM (the window-layer distribution: the hottest half of the items serves 72 percent of the reads, adv-cache-2) | an SRAM read for 72 percent of the reads, a DRAM read for the rest | 0.25 x 0.72 + 2.0 x 0.28 = 0.74 nJ per read on the memory side, for USD 250 per GiB of SRAM (about half the dataset) | a capex trade bounded by the GDDR7 row above and the SRAM-die row below; no new form |
|
||||
| An alternative layout (the 16 load sites and their era windows: bank the dataset by site) | nothing per read: every read is still an activate on the device that holds the item | 1x | 1x |
|
||||
|
||||
Reading. (1) The necessary state is the whole argument: a read returns 80 bits and the state that must meet it is
|
||||
2,100, so moving the computation to the data costs 26x the bits of moving the data to the computation, on every
|
||||
medium. On one die that is 2.7 nJ of wire against 0.10; on a package half a DRAM read; across boards five to ten DRAM
|
||||
reads. There is no memory-system saving for data-local execution to erase: the dependent read must be served by the
|
||||
device that holds the item in every form, and what the data's trip to the lane costs (0.10 nJ on a die, 0.04 on a
|
||||
package) is already the cheapest term in the model. So the ProgPoW audit's threat does not apply to a hash whose
|
||||
live state is 26x its read width; it applies to a hash whose state is a few words, which class v6 is not. (2)
|
||||
Memory sharing is the baseline, not an attack: the chip's lanes already share one dataset and one set-up; the
|
||||
per-hash share of the set-up is 10^-15 J and of the host USD 0.15 per machine. (3) Partial stores and recomputation
|
||||
are priced by the pebbling curve and lose at every point; the hot-item SRAM is a capex trade between the two rows
|
||||
of section 6, not a new form. (4) Cumulative memory complexity of one evaluation, stated as the bound a chip must
|
||||
pay: 128 dependent random reads, each an activate and a 32-byte sector, 4 KB of sector traffic and 128 x 2,100 bits
|
||||
of state carried in a lane (never moved); 256 nJ of memory energy per hash on GDDR7 (154 on HBM3, 32 on the SRAM
|
||||
die) plus 0.47 microjoules of shadow on this core at N5. Bandwidth hardness: the GDDR7 board is bound by activates
|
||||
(21.3 G per second over 16 devices, 166 MH/s per board) at 38 percent of its pin bandwidth (682 GB/s of sectors of
|
||||
1,792), so the hard quantity is the activate rate per dollar of devices, not bytes per second, and a chip cannot buy
|
||||
more activates per device than the card has. (5) Bounded by this cost model: data-local execution (26x by the bit
|
||||
count, no physical design needed), memory sharing (the baseline), partial stores (monotone), layouts (per read
|
||||
invariant). Needing physical design before a number is a bound: the HBM activate ceiling (unmeasured; the AWS F2
|
||||
hour), the SRAM die's wire term (0.5 to 2.0 nJ per 64 bytes until a placed macro array exists), the on-package hop
|
||||
(UCIe's 0.5 pJ per bit is a claim), and the data-local form on a 3D-stacked SRAM (a vertical hop at about 0.1 pJ per
|
||||
bit, approximate, would cut the one-die row to about 0.2 nJ, still 2x the baseline and still no saving).
|
||||
|
||||
## 12. The dataset comparison that decides class v7 (the third review, 17:5x BST)
|
||||
|
||||
Two datasets, each with the adversary's burden (a chip with a host keeping it current) and the commodity burden
|
||||
(what every honest node pays at the boundaries), priced on the record's figures:
|
||||
|
||||
| | State-coupled dataset (class v5 and v6: derived from the chain's execution state at the era cut) | Epoch-defined bounded dataset with a published support horizon (derived from the certified checkpoint's hash at the era cut; the size schedule published years ahead) |
|
||||
|---|---|---|
|
||||
| Adversary: sync bandwidth | the day stream and the leaves: 16.5 KB/s to 10,000 members, 45 MB per member per window today (the record 2a.2); grows with the state | the seed: 32 bytes per era; the schedule: a constant |
|
||||
| Adversary: update cost | the rebuild 0.7 J per machine per window (0.0001 percent of its energy); the host's execution of every block (85 W, USD 1,500 per 100 machines: 0.85 W and USD 15 per machine, under 1 percent) | the rebuild only (the same 0.7 J); a light client following checkpoints (bytes, watts of nothing) |
|
||||
| Adversary: storage | the state (hundreds of MB today, unbounded) on the host; the dataset on the machine | the dataset only |
|
||||
| Adversary: adversarial state growth | raises the HOST's cost (storage, re-execution), never the dataset's size (the atom folds the state into a dataset of the schedule's size); the growth is paid in gas by whoever causes it | none |
|
||||
| Adversary: what is excluded | the stale machine (its dataset wrong the moment its host is) and the recompute chip; not a specialised machine with a host (the review's rule 3), which pays under 1 percent | the stale machine (a chip missing the epoch flip mines a dead object, as a stale card does); the recompute chip is excluded by the pebbling curve, not by the dataset; the seed is unknown before the checkpoint, so nothing precomputes |
|
||||
| Commodity: the boundaries | today's eight crossing faults were all state boundaries: the crossing, the partition, the snapshot, the cold start (the 15:42Z deep-reorg reset, the pruning-point anchor, the snapshot resumed under the next object, the stale stream); every family epoch is a state-agreement event (layer 3, section 3) | one boundary kind: the checkpoint's hash, which every node holds by the chain's own rule; no stream, no snapshot, no re-execution on the mining path |
|
||||
| Commodity: re-execution per node | about 10 minutes from genesis on Devnet 3, hours on the shared devnet, and the dataset is wrong until it is done | none for mining (execution continues for the EVM, decoupled from the hash) |
|
||||
| Commodity: proving coexistence | unchanged by the coupling (set by the dataset's size and the prover's 8.6 GB) | the same |
|
||||
| Commodity: the cold start and the partition | a node that missed the lead serves nothing for a window; a partition's two sides derive two datasets | a node derives the dataset from any header it accepts; a partition's two sides derive the same dataset while they share the checkpoint |
|
||||
|
||||
The growth schedule 5.5 / 8.5 / 11.5 GiB, per step (the hash lane's VRAM rows: the dataset needs about 1.3x its size
|
||||
in device memory with the working set; the tier room 75 percent of card memory, 50 percent of Apple unified; the
|
||||
prover 8.6 GB beside it):
|
||||
|
||||
| Step | Adversary burden | Commodity burden: who leaves mining | Who loses proving coexistence | Owners' replacement cost (approximate street prices, October 2026) | The 24 Gb GDDR7 case |
|
||||
|---|---|---|---|---|---|
|
||||
| 5.5 GiB at the v6 epoch | GDDR7 board: 0 (16 x 2 GB holds it); SRAM die: 3 reticles, USD 1,500 of ticket | the 6 GB and 8 GB tiers (25 percent by count): about 7.3 GiB of device memory needed | the 12 GB tier (22 percent): 7.3 + 8.6 over 12; it mines or proves | an 8 GB card to a 16 GB card: USD 350 to 450 (a 16 GB 5060 Ti or 9060 XT class) | none needed |
|
||||
| 8.5 GiB two years on | GDDR7 board: 0 (the 32 GB board holds it); the SRAM die 5 reticles, USD 2,500 | the 10 and 11 GB tiers (9 percent) and the 16 GB unified Mac (about 8 GiB of room): about 11 GiB needed | the 16 GB tier (28 percent): 11 + 8.6 over 16; it mines or proves | a 10 or 12 GB card to a 16 GB card: USD 350 to 450; a 16 GB Mac has no upgrade | none needed |
|
||||
| 11.5 GiB at four years | GDDR7 board: 0; the SRAM die 6 reticles, USD 3,000 | the 12 GB tier (22 percent): about 15 GiB needed | the 24 GB tier mines and proves (15 + 8.6 under 24) | a 12 GB card to a 16 GB card: USD 350 to 450 | none needed |
|
||||
| The 24 Gb (3 GB) GDDR7 devices (Micron ended 2 GB production, 3 GB at USD 60 to 70 each, the TrendForce note) | a 16-device board at 48 GB for USD 1,000 of memory instead of 320: the adversary over-provisions ONCE for the whole horizon, USD 9.9 per MH/s instead of 4.84, still 1.6x the card per dollar and unchanged per joule | the 50-series cards on 3 GB devices (the 5090 32 GB, the 5080 Super class at 24 GB) carry the schedule to year 4 and beyond; the schedule retires the 8 and 12 GB tiers, not the chip | | | the schedule's one effect on the chip is a USD 700 memory ticket, paid once |
|
||||
|
||||
Reading: the schedule is a commodity-side cost at every step (a quarter of the cards by count at 5.5 GiB, a further
|
||||
third losing proving coexistence at 8.5, the 12 GB tier out at 11.5; USD 350 to 450 per displaced owner) and a chip-side
|
||||
cost only on the SRAM die, and the 24 Gb device removes even the DRAM board's small ticket. The state coupling buys,
|
||||
against the chip, the exclusion of a stale machine (which the epoch seed also buys) and a host at under 1 percent of
|
||||
the machine's cost; it costs the honest side every state boundary the devnet has crossed this week.
|
||||
|
||||
The recommendation, five lines:
|
||||
1. Class v7 derives the dataset from the certified checkpoint's hash at the era cut (an epoch-defined bounded
|
||||
dataset), with the size schedule published at genesis as the support horizon; the execution state stays in the
|
||||
headers for the EVM and leaves the hash.
|
||||
2. The dataset floor schedule stays (5.5 / 8.5 / 11.5 GiB, steps offset from the family flips), served as a ticket
|
||||
on an SRAM die and as a retirement of the 8 GB tier at the first step and the 12 GB tier at the third, with the
|
||||
replacement cost per owner stated; the 24 Gb device is priced in as the DRAM chip's one-time USD 700.
|
||||
3. The family bank and its 180-day calendar stay, served as response capability with a zero obsolescence credit
|
||||
per transition (section 8), never as energy or economic resistance.
|
||||
4. The chip line served is the complete machine's: 1.8x per joule node-for-node and 2.1x a node ahead on the
|
||||
GDDR7 board against the Blackwell tier at its lock, 3x against the cohort, 3.3x per dollar; the hold is the
|
||||
project cost against miner revenue (the profitability surface, the economics lane).
|
||||
5. The data-local threat is closed by the live state (26x the bits of the read; no saving to erase) and needs no
|
||||
design change; the three terms that still need physical design before they are bounds (the HBM activate
|
||||
ceiling, the SRAM die's wire, the on-package hop) are named in section 11 and owed.
|
||||
|
|
@ -439,7 +439,7 @@ The conditions, read off the surface, which are the economic-resistance statemen
|
|||
|
||||
#### 10.0h The served text (for the site audit lane, served as written; each number's label; what is withdrawn)
|
||||
|
||||
**The claim statement above the sentence (the third review, 10.0l item 3), served verbatim:** "Igneum remains competitive on accessible commodity GPUs even when specialised mining hardware is assumed to exist, remain compatible and seek profit; its security does not rely on identifying that hardware or retiring it through emergency changes." Under it the three statements (energy, economic, response capability), rotation named as an optional improvement to the baseline, not the mechanism; the economic statement reads from the coexistence model (10.0l item 2, by 21:00 UK) once it exists, and until then names the model as owed rather than any threshold.
|
||||
**The claim statement above the sentence (the third review, 10.0l item 3), served verbatim:** "Igneum remains competitive on accessible commodity GPUs even when specialised mining hardware is assumed to exist, remain compatible and seek profit; its security does not rely on identifying that hardware or retiring it through emergency changes." Under it the three statements (energy, economic, response capability), rotation named as an optional improvement to the baseline, not the mechanism; the energy statement reads whole-machine per tier (10.0n: the complete GDDR7 machine 1.8x the 5090 at its lock node-for-node, 2.1x a node ahead, 1.6x / 1.9x the 5080, 2.8x / 3.3x the cohort, synthesis, placed rows to follow; the core-only rows beside it marked); the economic statement reads from the coexistence model (10.0n; `docs/analysis/class-v6/coexistence-model.md`, landed, modelled): the GDDR7 board passes all six success conditions at a one to three year life, the N2 SRAM die fails five once built with development sunk, and the larger half of a chip's economic edge is capital cost per accepted hash, not joules; the response-capability statement says rotation costs a chip versatility, not life.
|
||||
|
||||
**The sentence, as the external review words it (10.0f item 5), served verbatim, with one word made honest (the site audit lane's read, 16:5x UK: class v4 and v5 have eight registers per lane and the window is class v6's new core shape, so "retains" is read as "retains across every rotation"):** "Class v6 adopts the 64-register window and retains it across every rotation. Current modelling estimates a 2.2x to 2.4x energy-efficiency advantage for the strongest specialised designs assessed against the GPU tier (2.0x on the GPU's own node). The long-program and select-tree proposals were rejected. Economic resistance depends on development cost, deployment economics and productive hardware lifetime; family transitions receive an obsolescence benefit only where a loss of competitiveness is demonstrated; programmable multi-epoch designs are included in the assessment."
|
||||
|
||||
|
|
@ -502,6 +502,31 @@ For the window line of the close: this is a second measured card row (after 10.0
|
|||
|
||||
**(3) The claim statement:** "Igneum remains competitive on accessible commodity GPUs even when specialised mining hardware is assumed to exist, remain compatible and seek profit; its security does not rely on identifying that hardware or retiring it through emergency changes." Under it the three statements of 10.0g item 1 (energy, economic, response capability), with rotation named as an optional improvement to the baseline, not the mechanism: 10.0d's schedule is what the rotation does when it runs, and the claim stands without it.
|
||||
|
||||
#### 10.0m Amendment (17:2x UK): the multi-family adversary lane's full rows (`docs/analysis/class-v6/multi-family-adversary.md` on class-v6-adversary; synthesis only, ASAP7, a gate-level VCD; the SRAM term modelled with its band; node factors claimed; board rows modelled; for class v7, and the answer to the review's obsolescence question)
|
||||
|
||||
**The answer to the review's question (10.0f item 1): the family bank costs the chip firmware plus 43 percent of its core cells and 11 percent of its shadow energy; no epoch costs it a part, so the credited obsolescence benefit is zero for every transition in the bank.** The 18-family core on the class v4 draw: 6.58 pJ per lane-op at ASAP7 (5.8 to 8.3), 4.61 at N5, 3.32 at N3; k 0.45 node-for-node and 0.32 a node ahead at the 5090's lock; the genesis-only core on the same microarchitecture 5.91 / 4.13 / 2.98 (k 0.40 / 0.29). The reserve families at the lock, k N5 / N3: shfla 0.37 / 0.26, prmt 0.25 / 0.18, popc 0.25 / 0.18, clz 0.23 / 0.17, bfe 0.25 / 0.18, shl and shr 0.57 / 0.41, sel 0.54 / 0.39, andn 0.44 / 0.32, mm8 0.44 / 0.32 (1.0 pJ per MAC against the card's 2.2), lop3 0.27 / 0.19; every live-family mix sits under the class v4 draw (all eight live -11 percent): **the card's dearest instructions are the genesis ones, so the reserve is datapath diversity, never joules.** Per-transition loss -3.7 to +0.7 percent of the shadow energy (the shape 64 -3.7, W = 4 +0.3, W = 8 +0.7, an atom 0 per hash and 0.7 J per window for the build).
|
||||
|
||||
The whole hash at N5 / N3: the GDDR7 board 2.5x / 2.9x against the 5090 at its lock, 2.2x / 2.6x the 5080, 1.5x / 1.7x the M5 Max (reported, not headlined); HBM3 one stack 2.9x / 3.5x; the N2 SRAM die at W = 1 4.6x / 6.2x. **The complete machine** (controller, host share, PSU, VRM, cooling in): the GDDR7 machine 1.28 microjoules and USD 4.84 per MH/s, **1.8x the 5090 at its lock per joule node-for-node (1.5x to 2.1x), 2.1x a node ahead, 1.6x / 1.9x the 5080, 2.8x / 3.3x the cohort card; 3.3x per dollar at MSRP**; HBM3 one stack no better per joule and 2x to 44x worse per dollar (the adversary's cheapest memory is the card's own); the SRAM die 3.1x / 4.2x per joule, 7x per dollar. The three-year cost per TH: USD 0.080 against the 5090's 0.220 and the cohort's 0.270; electricity a quarter of the machine's cost, so **capex per MH/s is the larger half of the economic edge**, which is the coexistence model's first input (10.0l). The k lane's gated 64-register flop core (6.2 pJ) and this lane's macro core (5.9) agree within 5 percent, which closes 10.0j's disagreement: **the window's form is not the lever; in the macro form its residual over 32 registers is 0.3 to 0.4 pJ** (about 0.05 of k at N3), so the window stays as a measured-cheap, small defence and is not the served edge. The placed and 32-lane rows follow; the two review additions by 21:30 UK.
|
||||
|
||||
#### 10.0n The three statements, re-read on the whole machine and the coexistence model (the coordinator, 17:3x UK; the multi-family adversary lane's whole-machine rows, synthesis with the SRAM band and the node factors claimed, placed rows to follow; the coexistence model's first run, `docs/analysis/class-v6/coexistence-model.md`, lane 3 at 7cf39665, 16:24 UK, every row modelled, landed with this amendment; the connected-state lane's kill, 17:27 UK)
|
||||
|
||||
**ENERGY resistance, whole-machine, per tier (the core-only rows of 10.0e and 10.0i beside it, marked):** the complete GDDR7 machine (controller, host share, PSU, VRM, cooling in; 1.28 microjoules per hash, USD 4.84 per MH/s) reads 1.8x the 5090 at its lock per joule node-for-node (1.5x to 2.1x), 2.1x a node ahead; 1.6x / 1.9x the 5080; 2.8x / 3.3x the Ada, Ampere and 9070 XT cohort; about 1.5x the Apple tier (reported, not headlined). Synthesis with the SRAM band, the node factors claimed, placed rows to follow (the k lane 17:30, the multi-family lane 21:00). The core-only rows beside it: the 8-lane gated core 2.9x to 3.3x at the lock a node ahead (10.0i), the 32-lane core 2.6x / 2.9x (10.0m).
|
||||
|
||||
**ECONOMIC resistance, from the coexistence model (lane 3's first run, modelled, the sunk-development case first as the mandatory one; the measure cost per accepted MH/s-hour, accepted = 97 percent of raw, a 1 percent pool fee; the population lane 4's class v5 table at street prices; the chip rows lane 4's columns at 3.2 pJ per forced op; three electricity prices 0.06 / 0.12 / 0.25 per kWh):**
|
||||
|
||||
| | at USD 0.06 / 0.12 / 0.25 per kWh, micro-USD per accepted MH/s-hour | Label |
|
||||
|---|---|---|
|
||||
| The best GPU owner (the 5070 Ti at its knee; the 5080) | 156 / 263 / 493; 204 / 332 / 611 | modelled on measured joules |
|
||||
| The best GPU entrant (a new 5070 Ti, two years, resold) | 415 / 521 / 751 (hardware 59 to 64 percent of it) | modelled |
|
||||
| The GDDR7 board, development sunk, life 1 y / 3 y, a farm | 750 / 799 / 906; 307 / 356 / 463 | modelled |
|
||||
| The N2 SRAM die, development sunk, life 1 y / 3 y | 134 / 163 / 225; 72 / 101 / 163 | modelled |
|
||||
|
||||
The advantage, separated (the chip at 0.06): the board over a Blackwell owner 0.5x to 1.1x and over the entrant 1.4x to 2.3x at three years (operating 2.2x to 5.2x, hardware 1.3x to 1.9x); the die 2.2x to 4.6x over the owner and 5.8x to 9.9x over the entrant (operating 3.7x to 9x, hardware 9x to 14x). A 2.6x chip at 0.06 against a GPU at 0.25 is 10.9x on power alone, as the review says, but power is 60 to 70 percent of the owner's cost and 30 to 40 of the entrant's, so the total is a third to a half of that. **The larger half of the chip's economic edge is capital cost per hash, not joules** (the multi-family lane: 3.3x per dollar at MSRP, 4.8x at street; the three-year cost per TH USD 0.080 against the 5090's 0.220; electricity a quarter of the machine's cost), so the coexistence model weighs capex per accepted hash as its main term, the sunk-cost case especially. Break-even electricity: a GPU owner with sunk hardware matches the board at three years up to 9 to 15 cents on Blackwell (4 to 6 on Ada and Ampere), at one year up to 27 to 40; matches the die at 0 to 5 cents (no grid price keeps a GPU level with a sunk die). Five years, three price paths, the sunk fleet entering in year 2, GPUs entering at the cheapest entrant's cost and exiting in cost order, a 1.5x GPU generation at year 3: the board at a sunk USD 10 M holds 10 to 31 percent of the chain at a 21 to 41 percent gross margin with GPU entrants setting the price in every path (coexistence); the die at USD 1 M holds 1 to 22 percent in growing and flat paths and takes the chain in year 5 of a shrinking one, at USD 10 M takes 37 to 73 percent on landing and 100 percent by year 3 on flat and shrinking paths, at USD 100 M takes 100 percent on landing in every path and then runs at a loss, with 0 of 17 GPU classes above water. The accessible supply: the chain's hash at the GPU entry equilibrium is 2.6 / 8.8 / 26 / 88 TH/s at IGN 0.03 / 0.10 / 0.30 / 1.00, which is 34,000 to 1.1 M 5070 Ti-class cards (under 4 percent of a generation's shipments) against 480 to 16,000 SRAM dies (8 to 270 N2 wafers) or 16,000 to 530,000 GDDR7 boards; the dependence on one supplier total for the die, partial for the board, nil for GPUs.
|
||||
|
||||
**The conditions under which the success statement holds (the model's section 9, six rows, each with its number): (a) the chip's all-in cost within about 1.5x of the best GPU owner at the GPU's electricity; (b) the chip's hardware per MH/s not under a quarter of the GPU entrant's; (c) a fleet above a third of the chain costing more than a year's miner revenue; (d) GPUs keeping resale and a use outside mining; (e) the per-joule gap at the honest knee under about 3x; (f) the supplier's margin a normal return that does not rise with the halvings. The GDDR7 board passes all six on today's rows at a one to three year life. The N2 SRAM die fails (a), (b), (c), (e) and (f) at every life, price path and electricity price once it exists with development sunk, and the only condition that holds it is that nobody pays to build it (the surface: IGN 0.73 for a USD 150 M project at a third of the chain, 0.22 taking the chain, 0.12 to 0.16 for a revision), which is a statement about an investor's decision, not a level the chain stays below.** What the chain controls: the honest side's efficiency (conditions a and e), the detector (c's visibility), the dataset floor as a ticket (c by 2x to 4x); nothing in the hash moves (a), (b) or (e) for the die by the factor they need. Unverified: every chip figure; the chip's USD per MH/s (within 2x; the term (b) and (c) rest on); street prices; the Ada and Ampere knees; the single-rule GPU reaction (a second cut adds a per-class supply curve, the installed base as a cap, re-entry on a price rise).
|
||||
|
||||
**RESPONSE capability:** rotation costs a chip versatility, not life. The 18-family bank costs a chip firmware plus 43 percent of its core cells and 11 percent of its shadow energy, with zero obsolescence credit on any transition in the bank (10.0m); a passed boundary proves the rotation works (10.0d). The window: its form is not the lever (the gated flop file and the macro file agree within 5 percent; the residual over 32 registers 0.3 to 0.4 pJ); its cost to the card is under 1 percent, measured on a 5090 (+0.6 percent) and a 4090 (-0.9 percent) at stock on 8 October (the connected-state lane, replacing "about 0, unmeasured"); and the connected-state class is KILLED (the connected window moves the chip's edge 1.10x and 1.08x against the 1.25x gate on the k lane's re-optimised core; necessity costs a clock-gated file nothing, since it pays per write, not per live register; only the window's width reaches the chip, +0.14 of k).
|
||||
|
||||
#### 10.0d The rotation schedule the close adopts (the rotation lane, `docs/design/class-rotation-four-layers.md` on class-v6-rotation at bd43f808, build-3, gate green, 14:4x UK; one line per layer; both of this document's constraints held: the 180-day family epoch not shorter, W = 4 not drawn)
|
||||
|
||||
| Layer | Boundaries a year | What it draws, from where | Exposure per boundary (this document's units) | Chip | Label |
|
||||
|
|
|
|||
|
|
@ -28,3 +28,4 @@ docs/plans/cryptanalysis
|
|||
docs/analysis/class-v6
|
||||
# 8 October 2026: the miner window audit (an internal design record: the founder's order, lane ids, the build plan)
|
||||
docs/design/app-audit-2026-10-08.md
|
||||
docs/analysis/class-v6/coexistence-model.md
|
||||
|
|
|
|||
Loading…
Reference in a new issue