multi-family adversary: the placed full-core rows, the board and D2(b) at the placed energy, the transition matrix, the one page

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 7e715b50b (10b901a46934c24ce9b44c74bc6fd4db7c40d2b0) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 16:09:55 +00:00
parent 74678180d4
commit 2b7ad8fd27

View file

@ -19,7 +19,23 @@ population, Apple reported and not headlined; every defence is scored after the
## 0. One page
(filled from the rows below at each cut; section 8 carries the three-part statement)
1. The calendar changes the chip's firmware, not its parts. One programmable core carries every bank entry: 18 op
families, the 3 atoms as programs, the fold with its constants in registers, the 3 shapes, both read widths, the
64-register window in an SRAM macro. Carrying the bank costs it 20 percent of its energy on the class v4 draw
(placed and routed) and 41 percent of its cells; no transition needs a part it lacks; the credited obsolescence
benefit is zero for every entry (section 7).
2. The rows (ASAP7, placed and routed, SPEF, gate-level VCD; the SRAM term modelled, bands shown; node factors
claimed): the full core 9.36 pJ per lane-op at ASAP7 on the class v4 draw (8.6 to 11.1), 6.55 at N5, 4.72 at N3;
k 0.64 node-for-node and 0.46 a node ahead at the 5090's lock; the reserve families k 0.33 to 0.86 placed; the
genesis-only core 7.78 (k 0.53 / 0.38); the 32-lane genesis core 10 percent under (synthesis).
3. The complete machine (memory, controller, host share, power train, cooling) on the card's own GDDR7: 1.5x the
5090 at its lock per joule node-for-node (1.3x to 1.7x), 1.8x a node ahead (1.5x to 2.0x); 3.3x per dollar; HBM3
buys it nothing; the N2 SRAM die 2.3x and 3.1x; the stored-half hybrid (section 13) 1.9x and 2.4x at USD 2.80 per
MH/s. Capex per MH/s is the larger half of the chip's edge; the project cost is its hold.
4. Data-local execution cannot pay (the live state is 26x the read, section 11); memory sharing is the baseline;
recomputation loses; selective participation gains 2 to 5 percent for 50 to 90 percent of revenue (section 13).
5. Class v7 should take the epoch-defined bounded dataset with a published support horizon and keep the floor
schedule and the bank; the bank is served as response capability, never as resistance (sections 8 and 12).
## 1. The design: what the adversary builds
@ -258,8 +274,7 @@ and detailed placement, CTS, global and detailed routing, OpenRCX parasitics; Op
SPEF under the gate-level VCD; the clock at 12,000 ps with leakage restated at the 1,500 ps equivalent, section 2.1;
the final PDN connectivity check reports two macro power pins unconnected, a floorplan artefact of the FakeRAM
pin geometry that does not touch the netlist, the parasitics or the power figure). The 8-lane genesis core (the
comparator) first; the 8-lane full core's routed power run is in its VCD sweep (11 of 58 tags at 17:4x BST) and
replaces the scaled estimate below when it lands.
comparator) first, then the 8-lane full core (routed, 58 tags, 18:0x BST).
The genesis core, 8 lanes, routed: 161,504 cells (the synthesis 66,973: placement adds fill, buffers and the clock
tree), 2 macros (the base core's instruction word fits 34 bits, so Yosys removed the second imem macro):
@ -280,28 +295,73 @@ tree), 2 macros (the base core's instruction word fits 34 bits, so Yosys removed
| mf8base | power | shfl | 161504 | 0.000658 | 0.000264 | 3.2 (0.89 / 1.1) | 0.95 / 1.69 / 3.38 | 4.89 (4.15 to 6.57) | 3.42 | 2.46 | 1.77 | 55.8 / 29.4 | 0.116 | 0.0838 | 0.0441 |
| mf8base | power | load | 161504 | 0.000926 | 0.000263 | 5.22 (0.89 / 3.1) | 0.95 / 1.69 / 3.38 | 6.91 (6.17 to 8.59) | 4.83 | 3.48 | 2.51 | 13.9 / 8.3 | 0.583 | 0.419 | 0.25 |
Reading. Placement and the clock tree add 32 percent to the class v4 draw on the genesis core (7.78 pJ per lane-op
routed against 5.91 synthesised at ASAP7; 5.44 against 4.13 at N5), inside the k lane's +20 to +40 percent
expectation; the sequential term barely moves (0.90 against 0.88 pJ) and the combinational and clock term carries
the increase (3.9 against 3.0), which is wire capacitance on the macro output nets and the result reduction. The
per-family order is unchanged. Node-for-node the placed genesis core reads k 0.53 at the 5090's lock on the class
v4 draw (0.38 a node ahead); the full core, scaled by the synthesised full-to-genesis ratio (1.11) until its routed
row lands, 8.66 pJ at ASAP7, 6.06 at N5 and 4.36 at N3, k 0.59 node-for-node and 0.42 a node ahead.
The full core (every bank entry), 8 lanes, routed: 227,069 cells, 3 macros:
| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| mf8full | power | mix | 227069 | 0.00136 | 0.000395 | 7.6 (0.9 / 5.2) | 0.99 / 1.76 / 3.52 | 9.36 (8.59 to 11.1) | 6.55 | 4.72 | 3.4 | 18.8 / 10.3 | 0.636 | 0.458 | 0.251 |
| mf8full | power | mixld | 227069 | 0.00138 | 0.000395 | 7.72 (0.9 / 5.2) | 0.986 / 1.75 / 3.5 | 9.47 (8.71 to 11.2) | 6.63 | 4.77 | 3.44 | 18.8 / 10.3 | 0.644 | 0.463 | 0.254 |
| mf8full | power | mix1 | 227069 | 0.00134 | 0.000395 | 7.45 (0.91 / 5) | 1.01 / 1.8 / 3.59 | 9.25 (8.47 to 11) | 6.48 | 4.66 | 3.36 | 18.8 / 10.3 | 0.629 | 0.453 | 0.248 |
| mf8full | power | mix2 | 227069 | 0.00124 | 0.000395 | 6.72 (0.9 / 4.3) | 1 / 1.78 / 3.56 | 8.5 (7.73 to 10.3) | 5.95 | 4.29 | 3.09 | 18.8 / 10.3 | 0.578 | 0.416 | 0.228 |
| mf8full | power | mixw4 | 227069 | 0.00136 | 0.000395 | 7.6 (0.9 / 5.1) | 0.977 / 1.73 / 3.47 | 9.34 (8.58 to 11.1) | 6.54 | 4.71 | 3.39 | 18.8 / 10.3 | 0.635 | 0.457 | 0.25 |
| mf8full | power | mix64 | 227069 | 0.00133 | 0.000395 | 7.38 (0.9 / 4.9) | 0.993 / 1.76 / 3.53 | 9.14 (8.37 to 10.9) | 6.4 | 4.61 | 3.32 | 18.8 / 10.3 | 0.621 | 0.447 | 0.245 |
| mf8full | power | add | 227069 | 0.00104 | 0.000396 | 5.17 (0.89 / 2.7) | 0.95 / 1.69 / 3.38 | 6.86 (6.12 to 8.54) | 4.8 | 3.46 | 2.49 | 11.3 / 6.2 | 0.774 | 0.557 | 0.306 |
| mf8full | power | sub | 227069 | 0.00104 | 0.000396 | 5.21 (0.89 / 2.7) | 0.95 / 1.69 / 3.38 | 6.9 (6.16 to 8.58) | 4.83 | 3.48 | 2.5 | 11.3 / 6.2 | 0.779 | 0.561 | 0.308 |
| mf8full | power | xor | 227069 | 0.00105 | 0.000396 | 5.24 (0.87 / 2.9) | 0.95 / 1.69 / 3.38 | 6.93 (6.19 to 8.62) | 4.85 | 3.49 | 2.51 | 11.3 / 6.2 | 0.782 | 0.563 | 0.309 |
| mf8full | power | or | 227069 | 0.000788 | 0.000397 | 3.31 (0.83 / 0.92) | 0.95 / 1.69 / 3.38 | 5 (4.26 to 6.68) | 3.5 | 2.52 | 1.81 | 11.3 / 6.2 | 0.564 | 0.406 | 0.223 |
| mf8full | power | rotl | 227069 | 0.00107 | 0.000396 | 5.43 (0.87 / 3) | 0.95 / 1.69 / 3.38 | 7.12 (6.38 to 8.81) | 4.99 | 3.59 | 2.58 | 11.3 / 6.2 | 0.804 | 0.579 | 0.318 |
| mf8full | power | rotr | 227069 | 0.00104 | 0.000396 | 5.17 (0.89 / 2.7) | 0.95 / 1.69 / 3.38 | 6.86 (6.12 to 8.54) | 4.8 | 3.46 | 2.49 | 11.3 / 6.2 | 0.774 | 0.557 | 0.306 |
| mf8full | power | mul | 227069 | 0.000932 | 0.000396 | 4.39 (0.84 / 1.9) | 0.95 / 1.69 / 3.38 | 6.08 (5.34 to 7.77) | 4.25 | 3.06 | 2.21 | 13.9 / 8.3 | 0.513 | 0.369 | 0.22 |
| mf8full | power | mulhi | 227069 | 0.000906 | 0.000396 | 4.2 (0.84 / 1.8) | 0.95 / 1.69 / 3.38 | 5.89 (5.15 to 7.57) | 4.12 | 2.97 | 2.14 | 39.6 / 21 | 0.196 | 0.141 | 0.0749 |
| mf8full | power | mad | 227069 | 0.00141 | 0.000395 | 7.95 (0.94 / 5.5) | 1.2 / 2.12 / 4.25 | 10.1 (9.15 to 12.2) | 7.05 | 5.08 | 3.65 | 13.9 / 8.3 | 0.849 | 0.611 | 0.365 |
| mf8full | power | shfl | 227069 | 0.00099 | 0.000397 | 4.82 (0.87 / 2.4) | 0.95 / 1.69 / 3.38 | 6.51 (5.77 to 8.2) | 4.56 | 3.28 | 2.36 | 55.8 / 29.4 | 0.155 | 0.112 | 0.0588 |
| mf8full | power | load | 227069 | 0.00125 | 0.000395 | 6.81 (0.87 / 4.4) | 0.95 / 1.69 / 3.38 | 8.49 (7.76 to 10.2) | 5.95 | 4.28 | 3.08 | 13.9 / 8.3 | 0.716 | 0.516 | 0.308 |
| mf8full | power | fwd | 227069 | 0.00116 | 0.000395 | 6.12 (0.87 / 3.7) | 0.95 / 1.69 / 3.38 | 7.81 (7.07 to 9.5) | 5.47 | 3.94 | 2.83 | 13.9 / 8.3 | 0.659 | 0.474 | 0.283 |
| mf8full | power | prmt | 227069 | 0.000965 | 0.000397 | 4.63 (0.87 / 2.2) | 0.95 / 1.69 / 3.38 | 6.32 (5.58 to 8) | 4.42 | 3.18 | 2.29 | 22.3 / 11.5 | 0.384 | 0.277 | 0.143 |
| mf8full | power | lop3 | 227069 | 0.00104 | 0.000396 | 5.17 (0.94 / 2.7) | 1.2 / 2.12 / 4.25 | 7.29 (6.37 to 9.42) | 5.1 | 3.68 | 2.65 | 24.1 / 13 | 0.393 | 0.283 | 0.152 |
| mf8full | power | shfla | 227069 | 0.00105 | 0.000396 | 5.24 (0.94 / 2.7) | 1.2 / 2.12 / 4.25 | 7.37 (6.44 to 9.49) | 5.16 | 3.71 | 2.67 | 17.3 / 9.49 | 0.544 | 0.391 | 0.215 |
| mf8full | power | popc | 227069 | 0.000765 | 0.000397 | 3.13 (0.83 / 0.74) | 0.95 / 1.69 / 3.38 | 4.82 (4.08 to 6.51) | 3.37 | 2.43 | 1.75 | 17 / 9.3 | 0.363 | 0.261 | 0.143 |
| mf8full | power | clz | 227069 | 0.000759 | 0.000397 | 3.09 (0.83 / 0.79) | 0.95 / 1.69 / 3.38 | 4.78 (4.04 to 6.47) | 3.34 | 2.41 | 1.73 | 18.4 / 10.1 | 0.331 | 0.238 | 0.131 |
| mf8full | power | bfe | 227069 | 0.000779 | 0.000397 | 3.24 (0.83 / 0.92) | 0.95 / 1.69 / 3.38 | 4.93 (4.19 to 6.62) | 3.45 | 2.48 | 1.79 | 17.4 / 9.55 | 0.361 | 0.26 | 0.143 |
| mf8full | power | shl | 227069 | 0.000886 | 0.000397 | 4.04 (0.85 / 1.7) | 0.95 / 1.69 / 3.38 | 5.73 (4.99 to 7.41) | 4.01 | 2.89 | 2.08 | 8.48 / 4.65 | 0.862 | 0.621 | 0.341 |
| mf8full | power | shr | 227069 | 0.000866 | 0.000397 | 3.89 (0.84 / 1.5) | 0.95 / 1.69 / 3.38 | 5.58 (4.84 to 7.26) | 3.9 | 2.81 | 2.02 | 8.48 / 4.65 | 0.84 | 0.604 | 0.332 |
| mf8full | power | sel | 227069 | 0.000985 | 0.000396 | 4.79 (0.94 / 2.4) | 1.2 / 2.12 / 4.25 | 6.91 (5.98 to 9.03) | 4.84 | 3.48 | 2.51 | 11.3 / 6.2 | 0.78 | 0.562 | 0.308 |
| mf8full | power | andn | 227069 | 0.000897 | 0.000397 | 4.12 (0.85 / 1.7) | 0.95 / 1.69 / 3.38 | 5.81 (5.07 to 7.5) | 4.07 | 2.93 | 2.11 | 11.3 / 6.2 | 0.656 | 0.472 | 0.259 |
| mf8full | power | mm8 | 227069 | 0.00112 | 0.000396 | 5.78 (0.94 / 3.3) | 1.2 / 2.12 / 4.25 | 7.9 (6.97 to 10) | 5.53 | 3.98 | 2.87 | 16.4 / 8.8 | 0.628 | 0.452 | 0.243 |
Reading. (1) Placement and the clock tree add 32 percent to the genesis core's class v4 draw (7.78 pJ per lane-op
routed against 5.91 synthesised at ASAP7; 5.44 against 4.13 at N5) and 42 percent to the full core's (9.36 against
6.58; 6.55 against 4.61 at N5), inside and at the top of the k lane's +20 to +40 percent expectation; the
sequential term barely moves (0.90 against 0.88 pJ) and the combinational and clock term carries the increase (5.2
against 3.6 on the full core), which is wire capacitance on the macro output nets, the wider result mux and the lane
reduction. (2) **The bank's cost after placement is 20 percent of the energy on the class v4 draw (9.36 against
7.78) and 41 percent of the cells**, against 11 and 43 percent at synthesis: the idle units' wires are not free. (3)
Node-for-node the placed full core reads k 0.64 at the 5090's lock on the class v4 draw (0.46 a node ahead); the
genesis core 0.53 and 0.38. Per family at the lock, placed, k N5 / N3: the ARX group 0.77 to 0.80 / 0.56 to 0.58,
or 0.56 / 0.41, mul 0.51 / 0.37, mulhi 0.20 / 0.14, mad 0.85 / 0.61, shfl 0.16 / 0.11, the load with the fold 0.72
/ 0.52, prmt 0.38 / 0.28, lop3 0.39 / 0.28, shfla 0.54 / 0.39, popc 0.36 / 0.26, clz 0.33 / 0.24, bfe 0.36 / 0.26,
shl and shr 0.84 to 0.86 / 0.60 to 0.62, sel 0.78 / 0.56, andn 0.66 / 0.47, mm8 0.63 / 0.45 per dp4a (1.38 pJ per
MAC at N5 against the card's 2.2). The per-family order is unchanged by placement. (4) The live-family mixes stay
under the class v4 draw (shfla and mm8 live -1.2 percent, all eight live -9.2 percent, the 64 shape -2.4, W = 4
-0.2, one load in 16 +1.2), with one correction the testbench cannot make: an mm8 instruction on the card is a
32-MAC tile and on this core eight dp4a ops, so at 4 points of 79 the chip executes 35 percent more lane-ops per
hash and pays +29 percent of shadow energy, and the card's premium rises by the same +29 percent (its tile at 70 pJ
per lane against the draw's 10.3): the ratio does not move, the lane count does (section 14, row 1).
The 32-lane genesis core (synthesis only; four window macros, one imem pair shared by 32 lanes): 261,440 cells,
5.29 pJ per lane-op at ASAP7 (4.65 to 6.77) on the class v4 draw, 3.70 at N5, k 0.36 at the lock: 10 percent
under the 8-lane core, the imem and the sequencer amortised over four times the lanes. The 32-lane full core
(synthesis) and its placed form are in the flow (adv-a and adv-f at 17:4x BST) and land as a delta.
The board rows of section 6 restated at the placed energy (the full core at 6.06 pJ per lane-op at N5, band 5.45
to 7.43; 4.36 at N3): the complete GDDR7 machine 1.44 microjoules per hash, **1.6x the 5090 at its lock per joule
node-for-node (1.4x to 1.8x), 1.9x a node ahead (1.5x to 2.1x)**, 1.4x and 1.7x against the 5080 at its lock,
2.5x and 2.9x against the cohort card, USD 4.84 per MH/s (3.3x per dollar at MSRP) unchanged; the N2 SRAM die 2.4x
and 3.3x; the stored-half hybrid of section 13 about 2.4x and 2.9x (the synthesised 2.69x and 3.33x less the
placement tenth). Placement moves every per-joule ratio down by about
a tenth and no per-dollar ratio, and the review's rule that synthesis is never a lower bound is read here the
other way too: the placed figure is the row, the synthesised one the floor it sat above by a third.
The board rows of section 6 restated at the placed energy (the full core at 6.55 pJ per lane-op at N5, band 6.01
to 7.77; 4.72 at N3): the complete GDDR7 machine 1.53 microjoules per hash, **1.5x the 5090 at its lock per joule
node-for-node (1.3x to 1.7x), 1.8x a node ahead (1.5x to 2.0x)**, 1.3x and 1.6x against the 5080 at its lock, 2.4x
and 2.8x against the cohort card, USD 4.84 per MH/s (3.3x per dollar at MSRP) unchanged; HBM3 one stack 1.7x and
2.1x at USD 9.91 per MH/s; the N2 SRAM die 2.3x and 3.1x at USD 2.94 and 2.14 per MH/s; the stored-half hybrid of
section 13 1.9x and 2.4x at the mean hit rate (2.1x and 2.65x on the p98 program). Placement moves every per-joule
ratio down by a sixth and no per-dollar ratio; the review's rule that synthesis is never a lower bound is read here
the other way too: the placed figure is the row, the synthesised one the floor it sat above by 42 percent. The
three-year cost per TH: the GDDR7 machine 0.085 USD against the 5090's 0.220 and the cohort's 0.270 (2.6x to 3.2x).
## 6. The board: joules per valid hash and USD per sustained MH/s for the complete machine
@ -399,7 +459,7 @@ program for an atom or a shape), from the rows of sections 3 and 4; the GPU's ow
| R2 perm live | a byte selector | yes | -1.9 percent | 1.30 | 0 |
| R3 popc and clz live | a popcount tree, a priority encoder | yes | -2.5 percent | 1.50 and 1.63 | 0 |
| R4 to R7 (bfe, shl and shr, sel, andn) live | a shifter, a mask, a select, an and-not | yes | -2.1 to -1.4 percent | 0.75 to 1.54 | 0 |
| R8 mm8 live | a u8 dot4 per lane (8 chip ops per card tile) | yes | -0.8 percent (3.9 pJ per dp4a at N5, 1.0 pJ per MAC, against the card's 2.2 pJ per MAC at the lock) | the tile: 2.43 the add step per card instruction (32 MACs) | 0 (the chip's MAC is cheaper than the card's by 4x to 30x on the public figures; this is the card's loss, not the chip's) |
| R8 mm8 live | a u8 dot4 per lane (8 chip ops per card tile) | yes | +29 percent of shadow energy and +35 percent of lane-ops per hash at 4 points (eight dp4a per card tile at 5.5 pJ placed, 1.38 pJ per MAC against the card's 2.2); the lane count, not the ratio, pays: USD 27 of N5 silicon per board | the tile: +29 percent of the premium (70 pJ per lane-instruction against the draw's 10.3) | 0 (both sides pay +29 percent; k 0.63 on the tile equals the draw's 0.64) |
| the op-mix band draw (B = 4 on injecting families) | nothing: the program | yes | within the family rows above | within 11 percent per instruction (shadow-k 6.2) | 0 |
| the fold constants draw | five registers | yes | 0 | 0 | 0 |
| the block shape draw (64, 128, 256) | the program-length register; 256 words of imem | yes | -3.7 percent at 64 (the imem term; 0 at 128 and 256) | 64 ran 2.5 to 3.5 percent faster than 256 on the 5090 and the M5 Max (measured) | 0 |
@ -431,17 +491,17 @@ and this is its price: under 1 W and USD 15 per machine.
Three separate things, each with its own number and its own holder.
**Energy resistance** is a property of the memory system and the shadow, not of the calendar. Against the
adversary's re-optimised chip the complete GDDR7 machine reads 1.8x the 5090 at its lock per joule node-for-node
(1.5x to 2.1x), 2.1x a node ahead; 1.6x and 1.9x against the 5080 at its lock; 2.8x and 3.3x against the cohort
card; the N2 SRAM die 3.1x and 4.2x. The bank moves these by 5 percent in the honest side's favour (the chip's
shadow energy +11 percent for carrying 18 families instead of 10) and no more; the register window moves them by
0.03 to 0.13 of k and no more. What holds the per-joule number is the shadow's size on the card's own operating
adversary's re-optimised chip, placed and routed, the complete GDDR7 machine reads 1.5x the 5090 at its lock per
joule node-for-node (1.3x to 1.7x), 1.8x a node ahead (1.5x to 2.0x); 1.3x and 1.6x against the 5080 at its lock;
2.4x and 2.8x against the cohort card; the stored-half hybrid 1.9x and 2.4x; the N2 SRAM die 2.3x and 3.1x. The
bank moves these by a sixth in the honest side's favour (the chip's shadow energy +20 percent placed for carrying
18 families instead of 10) and no more; the register window moves them by 0.03 to 0.13 of k and no more. What holds the per-joule number is the shadow's size on the card's own operating
point and the memory system's activate ceiling; what would move it is a memory arrangement the chip cannot buy, and
section 6 says there is none: the chip's cheapest memory is the card's own.
**Economic resistance** is capex per sustained MH/s and the project cost against the chain's revenue. The chip's
machine costs USD 4.84 per MH/s against the card's 16 to 18, so over a 3-year life it mines at 0.080 USD per TH
against 0.22 to 0.27: 2.7x to 3.4x, of which electricity is the smaller half. The calendar does not shorten that
machine costs USD 4.84 per MH/s against the card's 16 to 18 (the stored-half hybrid USD 2.80), so over a 3-year
life it mines at 0.085 USD per TH against 0.22 to 0.27: 2.6x to 3.2x, of which electricity is the smaller half. The calendar does not shorten that
life: no transition in the bank retires the chip, so the 3-year stress life holds in full and the withdrawn headline
("dies within an epoch, under 1x over its life") stays withdrawn. What holds the economics is the project (USD 20 M
to 75 M for the GDDR7-board chip, floor lane 5; USD 100 M to 500 M for the SRAM die, claimed) against the miner
@ -449,8 +509,8 @@ revenue the chain pays, which is the profitability surface the review asks for a
dataset floor is a ticket on the SRAM die only (USD 1,000 per step) and nothing on the DRAM board (USD 60 per step).
**Response capability** is what the bank actually buys: the chain can change its object every 180 days without a
release, and a chip that carries the bank follows by firmware at a per-transition loss of -3.7 to +0.7 percent of
its shadow energy (zero credited). A family outside the bank costs that chip an emulation penalty of 2 to 6 percent
release, and a chip that carries the bank follows by firmware at a per-transition loss of -9 to +1 percent of
its shadow energy on every entry but the tile, where both sides pay +29 percent (zero credited). A family outside the bank costs that chip an emulation penalty of 2 to 6 percent
of the shadow at 4 points, which every card that predates the release pays too; a structural change (a new read
atom, a new derivation) costs the chip a host update and the cards a release. So the bank is a response channel
whose value per event is the per-transition loss column, not a chip retirement; it is worth keeping for what it is
@ -605,36 +665,39 @@ The recommendation, five lines:
design change; the three terms that still need physical design before they are bounds (the HBM activate
ceiling, the SRAM die's wire, the on-package hop) are named in section 11 and owed.
## 13. D2(b): memory sharing, recomputation, data-local execution and selective participation, the cheapest combination priced (Igneum 2.0, 17:3x BST)
## 13. D2(b): memory sharing, recomputation, data-local execution and selective participation, the cheapest combination priced (Igneum 2.0; first run 17:3x BST on the synthesised rows, this run 18:0x BST on the placed rows)
The harness `tools/chip-model/mf/flow/d2b.py` (run on a rented host; build-2's pool held no free cores at 17:2x BST)
takes the mf core's rows at N5 node-for-node (the measured drawn-mix row as the anchor, the unit rows for the
per-draw movement), the card's measured 2.33 microjoules at its lock with its 0.652 premium moved by the draw's
measured per-op ratios, chip-model-v3's memory figures, adv-cache-2's measured window-layer hit rates and the layer
1 band, and prices every combination on the whole machine in the served convention (the honest 5090 at its lock
over the chip machine on its own GDDR7 board). Nothing here is a chip measurement; the bands are the memory and
SRAM bands of section 6.
The harness `tools/chip-model/mf/flow/d2b.py` (run on a rented host; build-2's pool held no free cores) takes the mf
core's placed rows at N5 node-for-node (the measured drawn-mix row as the anchor, the unit rows for the per-draw
movement), the card's measured 2.33 microjoules at its lock with its 0.652 premium moved by the draw's measured
per-op ratios, chip-model-v3's memory figures, the window-layer hit rates (the hash lane's reconcile of 17:11 BST:
the hottest half of the items serves 58 percent of the reads on the mean program, 72 percent on Devnet 3's p98
program; adv-cache-2's figure was the tail) and the layer 1 band, and prices every combination on the whole machine
in the served convention (the honest 5090 at its lock over the chip machine on its own GDDR7 board). Nothing here
is a chip measurement; the bands are the memory and SRAM bands of section 6.
D2(b) on the mf core at N5; the served convention: the honest 5090 at its lock over the chip machine on its own GDDR7 board
D2(b) on the mf core at N5 (power: placed and routed); the served convention: the honest 5090 at its lock over the chip machine on its own GDDR7 board
| Quantity (2,000 era draws under the layer 1 band, two reserve families live per epoch) | Nominal | Band |
|---|---|---|
| The whole-machine ratio, mean over epochs | 1.81x | 1.51 to 2.05 |
| 5th / 50th / 95th percentile over epochs | 1.73 / 1.81 / 1.89 | spread 9.0 percent |
| A specialist mining only its best 50 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.85x / 0.50 of full-time revenue | gain in ratio 2.1 percent, at a loss of 50 percent of revenue |
| A specialist mining only its best 25 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.87x / 0.25 of full-time revenue | gain in ratio 3.4 percent, at a loss of 75 percent of revenue |
| A specialist mining only its best 10 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.89x / 0.10 of full-time revenue | gain in ratio 4.7 percent, at a loss of 90 percent of revenue |
| The whole-machine ratio, mean over epochs | 1.51x | 1.30 to 1.68 |
| 5th / 50th / 95th percentile over epochs | 1.45 / 1.51 / 1.58 | spread 9.1 percent |
| A specialist mining only its best 50 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.55x / 0.50 of full-time revenue | gain in ratio 2.1 percent, at a loss of 50 percent of revenue |
| A specialist mining only its best 25 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.57x / 0.25 of full-time revenue | gain in ratio 3.4 percent, at a loss of 75 percent of revenue |
| A specialist mining only its best 10 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.59x / 0.10 of full-time revenue | gain in ratio 4.7 percent, at a loss of 90 percent of revenue |
| Combination (the class v4 draw) | Whole-machine ratio vs the 5090 at its lock (served convention) | Band | Machine microjoules per hash | Sustained MH/s per board | Extra capex USD | USD per MH/s (board 661 + extra) |
|---|---|---|---|---|---|---|
| The baseline machine: the state in its lane, the data to it, the full dataset on the board (section 6) | 1.82x | 1.52 to 2.06 | 1.280 | 136 | 0 | 4.84 |
| hottest 25 percent of the items in SRAM beside the DRAM, serving 42% of reads (the window-layer distribution on measured programs; 0.5 GiB of SRAM, USD 125, claimed) | 2.23x | 1.84 to 2.51 | 1.045 | 236 | 125 | 3.33 |
| the same at a uniform store (25% of reads: the lower bound) | 2.02x | 1.68 to 2.28 | 1.154 | 182 | 125 | 4.32 |
| hottest 50 percent of the items in SRAM beside the DRAM, serving 72% of reads (the window-layer distribution on measured programs; 1.0 GiB of SRAM, USD 250, claimed) | 2.69x | 2.20 to 3.03 | 0.865 | 485 | 250 | 1.88 |
| the same at a uniform store (50% of reads: the lower bound) | 2.30x | 1.90 to 2.59 | 1.012 | 273 | 250 | 3.34 |
| hottest 75 percent of the items in SRAM beside the DRAM, serving 89% of reads (the window-layer distribution on measured programs; 1.5 GiB of SRAM, USD 375, claimed) | 3.10x | 2.49 to 3.49 | 0.753 | 1248 | 375 | 0.83 |
| the same at a uniform store (75% of reads: the lower bound) | 2.73x | 2.23 to 3.07 | 0.852 | 546 | 375 | 1.90 |
| hottest 50 percent in SRAM and the other 28% of items RECOMPUTED on the core (9,360 ops each) instead of read from DRAM: no DRAM at all (the rate core-bound on the baseline board's lanes) | 0.78x | 0.63 to 0.86 | 3.000 | 136 | 250 | 6.68 |
| The baseline machine: the state in its lane, the data to it, the full dataset on the board (section 6) | 1.52x | 1.31 to 1.69 | 1.528 | 136 | 0 | 4.84 |
| hottest 25 percent of the items in SRAM beside the DRAM, serving 34% of reads (the mean over programs, the hash lane's reconcile); 0.5 GiB of SRAM, USD 125, claimed | 1.73x | 1.48 to 1.91 | 1.345 | 207 | 125 | 3.80 |
| the same at a uniform store (25% of reads: the lower bound) | 1.66x | 1.42 to 1.83 | 1.403 | 182 | 125 | 4.32 |
| hottest 50 percent of the items in SRAM beside the DRAM, serving 58% of reads (the mean over programs, the hash lane's reconcile); 1.0 GiB of SRAM, USD 250, claimed | 1.93x | 1.65 to 2.12 | 1.206 | 326 | 250 | 2.80 |
| the same at a uniform store (50% of reads: the lower bound) | 1.85x | 1.58 to 2.03 | 1.260 | 273 | 250 | 3.34 |
| hottest 50 percent of the items in SRAM beside the DRAM, serving 72% of reads (the p98 program, Devnet 3); 1.0 GiB of SRAM, USD 250, claimed | 2.09x | 1.78 to 2.29 | 1.113 | 485 | 250 | 1.88 |
| the same at a uniform store (50% of reads: the lower bound) | 1.85x | 1.58 to 2.03 | 1.260 | 273 | 250 | 3.34 |
| hottest 75 percent of the items in SRAM beside the DRAM, serving 72% of reads (the mean over programs, the hash lane's reconcile); 1.5 GiB of SRAM, USD 375, claimed | 2.08x | 1.77 to 2.27 | 1.122 | 487 | 375 | 2.13 |
| the same at a uniform store (75% of reads: the lower bound) | 2.12x | 1.80 to 2.32 | 1.101 | 546 | 375 | 1.90 |
| hottest 50 percent in SRAM and the other 42% of items RECOMPUTED on the core (9,360 ops each) instead of read from DRAM: no DRAM at all (the rate core-bound on the baseline board's lanes) | 0.43x | 0.37 to 0.47 | 5.410 | 136 | 250 | 6.68 |
| Form | Bits moved per dependent read | Against the baseline (80 bits to the lane) | Energy per read on a die at 1.3 pJ per bit / on a package at 0.5 / across boards at 5 to 10 |
|---|---|---|---|
@ -647,9 +710,11 @@ The served bracket (the complete GDDR7 machine 1.8x node-for-node, 1.5 to 2.1; 2
Reading, and whether the served bracket moves. (1) It moves, by one combination: a hot-item SRAM beside the DRAM.
The board is activate-bound, so every read an SRAM serves is a read the DRAM does not, and the same 16 devices run
more hashes: the hottest half of the items (1 GiB of N2 SRAM, USD 250 claimed, the per-program hot set refilled
from the board's own DRAM in about 2 ms per epoch) serves 72 percent of the reads on the measured window-layer
distribution and reads 2.69x per joule (2.20 to 3.03) and USD 1.88 per MH/s against the baseline's 1.82x and 4.84;
with no hot-set knowledge at all (a uniform half) 2.30x and USD 3.34; a node ahead 3.33x. The record's
from the board's own DRAM in about 2 ms per epoch) serves 58 percent of the reads on the mean program and reads
1.93x per joule (1.65 to 2.12) and USD 2.80 per MH/s against the baseline's 1.52x and 4.84 on the placed rows (on
the p98 program, 72 percent of reads, 2.09x and USD 1.88); with no hot-set knowledge at all (a uniform half) 1.85x
and USD 3.34; a node ahead 2.40x (the p98 2.65x). The first run of this section on the synthesised rows read 1.82x,
2.69x and 3.33x; the placed rows take a sixth off each. The record's
partial-store curve (chip-model-v3 5.4) priced the un-stored items as recomputed, which loses (0.78x here); stored
in SRAM they win, and the DRAM-board chip's cheapest form is this hybrid. The dataset floor raises its SRAM ticket
(USD 690 for the hot half at 5.5 GiB, USD 2.9 per MH/s) and not its ratio: the hit rate per fraction is scale-free.
@ -657,10 +722,36 @@ in SRAM they win, and the DRAM-board chip's cheapest form is this hybrid. The da
medium (section 11). (3) Memory sharing is the baseline: N engines over one dataset are the lanes in flight the
board already has; the set-up's share is 1.4e-12 J per hash. (4) Selective participation moves nothing: across
2,000 era draws under the layer 1 band with two reserve families live per epoch the whole-machine ratio spreads 9
percent (5th to 95th percentile 1.73x to 1.89x); a specialist mining only its best half, quarter or tenth of epochs
percent (5th to 95th percentile 1.45x to 1.58x on the placed rows); a specialist mining only its best half, quarter or tenth of epochs
gains 2, 3 or 5 percent of ratio and loses 50, 75 or 90 percent of revenue, because difficulty follows the fleet it
joins and it earns in proportion to the time it mines; downtime and re-entry cost it the DAA window's lag each
way and buy it nothing. The GPU-cost budget of 10 percent at the lock (set before these results) is untouched:
no honest energy is spent in this section. Served sentence for the bracket: against a chip on the same node the
honest floor is the stored-half hybrid at 2.3x to 2.7x per joule and 5x to 9x per dollar, not the DRAM board's
1.8x and 3.3x; the lever on it is the dataset floor as a capex ticket.
honest floor is the stored-half hybrid at 1.9x to 2.1x per joule and 6x to 9x per dollar (1.65x to 2.3x on the
band), not the DRAM board's 1.5x and 3.3x; the lever on it is the dataset floor as a capex ticket.
## 14. The adversary's transition matrix (Igneum 2.0 D3, the complete plan's page 12; 18:0x BST)
Each row a required result on the evidence standard (SRAM macros with ports and area from FakeRAM2.0, the time-
multiplexed single port, routed wiring and clock tree, the memory interface and board from chip-model-v3's rows,
switching activity from the gate-level VCD; same-node N5 and advanced-node N3 kept apart; a band on every row;
measured where the placed cores give it, modelled and labelled where they do not). The lifetime rule: the
programmable design survives; a retirement credit only where the cheapest adaptation loses competitiveness.
| Adaptation | Throughput | Energy | Cost | Survival through the schedule | Band and label | Credit |
|---|---|---|---|---|---|---|
| 1. Firmware update (a family transition inside the bank) | unchanged for 17 of 18 entries (one lane-op per slot each); the tile family at 4 points raises lane-ops per hash by 35 percent (eight dp4a per card tile), bought with USD 27 of lanes | -9 to +1 percent of the shadow energy per transition on the placed rows (all eight reserve families live -9.2, shfla and mm8 live -1.2, the 64 shape -2.4, W = 4 -0.2, one load in 16 +1.2); the tile +29 percent, the card's premium +29 percent too | USD 0 of silicon; the era's registers and the program | yes: every entry is in the die at genesis | measured on the placed 8-lane core (the SRAM term 0.95 to 3.4 pJ); the same-node ratio 1.5x (1.3 to 1.7) before and after every transition within 0.1x | 0 |
| 2. Emulation (a family outside the 18, the bank-refresh case) | the emulated op is 3 to 10 lane-ops (the cards' measured penalties for the reserve families emulated: 1.13x to 2.4x per op; a chip's sequence the same shape) at 4 points of 79: +10 to +45 percent of lane-ops, bought with lanes | +5 to +25 percent of the shadow energy while the family is live; the same-node ratio 1.5x to 1.3x to 1.45x; slower, not unprofitable (USD 4.84 to about 6 per MH/s) | USD 0 of silicon until a revision (row 6) | yes | modelled from the measured card penalties and the placed per-family rows; approximate | 0 (the cheapest adaptation keeps 2.5x to 3x per dollar) |
| 3. Memory expansion or over-provisioning | unchanged (the activate rate per device is the bound; more devices add rate one for one) | unchanged per read | the 16-device 2 GB board (32 GB) already holds the 11.5 GiB step: USD 0 through year 4; 24 Gb devices USD 1,000 of memory for 48 GB (USD 9.9 per MH/s, once); the SRAM die USD 1,000 per step; the stored-half hybrid USD 250 per GiB of hot half (USD 690 at 5.5 GiB) | yes on every arrangement; the schedule is a ticket on SRAM only | modelled on the record's memory rows and the device prices (claimed) | 0 |
| 4. Companion CPU, GPU or FPGA | the host serves 100 machines (one full node, the state-derived dataset's keeper); a GPU beside the chips proves (SP1 CUDA) and does not mine; an FPGA is 0.3x to 0.4x of the 5090 per watt on the measured HBM2 row and adds nothing | +0.85 W per machine for the host | +USD 15 per machine for the host; a 5090 prover per shard rate the operator chooses (the proving line is the GPU line, unchanged) | yes | the host figures approximate; the FPGA row the Horizon lane's measured public row | 0: the hybrid mining-and-proving operator mines at the chip's economics and proves at the GPU's; the 80/20 split caps the proving share |
| 5. Favourable-period mining | the ratio spreads 9 percent across 2,000 era draws under the layer 1 band (5th to 95th percentile 1.45x to 1.58x same-node); mining only the best half, quarter or tenth of epochs gains 2, 3 or 5 percent of ratio | the same | revenue 50, 75 or 90 percent lower; re-entry costs the DAA window's lag each way; difficulty follows the fleet it joins | yes | modelled on the placed rows and the band; the draws uniform within the band (approximate) | 0 (nothing to gain) |
| 6. Silicon revision (a new family the firmware cannot carry economically) | a native unit restores row 1's throughput | row 1's energy | a metal-only respin USD 1 M to 3 M and 3 to 4 months, a full respin USD 10 M to 20 M of masks and 6 to 9 months at N5 (claimed, the record's project figures); the whole core design reused | yes: a bank refresh takes effect at family epoch n + 2, 360 days after its tally, so the revision ships before the family is live | approximate | 0 (the delay fits inside the bank-refresh notice) |
Reading: no row loses competitiveness; the cheapest adaptation is firmware for every entry in the bank (row 1), a
sequence for an entry outside it (row 2, 1.3x to 1.45x same-node while live), and a respin only if the chain adds a
family the sequence cannot carry, which it has 360 days' notice of. The stress life of three years holds in full.
PENDING with clocks: the 32-lane full core's synthesis row (adv-a, in ABC at 18:0x BST; by 19:30) and its placed
row (adv-f, at floorplan; by 21:00); the k lane's crossbar, scratch and tile rows (its pod); a real memory
compiler's figure for the two macros (owed, no clock: FakeRAM gives area and pins only); the per-family rows on
the 32-lane placed core (by 21:30 if adv-f lands, else the next pass).