From 2b7ad8fd27eaa186ef578ad7fe0d906f5a6499ce Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 16:09:55 +0000 Subject: [PATCH] multi-family adversary: the placed full-core rows, the board and D2(b) at the placed energy, the transition matrix, the one page Co-Authored-By: Claude Fable 5.1 Documents-only replay of 7e715b50b (10b901a46934c24ce9b44c74bc6fd4db7c40d2b0) for the box mirror master --- .../class-v6/multi-family-adversary.md | 205 +++++++++++++----- 1 file changed, 148 insertions(+), 57 deletions(-) diff --git a/docs/analysis/class-v6/multi-family-adversary.md b/docs/analysis/class-v6/multi-family-adversary.md index be5ab49b1..5676e3aec 100644 --- a/docs/analysis/class-v6/multi-family-adversary.md +++ b/docs/analysis/class-v6/multi-family-adversary.md @@ -19,7 +19,23 @@ population, Apple reported and not headlined; every defence is scored after the ## 0. One page -(filled from the rows below at each cut; section 8 carries the three-part statement) +1. The calendar changes the chip's firmware, not its parts. One programmable core carries every bank entry: 18 op + families, the 3 atoms as programs, the fold with its constants in registers, the 3 shapes, both read widths, the + 64-register window in an SRAM macro. Carrying the bank costs it 20 percent of its energy on the class v4 draw + (placed and routed) and 41 percent of its cells; no transition needs a part it lacks; the credited obsolescence + benefit is zero for every entry (section 7). +2. The rows (ASAP7, placed and routed, SPEF, gate-level VCD; the SRAM term modelled, bands shown; node factors + claimed): the full core 9.36 pJ per lane-op at ASAP7 on the class v4 draw (8.6 to 11.1), 6.55 at N5, 4.72 at N3; + k 0.64 node-for-node and 0.46 a node ahead at the 5090's lock; the reserve families k 0.33 to 0.86 placed; the + genesis-only core 7.78 (k 0.53 / 0.38); the 32-lane genesis core 10 percent under (synthesis). +3. The complete machine (memory, controller, host share, power train, cooling) on the card's own GDDR7: 1.5x the + 5090 at its lock per joule node-for-node (1.3x to 1.7x), 1.8x a node ahead (1.5x to 2.0x); 3.3x per dollar; HBM3 + buys it nothing; the N2 SRAM die 2.3x and 3.1x; the stored-half hybrid (section 13) 1.9x and 2.4x at USD 2.80 per + MH/s. Capex per MH/s is the larger half of the chip's edge; the project cost is its hold. +4. Data-local execution cannot pay (the live state is 26x the read, section 11); memory sharing is the baseline; + recomputation loses; selective participation gains 2 to 5 percent for 50 to 90 percent of revenue (section 13). +5. Class v7 should take the epoch-defined bounded dataset with a published support horizon and keep the floor + schedule and the bank; the bank is served as response capability, never as resistance (sections 8 and 12). ## 1. The design: what the adversary builds @@ -258,8 +274,7 @@ and detailed placement, CTS, global and detailed routing, OpenRCX parasitics; Op SPEF under the gate-level VCD; the clock at 12,000 ps with leakage restated at the 1,500 ps equivalent, section 2.1; the final PDN connectivity check reports two macro power pins unconnected, a floorplan artefact of the FakeRAM pin geometry that does not touch the netlist, the parasitics or the power figure). The 8-lane genesis core (the -comparator) first; the 8-lane full core's routed power run is in its VCD sweep (11 of 58 tags at 17:4x BST) and -replaces the scaled estimate below when it lands. +comparator) first, then the 8-lane full core (routed, 58 tags, 18:0x BST). The genesis core, 8 lanes, routed: 161,504 cells (the synthesis 66,973: placement adds fill, buffers and the clock tree), 2 macros (the base core's instruction word fits 34 bits, so Yosys removed the second imem macro): @@ -280,28 +295,73 @@ tree), 2 macros (the base core's instruction word fits 34 bits, so Yosys removed | mf8base | power | shfl | 161504 | 0.000658 | 0.000264 | 3.2 (0.89 / 1.1) | 0.95 / 1.69 / 3.38 | 4.89 (4.15 to 6.57) | 3.42 | 2.46 | 1.77 | 55.8 / 29.4 | 0.116 | 0.0838 | 0.0441 | | mf8base | power | load | 161504 | 0.000926 | 0.000263 | 5.22 (0.89 / 3.1) | 0.95 / 1.69 / 3.38 | 6.91 (6.17 to 8.59) | 4.83 | 3.48 | 2.51 | 13.9 / 8.3 | 0.583 | 0.419 | 0.25 | -Reading. Placement and the clock tree add 32 percent to the class v4 draw on the genesis core (7.78 pJ per lane-op -routed against 5.91 synthesised at ASAP7; 5.44 against 4.13 at N5), inside the k lane's +20 to +40 percent -expectation; the sequential term barely moves (0.90 against 0.88 pJ) and the combinational and clock term carries -the increase (3.9 against 3.0), which is wire capacitance on the macro output nets and the result reduction. The -per-family order is unchanged. Node-for-node the placed genesis core reads k 0.53 at the 5090's lock on the class -v4 draw (0.38 a node ahead); the full core, scaled by the synthesised full-to-genesis ratio (1.11) until its routed -row lands, 8.66 pJ at ASAP7, 6.06 at N5 and 4.36 at N3, k 0.59 node-for-node and 0.42 a node ahead. +The full core (every bank entry), 8 lanes, routed: 227,069 cells, 3 macros: + +| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked | +|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| +| mf8full | power | mix | 227069 | 0.00136 | 0.000395 | 7.6 (0.9 / 5.2) | 0.99 / 1.76 / 3.52 | 9.36 (8.59 to 11.1) | 6.55 | 4.72 | 3.4 | 18.8 / 10.3 | 0.636 | 0.458 | 0.251 | +| mf8full | power | mixld | 227069 | 0.00138 | 0.000395 | 7.72 (0.9 / 5.2) | 0.986 / 1.75 / 3.5 | 9.47 (8.71 to 11.2) | 6.63 | 4.77 | 3.44 | 18.8 / 10.3 | 0.644 | 0.463 | 0.254 | +| mf8full | power | mix1 | 227069 | 0.00134 | 0.000395 | 7.45 (0.91 / 5) | 1.01 / 1.8 / 3.59 | 9.25 (8.47 to 11) | 6.48 | 4.66 | 3.36 | 18.8 / 10.3 | 0.629 | 0.453 | 0.248 | +| mf8full | power | mix2 | 227069 | 0.00124 | 0.000395 | 6.72 (0.9 / 4.3) | 1 / 1.78 / 3.56 | 8.5 (7.73 to 10.3) | 5.95 | 4.29 | 3.09 | 18.8 / 10.3 | 0.578 | 0.416 | 0.228 | +| mf8full | power | mixw4 | 227069 | 0.00136 | 0.000395 | 7.6 (0.9 / 5.1) | 0.977 / 1.73 / 3.47 | 9.34 (8.58 to 11.1) | 6.54 | 4.71 | 3.39 | 18.8 / 10.3 | 0.635 | 0.457 | 0.25 | +| mf8full | power | mix64 | 227069 | 0.00133 | 0.000395 | 7.38 (0.9 / 4.9) | 0.993 / 1.76 / 3.53 | 9.14 (8.37 to 10.9) | 6.4 | 4.61 | 3.32 | 18.8 / 10.3 | 0.621 | 0.447 | 0.245 | +| mf8full | power | add | 227069 | 0.00104 | 0.000396 | 5.17 (0.89 / 2.7) | 0.95 / 1.69 / 3.38 | 6.86 (6.12 to 8.54) | 4.8 | 3.46 | 2.49 | 11.3 / 6.2 | 0.774 | 0.557 | 0.306 | +| mf8full | power | sub | 227069 | 0.00104 | 0.000396 | 5.21 (0.89 / 2.7) | 0.95 / 1.69 / 3.38 | 6.9 (6.16 to 8.58) | 4.83 | 3.48 | 2.5 | 11.3 / 6.2 | 0.779 | 0.561 | 0.308 | +| mf8full | power | xor | 227069 | 0.00105 | 0.000396 | 5.24 (0.87 / 2.9) | 0.95 / 1.69 / 3.38 | 6.93 (6.19 to 8.62) | 4.85 | 3.49 | 2.51 | 11.3 / 6.2 | 0.782 | 0.563 | 0.309 | +| mf8full | power | or | 227069 | 0.000788 | 0.000397 | 3.31 (0.83 / 0.92) | 0.95 / 1.69 / 3.38 | 5 (4.26 to 6.68) | 3.5 | 2.52 | 1.81 | 11.3 / 6.2 | 0.564 | 0.406 | 0.223 | +| mf8full | power | rotl | 227069 | 0.00107 | 0.000396 | 5.43 (0.87 / 3) | 0.95 / 1.69 / 3.38 | 7.12 (6.38 to 8.81) | 4.99 | 3.59 | 2.58 | 11.3 / 6.2 | 0.804 | 0.579 | 0.318 | +| mf8full | power | rotr | 227069 | 0.00104 | 0.000396 | 5.17 (0.89 / 2.7) | 0.95 / 1.69 / 3.38 | 6.86 (6.12 to 8.54) | 4.8 | 3.46 | 2.49 | 11.3 / 6.2 | 0.774 | 0.557 | 0.306 | +| mf8full | power | mul | 227069 | 0.000932 | 0.000396 | 4.39 (0.84 / 1.9) | 0.95 / 1.69 / 3.38 | 6.08 (5.34 to 7.77) | 4.25 | 3.06 | 2.21 | 13.9 / 8.3 | 0.513 | 0.369 | 0.22 | +| mf8full | power | mulhi | 227069 | 0.000906 | 0.000396 | 4.2 (0.84 / 1.8) | 0.95 / 1.69 / 3.38 | 5.89 (5.15 to 7.57) | 4.12 | 2.97 | 2.14 | 39.6 / 21 | 0.196 | 0.141 | 0.0749 | +| mf8full | power | mad | 227069 | 0.00141 | 0.000395 | 7.95 (0.94 / 5.5) | 1.2 / 2.12 / 4.25 | 10.1 (9.15 to 12.2) | 7.05 | 5.08 | 3.65 | 13.9 / 8.3 | 0.849 | 0.611 | 0.365 | +| mf8full | power | shfl | 227069 | 0.00099 | 0.000397 | 4.82 (0.87 / 2.4) | 0.95 / 1.69 / 3.38 | 6.51 (5.77 to 8.2) | 4.56 | 3.28 | 2.36 | 55.8 / 29.4 | 0.155 | 0.112 | 0.0588 | +| mf8full | power | load | 227069 | 0.00125 | 0.000395 | 6.81 (0.87 / 4.4) | 0.95 / 1.69 / 3.38 | 8.49 (7.76 to 10.2) | 5.95 | 4.28 | 3.08 | 13.9 / 8.3 | 0.716 | 0.516 | 0.308 | +| mf8full | power | fwd | 227069 | 0.00116 | 0.000395 | 6.12 (0.87 / 3.7) | 0.95 / 1.69 / 3.38 | 7.81 (7.07 to 9.5) | 5.47 | 3.94 | 2.83 | 13.9 / 8.3 | 0.659 | 0.474 | 0.283 | +| mf8full | power | prmt | 227069 | 0.000965 | 0.000397 | 4.63 (0.87 / 2.2) | 0.95 / 1.69 / 3.38 | 6.32 (5.58 to 8) | 4.42 | 3.18 | 2.29 | 22.3 / 11.5 | 0.384 | 0.277 | 0.143 | +| mf8full | power | lop3 | 227069 | 0.00104 | 0.000396 | 5.17 (0.94 / 2.7) | 1.2 / 2.12 / 4.25 | 7.29 (6.37 to 9.42) | 5.1 | 3.68 | 2.65 | 24.1 / 13 | 0.393 | 0.283 | 0.152 | +| mf8full | power | shfla | 227069 | 0.00105 | 0.000396 | 5.24 (0.94 / 2.7) | 1.2 / 2.12 / 4.25 | 7.37 (6.44 to 9.49) | 5.16 | 3.71 | 2.67 | 17.3 / 9.49 | 0.544 | 0.391 | 0.215 | +| mf8full | power | popc | 227069 | 0.000765 | 0.000397 | 3.13 (0.83 / 0.74) | 0.95 / 1.69 / 3.38 | 4.82 (4.08 to 6.51) | 3.37 | 2.43 | 1.75 | 17 / 9.3 | 0.363 | 0.261 | 0.143 | +| mf8full | power | clz | 227069 | 0.000759 | 0.000397 | 3.09 (0.83 / 0.79) | 0.95 / 1.69 / 3.38 | 4.78 (4.04 to 6.47) | 3.34 | 2.41 | 1.73 | 18.4 / 10.1 | 0.331 | 0.238 | 0.131 | +| mf8full | power | bfe | 227069 | 0.000779 | 0.000397 | 3.24 (0.83 / 0.92) | 0.95 / 1.69 / 3.38 | 4.93 (4.19 to 6.62) | 3.45 | 2.48 | 1.79 | 17.4 / 9.55 | 0.361 | 0.26 | 0.143 | +| mf8full | power | shl | 227069 | 0.000886 | 0.000397 | 4.04 (0.85 / 1.7) | 0.95 / 1.69 / 3.38 | 5.73 (4.99 to 7.41) | 4.01 | 2.89 | 2.08 | 8.48 / 4.65 | 0.862 | 0.621 | 0.341 | +| mf8full | power | shr | 227069 | 0.000866 | 0.000397 | 3.89 (0.84 / 1.5) | 0.95 / 1.69 / 3.38 | 5.58 (4.84 to 7.26) | 3.9 | 2.81 | 2.02 | 8.48 / 4.65 | 0.84 | 0.604 | 0.332 | +| mf8full | power | sel | 227069 | 0.000985 | 0.000396 | 4.79 (0.94 / 2.4) | 1.2 / 2.12 / 4.25 | 6.91 (5.98 to 9.03) | 4.84 | 3.48 | 2.51 | 11.3 / 6.2 | 0.78 | 0.562 | 0.308 | +| mf8full | power | andn | 227069 | 0.000897 | 0.000397 | 4.12 (0.85 / 1.7) | 0.95 / 1.69 / 3.38 | 5.81 (5.07 to 7.5) | 4.07 | 2.93 | 2.11 | 11.3 / 6.2 | 0.656 | 0.472 | 0.259 | +| mf8full | power | mm8 | 227069 | 0.00112 | 0.000396 | 5.78 (0.94 / 3.3) | 1.2 / 2.12 / 4.25 | 7.9 (6.97 to 10) | 5.53 | 3.98 | 2.87 | 16.4 / 8.8 | 0.628 | 0.452 | 0.243 | + +Reading. (1) Placement and the clock tree add 32 percent to the genesis core's class v4 draw (7.78 pJ per lane-op +routed against 5.91 synthesised at ASAP7; 5.44 against 4.13 at N5) and 42 percent to the full core's (9.36 against +6.58; 6.55 against 4.61 at N5), inside and at the top of the k lane's +20 to +40 percent expectation; the +sequential term barely moves (0.90 against 0.88 pJ) and the combinational and clock term carries the increase (5.2 +against 3.6 on the full core), which is wire capacitance on the macro output nets, the wider result mux and the lane +reduction. (2) **The bank's cost after placement is 20 percent of the energy on the class v4 draw (9.36 against +7.78) and 41 percent of the cells**, against 11 and 43 percent at synthesis: the idle units' wires are not free. (3) +Node-for-node the placed full core reads k 0.64 at the 5090's lock on the class v4 draw (0.46 a node ahead); the +genesis core 0.53 and 0.38. Per family at the lock, placed, k N5 / N3: the ARX group 0.77 to 0.80 / 0.56 to 0.58, +or 0.56 / 0.41, mul 0.51 / 0.37, mulhi 0.20 / 0.14, mad 0.85 / 0.61, shfl 0.16 / 0.11, the load with the fold 0.72 +/ 0.52, prmt 0.38 / 0.28, lop3 0.39 / 0.28, shfla 0.54 / 0.39, popc 0.36 / 0.26, clz 0.33 / 0.24, bfe 0.36 / 0.26, +shl and shr 0.84 to 0.86 / 0.60 to 0.62, sel 0.78 / 0.56, andn 0.66 / 0.47, mm8 0.63 / 0.45 per dp4a (1.38 pJ per +MAC at N5 against the card's 2.2). The per-family order is unchanged by placement. (4) The live-family mixes stay +under the class v4 draw (shfla and mm8 live -1.2 percent, all eight live -9.2 percent, the 64 shape -2.4, W = 4 +-0.2, one load in 16 +1.2), with one correction the testbench cannot make: an mm8 instruction on the card is a +32-MAC tile and on this core eight dp4a ops, so at 4 points of 79 the chip executes 35 percent more lane-ops per +hash and pays +29 percent of shadow energy, and the card's premium rises by the same +29 percent (its tile at 70 pJ +per lane against the draw's 10.3): the ratio does not move, the lane count does (section 14, row 1). The 32-lane genesis core (synthesis only; four window macros, one imem pair shared by 32 lanes): 261,440 cells, 5.29 pJ per lane-op at ASAP7 (4.65 to 6.77) on the class v4 draw, 3.70 at N5, k 0.36 at the lock: 10 percent under the 8-lane core, the imem and the sequencer amortised over four times the lanes. The 32-lane full core (synthesis) and its placed form are in the flow (adv-a and adv-f at 17:4x BST) and land as a delta. -The board rows of section 6 restated at the placed energy (the full core at 6.06 pJ per lane-op at N5, band 5.45 -to 7.43; 4.36 at N3): the complete GDDR7 machine 1.44 microjoules per hash, **1.6x the 5090 at its lock per joule -node-for-node (1.4x to 1.8x), 1.9x a node ahead (1.5x to 2.1x)**, 1.4x and 1.7x against the 5080 at its lock, -2.5x and 2.9x against the cohort card, USD 4.84 per MH/s (3.3x per dollar at MSRP) unchanged; the N2 SRAM die 2.4x -and 3.3x; the stored-half hybrid of section 13 about 2.4x and 2.9x (the synthesised 2.69x and 3.33x less the -placement tenth). Placement moves every per-joule ratio down by about -a tenth and no per-dollar ratio, and the review's rule that synthesis is never a lower bound is read here the -other way too: the placed figure is the row, the synthesised one the floor it sat above by a third. - +The board rows of section 6 restated at the placed energy (the full core at 6.55 pJ per lane-op at N5, band 6.01 +to 7.77; 4.72 at N3): the complete GDDR7 machine 1.53 microjoules per hash, **1.5x the 5090 at its lock per joule +node-for-node (1.3x to 1.7x), 1.8x a node ahead (1.5x to 2.0x)**, 1.3x and 1.6x against the 5080 at its lock, 2.4x +and 2.8x against the cohort card, USD 4.84 per MH/s (3.3x per dollar at MSRP) unchanged; HBM3 one stack 1.7x and +2.1x at USD 9.91 per MH/s; the N2 SRAM die 2.3x and 3.1x at USD 2.94 and 2.14 per MH/s; the stored-half hybrid of +section 13 1.9x and 2.4x at the mean hit rate (2.1x and 2.65x on the p98 program). Placement moves every per-joule +ratio down by a sixth and no per-dollar ratio; the review's rule that synthesis is never a lower bound is read here +the other way too: the placed figure is the row, the synthesised one the floor it sat above by 42 percent. The +three-year cost per TH: the GDDR7 machine 0.085 USD against the 5090's 0.220 and the cohort's 0.270 (2.6x to 3.2x). ## 6. The board: joules per valid hash and USD per sustained MH/s for the complete machine @@ -399,7 +459,7 @@ program for an atom or a shape), from the rows of sections 3 and 4; the GPU's ow | R2 perm live | a byte selector | yes | -1.9 percent | 1.30 | 0 | | R3 popc and clz live | a popcount tree, a priority encoder | yes | -2.5 percent | 1.50 and 1.63 | 0 | | R4 to R7 (bfe, shl and shr, sel, andn) live | a shifter, a mask, a select, an and-not | yes | -2.1 to -1.4 percent | 0.75 to 1.54 | 0 | -| R8 mm8 live | a u8 dot4 per lane (8 chip ops per card tile) | yes | -0.8 percent (3.9 pJ per dp4a at N5, 1.0 pJ per MAC, against the card's 2.2 pJ per MAC at the lock) | the tile: 2.43 the add step per card instruction (32 MACs) | 0 (the chip's MAC is cheaper than the card's by 4x to 30x on the public figures; this is the card's loss, not the chip's) | +| R8 mm8 live | a u8 dot4 per lane (8 chip ops per card tile) | yes | +29 percent of shadow energy and +35 percent of lane-ops per hash at 4 points (eight dp4a per card tile at 5.5 pJ placed, 1.38 pJ per MAC against the card's 2.2); the lane count, not the ratio, pays: USD 27 of N5 silicon per board | the tile: +29 percent of the premium (70 pJ per lane-instruction against the draw's 10.3) | 0 (both sides pay +29 percent; k 0.63 on the tile equals the draw's 0.64) | | the op-mix band draw (B = 4 on injecting families) | nothing: the program | yes | within the family rows above | within 11 percent per instruction (shadow-k 6.2) | 0 | | the fold constants draw | five registers | yes | 0 | 0 | 0 | | the block shape draw (64, 128, 256) | the program-length register; 256 words of imem | yes | -3.7 percent at 64 (the imem term; 0 at 128 and 256) | 64 ran 2.5 to 3.5 percent faster than 256 on the 5090 and the M5 Max (measured) | 0 | @@ -431,17 +491,17 @@ and this is its price: under 1 W and USD 15 per machine. Three separate things, each with its own number and its own holder. **Energy resistance** is a property of the memory system and the shadow, not of the calendar. Against the -adversary's re-optimised chip the complete GDDR7 machine reads 1.8x the 5090 at its lock per joule node-for-node -(1.5x to 2.1x), 2.1x a node ahead; 1.6x and 1.9x against the 5080 at its lock; 2.8x and 3.3x against the cohort -card; the N2 SRAM die 3.1x and 4.2x. The bank moves these by 5 percent in the honest side's favour (the chip's -shadow energy +11 percent for carrying 18 families instead of 10) and no more; the register window moves them by -0.03 to 0.13 of k and no more. What holds the per-joule number is the shadow's size on the card's own operating +adversary's re-optimised chip, placed and routed, the complete GDDR7 machine reads 1.5x the 5090 at its lock per +joule node-for-node (1.3x to 1.7x), 1.8x a node ahead (1.5x to 2.0x); 1.3x and 1.6x against the 5080 at its lock; +2.4x and 2.8x against the cohort card; the stored-half hybrid 1.9x and 2.4x; the N2 SRAM die 2.3x and 3.1x. The +bank moves these by a sixth in the honest side's favour (the chip's shadow energy +20 percent placed for carrying +18 families instead of 10) and no more; the register window moves them by 0.03 to 0.13 of k and no more. What holds the per-joule number is the shadow's size on the card's own operating point and the memory system's activate ceiling; what would move it is a memory arrangement the chip cannot buy, and section 6 says there is none: the chip's cheapest memory is the card's own. **Economic resistance** is capex per sustained MH/s and the project cost against the chain's revenue. The chip's -machine costs USD 4.84 per MH/s against the card's 16 to 18, so over a 3-year life it mines at 0.080 USD per TH -against 0.22 to 0.27: 2.7x to 3.4x, of which electricity is the smaller half. The calendar does not shorten that +machine costs USD 4.84 per MH/s against the card's 16 to 18 (the stored-half hybrid USD 2.80), so over a 3-year +life it mines at 0.085 USD per TH against 0.22 to 0.27: 2.6x to 3.2x, of which electricity is the smaller half. The calendar does not shorten that life: no transition in the bank retires the chip, so the 3-year stress life holds in full and the withdrawn headline ("dies within an epoch, under 1x over its life") stays withdrawn. What holds the economics is the project (USD 20 M to 75 M for the GDDR7-board chip, floor lane 5; USD 100 M to 500 M for the SRAM die, claimed) against the miner @@ -449,8 +509,8 @@ revenue the chain pays, which is the profitability surface the review asks for a dataset floor is a ticket on the SRAM die only (USD 1,000 per step) and nothing on the DRAM board (USD 60 per step). **Response capability** is what the bank actually buys: the chain can change its object every 180 days without a -release, and a chip that carries the bank follows by firmware at a per-transition loss of -3.7 to +0.7 percent of -its shadow energy (zero credited). A family outside the bank costs that chip an emulation penalty of 2 to 6 percent +release, and a chip that carries the bank follows by firmware at a per-transition loss of -9 to +1 percent of +its shadow energy on every entry but the tile, where both sides pay +29 percent (zero credited). A family outside the bank costs that chip an emulation penalty of 2 to 6 percent of the shadow at 4 points, which every card that predates the release pays too; a structural change (a new read atom, a new derivation) costs the chip a host update and the cards a release. So the bank is a response channel whose value per event is the per-transition loss column, not a chip retirement; it is worth keeping for what it is @@ -605,36 +665,39 @@ The recommendation, five lines: design change; the three terms that still need physical design before they are bounds (the HBM activate ceiling, the SRAM die's wire, the on-package hop) are named in section 11 and owed. -## 13. D2(b): memory sharing, recomputation, data-local execution and selective participation, the cheapest combination priced (Igneum 2.0, 17:3x BST) +## 13. D2(b): memory sharing, recomputation, data-local execution and selective participation, the cheapest combination priced (Igneum 2.0; first run 17:3x BST on the synthesised rows, this run 18:0x BST on the placed rows) -The harness `tools/chip-model/mf/flow/d2b.py` (run on a rented host; build-2's pool held no free cores at 17:2x BST) -takes the mf core's rows at N5 node-for-node (the measured drawn-mix row as the anchor, the unit rows for the -per-draw movement), the card's measured 2.33 microjoules at its lock with its 0.652 premium moved by the draw's -measured per-op ratios, chip-model-v3's memory figures, adv-cache-2's measured window-layer hit rates and the layer -1 band, and prices every combination on the whole machine in the served convention (the honest 5090 at its lock -over the chip machine on its own GDDR7 board). Nothing here is a chip measurement; the bands are the memory and -SRAM bands of section 6. +The harness `tools/chip-model/mf/flow/d2b.py` (run on a rented host; build-2's pool held no free cores) takes the mf +core's placed rows at N5 node-for-node (the measured drawn-mix row as the anchor, the unit rows for the per-draw +movement), the card's measured 2.33 microjoules at its lock with its 0.652 premium moved by the draw's measured +per-op ratios, chip-model-v3's memory figures, the window-layer hit rates (the hash lane's reconcile of 17:11 BST: +the hottest half of the items serves 58 percent of the reads on the mean program, 72 percent on Devnet 3's p98 +program; adv-cache-2's figure was the tail) and the layer 1 band, and prices every combination on the whole machine +in the served convention (the honest 5090 at its lock over the chip machine on its own GDDR7 board). Nothing here +is a chip measurement; the bands are the memory and SRAM bands of section 6. -D2(b) on the mf core at N5; the served convention: the honest 5090 at its lock over the chip machine on its own GDDR7 board +D2(b) on the mf core at N5 (power: placed and routed); the served convention: the honest 5090 at its lock over the chip machine on its own GDDR7 board | Quantity (2,000 era draws under the layer 1 band, two reserve families live per epoch) | Nominal | Band | |---|---|---| -| The whole-machine ratio, mean over epochs | 1.81x | 1.51 to 2.05 | -| 5th / 50th / 95th percentile over epochs | 1.73 / 1.81 / 1.89 | spread 9.0 percent | -| A specialist mining only its best 50 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.85x / 0.50 of full-time revenue | gain in ratio 2.1 percent, at a loss of 50 percent of revenue | -| A specialist mining only its best 25 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.87x / 0.25 of full-time revenue | gain in ratio 3.4 percent, at a loss of 75 percent of revenue | -| A specialist mining only its best 10 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.89x / 0.10 of full-time revenue | gain in ratio 4.7 percent, at a loss of 90 percent of revenue | +| The whole-machine ratio, mean over epochs | 1.51x | 1.30 to 1.68 | +| 5th / 50th / 95th percentile over epochs | 1.45 / 1.51 / 1.58 | spread 9.1 percent | +| A specialist mining only its best 50 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.55x / 0.50 of full-time revenue | gain in ratio 2.1 percent, at a loss of 50 percent of revenue | +| A specialist mining only its best 25 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.57x / 0.25 of full-time revenue | gain in ratio 3.4 percent, at a loss of 75 percent of revenue | +| A specialist mining only its best 10 percent of epochs: its mean ratio while mining / its revenue against mining every epoch (difficulty adjusts to the fleet it joins; it earns in proportion to the time it mines) | 1.59x / 0.10 of full-time revenue | gain in ratio 4.7 percent, at a loss of 90 percent of revenue | | Combination (the class v4 draw) | Whole-machine ratio vs the 5090 at its lock (served convention) | Band | Machine microjoules per hash | Sustained MH/s per board | Extra capex USD | USD per MH/s (board 661 + extra) | |---|---|---|---|---|---|---| -| The baseline machine: the state in its lane, the data to it, the full dataset on the board (section 6) | 1.82x | 1.52 to 2.06 | 1.280 | 136 | 0 | 4.84 | -| hottest 25 percent of the items in SRAM beside the DRAM, serving 42% of reads (the window-layer distribution on measured programs; 0.5 GiB of SRAM, USD 125, claimed) | 2.23x | 1.84 to 2.51 | 1.045 | 236 | 125 | 3.33 | -| the same at a uniform store (25% of reads: the lower bound) | 2.02x | 1.68 to 2.28 | 1.154 | 182 | 125 | 4.32 | -| hottest 50 percent of the items in SRAM beside the DRAM, serving 72% of reads (the window-layer distribution on measured programs; 1.0 GiB of SRAM, USD 250, claimed) | 2.69x | 2.20 to 3.03 | 0.865 | 485 | 250 | 1.88 | -| the same at a uniform store (50% of reads: the lower bound) | 2.30x | 1.90 to 2.59 | 1.012 | 273 | 250 | 3.34 | -| hottest 75 percent of the items in SRAM beside the DRAM, serving 89% of reads (the window-layer distribution on measured programs; 1.5 GiB of SRAM, USD 375, claimed) | 3.10x | 2.49 to 3.49 | 0.753 | 1248 | 375 | 0.83 | -| the same at a uniform store (75% of reads: the lower bound) | 2.73x | 2.23 to 3.07 | 0.852 | 546 | 375 | 1.90 | -| hottest 50 percent in SRAM and the other 28% of items RECOMPUTED on the core (9,360 ops each) instead of read from DRAM: no DRAM at all (the rate core-bound on the baseline board's lanes) | 0.78x | 0.63 to 0.86 | 3.000 | 136 | 250 | 6.68 | +| The baseline machine: the state in its lane, the data to it, the full dataset on the board (section 6) | 1.52x | 1.31 to 1.69 | 1.528 | 136 | 0 | 4.84 | +| hottest 25 percent of the items in SRAM beside the DRAM, serving 34% of reads (the mean over programs, the hash lane's reconcile); 0.5 GiB of SRAM, USD 125, claimed | 1.73x | 1.48 to 1.91 | 1.345 | 207 | 125 | 3.80 | +| the same at a uniform store (25% of reads: the lower bound) | 1.66x | 1.42 to 1.83 | 1.403 | 182 | 125 | 4.32 | +| hottest 50 percent of the items in SRAM beside the DRAM, serving 58% of reads (the mean over programs, the hash lane's reconcile); 1.0 GiB of SRAM, USD 250, claimed | 1.93x | 1.65 to 2.12 | 1.206 | 326 | 250 | 2.80 | +| the same at a uniform store (50% of reads: the lower bound) | 1.85x | 1.58 to 2.03 | 1.260 | 273 | 250 | 3.34 | +| hottest 50 percent of the items in SRAM beside the DRAM, serving 72% of reads (the p98 program, Devnet 3); 1.0 GiB of SRAM, USD 250, claimed | 2.09x | 1.78 to 2.29 | 1.113 | 485 | 250 | 1.88 | +| the same at a uniform store (50% of reads: the lower bound) | 1.85x | 1.58 to 2.03 | 1.260 | 273 | 250 | 3.34 | +| hottest 75 percent of the items in SRAM beside the DRAM, serving 72% of reads (the mean over programs, the hash lane's reconcile); 1.5 GiB of SRAM, USD 375, claimed | 2.08x | 1.77 to 2.27 | 1.122 | 487 | 375 | 2.13 | +| the same at a uniform store (75% of reads: the lower bound) | 2.12x | 1.80 to 2.32 | 1.101 | 546 | 375 | 1.90 | +| hottest 50 percent in SRAM and the other 42% of items RECOMPUTED on the core (9,360 ops each) instead of read from DRAM: no DRAM at all (the rate core-bound on the baseline board's lanes) | 0.43x | 0.37 to 0.47 | 5.410 | 136 | 250 | 6.68 | | Form | Bits moved per dependent read | Against the baseline (80 bits to the lane) | Energy per read on a die at 1.3 pJ per bit / on a package at 0.5 / across boards at 5 to 10 | |---|---|---|---| @@ -647,9 +710,11 @@ The served bracket (the complete GDDR7 machine 1.8x node-for-node, 1.5 to 2.1; 2 Reading, and whether the served bracket moves. (1) It moves, by one combination: a hot-item SRAM beside the DRAM. The board is activate-bound, so every read an SRAM serves is a read the DRAM does not, and the same 16 devices run more hashes: the hottest half of the items (1 GiB of N2 SRAM, USD 250 claimed, the per-program hot set refilled -from the board's own DRAM in about 2 ms per epoch) serves 72 percent of the reads on the measured window-layer -distribution and reads 2.69x per joule (2.20 to 3.03) and USD 1.88 per MH/s against the baseline's 1.82x and 4.84; -with no hot-set knowledge at all (a uniform half) 2.30x and USD 3.34; a node ahead 3.33x. The record's +from the board's own DRAM in about 2 ms per epoch) serves 58 percent of the reads on the mean program and reads +1.93x per joule (1.65 to 2.12) and USD 2.80 per MH/s against the baseline's 1.52x and 4.84 on the placed rows (on +the p98 program, 72 percent of reads, 2.09x and USD 1.88); with no hot-set knowledge at all (a uniform half) 1.85x +and USD 3.34; a node ahead 2.40x (the p98 2.65x). The first run of this section on the synthesised rows read 1.82x, +2.69x and 3.33x; the placed rows take a sixth off each. The record's partial-store curve (chip-model-v3 5.4) priced the un-stored items as recomputed, which loses (0.78x here); stored in SRAM they win, and the DRAM-board chip's cheapest form is this hybrid. The dataset floor raises its SRAM ticket (USD 690 for the hot half at 5.5 GiB, USD 2.9 per MH/s) and not its ratio: the hit rate per fraction is scale-free. @@ -657,10 +722,36 @@ in SRAM they win, and the DRAM-board chip's cheapest form is this hybrid. The da medium (section 11). (3) Memory sharing is the baseline: N engines over one dataset are the lanes in flight the board already has; the set-up's share is 1.4e-12 J per hash. (4) Selective participation moves nothing: across 2,000 era draws under the layer 1 band with two reserve families live per epoch the whole-machine ratio spreads 9 -percent (5th to 95th percentile 1.73x to 1.89x); a specialist mining only its best half, quarter or tenth of epochs +percent (5th to 95th percentile 1.45x to 1.58x on the placed rows); a specialist mining only its best half, quarter or tenth of epochs gains 2, 3 or 5 percent of ratio and loses 50, 75 or 90 percent of revenue, because difficulty follows the fleet it joins and it earns in proportion to the time it mines; downtime and re-entry cost it the DAA window's lag each way and buy it nothing. The GPU-cost budget of 10 percent at the lock (set before these results) is untouched: no honest energy is spent in this section. Served sentence for the bracket: against a chip on the same node the -honest floor is the stored-half hybrid at 2.3x to 2.7x per joule and 5x to 9x per dollar, not the DRAM board's -1.8x and 3.3x; the lever on it is the dataset floor as a capex ticket. +honest floor is the stored-half hybrid at 1.9x to 2.1x per joule and 6x to 9x per dollar (1.65x to 2.3x on the +band), not the DRAM board's 1.5x and 3.3x; the lever on it is the dataset floor as a capex ticket. + +## 14. The adversary's transition matrix (Igneum 2.0 D3, the complete plan's page 12; 18:0x BST) + +Each row a required result on the evidence standard (SRAM macros with ports and area from FakeRAM2.0, the time- +multiplexed single port, routed wiring and clock tree, the memory interface and board from chip-model-v3's rows, +switching activity from the gate-level VCD; same-node N5 and advanced-node N3 kept apart; a band on every row; +measured where the placed cores give it, modelled and labelled where they do not). The lifetime rule: the +programmable design survives; a retirement credit only where the cheapest adaptation loses competitiveness. + +| Adaptation | Throughput | Energy | Cost | Survival through the schedule | Band and label | Credit | +|---|---|---|---|---|---|---| +| 1. Firmware update (a family transition inside the bank) | unchanged for 17 of 18 entries (one lane-op per slot each); the tile family at 4 points raises lane-ops per hash by 35 percent (eight dp4a per card tile), bought with USD 27 of lanes | -9 to +1 percent of the shadow energy per transition on the placed rows (all eight reserve families live -9.2, shfla and mm8 live -1.2, the 64 shape -2.4, W = 4 -0.2, one load in 16 +1.2); the tile +29 percent, the card's premium +29 percent too | USD 0 of silicon; the era's registers and the program | yes: every entry is in the die at genesis | measured on the placed 8-lane core (the SRAM term 0.95 to 3.4 pJ); the same-node ratio 1.5x (1.3 to 1.7) before and after every transition within 0.1x | 0 | +| 2. Emulation (a family outside the 18, the bank-refresh case) | the emulated op is 3 to 10 lane-ops (the cards' measured penalties for the reserve families emulated: 1.13x to 2.4x per op; a chip's sequence the same shape) at 4 points of 79: +10 to +45 percent of lane-ops, bought with lanes | +5 to +25 percent of the shadow energy while the family is live; the same-node ratio 1.5x to 1.3x to 1.45x; slower, not unprofitable (USD 4.84 to about 6 per MH/s) | USD 0 of silicon until a revision (row 6) | yes | modelled from the measured card penalties and the placed per-family rows; approximate | 0 (the cheapest adaptation keeps 2.5x to 3x per dollar) | +| 3. Memory expansion or over-provisioning | unchanged (the activate rate per device is the bound; more devices add rate one for one) | unchanged per read | the 16-device 2 GB board (32 GB) already holds the 11.5 GiB step: USD 0 through year 4; 24 Gb devices USD 1,000 of memory for 48 GB (USD 9.9 per MH/s, once); the SRAM die USD 1,000 per step; the stored-half hybrid USD 250 per GiB of hot half (USD 690 at 5.5 GiB) | yes on every arrangement; the schedule is a ticket on SRAM only | modelled on the record's memory rows and the device prices (claimed) | 0 | +| 4. Companion CPU, GPU or FPGA | the host serves 100 machines (one full node, the state-derived dataset's keeper); a GPU beside the chips proves (SP1 CUDA) and does not mine; an FPGA is 0.3x to 0.4x of the 5090 per watt on the measured HBM2 row and adds nothing | +0.85 W per machine for the host | +USD 15 per machine for the host; a 5090 prover per shard rate the operator chooses (the proving line is the GPU line, unchanged) | yes | the host figures approximate; the FPGA row the Horizon lane's measured public row | 0: the hybrid mining-and-proving operator mines at the chip's economics and proves at the GPU's; the 80/20 split caps the proving share | +| 5. Favourable-period mining | the ratio spreads 9 percent across 2,000 era draws under the layer 1 band (5th to 95th percentile 1.45x to 1.58x same-node); mining only the best half, quarter or tenth of epochs gains 2, 3 or 5 percent of ratio | the same | revenue 50, 75 or 90 percent lower; re-entry costs the DAA window's lag each way; difficulty follows the fleet it joins | yes | modelled on the placed rows and the band; the draws uniform within the band (approximate) | 0 (nothing to gain) | +| 6. Silicon revision (a new family the firmware cannot carry economically) | a native unit restores row 1's throughput | row 1's energy | a metal-only respin USD 1 M to 3 M and 3 to 4 months, a full respin USD 10 M to 20 M of masks and 6 to 9 months at N5 (claimed, the record's project figures); the whole core design reused | yes: a bank refresh takes effect at family epoch n + 2, 360 days after its tally, so the revision ships before the family is live | approximate | 0 (the delay fits inside the bank-refresh notice) | + +Reading: no row loses competitiveness; the cheapest adaptation is firmware for every entry in the bank (row 1), a +sequence for an entry outside it (row 2, 1.3x to 1.45x same-node while live), and a respin only if the chain adds a +family the sequence cannot carry, which it has 360 days' notice of. The stress life of three years holds in full. + +PENDING with clocks: the 32-lane full core's synthesis row (adv-a, in ABC at 18:0x BST; by 19:30) and its placed +row (adv-f, at floorplan; by 21:00); the k lane's crossbar, scratch and tile rows (its pod); a real memory +compiler's figure for the two macros (owed, no clock: FakeRAM gives area and pins only); the per-family rows on +the 32-lane placed core (by 21:30 if adv-f lands, else the next pass).