diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index 69240d31e..10c6c8797 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -441,7 +441,7 @@ The conditions, read off the surface, which are the economic-resistance statemen **The sentence, as the external review words it (10.0f item 5), served verbatim, with one word made honest (the site audit lane's read, 16:5x UK: class v4 and v5 have eight registers per lane and the window is class v6's new core shape, so "retains" is read as "retains across every rotation"):** "Class v6 adopts the 64-register window and retains it across every rotation. Current modelling estimates a 2.2x to 2.4x energy-efficiency advantage for the strongest specialised designs assessed against the GPU tier (2.0x on the GPU's own node). The long-program and select-tree proposals were rejected. Economic resistance depends on development cost, deployment economics and productive hardware lifetime; family transitions receive an obsolescence benefit only where a loss of competitiveness is demonstrated; programmable multi-epoch designs are included in the assessment." -The labels on its numbers: "2.2x to 2.4x" is modelled (the GPU side measured: the RTX 5080 at its 1,100 MHz lock 2.06 microjoules per hash and the RTX 5090 at its 1,300 MHz lock 2.33, both on class v4, PC 1 and rented pods, 8 October 2026; the chip side the k lane's synthesised 8-lane sequencer core with the 64-register window on ASAP7, scaled to N3 on TSMC's headline factors, claimed; the chip's memory the chip model's GDDR7 board, modelled; the card's cost of the window measured at stock on a rented 5090 and 4090 at 16:4x UK, within 5 percent per load with the liveness chain, no spill). "2.0x on the GPU's own node" is modelled (the same core node-for-node, k 1.09). Against the 32-lane window core the adversary would build the same figures read about 2.4x to 2.6x a node ahead and 2.0x to 2.2x node-for-node (synthesised, pending the re-optimised row by 18:00 UK); the served sentence's range is kept as the review wrote it and the 32-lane rows sit beside it on the page as the pending row. The window's k is synthesis-derived and not a lower bound. +**Where the figures come from (the coordinator, 16:1x UK): the served sentence's numbers are read from the k lane's PLACED GATED rows (17:30 UK; 10.0i), not from 725d2945's close and not from the 16:0x synthesis; until they land the figures sit in 10.0i's bracket, and the serve slips to 18:30 if the rows are late.** The labels on its numbers as first written: "2.2x to 2.4x" is modelled (the GPU side measured: the RTX 5080 at its 1,100 MHz lock 2.06 microjoules per hash and the RTX 5090 at its 1,300 MHz lock 2.33, both on class v4, PC 1 and rented pods, 8 October 2026; the chip side the k lane's synthesised 8-lane sequencer core with the 64-register window on ASAP7, scaled to N3 on TSMC's headline factors, claimed; the chip's memory the chip model's GDDR7 board, modelled; the card's cost of the window measured at stock on a rented 5090 and 4090 at 16:4x UK, within 5 percent per load with the liveness chain, no spill). "2.0x on the GPU's own node" is modelled (the same core node-for-node, k 1.09). Against the 32-lane window core the adversary would build the same figures read about 2.4x to 2.6x a node ahead and 2.0x to 2.2x node-for-node (synthesised, pending the re-optimised row by 18:00 UK); the served sentence's range is kept as the review wrote it and the 32-lane rows sit beside it on the page as the pending row. The window's k is synthesis-derived and not a lower bound. The lines the page carries beside it, each labelled: - The three statements, separate (10.0g item 1): energy resistance (the figures above); economic resistance (the profitability surface of 10.0f item 2, lane 3's first cut, modelled: p* scales as the project cost over the share times the discounted life, and under 5 percent with the per-joule edge; the cheapest attractive project is a USD 20 M DRAM-board design taking the whole chain for three years at about IGN 0.02 to 0.03, at a third 0.055 to 0.10; the SRAM die at N2 0.22 to 0.73; a fixed-lane chip under rotation needs 4x the price of a programmable one; stated as the conditions under which development is attractive); response capability (a passed rotation boundary proves the rotation works, not that hardware dies; the schedule of 10.0d: hourly, weekly, 180-day family, emergency vote; measured per boundary). @@ -453,6 +453,22 @@ The lines the page carries beside it, each labelled: Withdrawn, to be struck everywhere it appears: the lifetime claim ("under 1x per hash over its 180-day life"; 10.0f item 1); the USD 300 M and USD 340 M lines and any "the chain stays below X" wording (10.0f item 2); any chip-arrival probability (10.0g item 5); "Monero's hash never changes" (10.0g item 6); the 725d2945 sentence ("under 3x per joule at the knee, under 1x per hash over its life") and this lane's section 9 proposal; W = 8 as a next width (W = 4 at genesis, 8 measured not free, 16 never). +#### 10.0i The adversary's re-optimised core, and the bracket the served figures sit in (the k lane, 16:0x UK; the coordinator's order 16:1x: the placed gated rows at 17:30 UK are the figures to serve, not 725d2945's and not this synthesis; the 18:00 serve slips to 18:30 if they are late) + +The rows (synthesis-only unless marked, 8 lanes, ASAP7 TC, the register-file clock gated by inferred ICG cells, a gate-level random-input VCD with every pin annotated, the steady state solved from 150 and 600 run cycles; a MODEL of a chip core, never a lower bound; N3 and N2 claimed): + +| Core (8 lanes) | Cells | pJ per lane-op ASAP7 | of which sequential | N5 (the card's node) | N3 | N2 | k at the lock, N5 / N3 / N2 | k at stock, N3 | k against the M5 Max, N3 | Label | +|---|---|---|---|---|---|---|---|---|---|---| +| The GPU-shaped base, 32 registers, ungated (the 14:0x headline row) | 186,443 | 6.9 | 2.4 | 4.8 | 3.5 | 2.5 | 0.78 / 0.56 / 0.40 | 0.31 | 0.50 | synthesised | +| The adversary's base, 32 registers, the register file clock-gated | 156,833 | 4.5 | 0.15 | 3.2 | 2.3 | 1.6 | 0.51 / 0.37 / 0.26 | 0.20 | 0.33 | synthesised | +| The GPU-shaped 64-register window, ungated | 267,731 | 9.7 | 3.75 | 6.8 | 4.9 | 3.5 | 1.09 / 0.78 / 0.56 | 0.43 | 0.70 | synthesised | +| The adversary's 64-register window, clock-gated | 225,441 | 6.2 | 0.2 | 4.3 | 3.1 | 2.2 | 0.70 / 0.50 / 0.36 | 0.28 | 0.45 | synthesised | +| The GPU-shaped base, 32 registers, ungated, PLACED AND ROUTED with SPEF and a clock tree (one run length, the load phase subtracted, about plus or minus 10 percent) | 444,478 | 11.3 | 2.3 flops plus 2.5 clock tree | 7.9 | 5.6 | 4.1 | 1.27 / 0.91 / 0.66 | 0.50 | 0.82 | placed | + +The live-state analysis (`livestate.py`, 64 drawn programs, 1,024 waits): under a fold that reads every register, 61.7 of 64 window values are live at every wait and 61.0 are necessary (they reach a later address or the result transitively), dead writes 2.8 percent; so the adversary cannot shrink the state by liveness or recompute, only make its access cheaper. Its cheaper forms: clock gating (measured above: it removes the whole flop-clock term, 2.4 of 6.9 and 3.75 of 9.7 pJ); a latch file (about 0.7x the gated row, modelled); SRAM-banked time-multiplexed state (modelled 8.5 to 10.5 pJ: NOT cheaper at 256 bytes per lane; the RTL form in synthesis on a pod as a check). **So the defence that remains is the gated 64-register core against the gated 32-register core: +1.7 pJ per lane-op at ASAP7, +0.84 at N3, +0.13 of k at the lock (0.37 to 0.50 at N3; 0.51 to 0.70 node-for-node). With the GPU side measured (10.0e: no spill, 83 and 67 percent occupancy, at most 5 percent per load on either card), the window is a real but small defence: the GDDR7 board at the lock moves from 3.3x (the gated base, N3) to 2.9x (the gated window), from 2.9x to 2.5x node-for-node.** The window stays (its liveness measured, its GPU cost under 5 percent); the honest line is that it adds about 0.13 of k after the adversary re-optimises. + +Two corrections this forces on the served numbers: (1) the honest adversary's base core is the GATED one, k 0.37 at N3 and 0.51 node-for-node, below the 0.56 and 0.78 of 14:0x (those are the GPU-shaped core a maker would not build); (2) placement adds more than the +30 percent estimated at 14:1x: the ungated placed core reads 11.3 pJ against 6.9 synthesised (+64 percent: wires and a 2.5 pJ clock tree). **So until the placed gated rows land (in flight on a rented pod, 17:30 UK) the served figures sit in a bracket, from the synthesised gated rows (4.5 and 6.2 pJ per lane-op: the GDDR7 board 3.3x to 2.9x at the lock at N3, 2.9x to 2.5x node-for-node) to the placed ungated row (11.3 pJ: 2.3x at N3 and 2.0x node-for-node for the base core, the window below it), with the placed gated figure expected near 6 to 8 pJ (approximate): about 2.6x to 3.1x at the lock at N3 and 2.3x to 2.7x node-for-node, the window about 0.3x under the base.** The placed gated rows replace this bracket as the served number when they land, and 10.0h's figures are read from them. + #### 10.0d The rotation schedule the close adopts (the rotation lane, `docs/design/class-rotation-four-layers.md` on class-v6-rotation at bd43f808, build-3, gate green, 14:4x UK; one line per layer; both of this document's constraints held: the 180-day family epoch not shorter, W = 4 not drawn) | Layer | Boundaries a year | What it draws, from where | Exposure per boundary (this document's units) | Chip | Label |