diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index 10c6c8797..fd1731082 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -469,6 +469,12 @@ The live-state analysis (`livestate.py`, 64 drawn programs, 1,024 waits): under Two corrections this forces on the served numbers: (1) the honest adversary's base core is the GATED one, k 0.37 at N3 and 0.51 node-for-node, below the 0.56 and 0.78 of 14:0x (those are the GPU-shaped core a maker would not build); (2) placement adds more than the +30 percent estimated at 14:1x: the ungated placed core reads 11.3 pJ against 6.9 synthesised (+64 percent: wires and a 2.5 pJ clock tree). **So until the placed gated rows land (in flight on a rented pod, 17:30 UK) the served figures sit in a bracket, from the synthesised gated rows (4.5 and 6.2 pJ per lane-op: the GDDR7 board 3.3x to 2.9x at the lock at N3, 2.9x to 2.5x node-for-node) to the placed ungated row (11.3 pJ: 2.3x at N3 and 2.0x node-for-node for the base core, the window below it), with the placed gated figure expected near 6 to 8 pJ (approximate): about 2.6x to 3.1x at the lock at N3 and 2.3x to 2.7x node-for-node, the window about 0.3x under the base.** The placed gated rows replace this bracket as the served number when they land, and 10.0h's figures are read from them. +#### 10.0j Amendment after the landing (16:4x UK): the multi-family adversary lane's first core, and a disagreement between two models that the placed rows settle + +The multi-family adversary lane (a1a9876a88f5a72fc; synthesis only, ASAP7 TC, a gate-level random-input VCD; the SRAM macro energy modelled with a band; node factors claimed; for class v7, but it bears on the served window line): one in-order SIMD core with the 64-register window in a FakeRAM 64 x 256 macro per 8 lanes (one 256-bit access serves eight lanes), the imem in two 256 x 34 macros, every unit operand-isolated, a 5-phase single-port slot (throughput bought with lanes, not ports), every bank entry firmware. The genesis-only variant on the class v4 draw: 5.9 pJ per lane-op at ASAP7 (band 5.1 to 7.7), 4.1 at N5, 3.0 at N3; the card pays 10.3 pJ per op on the same draw at the 1,300 lock, so k = 0.40 node-for-node (N5), 0.29 a node ahead (N3), 0.22 / 0.16 at stock. Per family (k N5 / N3 at the lock): the add class 0.45 / 0.33, or 0.36 / 0.26, mul 0.32 / 0.23, mad 0.55 / 0.40, mulhi 0.12 / 0.09, shfl 0.083 / 0.060, the load with the fold 0.43 / 0.31. The whole-hash shadow 0.42 microjoules at N5 (0.31 at N3), so the GDDR7 board reads 2.6x against the 5090 at its lock node-for-node (2.9x a node ahead) and 2.3x / 2.6x against the 5080. **The lane's reading: a re-optimised core sits 15 percent under the k lane's 32-register flop core and 40 percent under its 64-register flop core at the same node, and the register-window knob buys the card nothing once the adversary puts the state in a macro.** + +The disagreement, carried as two models until the placed rows settle it: the k lane's SRAM-banked time-multiplexed window (10.0i) is modelled at 8.5 to 10.5 pJ per lane-op, NOT cheaper than its gated flop window (6.2), with the RTL form in synthesis on a pod as a check; the multi-family lane's macro window reads 5.9 (band 5.1 to 7.7) for the whole core. The two differ in the macro's energy model and the access pattern (one 256-bit access per eight lanes against a per-lane port), and in whether the window's whole state is live across every wait (10.0i: 61 of 64 values necessary, so no lane can skip the access). Until the k lane's placed gated rows (17:30 UK) and the multi-family lane's placed 8-lane core (18:30 UK) land, the served window line is read with this caveat: **the window's measured cost to the card is under 5 percent and its defence against a flop-file adversary is about 0.13 of k; against a macro-file adversary it may be close to nothing, which is why the next programme's multi-family adversary is the lane that decides whether the window stays a defence or only a cost.** The 18-family variant (+43 percent of cells on the same microarchitecture), the live-family mixes and the 32-lane cores follow by 18:30 with the board and lifetime rows. + #### 10.0d The rotation schedule the close adopts (the rotation lane, `docs/design/class-rotation-four-layers.md` on class-v6-rotation at bd43f808, build-3, gate green, 14:4x UK; one line per layer; both of this document's constraints held: the 180-day family epoch not shorter, W = 4 not drawn) | Layer | Boundaries a year | What it draws, from where | Exposure per boundary (this document's units) | Chip | Label |