Class v6 10.0h and 10.0i: the sentence served at 18:30 on the tightened bracket (2.5x to 3.0x a node ahead, 2.1x to 2.6x node-for-node; the placed gated row at about 18:30 narrows it)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
50674ea51e
commit
1ace22052f
1 changed files with 11 additions and 1 deletions
|
|
@ -443,7 +443,7 @@ The conditions, read off the surface, which are the economic-resistance statemen
|
|||
|
||||
**The sentence, as the external review words it (10.0f item 5), served verbatim, with one word made honest (the site audit lane's read, 16:5x UK: class v4 and v5 have eight registers per lane and the window is class v6's new core shape, so "retains" is read as "retains across every rotation"):** "Class v6 adopts the 64-register window and retains it across every rotation. Current modelling estimates a 2.2x to 2.4x energy-efficiency advantage for the strongest specialised designs assessed against the GPU tier (2.0x on the GPU's own node). The long-program and select-tree proposals were rejected. Economic resistance depends on development cost, deployment economics and productive hardware lifetime; family transitions receive an obsolescence benefit only where a loss of competitiveness is demonstrated; programmable multi-epoch designs are included in the assessment."
|
||||
|
||||
**Where the figures come from (the coordinator, 16:1x UK): the served sentence's numbers are read from the k lane's PLACED GATED rows (17:30 UK; 10.0i), not from 725d2945's close and not from the 16:0x synthesis; until they land the figures sit in 10.0i's bracket, and the serve slips to 18:30 if the rows are late.** The labels on its numbers as first written: "2.2x to 2.4x" is modelled (the GPU side measured: the RTX 5080 at its 1,100 MHz lock 2.06 microjoules per hash and the RTX 5090 at its 1,300 MHz lock 2.33, both on class v4, PC 1 and rented pods, 8 October 2026; the chip side the k lane's synthesised 8-lane sequencer core with the 64-register window on ASAP7, scaled to N3 on TSMC's headline factors, claimed; the chip's memory the chip model's GDDR7 board, modelled; the card's cost of the window measured at stock on a rented 5090 and 4090 at 16:4x UK, within 5 percent per load with the liveness chain, no spill). "2.0x on the GPU's own node" is modelled (the same core node-for-node, k 1.09). Against the 32-lane window core the adversary would build the same figures read about 2.4x to 2.6x a node ahead and 2.0x to 2.2x node-for-node (synthesised, pending the re-optimised row by 18:00 UK); the served sentence's range is kept as the review wrote it and the 32-lane rows sit beside it on the page as the pending row. The window's k is synthesis-derived and not a lower bound.
|
||||
**Where the figures come from (the coordinator, 16:1x UK): the served sentence's numbers are read from the k lane's PLACED GATED rows (10.0i), not from 725d2945's close and not from the 16:0x synthesis; the placement slipped to about 18:30 UK, so the sentence served at 18:30 carries 10.0i's tightened bracket, "estimates a 2.5x to 3.0x energy-efficiency advantage for the strongest specialised designs assessed against the GPU tier (2.1x to 2.6x on the GPU's own node)", and the placed row narrows it to one figure each on the next landing.** The labels on its numbers as first written: "2.2x to 2.4x" is modelled (the GPU side measured: the RTX 5080 at its 1,100 MHz lock 2.06 microjoules per hash and the RTX 5090 at its 1,300 MHz lock 2.33, both on class v4, PC 1 and rented pods, 8 October 2026; the chip side the k lane's synthesised 8-lane sequencer core with the 64-register window on ASAP7, scaled to N3 on TSMC's headline factors, claimed; the chip's memory the chip model's GDDR7 board, modelled; the card's cost of the window measured at stock on a rented 5090 and 4090 at 16:4x UK, within 5 percent per load with the liveness chain, no spill). "2.0x on the GPU's own node" is modelled (the same core node-for-node, k 1.09). Against the 32-lane window core the adversary would build the same figures read about 2.4x to 2.6x a node ahead and 2.0x to 2.2x node-for-node (synthesised, pending the re-optimised row by 18:00 UK); the served sentence's range is kept as the review wrote it and the 32-lane rows sit beside it on the page as the pending row. The window's k is synthesis-derived and not a lower bound.
|
||||
|
||||
The lines the page carries beside it, each labelled:
|
||||
- The three statements, separate (10.0g item 1): energy resistance (the figures above); economic resistance (the profitability surface of 10.0f item 2, lane 3's first cut, modelled: p* scales as the project cost over the share times the discounted life, and under 5 percent with the per-joule edge; the cheapest attractive project is a USD 20 M DRAM-board design taking the whole chain for three years at about IGN 0.02 to 0.03, at a third 0.055 to 0.10; the SRAM die at N2 0.22 to 0.73; a fixed-lane chip under rotation needs 4x the price of a programmable one; stated as the conditions under which development is attractive); response capability (a passed rotation boundary proves the rotation works, not that hardware dies; the schedule of 10.0d: hourly, weekly, 180-day family, emergency vote; measured per boundary).
|
||||
|
|
@ -471,6 +471,16 @@ The live-state analysis (`livestate.py`, 64 drawn programs, 1,024 waits): under
|
|||
|
||||
Two corrections this forces on the served numbers: (1) the honest adversary's base core is the GATED one, k 0.37 at N3 and 0.51 node-for-node, below the 0.56 and 0.78 of 14:0x (those are the GPU-shaped core a maker would not build); (2) placement adds more than the +30 percent estimated at 14:1x: the ungated placed core reads 11.3 pJ against 6.9 synthesised (+64 percent: wires and a 2.5 pJ clock tree). **So until the placed gated rows land (in flight on a rented pod, 17:30 UK) the served figures sit in a bracket, from the synthesised gated rows (4.5 and 6.2 pJ per lane-op: the GDDR7 board 3.3x to 2.9x at the lock at N3, 2.9x to 2.5x node-for-node) to the placed ungated row (11.3 pJ: 2.3x at N3 and 2.0x node-for-node for the base core, the window below it), with the placed gated figure expected near 6 to 8 pJ (approximate): about 2.6x to 3.1x at the lock at N3 and 2.3x to 2.7x node-for-node, the window about 0.3x under the base.** The placed gated rows replace this bracket as the served number when they land, and 10.0h's figures are read from them.
|
||||
|
||||
The row to serve at 18:30 UK (the k lane, 17:5x UK; the placement of the gated 64-register core slipped to about 18:30 on a floorplan timing repair, the other five placed rows by 21:00): on the gated 64-register core, synthesis-only, a model never a lower bound (the GDDR7 board at the 5090's 1,300 MHz lock, 2.33 microjoules; E_chip = 0.466 + 102,100 x e_chip; node factors claimed):
|
||||
|
||||
| Figure | Chip pJ per lane-op | E_chip microjoules | The edge at the lock | Label |
|
||||
|---|---|---|---|---|
|
||||
| node-for-node (N5, ASAP7 x0.70) | 4.3 | 0.905 | 2.6x | synthesised |
|
||||
| a node ahead (N3) | 3.1 | 0.783 | 3.0x | synthesised, scaling claimed |
|
||||
| two nodes ahead (N2) | 2.2 | 0.691 | 3.4x | synthesised, scaling claimed |
|
||||
|
||||
The placed figure runs 30 to 65 percent over synthesis on this flow (the ungated base came in 64 percent over, wires and a clock tree, which gating removes in part), so the placed gated core is expected at 7 to 9 pJ at ASAP7, which puts the served figures at 2.1x to 2.4x node-for-node and 2.5x to 2.8x a node ahead (approximate until the placed row). **So the served bracket of this section holds and tightens to its lower half, and the honest sentence until the placed row is: "estimates a 2.5x to 3.0x energy-efficiency advantage (2.1x to 2.6x on the GPU's own node)"**, the placed row narrowing it to one figure each. The GPU side of the window is measured (no spill, at most 5 percent per load); the connected-state and multi-family lanes' rows agree with this core within 5 percent.
|
||||
|
||||
#### 10.0j Amendment after the landing (16:4x UK): the multi-family adversary lane's first core, and a disagreement between two models that the placed rows settle
|
||||
|
||||
The multi-family adversary lane (a1a9876a88f5a72fc; synthesis only, ASAP7 TC, a gate-level random-input VCD; the SRAM macro energy modelled with a band; node factors claimed; for class v7, but it bears on the served window line): one in-order SIMD core with the 64-register window in a FakeRAM 64 x 256 macro per 8 lanes (one 256-bit access serves eight lanes), the imem in two 256 x 34 macros, every unit operand-isolated, a 5-phase single-port slot (throughput bought with lanes, not ports), every bank entry firmware. The genesis-only variant on the class v4 draw: 5.9 pJ per lane-op at ASAP7 (band 5.1 to 7.7), 4.1 at N5, 3.0 at N3; the card pays 10.3 pJ per op on the same draw at the 1,300 lock, so k = 0.40 node-for-node (N5), 0.29 a node ahead (N3), 0.22 / 0.16 at stock. Per family (k N5 / N3 at the lock): the add class 0.45 / 0.33, or 0.36 / 0.26, mul 0.32 / 0.23, mad 0.55 / 0.40, mulhi 0.12 / 0.09, shfl 0.083 / 0.060, the load with the fold 0.43 / 0.31. The whole-hash shadow 0.42 microjoules at N5 (0.31 at N3), so the GDDR7 board reads 2.6x against the 5090 at its lock node-for-node (2.9x a node ahead) and 2.3x / 2.6x against the 5080. **The lane's reading: a re-optimised core sits 15 percent under the k lane's 32-register flop core and 40 percent under its 64-register flop core at the same node, and the register-window knob buys the card nothing once the adversary puts the state in a macro.**
|
||||
|
|
|
|||
Loading…
Reference in a new issue