multi-family adversary: clock stamps read from the Mac
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of 10b901a46 (10b901a46934c24ce9b44c74bc6fd4db7c40d2b0) for the box mirror master
This commit is contained in:
parent
23f3073676
commit
0db43b6578
1 changed files with 7 additions and 7 deletions
|
|
@ -274,7 +274,7 @@ and detailed placement, CTS, global and detailed routing, OpenRCX parasitics; Op
|
|||
SPEF under the gate-level VCD; the clock at 12,000 ps with leakage restated at the 1,500 ps equivalent, section 2.1;
|
||||
the final PDN connectivity check reports two macro power pins unconnected, a floorplan artefact of the FakeRAM
|
||||
pin geometry that does not touch the netlist, the parasitics or the power figure). The 8-lane genesis core (the
|
||||
comparator) first, then the 8-lane full core (routed, 58 tags, 18:0x BST).
|
||||
comparator) first, then the 8-lane full core (routed, 58 tags, 17:0x BST).
|
||||
|
||||
The genesis core, 8 lanes, routed: 161,504 cells (the synthesis 66,973: placement adds fill, buffers and the clock
|
||||
tree), 2 macros (the base core's instruction word fits 34 bits, so Yosys removed the second imem macro):
|
||||
|
|
@ -351,7 +351,7 @@ per lane against the draw's 10.3): the ratio does not move, the lane count does
|
|||
The 32-lane genesis core (synthesis only; four window macros, one imem pair shared by 32 lanes): 261,440 cells,
|
||||
5.29 pJ per lane-op at ASAP7 (4.65 to 6.77) on the class v4 draw, 3.70 at N5, k 0.36 at the lock: 10 percent
|
||||
under the 8-lane core, the imem and the sequencer amortised over four times the lanes. The 32-lane full core
|
||||
(synthesis) and its placed form are in the flow (adv-a and adv-f at 17:4x BST) and land as a delta.
|
||||
(synthesis) and its placed form are in the flow (adv-a and adv-f at 16:4x BST) and land as a delta.
|
||||
|
||||
The board rows of section 6 restated at the placed energy (the full core at 6.55 pJ per lane-op at N5, band 6.01
|
||||
to 7.77; 4.72 at N3): the complete GDDR7 machine 1.53 microjoules per hash, **1.5x the 5090 at its lock per joule
|
||||
|
|
@ -559,7 +559,7 @@ the HBM3 activate ceiling (unmeasured, the AWS F2 hour); the profitability surfa
|
|||
lifetime rows of section 6, which this file states as cost per TH and leaves the NPV to the economics lane.
|
||||
|
||||
|
||||
## 11. The data-local and memory-sharing adversary (the ProgPoW audit's threat; the third review, 17:5x BST)
|
||||
## 11. The data-local and memory-sharing adversary (the ProgPoW audit's threat; the third review, 17:0x BST)
|
||||
|
||||
The threat: split the dataset across processors and move the intermediate computation to the processor nearest
|
||||
the next item, share datasets, keep partial caches, recompute, re-lay the dataset, run several engines off one
|
||||
|
|
@ -615,7 +615,7 @@ hour), the SRAM die's wire term (0.5 to 2.0 nJ per 64 bytes until a placed macro
|
|||
(UCIe's 0.5 pJ per bit is a claim), and the data-local form on a 3D-stacked SRAM (a vertical hop at about 0.1 pJ per
|
||||
bit, approximate, would cut the one-die row to about 0.2 nJ, still 2x the baseline and still no saving).
|
||||
|
||||
## 12. The dataset comparison that decides class v7 (the third review, 17:5x BST)
|
||||
## 12. The dataset comparison that decides class v7 (the third review, 17:0x BST)
|
||||
|
||||
Two datasets, each with the adversary's burden (a chip with a host keeping it current) and the commodity burden
|
||||
(what every honest node pays at the boundaries), priced on the record's figures:
|
||||
|
|
@ -665,7 +665,7 @@ The recommendation, five lines:
|
|||
design change; the three terms that still need physical design before they are bounds (the HBM activate
|
||||
ceiling, the SRAM die's wire, the on-package hop) are named in section 11 and owed.
|
||||
|
||||
## 13. D2(b): memory sharing, recomputation, data-local execution and selective participation, the cheapest combination priced (Igneum 2.0; first run 17:3x BST on the synthesised rows, this run 18:0x BST on the placed rows)
|
||||
## 13. D2(b): memory sharing, recomputation, data-local execution and selective participation, the cheapest combination priced (Igneum 2.0; first run 16:3x BST on the synthesised rows, this run 17:0x BST on the placed rows)
|
||||
|
||||
The harness `tools/chip-model/mf/flow/d2b.py` (run on a rented host; build-2's pool held no free cores) takes the mf
|
||||
core's placed rows at N5 node-for-node (the measured drawn-mix row as the anchor, the unit rows for the per-draw
|
||||
|
|
@ -730,7 +730,7 @@ no honest energy is spent in this section. Served sentence for the bracket: agai
|
|||
honest floor is the stored-half hybrid at 1.9x to 2.1x per joule and 6x to 9x per dollar (1.65x to 2.3x on the
|
||||
band), not the DRAM board's 1.5x and 3.3x; the lever on it is the dataset floor as a capex ticket.
|
||||
|
||||
## 14. The adversary's transition matrix (Igneum 2.0 D3, the complete plan's page 12; 18:0x BST)
|
||||
## 14. The adversary's transition matrix (Igneum 2.0 D3, the complete plan's page 12; 17:0x BST)
|
||||
|
||||
Each row a required result on the evidence standard (SRAM macros with ports and area from FakeRAM2.0, the time-
|
||||
multiplexed single port, routed wiring and clock tree, the memory interface and board from chip-model-v3's rows,
|
||||
|
|
@ -751,7 +751,7 @@ Reading: no row loses competitiveness; the cheapest adaptation is firmware for e
|
|||
sequence for an entry outside it (row 2, 1.3x to 1.45x same-node while live), and a respin only if the chain adds a
|
||||
family the sequence cannot carry, which it has 360 days' notice of. The stress life of three years holds in full.
|
||||
|
||||
PENDING with clocks: the 32-lane full core's synthesis row (adv-a, in ABC at 18:0x BST; by 19:30) and its placed
|
||||
PENDING with clocks: the 32-lane full core's synthesis row (adv-a, in ABC at 17:0x BST; by 19:30) and its placed
|
||||
row (adv-f, at floorplan; by 21:00); the k lane's crossbar, scratch and tile rows (its pod); a real memory
|
||||
compiler's figure for the two macros (owed, no clock: FakeRAM gives area and pins only); the per-family rows on
|
||||
the 32-lane placed core (by 21:30 if adv-f lands, else the next pass).
|
||||
|
|
|
|||
Loading…
Reference in a new issue