multi-family adversary: the routed clock relaxation, the tier table, the statement, the data-local cost model and the dataset comparison
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of 2db208912 (2db208912e5fc0656930aaabd7db74ec2a68ca8e) for the box mirror master; left on the branch: tools/chip-model/mf/Makefile tools/chip-model/mf/flow/collect.py tools/chip-model/mf/flow/mf32base.mk tools/chip-model/mf/flow/mf32base.sdc tools/chip-model/mf/flow/mf32full.mk tools/chip-model/mf/flow/mf32full.sdc tools/chip-model/mf/flow/mf8base.mk tools/chip-model/mf/flow/mf8base.sdc tools/chip-model/mf/flow/mf8full.mk tools/chip-model/mf/flow/mf8full.sdc tools/chip-model/mf/tb/tb_mf_common.vh
This commit is contained in:
parent
01452ff1ea
commit
ad5ca6e47c
1 changed files with 154 additions and 3 deletions
|
|
@ -82,7 +82,12 @@ utilisation (30 for the 32-lane core) with the macros placed by the flow's macro
|
|||
global and detailed placement, CTS, global and detailed routing, OpenRCX parasitics. Power is OpenSTA `report_power`
|
||||
under the VCD of a random-input gate-level simulation of the netlist (iverilog; every instruction field drawn by
|
||||
`$random`, the window initialised with random words, a random returned word on every load), with a propagated 0.5
|
||||
activity as the cross-check. Synthesis-only rows (no wires, no clock tree) are marked; placed rows carry the SPEF.
|
||||
activity as the cross-check. Synthesis-only rows (no wires, no clock tree) are marked; placed rows carry the SPEF. The routed runs clock at
|
||||
12,000 ps with the ABC target held at 1,500 ps: the unpipelined execute path (the multiplier, the fold and the
|
||||
lane reduction) is 10.6 ns at ASAP7 TC, and at 1,500 ps the flow's timing repair spent its time on a path the
|
||||
adversary would pipeline instead (two to three registers per lane, about 0.1 to 0.2 pJ per lane-op, inside the
|
||||
band). Energy per op does not depend on the period; the leakage term does, and the collector restates it at the
|
||||
1,500 ps equivalent for the routed rows (both are in `table.csv`).
|
||||
|
||||
### 2.2 The activity and the steady state
|
||||
|
||||
|
|
@ -375,9 +380,49 @@ and this is its price: under 1 W and USD 15 per machine.
|
|||
|
||||
## 8. Energy resistance, economic resistance and response capability, stated separately
|
||||
|
||||
STATEMENT
|
||||
Three separate things, each with its own number and its own holder.
|
||||
|
||||
## 9. Sources and what is owed
|
||||
**Energy resistance** is a property of the memory system and the shadow, not of the calendar. Against the
|
||||
adversary's re-optimised chip the complete GDDR7 machine reads 1.8x the 5090 at its lock per joule node-for-node
|
||||
(1.5x to 2.1x), 2.1x a node ahead; 1.6x and 1.9x against the 5080 at its lock; 2.8x and 3.3x against the cohort
|
||||
card; the N2 SRAM die 3.1x and 4.2x. The bank moves these by 5 percent in the honest side's favour (the chip's
|
||||
shadow energy +11 percent for carrying 18 families instead of 10) and no more; the register window moves them by
|
||||
0.03 to 0.13 of k and no more. What holds the per-joule number is the shadow's size on the card's own operating
|
||||
point and the memory system's activate ceiling; what would move it is a memory arrangement the chip cannot buy, and
|
||||
section 6 says there is none: the chip's cheapest memory is the card's own.
|
||||
|
||||
**Economic resistance** is capex per sustained MH/s and the project cost against the chain's revenue. The chip's
|
||||
machine costs USD 4.84 per MH/s against the card's 16 to 18, so over a 3-year life it mines at 0.080 USD per TH
|
||||
against 0.22 to 0.27: 2.7x to 3.4x, of which electricity is the smaller half. The calendar does not shorten that
|
||||
life: no transition in the bank retires the chip, so the 3-year stress life holds in full and the withdrawn headline
|
||||
("dies within an epoch, under 1x over its life") stays withdrawn. What holds the economics is the project (USD 20 M
|
||||
to 75 M for the GDDR7-board chip, floor lane 5; USD 100 M to 500 M for the SRAM die, claimed) against the miner
|
||||
revenue the chain pays, which is the profitability surface the review asks for and the economics lane owns; the
|
||||
dataset floor is a ticket on the SRAM die only (USD 1,000 per step) and nothing on the DRAM board (USD 60 per step).
|
||||
|
||||
**Response capability** is what the bank actually buys: the chain can change its object every 180 days without a
|
||||
release, and a chip that carries the bank follows by firmware at a per-transition loss of -3.7 to +0.7 percent of
|
||||
its shadow energy (zero credited). A family outside the bank costs that chip an emulation penalty of 2 to 6 percent
|
||||
of the shadow at 4 points, which every card that predates the release pays too; a structural change (a new read
|
||||
atom, a new derivation) costs the chip a host update and the cards a release. So the bank is a response channel
|
||||
whose value per event is the per-transition loss column, not a chip retirement; it is worth keeping for what it is
|
||||
(no fork for a family change, a fixed-function datapath dead on day one, section 7's table) and must not be served as
|
||||
energy or economic resistance.
|
||||
|
||||
|
||||
## 9. Consequences per tier (the standing rule)
|
||||
|
||||
| Tier | What the rows mean for it | What is being done |
|
||||
|---|---|---|
|
||||
| Home miner, one 8 GB card | Against the adversary's machine the cohort card sits 2.8x to 3.3x behind per joule and 3.7x per dollar; the bank does not change that, the operating-point lock does not reach this tier (no Blackwell lever on Ampere; Ada's lock is worth 1.2x); the first dataset step (5.5 GiB) retires this card from mining | the schedule's replacement cost (USD 350 to 450 to a 16 GB card) is stated beside the step; the 24 Gb device case shows the chip pays USD 700 once for the same horizon |
|
||||
| One 12 GB card | the same per joule; mines through the 8.5 GiB step, loses proving coexistence at the first step (7.3 + 8.6 GB over 12), leaves mining at 11.5 GiB | the mine-or-prove routing (0.3.21) and the replacement cost; the step offsets (section 11) |
|
||||
| One 16 GB card (the 5080 at its lock, the 9070 XT) | the 5080 at its lock is the honest NVIDIA floor: 1.6x behind the chip machine node-for-node, 1.9x a node ahead; the 9070 XT 4x to 5x behind; mines through every step, loses proving coexistence at 8.5 GiB | the per-dollar edge (3.3x) is the number the project-cost wall must answer; the lock stays the one lever |
|
||||
| One 24 or 32 GB card (the 5090 at its lock, the M5 Max) | 1.8x / 2.1x behind the chip machine; the M5 Max about 1.5x (reported); mines and proves through every step | nothing on the card side moves this; the rotation layers are response capability, not resistance |
|
||||
| A rig | per joule its cards; per dollar 3.3x behind the chip at MSRP, 4.8x at street; over 3 years 2.7x to 3.4x per TH | the issuance trigger and the share-pattern detector (the record's item 4) stay the instruments; the profitability surface is the economics lane's |
|
||||
| A pool user | a chip fleet is a few operators at 1.1 to 1.3 microjoules; the detector on the observer is what tells a pool a chip has arrived | unchanged |
|
||||
| The public claim | the chip line this file supports: "a chip that survives rotation is a GPU-shaped core on the card's own memory; it reads about 1.8x the best honest card per joule on the same node, 2.1x a node ahead, and about 3x per dollar; the rotation calendar changes its firmware, not its cost" | served only on the coordinator's word, after the placed rows |
|
||||
|
||||
## 10. Sources and what is owed
|
||||
|
||||
- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured);
|
||||
the floor programme's tier table (`docs/design/class-v6-rotating-family.md` 10.4: the 5090 at the 1,300 lock 2.33
|
||||
|
|
@ -405,3 +450,109 @@ Owed: the k lane's crossbar, scratch and tile rows (in place and route at 16:0x
|
|||
the HBM3 activate ceiling (unmeasured, the AWS F2 hour); the profitability surface of the review's rule (2) over the
|
||||
lifetime rows of section 6, which this file states as cost per TH and leaves the NPV to the economics lane.
|
||||
|
||||
|
||||
## 11. The data-local and memory-sharing adversary (the ProgPoW audit's threat; the third review, 17:5x BST)
|
||||
|
||||
The threat: split the dataset across processors and move the intermediate computation to the processor nearest
|
||||
the next item, share datasets, keep partial caches, recompute, re-lay the dataset, run several engines off one
|
||||
set-up; price the cheapest combination of moving state, moving data, recomputing and local resources against the
|
||||
live state of class v6, and say whether the necessary state makes data-local execution dear enough to erase any
|
||||
memory-system saving. The cost model, explicit, every term labelled:
|
||||
|
||||
| Term | Per dependent read | Source |
|
||||
|---|---|---|
|
||||
| The live state a hash carries across a read | 64 registers x 32 bits = 2,048 bits (61 of 64 necessary across the chain, the review's reading of the fold rule) plus pc, nonce and era pointer about 64 bits: **about 2,100 bits** | the class program (`verify::fold_words`: every register is consumed) |
|
||||
| The data a read returns, with its address and control | 32 bits of data, 32 of address, about 16 of control: **about 80 bits** at W = 1 (176 at W = 4, 304 at W = 8) | floor lane 3, section 2.1 (modelled) |
|
||||
| Moving bits on one die (global wire) | 1.3 pJ per bit across a 24 mm die (0.65 to 2.6) | floor lane 3 (approximate) |
|
||||
| Moving bits between dies on a package | UCIe 0.5 pJ per bit (0.25 to 0.5) | claimed (UCIe via SNIA), floor lane 3 |
|
||||
| Moving bits between packages on a board | 5 to 10 pJ per bit (a PCB SerDes link; approximate, from memory) | approximate |
|
||||
| A dependent random 32-byte read from GDDR7 / HBM3 | 2.0 / 1.2 nJ (1.5 to 2.6 / 1.0 to 1.5) | chip-model-v3 5.3 (modelled) |
|
||||
| An on-die SRAM read at W = 1 | 0.25 nJ (0.20 to 0.35) including the wire | floor lane 3 (modelled) |
|
||||
| Recomputing one item instead of reading it | 9,360 ops, about 43 nJ on this core at N5 (4.61 pJ per op) and 6.3 nJ on a wired mixer pipeline (chip-model-v3 5.2) | modelled |
|
||||
| The per-window set-up (the 1 GiB build and the host) | 0.7 J per machine per window, 0.85 W and USD 15 of host per machine | section 7.1 |
|
||||
|
||||
The four forms, priced per read at W = 1 (the hash's own width), nominal with the band:
|
||||
|
||||
| Form | What moves | Cost per read | Against the baseline (data to the lane, 80 bits) |
|
||||
|---|---|---|---|
|
||||
| Baseline: the state stays in its lane, the data travels to it | 80 bits of data, address and control | on a die 0.10 nJ (0.05 to 0.21); on a package 0.04; the DRAM read itself 2.0 nJ beside it | 1x |
|
||||
| Data-local on one die: the state travels to the macro holding the item | about 2,100 bits | 2.7 nJ (1.4 to 5.5) of wire per read: **27x the baseline's wire, and 11x the whole GDDR7 read** | 27x |
|
||||
| Data-local across dies on a package (the dataset split over chiplets) | about 2,100 bits over UCIe plus the local read | 1.05 nJ (0.5 to 1.05) of hop per read: half a GDDR7 read, four SRAM reads | 26x the baseline hop |
|
||||
| Data-local across packages on a board (the dataset split over boards) | about 2,100 bits over a SerDes link | 10 to 21 nJ per read: 5x to 10x a GDDR7 read | 260x |
|
||||
| Memory sharing: N engines over one dataset | nothing new: the engines are the lanes in flight the baseline already has (1,172 on the GDDR7 board at the activate ceiling); the set-up is shared at 0.7 J per window | 1x | 1x |
|
||||
| A partial store that recomputes the rest (the pebbling curve, chip-model-v3 5.4 with the two corrections) | a recomputed item costs 9,360 ops | 43 nJ on this core, 6.3 nJ wired, against 2.0 nJ read: **monotone, the full store is the cheapest point at every f** | 3x to 21x per recomputed item |
|
||||
| A hot-item SRAM beside the DRAM (the window-layer distribution: the hottest half of the items serves 72 percent of the reads, adv-cache-2) | an SRAM read for 72 percent of the reads, a DRAM read for the rest | 0.25 x 0.72 + 2.0 x 0.28 = 0.74 nJ per read on the memory side, for USD 250 per GiB of SRAM (about half the dataset) | a capex trade bounded by the GDDR7 row above and the SRAM-die row below; no new form |
|
||||
| An alternative layout (the 16 load sites and their era windows: bank the dataset by site) | nothing per read: every read is still an activate on the device that holds the item | 1x | 1x |
|
||||
|
||||
Reading. (1) The necessary state is the whole argument: a read returns 80 bits and the state that must meet it is
|
||||
2,100, so moving the computation to the data costs 26x the bits of moving the data to the computation, on every
|
||||
medium. On one die that is 2.7 nJ of wire against 0.10; on a package half a DRAM read; across boards five to ten DRAM
|
||||
reads. There is no memory-system saving for data-local execution to erase: the dependent read must be served by the
|
||||
device that holds the item in every form, and what the data's trip to the lane costs (0.10 nJ on a die, 0.04 on a
|
||||
package) is already the cheapest term in the model. So the ProgPoW audit's threat does not apply to a hash whose
|
||||
live state is 26x its read width; it applies to a hash whose state is a few words, which class v6 is not. (2)
|
||||
Memory sharing is the baseline, not an attack: the chip's lanes already share one dataset and one set-up; the
|
||||
per-hash share of the set-up is 10^-15 J and of the host USD 0.15 per machine. (3) Partial stores and recomputation
|
||||
are priced by the pebbling curve and lose at every point; the hot-item SRAM is a capex trade between the two rows
|
||||
of section 6, not a new form. (4) Cumulative memory complexity of one evaluation, stated as the bound a chip must
|
||||
pay: 128 dependent random reads, each an activate and a 32-byte sector, 4 KB of sector traffic and 128 x 2,100 bits
|
||||
of state carried in a lane (never moved); 256 nJ of memory energy per hash on GDDR7 (154 on HBM3, 32 on the SRAM
|
||||
die) plus 0.47 microjoules of shadow on this core at N5. Bandwidth hardness: the GDDR7 board is bound by activates
|
||||
(21.3 G per second over 16 devices, 166 MH/s per board) at 38 percent of its pin bandwidth (682 GB/s of sectors of
|
||||
1,792), so the hard quantity is the activate rate per dollar of devices, not bytes per second, and a chip cannot buy
|
||||
more activates per device than the card has. (5) Bounded by this cost model: data-local execution (26x by the bit
|
||||
count, no physical design needed), memory sharing (the baseline), partial stores (monotone), layouts (per read
|
||||
invariant). Needing physical design before a number is a bound: the HBM activate ceiling (unmeasured; the AWS F2
|
||||
hour), the SRAM die's wire term (0.5 to 2.0 nJ per 64 bytes until a placed macro array exists), the on-package hop
|
||||
(UCIe's 0.5 pJ per bit is a claim), and the data-local form on a 3D-stacked SRAM (a vertical hop at about 0.1 pJ per
|
||||
bit, approximate, would cut the one-die row to about 0.2 nJ, still 2x the baseline and still no saving).
|
||||
|
||||
## 12. The dataset comparison that decides class v7 (the third review, 17:5x BST)
|
||||
|
||||
Two datasets, each with the adversary's burden (a chip with a host keeping it current) and the commodity burden
|
||||
(what every honest node pays at the boundaries), priced on the record's figures:
|
||||
|
||||
| | State-coupled dataset (class v5 and v6: derived from the chain's execution state at the era cut) | Epoch-defined bounded dataset with a published support horizon (derived from the certified checkpoint's hash at the era cut; the size schedule published years ahead) |
|
||||
|---|---|---|
|
||||
| Adversary: sync bandwidth | the day stream and the leaves: 16.5 KB/s to 10,000 members, 45 MB per member per window today (the record 2a.2); grows with the state | the seed: 32 bytes per era; the schedule: a constant |
|
||||
| Adversary: update cost | the rebuild 0.7 J per machine per window (0.0001 percent of its energy); the host's execution of every block (85 W, USD 1,500 per 100 machines: 0.85 W and USD 15 per machine, under 1 percent) | the rebuild only (the same 0.7 J); a light client following checkpoints (bytes, watts of nothing) |
|
||||
| Adversary: storage | the state (hundreds of MB today, unbounded) on the host; the dataset on the machine | the dataset only |
|
||||
| Adversary: adversarial state growth | raises the HOST's cost (storage, re-execution), never the dataset's size (the atom folds the state into a dataset of the schedule's size); the growth is paid in gas by whoever causes it | none |
|
||||
| Adversary: what is excluded | the stale machine (its dataset wrong the moment its host is) and the recompute chip; not a specialised machine with a host (the review's rule 3), which pays under 1 percent | the stale machine (a chip missing the epoch flip mines a dead object, as a stale card does); the recompute chip is excluded by the pebbling curve, not by the dataset; the seed is unknown before the checkpoint, so nothing precomputes |
|
||||
| Commodity: the boundaries | today's eight crossing faults were all state boundaries: the crossing, the partition, the snapshot, the cold start (the 15:42Z deep-reorg reset, the pruning-point anchor, the snapshot resumed under the next object, the stale stream); every family epoch is a state-agreement event (layer 3, section 3) | one boundary kind: the checkpoint's hash, which every node holds by the chain's own rule; no stream, no snapshot, no re-execution on the mining path |
|
||||
| Commodity: re-execution per node | about 10 minutes from genesis on Devnet 3, hours on the shared devnet, and the dataset is wrong until it is done | none for mining (execution continues for the EVM, decoupled from the hash) |
|
||||
| Commodity: proving coexistence | unchanged by the coupling (set by the dataset's size and the prover's 8.6 GB) | the same |
|
||||
| Commodity: the cold start and the partition | a node that missed the lead serves nothing for a window; a partition's two sides derive two datasets | a node derives the dataset from any header it accepts; a partition's two sides derive the same dataset while they share the checkpoint |
|
||||
|
||||
The growth schedule 5.5 / 8.5 / 11.5 GiB, per step (the hash lane's VRAM rows: the dataset needs about 1.3x its size
|
||||
in device memory with the working set; the tier room 75 percent of card memory, 50 percent of Apple unified; the
|
||||
prover 8.6 GB beside it):
|
||||
|
||||
| Step | Adversary burden | Commodity burden: who leaves mining | Who loses proving coexistence | Owners' replacement cost (approximate street prices, October 2026) | The 24 Gb GDDR7 case |
|
||||
|---|---|---|---|---|---|
|
||||
| 5.5 GiB at the v6 epoch | GDDR7 board: 0 (16 x 2 GB holds it); SRAM die: 3 reticles, USD 1,500 of ticket | the 6 GB and 8 GB tiers (25 percent by count): about 7.3 GiB of device memory needed | the 12 GB tier (22 percent): 7.3 + 8.6 over 12; it mines or proves | an 8 GB card to a 16 GB card: USD 350 to 450 (a 16 GB 5060 Ti or 9060 XT class) | none needed |
|
||||
| 8.5 GiB two years on | GDDR7 board: 0 (the 32 GB board holds it); the SRAM die 5 reticles, USD 2,500 | the 10 and 11 GB tiers (9 percent) and the 16 GB unified Mac (about 8 GiB of room): about 11 GiB needed | the 16 GB tier (28 percent): 11 + 8.6 over 16; it mines or proves | a 10 or 12 GB card to a 16 GB card: USD 350 to 450; a 16 GB Mac has no upgrade | none needed |
|
||||
| 11.5 GiB at four years | GDDR7 board: 0; the SRAM die 6 reticles, USD 3,000 | the 12 GB tier (22 percent): about 15 GiB needed | the 24 GB tier mines and proves (15 + 8.6 under 24) | a 12 GB card to a 16 GB card: USD 350 to 450 | none needed |
|
||||
| The 24 Gb (3 GB) GDDR7 devices (Micron ended 2 GB production, 3 GB at USD 60 to 70 each, the TrendForce note) | a 16-device board at 48 GB for USD 1,000 of memory instead of 320: the adversary over-provisions ONCE for the whole horizon, USD 9.9 per MH/s instead of 4.84, still 1.6x the card per dollar and unchanged per joule | the 50-series cards on 3 GB devices (the 5090 32 GB, the 5080 Super class at 24 GB) carry the schedule to year 4 and beyond; the schedule retires the 8 and 12 GB tiers, not the chip | | | the schedule's one effect on the chip is a USD 700 memory ticket, paid once |
|
||||
|
||||
Reading: the schedule is a commodity-side cost at every step (a quarter of the cards by count at 5.5 GiB, a further
|
||||
third losing proving coexistence at 8.5, the 12 GB tier out at 11.5; USD 350 to 450 per displaced owner) and a chip-side
|
||||
cost only on the SRAM die, and the 24 Gb device removes even the DRAM board's small ticket. The state coupling buys,
|
||||
against the chip, the exclusion of a stale machine (which the epoch seed also buys) and a host at under 1 percent of
|
||||
the machine's cost; it costs the honest side every state boundary the devnet has crossed this week.
|
||||
|
||||
The recommendation, five lines:
|
||||
1. Class v7 derives the dataset from the certified checkpoint's hash at the era cut (an epoch-defined bounded
|
||||
dataset), with the size schedule published at genesis as the support horizon; the execution state stays in the
|
||||
headers for the EVM and leaves the hash.
|
||||
2. The dataset floor schedule stays (5.5 / 8.5 / 11.5 GiB, steps offset from the family flips), served as a ticket
|
||||
on an SRAM die and as a retirement of the 8 GB tier at the first step and the 12 GB tier at the third, with the
|
||||
replacement cost per owner stated; the 24 Gb device is priced in as the DRAM chip's one-time USD 700.
|
||||
3. The family bank and its 180-day calendar stay, served as response capability with a zero obsolescence credit
|
||||
per transition (section 8), never as energy or economic resistance.
|
||||
4. The chip line served is the complete machine's: 1.8x per joule node-for-node and 2.1x a node ahead on the
|
||||
GDDR7 board against the Blackwell tier at its lock, 3x against the cohort, 3.3x per dollar; the hold is the
|
||||
project cost against miner revenue (the profitability surface, the economics lane).
|
||||
5. The data-local threat is closed by the live state (26x the bits of the read; no saving to erase) and needs no
|
||||
design change; the three terms that still need physical design before they are bounds (the HBM activate
|
||||
ceiling, the SRAM die's wire, the on-package hop) are named in section 11 and owed.
|
||||
|
|
|
|||
Loading…
Reference in a new issue