multi-family adversary: the routed clock relaxation, the tier table, the statement, the data-local cost model and the dataset comparison

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 2db208912 (2db208912e5fc0656930aaabd7db74ec2a68ca8e) for the box mirror master; left on the branch: tools/chip-model/mf/Makefile tools/chip-model/mf/flow/collect.py tools/chip-model/mf/flow/mf32base.mk tools/chip-model/mf/flow/mf32base.sdc tools/chip-model/mf/flow/mf32full.mk tools/chip-model/mf/flow/mf32full.sdc tools/chip-model/mf/flow/mf8base.mk tools/chip-model/mf/flow/mf8base.sdc tools/chip-model/mf/flow/mf8full.mk tools/chip-model/mf/flow/mf8full.sdc tools/chip-model/mf/tb/tb_mf_common.vh
This commit is contained in:
igneum-labs 2026-10-08 15:26:45 +00:00
parent 30ba24bea1
commit d1d2946038

View file

@ -82,7 +82,12 @@ utilisation (30 for the 32-lane core) with the macros placed by the flow's macro
global and detailed placement, CTS, global and detailed routing, OpenRCX parasitics. Power is OpenSTA `report_power`
under the VCD of a random-input gate-level simulation of the netlist (iverilog; every instruction field drawn by
`$random`, the window initialised with random words, a random returned word on every load), with a propagated 0.5
activity as the cross-check. Synthesis-only rows (no wires, no clock tree) are marked; placed rows carry the SPEF.
activity as the cross-check. Synthesis-only rows (no wires, no clock tree) are marked; placed rows carry the SPEF. The routed runs clock at
12,000 ps with the ABC target held at 1,500 ps: the unpipelined execute path (the multiplier, the fold and the
lane reduction) is 10.6 ns at ASAP7 TC, and at 1,500 ps the flow's timing repair spent its time on a path the
adversary would pipeline instead (two to three registers per lane, about 0.1 to 0.2 pJ per lane-op, inside the
band). Energy per op does not depend on the period; the leakage term does, and the collector restates it at the
1,500 ps equivalent for the routed rows (both are in `table.csv`).
### 2.2 The activity and the steady state
@ -375,9 +380,49 @@ and this is its price: under 1 W and USD 15 per machine.
## 8. Energy resistance, economic resistance and response capability, stated separately
STATEMENT
Three separate things, each with its own number and its own holder.
## 9. Sources and what is owed
**Energy resistance** is a property of the memory system and the shadow, not of the calendar. Against the
adversary's re-optimised chip the complete GDDR7 machine reads 1.8x the 5090 at its lock per joule node-for-node
(1.5x to 2.1x), 2.1x a node ahead; 1.6x and 1.9x against the 5080 at its lock; 2.8x and 3.3x against the cohort
card; the N2 SRAM die 3.1x and 4.2x. The bank moves these by 5 percent in the honest side's favour (the chip's
shadow energy +11 percent for carrying 18 families instead of 10) and no more; the register window moves them by
0.03 to 0.13 of k and no more. What holds the per-joule number is the shadow's size on the card's own operating
point and the memory system's activate ceiling; what would move it is a memory arrangement the chip cannot buy, and
section 6 says there is none: the chip's cheapest memory is the card's own.
**Economic resistance** is capex per sustained MH/s and the project cost against the chain's revenue. The chip's
machine costs USD 4.84 per MH/s against the card's 16 to 18, so over a 3-year life it mines at 0.080 USD per TH
against 0.22 to 0.27: 2.7x to 3.4x, of which electricity is the smaller half. The calendar does not shorten that
life: no transition in the bank retires the chip, so the 3-year stress life holds in full and the withdrawn headline
("dies within an epoch, under 1x over its life") stays withdrawn. What holds the economics is the project (USD 20 M
to 75 M for the GDDR7-board chip, floor lane 5; USD 100 M to 500 M for the SRAM die, claimed) against the miner
revenue the chain pays, which is the profitability surface the review asks for and the economics lane owns; the
dataset floor is a ticket on the SRAM die only (USD 1,000 per step) and nothing on the DRAM board (USD 60 per step).
**Response capability** is what the bank actually buys: the chain can change its object every 180 days without a
release, and a chip that carries the bank follows by firmware at a per-transition loss of -3.7 to +0.7 percent of
its shadow energy (zero credited). A family outside the bank costs that chip an emulation penalty of 2 to 6 percent
of the shadow at 4 points, which every card that predates the release pays too; a structural change (a new read
atom, a new derivation) costs the chip a host update and the cards a release. So the bank is a response channel
whose value per event is the per-transition loss column, not a chip retirement; it is worth keeping for what it is
(no fork for a family change, a fixed-function datapath dead on day one, section 7's table) and must not be served as
energy or economic resistance.
## 9. Consequences per tier (the standing rule)
| Tier | What the rows mean for it | What is being done |
|---|---|---|
| Home miner, one 8 GB card | Against the adversary's machine the cohort card sits 2.8x to 3.3x behind per joule and 3.7x per dollar; the bank does not change that, the operating-point lock does not reach this tier (no Blackwell lever on Ampere; Ada's lock is worth 1.2x); the first dataset step (5.5 GiB) retires this card from mining | the schedule's replacement cost (USD 350 to 450 to a 16 GB card) is stated beside the step; the 24 Gb device case shows the chip pays USD 700 once for the same horizon |
| One 12 GB card | the same per joule; mines through the 8.5 GiB step, loses proving coexistence at the first step (7.3 + 8.6 GB over 12), leaves mining at 11.5 GiB | the mine-or-prove routing (0.3.21) and the replacement cost; the step offsets (section 11) |
| One 16 GB card (the 5080 at its lock, the 9070 XT) | the 5080 at its lock is the honest NVIDIA floor: 1.6x behind the chip machine node-for-node, 1.9x a node ahead; the 9070 XT 4x to 5x behind; mines through every step, loses proving coexistence at 8.5 GiB | the per-dollar edge (3.3x) is the number the project-cost wall must answer; the lock stays the one lever |
| One 24 or 32 GB card (the 5090 at its lock, the M5 Max) | 1.8x / 2.1x behind the chip machine; the M5 Max about 1.5x (reported); mines and proves through every step | nothing on the card side moves this; the rotation layers are response capability, not resistance |
| A rig | per joule its cards; per dollar 3.3x behind the chip at MSRP, 4.8x at street; over 3 years 2.7x to 3.4x per TH | the issuance trigger and the share-pattern detector (the record's item 4) stay the instruments; the profitability surface is the economics lane's |
| A pool user | a chip fleet is a few operators at 1.1 to 1.3 microjoules; the detector on the observer is what tells a pool a chip has arrived | unchanged |
| The public claim | the chip line this file supports: "a chip that survives rotation is a GPU-shaped core on the card's own memory; it reads about 1.8x the best honest card per joule on the same node, 2.1x a node ahead, and about 3x per dollar; the rotation calendar changes its firmware, not its cost" | served only on the coordinator's word, after the placed rows |
## 10. Sources and what is owed
- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured);
the floor programme's tier table (`docs/design/class-v6-rotating-family.md` 10.4: the 5090 at the 1,300 lock 2.33
@ -405,3 +450,109 @@ Owed: the k lane's crossbar, scratch and tile rows (in place and route at 16:0x
the HBM3 activate ceiling (unmeasured, the AWS F2 hour); the profitability surface of the review's rule (2) over the
lifetime rows of section 6, which this file states as cost per TH and leaves the NPV to the economics lane.
## 11. The data-local and memory-sharing adversary (the ProgPoW audit's threat; the third review, 17:5x BST)
The threat: split the dataset across processors and move the intermediate computation to the processor nearest
the next item, share datasets, keep partial caches, recompute, re-lay the dataset, run several engines off one
set-up; price the cheapest combination of moving state, moving data, recomputing and local resources against the
live state of class v6, and say whether the necessary state makes data-local execution dear enough to erase any
memory-system saving. The cost model, explicit, every term labelled:
| Term | Per dependent read | Source |
|---|---|---|
| The live state a hash carries across a read | 64 registers x 32 bits = 2,048 bits (61 of 64 necessary across the chain, the review's reading of the fold rule) plus pc, nonce and era pointer about 64 bits: **about 2,100 bits** | the class program (`verify::fold_words`: every register is consumed) |
| The data a read returns, with its address and control | 32 bits of data, 32 of address, about 16 of control: **about 80 bits** at W = 1 (176 at W = 4, 304 at W = 8) | floor lane 3, section 2.1 (modelled) |
| Moving bits on one die (global wire) | 1.3 pJ per bit across a 24 mm die (0.65 to 2.6) | floor lane 3 (approximate) |
| Moving bits between dies on a package | UCIe 0.5 pJ per bit (0.25 to 0.5) | claimed (UCIe via SNIA), floor lane 3 |
| Moving bits between packages on a board | 5 to 10 pJ per bit (a PCB SerDes link; approximate, from memory) | approximate |
| A dependent random 32-byte read from GDDR7 / HBM3 | 2.0 / 1.2 nJ (1.5 to 2.6 / 1.0 to 1.5) | chip-model-v3 5.3 (modelled) |
| An on-die SRAM read at W = 1 | 0.25 nJ (0.20 to 0.35) including the wire | floor lane 3 (modelled) |
| Recomputing one item instead of reading it | 9,360 ops, about 43 nJ on this core at N5 (4.61 pJ per op) and 6.3 nJ on a wired mixer pipeline (chip-model-v3 5.2) | modelled |
| The per-window set-up (the 1 GiB build and the host) | 0.7 J per machine per window, 0.85 W and USD 15 of host per machine | section 7.1 |
The four forms, priced per read at W = 1 (the hash's own width), nominal with the band:
| Form | What moves | Cost per read | Against the baseline (data to the lane, 80 bits) |
|---|---|---|---|
| Baseline: the state stays in its lane, the data travels to it | 80 bits of data, address and control | on a die 0.10 nJ (0.05 to 0.21); on a package 0.04; the DRAM read itself 2.0 nJ beside it | 1x |
| Data-local on one die: the state travels to the macro holding the item | about 2,100 bits | 2.7 nJ (1.4 to 5.5) of wire per read: **27x the baseline's wire, and 11x the whole GDDR7 read** | 27x |
| Data-local across dies on a package (the dataset split over chiplets) | about 2,100 bits over UCIe plus the local read | 1.05 nJ (0.5 to 1.05) of hop per read: half a GDDR7 read, four SRAM reads | 26x the baseline hop |
| Data-local across packages on a board (the dataset split over boards) | about 2,100 bits over a SerDes link | 10 to 21 nJ per read: 5x to 10x a GDDR7 read | 260x |
| Memory sharing: N engines over one dataset | nothing new: the engines are the lanes in flight the baseline already has (1,172 on the GDDR7 board at the activate ceiling); the set-up is shared at 0.7 J per window | 1x | 1x |
| A partial store that recomputes the rest (the pebbling curve, chip-model-v3 5.4 with the two corrections) | a recomputed item costs 9,360 ops | 43 nJ on this core, 6.3 nJ wired, against 2.0 nJ read: **monotone, the full store is the cheapest point at every f** | 3x to 21x per recomputed item |
| A hot-item SRAM beside the DRAM (the window-layer distribution: the hottest half of the items serves 72 percent of the reads, adv-cache-2) | an SRAM read for 72 percent of the reads, a DRAM read for the rest | 0.25 x 0.72 + 2.0 x 0.28 = 0.74 nJ per read on the memory side, for USD 250 per GiB of SRAM (about half the dataset) | a capex trade bounded by the GDDR7 row above and the SRAM-die row below; no new form |
| An alternative layout (the 16 load sites and their era windows: bank the dataset by site) | nothing per read: every read is still an activate on the device that holds the item | 1x | 1x |
Reading. (1) The necessary state is the whole argument: a read returns 80 bits and the state that must meet it is
2,100, so moving the computation to the data costs 26x the bits of moving the data to the computation, on every
medium. On one die that is 2.7 nJ of wire against 0.10; on a package half a DRAM read; across boards five to ten DRAM
reads. There is no memory-system saving for data-local execution to erase: the dependent read must be served by the
device that holds the item in every form, and what the data's trip to the lane costs (0.10 nJ on a die, 0.04 on a
package) is already the cheapest term in the model. So the ProgPoW audit's threat does not apply to a hash whose
live state is 26x its read width; it applies to a hash whose state is a few words, which class v6 is not. (2)
Memory sharing is the baseline, not an attack: the chip's lanes already share one dataset and one set-up; the
per-hash share of the set-up is 10^-15 J and of the host USD 0.15 per machine. (3) Partial stores and recomputation
are priced by the pebbling curve and lose at every point; the hot-item SRAM is a capex trade between the two rows
of section 6, not a new form. (4) Cumulative memory complexity of one evaluation, stated as the bound a chip must
pay: 128 dependent random reads, each an activate and a 32-byte sector, 4 KB of sector traffic and 128 x 2,100 bits
of state carried in a lane (never moved); 256 nJ of memory energy per hash on GDDR7 (154 on HBM3, 32 on the SRAM
die) plus 0.47 microjoules of shadow on this core at N5. Bandwidth hardness: the GDDR7 board is bound by activates
(21.3 G per second over 16 devices, 166 MH/s per board) at 38 percent of its pin bandwidth (682 GB/s of sectors of
1,792), so the hard quantity is the activate rate per dollar of devices, not bytes per second, and a chip cannot buy
more activates per device than the card has. (5) Bounded by this cost model: data-local execution (26x by the bit
count, no physical design needed), memory sharing (the baseline), partial stores (monotone), layouts (per read
invariant). Needing physical design before a number is a bound: the HBM activate ceiling (unmeasured; the AWS F2
hour), the SRAM die's wire term (0.5 to 2.0 nJ per 64 bytes until a placed macro array exists), the on-package hop
(UCIe's 0.5 pJ per bit is a claim), and the data-local form on a 3D-stacked SRAM (a vertical hop at about 0.1 pJ per
bit, approximate, would cut the one-die row to about 0.2 nJ, still 2x the baseline and still no saving).
## 12. The dataset comparison that decides class v7 (the third review, 17:5x BST)
Two datasets, each with the adversary's burden (a chip with a host keeping it current) and the commodity burden
(what every honest node pays at the boundaries), priced on the record's figures:
| | State-coupled dataset (class v5 and v6: derived from the chain's execution state at the era cut) | Epoch-defined bounded dataset with a published support horizon (derived from the certified checkpoint's hash at the era cut; the size schedule published years ahead) |
|---|---|---|
| Adversary: sync bandwidth | the day stream and the leaves: 16.5 KB/s to 10,000 members, 45 MB per member per window today (the record 2a.2); grows with the state | the seed: 32 bytes per era; the schedule: a constant |
| Adversary: update cost | the rebuild 0.7 J per machine per window (0.0001 percent of its energy); the host's execution of every block (85 W, USD 1,500 per 100 machines: 0.85 W and USD 15 per machine, under 1 percent) | the rebuild only (the same 0.7 J); a light client following checkpoints (bytes, watts of nothing) |
| Adversary: storage | the state (hundreds of MB today, unbounded) on the host; the dataset on the machine | the dataset only |
| Adversary: adversarial state growth | raises the HOST's cost (storage, re-execution), never the dataset's size (the atom folds the state into a dataset of the schedule's size); the growth is paid in gas by whoever causes it | none |
| Adversary: what is excluded | the stale machine (its dataset wrong the moment its host is) and the recompute chip; not a specialised machine with a host (the review's rule 3), which pays under 1 percent | the stale machine (a chip missing the epoch flip mines a dead object, as a stale card does); the recompute chip is excluded by the pebbling curve, not by the dataset; the seed is unknown before the checkpoint, so nothing precomputes |
| Commodity: the boundaries | today's eight crossing faults were all state boundaries: the crossing, the partition, the snapshot, the cold start (the 15:42Z deep-reorg reset, the pruning-point anchor, the snapshot resumed under the next object, the stale stream); every family epoch is a state-agreement event (layer 3, section 3) | one boundary kind: the checkpoint's hash, which every node holds by the chain's own rule; no stream, no snapshot, no re-execution on the mining path |
| Commodity: re-execution per node | about 10 minutes from genesis on Devnet 3, hours on the shared devnet, and the dataset is wrong until it is done | none for mining (execution continues for the EVM, decoupled from the hash) |
| Commodity: proving coexistence | unchanged by the coupling (set by the dataset's size and the prover's 8.6 GB) | the same |
| Commodity: the cold start and the partition | a node that missed the lead serves nothing for a window; a partition's two sides derive two datasets | a node derives the dataset from any header it accepts; a partition's two sides derive the same dataset while they share the checkpoint |
The growth schedule 5.5 / 8.5 / 11.5 GiB, per step (the hash lane's VRAM rows: the dataset needs about 1.3x its size
in device memory with the working set; the tier room 75 percent of card memory, 50 percent of Apple unified; the
prover 8.6 GB beside it):
| Step | Adversary burden | Commodity burden: who leaves mining | Who loses proving coexistence | Owners' replacement cost (approximate street prices, October 2026) | The 24 Gb GDDR7 case |
|---|---|---|---|---|---|
| 5.5 GiB at the v6 epoch | GDDR7 board: 0 (16 x 2 GB holds it); SRAM die: 3 reticles, USD 1,500 of ticket | the 6 GB and 8 GB tiers (25 percent by count): about 7.3 GiB of device memory needed | the 12 GB tier (22 percent): 7.3 + 8.6 over 12; it mines or proves | an 8 GB card to a 16 GB card: USD 350 to 450 (a 16 GB 5060 Ti or 9060 XT class) | none needed |
| 8.5 GiB two years on | GDDR7 board: 0 (the 32 GB board holds it); the SRAM die 5 reticles, USD 2,500 | the 10 and 11 GB tiers (9 percent) and the 16 GB unified Mac (about 8 GiB of room): about 11 GiB needed | the 16 GB tier (28 percent): 11 + 8.6 over 16; it mines or proves | a 10 or 12 GB card to a 16 GB card: USD 350 to 450; a 16 GB Mac has no upgrade | none needed |
| 11.5 GiB at four years | GDDR7 board: 0; the SRAM die 6 reticles, USD 3,000 | the 12 GB tier (22 percent): about 15 GiB needed | the 24 GB tier mines and proves (15 + 8.6 under 24) | a 12 GB card to a 16 GB card: USD 350 to 450 | none needed |
| The 24 Gb (3 GB) GDDR7 devices (Micron ended 2 GB production, 3 GB at USD 60 to 70 each, the TrendForce note) | a 16-device board at 48 GB for USD 1,000 of memory instead of 320: the adversary over-provisions ONCE for the whole horizon, USD 9.9 per MH/s instead of 4.84, still 1.6x the card per dollar and unchanged per joule | the 50-series cards on 3 GB devices (the 5090 32 GB, the 5080 Super class at 24 GB) carry the schedule to year 4 and beyond; the schedule retires the 8 and 12 GB tiers, not the chip | | | the schedule's one effect on the chip is a USD 700 memory ticket, paid once |
Reading: the schedule is a commodity-side cost at every step (a quarter of the cards by count at 5.5 GiB, a further
third losing proving coexistence at 8.5, the 12 GB tier out at 11.5; USD 350 to 450 per displaced owner) and a chip-side
cost only on the SRAM die, and the 24 Gb device removes even the DRAM board's small ticket. The state coupling buys,
against the chip, the exclusion of a stale machine (which the epoch seed also buys) and a host at under 1 percent of
the machine's cost; it costs the honest side every state boundary the devnet has crossed this week.
The recommendation, five lines:
1. Class v7 derives the dataset from the certified checkpoint's hash at the era cut (an epoch-defined bounded
dataset), with the size schedule published at genesis as the support horizon; the execution state stays in the
headers for the EVM and leaves the hash.
2. The dataset floor schedule stays (5.5 / 8.5 / 11.5 GiB, steps offset from the family flips), served as a ticket
on an SRAM die and as a retirement of the 8 GB tier at the first step and the 12 GB tier at the third, with the
replacement cost per owner stated; the 24 Gb device is priced in as the DRAM chip's one-time USD 700.
3. The family bank and its 180-day calendar stay, served as response capability with a zero obsolescence credit
per transition (section 8), never as energy or economic resistance.
4. The chip line served is the complete machine's: 1.8x per joule node-for-node and 2.1x a node ahead on the
GDDR7 board against the Blackwell tier at its lock, 3x against the cohort, 3.3x per dollar; the hold is the
project cost against miner revenue (the profitability surface, the economics lane).
5. The data-local threat is closed by the live state (26x the bits of the read; no saving to erase) and needs no
design change; the three terms that still need physical design before they are bounds (the HBM activate
ceiling, the SRAM die's wire, the on-package hop) are named in section 11 and owed.