Merge ca3-coord 42fd7c9b into master (gate: green on 42fd7c9b, recorded by tools/ci/pre-push.sh; landed on the build mirror)

This commit is contained in:
igneum-labs 2026-10-07 20:36:24 +00:00
commit f3bb75cad3

View file

@ -337,6 +337,26 @@ measurements land.
- The ALU budgets of the M5 Max and the 9070 XT, their power at the hash, and the verifier's cost at N = 100,000
program ops are estimates; the program-length lever is a design item with its own measurements, not a result.
### 5.10 Class v5 on, the shadow at zero: does the state-derived dataset make the shadow unnecessary? (7 October 2026, 21:3x UK, the founder's question "we need a solution, deep research, other methods")
The question: with class v5 (the dataset built from the chain's execution state, refreshed per window) and the latency-shadow work of class v4 set to zero (class v3 energy on every GPU), what edge does the strongest chip keep over an RTX 5090 per joule? If it were at or under about 2x the shadow could come off after v5 and the class v4 premium (145 W on a 5090 at the unlocked core, 88 W at the 1,400 MHz lock, measured 7 October 2026) would vanish.
What class v5 changes for the chip, from `docs/design/class-v5-stored-state.md` sections 2 and 2a: the recompute chip (`f = 0`, the dataset derived on the fly from the cache) and the stateless or stale chip are removed as categories, because every item takes a leaf of the state and the leaves refresh every window. What it does NOT change: the strongest chip was never one of those. It is the `f = 1` stored-dataset chip of 5.5, a GPU's memory system without the GPU, and under class v5 it needs one thing more, the window's leaves, which section 2a.2 prices honestly: one node serves a whole farm, the leaves ship at 16.5 KB/s to 10,000 members today (45 MB per member per window over a 1 Gbit/s WAN before the state is 7,500x today's), and the rebuild on the chip is the same 32 ms per window every GPU pays. The dataset is still derived from a seed and the state, so a central node compresses everything but the state's bytes, and the chip stores the result as before.
The arithmetic, on the model's own figures (5.3: the activate-bound ceilings, 2.0 nJ per random read on GDDR7 and 1.2 nJ on HBM3, the static and controller watts; the 5090 at 136.1 MH/s on 326 W, 0.417 MH/W; 128 loads per hash, the shadow at zero so no ALU beside the memory). The node is a desktop-class CPU with an NVMe and 32 GB at about 85 W (approximate, from memory of such machines; a full node with the EVM executor at 1 block/s), shared by a farm (100 chips: 0.85 W each) or carried by every chip (the attacker's worst case, 85 W each).
| Chip, class v5 on, shadow at zero | MH/s (model) | W with a farm-shared node (0.85 W) | MH/W | Edge over the 5090 per joule | W with a node per chip (85 W) | Edge |
|---|---|---|---|---|---|---|
| GDDR7 `f = 1`, 16 devices (the 5090's own memory without the GPU) | 166 | 78 | 2.12 | 5.1x | 163 | 2.5x |
| HBM3 `f = 1`, one stack | 84 | 28 | 3.02 | 7.2x | 112 | 1.8x |
| HBM3 `f = 1`, eight stacks (an H100-class package) | 666 | 175 | 3.80 | 9.1x | 259 | 6.2x |
Every figure is modelled (arithmetic on cited memory figures, approximate where 5.3 marks it); none is measured; the node's watts are an approximate from memory.
Reading: NO. Class v5 with the shadow at zero leaves the strongest chip at 5.1x (GDDR7) to 9.1x (HBM3, eight stacks) per joule, the class v3 figures of 5.6 less a rounding, because the node is a farm cost and not a chip cost; only a chip forced to carry its own node falls near 2x, and only the small ones (one HBM3 stack at 1.8x, the GDDR7 board at 2.5x), while the eight-stack package stays at 6.2x even with a node per chip. So the shadow (class v4's 100,000 ops per hash in the memory wait, which brings the chip to 2.1x at k = 1 and 3.9x on the claimed X9 core) stays the only lever in this model that reaches the memory-system chip, and the class v4 premium is the price of that lever on today's GPUs. What class v5 buys is different and real: the recompute chip and the stale chip are gone as categories, every miner must hold and follow the chain, and a chip's dataset is wrong the moment its node is. The premium itself has two measured levers tonight: the core-clock lock (57 of 145 W back on the 5090 at 1,400 MHz for 1.4 percent of rate; the knee below 1,400 is the second pass's) and the per-card tune the app lands by itself; what would remove it is a shadow whose work is cheaper per op on a GPU than on a chip core (the research lane's question: a shadow shaped for the GPU's idle datapath at low clock, or a memory-side cost the chip cannot amortise), not the state-derived dataset.
Per tier: a miner on class v4 pays the premium and gets the 2.1x to 3.9x chip ceiling in exchange; on class v5 with the shadow kept the ceiling stays and the dataset is the chain's; on class v5 with the shadow dropped the premium goes and the ceiling returns to 5x to 9x. The decision is the founder's; this section gives the number.
## 6. The per-day derivation (item 2)
6 October 2026, Counter ASIC 3.0 item 2, worker `derive` (`docs/plans/counter-asic-3-derivation.md`; everything