Horizon: lanes 2 (algorithm), 7 (frontier) and 8 (new-pow designs) land, with the summary skeleton

Josh's 6 October 2026 ask: deep backward and forward research across the hash, finality,
economy, network and every shipped surface. This commit carries the first three lanes.

- docs/analysis/horizon/algorithm.md: the chip model on the 6 October numbers (f = 1 GDDR7
  chip 5.7x per joule against the 5090 at class v3, 2.1x at class v4 with k = 1), the FPGA
  lane tightened to 0.30x to 0.47x per watt, the reserve R0 to R8, the reconciled shadow-N
  ladder (section 5.3a) with HBM4 and three verifier brackets, the first measured verifier
  proxy on igneum-build-1 (class v4 5.06 ms cold, dr736 10.51: out), the dataset schedule
  to 2030; model sim/horizon/algorithm/model.py.
- docs/analysis/horizon/frontier.md: sixteen ideas ranked by payoff over difficulty with the
  Monero and Kaspa attacks, prior art cited, the honest never column; model
  sim/horizon/frontier/frontier_model.py.
- docs/analysis/horizon/new-pow.md sections 0 to 4: three new proof-of-work schemes defined,
  reviewed in two personas, scheme A (mining is proving) ruled out on bytes and
  sampleability, B and C in prototype on two rented 4090s; measured rows follow.
- docs/analysis/horizon-2026-10.md: the summary skeleton and the lane table.

Every rental cost cites docs/bench-log.md "Rental cost of hash, 6 October 2026".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-06 20:51:23 +01:00
parent 3f4f71965b
commit 5ff73932e9
9 changed files with 2189 additions and 0 deletions

View file

@ -0,0 +1,37 @@
# Horizon, October 2026: the ranked research across every system
Written 6 to 7 October 2026 by the Horizon coordinator (branch `horizon`, worktree `igneum-wt-horizon`) on Josh's ask of 6 October 2026, 20:4x UK: "deep backward and forward predictive research and modelling for our algo and all of our systems: is there room for improvement, room for more coin utility, algo improvements, security improvements, anything we can do to make a 51% attack impossible, basically creating a level of polish that has not been seen before." Later the same evening: "research anything else that we can research too, predictions, forward thinking, what we can actually do that has not been done or applied, think outside the box", "be revolutionary", and "if we create a new way of hashing or a new way of proof of work to revolutionise the space then that's absolutely fine, I want you to deploy everything to create something that has not ever been done before."
The bar main set, and the bar this document holds every claim to: "impossible" is not available to any proof-of-work chain. The bar is that a majority of hash buys nothing: it cannot reverse what finality locked, cannot forge a proof the nodes re-execute, cannot change a rule without 95 percent signalling, and loses more than it earns. Every claim here is a number with a model or a simulation behind it, priced in rented hash at the measured 6 October rate (`docs/bench-log.md`, "Rental cost of hash, 6 October 2026": USD 0.0117 per MH/s-hour, the live devnet at 1.16 GH/s), with its consequence per tier and what to build.
## The lanes
| Lane | File | State |
|---|---|---|
| 1 consensus-security | `docs/analysis/horizon/consensus-security.md`, and the paper `docs/analysis/51-percent.md` | running |
| 2 algorithm | `docs/analysis/horizon/algorithm.md` | running |
| 3 finality-and-weight | `docs/analysis/horizon/finality-and-weight.md` | running |
| 4 economy-and-utility | `docs/analysis/horizon/economy-and-utility.md` | running |
| 5 network | `docs/analysis/horizon/network.md` | queued (waits for the 10 bps rows of `block-rate-devnet2.md`) |
| 6 polish | `docs/analysis/horizon/polish.md` | queued |
| 7 frontier | `docs/analysis/horizon/frontier.md`, model `sim/horizon/frontier/frontier_model.py` | landed 6 Oct 2026, 21:1x UK |
| 8 new-proof-of-work | `docs/analysis/horizon/new-pow.md`, prototypes in `proto-newpow/` | running (designs tonight, measured rows tomorrow afternoon UK) |
## 1. One page for Josh
(Written last, from the lanes.)
## 2. The ranked list: top 25 across every lane
| Rank | Item | Lane | Evidence | Model | Hours | Consequence per tier | What to build | Gate |
|---|---|---|---|---|---|---|---|---|
## 3. What a 51 percent attacker can and cannot do
(Summary of `docs/analysis/51-percent.md`.)
## 4. Per lane: the three biggest findings
## 5. What was not run, and why
## 6. Rules and corrections for main

View file

@ -0,0 +1,350 @@
# Horizon lane 2: the shipped hash and its class system, refined
6 October 2026, evening UK. Lane 2 (algorithm) of the Horizon programme. Worktree `/Users/joshm/Projects/igneum-wt-horizon`, branch `horizon` on master 3f4f719. Scope: the shipped lottery hash and its classes (v2, v3 live, the v4 candidate `mx8+sh256x27`, the reserve R0 to R8, the era draw, the dataset schedule). No new puzzle is proposed here (lane 8 owns that). Every chip figure is arithmetic on cited figures and is approximate; every GPU figure names its bench-log entry or analysis file; the one new measurement is the verifier proxy on igneum-build-1 (section 2).
What was read, in full unless marked: `docs/spec/01-lottery-hash.md`, `docs/spec/04-seeds-and-vdf.md`, `docs/analysis/chip-model-v3.md`, `docs/analysis/asic-resistance-history.md`, `docs/analysis/latency-shadow-2026-10-06.md`, `docs/analysis/m16-recompute-attacker-2026-10-05.md`, `docs/analysis/sram-mirror.md`, `docs/analysis/int8-matrix-family.md`, `docs/analysis/weak-program-census-2026-10-03.md` (section 1), `docs/analysis/scratch-soundness.md` (section 0), `docs/analysis/card-lifetime-2026-10-05.md`, `docs/plans/counter-asic-3-status.md`, `docs/plans/counter-asic-3-reserve.md`, `docs/plans/counter-asic-3-derivation.md`, `docs/plans/counter-asic-2.md`, `docs/plans/epoch-length.md` (sections 11 and 12), `docs/plans/era-layout.md`, `docs/plans/mixer-x4.md` (sections 6.2 to 6.5), `docs/plans/read-width.md`, `docs/fud-ledger.md` (M1 to M31 headings; M9, M16, M17, M18 in full; M28 heading), `docs/bench-log.md` (entries "Counter ASIC 2.0, the numbers", "the 9070 XT on the eGPU", the 6 October item 8, item 6 and gate-run entries, "Rental cost of hash, 6 October 2026"), `igneum-pow/src/generator.rs` (the class constants, `ShadowClass`, `LoadClass::{V2, MX4, MX8, DR736}`), `igneum-pow/src/memhard.rs` (`Shape`, the growth rule, `round_key`, `derive_items_mask`), `igneum-pow/src/main.rs` (`bench`), `docs/analysis/prover-tiers-real-cards.md` (the 11 rented cards), and `/Users/joshm/Projects/igneum-wt-gpu-fleet/docs/analysis/block-rate-devnet2.md` (RUN_A and RUN_B are still placeholders at 21:30 UK; nothing from it is used).
## 1. The six questions and the one-line answers
| # | Question | Answer in one line |
|---|---|---|
| 1 | Chip model on the 6 October numbers; the FPGA lane | The stored-dataset chip (f = 1, GDDR7) reads 5.7x per joule against the 5090 bench row and 1.7x against the M5 Max at class v3; at class v4 and k = 1 those fall to 2.1x and 0.9x. Per dollar it is 56x under rented hash and 5.4x under an owned 5090 per MH/s-hour. The FPGA soft overlay tightens to 0.30x to 0.47x per watt: the measured 2.4 G reads/s equals the JEDEC tFAW ceiling of a 2-stack HBM2 part, so the 1.9x bank-bound row is unreachable on any FPGA that exists. AWS F2 at USD 1.98 an hour carries the exact HBM2 subsystem and can measure it |
| 2 | Reserve R0 to R8 | Every reserve family together costs a chip about 8 to 14 adders per lane, about USD 4 of N5 silicon on a 14,000-lane array; none moves the per-joule edge by over 10 percent. The reserve's value is obsolescence of a datapath taped out against class v3, and since the order is public that value is zero against a chip that ships with every block. Recommended order: R0 derive at dr368 (dr736 fails the gate proxy), R1 shfla, R2 perm, R3 popc and clz, R4 bfe, R5 shifts, R6 sel, R7 andn, R8 mm8 at era 8 by the rule, no exception |
| 3 | A class v5 from the shadow | Option (i) is void: the v4 shadow already consumes loaded data (its registers hold the dataset words of the iteration), so a chip precomputes nothing today. Option (ii), a shuffle-heavy shadow mix, raises the attacker's k floor from about 0.32 to 0.46 (approx), not to 1. Option (iii): the one N every owned card holds within 5 percent is 130,000 counted ops (the M5 Max's point); it buys 2.1x to 1.7x against the 5090 at k = 1 for 4.8 percent of the Mac's rate; the verifier at 130,000 is 4.9 ms on the box proxy and 8.5 ms on the half-core proxy, inside the gate |
| 4 | The era draw's randomness | Forging the certified checkpoint the era VDF reads costs 20 days of 100 percent hash: USD 5,600 at 1 GH/s, USD 5.6M at 1 TH/s rented, and buys one draw of a space whose spread is 0.8 to 3.2 percent of hash rate per card and 0 percent for a chip. Re-rolling by withholding needs a 1,800x faster VDF. The draw buys nothing against a chip; the public reserve order means a chip is taped out with every block. What would cost a chip is work (N) and the honest card's own watts, not unpredictability |
| 5 | The 2019-class verifier | Measured proxy tonight: a Zen 4 core at 3.8 GHz with server DRAM reads 2.0x to 2.2x the M5 Max core (v4 4.90 ms steady, 5.06 cold; dr736 9.76 steady, 10.51 cold: FAIL); with both SMT siblings busy (the pessimistic bracket) v4 8.23 ms, dr368 8.16, dr736 15.5. The 2.5x rule holds within 15 percent. The gate protects a node at 1 to 10 percent of one core per block rate, an 18-minute IBD, and a header-flood cost of 100 bogus headers per second per core |
| 6 | Dataset growth to 2030 | Hold option (b): 2 GiB at genesis, 4 GiB at year 4, 8 GiB at year 12. The 8 GB tier (26.7 percent of Steam today, about 0 by 2030 on the trend) mines to year 12 anyway; the 12 GB tier loses mine-and-prove compressed at the year-4 step and the 8 GB tier loses core-only at genesis, and both are the prover's footprint, not the dataset's. Growth is for the SRAM reticle (1.6 GiB per reticle at N5) and GPU L2, not for the HBM chip |
## 2. Method
| Item | What was done | Where |
|---|---|---|
| The model | One Python script, every table in this file; inputs listed with source and label | `sim/horizon/algorithm/model.py`, `README.md` beside it |
| Verifier proxy | `igneum-pow` built on igneum-build-1 through `tools/build-remote.sh` from the crate directory (`IGNEUM_AGENT=horizon`, build slot build-0, 10 s wall, sccache miss 1, artefact 991,352 bytes, sha256 6d2867...); `igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <c> --warps 50` for v2, mx8, mx8+sh256x27, dr368, dr736 under `nice -n 19 taskset -c 40` (one core), then the same on cores 40 and 88 at once (the two SMT siblings of one physical core: `thread_siblings_list` 40,88); load average 1.3 before, 2.3 after; `scaling_cur_freq` read 3,799,885 kHz during a run (the governor is schedutil with `scaling_max_freq` 2,750,000 but the core boosted to 3.8 GHz; `cpupower` needs root, so no fixed low clock was possible); the cache fill 361 ms on that core | section 5.5 |
| Not run | A Mac packbench ladder for v5 (the data-dependent and shuffle-heavy shadow variants do not exist as kernel text, so nothing could be measured; the plan is in 5.3); any PC job; any FPGA | section 7 |
| External figures | Steam Hardware Survey September 2026 (store.steampowered.com/hwsurvey, read 6 October 2026); ICCAD 2021 "Demystifying the Characteristics of High Bandwidth Memory for Real-Time Systems" Table I (JEDEC HBM2 timings in cycles at 2,133 MT/s, text extracted with pdftotext); Intel HBM2 IP user guide (16 AXI pseudo-channel ports per stack); AWS F2 pricing (f2.6xlarge USD 1.98 an hour on-demand, USD 0.66 spot, one VU47P with 16 GB HBM2) | section 5.1 |
## 3. Evidence
### 3.1 Honest cards, measured (class v3 unless said)
| Card (tier) | MH/s | W | uJ per hash | Source |
|---|---|---|---|---|
| RTX 5090, PC 2 bench, 431 W cap, the control | 132.2 | 350 | 2.65 | `latency-shadow-2026-10-06.md` s5; bench-log item 8 |
| RTX 5090, PC 1 app, Ember run 5 | 127.4 | 310 | 2.43 | `counter-asic-3-status.md` s7 item 1 |
| RTX 5090, app 5 October | 124 | 290 | 2.34 | `miner-eff` record, cited in latency-shadow s3 |
| RTX 5090, rented Vast pod, untuned | 98.5 | 258 | 2.62 | `prover-tiers-real-cards.md` |
| RTX 4090, rented | 52.3 | 183 | 3.50 | same |
| RTX 3090, rented | 37.8 | 229 | 6.05 | same |
| RTX A5000, rented | 47.7 | 223 | 4.67 | same |
| RTX 4070, PC 1, 1,860 MHz lock, 160 W cap | 30.95 | 79.5 | 2.57 | status item 8, 4070 rows |
| RTX 4070, rented | 25.0 | 91 | 3.65 | prover-tiers |
| RTX 5070, rented | 41.9 | 137 | 3.27 | prover-tiers |
| RTX 3060, rented | 23.8 | 104 | 4.36 | prover-tiers |
| RTX 3080, rented | 40.8 | 205 | 5.02 | prover-tiers |
| RTX 4060 Ti 16 GB, rented | 17.6 | 72 | 4.11 | prover-tiers |
| RTX 4060 Ti 8 GB, rented | 19.1 | 73 | 3.81 | prover-tiers |
| RTX 4060, rented | 17.1 | no reading (0.0 W logged) | n/a | prover-tiers |
| RX 9070 XT, PC 1 | 18.9 | 199 to 203 | 10.6 | bench-log 9070 XT telemetry; status 3a |
| Apple M5 Max, GPU + DRAM channels | 27.08 | 21.0 | 0.78 | latency-shadow s3 |
| Apple M5 Max, package (approx) | 27.08 | 38 | 1.40 | `ember-tune.md`, approximate |
At class v4 (`mx8+sh256x27`, measured): the 5090 131.95 MH/s at 431 W (3.27 uJ, the cap binds), the M5 Max 26.67 at 37.2 W (1.39), the 4070 31.08 at 109 W (3.51), the 9070 XT 19.29 at owed watts. Every other card's v4 energy below is modelled as the card's marginal ALU energy times N (NVIDIA 11 pJ per counted op measured on the 5090, Apple 6.9 measured on the M5 Max, AMD taken as NVIDIA's, approximate).
### 3.2 Chips (model, approximate; `chip-model-v3.md` s5.4, latency-shadow s6)
| Chip class | MH/s | W at v3 | uJ, v3 | uJ at v4, k = 0.3 / 0.5 / 1 / 1.5 | $ silicon + memory, v3 / v4 | $ per MH/s |
|---|---|---|---|---|---|---|
| f = 0 on-die recompute, 256 MiB SRAM, x8 | 41.7 | 53.7 | 1.29 | 1.62 / 1.84 / 2.39 / 2.94 | 700 / 740 | 16.8 |
| f = 1 stored dataset, GDDR7, 16 devices | 166.4 | 77.6 | 0.466 | 0.80 / 1.02 / 1.57 / 2.12 | 470 / 510 | 2.8 |
| f = 1, HBM3 one stack | 83.6 | 26.8 | 0.321 | 0.65 / 0.87 / 1.42 / 1.97 | 550 / 590 | 6.6 |
| f = 1, HBM3 eight stacks | 666 | 174 | 0.262 | 0.59 / 0.81 / 1.36 / 1.91 | 2,650 / 2,690 | 4.0 |
`k` is the chip core's energy per counted op over the 5090's measured 11 pJ. The f = 0 chip's measured stand-in (M16 round 2, the inline kernel inside the 5090's L2) ran 33.9 MH/s at 431 W: 0.256x per chip and 5x worse per joule than honest, so the f = 0 row above is the model's ceiling for that class, not a measurement.
### 3.3 The verifier proxy, measured tonight (igneum-build-1, EPYC 9454P, one core at 3.8 GHz, nice 19)
| Class | M5 Max core, quiet | 2.5x rule (approx) | Box one core, steady / cold | Box over Mac | Box half-core (SMT sibling loaded) | 10 ms gate |
|---|---|---|---|---|---|---|
| v2 | 0.60 | 1.5 | 1.13 / 1.30 | 1.88x | not run | pass |
| mx8 (class v3) | 2.06 | 5.2 | 4.52 / 4.67 | 2.19x | 7.56 | pass |
| mx8+sh256x27 (class v4) | 2.33 | 5.8 | 4.90 / 5.06 | 2.10x | 8.23 | pass |
| dr368 | 2.69 | 6.7 | 4.72 / 5.32 | 1.75x | 8.16 | pass, thin on the half-core |
| dr736 | 4.88 | 12.2 | 9.76 / 10.51 | 2.00x | 15.49 | FAIL (cold run over 10 ms on one core; 15.5 on the half-core) |
Raw lines (the Mac figures are `counter-asic-3-derivation.md` 5.1 and the gate-run entry): box one core `CPU verify: 1.128 / 4.516 / 4.901 / 4.716 / 9.763 ms per 32-lane warp, avg of 50`; cold `warp base 0: 1.305 / 4.670 / 5.061 / 5.324 / 10.509 ms`; half-core `7.562 / 8.234 / 8.162 / 15.491` (cpu 40) and `7.556 / 8.227 / 8.168 / 15.479` (cpu 88); every lane-0 and lane-31 hash equal to the Mac's vectors (`42246ba99fc58e4f`, `19b56348bc85304d`, the dr736 `e23d389f3eea0c83`). Cache fill 360.5 to 364.7 ms on the box core against 175 to 181 on the M5 Max.
### 3.4 Installed base (Steam Hardware Survey, September 2026, cited)
| VRAM | Share | Trend (cited: 8 GB 35.03 percent in August 2025, 33.66 in September 2025; 12 GB 19.30 in August 2025) |
|---|---|---|
| 512 MB to 4 GB | 14.06% | falling |
| 6 GB | 5.12% | falling |
| 8 GB | 26.71% | about -7 points a year |
| 10 to 11 GB | 2.74% | flat |
| 12 GB | 13.06% | about -6 points a year |
| 16 GB | 27.21% | rising; overtook 8 GB in August 2026 |
| 20 to 24 GB | 7.93% | rising slowly |
| 32 GB | 1.41% | rising slowly |
| 64 GB | 0.50% | new row |
### 3.5 Rental and ownership cost (bench-log "Rental cost of hash, 6 October 2026", measured)
USD 0.0117 per MH/s-hour on RunPod community pods (1,748 MH/s for USD 20.44 an hour); the 8x 4090 rig USD 0.0129; a 5090 pod USD 0.41 to 0.74 an hour for 98 to 128 MH/s.
## 4. Model
| Formula | Inputs (label) |
|---|---|
| Energy per hash E = W / rate | measured watts and MH/s per card; chip W from `chip-model-v3.md` 5.4 (approx) |
| Chip energy at class v4: E_v3 + N x 11 pJ x k | N = 100,000 counted ops (measured class v4), 11 pJ = the 5090's marginal per op (measured), k free |
| Edge per joule = E_card / E_chip | both sides above |
| Hourly cost per MH/s = price / (2 years x rate) + E x 3.6e9 x USD 0.10 per kWh | prices approximate (launch list from memory, labelled); chip $ from 5.4; electricity approx |
| FPGA random-read ceiling = min(banks / tRC, channels x 4 / tFAW, channels / tRRD) | 2 stacks, 16 channels, 32 pseudo-channels, 16 half-banks each (approx); tRC 48 cycles, tFAW 30, tRRD 6 at 1,066 MHz (ICCAD 2021 Table I, JEDEC HBM2); latency 137.8 ns (Shuhai, measured) |
| Reads in flight = rate x latency; per watt = rate / board W | U55C 115 to 150 W, U280 225 W (datasheets) |
| Verifier add per warp = shadow instructions / 1,000 x slope | slope 3.2 us (M5 Max, measured), 7.0 us (box one core, this lane), 12.1 us (box half-core, this lane) |
| Checkpoint forgery cost = network MH/s x 480 h x USD 0.0117 | 20 days of 100 percent hash to 2/3 weight (CLAUDE.md headline, from the finality sim) |
| Tier fit at a dataset step: miner resident = dataset + 192 MiB; prover peak beside the miner from `prover-tiers-real-cards.md` scaled by the dataset's growth | usable VRAM 75 percent (mine-only), 98 percent headless (the prover rows) |
## 5. Results
### 5.1 Task 1: the chip model on the 6 October numbers, and the FPGA lane
Per joule and per dollar against every honest card (the full table with every card is the `chip` section of the model; the rows that decide things):
| Card (tier) | uJ v3 / v4 | f = 0 chip edge, v3 / v4 at k = 1 | f = 1 GDDR7 edge, v3 / v4 at k = 0.3 / 0.5 / 1 / 1.5 | f = 1 HBM3 one stack, v3 / v4 at k = 1 | $ per MH/s (card, approx) | USD per MH/s-hour owned |
|---|---|---|---|---|---|---|
| RTX 5090 bench (32 GB) | 2.65 / 3.27 | 2.1x / 1.4x | 5.7x / 4.1x / 3.2x / 2.1x / 1.5x | 8.3x / 2.3x | 15.1 | 0.00113 |
| RTX 5090 app, Ember (32 GB) | 2.43 / 3.53* | 1.9x / 1.5x | 5.2x / 4.4x / 3.5x / 2.3x / 1.7x | 7.6x / 2.5x | 15.7 | 0.00114 |
| RTX 4090 rented (24 GB) | 3.50 / 4.60* | 2.7x / 1.9x | 7.5x / 5.8x / 4.5x / 2.9x / 2.2x | 10.9x / 3.2x | 30.6 | 0.00210 |
| RTX 3090 rented (24 GB) | 6.05 / 7.15* | 4.7x / 3.0x | 13.0x / 9.0x / 7.0x / 4.6x / 3.4x | 18.9x / 5.0x | 39.7 | 0.00287 |
| RTX 4070 PC 1 tuned (12 GB) | 2.57 / 3.51 | 2.0x / 1.5x | 5.5x / 4.4x / 3.5x / 2.2x / 1.7x | 8.0x / 2.5x | 17.7 | 0.00127 |
| RTX 5070 rented (12 GB) | 3.27 / 4.37* | 2.5x / 1.8x | 7.0x / 5.5x / 4.3x / 2.8x / 2.1x | 10.2x / 3.1x | 13.1 | 0.00108 |
| RTX 3060 rented (12 GB) | 4.36 / 5.46* | 3.4x / 2.3x | 9.4x / 6.9x / 5.4x / 3.5x / 2.6x | 13.6x / 3.8x | 13.8 | 0.00123 |
| RTX 4060 Ti 8 GB rented (8 GB) | 3.81 / 4.91* | 3.0x / 2.1x | 8.2x / 6.2x / 4.8x / 3.1x / 2.3x | 11.9x / 3.5x | 20.9 | 0.00157 |
| RTX 4060 Ti 16 GB rented (16 GB) | 4.11 / 5.21* | 3.2x / 2.2x | 8.8x / 6.5x / 5.1x / 3.3x / 2.5x | 12.8x / 3.7x | 28.4 | 0.00203 |
| RX 9070 XT (16 GB AMD) | 10.6 / 11.7* | 8.3x / 4.9x | 22.8x / 14.7x / 11.5x / 7.5x / 5.5x | 33.2x / 8.3x | 31.7 | 0.00287 |
| Apple M5 Max, GPU + DRAM | 0.78 / 1.39 | 0.6x / 0.6x | 1.7x / 1.8x / 1.4x / 0.9x / 0.7x | 2.4x / 1.0x | 129 | 0.00745 |
| Apple M5 Max, package (approx) | 1.40 / 2.02 | 1.1x / 0.9x | 3.0x / 2.5x / 2.0x / 1.3x / 1.0x | 4.4x / 1.4x | 129 | 0.00752 |
`*` modelled v4 energy. The chip's hourly cost per MH/s (two-year amortisation plus electricity, approx): f = 1 GDDR7 USD 0.00021, HBM3 one stack 0.00041, eight stacks 0.00025, f = 0 recompute 0.00109; an owned 5090 0.00113; rented hash 0.0117. So the stored-dataset chip undercuts rented hash 56x and an owned 5090 5.4x per MH/s-hour, and the recompute chip matches the 5090 exactly, which is why nobody builds it.
What changed against the 5 October record: the honest denominators moved (the 5090 is 2.34 to 2.65 uJ by where it is measured, not 2.40), the marginal ALU energy is measured at 11 pJ (not the model's 5.5), the M5 Max at 0.78 uJ is the honest best per joule by 3x, and the rented fleet shows the untuned mid-tier (3060, 3080, 3090, A5000, 4060 Ti) at 3.8 to 6.1 uJ, 1.5 to 2.3x worse than the 5090 bench row: against those cards the GDDR7 chip reads 8x to 13x per joule at v3 and 3.1x to 4.6x at v4 with k = 1. The per-tier reading: the Apple tier is already inside 2x of the GDDR7 chip with no shadow and crosses under 1x at v4 and k = 1; the 5090 and the tuned 4070 reach about 2.1x to 2.2x at v4 and k = 1; the untuned mid-tier and AMD stay at 3x to 7.5x, which Ember tuning (the 4070 rows: 3.65 to 2.57 uJ) closes by about 30 percent and nothing in the hash closes further.
The FPGA lane. The brief's soft-overlay range was 0.30x to 0.39x per watt measured-basis and 0.7x to 1.9x at an unmeasured bank-bound ceiling. The reads-in-flight model with the JEDEC HBM2 timings:
| Ceiling | Formula | G reads/s per card | Reads in flight at 137.8 ns | Per W at 115 / 150 / 225 W (M/s/W) | Against the 5090 per W (53.7 M at 326 W) |
|---|---|---|---|---|---|
| Measured, Shuhai U280 default mapping (FCCM 2020, Fig 7) | 32 pc x 75 M | 2.4 | 331 | 21 / 16 / 11 | 0.39x to 0.30x (U55C); 0.20x (U280) |
| tFAW-bound (JEDEC HBM2, ICCAD 2021 Table I: 30 cycles at 1,066 MHz, 4 ACT per channel) | 16 ch x 4 / 28.1 ns | 2.3 | 313 | 20 / 15 / 10 | 0.37x to 0.28x; 0.19x |
| tRRD-bound (6 cycles) | 16 ch / 5.6 ns | 2.8 | 392 | 25 / 19 / 13 | 0.46x to 0.35x; 0.24x |
| Bank-bound, no activate window (the epoch-length 12.2 ceiling row) | 32 pc x 16 banks / 45 ns | 11.4 | 1,567 | 99 / 76 / 51 | 1.84x to 1.41x; 0.94x |
| O'Connor's HBM2 activate figure as carried by chip-model-v3 5.3 | 16 ch x 8 / 12 ns | 10.7 | 1,470 | 93 / 71 / 47 | 1.73x to 1.32x; 0.88x |
The reading: Shuhai's measured 2.4 G/s equals the tFAW ceiling at JEDEC timings (2.3 G/s). The measured row was read on 6 October as a mapping artefact ("the paper's point is that this mapping is the wrong one for random access"); the activate window says it is the DRAM's own limit, which a bank-interleaved mapping does not lift because tFAW is enforced per channel by the die. The 11.4 G bank-bound row needs tFAW gone; O'Connor's 12 ns figure (which the chip model carries for HBM2 and HBM3) is 2.3x shorter than the JEDEC HBM2 cycle count tabled by ICCAD 2021, and the difference is the whole 0.7x to 1.9x row. Tightened range for a 2-stack HBM2 FPGA (U55C, U280, F2's VU47P): 2.3 to 2.9 G reads/s, 0.30x to 0.47x of the 5090 per watt, in the RX 9070 XT's class (2.4 to 2.7 G/s measured). HBM2e parts (Versal HBM, Agilex 7 M) raise the pin rate, not tFAW in nanoseconds (approx), so they sit in the same band; no FPGA with HBM3 exists as a product. Approximate throughout: the 16 half-banks per pseudo-channel, the 1,066 MHz reading of the ICCAD table, the board watts under load.
Can a rented FPGA hour measure it? Vast.ai lists no FPGAs (GPU marketplace only, checked 6 October 2026). AWS F2 (f2.6xlarge: one Virtex UltraScale+ HBM VU47P, 16 GB HBM2 in 2 stacks, 32 pseudo-channels, USD 1.98 an hour on-demand in us-east-1, USD 0.66 spot) carries the same HBM2 subsystem as the U55C and U280, so yes. The gate, written as a measurement plan:
| Step | What | Hours (agent) | Pass line |
|---|---|---|---|
| 1 | Port the chase kernel of `docs/benchmarks/repro.md` 2.2 to a Vitis HLS AXI master over the HBM IP: 1 GiB working set across all 32 pseudo-channels, N dependent 4-byte-reads-in-flight lanes (N = 256, 1,024, 4,096), the HBM IP's address map set to bank-interleaved (RAMA or the IP's "random access" option), a second variant with 32-byte reads | 4 to 6 | the kernel reports reads per second and the chain's checksum equal to the CPU's |
| 2 | Build the AFI (the F2 shell flow; the Vivado licence rides with the instance), 2 to 4 hours of F2 time at USD 2 to 8 | 2 (mostly waiting) | an AFI that loads |
| 3 | Run the ladder; read board power through `xbutil examine --report electrical` (or the F2 shell's sensors) at 1 Hz; take the mean over each run | 1 | reads per second and watts per rung |
| 4 | Write the row into `epoch-length.md` 12.2 in place of the ceiling row | 1 | the public claim carries a measured FPGA number |
| Gate | reads per second per watt at 1 GiB | | expected 15 to 25 M/s/W (0.3x to 0.5x of the 5090); the alarm line is 27 M/s/W (0.5x); over 54 M/s/W (1.0x) the FPGA lane becomes a Counter ASIC 4.0 item |
Consequences per tier of the FPGA finding: none today (no FPGA mines); if the measured row holds, a soft-overlay FPGA at USD 4,000 to 5,000 a card (approx) mines at an RX 9070 XT's rate per watt for 7x the price, so no home or rig tier is displaced by it; the per-program bitstream lane stays answered by layer 9.
### 5.2 Task 2: the reserve R0 to R8
| Slot | Family | Chip block it adds (approx area in 32-bit adders per lane, `counter-asic-3-reserve.md` s3) | Chip datapath energy per op, N5 floor (pJ, approx) | Honest step cost, Apple / NVIDIA / AMD (measured, ratio to the add-xor-rotate chain) | Verifier cost | Chip edge per joule it moves (model) | Cost per tier |
|---|---|---|---|---|---|---|---|
| R0 | derive (per-day item program) | a sequencer: 60 KB instruction store, 16-register file, ALU with multiplier and rotator; removes the 3x fixed-function credit of the f = 0 chip | n/a (it is the f = 0 chip's item cost: 9,992 ops per item) | hash rate 0 on all three vendors; daily build +7 ms Mac, 0 on the 5090 and 9070 XT; compile +0.75 s Mac, +1.1 s NVRTC, +2.0 s AMD per day (a once-a-day module is a requirement) | dr736 +2.8 ms per warp on the Mac, +5.2 on the box core, over the gate cold; dr368 +0.6 Mac, +0.2 box, 8.16 on the half-core | f = 0 chip: 0.92x to 0.34x to 0.43x per chip; per joule from 1.86x to about 0.7x to 0.9x at the same allowance (approx); f = 1 chips: 0 | pool verifier cores x2.4 at dr736, x1.3 at dr368; every miner tier 0 |
| R1 | shfla (lane + delta) | 32-lane x 32-bit crossbar per warp, 32,768 mux bits, 3 to 6 adders per lane | 1.00 | 1.91 / 1.53 / 0.75 to 0.84 | one op per instruction, under 0.01 ms per warp at W_new = 4 | shadow floor +21 percent at 5 percent of instructions (approx); the GPU pays 1.5x the step too, so k unchanged | Apple under 1 percent of ALU time (argued), NVIDIA and AMD 0 |
| R2 | perm (byte permute) | 4x4 byte crossbar, 128 mux bits, 1 to 2 adders | 0.10 | 1.13 emulated / 1.30 / 1.73 to 1.93 emulated | negligible | under 1 percent | Apple and AMD emulate at 1.1x to 1.9x per op: under 1 percent at 4 points |
| R3 | popc and clz | popcount tree and priority encoder, 1.5 to 3 adders | 0.10 | 0.87 and 1.01 / 1.50 and 1.63 / 0.91 to 1.01 and 1.19 to 1.30 | negligible | under 1 percent | 0 |
| R4 | bfe | mask generator on the shifter, 0.3 adders | 0.06 | 0.77 / 1.54 / 1.00 to 1.09 | negligible | 0 | 0 |
| R5 | shl, shr | barrel shifter beside the rotator, 0.2 adders | 0.06 | 0.85 and 0.86 / 1.27 and 1.28 / 1.00 to 1.02 and 0.91 to 1.02 | negligible | 0 | 0 |
| R6 | sel | 32 muxes, 0.3 adders | 0.03 | 0.76 / 1.32 / 0.90 to 1.01 | negligible | 0 | 0 |
| R7 | andn | 32 inverters, under 0.1 adders | 0.02 | 0.75 / 1.26 / 0.91 to 1.01 | negligible | 0 | 0 |
| R8 | mm8 | 8x8x16 u8 MAC tile per warp, about 100 adders per lane; licensable IP at every node | 1.60 | owed (Metal 4 matmul2d) / 2.43 / 1.68 to 1.83 native, exactness unverified | 1,024 MACs per instruction per unit, about 10 us per warp at W_new = 4 | shadow floor +37 percent at 5 percent of instructions (approx); the GPU pays 2.4x the step: k unchanged or worse for the GPU | Apple pays the library path (owed); NVIDIA and AMD native |
The ordering argument. Against the f = 1 chip, which is the chip anyone builds, every family R1 to R8 moves the per-joule edge by under 10 percent, because the chip's cost at class v4 is N x k x 11 pJ and a family changes the per-op energy of 4 points in 79 of the shadow mix. Against the f = 0 chip only R0 matters (it is the item cost). So the reserve's function is not per-joule resistance; it is to make a datapath taped out against class v3's eleven families wrong at the first unlock. A chip maker who reads this document tapes out every block from day one: R1 to R7 together are about 8 to 14 adders per lane (a lane with a 32x32 multiplier is about 35), about 11 mm^2 on a 14,000-lane N5 array, about USD 4 of silicon (sram-mirror.md's USD 0.36 per mm^2, approx); R8 is a licensed tile. The era schedule therefore buys nothing against a chip and the order can be set by honest cost alone: the largest new structure first while every vendor's measured step cost is in (shfla: AMD measured cheap at 0.75 to 0.84, which the reserve document said was the one number that could move it to R1), the emulated families next, mm8 last at era 8 by the rule (no era-4 exception: mm8 removes no adversary class, `epoch-length.md` 12.4). R0 enters at dr368, not dr736: the box proxy puts dr736 over the gate on a cold run (10.5 ms) and at 15.5 ms on the half-core; dr368 reads 4.7 and 8.2.
Recommended text change to `counter-asic-3-reserve.md` section 5: R1 shfla, R2 perm, R3 popc and clz, R4 to R7 unchanged, R8 mm8 at era 8; R0 at `derive_len = 368` with 736 behind the O-1.14 measurement. The era schedule "family n at era n" stands; what it costs per tier at each unlock is the step-cost table above (Apple under 1 percent of ALU time per family, NVIDIA and AMD nothing measurable), and what it buys is stated honestly in the public text (section 6, proposal 4).
### 5.3 Task 3: a class v5 from the shadow
The three options, each priced:
(i) Shadow ops that depend on the loaded data. The v4 shadow block runs at the end of every iteration on the eight lane registers, and those registers hold the words the iteration's 16 loads XORed in (`generator.rs`: the block is executed after instruction 63 with the iteration's `sel`; `verify.rs` runs it with the same `step`). So every shadow op already consumes loaded data, and the next iteration's load addresses depend on the shadow's outputs. A chip cannot precompute any of it; what it can do is what the GPU does: overlap one lane's shadow with another lane's reads in flight. Option (i) therefore moves k by 0. A variant that draws the shadow's immediates from loaded words is M18's one-bit select at 32 bits: a chip wires the operand, and it is not a defence. Verdict: void; no measurement needed.
(ii) Shadow work that exercises GPU structures a chip lacks. Bank-conflict timing is not a value and cannot enter a bit-exact function. Register-file width (8 registers today) can be widened in the shadow to 16 or 32 (ProgPoW used 32): a chip's register file per lane in flight grows from 32 B to 128 B, 1.8 MB on a 14,000-lane array, trivial. Warp shuffles are the structure that costs: the xor-mask shuffle needs a 5-stage butterfly per warp (5,120 mux bits), the indexed shuffle a full crossbar (32,768). The chip datapath floor per op (N5, approx: add 0.06 pJ, mul 0.52, butterfly 0.30, crossbar 1.00, mm8 tile 1.60) against the GPU's measured step costs:
| Shadow op mix | Chip floor per op (pJ, approx) | k floor after the 2x pipeline and 8x register and wire overhead of latency-shadow s6 (approx) | GPU cost of the same mix on the 5090 (step ratios, measured) |
|---|---|---|---|
| Today's weights (add 32, mul 22, rot 13, shfl 8 of 75) | 0.221 | 0.32 | 1.0x to 1.5x per op |
| Shuffle-heavy (shfl 14, shfla 8, the rest scaled) | 0.315 | 0.46 | the shuffles cost the 5090 1.49x to 1.53x the chain step, so its own energy per op rises about 10 percent (approx) |
| With mm8 at 8 points | about 0.45 | about 0.65 | mm8 costs the 5090 2.43x the step |
So a shuffle-heavy shadow raises the attacker's claimed floor from about 0.3 to about 0.5 and costs the GPU about 10 percent more energy per shadow op; it does not reach k = 1. What a chip would need to add: the 32-lane crossbar per warp (R1's structure), nothing else new. The structure argument is bounded: a wide-SIMD array with a crossbar per 32 lanes is still a fixed datapath with no scheduler, and that is where the 0.3 came from.
(iii) One N for every card. Per-card bind points at the 5 percent rule (measured): M5 Max about 130,000 counted ops (interpolated between 102,100 at -1.5 percent and 150,800 at -7.3), RTX 5090 at its 431 W cap about 210,000 (199,600 at -2.7, 330,700 at -34.7), RTX 4070 at its 160 W cap about 250,000 (199,600 at +0.4, 330,700 at -12.3, approx), RX 9070 XT over 331,000 (holds at every rung; budget about 650,000). The consensus N is the minimum: 130,000, set by the Mac. The unmeasured rented cards by ALU budget (approx, cores x clock): 3060 about 270,000, 4060 about 440,000, 3080 about 360,000; their power caps are unmeasured and on NVIDIA the cap binds before the budget, so these are not pass marks. The ladder at every N the lane can compute:
| N counted ops | 5090 uJ (rate delta) | M5 Max uJ (delta) | 4070 uJ (delta) | f = 1 GDDR7 chip uJ at k = 0.3 / 0.5 / 1 / 1.5 | Edge over the 5090 | Edge over the M5 Max | Edge over the 4070 | Verifier add per warp: M5 Max / box core / box half-core (ms) |
|---|---|---|---|---|---|---|---|---|
| 930 (class v3) | 2.65 | 0.78 | 2.57 | 0.47 / 0.47 / 0.47 / 0.47 | 5.7x | 1.66x | 5.5x | 0 |
| 102,100 (class v4) | 3.27 (-0.2%) | 1.39 (-1.5%) | 3.51 (+0.4%) | 0.80 / 1.02 / 1.58 / 2.14 | 4.1x / 3.2x / 2.1x / 1.5x | 1.74x / 1.36x / 0.88x / 0.65x | 4.4x / 3.4x / 2.2x / 1.6x | 0.18 / 0.39 / 0.67 |
| 130,000 (the one-N candidate) | 3.27 (-0.3%) | 1.43 (-4.8%) | 3.78 (+0.4%) | 0.89 / 1.18 / 1.89 / 2.60 | 3.7x / 2.8x / 1.7x / 1.3x | 1.61x / 1.22x / 0.76x / 0.55x | 4.2x / 3.2x / 2.0x / 1.5x | 0.23 / 0.49 / 0.85 |
| 150,800 | 3.27 (-0.3%) | 1.46 (-7.3%) | 3.98 (+0.4%) | 0.96 / 1.29 / 2.11 / 2.94 | 3.4x / 2.5x / 1.5x / 1.1x | 1.52x / 1.13x / 0.69x / 0.50x | 4.1x / 3.1x / 1.9x / 1.4x | 0.26 / 0.57 / 0.99 |
| 199,600 | 3.35 (-2.7%) | 1.58 (-10.5%) | 4.45 (+0.4%) | 1.12 / 1.56 / 2.65 / 3.74 | 3.0x / 2.1x / 1.3x / 0.9x | 1.40x / 1.01x / 0.59x / 0.42x | 4.0x / 2.9x / 1.7x / 1.2x | 0.35 / 0.76 / 1.31 |
| 330,700 | 4.99 (-34.7%) | 1.87 (-21.0%) | 5.89 (-12.3%) | 1.55 / 2.28 / 4.09 / 5.91 | 3.2x / 2.2x / 1.2x / 0.8x | 1.21x / 0.82x / 0.46x / 0.32x | 3.8x / 2.6x / 1.4x / 1.0x | 0.58 / 1.26 / 2.18 |
Per-tier watts at N = 130,000 (measured rungs interpolated): the M5 Max 37 W GPU plus DRAM (from 21), the 5090 431 W (its cap, from 350), the 4070 about 118 W (from 79.5), the 9070 XT owed; a 5090 rig pays about 23 percent more electricity for 0.3 percent less rate, an Apple miner 1.8x the GPU watts for 4.8 percent less rate, a pool user nothing, every verifier tier +0.23 to +0.85 ms per warp. The verifier at 130,000 on the half-core proxy is 8.5 ms (8.23 + 0.85 - 0.67), inside the gate with 1.5 ms spare; the gate does not bind N before the Mac does.
#### 5.3a The reconciled N ladder (this lane's measured cards, lane 7's HBM4 inputs; `--section ladder`)
Lane 7 (`docs/analysis/horizon/frontier.md` 2.3, `sim/horizon/frontier/frontier_model.py` model 1.4) adds HBM4: JEDEC JESD270-4 doubles channels per stack (16 to 32, cited), so its one-stack ceiling is taken as 2 x 10.7 = 21.4 G reads/s at 1.0 nJ per read (approximate, unsourced), 5 W static, 10 W controller: 167 MH/s at 0.218 uJ bare. The table below carries that chip beside the three of chip-model-v3 5.4, every chip at N = bare + (N - 930) x 11 pJ x k, and every honest card at its measured rung.
| N counted ops | 5090 uJ, W (rate delta) | M5 Max uJ, W (delta) | 4070 uJ, W (delta) | 9070 XT rate delta (W owed) | 12 GB and 8 GB rented cards | Chip edge over the 5090 per joule at k = 0.5 / 1 / 1.5: GDDR7 | HBM3 one stack | HBM3 eight stacks | HBM4 one stack | Verifier ms per warp: M5 Max core / 2019-class by the 2.5x rule / measured half-core proxy | Card that binds first (5 percent rule) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 930 (class v3) | 2.65, 350 W (0) | 0.78, 21 W (0) | 2.57, 80 W (0) | 0 | measured at v3 only: 5070 3.27 uJ, 3060 4.36, 4070 untuned 3.65, 4060 Ti 8 GB 3.81 | 5.7x | 8.3x | 10.1x | 12.2x | 2.06 / 5.2 / 7.56 | none |
| 49,700 | 3.21, 425 W (+0.1%) | 1.16, 31 W (-1.2%) | 3.01, 94 W (+0.4%) | +2.1% | not measured at any N | 4.4x / 3.2x / 2.5x | 5.5x / 3.7x / 2.9x | 6.1x / 4.0x / 3.0x | 6.6x / 4.3x / 3.1x | 2.15 / 5.4 / 7.88 | none |
| 102,100 (class v4) | 3.27, 431 W (-0.2%) | 1.39, 37 W (-1.5%) | 3.51, 109 W (+0.4%) | +2.0% | not measured | 3.2x / 2.1x / 1.5x | 3.7x / 2.3x / 1.6x | 4.0x / 2.4x / 1.7x | 4.2x / 2.5x / 1.7x | 2.24 / 5.6 / 8.23 | none (M5 Max -1.5%) |
| 130,000 | 3.27, 431 W (-0.3%) | 1.43, 37 W (-4.8%) | 3.78, 117 W (+0.4%) | +1.7% | not measured | 2.8x / 1.7x / 1.3x | 3.2x / 1.9x / 1.3x | 3.4x / 1.9x / 1.4x | 3.5x / 2.0x / 1.4x | 2.29 / 5.7 / 8.41 | M5 Max at its 5 percent point |
| 150,800 | 3.27, 431 W (-0.3%) | 1.46, 37 W (-7.3%) | 3.98, 124 W (+0.4%) | +1.5% | not measured | 2.5x / 1.5x / 1.1x | 2.9x / 1.7x / 1.2x | 3.0x / 1.7x / 1.2x | 3.1x / 1.8x / 1.2x | 2.32 / 5.8 / 8.55 | M5 Max (-7.3%) |
| 199,600 | 3.35, 431 W (-2.7%) | 1.58, 38 W (-10.5%) | 4.45, 138 W (+0.4%) | +1.0% | not measured | 2.1x / 1.3x / 0.9x | 2.4x / 1.3x / 0.9x | 2.5x / 1.4x / 0.9x | 2.6x / 1.4x / 1.0x | 2.41 / 6.0 / 8.87 | M5 Max (-10%), 5090 (-2.7%) |
| 330,700 | 4.99, 431 W (-34.7%) | 1.87, 40 W (-21.0%) | 5.89, 160 W (-12.3%) | +3.6% | not measured | 2.2x / 1.2x / 0.8x | 2.3x / 1.3x / 0.9x | 2.4x / 1.3x / 0.9x | 2.5x / 1.3x / 0.9x | 2.64 / 6.6 / 9.74 | 5090 (-35%), M5 Max (-21%), 4070 (-12%) |
Sources per column: 5090 and M5 Max rungs `latency-shadow-2026-10-06.md` s3 and s5; 4070 and 9070 XT rungs `counter-asic-3-status.md` item 8 (the 9070 XT's watts owed: the ADLX sampler parsed 0 samples); the rented cards `prover-tiers-real-cards.md` (class v3 only); chip bare energies chip-model-v3 5.4 and lane 7 model 1.4; the 11 pJ unit latency-shadow s5; the verifier slopes 3.2 us per 1,000 shadow instructions (Mac, measured) and 12.1 us (box half-core, this lane); the 130,000 row interpolated. Every chip cell is approximate.
Disagreements with lane 7's model 1.4, named: (a) its honest card at N is the linear 326-to-575 W model of chip-model-v3 5.7 (2.95 uJ at N = 100,000, 3.50 at 200,000, 4.22 at 330,000); the measured 5090 under its 431 W cap reads 3.27, 3.35 and 4.99 uJ with -0.2, -2.7 and -34.7 percent of rate (the cap binds from 102,100 ops; `power.min_limit` is 400 W, so no cap below the one measured exists on a 5090). (b) Its shadow core is 150 W fixed at the 5090's 136 MH/s; on a 167 MH/s HBM4 chip that under-counts the core by 1.23x. Energy per hash is N x 11 pJ x k whatever the chip's rate, so HBM4 at N = 100,000 and k = 1 is 1.33 uJ and the edge 2.5x, not 1.11 uJ and 2.65x; at 200,000 1.4x (lane 7: 1.74x); at 330,000 1.3x (1.33x: agrees, because the 5090's own energy jumps to 4.99 at that rung). (c) Its HBM3 and HBM4 ceilings (10.7 and 21.4 G per stack) rest on the 8-activates-per-12-ns rate chip-model-v3 5.3 carried from O'Connor; the JEDEC HBM2 cycle table (ICCAD 2021 Table I) gives 4 per 28 ns per channel (section 5.1 of this file), and HBM3's own tFAW is behind the paywall. If HBM3 and HBM4 keep HBM2's window the one-stack ceilings are 2.3 and 4.6 G, HBM4's bare energy 0.55 uJ and its bare edge 4.9x, not 11x. The GDDR7 column is the one with a measured anchor (the 5090 reaches 82 percent of its ceiling) and is the column to quote in the summary; the HBM columns are the upper bound.
Verifier headroom for N (lane 7's "10x of headroom"): on the M5 Max core 10 - 2.33 = 7.67 ms buys 2.4 M shadow instructions, N about 4.5 M ops (19x); on the 2.5x rule 4.2 ms buys 525,000 instructions, N about 1.06 M (10x, lane 7's figure); on the measured half-core proxy 1.77 ms buys 146,000 instructions, N about 370,000 (3.7x). The 10x holds on the rule and not on the pessimistic measured bracket; O-1.14 decides which. Either way the cards bind first: M5 Max 130,000, 5090 at 431 W 210,000, 4070 at 160 W about 250,000, 9070 XT over 331,000. An unconditional doubling of N per era (lane 7's candidate) would take the Apple tier out at the first step (200,000: -10.5 percent) and the 5090 and 4070 at the second (400,000: compute-bound at their caps), which is why the proposal below steps N by signal, not by schedule alone.
The verdict on v5: the shadow lever is close to spent on the owned cards. Going from 100,000 to 130,000 buys 2.1x to 1.7x against the 5090 at k = 1 and costs the Mac its whole 5 percent allowance; a shuffle-heavy mix buys the k floor 0.3 to 0.5. Together they define one candidate, `v5 = mx8 + sh256x35 with the shuffle-heavy weight table`, worth measuring but not worth a cut on its own: the chip question is k, and no shadow design moves k past about 0.5 against a fixed-datapath array.
What measurement decides it (not run tonight; the shuffle-heavy weight table does not exist as a knob, and the Mac's miner state was not checked, so the plan stands in for the run): (1) add a shadow weight table to `ShadowClass` (`generator.rs`, 2 hours), draw the block from it, emit it in the three dialects as today; (2) export `mx8+sh256x35` at today's weights and at the shuffle-heavy table for seed igneum-genesis; (3) Mac: `with-lock.sh measure packbench --pack <dir> --batches 60 --batch-log2 24 --group 256` with the IOReport sampler, 2 packs plus the control, about 6 minutes, the miner paused first (`curl -s http://127.0.0.1:60030/.../api/state` from the app's `app.url`, then pause through the app, never from a script); (4) the same ladder as PC jobs on the 5090, 4070 and 9070 XT through the existing `tools/ca3-shadow` playbooks with the card off through the runner's `--cards-off`; (5) `igneum-pow bench` on the Mac core and the box proxy. Pass lines: every owned card within 5 percent of its class v4 rate; bit-exact fingerprints on Metal, CUDA and AMD OpenCL; verifier under 10 ms on the half-core proxy; the 5090's marginal pJ per op on the shuffle-heavy mix read on the three rungs under its cap.
### 5.4 Task 4: the era draw's randomness
How it picks. Era n's seed E_n is the 1-hour class-group VDF of `Hash(chain_id || n || blue block hashes of the day before C_era(n))`, where C_era(n) is the highest certified checkpoint at least 7,200 DAA s before the era (spec 04 s4.4). One SplitMix64 stream from E_n draws, in order: the ten op weights perturbed by -2..+2 points each, the output fold rotations, an unused draw for `epoch_len` (set by signal), then under class v3 a second stream draws the load width (pinned at 4 bytes: the draw is consumed), the stride multiplier M (odd) and rotation R, and the four interleave bit positions (spec 01 s1.13.1, `era-layout.md` 1.1). Not drawn: the load count (16), the mixer round count (8) and multiplier (8), the cache and dataset sizes, the item construction.
What an attacker can bias, priced. Two routes. (a) Forge the certified checkpoint: needs 2/3 of the 30-day blue-block weight, which is 20 days of 100 percent of the network's hash (the headline). Rented at the measured USD 0.0117 per MH/s-hour:
| Network hash | 20 days of 100 percent, rented | What it buys in the draw |
|---|---|---|
| 1 GH/s | USD 5,616 | one era's (weights +-2, fold, M, R, pos): a per-card hash-rate spread of 0.8 to 3.2 percent (six eras measured, bench-log "Counter ASIC 2.0, the numbers"), 0 chip effect |
| 10 GH/s | USD 56,160 | the same |
| 100 GH/s | USD 561,600 | the same; the rental market could not supply 20 pods of any card at 19:00Z on 6 October (bench-log), so a TH/s is not rentable at all |
| 1 TH/s | USD 5.6M (not supplied) | the same |
(b) Re-roll without weight: the miner of the last blue block before C_era(n)'s cut withholds or publishes to change the input set; this costs one block's reward and needs the 3,600-s VDF evaluated inside the 2-s publish window, a 1,800x faster evaluator (spec 04 s4.6: a 300x evaluator beats the epoch's 600 s, not the era's 3,600 s). Even free, one re-roll is one more sample of the same space. So the draw is unbiasable at any price that matters, and that is the honest answer to "what can be biased": nothing worth having.
Weak corners. The op-weight perturbation can move the multiply share (mul, mad, mulhi: 22 of 75) to 16 or 28 of 75; at the N5 datapath floor that is 0.158 to 0.232 pJ per shadow op (0.195 at the base), about +-20 percent of the shadow's datapath energy, and the GPU's energy moves the same way (its IMAD is the chain's own op). The stride and interleave cost a chip two integer operations per load and an address-line permute, nothing per joule. The fold rotations are a wire mux. The measured six-era spread (1.3 percent on the 5090, 3.2 on the 9070 XT, 0.8 on the M5 Max) is the whole of what the draw moves. There is no corner that makes a chip easier, because no drawn parameter touches the memory bound, the item derivation or N.
Predictability against the 32-month lead time. Everything a chip needs is public at genesis: the eleven live families, the reserve order R0 to R8 and its era schedule, the mixer shape and x8, the dataset and cache schedule, the class v4 shadow's size and weights. The draw hides (M, R, pos, weights +-2, fold rotations) until 2 hours before each era, and none of those needs silicon: an address decoder that permutes lines, a programmable rotator, an immediate table. So a chip taped out in month 0 against this spec runs every era for the chain's life, and the era draw buys nothing against it. What the draw does buy: the fork-fatigue lesson of the history (no human release, no vote), and a per-program hard-datapath FPGA cannot amortise a bitstream across eras (layer 9 handles the within-era case). Said plainly for the public text: the era draw and the reserve are automatic schedule changes against fixed datapaths and governance, not unpredictability against a chip. What would be unpredictable and costly to a chip is not available in a genesis-fixed rule set: a per-era draw among K reviewed item constructions (K mixers or K derive forms) is still K public blocks a chip carries; a per-era draw of the load count or the dependent-read depth changes the memory bound per era and fails the 5 percent rule on the honest cards (read-width: a per-program width mix spread 5.5 to 22.3 percent). What costs a chip is N (its k), the memory system (f = 1 is the ceiling and is a commodity controller), and the honest card's own watts (the M5 Max at 0.78 uJ is inside 2x of the GDDR7 chip with no shadow at all).
### 5.5 Task 5: the 2019-class verifier gate
Measured tonight (section 3.3): on igneum-build-1 one EPYC 9454P core boosted to 3.8 GHz (not a low clock: the governor's cap is 2.75 GHz but the core read 3,800 MHz under load, and `cpupower` needs root) with server DDR5 behind it reads 1.9x to 2.2x the quiet M5 Max core across five classes, and 2.0x on the cache fill. The half-core proxy (both SMT siblings busy on the same class) reads 3.4x to 3.7x the Mac. A 2019 laptop core (Skylake-class at 3.5 to 4.5 GHz with DDR4 at about 80 ns) sits between these brackets on the arithmetic (lower IPC than Zen 4, lower DRAM latency than the server), so the 2.5x rule is about right and the two proxies bracket it. Against the gate:
| Class | One box core, cold run | Half-core | Verdict at 10 ms | Headroom left for shadow on the half-core (ms) |
|---|---|---|---|---|
| class v3 (mx8) | 4.67 | 7.56 | pass | 2.4 |
| class v4 (mx8+sh256x27) | 5.06 | 8.23 | pass | 1.8 (about 150,000 more shadow instructions at 12.1 us per 1,000) |
| dr368 | 5.32 | 8.16 | pass | 1.8 |
| dr736 | 10.51 | 15.49 | FAIL on both proxies | none |
So dr736 is out as a genesis-live or near-term reserve length on measured evidence, not on the 2.5x rule; dr368 is in. The owed O-1.14 measurement on a real 2019 laptop (the US laptop's Windows `igneum-pow` build, main's decision 7) still closes the question; the box proxy is the stand-in until it lands, and the half-core row should be the standing pessimistic rule in place of "2.5x" (it is a measurement; 2.5x is a ratio from memory).
How to measure it tomorrow: `tools/cross-remote.sh` from `igneum-pow/` builds the Windows exe on the box (1 min 44 s measured for the node; the pow crate is one crate, under a minute), the relay carries it to the US laptop when it appears, `igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <c> --warps 50` for the five classes, ms per warp into `docs/bench-log.md` under O-1.14; 1 hour of agent work. A rented old CPU on Vast is the fallback (Vast lists CPU-only offers by core generation; a 2019 Xeon or i7 host at under USD 0.10 an hour, approx), same binary, same command.
What the gate protects, and what loosening it costs:
| | At 10 ms per warp | At 20 ms (loosened) | At 5 ms (tightened) |
|---|---|---|---|
| A node on a 2019 laptop at 1 bps | 1 percent of one core per block | 2 percent | 0.5 percent |
| At 10 bps (the Devnet 2 experiment) | 10 percent of one core | 20 percent | 5 percent |
| IBD over the 108,000-header pruning window, one core | 18 min | 36 min | 9 min |
| Header flood (M15 class): invalid headers per second that saturate one core | 100 | 50 | 200 |
| A pool core verifying shares | 100 per second | 50 | 200 |
| What the lever buys the hash | the x8 mixer, dr368, the v4 shadow all fit with 1.8 ms spare on the half-core | dr736 fits (15.5 on the half-core: no, still out), x16 mixer fits | nothing of class v4 fits on the half-core |
Loosening to 20 ms would admit the x16 mixer (about 7 ms on the Mac, 15 on the half-core: still out) and not dr736 on the half-core, so it buys little against a chip (the f = 1 chip derives no item) and doubles the header-flood and IBD costs on the weakest node. Keep 10 ms; measure the laptop; use the half-core row until then.
### 5.6 Task 6: the dataset schedule to 2030
The schedule (spec 1.13.3 option (b), decided for the cache on 5 October, recommended for the dataset): 2 GiB at genesis, 4 GiB at year 4 (day 1,460), 8 GiB at year 12, 16 GiB at year 28; the cache 256 MiB, 512 MiB, 1 GiB, 2 GiB on the same days (`memhard.rs` `growth_doublings`). The devnet packs run 1 GiB today.
The installed base against it (Steam September 2026, cited; the trend is approximate):
| Year (approx calendar) | 8 GB share | 12 GB share | Dataset | Who falls out of mine-only | Who falls out of mine-and-prove (the prover's measured peaks, `prover-tiers-real-cards.md`) |
|---|---|---|---|---|---|
| 2026 (devnet) | 27% | 13% | 1 GiB | nobody with 4 GB or more | 8 GB: compressed does not fit beside the miner (measured); core-only 2^25 fits with 1 GB spare |
| 2027 (genesis, year 0) | about 20% | about 10% | 2 GiB | 4 GB cards hold with 0.5 to 0.8 GB spare (card-lifetime) | 8 GB loses core-only beside the miner (the miner's resident set grows 1 GiB: 7.35 + 1.0 GB over 8.19) and becomes prove-alone; 12 GB holds compressed on headless Linux (10.2 + 1.0 of 12.3) |
| 2031 (year 4) | about 0% | about 0% | 4 GiB | 4 GB cards (5 percent of Steam today, about 0 by then) | 12 GB loses compressed (10.2 + 3.0 GB over 12.3), keeps core-only (7.2 + 3.0 of 12.3); 16 GB keeps compressed (9.2 + 3.0 of 16.4) |
| 2039 (year 12) | 0 | 0 | 8 GiB | 8 GB cards (26.7 percent of Steam today; the trend says under 1 percent by 2030) | 12 GB loses core-only; 16 GB loses compressed, keeps core-only (7.4 + 7.0 of 16.4); 24 GB keeps compressed (11.0 + 7.0 of 24.6) |
| 2055 (year 28) | 0 | 0 | 16 GiB | 12 and 16 GB | 24 GB loses compressed; 32 GB keeps everything |
Reading per tier: the 8 GB tier, a quarter of Steam today, is never a mine-and-prove card on the measured prover footprint whatever the dataset does, and it mines until year 12, by which time its share on the trend is nil; the 12 GB tier (13 percent, falling 6 points a year) mines to year 28 and loses mine-and-prove compressed at year 4 because the prover peaks at 10.2 GB beside a 1.4 GB miner; the 16 GB tier (27 percent and rising) mines and proves compressed to year 4 and core-only to year 12; 24 and 32 GB are unconstrained to year 28. The dataset is never the binding constraint on any tier before year 12; the prover's 5.6 to 10.7 GB footprint is. On the chip side the schedule changes nothing: one HBM3 stack holds 24 GB and the 5090's own board 32 GB (chip-model-v3 5.7), so f = 1 reads every step of the schedule to year 28 without a second stack.
What growth is for, then. Two things, both real and neither a chip: (1) the cache above every GPU's on-die cache (96 MB on the 5090, 128 MB on GB202; the spec's own rule), so the honest hash stays DRAM-latency-bound and no consumer GPU gains an L2 shortcut; (2) the dataset above one reticle of SRAM at the node of the day (1.6 GiB per reticle at N5 headline density, 1.9 at N2, sram-mirror.md s4), so an "f = 1 in SRAM" chip, which would read at SRAM latency and beat the DRAM activate ceiling by 10x, stays a multi-reticle part: at 2 GiB that is 2 dies at N2, at 4 GiB 3 dies, at 8 GiB 5 dies (approx, headline density; the lower-bound density halves these). Wafer-scale parts already hold more (a Cerebras WSE-3 carries 44 GB of on-wafer SRAM, approximate, from memory, unpriced here): the schedule does not price that device out and nothing in the hash can, but at the dependent-read pattern a wafer's cross-die hops cost latency that no one has measured for this hash (open, section 7).
Recommendation with numbers: hold the schedule as decided (2 GiB genesis, doublings at years 4, 12, 28). Do not slow it: slowing buys the 8 GB tier nothing (it mines to year 12 either way) and loses the SRAM-reticle margin (at a flat 2 GiB one N2 reticle holds 1.9 GiB today and about 3.4 GiB by 2036 on the 6 percent a year density trend, sram-mirror s6, so a flat dataset fits one reticle within a decade). Do not grow faster: the only tier a faster schedule costs is the 8 GB tier (year 12 to year 4) and the Apple 8 GB laptop, and it buys nothing against the HBM chip. One change: write the prover footprint, not the dataset, into the public card-lifetime sentence (litepaper line 560: "12 GB or more proves full shards" becomes "12 GB mines and proves on headless Linux until the year-4 dataset step, 16 GB until year 12, 24 GB beyond; 8 GB proves alone"), since that is the number that moves users.
## 6. Ranked proposals
| Rank | Proposal | Evidence | Model | Hours | Consequence per tier | Gate |
|---|---|---|---|---|---|---|
| 1 | Close O-1.14 with a real 2019 laptop run and adopt the half-core proxy as the standing stand-in; dr736 out, dr368 in as R0's length | box proxy: dr736 10.5 ms cold, 15.5 half-core; v4 5.06 / 8.23 | section 5.5 | 2 (Windows cross build on the box, relay to the laptop, bench, log) | miners 0; pool verifier cores x1.3 at dr368 instead of x2.4; a 2019 node keeps 1.8 ms of headroom under v4 | ms per warp under 10 cold on the laptop for v3, v4, dr368; dr736 recorded as the figure that fails |
| 2 | Measure the FPGA lane on AWS F2 (one VU47P, HBM2, USD 1.98 an hour) and replace the 12.2 ceiling row | the measured 2.4 G/s equals the JEDEC tFAW ceiling; the 1.9x row rests on a 12 ns tFAW the JEDEC cycles do not support | section 5.1 | 8 to 10 agent hours plus USD 2 to 8 of F2 time | none today; the public FPGA claim becomes a measured 0.3x to 0.5x per watt | reads per second per watt at 1 GiB; alarm at 27 M/s/W, Counter ASIC 4.0 at 54 |
| 3 | Public-text correction: the era draw and the reserve are automatic schedule changes against fixed datapaths and against forks, not unpredictability against a chip; publish the chip's USD per MH/s-hour beside the honest cards' | section 5.4: every drawn parameter needs no silicon; the reserve is about USD 4 of N5 silicon on a chip | sections 5.1, 5.2, 5.4 | 1 | holders and miners read a claim that survives review; nothing on the devnet changes | `docs/evidence.md` row with the two numbers (USD 0.00021 chip, 0.00113 owned 5090, 0.0117 rented) and the draw sentence |
| 4 | Reserve order: R1 shfla, R2 perm, R3 popc and clz, R4 to R7 unchanged, R8 mm8 at era 8 by the rule; R0 at dr368 | AMD step cost of shfla measured 0.75 to 0.84 (the one number the reserve document said could move R3); dr736 fails the proxy | section 5.2 | 1 (spec text in `counter-asic-3-reserve.md` s5 and s6) | Apple pays shfla's 1.91x per op first, under 1 percent of ALU time at 4 points (argued, measured at the unlock rehearsal); NVIDIA and AMD 0 | the family-live 5 percent run per vendor at each unlock rehearsal |
| 5 | **N grows by the era draw at genesis, inside a verifier-bounded ladder, each step taken by 90 percent miner signal.** The shadow size N becomes a genesis ladder indexed per era, {100,000, 130,000, 200,000, 330,000, 650,000, 1,000,000} counted ops (the measured rungs, then doublings), floor 100,000 and ceiling 1,000,000 fixed at genesis (the ceiling is the 10x verifier headroom on the 2.5x rule; 370,000 on the half-core proxy until O-1.14 lands, which then sets it), the era stream consuming one draw for it as it does for `epoch_len`, and the step up or down set by 90 percent of blue blocks over 7 days at a day boundary (spec 5.7's mechanism, P2's signalling code), never unconditionally | lane 7: HBM4 doubles the f = 1 chip's rate per stack, so the bare edge rises 5.7x to 12x (upper bound) and only N answers it; this lane: the chip's edge over the 5090 at k = 1 falls 2.1x (100,000) to 1.7x (130,000) to 1.3x (200,000 and 330,000); an unconditional doubling takes the M5 Max out at the first step | section 5.3a; `--section ladder` | 6 to 8 (the ladder field in `ShadowClass` and the era stream, the signal rule shared with `epoch_len`, a fast-time run across one step, packs and vectors per rung) | At each step, measured: 100,000 to 130,000 costs the M5 Max 3.3 points of rate and 0 W more, the 5090 and 4070 nothing, the 9070 XT nothing, every verifier +0.05 ms; 130,000 to 200,000 costs the M5 Max 6 more points and the 5090 2.7 at its cap, the 4070 +21 W, every verifier +0.12 ms (Mac) to +0.5 (half-core); 200,000 to 330,000 is compute-bound on every NVIDIA card at its cap (5090 -35 percent) and is a step the signal would refuse until cards change; a pool user nothing at any step; a chip's shadow core grows with N at k x 11 pJ per op | ONE gate per step, published before the project recommends the signal: at the step's N, every card of the public benchmark set (the four owned plus the eleven rented models) within 5 percent of its rate at the previous step, bit-exact fingerprints on Metal, CUDA and AMD OpenCL, and the verifier under 10 ms per warp cold on the O-1.14 core (the half-core proxy until then) |
| 5a | A class v5 candidate `mx8 + sh256x35` with a shuffle-heavy shadow weight table (shfl 14, shfla 8 of 75), measured on the four owned cards before any cut: the first rung of proposal 5's ladder, plus the k-floor lever | the Mac's 5 percent point is 130,000; the shuffle mix raises the k floor 0.32 to 0.46 (approx) | section 5.3 | 4 to build the weight-table knob and packs, 1 Mac measure session (about 6 min under the lock, miner paused), 3 PC jobs | M5 Max -4.8 percent of rate at 37 W; 5090 -0.3 percent at its 431 W cap (a rig +23 percent electricity); 4070 0 at about 118 W; 9070 XT 0; verifier +0.23 ms Mac, +0.85 half-core | every owned card within 5 percent; bit-exact on three vendors; half-core verifier under 10 ms; the 5090's marginal pJ on the new mix read on three rungs |
| 6 | Hold the dataset schedule (2 GiB, years 4, 12, 28); write the prover footprint into the card-lifetime sentence | Steam shares and the measured prover peaks; one HBM3 stack holds every step | section 5.6 | 1 | 8 GB: mines to year 12, proves alone; 12 GB: mine-and-prove compressed to year 4, core-only to year 12, mines to year 28; 16 GB: compressed to year 4, core-only to year 12; 24 and 32 GB unconstrained to year 28 | the litepaper sentence matches the table; `docs/evidence.md` row "card lifetime" labelled designed |
| 7 | Make the Ember tune the shipped default per card model (the honest card's watts are the lever that moves every chip row) | the 4070 at 3.65 uJ untuned and 2.57 tuned (-30 percent); the 5090 2.65 bench against 2.34 app | section 5.1 | 2 (defaults table in the app from the fleet priors; already measured) | every NVIDIA tier gains 10 to 30 percent per joule; the chip's edge over the mid-tier falls from 8x to 13x toward 5x to 9x at v3 | MH per W per card model on the fleet night against the untuned baseline |
| 8 | Fund the k question: the item 3 cryptanalysis brief gains a chip-design line (a 14,000-lane SIMD array's energy per op on a random 32-lane program with shuffles, at N5 and at 28 nm) | every chip row at class v4 turns on k; nothing in the project measures it | section 5.3 | 0 agent hours; Josh's money (part of the USD 80,000 to 160,000 brief) | none until the number lands; it decides whether 2x is reachable | a reviewed estimate of k with its range |
Paragraphs.
1. The verifier gate is the one place tonight produced a measurement instead of a rule. The box proxy brackets a 2019 laptop from both sides (a 2022 server core at full boost; the same core with its sibling busy), and dr736 fails both brackets while class v4 passes both with 1.8 ms to spare. The measurement is one Windows build and one bench on the laptop Josh already owns; until it lands, the half-core row replaces the "2.5x" from memory in every status file.
2. The FPGA lane's upper row was built on an activate rate (8 per 12 ns per channel) that the JEDEC HBM2 cycle table does not support (4 per 28 ns); the measured Shuhai rate sits exactly on the JEDEC ceiling. That reading can be wrong (the ICCAD table's clock interpretation, the half-bank count, the board watts are all approximate), which is why the F2 hour is the proposal and not the conclusion. It is cheap and it turns a public ceiling claim into a measured one.
3. The honest statement about the draw and the reserve is owed before the public testnet. The project has said the era draw and the family reserve are "automatic anti-ASIC escalators"; against the chip that the model says anyone would build they escalate nothing, because every parameter they move is firmware or a USD 4 block. They are good design against forks and against a hard-datapath FPGA, and that is what the text should say. The chip's USD per MH/s-hour (56x under rental, 5.4x under an owned 5090) belongs beside it, because it is the number a miner will compute on the day a chip appears.
4. The reserve order changes only where the measurements moved: shfla's AMD cost came in cheap, so the largest structure goes first; dr736 failed the proxy, so R0 is dr368. mm8 keeps no exception because `epoch-length.md` 12.4 showed it removes no adversary class.
5. N as a genesis ladder is the one structural change this lane proposes, and it is lane 7's idea with the cards' measured bind points written into it. The memory generation it answers (HBM4, 2028) arrives on a two-to-three-year cadence; the chain must answer without a release, which the era stream and the `epoch_len` signal rule already provide the shape for. What the measured rungs add: the ladder's steps are the cards' own bind points, the step is taken by the miners who pay for it, and the ceiling is the verifier's measured core, not a 10x from a ratio. An unconditional doubling per era would retire the Apple tier at era 1 and every capped NVIDIA card at era 2, so the schedule alone is not the proposal; the schedule plus the signal is.
5a. v5 as a class (`mx8 + sh256x35` with a shuffle-heavy mix) is measurable in an evening and is not a cut. Its honest ceiling is the Mac's 5 percent and a k floor of about 0.5; it buys 2.1x to 1.7x against the 5090 at k = 1. The design item that decides more than v5 is k itself (proposal 8).
6. The dataset schedule is right as decided and the public sentence about cards is wrong in kind: it talks about the dataset when the prover is what ends a tier's mine-and-prove life.
7. The one lever that moves every chip row and costs no consensus change is the honest card's watts. The fleet showed the untuned mid-tier at 1.5 to 2.3x the 5090's energy per hash; Ember's measured tune on the 4070 took 30 percent off. Shipping it as the default is a miner-app change with a measured gate.
8. k is the whole chip question at class v4 and nobody in the project can measure it; the external brief can estimate it.
## 7. Open questions and what could not be run
| Item | Why not | What would close it |
|---|---|---|
| The v5 Mac packbench ladder | the shuffle-heavy shadow weight table does not exist as a knob (the block uses the program's weights), and the Mac's miner state was not checked; a run without the knob would have measured v4 again | proposal 5, 4 hours of code then the 6-minute measure session |
| A fixed low clock on the box | `cpupower` needs root; the core boosted to 3.8 GHz under schedutil | the laptop measurement (proposal 1); the half-core row is the pessimistic stand-in |
| The 9070 XT watts at class v4 and its per-joule row | the AMD watts job failed on 6 October and the re-run waits on the runner's `--cards-off` (status 6a) | the next cut's PC 1 job |
| The RTX 4060's watts (logged 0.0 W on the rented box) | the sampler read nothing on that host | one re-rent |
| HBM2 tFAW and half-bank count on the U55C and F2 parts | behind the JEDEC paywall; the ICCAD table is a simulator's configuration (DRAMSim3), not a datasheet, and its clock interpretation (1,066 MHz) is mine | the F2 hour (proposal 2) |
| A wafer-scale SRAM dataset holder (Cerebras-class, 44 GB on-wafer, approximate) | not priced anywhere in the project; its cross-die hop latency on a dependent-read chain is unmeasured | a Counter ASIC 4.0 analysis item, not a lane 2 item |
| The Steam trend to 2030 | linear extrapolation of two points per tier | the survey itself, yearly |
| The era draw's cryptanalysis (the stride bijection, the ROT weak-key draw) | out of scope here and still open in spec 1.8.4 and era-layout s8 | the item 3 brief |
| `block-rate-devnet2.md` RUN_A and RUN_B | placeholders at 21:30 UK | nothing in this lane depends on them |
## 8. Summary for the coordinator
Lane 2 refined the shipped hash and its classes on the 6 October numbers and one new measurement. The chip model's answer does not change in kind: the stored-dataset chip is the chip, it reads 5.7x per joule against the 5090 bench row and 1.7x against the M5 Max at class v3, 2.1x and 0.9x at class v4 and k = 1, and it undercuts rented hash 56x and an owned 5090 5.4x per MH/s-hour; the FPGA lane tightens to 0.30x to 0.47x per watt because the measured random-read rate of an HBM2 FPGA is the JEDEC activate ceiling and not a mapping artefact, and AWS F2 can measure it for USD 2 an hour. The reserve and the era draw buy nothing against a chip (every drawn parameter is firmware; every reserve block is about USD 4 of silicon) and the public text should say what they do buy. The verifier gate got its first measured proxies: class v4 passes a 2022 server core (5.06 ms cold) and the same core with its SMT sibling busy (8.23 ms); dr736 fails both (10.5 and 15.5 ms), so R0 is dr368. The dataset schedule holds; the tier constraint to 2030 is the prover's footprint, not the dataset.
1. Class v4 verifier on the box proxy 4.90 ms steady, 5.06 cold, 8.23 on the half-core; dr736 9.76 / 10.51 / 15.49: dr736 is out on measurement, class v4 keeps 1.8 ms under the gate on the pessimistic bracket (section 5.5; `sim/horizon/algorithm/model.py --section verifier`).
2. The f = 1 GDDR7 chip's edge per joule: 5.7x (5090 bench), 1.7x (M5 Max) at v3; 2.1x and 0.9x at v4 with k = 1; 4.1x and 1.7x at k = 0.3; against the untuned rented mid-tier 8x to 13x at v3; USD 0.00021 per MH/s-hour against 0.0117 rented (section 5.1; `--section chip`).
3. The HBM2 FPGA soft overlay: 2.3 to 2.9 G reads/s per 2-stack card by tFAW and tRRD (measured 2.4), 0.30x to 0.47x of the 5090 per watt; the 1.9x ceiling needs a tFAW of 12 ns that the JEDEC HBM2 table (28 ns) does not give (section 5.1; `--section fpga`).

View file

@ -0,0 +1,675 @@
# Horizon lane 7: frontier. Predictions to 2030, what no proof-of-work chain has shipped, and what Igneum can
6 October 2026, evening UK. Lane 7 of the Horizon programme. Worktree `/Users/joshm/Projects/igneum-wt-horizon` (branch `horizon`, at origin/master 3f4f719). Josh's words: "research anything else that we can research too, predictions, forward thinking, what we can actually do that has not been done or applied, think outside the box", and "be revolutionary". Main's bar: not features, but ideas that change what a proof-of-work chain is or what a GPU owner is to the world, each with its evidence, cost, gate, and the attack a Monero or Kaspa core developer would mount. I argue each attack as the project's own four personas would hear it (`.claude/agents/cryptographer.md`, `consensus-engineer.md`, `execution-engineer.md`, `miner-community-lead.md`).
Nothing in this file is a prediction of the coin's price, an offer to sell anything, or a change to any consensus parameter. Every chip figure is arithmetic on cited memory and logic figures; every GPU figure names its bench entry; a figure from memory says approximate. No em dashes.
**Read for grounding:** `docs/spec/00-overview.md`, `05-fees-and-economics.md`, `07-execution.md`, `09-pool-protocol.md`, `10-light-client.md`; `site/litepaper.html` (whole page, including "What Igneum does not claim"); `docs/fud-ledger.md` sections 3 (P1 to P10 present in the file; P11 to P23 are cross-referenced from the spec and round-3 entries), 4 (E1 to E8), 6 (C1 to C12), 9 (D1 to D6); `docs/commercial/prover-customer-brief.md`; `docs/design/payment-routes.md`, `developer-adoption.md`, `execution-layer.md`; `docs/analysis/chip-model-v3.md` sections 5 to 6; `docs/plans/counter-asic-3-status.md` section 4; `docs/analysis/prover-tiers-real-cards.md` (eleven rented cards, 6 October); `docs/analysis/economy-2026-10-04.md`; `docs/analysis/security-budget.md`; `docs/bench-log.md` line 2582 (rental cost of hash, measured 6 October); `vendor/rusty-kaspa/consensus/core/src/config/params.rs` (main checkout). `block-rate-devnet2.md` was still a template at 22:00 UK (RUN_A and RUN_B empty); nothing here depends on it.
**Model:** `sim/horizon/frontier/frontier_model.py` (pure Python, no numpy, about 50 ms; `python3 sim/horizon/frontier/frontier_model.py > sim/horizon/frontier/out.md`). Every table below marked "model" is printed by it; its `INPUTS` block labels each input measured, cited, designed or approximate with the source. Not run under the lock: it is arithmetic, not a measurement.
---
## 0. Everything ranked by payoff over difficulty
Payoff 1 to 5 is what the idea does for the chain's security, the coin's utility or the GPU owner's position in the world, if it works. Difficulty is Claude-side hours to a measurable prototype (Josh's rule: hours, never weeks). Verdicts: do now, prototype, watch, never. The "never" rows carry a sharp reason so the rest are not fantasy.
| Rank | Idea | Payoff | Hours | Verdict | One line why |
|---|---|---|---|---|---|
| 1 | 3.3 The weight table carried inside the recursive segment proof: a consensus proof at mergeset cost, so a browser verifies finality from one proof and asks no node for the voter set | 5 | 60 | do now (design and guest prototype) | Phase two's hardest item becomes incremental: each segment proof updates W2 by its own blue blocks; the cost is one BLS aggregate verify per 30 s inside the zkVM, which SP1 has precompiles for (approximate) |
| 2 | 3.2 Work-stake: vote weight as the external-job bond | 5 | 24 | prototype | A bond nobody can buy: 30 days of blocks. The at-risk pool income is thousands of IGN against a 0.0015 IGN coin bond (model section 3); the design stays "no stake" because weight is work, not coins |
| 3 | 3.7 Reproducible-build attestations in a registry contract; Ember refuses a release under N of M attestations | 4 | 12 | do now | Bitcoin's guix.sigs with the chain as the sigs repo; closes the devnet's "release key acts as operator" sentence (litepaper, Governance) |
| 4 | 3.4 A WebAssembly verifier of the wrapped block proof in the tab | 4 | 16 | do now | Three working precedents (a16z Helios WASM, ProjectZKM ziren-wasm-verifier, xycloo wasm-groth16-verifier); the certificate half already runs at 58 to 155 ms (bench-log) |
| 5 | 3.8 Ember as node, wallet and light client for everyone: node count equals miner count | 4 | 20 | do now | 1,000 testnet miners become 1,000 verifying nodes; Monero shows about 5,000 peers in 72 h (monero.fail), Ethereum 8,136 execution nodes (ethernodes); Sybil counts are irrelevant here because nothing counts nodes |
| 6 | 3.15 A 30-s randomness beacon from the checkpoint VDF | 4 | 24 | prototype | The pipeline exists (10-min VDF, 4.47 ms verify); a drand-class beacon with no league and 120x drand's latency; honest limit is the class-group ASIC (Chia timelords) |
| 7 | 3.1 The reward rule that prices rented hash out (pay per block falls when hash arrives faster than the 30-day weight) | 4 | 40 | prototype | Doubles the renter's break-even price (model section 2) and routes the cut to the proving pool, not to incumbents; the cost is a 30-day income ramp for honest newcomers, which the vote already imposes |
| 8 | 3.5 Continuous miner-voted parameters, bounded per block like Ethereum's gas limit, for `B_p`, the floors and the window | 3 | 30 | prototype | Replaces two-week 60 percent proposals with a drift anyone can read on the chain; Kaspa's Crescendo was a fixed DAA score (`params.rs:648`), Bitcoin's BIP9 a 95 percent tally; neither moves a number continuously |
| 9 | 3.16 The hourly program swap as a research dataset and the fleet library as a product | 3 | 10 | do now | 8,760 random kernels a year, compiled on three vendors with per-variant timings; compiler and GPU-architecture researchers have no such corpus; income small, standing large |
| 10 | 3.13 Igneum as the settlement layer for GPU rental | 3 | 40 | prototype (escrow plus sampled verification) | The chain's fee is 3 to 5 orders under a 7 to 15 percent platform take (model section 5), but it can settle only what it can verify; the verifiable subsets are named |
| 11 | 3.6 Treasury-less audit funding: burn-redirect bounties by 60 percent signal, review escrow on upgrade proposals | 3 | 20 | watch | The money exists only when the chain is used (USD 1,600 to 16,000 per 30 days at launch traffic, model section 4), and it is the switch spec 5.5 removed, with a veto |
| 12 | 3.14 Proofs sold to AI labs for verifiable inference | 2 | 40 | watch | zkLLM: 803 s of proving per forward pass on LLaMA-2-13B (arXiv 2609.27367 citing 2404.16109); the competitor is a USD 0 TEE attestation on H100-class cards the fleet does not own |
| 13 | 3.9 Hardware wallets that verify proofs | 2 | 12 | never on the secure element; do the companion verify | A bn254 pairing on a Cortex-M-class secure element is seconds to minutes (approximate); Ledger and Trezor do their heavy work in the companion app, and Igneum Wallet already verifies certificates with the node's code |
| 14 | 3.12 The fleet as a public compute market (rendering, inference) priced in IGN | 2 | 60 | never for unverifiable work; do for the verifiable subsets | An escrow without a verifier is a trust-me payment with lower fees; Render uses result quorums, Akash reputation, io.net attestations (secondary), none of which a chain can check |
| 15 | 3.11 Miners paid for proving others' chains as the main income, the lottery a tiebreaker | 1 | 0 | never by 2030; watch | All of Ethereum L1's proving at the Sep 2026 tracker cost is USD 36 a day; Igneum's year-1 emission is USD 13,700 a day at 0.005 (model section 7); demand must grow 1,000x against a cost curve falling 3x to 30x a year |
| 16 | 3.10 Proof-of-useful-work: the lottery hash partly a proof | 1 | 0 | never | Every coupling of leader election to proving re-opens Aleo (fastest prover wins); Ball et al. 2017 and Ofelimos 2022 show the sampleability conditions, and zkVM proving meets none of them |
Predictions (section 2) are not ranked; they are inputs. The three ideas a Monero or Kaspa core developer would not have thought of are 3.2, 3.3 and 4.3 (the hourly program as a hardware census), argued in section 4.
---
## 1. Method
Three kinds of work:
1. **Trend lines with arithmetic.** Each 2030 prediction has a cited anchor (a product page, a tracker, a standards body, a secondary analysis labelled as such) and a formula in the model script. Where the trend is from memory (consumer VRAM generations) the row says approximate.
2. **Idea arithmetic.** For the five ideas whose value depends on numbers (the rental tax, work-stake, the burn bounty, rental settlement, the proving-income ceiling) the script prints the table and the document reads it. The attack costs use the measured rental entry (`docs/bench-log.md` line 2582: 1,748 MH/s for USD 20.44 an hour, USD 0.0117 per MH/s-hour, about USD 11.69 per GH/s-hour) and are given per GH/s and evaluated at 1, 10, 100 GH/s and 1 TH/s, as the brief asks.
3. **Prior art.** WebSearch on 6 October 2026 for every idea; the paper or repository that tried it is cited, or the idea is marked "new" when none was found. Claims about other chains cite the repository file when a clone exists in the main checkout (`vendor/rusty-kaspa`) and are marked "not cloned, approximate" otherwise.
No simulator in `sim/` was re-run and no node harness was started: every idea here is a design question whose first gate is a measurement named in its section, and tonight the fleet, PC 1, PC 2 and the devnet were in use for the class v4 rehearsal and the Devnet 2 block-rate runs. What I could not run is listed in section 6.
---
## 2. Predictions to 2030
Each row: the trend, its anchor, the arithmetic, and what it does to the chip model and the proving tiers. Tables marked model are `frontier_model.py` section 1.
### 2.1 GPU memory per card
| Year | Flagship consumer card | GB | Label |
|---|---|---|---|
| 2016 | GTX 1080 | 8 | approximate (from memory) |
| 2018 | RTX 2080 Ti | 11 | approximate |
| 2020 | RTX 3090 | 24 | approximate |
| 2022 | RTX 4090 | 24 | approximate |
| 2025 | RTX 5090 | 32 | cited (`chip-model-v3.md` 5.1: 16 x 2 GB GDDR7 on 512 bits) |
Compound growth 1.167 a year (4x in 9 years). Extrapolated: 51 GB in 2028, 69 GB in 2030 (model). The module arithmetic is sharper than the curve: a 512-bit board is 16 devices; Micron has ended 2 GB GDDR7 (TrendForce, 24 September 2026, via `chip-model-v3.md` 5.1) and 3 GB devices are USD 60 to 70, so the next flagship is 48 GB (16 x 3 GB) and 64 GB is the 2030 shape. The RTX 60 series on Rubin (GR20x) is reported for 2028 after two slips (kopite7kimi via videocardz.com and thepcenthusiast.com; rumour, not a product). Datacentre: HBM4 at 36 GB per 12-high stack, 288 GB per GPU on Rubin NVL72 (Wikipedia HBM page, Micron March 2026 production).
What it does to the chip model: nothing for the stored-dataset chip (f = 1), whose memory is already 24 to 32 GB against a dataset of 2 GiB growing to 4 GiB at year 4 (`chip-model-v3.md` 5.7: dataset size is "not a lever against this chip"). What it does for miners: the dataset schedule (2 GiB, doubling at years 4, 12, 28; litepaper) stays under every card from 8 GB for twelve years, and the proving side, not the mining side, is what wants VRAM (section 2.6).
### 2.2 Memory dollars per GB
| Point | USD per GB | Source |
|---|---|---|
| 2023, GDDR6 | 3.38 | Tom's Hardware, "GDDR6 VRAM prices plummet", USD 27 per 8 GB |
| 2025, GDDR6 | 2.50 | TechSpot, "AI is eating all the DRAM" (2026) |
| 2026, GDDR6 | 3.30 | TechSpot, same |
| Sep 2026, GDDR7 2 GB device | 10.00 | TrendForce via `chip-model-v3.md` 5.1 |
| Sep 2026, GDDR7 3 GB device | 21.67 | TrendForce (USD 60 to 70 per device) |
| Oct 2026, HBM3E 36 GB stack | about 8.3 | siliconanalysts.com/data/hbm-pricing (factory gate; contract about 2x), approximate |
| Oct 2026, HBM4 36 GB stack | about 15.3 | siliconanalysts.com (USD 550 per stack); Samsung quoting USD 4.50 to 4.90 per Gb for HBM4 against 1.50 for HBM3E (BigGo Finance), approximate |
GB per dollar fell in 2026 for the first time in a decade and DRAM supply is forecast tight through 2027 with new fabs in 2028 (SoftwareSeni "HBM4 delays and GDDR7 shortages"). Memory is reported at 70 to 80 percent of the bill of materials of high-VRAM consumer cards by late 2025 (BuySellRam, secondary).
What it does to the chip model: the f = 1 chip and the GPU buy the same devices, so the ratio of their memory bills is fixed; what moves is the share of each bill that is memory. The chip's bill is about 70 percent memory (USD 320 of USD 470, `chip-model-v3.md` 5.4) and the 5090's about 16 percent at MSRP (USD 320 of USD 1,999) or 9 percent at the 2026 street price of USD 3,695 (localaimaster.com). A doubling of device prices raises the chip's cost 1.7x and the card's 1.1x to 1.2x: **the stored-dataset chip gets dearer relative to the GPU through 2027**, and the dollars-per-MH/s row (USD 2.8 against 14.7 at MSRP, 5.4 against 27 at street prices) narrows a little and no more. The per-joule row does not move at all, and per joule is where the chip wins (section 2.3).
### 2.3 Random-read bandwidth: GDDR7, HBM3E, HBM4
The lottery is latency-bound: one hash advances one dependent 4-byte read per memory latency, so the number that matters is random reads per second per watt, not GB/s (`chip-model-v3.md` 5.3 and 5.5). That ceiling is set by bank count and activate windows (tRC, tFAW), not by pin speed, so 48 Gbps GDDR7 is 28 Gbps GDDR7 here, and HBM3E is HBM3.
| Memory system | Reads/s ceiling | Why | Label |
|---|---|---|---|
| GDDR7, 16 devices, 512-bit (the 5090 board) | 21.3 G | 4 activates per 12 ns per channel x 64 channels | approximate (`chip-model-v3.md` 5.3) |
| RTX 5090 measured | 17.5 G | 82 percent of the ceiling | measured (bench-log Counter ASIC 2.0) |
| HBM3 or HBM3E, one stack | 10.7 G | 16 channels | approximate |
| **HBM4, one stack** | **21.4 G** | JEDEC JESD270-4 raises channels per stack from 16 to 32, each with two pseudo-channels (allaboutcircuits.com, EDN), which doubles activate parallelism if tFAW per channel holds | approximate, derived (model 1.3) |
| A 48 GB GDDR7 board | 21.3 G | capacity does not add channels | approximate |
The f = 1 chip in 2028 on HBM4 (model 1.4; every figure arithmetic, approximate):
| Chip | MH/s per chip | W bare / with a 150 W shadow core at k = 1 | uJ per hash bare / shadow | Gain per joule vs the 5090 bare (2.40 uJ) | Gain under the class v4 shadow (card 2.95 uJ at N = 100,000) |
|---|---|---|---|---|---|
| GDDR7 f = 1 (today's row) | 166 | 78 / 228 | 0.47 / 1.37 | 5.1x | 2.2x |
| HBM3 one stack | 84 | 27 / 177 | 0.32 / 2.12 | 7.5x | 1.4x |
| HBM4 one stack (2028) | 167 | 36 / 186 | 0.22 / 1.11 | 11.0x | 2.6x |
**Prediction:** HBM4 doubles the stored-dataset chip's rate per stack at about the same watts, so its bare per-joule edge rises from about 7x to about 11x, and under the class v4 shadow from about 2.3x to about 2.7x. The lever that answers it is N, the program work in the latency shadow: at N = 200,000 the HBM4 chip reads 1.74x at k = 1, at N = 330,000 (the 5090's full ALU budget) 1.33x (model 1.4). The verifier cost is N x 32 ops per warp: about 1 ms at N = 100,000 on one M5 Max core (measured class, counter-asic-3-status item 8), about 3 ms at 330,000, inside the 10 ms gate; the 2019-class core is unmeasured (O-1.14).
**Consequence and proposal (for the coordinator, not a change tonight):** write the schedule for N into the era draw at genesis, the way the dataset size already is: a candidate is a doubling of N per era until the verifier gate binds (about 1,000,000 ops, 10x of headroom on the M5 Max). The memory generation it answers arrives every two to three years and the chain must not need a human release to answer it. Per tier: no hash-rate cost while cards stay latency-bound (the M5 Max binds at about 290,000, the 9070 XT at about 650,000, approximate), watts up toward TGP (a 5090 from 326 toward 575 W, which the Ember power cap already manages), a verifier cost nodes and pools pay in milliseconds.
### 2.4 Price per card
| Card | Launch MSRP | 2026 street | Source |
|---|---|---|---|
| RTX 5090 | USD 1,999 | USD 3,695 to over 5,000 | `chip-model-v3.md` 5.1; localaimaster.com; tech-insider.org ("RTX 5090 tops USD 5,000"), secondary |
| RTX 4090 | USD 1,599 (approximate) | rental USD 0.28 to 0.60 an hour (gpus.io median) | rental cited, MSRP from memory |
Prediction: consumer card prices track memory prices through 2027 and ease in 2028 when fab capacity lands. For Igneum the price per card matters twice: the honest fleet's capital cost (not in the security budget, which is power only, `security-budget.md` section 6) and the renter's hourly price, which fell to USD 0.21 to 0.44 per 5090-hour on Vast.ai (getdeploying.com, 6 October 2026) even as purchase prices rose, because rented supply is sunk capital. **The rental market, not the purchase market, prices the 51 percent attack**, and the measured entry is USD 11.69 per GH/s-hour at RunPod list prices with the market unable to supply 20 more pods when asked (bench-log line 2582). At a TH/s: USD 11,700 an hour, and no supply.
### 2.5 The chip-fab cost curve
| Node | Mask set | Source | Igneum reading |
|---|---|---|---|
| 28 nm | USD 1 to 3 M | TubeTime (3 M); VBsemi (over 1 M) | The f = 1 memory-controller chip: no mixer on the die, a USD 5 to 30 M project (`asic-resistance-history.md` 2.5) |
| 7 nm | USD 10 to 15 M | VBsemi; Hacker News thread | The f = 0 recompute chip with 256 MiB on die: USD 50 to 75 M |
| 5 nm | USD 6.5 M (2026 data) to 30 M (2023 estimate) | siliconanalysts; HN | The shadow core (30 mm^2 at N5 for N = 100,000) drags the f = 1 chip toward this node, or to a reticle-class 28 nm die |
| 3 nm | USD 15 to 22 M (Q4 2025), up to 40 M (older) | siliconanalysts; semianalysis | Not relevant to a chip whose cost is memory |
Prediction: mask cost at a fixed node falls (5 nm quoted at 30 M in 2023, 6.5 M in 2026) while the leading node rises. So the shadow lever's fab-bill teeth weaken about 4x over three years; what holds in 2030 is the rate and joule arithmetic of 2.3, not the bill. **The stored-dataset chip gets cheaper to design and dearer to populate** through 2027, and the net is roughly flat against the GPU on dollars; per joule it gains with each memory generation unless N grows with it.
### 2.6 zkVM proving speed per dollar
| Point | USD per Ethereum L1 block proof | Hardware | Source |
|---|---|---|---|
| Jan 2025 | 1.69 | about 160 RTX 4090s for 90 percent real-time, USD 300 to 400 K cluster | HackMD "Ethproofs 2025 review" (willcorcoran); Succinct SP1 Hypercube blog (May 2025), secondary |
| Dec 2025 | under 0.04 | 16 x RTX 5090 (SP1 Hypercube: 99.7 percent of blocks under 12 s; cluster under USD 100 K); Pico Prism 16 GPUs (Brevis blog, Feb 2026) | same, secondary |
| Apr 2026 | | Cysic Venus 7.4 s on 24 GPUs | bex.co, secondary |
| Aug 2026 | | ZisK p99 9.62 s on 4 x RTX 5090 | GitHub comparative analysis (Ricosworks1), secondary |
| Sep 2026 | about 0.005 | "sub-half-cent" fields on ethproofs | same, secondary |
The 20-month ratio is 338x, about 33x a year (model 1.6). That cannot continue: it is software catching up with hardware. The table below uses 1.5x, 3x and 10x a year from the shard times measured on eleven rented cards on 6 October (`prover-tiers-real-cards.md`, the v1 shard, 4.7 M cycles).
| Card | Beside the miner today, s | Alone today, s | 2028 at 1.5x a year (beside / alone) | 2028 at 3x | 2028 at 10x | Under 10 s beside the miner in 2028? |
|---|---|---|---|---|---|---|
| RTX 3060 12 GB | 37.5 | 14.4 | 16.7 / 6.4 | 4.2 / 1.6 | 0.4 / 0.1 | at 3x or more |
| RTX 4060 8 GB (core-only beside) | 22.1 | 18.4 | 9.8 / 8.2 | 2.5 / 2.0 | 0.2 / 0.2 | even at 1.5x |
| RTX 4070 12 GB | 27.3 | 12.1 | 12.1 / 5.4 | 3.0 / 1.3 | 0.3 / 0.1 | at 3x or more |
| RTX 4060 Ti 16 GB | 34.6 | 11.6 | 15.4 / 5.2 | 3.8 / 1.3 | 0.3 / 0.1 | at 3x or more |
| RTX 3080 10 GB | 25.6 | 7.1 | 11.4 / 3.2 | 2.8 / 0.8 | 0.3 / 0.1 | at 3x or more |
| RTX 3090 24 GB | 19.9 | 14.9 | 8.8 / 6.6 | 2.2 / 1.7 | 0.2 / 0.1 | even at 1.5x |
| RTX 4090 24 GB | 26.1 | 6.3 | 11.6 / 2.8 | 2.9 / 0.7 | 0.3 / 0.1 | at 3x or more |
| RTX 5070 12 GB | 37.2 | 4.8 | 16.5 / 2.1 | 4.1 / 0.5 | 0.4 / 0.0 | at 3x or more |
| RTX 5090 32 GB | 10.7 | 6.3 | 4.8 / 2.8 | 1.2 / 0.7 | 0.1 / 0.1 | even at 1.5x |
**Prediction:** at the floor rate (1.5x a year, which is the GPU hardware cadence alone) only the 24 GB and 32 GB cards mine and prove inside 10 s in 2028; at 3x a year (half the historical software rate) every card from the 3060 up does, and an 8 GB card alone proves in 2 s. **The block-proof target ("under 10 s as provers improve", CLAUDE.md) should be written as a function of the measured fleet median shard time, re-read each era, not as a date.** Per tier: a 12 GB desktop card is the swing tier; under 1.5x it proves alone in under 7 s but not beside its miner, so the hand-off profile (`prover-tiers-real-cards.md`) is the thing to ship, not a bigger card.
What the cost curve does to the economics: the dollars per proof on the open market fall as fast as the volume rises, which is the arithmetic behind the "never by 2030" of 3.11.
---
## 3. The ideas, one section each
Each section: the idea in a paragraph; why nobody shipped it (cited, or "new"); what Igneum already has; the model; hours; the gate; the per-tier consequence; the Monero attack; the Kaspa attack; the verdict.
### 3.1 The reward rule that prices rented hash out
**The idea.** The pulse attack M14 (a renter arrives, mines for an hour, leaves) is recorded in the ledger and the finality rule already denies rented hash a vote. The reward side is untouched: a renter is paid per block like anyone. The rule: the block subsidy paid to a producer is multiplied by `m = clamp(W30 / H_now, 0.25, 1)`, where `W30` is the 30-day work-weighted hash the finality window already computes (blue blocks per DAA second over the W2 window) and `H_now` the DAA-window estimate. The remainder `(1 - m)` of the subsidy goes to that block's proving-pool escrow, not to incumbents and not to a burn. Fees are untouched. Hash that arrives faster than the 30-day weight can follow is paid less per block until the weight catches up.
**Why nobody shipped it.** Bitcoin Cash's EDA and every emergency rule since adjusted difficulty, never pay (approximate; the consensus engineer's list). Kaspa's DAA retargets per block over a sampled window (`vendor/rusty-kaspa/consensus/src/processes/difficulty.rs`, main checkout) and pays per block. Monero's RandomX changes what hash is, not what it is paid. Ethash coins and Ergo pay per block. No chain ties the subsidy to the ratio of fresh to sustained hash, because no chain had a sustained-hash number in consensus; Igneum has it in W2. Prior art searched: none found. New.
**What Igneum already has.** W2 (blue blocks per key over 30 days) and the DAA estimate are both in the node; the proving-pool escrow exists in execution state (spec 7.7 item 6); the emission split is a state transition by rule (design 1.1).
**The model** (frontier_model.py section 2, the measured rental entry USD 11.69 per GH/s-hour):
| Network hash | Attacker adds | H_now / W30 | m | Attacker IGN per hour, no rule | With rule | Rent USD per hour | Break-even IGN price, no rule | With rule |
|---|---|---|---|---|---|---|---|---|
| 1 GH/s | 1 GH/s | 2.0 | 0.50 | 57,038 | 28,519 | 12 | 0.00020 | 0.00041 |
| 10 GH/s | 10 GH/s | 2.0 | 0.50 | 57,038 | 28,519 | 117 | 0.00205 | 0.00410 |
| 100 GH/s | 100 GH/s | 2.0 | 0.50 | 57,038 | 28,519 | 1,169 | 0.02049 | 0.04099 |
| 1 TH/s | 1 TH/s | 2.0 | 0.50 | 57,038 | 28,519 | 11,690 | 0.20495 | 0.40990 |
| 10 GH/s | 50 GH/s | 6.0 | 0.25 | 95,064 | 23,766 | 584 | 0.00615 | 0.02459 |
A renter who doubles the network needs twice the coin price to break even; one who sextuples it needs four times. The diverted subsidy (57,000 to 85,000 IGN an hour in these rows) reaches the provers of the same blocks, who are the sustained population by sortition weight (spec 7.2). The honest-growth cost: a listing that doubles honest hash overnight cuts every miner's subsidy per block in half on top of the halving difficulty already imposes, until W30 catches up (the 30-day ramp of ledger C7, 0.9x by day 28 to 31). The floor 0.25 bounds the worst case at 4x.
**Hours.** 40: the rule in the coinbase state transition (8), W30 as a consensus value from the window the finality module keeps (8), the economy simulator scenario f with and without the rule (8), the fast-time harness timestamp test (8), spec text and tests (8).
**The gate.** (a) `sim/economy/sim.py` scenario f (a pool with the network's own hash arriving on day 10): incumbents' income under the rule above the no-rule row for all 30 days, newcomers' below. (b) The fast-time harness with headers back-dated inside Kaspa's 132-s tolerance: `m` moves under 2 percent. (c) Scenario d (a 20 percent operator vanishes): `m` stays at 1 (hash fell, so H_now < W30 and nobody is cut).
**Per tier.**
| Tier | Consequence |
|---|---|
| Home miner, 8 to 32 GB | In a doubling month, 25 percent of the pre-event subsidy instead of 50; unchanged in a steady month; a newcomer earns half-rate for its first month, as it votes nothing for its first month already |
| Rig | The same per block; a rig that joins during a listing spike earns half for a month |
| Pool user | The same through PPLNS; the pool's dashboard should show `m` |
| Prover | Gains: the diverted share lands in the pool escrow and is paid by sortition weight |
| Holder | Emission schedule unchanged in total; a larger share of it reaches sustained keys during spikes |
| Rollup customer | Nothing |
| Node operator | One more consensus value (W30) and one multiplier in the coinbase rule |
**The Monero core developer's attack.** "You have built an incumbents' cartel. Every rule that pays old miners more than new miners entrenches whoever was there first; RandomX exists so that a newcomer with a laptop earns exactly what a veteran earns per hash. Your 'to the pool, not incumbents' is cosmetic: sortition is by weight, and weight is the incumbents. And your W30 is your own finality window, so a 30-day-old farm that goes dark and returns is 'sustained' while a thousand honest newcomers after a listing are taxed for a month. Qubic reached 23 to 34 percent of our hash for weeks in August 2025 (arXiv 2512.01437) by renting and by paying miners in its own token; your rule would have taxed the honest miners who moved to P2Pool to fight it, because they were new keys." Answer: the tax is per block not per key, so moving pools under the same key costs nothing (spec 9.6), and a returning farm's W30 is its own blocks, which it did not make while dark. The entrenchment point stands and is the cost the gate measures.
**The Kaspa core developer's attack.** "H_now is your DAA estimate and the DAA is manipulable by timestamps inside the tolerance; a producer can lower H_now for its own block by back-dating within 132 s and raise its own m. Second, W30 is a function of the DAG past and differs between two honest tips every second; a reward that depends on it makes two honest nodes disagree about the coinbase amount of the same block unless W30 is read at a fixed ancestor (the checkpoint), and then it lags. Third, you have made emission depend on a window of 2.6 million blocks: your pruning point must now keep that window's per-second counts, which Kaspa prunes." Answer: read both numbers at the block's selected parent's last certified checkpoint (deterministic, in every node's past), accept the 30-s lag, and the timestamp test is gate (b). The pruning point already keeps the W2 window for finality (spec 3), so no new retention.
**Verdict: prototype.** Doubles the renter's break-even and costs honest newcomers a month of half subsidy, which is the same month the vote already costs them; gate (a) decides whether miners will wear it.
### 3.2 Work-stake: vote weight as the external-job bond
**The idea.** The external job market needs a bond because a customer waits (spec 5.4, O-5.6: `maxPgas x f_p x 1.5` in IGN, slashed on a late or bad proof). Replace the coin bond with the key's 30-day vote weight: a key that claims an external job and delivers late or wrong loses a share `s` of its weight for 30 days, the way equivocation strips 100 percent (spec 3.6). Weight is blue blocks. It cannot be bought, borrowed or bridged; it can only be mined, in public, over 30 days. The job market then needs no IGN escrow from the prover, which removes the capital barrier that keeps home cards out of Boundless (ZKC collateral) and Succinct (PROVE staking, docs.succinct.xyz/docs/provers) while keeping a bond larger than either.
**Why nobody shipped it.** Every proving network bonds in its own token (Boundless: stake scales with aggregate proving work per epoch, docs.boundless.network/zkc/mining/overview; Succinct: stake required to bid, more stake more concurrent auctions). No proof-of-work chain had a non-transferable, slowly earned weight per key until Igneum's finality rule. Decred's tickets are bought with coins; Ethereum's slashing is coins. New.
**What Igneum already has.** W2 per key, the 30-day strip for equivocation, the sortition that already draws assignees by weight (spec 7.2), the job record format (design 6).
**The model** (frontier_model.py section 3):
| Key's hash share | Blocks per 30 days at 1 bps | 30-day pool income at risk, IGN | Lost at s = 25 percent | Lost at s = 100 percent | The coin bond for one 1 B-cycle job at the floor |
|---|---|---|---|---|---|
| 0.01 percent | 259 | 1,643 | 411 | 1,643 | 0.0015 IGN |
| 0.1 percent | 2,592 | 16,427 | 4,107 | 16,427 | 0.0015 IGN |
| 1 percent | 25,920 | 164,271 | 41,068 | 164,271 | 0.0015 IGN |
| 10 percent | 259,200 | 1,642,706 | 410,676 | 1,642,706 | 0.0015 IGN |
The at-risk amount is five to nine orders of magnitude above the designed coin bond, before counting the lost vote. A late proof must be defined in DAA time against the claim (the P9 decision's 120-s claim timeout is the starting value), with one strike of grace per 30 days so a partition does not strip an honest key on its first miss.
**Hours.** 24: the strip rule in the finality module keyed by a job-fault record (8), the fault record in the job contract (design 6) with the evidence a node checks (8), tests and a devnet injection script (8).
**The gate.** Phase 4 devnet: 1,000 jobs with a 10 percent injected late or wrong rate: every injected fault stripped; zero honest keys stripped across a 60-s partition; a customer's job never waits more than the claim timeout plus one open window.
**Per tier.**
| Tier | Consequence |
|---|---|
| 8 GB solo miner below dust (under 100 blocks in 30 days) | No weight, so no external jobs under this rule; shards (no bond) unchanged; the pool protocol gives it a route (the pool's key, the pool's weight) |
| 12 to 32 GB home miner above dust | Takes external jobs with no IGN locked; one bad job costs a quarter of a month's sortition income and a quarter of its vote for 30 days |
| Rig | The same, at the rig's weight; the rig operator's whole weight backs each job, so a rig claims only jobs it can finish |
| Pool user | The pool's weight is the bond; the member's share of pool income carries the pool's record |
| Prover | A reputation nobody can buy, visible on chain per key |
| Holder | No IGN is locked in bonds, so no bond capital sits idle |
| Rollup customer | A bond measured in 30 days of public mining instead of a token balance; the customer brief's "your chain's own bond and slashing apply" becomes "Igneum's weight is at stake" once jobs settle on Igneum |
| Node operator | One more strip condition in a module that already strips |
**The Monero core developer's attack.** "CLAUDE.md says no stake anywhere in consensus. You have just made the vote weight a stake: it is at risk for an execution-layer fault. Whatever you call it, a prover now rationally hedges by splitting its mining across two keys, one that votes and never proves, one that proves and holds dust weight, which your F17 says buys nothing for the vote but buys everything here: the proving key has nothing to lose. So the bond is only real for operators too small to split, which is backwards." Answer: the sortition draws assignees in proportion to weight (spec 7.2 step 2), so a dust key is drawn with probability near zero and the proving key must carry weight to be assigned at all; the hedge costs the prover its assignments. The attack is right that this is a stake of work; the design's "no stake" means no coin balance in consensus, and that still holds. The spec wording needs the distinction.
**The Kaspa core developer's attack.** "'Late' is not a fact on a DAG. A proof included in a block at DAA score D is late relative to the claim at D minus T only along a chain; a reorg moves D. You will strip a key on one chain and not on another, and the strip is a consensus input to finality. Also, the fault evidence rides in blocks, so a producer who dislikes a prover can withhold its proof for T seconds and then carry the fault record. That is a griefing vector you did not have when nothing waited on a prover (ledger P9)." Answer: the evidence rule must be relative to the carrying block's own chain (as proof records are, spec 7.2 item 5), the deadline must be long relative to merge depth, and the withholding vector is real: a proof gossips to every producer, so withholding needs a majority of producers for T seconds, and a strip is only applied if no block in the carrier's past carried the proof. The griefing cost is one window of a majority, the same bound the finality rule already lives with.
**Verdict: prototype.** The bond nobody can buy; the two attacks name the wording (work-stake is not coin-stake) and the rule (evidence relative to the carrier's chain) that the prototype must carry.
### 3.3 The weight table carried inside the recursive segment proof: finality attested by provers, a consensus proof at mergeset cost
**The idea.** Phase two's consensus proof is scoped as "a zkVM program over the 30-day header window and the vote certificates" (design 7; O-10.8), which is 2.6 million headers per proof and the reason it is phase two. But the segment proof already recurses: segment N verifies segment N minus 1 (spec 7.8 item 1, measured on the 5090). Carry the W2 weight table as a public commitment inside that recursion. Each segment's guest takes the previous segment's committed weight table, adds the blue blocks of its own mergeset per vote key hash (the segment is exactly the mergeset of its chain block, spec 7 terms), ages out the blocks that left the 30-day window, and commits the new table. Every 30 s, when a certificate exists for a checkpoint inside the segment, the guest verifies the BLS aggregate against the table it holds and emits "checkpoint i certified under rule v2 with x percent of total weight". The segment proof then attests both execution and finality, and a light client verifying one wrapped proof learns the certified checkpoint and the state root with no voter set fetched from any node. The provers are the attesters of finality, by construction, with no new role.
**Why nobody shipped it.** Ethereum's sync-committee light clients (Helios, a16zcrypto.com "Building Helios") trust a committee and fetch it; Succinct's eth-proof-of-consensus (github.com/succinctlabs/eth-proof-of-consensus) proves sync-committee signatures in a SNARK but over a fixed committee, not a weight table that moves with every block. No proof-of-work chain has a weight table to carry. Mina carries a recursive proof of the whole chain but its consensus is stake (approximate, not cloned). The incremental-weight-in-recursion form: new.
**What Igneum already has.** The aggregator guest with `chain_len` and `prev` (spec 7.8 item 1), the 340-byte `BlockOutput`, the canonical voter list and bitmap (spec 3.10 C3), SP1 with BLS12-381 precompiles (approximate: SP1's precompile set includes bls12-381 field operations; the pairing cost inside the guest is unmeasured), O-10.2 which already asks for a header commitment to the weight table.
**The model.** Cost per segment: the segment adds at most 180 blocks per chain block at 1 BPS (mergeset limit, spec 7.1) times `N = 8` chain blocks, so about 1,440 table updates (a hash-map add and an age-out) per segment, which is negligible beside the execution; plus one BLS aggregate verification per certificate, at most one per 30 s. The BLS verify is the cost: G1 aggregation of up to V keys and one pairing. In SP1 with the bls12-381 precompiles a pairing is of the order of tens of millions of cycles (approximate, from memory of the precompile benchmarks; unmeasured here), so at 1 pgas = 1,000 cycles it is tens of thousands of pgas, about one shard's budget (`S_p` 30,000 pgas) per 30 s. That is a real cost: about one extra shard per 30 blocks, 3 percent of proving capacity at launch traffic. The table commitment is 32 bytes in the public values; the voter list for a 10,000-key table is 10,000 x 60 bytes = 600 KB of witness per segment, which gossips with the shard witnesses (design 5.1: witnesses are not consensus data).
**Hours.** 60: the guest's table update and commitment (16), the BLS verify inside the guest and its cycle count on the 5090 (16, needs PC 2 or a fleet box), the light-client path that reads the certified index from the public values (8), the spec text for 7.8 and 10 (8), a fast-time run where a partition's two certificates are both presented to the guest and it accepts one (12).
**The gate.** (a) Cycle count of one certificate verification inside the guest under 50 M cycles on the pinned SP1 (so under two shards). (b) The browser card (spec 10.8) shows "voter set: verified" with no node asked for the set. (c) The fast-time C4 scenario (a certificate over a chain the node is not on, `docs/fud-ledger.md` C4 sweep): the proof refuses a certificate whose signers' weight at that block is under 2/3 of the table it carries.
**Per tier.**
| Tier | Consequence |
|---|---|
| Home miner, any card | Nothing changes in mining; a 12 GB prover's shard gets the certificate verification about once in 30 shards |
| Rig, prover | About 3 percent more proving work at launch traffic, paid from the same pool; the aggregator's record grows by 32 bytes |
| Pool user | Nothing |
| Holder, wallet user | A phone or tab verifies "locked" from one proof and trusts no node for the voter set; the spec 10.1 row "voter list from nodes" is deleted |
| Rollup customer | The bridge on Ethereum verifies one wrapped proof and needs no relayer or committee: this is the proof bridge of spec 7.3, delivered earlier |
| Node operator | The witness gossip carries the voter list per segment (600 KB at 10,000 keys) |
**The Monero core developer's attack.** "You have moved finality's safety from a BLS signature every node checks to a SNARK every node trusts. A soundness bug in SP1 (ledger P7) now forges not only a state root, which full nodes veto by native execution, but a certificate, which full nodes cannot veto because they verify the real BLS certificate separately and will disagree with the proof. Which do you believe? And your 2/3 test inside the proof is against a table the proof itself computed; a bug in the table update is a bug in finality, and it ships in a guest program, not in node code anyone reads." Answer: full nodes keep verifying the BLS certificate natively and the native-execution veto extends to the public values (a record whose certified index or table commitment differs from the node's own is invalid, spec 7.2 item 5 as written), so a forged certificate is, as for state, a light-client problem and never a chain split. The table update is a second implementation of W2 and must be differential-tested against the node's (gate c).
**The Kaspa core developer's attack.** "The weight table is defined over the block's DAG past (W2 counts blue blocks in the past of the chain block); your segment is the mergeset of chain block C in GHOSTDAG order, so the incremental update is only correct if every block in the window is in exactly one segment's mergeset, which holds for blue and red blocks of the selected chain's mergesets, but a reorg of the selected chain re-cuts the segments and the table must be re-derived from the fork point. Your proof chain breaks at every reorg deeper than one segment, and at 10 BPS with k = 124 a reorg of 8 chain blocks is ordinary." Answer: correct, and the recursion already restarts at an unproven segment (spec 7.8 item 7, the unproven rule); a reorg deeper than a segment invalidates the records of the abandoned chain as it does today. The table commitment must therefore be part of the statement per chain block, re-proven on the new chain, which is what re-proving the segment already does. The cost at 10 BPS is the open question for the gate.
**Verdict: do now (design and guest prototype).** It turns phase two's hardest item into an incremental one on code that exists, and it is the only road to "your browser verifies Igneum" with no node in the trust row.
### 3.4 A WebAssembly verifier of the wrapped block proof in the browser
**The idea.** The homepage card verifies a BLS certificate in JavaScript today (58 to 68 ms warm, 139 to 155 ms cold, bench-log round 6, P3). Ship the other half: the Groth16 or Plonk wrapper of the segment proof verified in WebAssembly in the tab, with the measured millisecond count shown.
**Why nobody shipped it on a proof-of-work chain.** Because no proof-of-work chain proves its blocks. The working precedents are rollup-side: ProjectZKM's `ziren-wasm-verifier` (github.com/ProjectZKM/ziren-wasm-verifier: "Verify STARK, Groth16 and Plonk proofs in browser", one Rust codebase to native and WASM), xycloo's `wasm-groth16-verifier` (github.com/Xycloo/wasm-groth16-verifier, with a live demo), a16z's Helios shipped as `@a16z/helios` on npm with WASM bindings (github.com/a16z/helios issue 76 and the npm package).
**What Igneum already has.** `site/verify/core.js` (BLAKE2b header hashes, canonical voter list, BLS aggregate over `@noble/curves`), the SP1 light verifier as a 58 MB native binary (bench-log, "the program id split"), the `wrap` step in the `ProofSystem` trait (design 5.6) unbuilt.
**The model.** A Groth16 proof over bn254 is three group elements, about 128 bytes compressed (spec 10.5, approximate); verification is one multi-pairing. In WASM a bn254 pairing is of the order of 10 to 50 ms on a laptop core (approximate, from the ziren and xycloo demos' order of magnitude; unmeasured here). Bytes per day in phase two on-demand mode: 800 bytes per open (spec 10.5).
**Hours.** 16: wrap the pinned aggregator proof to Groth16 with SP1's wrapper on a 24 GB fleet card (8, the P3 phase 2 benchmark brought forward), compile the verifier to WASM and wire it to the card (8). The wrapper's own cost on consumer hardware is the open measurement R4.
**The gate.** The card shows the wrapped proof verified in the tab, with bytes and milliseconds, on a phone-sized viewport, against the live devnet; the number lands in the bench-log with the browser and the device.
**Per tier.** A home miner's Ember node serves the proof to the tab; a holder with no node verifies state in the tab (the state root still rests on a certificate the client is given until 3.3 lands); a rollup customer sees the verifier it will run on its own chain; node operators serve one more 128-byte object.
**The Monero core developer's attack.** "A verifier in a tab served by your domain verifies whatever your domain says the verifying key is. Your 'no middleman' is your web server. Monero's answer to this class is: run a node." Answer: correct, which is why spec 10.8 already removed "no node, no trust, no middleman" and why the phone app with a pinned seed list is the client that meets 10.6; the tab is a demonstration with its trust row stated on the card.
**The Kaspa core developer's attack.** "Fine, it verifies a proof. Of which chain? The proof commits to a chain block hash; the tab needs to know that block is on the selected chain at or below a certified checkpoint, which it asks a node for (spec 10.4 item 4). You verified the arithmetic and trusted the topology." Answer: correct until 3.3 folds the certified index into the same proof.
**Verdict: do now.** Cheap, precedented, and the phase 2 wrapper measurement has to happen anyway.
### 3.5 Continuous miner-voted parameters, bounded per block, in place of two-week proposals
**The idea.** Spec 5.8 sets a parameter the genesis rules leave to miners by a registered proposal passing 60 percent of blue blocks over two weeks. For the handful of parameters that are dials rather than switches (`B_p`, `S_p`, the base-fee floors, the exclusive window, the external claim timeout) use Ethereum's gas-limit mechanism instead: each block carries the producer's vote for each dial, the value in force at a block is the median of the window's votes, and the median may move at most 1/1,024 of its value per block, within a hard range fixed at genesis. No proposal, no bit, no two-week window, no human. Switches (a new instruction family, a proof-system version) keep the 90 percent signal.
**Why nobody shipped it this way.** Ethereum moves its gas limit by producer vote, bounded to 1/1,024 of the parent's limit per block (geth `core/block_validator.go`, VerifyGaslimit; approximate, not cloned). Bitcoin's BIP9 is a 95 percent tally over 2,016 blocks with LOCKED_IN and one more retarget before activation (bips.dev/9); BIP 135 generalised the thresholds. Kaspa's Crescendo was a fixed DAA score: `crescendo_activation: ForkActivation::new(110_165_000)` for mainnet and `88_657_000` for testnet (`vendor/rusty-kaspa/consensus/core/src/config/params.rs` lines 648 and 704; the struct at line 28), with nodes connecting only to protocol version 7 peers from 24 hours before (docs/crescendo-guide.md at v1.0.0). Monero's upgrades are scheduled hard forks, formerly every six months, now every 9 to 12 months (getmonero.org). Igneum's own 6 October incident was a fixed-height activation crossing a half-updated fleet (CLAUDE.md, Devnet 2 rules). Nobody applied Ethereum's dial to a proof-of-work chain's economic parameters. The combination is new; the mechanism is Ethereum's.
**What Igneum already has.** The header's version bits (O-5.3 candidate), the 60 percent rule, the fee parameters as `Params.fees` per network (spec 5.11), the DAA window every node computes.
**The model.** At 1/1,024 per block and 1 BPS a dial can move 2.3x in a day (1.001^86,400) if every producer votes the same way, 1.07x if 51 percent do and 49 percent vote the other way (the median moves only when a majority agrees, and then one step per block). A hostile 51 percent can therefore walk a dial to the genesis bound in days; the bound is the defence, as it is on Ethereum (the gas limit has a hard floor and no cap besides the vote).
**Hours.** 30: the vote field and median rule (10), the clamp and bounds in `Params` (6), tests including the 51 percent walk (8), spec 5.8 text (6).
**The gate.** On the fast-time harness, 100 producers at 60/40 split: the dial moves toward the 60 side at the predicted rate and stops at the bound; with a 50/50 split it does not move; a producer that votes outside the range is invalid.
**Per tier.** Miners set the dials with their blocks (Ember shows the vote and defaults to "hold"); a pool votes for its members in mode A and B templates and the member sees it (spec 9.4); provers watch `B_p` and `S_p` move with the fleet's measured shard time instead of waiting for a human; holders and rollup customers see fee floors that track usage; node operators gain one field per header.
**The Monero core developer's attack.** "You have handed the fee floor to whoever has 51 percent of blocks, with no social veto. On Monero the dynamic block size has a penalty curve exactly so that a majority cannot cheaply walk it; your 1/1,024 is a speed limit, not a cost. A pool with 51 percent lowers `f_p` to its floor, bloats blocks with wash gas it no longer pays for, and the provers eat the backlog." Answer: the base fee is burned in full, so wash gas is never free (ledger E3), and the backlog rule halves `B_p` regardless of the vote (design 4.3); the range bound caps the walk. The point stands that a dial needs a cost curve, not only a speed limit: the prototype should add Monero's shape (a vote away from the median costs the producer a fraction of its subsidy).
**The Kaspa core developer's attack.** "A per-block vote on a DAG: which blocks vote? Blue blocks of the selected chain's mergesets, in order, and the median over a window is a function of the block's past, fine. But a parameter in force 'at a block' must be the same for every node validating that block: use the value at the block's selected parent's checkpoint, or two honest nodes meter the same transaction at two prices. You have the same determinism bug the proof-record rule had before P11." Answer: correct; the value in force is read at the last certified checkpoint in the block's past, as 3.1's W30 is.
**Verdict: prototype.** The chain's economic dials follow the fleet without a human; the two attacks give it the two rules (a cost curve, a checkpoint-anchored read) it needs.
### 3.6 Treasury-less audit funding: bounties from the burn, a review escrow on upgrades
**The idea.** Josh removed the dev fund (spec 5.5) and the project pays audits from the Ember dev fee and founders' mined coins (litepaper). The question: money for audits that comes from users paying for something, with no standing address. Two mechanisms. (a) **Burn redirect.** The base fee burns to nobody. A reproducible break submitted under spec 0.5 and accepted by 60 percent of blue blocks over a window redirects the base-fee burn of the next 7 (execution) or 30 (consensus) days to the submitter's address, once, then returns to burning. No address exists between events. (b) **Review escrow.** An upgrade proposal under 5.7 must escrow IGN in a contract that pays reviewers named in the proposal on a 60 percent "review complete" signal, or refunds on failure; the proposer pays, which is a user paying for a thing (the right to propose code).
**Why nobody shipped it.** Zcash funds development from the block subsidy (NU6: 8 percent to Zcash Community Grants, 12 percent to a protocol lockbox, ZIP 1015); Decred from a 10 percent treasury spent by stakeholder vote, capped at 4 percent of balance a month since January 2026 (DCP-0013); Monero from the CCS, donations off-chain; Optimism from an 850 M OP reserve for retro funding. Bug bounties pay 10 percent of funds at risk (Immunefi's standard) from the protocol's own treasury; Code4rena runs contests at zero platform fee since 2025. Nobody funds audits from a burn redirect, because a burn redirect is a subsidy to a payee by rule, and the chains that wanted that built a treasury. New in form; a dev fund in substance (see the Monero attack).
**What Igneum already has.** The base-fee burn in both dimensions, the 60 percent signalling, the ledger's break-submission rule (spec 0.5), the proposal registration transaction (spec 5.8).
**The model** (frontier_model.py section 4):
| Chain traffic (fraction of full blocks) | Base fee burned per day, IGN | 30-day redirect, IGN | USD at 0.02 | USD at 0.10 |
|---|---|---|---|---|
| 0.01 | 2,592 | 77,760 | 1,555 | 7,776 |
| 0.10 | 25,920 | 777,600 | 15,552 | 77,760 |
| 0.50 | 129,600 | 3,888,000 | 77,760 | 388,800 |
| 1.00 | 259,200 | 7,776,000 | 155,520 | 777,600 |
At launch traffic a 30-day redirect is under one audit contest; at half-full blocks it is a serious bounty. The money exists only once the chain is used.
**Hours.** 20 for the escrow contract and the redirect rule as a proposal kind; 0 for the honest alternative, which already exists.
**The gate.** None that a simulator settles; the gate is Josh's: does a per-event, miner-approved payee with no standing address pass the test that removed the dev fund?
**Per tier.** Miners vote on each payout with their blocks and can refuse all of them; holders see supply that would have burned paid to a named person; a prover, rig, pool user and rollup customer see nothing unless a break affects them; the node operator gains a proposal kind.
**The Monero core developer's attack.** "This is a dev fund with extra steps. You removed a 5 percent fund because 'a switch that routes money to an address somebody controls is the first thing a critic points at'; you have now written a switch that routes money to an address somebody controls, gated by the same miners who gate everything else, and you have made the miners the judge of which cryptographer gets paid. Our CCS works because it is off-chain and voluntary and no consensus rule touches it. Keep the burn a burn." The attack is right. Answer: there is no counter besides the honest one: the alternative is the one the litepaper already states (client fee, founders' mined coins, grants off-chain), and the ledger should carry this entry as considered and rejected on the same ground as E4.
**The Kaspa core developer's attack.** "Also a soft target: a miner cartel with 60 percent invents a break, 'accepts' it, and un-burns 30 days of fees to itself. Your spec 0.5 reproducibility rule is a human judgement; consensus cannot check it." Correct.
**Verdict: watch.** Design it, do not ship it; record it in the ledger beside E4 as the honest answer to "where do audits come from with no fund": from the company's dev fee and from the people who care, in the open.
### 3.7 Reproducible-build attestations in a registry contract; Ember refuses a release under N of M attestations
**The idea.** Bitcoin Core builds with Guix and independent builders publish signed attestations of the output hashes to the `guix.sigs` repository (`bitcoin/bitcoin` PR 21462 added `guix-attest` and `guix-verify`; bitcoinops.org reproducible builds). Put the attestation registry on Igneum: a contract where a builder set posts `(release tag, artefact hash, signature)`; the Ember updater refuses to install a release whose artefact hash has fewer than N attestations from M builders listed in the release key's policy, and shows the attesters. Windows exes are already reproducible here (`-Wl,--no-insert-timestamp`, CLAUDE.md), so the hash is well defined.
**Why nobody shipped it on chain.** Bitcoin keeps its sigs in a Git repository because Bitcoin has no contract state; Ethereum clients attest off-chain. Igneum has an EVM and a signed updater that already checks a manifest hash. The on-chain registry read by the updater: new in placement, not in idea.
**What Igneum already has.** Reproducible builds on the box (`tools/build-remote.sh`, `cross-remote.sh`), the signed manifest and updater in Ember (litepaper, Ember table), the commit-string check (`tools/ci/commit-string-check.sh`), the release key in genesis (spec 8).
**The model.** Trust goes from one key (the release key, which on the devnet "acts as the operator", litepaper Governance) to N of M builders, with M growing as outside builders arrive. The failure the rule catches: a release signed by a stolen key whose hash no independent builder reproduced. Cost per release: one transaction per builder.
**Hours.** 12: the registry contract (4), the updater's N-of-M check and the display (6), the CI step that posts the box's attestation (2).
**The gate.** A release whose binary is altered after signing is refused by Ember on three machines; a correct release with N attestations installs; the registry shows both.
**Per tier.** Every miner's updater refuses an unattested binary and names who attested; a pool operator and a node operator get a chain-readable answer to "is this the binary everyone runs"; holders and rollup customers see that the "release key as operator" sentence has a closing mechanism.
**The Monero core developer's attack.** "Who are the M builders at launch? The founder, under three names. Reproducible builds are only as good as the independence of the builders, and a pseudonymous one-founder project has one builder. Gitian and Guix were worth something because dozens of people with names attested. You are moving a sigs repo on chain; you are not adding a builder." Answer: correct, and the registry is what lets a second builder exist with a public record; the gate should be the first attestation from a machine the project does not own.
**The Kaspa core developer's attack.** "A node that reads a contract to decide whether to update is a node whose update path depends on the chain being live and unforked; during the 6 October two-sided chain you would have had two registries." Answer: the updater installs nothing while finality is paused, which is a rule worth adding anyway.
**Verdict: do now.** Twelve hours, and it closes a sentence the litepaper has to carry today.
### 3.8 Ember as node, wallet and light client for everyone
**The idea.** Ember already runs a node, mines, proves and keeps a key; Igneum Wallet reads Ember's node when present. Make the one app the client for everyone: a holder who does not mine runs Ember in "verify" mode (the light client of spec 10 inside the same binary, the node card to pin a seed), a miner runs it in full mode. Every miner is a node; every holder is at least a light client; nobody runs a browser wallet against someone else's RPC by default. The node count becomes the miner count plus the holders who chose full mode.
**Why nobody shipped it.** Bitcoin Core is a node and a wallet but not a miner; miners run separate software (approximate). Monero's GUI runs a node and a wallet and can mine on the CPU (approximate; the litepaper's reason for no CPU lane is botnets). Kaspa's miners run kaspad plus a separate miner. Igneum's Ember already supervises node, miner, prover and key (litepaper, Ember table), so the step is small. Not new; the combination with the light client and the vote key in one binary is Igneum's.
**What it does to node counts and Sybil counts.** Bitcoin: 24,682 reachable nodes (bitnodes.io, 5 October 2026); Ethereum: 8,136 execution clients on ethernodes.org, 11,781 on Etherscan the same day; Monero: about 5,000 peers seen in 72 hours by one tracker (monero.fail), all secondary. Igneum's gate 4 is 1,000 independent miners; with Ember as the node those are 1,000 full nodes on day one of the public testnet, each with a vote key. Sybil counts are irrelevant on Igneum by design: nothing in consensus counts nodes or keys (spec 3.1 W6, ledger F17), so a Sybil inflates a node map and nothing else. The one place a count matters, "no verifier, no vote" (spec 9.7 item 2), is served better: every member has a verifier because the signer is the verifier.
**Hours.** 20: the light-client engine inside Ember's process with a mode switch (12), the wallet reading it (4), the node card pinning (4).
**The gate.** A machine with no GPU runs Ember in verify mode, shows "locked" from a certificate it verified, and sends a transaction; a miner's Ember shows one key in the header of its blocks and the same key signing votes.
**Per tier.** A home miner runs one program; a holder with a laptop runs a verifier instead of trusting an RPC; a pool user's member process is Ember (spec 9.1); a rig runs one signer and many workers (spec 9.6); the node operator is now everyone.
**The Monero core developer's attack.** "A node that is also a hot wallet with a vote key and a miner is one process with every secret in it; one bug in the dashboard's local HTTP server (you serve it behind a per-launch token) and the key, the vote and the coins go together. Monero separates the daemon from the wallet for this reason." Answer: the separation stays at the process level (signer, workers, node, wallet are separate processes under one supervisor), and the vote key and the payout key are different keys; the attack names the test (the dashboard token's threat model) that gate 4 must include.
**The Kaspa core developer's attack.** "Node count is a vanity metric; what matters is who produces blocks and who the honest majority peers with. A thousand Ember nodes behind home NAT accept no inbound connections, so your reachable count is your seed list plus the rigs, and an eclipse of the seed list eclipses the fleet." Answer: correct; spec 10.6's pinned seed list with identity keys and the local peer set is the defence, and the eclipse test O-3.7 with Ember nodes is the gate.
**Verdict: do now.** Twenty hours; the testnet gate then measures nodes, not only miners.
### 3.9 Hardware wallets that verify proofs
**The idea.** A Ledger or Trezor that verifies the wrapped block proof (or the certificate) before signing, so "final" is checked on the device, not in the companion app.
**Why nobody shipped it.** Ledger's secure element is EAL6+ certified and Trezor's Safe 3 and 5 use an Infineon OPTIGA Trust M (trezor.io; ledger.com), both small, slow, memory-poor chips. The cryptographic work hardware wallets do is signing; the heavy lifting (sync, proofs, history) is in the companion. A bn254 pairing on a Cortex-M-class secure element is seconds to minutes (approximate, from memory; the Bulletproofs-on-Trezor paper, eprint 2020/281, shows what Micropython on a Trezor costs for range proofs, and it is slow). Igneum Wallet already verifies the finality certificate with the node's own code (litepaper, Wallet table), which is the companion doing it.
**The model.** Certificate verify: G1 aggregation of up to 10,000 keys plus one BLS12-381 pairing; on a laptop in JavaScript 58 to 155 ms (bench-log). A secure element is 100x to 1,000x slower on scalar arithmetic than a laptop core (approximate), so 6 to 150 s per certificate, every 30 s. Not shippable.
**Hours.** 12 for a companion-side integration (the wallet already does it); the on-device path is not worth hours.
**Per tier.** Holders get "final means final" in the companion today; nobody gets it on the secure element.
**The Monero core developer's attack.** "Monero's Ledger app exists and does nothing but sign; it took years and a custom protocol (eprint 2020/281). You will not get a proof verifier onto a secure element and you should not pretend to." Correct.
**The Kaspa core developer's attack.** "A device that verifies a proof still needs to know which chain tip the proof is of; it will take that from the companion, which is the thing you did not trust." Correct.
**Verdict: never on the secure element; do the companion verify** (already done in Igneum Wallet 0.1.1; the Ledger and Trezor apps when they exist should display the companion's verified state and sign).
### 3.10 Proof-of-useful-work: the lottery hash partly a proof
**The idea as asked.** Make the leader election depend in part on proving work, so the energy that picks the block maker is useful.
**Why every attempt failed, cited.** Primecoin (2013) found Cunningham and bi-twin prime chains that nobody uses (Bitcoin Magazine, July 2013). Gridcoin pays for BOINC work and stops if BOINC stops (gridcoin.us; the 2022 "Challenges of PoUW" survey, arXiv 2209.03865). Ball, Rosen, Sabin and Vasudevan (eprint 2017/203) gave proofs of useful work from fine-grained problems (Orthogonal Vectors, 3SUM, APSP) and state the conditions: the problem must be sampleable at a tunable hardness with instances the miner cannot choose, and the verifier must be cheaper than the work. Ofelimos (Fitzi, Kiayias, Panagiotakos, Russell, CRYPTO 2022, eprint 2021/1379) got a provably secure protocol by making the work a doubly efficient local search whose usefulness is a side effect and small. The 2026 "Economics of Proof-of-Useful-Work" (arXiv 2606.06700) and the empirical study of Pearl's cuPOW (arXiv 2606.04819, "The Usefulness Gap") find the same gap between the work paid for and the work anyone wanted. I found no Coinbase paper on the subject (searched 6 October 2026); if Josh has one in mind, its title is needed. Aleo ran proving as the consensus work and the fastest prover won (CLAUDE.md: the Aleo lesson; litepaper precedents table, approximate). Boundless's PoVW (docs.boundless.network/zkc/mining/overview) pays ZKC pro rata to cycles proven per epoch with a stake that scales with the work, which is a reward for proving, not a leader election, and it is on a proof-of-stake chain.
**The sampleability problem, plainly.** A lottery needs a puzzle whose instances are drawn at random from a distribution the miner cannot steer, whose hardness is tunable by a target, and whose solution is verifiable in milliseconds. zkVM proving has none of these: the instances (segments, jobs) are chosen by users and producers, the hardness is whatever the program is, and the verifier is tens of milliseconds to seconds. Any blend ("a miner's lottery target eases in proportion to its proven cycles last hour") gives the fastest prover more blocks, which is Aleo with a cap, and a cap small enough to be safe is a reward too small to be useful.
**What Igneum already has instead.** The separation (litepaper: "the lottery and the proving are kept separate on purpose"), the 20 percent pool paid by sortition by weight, PoVW-like cycle metering through pgas.
**Hours.** 0.
**Per tier.** Nothing changes; the 12 GB card still earns from proving through the pool.
**The Monero core developer's attack.** "RandomX's whole point is that the work has no second use, because any second use is a subsidy to whoever does the second thing best, and that is a specialist. The moment your hash is 'partly a proof', the best prover is the best miner, and the best prover is a datacentre. You know this; it is in your own CLAUDE.md."
**The Kaspa core developer's attack.** "Leader election on a DAG must be a memoryless Poisson process so that GHOSTDAG's k and the orphan analysis hold; a target that depends on the miner's past hour of proving is not memoryless and your blue-set bounds no longer apply."
**Verdict: never.** Both attacks are correct and the second is fatal to the DAG analysis. The honest version of "useful work" is the one Igneum has: the same card, two jobs, two payments, no coupling. Lane 8 may take one adjacent idea by name: **the shadow-useful puzzle**, in which the program work placed in the latency shadow (class v4, 100,000 ops per hash that cost the card nothing) is itself a small verifiable sub-computation drawn from chain state (a hash-based commitment to a sampled Merkle path of the segment's state witness), so the shadow ops have a second use that does not change who wins. It changes nothing about leader election because the shadow is free; whether a useful shadow program is as chip-hostile as a random one is lane 8's question.
### 3.11 Miners paid for proving others' chains as the main income, the lottery as the tiebreaker
**The idea as asked.** Invert the design: proving is the income, the lottery only orders.
**The arithmetic** (frontier_model.py section 7):
| Income line | USD per day | Basis |
|---|---|---|
| Proving every Ethereum L1 block at the Sep 2026 tracker cost | 36 | USD 0.005 x 7,200 blocks; a buyer pays above cost, call it 10x: 360 |
| The same at the Dec 2025 cost | 288 | under USD 0.04 per block |
| All rollup proving spend (customer brief) | 8,200 to 27,400 | "low millions a year", approximate |
| Boundless, trailing day in the explorer, 4 Oct 2026 | 2 | 8.4 T cycles at USD 0.21 per billion, `developer-adoption.md` 2b, approximate |
| Igneum year-1 emission at USD 0.005 per IGN | 13,700 | 31.688 IGN per block x 86,400 |
| At USD 0.02 | 54,800 | |
| At USD 0.10 | 273,800 | |
The whole public proving market is three to four orders of magnitude under year-1 emission at any price input. The cost curve (section 2.6) falls 3x to 30x a year, so dollars per proof fall as fast as volume rises; for proving to be the main income by 2030, paid demand must grow about 1,000x in dollars. The design's own claim is the defensible one: a second income that keeps cards on after the subsidy fades (spec 5.10.2).
**Why nobody shipped it.** Succinct and Boundless are exactly this (proving as the income) and have a token for the lottery's role; their provers are datacentre operators (ledger C10). Nobody has made it a GPU home-miner's main income because the market is this size.
**Hours.** 0.
**Per tier.** The 12 GB home card earns pool emission today and job income later; the number that matters to it is the pool share, not the market.
**The Monero core developer's attack.** "Your 'paid, useful, verifiable work' line implies the work pays. It does not and will not; say so in the litepaper's income table." Answer: the litepaper already says "small market today", "upside, not a promise" (ledger P6); the arithmetic above should join it.
**The Kaspa core developer's attack.** "If proving were the income, the lottery would be a cost centre miners minimise, hash would fall to the floor, and your 51 percent cost would be the cost of a few 5090s. Keep the lottery paid." Correct.
**Verdict: never by 2030 as the main income; watch the market yearly.**
### 3.12 The GPU fleet as a public compute market beyond proofs, priced in IGN
**The idea.** Rendering, inference, transcoding, simulation sold by Igneum miners for IGN, through the same client that switches between hashing and proving.
**The honest problem.** General compute is unverifiable: a renter cannot tell a rendered frame from a cheaper one, an inference from a smaller model's, without redoing the work. The existing markets answer with trust substitutes: Render uses result quorums for graphics, Akash provider auctions plus reputation, io.net proof-of-work-style attestations (all secondary, io.net's own comparison page and a 2026 DePIN survey). None of those is checkable by a chain.
**The verifiable subsets, named.**
| Work | How it is verified | Status for Igneum |
|---|---|---|
| ZK proving jobs | The proof | The precompile (design 6), Designed |
| Deterministic recompute with sampling | Commit to every intermediate, a verifier re-runs a random fraction (Statistical Proof of Execution, arXiv 2503.18899; sampled layerwise proofs for inference, arXiv 2609.27367) | Feasible as an app on the precompile: the sampled chunk is the job; the rest is a commitment |
| Rendering with result quorum | Two or three miners render the same frame; the chain pays on agreement (Render's approach, approximate) | An app; the chain pays per agreement, cannot judge quality |
| TEE-attested inference | NVIDIA confidential computing attestation on H100 and H200 (phala.com GPU TEE) | The fleet's cards have no TEE; not Igneum's |
| Bitwise-reproducible training | Verde-style proofs of learning on a rollup (secondary, io.net comparison page) | Research |
**Hours.** 60 for a sampled-recompute job type on top of the precompile; 0 for the general market.
**Per tier.** A 24 GB card could sell sampled-recompute work; an 8 GB card cannot hold most inference models; the rollup customer is unaffected; a holder sees IGN demand only for the verifiable subset.
**The Monero core developer's attack.** "You would be Golem, Render and Akash with a worse token story and a settlement layer nobody asked for. The honest answer to 'GPU owners should be paid for useful work' is a market with reputation, and reputation is not a consensus rule." Correct for the general case.
**The Kaspa core developer's attack.** "Every second a card spends on a render is a second off the lottery; the design's own economy model shows hash falling 14 percent when external pay rises 10x (scenario b). A compute market large enough to matter would empty the lottery." Correct, and it is the reason 3.11 is never.
**Verdict: never for unverifiable work; do the verifiable subsets as apps on the precompile** (the sampled-recompute job type is the one worth 60 hours).
### 3.13 Igneum as the settlement layer for GPU rental itself
**The idea.** Vast.ai and RunPod match renters and hosts and take a platform cut; the escrow, the metering and the payout could run on Igneum, where every miner is already a host with a funded wallet and a card that is on.
**The fee arithmetic** (frontier_model.py section 5):
| Card | Vast.ai on-demand USD/h | Platform take modelled | Host loses USD per card-year | Igneum settlement per rental (2 transfers at the floor) | At USD 0.02 / 0.10 per IGN |
|---|---|---|---|---|---|
| RTX 5090 | 0.44 (getdeploying.com, 6 Oct 2026) | Vast about 15 percent (secondary) | 579 | 0.0102 IGN | 0.0002 / 0.0010 |
| RTX 5090 | 0.44 | RunPod about 7 percent (secondary: hosts keep 93) | 270 | 0.0102 IGN | |
| RTX 4090 | 0.31 | Vast about 15 percent | 408 | 0.0102 IGN | |
| RTX 4090 | 0.31 | RunPod about 7 percent | 190 | 0.0102 IGN | |
Caveat on the takes: secondary comparisons put Vast at about 15 percent and RunPod at about 7 percent; Vast's own June 2024 product update says the host fee was removed and replaced by a surcharge it does not publish, so the 15 percent is a market estimate, not a fee page. The chain's fee is three to five orders of magnitude under either. The platform's take pays for matching, images, dispute, trust and the verification of delivered work, and the last is what the chain cannot do (3.12).
**What Igneum already has.** Funded miner wallets, the pool protocol's TLS transport and member identity (spec 9), the job escrow shape (design 6), the sampled-recompute path above.
**Hours.** 40: a rental escrow contract with hourly streaming and a sampled attestation of liveness (the host signs a challenge per minute with the vote key; proves possession of the card by running one lottery warp on it, which the CPU verifier checks in 0.44 ms) (24), a client-side matching list (16). It settles payment and liveness; it does not verify the renter's workload.
**The gate.** Ten rentals between fleet boxes with one host that goes dark: the escrow pays to the minute of the last valid challenge; the renter's refund is exact; the chain fee per rental under 0.02 IGN.
**Per tier.** A home miner rents out idle hours with no platform cut and a 30-day public record as a host; a rig lists eight cards; a pool user is unaffected; the prover role and the host role compete for the same seconds; a holder sees IGN demand per rental; a rollup customer is unaffected.
**The Monero core developer's attack.** "Escrow is 1 percent of a marketplace. The 15 percent is the other 99: the people who answer when a pod dies. You will have a cheaper escrow and no renters, and every renter you do get will be running the thing Vast bans. Also: a card that is rented is a card that is not mining, so you are paying people to leave your lottery." The last point is 3.12's and stands.
**The Kaspa core developer's attack.** "Streaming payments per minute at 1 BPS are 1,440 transactions a day per rental, each burning a base fee; at a thousand rentals that is your whole block budget. Use a channel, settle twice." Correct, and the model's two transfers assume exactly that.
**Verdict: prototype** the escrow with the liveness challenge, because it reuses the vote key and the CPU verifier in a way no other chain can, and because miners are hosts already; do not call it a marketplace.
### 3.14 Proofs sold to AI labs for verifiable inference
**The state of the art, cited.** zkLLM (arXiv 2404.16109, CCS 2024) proves a 13 B-parameter LLM's inference in under 15 minutes with proofs under 200 kB, verified in 1 to 3 s; the 2026 sampled-layerwise paper (arXiv 2609.27367) measures 803 s of proving per forward pass on LLaMA-2-13B and extrapolates about 18 days per 2,000-token generation under full ZK. EZKL's median proof time on small workloads is about 8.2 s and a 100 M-parameter model is about 10,000 s per proof at today's throughput (proofoftech.org, secondary). Modulus Labs' Remainder prover was benchmarked at USD 0.085 per proof to verify on Base; the team joined Tools for Humanity in late 2024 and no longer sells (proofoftech.org). The competitor is a TEE: NVIDIA confidential computing on H100 and H200 with remote attestation, sold today at near-zero overhead (phala.com; arXiv 2607.19353 benchmarks), and sampling schemes (SPEX, arXiv 2503.18899) that are statistical, not cryptographic.
**Cost per token, approximate.** 803 s of one GPU per forward pass on a 13 B model at a USD 0.44 5090-hour is about USD 0.10 per token proven. An unverified 13 B token is of the order of USD 0.0000002 (secondary inference pricing pages, 2026). The gap is five to six orders of magnitude.
**What Igneum could sell by 2030.** Not inference proofs for frontier models. Proofs that a committed small model (under 100 M parameters) produced an output from a committed input, batched; proofs of aggregation over many small inferences; proofs of a sampled layer (the hybrid in arXiv 2609.27367) as a job type. `developer-adoption.md` 2b already draws the line at "verifiable compute, not verifiable AI".
**Hours.** 40 for a sampled-layer job type once the precompile exists; 0 today.
**Per tier.** A 24 GB card could prove a small model's inference as a job; nothing for smaller cards; a rollup customer is unaffected.
**The Monero core developer's attack.** "A lab that wants verifiable inference buys an H100 with a TEE and gets an attestation for free. Your 100,000x-slower proof is for people who do not trust NVIDIA's attestation key, and those people are not buying GPU time from strangers." Fair for 2026 to 2028.
**The Kaspa core developer's attack.** "Nothing here touches consensus; it is an app on the precompile. Stop listing apps as protocol ideas." Fair.
**Verdict: watch** the cost curve yearly; the crossing where ZK beats a TEE on cost per token is not in sight by 2030 on the cited numbers.
### 3.15 The Igneum program pipeline as a verifiable randomness beacon
**The idea.** The chain already derives an unbiasable seed once an hour: a certified checkpoint, through a 10-minute class-group VDF (spec 04; 516-byte proof, 4.47 ms verify). Run the same VDF on every certified checkpoint hash at a 30-s delay and publish the output: a public randomness beacon at 30-s cadence with no league, no threshold key and no trusted set.
**What drand is, cited.** The League of Entropy runs drand: threshold BLS over `H(round)` in unchained mode, a 2/3 threshold of a fixed set of organisations (the threshold must exceed 50 percent), quicknet at 3-s rounds since October 2023, timelock encryption built on it (docs.drand.love quicknet post and cryptography page). Its trust assumption is that under a third of a named set collude.
**The model** (frontier_model.py section 6):
| Beacon | Period | Latency | Unbiasability | Trust |
|---|---|---|---|---|
| drand quicknet | 3 s | about 3 s | threshold BLS, under 1/3 of about 20 organisations collude | a league |
| Igneum epoch seed today | 3,600 s | 600 s | certified checkpoint plus a VDF the last producer cannot evaluate in time | nobody |
| Proposed per-checkpoint beacon | 30 s | 30 to 60 s | the checkpoint is locked by 2/3 of 30-day weight before the VDF starts; a last-block grind costs a block's subsidy per try and buys a bit only if the attacker evaluates the VDF faster than the chain | nobody; the honest limit is the class-group ASIC (Chia's timelords are software or ASIC, docs.chia.net) |
Chia's hardware timelords are the precedent for "the fastest squarer learns the value first" (Boneh, Bonneau, Bünz, Fisch, eprint 2018/601 for the VDF; Chia's class-group VDF competition repository for the implementation lineage). That is a front-running edge measured in seconds, not a bias.
**What Igneum already has.** The VDF prototype (`proto-vdf/`), `seed_source` in headers, PREVRANDAO already defined from the epoch VDF (spec 7.1), the certificate every 30 s.
**Hours.** 24: a 30-s VDF parameter set and the proof relay per checkpoint (12), an RPC and a `wss` feed (6), a contract exposing the latest value and a verify function (6).
**The gate.** 2,880 values a day on the devnet for a week; every value verified by an independent client in under 5 ms; no value published before its checkpoint locked; a deliberate withholding of the last block before a checkpoint measured for its effect on the output (none, because the checkpoint is what is locked).
**Per tier.** A node operator evaluates one 30-s VDF per checkpoint (one core); a miner does nothing new; an app developer gets a 30-s beacon and timelock encryption; a holder sees a product that drand's users (lotteries, raffles on Sui, approximate) might pay gas for; a rollup customer could read it through the proof bridge.
**The Monero core developer's attack.** "Your beacon is only as unbiasable as your finality, and your finality pauses whenever under 2/3 of weight is connected (spec 03). A beacon that stops when the chain is partitioned is not a beacon; drand ran through every outage its members had because it needs a threshold, not a supermajority of all." Answer: correct; the beacon publishes nothing during a pause and must say so, which is still a stronger statement than a league's liveness.
**The Kaspa core developer's attack.** "A 30-s VDF on a 1-BPS chain is fine; at 10 BPS your checkpoints are still 30 s of DAA time, fine; but the VDF input must be the checkpoint hash as every node agrees it, and your C4 finding showed two honest nodes can hold two certified checkpoints at one index for a window. Two beacons." Answer: the beacon for index i is published only when a single certificate for i is in the past of the next certified checkpoint, which is the F24 re-determination path; one window of delay in the worst case.
**Verdict: prototype.** Twenty-four hours on code that exists, and a product no proof-of-work chain offers.
### 3.16 The hourly program swap as a research dataset; the fleet library as a product
**The idea.** Igneum generates 8,760 random GPU kernels a year, compiles each on Metal, CUDA and OpenCL, races up to 17 variants per card (lever 1, measured +17 to +21 percent on the M5 Max), and logs per-card per-variant timings to the fleet log (lever 2). That corpus does not exist anywhere: a continuous stream of random, bit-exact-across-vendors integer kernels with measured performance on every consumer GPU, under a fixed memory footprint. Publish it (the generator is public with the spec; the timings are the product) and the fleet library (the per-card best-variant table) as a dataset.
**Who would pay, what for.** Compiler teams (LLVM's NVPTX and AMDGPU backends, Apple's Metal compiler) for a regression corpus with ground truth across vendors; GPU microarchitecture researchers for a latency-bound random-read benchmark across generations (the dependent-read ceilings of `chip-model-v3.md` 5.3 are exactly what such a corpus measures); the project's own cryptanalysts (the weak-program census, `weak-program-census-2026-10-03.md`) for the distribution of program properties. Money: small (research datasets are grants and goodwill, not revenue); standing: large, and it is the public benchmark the litepaper promises for January 2027 made continuous.
**Why nobody shipped it.** RandomX programs are per hash, interpreted, and never logged; ProgPoW's period changes were never published as a corpus (approximate). New as a dataset.
**Hours.** 10: a daily export of the fleet log and the generator seed list to a public bucket with a schema (6), a README with the citation form (4).
**The gate.** One outside group cites it.
**Per tier.** Every miner's timings are in it (anonymised to card model); a 9070 XT owner sees why their card is 7x worse per joule than a 5090 on dependent reads (`chip-model-v3.md` 5.8); nothing else changes.
**The Monero core developer's attack.** "A public corpus of your programs with timings is the chip designer's training set." Answer: the generator is public already (github.com/igneum-network/spec) and a chip must run next hour's program, not last year's; what the corpus gives a chip designer is the distribution, which the spec gives too.
**The Kaspa core developer's attack.** "Not a consensus matter." Correct.
**Verdict: do now.** Ten hours and it makes the benchmark promise continuous.
---
## 4. What would make a Monero or Kaspa core developer say "I had not thought of that"
Three, with the exact reasoning each would use to attack it. The first two are 3.2 and 3.3 restated as the thing that is new; the third is new in this file.
### 4.1 Work as the only stake, and it is slashable
Monero's and Kaspa's shared premise: in proof of work nothing is at stake except the block you are mining, so misbehaviour by a miner outside block production (a bad job, a withheld proof) cannot be punished, only priced. Igneum's finality weight is a quantity that is at stake, is earned by work alone over 30 days, cannot be transferred, and is already stripped for equivocation. Extending the strip to execution-layer faults (3.2) gives proof of work a slashable bond with no coin and no stake class.
**The Monero developer's attack, verbatim form.** "Then it is stake. You have a class of participants with something to lose that others do not, and a rule that takes it from them for a judgement call. Every argument you make against proof of stake (capture, cartels, nothing-at-stake inverted into everything-at-stake) applies to a stake made of blocks. Worse, your stake depreciates on its own in 30 days, so the rational prover front-loads bad behaviour in the last days of its weight." Answer: the weight is not transferable and not purchasable, which removes capture by capital; the last-days attack is bounded by the 30-day re-earn, and the sortition is proportional to current weight, so a depreciating key is drawn less. The concession: the spec must stop saying "no stake" and say "no coin stake; the only thing at stake is 30 days of public work".
**The Kaspa developer's attack.** "Any slashing condition needs an objective, deterministic fault; on a DAG 'late' needs a clock, and your clock is DAA score along the carrier's chain, which is deterministic. Fine. But you now have a second use for the weight table that the finality module computes, and the two uses must read the same table at the same block or two honest nodes strip differently. Your proof-record rule needed P11 for this; write the same sentence now." Accepted.
### 4.2 Finality carried forward inside the execution proof
The Kaspa premise: finality on a DAG is a fork-choice property computed by every node from the DAG it holds; it cannot be a proof. The Monero premise: a light client trusts whatever gave it the checkpoint. Igneum's segment proof already recurses from genesis; carrying the weight table in it (3.3) makes "certified under rule v2" a public output of the same proof that attests the state root, with the update costing one mergeset per segment, not a 30-day window per proof.
**The Kaspa developer's attack.** "The proof attests a chain; finality is about the DAG. Your W2 counts blue blocks in the chain block's past, and 'blue' is GHOSTDAG's judgement, which the proof does not recompute (it would have to run GHOSTDAG over k = 18 or 124 anticone sets inside a zkVM). So the proof takes blueness as a witness from the node, and a node that lies about which blocks are blue gives the proof a wrong table. You have proven the arithmetic and trusted the colouring." This is the sharp one. Answer: the colouring is committed by the header (the mergeset and blue set are determined by the parents, which the header commits to), so the witness is checkable against headers the proof also carries; but checking it means running GHOSTDAG's blue-set rule for each merged block inside the guest, which is bounded (anticone size at most k) and unmeasured. The gate for 3.3 must add: cycle count of the GHOSTDAG colouring check per mergeset inside the guest, and if it is too heavy, the colouring stays a witness and the light client's trust row says "blue set from nodes" until it is not.
**The Monero developer's attack.** "You have made finality depend on your proof system's soundness in the light client. Say so on the card." Already in spec 10.1 for the proof system row; the row must now name finality too.
### 4.3 The hourly program as an 8,760-question hardware census
**The idea.** Every hour the chain hands every card a new random program and every card races 17 compiled variants of it and reports which won and how fast (lever 1, measured; lever 2, shipped). A chip built for the lottery cannot look like a GPU on 8,760 different programs a year: its best variant, its timing distribution across programs, its sensitivity to instruction mix are a fingerprint. Make the fingerprint part of the share protocol: a pool records, per member and per epoch, the variant that won and the share-rate ratio between consecutive programs; the chain's observer publishes the distribution per card model from the fleet library; a key whose ratio pattern sits outside every known card's envelope for N epochs is flagged publicly (the share-pattern detector of Counter ASIC item 4, which found Monero's chips by nonce patterns, now with a per-program timing axis a chip must fake 24 times a day).
**Why it is new.** RandomX programs are per hash and no pool sees their timing; Monero's chip detection used nonce distributions (MoneroCrusher, approximate); ProgPoW audits priced the chip but had no running census. Igneum's hourly swap with per-card racing produces the census as a by-product. New.
**The Monero developer's attack.** "Timing is self-reported. A chip reports whatever a 4090 would report; it has the 4090's published envelope from your own dataset (3.16). And MoneroCrusher found us the chips not by timing but by nonce patterns, which a chip emulates trivially once it knows you look. Detection that depends on the attacker's cooperation is theatre." Answer: the share rate per epoch is not self-reported; it is the pool's count of verified shares, and a chip that throttles itself to a 4090's per-program envelope on every program forfeits its edge on the programs where it is strong, which is a cost measured in hash. The detector cannot prove a chip; it can price the chip's camouflage. That is the honest claim.
**The Kaspa developer's attack.** "We welcomed chips, so nothing here is for us. But as engineering: your per-epoch ratio depends on the pool's vardiff and on network luck; the envelope for one card model will be wide, and a 2x chip sits inside it. Your detector finds a 10x chip and misses the 2x one your model says is the threat." Fair: the detector's resolution is the gate (one epoch's share-rate variance per member at one share per 10 s is about 5 percent over an hour; a 2x step is 40 standard deviations, a 1.2x step 4; so it resolves 1.2x in a day and 1.05x in a month, approximate).
**Hours.** 16: the per-epoch ratio in the pool protocol's `stats` (spec 9.5) and the observer's envelope per card model (12), the public page (4).
**The gate.** The observer flags a deliberately throttled fleet box (a 5090 capped to a 4070's rate) within 24 epochs, and flags no honest card over a week.
**Per tier.** Every miner's card model gets an envelope; a home miner on an unusual card (Apple, Intel) must be in the library or will be flagged; pools carry one more statistic; nothing in consensus.
**Verdict: watch,** then do once the pool protocol exists: it is the cheapest instrument the chain has for the question the chip model cannot answer from a spreadsheet.
---
## 5. The incremental list
Smaller than the sections above; each with hours and a gate.
| # | Item | Hours | Gate | Why now |
|---|---|---|---|---|
| I1 | Expose per-key 30-day weight and blue-block count in `IgneumInfo` so hashrate forwards and hardware-finance contracts settle from chain state (`developer-adoption.md` 2c) | 6 | A forward contract settles on the devnet against `getFinalityWeights` with no oracle | The data is already maintained for finality |
| I2 | Write the N schedule (latency-shadow program length) into the era draw at genesis, a doubling per era until the verifier gate binds (section 2.3) | 8 (spec text and the draw) | Verifier under 10 ms on a 2019-class core at the year-6 N | HBM4 arrives in 2027 to 2028 and the chain must answer it without a release |
| I3 | Define the block-proof target as a function of the fleet's measured median shard time, published per era, not as "under 10 s" (section 2.6) | 4 | The site reads it from the bench table | Honesty about the 12 GB tier |
| I4 | A second zkVM implementation of the `ProofSystem` trait (RISC Zero or OpenVM) running on one fleet box as a shadow verifier, so a soundness bug in one system is detected by disagreement before it reaches a light client | 24 | 1,000 segments agree across both; one injected bad proof disagrees | Ledger P7, D6: the veto protects full nodes, nothing protects light clients today |
| I5 | The Ember updater installs nothing while finality is paused (3.7's Kaspa attack) | 2 | A paused devnet, a published release, no install | Free |
| I6 | A finality-pause page on the site that shows the connected weight fraction live, so the "node reports the pause" sentence has a public face | 4 | Shows tonight's 18:42Z pause from the observer's data | Tonight's incident |
| I7 | Equivocation-evidence bounty paid in sortition slots: the key that first carries valid evidence inherits the stripped key's shard assignments for 30 days (no coins move; weight is reassigned, not created) | 12 | Two signers under one key on the fast-time harness; the evidence carrier wins the stripped key's draws | Makes watching for equivocation pay without a treasury |
| I8 | Mandatory proofs activation height set from a measured coverage share (spec 7.8 item 10) | 4 | Coverage above 99 percent for 7 days on the devnet | The rule is written and off |
| I9 | The exclusive window at 25 s and the claim timeout at 120 s on the phase 4 devnet (decided by Josh, P9) with the economy simulator re-run at the measured shard times from `prover-tiers-real-cards.md` instead of the 20-s target | 6 | The 3060 class's shard share within 5 points of its weight share | The inputs changed today |
| I10 | `eth_getProof`, `debug_traceTransaction`, `eth_subscribe` (D5 step 2) before any outside team | 24 | Foundry's debugger and the Blockscout fork run against a devnet node | The light client and every tool depend on `eth_getProof` |
| I11 | Register chain ids 4461 to 4463 on ethereum-lists/chains before the public testnet (spec 7.1) | 1 | The PR merged | Wallets |
| I12 | Publish the 2028 tier table (section 2.6) on the miner page with its three rates, so no card owner buys on a promise | 2 | Live | Josh's consequences rule |
| I13 | A spec sentence in 03 and 05: "no coin stake; the only thing at stake is 30 days of public work" (4.1) | 1 | Text | Before 3.2 is prototyped |
| I14 | The litepaper's income table gains the proving-market arithmetic of 3.11 in one line | 1 | Text | Ledger P6 asked for honesty; the number makes it concrete |
| I15 | A ledger entry beside E4 recording 3.6 as considered and rejected on E4's ground | 1 | Text | So the question is not re-asked |
---
## 6. Open questions and what I could not run
- **The BLS verification cycle count inside the SP1 guest** (3.3's gate a) needs a 24 GB card; PC 2 and the fleet were on the class v4 rehearsal and the Devnet 2 block-rate runs tonight. Without it, the 3 percent proving-capacity cost is an estimate.
- **The GHOSTDAG colouring check inside the guest** (4.2's Kaspa attack) is unmeasured and may be the real cost of 3.3; it is added to that gate.
- **The wrapper** (3.4) does not exist in the repository (ledger P3, R4); the 16 hours include building it on a fleet card.
- **HBM4 energy per random read** (section 2.3) is an unsourced estimate (1.0 nJ); JEDEC timing is behind the paywall; the 2x channel count is cited, the tFAW-per-channel assumption is mine.
- **The rental-tax rule's determinism** (3.1) depends on reading W30 and H_now at a checkpoint in the block's past; the lag's effect on the renter's first 30 s is unmodelled.
- **Vast.ai's actual take** is unpublished; the 15 percent is secondary.
- **No Coinbase paper on useful work was found**; if one exists its title is needed to cite it.
- **`block-rate-devnet2.md`** was a template at writing time; the 10 BPS question matters for 3.3 (segments re-cut at reorgs) and 3.5 (votes per block), and should be re-read when RUN_A lands.
- **Lane 8** holds the shadow-useful puzzle (3.10's handover) and anything about new puzzle shapes; nothing here designs a puzzle.
---
## 7. Summary for the coordinator
Lane 7 read the spec, the litepaper, the ledger sections asked, the design files, the chip model and the fleet's eleven-card table, searched prior art for sixteen ideas, and wrote one arithmetic model (`sim/horizon/frontier/frontier_model.py`) behind every number. Three findings:
1. **HBM4 raises the stored-dataset chip's per-joule edge from about 7x to about 11x bare and from about 2.3x to about 2.7x under the class v4 latency shadow at N = 100,000 (model 1.4, every chip figure arithmetic), because JEDEC doubled channels per stack (16 to 32); N = 200,000 brings it to 1.7x and N = 330,000 to 1.3x at k = 1. The N schedule belongs in the era draw at genesis (I2), with 10x of verifier headroom.**
2. **Vote weight is a slashable, non-purchasable bond (3.2): a key with 0.1 percent of hash has 16,427 IGN of 30-day pool income and its vote at risk against a designed coin bond of 0.0015 IGN per job (model section 3). The design's "no stake" must become "no coin stake".**
3. **The consensus proof can be incremental (3.3): carry the W2 table inside the recursive segment proof and update it by one mergeset per segment, with one BLS verify per 30 s (about one shard's budget, approximate, unmeasured). It is the only road to a browser that trusts no node for the voter set, and its real cost is the GHOSTDAG colouring check inside the guest (4.2), which is the first measurement to run.**
Two honest nevers with arithmetic: proving others' chains cannot be the main income by 2030 (all of Ethereum L1's proving is USD 36 a day at the Sep 2026 tracker cost against USD 13,700 a day of year-1 emission at USD 0.005; 3.11), and the lottery hash cannot be partly a proof without re-opening Aleo and breaking the DAG's memoryless election (3.10). One rule for main: the litepaper's "proving: a second income" line should carry the 3.11 arithmetic (I14), and the spec should carry the "no coin stake" sentence (I13) before any work-stake prototype starts.

View file

@ -0,0 +1,239 @@
# Horizon lane 8: a new proof of work (three candidate schemes, reviews, prototypes, verdicts)
6 October 2026, evening UK, lane `new-proof-of-work`, worktree `/Users/joshm/Projects/igneum-wt-horizon` (branch `horizon` from master, at 3f4f719). Output of this lane: this file and `proto-newpow/<scheme>/`. Nothing here touches the shipped hash, `igneum-pow`, the node, the manifest or the live devnet; every prototype is a benchmark beside the worker, never inside it.
Josh's mandate, verbatim: "if we create a new way of hashing or a new way of proof of work to revolutionise the space then that's absolutely fine, I want you to deploy everything to create something that has not ever been done before."
## 0. Progress (kept current for the coordinator)
| Time (UTC) | State |
|---|---|
| 19:05 | Lane started. Read: preamble, CLAUDE.md, the two personas, spec 01 (whole), 04, 07, chip-model-v3 (whole), asic-resistance-history (sections 0 to 3 and 4.3, 5), latency-shadow-2026-10-06 (whole), counter-asic-3-status sections 1 to 5, counter-asic-3-node section 6 (the P2 signalling rule), int8-matrix-family sections 1 to 3, scratch-soundness verdict, proving-methods (whole), fud-ledger M1, M7, M16, M22, M28, P2, F13, proto-cuda host.cu and the mx8-genesis pack (kernel.cu, memhard.h, program.h, vectors.h), proto-cuda/emu, family-probe.cu, the fleet's prover-tiers-real-cards.md, bench-log line 2582 (rental cost) |
| 19:25 | Fleet agent asked for two boxes; answered at 19:29: two quiet RTX 4090s (RunPod, nvcc 12.8 at /usr/local/cuda/bin, directory /root/horizon-newpow, until 22:30Z). No quiet Ampere card exists tonight; a loaded 3090 is offered. Main's note: cost rows use bench-log 2582 (USD 0.0117 per MH/s-hour) |
| 19:35 | File skeleton written. Two prototype sub-agents launched (budget two at once): `mma-shadow` on box 1 (47.47.180.77), `state-dataset` on box 2 (213.173.98.36) plus CPU rows on igneum-build-1. Designs being written in this file meanwhile |
| 19:41 | Section 3 complete: the three designs, the one-table comparison, the migration path. Scheme A's verdict is already visible in its own numbers (A1 dead on 2.9 MB of openings per block, A2 dead on sampleability and a 32 to 40 ms proof verify; A0 is scheme C with the trace as state). Prototypes running: `mma-shadow` (box 1) and `state-dataset` (box 2 and igneum-build-1). mm8 two-output correction sent to the prototype |
| (next) | Section 4 reviews (cryptographer, consensus engineer, two per scheme); the pick; section 5 measured rows as they land; section 6 verdicts; section 7 ranked next steps |
## 1. What was read and the facts this lane stands on
Every figure below is from the named file; "approximate" marks a figure from memory.
| Fact | Value | Source |
|---|---|---|
| The shipped hash | 64 instructions x 8 iterations, 16 loads per program (128 dependent 4-byte reads per hash), 8 registers, 32-lane unit with xor shuffles, class v3 = mixer x8 item derivation over a 256 MiB ChaCha12 cache, 1 GiB dataset in the packs (2 GiB designed), era draws, VDF seeds | `docs/spec/01-lottery-hash.md` 1.4 to 1.13 |
| Verifier today | 2.06 ms per unit on one M5 Max core (class v3, 4,096 item derivations), about 5.2 ms on a 2019-class core by the 2.5x rule; the 10 ms gate | `docs/plans/counter-asic-3-status.md` section 3, `latency-shadow-2026-10-06.md` section 4 |
| RTX 5090 at the hash | 136.1 MH/s (readwidth), 132.2 (shadow control), 290 W in the app, 350 W in the bench; 2.34 to 2.65 microjoules per hash; 17.5 G dependent reads per second, 82 percent of the GDDR7 activate ceiling; 45.2 T int op/s; marginal ALU energy 10 to 13 pJ per counted op | `chip-model-v3.md` 5.1, `latency-shadow-2026-10-06.md` 5 |
| The chip that matters | the f = 1 stored-dataset memory-controller chip: 5.1x per joule on GDDR7, 7.5x to 9.2x on HBM3 in the model; 2.1x to 4.8x by the Ethash precedent; the recompute chip (f = 0) 0.31x per chip, 1.86x per joule | `chip-model-v3.md` 5.4 to 5.6 |
| The one lever against it | program work in the latency shadow: at N = 100,000 ops per hash the chip's edge over the 5090 falls from 5.6x to 2.1x at k = 1 (chip core energy per op equal to the GPU's 11 pJ), to 3.2x at k = 0.5, 4.1x at k = 0.3; class v4 candidate `mx8+sh256x27` | `latency-shadow-2026-10-06.md` 6 and 10 |
| Step costs per family on the 5090 (ratio to the add-xor-rotate chain, 7,941 G lane-steps/s) | rotr 1.32, shflx 1.49, shfla 1.53, dot4 1.16, mm8 (`mma.m8n8k16.u8`, bit-exact) 2.43 | `counter-asic-3-status.md` section 3, item 6 |
| mm8 on other vendors | AMD RDNA 4: WMMA iu8 builtin reaches gfx12, 1.68 to 1.83 per step, fragment layout UNVERIFIED (exactness not attempted); Apple: no integer simdgroup matrix in MSL, Metal 4 `matmul2d` uchar x uchar into int exists but not from the Swift toolchain used here, per-lane dot4 emulation 1.6x unsigned, 4.7x signed | `counter-asic-3-status.md` item 6 AMD column, `int8-matrix-family.md` 1 and 4 |
| Proving today | SP1 6.8.1 Hypercube; the v1 shard (4.7 M cycles) proves in 4.8 to 18 s on 12 to 32 GB cards with the patched server, 7.4 to 8.0 GB alone; compressed proof 1.27 MB, verified in 32 to 40 ms; the aggregator 2.2 to 9.7 s per block | `prover-tiers-real-cards.md`, `proving-methods.md` 1 and 2.1 |
| Proof payment | 80/20 lottery/proving split of the subsidy; shards by weighted sortition (8 assignees, 10 s window), aggregator share 1,000 bps, unproven deadline 600 DAA s; the native-execution veto: a record whose statement differs from the node's own execution pays nothing | `docs/spec/07-execution.md` 7.2, 7.7, 7.8 |
| Class activation | P2: a class flips when 95 percent of blue blocks over a one-day window carry the object byte (header version high byte), with a fixed-height floor; one-sweep binary rollout; Devnet 2 gate first | `docs/plans/counter-asic-3-node.md` section 6, CLAUDE.md 6 Oct rules |
| Rented hash | USD 0.0117 per MH/s-hour (1,748 MH/s for USD 20.44 per hour on RunPod community pods, 18:45Z); the live devnet 1.16 GH/s | `docs/bench-log.md` line 2582 |
## 2. Method
Designs first (section 3), each reviewed in two personas (section 4), two picked, two prototypes measured on real cards (section 5), verdicts (section 6). Prototype shape: the mx8-genesis pack's own kernel text (`proto-cuda/packs-ca2-mixer/mx8-genesis/kernel.cu`, `memhard.h`) modified by the smallest change each scheme needs, compiled with nvcc 12.8 on the two RunPod 4090 boxes, timed by CUDA events over 2^24-nonce batches with `nvidia-smi` at 1 Hz, and checked bit for bit against a C CPU reference that interprets a 32-lane unit register-major with lazy item derivation (the verifier's shape) on 1,024 random lanes. CPU rows on igneum-build-1 (EPYC 9454P, Zen 4, one core at up to 3.8 GHz, which is NOT a 2019 core; the 2.5x rule of the project stands in for that core, labelled). Chip rows are the chip-model-v3 method (section 5 of that file) applied to each scheme, approximate where that file is approximate.
## 3. The three candidate schemes
Shared notation: `K_d` the day key (spec 1.8.1), `S_e` the epoch seed words (spec 1.3), the unit = 32 aligned nonces (spec 1.9), `item(t)` the 16-word class v3 derivation of spec 1.8.5, `target64` as spec 1.10. Difficulty in every scheme is the unchanged 64-bit target comparison on the unchanged DAA (spec 2.3): none of the three changes what a block's work unit is worth, only what the unit of work consists of, so the difficulty controller sees the same statistics. Each scheme is a program class in the sense of spec 1.4.5 (a `generator` number and a class byte), so activation runs through P2 (section 3.5).
### 3.1 Scheme A: mining is proving ("proof of committed trace")
**The claim to test.** The lottery's work is a bounded piece of the chain's own proving, so the 80/20 split collapses into one payment and the hash rate is the proving capacity.
**The puzzle, in its most favourable form.** Every full node already runs the segment natively (spec 7, the native-execution veto). Add one step to that: every node also runs the zkVM executor (SP1's RISC-V executor, CPU, no proving) on the segment's shard inputs and keeps the shard TRACE: the cells of every table the shard touched, about 280 M cells for the adopted v1 shard, 1.1 GB (`proving-methods.md` 1.4). The lottery of epoch `e` then uses as its dataset the trace of the last segment whose last chain block has DAA score at most `3,600 e - 1,200` (the same 20-minute lead as the epoch seed, spec 4.3), serialised row-major and padded by zero to the dataset size, and the item derivation becomes `item_A(t) = class v3 derivation with s[i] ^= row(t)[i]` for the 16 words of trace row `t` (64 bytes of trace per item). Everything else is the shipped hash: 128 dependent 4-byte reads, the 32-lane unit, the fold, `target64`. Header commitment: nothing new; the trace is a function of the segment, the segment of the chain, the chain of the header's past, exactly as the epoch seed is (spec 4.3 item 4); the class byte 5 of P2 marks the object. Difficulty: unchanged. Verifier: the node derives up to 4,096 items per unit from the 256 MiB cache (as today) plus 4,096 reads of the trace rows it holds in RAM (1.1 GB): the measured cost of exactly this read pattern is scheme C's row in section 5 (the two schemes share the verifier shape). Under 10 ms on a 2019 core if scheme C's row is.
**Two stronger forms, and why each dies on a number.**
| Form | What the miner must hold or do | Verifier | Why it dies |
|---|---|---|---|
| A1, "proof of committed codeword": the dataset is the Reed-Solomon codeword of the trace (BaseFold's stacked encoding at blowup 4, 4.5 GB, `proving-methods.md` 1.1 and 1.4), whose Merkle root the shard record already publishes; the miner must have done the COMMIT stage of the proof (encode and hash) to mine | the LDE and the Merkle tree: the first stage of every STARK or BaseFold prover, the part a GPU spends a large share of its proving time on (approximate: 30 to 50 percent of the stage time, unmeasured here) | cannot derive a codeword word locally: one evaluation of a stacked column at one point is O(2^21) field operations (the stacking height, `proving-methods.md` 1.1), so 4,096 words per unit is about 8 G operations, about 1 s on a core, 100x over the gate. The block must therefore carry Merkle openings: 4,096 per unit (every lane's loads feed every other lane through `shfl`) x 22 levels x 32 bytes = 2.9 MB per block; with a 1-lane unit (no `shfl`) 128 x 22 x 32 = 90 KB per block, 7.8 GB per day of PoW witness at one block a second against about 200 B per header today (Kaspa header, approximate). Headers-first validation, the pruning proof and the light client (spec 10) all break on the bytes |
| A2, "mine a proving step": each nonce selects a random piece of the pending proof work (a FRI fold of one column, a Poseidon2 Merkle layer, a zerocheck round) and the hash is of that piece's output | that piece | the verifier either recomputes the piece (then the verifier did the useful work, and the piece bought nothing: the definition of useless) or verifies a proof of it (an SP1 compressed proof verifies in 32 to 40 ms, `proving-methods.md` 2.1, four times the whole gate, and a per-piece proof does not exist). And the pieces run out: a segment's proof is about 5 s of one 5090 (`prover-tiers-real-cards.md`: 4.8 to 18 s per shard on 12 to 32 GB cards, 2.2 to 9.7 s of aggregation) against 8 G hashes per 8-s segment at 1 GH/s; the useful fraction of the lottery's work is bounded by (proving work per segment) / (network hashes per segment), about 8 percent at 1 GH/s and 0.08 percent at 100 GH/s (arithmetic on the cited figures), because proving work is set by gas and lottery work by the security budget. They are the same quantity only by coincidence at one network size |
**The sampleability objection, stated and answered.** Ball, Rosen, Sabin and Vasudevan, "Proofs of Useful Work" (ePrint 2017/203) build PoUW for problems with a random self-reduction (orthogonal vectors, 3SUM, all-pairs shortest paths): a useful instance is embedded into a random challenge so that the challenge is hard on average, solving it solves the instance, and verification is fast; the price is a polynomial blow-up and a problem class with that structure. Ofelimos (Fitzi, Kiayias, Panagiotakos, Russell, CRYPTO 2022) makes the useful work a doubly-parallel local search whose QUALITY improves with more work, so more hash rate yields better solutions. Primecoin (2013) mined Cunningham chains (useless, but sampleable: the nonce picks the chain's origin). Gridcoin pays BOINC credit through a trusted whitelist, not a puzzle (approximate, from memory for the last two). Proof generation has none of the three properties these constructions need: (1) a segment has one proof, not a distribution of instances, so there is nothing for the nonce to sample; (2) partial progress has no verifiable value under the gate without a proof, and a proof verify costs 32 to 40 ms; (3) the quantity of useful work is fixed by demand (gas), the quantity of lottery work by the security budget (hash rate), and a puzzle whose per-solution work is fixed by demand is not a difficulty-adjustable lottery. The answer this design gives is therefore no: the only form that survives the gate and the bytes is A0 above, in which the miner must HOLD the trace, which is "proof of stored state" with the trace as the state, and the useful work gained per hash is zero (every node computes the trace natively anyway; ledger F13). Mining does not become proving; it becomes proof that the miner executes.
**Chip edge per joule (chip-model-v3 method).** A0 changes the dataset's contents, not its size or read pattern; the f = 1 stored-dataset chip stores whatever the items are: 5.1x per joule on GDDR7, 7.5x to 9.2x on HBM3, unchanged from `chip-model-v3.md` 5.4. The f = 0 recompute chip must now hold the 1.1 GB trace somewhere (it cannot derive an item without row `t`), so it needs DRAM beside its SRAM and becomes an f = 1 chip; the 0.31x row disappears, which is a small gain since that chip was never the threat. A1 and A2 are not priced: they fail before a chip is drawn.
**Verifier cost.** A0: the class v3 derivation (2.06 ms per unit on an M5 Max core, measured) plus 4,096 random 64-byte reads from a 1.1 GB array; the read row is measured in section 5 (scheme C's CPU rows, the same pattern). A1: 2.9 MB of openings or a 1-lane unit (see table). A2: a proof verify (32 to 40 ms, measured) or nothing useful.
**Bit-exactness.** A0: the trace rows are bytes; the derivation is integer; nothing vendor-specific is added. The zkVM executor's trace must itself be deterministic across platforms (SP1's executor is Rust, no floating point in the trace path, approximate: not audited here); any disagreement about a trace cell is a consensus split, so A0 adds the SP1 executor to the consensus-critical code, which ledger P7 already names as a risk class for the proofs and which here would extend to the lottery.
**Known attacks.** Grinding: a block producer influences the trace through the transactions it includes, but the trace used is 20 minutes old and keyed by `K_d` through the mixer; a producer cannot predict a useful bias in 128 dependent reads of a keyed derivation (the MTP lesson, `asic-resistance-history.md` 2.3: Dinur and Nadler controlled addresses by controlling contents; here the contents enter only through a keyed, chained derivation whose addresses are the cache-line indices of the mixer state, not the trace). Outsourcing: pools serve the dataset (1 to 2 GiB per miner per day), so "the miner holds the trace" becomes "someone in the pool holds it". Precomputation: the dataset is computable 20 minutes early, as today. Light evaluation: as the f = 0 row above. Sampleability: the paragraph above. Empty segments: a day or an epoch with no transactions has a near-empty trace (the shard statement still applies rewards, spec 7.7 item 8), so the dataset is mostly padding and the hash degrades to today's class v3: graceful, not an attack.
**Game theory of the collapsed payment (if A1 or A2 had worked).** The 20 percent pool would go: provers are miners and the proof is the by-product of the lottery. Who is paid for what: the block reward pays for the block and the proof piece in it; a prover that mines earns exactly a miner that proves. Pools: the pool does the useful work and sells shares, as today, so proving centralises exactly as mining does (Aleo's proof of succinct work centralised on the same path: the coinbase puzzle was a synthetic circuit proven by whoever had the most GPUs; approximate, from memory, the cryptographer persona's own list). Proving demand at zero: the puzzle has no useful content and must fall back to a synthetic dataset, so the chain carries two puzzle modes and a mode switch, a consensus rule and a grinding surface. External jobs: a customer's proof could only be mined if its segment entered the puzzle, so the security budget would be spent on the customer's work at the marginal cost of inclusion, a subsidy from holders to customers unless priced by a burned fee. These are the reasons CLAUDE.md records the lottery and the proving as separate on purpose; nothing found tonight overturns them.
**Migration.** A0 is a class byte (P2): class v5 by the 95 percent signal with the floor height, one-sweep binary rollout, Devnet 2 gate first (CLAUDE.md 6 Oct). Every node must run the zkVM executor on every segment before the flip (CPU cost: the v1 shard is 4.7 M cycles; SP1's executor runs at tens of MHz on a CPU, approximate, so under a second per 8-s segment); a node without it cannot validate PoW after the flip, which is what the floor height is for.
**Verdict line (section 6 has the reasoning):** NEVER as "mining is proving" (A1, A2); A0 is scheme C with the trace as the state and is folded into C.
### 3.2 Scheme B: a GPU-structure-bound puzzle ("mx8+mm8xR", the tensor-shaped integer shadow)
**The claim to test.** Fill the latency shadow (the only lever that moves the f = 1 chip, `chip-model-v3.md` 5.7, `latency-shadow-2026-10-06.md`) with work shaped like a GPU's own tensor datapath rather than its scalar ALUs, so that a chip must carry GPU-class matrix units whose energy per operation NVIDIA's own silicon already sits near the floor of, and the chip's residual edge `k` (its energy per op divided by the GPU's) cannot fall far under 1. The brief's other candidates are rejected on definition: shared-memory bank timing and register-file width are timings and capacities, not values, and a consensus rule can only check values; a per-warp scratchpad was measured and found not sound as a chip layer (`docs/analysis/scratch-soundness.md`: the live state is bounded by the read-modify-write count, 64 to 320 bytes per lane, which a chip keeps in SRAM at under 5 percent of its mirror).
**The puzzle.** Class v3 (mixer x8, 16 loads, the 64 base instructions, the era draws) plus a block of `R` `mm8` steps executed at the end of every iteration, after instruction 63 and before the next iteration samples `sel`; `8 R` steps per hash, no load in the block, the base program untouched (the same seam the class v4 shadow uses, `latency-shadow-2026-10-06.md` section 2, so the acceptance rule of 1.4.6 keeps its verdict draw for draw). Step `k` has four draws from the program stream after the base and shadow draws: `a = below(8)`, `b = below(7) + (b >= a)`, `c = below(8)`, `c2 = below(7) + (c2 >= c)`. Its semantics over the unit:
```
A: 8 x 16 uint8 lane l holds A[l >> 2][4 (l & 3) .. 4 (l & 3) + 3] = the 4 bytes of r[a] (byte 0 = lowest k)
B: 16 x 8 uint8 lane l holds B[4 (l & 3) .. +3][l >> 2] = the 4 bytes of r[b]
C = A x B C[i][j] = sum over k of A[i][k] B[k][j], exact in int32 (at most 1,040,400)
r[c] = r[c] + C[l >> 2][2 (l & 3)] (mod 2^32)
r[c2] = r[c2] + C[l >> 2][2 (l & 3) + 1] (mod 2^32)
```
This is the fragment layout of PTX `mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32` with the two `.s32` outputs `d0`, `d1` of each lane (PTX ISA, "Matrix Fragments for mma.m8n8k16", the integer layout; `docs/analysis/int8-matrix-family.md` 2.2 quotes it and `proto-cuda/family-probe.cu` carries the CPU reference that was bit-exact on the RTX 5090 on 5 October 2026, `counter-asic-3-status.md` item 6). The spec defines the matrices and the lane ownership, not the instruction: a vendor permutes its own fragment layout into this one (a shuffle) and the result is defined whatever the hardware. Why TWO destinations where the reserve entry R8 of 1.13.2 has one: with one element per lane only 32 of the 64 products of C are consumed, so a chip does 512 multiply-adds per step where the GPU does 1,024 and the GPU hands the chip a free 2x on the block; with both outputs consumed the whole tile is load-bearing. This is a correction to the reserve text as well (section 7). Unit of work: still the unit. Header, difficulty: unchanged; the class byte 5.
**Parameters and the verifier's law.** Per unit the verifier adds `8 R x 1,024` unsigned byte multiply-adds (the register-major interpreter already holds the 32 lanes' registers, so A and B are in hand). Plain scalar C: about 2 ops per multiply-add, 16,400 ops per step per unit; at the 18 G op/s the x8 verifier shows on the M5 Max core (`chip-model-v3.md` 5.7, measured) that is 0.9 microseconds per step per unit: `R = 128` adds about 0.9 ms, `R = 512` about 3.7 ms. With a byte-dot instruction (AVX-VNNI `vpdpbusd`, 64 multiply-adds per instruction; NEON `udot`, 16 per instruction) the same work is 4x to 16x fewer instructions (approximate). The headroom is the class v3 margin: 7.9 ms steady on the M5 Max core, 4.6 ms on a 2019-class core by the 2.5x rule (`latency-shadow-2026-10-06.md` section 4), so scalar `R` is bounded near 500 on the 2019 core and the measured rows of section 5 set it. On the GPU side the 5090's dependent-chain probe runs `mm8` at 2.43x the add-xor-rotate step, 7,941 / 2.43 G lane-steps per second = 102 G `mm8` per second per card, 104 T multiply-adds per second (the chain probe, not the tensor peak); at 4.25 M units per second (136 MH/s) the card could hide about 24,000 `mm8` per hash before the probe rate binds, `R` about 3,000, far above what the verifier allows. So the verifier binds first, at about `R = 500` scalar on a 2019 core (approximate until section 5) and perhaps 4x higher with VNNI: the honest card never leaves the latency bound on the candidate `R`.
**Chip edge per joule (chip-model-v3 method).** The f = 1 chip (memory, controller, static: 0.466 microjoules per hash on GDDR7, 0.321 on one HBM3 stack, `chip-model-v3.md` 5.4) plus a tensor array: energy per hash = memory + `8 R x 1,024 x e_mac x k_mma`, where `e_mac` is the honest card's marginal energy per multiply-add at the block (measured in section 5 as watts delta over MACs per second) and `k_mma` the chip's ratio to it. The difference from the ALU shadow (`k` down to about 0.3 for a fixed-datapath array at N5, `latency-shadow-2026-10-06.md` 6) is where the honest card's engine sits: NVIDIA's tensor cores are int8 multiply-add arrays at N4/N5-class density already, so a chip's array is the same circuit (approximate: int8 MAC datapath about 0.05 to 0.1 pJ at N5, the movement of fragments through the register file the larger term on both sides; from memory) and `k_mma` is near 1 with a floor near 0.5 for a chip that keeps the fragments in a local register file instead of the GPU's banked one. Rows at `k_mma` = 1, 0.5 and the measured `e_mac` are filled in section 5.3 from the 4090 measurement; the shape of the result is already clear: the block lowers the chip's edge by the same mechanism as the ALU shadow and the chip's best case is better bounded, at the price that the honest card's own watts rise by the block's energy (the tensor path is efficient, so the rise per unit of chip-forcing work should be smaller than the ALU shadow's 11 pJ per op; the measurement says).
**Bit-exactness across vendors.**
| Vendor | Path | State |
|---|---|---|
| NVIDIA sm_75 and later (Turing, Ampere, Ada, Blackwell) | one `mma.sync.m8n8k16.u8` per step, identity permutation, wrap on the `.s32` accumulate (the wrap edge vector of the R8 entry was bit-exact on the 5090, `int8-matrix-family.md` 3) | measured bit-exact on the 5090 (5 October) and on the 4090 tonight (section 5) |
| NVIDIA sm_61 to sm_72 (Pascal GTX 10 series, Volta) | no `mma` with `.u8`; emulation: 8 `__shfl_sync` gathers plus 8 `dp4a` per lane per step (`dp4a.u32.u32`, sm_61+) | unmeasured; cost about 16 steps per `mm8` against 2.43 native, approximate; within the 8x emulation bound of 1.13.2 on a per-op basis only if the shuffles are cheap; a GTX 1080 owner is the first NVIDIA tier to pay |
| AMD RDNA 3 and 4 (gfx11, gfx12) | `V_WMMA_I32_16X16X16_IU8` with the 8 x 16 and 16 x 8 tiles zero-padded to 16 x 16 and a fixed lane permutation (`ds_bpermute`) into the spec layout; the RDNA 4 builtin takes 2 ints per lane for A and B (`int8-matrix-family.md` 1) | the fragment layout is UNVERIFIED (status item 6: not in any source at hand; the CPU reference was not attempted rather than guessed). This is the gate for AMD: a PC 1 job on the 9070 XT with the spec reference. Until it passes, AMD runs the emulation path (`v_dot4_u32_u8` is native on gfx11 and gfx12, 1.06x per op measured) at about 12 to 16 steps per `mm8`, approximate |
| AMD RDNA 2 and older, CDNA | no WMMA on RDNA 2; `v_dot4_i32_i8` exists (`dot1-insts`, signed only, approximate); CDNA 3 has `V_MFMA_I32_16X16X32_I8` with its own layout | emulation; unmeasured |
| Apple (M-series, Metal) | no integer `simdgroup_matrix` in MSL; Metal 4 `mpp::tensor_ops::matmul2d` has `uchar x uchar -> int` (table 7.3, OS 26.4) through a tensor API not reachable from the Swift toolchain this project uses, and its wrap semantics are unverified (`int8-matrix-family.md` 1). The emulation: 4 shuffles plus 4 unsigned `dot4` emulations at 1.6x per op (measured 5 October), about 10 ALU steps per `mm8` | an Apple miner pays about 4x NVIDIA's per-step cost on the block (approximate); the M5 Max is latency-bound to about 130,000 counted ops per hash (measured), so `R = 128` (1,024 `mm8` per hash, about 10,000 steps emulated) fits inside its shadow and `R = 512` (about 41,000) still does by the ops count, with the hash-rate cost owed to a Metal measurement. Apple's path is the honest card's worst and the chip's argument does not depend on it |
| Intel Arc | XMX through `cl_intel_subgroup_matrix_multiply_accumulate` (approximate, unverified); dp4a-class `dot` otherwise | unmeasured |
So the fleet splits by generation: Turing-and-later NVIDIA and RDNA 3-and-later AMD run the block natively; everything older and Apple emulate. The 5 percent rule (`counter-asic-2-public.md`) is checked per card with the block live, as 1.13.2 requires for an emulating vendor.
**Known attacks.** Grinding: none new; the block has no data-dependent control flow and reads no memory. Outsourcing, precomputation: as today (the draws are public per epoch; there is nothing to precompute because the inputs are the per-nonce registers). Light evaluation (the f = 0 chip): untouched; the block never reads the dataset. MTP-class content control: not applicable (no attacker-chosen memory). Sampleability: not applicable. New: (i) the half-tile shortcut, closed by the two-destination form above; (ii) a zero or low-entropy fragment (if `r[a]` is zero in every lane the step is free): the acceptance rule's register-saturation test (1.4.6 (c)) already rejects programs with stuck registers, and the base program's loads re-randomise every register every iteration; (iii) the trailing-step contraction: the last `mm8` of the last iteration writes two registers that feed only the fold; a chip could skip nothing because the fold reads all eight, but a step whose destinations are both never read again before the fold still costs the GPU a full tile; the draw rule should forbid `c, c2` outside the fold's register set, which is every register, so there is no such step. (iv) The licensable-IP objection (the history's reason for ranking `mm8` last in the reserve, `asic-resistance-history.md` 4.3 row 6): a chip maker licenses an int8 MMA block at any node. True, and it is the point of the design: the chip must then carry a GPU-class tensor array per 32 lanes in flight at the memory's activate ceiling, 1,172 lanes on GDDR7 (`chip-model-v3.md` 5.5), 37 tiles in flight, and its edge is `k_mma`, a ratio of two copies of the same circuit; the design does not claim the chip cannot be built, it claims the chip is a GPU.
**Game theory.** None of the payment changes; this is a hash change. A pool user sees nothing. A prover that mines pays the block's watts on the same card it proves on; the tensor path is also the proving path's (SP1's Poseidon2 and NTT kernels are integer, not tensor, so there is no contention beyond the power limit, approximate). When proving demand is zero nothing changes.
**Migration.** Class v5 by P2: the object byte 5, the 95 percent one-day window, the floor height; one-sweep binary rollout because the digest flips; Devnet 2 gate first. The kernel emitter (`igneum-pow/src/emit.rs`) gains the block in its three dialects, the CUDA one with inline PTX and the `IGNEUM_MM8_REF` fallback for sm_61 to sm_72, the OpenCL one with the AMD builtin behind a feature test and the emulation otherwise, the Metal one with the emulation; the one-click workers compile the text as they do today (NVRTC accepts inline PTX). Gates before a cut: the six gates of `counter-asic-2-rollout.md` section 7 plus the AMD layout verification and the Metal emulation's hash-rate cost on the M5 Max.
### 3.3 Scheme C: proof of stored state ("sd1", the dataset is the chain)
**The claim to test.** The dataset is the recent chain state, so every hash proves the miner holds the chain, and the lottery's reads double as a verifiable random sample of state for light clients.
**The puzzle.** Day `d`'s snapshot `SS(d)` is the execution state at the state root `R_d` of the certified checkpoint `C_day(d)`, the highest-index checkpoint whose block has DAA score at most `86,400 d - 1,200` (the epoch seed's lead, spec 4.3). Its leaves: the `(key, value)` pairs of the execution state trie in key order (storage slots, account records, 64-byte chunks of code), leaf `t` serialised as `leaf(t) = Blake2b-512(R_d || t_le32 || key_t || value_t)` (the chain's own hash, spec 0.6), 64 bytes each, so every leaf carries full entropy whatever its content and no two days share a leaf; for `t` beyond the state's leaf count `leaf(t) = 0`. When the state has more leaves than the dataset has items, the dataset holds the first `2^(D-4)` leaves in the order of `Blake2b-256(K_d || key)`: a keyed sample that cannot be chosen without the whole state. The item derivation is class v3's with one line added before the round loop: `s[i] ^= leaf(t)[i]` for `i` in 0..15. The hash kernel is byte for byte the shipped one; only the daily build changes. Header commitment: none new. `R_d` is the state root of a block in the header's own past at a fixed blue score, so "which snapshot was this block mined under" is a function of the header alone once the chain is known, as the program is (spec 1.12). The class byte 5 marks the object. Difficulty: unchanged. Unit of work: unchanged.
**The verifier.** A node holds the 256 MiB cache (as today) and `SS(d)` in RAM or mmap (2 GiB at the genesis dataset size, growing on the 1.13.3 schedule), derives up to 4,096 items per unit lazily as today and reads `leaf(t)` for each: 4,096 random 64-byte reads. The measured cost of those reads and of the derivation is section 5.2. Verifier memory: +2 GiB (+4 GiB at year 4). The daily snapshot build: one pass over the state trie, hashed per leaf (one Blake2b-512 per 64 bytes: about 2^25 hashes for 2 GiB, seconds on a core, section 5.2 measures the stand-in).
**What is new, and the prior art.** Permacoin (Miller, Juels, Shi, Parno, Katz, IEEE S&P 2014) made the puzzle a proof of retrievability over a large PUBLIC FILE chosen by a dealer, with Merkle openings in each block; the file was external to the chain and the openings were the bytes. Spacemesh (proof of space-time over a plotted file of random data) and Chia (Abusalah, Alwen, Cohen, Khilko, Pietrzak, Reyzin, "Beyond Hellman's time-memory trade-offs with applications to proofs of space", ASIACRYPT 2017; Chia's plots) prove storage of USELESS data. Verthash (Vertcoin, January 2021; `asic-resistance-history.md` row 8) is the nearest: a 1.2 GB file generated from the chain's own block headers, random reads, no chip after 69 months on a small prize; its data is headers (low entropy per byte, static once written) and it proves nothing about state. Ethash's DAG is from the epoch seed (random). What Igneum C adds: (1) the dataset is the EXECUTION STATE, keyed per day by `K_d` through the memory-hard derivation, so it is never easier than today's dataset (the leaf is one more 64-byte input to a 9,360-op chain) and it cannot be built from the day key alone: whoever builds it holds the state; (2) every block's 128 loads per lane name 128 keyed items whose leaves are a uniformly random sample of state (the addresses are the mixer-state cache-line indices through 128 dependent reads, unbiasable by the producer at a cost below a block), so a light client that asks any full node for the block's `(key, value)` leaves with their openings against `R_d` gets a free daily spot check of state availability; (3) the beacon: the lottery output is already public randomness, biasable by withholding at the cost of a block, as every PoW; C adds nothing there and the design says so. Honest limit: a pool can ship the 2 GiB dataset or the snapshot to its miners once a day (2 GiB per miner per day, 23 MB/s for a thousand miners), so "every miner holds the chain" is really "every mining OPERATION holds the state", which is still a change: today a pool miner needs nothing but the day key.
**Chip edge per joule.** Unchanged against the f = 1 chip (it stores items whatever they are): 5.1x on GDDR7, 7.5x to 9.2x on HBM3 in the model, 2.1x to 4.8x by the Ethash precedent. The f = 0 recompute chip must hold the leaves (2 GiB) in DRAM to derive anything, so it becomes an f = 1 chip and the 0.31x row disappears. C is therefore not an anti-chip scheme and does not claim to be; it composes with B (the shadow is in the kernel, the state is in the build).
**Bit-exactness.** The leaf is the output of the chain's own hash over bytes every node agrees on by consensus; the derivation is integer; nothing vendor-specific is added. The one new consensus-critical function is the canonical serialisation of state (key order, chunking of code), which every node must compute identically: a bug there splits the chain on a day boundary, the class of M20 and the DAA 198,000 incident. The Devnet 2 gate exists for exactly this.
**Known attacks.** State grinding: a producer can write state (pay gas) to influence leaves; the leaf is hashed with `R_d`, which depends on every leaf, and enters a keyed chained derivation whose read addresses are mixer state, so no bias on 128 dependent reads is reachable at a cost below a block (the MTP lesson, `asic-resistance-history.md` 2.3, is the reason for the hash and the key, not the plain bytes). Compressible state: an attacker fills state with zeros hoping a chip stores it compressed; the hashed leaves are full-entropy, and the padding region (`leaf = 0`) is today's dataset, which the chip already stores at 64 B per item. Precomputation: the snapshot is fixed when `C_day(d)` is certified and `K_d` is known, 20 minutes before the day (the same lead as the epoch seed), and the build is 13 to 77 ms on the GPUs measured (`counter-asic-3-status.md`) plus the leaf pass; a reorg across `C_day(d)` is a merge-depth-scale event, accepted as for the epoch seed (spec 4.3 item 4, O-4.3). Outsourcing: the pool ships the dataset (above). Light evaluation: the f = 0 row above. Long-range: an attacker building an alternative history must build its alternative state snapshots to mine on it, which it does anyway; no change. Finality pause: `C_day(d)` must be certified; if finality is paused for more than the lead the day's snapshot is not derivable and mining would stop, the coupling spec 4.3 argues against for the epoch seed. Rule, as there: take the selected-chain block at that blue score certified or not, deep enough that a reorg across it is a merge-depth event. Empty state at launch: every leaf is `Blake2b(R_d || t || empty)`, full entropy, the dataset as good as today's: graceful.
**Game theory.** No payment changes. A miner must run or rent a node (or trust a pool's dataset); the solo-mining floor rises by a full node's state (today's devnet: megabytes; a used chain: gigabytes). A pool user sees a 2 GiB daily download or nothing (the pool serves the dataset). A prover that mines already holds state. A holder gains a daily sample of state availability per block, for free. A rollup customer gains nothing directly.
**Migration.** Class v5 by P2 (object byte 5, 95 percent over a day, floor height, one-sweep rollout, Devnet 2 first). Every node needs the snapshot builder before the flip (a node without it cannot validate PoW after the flip: the floor height's job). Workers need nothing new: the kernel text is unchanged, the dataset arrives from the node's `prepare` line as today, built on the GPU from the cache plus a leaf array the node hands over (2 GiB per day over the local socket) or built on the node's CPU and uploaded. The one-click miner's "nothing to install but the driver" line holds; "nothing to download but the day key" does not.
### 3.4 The three schemes in one table
| Scheme | What it is | Chip edge per joule vs the 5090 (f = 1 chip, chip-model-v3 method) | Verifier ms per unit (model, then measured in section 5) | Vendor bit-exactness | Ships as class v5? |
|---|---|---|---|---|---|
| A, mining is proving | A1 committed codeword, A2 proving steps: dead on bytes and on sampleability; A0 trace-as-dataset survives and is C with the trace as state | A0 unchanged (5.1x GDDR7); A1, A2 not priced | A0: 2.06 + the leaf-read row; A1: 2.9 MB of openings; A2: 32 to 40 ms | A0 adds the zkVM executor to consensus | NEVER as mining = proving; A0 folds into C |
| B, tensor-shaped shadow | class v3 plus `8 R` int8 8x8x16 tile steps per hash in the PTX fragment layout, two outputs per lane | memory + `8 R x 1,024 x e_mac x k_mma`; `k_mma` near 1 with a floor near 0.5 (approximate); rows from the measured `e_mac` in 5.3 | 2.06 + about 0.9 microseconds per step per unit scalar (R = 128: +0.9 ms; R = 512: +3.7 ms); VNNI 4x to 16x less | native on sm_75+ and RDNA 3+ (AMD layout unverified); emulated on Pascal, RDNA 2, Apple | prototype further; a class v5 candidate after the AMD gate |
| C, stored state | the daily dataset derives from the execution state snapshot; kernel unchanged; a state sample per block | unchanged (5.1x GDDR7, 7.5x to 9.2x HBM3); the f = 0 chip disappears | 2.06 + 4,096 leaf reads (section 5.2) | nothing vendor-specific; the serialisation is the consensus risk | prototype further; a class v5 candidate on its own or beside B |
### 3.5 Migration through the class system (common to B and C)
The P2 rule as designed (`docs/plans/counter-asic-3-node.md` section 6): the header version's high byte carries the producer's object version; epoch `e` is the new class when the window of one day ending at its seed block has at least 9,500 bps of blue blocks at or above the byte, or when the floor height `N` is reached, or when epoch `e - 1` already was; the rule answers v5 only where it would answer v4. For B and C the object byte is 5 and the floor is set at the publish as DAA + 14,400 rounded up to the epoch boundary. The order: Devnet 2 crossing with `tools/fleet/devnet2-gate.sh` (zero rejected blocks across the flip, no reorg over depth 3, exec roots agreeing, a segment record paid, every node on the new version), then the live devnet in one binary sweep (the digest flips), then the flip by signal. A node that synced from a pruning proof takes the floor rule for epochs whose window reaches below its pruning point (the same class as the era witness, status item). For C the floor also bounds how long a non-upgraded node can keep validating PoW: none after the flip, so the sweep must be complete before the floor, which is the 10,800-DAA check already in the rule.
## 4. Reviews: two personas, two short reviews per scheme
Written by this lane in the persona files' voices (`.claude/agents/cryptographer.md`, `.claude/agents/consensus-engineer.md`): every design claim with its attack, every claim about another chain with its file, and the smallest change the code allows. Each review names the break if there is one.
### 4.1 Scheme A, mining is proving
**Cryptographer.** The break is structural and has a name: a lottery needs a distribution of instances and proving has one instance per segment. A1 moves the work into the commit phase and pays for it in witness bytes (2.9 MB per block, or 90 KB with the unit cut to one lane, which also removes `shfl` and with it the only thing that makes the 32-lane unit a unit; `docs/spec/01-lottery-hash.md` 1.9). A2 either recomputes (useless by definition) or verifies a proof (32 to 40 ms measured, `proving-methods.md` 2.1; the gate is 10 ms). The useful fraction bound (proving work per segment over network hashes per segment) is the argument Ball, Rosen, Sabin and Vasudevan make in the negative direction: without a random self-reduction the embedded instance is a constant, and a constant is amortised to zero by the first miner who computes it. A0 is sound as far as it goes and is scheme C. Second attack on A0 that C does not have: the SP1 executor enters the consensus path for the lottery, so an executor bug that produces a different trace on one platform (an undefined-behaviour corner in a precompile patch, a `sha3` or `k256` version skew) splits the chain at the PoW, not at the proof; ledger P7 priced that class for proofs where the native veto bounds the damage to one payout; here nothing bounds it. Verdict: never for A1 and A2; A0 only as C with the trace, and then the state is the better choice of data because every node already agrees on it without a second executor.
**Consensus engineer.** The code says the same thing from the other side. Block validation in the fork validates the header's PoW before the body is fetched and before execution (`check_pow` on the header path, rusty-kaspa's `header_processor`; the fork's `igneum/exec` runs after the block is accepted into the DAG). A puzzle whose verification needs the segment's trace makes header validation wait on the executor of a segment 20 minutes old, which is fine for a synced node and fatal for IBD: a syncing node must execute every segment of history, in order, to validate the headers of history, so headers-first sync and the pruning proof (which validates headers without bodies) are gone; Kaspa's pruning-point sync (`consensus/src/pipeline/pruning_processor`, approximate location) assumes PoW is a function of the header and a small amount of context. A0 shares this break with C and C answers it in 4.3 by making the snapshot a function of a certified checkpoint's state root plus the state itself, which a pruned node fetches as a snapshot (the exec snapshot path the node already has, CLAUDE.md 6 Oct rule "the p2p snapshot path refuses a snapshot below the node's tip"). For A1 the 90 KB per header kills the header relay and the 600-block record window arithmetic alike. Verdict: never for A1 and A2; A0 is C with a worse data source.
### 4.2 Scheme B, the tensor-shaped shadow
**Cryptographer.** No break in the puzzle's soundness: the block is a straight-line integer map with no memory and no data-dependent control flow, its inputs are the per-nonce registers, and the two-output form makes the whole tile load-bearing. Two things to name. (i) The claim that `k_mma` cannot fall far under 1 rests on the honest card's tensor path being near the floor of int8 multiply-add energy; that is an engineering judgement, approximate, and the external chip review (status item 3, `funding.md`) is where it is tested, with the ALU shadow's `k` beside it. The design's advantage over the ALU shadow is bounded, not proven: it narrows the chip's best case from about 0.3 to about 0.5 (approximate) and does not remove the edge. (ii) The fragment layout is a specification of lane ownership; every emulating vendor must reproduce it exactly, and the AMD WMMA layout is unverified (status item 6). Until a 9070 XT run with the spec reference passes, B is bit-exact on one vendor's hardware and in every emulation, which is the state class v2 was in on 3 October (ledger M8) and not a state to cut from. No grinding, outsourcing or precomputation surface is added. A statistical point: the mm8 step's outputs are sums of 16 byte products, so each added value is at most 1,040,400 and its top 12 bits are zero; added into a register it changes the low 20 bits in a structured way. That is fine inside a chain of multiplies and rotates, and the stats run of `TESTS.md` section 3 should be run on the class before any vector is frozen. Verdict: prototype further; a class v5 candidate after the AMD gate and the stats run.
**Consensus engineer.** No consensus change beyond the class byte and the emitter: the block is kernel text (`igneum-pow/src/emit.rs` gains the step in the three dialects), the verifier (`verify.rs`) gains the step in the register-major loop, the acceptance rule is untouched by construction, and the node's seam is the v5 switch beside the v4 one (`counter-asic-3-node.md` section 1 lists every file the v4 switch touched; v5 is the same list with one more number). The break to name is operational: the fleet splits by hardware generation. Pascal, Volta, RDNA 2 and every Apple card emulate, and the one-click workers compile three paths where they compile one today; the OpenCL path on AMD needs a feature test at compile time (the WMMA builtin exists on gfx11 and gfx12 and not on gfx10, `int8-matrix-family.md` 1), which the pack text must carry as a preprocessor branch, and a wrong branch is a wrong hash, which the self-test catches before the worker serves (`packfile.h`, M28's fix). The 5 percent rule must be checked per emulating card with the block live, and the Apple row is the one most likely to fail it at high `R`. Verdict: prototype further; set `R` from the measured rows with the Apple emulation measured on the M5 Max before any cut; the AMD layout is a hard gate.
### 4.3 Scheme C, proof of stored state
**Cryptographer.** The construction is never weaker than today's: the leaf is one more input to the same chained derivation, keyed by the day key through the first mixer, and the f = 1 chip's row is unchanged, which the design says. The break to name is the one Permacoin and MTP both met: who controls the data controls the addresses, unless the data enters through a key the controller does not have. Here the controller of state content (anyone paying gas) does not control `K_d` (a VDF output fixed 20 minutes before the day, spec 4) and the leaf is `Blake2b(R_d || t || key || value)`, so the only lever is choosing content before `R_d` is known, which affects every leaf through `R_d` and none in a predictable way. I find no bias at a cost below a block. The second point: the "verifiable sample of state" is real but modest. A block commits to 128 leaves per lane through 128 dependent reads; a light client checking them needs openings against `R_d` from a full node (depth about 25 at 2^25 leaves, 128 x 25 x 32 B = 102 KB per block if fetched, arithmetic), which is a spot check of availability, not a proof of state correctness; the consensus proof of spec 10 and ledger P4 is still what a light client needs for correctness. The design says this. Third: the canonical serialisation is new consensus-critical code, and the day-boundary flip is the moment it bites (every node rebuilds at once). Verdict: prototype further; a class v5 candidate; the serialisation needs the same test discipline as the DAA switch (a Devnet 2 crossing over a day boundary with a non-trivial state).
**Consensus engineer.** The break I would have named, the IBD and pruning-proof problem of A0, C answers: the snapshot is a function of `R_d`, a state root at a certified checkpoint, and the node already carries an exec snapshot path (`exec_restart_*` fields, the p2p snapshot message, CLAUDE.md 6 Oct). A syncing node validates historical PoW only by holding every day's snapshot, which is 2 GiB per day of history, which is NOT acceptable for IBD. Fix, the smallest I can see: PoW of blocks below the pruning point is not re-validated (Kaspa's pruning proof validates the proof's headers' PoW; rusty-kaspa `consensus/src/processes/pruning_proof`, approximate), so historical snapshots are needed only for the headers inside the proof, which are a bounded set per level; and for them the proof carries the day's `R_d` and the node either holds that day's state (recent days) or trusts the certificate (the finality rule's lock already makes those headers irreversible). That is a spec item for `docs/spec/10-light-client.md` and the pruning section of spec 02, and it is the same shape as the era-seed witness (status item). Second: the verifier's RAM (+2 GiB, +4 GiB at year 4) and the daily build on a node without a GPU (a seed node, a Hetzner box: the leaf pass is CPU work, measured in 5.2) must stay inside the node's budget; the devnet hands on igneum-build-1 have 128 GB, a home node has 16. Third: a day boundary is now a consensus event that depends on a certified checkpoint 20 minutes before it; under a finality pause (6 October, 18:42Z) the day's snapshot falls back to the uncertified selected-chain block, as spec 4.3 argues for the epoch seed; the rule must be written once and tested on the fast-time harness across a pause. Verdict: prototype further; a class v5 candidate; three spec items (pruning-proof witness, the pause rule, the node RAM budget) before a cut.
### 4.4 The pick
B and C are the two to prototype: both are class objects on the shipped hash, both leave the dataset's memory bound untouched, and they compose (B is kernel text, C is the daily build). A is not prototyped: A1 and A2 fail on bytes and on sampleability before any kernel, and A0 is C with a worse data source. The order of merit at this point, before measurement: C first (no vendor risk, unchanged hash rate by construction, a real new property per block, the pool caveat stated), B second (a real lever against the f = 1 chip with a bounded `k`, a vendor split and an unverified AMD layout). Section 6 revisits the order on the measured rows.
## 5. The prototypes and the measured rows
Both prototypes live under `proto-newpow/` with a README carrying the exact commands, the card, the driver and the RESULTS table; this section carries the rows and their consequences. Boxes: two RunPod RTX 4090 24 GB (driver 595.91, nvcc 12.8, `-arch=sm_89`), quiet, the card to itself; CPU rows on igneum-build-1 (EPYC 9454P, one pinned core, `nice -n 19`). Power by `nvidia-smi` at 1 Hz, the mean after the first 10 s of each timed run. Bit-exactness: the PTX path against the reference path on 2^24 lanes (fingerprint), and the GPU against the C CPU reference on 1,024 random lanes.
### 5.1 `mma-shadow` (scheme B), box 1
ROWS_B
### 5.2 `state-dataset` (scheme C), box 2 and igneum-build-1
ROWS_C
### 5.3 The chip rows on the measured numbers
CHIP_ROWS
## 6. Verdicts
| Scheme | Verdict | Why, in one line |
|---|---|---|
| A, mining is proving | NEVER (A1, A2); A0 folds into C | one proof per segment is not a distribution of puzzles; the bytes (2.9 MB of openings per block) or the verify (32 to 40 ms) kill every form that is not "hold the trace", and holding the trace is C with a worse data source |
| B, tensor-shaped shadow | VERDICT_B |
| C, stored state | VERDICT_C |
**A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.
VERDICT_BC_PROSE
## 7. Ranked next steps
Hours are agent hours (Josh's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own.
| Rank | Proposal | Evidence | Model | Hours | Consequence per tier | Gate |
|---|---|---|---|---|---|---|
| 1 | Scheme C as class v5 content: the daily dataset derives from the state snapshot (`sd1`), spec text for 1.8.5 (the leaf line), 1.12 (the snapshot's checkpoint and lead), 10 (the pruning-proof witness), 4.3's pause rule applied to the day boundary | section 5.2: the hash kernel and rate are unchanged by construction, the build and verifier costs are the measured rows; every miner operation must hold state | the verifier row: 2.06 ms + the leaf-read row per unit; node RAM + the dataset size | spec 3 h; emitter and `memhard.rs` leaf line 2 h; node snapshot builder (canonical serialisation, one pass per day, the `prepare` hand-over) 6 h; fast-time harness across a day boundary and a finality pause 3 h; Devnet 2 crossing 2 h | 8 to 32 GB cards: no hash-rate change, +2 GiB device memory during the build only (streamable); rig: the same per card; pool user: a 2 GiB daily download or nothing; solo miner: a full node's state; node operator: +2 GiB RAM (+4 at year 4) and a daily leaf pass; holder: a daily random sample of state per block; prover, rollup customer: nothing | the fast-time 3-node network crosses a day boundary with a non-trivial state and no fork; verifier under 10 ms on a 2019-class core with the snapshot in RAM; Devnet 2 PASS over a day boundary |
| 2 | Scheme B's AMD gate: the WMMA iu8 fragment layout on the RX 9070 XT against the spec reference (the two-output tile), a PC 1 job with the card alone | status item 6 (layout unverified); section 5.1's bit-exact rows on NVIDIA | the spec layout as the definition; a lane permutation per vendor | 3 h (the probe exists: `family-probe.cu` mm8 row; the OpenCL twin needs the builtin path and the permutation) | AMD 16 GB tier: decides whether RDNA 3 and 4 run B natively or emulate at about 12 to 16 steps per mm8 | 1,024 random units bit-exact on the 9070 XT, both tile outputs |
| 3 | Scheme B's Apple row: the emulated mm8 block in Metal on the M5 Max at R = 32, 128, 512, hash rate and watts under `with-lock.sh measure` | section 3.2's vendor table: Apple is the honest card's worst case; the M5 Max binds at about 130,000 counted ops | 4 shuffles plus 4 dot4 emulations per step, about 10 steps per mm8 (approximate) | 3 h | Apple tier: the 5 percent rule decides the largest R the class can carry | rate within 5 percent of mx8 at the chosen R; bit-exact against the C reference |
| 4 | Scheme B as class v5 content beside C (`mx8+sd1+mm8xR`) once ranks 2 and 3 pass: the emitter's three dialects (inline PTX with the `IGNEUM_MM8_REF` fallback for sm_61 to sm_72, the AMD builtin behind a feature test, the Metal emulation), the verifier step in `verify.rs`, the stats and fuzz runs of `TESTS.md` on the class, the v5 switch beside v4 | section 5.1 and 5.3: the block's measured watts and the chip rows at k | the chip rows of 5.3 | emitter 4 h; verifier 1 h; suites 2 h; node seam 2 h; Devnet 2 2 h | NVIDIA Turing and later, AMD RDNA 3 and later: native, the measured watts; Pascal, RDNA 2, Apple: emulation at the measured 5 percent check; pool user: nothing; chip: a tensor array per 32 lanes in flight | the six gates of `counter-asic-2-rollout.md` 7 plus the AMD and Apple rows; verifier under 10 ms on a 2019-class core at the chosen R |
| 5 | The reserve text correction: R8 `mm8` consumes both tile outputs (two destinations) so a chip cannot halve the work | section 3.2 (the half-tile shortcut) | arithmetic: 32 of 64 products consumed in the single-output form | 1 h (spec 1.13.2 text and the edge vectors) | none until era 4 or a signal | the R8 edge vectors re-cut for two outputs, bit-exact on the three vendors |
| 6 | The light-client state sample: a node RPC that returns, for a block, the 128 leaves of lane n with openings against R_d, and a client check | section 3.3 (the sample is real but modest: availability, not correctness) | 128 x 25 x 32 B = 102 KB per block fetched on demand (arithmetic) | 4 h | holder and light client: a spot check of state availability per block; node: one RPC | 128 openings verify against R_d on the fast-time network |
| 7 | The 2019-class core measurement (O-1.14) for the verifier under C and under B at the chosen R | every verifier row here is M5 Max or Zen 4 plus the 2.5x rule | the 2.5x rule stands in | 1 h once a core is found (a 2019 laptop or a rented older CPU box) | every tier: the gate that fixes R and the snapshot read budget | under 10 ms per unit, worst cold |
| 8 | Do not build scheme A (mining is proving) in any form; record the sampleability bound (proving work per segment over network hashes per segment: 8 percent at 1 GH/s, 0.08 percent at 100 GH/s) in the ledger beside F13 | section 3.1 and 4.1 | the bound's arithmetic | 0.5 h (a ledger row) | none | none |
One paragraph each.
**1. Scheme C first.** It is the cheapest change with a new property: the kernel text, the hash rate and the chip rows are untouched, so every measured number of Counter ASIC 2.0 and 3.0 stands; what changes is the daily build and the node. The cost is a canonical state serialisation in consensus, which is why the fast-time harness must cross a day boundary with state and a finality pause before the Devnet 2 crossing. The pool caveat is stated: the design forces the operation, not the card, to hold state.
**2 and 3. B's two vendor gates.** B is bit-exact tonight on NVIDIA (section 5.1) and in every emulation; it is not a class candidate until an AMD card reproduces the spec layout and the Apple emulation's hash-rate cost is measured. Both are short jobs with the card alone.
**4. B beside C.** The order matters: C lands first because it needs no vendor work; B joins the same class byte or the next one once ranks 2 and 3 pass, with R set from 5.1 and the 2019-core row.
**5 to 8.** Small, named, and each closes a thread this lane opened.
## 8. Open questions and what could not be run
| Question | Why it could not be closed tonight | What closes it |
|---|---|---|
| The 2019-class core (O-1.14) | no such core in the fleet; igneum-build-1 is Zen 4, the Mac is M5 Max; the 2.5x rule stands in | rank 7 |
| The AMD WMMA fragment layout | PC 1 is Josh's desk and the AMD rows were owed all day (status file); no AMD card on RunPod or Vast tonight (fleet agent) | rank 2 |
| The Apple emulation of mm8 | a Metal emulation kernel is a 3-hour job and the Mac measure lock was free; not started because the AMD gate decides first whether B proceeds | rank 3 |
| The canonical state serialisation for C | a design item that touches the exec layer (`igneum/exec`), out of this lane's files | rank 1 |
| The pruning-proof witness for C's historical PoW | spec 02 and 10 items | rank 1 |
| Whether the tensor path's marginal energy on the 5090 differs from the 4090's | one card measured (box 1); the 5090 is on Josh's desk | a PC 2 job with the same `run.sh` |
| The Ampere row (3060, 3080, 3090) | every Ampere card of the fleet was mining and proving the live devnet; a loaded 3090 was offered and declined (a loaded card's rate is not a number) | one quiet Ampere pod |
| The verifier with a byte-dot instruction (VNNI, NEON udot) | the C reference is scalar | 1 h: an AVX-VNNI and a NEON path in `verify_ref.c`, measured on both cores |
| Scheme C's leaf array on an 8 GB card at the 2 GiB design size | the build holds dataset + cache + leaves on the device (4.3 GiB) unless chunked | the chunked build (section 5.2 says whether it is trivial) |
## 9. Summary for the coordinator
SUMMARY

View file

@ -0,0 +1,12 @@
# Horizon lane 2 (algorithm): the model behind docs/analysis/horizon/algorithm.md
Run (plain Python 3, no numpy, under a second):
python3 sim/horizon/algorithm/model.py # every table, markdown on stdout
python3 sim/horizon/algorithm/model.py --section chip # chip | fpga | v5 | era | verifier | schedule
Every input is listed at the top of `model.py` with its source file or URL and its label (measured, cited,
approximate). The verifier rows are this lane's own measurement of 6 October 2026 on igneum-build-1 (the
`igneum-pow` crate built there through `tools/build-remote.sh`, `IGNEUM_AGENT=horizon`, one slot, 10 s; the bench
pinned with `taskset -c 40` under `nice -n 19`, then both SMT siblings 40 and 88 at once); the commands and raw lines
are in section 2 of the document.

View file

@ -0,0 +1,392 @@
#!/usr/bin/env python3
"""Horizon lane 2 (algorithm): the arithmetic behind docs/analysis/horizon/algorithm.md.
Every number printed here is arithmetic on a cited or measured input; the inputs are listed with their source
in INPUTS below and in the document. Nothing here measures a chip. Run:
python3 sim/horizon/algorithm/model.py # prints every table as markdown
python3 sim/horizon/algorithm/model.py --section chip|fpga|v5|era|verifier|schedule
No numpy needed (plain Python 3). Sections:
chip per-joule and per-dollar edge of the three chip classes against each honest card, class v3 and v4
fpga HBM2 FPGA random-read ceilings from bank count, tRC, tFAW, tRRD; reads in flight; per watt
v5 shadow N options: chip edge per joule at k = 0.3, 0.5, 1, 1.5; verifier cost; per-tier watts
era the era draw: what grinding a certified checkpoint costs, and what a draw can move
verifier the 2019-class gate: the box proxy rows against the M5 Max and the 2.5x rule
schedule dataset growth against the Steam VRAM shares and the prover footprint per tier
"""
import sys
# ----------------------------------------------------------------------------------------------------------------
# INPUTS (measured = a bench-log or analysis entry; cited = a URL or paper; approx = from memory or an estimate)
# ----------------------------------------------------------------------------------------------------------------
PJ = 1e-12
UJ = 1e-6
# Honest cards: name, MH/s, watts, source, tier. The v3 class (mx8) unless said otherwise.
CARDS = [
# owned, measured 5 and 6 October 2026
("RTX 5090 (PC 2 bench, 431 W cap, control)", 132.2, 350.0, "latency-shadow-2026-10-06.md s5 (bench-log item 8)", "32 GB"),
("RTX 5090 (PC 1 app, Ember run 5)", 127.4, 310.0, "counter-asic-3-status.md s7 item 1", "32 GB"),
("RTX 5090 (rented Vast pod, untuned)", 98.48, 258.2, "prover-tiers-real-cards.md", "32 GB"),
("RTX 4090 (rented)", 52.25, 183.1, "prover-tiers-real-cards.md", "24 GB"),
("RTX 3090 (rented)", 37.79, 228.8, "prover-tiers-real-cards.md", "24 GB"),
("RTX A5000 (rented)", 47.7, 222.7, "prover-tiers-real-cards.md", "24 GB"),
("RTX 4070 (PC 1, 1,860 MHz lock, 160 W cap)", 30.95, 79.5, "counter-asic-3-status.md item 8 4070 rows", "12 GB"),
("RTX 4070 (rented, untuned)", 24.99, 91.1, "prover-tiers-real-cards.md", "12 GB"),
("RTX 5070 (rented)", 41.89, 137.0, "prover-tiers-real-cards.md", "12 GB"),
("RTX 3060 (rented)", 23.78, 103.7, "prover-tiers-real-cards.md", "12 GB"),
("RTX 3080 (rented)", 40.82, 204.9, "prover-tiers-real-cards.md", "10 GB"),
("RTX 4060 Ti 16 GB (rented)", 17.58, 72.3, "prover-tiers-real-cards.md", "16 GB"),
("RTX 4060 Ti 8 GB (rented)", 19.07, 72.6, "prover-tiers-real-cards.md", "8 GB"),
("RX 9070 XT (PC 1, ADLX 199 to 203 W)", 18.9, 201.0, "bench-log 9070 XT telemetry; status 3a watts job", "16 GB AMD"),
("Apple M5 Max (GPU + DRAM channels)", 27.08, 21.0, "latency-shadow-2026-10-06.md s3", "Apple"),
("Apple M5 Max (package, approx 38 W)", 27.08, 38.0, "ember-tune.md Apple row, approximate", "Apple"),
]
# the RTX 4060 (rented) logged 0.0 W: no watt reading, left out of the per-joule rows
# Marginal ALU energy per counted op at the shadow (measured): 5090 11 pJ (10.2 to 13.2), M5 Max 6.9 pJ
MARGINAL_PJ = {"NVIDIA": 11.0, "Apple": 6.9, "AMD": 11.0} # AMD unmeasured: NVIDIA's figure, approx
N_V4 = 100_000 # counted ops per hash in the class v4 shadow (sh256x27: 102,100 counted, 55,296 instrs)
# Measured v4 cards (whole card): 5090 3.27 uJ at 431 W; M5 Max 1.40 uJ at 37 W; 4070 3.51 uJ at 109 W
MEASURED_V4 = {
"RTX 5090 (PC 2 bench, 431 W cap, control)": (131.95, 431.0),
"RTX 4070 (PC 1, 1,860 MHz lock, 160 W cap)": (31.08, 109.0),
"Apple M5 Max (GPU + DRAM channels)": (26.67, 37.2),
"Apple M5 Max (package, approx 38 W)": (26.67, 54.0), # 37.2 + the same 17 W of package, approx
}
# Chips (chip-model-v3.md s5.4 and latency-shadow s6; all approximate arithmetic on cited memory figures)
CHIPS = [
# name, MH/s, watts at v3, silicon+memory dollars at v3, extra dollars for the v4 ALU array
("f = 0 on-die recompute (256 MiB SRAM, x8 mixer)", 41.7, 53.7, 700.0, 40.0),
("f = 1 stored dataset, GDDR7 (16 devices, 512-bit)", 166.4, 77.6, 470.0, 40.0),
("f = 1 stored dataset, HBM3 one stack", 83.6, 26.8, 550.0, 40.0),
("f = 1 stored dataset, HBM3 eight stacks", 666.4, 174.4, 2650.0, 40.0),
]
K_LIST = [0.3, 0.5, 1.0, 1.5]
# Card list prices, approximate (launch MSRP from memory unless cited), USD
CARD_PRICE = {
"RTX 5090": 1999.0, "RTX 4090": 1599.0, "RTX 3090": 1499.0, "RTX A5000": 2250.0, "RTX 4070": 549.0,
"RTX 5070": 549.0, "RTX 3060": 329.0, "RTX 3080": 699.0, "RTX 4060 Ti 16 GB": 499.0, "RTX 4060 Ti 8 GB": 399.0,
"RX 9070 XT": 599.0, "Apple M5 Max": 3500.0,
}
RENTAL_USD_PER_MHS_HOUR = 0.0117 # bench-log "Rental cost of hash, 6 October 2026" (measured, RunPod list)
ELECTRICITY_USD_PER_KWH = 0.10 # approx
AMORTISE_HOURS = 2 * 365 * 24 # two years, approx
def uj(mhs, watts):
return watts / (mhs * 1e6) / UJ
def card_price(name):
for k, v in CARD_PRICE.items():
if name.startswith(k):
return v
return None
def hourly_cost_per_mhs(price, uj_per_hash):
cap = price / AMORTISE_HOURS # USD per hour for the whole device
return cap, uj_per_hash * 3.6e9 * UJ / 3.6e6 * ELECTRICITY_USD_PER_KWH # (capex/h total, energy USD per MH/s-h)
def section_chip():
print("## Chip edge per joule and per dollar, class v3 and class v4\n")
print("Chip rows (model, approximate): energy per hash at v3 = watts / rate; at v4 the chip adds N x 11 pJ x k for its")
print(f"shadow core (N = {N_V4:,} counted ops, k = chip core energy per op over the 5090's measured 11 pJ).\n")
print("| Chip class | MH/s | W at v3 | uJ per hash, v3 | uJ at v4, k = 0.3 / 0.5 / 1 / 1.5 | $ silicon + memory, v3 / v4 | $ per MH/s, v3 |")
print("|---|---|---|---|---|---|---|")
chip_rows = []
for name, mhs, w, usd, usd_v4 in CHIPS:
e3 = uj(mhs, w)
e4 = [e3 + N_V4 * 11.0 * k * PJ / UJ for k in K_LIST]
chip_rows.append((name, e3, e4, usd, usd_v4))
print(f"| {name} | {mhs:.1f} | {w:.1f} | {e3:.3f} | " + " / ".join(f"{x:.2f}" for x in e4) +
f" | {usd:.0f} / {usd + usd_v4:.0f} | {usd / mhs:.1f} |")
print()
print("Honest cards at v3 (measured) and at v4 (measured where a row exists, else the card's marginal ALU energy times N;")
print("a card under a power cap that binds at v4 pays less energy and loses rate instead, as the 5090 did).\n")
print("| Card (tier) | MH/s | W | uJ, v3 | uJ, v4 | Edge of f=0 chip per joule, v3 / v4 at k=1 | f=1 GDDR7, v3 / v4 at k=0.3 / 0.5 / 1 / 1.5 | f=1 HBM3 one stack, v3 / v4 at k=1 | $ list (approx) | $ per MH/s | USD per MH/s-hour owned (capex 2 y + 0.10/kWh) |")
print("|---|---|---|---|---|---|---|---|---|---|---|")
for name, mhs, w, src, tier in CARDS:
e3 = uj(mhs, w)
vendor = "Apple" if "Apple" in name else ("AMD" if "RX " in name else "NVIDIA")
if name in MEASURED_V4:
m4, w4 = MEASURED_V4[name]
e4 = uj(m4, w4)
tag = ""
else:
e4 = e3 + N_V4 * MARGINAL_PJ[vendor] * PJ / UJ
tag = "*"
f0, f1g, f1h = chip_rows[0], chip_rows[1], chip_rows[2]
price = card_price(name)
pm = f"{price:.0f}" if price else "n/a"
ppm = f"{price / mhs:.1f}" if price else "n/a"
if price:
cap, en = hourly_cost_per_mhs(price, e3)
own = f"{cap / mhs + en:.5f}"
else:
own = "n/a"
print(f"| {name} ({tier}) | {mhs:.2f} | {w:.1f} | {e3:.2f} | {e4:.2f}{tag} | {e3 / f0[1]:.2f}x / {e4 / f0[2][2]:.2f}x | "
f"{e3 / f1g[1]:.1f}x / " + " / ".join(f"{e4 / x:.1f}x" for x in f1g[2]) +
f" | {e3 / f1h[1]:.1f}x / {e4 / f1h[2][2]:.1f}x | {pm} | {ppm} | {own} |")
print("\n`*` = modelled v4 energy (marginal pJ x N), not measured. Rental hash for comparison: USD 0.0117 per MH/s-hour (measured, RunPod).")
# chip hourly cost
print("\nChip hourly cost per MH/s (amortised over two years plus electricity at USD 0.10 per kWh, approximate):\n")
print("| Chip | uJ, v3 | USD per MH/s-hour | Against rental (0.0117) | Against an owned 5090 (bench row) |")
print("|---|---|---|---|---|")
p5090 = CARD_PRICE["RTX 5090"]
e5090 = uj(132.2, 350.0)
cap5, en5 = hourly_cost_per_mhs(p5090, e5090)
own5090 = cap5 / 132.2 + en5
for name, mhs, w, usd, usd_v4 in CHIPS:
e3 = uj(mhs, w)
cap, en = hourly_cost_per_mhs(usd, e3)
c = cap / mhs + en
print(f"| {name} | {e3:.3f} | {c:.6f} | {RENTAL_USD_PER_MHS_HOUR / c:.0f}x cheaper | {own5090 / c:.1f}x cheaper |")
print(f"\nOwned 5090 (bench row): USD {own5090:.5f} per MH/s-hour; rental is {RENTAL_USD_PER_MHS_HOUR / own5090:.0f}x that.")
def section_fpga():
print("## HBM2 FPGA random-read ceilings (U55C, U280, AWS F2 VU47P: 2 stacks, 16 channels, 32 pseudo-channels)\n")
channels = 16
pcs = 32
banks_per_pc = 16 # approx: HBM2 16 banks per channel, each split into half banks per pseudo-channel
clk_mhz = 1066.0 # 2,133 MT/s data rate, JEDEC HBM2 as tabled by ICCAD 2021 Table I (cycles)
ns = 1000.0 / clk_mhz
tRC = (14 + 34) * ns # tRCD + tRAS cycles (ICCAD 2021 Table I): 48 cycles
tFAW = 30 * ns # 4 activates per channel per tFAW (Table I)
tRRD = 6 * ns # ACT to ACT, different bank (Table I)
bank_bound = pcs * banks_per_pc / (tRC * 1e-9)
faw_bound = channels * 4 / (tFAW * 1e-9)
rrd_bound = channels / (tRRD * 1e-9)
oconnor_bound = channels * 8 / (12e-9) # the figure chip-model-v3.md s5.3 carried for HBM2/3 (O'Connor Table 2)
lat = 137.8 # ns, Shuhai Table IV page miss on the U280
print(f"tRC {tRC:.1f} ns, tFAW {tFAW:.1f} ns (4 ACT per channel), tRRD {tRRD:.1f} ns, latency {lat} ns (page miss, measured).\n")
print("| Ceiling | Formula | G reads/s per card | Reads in flight at 137.8 ns | Per W at 115 / 150 / 225 W (M reads/s/W) | Against the 5090 per W (53.7 M at 326 W; 50.0 M at 350 W) |")
print("|---|---|---|---|---|---|")
rows = [
("Measured, Shuhai U280 default mapping", "2.4 G (FCCM 2020, Fig 7)", 2.4e9),
("tFAW-bound (JEDEC HBM2 cycles, ICCAD 2021 Table I)", "16 ch x 4 / 28.1 ns", faw_bound),
("tRRD-bound", "16 ch / 5.6 ns", rrd_bound),
("Bank-bound, no activate window (the 12.2 ceiling row)", "32 pc x 16 banks / 45 ns", bank_bound),
("O'Connor HBM2 activate figure as carried by chip-model-v3 s5.3", "16 ch x 8 / 12 ns", oconnor_bound),
]
for name, f, g in rows:
flight = g * lat * 1e-9
pw = [g / w / 1e6 for w in (115, 150, 225)]
print(f"| {name} | {f} | {g / 1e9:.1f} | {flight:.0f} | " + " / ".join(f"{x:.0f}" for x in pw) +
f" | {pw[0] / 53.7:.2f}x to {pw[1] / 53.7:.2f}x (U55C); {pw[2] / 53.7:.2f}x (U280) |")
print("\nReading: the measured 2.4 G/s equals the tFAW-bound ceiling at JEDEC HBM2 timings (2.3 G/s), so the measured row is")
print("not a mapping artefact: it is the activate window. The bank-bound 11.4 G/s needs tFAW removed, which the DRAM die")
print("enforces and no controller overrides. The tightened range for a 2-stack HBM2 FPGA: 2.3 to 2.9 G/s, 0.30x to 0.47x of")
print("the 5090 per watt at the U55C's 115 to 150 W; the 0.7x to 1.9x row survives only if HBM3-class parts (tFAW about half,")
print("4 bank groups, approximate) reach an FPGA, and none has (Versal HBM and Agilex M are HBM2e).")
def section_v5():
print("## Class v5 options: the shadow N against the chip's k\n")
# honest cards at N: measured ladders where they exist, else model
# 5090 measured (431 W cap): N -> (MH/s, W)
r5090 = {930: (132.2, 350), 49700: (132.28, 424.6), 102100: (131.95, 431), 150800: (131.75, 431), 199600: (128.67, 431), 330700: (86.39, 431)}
rmac = {930: (27.08, 21.0), 49700: (26.75, 31.1), 102100: (26.67, 37.2), 150800: (25.10, 36.7), 199600: (24.25, 38.2), 330700: (21.39, 40.1)}
r4070 = {930: (30.95, 79.5), 49700: (31.08, 93.5), 102100: (31.08, 109.0), 199600: (31.07, 138.2), 330700: (27.14, 159.9)}
gddr7 = uj(166.4, 77.6)
hbm3 = uj(83.6, 26.8)
print("| N counted ops | 5090 uJ (rate delta) | M5 Max uJ (delta) | 4070 uJ (delta) | 9070 XT | f=1 GDDR7 chip uJ at k = 0.3 / 0.5 / 1 / 1.5 | Edge over the 5090 | Edge over the M5 Max | Edge over the 4070 | Verifier add, M5 Max core / box core / box half-core (ms per warp) |")
print("|---|---|---|---|---|---|---|---|---|---|")
# verifier slopes: Mac 3.2 us per 1,000 shadow instrs; box one core 7.0; box half core 12.1 (measured 6 Oct, this lane)
for N in (930, 49700, 102100, 130000, 150800, 199600, 330700):
def interp(tbl):
ks = sorted(tbl)
if N in tbl:
return tbl[N]
lo = max(k for k in ks if k < N); hi = min(k for k in ks if k > N)
t = (N - lo) / (hi - lo)
return (tbl[lo][0] + t * (tbl[hi][0] - tbl[lo][0]), tbl[lo][1] + t * (tbl[hi][1] - tbl[lo][1]))
m5, w5 = interp(r5090); mm, wm = interp(rmac); m4, w4 = interp(r4070)
e5, em, e4 = uj(m5, w5), uj(mm, wm), uj(m4, w4)
shadow = max(N - 930, 0)
chip = [gddr7 + shadow * 11.0 * k * PJ / UJ for k in K_LIST]
instrs = shadow / 1.83
vadd = (instrs / 1000 * 3.2e-3, instrs / 1000 * 7.0e-3, instrs / 1000 * 12.1e-3)
print(f"| {N:,} | {e5:.2f} ({(m5 / 132.2 - 1) * 100:+.1f}%) | {em:.2f} ({(mm / 27.08 - 1) * 100:+.1f}%) | {e4:.2f} ({(m4 / 30.95 - 1) * 100:+.1f}%) | holds to 331k (watts owed) | "
+ " / ".join(f"{c:.2f}" for c in chip) + " | " + " / ".join(f"{e5 / c:.1f}x" for c in chip) + " | "
+ " / ".join(f"{em / c:.2f}x" for c in chip) + " | " + " / ".join(f"{e4 / c:.1f}x" for c in chip)
+ f" | {vadd[0]:.2f} / {vadd[1]:.2f} / {vadd[2]:.2f} |")
print("\nThe 130,000 row is interpolated between the measured 102,100 and 150,800 rungs (the Mac's 5 percent point).")
print("\nOption (ii), a shuffle-heavy shadow mix: the chip's datapath floor per op (N5, approx) by op class and the GPU's measured step cost.\n")
print("| Op class | Chip datapath energy per op, N5 floor (pJ, approx) | GPU step cost on the 5090 (ratio to the add-xor-rotate chain, measured) | Structure a chip must add |")
print("|---|---|---|---|")
ops = [("add, sub, xor, or (weights 32 of 75)", 0.06, "1.00", "adder, gates"),
("mul, mulhi, mad (22 of 75)", 0.52, "1.00 (IMAD is the chain's own op)", "32x32 multiplier"),
("rotl, rotr (13 of 75)", 0.06, "1.32", "barrel rotator"),
("shfl xor-mask (8 of 75)", 0.30, "1.49", "5-stage butterfly per warp (5,120 mux bits)"),
("shfla lane+delta (reserve R3)", 1.00, "1.53", "32x32x32-bit crossbar per warp (32,768 mux bits)"),
("mm8 (reserve R8)", 1.60, "2.43", "8x8x16 u8 MAC tile per warp (1,024 MACs)")]
for n, e, g, s in ops:
print(f"| {n} | {e:.2f} | {g} | {s} |")
def mix_floor(w_alu, w_mul, w_rot, w_shfl, w_shfla):
tot = w_alu + w_mul + w_rot + w_shfl + w_shfla
return (w_alu * 0.06 + w_mul * 0.52 + w_rot * 0.06 + w_shfl * 0.30 + w_shfla * 1.00) / tot
base = mix_floor(32, 22, 13, 8, 0)
heavy = mix_floor(26, 18, 9, 14, 8)
print(f"\nShadow mix energy floor per op: today's weights {base:.3f} pJ; a shuffle-heavy v5 mix (shfl 14, shfla 8 of 75, the")
print(f"others scaled) {heavy:.3f} pJ, {heavy / base:.2f}x. With the 2x pipeline and 8x register-file and wire overhead of")
print(f"latency-shadow s6 the k floor moves from {base * 16 / 11:.2f} to {heavy * 16 / 11:.2f} (approx): the shuffle-heavy mix")
print("raises the attacker's claimed floor, it does not reach k = 1.")
def section_era():
print("## The era draw: what grinding costs and what a draw can move\n")
print("Forging the certified checkpoint the era VDF reads needs 2/3 of the 30-day blue-block weight: 20 days of 100 percent")
print("hash (CLAUDE.md headline). Rented at USD 0.0117 per MH/s-hour (measured, RunPod list, 6 October 2026):\n")
print("| Network hash | 20 days of 100 percent, rented | What the forged checkpoint buys in the draw |")
print("|---|---|---|")
for g in (1, 10, 100, 1000):
usd = g * 1000 * 480 * RENTAL_USD_PER_MHS_HOUR
print(f"| {g} GH/s | USD {usd:,.0f} | one era's (M, R, pos, weights +-2, fold rotations): 0.8 to 3.2 percent hash-rate spread per card (measured six eras), 0 chip effect |")
print("\nWithholding the last blue block before C_era(n) to re-roll the input costs one block and needs the 3,600 s VDF")
print("evaluated inside the 2 s publish window: a 1,800x faster evaluator (spec 04 s4.6 margin table: 300x beats the epoch,")
print("not the era). Gain of the re-roll even if free: one draw of the same space.\n")
print("Op-weight corners: each of the ten non-load weights is perturbed by -2..+2 points; the multiply share (mul, mad,")
print("mulhi: 22 of 75) can move to 16 or 28 of 75. Chip datapath energy per shadow op at the N5 floor (0.06 add, 0.52 mul):")
for share in (16, 22, 28):
e = (share * 0.52 + (75 - share) * 0.06) / 75
print(f" multiply share {share}/75: {e:.3f} pJ per op")
print("so the weakest corner for a chip is a draw with the fewest multiplies, worth about 20 percent of the shadow's")
print("datapath energy (approx) and nothing on the memory side; the GPU's cost moves the same way.")
def section_verifier():
print("## The 2019-class verifier gate: measured proxies (this lane, 6 October 2026, igneum-build-1)\n")
mac = {"v2": 0.60, "mx8 (class v3)": 2.06, "mx8+sh256x27 (class v4)": 2.33, "dr368": 2.69, "dr736": 4.88}
box1 = {"v2": 1.128, "mx8 (class v3)": 4.516, "mx8+sh256x27 (class v4)": 4.901, "dr368": 4.716, "dr736": 9.763}
box1cold = {"v2": 1.305, "mx8 (class v3)": 4.670, "mx8+sh256x27 (class v4)": 5.061, "dr368": 5.324, "dr736": 10.509}
half = {"mx8 (class v3)": 7.56, "mx8+sh256x27 (class v4)": 8.23, "dr368": 8.16, "dr736": 15.49}
print("| Class | M5 Max core, quiet (ms per warp) | 2.5x rule (approx) | Box EPYC 9454P one core at 3.8 GHz, nice 19, taskset (steady / cold) | Box over Mac | Box half-core (SMT sibling loaded) | Gate 10 ms |")
print("|---|---|---|---|---|---|---|")
for c in mac:
h = half.get(c)
verdict = "pass" if box1cold[c] < 10 and (h is None or h < 10) else ("FAIL cold on one core" if box1cold[c] >= 10 else "fails the half-core proxy")
print(f"| {c} | {mac[c]:.2f} | {mac[c] * 2.5:.1f} | {box1[c]:.2f} / {box1cold[c]:.2f} | {box1[c] / mac[c]:.2f}x | {('%.2f' % h) if h else 'n/a'} | {verdict} |")
print("\nCache fill on the box core: 361 ms (M5 Max 175 to 181 ms): 2.0x. The box clock read 3,800 MHz during the run")
print("(scaling_cur_freq), so this is a 2022 server core at full boost with server DDR5 latency, not a 2019 laptop; the")
print("half-core row (both SMT siblings busy) is the pessimistic bracket. What the gate protects at 10 ms per warp:\n")
for ms in (10.0, 20.0):
print(f"| at {ms:.0f} ms per warp | 1 bps: {ms / 1000 * 100:.0f}% of one core | 10 bps: {ms / 1000 * 10 * 100:.0f}% of one core | IBD 108,000 headers: {108000 * ms / 1000 / 60:.0f} min | header flood to saturate one core: {1000 / ms:.0f} invalid headers per second | pool: {1000 / ms:.0f} shares per second per core |")
def section_schedule():
print("## Dataset growth against the installed base and the prover footprint\n")
steam = [("512 MB to 4 GB", 2.60 + 1.44 + 3.69 + 1.20 + 5.13), ("6 GB", 5.12), ("8 GB", 26.71), ("10 to 11 GB", 1.91 + 0.83),
("12 GB", 13.06), ("16 GB", 27.21), ("20 to 24 GB", 1.39 + 1.04 + 5.50), ("32 GB", 1.41), ("64 GB", 0.50), ("other", 1.26)]
print("Steam Hardware Survey, September 2026 (cited, store.steampowered.com/hwsurvey; 8 GB was 35.03 percent in August 2025")
print("and 33.66 in September 2025, 12 GB 19.30 in August 2025: videocardz, wccftech, pcguide, cited):\n")
print("| VRAM | Share of Steam users, Sep 2026 |")
print("|---|---|")
for n, s in steam:
print(f"| {n} | {s:.2f}% |")
# trend: 8 GB -7.3 points per year (33.66 -> 26.71 Sep 2025 to Sep 2026); 12 GB -6.2 (19.30 Aug 2025 -> 13.06)
print("\nLinear extrapolation of the 8 GB and 12 GB shares (approx; the 16 GB and 24 GB tiers absorb them):\n")
print("| Year | 8 GB share | 12 GB share | Dataset (option b) | Cache |")
print("|---|---|---|---|---|")
for y, ds, cache in ((2026, "1 GiB (devnet packs)", "256 MiB"), (2027, "2 GiB (genesis)", "256 MiB"), (2028, "2 GiB", "256 MiB"), (2029, "2 GiB", "256 MiB"),
(2030, "2 GiB", "256 MiB"), (2031, "4 GiB (year 4)", "512 MiB"), (2035, "4 GiB", "512 MiB"), (2039, "8 GiB (year 12)", "1 GiB")):
s8 = max(0.0, 26.71 - 7.0 * (y - 2026)); s12 = max(0.0, 13.06 - 6.2 * (y - 2026) * 0.5)
print(f"| {y} | {s8:.0f}% | {s12:.0f}% | {ds} | {cache} |")
print("\nMemory per tier at each step (MiB): miner resident = dataset + 128 output buffers + scratch-free kernel (cache freed")
print("after the build, the decided reading); prover beside the miner from prover-tiers-real-cards.md (peak of the")
print("compressed 2^26 and core-only 2^25 profiles, both with a 1.4 GB miner resident at the 1 GiB dataset):\n")
print("| Tier (usable = 75% of VRAM, Apple 50%) | Usable MiB | Mine-only: last step that fits | Mine + prove compressed (peak 8.9 to 10.7 GB at 1 GiB): last step | Mine + prove core-only (peak 7.1 to 7.4 GB at 1 GiB): last step |")
print("|---|---|---|---|---|")
steps = [("1 GiB (today)", 1024), ("2 GiB (genesis)", 2048), ("4 GiB (year 4)", 4096), ("8 GiB (year 12)", 8192), ("16 GiB (year 28)", 16384)]
# (tier, VRAM MiB, measured compressed peak beside the miner at the 1 GiB dataset in MiB or None, measured core-only peak)
tiers = [("8 GB (4060 Ti 8 GB)", 8192, None, 7352), ("12 GB (4070, 5070)", 12288, 10240, 5734 + 1434), ("16 GB (4060 Ti 16 GB)", 16384, 9216, 5939 + 1434),
("24 GB (4090, 3090, A5000)", 24576, 10956, 6246 + 1434), ("32 GB (5090)", 32768, 10138, 6451 + 1434)]
for name, vram, comp_peak1, core_peak1 in tiers:
usable = vram * 0.75
mine_last = "none"; comp_last = "none"; core_last = "none"
for sname, ds in steps:
if ds + 128 + 64 <= usable:
mine_last = sname
grow = ds - 1024 # the miner's resident set grows with the dataset; the prover's part does not
if comp_peak1 is not None and comp_peak1 + grow <= vram * 0.98: # headless Linux: the rows used nearly the whole card
comp_last = sname
if core_peak1 + grow <= vram * 0.98:
core_last = sname
print(f"| {name} | {usable:.0f} | {mine_last} | {comp_last} | {core_last} |")
print("\nReading: the 8 GB tier never mines and proves compressed at any dataset size (measured: does not fit at 1 GiB); core-only")
print("fits at 1 GiB with 1 GB spare and loses that at the 2 GiB genesis step. The 12 GB tier mines and proves compressed on")
print("headless Linux at 1 and 2 GiB and loses it at the year-4 step (4 GiB). The 16 GB tier holds compressed to year 4 and")
print("core-only to year 12. Mine-only: 8 GB to year 12, 12 and 16 GB to year 28, as card-lifetime-2026-10-05.md says.")
def section_ladder():
"""The reconciled N ladder asked for by the coordinator (lane 7's HBM4 inputs, this lane's measured cards)."""
print("## The reconciled N ladder (lane 2 measured cards, lane 7 HBM4 inputs)\n")
# chips: bare uJ per hash (memory-bound f = 1): 128 x E_read + static / rate. Lane 7's HBM4 inputs (approximate):
# one stack, 32 channels, ceiling 2 x 10.7 G = 21.4 G reads/s, 1.0 nJ per read, 5 W static, 10 W controller.
hbm4_rate = 21.4e9 / 128
hbm4_bare = (128 * 1.0e-9 * hbm4_rate + 15.0) / hbm4_rate / UJ
chips = [("GDDR7", uj(166.4, 77.6)), ("HBM3 1 stack", uj(83.6, 26.8)), ("HBM3 8 stacks", uj(666.4, 174.4)), ("HBM4 1 stack (lane 7)", hbm4_bare)]
r5090 = {930: (132.2, 350), 49700: (132.28, 424.6), 102100: (131.95, 431), 150800: (131.75, 431), 199600: (128.67, 431), 330700: (86.39, 431)}
rmac = {930: (27.08, 21.0), 49700: (26.75, 31.1), 102100: (26.67, 37.2), 150800: (25.10, 36.7), 199600: (24.25, 38.2), 330700: (21.39, 40.1)}
r4070 = {930: (30.95, 79.5), 49700: (31.08, 93.5), 102100: (31.08, 109.0), 199600: (31.07, 138.2), 330700: (27.14, 159.9)}
r9070 = {930: 18.92, 49700: 19.31, 102100: 19.29, 199600: 19.11, 330700: 19.60}
def interp(tbl, N):
ks = sorted(tbl)
if N in tbl:
return tbl[N]
lo = max(k for k in ks if k < N); hi = min(k for k in ks if k > N)
t = (N - lo) / (hi - lo)
a, b = tbl[lo], tbl[hi]
return tuple(x + t * (y - x) for x, y in zip(a, b)) if isinstance(a, tuple) else a + t * (b - a)
print("Chip bare uJ per hash: " + ", ".join(f"{n} {e:.3f}" for n, e in chips) + ". Chip at N: bare + (N - 930) x 11 pJ x k.")
print("Card energy: measured ladders (5090 under its 431 W cap; M5 Max GPU + DRAM; 4070 at its 160 W cap); 130,000 interpolated.\n")
print("| N counted ops | 5090 uJ, W (rate delta) | M5 Max uJ, W (delta) | 4070 uJ, W (delta) | 9070 XT rate delta (W owed) | 12 GB and 8 GB rented cards | Chip edge over the 5090 per joule, k = 0.5 / 1 / 1.5: GDDR7 | HBM3 one stack | HBM3 eight stacks | HBM4 one stack | Verifier ms per warp: M5 Max core / 2019-class (2.5x rule) / half-core proxy | Card that binds first (5 percent rule) |")
print("|---|---|---|---|---|---|---|---|---|---|---|---|")
binds = {930: "none", 49700: "none", 102100: "none (M5 Max -1.5%)", 130000: "M5 Max at its 5% point", 150800: "M5 Max (-7.3%)", 199600: "M5 Max (-10%), 5090 (-2.7%)", 330700: "5090 (-35%), M5 Max (-21%), 4070 (-12%)"}
for N in (930, 49700, 102100, 130000, 150800, 199600, 330700):
m5, w5 = interp(r5090, N); mm, wm = interp(rmac, N); m4, w4 = interp(r4070, N); m9 = interp(r9070, N)
e5 = uj(m5, w5)
instrs = max(N - 930, 0) / 1.83
vmac = 2.06 + instrs / 1000 * 3.2e-3
vrule = vmac * 2.5
vhalf = 7.56 + instrs / 1000 * 12.1e-3
cells = []
for name, bare in chips:
cells.append(" / ".join(f"{e5 / (bare + max(N - 930, 0) * 11.0 * k * PJ / UJ):.1f}x" for k in (0.5, 1.0, 1.5)))
print(f"| {N:,} | {e5:.2f}, {w5:.0f} W ({(m5 / 132.2 - 1) * 100:+.1f}%) | {uj(mm, wm):.2f}, {wm:.0f} W ({(mm / 27.08 - 1) * 100:+.1f}%) | {uj(m4, w4):.2f}, {w4:.0f} W ({(m4 / 30.95 - 1) * 100:+.1f}%) | {(m9 / 18.92 - 1) * 100:+.1f}% | not measured at any N (v3 only) | "
+ " | ".join(cells) + f" | {vmac:.2f} / {vrule:.1f} / {vhalf:.2f} | {binds[N]} |")
print("\nDisagreements with lane 7's model 1.4, named: (a) its honest card at N is the linear 326 to 575 W model (2.95 uJ at")
print("N = 100,000, 3.50 at 200,000, 4.22 at 330,000); the measured 5090 under its 431 W cap reads 3.27, 3.35 and 4.99 uJ with")
print("-0.2, -2.7 and -34.7 percent of rate (the cap binds from 102,100 ops; power.min_limit is 400 W so no cap under it exists);")
print("(b) its shadow core is 150 W fixed at the 5090's 136 MH/s, so on a 167 MH/s HBM4 chip it under-counts the core by 1.23x;")
print("energy per hash is N x 11 pJ x k whatever the chip's rate, which is what this table uses (HBM4 at N = 100,000, k = 1: 1.33 uJ,")
print(f"not 1.11; edge {uj(131.95, 431) / (hbm4_bare + 101170 * 11e-12 / UJ):.1f}x, not 2.65x); (c) its HBM3 and HBM4 ceilings (10.7 and 21.4 G) rest on 8 activates per 12 ns")
print("per channel; the JEDEC HBM2 cycle table gives 4 per 28 ns (this file, section 5.1), and HBM3's own tFAW is behind the paywall.")
hbm4_low = (128 * 1.0e-9 * (4.6e9 / 128) + 15.0) / (4.6e9 / 128) / UJ
print(f"If HBM3 and HBM4 carry HBM2's activate window the one-stack ceilings are 2.3 and 4.6 G and the HBM4 bare energy {hbm4_low:.2f} uJ,")
print(f"an edge of {2.65 / hbm4_low:.1f}x bare over the 5090 instead of 11x; the GDDR7 column (the 5090 measures 82 percent of its ceiling) is")
print("the one with a measured anchor and is the column to quote.")
print("\nVerifier headroom for N (the \"10x\" claim): on the M5 Max core 10 - 2.33 = 7.67 ms buys 2.4 M shadow instructions, N about")
print("4.5 M ops (19x); on the 2.5x rule 4.2 ms buys 525,000 instructions, N about 1.06 M (10x); on the measured half-core proxy 1.77 ms")
print("buys 146,000 instructions, N about 370,000 (3.7x). The cards bind first at every bracket: M5 Max 130,000, 5090 at 431 W")
print("210,000, 4070 at 160 W about 250,000, 9070 XT over 331,000.")
SECTIONS = {"chip": section_chip, "fpga": section_fpga, "v5": section_v5, "ladder": section_ladder, "era": section_era, "verifier": section_verifier, "schedule": section_schedule}
if __name__ == "__main__":
want = None
if len(sys.argv) > 2 and sys.argv[1] == "--section":
want = sys.argv[2]
for k, f in SECTIONS.items():
if want is None or want == k:
f()
print()

View file

@ -0,0 +1,13 @@
# sim/horizon/frontier
Lane 7 (frontier) of the Horizon programme, 6 October 2026. One script, pure Python 3 (no numpy):
python3 sim/horizon/frontier/frontier_model.py > sim/horizon/frontier/out.md
It prints the arithmetic behind `docs/analysis/horizon/frontier.md`: the 2030 predictions (VRAM, memory
dollars, random-read ceilings for GDDR7, HBM3 and HBM4, the stored-dataset chip under the latency shadow,
zkVM cost per Ethereum block, the 2028 prover tiers), the rental-tax reward rule, the work-stake bond,
the burn-funded bounty, the GPU-rental settlement comparison, the 30-s beacon and the proving-income
ceiling. Every input carries a label (measured, cited, designed, approximate) and its source in the
`INPUTS` table at the top of the script and of the output. Run on the Mac without the lock: it is
arithmetic, not a measurement (about 50 ms).

View file

@ -0,0 +1,251 @@
#!/usr/bin/env python3
"""Horizon lane 7 (frontier): the arithmetic behind docs/analysis/horizon/frontier.md.
Every input is labelled in the INPUTS dict: measured (a repo bench entry), cited (a URL
or paper named in frontier.md) or approximate (from memory, or an estimate). Nothing
here is a prediction of a coin price. Run: python3 frontier_model.py (pure Python 3,
no numpy; about 50 ms).
Sections (one function each, printed as markdown tables):
1. predictions_2030 VRAM, GB per dollar, random-read ceilings (GDDR7, HBM3, HBM4),
the f = 1 chip on HBM4 with and without the latency shadow,
zkVM cost per Ethereum block, prover tiers in 2028
2. rental_tax the reward rule: pay per block falls when hash arrives faster
than the 30-day weight can follow (idea 1)
3. work_stake vote weight as the external-job bond (idea 2)
4. burn_bounty audits paid from the base-fee burn by 60% signal (idea 6)
5. rental_settlement Igneum as the settlement layer for GPU rental (idea 10)
6. beacon a 30-s VDF beacon against drand quicknet (idea 12)
7. proving_income could proving others' chains be the main income by 2030 (idea 14)
"""
import math
INPUTS = {
# label, value, source
"rtx5090_mhs_v3": ("measured", 136.1, "docs/analysis/chip-model-v3.md 5.1 (bench-log Counter ASIC 2.0)"),
"rtx5090_w_v3": ("measured", 326.0, "chip-model-v3.md 5.1 (peak with prover on; 290 W in the app, 350 W bench)"),
"rtx5090_uj_per_hash": ("measured", 2.40, "chip-model-v3.md 5.1"),
"rtx5090_reads_per_s": ("measured", 17.5e9, "chip-model-v3.md 5.1 (CUDA wall)"),
"gddr7_ceiling_reads": ("approximate", 21.3e9, "chip-model-v3.md 5.3 activate-bound ceiling, 16 devices"),
"hbm3_ceiling_reads_per_stack": ("approximate", 10.7e9, "chip-model-v3.md 5.3 (8 activates per 12 ns per channel, 16 channels)"),
"hbm4_channels_per_stack": ("cited", 32, "JEDEC JESD270-4 via allaboutcircuits.com: channels 16 to 32, each with two pseudo-channels"),
"hbm3_channels_per_stack": ("cited", 16, "chip-model-v3.md 5.1 (Synopsys HBM3 glossary)"),
"gddr7_read_nj": ("approximate", 2.0, "chip-model-v3.md 5.3"),
"hbm3_read_nj": ("approximate", 1.2, "chip-model-v3.md 5.3"),
"hbm4_read_nj": ("approximate", 1.0, "estimate: 15 percent under HBM3 on a shorter interposer path; unsourced"),
"hbm3_stack_usd": ("approximate", 200.0, "chip-model-v3.md 5.1 (siliconanalysts, 24 GB factory gate)"),
"hbm4_stack_usd": ("approximate", 550.0, "siliconanalysts.com/data/hbm-pricing, 36 GB 12-high, October 2026"),
"gddr7_2gb_usd": ("cited", 20.0, "TrendForce 24 Sep 2026 via chip-model-v3.md 5.1"),
"gddr6_8gb_usd_2023": ("cited", 27.0, "Tom's Hardware 'GDDR6 VRAM prices plummet' (2023)"),
"gddr6_usd_per_gb_2025": ("cited", 2.50, "TechSpot 'AI is eating all the DRAM' (2026)"),
"gddr6_usd_per_gb_2026": ("cited", 3.30, "TechSpot, same article"),
"rtx5090_msrp": ("cited", 1999.0, "chip-model-v3.md 5.1"),
"rtx5090_street_2026": ("cited", 3695.0, "localaimaster.com GPU price-per-GB table, 2026"),
"shadow_N": ("measured", 100000, "class v4 candidate mx8+sh256x27, counter-asic-3-status.md section 4"),
"shadow_chip_core_w_at_k1": ("approximate", 150.0, "latency-shadow-2026-10-06.md via counter-asic-3-status.md section 4 (14,000-lane array, N5)"),
"card_w_at_N100k": ("approximate", 401.0, "chip-model-v3.md 5.7 (linear toward 575 W TGP)"),
"card_uj_at_N100k": ("approximate", 2.95, "chip-model-v3.md 5.7"),
"ethproofs_usd_per_block_jan2025": ("cited", 1.69, "HackMD 'Ethproofs 2025 review' (willcorcoran), secondary"),
"ethproofs_usd_per_block_sep2026": ("cited", 0.005, "ethproofs via the Sept 2026 comparative analysis (GitHub Ricosworks1), secondary; 'under 4 cents' by Dec 2025 per HackMD"),
"eth_blocks_per_day": ("cited", 7200, "12-s slots"),
"emission_ign_per_block": ("designed", 31.688, "sim/economy assumptions, spec 2.5, 1 block/s, pre-halving"),
"rented_usd_per_gh_hour": ("measured", 11.69, "docs/bench-log.md line 2582, Rental cost of hash, 6 October 2026: 1,748 MH/s for USD 20.44 per hour on RunPod community pods, USD 0.0117 per MH/s-hour; the live devnet 1.16 GH/s"),
"base_fee_full_block_ign": ("designed", 3.0, "spec 5.11: a full block burns 3 IGN at the floor"),
"base_fee_full_day_ign": ("designed", 259200.0, "spec 5.11"),
"vast_take": ("approximate", 0.15, "secondary comparisons (spheron, miningboard) say about 15 percent; Vast's own June 2024 update says the host fee was removed and replaced by a surcharge it does not publish"),
"runpod_take": ("approximate", 0.07, "secondary (miningboard): hosts keep 93 percent"),
"rtx5090_rent_usd_h": ("cited", 0.44, "Vast.ai on-demand, getdeploying.com 6 Oct 2026; RunPod secure cloud 0.99"),
"rtx4090_rent_usd_h": ("cited", 0.31, "Vast.ai low, gpuperhour/runcrate 2026"),
"plain_transfer_ign": ("designed", 0.0051, "spec 5.11"),
"vdf_10min_s": ("measured", 600, "spec 04 epoch VDF; 4.47 ms verify, 516 B proof"),
"drand_quicknet_period_s": ("cited", 3, "docs.drand.love quicknet, unchained"),
"checkpoint_period_s": ("designed", 30, "spec 03"),
}
def v(k):
return INPUTS[k][1]
def row(*cells):
print("| " + " | ".join(str(c) for c in cells) + " |")
def header(*cells):
row(*cells)
print("|" + "---|" * len(cells))
# ---------------------------------------------------------------- 1. predictions
def predictions_2030():
print("\n## 1. Predictions to 2030\n")
print("### 1.1 Flagship consumer VRAM (approximate: generations from memory, the 5090 cited)\n")
gens = [(2016, "GTX 1080", 8), (2018, "RTX 2080 Ti", 11), (2020, "RTX 3090", 24), (2022, "RTX 4090", 24), (2025, "RTX 5090", 32)]
header("Year", "Card", "GB", "Years since 2016", "GB growth per year (compound)")
for y, c, gb in gens:
n = y - 2016
g = (gb / 8) ** (1 / n) if n else float("nan")
row(y, c, gb, n, f"{g:.3f}" if n else "")
cagr = (32 / 8) ** (1 / 9)
print(f"\nCompound growth 2016 to 2025: {cagr:.3f} per year (4x in 9 years). Extrapolated: 2028 {32*cagr**3:.0f} GB, 2030 {32*cagr**5:.0f} GB (approximate).")
print("Module arithmetic: a 512-bit board is 16 devices; 2 GB devices give 32 GB, 3 GB devices 48 GB (Micron ends 2 GB GDDR7, TrendForce Sep 2026), 4 GB devices 64 GB. So the 2028 flagship is 48 GB if the RTX 60 series (Rubin GR20x, rumoured 2028, kopite7kimi via videocardz) ships 3 GB GDDR7, and 64 GB is the 2030 shape.\n")
print("### 1.2 Memory dollars per GB (consumer GDDR)\n")
header("Point", "USD per GB", "Source label")
row("2023 GDDR6", f"{v('gddr6_8gb_usd_2023')/8:.2f}", "cited")
row("2025 GDDR6", f"{v('gddr6_usd_per_gb_2025'):.2f}", "cited")
row("2026 GDDR6", f"{v('gddr6_usd_per_gb_2026'):.2f}", "cited")
row("Sep 2026 GDDR7 2 GB device", f"{v('gddr7_2gb_usd')/2:.2f}", "cited")
row("Sep 2026 GDDR7 3 GB device", f"{65/3:.2f}", "cited (60 to 70 USD per device)")
print("\nDirection: GB per dollar fell in 2026 for the first time in a decade (DRAM shortage, forecast tight through 2027). The chip model's f = 1 chip pays the same device price the GPU does, so the ratio of chip memory cost to GPU memory cost is unchanged; what changes is the share of each bill of materials that is memory.\n")
print("### 1.3 Random-read ceilings per memory system (reads per second; the lottery is latency-bound, so this is the number that matters, not GB/s)\n")
hbm4_ceiling = v("hbm3_ceiling_reads_per_stack") * v("hbm4_channels_per_stack") / v("hbm3_channels_per_stack")
header("Memory system", "Reads/s ceiling", "Scaling rule", "Label")
row("GDDR7, 16 devices, 512-bit (RTX 5090 board)", f"{v('gddr7_ceiling_reads')/1e9:.1f} G", "activates per tFAW per channel x 64 channels; pin rate irrelevant (28 Gbps = 48 Gbps)", "approximate")
row("RTX 5090 measured", f"{v('rtx5090_reads_per_s')/1e9:.1f} G", "82 percent of the ceiling", "measured")
row("HBM3 or HBM3E, one stack", f"{v('hbm3_ceiling_reads_per_stack')/1e9:.1f} G", "16 channels", "approximate")
row("HBM4, one stack", f"{hbm4_ceiling/1e9:.1f} G", "32 channels (JEDEC): 2x the activate parallelism per stack if tFAW per channel holds", "approximate, derived")
row("A 48 GB GDDR7 board (16 x 3 GB)", f"{v('gddr7_ceiling_reads')/1e9:.1f} G", "same channel count; capacity does not add channels", "approximate")
print("\n### 1.4 The stored-dataset (f = 1) chip in 2028 on HBM4, bare and under the latency shadow\n")
def chip(ceiling, read_nj, static_w, ctrl_w, mem_usd, extra_w=0.0, label=""):
rate = ceiling / 128
p = rate * 128 * read_nj * 1e-9 + static_w + ctrl_w + extra_w
uj = p / rate * 1e6
return rate, p, uj
header("Chip", "MH/s per chip", "W", "uJ per hash", "Gain per joule vs 5090 at 2.40 uJ", "Gain per joule vs 5090 under the shadow (2.95 uJ at N = 100k)", "Memory USD")
for name, ceil, nj, st, ct, usd in [
("GDDR7 f=1 (today's model row)", v("gddr7_ceiling_reads"), v("gddr7_read_nj"), 20, 15, 320),
("HBM3 one stack f=1", v("hbm3_ceiling_reads_per_stack"), v("hbm3_read_nj"), 4, 10, 400),
("HBM4 one stack f=1 (2028)", hbm4_ceiling, v("hbm4_read_nj"), 5, 10, v("hbm4_stack_usd") + 200),
]:
r, p, uj = chip(ceil, nj, st, ct, usd)
r2, p2, uj2 = chip(ceil, nj, st, ct, usd, extra_w=v("shadow_chip_core_w_at_k1"))
row(name, f"{r/1e6:.0f}", f"{p:.0f} (bare) / {p2:.0f} (with a 150 W shadow core at k = 1)", f"{uj:.2f} / {uj2:.2f}", f"{2.40/uj:.1f}x", f"{v('card_uj_at_N100k')/uj2:.1f}x", f"{usd:.0f}")
print("\nReading: HBM4's doubled channel count doubles the chip's rate per stack at about the same watts, so the bare per-joule edge rises from about 7x to about 11x, and under the class v4 shadow (N = 100,000, k = 1) from about 2.3x to about 2.7x (approximate; every chip figure is arithmetic). The lever that answers it is N: the chip's shadow core scales with N while the card's spare ALU budget is 330,000 ops per hash on the 5090. Verifier cost is N x 32 ops per warp: 3.2 M ops at N = 100k (about 1 ms on one M5 Max core, measured class), 10 M at N = 330k (about 3 ms), inside the 10 ms gate; the 2019-class core is unmeasured.\n")
# N needed to hold 2x on HBM4
for N in (100000, 200000, 330000):
card_w = 326 + (575 - 326) * N / 330000
card_uj = card_w / 136.1
core_w = v("shadow_chip_core_w_at_k1") * N / 100000
r, p, uj = chip(hbm4_ceiling, v("hbm4_read_nj"), 5, 10, 0, extra_w=core_w)
print(f"- N = {N:,}: card {card_w:.0f} W, {card_uj:.2f} uJ; HBM4 chip {p:.0f} W, {uj:.2f} uJ; gain {card_uj/uj:.2f}x at k = 1")
print("So the schedule for N should be written into the era draw at genesis (a doubling per era is the candidate), because the memory generation it answers arrives every two to three years and the verifier has 10x of headroom.\n")
print("### 1.5 Chip fabrication cost curve (mask sets, cited; project totals approximate)\n")
header("Node", "Mask set USD", "Source", "What it means for Igneum")
row("28 nm", "1 to 3 M", "TubeTime (3 M), VBsemi (over 1 M)", "The f = 1 memory-controller chip lives here: no mixer on the die. Project 5 to 30 M (history 2.5)")
row("7 nm", "10 to 15 M", "VBsemi, HN thread", "The f = 0 recompute chip with the 256 MiB cache on die. Project 50 to 75 M")
row("5 nm", "6.5 M (2026 data) to 30 M (2023 estimate)", "siliconanalysts, HN", "The shadow core at N5 (30 mm^2 at N = 100k) pushes the f = 1 chip from a 28 nm project to a 5 nm one, or to a reticle-class 28 nm die")
row("3 nm", "15 to 22 M (Q4 2025), up to 40 M (older estimate)", "siliconanalysts, semianalysis", "Not relevant to a chip whose cost is memory")
print("\nThe curve is falling at a given node (5 nm masks quoted at 30 M in 2023 and 6.5 M in 2026) while the leading node's cost rises. Consequence: the shadow lever's economic teeth (forcing an advanced-node core onto a memory-controller chip) weaken by about 4x in mask cost over three years; the rate and joule arithmetic above, not the fab bill, is what holds in 2030.\n")
print("### 1.6 zkVM proving cost per Ethereum block (public tracker, secondary sources, approximate)\n")
a, b = v("ethproofs_usd_per_block_jan2025"), v("ethproofs_usd_per_block_sep2026")
months = 20
per_year = (b / a) ** (12 / months)
header("Point", "USD per Ethereum block proof", "Hardware named")
row("Jan 2025", f"{a:.2f}", "about 160 RTX 4090s for 90 percent real-time (Succinct, May 2025 estimate)")
row("Dec 2025", "under 0.04", "16 x RTX 5090 (SP1 Hypercube 99.7 percent under 12 s); Pico Prism 16 GPUs")
row("Sep 2026", f"{b:.3f}", "ZisK 4 x RTX 5090 p99 9.62 s (Aug 2026); Cysic Venus 7.4 s on 24 GPUs (Apr 2026)")
print(f"\nThe 20-month ratio is {a/b:.0f}x, which is {1/per_year:.0f}x per year. That rate cannot hold (it is software catching up with hardware), so the table below uses 1.5x, 3x and 10x per year from today's measured Igneum shard times.\n")
print("### 1.7 Which card tier proves a v1 shard in under 10 s in 2028 (measured 6 Oct 2026 times, prover-tiers-real-cards.md, divided by two years of software gain)\n")
tiers = [("RTX 3060 12 GB", 37.5, 14.4), ("RTX 4060 8 GB (core-only beside, alone)", 22.1, 18.4), ("RTX 4070 12 GB", 27.3, 12.1), ("RTX 4060 Ti 16 GB", 34.6, 11.6), ("RTX 3080 10 GB", 25.6, 7.1), ("RTX 3090 24 GB", 19.9, 14.9), ("RTX 4090 24 GB", 26.1, 6.3), ("RTX 5070 12 GB", 37.2, 4.8), ("RTX 5090 32 GB", 10.7, 6.3)]
header("Card", "Beside the miner today, s", "Alone today, s", "2028 at 1.5x/yr (beside / alone)", "2028 at 3x/yr", "2028 at 10x/yr", "Under 10 s beside the miner in 2028?")
for name, beside, alone in tiers:
cells = []
for g in (1.5, 3, 10):
cells.append(f"{beside/g**2:.1f} / {alone/g**2:.1f}")
verdict = "yes at 3x or more" if beside / 9 < 10 else "no"
if beside / 2.25 < 10:
verdict = "yes even at 1.5x"
row(name, beside, alone, *cells, verdict)
print("\nConsequence per tier: at the floor rate (1.5x a year) only the 32 GB card mines and proves inside 10 s in 2028, so a 10-s proof lag at launch is a 24 GB and 32 GB story; at 3x a year every card from the 3060 up does it, and the 8 GB card alone proves in 2 s. The block-proof target (under 10 s behind the tip) should be written as a function of the measured fleet median, re-read each era, not as a date.\n")
# ---------------------------------------------------------------- 2. rental tax
def rental_tax():
print("\n## 2. Idea 1: the reward rule that prices rented hash out\n")
print("Rule modelled: m = clamp(W30 / H_now, m_min, 1) where W30 is the 30-day work-weighted hash (the finality window's blue blocks per DAA second, which every node already computes for W2) and H_now the DAA-window estimate. The block subsidy paid to the producer is m x the schedule; the remainder (1 - m) x subsidy goes to the proving pool escrow of that block (not to incumbents, to avoid the cartel transfer; see the Monero attack in the text). Fees are untouched.\n")
E = v("emission_ign_per_block") * 3600 # IGN per hour
rent = v("rented_usd_per_gh_hour")
header("Network hash", "Attacker adds", "H_now / W30", "m", "Attacker's share of blocks", "Attacker IGN per hour, no rule", "With rule", "Rent USD per hour", "Break-even IGN price, no rule", "With rule", "To the pool per hour, IGN")
for net in (1, 10, 100, 1000):
for mult in (1.0, 2.0, 5.0):
add = net * mult
ratio = (net + add) / net
m = max(0.25, min(1.0, 1 / ratio))
share = add / (net + add)
no_rule = E * share
with_rule = no_rule * m
cost = add * rent
pool = E * (1 - m)
row(f"{net} GH/s", f"{add:.0f} GH/s", f"{ratio:.1f}", f"{m:.2f}", f"{share:.2f}", f"{no_rule:,.0f}", f"{with_rule:,.0f}", f"{cost:,.0f}", f"{cost/no_rule:.5f}", f"{cost/with_rule:.5f}", f"{pool:,.0f}")
print("\nHonest-growth cost: a listing that doubles honest hash overnight halves every miner's subsidy per block (not per hash: difficulty halves the per-hash rate anyway; the rule halves it again) until W30 catches up, which is the 30-day ramp of ledger C7 (0.9x on day 28 to 31). With the floor m_min = 0.25 the worst case is a 4x cut, and the money is not lost to the chain: it reaches the provers of the same block, who are the same population. Per tier: a home miner's monthly income during a doubling month falls 50 percent under the rule against 50 percent already from difficulty (so 25 percent of the pre-event figure); a pool user sees the same through PPLNS; a prover with weight gains the diverted share. The gate: in the economy simulator (sim/economy/sim.py scenario f, a pool with the network's hash arriving on day 10) the incumbents' income under the rule must stay above the no-rule row for the 30 days and the newcomers' under; and in the fast-time harness a timestamp-manipulated H_now (headers inside the 132-s tolerance) must move m by under 2 percent.\n")
# ---------------------------------------------------------------- 3. work stake
def work_stake():
print("\n## 3. Idea 2: work-stake, vote weight as the external-job bond\n")
print("A key that claims an external job and delivers late or wrong loses s of its 30-day weight for 30 days (as equivocation strips 100 percent, spec 3.6). Weight is blue blocks; it cannot be bought, only mined. Arithmetic: what a stripped key forgoes.\n")
E = v("emission_ign_per_block")
header("Key's hash share", "Blocks per 30 days at 1 bps", "Weight share", "Shard sortition income per 30 days (20 percent pool, pro rata), IGN", "Stripped at s = 25 percent: lost pool income over 30 days, IGN", "Stripped at s = 100 percent", "IGN bond that would match (design 4.6: maxPgas x f_p x 1.5 for a 1 B-cycle job at the floor)")
for share in (0.0001, 0.001, 0.01, 0.1):
blocks = share * 86400 * 30
pool_30d = 0.2 * E * 86400 * 30 * share
row(f"{share*100:.2f} percent", f"{blocks:,.0f}", f"{share*100:.2f} percent", f"{pool_30d:,.0f}", f"{pool_30d*0.25:,.0f}", f"{pool_30d:,.0f}", "0.0015 IGN")
print("\nReading: for every key above dust the 30-day pool income at risk is many orders above the designed IGN bond for one job, so weight is a far larger bond than coins, and it is a bond nobody can buy on a market. It also strips the key's vote for 30 days, which is the sentence the finality rule already hands out for equivocation. The cost: a false positive (a partition that makes an honest proof late) strips an honest voter; so the rule must use DAA time, a long deadline (the 120-s claim timeout of P9 decision, or longer), and a one-strike grace per 30 days. Per tier: a solo 8 GB miner below dust has no weight and so cannot take external jobs at all under this rule (it can still prove shards, which carry no bond); a pool user's jobs are the pool's and the pool's weight is at risk, which is what a pool operator wants priced. Gate: on the phase 4 devnet, 1,000 jobs with a 10 percent injected late rate: every injected fault stripped, zero honest keys stripped across a 60-s partition.\n")
# ---------------------------------------------------------------- 4. burn bounty
def burn_bounty():
print("\n## 4. Idea 6: audits paid from the burn, by 60 percent signal, no standing address\n")
full_day = v("base_fee_full_day_ign")
header("Chain traffic (fraction of full blocks)", "Base fee burned per day, IGN", "7-day redirect, IGN", "30-day redirect, IGN", "USD at 0.02 (7 d / 30 d)", "USD at 0.10 (7 d / 30 d)")
for frac in (0.01, 0.1, 0.5, 1.0):
d = full_day * frac
row(f"{frac:.2f}", f"{d:,.0f}", f"{7*d:,.0f}", f"{30*d:,.0f}", f"{7*d*0.02:,.0f} / {30*d*0.02:,.0f}", f"{7*d*0.10:,.0f} / {30*d*0.10:,.0f}")
print("\nReading: at launch traffic (1 to 10 percent of full blocks) a 30-day redirect is USD 1,600 to 16,000 at 0.02 per IGN, under one Code4rena contest (base pricing from USD 6,500 before the 2025 zero-fee change; Immunefi's standard pays 10 percent of funds at risk). At half-full blocks it reaches a serious bounty (USD 78,000 for 30 days at 0.02). So the burn can fund audits only once the chain is used; before that the only money is the client's 1 percent dev fee and the founders' mined coins (litepaper: grants from founders' mined coins). The text gives the attack: this is a dev fund with a 60 percent veto and a per-event payee, which is exactly the switch spec 5.5 removed.\n")
# ---------------------------------------------------------------- 5. rental settlement
def rental_settlement():
print("\n## 5. Idea 10: Igneum as the settlement layer for GPU rental\n")
header("Card", "Vast.ai on-demand USD/h (cited)", "Platform take modelled", "Host loses USD per card-year", "Igneum settlement cost per rental (2 transfers at the floor), IGN", "At 0.02 and 0.10 USD per IGN", "Hours of rental to pay 1 USD of chain fees at 0.02")
for name, price in (("RTX 5090", v("rtx5090_rent_usd_h")), ("RTX 4090", v("rtx4090_rent_usd_h"))):
for take_name, take in (("Vast about 15 percent", v("vast_take")), ("RunPod about 7 percent", v("runpod_take"))):
lost = price * take * 8766
fee = 2 * v("plain_transfer_ign")
row(name, f"{price:.2f}", take_name, f"{lost:,.0f}", f"{fee:.4f}", f"{fee*0.02:.5f} / {fee*0.10:.4f}", f"{1/(fee*0.02)/ (1/1):,.0f} rentals")
print("\nReading: the chain's fee is three to five orders of magnitude under the platform take. The platform's take pays for what the chain cannot do: matching, trust, dispute, image hosting, and the verification of delivered work. The honest problem is the last one: a rented hour of general compute is unverifiable, so an on-chain escrow without a verifier is a trust-me payment with lower fees. What is verifiable on this chain today: ZK proving jobs (the precompile); what is verifiable with sampling: deterministic recompute checked on a sampled fraction (SPEX, arXiv 2503.18899; Render's result-quorum, approximate); what needs hardware the fleet does not have: TEE attestation (NVIDIA confidential computing is H100 and H200 class, phala.com; no consumer card has it).\n")
# ---------------------------------------------------------------- 6. beacon
def beacon():
print("\n## 6. Idea 12: a randomness beacon from the checkpoint VDF\n")
header("Beacon", "Period", "Latency to a value", "Unbiasability argument", "Verify cost", "Who runs it")
row("drand quicknet (League of Entropy)", f"{v('drand_quicknet_period_s')} s", "about 3 s", "threshold BLS over H(round) with a 2/3 threshold of about 20 named organisations; unbiasable while under 1/3 collude; unchained", "one BLS verify", "a league, trusted set")
row("Igneum epoch seed today", "3,600 s", f"{v('vdf_10min_s')} s (10-min VDF)", "a certified checkpoint 10 minutes before use; the VDF makes the last block producer's choice useless because it cannot see the output in time", "4.47 ms (516 B)", "nobody: any node evaluates")
row("Proposed: a per-checkpoint VDF beacon", f"{v('checkpoint_period_s')} s", "30 to 60 s (a 30-s VDF of each certified checkpoint hash)", "the checkpoint is locked by 2/3 of 30-day weight before the VDF starts, so no single party chooses the input; a last-block grind costs a block's subsidy per try and buys one bit of influence only if the attacker can evaluate the VDF faster than the chain, which is the class-group ASIC question (Chia timelords)", "4.47 ms per value, 2,880 values a day", "nobody: the epoch pipeline already exists")
print("\nReading: the chain already produces an unbiasable value once an hour with a 10-minute delay. A 30-s beacon is the same code at 120x the cadence, and its honest limit is the one Chia carries: the fastest class-group squarer sets the floor on 'delay', so a timelord-class ASIC owner can learn the value earlier than everyone else (Chia docs: timelords are software or ASIC). That earlier knowledge is a front-running edge, not a bias. The product: PREVRANDAO per block already comes from this pipeline (spec 7.1); the beacon makes it a 30-s value usable off-chain (lotteries, shuffles, timelock encryption as drand does).\n")
# ---------------------------------------------------------------- 7. proving income
def proving_income():
print("\n## 7. Idea 14: proving other chains as the main income by 2030\n")
eth_day = v("ethproofs_usd_per_block_sep2026") * v("eth_blocks_per_day")
emission_day_ign = v("emission_ign_per_block") * 86400
header("Income line", "USD per day", "Basis")
row("Proving every Ethereum L1 block at the Sep 2026 tracker cost", f"{eth_day:,.0f}", "0.005 USD x 7,200 blocks; the price a buyer pays is above cost, call it 10x: 360")
row("The same at the Dec 2025 cost (under 0.04)", f"{0.04*7200:,.0f}", "secondary")
for p in (0.005, 0.02, 0.10):
row(f"Igneum year-1 emission at {p} USD per IGN", f"{emission_day_ign*p:,.0f}", "31.688 IGN per block x 86,400")
row("Rollup proving spend, all rollups (customer brief)", f"{3e6/365:,.0f} to {10e6/365:,.0f}", "low millions a year, approximate")
row("Boundless trailing day in the explorer (4 Oct 2026)", f"{8.4*0.21:,.0f}", "8.4 T cycles at a 0.21 USD per B-cycle median, developer-adoption.md 2b, approximate")
print("\nReading: the whole public proving market is three to four orders of magnitude under year-1 emission at any price input. For proving to be the main income by 2030, demand must grow about 1,000x while cost per proof keeps falling 3x to 30x a year, which pushes dollars per proof down as fast as volume rises. The arithmetic says never by 2030 for 'main income'; it says 'yes' for 'a second income that keeps cards on after the subsidy fades' (spec 5.10.2), which is the design's own claim.\n")
if __name__ == "__main__":
print("# frontier_model.py output (Horizon lane 7), run on " + __import__("datetime").date.today().isoformat())
print("\nInputs and labels:\n")
header("Key", "Label", "Value", "Source")
for k, (lab, val, src) in INPUTS.items():
row(k, lab, val, src)
predictions_2030()
rental_tax()
work_stake()
burn_bounty()
rental_settlement()
beacon()
proving_income()

220
sim/horizon/frontier/out.md Normal file
View file

@ -0,0 +1,220 @@
# frontier_model.py output (Horizon lane 7), run on 2026-10-06
Inputs and labels:
| Key | Label | Value | Source |
|---|---|---|---|
| rtx5090_mhs_v3 | measured | 136.1 | docs/analysis/chip-model-v3.md 5.1 (bench-log Counter ASIC 2.0) |
| rtx5090_w_v3 | measured | 326.0 | chip-model-v3.md 5.1 (peak with prover on; 290 W in the app, 350 W bench) |
| rtx5090_uj_per_hash | measured | 2.4 | chip-model-v3.md 5.1 |
| rtx5090_reads_per_s | measured | 17500000000.0 | chip-model-v3.md 5.1 (CUDA wall) |
| gddr7_ceiling_reads | approximate | 21300000000.0 | chip-model-v3.md 5.3 activate-bound ceiling, 16 devices |
| hbm3_ceiling_reads_per_stack | approximate | 10700000000.0 | chip-model-v3.md 5.3 (8 activates per 12 ns per channel, 16 channels) |
| hbm4_channels_per_stack | cited | 32 | JEDEC JESD270-4 via allaboutcircuits.com: channels 16 to 32, each with two pseudo-channels |
| hbm3_channels_per_stack | cited | 16 | chip-model-v3.md 5.1 (Synopsys HBM3 glossary) |
| gddr7_read_nj | approximate | 2.0 | chip-model-v3.md 5.3 |
| hbm3_read_nj | approximate | 1.2 | chip-model-v3.md 5.3 |
| hbm4_read_nj | approximate | 1.0 | estimate: 15 percent under HBM3 on a shorter interposer path; unsourced |
| hbm3_stack_usd | approximate | 200.0 | chip-model-v3.md 5.1 (siliconanalysts, 24 GB factory gate) |
| hbm4_stack_usd | approximate | 550.0 | siliconanalysts.com/data/hbm-pricing, 36 GB 12-high, October 2026 |
| gddr7_2gb_usd | cited | 20.0 | TrendForce 24 Sep 2026 via chip-model-v3.md 5.1 |
| gddr6_8gb_usd_2023 | cited | 27.0 | Tom's Hardware 'GDDR6 VRAM prices plummet' (2023) |
| gddr6_usd_per_gb_2025 | cited | 2.5 | TechSpot 'AI is eating all the DRAM' (2026) |
| gddr6_usd_per_gb_2026 | cited | 3.3 | TechSpot, same article |
| rtx5090_msrp | cited | 1999.0 | chip-model-v3.md 5.1 |
| rtx5090_street_2026 | cited | 3695.0 | localaimaster.com GPU price-per-GB table, 2026 |
| shadow_N | measured | 100000 | class v4 candidate mx8+sh256x27, counter-asic-3-status.md section 4 |
| shadow_chip_core_w_at_k1 | approximate | 150.0 | latency-shadow-2026-10-06.md via counter-asic-3-status.md section 4 (14,000-lane array, N5) |
| card_w_at_N100k | approximate | 401.0 | chip-model-v3.md 5.7 (linear toward 575 W TGP) |
| card_uj_at_N100k | approximate | 2.95 | chip-model-v3.md 5.7 |
| ethproofs_usd_per_block_jan2025 | cited | 1.69 | HackMD 'Ethproofs 2025 review' (willcorcoran), secondary |
| ethproofs_usd_per_block_sep2026 | cited | 0.005 | ethproofs via the Sept 2026 comparative analysis (GitHub Ricosworks1), secondary; 'under 4 cents' by Dec 2025 per HackMD |
| eth_blocks_per_day | cited | 7200 | 12-s slots |
| emission_ign_per_block | designed | 31.688 | sim/economy assumptions, spec 2.5, 1 block/s, pre-halving |
| rented_usd_per_gh_hour | measured | 11.69 | docs/bench-log.md line 2582, Rental cost of hash, 6 October 2026: 1,748 MH/s for USD 20.44 per hour on RunPod community pods, USD 0.0117 per MH/s-hour; the live devnet 1.16 GH/s |
| base_fee_full_block_ign | designed | 3.0 | spec 5.11: a full block burns 3 IGN at the floor |
| base_fee_full_day_ign | designed | 259200.0 | spec 5.11 |
| vast_take | approximate | 0.15 | secondary comparisons (spheron, miningboard) say about 15 percent; Vast's own June 2024 update says the host fee was removed and replaced by a surcharge it does not publish |
| runpod_take | approximate | 0.07 | secondary (miningboard): hosts keep 93 percent |
| rtx5090_rent_usd_h | cited | 0.44 | Vast.ai on-demand, getdeploying.com 6 Oct 2026; RunPod secure cloud 0.99 |
| rtx4090_rent_usd_h | cited | 0.31 | Vast.ai low, gpuperhour/runcrate 2026 |
| plain_transfer_ign | designed | 0.0051 | spec 5.11 |
| vdf_10min_s | measured | 600 | spec 04 epoch VDF; 4.47 ms verify, 516 B proof |
| drand_quicknet_period_s | cited | 3 | docs.drand.love quicknet, unchained |
| checkpoint_period_s | designed | 30 | spec 03 |
## 1. Predictions to 2030
### 1.1 Flagship consumer VRAM (approximate: generations from memory, the 5090 cited)
| Year | Card | GB | Years since 2016 | GB growth per year (compound) |
|---|---|---|---|---|
| 2016 | GTX 1080 | 8 | 0 | |
| 2018 | RTX 2080 Ti | 11 | 2 | 1.173 |
| 2020 | RTX 3090 | 24 | 4 | 1.316 |
| 2022 | RTX 4090 | 24 | 6 | 1.201 |
| 2025 | RTX 5090 | 32 | 9 | 1.167 |
Compound growth 2016 to 2025: 1.167 per year (4x in 9 years). Extrapolated: 2028 51 GB, 2030 69 GB (approximate).
Module arithmetic: a 512-bit board is 16 devices; 2 GB devices give 32 GB, 3 GB devices 48 GB (Micron ends 2 GB GDDR7, TrendForce Sep 2026), 4 GB devices 64 GB. So the 2028 flagship is 48 GB if the RTX 60 series (Rubin GR20x, rumoured 2028, kopite7kimi via videocardz) ships 3 GB GDDR7, and 64 GB is the 2030 shape.
### 1.2 Memory dollars per GB (consumer GDDR)
| Point | USD per GB | Source label |
|---|---|---|
| 2023 GDDR6 | 3.38 | cited |
| 2025 GDDR6 | 2.50 | cited |
| 2026 GDDR6 | 3.30 | cited |
| Sep 2026 GDDR7 2 GB device | 10.00 | cited |
| Sep 2026 GDDR7 3 GB device | 21.67 | cited (60 to 70 USD per device) |
Direction: GB per dollar fell in 2026 for the first time in a decade (DRAM shortage, forecast tight through 2027). The chip model's f = 1 chip pays the same device price the GPU does, so the ratio of chip memory cost to GPU memory cost is unchanged; what changes is the share of each bill of materials that is memory.
### 1.3 Random-read ceilings per memory system (reads per second; the lottery is latency-bound, so this is the number that matters, not GB/s)
| Memory system | Reads/s ceiling | Scaling rule | Label |
|---|---|---|---|
| GDDR7, 16 devices, 512-bit (RTX 5090 board) | 21.3 G | activates per tFAW per channel x 64 channels; pin rate irrelevant (28 Gbps = 48 Gbps) | approximate |
| RTX 5090 measured | 17.5 G | 82 percent of the ceiling | measured |
| HBM3 or HBM3E, one stack | 10.7 G | 16 channels | approximate |
| HBM4, one stack | 21.4 G | 32 channels (JEDEC): 2x the activate parallelism per stack if tFAW per channel holds | approximate, derived |
| A 48 GB GDDR7 board (16 x 3 GB) | 21.3 G | same channel count; capacity does not add channels | approximate |
### 1.4 The stored-dataset (f = 1) chip in 2028 on HBM4, bare and under the latency shadow
| Chip | MH/s per chip | W | uJ per hash | Gain per joule vs 5090 at 2.40 uJ | Gain per joule vs 5090 under the shadow (2.95 uJ at N = 100k) | Memory USD |
|---|---|---|---|---|---|---|
| GDDR7 f=1 (today's model row) | 166 | 78 (bare) / 228 (with a 150 W shadow core at k = 1) | 0.47 / 1.37 | 5.1x | 2.2x | 320 |
| HBM3 one stack f=1 | 84 | 27 (bare) / 177 (with a 150 W shadow core at k = 1) | 0.32 / 2.12 | 7.5x | 1.4x | 400 |
| HBM4 one stack f=1 (2028) | 167 | 36 (bare) / 186 (with a 150 W shadow core at k = 1) | 0.22 / 1.11 | 11.0x | 2.6x | 750 |
Reading: HBM4's doubled channel count doubles the chip's rate per stack at about the same watts, so the bare per-joule edge rises from about 7x to about 11x, and under the class v4 shadow (N = 100,000, k = 1) from about 2.3x to about 2.7x (approximate; every chip figure is arithmetic). The lever that answers it is N: the chip's shadow core scales with N while the card's spare ALU budget is 330,000 ops per hash on the 5090. Verifier cost is N x 32 ops per warp: 3.2 M ops at N = 100k (about 1 ms on one M5 Max core, measured class), 10 M at N = 330k (about 3 ms), inside the 10 ms gate; the 2019-class core is unmeasured.
- N = 100,000: card 401 W, 2.95 uJ; HBM4 chip 186 W, 1.11 uJ; gain 2.65x at k = 1
- N = 200,000: card 477 W, 3.50 uJ; HBM4 chip 336 W, 2.01 uJ; gain 1.74x at k = 1
- N = 330,000: card 575 W, 4.22 uJ; HBM4 chip 531 W, 3.18 uJ; gain 1.33x at k = 1
So the schedule for N should be written into the era draw at genesis (a doubling per era is the candidate), because the memory generation it answers arrives every two to three years and the verifier has 10x of headroom.
### 1.5 Chip fabrication cost curve (mask sets, cited; project totals approximate)
| Node | Mask set USD | Source | What it means for Igneum |
|---|---|---|---|
| 28 nm | 1 to 3 M | TubeTime (3 M), VBsemi (over 1 M) | The f = 1 memory-controller chip lives here: no mixer on the die. Project 5 to 30 M (history 2.5) |
| 7 nm | 10 to 15 M | VBsemi, HN thread | The f = 0 recompute chip with the 256 MiB cache on die. Project 50 to 75 M |
| 5 nm | 6.5 M (2026 data) to 30 M (2023 estimate) | siliconanalysts, HN | The shadow core at N5 (30 mm^2 at N = 100k) pushes the f = 1 chip from a 28 nm project to a 5 nm one, or to a reticle-class 28 nm die |
| 3 nm | 15 to 22 M (Q4 2025), up to 40 M (older estimate) | siliconanalysts, semianalysis | Not relevant to a chip whose cost is memory |
The curve is falling at a given node (5 nm masks quoted at 30 M in 2023 and 6.5 M in 2026) while the leading node's cost rises. Consequence: the shadow lever's economic teeth (forcing an advanced-node core onto a memory-controller chip) weaken by about 4x in mask cost over three years; the rate and joule arithmetic above, not the fab bill, is what holds in 2030.
### 1.6 zkVM proving cost per Ethereum block (public tracker, secondary sources, approximate)
| Point | USD per Ethereum block proof | Hardware named |
|---|---|---|
| Jan 2025 | 1.69 | about 160 RTX 4090s for 90 percent real-time (Succinct, May 2025 estimate) |
| Dec 2025 | under 0.04 | 16 x RTX 5090 (SP1 Hypercube 99.7 percent under 12 s); Pico Prism 16 GPUs |
| Sep 2026 | 0.005 | ZisK 4 x RTX 5090 p99 9.62 s (Aug 2026); Cysic Venus 7.4 s on 24 GPUs (Apr 2026) |
The 20-month ratio is 338x, which is 33x per year. That rate cannot hold (it is software catching up with hardware), so the table below uses 1.5x, 3x and 10x per year from today's measured Igneum shard times.
### 1.7 Which card tier proves a v1 shard in under 10 s in 2028 (measured 6 Oct 2026 times, prover-tiers-real-cards.md, divided by two years of software gain)
| Card | Beside the miner today, s | Alone today, s | 2028 at 1.5x/yr (beside / alone) | 2028 at 3x/yr | 2028 at 10x/yr | Under 10 s beside the miner in 2028? |
|---|---|---|---|---|---|---|
| RTX 3060 12 GB | 37.5 | 14.4 | 16.7 / 6.4 | 4.2 / 1.6 | 0.4 / 0.1 | yes at 3x or more |
| RTX 4060 8 GB (core-only beside, alone) | 22.1 | 18.4 | 9.8 / 8.2 | 2.5 / 2.0 | 0.2 / 0.2 | yes even at 1.5x |
| RTX 4070 12 GB | 27.3 | 12.1 | 12.1 / 5.4 | 3.0 / 1.3 | 0.3 / 0.1 | yes at 3x or more |
| RTX 4060 Ti 16 GB | 34.6 | 11.6 | 15.4 / 5.2 | 3.8 / 1.3 | 0.3 / 0.1 | yes at 3x or more |
| RTX 3080 10 GB | 25.6 | 7.1 | 11.4 / 3.2 | 2.8 / 0.8 | 0.3 / 0.1 | yes at 3x or more |
| RTX 3090 24 GB | 19.9 | 14.9 | 8.8 / 6.6 | 2.2 / 1.7 | 0.2 / 0.1 | yes even at 1.5x |
| RTX 4090 24 GB | 26.1 | 6.3 | 11.6 / 2.8 | 2.9 / 0.7 | 0.3 / 0.1 | yes at 3x or more |
| RTX 5070 12 GB | 37.2 | 4.8 | 16.5 / 2.1 | 4.1 / 0.5 | 0.4 / 0.0 | yes at 3x or more |
| RTX 5090 32 GB | 10.7 | 6.3 | 4.8 / 2.8 | 1.2 / 0.7 | 0.1 / 0.1 | yes even at 1.5x |
Consequence per tier: at the floor rate (1.5x a year) only the 32 GB card mines and proves inside 10 s in 2028, so a 10-s proof lag at launch is a 24 GB and 32 GB story; at 3x a year every card from the 3060 up does it, and the 8 GB card alone proves in 2 s. The block-proof target (under 10 s behind the tip) should be written as a function of the measured fleet median, re-read each era, not as a date.
## 2. Idea 1: the reward rule that prices rented hash out
Rule modelled: m = clamp(W30 / H_now, m_min, 1) where W30 is the 30-day work-weighted hash (the finality window's blue blocks per DAA second, which every node already computes for W2) and H_now the DAA-window estimate. The block subsidy paid to the producer is m x the schedule; the remainder (1 - m) x subsidy goes to the proving pool escrow of that block (not to incumbents, to avoid the cartel transfer; see the Monero attack in the text). Fees are untouched.
| Network hash | Attacker adds | H_now / W30 | m | Attacker's share of blocks | Attacker IGN per hour, no rule | With rule | Rent USD per hour | Break-even IGN price, no rule | With rule | To the pool per hour, IGN |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 GH/s | 1 GH/s | 2.0 | 0.50 | 0.50 | 57,038 | 28,519 | 12 | 0.00020 | 0.00041 | 57,038 |
| 1 GH/s | 2 GH/s | 3.0 | 0.33 | 0.67 | 76,051 | 25,350 | 23 | 0.00031 | 0.00092 | 76,051 |
| 1 GH/s | 5 GH/s | 6.0 | 0.25 | 0.83 | 95,064 | 23,766 | 58 | 0.00061 | 0.00246 | 85,558 |
| 10 GH/s | 10 GH/s | 2.0 | 0.50 | 0.50 | 57,038 | 28,519 | 117 | 0.00205 | 0.00410 | 57,038 |
| 10 GH/s | 20 GH/s | 3.0 | 0.33 | 0.67 | 76,051 | 25,350 | 234 | 0.00307 | 0.00922 | 76,051 |
| 10 GH/s | 50 GH/s | 6.0 | 0.25 | 0.83 | 95,064 | 23,766 | 584 | 0.00615 | 0.02459 | 85,558 |
| 100 GH/s | 100 GH/s | 2.0 | 0.50 | 0.50 | 57,038 | 28,519 | 1,169 | 0.02049 | 0.04099 | 57,038 |
| 100 GH/s | 200 GH/s | 3.0 | 0.33 | 0.67 | 76,051 | 25,350 | 2,338 | 0.03074 | 0.09223 | 76,051 |
| 100 GH/s | 500 GH/s | 6.0 | 0.25 | 0.83 | 95,064 | 23,766 | 5,845 | 0.06148 | 0.24594 | 85,558 |
| 1000 GH/s | 1000 GH/s | 2.0 | 0.50 | 0.50 | 57,038 | 28,519 | 11,690 | 0.20495 | 0.40990 | 57,038 |
| 1000 GH/s | 2000 GH/s | 3.0 | 0.33 | 0.67 | 76,051 | 25,350 | 23,380 | 0.30742 | 0.92227 | 76,051 |
| 1000 GH/s | 5000 GH/s | 6.0 | 0.25 | 0.83 | 95,064 | 23,766 | 58,450 | 0.61485 | 2.45940 | 85,558 |
Honest-growth cost: a listing that doubles honest hash overnight halves every miner's subsidy per block (not per hash: difficulty halves the per-hash rate anyway; the rule halves it again) until W30 catches up, which is the 30-day ramp of ledger C7 (0.9x on day 28 to 31). With the floor m_min = 0.25 the worst case is a 4x cut, and the money is not lost to the chain: it reaches the provers of the same block, who are the same population. Per tier: a home miner's monthly income during a doubling month falls 50 percent under the rule against 50 percent already from difficulty (so 25 percent of the pre-event figure); a pool user sees the same through PPLNS; a prover with weight gains the diverted share. The gate: in the economy simulator (sim/economy/sim.py scenario f, a pool with the network's hash arriving on day 10) the incumbents' income under the rule must stay above the no-rule row for the 30 days and the newcomers' under; and in the fast-time harness a timestamp-manipulated H_now (headers inside the 132-s tolerance) must move m by under 2 percent.
## 3. Idea 2: work-stake, vote weight as the external-job bond
A key that claims an external job and delivers late or wrong loses s of its 30-day weight for 30 days (as equivocation strips 100 percent, spec 3.6). Weight is blue blocks; it cannot be bought, only mined. Arithmetic: what a stripped key forgoes.
| Key's hash share | Blocks per 30 days at 1 bps | Weight share | Shard sortition income per 30 days (20 percent pool, pro rata), IGN | Stripped at s = 25 percent: lost pool income over 30 days, IGN | Stripped at s = 100 percent | IGN bond that would match (design 4.6: maxPgas x f_p x 1.5 for a 1 B-cycle job at the floor) |
|---|---|---|---|---|---|---|
| 0.01 percent | 259 | 0.01 percent | 1,643 | 411 | 1,643 | 0.0015 IGN |
| 0.10 percent | 2,592 | 0.10 percent | 16,427 | 4,107 | 16,427 | 0.0015 IGN |
| 1.00 percent | 25,920 | 1.00 percent | 164,271 | 41,068 | 164,271 | 0.0015 IGN |
| 10.00 percent | 259,200 | 10.00 percent | 1,642,706 | 410,676 | 1,642,706 | 0.0015 IGN |
Reading: for every key above dust the 30-day pool income at risk is many orders above the designed IGN bond for one job, so weight is a far larger bond than coins, and it is a bond nobody can buy on a market. It also strips the key's vote for 30 days, which is the sentence the finality rule already hands out for equivocation. The cost: a false positive (a partition that makes an honest proof late) strips an honest voter; so the rule must use DAA time, a long deadline (the 120-s claim timeout of P9 decision, or longer), and a one-strike grace per 30 days. Per tier: a solo 8 GB miner below dust has no weight and so cannot take external jobs at all under this rule (it can still prove shards, which carry no bond); a pool user's jobs are the pool's and the pool's weight is at risk, which is what a pool operator wants priced. Gate: on the phase 4 devnet, 1,000 jobs with a 10 percent injected late rate: every injected fault stripped, zero honest keys stripped across a 60-s partition.
## 4. Idea 6: audits paid from the burn, by 60 percent signal, no standing address
| Chain traffic (fraction of full blocks) | Base fee burned per day, IGN | 7-day redirect, IGN | 30-day redirect, IGN | USD at 0.02 (7 d / 30 d) | USD at 0.10 (7 d / 30 d) |
|---|---|---|---|---|---|
| 0.01 | 2,592 | 18,144 | 77,760 | 363 / 1,555 | 1,814 / 7,776 |
| 0.10 | 25,920 | 181,440 | 777,600 | 3,629 / 15,552 | 18,144 / 77,760 |
| 0.50 | 129,600 | 907,200 | 3,888,000 | 18,144 / 77,760 | 90,720 / 388,800 |
| 1.00 | 259,200 | 1,814,400 | 7,776,000 | 36,288 / 155,520 | 181,440 / 777,600 |
Reading: at launch traffic (1 to 10 percent of full blocks) a 30-day redirect is USD 1,600 to 16,000 at 0.02 per IGN, under one Code4rena contest (base pricing from USD 6,500 before the 2025 zero-fee change; Immunefi's standard pays 10 percent of funds at risk). At half-full blocks it reaches a serious bounty (USD 78,000 for 30 days at 0.02). So the burn can fund audits only once the chain is used; before that the only money is the client's 1 percent dev fee and the founders' mined coins (litepaper: grants from founders' mined coins). The text gives the attack: this is a dev fund with a 60 percent veto and a per-event payee, which is exactly the switch spec 5.5 removed.
## 5. Idea 10: Igneum as the settlement layer for GPU rental
| Card | Vast.ai on-demand USD/h (cited) | Platform take modelled | Host loses USD per card-year | Igneum settlement cost per rental (2 transfers at the floor), IGN | At 0.02 and 0.10 USD per IGN | Hours of rental to pay 1 USD of chain fees at 0.02 |
|---|---|---|---|---|---|---|
| RTX 5090 | 0.44 | Vast about 15 percent | 579 | 0.0102 | 0.00020 / 0.0010 | 4,902 rentals |
| RTX 5090 | 0.44 | RunPod about 7 percent | 270 | 0.0102 | 0.00020 / 0.0010 | 4,902 rentals |
| RTX 4090 | 0.31 | Vast about 15 percent | 408 | 0.0102 | 0.00020 / 0.0010 | 4,902 rentals |
| RTX 4090 | 0.31 | RunPod about 7 percent | 190 | 0.0102 | 0.00020 / 0.0010 | 4,902 rentals |
Reading: the chain's fee is three to five orders of magnitude under the platform take. The platform's take pays for what the chain cannot do: matching, trust, dispute, image hosting, and the verification of delivered work. The honest problem is the last one: a rented hour of general compute is unverifiable, so an on-chain escrow without a verifier is a trust-me payment with lower fees. What is verifiable on this chain today: ZK proving jobs (the precompile); what is verifiable with sampling: deterministic recompute checked on a sampled fraction (SPEX, arXiv 2503.18899; Render's result-quorum, approximate); what needs hardware the fleet does not have: TEE attestation (NVIDIA confidential computing is H100 and H200 class, phala.com; no consumer card has it).
## 6. Idea 12: a randomness beacon from the checkpoint VDF
| Beacon | Period | Latency to a value | Unbiasability argument | Verify cost | Who runs it |
|---|---|---|---|---|---|
| drand quicknet (League of Entropy) | 3 s | about 3 s | threshold BLS over H(round) with a 2/3 threshold of about 20 named organisations; unbiasable while under 1/3 collude; unchained | one BLS verify | a league, trusted set |
| Igneum epoch seed today | 3,600 s | 600 s (10-min VDF) | a certified checkpoint 10 minutes before use; the VDF makes the last block producer's choice useless because it cannot see the output in time | 4.47 ms (516 B) | nobody: any node evaluates |
| Proposed: a per-checkpoint VDF beacon | 30 s | 30 to 60 s (a 30-s VDF of each certified checkpoint hash) | the checkpoint is locked by 2/3 of 30-day weight before the VDF starts, so no single party chooses the input; a last-block grind costs a block's subsidy per try and buys one bit of influence only if the attacker can evaluate the VDF faster than the chain, which is the class-group ASIC question (Chia timelords) | 4.47 ms per value, 2,880 values a day | nobody: the epoch pipeline already exists |
Reading: the chain already produces an unbiasable value once an hour with a 10-minute delay. A 30-s beacon is the same code at 120x the cadence, and its honest limit is the one Chia carries: the fastest class-group squarer sets the floor on 'delay', so a timelord-class ASIC owner can learn the value earlier than everyone else (Chia docs: timelords are software or ASIC). That earlier knowledge is a front-running edge, not a bias. The product: PREVRANDAO per block already comes from this pipeline (spec 7.1); the beacon makes it a 30-s value usable off-chain (lotteries, shuffles, timelock encryption as drand does).
## 7. Idea 14: proving other chains as the main income by 2030
| Income line | USD per day | Basis |
|---|---|---|
| Proving every Ethereum L1 block at the Sep 2026 tracker cost | 36 | 0.005 USD x 7,200 blocks; the price a buyer pays is above cost, call it 10x: 360 |
| The same at the Dec 2025 cost (under 0.04) | 288 | secondary |
| Igneum year-1 emission at 0.005 USD per IGN | 13,689 | 31.688 IGN per block x 86,400 |
| Igneum year-1 emission at 0.02 USD per IGN | 54,757 | 31.688 IGN per block x 86,400 |
| Igneum year-1 emission at 0.1 USD per IGN | 273,784 | 31.688 IGN per block x 86,400 |
| Rollup proving spend, all rollups (customer brief) | 8,219 to 27,397 | low millions a year, approximate |
| Boundless trailing day in the explorer (4 Oct 2026) | 2 | 8.4 T cycles at a 0.21 USD per B-cycle median, developer-adoption.md 2b, approximate |
Reading: the whole public proving market is three to four orders of magnitude under year-1 emission at any price input. For proving to be the main income by 2030, demand must grow about 1,000x while cost per proof keeps falling 3x to 30x a year, which pushes dollars per proof down as fast as volume rises. The arithmetic says never by 2030 for 'main income'; it says 'yes' for 'a second income that keeps cards on after the subsidy fades' (spec 5.10.2), which is the design's own claim.