igneum/docs/analysis/horizon/algorithm.md
igneum-josh 5ff73932e9 Horizon: lanes 2 (algorithm), 7 (frontier) and 8 (new-pow designs) land, with the summary skeleton
Josh's 6 October 2026 ask: deep backward and forward research across the hash, finality,
economy, network and every shipped surface. This commit carries the first three lanes.

- docs/analysis/horizon/algorithm.md: the chip model on the 6 October numbers (f = 1 GDDR7
  chip 5.7x per joule against the 5090 at class v3, 2.1x at class v4 with k = 1), the FPGA
  lane tightened to 0.30x to 0.47x per watt, the reserve R0 to R8, the reconciled shadow-N
  ladder (section 5.3a) with HBM4 and three verifier brackets, the first measured verifier
  proxy on igneum-build-1 (class v4 5.06 ms cold, dr736 10.51: out), the dataset schedule
  to 2030; model sim/horizon/algorithm/model.py.
- docs/analysis/horizon/frontier.md: sixteen ideas ranked by payoff over difficulty with the
  Monero and Kaspa attacks, prior art cited, the honest never column; model
  sim/horizon/frontier/frontier_model.py.
- docs/analysis/horizon/new-pow.md sections 0 to 4: three new proof-of-work schemes defined,
  reviewed in two personas, scheme A (mining is proving) ruled out on bytes and
  sampleability, B and C in prototype on two rented 4090s; measured rows follow.
- docs/analysis/horizon-2026-10.md: the summary skeleton and the lane table.

Every rental cost cites docs/bench-log.md "Rental cost of hash, 6 October 2026".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 20:51:23 +01:00

63 KiB

Horizon lane 2: the shipped hash and its class system, refined

6 October 2026, evening UK. Lane 2 (algorithm) of the Horizon programme. Worktree /Users/joshm/Projects/igneum-wt-horizon, branch horizon on master 3f4f719. Scope: the shipped lottery hash and its classes (v2, v3 live, the v4 candidate mx8+sh256x27, the reserve R0 to R8, the era draw, the dataset schedule). No new puzzle is proposed here (lane 8 owns that). Every chip figure is arithmetic on cited figures and is approximate; every GPU figure names its bench-log entry or analysis file; the one new measurement is the verifier proxy on igneum-build-1 (section 2).

What was read, in full unless marked: docs/spec/01-lottery-hash.md, docs/spec/04-seeds-and-vdf.md, docs/analysis/chip-model-v3.md, docs/analysis/asic-resistance-history.md, docs/analysis/latency-shadow-2026-10-06.md, docs/analysis/m16-recompute-attacker-2026-10-05.md, docs/analysis/sram-mirror.md, docs/analysis/int8-matrix-family.md, docs/analysis/weak-program-census-2026-10-03.md (section 1), docs/analysis/scratch-soundness.md (section 0), docs/analysis/card-lifetime-2026-10-05.md, docs/plans/counter-asic-3-status.md, docs/plans/counter-asic-3-reserve.md, docs/plans/counter-asic-3-derivation.md, docs/plans/counter-asic-2.md, docs/plans/epoch-length.md (sections 11 and 12), docs/plans/era-layout.md, docs/plans/mixer-x4.md (sections 6.2 to 6.5), docs/plans/read-width.md, docs/fud-ledger.md (M1 to M31 headings; M9, M16, M17, M18 in full; M28 heading), docs/bench-log.md (entries "Counter ASIC 2.0, the numbers", "the 9070 XT on the eGPU", the 6 October item 8, item 6 and gate-run entries, "Rental cost of hash, 6 October 2026"), igneum-pow/src/generator.rs (the class constants, ShadowClass, LoadClass::{V2, MX4, MX8, DR736}), igneum-pow/src/memhard.rs (Shape, the growth rule, round_key, derive_items_mask), igneum-pow/src/main.rs (bench), docs/analysis/prover-tiers-real-cards.md (the 11 rented cards), and /Users/joshm/Projects/igneum-wt-gpu-fleet/docs/analysis/block-rate-devnet2.md (RUN_A and RUN_B are still placeholders at 21:30 UK; nothing from it is used).

1. The six questions and the one-line answers

# Question Answer in one line
1 Chip model on the 6 October numbers; the FPGA lane The stored-dataset chip (f = 1, GDDR7) reads 5.7x per joule against the 5090 bench row and 1.7x against the M5 Max at class v3; at class v4 and k = 1 those fall to 2.1x and 0.9x. Per dollar it is 56x under rented hash and 5.4x under an owned 5090 per MH/s-hour. The FPGA soft overlay tightens to 0.30x to 0.47x per watt: the measured 2.4 G reads/s equals the JEDEC tFAW ceiling of a 2-stack HBM2 part, so the 1.9x bank-bound row is unreachable on any FPGA that exists. AWS F2 at USD 1.98 an hour carries the exact HBM2 subsystem and can measure it
2 Reserve R0 to R8 Every reserve family together costs a chip about 8 to 14 adders per lane, about USD 4 of N5 silicon on a 14,000-lane array; none moves the per-joule edge by over 10 percent. The reserve's value is obsolescence of a datapath taped out against class v3, and since the order is public that value is zero against a chip that ships with every block. Recommended order: R0 derive at dr368 (dr736 fails the gate proxy), R1 shfla, R2 perm, R3 popc and clz, R4 bfe, R5 shifts, R6 sel, R7 andn, R8 mm8 at era 8 by the rule, no exception
3 A class v5 from the shadow Option (i) is void: the v4 shadow already consumes loaded data (its registers hold the dataset words of the iteration), so a chip precomputes nothing today. Option (ii), a shuffle-heavy shadow mix, raises the attacker's k floor from about 0.32 to 0.46 (approx), not to 1. Option (iii): the one N every owned card holds within 5 percent is 130,000 counted ops (the M5 Max's point); it buys 2.1x to 1.7x against the 5090 at k = 1 for 4.8 percent of the Mac's rate; the verifier at 130,000 is 4.9 ms on the box proxy and 8.5 ms on the half-core proxy, inside the gate
4 The era draw's randomness Forging the certified checkpoint the era VDF reads costs 20 days of 100 percent hash: USD 5,600 at 1 GH/s, USD 5.6M at 1 TH/s rented, and buys one draw of a space whose spread is 0.8 to 3.2 percent of hash rate per card and 0 percent for a chip. Re-rolling by withholding needs a 1,800x faster VDF. The draw buys nothing against a chip; the public reserve order means a chip is taped out with every block. What would cost a chip is work (N) and the honest card's own watts, not unpredictability
5 The 2019-class verifier Measured proxy tonight: a Zen 4 core at 3.8 GHz with server DRAM reads 2.0x to 2.2x the M5 Max core (v4 4.90 ms steady, 5.06 cold; dr736 9.76 steady, 10.51 cold: FAIL); with both SMT siblings busy (the pessimistic bracket) v4 8.23 ms, dr368 8.16, dr736 15.5. The 2.5x rule holds within 15 percent. The gate protects a node at 1 to 10 percent of one core per block rate, an 18-minute IBD, and a header-flood cost of 100 bogus headers per second per core
6 Dataset growth to 2030 Hold option (b): 2 GiB at genesis, 4 GiB at year 4, 8 GiB at year 12. The 8 GB tier (26.7 percent of Steam today, about 0 by 2030 on the trend) mines to year 12 anyway; the 12 GB tier loses mine-and-prove compressed at the year-4 step and the 8 GB tier loses core-only at genesis, and both are the prover's footprint, not the dataset's. Growth is for the SRAM reticle (1.6 GiB per reticle at N5) and GPU L2, not for the HBM chip

2. Method

Item What was done Where
The model One Python script, every table in this file; inputs listed with source and label sim/horizon/algorithm/model.py, README.md beside it
Verifier proxy igneum-pow built on igneum-build-1 through tools/build-remote.sh from the crate directory (IGNEUM_AGENT=horizon, build slot build-0, 10 s wall, sccache miss 1, artefact 991,352 bytes, sha256 6d2867...); igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <c> --warps 50 for v2, mx8, mx8+sh256x27, dr368, dr736 under nice -n 19 taskset -c 40 (one core), then the same on cores 40 and 88 at once (the two SMT siblings of one physical core: thread_siblings_list 40,88); load average 1.3 before, 2.3 after; scaling_cur_freq read 3,799,885 kHz during a run (the governor is schedutil with scaling_max_freq 2,750,000 but the core boosted to 3.8 GHz; cpupower needs root, so no fixed low clock was possible); the cache fill 361 ms on that core section 5.5
Not run A Mac packbench ladder for v5 (the data-dependent and shuffle-heavy shadow variants do not exist as kernel text, so nothing could be measured; the plan is in 5.3); any PC job; any FPGA section 7
External figures Steam Hardware Survey September 2026 (store.steampowered.com/hwsurvey, read 6 October 2026); ICCAD 2021 "Demystifying the Characteristics of High Bandwidth Memory for Real-Time Systems" Table I (JEDEC HBM2 timings in cycles at 2,133 MT/s, text extracted with pdftotext); Intel HBM2 IP user guide (16 AXI pseudo-channel ports per stack); AWS F2 pricing (f2.6xlarge USD 1.98 an hour on-demand, USD 0.66 spot, one VU47P with 16 GB HBM2) section 5.1

3. Evidence

3.1 Honest cards, measured (class v3 unless said)

Card (tier) MH/s W uJ per hash Source
RTX 5090, PC 2 bench, 431 W cap, the control 132.2 350 2.65 latency-shadow-2026-10-06.md s5; bench-log item 8
RTX 5090, PC 1 app, Ember run 5 127.4 310 2.43 counter-asic-3-status.md s7 item 1
RTX 5090, app 5 October 124 290 2.34 miner-eff record, cited in latency-shadow s3
RTX 5090, rented Vast pod, untuned 98.5 258 2.62 prover-tiers-real-cards.md
RTX 4090, rented 52.3 183 3.50 same
RTX 3090, rented 37.8 229 6.05 same
RTX A5000, rented 47.7 223 4.67 same
RTX 4070, PC 1, 1,860 MHz lock, 160 W cap 30.95 79.5 2.57 status item 8, 4070 rows
RTX 4070, rented 25.0 91 3.65 prover-tiers
RTX 5070, rented 41.9 137 3.27 prover-tiers
RTX 3060, rented 23.8 104 4.36 prover-tiers
RTX 3080, rented 40.8 205 5.02 prover-tiers
RTX 4060 Ti 16 GB, rented 17.6 72 4.11 prover-tiers
RTX 4060 Ti 8 GB, rented 19.1 73 3.81 prover-tiers
RTX 4060, rented 17.1 no reading (0.0 W logged) n/a prover-tiers
RX 9070 XT, PC 1 18.9 199 to 203 10.6 bench-log 9070 XT telemetry; status 3a
Apple M5 Max, GPU + DRAM channels 27.08 21.0 0.78 latency-shadow s3
Apple M5 Max, package (approx) 27.08 38 1.40 ember-tune.md, approximate

At class v4 (mx8+sh256x27, measured): the 5090 131.95 MH/s at 431 W (3.27 uJ, the cap binds), the M5 Max 26.67 at 37.2 W (1.39), the 4070 31.08 at 109 W (3.51), the 9070 XT 19.29 at owed watts. Every other card's v4 energy below is modelled as the card's marginal ALU energy times N (NVIDIA 11 pJ per counted op measured on the 5090, Apple 6.9 measured on the M5 Max, AMD taken as NVIDIA's, approximate).

3.2 Chips (model, approximate; chip-model-v3.md s5.4, latency-shadow s6)

Chip class MH/s W at v3 uJ, v3 uJ at v4, k = 0.3 / 0.5 / 1 / 1.5 $ silicon + memory, v3 / v4 $ per MH/s
f = 0 on-die recompute, 256 MiB SRAM, x8 41.7 53.7 1.29 1.62 / 1.84 / 2.39 / 2.94 700 / 740 16.8
f = 1 stored dataset, GDDR7, 16 devices 166.4 77.6 0.466 0.80 / 1.02 / 1.57 / 2.12 470 / 510 2.8
f = 1, HBM3 one stack 83.6 26.8 0.321 0.65 / 0.87 / 1.42 / 1.97 550 / 590 6.6
f = 1, HBM3 eight stacks 666 174 0.262 0.59 / 0.81 / 1.36 / 1.91 2,650 / 2,690 4.0

k is the chip core's energy per counted op over the 5090's measured 11 pJ. The f = 0 chip's measured stand-in (M16 round 2, the inline kernel inside the 5090's L2) ran 33.9 MH/s at 431 W: 0.256x per chip and 5x worse per joule than honest, so the f = 0 row above is the model's ceiling for that class, not a measurement.

3.3 The verifier proxy, measured tonight (igneum-build-1, EPYC 9454P, one core at 3.8 GHz, nice 19)

Class M5 Max core, quiet 2.5x rule (approx) Box one core, steady / cold Box over Mac Box half-core (SMT sibling loaded) 10 ms gate
v2 0.60 1.5 1.13 / 1.30 1.88x not run pass
mx8 (class v3) 2.06 5.2 4.52 / 4.67 2.19x 7.56 pass
mx8+sh256x27 (class v4) 2.33 5.8 4.90 / 5.06 2.10x 8.23 pass
dr368 2.69 6.7 4.72 / 5.32 1.75x 8.16 pass, thin on the half-core
dr736 4.88 12.2 9.76 / 10.51 2.00x 15.49 FAIL (cold run over 10 ms on one core; 15.5 on the half-core)

Raw lines (the Mac figures are counter-asic-3-derivation.md 5.1 and the gate-run entry): box one core CPU verify: 1.128 / 4.516 / 4.901 / 4.716 / 9.763 ms per 32-lane warp, avg of 50; cold warp base 0: 1.305 / 4.670 / 5.061 / 5.324 / 10.509 ms; half-core 7.562 / 8.234 / 8.162 / 15.491 (cpu 40) and 7.556 / 8.227 / 8.168 / 15.479 (cpu 88); every lane-0 and lane-31 hash equal to the Mac's vectors (42246ba99fc58e4f, 19b56348bc85304d, the dr736 e23d389f3eea0c83). Cache fill 360.5 to 364.7 ms on the box core against 175 to 181 on the M5 Max.

3.4 Installed base (Steam Hardware Survey, September 2026, cited)

VRAM Share Trend (cited: 8 GB 35.03 percent in August 2025, 33.66 in September 2025; 12 GB 19.30 in August 2025)
512 MB to 4 GB 14.06% falling
6 GB 5.12% falling
8 GB 26.71% about -7 points a year
10 to 11 GB 2.74% flat
12 GB 13.06% about -6 points a year
16 GB 27.21% rising; overtook 8 GB in August 2026
20 to 24 GB 7.93% rising slowly
32 GB 1.41% rising slowly
64 GB 0.50% new row

3.5 Rental and ownership cost (bench-log "Rental cost of hash, 6 October 2026", measured)

USD 0.0117 per MH/s-hour on RunPod community pods (1,748 MH/s for USD 20.44 an hour); the 8x 4090 rig USD 0.0129; a 5090 pod USD 0.41 to 0.74 an hour for 98 to 128 MH/s.

4. Model

Formula Inputs (label)
Energy per hash E = W / rate measured watts and MH/s per card; chip W from chip-model-v3.md 5.4 (approx)
Chip energy at class v4: E_v3 + N x 11 pJ x k N = 100,000 counted ops (measured class v4), 11 pJ = the 5090's marginal per op (measured), k free
Edge per joule = E_card / E_chip both sides above
Hourly cost per MH/s = price / (2 years x rate) + E x 3.6e9 x USD 0.10 per kWh prices approximate (launch list from memory, labelled); chip $ from 5.4; electricity approx
FPGA random-read ceiling = min(banks / tRC, channels x 4 / tFAW, channels / tRRD) 2 stacks, 16 channels, 32 pseudo-channels, 16 half-banks each (approx); tRC 48 cycles, tFAW 30, tRRD 6 at 1,066 MHz (ICCAD 2021 Table I, JEDEC HBM2); latency 137.8 ns (Shuhai, measured)
Reads in flight = rate x latency; per watt = rate / board W U55C 115 to 150 W, U280 225 W (datasheets)
Verifier add per warp = shadow instructions / 1,000 x slope slope 3.2 us (M5 Max, measured), 7.0 us (box one core, this lane), 12.1 us (box half-core, this lane)
Checkpoint forgery cost = network MH/s x 480 h x USD 0.0117 20 days of 100 percent hash to 2/3 weight (CLAUDE.md headline, from the finality sim)
Tier fit at a dataset step: miner resident = dataset + 192 MiB; prover peak beside the miner from prover-tiers-real-cards.md scaled by the dataset's growth usable VRAM 75 percent (mine-only), 98 percent headless (the prover rows)

5. Results

5.1 Task 1: the chip model on the 6 October numbers, and the FPGA lane

Per joule and per dollar against every honest card (the full table with every card is the chip section of the model; the rows that decide things):

Card (tier) uJ v3 / v4 f = 0 chip edge, v3 / v4 at k = 1 f = 1 GDDR7 edge, v3 / v4 at k = 0.3 / 0.5 / 1 / 1.5 f = 1 HBM3 one stack, v3 / v4 at k = 1 $ per MH/s (card, approx) USD per MH/s-hour owned
RTX 5090 bench (32 GB) 2.65 / 3.27 2.1x / 1.4x 5.7x / 4.1x / 3.2x / 2.1x / 1.5x 8.3x / 2.3x 15.1 0.00113
RTX 5090 app, Ember (32 GB) 2.43 / 3.53* 1.9x / 1.5x 5.2x / 4.4x / 3.5x / 2.3x / 1.7x 7.6x / 2.5x 15.7 0.00114
RTX 4090 rented (24 GB) 3.50 / 4.60* 2.7x / 1.9x 7.5x / 5.8x / 4.5x / 2.9x / 2.2x 10.9x / 3.2x 30.6 0.00210
RTX 3090 rented (24 GB) 6.05 / 7.15* 4.7x / 3.0x 13.0x / 9.0x / 7.0x / 4.6x / 3.4x 18.9x / 5.0x 39.7 0.00287
RTX 4070 PC 1 tuned (12 GB) 2.57 / 3.51 2.0x / 1.5x 5.5x / 4.4x / 3.5x / 2.2x / 1.7x 8.0x / 2.5x 17.7 0.00127
RTX 5070 rented (12 GB) 3.27 / 4.37* 2.5x / 1.8x 7.0x / 5.5x / 4.3x / 2.8x / 2.1x 10.2x / 3.1x 13.1 0.00108
RTX 3060 rented (12 GB) 4.36 / 5.46* 3.4x / 2.3x 9.4x / 6.9x / 5.4x / 3.5x / 2.6x 13.6x / 3.8x 13.8 0.00123
RTX 4060 Ti 8 GB rented (8 GB) 3.81 / 4.91* 3.0x / 2.1x 8.2x / 6.2x / 4.8x / 3.1x / 2.3x 11.9x / 3.5x 20.9 0.00157
RTX 4060 Ti 16 GB rented (16 GB) 4.11 / 5.21* 3.2x / 2.2x 8.8x / 6.5x / 5.1x / 3.3x / 2.5x 12.8x / 3.7x 28.4 0.00203
RX 9070 XT (16 GB AMD) 10.6 / 11.7* 8.3x / 4.9x 22.8x / 14.7x / 11.5x / 7.5x / 5.5x 33.2x / 8.3x 31.7 0.00287
Apple M5 Max, GPU + DRAM 0.78 / 1.39 0.6x / 0.6x 1.7x / 1.8x / 1.4x / 0.9x / 0.7x 2.4x / 1.0x 129 0.00745
Apple M5 Max, package (approx) 1.40 / 2.02 1.1x / 0.9x 3.0x / 2.5x / 2.0x / 1.3x / 1.0x 4.4x / 1.4x 129 0.00752

* modelled v4 energy. The chip's hourly cost per MH/s (two-year amortisation plus electricity, approx): f = 1 GDDR7 USD 0.00021, HBM3 one stack 0.00041, eight stacks 0.00025, f = 0 recompute 0.00109; an owned 5090 0.00113; rented hash 0.0117. So the stored-dataset chip undercuts rented hash 56x and an owned 5090 5.4x per MH/s-hour, and the recompute chip matches the 5090 exactly, which is why nobody builds it.

What changed against the 5 October record: the honest denominators moved (the 5090 is 2.34 to 2.65 uJ by where it is measured, not 2.40), the marginal ALU energy is measured at 11 pJ (not the model's 5.5), the M5 Max at 0.78 uJ is the honest best per joule by 3x, and the rented fleet shows the untuned mid-tier (3060, 3080, 3090, A5000, 4060 Ti) at 3.8 to 6.1 uJ, 1.5 to 2.3x worse than the 5090 bench row: against those cards the GDDR7 chip reads 8x to 13x per joule at v3 and 3.1x to 4.6x at v4 with k = 1. The per-tier reading: the Apple tier is already inside 2x of the GDDR7 chip with no shadow and crosses under 1x at v4 and k = 1; the 5090 and the tuned 4070 reach about 2.1x to 2.2x at v4 and k = 1; the untuned mid-tier and AMD stay at 3x to 7.5x, which Ember tuning (the 4070 rows: 3.65 to 2.57 uJ) closes by about 30 percent and nothing in the hash closes further.

The FPGA lane. The brief's soft-overlay range was 0.30x to 0.39x per watt measured-basis and 0.7x to 1.9x at an unmeasured bank-bound ceiling. The reads-in-flight model with the JEDEC HBM2 timings:

Ceiling Formula G reads/s per card Reads in flight at 137.8 ns Per W at 115 / 150 / 225 W (M/s/W) Against the 5090 per W (53.7 M at 326 W)
Measured, Shuhai U280 default mapping (FCCM 2020, Fig 7) 32 pc x 75 M 2.4 331 21 / 16 / 11 0.39x to 0.30x (U55C); 0.20x (U280)
tFAW-bound (JEDEC HBM2, ICCAD 2021 Table I: 30 cycles at 1,066 MHz, 4 ACT per channel) 16 ch x 4 / 28.1 ns 2.3 313 20 / 15 / 10 0.37x to 0.28x; 0.19x
tRRD-bound (6 cycles) 16 ch / 5.6 ns 2.8 392 25 / 19 / 13 0.46x to 0.35x; 0.24x
Bank-bound, no activate window (the epoch-length 12.2 ceiling row) 32 pc x 16 banks / 45 ns 11.4 1,567 99 / 76 / 51 1.84x to 1.41x; 0.94x
O'Connor's HBM2 activate figure as carried by chip-model-v3 5.3 16 ch x 8 / 12 ns 10.7 1,470 93 / 71 / 47 1.73x to 1.32x; 0.88x

The reading: Shuhai's measured 2.4 G/s equals the tFAW ceiling at JEDEC timings (2.3 G/s). The measured row was read on 6 October as a mapping artefact ("the paper's point is that this mapping is the wrong one for random access"); the activate window says it is the DRAM's own limit, which a bank-interleaved mapping does not lift because tFAW is enforced per channel by the die. The 11.4 G bank-bound row needs tFAW gone; O'Connor's 12 ns figure (which the chip model carries for HBM2 and HBM3) is 2.3x shorter than the JEDEC HBM2 cycle count tabled by ICCAD 2021, and the difference is the whole 0.7x to 1.9x row. Tightened range for a 2-stack HBM2 FPGA (U55C, U280, F2's VU47P): 2.3 to 2.9 G reads/s, 0.30x to 0.47x of the 5090 per watt, in the RX 9070 XT's class (2.4 to 2.7 G/s measured). HBM2e parts (Versal HBM, Agilex 7 M) raise the pin rate, not tFAW in nanoseconds (approx), so they sit in the same band; no FPGA with HBM3 exists as a product. Approximate throughout: the 16 half-banks per pseudo-channel, the 1,066 MHz reading of the ICCAD table, the board watts under load.

Can a rented FPGA hour measure it? Vast.ai lists no FPGAs (GPU marketplace only, checked 6 October 2026). AWS F2 (f2.6xlarge: one Virtex UltraScale+ HBM VU47P, 16 GB HBM2 in 2 stacks, 32 pseudo-channels, USD 1.98 an hour on-demand in us-east-1, USD 0.66 spot) carries the same HBM2 subsystem as the U55C and U280, so yes. The gate, written as a measurement plan:

Step What Hours (agent) Pass line
1 Port the chase kernel of docs/benchmarks/repro.md 2.2 to a Vitis HLS AXI master over the HBM IP: 1 GiB working set across all 32 pseudo-channels, N dependent 4-byte-reads-in-flight lanes (N = 256, 1,024, 4,096), the HBM IP's address map set to bank-interleaved (RAMA or the IP's "random access" option), a second variant with 32-byte reads 4 to 6 the kernel reports reads per second and the chain's checksum equal to the CPU's
2 Build the AFI (the F2 shell flow; the Vivado licence rides with the instance), 2 to 4 hours of F2 time at USD 2 to 8 2 (mostly waiting) an AFI that loads
3 Run the ladder; read board power through xbutil examine --report electrical (or the F2 shell's sensors) at 1 Hz; take the mean over each run 1 reads per second and watts per rung
4 Write the row into epoch-length.md 12.2 in place of the ceiling row 1 the public claim carries a measured FPGA number
Gate reads per second per watt at 1 GiB expected 15 to 25 M/s/W (0.3x to 0.5x of the 5090); the alarm line is 27 M/s/W (0.5x); over 54 M/s/W (1.0x) the FPGA lane becomes a Counter ASIC 4.0 item

Consequences per tier of the FPGA finding: none today (no FPGA mines); if the measured row holds, a soft-overlay FPGA at USD 4,000 to 5,000 a card (approx) mines at an RX 9070 XT's rate per watt for 7x the price, so no home or rig tier is displaced by it; the per-program bitstream lane stays answered by layer 9.

5.2 Task 2: the reserve R0 to R8

Slot Family Chip block it adds (approx area in 32-bit adders per lane, counter-asic-3-reserve.md s3) Chip datapath energy per op, N5 floor (pJ, approx) Honest step cost, Apple / NVIDIA / AMD (measured, ratio to the add-xor-rotate chain) Verifier cost Chip edge per joule it moves (model) Cost per tier
R0 derive (per-day item program) a sequencer: 60 KB instruction store, 16-register file, ALU with multiplier and rotator; removes the 3x fixed-function credit of the f = 0 chip n/a (it is the f = 0 chip's item cost: 9,992 ops per item) hash rate 0 on all three vendors; daily build +7 ms Mac, 0 on the 5090 and 9070 XT; compile +0.75 s Mac, +1.1 s NVRTC, +2.0 s AMD per day (a once-a-day module is a requirement) dr736 +2.8 ms per warp on the Mac, +5.2 on the box core, over the gate cold; dr368 +0.6 Mac, +0.2 box, 8.16 on the half-core f = 0 chip: 0.92x to 0.34x to 0.43x per chip; per joule from 1.86x to about 0.7x to 0.9x at the same allowance (approx); f = 1 chips: 0 pool verifier cores x2.4 at dr736, x1.3 at dr368; every miner tier 0
R1 shfla (lane + delta) 32-lane x 32-bit crossbar per warp, 32,768 mux bits, 3 to 6 adders per lane 1.00 1.91 / 1.53 / 0.75 to 0.84 one op per instruction, under 0.01 ms per warp at W_new = 4 shadow floor +21 percent at 5 percent of instructions (approx); the GPU pays 1.5x the step too, so k unchanged Apple under 1 percent of ALU time (argued), NVIDIA and AMD 0
R2 perm (byte permute) 4x4 byte crossbar, 128 mux bits, 1 to 2 adders 0.10 1.13 emulated / 1.30 / 1.73 to 1.93 emulated negligible under 1 percent Apple and AMD emulate at 1.1x to 1.9x per op: under 1 percent at 4 points
R3 popc and clz popcount tree and priority encoder, 1.5 to 3 adders 0.10 0.87 and 1.01 / 1.50 and 1.63 / 0.91 to 1.01 and 1.19 to 1.30 negligible under 1 percent 0
R4 bfe mask generator on the shifter, 0.3 adders 0.06 0.77 / 1.54 / 1.00 to 1.09 negligible 0 0
R5 shl, shr barrel shifter beside the rotator, 0.2 adders 0.06 0.85 and 0.86 / 1.27 and 1.28 / 1.00 to 1.02 and 0.91 to 1.02 negligible 0 0
R6 sel 32 muxes, 0.3 adders 0.03 0.76 / 1.32 / 0.90 to 1.01 negligible 0 0
R7 andn 32 inverters, under 0.1 adders 0.02 0.75 / 1.26 / 0.91 to 1.01 negligible 0 0
R8 mm8 8x8x16 u8 MAC tile per warp, about 100 adders per lane; licensable IP at every node 1.60 owed (Metal 4 matmul2d) / 2.43 / 1.68 to 1.83 native, exactness unverified 1,024 MACs per instruction per unit, about 10 us per warp at W_new = 4 shadow floor +37 percent at 5 percent of instructions (approx); the GPU pays 2.4x the step: k unchanged or worse for the GPU Apple pays the library path (owed); NVIDIA and AMD native

The ordering argument. Against the f = 1 chip, which is the chip anyone builds, every family R1 to R8 moves the per-joule edge by under 10 percent, because the chip's cost at class v4 is N x k x 11 pJ and a family changes the per-op energy of 4 points in 79 of the shadow mix. Against the f = 0 chip only R0 matters (it is the item cost). So the reserve's function is not per-joule resistance; it is to make a datapath taped out against class v3's eleven families wrong at the first unlock. A chip maker who reads this document tapes out every block from day one: R1 to R7 together are about 8 to 14 adders per lane (a lane with a 32x32 multiplier is about 35), about 11 mm^2 on a 14,000-lane N5 array, about USD 4 of silicon (sram-mirror.md's USD 0.36 per mm^2, approx); R8 is a licensed tile. The era schedule therefore buys nothing against a chip and the order can be set by honest cost alone: the largest new structure first while every vendor's measured step cost is in (shfla: AMD measured cheap at 0.75 to 0.84, which the reserve document said was the one number that could move it to R1), the emulated families next, mm8 last at era 8 by the rule (no era-4 exception: mm8 removes no adversary class, epoch-length.md 12.4). R0 enters at dr368, not dr736: the box proxy puts dr736 over the gate on a cold run (10.5 ms) and at 15.5 ms on the half-core; dr368 reads 4.7 and 8.2.

Recommended text change to counter-asic-3-reserve.md section 5: R1 shfla, R2 perm, R3 popc and clz, R4 to R7 unchanged, R8 mm8 at era 8; R0 at derive_len = 368 with 736 behind the O-1.14 measurement. The era schedule "family n at era n" stands; what it costs per tier at each unlock is the step-cost table above (Apple under 1 percent of ALU time per family, NVIDIA and AMD nothing measurable), and what it buys is stated honestly in the public text (section 6, proposal 4).

5.3 Task 3: a class v5 from the shadow

The three options, each priced:

(i) Shadow ops that depend on the loaded data. The v4 shadow block runs at the end of every iteration on the eight lane registers, and those registers hold the words the iteration's 16 loads XORed in (generator.rs: the block is executed after instruction 63 with the iteration's sel; verify.rs runs it with the same step). So every shadow op already consumes loaded data, and the next iteration's load addresses depend on the shadow's outputs. A chip cannot precompute any of it; what it can do is what the GPU does: overlap one lane's shadow with another lane's reads in flight. Option (i) therefore moves k by 0. A variant that draws the shadow's immediates from loaded words is M18's one-bit select at 32 bits: a chip wires the operand, and it is not a defence. Verdict: void; no measurement needed.

(ii) Shadow work that exercises GPU structures a chip lacks. Bank-conflict timing is not a value and cannot enter a bit-exact function. Register-file width (8 registers today) can be widened in the shadow to 16 or 32 (ProgPoW used 32): a chip's register file per lane in flight grows from 32 B to 128 B, 1.8 MB on a 14,000-lane array, trivial. Warp shuffles are the structure that costs: the xor-mask shuffle needs a 5-stage butterfly per warp (5,120 mux bits), the indexed shuffle a full crossbar (32,768). The chip datapath floor per op (N5, approx: add 0.06 pJ, mul 0.52, butterfly 0.30, crossbar 1.00, mm8 tile 1.60) against the GPU's measured step costs:

Shadow op mix Chip floor per op (pJ, approx) k floor after the 2x pipeline and 8x register and wire overhead of latency-shadow s6 (approx) GPU cost of the same mix on the 5090 (step ratios, measured)
Today's weights (add 32, mul 22, rot 13, shfl 8 of 75) 0.221 0.32 1.0x to 1.5x per op
Shuffle-heavy (shfl 14, shfla 8, the rest scaled) 0.315 0.46 the shuffles cost the 5090 1.49x to 1.53x the chain step, so its own energy per op rises about 10 percent (approx)
With mm8 at 8 points about 0.45 about 0.65 mm8 costs the 5090 2.43x the step

So a shuffle-heavy shadow raises the attacker's claimed floor from about 0.3 to about 0.5 and costs the GPU about 10 percent more energy per shadow op; it does not reach k = 1. What a chip would need to add: the 32-lane crossbar per warp (R1's structure), nothing else new. The structure argument is bounded: a wide-SIMD array with a crossbar per 32 lanes is still a fixed datapath with no scheduler, and that is where the 0.3 came from.

(iii) One N for every card. Per-card bind points at the 5 percent rule (measured): M5 Max about 130,000 counted ops (interpolated between 102,100 at -1.5 percent and 150,800 at -7.3), RTX 5090 at its 431 W cap about 210,000 (199,600 at -2.7, 330,700 at -34.7), RTX 4070 at its 160 W cap about 250,000 (199,600 at +0.4, 330,700 at -12.3, approx), RX 9070 XT over 331,000 (holds at every rung; budget about 650,000). The consensus N is the minimum: 130,000, set by the Mac. The unmeasured rented cards by ALU budget (approx, cores x clock): 3060 about 270,000, 4060 about 440,000, 3080 about 360,000; their power caps are unmeasured and on NVIDIA the cap binds before the budget, so these are not pass marks. The ladder at every N the lane can compute:

N counted ops 5090 uJ (rate delta) M5 Max uJ (delta) 4070 uJ (delta) f = 1 GDDR7 chip uJ at k = 0.3 / 0.5 / 1 / 1.5 Edge over the 5090 Edge over the M5 Max Edge over the 4070 Verifier add per warp: M5 Max / box core / box half-core (ms)
930 (class v3) 2.65 0.78 2.57 0.47 / 0.47 / 0.47 / 0.47 5.7x 1.66x 5.5x 0
102,100 (class v4) 3.27 (-0.2%) 1.39 (-1.5%) 3.51 (+0.4%) 0.80 / 1.02 / 1.58 / 2.14 4.1x / 3.2x / 2.1x / 1.5x 1.74x / 1.36x / 0.88x / 0.65x 4.4x / 3.4x / 2.2x / 1.6x 0.18 / 0.39 / 0.67
130,000 (the one-N candidate) 3.27 (-0.3%) 1.43 (-4.8%) 3.78 (+0.4%) 0.89 / 1.18 / 1.89 / 2.60 3.7x / 2.8x / 1.7x / 1.3x 1.61x / 1.22x / 0.76x / 0.55x 4.2x / 3.2x / 2.0x / 1.5x 0.23 / 0.49 / 0.85
150,800 3.27 (-0.3%) 1.46 (-7.3%) 3.98 (+0.4%) 0.96 / 1.29 / 2.11 / 2.94 3.4x / 2.5x / 1.5x / 1.1x 1.52x / 1.13x / 0.69x / 0.50x 4.1x / 3.1x / 1.9x / 1.4x 0.26 / 0.57 / 0.99
199,600 3.35 (-2.7%) 1.58 (-10.5%) 4.45 (+0.4%) 1.12 / 1.56 / 2.65 / 3.74 3.0x / 2.1x / 1.3x / 0.9x 1.40x / 1.01x / 0.59x / 0.42x 4.0x / 2.9x / 1.7x / 1.2x 0.35 / 0.76 / 1.31
330,700 4.99 (-34.7%) 1.87 (-21.0%) 5.89 (-12.3%) 1.55 / 2.28 / 4.09 / 5.91 3.2x / 2.2x / 1.2x / 0.8x 1.21x / 0.82x / 0.46x / 0.32x 3.8x / 2.6x / 1.4x / 1.0x 0.58 / 1.26 / 2.18

Per-tier watts at N = 130,000 (measured rungs interpolated): the M5 Max 37 W GPU plus DRAM (from 21), the 5090 431 W (its cap, from 350), the 4070 about 118 W (from 79.5), the 9070 XT owed; a 5090 rig pays about 23 percent more electricity for 0.3 percent less rate, an Apple miner 1.8x the GPU watts for 4.8 percent less rate, a pool user nothing, every verifier tier +0.23 to +0.85 ms per warp. The verifier at 130,000 on the half-core proxy is 8.5 ms (8.23 + 0.85 - 0.67), inside the gate with 1.5 ms spare; the gate does not bind N before the Mac does.

5.3a The reconciled N ladder (this lane's measured cards, lane 7's HBM4 inputs; --section ladder)

Lane 7 (docs/analysis/horizon/frontier.md 2.3, sim/horizon/frontier/frontier_model.py model 1.4) adds HBM4: JEDEC JESD270-4 doubles channels per stack (16 to 32, cited), so its one-stack ceiling is taken as 2 x 10.7 = 21.4 G reads/s at 1.0 nJ per read (approximate, unsourced), 5 W static, 10 W controller: 167 MH/s at 0.218 uJ bare. The table below carries that chip beside the three of chip-model-v3 5.4, every chip at N = bare + (N - 930) x 11 pJ x k, and every honest card at its measured rung.

N counted ops 5090 uJ, W (rate delta) M5 Max uJ, W (delta) 4070 uJ, W (delta) 9070 XT rate delta (W owed) 12 GB and 8 GB rented cards Chip edge over the 5090 per joule at k = 0.5 / 1 / 1.5: GDDR7 HBM3 one stack HBM3 eight stacks HBM4 one stack Verifier ms per warp: M5 Max core / 2019-class by the 2.5x rule / measured half-core proxy Card that binds first (5 percent rule)
930 (class v3) 2.65, 350 W (0) 0.78, 21 W (0) 2.57, 80 W (0) 0 measured at v3 only: 5070 3.27 uJ, 3060 4.36, 4070 untuned 3.65, 4060 Ti 8 GB 3.81 5.7x 8.3x 10.1x 12.2x 2.06 / 5.2 / 7.56 none
49,700 3.21, 425 W (+0.1%) 1.16, 31 W (-1.2%) 3.01, 94 W (+0.4%) +2.1% not measured at any N 4.4x / 3.2x / 2.5x 5.5x / 3.7x / 2.9x 6.1x / 4.0x / 3.0x 6.6x / 4.3x / 3.1x 2.15 / 5.4 / 7.88 none
102,100 (class v4) 3.27, 431 W (-0.2%) 1.39, 37 W (-1.5%) 3.51, 109 W (+0.4%) +2.0% not measured 3.2x / 2.1x / 1.5x 3.7x / 2.3x / 1.6x 4.0x / 2.4x / 1.7x 4.2x / 2.5x / 1.7x 2.24 / 5.6 / 8.23 none (M5 Max -1.5%)
130,000 3.27, 431 W (-0.3%) 1.43, 37 W (-4.8%) 3.78, 117 W (+0.4%) +1.7% not measured 2.8x / 1.7x / 1.3x 3.2x / 1.9x / 1.3x 3.4x / 1.9x / 1.4x 3.5x / 2.0x / 1.4x 2.29 / 5.7 / 8.41 M5 Max at its 5 percent point
150,800 3.27, 431 W (-0.3%) 1.46, 37 W (-7.3%) 3.98, 124 W (+0.4%) +1.5% not measured 2.5x / 1.5x / 1.1x 2.9x / 1.7x / 1.2x 3.0x / 1.7x / 1.2x 3.1x / 1.8x / 1.2x 2.32 / 5.8 / 8.55 M5 Max (-7.3%)
199,600 3.35, 431 W (-2.7%) 1.58, 38 W (-10.5%) 4.45, 138 W (+0.4%) +1.0% not measured 2.1x / 1.3x / 0.9x 2.4x / 1.3x / 0.9x 2.5x / 1.4x / 0.9x 2.6x / 1.4x / 1.0x 2.41 / 6.0 / 8.87 M5 Max (-10%), 5090 (-2.7%)
330,700 4.99, 431 W (-34.7%) 1.87, 40 W (-21.0%) 5.89, 160 W (-12.3%) +3.6% not measured 2.2x / 1.2x / 0.8x 2.3x / 1.3x / 0.9x 2.4x / 1.3x / 0.9x 2.5x / 1.3x / 0.9x 2.64 / 6.6 / 9.74 5090 (-35%), M5 Max (-21%), 4070 (-12%)

Sources per column: 5090 and M5 Max rungs latency-shadow-2026-10-06.md s3 and s5; 4070 and 9070 XT rungs counter-asic-3-status.md item 8 (the 9070 XT's watts owed: the ADLX sampler parsed 0 samples); the rented cards prover-tiers-real-cards.md (class v3 only); chip bare energies chip-model-v3 5.4 and lane 7 model 1.4; the 11 pJ unit latency-shadow s5; the verifier slopes 3.2 us per 1,000 shadow instructions (Mac, measured) and 12.1 us (box half-core, this lane); the 130,000 row interpolated. Every chip cell is approximate.

Disagreements with lane 7's model 1.4, named: (a) its honest card at N is the linear 326-to-575 W model of chip-model-v3 5.7 (2.95 uJ at N = 100,000, 3.50 at 200,000, 4.22 at 330,000); the measured 5090 under its 431 W cap reads 3.27, 3.35 and 4.99 uJ with -0.2, -2.7 and -34.7 percent of rate (the cap binds from 102,100 ops; power.min_limit is 400 W, so no cap below the one measured exists on a 5090). (b) Its shadow core is 150 W fixed at the 5090's 136 MH/s; on a 167 MH/s HBM4 chip that under-counts the core by 1.23x. Energy per hash is N x 11 pJ x k whatever the chip's rate, so HBM4 at N = 100,000 and k = 1 is 1.33 uJ and the edge 2.5x, not 1.11 uJ and 2.65x; at 200,000 1.4x (lane 7: 1.74x); at 330,000 1.3x (1.33x: agrees, because the 5090's own energy jumps to 4.99 at that rung). (c) Its HBM3 and HBM4 ceilings (10.7 and 21.4 G per stack) rest on the 8-activates-per-12-ns rate chip-model-v3 5.3 carried from O'Connor; the JEDEC HBM2 cycle table (ICCAD 2021 Table I) gives 4 per 28 ns per channel (section 5.1 of this file), and HBM3's own tFAW is behind the paywall. If HBM3 and HBM4 keep HBM2's window the one-stack ceilings are 2.3 and 4.6 G, HBM4's bare energy 0.55 uJ and its bare edge 4.9x, not 11x. The GDDR7 column is the one with a measured anchor (the 5090 reaches 82 percent of its ceiling) and is the column to quote in the summary; the HBM columns are the upper bound.

Verifier headroom for N (lane 7's "10x of headroom"): on the M5 Max core 10 - 2.33 = 7.67 ms buys 2.4 M shadow instructions, N about 4.5 M ops (19x); on the 2.5x rule 4.2 ms buys 525,000 instructions, N about 1.06 M (10x, lane 7's figure); on the measured half-core proxy 1.77 ms buys 146,000 instructions, N about 370,000 (3.7x). The 10x holds on the rule and not on the pessimistic measured bracket; O-1.14 decides which. Either way the cards bind first: M5 Max 130,000, 5090 at 431 W 210,000, 4070 at 160 W about 250,000, 9070 XT over 331,000. An unconditional doubling of N per era (lane 7's candidate) would take the Apple tier out at the first step (200,000: -10.5 percent) and the 5090 and 4070 at the second (400,000: compute-bound at their caps), which is why the proposal below steps N by signal, not by schedule alone.

The verdict on v5: the shadow lever is close to spent on the owned cards. Going from 100,000 to 130,000 buys 2.1x to 1.7x against the 5090 at k = 1 and costs the Mac its whole 5 percent allowance; a shuffle-heavy mix buys the k floor 0.3 to 0.5. Together they define one candidate, v5 = mx8 + sh256x35 with the shuffle-heavy weight table, worth measuring but not worth a cut on its own: the chip question is k, and no shadow design moves k past about 0.5 against a fixed-datapath array.

What measurement decides it (not run tonight; the shuffle-heavy weight table does not exist as a knob, and the Mac's miner state was not checked, so the plan stands in for the run): (1) add a shadow weight table to ShadowClass (generator.rs, 2 hours), draw the block from it, emit it in the three dialects as today; (2) export mx8+sh256x35 at today's weights and at the shuffle-heavy table for seed igneum-genesis; (3) Mac: with-lock.sh measure packbench --pack <dir> --batches 60 --batch-log2 24 --group 256 with the IOReport sampler, 2 packs plus the control, about 6 minutes, the miner paused first (curl -s http://127.0.0.1:60030/.../api/state from the app's app.url, then pause through the app, never from a script); (4) the same ladder as PC jobs on the 5090, 4070 and 9070 XT through the existing tools/ca3-shadow playbooks with the card off through the runner's --cards-off; (5) igneum-pow bench on the Mac core and the box proxy. Pass lines: every owned card within 5 percent of its class v4 rate; bit-exact fingerprints on Metal, CUDA and AMD OpenCL; verifier under 10 ms on the half-core proxy; the 5090's marginal pJ per op on the shuffle-heavy mix read on the three rungs under its cap.

5.4 Task 4: the era draw's randomness

How it picks. Era n's seed E_n is the 1-hour class-group VDF of Hash(chain_id || n || blue block hashes of the day before C_era(n)), where C_era(n) is the highest certified checkpoint at least 7,200 DAA s before the era (spec 04 s4.4). One SplitMix64 stream from E_n draws, in order: the ten op weights perturbed by -2..+2 points each, the output fold rotations, an unused draw for epoch_len (set by signal), then under class v3 a second stream draws the load width (pinned at 4 bytes: the draw is consumed), the stride multiplier M (odd) and rotation R, and the four interleave bit positions (spec 01 s1.13.1, era-layout.md 1.1). Not drawn: the load count (16), the mixer round count (8) and multiplier (8), the cache and dataset sizes, the item construction.

What an attacker can bias, priced. Two routes. (a) Forge the certified checkpoint: needs 2/3 of the 30-day blue-block weight, which is 20 days of 100 percent of the network's hash (the headline). Rented at the measured USD 0.0117 per MH/s-hour:

Network hash 20 days of 100 percent, rented What it buys in the draw
1 GH/s USD 5,616 one era's (weights +-2, fold, M, R, pos): a per-card hash-rate spread of 0.8 to 3.2 percent (six eras measured, bench-log "Counter ASIC 2.0, the numbers"), 0 chip effect
10 GH/s USD 56,160 the same
100 GH/s USD 561,600 the same; the rental market could not supply 20 pods of any card at 19:00Z on 6 October (bench-log), so a TH/s is not rentable at all
1 TH/s USD 5.6M (not supplied) the same

(b) Re-roll without weight: the miner of the last blue block before C_era(n)'s cut withholds or publishes to change the input set; this costs one block's reward and needs the 3,600-s VDF evaluated inside the 2-s publish window, a 1,800x faster evaluator (spec 04 s4.6: a 300x evaluator beats the epoch's 600 s, not the era's 3,600 s). Even free, one re-roll is one more sample of the same space. So the draw is unbiasable at any price that matters, and that is the honest answer to "what can be biased": nothing worth having.

Weak corners. The op-weight perturbation can move the multiply share (mul, mad, mulhi: 22 of 75) to 16 or 28 of 75; at the N5 datapath floor that is 0.158 to 0.232 pJ per shadow op (0.195 at the base), about +-20 percent of the shadow's datapath energy, and the GPU's energy moves the same way (its IMAD is the chain's own op). The stride and interleave cost a chip two integer operations per load and an address-line permute, nothing per joule. The fold rotations are a wire mux. The measured six-era spread (1.3 percent on the 5090, 3.2 on the 9070 XT, 0.8 on the M5 Max) is the whole of what the draw moves. There is no corner that makes a chip easier, because no drawn parameter touches the memory bound, the item derivation or N.

Predictability against the 32-month lead time. Everything a chip needs is public at genesis: the eleven live families, the reserve order R0 to R8 and its era schedule, the mixer shape and x8, the dataset and cache schedule, the class v4 shadow's size and weights. The draw hides (M, R, pos, weights +-2, fold rotations) until 2 hours before each era, and none of those needs silicon: an address decoder that permutes lines, a programmable rotator, an immediate table. So a chip taped out in month 0 against this spec runs every era for the chain's life, and the era draw buys nothing against it. What the draw does buy: the fork-fatigue lesson of the history (no human release, no vote), and a per-program hard-datapath FPGA cannot amortise a bitstream across eras (layer 9 handles the within-era case). Said plainly for the public text: the era draw and the reserve are automatic schedule changes against fixed datapaths and governance, not unpredictability against a chip. What would be unpredictable and costly to a chip is not available in a genesis-fixed rule set: a per-era draw among K reviewed item constructions (K mixers or K derive forms) is still K public blocks a chip carries; a per-era draw of the load count or the dependent-read depth changes the memory bound per era and fails the 5 percent rule on the honest cards (read-width: a per-program width mix spread 5.5 to 22.3 percent). What costs a chip is N (its k), the memory system (f = 1 is the ceiling and is a commodity controller), and the honest card's own watts (the M5 Max at 0.78 uJ is inside 2x of the GDDR7 chip with no shadow at all).

5.5 Task 5: the 2019-class verifier gate

Measured tonight (section 3.3): on igneum-build-1 one EPYC 9454P core boosted to 3.8 GHz (not a low clock: the governor's cap is 2.75 GHz but the core read 3,800 MHz under load, and cpupower needs root) with server DDR5 behind it reads 1.9x to 2.2x the quiet M5 Max core across five classes, and 2.0x on the cache fill. The half-core proxy (both SMT siblings busy on the same class) reads 3.4x to 3.7x the Mac. A 2019 laptop core (Skylake-class at 3.5 to 4.5 GHz with DDR4 at about 80 ns) sits between these brackets on the arithmetic (lower IPC than Zen 4, lower DRAM latency than the server), so the 2.5x rule is about right and the two proxies bracket it. Against the gate:

Class One box core, cold run Half-core Verdict at 10 ms Headroom left for shadow on the half-core (ms)
class v3 (mx8) 4.67 7.56 pass 2.4
class v4 (mx8+sh256x27) 5.06 8.23 pass 1.8 (about 150,000 more shadow instructions at 12.1 us per 1,000)
dr368 5.32 8.16 pass 1.8
dr736 10.51 15.49 FAIL on both proxies none

So dr736 is out as a genesis-live or near-term reserve length on measured evidence, not on the 2.5x rule; dr368 is in. The owed O-1.14 measurement on a real 2019 laptop (the US laptop's Windows igneum-pow build, main's decision 7) still closes the question; the box proxy is the stand-in until it lands, and the half-core row should be the standing pessimistic rule in place of "2.5x" (it is a measurement; 2.5x is a ratio from memory).

How to measure it tomorrow: tools/cross-remote.sh from igneum-pow/ builds the Windows exe on the box (1 min 44 s measured for the node; the pow crate is one crate, under a minute), the relay carries it to the US laptop when it appears, igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <c> --warps 50 for the five classes, ms per warp into docs/bench-log.md under O-1.14; 1 hour of agent work. A rented old CPU on Vast is the fallback (Vast lists CPU-only offers by core generation; a 2019 Xeon or i7 host at under USD 0.10 an hour, approx), same binary, same command.

What the gate protects, and what loosening it costs:

At 10 ms per warp At 20 ms (loosened) At 5 ms (tightened)
A node on a 2019 laptop at 1 bps 1 percent of one core per block 2 percent 0.5 percent
At 10 bps (the Devnet 2 experiment) 10 percent of one core 20 percent 5 percent
IBD over the 108,000-header pruning window, one core 18 min 36 min 9 min
Header flood (M15 class): invalid headers per second that saturate one core 100 50 200
A pool core verifying shares 100 per second 50 200
What the lever buys the hash the x8 mixer, dr368, the v4 shadow all fit with 1.8 ms spare on the half-core dr736 fits (15.5 on the half-core: no, still out), x16 mixer fits nothing of class v4 fits on the half-core

Loosening to 20 ms would admit the x16 mixer (about 7 ms on the Mac, 15 on the half-core: still out) and not dr736 on the half-core, so it buys little against a chip (the f = 1 chip derives no item) and doubles the header-flood and IBD costs on the weakest node. Keep 10 ms; measure the laptop; use the half-core row until then.

5.6 Task 6: the dataset schedule to 2030

The schedule (spec 1.13.3 option (b), decided for the cache on 5 October, recommended for the dataset): 2 GiB at genesis, 4 GiB at year 4 (day 1,460), 8 GiB at year 12, 16 GiB at year 28; the cache 256 MiB, 512 MiB, 1 GiB, 2 GiB on the same days (memhard.rs growth_doublings). The devnet packs run 1 GiB today.

The installed base against it (Steam September 2026, cited; the trend is approximate):

Year (approx calendar) 8 GB share 12 GB share Dataset Who falls out of mine-only Who falls out of mine-and-prove (the prover's measured peaks, prover-tiers-real-cards.md)
2026 (devnet) 27% 13% 1 GiB nobody with 4 GB or more 8 GB: compressed does not fit beside the miner (measured); core-only 2^25 fits with 1 GB spare
2027 (genesis, year 0) about 20% about 10% 2 GiB 4 GB cards hold with 0.5 to 0.8 GB spare (card-lifetime) 8 GB loses core-only beside the miner (the miner's resident set grows 1 GiB: 7.35 + 1.0 GB over 8.19) and becomes prove-alone; 12 GB holds compressed on headless Linux (10.2 + 1.0 of 12.3)
2031 (year 4) about 0% about 0% 4 GiB 4 GB cards (5 percent of Steam today, about 0 by then) 12 GB loses compressed (10.2 + 3.0 GB over 12.3), keeps core-only (7.2 + 3.0 of 12.3); 16 GB keeps compressed (9.2 + 3.0 of 16.4)
2039 (year 12) 0 0 8 GiB 8 GB cards (26.7 percent of Steam today; the trend says under 1 percent by 2030) 12 GB loses core-only; 16 GB loses compressed, keeps core-only (7.4 + 7.0 of 16.4); 24 GB keeps compressed (11.0 + 7.0 of 24.6)
2055 (year 28) 0 0 16 GiB 12 and 16 GB 24 GB loses compressed; 32 GB keeps everything

Reading per tier: the 8 GB tier, a quarter of Steam today, is never a mine-and-prove card on the measured prover footprint whatever the dataset does, and it mines until year 12, by which time its share on the trend is nil; the 12 GB tier (13 percent, falling 6 points a year) mines to year 28 and loses mine-and-prove compressed at year 4 because the prover peaks at 10.2 GB beside a 1.4 GB miner; the 16 GB tier (27 percent and rising) mines and proves compressed to year 4 and core-only to year 12; 24 and 32 GB are unconstrained to year 28. The dataset is never the binding constraint on any tier before year 12; the prover's 5.6 to 10.7 GB footprint is. On the chip side the schedule changes nothing: one HBM3 stack holds 24 GB and the 5090's own board 32 GB (chip-model-v3 5.7), so f = 1 reads every step of the schedule to year 28 without a second stack.

What growth is for, then. Two things, both real and neither a chip: (1) the cache above every GPU's on-die cache (96 MB on the 5090, 128 MB on GB202; the spec's own rule), so the honest hash stays DRAM-latency-bound and no consumer GPU gains an L2 shortcut; (2) the dataset above one reticle of SRAM at the node of the day (1.6 GiB per reticle at N5 headline density, 1.9 at N2, sram-mirror.md s4), so an "f = 1 in SRAM" chip, which would read at SRAM latency and beat the DRAM activate ceiling by 10x, stays a multi-reticle part: at 2 GiB that is 2 dies at N2, at 4 GiB 3 dies, at 8 GiB 5 dies (approx, headline density; the lower-bound density halves these). Wafer-scale parts already hold more (a Cerebras WSE-3 carries 44 GB of on-wafer SRAM, approximate, from memory, unpriced here): the schedule does not price that device out and nothing in the hash can, but at the dependent-read pattern a wafer's cross-die hops cost latency that no one has measured for this hash (open, section 7).

Recommendation with numbers: hold the schedule as decided (2 GiB genesis, doublings at years 4, 12, 28). Do not slow it: slowing buys the 8 GB tier nothing (it mines to year 12 either way) and loses the SRAM-reticle margin (at a flat 2 GiB one N2 reticle holds 1.9 GiB today and about 3.4 GiB by 2036 on the 6 percent a year density trend, sram-mirror s6, so a flat dataset fits one reticle within a decade). Do not grow faster: the only tier a faster schedule costs is the 8 GB tier (year 12 to year 4) and the Apple 8 GB laptop, and it buys nothing against the HBM chip. One change: write the prover footprint, not the dataset, into the public card-lifetime sentence (litepaper line 560: "12 GB or more proves full shards" becomes "12 GB mines and proves on headless Linux until the year-4 dataset step, 16 GB until year 12, 24 GB beyond; 8 GB proves alone"), since that is the number that moves users.

6. Ranked proposals

Rank Proposal Evidence Model Hours Consequence per tier Gate
1 Close O-1.14 with a real 2019 laptop run and adopt the half-core proxy as the standing stand-in; dr736 out, dr368 in as R0's length box proxy: dr736 10.5 ms cold, 15.5 half-core; v4 5.06 / 8.23 section 5.5 2 (Windows cross build on the box, relay to the laptop, bench, log) miners 0; pool verifier cores x1.3 at dr368 instead of x2.4; a 2019 node keeps 1.8 ms of headroom under v4 ms per warp under 10 cold on the laptop for v3, v4, dr368; dr736 recorded as the figure that fails
2 Measure the FPGA lane on AWS F2 (one VU47P, HBM2, USD 1.98 an hour) and replace the 12.2 ceiling row the measured 2.4 G/s equals the JEDEC tFAW ceiling; the 1.9x row rests on a 12 ns tFAW the JEDEC cycles do not support section 5.1 8 to 10 agent hours plus USD 2 to 8 of F2 time none today; the public FPGA claim becomes a measured 0.3x to 0.5x per watt reads per second per watt at 1 GiB; alarm at 27 M/s/W, Counter ASIC 4.0 at 54
3 Public-text correction: the era draw and the reserve are automatic schedule changes against fixed datapaths and against forks, not unpredictability against a chip; publish the chip's USD per MH/s-hour beside the honest cards' section 5.4: every drawn parameter needs no silicon; the reserve is about USD 4 of N5 silicon on a chip sections 5.1, 5.2, 5.4 1 holders and miners read a claim that survives review; nothing on the devnet changes docs/evidence.md row with the two numbers (USD 0.00021 chip, 0.00113 owned 5090, 0.0117 rented) and the draw sentence
4 Reserve order: R1 shfla, R2 perm, R3 popc and clz, R4 to R7 unchanged, R8 mm8 at era 8 by the rule; R0 at dr368 AMD step cost of shfla measured 0.75 to 0.84 (the one number the reserve document said could move R3); dr736 fails the proxy section 5.2 1 (spec text in counter-asic-3-reserve.md s5 and s6) Apple pays shfla's 1.91x per op first, under 1 percent of ALU time at 4 points (argued, measured at the unlock rehearsal); NVIDIA and AMD 0 the family-live 5 percent run per vendor at each unlock rehearsal
5 N grows by the era draw at genesis, inside a verifier-bounded ladder, each step taken by 90 percent miner signal. The shadow size N becomes a genesis ladder indexed per era, {100,000, 130,000, 200,000, 330,000, 650,000, 1,000,000} counted ops (the measured rungs, then doublings), floor 100,000 and ceiling 1,000,000 fixed at genesis (the ceiling is the 10x verifier headroom on the 2.5x rule; 370,000 on the half-core proxy until O-1.14 lands, which then sets it), the era stream consuming one draw for it as it does for epoch_len, and the step up or down set by 90 percent of blue blocks over 7 days at a day boundary (spec 5.7's mechanism, P2's signalling code), never unconditionally lane 7: HBM4 doubles the f = 1 chip's rate per stack, so the bare edge rises 5.7x to 12x (upper bound) and only N answers it; this lane: the chip's edge over the 5090 at k = 1 falls 2.1x (100,000) to 1.7x (130,000) to 1.3x (200,000 and 330,000); an unconditional doubling takes the M5 Max out at the first step section 5.3a; --section ladder 6 to 8 (the ladder field in ShadowClass and the era stream, the signal rule shared with epoch_len, a fast-time run across one step, packs and vectors per rung) At each step, measured: 100,000 to 130,000 costs the M5 Max 3.3 points of rate and 0 W more, the 5090 and 4070 nothing, the 9070 XT nothing, every verifier +0.05 ms; 130,000 to 200,000 costs the M5 Max 6 more points and the 5090 2.7 at its cap, the 4070 +21 W, every verifier +0.12 ms (Mac) to +0.5 (half-core); 200,000 to 330,000 is compute-bound on every NVIDIA card at its cap (5090 -35 percent) and is a step the signal would refuse until cards change; a pool user nothing at any step; a chip's shadow core grows with N at k x 11 pJ per op ONE gate per step, published before the project recommends the signal: at the step's N, every card of the public benchmark set (the four owned plus the eleven rented models) within 5 percent of its rate at the previous step, bit-exact fingerprints on Metal, CUDA and AMD OpenCL, and the verifier under 10 ms per warp cold on the O-1.14 core (the half-core proxy until then)
5a A class v5 candidate mx8 + sh256x35 with a shuffle-heavy shadow weight table (shfl 14, shfla 8 of 75), measured on the four owned cards before any cut: the first rung of proposal 5's ladder, plus the k-floor lever the Mac's 5 percent point is 130,000; the shuffle mix raises the k floor 0.32 to 0.46 (approx) section 5.3 4 to build the weight-table knob and packs, 1 Mac measure session (about 6 min under the lock, miner paused), 3 PC jobs M5 Max -4.8 percent of rate at 37 W; 5090 -0.3 percent at its 431 W cap (a rig +23 percent electricity); 4070 0 at about 118 W; 9070 XT 0; verifier +0.23 ms Mac, +0.85 half-core every owned card within 5 percent; bit-exact on three vendors; half-core verifier under 10 ms; the 5090's marginal pJ on the new mix read on three rungs
6 Hold the dataset schedule (2 GiB, years 4, 12, 28); write the prover footprint into the card-lifetime sentence Steam shares and the measured prover peaks; one HBM3 stack holds every step section 5.6 1 8 GB: mines to year 12, proves alone; 12 GB: mine-and-prove compressed to year 4, core-only to year 12, mines to year 28; 16 GB: compressed to year 4, core-only to year 12; 24 and 32 GB unconstrained to year 28 the litepaper sentence matches the table; docs/evidence.md row "card lifetime" labelled designed
7 Make the Ember tune the shipped default per card model (the honest card's watts are the lever that moves every chip row) the 4070 at 3.65 uJ untuned and 2.57 tuned (-30 percent); the 5090 2.65 bench against 2.34 app section 5.1 2 (defaults table in the app from the fleet priors; already measured) every NVIDIA tier gains 10 to 30 percent per joule; the chip's edge over the mid-tier falls from 8x to 13x toward 5x to 9x at v3 MH per W per card model on the fleet night against the untuned baseline
8 Fund the k question: the item 3 cryptanalysis brief gains a chip-design line (a 14,000-lane SIMD array's energy per op on a random 32-lane program with shuffles, at N5 and at 28 nm) every chip row at class v4 turns on k; nothing in the project measures it section 5.3 0 agent hours; Josh's money (part of the USD 80,000 to 160,000 brief) none until the number lands; it decides whether 2x is reachable a reviewed estimate of k with its range

Paragraphs.

  1. The verifier gate is the one place tonight produced a measurement instead of a rule. The box proxy brackets a 2019 laptop from both sides (a 2022 server core at full boost; the same core with its sibling busy), and dr736 fails both brackets while class v4 passes both with 1.8 ms to spare. The measurement is one Windows build and one bench on the laptop Josh already owns; until it lands, the half-core row replaces the "2.5x" from memory in every status file.

  2. The FPGA lane's upper row was built on an activate rate (8 per 12 ns per channel) that the JEDEC HBM2 cycle table does not support (4 per 28 ns); the measured Shuhai rate sits exactly on the JEDEC ceiling. That reading can be wrong (the ICCAD table's clock interpretation, the half-bank count, the board watts are all approximate), which is why the F2 hour is the proposal and not the conclusion. It is cheap and it turns a public ceiling claim into a measured one.

  3. The honest statement about the draw and the reserve is owed before the public testnet. The project has said the era draw and the family reserve are "automatic anti-ASIC escalators"; against the chip that the model says anyone would build they escalate nothing, because every parameter they move is firmware or a USD 4 block. They are good design against forks and against a hard-datapath FPGA, and that is what the text should say. The chip's USD per MH/s-hour (56x under rental, 5.4x under an owned 5090) belongs beside it, because it is the number a miner will compute on the day a chip appears.

  4. The reserve order changes only where the measurements moved: shfla's AMD cost came in cheap, so the largest structure goes first; dr736 failed the proxy, so R0 is dr368. mm8 keeps no exception because epoch-length.md 12.4 showed it removes no adversary class.

  5. N as a genesis ladder is the one structural change this lane proposes, and it is lane 7's idea with the cards' measured bind points written into it. The memory generation it answers (HBM4, 2028) arrives on a two-to-three-year cadence; the chain must answer without a release, which the era stream and the epoch_len signal rule already provide the shape for. What the measured rungs add: the ladder's steps are the cards' own bind points, the step is taken by the miners who pay for it, and the ceiling is the verifier's measured core, not a 10x from a ratio. An unconditional doubling per era would retire the Apple tier at era 1 and every capped NVIDIA card at era 2, so the schedule alone is not the proposal; the schedule plus the signal is.

5a. v5 as a class (mx8 + sh256x35 with a shuffle-heavy mix) is measurable in an evening and is not a cut. Its honest ceiling is the Mac's 5 percent and a k floor of about 0.5; it buys 2.1x to 1.7x against the 5090 at k = 1. The design item that decides more than v5 is k itself (proposal 8).

  1. The dataset schedule is right as decided and the public sentence about cards is wrong in kind: it talks about the dataset when the prover is what ends a tier's mine-and-prove life.

  2. The one lever that moves every chip row and costs no consensus change is the honest card's watts. The fleet showed the untuned mid-tier at 1.5 to 2.3x the 5090's energy per hash; Ember's measured tune on the 4070 took 30 percent off. Shipping it as the default is a miner-app change with a measured gate.

  3. k is the whole chip question at class v4 and nobody in the project can measure it; the external brief can estimate it.

7. Open questions and what could not be run

Item Why not What would close it
The v5 Mac packbench ladder the shuffle-heavy shadow weight table does not exist as a knob (the block uses the program's weights), and the Mac's miner state was not checked; a run without the knob would have measured v4 again proposal 5, 4 hours of code then the 6-minute measure session
A fixed low clock on the box cpupower needs root; the core boosted to 3.8 GHz under schedutil the laptop measurement (proposal 1); the half-core row is the pessimistic stand-in
The 9070 XT watts at class v4 and its per-joule row the AMD watts job failed on 6 October and the re-run waits on the runner's --cards-off (status 6a) the next cut's PC 1 job
The RTX 4060's watts (logged 0.0 W on the rented box) the sampler read nothing on that host one re-rent
HBM2 tFAW and half-bank count on the U55C and F2 parts behind the JEDEC paywall; the ICCAD table is a simulator's configuration (DRAMSim3), not a datasheet, and its clock interpretation (1,066 MHz) is mine the F2 hour (proposal 2)
A wafer-scale SRAM dataset holder (Cerebras-class, 44 GB on-wafer, approximate) not priced anywhere in the project; its cross-die hop latency on a dependent-read chain is unmeasured a Counter ASIC 4.0 analysis item, not a lane 2 item
The Steam trend to 2030 linear extrapolation of two points per tier the survey itself, yearly
The era draw's cryptanalysis (the stride bijection, the ROT weak-key draw) out of scope here and still open in spec 1.8.4 and era-layout s8 the item 3 brief
block-rate-devnet2.md RUN_A and RUN_B placeholders at 21:30 UK nothing in this lane depends on them

8. Summary for the coordinator

Lane 2 refined the shipped hash and its classes on the 6 October numbers and one new measurement. The chip model's answer does not change in kind: the stored-dataset chip is the chip, it reads 5.7x per joule against the 5090 bench row and 1.7x against the M5 Max at class v3, 2.1x and 0.9x at class v4 and k = 1, and it undercuts rented hash 56x and an owned 5090 5.4x per MH/s-hour; the FPGA lane tightens to 0.30x to 0.47x per watt because the measured random-read rate of an HBM2 FPGA is the JEDEC activate ceiling and not a mapping artefact, and AWS F2 can measure it for USD 2 an hour. The reserve and the era draw buy nothing against a chip (every drawn parameter is firmware; every reserve block is about USD 4 of silicon) and the public text should say what they do buy. The verifier gate got its first measured proxies: class v4 passes a 2022 server core (5.06 ms cold) and the same core with its SMT sibling busy (8.23 ms); dr736 fails both (10.5 and 15.5 ms), so R0 is dr368. The dataset schedule holds; the tier constraint to 2030 is the prover's footprint, not the dataset.

  1. Class v4 verifier on the box proxy 4.90 ms steady, 5.06 cold, 8.23 on the half-core; dr736 9.76 / 10.51 / 15.49: dr736 is out on measurement, class v4 keeps 1.8 ms under the gate on the pessimistic bracket (section 5.5; sim/horizon/algorithm/model.py --section verifier).
  2. The f = 1 GDDR7 chip's edge per joule: 5.7x (5090 bench), 1.7x (M5 Max) at v3; 2.1x and 0.9x at v4 with k = 1; 4.1x and 1.7x at k = 0.3; against the untuned rented mid-tier 8x to 13x at v3; USD 0.00021 per MH/s-hour against 0.0117 rented (section 5.1; --section chip).
  3. The HBM2 FPGA soft overlay: 2.3 to 2.9 G reads/s per 2-stack card by tFAW and tRRD (measured 2.4), 0.30x to 0.47x of the 5090 per watt; the 1.9x ceiling needs a tFAW of 12 ns that the JEDEC HBM2 table (28 ns) does not give (section 5.1; --section fpga).