igneum/sim/horizon/algorithm/model.py
igneum-labs 1ac557cca8 Horizon: lanes 2 (algorithm), 7 (frontier) and 8 (new-pow designs) land, with the summary skeleton
the project lead's 6 October 2026 ask: deep backward and forward research across the hash, finality,
economy, network and every shipped surface. This commit carries the first three lanes.

- docs/analysis/horizon/algorithm.md: the chip model on the 6 October numbers (f = 1 GDDR7
  chip 5.7x per joule against the 5090 at class v3, 2.1x at class v4 with k = 1), the FPGA
  lane tightened to 0.30x to 0.47x per watt, the reserve R0 to R8, the reconciled shadow-N
  ladder (section 5.3a) with HBM4 and three verifier brackets, the first measured verifier
  proxy on igneum-build-1 (class v4 5.06 ms cold, dr736 10.51: out), the dataset schedule
  to 2030; model sim/horizon/algorithm/model.py.
- docs/analysis/horizon/frontier.md: sixteen ideas ranked by payoff over difficulty with the
  Monero and Kaspa attacks, prior art cited, the honest never column; model
  sim/horizon/frontier/frontier_model.py.
- docs/analysis/horizon/new-pow.md sections 0 to 4: three new proof-of-work schemes defined,
  reviewed in two personas, scheme A (mining is proving) ruled out on bytes and
  sampleability, B and C in prototype on two rented 4090s; measured rows follow.
- docs/analysis/horizon-2026-10.md: the summary skeleton and the lane table.

Every rental cost cites docs/bench-log.md "Rental cost of hash, 6 October 2026".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 19:51:23 +00:00

392 lines
29 KiB
Python

#!/usr/bin/env python3
"""Horizon lane 2 (algorithm): the arithmetic behind docs/analysis/horizon/algorithm.md.
Every number printed here is arithmetic on a cited or measured input; the inputs are listed with their source
in INPUTS below and in the document. Nothing here measures a chip. Run:
python3 sim/horizon/algorithm/model.py # prints every table as markdown
python3 sim/horizon/algorithm/model.py --section chip|fpga|v5|era|verifier|schedule
No numpy needed (plain Python 3). Sections:
chip per-joule and per-dollar edge of the three chip classes against each honest card, class v3 and v4
fpga HBM2 FPGA random-read ceilings from bank count, tRC, tFAW, tRRD; reads in flight; per watt
v5 shadow N options: chip edge per joule at k = 0.3, 0.5, 1, 1.5; verifier cost; per-tier watts
era the era draw: what grinding a certified checkpoint costs, and what a draw can move
verifier the 2019-class gate: the box proxy rows against the M5 Max and the 2.5x rule
schedule dataset growth against the Steam VRAM shares and the prover footprint per tier
"""
import sys
# ----------------------------------------------------------------------------------------------------------------
# INPUTS (measured = a bench-log or analysis entry; cited = a URL or paper; approx = from memory or an estimate)
# ----------------------------------------------------------------------------------------------------------------
PJ = 1e-12
UJ = 1e-6
# Honest cards: name, MH/s, watts, source, tier. The v3 class (mx8) unless said otherwise.
CARDS = [
# owned, measured 5 and 6 October 2026
("RTX 5090 (PC 2 bench, 431 W cap, control)", 132.2, 350.0, "latency-shadow-2026-10-06.md s5 (bench-log item 8)", "32 GB"),
("RTX 5090 (PC 1 app, Ember run 5)", 127.4, 310.0, "counter-asic-3-status.md s7 item 1", "32 GB"),
("RTX 5090 (rented Vast pod, untuned)", 98.48, 258.2, "prover-tiers-real-cards.md", "32 GB"),
("RTX 4090 (rented)", 52.25, 183.1, "prover-tiers-real-cards.md", "24 GB"),
("RTX 3090 (rented)", 37.79, 228.8, "prover-tiers-real-cards.md", "24 GB"),
("RTX A5000 (rented)", 47.7, 222.7, "prover-tiers-real-cards.md", "24 GB"),
("RTX 4070 (PC 1, 1,860 MHz lock, 160 W cap)", 30.95, 79.5, "counter-asic-3-status.md item 8 4070 rows", "12 GB"),
("RTX 4070 (rented, untuned)", 24.99, 91.1, "prover-tiers-real-cards.md", "12 GB"),
("RTX 5070 (rented)", 41.89, 137.0, "prover-tiers-real-cards.md", "12 GB"),
("RTX 3060 (rented)", 23.78, 103.7, "prover-tiers-real-cards.md", "12 GB"),
("RTX 3080 (rented)", 40.82, 204.9, "prover-tiers-real-cards.md", "10 GB"),
("RTX 4060 Ti 16 GB (rented)", 17.58, 72.3, "prover-tiers-real-cards.md", "16 GB"),
("RTX 4060 Ti 8 GB (rented)", 19.07, 72.6, "prover-tiers-real-cards.md", "8 GB"),
("RX 9070 XT (PC 1, ADLX 199 to 203 W)", 18.9, 201.0, "bench-log 9070 XT telemetry; status 3a watts job", "16 GB AMD"),
("Apple M5 Max (GPU + DRAM channels)", 27.08, 21.0, "latency-shadow-2026-10-06.md s3", "Apple"),
("Apple M5 Max (package, approx 38 W)", 27.08, 38.0, "ember-tune.md Apple row, approximate", "Apple"),
]
# the RTX 4060 (rented) logged 0.0 W: no watt reading, left out of the per-joule rows
# Marginal ALU energy per counted op at the shadow (measured): 5090 11 pJ (10.2 to 13.2), M5 Max 6.9 pJ
MARGINAL_PJ = {"NVIDIA": 11.0, "Apple": 6.9, "AMD": 11.0} # AMD unmeasured: NVIDIA's figure, approx
N_V4 = 100_000 # counted ops per hash in the class v4 shadow (sh256x27: 102,100 counted, 55,296 instrs)
# Measured v4 cards (whole card): 5090 3.27 uJ at 431 W; M5 Max 1.40 uJ at 37 W; 4070 3.51 uJ at 109 W
MEASURED_V4 = {
"RTX 5090 (PC 2 bench, 431 W cap, control)": (131.95, 431.0),
"RTX 4070 (PC 1, 1,860 MHz lock, 160 W cap)": (31.08, 109.0),
"Apple M5 Max (GPU + DRAM channels)": (26.67, 37.2),
"Apple M5 Max (package, approx 38 W)": (26.67, 54.0), # 37.2 + the same 17 W of package, approx
}
# Chips (chip-model-v3.md s5.4 and latency-shadow s6; all approximate arithmetic on cited memory figures)
CHIPS = [
# name, MH/s, watts at v3, silicon+memory dollars at v3, extra dollars for the v4 ALU array
("f = 0 on-die recompute (256 MiB SRAM, x8 mixer)", 41.7, 53.7, 700.0, 40.0),
("f = 1 stored dataset, GDDR7 (16 devices, 512-bit)", 166.4, 77.6, 470.0, 40.0),
("f = 1 stored dataset, HBM3 one stack", 83.6, 26.8, 550.0, 40.0),
("f = 1 stored dataset, HBM3 eight stacks", 666.4, 174.4, 2650.0, 40.0),
]
K_LIST = [0.3, 0.5, 1.0, 1.5]
# Card list prices, approximate (launch MSRP from memory unless cited), USD
CARD_PRICE = {
"RTX 5090": 1999.0, "RTX 4090": 1599.0, "RTX 3090": 1499.0, "RTX A5000": 2250.0, "RTX 4070": 549.0,
"RTX 5070": 549.0, "RTX 3060": 329.0, "RTX 3080": 699.0, "RTX 4060 Ti 16 GB": 499.0, "RTX 4060 Ti 8 GB": 399.0,
"RX 9070 XT": 599.0, "Apple M5 Max": 3500.0,
}
RENTAL_USD_PER_MHS_HOUR = 0.0117 # bench-log "Rental cost of hash, 6 October 2026" (measured, RunPod list)
ELECTRICITY_USD_PER_KWH = 0.10 # approx
AMORTISE_HOURS = 2 * 365 * 24 # two years, approx
def uj(mhs, watts):
return watts / (mhs * 1e6) / UJ
def card_price(name):
for k, v in CARD_PRICE.items():
if name.startswith(k):
return v
return None
def hourly_cost_per_mhs(price, uj_per_hash):
cap = price / AMORTISE_HOURS # USD per hour for the whole device
return cap, uj_per_hash * 3.6e9 * UJ / 3.6e6 * ELECTRICITY_USD_PER_KWH # (capex/h total, energy USD per MH/s-h)
def section_chip():
print("## Chip edge per joule and per dollar, class v3 and class v4\n")
print("Chip rows (model, approximate): energy per hash at v3 = watts / rate; at v4 the chip adds N x 11 pJ x k for its")
print(f"shadow core (N = {N_V4:,} counted ops, k = chip core energy per op over the 5090's measured 11 pJ).\n")
print("| Chip class | MH/s | W at v3 | uJ per hash, v3 | uJ at v4, k = 0.3 / 0.5 / 1 / 1.5 | $ silicon + memory, v3 / v4 | $ per MH/s, v3 |")
print("|---|---|---|---|---|---|---|")
chip_rows = []
for name, mhs, w, usd, usd_v4 in CHIPS:
e3 = uj(mhs, w)
e4 = [e3 + N_V4 * 11.0 * k * PJ / UJ for k in K_LIST]
chip_rows.append((name, e3, e4, usd, usd_v4))
print(f"| {name} | {mhs:.1f} | {w:.1f} | {e3:.3f} | " + " / ".join(f"{x:.2f}" for x in e4) +
f" | {usd:.0f} / {usd + usd_v4:.0f} | {usd / mhs:.1f} |")
print()
print("Honest cards at v3 (measured) and at v4 (measured where a row exists, else the card's marginal ALU energy times N;")
print("a card under a power cap that binds at v4 pays less energy and loses rate instead, as the 5090 did).\n")
print("| Card (tier) | MH/s | W | uJ, v3 | uJ, v4 | Edge of f=0 chip per joule, v3 / v4 at k=1 | f=1 GDDR7, v3 / v4 at k=0.3 / 0.5 / 1 / 1.5 | f=1 HBM3 one stack, v3 / v4 at k=1 | $ list (approx) | $ per MH/s | USD per MH/s-hour owned (capex 2 y + 0.10/kWh) |")
print("|---|---|---|---|---|---|---|---|---|---|---|")
for name, mhs, w, src, tier in CARDS:
e3 = uj(mhs, w)
vendor = "Apple" if "Apple" in name else ("AMD" if "RX " in name else "NVIDIA")
if name in MEASURED_V4:
m4, w4 = MEASURED_V4[name]
e4 = uj(m4, w4)
tag = ""
else:
e4 = e3 + N_V4 * MARGINAL_PJ[vendor] * PJ / UJ
tag = "*"
f0, f1g, f1h = chip_rows[0], chip_rows[1], chip_rows[2]
price = card_price(name)
pm = f"{price:.0f}" if price else "n/a"
ppm = f"{price / mhs:.1f}" if price else "n/a"
if price:
cap, en = hourly_cost_per_mhs(price, e3)
own = f"{cap / mhs + en:.5f}"
else:
own = "n/a"
print(f"| {name} ({tier}) | {mhs:.2f} | {w:.1f} | {e3:.2f} | {e4:.2f}{tag} | {e3 / f0[1]:.2f}x / {e4 / f0[2][2]:.2f}x | "
f"{e3 / f1g[1]:.1f}x / " + " / ".join(f"{e4 / x:.1f}x" for x in f1g[2]) +
f" | {e3 / f1h[1]:.1f}x / {e4 / f1h[2][2]:.1f}x | {pm} | {ppm} | {own} |")
print("\n`*` = modelled v4 energy (marginal pJ x N), not measured. Rental hash for comparison: USD 0.0117 per MH/s-hour (measured, RunPod).")
# chip hourly cost
print("\nChip hourly cost per MH/s (amortised over two years plus electricity at USD 0.10 per kWh, approximate):\n")
print("| Chip | uJ, v3 | USD per MH/s-hour | Against rental (0.0117) | Against an owned 5090 (bench row) |")
print("|---|---|---|---|---|")
p5090 = CARD_PRICE["RTX 5090"]
e5090 = uj(132.2, 350.0)
cap5, en5 = hourly_cost_per_mhs(p5090, e5090)
own5090 = cap5 / 132.2 + en5
for name, mhs, w, usd, usd_v4 in CHIPS:
e3 = uj(mhs, w)
cap, en = hourly_cost_per_mhs(usd, e3)
c = cap / mhs + en
print(f"| {name} | {e3:.3f} | {c:.6f} | {RENTAL_USD_PER_MHS_HOUR / c:.0f}x cheaper | {own5090 / c:.1f}x cheaper |")
print(f"\nOwned 5090 (bench row): USD {own5090:.5f} per MH/s-hour; rental is {RENTAL_USD_PER_MHS_HOUR / own5090:.0f}x that.")
def section_fpga():
print("## HBM2 FPGA random-read ceilings (U55C, U280, AWS F2 VU47P: 2 stacks, 16 channels, 32 pseudo-channels)\n")
channels = 16
pcs = 32
banks_per_pc = 16 # approx: HBM2 16 banks per channel, each split into half banks per pseudo-channel
clk_mhz = 1066.0 # 2,133 MT/s data rate, JEDEC HBM2 as tabled by ICCAD 2021 Table I (cycles)
ns = 1000.0 / clk_mhz
tRC = (14 + 34) * ns # tRCD + tRAS cycles (ICCAD 2021 Table I): 48 cycles
tFAW = 30 * ns # 4 activates per channel per tFAW (Table I)
tRRD = 6 * ns # ACT to ACT, different bank (Table I)
bank_bound = pcs * banks_per_pc / (tRC * 1e-9)
faw_bound = channels * 4 / (tFAW * 1e-9)
rrd_bound = channels / (tRRD * 1e-9)
oconnor_bound = channels * 8 / (12e-9) # the figure chip-model-v3.md s5.3 carried for HBM2/3 (O'Connor Table 2)
lat = 137.8 # ns, Shuhai Table IV page miss on the U280
print(f"tRC {tRC:.1f} ns, tFAW {tFAW:.1f} ns (4 ACT per channel), tRRD {tRRD:.1f} ns, latency {lat} ns (page miss, measured).\n")
print("| Ceiling | Formula | G reads/s per card | Reads in flight at 137.8 ns | Per W at 115 / 150 / 225 W (M reads/s/W) | Against the 5090 per W (53.7 M at 326 W; 50.0 M at 350 W) |")
print("|---|---|---|---|---|---|")
rows = [
("Measured, Shuhai U280 default mapping", "2.4 G (FCCM 2020, Fig 7)", 2.4e9),
("tFAW-bound (JEDEC HBM2 cycles, ICCAD 2021 Table I)", "16 ch x 4 / 28.1 ns", faw_bound),
("tRRD-bound", "16 ch / 5.6 ns", rrd_bound),
("Bank-bound, no activate window (the 12.2 ceiling row)", "32 pc x 16 banks / 45 ns", bank_bound),
("O'Connor HBM2 activate figure as carried by chip-model-v3 s5.3", "16 ch x 8 / 12 ns", oconnor_bound),
]
for name, f, g in rows:
flight = g * lat * 1e-9
pw = [g / w / 1e6 for w in (115, 150, 225)]
print(f"| {name} | {f} | {g / 1e9:.1f} | {flight:.0f} | " + " / ".join(f"{x:.0f}" for x in pw) +
f" | {pw[0] / 53.7:.2f}x to {pw[1] / 53.7:.2f}x (U55C); {pw[2] / 53.7:.2f}x (U280) |")
print("\nReading: the measured 2.4 G/s equals the tFAW-bound ceiling at JEDEC HBM2 timings (2.3 G/s), so the measured row is")
print("not a mapping artefact: it is the activate window. The bank-bound 11.4 G/s needs tFAW removed, which the DRAM die")
print("enforces and no controller overrides. The tightened range for a 2-stack HBM2 FPGA: 2.3 to 2.9 G/s, 0.30x to 0.47x of")
print("the 5090 per watt at the U55C's 115 to 150 W; the 0.7x to 1.9x row survives only if HBM3-class parts (tFAW about half,")
print("4 bank groups, approximate) reach an FPGA, and none has (Versal HBM and Agilex M are HBM2e).")
def section_v5():
print("## Class v5 options: the shadow N against the chip's k\n")
# honest cards at N: measured ladders where they exist, else model
# 5090 measured (431 W cap): N -> (MH/s, W)
r5090 = {930: (132.2, 350), 49700: (132.28, 424.6), 102100: (131.95, 431), 150800: (131.75, 431), 199600: (128.67, 431), 330700: (86.39, 431)}
rmac = {930: (27.08, 21.0), 49700: (26.75, 31.1), 102100: (26.67, 37.2), 150800: (25.10, 36.7), 199600: (24.25, 38.2), 330700: (21.39, 40.1)}
r4070 = {930: (30.95, 79.5), 49700: (31.08, 93.5), 102100: (31.08, 109.0), 199600: (31.07, 138.2), 330700: (27.14, 159.9)}
gddr7 = uj(166.4, 77.6)
hbm3 = uj(83.6, 26.8)
print("| N counted ops | 5090 uJ (rate delta) | M5 Max uJ (delta) | 4070 uJ (delta) | 9070 XT | f=1 GDDR7 chip uJ at k = 0.3 / 0.5 / 1 / 1.5 | Edge over the 5090 | Edge over the M5 Max | Edge over the 4070 | Verifier add, M5 Max core / box core / box half-core (ms per warp) |")
print("|---|---|---|---|---|---|---|---|---|---|")
# verifier slopes: Mac 3.2 us per 1,000 shadow instrs; box one core 7.0; box half core 12.1 (measured 6 Oct, this lane)
for N in (930, 49700, 102100, 130000, 150800, 199600, 330700):
def interp(tbl):
ks = sorted(tbl)
if N in tbl:
return tbl[N]
lo = max(k for k in ks if k < N); hi = min(k for k in ks if k > N)
t = (N - lo) / (hi - lo)
return (tbl[lo][0] + t * (tbl[hi][0] - tbl[lo][0]), tbl[lo][1] + t * (tbl[hi][1] - tbl[lo][1]))
m5, w5 = interp(r5090); mm, wm = interp(rmac); m4, w4 = interp(r4070)
e5, em, e4 = uj(m5, w5), uj(mm, wm), uj(m4, w4)
shadow = max(N - 930, 0)
chip = [gddr7 + shadow * 11.0 * k * PJ / UJ for k in K_LIST]
instrs = shadow / 1.83
vadd = (instrs / 1000 * 3.2e-3, instrs / 1000 * 7.0e-3, instrs / 1000 * 12.1e-3)
print(f"| {N:,} | {e5:.2f} ({(m5 / 132.2 - 1) * 100:+.1f}%) | {em:.2f} ({(mm / 27.08 - 1) * 100:+.1f}%) | {e4:.2f} ({(m4 / 30.95 - 1) * 100:+.1f}%) | holds to 331k (watts owed) | "
+ " / ".join(f"{c:.2f}" for c in chip) + " | " + " / ".join(f"{e5 / c:.1f}x" for c in chip) + " | "
+ " / ".join(f"{em / c:.2f}x" for c in chip) + " | " + " / ".join(f"{e4 / c:.1f}x" for c in chip)
+ f" | {vadd[0]:.2f} / {vadd[1]:.2f} / {vadd[2]:.2f} |")
print("\nThe 130,000 row is interpolated between the measured 102,100 and 150,800 rungs (the Mac's 5 percent point).")
print("\nOption (ii), a shuffle-heavy shadow mix: the chip's datapath floor per op (N5, approx) by op class and the GPU's measured step cost.\n")
print("| Op class | Chip datapath energy per op, N5 floor (pJ, approx) | GPU step cost on the 5090 (ratio to the add-xor-rotate chain, measured) | Structure a chip must add |")
print("|---|---|---|---|")
ops = [("add, sub, xor, or (weights 32 of 75)", 0.06, "1.00", "adder, gates"),
("mul, mulhi, mad (22 of 75)", 0.52, "1.00 (IMAD is the chain's own op)", "32x32 multiplier"),
("rotl, rotr (13 of 75)", 0.06, "1.32", "barrel rotator"),
("shfl xor-mask (8 of 75)", 0.30, "1.49", "5-stage butterfly per warp (5,120 mux bits)"),
("shfla lane+delta (reserve R3)", 1.00, "1.53", "32x32x32-bit crossbar per warp (32,768 mux bits)"),
("mm8 (reserve R8)", 1.60, "2.43", "8x8x16 u8 MAC tile per warp (1,024 MACs)")]
for n, e, g, s in ops:
print(f"| {n} | {e:.2f} | {g} | {s} |")
def mix_floor(w_alu, w_mul, w_rot, w_shfl, w_shfla):
tot = w_alu + w_mul + w_rot + w_shfl + w_shfla
return (w_alu * 0.06 + w_mul * 0.52 + w_rot * 0.06 + w_shfl * 0.30 + w_shfla * 1.00) / tot
base = mix_floor(32, 22, 13, 8, 0)
heavy = mix_floor(26, 18, 9, 14, 8)
print(f"\nShadow mix energy floor per op: today's weights {base:.3f} pJ; a shuffle-heavy v5 mix (shfl 14, shfla 8 of 75, the")
print(f"others scaled) {heavy:.3f} pJ, {heavy / base:.2f}x. With the 2x pipeline and 8x register-file and wire overhead of")
print(f"latency-shadow s6 the k floor moves from {base * 16 / 11:.2f} to {heavy * 16 / 11:.2f} (approx): the shuffle-heavy mix")
print("raises the attacker's claimed floor, it does not reach k = 1.")
def section_era():
print("## The era draw: what grinding costs and what a draw can move\n")
print("Forging the certified checkpoint the era VDF reads needs 2/3 of the 30-day blue-block weight: 20 days of 100 percent")
print("hash (CLAUDE.md headline). Rented at USD 0.0117 per MH/s-hour (measured, RunPod list, 6 October 2026):\n")
print("| Network hash | 20 days of 100 percent, rented | What the forged checkpoint buys in the draw |")
print("|---|---|---|")
for g in (1, 10, 100, 1000):
usd = g * 1000 * 480 * RENTAL_USD_PER_MHS_HOUR
print(f"| {g} GH/s | USD {usd:,.0f} | one era's (M, R, pos, weights +-2, fold rotations): 0.8 to 3.2 percent hash-rate spread per card (measured six eras), 0 chip effect |")
print("\nWithholding the last blue block before C_era(n) to re-roll the input costs one block and needs the 3,600 s VDF")
print("evaluated inside the 2 s publish window: a 1,800x faster evaluator (spec 04 s4.6 margin table: 300x beats the epoch,")
print("not the era). Gain of the re-roll even if free: one draw of the same space.\n")
print("Op-weight corners: each of the ten non-load weights is perturbed by -2..+2 points; the multiply share (mul, mad,")
print("mulhi: 22 of 75) can move to 16 or 28 of 75. Chip datapath energy per shadow op at the N5 floor (0.06 add, 0.52 mul):")
for share in (16, 22, 28):
e = (share * 0.52 + (75 - share) * 0.06) / 75
print(f" multiply share {share}/75: {e:.3f} pJ per op")
print("so the weakest corner for a chip is a draw with the fewest multiplies, worth about 20 percent of the shadow's")
print("datapath energy (approx) and nothing on the memory side; the GPU's cost moves the same way.")
def section_verifier():
print("## The 2019-class verifier gate: measured proxies (this lane, 6 October 2026, igneum-build-1)\n")
mac = {"v2": 0.60, "mx8 (class v3)": 2.06, "mx8+sh256x27 (class v4)": 2.33, "dr368": 2.69, "dr736": 4.88}
box1 = {"v2": 1.128, "mx8 (class v3)": 4.516, "mx8+sh256x27 (class v4)": 4.901, "dr368": 4.716, "dr736": 9.763}
box1cold = {"v2": 1.305, "mx8 (class v3)": 4.670, "mx8+sh256x27 (class v4)": 5.061, "dr368": 5.324, "dr736": 10.509}
half = {"mx8 (class v3)": 7.56, "mx8+sh256x27 (class v4)": 8.23, "dr368": 8.16, "dr736": 15.49}
print("| Class | M5 Max core, quiet (ms per warp) | 2.5x rule (approx) | Box EPYC 9454P one core at 3.8 GHz, nice 19, taskset (steady / cold) | Box over Mac | Box half-core (SMT sibling loaded) | Gate 10 ms |")
print("|---|---|---|---|---|---|---|")
for c in mac:
h = half.get(c)
verdict = "pass" if box1cold[c] < 10 and (h is None or h < 10) else ("FAIL cold on one core" if box1cold[c] >= 10 else "fails the half-core proxy")
print(f"| {c} | {mac[c]:.2f} | {mac[c] * 2.5:.1f} | {box1[c]:.2f} / {box1cold[c]:.2f} | {box1[c] / mac[c]:.2f}x | {('%.2f' % h) if h else 'n/a'} | {verdict} |")
print("\nCache fill on the box core: 361 ms (M5 Max 175 to 181 ms): 2.0x. The box clock read 3,800 MHz during the run")
print("(scaling_cur_freq), so this is a 2022 server core at full boost with server DDR5 latency, not a 2019 laptop; the")
print("half-core row (both SMT siblings busy) is the pessimistic bracket. What the gate protects at 10 ms per warp:\n")
for ms in (10.0, 20.0):
print(f"| at {ms:.0f} ms per warp | 1 bps: {ms / 1000 * 100:.0f}% of one core | 10 bps: {ms / 1000 * 10 * 100:.0f}% of one core | IBD 108,000 headers: {108000 * ms / 1000 / 60:.0f} min | header flood to saturate one core: {1000 / ms:.0f} invalid headers per second | pool: {1000 / ms:.0f} shares per second per core |")
def section_schedule():
print("## Dataset growth against the installed base and the prover footprint\n")
steam = [("512 MB to 4 GB", 2.60 + 1.44 + 3.69 + 1.20 + 5.13), ("6 GB", 5.12), ("8 GB", 26.71), ("10 to 11 GB", 1.91 + 0.83),
("12 GB", 13.06), ("16 GB", 27.21), ("20 to 24 GB", 1.39 + 1.04 + 5.50), ("32 GB", 1.41), ("64 GB", 0.50), ("other", 1.26)]
print("Steam Hardware Survey, September 2026 (cited, store.steampowered.com/hwsurvey; 8 GB was 35.03 percent in August 2025")
print("and 33.66 in September 2025, 12 GB 19.30 in August 2025: videocardz, wccftech, pcguide, cited):\n")
print("| VRAM | Share of Steam users, Sep 2026 |")
print("|---|---|")
for n, s in steam:
print(f"| {n} | {s:.2f}% |")
# trend: 8 GB -7.3 points per year (33.66 -> 26.71 Sep 2025 to Sep 2026); 12 GB -6.2 (19.30 Aug 2025 -> 13.06)
print("\nLinear extrapolation of the 8 GB and 12 GB shares (approx; the 16 GB and 24 GB tiers absorb them):\n")
print("| Year | 8 GB share | 12 GB share | Dataset (option b) | Cache |")
print("|---|---|---|---|---|")
for y, ds, cache in ((2026, "1 GiB (devnet packs)", "256 MiB"), (2027, "2 GiB (genesis)", "256 MiB"), (2028, "2 GiB", "256 MiB"), (2029, "2 GiB", "256 MiB"),
(2030, "2 GiB", "256 MiB"), (2031, "4 GiB (year 4)", "512 MiB"), (2035, "4 GiB", "512 MiB"), (2039, "8 GiB (year 12)", "1 GiB")):
s8 = max(0.0, 26.71 - 7.0 * (y - 2026)); s12 = max(0.0, 13.06 - 6.2 * (y - 2026) * 0.5)
print(f"| {y} | {s8:.0f}% | {s12:.0f}% | {ds} | {cache} |")
print("\nMemory per tier at each step (MiB): miner resident = dataset + 128 output buffers + scratch-free kernel (cache freed")
print("after the build, the decided reading); prover beside the miner from prover-tiers-real-cards.md (peak of the")
print("compressed 2^26 and core-only 2^25 profiles, both with a 1.4 GB miner resident at the 1 GiB dataset):\n")
print("| Tier (usable = 75% of VRAM, Apple 50%) | Usable MiB | Mine-only: last step that fits | Mine + prove compressed (peak 8.9 to 10.7 GB at 1 GiB): last step | Mine + prove core-only (peak 7.1 to 7.4 GB at 1 GiB): last step |")
print("|---|---|---|---|---|")
steps = [("1 GiB (today)", 1024), ("2 GiB (genesis)", 2048), ("4 GiB (year 4)", 4096), ("8 GiB (year 12)", 8192), ("16 GiB (year 28)", 16384)]
# (tier, VRAM MiB, measured compressed peak beside the miner at the 1 GiB dataset in MiB or None, measured core-only peak)
tiers = [("8 GB (4060 Ti 8 GB)", 8192, None, 7352), ("12 GB (4070, 5070)", 12288, 10240, 5734 + 1434), ("16 GB (4060 Ti 16 GB)", 16384, 9216, 5939 + 1434),
("24 GB (4090, 3090, A5000)", 24576, 10956, 6246 + 1434), ("32 GB (5090)", 32768, 10138, 6451 + 1434)]
for name, vram, comp_peak1, core_peak1 in tiers:
usable = vram * 0.75
mine_last = "none"; comp_last = "none"; core_last = "none"
for sname, ds in steps:
if ds + 128 + 64 <= usable:
mine_last = sname
grow = ds - 1024 # the miner's resident set grows with the dataset; the prover's part does not
if comp_peak1 is not None and comp_peak1 + grow <= vram * 0.98: # headless Linux: the rows used nearly the whole card
comp_last = sname
if core_peak1 + grow <= vram * 0.98:
core_last = sname
print(f"| {name} | {usable:.0f} | {mine_last} | {comp_last} | {core_last} |")
print("\nReading: the 8 GB tier never mines and proves compressed at any dataset size (measured: does not fit at 1 GiB); core-only")
print("fits at 1 GiB with 1 GB spare and loses that at the 2 GiB genesis step. The 12 GB tier mines and proves compressed on")
print("headless Linux at 1 and 2 GiB and loses it at the year-4 step (4 GiB). The 16 GB tier holds compressed to year 4 and")
print("core-only to year 12. Mine-only: 8 GB to year 12, 12 and 16 GB to year 28, as card-lifetime-2026-10-05.md says.")
def section_ladder():
"""The reconciled N ladder asked for by the coordinator (lane 7's HBM4 inputs, this lane's measured cards)."""
print("## The reconciled N ladder (lane 2 measured cards, lane 7 HBM4 inputs)\n")
# chips: bare uJ per hash (memory-bound f = 1): 128 x E_read + static / rate. Lane 7's HBM4 inputs (approximate):
# one stack, 32 channels, ceiling 2 x 10.7 G = 21.4 G reads/s, 1.0 nJ per read, 5 W static, 10 W controller.
hbm4_rate = 21.4e9 / 128
hbm4_bare = (128 * 1.0e-9 * hbm4_rate + 15.0) / hbm4_rate / UJ
chips = [("GDDR7", uj(166.4, 77.6)), ("HBM3 1 stack", uj(83.6, 26.8)), ("HBM3 8 stacks", uj(666.4, 174.4)), ("HBM4 1 stack (lane 7)", hbm4_bare)]
r5090 = {930: (132.2, 350), 49700: (132.28, 424.6), 102100: (131.95, 431), 150800: (131.75, 431), 199600: (128.67, 431), 330700: (86.39, 431)}
rmac = {930: (27.08, 21.0), 49700: (26.75, 31.1), 102100: (26.67, 37.2), 150800: (25.10, 36.7), 199600: (24.25, 38.2), 330700: (21.39, 40.1)}
r4070 = {930: (30.95, 79.5), 49700: (31.08, 93.5), 102100: (31.08, 109.0), 199600: (31.07, 138.2), 330700: (27.14, 159.9)}
r9070 = {930: 18.92, 49700: 19.31, 102100: 19.29, 199600: 19.11, 330700: 19.60}
def interp(tbl, N):
ks = sorted(tbl)
if N in tbl:
return tbl[N]
lo = max(k for k in ks if k < N); hi = min(k for k in ks if k > N)
t = (N - lo) / (hi - lo)
a, b = tbl[lo], tbl[hi]
return tuple(x + t * (y - x) for x, y in zip(a, b)) if isinstance(a, tuple) else a + t * (b - a)
print("Chip bare uJ per hash: " + ", ".join(f"{n} {e:.3f}" for n, e in chips) + ". Chip at N: bare + (N - 930) x 11 pJ x k.")
print("Card energy: measured ladders (5090 under its 431 W cap; M5 Max GPU + DRAM; 4070 at its 160 W cap); 130,000 interpolated.\n")
print("| N counted ops | 5090 uJ, W (rate delta) | M5 Max uJ, W (delta) | 4070 uJ, W (delta) | 9070 XT rate delta (W owed) | 12 GB and 8 GB rented cards | Chip edge over the 5090 per joule, k = 0.5 / 1 / 1.5: GDDR7 | HBM3 one stack | HBM3 eight stacks | HBM4 one stack | Verifier ms per warp: M5 Max core / 2019-class (2.5x rule) / half-core proxy | Card that binds first (5 percent rule) |")
print("|---|---|---|---|---|---|---|---|---|---|---|---|")
binds = {930: "none", 49700: "none", 102100: "none (M5 Max -1.5%)", 130000: "M5 Max at its 5% point", 150800: "M5 Max (-7.3%)", 199600: "M5 Max (-10%), 5090 (-2.7%)", 330700: "5090 (-35%), M5 Max (-21%), 4070 (-12%)"}
for N in (930, 49700, 102100, 130000, 150800, 199600, 330700):
m5, w5 = interp(r5090, N); mm, wm = interp(rmac, N); m4, w4 = interp(r4070, N); m9 = interp(r9070, N)
e5 = uj(m5, w5)
instrs = max(N - 930, 0) / 1.83
vmac = 2.06 + instrs / 1000 * 3.2e-3
vrule = vmac * 2.5
vhalf = 7.56 + instrs / 1000 * 12.1e-3
cells = []
for name, bare in chips:
cells.append(" / ".join(f"{e5 / (bare + max(N - 930, 0) * 11.0 * k * PJ / UJ):.1f}x" for k in (0.5, 1.0, 1.5)))
print(f"| {N:,} | {e5:.2f}, {w5:.0f} W ({(m5 / 132.2 - 1) * 100:+.1f}%) | {uj(mm, wm):.2f}, {wm:.0f} W ({(mm / 27.08 - 1) * 100:+.1f}%) | {uj(m4, w4):.2f}, {w4:.0f} W ({(m4 / 30.95 - 1) * 100:+.1f}%) | {(m9 / 18.92 - 1) * 100:+.1f}% | not measured at any N (v3 only) | "
+ " | ".join(cells) + f" | {vmac:.2f} / {vrule:.1f} / {vhalf:.2f} | {binds[N]} |")
print("\nDisagreements with lane 7's model 1.4, named: (a) its honest card at N is the linear 326 to 575 W model (2.95 uJ at")
print("N = 100,000, 3.50 at 200,000, 4.22 at 330,000); the measured 5090 under its 431 W cap reads 3.27, 3.35 and 4.99 uJ with")
print("-0.2, -2.7 and -34.7 percent of rate (the cap binds from 102,100 ops; power.min_limit is 400 W so no cap under it exists);")
print("(b) its shadow core is 150 W fixed at the 5090's 136 MH/s, so on a 167 MH/s HBM4 chip it under-counts the core by 1.23x;")
print("energy per hash is N x 11 pJ x k whatever the chip's rate, which is what this table uses (HBM4 at N = 100,000, k = 1: 1.33 uJ,")
print(f"not 1.11; edge {uj(131.95, 431) / (hbm4_bare + 101170 * 11e-12 / UJ):.1f}x, not 2.65x); (c) its HBM3 and HBM4 ceilings (10.7 and 21.4 G) rest on 8 activates per 12 ns")
print("per channel; the JEDEC HBM2 cycle table gives 4 per 28 ns (this file, section 5.1), and HBM3's own tFAW is behind the paywall.")
hbm4_low = (128 * 1.0e-9 * (4.6e9 / 128) + 15.0) / (4.6e9 / 128) / UJ
print(f"If HBM3 and HBM4 carry HBM2's activate window the one-stack ceilings are 2.3 and 4.6 G and the HBM4 bare energy {hbm4_low:.2f} uJ,")
print(f"an edge of {2.65 / hbm4_low:.1f}x bare over the 5090 instead of 11x; the GDDR7 column (the 5090 measures 82 percent of its ceiling) is")
print("the one with a measured anchor and is the column to quote.")
print("\nVerifier headroom for N (the \"10x\" claim): on the M5 Max core 10 - 2.33 = 7.67 ms buys 2.4 M shadow instructions, N about")
print("4.5 M ops (19x); on the 2.5x rule 4.2 ms buys 525,000 instructions, N about 1.06 M (10x); on the measured half-core proxy 1.77 ms")
print("buys 146,000 instructions, N about 370,000 (3.7x). The cards bind first at every bracket: M5 Max 130,000, 5090 at 431 W")
print("210,000, 4070 at 160 W about 250,000, 9070 XT over 331,000.")
SECTIONS = {"chip": section_chip, "fpga": section_fpga, "v5": section_v5, "ladder": section_ladder, "era": section_era, "verifier": section_verifier, "schedule": section_schedule}
if __name__ == "__main__":
want = None
if len(sys.argv) > 2 and sys.argv[1] == "--section":
want = sys.argv[2]
for k, f in SECTIONS.items():
if want is None or want == k:
f()
print()