Josh's 6 October 2026 ask: deep backward and forward research across the hash, finality, economy, network and every shipped surface. This commit carries the first three lanes. - docs/analysis/horizon/algorithm.md: the chip model on the 6 October numbers (f = 1 GDDR7 chip 5.7x per joule against the 5090 at class v3, 2.1x at class v4 with k = 1), the FPGA lane tightened to 0.30x to 0.47x per watt, the reserve R0 to R8, the reconciled shadow-N ladder (section 5.3a) with HBM4 and three verifier brackets, the first measured verifier proxy on igneum-build-1 (class v4 5.06 ms cold, dr736 10.51: out), the dataset schedule to 2030; model sim/horizon/algorithm/model.py. - docs/analysis/horizon/frontier.md: sixteen ideas ranked by payoff over difficulty with the Monero and Kaspa attacks, prior art cited, the honest never column; model sim/horizon/frontier/frontier_model.py. - docs/analysis/horizon/new-pow.md sections 0 to 4: three new proof-of-work schemes defined, reviewed in two personas, scheme A (mining is proving) ruled out on bytes and sampleability, B and C in prototype on two rented 4090s; measured rows follow. - docs/analysis/horizon-2026-10.md: the summary skeleton and the lane table. Every rental cost cites docs/bench-log.md "Rental cost of hash, 6 October 2026". Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
392 lines
29 KiB
Python
392 lines
29 KiB
Python
#!/usr/bin/env python3
|
|
"""Horizon lane 2 (algorithm): the arithmetic behind docs/analysis/horizon/algorithm.md.
|
|
|
|
Every number printed here is arithmetic on a cited or measured input; the inputs are listed with their source
|
|
in INPUTS below and in the document. Nothing here measures a chip. Run:
|
|
|
|
python3 sim/horizon/algorithm/model.py # prints every table as markdown
|
|
python3 sim/horizon/algorithm/model.py --section chip|fpga|v5|era|verifier|schedule
|
|
|
|
No numpy needed (plain Python 3). Sections:
|
|
chip per-joule and per-dollar edge of the three chip classes against each honest card, class v3 and v4
|
|
fpga HBM2 FPGA random-read ceilings from bank count, tRC, tFAW, tRRD; reads in flight; per watt
|
|
v5 shadow N options: chip edge per joule at k = 0.3, 0.5, 1, 1.5; verifier cost; per-tier watts
|
|
era the era draw: what grinding a certified checkpoint costs, and what a draw can move
|
|
verifier the 2019-class gate: the box proxy rows against the M5 Max and the 2.5x rule
|
|
schedule dataset growth against the Steam VRAM shares and the prover footprint per tier
|
|
"""
|
|
import sys
|
|
|
|
# ----------------------------------------------------------------------------------------------------------------
|
|
# INPUTS (measured = a bench-log or analysis entry; cited = a URL or paper; approx = from memory or an estimate)
|
|
# ----------------------------------------------------------------------------------------------------------------
|
|
PJ = 1e-12
|
|
UJ = 1e-6
|
|
|
|
# Honest cards: name, MH/s, watts, source, tier. The v3 class (mx8) unless said otherwise.
|
|
CARDS = [
|
|
# owned, measured 5 and 6 October 2026
|
|
("RTX 5090 (PC 2 bench, 431 W cap, control)", 132.2, 350.0, "latency-shadow-2026-10-06.md s5 (bench-log item 8)", "32 GB"),
|
|
("RTX 5090 (PC 1 app, Ember run 5)", 127.4, 310.0, "counter-asic-3-status.md s7 item 1", "32 GB"),
|
|
("RTX 5090 (rented Vast pod, untuned)", 98.48, 258.2, "prover-tiers-real-cards.md", "32 GB"),
|
|
("RTX 4090 (rented)", 52.25, 183.1, "prover-tiers-real-cards.md", "24 GB"),
|
|
("RTX 3090 (rented)", 37.79, 228.8, "prover-tiers-real-cards.md", "24 GB"),
|
|
("RTX A5000 (rented)", 47.7, 222.7, "prover-tiers-real-cards.md", "24 GB"),
|
|
("RTX 4070 (PC 1, 1,860 MHz lock, 160 W cap)", 30.95, 79.5, "counter-asic-3-status.md item 8 4070 rows", "12 GB"),
|
|
("RTX 4070 (rented, untuned)", 24.99, 91.1, "prover-tiers-real-cards.md", "12 GB"),
|
|
("RTX 5070 (rented)", 41.89, 137.0, "prover-tiers-real-cards.md", "12 GB"),
|
|
("RTX 3060 (rented)", 23.78, 103.7, "prover-tiers-real-cards.md", "12 GB"),
|
|
("RTX 3080 (rented)", 40.82, 204.9, "prover-tiers-real-cards.md", "10 GB"),
|
|
("RTX 4060 Ti 16 GB (rented)", 17.58, 72.3, "prover-tiers-real-cards.md", "16 GB"),
|
|
("RTX 4060 Ti 8 GB (rented)", 19.07, 72.6, "prover-tiers-real-cards.md", "8 GB"),
|
|
("RX 9070 XT (PC 1, ADLX 199 to 203 W)", 18.9, 201.0, "bench-log 9070 XT telemetry; status 3a watts job", "16 GB AMD"),
|
|
("Apple M5 Max (GPU + DRAM channels)", 27.08, 21.0, "latency-shadow-2026-10-06.md s3", "Apple"),
|
|
("Apple M5 Max (package, approx 38 W)", 27.08, 38.0, "ember-tune.md Apple row, approximate", "Apple"),
|
|
]
|
|
# the RTX 4060 (rented) logged 0.0 W: no watt reading, left out of the per-joule rows
|
|
|
|
# Marginal ALU energy per counted op at the shadow (measured): 5090 11 pJ (10.2 to 13.2), M5 Max 6.9 pJ
|
|
MARGINAL_PJ = {"NVIDIA": 11.0, "Apple": 6.9, "AMD": 11.0} # AMD unmeasured: NVIDIA's figure, approx
|
|
N_V4 = 100_000 # counted ops per hash in the class v4 shadow (sh256x27: 102,100 counted, 55,296 instrs)
|
|
|
|
# Measured v4 cards (whole card): 5090 3.27 uJ at 431 W; M5 Max 1.40 uJ at 37 W; 4070 3.51 uJ at 109 W
|
|
MEASURED_V4 = {
|
|
"RTX 5090 (PC 2 bench, 431 W cap, control)": (131.95, 431.0),
|
|
"RTX 4070 (PC 1, 1,860 MHz lock, 160 W cap)": (31.08, 109.0),
|
|
"Apple M5 Max (GPU + DRAM channels)": (26.67, 37.2),
|
|
"Apple M5 Max (package, approx 38 W)": (26.67, 54.0), # 37.2 + the same 17 W of package, approx
|
|
}
|
|
|
|
# Chips (chip-model-v3.md s5.4 and latency-shadow s6; all approximate arithmetic on cited memory figures)
|
|
CHIPS = [
|
|
# name, MH/s, watts at v3, silicon+memory dollars at v3, extra dollars for the v4 ALU array
|
|
("f = 0 on-die recompute (256 MiB SRAM, x8 mixer)", 41.7, 53.7, 700.0, 40.0),
|
|
("f = 1 stored dataset, GDDR7 (16 devices, 512-bit)", 166.4, 77.6, 470.0, 40.0),
|
|
("f = 1 stored dataset, HBM3 one stack", 83.6, 26.8, 550.0, 40.0),
|
|
("f = 1 stored dataset, HBM3 eight stacks", 666.4, 174.4, 2650.0, 40.0),
|
|
]
|
|
K_LIST = [0.3, 0.5, 1.0, 1.5]
|
|
|
|
# Card list prices, approximate (launch MSRP from memory unless cited), USD
|
|
CARD_PRICE = {
|
|
"RTX 5090": 1999.0, "RTX 4090": 1599.0, "RTX 3090": 1499.0, "RTX A5000": 2250.0, "RTX 4070": 549.0,
|
|
"RTX 5070": 549.0, "RTX 3060": 329.0, "RTX 3080": 699.0, "RTX 4060 Ti 16 GB": 499.0, "RTX 4060 Ti 8 GB": 399.0,
|
|
"RX 9070 XT": 599.0, "Apple M5 Max": 3500.0,
|
|
}
|
|
RENTAL_USD_PER_MHS_HOUR = 0.0117 # bench-log "Rental cost of hash, 6 October 2026" (measured, RunPod list)
|
|
ELECTRICITY_USD_PER_KWH = 0.10 # approx
|
|
AMORTISE_HOURS = 2 * 365 * 24 # two years, approx
|
|
|
|
|
|
def uj(mhs, watts):
|
|
return watts / (mhs * 1e6) / UJ
|
|
|
|
|
|
def card_price(name):
|
|
for k, v in CARD_PRICE.items():
|
|
if name.startswith(k):
|
|
return v
|
|
return None
|
|
|
|
|
|
def hourly_cost_per_mhs(price, uj_per_hash):
|
|
cap = price / AMORTISE_HOURS # USD per hour for the whole device
|
|
return cap, uj_per_hash * 3.6e9 * UJ / 3.6e6 * ELECTRICITY_USD_PER_KWH # (capex/h total, energy USD per MH/s-h)
|
|
|
|
|
|
def section_chip():
|
|
print("## Chip edge per joule and per dollar, class v3 and class v4\n")
|
|
print("Chip rows (model, approximate): energy per hash at v3 = watts / rate; at v4 the chip adds N x 11 pJ x k for its")
|
|
print(f"shadow core (N = {N_V4:,} counted ops, k = chip core energy per op over the 5090's measured 11 pJ).\n")
|
|
print("| Chip class | MH/s | W at v3 | uJ per hash, v3 | uJ at v4, k = 0.3 / 0.5 / 1 / 1.5 | $ silicon + memory, v3 / v4 | $ per MH/s, v3 |")
|
|
print("|---|---|---|---|---|---|---|")
|
|
chip_rows = []
|
|
for name, mhs, w, usd, usd_v4 in CHIPS:
|
|
e3 = uj(mhs, w)
|
|
e4 = [e3 + N_V4 * 11.0 * k * PJ / UJ for k in K_LIST]
|
|
chip_rows.append((name, e3, e4, usd, usd_v4))
|
|
print(f"| {name} | {mhs:.1f} | {w:.1f} | {e3:.3f} | " + " / ".join(f"{x:.2f}" for x in e4) +
|
|
f" | {usd:.0f} / {usd + usd_v4:.0f} | {usd / mhs:.1f} |")
|
|
print()
|
|
print("Honest cards at v3 (measured) and at v4 (measured where a row exists, else the card's marginal ALU energy times N;")
|
|
print("a card under a power cap that binds at v4 pays less energy and loses rate instead, as the 5090 did).\n")
|
|
print("| Card (tier) | MH/s | W | uJ, v3 | uJ, v4 | Edge of f=0 chip per joule, v3 / v4 at k=1 | f=1 GDDR7, v3 / v4 at k=0.3 / 0.5 / 1 / 1.5 | f=1 HBM3 one stack, v3 / v4 at k=1 | $ list (approx) | $ per MH/s | USD per MH/s-hour owned (capex 2 y + 0.10/kWh) |")
|
|
print("|---|---|---|---|---|---|---|---|---|---|---|")
|
|
for name, mhs, w, src, tier in CARDS:
|
|
e3 = uj(mhs, w)
|
|
vendor = "Apple" if "Apple" in name else ("AMD" if "RX " in name else "NVIDIA")
|
|
if name in MEASURED_V4:
|
|
m4, w4 = MEASURED_V4[name]
|
|
e4 = uj(m4, w4)
|
|
tag = ""
|
|
else:
|
|
e4 = e3 + N_V4 * MARGINAL_PJ[vendor] * PJ / UJ
|
|
tag = "*"
|
|
f0, f1g, f1h = chip_rows[0], chip_rows[1], chip_rows[2]
|
|
price = card_price(name)
|
|
pm = f"{price:.0f}" if price else "n/a"
|
|
ppm = f"{price / mhs:.1f}" if price else "n/a"
|
|
if price:
|
|
cap, en = hourly_cost_per_mhs(price, e3)
|
|
own = f"{cap / mhs + en:.5f}"
|
|
else:
|
|
own = "n/a"
|
|
print(f"| {name} ({tier}) | {mhs:.2f} | {w:.1f} | {e3:.2f} | {e4:.2f}{tag} | {e3 / f0[1]:.2f}x / {e4 / f0[2][2]:.2f}x | "
|
|
f"{e3 / f1g[1]:.1f}x / " + " / ".join(f"{e4 / x:.1f}x" for x in f1g[2]) +
|
|
f" | {e3 / f1h[1]:.1f}x / {e4 / f1h[2][2]:.1f}x | {pm} | {ppm} | {own} |")
|
|
print("\n`*` = modelled v4 energy (marginal pJ x N), not measured. Rental hash for comparison: USD 0.0117 per MH/s-hour (measured, RunPod).")
|
|
# chip hourly cost
|
|
print("\nChip hourly cost per MH/s (amortised over two years plus electricity at USD 0.10 per kWh, approximate):\n")
|
|
print("| Chip | uJ, v3 | USD per MH/s-hour | Against rental (0.0117) | Against an owned 5090 (bench row) |")
|
|
print("|---|---|---|---|---|")
|
|
p5090 = CARD_PRICE["RTX 5090"]
|
|
e5090 = uj(132.2, 350.0)
|
|
cap5, en5 = hourly_cost_per_mhs(p5090, e5090)
|
|
own5090 = cap5 / 132.2 + en5
|
|
for name, mhs, w, usd, usd_v4 in CHIPS:
|
|
e3 = uj(mhs, w)
|
|
cap, en = hourly_cost_per_mhs(usd, e3)
|
|
c = cap / mhs + en
|
|
print(f"| {name} | {e3:.3f} | {c:.6f} | {RENTAL_USD_PER_MHS_HOUR / c:.0f}x cheaper | {own5090 / c:.1f}x cheaper |")
|
|
print(f"\nOwned 5090 (bench row): USD {own5090:.5f} per MH/s-hour; rental is {RENTAL_USD_PER_MHS_HOUR / own5090:.0f}x that.")
|
|
|
|
|
|
def section_fpga():
|
|
print("## HBM2 FPGA random-read ceilings (U55C, U280, AWS F2 VU47P: 2 stacks, 16 channels, 32 pseudo-channels)\n")
|
|
channels = 16
|
|
pcs = 32
|
|
banks_per_pc = 16 # approx: HBM2 16 banks per channel, each split into half banks per pseudo-channel
|
|
clk_mhz = 1066.0 # 2,133 MT/s data rate, JEDEC HBM2 as tabled by ICCAD 2021 Table I (cycles)
|
|
ns = 1000.0 / clk_mhz
|
|
tRC = (14 + 34) * ns # tRCD + tRAS cycles (ICCAD 2021 Table I): 48 cycles
|
|
tFAW = 30 * ns # 4 activates per channel per tFAW (Table I)
|
|
tRRD = 6 * ns # ACT to ACT, different bank (Table I)
|
|
bank_bound = pcs * banks_per_pc / (tRC * 1e-9)
|
|
faw_bound = channels * 4 / (tFAW * 1e-9)
|
|
rrd_bound = channels / (tRRD * 1e-9)
|
|
oconnor_bound = channels * 8 / (12e-9) # the figure chip-model-v3.md s5.3 carried for HBM2/3 (O'Connor Table 2)
|
|
lat = 137.8 # ns, Shuhai Table IV page miss on the U280
|
|
print(f"tRC {tRC:.1f} ns, tFAW {tFAW:.1f} ns (4 ACT per channel), tRRD {tRRD:.1f} ns, latency {lat} ns (page miss, measured).\n")
|
|
print("| Ceiling | Formula | G reads/s per card | Reads in flight at 137.8 ns | Per W at 115 / 150 / 225 W (M reads/s/W) | Against the 5090 per W (53.7 M at 326 W; 50.0 M at 350 W) |")
|
|
print("|---|---|---|---|---|---|")
|
|
rows = [
|
|
("Measured, Shuhai U280 default mapping", "2.4 G (FCCM 2020, Fig 7)", 2.4e9),
|
|
("tFAW-bound (JEDEC HBM2 cycles, ICCAD 2021 Table I)", "16 ch x 4 / 28.1 ns", faw_bound),
|
|
("tRRD-bound", "16 ch / 5.6 ns", rrd_bound),
|
|
("Bank-bound, no activate window (the 12.2 ceiling row)", "32 pc x 16 banks / 45 ns", bank_bound),
|
|
("O'Connor HBM2 activate figure as carried by chip-model-v3 s5.3", "16 ch x 8 / 12 ns", oconnor_bound),
|
|
]
|
|
for name, f, g in rows:
|
|
flight = g * lat * 1e-9
|
|
pw = [g / w / 1e6 for w in (115, 150, 225)]
|
|
print(f"| {name} | {f} | {g / 1e9:.1f} | {flight:.0f} | " + " / ".join(f"{x:.0f}" for x in pw) +
|
|
f" | {pw[0] / 53.7:.2f}x to {pw[1] / 53.7:.2f}x (U55C); {pw[2] / 53.7:.2f}x (U280) |")
|
|
print("\nReading: the measured 2.4 G/s equals the tFAW-bound ceiling at JEDEC HBM2 timings (2.3 G/s), so the measured row is")
|
|
print("not a mapping artefact: it is the activate window. The bank-bound 11.4 G/s needs tFAW removed, which the DRAM die")
|
|
print("enforces and no controller overrides. The tightened range for a 2-stack HBM2 FPGA: 2.3 to 2.9 G/s, 0.30x to 0.47x of")
|
|
print("the 5090 per watt at the U55C's 115 to 150 W; the 0.7x to 1.9x row survives only if HBM3-class parts (tFAW about half,")
|
|
print("4 bank groups, approximate) reach an FPGA, and none has (Versal HBM and Agilex M are HBM2e).")
|
|
|
|
|
|
def section_v5():
|
|
print("## Class v5 options: the shadow N against the chip's k\n")
|
|
# honest cards at N: measured ladders where they exist, else model
|
|
# 5090 measured (431 W cap): N -> (MH/s, W)
|
|
r5090 = {930: (132.2, 350), 49700: (132.28, 424.6), 102100: (131.95, 431), 150800: (131.75, 431), 199600: (128.67, 431), 330700: (86.39, 431)}
|
|
rmac = {930: (27.08, 21.0), 49700: (26.75, 31.1), 102100: (26.67, 37.2), 150800: (25.10, 36.7), 199600: (24.25, 38.2), 330700: (21.39, 40.1)}
|
|
r4070 = {930: (30.95, 79.5), 49700: (31.08, 93.5), 102100: (31.08, 109.0), 199600: (31.07, 138.2), 330700: (27.14, 159.9)}
|
|
gddr7 = uj(166.4, 77.6)
|
|
hbm3 = uj(83.6, 26.8)
|
|
print("| N counted ops | 5090 uJ (rate delta) | M5 Max uJ (delta) | 4070 uJ (delta) | 9070 XT | f=1 GDDR7 chip uJ at k = 0.3 / 0.5 / 1 / 1.5 | Edge over the 5090 | Edge over the M5 Max | Edge over the 4070 | Verifier add, M5 Max core / box core / box half-core (ms per warp) |")
|
|
print("|---|---|---|---|---|---|---|---|---|---|")
|
|
# verifier slopes: Mac 3.2 us per 1,000 shadow instrs; box one core 7.0; box half core 12.1 (measured 6 Oct, this lane)
|
|
for N in (930, 49700, 102100, 130000, 150800, 199600, 330700):
|
|
def interp(tbl):
|
|
ks = sorted(tbl)
|
|
if N in tbl:
|
|
return tbl[N]
|
|
lo = max(k for k in ks if k < N); hi = min(k for k in ks if k > N)
|
|
t = (N - lo) / (hi - lo)
|
|
return (tbl[lo][0] + t * (tbl[hi][0] - tbl[lo][0]), tbl[lo][1] + t * (tbl[hi][1] - tbl[lo][1]))
|
|
m5, w5 = interp(r5090); mm, wm = interp(rmac); m4, w4 = interp(r4070)
|
|
e5, em, e4 = uj(m5, w5), uj(mm, wm), uj(m4, w4)
|
|
shadow = max(N - 930, 0)
|
|
chip = [gddr7 + shadow * 11.0 * k * PJ / UJ for k in K_LIST]
|
|
instrs = shadow / 1.83
|
|
vadd = (instrs / 1000 * 3.2e-3, instrs / 1000 * 7.0e-3, instrs / 1000 * 12.1e-3)
|
|
print(f"| {N:,} | {e5:.2f} ({(m5 / 132.2 - 1) * 100:+.1f}%) | {em:.2f} ({(mm / 27.08 - 1) * 100:+.1f}%) | {e4:.2f} ({(m4 / 30.95 - 1) * 100:+.1f}%) | holds to 331k (watts owed) | "
|
|
+ " / ".join(f"{c:.2f}" for c in chip) + " | " + " / ".join(f"{e5 / c:.1f}x" for c in chip) + " | "
|
|
+ " / ".join(f"{em / c:.2f}x" for c in chip) + " | " + " / ".join(f"{e4 / c:.1f}x" for c in chip)
|
|
+ f" | {vadd[0]:.2f} / {vadd[1]:.2f} / {vadd[2]:.2f} |")
|
|
print("\nThe 130,000 row is interpolated between the measured 102,100 and 150,800 rungs (the Mac's 5 percent point).")
|
|
print("\nOption (ii), a shuffle-heavy shadow mix: the chip's datapath floor per op (N5, approx) by op class and the GPU's measured step cost.\n")
|
|
print("| Op class | Chip datapath energy per op, N5 floor (pJ, approx) | GPU step cost on the 5090 (ratio to the add-xor-rotate chain, measured) | Structure a chip must add |")
|
|
print("|---|---|---|---|")
|
|
ops = [("add, sub, xor, or (weights 32 of 75)", 0.06, "1.00", "adder, gates"),
|
|
("mul, mulhi, mad (22 of 75)", 0.52, "1.00 (IMAD is the chain's own op)", "32x32 multiplier"),
|
|
("rotl, rotr (13 of 75)", 0.06, "1.32", "barrel rotator"),
|
|
("shfl xor-mask (8 of 75)", 0.30, "1.49", "5-stage butterfly per warp (5,120 mux bits)"),
|
|
("shfla lane+delta (reserve R3)", 1.00, "1.53", "32x32x32-bit crossbar per warp (32,768 mux bits)"),
|
|
("mm8 (reserve R8)", 1.60, "2.43", "8x8x16 u8 MAC tile per warp (1,024 MACs)")]
|
|
for n, e, g, s in ops:
|
|
print(f"| {n} | {e:.2f} | {g} | {s} |")
|
|
def mix_floor(w_alu, w_mul, w_rot, w_shfl, w_shfla):
|
|
tot = w_alu + w_mul + w_rot + w_shfl + w_shfla
|
|
return (w_alu * 0.06 + w_mul * 0.52 + w_rot * 0.06 + w_shfl * 0.30 + w_shfla * 1.00) / tot
|
|
base = mix_floor(32, 22, 13, 8, 0)
|
|
heavy = mix_floor(26, 18, 9, 14, 8)
|
|
print(f"\nShadow mix energy floor per op: today's weights {base:.3f} pJ; a shuffle-heavy v5 mix (shfl 14, shfla 8 of 75, the")
|
|
print(f"others scaled) {heavy:.3f} pJ, {heavy / base:.2f}x. With the 2x pipeline and 8x register-file and wire overhead of")
|
|
print(f"latency-shadow s6 the k floor moves from {base * 16 / 11:.2f} to {heavy * 16 / 11:.2f} (approx): the shuffle-heavy mix")
|
|
print("raises the attacker's claimed floor, it does not reach k = 1.")
|
|
|
|
|
|
def section_era():
|
|
print("## The era draw: what grinding costs and what a draw can move\n")
|
|
print("Forging the certified checkpoint the era VDF reads needs 2/3 of the 30-day blue-block weight: 20 days of 100 percent")
|
|
print("hash (CLAUDE.md headline). Rented at USD 0.0117 per MH/s-hour (measured, RunPod list, 6 October 2026):\n")
|
|
print("| Network hash | 20 days of 100 percent, rented | What the forged checkpoint buys in the draw |")
|
|
print("|---|---|---|")
|
|
for g in (1, 10, 100, 1000):
|
|
usd = g * 1000 * 480 * RENTAL_USD_PER_MHS_HOUR
|
|
print(f"| {g} GH/s | USD {usd:,.0f} | one era's (M, R, pos, weights +-2, fold rotations): 0.8 to 3.2 percent hash-rate spread per card (measured six eras), 0 chip effect |")
|
|
print("\nWithholding the last blue block before C_era(n) to re-roll the input costs one block and needs the 3,600 s VDF")
|
|
print("evaluated inside the 2 s publish window: a 1,800x faster evaluator (spec 04 s4.6 margin table: 300x beats the epoch,")
|
|
print("not the era). Gain of the re-roll even if free: one draw of the same space.\n")
|
|
print("Op-weight corners: each of the ten non-load weights is perturbed by -2..+2 points; the multiply share (mul, mad,")
|
|
print("mulhi: 22 of 75) can move to 16 or 28 of 75. Chip datapath energy per shadow op at the N5 floor (0.06 add, 0.52 mul):")
|
|
for share in (16, 22, 28):
|
|
e = (share * 0.52 + (75 - share) * 0.06) / 75
|
|
print(f" multiply share {share}/75: {e:.3f} pJ per op")
|
|
print("so the weakest corner for a chip is a draw with the fewest multiplies, worth about 20 percent of the shadow's")
|
|
print("datapath energy (approx) and nothing on the memory side; the GPU's cost moves the same way.")
|
|
|
|
|
|
def section_verifier():
|
|
print("## The 2019-class verifier gate: measured proxies (this lane, 6 October 2026, igneum-build-1)\n")
|
|
mac = {"v2": 0.60, "mx8 (class v3)": 2.06, "mx8+sh256x27 (class v4)": 2.33, "dr368": 2.69, "dr736": 4.88}
|
|
box1 = {"v2": 1.128, "mx8 (class v3)": 4.516, "mx8+sh256x27 (class v4)": 4.901, "dr368": 4.716, "dr736": 9.763}
|
|
box1cold = {"v2": 1.305, "mx8 (class v3)": 4.670, "mx8+sh256x27 (class v4)": 5.061, "dr368": 5.324, "dr736": 10.509}
|
|
half = {"mx8 (class v3)": 7.56, "mx8+sh256x27 (class v4)": 8.23, "dr368": 8.16, "dr736": 15.49}
|
|
print("| Class | M5 Max core, quiet (ms per warp) | 2.5x rule (approx) | Box EPYC 9454P one core at 3.8 GHz, nice 19, taskset (steady / cold) | Box over Mac | Box half-core (SMT sibling loaded) | Gate 10 ms |")
|
|
print("|---|---|---|---|---|---|---|")
|
|
for c in mac:
|
|
h = half.get(c)
|
|
verdict = "pass" if box1cold[c] < 10 and (h is None or h < 10) else ("FAIL cold on one core" if box1cold[c] >= 10 else "fails the half-core proxy")
|
|
print(f"| {c} | {mac[c]:.2f} | {mac[c] * 2.5:.1f} | {box1[c]:.2f} / {box1cold[c]:.2f} | {box1[c] / mac[c]:.2f}x | {('%.2f' % h) if h else 'n/a'} | {verdict} |")
|
|
print("\nCache fill on the box core: 361 ms (M5 Max 175 to 181 ms): 2.0x. The box clock read 3,800 MHz during the run")
|
|
print("(scaling_cur_freq), so this is a 2022 server core at full boost with server DDR5 latency, not a 2019 laptop; the")
|
|
print("half-core row (both SMT siblings busy) is the pessimistic bracket. What the gate protects at 10 ms per warp:\n")
|
|
for ms in (10.0, 20.0):
|
|
print(f"| at {ms:.0f} ms per warp | 1 bps: {ms / 1000 * 100:.0f}% of one core | 10 bps: {ms / 1000 * 10 * 100:.0f}% of one core | IBD 108,000 headers: {108000 * ms / 1000 / 60:.0f} min | header flood to saturate one core: {1000 / ms:.0f} invalid headers per second | pool: {1000 / ms:.0f} shares per second per core |")
|
|
|
|
|
|
def section_schedule():
|
|
print("## Dataset growth against the installed base and the prover footprint\n")
|
|
steam = [("512 MB to 4 GB", 2.60 + 1.44 + 3.69 + 1.20 + 5.13), ("6 GB", 5.12), ("8 GB", 26.71), ("10 to 11 GB", 1.91 + 0.83),
|
|
("12 GB", 13.06), ("16 GB", 27.21), ("20 to 24 GB", 1.39 + 1.04 + 5.50), ("32 GB", 1.41), ("64 GB", 0.50), ("other", 1.26)]
|
|
print("Steam Hardware Survey, September 2026 (cited, store.steampowered.com/hwsurvey; 8 GB was 35.03 percent in August 2025")
|
|
print("and 33.66 in September 2025, 12 GB 19.30 in August 2025: videocardz, wccftech, pcguide, cited):\n")
|
|
print("| VRAM | Share of Steam users, Sep 2026 |")
|
|
print("|---|---|")
|
|
for n, s in steam:
|
|
print(f"| {n} | {s:.2f}% |")
|
|
# trend: 8 GB -7.3 points per year (33.66 -> 26.71 Sep 2025 to Sep 2026); 12 GB -6.2 (19.30 Aug 2025 -> 13.06)
|
|
print("\nLinear extrapolation of the 8 GB and 12 GB shares (approx; the 16 GB and 24 GB tiers absorb them):\n")
|
|
print("| Year | 8 GB share | 12 GB share | Dataset (option b) | Cache |")
|
|
print("|---|---|---|---|---|")
|
|
for y, ds, cache in ((2026, "1 GiB (devnet packs)", "256 MiB"), (2027, "2 GiB (genesis)", "256 MiB"), (2028, "2 GiB", "256 MiB"), (2029, "2 GiB", "256 MiB"),
|
|
(2030, "2 GiB", "256 MiB"), (2031, "4 GiB (year 4)", "512 MiB"), (2035, "4 GiB", "512 MiB"), (2039, "8 GiB (year 12)", "1 GiB")):
|
|
s8 = max(0.0, 26.71 - 7.0 * (y - 2026)); s12 = max(0.0, 13.06 - 6.2 * (y - 2026) * 0.5)
|
|
print(f"| {y} | {s8:.0f}% | {s12:.0f}% | {ds} | {cache} |")
|
|
print("\nMemory per tier at each step (MiB): miner resident = dataset + 128 output buffers + scratch-free kernel (cache freed")
|
|
print("after the build, the decided reading); prover beside the miner from prover-tiers-real-cards.md (peak of the")
|
|
print("compressed 2^26 and core-only 2^25 profiles, both with a 1.4 GB miner resident at the 1 GiB dataset):\n")
|
|
print("| Tier (usable = 75% of VRAM, Apple 50%) | Usable MiB | Mine-only: last step that fits | Mine + prove compressed (peak 8.9 to 10.7 GB at 1 GiB): last step | Mine + prove core-only (peak 7.1 to 7.4 GB at 1 GiB): last step |")
|
|
print("|---|---|---|---|---|")
|
|
steps = [("1 GiB (today)", 1024), ("2 GiB (genesis)", 2048), ("4 GiB (year 4)", 4096), ("8 GiB (year 12)", 8192), ("16 GiB (year 28)", 16384)]
|
|
# (tier, VRAM MiB, measured compressed peak beside the miner at the 1 GiB dataset in MiB or None, measured core-only peak)
|
|
tiers = [("8 GB (4060 Ti 8 GB)", 8192, None, 7352), ("12 GB (4070, 5070)", 12288, 10240, 5734 + 1434), ("16 GB (4060 Ti 16 GB)", 16384, 9216, 5939 + 1434),
|
|
("24 GB (4090, 3090, A5000)", 24576, 10956, 6246 + 1434), ("32 GB (5090)", 32768, 10138, 6451 + 1434)]
|
|
for name, vram, comp_peak1, core_peak1 in tiers:
|
|
usable = vram * 0.75
|
|
mine_last = "none"; comp_last = "none"; core_last = "none"
|
|
for sname, ds in steps:
|
|
if ds + 128 + 64 <= usable:
|
|
mine_last = sname
|
|
grow = ds - 1024 # the miner's resident set grows with the dataset; the prover's part does not
|
|
if comp_peak1 is not None and comp_peak1 + grow <= vram * 0.98: # headless Linux: the rows used nearly the whole card
|
|
comp_last = sname
|
|
if core_peak1 + grow <= vram * 0.98:
|
|
core_last = sname
|
|
print(f"| {name} | {usable:.0f} | {mine_last} | {comp_last} | {core_last} |")
|
|
print("\nReading: the 8 GB tier never mines and proves compressed at any dataset size (measured: does not fit at 1 GiB); core-only")
|
|
print("fits at 1 GiB with 1 GB spare and loses that at the 2 GiB genesis step. The 12 GB tier mines and proves compressed on")
|
|
print("headless Linux at 1 and 2 GiB and loses it at the year-4 step (4 GiB). The 16 GB tier holds compressed to year 4 and")
|
|
print("core-only to year 12. Mine-only: 8 GB to year 12, 12 and 16 GB to year 28, as card-lifetime-2026-10-05.md says.")
|
|
|
|
|
|
def section_ladder():
|
|
"""The reconciled N ladder asked for by the coordinator (lane 7's HBM4 inputs, this lane's measured cards)."""
|
|
print("## The reconciled N ladder (lane 2 measured cards, lane 7 HBM4 inputs)\n")
|
|
# chips: bare uJ per hash (memory-bound f = 1): 128 x E_read + static / rate. Lane 7's HBM4 inputs (approximate):
|
|
# one stack, 32 channels, ceiling 2 x 10.7 G = 21.4 G reads/s, 1.0 nJ per read, 5 W static, 10 W controller.
|
|
hbm4_rate = 21.4e9 / 128
|
|
hbm4_bare = (128 * 1.0e-9 * hbm4_rate + 15.0) / hbm4_rate / UJ
|
|
chips = [("GDDR7", uj(166.4, 77.6)), ("HBM3 1 stack", uj(83.6, 26.8)), ("HBM3 8 stacks", uj(666.4, 174.4)), ("HBM4 1 stack (lane 7)", hbm4_bare)]
|
|
r5090 = {930: (132.2, 350), 49700: (132.28, 424.6), 102100: (131.95, 431), 150800: (131.75, 431), 199600: (128.67, 431), 330700: (86.39, 431)}
|
|
rmac = {930: (27.08, 21.0), 49700: (26.75, 31.1), 102100: (26.67, 37.2), 150800: (25.10, 36.7), 199600: (24.25, 38.2), 330700: (21.39, 40.1)}
|
|
r4070 = {930: (30.95, 79.5), 49700: (31.08, 93.5), 102100: (31.08, 109.0), 199600: (31.07, 138.2), 330700: (27.14, 159.9)}
|
|
r9070 = {930: 18.92, 49700: 19.31, 102100: 19.29, 199600: 19.11, 330700: 19.60}
|
|
def interp(tbl, N):
|
|
ks = sorted(tbl)
|
|
if N in tbl:
|
|
return tbl[N]
|
|
lo = max(k for k in ks if k < N); hi = min(k for k in ks if k > N)
|
|
t = (N - lo) / (hi - lo)
|
|
a, b = tbl[lo], tbl[hi]
|
|
return tuple(x + t * (y - x) for x, y in zip(a, b)) if isinstance(a, tuple) else a + t * (b - a)
|
|
print("Chip bare uJ per hash: " + ", ".join(f"{n} {e:.3f}" for n, e in chips) + ". Chip at N: bare + (N - 930) x 11 pJ x k.")
|
|
print("Card energy: measured ladders (5090 under its 431 W cap; M5 Max GPU + DRAM; 4070 at its 160 W cap); 130,000 interpolated.\n")
|
|
print("| N counted ops | 5090 uJ, W (rate delta) | M5 Max uJ, W (delta) | 4070 uJ, W (delta) | 9070 XT rate delta (W owed) | 12 GB and 8 GB rented cards | Chip edge over the 5090 per joule, k = 0.5 / 1 / 1.5: GDDR7 | HBM3 one stack | HBM3 eight stacks | HBM4 one stack | Verifier ms per warp: M5 Max core / 2019-class (2.5x rule) / half-core proxy | Card that binds first (5 percent rule) |")
|
|
print("|---|---|---|---|---|---|---|---|---|---|---|---|")
|
|
binds = {930: "none", 49700: "none", 102100: "none (M5 Max -1.5%)", 130000: "M5 Max at its 5% point", 150800: "M5 Max (-7.3%)", 199600: "M5 Max (-10%), 5090 (-2.7%)", 330700: "5090 (-35%), M5 Max (-21%), 4070 (-12%)"}
|
|
for N in (930, 49700, 102100, 130000, 150800, 199600, 330700):
|
|
m5, w5 = interp(r5090, N); mm, wm = interp(rmac, N); m4, w4 = interp(r4070, N); m9 = interp(r9070, N)
|
|
e5 = uj(m5, w5)
|
|
instrs = max(N - 930, 0) / 1.83
|
|
vmac = 2.06 + instrs / 1000 * 3.2e-3
|
|
vrule = vmac * 2.5
|
|
vhalf = 7.56 + instrs / 1000 * 12.1e-3
|
|
cells = []
|
|
for name, bare in chips:
|
|
cells.append(" / ".join(f"{e5 / (bare + max(N - 930, 0) * 11.0 * k * PJ / UJ):.1f}x" for k in (0.5, 1.0, 1.5)))
|
|
print(f"| {N:,} | {e5:.2f}, {w5:.0f} W ({(m5 / 132.2 - 1) * 100:+.1f}%) | {uj(mm, wm):.2f}, {wm:.0f} W ({(mm / 27.08 - 1) * 100:+.1f}%) | {uj(m4, w4):.2f}, {w4:.0f} W ({(m4 / 30.95 - 1) * 100:+.1f}%) | {(m9 / 18.92 - 1) * 100:+.1f}% | not measured at any N (v3 only) | "
|
|
+ " | ".join(cells) + f" | {vmac:.2f} / {vrule:.1f} / {vhalf:.2f} | {binds[N]} |")
|
|
print("\nDisagreements with lane 7's model 1.4, named: (a) its honest card at N is the linear 326 to 575 W model (2.95 uJ at")
|
|
print("N = 100,000, 3.50 at 200,000, 4.22 at 330,000); the measured 5090 under its 431 W cap reads 3.27, 3.35 and 4.99 uJ with")
|
|
print("-0.2, -2.7 and -34.7 percent of rate (the cap binds from 102,100 ops; power.min_limit is 400 W so no cap under it exists);")
|
|
print("(b) its shadow core is 150 W fixed at the 5090's 136 MH/s, so on a 167 MH/s HBM4 chip it under-counts the core by 1.23x;")
|
|
print("energy per hash is N x 11 pJ x k whatever the chip's rate, which is what this table uses (HBM4 at N = 100,000, k = 1: 1.33 uJ,")
|
|
print(f"not 1.11; edge {uj(131.95, 431) / (hbm4_bare + 101170 * 11e-12 / UJ):.1f}x, not 2.65x); (c) its HBM3 and HBM4 ceilings (10.7 and 21.4 G) rest on 8 activates per 12 ns")
|
|
print("per channel; the JEDEC HBM2 cycle table gives 4 per 28 ns (this file, section 5.1), and HBM3's own tFAW is behind the paywall.")
|
|
hbm4_low = (128 * 1.0e-9 * (4.6e9 / 128) + 15.0) / (4.6e9 / 128) / UJ
|
|
print(f"If HBM3 and HBM4 carry HBM2's activate window the one-stack ceilings are 2.3 and 4.6 G and the HBM4 bare energy {hbm4_low:.2f} uJ,")
|
|
print(f"an edge of {2.65 / hbm4_low:.1f}x bare over the 5090 instead of 11x; the GDDR7 column (the 5090 measures 82 percent of its ceiling) is")
|
|
print("the one with a measured anchor and is the column to quote.")
|
|
print("\nVerifier headroom for N (the \"10x\" claim): on the M5 Max core 10 - 2.33 = 7.67 ms buys 2.4 M shadow instructions, N about")
|
|
print("4.5 M ops (19x); on the 2.5x rule 4.2 ms buys 525,000 instructions, N about 1.06 M (10x); on the measured half-core proxy 1.77 ms")
|
|
print("buys 146,000 instructions, N about 370,000 (3.7x). The cards bind first at every bracket: M5 Max 130,000, 5090 at 431 W")
|
|
print("210,000, 4070 at 160 W about 250,000, 9070 XT over 331,000.")
|
|
|
|
|
|
SECTIONS = {"chip": section_chip, "fpga": section_fpga, "v5": section_v5, "ladder": section_ladder, "era": section_era, "verifier": section_verifier, "schedule": section_schedule}
|
|
|
|
if __name__ == "__main__":
|
|
want = None
|
|
if len(sys.argv) > 2 and sys.argv[1] == "--section":
|
|
want = sys.argv[2]
|
|
for k, f in SECTIONS.items():
|
|
if want is None or want == k:
|
|
f()
|
|
print()
|